Skip to content

feat(regen): add Flux TTS Controls - #797

Merged
GregHolmes merged 15 commits into
mainfrom
gh/sdk-gen-2026-09-30
Oct 1, 2026
Merged

GregHolmes merged 15 commits into
mainfrom
gh/sdk-gen-2026-09-30

Conversation

@GregHolmes

@GregHolmes GregHolmes commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Regenerate the Python SDK with fernapi/fern-python-sdk 5.27.1 from the Flux TTS Controls specification branch.
  • Add Flux TTS Controls behavior, SDK-local documentation, deterministic coverage, and a live smoke script.
  • Retain existing 7.x compatibility, security, websocket, transport, and wire-test patches while preserving generated Listen V2 Warning and Configure additions.

v7 Compatibility

  • SharedTopics and SharedIntents expose the corrected direct segments shape while retaining deprecated results facades, legacy type/Params modules, legacy construction, and root exports.
  • Direct models normalize legacy results input into generated segments; absent legacy results remains None.
  • Existing dict-style Listen V2 response access remains available, including the new typed Warning model.

Flux TTS Controls

  • Add docs/FluxTtsControls.md, reference.md coverage for speak.v2.audio.generate, and README links. Controls are English-only: pronunciation works with Flux batch, Flux WebSocket, and Aura-2; pauses are Flux batch-only.
  • Standardize the canonical pause marker as \{pause:500ms\} and validate the 500-3000 ms range, 100 ms increments, and eight-pause maximum.
  • TextBuilder.build() rejects mixed pronunciation and pause controls before a batch request; Flux would reject that pair with CONTROL_COMBINATION_INVALID.
  • Add deterministic Controls coverage plus tests/manual/speak/v2/controls/main.py, which defaults to production and uses DEEPGRAM_BASE_URL for staging/custom environments.
  • Update the live TextBuilder examples to honor DEEPGRAM_BASE_URL; examples 22-24 passed against staging.
  • Restore rationale/JSDoc on retained agent sanitization, query encoding, and language_hint compatibility shims.

Release

  • Keep committed version sources at the published 7.11.0; this feat(regen) PR should cause Release Please to create 7.12.0.
  • TextBuilder: pause controls cap at 3000 ms and eight per request, emit \{pause:<n>ms\}, and validate embedded markers. Pronunciation controls emit escaped braces. from_ssml() raises for out-of-range breaks instead of rounding.
  • Publish only after public Flux Controls documentation matches the shipped English-only batch/WebSocket behavior.

Validation

  • poetry run pytest - 1,065 passed, 4 skipped
  • poetry run mypy src tests/typecheck
  • poetry check --lock
  • Ruff over changed Python files
  • Staging Controls smoke passed with DEEPGRAM_BASE_URL: batch pause/pronunciation, expected 400 speed-plus-pronunciation rejection, streaming pronunciation, and WebSocket pause DATA-0002
  • Updated TextBuilder examples 22-24 passed against staging

feat(textbuilder): emit escaped Flux pause and pronunciation markers, enforce 500-3000 ms and eight-pause limits, and raise ValueError for mixed or out-of-range controls

@github-actions

github-actions Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Code Coverage

Package Line Rate Branch Rate Complexity Health
src.deepgram 97% 94% 0 ✔
src.deepgram.agent 100% 100% 0 ✔
src.deepgram.agent.v1 98% 100% 0 ✔
src.deepgram.agent.v1.settings 100% 100% 0 ✔
src.deepgram.agent.v1.settings.think 100% 100% 0 ✔
src.deepgram.agent.v1.settings.think.models 97% 100% 0 ✔
src.deepgram.auth 100% 100% 0 ✔
src.deepgram.auth.v1 100% 100% 0 ✔
src.deepgram.auth.v1.tokens 97% 100% 0 ✔
src.deepgram.core 88% 81% 0 ➖
src.deepgram.errors 100% 100% 0 ✔
src.deepgram.helpers 98% 92% 0 ✔
src.deepgram.listen 100% 100% 0 ✔
src.deepgram.listen.v1 98% 93% 0 ✔
src.deepgram.listen.v1.media 97% 100% 0 ✔
src.deepgram.listen.v2 98% 93% 0 ✔
src.deepgram.manage 100% 100% 0 ✔
src.deepgram.manage.v1 100% 100% 0 ✔
src.deepgram.manage.v1.models 96% 100% 0 ✔
src.deepgram.manage.v1.projects 97% 100% 0 ✔
src.deepgram.manage.v1.projects.billing 100% 100% 0 ✔
src.deepgram.manage.v1.projects.billing.balances 96% 100% 0 ✔
src.deepgram.manage.v1.projects.billing.breakdown 97% 100% 0 ✔
src.deepgram.manage.v1.projects.billing.fields 97% 100% 0 ✔
src.deepgram.manage.v1.projects.billing.purchases 97% 100% 0 ✔
src.deepgram.manage.v1.projects.keys 96% 100% 0 ✔
src.deepgram.manage.v1.projects.members 97% 100% 0 ✔
src.deepgram.manage.v1.projects.members.invites 96% 100% 0 ✔
src.deepgram.manage.v1.projects.members.scopes 96% 100% 0 ✔
src.deepgram.manage.v1.projects.models 96% 100% 0 ✔
src.deepgram.manage.v1.projects.usage 98% 100% 0 ✔
src.deepgram.manage.v1.projects.usage.breakdown 97% 100% 0 ✔
src.deepgram.manage.v1.projects.usage.fields 97% 100% 0 ✔
src.deepgram.read 100% 100% 0 ✔
src.deepgram.read.v1 100% 100% 0 ✔
src.deepgram.read.v1.text 98% 100% 0 ✔
src.deepgram.self_hosted 100% 100% 0 ✔
src.deepgram.self_hosted.v1 100% 100% 0 ✔
src.deepgram.self_hosted.v1.distribution_credentials 96% 100% 0 ✔
src.deepgram.speak 100% 100% 0 ✔
src.deepgram.speak.v1 98% 97% 0 ✔
src.deepgram.speak.v1.audio 91% 80% 0 ✔
src.deepgram.speak.v2 98% 93% 0 ✔
src.deepgram.speak.v2.audio 100% 100% 0 ✔
src.deepgram.voice_agent 100% 100% 0 ✔
src.deepgram.voice_agent.configurations 95% 100% 0 ✔
src.deepgram.voice_agent.variables 95% 100% 0 ✔
Summary 95% (6538 / 6855) 91% (1438 / 1574) 0 ✔

Scope: hand-maintained SDK logic. Fern-generated data models (types/, requests/), package __init__.py files, version.py, and the unused core/http_sse/ scaffolding are excluded — see .coveragerc. Unscoped whole-package coverage is ~70%.

@GregHolmes GregHolmes changed the title chore: SDK regeneration 2026-09-30 feat(regen): add Flux TTS Controls Oct 1, 2026

@dg-coreylweathers dg-coreylweathers left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What this PR does
Regenerates the SDK from the Flux TTS Controls spec: adds the pause and pronunciation control surface for /v2/speak (batch headers, WebSocket controls_applied, Warning, ConfigureFailure codes), retypes Topics/Intents to the direct segments shape the API returns while keeping the old results.* paths as facades, adds a Listen V2 Warning type, moves TextBuilder to the Flux batch syntax and limits, and re-applies every 7.x manual patch. Version sources stay at 7.11.0 so release-please cuts 7.12.0.

What I checked

  • Did every manual patch survive? Yes: sanitizer, optional control-send defaults, broad excepts, dict-tolerant send_configure, bool query coercion, header redaction, language_hint shims, update-listen coercion, dict-compat on all Listen V2 responses including the new Warning, StrictInt expressivity, legacy re-exports, _TARGET_MODULES.
  • Is the spec it was generated from what main has? Yes: git diff e252995 origin/main -- api/specs in deepgram-docs is empty.
  • Do the runtime claims hold? Yes, on prod and staging: batch pause OK, pronunciation+speed 1.1 → 400 CONTROL_COMBINATION_INVALID, WebSocket pronunciation → pronunciations_applied=1, WebSocket pause → DATA-0002, pause+speed 1.2 → PAUSE_SPEED_CAP_EXCEEDED, 9 pauses → BREAKS_LIMIT_EXCEEDED, 3100/400 ms → BREAK_OUT_OF_RANGE.
  • Gates in Docker: mypy clean, ruff clean on changed files, pytest 1064 passed, Pydantic 1.10 pass on the compat tests.

Blocking

  • .agents/skills/deepgram-python-text-to-speech/SKILL.md:85-93 chains .pronunciation(...) and .pause(500) and tells the reader to send it to speak.v2.audio.generate. Sent to prod as built, that string returns 400 CONTROL_COMBINATION_INVALID (request_id 01a0f74d-717b-7f91-83c5-e1edb67a498c); docs/FluxTtsControls.md:9 states the rule it breaks. Fix: drop .pause(500) from the example and add one sentence: "Do not combine .pronunciation() and .pause() in one request; Flux rejects the pair with CONTROL_COMBINATION_INVALID."

Should-fix

  • TextBuilder validates counts and ranges but not this combination; add a build() check that raises when both a pause and a pronunciation are present, with a test.
  • Positioning: docs/FluxTtsControls.md:3, helpers/README.md:7, and SKILL.md:80 frame TextBuilder as "English Flux batch", but pronunciation also works over the Flux WebSocket and on Aura-2 /v1/speak (same syntax per deepgram-docs; a live Aura-2 call with TextBuilder output returned dg-pronunciations-applied: 1). README.md:32 reads as if streaming supports pauses. Say pronunciation (batch, WebSocket, Aura-2) and pause (Flux batch only).
  • Release notes: no commit title mentions TextBuilder, so the 7.12.0 changelog will not say that .pause() now caps at 3000 ms and 8 per request, emits \{pause:<n>ms\}, .pronunciation() emits escaped braces, from_ssml raises on out-of-range breaks instead of rounding, and .text() validates embedded markers. Add a TextBuilder bullet.
  • Re-apply lost comments: _sanitize_numeric_types docstring and spacing in agent/v1/socket_client.py, the urlencode comment in query_encoder.py:9, the language_hint deprecation docstring in deepgram_listen_provider_v2.py:31.

Nits: examples/23:6,31 says "pronunciations and pauses" but uses no pause; examples/24:3 docstring not aligned with the README retitle; reference.md still has no speak.v2.audio.generate section (pre-existing).

Merge this before #788 and #789; both already conflict with main and touch the same ledger files.

@GregHolmes

Copy link
Copy Markdown
Contributor Author

Resolved all current review feedback in 2061a9e:

  • TextBuilder now rejects mixed pause and pronunciation controls before the batch request, with regression coverage.
  • Corrected the skill/example/docs positioning: pronunciation spans Flux batch, Flux WebSocket, and Aura-2; pauses are Flux batch-only.
  • Restored frozen compatibility-shim rationale/JSDoc, added maintained speak.v2.audio.generate reference coverage, and aligned examples.

Validation: poetry run pytest (1,065 passed, 4 skipped), mypy, changed-file Ruff, lock validation, Controls staging smoke, and examples 22-24 against staging.

@GregHolmes

Copy link
Copy Markdown
Contributor Author

Resolved the independent-review blocker in HEAD:

  • ssml_to_deepgram() now rejects SSML that combines <phoneme> and <break> controls with the same ValueError as TextBuilder.build().
  • Updated the obsolete mixed-control conversion tests and public helper example.

Validation remains green: poetry run pytest (1,065 passed, 4 skipped), mypy, changed-file Ruff, and lock validation.

@GregHolmes

Copy link
Copy Markdown
Contributor Author

Resolved the final independent-review findings in HEAD:

  • add_pronunciation() now rejects an existing pause marker before emitting a server-rejected mixed-control batch payload; regression coverage added.
  • The central Controls guide now states the English-only launch boundary.

Validation: poetry run pytest (1,066 passed, 4 skipped), mypy, changed-file Ruff, and lock validation.

TextBuilder emits escaped \{pause:<n>ms\} markers and enforces the 500-3000 ms range with no more than eight pauses per batch request.\n\nIt raises ValueError for invalid SSML breaks and mixed pronunciation/pause controls. Flux controls are English-only at launch; pauses are Flux batch-only.
@GregHolmes

Copy link
Copy Markdown
Contributor Author

Added a Release Please-visible feat(textbuilder) migration note in HEAD. Its commit body and the Controls guide cover escaped markers, 500-3000 ms/eight-pause limits, local ValueError validation for invalid SSML and mixed controls, English-only launch scope, and Flux-batch-only pauses.

@GregHolmes

Copy link
Copy Markdown
Contributor Author

Resolved the final review nit in HEAD: the TTS skill index now says only Flux pause controls are batch-only, consistent with Flux WebSocket pronunciation support.

@dg-coreylweathers dg-coreylweathers left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed at 1f17cf8. Approving with one nit.

What changed since the request-changes

  • The skill-file example now builds pronunciation only and states the no-mixing rule. Built with this branch's TextBuilder and sent to prod /v2/speak, it returns 200 with dg-pronunciations-applied: 1.
  • build(), add_pronunciation(), and ssml_to_deepgram() raise ValueError on a pronunciation plus pause mix, each with a test.
  • Docs, helpers README, skill file, and README now say pronunciation works on Flux batch, Flux WebSocket, and Aura-2; pauses are Flux batch only; English-only at launch. reference.md gained a Speak V2 Audio section.
  • The sanitizer docstring, urlencode comment, and language_hint deprecation docstring are restored.

Gates on this head (Docker): ruff clean on changed files, mypy clean, pytest 1066 passed / 4 skipped with WireMock, Pydantic 1.10 pass on the TextBuilder, controls, and compat tests.

Nit before merge: this repo squash-merges with the PR body as the commit message, and release-please only picks up conventional-commit header lines from it, so the feat(textbuilder) commit body and the "TextBuilder" bullet under "Release" will not appear in the 7.12.0 CHANGELOG. Add this line to the PR description on its own line:

feat(textbuilder): emit escaped Flux pause and pronunciation markers, enforce 500-3000 ms and eight-pause limits, and raise ValueError for mixed or out-of-range controls

Merge this before #788 and #789, which both need a rebase.

@GregHolmes
GregHolmes merged commit d2b5514 into main Oct 1, 2026
11 checks passed
@GregHolmes
GregHolmes deleted the gh/sdk-gen-2026-09-30 branch October 1, 2026 15:43
GregHolmes added a commit that referenced this pull request Oct 2, 2026
🤖 I have created a release *beep* *boop*

---

##
[7.12.0](v7.11.0...v7.12.0)
(2026-10-01)

### Features

* **Flux TTS controls:** Add inline pause markers for Flux batch
synthesis and IPA pronunciation overrides for Flux batch, Flux
WebSocket, and Aura-2. Pronunciation is Early Access; pauses are Flux
batch-only. See [Speed, Pause,
Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls).
([#797](#797))
([d2b5514](d2b5514))
* **TextBuilder migration:** Emit the escaped marker syntax Flux
accepts: pauses as `\{pause:500ms\}` and pronunciations as `\{"word":
"...", "pronounce": "<IPA>"\}`. The 7.11.0 forms are rejected or ignored
by `/v2/speak`, so upgrade when using TextBuilder with Flux TTS.
* **TextBuilder validation:** Match Flux batch limits: pauses must be
500-3000 ms in 100 ms increments, with at most eight per request.
`pause()`, `from_ssml()`, and `build()` now raise `ValueError` for
malformed, out-of-range, off-grid, or mixed pause-and-pronunciation
controls before sending a request the API would reject.
* **Listen v2 (Flux):** Add typed `Warning` responses while retaining
dict-style response access during the v7 transition.

### Compatibility

* **Topics and Intents:** Type the direct API `segments` response shape
while retaining deprecated `results` facades, legacy import paths, and
legacy construction throughout v7.

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: Corey Weathers <corey.weathers@deepgram.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants