feat(regen): add Flux TTS Controls - #797
Conversation
|
dg-coreylweathers
left a comment
There was a problem hiding this comment.
What this PR does
Regenerates the SDK from the Flux TTS Controls spec: adds the pause and pronunciation control surface for /v2/speak (batch headers, WebSocket controls_applied, Warning, ConfigureFailure codes), retypes Topics/Intents to the direct segments shape the API returns while keeping the old results.* paths as facades, adds a Listen V2 Warning type, moves TextBuilder to the Flux batch syntax and limits, and re-applies every 7.x manual patch. Version sources stay at 7.11.0 so release-please cuts 7.12.0.
What I checked
- Did every manual patch survive? Yes: sanitizer, optional control-send defaults, broad excepts, dict-tolerant
send_configure, bool query coercion, header redaction,language_hintshims, update-listen coercion, dict-compat on all Listen V2 responses including the new Warning, StrictInt expressivity, legacy re-exports,_TARGET_MODULES. - Is the spec it was generated from what main has? Yes:
git diff e252995 origin/main -- api/specsin deepgram-docs is empty. - Do the runtime claims hold? Yes, on prod and staging: batch pause OK, pronunciation+speed 1.1 → 400 CONTROL_COMBINATION_INVALID, WebSocket pronunciation →
pronunciations_applied=1, WebSocket pause → DATA-0002, pause+speed 1.2 → PAUSE_SPEED_CAP_EXCEEDED, 9 pauses → BREAKS_LIMIT_EXCEEDED, 3100/400 ms → BREAK_OUT_OF_RANGE. - Gates in Docker: mypy clean, ruff clean on changed files, pytest 1064 passed, Pydantic 1.10 pass on the compat tests.
Blocking
.agents/skills/deepgram-python-text-to-speech/SKILL.md:85-93chains.pronunciation(...)and.pause(500)and tells the reader to send it tospeak.v2.audio.generate. Sent to prod as built, that string returns 400CONTROL_COMBINATION_INVALID(request_id 01a0f74d-717b-7f91-83c5-e1edb67a498c);docs/FluxTtsControls.md:9states the rule it breaks. Fix: drop.pause(500)from the example and add one sentence: "Do not combine.pronunciation()and.pause()in one request; Flux rejects the pair withCONTROL_COMBINATION_INVALID."
Should-fix
- TextBuilder validates counts and ranges but not this combination; add a
build()check that raises when both a pause and a pronunciation are present, with a test. - Positioning:
docs/FluxTtsControls.md:3,helpers/README.md:7, andSKILL.md:80frame TextBuilder as "English Flux batch", but pronunciation also works over the Flux WebSocket and on Aura-2/v1/speak(same syntax per deepgram-docs; a live Aura-2 call with TextBuilder output returneddg-pronunciations-applied: 1).README.md:32reads as if streaming supports pauses. Say pronunciation (batch, WebSocket, Aura-2) and pause (Flux batch only). - Release notes: no commit title mentions TextBuilder, so the 7.12.0 changelog will not say that
.pause()now caps at 3000 ms and 8 per request, emits\{pause:<n>ms\},.pronunciation()emits escaped braces,from_ssmlraises on out-of-range breaks instead of rounding, and.text()validates embedded markers. Add a TextBuilder bullet. - Re-apply lost comments:
_sanitize_numeric_typesdocstring and spacing inagent/v1/socket_client.py, the urlencode comment inquery_encoder.py:9, thelanguage_hintdeprecation docstring indeepgram_listen_provider_v2.py:31.
Nits: examples/23:6,31 says "pronunciations and pauses" but uses no pause; examples/24:3 docstring not aligned with the README retitle; reference.md still has no speak.v2.audio.generate section (pre-existing).
Merge this before #788 and #789; both already conflict with main and touch the same ledger files.
|
Resolved all current review feedback in
Validation: |
|
Resolved the independent-review blocker in
Validation remains green: |
|
Resolved the final independent-review findings in
Validation: |
TextBuilder emits escaped \{pause:<n>ms\} markers and enforces the 500-3000 ms range with no more than eight pauses per batch request.\n\nIt raises ValueError for invalid SSML breaks and mixed pronunciation/pause controls. Flux controls are English-only at launch; pauses are Flux batch-only.
|
Added a Release Please-visible |
|
Resolved the final review nit in |
dg-coreylweathers
left a comment
There was a problem hiding this comment.
Re-reviewed at 1f17cf8. Approving with one nit.
What changed since the request-changes
- The skill-file example now builds pronunciation only and states the no-mixing rule. Built with this branch's TextBuilder and sent to prod
/v2/speak, it returns 200 withdg-pronunciations-applied: 1. build(),add_pronunciation(), andssml_to_deepgram()raiseValueErroron a pronunciation plus pause mix, each with a test.- Docs, helpers README, skill file, and README now say pronunciation works on Flux batch, Flux WebSocket, and Aura-2; pauses are Flux batch only; English-only at launch.
reference.mdgained a Speak V2 Audio section. - The sanitizer docstring, urlencode comment, and
language_hintdeprecation docstring are restored.
Gates on this head (Docker): ruff clean on changed files, mypy clean, pytest 1066 passed / 4 skipped with WireMock, Pydantic 1.10 pass on the TextBuilder, controls, and compat tests.
Nit before merge: this repo squash-merges with the PR body as the commit message, and release-please only picks up conventional-commit header lines from it, so the feat(textbuilder) commit body and the "TextBuilder" bullet under "Release" will not appear in the 7.12.0 CHANGELOG. Add this line to the PR description on its own line:
feat(textbuilder): emit escaped Flux pause and pronunciation markers, enforce 500-3000 ms and eight-pause limits, and raise ValueError for mixed or out-of-range controls
🤖 I have created a release *beep* *boop* --- ## [7.12.0](v7.11.0...v7.12.0) (2026-10-01) ### Features * **Flux TTS controls:** Add inline pause markers for Flux batch synthesis and IPA pronunciation overrides for Flux batch, Flux WebSocket, and Aura-2. Pronunciation is Early Access; pauses are Flux batch-only. See [Speed, Pause, Pronunciation](https://developers.deepgram.com/docs/tts-voice-controls). ([#797](#797)) ([d2b5514](d2b5514)) * **TextBuilder migration:** Emit the escaped marker syntax Flux accepts: pauses as `\{pause:500ms\}` and pronunciations as `\{"word": "...", "pronounce": "<IPA>"\}`. The 7.11.0 forms are rejected or ignored by `/v2/speak`, so upgrade when using TextBuilder with Flux TTS. * **TextBuilder validation:** Match Flux batch limits: pauses must be 500-3000 ms in 100 ms increments, with at most eight per request. `pause()`, `from_ssml()`, and `build()` now raise `ValueError` for malformed, out-of-range, off-grid, or mixed pause-and-pronunciation controls before sending a request the API would reject. * **Listen v2 (Flux):** Add typed `Warning` responses while retaining dict-style response access during the v7 transition. ### Compatibility * **Topics and Intents:** Type the direct API `segments` response shape while retaining deprecated `results` facades, legacy import paths, and legacy construction throughout v7. --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: Corey Weathers <corey.weathers@deepgram.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Summary
fernapi/fern-python-sdk5.27.1 from the Flux TTS Controls specification branch.v7 Compatibility
SharedTopicsandSharedIntentsexpose the corrected directsegmentsshape while retaining deprecatedresultsfacades, legacy type/Params modules, legacy construction, and root exports.resultsinput into generatedsegments; absent legacyresultsremainsNone.Flux TTS Controls
docs/FluxTtsControls.md,reference.mdcoverage forspeak.v2.audio.generate, and README links. Controls are English-only: pronunciation works with Flux batch, Flux WebSocket, and Aura-2; pauses are Flux batch-only.\{pause:500ms\}and validate the 500-3000 ms range, 100 ms increments, and eight-pause maximum.TextBuilder.build()rejects mixed pronunciation and pause controls before a batch request; Flux would reject that pair withCONTROL_COMBINATION_INVALID.tests/manual/speak/v2/controls/main.py, which defaults to production and usesDEEPGRAM_BASE_URLfor staging/custom environments.DEEPGRAM_BASE_URL; examples 22-24 passed against staging.language_hintcompatibility shims.Release
7.11.0; thisfeat(regen)PR should cause Release Please to create7.12.0.\{pause:<n>ms\}, and validate embedded markers. Pronunciation controls emit escaped braces.from_ssml()raises for out-of-range breaks instead of rounding.Validation
poetry run pytest- 1,065 passed, 4 skippedpoetry run mypy src tests/typecheckpoetry check --lockDEEPGRAM_BASE_URL: batch pause/pronunciation, expected400speed-plus-pronunciation rejection, streaming pronunciation, and WebSocket pauseDATA-0002feat(textbuilder): emit escaped Flux pause and pronunciation markers, enforce 500-3000 ms and eight-pause limits, and raise ValueError for mixed or out-of-range controls