Repository navigation
feat(irodori_tts): cap reference audio length like Python (max_ref_sec) - #854
Merged
Merged
Conversation
…conds) Python Irodori-TTS cuts a single reference clip to the checkpoint's ref_max_seconds before encoding it (120 s for v4; 30 s for checkpoints that do not state it, such as v3), and trims the latent to ceil(seconds * 48000 / 1920) frames. audio.cpp encoded the whole clip, so a long reference gave a different speaker condition than Python, and its encode memory kept growing with the clip length. The session now applies the same cap by default. The checkpoint value is read from model_config.json. A new request option, max_ref_seconds, sets another length like infer.py --max-ref-seconds; 0 keeps the whole reference. The reference cache key is taken from the trimmed audio plus the latent cap. Speaker embeddings and no-reference requests are unaffected. Published GGUFs embed a contract without the option, so it is dropped from the validation copy like speaker_embedding_path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Owner
|
@KKTTSPJ Could you normalize the option name to |
Follows the *_sec naming of the other duration options (duration_sec, min_duration_sec, max_duration_sec). The docs still name Python's max_ref_seconds for reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Renamed to |
Owner
|
@KKTTSPJ Thanks! PR merged. |
0xShug0
pushed a commit
that referenced
this pull request
Oct 10, 2026
* webui: let a studio panel run the page's request Model panels are registered per family, and the page always builds and sends the request itself. A conversation view has to send its own request and show the reply as it streams in, while it keeps the page's Run button, Ctrl+Enter, Cancel, elapsed time and status line. This adds that hook without changing what any existing panel does: - A panel entry may list the tasks it covers, such as tasks: ['s2s']. An entry without the list covers every task of its family, as before. - With requestMode 'panel', run() hands the request to the runner the panel registers through the page's registerPanelRunner, which the next commit passes to panels as setPanelRunner; registering returns the function that removes it. It passes the source audio, the resolved seed, Max tokens, Language, the request options and the abort signal that Cancel uses. The hand-off comes before anything is awaited, so the panel can start audio inside the Run click or key press. The reply the panel returns is shown in the Result column; these runs are not added to Run history. Missing source audio, a recording still running, or Run from a tab other than Studio is reported as a warning, like the page's other checks before a run. - With UI management the panel can ask for the entry in a given mode. This uses the page's ensureLoadedMode unchanged, after reloading an entry that is resident with another package or with imported settings that differ, as ensureLoaded does for the other runs. - api.ts gets taskStreamEvents, a reader for /v1/tasks/stream with "stream_format": "sse". It yields each event and then the result, throws the server's error message, and throws TaskStreamClosedError when the stream ends before task.stream.done. The done message carries the whole reply again and comes over many reads, so the reader keeps its pieces and joins them once, when the message ends. No panel uses requestMode 'panel' yet. * webui: LFM2.5-Audio conversation panel for speech-to-speech The LFM2.5-Audio speech-to-speech entries ran one question per Run, as a new conversation each time, and showed the reply only when it was done. The server can now carry a conversation (earlier turns sent as request artifacts) and stream a reply as server-sent events, so this gives the two S2S entries a page of their own, as asked on #828. The panel covers only the S2S entries; ASR and TTS keep the generic controls. It adds a conversation above the page's own Language, Seed, Max tokens, Source audio and Model parameters, which it uses as they are. Each Run: - uploads the recorded or chosen question, and sends it to /v1/tasks/stream with "stream_format": "sse" and return_codes, with each earlier turn as its question's upload path and the reply artifact that turn returned; - plays the reply as it streams in, its 80 ms chunks scheduled back to back on one Web Audio clock, and shows the reply text as it arrives; - adds the turn to a list with the question, the reply text and the reply audio, and clears the source picker for the next question. New conversation starts over, and Leave out drops an old turn from the next request, which is what the server asks for past its 8192-step limit. A failed or stopped turn, or one whose stream closed before its result, is never sent as history; the next Run takes its place. Starting a turn stops the previous reply, and Stop audio stops a reply that is still playing. If the browser does not let live audio start, the panel says so and the reply is in its player when it finishes. Finished turns are kept when the user opens another tab or entry and comes back (leaving the panel stops a turn still running), and a turn on the other checkpoint starts a new conversation. With UI management the entry is loaded in streaming mode when needed; on a server with a config file, an entry in offline mode is run through /v1/tasks/run and its reply shows when it is done. An S2S entry configured under an id that is not the catalog's gets the family's ASR parameters in Model parameters; turns leave those out and send the page's Max tokens. Each turn's players and its Leave out button are named with the turn number for screen readers, and its state badge is a polite live region. The page passes panels three more props for this: modelId, setPanelRunner, which is its registerPanelRunner, and clearSource, its own clearSourceFile. The other panels do not declare them and ignore them. model_params.json gets text_temperature and text_top_k for the two S2S entries, with the server's defaults, and the S2S hints and the model docs describe the conversation. * webui: rebuild the page Rebuilt from this branch's source with the lockfile (node 24.19.0, npm 11.17.0): cd webui/native npm ci npm run build The page on main was built from an older tree. It lacks changes to the model specs or the page from #824, #829, #830, #833, #834, #835, #846, #851, #853 and #854, so this rebuild also brings those into the page. Rebuilding main's own source the same way gives a page that differs from this one only in this branch's changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Python Irodori-TTS cuts a reference clip to the checkpoint's
ref_max_secondsbefore encoding it. That is 120 seconds for v4. Checkpoints that do not state a value, such as v3, get 30 seconds (_default_max_ref_seconds/_LEGACY_MAX_REF_SECONDSinirodori_tts/inference_runtime.py).infer.py --max-ref-secondssets another length per request, and0turns the cap off.audio.cpp encodes the whole clip. So for a reference longer than the cap, the speaker condition is not the one Python computes, and the reference encode needs more memory as the clip gets longer.
This PR applies the same cap by default.
Changes
IrodoriModelConfig::ref_max_secondsis read frommodel_config.json. If the value is missing or not positive, it falls back to 30, as in Python. The published GGUFs:120.int(seconds * sample_rate)frames, at the clip's own sample rate. After encoding, it trims the latent toceil(seconds * 48000 / 1920)steps. Python does both steps the same way.max_ref_sec(float,>= 0) sets another length. It is named likeduration_sec; Python calls itmax_ref_seconds.0keeps the whole clip, which is the old behaviour. Likeinfer.py, it is set per request, not per session.speaker_embedding_pathis.docs/models/irodori_tts.mdgets a short "Reference Length" section and an option row.Behaviour change: only references longer than the cap are affected (over 120 s on v4, over 30 s on v3).
max_ref_sec=0restores the previous output exactly (see below).Validation
Setup:
Clips:
This PR against
main(CUDA):max_ref_sec=0Trimming equals cutting the file beforehand:
max_ref_sec=30(also as the string"30")max_ref_sec=120Python behaves the same way (
infer.py, fp32, CUDA, same text/seed/steps; repeated runs are byte-identical):--max-ref-seconds 30So both implementations now give the codec the same samples. This PR does not change how far the two implementations are apart for reference audio. On v4 with the 34.7 s clip the mel cosine is 0.9724, the same value as the reference-WAV row in #829.
Other checks:
max_ref_sec=30→ default gave each request the same output as when it ran alone.model_specs/(the GGUF's embedded contract), the option is accepted and the outputs are the same.max_ref_sec; the code is otherwise unchanged):max_ref_sec=30: identical to the 30 s pre-cut clip, which is also identical on main.audiocpp_cli --request-option max_ref_sec=30gives the same output as the server.-1→Irodori-TTS max_ref_sec must be non-negative.inf→max_ref_sec must be a finite float.abc→ the float parser's error.max_ref_secondsis rejected as an unknown option.mainrejects the option as unknown.🤖 Generated with Claude Code