Skip to content

feat(irodori_tts): cap reference audio length like Python (max_ref_sec) - #854

Merged
0xShug0 merged 2 commits into
0xShug0:mainfrom
KKTTSPJ:pr/irodori-max-ref-seconds
Oct 10, 2026
Merged

0xShug0 merged 2 commits into
0xShug0:mainfrom
KKTTSPJ:pr/irodori-max-ref-seconds

Conversation

@KKTTSPJ

@KKTTSPJ KKTTSPJ commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Python Irodori-TTS cuts a reference clip to the checkpoint's ref_max_seconds before encoding it. That is 120 seconds for v4. Checkpoints that do not state a value, such as v3, get 30 seconds (_default_max_ref_seconds / _LEGACY_MAX_REF_SECONDS in irodori_tts/inference_runtime.py). infer.py --max-ref-seconds sets another length per request, and 0 turns the cap off.

audio.cpp encodes the whole clip. So for a reference longer than the cap, the speaker condition is not the one Python computes, and the reference encode needs more memory as the clip gets longer.

This PR applies the same cap by default.

Changes

  • IrodoriModelConfig::ref_max_seconds is read from model_config.json. If the value is missing or not positive, it falls back to 30, as in Python. The published GGUFs:
    • v4 Small, v4.1 Small and v4.1 Anime carry 120.
    • The 500M v3 and 600M v3 VoiceDesign GGUFs carry none, so they get 30.
  • Trimming before encode: the session cuts the reference to its first int(seconds * sample_rate) frames, at the clip's own sample rate. After encoding, it trims the latent to ceil(seconds * 48000 / 1920) steps. Python does both steps the same way.
  • New request option max_ref_sec (float, >= 0) sets another length. It is named like duration_sec; Python calls it max_ref_seconds. 0 keeps the whole clip, which is the old behaviour. Like infer.py, it is set per request, not per session.
  • Reference cache key: it is now computed from the trimmed audio plus the latent cap. So the same clip sent with different caps does not reuse a cache entry.
  • Published GGUFs embed a contract that does not declare the option. In that case it is dropped from the validation copy, as speaker_embedding_path is.
  • Unchanged: speaker embeddings and no-reference requests.
  • Docs: docs/models/irodori_tts.md gets a short "Reference Length" section and an option row.

Behaviour change: only references longer than the cap are affected (over 120 s on v4, over 30 s on v3). max_ref_sec=0 restores the previous output exactly (see below).

Validation

Setup:

  • Models: v4 Small q8_0 GGUF and 500M v3 q8_0 GGUF.
  • Generation: seed 42, 48 steps, the same text throughout.
  • Backends: CUDA (RTX 5060 Ti) and CPU, through the server unless noted.
  • "Identical" means a byte-identical WAV.

Clips:

  • A 34.7 s clip (44.1 kHz, stereo, 16-bit).
  • A 138.8 s clip made by repeating it 4 times.
  • Copies of each cut beforehand to their first 30 s / 120 s (the same frames the trim keeps).

This PR against main (CUDA):

model reference default max_ref_sec=0
v4 34.7 s clip (under the cap) identical to main identical to main
v4 138.8 s clip differs from main; identical to the 120 s pre-cut clip (on this PR and on main) identical to main
v3 34.7 s clip differs from main; identical to the 30 s pre-cut clip (on this PR and on main) identical to main
v4 speaker embedding / no reference identical to main identical (option has no effect)

Trimming equals cutting the file beforehand:

request result
v4, 34.7 s clip, max_ref_sec=30 (also as the string "30") identical to the 30 s pre-cut clip
v4, 138.8 s clip, max_ref_sec=120 identical to the default

Python behaves the same way (infer.py, fp32, CUDA, same text/seed/steps; repeated runs are byte-identical):

model default output identical to
v4 138.8 s clip the 120 s pre-cut clip
v3 34.7 s clip the 30 s pre-cut clip
v4 34.7 s clip with --max-ref-seconds 30 the 30 s pre-cut clip

So both implementations now give the codec the same samples. This PR does not change how far the two implementations are apart for reference audio. On v4 with the 34.7 s clip the mel cosine is 0.9724, the same value as the reference-WAV row in #829.

Other checks:

  • Cache: in one server, the same clip sent as default → max_ref_sec=30 → default gave each request the same output as when it ran alone.
  • Embedded contract: with no model_specs/ (the GGUF's embedded contract), the option is accepted and the outputs are the same.
  • CPU (run before the option was renamed to max_ref_sec; the code is otherwise unchanged):
    • v4, 34.7 s clip, default: identical to main.
    • max_ref_sec=30: identical to the 30 s pre-cut clip, which is also identical on main.
    • Speaker embedding: identical to main.
  • CLI: audiocpp_cli --request-option max_ref_sec=30 gives the same output as the server.
  • Invalid values:
    • -1 → Irodori-TTS max_ref_sec must be non-negative.
    • inf → max_ref_sec must be a finite float.
    • abc → the float parser's error.
    • The Python-style name max_ref_seconds is rejected as an unknown option.
    • Before this PR, main rejects the option as unknown.

🤖 Generated with Claude Code

…conds)

Python Irodori-TTS cuts a single reference clip to the checkpoint's
ref_max_seconds before encoding it (120 s for v4; 30 s for checkpoints that
do not state it, such as v3), and trims the latent to
ceil(seconds * 48000 / 1920) frames. audio.cpp encoded the whole clip, so a
long reference gave a different speaker condition than Python, and its
encode memory kept growing with the clip length.

The session now applies the same cap by default. The checkpoint value is
read from model_config.json. A new request option, max_ref_seconds, sets
another length like infer.py --max-ref-seconds; 0 keeps the whole
reference. The reference cache key is taken from the trimmed audio plus the
latent cap. Speaker embeddings and no-reference requests are unaffected.
Published GGUFs embed a contract without the option, so it is dropped from
the validation copy like speaker_embedding_path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@0xShug0

0xShug0 commented Oct 9, 2026

Copy link
Copy Markdown
Owner

@KKTTSPJ Could you normalize the option name to max_ref_sec? Thanks!

Follows the *_sec naming of the other duration options (duration_sec,
min_duration_sec, max_duration_sec). The docs still name Python's
max_ref_seconds for reference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@KKTTSPJ KKTTSPJ changed the title feat(irodori_tts): cap reference audio length like Python (max_ref_seconds) feat(irodori_tts): cap reference audio length like Python (max_ref_sec) Oct 10, 2026
@KKTTSPJ

KKTTSPJ commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

Renamed to max_ref_sec (request option, spec, docs and error messages; the docs still mention Python's max_ref_seconds). Outputs are unchanged, and the old name is now rejected as an unknown option. Thanks!

@0xShug0
0xShug0 merged commit 4f92afe into 0xShug0:main Oct 10, 2026
6 checks passed
@0xShug0

0xShug0 commented Oct 10, 2026

Copy link
Copy Markdown
Owner

@KKTTSPJ Thanks! PR merged.

@KKTTSPJ
KKTTSPJ deleted the pr/irodori-max-ref-seconds branch October 10, 2026 00:27
0xShug0 pushed a commit that referenced this pull request Oct 10, 2026
* webui: let a studio panel run the page's request

Model panels are registered per family, and the page always builds and
sends the request itself. A conversation view has to send its own
request and show the reply as it streams in, while it keeps the page's
Run button, Ctrl+Enter, Cancel, elapsed time and status line.

This adds that hook without changing what any existing panel does:

- A panel entry may list the tasks it covers, such as tasks: ['s2s'].
  An entry without the list covers every task of its family, as before.
- With requestMode 'panel', run() hands the request to the runner the
  panel registers through the page's registerPanelRunner, which the
  next commit passes to panels as setPanelRunner; registering returns
  the function that removes it. It passes the source audio, the
  resolved seed, Max tokens, Language, the request options and the
  abort signal that Cancel uses. The hand-off comes before anything is
  awaited, so the panel can start audio inside the Run click or key
  press. The reply the panel returns is shown in the Result column;
  these runs are not added to Run history. Missing source audio, a
  recording still running, or Run from a tab other than Studio is
  reported as a warning, like the page's other checks before a run.
- With UI management the panel can ask for the entry in a given mode.
  This uses the page's ensureLoadedMode unchanged, after reloading an
  entry that is resident with another package or with imported
  settings that differ, as ensureLoaded does for the other runs.
- api.ts gets taskStreamEvents, a reader for /v1/tasks/stream with
  "stream_format": "sse". It yields each event and then the result,
  throws the server's error message, and throws TaskStreamClosedError
  when the stream ends before task.stream.done. The done message
  carries the whole reply again and comes over many reads, so the
  reader keeps its pieces and joins them once, when the message ends.

No panel uses requestMode 'panel' yet.

* webui: LFM2.5-Audio conversation panel for speech-to-speech

The LFM2.5-Audio speech-to-speech entries ran one question per Run, as a
new conversation each time, and showed the reply only when it was done.
The server can now carry a conversation (earlier turns sent as request
artifacts) and stream a reply as server-sent events, so this gives the
two S2S entries a page of their own, as asked on #828.

The panel covers only the S2S entries; ASR and TTS keep the generic
controls. It adds a conversation above the page's own Language, Seed,
Max tokens, Source audio and Model parameters, which it uses as they
are. Each Run:

- uploads the recorded or chosen question, and sends it to
  /v1/tasks/stream with "stream_format": "sse" and return_codes, with
  each earlier turn as its question's upload path and the reply
  artifact that turn returned;
- plays the reply as it streams in, its 80 ms chunks scheduled back to
  back on one Web Audio clock, and shows the reply text as it arrives;
- adds the turn to a list with the question, the reply text and the
  reply audio, and clears the source picker for the next question.

New conversation starts over, and Leave out drops an old turn from the
next request, which is what the server asks for past its 8192-step
limit. A failed or stopped turn, or one whose stream closed before its
result, is never sent as history; the next Run takes its place.
Starting a turn stops the previous reply, and Stop audio stops a reply
that is still playing. If the browser does not let live audio start,
the panel says so and the reply is in its player when it finishes.
Finished turns are kept when the user opens another tab or entry and
comes back (leaving the panel stops a turn still running), and a turn
on the other checkpoint starts a new conversation. With UI management
the entry is loaded in streaming mode when needed; on a server with a
config file, an entry in offline mode is run through /v1/tasks/run and
its reply shows when it is done. An S2S entry configured under an id
that is not the catalog's gets the family's ASR parameters in Model
parameters; turns leave those out and send the page's Max tokens.

Each turn's players and its Leave out button are named with the turn
number for screen readers, and its state badge is a polite live
region.

The page passes panels three more props for this: modelId,
setPanelRunner, which is its registerPanelRunner, and clearSource, its
own clearSourceFile. The other panels do not declare them and ignore
them. model_params.json gets text_temperature and text_top_k for the
two S2S entries, with the server's defaults, and the S2S hints and the
model docs describe the conversation.

* webui: rebuild the page

Rebuilt from this branch's source with the lockfile (node 24.19.0,
npm 11.17.0):

  cd webui/native
  npm ci
  npm run build

The page on main was built from an older tree. It lacks changes to the
model specs or the page from #824, #829, #830, #833, #834, #835, #846,
#851, #853 and #854, so this rebuild also brings those into the page.
Rebuilding main's own source the same way gives a page that differs
from this one only in this branch's changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants