Repository navigation
feat(webui): LFM2.5-Audio conversation panel for speech-to-speech - #857
Merged
Merged
Conversation
Model panels are registered per family, and the page always builds and sends the request itself. A conversation view has to send its own request and show the reply as it streams in, while it keeps the page's Run button, Ctrl+Enter, Cancel, elapsed time and status line. This adds that hook without changing what any existing panel does: - A panel entry may list the tasks it covers, such as tasks: ['s2s']. An entry without the list covers every task of its family, as before. - With requestMode 'panel', run() hands the request to the runner the panel registers through the page's registerPanelRunner, which the next commit passes to panels as setPanelRunner; registering returns the function that removes it. It passes the source audio, the resolved seed, Max tokens, Language, the request options and the abort signal that Cancel uses. The hand-off comes before anything is awaited, so the panel can start audio inside the Run click or key press. The reply the panel returns is shown in the Result column; these runs are not added to Run history. Missing source audio, a recording still running, or Run from a tab other than Studio is reported as a warning, like the page's other checks before a run. - With UI management the panel can ask for the entry in a given mode. This uses the page's ensureLoadedMode unchanged, after reloading an entry that is resident with another package or with imported settings that differ, as ensureLoaded does for the other runs. - api.ts gets taskStreamEvents, a reader for /v1/tasks/stream with "stream_format": "sse". It yields each event and then the result, throws the server's error message, and throws TaskStreamClosedError when the stream ends before task.stream.done. The done message carries the whole reply again and comes over many reads, so the reader keeps its pieces and joins them once, when the message ends. No panel uses requestMode 'panel' yet.
The LFM2.5-Audio speech-to-speech entries ran one question per Run, as a new conversation each time, and showed the reply only when it was done. The server can now carry a conversation (earlier turns sent as request artifacts) and stream a reply as server-sent events, so this gives the two S2S entries a page of their own, as asked on 0xShug0#828. The panel covers only the S2S entries; ASR and TTS keep the generic controls. It adds a conversation above the page's own Language, Seed, Max tokens, Source audio and Model parameters, which it uses as they are. Each Run: - uploads the recorded or chosen question, and sends it to /v1/tasks/stream with "stream_format": "sse" and return_codes, with each earlier turn as its question's upload path and the reply artifact that turn returned; - plays the reply as it streams in, its 80 ms chunks scheduled back to back on one Web Audio clock, and shows the reply text as it arrives; - adds the turn to a list with the question, the reply text and the reply audio, and clears the source picker for the next question. New conversation starts over, and Leave out drops an old turn from the next request, which is what the server asks for past its 8192-step limit. A failed or stopped turn, or one whose stream closed before its result, is never sent as history; the next Run takes its place. Starting a turn stops the previous reply, and Stop audio stops a reply that is still playing. If the browser does not let live audio start, the panel says so and the reply is in its player when it finishes. Finished turns are kept when the user opens another tab or entry and comes back (leaving the panel stops a turn still running), and a turn on the other checkpoint starts a new conversation. With UI management the entry is loaded in streaming mode when needed; on a server with a config file, an entry in offline mode is run through /v1/tasks/run and its reply shows when it is done. An S2S entry configured under an id that is not the catalog's gets the family's ASR parameters in Model parameters; turns leave those out and send the page's Max tokens. Each turn's players and its Leave out button are named with the turn number for screen readers, and its state badge is a polite live region. The page passes panels three more props for this: modelId, setPanelRunner, which is its registerPanelRunner, and clearSource, its own clearSourceFile. The other panels do not declare them and ignore them. model_params.json gets text_temperature and text_top_k for the two S2S entries, with the server's defaults, and the S2S hints and the model docs describe the conversation.
Rebuilt from this branch's source with the lockfile (node 24.19.0, npm 11.17.0): cd webui/native npm ci npm run build The page on main was built from an older tree. It lacks changes to the model specs or the page from 0xShug0#824, 0xShug0#829, 0xShug0#830, 0xShug0#833, 0xShug0#834, 0xShug0#835, 0xShug0#846, 0xShug0#851, 0xShug0#853 and 0xShug0#854, so this rebuild also brings those into the page. Rebuilding main's own source the same way gives a page that differs from this one only in this branch's changes.
Owner
|
@ykhrustalev Merged 🎉 ! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #828, where you asked for a dedicated LFM page like YuE2. The LFM2.5-Audio speech-to-speech entries (EN and JP) now get a Studio panel that holds a conversation: you record or pick a question, hear the reply as it streams in, and keep talking. ASR, TTS and other models work as before.
Problem
What this PR changes
tasks: ['s2s']) and, withrequestMode: 'panel', send the request itself while the page's Run, Ctrl+Enter, Cancel and status line keep working.api.tsgets an SSE reader for/v1/tasks/streamwebui/native/src/lib/models/lfm2_audio/). Each Run uploads the question, sends it with the earlier turns as artifacts, and plays the reply audio and text as they arrive. It adds New conversation, Leave out (drops an old turn, e.g. near the step limit) and Stop audio, and reuses the page's Seed, Max tokens, Source audio and Model parameters, where the S2S entries gaintext_temperatureandtext_top_kwebui/native/dist/index.html(node 24.19.0, npm 11.17.0). Main's committed page is stale (built from an older tree), so this also brings in the UI and model-spec changes of feat(asr): add experimental Whistle community model #824, irodori_tts: accept Speaker Inversion embeddings as the speaker condition #829, Add TF-GridNet and MossFormer2 speech separation #830, Add Hviske v6 Danish ASR support #833, perf(lfm2_audio): use ggml's repacked CPU kernels for quantized weights #834, irodori_tts: opt-in chunked codec decode and reference encode to bound peak memory #835, Add FireRed VAD and streaming support #846, Add native wake word and audio classification models #851, perf(lfm2_audio): keep several conversations in an S2S session #853 and feat(irodori_tts): cap reference audio length like Python (max_ref_sec) #854. If it conflicts, drop it and rebuild with the commands belowWhat changes for users
--ui-management, the panel loads the entry in streaming mode. With a config file, live playback needs"mode": "streaming"; an offline entry works, but replies show only when doneLeft for later
"mode": "streaming"for the S2S entries, so Load does not mean a reload on the first RunTesting
All on an M3 Ultra (Metal, Q8_0): the server built from this branch, and headless Chromium 153 driven by Playwright with real clicks.
npm run check(0 errors, 0 warnings) andnpm run buildpass at each commit. Building main twice gave the same page byte for byte, so commit 3 can be checked by rebuilding ittext_temperature0.5 andtext_top_k20. The history sent back was byte-equal to what the server returned, the played audio matched the result sample for sample, and a plain Python client replaying the turns got the same replies--ui-management: Run with nothing loaded loads the entry once, in streaming mode, and an entry loaded offline or with another package is reloadedlfm2-audio-s2s-stream(the README's SSE example) works with default settings; main's page fails on it with a max_tokens conflictBuild and run
OpenMP was off only because AppleClang has none. Without
--configthe backend defaults to CUDA, so pass--backend metal. Exampleserver.json:{"backend": "metal", "threads": 8, "models": [ {"id": "lfm2-audio-s2s", "family": "lfm2_audio", "task": "s2s", "mode": "streaming", "path": "models/LFM2.5-Audio-1.5B-GGUF", "session_options": {"lfm2_audio.model_gguf": "LFM2.5-Audio-1.5B-Q8_0.gguf"}}, {"id": "lfm2-audio-jp-s2s", "family": "lfm2_audio", "task": "s2s", "mode": "streaming", "path": "models/LFM2.5-Audio-1.5B-JP-GGUF", "session_options": {"lfm2_audio.model_gguf": "LFM2.5-Audio-1.5B-JP-Q8_0.gguf"}}]}Then choose an LFM2.5-Audio S2S entry under Voice Conversion / S2S, record or pick a question, and press Run.