Skip to content

feat(webui): LFM2.5-Audio conversation panel for speech-to-speech - #857

Merged
0xShug0 merged 3 commits into
0xShug0:mainfrom
Liquid4All:webui-lfm2-audio-panel
Oct 10, 2026
Merged

0xShug0 merged 3 commits into
0xShug0:mainfrom
Liquid4All:webui-lfm2-audio-panel

Conversation

@ykhrustalev

Copy link
Copy Markdown
Contributor

Follow-up to #828, where you asked for a dedicated LFM page like YuE2. The LFM2.5-Audio speech-to-speech entries (EN and JP) now get a Studio panel that holds a conversation: you record or pick a question, hear the reply as it streams in, and keep talking. ASR, TTS and other models work as before.

Problem

What this PR changes

What changes for users

  • Replies stream in as speech and text. On an M3 Ultra with Q8_0 and the server on the same machine, the first audio played 0.5 to 0.8 s after pressing Run
  • A failed turn (step limit, stream closed early) is not sent as history; the next Run takes its place. After a server restart the uploaded questions are gone, so the panel asks you to start a new conversation
  • Finished turns are kept when you switch tabs or entries (a reply still streaming stops). Reloading the page, or running a turn on the other checkpoint (EN or JP), starts a new conversation
  • With --ui-management, the panel loads the entry in streaming mode. With a config file, live playback needs "mode": "streaming"; an offline entry works, but replies show only when done
  • S2S turns no longer go to Run history and have no Save settings JSON

Left for later

  • A wider LFM2.5-Audio page for TTS and ASR, with a voice picker and per-task defaults (the generic controls already list the TTS voices)
  • Catalog "mode": "streaming" for the S2S entries, so Load does not mean a reload on the first Run
  • Saving and reopening a conversation, and recovering one after a server restart
  • Live push-to-talk from the microphone (the live route takes no history yet)
  • Translations of the panel's text

Testing
All on an M3 Ultra (Metal, Q8_0): the server built from this branch, and headless Chromium 153 driven by Playwright with real clicks.

  • npm run check (0 errors, 0 warnings) and npm run build pass at each commit. Building main twice gave the same page byte for byte, so commit 3 can be checked by rebuilding it
  • Conversations: EN with 5 turns, EN with 2 (the first recorded through Chromium's fake microphone, the second from a file), and JP with 2 at text_temperature 0.5 and text_top_k 20. The history sent back was byte-equal to what the server returned, the played audio matched the result sample for sample, and a plain Python client replaying the turns got the same replies
  • The step limit (Max tokens 8192), a failed upload, Cancel then Run, and a reply cut at max_tokens sent back as history all left the conversation usable. With the server stopped mid-reply, the turn showed as stopped and was not sent as history. After the restart the next turn said the uploads were gone, and New conversation worked
  • With autoplay blocked, the panel says so; the reply is in its player when done
  • --ui-management: Run with nothing loaded loads the entry once, in streaming mode, and an entry loaded offline or with another package is reloaded
  • A config entry named lfm2-audio-s2s-stream (the README's SSE example) works with default settings; main's page fails on it with a max_tokens conflict
  • Firefox 155 gives the same replies as Chromium
  • Regression: TTS and ASR requests and outputs match main's page, and YuE2, AuK, PersonaPlex and the LFM TTS/ASR entries show the same controls
  • Not tested: Safari, a real microphone, CPU and CUDA backends, conversations longer than 5 turns, the parallel runtime

Build and run

cd webui/native && npm ci && npm run check && npm run build && cd ../..
cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release \
  -DAUDIOCPP_MODEL_SET=custom -DAUDIOCPP_MODELS=lfm2_audio \
  -DAUDIOCPP_BUILD_NATIVE_MODEL_MANAGER=ON -DENGINE_ENABLE_OPENMP=OFF
cmake --build build --target audiocpp_server -j 16
./build/bin/audiocpp_server --config server.json --ui --backend metal --threads 8 --log
# the --ui-management runs, with no config file:
./build/bin/audiocpp_server --ui --ui-management --backend metal --threads 8 --log

OpenMP was off only because AppleClang has none. Without --config the backend defaults to CUDA, so pass --backend metal. Example server.json:

{"backend": "metal", "threads": 8,
 "models": [
  {"id": "lfm2-audio-s2s", "family": "lfm2_audio", "task": "s2s", "mode": "streaming",
   "path": "models/LFM2.5-Audio-1.5B-GGUF",
   "session_options": {"lfm2_audio.model_gguf": "LFM2.5-Audio-1.5B-Q8_0.gguf"}},
  {"id": "lfm2-audio-jp-s2s", "family": "lfm2_audio", "task": "s2s", "mode": "streaming",
   "path": "models/LFM2.5-Audio-1.5B-JP-GGUF",
   "session_options": {"lfm2_audio.model_gguf": "LFM2.5-Audio-1.5B-JP-Q8_0.gguf"}}]}

Then choose an LFM2.5-Audio S2S entry under Voice Conversion / S2S, record or pick a question, and press Run.

Model panels are registered per family, and the page always builds and
sends the request itself. A conversation view has to send its own
request and show the reply as it streams in, while it keeps the page's
Run button, Ctrl+Enter, Cancel, elapsed time and status line.

This adds that hook without changing what any existing panel does:

- A panel entry may list the tasks it covers, such as tasks: ['s2s'].
  An entry without the list covers every task of its family, as before.
- With requestMode 'panel', run() hands the request to the runner the
  panel registers through the page's registerPanelRunner, which the
  next commit passes to panels as setPanelRunner; registering returns
  the function that removes it. It passes the source audio, the
  resolved seed, Max tokens, Language, the request options and the
  abort signal that Cancel uses. The hand-off comes before anything is
  awaited, so the panel can start audio inside the Run click or key
  press. The reply the panel returns is shown in the Result column;
  these runs are not added to Run history. Missing source audio, a
  recording still running, or Run from a tab other than Studio is
  reported as a warning, like the page's other checks before a run.
- With UI management the panel can ask for the entry in a given mode.
  This uses the page's ensureLoadedMode unchanged, after reloading an
  entry that is resident with another package or with imported
  settings that differ, as ensureLoaded does for the other runs.
- api.ts gets taskStreamEvents, a reader for /v1/tasks/stream with
  "stream_format": "sse". It yields each event and then the result,
  throws the server's error message, and throws TaskStreamClosedError
  when the stream ends before task.stream.done. The done message
  carries the whole reply again and comes over many reads, so the
  reader keeps its pieces and joins them once, when the message ends.

No panel uses requestMode 'panel' yet.
The LFM2.5-Audio speech-to-speech entries ran one question per Run, as a
new conversation each time, and showed the reply only when it was done.
The server can now carry a conversation (earlier turns sent as request
artifacts) and stream a reply as server-sent events, so this gives the
two S2S entries a page of their own, as asked on 0xShug0#828.

The panel covers only the S2S entries; ASR and TTS keep the generic
controls. It adds a conversation above the page's own Language, Seed,
Max tokens, Source audio and Model parameters, which it uses as they
are. Each Run:

- uploads the recorded or chosen question, and sends it to
  /v1/tasks/stream with "stream_format": "sse" and return_codes, with
  each earlier turn as its question's upload path and the reply
  artifact that turn returned;
- plays the reply as it streams in, its 80 ms chunks scheduled back to
  back on one Web Audio clock, and shows the reply text as it arrives;
- adds the turn to a list with the question, the reply text and the
  reply audio, and clears the source picker for the next question.

New conversation starts over, and Leave out drops an old turn from the
next request, which is what the server asks for past its 8192-step
limit. A failed or stopped turn, or one whose stream closed before its
result, is never sent as history; the next Run takes its place.
Starting a turn stops the previous reply, and Stop audio stops a reply
that is still playing. If the browser does not let live audio start,
the panel says so and the reply is in its player when it finishes.
Finished turns are kept when the user opens another tab or entry and
comes back (leaving the panel stops a turn still running), and a turn
on the other checkpoint starts a new conversation. With UI management
the entry is loaded in streaming mode when needed; on a server with a
config file, an entry in offline mode is run through /v1/tasks/run and
its reply shows when it is done. An S2S entry configured under an id
that is not the catalog's gets the family's ASR parameters in Model
parameters; turns leave those out and send the page's Max tokens.

Each turn's players and its Leave out button are named with the turn
number for screen readers, and its state badge is a polite live
region.

The page passes panels three more props for this: modelId,
setPanelRunner, which is its registerPanelRunner, and clearSource, its
own clearSourceFile. The other panels do not declare them and ignore
them. model_params.json gets text_temperature and text_top_k for the
two S2S entries, with the server's defaults, and the S2S hints and the
model docs describe the conversation.
Rebuilt from this branch's source with the lockfile (node 24.19.0,
npm 11.17.0):

  cd webui/native
  npm ci
  npm run build

The page on main was built from an older tree. It lacks changes to the
model specs or the page from 0xShug0#824, 0xShug0#829, 0xShug0#830, 0xShug0#833, 0xShug0#834, 0xShug0#835, 0xShug0#846,
0xShug0#851, 0xShug0#853 and 0xShug0#854, so this rebuild also brings those into the page.
Rebuilding main's own source the same way gives a page that differs
from this one only in this branch's changes.
@0xShug0
0xShug0 merged commit ef5c685 into 0xShug0:main Oct 10, 2026
6 checks passed
@0xShug0

0xShug0 commented Oct 10, 2026

Copy link
Copy Markdown
Owner

@ykhrustalev Merged 🎉 !

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants