Skip to content

Features/gpt live provider - #1449

Merged
iceljc merged 10 commits into
SciSharp:masterfrom
iceljc:features/gpt-live-provider
Sep 23, 2026
Merged

iceljc merged 10 commits into
SciSharp:masterfrom
iceljc:features/gpt-live-provider

Conversation

@iceljc

@iceljc iceljc commented Sep 22, 2026

Copy link
Copy Markdown
Collaborator

No description provided.

jichenglu and others added 10 commits September 16, 2026 17:46
Integrates OpenAI's full duplex voice endpoint, wss://api.openai.com/v1/live/sessions,
as a new IRealTimeCompletion provider under Providers/Live. It registers under the
provider key "openai-live" so the existing gpt-realtime provider is untouched, and
falls back to the "openai" credentials when no "openai-live" section is configured.

Live is a different protocol from the realtime API, and that shapes the implementation:

- No turn completion event. A turn is closed by alternation (the other speaker
  starting), by the model's audio going idle, or by an idle backstop timer.
- No interruption callback. Live arbitrates barge-in inside the model, so
  CancelModelResponse is a no-op and queued playback is never discarded locally.
- Reasoning and tools are delegated to a separate backend model. Responses
  delegation maps tool calls onto the existing routing pipeline; client delegation
  runs the BotSharp agent itself and answers with session.commentary.append.
- The prompt is split in two: conversation style and a delegation policy go to the
  voice model, the agent instruction goes to the backend handler.
- session.*.append is capped at 500 tokens. Appends concatenate, so over-budget
  content is split on sentence boundaries rather than truncated; only an
  instruction too large for five appends is dropped, since half a directive can
  mean the opposite of the whole.
- The model opens the call, greeting from the agent's .welcome template when it
  has one and from configuration otherwise.

Also adds a display-only transcript delta channel to RealtimeHubConnection, and makes
ChatStreamMiddleware disconnect the model however its receive loop ends. The browser
sends "disconnect" and closes the socket in the same breath, so the close often won
the race and left the final turn stranded in the provider's buffer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Responses delegation is the only mode now: Live calls the backend model itself and
feeds the results back into the call.

Client delegation was scaffolded rather than finished. session.delegation.created
carries only an id and an offset, never the task text, so answering one meant pairing
the delegation with an utterance reconstructed from the transcript stream and guessing
at the boundary. That machinery, and its two pairing windows, are not worth carrying
while nothing uses it.

Removes LiveCompletionProvider.ClientDelegation.cs, the DelegationType setting and its
branches, DelegationWaitMs and UtteranceClaimMs, and LiveDelegationType.Client.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a "while the backend is working" section to the default voice prompt. Staying in
the conversation while the backend thinks is the point of a full duplex model, but the
model has to be told how: acknowledge and stay present rather than fall silent, keep
answering whatever it already knows, keep listening in case the request changes, and
never narrate the mechanics - no tools, no systems, no "backend".

OpenAI's prompting guide covers when to delegate and says not to guess the result while
waiting, but leaves conversational bridging to the product, so these rules are ours.

Also moves the DefaultVoiceInstruction doc comment back onto its own member; adding
DefaultGreeting had left it stranded on the wrong one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Delegating to the backend model made the voice disfluent: the caller heard "Okay,
checking", then ten seconds of nothing, then "that now." Not delegating was fine.

The socket has one consumer, and a tool call was awaited inside it - through
routing.InvokeFunction, the conversation hooks, and the session update that follows.
Nothing was read off the socket for as long as that took. On a half duplex session
that costs nothing, because the model is silent while a tool runs; on Live the model
keeps talking through the delegation, so its audio frames piled up unread and arrived
in a burst once the function returned. Telling the model to hold the call made this
visible rather than causing it.

So the receive loop now parses an event and moves on. Conversation work - invoking a
tool, recording a turn - goes to a SerialWorkQueue and runs on one background worker
in the order it was queued. Serial rather than fire and forget: two tool calls must
not interleave their session updates, and the conversation state this work touches is
not thread safe. Audio and live transcript deltas stay on the loop, since queueing
them behind a tool call is the bug again.

Turn boundaries are still decided on the loop. The flushes take the buffer there and
queue only its delivery, so a turn held up by a slow tool cannot pick up the words the
model spoke after it, or be filed under the item id it had moved on to.

Two hazards that came with it:

- Queued work can tear the session down - the hub reconnects from inside a tool call -
  and draining would then wait on the worker running it. DrainAsync spots that through
  an AsyncLocal marker, closes the queue and returns.
- The old receive loop unwinds after a reconnect has already installed new fields, so
  its teardown reached the new session and would have closed the new queue outright.
  ReceiveMessage captures its session and queue up front. This predates the queue: the
  same loop could dispose the new session, saved only by the delay in Reconnect.

Drops the 300ms settle delay after session.update, copied from the realtime provider.
The caller's next act is response.create on the same socket and the server applies
what it is sent in order, so there was nothing for it to win.

Also handles message items from the backend turn, which were falling through the
function_call check: the answer it settled on is logged, not recorded, since the
assistant turn is still taken from the spoken transcript.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Live and realtime are different kinds of voice session: a live model runs the
conversation itself and delegates the thinking to a backend model, so its
provider and model are not interchangeable with a realtime one.

- ILiveCompletion, deriving from IRealTimeCompletion so a live completer still
  drives the same hub, middleware and hooks. The separate interface keeps the
  families apart in DI: GetServices<IRealTimeCompletion>() never returns a live
  provider, so it is reachable only when an agent asks for one.
- llm_config.live on the agent, a sibling of llm_config.realtime, with Mongo
  mapping. RealtimeHub picks the family from which config is set and resolves
  the provider name in that family's list alone.
- LlmModelType.Live and LlmModelCapability.Live, so live models are typed as
  their own kind and stay out of realtime lookups. gpt-live-1 is listed under
  the openai provider and the openai-live entry is gone.
- Restores the RealtimeModelSettings fallback that was hardcoded to
  openai-live/gpt-live-1 while the provider was being built.

Also: the voice prompt now comes from the agent's "live" channel instruction
rather than a setting, delegation reasoning effort comes from the agent's own
llm_config rather than the live one, the socket address can be configured per
model entry, and each SerialWorkQueue carries an id so concurrent calls can be
told apart in the logs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Critical in a local debug build, Information otherwise.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@qodo-code-review

Copy link
Copy Markdown
Contributor

Qodo reviews are paused for this user.

Troubleshooting steps vary by plan Learn more →

On a Teams plan?
Reviews resume once this user has a paid seat and their Git account is linked in Qodo.
Link Git account →

Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center?
These require an Enterprise plan - Contact us
Contact us →

@iceljc
iceljc merged commit eb23532 into SciSharp:master Sep 23, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants