Features/gpt live provider - #1449
Merged
Merged
Conversation
Integrates OpenAI's full duplex voice endpoint, wss://api.openai.com/v1/live/sessions, as a new IRealTimeCompletion provider under Providers/Live. It registers under the provider key "openai-live" so the existing gpt-realtime provider is untouched, and falls back to the "openai" credentials when no "openai-live" section is configured. Live is a different protocol from the realtime API, and that shapes the implementation: - No turn completion event. A turn is closed by alternation (the other speaker starting), by the model's audio going idle, or by an idle backstop timer. - No interruption callback. Live arbitrates barge-in inside the model, so CancelModelResponse is a no-op and queued playback is never discarded locally. - Reasoning and tools are delegated to a separate backend model. Responses delegation maps tool calls onto the existing routing pipeline; client delegation runs the BotSharp agent itself and answers with session.commentary.append. - The prompt is split in two: conversation style and a delegation policy go to the voice model, the agent instruction goes to the backend handler. - session.*.append is capped at 500 tokens. Appends concatenate, so over-budget content is split on sentence boundaries rather than truncated; only an instruction too large for five appends is dropped, since half a directive can mean the opposite of the whole. - The model opens the call, greeting from the agent's .welcome template when it has one and from configuration otherwise. Also adds a display-only transcript delta channel to RealtimeHubConnection, and makes ChatStreamMiddleware disconnect the model however its receive loop ends. The browser sends "disconnect" and closes the socket in the same breath, so the close often won the race and left the final turn stranded in the provider's buffer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Responses delegation is the only mode now: Live calls the backend model itself and feeds the results back into the call. Client delegation was scaffolded rather than finished. session.delegation.created carries only an id and an offset, never the task text, so answering one meant pairing the delegation with an utterance reconstructed from the transcript stream and guessing at the boundary. That machinery, and its two pairing windows, are not worth carrying while nothing uses it. Removes LiveCompletionProvider.ClientDelegation.cs, the DelegationType setting and its branches, DelegationWaitMs and UtteranceClaimMs, and LiveDelegationType.Client. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a "while the backend is working" section to the default voice prompt. Staying in the conversation while the backend thinks is the point of a full duplex model, but the model has to be told how: acknowledge and stay present rather than fall silent, keep answering whatever it already knows, keep listening in case the request changes, and never narrate the mechanics - no tools, no systems, no "backend". OpenAI's prompting guide covers when to delegate and says not to guess the result while waiting, but leaves conversational bridging to the product, so these rules are ours. Also moves the DefaultVoiceInstruction doc comment back onto its own member; adding DefaultGreeting had left it stranded on the wrong one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Delegating to the backend model made the voice disfluent: the caller heard "Okay, checking", then ten seconds of nothing, then "that now." Not delegating was fine. The socket has one consumer, and a tool call was awaited inside it - through routing.InvokeFunction, the conversation hooks, and the session update that follows. Nothing was read off the socket for as long as that took. On a half duplex session that costs nothing, because the model is silent while a tool runs; on Live the model keeps talking through the delegation, so its audio frames piled up unread and arrived in a burst once the function returned. Telling the model to hold the call made this visible rather than causing it. So the receive loop now parses an event and moves on. Conversation work - invoking a tool, recording a turn - goes to a SerialWorkQueue and runs on one background worker in the order it was queued. Serial rather than fire and forget: two tool calls must not interleave their session updates, and the conversation state this work touches is not thread safe. Audio and live transcript deltas stay on the loop, since queueing them behind a tool call is the bug again. Turn boundaries are still decided on the loop. The flushes take the buffer there and queue only its delivery, so a turn held up by a slow tool cannot pick up the words the model spoke after it, or be filed under the item id it had moved on to. Two hazards that came with it: - Queued work can tear the session down - the hub reconnects from inside a tool call - and draining would then wait on the worker running it. DrainAsync spots that through an AsyncLocal marker, closes the queue and returns. - The old receive loop unwinds after a reconnect has already installed new fields, so its teardown reached the new session and would have closed the new queue outright. ReceiveMessage captures its session and queue up front. This predates the queue: the same loop could dispose the new session, saved only by the delay in Reconnect. Drops the 300ms settle delay after session.update, copied from the realtime provider. The caller's next act is response.create on the same socket and the server applies what it is sent in order, so there was nothing for it to win. Also handles message items from the backend turn, which were falling through the function_call check: the answer it settled on is logged, not recorded, since the assistant turn is still taken from the spoken transcript. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Live and realtime are different kinds of voice session: a live model runs the conversation itself and delegates the thinking to a backend model, so its provider and model are not interchangeable with a realtime one. - ILiveCompletion, deriving from IRealTimeCompletion so a live completer still drives the same hub, middleware and hooks. The separate interface keeps the families apart in DI: GetServices<IRealTimeCompletion>() never returns a live provider, so it is reachable only when an agent asks for one. - llm_config.live on the agent, a sibling of llm_config.realtime, with Mongo mapping. RealtimeHub picks the family from which config is set and resolves the provider name in that family's list alone. - LlmModelType.Live and LlmModelCapability.Live, so live models are typed as their own kind and stay out of realtime lookups. gpt-live-1 is listed under the openai provider and the openai-live entry is gone. - Restores the RealtimeModelSettings fallback that was hardcoded to openai-live/gpt-live-1 while the provider was being built. Also: the voice prompt now comes from the agent's "live" channel instruction rather than a setting, delegation reasoning effort comes from the agent's own llm_config rather than the live one, the socket address can be configured per model entry, and each SerialWorkQueue carries an id so concurrent calls can be told apart in the logs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…atures/gpt-live-provider
Critical in a local debug build, Information otherwise. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Contributor
Qodo reviews are paused for this user.Troubleshooting steps vary by plan Learn more → On a Teams plan? Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center? |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.