Conversation
Issue #188: usage statistics recorded estimated token counts instead of real backend-reported values. Proxy streaming paths did not request include_usage; local paths ignored engine-reported tokens; users could not distinguish real from estimated. - Add OnUsage callback in Options to thread real prompt/completion tokens from llama-server and OpenAI-compatible engines to handlers - Add stream_options.include_usage to proxy request bodies so backends report real usage in stream tails - Add resolveUsageTokens with per-side sanity validation: use real values when available, fall back to estimation otherwise, mark mixed rows as estimated - Improve CJK token estimation (0.6 tokens/char vs flat 4 chars/token) - Capture partial stream content for fallback estimation when upstream reports no usage; record even on mid-stream disconnect - Add estimated_requests column (PRAGMA-checked migration), expose in API/OpenAPI/frontend with amber badge - Move Anthropic native stream recording outside err==nil guard - Add tests for proxy stream usage, fallback estimation, and mid-stream error recording 22 files, +1042/-206
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
用量统计记录后端真实 token 数,并标记估算值
Fixes #188
问题
用量统计页面的 token 数全部是估算值,不是后端真实上报的。代理流式路径没有要求后端附带
usage;本地路径忽略了引擎已上报的真实 token;用户无法区分真实与估算。改动
引擎真值穿透:
Options新增OnUsage回调,llama-server / OpenAI 兼容引擎在响应中拿到的prompt_tokens/completion_tokens通过回调传给 handler,handler 用resolveUsageTokens逐侧判定——有真值用真值,无真值用估算,混合则打"估算"标记。代理流式补 usage:
openAIChatRequestToProxyBody和anthropicRequestToProxyBody流式时加stream_options.include_usage,让后端在流尾返回真实 usage。流中断时仍记录已捕获的部分。回退估算改进:CJK 文本按 0.6 tokens/char 估算(原来一律 4 chars/token,中文严重低估);流式无 usage 时从已捕获的 SSE 内容提取文本做估算并打标记;非流式、工具路径同理补全。
估算标记端到端:DB schema 加
estimated_requests列(用 PRAGMA 检查是否已存在,不盲跑 ALTER);API 响应、OpenAPI 规范、前端类型同步加字段;用量表格中估算行显示 amber "估算" 标签。Anthropic 原生流式:移出
if err == nil守卫,客户端中途断开也记录。不改的项
验证
gofmt·go vet·go build ./...·go test ./...· OpenAPI sync ·npm run build全部通过统计
22 files, +1042 / -206