Skip to content

fix(usage): record real backend token counts and mark estimated rows - #201

Open
apapi wants to merge 1 commit into
mainfrom
fix/issue-188-real-usage-tokens
Open

apapi wants to merge 1 commit into
mainfrom
fix/issue-188-real-usage-tokens

Conversation

@apapi

@apapi apapi commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator

用量统计记录后端真实 token 数,并标记估算值

Fixes #188

问题

用量统计页面的 token 数全部是估算值,不是后端真实上报的。代理流式路径没有要求后端附带 usage;本地路径忽略了引擎已上报的真实 token;用户无法区分真实与估算。

改动

  1. 引擎真值穿透:Options 新增 OnUsage 回调,llama-server / OpenAI 兼容引擎在响应中拿到的 prompt_tokens / completion_tokens 通过回调传给 handler,handler 用 resolveUsageTokens 逐侧判定——有真值用真值,无真值用估算,混合则打"估算"标记。

  2. 代理流式补 usage:openAIChatRequestToProxyBody 和 anthropicRequestToProxyBody 流式时加 stream_options.include_usage,让后端在流尾返回真实 usage。流中断时仍记录已捕获的部分。

  3. 回退估算改进:CJK 文本按 0.6 tokens/char 估算(原来一律 4 chars/token,中文严重低估);流式无 usage 时从已捕获的 SSE 内容提取文本做估算并打标记;非流式、工具路径同理补全。

  4. 估算标记端到端:DB schema 加 estimated_requests 列(用 PRAGMA 检查是否已存在,不盲跑 ALTER);API 响应、OpenAPI 规范、前端类型同步加字段;用量表格中估算行显示 amber "估算" 标签。

  5. Anthropic 原生流式:移出 if err == nil 守卫,客户端中途断开也记录。

不改的项

  • 三处非代理 OnUsage 对真实引擎是死代码(所有 Engine 都实现了 ChatCompletionProxier),保留作 safety net
  • 64KB tail 限制(只影响无 usage 时的回退估算精度)
  • 内容提取不含 reasoning_content / thinking_delta(低优先级)
  • estimated_requests 不在 key/source/pool 汇总层面暴露(后续迭代)
  • 历史数据不回填

验证

gofmt · go vet · go build ./... · go test ./... · OpenAPI sync · npm run build 全部通过

统计

22 files, +1042 / -206

Issue #188: usage statistics recorded estimated token counts instead of
real backend-reported values. Proxy streaming paths did not request
include_usage; local paths ignored engine-reported tokens; users could
not distinguish real from estimated.

- Add OnUsage callback in Options to thread real prompt/completion
  tokens from llama-server and OpenAI-compatible engines to handlers
- Add stream_options.include_usage to proxy request bodies so backends
  report real usage in stream tails
- Add resolveUsageTokens with per-side sanity validation: use real
  values when available, fall back to estimation otherwise, mark mixed
  rows as estimated
- Improve CJK token estimation (0.6 tokens/char vs flat 4 chars/token)
- Capture partial stream content for fallback estimation when upstream
  reports no usage; record even on mid-stream disconnect
- Add estimated_requests column (PRAGMA-checked migration), expose in
  API/OpenAPI/frontend with amber badge
- Move Anthropic native stream recording outside err==nil guard
- Add tests for proxy stream usage, fallback estimation, and mid-stream
  error recording

22 files, +1042/-206

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

API_KEY的Token 用量统计异常:输入 Tokens 显示为请求次数,输出 Tokens 始终为 0

1 participant