Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 17 additions & 4 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -127,7 +127,7 @@ MULTICA_TRUSTED_PROXIES=
# URL, task token — are injected by the daemon separately.) Nothing below
# governs that path.
#
# Once configured, the layer has exactly two consumers, and this is what each
# Once configured, the layer has three consumers, and this is what each
# one sends upstream:
# - Chat auto-titling: the first user message of a new chat session, verbatim
# and uncapped. Attachments are never included.
Expand All @@ -136,13 +136,19 @@ MULTICA_TRUSTED_PROXIES=
# answered capped at 3000 characters (2000 head + 1000 tail) and each
# older message at 800.
#
# - Opt-in Lark group knowledge: filename and up to 256 KiB extracted PDF/DOCX
# text for classification/assessment; an 8 KiB question and up to six
# 5000-rune group sources for Q&A. No private assessments in group answers.
# Capture/reconciliation makes no model calls. Process mode requires this
# layer enabled, an explicit model and daily invocation cap (not dollar cap).
#
# Leaving BOTH the API key and the base URL empty is a fully supported
# configuration, and the right one when policy forbids this layer sending chat
# content anywhere: it is disabled, makes zero upstream requests, and neither
# feature above sends anything. Nothing breaks — chats keep the title the
# content anywhere: it is disabled, makes zero upstream requests, and no
# consumer sends anything. With Lark knowledge processing disabled, nothing breaks — chats keep the title the
# client derives from the first message (30 characters, no model involved) and
# the follow-up question buttons simply never appear. server/pkg/llm asserts
# both the zero-request behaviour and the two-consumer list above, so neither
# both the zero-request behaviour and the consumer list above, so neither
# can drift from this comment silently.
# - API key for the upstream (OpenAI or any OpenAI-compatible gateway).
MULTICA_LLM_API_KEY=
Expand Down Expand Up @@ -636,6 +642,13 @@ MULTICA_LARK_CALLBACK_BASE_URL=
# environment handling.
MULTICA_LARK_WS_PROXY_URL=

# Optional single-group durable knowledge policy. Blank disables it.
# JSON keys: mode (capture|process), installation_id, app_id, workspace_id,
# agent_id, project_id, owner_id, chat_id, start_time (RFC3339), model,
# daily_model_calls (1..100). No credentials belong in this policy.
# See docs/runbooks/lark-group-knowledge.md for boundaries and cutover.
MULTICA_LARK_KNOWLEDGE_POLICY=

# DingTalk bot integration (Settings → Integrations "Bind to DingTalk")
# Off until MULTICA_DINGTALK_SECRET_KEY is set — a base64-encoded 32-byte key
# that encrypts each Bot's AppSecret at rest. Leave empty to disable.
Expand Down
2 changes: 1 addition & 1 deletion Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ RUN cd server && CGO_ENABLED=0 go build -ldflags "-s -w" -o bin/backfill_codex_u
# --- Runtime stage ---
FROM alpine:3.21

RUN apk add --no-cache ca-certificates tzdata
RUN apk add --no-cache ca-certificates tzdata poppler-utils tesseract-ocr tesseract-ocr-data-eng tesseract-ocr-data-vie

WORKDIR /app

Expand Down
6 changes: 4 additions & 2 deletions apps/docs/content/docs/environment-variables.fr.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -209,12 +209,14 @@ Ce groupe configure la génération d'assistance côté serveur, comme les titre

`MULTICA_LLM_MAX_RETRIES` est la seule source de la politique de nouvelles tentatives. Laissez-la non définie pour la valeur par défaut de 2, définissez `0` pour envoyer exactement une requête par appel, ou une valeur de 1 à 5 pour plafonner les nouvelles tentatives à ce nombre. C'est un plafond, pas un quota : seuls les échecs réessayables le consomment, et un succès ou l'échéance propre de l'appelant peut mettre fin à l'appel plus tôt. Toute autre valeur — négative, non numérique ou supérieure à 5 — fait échouer le démarrage au lieu d'être corrigée silencieusement. Ce plafond est un budget de latence : le délai entre deux tentatives commence à 0,5 s et double jusqu'à un maximum de 8 s, si bien qu'un budget plus élevé dépasserait les échéances des appelants et transformerait un échec réessayable en expiration de délai. Les nouvelles tentatives couvrent les échecs de connexion et les codes HTTP 408, 409, 429 et 5xx ; toutes les autres réponses 4xx sont renvoyées telles quelles. Au démarrage, le serveur journalise la politique effective sous la forme `llm retry policy`, sans aucun identifiant dans la ligne.

Deux fonctionnalités utilisent cette couche, et chacune envoie du contenu de discussion au point de terminaison que vous configurez :
Trois fonctionnalités utilisent cette couche, et chacune envoie du contenu de discussion au point de terminaison que vous configurez :

- **Titre automatique des discussions** — le premier message de l'utilisateur dans une nouvelle session de discussion, envoyé tel quel. Les pièces jointes ne sont jamais incluses.
- **Questions de suivi** (les boutons de suggestion sous la réponse d'un agent) — la fin de la conversation : jusqu'à 6 messages, la réponse à laquelle les questions font suite étant tronquée à 3 000 caractères et chaque message plus ancien à 800.

Lorsque la clé d'API et l'URL de base sont toutes deux vides, cette couche est désactivée et n'effectue aucune requête en amont — aucune des deux fonctionnalités ci-dessus n'envoie quoi que ce soit. C'est la configuration prise en charge lorsque votre politique n'autorise pas cette couche à envoyer du contenu de discussion hors du déploiement : les sessions de discussion conservent le titre que le client dérive du premier message, les boutons de questions de suivi n'apparaissent pas, et tout le reste fonctionne normalement.
- **Connaissances de groupe Lark, sur activation explicite** — envoie un nom de fichier et jusqu’à 256 Kio de texte extrait de PDF/DOCX pour classification et évaluation. Les réponses aux questions utilisent une question de 8 Kio maximum et jusqu’à six sources du groupe de 5000 caractères chacune, sans évaluations privées. La capture et la réconciliation ordinaires ne font aucun appel au modèle. `MULTICA_LARK_KNOWLEDGE_POLICY` est vide par défaut (désactivé) ; le mode `process` exige un modèle explicite et un plafond quotidien d’appels. Les nouvelles tentatives du SDK ajoutent des requêtes ; ce plafond ne limite pas les dépenses. Le traitement refuse de démarrer si cette couche est désactivée ; le mode `capture` reste disponible.

Lorsque la clé d'API et l'URL de base sont toutes deux vides, cette couche est désactivée et n'effectue aucune requête en amont — aucune des fonctionnalités ci-dessus n'envoie quoi que ce soit. C'est la configuration prise en charge lorsque votre politique n'autorise pas cette couche à envoyer du contenu de discussion hors du déploiement : les sessions de discussion conservent le titre que le client dérive du premier message, les boutons de questions de suivi n'apparaissent pas, et tout le reste fonctionne normalement.

<Callout type="info">
Ceci ne concerne que la couche d'assistance. L'exécution d'un agent suit un chemin de données distinct : lorsqu'un agent répond dans une discussion, votre daemon exécute l'outil de codage IA de cet agent avec les propres identifiants de l'outil, et ne lui transmet pas les paramètres `MULTICA_LLM_*` ci-dessus. (Les variables de connexion à Multica propres à l'exécution, dont l'agent a besoin, sont injectées séparément par le daemon.) Vider les variables ci-dessus n'a aucun effet sur ce chemin — contrôlez-le via la configuration du runtime de l'agent.
Expand Down
6 changes: 4 additions & 2 deletions apps/docs/content/docs/environment-variables.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -212,12 +212,14 @@ This group configures server-side assist generation, such as conversation titles

`MULTICA_LLM_DISABLE_THINKING` is for gateways whose models spend a reasoning ("thinking") pass that dominates the latency budget of the assist calls. When set to `true` (or `1`), the server adds `chat_template_kwargs: {"enable_thinking": false}` to every request body; gateways that forward the field — vLLM/sglang-style deployments and some GLM/Qwen-style LiteLLM routes — then skip the thinking pass. Standard OpenAI endpoints reject unknown body fields, so only enable this when your upstream accepts it. GPT-5.6-family models already get `reasoning_effort: none` on the follow-up questions request. Chat auto-titling sends no reasoning field; this switch affects it only when the configured upstream accepts `chat_template_kwargs`. Accepted values are `true`/`false` and `1`/`0` (case-insensitive); anything else fails startup. The startup log reports the effective state alongside the retry policy.

Two features use this layer, and each one sends chat content to the endpoint you configure:
Three features use this layer, and each one sends chat content to the endpoint you configure:

- **Chat auto-titling** — the first user message of a new chat session, sent verbatim. Attachments are never included.
- **Follow-up questions** (the suggestion buttons under an agent reply) — the tail of the conversation: up to 6 messages, with the reply being answered capped at 3000 characters and each older message at 800.

When both the API key and the base URL are empty, this layer is disabled and makes no upstream request at all — neither feature above sends anything. That is the supported configuration when your policy does not allow this layer to send chat content off the deployment: chat sessions keep the title the client derives from the first message, the follow-up question buttons do not appear, and everything else is unaffected.
- **Opt-in Lark group knowledge** — sends a filename and up to 256 KiB of extracted PDF/DOCX text for classification and assessment. Group Q&A sends a question up to 8 KiB and up to six group sources capped at 5000 runes each; private assessments are excluded. Ordinary capture and reconciliation make no model calls. `MULTICA_LARK_KNOWLEDGE_POLICY` defaults to empty (disabled); its explicit model and daily invocation cap apply only in `process` mode. SDK retries can add requests; this is not a monetary cap. Processing refuses startup when the assist layer is disabled; `capture` mode can still store sources.

When both the API key and the base URL are empty, this layer is disabled and makes no upstream request at all — none of these features sends anything. That is the supported configuration when your policy does not allow this layer to send chat content off the deployment: chat sessions keep the title the client derives from the first message, the follow-up question buttons do not appear, and everything else is unaffected.

<Callout type="info">
This covers the assist layer only. Running an agent is a separate data path: when an agent answers a chat, your daemon executes that agent's AI coding tool using the tool's own credentials, and does not forward the `MULTICA_LLM_*` settings above to it. (The run-scoped Multica connection variables the agent needs are injected by the daemon separately.) Emptying the variables above does not affect that path — govern it through the agent's runtime configuration.
Expand Down
6 changes: 4 additions & 2 deletions apps/docs/content/docs/environment-variables.zh.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -212,12 +212,14 @@ MULTICA_PUBLIC_URL=https://api.multica.example.com

`MULTICA_LLM_DISABLE_THINKING` 面向"模型会先跑一段推理(thinking)并吃掉辅助调用延迟预算"的网关。设为 `true`(或 `1`)时,服务端会在每个请求体中附带 `chat_template_kwargs: {"enable_thinking": false}`;支持透传该字段的网关(vLLM/sglang 类部署、部分 GLM/Qwen 系的 LiteLLM 路由)会因此跳过思考阶段。标准 OpenAI 端点会拒绝未知的请求体字段,所以仅在上游接受时启用。GPT-5.6 系模型在后续提问建议请求中已由服务端附带 `reasoning_effort: none`。会话自动命名不发送推理字段;只有当所配置的上游接受 `chat_template_kwargs` 时,此项才会影响它。可接受的取值为 `true`/`false` 和 `1`/`0`(不区分大小写);其它取值会让服务启动失败。启动日志会在重试策略旁打印生效状态。

有两个功能会使用这一层,它们都会把聊天内容发送到你配置的 endpoint:
有三个功能会使用这一层,它们都会把聊天内容发送到你配置的 endpoint:

- **对话标题自动生成** —— 发送新对话中用户的第一条消息,原文发送。附件不会包含在内。
- **后续提问建议**(智能体回复下方的按钮)—— 发送对话末尾的内容:最多 6 条消息,其中被追问的那条回复上限 3000 字符,更早的每条上限 800 字符。

API key 与 base URL 都为空时,这一层关闭,不会发出任何上游请求——上面两个功能都不再发送内容。如果你的政策不允许这一层把聊天内容发到部署之外,这就是受支持的配置方式:对话继续使用客户端根据第一条消息生成的标题,后续提问按钮不再出现,其余功能不受影响。
- **显式启用的 Lark 群知识功能** —— 分类与评估发送文件名和最多 256 KiB 的 PDF/DOCX 提取文本。群问答发送最多 8 KiB 的问题及最多 6 个群内来源,每个来源最多 5000 个字符,不包含私有评估。普通消息采集与数据核对不调用模型。`MULTICA_LARK_KNOWLEDGE_POLICY` 默认为空(关闭);`process` 模式要求明确指定模型和每日调用上限。SDK 重试可能增加请求次数,该上限不是费用上限。辅助层关闭时,处理模式拒绝启动;`capture` 模式仍可采集来源。

API key 与 base URL 都为空时,这一层关闭,不会发出任何上游请求——上述功能都不再发送内容。如果你的政策不允许这一层把聊天内容发到部署之外,这就是受支持的配置方式:对话继续使用客户端根据第一条消息生成的标题,后续提问按钮不再出现,其余功能不受影响。

<Callout type="info">
这只覆盖辅助生成这一层。智能体执行是另一条数据路径:智能体回复对话时,由你的守护进程调用该智能体的 AI 编程工具,用的是那个工具自己的凭据,守护进程不会把上面的 `MULTICA_LLM_*` 配置传给它。(智能体自身需要的运行级 Multica 连接变量由守护进程单独注入。)把上面的变量留空不影响这条路径——它由智能体的运行时配置决定。
Expand Down
Loading
Loading