Skip to content

fix(tokenizer): accept the OpenAI developer role (map to system for templates without it) - #391

Draft
gdevenyi wants to merge 1 commit into
FlashML-org:mainfrom
gdevenyi:fix/developer-role
Draft

fix(tokenizer): accept the OpenAI developer role (map to system for templates without it)#391
gdevenyi wants to merge 1 commit into
FlashML-org:mainfrom
gdevenyi:fix/developer-role

Conversation

@gdevenyi

@gdevenyi gdevenyi commented Sep 5, 2026

Copy link
Copy Markdown

What this fixes

A request whose system prompt uses the OpenAI developer role (the current spelling of system; pi sends it by default for openai-completions providers, compat.supportsDeveloperRole: true) fails with 400 could not encode request: Unexpected message role. on every model whose chat template does not spell the role, which is most of them (Qwen3.8-Flash-Next, Qwen3.x, GLM, ...). The template raises, the server relays it.

TokenizeManager._render now maps developer to system before apply_chat_template, unless the template mentions developer itself (gpt-oss renders the two roles differently and keeps its input). The mapping copies the message; the request is not mutated.

Testing

  • tests/tokenizer/test_tokenize.py::test_developer_role_maps_to_system_unless_the_template_knows_it.
  • Live on a Qwen3.8-Flash-Next TP=2 server: {"role": "developer", ...} went from the 400 above to the same prompt_tokens as the system spelling.

🤖 Generated with Claude Code

https://claude.ai/code/session_0173pf9k9fSVtwbm3f898HDt

…emplates without it)

Clients such as pi send the system prompt with role "developer" by default
(OpenAI's current spelling of "system"). Most chat templates do not know the
role and raise "Unexpected message role", which the server returns as a 400.
Map developer to system before rendering unless the template spells the role
itself (gpt-oss does).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0173pf9k9fSVtwbm3f898HDt
@gdevenyi

Copy link
Copy Markdown
Author

Re-tested against current main (fb7f732, i.e. after the #418 / #427 / #426 quantization refactor) on 2 x RTX 6000 Ada, TP=2 box.

Method. This PR's head merged onto main, then the full pytest tests suite. The run is CUDA-hidden (CUDA_VISIBLE_DEVICES="") on purpose: this box is serving a model on both GPUs, and with them visible 65-95 GPU tests fail on main itself with AcceleratorError: out of memory, with the count swinging ~10 between identical runs. Hiding CUDA makes the result deterministic, so a failure-set difference against main means something. Baseline: main = 1205 passed, 350 skipped, 0 failed.

Result: 1206 passed, 350 skipped, no new failures.

The +1 over main are this PR's own tests, and they ran (not skipped).

🤖 Generated with Claude Code

https://claude.ai/code/session_0173pf9k9fSVtwbm3f898HDt

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant