Skip to content

fix(server): apply default_instruction to qwen2_flash queries - #621

Merged
svonava merged 1 commit into
mainfrom
fix/qwen2-flash-default-instruction
Oct 8, 2026
Merged

svonava merged 1 commit into
mainfrom
fix/qwen2-flash-default-instruction

Conversation

@svonava

@svonava svonava commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Qwen2FlashAdapter.encode never reads the profile's default_instruction runtime option. It only uses the request's instruction (params.instruction, or options.instruction on the HTTP path). A query sent with is_query=true and no explicit instruction is therefore formatted with an empty instruction:

Instruct: 
Query:what is sie?

instead of

Instruct: Given a web search query, retrieve relevant passages that answer the query
Query:what is sie?

User impact

Every query to a model served by the qwen2_flash adapter on CUDA without an explicit instruction is embedded with an empty instruction. That is the default way most clients call these models. Instruction-tuned embedders expect the instruction on the query side, so these query embeddings are degraded. Documents are not affected. The shipped configs that use this adapter, all with a query_template and a default_instruction, are:

  • Qwen/Qwen3-Embedding-8B
  • Qwen/Qwen3-Embedding-0.6B
  • NovaSearch/stella_en_1.5B_v5
  • tencent/R3-embedding-0.6b

The behaviour also differed across adapters for the same models. The SGLang embedding adapter (e.g. Qwen/Qwen3-Embedding-4B) applies default_instruction, and so does qwen2_flash's own CPU fallback, SentenceTransformerDenseAdapter. The same model could therefore embed queries differently on CPU and on CUDA.

Fix

When is_query is true and the request gives no instruction (instruction is None), use options["default_instruction"]. The options have already been merged from the profile's adapter_options.runtime and the request, so a request-level default_instruction still overrides the profile. The precedence matches XLMRobertaFlashAdapter and SentenceTransformerDenseAdapter:

  • an explicit instruction wins;
  • an explicit empty-string instruction is kept, because the API types allow "" and the sibling adapters keep it;
  • documents never get the default instruction.

Note that the SGLang and PyTorch embedding adapters use instruction or default_instruction, so for them "" falls back to the default. This PR mirrors the flash and SentenceTransformer semantics instead, because SentenceTransformer is this adapter's CPU fallback and the two should format text the same way.

Tests

The new file packages/sie_server/tests/adapters/test_qwen2_flash.py runs on CPU and loads no models. It stubs the flash transformer stack, captures the exact text handed to the tokenizer, and builds options from the shipped model YAMLs through merge_runtime_options, the same merge the HTTP and queue paths use.

  • The Qwen3-Embedding-8B query uses the profile default instruction (exact string).
  • Every shipped qwen2_flash model config applies its own default_instruction (parametrized over the YAMLs).
  • An explicit instruction wins over the default.
  • An explicit "" instruction is kept.
  • A request-level default_instruction overrides the profile.
  • Documents get no instruction.
  • A query with no instruction and no default keeps the previous empty slot.
  • Metered input_token_counts match the formatted text, including the instruction tokens.

Against main before the fix, 6 of these tests fail with 'Instruct: \nQuery:…'. All pass with the fix.

Validation

  • mise run lint: passed
  • mise run typecheck: passed
  • mise run test -- packages/sie_server/tests/adapters/: 4542 passed, 57 skipped
  • mise run test: 12627 passed, 564 skipped, 5 failed. The 5 failures are macOS-local PermissionErrors from renaming read-only directories in config/test_serving_artifacts.py and core/test_model_loader_cache.py. They fail the same way with this change reverted, so they are unrelated to it.

Follow-up (not in this PR)

RoPEFlashAdapter has the same gap. It is used by NovaSearch/stella_en_400M_v5, which also configures a default_instruction. This PR keeps to qwen2_flash, and the same change can follow separately.

Compatibility

No public API, wire or config changes. This is a runtime behaviour fix for queries to the models above. Query embeddings for these models change after the fix and now match the published recipe. Document embeddings are unchanged, so existing indexes do not need to be re-embedded.

Summary by CodeRabbit

  • Bug Fixes
    • Query formatting now uses the configured default instruction when no instruction is provided. Explicit instructions—including an empty instruction—are preserved, and document formatting remains unchanged.
  • Tests
    • Added coverage for query and document formatting, instruction precedence, and token counts.

Qwen2FlashAdapter.encode never read the profile's `default_instruction`
runtime option. A query without an explicit instruction was formatted as
"Instruct: \nQuery:{text}" for Qwen3-Embedding-0.6B/8B, stella_en_1.5B_v5
and R3-embedding-0.6b on the CUDA flash path, while the SentenceTransformer
CPU fallback and the SGLang adapter applied the instruction.

Fall back to `default_instruction` for queries when the request gives no
instruction, matching XLMRobertaFlashAdapter and the SentenceTransformer
fallback: an explicit instruction (including "") wins, and documents get
no instruction.
@svonava
svonava requested a review from a team as a code owner October 8, 2026 04:39
@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Team
  • Run ID: eae6b7d9-35e0-413a-8daf-94b54717408b
📥 Commits

Reviewing files that changed from the base of the PR and between 6163120 and 7f70319.

📒 Files selected for processing (2)
  • packages/sie_server/src/sie_server/adapters/qwen2_flash/__init__.py
  • packages/sie_server/tests/adapters/test_qwen2_flash.py

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 7 remain after this review.


📝 Walkthrough

Walkthrough

Qwen2FlashAdapter.encode now uses default_instruction for queries when instruction is None. New CPU-only tests check default and explicit instructions, document formatting, and token counts.

Changes

Query instruction handling

Layer / File(s) Summary
Query instruction fallback and tests
packages/sie_server/src/sie_server/adapters/qwen2_flash/__init__.py, packages/sie_server/tests/adapters/test_qwen2_flash.py
When a query has no instruction, encode passes default_instruction to extract_texts. Explicit instructions, including an empty string, remain unchanged, and non-query calls do not use the fallback. Tests check profile and request defaults, query templates, document formatting, and token counts.

Suggested reviewers: dragosboca

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 7f703

No actionable issue is identified that would prevent merging after normal checks.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 6.67% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 15 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: applying default_instruction to qwen2_flash queries.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@svonava
svonava merged commit db4c0a8 into main Oct 8, 2026
21 checks passed
@svonava
svonava deleted the fix/qwen2-flash-default-instruction branch October 8, 2026 04:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant