Skip to content

chore(main): release 0.10.0 - #660

Merged
svonava merged 2 commits into
mainfrom
release-please--branches--main
Oct 9, 2026
Merged

svonava merged 2 commits into
mainfrom
release-please--branches--main

Conversation

@sie-release-automation

Copy link
Copy Markdown
Contributor

🤖 I have created a release beep boop

0.10.0 (2026-10-09)

⚠ BREAKING CHANGES

  • server: a LoRA whose PEFT config sets bias to "all" or "lora_only" no longer loads. A profile that declares one still loads its model without that LoRA and logs the refusal. A request that names one gets the same retryable LORA_LOADING response as for any other LoRA that fails to load in the background, and the server log carries the refusal. Retrain or export the LoRA with bias="none".

Features

  • config: accept hybrid encode and score routing in a cluster (#576) (96de71c), closes #415
  • config: refuse a remote profile that names an undefined upstream (#500) (4977097)
  • examples: add bounded agent stage latency trials (2326f19)
  • examples: add custom-entity-types, caller-named entity types against LLMs, Comprehend and spaCy (#468) (7860d1f)
  • examples: add reproducible agent stage latency evaluation (aa20315)
  • examples: add reproducible photo transcription study (#630) (a826a95)
  • examples: add structured-output-accuracy, strict structured outputs scored on the Structured Output Benchmark (#474) (96b7f9b)
  • examples: compare matched short-verdict guardrails configurations (#479) (3ebda81)
  • examples: document-to-markdown-olmocr prints the score without the math files (#459) (37e1380)
  • examples: evaluate ordinary named entity tags (#629) (4ecec8a)
  • examples: rebuild guardrails as a harmful-prompt screening study (#456) (01e606e)
  • examples: rebuild image-search on a held-out Amazon Berkeley Objects product catalogue (#470) (8041b15)
  • examples: rebuild redact on the pre-registered PII coverage run (#452) (f99850a)
  • examples: rebuild visual-document-search on the ViDoRe v3 head-to-head (#472) (832f561)
  • examples: replay document-field extraction evidence (#511) (971f2d4)
  • examples: reproduce caller-rule rerank study (#485) (c5df8f9)
  • examples: reproduce GLM-OCR receipt and handwriting results (#486) (ea15097)
  • examples: reproduce RF100-VL named-object detection (#487) (56d932e)
  • examples: reproduce the image retrieval pilot (#602) (ed79362)
  • gateway: bridge a lane its transport reports cold (#592) (b5d6253)
  • gateway: bridge buffered cluster generation refusals (#547) (6ec8dec)
  • gateway: bridge cluster extraction and audio refusals (#549) (b009e9f)
  • gateway: bridge encode and score under a numerical admission (#587) (f6e0820), closes #415
  • gateway: bridge governed generation to the remote route a policy names (#607) (e875a65)
  • gateway: carry generation GPU time (11612e4)
  • gateway: carry sealed generation GPU time (2a4a3a1)
  • gateway: carry the fallback reason on every bridged work item (#567) (8823a89)
  • gateway: carry the requested model on queued work (#504) (c98bfbb)
  • gateway: coordinate threshold demand with bounded leases (#554) (6e1424a)
  • gateway: disclose which side served on gateway responses (#502) (04e6159)
  • gateway: enforce remote-forbid with verified worker dispatch (#546) (19be847)
  • gateway: give a 503 serving refusal its own admission outcome (#493) (39a1aa9)
  • gateway: let a model access policy admit remote routes (#591) (7e7050f)
  • gateway: opt in to pre-acceptance remote spill (#550) (4f3a62e)
  • gateway: preserve validated remote routing policies (#542) (f05efe5)
  • gateway: publish load-only model readiness work (#541) (bd945d6)
  • gateway: restore cluster stream refusals before output (#548) (0ebc55c)
  • helm: add a remote worker pool that alone holds upstream credentials (#496) (4747581)
  • helm: give sie-config the upstream names the remote lanes define (#508) (72b48d6)
  • helm: limit the remote lanes' network access with a NetworkPolicy (#498) (f81fab3)
  • helm: mount equivalence evidence into remote lanes (#577) (738cf38)
  • models: add a 1120 soft-token image profile for Gemma 4 31B on H100 (#633) (d97dcf2)
  • models: add a compact 768-visual-token profile to tomoro-colqwen3-embed-4b (#469) (92a2a85)
  • models: add an 8,192-token output Gemma 4 31B hi-res profile (#635) (b9231ce)
  • models: add Qdrant/splade-ecommerce-esci sparse model (#631) (842c3b0)
  • models: refresh Qwen3.5-122B candidate (44047bd)
  • models: refresh Qwen3.5-122B candidate (e72f50f)
  • queue: support load-only work and reply to backend fallback refusals (#525) (1aa45ce)
  • remote: admit exact fresh OpenAI equivalence evidence (#535) (c6f077f)
  • remote: admit fresh matching SIE profile identity (#540) (98b55ec)
  • remote: bind equivalence evidence to execution identity, not one process (#572) (21f7c0b), closes #415
  • remote: expose bounded numerical process diagnostics (#561) (9f966be)
  • remote: expose diagnostic numerical process snapshots (#557) (6113be1)
  • remote: measure equivalence through a cluster gateway (#579) (e964598)
  • remote: measure the local serving envelope in the equivalence probe (#588) (43157ba)
  • remote: re-verify the numerical admission before an admitted item runs (#586) (66f5635), closes #415
  • remote: report remote admissions in worker health (#578) (a85c04f)
  • route remote profiles to a dedicated remote worker bundle (#492) (ab5d4a9)
  • routing: enable shared threshold routing behind cluster opt-in (#555) (93d5660)
  • sdk: accept origin-confined configured HTTP clients (#539) (dea0ca1)
  • sdk: forbid remote serving and report which side served (#495) (17ce8b4)
  • sdk: pass a target language to recommend (#611) (47cf609)
  • server: add a routing block to the model config (#490) (c02785e)
  • server: add native GLiNER2.5 and Privacy Filter extraction (#594) (e01ad1e)
  • server: add native ZeRank 2 scoring (#596) (be8a688)
  • server: add optional cuDNN SDPA startup control (#650) (8770c72)
  • server: add Parakeet-TDT speech-to-text (nvidia/parakeet-tdt-0.6b-v3) (#608) (7e2a1d3)
  • server: add Qwen3Guard-Gen-4B and Qwen3Guard-Gen-0.6B guard models (#448) (471189c)
  • server: add SAM 3 open-vocabulary detection (#614) (7e7d0f5)
  • server: answer an upstream that cannot serve yet with a retryable 503 (#494) (ab0217b)
  • server: bridge cold single-node generation before output (#533) (bab928f)
  • server: cap the calls to each upstream and stop calling one that keeps failing (#506) (1a259cc)
  • server: collect process-bound numerical fleet evidence (#556) (7d0f487)
  • server: count hybrid upstream generation with the model's tokenizer (#620) (2f9c107)
  • server: declare openai upstream endpoints and operator request fields (#499) (2292823)
  • server: default gliner-biomed-large-v1.0 to threshold 0.8, chosen on dev data (#461) (6c13f28)
  • server: disclose which side served and report routing in /v1/models (#445) (946b6fb)
  • server: expose conservative immutable profile identity (#530) (1c62676)
  • server: identify flash BGE-M3 profiles with exact library builds (#575) (de82bad), closes #415
  • server: identify the BERT, cross-encoder and Nomic flash adapters (#581) (cd3cb91)
  • server: load a remote profile without waiting for a local load (#491) (f53883b)
  • server: measure remote equivalence against local noise (#531) (e5cc5c6)
  • server: preserve onboarded templates for direct remote chat (#537) (5581a23)
  • server: run the remote bundle on transformers 5 so hybrid counting loads transformers-5 tokenizers (#625) (577f3cf)
  • server: serve a cold model through its remote profile on a single node (#503) (3742178)
  • server: serve a remote-backed embedding model through an SIE upstream (#437) (7a56b3c)
  • server: serve embeddings and rerank through an openai upstream (#505) (23cc6df)
  • server: serve native SIE chat through remote profiles (#523) (d0869f4)
  • server: serve OpenAI upstream chat and raw generation (#528) (21e2ab3)
  • server: serve remote chat through queued generation (#529) (a8bbfa4)
  • server: serve score, extract and every encode output through an SIE upstream (#501) (6344850)
  • server: SMVE sparse multi-vector encoding, with smve profiles for TopK models (#518) (9eeff51)
  • server: stream generation through a remote SIE upstream (#514) (d3fa3c1)
  • server: support native GLiNER2 entity descriptions (#605) (461f8dc)
  • server: support native LightOnOCR-3 extraction (#632) (5e60012)
  • server: validate shared remote chat responses (#513) (42897e9)
  • telemetry: observe cluster remote fallback persistence (#551) (8e2032b)
  • worker: add versioned execution authority IPC methods (#544) (5f832a4)
  • worker: fence verified dispatch with versioned queues (#545) (2ba3d6e)

Bug Fixes

  • answer /v1/rerank without usage when the score has none (#569) (3459946)
  • batcher: serve long-form audio in its own lane, alternating with clips (#601) (45b6ffe), closes #585
  • examples: align graph evidence with recorded calls (#510) (ef90dd6)
  • examples: image-search scores SigLIP so400m-384, the model superlinked.com/image-search sells (#476) (395366c)
  • examples: keep journal replay outside latency timing (10564f8)
  • examples: redact credentials in reply keys (b57c033)
  • examples: reject malformed latency reply members (e068b1d)
  • examples: rerun and grade document-field extraction with the measured recipe (#622) (557f856)
  • examples: retain owned replies in timed event capture (ab247dc)
  • examples: scope OCR quality claims and support self-hosting (#483) (6cccc04)
  • examples: sync latency outputs and retain deadline reasons (c4a9274)
  • examples: validate latency trial execution evidence (c4e94a3)
  • gateway: answer a queued generation INPUT_TOO_LONG with 400 (#497) (0cb086d)
  • gateway: derive a failed bridge's fallback error as the single server does (#580) (70b3181)
  • gateway: dispatch a forbid request normally when its model has no remote route (#564) (2600ded)
  • gateway: keep a hidden remote profile out of the model listing (#598) (7a43250)
  • gateway: keep json_schema property order through to the worker (#455) (571f45b)
  • gateway: keep requests for undeclared outputs off numerical bridges (#597) (b6f1702)
  • helm: require pytorch for remote lanes (2e1aef2)
  • helm: require pytorch for remote lanes (1545d61)
  • match identifier patterns against the whole value (#568) (6e5ce3d)
  • models: pin GLiNER v2.5 tokenizer dependencies (#634) (a52cdbb)
  • models: raise the Hy-MT2-1.8B output cap to 1024 tokens (#610) (e831d4d)
  • models: validate Qwen runtime defaults (8fae464)
  • parakeet: decode long audio in pieces cut at pauses (#624) (4d0e3c6)
  • queue: answer a remote profile's upstream refusal at once with its wait (#570) (22aa825)
  • remote: bind profile identity to observed hardware and BLAS (#536) (89640b0)
  • remote: bound configuration waits and settle grammar refusals (#558) (b6bbf38)
  • remote: carry admitted numerical work on its own subject (#589) (6d4d74b)
  • remote: keep instructions and unadmitted plans off numerical bridges (#600) (ab016b1)
  • remote: name both identities when SIE identity admission refuses a mismatch (#563) (52c6c6d)
  • sdk: preserve HTTP retry hints in terminal server errors (#560) (7264193)
  • sdk: serialize encode query roles in nested options (#617) (812962d)
  • server: apply default_instruction in the remaining flash embedders (#623) (07983db)
  • server: apply default_instruction to qwen2_flash queries (#621) (db4c0a8)
  • server: classify all of a long text with GLiNER2, not only its first 512 words (#449) (82cbf2e)
  • server: convert queued tool-call arguments by the declared schema (#609) (5f00470)
  • server: default grammar-constrained generation to greedy sampling (#453) (b1c8887)
  • server: exit on SIGTERM after starting the drain (#637) (ecb957d)
  • server: give background identity refreshes their own budget (#618) (c998453)
  • server: honor explicitly selected default profiles (#534) (87a0759)
  • server: ignore optional triton type imports (10cac53)
  • server: ignore optional triton type imports (f9ab207)
  • server: let GLiNER checkpoints on byte-level BPE encoders read split words (#460) (d214552)
  • server: meter ColQwen3 text from processed inputs (#509) (d787725)
  • server: pin GLiFormer decoding probe to float32 (#562) (df09d4b)
  • server: read long documents whole in GLiNER extract (#451) (2f97d02)
  • server: refuse LoRAs that train biases (#583) (dcfdf51)
  • server: refuse remote generation before streaming headers (#532) (9719ae1)
  • server: reject a config delta per model instead of stalling its bundle (#489) (81cc427)
  • server: reject incomplete GLiNER2 extraction items (#552) (689e4d8)
  • server: report GLiNER2 caller errors as invalid input (#526) (96066a3)
  • server: report NLI zero-shot extract usage and enforce the extract label limit on the queue path (#447) (2cf3d89)
  • server: score a lone GLiClass label on its own in single-label mode (#616) (45c3fee)
  • server: score NLI zero-shot labels like the transformers pipeline (#615) (d134ab2)
  • server: score Qwen3 text rerankers in float32 so top candidates stop tying (#440) (8de63e3)
  • server: size NLI zero-shot extract batches by the rows each item runs (#465) (57a9856)
  • server: skip remote profiles a worker refuses in pool isolation (#627) (cb79fc1)
  • server: suppress private chat reasoning logprobs independently of thinking policy (#527) (d9d63dc)
  • sglang: bound JSON number digits in XGrammar json_schema grammars (#467) (7f3666d)
  • validate remote routing before configuration writes (#512) (cb95d4a)
  • whisper: transcribe long-form audio one recording per pipeline call (#582) (8bdaeba), closes #574
  • worker: enforce live configuration authority through execution (#543) (f6c530c)

Performance Improvements

  • colbert: tokenize a ModernColBERT batch in one call and pack it on the host (#646) (e15d014)
  • gliner: run 32 rows per forward pass instead of gliner's default 8 (#638) (e7622c3)
  • gliner: turn the adaptive batching controller off on the GLiNER profiles (#648) (cfaafe3)
  • models: add a 16-request CUDA-graph H100 profile for Qwen3.8-27B-FP8 (#454) (409fee6)
  • models: compile compact JSON grammars on every Qwen3.6-35B-A3B profile (#636) (d8f58cb)
  • models: compile compact, digit-bounded JSON grammars on the Gemma 4 31B hi-res profiles (#641) (a8eb93d)
  • models: enable decode CUDA graphs on the Qwen3.8-27B-FP8 bare route and overlap the batch lane (#640) (0f0d691)
  • models: read a single Qwen3.8-27B-FP8 page image at 3.2 megapixels and compile compact JSON grammars (#466) (329df87)
  • models: serve Qwen3.6-35B-A3B default on the CUDA 13 engine with CUDA graphs (#639) (2f7fcbd)
  • qwen3-embedding: fuse the 8B's elementwise layer work into exact Triton kernels (#645) (6949a8b)
  • qwen3-embedding: keep up to four Qwen3-Embedding-4B batches in flight to SGLang (#647) (99140a4)
  • sam3: skip the unused mask decoder, encode images one at a time and cache label encodings (#643) (06bc893)
  • server: run SGLang OCR batches concurrently instead of one at a time (#462) (dce7611)
  • server: serve OWLv2 with the fast image processor and shrink oversized photos (#473) (b01062d)
  • sparse: pool before the sparse activation and keep doc-v3 batches at full size (#642) (9075c1f)

This PR was generated with Release Please. See documentation.

@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Bot user detected.

To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 5ebd44cb-efd9-400c-aa25-9c9963a2730f

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@svonava
svonava merged commit addfe02 into main Oct 9, 2026
45 checks passed
@svonava
svonava deleted the release-please--branches--main branch October 9, 2026 20:18
@sie-release-automation

Copy link
Copy Markdown
Contributor Author

🤖 Created releases:

🌻

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

1 participant