Skip to content

fix: use tokenizer-specific pre-tokenization rules - #1975

Merged
leejet merged 1 commit into
masterfrom
fix/tokenizer-specific-pretokenization
Sep 14, 2026
Merged

leejet merged 1 commit into
masterfrom
fix/tokenizer-specific-pretokenization

Conversation

@leejet

@leejet leejet commented Sep 14, 2026

Copy link
Copy Markdown
Owner

Summary

  • Replace the handwritten splitter with precompiled Unicode regexes for CLIP, Qwen/SenseNova, and Mistral.
  • Remove incorrect regex boundaries from Gemma while preserving its normalization flow.

Related Issue / Discussion

N/A

Additional Information

N/A

Checklist

@leejet
leejet merged commit 59c23bc into master Sep 14, 2026
9 checks passed
@leejet
leejet deleted the fix/tokenizer-specific-pretokenization branch September 14, 2026 18:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant