Conversation
Tokenization already runs across threads, but pooling ran one text at a time afterwards, so it became the bottleneck on a machine with cores to spare. Pooling now happens inside the same parallel iterator, using the tokenizers crate's own helper, so TOKENIZERS_PARALLELISM keeps controlling both and wasm still builds without threads. pool_ids also sums over a contiguous row slice instead of an ndarray row iterator, which lets the loop vectorize. On 50k Wikipedia paragraphs, samples per second before -> after: potion-base-8M 1 thread 18,845 -> 23,193 all cores 64,306 -> 98,176 potion-code-16M-v2 1 thread 9,712 -> 18,640 all cores 15,830 -> 124,592 potion-multilingual-128M 1 thread 9,759 -> 9,840 all cores 49,704 -> 96,951 Embeddings are unchanged: identical bytes for potion-base-8M, science-32M, code-16M-v2 and multilingual-128M and the float16, int8 and vocab-quantized fixtures, over Wikipedia, ag_news, long and adversarial texts.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tokenization already runs across threads, but pooling ran one text at a time after it. This pools in the same parallel iterator, and sums over a contiguous row slice so the loop vectorizes.
50k Wikipedia paragraphs, samples/s:
The single-thread gain comes from the slice change only.