docs(streaming): clarify warm_willneed call paths - #120
Open
agourakis82 wants to merge 2 commits into
Open
agourakis82 wants to merge 2 commits into
agourakis82 wants to merge 2 commits into
Conversation
While investigating Edge0-AI#110 (edge0-8b first-token latency), traced why toggling warm_willneed had no visible effect on that tier: it's only read by StreamingSwitchGLU.prefetch() and stage_experts(), and stage_experts() early-returns immediately when staged is False. prod_k8() (edge0-8b's production profile) sets staged=False and history_prefetch=False, so neither call site that reads warm_willneed ever executes -- the flag is dead code for that tier's default config, not merely ineffective in this instance. Documents this in both the class docstring (options.py) and the streaming.md options table, so the next person investigating slow first-token latency on edge0-8b doesn't spend time on the same dead end. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Preserve the PR history by merging current upstream main. Adapt the documentation to the python/ layout and describe explicit prefetch and current staged consumer behavior without changing runtime options. Assisted-by: Codex
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The documentation-only changes have one minor clarification remaining and no blocking issues.
Review effort: Balanced
Findings: 1
What changed in this PR
Clarifies when streaming readahead applies, helping users investigating first-token latency in #110.
Changes:
- Documents
warm_willneedcall paths and loading limitations. - Updates the
prod_k8()description to reflect staged prerouter consumer layers.
| File | Description |
|---|---|
| python/src/edge0/streaming/options.py | Expands the warm_willneed field documentation. |
| docs/streaming.md | Clarifies readahead behavior and the production preset. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+95
to
+97
| #: or affect plain on-demand / whole-layer loading. Explicit prefetch | ||
| #: calls, including ``prefetch_from_prefill()``, can still consult it | ||
| #: when automatic history prefetch and staging are disabled. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Updates the streaming documentation related to #110 for the current
python/layout and production profile.
warm_willneedis consulted byprefetch()for missing experts and bystage_experts()on staged layers. It does not enable either path or warmplain on-demand or whole-layer loads by itself. Explicit prefetch calls,
including
prefetch_from_prefill(), can still consult the flag when automatichistory prefetch and staging are disabled.
The current
prod_k8()enables staging for prerouter consumer layers. Thissupersedes the original PR description's blanket claim that the default
profile makes the flag inert.
The update merges current upstream main without rewriting the PR history.
Relative to main, only
docs/streaming.mdand documentation comments inpython/src/edge0/streaming/options.pychange. Runtime settings and code areunchanged.
Validation
current
prod_k8()profile.git diff --checkagainst current main passes.documentation-only update. Historical test results do not establish
validation of the new head.
Assisted-by: Codex