[Klaud Cold] Update dsr1-fp4-b200-dynamo-sglang SGLang image to v0.5.19-cu130-runtime (blocked: recipe-pinned ai-dynamo 0.8.1 incompatible) / 将 dsr1-fp4-b200-dynamo-sglang 的 SGLang 镜像更新至 v0.5.19-cu130-runtime(受阻:配方固定的 ai-dynamo 0.8.1 不兼容) - #2907
Conversation
…19-cu130-runtime Update the dsr1-fp4-b200-dynamo-sglang master image from lmsysorg/sglang:v0.5.8.post1-cu130-runtime to lmsysorg/sglang:v0.5.19-cu130-runtime, pinned by manifest digest. Model, precision, topology, workloads, recipe references and evals are unchanged. Klaud Cold candidate 906fff6c74191d2d-fa39a9e32aba886a. 将 dsr1-fp4-b200-dynamo-sglang 主配置镜像从 lmsysorg/sglang:v0.5.8.post1-cu130-runtime 更新为 lmsysorg/sglang:v0.5.19-cu130-runtime(按 manifest 摘要固定)。 模型、精度、拓扑、负载、配方引用与评测均保持不变。 Klaud Cold 候选 906fff6c74191d2d-fa39a9e32aba886a。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
|
Status: terminated before any GPU dispatch; PR closed as draft, branch retained. Confirmed finding: this family's serving stack is installed at job start by srt-slurm pin Action / next step: no Links: image commit 8ab3bdf; baseline producer https://github.com/SemiAnalysisAI/InferenceX/actions/runs/21888944497/attempts/1. 状态: 在任何 GPU 派发之前终止;PR 以草稿状态关闭,分支保留。 已确认的发现: 该家族的服务栈由 srt-slurm 固定提交 操作 / 下一步: 未派发任何 链接: 镜像提交 8ab3bdfbd4078d03018188fe5470dcdea019a850;基线生产运行 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/21888944497/attempts/1。 |
Current status / 当前状态
Status: stopped before any GPU dispatch — confirmed incompatibility between the candidate image and this family's recipe-pinned serving stack. PR closed as draft; branch retained so this exact candidate is not reselected.
Next step (manual, out of Klaud Cold scope): raise the
dynamo.versionpin used by this family's upstream srt-slurm recipe, or move the family onto checked-in recipes with a compatibledynamo.hash(as the MTP sibling already does), then retry the image refresh.The one real change in this PR is the master image bump for
dsr1-fp4-b200-dynamo-sglanginconfigs/nvidia-master.yaml:lmsysorg/sglang:v0.5.8.post1-cu130-runtime→lmsysorg/sglang:v0.5.19-cu130-runtime@sha256:710bc11443a7b1807d69803386468101bcfced35f86bf8fe92a8209e05a2f052(manifest-list digest, same pinning style as the existingv0.5.14-cu130@sha256:…entries). Model, precision, topology, workloads, recipe references, router metadata and eval selection are unchanged; the generated matrix (5 points, 2/6/6/2/6 nodes, evals at conc 128/128/64/2048) differs from base only inimage. Noperf-changelog.yamlentry was added because targeted validation never passed. Green benchmarks would not prove global checks pass; none were run here.状态:在任何 GPU 派发之前停止 —— 已确认候选镜像与该家族配方固定的服务栈不兼容。PR 以草稿状态关闭;保留分支以避免再次选中同一候选。
下一步(需人工处理,超出 Klaud Cold 范围):提高该家族上游 srt-slurm 配方所固定的
dynamo.version,或将该家族迁移到带有兼容dynamo.hash的仓库内配方(MTP 兄弟家族已采用此方式),然后重试镜像刷新。本 PR 唯一的实际改动是
configs/nvidia-master.yaml中dsr1-fp4-b200-dynamo-sglang的主镜像:lmsysorg/sglang:v0.5.8.post1-cu130-runtime→lmsysorg/sglang:v0.5.19-cu130-runtime@sha256:710bc11443a7b1807d69803386468101bcfced35f86bf8fe92a8209e05a2f052(manifest-list 摘要,与现有v0.5.14-cu130@sha256:…条目相同的固定方式)。模型、精度、拓扑、负载、配方引用、router 元数据与评测选择均未改变;生成矩阵(5 个点,节点数 2/6/6/2/6,评测并发 128/128/64/2048)与基线仅image不同。由于定向验证未通过,未追加perf-changelog.yaml条目。基准通过不代表全局检查通过;此处未运行任何基准。Candidate evidence / 候选证据
906fff6c74191d2d-fa39a9e32aba886a; base SHAcc53d2341f5ba0374c76c592f95cf72c16bfa900; branchklaud/auto-906fff6c74191d2d-fa39a9e32aba886a; head8ab3bdfbd4078d03018188fe5470dcdea019a850./api/v1/latest-images): dsr1 / b200 / dynamo-sglang / fp4 / spec none / disagg / 8192→1024 / single_turn onlmsysorg/sglang:v0.5.8.post1-cu130-runtime, published 2026-02-11. Release feed (/api/v1/framework-releases): sglangv0.5.19. Review reason: release-string-mismatch.configs/nvidia-master.yaml:dsr1-fp4-b200-dynamo-sglang(runnercluster:b200-nscale,configs/runners.yamlmaps it only tob200-nscale-slurm_00..09). The MTP variant is the separate keydsr1-fp4-b200-dynamo-sglang-mtpand is untouched.lmsysorg/sglang:v0.5.19-cu130-runtimeexists on Docker Hub (pushed 2026-09-04T21:25Z; manifest digestsha256:710bc1…a2f052; linux/amd64 imagesha256:c31ecf…5fa476). Same CUDA 13.0 runtime line as the currentcu130-runtimepin; B200 (sm_100) is a supported Blackwell target of the cu130 build.check-capacity --cluster b200-nscaleexit 0 before edits and again before branch/PR creation (01:51Z and 01:59Z UTC, 2026-09-09).configs/nvidia-master.yamlor the B200 nscale launchers (Validate vLLM Router on GB200: DEP4, DEP8, and 1P/2D #2549, Validate official vLLM Router with srt-slurm #2543, test: validate upstream-native vLLM Router topologies #2731, perf(agentx): refresh B200 vLLM MTP recipe / 刷新 B200 vLLM MTP 配置 #2646, Add DeepSeek-V4-Pro FP4 B200 Dynamo TensorRT-LLM recipes / 添加 DeepSeek-V4-Pro FP4 B200 Dynamo TensorRT-LLM 配置 #2552, [WIP] Test DeepSeek-V4 GB300 Dynamo AgentX recipes / 测试 DeepSeek-V4 GB300 Dynamo AgentX 配置 #2302, Add dsv4f-fp8-b200-vllm: DeepSeek-V4-Flash-0731 FP8 B200 vLLM single-node recipe #2535, feat(config): add GLM-5.2 NVFP4 B200 Dynamo-SGLang AgentX recipes / 新增 GLM-5.2-NVFP4 B200 Dynamo-SGLang AgentX 配方 #2673, feat(glm52-agentx): Add B300 Dynamo+TRT-LLM AgentX recipes #2666, refactor: share AgentX matrix generation / 共用 AgentX 矩阵生成逻辑 #2903) modify this key, its image,recipes/b200-fp4/8k1k.yamlor the srt-slurm pin; noklaud/auto-906fff…branch existed.Baseline (published 2026-02-11) / 基线(发布于 2026-02-11)
GET /api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-02-11&exact=truefiltered to hardware=b200, framework=dynamo-sglang, model=dsr1, precision=fp4, spec_method=none, disagg=true, isl=8192, osl=1024, image=lmsysorg/sglang:v0.5.8.post1-cu130-runtime(11 rows);GET /api/v1/workflow-info?date=2026-02-11;GET /api/v1/evaluations?model=DeepSeek-R1-0528&date=2026-02-11&exact=true.46e0b5ff02e26f07738eab5105c3c1da6123e8a9). Curve snapshot id274is not a producer id. Frozen for all attempts; the old image was never dispatched.Initial attempt / 初次尝试
Image / commit:
lmsysorg/sglang:v0.5.19-cu130-runtime@sha256:710bc1…a2f052@8ab3bdfbd4078d03018188fe5470dcdea019a850.Run links: N/A — no
e2e-tests.ymlrun was dispatched (reason below). No GPU time was used.Benchmark / eval results: N/A — not run.
Per-point deltas vs baseline: N/A — not run.
Diagnosis (pre-dispatch, evidence-backed):
runners/launch_b200-nscale-compat.sh, which for dynamo-sglang + dsr1 + fp4 clonesNVIDIA/srt-slurmat pina98738de9b2233459b5456e9ed71af09ce893f92and applies the upstream reciperecipes/b200-fp4/8k1k.yaml(not checked into InferenceX). That recipe setsdynamo.version: 0.8.1andmodel.container: dynamo-sglang; the launcher maps that alias to the master image's squash file, so the image bump does reach the containers.pip install … ai-dynamo-runtime==0.8.1 ai-dynamo==0.8.1(src/srtctl/core/schema.py,DynamoConfig.get_install_commands). ai-dynamo 0.8.1 (PyPI, 2026-01-23) declaressglang==0.5.6.post2and was validated in production against the current v0.5.8.post1 image.dynamo.sglangimports, unguarded and at module top level:from sglang.srt.server_args_config_parser import ConfigArgumentMerger(dynamo/sglang/args.py:19),from sglang.srt.tracing import trace(dynamo/sglang/request_handlers/handler_base.py:15) andfrom sglang.srt.utils import get_local_ip_auto, get_zmq_socket, maybe_wrap_ipv6_address(dynamo/sglang/publisher.py:12); all three modules are imported bydynamo/sglang/main.py.sglang==0.5.19wheel (same version thev0.5.19-cu130-runtimeDockerfile installs),sglang/srt/server_args_config_parser.pyno longer exists (moved tosglang/srt/utils/server_args_config_parser.py), thesglang/srt/tracing/package no longer exists (moved tosglang/srt/observability/trace.py), andmaybe_wrap_ipv6_addressis not defined anywhere in the package. All three existed inv0.5.8.post1. Every prefill/decode worker would therefore exit withModuleNotFoundErrorbefore loading weights; the recipe's health check (max_attemptsraised to 720 × 10 s by the launcher) would then hold the 2- and 6-node allocations for up to two hours per job.router.version: "0.8.0"is display metadata only; no launcher or workflow consumes it, so it cannot change the installed dynamo version.dynamo.version/dynamo.hashin the recipe or a different recipe path/launcher pin. The recipe is external and shared, the master recipe reference must be preserved, and the launcher is shared code, so no repair exists within the Klaud Cold edit scope. The MTP siblingdsr1-fp4-b200-dynamo-sglang-mtpalready runslmsysorg/sglang:v0.5.12.post1with checked-in recipes anddynamo.hash: 5b4bc1dd70965017a737c71b19db5a0aeaa88727; the same pattern is the natural manual fix.Outcome: not dispatched — confirmed incompatibility between the candidate image and the family's recipe-pinned ai-dynamo 0.8.1, deliberately not confirmed by burning B200 nodes on a deterministic import failure. Repairs used: 0/5.
Next step: manual recipe/launcher change by a maintainer, then a new image-refresh candidate.
镜像 / 提交:
lmsysorg/sglang:v0.5.19-cu130-runtime@sha256:710bc1…a2f052@8ab3bdfbd4078d03018188fe5470dcdea019a850。运行链接:不适用 —— 未派发任何
e2e-tests.yml运行(原因见下)。未使用 GPU 时间。基准 / 评测结果:不适用 —— 未运行。
逐点相对基线变化:不适用 —— 未运行。
诊断(派发前,有证据支持):
runners/launch_b200-nscale-compat.sh运行;对于 dynamo-sglang + dsr1 + fp4,它会以固定提交a98738de9b2233459b5456e9ed71af09ce893f92克隆NVIDIA/srt-slurm,并使用上游配方recipes/b200-fp4/8k1k.yaml(未纳入 InferenceX 仓库)。该配方设置dynamo.version: 0.8.1与model.container: dynamo-sglang;启动器把该别名映射到主镜像的 squash 文件,因此镜像更新确实会进入容器。pip install … ai-dynamo-runtime==0.8.1 ai-dynamo==0.8.1安装服务栈(src/srtctl/core/schema.py的DynamoConfig.get_install_commands)。ai-dynamo 0.8.1(PyPI,2026-01-23)声明依赖sglang==0.5.6.post2,并在生产中已与当前 v0.5.8.post1 镜像验证通过。dynamo.sglang在模块顶层无保护地导入:from sglang.srt.server_args_config_parser import ConfigArgumentMerger(dynamo/sglang/args.py:19)、from sglang.srt.tracing import trace(dynamo/sglang/request_handlers/handler_base.py:15)以及from sglang.srt.utils import get_local_ip_auto, get_zmq_socket, maybe_wrap_ipv6_address(dynamo/sglang/publisher.py:12);这三个模块均由dynamo/sglang/main.py导入。sglang==0.5.19wheel(与v0.5.19-cu130-runtimeDockerfile 安装的版本一致)中,sglang/srt/server_args_config_parser.py已不存在(迁移至sglang/srt/utils/server_args_config_parser.py),sglang/srt/tracing/包已不存在(迁移至sglang/srt/observability/trace.py),且maybe_wrap_ipv6_address在整个包中均无定义。三者在v0.5.8.post1中都存在。因此每个 prefill/decode worker 都会在加载权重之前因ModuleNotFoundError退出;配方的健康检查(启动器将max_attempts提高到 720 × 10 秒)随后会让 2 节点与 6 节点分配每个作业最多占用两小时。router.version: "0.8.0"仅为展示元数据;没有启动器或工作流消费它,因此无法改变实际安装的 dynamo 版本。dynamo.version/dynamo.hash,或更换配方路径 / 启动器固定提交。该配方位于外部且共享,主配置的配方引用必须保留,启动器属于共享代码,因此 Klaud Cold 的编辑范围内不存在可行修复。MTP 兄弟家族dsr1-fp4-b200-dynamo-sglang-mtp已使用仓库内配方与dynamo.hash: 5b4bc1dd70965017a737c71b19db5a0aeaa88727运行lmsysorg/sglang:v0.5.12.post1;相同模式是自然的人工修复方案。结果:未派发 —— 已确认候选镜像与该家族配方固定的 ai-dynamo 0.8.1 不兼容,有意不通过占用 B200 节点来复现一个确定性的导入失败。已使用修复次数:0/5。
下一步:由维护者人工修改配方 / 启动器,然后产生新的镜像刷新候选。
Final full sweep / 最终全量 sweep
perf-changelog.yamlentry was appended, the PR was never marked ready andfull-sweep-enabledwas never applied. Norun-sweep.ymlrun exists for this head.perf-changelog.yaml条目,PR 未转为 ready,也未添加full-sweep-enabled标签。该 head 不存在任何run-sweep.yml运行。Disposition / 处置
klaud/auto-906fff6c74191d2d-fa39a9e32aba886aretained so only this exact candidate is blocked. A changed source image or release (or a fixed recipe/launcher pin) can be selected again.klaud/auto-906fff6c74191d2d-fa39a9e32aba886a,仅阻止这一确切候选。源镜像 / 发布版本变化(或配方 / 启动器固定提交修复后)可再次被选中。🤖 Generated with Claude Code