results: Arc Pro B70 Qwen3.8-Flash-Next Q2_0 (llama.cpp sycl + vulkan) - #48
Merged
Merged
Conversation
Two llama.cpp b11026 runs of ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Q2_0 on a single Arc Pro B70, sycl and vulkan. The model is 176.9B parameters and its smallest full-weight GGUF is 61.9GiB, so neither run is fully VRAM-resident: expert tensors of the first N of 48 layers stay in system RAM via -ncmoe (8 for sycl, 12 for vulkan). Both notes record the offload split, since the numbers are not comparable to a fully resident result.
|
@shacortes is attempting to deploy a commit to the Community Labs Team on Vercel. A member of the Team first needs to authorize it. |
- results/shacortes/qwen3-8-flash-next-llamacpp-q2_0-sycl.json: imported (result 86) - results/shacortes/qwen3-8-flash-next-llamacpp-q2_0-vulkan.json: imported (result 87) Merge ingestion completed. Rerunning this workflow will not duplicate these results. |
The notes described the -ncmoe offload as expert tensors being "held in system RAM", which covers residency but not execution. -ncmoe overrides those tensors to ggml_backend_cpu_buffer_type, so their matmuls run on the CPU backend, not just their storage. Says so explicitly on both runs, and states what stays on the GPU: all 48 layers via -ngl 99, attention, the KV cache, every non-expert tensor, and the experts of the remaining layers. Drops the vulkan note's claim that the remainder "oversubscribes" the 32GB card - at -ncmoe 12 less is resident than the sycl run's 24.2GiB at -ncmoe 8, and an oversubscribed allocation would have aborted.
|
nice PR |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two llama.cpp runs of Qwen3.8-Flash-Next at Q2_0 on a single Intel Arc Pro B70, against the Q2_0 board added in #42.
llama-bench
tg256/pp512, 5 runs, single request, negligible context depth — stock upstream ggml-org/llama.cpp b11026 (b49650adb), clean tree.These are hybrid runs, not pure-GPU ones
Worth being precise about, since it affects how the numbers compare to a fully resident board.
-ngl 99offloads all 48 layers to the card.-ncmoe Nthen overrides the MoE expert FFN tensors of the first N layers toggml_backend_cpu_buffer_type(), so those tensors both live in system RAM and have their matmuls executed on the CPU backend.Everything else is on the Arc: attention, the KV cache, every non-expert tensor, and the experts of the remaining 40 (sycl) / 36 (vulkan) layers — 24.2GiB of weights resident on the sycl run. Because the model is MoE, only the experts actually routed per token in the offloaded layers ever reach the CPU.
The offload isn't optional: the model is 176.9B parameters (~3B active per token) and its smallest full-weight GGUF — ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, Q2_0 — is 61.9GiB, which cannot fit in 32GB. On sycl,
-ncmoe 8is both the optimum and the floor;-ncmoe 4aborts out of VRAM. The two backends needed different splits, so they aren't directly comparable to each other either.Context depth, not offload, drives the serving gap
The sycl note records live serving figures from the same machine and weights at
-ncmoe 16, labelled as serving rather than benchmark: 19.4 tok/s decode on a 41,415-token cold-cache prompt at 263.4 tok/s prefill, sustained 19.5–21.7 tok/s across 22k–41k context. Most of the gap from 28.15 is depth, not offload — atg256sweep at depth 0 gives 28.13 at-ncmoe 8versus 27.13 at-ncmoe 16, about 1 tok/s. Prefill holds up with length where decode does not.Both files pass
frontend/scripts/validate-results.mjs --author shacortes.