Skip to content

results: Arc Pro B70 Qwen3.8-Flash-Next Q2_0 (llama.cpp sycl + vulkan) - #48

Merged
jackwsmth merged 2 commits into
labscommunity:mainfrom
shacortes:results/flash-next-q2_0
Sep 26, 2026
Merged

jackwsmth merged 2 commits into
labscommunity:mainfrom
shacortes:results/flash-next-q2_0

Conversation

@shacortes

@shacortes shacortes commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Two llama.cpp runs of Qwen3.8-Flash-Next at Q2_0 on a single Intel Arc Pro B70, against the Q2_0 board added in #42.

backend decode tok/s prefill tok/s MoE experts on CPU
sycl 28.15 453.66 first 8 of 48 layers
vulkan 19.40 202.97 first 12 of 48 layers

llama-bench tg256/pp512, 5 runs, single request, negligible context depth — stock upstream ggml-org/llama.cpp b11026 (b49650adb), clean tree.

These are hybrid runs, not pure-GPU ones

Worth being precise about, since it affects how the numbers compare to a fully resident board. -ngl 99 offloads all 48 layers to the card. -ncmoe N then overrides the MoE expert FFN tensors of the first N layers to ggml_backend_cpu_buffer_type(), so those tensors both live in system RAM and have their matmuls executed on the CPU backend.

Everything else is on the Arc: attention, the KV cache, every non-expert tensor, and the experts of the remaining 40 (sycl) / 36 (vulkan) layers — 24.2GiB of weights resident on the sycl run. Because the model is MoE, only the experts actually routed per token in the offloaded layers ever reach the CPU.

The offload isn't optional: the model is 176.9B parameters (~3B active per token) and its smallest full-weight GGUF — ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, Q2_0 — is 61.9GiB, which cannot fit in 32GB. On sycl, -ncmoe 8 is both the optimum and the floor; -ncmoe 4 aborts out of VRAM. The two backends needed different splits, so they aren't directly comparable to each other either.

Context depth, not offload, drives the serving gap

The sycl note records live serving figures from the same machine and weights at -ncmoe 16, labelled as serving rather than benchmark: 19.4 tok/s decode on a 41,415-token cold-cache prompt at 263.4 tok/s prefill, sustained 19.5–21.7 tok/s across 22k–41k context. Most of the gap from 28.15 is depth, not offload — a tg256 sweep at depth 0 gives 28.13 at -ncmoe 8 versus 27.13 at -ncmoe 16, about 1 tok/s. Prefill holds up with length where decode does not.

Both files pass frontend/scripts/validate-results.mjs --author shacortes.

Two llama.cpp b11026 runs of ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
Q2_0 on a single Arc Pro B70, sycl and vulkan.

The model is 176.9B parameters and its smallest full-weight GGUF is
61.9GiB, so neither run is fully VRAM-resident: expert tensors of the
first N of 48 layers stay in system RAM via -ncmoe (8 for sycl, 12 for
vulkan). Both notes record the offload split, since the numbers are not
comparable to a fully resident result.
@vercel

vercel Bot commented Sep 26, 2026

Copy link
Copy Markdown

@shacortes is attempting to deploy a commit to the Community Labs Team on Vercel.

A member of the Team first needs to authorize it.

@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
- results/shacortes/qwen3-8-flash-next-llamacpp-q2_0-sycl.json: imported (result 86)
- results/shacortes/qwen3-8-flash-next-llamacpp-q2_0-vulkan.json: imported (result 87)
Merge ingestion completed. Rerunning this workflow will not duplicate these results.

Site / sign up · Workflow details and retry

The notes described the -ncmoe offload as expert tensors being "held in
system RAM", which covers residency but not execution. -ncmoe overrides
those tensors to ggml_backend_cpu_buffer_type, so their matmuls run on
the CPU backend, not just their storage.

Says so explicitly on both runs, and states what stays on the GPU: all
48 layers via -ngl 99, attention, the KV cache, every non-expert tensor,
and the experts of the remaining layers.

Drops the vulkan note's claim that the remainder "oversubscribes" the
32GB card - at -ncmoe 12 less is resident than the sycl run's 24.2GiB
at -ncmoe 8, and an oversubscribed allocation would have aborted.
@jackwsmth
jackwsmth merged commit 81bd31a into labscommunity:main Sep 26, 2026
3 of 4 checks passed
@ByGamer01

Copy link
Copy Markdown

nice PR

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants