This repository accompanies CacheSched: Efficient LLM Request Scheduling for Hierarchical Prefix Caching. It provides the serving project, simulator, performance profiles, experiment workflows, and paper data.
There are two ways to evaluate this artifact. Both reach the same figures and tables.
- Option A — your own machines. Follow the standard flow below on hardware that meets the requirements.
- Option B — our machines. Request temporary SSH access through HotCRP; we open an account for the time window you ask for and close it when you are done.
Reviewers who only want to check the plots against the paper data need neither option: the quick start runs on any CPU-only laptop.
Base requirements for every GPU experiment:
| Item | Requirement |
|---|---|
| OS and arch | Linux x86_64 |
| Python | 3.10 |
| Toolchain | Rust 1.85.1, standard build tools (gcc, make, cmake) |
| GPU stack | NVIDIA driver supporting CUDA 12.4 |
| Disk | ~130 GiB for model weights; allow 250 GiB for models, environments, caches, and results |
| Network | Hugging Face access; meta-llama/Llama-3.1-8B-Instruct is gated and needs an approved account |
| Extra (only for the vLLM and Preble groups) | Docker, NVIDIA Container Toolkit, Kubernetes |
GPUs needed per target. Every GPU experiment runs on a single node; the parallelism in the table is what the supplied recipes configure.
| Target | Recipe | GPUs |
|---|---|---|
| All figures and Table 3 from supplied data | — | None (CPU only) |
| Installation smoke test | gpu_smoke |
1 GPU, 48 GiB class (e.g. one L40S) |
| Figure 10, Figure 11 | overall_synthetic, overall_sharegpt |
8x L40S 48 GiB, one node (dp_size 8) |
| Figure 14, Table 3 | component_ablation |
8x A100 80 GiB, one node (Llama dp_size 8; Mixtral dp_size 4, tp_size 2) |
| Preble comparison | preble |
8x A100 80 GiB plus a Kubernetes cluster (AIBrix setup) |
| vLLM group inside Figures 10-11 | overall_* with --methods vLLM |
Same 8x L40S node plus Production Stack |
| Figure 15 | sensitivity_skewness |
None (CPU simulator on the simulator branch) |
Hardware of a different class still runs, but the absolute numbers will not match the paper; the supplied performance parameters in profile/configs/ are fitted for L40S and A100. To use other GPUs, generate new parameters with profiling.
Then follow, in order: installation (environment, kernels, model downloads), the minimal traffic test to confirm the install, the experiment guide for the recipes above, and plotting.
We keep the evaluation nodes (8x L40S and 8x A100) installed and ready, with the repository, Python environments, model weights, and ShareGPT4 data already in place, so no setup step is needed.
-
Post a comment on the artifact submission in HotCRP containing:
- your SSH public key (the full contents of
~/.ssh/id_ed25519.pub, one line starting withssh-ed25519orssh-rsa; never send a private key); - the time window you want, with time zone and date, for example
2026-04-12 09:00 - 2026-04-14 18:00 UTC; - which nodes you need (L40S for Figures 10-11, A100 for Figure 14 and Table 3).
Please request the window at least 12 hours ahead.
- your SSH public key (the full contents of
-
We install the key and confirm on HotCRP. Before the window opens we add your key to
authorized_keyson the requested node and reply with the exact login command, the checkout path, the environment activation command, and the model paths file to pass to--paths. -
Run the experiments during your window. The commands are the ones in the experiment guide; the installation and download steps are already done. Write results to your own directory under your home, for example
--output ~/results/ablation. Tell us on HotCRP if you need the window extended; we will try our best to extend it. -
Copy out anything you want to keep before the window ends, for example with
scp -r <user>@<host>:~/results/ablation ./. -
We remove the key when the window ends or as soon as you tell us you are finished, and the account is closed. The key is used only for this access and is not shared with anyone else. All communication stays inside HotCRP, so reviewer identities remain anonymous to us.
If a run fails or a node is unavailable, please let us know in HotCRP with the failing command and the contents of server.log and benchmark.log from the affected point; we will respond there.
Clone the repository URL supplied with the artifact submission:
git clone --branch main <repository-url> cachesched-ae
cd cachesched-ae| Branch | Purpose |
|---|---|
main |
Complete CacheSched GPU serving project and experiment tools. Start here. |
simulator |
The same project layout with the original simulator backend. Use it for simulation experiments. |
v0.4.5 |
Original SGLang baseline, installed separately by the setup script. |
Switch between the GPU and simulator source:
git switch simulator
git switch mainAll commands run from the repository root using Python 3.10. Plotting uses the active experiment environment.
| Directory | Purpose | Instructions |
|---|---|---|
code/ |
Serving engine, router, original clients, and kernel source | Engine |
setup/ |
Environment and dependencies | Installation |
profile/ |
Use saved performance parameters or sample new ones | Profiling |
benchmark/ |
Run experiments and collect results | Experiments |
plotting/ |
Plot supplied data or new measurements | Plotting |
To run experiments, read installation, then experiments, then plotting. The experiment guide covers the initial traffic test, experiment and comparison-group selection, configuration files, and retries. To use supplied data, follow the quick start below and continue directly to plotting.
Read engine entry points to see the underlying launch and client commands. Profiling is needed only to select or create performance parameters; Production Stack covers the vLLM group. Plotting data lists data locations and how to rebuild the plotting CSV.
Create the Python environment if needed, then install the plotting dependencies. The same environment can later be extended with the GPU or simulator dependencies in the installation guide.
python3.10 -m venv .venv
. .venv/bin/activate
python -m pip install -r setup/requirements-plot.txt
python plotting/plot.py
python plotting/table3.pyThis produces six experimental figures as PDF under results/paper/figures/ and Table 3 under results/paper/tables/. Supplied-data plotting needs only a CPU.
Install the selected backend using setup/README.md, then configure local paths and follow benchmark/README.md.
| Experiment | Recipe | Estimated minimum runtime |
|---|---|---|
| Figure 10: synthetic conversation performance | overall_synthetic |
6 h |
| Figure 11: ShareGPT4 conversation performance | overall_sharegpt |
10 h |
| Figure 14: router and engine component comparison on Llama and Mixtral | component_ablation |
4 h |
| Table 3: throughput and goodput on Llama and Mixtral | component_ablation |
2 h |
| Figure 15: conversation skewness | sensitivity_skewness |
5 h |
These estimated times are based on our full experiments. Low-QPS runs account for a substantial share of the time; a partial selection may be considered when time is limited. Table 3 includes only the additional Preble runtime, reusing the Figure 14 results.
The experiment guide also provides minimal examples and commands to retry selected request rates. Figures 12–13 are available through supplied-data plotting. Saved performance parameters are ready to use; profiling is optional.
To plot newly collected data, pass the experiment directory or CSV:
python plotting/plot.py --data results/synthetic
python plotting/plot.py --data results/synthetic/plot_data.csvWithout --data, output goes to results/paper/figures/. With --data, output goes to figures/ alongside the CSV. Use --output to choose another destination.