Skip to content

Repository files navigation

CacheSched: Artifact Evaluation

This repository accompanies CacheSched: Efficient LLM Request Scheduling for Hierarchical Prefix Caching. It provides the serving project, simulator, performance profiles, experiment workflows, and paper data.

For artifact reviewers

There are two ways to evaluate this artifact. Both reach the same figures and tables.

  • Option A — your own machines. Follow the standard flow below on hardware that meets the requirements.
  • Option B — our machines. Request temporary SSH access through HotCRP; we open an account for the time window you ask for and close it when you are done.

Reviewers who only want to check the plots against the paper data need neither option: the quick start runs on any CPU-only laptop.

Option A: run on your own machines

Base requirements for every GPU experiment:

Item Requirement
OS and arch Linux x86_64
Python 3.10
Toolchain Rust 1.85.1, standard build tools (gcc, make, cmake)
GPU stack NVIDIA driver supporting CUDA 12.4
Disk ~130 GiB for model weights; allow 250 GiB for models, environments, caches, and results
Network Hugging Face access; meta-llama/Llama-3.1-8B-Instruct is gated and needs an approved account
Extra (only for the vLLM and Preble groups) Docker, NVIDIA Container Toolkit, Kubernetes

GPUs needed per target. Every GPU experiment runs on a single node; the parallelism in the table is what the supplied recipes configure.

Target Recipe GPUs
All figures and Table 3 from supplied data — None (CPU only)
Installation smoke test gpu_smoke 1 GPU, 48 GiB class (e.g. one L40S)
Figure 10, Figure 11 overall_synthetic, overall_sharegpt 8x L40S 48 GiB, one node (dp_size 8)
Figure 14, Table 3 component_ablation 8x A100 80 GiB, one node (Llama dp_size 8; Mixtral dp_size 4, tp_size 2)
Preble comparison preble 8x A100 80 GiB plus a Kubernetes cluster (AIBrix setup)
vLLM group inside Figures 10-11 overall_* with --methods vLLM Same 8x L40S node plus Production Stack
Figure 15 sensitivity_skewness None (CPU simulator on the simulator branch)

Hardware of a different class still runs, but the absolute numbers will not match the paper; the supplied performance parameters in profile/configs/ are fitted for L40S and A100. To use other GPUs, generate new parameters with profiling.

Then follow, in order: installation (environment, kernels, model downloads), the minimal traffic test to confirm the install, the experiment guide for the recipes above, and plotting.

Option B: request access to our machines through HotCRP

We keep the evaluation nodes (8x L40S and 8x A100) installed and ready, with the repository, Python environments, model weights, and ShareGPT4 data already in place, so no setup step is needed.

  1. Post a comment on the artifact submission in HotCRP containing:

    • your SSH public key (the full contents of ~/.ssh/id_ed25519.pub, one line starting with ssh-ed25519 or ssh-rsa; never send a private key);
    • the time window you want, with time zone and date, for example 2026-04-12 09:00 - 2026-04-14 18:00 UTC;
    • which nodes you need (L40S for Figures 10-11, A100 for Figure 14 and Table 3).

    Please request the window at least 12 hours ahead.

  2. We install the key and confirm on HotCRP. Before the window opens we add your key to authorized_keys on the requested node and reply with the exact login command, the checkout path, the environment activation command, and the model paths file to pass to --paths.

  3. Run the experiments during your window. The commands are the ones in the experiment guide; the installation and download steps are already done. Write results to your own directory under your home, for example --output ~/results/ablation. Tell us on HotCRP if you need the window extended; we will try our best to extend it.

  4. Copy out anything you want to keep before the window ends, for example with scp -r <user>@<host>:~/results/ablation ./.

  5. We remove the key when the window ends or as soon as you tell us you are finished, and the account is closed. The key is used only for this access and is not shared with anyone else. All communication stays inside HotCRP, so reviewer identities remain anonymous to us.

If a run fails or a node is unavailable, please let us know in HotCRP with the failing command and the contents of server.log and benchmark.log from the affected point; we will respond there.

Repository and branches

Clone the repository URL supplied with the artifact submission:

git clone --branch main <repository-url> cachesched-ae
cd cachesched-ae
Branch Purpose
main Complete CacheSched GPU serving project and experiment tools. Start here.
simulator The same project layout with the original simulator backend. Use it for simulation experiments.
v0.4.5 Original SGLang baseline, installed separately by the setup script.

Switch between the GPU and simulator source:

git switch simulator
git switch main

All commands run from the repository root using Python 3.10. Plotting uses the active experiment environment.

Components

Directory Purpose Instructions
code/ Serving engine, router, original clients, and kernel source Engine
setup/ Environment and dependencies Installation
profile/ Use saved performance parameters or sample new ones Profiling
benchmark/ Run experiments and collect results Experiments
plotting/ Plot supplied data or new measurements Plotting

Reading order

To run experiments, read installation, then experiments, then plotting. The experiment guide covers the initial traffic test, experiment and comparison-group selection, configuration files, and retries. To use supplied data, follow the quick start below and continue directly to plotting.

Read engine entry points to see the underlying launch and client commands. Profiling is needed only to select or create performance parameters; Production Stack covers the vLLM group. Plotting data lists data locations and how to rebuild the plotting CSV.

Quick start: plot supplied data

Create the Python environment if needed, then install the plotting dependencies. The same environment can later be extended with the GPU or simulator dependencies in the installation guide.

python3.10 -m venv .venv
. .venv/bin/activate
python -m pip install -r setup/requirements-plot.txt
python plotting/plot.py
python plotting/table3.py

This produces six experimental figures as PDF under results/paper/figures/ and Table 3 under results/paper/tables/. Supplied-data plotting needs only a CPU.

Run experiments

Install the selected backend using setup/README.md, then configure local paths and follow benchmark/README.md.

Experiment Recipe Estimated minimum runtime
Figure 10: synthetic conversation performance overall_synthetic 6 h
Figure 11: ShareGPT4 conversation performance overall_sharegpt 10 h
Figure 14: router and engine component comparison on Llama and Mixtral component_ablation 4 h
Table 3: throughput and goodput on Llama and Mixtral component_ablation 2 h
Figure 15: conversation skewness sensitivity_skewness 5 h

These estimated times are based on our full experiments. Low-QPS runs account for a substantial share of the time; a partial selection may be considered when time is limited. Table 3 includes only the additional Preble runtime, reusing the Figure 14 results.

The experiment guide also provides minimal examples and commands to retry selected request rates. Figures 12–13 are available through supplied-data plotting. Saved performance parameters are ready to use; profiling is optional.

To plot newly collected data, pass the experiment directory or CSV:

python plotting/plot.py --data results/synthetic
python plotting/plot.py --data results/synthetic/plot_data.csv

Without --data, output goes to results/paper/figures/. With --data, output goes to figures/ alongside the CSV. Use --output to choose another destination.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages