Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,15 @@
# Changelog

## Unreleased

- Add a budget or penalty on case churn or expected negative flips, and match comparisons on
both queue movement and expected flips.
- Price movement by queue (`queue_costs`), so budgets can count cases, hours, or money.
- Return per-queue dual prices from the flow solver.
- Breaking: `solve_assignment` takes incumbent labels instead of incumbent loads.
- Add a regime simulation mapping when holding queue totals is costly and how learned prices
compare with exact per-batch assignment.

## 0.1.0 - 2026-08-17

- Define queue shift as incumbent-relative workload movement across predicted queues.
Expand Down
28 changes: 25 additions & 3 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@ PYTHON := uv run python
SOURCE_DATE_EPOCH := 1786960800
export SOURCE_DATE_EPOCH

.PHONY: sync format lint test build theory oracle estimated analyze reproduce check-generated paper research-check check clean
.PHONY: sync format lint test build theory oracle estimated regimes analyze reproduce check-generated paper research-check check clean

sync:
uv sync --all-groups --all-extras
Expand Down Expand Up @@ -32,6 +32,9 @@ oracle:
estimated:
$(PYTHON) -m experiments.run_estimated --outdir results/estimated

regimes:
$(PYTHON) -m experiments.run_regimes --outdir results/regimes

analyze:
MPLCONFIGDIR=/tmp/queue-shift-matplotlib $(PYTHON) -m experiments.analyze_oracle \
results/oracle \
Expand All @@ -45,8 +48,17 @@ analyze:
--tex paper/generated/estimated_summary.tex \
--figure paper/figures/estimated_validation.pdf \
--macros paper/generated/estimated_numbers.tex

reproduce: oracle estimated analyze
MPLCONFIGDIR=/tmp/queue-shift-matplotlib $(PYTHON) -m experiments.analyze_regimes \
results/regimes \
--out results/regimes_summary.csv \
--lock-out results/regimes_lock_cost.csv \
--lock-tex paper/generated/regimes_lock_cost.tex \
--tex paper/generated/regimes_offsets.tex \
--frontier paper/figures/regimes_frontier.pdf \
--mix-gap paper/figures/regimes_mix_gap.pdf \
--macros paper/generated/regimes_numbers.tex

reproduce: oracle estimated regimes analyze

check-generated:
@tmp=$$(mktemp -d) && \
Expand All @@ -56,12 +68,22 @@ check-generated:
MPLCONFIGDIR=/tmp/queue-shift-matplotlib $(PYTHON) -m experiments.analyze_estimated \
results/estimated --out $$tmp/estimated_summary.csv --tex $$tmp/estimated_summary.tex \
--figure $$tmp/estimated_validation.pdf --macros $$tmp/estimated_numbers.tex >/dev/null && \
MPLCONFIGDIR=/tmp/queue-shift-matplotlib $(PYTHON) -m experiments.analyze_regimes \
results/regimes --out $$tmp/regimes_summary.csv --lock-out $$tmp/regimes_lock_cost.csv \
--lock-tex $$tmp/regimes_lock_cost.tex \
--tex $$tmp/regimes_offsets.tex --frontier $$tmp/regimes_frontier.pdf \
--mix-gap $$tmp/regimes_mix_gap.pdf --macros $$tmp/regimes_numbers.tex >/dev/null && \
diff -q results/oracle_summary.csv $$tmp/oracle_summary.csv && \
diff -q results/estimated_summary.csv $$tmp/estimated_summary.csv && \
diff -q paper/generated/oracle_summary.tex $$tmp/oracle_summary.tex && \
diff -q paper/generated/estimated_summary.tex $$tmp/estimated_summary.tex && \
diff -q paper/generated/oracle_numbers.tex $$tmp/oracle_numbers.tex && \
diff -q paper/generated/estimated_numbers.tex $$tmp/estimated_numbers.tex && \
diff -q results/regimes_summary.csv $$tmp/regimes_summary.csv && \
diff -q results/regimes_lock_cost.csv $$tmp/regimes_lock_cost.csv && \
diff -q paper/generated/regimes_offsets.tex $$tmp/regimes_offsets.tex && \
diff -q paper/generated/regimes_lock_cost.tex $$tmp/regimes_lock_cost.tex && \
diff -q paper/generated/regimes_numbers.tex $$tmp/regimes_numbers.tex && \
rm -rf $$tmp && echo "generated results are synchronized"

paper:
Expand Down
72 changes: 54 additions & 18 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,16 +4,19 @@ Classifier updates should respect the cost they actually create. When prediction
specialized queues, that cost is the change in queue totals, not the number of individual
predictions that change.

Queue Shift finds the highest-value batch assignment under a limit on queue-load movement. The
problem is a minimum-cost flow, so the solver returns an integral global optimum. At the movement
budget produced by any comparison rule, including a negative-flip method, Queue Shift weakly
improves the accepted model's total probability score. If those probabilities are the true
conditional probabilities, the result is dominance in conditional expected accuracy.

The paper proves the result and reports two stable-distribution experiments. The main validation
estimates both classifiers on finite historical samples, releases the candidate only after an
independent accuracy win, and tests small samples, weak innovation, overconfidence, and
underconfidence.
Queue Shift finds the highest-value batch assignment under two limits: on work moved between
queues, counted in cases, hours, or money, and on expected negative flips. With only the movement
limit, or with a penalty on flips, the problem is a minimum-cost flow with an integral global
optimum and per-queue prices that reproduce it case by case; hard flip budgets and per-queue
prices are solved exactly as small integer programs. At the movement and expected flips of any
comparison rule, including a negative-flip interpolation, Queue Shift has at least the rule's total
probability score. If those probabilities are the true conditional probabilities, the result is
dominance in conditional expected accuracy.

The paper proves the result and uses simulation to map where the constraint matters. Holding queue
totals costs little when an update mostly reorders cases and a lot when it moves volume between
queues. Learned queue prices nearly match exact assignment in accuracy but miss the movement budget
in most batches.

Read the current manuscript: [Queue Shift: Updating Classifiers Under a Staffing
Budget](paper/main.pdf).
Expand All @@ -29,7 +32,7 @@ uv sync --all-groups --all-extras
```python
import numpy as np

from queue_shift import solve_assignment
from queue_shift import case_flip_weights, solve_assignment

candidate_probability = np.array(
[
Expand All @@ -38,19 +41,52 @@ candidate_probability = np.array(
[0.20, 0.30, 0.50],
]
)
incumbent_loads = np.array([1, 1, 1])
incumbent_labels = np.array([0, 2, 1])

shift_only = solve_assignment(
costs=1 - candidate_probability,
incumbent_labels=incumbent_labels,
move_budget=0,
)
print(shift_only.labels, shift_only.churn) # [0 1 2] 2: a swap keeps totals

result = solve_assignment(
weights = case_flip_weights(
candidate_probability, incumbent_labels, "expected_negative"
)
flip_limited = solve_assignment(
costs=1 - candidate_probability,
incumbent_loads=incumbent_loads,
incumbent_labels=incumbent_labels,
move_budget=0,
flip_weights=weights,
flip_budget=0.2,
)
# [0 2 1] 0.0: the swap would risk 0.1 + 0.3 = 0.4 expected negative flips
print(flip_limited.labels, flip_limited.flip_load)
```

The movement budget limits work added to queues beyond their incumbent loads. By default work is
counted in cases, and the budget equals half the L1 distance between the proposed and incumbent
queue-load vectors. A budget of zero preserves every incumbent queue total while allowing different
cases to fill those slots. Pass `queue_costs` to price a case in each queue, for example in average
handle hours or cost per case; the budget is then in hours or money:

```python
hours_per_case = np.array([0.5, 1.0, 4.0])
by_hours = solve_assignment(
costs=1 - candidate_probability,
incumbent_labels=incumbent_labels,
move_budget=2.0,
queue_costs=hours_per_case,
)
print(result.labels)
print(by_hours.labels, by_hours.workload_shift) # [0 1 2] 0.0
```

The movement budget is half the L1 distance between the proposed and incumbent queue-load
vectors. A budget of zero preserves every incumbent queue total while allowing different cases to
fill those slots.
Flip weights price a case leaving its incumbent queue. `"expected_negative"` weights a move by
the candidate's probability that the incumbent was right, so the weighted total is the expected
number of negative flips; `"churn"` counts every changed label. `flip_budget` caps the total
exactly (solved as a mixed-integer program). `flip_penalty` instead charges for it, which keeps
the problem a minimum-cost flow, scales to large batches, and returns per-queue prices
(`queue_prices`) that reproduce the assignment case by case.

## Reproduce the paper

Expand Down
47 changes: 33 additions & 14 deletions experiments/analyze_estimated.py
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,13 @@ def summarize(release: pd.DataFrame, results: pd.DataFrame) -> pd.DataFrame:
raise RuntimeError("an exact assignment exceeded its matched movement budget")
if (results["predicted_gain_pp"] < -1e-9).any():
raise RuntimeError("supplied-score dominance failed")
joint_budget_slack = 2e-6 * (1 + results["n_cases"])
if (
results["joint_expected_flips"] > results["flip_budget"] + joint_budget_slack
).any():
raise RuntimeError("a joint assignment exceeded its matched flip budget")
if (results["joint_predicted_gain_pp"] < -1e-9).any():
raise RuntimeError("joint supplied-score dominance failed")
if (results["conditional_gain_pp"] < results["robust_lower_bound_pp"] - 1e-8).any():
raise RuntimeError("the probability-error lower bound failed")
endpoint = results[results["alpha"] == 1.0]
Expand All @@ -54,6 +61,9 @@ def summarize(release: pd.DataFrame, results: pd.DataFrame) -> pd.DataFrame:
).agg(
move_share=("baseline_move_share", "mean"),
nfr=("baseline_nfr", "mean"),
operational_nfr=("operational_nfr", "mean"),
joint_nfr=("joint_nfr", "mean"),
joint_conditional_gain_pp=("joint_conditional_gain_pp", "mean"),
predicted_gain_pp=("predicted_gain_pp", "mean"),
conditional_gain_pp=("conditional_gain_pp", "mean"),
realized_gain_pp=("realized_gain_pp", "mean"),
Expand All @@ -63,6 +73,10 @@ def summarize(release: pd.DataFrame, results: pd.DataFrame) -> pd.DataFrame:
repetitions=("realized_gain_pp", "size"),
move_share=("move_share", "mean"),
nfr=("nfr", "mean"),
operational_nfr=("operational_nfr", "mean"),
joint_nfr=("joint_nfr", "mean"),
joint_conditional_gain_pp=("joint_conditional_gain_pp", "mean"),
joint_conditional_gain_se=("joint_conditional_gain_pp", "sem"),
predicted_gain_pp=("predicted_gain_pp", "mean"),
conditional_gain_pp=("conditional_gain_pp", "mean"),
conditional_gain_se=("conditional_gain_pp", "sem"),
Expand All @@ -89,17 +103,20 @@ def write_table(summary: pd.DataFrame, path: Path) -> None:
selected = summary[summary["alpha"].isin([0.0, 0.5, 1.0])]
rows = []
for row in selected.itertuples(index=False):
conditional_low = row.conditional_gain_pp - 1.96 * row.conditional_gain_se
conditional_high = row.conditional_gain_pp + 1.96 * row.conditional_gain_se
realized_low = row.realized_gain_pp - 1.96 * row.realized_gain_se
realized_high = row.realized_gain_pp + 1.96 * row.realized_gain_se
qs_low = row.conditional_gain_pp - 1.96 * row.conditional_gain_se
qs_high = row.conditional_gain_pp + 1.96 * row.conditional_gain_se
joint_low = row.joint_conditional_gain_pp - 1.96 * row.joint_conditional_gain_se
joint_high = (
row.joint_conditional_gain_pp + 1.96 * row.joint_conditional_gain_se
)
rows.append(
f"{SCENARIO_LABELS[row.scenario]} & {row.alpha:.1f} & "
f"{100 * row.move_share:.2f} & "
f"{row.predicted_gain_pp:.2f} & "
f"{row.conditional_gain_pp:.2f} "
f"[{conditional_low:.2f}, {conditional_high:.2f}] & "
f"{row.realized_gain_pp:.2f} [{realized_low:.2f}, {realized_high:.2f}] \\\\"
f"{100 * row.nfr:.1f} & {100 * row.operational_nfr:.1f} & "
f"{100 * row.joint_nfr:.1f} & "
f"{row.conditional_gain_pp:.2f} [{qs_low:.2f}, {qs_high:.2f}] & "
f"{row.joint_conditional_gain_pp:.2f} "
f"[{joint_low:.2f}, {joint_high:.2f}] \\\\"
)
content = "\n".join(
[
Expand All @@ -108,13 +125,14 @@ def write_table(summary: pd.DataFrame, path: Path) -> None:
r"\caption{Estimated-model validation against NFR interpolation}",
r"\label{tab:estimated-validation}",
r"\scriptsize",
r"\begin{tabular}{lrrrrr}",
r"\begin{tabular}{lrrrrrrr}",
r"\toprule",
(
r"Scenario & $\alpha$ & Shift (\%) & Score gain & "
r"True gain [95\% CI] & Realized gain [95\% CI] \\"
r" & & Shift & \multicolumn{3}{c}{NFR (\%)} & "
r"\multicolumn{2}{c}{True gain over interp. (pp) [95\% CI]} \\"
),
r" & & & (pp) & (pp) & (pp) \\",
r"\cmidrule(lr){4-6}\cmidrule(lr){7-8}",
r"Scenario & $\alpha$ & (\%) & Interp. & QS & QS+flips & QS & QS+flips \\",
r"\midrule",
*rows,
r"\bottomrule",
Expand All @@ -125,8 +143,9 @@ def write_table(summary: pd.DataFrame, path: Path) -> None:
r"estimated on historical samples and released only after the candidate "
r"wins on an independent labeled sample. Each row compares exact "
r"assignment with probability interpolation at the same realized queue "
r"shift. True gain uses the simulation's conditional probabilities and "
r"is unavailable in applications. Results average four planning batches "
r"shift (QS), and at both its queue shift and its expected negative flips "
r"(QS+flips). True gain uses the simulation's conditional "
r"probabilities and is unavailable in applications. Results average four planning batches "
r"within each independently trained accepted update before averaging "
r"across updates."
),
Expand Down
Loading
Loading