Skip to content

ADD: SVRIMG convective mode benchmark (Cellular vs QLCS) - #37

Merged
rcjackson merged 1 commit into
ARM-Development:mainfrom
rcjackson:svrimg
Sep 14, 2026
Merged

rcjackson merged 1 commit into
ARM-Development:mainfrom
rcjackson:svrimg

Conversation

@rcjackson

Copy link
Copy Markdown
Collaborator

Why

The labeller has only ever been scored against in-house CSAPR2 hand labels, so there is no evidence it generalises beyond Bankhead National Forest, and no way to compare it to published human annotation.

SVRIMG (Haberlie, Ashley & Karpinski 2020) supplies 1,741 hand-classified GridRad/NEXRAD scenes centred on SPC severe reports. Collapsed to the two convective mode classes, Cellular vs QLCS, that is 1,315 scenes with a published split:

Split Years N Cellular QLCS
train 1996–2011 997 541 456
val 2012–2013 88 45 43
test 2014–2017 230 78 152

What

  • lars/preprocessing/svrimg.py — fetch_svrimg_labels / preprocess_svrimg_data, mirroring preprocess_radar_data's CSV schema so the existing inference, Dawid-Skene and kappa tooling runs unchanged. Adds unid and split.
  • example_codebooks/CODEBOOK_SVRIMG.md — Cellular / QLCS / Ambiguous, same ten-section layout the parsers expect.
  • notebooks/reports/calibrate_svrimg_thresholds.py + results CSV.
  • tests/test_svrimg.py — 15 tests, no network.
  • Two-line additions to lars/preprocessing/__init__.py and STANDARD_LABEL_MAP.

Two format details worth knowing

The PNGs are data, not pictures. They are palette-mode images whose palette is a dummy linear ramp (i -> (3i, 3i+1, 3i+2)), so they render greyscale in any ordinary viewer. The palette indices carry the values. read_svrimg_png reads the indices and refuses a non-palette file, so an accidentally RGB-converted image fails loudly rather than producing plausible-looking nonsense.

Those indices are dBZ, integer-truncated and floored at zero — confirmed against the upstream GridRad file, where REFC_MAX is itself uint8 with units: dBZ. There are no negative values, so vmin=0, not the -20 used for CSAPR2.

Thresholds were recalibrated, not inherited

CODEBOOK_NWSRef.md's area fractions were tuned for 1 km CSAPR2 on a 2048×2048 polar grid; they do not transfer to 3.75 km column-max on a 136×136 Cartesian grid. Calibrated on the training split only:

Threshold AUC Youden cut Sens / Spec Adopted
pct_gates_10dBZ 0.903 29.98% 0.91 / 0.75 no
pct_gates_20dBZ 0.929 18.00% 0.87 / 0.83 no
pct_gates_30dBZ 0.935 6.48% 0.89 / 0.82 yes, as 6.5%
pct_gates_40dBZ 0.889 2.35% 0.78 / 0.83 no
pct_gates_50dBZ 0.601 — — no
pct_gates_60dBZ 0.684 0.0% (degenerate) — no

Only the strongest is adopted; the others are correlated and would stack redundant reclassification triggers.

⚠️ Caveat, also recorded in the codebook changelog: these fractions largely measure echo extent, not linearity. Expect them to misfire on large non-linear convective clusters, which SVRIMG annotators often labelled Cellular. The rule is a useful shortcut on this distribution, not a definition of convective mode.

Reviewer notes

  • The codebook describes classes in reflectivity and topology rather than colour, because COLOR_DBZ_RANGE in lars/nepho/inference.py is calibrated to a ChaseSpectral-like ramp (green 10–30, yellow 30–40, red 40–50) that does not match the NWSRef bands these scenes render with. Worth a separate look at whether CODEBOOK_NWSRef.md has the same mismatch.
  • The published split carries a prior shift — train is 54% Cellular, test is 66% QLCS. Report balanced accuracy and per-class recall; raw accuracy will flatter a model that merely learned the training prior.
  • No-echo gates are masked before rendering, mirroring preprocess_radar_data's where(field > min_ref). Without this they paint as the bottom of the ramp and fill the frame with apparent echo.

Testing

pytest tests/ --ignore=tests/integration → 116 passed on this branch, verified in a clean worktree off main so none of the in-flight class_scores work is mixed in.

Unrelated pre-existing issue: a full pytest tests/ run fails during collection because tests/integration/conftest.py globally evicts the pip_system_certs mock from sys.modules, breaking every top-level test file collected after it. Reproduces on main without this branch; not addressed here.

Not done: end-to-end labelling of the 230-image test split with a live model needs API credentials, so κ against SVRIMG ground truth is still to be measured.

🤖 Generated with Claude Code

Adds an external, human-labelled benchmark so the labeller can be scored
against published ground truth rather than only in-house CSAPR2 labels.

SVRIMG (Haberlie, Ashley and Karpinski 2020) provides 1,741 hand-classified
GridRad/NEXRAD scenes centred on SPC severe reports. Collapsed to the two
convective mode classes, Cellular and QLCS, that is 1,315 scenes across the
published split: 997 train, 88 validation, 230 test.

lars/preprocessing/svrimg.py mirrors preprocess_radar_data's CSV schema, so
the existing inference, Dawid-Skene and kappa tooling runs unchanged; 'unid'
and 'split' are added.

Two properties of the upstream format drive the implementation:

- The distributed PNGs are palette-mode images whose palette is a dummy
  linear ramp, so they look greyscale in an ordinary viewer. The palette
  indices are the data. read_svrimg_png reads the indices and refuses any
  non-palette file, so a silently RGB-converted image fails loudly instead
  of producing plausible-looking nonsense.
- Those indices are reflectivity in dBZ, integer-truncated and floored at
  zero (upstream GridRad REFC_MAX is itself uint8 with units dBZ). There are
  no negative values, hence vmin=0 rather than the -20 used for CSAPR2.

Thresholds are calibrated on the training split only, not carried over from
CODEBOOK_NWSRef.md, whose area fractions were tuned for 1 km CSAPR2 data on
a 2048x2048 polar grid and do not transfer to 3.75 km column-max on a
136x136 Cartesian grid. pct_gates_30dBZ is the strongest single
discriminator at AUC 0.935, Youden cut 6.48 percent, adopted as 6.5.
pct_gates_20/10/40dBZ also separate (AUC 0.929/0.903/0.889) but are omitted
to avoid stacking redundant correlated reclassification rules;
pct_gates_50dBZ does not discriminate (AUC 0.601).

Caveat recorded in the codebook changelog: these coverage fractions largely
measure echo extent rather than linearity, so they are expected to misfire
on large non-linear convective clusters.

The codebook describes classes in reflectivity and topology rather than
colour, because COLOR_DBZ_RANGE in lars/nepho/inference.py is calibrated to
a ChaseSpectral-like ramp that does not match the NWSRef bands these scenes
render with.

Note for reviewers: the published split carries a prior shift (train 54%
Cellular, test 66% QLCS), so report balanced accuracy and per-class recall
rather than raw accuracy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@rcjackson
rcjackson merged commit 1444aa7 into ARM-Development:main Sep 14, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant