Repository navigation
ADD: SVRIMG convective mode benchmark (Cellular vs QLCS) - #37
Merged
Merged
Conversation
Adds an external, human-labelled benchmark so the labeller can be scored against published ground truth rather than only in-house CSAPR2 labels. SVRIMG (Haberlie, Ashley and Karpinski 2020) provides 1,741 hand-classified GridRad/NEXRAD scenes centred on SPC severe reports. Collapsed to the two convective mode classes, Cellular and QLCS, that is 1,315 scenes across the published split: 997 train, 88 validation, 230 test. lars/preprocessing/svrimg.py mirrors preprocess_radar_data's CSV schema, so the existing inference, Dawid-Skene and kappa tooling runs unchanged; 'unid' and 'split' are added. Two properties of the upstream format drive the implementation: - The distributed PNGs are palette-mode images whose palette is a dummy linear ramp, so they look greyscale in an ordinary viewer. The palette indices are the data. read_svrimg_png reads the indices and refuses any non-palette file, so a silently RGB-converted image fails loudly instead of producing plausible-looking nonsense. - Those indices are reflectivity in dBZ, integer-truncated and floored at zero (upstream GridRad REFC_MAX is itself uint8 with units dBZ). There are no negative values, hence vmin=0 rather than the -20 used for CSAPR2. Thresholds are calibrated on the training split only, not carried over from CODEBOOK_NWSRef.md, whose area fractions were tuned for 1 km CSAPR2 data on a 2048x2048 polar grid and do not transfer to 3.75 km column-max on a 136x136 Cartesian grid. pct_gates_30dBZ is the strongest single discriminator at AUC 0.935, Youden cut 6.48 percent, adopted as 6.5. pct_gates_20/10/40dBZ also separate (AUC 0.929/0.903/0.889) but are omitted to avoid stacking redundant correlated reclassification rules; pct_gates_50dBZ does not discriminate (AUC 0.601). Caveat recorded in the codebook changelog: these coverage fractions largely measure echo extent rather than linearity, so they are expected to misfire on large non-linear convective clusters. The codebook describes classes in reflectivity and topology rather than colour, because COLOR_DBZ_RANGE in lars/nepho/inference.py is calibrated to a ChaseSpectral-like ramp that does not match the NWSRef bands these scenes render with. Note for reviewers: the published split carries a prior shift (train 54% Cellular, test 66% QLCS), so report balanced accuracy and per-class recall rather than raw accuracy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The labeller has only ever been scored against in-house CSAPR2 hand labels, so there is no evidence it generalises beyond Bankhead National Forest, and no way to compare it to published human annotation.
SVRIMG (Haberlie, Ashley & Karpinski 2020) supplies 1,741 hand-classified GridRad/NEXRAD scenes centred on SPC severe reports. Collapsed to the two convective mode classes, Cellular vs QLCS, that is 1,315 scenes with a published split:
What
lars/preprocessing/svrimg.py—fetch_svrimg_labels/preprocess_svrimg_data, mirroringpreprocess_radar_data's CSV schema so the existing inference, Dawid-Skene and kappa tooling runs unchanged. Addsunidandsplit.example_codebooks/CODEBOOK_SVRIMG.md— Cellular / QLCS / Ambiguous, same ten-section layout the parsers expect.notebooks/reports/calibrate_svrimg_thresholds.py+ results CSV.tests/test_svrimg.py— 15 tests, no network.lars/preprocessing/__init__.pyandSTANDARD_LABEL_MAP.Two format details worth knowing
The PNGs are data, not pictures. They are palette-mode images whose palette is a dummy linear ramp (
i -> (3i, 3i+1, 3i+2)), so they render greyscale in any ordinary viewer. The palette indices carry the values.read_svrimg_pngreads the indices and refuses a non-palette file, so an accidentally RGB-converted image fails loudly rather than producing plausible-looking nonsense.Those indices are dBZ, integer-truncated and floored at zero — confirmed against the upstream GridRad file, where
REFC_MAXis itselfuint8withunits: dBZ. There are no negative values, sovmin=0, not the-20used for CSAPR2.Thresholds were recalibrated, not inherited
CODEBOOK_NWSRef.md's area fractions were tuned for 1 km CSAPR2 on a 2048×2048 polar grid; they do not transfer to 3.75 km column-max on a 136×136 Cartesian grid. Calibrated on the training split only:Only the strongest is adopted; the others are correlated and would stack redundant reclassification triggers.
Reviewer notes
COLOR_DBZ_RANGEinlars/nepho/inference.pyis calibrated to a ChaseSpectral-like ramp (green 10–30, yellow 30–40, red 40–50) that does not match the NWSRef bands these scenes render with. Worth a separate look at whetherCODEBOOK_NWSRef.mdhas the same mismatch.preprocess_radar_data'swhere(field > min_ref). Without this they paint as the bottom of the ramp and fill the frame with apparent echo.Testing
pytest tests/ --ignore=tests/integration→ 116 passed on this branch, verified in a clean worktree offmainso none of the in-flightclass_scoreswork is mixed in.Unrelated pre-existing issue: a full
pytest tests/run fails during collection becausetests/integration/conftest.pyglobally evicts thepip_system_certsmock fromsys.modules, breaking every top-level test file collected after it. Reproduces onmainwithout this branch; not addressed here.Not done: end-to-end labelling of the 230-image test split with a live model needs API credentials, so κ against SVRIMG ground truth is still to be measured.
🤖 Generated with Claude Code