Skip to content

Add functionality from radar palette including an new Gate ID function - #138

Open
scollis wants to merge 9 commits into
ARM-Development:mainfrom
scollis:gateid-backend
Open

scollis wants to merge 9 commits into
ARM-Development:mainfrom
scollis:gateid-backend

Conversation

@scollis

@scollis scollis commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

New Gate ID functionality. Mainly based on Park, Ryzhkov, Zrnić and Kim, 2009, Wea. Forecasting 24, 730–748 (doi:10.1175/2008WAF2222205.1) for the fuzzy logic classifier and Ryzhkov et al., 2017, J. Appl. Meteor. Climatol. 56, 1797–1816 (doi:10.1175/JAMC-D-16-0098.1) for the circular depolarization ratio.

Adds a new biological classification and also adds a geometric constraint to second trips. The ID creates multiple IDs which could be valid returns so collapses these down into the standard CMAC gate IDs of rain, melt, ice, no scatter and second trip.

image image

Example yaml for configuration

  bnf_csapr2_ppi:
    gate_id_method: radar_palette     # default 'cmac_fuzzy' for every radar
    gate_id_class_map: {graupel_hail: rain}   # partial maps merge over the default
    gate_id_snr_min: 3.0
    gate_id_incoherent_frac: 0.35
    gate_id_apply_offsets: true
    gate_id_publish_margin: false```

This was co-written by Claude. Some testing will be needed. 

CMAC classifies the dominant scatterer with its own five-class fuzzy
scheme (cmac_processing.do_my_fuzz). radar_palette.gateid is a second
classifier over the same moments with eleven classes, its own noise-floor,
second-trip and despeckle logic, and a documented evaluation across BNF
C-SAPR2, SGP C-SAPR and NEXRAD volumes. This adds it as an alternative
backend so the two can be compared on real data, and so a site can pick
the one that behaves better on its radar.

The new module classifies with radar_palette and folds the result back
onto CMAC's five categories. It does not widen the gate_id code space,
because that space is a contract: everything downstream reads the
category vocabulary out of the field's notes attribute,
gate_id.get_gate_id_categories derives each code from its *position* in
that string, and the Z-PHI phase-processing gate filter reads codes 1
and 2 as bare literals. Widening or reordering the codes would silently
repoint those literals at different classes. The eleven-class result is
published in full alongside the folded field as
scatterer_classification, with CF flag_values/flag_meanings, so nothing
is lost and nothing downstream changes meaning.

Two integration details are handled deliberately rather than inherited.

First, the classification input is built from CMAC's own per-radar
field_names configuration rather than letting radar_palette resolve
field names itself. Its resolver walks an ordered preference list that
puts the uncorrected_* moments first, and on an ARM a1 volume those sit
alongside the corrected ones, so it would score gates on reflectivity
and ZDR that have not had the per-radar ref_offset and zdr_offset
applied while every other CMAC product uses the offset fields. Handing
it a view carrying exactly one candidate per moment removes the
ambiguity and keeps field naming in the config, where CMAC keeps it.

Second, the backend requires at least one finite measured moment at a
gate before a hydrometeor label survives, and forces no_scatter
elsewhere. radar_palette's own evidence test asks whether a feature key
is present, not whether its value at a gate is finite; gate temperature
and height over the freezing level are derived from gate geometry and
are finite everywhere, so an empty-sky gate can score above zero on
temperature alone. Measured on real volumes that reached 1,073,535
labelled empty gates on one BNF C-SAPR2 PPI and 8.4 million on a NEXRAD
surveillance volume. In CMAC that would not be cosmetic: gate_id is the
single authoritative mask behind velocity dealiasing, the Z-PHI
attenuation and KDP path, and all three rain-rate estimators, so a label
on an unmeasured gate would propagate into a rainfall total. The count
of gates demoted is reported as n_no_evidence.

The freezing level the melting-layer constraints want is derived from
the mapped sounding before classification, since get_melt cannot run
until a gate_id field exists. radar_palette is imported only when a
configuration selects this backend, so it remains an optional
dependency, and a missing install raises an error that names the fix.

Tests build their radars with pyart.testing and run offline; the ones
that need the classifier skip when radar_palette is absent.
Adds gate_id_method to the per-radar processing tunables and dispatches on
it in cmac(). It defaults to 'cmac_fuzzy' for every radar, so behaviour is
unchanged unless a configuration asks for the radar_palette backend, and
an unrecognised value raises rather than silently falling through to one
of the two.

The remaining gate_id_* keys tune the radar_palette backend only. Each is
left None where the intent is "take that classifier's own default", which
is not the same as passing None through to it: gate_id_incoherent_frac in
particular switches the phase-coherence test off entirely when passed as
None, so None is resolved to the classifier's constant and 0 is the
documented way to disable the test. gate_id_incoherent_frac is exposed
per radar rather than fixed because it was measured not to transfer
between instruments: the fraction of precipitation gates it rejects runs
from under 1% on weak-echo C-band RHIs to 24.6% on an S-band NEXRAD
volume.

Which classifier produced gate_id is not recoverable from the field once
the categories are folded, and every masked product in the file depends
on it, so cmac() records gate_id_method, the temperature source and the
count of gates demoted by the evidence guard as file metadata.
Adds the gate_id_method key and the radar_palette tuning keys to the YAML
configuration reference, spells out the five-category notes contract both
backends honour and the eleven-to-five fold between them, and records the
two integration behaviours a user would otherwise have to read the source
to discover: that the classification input is built from the radar's own
field_names rather than the classifier's preference order, and that a gate
with no finite measured moment cannot carry a hydrometeor label.
get_gate_id_categories derives a category's integer code from where its
label sits in the field's notes string, not from the integer printed next
to it. cmac() appended ',5:clutter' once for a ground_clutter field and
again for a vendor classification_mask, so a volume carrying both -- which
is every BNF, CACTI and TRACER C-SAPR2 a1 volume -- ended up with clutter
listed twice. Anything reading the categories back then resolved clutter
to 6 while the gates had been set to 5, and a terrain_blockage category
appended afterwards landed on 7 with nothing at that code.

Both categories are now appended through gate_id.append_gate_id_category,
which is idempotent, keeps valid_min/valid_max consistent, and returns the
code to write so the caller does not hard-code it -- hard-coding is what
allowed the notes string and the gate values to disagree.

Also makes the classification_mask overlay opt-out per radar. Its
semantics differ across instrument generations: on TRACER C-SAPR2 a1
volumes the mask equals 8, its clutter bit alone, at 171,560 of 171,600
gates, so the overlay relabels the entire volume as clutter and no
meteorological gate survives into the dealiasing, KDP or rain-rate steps.
ARM's production CMAC output for the same volume carries no clutter code
at all, so that generation of data was never processed with this overlay.
On CACTI C-SAPR2 the same field is 0 at 94% of gates and the overlay
behaves as intended, so use_classification_mask defaults to True and
nothing changes for the sites where it works.
Deciding what a gate measured is a judgement about the measurement, so the
radar_palette backend now takes each moment from the uncorrected field
where the volume publishes one and falls back to the radar's field_names
entry otherwise.

The unprefixed fields on an ARM a1 volume are not raw: they carry vendor
clutter filtering and thresholding, and on the CACTI and TRACER
generations the volume also ships attenuation_corrected_* variants, so
what "reflectivity" has already been through differs by site and by
instrument generation. The uncorrected moments are the same measurement
everywhere. Every ARM C-SAPR2 a1 volume examined -- BNF 2026, CACTI 2018,
TRACER 2022 -- publishes uncorrected reflectivity, ZDR, RhoHV, PhiDP,
Doppler velocity and spectrum width; BNF additionally publishes
uncorrected NCP and SNR. Legacy MDV-era X-SAPR and NEXRAD volumes publish
none of them, which is why every moment keeps a documented fallback to the
configured field rather than becoming a hard requirement.

Exactly one candidate per logical moment is still handed to the
classifier, so radar_palette's own preference order cannot resolve to a
field other than the intended one; what changed is which field that is.
The resolved source per moment is returned, printed under verbose, and
recorded in the classifier metadata, so a run says what it classified on
rather than leaving it to be inferred.

cmac() applies ref_offset and zdr_offset in place to the *configured*
reflectivity and ZDR before classification, and an uncorrected moment has
not been through that, so the backend adds the offsets to its own view of
them -- a calibration constant describes the instrument, not a correction
to the measurement, and without it the classifier would score an
uncalibrated radar against absolute dBZ and dB thresholds. A moment that
fell back to the configured field is not offset again. Set
gate_id_apply_offsets false to classify on raw values.
The melting layer is a valid radar return and belongs in the corrected
reflectivity, but the gate filter handed to calculate_attenuation_zphi
admitted only gate_id 1 and 2, i.e. rain and snow, so melting-layer gates
were excluded from the correction and from corrected_reflectivity,
specific_attenuation and the KDP path that follow it. The velocity
dealiasing filter a few lines earlier already includes rain, melting and
snow, so the two filters disagreed about what counts as a return.

Those two codes were also bare literals: they are correct only for the
fuzzy classifier's category order, and any backend or configuration that
ordered the categories differently would have had the filter silently
select the wrong classes. The codes now come from the field's own notes
string through get_gate_id_categories, and a category absent from a
particular run is skipped rather than assumed.

This changes the default cmac_fuzzy output: corrected_reflectivity and
everything derived from it now cover melting-layer gates, which is the
intended behaviour and not a byte-identical change.
Every ARM volume publishes several candidates per moment -- uncorrected,
vendor-processed, and on some generations attenuation-corrected -- so which
one the classifier saw is not recoverable from the output file. cmac() now
writes it alongside gate_id_method and the temperature source.
radar_palette skips sweeps it judges unsuitable and returns 'unclassified'
for every gate in them. Folded onto CMAC's five categories that becomes
no_scatter, which is indistinguishable from empty sky: the VAP would write
a file whose corrected velocity, KDP and all three rain rates are empty,
with nothing in it saying the classifier had declined.

Found by running the TRACER C-SAPR2 cell-tracking sequence of 2022-06-17:
four RHIs of 16 to 76 rays, elevation spans near 19 degrees, came back
entirely unclassified while the fuzzy classifier found 1,186 to 15,434
rain gates in the same volumes. The classifier's own metadata names the
reason -- skipped_sweeps carries the sweep index, mode and elevation
span -- so nothing had to be inferred; the backend was simply discarding
it. Tuning snr_min, texture_window, min_run or incoherent_frac does not
change the outcome, because the sweep is rejected before scoring.

The backend now warns, with the sweep numbers and elevation spans, when
any sweep is skipped or more than half the gates come back unclassified,
and writes gate_id_unclassified_gates and gate_id_skipped_sweeps into the
output file. gate_id_unclassified_policy='error' turns it into a hard
failure, which is what a production run that must not emit an empty mask
should set. The count is taken before the evidence guard and the fold,
both of which would otherwise hide it.
The documentation rounded all four skipped sweeps to one elevation span.
They are 3.7, 8.7, 13.8 and 18.8 degrees, and the fuzzy classifier labels
1,186 to 15,434 gates rain on the three wider ones, so the spread matters
to anyone judging whether their own scan strategy is affected.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant