TrueFalseInverterScorer and multi-scorer TrueFalseCompositeScorer results can normalize to a float that contradicts their own verdict.
FloatScaleThresholdScorer stores its raw value as original_float_value in the score metadata, and normalize_score_to_float prefers that over the verdict. The inverter flips score_value but keeps the metadata, and the composite merges the child metadata as is:
threshold : value=True normalized=1.0
inverted : value=False normalized=1.0
AND composite : value=False normalized=1.0 metadata={'original_float_value': 1.0}
This matters for Crescendo with an objective scorer like AND(threshold(harm), Inverter(refusal)). With use_score_as_feedback, the adversarial chat gets told a refused turn scored "1.00 on a scale of 0.0 to 1.0". TAP node ranking and logs use the same helper.
Fix: the inverter stores 1 - original_float_value, and a composite over more than one scorer drops it, since one child's float doesn't describe the combined verdict.
TrueFalseInverterScorerand multi-scorerTrueFalseCompositeScorerresults can normalize to a float that contradicts their own verdict.FloatScaleThresholdScorerstores its raw value asoriginal_float_valuein the score metadata, andnormalize_score_to_floatprefers that over the verdict. The inverter flipsscore_valuebut keeps the metadata, and the composite merges the child metadata as is:This matters for Crescendo with an objective scorer like
AND(threshold(harm), Inverter(refusal)). Withuse_score_as_feedback, the adversarial chat gets told a refused turn scored "1.00 on a scale of 0.0 to 1.0". TAP node ranking and logs use the same helper.Fix: the inverter stores
1 - original_float_value, and a composite over more than one scorer drops it, since one child's float doesn't describe the combined verdict.