Skip to content

Should the true/false aggregators get a strict-empty variant, like the float ones? #2883

Description

@feiiiiii5

The float family has a way to refuse the "no data" default; the true/false family does not, and its default reads as a verdict.

ScoreAggregatorResult reserves None for exactly one meaning:

None means no verdict was reachable, because at least one constituent score was undetermined and the aggregation could not settle without it.

But with no constituents at all, TrueFalseScoreAggregator.OR (and AND, MAJORITY) return a completed False:

# true_false_score_aggregator.py
if not scores_list:
    # No scores; return a neutral result
    return ScoreAggregatorResult(
        value=False,
        description=f"No scores provided to {name} composite scorer.",
        ...

False is not neutral here — it is the same value OR returns for a genuine refutation, so a caller that filtered to zero applicable children gets Score(score_value="false") with status=COMPLETE, i.e. "this prompt is not harmful" rather than "nothing was evaluated". _and makes the same choice deliberately for the one-observation case ("One definite False settles an AND regardless of what else could not be observed"), which is a different question from having no observations at all.

The float family solved exactly this: _create_aggregator takes raise_on_empty: bool = False, and it ships AVERAGE_RAISE_ON_EMPTY, MAX_RAISE_ON_EMPTY and MIN_RAISE_ON_EMPTY so a user can refuse the "0.0 from nothing" default. tests/unit/score/test_float_scale_score_aggregator.py covers them.

Question

Should the true/false family get the same opt-in variants (OR_RAISE_ON_EMPTY and friends), or is False on an empty aggregate meant to stand?

I lean towards adding them, for the same reason the float variants exist: the default stays convenient, and a caller who needs "no evidence" to be distinguishable from "no harm" gets a way to say so. I am not proposing to change the default, since the existing tests pin it and it may well be deliberate for the composite scorers that call it.

If the answer is no, I would also suggest the description say so more loudly than "return a neutral result", since value=False is what the rest of the pipeline reads.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions