Skip to content

fix(sklearn-metrics): validate priv_group matches at least one sample - #580

Open
aniruddhaadak80 wants to merge 1 commit into
Trusted-AI:mainfrom
aniruddhaadak80:fix/sklearn-metrics-priv-group-validation
Open

aniruddhaadak80 wants to merge 1 commit into
Trusted-AI:mainfrom
aniruddhaadak80:fix/sklearn-metrics-priv-group-validation

Conversation

@aniruddhaadak80

@aniruddhaadak80 aniruddhaadak80 commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

difference() and ratio() in aif360/sklearn/metrics/metrics.py never checked that priv_group actually matches at least one sample, so a priv_group outside the observed groups silently produced a wrong fairness value instead of an error. This affects statistical_parity_difference, equal_opportunity_difference, disparate_impact_ratio and the other metrics built on those two helpers.

Root cause

Both helpers compute idx = (groups == priv_group) and then slice the privileged and unprivileged arrays with it. When nothing matches, idx is all-False, the privileged slice is empty, and the metric is evaluated over an empty array. np.mean([]) is nan and an empty ratio is 0.0, so statistical_parity_difference returns nan and disparate_impact_ratio returns 0.0 -- both of which look like plausible fairness numbers and will happily be reported or asserted on.

The most common way to reach this is the default priv_group=1 with more than one protected attribute. check_groups then returns tuples such as (1, 0), and (1, 0) == 1 is always False, so every call on a multi-attribute dataset silently yields nan/0.0 with no indication that priv_group needs to be a tuple.

Changes

  • aif360/sklearn/metrics/metrics.py - in both difference() and ratio(), raise ValueError when (groups == priv_group).any() is False. The message includes the value that was passed and the groups that are actually available, so the caller can see what to pass instead.
  • tests/sklearn/test_metrics.py - three regression tests:
    • test_priv_group_no_match_raises: an out-of-range priv_group raises for both metrics;
    • test_priv_group_default_with_multiple_prot_attrs_raises: the default priv_group=1 with prot_attr=['sex', 'age'] raises for both;
    • test_priv_group_valid_still_works: a valid priv_group still returns a float, guarding against over-eager validation.

Testing

python -m pytest tests/sklearn/test_metrics.py does not run in my environment: that module constructs AdultDataset(...) at module level, so collection aborts offline with SystemExit: 1 before any test is collected. That is pre-existing and unrelated to this change, but it does mean I could not run the file itself.

I verified the change with a standalone script driving the public API, covering the same three cases as the new tests:

  • priv_group=999 on sex in {0, 1} -> both metrics raise ValueError: priv_group=999 does not match any sample in the protected attribute(s). Available groups: [1, 0]
  • default priv_group=1 with prot_attr=['sex', 'age'] -> both raise ValueError: ... Available groups: [(1, 0), (1, 1), (0, 0), (0, 1)]
  • valid priv_group=1 still returns floats (-1.0 for the difference, 0.0 for the ratio)

I also confirmed the pre-fix behaviour by replaying the original unvalidated path against the same inputs: idx.any() is False, the privileged slice has length 0, and np.mean of it is nan. That is the silently wrong value this change now turns into an actionable error.

Fixes #578

difference() and ratio() did not check that priv_group matches any sample
in the protected attribute(s). With multiple protected attributes, the
default priv_group=1 never matches any intersectional group tuple,
causing the privileged subset to be empty and the metric to silently
return NaN or a wrong value.

Add validation that raises a clear ValueError when priv_group does not
match any sample, listing the available groups to help users identify
the correct value.

Fixes Trusted-AI#578

Signed-off-by: Aniruddha Adak <aniruddhaadak80@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant