Skip to content

PartialDependence crashes with ValueError on datasets with fewer than 10 rows #681

Description

@aniruddhaadak80

Summary

interpret.blackbox.PartialDependence raises ValueError for any dataset with fewer than 10 rows:

ValueError: Cannot take a larger sample than population when 'replace=False'

Steps to reproduce

import numpy as np
from sklearn.linear_model import LinearRegression
from interpret.blackbox import PartialDependence

rng = np.random.default_rng(0)
X = rng.normal(size=(3, 2))   # any n < 10 triggers this
y = X[:, 0] * 2.0 - X[:, 1]

model = LinearRegression().fit(X, y)
PartialDependence(model, X, feature_types=["continuous", "continuous"])

Traceback

File ".../interpret/blackbox/_partialdependence.py", line 119, in __init__
    pdp = _gen_pdp(
File ".../interpret/blackbox/_partialdependence.py", line 55, in _gen_pdp
    np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), :
ValueError: Cannot take a larger sample than population when 'replace=False'

Root cause

_gen_pdp in python/interpret-core/interpret/blackbox/_partialdependence.py builds an
individual conditional expectation (ICE) line for every row of the input data, then keeps a
random subsample of them as background_scores for plotting:

ice_lines = ice_lines[
    np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), :
]

num_ice_samples defaults to 10 and is not exposed on the PartialDependence constructor.
The subsample is drawn with replace=False, so NumPy refuses to draw more items than the
population size. Any dataset with fewer than num_ice_samples rows therefore raises, and the
underlying data size is never checked or surfaced to the caller.

Suggested fix

Clamp the requested sample count to the available number of rows before sampling, so datasets
smaller than num_ice_samples produce one background line per available row instead of
crashing:

num_ice_samples = min(num_ice_samples, ice_lines.shape[0])

Downstream consumers only iterate background_scores by row
(interpret/visual/plot.py loops for i in range(background_lines.shape[0])), so a smaller
subsample renders correctly without any other change.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions