Summary
interpret.blackbox.PartialDependence raises ValueError for any dataset with fewer than 10 rows:
ValueError: Cannot take a larger sample than population when 'replace=False'
Steps to reproduce
import numpy as np
from sklearn.linear_model import LinearRegression
from interpret.blackbox import PartialDependence
rng = np.random.default_rng(0)
X = rng.normal(size=(3, 2)) # any n < 10 triggers this
y = X[:, 0] * 2.0 - X[:, 1]
model = LinearRegression().fit(X, y)
PartialDependence(model, X, feature_types=["continuous", "continuous"])
Traceback
File ".../interpret/blackbox/_partialdependence.py", line 119, in __init__
pdp = _gen_pdp(
File ".../interpret/blackbox/_partialdependence.py", line 55, in _gen_pdp
np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), :
ValueError: Cannot take a larger sample than population when 'replace=False'
Root cause
_gen_pdp in python/interpret-core/interpret/blackbox/_partialdependence.py builds an
individual conditional expectation (ICE) line for every row of the input data, then keeps a
random subsample of them as background_scores for plotting:
ice_lines = ice_lines[
np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), :
]
num_ice_samples defaults to 10 and is not exposed on the PartialDependence constructor.
The subsample is drawn with replace=False, so NumPy refuses to draw more items than the
population size. Any dataset with fewer than num_ice_samples rows therefore raises, and the
underlying data size is never checked or surfaced to the caller.
Suggested fix
Clamp the requested sample count to the available number of rows before sampling, so datasets
smaller than num_ice_samples produce one background line per available row instead of
crashing:
num_ice_samples = min(num_ice_samples, ice_lines.shape[0])
Downstream consumers only iterate background_scores by row
(interpret/visual/plot.py loops for i in range(background_lines.shape[0])), so a smaller
subsample renders correctly without any other change.
Summary
interpret.blackbox.PartialDependenceraisesValueErrorfor any dataset with fewer than 10 rows:Steps to reproduce
Traceback
Root cause
_gen_pdpinpython/interpret-core/interpret/blackbox/_partialdependence.pybuilds anindividual conditional expectation (ICE) line for every row of the input data, then keeps a
random subsample of them as
background_scoresfor plotting:num_ice_samplesdefaults to10and is not exposed on thePartialDependenceconstructor.The subsample is drawn with
replace=False, so NumPy refuses to draw more items than thepopulation size. Any dataset with fewer than
num_ice_samplesrows therefore raises, and theunderlying data size is never checked or surfaced to the caller.
Suggested fix
Clamp the requested sample count to the available number of rows before sampling, so datasets
smaller than
num_ice_samplesproduce one background line per available row instead ofcrashing:
Downstream consumers only iterate
background_scoresby row(
interpret/visual/plot.pyloopsfor i in range(background_lines.shape[0])), so a smallersubsample renders correctly without any other change.