Skip to content

data: fix PartialDependence crash on datasets with fewer than 10 rows - #683

Open
aniruddhaadak80 wants to merge 1 commit into
interpretml:mainfrom
aniruddhaadak80:fix/pdp-ice-samples-clamp
Open

aniruddhaadak80 wants to merge 1 commit into
interpretml:mainfrom
aniruddhaadak80:fix/pdp-ice-samples-clamp

Conversation

@aniruddhaadak80

Copy link
Copy Markdown

Summary

interpret.blackbox.PartialDependence raised ValueError: Cannot take a larger sample than population when 'replace=False' for any dataset with fewer than 10 rows, because the ICE background subsample was drawn without replacement using a hardcoded default of 10. This clamps the requested sample count to the number of available rows.

Root cause

In python/interpret-core/interpret/blackbox/_partialdependence.py, _gen_pdp computes one individual conditional expectation line per input row, then keeps a random subsample of them as the background_scores series used for plotting:

num_ice_samples=10,   # default, not exposed on the PartialDependence constructor
...
ice_lines = ice_lines[
    np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), :
]

num_ice_samples defaults to 10 and _gen_pdp is only called from PartialDependence.__init__ (line 119), which does not thread this parameter through, so every caller gets 10. Because the draw uses replace=False, NumPy refuses to select more items than the population size, so any dataset with fewer than 10 rows raised. Nothing in __init__ validated or surfaced the data size, and the failure surfaced as a raw NumPy error from deep inside the explainer.

The only consumer of background_scores is interpret/visual/plot.py:390, which iterates it by row (for i in range(background_lines.shape[0])), so returning fewer lines than requested is handled correctly and needs no downstream change.

Changes

  • python/interpret-core/interpret/blackbox/_partialdependence.py — clamp num_ice_samples to ice_lines.shape[0] before the np.random.choice draw in _gen_pdp. Datasets at or above the threshold are unaffected because min() resolves back to the original value.
  • python/interpret-core/tests/blackbox/test_partialdependence.py — new regression tests: a parametrized case over 1/2/5/9 rows asserting the crash is gone and that exactly one background line per row is returned, an end-to-end PartialDependence + explain_global() construction on a 3-row dataset, and a 25-row case asserting the existing cap of num_ice_samples still holds.

Testing

Run from python/interpret-core with PYTHONPATH pointed at that directory.

Before the fix — new tests fail, existing behaviour for large data already passes:

$ python -m pytest tests/blackbox/test_partialdependence.py -v
collected 6 items

tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[1] FAILED [ 16%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[2] FAILED [ 33%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[5] FAILED [ 50%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[9] FAILED [ 66%]
tests/blackbox/test_partialdependence.py::test_partial_dependence_small_dataset_does_not_crash FAILED [ 83%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_enough_rows_still_caps_at_num_ice_samples PASSED [100%]

_ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _
interpret\blackbox\_partialdependence.py:55: in _gen_pdp
    np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), :
E   ValueError: Cannot take a larger sample than population when 'replace=False'

================== 5 failed, 1 passed in 141.23s (0:02:21) ===================

After the fix — all 6 pass:

$ python -m pytest tests/blackbox/test_partialdependence.py -v
collected 6 items

tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[1] PASSED [ 16%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[2] PASSED [ 33%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[5] PASSED [ 50%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_fewer_rows_than_ice_samples[9] PASSED [ 66%]
tests/blackbox/test_partialdependence.py::test_partial_dependence_small_dataset_does_not_crash PASSED [ 83%]
tests/blackbox/test_partialdependence.py::test_gen_pdp_enough_rows_still_caps_at_num_ice_samples PASSED [100%]

======================== 6 passed in 120.40s (0:02:00) ========================

Existing blackbox test suite (including the pre-existing test_sensitivity.py) is green:

$ python -m pytest tests/blackbox/
collected 7 items

tests\blackbox\test_partialdependence.py ......                          [ 85%]
tests\blackbox\test_sensitivity.py .                                     [100%]

======================== 7 passed in 210.39s (0:03:30) =========================

Formatting matches the ruff-format hook configured in .pre-commit-config.yaml:

$ python -m ruff format --check python/interpret-core/interpret/blackbox/_partialdependence.py python/interpret-core/tests/blackbox/test_partialdependence.py
2 files already formatted

ruff check is clean on the new test file. On the modified _partialdependence.py it reports I001 (import sorting), NPY002 (legacy np.random.choice) and PLC0415 (deferred import) — all three are pre-existing on main and unchanged by this PR, and this repo's pre-commit only runs ruff-format, so they are left alone.

Fixes #681

…etml#681)

_gen_pdp built an individual conditional expectation line for every
row of the input data and then kept a random subsample of them as
"background_scores" for plotting:

    ice_lines = ice_lines[
        np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), :
    ]

num_ice_samples defaults to 10 and is not exposed on the
PartialDependence constructor. Because the subsample is drawn with
replace=False, NumPy refuses to draw more items than the population
size, so every dataset with fewer than num_ice_samples rows raised
"ValueError: Cannot take a larger sample than population when
'replace=False'" instead of producing an explanation. The underlying
data size was never checked or surfaced to the caller.

Clamp the requested count to the number of available rows so smaller
datasets yield one background line per row rather than crashing.
Datasets at or above num_ice_samples are unaffected, since the min()
resolves back to the original value.

The only consumer of background_scores is interpret/visual/plot.py,
which iterates it by row ("for i in range(background_lines.shape[0])"),
so a smaller subsample renders correctly with no other change.

Add tests/blackbox/test_partialdependence.py with a parametrized case
over 1/2/5/9 rows, an end-to-end PartialDependence construction on a
3-row dataset, and a case asserting that a 25-row dataset still keeps
exactly num_ice_samples background lines.

Fixes interpretml#681

Signed-off-by: Aniruddha Adak <aniruddhaadak80@users.noreply.github.com>
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 22.23%. Comparing base (560f8dd) to head (595c705).

❗ There is a different number of reports uploaded between BASE (560f8dd) and HEAD (595c705). Click for more details.

HEAD has 791 uploads less than BASE
Flag BASE (560f8dd) HEAD (595c705)
sdist_linuxarm_311_python 44 0
sdist_linuxarm_313_python 44 1
sdist_linuxarm_314_python 44 1
sdist_linuxarm_312_python 44 1
sdist_mac_314_python 39 1
sdist_mac_311_python 11 1
sdist_mac_313_python 43 1
sdist_mac_312_python 38 1
sdist_linux_312_python 43 1
sdist_linux_314_python 40 1
sdist_linux_313_python 43 1
sdist_linux_311_python 44 1
bdist_mac_313_python 17 1
sdist_win_314_python 7 0
bdist_mac_311_python 8 1
bdist_linuxarm_314_python 31 1
sdist_win_313_python 7 0
bdist_linuxarm_311_python 30 0
sdist_win_311_python 7 0
bdist_linuxarm_313_python 32 0
bdist_mac_312_python 19 1
bdist_mac_314_python 21 0
bdist_win_311_python 10 0
bdist_linuxarm_312_python 34 1
bdist_linux_311_python 20 0
bdist_linux_314_python 18 0
sdist_win_312_python 7 0
bdist_linux_313_python 18 0
bdist_linux_312_python 21 1
bdist_win_314_python 8 0
bdist_win_313_python 10 0
bdist_win_312_python 6 0
Additional details and impacted files
@@             Coverage Diff             @@
##             main     #683       +/-   ##
===========================================
- Coverage   67.21%   22.23%   -44.99%     
===========================================
  Files          77       77               
  Lines       11735    11736        +1     
===========================================
- Hits         7888     2609     -5279     
- Misses       3847     9127     +5280     
Flag Coverage Δ
bdist_linux_311_python ?
bdist_linux_312_python 17.02% <100.00%> (-49.96%) ⬇️
bdist_linux_313_python ?
bdist_linux_314_python ?
bdist_linuxarm_311_python ?
bdist_linuxarm_312_python 21.03% <100.00%> (-45.95%) ⬇️
bdist_linuxarm_313_python ?
bdist_linuxarm_314_python 21.96% <100.00%> (-44.93%) ⬇️
bdist_mac_311_python 19.81% <100.00%> (-47.32%) ⬇️
bdist_mac_312_python 22.15% <100.00%> (-44.98%) ⬇️
bdist_mac_313_python 22.15% <100.00%> (-44.98%) ⬇️
bdist_mac_314_python ?
bdist_win_311_python ?
bdist_win_312_python ?
bdist_win_313_python ?
bdist_win_314_python ?
sdist_linux_311_python 16.97% <100.00%> (-49.96%) ⬇️
sdist_linux_312_python 16.97% <100.00%> (-49.96%) ⬇️
sdist_linux_313_python 16.97% <100.00%> (-49.96%) ⬇️
sdist_linux_314_python 16.75% <100.00%> (-50.08%) ⬇️
sdist_linuxarm_311_python ?
sdist_linuxarm_312_python 19.85% <100.00%> (-47.08%) ⬇️
sdist_linuxarm_313_python 22.11% <100.00%> (-44.82%) ⬇️
sdist_linuxarm_314_python 21.02% <100.00%> (-45.81%) ⬇️
sdist_mac_311_python 18.63% <0.00%> (-48.42%) ⬇️
sdist_mac_312_python 22.06% <100.00%> (-44.98%) ⬇️
sdist_mac_313_python 18.26% <0.00%> (-48.79%) ⬇️
sdist_mac_314_python 21.86% <100.00%> (-45.09%) ⬇️
sdist_win_311_python ?
sdist_win_312_python ?
sdist_win_313_python ?
sdist_win_314_python ?

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@aniruddhaadak80

Copy link
Copy Markdown
Author

Four checks are red here, and none of them are in this diff - they all fail while importing the environment, before the clamped-sample code is reached. Recording the diagnosis so the red is not mistaken for the change.

1. testB / testS on mac_312 - the root failure (everything else is cancelled fail-fast)

ext/test_examples.py:13: in <module>
    from interpret.glassbox import LinearRegression
/Library/Frameworks/Python.framework/Versions/3.12/.../interpret/glassbox/_aplr.py:54: in <module>
    class APLRRegressor(
<frozen abc>:106: in __new__
E   TypeError: Cannot create a consistent method resolution
    order (MRO) for bases RegressorMixin, LocalExplainer, GlobalExplainer, BaseEstimator, APLRRegressor

Note the path: this is the installed wheel in site-packages, not the working tree. The import of interpret itself raises on that one interpreter/OS combination, so pytest fails at collection - 90 collection errors across test_selenium.py, utils/test_clean_x.py, visual/test_interactive.py and others, with 95 passed, 5 skipped, 90 errors. Every other job reports The strategy configuration was canceled because "testB.mac_312_python_3_12_macos" failed, i.e. they are collateral, not independent.

2. test_powerlift on linux_312

ImportError while loading conftest 'tests/conftest.py'.
E   AttributeError: module 'sqlalchemy.orm.attributes' has no attribute 'ScalarAttributeImpl'. Did you mean: '_ScalarAttributeImpl'?

SQLAlchemy 2.x removed the public ScalarAttributeImpl name; anything importing it needs the private _ScalarAttributeImpl or a pinned SQLAlchemy. Unrelated to the PDP change.

3. docs

mystnb fails to execute the example notebooks (custom-interactions, differential-privacy, group-importances, interpretable-classification, interpretable-regression, interpretable-regression-synthetic, merge-ebms, quantile-regression - CellExecutionError), giving build finished with problems, 48 warnings and exit 78. The job then also cannot report its own result on a fork PR: RequestError [HttpError]: Resource not accessible by integration - .../checks/runs#create-a-check-run, because the fork token may not create check runs in the base repo.

4. Everything else - cancelled by fail-fast, not independently broken.

Given the change itself is a two-line clamp plus tests, and the 40 passing checks cover the touched path, I would rather ask than paper over this: is there a preferred way to get the mac_312 / SQLAlchemy-2.x environment repaired? I have not touched the workflow or pinned anything, since that is a maintainer call and would widen this PR well beyond its purpose.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

PartialDependence crashes with ValueError on datasets with fewer than 10 rows

1 participant