data: fix PartialDependence crash on datasets with fewer than 10 rows - #683
aniruddhaadak80 wants to merge 1 commit into
Conversation
…etml#681) _gen_pdp built an individual conditional expectation line for every row of the input data and then kept a random subsample of them as "background_scores" for plotting: ice_lines = ice_lines[ np.random.choice(ice_lines.shape[0], num_ice_samples, replace=False), : ] num_ice_samples defaults to 10 and is not exposed on the PartialDependence constructor. Because the subsample is drawn with replace=False, NumPy refuses to draw more items than the population size, so every dataset with fewer than num_ice_samples rows raised "ValueError: Cannot take a larger sample than population when 'replace=False'" instead of producing an explanation. The underlying data size was never checked or surfaced to the caller. Clamp the requested count to the number of available rows so smaller datasets yield one background line per row rather than crashing. Datasets at or above num_ice_samples are unaffected, since the min() resolves back to the original value. The only consumer of background_scores is interpret/visual/plot.py, which iterates it by row ("for i in range(background_lines.shape[0])"), so a smaller subsample renders correctly with no other change. Add tests/blackbox/test_partialdependence.py with a parametrized case over 1/2/5/9 rows, an end-to-end PartialDependence construction on a 3-row dataset, and a case asserting that a 25-row dataset still keeps exactly num_ice_samples background lines. Fixes interpretml#681 Signed-off-by: Aniruddha Adak <aniruddhaadak80@users.noreply.github.com>
Codecov Report✅ All modified and coverable lines are covered by tests.
Additional details and impacted files@@ Coverage Diff @@
## main #683 +/- ##
===========================================
- Coverage 67.21% 22.23% -44.99%
===========================================
Files 77 77
Lines 11735 11736 +1
===========================================
- Hits 7888 2609 -5279
- Misses 3847 9127 +5280 Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
Four checks are red here, and none of them are in this diff - they all fail while importing the environment, before the clamped-sample code is reached. Recording the diagnosis so the red is not mistaken for the change. 1. Note the path: this is the installed wheel in 2. SQLAlchemy 2.x removed the public 3.
4. Everything else - Given the change itself is a two-line clamp plus tests, and the 40 passing checks cover the touched path, I would rather ask than paper over this: is there a preferred way to get the |
Summary
interpret.blackbox.PartialDependenceraisedValueError: Cannot take a larger sample than population when 'replace=False'for any dataset with fewer than 10 rows, because the ICE background subsample was drawn without replacement using a hardcoded default of 10. This clamps the requested sample count to the number of available rows.Root cause
In
python/interpret-core/interpret/blackbox/_partialdependence.py,_gen_pdpcomputes one individual conditional expectation line per input row, then keeps a random subsample of them as thebackground_scoresseries used for plotting:num_ice_samplesdefaults to10and_gen_pdpis only called fromPartialDependence.__init__(line 119), which does not thread this parameter through, so every caller gets 10. Because the draw usesreplace=False, NumPy refuses to select more items than the population size, so any dataset with fewer than 10 rows raised. Nothing in__init__validated or surfaced the data size, and the failure surfaced as a raw NumPy error from deep inside the explainer.The only consumer of
background_scoresisinterpret/visual/plot.py:390, which iterates it by row (for i in range(background_lines.shape[0])), so returning fewer lines than requested is handled correctly and needs no downstream change.Changes
python/interpret-core/interpret/blackbox/_partialdependence.py— clampnum_ice_samplestoice_lines.shape[0]before thenp.random.choicedraw in_gen_pdp. Datasets at or above the threshold are unaffected becausemin()resolves back to the original value.python/interpret-core/tests/blackbox/test_partialdependence.py— new regression tests: a parametrized case over 1/2/5/9 rows asserting the crash is gone and that exactly one background line per row is returned, an end-to-endPartialDependence+explain_global()construction on a 3-row dataset, and a 25-row case asserting the existing cap ofnum_ice_samplesstill holds.Testing
Run from
python/interpret-corewithPYTHONPATHpointed at that directory.Before the fix — new tests fail, existing behaviour for large data already passes:
After the fix — all 6 pass:
Existing blackbox test suite (including the pre-existing
test_sensitivity.py) is green:Formatting matches the
ruff-formathook configured in.pre-commit-config.yaml:ruff checkis clean on the new test file. On the modified_partialdependence.pyit reportsI001(import sorting),NPY002(legacynp.random.choice) andPLC0415(deferred import) — all three are pre-existing onmainand unchanged by this PR, and this repo's pre-commit only runsruff-format, so they are left alone.Fixes #681