Skip to content

Reduce pinned allocation rounding for fitted caches - #1057

Draft
autumnust wants to merge 3 commits into
mainfrom
codex/pinned-cache-allocation-guide
Draft

autumnust wants to merge 3 commits into
mainfrom
codex/pinned-cache-allocation-guide

Conversation

@autumnust

@autumnust autumnust commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

CUDA fits with multiple estimators can allocate substantially more pinned host memory than their cache tensors require. Automatically apply pinned_max_round_threshold_mb:1 immediately before offloading fitted state on PyTorch 2.13 or later.

Preserve the allocator's current configuration and any explicit rounding threshold, including settings applied after CUDA initialization. Keep freed blocks available for reuse. Older PyTorch versions retain their existing behavior. The setting applies process-wide when cached fitting first needs pinned storage. Update the inference guide with the automatic behavior and caller overrides.

Related evidence: TabArena PR #644 uses pinned_max_cached_size_mb:64 and avoids full-cache serialization copies. Its reported results support reducing allocation waste. That combined retention and serialization change does not isolate this rounding-only optimization.

Validation outside CI: 43 focused checks passed, including fresh-process CUDA fits, environment and runtime overrides, and simulated older-version compatibility. In two protected KumoTabular Large runs on the full airfoil_self_noise benchmark split with sixteen estimators, automatic configuration reduced pinned allocator storage from 384 to 308 MiB. All 1,002 validation and 501 test predictions matched exactly. This single pair does not establish stable timing gains. Full pre-commit checks and the documentation build passed.

@copy-pr-bot

copy-pr-bot Bot commented Oct 7, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@autumnust autumnust changed the title Document pinned allocation tuning for fitted caches Reduce pinned allocation rounding for fitted caches Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant