Skip to content

Performance prototypes: partitioned-inverse solve, split/fused factorization, ordering study (experiments 0/5/6) - #107

Merged
github-actions[bot] merged 17 commits into
mainfrom
proto/partitioned-inverse-solve
Oct 9, 2026
Merged

github-actions[bot] merged 17 commits into
mainfrom
proto/partitioned-inverse-solve

Conversation

@sshin23

@sshin23 sshin23 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Bench-level prototype campaign for the performance plan (PERFORMANCE.md experiments 0, 5, 6; issues #25, #75, #82), developed against the pglib ACOPF condensed KKT systems that MadNLP produces. Everything was measured on a Quadro GV100 against cuDSS 0.8.0 on the same device, with per-trial correctness gates (factor comparison against stock, relative residual, zero failing fronts) — several intermediate states were faster and wrong, so no number below is reported without its gate.

Results (78k-bus condensed KKT, n = 674,562, Float64, median of ≥10 warm)

phase stock defaults this branch (best measured) cuDSS
solve (1 RHS) 18.0 ms 3.5 ms 3.8 ms
refactorization 167.7 ms 31.9 ms 24.3 ms
per IPM iteration 185 ms ~35.4 ms 28.1 ms

Everything is cuDSS-free (native reordering_alg = "algo4"). On AMDGPU (Radeon VII, gfx906), where cuDSS does not exist, the solve prototype runs with three mechanical shims: 20.5 → 8.05 ms.

What is in the branch

src (the only reviewable library change, deliberately tiny): subtree_max_fronts analysis tuning option — one Options field plus one count in build_schedule's eligibility pass, default off. Measured near-neutral on the GV100 (it falsified the longest-walk hypothesis: the subtree launch is throughput-bound at ~3.3k subtrees) and kept as the instrument to re-test on larger-SM devices.

bench/ (prototypes and experiment records, not library code):

  • solve_proto_78k.jl, solve_proto_all.jl — partitioned-inverse solve (per-front L11 inverses, TRSV→GEMV) with fused dependency-counter sweeps and a single-block chain kernel; the full SPD-harness table including the matrices where it regresses (lap3d_40, apache2) and the per-schedule strategy pick that implies.
  • fact_split.jl — split factorization: stock assembly + tiled syrk (32×32, disjoint writes, deterministic, bit-identical) + chunked trsm for m > 64; the routing study (regime_c_rows, width thresholds).
  • fact_fused.jl — segmented dependency-counter factorization: one mega-kernel per segment with four workgroup roles (extend-add, panel, trsm, syrk tiles), private contribution blocks, wides between segments; plus the write-through extend-add and the stream/budget knobs.
  • order_search.jl, solve_nd.jl, get_cudss_perm.jl — the ordering study.
  • amd_proto.jl, amd/Project.toml — the AMDGPU validation.
  • final_sweep.jl — solve configuration sweeps.

Findings a future implementation needs (each measured, details in the commit messages and comments below)

  1. Ordering: the automatic AMD/ND chooser picks AMD here because ordering_cost weighs depth at ~0.2% for n = 674k and scores column-etree depth, which does not predict supernodal schedule depth (cuDSS perm: 1279 column levels → 30 schedule levels; AMD: competitive column metrics → 65). Forcing ND (algo4) is worth 2× schedule depth, 36% on the solve, 11–17% on factorization. Recommendation: score supernodal schedule depth, or prefer ND on GPU backends.
  2. The update stack is incompatible with out-of-level-order execution: CB slot placement (PR layout: place contribution blocks offline over their lifetimes (issue #48) #54) assumes level-synchronous lifetimes; counters/streams clobber live CBs silently. Private CB slots cost 0.14 GB here vs the 34 MB reusing stack — the memory price of experiment 5 unless placement learns DAG lifetimes.
  3. Full cross-stream overlap deadlocks structurally on this tree (62 wides from level ~3; 31.5k dependent blocks saturate SMs ahead of the vendor stream). The segmented form is the deadlock-free shape. Backward counter kernels must dispatch parents-first or they deadlock outright.
  4. The regime-B width boundary at 64 is correct: replacing the vendor path for 64 < w ≤ 96 with local-memory kernels regresses 32→42 ms — cuSOLVER's blocked potrf wins above w ≈ 64. Conversely regime_c_rows matters as much as width: tall-narrow fronts belong on the split path (the solve's dense-path rows trigger is a third of its win).
  5. Subtree kernel occupancy is local-memory-bound: dropping the 48 KB budget class (subtree_budgets = [16384]) is worth 13.3 → 8.4 ms on that bucket.
  6. FP32 factorization is unusable on these matrices (κ ≈ 1e14; FP64's 1.4e-2 single-solve residual is itself the κ·ε floor), consistent with cuDSS's documented FP32 failures — the kernels are element-type generic now regardless, and the attempt exposed a Float64 assumption in subtree local-memory sizing (fixed on this branch in the bench launcher; the src launcher sizes by sizeof(T) already).
  7. Smaller mechanisms worth keeping: concurrent extend-add needs atomic adds; packed-triangle iteration is a loss inside element loops (serial O(m) unpack); _factor_panel_c!'s info check host-synchronizes, blocking stream overlap of independent wide fronts; CUDA-graph replay is neutral at this scale (busy-bound) and only pays below n ≈ 50k.

Review guidance

The src diff is the only thing proposed for main as-is. The bench prototypes are experiment records: deliberately CUDA-only, SPD-only, nbatch = 1, and kept exactly as measured (including dead ends) so the numbers in this PR remain reproducible — bench/solve_proto_78k.jl and bench/fact_fused.jl each reproduce their headline table in one run given the KKT dump (via bench/dump_madnlp_kkt.jl). Where a prototype contradicts a PLAN.md design assumption (stack placement, ordering cost), the finding is flagged in the comments rather than edited into PLAN.md, per AGENTS.md.

🤖 Generated with Claude Code

sshin23 and others added 2 commits October 8, 2026 10:43
…ter sweeps

Experiments 5/6 of PERFORMANCE.md (issues #82, #25), outside src/: after
factorization every front's L11 (w <= 256) is inverted into a side buffer so
each front's TRSV becomes one GEMV; the forward and backward sweeps then run
as one fused kernel (regime-A subtrees merged as leading/trailing blocks,
etree child-counters with spin/threadfence signaling) plus one single-block
chain kernel for the near-serial tall fronts at the top, with the vendor
dense path only on the remaining wide roots. CUDA-only, SPD, single RHS.

On a Quadro GV100 the pglib_opf_case78484_epigrids condensed KKT solve goes
from 18.0 ms (stock sweeps) to 6.3 ms, against 3.8 ms for cuDSS; the full
SPD harness table is in the PR description, including the matrices where the
approach regresses and a per-schedule strategy pick is needed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Extracts cuDSS's reordering permutation (raw cudssDataGet, the high-level
getter refuses it) and replays the prototype solve under user_perm, plus an
amalgamation sweep. On the 78k condensed KKT the cuDSS ordering halves the
elimination-tree depth (65 -> 30 levels) and takes the prototype solve from
6.35 ms to 4.81 ms on a GV100; amalgamation changes only regress.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up: the reordering is a third of the remaining gap

Extracting cuDSS's reordering permutation (raw cudssDataGet with CUDSS_DATA_PERM_REORDER_ROW; the CUDSS.jl getter refuses the vector parameters) and replaying the prototype solve under user_perm, order-controlled A/B on the 78k condensed KKT (GV100, median of 20 warm):

ordering etree levels nnz(L) supernodes analysis prototype solve
METIS ND (SDS default) 65 13.2M 92,878 12 s 6.35 ms
cuDSS reordering via user_perm 30 14.8M 92,220 4.3 s 4.81 ms

cuDSS's custom ND halves the elimination-tree depth at +12% fill, and on a latency-bound solve that is worth 1.3×: the prototype goes to 4.81 ms against cuDSS's 3.77 ms (1.27× remaining). Supernode counts are nearly identical, so this is separator/recursion shape, not amalgamation — and indeed an amalgamation sweep ((64, 0.5, 16) … (256, 1.0, 64)) only regresses (8.7–28 ms, table in bench/partition_sweep.jl output).

Two implications for the plan:

  • Reordering depth belongs next to fill in the ordering objective (evaluate_ordering already computes nlevels — it could weigh it). METIS knobs (nd_nlevels, nd_ubfactor, nseps) may close part of this without importing a permutation.
  • When the pattern repeats (MadNLP), importing the cuDSS permutation once is already practical: analysis under user_perm is 4.3 s, faster than both SDS-with-METIS (12 s) and cuDSS's own analysis (6.5 s).

🤖 Generated with Claude Code

…n the 78k KKT

With the cuDSS reordering (depth-30 etree) the two-tier split is obsolete:
one fused dependency-counter kernel per direction covering every front of
width <= 256 solves the 78k condensed KKT in 3.61 ms on a GV100, against
3.77 ms for cuDSS and 18.0 ms for the stock sweeps. Under the METIS ordering
the same configuration reads 5.64 ms, so the shallow ordering and the
single-tier structure are both required.

Also fixes a latent correctness bug in the prototypes: build_fused keyed the
live-subtree-root accounting on the width cap rather than on whether the
segment is the merged tier, so a tier-1 cap wider than 128 double-counted
subtree-root signals and released waiters early (relres 5e+11).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Final round: single-tier fused solve under the cuDSS ordering — 3.61 ms vs cuDSS 3.77 ms

The two-tier structure (fused counters below, single-block chain above) was tuned on the depth-65 METIS tree. Under the cuDSS ordering (depth 30) the chain is obsolete: one fused dependency-counter kernel per direction covering every front of width ≤ 256 is both simpler and faster. Order-controlled, 3 repeats, relative residual equal to stock/cuDSS:

configuration (GV100, 78k condensed KKT, 1 RHS) solve
stock level-batched sweeps, METIS 18.0 ms
two-tier prototype, METIS 6.3 ms
two-tier prototype, cuDSS ordering 4.81 ms
single-tier prototype, cuDSS ordering 3.61 ms
cuDSS 0.8.0 3.77 ms

Single-tier under METIS reads 5.64 ms — the shallow ordering and the single-tier structure are only fast together. The entire solve is now two custom kernels + the vendor root + two permute kernels.

The sweep also flushed out a latent correctness bug, fixed in this branch: build_fused keyed its live-subtree-root accounting on the width cap (wcap ≤ 128) as a proxy for "this is the merged tier". With a wider tier-1 cap, subtree-root signals were double-counted and waiters released early — a silently wrong solve (relres 5e+11) that was faster than the correct one, i.e. exactly the kind of fail-open a fixed-seed residual check catches and a timing comparison would have happily reported.

Caveats as before: tuned for KKT-shaped trees; the ordering is imported from a one-time cuDSS analysis of the same pattern (bench/get_cudss_perm.jl), after which SDS's own analysis under user_perm is 3.6 s — faster than both SDS-with-METIS (12 s) and cuDSS's analysis (6.5 s).

🤖 Generated with Claude Code

…x906)

Portability shims only: ROCArray + the CSR-triplet DirectSolver constructor
(no AMDGPU convenience extension exists yet, T23), and the device-scope fence
as Threads.atomic_fence(), which both GPU backends lower to an LLVM fence.
Atomix atomics, the dependency-counter spinning and the dispatch-order
deadlock-freedom assumption all hold on gfx906.

Radeon VII, 78k condensed KKT, 1 RHS, median of 20 warm, residuals matching
the stock sweeps: stock 20.5 ms -> 8.05 ms (single tier, cuDSS ordering);
under METIS 28.5 -> 12.2 ms. Same structural ranking as on the GV100, where
cuDSS does not exist as a comparison point at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

AMDGPU: it runs, first try

Radeon VII (gfx906, ROCm), same matrix, same protocol, residuals matching stock:

config stock sweeps prototype
METIS ordering, two-tier 28.5 ms 12.2 ms
cuDSS ordering, two-tier 20.7 ms 8.80 ms
cuDSS ordering, single-tier 20.5 ms 8.05 ms

Portability cost was three mechanical substitutions (bench/amd_proto.jl): ROCArray, the CSR-triplet DirectSolver constructor (no AMDGPU extension yet, T23), and CUDA.threadfence() → Threads.atomic_fence(), which both backends lower to an LLVM fence. Atomix atomics, the spin-counters and the dispatch-order deadlock-freedom assumption all held on gfx906 at 1024-thread workgroups. The wide-root vendor path fell back to the generic kernels, so T23's rocBLAS bindings would narrow the remaining stock-vs-prototype gap further.

There is no cuDSS on this hardware — this is the portability case of PLAN/PERFORMANCE.md made concrete.

🤖 Generated with Claude Code

… 78k KKT

The stock fused regime-B kernel keeps each front's trsm and contribution-block
update inside the front's single workgroup, which serializes 1-2 ms per fat
front near the root. This prototype keeps assembly + F11 Cholesky one-block
and moves the trsm (128-row chunks) and the syrk (32x32 tiles, one workgroup
each, disjoint writes, deterministic) of fronts with m > 64 to separate
many-block kernels; the factor is bit-identical to stock on every config.

With the routing knobs that fall out of the same measurements (fat fronts to
the vendor path via regime_c_rows = 256, subtree serial walks capped via
subtree_parallelism = 16384, cuDSS ordering), refactorization of the 78k
condensed KKT goes 167.7 -> 64.0 ms on a GV100 (cuDSS: 24.3 ms). The tiled
syrk does the dominant flops in 1.2 ms where the fused kernel spent ~12.
CUDA-graph replay of the whole sequence is validated bit-identical and
timing-neutral: the remainder is kernel-busy time (vendor dense calls for the
303 C fronts, the one-block panel kernels, the longest regime-A subtrees and
the level-batched extend-add), not launch overhead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Factorization round: 167.7 → 64.0 ms (cuDSS 24.3), factor bit-identical

Same method as the solve: measure, route, split. bench/fact_split.jl.

refactorization (GV100, 78k condensed KKT) time
stock, METIS, defaults 167.7 ms
+ cuDSS ordering 138.8 ms
+ regime_c_rows = 256, subtree_parallelism = 16384 72.7 ms
+ split kernels (tiled syrk + chunked trsm for m > 64) 64.0 ms
cuDSS 24.3 ms

Findings:

  • The fused regime-B kernel's one-workgroup-per-front design costs 1–2 ms per fat front near the root; its syrk is ~12 ms of that and the tiled replacement does the same flops in 1.2 ms, bit-identical (disjoint 32×32 tile writes, no atomics — stays inside the determinism contract).
  • regime_c_rows matters as much as regime_c_width: tall-narrow fronts are the expensive ones, and the measured optimum keeps m > 256 on the vendor path even after the splits (pulling them into regime B trades the vendor launch storm for one-block assembly serialization and loses).
  • subtree_parallelism = 16384 halves the subtree term the same way it helped the solve: the floor is the longest serial subtree walk.
  • CUDA-graph replay of the whole ~1600-launch sequence: validated bit-identical, timing-neutral — the remaining 64 ms is kernel-busy, not launch gaps. The next levers are all kernel-shape work: tiled extend-add, split assembly for tall fronts, batched vendor calls for the C fronts.

Per IPM iteration the prototype stack now reads: refactorization + solve = 64.0 + 3.6 = 67.6 ms vs cuDSS 24.3 + 3.8 = 28.1 ms, from 185 ms stock.

🤖 Generated with Claude Code

…-> 40.6 ms

The vendor dense path cost ~26 ms in launch-and-API overhead for 303
tall-narrow fronts (~8 launches each). Those fronts keep the stock multi-block
assembly kernels but now factor through the split kernels batched per level: a
no-assembly variant of the panel kernel (w <= 64), the 128-row trsm chunks and
the 32x32 syrk tiles; the vendor path keeps only fronts wider than 64. With
that split, lowering regime_c_rows monotonically helps (64 is the floor), and
subtree launches run longest-first through a replicated launcher. A tiled
extend-add was implemented and measured neutral: that bucket is scatter
bandwidth, not parallelism, and is at its floor.

GV100, 78k condensed KKT: refactorization 40.6 ms (was 64.0 after round 1,
167.7 stock; cuDSS 24.3). Relative factor difference to stock is ~3e-10 where
vendor arithmetic was replaced, residuals unchanged. Per IPM iteration the
prototype stack is now 40.6 + 3.6 = 44.2 ms against cuDSS's 28.1, from 185 ms
stock.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Factorization round 2: 64.0 → 40.6 ms (cuDSS 24.3)

refactorization (GV100, 78k condensed KKT) time
stock best (round-1 knobs) 72.7 ms
round 1 split (tiled syrk + chunked trsm) 64.0 ms
+ C-front split: stock assembly, split factor kernels, vendor only for w > 64 45.1–54.9 ms
+ regime_c_rows → 64 (monotone once C fronts are cheap) 41.7 ms
+ longest-first subtree launches 40.6 ms
cuDSS 24.3 ms

The C-front finding is the transferable one: the vendor dense path's cost on tall-narrow fronts is ~8 launches of API overhead per front, while its assembly kernels (multi-block zero/scatter/extend-add) are the efficient part — the exact inverse of regime B, whose fused kernel has cheap assembly access but serializes the factor math. The split takes the best half of each. Also measured and worth recording as negatives: a tiled extend-add is timing-neutral (that bucket is scatter bandwidth at ~10 ms, a real floor), and CUDA-graph replay remains neutral throughout.

Factor matches stock to ~3e-10 relative (vendor arithmetic replaced for ex-C fronts; bit-identical on B-only configs). Remaining buckets: subtrees ~12 ms (longest-walk bound), panel ~8, wide-root vendor ~8, extend-add ~9 — in-src territory (persistent kernels, batched wide fronts).

Prototype stack per IPM iteration: 44.2 ms vs cuDSS 28.1 ms, from 185 ms stock (4.2×).

🤖 Generated with Claude Code

Four directions measured and closed on the 78k condensed KKT (GV100):
- amalgamation under the split kernels: all variants regress (default 32/0.25/8
  optimal, as under the stock kernels);
- overlapping the subtree budget-class launches on a second stream: neutral
  (the large class saturates the device) and it invalidates graph capture;
- Float32 factorization: the kernels are generic now and run 18% faster, but
  the matrix has kappa ~ 1e14 (FP64 relres 1.4e-2 ~ kappa * eps), so FP32
  factors are unusable, Jacobi scaling does not change the pivot failure
  (same failed column), and even delta = 1e-6 regularization in FP64 is
  unusable for the original system (relres 2.1) - consistent with cuDSS's
  documented FP32 failures on the condensed dumps;
- the subtree launcher now sizes local memory by sizeof(eltype) (a latent
  Float64 assumption that produced illegal addresses under Float32).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Factorization round 3: boundary mapping — 40.6 ms stands

Four directions measured to closure (all in bench/fact_split.jl history):

direction result
amalgamation under the split kernels all variants regress; default (32, 0.25, 8) is optimal
subtree classes on concurrent streams neutral (big class saturates the GPU); breaks graph capture
Float32 factorization kernels generic + run at 33.3 ms (−18%), but κ ≈ 1e14 makes FP32 factors unusable; Jacobi scaling leaves the same pivot failure; δ-regularization unusable even in FP64 (relres 2.1) — matches cuDSS's documented FP32 failures on condensed dumps
(fixed en route) subtree launcher local-memory sizing assumed Float64 — illegal address under FP32; now sizeof(eltype)

The conditioning argument is worth keeping: FP64's single-solve relres of 1.4e-2 on these matrices is itself the κ·ε floor, so no precision trick at the solver level can help — accuracy beyond that needs refinement against an FP64 factor, which is already the design (T16/T18).

Final prototype scoreboard (GV100, 78k condensed KKT, per IPM iteration): 44.2 ms (refactorization 40.6 + solve 3.6) vs cuDSS 28.1 ms, from 185 ms stock — 4.2×, remaining gap 1.57×, with the solve itself faster than cuDSS. Remaining distance is in-src work: persistent/merged kernels for the subtree and panel buckets, batched wide-front processing, and the scatter-bandwidth extend-add floor.

🤖 Generated with Claude Code

One fused kernel per segment whose workgroups play four roles in a
topologically ordered grid (extend-add chunks, panels, trsm chunks, syrk
tiles) with etree counters; segments split at the groups containing wide
fronts, which run between segments through _factor_panel_c!. Verified: zero
failing fronts, factor matches the split pipeline (~3e-10, vendor-arithmetic
class), residuals unchanged, stable over repeats and rows=64/96.

Two findings that outlive the prototype:
- Full overlap (variant i: wides on a concurrent stream behind spin-waiters)
  DEADLOCKS structurally on this tree: 62 wide fronts appear early, 31.5k of
  34.7k blocks depend on them, and spinning dependents saturate every SM
  before the vendor stream can run. Measured, not hypothesized; the segmented
  form is the deadlock-free shape of this idea.
- The update stack's contribution-block placement (PR #54 lifetimes) assumes
  level-synchronous execution. Any out-of-level-order execution -- counters,
  streams, or persistent kernels -- clobbers live CBs through slot reuse, and
  a global zero pass does the same. The prototype gives every non-subtree
  front a private CB slot instead: 0.14 GB here (~4x the reusing stack),
  which is the memory price of experiment 5 unless placement learns DAG
  lifetimes. Concurrent extend-add children additionally need atomic adds
  (the stock owner-pull order serializes children; the fused grid does not).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Factorization round 3 (variant i → segmented): 40.6 → 33.5 ms (cuDSS 24.3)

Built at the owner's request as the full-overlap experiment; the honest outcome is a working segmented form plus two structural findings that any in-src experiment-5 implementation will need:

refactorization (GV100, 78k condensed KKT) time
split kernels (round 2) 40.6 ms
segmented dependency-counter mega-kernel 33.5 ms
cuDSS 24.3 ms
  • Variant i (full overlap) deadlocks structurally on this tree: 62 wide fronts appear from level ~3 up, 31.5k of 34.7k blocks transitively depend on them, and spinning dependents saturate the SMs before the concurrent vendor stream can schedule. The segmented form (split at wide groups, vendor between segments) is the deadlock-free shape; fusion within segments keeps most of the win.
  • The update stack is incompatible with out-of-level-order execution. CB slot reuse (PR layout: place contribution blocks offline over their lifetimes (issue #48) #54) assumes level-synchronous lifetimes; counters/streams/persistent kernels clobber live CBs, and failures surface as mid-column pivot errors far from the cause. The prototype gives every non-subtree front a private CB slot — 0.14 GB on this matrix vs the 34 MB reusing stack — which is the real memory price of experiment 5 unless placement learns DAG lifetimes.
  • Concurrent extend-add requires atomic adds (children of one parent race on shared slots once the serial owner-pull order is gone); a fail-open reminder that the broken runs were faster — every timing here is gated on zero failing fronts, Δfactor ≈ 3e-10, and matching residuals.

Stack now: refactorization 33.5 + solve 3.6 = 37.1 ms per IPM iteration vs cuDSS 28.1, from 185 ms stock — 5.0×, remaining gap 1.32×.

🤖 Generated with Claude Code

The cuDSS-ordering advantage was never about cuDSS's partitioning: forcing
SDS's own nested dissection (reordering_alg = "algo4") yields schedule depth
29 vs the import's 30, factorization 33.1 ms vs 34.8, solve 3.5 vs 3.7 on the
78k condensed KKT. The depth-65 default came from the automatic AMD/ND
chooser: ordering_cost = flops * (1 + nlevels/n) weighs depth at ~0.2% for
n = 674562, so AMD wins on flops — and nlevels is COLUMN-etree depth, which
does not predict the supernodal schedule depth that governs GPU time (cuDSS
perm: 1279 column levels -> 30 schedule levels; METIS: 1216 -> 29; both
shallower at column level than their schedule ratio suggests). A METIS knob
sweep (seed, ufactor, nseps, ccorder, pfactor) moves schedule depth 28-41
with seed-dependent noise; nothing beats plain ND meaningfully.

Consequences: the prototype stack no longer depends on cuDSS for anything
(relevant for AMDGPU, where cuDSS does not exist), and the ordering chooser
should score supernodal schedule depth (or simply prefer ND on GPU backends).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Ordering study (proposal E): reordering_alg = "algo4" replaces the cuDSS import

ordering (78k condensed KKT, GV100) schedule depth refactorization solve
SDS default (auto → AMD) 65 — baseline configs — —
imported cuDSS reordering 30 34.8 ms 3.7 ms
native ND (algo4) 29 33.1 ms 3.5 ms
METIS seed sweep best 28 33.2 ms 3.6 ms

Root cause of the depth-65 default, in two parts: (1) ordering_cost = flops × (1 + nlevels/n) weighs depth at ~0.2% for n = 674k, so the AMD/ND chooser is a pure flop contest that AMD wins; (2) nlevels is column-etree depth, which doesn't predict the supernodal schedule depth that governs GPU time — the cuDSS perm has more column levels than METIS (1279 vs 1216) yet the same schedule depth, and AMD's schedule is 2× deeper despite competitive column metrics. Recommendation: score candidate orderings by supernodal schedule depth (cheap to compute from the supernode partition), or simply prefer ND on GPU backends.

Practical consequences: the whole prototype stack is now cuDSS-free — one setparam!(solver, "reordering_alg", "algo4") — which also makes the AMDGPU path self-contained. user_nd_partition_tree import was checked and is T24-pending in SDS; the search shows it isn't needed. (En route, the per-trial residual gate caught a stale copy of the early-release counter bug a second time — fast-but-wrong remains this work's defining hazard.)

Scoreboard, now cuDSS-free: refactorization 33.1 + solve 3.5 ≈ 36.6 ms per IPM iteration vs cuDSS 28.1, from 185 ms stock — 5.1×.

🤖 Generated with Claude Code

A child whose contribution block has exactly one consumer writes it straight
into the parent's frontal matrix from its syrk tiles, skipping the CB
round-trip. Safety conditions measured out stricter than estimated: the
parent must be pre-assembled (C-group; B parents zero in-kernel and would
wipe the contribution), ALL CB children count as siblings (subtree-root
children extend-add concurrently, so 'only child' must include them), and
wide children are excluded (they have no tiles to settle the ea_left debt --
the first version deadlocked on exactly that under one ordering). The
eligible set is then ~80 fronts / ~1.4M CB entries (~8% of CB traffic, not
the ~30% a looser census suggested), worth ~0.1 ms: 33.07 -> 32.94 ms under
native ND, within noise. Kept because it is correct and free; recorded
because the census-first discipline is what kept it from being sold as a win.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Proposal B (write-through extend-add): correct, kept, marginal

Implemented in the fused kernel: a child with a single CB consumer writes its finished contribution straight into the parent's frontal matrix from its syrk tiles. The safety conditions turned out stricter than the back-of-envelope: parent pre-assembled (C-group only), siblings counted including subtree-root children (they extend-add concurrently), and wide children excluded — a wide child has no tiles to settle the dependency debt, and the first version deadlocked on exactly that under one ordering (caught by the timeout guard, fixed by one condition).

Eligible set after the honest conditions: ~80 fronts, ~1.4M CB entries ≈ 8% of CB traffic (not the ~30% a looser census suggested) → 33.07 → 32.94 ms under native ND, within noise. Kept because it's correct and free.

Running scoreboard (GV100, 78k condensed KKT, cuDSS-free native-ND stack): refactorization 32.9 ms + solve 3.5 ms ≈ 36.4 ms per IPM iteration vs cuDSS 28.1, from 185 ms stock.

🤖 Generated with Claude Code

…-neutral

Adds an analysis tuning option capping the number of fronts a regime-A
subtree may contain (0 = no limit, the default; the serial walk of a subtree
is its workgroup's critical path). One field in Options, one count in the
eligibility pass of build_schedule, validated like subtree_parallelism.

Measured on the 78k condensed KKT (GV100, native ND, fused prototype):
33.03 ms uncapped vs 32.77 at cap 32 — within noise. The sweep disproves the
longest-walk-floor hypothesis for this configuration: with ~3.3k subtrees the
launch is throughput-bound, and the 188-front walk hides inside the waves.
The knob is kept because it is the right instrument to re-test that on
machines with more SMs or trees with fewer subtrees, where the balance tips.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Proposal C (subtree front-count cap): implemented as subtree_max_fronts, measured near-neutral

Small src addition on this branch: Options(subtree_max_fronts = N) caps regime-A subtree size by front count (0 = off), one count in build_schedule's eligibility pass, validated like subtree_parallelism.

Sweep on the 78k condensed KKT (GV100, native ND, fused prototype): uncapped 33.03 ms; caps 96/64/48/32 → 32.97/32.96/32.99/32.77 ms. Within noise — which falsifies the longest-walk-floor hypothesis for this configuration: with ~3,316 subtrees the launch is throughput-bound and the 188-front serial walk hides inside the occupancy waves. The knob stays because it's the right instrument on devices with more SMs (or schedules with fewer subtrees), where the critical path can surface.

Scoreboard unchanged within noise: ~36.3 ms per IPM iteration, cuDSS-free.

🤖 Generated with Claude Code

…s mapped

Gain: subtree_budgets = [16384] raises the subtree kernel's occupancy (the
48 KB class caps it at 1 block/SM) and is the new best, 33.0 -> 31.9 ms
(single-class launch, 8.4 ms bucket from 13.3).

Three negatives with mechanisms, kept as knobs/record:
- packed-triangle extend-add iteration regresses 33 -> 42: the per-element
  column unpack is a serial O(m) walk, costlier than the idle upper-half
  threads it removes (fine once per thread in the trsm chunk kernel, fatal
  inside the element loop);
- replacing the vendor path for 64 < w <= 96 fronts with 96-class localmem
  kernels regresses 32 -> 42: cuSOLVER's blocked potrf beats a column-serial
  localmem Cholesky above w ~ 64, i.e. the existing regime-B width boundary
  is in exactly the right place;
- a stream pool for the independent boundary wides is neutral: the
  _factor_panel_c! info check synchronizes the host per call, so overlap
  needs an async-info vendor path in src, not more streams.

Remaining to cuDSS (24.3): mega-kernel internals (11.0), subtrees (8.4),
vendor dense (10.1) — each now requires src-grade work (role rebalancing,
subtree walk redesign, async potrf info).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

33 → 24 push: landed 31.9 ms; the rest is src-grade

One gain: subtree_budgets = [16384] — the 48 KB local-memory class caps the subtree kernel at 1 block/SM, so dropping it raises occupancy: 33.0 → 31.9 ms (subtree bucket 13.3 → 8.4, single-class launch).

Three negatives, each with its mechanism (details in bench/fact_fused.jl history): packed-triangle extend-add iteration (serial O(m) unpack per element, 33→42), 96-class localmem kernels for 64<w≤96 wides (cuSOLVER's blocked potrf wins above w≈64 — the regime-B width boundary at 64 is in exactly the right place), and a stream pool for independent boundary wides (neutral: _factor_panel_c!'s info check host-synchronizes per call — overlap needs an async-info vendor path).

Remaining decomposition at 31.9: mega-kernel 11.0 + subtrees 8.4 + vendor dense ~10 + stats 0.6. Closing to 24.3 needs src work on all three: role-level rebalancing inside the fused kernel, the subtree walk's per-front efficiency, and async vendor info handling. Bench-level levers are exhausted — effect sizes are now inside run-to-run noise.

Final factorization: 167.7 → 31.9 ms (5.3×), 1.31× from cuDSS, cuDSS-free. Per IPM iteration with the solve: ~35.4 ms vs cuDSS 28.1, from 185 stock.

🤖 Generated with Claude Code

Its copy of build_fused predates the merged-tier fix (the early-release
counter bug caught twice by the residual gate); solve_nd.jl carries the
same experiment on the fixed machinery.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23 sshin23 changed the title Prototype: partitioned-inverse solve with fused dependency-counter sweeps (experiments 5/6) Performance prototypes: partitioned-inverse solve, split/fused factorization, ordering study (experiments 0/5/6) Oct 8, 2026
@sshin23
sshin23 marked this pull request as ready for review October 8, 2026 23:58
@sshin23
sshin23 requested a review from michel2323 October 9, 2026 00:02
…(CUDA and AMDGPU)

MadNLPSDS.jl implements MadNLP.AbstractLinearSolver over the public SDS
handle API, modeled on MadNLPGPU 0.10's CUDSSSolver: device-CSC input read
as the CSR of the upper triangle (view 'U'), values aliased with an update!
fallback, in-place solve_linear_system!, Cholesky info mapped to inertia.
Backend-agnostic: the same ~90 lines drive CUSPARSE and rocSPARSE matrices.

Full MadNLP on the 78k-bus ACOPF (SparseCondensedKKTSystem, tol 1e-6):
identical iteration count (101), objective and inertia behavior across all
three configurations. GV100: cuDSS 16.8 s wall (9.7 s iteration loop) vs
SDS-with-stock-kernels 38.8 s (26.1 s loop) — the 2.7x is the stock-kernel
ratio documented in this PR; the prototype kernels are not wired into src.
AMD Radeon VII: 71.0 s wall (61.3 s loop) — to our knowledge the first
GPU-resident sparse-direct ACOPF interior-point run on AMD hardware, where
no cuDSS exists; this wrapper is the missing sparse solver for MadNLPGPU's
AMDGPU extension. One pitfall recorded: linear-solver tuning knobs are
phase-coupled (regime_c_rows = 64 is right for the prototype factorization
kernels but puts ~300 fronts on the stock solve's vendor path, 215 ms per
backsolve; 256 is the stock-balanced setting).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

End-to-end: full MadNLP ACOPF with SDS, on CUDA and on AMD

bench/e2e/MadNLPSDS.jl is a ~90-line backend-agnostic MadNLP.AbstractLinearSolver over the public handle API (modeled on CUDSSSolver). Full interior-point runs, 78k-bus ACOPF, SparseCondensedKKTSystem, tol 1e-6 — identical iterations (101), objective, and inertia behavior in all three configurations:

cuDSS / GV100 SDS / GV100 (stock kernels) SDS / Radeon VII
wall 16.8 s 38.8 s 71.0 s
analysis (init) 7.1 s 12.7 s 9.6 s
iteration loop 9.7 s 26.1 s 61.3 s
factorizations / backsolves 198 / 164 194 / 190 195 / 182

The GV100 gap is exactly the stock-kernel ratio measured throughout this PR — the prototype kernels live in bench/, so an in-src adoption of the split/fused factorization and algo1 solve closes most of the loop difference. The AMD column is, to our knowledge, the first GPU-resident sparse-direct ACOPF IPM on AMD hardware — there is no cuDSS there, and this wrapper is the missing sparse solver for the AMDGPU extension.

One operational finding: the tuning knobs are phase-coupled — regime_c_rows = 64 (right for the prototype factorization) puts ~300 fronts on the stock solve's per-front vendor path (215 ms/backsolve, a 6× end-to-end regression caught by the counters); 256 balances stock factorization and solve.

🤖 Generated with Claude Code

SDSProto.jl assembles this PR's solve and factorization prototypes
(partitioned-inverse + fused counter sweeps; segmented fused factorization
with private contribution blocks) into one backend-portable module
(KernelAbstractions allocation, Threads.atomic_fence), and SDSProtoSolver
drives them from MadNLP. Because both phases are prototype kernels, the
phase-coupled routing conflict disappears (rows = 64 is optimal for both).

Full MadNLP, 78k-bus ACOPF, identical 101 iterations and objective in every
configuration:

                      wall      iteration loop
  cuDSS   / GV100     16.9 s     9.8 s
  stock   / GV100     38.8 s    26.1 s
  PROTO   / GV100     25.5 s    11.7 s   (1.2x from cuDSS end-to-end)
  stock   / RadeonVII 72.8 s    63.2 s
  PROTO   / RadeonVII 43.6 s    33.9 s   (1.9x over stock on AMD)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

Capstone: the PR's prototype kernels inside MadNLP, on both vendors

bench/e2e/SDSProto.jl packages the prototype solve (partitioned inverse + fused counter sweeps) and factorization (segmented counters, private CBs) as one backend-portable module; SDSProtoSolver drives it from MadNLP. With both phases on prototype kernels the phase-coupled knob conflict disappears.

Full interior-point, 78k-bus ACOPF, identical 101 iterations and objective everywhere:

wall iteration loop
cuDSS / GV100 16.9 s 9.8 s
SDS stock / GV100 38.8 s 26.1 s
SDS prototype / GV100 25.5 s 11.7 s (1.2× from cuDSS)
SDS stock / Radeon VII 72.8 s 63.2 s
SDS prototype / Radeon VII 43.6 s 33.9 s

The loop arithmetic closes: 193 factorizations × ~36 ms + 202 solves × ~4 ms + wrapper sync ≈ 11.7 s. The remaining GV100 distance to cuDSS is the documented factorization gap (31.9 vs 24.3 ms) plus per-call synchronization that an in-src integration does asynchronously. On AMD this is now a practical end-to-end path — 34 s of iteration loop for a 78k-bus ACOPF on a 2019 consumer card, with no vendor sparse solver in existence there.

🤖 Generated with Claude Code

sshin23 and others added 2 commits October 8, 2026 22:01
…lity

The src knob from this branch now carries its test, in the house style of
the subtree_parallelism testset: Options validation (default, copy, range,
not-a-setparam-string), every subtree within the cap as counted from the
schedule's own subtree_ptr spans and matching the independently accumulated
etree front counts, regime A only shrinking, and maximality (a subtree
root's parent is over the cap or was outside regime A before it). Green on
the CPU suite locally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The test suite already includes ROCBackend when AMDGPU.jl is functional
(test/backends.jl), and the job template already installs AMDGPU.jl for the
amdgpu matrix entry — only the matrix entry itself was disabled. Run it as a
non-required continue-on-error leg: it exercises the KernelAbstractions
kernels on AMD hardware today (the bench prototypes on this branch validated
that path end-to-end on a gfx906), and T23 promotes it to a required check.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshin23

sshin23 commented Oct 9, 2026

Copy link
Copy Markdown
Contributor Author

First AMD CI run: 99.4% pass; T23 is ~a constructor extension

The amdgpu matrix leg (enabled continue-on-error on this branch) was picked up by the pinwheel runner and ran the full suite: 29,724 pass / 2 fail / 172 error of 29,898, in 9m52s.

The errors are essentially one missing method, repeated: 158× MethodError: no method matching DirectSolver(::rocSPARSE.ROCSparseMatrixCSR…), plus the matching CSR(::ROCSparse…) and cholesky(::ROCSparseMatrixCSR) forwarders — i.e., exactly the convenience-constructor surface of the CUDA extension, ROC-flavored. Everything reachable through the raw-vector constructors (the kernels, both backends' numeric and solve paths) passes. T23's required core is therefore ~30 lines of matrix-wrapper forwarding; rocBLAS/rocSOLVER dense bindings are the optional second half (the KA fallbacks carried the dense path through this run).

The leg stays continue-on-error on this branch; with T23's constructors it should go fully green and can be promoted to a required check.

🤖 Generated with Claude Code

@michel2323 michel2323 added the claude:pr PR opened by the Claude implementer; handled by the Claude pipeline label Oct 9, 2026
Comment thread bench/amd_proto.jl
Comment on lines +1 to +19
# Prototype (PERFORMANCE.md experiments 5/6, issues #82/#25): partitioned-inverse
# solve with fused dependency-counter sweeps, CUDA only, SPD, 1 RHS, nbatch = 1.
#
# After factorization, invert every front's L11 (w <= MW) into a side buffer;
# each front's TRSV becomes one GEMV. The sweeps then run as:
# fwd: ONE fused kernel (regime-A subtrees as leading blocks + every front of
# width <= 128 below the first wide front, etree child-counters with
# spin/threadfence signaling, atomic L21 scatter)
# + ONE single-workgroup "chain" kernel walking the near-serial tall
# fronts (w <= 256) at the top with plain block barriers
# + the stock vendor dense path for the remaining wide roots.
# bwd: the same in reverse (chain waits dispatch parents-first: fronts MUST
# get reverse-topological block ids or the kernel deadlocks).
# Measured on a Quadro GV100, pglib_opf_case78484_epigrids condensed KKT
# (n = 674562): stock sweeps 18.0 ms -> 6.3 ms; cuDSS 3.8 ms. See the PR body
# for the full matrix table (wide-front matrices regress; pick per schedule).
# Constraints: CHWG >= NSTRIP * MW (strip indexing); NSTRIP/MW compile-time.
#
# CUDA_VISIBLE_DEVICES=1 julia +1.13 --project=bench bench/solve_proto_78k.jl

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit (non-blocking): this header is a verbatim copy of bench/solve_proto_78k.jl's. Line 2 says "CUDA only" and line 19 gives the CUDA script's invocation (--project=bench bench/solve_proto_78k.jl), but this file is the ROCm port that runs under bench/amd/Project.toml. Suggest replacing the two lines with something like "ROCm port of solve_proto_78k.jl (ROCArray, CSR-triplet constructor, KA fence), SPD, 1 RHS" and julia +1.13 --project=bench/amd bench/amd_proto.jl, so the AMD numbers in the PR body stay reproducible from the file alone.

Comment thread src/symbolic/schedule.jl
Comment on lines +358 to +359
eligible[s] = !big[s] && all(c -> eligible[c], kids) && peak[s] <= maxcap && work[s] <= maxwork &&
nfronts[s] <= maxfronts

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit (non-blocking): the logic is right (nfronts is monotone up the tree, so eligible_cap[s] == eligible[s] && nfronts[s] <= maxfronts and the maximal subtrees stay well defined, which is exactly what the new test checks). But the build_schedule docstring above (item 2, "regime A: a front is eligible when …") still lists only the C-threshold, the stack-peak and the flop-limit conditions. Add the front-count limit there, e.g. "… and its subtree has at most opts.subtree_max_fronts fronts (0: no limit)", so the docstring and the Options table agree on what makes a front eligible.

Comment thread .github/workflows/ci.yml
Comment on lines +75 to +82
# 'amdgpu' runs as a non-required, continue-on-error leg until the
# AMDGPU extension exists (T23): the test suite already includes
# ROCBackend when AMDGPU is functional (test/backends.jl), so this
# exercises the KernelAbstractions kernels on AMD hardware today.
# Promote it to the required checks of the branch ruleset with T23.
backend: ['cuda', 'amdgpu']
runs-on: [self-hosted, linux, X64, gpu, '${{ matrix.backend }}']
continue-on-error: ${{ matrix.backend == 'amdgpu' }}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit (non-blocking): the leg works as intended — I checked that run 37874446430 concluded success with the amdgpu job red, so claude-pipeline.yml's all_green/ci_failed logic (workflow-run conclusion) is unaffected and no CI-fix round is triggered by the AMD leg. Two follow-ups for whoever touches this next: (1) AGENTS.md ("Environment → GitHub Actions", the sentence "the amdgpu runner is disabled in ci.yml until the AMDGPU extension exists, T23") is now stale and should say the leg runs non-required/continue-on-error and only needs promoting to a required check with T23; (2) the 2 failures + 172 errors of this run are all the missing ROC constructor surface — the two test_ubatch.jl:245 failures catch a MethodError from cholesky(::ROCSparseMatrixCSR; view, check) instead of a FactorizationError, same cause as the errors — so the PR comment's "T23 is ~a constructor extension" reading is confirmed by the log.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

VERDICT: APPROVE (head f55e3a8)

What this PR is. Not a TNN task of TASKS.md: an owner-labelled (claude:pr by michel2323) prototype campaign from a human collaborator. The task-scope and Report-block criteria therefore do not apply literally; I reviewed the diff against the scope the PR body declares (one tiny src change, bench/ experiment records, one ci.yml leg) and against AGENTS.md. The pipeline already handles non-task PRs after merge (claude-pipeline.yml, "Not a task PR").

Checked.

  • Protected files: PLAN.md, TASKS.md, Manifest.toml untouched. No public parameter or phase string changed; subtree_max_fronts joins TUNING_OPTIONS (an Options keyword, not a setparam! string, like subtree_parallelism; the test asserts setparam! rejects it with ArgumentError).
  • src: the diff is exactly what the PR claims: one Options field (default 0 = off, validated by _parse_int >= 0), one subtree front count in build_schedule eligibility pass. nfronts is monotone up the tree, so capped eligibility is (eligible && nfronts <= cap) and the maximal-subtree construction and regime assignment are unchanged. Host-side Int, no kernel or layout touched. copy/show iterate fieldnames, so the new field is covered.
  • Test (test_symbolic_schedule.jl): option validation, cap enforcement counted from subtree_ptr, subtree sizes equal to the tree counts, regime A only shrinks, maximality (parent over the cap or outside A), and a monotone subtree-count ordering. Uses the shared laplacian2d/thrown/check_schedule helpers; the file is not a SPLIT_FILES member, so no RUN_SHARED is needed; the children-before-parents numbering it relies on is a documented property of snparent.
  • CI: test-github-cpuonly and test-gpu (cuda) green on this head; Aqua and docs green (the docstring table is the documentation). The amdgpu leg is red under continue-on-error and the workflow run concluded success, so all_green in the pipeline is unaffected. From its log: the 2 failures (test_ubatch.jl:245, both element types) catch a MethodError from the missing cholesky(::ROCSparseMatrixCSR) forwarder, the same cause as the 172 errors, consistent with the PR comment.
  • bench/: CUDA/ROCm-only prototypes, as declared; separate environments (bench/amd, bench/e2e) with the package as a path source (same as the existing bench/Project.toml), no monkey-patching of the library, no machine-specific paths beyond the documented data dump. Not reviewed line-by-line as library code, per the PR framing as experiment records.

Blocking findings: none.

Non-blocking (inline): stale copied header in bench/amd_proto.jl; the build_schedule docstring misses the new eligibility condition; the AGENTS.md sentence saying the amdgpu runner is disabled is now stale.

For the owner (outside the diff): the PLAN.md regime-threshold row (the Options-keyword list) does not list subtree_max_fronts; and findings 1-3 and 7 of the PR body (ordering-cost scoring vs schedule depth, CB-slot lifetimes vs out-of-level-order execution, the host sync in _factor_panel_c!) live only in PR comments. AGENTS.md routes such findings to found-by-agent issues so the triage gate sees them before the next task starts.

@github-actions
github-actions Bot merged commit 2b031f3 into main Oct 9, 2026
9 of 10 checks passed
@github-actions
github-actions Bot deleted the proto/partitioned-inverse-solve branch October 9, 2026 12:02
michel2323 pushed a commit that referenced this pull request Oct 9, 2026
…in the plan, tasks and performance notes [skip ci]

PLAN.md: subtree_max_fronts in the regime-threshold row with the measured
knobs; the ordering cost model's failure on large KKT systems (#108); the
level-synchronous assumption of the stack placement (#109); the
partitioned-inverse / fused-solve prototype results in the solve section;
status paragraph and backends bullet (AMD leg, e2e run, 1.3x cuDSS prototypes).
TASKS.md: owner notes under T23 (AMD leg, constructor surface), T24 (native
ND vs the imported cuDSS ordering), T25 (ordered deliverables from #107 and
the negative results) and the External task (bench/e2e as starting point).
PERFORMANCE.md: tracked-issue rows for #107-#110, the "Experiments 5-6
prototypes" results section, status on plan rows 4-6, revised priority.
AGENTS.md: the amdgpu runner sentence (continue-on-error leg since #107).
README.md, bench/README.md, STATE.md: AMD/e2e status, the new bench scripts,
addendum on the pause's merged PRs.

Refs #107, #108, #109, #110.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014NfiFN5Ukdcnbh87uff2Dg
michel2323 added a commit that referenced this pull request Oct 9, 2026
…in the plan, tasks and performance notes [skip ci] (#111)

PLAN.md: subtree_max_fronts in the regime-threshold row with the measured
knobs; the ordering cost model's failure on large KKT systems (#108); the
level-synchronous assumption of the stack placement (#109); the
partitioned-inverse / fused-solve prototype results in the solve section;
status paragraph and backends bullet (AMD leg, e2e run, 1.3x cuDSS prototypes).
TASKS.md: owner notes under T23 (AMD leg, constructor surface), T24 (native
ND vs the imported cuDSS ordering), T25 (ordered deliverables from #107 and
the negative results) and the External task (bench/e2e as starting point).
PERFORMANCE.md: tracked-issue rows for #107-#110, the "Experiments 5-6
prototypes" results section, status on plan rows 4-6, revised priority.
AGENTS.md: the amdgpu runner sentence (continue-on-error leg since #107).
README.md, bench/README.md, STATE.md: AMD/e2e status, the new bench scripts,
addendum on the pause's merged PRs.

Refs #107, #108, #109, #110.


Claude-Session: https://claude.ai/code/session_014NfiFN5Ukdcnbh87uff2Dg

Co-authored-by: Claude <noreply@anthropic.com>
michel2323 added a commit that referenced this pull request Oct 9, 2026
…pe panel [skip ci]

Re-ran the SDS side of compare.jl on the RTX 4080 at ebfb730 (89 rows, all
ok; cuDSS rows unchanged from 2026-10-07). src/ is unchanged in effect since
#102, so the per-feature bars move only by run-to-run noise.

compare_report.jl adds a second panel and a comparison.md table for the PR #107
prototypes on the 78k-bus condensed KKT (GV100): solve 4.74x -> 0.92x cuDSS,
refactorization 6.90x -> 1.31x, MadNLP iteration loop 2.66x -> 1.19x. The
numbers are quoted from PERFORMANCE.md into comparison/prototype_pr107.csv, not
measured by compare.jl; the panel and README caption say so. Validated
colorblind-safe phase colours, device and SHA in the plot title.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y39AMnnixrVPFivvwBLFpP
michel2323 added a commit that referenced this pull request Oct 9, 2026
…tasks T22-T30 [skip ci]

The former T25 (performance pass) carried six deliverables from the #107
prototypes, several sessions of work under the one-task-per-session rule.
TASKS.md now has: T22 ordering chooser scoring the supernodal schedule depth
(#108, host only, precondition for the fused solve); T23 backends (unchanged);
T24 partitioned-inverse fused solve, solve_alg = "algo1" (#82 solve half);
T25 split factorization (tiled SYRK, chunked TRSM, split regime-C fronts,
longest-first regime A); T26 segmented fused factorization with dependency
counters, with #109 and #110 inside it plus the #75 remainder and the #96
re-measurement; T27 hybrid memory (was T26); T28 robustness extras (was T27,
with the FP32 finding of #107); T29 non-uniform batch (was T22); T30 ND
partition-tree export and ordering cache (was T24); External unchanged.
Each new task names the bench prototype it adopts, the measured numbers, the
negative results not to re-run, the tests and the Report asks.

PLAN.md, PERFORMANCE.md, bench/README.md: task-id references follow the
renumbering; the stale "empty repository" status line and the M1 row's
ordering sentence are updated. STATE.md: section 8 records the restructuring
and the GitHub steps (issue renames, triage labels, chain resumes with T22).

Refs #96, #108, #109, #110.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wzx8WtMEXnAfyzyujdySFh
michel2323 added a commit that referenced this pull request Oct 9, 2026
…tasks T22–T31 [skip ci] (#114)

* Docs: split T25 into T22/T24-T26 after PR #107, renumber the post-v1 tasks T22-T30 [skip ci]

The former T25 (performance pass) carried six deliverables from the #107
prototypes, several sessions of work under the one-task-per-session rule.
TASKS.md now has: T22 ordering chooser scoring the supernodal schedule depth
(#108, host only, precondition for the fused solve); T23 backends (unchanged);
T24 partitioned-inverse fused solve, solve_alg = "algo1" (#82 solve half);
T25 split factorization (tiled SYRK, chunked TRSM, split regime-C fronts,
longest-first regime A); T26 segmented fused factorization with dependency
counters, with #109 and #110 inside it plus the #75 remainder and the #96
re-measurement; T27 hybrid memory (was T26); T28 robustness extras (was T27,
with the FP32 finding of #107); T29 non-uniform batch (was T22); T30 ND
partition-tree export and ordering cache (was T24); External unchanged.
Each new task names the bench prototype it adopts, the measured numbers, the
negative results not to re-run, the tests and the Report asks.

PLAN.md, PERFORMANCE.md, bench/README.md: task-id references follow the
renumbering; the stale "empty repository" status line and the M1 row's
ordering sentence are updated. STATE.md: section 8 records the restructuring
and the GitHub steps (issue renames, triage labels, chain resumes with T22).

Refs #96, #108, #109, #110.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wzx8WtMEXnAfyzyujdySFh

* Docs: T31 for the oneAPI/Metal remainder of T23 (#112, #53) [skip ci]

PR #113 delivers the AMDGPU extension and the local-memory cap alone and
opened #112 for the rest of T23. The remainder becomes T31 at the end of the
chain: oneAPI and Metal extensions, the #53 allocation-free :ka fallbacks and
the select_impl order noted in #112. STATE.md section 8 updated.

Refs #112, #53.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wzx8WtMEXnAfyzyujdySFh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
sshin23 added a commit that referenced this pull request Oct 10, 2026
METIS_OPTION_NSEPS (candidate separators per split) is the one NodeND knob
that reproducibly improves the ACOPF KKT orderings: on
kkt_pglib_opf_case78484_epigrids_condensed_10 (n = 157k), nd_nseps = 4
gives nnz(L) 14.42M vs 14.76M (-2.3%), factorization flops -6.7%, and the
PR #107 fused refactorization 30.1 ms vs 32.1 ms (-6.2%, interleaved A/B
on a GV100, correctness-gated) at the same schedule depth. The ordering
itself runs ~2x slower, so the default stays at the METIS default (-1);
the option pays off when one symbolic analysis is reused across many
factorizations and is exposed like nd_ubfactor.

ND_PROVIDER gains a third argument (values < 1 fall back to the library
default); validated like the other config ints and covered in
test_options and test_symbolic_etree.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

claude:pr PR opened by the Claude implementer; handled by the Claude pipeline

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants