Repository navigation
Performance prototypes: partitioned-inverse solve, split/fused factorization, ordering study (experiments 0/5/6) - #107
Conversation
…ter sweeps Experiments 5/6 of PERFORMANCE.md (issues #82, #25), outside src/: after factorization every front's L11 (w <= 256) is inverted into a side buffer so each front's TRSV becomes one GEMV; the forward and backward sweeps then run as one fused kernel (regime-A subtrees merged as leading/trailing blocks, etree child-counters with spin/threadfence signaling) plus one single-block chain kernel for the near-serial tall fronts at the top, with the vendor dense path only on the remaining wide roots. CUDA-only, SPD, single RHS. On a Quadro GV100 the pglib_opf_case78484_epigrids condensed KKT solve goes from 18.0 ms (stock sweeps) to 6.3 ms, against 3.8 ms for cuDSS; the full SPD harness table is in the PR description, including the matrices where the approach regresses and a per-schedule strategy pick is needed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Extracts cuDSS's reordering permutation (raw cudssDataGet, the high-level getter refuses it) and replays the prototype solve under user_perm, plus an amalgamation sweep. On the 78k condensed KKT the cuDSS ordering halves the elimination-tree depth (65 -> 30 levels) and takes the prototype solve from 6.35 ms to 4.81 ms on a GV100; amalgamation changes only regress. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-up: the reordering is a third of the remaining gapExtracting cuDSS's reordering permutation (raw
cuDSS's custom ND halves the elimination-tree depth at +12% fill, and on a latency-bound solve that is worth 1.3×: the prototype goes to 4.81 ms against cuDSS's 3.77 ms (1.27× remaining). Supernode counts are nearly identical, so this is separator/recursion shape, not amalgamation — and indeed an amalgamation sweep ( Two implications for the plan:
🤖 Generated with Claude Code |
…n the 78k KKT With the cuDSS reordering (depth-30 etree) the two-tier split is obsolete: one fused dependency-counter kernel per direction covering every front of width <= 256 solves the 78k condensed KKT in 3.61 ms on a GV100, against 3.77 ms for cuDSS and 18.0 ms for the stock sweeps. Under the METIS ordering the same configuration reads 5.64 ms, so the shallow ordering and the single-tier structure are both required. Also fixes a latent correctness bug in the prototypes: build_fused keyed the live-subtree-root accounting on the width cap rather than on whether the segment is the merged tier, so a tier-1 cap wider than 128 double-counted subtree-root signals and released waiters early (relres 5e+11). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Final round: single-tier fused solve under the cuDSS ordering — 3.61 ms vs cuDSS 3.77 msThe two-tier structure (fused counters below, single-block chain above) was tuned on the depth-65 METIS tree. Under the cuDSS ordering (depth 30) the chain is obsolete: one fused dependency-counter kernel per direction covering every front of width ≤ 256 is both simpler and faster. Order-controlled, 3 repeats, relative residual equal to stock/cuDSS:
Single-tier under METIS reads 5.64 ms — the shallow ordering and the single-tier structure are only fast together. The entire solve is now two custom kernels + the vendor root + two permute kernels. The sweep also flushed out a latent correctness bug, fixed in this branch: Caveats as before: tuned for KKT-shaped trees; the ordering is imported from a one-time cuDSS analysis of the same pattern ( 🤖 Generated with Claude Code |
…x906) Portability shims only: ROCArray + the CSR-triplet DirectSolver constructor (no AMDGPU convenience extension exists yet, T23), and the device-scope fence as Threads.atomic_fence(), which both GPU backends lower to an LLVM fence. Atomix atomics, the dependency-counter spinning and the dispatch-order deadlock-freedom assumption all hold on gfx906. Radeon VII, 78k condensed KKT, 1 RHS, median of 20 warm, residuals matching the stock sweeps: stock 20.5 ms -> 8.05 ms (single tier, cuDSS ordering); under METIS 28.5 -> 12.2 ms. Same structural ranking as on the GV100, where cuDSS does not exist as a comparison point at all. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
AMDGPU: it runs, first tryRadeon VII (gfx906, ROCm), same matrix, same protocol, residuals matching stock:
Portability cost was three mechanical substitutions ( There is no cuDSS on this hardware — this is the portability case of PLAN/PERFORMANCE.md made concrete. 🤖 Generated with Claude Code |
… 78k KKT The stock fused regime-B kernel keeps each front's trsm and contribution-block update inside the front's single workgroup, which serializes 1-2 ms per fat front near the root. This prototype keeps assembly + F11 Cholesky one-block and moves the trsm (128-row chunks) and the syrk (32x32 tiles, one workgroup each, disjoint writes, deterministic) of fronts with m > 64 to separate many-block kernels; the factor is bit-identical to stock on every config. With the routing knobs that fall out of the same measurements (fat fronts to the vendor path via regime_c_rows = 256, subtree serial walks capped via subtree_parallelism = 16384, cuDSS ordering), refactorization of the 78k condensed KKT goes 167.7 -> 64.0 ms on a GV100 (cuDSS: 24.3 ms). The tiled syrk does the dominant flops in 1.2 ms where the fused kernel spent ~12. CUDA-graph replay of the whole sequence is validated bit-identical and timing-neutral: the remainder is kernel-busy time (vendor dense calls for the 303 C fronts, the one-block panel kernels, the longest regime-A subtrees and the level-batched extend-add), not launch overhead. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Factorization round: 167.7 → 64.0 ms (cuDSS 24.3), factor bit-identicalSame method as the solve: measure, route, split.
Findings:
Per IPM iteration the prototype stack now reads: refactorization + solve = 64.0 + 3.6 = 67.6 ms vs cuDSS 24.3 + 3.8 = 28.1 ms, from 185 ms stock. 🤖 Generated with Claude Code |
…-> 40.6 ms The vendor dense path cost ~26 ms in launch-and-API overhead for 303 tall-narrow fronts (~8 launches each). Those fronts keep the stock multi-block assembly kernels but now factor through the split kernels batched per level: a no-assembly variant of the panel kernel (w <= 64), the 128-row trsm chunks and the 32x32 syrk tiles; the vendor path keeps only fronts wider than 64. With that split, lowering regime_c_rows monotonically helps (64 is the floor), and subtree launches run longest-first through a replicated launcher. A tiled extend-add was implemented and measured neutral: that bucket is scatter bandwidth, not parallelism, and is at its floor. GV100, 78k condensed KKT: refactorization 40.6 ms (was 64.0 after round 1, 167.7 stock; cuDSS 24.3). Relative factor difference to stock is ~3e-10 where vendor arithmetic was replaced, residuals unchanged. Per IPM iteration the prototype stack is now 40.6 + 3.6 = 44.2 ms against cuDSS's 28.1, from 185 ms stock. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Factorization round 2: 64.0 → 40.6 ms (cuDSS 24.3)
The C-front finding is the transferable one: the vendor dense path's cost on tall-narrow fronts is ~8 launches of API overhead per front, while its assembly kernels (multi-block zero/scatter/extend-add) are the efficient part — the exact inverse of regime B, whose fused kernel has cheap assembly access but serializes the factor math. The split takes the best half of each. Also measured and worth recording as negatives: a tiled extend-add is timing-neutral (that bucket is scatter bandwidth at ~10 ms, a real floor), and CUDA-graph replay remains neutral throughout. Factor matches stock to ~3e-10 relative (vendor arithmetic replaced for ex-C fronts; bit-identical on B-only configs). Remaining buckets: subtrees ~12 ms (longest-walk bound), panel ~8, wide-root vendor ~8, extend-add ~9 — in-src territory (persistent kernels, batched wide fronts). Prototype stack per IPM iteration: 44.2 ms vs cuDSS 28.1 ms, from 185 ms stock (4.2×). 🤖 Generated with Claude Code |
Four directions measured and closed on the 78k condensed KKT (GV100): - amalgamation under the split kernels: all variants regress (default 32/0.25/8 optimal, as under the stock kernels); - overlapping the subtree budget-class launches on a second stream: neutral (the large class saturates the device) and it invalidates graph capture; - Float32 factorization: the kernels are generic now and run 18% faster, but the matrix has kappa ~ 1e14 (FP64 relres 1.4e-2 ~ kappa * eps), so FP32 factors are unusable, Jacobi scaling does not change the pivot failure (same failed column), and even delta = 1e-6 regularization in FP64 is unusable for the original system (relres 2.1) - consistent with cuDSS's documented FP32 failures on the condensed dumps; - the subtree launcher now sizes local memory by sizeof(eltype) (a latent Float64 assumption that produced illegal addresses under Float32). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Factorization round 3: boundary mapping — 40.6 ms standsFour directions measured to closure (all in
The conditioning argument is worth keeping: FP64's single-solve relres of 1.4e-2 on these matrices is itself the κ·ε floor, so no precision trick at the solver level can help — accuracy beyond that needs refinement against an FP64 factor, which is already the design (T16/T18). Final prototype scoreboard (GV100, 78k condensed KKT, per IPM iteration): 44.2 ms (refactorization 40.6 + solve 3.6) vs cuDSS 28.1 ms, from 185 ms stock — 4.2×, remaining gap 1.57×, with the solve itself faster than cuDSS. Remaining distance is in-src work: persistent/merged kernels for the subtree and panel buckets, batched wide-front processing, and the scatter-bandwidth extend-add floor. 🤖 Generated with Claude Code |
One fused kernel per segment whose workgroups play four roles in a topologically ordered grid (extend-add chunks, panels, trsm chunks, syrk tiles) with etree counters; segments split at the groups containing wide fronts, which run between segments through _factor_panel_c!. Verified: zero failing fronts, factor matches the split pipeline (~3e-10, vendor-arithmetic class), residuals unchanged, stable over repeats and rows=64/96. Two findings that outlive the prototype: - Full overlap (variant i: wides on a concurrent stream behind spin-waiters) DEADLOCKS structurally on this tree: 62 wide fronts appear early, 31.5k of 34.7k blocks depend on them, and spinning dependents saturate every SM before the vendor stream can run. Measured, not hypothesized; the segmented form is the deadlock-free shape of this idea. - The update stack's contribution-block placement (PR #54 lifetimes) assumes level-synchronous execution. Any out-of-level-order execution -- counters, streams, or persistent kernels -- clobbers live CBs through slot reuse, and a global zero pass does the same. The prototype gives every non-subtree front a private CB slot instead: 0.14 GB here (~4x the reusing stack), which is the memory price of experiment 5 unless placement learns DAG lifetimes. Concurrent extend-add children additionally need atomic adds (the stock owner-pull order serializes children; the fused grid does not). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Factorization round 3 (variant i → segmented): 40.6 → 33.5 ms (cuDSS 24.3)Built at the owner's request as the full-overlap experiment; the honest outcome is a working segmented form plus two structural findings that any in-src experiment-5 implementation will need:
Stack now: refactorization 33.5 + solve 3.6 = 37.1 ms per IPM iteration vs cuDSS 28.1, from 185 ms stock — 5.0×, remaining gap 1.32×. 🤖 Generated with Claude Code |
The cuDSS-ordering advantage was never about cuDSS's partitioning: forcing SDS's own nested dissection (reordering_alg = "algo4") yields schedule depth 29 vs the import's 30, factorization 33.1 ms vs 34.8, solve 3.5 vs 3.7 on the 78k condensed KKT. The depth-65 default came from the automatic AMD/ND chooser: ordering_cost = flops * (1 + nlevels/n) weighs depth at ~0.2% for n = 674562, so AMD wins on flops — and nlevels is COLUMN-etree depth, which does not predict the supernodal schedule depth that governs GPU time (cuDSS perm: 1279 column levels -> 30 schedule levels; METIS: 1216 -> 29; both shallower at column level than their schedule ratio suggests). A METIS knob sweep (seed, ufactor, nseps, ccorder, pfactor) moves schedule depth 28-41 with seed-dependent noise; nothing beats plain ND meaningfully. Consequences: the prototype stack no longer depends on cuDSS for anything (relevant for AMDGPU, where cuDSS does not exist), and the ordering chooser should score supernodal schedule depth (or simply prefer ND on GPU backends). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Ordering study (proposal E):
|
| ordering (78k condensed KKT, GV100) | schedule depth | refactorization | solve |
|---|---|---|---|
| SDS default (auto → AMD) | 65 | — baseline configs — | — |
| imported cuDSS reordering | 30 | 34.8 ms | 3.7 ms |
native ND (algo4) |
29 | 33.1 ms | 3.5 ms |
| METIS seed sweep best | 28 | 33.2 ms | 3.6 ms |
Root cause of the depth-65 default, in two parts: (1) ordering_cost = flops × (1 + nlevels/n) weighs depth at ~0.2% for n = 674k, so the AMD/ND chooser is a pure flop contest that AMD wins; (2) nlevels is column-etree depth, which doesn't predict the supernodal schedule depth that governs GPU time — the cuDSS perm has more column levels than METIS (1279 vs 1216) yet the same schedule depth, and AMD's schedule is 2× deeper despite competitive column metrics. Recommendation: score candidate orderings by supernodal schedule depth (cheap to compute from the supernode partition), or simply prefer ND on GPU backends.
Practical consequences: the whole prototype stack is now cuDSS-free — one setparam!(solver, "reordering_alg", "algo4") — which also makes the AMDGPU path self-contained. user_nd_partition_tree import was checked and is T24-pending in SDS; the search shows it isn't needed. (En route, the per-trial residual gate caught a stale copy of the early-release counter bug a second time — fast-but-wrong remains this work's defining hazard.)
Scoreboard, now cuDSS-free: refactorization 33.1 + solve 3.5 ≈ 36.6 ms per IPM iteration vs cuDSS 28.1, from 185 ms stock — 5.1×.
🤖 Generated with Claude Code
A child whose contribution block has exactly one consumer writes it straight into the parent's frontal matrix from its syrk tiles, skipping the CB round-trip. Safety conditions measured out stricter than estimated: the parent must be pre-assembled (C-group; B parents zero in-kernel and would wipe the contribution), ALL CB children count as siblings (subtree-root children extend-add concurrently, so 'only child' must include them), and wide children are excluded (they have no tiles to settle the ea_left debt -- the first version deadlocked on exactly that under one ordering). The eligible set is then ~80 fronts / ~1.4M CB entries (~8% of CB traffic, not the ~30% a looser census suggested), worth ~0.1 ms: 33.07 -> 32.94 ms under native ND, within noise. Kept because it is correct and free; recorded because the census-first discipline is what kept it from being sold as a win. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Proposal B (write-through extend-add): correct, kept, marginalImplemented in the fused kernel: a child with a single CB consumer writes its finished contribution straight into the parent's frontal matrix from its syrk tiles. The safety conditions turned out stricter than the back-of-envelope: parent pre-assembled (C-group only), siblings counted including subtree-root children (they extend-add concurrently), and wide children excluded — a wide child has no tiles to settle the dependency debt, and the first version deadlocked on exactly that under one ordering (caught by the timeout guard, fixed by one condition). Eligible set after the honest conditions: ~80 fronts, ~1.4M CB entries ≈ 8% of CB traffic (not the ~30% a looser census suggested) → 33.07 → 32.94 ms under native ND, within noise. Kept because it's correct and free. Running scoreboard (GV100, 78k condensed KKT, cuDSS-free native-ND stack): refactorization 32.9 ms + solve 3.5 ms ≈ 36.4 ms per IPM iteration vs cuDSS 28.1, from 185 ms stock. 🤖 Generated with Claude Code |
…-neutral Adds an analysis tuning option capping the number of fronts a regime-A subtree may contain (0 = no limit, the default; the serial walk of a subtree is its workgroup's critical path). One field in Options, one count in the eligibility pass of build_schedule, validated like subtree_parallelism. Measured on the 78k condensed KKT (GV100, native ND, fused prototype): 33.03 ms uncapped vs 32.77 at cap 32 — within noise. The sweep disproves the longest-walk-floor hypothesis for this configuration: with ~3.3k subtrees the launch is throughput-bound, and the 188-front walk hides inside the waves. The knob is kept because it is the right instrument to re-test that on machines with more SMs or trees with fewer subtrees, where the balance tips. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Proposal C (subtree front-count cap): implemented as
|
…s mapped Gain: subtree_budgets = [16384] raises the subtree kernel's occupancy (the 48 KB class caps it at 1 block/SM) and is the new best, 33.0 -> 31.9 ms (single-class launch, 8.4 ms bucket from 13.3). Three negatives with mechanisms, kept as knobs/record: - packed-triangle extend-add iteration regresses 33 -> 42: the per-element column unpack is a serial O(m) walk, costlier than the idle upper-half threads it removes (fine once per thread in the trsm chunk kernel, fatal inside the element loop); - replacing the vendor path for 64 < w <= 96 fronts with 96-class localmem kernels regresses 32 -> 42: cuSOLVER's blocked potrf beats a column-serial localmem Cholesky above w ~ 64, i.e. the existing regime-B width boundary is in exactly the right place; - a stream pool for the independent boundary wides is neutral: the _factor_panel_c! info check synchronizes the host per call, so overlap needs an async-info vendor path in src, not more streams. Remaining to cuDSS (24.3): mega-kernel internals (11.0), subtrees (8.4), vendor dense (10.1) — each now requires src-grade work (role rebalancing, subtree walk redesign, async potrf info). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
33 → 24 push: landed 31.9 ms; the rest is src-gradeOne gain: Three negatives, each with its mechanism (details in Remaining decomposition at 31.9: mega-kernel 11.0 + subtrees 8.4 + vendor dense ~10 + stats 0.6. Closing to 24.3 needs src work on all three: role-level rebalancing inside the fused kernel, the subtree walk's per-front efficiency, and async vendor info handling. Bench-level levers are exhausted — effect sizes are now inside run-to-run noise. Final factorization: 167.7 → 31.9 ms (5.3×), 1.31× from cuDSS, cuDSS-free. Per IPM iteration with the solve: ~35.4 ms vs cuDSS 28.1, from 185 stock. 🤖 Generated with Claude Code |
Its copy of build_fused predates the merged-tier fix (the early-release counter bug caught twice by the residual gate); solve_nd.jl carries the same experiment on the fixed machinery. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…(CUDA and AMDGPU) MadNLPSDS.jl implements MadNLP.AbstractLinearSolver over the public SDS handle API, modeled on MadNLPGPU 0.10's CUDSSSolver: device-CSC input read as the CSR of the upper triangle (view 'U'), values aliased with an update! fallback, in-place solve_linear_system!, Cholesky info mapped to inertia. Backend-agnostic: the same ~90 lines drive CUSPARSE and rocSPARSE matrices. Full MadNLP on the 78k-bus ACOPF (SparseCondensedKKTSystem, tol 1e-6): identical iteration count (101), objective and inertia behavior across all three configurations. GV100: cuDSS 16.8 s wall (9.7 s iteration loop) vs SDS-with-stock-kernels 38.8 s (26.1 s loop) — the 2.7x is the stock-kernel ratio documented in this PR; the prototype kernels are not wired into src. AMD Radeon VII: 71.0 s wall (61.3 s loop) — to our knowledge the first GPU-resident sparse-direct ACOPF interior-point run on AMD hardware, where no cuDSS exists; this wrapper is the missing sparse solver for MadNLPGPU's AMDGPU extension. One pitfall recorded: linear-solver tuning knobs are phase-coupled (regime_c_rows = 64 is right for the prototype factorization kernels but puts ~300 fronts on the stock solve's vendor path, 215 ms per backsolve; 256 is the stock-balanced setting). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
End-to-end: full MadNLP ACOPF with SDS, on CUDA and on AMD
The GV100 gap is exactly the stock-kernel ratio measured throughout this PR — the prototype kernels live in One operational finding: the tuning knobs are phase-coupled — 🤖 Generated with Claude Code |
SDSProto.jl assembles this PR's solve and factorization prototypes
(partitioned-inverse + fused counter sweeps; segmented fused factorization
with private contribution blocks) into one backend-portable module
(KernelAbstractions allocation, Threads.atomic_fence), and SDSProtoSolver
drives them from MadNLP. Because both phases are prototype kernels, the
phase-coupled routing conflict disappears (rows = 64 is optimal for both).
Full MadNLP, 78k-bus ACOPF, identical 101 iterations and objective in every
configuration:
wall iteration loop
cuDSS / GV100 16.9 s 9.8 s
stock / GV100 38.8 s 26.1 s
PROTO / GV100 25.5 s 11.7 s (1.2x from cuDSS end-to-end)
stock / RadeonVII 72.8 s 63.2 s
PROTO / RadeonVII 43.6 s 33.9 s (1.9x over stock on AMD)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Capstone: the PR's prototype kernels inside MadNLP, on both vendors
Full interior-point, 78k-bus ACOPF, identical 101 iterations and objective everywhere:
The loop arithmetic closes: 193 factorizations × ~36 ms + 202 solves × ~4 ms + wrapper sync ≈ 11.7 s. The remaining GV100 distance to cuDSS is the documented factorization gap (31.9 vs 24.3 ms) plus per-call synchronization that an in-src integration does asynchronously. On AMD this is now a practical end-to-end path — 34 s of iteration loop for a 78k-bus ACOPF on a 2019 consumer card, with no vendor sparse solver in existence there. 🤖 Generated with Claude Code |
…lity The src knob from this branch now carries its test, in the house style of the subtree_parallelism testset: Options validation (default, copy, range, not-a-setparam-string), every subtree within the cap as counted from the schedule's own subtree_ptr spans and matching the independently accumulated etree front counts, regime A only shrinking, and maximality (a subtree root's parent is over the cap or was outside regime A before it). Green on the CPU suite locally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The test suite already includes ROCBackend when AMDGPU.jl is functional (test/backends.jl), and the job template already installs AMDGPU.jl for the amdgpu matrix entry — only the matrix entry itself was disabled. Run it as a non-required continue-on-error leg: it exercises the KernelAbstractions kernels on AMD hardware today (the bench prototypes on this branch validated that path end-to-end on a gfx906), and T23 promotes it to a required check. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First AMD CI run: 99.4% pass; T23 is ~a constructor extensionThe The errors are essentially one missing method, repeated: 158× The leg stays continue-on-error on this branch; with T23's constructors it should go fully green and can be promoted to a required check. 🤖 Generated with Claude Code |
| # Prototype (PERFORMANCE.md experiments 5/6, issues #82/#25): partitioned-inverse | ||
| # solve with fused dependency-counter sweeps, CUDA only, SPD, 1 RHS, nbatch = 1. | ||
| # | ||
| # After factorization, invert every front's L11 (w <= MW) into a side buffer; | ||
| # each front's TRSV becomes one GEMV. The sweeps then run as: | ||
| # fwd: ONE fused kernel (regime-A subtrees as leading blocks + every front of | ||
| # width <= 128 below the first wide front, etree child-counters with | ||
| # spin/threadfence signaling, atomic L21 scatter) | ||
| # + ONE single-workgroup "chain" kernel walking the near-serial tall | ||
| # fronts (w <= 256) at the top with plain block barriers | ||
| # + the stock vendor dense path for the remaining wide roots. | ||
| # bwd: the same in reverse (chain waits dispatch parents-first: fronts MUST | ||
| # get reverse-topological block ids or the kernel deadlocks). | ||
| # Measured on a Quadro GV100, pglib_opf_case78484_epigrids condensed KKT | ||
| # (n = 674562): stock sweeps 18.0 ms -> 6.3 ms; cuDSS 3.8 ms. See the PR body | ||
| # for the full matrix table (wide-front matrices regress; pick per schedule). | ||
| # Constraints: CHWG >= NSTRIP * MW (strip indexing); NSTRIP/MW compile-time. | ||
| # | ||
| # CUDA_VISIBLE_DEVICES=1 julia +1.13 --project=bench bench/solve_proto_78k.jl |
There was a problem hiding this comment.
nit (non-blocking): this header is a verbatim copy of bench/solve_proto_78k.jl's. Line 2 says "CUDA only" and line 19 gives the CUDA script's invocation (--project=bench bench/solve_proto_78k.jl), but this file is the ROCm port that runs under bench/amd/Project.toml. Suggest replacing the two lines with something like "ROCm port of solve_proto_78k.jl (ROCArray, CSR-triplet constructor, KA fence), SPD, 1 RHS" and julia +1.13 --project=bench/amd bench/amd_proto.jl, so the AMD numbers in the PR body stay reproducible from the file alone.
| eligible[s] = !big[s] && all(c -> eligible[c], kids) && peak[s] <= maxcap && work[s] <= maxwork && | ||
| nfronts[s] <= maxfronts |
There was a problem hiding this comment.
nit (non-blocking): the logic is right (nfronts is monotone up the tree, so eligible_cap[s] == eligible[s] && nfronts[s] <= maxfronts and the maximal subtrees stay well defined, which is exactly what the new test checks). But the build_schedule docstring above (item 2, "regime A: a front is eligible when …") still lists only the C-threshold, the stack-peak and the flop-limit conditions. Add the front-count limit there, e.g. "… and its subtree has at most opts.subtree_max_fronts fronts (0: no limit)", so the docstring and the Options table agree on what makes a front eligible.
| # 'amdgpu' runs as a non-required, continue-on-error leg until the | ||
| # AMDGPU extension exists (T23): the test suite already includes | ||
| # ROCBackend when AMDGPU is functional (test/backends.jl), so this | ||
| # exercises the KernelAbstractions kernels on AMD hardware today. | ||
| # Promote it to the required checks of the branch ruleset with T23. | ||
| backend: ['cuda', 'amdgpu'] | ||
| runs-on: [self-hosted, linux, X64, gpu, '${{ matrix.backend }}'] | ||
| continue-on-error: ${{ matrix.backend == 'amdgpu' }} |
There was a problem hiding this comment.
nit (non-blocking): the leg works as intended — I checked that run 37874446430 concluded success with the amdgpu job red, so claude-pipeline.yml's all_green/ci_failed logic (workflow-run conclusion) is unaffected and no CI-fix round is triggered by the AMD leg. Two follow-ups for whoever touches this next: (1) AGENTS.md ("Environment → GitHub Actions", the sentence "the amdgpu runner is disabled in ci.yml until the AMDGPU extension exists, T23") is now stale and should say the leg runs non-required/continue-on-error and only needs promoting to a required check with T23; (2) the 2 failures + 172 errors of this run are all the missing ROC constructor surface — the two test_ubatch.jl:245 failures catch a MethodError from cholesky(::ROCSparseMatrixCSR; view, check) instead of a FactorizationError, same cause as the errors — so the PR comment's "T23 is ~a constructor extension" reading is confirmed by the log.
There was a problem hiding this comment.
VERDICT: APPROVE (head f55e3a8)
What this PR is. Not a TNN task of TASKS.md: an owner-labelled (claude:pr by michel2323) prototype campaign from a human collaborator. The task-scope and Report-block criteria therefore do not apply literally; I reviewed the diff against the scope the PR body declares (one tiny src change, bench/ experiment records, one ci.yml leg) and against AGENTS.md. The pipeline already handles non-task PRs after merge (claude-pipeline.yml, "Not a task PR").
Checked.
- Protected files: PLAN.md, TASKS.md, Manifest.toml untouched. No public parameter or phase string changed; subtree_max_fronts joins TUNING_OPTIONS (an Options keyword, not a setparam! string, like subtree_parallelism; the test asserts setparam! rejects it with ArgumentError).
- src: the diff is exactly what the PR claims: one Options field (default 0 = off, validated by _parse_int >= 0), one subtree front count in build_schedule eligibility pass. nfronts is monotone up the tree, so capped eligibility is (eligible && nfronts <= cap) and the maximal-subtree construction and regime assignment are unchanged. Host-side Int, no kernel or layout touched. copy/show iterate fieldnames, so the new field is covered.
- Test (test_symbolic_schedule.jl): option validation, cap enforcement counted from subtree_ptr, subtree sizes equal to the tree counts, regime A only shrinks, maximality (parent over the cap or outside A), and a monotone subtree-count ordering. Uses the shared laplacian2d/thrown/check_schedule helpers; the file is not a SPLIT_FILES member, so no RUN_SHARED is needed; the children-before-parents numbering it relies on is a documented property of snparent.
- CI: test-github-cpuonly and test-gpu (cuda) green on this head; Aqua and docs green (the docstring table is the documentation). The amdgpu leg is red under continue-on-error and the workflow run concluded success, so all_green in the pipeline is unaffected. From its log: the 2 failures (test_ubatch.jl:245, both element types) catch a MethodError from the missing cholesky(::ROCSparseMatrixCSR) forwarder, the same cause as the 172 errors, consistent with the PR comment.
- bench/: CUDA/ROCm-only prototypes, as declared; separate environments (bench/amd, bench/e2e) with the package as a path source (same as the existing bench/Project.toml), no monkey-patching of the library, no machine-specific paths beyond the documented data dump. Not reviewed line-by-line as library code, per the PR framing as experiment records.
Blocking findings: none.
Non-blocking (inline): stale copied header in bench/amd_proto.jl; the build_schedule docstring misses the new eligibility condition; the AGENTS.md sentence saying the amdgpu runner is disabled is now stale.
For the owner (outside the diff): the PLAN.md regime-threshold row (the Options-keyword list) does not list subtree_max_fronts; and findings 1-3 and 7 of the PR body (ordering-cost scoring vs schedule depth, CB-slot lifetimes vs out-of-level-order execution, the host sync in _factor_panel_c!) live only in PR comments. AGENTS.md routes such findings to found-by-agent issues so the triage gate sees them before the next task starts.
…in the plan, tasks and performance notes [skip ci] PLAN.md: subtree_max_fronts in the regime-threshold row with the measured knobs; the ordering cost model's failure on large KKT systems (#108); the level-synchronous assumption of the stack placement (#109); the partitioned-inverse / fused-solve prototype results in the solve section; status paragraph and backends bullet (AMD leg, e2e run, 1.3x cuDSS prototypes). TASKS.md: owner notes under T23 (AMD leg, constructor surface), T24 (native ND vs the imported cuDSS ordering), T25 (ordered deliverables from #107 and the negative results) and the External task (bench/e2e as starting point). PERFORMANCE.md: tracked-issue rows for #107-#110, the "Experiments 5-6 prototypes" results section, status on plan rows 4-6, revised priority. AGENTS.md: the amdgpu runner sentence (continue-on-error leg since #107). README.md, bench/README.md, STATE.md: AMD/e2e status, the new bench scripts, addendum on the pause's merged PRs. Refs #107, #108, #109, #110. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014NfiFN5Ukdcnbh87uff2Dg
…in the plan, tasks and performance notes [skip ci] (#111) PLAN.md: subtree_max_fronts in the regime-threshold row with the measured knobs; the ordering cost model's failure on large KKT systems (#108); the level-synchronous assumption of the stack placement (#109); the partitioned-inverse / fused-solve prototype results in the solve section; status paragraph and backends bullet (AMD leg, e2e run, 1.3x cuDSS prototypes). TASKS.md: owner notes under T23 (AMD leg, constructor surface), T24 (native ND vs the imported cuDSS ordering), T25 (ordered deliverables from #107 and the negative results) and the External task (bench/e2e as starting point). PERFORMANCE.md: tracked-issue rows for #107-#110, the "Experiments 5-6 prototypes" results section, status on plan rows 4-6, revised priority. AGENTS.md: the amdgpu runner sentence (continue-on-error leg since #107). README.md, bench/README.md, STATE.md: AMD/e2e status, the new bench scripts, addendum on the pause's merged PRs. Refs #107, #108, #109, #110. Claude-Session: https://claude.ai/code/session_014NfiFN5Ukdcnbh87uff2Dg Co-authored-by: Claude <noreply@anthropic.com>
…pe panel [skip ci] Re-ran the SDS side of compare.jl on the RTX 4080 at ebfb730 (89 rows, all ok; cuDSS rows unchanged from 2026-10-07). src/ is unchanged in effect since #102, so the per-feature bars move only by run-to-run noise. compare_report.jl adds a second panel and a comparison.md table for the PR #107 prototypes on the 78k-bus condensed KKT (GV100): solve 4.74x -> 0.92x cuDSS, refactorization 6.90x -> 1.31x, MadNLP iteration loop 2.66x -> 1.19x. The numbers are quoted from PERFORMANCE.md into comparison/prototype_pr107.csv, not measured by compare.jl; the panel and README caption say so. Validated colorblind-safe phase colours, device and SHA in the plot title. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y39AMnnixrVPFivvwBLFpP
…tasks T22-T30 [skip ci] The former T25 (performance pass) carried six deliverables from the #107 prototypes, several sessions of work under the one-task-per-session rule. TASKS.md now has: T22 ordering chooser scoring the supernodal schedule depth (#108, host only, precondition for the fused solve); T23 backends (unchanged); T24 partitioned-inverse fused solve, solve_alg = "algo1" (#82 solve half); T25 split factorization (tiled SYRK, chunked TRSM, split regime-C fronts, longest-first regime A); T26 segmented fused factorization with dependency counters, with #109 and #110 inside it plus the #75 remainder and the #96 re-measurement; T27 hybrid memory (was T26); T28 robustness extras (was T27, with the FP32 finding of #107); T29 non-uniform batch (was T22); T30 ND partition-tree export and ordering cache (was T24); External unchanged. Each new task names the bench prototype it adopts, the measured numbers, the negative results not to re-run, the tests and the Report asks. PLAN.md, PERFORMANCE.md, bench/README.md: task-id references follow the renumbering; the stale "empty repository" status line and the M1 row's ordering sentence are updated. STATE.md: section 8 records the restructuring and the GitHub steps (issue renames, triage labels, chain resumes with T22). Refs #96, #108, #109, #110. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wzx8WtMEXnAfyzyujdySFh
…tasks T22–T31 [skip ci] (#114) * Docs: split T25 into T22/T24-T26 after PR #107, renumber the post-v1 tasks T22-T30 [skip ci] The former T25 (performance pass) carried six deliverables from the #107 prototypes, several sessions of work under the one-task-per-session rule. TASKS.md now has: T22 ordering chooser scoring the supernodal schedule depth (#108, host only, precondition for the fused solve); T23 backends (unchanged); T24 partitioned-inverse fused solve, solve_alg = "algo1" (#82 solve half); T25 split factorization (tiled SYRK, chunked TRSM, split regime-C fronts, longest-first regime A); T26 segmented fused factorization with dependency counters, with #109 and #110 inside it plus the #75 remainder and the #96 re-measurement; T27 hybrid memory (was T26); T28 robustness extras (was T27, with the FP32 finding of #107); T29 non-uniform batch (was T22); T30 ND partition-tree export and ordering cache (was T24); External unchanged. Each new task names the bench prototype it adopts, the measured numbers, the negative results not to re-run, the tests and the Report asks. PLAN.md, PERFORMANCE.md, bench/README.md: task-id references follow the renumbering; the stale "empty repository" status line and the M1 row's ordering sentence are updated. STATE.md: section 8 records the restructuring and the GitHub steps (issue renames, triage labels, chain resumes with T22). Refs #96, #108, #109, #110. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wzx8WtMEXnAfyzyujdySFh * Docs: T31 for the oneAPI/Metal remainder of T23 (#112, #53) [skip ci] PR #113 delivers the AMDGPU extension and the local-memory cap alone and opened #112 for the rest of T23. The remainder becomes T31 at the end of the chain: oneAPI and Metal extensions, the #53 allocation-free :ka fallbacks and the select_impl order noted in #112. STATE.md section 8 updated. Refs #112, #53. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wzx8WtMEXnAfyzyujdySFh --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
METIS_OPTION_NSEPS (candidate separators per split) is the one NodeND knob that reproducibly improves the ACOPF KKT orderings: on kkt_pglib_opf_case78484_epigrids_condensed_10 (n = 157k), nd_nseps = 4 gives nnz(L) 14.42M vs 14.76M (-2.3%), factorization flops -6.7%, and the PR #107 fused refactorization 30.1 ms vs 32.1 ms (-6.2%, interleaved A/B on a GV100, correctness-gated) at the same schedule depth. The ordering itself runs ~2x slower, so the default stays at the METIS default (-1); the option pays off when one symbolic analysis is reused across many factorizations and is exposed like nd_ubfactor. ND_PROVIDER gains a third argument (values < 1 fall back to the library default); validated like the other config ints and covered in test_options and test_symbolic_etree.
Bench-level prototype campaign for the performance plan (PERFORMANCE.md experiments 0, 5, 6; issues #25, #75, #82), developed against the pglib ACOPF condensed KKT systems that MadNLP produces. Everything was measured on a Quadro GV100 against cuDSS 0.8.0 on the same device, with per-trial correctness gates (factor comparison against stock, relative residual, zero failing fronts) — several intermediate states were faster and wrong, so no number below is reported without its gate.
Results (78k-bus condensed KKT, n = 674,562, Float64, median of ≥10 warm)
Everything is cuDSS-free (native
reordering_alg = "algo4"). On AMDGPU (Radeon VII, gfx906), where cuDSS does not exist, the solve prototype runs with three mechanical shims: 20.5 → 8.05 ms.What is in the branch
src (the only reviewable library change, deliberately tiny):
subtree_max_frontsanalysis tuning option — oneOptionsfield plus one count inbuild_schedule's eligibility pass, default off. Measured near-neutral on the GV100 (it falsified the longest-walk hypothesis: the subtree launch is throughput-bound at ~3.3k subtrees) and kept as the instrument to re-test on larger-SM devices.bench/ (prototypes and experiment records, not library code):
solve_proto_78k.jl,solve_proto_all.jl— partitioned-inverse solve (per-front L11 inverses, TRSV→GEMV) with fused dependency-counter sweeps and a single-block chain kernel; the full SPD-harness table including the matrices where it regresses (lap3d_40, apache2) and the per-schedule strategy pick that implies.fact_split.jl— split factorization: stock assembly + tiled syrk (32×32, disjoint writes, deterministic, bit-identical) + chunked trsm for m > 64; the routing study (regime_c_rows, width thresholds).fact_fused.jl— segmented dependency-counter factorization: one mega-kernel per segment with four workgroup roles (extend-add, panel, trsm, syrk tiles), private contribution blocks, wides between segments; plus the write-through extend-add and the stream/budget knobs.order_search.jl,solve_nd.jl,get_cudss_perm.jl— the ordering study.amd_proto.jl,amd/Project.toml— the AMDGPU validation.final_sweep.jl— solve configuration sweeps.Findings a future implementation needs (each measured, details in the commit messages and comments below)
ordering_costweighs depth at ~0.2% for n = 674k and scores column-etree depth, which does not predict supernodal schedule depth (cuDSS perm: 1279 column levels → 30 schedule levels; AMD: competitive column metrics → 65). Forcing ND (algo4) is worth 2× schedule depth, 36% on the solve, 11–17% on factorization. Recommendation: score supernodal schedule depth, or prefer ND on GPU backends.regime_c_rowsmatters as much as width: tall-narrow fronts belong on the split path (the solve's dense-pathrowstrigger is a third of its win).subtree_budgets = [16384]) is worth 13.3 → 8.4 ms on that bucket.sizeof(T)already)._factor_panel_c!'s info check host-synchronizes, blocking stream overlap of independent wide fronts; CUDA-graph replay is neutral at this scale (busy-bound) and only pays below n ≈ 50k.Review guidance
The src diff is the only thing proposed for
mainas-is. The bench prototypes are experiment records: deliberately CUDA-only, SPD-only, nbatch = 1, and kept exactly as measured (including dead ends) so the numbers in this PR remain reproducible —bench/solve_proto_78k.jlandbench/fact_fused.jleach reproduce their headline table in one run given the KKT dump (viabench/dump_madnlp_kkt.jl). Where a prototype contradicts a PLAN.md design assumption (stack placement, ordering cost), the finding is flagged in the comments rather than edited into PLAN.md, per AGENTS.md.🤖 Generated with Claude Code