Repository navigation
Conversation
On a GPU backend the analysis builds three data-parallel pieces through KernelAbstractions kernels with element-identical results to the host originals: the SymmetricPattern adjacency (packed-key sort + dedup), the assembly map (binary search per stored entry), and the unsigned amap grouping (packed-key sort reproducing the host (abs, neg, p) order). Gated by device_maps_supported (symmetric structures, triangular stored views, packed-key size bounds); everything else keeps the host route. Measured on a 674k condensed ACOPF KKT (GV100): assembly map 0.37 -> 0.03 s, grouping 0.68 -> 0.04 s, pattern build 0.83 -> 0.21 s; MadNLP SparseCondensedKKTSystem init 6.4 -> 3.3 s with a given permutation, same iterations and objective. Tests: test/test_symbolic_device.jl asserts host/device equality for all three routes on every backend (sizes 0..1500, views L/U, structures SPD/S) plus an end-to-end analyze/factorize/solve. CPU suite 41389 pass; CUDA leg of the new file 121 pass locally (GV100). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Device construction of the symbolic maps
Builds three data-parallel pieces of the host analysis on the solver's GPU
backend, with element-identical results to the host originals:
SymmetricPattern— the both-triangle adjacency by packed-key device sortand deduplication (
device_symmetric_pattern);assembly_map— one thread per CSR row, binary search per entry(
device_assembly_map);_group_amap(unsigned) — packed-key device sort reproducing the host(abs, neg, p)order exactly (device_group_amap).The route is gated by
device_maps_supported: a GPU backend, a symmetricstructure (
"SPD","HPD","S","H") with a triangular stored view, andsizes that fit the packed 64-bit keys (
n,nnz< 2^32, factor offsets< 2^31). Everything else —
"G", full views, the CPU backend, oversizedproblems — takes the host route unchanged.
DirectSolverpasses its backendthrough
_reorder!/_symbolic!;symbolic_analysisacceptsdeviceas anopt-in keyword.
Why
On a 674k-row condensed ACOPF KKT (pglib 78484, GV100), these three pieces
are the bulk of the analysis time once the ordering is given:
End to end (MadNLP
SparseCondensedKKTSystem,user_perm, warm paired runsin one session): analysis-dominated init 6.4 s → 3.3 s, wall 33.0 s → 30.1 s,
identical iteration count and objective.
Determinism and exactness
bit-identical to the host construction (the new tests assert equality, with
explicit non-vacuity checks).
gated triangle views each stored entry maps to a distinct offset, so the
host tie order (
ponly) is reproduced exactly.Conventions notes (deviations flagged rather than hidden)
_group_count_kernel!and the pattern/grouping reductions use integeratomics during the analysis phase. PLAN §2.4's no-atomics rule is written
for assembly/extend-add in the numeric phase; analysis-side integer counts
are order-independent and deterministic. Flagging since the rule's wording
is broader than its motivation.
_host_pattern's eager range validation(the
SymmetricPatterninner constructor still validates the result);malformed input fails with the inner constructor's error rather than
_host_pattern's.adapt(backend, symbolic, INT); deduplicating the two uploads is left to afollow-up since it touches the
Symbolic/adaptcontract.Tests
test/test_symbolic_device.jl: host/device equality for all three routesover sizes (0, 1, 40, 300, 1500) × views (
'L','U') × structures(
"SPD","S") on every backend (the kernels run on the CPU backend too,so CPU CI exercises them), plus an end-to-end solver check that the stored
maps equal recomputed host maps and the factorization/solve still passes.
🤖 Generated with Claude Code