Skip to content

symbolic: device construction of the pattern and assembly maps - #122

Open
sshin23 wants to merge 2 commits into
mainfrom
perf/device-symbolic-maps
Open

sshin23 wants to merge 2 commits into
mainfrom
perf/device-symbolic-maps

Conversation

@sshin23

@sshin23 sshin23 commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Device construction of the symbolic maps

Builds three data-parallel pieces of the host analysis on the solver's GPU
backend, with element-identical results to the host originals:

  • SymmetricPattern — the both-triangle adjacency by packed-key device sort
    and deduplication (device_symmetric_pattern);
  • assembly_map — one thread per CSR row, binary search per entry
    (device_assembly_map);
  • _group_amap (unsigned) — packed-key device sort reproducing the host
    (abs, neg, p) order exactly (device_group_amap).

The route is gated by device_maps_supported: a GPU backend, a symmetric
structure ("SPD", "HPD", "S", "H") with a triangular stored view, and
sizes that fit the packed 64-bit keys (n, nnz < 2^32, factor offsets
< 2^31). Everything else — "G", full views, the CPU backend, oversized
problems — takes the host route unchanged. DirectSolver passes its backend
through _reorder!/_symbolic!; symbolic_analysis accepts device as an
opt-in keyword.

Why

On a 674k-row condensed ACOPF KKT (pglib 78484, GV100), these three pieces
are the bulk of the analysis time once the ordering is given:

piece host device
assembly map (4.3M entries) 0.37 s 0.03 s
amap grouping 0.68 s 0.04 s
symmetric pattern build 0.83 s 0.21 s

End to end (MadNLP SparseCondensedKKTSystem, user_perm, warm paired runs
in one session): analysis-dominated init 6.4 s → 3.3 s, wall 33.0 s → 30.1 s,
identical iteration count and objective.

Determinism and exactness

  • Integer sort keys and integer atomic counts only; every output is
    bit-identical to the host construction (the new tests assert equality, with
    explicit non-vacuity checks).
  • The grouping comparator folds the sign bit below the offset; within the
    gated triangle views each stored entry maps to a distinct offset, so the
    host tie order (p only) is reproduced exactly.

Conventions notes (deviations flagged rather than hidden)

  • _group_count_kernel! and the pattern/grouping reductions use integer
    atomics during the analysis phase. PLAN §2.4's no-atomics rule is written
    for assembly/extend-add in the numeric phase; analysis-side integer counts
    are order-independent and deterministic. Flagging since the rule's wording
    is broader than its motivation.
  • The device pattern route skips _host_pattern's eager range validation
    (the SymmetricPattern inner constructor still validates the result);
    malformed input fails with the inner constructor's error rather than
    _host_pattern's.
  • The maps are uploaded for the kernels and later again by
    adapt(backend, symbolic, INT); deduplicating the two uploads is left to a
    follow-up since it touches the Symbolic/adapt contract.

Tests

test/test_symbolic_device.jl: host/device equality for all three routes
over sizes (0, 1, 40, 300, 1500) × views ('L', 'U') × structures
("SPD", "S") on every backend (the kernels run on the CPU backend too,
so CPU CI exercises them), plus an end-to-end solver check that the stored
maps equal recomputed host maps and the factorization/solve still passes.

🤖 Generated with Claude Code

sshin23 and others added 2 commits October 9, 2026 11:20
On a GPU backend the analysis builds three data-parallel pieces through
KernelAbstractions kernels with element-identical results to the host
originals: the SymmetricPattern adjacency (packed-key sort + dedup), the
assembly map (binary search per stored entry), and the unsigned amap
grouping (packed-key sort reproducing the host (abs, neg, p) order).
Gated by device_maps_supported (symmetric structures, triangular stored
views, packed-key size bounds); everything else keeps the host route.

Measured on a 674k condensed ACOPF KKT (GV100): assembly map 0.37 -> 0.03 s,
grouping 0.68 -> 0.04 s, pattern build 0.83 -> 0.21 s; MadNLP
SparseCondensedKKTSystem init 6.4 -> 3.3 s with a given permutation, same
iterations and objective.

Tests: test/test_symbolic_device.jl asserts host/device equality for all
three routes on every backend (sizes 0..1500, views L/U, structures SPD/S)
plus an end-to-end analyze/factorize/solve. CPU suite 41389 pass; CUDA leg
of the new file 121 pass locally (GV100).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
_group_amap has no docstring, so the internals @autodocs page has no target
for it; Documenter cannot resolve a [`_group_amap`](@ref). Per AGENTS.md,
@ref is for documented names only, plain backticks otherwise.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant