You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
T23 remainder: oneAPI and Metal extensions, allocation-free KA dense fallbacks (#53) #112
What: by owner decision in the implementing session, the T23 PR delivers only the AMDGPU extension (ext/SparseDirectSolverAMDGPUExt.jl: rocSPARSE adapters, the public API, rocBLAS/rocSOLVER vendor bindings, tested on an MI300X) and item 2 of #60 (max_local_bytes(backend), the 64 KiB regime-A class, resolve_subtree_budgets). Three parts of the T23 text are not done and are no longer tracked by an open task once #23 closes:
the oneAPI extension (oneMKL dense bindings, oneSparseMatrixCSR adapters), precompile-checked only;
the Metal extension (MPS bindings or KA only; max_local_bytes(::MetalBackend) = 32768, which the new capability already supports), by review only;
For (3), the T23 survey found no allocation calls in src/dense/fallback/*.jl: the per-call cost is the launch configuration (Vals built from runtime flags, e.g. Val(ul == 'L') in potrf.jl:79, Val(tA)/Val(tB) in gemm.jl:58, the Union-typed tile of _default_tile in gemm.jl:63; kernel objects built per call; keyword calls). Also note that select_impl resolves :auto to :generic (the allocating LinearAlgebra path) on a backend without vendor_* bindings whose mul!/cholesky! probes pass, so on oneAPI/Metal :auto would not reach the :ka fallbacks unless that order changes.
Why it is out of scope: T23 as written owns it; the session split it off to land the AMDGPU part in a reviewable PR. It does not block T24–T27. It blocks running the suite on oneAPI or Metal.
Suggested fix or plan change: a new task "T23b — oneAPI and Metal extensions, allocation-free KA fallbacks (#53)" after T23, with the T23 owner notes for #53 moved there.
Triaged: becomes task T31 (oneAPI and Metal extensions, allocation-free KA dense fallbacks, #53) at the end of the chain in TASKS.md (PR #114), since it blocks nothing and no oneAPI/Metal hardware is available. Closes with that task's PR.
Found while working on: T23 (PR #113)
What: by owner decision in the implementing session, the T23 PR delivers only the AMDGPU extension (
ext/SparseDirectSolverAMDGPUExt.jl: rocSPARSE adapters, the public API, rocBLAS/rocSOLVER vendor bindings, tested on an MI300X) and item 2 of #60 (max_local_bytes(backend), the 64 KiB regime-A class,resolve_subtree_budgets). Three parts of the T23 text are not done and are no longer tracked by an open task once #23 closes:oneSparseMatrixCSRadapters), precompile-checked only;max_local_bytes(::MetalBackend) = 32768, which the new capability already supports), by review only;:kadense fallbacks (ka_potrf!,ka_trsm!,ka_gemm!and batched variants), which are the only dense path on oneAPI/Metal.For (3), the T23 survey found no allocation calls in
src/dense/fallback/*.jl: the per-call cost is the launch configuration (Vals built from runtime flags, e.g.Val(ul == 'L')inpotrf.jl:79,Val(tA)/Val(tB)ingemm.jl:58, theUnion-typed tile of_default_tileingemm.jl:63; kernel objects built per call; keyword calls). Also note thatselect_implresolves:autoto:generic(the allocating LinearAlgebra path) on a backend withoutvendor_*bindings whosemul!/cholesky!probes pass, so on oneAPI/Metal:autowould not reach the:kafallbacks unless that order changes.Why it is out of scope: T23 as written owns it; the session split it off to land the AMDGPU part in a reviewable PR. It does not block T24–T27. It blocks running the suite on oneAPI or Metal.
Suggested fix or plan change: a new task "T23b — oneAPI and Metal extensions, allocation-free KA fallbacks (#53)" after T23, with the T23 owner notes for #53 moved there.