eBPF-based Linux profiler — CPU, GPU and off-CPU in one flame graph, with CUDA kernels correlated to the launching stack. System-wide or per-PID, pprof output.
One capture of a live PyTorch step: pa_train_step (torch_workload.py:76) under
Thread.run, down through torch::autograd and cublasSgemm_v2, across the
[gpu:launch] boundary into the CUDA kernel — Python, C++ and GPU in one tree.
Interactive version.
- CUDA kernels under the stack that launched them, joined on CUPTI's correlation id — a slow kernel points at the call path responsible, Python frames included. How the adapter is loaded →
- Python frames walked from BPF, out of the interpreter's own frame chain.
No injection into the target, no
CAP_SYS_PTRACE, nothing mutated. CPython 3.12–3.14. How the walk works → - Off-CPU stalls:
--offcpuhookssched_switchand accumulates blocking time per call site — lock waits, syscall blocks, mutex contention. - Release C++/Rust without frame pointers, via a hybrid FP +
.eh_frameCFI walker. Node.js, Go and any runtime writing/tmp/perf-<pid>.maptoo. - Stripped production binaries, symbolized from
.gnu_debugdataor fetched off-box. debuginfod setup → - Hardware counters and Kubernetes labels, for PMU investigations and per-pod attribution.
Hot-attach to a running process — no restart, no preinstalled agent. The agent never writes to the process it measures.
# Build (one-time; see BUILDING.md for the toolchain)
make build
# Grant capabilities once so later runs don't need sudo
sudo setcap cap_bpf,cap_perfmon,cap_sys_ptrace,cap_checkpoint_restore,cap_syslog+ep ./perf-agent
# Capture a 30-second CPU profile of one process
./perf-agent --profile --pid <PID> --duration 30s
# Inspect
go tool pprof <output>.pb.gzcap_syslog is what makes /proc/kallsyms return real addresses; without it
kernel frames stay unsymbolized with no error. On kernels older than 5.9, add
cap_sys_admin.
For CPU + GPU in one tree, see Usage.
| Usage and flags | every flag, with the examples |
| Requirements | kernel, capabilities, distro notes |
| Architecture | how a sample becomes a profile |
| Output | file naming, pprof fidelity, PMU format |
| What you can do | the capabilities above, in full |
| Library usage | embedding the agent in Go |
| Building · Testing | toolchain and test gates |
| Releases · Changelog | what shipped when |
Deeper notes: GPU injection · CPython frame walking · debuginfod symbolization · perf.data output
Contributing · Code of conduct · Security · License (Apache-2.0)
