Skip to content

About

eBPF-based Linux profiler — CPU, GPU and off-CPU in one flame graph, with CUDA kernels correlated to the launching stack. System-wide or per-PID, pprof output.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Latest commit

 

History

523 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

perf-agent

eBPF-based Linux profiler — CPU, GPU and off-CPU in one flame graph, with CUDA kernels correlated to the launching stack. System-wide or per-PID, pprof output.

CI Tests Go Reference Go Version License


PyTorch training: Python, torch's C++ autograd, and CUDA kernels in one flame graph

One capture of a live PyTorch step: pa_train_step (torch_workload.py:76) under Thread.run, down through torch::autograd and cublasSgemm_v2, across the [gpu:launch] boundary into the CUDA kernel — Python, C++ and GPU in one tree. Interactive version.


What it does

  • CUDA kernels under the stack that launched them, joined on CUPTI's correlation id — a slow kernel points at the call path responsible, Python frames included. How the adapter is loaded →
  • Python frames walked from BPF, out of the interpreter's own frame chain. No injection into the target, no CAP_SYS_PTRACE, nothing mutated. CPython 3.12–3.14. How the walk works →
  • Off-CPU stalls: --offcpu hooks sched_switch and accumulates blocking time per call site — lock waits, syscall blocks, mutex contention.
  • Release C++/Rust without frame pointers, via a hybrid FP + .eh_frame CFI walker. Node.js, Go and any runtime writing /tmp/perf-<pid>.map too.
  • Stripped production binaries, symbolized from .gnu_debugdata or fetched off-box. debuginfod setup →
  • Hardware counters and Kubernetes labels, for PMU investigations and per-pod attribution.

Hot-attach to a running process — no restart, no preinstalled agent. The agent never writes to the process it measures.

Quickstart

# Build (one-time; see BUILDING.md for the toolchain)
make build

# Grant capabilities once so later runs don't need sudo
sudo setcap cap_bpf,cap_perfmon,cap_sys_ptrace,cap_checkpoint_restore,cap_syslog+ep ./perf-agent

# Capture a 30-second CPU profile of one process
./perf-agent --profile --pid <PID> --duration 30s

# Inspect
go tool pprof <output>.pb.gz

cap_syslog is what makes /proc/kallsyms return real addresses; without it kernel frames stay unsymbolized with no error. On kernels older than 5.9, add cap_sys_admin.

For CPU + GPU in one tree, see Usage.

Documentation

Usage and flags every flag, with the examples
Requirements kernel, capabilities, distro notes
Architecture how a sample becomes a profile
Output file naming, pprof fidelity, PMU format
What you can do the capabilities above, in full
Library usage embedding the agent in Go
Building · Testing toolchain and test gates
Releases · Changelog what shipped when

Deeper notes: GPU injection · CPython frame walking · debuginfod symbolization · perf.data output


Contributing · Code of conduct · Security · License (Apache-2.0)

About

eBPF-based Linux profiler — CPU, GPU and off-CPU in one flame graph, with CUDA kernels correlated to the launching stack. System-wide or per-PID, pprof output.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages