Repository navigation
Conversation
262 commits since v1.2.1. The GPU and Python release.
The earlier draft on docs/prepare-v1.3.0 is not reused as-is: it was written
on 2026-09-20, sat 30 commits behind main, and three of its claims had since
become false.
- It said Python frames are "not symbolized", each rendering as
python:0x<address>. #130 landed after it was written; frames now carry
qualname and file:line, and python:0x is only the fallback when a code
object cannot be resolved.
- It said launches are sampled "one launch in N (-period)". There is no
-period on perf-agent. The adapter reads PERFAGENT_GPU_SAMPLE_PERIOD from
the TARGET's environment, which must be set before the workload starts;
-period exists only on cmd/gpu-cuda-profile, for runs it launches itself.
An operator following the old sentence would have looked for a flag that
does not exist on the binary they were holding.
- It credited Kubernetes pod labels to this release. They shipped in #14,
before v1.2.1. Only the per-sample label machinery (#28) is new.
Added since the draft: --gpu-ring-bytes and the gpu: LOST reporting (#204),
MiniDebugInfo eval-loop reading and its cache (#190, #197), the separately
bounded debuginfod dial (#199), the bounded ptrace-stop spin (#198), -fuse
(#158), Go PGO via Function.start_line (#180), and the correctness fixes
(#164, #172, #185, #187, #192, #195, #196).
A Known limitations section replaces the limits that were scattered through
the prose: per-PID Python enrolment (#194), amd64 + glibc only, off-CPU
stacks staying native, Python+GPU not covered by CI, CUPTI untested on ARM,
and the undecided DEBUGINFOD_URLS default (#111).
Every flag named was checked to exist, and every issue cited to resolve.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
262 commits since v1.2.1. The GPU and Python release.
Why this is not the old draft
docs/prepare-v1.3.0was written on 2026-09-20 and sat 30 commits behindmain. Three of its claims had since become false, and one of them wouldhave actively misled an operator:
It said Python frames are "not symbolized", each rendering as
python:0x<address>. #130landed after it was written. Frames now carry qualname and
file:line—pa_train_step (torch_workload.py:76)— andpython:0x…is only thefallback when a code object cannot be resolved.
It said launches are sampled "one launch in N (
-period)". There is no-periodonperf-agent. The adapter readsPERFAGENT_GPU_SAMPLE_PERIODfrom the target's environment, which has to be set before the workload
starts, next to
CUDA_INJECTION64_PATH.-periodexists only oncmd/gpu-cuda-profile, for runs it launches itself. Someone following theold sentence would have gone looking for a flag the binary in their hand
does not have.
It credited Kubernetes pod labels to this release. They shipped in feat(perfagent): namespace-aware --pid + k8s pprof labels #14,
before v1.2.1. Only the per-sample label machinery
(#28) is new.
What the draft was missing
Everything merged after 2026-09-20:
--gpu-ring-bytesand thegpu: LOSTreporting (#204), MiniDebugInfo eval-loop reading and its cache (#190, #197),
the separately bounded debuginfod dial (#199), the bounded ptrace-stop spin
(#198),
-fuse(#158), Go PGO viaFunction.start_line(#180), and thecorrectness fixes (#164, #172, #185, #187, #192, #195, #196).
Known limitations
The limits were scattered through the draft's prose; they are now one
section, so a reader deciding whether this release fits their fleet does not
have to assemble them:
-acapture gets native-onlystacks (#194)
DEBUGINFOD_URLSstill enables network symbolization implicitly(#111)
Checked, not assumed
Every flag named in the notes was verified to exist in the source
(
-gpu-pc-sampling,-mode,-discover,-keep-instrumentation-frames,--no-debuginfod,--gpu-ring-bytes,-fuse,-period, …), and everyissue cited was verified to resolve.
One claim I expected to have to soften turned out to be true: release tags
do publish the images.
release.ymlpushes per-arch tags and builds themanifests, so #139 — which
says nothing pushes them — is stale and can be closed by
#151.
One caveat on
--gpu-ring-bytesThe notes describe what it does and the loss it addresses, both measured. The
recovery figure (4 MB → 64 MB taking
kernel_droppedfrom 41,149 to 0) is asingle clean pair; the confirmation run is still owed on a healthy GPU. The
notes therefore state the loss and the knob, and claim no recovery number.