Skip to content

docs: release notes for 1.3.0 - #205

Open
dpsoft wants to merge 1 commit into
mainfrom
docs/release-1.3.0
Open

dpsoft wants to merge 1 commit into
mainfrom
docs/release-1.3.0

Conversation

@dpsoft

@dpsoft dpsoft commented Oct 9, 2026

Copy link
Copy Markdown
Owner

262 commits since v1.2.1. The GPU and Python release.

Why this is not the old draft

docs/prepare-v1.3.0 was written on 2026-09-20 and sat 30 commits behind
main. Three of its claims had since become false, and one of them would
have actively misled an operator:

  • It said Python frames are "not symbolized", each rendering as
    python:0x<address>. #130
    landed after it was written. Frames now carry qualname and file:line —
    pa_train_step (torch_workload.py:76) — and python:0x… is only the
    fallback when a code object cannot be resolved.

  • It said launches are sampled "one launch in N (-period)". There is no
    -period on perf-agent.
    The adapter reads PERFAGENT_GPU_SAMPLE_PERIOD
    from the target's environment, which has to be set before the workload
    starts, next to CUDA_INJECTION64_PATH. -period exists only on
    cmd/gpu-cuda-profile, for runs it launches itself. Someone following the
    old sentence would have gone looking for a flag the binary in their hand
    does not have.

  • It credited Kubernetes pod labels to this release. They shipped in feat(perfagent): namespace-aware --pid + k8s pprof labels #14,
    before v1.2.1. Only the per-sample label machinery
    (#28) is new.

What the draft was missing

Everything merged after 2026-09-20: --gpu-ring-bytes and the gpu: LOST
reporting (#204), MiniDebugInfo eval-loop reading and its cache (#190, #197),
the separately bounded debuginfod dial (#199), the bounded ptrace-stop spin
(#198), -fuse (#158), Go PGO via Function.start_line (#180), and the
correctness fixes (#164, #172, #185, #187, #192, #195, #196).

Known limitations

The limits were scattered through the draft's prose; they are now one
section, so a reader deciding whether this release fits their fleet does not
have to assemble them:

  • Python enrolment is per-PID — a system-wide -a capture gets native-only
    stacks (#194)
  • Python frames are amd64 + glibc only
  • off-CPU stacks stay native
  • Python + GPU is not covered by CI (no CI machine has a GPU)
  • the CUPTI integration is untested on ARM
  • DEBUGINFOD_URLS still enables network symbolization implicitly
    (#111)

Checked, not assumed

Every flag named in the notes was verified to exist in the source
(-gpu-pc-sampling, -mode, -discover, -keep-instrumentation-frames,
--no-debuginfod, --gpu-ring-bytes, -fuse, -period, …), and every
issue cited was verified to resolve.

One claim I expected to have to soften turned out to be true: release tags
do publish the images. release.yml pushes per-arch tags and builds the
manifests, so #139 — which
says nothing pushes them — is stale and can be closed by
#151.

One caveat on --gpu-ring-bytes

The notes describe what it does and the loss it addresses, both measured. The
recovery figure (4 MB → 64 MB taking kernel_dropped from 41,149 to 0) is a
single clean pair; the confirmation run is still owed on a healthy GPU. The
notes therefore state the loss and the knob, and claim no recovery number.

262 commits since v1.2.1. The GPU and Python release.

The earlier draft on docs/prepare-v1.3.0 is not reused as-is: it was written
on 2026-09-20, sat 30 commits behind main, and three of its claims had since
become false.

  - It said Python frames are "not symbolized", each rendering as
    python:0x<address>. #130 landed after it was written; frames now carry
    qualname and file:line, and python:0x is only the fallback when a code
    object cannot be resolved.

  - It said launches are sampled "one launch in N (-period)". There is no
    -period on perf-agent. The adapter reads PERFAGENT_GPU_SAMPLE_PERIOD from
    the TARGET's environment, which must be set before the workload starts;
    -period exists only on cmd/gpu-cuda-profile, for runs it launches itself.
    An operator following the old sentence would have looked for a flag that
    does not exist on the binary they were holding.

  - It credited Kubernetes pod labels to this release. They shipped in #14,
    before v1.2.1. Only the per-sample label machinery (#28) is new.

Added since the draft: --gpu-ring-bytes and the gpu: LOST reporting (#204),
MiniDebugInfo eval-loop reading and its cache (#190, #197), the separately
bounded debuginfod dial (#199), the bounded ptrace-stop spin (#198), -fuse
(#158), Go PGO via Function.start_line (#180), and the correctness fixes
(#164, #172, #185, #187, #192, #195, #196).

A Known limitations section replaces the limits that were scattered through
the prose: per-PID Python enrolment (#194), amd64 + glibc only, off-CPU
stacks staying native, Python+GPU not covered by CI, CUPTI untested on ARM,
and the undecided DEBUGINFOD_URLS default (#111).

Every flag named was checked to exist, and every issue cited to resolve.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant