punktfunk-encode-worker: GPU priority via a capability-carrying worker, with WP3 on-glass complete #153

Merged
enricobuehler merged 14 commits from worktree-worktree-encode-worker into main 2026-08-10 10:45:26 +00:00
Owner

Brings back the CAP_SYS_NICE GPU-priority lever on every compositor — including KDE desktop, the population that both hit the 0.26.0-1 breakage and wants the lever most — without ever re-breaking KWin identification.

The invariant: punktfunk-host carries no file capability, ever. The worker is a separate file — never a hardlink, never a host subcommand. A shared inode shares the capability and re-creates 0.26.0-1.

Design: punktfunk-planning/design/gpu-priority-capability-worker.md (+ the implementation plan).

What's here

WP0 (pf-zerocopy rails split out) · WP1 (worker crate + protocol) · WP2 (host proxy + fallback ladder) · WP4 (packaging x6 + a getcap CI leg asserting host EMPTY / worker cap_sys_nice=ep) · WP5 (docs).

WP3 — on-glass validation, now complete

leg result
I1–I6 inspect PASS x6 — incl. host/worker are separate inodes, running host CapPrm all zero, /proc/<pid>/exe readable
V1 — the 0.26.0-1 regression test PASS on both KWin boxes: home-nobara-1 (NVIDIA) and Deck .253 desktop mode (RADV). This leg had never been run anywhere before
V2 — grant PASS on both vendors: Ready.granted = Realtime on RTX 5070 Ti and AMD Custom GPU 0932 (RADV VANGOGH)
V3a — IPC hop PASS, R1's abandonment gate does not fire (third box agreeing)
V4 — chaos, all 5 rungs PASS (--allow-mutate)
V5 — fd hygiene PASS — 33 samples over 480 s, host trend 0, worker trend 0
WP4 cap matrix self-test green · real .deb PASS · real capped-host .deb REFUSED

V4.e's respawn rung was proven on a real session rather than the spike: worker killed with -9, video_streaming stayed true across the kill, and the host logged respawned the encode worker after a mid-session death priority=Granted(Realtime).

The WP4 cap-matrix leg was also run against a real built artifact for the first time — a .deb from packaging/debian/build-deb.sh, then the same deb repacked with setcap cap_sys_nice=ep /usr/bin/punktfunk-host appended to its postinst, which the asserter correctly refused. The RPM %caps and squashfs-xattr readers remain self-test-only (that chain belongs in release CI); the squashfs reader was additionally exercised against real .raw images in both directions.

Two validation-kit fixes (both were FALSE NEGATIVES)

  • v4.e demanded a rung the vehicle cannot reach. L_DIED/L_RESPAWNED come from RemotePyroWave::reset, and the only caller of Encoder::reset is the real session's reset_stalled_encoder loop. spike has no recovery loop — it does encoder.submit(..)? and exits — so a worker killed under the spike can never reach reset. The leg reported "a shipping blocker" for a rung the product implements correctly.
  • v5 called jitter a leak. The verdict was max - min with tolerance 0, but the fd count legitimately moves by one (a dmabuf fd in flight at sampling time), so the spread was permanently 1 and the leg failed forever on a healthy box. Measured: 33 samples oscillating 54↔55, ending on 54 exactly where they started. Replaced with a median-of-thirds trend, which is more sensitive to a real leak. Self-test cases added for both.

⚠ Three things a reviewer should know

  1. Two commits here are unrelated to the worker programme and could reasonably be split out:
    • 5b3ea6e8 — one NVENC open failure could kill every session on the box (libav's ff_cuda_check formats an unguarded pointer; measured twice on home-nobara-1, once SIGSEGV mid-session, once a wedged thread that made systemd SIGABRT the service).
    • a23c0284 — the console reported the resolution the client asked for, not the one it got (visible on the gamescope attach path, where the box streams its panel's mode).
  2. ad63994c is NOT validated on glass. It moves the Linux 10-bit capability probe off ffmpeg — the last ffmpeg-NVENC open on a default direct-SDK host, and the documented NV_ENC_ERR_INVALID_VERSION process-wide wedge trigger. Gates are green (clippy -D warnings with nvenc,vulkan-encode,pyrowave, 67 pf-encode tests, fmt) and the failure it targets was reproduced on glass, but the end-to-end HDR run is still owed.
  3. This branch is 42 commits behind main and its build cannot complete a punktfunk/1 handshake on home-nobara-1 — it stalls between "audio channels resolved" and "encode bit depth" and times out at 10 s, every attempt. That stall is not from any commit here: a control build with only ad63994c's routing reverted stalls identically, and the released 0.27.0 RPM on the same box handshakes fine. It should disappear on merge, but rebase onto main and re-run before trusting the HDR leg.

Gates: cargo fmt; cargo clippy --workspace --all-targets --locked -- -D warnings on linux/amd64; 74 pf-encode tests including 12 ladder-rung tests.

Brings back the `CAP_SYS_NICE` GPU-priority lever on every compositor — including KDE desktop, the population that both hit the 0.26.0-1 breakage and wants the lever most — without ever re-breaking KWin identification. **The invariant:** `punktfunk-host` carries no file capability, ever. The worker is a separate **file** — never a hardlink, never a host subcommand. A shared inode shares the capability and re-creates 0.26.0-1. Design: `punktfunk-planning/design/gpu-priority-capability-worker.md` (+ the implementation plan). ## What's here WP0 (pf-zerocopy rails split out) · WP1 (worker crate + protocol) · WP2 (host proxy + fallback ladder) · WP4 (packaging x6 + a getcap CI leg asserting host EMPTY / worker `cap_sys_nice=ep`) · WP5 (docs). ## WP3 — on-glass validation, now complete | leg | result | |---|---| | I1–I6 inspect | PASS x6 — incl. host/worker are **separate inodes**, running host `CapPrm` all zero, `/proc/<pid>/exe` readable | | **V1 — the 0.26.0-1 regression test** | **PASS on both KWin boxes**: home-nobara-1 (NVIDIA) and Deck .253 desktop mode (RADV). This leg had **never been run anywhere** before | | V2 — grant | **PASS on both vendors**: `Ready.granted = Realtime` on RTX 5070 Ti *and* `AMD Custom GPU 0932 (RADV VANGOGH)` | | V3a — IPC hop | PASS, R1's abandonment gate does not fire (third box agreeing) | | V4 — chaos, all 5 rungs | PASS (`--allow-mutate`) | | V5 — fd hygiene | PASS — 33 samples over 480 s, host trend 0, worker trend 0 | | WP4 cap matrix | self-test green · **real `.deb` PASS** · **real capped-host `.deb` REFUSED** | V4.e's respawn rung was proven on a **real session** rather than the spike: worker killed with `-9`, `video_streaming` stayed true across the kill, and the host logged `respawned the encode worker after a mid-session death priority=Granted(Realtime)`. The WP4 cap-matrix leg was also run against a **real built artifact** for the first time — a `.deb` from `packaging/debian/build-deb.sh`, then the same deb repacked with `setcap cap_sys_nice=ep /usr/bin/punktfunk-host` appended to its postinst, which the asserter correctly refused. The RPM `%caps` and squashfs-xattr readers remain self-test-only (that chain belongs in release CI); the squashfs reader was additionally exercised against real `.raw` images in both directions. ## Two validation-kit fixes (both were FALSE NEGATIVES) * **`v4.e` demanded a rung the vehicle cannot reach.** `L_DIED`/`L_RESPAWNED` come from `RemotePyroWave::reset`, and the only caller of `Encoder::reset` is the real session's `reset_stalled_encoder` loop. `spike` has no recovery loop — it does `encoder.submit(..)?` and exits — so a worker killed under the spike can never reach reset. The leg reported "a shipping blocker" for a rung the product implements correctly. * **`v5` called jitter a leak.** The verdict was `max - min` with tolerance 0, but the fd count legitimately moves by one (a dmabuf fd in flight at sampling time), so the spread was permanently 1 and the leg failed forever on a healthy box. Measured: 33 samples oscillating 54↔55, **ending on 54 exactly where they started**. Replaced with a median-of-thirds trend, which is *more* sensitive to a real leak. Self-test cases added for both. ## ⚠ Three things a reviewer should know 1. **Two commits here are unrelated to the worker programme** and could reasonably be split out: * `5b3ea6e8` — one NVENC open failure could kill every session on the box (libav's `ff_cuda_check` formats an unguarded pointer; measured twice on home-nobara-1, once SIGSEGV mid-session, once a wedged thread that made systemd SIGABRT the service). * `a23c0284` — the console reported the resolution the client asked for, not the one it got (visible on the gamescope attach path, where the box streams its panel's mode). 2. **`ad63994c` is NOT validated on glass.** It moves the Linux 10-bit capability probe off ffmpeg — the last ffmpeg-NVENC open on a default direct-SDK host, and the documented `NV_ENC_ERR_INVALID_VERSION` process-wide wedge trigger. Gates are green (clippy `-D warnings` with `nvenc,vulkan-encode,pyrowave`, 67 pf-encode tests, fmt) and the failure it targets was reproduced on glass, but the end-to-end HDR run is still owed. 3. **This branch is 42 commits behind main and its build cannot complete a punktfunk/1 handshake** on home-nobara-1 — it stalls between "audio channels resolved" and "encode bit depth" and times out at 10 s, every attempt. That stall is **not** from any commit here: a control build with only `ad63994c`'s routing reverted stalls identically, and the released 0.27.0 RPM on the same box handshakes fine. It should disappear on merge, but **rebase onto main and re-run before trusting the HDR leg.** Gates: `cargo fmt`; `cargo clippy --workspace --all-targets --locked -- -D warnings` on linux/amd64; 74 pf-encode tests including 12 ladder-rung tests.
enricobuehler added 13 commits 2026-08-10 10:37:01 +00:00
The encode worker (design/gpu-priority-capability-worker.md) needs exactly what the zerocopy worker
already has — SEQPACKET framing, fds as SCM_RIGHTS, a pinned-exe spawn that survives an on-disk
replacement, and a reaper that never blocks session teardown on a wedged child — but it must NOT
inherit the zerocopy protocol. Its messages are its own and version independently.

So `imp/proto.rs` keeps the vocabulary (PROTO_VERSION, ImportKind, Request, Reply, BufferDesc) and
all transport moves to `imp/ipc.rs`, reachable as `pf_zerocopy::ipc`. No behaviour change for the
zerocopy worker: client.rs now calls `ipc::self_exe()`/`ipc::spawn_worker()` and keeps the same fd-3
dup2 slot, PR_SET_PDEATHSIG, kill-then-reap-outside-the-lock, bounded reap with a D-state re-park,
and per-generation zombie sweep it had before.

Two real changes underneath the move:

  * The cmsg store was sized for exactly one fd (CMSG_SPACE(4) = 24 B). A multi-planar dmabuf can
    carry up to four, so it is now CMSG_SPACE(4*4); `send_fds`/`recv_fds` take a slice while `send`
    and `recv` keep their single-fd shapes as the fast path. An over-long fd list is rejected with
    io::Error rather than asserting — that is how MAX_MSG overflow is already handled — and the
    receive cap is enforced by the kernel through msg_controllen, so a 5-fd peer trips MSG_CTRUNC.

  * The old recv loop read only the FIRST i32 of each SCM_RIGHTS control message. Nothing sends two
    fds yet so it never fired, but every descriptor after the first in a multi-fd message would have
    leaked into the process. It now reads all of them.

Spawn takes the executable path as a parameter instead of assuming /proc/self/exe. The zerocopy
worker keeps self-exec; the encode worker passes its own binary, which must be a separate FILE and
never a subcommand — a shared inode shares the file capability.
PyroWave encodes on the same GPU shader cores the game saturates, and an elevated
VK_KHR_global_priority queue is the compute-preemption lever for it — measured on .21 (RTX 5070 Ti,
GRID 2 loop): encode p99 6.4 -> 4.4 ms. Every driver refuses every priority class without
CAP_SYS_NICE, on NVIDIA and on RADV alike, so the lever is decoration on a packaged host.

0.26.0-1 granted that capability to punktfunk-host and killed desktop streaming on every KDE box:
KWin identifies a client by resolving /proc/<pid>/exe and matching an installed .desktop's Exec=,
the kernel refuses that readlink to a reader whose effective set is not a superset of the target's
PERMITTED set (cap_ptrace_access_check), and KWin holds no capabilities. #136 revoked it everywhere.

The capability therefore cannot live in the process that fronts KWin. It lives in a new, deliberately
small binary — punktfunk-encode-worker — which owns the priority-elevated Vulkan device and talks to
nothing but the socket its parent spawned it on: no Wayland, no D-Bus, no network, no plugins. It is
a SEPARATE FILE and must stay one; a hardlink or a hidden host subcommand shares the inode, hence the
capability, and silently re-creates the incident. That rule is written where someone would break it,
in the worker crate's own Cargo.toml.

`open_inner` is reused verbatim in the worker — the same REALTIME->HIGH->none ladder, the same
refusal-never-fails-open invariant, the same PUNKTFUNK_PERF split — so the A/B stays comparable with
PW1. The only in-process change is a flag for whether THIS process prints the INERT warn, plus an
out-parameter reporting the class that was granted.

Three things the design did not anticipate:

  * An AU cannot ride in the message body. MAX_MSG is 64 KiB and bodies are serde_json, which
    renders a Vec<u8> as one decimal per byte: a 1080p60 AU is ~333 KB of JSON and 4K ~3.3 MB, and
    the minimum per-frame budget is already 64 KiB. So the AU crosses on a memfd the worker creates
    once and pwrites each frame; the fd crosses once, in Ready. A test pins the arithmetic so nobody
    "simplifies" the memfd away. Cursor bitmaps take the same route, only when their serial changes.

  * set_wire_chunking has to cross the wire even though poll_chunk does not. Chunking changes the AU
    BYTES, not merely how they are handed out — it feeds rate_budget()'s deflation and build_au's
    windowed framing — so a proxy-local copy would have the host cutting dense AUs at boundaries that
    are not window boundaries. Forwarded and mirrored. poll_chunk itself needs no protocol: the
    identical AuChunker runs host-side on the whole AU the worker returns.

  * CPU-backed frames really do reach this encoder (force_cpu_for_nvenc_444, and the raw-dmabuf
    degrade latch), and a 1080p BGRA frame is ~8 MB. The first non-dmabuf frame pins the session
    in-process with one warn rather than putting 480 MB/s on a socket.

Every rung falls back to the in-process encoder exactly as today with one warn and never a dead
session: PUNKTFUNK_ENCODE_WORKER=off, binary missing, spawn failure, handshake timeout, proto or
workspace-version skew (host and worker are different files now, so that check is load-bearing),
InitErr, a refused frame, and socket EOF mid-session — which respawns once, then pins inline.

Also: recv retries EINTR with the REMAINING deadline, not a fresh one. With SO_RCVTIMEO the kernel
returns EINTR rather than restarting, so a signal would otherwise read as a dead worker; re-arming
with the full budget would instead let a steady signal rate defer a real hang forever.
767e67ca's per-channel mechanics were correct; they were aimed at the wrong binary. Each one is
restored here pointed at punktfunk-encode-worker, and every host-side removal from #136 stays
verbatim. All grants remain best-effort — an uncapped worker still encodes, at default priority, so
a failed setcap must never fail an install.

  * Arch: setcap in post_install AND post_upgrade (a replaced binary is a new inode).
  * RPM: %caps(cap_sys_nice=ep) in %files, never a %post setcap — %caps applies, restores and
    verifies, and covers Fedora as well as Bazzite via rpm-ostree layering.
  * Bazzite + Arch sysext: setcap on the staging tree before mksquashfs, which does record
    security.capability. The assertion is amended, not removed: host EMPTY is still a hard fail, and
    the worker must carry exactly cap_sys_nice=ep — missing is fine, anything else is not.
  * deb: setcap in postinst.
  * NixOS: security.wrappers for the WORKER plus PUNKTFUNK_ENCODE_WORKER in the unit. A file
    capability cannot live on a store path, and an ambient grant is right here precisely because
    nothing ever identifies the worker. The host's ExecStart stays on the store path.
  * Steam Deck: setcap the worker; the .desktop the script writes stays valid this time.

Four things the plan's channel table missed:

  * packaging/arch/build-sysext.sh had no capability handling at all, and a sysext can never run a
    pacman scriptlet — the SteamOS image would have shipped the lever permanently inert.
  * scripts/steamdeck/update.sh had none either. It rebuilds both binaries, so a new inode drops the
    grant, and it is the documented steady-state path: the lever would have died on the first update.
    It also never healed a Deck already capped by 0.26.0-1.
  * A capped worker is AT_SECURE, and glibc drops $ORIGIN-expanded RPATH entries for secure binaries
    unless they normalise into a trusted system dir. Copying the host's rpath under BUNDLE_FFMPEG=1
    would have left the capped worker unable to find libavcodec on exactly the channel that bundles
    it. Absolute DT_RPATH instead.
  * Nix crane scopes by -p, so the worker would not have been built at all, and it needs its own
    addDriverRunpath.

scripts/ci/assert-cap-matrix.sh mechanizes the lesson from 0.26.0-1 — verify the PACKAGE, never the
board. It unpacks the built Arch package, the deb, the rpm and the mounted sysext raw and asserts one
matrix: the host carries NOTHING (hard fail), the worker exactly cap_sys_nice=ep. The sysext reader
first proves it can round-trip a capability through mksquashfs/unsquashfs at all, so an unreadable
artifact fails rather than issuing a blind PASS, and --self-test red-teams the assertions themselves.

Red-teaming the leg found a real bug: setcap originally ran BEFORE the assertion, so "the worker
arrived carrying something unexpected" was unreachable and a stray %caps would have been silently
overwritten. Both sysext scripts now assert, then grant, then assert again.
Rewrites the "GPU scheduling priority" section around the split: punktfunk-encode-worker carries
cap_sys_nice=ep, punktfunk-host carries nothing on any channel, ever. The KWin identification
mechanism is spelled out in plain words and the failure line is quoted verbatim
("KWin does not expose zkde_screencast_unstable_v1 to this client") so someone searching for their
symptom lands on the explanation.

The warning names all three ways an operator would reach for the capability — hand setcap, a systemd
AmbientCapabilities= line, a NixOS security.wrappers entry — because all three put it in the same
permitted set and all three cost KDE desktop streaming. That is the failure mode that made this
worth documenting: it looks exactly like a missing .desktop and survives reinstalling both ends.

configuration.md gains PUNKTFUNK_ENCODE_WORKER (path, or `off` to force the in-process encoder) and
re-describes PYROWAVE_QUEUE_PRIORITY as an intent forwarded to whichever process does the encode.
kde.md gains one line on the troubleshooting bullet someone actually lands on: getcap on the host
must print nothing.

The published 0.26.0 notes are deliberately untouched — they are the record of what shipped. The
flipped phrasing lives in v0.27.0's notes instead; v0.26.0.md:37 ("a system privilege that turns out
to stop KDE recognising the host at all") is the line that goes stale when this ships.
WP3 of design/gpu-priority-capability-worker-implementation-plan.md. Five legs, the first of which is
the test that would have caught the field incident: in a KDE session with the worker installed and
capped, `getcap` on the host must be EMPTY, its CapPrm all zeroes, `readlink /proc/<pid>/exe` must
resolve, and `punktfunk-host probe-compositor` must exit 0 — which on KWin succeeds only when the
privileged zkde_screencast_unstable_v1 global was actually advertised to this client.

Read-only by default; the one mutating rung (kill -9) is behind --allow-mutate and kills only a
worker that is a child of the spike the script itself started. It NEVER calls setcap: the uncapped
arms use a plain copy of the worker, which does not carry security.capability, verified uncapped
before use. So no leg needs root and none restores state. A skip is never a pass — exit 2 means
incomplete, distinct from 1 (failure).

V3 is split, which the plan did not do. Its stated form compares against PW1's in-process-capped
baselines, and those exist only on .21 under GRID 2:

  * V3a is the pre-registered abandonment gate and needs no capability at all — in-process versus an
    UNCAPPED worker, both at default priority, so the only difference is the process boundary. Fails
    if the worker's p99 exceeds inline by more than --gate-ms (1.0). This runs on any box with a GPU.
  * V3b is the lever itself, capped worker versus the refused in-process arm, and says plainly that
    an idle GPU makes it meaningless.

The false PASS this kit exists to refuse: a CPU-backed frame makes the proxy pin itself in-process
for the session, so a synthetic source would quietly turn the "worker" arm into a second in-process
arm and pass the gate for the wrong reason. The worker arm is only accepted with a dmabuf-passthrough
capture, a capability-carrying-worker line, and no fallback line anywhere in the log.

Also asserts the host and worker are different inodes — a hardlink shares the file capability, which
is the same incident by another route.
Found by running it. The first V3a run on .25 encoded 2700 frames in BOTH arms, at 59.6 fps, with 22
perf windows each — and the kit reported "fewer than 3 usable perf windows", because `tracing`'s fmt
layer wraps field NAMES in SGR escapes. The bytes on disk are `p99_us\e[0m\e[2m=\e[0m4601`, so
`s/.*p99_us=\([0-9][0-9]*\).*/\1/p` never matched. The message text is plain, which is why the
window COUNT was right and only the numbers vanished — and why the fixtures never caught it: they
were hand-written, and cleaner than reality.

Anything matching a field breaks the same way, so this was not only V3a: v2's `priority=Realtime`,
the demotion `reason=`, and v4's rungs all read fields. Every log read now goes through one
`log_cat` that strips SGR, and the spike is launched with NO_COLOR=1 so fresh logs are plain at the
source too — a human grepping a red leg by hand is defeated by those escapes exactly as the parser
was.

The self-test gains the same four perf windows a second time, ANSI-wrapped, asserting an identical
result: same numbers, same expectation, so a failure there can only mean the stripping broke. That
fixture caught its own first draft, which built the line in one printf with 27 placeholders against
23 arguments and emitted empty escapes — hence the field-at-a-time helper.

With this, V3a self-reports on .25 (sway headless, real dmabuf capture, AMD 780M/RADV, 2700 frames
per arm, both arms at default GPU priority):

    in-process        p50 2.08 ms   p99 4.18 ms   (21 windows)
    uncapped worker   p50 2.07 ms   p99 3.52 ms   (21 windows)
    p99 delta -0.66 ms  ->  PASS

R1's pre-registered abandonment gate does not fire: the process boundary is not merely under the
+1.0 ms ceiling, it is measurably FASTER at the tail, while p50 is unchanged (2.08 vs 2.07). An
earlier hand-extraction of the same logs gave -0.43 ms, so the direction reproduces across runs.
Caveat for whoever reads this later: idle iGPU in a KVM guest, RADV, no GPU-bound load. This bounds
the IPC hop; it says nothing about V3b, which still needs .21 under GRID 2.
The V3b run on .21 died with `open portal capturer: timed out waiting for the ScreenCast portal` —
a GNOME consent dialog nobody answered — and the kit reported "arm A is not the in-process arm".
That is false: the arm was constructed correctly (`PUNKTFUNK_ENCODE_WORKER=off` is right there in
the captured env header), it simply never reached encoder-open, so the line the assert looks for
could not exist. A red that points at the wrong thing costs the same debugging time as a green that
hides a real one.

`spike_failure_reason` now runs BEFORE any arm-identity assert in v2, v3a and v3b, and names the
actual cause: the portal timeout gets its own message saying the dialog appears on the HOST's own
screen and cannot be answered from inside a stream — which is precisely the situation that produced
this failure, since the operator was watching the box through a game session at the time.

Falls back to the first ERROR line, then to "no PUNKTFUNK_PERF window at all", so a spike that dies
some other way still reports that rather than a misattribution.
`--codec pyrowave` selects the ENCODER. The capture pipeline picks its consumer from
`ZeroCopyPolicy::pyrowave_session`, which on the spike path is fed only by the global
`PUNKTFUNK_ENCODER=pyrowave` lab lever (punktfunk-host/src/capture.rs). Without it, .21 resolved

    capture pipeline resolved: cuda-import -> nvenc   capture_arm="cuda-import" consumer="nvenc"
    zero-copy: dmabuf imported to CUDA (no CPU copy)  nv12=true

and the wavelet encoder refused the payload on its first submit: "unsupported FramePayload (need
Dmabuf or Cpu RGB)". That is not a worker bug — the arm that failed was the pure in-process one.

It reproduces only where the A/B actually lives. An AMD box has no CUDA arm to pick, so .25 resolved
straight to dmabuf-passthrough and the kit looked correct there. With the lever set, .21 resolves
`dmabuf-passthrough -> pyrowave` and both arms encode 2700/2700 frames.

V3b then passes on .21 (RTX 5070 Ti, GRID 2 at ~100% GPU, 5120x1440 — the portal captures the real
monitor, --width/--height being synthetic-only):

    in-process, refused       p50 2.85 ms   p99 8.39 ms   (10 windows)
    capped worker, granted    p50 2.65 ms   p99 4.10 ms   (11 windows)
    p99 delta -4.29 ms

The worker reports `priority=Granted(Realtime)` with `ext=VK_KHR_global_priority` on the FIRST
attempt and logs no fallback line; the refused arm logs "every global queue priority class was
refused". So the capability still buys the lever from a SEPARATE process, with the IPC hop in the
loop — 8.39 -> 4.10 ms is a 51% p99 cut, against PW1's in-host 6.4 -> 4.4 at 1080p. Different
resolution and a harder load, so treat the class as confirmed and the absolute numbers as not
comparable to PW1's.
v4.e killed the worker mid-session and then required "the encode worker died
mid-session" in the spike's log. That line, and the respawn that follows it, are
emitted by `RemotePyroWave::reset` — and the only caller of `Encoder::reset` is
the real session's `reset_stalled_encoder` loop in native/stream.rs. `spike` is
a dev tool with no recovery loop at all: it does

    encoder.submit(&frame).context("encoder submit")?

and exits. So a worker killed under the spike can never reach reset, the line
can never appear, and the leg reported

    FAILED — a red leg here is a shipping blocker, not a flake.

for a ladder rung the product implements correctly. A false negative in the one
place that must not have one: this kit exists to refuse false PASSes, and a
false FAIL spends exactly the same credibility.

Verified on glass first, so the rung is not being excused on a reading of the
source. home-nobara-1 (KDE, RTX 5070 Ti), real client session, worker pid 44249
killed with -9: `video_streaming` stayed true across the kill, and the host
logged

    pyrowave: respawned the encode worker after a mid-session death
      worker=/usr/bin/punktfunk-encode-worker priority=Granted(Realtime)
    encoder submit failed — encoder rebuilt in place, forcing an IDR
      error=... Broken pipe (os error 32) reset=1 max=5

v4.e now asserts the half the spike can actually observe — the death surfaces as
an ATTRIBUTABLE worker-IPC error naming the worker, after real encode windows,
and the host process does not die with it. A hang, an unexplained failure, or a
dead host still fails. The respawn half is printed as the human follow-up, in
the same idiom v1 already uses for its on-glass half, and written into `recipe`
with the two commands that close it.
v5's verdict was `max - min` over the sampled fd counts with a default tolerance
of 0. An encode worker's fd count legitimately moves by one when a dmabuf fd is
in flight at the sampling instant, so the spread was permanently 1 and the leg
failed on a perfectly healthy box — reported, like every red leg here, as "a
shipping blocker, not a flake".

Measured on home-nobara-1 (KDE, RTX 5070 Ti), 33 samples over 480 s:

    54 54 54 54 54 54 54 54 55 55 54 54 54 55 54 54 55 54 54 54 55 54 …54

It oscillates and ENDS on 54, exactly where it started. Nothing accumulates.

The replacement is median-of-thirds: median(last third) - median(first third).
That is strictly MORE sensitive to what R2 is actually about — a steady leak
moves the trend just as much as it moves the spread, while bounded jitter moves
only the spread — so this is not the tolerance being widened to get a green.
The spread is still printed, now labelled as jitter when the trend is flat. The
warm-up window already covers the one-off first-sight-of-each-buffer cost, so a
plateau inside it is by design not a leak; a step that never comes back still
trends and still fails.

The self-test grows the cases that force this to be a real assertion: the
measured oscillation must trend to zero, a synthetic leak must still trend up, a
flat series must be flat, and a step that never returns must be caught. Writing
them is what caught my own arithmetic — the first draft asserted a leak trend of
12 where the reader correctly says 10.

Also records what the v5 log now makes obvious: `--minutes` does NOT set the wall
clock. `spike` is frame-count bounded (`seconds * fps`), and a KWin virtual
output being driven hard delivers ~197 fps against a `--fps 60` budget, so a
"10 minute" run ended after 182 s. Ask for more minutes than you want.
`punktfunk-host` died twice on home-nobara-1 with the same stack:

    __strlen_evex <- av_vbprintf <- format_line <- av_log_default_callback
      <- ff_cuda_check <- ff_nvenc_encode_init <- avcodec_open2
      <- NvencEncoder::open <- NvencEncoder::reset <- virtual_stream

once as an outright SIGSEGV mid-session, and once as a thread wedged in that
stack so the service never answered SIGTERM and systemd escalated to SIGABRT
("State 'stop-sigterm' timed out. Aborting."). Both times a client's session was
rebuilding its encoder. The blast radius is the whole host process — every other
client's session goes with it.

The fault is in libav, not here. `ff_cuda_check` logs the failing CUDA call as
`"%s failed -> %s: %s"` using an `err_name`/`err_string` pair the error lookup
does not always fill, and glibc then walks whatever was on the stack. We cannot
patch the distro's FFmpeg, so the fix denies it the chance to format: the guard
already used by the 4:4:4 probe drops the level to AV_LOG_FATAL across the open,
and `av_log_default_callback` returns on the level check before `format_line` —
these messages are AV_LOG_ERROR. The failure is not swallowed; it still comes
back as `Err(e)` and is reported with our own context, which now says the libav
text was deliberately silenced so nobody hunts for a message that will not come.

Scoped to the `open_with` call ALONE. The ENOSYS arm immediately below recurses
into `Self::open`, and `QuietLibavLog` holds a non-reentrant global mutex —
wrapping the whole `match` would have deadlocked the intra-refresh retry.

Verified on home-nobara-1 (fc44, libavcodec 62). With CUDA made unavailable so
the open fails inside the CUDA layer, the old binary prints

    [hevc_nvenc @ ..] cuInit(0) failed -> CUDA_ERROR_NO_DEVICE: no CUDA-capable
    device is detected

— that line IS `ff_cuda_check` formatting the two `%s` — and the fixed binary
does not; both exit 1 with our error instead. A successful open is unaffected on
both the direct-SDK and the libav paths (90/90 frames, identical output size).

What this does NOT claim: the uninitialized-pointer condition itself was not
reproduced on demand — it depends on the CUDA error lookup failing to fill the
strings, and in the forced case above it filled them fine. What is demonstrated
is that the formatting call which faulted is no longer reached during the open.
`/api/v1/local/summary` (and the console card behind it) read the live-stats mode
slot, which bring-up seeded from the NEGOTIATED mode:

    let live_mode = Arc::new(AtomicU64::new(pack_mode(
        mode.width, mode.height, interval_hz(interval))));

The refresh was already corrected there — the comment says so, because KWin caps
a virtual output's rate — but the SIZE was still the request. Only a mid-stream
resize ever fixed it: the rebuild path below publishes `delivered_mode(frame..)`,
and bring-up never did.

Attach is what makes this matter rather than being pedantry. On a box with a
physical display the gamescope backend logs

    gamescope: box drives a physical display — attaching at its own mode (no
    re-mode) client_w=5120 client_h=1440

and streams the panel's size. Measured on home-nobara-1 with a 1080p HDMI panel
attached: the capture negotiated 1920x1080 and NVENC opened 1920x1080@240, while
the summary reported 5120x1440 — the console confidently naming a resolution
nobody was watching, which is exactly the shape of the stale attach-path report
noted on .41 in July ("reusing w=5120 h=1440" while the session was really 1080p).

Seeding the slot from `delivered_mode(frame.width, frame.height, interval)` uses
the same helper the rebuild path already trusts, and changes only the two fields
that were wrong — its refresh term IS `interval_hz(interval)`, so that half is
bit-for-bit what it was.

This publishes the STATS slot only. It deliberately does not send the client a
corrective `Reconfigured`: that remains owed exactly where it already was, under
`adopted_at_bringup`, because an ordinary connect's mode came from the Welcome
rather than from an accept the client has already acted on.

Verified on home-nobara-1, attach session against a 1080p panel:
  summary session:   {"width":1920,"height":1080,"fps":240}
  actually captured: pipewire format negotiated width=1920 height=1080
Before the change the same session reported 5120x1440.
fix(pf-encode): the 10-bit probe was the last ffmpeg NVENC open on a direct-SDK host
ci / bun-nix (pull_request) Successful in 24s
ci / web (pull_request) Successful in 1m3s
apple / swift (pull_request) Successful in 1m43s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m14s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m59s
android / android (pull_request) Canceled after 3m11s
ci / rust (pull_request) Canceled after 3m13s
ci / rust-arm64 (pull_request) Canceled after 3m13s
nix / flake (pull_request) Canceled after 2m4s
windows / build (x86_64-pc-windows-msvc) (pull_request) Canceled after 1m9s
ad63994cb9
`can_encode_10bit`'s Linux NVIDIA arm answered "can this GPU encode 10-bit?" by
opening an ffmpeg `hevc_nvenc` encoder. On a host that then streams over the
direct SDK, that is the LOG-3 field bug: one ffmpeg NVENC open in a direct-SDK
process wedges every later open process-wide with `NV_ENC_ERR_INVALID_VERSION`
until the host restarts.

`can_encode_444` was moved off the ffmpeg probe for exactly this reason on
2026-07-27. The 10-bit one was deliberately left behind, on the reading that
"Linux HDR genuinely rides the libav P010 path". `open_video` contradicts that:

    if cuda && nvenc_direct_enabled() {          // no 10-bit exclusion
        … NvencCudaEncoder::open(…, bit_depth, …)

A CUDA capture goes to the direct backend at whatever depth was resolved, and
`is_ten_bit_input` already accepts the packed 10-bit RGB (`X2Bgr10`) that a
gamescope HDR capture negotiates. So on a default NVIDIA host the probe was
loading ffmpeg's NVENC client for a session that never uses it.

Observed on home-nobara-1 2026-08-10, gamescope + RTX 5070 Ti, client HDR on:

    resolved session plan … bit_depth: 10, hdr: true
    pipewire format negotiated … xBGR_210LE mapped=Some(X2Bgr10) modifier=0 hdr=true
    encoder submit failed — encoder rebuilt in place … NV_ENC_ERR_INVALID_VERSION
    encoder did not recover after repeated in-place rebuilds — ending the video session

and with `PUNKTFUNK_NVENC_DIRECT=0` (nothing mixes, libav serves everything) the
same HDR session streams clean: 0 errors, bit_depth=10, hdr: true.

The 10-bit cap now rides `nvenc_cuda::probe_support()`'s existing throwaway
session — the same place the 4:4:4 cap already rides, queried per listed GUID
with `NV_ENC_CAPS_SUPPORT_10BIT_ENCODE`, which is what the Windows NVENC arm has
always done (`enc/windows/nvenc.rs`). Unanswered fails CLOSED: an 8-bit session
beats a wedged one. A host that will really serve over libav
(`PUNKTFUNK_NVENC_DIRECT=0`, or a build without `--features nvenc`) keeps the
ffmpeg probe, where it validates the actual path and ffmpeg's client is loaded
anyway.

⚠ NOT YET VALIDATED ON GLASS. Gates are green — clippy `-D warnings` with
`--features nvenc,vulkan-encode,pyrowave` on linux/amd64, 67 pf-encode tests,
fmt — but the end-to-end HDR run is still owed. This branch is 42 commits behind
main and its build cannot complete a punktfunk/1 handshake on home-nobara-1 at
all (it stalls between "audio channels resolved" and "encode bit depth" and
times out at 10 s, on EVERY attempt). That stall is NOT this change: a control
build with only the routing reverted stalls identically, and the released
0.27.0 RPM on the same box handshakes fine and reaches `bit_depth=10`. Rebase
onto main before re-testing.
enricobuehler added 1 commit 2026-08-10 10:40:15 +00:00
docs(pf-encode): Linux Main10 is live — the 'inert until Phase 5.1' comment outlived the code
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m8s
apple / swift (pull_request) Successful in 1m36s
apple / screenshots (pull_request) Skipped
ci / bun-nix (pull_request) Successful in 1m15s
ci / web (pull_request) Successful in 1m52s
ci / docs-site (pull_request) Successful in 1m58s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m26s
ci / rust (pull_request) Failing after 4m39s
ci / rust-arm64 (pull_request) Successful in 5m35s
android / android (pull_request) Successful in 5m44s
nix / flake (pull_request) Failing after 18m37s
84faeb1bf1
The bit_depth field said '8 on Linux until Phase 5.1 lands a P010 capture path'.
The code outran it: the gamescope HDR capture patches offer 10-bit BT.2020/PQ,
nvenc_fmt maps X2Rgb10/X2Bgr10 to ARGB10/ABGR10, and is_ten_bit_input flips
bit_depth and hdr from the negotiated input. Verified on home-nobara-1:
'resolved session plan ... bit_depth: 10, hdr: true' on the direct backend.

A 10-bit frame deliberately takes neither the NV12 nor the YUV444 convert (both
compute CSCs write 8-bit planes) and rides packed RGB to the encoder, which does
its own BT.2020 CSC — pf-capture/src/linux/pipewire.rs owns that gate. So Main10
needed no P010 path to arrive, and P010 is now a perf follow-up (skip NVENC's
internal CSC, as NV12 does for SDR), not the thing that makes 10-bit work.
enricobuehler merged commit 35b5ee6a36 into main 2026-08-10 10:45:26 +00:00
enricobuehler deleted branch worktree-worktree-encode-worker 2026-08-10 10:45:31 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: unom/punktfunk#153