The host cursor stops riding the video and becomes a real OS cursor on
the client (the Parsec/RDP model): pointer feel no longer pays the
capture→encode→network→decode→present round trip.
Wire (M2a):
- Hello grows a client_caps trailing byte (CLIENT_CAP_CURSOR) after the
fixed display_hdr block — presence disambiguated by remaining length,
which caps the post-HDR tail at 27 bytes (documented); Welcome answers
HOST_CAP_CURSOR (capable-and-asked, the 444/clipboard precedent).
- CursorShape (0x50, control stream): serial + dims + hotspot + straight
RGBA, ≤120px/side so the u16 frame always fits (128² would overshoot);
client caches by serial — re-showing a known shape costs 14 bytes, not
a bitmap (RDP pointer-cache for free).
- CursorState (0xD0 datagram): serial + visible/relative_hint flags +
position, sent once per encode-loop tick — latest-wins, self-healing
under loss, no refresh timer. relative_hint is reserved for M3.
- Client core: two new planes (control-task + datagram-task arms) →
next_cursor_shape/next_cursor_state; connect() grows client_caps
(C ABI passes 0 until the v11 cursor poll fns exist).
Host (M2b, Linux portal only):
- handshake::cursor_forward is THE predicate (client asked ∧ Linux ∧
compositor ≠ gamescope) — Welcome bit and session wiring both read it.
- SessionPlan.cursor_blend goes false for a forwarding session; the
encode loop ticks a CursorForwarder every iteration: shape-serial diff
→ control-task bridge (mirrors probe_result), state datagram → conn.
- CursorOverlay/capture CursorState carry the hotspot through
(nearest-neighbor downscale backstop for XL cursors, unit-tested).
Presenter:
- CursorChannel drains both planes per loop iteration; shapes become
SDL color cursors (from_surface + hotspot), applied while the desktop
mouse model is engaged; visibility follows the host; capture/released
hands back the system cursor. Sessions advertise the cap when they
START in desktop mode.
Verified on .21: fmt + clippy -D warnings (7 crates) + tests green
(core 218 incl. new wire roundtrips, host 245 incl. e2e + forwarder
downscale tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Flip both zero-copy levers from opt-in to default:
- The capture negotiation now PREFERS gamescope's producer-side NV12 pod by
default (PUNKTFUNK_PIPEWIRE_NV12=0 restores the packed-RGB negotiation).
Codec-aware gating rides a new OutputFormat::nv12_native ->
ZeroCopyPolicy::native_nv12_session edge resolved by the host facade from
the session plan (pf_encode::linux_native_nv12_ok): only H265/AV1 sessions
whose backend can open the raw Vulkan Video encoder ever see the NV12 pod
-- an H264/Moonlight session (libav VAAPI, which would misread the
two-plane buffer) keeps today's BGRx negotiation, as do the
GameStream-resolve and portal-mirror paths, PyroWave, and NVENC prefs.
- RGB-direct's unaligned modes default to the true-extent direct import
(PUNKTFUNK_VULKAN_RGB_TRUE_EXTENT=0 restores the padded-copy staging).
Guarded-tested on Van Gogh with the kernel journal watched: clean, and at
5.38 ms p50 the fastest 1080p encode path measured on that hardware. The
EFC only exists on Mesa >= 26, where the codedExtent-driven session_init
padding is guaranteed (verified back to Mesa 24.2).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
With PUNKTFUNK_PIPEWIRE_NV12=1 (bring-up gate), the PipeWire negotiation
offers an NV12 LINEAR DMA-BUF pod (BT.709 limited pinned MANDATORY) ahead
of the BGRx one, and gamescope's producer-side RGB->YUV pass replaces the
host CSC entirely: the encoder imports the two-plane buffer as a profiled
VIDEO_ENCODE_SRC image and the VCN encodes it directly. Contributed
measurements: encode p99 2.9 ms (from ~4-4.5 ms via EFC RGB-direct),
60 fps capture, 0 send drops.
Hardening on top of the contributed patch:
- unaligned modes (1080p!) stage through a padded aligned NV12 copy (edge
rows/columns duplicated, transfer-only) instead of direct-importing the
visible-size buffer -- a direct import would make the VCN read past the
producer allocation, the exact OOB class behind the 2026-07-20 field GPU
reset; the encode extents return to the aligned coded extent everywhere
- the UV plane layout honors the producer's plane-1 chunk (offset/stride)
when the SPA buffer carries one (same-BO verified by inode), with the
contiguous-plane contract as fallback
- PyroWave sessions are excluded from the gate (their Vulkan compute CSC
ingests packed RGB), and a native-NV12 session that resolves to libav
VAAPI (H264 codec, PUNKTFUNK_VULKAN_ENCODE=0, feature off, or a failed
Vulkan open) refuses at open instead of streaming garbage chroma
- pad staging images carry TRANSFER_SRC (the width-padding pass self-copies
the staging image -- previously missing on 1366-wide modes)
- metadata-cursor one-shot warn (parity with RGB-direct) and a padded-NV12
PUNKTFUNK_PERF split label; the padded RGB copy gets its timestamps too
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two session-transition stalls found live on a SteamOS Deck host, one cause:
the pipeline retry loop couldn't tell a wait-it-out failure from a
retry-now-and-it-works one.
- Gamescope first-frame race: a PipeWire stream connected while gamescope
re-initializes its headless takeover negotiates a format, reaches
Streaming, and never receives a buffer — while a fresh connect delivers
within ~0.5 s. Every gamescope bring-up ate the full 10 s first-frame
budget on attempt 1 (17 s bring-ups; KWin: 1.2 s). The retry loop's FIRST
attempt now waits 2.5 s (Capturer::next_frame_within), so the reconnect
that fixes the race happens at ~3 s. Later attempts keep the patient 10 s
— the documented 30-60 s Big Picture cold start still fits the budget.
- Capture-loss rebuild vs session switch: the rebuild loop re-detects the
active session between build_pipeline_with_retry calls, but each call
burned 8 attempts (~13 s) against a compositor that no longer exists (a
Desktop→Gaming switch spent 13 of its 27 s retrying gone-KWin). The
capture-loss path now passes max_attempts=2, turning the outer loop into
~1 s detect-and-retry cycles that follow the box to the new session.
The resize path's direct build_pipeline call keeps the default budget (no
retry wrapper there to absorb an early bail).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`ci.yml`'s clippy gate has been failing on main since the pre-0.16.0 sweep landed, and
because clippy stops at the first crate it can't compile, the visible error was only
ever the first of four. `ci.yml` is not tag-triggered, so v0.16.0 cut and shipped over
a red main; none of this reaches the release artifacts (the two E0308s are in a test
binary and a dev tool, and the rest are lints) but the gate has been blind since.
Fixed, in the order clippy surfaced them:
- clients/probe failed to COMPILE (E0308, x2). `810d918d` moved `clock_sync` onto the
resumable `io::MsgReader` but only updated the client pump, leaving the probe passing
a bare `&mut RecvStream`. The probe now wraps the control stream in a `MsgReader` at
`open_bi` and threads that everywhere — Welcome, the --remode and --bitrate watchers,
and the speed-test result read — which is also what the refactor was for: those reads
sit behind timeouts and `select!`, exactly where a straddling frame would desync the
stream for the rest of the run.
- pf-capture: `SPA_META_Cursor as u32` is a `u32 -> u32` no-op
(`clippy::unnecessary_cast`). Line 1274 already passes the same constant uncast, so
the type is not in question.
- pf-inject: `noop as usize` on the SIGUSR1 wake handler added in `986402f7` trips
`clippy::function_casts_as_integer`; goes via `*const ()` as the lint asks. Same value
in the `usize`-typed `sa_sigaction` slot.
- punktfunk-core: the `ctrl_framing` test module's `use super::*` is unused, which
`-D warnings` promotes to an error.
Verified with CI's own commands on a Linux box (this is all Linux-gated code, so a Mac
cannot check it): `cargo clippy --workspace --all-targets --locked -- -D warnings`
finishes clean, `cargo build --workspace --locked` succeeds, and
`cargo test --workspace --locked` exits 0 with no failures — including punktfunk-core's
196-test suite, which covers the control-stream framing the probe change touches.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two medium findings from the round-1 sweep.
- The `.process` callback dequeued a buffer INSIDE its `catch_unwind`, and every
requeue site was inside too. `newest` is a raw pointer with no Drop, so any
caught panic (update_cursor_meta / consume_frame) unwound past all three
requeues and permanently stranded that buffer. Once the stream's fixed pool
drained, `dequeue_raw_buffer` returned null every call and capture silently
wedged while still reporting negotiated/active — defeating the very
catch_unwind that was meant to keep a panic survivable. The drain loop now runs
OUTSIDE the catch (dequeue/queue are non-panicking C FFI pointer ops) and
`newest` is requeued exactly once after it, on every path: normal,
corrupted-skip, or caught panic.
- `recreate_ring` invalidated `video_conv` and `hdr_p010_conv` but not
`pyro_conv`. That converter is mode-baked — BgraToYuvPlanes selects entirely
different shaders and output formats for SDR (8-bit BT.709 → R8/R8G8) vs HDR
(scRGB→PQ BT.2020 → R16/R16G16) — and `ensure_pyro_conv` only builds when None,
so a display_hdr flip reused the stale SDR converter against a freshly
HDR-formatted pyro ring, corrupting every frame. Reachable at the documented
"Downgrade point D": a PyroWave session with client_10bit=true that opens on a
box where HDR can't enable, then flips once the display comes up.
Linux .21: pf-capture 1/0. Windows .173: `cargo check -p pf-capture` clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The three high-severity defects from the round-1 sweep of pf-inject /
pf-zerocopy / pf-capture (adjudicated against source — all 7 reported criticals
in these crates downgraded; these were the real highs).
- pf-inject steam_gadget: `SteamDeckGadget::drop` set `running=false` then joined
the control thread, which spends steady state parked in a blocking, no-timeout
`EVENT_FETCH` ioctl that only tests `running` at its loop top. The flag never
reaches it, closing the fd can't wake an in-flight ioctl (the syscall holds a
file reference, and the fd is shared via Arc by the very threads being joined),
so the join hung — and it runs on the session input thread via
`PadSlots::sweep`, driven by the client's `active_mask`, so a remote peer
clearing its pad bit could freeze all session input. Now wakes the parked
threads with SIGUSR1 (no-op, non-SA_RESTART handler → the ioctl returns EINTR
and the loop exits), retried until each reports done and bounded (~1s).
- pf-zerocopy cuda: the GL→CUDA "sync point" was never established for the copy.
`cuGraphicsMapResources`/`UnmapResources` were issued on the NULL stream, but
the D2D copy runs on `copy_stream()`, a `CU_STREAM_NON_BLOCKING` stream exempt
from implicit NULL-stream ordering — and the GL de-tile/CSC that produced the
texture ends with only `glFlush` (no fence). So the copy could race ahead of
the not-yet-retired GL draw: intermittent stale/torn frames under GPU load, on
the default NVIDIA capture→encode path. Map, copy, and unmap now share
`copy_stream()`, so map's device-side guarantee orders the GL work before the
copy. Zero-copy preserved (no GPU→CPU→GPU roundtrip).
- pf-capture cursor meta: `update_cursor_meta` trusted three producer-written
fields (bitmap_offset, pixel offset, stride) with no bound against the metadata
region, driving OOB pointer arithmetic and an oversized `from_raw_parts` — an
OOB read that SIGSEGVs inside the PipeWire `.process` callback (uncatchable by
the surrounding `catch_unwind`). Switched to `spa_buffer_find_meta` to obtain
the region's real `size` and validate every offset with checked arithmetic
before each deref/slice, mirroring the fd-length guard the main frame path
already applies.
Compile + existing tests green on Linux .21 (real RTX 5070 Ti): pf-inject 74/0,
pf-zerocopy 17/0, pf-capture 1/0. The gadget deadlock path only executes on a
SteamOS host with raw_gadget/dummy_hcd (not reproducible on the CachyOS box), so
that fix is reasoned + compile-verified, not runtime-exercised.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A PyroWave session on an NVIDIA-auto host was forced onto CPU-RGB capture
(session_plan flipped gpu=false): Mutter blits tiled->LINEAR, we mmap +
de-pad ~30 MB, the encoder re-uploads it - three full-frame CPU touches
per frame at 5120x1440 while an HEVC session on the same box rides the
tiled EGL/CUDA zero-copy. The dmabuf passthrough + Vulkan tiled import
were already validated (8dc5d672) but only reachable via the global
PUNKTFUNK_ENCODER=pyrowave lab policy.
ZeroCopyPolicy gains pyrowave_session (from OutputFormat.pyrowave, i.e.
the negotiated codec): the capturer skips the NVENC-only EGL->CUDA
importer, takes the raw-dmabuf passthrough, and advertises the wavelet
encoder's Vulkan-importable modifiers so Mutter+NVIDIA negotiates tiled
zero-copy. The forced-CPU flip in session_plan is gone.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The GNOME 50 HDR format offer took the PQ transfer id straight from
pw::spa::sys, which only exists on libspa new enough to carry the
BT2020_10/SMPTE2084/ARIB_STD_B67 block. Ubuntu 24.04 (noble) — the .deb host
builder — ships older headers, so bindgen emitted no such constant and the
host failed to compile there:
error[E0425]: cannot find value `SPA_VIDEO_TRANSFER_SMPTE2084`
in crate `pw::spa::sys`
This never showed up locally or on the Linux CI: both run a current PipeWire,
where the binding is present. It broke deb.yml's build-publish-host job, so
v0.14.0 published its client .debs but no host .deb.
Spell the id out (14) instead. It's wire ABI, not a private detail — SPA
mirrors GStreamer's GstVideoTransferFunction and that block was added as a
unit, so the value is the same on every libspa that has the symbol. On one
that doesn't, PipeWire fails to intersect the offer and the session
negotiates SDR, which is what an HDR-incapable host should do anyway (the
path needs GNOME 50+ regardless).
A test pins our value against pw::spa::sys wherever the symbol exists, so a
renumbered enum fails loudly instead of silently mis-tagging the transfer
function. It only builds where tests are compiled — the .deb/.rpm builders
run plain `cargo build`, so it can't reintroduce the failure it guards.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
`feat(hdr): GNOME 50 HDR screencast capture` (0e977817) landed with rustfmt
drift — six files were not clean under the pinned 1.96.0 toolchain, so
`cargo fmt --all --check` (ci.yml "Format") is red on main. Pure whitespace/
wrapping from `cargo fmt --all`; no semantic change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GNOME 50 (Mutter MR 4928, PipeWire >= 1.6) added HDR screen sharing for
monitor streams: 10-bit PQ formats (xRGB_210LE/xBGR_210LE) with MANDATORY
BT.2020 + SMPTE-2084 colorimetry props, advertised while the mirrored
monitor is in BT.2100 colour mode. Wire the Linux host into it end-to-end
on the GameStream desktop-mirror path (PUNKTFUNK_VIDEO_SOURCE=portal):
* pf-frame: PixelFormat::X2Rgb10/X2Bgr10 (DRM XR30/XB30; X2Bgr10 is the
Windows Rgb10a2 layout) + fourccs.
* pf-capture: want_hdr portal offer — HDR-only LINEAR-dmabuf pods with
MANDATORY PQ/BT.2020 props (SHM excluded: Mutter's SHM record path
paints 8-bit ARGB32 regardless of format; tiled excluded: the EGL
de-tile blit is 8-bit RGBA8), negotiated-colorimetry parse, generic
HDR10 hdr_meta(), packed-10-bit CPU cursor blend, a process-wide SDR
downgrade latch on negotiation timeout, and a DisplayConfig BT.2100
colour-mode probe (gnome_hdr_monitor_active).
* pf-encode: libav NVENC X2RGB10->P010 swscale (BT.2020 limited) ->
HEVC Main10 / 10-bit AV1 with PQ VUI; VAAPI 10-bit on both paths (CPU
P010 upload + dmabuf XR30 scale_vaapi p010/bt2020); can_encode_10bit
now probes for real on Linux; 10-bit sessions route around the
8-bit-only Vulkan-video/direct-NVENC backends.
* GameStream: host_hdr_capable() Linux arm, live monitor-HDR check at
RTSP honor time, capturer-pool reuse keyed on HDR-ness, gs_bit_depth
covers the new formats. New `punktfunk-host hdr-probe` diagnostic and
a PUNKTFUNK_SPIKE_HDR spike lever.
* Native plane stays honestly 8-bit via capturer_supports_hdr(): Mutter
RecordVirtual streams are SDR-only upstream (GNOME 50 and 51-dev), so
virtual-display sources cannot deliver HDR yet.
Validated on the RTX 5070 Ti (GNOME 50.3 / PipeWire 1.6.8): the Main10
probes pass and the ignored nvenc_hdr10_smoke GPU test emits an IDR that
ffprobe reads as Main 10 / yuv420p10le / bt2020nc / smpte2084 / limited.
Live HDR capture negotiation still needs an HDR monitor on glass; VAAPI
10-bit needs the AMD box.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
design/latency-reduction-2026-07.md T2.5's Linux half: the LINEAR dmabuf path
(gamescope's only offer) fed NVENC RGB, paying its internal RGB->YUV CSC on
the SM the game is saturating — the exact contention §5.A removed everywhere
else. The Vulkan bridge now carries a buffer-to-buffer RGB->NV12 compute
shader (rgb2nv12_buf.comp, BT.709 limited, coefficient-identical to
pf-encode's rgb2yuv.comp; whole-word writes so no 8-bit-storage feature is
needed): import dmabuf -> dispatch CSC into the exportable buffer -> CUDA
de-strides both planes into a pooled two-plane NV12 buffer. PUNKTFUNK_NV12
(default-on) now covers LINEAR; a CSC failure latches RGB for the stream
(mid-frame fallback, no dropped frame); 4:4:4 LINEAR sessions stay RGB (never
silently subsample). New ImportKind::LinearNv12 rides the existing worker IPC
(appended last per the wire-tag rule); cursor stays downstream (blend_nv12).
Validated: .21 clippy -D warnings (pf-zerocopy/pf-capture/host+nvenc) + 17
zero-copy tests. Owed: on-glass gamescope session (visual + dmon sm% check).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
design/latency-reduction-2026-07.md tier 1, remaining halves:
- T1.1: the native encode loop wakes on the capture's ACTUAL arrival instead of
sampling at a free-running tick — deletes the sample-and-hold (~half a frame
interval on average, a full one worst-case: ~8ms avg @60fps). New
Capturer::supports_arrival_wait/wait_arrival pair (IDD-push waits its
frame-ready event against the shared-header token; the PipeWire portal blocks
its channel with a pending stash); backends without an arrival signal — and
PUNKTFUNK_FRAME_DRIVEN=0 — keep the legacy tick bit-identically. A
0.9x-interval rate floor caps encode at ~1.11x target when the compositor
outruns the session; a +0.5x-interval keepalive keeps static desktops
re-encoding at 1.5x-interval cadence. Pacing deadlines re-anchor to the
actual submit so they can't drift against the arrival clock. GameStream
plane untouched.
- T1.4: the jump-to-live detectors run on WALL-CLOCK now (STANDING_TIME /
FLUSH_AFTER = 250ms) instead of 30-frame counts whose meaning scaled with
fps (500ms @60 but 125ms @240 — and stretching further under T1.1's slower
static-scene repeats). The queue trip also requires depth still >= high, so
a hysteresis-band hover can't fire on elapsed time alone.
Validated: .21 Linux 185 core + 177 host + pf-capture tests, clippy
-D warnings; .133 Windows cargo check of pf-capture + punktfunk-host green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
capture/linux (PipeWire portal) + capture/windows (IDD direct-push: dxgi
mechanics, idd_push + submodules, synthetic_nv12) + pwinit move into
crates/pf-capture behind the Capturer trait + synthetic sources (plan §W6).
The crate speaks pf-frame (CapturedFrame/PixelFormat + the DXGI identity),
pf-zerocopy (CUDA import), and the pf-win-display leaves, and NEVER pf-encode —
the capture->encode edge is one-way. This completes the deliberate capture/encode
crate split (the invasive path the plan had merged into one pf-media): capture
and encode are now separate subsystem crates sharing only pf-frame.
Four seams keep the capturer off the orchestrator:
- VirtualOutput is EXPLODED into primitives (remote_fd/node_id/preferred_mode/
keepalive) by the host facade, so pf-capture never depends on the vdisplay type;
- FrameChannelSender: the sealed-channel delivery is a Send+Sync closure the host
facade builds from the pf-vdisplay control device + send_frame_channel IOCTL and
hands in; ChannelBroker holds the closure instead of the control HANDLE (the
whole-desktop handle-duplication security boundary is byte-for-byte unchanged);
- console_session_mismatch + desktop_bounds live in pf-win-display (leaf peers);
- pwinit moves here (audio caller -> pf_capture::pwinit).
The host keeps capture.rs as a thin BRIDGE: it re-exports the vocabulary + capturer
types (every crate::capture::* path is unchanged) and keeps open_portal_monitor /
capture_virtual_output, which resolve the ZeroCopyPolicy + FrameChannelSender and
call into pf-capture. verify_is_wudfhost + install_gpu_pref_hook are re-exported
(the gamepad-channel bootstrap + the main.rs subcommand consume them).
Co-developed: the resident-HID-mouse compose-kick hook (HID_COMPOSE_KICK + the
HID-first cursor kick + _display_wake) rides this commit into pf-capture; the host
mouse_windows registration side lands separately on top.
Verified: Linux clippy -D warnings (pf-capture + host nvenc,vulkan-encode,pyrowave
--all-targets) + host tests 299/299; Windows clippy -D warnings (pf-capture
--all-targets + host nvenc,amf-qsv --all-targets) Finished exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>