PUNKTFUNK_NVENC_ASYNC gains a real tri-state: 1 = always (as before, now
documented as ~+1 tick at depth-1), 0 = never (vetoes escalation), unset =
ADAPTIVE — off until the session loop's cadence-overrun detector escalates.
The host loop's adaptive-depth leaky bucket grows a second stage: once the
capturer's depth is maxed (Linux portal is permanently depth-1), it asks
the encoder for pipelined retrieve via the new Encoder::set_pipelined hook
(asked exactly once; default impl declines, Windows untouched).
nvenc_cuda engages at a safe point via a clean session rebuild WITHOUT the
IO-stream binding: with input==output stream bound, later stream work
waits on prior encode completions and would serialize a pipelined session
— stream-ordered submit and two-thread retrieve are mutually exclusive.
The ordered gate now also requires async_rt absence (belt-and-braces for
the runtime switch). Re-open's first frame is the standard session IDR.
On-hardware test: escalate mid-session → retrieve thread live, binding
gone, all AUs deliver, first post-escalation AU is the IDR.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
design/vulkan-rgb-direct-encode.md B1: when the probe passes AND
PUNKTFUNK_VULKAN_RGB_DIRECT=1 (default OFF until the B2 on-glass A/B),
the session opens with pictureFormat=B8G8R8A8 and the captured RGB frame
becomes the encode source — the VCN EFC front-end does the 709-narrow CSC
inline during encode. Deleted from the per-frame path: the compute CSC
dispatch, both NV12 plane copies, the semaphore hop and one queue submit
(dmabuf frames are ONE encode-queue submit; ~17 MB/frame of GPU traffic
gone). The DPB stays NV12; RFI/RC/quality-level machinery is untouched.
Shape:
- probe_rgb_direct now returns the chroma-siting bits to create the
session with (midpoint preferred, else cosited-even per axis); the open
verdict line gains "active" / "available(off; ...)" states.
- RgbProfileStack: the rgb-chained video profile rebuilt on the stack per
profiled-image creation after open (profile identity is by value) —
dmabuf imports become profiled VIDEO_ENCODE_SRC images (import cache
unchanged), the CPU staging image likewise (concurrent encode+compute).
- record_submit_rgb: steps 2–4 twin (dmabuf: single submit; CPU: staging
copy on the compute queue, semaphore-ordered); shared step-1/bookkeeping.
- begin_encode_cmd takes a SrcAcquire (CSC general / fresh-import FOREIGN
QFOT / cached visibility-only / staging TRANSFER_DST) so both paths
share the encode recording.
- RGB-direct frames skip the CSC per-slot resources entirely (make_frame
split into csc/common halves); cursor bitmaps warn once (EFC cannot
composite; gamescope — the flagship — embeds the cursor itself).
On-glass (780M RADV PHOENIX, host Mesa 26.0.4): all four smokes pass
(vulkan_smoke{,_av1,_rgb,_rgb_av1}); the EFC-encoded H265 stream decodes
clean and matches the CSC-encoded stream at 49.9 dB average PSNR (min
48.9) — within the design's ±1-code-value tolerance. (Raw-OBU ffmpeg
probing of the tiny AV1 dumps fails identically for BOTH paths —
pre-existing dump quirk, not a stream defect.) check+clippy+full unit
suite green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
PUNKTFUNK_NVENC_SLICES=N (2..=32, default off) splits H.264/HEVC frames
into N slices (sliceMode 3); PUNKTFUNK_NVENC_SUBFRAME=1 (default off)
arms enableSubFrameWrite + reportSliceOffsets on sync sessions only.
Both experimental groundwork for sub-frame slice output (plan §7 LN1).
nvenc_cuda_subframe_slice_probe (on-hardware, ignored) answers the LN1
go/no-go: spins lock_bitstream(doNotWait) against an in-flight frame and
prints the (t_us, status, numSlices, bytes) timeline — incremental slice
availability and its spacing, or all-at-completion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
On-glass B0 result from the 780M (RADV PHOENIX, host Mesa 26.0.4): the
probe returned no-709-narrow-midpoint because RADV advertises
xChromaOffsets = COSITED_EVEN only — the VCN EFC does canonical H.26x
left-cosited x-chroma, while the requirement was written to bit-match our
2x2-average (midpoint) shader. The delta is a half-pel chroma-x phase,
imperceptible and unsignalled in our bitstream; EFC's siting is arguably
the more correct one. Accept either siting bit per axis (model + range
stay exact: 709 narrow); B1 will pass the preferred available bit.
With this the 780M probes rgb_direct=available; both GPU smoke tests
(H265 + AV1) pass on that box against host Mesa 26.0.4 — the first
real-RADV validation of the poll fix, the quality-level control, and B0.
(Container mesa alone can't run the H265 smoke: Fedora ships RADV with
H264/H265 encode compiled out — AV1 only. Use the host ICD via
VK_ICD_FILENAMES on Atomic boxes.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bind the session's IO streams to the encode thread's high-priority copy
stream: in sync-retrieve depth-1 use the per-frame input copy and cursor
blend now enqueue with NO cuStreamSynchronize and encode_picture orders
after them on the stream. Same stream both directions, so the encode's
completion is inserted into the stream and later work (the next frame's
copy into a reused ring slot) waits for it.
Soundness gate: the fast path engages only when pending is empty (true
depth-1 usage) — every prior encode was drained by a blocking poll, and
the caller holds the frame payload across the matching poll (contract now
documented on Encoder::submit; both host loops already comply). Pipelined
callers and PUNKTFUNK_NVENC_ASYNC mode keep the blocking copies.
True zero-copy input registration (registering the worker-owned IPC
buffer directly) stays the LN2 v2 follow-up — it needs a contiguous
worker-pool NV12 layout and a registration<->IPC-mapping lifetime tie.
PUNKTFUNK_NVENC_STREAM_ORDERED=0 restores the old blocking behavior.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Latency plan §7 LN0 (Linux/NVIDIA encode follow-on):
- sampled PUNKTFUNK_PERF submit split (copy/blend/map/pic) in nvenc_cuda —
the host loop's submit_us folds all four together; the D2D input copy is
the LN2 zero-copy target and now measurable on its own
- sampled blocking lock_bitstream timing on pf-nvenc-out — in two-thread
mode the host loop's wait_us wraps a non-blocking poll, so the real
encode wait was measured by no timer
- caps probe + log SUPPORT_SUBFRAME_READBACK / SUPPORT_DYNAMIC_SLICE_MODE
(LN1 sub-frame slice-output prerequisites, fleet visibility)
- explicit zeroReorderDelay=1 in the shared low-latency config (P-only +
no lookahead has no reordering anyway; pins the bit against preset or
driver drift; shared with the Windows backend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
First phase of design/vulkan-rgb-direct-encode.md (punktfunk-planning): make
the captured BGRx dmabuf the direct encode source with the VCN EFC front-end
doing the 709-narrow CSC — deleting the per-frame compute CSC, both plane
copies, the semaphore hop and one queue submit.
B0 changes nothing about the encode path. It vendors the extension surface
(vk_valve_rgb.rs — ash 0.38 predates it; same rationale and style as the
vk_av1_encode module) and probes at open whether this host qualifies:
extension present (Mesa >= 26.0 + EFC hardware) → feature bit → conversion
caps cover the compute shader's exact math (709 / narrow / midpoint both
axes) → encode-src format set offers B8G8R8A8 with DRM-modifier tiling.
The verdict is one INFO line (rgb_direct=available | first missing
requirement) — the field telemetry that decides where B1 can default on.
Verified: cargo check + clippy clean + unit tests green (Linux box); the
probe short-circuits to no-ext on non-RADV drivers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Field follow-up to 6dc195f9 (the AMD host whose H265 encode sits at ~5 ms):
the remaining time is the VCN ASIC preset, and we never chose one.
1. Explicit encode quality level. RADV only emits a VCN preset op when the
app issues ENCODE_QUALITY_LEVEL — we never did, so every session ran the
firmware's default preset. Install the resolved level on the first
frame's RESET control (chained ahead of the CBR state, which satisfies
the spec's quality-change-carries-rate-control rule) and bake the same
level into the session-parameters object (the spec requires the match).
Default 0 = the fastest tier the driver exposes (RADV: SPEED, except
H265 on pre-RDNA4 which the driver pins to BALANCE); clamped to the
profile's maxQualityLevels. PUNKTFUNK_VULKAN_QUALITY=0..3 overrides for
quality-biased setups. Logged at open so field logs show the tier.
2. Persistent bitstream mapping. read_slot vkMapMemory/vkUnmapMemory'd the
host-visible bitstream buffer every frame; map each ring slot once at
build instead (coherent memory; vkFreeMemory implicitly unmaps).
3. PUNKTFUNK_PERF CSC split. GPU timestamps bracket the compute batch
(import barriers + cursor prep + CSC dispatch + plane copies) and a
sampled csc_us line logs every ~2 s — separating shader/copy time from
the ASIC encode inside the pump's wait_us, the VAAPI submit-split
analog for this backend. Off (and unrecorded) unless PUNKTFUNK_PERF.
Verified: cargo check + clippy + DPB/RPS unit tests green (Linux box).
The GPU smoke tests still need a RADV host (NVIDIA driver fails session
open on the build box, pre-existing).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The session pump's depth-1 loop (capture → submit → poll) treats a poll()
None as "the backend holds the frame internally, re-poll next tick" —
true only of the libav AMF/QSV wrappers, whose codecs genuinely hold ~2
frames. The Vulkan Video backend probed its slot fence with
get_fence_status instead, so the AU of every frame — finished on the
ASIC in ~5 ms — sat unharvested until the tick AFTER the next frame was
submitted: encode_us read as one full frame period (~17 ms at 60 Hz vs
VAAPI's ~5.3 ms on the same VCN, the AMD field report) and every frame
shipped a frame period late — one frame of avoidable glass-to-glass
latency on the path that exists to beat libav VAAPI.
Wait the oldest in-flight slot's fence in poll() instead, matching the
sync NVENC backend's blocking lock_bitstream and the documented
Capturer::pipeline_depth contract ("capture → submit → poll-blocks").
Bounded by ENCODE_FENCE_TIMEOUT_NS like enqueue()'s backpressure wait,
for the same reason: poll runs on the thread the stall recovery runs on,
so a wedged GPU must surface as an error, not park it. None now only
means "nothing submitted". Covers H265 and AV1 (shared poll).
Verified: cargo check + DPB/RPS unit tests green (home-worker-5); the
GPU smoke tests fail at session open on that box's NVIDIA driver
identically on unmodified origin/main — pre-existing, needs the RADV
box for on-glass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
VulkanVideoEncoder::open_inner creates ~20 Vulkan objects across ~15
fallible steps, but all cleanup lived in the encoder's Drop — which only
runs once the value exists at the final Ok(Self). Any earlier ?/bail!
leaked everything built so far (a VkDevice + GPU memory per retried
open, and this backend is the default encode path on AMD/Intel Linux
hosts where open can fail transiently).
Factor the entire teardown sequence — unchanged — into a VkTeardown
guard whose Drop destroys any prefix of the build (vkDestroy*/vkFree*
are defined no-ops on VK_NULL_HANDLE): open_inner mirrors each object
into the guard as it is created and disarms it only at Ok(Self); the
encoder's own Drop rebuilds one from its fields, so both paths share
one sequence and cannot drift. make_frame now builds in place into a
guard-parked null Frame so a mid-build failure unwinds its partial
handles too, and make_video_image / vk_util::make_plain_image (also
used by the PyroWave backend) / build_parameters_h265 no longer leak
their own partially-created objects on failure.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Completes the partial fix from the previous commit. The Windows PyroWave backend
caches its imported plane images keyed on the D3D11 texture's COM address, and
holds no reference on that texture — so once the capturer recreates its ring,
those addresses can be handed straight back out by the allocator and a
pointer-keyed cache hit returns an image bound to a texture that no longer
exists. Adding the extent to the key ruled out same-address-different-size
aliasing, but a recycle at identical dimensions still aliased.
The capturer already tracks exactly the value needed: `generation`, bumped on
every ring recreate. Plumbed it onto `PyroFrameShare` and the encoder now
flushes every cached import when it changes, which makes cache identity
independent of allocator behaviour rather than a bet against pointer reuse.
Validated on the RTX box: `pyrowave_win_smoke` (forced with `--ignored`, the
only test that actually exercises this path on real hardware) passes all ten
configurations — 1024²/720p/1080p/1440p across SDR/HDR and 4:2:0/4:4:4 — with
correct decoded chroma means, so the steady-state cache-hit path still works.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The Windows direct-NVENC backend registers and encodes the capturer's textures
IN PLACE (no CopyResource), so how deep it may pipeline is a property of the
CAPTURER, not of the encoder. It was bounded only by `async_inflight_cap()` —
`PUNKTFUNK_NVENC_ASYNC_DEPTH`, default 4, clamped to the output-bitstream pool
— which consults nothing about the capturer, while the comment at the
backpressure loop claimed it "keep[s] in-flight depth within the capturer's
texture ring". It never did.
The IDD-push capturer rotates `OUT_RING = 3` per delivered frame with no regard
for encode completion (its own invariant note says OUT_RING(3) > max
pipeline_depth(2)). With the default async depth of 4 the encoder can therefore
still be reading a texture the capturer has already handed out again and
overwritten: torn or mixed frames. It is visual corruption rather than UB, so it
fails silently and intermittently — the worst shape to diagnose from a field
report.
Adds `Encoder::set_input_ring_depth`, reported from `Capturer::pipeline_depth`,
and bounds the async backpressure loop by `min(async_inflight_cap(), depth)`.
For IDD-push that yields 2, matching the capturer's stated contract; backends
that copy their input, or are synchronous, ignore it.
Wired at ALL THREE encoder-creation sites (initial open, stall/resize rebuild,
ABR rebuild) and forwarded through `TrackedEncoder` — this crate has a
documented trap where an unforwarded defaulted trait method silently no-ops
through that wrapper, which has already bitten the direct-NVENC work once and
the wire-chunking probe once.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Five of the nine medium findings from the pf-encode sweep. The remaining four
need cross-crate plumbing or an unwind refactor and are deliberately left out.
- vulkan_video `enqueue` waited on the backpressure fence with `u64::MAX`. That
wait runs ON the host encode thread — the same thread the stall watchdog's
`reset()` would run on — so a wedged GPU parked the one thread that could
recover the session: no error, no reset, and teardown blocking on the join.
This is the DEFAULT encode path for AMD/Intel Linux hosts (both shipped build
recipes enable `vulkan-encode` and `vulkan_encode_enabled()` defaults true).
Now bounded by ENCODE_FENCE_TIMEOUT_NS with expiry surfaced as an error.
- vulkan_video `import_cached` evicted a cached dmabuf import and destroyed its
image/view/memory with no fence wait, while up to `ring_depth - 1` submitted
frames may still reference it — a GPU-side use-after-free. `Drop` and `reset`
both idle first; this was the one unguarded destroy. Now idles before the
eviction loop, guarded on the length test so the steady state pays nothing.
- vulkan_video `read_slot` never asked for the encode's operation status, so a
FAILED encode was indistinguishable from a successful one and its feedback
was read as if it described real bitstream. Now requests WITH_STATUS_KHR and
refuses anything that is not COMPLETE.
- linux/pyrowave `encode_frame` opens its recording window early and has six
fallible steps inside it; every one returned with `cmd` still RECORDING, and
nothing repaired it (one `begin_command_buffer` in the file, and neither
`reset()` nor `Drop` touches `cmd`), so the next frame called `begin` on a
recording buffer — invalid usage. `submit` now resets the buffer on error;
legal on all these paths since the pool carries RESET_COMMAND_BUFFER and the
buffer is not pending.
- windows/pyrowave `encode_frame` ignored `frame.width`/`height` and imported
planes at the encoder's configured extent, so a ring recreate at a new mode
(the IDD capturer does this autonomously on a confirmed descriptor change)
read the planes under a stale VkImageCreateInfo. Added the size guard its QSV
and AMF siblings already carry, and keyed the plane cache on
(address, width, height) so a recycled COM address cannot resurrect an import
of a different size. NOTE: a recycle at the SAME size is still theoretically
possible; the complete fix keys on the capturer's ring generation and needs
that plumbed onto `PyroFrameShare`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Three defects from the pf-encode sweep, each adjudicated against source.
- Vulkan Video never received fecbec2d's taint sweep (it was carved out one
commit later). `pick_recovery_slot` accepts any resident slot whose wire is
below the CURRENT loss start, but "resident and older than this loss" is not
"the client decoded it": after an earlier loss [a,b] recovered at wire r,
everything in [a, r-1] is undecodable at the client — the lost frames plus
every frame that predicted through the gap — and those wires stay eligible
until the 8-slot ring rolls them out. A later loss could therefore anchor on
one and ship it tagged `recovery_anchor`, which is the client's definitive
re-anchor signal (punktfunk-core/src/reanchor.rs): the host lifts the client's
post-loss freeze onto a picture built from a reference it never had. Swept
before anchor selection, matching AMF/QSV.
`slot_wire` is blanked and `slot_poc` deliberately is NOT: `slot_poc` feeds
`build_h265_rps_s0`, which must keep naming every physically-resident DPB
picture or a conforming decoder evicts them and the anchor then references a
picture the client already dropped.
- QSV's sweep was incomplete, and in its MODAL case. `ltr_slots` mirrors the
hardware DPB, but nulling an entry issues no VPL call — the frame stays marked
long-term until that LongTermIdx is re-marked or an IDR flushes it (amf.rs
states this verbatim). The rejection loop iterates the post-sweep mirror and
only rejects `Some` slots, so it silently skipped the single entry the sweep
exists to distrust, leaving the recovery frame free to predict from it. With
NUM_LTR_SLOTS=2 the "exactly one slot swept" case is the common one, and the
two existing tests cover only the both-survive and both-swept cases. Taint is
now recorded in `ltr_tainted` with the FrameOrder left in place, so anchor
selection and the queued-force guard skip it while the rejection list still
names it.
- `read_slot` built a slice from the driver-reported (offset, bytes-written)
encode feedback with no validation against `bs_size`, so a driver reporting a
range outside the bitstream buffer produced an out-of-bounds read shipped
straight onto the wire. Checked in u64 (so the add cannot wrap) before
`map_memory`, so the error path has no unmap to unwind.
Adds `taint_sweep_excludes_slots_from_an_earlier_loss` covering the two-loss
case the existing single-loss test does not reach.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
All four found in the pf-encode quality sweep and verified against source.
- NVENC partial-init leak (BOTH platforms, high): `init_session` publishes
`self.encoder` — and on Windows charges LIVE_SESSION_UNITS — *before* its
remaining fallible steps (bitstream buffers; on Linux also the input-surface
alloc and `register_resource`). A failure there left a live session with
`inited == false`, and every guard on the re-init path keys off `inited`, so
the next submit skipped teardown and overwrote `self.encoder`: the session
leaked permanently toward the driver's per-process cap, and its budget units
never returned, progressively starving parallel-display admission. `teardown`
already keys off `encoder.is_null()` rather than `inited`, so it cleans up
exactly this half-built state — it just was never called. Now invoked on the
`init_session` error path on both platforms.
- `can_encode_10bit` asked the wrong backend (medium): it resolved via
`linux_auto_is_vaapi`, which ignores `encoder_pref`, while `can_encode_444`
and `open_video` honour it. On a host that forces a backend (e.g.
`encoder_pref = "vaapi"` on an NVIDIA box) the probe answered for NVENC while
the session opened VAAPI, so the negotiated bit depth — and the HDR/SDR colour
label derived from it — described a backend that never ran. Now uses the same
`linux_zero_copy_is_vaapi` mirror, and `linux_auto_is_vaapi` carries a warning
that it resolves the `auto` case only and is not a dispatch mirror.
- Linux software arm ignored SW_BITRATE_CEIL (low): the Windows arm clamped
openh264 to 100 Mbps, the Linux arm passed the full negotiated rate. The
constant is now module-scope so both arms share one value.
- QSV/AMF env-parity (low): `PUNKTFUNK_IR_PERIOD_FRAMES` was a no-op on QSV
despite the comment claiming parity with AMF, and `PUNKTFUNK_NO_QSV_LTR` /
`PUNKTFUNK_INTRA_REFRESH` had dropped AMF's `trim()` and `yes`/`on` spellings,
so a value with stray whitespace silently did nothing on Intel while the same
value worked on AMD.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Both surfaced in a post-refactor quality sweep of pf-encode and were then
verified against the source (and, for pyrowave, against the C side).
- PyroWave (BOTH platforms): `reset()` destroyed the encoder and, when the
rebuild failed, returned `false` leaving `pw_enc` pointing at the freed
object — `Drop` then destroyed it a second time. `pyrowave_encoder_destroy`
is a plain `delete` (pyrowave_c.cpp:1184, which also reads `encoder->device`
afterwards) with no null check, so this is a real double free. The failure
branch is not vacuous: the rebuild fails when the device is lost/OOM, which
is exactly the state that makes the stall watchdog call `reset()` in the
first place, so the host corrupts its heap on the path that runs when things
are already going wrong. Now nulls `pw_enc` before the fallible create,
publishes only on success, and null-guards both `Drop` and `encode_frame`
(the Windows `Drop` already guarded `sync` this way).
- QSV: `reset()` dropped `pending` — each entry owning the `Box<BsBuf>` the
runtime writes into asynchronously — BEFORE `MFXVideoENCODE_Close` aborted
those operations, so the VPL runtime could write into freed heap. The
preceding drain is best-effort and bails on the first `Err`, i.e. precisely
the wedged-encoder case that triggers the reset. Fixed by ordering: Close,
then clear. The full-teardown path was already correct (`Inner` declares
`session` before `pending`, and fields drop in declaration order).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Field report (lid-closed Intel laptop, ~6-19% sustained loss): the stream
never healed — permanent macroblock soup. Three stacked bugs:
- QSV answered RFI with PreferredRefList only, a reorder HINT per the VPL
spec — the recovery frame could keep predicting from tainted short-term
refs. Now rejects every other DPB candidate (RejectedRefList) and caps
L0 at one active entry (AVC/HEVC), matching AMF's hard
ForceLTRReferenceBitfield / NVENC invalidation semantics.
- Neither QSV nor AMF taint-swept LTR slots across losses: a slot marked
inside the client's corrupt window became the "known-good" anchor of the
NEXT loss, propagating corruption through every recovery. Both now drop
slots at-or-after the loss start before picking an anchor, and guard a
queued force whose slot the sweep emptied (no false recovery_anchor tag).
- The native plane re-anchored the FULL IDR cooldown on every successful
RFI, so under sustained loss the client's escalating keyframe requests
were coalesced away indefinitely (field log: dozens swallowed, one IDR
per ~8 s). RFI now anchors a 300 ms echo window with a 2-swallow budget
per loss episode; a client still asking past that gets its IDR.
Live-validated on Arc (qsv feature): 6/6 including the new
qsv_live_ltr_rfi_taint_sweep_declines (a loss covering every live mark
declines the RFI and falls back to IDR recovery).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The GNOME 50 HDR work added PixelFormat::X2Rgb10 / X2Bgr10 but only taught the
Linux encoders about them. `sws_src` in the Windows-gated ffmpeg_win.rs matches
PixelFormat exhaustively, so the Windows host stopped compiling:
error[E0004]: non-exhaustive patterns: `X2Rgb10` and `X2Bgr10` not covered
--> crates\pf-encode\src\enc\windows\ffmpeg_win.rs:132:14
Linux CI never caught it — the file is cfg(windows), so `cargo clippy
--workspace` on the Linux runner never compiles it.
Both are Linux-only screencast formats (the Windows HDR path stays
Rgb10a2/P010, per the PixelFormat docs), so they join the existing bail arm.
Spelled out rather than folded into a `_` catch-all so the next PixelFormat
addition breaks this match again on purpose.
Verified: cargo check --workspace --all-targets --features nvenc,amf-qsv on the
Windows box (192.168.1.173).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
`cargo clippy --workspace --all-targets --locked -- -D warnings` was red on
main — three lints landed with the GNOME 50 HDR + PyroWave 4:4:4 work:
* pyrowave_wire.rs: `aw / 2 >> level` tripped clippy::precedence. Rust already
binds `/` tighter than `>>`, so this always parsed as `(aw / 2) >> level`
(subband dim at half res, then one halving per DWT level) — the parens are
purely explicit, no change in behaviour.
* linux/mod.rs: `probe_can_encode_10bit` sat after `mod hdr_tests`
(clippy::items_after_test_module) — moved above the test module, unchanged.
Lint-only; no functional change. fmt/clippy/test all green afterwards.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
`feat(hdr): GNOME 50 HDR screencast capture` (0e977817) landed with rustfmt
drift — six files were not clean under the pinned 1.96.0 toolchain, so
`cargo fmt --all --check` (ci.yml "Format") is red on main. Pure whitespace/
wrapping from `cargo fmt --all`; no semantic change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GNOME 50 (Mutter MR 4928, PipeWire >= 1.6) added HDR screen sharing for
monitor streams: 10-bit PQ formats (xRGB_210LE/xBGR_210LE) with MANDATORY
BT.2020 + SMPTE-2084 colorimetry props, advertised while the mirrored
monitor is in BT.2100 colour mode. Wire the Linux host into it end-to-end
on the GameStream desktop-mirror path (PUNKTFUNK_VIDEO_SOURCE=portal):
* pf-frame: PixelFormat::X2Rgb10/X2Bgr10 (DRM XR30/XB30; X2Bgr10 is the
Windows Rgb10a2 layout) + fourccs.
* pf-capture: want_hdr portal offer — HDR-only LINEAR-dmabuf pods with
MANDATORY PQ/BT.2020 props (SHM excluded: Mutter's SHM record path
paints 8-bit ARGB32 regardless of format; tiled excluded: the EGL
de-tile blit is 8-bit RGBA8), negotiated-colorimetry parse, generic
HDR10 hdr_meta(), packed-10-bit CPU cursor blend, a process-wide SDR
downgrade latch on negotiation timeout, and a DisplayConfig BT.2100
colour-mode probe (gnome_hdr_monitor_active).
* pf-encode: libav NVENC X2RGB10->P010 swscale (BT.2020 limited) ->
HEVC Main10 / 10-bit AV1 with PQ VUI; VAAPI 10-bit on both paths (CPU
P010 upload + dmabuf XR30 scale_vaapi p010/bt2020); can_encode_10bit
now probes for real on Linux; 10-bit sessions route around the
8-bit-only Vulkan-video/direct-NVENC backends.
* GameStream: host_hdr_capable() Linux arm, live monitor-HDR check at
RTSP honor time, capturer-pool reuse keyed on HDR-ness, gs_bit_depth
covers the new formats. New `punktfunk-host hdr-probe` diagnostic and
a PUNKTFUNK_SPIKE_HDR spike lever.
* Native plane stays honestly 8-bit via capturer_supports_hdr(): Mutter
RecordVirtual streams are SDR-only upstream (GNOME 50 and 51-dev), so
virtual-display sources cannot deliver HDR yet.
Validated on the RTX 5070 Ti (GNOME 50.3 / PipeWire 1.6.8): the Main10
probes pass and the ignored nvenc_hdr10_smoke GPU test emits an IDR that
ffprobe reads as Main 10 / yuv420p10le / bt2020nc / smpte2084 / limited.
Live HDR capture negotiation still needs an HDR monitor on glass; VAAPI
10-bit needs the AMD box.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The vendored rate controller packs its wavelet block index into 16 bits
(RDOperation.block_offset_saving), so a mode whose 32x32-block count exceeds
u16::MAX wraps inside the controller and corrupts the bitstream — ~8K-class
4:4:4 territory. Compute the exact count (`block_count_32x32`, the counting walk
of upstream init_block_meta, pinned against the validated Apple WaveletLayout)
and expose `pyrowave_mode_fits_rdo`; the negotiator downgrades such a session to
4:2:0 before the Welcome (the honest-downgrade channel), and both encoders
refuse outright if one slips through rather than emit a wrapped stream.
Vendor patches: 0002-rdo-saving-clamp (analyze_rate_control.comp clamps the
saving accumulation to the target, same overrun class as 0001; slangmosh.hpp
regenerated), 0003-devel-encode-16bit-read (devel tool y4m 16-bit plane reads;
tool-only, kept so the vendored source stays honest).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PyroWave's wavelet encode runs on the GPU's compute/shader cores, so a GPU-bound
game starves it: submit spikes from ~2 ms to ~15 ms under a 95%+ game load and
the stream fps collapses. NVENC is immune (separate encoder ASIC). Two levers to
let the encode get scheduled ahead of the game's rendering:
- Windows process GPU scheduling: D3DKMTSetProcessSchedulingPriorityClass, env
PUNKTFUNK_GPU_PRIORITY = off|above-normal|high (default)|realtime. Best-effort,
once per process, non-fatal on refusal (enc/windows/pyrowave.rs).
- Global-priority Vulkan encode queue (Granite patch 0005): request a
VK_KHR_global_priority queue (PYROWAVE_QUEUE_PRIORITY = off|high|realtime,
default realtime), downgrading REALTIME→HIGH→none on NOT_PERMITTED so a refused
class never regresses the encoder to HEVC.
HONEST STATUS: on an RTX 4090 / Windows / WDDM neither moved the ~15 ms spikes —
the graphics-vs-compute preemption granularity is the wall, not the priority
level. Kept because both are correct, harmless (graceful fallback), and may help
other GPUs/drivers. For a GPU-saturated game the working levers are reducing the
encode's GPU cost (4:2:0/8-bit) or H.265; PyroWave holds full rate on the desktop
and in games that leave the GPU headroom.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Real-world PyroWave streaming maxed ~2.5 Gbps with sagging fps while raw
transport does 4.8. Root-caused to serial per-frame paths at BOTH ends
(the transport was never the limit); this fixes the two dominant ones.
Host (vendored shim, patch 0004): pyrowave_encoder_encode_gpu_synchronous
allocated four Vulkan buffers (meta + bitstream, Device + CachedHost) on
EVERY frame. At 240 fps with MB-scale bitstreams that per-frame allocator
churn stalled the encode itself. Pool them on the encoder and reuse across
frames (recreate only on a size grow); the sizes are session-fixed, so it
is pure reuse after frame 1. On an RTX 4090 the 5120x1440 submit+fence-wait
drops ~15 ms -> ~1 ms, i.e. the host serial ceiling goes 64 -> 1025 fps
(444+HDR 44 -> 614). Safe under the synchronous encode model; re-validated
by pyrowave_win_smoke (Windows) and pyrowave_smoke/_444 (Linux). Applies to
both host encoder paths (they share the shim).
Client (Apple Metal decoder): WaveletBitstream.parse reserved the payload
buffer per packet (reserveCapacity(count + words), an exact realloc each of
~3000 packets/frame => O(n²)) and copied word-by-word. Reserve once up
front and memcpy each packet's coefficients in one shot (all Apple
platforms are little-endian, so the wire's LE u32s land verbatim; memcpy is
alignment-free). 5.44 ms -> 0.055 ms per 1.44 MB frame (25x); byte-identical
(parser unit tests + golden-frame PSNR unchanged).
Also:
- native.rs: PUNKTFUNK_PYROWAVE_MAX_MBPS caps PyroWave's open-loop Automatic
bitrate pin for hosts on a constrained link (unset => no cap; an explicit
client rate bypasses it). The pin is all-intra + ABR-off, so at a high
pixel rate it can outrun the fabric (4:4:4+HDR 5120x1440@240 pins ~5.3
Gbps, over a 5 GbE link) and the overshoot just becomes loss.
- pf-encode caps(): report the real opened chroma instead of a hardcoded
4:2:0 default, so a genuine 4:4:4 session no longer trips the spurious
"encoder chroma disagrees with the negotiated Welcome" warn. Also fix a
latent Windows reset() that rebuilt at 4:2:0 for a 4:4:4 session.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phase 3 of design/pyrowave-444-hdr.md. A PyroWave session now negotiates HDR
(10-bit) and 4:4:4 on a Windows host exactly like HEVC/AV1, and the Linux
client presents it through the real HDR10 path.
Host (Windows): BgraToYuvPlanes becomes mode-aware — SDR/BGRA and HDR/scRGB
variants at half- or full-res chroma. The HDR passes reuse HdrP010Converter's
exact colour math (scRGB -> PQ BT.2020 limited studio codes, verified by
hdr_p010_selftest) but write P010-style MSB-packed codes into two separate
shareable R16_UNORM/R16G16_UNORM textures; chroma keeps the pyrowave family's
centre-sited 2x2 box. idd_push pins the composition to the NEGOTIATED depth
(SDR sessions force advanced color off as before; 10-bit sessions enable it
and ride the FP16 ring), and the descriptor poller re-asserts that state
instead of following display flips the fixed-format encoder can't. The
encoder imports 8/16-bit planes per session and stamps the sequence header's
BT.2020/PQ/matrix bits on HDR (stamp_color_bits, extending 574e3e4e's range
stamp); supports_10bit/can_encode_10bit/can_encode_444 gates open (HDR
Windows-only — Linux capture has no HDR source).
Client: the plane ring becomes R16_UNORM for 10-bit sessions (with a
STORAGE_IMAGE format probe), the planar CSC pass joins the HDR10 swapchain
rebuild (set_hdr_mode previously destroyed it without rebuilding — latent),
st.hdr follows frame.color.is_pq(), and the planar push constants carry
depth-10 MSB-packed rows + the PQ tonemap mode, identical to the NV12 arm.
Verified: .173 (RTX 4090) deploy-config clippy + fmt + wire tests + the
extended pyrowave_win_smoke (10-case {SDR,HDR}x{420,444} matrix incl. R16
imports and header stamps); .21 (RTX 5070 Ti) clippy across 4 crates, host
186 tests, client/presenter/encode tests, both Linux GPU smokes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Phase 2 of design/pyrowave-444-hdr.md. can_encode_444(PyroWave) now returns
true on Linux, so a client that advertises VIDEO_CAP_444 (the 4:4:4 setting)
and prefers PyroWave negotiates a full-chroma wavelet session end to end
(the client decoder side landed in 5eb930e7).
- rgb2yuv444.comp: the 4:4:4 twin of rgb2yuv.comp — one invocation per pixel,
full-res interleaved RG8 CbCr, no box filter/siting, byte-identical BT.709
limited coefficients; compiled .spv committed (glslangValidator -V, matches
the existing shader's toolchain).
- Encoder: chroma-conditional pyrowave create (open + reset), full-res chroma
plane + views, per-pixel dispatch, 4:2:0-only even-dims check.
- Tests: decode oracle grows a 4:4:4 mode (YUV444P CPU readback);
pyrowave_smoke_444 round-trips plane means AND drives the busy test card at
the ~2.6 bpp operating point asserting in-budget + run-to-run deterministic
AU sizes — the exact regime that silently corrupted before the vendored
payload_data fix (patches/0001), so this doubles as its regression test.
Verified on .21 (RTX 5070 Ti): clippy -D warnings, host tests, and both GPU
smokes (pyrowave_smoke + pyrowave_smoke_444) green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Phase 1 of design/pyrowave-444-hdr.md. No behavior change yet: the handshake's
4:4:4 gate now admits PyroWave (probe = can_encode_444(codec), capture gate
inherently satisfied — the wavelet path always ingests an RGB source and does
its own CSC), but can_encode_444 stays false for PyroWave until the per-OS
full-res-chroma CSC variants land (Phase 2 Linux, Phase 3 Windows), so every
session still resolves 4:2:0/8-bit.
- Both host encoders take the negotiated ChromaFormat (bail on 444 for now);
the PUNKTFUNK_ENCODER=pyrowave lab override pins 4:2:0.
- Bitrate: the automatic ~1.6 bpp pin resolves AFTER depth+chroma and scales
x1.625 for 4:4:4 / x1.15 for 10-bit (factors from the Phase-0 fixture
matrix); the mid-stream mode-switch re-resolve threads the session's values.
- Client: PyroWaveDecoder builds its plane ring (full-res chroma when 444) and
creates the upstream decoder from the negotiated chroma, keeps chroma fixed
across mid-stream resizes, drops the even-dims requirement for 444, and
returns the negotiated Welcome ColorInfo as the frame colour contract
instead of hardcoded BT.709 (the wavelet bitstream has no VUI).
Verified on .21 (RTX 5070 Ti): clippy -D warnings (host+client+encode), host
186 tests, client + pf-encode tests, fmt, and the pyrowave_smoke GPU
round-trip through the patched vendored lib (97cf15e3).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
pyrowave's encoder fills the BitstreamSequenceHeader with `= {}` and its C API
offers no way to set colour/range, so it signals ycbcr_range=0=FULL — but both
host CSCs (rgb2yuv.comp on Linux, BgraToYuvPlanes on Windows) always emit BT.709
LIMITED Y'CbCr (black = Y'16). A client that honours the VUI (the Apple wavelet
decoder reads bit 30 of word1) then skips the limited→full expansion and shows
washed-out, raised blacks — reported on both Linux and Windows hosts.
Patch the range bit HONEST (mark_limited_range in the shared pyrowave_wire, called
by both encoders after packetize). Clients that hardcode limited (the Vulkan
video_pyrowave path) are unaffected, and pyrowave's own decode ignores the flag
(raw Y'CbCr reconstruction). No client rebuild needed. Unit-tested + asserted in
pyrowave_win_smoke; the smoke decode still round-trips 100/180/60.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Wire PyroWave into the Windows host (design/pyrowave-windows-host-zerocopy.md).
Before this a macOS client + Windows host that both selected PyroWave silently ran
HEVC: the host never advertised CODEC_PYROWAVE and open_video_backend bailed.
Approach (zero-copy, no GPU→CPU→GPU): pyrowave owns its own Vulkan device
(create_device_by_compat, by render-GPU vendor/device-id — NOT LUID, invalid in
Session 0). The capturer runs a BGRA→YUV BT.709-limited CSC (matching rgb2yuv.comp)
into TWO SEPARATE shareable plane textures — full-res R8 Y + half-res R8G8 CbCr —
which the encoder imports into pyrowave's device. Separate single/two-component
textures import reliably on NVIDIA at any size; a single planar NV12 import does NOT
(the vendored interop test: "only very specific resource sizes" — confirmed on-glass:
1024² fine, 720p/1080p/1440p garbage). A shared D3D11 fence, signalled after the CSC,
is imported as a Vulkan timeline semaphore so the wavelet read is ordered after it.
- pf-encode: enc/windows/pyrowave.rs (Encoder impl, two-plane import + Linux-style
plane views); host_wire_caps advertises CODEC_PYROWAVE on Windows when the backend
isn't Software; open_video_backend routes a negotiated PyroWave session first;
pyrowave-sys on the Windows target; interop confirmed at open → clean HEVC fallback.
- pf-encode: shared, unit-tested enc/pyrowave_wire.rs (single source of truth for the
client-facing AU framing); Linux encoder uses it too.
- pf-capture: dxgi.rs BgraToYuvPlanes CSC; idd_push.rs pyrowave mode — forces the
virtual display SDR (the VideoProcessor can't ingest the FP16 HDR ring), a
two-plane shareable out-ring, a shared fence passed every frame (so a rebuilt
encoder re-imports it). Threaded via OutputFormat::pyrowave.
- pf-frame: D3d11Frame::pyro carries the CbCr plane + fence; OutputFormat::pyrowave.
Verified on .173 (RTX 4090): full-host build + clippy -D warnings (nvenc,amf-qsv) +
fmt --all --check; pyrowave_wire unit tests; pyrowave_win_smoke GPU test round-trips
distinct Y/Cb/Cr (100/180/60) exactly at 1024²/720p/1080p/1440p; Stage-0 interop
validated in the real Session-0 service context on-glass. Deployed to the box.
Owed: final on-glass picture/latency confirmation.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
design/latency-reduction-2026-07.md tier 2, the two code-side halves:
- T2.2: the Linux direct-NVENC backend gains the two-thread retrieve
(PUNKTFUNK_NVENC_ASYNC, the same opt-in knob as Windows): the session stays
sync-mode (async events are Windows-only) but the blocking lock_bitstream
moves to a dedicated pf-nvenc-out thread — the NVENC guide's sanctioned
submit-thread/output-thread split. poll() drains completions non-blocking,
submit() backpressures at PUNKTFUNK_NVENC_ASYNC_DEPTH (default 4) in-flight;
map/unmap and every other session call stay on the encode thread; teardown
joins the thread before destroying the session. Under a GPU-saturating game
completed frames queue instead of serializing capture on the encode wait.
- T2.3: PUNKTFUNK_GPU_PRIORITY_CLASS gains 'auto' AND IT IS THE NEW DEFAULT
(gpu-contention §5.C): HIGH immediately, then REALTIME where the documented
NVIDIA+HAGS+near-full-VRAM NVENC hang cannot bite — HAGS probed once via
D3DKMT WDDM_2_7_CAPS (off => REALTIME outright); HAGS on => a pf-gpu-prio
monitor flips REALTIME<->HIGH on LOCAL-segment VRAM headroom (downgrade
>92% of budget, restore <=85% for 3x2s polls). 'high' restores the old
static default; 'realtime' pins it (operator owns the hazard).
Validated: .21 clippy -D warnings (punktfunk-host --features nvenc) against
the QSV-merged main; .133 Windows cargo check of pf-frame + punktfunk-host.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The eight W6 leaf crates hardcoded 0.12.0 instead of inheriting the
workspace version — switched to version.workspace = true so the next bump
is one line again.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
encode.rs + encode/* (NVENC, VAAPI, native AMF, AMF/QSV ffmpeg, direct-SDK
NVENC/CUDA, raw Vulkan-Video, PyroWave, openh264) move into crates/pf-encode
behind one Encoder trait + open_video selector (plan §W6). The crate speaks the
shared frame vocabulary (pf-frame: CapturedFrame/PixelFormat + the DXGI identity
D3d11Frame/make_device) and pf-zerocopy (CUDA context/buffers), and NEVER
pf-capture — the capture→encode edge is one-way (ZeroCopyPolicy, prior commit).
Dep moves: the heavy encoder deps (ffmpeg-next, the NVENC SDK, openh264,
pyrowave-sys) move from the host to pf-encode; the host's
nvenc/amf-qsv/vulkan-encode/pyrowave features now FORWARD to pf-encode/*. The
host keeps a mod-encode shim (pub use pf_encode) so every crate::encode::* path
(negotiator + GameStream/native/mgmt planes) is unchanged.
resolve_render_adapter_luid moves from the host's windows/win_adapter.rs into
pf-gpu (both pf-encode and pf-capture need it as a peer of GPU selection); its 5
call sites (encode amf/nvenc, capture idd_push/synthetic_nv12, vdisplay manager)
rewire to pf_gpu::resolve_render_adapter_luid and win_adapter.rs is deleted.
pf-frame's make_device gains a # Safety section (public-unsafe-fn lint, latent
since the pf-frame carve — a full-workspace -D warnings clippy catches it).
Verified: Linux clippy -D warnings (pf-encode + host nvenc,vulkan-encode,pyrowave
--all-targets) + 13/13 pf-encode + 299/299 host tests; Windows clippy -D warnings
(pf-encode nvenc,amf-qsv --all-targets + host nvenc,amf-qsv --all-targets)
Finished exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>