The two direct-NVENC backends carried hand-copied twins of the same
loss-recovery decision: range validity, covering-range dedup, DPB
window, clamp — ~30 duplicated lines each. The decision now lives once
as nvenc_core::plan_range_recovery (the range half of WP7.2; the slot
half is enc/rfi.rs), pure and unit-tested; each backend keeps its
session gate, its unsafe per-timestamp driver loop, and its state
stores.
The step order is load-bearing and now pinned by tests: the covering
dedup runs with the UNCLAMPED last and BEFORE the DPB window (a covered
re-ask never touches the driver even when the range has since aged out
of the DPB), the boundary at next_ts - RFI_DPB is inclusive, and the
Invalidate carries the CLAMPED last — which is also what the caller
records in last_rfi_range, exactly as the inline code stored it. A
driver failure mid-loop still returns false with NO range recorded and
no anchor armed. Decline deliberately clears nothing (neither twin
touched pending_anchor on decline — same shape as Vulkan's non-clear,
opposite of AMF/QSV; do not harmonize).
The exact-cover → Covered test records EXISTING behavior including that
a covered range survives a forced IDR with zero driver calls — a
recorded fact, not an endorsement. RFI_DPB's import leaves both twins:
its only per-backend use was the arithmetic that moved.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`std::env::var` was on three per-frame paths. Measured on `.173` (Windows, 57
environment variables, 2M iterations): **121.9 ns** per call for the NVENC
in-flight cap and **114.9 ns** per call for the ffmpeg poll spin, against
**0.9 ns** once memoized. On Linux, 32 ns → 1.5 ns.
⚠ Those numbers deflate the filing, and that is worth recording: the audit
ranked this as a hot-path defect, but ~120 ns/frame is ~0.003% of a frame
budget. The fix is still right — it is free, and it takes a global environment
lock off the encode thread — but nobody should schedule it ahead of anything on
the strength of the "hot path" framing.
The severity RANKING was also inverted, and the measurement confirms why. The
site the audit called worst — `nvenc_cuda`'s backpressure loop condition — costs
a default session nothing, because the condition short-circuits on
`async_rt.is_some()` and the default session never engages the two-thread
retrieve. The one that actually pays every frame is Windows `submit`, which
consults `async_inflight_cap()` in BOTH arms of the ring-depth match, sync mode
included, where the result is then thrown away. Windows `poll` is the second
unconditional one, and the audit ranked it last.
Memoized inside each helper rather than latched into a session field. Nothing in
the workspace mutates these variables at runtime (enumerated: no `set_var` for
either key anywhere in first-party code, and the Windows service's arbitrary-key
`host.env` loader runs in the SCM supervisor, which re-execs the host as a child
and never opens an encoder itself). A field would instead change WHEN the value
is read — and the Windows `cap` composes the env with `input_ring_depth`, which
`set_input_ring_depth` may change after open, so freezing that half would
reintroduce the in-place-overwrite bug the ring term exists to prevent.
Two deliberate behaviour changes ride along on `PUNKTFUNK_FFWIN_POLL_MS`, so
"behaviour-preserving" describes the memoization only, not this whole commit:
- **A 1000 ms ceiling.** The reachable hazard was never the overflow — that
needed `ms >= 1.8e16` — it was a slipped digit: `=100000000` was a 27.7-hour
spin of the encode thread.
- **`.trim()`, now on all three parsers.** An earlier draft applied the house
rule (WP7.8) to one of the three, which left `PUNKTFUNK_NVENC_ASYNC=" 1 "`
working while `PUNKTFUNK_NVENC_ASYNC_DEPTH=" 6 "` silently fell back to 4 —
and memoization would have frozen that silent fallback for the process
lifetime. Reachable: the Windows `host.env` loader trims around `=` before
stripping quotes, so `=" 2 "` yields a value with inner spaces.
⚠ Correction to an earlier draft of this message, which asserted that the audit's
`saturating_mul` proposal would relocate an overflow panic into release builds.
That is FALSE and the code comment now says so: `Duration::from_micros(u64::MAX)`
is ~1.8e13 seconds, six orders of magnitude below `Duration`'s ceiling, so
`Instant + Duration` neither overflows nor panics (measured, with and without
debug assertions). `saturating_mul` is still wrong, for a different reason — it
sets a deadline ~584,000 years out on a spin that provably never produces the
owed AU, i.e. it wedges the encode thread permanently. A hang, not a panic.
WP6.1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`NV_ENC_ERR_INVALID_VERSION` is how the driver reports two opposite failures, and we only
ever explained one of them. The message told the operator to update the NVIDIA driver or
reboot — correct for a genuine header/kernel-module skew, and actively misleading for the
field case behind it: a host that streams once per boot and then fails every later session
at the caps probe, until the *process* restarts. A version skew is static; it cannot come
and go inside one process, so that advice cost the reporter a reboot per stream.
Split the two on the only fact that tells them apart: whether a session has already opened
in this process. `nvenc_status` gains a `SESSION_OPENED` latch set right after every
successful `open_encode_session_ex` — both backends' caps probe and real open, plus the
Windows availability probe. No session yet, the version word really is in question and the
existing skew advice stands. A session already opened, and the kernel module demonstrably
accepted this build's version word, so the message now names per-process driver state and
points at the cheap fix (restart the host service, no reboot) plus a request for the log.
The load-time gate cannot serve as this discriminator: `NvEncodeAPIGetMaxSupportedVersion`
is a pure userspace query, so the classic "updated the driver, didn't reboot" skew sails
through it and only fails later at the open. Only a session that actually opened proves the
kernel module agreed.
The split lives in a pure `invalid_version(bool)` so both halves are unit-tested without
touching the process-wide latch. This is a diagnosis change only — it does not fix the
underlying field bug, whose root cause is still open.
Verified on Linux (192.168.1.25): clippy `-p pf-encode --features nvenc` clean, `cargo test
-p pf-encode --features nvenc --lib` 15 passed, rustfmt clean.
`33121ece` added a block explicitly labelled "TEMP KWin composite probe (DROP BEFORE
MERGE)" to prove the cursor blend was dispatching. It merged, and has been running on
every blended frame since: an atomic fetch_add per frame plus a `tracing::info!` with
seven computed fields every 512 frames, on the submit hot path of the default-on
Linux direct-NVENC backend.
The signal it existed for is already covered — the failure arm above it warns once
with the dispatch error, and the `else` arm warns when an overlay arrives with no
blend at all. What is deleted is only the success-path telemetry.
Verified `-D warnings --all-targets` + tests on Linux (default; shipped
nvenc,vulkan-encode,pyrowave), and `clippy --features vulkan-encode,pyrowave` on real
AMD RDNA3 hardware.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`open_video`'s `cursor_blend` argument was a request with no answer: lib.rs did
`let _ = cursor_blend;` and only three backends ever read `CapturedFrame::cursor`.
So a session could ask for a composited pointer, get a backend that silently
discards it, and stream with no mouse cursor and nothing in the logs. Two
separately-confirmed audit findings — the VAAPI dmabuf path and the libav-NVENC CUDA
path — are symptoms of that one hole.
`EncoderCaps::blends_cursor` makes it a fact each backend states. The four exhaustive
`EncoderCaps { .. }` constructors mean adding the field is a compile error until every
backend answers, which is the enforcement mechanism for future backends rather than a
side effect. Vulkan Video answers from its ACTUAL configured source rather than
statically: only the CSC path composites (`prep_cursor` feeds the compute shader),
while the RGB-direct/EFC front-end and the native-NV12 source have no compositing
stage at all and merely warn once that the pointer is being dropped.
`open_video` warns when a session asked for blending and the opened backend cannot
deliver it. A warning is deliberately all it does: `open_video` cannot re-plan
capture, so refusing would trade a missing pointer for a dead session. The host owns
`plan.cursor_blend` and is the only layer that can fall back to capturer-side
compositing — this gives it something to base that on.
Enforcement is NOT included. The reviewed design proposed refusing the client's
host-composite flip to keep the client drawing its own pointer, but `CursorRenderMode`
is client->host only: there is no host->client counterpart, so refusing yields no
pointer at all — the same failure it claimed to prevent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Escalate-and-hold's missing half. The contention escalation (capture depth,
then the NVENC pipelined retrieve) was permanent: one sustained overrun — even
one CAUSED by the ABR overdrive's rebuild storms — cost the session its
depth-1 latency and its sub-frame streaming forever, and `encode_us` reported
queue depth instead of ASIC time for the rest of the session.
- The leaky bucket now keeps scoring after escalation. A sustained
every-frame-on-cadence run (~5 s at 120 fps) winds back one stage in reverse
order: pipelined retrieve first (its rebuild restores the IO-stream binding
and sub-frame chunked streaming), then capture depth back to 1. Attempts are
paced by an exponential backoff (1 → 5 → 25 min, capped) — a workload that
truly needs the escalation converges to keeping it, but never a permanent
latch.
- NVENC (Linux) implements `set_pipelined(false)`: a `want_sync` latch handled
at the same drained safe point as the engage side (`maybe_disengage_async`
mirrors `maybe_engage_async`); the lazy sync re-init re-arms everything and
opens on an IDR. The stream loop polls until the switch lands, then re-runs
the escalation warmup so the wind-back's own stall can't re-escalate it.
`PUNKTFUNK_NVENC_ASYNC=1` (operator-pinned async) refuses the wind-back;
the trait doc now specifies the two-way contract.
- While escalated, `cadence_degraded` stays latched (bitrate climbs refused)
even with the bucket drained: the headroom is spent, and climbs resuming
mid-escalation would saw against it and starve the clean run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three encode-time fixes from the 4K120 field analysis (sessions pinned ~107 of
120 fps, 14-15 ms reported encode):
- Ceiling truth: the codec-level bitrate clamp (binary search at session open)
now records its result in a process-lifetime advisory cache keyed by
GPU/config, so an ABR overshoot opens — or in-place reconfigures — straight
AT the ceiling instead of re-running the ~6-open search (and a full rebuild
+ IDR) on every overshoot. New `Encoder::applied_bitrate_bps()` exposes the
post-clamp rate so the session loop can stop pacing/acking a phantom
requested rate (consumed host-side in a follow-up).
- Clamp-search hygiene: only NVENC parameter/caps rejections steer the search
now (`NvCallError` keeps the raw status downcastable); a transient failure
(busy engine, session limit, OOM, driver skew) propagates instead of
shrinking into — and now caching — a bogus ceiling. The floor fallback also
records the split mode it actually opened with, so a later reconfigure
re-presents the real session params.
- Split-frame selection: one shared `resolve_split_mode` replaces the two
byte-identical direct-SDK copies; the force-2-way threshold moves to
`SPLIT_FORCE_PIXEL_RATE` (950 Mpix/s, shared with the libav path) because
4K120 = 995,328,000 px/s missed the old `> 1e9` gate by 0.47% and stayed on
AUTO — which never engages at 2160 px height, leaving the second NVENC
engine idle in exactly the mode the threshold existed for. The Linux
session-ready info! line now carries the final split mode (journals are
INFO+; this was undiagnosable from user logs).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
All ten `NvencCudaEncoder::open` call sites in the `#[cfg(test)]` module passed 9
arguments to a 10-parameter function: `cursor_blend` was added to `open` and the
tests were never updated. The module has been uncompilable ever since —
`cargo clippy -p pf-encode --features nvenc --all-targets` fails with 10x E0061.
Nothing in CI noticed because no job compiles this crate's feature-gated test
targets: ci.yml lints/tests with default features (which do not include `nvenc`),
and windows-host.yml lints `-p punktfunk-host`, which builds pf-encode as a
dependency and so never builds its test targets. The CI commit below closes that
gap; this has to land first or that leg starts red.
`false` is what these tests exercised before `cursor_blend` existed: every frame
they build sets `cursor: None`, and there is no cursor-blend test.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Since the cursor-channel work, every Linux virtual output was created in
metadata pointer mode — making ALL sessions depend on host-side cursor
compositing, including the ones that can never use the channel (Moonlight/
GameStream, legacy clients, capture-mode starts). Those sessions paid the
blend bring-up per session and, whenever a visible cursor was composited,
the loss of NVENC's stream-ordered submit — for a strictly worse cursor
than the compositor's own.
Mirror the Windows no-regression gate: the already-wired per-session
set_hw_cursor(cursor_forward) now drives the Linux backends too. A
cursor-channel session gets metadata (shapes forwarded, composite flip
blends host-side — today's validated path, unchanged); every other session
gets the pointer compositor-EMBEDDED at creation (KWin zkde pointer=2,
Mutter cursor-mode=1, wlroots/hyprland portal CursorMode::Embedded — their
pre-channel default; the portal pair also gains the Metadata arm for
channel sessions, wired but untested on-glass). SessionPlan.cursor_blend
narrows to cursor-forward sessions and rides into the direct-SDK NVENC
open, which skips the Vulkan slot-blend bring-up entirely when off —
embedded sessions ring on plain CUDA surfaces and pay zero cursor cost,
per-session or per-frame.
The keep-alive registry's reuse key grows the created pointer mode (new
VirtualDisplay::hw_cursor getter): a kept embedded display has no cursor
metadata for a channel session to forward, and a kept metadata display
would leave a channel-less session with no pointer in its frames.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A vendored PTX blob is JIT'd against the driver's ISA ceiling, so the
cursor-blend module silently dies on drivers older than the generating
toolkit (CUDA_ERROR_UNSUPPORTED_PTX_VERSION/INVALID_PTX, 222/218 — the
KWin leg's invisible composite cursor on driver 595/CUDA 13.2 vs a
CUDA 13.3 blob). SPIR-V has no such coupling, and the in-tree precedent
already exists twice (vulkan_video's CSC blend, VkBridge's exportable
OPAQUE_FD → cuImportExternalMemory bridge).
New pf_zerocopy::vkslot::VkSlotBlend: the direct-SDK NVENC encoder now
allocates its input ring as exportable Vulkan buffers CUDA-imports (same
contiguous InputSurface layouts, pitch = row bytes rounded to 256), and
the cursor composite is a compute dispatch over the cursor's rectangle
(cursor_blend.comp, vendored .spv; spec-constant selects ARGB/NV12/YUV444;
BT.709 limited, matching the retired .cu). The surface SSBO is uint[] with
every invocation owning whole words — no 8-bit-storage device dependency.
Cursor-bearing frames force the existing CPU-synced submit path so the
CUDA copy → Vulkan dispatch (fence-waited) → NVENC encode ordering is
CPU-established; cursorless frames keep the stream-ordered fast path
untouched. Any bring-up/alloc/registration failure falls back wholesale
to plain pitched CUDA surfaces (never a mixed or short ring): sessions
always encode, composite mode just loses the cursor, warned once.
cursor_blend.cu / cursor_blend.ptx and the CursorBlend PTX loader are
deleted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Rate-limited journal logging bracketing the silent segment of the Linux
composite path: what overlay (if any) the encode loop's composite arm
hands the encoder, and whether the NVENC cursor-blend kernel launches and
with what geometry. On-glass (KWin leg) the blend module loads yet the
composite cursor stays invisible — these two probes attribute it to
no-overlay / stripped-en-route / kernel-draws-nothing. The module-loaded
INFO line is permanent (success must be as attributable as failure);
everything else in this commit is temporary.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The vendored cursor_blend.ptx was regenerated with the CUDA 13.3 toolkit,
which stamps '.version 9.3' — and a driver whose JIT predates that ISA
refuses the whole module with CUDA_ERROR_UNSUPPORTED_PTX_VERSION (222),
silently killing host-side cursor compositing (on-glass: composite-mode
cursor invisible on driver 595.58, KWin leg). The kernels use only baseline
arithmetic, so hand-lower the version directive to 8.0 (CUDA 12.0), which
every supported Turing+ driver JITs. Both files now warn that an nvcc
regeneration re-stamps the version and must be re-lowered.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Gates recorded (nvenc-subframe-slice-output.md Phase 3): loss-harness curve
identical streamed-vs-whole; adversarial security review clean; live .21
3-leg A/B — streamed 0 corrupted frames, p99 8527→5363 µs vs sliced-whole,
first_slice_us ~500 µs of a ~1250 µs encode reaching the send thread; wire
bytes rate-governed under CBR + 1-frame VBV (the ~1-2 % slice-header cost is
absorbed by rate control, not added to the wire). GameStream: plane has zero
code diff (own RTP packetizer; shares only untouched send_pacing); a live
Moonlight re-test joins the standing owed Moonlight item.
- resolve_slices(codec, default): env 1..=32 wins (1 = the NEW explicit
single-slice escape), else the backend default — 4 on Linux direct-NVENC,
1 elsewhere. resolve_subframe(default_on): PUNKTFUNK_NVENC_SUBFRAME
tri-state (0 = never, 1 = force, unset = backend default).
- Linux nvenc_cuda: resolves both ONCE in query_caps — the sub-frame default
is gated on the GPU's SUBFRAME_READBACK cap so an unsupporting GPU never
has enableSubFrameWrite forced into its init params — and feeds the same
resolved values to build_config, build_init_params (open AND in-place
reconfigure) and the chunked-poll latch.
- Windows: env-only as before (async path untouched, byte-identical config
for unset env); LowLatencyConfig.slices + the explicit subframe param
replace the shared env reads.
- Tests: chunked e2e now runs at the DEFAULTS (no knobs); the fallback test
becomes the escape test (SLICES=1 disarms; SUBFRAME=0 disarms while
4-slice encode continues on the plain poll path).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The encoder half of sub-frame slice output (latency §7 LN1, planning
design/nvenc-subframe-slice-output.md Phase 1): with PUNKTFUNK_NVENC_SLICES=N +
PUNKTFUNK_NVENC_SUBFRAME=1 on a sync depth-1 session, Encoder::poll_chunk hands
the in-flight AU out as slice-boundary chunks read through doNotWait sub-frame
locks while the tail is still encoding — the readback loop the on-hw probe
validated (~200 µs slice spacing on the 5070 Ti), productionized.
- codec.rs: AuChunk (AU metadata on the first chunk, `last` closes the AU;
chunks concatenate to exactly the poll() bytes) + supports_chunked_poll /
poll_chunk trait surface. Default impl wraps poll() as one self-closing
chunk, so a chunk-driven consumer works against every backend.
- nvenc_cuda: chunked readback cut at slice boundaries only (bitstream size at
n reported slices = end of slice n, Annex-B contiguous); completion is NEVER
numSlices alone — one finishing BLOCKING lock is the authority and the wedge
watchdog, so the final chunk blocks exactly like sync poll (the depth-1 pump
contract; 6dc195f9 bug class). Keyframe on early chunks is the submit-time
IDR prediction (exact under P-only/infinite-GOP), cross-checked at finish.
Debug builds shadow-check emitted chunks against the finished AU prefix.
Mutually exclusive with pipelined retrieve (gated off when async_rt exists,
dropped by the escalation rebuild); composes with stream-ordered submit.
- nvenc_core: slices_env/subframe_requested shared parses so the config author
and the chunked-poll arming can't disagree.
- TrackedEncoder forwards both new methods (the set_wire_chunking trap class).
Host loop untouched — Phase 2 (VIDEO_CAP_STREAMED_AU seal/send) consumes this.
On-hw: nvenc_cuda_chunked_poll_end_to_end + nvenc_cuda_chunked_poll_fallback_whole_au.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PUNKTFUNK_NVENC_ASYNC gains a real tri-state: 1 = always (as before, now
documented as ~+1 tick at depth-1), 0 = never (vetoes escalation), unset =
ADAPTIVE — off until the session loop's cadence-overrun detector escalates.
The host loop's adaptive-depth leaky bucket grows a second stage: once the
capturer's depth is maxed (Linux portal is permanently depth-1), it asks
the encoder for pipelined retrieve via the new Encoder::set_pipelined hook
(asked exactly once; default impl declines, Windows untouched).
nvenc_cuda engages at a safe point via a clean session rebuild WITHOUT the
IO-stream binding: with input==output stream bound, later stream work
waits on prior encode completions and would serialize a pipelined session
— stream-ordered submit and two-thread retrieve are mutually exclusive.
The ordered gate now also requires async_rt absence (belt-and-braces for
the runtime switch). Re-open's first frame is the standard session IDR.
On-hardware test: escalate mid-session → retrieve thread live, binding
gone, all AUs deliver, first post-escalation AU is the IDR.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
PUNKTFUNK_NVENC_SLICES=N (2..=32, default off) splits H.264/HEVC frames
into N slices (sliceMode 3); PUNKTFUNK_NVENC_SUBFRAME=1 (default off)
arms enableSubFrameWrite + reportSliceOffsets on sync sessions only.
Both experimental groundwork for sub-frame slice output (plan §7 LN1).
nvenc_cuda_subframe_slice_probe (on-hardware, ignored) answers the LN1
go/no-go: spins lock_bitstream(doNotWait) against an in-flight frame and
prints the (t_us, status, numSlices, bytes) timeline — incremental slice
availability and its spacing, or all-at-completion.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bind the session's IO streams to the encode thread's high-priority copy
stream: in sync-retrieve depth-1 use the per-frame input copy and cursor
blend now enqueue with NO cuStreamSynchronize and encode_picture orders
after them on the stream. Same stream both directions, so the encode's
completion is inserted into the stream and later work (the next frame's
copy into a reused ring slot) waits for it.
Soundness gate: the fast path engages only when pending is empty (true
depth-1 usage) — every prior encode was drained by a blocking poll, and
the caller holds the frame payload across the matching poll (contract now
documented on Encoder::submit; both host loops already comply). Pipelined
callers and PUNKTFUNK_NVENC_ASYNC mode keep the blocking copies.
True zero-copy input registration (registering the worker-owned IPC
buffer directly) stays the LN2 v2 follow-up — it needs a contiguous
worker-pool NV12 layout and a registration<->IPC-mapping lifetime tie.
PUNKTFUNK_NVENC_STREAM_ORDERED=0 restores the old blocking behavior.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Latency plan §7 LN0 (Linux/NVIDIA encode follow-on):
- sampled PUNKTFUNK_PERF submit split (copy/blend/map/pic) in nvenc_cuda —
the host loop's submit_us folds all four together; the D2D input copy is
the LN2 zero-copy target and now measurable on its own
- sampled blocking lock_bitstream timing on pf-nvenc-out — in two-thread
mode the host loop's wait_us wraps a non-blocking poll, so the real
encode wait was measured by no timer
- caps probe + log SUPPORT_SUBFRAME_READBACK / SUPPORT_DYNAMIC_SLICE_MODE
(LN1 sub-frame slice-output prerequisites, fleet visibility)
- explicit zeroReorderDelay=1 in the shared low-latency config (P-only +
no lookahead has no reordering anyway; pins the bit against preset or
driver drift; shared with the Windows backend)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
All four found in the pf-encode quality sweep and verified against source.
- NVENC partial-init leak (BOTH platforms, high): `init_session` publishes
`self.encoder` — and on Windows charges LIVE_SESSION_UNITS — *before* its
remaining fallible steps (bitstream buffers; on Linux also the input-surface
alloc and `register_resource`). A failure there left a live session with
`inited == false`, and every guard on the re-init path keys off `inited`, so
the next submit skipped teardown and overwrote `self.encoder`: the session
leaked permanently toward the driver's per-process cap, and its budget units
never returned, progressively starving parallel-display admission. `teardown`
already keys off `encoder.is_null()` rather than `inited`, so it cleans up
exactly this half-built state — it just was never called. Now invoked on the
`init_session` error path on both platforms.
- `can_encode_10bit` asked the wrong backend (medium): it resolved via
`linux_auto_is_vaapi`, which ignores `encoder_pref`, while `can_encode_444`
and `open_video` honour it. On a host that forces a backend (e.g.
`encoder_pref = "vaapi"` on an NVIDIA box) the probe answered for NVENC while
the session opened VAAPI, so the negotiated bit depth — and the HDR/SDR colour
label derived from it — described a backend that never ran. Now uses the same
`linux_zero_copy_is_vaapi` mirror, and `linux_auto_is_vaapi` carries a warning
that it resolves the `auto` case only and is not a dispatch mirror.
- Linux software arm ignored SW_BITRATE_CEIL (low): the Windows arm clamped
openh264 to 100 Mbps, the Linux arm passed the full negotiated rate. The
constant is now module-scope so both arms share one value.
- QSV/AMF env-parity (low): `PUNKTFUNK_IR_PERIOD_FRAMES` was a no-op on QSV
despite the comment claiming parity with AMF, and `PUNKTFUNK_NO_QSV_LTR` /
`PUNKTFUNK_INTRA_REFRESH` had dropped AMF's `trim()` and `yes`/`on` spellings,
so a value with stray whitespace silently did nothing on Intel while the same
value worked on AMD.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GNOME 50 (Mutter MR 4928, PipeWire >= 1.6) added HDR screen sharing for
monitor streams: 10-bit PQ formats (xRGB_210LE/xBGR_210LE) with MANDATORY
BT.2020 + SMPTE-2084 colorimetry props, advertised while the mirrored
monitor is in BT.2100 colour mode. Wire the Linux host into it end-to-end
on the GameStream desktop-mirror path (PUNKTFUNK_VIDEO_SOURCE=portal):
* pf-frame: PixelFormat::X2Rgb10/X2Bgr10 (DRM XR30/XB30; X2Bgr10 is the
Windows Rgb10a2 layout) + fourccs.
* pf-capture: want_hdr portal offer — HDR-only LINEAR-dmabuf pods with
MANDATORY PQ/BT.2020 props (SHM excluded: Mutter's SHM record path
paints 8-bit ARGB32 regardless of format; tiled excluded: the EGL
de-tile blit is 8-bit RGBA8), negotiated-colorimetry parse, generic
HDR10 hdr_meta(), packed-10-bit CPU cursor blend, a process-wide SDR
downgrade latch on negotiation timeout, and a DisplayConfig BT.2100
colour-mode probe (gnome_hdr_monitor_active).
* pf-encode: libav NVENC X2RGB10->P010 swscale (BT.2020 limited) ->
HEVC Main10 / 10-bit AV1 with PQ VUI; VAAPI 10-bit on both paths (CPU
P010 upload + dmabuf XR30 scale_vaapi p010/bt2020); can_encode_10bit
now probes for real on Linux; 10-bit sessions route around the
8-bit-only Vulkan-video/direct-NVENC backends.
* GameStream: host_hdr_capable() Linux arm, live monitor-HDR check at
RTSP honor time, capturer-pool reuse keyed on HDR-ness, gs_bit_depth
covers the new formats. New `punktfunk-host hdr-probe` diagnostic and
a PUNKTFUNK_SPIKE_HDR spike lever.
* Native plane stays honestly 8-bit via capturer_supports_hdr(): Mutter
RecordVirtual streams are SDR-only upstream (GNOME 50 and 51-dev), so
virtual-display sources cannot deliver HDR yet.
Validated on the RTX 5070 Ti (GNOME 50.3 / PipeWire 1.6.8): the Main10
probes pass and the ignored nvenc_hdr10_smoke GPU test emits an IDR that
ffprobe reads as Main 10 / yuv420p10le / bt2020nc / smpte2084 / limited.
Live HDR capture negotiation still needs an HDR monitor on glass; VAAPI
10-bit needs the AMD box.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
design/latency-reduction-2026-07.md tier 2, the two code-side halves:
- T2.2: the Linux direct-NVENC backend gains the two-thread retrieve
(PUNKTFUNK_NVENC_ASYNC, the same opt-in knob as Windows): the session stays
sync-mode (async events are Windows-only) but the blocking lock_bitstream
moves to a dedicated pf-nvenc-out thread — the NVENC guide's sanctioned
submit-thread/output-thread split. poll() drains completions non-blocking,
submit() backpressures at PUNKTFUNK_NVENC_ASYNC_DEPTH (default 4) in-flight;
map/unmap and every other session call stay on the encode thread; teardown
joins the thread before destroying the session. Under a GPU-saturating game
completed frames queue instead of serializing capture on the encode wait.
- T2.3: PUNKTFUNK_GPU_PRIORITY_CLASS gains 'auto' AND IT IS THE NEW DEFAULT
(gpu-contention §5.C): HIGH immediately, then REALTIME where the documented
NVIDIA+HAGS+near-full-VRAM NVENC hang cannot bite — HAGS probed once via
D3DKMT WDDM_2_7_CAPS (off => REALTIME outright); HAGS on => a pf-gpu-prio
monitor flips REALTIME<->HIGH on LOCAL-segment VRAM headroom (downgrade
>92% of budget, restore <=85% for 3x2s polls). 'high' restores the old
static default; 'realtime' pins it (operator owns the hazard).
Validated: .21 clippy -D warnings (punktfunk-host --features nvenc) against
the QSV-merged main; .133 Windows cargo check of pf-frame + punktfunk-host.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
encode.rs + encode/* (NVENC, VAAPI, native AMF, AMF/QSV ffmpeg, direct-SDK
NVENC/CUDA, raw Vulkan-Video, PyroWave, openh264) move into crates/pf-encode
behind one Encoder trait + open_video selector (plan §W6). The crate speaks the
shared frame vocabulary (pf-frame: CapturedFrame/PixelFormat + the DXGI identity
D3d11Frame/make_device) and pf-zerocopy (CUDA context/buffers), and NEVER
pf-capture — the capture→encode edge is one-way (ZeroCopyPolicy, prior commit).
Dep moves: the heavy encoder deps (ffmpeg-next, the NVENC SDK, openh264,
pyrowave-sys) move from the host to pf-encode; the host's
nvenc/amf-qsv/vulkan-encode/pyrowave features now FORWARD to pf-encode/*. The
host keeps a mod-encode shim (pub use pf_encode) so every crate::encode::* path
(negotiator + GameStream/native/mgmt planes) is unchanged.
resolve_render_adapter_luid moves from the host's windows/win_adapter.rs into
pf-gpu (both pf-encode and pf-capture need it as a peer of GPU selection); its 5
call sites (encode amf/nvenc, capture idd_push/synthetic_nv12, vdisplay manager)
rewire to pf_gpu::resolve_render_adapter_luid and win_adapter.rs is deleted.
pf-frame's make_device gains a # Safety section (public-unsafe-fn lint, latent
since the pf-frame carve — a full-workspace -D warnings clippy catches it).
Verified: Linux clippy -D warnings (pf-encode + host nvenc,vulkan-encode,pyrowave
--all-targets) + 13/13 pf-encode + 299/299 host tests; Windows clippy -D warnings
(pf-encode nvenc,amf-qsv --all-targets + host nvenc,amf-qsv --all-targets)
Finished exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>