Rides on 25765c53 (the concurrent session landed the same typed-errno
fix first — this branch's twin commit was dropped in the rebase; these
are its surviving residuals):
- Two tests pinning what the ladder's step-down rests on: the typed
ffmpeg::Error survives with_context layers as a downcastable source
(an eager format! anywhere between open_with and the ladder would
silently break it), and the phrase WITHOUT the type no longer
classifies.
- The CudaHw::new eager-format guard note: those bail!s must NOT be
converted to typed errors — a hwdevice/hwframes EINVAL is a config
error no bitrate can fix, and enrolling it would burn ~10 doomed
opens before surfacing the real failure.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Verified against nvEncodeAPI.h's own splitEncodeMode doc (user-prompted
— the audit's 'exclusive for HEVC' one-liner deserved checking):
- H.264: split 'is not applicable' — hard-DISABLE the mode so the
written config, CeilingKey, the split diagnostic log and the
rejection-retry stay truthful (the retry used to re-open a
byte-identical session after an H.264 'split rejection'). The libav
path's operator arm gains the codec gate its auto arm always had.
- HEVC: split 'not supported if … subframe mode' — when WE force split
(TWO/THREE/AUTO_FORCED, the 4K120 throughput lever), sub-frame yields
with a logged escape (PUNKTFUNK_SPLIT_ENCODE=0 chooses sub-frame).
⚠ Keyed on FORCED modes only, never != DISABLE: AUTO(0) is the
resolver's fallthrough for every sub-950Mpix session, and the wider
key would have disarmed the Phase-3 chunked-poll feature fleet-wide
(critic catch). Under AUTO the driver arbitrates — the shipped state.
- AV1: untouched — per-tile sub-frame + split are legal together.
The arbitration is a pure nvenc_core fn called by each backend BEFORE
the ladder, the ceiling key and the chunked-poll latch — all three see
the post-arbitration truth. A drop inside build_init_params would have
left poll_chunk busy-polling its whole budget every AU (numSlices stays
0 without reportSliceOffsets; both loop exits dead — critic catch).
Linux latches subframe_forced beside subframe_on at query_caps (no env
re-reads after open); Windows records the arbitrated state so
reconfigure presents exactly the params the open had (also closes the
pre-existing mid-session env-flip hazard there).
Truth-table tests in nvenc_core; PUNKTFUNK_NVENC_SUBFRAME documented
(it never was); PUNKTFUNK_SPLIT_ENCODE row updated.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bitrate-probe ladder stepped down on format!("{e:#}").contains(
"Invalid argument") — an English substring over the WHOLE context chain,
which also fired on any other wrapped EINVAL (e.g. a CUDA-context errno)
and gated a ~10-step ladder on strerror wording. The root ffmpeg::Error
survives the anyhow chain; downcast and match Error::Other{errno:EINVAL}
instead. Same fix for the intra-refresh ENOSYS probe in the open path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Every capability probe drops libav's log level to AV_LOG_FATAL around an encoder open it
expects to fail, then restores what was there before. That level is one process-global int
and the probes genuinely overlap — they are reached both from `/serverinfo` and from session
bring-up. Two overlapping save/restore pairs interleave as get(INFO) → get(FATAL) →
set(INFO) → set(FATAL), and the process is then pinned at AV_LOG_FATAL permanently: every
later libav diagnostic silently dropped, including the ones you want when a stream fails to
open later on.
Replaces the four hand-rolled save/restore pairs with one RAII guard over a single shared
mutex. The audit filed two sites; there are four — the NVENC 4:4:4 and 10-bit probes in
`linux/mod.rs` and both VAAPI probes, which is why the lock has to be shared across the two
modules rather than kept local to either.
Being RAII also closes a smaller hole the old shape had: the restore was a statement after
the open, so any early return added to those functions later would have leaked the quiet
level. Now it cannot.
The guard is poison-tolerant — a probe that panicked mid-window has already restored the
level through `Drop`, so refusing the lock forever afterwards would be strictly worse than
proceeding. It is not re-entrant, which is safe here because no probe body reaches another
probe (they only call encoder-open helpers); that invariant is written down at the type.
Serialising the probes costs nothing measurable: they are process-once behind their caches
and each already pays for a real encoder open.
Verified with the canonical Linux gate (docker linux/amd64): fmt, clippy --all-targets at
default and at nvenc,vulkan-encode,pyrowave, and the pf-encode test leg (37 passed).
`open_video`'s `cursor_blend` argument was a request with no answer: lib.rs did
`let _ = cursor_blend;` and only three backends ever read `CapturedFrame::cursor`.
So a session could ask for a composited pointer, get a backend that silently
discards it, and stream with no mouse cursor and nothing in the logs. Two
separately-confirmed audit findings — the VAAPI dmabuf path and the libav-NVENC CUDA
path — are symptoms of that one hole.
`EncoderCaps::blends_cursor` makes it a fact each backend states. The four exhaustive
`EncoderCaps { .. }` constructors mean adding the field is a compile error until every
backend answers, which is the enforcement mechanism for future backends rather than a
side effect. Vulkan Video answers from its ACTUAL configured source rather than
statically: only the CSC path composites (`prep_cursor` feeds the compute shader),
while the RGB-direct/EFC front-end and the native-NV12 source have no compositing
stage at all and merely warn once that the pointer is being dropped.
`open_video` warns when a session asked for blending and the opened backend cannot
deliver it. A warning is deliberately all it does: `open_video` cannot re-plan
capture, so refusing would trade a missing pointer for a dead session. The host owns
`plan.cursor_blend` and is the only layer that can fall back to capturer-side
compositing — this gives it something to base that on.
Enforcement is NOT included. The reviewed design proposed refusing the client's
host-composite flip to keep the client drawing its own pointer, but `CursorRenderMode`
is client->host only: there is no host->client counterpart, so refusing yields no
pointer at all — the same failure it claimed to prevent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three encode-time fixes from the 4K120 field analysis (sessions pinned ~107 of
120 fps, 14-15 ms reported encode):
- Ceiling truth: the codec-level bitrate clamp (binary search at session open)
now records its result in a process-lifetime advisory cache keyed by
GPU/config, so an ABR overshoot opens — or in-place reconfigures — straight
AT the ceiling instead of re-running the ~6-open search (and a full rebuild
+ IDR) on every overshoot. New `Encoder::applied_bitrate_bps()` exposes the
post-clamp rate so the session loop can stop pacing/acking a phantom
requested rate (consumed host-side in a follow-up).
- Clamp-search hygiene: only NVENC parameter/caps rejections steer the search
now (`NvCallError` keeps the raw status downcastable); a transient failure
(busy engine, session limit, OOM, driver skew) propagates instead of
shrinking into — and now caching — a bogus ceiling. The floor fallback also
records the split mode it actually opened with, so a later reconfigure
re-presents the real session params.
- Split-frame selection: one shared `resolve_split_mode` replaces the two
byte-identical direct-SDK copies; the force-2-way threshold moves to
`SPLIT_FORCE_PIXEL_RATE` (950 Mpix/s, shared with the libav path) because
4K120 = 995,328,000 px/s missed the old `> 1e9` gate by 0.47% and stayed on
AUTO — which never engages at 2160 px height, leaving the second NVENC
engine idle in exactly the mode the threshold existed for. The Linux
session-ready info! line now carries the final split mode (journals are
INFO+; this was undiagnosable from user logs).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`submit_cpu`'s RGB24/BGR24 → rgb0/bgr0 expand ran `w*h` iterations per frame, each
building two bounds-checked sub-slices for a 3-byte copy plus a pad-byte store — a
shape LLVM will not vectorise into the byte shuffle it actually is. At 3840x2160
that is 8.3M iterations on the encode thread, per frame.
It is not an edge case. The module header says the portal commonly negotiates
packed 24-bit RGB, and pf-capture offers `VideoFormat::RGB` first because wlroots
commonly fixates it — so this was the mainstream CPU path, running roughly an order
of magnitude slower than the 4-bpp sibling branch (a plain row memcpy) purely
because of how it was written.
swscale's packed-RGB expanders are SIMD, the sibling VAAPI backend already routes
RGB24/BGR24 through them, and this file already owned the `sws_getContext` /
`sws_freeContext` lifecycle — so this is net subtractive rather than new machinery:
the existing CSC condition widens to include `expand`, and the hand-written loop
goes away. No new field, no second context, nothing added to `Drop`: `expand` is
false whenever `want_444` (see the `nvenc_pixel`/`expand` binding), so the three
users are mutually exclusive and one context serves whichever applies. The `expand`
field itself is gone with its only reader.
`sws_setColorspaceDetails` is now applied ONLY for the 4:4:4/HDR users. The expand
is a pure byte shuffle — NVENC does the RGB→YUV itself downstream — and handing it
a matrix and range would have silently range-converted every packed-RGB session,
which is exactly what the module header promises does not happen.
Verified `-D warnings --all-targets` + tests on Linux (default; shipped
nvenc+vulkan-encode+pyrowave). The pixel path itself is unverifiable without an
NVIDIA box: `nvenc_hdr10_smoke` and friends are #[ignore]d, so the channel order
and the pad byte want an on-glass check on the CachyOS 5070 Ti.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`sws_getContext` hands back a raw pointer whose only `sws_freeContext` lives in
`Drop for NvencEncoder` — and `Drop` needs a CONSTRUCTED `Self`, which does not
exist on either of `open`'s early returns: the intra-refresh-unsupported path
(which recurses into `Self::open`) and the plain error return. Both sit between
the context's creation and the `Ok(NvencEncoder { … })` that would give it an
owner, so every failed open leaked one.
That is not a once-per-process wart. `open_nvenc_probed` walks an EINVAL bitrate
ladder — requested rate, then the codec-level cap, then ×3/4 down to 50 Mb/s —
re-entering `open` at each step, so a GPU that refuses the requested bitrate leaks
a context per step (up to ~10). The intra-refresh retry adds one more.
Fixed structurally rather than with a guard: the block depends only on values known
before the encoder open (`want_444`, `want_hdr10`, `cuda`, `format`, `nvenc_pixel`,
the dimensions, `full_range_444`), so moving it BELOW `video.open_with(opts)` — to
just above the struct construction, with nothing fallible in between — makes the
leak unrepresentable instead of merely handled. No behaviour change: same context,
same colourspace details, same field.
Found while designing the WP1.4 swscale expand, which would have promoted this from
a 4:4:4/HDR-only leak (the only sessions that build a context today) to one on every
packed-RGB session. Fixing it first is the prerequisite for that change.
Verified `-D warnings --all-targets` + tests on Linux (default; shipped
nvenc+vulkan-encode+pyrowave).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bind the session's IO streams to the encode thread's high-priority copy
stream: in sync-retrieve depth-1 use the per-frame input copy and cursor
blend now enqueue with NO cuStreamSynchronize and encode_picture orders
after them on the stream. Same stream both directions, so the encode's
completion is inserted into the stream and later work (the next frame's
copy into a reused ring slot) waits for it.
Soundness gate: the fast path engages only when pending is empty (true
depth-1 usage) — every prior encode was drained by a blocking poll, and
the caller holds the frame payload across the matching poll (contract now
documented on Encoder::submit; both host loops already comply). Pipelined
callers and PUNKTFUNK_NVENC_ASYNC mode keep the blocking copies.
True zero-copy input registration (registering the worker-owned IPC
buffer directly) stays the LN2 v2 follow-up — it needs a contiguous
worker-pool NV12 layout and a registration<->IPC-mapping lifetime tie.
PUNKTFUNK_NVENC_STREAM_ORDERED=0 restores the old blocking behavior.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`cargo clippy --workspace --all-targets --locked -- -D warnings` was red on
main — three lints landed with the GNOME 50 HDR + PyroWave 4:4:4 work:
* pyrowave_wire.rs: `aw / 2 >> level` tripped clippy::precedence. Rust already
binds `/` tighter than `>>`, so this always parsed as `(aw / 2) >> level`
(subband dim at half res, then one halving per DWT level) — the parens are
purely explicit, no change in behaviour.
* linux/mod.rs: `probe_can_encode_10bit` sat after `mod hdr_tests`
(clippy::items_after_test_module) — moved above the test module, unchanged.
Lint-only; no functional change. fmt/clippy/test all green afterwards.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GNOME 50 (Mutter MR 4928, PipeWire >= 1.6) added HDR screen sharing for
monitor streams: 10-bit PQ formats (xRGB_210LE/xBGR_210LE) with MANDATORY
BT.2020 + SMPTE-2084 colorimetry props, advertised while the mirrored
monitor is in BT.2100 colour mode. Wire the Linux host into it end-to-end
on the GameStream desktop-mirror path (PUNKTFUNK_VIDEO_SOURCE=portal):
* pf-frame: PixelFormat::X2Rgb10/X2Bgr10 (DRM XR30/XB30; X2Bgr10 is the
Windows Rgb10a2 layout) + fourccs.
* pf-capture: want_hdr portal offer — HDR-only LINEAR-dmabuf pods with
MANDATORY PQ/BT.2020 props (SHM excluded: Mutter's SHM record path
paints 8-bit ARGB32 regardless of format; tiled excluded: the EGL
de-tile blit is 8-bit RGBA8), negotiated-colorimetry parse, generic
HDR10 hdr_meta(), packed-10-bit CPU cursor blend, a process-wide SDR
downgrade latch on negotiation timeout, and a DisplayConfig BT.2100
colour-mode probe (gnome_hdr_monitor_active).
* pf-encode: libav NVENC X2RGB10->P010 swscale (BT.2020 limited) ->
HEVC Main10 / 10-bit AV1 with PQ VUI; VAAPI 10-bit on both paths (CPU
P010 upload + dmabuf XR30 scale_vaapi p010/bt2020); can_encode_10bit
now probes for real on Linux; 10-bit sessions route around the
8-bit-only Vulkan-video/direct-NVENC backends.
* GameStream: host_hdr_capable() Linux arm, live monitor-HDR check at
RTSP honor time, capturer-pool reuse keyed on HDR-ness, gs_bit_depth
covers the new formats. New `punktfunk-host hdr-probe` diagnostic and
a PUNKTFUNK_SPIKE_HDR spike lever.
* Native plane stays honestly 8-bit via capturer_supports_hdr(): Mutter
RecordVirtual streams are SDR-only upstream (GNOME 50 and 51-dev), so
virtual-display sources cannot deliver HDR yet.
Validated on the RTX 5070 Ti (GNOME 50.3 / PipeWire 1.6.8): the Main10
probes pass and the ignored nvenc_hdr10_smoke GPU test emits an IDR that
ffprobe reads as Main 10 / yuv420p10le / bt2020nc / smpte2084 / limited.
Live HDR capture negotiation still needs an HDR monitor on glass; VAAPI
10-bit needs the AMD box.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
encode.rs + encode/* (NVENC, VAAPI, native AMF, AMF/QSV ffmpeg, direct-SDK
NVENC/CUDA, raw Vulkan-Video, PyroWave, openh264) move into crates/pf-encode
behind one Encoder trait + open_video selector (plan §W6). The crate speaks the
shared frame vocabulary (pf-frame: CapturedFrame/PixelFormat + the DXGI identity
D3d11Frame/make_device) and pf-zerocopy (CUDA context/buffers), and NEVER
pf-capture — the capture→encode edge is one-way (ZeroCopyPolicy, prior commit).
Dep moves: the heavy encoder deps (ffmpeg-next, the NVENC SDK, openh264,
pyrowave-sys) move from the host to pf-encode; the host's
nvenc/amf-qsv/vulkan-encode/pyrowave features now FORWARD to pf-encode/*. The
host keeps a mod-encode shim (pub use pf_encode) so every crate::encode::* path
(negotiator + GameStream/native/mgmt planes) is unchanged.
resolve_render_adapter_luid moves from the host's windows/win_adapter.rs into
pf-gpu (both pf-encode and pf-capture need it as a peer of GPU selection); its 5
call sites (encode amf/nvenc, capture idd_push/synthetic_nv12, vdisplay manager)
rewire to pf_gpu::resolve_render_adapter_luid and win_adapter.rs is deleted.
pf-frame's make_device gains a # Safety section (public-unsafe-fn lint, latent
since the pf-frame carve — a full-workspace -D warnings clippy catches it).
Verified: Linux clippy -D warnings (pf-encode + host nvenc,vulkan-encode,pyrowave
--all-targets) + 13/13 pf-encode + 299/299 host tests; Windows clippy -D warnings
(pf-encode nvenc,amf-qsv --all-targets + host nvenc,amf-qsv --all-targets)
Finished exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>