6c2183ceec163f75e76cfdf6ff11c44478a29b47
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6c2183ceec |
fix(nvenc): stop telling operators to reboot when the driver version is fine
`NV_ENC_ERR_INVALID_VERSION` is how the driver reports two opposite failures, and we only ever explained one of them. The message told the operator to update the NVIDIA driver or reboot — correct for a genuine header/kernel-module skew, and actively misleading for the field case behind it: a host that streams once per boot and then fails every later session at the caps probe, until the *process* restarts. A version skew is static; it cannot come and go inside one process, so that advice cost the reporter a reboot per stream. Split the two on the only fact that tells them apart: whether a session has already opened in this process. `nvenc_status` gains a `SESSION_OPENED` latch set right after every successful `open_encode_session_ex` — both backends' caps probe and real open, plus the Windows availability probe. No session yet, the version word really is in question and the existing skew advice stands. A session already opened, and the kernel module demonstrably accepted this build's version word, so the message now names per-process driver state and points at the cheap fix (restart the host service, no reboot) plus a request for the log. The load-time gate cannot serve as this discriminator: `NvEncodeAPIGetMaxSupportedVersion` is a pure userspace query, so the classic "updated the driver, didn't reboot" skew sails through it and only fails later at the open. Only a session that actually opened proves the kernel module agreed. The split lives in a pure `invalid_version(bool)` so both halves are unit-tested without touching the process-wide latch. This is a diagnosis change only — it does not fix the underlying field bug, whose root cause is still open. Verified on Linux (192.168.1.25): clippy `-p pf-encode --features nvenc` clean, `cargo test -p pf-encode --features nvenc --lib` 15 passed, rustfmt clean. |
||
|
|
ab63e0dad3 |
fix(encode/nvenc-windows): fail safe when nobody sets the input ring depth
`set_input_ring_depth` is a DEFAULTED trait method, so a caller that forgets it fails silently — and one did. The GameStream loop opens an encoder (gamestream/stream.rs:671 and the rebuild at :831) and never called it, while the Windows IDD-push capturer declares a ring of 2 and `async_inflight_cap()` defaults to 4. With PUNKTFUNK_NVENC_ASYNC=1 a Moonlight session therefore pipelined four encodes against a two-texture ring, letting the capturer rotate a texture out from under a live encode: torn or mixed frames, never an error — precisely the corruption the cap exists to prevent, and precisely what the trait doc warns "fails silently and intermittently". Guarded at the single point of CONSUMPTION rather than by plumbing the setter into every loop: `input_ring_depth` has exactly one reader in the workspace, so one fail-safe covers every caller including ones not yet written, whereas fixing N call sites only fixes the N someone remembered. An unconfigured ring is now treated as the shallowest any capturer here declares, so the unconfigured path degrades to less pipelining — a latency cost, not corruption. The two GameStream sites also pass the REAL depth, because the fail-safe is a floor, not a substitute: `idd_depth` is configurable and a deeper ring is free pipelining the fallback would forfeit. Default (sync) sessions are unaffected — the cap is only read while the async retrieve thread exists, and that is opt-in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c1d54b835b |
fix(encode/nvenc): cache the level ceiling, expose the applied bitrate, force 2-way split at 4K120
Three encode-time fixes from the 4K120 field analysis (sessions pinned ~107 of 120 fps, 14-15 ms reported encode): - Ceiling truth: the codec-level bitrate clamp (binary search at session open) now records its result in a process-lifetime advisory cache keyed by GPU/config, so an ABR overshoot opens — or in-place reconfigures — straight AT the ceiling instead of re-running the ~6-open search (and a full rebuild + IDR) on every overshoot. New `Encoder::applied_bitrate_bps()` exposes the post-clamp rate so the session loop can stop pacing/acking a phantom requested rate (consumed host-side in a follow-up). - Clamp-search hygiene: only NVENC parameter/caps rejections steer the search now (`NvCallError` keeps the raw status downcastable); a transient failure (busy engine, session limit, OOM, driver skew) propagates instead of shrinking into — and now caching — a bogus ceiling. The floor fallback also records the split mode it actually opened with, so a later reconfigure re-presents the real session params. - Split-frame selection: one shared `resolve_split_mode` replaces the two byte-identical direct-SDK copies; the force-2-way threshold moves to `SPLIT_FORCE_PIXEL_RATE` (950 Mpix/s, shared with the libav path) because 4K120 = 995,328,000 px/s missed the old `> 1e9` gate by 0.47% and stayed on AUTO — which never engages at 2160 px height, leaving the second NVENC engine idle in exactly the mode the threshold existed for. The Linux session-ready info! line now carries the final split mode (journals are INFO+; this was undiagnosable from user logs). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
67b7981096 |
feat(encode/nvenc-linux): LN1 phase-3 — multi-slice sub-frame encode DEFAULT-ON for Linux direct-NVENC
apple / swift (push) Successful in 1m16s
apple / screenshots (push) Successful in 6m15s
windows-host / package (push) Successful in 9m39s
ci / web (push) Successful in 1m10s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
Gates recorded (nvenc-subframe-slice-output.md Phase 3): loss-harness curve identical streamed-vs-whole; adversarial security review clean; live .21 3-leg A/B — streamed 0 corrupted frames, p99 8527→5363 µs vs sliced-whole, first_slice_us ~500 µs of a ~1250 µs encode reaching the send thread; wire bytes rate-governed under CBR + 1-frame VBV (the ~1-2 % slice-header cost is absorbed by rate control, not added to the wire). GameStream: plane has zero code diff (own RTP packetizer; shares only untouched send_pacing); a live Moonlight re-test joins the standing owed Moonlight item. - resolve_slices(codec, default): env 1..=32 wins (1 = the NEW explicit single-slice escape), else the backend default — 4 on Linux direct-NVENC, 1 elsewhere. resolve_subframe(default_on): PUNKTFUNK_NVENC_SUBFRAME tri-state (0 = never, 1 = force, unset = backend default). - Linux nvenc_cuda: resolves both ONCE in query_caps — the sub-frame default is gated on the GPU's SUBFRAME_READBACK cap so an unsupporting GPU never has enableSubFrameWrite forced into its init params — and feeds the same resolved values to build_config, build_init_params (open AND in-place reconfigure) and the chunked-poll latch. - Windows: env-only as before (async path untouched, byte-identical config for unset env); LowLatencyConfig.slices + the explicit subframe param replace the shared env reads. - Tests: chunked e2e now runs at the DEFAULTS (no knobs); the fallback test becomes the escape test (SLICES=1 disarms; SUBFRAME=0 disarms while 4-slice encode continues on the plain poll path). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
6f52397342 |
fix(encode): bound NVENC async pipelining by the capturer's texture ring
The Windows direct-NVENC backend registers and encodes the capturer's textures IN PLACE (no CopyResource), so how deep it may pipeline is a property of the CAPTURER, not of the encoder. It was bounded only by `async_inflight_cap()` — `PUNKTFUNK_NVENC_ASYNC_DEPTH`, default 4, clamped to the output-bitstream pool — which consults nothing about the capturer, while the comment at the backpressure loop claimed it "keep[s] in-flight depth within the capturer's texture ring". It never did. The IDD-push capturer rotates `OUT_RING = 3` per delivered frame with no regard for encode completion (its own invariant note says OUT_RING(3) > max pipeline_depth(2)). With the default async depth of 4 the encoder can therefore still be reading a texture the capturer has already handed out again and overwritten: torn or mixed frames. It is visual corruption rather than UB, so it fails silently and intermittently — the worst shape to diagnose from a field report. Adds `Encoder::set_input_ring_depth`, reported from `Capturer::pipeline_depth`, and bounds the async backpressure loop by `min(async_inflight_cap(), depth)`. For IDD-push that yields 2, matching the capturer's stated contract; backends that copy their input, or are synchronous, ignore it. Wired at ALL THREE encoder-creation sites (initial open, stall/resize rebuild, ABR rebuild) and forwarded through `TrackedEncoder` — this crate has a documented trap where an unforwarded defaulted trait method silently no-ops through that wrapper, which has already bitten the direct-NVENC work once and the wire-chunking probe once. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f9668b16a1 |
fix(encode): NVENC partial-init session leak + three backend-parity gaps
apple / swift (push) Successful in 1m18s
apple / screenshots (push) Successful in 4m33s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
ci / rust (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
windows-host / package (push) Has been cancelled
All four found in the pf-encode quality sweep and verified against source. - NVENC partial-init leak (BOTH platforms, high): `init_session` publishes `self.encoder` — and on Windows charges LIVE_SESSION_UNITS — *before* its remaining fallible steps (bitstream buffers; on Linux also the input-surface alloc and `register_resource`). A failure there left a live session with `inited == false`, and every guard on the re-init path keys off `inited`, so the next submit skipped teardown and overwrote `self.encoder`: the session leaked permanently toward the driver's per-process cap, and its budget units never returned, progressively starving parallel-display admission. `teardown` already keys off `encoder.is_null()` rather than `inited`, so it cleans up exactly this half-built state — it just was never called. Now invoked on the `init_session` error path on both platforms. - `can_encode_10bit` asked the wrong backend (medium): it resolved via `linux_auto_is_vaapi`, which ignores `encoder_pref`, while `can_encode_444` and `open_video` honour it. On a host that forces a backend (e.g. `encoder_pref = "vaapi"` on an NVIDIA box) the probe answered for NVENC while the session opened VAAPI, so the negotiated bit depth — and the HDR/SDR colour label derived from it — described a backend that never ran. Now uses the same `linux_zero_copy_is_vaapi` mirror, and `linux_auto_is_vaapi` carries a warning that it resolves the `auto` case only and is not a dispatch mirror. - Linux software arm ignored SW_BITRATE_CEIL (low): the Windows arm clamped openh264 to 100 Mbps, the Linux arm passed the full negotiated rate. The constant is now module-scope so both arms share one value. - QSV/AMF env-parity (low): `PUNKTFUNK_IR_PERIOD_FRAMES` was a no-op on QSV despite the comment claiming parity with AMF, and `PUNKTFUNK_NO_QSV_LTR` / `PUNKTFUNK_INTRA_REFRESH` had dropped AMF's `trim()` and `yes`/`on` spellings, so a value with stray whitespace silently did nothing on Intel while the same value worked on AMD. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ebd9967547 |
feat(pyrowave): Windows host encoder — separate-plane zero-copy D3D11→Vulkan
Wire PyroWave into the Windows host (design/pyrowave-windows-host-zerocopy.md). Before this a macOS client + Windows host that both selected PyroWave silently ran HEVC: the host never advertised CODEC_PYROWAVE and open_video_backend bailed. Approach (zero-copy, no GPU→CPU→GPU): pyrowave owns its own Vulkan device (create_device_by_compat, by render-GPU vendor/device-id — NOT LUID, invalid in Session 0). The capturer runs a BGRA→YUV BT.709-limited CSC (matching rgb2yuv.comp) into TWO SEPARATE shareable plane textures — full-res R8 Y + half-res R8G8 CbCr — which the encoder imports into pyrowave's device. Separate single/two-component textures import reliably on NVIDIA at any size; a single planar NV12 import does NOT (the vendored interop test: "only very specific resource sizes" — confirmed on-glass: 1024² fine, 720p/1080p/1440p garbage). A shared D3D11 fence, signalled after the CSC, is imported as a Vulkan timeline semaphore so the wavelet read is ordered after it. - pf-encode: enc/windows/pyrowave.rs (Encoder impl, two-plane import + Linux-style plane views); host_wire_caps advertises CODEC_PYROWAVE on Windows when the backend isn't Software; open_video_backend routes a negotiated PyroWave session first; pyrowave-sys on the Windows target; interop confirmed at open → clean HEVC fallback. - pf-encode: shared, unit-tested enc/pyrowave_wire.rs (single source of truth for the client-facing AU framing); Linux encoder uses it too. - pf-capture: dxgi.rs BgraToYuvPlanes CSC; idd_push.rs pyrowave mode — forces the virtual display SDR (the VideoProcessor can't ingest the FP16 HDR ring), a two-plane shareable out-ring, a shared fence passed every frame (so a rebuilt encoder re-imports it). Threaded via OutputFormat::pyrowave. - pf-frame: D3d11Frame::pyro carries the CbCr plane + fence; OutputFormat::pyrowave. Verified on .173 (RTX 4090): full-host build + clippy -D warnings (nvenc,amf-qsv) + fmt --all --check; pyrowave_wire unit tests; pyrowave_win_smoke GPU test round-trips distinct Y/Cb/Cr (100/180/60) exactly at 1024²/720p/1080p/1440p; Stage-0 interop validated in the real Session-0 service context on-glass. Deployed to the box. Owed: final on-glass picture/latency confirmation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9a36ea2132 |
refactor(host/W6.2): extract the video encode backends into the pf-encode crate
encode.rs + encode/* (NVENC, VAAPI, native AMF, AMF/QSV ffmpeg, direct-SDK NVENC/CUDA, raw Vulkan-Video, PyroWave, openh264) move into crates/pf-encode behind one Encoder trait + open_video selector (plan §W6). The crate speaks the shared frame vocabulary (pf-frame: CapturedFrame/PixelFormat + the DXGI identity D3d11Frame/make_device) and pf-zerocopy (CUDA context/buffers), and NEVER pf-capture — the capture→encode edge is one-way (ZeroCopyPolicy, prior commit). Dep moves: the heavy encoder deps (ffmpeg-next, the NVENC SDK, openh264, pyrowave-sys) move from the host to pf-encode; the host's nvenc/amf-qsv/vulkan-encode/pyrowave features now FORWARD to pf-encode/*. The host keeps a mod-encode shim (pub use pf_encode) so every crate::encode::* path (negotiator + GameStream/native/mgmt planes) is unchanged. resolve_render_adapter_luid moves from the host's windows/win_adapter.rs into pf-gpu (both pf-encode and pf-capture need it as a peer of GPU selection); its 5 call sites (encode amf/nvenc, capture idd_push/synthetic_nv12, vdisplay manager) rewire to pf_gpu::resolve_render_adapter_luid and win_adapter.rs is deleted. pf-frame's make_device gains a # Safety section (public-unsafe-fn lint, latent since the pf-frame carve — a full-workspace -D warnings clippy catches it). Verified: Linux clippy -D warnings (pf-encode + host nvenc,vulkan-encode,pyrowave --all-targets) + 13/13 pf-encode + 299/299 host tests; Windows clippy -D warnings (pf-encode nvenc,amf-qsv --all-targets + host nvenc,amf-qsv --all-targets) Finished exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |