0d9d78398c0465da34c1ffb734d51ecc77629a13
44
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
751d1de506 |
fix(encode/windows): a stale PUNKTFUNK_ENCODER pin no longer wedges a conflicting GPU selection
On a hybrid box, picking a GPU in the web console whose vendor contradicts a host.env PUNKTFUNK_ENCODER pin produced an unrecoverable session: the pin won backend selection while the adapter followed the console, the wrong-vendor encoder failed deterministically at submit, the reset ladder burned its 5 in-place rebuilds on it, and the client reconnected into the identical wall forever (~10 s per cycle, no visible reason). Three legs: - windows_resolved_backend() now reconciles: a hardware pin whose vendor contradicts the selected GPU is overridden by the adapter-derived backend (capture + encode share one adapter, so honoring the pin can only fail); open_video warns loudly when a pin loses. The reconciliation is a pure, unit-tested table (resolve_windows_backend). - the QSV wrong-adapter bind is typed TerminalEncoderError, and the stream loop's reset ladder ends the session immediately on it instead of feeding a deterministic config error 5 futile rebuilds. - the un-pinned path was already correct (verified on-glass: NVIDIA preference + no pin = stable AV1 10-bit HDR on NVENC), so the override simply takes the conflict case onto that path. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3f4ad08869 |
test(encode/ffmpeg_win): extract the decision logic and pin it with unit tests
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
decky / build-publish (push) Successful in 20s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
deb / build-publish (push) Successful in 9m9s
docker / build-push-arm64cross (push) Successful in 10s
docker / deploy-docs (push) Successful in 24s
deb / build-publish-host (push) Successful in 10m47s
android / android (push) Successful in 14m53s
deb / build-publish-client-arm64 (push) Successful in 10m38s
windows-host / package (push) Successful in 17m40s
ci / web (push) Successful in 1m46s
ci / docs-site (push) Successful in 1m50s
arch / build-publish (push) Successful in 11m37s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 21m11s
apple / swift (push) Successful in 6m4s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 21m15s
ci / bench (push) Successful in 5m47s
ci / rust-arm64 (push) Successful in 9m11s
ci / rust (push) Successful in 21m50s
apple / screenshots (push) Successful in 25m14s
The QSV open-failure fallback (1,400 lines, 23 unsafe, 0 tests) follows the vaapi.rs treatment: the device-free decisions now live in named functions with their contracts pinned — the per-vendor zero-copy default matrix (AMF on-glass-validated on, QSV opt-in), the PUNKTFUNK_FFWIN_POLL_MS clamp-before-µs-conversion (the 27.7-hour-spin class), the readback routing table with its mid-stream depth-change guard, the swscale source map, the QSV display-remoting latency contract (async_depth=1/low_power=1/ look_ahead=0/forced_idr=1/scenario) and the AMF no-B-frames contract, and the per-vendor zero-copy pool bind flags. A probe smoke rides along #[ignore]d for the runner. No FFI plumbing chased; no behavior changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f495b201e1 |
fix(encode): delete the write-only EncoderCaps::supports_hdr_metadata
A caps field nothing reads is a contract nobody honors — and this one shipped write-only: its single reader anywhere in the workspace was a hardware-gated assertion inside pf-encode's own AMF smoke test. Both planes send the static HDR grade out-of-band unconditionally (the native 0xCE datagram per keyframe, the GameStream 0x010e control message), every first-party client reads exclusively that path, and none parse in-band SEI — so the host decision the field was reserved for (suppress out-of-band when the encoder embeds) can never validly exist. The field's doc contract had also rotted in two directions: it claimed set_hdr_meta no-ops when false (native AMF and QSV consume it regardless) and that only Windows direct-NVENC attaches in-band metadata (AMF and QSV do too). The in-band SEI/OBU emission itself is untouched — it stays a bonus for stock decoders, documented at the emit sites; the trait docs now describe the real routing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1cefd37603 |
fix(pf-encode): arbitrate NVENC split-encode vs sub-frame readback (Phase 8)
Verified against nvEncodeAPI.h's own splitEncodeMode doc (user-prompted — the audit's 'exclusive for HEVC' one-liner deserved checking): - H.264: split 'is not applicable' — hard-DISABLE the mode so the written config, CeilingKey, the split diagnostic log and the rejection-retry stay truthful (the retry used to re-open a byte-identical session after an H.264 'split rejection'). The libav path's operator arm gains the codec gate its auto arm always had. - HEVC: split 'not supported if … subframe mode' — when WE force split (TWO/THREE/AUTO_FORCED, the 4K120 throughput lever), sub-frame yields with a logged escape (PUNKTFUNK_SPLIT_ENCODE=0 chooses sub-frame). ⚠ Keyed on FORCED modes only, never != DISABLE: AUTO(0) is the resolver's fallthrough for every sub-950Mpix session, and the wider key would have disarmed the Phase-3 chunked-poll feature fleet-wide (critic catch). Under AUTO the driver arbitrates — the shipped state. - AV1: untouched — per-tile sub-frame + split are legal together. The arbitration is a pure nvenc_core fn called by each backend BEFORE the ladder, the ceiling key and the chunked-poll latch — all three see the post-arbitration truth. A drop inside build_init_params would have left poll_chunk busy-polling its whole budget every AU (numSlices stays 0 without reportSliceOffsets; both loop exits dead — critic catch). Linux latches subframe_forced beside subframe_on at query_caps (no env re-reads after open); Windows records the arbitrated state so reconfigure presents exactly the params the open had (also closes the pre-existing mid-session env-flip hazard there). Truth-table tests in nvenc_core; PUNKTFUNK_NVENC_SUBFRAME documented (it never was); PUNKTFUNK_SPLIT_ENCODE row updated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cb69cd3b8c |
refactor(pf-encode): move the AMF C-ABI mirror to its own file (WP7.4)
The ~410-line hand-mirrored AMF vtable ABI was already isolated in an inline `mod sys` — the remaining value of the audit's 'best split in the crate' is the FILE boundary: amf.rs drops 3379 → 2965 lines and the pure unsafe-FFI surface (25 unsafe fns, zero policy) is reviewable in isolation, which is the crate's stated review goal. Pure move: `#[path = "amf_sys.rs"] mod sys;` keeps the module name and every `sys::` call site byte-identical; the module was self-contained (only `use std::ffi::c_void`, no super:: references). The banner became the file's module docs; contents de-indented one level — rustfmt-clean on the first check, which is what proves the move byte-exact. Gated on .173: all five clippy combos + the amf-qsv,qsv test leg (38 passed — the live AMF matrix on real VCN silicon among them). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9491cc8759 |
refactor(pf-encode): extract the range-family RFI recovery policy (WP7.2)
The two direct-NVENC backends carried hand-copied twins of the same loss-recovery decision: range validity, covering-range dedup, DPB window, clamp — ~30 duplicated lines each. The decision now lives once as nvenc_core::plan_range_recovery (the range half of WP7.2; the slot half is enc/rfi.rs), pure and unit-tested; each backend keeps its session gate, its unsafe per-timestamp driver loop, and its state stores. The step order is load-bearing and now pinned by tests: the covering dedup runs with the UNCLAMPED last and BEFORE the DPB window (a covered re-ask never touches the driver even when the range has since aged out of the DPB), the boundary at next_ts - RFI_DPB is inclusive, and the Invalidate carries the CLAMPED last — which is also what the caller records in last_rfi_range, exactly as the inline code stored it. A driver failure mid-loop still returns false with NO range recorded and no anchor armed. Decline deliberately clears nothing (neither twin touched pending_anchor on decline — same shape as Vulkan's non-clear, opposite of AMF/QSV; do not harmonize). The exact-cover → Covered test records EXISTING behavior including that a covered range survives a forced IDR with zero driver calls — a recorded fact, not an endorsement. RFI_DPB's import leaves both twins: its only per-backend use was the arithmetic that moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9f1e648e4e |
refactor(pf-encode): extract the slot-family RFI recovery policy (WP7.2)
AMF (user-LTR bitfield), QSV (mfxExtRefListCtrl) and Vulkan Video (the
app-owned DPB slot table) each hand-implemented the same loss-recovery
decision: distrust every reference encoded at-or-after the loss start,
anchor on the newest one strictly older. Three copies had already
diverged once — the
|
||
|
|
e680096c6a |
perf(encode): stop re-reading the environment on every submit and poll
`std::env::var` was on three per-frame paths. Measured on `.173` (Windows, 57 environment variables, 2M iterations): **121.9 ns** per call for the NVENC in-flight cap and **114.9 ns** per call for the ffmpeg poll spin, against **0.9 ns** once memoized. On Linux, 32 ns → 1.5 ns. ⚠ Those numbers deflate the filing, and that is worth recording: the audit ranked this as a hot-path defect, but ~120 ns/frame is ~0.003% of a frame budget. The fix is still right — it is free, and it takes a global environment lock off the encode thread — but nobody should schedule it ahead of anything on the strength of the "hot path" framing. The severity RANKING was also inverted, and the measurement confirms why. The site the audit called worst — `nvenc_cuda`'s backpressure loop condition — costs a default session nothing, because the condition short-circuits on `async_rt.is_some()` and the default session never engages the two-thread retrieve. The one that actually pays every frame is Windows `submit`, which consults `async_inflight_cap()` in BOTH arms of the ring-depth match, sync mode included, where the result is then thrown away. Windows `poll` is the second unconditional one, and the audit ranked it last. Memoized inside each helper rather than latched into a session field. Nothing in the workspace mutates these variables at runtime (enumerated: no `set_var` for either key anywhere in first-party code, and the Windows service's arbitrary-key `host.env` loader runs in the SCM supervisor, which re-execs the host as a child and never opens an encoder itself). A field would instead change WHEN the value is read — and the Windows `cap` composes the env with `input_ring_depth`, which `set_input_ring_depth` may change after open, so freezing that half would reintroduce the in-place-overwrite bug the ring term exists to prevent. Two deliberate behaviour changes ride along on `PUNKTFUNK_FFWIN_POLL_MS`, so "behaviour-preserving" describes the memoization only, not this whole commit: - **A 1000 ms ceiling.** The reachable hazard was never the overflow — that needed `ms >= 1.8e16` — it was a slipped digit: `=100000000` was a 27.7-hour spin of the encode thread. - **`.trim()`, now on all three parsers.** An earlier draft applied the house rule (WP7.8) to one of the three, which left `PUNKTFUNK_NVENC_ASYNC=" 1 "` working while `PUNKTFUNK_NVENC_ASYNC_DEPTH=" 6 "` silently fell back to 4 — and memoization would have frozen that silent fallback for the process lifetime. Reachable: the Windows `host.env` loader trims around `=` before stripping quotes, so `=" 2 "` yields a value with inner spaces. ⚠ Correction to an earlier draft of this message, which asserted that the audit's `saturating_mul` proposal would relocate an overflow panic into release builds. That is FALSE and the code comment now says so: `Duration::from_micros(u64::MAX)` is ~1.8e13 seconds, six orders of magnitude below `Duration`'s ceiling, so `Instant + Duration` neither overflows nor panics (measured, with and without debug assertions). `saturating_mul` is still wrong, for a different reason — it sets a deadline ~584,000 years out on a spin that provably never produces the owed AU, i.e. it wedges the encode thread permanently. A hang, not a panic. WP6.1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b3f8803ea3 |
fix(encode/qsv): stop advertising intra-refresh the driver silently dropped
`ir_active` was `cfg.intra_refresh && set.co2.is_some()` — i.e. "we asked for it", not "we got it". Both `Query` and `Init` can return MFX_WRN_INCOMPATIBLE_VIDEO_PARAM, a WARNING the open path deliberately accepts, while dropping the intra-refresh wave on the floor. That lie is not cosmetic. `ir_active` feeds `EncoderCaps::intra_refresh`, so the session advertises gradual refresh to the client; the client then stops asking for the IDRs it would otherwise request on packet loss, and a lost frame gets concealed instead of repaired. The stream degrades exactly when recovery matters. Now confirmed against the driver: a best-effort `GetVideoParam` with a CodingOption2 buffer chained on, read back for `IntRefType`. Deliberately a SEPARATE query rather than a buffer attached to the existing `BufferSizeInKB` call (which is what the audit finding suggested) — that value backs every bitstream allocation, and a runtime that disliked the chained buffer would take it down with the readback. A failure here costs only the verdict and resolves conservatively (trust the request), so the worst case is the previous behaviour. Verified on Intel UHD 750 (.42), which is where the Intel on-glass run originally caught this: with PUNKTFUNK_INTRA_REFRESH=1, H.264 now logs "silently dropped intra-refresh (GetVideoParam reports IntRefType=0)" and advertises it OFF, while H.265 on the same GPU stays ON — it is genuinely active there. Per-codec, so it could never have been decided statically. 8 live QSV tests pass. Gated on .173: clippy -D warnings at nvenc,amf-qsv,qsv (host + pf-encode --all-targets), amf-qsv without qsv, qsv alone, no-features, cargo test --features qsv, rustfmt — all green. qsv.rs is Windows-only so the Linux legs do not compile it. Closes WP3.2(c), the last Phase 3 item that was hardware-blocked. |
||
|
|
e0e845b7df |
fix(encode/pyrowave): pin the NT-handle import contract to consume-on-success-only
import_plane closed the shared NT handle on every pyrowave_image_create
failure, believing pyrowave only consumes it on success. The vendored
truth was messier: Granite's allocator closed by-reference handles
UNCONDITIONALLY after the first vkAllocateMemory — success AND failure —
while failures before the allocator (validation, vkCreateImage, no
memory type) left the handle open, and both classes surface as the same
error code. So the call site could not know whether to close, and an
allocate-stage failure double-closed a possibly-recycled handle value
(audit WP5.3). The filed fix (DuplicateHandle + close the original
unconditionally) traded the double close for a guaranteed leak on every
pre-allocate failure and left the ambiguity in place.
Patch 0006 fixes the callee instead: the allocator never closes a
caller's handle, and pyrowave_image_create consumes it exactly at its
success return — which is what pyrowave.h ("take ownership and close the
HANDLE on import") documents, and what Granite's semaphore import
already did, so import_fence was correct as written and the two paths
now share one contract. The close-on-failure in import_plane is thereby
correct on EVERY path, including a vkBindImageMemory failure after a
successful allocate. Bonus fix recorded in the patch: the allocator's
block-recycling retry loop used to re-run vkAllocateMemory with
import_info still pointing at the just-closed handle.
Both-platform gates green; pyrowave_win_smoke (.173, 34/34) re-validates
the live import path against the rebuilt vendored C++.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d6d4408c8e |
fix(encode/nvenc): teardown that tells wedged from routine, and a session budget that can heal
Three Windows NVENC teardown/accounting defects (audit WP5.2, marked BLOCKING — the obvious fixes were reviewed and rejected before this): Event-handle leak: init_session pushed each completion event to self.events only AFTER nvEncRegisterAsyncEvent, so a failed registration leaked the Win32 handle. Push first — the init-failure path already runs teardown, which closes everything in the list (unregistering the one never-registered event is a harmless error return). Wedged drain: teardown drains the retrieve thread's backlog, and each queued job could burn its full 5 s completion wait — cap x 5 s on the encode thread, paid again on every rebuild of the stall-recovery ladder. The rejected fix (a shutdown flag) abandoned the backlog on EVERY teardown, turning the routine drain into destroy-while-encoding. Instead the wedge evidence stays where it lives: after one full-budget timeout in the retrieve loop, later jobs wait a 250 ms slice; a success resets the latch. Routine teardown is byte-identical to before — nothing is ever abandoned — and a first timeout is already encoder-fatal, so short slices behind an established wedge change only how long teardown blocks. Session-unit accounting: LIVE_SESSION_UNITS was refunded even when destroy_encoder FAILED, drifting the budget low and over-admitting parallel displays. The rejected fix (never refund on failure) let one transient wedge episode poison admission until a host restart. Now the refund follows PROOF: destroy succeeded, or its status says the driver holds no session (device gone/TDR). Ambiguous failures park the handle — units still charged, fail-closed, pinning the D3D11 device the driver session references — and init_session retries the destroy while zero sessions are live, under a gate that serializes every session open against the reap so a recycled handle address can never alias a live session. The charge lasts exactly as long as the driver keeps refusing the destroy, which is the definition of still-leaked. Admission stays a lock-free atomic load. Both-platform gates green (docker linux/amd64 4 legs; .173 all 7 legs). The retry-destroy-on-a-failed-handle assumption is documented at the reap: the SDK defines no retry semantics either way. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
87863c12a9 |
fix(encode/windows): clear the Windows Phase 4 backlog — NVENC sticky state, QSV HDR + ABR
arch / build-publish (push) Failing after 5m39s
ci / web (push) Successful in 1m2s
ci / docs-site (push) Successful in 1m3s
android / android (push) Successful in 11m50s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
ci / bench (push) Successful in 6m25s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
ci / rust-arm64 (push) Successful in 10m5s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 5m22s
deb / build-publish-client-arm64 (push) Successful in 9m34s
deb / build-publish (push) Successful in 10m58s
docker / deploy-docs (push) Successful in 12s
deb / build-publish-host (push) Successful in 11m51s
docker / build-push-arm64cross (push) Successful in 6m20s
apple / swift (push) Successful in 5m22s
ci / rust (push) Successful in 22m30s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m2s
windows-host / package (push) Successful in 10m48s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m13s
apple / screenshots (push) Canceled after 16m50s
Four items, gated on the RTX box (.173, all 7 Windows legs) and the Linux docker gate, with
the QSV HDR one proven by an A/B on real Intel silicon (.42, UHD 750).
WP4.1 — Windows NVENC sticky state. Both halves were the same mistake: a single field doing
duty as both "what the session negotiated" and "what this session can actually do", so the
first downgrade destroyed the negotiation.
- `query_caps` clears `self.hdr` on a GPU without 10-bit encode, but `submit` re-derives the
requested HDR state from every frame's pixel format and compared it against that cleared
value. On such a GPU a P010 capturer therefore reported "HDR changed" on EVERY FRAME and
tore down and rebuilt the whole encode session each time. The trigger now compares against
`hdr_requested`, and the downgrade is latched in `hdr_unsupported` so it is remembered
(and warned once) instead of rediscovered per init.
- One subsampled-YUV frame cleared `chroma_444` permanently, so 4:4:4 never returned when
the capturer went back to RGB. The negotiation now lives in `chroma_444_requested` and the
effective value is recomputed at every init; the per-session downgrade still happens and
is still reported honestly through `caps()`, it just is not destructive any more.
WP4.3(a) — QSV HDR mastering luminance was divided by 10,000. The old comment justified it as
"VPL wants whole cd/m²", which is true of the VIDEO-PROCESSING unit but not the encode path:
the runtime passes the value straight into the ITU-T H.265 Annex D mastering-display SEI,
whose unit is 0.0001 cd/m². Verified end to end on Intel UHD 750 by dumping the bitstream and
reading SEI 137 with trace_headers, before and after:
before: max_display_mastering_luminance = 1000 (0.1 nits)
after: max_display_mastering_luminance = 10000000 (1000 nits)
with `min` (500 = 0.05 nits) and both content-light-level fields unchanged. That asymmetry —
a correct min beside a 10,000x-low max — was the original tell.
WP4.3(b) — `reconfigure_bitrate` never re-read `BufferSizeInKB`, so an ABR step-up could hand
the runtime an AU buffer sized for the old, lower bitrate. Fixed in both places it was broken:
re-read the driver's answer via `GetVideoParam` after the `Reset` (mirroring `init_encode`),
AND stop `take_bs` recycling pooled buffers without checking their capacity — the pool would
otherwise keep serving short buffers even once `bs_bytes` was correct.
WP7.8 — the `PUNKTFUNK_QSV_FFMPEG` half, finally landable now that a Windows gate exists (it
sits behind `#[cfg(feature = "qsv")]`, so nothing in a Linux-only matrix compiles it — it was
deliberately held back for exactly that reason). Trimmed like its siblings, and kept
case-insensitive: the knob already accepted `TRUE`, and the bare house `matches!` would have
silently dropped that spelling.
Phase 4 of the audit is now complete, and WP7.8 with it.
Verified: Windows gate on .173 — clippy -D warnings at nvenc,amf-qsv,qsv (host + pf-encode
--all-targets), amf-qsv without qsv, qsv alone, no-features, `cargo test --features qsv`
(34 passed), rustfmt. Linux docker gate L1-L4 green. On-glass: the QSV A/B above on .42.
Owed: NVENC sticky-state on-glass on .173 (needs a live 10-bit-capable capture toggle).
|
||
|
|
6c2183ceec |
fix(nvenc): stop telling operators to reboot when the driver version is fine
`NV_ENC_ERR_INVALID_VERSION` is how the driver reports two opposite failures, and we only ever explained one of them. The message told the operator to update the NVIDIA driver or reboot — correct for a genuine header/kernel-module skew, and actively misleading for the field case behind it: a host that streams once per boot and then fails every later session at the caps probe, until the *process* restarts. A version skew is static; it cannot come and go inside one process, so that advice cost the reporter a reboot per stream. Split the two on the only fact that tells them apart: whether a session has already opened in this process. `nvenc_status` gains a `SESSION_OPENED` latch set right after every successful `open_encode_session_ex` — both backends' caps probe and real open, plus the Windows availability probe. No session yet, the version word really is in question and the existing skew advice stands. A session already opened, and the kernel module demonstrably accepted this build's version word, so the message now names per-process driver state and points at the cheap fix (restart the host service, no reboot) plus a request for the log. The load-time gate cannot serve as this discriminator: `NvEncodeAPIGetMaxSupportedVersion` is a pure userspace query, so the classic "updated the driver, didn't reboot" skew sails through it and only fails later at the open. Only a session that actually opened proves the kernel module agreed. The split lives in a pure `invalid_version(bool)` so both halves are unit-tested without touching the process-wide latch. This is a diagnosis change only — it does not fix the underlying field bug, whose root cause is still open. Verified on Linux (192.168.1.25): clippy `-p pf-encode --features nvenc` clean, `cargo test -p pf-encode --features nvenc --lib` 15 passed, rustfmt clean. |
||
|
|
ef239691df |
feat(encode): make cursor blending a queryable capability, not an assumption
`open_video`'s `cursor_blend` argument was a request with no answer: lib.rs did
`let _ = cursor_blend;` and only three backends ever read `CapturedFrame::cursor`.
So a session could ask for a composited pointer, get a backend that silently
discards it, and stream with no mouse cursor and nothing in the logs. Two
separately-confirmed audit findings — the VAAPI dmabuf path and the libav-NVENC CUDA
path — are symptoms of that one hole.
`EncoderCaps::blends_cursor` makes it a fact each backend states. The four exhaustive
`EncoderCaps { .. }` constructors mean adding the field is a compile error until every
backend answers, which is the enforcement mechanism for future backends rather than a
side effect. Vulkan Video answers from its ACTUAL configured source rather than
statically: only the CSC path composites (`prep_cursor` feeds the compute shader),
while the RGB-direct/EFC front-end and the native-NV12 source have no compositing
stage at all and merely warn once that the pointer is being dropped.
`open_video` warns when a session asked for blending and the opened backend cannot
deliver it. A warning is deliberately all it does: `open_video` cannot re-plan
capture, so refusing would trade a missing pointer for a dead session. The host owns
`plan.cursor_blend` and is the only layer that can fall back to capturer-side
compositing — this gives it something to base that on.
Enforcement is NOT included. The reviewed design proposed refusing the client's
host-composite flip to keep the client drawing its own pointer, but `CursorRenderMode`
is client->host only: there is no host->client counterpart, so refusing yields no
pointer at all — the same failure it claimed to prevent.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
ab63e0dad3 |
fix(encode/nvenc-windows): fail safe when nobody sets the input ring depth
`set_input_ring_depth` is a DEFAULTED trait method, so a caller that forgets it fails silently — and one did. The GameStream loop opens an encoder (gamestream/stream.rs:671 and the rebuild at :831) and never called it, while the Windows IDD-push capturer declares a ring of 2 and `async_inflight_cap()` defaults to 4. With PUNKTFUNK_NVENC_ASYNC=1 a Moonlight session therefore pipelined four encodes against a two-texture ring, letting the capturer rotate a texture out from under a live encode: torn or mixed frames, never an error — precisely the corruption the cap exists to prevent, and precisely what the trait doc warns "fails silently and intermittently". Guarded at the single point of CONSUMPTION rather than by plumbing the setter into every loop: `input_ring_depth` has exactly one reader in the workspace, so one fail-safe covers every caller including ones not yet written, whereas fixing N call sites only fixes the N someone remembered. An unconfigured ring is now treated as the shallowest any capturer here declares, so the unconfigured path degrades to less pipelining — a latency cost, not corruption. The two GameStream sites also pass the REAL depth, because the fail-safe is a floor, not a substitute: `idd_depth` is configurable and a deeper ring is free pipelining the fallback would forfeit. Default (sync) sessions are unaffected — the cap is only read while the async retrieve thread exists, and that is opt-in. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c1d54b835b |
fix(encode/nvenc): cache the level ceiling, expose the applied bitrate, force 2-way split at 4K120
Three encode-time fixes from the 4K120 field analysis (sessions pinned ~107 of 120 fps, 14-15 ms reported encode): - Ceiling truth: the codec-level bitrate clamp (binary search at session open) now records its result in a process-lifetime advisory cache keyed by GPU/config, so an ABR overshoot opens — or in-place reconfigures — straight AT the ceiling instead of re-running the ~6-open search (and a full rebuild + IDR) on every overshoot. New `Encoder::applied_bitrate_bps()` exposes the post-clamp rate so the session loop can stop pacing/acking a phantom requested rate (consumed host-side in a follow-up). - Clamp-search hygiene: only NVENC parameter/caps rejections steer the search now (`NvCallError` keeps the raw status downcastable); a transient failure (busy engine, session limit, OOM, driver skew) propagates instead of shrinking into — and now caching — a bogus ceiling. The floor fallback also records the split mode it actually opened with, so a later reconfigure re-presents the real session params. - Split-frame selection: one shared `resolve_split_mode` replaces the two byte-identical direct-SDK copies; the force-2-way threshold moves to `SPLIT_FORCE_PIXEL_RATE` (950 Mpix/s, shared with the libav path) because 4K120 = 995,328,000 px/s missed the old `> 1e9` gate by 0.47% and stayed on AUTO — which never engages at 2160 px height, leaving the second NVENC engine idle in exactly the mode the threshold existed for. The Linux session-ready info! line now carries the final split mode (journals are INFO+; this was undiagnosable from user logs). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
42e5f5ad1e |
fix(encode/windows): two dead items the new lint leg exposed in the shipped combo
windows-host / package (push) Successful in 11m35s
ci / web (push) Successful in 52s
ci / docs-site (push) Successful in 1m0s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 13s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 8s
ci / bench (push) Successful in 6m3s
arch / build-publish (push) Successful in 12m26s
android / android (push) Successful in 15m11s
docker / deploy-docs (push) Successful in 11s
deb / build-publish (push) Successful in 11m16s
deb / build-publish-host (push) Successful in 12m13s
apple / swift (push) Successful in 5m48s
ci / rust (push) Successful in 21m49s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m43s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 14m52s
apple / screenshots (push) Successful in 24m41s
|
||
|
|
6be174dc9c |
fix(encode/windows): the encode bit depth must follow the pixels, not the negotiation
All three Windows backends derived it the same wrong way —
`bit_depth >= 10 || matches!(format, P010 | Rgb10a2)` at ffmpeg_win.rs:154,
amf.rs:1275 and qsv.rs:777 — so the NEGOTIATED depth could force a 10-bit encoder
over an 8-bit capture.
That combination is not hypothetical. A client advertises 10-bit, the handshake
negotiates bit_depth=10, and then enabling advanced colour on the IDD virtual
display fails — at which point the capturer says exactly that and delivers 8-bit
NV12 anyway ("10-bit HDR was negotiated but enabling advanced color on the virtual
display FAILED — encoding 8-bit SDR"). Every vendor then lost the session, each in
its own way:
- native AMF and native QSV derived `expected = P010`, saw Nv12, and bail!'d at
open;
- the libavcodec path accepted the open, built a P010 encoder, and then failed
EVERY submit forever, because its per-frame check recomputes the depth from the
frame and never matches. reset() could not help: the rebuild re-derived the
same wrong depth from the same stored bit_depth. The host burned
MAX_ENCODER_RESETS and ended the session.
The last one is the worst of the three because it IS the fallback: native QSV
correctly refuses this input at open, and lib.rs then falls back to the ffmpeg path
"for robustness" — which accepted it and died per frame instead.
Since the duplication was the bug, the fix is one shared `ten_bit_input()` in
codec.rs that all three call, and it follows the delivered pixels. That also keeps
the stream HONEST rather than merely alive: the depth selects the colour signalling
(BT.2020 PQ vs BT.709) and the staging surface format, so an 8-bit capture now
produces an 8-bit stream that says it is SDR — which is what the capturer already
reported it was sending. It warns when the negotiated depth is discarded.
The negotiated depth stays an upper bound. The session LABEL may still claim HDR;
that mismatch lives in the negotiation, not in the encoder, and is not addressed
here.
`is_10bit_format` keeps its (now format-only) definition and is used by
submit_d3d11's per-frame check, so the predicate the encoder was built from and the
one it re-checks per frame cannot drift. That check can now only fire on a genuine
mid-stream depth change.
Verified `--all-targets -D warnings` on Windows (no features / pyrowave / qsv /
nvenc,qsv; 31 tests) and Linux (default; shipped nvenc+vulkan-encode+pyrowave).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
6c97c00add |
fix(encode/pyrowave): the RDO block-index cap must cover 4:2:0, not just 4:4:4
The vendored rate controller packs the 32x32 block index into the low 16 bits of `RDOperation::block_offset_saving` (pyrowave-sys patches/0002). Past u16::MAX the index collides with the `saving` field, the resolve over-credits, and the emitted payload can overshoot the buffer `pyrowave_encoder_packetize` writes into — whose only bounds check is an `assert` that the Release (NDEBUG) vendored build compiles out. There is no Rust-side guard behind it. All three callers of `pyrowave_mode_fits_rdo` hardcoded 4:4:4, leaving 4:2:0 unchecked — but 4:2:0 overflows too, just later: 8192x6144 is 73728 blocks and 8192x8192 is 98304, while `Codec::PyroWave.max_dimension()` permits 8192 per axis, so both are reachable from a client-requested `mode=WxHxFPS`. Worse, the host's only use of the helper (native/handshake.rs) is a 4:4:4 -> 4:2:0 downgrade, so an oversized mode was actively routed INTO the unguarded branch. The cap now lives in `validate_dimensions`, checked against 4:2:0 — the most permissive chroma, so a mode that cannot fit there fits no PyroWave session at any chroma. That is the single chokepoint both the negotiator (handshake.rs:180) and `open_video_backend` already run. Both encoder opens additionally check the chroma actually being opened: the 4:4:4 half, plus defence in depth for the `PUNKTFUNK_ENCODER=pyrowave` lab override. Not dormant: punktfunk-host has `default = ["pyrowave"]` and no build passes `--no-default-features`, so these backends ship in every host binary. Tests pin the arithmetic (8192x6144 = 73728, 8192x8192 = 98304, 7680x4320 still fits) and assert the cap does not leak to H.265/AV1, which have no such controller. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
07348c0175 |
refactor(encode): drop the crate-wide allow(dead_code)
Inherited from the pre-extraction host crate root as scaffolding for backend
paths defined ahead of the build that used them. A census across every feature
combination on both platforms (flip it to `warn`, rebuild) found it was hiding
exactly two items — so it bought nothing while blinding the crate to future rot:
- `vaapi::fourcc` — superseded by `pf_frame::drm_fourcc`; no call sites left.
- `vulkan_video::open_opts` — test-only (the smoke tests use it to pass the
RGB-direct request explicitly rather than through the env), now `#[cfg(test)]`.
Removing it surfaced a real latent defect the Linux census could not see:
`amf.rs`'s `percentile` / `drive_and_measure` helpers are used only by the
`#[cfg(feature = "amf-qsv")]` latency A/B benchmark, so a `--features nvenc,qsv`
build compiled the helpers with their caller gated out. They now carry the same
gate as their caller.
Verified `--all-targets -D warnings` on Linux (no features; shipped
nvenc+vulkan-encode+pyrowave) and Windows (no features; pyrowave; qsv; nvenc,qsv).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
428bd2519c |
test(qsv): converter→ring-profile-P010→encoder live e2e (the RTV-written seam)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
87d7eceabf |
fix(encode/qsv): write the colour VUI unconditionally + Intel-lane colour diagnostics
The native QSV encoder only attached mfxExtVideoSignalInfo on HDR sessions — an SDR stream carried no colour description at all (NVENC writes 709-limited unconditionally; qsv.rs now mirrors it). Diagnostics for the field-reported blue/magenta Intel-host colours: - hdr-p010-selftest takes an optional WxH + GPU vendor (intel|nvidia|amd) and prints which adapter it tested — 64x64 on the default GPU proved nothing on dual-GPU boxes, and the field heights (1080/1400) are not 16-aligned. - pf-capture: ignored live test running the converter at 1920x1080 pinned to the Intel adapter. - pf-encode: qsv_live_p010_1080_colorbars_dump — known P010 bars through the UNALIGNED-height ingest copy (1080 src -> align16 1088 pool, the seam no 640x480 test exercises) to Main10 HEVC, dumped for off-box decode checks. - dxgi.rs: drop the stale 'falls back to the R10 path' claim (no such path). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
67b7981096 |
feat(encode/nvenc-linux): LN1 phase-3 — multi-slice sub-frame encode DEFAULT-ON for Linux direct-NVENC
apple / swift (push) Successful in 1m16s
apple / screenshots (push) Successful in 6m15s
windows-host / package (push) Successful in 9m39s
ci / web (push) Successful in 1m10s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
Gates recorded (nvenc-subframe-slice-output.md Phase 3): loss-harness curve identical streamed-vs-whole; adversarial security review clean; live .21 3-leg A/B — streamed 0 corrupted frames, p99 8527→5363 µs vs sliced-whole, first_slice_us ~500 µs of a ~1250 µs encode reaching the send thread; wire bytes rate-governed under CBR + 1-frame VBV (the ~1-2 % slice-header cost is absorbed by rate control, not added to the wire). GameStream: plane has zero code diff (own RTP packetizer; shares only untouched send_pacing); a live Moonlight re-test joins the standing owed Moonlight item. - resolve_slices(codec, default): env 1..=32 wins (1 = the NEW explicit single-slice escape), else the backend default — 4 on Linux direct-NVENC, 1 elsewhere. resolve_subframe(default_on): PUNKTFUNK_NVENC_SUBFRAME tri-state (0 = never, 1 = force, unset = backend default). - Linux nvenc_cuda: resolves both ONCE in query_caps — the sub-frame default is gated on the GPU's SUBFRAME_READBACK cap so an unsupporting GPU never has enableSubFrameWrite forced into its init params — and feeds the same resolved values to build_config, build_init_params (open AND in-place reconfigure) and the chunked-poll latch. - Windows: env-only as before (async path untouched, byte-identical config for unset env); LowLatencyConfig.slices + the explicit subframe param replace the shared env reads. - Tests: chunked e2e now runs at the DEFAULTS (no knobs); the fallback test becomes the escape test (SLICES=1 disarms; SUBFRAME=0 disarms while 4-slice encode continues on the plain poll path). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b0fbb80fd5 |
fix(encode): key the PyroWave plane-import cache on the capturer's ring generation
Completes the partial fix from the previous commit. The Windows PyroWave backend caches its imported plane images keyed on the D3D11 texture's COM address, and holds no reference on that texture — so once the capturer recreates its ring, those addresses can be handed straight back out by the allocator and a pointer-keyed cache hit returns an image bound to a texture that no longer exists. Adding the extent to the key ruled out same-address-different-size aliasing, but a recycle at identical dimensions still aliased. The capturer already tracks exactly the value needed: `generation`, bumped on every ring recreate. Plumbed it onto `PyroFrameShare` and the encoder now flushes every cached import when it changes, which makes cache identity independent of allocator behaviour rather than a bet against pointer reuse. Validated on the RTX box: `pyrowave_win_smoke` (forced with `--ignored`, the only test that actually exercises this path on real hardware) passes all ten configurations — 1024²/720p/1080p/1440p across SDR/HDR and 4:2:0/4:4:4 — with correct decoded chroma means, so the steady-state cache-hit path still works. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
6f52397342 |
fix(encode): bound NVENC async pipelining by the capturer's texture ring
The Windows direct-NVENC backend registers and encodes the capturer's textures IN PLACE (no CopyResource), so how deep it may pipeline is a property of the CAPTURER, not of the encoder. It was bounded only by `async_inflight_cap()` — `PUNKTFUNK_NVENC_ASYNC_DEPTH`, default 4, clamped to the output-bitstream pool — which consults nothing about the capturer, while the comment at the backpressure loop claimed it "keep[s] in-flight depth within the capturer's texture ring". It never did. The IDD-push capturer rotates `OUT_RING = 3` per delivered frame with no regard for encode completion (its own invariant note says OUT_RING(3) > max pipeline_depth(2)). With the default async depth of 4 the encoder can therefore still be reading a texture the capturer has already handed out again and overwritten: torn or mixed frames. It is visual corruption rather than UB, so it fails silently and intermittently — the worst shape to diagnose from a field report. Adds `Encoder::set_input_ring_depth`, reported from `Capturer::pipeline_depth`, and bounds the async backpressure loop by `min(async_inflight_cap(), depth)`. For IDD-push that yields 2, matching the capturer's stated contract; backends that copy their input, or are synchronous, ignore it. Wired at ALL THREE encoder-creation sites (initial open, stall/resize rebuild, ABR rebuild) and forwarded through `TrackedEncoder` — this crate has a documented trap where an unforwarded defaulted trait method silently no-ops through that wrapper, which has already bitten the direct-NVENC work once and the wire-chunking probe once. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a38adad943 |
fix(encode): bound GPU waits, validate encode status, repair command-buffer and cache invariants
Five of the nine medium findings from the pf-encode sweep. The remaining four need cross-crate plumbing or an unwind refactor and are deliberately left out. - vulkan_video `enqueue` waited on the backpressure fence with `u64::MAX`. That wait runs ON the host encode thread — the same thread the stall watchdog's `reset()` would run on — so a wedged GPU parked the one thread that could recover the session: no error, no reset, and teardown blocking on the join. This is the DEFAULT encode path for AMD/Intel Linux hosts (both shipped build recipes enable `vulkan-encode` and `vulkan_encode_enabled()` defaults true). Now bounded by ENCODE_FENCE_TIMEOUT_NS with expiry surfaced as an error. - vulkan_video `import_cached` evicted a cached dmabuf import and destroyed its image/view/memory with no fence wait, while up to `ring_depth - 1` submitted frames may still reference it — a GPU-side use-after-free. `Drop` and `reset` both idle first; this was the one unguarded destroy. Now idles before the eviction loop, guarded on the length test so the steady state pays nothing. - vulkan_video `read_slot` never asked for the encode's operation status, so a FAILED encode was indistinguishable from a successful one and its feedback was read as if it described real bitstream. Now requests WITH_STATUS_KHR and refuses anything that is not COMPLETE. - linux/pyrowave `encode_frame` opens its recording window early and has six fallible steps inside it; every one returned with `cmd` still RECORDING, and nothing repaired it (one `begin_command_buffer` in the file, and neither `reset()` nor `Drop` touches `cmd`), so the next frame called `begin` on a recording buffer — invalid usage. `submit` now resets the buffer on error; legal on all these paths since the pool carries RESET_COMMAND_BUFFER and the buffer is not pending. - windows/pyrowave `encode_frame` ignored `frame.width`/`height` and imported planes at the encoder's configured extent, so a ring recreate at a new mode (the IDD capturer does this autonomously on a confirmed descriptor change) read the planes under a stale VkImageCreateInfo. Added the size guard its QSV and AMF siblings already carry, and keyed the plane cache on (address, width, height) so a recycled COM address cannot resurrect an import of a different size. NOTE: a recycle at the SAME size is still theoretically possible; the complete fix keys on the capturer's ring generation and needs that plumbed onto `PyroFrameShare`. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b22d0da75b |
fix(encode): port the RFI taint sweep to Vulkan Video, close the QSV sweep hole, bounds-check encode feedback
apple / swift (push) Successful in 1m19s
apple / screenshots (push) Successful in 6m24s
ci / web (push) Successful in 49s
ci / docs-site (push) Successful in 51s
ci / bench (push) Successful in 6m55s
deb / build-publish (push) Successful in 9m12s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 34s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
arch / build-publish (push) Successful in 17m2s
docker / deploy-docs (push) Successful in 14s
android / android (push) Successful in 17m59s
deb / build-publish-host (push) Successful in 11m19s
ci / rust (push) Successful in 25m3s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m52s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m54s
windows-host / package (push) Failing after 13m20s
Three defects from the pf-encode sweep, each adjudicated against source. - Vulkan Video never received fecbec2d's taint sweep (it was carved out one commit later). `pick_recovery_slot` accepts any resident slot whose wire is below the CURRENT loss start, but "resident and older than this loss" is not "the client decoded it": after an earlier loss [a,b] recovered at wire r, everything in [a, r-1] is undecodable at the client — the lost frames plus every frame that predicted through the gap — and those wires stay eligible until the 8-slot ring rolls them out. A later loss could therefore anchor on one and ship it tagged `recovery_anchor`, which is the client's definitive re-anchor signal (punktfunk-core/src/reanchor.rs): the host lifts the client's post-loss freeze onto a picture built from a reference it never had. Swept before anchor selection, matching AMF/QSV. `slot_wire` is blanked and `slot_poc` deliberately is NOT: `slot_poc` feeds `build_h265_rps_s0`, which must keep naming every physically-resident DPB picture or a conforming decoder evicts them and the anchor then references a picture the client already dropped. - QSV's sweep was incomplete, and in its MODAL case. `ltr_slots` mirrors the hardware DPB, but nulling an entry issues no VPL call — the frame stays marked long-term until that LongTermIdx is re-marked or an IDR flushes it (amf.rs states this verbatim). The rejection loop iterates the post-sweep mirror and only rejects `Some` slots, so it silently skipped the single entry the sweep exists to distrust, leaving the recovery frame free to predict from it. With NUM_LTR_SLOTS=2 the "exactly one slot swept" case is the common one, and the two existing tests cover only the both-survive and both-swept cases. Taint is now recorded in `ltr_tainted` with the FrameOrder left in place, so anchor selection and the queued-force guard skip it while the rejection list still names it. - `read_slot` built a slice from the driver-reported (offset, bytes-written) encode feedback with no validation against `bs_size`, so a driver reporting a range outside the bitstream buffer produced an out-of-bounds read shipped straight onto the wire. Checked in u64 (so the add cannot wrap) before `map_memory`, so the error path has no unmap to unwind. Adds `taint_sweep_excludes_slots_from_an_earlier_loss` covering the two-loss case the existing single-loss test does not reach. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f9668b16a1 |
fix(encode): NVENC partial-init session leak + three backend-parity gaps
apple / swift (push) Successful in 1m18s
apple / screenshots (push) Successful in 4m33s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
ci / rust (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
windows-host / package (push) Has been cancelled
All four found in the pf-encode quality sweep and verified against source. - NVENC partial-init leak (BOTH platforms, high): `init_session` publishes `self.encoder` — and on Windows charges LIVE_SESSION_UNITS — *before* its remaining fallible steps (bitstream buffers; on Linux also the input-surface alloc and `register_resource`). A failure there left a live session with `inited == false`, and every guard on the re-init path keys off `inited`, so the next submit skipped teardown and overwrote `self.encoder`: the session leaked permanently toward the driver's per-process cap, and its budget units never returned, progressively starving parallel-display admission. `teardown` already keys off `encoder.is_null()` rather than `inited`, so it cleans up exactly this half-built state — it just was never called. Now invoked on the `init_session` error path on both platforms. - `can_encode_10bit` asked the wrong backend (medium): it resolved via `linux_auto_is_vaapi`, which ignores `encoder_pref`, while `can_encode_444` and `open_video` honour it. On a host that forces a backend (e.g. `encoder_pref = "vaapi"` on an NVIDIA box) the probe answered for NVENC while the session opened VAAPI, so the negotiated bit depth — and the HDR/SDR colour label derived from it — described a backend that never ran. Now uses the same `linux_zero_copy_is_vaapi` mirror, and `linux_auto_is_vaapi` carries a warning that it resolves the `auto` case only and is not a dispatch mirror. - Linux software arm ignored SW_BITRATE_CEIL (low): the Windows arm clamped openh264 to 100 Mbps, the Linux arm passed the full negotiated rate. The constant is now module-scope so both arms share one value. - QSV/AMF env-parity (low): `PUNKTFUNK_IR_PERIOD_FRAMES` was a no-op on QSV despite the comment claiming parity with AMF, and `PUNKTFUNK_NO_QSV_LTR` / `PUNKTFUNK_INTRA_REFRESH` had dropped AMF's `trim()` and `yes`/`on` spellings, so a value with stray whitespace silently did nothing on Intel while the same value worked on AMD. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
04e4394ee0 |
fix(encode): close the two teardown memory-safety holes in the reset paths
apple / swift (push) Successful in 1m17s
apple / screenshots (push) Successful in 6m14s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / web (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
deb / build-publish (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
windows-host / package (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
Both surfaced in a post-refactor quality sweep of pf-encode and were then verified against the source (and, for pyrowave, against the C side). - PyroWave (BOTH platforms): `reset()` destroyed the encoder and, when the rebuild failed, returned `false` leaving `pw_enc` pointing at the freed object — `Drop` then destroyed it a second time. `pyrowave_encoder_destroy` is a plain `delete` (pyrowave_c.cpp:1184, which also reads `encoder->device` afterwards) with no null check, so this is a real double free. The failure branch is not vacuous: the rebuild fails when the device is lost/OOM, which is exactly the state that makes the stall watchdog call `reset()` in the first place, so the host corrupts its heap on the path that runs when things are already going wrong. Now nulls `pw_enc` before the fallible create, publishes only on success, and null-guards both `Drop` and `encode_frame` (the Windows `Drop` already guarded `sync` this way). - QSV: `reset()` dropped `pending` — each entry owning the `Box<BsBuf>` the runtime writes into asynchronously — BEFORE `MFXVideoENCODE_Close` aborted those operations, so the VPL runtime could write into freed heap. The preceding drain is best-effort and bails on the first `Err`, i.e. precisely the wedged-encoder case that triggers the reset. Fixed by ordering: Close, then clear. The full-teardown path was already correct (`Inner` declares `session` before `pending`, and fields drop in declaration order). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
fecbec2daf |
fix(encode): make LTR-RFI loss recovery sound under sustained loss
ci / web (push) Successful in 53s
ci / docs-site (push) Successful in 1m6s
apple / swift (push) Successful in 1m17s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 5m40s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 5m6s
apple / screenshots (push) Successful in 6m12s
deb / build-publish (push) Successful in 9m5s
docker / deploy-docs (push) Successful in 24s
deb / build-publish-host (push) Successful in 13m49s
windows-host / package (push) Successful in 16m17s
arch / build-publish (push) Successful in 17m24s
android / android (push) Successful in 12m53s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 14m31s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m5s
ci / rust (push) Successful in 25m35s
Field report (lid-closed Intel laptop, ~6-19% sustained loss): the stream never healed — permanent macroblock soup. Three stacked bugs: - QSV answered RFI with PreferredRefList only, a reorder HINT per the VPL spec — the recovery frame could keep predicting from tainted short-term refs. Now rejects every other DPB candidate (RejectedRefList) and caps L0 at one active entry (AVC/HEVC), matching AMF's hard ForceLTRReferenceBitfield / NVENC invalidation semantics. - Neither QSV nor AMF taint-swept LTR slots across losses: a slot marked inside the client's corrupt window became the "known-good" anchor of the NEXT loss, propagating corruption through every recovery. Both now drop slots at-or-after the loss start before picking an anchor, and guard a queued force whose slot the sweep emptied (no false recovery_anchor tag). - The native plane re-anchored the FULL IDR cooldown on every successful RFI, so under sustained loss the client's escalating keyframe requests were coalesced away indefinitely (field log: dozens swallowed, one IDR per ~8 s). RFI now anchors a 300 ms echo window with a 2-swallow budget per loss episode; a client still asking past that gets its IDR. Live-validated on Arc (qsv feature): 6/6 including the new qsv_live_ltr_rfi_taint_sweep_declines (a loss covering every live mark declines the RFI and falls back to IDR recovery). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
232b41e88b |
fix(pf-encode/windows): cover the Linux HDR pixel formats in the win swscale match
The GNOME 50 HDR work added PixelFormat::X2Rgb10 / X2Bgr10 but only taught the
Linux encoders about them. `sws_src` in the Windows-gated ffmpeg_win.rs matches
PixelFormat exhaustively, so the Windows host stopped compiling:
error[E0004]: non-exhaustive patterns: `X2Rgb10` and `X2Bgr10` not covered
--> crates\pf-encode\src\enc\windows\ffmpeg_win.rs:132:14
Linux CI never caught it — the file is cfg(windows), so `cargo clippy
--workspace` on the Linux runner never compiles it.
Both are Linux-only screencast formats (the Windows HDR path stays
Rgb10a2/P010, per the PixelFormat docs), so they join the existing bail arm.
Spelled out rather than folded into a `_` catch-all so the next PixelFormat
addition breaks this match again on purpose.
Verified: cargo check --workspace --all-targets --features nvenc,amf-qsv on the
Windows box (192.168.1.173).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
75474dcc90 |
fix(pyrowave): guard 4:4:4 modes that overflow the rate controller's block index
The vendored rate controller packs its wavelet block index into 16 bits (RDOperation.block_offset_saving), so a mode whose 32x32-block count exceeds u16::MAX wraps inside the controller and corrupts the bitstream — ~8K-class 4:4:4 territory. Compute the exact count (`block_count_32x32`, the counting walk of upstream init_block_meta, pinned against the validated Apple WaveletLayout) and expose `pyrowave_mode_fits_rdo`; the negotiator downgrades such a session to 4:2:0 before the Welcome (the honest-downgrade channel), and both encoders refuse outright if one slips through rather than emit a wrapped stream. Vendor patches: 0002-rdo-saving-clamp (analyze_rate_control.comp clamps the saving accumulation to the target, same overrun class as 0001; slangmosh.hpp regenerated), 0003-devel-encode-16bit-read (devel tool y4m 16-bit plane reads; tool-only, kept so the vendored source stays honest). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ac0e73321c |
perf(pyrowave): elevated GPU scheduling + global-priority encode queue
apple / screenshots (push) Successful in 6m18s
ci / web (push) Successful in 1m25s
ci / docs-site (push) Successful in 1m5s
android / android (push) Successful in 13m2s
arch / build-publish (push) Successful in 12m39s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
ci / bench (push) Successful in 5m36s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
windows-host / package (push) Successful in 16m23s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m8s
deb / build-publish-host (push) Successful in 10m16s
deb / build-publish (push) Successful in 12m30s
ci / rust (push) Successful in 19m26s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 22m20s
docker / deploy-docs (push) Successful in 23s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 24m32s
apple / swift (push) Successful in 1m17s
PyroWave's wavelet encode runs on the GPU's compute/shader cores, so a GPU-bound game starves it: submit spikes from ~2 ms to ~15 ms under a 95%+ game load and the stream fps collapses. NVENC is immune (separate encoder ASIC). Two levers to let the encode get scheduled ahead of the game's rendering: - Windows process GPU scheduling: D3DKMTSetProcessSchedulingPriorityClass, env PUNKTFUNK_GPU_PRIORITY = off|above-normal|high (default)|realtime. Best-effort, once per process, non-fatal on refusal (enc/windows/pyrowave.rs). - Global-priority Vulkan encode queue (Granite patch 0005): request a VK_KHR_global_priority queue (PYROWAVE_QUEUE_PRIORITY = off|high|realtime, default realtime), downgrading REALTIME→HIGH→none on NOT_PERMITTED so a refused class never regresses the encoder to HEVC. HONEST STATUS: on an RTX 4090 / Windows / WDDM neither moved the ~15 ms spikes — the graphics-vs-compute preemption granularity is the wall, not the priority level. Kept because both are correct, harmless (graceful fallback), and may help other GPUs/drivers. For a GPU-saturated game the working levers are reducing the encode's GPU cost (4:2:0/8-bit) or H.265; PyroWave holds full rate on the desktop and in games that leave the GPU headroom. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9fe9c451dc |
perf(pyrowave): pool encoder scratch buffers + fix client parser O(n²) — lift the 2.5 Gbps wall
ci / web (push) Successful in 47s
ci / docs-site (push) Successful in 1m7s
apple / swift (push) Successful in 1m17s
decky / build-publish (push) Successful in 25s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
ci / bench (push) Successful in 5m51s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 4m51s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
release / apple (push) Successful in 9m11s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m6s
deb / build-publish (push) Successful in 12m55s
android / android (push) Successful in 13m15s
docker / deploy-docs (push) Successful in 25s
deb / build-publish-host (push) Successful in 13m7s
arch / build-publish (push) Successful in 16m4s
windows-host / package (push) Successful in 16m19s
apple / screenshots (push) Successful in 6m34s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 14m47s
ci / rust (push) Successful in 24m17s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 14m31s
Real-world PyroWave streaming maxed ~2.5 Gbps with sagging fps while raw transport does 4.8. Root-caused to serial per-frame paths at BOTH ends (the transport was never the limit); this fixes the two dominant ones. Host (vendored shim, patch 0004): pyrowave_encoder_encode_gpu_synchronous allocated four Vulkan buffers (meta + bitstream, Device + CachedHost) on EVERY frame. At 240 fps with MB-scale bitstreams that per-frame allocator churn stalled the encode itself. Pool them on the encoder and reuse across frames (recreate only on a size grow); the sizes are session-fixed, so it is pure reuse after frame 1. On an RTX 4090 the 5120x1440 submit+fence-wait drops ~15 ms -> ~1 ms, i.e. the host serial ceiling goes 64 -> 1025 fps (444+HDR 44 -> 614). Safe under the synchronous encode model; re-validated by pyrowave_win_smoke (Windows) and pyrowave_smoke/_444 (Linux). Applies to both host encoder paths (they share the shim). Client (Apple Metal decoder): WaveletBitstream.parse reserved the payload buffer per packet (reserveCapacity(count + words), an exact realloc each of ~3000 packets/frame => O(n²)) and copied word-by-word. Reserve once up front and memcpy each packet's coefficients in one shot (all Apple platforms are little-endian, so the wire's LE u32s land verbatim; memcpy is alignment-free). 5.44 ms -> 0.055 ms per 1.44 MB frame (25x); byte-identical (parser unit tests + golden-frame PSNR unchanged). Also: - native.rs: PUNKTFUNK_PYROWAVE_MAX_MBPS caps PyroWave's open-loop Automatic bitrate pin for hosts on a constrained link (unset => no cap; an explicit client rate bypasses it). The pin is all-intra + ABR-off, so at a high pixel rate it can outrun the fabric (4:4:4+HDR 5120x1440@240 pins ~5.3 Gbps, over a 5 GbE link) and the overshoot just becomes loss. - pf-encode caps(): report the real opened chroma instead of a hardcoded 4:2:0 default, so a genuine 4:4:4 session no longer trips the spurious "encoder chroma disagrees with the negotiated Welcome" warn. Also fix a latent Windows reset() that rebuilt at 4:2:0 for a 4:4:4 session. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
188edde2b3 |
feat(pyrowave): Windows host HDR + 4:4:4, Rust client HDR present
Phase 3 of design/pyrowave-444-hdr.md. A PyroWave session now negotiates HDR
(10-bit) and 4:4:4 on a Windows host exactly like HEVC/AV1, and the Linux
client presents it through the real HDR10 path.
Host (Windows): BgraToYuvPlanes becomes mode-aware — SDR/BGRA and HDR/scRGB
variants at half- or full-res chroma. The HDR passes reuse HdrP010Converter's
exact colour math (scRGB -> PQ BT.2020 limited studio codes, verified by
hdr_p010_selftest) but write P010-style MSB-packed codes into two separate
shareable R16_UNORM/R16G16_UNORM textures; chroma keeps the pyrowave family's
centre-sited 2x2 box. idd_push pins the composition to the NEGOTIATED depth
(SDR sessions force advanced color off as before; 10-bit sessions enable it
and ride the FP16 ring), and the descriptor poller re-asserts that state
instead of following display flips the fixed-format encoder can't. The
encoder imports 8/16-bit planes per session and stamps the sequence header's
BT.2020/PQ/matrix bits on HDR (stamp_color_bits, extending 574e3e4e's range
stamp); supports_10bit/can_encode_10bit/can_encode_444 gates open (HDR
Windows-only — Linux capture has no HDR source).
Client: the plane ring becomes R16_UNORM for 10-bit sessions (with a
STORAGE_IMAGE format probe), the planar CSC pass joins the HDR10 swapchain
rebuild (set_hdr_mode previously destroyed it without rebuilding — latent),
st.hdr follows frame.color.is_pq(), and the planar push constants carry
depth-10 MSB-packed rows + the PQ tonemap mode, identical to the NV12 arm.
Verified: .173 (RTX 4090) deploy-config clippy + fmt + wire tests + the
extended pyrowave_win_smoke (10-case {SDR,HDR}x{420,444} matrix incl. R16
imports and header stamps); .21 (RTX 5070 Ti) clippy across 4 crates, host
186 tests, client/presenter/encode tests, both Linux GPU smokes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
5eb930e71d |
feat(pyrowave): negotiation plumbing for 4:4:4 + HDR — thread chroma/depth/ColorInfo end to end
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 6m49s
ci / web (push) Successful in 58s
ci / docs-site (push) Successful in 1m10s
arch / build-publish (push) Successful in 11m34s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 6m43s
ci / bench (push) Successful in 5m21s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 8m9s
deb / build-publish-host (push) Successful in 9m20s
android / android (push) Successful in 22m17s
flatpak / build-publish (push) Successful in 6m12s
ci / rust (push) Successful in 19m3s
deb / build-publish (push) Successful in 12m40s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 9m8s
apple / swift (push) Successful in 1m18s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m17s
docker / deploy-docs (push) Successful in 11s
apple / screenshots (push) Successful in 6m14s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 23m5s
windows-host / package (push) Successful in 16m27s
Phase 1 of design/pyrowave-444-hdr.md. No behavior change yet: the handshake's
4:4:4 gate now admits PyroWave (probe = can_encode_444(codec), capture gate
inherently satisfied — the wavelet path always ingests an RGB source and does
its own CSC), but can_encode_444 stays false for PyroWave until the per-OS
full-res-chroma CSC variants land (Phase 2 Linux, Phase 3 Windows), so every
session still resolves 4:2:0/8-bit.
- Both host encoders take the negotiated ChromaFormat (bail on 444 for now);
the PUNKTFUNK_ENCODER=pyrowave lab override pins 4:2:0.
- Bitrate: the automatic ~1.6 bpp pin resolves AFTER depth+chroma and scales
x1.625 for 4:4:4 / x1.15 for 10-bit (factors from the Phase-0 fixture
matrix); the mid-stream mode-switch re-resolve threads the session's values.
- Client: PyroWaveDecoder builds its plane ring (full-res chroma when 444) and
creates the upstream decoder from the negotiated chroma, keeps chroma fixed
across mid-stream resizes, drops the even-dims requirement for 444, and
returns the negotiated Welcome ColorInfo as the frame colour contract
instead of hardcoded BT.709 (the wavelet bitstream has no VUI).
Verified on .21 (RTX 5070 Ti): clippy -D warnings (host+client+encode), host
186 tests, client + pf-encode tests, fmt, and the pyrowave_smoke GPU
round-trip through the patched vendored lib (
|
||
|
|
574e3e4e3f |
fix(pyrowave): signal ycbcr_range=LIMITED in the sequence header
docker / deploy-docs (push) Successful in 23s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 22m2s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 22m8s
apple / swift (push) Successful in 1m23s
apple / screenshots (push) Successful in 6m20s
ci / docs-site (push) Successful in 1m10s
ci / web (push) Successful in 1m35s
android / android (push) Successful in 13m8s
windows-host / package (push) Successful in 16m16s
ci / bench (push) Successful in 5m21s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 5m28s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
arch / build-publish (push) Successful in 18m52s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
deb / build-publish-host (push) Successful in 9m38s
deb / build-publish (push) Successful in 12m35s
ci / rust (push) Successful in 19m12s
pyrowave's encoder fills the BitstreamSequenceHeader with `= {}` and its C API
offers no way to set colour/range, so it signals ycbcr_range=0=FULL — but both
host CSCs (rgb2yuv.comp on Linux, BgraToYuvPlanes on Windows) always emit BT.709
LIMITED Y'CbCr (black = Y'16). A client that honours the VUI (the Apple wavelet
decoder reads bit 30 of word1) then skips the limited→full expansion and shows
washed-out, raised blacks — reported on both Linux and Windows hosts.
Patch the range bit HONEST (mark_limited_range in the shared pyrowave_wire, called
by both encoders after packetize). Clients that hardcode limited (the Vulkan
video_pyrowave path) are unaffected, and pyrowave's own decode ignores the flag
(raw Y'CbCr reconstruction). No client rebuild needed. Unit-tested + asserted in
pyrowave_win_smoke; the smoke decode still round-trips 100/180/60.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
ebd9967547 |
feat(pyrowave): Windows host encoder — separate-plane zero-copy D3D11→Vulkan
Wire PyroWave into the Windows host (design/pyrowave-windows-host-zerocopy.md). Before this a macOS client + Windows host that both selected PyroWave silently ran HEVC: the host never advertised CODEC_PYROWAVE and open_video_backend bailed. Approach (zero-copy, no GPU→CPU→GPU): pyrowave owns its own Vulkan device (create_device_by_compat, by render-GPU vendor/device-id — NOT LUID, invalid in Session 0). The capturer runs a BGRA→YUV BT.709-limited CSC (matching rgb2yuv.comp) into TWO SEPARATE shareable plane textures — full-res R8 Y + half-res R8G8 CbCr — which the encoder imports into pyrowave's device. Separate single/two-component textures import reliably on NVIDIA at any size; a single planar NV12 import does NOT (the vendored interop test: "only very specific resource sizes" — confirmed on-glass: 1024² fine, 720p/1080p/1440p garbage). A shared D3D11 fence, signalled after the CSC, is imported as a Vulkan timeline semaphore so the wavelet read is ordered after it. - pf-encode: enc/windows/pyrowave.rs (Encoder impl, two-plane import + Linux-style plane views); host_wire_caps advertises CODEC_PYROWAVE on Windows when the backend isn't Software; open_video_backend routes a negotiated PyroWave session first; pyrowave-sys on the Windows target; interop confirmed at open → clean HEVC fallback. - pf-encode: shared, unit-tested enc/pyrowave_wire.rs (single source of truth for the client-facing AU framing); Linux encoder uses it too. - pf-capture: dxgi.rs BgraToYuvPlanes CSC; idd_push.rs pyrowave mode — forces the virtual display SDR (the VideoProcessor can't ingest the FP16 HDR ring), a two-plane shareable out-ring, a shared fence passed every frame (so a rebuilt encoder re-imports it). Threaded via OutputFormat::pyrowave. - pf-frame: D3d11Frame::pyro carries the CbCr plane + fence; OutputFormat::pyrowave. Verified on .173 (RTX 4090): full-host build + clippy -D warnings (nvenc,amf-qsv) + fmt --all --check; pyrowave_wire unit tests; pyrowave_win_smoke GPU test round-trips distinct Y/Cb/Cr (100/180/60) exactly at 1024²/720p/1080p/1440p; Stage-0 interop validated in the real Session-0 service context on-glass. Deployed to the box. Owed: final on-glass picture/latency confirmation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
df70bf00f9 |
test(qsv): live on-glass matrix — HEVC, Main10+HDR, AV1 10-bit, LTR-RFI, no-IDR retarget
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9e6ab9b94d |
fix(qsv): enable D3D11 multithread protection before SetHandle (on-glass -16 on Arc)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
0a617e7779 |
fix(qsv): clippy — per-block SAFETY comments, div_ceil, clamp, vec_box allow
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
75d1322b8a |
fix(qsv): end the inner borrow before the retarget param rebuild
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
55e7f3fca9 |
feat(encode): native QSV backend — libvpl-sys + qsv.rs (Phases 0-3 of design/native-qsv-encoder.md)
Vendored MIT VPL dispatcher (static, trimmed tree, pin 674d015b/v2.17.0) built via cmake+bindgen behind new feature 'qsv' (pf-encode + punktfunk-host forward). qsv.rs: dispatcher session on the capture adapter (LUID-matched), SetHandle D3D11, AsyncDepth=1/GopRefDist=1/VDEnc/CBR + HRD-off low-latency config, GetSurfaceForEncode + GPU CopySubresourceRegion input (zero-copy, no readback path), bounded sync-point poll, in-place reset with teardown escalation, no-IDR bitrate retarget (Reset + StartNewSequence=OFF), 10-bit P010 HEVC-Main10/AV1, HDR mastering/CLL SEI-OBU at IDR + BT.2020/PQ VSI, LTR-RFI via mfxExtRefListCtrl (AMF slot policy port, Query-gated per codec, wire-index FrameOrder pinning). Dispatch: native-first with ffmpeg fallback + PUNKTFUNK_QSV_FFMPEG hatch; probes (can_encode_10bit / windows_codec_support / windows_backend_is_probed) now answer natively for QSV. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9a36ea2132 |
refactor(host/W6.2): extract the video encode backends into the pf-encode crate
encode.rs + encode/* (NVENC, VAAPI, native AMF, AMF/QSV ffmpeg, direct-SDK NVENC/CUDA, raw Vulkan-Video, PyroWave, openh264) move into crates/pf-encode behind one Encoder trait + open_video selector (plan §W6). The crate speaks the shared frame vocabulary (pf-frame: CapturedFrame/PixelFormat + the DXGI identity D3d11Frame/make_device) and pf-zerocopy (CUDA context/buffers), and NEVER pf-capture — the capture→encode edge is one-way (ZeroCopyPolicy, prior commit). Dep moves: the heavy encoder deps (ffmpeg-next, the NVENC SDK, openh264, pyrowave-sys) move from the host to pf-encode; the host's nvenc/amf-qsv/vulkan-encode/pyrowave features now FORWARD to pf-encode/*. The host keeps a mod-encode shim (pub use pf_encode) so every crate::encode::* path (negotiator + GameStream/native/mgmt planes) is unchanged. resolve_render_adapter_luid moves from the host's windows/win_adapter.rs into pf-gpu (both pf-encode and pf-capture need it as a peer of GPU selection); its 5 call sites (encode amf/nvenc, capture idd_push/synthetic_nv12, vdisplay manager) rewire to pf_gpu::resolve_render_adapter_luid and win_adapter.rs is deleted. pf-frame's make_device gains a # Safety section (public-unsafe-fn lint, latent since the pf-frame carve — a full-workspace -D warnings clippy catches it). Verified: Linux clippy -D warnings (pf-encode + host nvenc,vulkan-encode,pyrowave --all-targets) + 13/13 pf-encode + 299/299 host tests; Windows clippy -D warnings (pf-encode nvenc,amf-qsv --all-targets + host nvenc,amf-qsv --all-targets) Finished exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |