3f4ad088695c65b38f6f9374b1fbc1ebbb13994d
1785
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3f4ad08869 |
test(encode/ffmpeg_win): extract the decision logic and pin it with unit tests
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
decky / build-publish (push) Successful in 20s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
deb / build-publish (push) Successful in 9m9s
docker / build-push-arm64cross (push) Successful in 10s
docker / deploy-docs (push) Successful in 24s
deb / build-publish-host (push) Successful in 10m47s
android / android (push) Successful in 14m53s
deb / build-publish-client-arm64 (push) Successful in 10m38s
windows-host / package (push) Successful in 17m40s
ci / web (push) Successful in 1m46s
ci / docs-site (push) Successful in 1m50s
arch / build-publish (push) Successful in 11m37s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 21m11s
apple / swift (push) Successful in 6m4s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 21m15s
ci / bench (push) Successful in 5m47s
ci / rust-arm64 (push) Successful in 9m11s
ci / rust (push) Successful in 21m50s
apple / screenshots (push) Successful in 25m14s
The QSV open-failure fallback (1,400 lines, 23 unsafe, 0 tests) follows the vaapi.rs treatment: the device-free decisions now live in named functions with their contracts pinned — the per-vendor zero-copy default matrix (AMF on-glass-validated on, QSV opt-in), the PUNKTFUNK_FFWIN_POLL_MS clamp-before-µs-conversion (the 27.7-hour-spin class), the readback routing table with its mid-stream depth-change guard, the swscale source map, the QSV display-remoting latency contract (async_depth=1/low_power=1/ look_ahead=0/forced_idr=1/scenario) and the AMF no-B-frames contract, and the per-vendor zero-copy pool bind flags. A probe smoke rides along #[ignore]d for the runner. No FFI plumbing chased; no behavior changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fc335b39e9 |
fix(host/encode): negotiate the cursor around what the encoder can blend
EncoderCaps::blends_cursor's contract said the HOST must fall back to capturer-side compositing when a cursor-as-metadata session lands on an encoder that can't composite — but that host half was never built: open_video warned and the session streamed WITHOUT a pointer (confirmed on the VAAPI dmabuf and libav-NVENC CUDA paths; latent on vulkan RGB-direct/native-NV12). The negotiation is now caps-aware, ahead of capture, on both planes: * pf-encode grows cursor_blend_capable() — the pre-open dispatch mirror (sibling of linux_native_nv12_ok) answering whether the resolved backend composites frame.cursor; its pure core is test-pinned arm by arm. * Native plane: handshake::cursor_forward grants the cursor channel only where the resolved backend can blend (the capture-mouse flip makes the host draw the pointer on demand); denied sessions keep the pre-channel path — the compositor EMBEDS the pointer, never cursorless, never doubled. The Welcome's HOST_CAP_CURSOR bit is computed once and read back at both session-wiring sites instead of recomputed. SessionPlan::output_format additionally keeps every cursor-blend session off producer-native NV12 (the arm with no CSC to fold a cursor into), and vulkan RGB-direct now yields to a cursor-blend session even when pinned (EFC cannot composite; the open logs the override). Windows plans cursor_blend=false via the new shared cursor_blend_for() rule — the IDD capturer composites the pointer itself, and asking the encoder anyway fired the blends-cursor warn spuriously on every cursor-channel session. * GameStream plane: the hardcoded cursor_blend=true is gone. The portal source asks for cursor-as-metadata only when the resolved backend blends, otherwise negotiates an Embedded pointer (choose_cursor_mode's new ladder); the capturer pool now also keys on that mode. The virtual-output source passes false — its capture embeds the pointer where it can. The per-arm warns in vulkan_video (RGB-direct, native-NV12) are now structurally unreachable and removed. open_video's post-open check stays as the single backstop for what planning cannot see: a Vulkan-open falling back to VAAPI mid-session, and the gamescope residual (no embedded mode exists there, so a never-blending backend — H.264-on-AMD VAAPI, software — still streams cursorless; fixing that needs a compositing stage, deliberately not built in this pass). Zero-copy is preserved throughout — every fallback is a capture-negotiation change, never a readback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f495b201e1 |
fix(encode): delete the write-only EncoderCaps::supports_hdr_metadata
A caps field nothing reads is a contract nobody honors — and this one shipped write-only: its single reader anywhere in the workspace was a hardware-gated assertion inside pf-encode's own AMF smoke test. Both planes send the static HDR grade out-of-band unconditionally (the native 0xCE datagram per keyframe, the GameStream 0x010e control message), every first-party client reads exclusively that path, and none parse in-band SEI — so the host decision the field was reserved for (suppress out-of-band when the encoder embeds) can never validly exist. The field's doc contract had also rotted in two directions: it claimed set_hdr_meta no-ops when false (native AMF and QSV consume it regardless) and that only Windows direct-NVENC attaches in-band metadata (AMF and QSV do too). The in-band SEI/OBU emission itself is untouched — it stays a bonus for stock decoders, documented at the emit sites; the trait docs now describe the real routing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
232b6d6be2 |
test(encode/vaapi): extract the open-time decision logic and pin it with unit tests
The fallback backend under Vulkan Video — and the only AMD/Intel H.264 and 10-bit/HDR encoder — had 1,300 lines, 26 unsafe blocks, and zero tests. The device-free decisions now live in named functions with their contracts pinned: the entrypoint ladder + LP_MODE latch round-trip (the cross-GPU session-killer and the 8-bit-pins-10-bit under-advertisement are both key'd tests now), the PUNKTFUNK_VAAPI_LOW_POWER / _ASYNC_DEPTH grammars, the VUI ↔ scale_vaapi colour agreement (the Mesa-BT.601 hue-shift pin), the honest-downgrade depth table, the HEVC-Main10-only explicit profile, and the 10-bit probe gate. Probe + CPU-path encode round-trip ride along as #[ignore]d hardware smokes in the house style. No FFI plumbing was chased; no behavior changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cc848479c4 |
test(pf-encode): re-point the cpu_img size-change smoke at the CSC guard's contract
arch / build-publish (push) Failing after 58s
ci / web (push) Successful in 51s
ci / docs-site (push) Successful in 53s
android / android (push) Failing after 7m59s
ci / bench (push) Failing after 4m9s
ci / rust-arm64 (push) Failing after 6m44s
ci / rust (push) Failing after 6m45s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 24s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 12s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 7s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 47s
apple / swift (push) Successful in 5m28s
deb / build-publish (push) Successful in 9m25s
windows-host / package (push) Successful in 10m44s
docker / deploy-docs (push) Successful in 31s
docker / build-push-arm64cross (push) Successful in 5m33s
deb / build-publish-client-arm64 (push) Successful in 8m20s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 7m41s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 7m4s
deb / build-publish-host (push) Successful in 12m7s
apple / screenshots (push) Successful in 25m11s
vulkan_cpu_img_survives_a_source_size_change drove MISMATCHED source sizes through the then-lenient CSC arm as its vehicle for the staging cache hazard (format-only-keyed cpu_img → OOB copy while submit said Ok). e3354b6d's guard — correctly the equality check every sibling arm always had, against the MODE, not the coded extent, so the padded render-vs-coded tolerance in the direct arms is untouched — makes that scenario unrepresentable through submit and broke the test on main (.25 layers baseline read 13/13 instead of 14/12). Replaced by vulkan_csc_refuses_a_mismatched_source: refusal pinned in BOTH directions, plus the property that actually needs proving — a refused submit does not WEDGE the session (the bail lands after step 1's frame-type bookkeeping; the next well-sized frame must still encode, and an AU must come out). Verified on the 780M under validation layers: 8/8 vulkan tests, full-suite baseline restored to 14/12. The WP4.2 size-keyed staging stays as belt-and-braces; the hazard it fixed is now structurally unreachable through submit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bf9386ecb2 |
test(pf-encode): pin the typed-EINVAL classifier's chain-survival contract (Phase 8)
ci / docs-site (push) Failing after 52s
android / android (push) Failing after 6m28s
ci / web (push) Successful in 1m5s
ci / bench (push) Successful in 6m40s
decky / build-publish (push) Successful in 26s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 16s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
arch / build-publish (push) Successful in 12m51s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
ci / rust-arm64 (push) Successful in 9m51s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Failing after 42s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 17s
deb / build-publish-client-arm64 (push) Successful in 8m54s
deb / build-publish-host (push) Successful in 10m31s
deb / build-publish (push) Successful in 11m17s
windows-host / package (push) Successful in 10m41s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m25s
docker / build-push-arm64cross (push) Skipped
docker / deploy-docs (push) Skipped
apple / swift (push) Successful in 5m25s
ci / rust (push) Successful in 22m21s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m7s
apple / screenshots (push) Successful in 23m50s
Rides on
|
||
|
|
bf9fb3fb22 |
fix(pf-encode): enable VK_EXT_queue_family_foreign for the dmabuf acquires (Phase 8)
Both Linux Vulkan encode backends named QUEUE_FAMILY_FOREIGN_EXT as the acquire barriers' src family without ever enabling the extension — spec-invalid on every device, tolerated by RADV. The audit filed vulkan_video's three sites; pyrowave's fresh-import acquire had the identical defect on its own device (critic catch). Enable when advertised (a fresh open-time enumerate — the rgb probe's is a probe-local and skipped entirely on native-NV12, so there was nothing to reuse; pf-presenter/dmabuf.rs is the in-repo precedent that already enables this extension). Not advertised → the core-1.1 QUEUE_FAMILY_EXTERNAL conservative substitute, chosen once at open and warn-logged (no fleet hardware takes that arm; such devices were never valid targets before). All four sites are acquire-only (src=FOREIGN, EXCLUSIVE images, oldLayout=UNDEFINED) — the swap is index-only. On-glass: 780M under validation layers — vulkan smokes + pyrowave smokes green, FOREIGN advertised and enabled, no fallback engaged. Shared ext_advertised helper in vk_util (cfg = the union of both consumers) with a unit test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1cefd37603 |
fix(pf-encode): arbitrate NVENC split-encode vs sub-frame readback (Phase 8)
Verified against nvEncodeAPI.h's own splitEncodeMode doc (user-prompted — the audit's 'exclusive for HEVC' one-liner deserved checking): - H.264: split 'is not applicable' — hard-DISABLE the mode so the written config, CeilingKey, the split diagnostic log and the rejection-retry stay truthful (the retry used to re-open a byte-identical session after an H.264 'split rejection'). The libav path's operator arm gains the codec gate its auto arm always had. - HEVC: split 'not supported if … subframe mode' — when WE force split (TWO/THREE/AUTO_FORCED, the 4K120 throughput lever), sub-frame yields with a logged escape (PUNKTFUNK_SPLIT_ENCODE=0 chooses sub-frame). ⚠ Keyed on FORCED modes only, never != DISABLE: AUTO(0) is the resolver's fallthrough for every sub-950Mpix session, and the wider key would have disarmed the Phase-3 chunked-poll feature fleet-wide (critic catch). Under AUTO the driver arbitrates — the shipped state. - AV1: untouched — per-tile sub-frame + split are legal together. The arbitration is a pure nvenc_core fn called by each backend BEFORE the ladder, the ceiling key and the chunked-poll latch — all three see the post-arbitration truth. A drop inside build_init_params would have left poll_chunk busy-polling its whole budget every AU (numSlices stays 0 without reportSliceOffsets; both loop exits dead — critic catch). Linux latches subframe_forced beside subframe_on at query_caps (no env re-reads after open); Windows records the arbitrated state so reconfigure presents exactly the params the open had (also closes the pre-existing mid-session env-flip hazard there). Truth-table tests in nvenc_core; PUNKTFUNK_NVENC_SUBFRAME documented (it never was); PUNKTFUNK_SPLIT_ENCODE row updated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
25765c53ec |
fix(encode): classify libav-NVENC open failures by errno, not English strerror text
ci / docs-site (push) Successful in 58s
ci / web (push) Successful in 59s
apple / swift (push) Successful in 5m17s
ci / bench (push) Successful in 7m27s
deb / build-publish (push) Failing after 8m0s
ci / rust-arm64 (push) Failing after 9m1s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 22s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
ci / rust (push) Failing after 9m31s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
deb / build-publish-host (push) Successful in 10m23s
arch / build-publish (push) Successful in 12m54s
android / android (push) Successful in 15m8s
deb / build-publish-client-arm64 (push) Successful in 7m54s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 5m43s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 6m35s
windows-host / package (push) Successful in 19m10s
docker / build-push-arm64cross (push) Successful in 8s
docker / deploy-docs (push) Successful in 24s
apple / screenshots (push) Successful in 23m39s
The bitrate-probe ladder stepped down on format!("{e:#}").contains(
"Invalid argument") — an English substring over the WHOLE context chain,
which also fired on any other wrapped EINVAL (e.g. a CUDA-context errno)
and gated a ~10-step ladder on strerror wording. The root ffmpeg::Error
survives the anyhow chain; downcast and match Error::Other{errno:EINVAL}
instead. Same fix for the intra-refresh ENOSYS probe in the open path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
36589c39da |
fix(gamestream): stall-recovery ladder — a wedged encoder no longer ends the stream
The backends deliberately turn a wedged GPU into a bounded error so the caller can reset in place, but only the native plane ever did: on the GameStream path every submit/poll error propagated straight out of the video loop, costing Moonlight clients a full disconnect/reconnect. Port the native ladder (bounded resets + backoff + silent-wedge watchdog, same field-lesson backoff curve). Also keep the drop-recovery keyframe armed until it actually EMITS: it was consumed on read, so the coalesce gate could swallow it for good — leaving duplicate wire indices in the encoder's reference table for a later RFI to anchor on, exactly the stale-anchor case rfi.rs exists to prevent, and it opened only under congestion + loss, which is when RFI fires. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e3354b6d5d |
fix(encode/vulkan): guard the CSC source dimensions, and bound reset()'s wait
The CSC path was the only backend arm that took frame.width/height on trust: the shader samples with clamped 1:1 texelFetch, so a mismatched frame silently streamed a cropped/edge-padded picture where every sibling errors into the encoder-rebuild path. The import cache now also carries the extent it imported at (a (st_dev, st_ino) hit alone doesn't prove the allocation still matches) and is dropped on reset(). reset() opened with an untimed device_wait_idle on the one thread whose every other wait is capped at ENCODE_FENCE_TIMEOUT_NS for exactly this reason — reset() runs BECAUSE the GPU looks wedged. Both vulkan-video and pyrowave now bound the wait and report "no in-place rebuild" on timeout instead of parking recovery on the suspect device; Drop keeps the unbounded wait (teardown must stay memory-safe against a wedged device). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
28f8fc71c4 |
refactor(pf-encode): split vulkan_video's construction tail into vk_build.rs (WP7.5)
ci / web (push) Successful in 56s
ci / rust-arm64 (push) Failing after 2m18s
ci / docs-site (push) Successful in 1m2s
android / android (push) Failing after 4m7s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 9s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
apple / swift (push) Successful in 5m21s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 6m10s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 5m40s
deb / build-publish-client-arm64 (push) Successful in 7m57s
docker / build-push-arm64cross (push) Successful in 10s
deb / build-publish (push) Successful in 10m6s
docker / deploy-docs (push) Successful in 24s
deb / build-publish-host (push) Successful in 10m30s
arch / build-publish (push) Successful in 13m57s
windows-host / package (push) Successful in 14m6s
ci / rust (push) Successful in 22m9s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m30s
apple / screenshots (push) Successful in 24m30s
The ~820-line tail of free builders — make_frame/make_frame_csc/ make_frame_common, make_video_image, probe_rgb_direct, the H.265/AV1 parameter-set writers and the AV1 bit-writer — moves to a #[path] child module, the amf_sys.rs shape: the child sees the parent's private items (Frame and friends), so the split costs zero visibility churn. Six parent-called items went pub(super); five stay child-private (dead_code is per-item and each is used within the child). vulkan_video.rs drops 5,292 → 4,489 lines and the construction unsafe gets its own review surface; steady-state encode logic stays in the parent. ⚠ Trap recorded for future child-module splits: inline `use super::X` statements INSIDE moved fn bodies silently change meaning (super shifts one level) — vk_av1_encode/vk_valve_rgb imports needed crate:: paths. Proven on-glass, not just compiled: all 8 vulkan GPU smokes green under the validation layers on the 780M post-split (H.265 + AV1, RGB-direct, CSC, CPU paths — every moved constructor exercised). nvenc_cuda.rs and qsv.rs are DECLINED the same treatment, with evidence in the handoff doc: no equivalent self-contained seam — their candidate regions are ~150-line loader/accessor clusters interleaved with the encoders' own state types, and a thin-forwarder impl split is exactly the churn a no-defect phase penalizes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bc5ec5105c |
refactor(pf-encode): one Linux backend resolver, consumed by dispatch AND mirrors (WP7.6)
The Linux backend decision existed as open_video_backend's string match plus five partial hand-copies (the zero-copy-plane gate, the pyro advertisement gate, the software advertisement pin, the vulkan pref ceiling, resolved_backend_is_gpu). Windows solved this long ago — windows_resolved_backend() is consumed by its dispatch AND its mirrors, with labels still stamped at the open sites. Linux now has the twin: resolve_linux_backend (pure, lazy auto probe) + linux_resolved_backend (config wrapper, unknown→auto exactly as every mirror's old `_` arm). ⚠ This deliberately DEVIATES from the audit's shadow-assertion prescription, on both critics' findings: the shadow had no execution venue (open_video had ZERO test call sites; shipped hosts are --release) — unfalsifiable ceremony — and the file's own Windows half proves dispatch-consumes-resolver is safe: the mgmt record is protected by the label-from-the-open-site convention, not by resolver avoidance. One consumed table beats two tables plus an inert cross-check. Preservation riders from the critique, all applied: - pref_ceiling KEEPS its cfg!(vulkan-encode) branch (the resolver is feature-blind; dropping it re-creates advertise-then-die-at-open). - linux_zero_copy_is_vaapi's Vulkan|Software arm preserves the old `_` fallthrough EXACTLY — the vulkan-on-NVIDIA capture-plane mismatch and the software twin are FILED in the design doc, not fixed here. - resolved_backend_is_gpu splits as linux + not(any(windows, linux)) — never a target list, so no exotic target loses the fn. - The auto probe is lazy (impl FnOnce) — explicit prefs stay zero-probe (/serverinfo polls through these mirrors), pinned by a panicking closure in the resolver test. New coverage: the full alias table pinned, and a GPU-free dispatch test through the real open_video_backend (software arm, vendored openh264) — its first test call site anywhere. Config-latch seam guarded loudly. Cross-crate mirrors recorded out of scope in the design doc: session_plan::resolve_encoder, gamestream/serverinfo.rs base_codec_mode_support, capture.rs's "pyrowave" string. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cb69cd3b8c |
refactor(pf-encode): move the AMF C-ABI mirror to its own file (WP7.4)
The ~410-line hand-mirrored AMF vtable ABI was already isolated in an inline `mod sys` — the remaining value of the audit's 'best split in the crate' is the FILE boundary: amf.rs drops 3379 → 2965 lines and the pure unsafe-FFI surface (25 unsafe fns, zero policy) is reviewable in isolation, which is the crate's stated review goal. Pure move: `#[path = "amf_sys.rs"] mod sys;` keeps the module name and every `sys::` call site byte-identical; the module was self-contained (only `use std::ffi::c_void`, no super:: references). The banner became the file's module docs; contents de-indented one level — rustfmt-clean on the first check, which is what proves the move byte-exact. Gated on .173: all five clippy combos + the amf-qsv,qsv test leg (38 passed — the live AMF matrix on real VCN silicon among them). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9491cc8759 |
refactor(pf-encode): extract the range-family RFI recovery policy (WP7.2)
The two direct-NVENC backends carried hand-copied twins of the same loss-recovery decision: range validity, covering-range dedup, DPB window, clamp — ~30 duplicated lines each. The decision now lives once as nvenc_core::plan_range_recovery (the range half of WP7.2; the slot half is enc/rfi.rs), pure and unit-tested; each backend keeps its session gate, its unsafe per-timestamp driver loop, and its state stores. The step order is load-bearing and now pinned by tests: the covering dedup runs with the UNCLAMPED last and BEFORE the DPB window (a covered re-ask never touches the driver even when the range has since aged out of the DPB), the boundary at next_ts - RFI_DPB is inclusive, and the Invalidate carries the CLAMPED last — which is also what the caller records in last_rfi_range, exactly as the inline code stored it. A driver failure mid-loop still returns false with NO range recorded and no anchor armed. Decline deliberately clears nothing (neither twin touched pending_anchor on decline — same shape as Vulkan's non-clear, opposite of AMF/QSV; do not harmonize). The exact-cover → Covered test records EXISTING behavior including that a covered range survives a forced IDR with zero driver calls — a recorded fact, not an endorsement. RFI_DPB's import leaves both twins: its only per-backend use was the arithmetic that moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9f1e648e4e |
refactor(pf-encode): extract the slot-family RFI recovery policy (WP7.2)
AMF (user-LTR bitfield), QSV (mfxExtRefListCtrl) and Vulkan Video (the
app-owned DPB slot table) each hand-implemented the same loss-recovery
decision: distrust every reference encoded at-or-after the loss start,
anchor on the newest one strictly older. Three copies had already
diverged once — the
|
||
|
|
d28fb1282b |
test(pf-encode): guard TrackedEncoder's forwarding completeness (WP7.7, cheap half)
A defaulted Encoder method that TrackedEncoder doesn't forward silently no-ops through the wrapper — the host loop only ever holds the wrapped box, so the feature dies for every session with nothing in the logs. The trap has bitten three times (set_wire_chunking's §4.4 chunking probe, set_pipelined's LN3 escalation, applied_bitrate_bps's ABR truth), and every Phase 7 consolidation that adds a trait method re-arms it. Source-text parse of the trait and the forwarding impl (both top-level rustfmt items: block ends at the first column-0 brace, method names sit on 'fn '-prefixed lines), then set equality — the reverse direction is already a compile error, so equality == completeness. Mutation-verified: removing set_wire_chunking's forward fails naming exactly that method. 16/16 today. Limit stated in the test doc: name-set equality only — a forward whose body delegates to the WRONG inner method still passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
94b818da28 |
fix(client/pyrowave): declare the decode planes' real format — 10-bit sessions decoded through an R8 view of R16 planes
apple / swift (push) Successful in 5m14s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 6m32s
arch / build-publish (push) Failing after 7m6s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
decky / build-publish (push) Successful in 18s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 8s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m53s
android / android (push) Successful in 12m56s
docker / deploy-docs (push) Successful in 24s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 5m41s
docker / build-push-arm64cross (push) Successful in 4m15s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 5m40s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 7m44s
deb / build-publish-host (push) Successful in 10m43s
deb / build-publish-client-arm64 (push) Successful in 9m8s
apple / screenshots (push) Successful in 22m58s
flatpak / build-publish (push) Successful in 6m56s
deb / build-publish (push) Successful in 8m58s
ci / rust (push) Successful in 22m10s
ci / rust-arm64 (push) Successful in 9m22s
ci / web (push) Successful in 48s
ci / docs-site (push) Successful in 1m13s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 12m48s
ci / bench (push) Successful in 6m3s
The HDR/4:4:4 leg (
|
||
|
|
3e7828522d |
fix(tray): drop the conflicting-host warning — the tray is the wrong place for it
ci / web (push) Successful in 59s
ci / docs-site (push) Successful in 1m6s
ci / bench (push) Successful in 7m0s
deb / build-publish (push) Failing after 7m35s
decky / build-publish (push) Successful in 22s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
deb / build-publish-host (push) Failing after 8m6s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 8s
android / android (push) Successful in 13m38s
ci / rust-arm64 (push) Successful in 12m28s
docker / deploy-docs (push) Successful in 24s
arch / build-publish (push) Successful in 14m10s
windows-host / package (push) Successful in 10m56s
deb / build-publish-client-arm64 (push) Successful in 8m36s
docker / build-push-arm64cross (push) Successful in 4m37s
apple / swift (push) Successful in 5m28s
ci / rust (push) Successful in 22m58s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 13m9s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m0s
apple / screenshots (push) Successful in 23m51s
The tray led its tooltip with "⚠ conflicting host: Sunshine" and (on Linux) went NeedsAttention for a Sunshine that was merely *installed* — never started, never listening, harmless. An always-on warning over a dormant package is a false alarm on every poll, and the tray is the wrong surface for a one-shot install-time observation anyway. Removes the `conflicts` field from the tray's `Summary`, the tooltip prefix, `has_conflicts()`, and the ksni attention branch that fed on it. The host still detects and reports conflicts where that belongs — the startup `punktfunk::detect` warning, the `detect-conflicts` subcommand, and the `/api/v1/local/summary` field, which the tray must (and does, per the new test) keep tolerating on the wire. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ffc7aec91a |
feat(encode/pyrowave): log which GPU pyrowave picked — the selection stays put, by decision
ci / web (push) Successful in 55s
ci / docs-site (push) Successful in 1m12s
apple / swift (push) Successful in 5m20s
ci / rust (push) Failing after 5m55s
android / android (push) Failing after 5m59s
deb / build-publish (push) Failing after 5m1s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
decky / build-publish (push) Successful in 17s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
ci / bench (push) Successful in 7m17s
ci / rust-arm64 (push) Successful in 9m40s
docker / deploy-docs (push) Successful in 26s
docker / build-push-arm64cross (push) Successful in 4m9s
deb / build-publish-host (push) Successful in 10m32s
arch / build-publish (push) Failing after 12m34s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 5m58s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 5m50s
deb / build-publish-client-arm64 (push) Failing after 7m10s
windows-host / package (push) Successful in 17m27s
apple / screenshots (push) Successful in 23m26s
WP4.5's device-selection half, closed as the observability intermediate after TWO selection designs died in adversarial review: - Attempt #1 (d26bcf05, withdrawn): match pf_gpu::selected_gpu(). Its Linux auto arm answers "the NVIDIA GPU" whenever /dev/nvidiactl exists, moving the encoder off the iGPU that can import the compositor's dmabufs on an Intel-compositor + NVIDIA-present laptop — import failures feed the process-wide raw-dmabuf latch, which never un-latches. - Attempt #2 (this session, withdrawn before commit): anchor on the PUNKTFUNK_RENDER_NODE-else-renderD128 node via VK_EXT_physical_device_drm. Render minors are driver-BIND-ORDER artifacts, not display topology: on the common AMD-iGPU + NVIDIA-display desktop, in-tree amdgpu binds before out-of-tree nvidia, so the anchor deterministically picks the idle iGPU while the compositor allocates on NVIDIA — the same latch, opposite polarity, behind a success-looking log. The correct oracle is evidence of which device ALLOCATED the capture buffers — producer identity from the capture negotiation, threaded per session into this open. Until that plumbing exists, selection stays first-usable, both call sites still share one selector (pure over the device list, so capture_modifiers and open_inner cannot diverge — including across an in-place resize's re-open, which does not renegotiate capture), and the open logs ONE greppable line: picked vendor/device, the anchor node and its owner (DRM render major/minor, VK_EXT_pci_bus_info fallback), and the console's selected GPU. A wrong-device session on a multi-GPU host used to be completely invisible; a field report can now show it. No WARN arm on purpose: the wrong-pick direction inverts between the laptop and desktop topologies, so a mismatch is not evidence of a wrong pick, and a warning that fires forever on healthy hosts teaches people to ignore warnings. Decision recorded against the audit's framing: manual console GPU selection stays unhonored by pyrowave on Linux (the Windows twin honors it) — honoring console-mutable state without per-session threading is what made attempt #1 unsafe. Verified on the 780M: the line resolves all three identities (1002:15bf x3, DRM-props match live); full pyrowave on-glass suite green; selection behaviour byte-for-byte unchanged. WP4.5 (device half). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
46935bf0a5 |
fix(encode/vulkan): VBR instead of CBR — the driver was stuffing ~98% filler into every calm stream
Second attempt at WP6.3; the first (ce543668) was withdrawn after its tight CBR window measured as a 36x bandwidth regression (97% filler NALs). The correction that unlocked this one: that measurement's BASELINE row was an 8-frame artifact. Under the shipped 1000ms/500ms CBR window a calm stream overflows the CPB once the initial fill drains (~30 frames at 10Mbps/60fps) and RADV then pads every frame to the exact rate share, forever. Measured on the 780M (1280x720@60, 10Mbps, calm content): 64 frames = 97.5% filler; 300 frames = 5.63MB at 98.5% filler where this commit ships 83KB at 0%. AV1: 99.6% -> 0%. The status quo was ~the full target bitrate of zeros on every idle AMD/Intel Vulkan-encode desktop and Steam Deck — and stuffing to exactly the target permanently satisfied the ABR calm brake (actual >= 3/4 * current), the ratchet WP6.3's withdrawal feared from the tight window, live in the shipped code all along. The fix reads VkVideoEncodeCapabilitiesKHR::rateControlModes (previously ignored — rateControlMode was hardcoded CBR with no capability check) and installs VBR with average == max plus the house ~1-frame window (vbv_window_ms, PUNKTFUNK_VBV_FRAMES-scaled) when the driver advertises VBR. VBR permits underspend, the exact missing degree of freedom: Vulkan exposes no filler-suppression control (AMF's filler_data=false / NVENC's default-off have no VK equivalent), so the MODE is the only lever. CBR-only drivers keep the loose window untouched — tightening it under CBR just starts the stuffing 30 frames earlier. Drivers advertising neither mode (ANV per current Mesa) keep the pre-existing CBR install, now WARN-logged. No pacing claim, deliberately: burst A/B on the 780M is byte-identical between 1000ms CBR and 17ms VBR (max AU 1.19MB in both) — this firmware ignores the window for QP decisions entirely. The payload is filler elimination. PUNKTFUNK_VULKAN_RC=cbr|vbr is the field escape hatch and the on-box A/B control (two withdrawn attempts bought that insurance). Also on the same caps struct: maxBitrate is now read and clamps open + retarget (RADV reports 1 Gbps — within 5% of the 4K120 ABR targets), and applied_bitrate_bps() reports the encoder-side truth (pending-first, so the session loop's read right after reconfigure_bitrate sees the clamp) — without it a binding clamp would feed the ABR a phantom base, the trap the trait doc names. And the one-frame VUID-vkCmdBeginVideoCodingKHR-pBeginInfo-08254 violation found in the withdrawal review: record_submit promoted a pending retarget into self.bitrate BEFORE recording whenever first_frame was set, so after a mid-stream reset() (which preserves the pending rate and rc_installed) the begin-coding declaration named a rate the session had not installed — and the two triggers, ABR retarget and the stall watchdog, correlate. Now the declaration always names the session's current rate and the RESET install carries the pending one via its own struct; promotion stays in post_submit_bookkeeping. The extended validation-layer test reproduces the retarget-then-reset coincidence: exactly one 08254 on the pre-fix build, zero on this one (RADV PHOENIX, on glass). Gates: docker amd64 legs green; Windows .173 seven legs (34 passed); .25 full-suite parity vs origin/main (identical CUDA-only failures) + all 9 vulkan on-glass tests + validation layers clean. WP6.3, plus WP7.1's ms-form half (vbv_window_ms). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fdded5b8c3 |
fix(encode/pyrowave): refuse a frame that isn't the session's mode
ci / web (push) Successful in 54s
ci / docs-site (push) Successful in 55s
ci / bench (push) Failing after 6m37s
ci / rust (push) Failing after 6m43s
deb / build-publish-host (push) Failing after 5m44s
arch / build-publish (push) Failing after 6m45s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 12s
decky / build-publish (push) Successful in 18s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 7s
docker / build-push-arm64cross (push) Successful in 9s
docker / deploy-docs (push) Successful in 23s
ci / rust-arm64 (push) Successful in 9m47s
deb / build-publish (push) Successful in 9m9s
android / android (push) Successful in 11m59s
deb / build-publish-client-arm64 (push) Successful in 8m44s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 8m27s
windows-host / package (push) Successful in 17m35s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m58s
apple / swift (push) Successful in 6m51s
apple / screenshots (push) Successful in 26m43s
PyroWave never checked frame dimensions against the session, and it applies no alignment — `width`/`height` are the negotiated mode verbatim — so a mismatched frame was encoded edge-smeared or cropped, silently, forever. Every other Linux backend already refuses exactly this, with this shape, in `submit`: libav-NVENC (`linux/mod.rs`), VAAPI (`vaapi.rs`) and openh264 (`sw.rs`) all carry the same `ensure!`. PyroWave was the only one that didn't. That is the justification; an earlier draft cited `vulkan_video.rs` instead, which is the weaker precedent — it bails in its Dmabuf arms only, and its CSC path and CPU arm have no dimension guard at all. Mostly this is a wrong-picture bug and not a memory-safety one: `rgb2yuv.comp` clamps every fetch with `min(p, textureSize - 1)` and the CPU arm uploads `min(len, need)` into a session-sized image. But it also closes a narrow real hazard that was not in the filing: `import_cached` keys on `(st_dev, st_ino)` and returns the cached `VkImage` on a hit WITHOUT rechecking the extent, and unlike the capture side it is never cleared on a renegotiation — so a dmabuf inode recycled across a shrinking renegotiation would hand the encoder an image sized for the old, larger allocation. This check closes that route. ⚠ Recorded at the code because it changes the failure mode, not just the detection: a mismatch is NOT always transient. A compositor-initiated PipeWire renegotiation updates the capturer's size in place and signals nothing the encode loop reads, so it can be a permanent new steady state — and `reset()` reopens at the same dimensions by construction, so the host's five-reset budget cannot recover (~3.1 s of frozen stream, then the session ends) where before it would have streamed on with a wrong picture. For a real mode change that trade is clearly right; for a 16-row KWin mismatch it is not, and the proper fix is for the host to classify this error as a PIPELINE rebuild rather than an encoder reset. Filed, not done here. The device-selection half of WP4.5 was written, reviewed and WITHDRAWN — see the handoff doc. Matching the selected render GPU regresses the hybrid Intel-compositor + NVIDIA-present topology this project has a live field report for, because `selected_gpu()` answers NVIDIA whenever `/dev/nvidiactl` exists regardless of where capture actually runs. WP4.5 (dimension half). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9fe9cbbf07 |
perf(encode/vulkan): build the padded RGB frame in the staging memory, not beside it
The RGB-direct CPU-upload path allocated and zero-filled a whole padded frame on every submit, filled it row by row, then memcpy'd the whole thing into the mapped staging buffer. The zero-fill was entirely dead: the row loop writes every byte of every row, and rows past the source re-copy the last source row. So the frame paid for an allocation, a page-fault storm over fresh pages, a full zero-fill, and two full-frame copies where one would do. Map first, write the padded rows straight into the mapping. Nothing reads back from the destination — the row source is always the caller's buffer, never the staging memory — so writing into (write-combined) host memory costs nothing extra, and `make_host_buffer` allocates HOST_COHERENT, so the writes need no flush before the transfer reads them. Measured on real RDNA3 780M silicon, release build, `submit` alone, 400 frames x 3 rounds: - **1920x1080** (rows padded to 1088): p50 2315 -> **1725 us** (-25%), p99 2819 -> **1972 us** (-30%), and the spread (p99-min) tightens 739 -> 483 us. - **1366x768 -> 1408x768** (BOTH axes padded, so the column tail loop runs): p50 1013 -> **926 us** (-8%), p99 1303 -> **1110 us** (-15%). This is the one case where the new code could have been slower — 4-byte stores straight into write-combined memory — and it is not. - **CONTROL, 1280x720** (64x16-aligned, so the branch is never entered): p50 692 vs 692 us. No delta, which is what makes the two above attributable to this change rather than to anything else on the branch. The branch is reached by any RGB-direct session at a mode that is not 64x16 aligned, whenever capture delivers CPU frames (the default Linux capture path is dmabuf and never enters it). 1080p qualifies, since 1080 aligns to 1088. ⚠ An earlier draft of this message claimed "33 MB at 4K" — that is wrong: 3840 and 2160 are both already aligned, so 4K UHD never enters this branch at all. The large case is an ultrawide like 3440x1440 -> 3456x1440, ~20 MB. Three things this deliberately does NOT do: - The extent guards stay scoped to the `pad` branch. The CSC path deliberately supports a source SMALLER than the encode extent — its shader clamps at sample time — so a guard hoisted above the branch would break it. - Every fallible step stays above the map. An error raised between `map_memory` and `unmap_memory` would strand the mapping for the life of the slot's staging buffer, and the next frame's `map_memory` on it then violates VUID-vkMapMemory-memory-00678. (The filed "two exits above the map leak Vulkan objects" hazard is separately already gone: `47a23bec` moved that unwind into `make_host_buffer`.) - It adds guards rather than removing them: a zero source axis made `sh - 1` underflow, and "cannot fail after the map" has to be true by construction. Three corrections to an earlier draft, all found by review of that draft: - The `dw*dh*4 == need` precondition was a `debug_assert!` placed BELOW the map. That is wrong twice: it made the only check on the slice length vanish from the builds that ship, and a fired assert would have unwound past `unmap_memory` — the exact failure the bullet above says was designed out. It is now a real checked `?` above the map. - `need` was `(iw * ih * 4) as u64`: a u32 multiply widened after the fact, which agreed with the usize slice length only up to ~32768x32768. The old code's `min(need)` was a hard backstop against exactly that and the rewrite dropped it. Now widened before the multiply, matching `read_slot`'s existing discipline. - The extent guard now runs BEFORE the payload-length guard, so the usize `sw * sh * 4` cannot overflow on a garbage frame header. WP6.2(a). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e680096c6a |
perf(encode): stop re-reading the environment on every submit and poll
`std::env::var` was on three per-frame paths. Measured on `.173` (Windows, 57 environment variables, 2M iterations): **121.9 ns** per call for the NVENC in-flight cap and **114.9 ns** per call for the ffmpeg poll spin, against **0.9 ns** once memoized. On Linux, 32 ns → 1.5 ns. ⚠ Those numbers deflate the filing, and that is worth recording: the audit ranked this as a hot-path defect, but ~120 ns/frame is ~0.003% of a frame budget. The fix is still right — it is free, and it takes a global environment lock off the encode thread — but nobody should schedule it ahead of anything on the strength of the "hot path" framing. The severity RANKING was also inverted, and the measurement confirms why. The site the audit called worst — `nvenc_cuda`'s backpressure loop condition — costs a default session nothing, because the condition short-circuits on `async_rt.is_some()` and the default session never engages the two-thread retrieve. The one that actually pays every frame is Windows `submit`, which consults `async_inflight_cap()` in BOTH arms of the ring-depth match, sync mode included, where the result is then thrown away. Windows `poll` is the second unconditional one, and the audit ranked it last. Memoized inside each helper rather than latched into a session field. Nothing in the workspace mutates these variables at runtime (enumerated: no `set_var` for either key anywhere in first-party code, and the Windows service's arbitrary-key `host.env` loader runs in the SCM supervisor, which re-execs the host as a child and never opens an encoder itself). A field would instead change WHEN the value is read — and the Windows `cap` composes the env with `input_ring_depth`, which `set_input_ring_depth` may change after open, so freezing that half would reintroduce the in-place-overwrite bug the ring term exists to prevent. Two deliberate behaviour changes ride along on `PUNKTFUNK_FFWIN_POLL_MS`, so "behaviour-preserving" describes the memoization only, not this whole commit: - **A 1000 ms ceiling.** The reachable hazard was never the overflow — that needed `ms >= 1.8e16` — it was a slipped digit: `=100000000` was a 27.7-hour spin of the encode thread. - **`.trim()`, now on all three parsers.** An earlier draft applied the house rule (WP7.8) to one of the three, which left `PUNKTFUNK_NVENC_ASYNC=" 1 "` working while `PUNKTFUNK_NVENC_ASYNC_DEPTH=" 6 "` silently fell back to 4 — and memoization would have frozen that silent fallback for the process lifetime. Reachable: the Windows `host.env` loader trims around `=` before stripping quotes, so `=" 2 "` yields a value with inner spaces. ⚠ Correction to an earlier draft of this message, which asserted that the audit's `saturating_mul` proposal would relocate an overflow panic into release builds. That is FALSE and the code comment now says so: `Duration::from_micros(u64::MAX)` is ~1.8e13 seconds, six orders of magnitude below `Duration`'s ceiling, so `Instant + Duration` neither overflows nor panics (measured, with and without debug assertions). `saturating_mul` is still wrong, for a different reason — it sets a deadline ~584,000 years out on a spin that provably never produces the owed AU, i.e. it wedges the encode thread permanently. A hang, not a panic. WP6.1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0044649b19 |
fix(encode/nvenc): codec-gate the HEVC 4:4:4 union write, and stop it eating the 10-bit arm
android / android (push) Successful in 12m39s
windows-host / package (push) Successful in 10m31s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
decky / build-publish (push) Successful in 19s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 12s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
arch / build-publish (push) Successful in 18m50s
docker / deploy-docs (push) Successful in 12s
docker / build-push-arm64cross (push) Successful in 5m55s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m21s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m29s
ci / docs-site (push) Successful in 1m11s
ci / web (push) Successful in 1m11s
apple / swift (push) Successful in 5m39s
ci / bench (push) Successful in 6m55s
ci / rust-arm64 (push) Failing after 8m55s
ci / rust (push) Successful in 21m27s
deb / build-publish (push) Successful in 8m45s
deb / build-publish-client-arm64 (push) Successful in 9m9s
deb / build-publish-host (push) Successful in 11m26s
apple / screenshots (push) Successful in 26m31s
`encodeCodecConfig` is a C union, so the `hevcConfig` writes in the 4:4:4 branch are only meaningful on an HEVC session — on H.264 or AV1 they reinterpret that codec's own config bytes. The branch was gated purely on `chroma_444 && full_chroma_input` with no codec test, and stayed non-UB only because `lib.rs` degrades 4:4:4 for non-HEVC codecs: a two-file invariant with nothing asserting it, on the path BOTH direct-NVENC backends take. The audit filed that much. What it did not note is that the same shape hides a second bug: this is an `if`/`else if`, so a non-HEVC session that arrived with `chroma_444` set took the HEVC branch and skipped the per-codec bit-depth arm entirely — ending up with neither HEVC 4:4:4 (wrong for it) nor its own 10-bit configuration (simply absent). AV1 would have lost `pixelBitDepthMinus8`/`inputPixelBitDepthMinus8` silently. Non-HEVC now falls through to the arm that knows what to do with it, and an unexpected request is logged rather than swallowed. Adds two tests that need no GPU — `apply_low_latency_config` is pure config authoring, so an AV1 session can assert it never receives the HEVC FREXT profile GUID (an INVALID_PARAM at open) and that its own depth still lands, with an HEVC case guarding the good path. ⚠ `NV_ENC_CONFIG` must NOT be `mem::zeroed` in those tests: `frameFieldMode`/`mvPrecision` are C enums whose discriminants start at 1, so all-zero is not a valid value and Rust's zero-init check ABORTS the process (SIGABRT, caught by the Linux gate). They seed it from `Default` the same way `build_config` does before overwriting from the driver preset. Worth knowing before writing any further test against these SDK types. Verified: Linux gate L1-L4 green (39 tests, the two new ones among them) and the full Windows gate on .173 — 7 legs. Note the Windows test leg runs `--features qsv`, so these tests only execute on the Linux leg; the Windows legs prove the change compiles in all five feature combinations. |
||
|
|
b3f8803ea3 |
fix(encode/qsv): stop advertising intra-refresh the driver silently dropped
`ir_active` was `cfg.intra_refresh && set.co2.is_some()` — i.e. "we asked for it", not "we got it". Both `Query` and `Init` can return MFX_WRN_INCOMPATIBLE_VIDEO_PARAM, a WARNING the open path deliberately accepts, while dropping the intra-refresh wave on the floor. That lie is not cosmetic. `ir_active` feeds `EncoderCaps::intra_refresh`, so the session advertises gradual refresh to the client; the client then stops asking for the IDRs it would otherwise request on packet loss, and a lost frame gets concealed instead of repaired. The stream degrades exactly when recovery matters. Now confirmed against the driver: a best-effort `GetVideoParam` with a CodingOption2 buffer chained on, read back for `IntRefType`. Deliberately a SEPARATE query rather than a buffer attached to the existing `BufferSizeInKB` call (which is what the audit finding suggested) — that value backs every bitstream allocation, and a runtime that disliked the chained buffer would take it down with the readback. A failure here costs only the verdict and resolves conservatively (trust the request), so the worst case is the previous behaviour. Verified on Intel UHD 750 (.42), which is where the Intel on-glass run originally caught this: with PUNKTFUNK_INTRA_REFRESH=1, H.264 now logs "silently dropped intra-refresh (GetVideoParam reports IntRefType=0)" and advertises it OFF, while H.265 on the same GPU stays ON — it is genuinely active there. Per-codec, so it could never have been decided statically. 8 live QSV tests pass. Gated on .173: clippy -D warnings at nvenc,amf-qsv,qsv (host + pf-encode --all-targets), amf-qsv without qsv, qsv alone, no-features, cargo test --features qsv, rustfmt — all green. qsv.rs is Windows-only so the Linux legs do not compile it. Closes WP3.2(c), the last Phase 3 item that was hardware-blocked. |
||
|
|
47a23bec12 |
fix(encode/vulkan): unwind every open/import leak, and serve 24-bpp CPU instead of dying on it
decky / build-publish (push) Successful in 23s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
arch / build-publish (push) Successful in 12m42s
windows-host / package (push) Successful in 18m11s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m14s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m6s
docker / deploy-docs (push) Successful in 32s
docker / build-push-arm64cross (push) Successful in 6m26s
apple / swift (push) Successful in 6m42s
deb / build-publish-host (push) Failing after 5m39s
android / android (push) Failing after 6m36s
deb / build-publish-client-arm64 (push) Failing after 5m35s
deb / build-publish (push) Successful in 9m40s
ci / rust (push) Successful in 22m26s
ci / web (push) Successful in 52s
ci / docs-site (push) Successful in 55s
ci / bench (push) Successful in 5m54s
ci / rust-arm64 (push) Successful in 11m18s
apple / screenshots (push) Successful in 27m18s
Phase 5's Linux half (audit WP5.1 + WP5.4), each item shaped by the
review that rejected the obvious fix:
Dmabuf import unwind (vk_util): every failure after create_image leaked
the VkImage, and the dup'd dmabuf fd leaked as a raw i32. The sharp edge
is that a SUCCESSFUL vkAllocateMemory transfers fd ownership to Vulkan
(vkFreeMemory closes it), so the naive close-on-error is a double close
that clobbers whatever unrelated descriptor recycled the number. The dup
now lives in an OwnedFd released exactly in the allocate-success arm;
every other path drops it once, and bind/view failures free image+memory.
PyroWave open unwind: open_inner had ~20 fallible steps that each leaked
everything before them (instance, device, pyrowave objects, the whole
CSC pipeline). Rather than a parallel teardown guard — whose reviewed
hazards were a null-unsafe pyrowave_encoder_destroy and a drifting
duplicate of Drop — Self is now constructed right after create_device
with every later resource null, and the existing Drop (wait-idle first,
pw_enc null-guarded, delete-nullptr and VK_NULL_HANDLE destroys are
no-ops) is the single unwind path for error and normal teardown alike.
The ensure_cpu_rgb staging twins (create/allocate/bind, both backends)
and the RGB-direct make_view pair get the same discipline via a shared
make_host_buffer. Observed on hardware: 32 forced import failures, zero
fd drift (the new import_failure_leaks_no_fds smoke on RADV).
24-bpp CPU service (WP5.4): pixel_to_vk had no mapping for the packed
Rgb/Bgr the PipeWire portal negotiates, so a session committed to a path
the backend could not serve and died at its first frame. The filed
open-gate was rejected as a half-mirror — the dmabuf axis is keyed by
fourcc at submit, unknowable at open — so instead the CPU axis is
SERVED: a 3-to-4 expand at the staging upload (normalize_cpu_rgb, the
CPU twin of WP1.4's swscale expand), order-preserving for the CSC
samplers and BGRA-forced for the RGB-direct encode source, whose session
pictureFormat is B8G8R8A8 — the on-glass run caught R-first sources
violating VUID-vkCmdEncodeVideoKHR-pEncodeInfo-08207, a mismatch that
predates this change for plain Rgbx CPU sources. The dmabuf axis feeds
pf-zerocopy's raw-dmabuf degrade latch (
|
||
|
|
e0e845b7df |
fix(encode/pyrowave): pin the NT-handle import contract to consume-on-success-only
import_plane closed the shared NT handle on every pyrowave_image_create
failure, believing pyrowave only consumes it on success. The vendored
truth was messier: Granite's allocator closed by-reference handles
UNCONDITIONALLY after the first vkAllocateMemory — success AND failure —
while failures before the allocator (validation, vkCreateImage, no
memory type) left the handle open, and both classes surface as the same
error code. So the call site could not know whether to close, and an
allocate-stage failure double-closed a possibly-recycled handle value
(audit WP5.3). The filed fix (DuplicateHandle + close the original
unconditionally) traded the double close for a guaranteed leak on every
pre-allocate failure and left the ambiguity in place.
Patch 0006 fixes the callee instead: the allocator never closes a
caller's handle, and pyrowave_image_create consumes it exactly at its
success return — which is what pyrowave.h ("take ownership and close the
HANDLE on import") documents, and what Granite's semaphore import
already did, so import_fence was correct as written and the two paths
now share one contract. The close-on-failure in import_plane is thereby
correct on EVERY path, including a vkBindImageMemory failure after a
successful allocate. Bonus fix recorded in the patch: the allocator's
block-recycling retry loop used to re-run vkAllocateMemory with
import_info still pointing at the just-closed handle.
Both-platform gates green; pyrowave_win_smoke (.173, 34/34) re-validates
the live import path against the rebuilt vendored C++.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d6d4408c8e |
fix(encode/nvenc): teardown that tells wedged from routine, and a session budget that can heal
Three Windows NVENC teardown/accounting defects (audit WP5.2, marked BLOCKING — the obvious fixes were reviewed and rejected before this): Event-handle leak: init_session pushed each completion event to self.events only AFTER nvEncRegisterAsyncEvent, so a failed registration leaked the Win32 handle. Push first — the init-failure path already runs teardown, which closes everything in the list (unregistering the one never-registered event is a harmless error return). Wedged drain: teardown drains the retrieve thread's backlog, and each queued job could burn its full 5 s completion wait — cap x 5 s on the encode thread, paid again on every rebuild of the stall-recovery ladder. The rejected fix (a shutdown flag) abandoned the backlog on EVERY teardown, turning the routine drain into destroy-while-encoding. Instead the wedge evidence stays where it lives: after one full-budget timeout in the retrieve loop, later jobs wait a 250 ms slice; a success resets the latch. Routine teardown is byte-identical to before — nothing is ever abandoned — and a first timeout is already encoder-fatal, so short slices behind an established wedge change only how long teardown blocks. Session-unit accounting: LIVE_SESSION_UNITS was refunded even when destroy_encoder FAILED, drifting the budget low and over-admitting parallel displays. The rejected fix (never refund on failure) let one transient wedge episode poison admission until a host restart. Now the refund follows PROOF: destroy succeeded, or its status says the driver holds no session (device gone/TDR). Ambiguous failures park the handle — units still charged, fail-closed, pinning the D3D11 device the driver session references — and init_session retries the destroy while zero sessions are live, under a gate that serializes every session open against the reap so a recycled handle address can never alias a live session. The charge lasts exactly as long as the driver keeps refusing the destroy, which is the definition of still-leaked. Admission stays a lock-free atomic load. Both-platform gates green (docker linux/amd64 4 legs; .173 all 7 legs). The retry-destroy-on-a-failed-handle assumption is documented at the reap: the SDK defines no retry semantics either way. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3efbe4164e |
fix(capture): capture CPU frames once the encoder proves it can't import dmabufs
audit / cargo-audit (push) Successful in 2m13s
audit / bun-audit (push) Failing after 13s
ci / web (push) Successful in 57s
ci / docs-site (push) Successful in 1m3s
windows-host / package (push) Successful in 10m18s
ci / bench (push) Successful in 6m41s
android / android (push) Successful in 12m26s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 14s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
arch / build-publish (push) Successful in 12m30s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m10s
ci / rust-arm64 (push) Successful in 11m46s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m17s
deb / build-publish (push) Successful in 11m17s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m13s
deb / build-publish-host (push) Successful in 12m35s
deb / build-publish-client-arm64 (push) Successful in 8m23s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m5s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 7m40s
apple / swift (push) Successful in 5m28s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8m6s
docker / deploy-docs (push) Successful in 14s
ci / rust (push) Successful in 21m44s
docker / build-push-arm64cross (push) Successful in 7m18s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m37s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m14s
flatpak / build-publish (push) Successful in 6m38s
release / apple (push) Successful in 29m56s
apple / screenshots (push) Successful in 25m36s
A dmabuf import the GPU driver refuses is refused identically on every retry, but the only recovery above it was the encode-stall ladder: five in-place encoder rebuilds, then the video session ends. So a host whose driver will not take what its compositor allocates lost every session on its first frame, and every reconnect repeated it — while the very same host streamed fine with `PUNKTFUNK_ZEROCOPY=0`. The software knew how to run that machine and never chose to. Latch it, exactly as the sibling CUDA-import path already does after repeated worker deaths: three consecutive import failures with no frame in between disable the raw-dmabuf passthrough for the host process, and capture negotiates CPU frames from the next session on. Three sits below the encoder's rebuild budget, so the latch is set before the session it doomed ends — one bad session, then a working (if slower) host, with a log line saying which and why instead of an operator having to find an environment variable. Only the two stages that ARE the import are counted — the buffersrc push of the DRM-PRIME descriptor and the buffersink pull where `hwmap` maps it into a VA surface. `avcodec_send_frame` is deliberately left out: that one is the encoder stalling, which the in-place rebuild exists to recover, and taking zero-copy away permanently over a transient fault would be a bad trade. The latch lives in pf-zerocopy because it is the leaf both sides can see — the capture→encode edge is one-way by design, so pf-capture cannot ask pf-encode anything. |
||
|
|
f32207a6ba |
fix(encode/vaapi): tell libva the dmabuf's real size, not zero
The zero-copy DRM-PRIME descriptor declared `objects[0].size = 0`, on the belief that ffmpeg would work the real size out. It does not: both of its import paths pass the value straight through to libva — `prime_desc.objects[i].size` on the PRIME_2 path, `buffer_desc.data_size` on the legacy fallback it tries next — so every VA driver we have ever handed a dmabuf to was told the backing object was empty and left to derive the size itself. The drivers this path has run on (radeonsi, modern Intel iHD) derive it correctly, which is why nobody noticed. A Gen9 Intel host does not get that far: `vaCreateSurfaces` answers VA_STATUS_ERROR_ALLOCATION_FAILED on the first frame and every frame after it, `av_buffersink_get_frame` returns EIO, and five in-place encoder rebuilds later the video session is over. That host cannot stream at all with zero-copy on. `lseek(SEEK_END)` is the standard dma-buf size query and the same one this tree's Vulkan bridge already performs on these very fds. A kernel that refuses it leaves the old 0 rather than costing a frame that might still have encoded. Whether this alone fixes that host is unconfirmed — the descriptor was wrong either way, and it is the first thing the driver reads. |
||
|
|
7c82a72ecd |
fix(zerocopy): find the NVIDIA render node instead of assuming renderD128
The EGL importer — the head of the CUDA/NVENC zero-copy path — opened
`/dev/dri/renderD128` and hoped. On a single-GPU host that is the
NVIDIA node and it always worked. On a hybrid laptop the iGPU is bound
first, so renderD128 is Intel: we built a GBM device on Mesa, asked it
for a pbuffer-capable OpenGL config, got none (Mesa's GBM platform
advertises only window-capable ones — gbm_surface is the point of that
platform there), and reported
zerocopy worker init failed: no EGL config for OpenGL
on a machine whose NVIDIA EGL stack was demonstrably fine, naming
neither the device nor the reason. `PUNKTFUNK_RENDER_NODE` was no
escape either: this path never read it.
Pick the node by what it is — the first `/dev/dri/renderD*` whose
sysfs PCI vendor is 0x10de, the same identification the crate's Vulkan
bridge already uses for its physical device. A local scan rather than a
`pf_gpu` dependency, because this crate is a leaf whose worker is its
own process, and `pf_gpu::linux_render_node` answers a different
question anyway: it follows the operator's GPU *preference*, which may
legitimately name the iGPU. What is needed here is not which GPU to use
but where CUDA is. `PUNKTFUNK_ZEROCOPY_RENDER_NODE` overrides, and no
NVIDIA node found keeps the historical guess — sysfs may just not be
mounted, and the CUDA context is what properly fails on a host with no
NVIDIA at all.
Then stop the config query being able to fail this way: retry without
the surface-type constraint, which is honest because we never create an
EGLSurface — `eglMakeCurrent` runs surfaceless. And name the node in
every error and in the ready line, so the next hybrid host diagnoses
itself.
|
||
|
|
6571a26aa9 |
fix(scripting): park the idle runner instead of pinning a core
`punktfunk-scripting` with nothing to run pinned a full CPU core
indefinitely — in exactly the state its own systemd unit documents as
inert ("the runner does nothing until you add scripts or install
plugins"). A field report on a 7.7 GB laptop had it at 99.9% CPU for
two hours straight, `strace` showing a bare `clock_gettime` loop and no
other syscall at all, which starved the compositor badly enough to
blank the physical display.
Nothing in the runner was looping. `Effect.runFork` registers nothing
with the event loop — effect's keep-alive timer lives in
`Runtime.makeRunMain`, which runner-cli deliberately doesn't use
because its shutdown is the bespoke two-signal one — and a bun process
whose only pending work is an unresolved promise busy-spins rather than
blocking. With units running there is normally a socket or timer to
park on; with none there is nothing at all, so the no-op state was the
one state that span.
Hold one idle handle for the runner's lifetime and drop it when the
unit tree ends on its own. Measured on the packaged bun (1.3.14), three
seconds idle with empty scripts/plugins dirs: 3098 ms of CPU before,
280 ms after (bun's startup, then nothing).
The regression is a property of the process rather than of any function
in it, so the test spawns the real CLI and reads the child's CPU time,
asserting it also logged and stayed alive — a crashed runner burns no
CPU either and must not pass.
|
||
|
|
8578141d43 |
fix(encode/vulkan): AV1 at unaligned modes was violating two VUIDs on every frame
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
android / android (push) Successful in 15m38s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 13s
arch / build-publish (push) Successful in 16m13s
windows-host / package (push) Successful in 10m20s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m20s
apple / swift (push) Successful in 5m36s
docker / deploy-docs (push) Successful in 23s
docker / build-push-arm64cross (push) Successful in 8m41s
ci / web (push) Successful in 48s
ci / docs-site (push) Successful in 1m8s
ci / bench (push) Successful in 6m49s
deb / build-publish-client-arm64 (push) Successful in 7m31s
deb / build-publish (push) Successful in 11m40s
ci / rust-arm64 (push) Successful in 12m11s
deb / build-publish-host (push) Successful in 12m15s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 14m59s
ci / rust (push) Successful in 21m45s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m51s
apple / screenshots (push) Successful in 25m22s
Found by running the smokes at 1920x1080 under the validation layers while chasing the three
items left over from WP4.2. AV1 forbids the encode source's `codedExtent` differing from the
sequence header without `FRAME_SIZE_OVERRIDE`
(VUID-vkCmdEncodeVideoKHR-flags-10324), or from the reference slots without
`MOTION_VECTOR_SCALING` (`-10325`). RADV PHOENIX advertises NEITHER — and RGB-direct is the
default on EFC hosts with true-extent the default at unaligned modes, so plain 1080p AV1 was
tripping both on every frame: source 1920x1080 against an app-aligned 1920x1088 header and
DPB. Measured 16 violations per 8-frame run; the CSC path had none.
The fix is to make all three agree at the RENDER size rather than the aligned one — an
unpadded coded size is valid on this hardware, so the coded frame simply IS the visible
frame:
- the AV1 sequence header follows the render size when true-extent is active (joining
native NV12, which already authors true-size headers for the same reason);
- the DPB setup and reference slots carry `src_extent` instead of the aligned `ext2d`.
`src_extent` already collapses to `ext2d` whenever true-extent is off, so every other
configuration is untouched;
- `render_and_frame_size_different` now compares render against the DECLARED source extent
instead of the aligned size, or true-extent would have claimed a mismatch that no longer
exists.
Fixing only the header is not enough and is actively misleading: it clears -10324 and
immediately exposes -10325, because the mismatch has moved to the reference slots rather than
gone. Both had to move together.
This keeps the EFC fast path. Two alternatives were implemented, measured and rejected on the
way here: falling back to the compute CSC costs the zero-copy the B2 work existed to deliver,
and routing to the padded-copy staging trades these two VUIDs for
VUID-VkImageCreateInfo-pNext-06811 — `pad_img`'s extra TRANSFER_SRC usage is not in the
profile-advertised set, measured 8x per session on HEVC-padded too, so that is a pre-existing
defect of the padded path and not somewhere to route a default session.
HEVC is deliberately untouched: it has no equivalent constraint (its crop rides the
conformance window), it measures zero violations, and its aligned-SPS path is the validated
one.
On RADV PHOENIX (780M, Mesa 26.0.4) the AV1 stream now decodes as coded 1920x1080 / render
1920x1080 — genuinely unpadded, where before the alignment rows were encoded and cropped
back out. The CSC path still reports coded 1920x1088 / render 1920x1080, which is correct for
it. All four `vulkan_smoke*` pass at 256x256 and 1920x1080.
Two validation errors remain on the RGB-direct path and are NOT ours, now with evidence
rather than assumption: `VUID-VkImageViewCreateInfo-image-08336` uses the PROFILE-BLIND format
query, so it cannot see that RGB conversion legalises BGRA as an encode source — the
profile-aware query used by -06811 accepts the very same image; and
`VUID-VkQueryPoolCreateInfo-pNext-pNext` rejects
`VkVideoEncodeProfileRgbConversionInfoVALVE`, which the VALVE extension REQUIRES for profile
identity, and the layer diagnoses itself as "a struct from an extension added to a later
version of the Vulkan header".
Verified: canonical Linux gate (docker linux/amd64) L1-L4 green; on-glass on RADV PHOENIX
under `VK_LOADER_LAYERS_ENABLE='*validation*'`, with the bitstreams read back through
libdav1d/trace_headers.
|
||
|
|
87863c12a9 |
fix(encode/windows): clear the Windows Phase 4 backlog — NVENC sticky state, QSV HDR + ABR
arch / build-publish (push) Failing after 5m39s
ci / web (push) Successful in 1m2s
ci / docs-site (push) Successful in 1m3s
android / android (push) Successful in 11m50s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
ci / bench (push) Successful in 6m25s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
ci / rust-arm64 (push) Successful in 10m5s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 5m22s
deb / build-publish-client-arm64 (push) Successful in 9m34s
deb / build-publish (push) Successful in 10m58s
docker / deploy-docs (push) Successful in 12s
deb / build-publish-host (push) Successful in 11m51s
docker / build-push-arm64cross (push) Successful in 6m20s
apple / swift (push) Successful in 5m22s
ci / rust (push) Successful in 22m30s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m2s
windows-host / package (push) Successful in 10m48s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m13s
apple / screenshots (push) Canceled after 16m50s
Four items, gated on the RTX box (.173, all 7 Windows legs) and the Linux docker gate, with
the QSV HDR one proven by an A/B on real Intel silicon (.42, UHD 750).
WP4.1 — Windows NVENC sticky state. Both halves were the same mistake: a single field doing
duty as both "what the session negotiated" and "what this session can actually do", so the
first downgrade destroyed the negotiation.
- `query_caps` clears `self.hdr` on a GPU without 10-bit encode, but `submit` re-derives the
requested HDR state from every frame's pixel format and compared it against that cleared
value. On such a GPU a P010 capturer therefore reported "HDR changed" on EVERY FRAME and
tore down and rebuilt the whole encode session each time. The trigger now compares against
`hdr_requested`, and the downgrade is latched in `hdr_unsupported` so it is remembered
(and warned once) instead of rediscovered per init.
- One subsampled-YUV frame cleared `chroma_444` permanently, so 4:4:4 never returned when
the capturer went back to RGB. The negotiation now lives in `chroma_444_requested` and the
effective value is recomputed at every init; the per-session downgrade still happens and
is still reported honestly through `caps()`, it just is not destructive any more.
WP4.3(a) — QSV HDR mastering luminance was divided by 10,000. The old comment justified it as
"VPL wants whole cd/m²", which is true of the VIDEO-PROCESSING unit but not the encode path:
the runtime passes the value straight into the ITU-T H.265 Annex D mastering-display SEI,
whose unit is 0.0001 cd/m². Verified end to end on Intel UHD 750 by dumping the bitstream and
reading SEI 137 with trace_headers, before and after:
before: max_display_mastering_luminance = 1000 (0.1 nits)
after: max_display_mastering_luminance = 10000000 (1000 nits)
with `min` (500 = 0.05 nits) and both content-light-level fields unchanged. That asymmetry —
a correct min beside a 10,000x-low max — was the original tell.
WP4.3(b) — `reconfigure_bitrate` never re-read `BufferSizeInKB`, so an ABR step-up could hand
the runtime an AU buffer sized for the old, lower bitrate. Fixed in both places it was broken:
re-read the driver's answer via `GetVideoParam` after the `Reset` (mirroring `init_encode`),
AND stop `take_bs` recycling pooled buffers without checking their capacity — the pool would
otherwise keep serving short buffers even once `bs_bytes` was correct.
WP7.8 — the `PUNKTFUNK_QSV_FFMPEG` half, finally landable now that a Windows gate exists (it
sits behind `#[cfg(feature = "qsv")]`, so nothing in a Linux-only matrix compiles it — it was
deliberately held back for exactly that reason). Trimmed like its siblings, and kept
case-insensitive: the knob already accepted `TRUE`, and the bare house `matches!` would have
silently dropped that spelling.
Phase 4 of the audit is now complete, and WP7.8 with it.
Verified: Windows gate on .173 — clippy -D warnings at nvenc,amf-qsv,qsv (host + pf-encode
--all-targets), amf-qsv without qsv, qsv alone, no-features, `cargo test --features qsv`
(34 passed), rustfmt. Linux docker gate L1-L4 green. On-glass: the QSV A/B above on .42.
Owed: NVENC sticky-state on-glass on .173 (needs a live 10-bit-capable capture toggle).
|
||
|
|
bab309017d |
fix(encode/vulkan): five mode-specific correctness bugs, three caught on real silicon
WP4.2 of the pf-encode audit, validated on RADV PHOENIX (780M, Mesa 26.0.4) under the Vulkan
validation layers rather than by inspection alone. Three of the five reproduce with the
driver naming the violation; one turned out milder than filed and is recorded as such.
1. `ensure_cpu_rgb` cached the staging image on FORMAT ONLY while sizing it to the SOURCE
frame, and the copy uses the CURRENT frame's extent — so a same-format source-size
increase copied past the allocation. Now keyed on (format, width, height). This is the
memory-safety one: on hardware it was 8 x
VUID-vkCmdCopyBufferToImage-imageSubresource-07971 ("extent.width (512) exceeds
imageSubresource width extent (128)") and `submit` returned Ok every single time, so
nothing upstream could notice. Zero after the fix. Note the trigger for anyone re-testing:
the image is cached PER RING SLOT, so the ring must wrap before a slot sees a larger frame
— two differently-sized submits in a row land on different slots and look fine.
2. The RGB-direct CPU padding loop only ever grows a source into the aligned extent; a source
LARGER than the encode extent panicked on the row `copy_from_slice`. Now refused by name,
matching what the DMA-BUF arms already do, so a bad frame takes the encoder-rebuild path
instead of unwinding out of the encode thread with an index message.
3. AV1 wrote `render_width/height_minus_1` but never set
`render_and_frame_size_different`, and AV1 ignores the render size unless that flag says it
differs — so every mode needing 64x16 alignment presented the padding. libdav1d on the
encoded stream now reports `size 1920x1088 ... render 1920x1080`; before, the flag being
unset meant render defaulted to the coded 1088, i.e. 8 rows of duplicated edge pixels
shipped to the client.
4. `read_slot` read the PERF timestamp pool whenever the pool merely EXISTED, but the
padded-RGB path's CPU-upload arm records its own command buffer and writes no timestamps —
reading an unreset query with WAIT is undefined. Now gated on a per-slot `ts_written`,
because "a pool exists" was being used as proof of "the pool was written". SEVERITY
CORRECTION vs the audit: this is not a hang in practice. It needs PUNKTFUNK_PERF *and*
RGB-direct *and* the non-default PUNKTFUNK_VULKAN_RGB_TRUE_EXTENT=0 *and* an unaligned mode
*and* CPU capture, and RADV then fails the read rather than blocking. Another driver may
block, hence the fix.
5. `reset()` re-arms `first_frame`, which was also what decided whether begin-video-coding
declares the rate-control state. But a reset does not rebuild the session, so the CBR
installed earlier is still current when the next frame opens its coding scope: the
declaration was omitted exactly when it was required. Split into `rc_installed`, which
survives reset. On hardware this was
VUID-vkCmdBeginVideoCodingKHR-pBeginInfo-08253 ("no VkVideoEncodeRateControlInfoKHR ... but
the currently set mode is CBR"); zero after.
Also fixes two pre-existing AV1 spec violations that were firing on EVERY frame of a shipped
path and are not in the audit at all — found only because the validation layers were on, and
confirmed untouched by this branch before being attributed as pre-existing:
`constantQIndex` must be 0 unless rate control is DISABLED (VUID-...-10320; this session is
always CBR, so the driver owns Q), and `pStdPictureInfo->pSegmentation` must be NULL
(VUID-...-10350; it pointed at a zeroed struct, and leaving it unset means the same
"segmentation disabled" while being spec-correct).
Items 1 and 5 keep their reproductions as `#[ignore]`d hardware tests, documented as
meaningful only under `VK_LOADER_LAYERS_ENABLE='*validation*'` since the violations are
invisible to the API's return values.
STILL OPEN, filed not fixed (three more pre-existing validation errors the cleanup exposed):
VUID-VkQueryPoolCreateInfo-pNext-pNext on the AV1 and RGB paths, and 2 x
VUID-VkImageViewCreateInfo-image-08336 (the RGB-direct encode-src view's format features lack
VIDEO_ENCODE_INPUT) — the latter may be a validation-layer gap around the VALVE extension
rather than our bug, and wants real analysis rather than a quick patch.
Verified: canonical Linux gate (docker linux/amd64) fmt + clippy --all-targets at default and
at nvenc,vulkan-encode,pyrowave + both test legs; and on RADV PHOENIX all four `vulkan_smoke*`
tests pass at 256x256 and at 1920x1080 with zero validation errors on the paths touched here.
|
||
|
|
e5e68f1d24 |
feat(presenter): DRM card selection + a usable kmsdrm swapchain error
ci / web (push) Successful in 1m2s
ci / docs-site (push) Successful in 1m15s
apple / swift (push) Successful in 5m25s
ci / bench (push) Successful in 7m44s
deb / build-publish-host (push) Successful in 10m30s
decky / build-publish (push) Successful in 30s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
android / android (push) Successful in 12m52s
deb / build-publish (push) Successful in 11m46s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
arch / build-publish (push) Successful in 13m6s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 45s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 12s
ci / rust-arm64 (push) Successful in 13m32s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m1s
deb / build-publish-client-arm64 (push) Successful in 9m12s
windows-host / package (push) Successful in 18m4s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 7m1s
flatpak / build-publish (push) Successful in 7m25s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m46s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 9m7s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 9m24s
ci / rust (push) Successful in 23m48s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m28s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 4m46s
docker / deploy-docs (push) Successful in 22s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 5m42s
release / apple (push) Successful in 31m30s
docker / build-push-arm64cross (push) Successful in 7m6s
apple / screenshots (push) Successful in 23m4s
Both come straight out of the first compositor-less bring-up (P3), on a two-GPU box: PUNKTFUNK_DRM_CARD=<n> pins SDL's KMSDRM device index. SDL enumerates /dev/dri/card* and takes the first it can open, which is regularly the wrong one: it chose the card a live compositor already held DRM master on and died at swapchain creation, while the idle card with the connected display sat unused. Kept an explicit operator choice rather than auto-detection — deciding "is this card already mastered" needs the very ioctl that taking master IS, so any in-process guess would be fragile. The swapchain error now carries what to actually check. "vkCreateSwapchainKHR: Initialization of an object has failed" is useless to someone bringing up a kiosk; on the kmsdrm backend it now names the pinned card and lists the three real causes in order (no connected connector / another DRM master / NVIDIA). Empty on every other backend, where it would be noise. The NVIDIA note is measured, not guessed: on NVIDIA proprietary + kmsdrm, Vulkan enumerates the GPU AND the display (VK_KHR_display reports the connected HDMI connector) and still fails — as root, with nvidia_drm.modeset=Y, on a card no compositor was using. Not permissions, not DRM master; their direct-display path wants the display leased via vkAcquireDrmDisplayEXT and SDL's kmsdrm surface path does not do that. Plan: punktfunk-planning design/embedded-arm64-client.md §P3 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
aff169e131 |
build(flatpak): make the manifest architecture-generic
Two things were x86-specific, both now expressed properly instead of
hardcoded:
* PKG_CONFIG_PATH carries the runtime's multiarch directory.
flatpak-builder does NOT shell-expand `env` values, so ${FLATPAK_ARCH}
would be taken literally — a build-options.arch override supplies the
aarch64 string and inherits the rest.
* The prebuilt Skia archive is per-target and pinned by sha256. Two
`type: file` sources discriminated by only-arches now share one
dest-filename, so SKIA_BINARIES_URL stays a single literal path.
Upstream publishes the aarch64 archive under the same skia commit hash
and the same resolved-feature key, so the two stay in lockstep on a
skia-safe bump.
build-flatpak.sh gains ARCH (default: this machine's), passes --arch to
both flatpak-builder and build-bundle, and arch-suffixes the bundle name so
an x86_64 and an aarch64 build can coexist in dist/ instead of the second
silently overwriting the first. CI composes its own published filename, so
nothing downstream changes.
The aarch64 Skia sha256 was verified by downloading the published archive,
not taken from the API. No aarch64 flatpak has been built — flatpak-builder
builds in a sandbox for the target arch, which needs an arm64 machine.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4dc078692d |
build(arch): build the client for aarch64 (Arch Linux ARM)
arch=('x86_64' 'aarch64'), and on aarch64 `pkgname` drops punktfunk-host
so makepkg never enters the host's build or package path; build() takes a
client-only cargo invocation without the NVENC/Vulkan-encode features.
Dropping the host from pkgname rather than giving package_punktfunk-host a
per-package arch is what makes the skip unambiguous, and it mirrors how
PF_WITH_WEB already extends that array.
The host stays x86-only because its encode stack (NVENC/QSV/AMF) is.
Verified by sourcing the PKGBUILD under both CARCH values: x86_64 yields
`punktfunk-host punktfunk-client`, aarch64 yields `punktfunk-client`. NOT
verified by a real build — there is no official arm64 Arch container, and
the README says so rather than implying it is tested.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
1c76b8c7b4 |
build(rpm): let the spec build a client-only aarch64 RPM
The blocker was never ExclusiveArch — it was that %build builds host + tray + client together and %install lays down the host unconditionally, with no way to express "client only". `%bcond_without host` (default ON, so an x86_64 build is unchanged) now gates the host binary, the tray, the headless-session data, the firewalld services, the main %post and the bare %files section. Omitting %files for the MAIN package is the load-bearing part: it is what stops rpm emitting an empty `punktfunk` alongside punktfunk-client. build-rpm.sh exposes it as PF_WITHOUT_HOST=1 and skips the libcuda stub regeneration and the libcuda leak check, neither of which means anything without the host. Verified two ways. `rpmspec -P` shows the default expansion still carrying 3 host install lines, the tray build and one bare %files, while --without host carries none of those and one %files client. And a real rpmbuild in a native aarch64 Fedora 43 container produced punktfunk-client-*.aarch64.rpm with correct aarch64 sonames (libc.so.6(GLIBC_2.17), libSDL3.so.0, libavcodec.so.61) and no main package. Not a cross-compile: %build runs cargo for the build machine's arch, so this wants an arm64 builder. No CI leg yet — deliberate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cac82d13df |
ci: check the client for aarch64 on every push
Adds a `rust-arm64` job: clippy -D warnings for the client crates at --target aarch64-unknown-linux-gnu, plus a --no-default-features session build so the minimal embedded configuration keeps compiling in its own right rather than only as a subset of the default one. Runs in the cross image on the ordinary amd64 runner; client crates are listed explicitly because --workspace would drag in the x86-only host encode stack. This is a regression guard, not an artifact leg (deb.yml ships those). c_char signedness cuts both ways and neither direction is visible from an x86-only CI: hardcoding i8 fails to COMPILE on ARM, while casting to u8 fails the LINT on ARM. Two such defects were already found this way. Also documents the arm64 client .deb recipe in packaging/debian/README.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ab3884c763 |
fix(core/abi): cast identity PEM pointers with .cast(), not as *mut u8
`c_char` is i8 on x86_64 and u8 on aarch64, so `cert_pem_out as *mut u8` is a REQUIRED conversion on one target and a no-op on the other — where clippy::unnecessary_cast then denies it. `.cast::<u8>()` is correct and lint-clean on both. Caught by the new aarch64 clippy leg on its first run. A sweep of the remaining `as *mut u8` / `as *const u8` sites in the client crates found no other c_char-derived casts; the rest convert from c_void or u16. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
000f8b85b0 |
build(ci): cross-compile and publish an aarch64 Linux client
Adds punktfunk-rust-ci-arm64cross (the rust-ci toolchain + an Ubuntu ports
arm64 multiarch sysroot) and a deb.yml job that cross-builds the client
package on the ordinary amd64 runner — no arm64 runner in the fleet and
none needed. Client only: the Linux host encodes with NVENC/QSV/AMF, all
x86. The apt registry keys pool entries by the .deb's own arch, so an arm64
box picks the package up with no client-side configuration.
Three things this had to get right, each of which cost a failed build:
* rustup target add must run against the toolchain rust-toolchain.toml
PINS, not the image's default `stable` — otherwise the pinned toolchain
has no aarch64 std and the build dies ~230 crates in on "can't find
crate for core". Copying the pin file in also pre-downloads the
toolchain every CI job currently re-fetches on first use.
* ffmpeg-sys-next compiles a probe it intends to RUN, so it forces
.target(HOST) while still passing the TARGET's include paths; the arm64
dir then shadows the host's own libc headers and the x86 compiler dies
in bits/math-vector.h on NEON/SVE types. Prepending the host dir does
NOT help — GCC drops a -I that duplicates a system directory, keeping
its original later position — so ci/pf-host-cc strips target include
dirs from host compiles instead.
* the job hard-fails on a non-AArch64 session binary, rather than
shipping an amd64 payload under an arm64 package name.
skia needs no special handling: the aarch64-linux-gnu textlayout+vulkan
prebuilt resolves, so arm64 ships the FULL client, OSD included — unlike
Windows ARM64, which still builds --no-default-features.
The cross image builds in its own docker.yml job (needs: build-push) since
it is FROM punktfunk-rust-ci:latest and must not race the entry that
publishes that base. rpm/Arch/Flatpak stay x86_64 for now — their builders
have no arm64 sysroot, which is its own piece of work.
Plan: punktfunk-planning design/embedded-arm64-client.md
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
915fc3ef9e |
feat(session): headless --pair, so a box with only SSH can enrol
`punktfunk-session --pair <PIN> --connect host[:port]` runs the SPAKE2 ceremony with no window and no toolkit, prints the same `paired <addr>:<port> fp=<hex>` line as `punktfunk-client --pair`, and exits. Until now the PIN ceremony lived only in the GTK shell or the Skia console, so enrolling an embedded/kiosk client meant installing a desktop on it — or copying the identity store by hand. Dispatches above every graphics call (the machine may have no display) and is in the `--no-default-features` build: enrolling must never be the reason a minimal image pulls in Skia. `--name` defaults to the hostname instead of the desktop path's hardcoded "Steam Deck". `forget_placeholder` and `device_name` move into pf_client_core::trust so the two binaries share one implementation rather than the session re-deriving them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b7cea71bbd |
fix(presenter): type the Vulkan extension pointers as c_char, not i8
`char` is signed on x86_64 but UNSIGNED on aarch64, so the hardcoded `Vec<*const i8>` compiles on every desktop target and then fails to match ash's `&[*const c_char]` when cross-compiling the client for ARM. Found by the aarch64 cross-build spike (punktfunk-planning: embedded-arm64-client.md) — the only portability defect in the client stack; nothing else in the client crates assumes an architecture. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
07f8e37c85 |
fix(encode/sw): refuse a mode openh264 cannot encode at open, not at every submit
ci / docs-site (push) Successful in 1m2s
ci / web (push) Successful in 1m10s
decky / build-publish (push) Successful in 18s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 8s
apple / swift (push) Successful in 4m12s
ci / bench (push) Successful in 7m0s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 7m3s
android / android (push) Successful in 12m29s
docker / deploy-docs (push) Successful in 28s
deb / build-publish-host (push) Successful in 13m15s
deb / build-publish (push) Successful in 11m37s
arch / build-publish (push) Successful in 15m59s
windows-host / package (push) Successful in 18m2s
ci / rust (push) Successful in 22m20s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 14m44s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m53s
apple / screenshots (push) Successful in 24m23s
openh264 tops out at level 5.2 — 3840x2160 landscape or 2160x3840 portrait — and enforces that ceiling inside `reinit`, which the crate calls on the FIRST ENCODE rather than at encoder construction. So an oversized mode built a perfectly healthy-looking encoder and then failed every single submit: the session connects, negotiates, and never delivers a frame, with the real reason buried in a per-frame error rather than at the open. `validate_dimensions` does not cover this. It is keyed on the codec, and H.264 legitimately reaches 4096 on every hardware backend — this ceiling belongs to the software backend alone, which is exactly the path a GPU-less host falls back to. Rejects at open instead, mirroring the rule from the openh264 version we actually ship (0.9.3) rather than from its docs — including the orientation-aware shape, since a naive per-axis `w <= 3840 && h <= 2160` would wrongly refuse a legal 2160x3840 portrait session. Three tests: the accepted modes in both orientations, the modes `validate_dimensions` lets through but openh264 cannot serve, and that `open` itself refuses rather than deferring to submit. They cost nothing to run — the guard fires before openh264 is initialised at all. Verified with the canonical Linux gate (docker linux/amd64): fmt, clippy --all-targets at default and at nvenc,vulkan-encode,pyrowave, and the pf-encode test leg (37 passed). |
||
|
|
ed525c4c73 |
fix(encode/vulkan): PUNKTFUNK_VULKAN_RGB_DIRECT read "0 " as force-ON
The knob tested `v == "0"` for the disable case and treated everything else that was set as a force-enable. So a trailing space — the kind a `.env` line or a shell script leaves behind — made `PUNKTFUNK_VULKAN_RGB_DIRECT=0 ` mean the exact opposite of what the operator wrote, forcing the RGB-direct encode source on. On a cursor-blend session that is not a subtle difference: the EFC front-end cannot blend, so the pointer disappears from the stream. An empty value and any typo did the same thing. Now parsed like every sibling knob in the crate — `matches!(v.trim(), "1"|"true"|"yes"|"on")` and the matching false spellings — so unrecognised input falls back to the default instead of being read as a force-on. `=0` and `=1` keep working exactly as documented. Splits the parse into a pure function so the accepted spellings are covered by tests. That matters more than it looks: the env-var read itself is untestable in a parallel test binary (the process environment is shared), which is precisely why this knob's behaviour had no coverage and the inverted case survived. Verified with the canonical Linux gate (docker linux/amd64): fmt, clippy --all-targets at default and at nvenc,vulkan-encode,pyrowave, and the pf-encode test leg (37 passed). |
||
|
|
3fac1a6da1 |
fix(encode/linux): stop the capability probes racing each other's libav log level
Every capability probe drops libav's log level to AV_LOG_FATAL around an encoder open it expects to fail, then restores what was there before. That level is one process-global int and the probes genuinely overlap — they are reached both from `/serverinfo` and from session bring-up. Two overlapping save/restore pairs interleave as get(INFO) → get(FATAL) → set(INFO) → set(FATAL), and the process is then pinned at AV_LOG_FATAL permanently: every later libav diagnostic silently dropped, including the ones you want when a stream fails to open later on. Replaces the four hand-rolled save/restore pairs with one RAII guard over a single shared mutex. The audit filed two sites; there are four — the NVENC 4:4:4 and 10-bit probes in `linux/mod.rs` and both VAAPI probes, which is why the lock has to be shared across the two modules rather than kept local to either. Being RAII also closes a smaller hole the old shape had: the restore was a statement after the open, so any early return added to those functions later would have leaked the quiet level. Now it cannot. The guard is poison-tolerant — a probe that panicked mid-window has already restored the level through `Drop`, so refusing the lock forever afterwards would be strictly worse than proceeding. It is not re-entrant, which is safe here because no probe body reaches another probe (they only call encoder-open helpers); that invariant is written down at the type. Serialising the probes costs nothing measurable: they are process-once behind their caches and each already pays for a real encoder open. Verified with the canonical Linux gate (docker linux/amd64): fmt, clippy --all-targets at default and at nvenc,vulkan-encode,pyrowave, and the pf-encode test leg (37 passed). |
||
|
|
a17413327b |
fix(encode/vulkan): give the vendored Vulkan ABI the layout guard it never had
`vk_av1_encode.rs` and `vk_valve_rgb.rs` are hand-copied `#[repr(C)]` structs handed to the driver through raw `p_next` chains. Nothing in the type system relates them to the C definitions any more, so an edit that inserts, drops, widens or re-pads a field is not a compile error — it is the driver reading our bytes at the wrong offsets, silently. The sibling vendored ABI in `amf.rs` has carried assertions for exactly this reason; these two had none at all. Adds size, alignment and per-field offset assertions for all 19 vendored structs. They are `const` rather than `#[cfg(test)]` (the shape `amf.rs` uses) so they hold in every build including the shipped one, and on any target the modules compile for. Assertions cannot catch a swap of two same-typed fields — the offsets are unchanged — so the field order was diffed field-by-field against the authoritative headers while writing them: `vulkan_core.h` and `vk_video/vulkan_video_codec_av1std_encode.h` from Vulkan-Headers `main` as of 2026-07-25. That diff covered every struct, every `ST_*` structure-type value, every flag bit and both enum groups, and found no drift — the vendored copies are faithful. The one remaining hand-copied table the compiler still cannot see is the bitfield member order inside the three `*Flags` words, where a wrong index means the driver reads `use_superres` where we meant `render_and_frame_size_different`. Three tests pin those to the header by listing the members in C declaration order and asserting the Nth setter writes bit N, with the trailing `reserved` field checked too — a dropped member shifts `reserved` down and fails. Verified with the canonical Linux gate (docker linux/amd64): fmt, clippy `--all-targets` default and `nvenc,vulkan-encode,pyrowave`, and both test legs. |
||
|
|
6c2183ceec |
fix(nvenc): stop telling operators to reboot when the driver version is fine
`NV_ENC_ERR_INVALID_VERSION` is how the driver reports two opposite failures, and we only ever explained one of them. The message told the operator to update the NVIDIA driver or reboot — correct for a genuine header/kernel-module skew, and actively misleading for the field case behind it: a host that streams once per boot and then fails every later session at the caps probe, until the *process* restarts. A version skew is static; it cannot come and go inside one process, so that advice cost the reporter a reboot per stream. Split the two on the only fact that tells them apart: whether a session has already opened in this process. `nvenc_status` gains a `SESSION_OPENED` latch set right after every successful `open_encode_session_ex` — both backends' caps probe and real open, plus the Windows availability probe. No session yet, the version word really is in question and the existing skew advice stands. A session already opened, and the kernel module demonstrably accepted this build's version word, so the message now names per-process driver state and points at the cheap fix (restart the host service, no reboot) plus a request for the log. The load-time gate cannot serve as this discriminator: `NvEncodeAPIGetMaxSupportedVersion` is a pure userspace query, so the classic "updated the driver, didn't reboot" skew sails through it and only fails later at the open. Only a session that actually opened proves the kernel module agreed. The split lives in a pure `invalid_version(bool)` so both halves are unit-tested without touching the process-wide latch. This is a diagnosis change only — it does not fix the underlying field bug, whose root cause is still open. Verified on Linux (192.168.1.25): clippy `-p pf-encode --features nvenc` clean, `cargo test -p pf-encode --features nvenc --lib` 15 passed, rustfmt clean. |