c5eea458b4d4719f8e1263dd6d189e35a77c8d1b
107
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
deeb8b6700 |
feat(pf-encode): build against FFmpeg 9
apple / swift (pull_request) Successful in 1m53s
apple / screenshots (pull_request) Skipped
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 2m34s
ci / web (pull_request) Successful in 2m32s
ci / docs-site (pull_request) Successful in 1m25s
ci / bun-nix (pull_request) Successful in 26s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 3m23s
android / android (pull_request) Successful in 6m47s
ci / rust-arm64 (pull_request) Successful in 8m49s
nix / flake (pull_request) Failing after 16m7s
ci / rust (pull_request) Successful in 23m39s
ffmpeg-next 8.1.0 could not accept FFmpeg 9 at all: ffmpeg-sys-next's version probe
covered avcodec majors 56..62 (the range is exclusive of its end), so libavcodec 63 fell
outside what it knew how to bind. 9.0.0 widens that to 56..63, which is what actually
unblocks Arch. Bump both pins — the unconditional Linux dep and the optional Windows
amf-qsv one — and the lock with them.
No API drift to fix. The crate major is a CEILING, not a target: one source tree still
spans FFmpeg 7.x/libavcodec 61, 8.x/62 and 9.x/63 via per-version cfgs, and every wrapper
symbol the NVENC-libav, VAAPI and amf-qsv backends name survives 8.1.0 -> 9.0.0
unchanged. The three hand-written #[repr(C)] hwcontext mirrors are the parts no compiler
checks, so they were re-read against the real headers rather than trusted:
AVCUDADeviceContext and AVD3D11VAFramesContext are byte-identical across 7.1/8/9, and
AVD3D11VADeviceContext gained two trailing UINTs in 8 that 7.1 lacks — which is why that
mirror deliberately stops at the common prefix, and why its assertions now say what they
do and do not buy you. They pin our layout, not libav's; a green build is not evidence.
The CI image is the step that makes this reach users. arch.yml deliberately runs no -Syu
("the image's snapshot IS the build environment"), so the builder stayed frozen on ffmpeg
8 no matter what Arch shipped, and a canary built from that snapshot could not satisfy the
soname dep the PKGBUILD now derives. Re-keying ci/ rebuilds it against ffmpeg 9.
Ubuntu and Windows deliberately stay put: the noble .deb bundles its own FFmpeg 8 behind
an rpath and strips the libav sonames from its Depends, and Windows bundles BtbN DLLs into
the signed installer — neither is exposed to the break, BtbN publishes no FFmpeg 9 build,
and moving either would re-qualify an encode stack to buy nothing.
Verified end to end on 192.168.1.21 (CachyOS, system ffmpeg 2:9.0-5, RTX 5070 Ti): host
builds clean and links libavcodec.so.63/libavutil.so.61/libavfilter.so.12/libswscale.so.10
with no unresolved sonames; the ffmpeg-8 compat shim is gone and the service runs with
NRestarts=0 and answers 401 on :47990; pf-encode's 67 tests pass; and a live synthetic
encode drives real NVENC hardware through FFmpeg 9's libavcodec to a decodable 1080p HEVC
stream (180/180 frames, FEC loopback 0 mismatches) with libavcodec.so.63 and
libnvidia-encode both mapped into the encoding process.
|
||
|
|
c16e07d746 |
fix(encode/nvenc): AV1 stops shipping half a frame
Every 4K AV1 frame this host encoded reached the wire truncated to its first tile, and had since AV1 was wired up. Measured on .21 (RTX 5070 Ti, 4K60, split AUTO): each access unit carried a frame header declaring two tile rows and a single Tile Group OBU with tg_start = tg_end = 0, so libdav1d rejected 835 of 836 AUs with "Error parsing frame header". NVIDIA's hardware decoder accepts the truncated stream, which is why native Vulkan Video looked healthy at 60 fps while both conformant software decoders — rav1d in-tree and libdav1d out-of-tree — refused every frame and clients fell to a black screen. The two halves of sub-frame readback are armed by different conditions. build_init_params arms the WRITER (enableSubFrameWrite + reportSliceOffsets) from subframe_on alone; the chunked READER additionally requires slices >= 2, and resolve_slices returns 1 for AV1 unconditionally — before the PUNKTFUNK_NVENC_SLICES override is even read, because AV1 partitions via tiles rather than slices. So an AV1 session asked the driver to publish its output tile by tile and then took only the first tile with one blocking lock_bitstream. resolve_split_subframe — the one arbitration point both direct-SDK backends already call — now disarms sub-frame for AV1 and returns split_mode untouched, so AV1 keeps every engine split encode gives it. Arming the reader instead is not a drop-in alternative: poll_chunk cuts at bitstreamSizeInBytes on the reasoning that "slices are contiguous Annex-B", which AV1's OBUs are not. With sub-frame disarmed and split still AUTO, the same session decodes 654/654 frames clean through libdav1d. The test that pinned this as correct (av1_untouched, "both features are legal together") is replaced by one that pins the disarm, and by one that checks the reader's gate against the writer's — the comparison nothing made. The Linux latch comment claiming the two "can't disagree" is corrected; that claim is what made this invisible. |
||
|
|
01294e3a53 |
refactor(pf-encode): WP4 — one split policy, shared with the libav path
The libav NVENC path carried its own inline copy of the split decision and had
already drifted from the direct-SDK selector: it hard-coded a 2-way split
regardless of engine count, and had no depth rule at all. That is the drift the
shared resolver was extracted to prevent, and the copy quietly reintroduced it.
Routing it through `resolve_split_mode` needed the policy to MOVE. `nvenc_core`
is gated on `feature = "nvenc"`, but the libav path is precisely the build where
that feature is OFF (`PUNKTFUNK_NVENC_DIRECT=0`, and the featureless packages --
the packaging gap this project has been bitten by before). So
resolve_split_mode / max_forced_split_mode / clamp_to_engines, plus a new
`forced_split_width`, now live in `codec.rs`, which is always compiled and
already owned SPLIT_FORCE_PIXEL_RATE.
That means the NV_ENC_SPLIT_ENCODE_MODE values had to be hand-written as plain
constants, since the SDK enum does not exist without the feature. They are
therefore pinned: `nvenc_split_constants_match_the_sdk` (feature-gated, the only
place both are visible at once) asserts all five against the real enum, so the
copies cannot rot.
⚠ Only the FORCED outcomes are actionable on the libav side -- libavcodec's
`split_encode_mode` AVOption is its own vocabulary and our DISABLE is the NVENC
enum's 15, which would be meaningless there. DISABLE/AUTO both map to "leave the
option unset", which is exactly today's behaviour (unset = the driver's auto).
`engines = 0` ("not probed") maps to 2-way, preserving what that site always did;
a 3-NVENC part gets the wider split only on the direct-SDK path, which is the one
that actually probes.
⚠⚠ VERIFICATION GAP: .133 went down mid-change (no ping), so the WINDOWS leg is
UNVERIFIED. This matters more than usual -- the Windows backend imported
resolve_split_mode from nvenc_core and that import had to move too, which a grep
caught rather than a compiler. Re-run before trusting it:
cargo clippy -p pf-encode --features nvenc --all-targets -- -D warnings
Verified .21: clippy -D warnings clean BOTH with and without the nvenc feature
(the featureless build is the whole point of the move) and with
nvenc,vulkan-encode; 65 unit tests incl. the new constant-parity test; 25/25
NVENC on-hardware; punktfunk-host clippy clean. fmt clean.
|
||
|
|
1062aa780f |
test(pf-encode): measure the bits/frame curve — no crossover, split always wins
WP0's real deliverable, and the hole every previous measurement in this
programme had. All prior timings ran against driver-zeroed buffers, so rate
control had nothing to code (~300 B/AU against an 833 KB quota) and only the
PIXEL-proportional half of the encode cost was ever exercised -- while the 4K60
HDR field report was a BITS/FRAME problem at 6.8 Mbit/frame.
Adds `pf_zerocopy::cuda::write_plane_from_host`, the exact mirror of the existing
read_plane_to_host. No new loader entry was needed: cuMemcpy2DAsync_v2 was
already in the table and CUDA_MEMCPY2D just needed the reverse memory types.
Linux-only by construction (pf-zerocopy's `imp` is cfg'd to linux).
⚠ Two harness mistakes found and fixed by looking at bytes/AU rather than
trusting the knob:
- Pure per-pixel noise is INCOMPRESSIBLE, so a low bitrate target does not
produce low bits/frame -- it OVERSHOOTS. At a nominal 50 Mbps the encoder
emitted 719 KB/AU against a 104 KB quota, and the three lowest rows of the
first sweep all sat at the same ~5.7 Mbit/frame. Sweeping nominal bitrate
measures nothing.
- So the sweep moves CONTENT DETAIL (block size) instead, and the x-axis is the
bits/frame the encoder ACTUALLY produced, never the one requested.
4K60 HEVC 8-bit, real content, single-engine vs forced-2:
bits/frame Ada 4090 Blackwell 5070 Ti
0.2-0.3 Mb 4567 -> 2381 1.92x 5549 -> 3552 1.56x
~1.1-1.2 Mb 5060 -> 2626 1.93x 5867 -> 4082 1.44x
~3.3 Mb 8478 -> 4455 1.90x 9286 -> 5862 1.58x
~9.6 Mb 16237 -> 8114 2.00x 16435 -> 9275 1.77x
RESULTS. (1) Encode time scales strongly with bits/frame -- 4.6 ms to 16.2 ms
across the range on Ada -- confirming the hypothesis' core claim. (2) There is NO
CROSSOVER: split wins at every point on both architectures (Ada ~1.9-2.0x and
notably flat, Blackwell 1.44-1.77x). So the arbitration's encode-side answer is
essentially always "split", which makes the sub-frame handicap the only decision
that actually matters -- exactly the part already built and unit-pinned.
(3) It corroborates the field capture: at ~6.8 Mbit/frame these curves put
single-engine 4K60 around 10-13 ms, and the field report was 10.3 ms on a 4090.
That reads as real ASIC time, not the retrieve-queue inflation it might have been.
⚠ Caveat the data itself shows: cost is NOT monotonic in bits/frame alone. The
1px row lands at the HIGHEST bits/frame yet encodes FASTER than the 4px row on
both boxes (Ada 10148 vs 16237 us) -- pure noise defeats motion estimation, which
gives up early, where semi-structured content makes it search hard. Content
structure is a real term, so "bits/frame" is a good axis but not a complete cost
model.
Verified .21: clippy -D warnings clean (pf-encode + pf-zerocopy), 64 unit tests,
25/25 NVENC on-hardware. Curves run on both Ada and Blackwell. fmt clean.
|
||
|
|
50b3fd1012 |
fix(pf-encode): drop the 10-bit short circuit — measured wrong on Ada, twice
WP1.3, and the measurement that justifies it. `resolve_split_mode`'s 10-bit rule sat ABOVE the pixel-rate arm and took no codec, so it (D1) vetoed 10-bit 4K120 -- the very case the pixel-rate arm exists for -- and (D2) applied an HEVC-Main10-on- Ada result to AV1 10-bit, which has no such measurement. Both fixed: the pixel-rate arm now comes first, and what remains is codec-scoped to HEVC and only applies BELOW that bar, where a second engine buys nothing anyway. The rule rested on one datapoint: 5120x1440@240 Main10 on Ada, forced-2 7.6 ms vs 2.8 ms single-engine -- split 2.7x SLOWER. Dropping the short circuit flips that exact configuration's behaviour, so it was re-measured on a 4090 (AD102, driver 610.43.03), 400 Mbps, sub-frame pinned off, via a new mode-parameterizable Main10 A/B test (PF_AB_MODE=WxHxFPS reproduces the original operating point). Ada 4090 single forced-2 ratio 3840x2160@60 4483 us 2178 us 2.06x split WINS 5120x1440@240 3689 us 2813 us 1.31x split WINS <- the veto's origin 3840x2160@120 4148 us 2189 us 1.89x split WINS Blackwell 5070 Ti 3840x2160@60 4216 us 2477 us 1.70x split WINS 5120x1440@240 4651 us 3894 us 1.19x split WINS Split wins for Main10 at every mode on BOTH architectures, including the config the veto came from. The original number does not reproduce. ⚠ Caveats, unchanged from the rest of this work: content is trivial (297-300 B/AU against an 833 KB CBR quota -- zeroed VRAM), so this is the pixel-proportional term and the bits/frame regime is still unmeasured; debug build; and the driver differs from whenever the original was taken. Also validated on Ada in the same session -- the whole spike set reproduces on a SECOND architecture and an OLDER driver (610.43.03 vs 610.57.04): S1a in-place split switch accepted with zero IDRs both directions; S1b takes effect (|C-B|=12 vs |C-A|=1921, the cleanest run yet); S1c pair flip passes; D5 confirmed (AUTO+sub-frame 4424 vs DISABLE 4409, 15 us apart -- and AUTO without sub-frame 2310 ~= TWO_FORCED 2314, so the arm stays); engines=2 with THREE_FORCED correctly clamped to mode 2; arbitration converged with exactly 1 keyframe. Verified: .21 clippy -D warnings clean + 64 unit tests; .133 Windows clippy -D warnings clean (the resolver signature grew a `codec` param, so both backends moved); Ada + Blackwell on-hardware as above. fmt clean. |
||
|
|
2366c4fe31 |
feat(pf-encode,host): price the HEVC sub-frame trade so arbitration can cover it
The named next step after WP3's first increment. That increment deliberately REFUSED to arbitrate HEVC-with-sub-frame -- the fleet default, and the reported field case -- because engaging split there gives up sub-frame readback, whose whole value is that the send overlaps the encode. An encoder measuring only encode time would see split as ~2x faster, take it, and make end-to-end latency worse while reporting a win. This supplies the missing number. The real comparison is encode_1eng + send_of_last_slice against encode_2eng + send_of_whole_AU, so the challenger owes roughly spread x (slices-1)/slices. Split across the two sides that can each see half: - Host: new `Encoder::set_send_spread_us` (defaulted, forwarded by TrackedEncoder -- same trap class as set_wire_chunking, and unforwarded it would fail SILENTLY IN THE SAFE DIRECTION, which is the hardest kind to notice). The send thread is the only place a paced send is observed and the encode loop the only place the encoder can be touched, so it goes over an AtomicU32 like encoder_ceiling_kbps, EWMA-smoothed 3:1 per completed AU: one content spike must not flip a verdict that then gets cached. - Encoder: turns the raw spread into the handicap, because only it knows `slices`. SplitArbiter::with_handicap charges it to the challenger before the comparison. A unit test runs identical encode numbers with a cheap and an expensive send and asserts the verdict REVERSES -- with an expensive send the arm that looks twice as fast is a loss end to end, and the incumbent must hold. That is precisely the regression an encode-only arbiter ships. Gate now opens for HEVC+sub-frame only when a spread has actually been reported (and slices >= 2); with no hint it still refuses, so behaviour is unchanged until the host feeds it. Two mechanics this needed: - apply_split_mode became a PAIR flip (split + sub-frame), routed through resolve_split_subframe and restoring from `subframe_opened_with` so a session that never had sub-frame can never gain it. It also recomputes `subframe_chunks`, which reconfigure_bitrate does NOT -- spike S1c's finding; leave it stale and supports_chunked_poll keeps saying yes while numSlices never advances, so poll_chunk busy-polls its whole budget every AU. - The arbiter is now fed from BOTH completion points. A sub-frame session finishes through poll_chunk, so the incumbent arm of an HEVC experiment would otherwise never deliver a sample -- only the challenger, with sub-frame dropped, comes through poll. Verified .21: clippy -D warnings clean for pf-encode AND punktfunk-host with nvenc, 63 unit tests (1 new), 23/23 NVENC on-hardware green. Verified .133: Windows clippy -D warnings clean, zero dead_code. fmt clean. |
||
|
|
3b283dc26e |
feat(pf-encode): WP3 — live split arbitration, measured on the session, no IDR
The fix S1 unlocked. Rather than predict the right split mode at open — which
cannot work, because the decision depends on bits/frame and an Automatic client's
steady-state bitrate is unknown at open (ABR climbs in place afterwards) — the
encoder now measures both arms on the live session and keeps the winner. S1
proved nvEncReconfigureEncoder takes a changed splitEncodeMode with
resetEncoder=0, emits no IDR, and actually applies it, so the experiment is
invisible on the wire.
Deliberately measures instead of modelling: hard-coded per-arch constants are
exactly how the rule this replaces went wrong (one 5120x1440@240 Ada datapoint
generalised into a fleet-wide 10-bit veto). A measurement tracks driver updates
for free.
`SplitArbiter` (pure state machine, unit-tested without a GPU): measure incumbent
-> switch -> SETTLE -> measure challenger -> keep the winner, else switch back.
Verdicts cache per (gpu, codec, mode, depth, chroma) so later sessions open
straight into the winning arm; the key is CeilingKey minus split_mode, since the
split mode is the thing being decided.
⚠ SETTLE_FRAMES=16 is load-bearing, not padding: split-encode does not reach
steady state on the first frame (a FRESH TWO_FORCED session measured early-half
3280us vs late-half 1996), so judging an arm right after switching reads the
transient — intermittently, which would then be cached. A unit test feeds exactly
that transient and asserts the arbiter still sees the steady state.
Safety gates, all correctness conditions rather than preferences: opt-in
(PUNKTFUNK_NVENC_SPLIT_ARBITRATE=1) while it earns trust; an operator
PUNKTFUNK_SPLIT_ENCODE pin always wins; skip if a verdict is already cached; sync
depth-1 only (async_rt.is_none(), same gate chunked poll uses — under pipelined
retrieve the submit->AU span includes queue depth and the comparison is noise);
needs >=2 engines; never H.264.
⚠ And the one that bounds this increment: NO SUB-FRAME TRADE. For HEVC, forcing
split gives up sub-frame readback, which costs send/encode overlap the ENCODER
CANNOT SEE — it measures encode time only, so it would reliably prefer split and
silently make end-to-end latency worse. So arbitration runs only where nothing is
traded: sub-frame already off, or AV1 (both features legal). Pricing that trade
needs the host's send cost and is the next work package.
Challenger choice tests the question worth asking — anything not already the
widest forced split is challenged BY the widest ("are we leaving engines idle?").
The naive "challenge whatever we are not" spent the experiment re-proving that
splitting beats not-splitting, while parking the session on the slow arm to do
it, because 4K60 sits on the fallthrough AUTO.
⚠ Every new nvenc_core item is linux-gated: the arbiter is wired into the Linux
backend only for now and nvenc_core compiles on Windows too. Caught by the .133
check, not by reasoning — the first cut failed Windows clippy with 12 dead_code
errors, the exact item-level trap this file already carries a scar from.
Verified .21: clippy --features nvenc --all-targets -D warnings clean, 62 unit
tests (4 new arbiter tests), 23/23 NVENC on-hardware green including a new
end-to-end convergence test asserting ZERO extra IDRs and a cached verdict.
Verified .133: Windows clippy -D warnings clean, zero dead_code. fmt clean.
|
||
|
|
9a1d8be4cc |
fix(pf-encode): AUTO split is conditional on sub-frame — do NOT retire the arm
Last change's docs concluded "AUTO never splits, retire the arm" from the sub-frame-ON measurement alone. Measured the missing leg before implementing it, and the conclusion was wrong. On .21 at 4K, plain AUTO (env unset, the resolver's fallthrough): sub-frame ON -> 5023/5157 us/frame ~= DISABLE 4979/5000 (does NOT split) sub-frame OFF -> 2401/2352 us/frame ~= TWO_FORCED 2319/2378 (DOES split) So AUTO is CONDITIONAL, not dead. Retiring it would have silently cost every sub-frame-off session its second engine -- a regression introduced while "cleaning up" an arm that looked inert. Split and sub-frame are mutually unsupported for HEVC, so the driver resolves AUTO to no-split only in that combination. Fix is disclosure, not removal: - resolve_split_subframe debug-logs the inert HEVC + AUTO + sub-frame case, which is the fleet default shape: "split_mode=AUTO" has meant "no split" for every default session and nothing said so. Deliberately NOT rewritten to DISABLE -- the mode we pass is what the driver was actually given, and the ceiling-cache key must keep describing that. - New unit test `auto_survives_the_arbitration_in_both_subframe_states` pins the contract so the arm cannot be simplified away later. - The resolver doc now records both measured legs instead of "AUTO is dead". Also in this change: - WP1.6: `resolve_subframe`'s doc said "Windows passes `false`". Stale since the 2026-07-31 .173 A/B flipped Windows to caps-gated default-on. It mattered: it made the AUTO-plus-sub-frame dead combination look Linux-only when it is fleet-wide. - Windows session-ready log parity: split_mode + engines + subframe. The Windows line had no split_mode at all, so a Windows field report could not answer "did this session actually split?" -- the question that started this whole thread. Verified: fmt clean; .21 clippy -p pf-encode --features nvenc --all-targets -D warnings clean, 58 unit tests (1 new), 22/22 NVENC on-hardware tests green; .133 Windows clippy --features nvenc --all-targets -D warnings clean (15m cold, zero errors or warnings) -- the Windows backend is cfg'd out on both macOS and the Linux box, so that leg needed a real Windows host. |
||
|
|
88f29a9411 |
feat(pf-encode): use every NVENC engine the GPU has, not a hard-coded two
WP1.1 plus the engine-count fix. `resolve_split_mode` forced TWO_FORCED at high pixel rate regardless of hardware, so a 3-NVENC part (GB202, AD102 workstation) left a third of its encode silicon idle, and a 1-NVENC part paid a wasted session open to discover it could not split. Probes NV_ENC_CAPS_NUM_ENCODER_ENGINES in both direct-SDK backends' query_caps (the cap is `= 49` in both linux_sys and windows_sys of the vendored SDK 0.4.0 -- the caps enum is cfg-selected per-OS, so that was checked) and latches it on a backend field. NOT on EncoderCaps: nine backends construct that struct as exhaustive literals, so a new field would be a 9-site change of which 7 are unrelated codecs passing a meaningless value, and the only consumer is the resolver. New `max_forced_split_mode(engines)`: 1 -> DISABLE, 2 -> TWO, 3 -> THREE, and >3 -> AUTO_FORCED, because NV_ENC_SPLIT_ENCODE_MODE cannot NAME more than three (NVENCAPI 12.1; values 4..14 are unallocated, so a future API may extend it) and AUTO_FORCED = "split, driver picks how many" is measurably a real split (2.01x vs disabled on .21). 0 = unprobed keeps the historical two-engine assumption. ⚠ WHY THE CLAMP EXISTS, measured on .21 (RTX 5070 Ti, 2 NVENC, 4K HEVC): requesting THREE_FORCED was HONOURED -- session opened in mode 3 -- and ran at 2303 us/frame, identical to TWO_FORCED's 2308. The driver does not reject an over-ask; it silently encodes narrower. So the rejection fallback cannot find the ceiling and PUNKTFUNK_SPLIT_ENCODE=3 on a 2-engine card would have logged a 3-way split over a 2-way encode. Operator overrides are now clamped with a warn. The ordering trap is covered by a test: on a >3-engine part hw_max is AUTO_FORCED (1), which is not "narrower than" TWO_FORCED (2) despite comparing smaller, so a naive min() would collapse a legitimate 3-way request to AUTO. Also adds `engines` and `subframe` to the Linux session-ready log: split_mode alone is ambiguous between "used both engines" and "left a third idle", and since the driver honours an over-wide request the mode cannot be read without the ceiling it was chosen from. This is the line a field report needs. --- and a correction to S1b, in the same change --- Re-running S1b afterwards flipped its verdict to "the driver appears to have IGNORED the in-place split change", contradicting the isolated runs that produced the |C-B|=34 figure already written into the design docs. Investigated rather than re-rolled. The switched leg was landing MIDWAY between the arms (~3600 us against A~5050, B~2300) and the nearest-neighbour verdict flipped on noise. Cause: split-encode does not reach steady state on the first frame -- a FRESH TWO_FORCED session shows it too (early-half 3280 us vs late-half 1996 in one run), so it is split warmup generally, not something specific to reconfiguring in place. A single median over the whole window cannot see that. The test now reports early-half vs late-half and gives a switched leg SETTLE=16 frames before its window opens, every leg the same length. With that, 4/4 runs agree: the switched leg reaches ~2030 us against a fresh-split ~2000 and a single-engine ~4900. ⚠ S1b's CONCLUSION stands (the switch does take effect) but the evidence behind the committed number did not reproduce; the docs are corrected rather than left implying a cleaner result than the harness could support. ⚠⚠ This is a WP3 REQUIREMENT, not just a test fix: a live-session arbitration that switches arms and immediately measures will misjudge the arm it just chose, because the encoder needs ~16 frames to settle. The settle window has to be part of the arbitration, and it is now a measured number rather than a guess. Verified on .21: clippy --features nvenc --all-targets -D warnings clean, 57 unit tests (3 new), all 23 NVENC on-hardware tests green, fmt clean. The 3 failing on-hw tests in a full --ignored run are VAAPI (no AMD/Intel GPU on that box -- their own ignore reason says so), pre-existing and unrelated. |
||
|
|
70b81ac3d7 |
test(pf-encode): S1c + the D5 confirm — pair flips in place, AUTO really is dead
S1c `nvenc_cuda_split_subframe_pair_reconfigure`: the leg S1a/S1b excluded. Both pinned sub-frame OFF to isolate the split variable, but a real HEVC arbitration cannot -- split and sub-frame are mutually unsupported there, so engaging split means flipping enableSubFrameWrite in the same breath, a second init param and the one the reconfigure path deliberately pins. RESULT on .21: the PAIR moves in place, accepted, ZERO IDRs, both directions. It also pins the invariant that makes this safe to build on: `subframe_chunks` is latched ONLY in the init path (~line 1625) and is NOT recomputed by reconfigure_bitrate, so a caller flipping sub-frame in place must clear it too or supports_chunked_poll keeps reporting true and poll_chunk busy-polls its whole budget every AU against a numSlices that never advances. The test performs the correct sequence and asserts the state stays coherent, so WP3 has a worked example rather than a warning. `nvenc_cuda_auto_split_with_subframe`: the D5 confirm -- the one claim in the design's defect list that was only ever inferred. The driver reports no "mode I actually chose", so it is settled by timing, at 4K where the gap is ~2x. RESULT: AUTO (env unset) + sub-frame 4904 us/frame, DISABLE + sub-frame 5062, TWO_FORCED without sub-frame 3464. AUTO sits 158 us from DISABLE and 1440 from TWO_FORCED ⇒ D5 CONFIRMED: plain AUTO does not split while sub-frame is on, so the resolver's AUTO fallthrough reads as "let the driver decide" and means "never split". ⚠ TRAP, hit on this test's first run and now documented in it: the env knob CANNOT express plain AUTO. `0` is DISABLE and `1` is AUTO_FORCED, and resolve_split_subframe counts AUTO_FORCED as forced, so passing `1` silently disarms sub-frame and measures a different configuration entirely -- which produced a spurious "D5 REFUTED". Plain AUTO is only reachable as the resolver's fallthrough with the env unset. The leg now asserts sub-frame resolved TRUE, so the test can no longer answer the wrong question quietly. Verified on .21: clippy --features nvenc --all-targets -D warnings clean, all 4 spikes green, the normal 54-test suite unaffected, cargo fmt --all --check clean. |
||
|
|
4b57d11dd8 |
test(pf-encode): S1 spike — splitEncodeMode CAN change in place, no IDR
Two on-hardware spikes answering the gate on the split-encode engagement
program (design/nvenc-split-encode-engagement-implementation-plan.md).
S1a `nvenc_cuda_split_reconfigure_in_place`: can splitEncodeMode change via
nvEncReconfigureEncoder with resetEncoder=0, without an IDR? Our "reconfigure
must present the SAME init params as the open" rule (windows/nvenc.rs:620) is
our own invariant and had never been tested against a driver. It reports rather
than asserts the verdict -- both outcomes are legitimate findings -- and only
asserts what would invalidate the measurement (session live, engines >= 2, the
arms actually differ). Sub-frame is pinned off so the driver can't reject for
the wrong reason (HEVC forced-split and sub-frame are mutually unsupported).
S1b `nvenc_cuda_split_reconfigure_takes_effect`: the other half -- a driver that
accepts the parameter and quietly ignores it looks identical to one that honours
it. Three legs at 4K (fresh DISABLE / fresh TWO_FORCED / DISABLE->TWO in place);
if C tracks B and not A, the switch is real.
RESULT on .21 (RTX 5070 Ti, GB203 Blackwell, driver 610.57.04):
NV_ENC_CAPS_NUM_ENCODER_ENGINES = 2
S1a: accepted, ZERO IDRs, both directions.
S1b: A fresh DISABLE 5054 us/frame, B fresh TWO_FORCED 2453,
C switched in place 2419 -- |C-B|=34 vs |C-A|=2635. It takes effect,
and split is a clean ~2x at 4K.
Two limits, both recorded in the test docs rather than the commit only. The
frames come out at 427 B/AU against an 833 KB CBR quota: the driver hands back
zeroed VRAM, so the rotated buffers are identical and rate control skip-codes
everything. So this measures the PIXEL-proportional half of the cost only --
the bits/frame regime the field case lives in is untested here, and the test
prints an explicit INCONCLUSIVE-on-content line when it detects that. And this
is Blackwell 8-bit; the Ada Main10 question is untouched.
Verified on .21: clippy -p pf-encode --features nvenc --all-targets -D warnings
clean, both spikes green, cargo fmt --all --check clean.
|
||
|
|
f87c1e6cec |
fix(encode): gate multi-slice frames on the client's decoder — the 0.17.0 Chromecast crash
ci / docs-site (push) Successful in 1m9s
apple / swift (push) Successful in 1m15s
ci / web (push) Successful in 3m16s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 7s
deb / build-publish-client-arm64 (push) Successful in 2m13s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 5s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 7s
ci / rust-arm64 (push) Successful in 3m34s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 5s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 13s
deb / build-publish-host (push) Successful in 4m20s
docker / builders-arm64cross (push) Successful in 6s
docker / deploy-docs (push) Successful in 28s
deb / build-publish (push) Successful in 6m17s
arch / build-publish (push) Failing after 7m58s
flatpak / build-publish (push) Successful in 9m6s
ci / rust (push) Successful in 13m37s
release / apple (push) Successful in 16m32s
windows-host / package (push) Successful in 18m21s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 32s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m21s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m44s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m24s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m27s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 2m56s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 3m54s
apple / screenshots (push) Successful in 20m45s
android / android (push) Successful in 5m0s
Field report: since 0.17.0 a stream to a Chromecast with Google TV 4K freezes on the first frame and ~80% of the time crashes + reboots the DEVICE — with both the Punktfunk app and Moonlight, while an Xbox Series S is fine. Root cause: LN1 Phase 3 ( |
||
|
|
aa070f2d55 |
feat(ffi): hand-mirrored C structs are now layout-checked at compile time
The sharpest memory-safety risk left in this codebase is not an `unsafe` block — it is a hand-written `#[repr(C)]` mirror of an external C struct. Get a field offset wrong and nothing fails to compile and nothing reliably crashes: the library reads a pointer, a length or a pitch out of the wrong bytes. Eleven such structs across five files had NO check at all. Guarded here, each next to the struct it protects: * `AVCUDADeviceContext`, and `AVD3D11VADeviceContext`/`AVD3D11VAFramesContext`. `ffmpeg-sys-next` binds none of them, so these mirrors are the only definitions — and we WRITE through them (`cuda_ctx`, `device`, `bind_flags`). ⚠ The D3D11VA pair is duplicated VERBATIM in two crates (pf-encode's `ffmpeg_win.rs`, pf-client-core's `video_d3d11.rs`) because neither can depend on the other; they must agree with libav and with each other, and now a drift in either is a build error. * The six cuda.h structs. Three were already asserted — but only in `#[cfg(test)]`, so the check ran when someone ran the tests and never in a release build. They are `const` now. The other three, including `CUDA_MEMCPY2D` which is filled on EVERY zero-copy frame, had nothing. * `MsghdrX`, Darwin's `msghdr_x`, which `libc` does not expose. Its layout is not reviewable by eye: the 32-bit fields force padding before each following pointer, so `msg_iov` sits at 16 and not 12. `sendmsg_x`/`recvmsg_x` take the pointer and length from it. * `IPolicyConfigVtbl` — the sharpest of the set. It mirrors an UNDOCUMENTED COM interface, and `set_default_endpoint` is called by SLOT INDEX through a ten-entry `_reserved` gap that carries no names to anchor a review. A field added or resized above it does not break the build; it calls a different function pointer through a mismatched signature. Every assertion is `const _: () = assert!(..)`, so it holds on every build including release and cannot be skipped. The compiler verified the numbers — the sizes and offsets asserted here are the ones the target actually produces, on each platform that compiles the struct. Verified: Linux .21 fmt + both CI clippy steps rc=0 (CUDA + libav CUDA mirrors); Windows .47 full CI clippy set rc=0 + pf-capture tests (D3D11VA pair, COM vtable); macOS `cargo check -p punktfunk-core` (MsghdrX — the only platform that compiles it). |
||
|
|
22936bbc89 |
fix(vaapi): use FFmpeg bt2020nc matrix name
ci / rust (pull_request) Canceled after 0s
ci / rust-arm64 (pull_request) Canceled after 0s
ci / web (pull_request) Canceled after 0s
ci / docs-site (pull_request) Canceled after 0s
ci / bench (pull_request) Canceled after 0s
apple / swift (pull_request) Canceled after 0s
apple / screenshots (pull_request) Canceled after 0s
android / android (pull_request) Canceled after 0s
FFmpeg rejects bt2020 as an out_color_matrix value. Use its canonical bt2020nc name for non-constant-luminance BT.2020, matching the existing AVCOL_SPC_BT2020_NCL encoder VUI. Co-Authored-By: OpenAI Codex <noreply@openai.com> |
||
|
|
8e7ba00d2d |
refactor(encode/linux): third fence off — vaapi.rs, and VaapiHw::new needed no marker at all
Same criterion as the last two: `vaapi.rs`'s sites are pointer dereferences and libav ctx calls,
not ash, so a proof here carries an argument. Three regions, three arguments.
`VaapiHw::new` also loses its `unsafe fn` outright, and the contrast with its CUDA twin is the whole
point: `CudaHw::new` keeps the marker because it is HANDED a `CUcontext` the caller must vouch for,
while this one takes four scalars and opens the VAAPI device itself. Two functions of near-identical
shape, opposite answers, decided by whether a caller can supply something broken.
The other two regions keep their markers (`open_vaapi_encoder`/`_mode` are handed raw
`*mut AVBufferRef`s) and gain the proofs they lacked. The encoder-config block now records the fact
that makes it sound rather than obvious: `av_buffer_ref` returns a NEW reference the codec context
adopts, so the callee shares the caller's device/frames buffers instead of consuming them — which is
also why the low-power entrypoint ladder can retry with the same two pointers after a failed attempt.
The frames-pool block reuses the `CudaHw::new` argument: alloc returns null-or-initialized and
`AvBuffer::from_raw` rejects null, so the `?` leaves before any field store can run.
One call site dropped its `unsafe {}` and its comment with it — the comment argued that libav was
initialized, which is a real precondition of the call but not one a caller can violate, so it is now
a note rather than a contract.
14 fenced files -> 11. Verified on .21 (fmt + both CI clippy steps rc=0) and, because this is the
live VAAPI construct path rather than dead code, ON THE AMD 780M (.116, pf-build distrobox):
`vaapi_cpu_encode_smoke`, `dmabuf_inner_alloc_drop_cycles` and `vaapi_probe_smoke` all pass —
3 passed / 0 failed, H265 + AV1 probes still true in both 8- and 10-bit.
|
||
|
|
d72822ced7 |
refactor(encode/linux): second fence off — CudaHw::new's pointer walk gets its proof
`linux/mod.rs`'s fifteen sites are the same kind as `video_vulkan.rs`'s, not the ash kind: nine raw pointer dereferences and six libav calls, all inside `CudaHw::new`, which had no `unsafe` block and therefore no proof of the one thing worth proving here — that the pointer chain it walks is live. The marker STAYS (`cu_ctx: *mut c_void` is a `CUcontext` the caller must supply valid). The body is now two blocks, one per phase, because there are two distinct arguments to make. Both turn on the same non-obvious fact: `av_hwdevice_ctx_alloc`/`av_hwframe_ctx_alloc` return null or a ref whose `data` libav has ALREADY initialized, and `AvBuffer::from_raw` rejects null — so the `?` leaves before any of the field stores below it can run. That is what makes the `(*dev_ctx)`/`(*fc)` writes in-bounds stores on live allocations rather than a hope, and it is exactly the reasoning that was missing. The device block also records the ordering constraint that was implicit: `cuda_ctx` must be stored BEFORE `av_hwdevice_ctx_init`, which reads it. Two files now need no exemption: 14 fenced -> 12. Both were removable for the same reason — their sites are pointer dereferences, where a proof carries an argument, unlike the ash backends where it could only restate the call. That is the criterion for which fence to attack next, not file size. Verified on .21: fmt + `clippy --workspace --all-targets -- -D warnings` + the feature-gated `-p pf-encode --features nvenc,vulkan-encode,pyrowave` step, all rc=0 with no allow in either file. |
||
|
|
6de325a6b6 |
fix(ci): the unsafe lint said warn while CI enforced it as deny, and main went red
windows-host / package (push) Failing after 13m8s
windows-host / winget-source (push) Skipped
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m43s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m31s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 4m48s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 5m55s
ci / web (push) Successful in 1m21s
ci / docs-site (push) Successful in 1m5s
android / android (push) Failing after 6m25s
ci / bench (push) Successful in 8m21s
deb / build-publish (push) Successful in 8m44s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
decky / build-publish (push) Successful in 32s
deb / build-publish-host (push) Successful in 10m2s
ci / rust (push) Failing after 17m18s
ci / rust-arm64 (push) Successful in 16m6s
arch / build-publish (push) Failing after 17m19s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 24s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 7m54s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 7m44s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11m12s
deb / build-publish-client-arm64 (push) Successful in 14m5s
flatpak / build-publish (push) Failing after 8m59s
apple / swift (push) Failing after 7m20s
apple / screenshots (push) Skipped
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 17m16s
docker / build-push-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 26m3s
`unsafe_op_in_unsafe_fn = "warn"` was adopted workspace-wide in 39513528 on the assumption that
`warn` is a soft setting you can clear at leisure. It is not: ci.yml runs `cargo clippy … -D
warnings`, which promotes it to a hard error, so main has failed on EVERY commit since — Linux
`rust` and `rust-arm64` both dying on `pf-client-core` with 70 E0133 errors, and windows-host.yml
alongside them. A lint level that understates its own severity is worse than a strict one, so this
states what CI already does — `deny` — and writes the exemptions down instead.
Fourteen GPU/FFI backend files take `#![allow(unsafe_op_in_unsafe_fn)]`, each with its reason and
the workspace Cargo.toml carrying the argument in full. They are not "not done yet": measured
across them, 64% of the sites are a single third-party FFI call (ash, pyrowave-sys, libav, the
NVENC/AMF entry tables), and of the 44 `unsafe fn`s only 4 have a body containing no unsafe
operation at all. Since pf-encode also denies `undocumented_unsafe_blocks`, narrowing them means a
hand-written SAFETY comment per line that could only restate the signature — the exact noise that
made `unsafe` stop meaning anything here before. Everything else stays at zero and enforced, and
each allow is removable on its own terms.
Two smaller things this had to clear, both invisible to the job that would have caught them:
- `service.rs`: `undocumented_unsafe_blocks` wants the proof on EACH block, and a comment covering
a group of consecutive `unsafe` statements only credits the first — so the two `OwnedHandle`
wraps became their own statements. Windows-gated, so only the .47 gate sees it.
|
||
|
|
a3843b6996 |
test(encode): the VAAPI dmabuf path verified on AMD silicon, and its owner-only fields say so
Closes the last unverified leg of the AvBuffer work. `DmabufInner` was the heaviest
ownership change in the crate — four owned objects (DRM device, derived VAAPI device,
DRM-PRIME frames ctx, filter graph) whose eight failure branches each repeated the
same four-line unwind, once inside a macro, plus a ninth copy in `Drop` — and it had
no coverage at all. `vaapi_cpu_encode_smoke` does not reach it: that drives the
swscale/CPU-upload path, which uses `VaapiHw` and never builds a graph.
`dmabuf_inner_alloc_drop_cycles` loops construct/drop, which is the whole contract
now that every handle releases itself. **It passes on a Radeon 780M**, as do the two
pre-existing VAAPI tests, so `VaapiHw` and `DmabufInner` are both hardware-verified
rather than compile-only.
Also fixes three more owner-only fields the AMD box surfaced: `graph`,
`vaapi_device` and `drm_device` are `never read` since the hand-written `Drop` that
used to read them is gone. Same call as the decoders and the QSV pair — annotate
rather than delete (removing `graph` would free it while `src`/`sink` still point
into it) or underscore-rename (which hides what they hold). `drm_frames` is NOT
annotated: `submit` genuinely reads it per frame.
I had missed these twice: my Linux log greps searched "never constructed|never used",
which does not match "never read". The check now includes all three spellings.
Two environment notes worth keeping, both diagnosed by testing OUTSIDE our code
first (ffmpeg's own CLI reproduced each):
* Inside a distrobox on an immutable host, VAAPI needs
`LIBVA_DRIVERS_PATH=/run/host/usr/lib64/dri`. The container's mesa (25.3.6) is
older than the host's (26.0.4) and every encoder open fails with a bare ENOSYS —
"Function not implemented", naming nothing. The test doc says so.
* `cargo check --workspace` in that container fails on `glib-sys` (no GTK dev
headers). Unrelated to this branch; .21 checks the full workspace clean.
Verified: AMD .116 (Radeon 780M, Mesa 26.0.4) full pf-encode suite 33 passed / 0
failed plus all 3 ignored VAAPI hardware tests green. Linux .21 workspace check exit
0 / zero errors, pf-encode 33/0, pf-client-core 34/0, and zero dead-code warnings
across all three spellings.
|
||
|
|
eb9c5be20d |
refactor(encode): the VAAPI dmabuf path stops unwinding by hand
`DmabufInner::open` builds four owned objects — a DRM device, a VAAPI device derived from it, a DRM-PRIME frames context, and a filter graph — and every one of its eight failure branches unwound them by hand. The same four-line block (`avfilter_graph_free` + three `av_buffer_unref`s) appeared eight times, once *inside a macro*, plus a ninth copy in `Drop`. Adding a step to that function meant remembering to extend the unwind at exactly the right depth; getting it wrong leaks a device per failed session (the persistent listener accumulates them) or frees one twice. All eight are gone. `AvBuffer` already owned the buffer refs; `AvFilterGraph` now does the same for the graph, so each handle is owned the moment it exists and an early `bail!` releases whatever was built so far. `open` lost ~40 lines of cleanup and gained none. Field order in `DmabufInner` is load-bearing and says so: graph, frames, VAAPI device, DRM device, then `enc` LAST. Fields drop in declaration order, and that sequence reproduces the old hand-written `Drop` exactly — including that it ran ahead of every field, so all four were released before ffmpeg-next dropped the encoder. Everything here holds its own reference, so refcounting makes any order sound; the ordering is pinned so a future reorder cannot quietly change what ships. One subtlety preserved deliberately: the buffersrc parameters take `drm_frames` BORROWED, not ref'd (`av_buffersrc_parameters_set` takes its own ref). That is now `drm_frames.as_ptr()` — same borrow, same single owned ref, no new leak. Verified on .21 (CachyOS, RTX 5070 Ti, FFmpeg 62): `cargo check -p pf-encode --all-targets` clean at exit 0, `cargo test -p pf-encode` 33 passed / 0 failed, and `cuda_hw_alloc_drop_cycles` still passes against real CUDA. vaapi.rs now contains zero `av_buffer_unref` and zero `avfilter_graph_free` calls, down from 40 and 8. The dmabuf path itself still needs AMD/Intel silicon to exercise end to end. |
||
|
|
7adc1db672 |
refactor(encode): AVBufferRef ownership moves into an RAII handle
`CudaHw::new` and `VaapiHw::new` are the same shape — alloc a hwdevice, deref it to fill fields, init, alloc a frames ctx, deref, init — and both unwound by hand: `av_buffer_unref` on every failure branch (three in CUDA, two in VAAPI) plus a hand-written `Drop` repeating the pair. That shape has two failure modes and the compiler can see neither: add a branch and forget the cleanup (leak), or let two cleanup paths run (double-unref, an abort inside glibc). `libav::AvBuffer` owns the ref instead. It null-checks on the way in — the check each caller open-coded — and unrefs exactly once on drop, so an early `?` releases whatever was built so far and the failure branches carry no cleanup at all. Both hand-written `Drop` impls are gone, and `linux/mod.rs` goes from seven `av_buffer_unref` calls to zero. The one subtlety, called out at both structs: these two fields must stay declared frames-BEFORE-device. Fields drop in declaration order, and the code being replaced deliberately unref'd frames first (a frames ctx holds its own reference on its device). Refcounting makes either order sound, but a field reorder should not silently change what ships, so the ordering is load-bearing and commented as such. Also adds `cuda_hw_alloc_drop_cycles` — the RAII path had NO test coverage: the NVENC smoke tests take the CPU path and never construct a `CudaHw`, and the VAAPI twin's tests need AMD/Intel silicon. Looping construct/drop is what catches the double-unref (abort) and the leak (allocator growth) this refactor is about. Verified on Nobara (RTX 5070 Ti): `cargo check -p pf-encode --all-targets` clean at exit 0, `cargo test -p pf-encode` 33 passed / 0 failed, and `cuda_hw_alloc_drop_cycles` passes against a real CUDA device — eight construct/drop cycles, no abort. `VaapiHw` is compile-verified only; it is the identical shape but no AMD/Intel box was reachable to run its two ignored tests. |
||
|
|
3d3ecf1e82 |
test(encode/nvenc): the NVIDIA HDR leg, verified on an RTX 5070 Ti
The NVIDIA half of "zero-copy, no host CSC, 10-bit HDR" had never met a GPU. Two `#[ignore]`d smokes now drive it, run on `.41` (Bazzite f43, RTX 5070 Ti, driver 595.58.03): * `nvenc_cuda_hdr10_packed_rgb` — a packed 2:10:10:10 CUDA payload straight into NVENC as `ARGB10`, HEVC **and** AV1. Asserts what would catch a mislabelled stream: the encoder DERIVED 10-bit and HDR from the input format rather than being told, and picked `ARGB10` for `X2Rgb10`. That derivation is what selects Main10 / AV1-at-10 and the BT.2020 PQ signalling. * `nvenc_cuda_hdr10_cursor_blend` — `cursor_blend.comp` MODE 3/4, the 10-bit channel-unpacking twin. Asserts the blend targets `SlotFormat::X2Rgb10` and not the 8-bit layout, which would tint the pointer and shift its channels. Both green, and the dumped bitstreams decode 0-error reporting `Main 10` / AV1 `Main`, `yuv420p10le`, `bt2020nc` / `smpte2084` / `bt2020`. There is no host colour conversion anywhere on this path: the frame arrives as a LINEAR dmabuf, crosses to CUDA through the Vulkan bridge, and NVENC's ASIC does the BT.2020 conversion following the VUI the session configured. The "no CSC" property AMD gets from the EFC, NVIDIA gets from the encoder itself. The whole pre-existing `nvenc_` suite was re-run alongside them — 16/16, no regression from the depth-derivation and buffer-format changes underneath it. Capture half checked too: the patched gamescope binary built on the AMD box runs unmodified on the NVIDIA one (same Fedora 43 base) and its node offers the same four formats with the colorimetry props — so headless gamescope + 10-bit PQ is not AMD-specific. |
||
|
|
576bf7e294 |
test(encode/vulkan): 10-bit smokes — and they pass on a real AMD GPU
The Vulkan Video 10-bit path was the least-exercised code on this branch:
nothing had ever created a Main10 video session, so the profile query, the
`G10X6…3PACK16` picture allocation, the scratch→plane copy, the hand-packed
AV1 sequence header and the colour signalling were all first-run-on-glass.
Three `#[ignore]`d smokes now drive them, and they were run.
`.116` (Bazzite f43, AMD 780M, RADV / Mesa 26.0.4), all three green:
* `vulkan_smoke_10bit` — HEVC Main10 through the compute CSC;
* `vulkan_smoke_10bit_av1` — AV1 at 10 bits;
* `vulkan_smoke_rgb_10bit` — HDR with NO host CSC, the EFC converting BT.2020
off the packed 10-bit source. It did **not** soft-skip, which answers the one
capability in this whole feature I had only ever read in a registry: RADV's
VCN EFC really does advertise `MODEL_YCBCR_2020` and accept a 10-bit
packed-RGB encode source.
Every stream decodes 0-error and reports `yuv420p10le` + `bt2020nc` /
`smpte2084` / `bt2020`. The AV1 one parsing at all is the load-bearing result
there: `high_bitdepth` precedes the CICP bytes in `color_config()`, so a wrong
bit would have thrown every later field out of phase rather than merely
mislabelling the depth.
Round-trip on solid frames, fed (160,160,800) as 10-bit codes:
compute CSC -> (159, 158, 796)
EFC -> (159, 158, 796)
AV1 -> (158, 159, 794)
Under 0.5%, all of it limited-range quantisation and lossy encode. The compute
and EFC results being IDENTICAL is the strongest check available: two
independent BT.2020 NCL implementations — my shader and AMD's fixed-function
block — agreeing to within rounding.
The smokes deliberately assert structure (submits encode, AU count, the depth
the encoder settled on), not colour: a shader writing the 10 bits into the
wrong end of the word still produces a decodable stream. Colour is the dump +
ffmpeg round-trip above, which is what actually caught nothing this time.
Also corrects a comment: I claimed a wrong store factor was a "~1.6% luminance
error". It is not. The `<< 6` PLACEMENT is the load-bearing part (dropping it
is 64x too dark); `1/1023` vs `64/65535` is ~0.1% and harmless.
|
||
|
|
cb4690f216 |
feat(encode/vulkan): probe the device instead of guessing — AV1 10-bit + zero-CSC HDR
Three fixes to the same mistake: deciding what the Vulkan Video backend can do
from a table in our heads rather than from the driver, and routing everything
that didn't fit to libav VAAPI — where a session loses real RFI recovery and
the cursor blend for no reason the hardware asked for.
**Capability probe, per codec AND depth.** `probe_encode_support`'s "is there
an encode queue" boolean becomes `VulkanEncodeCaps { supported, eight_bit,
ten_bit }`, answered by `vkGetPhysicalDeviceVideoCapabilitiesKHR` against the
very profile chain the session open builds. So the dispatcher's prediction
cannot disagree with reality: a capable device keeps the Vulkan path, an
incapable one routes to VAAPI BEFORE burning a failed open, and the
cursor-blend mirror stays honest for free. This is the shape the direct-SDK
NVENC path already uses for its codec GUIDs.
**AV1 10-bit.** It was excluded on a guess about driver coverage; now the
device answers. `color_config()` carries `high_bitdepth` + the BT.2020/PQ CICP
triplet in both the `StdVideoAV1ColorConfig` and the sequence-header OBU we
bit-pack ourselves — they must stay identical or the driver's frame OBUs parse
against a header we didn't write. `high_bitdepth` sits BEFORE the CICP bytes,
so getting it wrong doesn't just mislabel the depth, it puts every following
field one bit out of phase; the new test reads the packed bits back.
**Zero-CSC RGB-direct in HDR.** The EFC probe assumed BT.709 and BGRA. It now
asks for the model this session's colourimetry needs (`MODEL_YCBCR_2020` for
10-bit — the extension has always had it) and for the CAPTURED format as an
encode-source format, and the session create-info selects the matching model.
An HDR session with no pointer to composite therefore hands the captured
buffer straight to the fixed-function front end and runs no host CSC at all.
Sessions that DO composite a pointer keep the compute CSC, unchanged: the EFC
cannot blend, and that rule outranks everything.
Also: `can_encode_10bit` on AMD/Intel now reports the union of VAAPI's and
Vulkan Video's answers instead of VAAPI's alone. `open_amd_intel` tries Vulkan
first and falls back, so either one being able to encode Main10 makes the
session 10-bit-capable — answering `false` because only one of them said yes
stranded encodable HDR sessions at 8 bits.
|
||
|
|
479f0965ee |
feat(encode/vulkan): Vulkan Video encodes 10-bit, so AMD/Intel HDR keeps the good path
The Vulkan Video backend was 8-bit for no structural reason — the API has `VK_VIDEO_COMPONENT_BIT_DEPTH_10_BIT` and `PROFILE_IDC_MAIN_10` in the very fields this pinned to 8 and MAIN, and AMD VCN and Intel both encode Main10. It was six hardcoded sites, and the cost of leaving them was paid twice over: an HDR session had to take libav VAAPI, losing real RFI loss recovery AND the compute CSC's cursor blend — which on gamescope is the only way the pointer reaches the stream at all, since gamescope has no embedded-cursor mode. An HDR session now opens a Main10 profile with 10-bit component depths, a `G10X6_B10X6R10X6_2PLANE_420_UNORM_3PACK16` picture + DPB, and an SPS carrying `bit_depth_*_minus8 = 2` with the BT.2020/PQ CICP triplet instead of BT.709. `rgb2yuv10.comp` is the CSC's twin, and the two interesting parts of it are: * it is a PURE 3x3 matrix. The samples arrive already PQ-encoded (gamescope composites into the PQ container), so BT.2020 NCL applies to the code values as they are — there is no transfer function to apply here and applying one would be wrong; * the scratch planes are `R16`/`RG16`, not the picture's plane formats. The 10-bit ycbcr plane formats are not storage-image formats, so the shader writes the value into the HIGH bits by hand (`code10 << 6`, hence the `64/65535` factor and not `1/1023`) into planes that are merely SIZE-compatible with the picture's — which is all `vkCmdCopyImage` requires. Scope and safety: * HEVC only. AV1 10-bit encode has far thinner driver coverage, and a session open is not the place to gamble on it — those stay on VAAPI, as does a device that fails the Main10 profile query inside the open (the pre-existing "failed Vulkan open falls back to VAAPI" net, no new probe needed). * HDR pins the compute-CSC arm over the EFC RGB-direct one, which the EFC could not serve anyway: its fixed-function conversion is 8-bit BT.709 narrow with no knob for BT.2020. * `open_inner` binds `hdr` to the parameter-set HEADER bytes, so the depth flag is `ten_bit` there — the one name collision this change had to route around. |
||
|
|
17f824c3e9 |
feat(encode/nvenc): an HDR capture stays zero-copy on NVIDIA
AMD/Intel needed no new encoder code for HDR — the VAAPI path already ingests an XR30 dmabuf into `format=p010:out_color_matrix=bt2020`. NVIDIA did: the 10-bit formats were excluded from the GPU import outright, so an HDR session fell back to a CPU readback plus swscale, which is the one thing the capture path is not allowed to ship. It turns out no CSC kernel is needed. NVENC ingests packed 10-bit RGB natively as `ARGB10`/`ABGR10` and does the conversion itself following the configured VUI matrix — which `apply_low_latency_config` already sets to BT.2020 NCL for an HDR session. So the frame travels LINEAR dmabuf → Vulkan bridge → CUDA → NVENC unconverted: no host CSC pass, no depth loss, no extra work on a contended SM. * invariant 1 is restated rather than dropped: HDR must never take the TILED EGL de-tile blit (it renders into an 8-bit `GL_RGBA8` texture). The HDR pods are LINEAR-only by construction, so the plan may build the importer; the per-frame gate — which sees the negotiated modifier the plan cannot — is what enforces the tiled half, and falls back to the CPU path if a producer ever ignores our offer. * …but only where the encoder can actually take the payload (`linux_hdr_cuda_ok`). libav's HDR route builds a P010 hardware frames context and swscales into it, so on a host without the direct-SDK backend a packed-2:10:10:10 CUDA buffer would land in a P010 surface as garbage. Those keep the CPU path. * `nvenc_cuda` stops pinning 8-bit/SDR. Depth and HDR now follow the INPUT format, like the Windows backend: a 10-bit session whose capture came back 8-bit encodes AND labels 8-bit rather than mislabelling. * the cursor-blend compute shader gains two 10-bit modes, so the pointer gamescope leaves out of its node survives the HDR path. Same display-referred blend the CPU path's `composite_cursor_rgb10` already does — the samples are PQ, and a real sRGB→PQ cursor LUT is polish, not correctness for a pointer. |
||
|
|
28c50d1c5b |
fix(encode): every backend signals its colour, so no decoder has to guess
audit / cargo-audit (push) Successful in 2m38s
audit / bun-audit (push) Successful in 13s
ci / rust (push) Failing after 12s
ci / web (push) Successful in 1m4s
ci / docs-site (push) Successful in 1m8s
ci / bench (push) Successful in 6m59s
ci / rust-arm64 (push) Successful in 10m2s
android / android (push) Successful in 13m6s
decky / build-publish (push) Successful in 23s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
arch / build-publish (push) Failing after 14m19s
deb / build-publish (push) Successful in 11m13s
deb / build-publish-host (push) Failing after 4m36s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 44s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m6s
docker / build-push-arm64cross (push) Successful in 11s
docker / deploy-docs (push) Successful in 35s
windows-host / package (push) Failing after 7m59s
windows-host / winget-source (push) Skipped
deb / build-publish-client-arm64 (push) Successful in 7m11s
apple / swift (push) Successful in 5m17s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m21s
flatpak / build-publish (push) Successful in 6m43s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m30s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 4m57s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 6m23s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m6s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m14s
apple / screenshots (push) Successful in 22m57s
Three encode paths shipped a bitstream with no colour description at all, leaving primaries/transfer/matrix/range "unspecified": - Vulkan Video HEVC (`vk_build.rs`) built an SPS with no VUI whatsoever. This is the DEFAULT backend for AMD/Intel Linux hosts on HEVC/AV1. - Vulkan Video AV1 packed `color_description_present_flag = 0`. - The openh264 software path wrote nothing (it converts BT.709 limited and relied on decoders defaulting to that). - The libav-NVENC Linux path excluded packed-RGB 4:2:0, on the belief that "NVENC's internal CSC writes its own VUI". It doesn't: libavcodec derives `colourDescriptionPresentFlag` from the AVCodecContext colour fields, so leaving them unspecified emits none. Reachable on a CPU/dmabuf capture, a build without `--features nvenc`, or PUNKTFUNK_NVENC_DIRECT=0. Unsignalled looks fine on every punktfunk client — `csc_rows` falls back to BT.709 on "unspecified" — which is why this survived. Vendor TV decoders do not: they guess colorimetry from RESOLUTION, and an LG webOS panel reads a 4K SDR stream as BT.2020 and renders it visibly washed out. All four now signal BT.709 limited, which is what every host CSC actually produces (`rgb2yuv.comp`, `convert_bt709`, the swscale paths) and what the Welcome's `ColorInfo::SDR_BT709` already advertises out-of-band. NVENC, VAAPI, QSV, AMF and the Windows libav path were already correct. Two tests, both parsing the REAL emitted bitstream rather than re-asserting the constants: an independent bit-walk of the AV1 sequence header (the packed OBU must stay identical to the `StdVideoAV1ColorConfig` handed to the driver), and an H.264 SPS/VUI parse proving openh264 honours the request instead of dropping it. Not yet verified on hardware: the HEVC VUI depends on the driver's SPS writer emitting `vui_parameters()`. PUNKTFUNK_VULKAN_ENCODE=0 falls back to VAAPI if a driver mishandles it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> (cherry picked from commit 3c56ff5717b2c9a0871953127da3dadd6a84220d) |
||
|
|
188f55d3b1 |
fix(encode/nvenc): the host advertises what the driver lists, not a superset
Every NVIDIA host advertised a static H.264|HEVC|AV1 superset, so a 1st-gen Maxwell (GTX 960M, no HEVC/AV1 encode) offered HEVC — a client that believed it got ~15 s of blank video and a disconnect instead of a stream. Both OSes now ask the driver itself (nvEncGetEncodeGUIDs) on one throwaway direct-SDK session: Linux on the shared CUDA context, Windows on the selected render adapter, wired into host_wire_caps AND the GameStream serverinfo mask (which had been left on the superset for NVIDIA on both OSes). Fails open — an unanswerable probe keeps the historical superset, so it can only ever narrow the advertisement to codecs the GPU really encodes. The HEVC 4:4:4 answer rides the same session on Linux instead of opening a libav hevc_nvenc FREXT probe: that open is the prime suspect for the field bug where one probe wedges NVENC process-wide (NV_ENC_ERR_INVALID_VERSION on every later session until a host restart), and the direct backend re-checks the same caps bit at session open anyway. The ffmpeg probe remains only for hosts that really stream over libav (PUNKTFUNK_NVENC_DIRECT=0 or a build without the nvenc feature), where ffmpeg's NVENC client runs regardless. The 10-bit probe deliberately stays libav — Linux HDR rides the libav P010 path. On-hardware: .136 (RTX 5070 Ti) 14/14 nvenc tests in one process incl. the probe followed by real sessions and dirty teardown; .173 (Windows RTX) probe + 47 release lib tests. The Windows probe test documents the pre-existing MSVC debug-link failure (LNK2019 via the sdk crate's unused lazy loader) — run it with --release, the same reason windows-host.yml gates with clippy. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> (cherry picked from commit 0346ec8090568eb499e8cb7d735305b28471185e) |
||
|
|
6d1baa0add |
fix(pf-encode/pyrowave): the bitrate pin holds on the WIRE, not the raw bitstream
A datagram-aligned PyroWave session inflates the codec bitstream ×1.2–1.3 on its way to the wire — greedy packing of few-hundred-byte atomic block packets into 1408 B windows zero-pads most window tails, plus the 4-byte prefixes and FRAG chains. The 2026-07 field report's 1440p60 10-bit "Automatic" pin of 407 Mb/s put a measured 550 Mb/s on a 1 GbE link; nothing enforced the pin past the rate controller. New shared WireBudget (pyrowave_wire.rs, both backends): tracks the real per-frame AU/bitstream ratio as a ×1024 fixed-point EMA (prior ×1.25, weight 1/8, clamped ×1.0–×2.0) and deflates the budget handed to pyrowave's rate control by it, so the windowed AU lands on the configured rate. Sealed-datagram framing (+4.5%) and FEC parity stay uncompensated — H.26x sessions carry those on top of the configured bitrate too, and the pin must mean the same thing for every codec. Dense (non-chunked) sessions are untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
dff63b2a29 |
fix(encode/nvenc): a visible cursor no longer serializes submit — the blend goes stream-ordered
windows-host / winget-source (push) Canceled after 0s
windows-host / package (push) Canceled after 2m32s
android / android (push) Canceled after 39s
apple / swift (push) Canceled after 0s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 34s
ci / rust (push) Canceled after 24s
ci / rust-arm64 (push) Canceled after 24s
ci / bench (push) Canceled after 5s
ci / web (push) Canceled after 24s
ci / docs-site (push) Canceled after 8s
deb / build-publish-client-arm64 (push) Canceled after 0s
deb / build-publish (push) Canceled after 2s
deb / build-publish-host (push) Canceled after 1s
decky / build-publish (push) Canceled after 4s
docker / build-push-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Canceled after 5s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Canceled after 5s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 5s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 2s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 2s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 2s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 1s
Report: iPad on a gamescope/NVIDIA 120 fps session capped at ~80 fps with repeat_fps 0, zero loss, capture 0 µs, ASIC 15 µs — and submit p50 at 10.2 ms, ~81 % of the loop period. Under gamescope the host composites the live pointer into EVERY frame, and a cursor-bearing frame forced the CPU-synced submit path: a blocking CUDA copy plus a fence-waited Vulkan blend, both exposed to the running game's GPU load. The "games hide the cursor" assumption the gate relied on does not hold under gamescope. The blend is now stream-ordered end to end. VkSlotBlend exports a timeline semaphore (VK_KHR_timeline_semaphore + external_semaphore_fd) into CUDA (cuImportExternalSemaphore, new dlopen entries): the enqueued copy signals it on the encode thread's copy stream, the blend submission waits for and advances it on the Vulkan queue, and a CUDA-side wait orders the encode after the blend on the session's bound IO stream — no CPU sync anywhere, so cursor frames keep the stream-ordered fast path. Each ring slot gets its own command buffer + descriptor set (written once) so several ordered blends can be in flight; cursor-bitmap uploads and teardown quiesce through the timeline. Drivers without the timeline export keep the previous CPU-synced blend, and any bring-up or per-frame failure still degrades to "no cursor", never a dropped frame. Also: the blocking multi-plane copies (the escalated/pipelined mode and the non-stream-ordered fallback) now enqueue every plane and pay ONE stream sync instead of one per plane (NV12 2→1, YUV444 3→1) — each exposed wait costs scheduling latency under GPU contention, which is what makes the escalation's blocking copies self-reinforcing. Verified on the RTX 5070 Ti box (driver 610.43.03): all 12 nvenc_cuda on-hardware smokes green, including the new nvenc_cuda_cursor_blend_stream_ordered (6 cursor AUs, all ordered, across a bitmap-serial flip); host suite 301/301; clippy --all-targets -D warnings clean; struct layouts of the hand-flattened cuda.h params asserted in tests. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fc335b39e9 |
fix(host/encode): negotiate the cursor around what the encoder can blend
EncoderCaps::blends_cursor's contract said the HOST must fall back to capturer-side compositing when a cursor-as-metadata session lands on an encoder that can't composite — but that host half was never built: open_video warned and the session streamed WITHOUT a pointer (confirmed on the VAAPI dmabuf and libav-NVENC CUDA paths; latent on vulkan RGB-direct/native-NV12). The negotiation is now caps-aware, ahead of capture, on both planes: * pf-encode grows cursor_blend_capable() — the pre-open dispatch mirror (sibling of linux_native_nv12_ok) answering whether the resolved backend composites frame.cursor; its pure core is test-pinned arm by arm. * Native plane: handshake::cursor_forward grants the cursor channel only where the resolved backend can blend (the capture-mouse flip makes the host draw the pointer on demand); denied sessions keep the pre-channel path — the compositor EMBEDS the pointer, never cursorless, never doubled. The Welcome's HOST_CAP_CURSOR bit is computed once and read back at both session-wiring sites instead of recomputed. SessionPlan::output_format additionally keeps every cursor-blend session off producer-native NV12 (the arm with no CSC to fold a cursor into), and vulkan RGB-direct now yields to a cursor-blend session even when pinned (EFC cannot composite; the open logs the override). Windows plans cursor_blend=false via the new shared cursor_blend_for() rule — the IDD capturer composites the pointer itself, and asking the encoder anyway fired the blends-cursor warn spuriously on every cursor-channel session. * GameStream plane: the hardcoded cursor_blend=true is gone. The portal source asks for cursor-as-metadata only when the resolved backend blends, otherwise negotiates an Embedded pointer (choose_cursor_mode's new ladder); the capturer pool now also keys on that mode. The virtual-output source passes false — its capture embeds the pointer where it can. The per-arm warns in vulkan_video (RGB-direct, native-NV12) are now structurally unreachable and removed. open_video's post-open check stays as the single backstop for what planning cannot see: a Vulkan-open falling back to VAAPI mid-session, and the gamescope residual (no embedded mode exists there, so a never-blending backend — H.264-on-AMD VAAPI, software — still streams cursorless; fixing that needs a compositing stage, deliberately not built in this pass). Zero-copy is preserved throughout — every fallback is a capture-negotiation change, never a readback. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f495b201e1 |
fix(encode): delete the write-only EncoderCaps::supports_hdr_metadata
A caps field nothing reads is a contract nobody honors — and this one shipped write-only: its single reader anywhere in the workspace was a hardware-gated assertion inside pf-encode's own AMF smoke test. Both planes send the static HDR grade out-of-band unconditionally (the native 0xCE datagram per keyframe, the GameStream 0x010e control message), every first-party client reads exclusively that path, and none parse in-band SEI — so the host decision the field was reserved for (suppress out-of-band when the encoder embeds) can never validly exist. The field's doc contract had also rotted in two directions: it claimed set_hdr_meta no-ops when false (native AMF and QSV consume it regardless) and that only Windows direct-NVENC attaches in-band metadata (AMF and QSV do too). The in-band SEI/OBU emission itself is untouched — it stays a bonus for stock decoders, documented at the emit sites; the trait docs now describe the real routing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
232b6d6be2 |
test(encode/vaapi): extract the open-time decision logic and pin it with unit tests
The fallback backend under Vulkan Video — and the only AMD/Intel H.264 and 10-bit/HDR encoder — had 1,300 lines, 26 unsafe blocks, and zero tests. The device-free decisions now live in named functions with their contracts pinned: the entrypoint ladder + LP_MODE latch round-trip (the cross-GPU session-killer and the 8-bit-pins-10-bit under-advertisement are both key'd tests now), the PUNKTFUNK_VAAPI_LOW_POWER / _ASYNC_DEPTH grammars, the VUI ↔ scale_vaapi colour agreement (the Mesa-BT.601 hue-shift pin), the honest-downgrade depth table, the HEVC-Main10-only explicit profile, and the 10-bit probe gate. Probe + CPU-path encode round-trip ride along as #[ignore]d hardware smokes in the house style. No FFI plumbing was chased; no behavior changes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
cc848479c4 |
test(pf-encode): re-point the cpu_img size-change smoke at the CSC guard's contract
arch / build-publish (push) Failing after 58s
ci / web (push) Successful in 51s
ci / docs-site (push) Successful in 53s
android / android (push) Failing after 7m59s
ci / bench (push) Failing after 4m9s
ci / rust-arm64 (push) Failing after 6m44s
ci / rust (push) Failing after 6m45s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 24s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 12s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 7s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 47s
apple / swift (push) Successful in 5m28s
deb / build-publish (push) Successful in 9m25s
windows-host / package (push) Successful in 10m44s
docker / deploy-docs (push) Successful in 31s
docker / build-push-arm64cross (push) Successful in 5m33s
deb / build-publish-client-arm64 (push) Successful in 8m20s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 7m41s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 7m4s
deb / build-publish-host (push) Successful in 12m7s
apple / screenshots (push) Successful in 25m11s
vulkan_cpu_img_survives_a_source_size_change drove MISMATCHED source sizes through the then-lenient CSC arm as its vehicle for the staging cache hazard (format-only-keyed cpu_img → OOB copy while submit said Ok). e3354b6d's guard — correctly the equality check every sibling arm always had, against the MODE, not the coded extent, so the padded render-vs-coded tolerance in the direct arms is untouched — makes that scenario unrepresentable through submit and broke the test on main (.25 layers baseline read 13/13 instead of 14/12). Replaced by vulkan_csc_refuses_a_mismatched_source: refusal pinned in BOTH directions, plus the property that actually needs proving — a refused submit does not WEDGE the session (the bail lands after step 1's frame-type bookkeeping; the next well-sized frame must still encode, and an AU must come out). Verified on the 780M under validation layers: 8/8 vulkan tests, full-suite baseline restored to 14/12. The WP4.2 size-keyed staging stays as belt-and-braces; the hazard it fixed is now structurally unreachable through submit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
bf9386ecb2 |
test(pf-encode): pin the typed-EINVAL classifier's chain-survival contract (Phase 8)
ci / docs-site (push) Failing after 52s
android / android (push) Failing after 6m28s
ci / web (push) Successful in 1m5s
ci / bench (push) Successful in 6m40s
decky / build-publish (push) Successful in 26s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 16s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
arch / build-publish (push) Successful in 12m51s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
ci / rust-arm64 (push) Successful in 9m51s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Failing after 42s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 17s
deb / build-publish-client-arm64 (push) Successful in 8m54s
deb / build-publish-host (push) Successful in 10m31s
deb / build-publish (push) Successful in 11m17s
windows-host / package (push) Successful in 10m41s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m25s
docker / build-push-arm64cross (push) Skipped
docker / deploy-docs (push) Skipped
apple / swift (push) Successful in 5m25s
ci / rust (push) Successful in 22m21s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m7s
apple / screenshots (push) Successful in 23m50s
Rides on
|
||
|
|
bf9fb3fb22 |
fix(pf-encode): enable VK_EXT_queue_family_foreign for the dmabuf acquires (Phase 8)
Both Linux Vulkan encode backends named QUEUE_FAMILY_FOREIGN_EXT as the acquire barriers' src family without ever enabling the extension — spec-invalid on every device, tolerated by RADV. The audit filed vulkan_video's three sites; pyrowave's fresh-import acquire had the identical defect on its own device (critic catch). Enable when advertised (a fresh open-time enumerate — the rgb probe's is a probe-local and skipped entirely on native-NV12, so there was nothing to reuse; pf-presenter/dmabuf.rs is the in-repo precedent that already enables this extension). Not advertised → the core-1.1 QUEUE_FAMILY_EXTERNAL conservative substitute, chosen once at open and warn-logged (no fleet hardware takes that arm; such devices were never valid targets before). All four sites are acquire-only (src=FOREIGN, EXCLUSIVE images, oldLayout=UNDEFINED) — the swap is index-only. On-glass: 780M under validation layers — vulkan smokes + pyrowave smokes green, FOREIGN advertised and enabled, no fallback engaged. Shared ext_advertised helper in vk_util (cfg = the union of both consumers) with a unit test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1cefd37603 |
fix(pf-encode): arbitrate NVENC split-encode vs sub-frame readback (Phase 8)
Verified against nvEncodeAPI.h's own splitEncodeMode doc (user-prompted — the audit's 'exclusive for HEVC' one-liner deserved checking): - H.264: split 'is not applicable' — hard-DISABLE the mode so the written config, CeilingKey, the split diagnostic log and the rejection-retry stay truthful (the retry used to re-open a byte-identical session after an H.264 'split rejection'). The libav path's operator arm gains the codec gate its auto arm always had. - HEVC: split 'not supported if … subframe mode' — when WE force split (TWO/THREE/AUTO_FORCED, the 4K120 throughput lever), sub-frame yields with a logged escape (PUNKTFUNK_SPLIT_ENCODE=0 chooses sub-frame). ⚠ Keyed on FORCED modes only, never != DISABLE: AUTO(0) is the resolver's fallthrough for every sub-950Mpix session, and the wider key would have disarmed the Phase-3 chunked-poll feature fleet-wide (critic catch). Under AUTO the driver arbitrates — the shipped state. - AV1: untouched — per-tile sub-frame + split are legal together. The arbitration is a pure nvenc_core fn called by each backend BEFORE the ladder, the ceiling key and the chunked-poll latch — all three see the post-arbitration truth. A drop inside build_init_params would have left poll_chunk busy-polling its whole budget every AU (numSlices stays 0 without reportSliceOffsets; both loop exits dead — critic catch). Linux latches subframe_forced beside subframe_on at query_caps (no env re-reads after open); Windows records the arbitrated state so reconfigure presents exactly the params the open had (also closes the pre-existing mid-session env-flip hazard there). Truth-table tests in nvenc_core; PUNKTFUNK_NVENC_SUBFRAME documented (it never was); PUNKTFUNK_SPLIT_ENCODE row updated. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
25765c53ec |
fix(encode): classify libav-NVENC open failures by errno, not English strerror text
ci / docs-site (push) Successful in 58s
ci / web (push) Successful in 59s
apple / swift (push) Successful in 5m17s
ci / bench (push) Successful in 7m27s
deb / build-publish (push) Failing after 8m0s
ci / rust-arm64 (push) Failing after 9m1s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 22s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
ci / rust (push) Failing after 9m31s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
deb / build-publish-host (push) Successful in 10m23s
arch / build-publish (push) Successful in 12m54s
android / android (push) Successful in 15m8s
deb / build-publish-client-arm64 (push) Successful in 7m54s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 5m43s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 6m35s
windows-host / package (push) Successful in 19m10s
docker / build-push-arm64cross (push) Successful in 8s
docker / deploy-docs (push) Successful in 24s
apple / screenshots (push) Successful in 23m39s
The bitrate-probe ladder stepped down on format!("{e:#}").contains(
"Invalid argument") — an English substring over the WHOLE context chain,
which also fired on any other wrapped EINVAL (e.g. a CUDA-context errno)
and gated a ~10-step ladder on strerror wording. The root ffmpeg::Error
survives the anyhow chain; downcast and match Error::Other{errno:EINVAL}
instead. Same fix for the intra-refresh ENOSYS probe in the open path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
e3354b6d5d |
fix(encode/vulkan): guard the CSC source dimensions, and bound reset()'s wait
The CSC path was the only backend arm that took frame.width/height on trust: the shader samples with clamped 1:1 texelFetch, so a mismatched frame silently streamed a cropped/edge-padded picture where every sibling errors into the encoder-rebuild path. The import cache now also carries the extent it imported at (a (st_dev, st_ino) hit alone doesn't prove the allocation still matches) and is dropped on reset(). reset() opened with an untimed device_wait_idle on the one thread whose every other wait is capped at ENCODE_FENCE_TIMEOUT_NS for exactly this reason — reset() runs BECAUSE the GPU looks wedged. Both vulkan-video and pyrowave now bound the wait and report "no in-place rebuild" on timeout instead of parking recovery on the suspect device; Drop keeps the unbounded wait (teardown must stay memory-safe against a wedged device). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
28f8fc71c4 |
refactor(pf-encode): split vulkan_video's construction tail into vk_build.rs (WP7.5)
ci / web (push) Successful in 56s
ci / rust-arm64 (push) Failing after 2m18s
ci / docs-site (push) Successful in 1m2s
android / android (push) Failing after 4m7s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 9s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
apple / swift (push) Successful in 5m21s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 6m10s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 5m40s
deb / build-publish-client-arm64 (push) Successful in 7m57s
docker / build-push-arm64cross (push) Successful in 10s
deb / build-publish (push) Successful in 10m6s
docker / deploy-docs (push) Successful in 24s
deb / build-publish-host (push) Successful in 10m30s
arch / build-publish (push) Successful in 13m57s
windows-host / package (push) Successful in 14m6s
ci / rust (push) Successful in 22m9s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m30s
apple / screenshots (push) Successful in 24m30s
The ~820-line tail of free builders — make_frame/make_frame_csc/ make_frame_common, make_video_image, probe_rgb_direct, the H.265/AV1 parameter-set writers and the AV1 bit-writer — moves to a #[path] child module, the amf_sys.rs shape: the child sees the parent's private items (Frame and friends), so the split costs zero visibility churn. Six parent-called items went pub(super); five stay child-private (dead_code is per-item and each is used within the child). vulkan_video.rs drops 5,292 → 4,489 lines and the construction unsafe gets its own review surface; steady-state encode logic stays in the parent. ⚠ Trap recorded for future child-module splits: inline `use super::X` statements INSIDE moved fn bodies silently change meaning (super shifts one level) — vk_av1_encode/vk_valve_rgb imports needed crate:: paths. Proven on-glass, not just compiled: all 8 vulkan GPU smokes green under the validation layers on the 780M post-split (H.265 + AV1, RGB-direct, CSC, CPU paths — every moved constructor exercised). nvenc_cuda.rs and qsv.rs are DECLINED the same treatment, with evidence in the handoff doc: no equivalent self-contained seam — their candidate regions are ~150-line loader/accessor clusters interleaved with the encoders' own state types, and a thin-forwarder impl split is exactly the churn a no-defect phase penalizes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9491cc8759 |
refactor(pf-encode): extract the range-family RFI recovery policy (WP7.2)
The two direct-NVENC backends carried hand-copied twins of the same loss-recovery decision: range validity, covering-range dedup, DPB window, clamp — ~30 duplicated lines each. The decision now lives once as nvenc_core::plan_range_recovery (the range half of WP7.2; the slot half is enc/rfi.rs), pure and unit-tested; each backend keeps its session gate, its unsafe per-timestamp driver loop, and its state stores. The step order is load-bearing and now pinned by tests: the covering dedup runs with the UNCLAMPED last and BEFORE the DPB window (a covered re-ask never touches the driver even when the range has since aged out of the DPB), the boundary at next_ts - RFI_DPB is inclusive, and the Invalidate carries the CLAMPED last — which is also what the caller records in last_rfi_range, exactly as the inline code stored it. A driver failure mid-loop still returns false with NO range recorded and no anchor armed. Decline deliberately clears nothing (neither twin touched pending_anchor on decline — same shape as Vulkan's non-clear, opposite of AMF/QSV; do not harmonize). The exact-cover → Covered test records EXISTING behavior including that a covered range survives a forced IDR with zero driver calls — a recorded fact, not an endorsement. RFI_DPB's import leaves both twins: its only per-backend use was the arithmetic that moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9f1e648e4e |
refactor(pf-encode): extract the slot-family RFI recovery policy (WP7.2)
AMF (user-LTR bitfield), QSV (mfxExtRefListCtrl) and Vulkan Video (the
app-owned DPB slot table) each hand-implemented the same loss-recovery
decision: distrust every reference encoded at-or-after the loss start,
anchor on the newest one strictly older. Three copies had already
diverged once — the
|
||
|
|
ffc7aec91a |
feat(encode/pyrowave): log which GPU pyrowave picked — the selection stays put, by decision
ci / web (push) Successful in 55s
ci / docs-site (push) Successful in 1m12s
apple / swift (push) Successful in 5m20s
ci / rust (push) Failing after 5m55s
android / android (push) Failing after 5m59s
deb / build-publish (push) Failing after 5m1s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
decky / build-publish (push) Successful in 17s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
ci / bench (push) Successful in 7m17s
ci / rust-arm64 (push) Successful in 9m40s
docker / deploy-docs (push) Successful in 26s
docker / build-push-arm64cross (push) Successful in 4m9s
deb / build-publish-host (push) Successful in 10m32s
arch / build-publish (push) Failing after 12m34s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 5m58s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 5m50s
deb / build-publish-client-arm64 (push) Failing after 7m10s
windows-host / package (push) Successful in 17m27s
apple / screenshots (push) Successful in 23m26s
WP4.5's device-selection half, closed as the observability intermediate after TWO selection designs died in adversarial review: - Attempt #1 (d26bcf05, withdrawn): match pf_gpu::selected_gpu(). Its Linux auto arm answers "the NVIDIA GPU" whenever /dev/nvidiactl exists, moving the encoder off the iGPU that can import the compositor's dmabufs on an Intel-compositor + NVIDIA-present laptop — import failures feed the process-wide raw-dmabuf latch, which never un-latches. - Attempt #2 (this session, withdrawn before commit): anchor on the PUNKTFUNK_RENDER_NODE-else-renderD128 node via VK_EXT_physical_device_drm. Render minors are driver-BIND-ORDER artifacts, not display topology: on the common AMD-iGPU + NVIDIA-display desktop, in-tree amdgpu binds before out-of-tree nvidia, so the anchor deterministically picks the idle iGPU while the compositor allocates on NVIDIA — the same latch, opposite polarity, behind a success-looking log. The correct oracle is evidence of which device ALLOCATED the capture buffers — producer identity from the capture negotiation, threaded per session into this open. Until that plumbing exists, selection stays first-usable, both call sites still share one selector (pure over the device list, so capture_modifiers and open_inner cannot diverge — including across an in-place resize's re-open, which does not renegotiate capture), and the open logs ONE greppable line: picked vendor/device, the anchor node and its owner (DRM render major/minor, VK_EXT_pci_bus_info fallback), and the console's selected GPU. A wrong-device session on a multi-GPU host used to be completely invisible; a field report can now show it. No WARN arm on purpose: the wrong-pick direction inverts between the laptop and desktop topologies, so a mismatch is not evidence of a wrong pick, and a warning that fires forever on healthy hosts teaches people to ignore warnings. Decision recorded against the audit's framing: manual console GPU selection stays unhonored by pyrowave on Linux (the Windows twin honors it) — honoring console-mutable state without per-session threading is what made attempt #1 unsafe. Verified on the 780M: the line resolves all three identities (1002:15bf x3, DRM-props match live); full pyrowave on-glass suite green; selection behaviour byte-for-byte unchanged. WP4.5 (device half). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
46935bf0a5 |
fix(encode/vulkan): VBR instead of CBR — the driver was stuffing ~98% filler into every calm stream
Second attempt at WP6.3; the first (ce543668) was withdrawn after its tight CBR window measured as a 36x bandwidth regression (97% filler NALs). The correction that unlocked this one: that measurement's BASELINE row was an 8-frame artifact. Under the shipped 1000ms/500ms CBR window a calm stream overflows the CPB once the initial fill drains (~30 frames at 10Mbps/60fps) and RADV then pads every frame to the exact rate share, forever. Measured on the 780M (1280x720@60, 10Mbps, calm content): 64 frames = 97.5% filler; 300 frames = 5.63MB at 98.5% filler where this commit ships 83KB at 0%. AV1: 99.6% -> 0%. The status quo was ~the full target bitrate of zeros on every idle AMD/Intel Vulkan-encode desktop and Steam Deck — and stuffing to exactly the target permanently satisfied the ABR calm brake (actual >= 3/4 * current), the ratchet WP6.3's withdrawal feared from the tight window, live in the shipped code all along. The fix reads VkVideoEncodeCapabilitiesKHR::rateControlModes (previously ignored — rateControlMode was hardcoded CBR with no capability check) and installs VBR with average == max plus the house ~1-frame window (vbv_window_ms, PUNKTFUNK_VBV_FRAMES-scaled) when the driver advertises VBR. VBR permits underspend, the exact missing degree of freedom: Vulkan exposes no filler-suppression control (AMF's filler_data=false / NVENC's default-off have no VK equivalent), so the MODE is the only lever. CBR-only drivers keep the loose window untouched — tightening it under CBR just starts the stuffing 30 frames earlier. Drivers advertising neither mode (ANV per current Mesa) keep the pre-existing CBR install, now WARN-logged. No pacing claim, deliberately: burst A/B on the 780M is byte-identical between 1000ms CBR and 17ms VBR (max AU 1.19MB in both) — this firmware ignores the window for QP decisions entirely. The payload is filler elimination. PUNKTFUNK_VULKAN_RC=cbr|vbr is the field escape hatch and the on-box A/B control (two withdrawn attempts bought that insurance). Also on the same caps struct: maxBitrate is now read and clamps open + retarget (RADV reports 1 Gbps — within 5% of the 4K120 ABR targets), and applied_bitrate_bps() reports the encoder-side truth (pending-first, so the session loop's read right after reconfigure_bitrate sees the clamp) — without it a binding clamp would feed the ABR a phantom base, the trap the trait doc names. And the one-frame VUID-vkCmdBeginVideoCodingKHR-pBeginInfo-08254 violation found in the withdrawal review: record_submit promoted a pending retarget into self.bitrate BEFORE recording whenever first_frame was set, so after a mid-stream reset() (which preserves the pending rate and rc_installed) the begin-coding declaration named a rate the session had not installed — and the two triggers, ABR retarget and the stall watchdog, correlate. Now the declaration always names the session's current rate and the RESET install carries the pending one via its own struct; promotion stays in post_submit_bookkeeping. The extended validation-layer test reproduces the retarget-then-reset coincidence: exactly one 08254 on the pre-fix build, zero on this one (RADV PHOENIX, on glass). Gates: docker amd64 legs green; Windows .173 seven legs (34 passed); .25 full-suite parity vs origin/main (identical CUDA-only failures) + all 9 vulkan on-glass tests + validation layers clean. WP6.3, plus WP7.1's ms-form half (vbv_window_ms). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fdded5b8c3 |
fix(encode/pyrowave): refuse a frame that isn't the session's mode
ci / web (push) Successful in 54s
ci / docs-site (push) Successful in 55s
ci / bench (push) Failing after 6m37s
ci / rust (push) Failing after 6m43s
deb / build-publish-host (push) Failing after 5m44s
arch / build-publish (push) Failing after 6m45s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 12s
decky / build-publish (push) Successful in 18s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 7s
docker / build-push-arm64cross (push) Successful in 9s
docker / deploy-docs (push) Successful in 23s
ci / rust-arm64 (push) Successful in 9m47s
deb / build-publish (push) Successful in 9m9s
android / android (push) Successful in 11m59s
deb / build-publish-client-arm64 (push) Successful in 8m44s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 8m27s
windows-host / package (push) Successful in 17m35s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m58s
apple / swift (push) Successful in 6m51s
apple / screenshots (push) Successful in 26m43s
PyroWave never checked frame dimensions against the session, and it applies no alignment — `width`/`height` are the negotiated mode verbatim — so a mismatched frame was encoded edge-smeared or cropped, silently, forever. Every other Linux backend already refuses exactly this, with this shape, in `submit`: libav-NVENC (`linux/mod.rs`), VAAPI (`vaapi.rs`) and openh264 (`sw.rs`) all carry the same `ensure!`. PyroWave was the only one that didn't. That is the justification; an earlier draft cited `vulkan_video.rs` instead, which is the weaker precedent — it bails in its Dmabuf arms only, and its CSC path and CPU arm have no dimension guard at all. Mostly this is a wrong-picture bug and not a memory-safety one: `rgb2yuv.comp` clamps every fetch with `min(p, textureSize - 1)` and the CPU arm uploads `min(len, need)` into a session-sized image. But it also closes a narrow real hazard that was not in the filing: `import_cached` keys on `(st_dev, st_ino)` and returns the cached `VkImage` on a hit WITHOUT rechecking the extent, and unlike the capture side it is never cleared on a renegotiation — so a dmabuf inode recycled across a shrinking renegotiation would hand the encoder an image sized for the old, larger allocation. This check closes that route. ⚠ Recorded at the code because it changes the failure mode, not just the detection: a mismatch is NOT always transient. A compositor-initiated PipeWire renegotiation updates the capturer's size in place and signals nothing the encode loop reads, so it can be a permanent new steady state — and `reset()` reopens at the same dimensions by construction, so the host's five-reset budget cannot recover (~3.1 s of frozen stream, then the session ends) where before it would have streamed on with a wrong picture. For a real mode change that trade is clearly right; for a 16-row KWin mismatch it is not, and the proper fix is for the host to classify this error as a PIPELINE rebuild rather than an encoder reset. Filed, not done here. The device-selection half of WP4.5 was written, reviewed and WITHDRAWN — see the handoff doc. Matching the selected render GPU regresses the hybrid Intel-compositor + NVIDIA-present topology this project has a live field report for, because `selected_gpu()` answers NVIDIA whenever `/dev/nvidiactl` exists regardless of where capture actually runs. WP4.5 (dimension half). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9fe9cbbf07 |
perf(encode/vulkan): build the padded RGB frame in the staging memory, not beside it
The RGB-direct CPU-upload path allocated and zero-filled a whole padded frame on every submit, filled it row by row, then memcpy'd the whole thing into the mapped staging buffer. The zero-fill was entirely dead: the row loop writes every byte of every row, and rows past the source re-copy the last source row. So the frame paid for an allocation, a page-fault storm over fresh pages, a full zero-fill, and two full-frame copies where one would do. Map first, write the padded rows straight into the mapping. Nothing reads back from the destination — the row source is always the caller's buffer, never the staging memory — so writing into (write-combined) host memory costs nothing extra, and `make_host_buffer` allocates HOST_COHERENT, so the writes need no flush before the transfer reads them. Measured on real RDNA3 780M silicon, release build, `submit` alone, 400 frames x 3 rounds: - **1920x1080** (rows padded to 1088): p50 2315 -> **1725 us** (-25%), p99 2819 -> **1972 us** (-30%), and the spread (p99-min) tightens 739 -> 483 us. - **1366x768 -> 1408x768** (BOTH axes padded, so the column tail loop runs): p50 1013 -> **926 us** (-8%), p99 1303 -> **1110 us** (-15%). This is the one case where the new code could have been slower — 4-byte stores straight into write-combined memory — and it is not. - **CONTROL, 1280x720** (64x16-aligned, so the branch is never entered): p50 692 vs 692 us. No delta, which is what makes the two above attributable to this change rather than to anything else on the branch. The branch is reached by any RGB-direct session at a mode that is not 64x16 aligned, whenever capture delivers CPU frames (the default Linux capture path is dmabuf and never enters it). 1080p qualifies, since 1080 aligns to 1088. ⚠ An earlier draft of this message claimed "33 MB at 4K" — that is wrong: 3840 and 2160 are both already aligned, so 4K UHD never enters this branch at all. The large case is an ultrawide like 3440x1440 -> 3456x1440, ~20 MB. Three things this deliberately does NOT do: - The extent guards stay scoped to the `pad` branch. The CSC path deliberately supports a source SMALLER than the encode extent — its shader clamps at sample time — so a guard hoisted above the branch would break it. - Every fallible step stays above the map. An error raised between `map_memory` and `unmap_memory` would strand the mapping for the life of the slot's staging buffer, and the next frame's `map_memory` on it then violates VUID-vkMapMemory-memory-00678. (The filed "two exits above the map leak Vulkan objects" hazard is separately already gone: `47a23bec` moved that unwind into `make_host_buffer`.) - It adds guards rather than removing them: a zero source axis made `sh - 1` underflow, and "cannot fail after the map" has to be true by construction. Three corrections to an earlier draft, all found by review of that draft: - The `dw*dh*4 == need` precondition was a `debug_assert!` placed BELOW the map. That is wrong twice: it made the only check on the slice length vanish from the builds that ship, and a fired assert would have unwound past `unmap_memory` — the exact failure the bullet above says was designed out. It is now a real checked `?` above the map. - `need` was `(iw * ih * 4) as u64`: a u32 multiply widened after the fact, which agreed with the usize slice length only up to ~32768x32768. The old code's `min(need)` was a hard backstop against exactly that and the rewrite dropped it. Now widened before the multiply, matching `read_slot`'s existing discipline. - The extent guard now runs BEFORE the payload-length guard, so the usize `sw * sh * 4` cannot overflow on a garbage frame header. WP6.2(a). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e680096c6a |
perf(encode): stop re-reading the environment on every submit and poll
`std::env::var` was on three per-frame paths. Measured on `.173` (Windows, 57 environment variables, 2M iterations): **121.9 ns** per call for the NVENC in-flight cap and **114.9 ns** per call for the ffmpeg poll spin, against **0.9 ns** once memoized. On Linux, 32 ns → 1.5 ns. ⚠ Those numbers deflate the filing, and that is worth recording: the audit ranked this as a hot-path defect, but ~120 ns/frame is ~0.003% of a frame budget. The fix is still right — it is free, and it takes a global environment lock off the encode thread — but nobody should schedule it ahead of anything on the strength of the "hot path" framing. The severity RANKING was also inverted, and the measurement confirms why. The site the audit called worst — `nvenc_cuda`'s backpressure loop condition — costs a default session nothing, because the condition short-circuits on `async_rt.is_some()` and the default session never engages the two-thread retrieve. The one that actually pays every frame is Windows `submit`, which consults `async_inflight_cap()` in BOTH arms of the ring-depth match, sync mode included, where the result is then thrown away. Windows `poll` is the second unconditional one, and the audit ranked it last. Memoized inside each helper rather than latched into a session field. Nothing in the workspace mutates these variables at runtime (enumerated: no `set_var` for either key anywhere in first-party code, and the Windows service's arbitrary-key `host.env` loader runs in the SCM supervisor, which re-execs the host as a child and never opens an encoder itself). A field would instead change WHEN the value is read — and the Windows `cap` composes the env with `input_ring_depth`, which `set_input_ring_depth` may change after open, so freezing that half would reintroduce the in-place-overwrite bug the ring term exists to prevent. Two deliberate behaviour changes ride along on `PUNKTFUNK_FFWIN_POLL_MS`, so "behaviour-preserving" describes the memoization only, not this whole commit: - **A 1000 ms ceiling.** The reachable hazard was never the overflow — that needed `ms >= 1.8e16` — it was a slipped digit: `=100000000` was a 27.7-hour spin of the encode thread. - **`.trim()`, now on all three parsers.** An earlier draft applied the house rule (WP7.8) to one of the three, which left `PUNKTFUNK_NVENC_ASYNC=" 1 "` working while `PUNKTFUNK_NVENC_ASYNC_DEPTH=" 6 "` silently fell back to 4 — and memoization would have frozen that silent fallback for the process lifetime. Reachable: the Windows `host.env` loader trims around `=` before stripping quotes, so `=" 2 "` yields a value with inner spaces. ⚠ Correction to an earlier draft of this message, which asserted that the audit's `saturating_mul` proposal would relocate an overflow panic into release builds. That is FALSE and the code comment now says so: `Duration::from_micros(u64::MAX)` is ~1.8e13 seconds, six orders of magnitude below `Duration`'s ceiling, so `Instant + Duration` neither overflows nor panics (measured, with and without debug assertions). `saturating_mul` is still wrong, for a different reason — it sets a deadline ~584,000 years out on a spin that provably never produces the owed AU, i.e. it wedges the encode thread permanently. A hang, not a panic. WP6.1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
47a23bec12 |
fix(encode/vulkan): unwind every open/import leak, and serve 24-bpp CPU instead of dying on it
decky / build-publish (push) Successful in 23s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
arch / build-publish (push) Successful in 12m42s
windows-host / package (push) Successful in 18m11s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m14s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m6s
docker / deploy-docs (push) Successful in 32s
docker / build-push-arm64cross (push) Successful in 6m26s
apple / swift (push) Successful in 6m42s
deb / build-publish-host (push) Failing after 5m39s
android / android (push) Failing after 6m36s
deb / build-publish-client-arm64 (push) Failing after 5m35s
deb / build-publish (push) Successful in 9m40s
ci / rust (push) Successful in 22m26s
ci / web (push) Successful in 52s
ci / docs-site (push) Successful in 55s
ci / bench (push) Successful in 5m54s
ci / rust-arm64 (push) Successful in 11m18s
apple / screenshots (push) Successful in 27m18s
Phase 5's Linux half (audit WP5.1 + WP5.4), each item shaped by the
review that rejected the obvious fix:
Dmabuf import unwind (vk_util): every failure after create_image leaked
the VkImage, and the dup'd dmabuf fd leaked as a raw i32. The sharp edge
is that a SUCCESSFUL vkAllocateMemory transfers fd ownership to Vulkan
(vkFreeMemory closes it), so the naive close-on-error is a double close
that clobbers whatever unrelated descriptor recycled the number. The dup
now lives in an OwnedFd released exactly in the allocate-success arm;
every other path drops it once, and bind/view failures free image+memory.
PyroWave open unwind: open_inner had ~20 fallible steps that each leaked
everything before them (instance, device, pyrowave objects, the whole
CSC pipeline). Rather than a parallel teardown guard — whose reviewed
hazards were a null-unsafe pyrowave_encoder_destroy and a drifting
duplicate of Drop — Self is now constructed right after create_device
with every later resource null, and the existing Drop (wait-idle first,
pw_enc null-guarded, delete-nullptr and VK_NULL_HANDLE destroys are
no-ops) is the single unwind path for error and normal teardown alike.
The ensure_cpu_rgb staging twins (create/allocate/bind, both backends)
and the RGB-direct make_view pair get the same discipline via a shared
make_host_buffer. Observed on hardware: 32 forced import failures, zero
fd drift (the new import_failure_leaks_no_fds smoke on RADV).
24-bpp CPU service (WP5.4): pixel_to_vk had no mapping for the packed
Rgb/Bgr the PipeWire portal negotiates, so a session committed to a path
the backend could not serve and died at its first frame. The filed
open-gate was rejected as a half-mirror — the dmabuf axis is keyed by
fourcc at submit, unknowable at open — so instead the CPU axis is
SERVED: a 3-to-4 expand at the staging upload (normalize_cpu_rgb, the
CPU twin of WP1.4's swscale expand), order-preserving for the CSC
samplers and BGRA-forced for the RGB-direct encode source, whose session
pictureFormat is B8G8R8A8 — the on-glass run caught R-first sources
violating VUID-vkCmdEncodeVideoKHR-pEncodeInfo-08207, a mismatch that
predates this change for plain Rgbx CPU sources. The dmabuf axis feeds
pf-zerocopy's raw-dmabuf degrade latch (
|
||
|
|
3efbe4164e |
fix(capture): capture CPU frames once the encoder proves it can't import dmabufs
audit / cargo-audit (push) Successful in 2m13s
audit / bun-audit (push) Failing after 13s
ci / web (push) Successful in 57s
ci / docs-site (push) Successful in 1m3s
windows-host / package (push) Successful in 10m18s
ci / bench (push) Successful in 6m41s
android / android (push) Successful in 12m26s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 14s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
arch / build-publish (push) Successful in 12m30s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m10s
ci / rust-arm64 (push) Successful in 11m46s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m17s
deb / build-publish (push) Successful in 11m17s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m13s
deb / build-publish-host (push) Successful in 12m35s
deb / build-publish-client-arm64 (push) Successful in 8m23s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m5s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 7m40s
apple / swift (push) Successful in 5m28s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8m6s
docker / deploy-docs (push) Successful in 14s
ci / rust (push) Successful in 21m44s
docker / build-push-arm64cross (push) Successful in 7m18s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m37s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m14s
flatpak / build-publish (push) Successful in 6m38s
release / apple (push) Successful in 29m56s
apple / screenshots (push) Successful in 25m36s
A dmabuf import the GPU driver refuses is refused identically on every retry, but the only recovery above it was the encode-stall ladder: five in-place encoder rebuilds, then the video session ends. So a host whose driver will not take what its compositor allocates lost every session on its first frame, and every reconnect repeated it — while the very same host streamed fine with `PUNKTFUNK_ZEROCOPY=0`. The software knew how to run that machine and never chose to. Latch it, exactly as the sibling CUDA-import path already does after repeated worker deaths: three consecutive import failures with no frame in between disable the raw-dmabuf passthrough for the host process, and capture negotiates CPU frames from the next session on. Three sits below the encoder's rebuild budget, so the latch is set before the session it doomed ends — one bad session, then a working (if slower) host, with a log line saying which and why instead of an operator having to find an environment variable. Only the two stages that ARE the import are counted — the buffersrc push of the DRM-PRIME descriptor and the buffersink pull where `hwmap` maps it into a VA surface. `avcodec_send_frame` is deliberately left out: that one is the encoder stalling, which the in-place rebuild exists to recover, and taking zero-copy away permanently over a transient fault would be a bad trade. The latch lives in pf-zerocopy because it is the leaf both sides can see — the capture→encode edge is one-way by design, so pf-capture cannot ask pf-encode anything. |
||
|
|
f32207a6ba |
fix(encode/vaapi): tell libva the dmabuf's real size, not zero
The zero-copy DRM-PRIME descriptor declared `objects[0].size = 0`, on the belief that ffmpeg would work the real size out. It does not: both of its import paths pass the value straight through to libva — `prime_desc.objects[i].size` on the PRIME_2 path, `buffer_desc.data_size` on the legacy fallback it tries next — so every VA driver we have ever handed a dmabuf to was told the backing object was empty and left to derive the size itself. The drivers this path has run on (radeonsi, modern Intel iHD) derive it correctly, which is why nobody noticed. A Gen9 Intel host does not get that far: `vaCreateSurfaces` answers VA_STATUS_ERROR_ALLOCATION_FAILED on the first frame and every frame after it, `av_buffersink_get_frame` returns EIO, and five in-place encoder rebuilds later the video session is over. That host cannot stream at all with zero-copy on. `lseek(SEEK_END)` is the standard dma-buf size query and the same one this tree's Vulkan bridge already performs on these very fds. A kernel that refuses it leaves the old 0 rather than costing a frame that might still have encoded. Whether this alone fixes that host is unconfirmed — the descriptor was wrong either way, and it is the first thing the driver reads. |
||
|
|
8578141d43 |
fix(encode/vulkan): AV1 at unaligned modes was violating two VUIDs on every frame
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
android / android (push) Successful in 15m38s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 13s
arch / build-publish (push) Successful in 16m13s
windows-host / package (push) Successful in 10m20s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m20s
apple / swift (push) Successful in 5m36s
docker / deploy-docs (push) Successful in 23s
docker / build-push-arm64cross (push) Successful in 8m41s
ci / web (push) Successful in 48s
ci / docs-site (push) Successful in 1m8s
ci / bench (push) Successful in 6m49s
deb / build-publish-client-arm64 (push) Successful in 7m31s
deb / build-publish (push) Successful in 11m40s
ci / rust-arm64 (push) Successful in 12m11s
deb / build-publish-host (push) Successful in 12m15s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 14m59s
ci / rust (push) Successful in 21m45s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m51s
apple / screenshots (push) Successful in 25m22s
Found by running the smokes at 1920x1080 under the validation layers while chasing the three
items left over from WP4.2. AV1 forbids the encode source's `codedExtent` differing from the
sequence header without `FRAME_SIZE_OVERRIDE`
(VUID-vkCmdEncodeVideoKHR-flags-10324), or from the reference slots without
`MOTION_VECTOR_SCALING` (`-10325`). RADV PHOENIX advertises NEITHER — and RGB-direct is the
default on EFC hosts with true-extent the default at unaligned modes, so plain 1080p AV1 was
tripping both on every frame: source 1920x1080 against an app-aligned 1920x1088 header and
DPB. Measured 16 violations per 8-frame run; the CSC path had none.
The fix is to make all three agree at the RENDER size rather than the aligned one — an
unpadded coded size is valid on this hardware, so the coded frame simply IS the visible
frame:
- the AV1 sequence header follows the render size when true-extent is active (joining
native NV12, which already authors true-size headers for the same reason);
- the DPB setup and reference slots carry `src_extent` instead of the aligned `ext2d`.
`src_extent` already collapses to `ext2d` whenever true-extent is off, so every other
configuration is untouched;
- `render_and_frame_size_different` now compares render against the DECLARED source extent
instead of the aligned size, or true-extent would have claimed a mismatch that no longer
exists.
Fixing only the header is not enough and is actively misleading: it clears -10324 and
immediately exposes -10325, because the mismatch has moved to the reference slots rather than
gone. Both had to move together.
This keeps the EFC fast path. Two alternatives were implemented, measured and rejected on the
way here: falling back to the compute CSC costs the zero-copy the B2 work existed to deliver,
and routing to the padded-copy staging trades these two VUIDs for
VUID-VkImageCreateInfo-pNext-06811 — `pad_img`'s extra TRANSFER_SRC usage is not in the
profile-advertised set, measured 8x per session on HEVC-padded too, so that is a pre-existing
defect of the padded path and not somewhere to route a default session.
HEVC is deliberately untouched: it has no equivalent constraint (its crop rides the
conformance window), it measures zero violations, and its aligned-SPS path is the validated
one.
On RADV PHOENIX (780M, Mesa 26.0.4) the AV1 stream now decodes as coded 1920x1080 / render
1920x1080 — genuinely unpadded, where before the alignment rows were encoded and cropped
back out. The CSC path still reports coded 1920x1088 / render 1920x1080, which is correct for
it. All four `vulkan_smoke*` pass at 256x256 and 1920x1080.
Two validation errors remain on the RGB-direct path and are NOT ours, now with evidence
rather than assumption: `VUID-VkImageViewCreateInfo-image-08336` uses the PROFILE-BLIND format
query, so it cannot see that RGB conversion legalises BGRA as an encode source — the
profile-aware query used by -06811 accepts the very same image; and
`VUID-VkQueryPoolCreateInfo-pNext-pNext` rejects
`VkVideoEncodeProfileRgbConversionInfoVALVE`, which the VALVE extension REQUIRES for profile
identity, and the layer diagnoses itself as "a struct from an extension added to a later
version of the Vulkan header".
Verified: canonical Linux gate (docker linux/amd64) L1-L4 green; on-glass on RADV PHOENIX
under `VK_LOADER_LAYERS_ENABLE='*validation*'`, with the bitstreams read back through
libdav1d/trace_headers.
|