faefbae830a8e8d1d1595dc317f2b49ffc19bb6f
16
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f373dffb5e |
chore: migrate the main workspace and pf-vkhdr-layer to edition 2024 (WP20)
The safety half of the rust-safety programme's §8.4: `std::env::set_var`/`remove_var` are
`unsafe fn` in edition 2024, converting the class of bug the programme found the hard way
(the
|
||
|
|
d25a20a233 |
feat(vkdecode): the AV1 rungs meet a second tile for the first time
Every AV1 frame either decode rung has ever been measured against is `tile_cols = tile_rows = 1`. The vendored vector is single-tile on all 274 of its frames, so every tile array the conversions fill — `tiles.widths`, `tiles.heights`, the per-tile records — had only ever been written at index 0, and a conversion that wrote tile 0 and left the rest zero would pass the whole suite. Our encoder splits 4K into TWO TILE ROWS. **The fixture.** `lowdelay-3840x2160.ivf.av1`, 261 KB, 60 frames — `punktfunk-host spike --source synthetic --codec av1 --width 3840 --height 2160 --fps 60 --seconds 1 --bitrate 1` on .21 (NVENC, RTX 5070 Ti), wrapped to IVF with `ffmpeg -f obu … -c copy` so `common::split_av1_aus` (the vendored parser's own `IvfIterator`) frames it exactly as it frames the vector, with no second splitter that could disagree. **4K is not a size choice, it is the only shape with the property.** Measured on the same box with the same command: 1280x720, 1920x1080 and 2560x1440 all give `tile_cols = tile_rows = 1`; 3840x2160 gives `tile_cols = 1, tile_rows = 2` with `width_in_sbs_minus_1 = [59]`, `height_in_sbs_minus_1 = [16, 16]`, and both tiles in ONE Tile Group OBU. 60 frames instead of 120 pays for the resolution: 261 KB, under both the 282 KB H.264 and 270 KB H.265 low-delay fixtures. Goldens are libavcodec's software decode, cross-checked between ffmpeg n8.1.2 (Arch x86_64, libdav1d) and 8.1.1 (Homebrew, macOS arm64, libdav1d) whose 746,496,000-byte raw outputs are BYTE-IDENTICAL, not merely equal per frame. 60 of 60 digests distinct. **AV1's frame accounting is asserted, never derived.** The vendored vector is 250 temporal units carrying 274 coded frames of which 24 are hidden; this stream is 60 units, 60 coded, 60 shown, 0 hidden, 0 `show_existing_frame`, 1 key frame. Neither is the general case, so both parity harnesses now take units / decoded / shown as three independent parameters instead of computing one from another, and the CPU guard states all six numbers. **A CPU gate that needed no hardware at all.** `pic_av1`'s new `a_two_tile_frame_fills_both_row_entries_and_leaves_the_rest_zero` pins the second row entry against its OWN `height_in_sbs_minus_1`, requires the two rows to tile the frame exactly, and requires TWO tile RECORDS out of ONE tile group with rows (0,0) and (1,0) — the transposition a square grid could never reveal — each spanning real bytes. The existing one-tile test asserts index 0 is right and `1..` are zero, which a broken multi-tile conversion also satisfies. ⚠⚠ **This is a file, and on AV1 that distinction has already cost a release.** "250/250 delivered frames bit-identical to libavcodec" was true for the entire period the host was shipping only the FIRST TILE of every 4K frame: the verification ran against a vendored file while the truncation lived in packetisation, and the suite stayed green throughout. This fixture closes the multi-tile gap on the DECODE rungs and closes nothing about fragmentation, reassembly, loss or AU boundaries — the golden header, both module docs and the leg docs all say so, at length, so the next reader does not inherit the same false confidence. Legs: `low_delay_host_av1_every_frame_hashes_bit_identical_to_libavcodec` on the Vulkan rung (11 ignored legs now) and on the D3D11VA rung, plus two non-ignored CPU tests. Verified: 11/11 Vulkan parity legs on .21 (RTX 5070 Ti, 610.57.04), the new one 60/60 bit-identical; workspace clippy `-D warnings` and `cargo fmt --all --check` clean on .21. |
||
|
|
f0702f3e06 |
feat(vkdecode): HEVC's exemption stops being an argument and becomes a vendored stream
`fd6241a2` made HEVC's freedom from the release-ordering defect falsifiable on CPU and
recorded what was still missing: no low-delay HEVC stream was vendored, so the exemption
rested on a structural argument plus one throwaway measurement. This vendors the stream,
and the exemption HELD.
**The fixture.** `lowdelay-640x480.h265`, 270 KB, 120 pictures — `punktfunk-host spike
--source synthetic --codec h265 --width 640 --height 480 --fps 60 --seconds 2 --bitrate 1`
on .21 (NVENC, RTX 5070 Ti, driver 610.57.04). Deliberately the H.264 sibling's resolution
and frame count: the two are then directly comparable, 640 and 480 are both multiples of
MinCbSizeY so there is no conformance window and a hash mismatch can only be decode rather
than readback geometry, and 270 KB sits alongside the 282 KB already accepted for H.264.
Goldens are libavcodec's software decode, cross-checked BIT-IDENTICAL across ffmpeg n8.1.2
(Arch, x86_64) and 8.1.1 (Homebrew, macOS arm64), 120 of 120 digests distinct.
**The exemption held, measured rather than argued.** `sps_max_dec_pic_buffering_minus1 = 4`
against the four pictures 8.3.2 keeps marked in steady state, `sps_max_num_reorder_pics = 0`,
`numRefL0 = 1` — a five-picture DPB filled exactly by four references plus the current
picture. 115 of the 120 access units retire a picture, and `removed ∩ dpb_refs` is **0 of
120**. A 300-picture 1080p stream from the same host reports the same shape: 295
retirements, 0 intersections. It is the encoder and not the resolution, exactly as for
H.264.
**A zero proves nothing on its own, so the fixture is pinned by its counterfactual.**
`test-25fps.h264` reported zero for two milestones while every stream we ship aliased on
99% of its frames. So the guarantee here is not "we looked and it was fine": hand
`plan_to_dxva_h265` the marked DPB as it stood BEFORE `decode_rps` — the mutation a
snapshot move would cause, reconstructed exactly as `dpb_refs(N-1) ∪ {stored(N-1)}` — and
the alias appears on **115 of 120** access units, driven through the real conversion rather
than through planner arithmetic. If a regeneration ever produced a stream that reordered,
or a DPB deeper than its reference count, that 115 collapses to 0 and the tests say so
instead of continuing to pass.
**The two rungs are exempt for different reasons, and the asymmetry is now a gate.** DXVA
binds the whole marked DPB — `RefPicList` is spec-defined that way, and an RFI long-term
anchor has to survive in it — so its exemption really is `H265Planner`'s snapshot ordering,
one call away from being untrue. `plan_to_vk_h265` never reads `dpb_refs` at all:
`pReferenceSlots` is the slots the operation uses, so it binds the current RPS sets, which
`decode_rps` itself derives and which therefore cannot name a picture that same RPS just
dropped. A new test feeds that conversion the identical widened snapshot and asserts
nothing changes, so a future change making the Vulkan rung bind the marked DPB — a
legitimate thing to want, since a *Foll* anchor invisible to the hardware is the RFI
failure shape — fails loudly instead of silently acquiring the defect.
What the Vulkan pixel leg adds is therefore NOT aliasing coverage, and its docs say so:
it is the first HEVC frame either rung has decoded from our own encoder, under a DPB that
retires and reissues a slot on 115 of 120 access units back to back, where the vendored
vector's reordering keeps that eviction slack.
Legs: `low_delay_host_h265_every_frame_hashes_bit_identical_to_libavcodec` on the Vulkan
rung (10 ignored legs now, up from 9) and on the D3D11VA rung, plus three non-ignored CPU
guards that run in ordinary CI.
Verified: 10/10 Vulkan parity legs on .21 (RTX 5070 Ti, 610.57.04), the new one 120/120
bit-identical; workspace clippy `-D warnings` and `cargo fmt --all --check` clean on .21.
|
||
|
|
fd6241a24f |
fix(dxvadec): the review round — a doc that had become false, a warn-storm on renegotiation, and HEVC's exemption made falsifiable
Four findings, all real. **`SlotMap`'s own docs had become false.** "feed it every `DpbUpdate` in decode order (via `Self::apply` or `plan_to_vk`, which applies internally)" — `plan_to_vk` no longer applies internally, which is the entire point of the change, and `release`'s docs named it as one of the two things that may free a slot. A reader following those docs would build the next caller wrong in exactly the way this commit's parent fixed. Both now say which conversions defer, which one does not, and why H.265 is the one that does not. **The deferred release warned on a legitimate event.** `release_deferred` warned per id when a deferred release found no slot — but a renegotiation replaces the whole `Session`, and with it the slot map, INSIDE `plan`, while the planner's own drain reports every drained picture in that same access unit's `removed`. Every one of those ids then misses, and nothing is wrong. `debug!`, with the legitimate cause named so the illegitimate one stays diagnosable. **HEVC's exemption was asserted only in its consequence.** `the_current_picture_is_ named_by_curr_pic_and_never_aliases_a_reference` checked that no reference shares the decode target's slot — which on the vendored vector holds whether or not the reasoning behind it does. That is precisely how the H.264 leg passed for two milestones. The test now also asserts the PLANNER property the exemption rests on (`removed ∩ dpb_refs = ∅`, falsified by moving `dpb_snapshot()` above `decode_rps`), and records that the low-delay measurement was 0 of 300 against H.264's 297 of 300 from the same host and the same run. It also records what is still missing: no low-delay HEVC stream is vendored, so HEVC's freedom is a re-derivable argument plus one measurement, not a standing hardware leg. **Two stale cross-references.** Both AV1 conversions told the reader the H.264/H.265 zero was "measured on reordering vectors and not a proof" — the open question this commit's parent closed. They now say what the answer was. |
||
|
|
834b244301 |
fix(client): the H.264 twin was real — every low-delay picture decoded into a surface it predicted from
The AV1 review round flagged the H.264 leg as "plausibly the same defect, traced in source, not reproduced" and deliberately did not touch it. It is reproduced now, and it is worse than the AV1 one: it fires on 297 of 300 access units of every stream a punktfunk host emits, at 720p, 1080p and 2160p alike, on BOTH the DXVA rung and the Vulkan one. **Decided on the CPU, no GPU needed.** `H264Planner` snapshots `dpb_refs` in `begin_picture`, BEFORE `finish_picture` runs 8.2.5's marking and C.4.5.3's bump, so a picture the sliding window unmarks and the bump then evicts lands in both `dpb_refs` (which `RefFrameList` is built from) and `dpb.removed`. The conversion released the whole `removed` list and then assigned the decode target a slot; `SlotMap::assign` takes the lowest free slot, which is the one just vacated. `CurrPic = N` and `RefFrameList[k] = N`, in one submission. The two conditions have to coincide in ONE access unit, and low-delay H.264 is exactly what makes them: `max_num_reorder_frames = 0` means the evicted picture has already been output, which is what makes it evictable at all. NVENC seals it by writing `max_num_ref_frames = 3` ALONGSIDE `max_dec_frame_buffering = 3` — a DPB exactly as deep as its reference count — so the window unmarks the oldest reference in the very unit whose bump drops it. The aliased picture is `ref_idx 2` of a three-entry `num_ref_idx_l0_active` list: addressable by any macroblock, not a spare. **Why two hardware-proven codecs and four GPUs never saw it.** `test-25fps.h264` is level 1.3 with no VUI `bitstream_restriction`, so `dpb_limit` falls back to A.3.1's level ceiling and gives a 7-frame DPB against 2 reference frames — the window unmarks two units before the bump can evict — and it REORDERS, which keeps an unmarked picture alive past the unit that unmarked it. Two independent reasons, both properties of that vector rather than of H.264. It measured zero and passed 250/250 throughout. `data/lowdelay-640x480.h264` is vendored to close exactly that: our own host's output, 120 pictures, goldens from libavcodec cross-checked bit-identical across two ffmpeg builds on two architectures. **The fix is the AV1 fix.** `DecodePlanDxva` and `DecodePlanVk` grow `release_after_decode`, the conversions hand the removals back instead of applying them, and the callers release them once the decode op is issued. It costs no slot the map does not have: `SlotMap::new` allocates `max_dpb_frames + 1` and the DPB never exceeds `max_dpb_frames`, so a free slot always exists with the whole `removed` list still held — measured, peak 4 of 4 on the stream that defers on 117 of 120 units. The Vulkan rung breaks on it in both DPB modes and neither loudly: DISTINCT hands the aliased reference the same array layer the setup writes; COINCIDE clears `slot_image[setup]` in the binding sync and the reference then resolves to no bound image, dropping out of `pReferenceSlots` with a `trace!`. Its deferred release runs on the FAILURE paths too — the fallible region's Result is held rather than `?`-ed, because seven exits sat between the conversion and the release and each would have leaked a slot. `a_full_dpb_bump_reuses_the_slot_but_the_pool_model_binds_a_fresh_image` asserted the aliasing as "the planner's normal behaviour": an authored depth-1 stream whose AU1 references the picture it evicts. It now asserts the opposite, which is the defect in two lines. New evidence, all of it runnable: the CPU proof pins BOTH numbers (0 on the vector, 117 of 120 on the low-delay stream) so neither can drift silently; the ledger-pressure test measures the peak; and a low-delay parity leg is added to `pf-vkdecode`'s `gpu_parity` and `pf-client-core`'s `video_d3d11_native::parity` so both rungs are held to what they stream rather than only to what they conform to. |
||
|
|
5aeb8d2552 |
fix(client): a failed AV1 decode left the surface's facts saying it holds the last picture
The `damaged` path has cleared `Session::held[setup_slot]` since M7, for a reason that now applies to the failure path too: the slot map says the surface holds THIS picture while the surface still carries whatever the previous occupant decoded, so a later `show_existing_frame` naming it blits the old picture's pixels with the old picture's geometry and colour. The failure path never reached that far before — `decode_into`'s error returned straight out of `frame_av1` — and the previous commit made it continue so the slot releases could run. |
||
|
|
3a4c94ad79 |
fix(dxvadec): the review round — a vacuous predicate, an overstated claim, and the H.264 twin of this defect
Five findings from the adversarial pass, all real. **The deferral predicate was vacuous.** `plan.dpb.removed` is ALWAYS a subset of `plan.dpb_refs`: `Av1Planner::plan_frame` snapshots `dpb_refs` before any mutation and `refresh_slots` can only report a picture that was in `self.slots` at that moment. So `filter(|id| dpb_refs.contains(id))` was a condition that is never false, the eager-release loop beside it could never release anything, and the test assertion "only a picture the submission points at earns the reprieve" could never fire. Now: defer every removal, say why in terms of the planner, and assert the PLANNER's property (`removed ⊆ dpb_refs`) — which is falsifiable, and whose failure would mean the conversion is releasing a surface `ref_frame_map` points at. **The failure-path claim was overstated.** Holding the decode's `Result` closes this frame's leak, not the unit's: `decode_av1` returns on the first failing frame and abandons the rest of the temporal unit's plans, so their removals are never released. 24 of 250 units carry a second frame. Named rather than fixed — what to do with the frames after a failure is the pump's question. **⚠⚠ The H.264 leg plausibly has the same defect, and the comment this change added said it could not.** `pic.rs` builds `RefFrameList` from `plan.dpb_refs`, and `H264Planner` snapshots that in `begin_picture` — BEFORE 8.2.5 marking and the DPB bump. The vendored bump drops a picture the sliding window just unmarked once it has been output, so a picture can land in both `RefFrameList` and `dpb.removed`: the AV1 aliasing shape exactly. Measured zero on the vendored vector — but that vector REORDERS, which is precisely what keeps an unmarked picture alive past the AU that unmarked it. A punktfunk host emits LOW-DELAY H.264, where output happens as each picture is decoded, which is the condition that makes eviction and unmarking land in the same access unit. Traced end to end in source, not reproduced (no low-delay vector). NOT fixed: changing a hardware-proven codec on an unreproduced suspicion is the worse risk two commits before a release. Instead `no_au_removes_a_picture_its_own_reference_list_names` makes the assumption falsifiable, and its message says what to do when it fires. HEVC is structurally safe and now says why: `H265Planner` snapshots `dpb_refs` AFTER `decode_rps`. **Four more stale promotion sites**, past the four already fixed: `Backend:: NativeD3d11va`'s variant doc, `Decoder::new`'s Windows rung comment, `lib.rs`'s module note and `clients/session/README.md`. Two sites that used the AV1 leg as the live EXAMPLE of an unproven rung are marked as expired rather than deleted — the reasoning is what the next bad-evidence leg will need. **The AV1 dump was missing.** `PF_DXVA_DUMP` wrote h264 and hevc only, for the one codec whose libavcodec capture has never been taken and where the dump is therefore the only tool. |
||
|
|
f4dda9074b |
feat(dxvadec): the AV1 picparams harness AV1 forgot, and the D3D11VA AV1 rung is promoted
Two halves. **The harness.** `libav_picparams_parity` covered H.264 and HEVC only, which is exactly the gap that let a wrong AV1 submission ship. It now plans, converts and packs all 274 frames of the vendored AV1 vector and checks what needs no capture: the three-buffer descriptor set with no quantization matrix (AV1's matrices are selected by index, so `dxva2_av1_end_frame` passes NULL/0 and there is no buffer to submit), no macroblock count anywhere, the 912-byte picture-parameter buffer, and the tile records — which unlike H.264/HEVC slice records do NOT abut, because a `DXVA_Tile_AV1` addresses a tile PAYLOAD and consecutive payloads are separated by their `tile_size_minus_1` fields. The one that matters most is `no_av1_submission_names_its_decode_surface_in_the_ reference_store`: the invariant the previous commit fixed, over the submitted BYTES rather than over the plan. libavcodec cannot produce that shape — it fills `RefFrameMapTextureIndex` from the pre-refresh store and takes `CurrPicTextureIndex` from a frame the reference update has not run on — which is the argument for calling it a defect rather than a convention. `AV1_FIELDS` reaches into the eight nested blocks (`tiles.widths`, `segmentation.feature_data`, …) so a future capture reports a field and not "260 bytes of tiles differ"; `field_table!` grew nested-path support for it. The `#[ignore]`d `our_av1_picture_parameters_match_libavcodecs` and the capture recipe are in place, and `the_dump_and_the_parser_agree…` now self-compares AV1 too. ⚠ NO libavcodec AV1 capture was taken and the module docs say so rather than leaving an absent result to be read as a pass: `.221` has no MSYS2, no gcc and no make, so a patched FFmpeg there is a toolchain bring-up, not a build. Everything this file claims about libavcodec's AV1 side is READ out of `dxva2_av1.c` (n8.1). That reading did turn up one live divergence, recorded at `pic_av1.rs`'s `pp.width` and deliberately NOT changed: libavcodec sends `avctx->width`, which is FrameWidth (pre-superres), where this crate sends UpscaledWidth. The two are equal whenever superres is off, which is every stream that exists here, so the 250/250 result says nothing either way and a blind change would be unmeasured. **The promotion.** `(D3d11va, CODEC_AV1)` is `verified` — 250/250 delivered frames bit-identical to libavcodec on an RTX 3500 Ada AND an Intel Arc. All three places move together: the evidence arm, the module table and `every_rung_runs_and_the_unproven_ones_are_named`, whose `unproven` array loses the pair and whose proven list gains it. ⚠ This changes rung SELECTION, not just a label. `verified` is what lets `auto` pick D3D11VA ahead of Vulkan Video, so Windows Intel and unknown-vendor boxes — where the ladder is `native-d3d11va → native-vk → sw` — now decode AV1 on D3D11VA where they previously fell to Vulkan. Taken deliberately: ~10x the Vulkan leg's speed, and the parity that promoted it was measured on an Intel Arc, which is the vendor family the change moves. Still no soak on the goldens, and the notes say so. Also: `frame_av1` holds the decode's `Result` instead of `?`-ing it, so both slot releases run on the failure path. `decode_av1` notes an error and keeps the session rather than rebuilding the slot map, so an early return leaked a surface per failed frame and hit `SlotError::Full` after nine. |
||
|
|
1c54d0999b |
fix(client): the D3D11VA AV1 rung decoded every inter frame into a surface it was predicting from
AV1 applies `refresh_frame_flags` AFTER the frame is decoded (7.20), so a frame that reads a reference slot and then overwrites it is the ORDINARY case, not an exotic one: 268 of the vendored vector's 274 frames do it, first at frame 6. `plan_to_dxva_av1` released every displaced picture inside the conversion — which is what the H.264 and H.265 siblings do with their whole `removed` list — and then assigned the decode target a slot. `SlotMap::assign` takes the lowest free slot, and the lowest free slot is the one just vacated. So the submission said `CurrPicTextureIndex = N` and `RefFrameMapTextureIndex[k] = N` in the same breath, on 268 of 274 frames: decode into the surface you predict from. Neither vendored H.264 nor H.265 vector ever produces that shape (measured: zero on 250 AUs), which is why an eager release survived two hardware-proven codecs and opened on the first AV1 frame past the key frame's neighbourhood. HEVC even has the invariant under test already — `the_current_picture_is_named_by_curr_pic_and_ never_aliases_a_reference` — and AV1 had nothing. The Vulkan rung already carries the fix; this is the same contract, and the DXVA constraint is the STRICTER of the two: Vulkan binds only the references a frame names, while `RefFrameMapTextureIndex` declares the whole store, so every picture the store still names has to survive the conversion. `DecodePlanDxvaAv1` grows `release_after_decode` and `frame_av1` applies it once the decode op is issued — next to the `refresh_frame_flags == 0` release that already waits for the same reason. Peak surfaces held goes 7 of the 9 the pool allocates, so the spare slot `SlotMap::new` adds is doing exactly the job it exists for. Measured on hardware before the fix: Intel Arc got 245 of 250 delivered frames wrong — 47% of luma at the first bad frame, max |delta| 242, chroma wrong too, a frame predicted from the wrong picture — and the only late frame it got right was the one intra frame, which names no reference and so could not alias. That reads as a `primary_ref_frame` defect and is not one: PRIMARY_REF_NONE and "has no references to alias" are the same frames. |
||
|
|
6d0a389dd2 |
fix(client): the D3D11VA AV1 rung decodes wrong pixels — the parity harness existed all along
The follow-up was framed as "build the frame-hash parity harness the D3D11VA AV1 rung is missing, then flip hardware_verified to true". Both halves were wrong. The harness was never missing. `video_d3d11_native`'s `parity` module has carried `av1_every_delivered_frame_hashes_bit_identical_to_libavcodec` since M7 wired the rung — wired to the SAME libavcodec goldens the Vulkan AV1 leg passes against, with the display-order model that handles the vector's 24 hidden frames, sitting `#[ignore]`d beside the H.264/H.265/Main10 legs. It had simply never been run on a device; .173 was powered off the day it was written. What the old evidence note called a missing harness is real about pf-dxvadec the CRATE, which cannot host one — it links no D3D11 — but the device half lives here and was already done. Run on .221, it FAILS, on both GPUs, deterministically (three runs each, identical first-divergent frame and identical hashes): 186/250 diverging display frames on an RTX 3500 Ada, 245/250 on an Intel Arc. It is the decode that is wrong, not the measurement, and three independent checks say so. H.264 and H.265 pass 250/250 and HEVC Main 10 50/50 through the same harness, the same readback geometry, the same crop and the same slot map on those same two GPUs. pf-vkdecode's Vulkan AV1 leg reproduces the same golden file 250/250 on the same box. And the goldens regenerate byte-for-byte from the ffmpeg build their own header names. Two signatures, and they are not one defect wearing two faces. NVIDIA is bit-exact for display frames 0..=63 and then loses ONE 16x24 luma block — 174 pixels, max |delta| 8, chroma untouched — on the frame whose order_hint first reaches 64, after which every remaining frame is downstream of it through prediction. The stream parks the key frame (order_hint 0) in BWDREF and ALTREF2 for its whole length, so 64 is where the distance to it reaches the edge of what get_relative_dist can represent at OrderHintBits = 7. Intel is structurally wrong from display frame 4 — 47% of luma, max |delta| 242, chroma wrong too, a frame predicted from the wrong picture — and the only later frame it gets right is the one whose primary_ref_frame is PRIMARY_REF_NONE. None of this is visible on glass, which is the whole argument for goldens: the rung streams 4K60 on both parts with a clean five-minute soak at roughly ten times the Vulkan leg's decode time. The 2026-08-07 field sessions that looked clean were looking at wrong pixels. So hardware_verified stays false, and the note now says why in the strongest available terms — it prints at warn on every session that lands here, and "decodes AV1 to wrong pixels" is what a support engineer needs to read. The pair stays in `every_rung_runs_and_the_unproven_ones_are_named`'s unproven array; its note still contains NEVER, because the pair has never PASSED parity, which is now a measured statement rather than an absence. Left deliberately unchanged: `auto` on Windows can still reach this rung for AV1, and on Intel it is the arm that fires, because that vendor advertises no SAMPLED usage on any decode profile so zero-copy Vulkan Video cannot run there. Barring it trades visibly-wrong AV1 for the software rung, which cannot keep up at 4K and is itself unproven. Which way that trade goes is a product call, so it is recorded at the admission site rather than made silently here. `av1_divergence_map` is kept, cleaned up and documented: it is what turned "186 frames differ" into a lead — one line per display frame, its verdict beside the plan facts that could explain it, and an opt-in raw-NV12 dump. At a frame where one vendor hashes correctly, that vendor's bytes ARE libavcodec's bytes and so a valid reference for the other's, which is how "how badly" was answered without new goldens. The tool that would localise the rest does not exist: pf-dxvadec's libav_picparams_parity covers H.264 and HEVC only, so the AV1 conversion has never been compared against libavcodec at the picture-parameter level either. That is the next step, not another session. Also in this file, since it is the same table and the same day: the VAAPI rung's AV1 leg has now decoded 250/250 of the vendored vector on RDNA3 and its arm is split from the H.264/H.265 ones, which genuinely have still never decoded anything. It is unverified for the same reason as ever — no parity — and the D3D11VA row above is exactly why that distinction is worth keeping: a rung can decode 250 frames and still be wrong. |
||
|
|
5c05246098 |
feat: M10 — FFmpeg is gone from the client
cargo tree -p punktfunk-client-session finds no ffmpeg. The host still does, which is the whole point: pf-encode keeps libavcodec unconditionally and no host workflow, packaging script or licence file was touched. Deleted: crates/pf-ffvk, video_vulkan.rs, video_vaapi.rs, video_libav.rs, the libavcodec half of video_d3d11.rs, the av_log machinery, ffmpeg::codec::Id as the decoder's vocabulary (the quic CODEC_* wire constants now serve, which is why the evidence table was keyed on them), DecodedImage::VkFrame and ::Dmabuf, the presenter's AVVkFrame lane, and the ffmpeg-fallback feature with everything behind it. DrmFrameGuard collapses from an enum to a newtype, which removes an unsafe impl Send. Roughly 25,000 lines. Then the CI, packaging, licensing and docs work the plan's §6 lists: the Windows workflows lose FFMPEG_DIR, PF_FFVK_VULKAN_INCLUDE and their PATH prepend; the MSIX loses its DLL wildcard; the client .deb stops emitting libav sonames on its own because depends come from dpkg-shlibdeps; arch, flatpak and nix drop the dependency; and the README's "FFmpeg 7 or 8" contract narrows to the host. Three defects reached users' machines in the first cut, and none was in the deletion itself. All three desktop Settings UIs offer vulkan, vaapi and d3d11va as stored decoder values, so those strings sit in shipped settings files today. Refusing them by name — which is the correct rule for a stale pin — would have bricked every upgraded client whose owner ever touched that dropdown. They now migrate onto the native rung for the same hardware family, at decoder construction AND at each dialog's lookup, because a legacy value that matches no preset displays as "Automatic" and silently rewrites the user's preference on the next save. M9's evidence filter was deleted on the argument that with no libavcodec twin below, barring an unproven rung removes hardware decode rather than moving down one rung. That is true on Windows and false on Linux for Intel and every unknown vendor id, where prefer_vulkan_first is false and the order is native-vaapi → native-vk: a rung that has decoded nothing anywhere sitting above one that is 250/250 on three drivers. Every Intel Linux desktop would have moved from libavcodec VAAPI, shipping for years, onto pf-vaadec by default — and a rung that constructs and then produces wrong pixels leaves only by the error-streak demotion, which this codebase already documents as not tripping on the B580's strobing. The filter is restored as a narrow, pure, testable rule: an unproven rung yields to a proven one, and to nothing else. Windows deliberately passes no rung below, because that vendor family is the one with a measured wrong-pixel report against Vulkan decode, and trading no evidence for evidence of corruption is the wrong direction. And the notices still said FFmpeg was bundled. The root file is what both desktop clients include_str! and what the MSIX ships, three lines under the new card saying no FFmpeg is bundled; Apple's Acknowledgements said it too, on iOS, tvOS and macOS. The generator now emits four per-client files scoped by transitive closure — 0 FFmpeg mentions in each, verified — while the root file keeps it for the host. That also ends the standing false attribution of ffmpeg-next, GTK4, windows-rs and the NVENC SDK to an iPhone. Windows has no reachable box, so it was compiled instead: a cross clippy at -D warnings on x86_64 and aarch64-pc-windows-msvc with the C toolchain stubbed so build scripts run without linking. That gate immediately caught an include_str! path one directory too deep, which nothing else could have. Gates: container clippy -D warnings, 160 tests, workspace check, both Windows targets clean, client ffmpeg count 0 and host 2. The four decode crates are untouched, so the hardware rungs' 250/250 stands. ⚠ Owed and unrun: no GPU has executed any of this milestone. M8's on-glass software check, M7's D3D11 and VAAPI AV1 hardware legs, and M9's field bake all still want hardware, and the bake window and criteria remain the user's. |
||
|
|
38554c1c6e |
feat(client): M9's code half — native first, FFmpeg behind an off-by-default feature
`ffmpeg-fallback` on pf-client-core, default off on the crate. With it off the libavcodec rungs are not compiled, pf-ffvk leaves the dependency graph, and no ladder or demotion arm names them; with it on each sits exactly where it sits today, directly below its native twin. That is the switch which makes M10 a deletion rather than a redesign. The bake window and the regression criteria are the user's, per the plan, and nothing here claims the M9 gate is met. The hard part was not the feature, it was honesty. Two of the four native rungs have never decoded a frame on any hardware — native VAAPI at all, and native D3D11VA's AV1 leg — and making those the default would assert evidence that does not exist. So admission is per rung and per codec: a pair with hardware evidence joins `auto` always; a pair without it joins only when nothing proven is left below it (a build with no FFmpeg twin, where the alternative is not a proven rung but the CPU) or when the user asks with PUNKTFUNK_NATIVE_FIRST=1. Pins bypass it, so a lab run can still reach any rung. The shipping default therefore changes in exactly three ways, all evidence-backed: AV1 `auto` takes native Vulkan (250/250 bit-identical on an RTX 5070 Ti), Windows H.264/H.265 `auto` takes native D3D11VA above its FFmpeg twin (parity on two GPUs plus a 30-minute soak), and a failing Vulkan rung on Windows demotes to native D3D11VA first. Everything unproven is byte-for-byte as it was. The evidence state is written where it cannot rot: a table in video.rs's module docs, the same facts in code as `native_evidence()`, a test asserting them in both feature states, and a per-session log line carrying the rung, the codec, whether hardware has verified that pair and the evidence string — at WARN when it has not. A support engineer reading a log can now tell proven from assumed without asking anyone. Termination needed a new guarantee. With the FFmpeg twins gone, two native rungs in opposite per-vendor orders could hand a session back and forth forever, so a rung once entered is never re-entered and the walk is monotone to software. The never-delivered fall-through still works: with the feature on it is unchanged, and with it off it is redundant, because the next candidate already IS the rung below. ⚠ ffmpeg-next remains a hard dependency of pf-client-core, deliberately. What is left off-feature is three type-level residues — the codec-id vocabulary, the AVVkFrame guard that is pf-presenter's public import, and a pixel-format in one signature — every one of them an M10 §6 line item. Deleting them here would mean deleting the presenter's FFmpeg lane, 55 call sites, in a milestone whose gates cannot run a GPU. No libavcodec decoder is opened in a default build. ⚠ video_d3d11.rs was gated item by item rather than wholesale, and nothing in this tree compiles it — it needs a Windows check before anyone trusts it. Gates: both feature states, container clippy -D warnings and 158/159 tests, workspace check. The four decode crates are untouched, so the hardware rungs' 250/250 stands. |
||
|
|
ef40890c80 |
feat(client): native D3D11VA AV1 — wired, and four defects it exposed
The AV1 arm of the native D3D11VA rung, parity-required because today's FFmpeg d3d11va rung already decodes AV1 Profile 0 and the excision must not silently drop it. Pin-only, as that rung is today. decode() walks the temporal unit frame by frame; submit() splits into decode_into and present, because AV1 decodes frames that are never shown. The proven H.264/H.265 body is byte-for-byte unchanged — review diffed it against HEAD mechanically and found only a rename plus one refusal arm — and the VideoProcessorBlt hand-off is untouched. That mattered more than anything else here: those two codecs are hardware-proven, .173 is powered off, and no gate that runs could have caught a regression in them. Every descriptor value comes from libavcodec's dxva2_av1.c read verbatim, not from symmetry with the other codecs: three buffers and no qmatrix (AV1 transmits none), NumMBsInBuffer zero on all three, ConfigBitstreamRaw 1, surface alignment 128, pool +8, and the session sized from the SEQUENCE header's max frame size — sizing from the frame would rebuild the decoder and drop every reference the first time a stream legally resized downward. Two places where following the H.264/HEVC pattern would have been wrong. libav pads the bitstream buffer and grows only its descriptor's DataSize, never a tile's, because a tile's size is exact — charging padding to the last record is corruption, not filler. And the committed tile records were one per tile GROUP spanning the whole OBU, header and frame header included, where libav emits one per TILE addressing the payload past its tile_size_minus_1; the vendored vector is single-tile, so the old tests passed either way. Review then found four more defects in the already-committed conversion, each confirmed against libavcodec AND Chromium's D3D11 AV1 accelerator: Tile widths and heights were the coded minus-1 where the field is a superblock COUNT — every tile declared one superblock short, on every frame, with a comment asserting the opposite of the truth. StatusReportFeedbackNumber must be zero for AV1. Both reference implementations disable it specifically for this codec — libav's note reads "breaks decoding on some drivers (tested on NVIDIA 457.09)", Chromium's "it crashes :|" — while both set it for H.264 and HEVC, which is why this rung's proven codecs never showed it. It would likely have presented as a hang or a rejected submission rather than bad pixels, sending the next session after the tile records instead. frame_refs[].Index is an index INTO RefFrameMapTextureIndex, not a surface index; the neighbouring line already filled that map correctly. Measured: 1636 reference entries on the vendored vector where the two differ. qm_y/u/v need the 0xFF "no matrix" sentinel — 0 is a valid matrix index, and 274 of 274 frames transmit no quantiser matrix, so every one was being dequantized against matrix 0. Also closed: the slot leak the Vulkan rung had already found and documented (a frame refreshing no slot is never reported removed, so nine of them exhaust the ledger); a tile-grid check that could not fire, replaced with libav's own cols*rows guard; per-reference sizes now taken from the reference's own header via RefState rather than the current frame's; and the render size clamped against the decoded picture in both rungs, since AV1 permits a render size larger than the frame. The parity leg was rewired through the real decode path — it previously called the internals directly, so its hidden-frame assertion described the harness's own counter rather than production withholding anything. Gates: macOS fmt/clippy/383 tests, container clippy -D warnings over four crates and 499 tests, and on Windows .133 (.173 is powered off) clean checks plus 97 pf-dxvadec tests. All 8 Vulkan gpu_parity legs re-verified bit-exact on the RTX 5070 Ti after the shared-code change. No AV1 frame has been decoded through this rung anywhere: it needs .173 back. |
||
|
|
509843f0d4 |
test(client): the D3D11VA rung's ten-bit path, measured too
The companion to the Vulkan ten-bit leg, over the same vector and the same P010 goldens — one golden file serves both rungs because a D3D11 P010 surface and Vulkan's 3PACK16 family hold the ten bits in the same place. This is the rung where the gap mattered most. D3D11VA exposes no per-picture status query at all, so its HDR evidence was a session that built a Main10 decoder and streamed without complaint — which is precisely what a Main10 stream decoding to garbage would also produce. Now there is a number. It exercises geometry the eight-bit legs cannot reach: P010 samples are two bytes, so a row is width * 2 rather than width, and HEVC's 128-line granule pads a 240-line picture to a 256-line surface — so the chroma plane starts a long way from where the display height alone would put it. Getting either wrong is the smeared-rows failure this project has already paid for once, and it would have looked like a decoder fault. The run body now takes the stream format and the expected access-unit count rather than assuming the eight-bit envelope and 250 frames. A CPU guard pins the vector at ten bits — 4:2:0, both depths minus8 == 2, 320x240, 50 access units. A regenerated eight-bit vector would otherwise turn this into a second run of the eight-bit path under a ten-bit name, passing, because its goldens would have been regenerated with it. Hardware: HEVC Main 10 50/50 bit-identical on the RTX 4090 and on the AMD Radeon iGPU, alongside the unchanged eight-bit legs at 250/250 on both. With the Vulkan leg's two drivers that is four independent drivers across two rungs for the ten-bit path, where yesterday there were none. |
||
|
|
3041d9bb43 |
test(client): M5's decoded pixels now answer to libavcodec's
The native D3D11VA rung had no pixel evidence at all. Its DXVA bytes were checked against libavcodec's own captured bytes, and its Intel bring-up proved the driver accepts the submission — but nothing had ever compared what came out. This is that comparison, against the same goldens and the same reference the Vulkan rung was held to: libavcodec's SOFTWARE decode, which is ground truth rather than a peer implementation, so the two rungs' verdicts are now directly comparable numbers. It reads back the DECODE surface, before the VideoProcessorBlt, so what is hashed is the half this rung is responsible for; the hand-off is the shared, field-proven half and is deliberately not in the measurement. Finding, recorded rather than papered over: this rung presents in DECODE order. It never consults AuPlan::dpb.outputs — submit blits setup_slot and returns. The native Vulkan rung keeps a display-order queue for exactly that reason, and libavcodec's D3D11VA rung reorders internally, so this rung differs from both. It cannot bite on punktfunk streams, which are zero-reorder and carry no B pictures, but that is a convention of our hosts rather than a structural guarantee, and a stream that did reorder would present out of order with nothing to say so. Both vendored vectors DO reorder — the H.265 one's first B picture at AU 3 is what localised the RPS slot defect — so a harness hashing in decode order would report a permutation against display-order goldens and read like a decoder fault. Instead each decoded surface is hashed against the PicId the planner gave it and the hashes are emitted in the planner's own output order. The reordering is the test's, done by the planner the rung already trusts, and `both_vendored_vectors_really_do_reorder` asserts the reason so the docs cannot go stale silently. The crop reads the chroma plane at RowPitch * texture height, not display height: the decode pool is aligned to the codec's granule and is taller than the picture. That is the 1088-row smear this project has already paid for. Two CPU guards run in ordinary CI. This file needs its own Annex-B splitter (pf-client-core does not depend on the vendored parser), and a splitter that disagreed with pf-bitstream's would fail on hardware as a frame-count mismatch that reads like a decoder defect; instead it fails on CPU, saying so. PF_DXVA_ADAPTER pins a GPU by description substring and every run prints the adapters it saw — .173 enumerates its AMD iGPU alongside the 4090, and which one answered is a fact worth printing rather than inferring. Hardware: H.264 and H.265 both 250/250 bit-identical on NVIDIA GeForce RTX 4090 and on the AMD Radeon iGPU, Windows. Gates: clippy -D warnings and the lib tests on Windows, the Linux container's clippy/tests/workspace check, and rustfmt. |
||
|
|
31087697a9 |
feat(client): M5 — native D3D11VA decode, pin-only pending hardware
The Windows fallback rung, and auto's first choice on Intel, now has a native implementation driven by pf-bitstream's plans instead of libavcodec. New crate pf-dxvadec holds everything that can be a pure function — the DXVA structure layouts, both codec conversions, bitstream packing, config selection — deliberately CROSS-PLATFORM, because a cfg(windows) module is verified by a remote cargo check and nothing else, and this milestone's riskiest code is exactly the part no local test can see. Only the FFI lives in video_d3d11_native.rs. windows-rs does not generate dxva.h at the pinned rev, so the DXVA structures are hand-declared: compile-time assertions on every struct size AND every field offset, packed bitfield words as plain integers with named builders and the bit positions written beside the C declaration, and a const zeroed() per struct so construction needs no unsafe at all. The crate's only unsafe is a sealed byte view over those PODs. Review round 13 checked all seven layouts field by field in declaration order — sizes, widths, array lengths, the PicEntry index/flag packing, and every named bit's position and width. The decode pool reproduces libavcodec's rather than inventing one: ONE texture with ArraySize = pool size, BIND_DECODER and nothing else, MiscFlags 0, aligned 16 for H.264 and 128 for HEVC. That is deliberate. This rung's predecessor records that a hand-built pool which validated on NVIDIA was rejected by Intel at the first SubmitDecoderBuffers — and Intel is the vendor this rung exists for. The VideoProcessorBlt into shareable RGBA is untouched: importing a multiplanar NV12 D3D11 texture into Vulkan device-losts on NVIDIA, so that hand-off is load-bearing field-proven code. It was extracted into a shared HandoffRing so both rungs fill one implementation; the review diffed the blit statement by statement, including the keyed-mutex pairing. Review round 13's four defects are fixed. The blocking one: the HEVC quantisation matrix was submitted unconditionally, and the vendored parser leaves it ALL ZEROS unless the stream codes one — unlike FFmpeg, which seeds the spec defaults. On a stream saying 'use the default matrices' the driver is obliged to apply what it is handed, so every residual would dequantise to zero and the picture would drift to flat prediction. It is now gated on scaling_list_enabled_flag exactly as libav gates it, with the Table 7-5/7-6 defaults supplied when enabled but uncoded. Second: NumMBsInBuffer was 0 where libav's H.264 path sets mb_width * mb_height. This module's whole method is verbatim reproduction on precisely the call that once failed for Intel, so an omitted descriptor field is the same class of bug as the pool. Third, and the one to watch on hardware: RefFrameList carried the frame's reference set rather than the pictures marked used for reference. Vulkan defines pReferenceSlots as the slots this operation uses, so a subset is correct there; DXVA defines RefFrameList as a statement about the DPB. The list DERIVATION survives a subset — which is exactly why a smoke test would have passed — but a long-term reference held across frames that none of them name would vanish and reappear, and a driver keeping per-reference state is entitled to discard it in between. That is the Ally X symptom shape. pf-bitstream now exposes a per-AU DPB snapshot for both codecs and the converters build the array from it, frame references first, marked tail appended. 121 of the 250 vendored AUs carry a marked picture the frame never names, so this is exercised, not theoretical. Fourth: the session identity omitted bit depth and chroma, while the Windows host flips an HDR desktop to PQ in-band with a new SPS — a depth change at unchanged size would have decoded 10-bit samples into an NV12 pool. Identity now derives from the SPS per AU and rebuilds. Wired PIN-ONLY (PUNKTFUNK_DECODER=native-d3d11va), absent from every auto arm. Nothing has decoded a frame yet, and M2's discipline was that auto admission comes only after hardware parity. A runtime streak demotes to the FFmpeg D3D11VA rung first, then software. Also scaffolded: a byte-diff harness against libavcodec's own DXVA picture parameters, with the FFmpeg patch and capture recipe in its docs. Nothing here is checked against libav's actual bytes the way M3 was checked against its pixels, and that is the cheap way to buy the confidence before hardware. Gates: fmt clean; container clippy -D warnings zero across pf-client-core + pf-presenter + pf-vkdecode + pf-dxvadec + punktfunk-core; tests 73/131/63/129/354 green; cargo check --workspace clean; Windows cargo check and clippy -D warnings clean on .173. |