deeb8b67000e4faa4939f736a3a21004b47e249a
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ef40890c80 |
feat(client): native D3D11VA AV1 — wired, and four defects it exposed
The AV1 arm of the native D3D11VA rung, parity-required because today's FFmpeg d3d11va rung already decodes AV1 Profile 0 and the excision must not silently drop it. Pin-only, as that rung is today. decode() walks the temporal unit frame by frame; submit() splits into decode_into and present, because AV1 decodes frames that are never shown. The proven H.264/H.265 body is byte-for-byte unchanged — review diffed it against HEAD mechanically and found only a rename plus one refusal arm — and the VideoProcessorBlt hand-off is untouched. That mattered more than anything else here: those two codecs are hardware-proven, .173 is powered off, and no gate that runs could have caught a regression in them. Every descriptor value comes from libavcodec's dxva2_av1.c read verbatim, not from symmetry with the other codecs: three buffers and no qmatrix (AV1 transmits none), NumMBsInBuffer zero on all three, ConfigBitstreamRaw 1, surface alignment 128, pool +8, and the session sized from the SEQUENCE header's max frame size — sizing from the frame would rebuild the decoder and drop every reference the first time a stream legally resized downward. Two places where following the H.264/HEVC pattern would have been wrong. libav pads the bitstream buffer and grows only its descriptor's DataSize, never a tile's, because a tile's size is exact — charging padding to the last record is corruption, not filler. And the committed tile records were one per tile GROUP spanning the whole OBU, header and frame header included, where libav emits one per TILE addressing the payload past its tile_size_minus_1; the vendored vector is single-tile, so the old tests passed either way. Review then found four more defects in the already-committed conversion, each confirmed against libavcodec AND Chromium's D3D11 AV1 accelerator: Tile widths and heights were the coded minus-1 where the field is a superblock COUNT — every tile declared one superblock short, on every frame, with a comment asserting the opposite of the truth. StatusReportFeedbackNumber must be zero for AV1. Both reference implementations disable it specifically for this codec — libav's note reads "breaks decoding on some drivers (tested on NVIDIA 457.09)", Chromium's "it crashes :|" — while both set it for H.264 and HEVC, which is why this rung's proven codecs never showed it. It would likely have presented as a hang or a rejected submission rather than bad pixels, sending the next session after the tile records instead. frame_refs[].Index is an index INTO RefFrameMapTextureIndex, not a surface index; the neighbouring line already filled that map correctly. Measured: 1636 reference entries on the vendored vector where the two differ. qm_y/u/v need the 0xFF "no matrix" sentinel — 0 is a valid matrix index, and 274 of 274 frames transmit no quantiser matrix, so every one was being dequantized against matrix 0. Also closed: the slot leak the Vulkan rung had already found and documented (a frame refreshing no slot is never reported removed, so nine of them exhaust the ledger); a tile-grid check that could not fire, replaced with libav's own cols*rows guard; per-reference sizes now taken from the reference's own header via RefState rather than the current frame's; and the render size clamped against the decoded picture in both rungs, since AV1 permits a render size larger than the frame. The parity leg was rewired through the real decode path — it previously called the internals directly, so its hidden-frame assertion described the harness's own counter rather than production withholding anything. Gates: macOS fmt/clippy/383 tests, container clippy -D warnings over four crates and 499 tests, and on Windows .133 (.173 is powered off) clean checks plus 97 pf-dxvadec tests. All 8 Vulkan gpu_parity legs re-verified bit-exact on the RTX 5070 Ti after the shared-code change. No AV1 frame has been decoded through this rung anywhere: it needs .173 back. |
||
|
|
cdd1f3efce |
fix(vkdecode): AV1 is bit-exact — the bug was a use-after-free, not the driver
250/250 frames bit-identical to libavcodec on NVIDIA 610.57.04, and all four other parity legs (H.264, H.265, Main 10, both four-byte-prefix twins) still green. session_av1 built the sequence header, handed pStdSequenceHeader to vkCreateVideoSessionParametersKHR, and dropped the backing the instant the call returned — on the documented assumption that Vulkan copies parameter data before returning. NVIDIA does not. It keeps the pointer and dereferences pColorConfig when a decode is RECORDED. The freed block became our own next allocation, whose bytes read back as mono_chrome = 1, and a monochrome frame skips exactly loop_filter_level[2..3] (AV1 7.14). That is the whole fingerprint two earlier rounds chased: luma bit-exact, chroma off by small amounts, and rewriting the chroma levels in the bitstream changing nothing — the driver read them correctly and then discarded them, because it believed the stream had no chroma. StoredParamsAv1 now holds the parameters object and its Std backing in one value, so an object whose backing is gone is unrepresentable. The road there is worth recording, because two well-evidenced conclusions were wrong before this one was right. A software oracle reproduced the divergence exactly by disabling chroma deblocking, and a GPU probe showed chroma levels [8,12] and [63,63] producing byte-identical output — which looked conclusive and was not. libavcodec's own Vulkan AV1 hwaccel is bit-exact on this same driver, which proved the hardware fine and the defect ours. ffmpeg never hits it: with VK_KHR_video_maintenance2 it uses inline session parameters and never creates a parameters object at all. The proof is direct rather than inferred: a throwaway Vulkan capture layer dumped both submissions and every byte of our AV1 picture info already matched libavcodec's, including the loop filter block; only the session parameters layer differed. Watching the block's address showed correct bytes at create and our next allocation at decode. Ruled out on hardware, so nobody re-tests them: filmGrainSupport, maxCodedExtent, maxDpbSlots/maxActiveReferences, VkVideoDecodeUsageInfoKHR, the tile-start sentinel, the setup slot's SavedOrderHints, a NULL pTimingInfo, and heap luck. Two earlier fixes are confirmed against libavcodec's captured wire bytes and kept: CDEF secondary strengths carry the coded value rather than the spec's in-place fixup, and LoopRestorationSize is log2-based. The refuted driver-ignores-chroma-levels claim is corrected everywhere it was written down, and that probe test now passes and points at the lifetime of everything a submission points at before blaming a vendor. ⚠ Adjacent and NOT fixed: session.rs and session_h265.rs drop their Std backings the same way, and those sets carry embedded pointers too. Both are measured bit-exact on four drivers, so nothing is known to be wrong — but the contract now rests on a driver behaviour measured FALSE for AV1 on a shipping driver. The SAFETY comments asserting it have been corrected; the structure is deliberately untouched pending its own pass. Gates: macOS fmt/clippy/336 tests, container clippy -D warnings, all green; 8/8 gpu_parity and 3/3 gpu_smoke legs verified on the RTX 5070 Ti. |
||
|
|
a404830456 |
feat(client): wire AV1 into the native Vulkan rung, pin-only
The third codec arm in video_vk_native, AV1 admitted to native_codec and to native_vulkan_gate by pin only. It stays out of `auto` on the same rule M5's D3D11VA rung follows: `auto` admission is earned with hardware evidence, and this has decoded nothing on a device. is_integrity_warning_av1 did not exist, so the client could not have concealed AV1 damage at all. Added, exhaustive, no wildcard: all three AV1 warnings really are damage, because AV1 has no spec-legal-but-noisy signal to mis-classify — no reorder envelope to announce, no MMCO to rebase — and the exhaustive match is what stops a future variant defaulting to clean. The blocking defect review found was two safety mechanisms cancelling each other. After a failure the decoder skipped to the next key frame answering Ok(None), and because AV1's planner has no flush its store kept planning cleanly, so those AUs carried no warnings and the client read them as proof the rung works — clearing the demotion streak and resetting its clock on every one. The streak could then never reach the threshold, which made the never-delivered fall-through to FFmpeg-Vulkan unreachable, which is the documented backstop for exactly three things: a level above maxLevelIdc, a sequence header disagreeing with the Welcome, and film grain. Film grain is the probe's own admitted assumption, so a grain stream would have frozen the screen for the session while DecodeHealth reported run 0 — recovered. AV1 now answers the wait with an error, as H.264 and H.265 already do through AwaitingIdr, so all three codecs are indistinguishable to the demotion machinery. That matters more than the extra precision of a third state: only the H.26x paths have hardware evidence, and they are proven WITH that behaviour. The obvious form of that fix would have wedged the decoder. A key frame can sit behind a skipped frame inside the same temporal unit — the vendored vector has 24 two-frame units — so erroring out of the per-plan loop would never reach it and the wait would never end. Skips are therefore counted per frame and the error raised only when the whole unit was skipped, with the metadata-only unit staying a clean Ok(None). Also closed: a refused temporal unit left an already-decoded frame in the ready queue, which shipped on the next AU as a clean success — putting a picture from a refused AU on screen, clearing the streak again, and latching delivered so the fall-through was disabled for good. The error arm now drains and releases unshown. MAX_DELIVERABLE is derived rather than picked: HOLD_HEADROOM minus the pipeline's own hold, pinned to pf-vkdecode's constant so a hardcoded depth fails the build. At the previous 8 the queue plus the presenter's 4-7 stood against a headroom of 8, so it capped memory without preventing the exhaustion it named, and a frame waiting 8 AUs burned 16 of the 17 query slots — where a re-armed slot reads as Failed and becomes a fabricated driver-corruption verdict in the very counter the Ally X signal lives in. The trim now runs after this AU's frame is taken, or at the derived depth it would drop a two-output unit's first frame and invert display order inside one AU. Its justification was also wrong: the claim that a temporal unit may carry a show_existing_frame alongside a shown frame is disproved by this repo's own golden — 250 units, 250 shown, zero show_existing. The bound is kept as defence in depth against a non-conformant or multi-operating-point stream, and now says so. Gates: macOS fmt/clippy/392 tests, container clippy -D warnings over six crates, 851 tests, workspace check. No hardware: the rung is pin-only and has still never decoded a frame on a device. |
||
|
|
cab3aa1726 |
feat(vkdecode): M7's Vulkan AV1 rung — GPU half, and the review that saved it
caps_av1 / session_av1 / decoder_av1, over the CPU half already committed, sharing the picture pool, bitstream ring, op ring, DPB settling and frame delivery with H.264 and H.265 rather than forking them. AV1 session parameters carry exactly one sequence header — no PPS, no VPS — so the parameters ledger is two-state: current, or recreate. The GPU plumbing came through review clean. The damage was all in the conversion committed two rounds ago, which nothing tested against a reference, and none of it would have failed a gate: clippy was clean, the tests were green, and the rung would have decoded its own conformance vector wrong on essentially every frame on AMD, silently. Four blocking defects, each measured on the vendored vector rather than argued: Nine StdVideoDecodeAV1PictureInfo flags were never set. Four change reconstruction — allow_screen_content_tools on 274 frames of 274, allow_warped_motion on 273, is_filter_switchable on 172, force_integer_mv on 1 — and RADV reads three of them directly. The block already set allow_intrabc, which is only codeable when screen-content tools are on, so it contradicted itself. LoopRestorationSize sent the pixel size where the field is log2(size) - 5. cros-codecs stores 64/128/256; RADV names its destination log2_restoration_size_minus5 and reads 1/2/3. Nothing truncates, nothing errors, and every frame with loop restoration reconstructs against a nonsense unit size. Per-reference Std info answered questions about the wrong picture: every reference carried the CURRENT frame's type, and RefFrameSignBias was never set at all. Sign bias is what tells a decoder a reference lies in the future, and this vector is the hidden-ALTREF one, so all-zero meant every reference was treated as past. Fixed at the source: pf-bitstream now records a RefState when a picture is stored — its own frame type, sign-bias mask, saved order hints — and carries it on the slot, so all three backends get answers about the reference rather than about the frame reading it. Film grain's six chroma-scaling fields were zero, which defeats the profile machinery that exists to refuse devices unable to synthesise grain. The reference-name compaction is fixed in the PLANNER, once. AuPlan::refs is now name-indexed with holes preserved, so a lost reference can no longer renumber every later AV1 reference name — a class that was live in both conversions and armed for the VAAPI rung that does not exist yet. The DXVA twin had a second name-versus-slot confusion: it read global motion by DPB slot from an array the spec indexes by reference name, and slot 0's matrix is all-zero rather than identity, so 273 references were given a zero warp. Also closed: pTileOffsets/pTileSizes were sized to tileCount while RADV reads AV1_MAX_NUM_TILES entries unconditionally — a 4-byte allocation read a kilobyte deep — now fixed 256-entry arrays with zeroed tails. And the test guarding the lost-reference refusal re-implemented the predicate inline, so deleting the guard left it green; both now call one named function. The bitstream layout now matches libavcodec: raw tile payloads only, frameHeaderOffset 0. The review established the spec-literal layout was NOT wrong — AV1 has no start-code scanning, so the 3-versus-4-byte and slices-only scars do not transfer, and no driver in the fleet reads frameHeaderOffset — but matching the validated reference deletes code, uploads 5835 fewer bytes over the vector, and removes the untested-driver tail. Upstream, and the third of its kind: the vendored parser writes ref_frame_sign_bias[i] in the same loop body where it writes order_hints[LAST_FRAME + i], so its array is shifted one down and index 7 is never written. Corrected in RefState::of with the shift documented, the vendored tree untouched, and pinned by a test that recomputes the bias from order_hints through the parser's own get_relative_dist. Gates: macOS fmt/clippy/tests, container clippy -D warnings over six crates, 845 tests, workspace check. No hardware: nothing here has reached a driver. |