feat(dxvadec): the AV1 picparams harness AV1 forgot, and the D3D11VA AV1 rung is promoted

Two halves.

**The harness.** `libav_picparams_parity` covered H.264 and HEVC only, which is
exactly the gap that let a wrong AV1 submission ship. It now plans, converts and
packs all 274 frames of the vendored AV1 vector and checks what needs no capture:
the three-buffer descriptor set with no quantization matrix (AV1's matrices are
selected by index, so `dxva2_av1_end_frame` passes NULL/0 and there is no buffer
to submit), no macroblock count anywhere, the 912-byte picture-parameter buffer,
and the tile records — which unlike H.264/HEVC slice records do NOT abut, because
a `DXVA_Tile_AV1` addresses a tile PAYLOAD and consecutive payloads are separated
by their `tile_size_minus_1` fields.

The one that matters most is `no_av1_submission_names_its_decode_surface_in_the_
reference_store`: the invariant the previous commit fixed, over the submitted
BYTES rather than over the plan. libavcodec cannot produce that shape — it fills
`RefFrameMapTextureIndex` from the pre-refresh store and takes
`CurrPicTextureIndex` from a frame the reference update has not run on — which is
the argument for calling it a defect rather than a convention.

`AV1_FIELDS` reaches into the eight nested blocks (`tiles.widths`,
`segmentation.feature_data`, …) so a future capture reports a field and not "260
bytes of tiles differ"; `field_table!` grew nested-path support for it. The
`#[ignore]`d `our_av1_picture_parameters_match_libavcodecs` and the capture recipe
are in place, and `the_dump_and_the_parser_agree…` now self-compares AV1 too.

⚠ NO libavcodec AV1 capture was taken and the module docs say so rather than
leaving an absent result to be read as a pass: `.221` has no MSYS2, no gcc and no
make, so a patched FFmpeg there is a toolchain bring-up, not a build. Everything
this file claims about libavcodec's AV1 side is READ out of `dxva2_av1.c` (n8.1).
That reading did turn up one live divergence, recorded at `pic_av1.rs`'s
`pp.width` and deliberately NOT changed: libavcodec sends `avctx->width`, which is
FrameWidth (pre-superres), where this crate sends UpscaledWidth. The two are equal
whenever superres is off, which is every stream that exists here, so the 250/250
result says nothing either way and a blind change would be unmeasured.

**The promotion.** `(D3d11va, CODEC_AV1)` is `verified` — 250/250 delivered frames
bit-identical to libavcodec on an RTX 3500 Ada AND an Intel Arc. All three places
move together: the evidence arm, the module table and
`every_rung_runs_and_the_unproven_ones_are_named`, whose `unproven` array loses the
pair and whose proven list gains it.

⚠ This changes rung SELECTION, not just a label. `verified` is what lets `auto`
pick D3D11VA ahead of Vulkan Video, so Windows Intel and unknown-vendor boxes —
where the ladder is `native-d3d11va → native-vk → sw` — now decode AV1 on D3D11VA
where they previously fell to Vulkan. Taken deliberately: ~10x the Vulkan leg's
speed, and the parity that promoted it was measured on an Intel Arc, which is the
vendor family the change moves. Still no soak on the goldens, and the notes say so.

Also: `frame_av1` holds the decode's `Result` instead of `?`-ing it, so both slot
releases run on the failure path. `decode_av1` notes an error and keeps the
session rather than rebuilding the slot map, so an early return leaked a surface
per failed frame and hit `SlotError::Full` after nine.
This commit is contained in:
2026-08-07 21:07:23 +02:00
parent 1c54d0999b
commit f4dda9074b
4 changed files with 763 additions and 118 deletions
+48 -58
View File
@@ -49,7 +49,7 @@
//! | native Vulkan Video | | H.265 (Main / Main10 / 4:4:4) | **yes** — same parity run + HDR chain and Deck/VanGogh legs (M3) |
//! | native Vulkan Video | | AV1 | **yes** — 250/250 bit-identical to libavcodec on an RTX 5070 Ti (M7); ONE vendor, no soak |
//! | native D3D11VA | [`crate::video_d3d11_native`] | H.264, H.265 | **yes** — frame-hash parity on an RTX 4090 and an AMD iGPU + a 30-minute soak (M5), re-confirmed 250/250 (+ 50/50 Main 10) on an RTX 3500 Ada and an Intel Arc on 2026-08-07 |
//! | native D3D11VA | | AV1 | **NO — it decodes WRONG PIXELS.** The parity harness that was missing turned out to exist (`video_d3d11_native`'s `parity` module, written by M7 and never run); running it on 2026-08-07 failed on BOTH GPUs of `.221`, deterministically: 186/250 diverging frames on an RTX 3500 Ada and 245/250 on an Intel Arc. Not the environment — H.264, H.265 and HEVC Main 10 pass 250/250/50 through the SAME harness on the same two GPUs, and pf-vkdecode's Vulkan AV1 leg reproduces the SAME goldens 250/250 on the same box. Two unlike signatures: NVIDIA is bit-exact for 63 frames and then loses ONE 16x24 luma block (174 px, max |delta| 8) on the frame whose `order_hint` first reaches 64, which then propagates; Intel is structurally wrong from display frame 4 (47% of luma, max |delta| 242). It still streams — 4K60 on both, a clean 5-minute Arc soak, ~10x the Vulkan leg's speed — which is exactly why the picture looked fine and only the goldens caught it. See `av1_divergence_map` |
//! | native D3D11VA | | AV1 | **yes** — 250/250 delivered frames bit-identical to libavcodec on an RTX 3500 Ada AND an Intel Arc (2026-08-07). It got there from 186/250 and 245/250 DIVERGING frames on those same two GPUs: `plan_to_dxva_av1` released the picture this frame's own refresh displaces before assigning the decode target its slot, and `SlotMap::assign` hands back the slot just vacated — so 268 of the vector's 274 frames named one surface as both `CurrPicTextureIndex` and a `RefFrameMapTextureIndex` entry. Intel followed the aliased surface (structurally wrong from display frame 4); NVIDIA tolerated it until the `order_hint` wrap at 64 made one 16x24 luma block depend on it. ONE defect, two driver tolerances — the two unlike signatures were not two bugs. TWO vendors, still NO soak on the goldens: the 5-minute 4K60 soak this row used to cite measured throughput, and "streams cleanly" was true throughout the failure |
//! | native VAAPI | [`crate::video_vaapi_native`] | AV1 | **not proven** — but it has now DECODED: 250/250 frames of the vendored AV1 vector on `.25` (Radeon 780M, RDNA3, Mesa 26.0.3) on 2026-08-07, NV12 on a tiled AMD modifier, and `probe_this_machines_libva` reports `AV1 Profile 0: VLD decode`. Never frame-hash parity-checked: the rung exports a tiled dmabuf with no CPU-readable image, so parity needs a readback path that does not exist yet |
//! | native VAAPI | | H.264, H.265 | **NO** — these two legs have still never decoded a frame anywhere (M6/M7) |
//! | software | `video_software` | H.264, AV1 | **not proven** — openh264 has never run on glass; rav1d HAS now decoded 1080p and 4K60 AV1 there (2026-08-07, .21) and recovers in-session from a mid-stream reference loss, but with no parity check and no soak. Its 4K "abort" was never about 4K: rav1d 1.1.0 kills the process on ANY decode error while it holds a single frame context, so `video_software` opens it with two — see [`crate::video_software`] |
@@ -1111,46 +1111,38 @@ pub fn native_evidence(rung: NativeRung, wire: u8) -> RungEvidence {
true,
"frame-hash parity on an RTX 4090 and an AMD iGPU + 30-min soak (M5)",
),
// 2026-08-07: this pair was PARITY-CHECKED for the first time, and it FAILED.
// 2026-08-07: this pair failed its first parity check and now PASSES it, on both of
// the box's GPUs, after one defect was fixed.
//
// The harness the previous note said did not exist did exist — `video_d3d11_native`'s
// `parity` module, written by M7 against the same libavcodec goldens the Vulkan rung
// uses, `#[ignore]`d and never once run on a device. Running it on `.221` failed on
// BOTH GPUs and did so deterministically (three runs each, identical first-divergent
// frame and identical hashes): 186/250 diverging display frames on an RTX 3500 Ada,
// 245/250 on an Intel Arc.
// The failure was 186/250 diverging display frames on an RTX 3500 Ada and 245/250 on
// an Intel Arc, deterministic on three runs each. The two signatures looked like two
// defects — NVIDIA bit-exact through display frame 63 and then one 16x24 luma block
// (max |delta| 8, chroma untouched) at the frame whose `order_hint` first reaches 64;
// Intel structurally wrong from display frame 4 (47% of luma, max |delta| 242, chroma
// wrong too) with only its one PRIMARY_REF_NONE frame right. They were ONE defect and
// two driver tolerances.
//
// It is the DECODE that is wrong, not the measurement. Three things rule the harness
// and the box out: H.264 and H.265 pass 250/250 and HEVC Main 10 50/50 through the
// SAME harness, the same readback geometry and the same slot map on those same two
// GPUs; pf-vkdecode's Vulkan AV1 leg reproduces the SAME golden file 250/250 on the
// same box; and the goldens themselves reproduce byte-for-byte from ffmpeg 8.1.1.
// `plan_to_dxva_av1` released the pictures this frame's own `refresh_frame_flags`
// displaces INSIDE the conversion, then assigned the decode target a slot — and
// `SlotMap::assign` takes the lowest free slot, which is the one just vacated. So the
// submission named the same surface as `CurrPicTextureIndex` and as a
// `RefFrameMapTextureIndex` entry, on 268 of the vector's 274 frames: decode into the
// surface you are predicting from. AV1 applies `refresh_frame_flags` AFTER the frame
// is decoded (7.20), so that shape is ordinary rather than exotic; neither vendored
// H.264 nor H.265 vector ever produces it, which is why an eager release survived two
// hardware-proven codecs. Intel followed the aliased surface, NVIDIA tolerated it
// until the order-hint wrap made one block's prediction depend on it. `pf_dxvadec`'s
// `the_decode_target_never_aliases_a_surface_the_submission_names` is the CPU guard.
//
// Two signatures, and they are not the same defect wearing two faces:
// * NVIDIA is bit-exact for display frames 0..=63 and then loses ONE 16x24 luma
// block — 174 pixels, max |delta| 8, chroma untouched — on the frame whose
// `order_hint` first reaches 64, after which every remaining frame is downstream
// of it. The stream parks the key frame (`order_hint` 0) in BWDREF and ALTREF2 for
// its whole length, so 64 is where the distance to it reaches the edge of
// `get_relative_dist`'s range at `OrderHintBits = 7`.
// * Intel is structurally wrong from display frame 4 — 47% of luma, max |delta| 242,
// chroma wrong too, i.e. predicted from the wrong picture — and the only later
// frame it gets right is the one whose `primary_ref_frame` is PRIMARY_REF_NONE.
//
// None of this shows on glass: the rung streams 4K60 on both parts with a clean
// 5-minute soak at ~10x the Vulkan leg's speed. That is the point of a golden.
//
// What would localise it does not exist: pf-dxvadec's `libav_picparams_parity`
// covers H.264 and HEVC only, so the AV1 conversion has never been compared against
// libavcodec at the picture-parameter level either. `av1_divergence_map` in
// `video_d3d11_native` carries the per-frame evidence.
// ⚠ What this pair still does NOT have, unlike the H.264/HEVC one above: a soak on
// the goldens. The 5-minute 4K60 soak this note used to lean on measured throughput,
// not pixels, and "streams cleanly" is exactly what was true while 186 frames were
// wrong.
(NativeRung::D3d11va, CODEC_AV1) => (
false,
"streams 4K60 on an RTX 3500 Ada AND an Intel Arc with a clean 5-min soak, but \
has NEVER passed frame-hash parity and now measurably FAILS it: 186/250 and \
245/250 display frames diverge from libavcodec on those two GPUs (2026-08-07), \
while H.264/H.265/Main10 pass through the same harness - this rung decodes AV1 \
to wrong pixels (M7)",
true,
"250/250 delivered frames bit-identical to libavcodec on an RTX 3500 Ada AND an \
Intel Arc (2026-08-07), after fixing a decode target that aliased a reference \
surface on 268 of 274 frames - two vendors, no soak (M7)",
),
// 2026-08-07: the VAAPI rung decoded its first frames ever — 250/250 of the vendored
// AV1 vector on `.25` (Radeon 780M, RDNA3, Mesa 26.0.3), NV12 on a tiled AMD
@@ -3397,20 +3389,20 @@ mod tests {
/// read against. Which of them `auto` may pick FIRST is
/// [`native_rung_admitted`]'s decision, asserted in the test after this one.
///
/// `(D3d11va, AV1)` stays here after 2026-08-07 even though it has now decoded on
/// hardware, and its warn line is why: it named the rung as unproven moments before that
/// rung failed 72 access units running, which is exactly the job this test protects. (The
/// cause was the HOST shipping half of every AV1 frame — `pf_encode`'s
/// `resolve_split_subframe` — not the rung.) Unproven is about EVIDENCE, not about whether
/// it has ever worked: one 25-second session with no parity check and no soak must not
/// promote a rung past Vulkan Video in the admission filter.
/// `(D3d11va, AV1)` LEFT this list on 2026-08-07, and the bar it had to clear is the
/// point. It had already decoded on hardware twice over — a 25-second session and a
/// 5-minute 4K60 soak on two GPUs — and it stayed unproven through both, because
/// unproven is about EVIDENCE and neither run looked at a pixel. What moved it was the
/// frame-hash parity check: 250/250 delivered frames bit-identical to libavcodec on an
/// RTX 3500 Ada AND an Intel Arc. Its first run of that check FAILED on both (186/250 and
/// 245/250 diverging), while the rung streamed 4K60 the whole time — so the warn line
/// this test protects was telling the truth right up to the run that retired it.
#[test]
fn every_rung_runs_and_the_unproven_ones_are_named() {
let unproven = [
(NativeRung::Vaapi, CODEC_H264),
(NativeRung::Vaapi, CODEC_HEVC),
(NativeRung::Vaapi, CODEC_AV1),
(NativeRung::D3d11va, CODEC_AV1),
];
for (rung, codec) in unproven {
let e = native_evidence(rung, codec);
@@ -3425,8 +3417,7 @@ mod tests {
"{} / {codec:#x}: the note is what the session log prints at warn — it \
must name plainly what this pair has NEVER had, whether that is a \
hardware run at all (VAAPI H.264/H.265) or the parity check that would \
promote it (VAAPI AV1, which HAS decoded, and D3D11VA AV1, which has \
decoded and then FAILED that check on two GPUs), got {:?}",
promote it (VAAPI AV1, which HAS decoded), got {:?}",
rung.name(),
e.note
);
@@ -3439,6 +3430,7 @@ mod tests {
(NativeRung::Vulkan, CODEC_AV1),
(NativeRung::D3d11va, CODEC_H264),
(NativeRung::D3d11va, CODEC_HEVC),
(NativeRung::D3d11va, CODEC_AV1),
] {
assert!(native_evidence(rung, codec).verified, "{}", rung.name());
}
@@ -3491,6 +3483,15 @@ mod tests {
(NativeRung::Vulkan, CODEC_AV1),
(NativeRung::D3d11va, CODEC_H264),
(NativeRung::D3d11va, CODEC_HEVC),
// Joined this list on 2026-08-07 with 250/250 on two vendors. ⚠ It is a
// BEHAVIOUR change on real machines and not only a label: this pair was the
// one the filter barred under `Some(Vulkan)`, so on Windows Intel and
// unknown-vendor boxes — where the ladder is `native-d3d11va → native-vk → sw`
// — `auto` now decodes AV1 on D3D11VA where it previously fell to Vulkan
// Video. Taken deliberately: the rung is ~10x the Vulkan leg's speed, and the
// parity that promoted it was measured on an Intel Arc, which is exactly the
// vendor family that change moves.
(NativeRung::D3d11va, CODEC_AV1),
] {
for below in [
None,
@@ -3505,17 +3506,6 @@ mod tests {
);
}
}
// Windows, Intel/unknown auto: the DXVA AV1 leg has never run, and the ladder
// passes `None` there on purpose — that vendor family is the one with a measured
// wrong-pixel report against Vulkan decode, so what is really below it is the CPU.
// This asserts the ARGUMENT the call site passes, which is where the judgement
// lives; `Some(Vulkan)` would bar it, and that is deliberately not what it passes.
assert!(native_rung_admitted(NativeRung::D3d11va, CODEC_AV1, None));
assert!(!native_rung_admitted(
NativeRung::D3d11va,
CODEC_AV1,
Some(NativeRung::Vulkan)
));
// The CPU rung is last everywhere, so nothing is ever below it and it always runs
// — including for a codec it has no decoder for, which is `last_rung_verdict`'s
// problem and not the filter's.
+90 -53
View File
@@ -20,17 +20,21 @@
//! * **H.264 and H.265** — frame-hash parity against libavcodec on an RTX 4090 and an AMD
//! iGPU plus a 30-minute soak (M5), re-confirmed on an RTX 3500 Ada and an Intel Arc on
//! 2026-08-07 (250/250 both codecs, plus 50/50 HEVC Main 10 on both).
//! * **AV1** — wired in M7. It streams: 4K60 on an RTX 3500 Ada and on an Intel Arc, with a
//! clean 5-minute soak. But it **fails frame-hash parity on both of those GPUs**, measured
//! 2026-08-07 — 186/250 diverging frames on the NVIDIA part and 245/250 on the Intel one,
//! deterministically, against the same libavcodec goldens the Vulkan rung reproduces
//! 250/250 on the SAME box. So this rung's AV1 leg produces wrong pixels and the session
//! log says so at `warn`. `av1_divergence_map` (below) carries the two signatures; the
//! tool that would localise it — an AV1 leg for pf-dxvadec's `libav_picparams_parity`,
//! which covers only H.264 and HEVC — does not exist yet.
//! * **AV1** — wired in M7, and frame-hash parity on the SAME two GPUs since 2026-08-07:
//! 250/250 delivered frames bit-identical to libavcodec on the RTX 3500 Ada and on the
//! Intel Arc. It streams 4K60 on both with a clean 5-minute soak, but that is throughput
//! and not pixels — the leg streamed exactly as cleanly while 186 and 245 of those 250
//! frames were WRONG, which is what the first run of this harness measured on 2026-08-07
//! and what `av1_divergence_map` (below) records. The defect was one line of DPB
//! bookkeeping in [`pf_dxvadec::plan_to_dxva_av1`]: it released the picture this frame's
//! own `refresh_frame_flags` displaces before assigning the decode target a slot, and
//! `SlotMap::assign` hands back the slot just vacated, so 268 of the vector's 274 frames
//! named one surface as both `CurrPicTextureIndex` and a `RefFrameMapTextureIndex` entry.
//! [`NativeD3d11Decoder::frame_av1`] now applies the conversion's
//! `release_after_decode` once the decode op is issued.
//!
//! Until M10 `auto` skipped AV1 here in favour of the libavcodec rung below; with that
//! gone the alternative is the CPU, so it still runs.
//! ⚠ Still no SOAK on the goldens, so this leg's evidence is one 250-frame vector on two
//! vendors — narrower than the H.264/H.265 legs above.
//!
//! A refusal or an init failure logs and falls through to the standard ladder, so neither the
//! pin nor the `auto` admission can cost a session its decoder.
@@ -572,43 +576,19 @@ impl NativeD3d11Decoder {
return self.show_existing_av1(plan);
}
let sub = self.plan_frame_av1(au, plan)?;
let shown = if damaged {
// Converted (so the slot map stayed in step with the planner's store),
// deliberately not submitted (fn docs).
//
// ⚠ And the surface's `held` entry is CLEARED rather than left. The slot
// map now says this slot holds THIS picture, while the surface still
// carries whatever the previous occupant decoded; a later
// `show_existing_frame` naming it would find the old picture's facts and
// blit the old picture's pixels. `None` makes that path return
// `Ok(None)` — nothing shown — which is what the unit's concealment
// already asked for.
if let Some(session) = self.session.as_mut() {
if let Some(held) = session.held.get_mut(usize::from(sub.setup_slot)) {
*held = None;
}
}
None
} else {
self.decode_into(au, &sub)?;
if let Some(session) = self.session.as_mut() {
// What this surface now holds, for a later `show_existing_frame`.
if let Some(held) = session.held.get_mut(usize::from(sub.setup_slot)) {
*held = Some(sub.facts);
}
}
if sub.show {
Some(self.present(sub.setup_slot, sub.facts)?)
} else {
None
}
};
// ⚠ The decode's `Result` is held rather than `?`-ed, so that the two slot
// releases below run on the FAILURE path too. `decode_av1` treats an error
// here as a health note and keeps the session — it does not rebuild the slot
// map — so an early return would leak a surface per failed frame and reach
// `SlotError::Full` after nine, which is a session that dies of an error it
// had already recovered from.
let shown = self.decode_and_present_av1(au, &sub, damaged);
// The surfaces this frame's own refresh displaced while its submission still
// NAMED them (fn docs). Released here for the same reason the block below
// waits: the decode op has been issued, so nothing can be assigned them
// until the next frame — and on the `damaged` path there is no op at all,
// where dropping the release would leak a surface just the same.
// until the next frame — and on the `damaged` and failed paths there is no
// op at all, where dropping the release would leak a surface just the same.
if let Some(session) = self.session.as_mut() {
for &id in &sub.release_after_decode {
if !session.slots.release(id) {
@@ -640,6 +620,49 @@ impl NativeD3d11Decoder {
}
}
}
shown
}
/// Submit one converted AV1 frame and blit it if it displays — the part of
/// [`Self::frame_av1`] that can fail, split out so its caller can run the slot
/// releases on the failure path as well as on the two clean ones.
fn decode_and_present_av1(
&mut self,
au: &[u8],
sub: &Submission,
damaged: bool,
) -> Result<Option<D3d11Frame>> {
let shown = if damaged {
// Converted (so the slot map stayed in step with the planner's store),
// deliberately not submitted (fn docs).
//
// ⚠ And the surface's `held` entry is CLEARED rather than left. The slot
// map now says this slot holds THIS picture, while the surface still
// carries whatever the previous occupant decoded; a later
// `show_existing_frame` naming it would find the old picture's facts and
// blit the old picture's pixels. `None` makes that path return
// `Ok(None)` — nothing shown — which is what the unit's concealment
// already asked for.
if let Some(session) = self.session.as_mut() {
if let Some(held) = session.held.get_mut(usize::from(sub.setup_slot)) {
*held = None;
}
}
None
} else {
self.decode_into(au, sub)?;
if let Some(session) = self.session.as_mut() {
// What this surface now holds, for a later `show_existing_frame`.
if let Some(held) = session.held.get_mut(usize::from(sub.setup_slot)) {
*held = Some(sub.facts);
}
}
if sub.show {
Some(self.present(sub.setup_slot, sub.facts)?)
} else {
None
}
};
Ok(shown)
}
@@ -2280,25 +2303,39 @@ mod parity {
/// goldens beside the plan facts that could explain it.
///
/// Not a gate — it asserts nothing and always "passes". It exists because
/// [`av1_every_delivered_frame_hashes_bit_identical_to_libavcodec`] FAILS on
/// every device tried so far, and a count of diverging frames is not a lead. This
/// is what turned that count into one, on 2026-08-07:
/// [`av1_every_delivered_frame_hashes_bit_identical_to_libavcodec`] FAILED on both
/// GPUs of `.221` the first time it was ever run, and a count of diverging frames
/// is not a lead. This is what turned that count into one, on 2026-08-07:
///
/// * **NVIDIA RTX 3500 Ada** — display frames 0..=63 bit-identical, then every one
/// of the remaining 186 diverges. The first bad frame is the one whose
/// `order_hint` first reaches **64**, and its error is 174 luma pixels in a
/// single 16x24 block (max |delta| 8, chroma untouched) which then propagates
/// of the remaining 186 diverged. The first bad frame was the one whose
/// `order_hint` first reaches **64**, and its error was 174 luma pixels in a
/// single 16x24 block (max |delta| 8, chroma untouched) which then propagated
/// through prediction. The stream keeps the key frame (`order_hint` 0) in the
/// BWDREF and ALTREF2 slots for its whole length, so 64 is where the distance to
/// it reaches the edge of what `get_relative_dist` can represent at
/// `OrderHintBits = 7`.
/// * **Intel Arc** — only display frames 0, 1, 2, 3 and 10 are bit-identical, and
/// the divergence is STRUCTURAL rather than marginal (47% of luma at the first
/// * **Intel Arc** — only display frames 0, 1, 2, 3 and 10 were bit-identical, and
/// the divergence was STRUCTURAL rather than marginal (47% of luma at the first
/// bad frame, max |delta| 242, chroma wrong too): a frame predicted from the
/// wrong picture, not a filter rounding.
///
/// Both are deterministic — three runs each, identical first-divergent frame and
/// identical hashes — so neither is a race against the decode queue.
/// Both were deterministic — three runs each, identical first-divergent frame and
/// identical hashes — so neither was a race against the decode queue.
///
/// **⚠ Both were ONE defect, and the two unlike signatures argued for two.** The
/// submission named a single surface as `CurrPicTextureIndex` and as a
/// `RefFrameMapTextureIndex` entry on 268 of the vector's 274 frames — decode into
/// the picture you predict from — because [`pf_dxvadec::plan_to_dxva_av1`] released
/// the displaced reference before assigning the decode target its slot. Intel
/// followed the aliased surface immediately; NVIDIA tolerated it until the order-hint
/// wrap put one block's prediction on the far side of it. Fixing that one thing took
/// BOTH vendors to 250/250. Two readings this map invited and that were wrong:
/// "`primary_ref_frame` or its resolution" (Intel's one correct late frame is
/// PRIMARY_REF_NONE **because** it is the intra frame, which names no reference and
/// so cannot alias) and "motion-field projection at the `get_relative_dist` sign
/// flip" (the wrap is where an already-aliased surface first mattered on NVIDIA, not
/// what was wrong). Read a signature as evidence about WHERE, not about WHAT.
///
/// Set `PF_AV1_DUMP=<tag>` to also write a few frames' raw NV12 to the temp
/// directory. That is how "how badly" was answered: at a frame where ONE vendor
+14
View File
@@ -687,6 +687,20 @@ pub fn plan_to_dxva_av1(
let color = &seq.color_config;
let mut pic_params = PicParamsAv1::zeroed();
// ⚠ UPSCALED width, where libavcodec sends the CODED one — a divergence that is
// inert on every stream that exists here and is written down rather than
// "fixed" because nothing can measure it.
//
// `dxva2_av1.c` sends `avctx->width`, and `update_context_with_frame_header`
// sets that from `frame_width_minus_1 + 1` — FrameWidth, the pre-superres coded
// width. The same goes for `frame_refs[i].width`, which libav reads off the
// reference's `AVFrame`. With superres OFF the two are equal by definition
// (7.20: `UpscaledWidth = FrameWidth` when `use_superres` is 0), which is every
// frame of both vendored vectors and every frame a punktfunk host emits — no
// encoder in this program codes superres. So the 250/250 parity result on two
// vendors says nothing either way about which is right, and changing it would
// be an unmeasured change to a rung that is finally proven. Revisit with a
// superres vector and a driver-by-driver measurement, not by reading.
pic_params.width = h.upscaled_width;
pic_params.height = h.frame_height;
pic_params.max_width = u32::from(seq.max_frame_width_minus_1) + 1;
@@ -145,6 +145,47 @@
//! while every other byte still looks right. If the macro is spelled differently in the tree,
//! any expression yielding the negotiated config's `ConfigBitstreamRaw` will do.
//!
//! **5b. AV1.** The identical `PFPP` block goes at the very END of
//! `ff_dxva2_av1_fill_picture_parameters` (`dxva2_av1.c:60`), with `h264` replaced by `av1` —
//! after the film-grain block, so every field is final. Three things differ from the two codecs
//! above and each of them changes what a capture MEANS:
//!
//! * **The AU index is a FRAME, not a temporal unit.** `ff_dxva2_common_end_frame` runs once per
//! submitted picture and an AV1 temporal unit may decode several, so `pf_au_index` walks
//! decoded frames. This crate's [`our_av1_submissions`] emits one entry per decoded frame for
//! the same reason, and the vendored vector is **274** frames in 250 units — a capture with
//! 250 `PFPP av1` lines is a capture of something else. (A `show_existing_frame` unit submits
//! nothing on either side; this vector has none.)
//! * **No `PFQM` line and no matrix buffer.** `dxva2_av1_end_frame` passes `NULL, 0` for the qm
//! pair, so the `qm_size > 0` branch of the block in step 4 logs `absent` on every frame. That
//! is the expected reading, not a missed patch site.
//! * **No `PFCFG` check.** [`preflight`]'s `ConfigBitstreamRaw` assertion is about the two short
//! slice-control formats; AV1's slice-control record is `DXVA_Tile_AV1` and has no short/long
//! pair, so a captured `PFCFG av1` line is ignored rather than compared.
//!
//! The stream is the vendored IVF at
//! `crates/pf-bitstream/vendor/cros-codecs/src/codec/av1/test_data/test-25fps.ivf.av1`:
//!
//! ```text
//! ffmpeg -hwaccel d3d11va -hwaccel_output_format d3d11 -i test-25fps.ivf.av1 -f null - 2> av1.log
//! grep -oE 'PF(PP|QM|BD|CFG) .*' av1.log > libav-av1.capture
//! ```
//!
//! ⚠ Add `-export_side_data +film_grain` to NOTHING: `apply_grain` and `pp->coding.film_grain`
//! both turn OFF when film grain is exported as side data, and this crate always applies it in
//! the decoder. The vendored vector codes no grain either way.
//!
//! **No such capture has been taken.** As of 2026-08-07 the AV1 comparison
//! (`our_av1_picture_parameters_match_libavcodecs`) has never run against libavcodec's bytes,
//! and the reason is stated rather than left as an absent result: `.221`, the only box in this
//! fleet with a D3D11VA GPU to spare, has no MSYS2, no gcc and no make, so producing a patched
//! FFmpeg there is a toolchain bring-up rather than a build. What DID localise the AV1 defect
//! of 2026-08-07 was `video_d3d11_native`'s frame-hash parity harness plus a CPU invariant
//! (`no_av1_submission_names_its_decode_surface_in_the_reference_store`, below) — so this file's
//! AV1 half is currently the no-capture half only, and every claim it makes about libavcodec's
//! AV1 side is READ out of `dxva2_av1.c` (n8.1) rather than measured. Tier one, in the
//! provenance section's terms, for all of it.
//!
//! **6. Run it.** `--enable-d3d11va` is on by default on Windows. Decode the SAME elementary
//! streams this test plans — the vendored vectors, in the repository at
//! `crates/pf-bitstream/vendor/cros-codecs/src/codec/{h264,h265}/test_data/test-25fps.{h264,h265}`:
@@ -167,6 +208,7 @@
//!
//! ```text
//! PF_LIBAV_CAPTURE_H264=libav-h264.capture PF_LIBAV_CAPTURE_HEVC=libav-hevc.capture \
//! PF_LIBAV_CAPTURE_AV1=libav-av1.capture \
//! cargo test -p pf-dxvadec --test libav_picparams_parity -- --ignored --nocapture
//! ```
//!
@@ -352,13 +394,19 @@ use pf_dxvadec::dxva::QmatrixHevc;
use pf_dxvadec::dxva::SliceH264Short;
use pf_dxvadec::dxva::SliceHevcShort;
use pf_dxvadec::dxva::UNUSED_ENTRY;
use pf_dxvadec::dxva_av1::PicEntryAv1;
use pf_dxvadec::dxva_av1::UNUSED_INDEX;
use pf_dxvadec::AuPlan;
use pf_dxvadec::Av1Planner;
use pf_dxvadec::BufferDescriptor;
use pf_dxvadec::Codec;
use pf_dxvadec::H264Planner;
use pf_dxvadec::H265Planner;
use pf_dxvadec::PicParamsAv1;
use pf_dxvadec::SliceRecord;
use pf_dxvadec::SlotMap;
use pf_dxvadec::TileAv1;
use pf_dxvadec::NUM_REF_SLOTS;
const TEST_25FPS_H264: &[u8] = include_bytes!(
"../../pf-bitstream/vendor/cros-codecs/src/codec/h264/test_data/test-25fps.h264"
@@ -366,6 +414,11 @@ const TEST_25FPS_H264: &[u8] = include_bytes!(
const TEST_25FPS_H265: &[u8] = include_bytes!(
"../../pf-bitstream/vendor/cros-codecs/src/codec/h265/test_data/test-25fps.h265"
);
/// The AV1 vector is an IVF container, and its unit of comparison is the TEMPORAL UNIT rather
/// than the access unit: one IVF packet may decode several frames, of which at most one shows.
const TEST_25FPS_AV1: &[u8] = include_bytes!(
"../../pf-bitstream/vendor/cros-codecs/src/codec/av1/test_data/test-25fps.ivf.av1"
);
/// Both vendored vectors carry exactly this many access units — pf-bitstream's own golden, and
/// the number of `PFPP` lines a valid capture holds.
@@ -453,8 +506,12 @@ struct OurSubmission {
qmatrix: Option<Vec<u8>>,
/// The descriptor set, in submission order.
descriptors: Vec<BufferDescriptor>,
/// The packer's slice records, for the internal-consistency checks.
/// The packer's slice records, for the internal-consistency checks. Empty on AV1, whose
/// slice-control buffer holds [`Self::tiles`] instead.
records: Vec<SliceRecord>,
/// AV1's slice-control records — one `DXVA_Tile_AV1` per TILE, not per tile group. Empty
/// on H.264 and H.265.
tiles: Vec<TileAv1>,
/// Bytes the packer wrote BEFORE the tail padding.
unpadded: u32,
/// `mb_width * mb_height` (H.264) or 0 (HEVC) — the value the descriptors must carry.
@@ -491,6 +548,7 @@ fn our_h264_submissions() -> Vec<OurSubmission> {
qmatrix: Some(pf_dxvadec::as_bytes(&dxva.qmatrix).to_vec()),
descriptors: pf_dxvadec::descriptors_h264(&dxva, &packed),
records: packed.records,
tiles: Vec::new(),
unpadded,
mb_count: dxva.mb_count,
});
@@ -527,6 +585,7 @@ fn our_hevc_submissions() -> Vec<OurSubmission> {
.map(|qm| pf_dxvadec::as_bytes(qm).to_vec()),
descriptors: pf_dxvadec::descriptors_h265(&dxva, &packed),
records: packed.records,
tiles: Vec::new(),
unpadded,
mb_count: 0,
});
@@ -535,6 +594,85 @@ fn our_hevc_submissions() -> Vec<OurSubmission> {
out
}
/// Every AV1 FRAME the vendored vector decodes — 274, of which 250 are displayed.
///
/// The unit of comparison is the frame and not the temporal unit, because that is what
/// libavcodec's hwaccel counts: `ff_dxva2_common_end_frame` runs once per submitted PICTURE, so
/// a capture's AU index walks decoded frames. A `show_existing_frame` unit submits nothing and
/// appears on neither side; this vector has none.
const VENDORED_AV1_FRAMES: usize = 274;
/// Plan, convert and pack the whole vendored AV1 vector, one entry per decoded FRAME.
///
/// ⚠ This is the only one of the three that has to speak the conversion's DEFERRED RELEASE
/// contract ([`pf_dxvadec::DecodePlanDxvaAv1::release_after_decode`]). A loop that converts
/// without it holds a surface on 268 of these 274 frames and runs the nine-slot ledger dry
/// inside ten — and, worse for a harness, it would compare a submission built by a caller that
/// is not the rung.
fn our_av1_submissions() -> Vec<OurSubmission> {
let mut planner = Av1Planner::new();
let mut slots = SlotMap::new(NUM_REF_SLOTS);
let mut mapping = vec![0u8; MAPPING_BYTES];
let mut out = Vec::new();
for (i, unit) in split_ivf(TEST_25FPS_AV1).into_iter().enumerate() {
let plans = planner
.plan_au(unit)
.unwrap_or_else(|e| panic!("unit {i} of the vendored AV1 vector must plan: {e}"));
for plan in &plans {
if plan.dpb.stored.is_none() {
continue; // `show_existing_frame`: no submission at all
}
let dxva = pf_dxvadec::plan_to_dxva_av1(unit, plan, &mut slots)
.unwrap_or_else(|e| panic!("unit {i} must convert: {e}"));
let packed = pf_dxvadec::pack_av1(unit, &dxva.bitstream, &dxva.tiles, &mut mapping)
.unwrap_or_else(|e| panic!("unit {i} must pack: {e}"));
let unpadded = pf_dxvadec::packed_size_av1(&dxva.bitstream) as u32;
out.push(OurSubmission {
pic_params: pf_dxvadec::as_bytes(&dxva.pic_params).to_vec(),
// AV1 transmits no quantization matrix at all: its matrices are SELECTED by
// index out of tables the decoder already has, and `dxva2_av1_end_frame`
// passes `NULL, 0` for the pair. `None` here is a fact about the codec, not a
// condition on the stream the way HEVC's is.
qmatrix: None,
descriptors: pf_dxvadec::descriptors_av1(&packed),
// AV1's slice-control records are `DXVA_Tile_AV1`, a different struct with a
// different size; the shared `SliceRecord` checks do not apply to them, and
// the tile records get their own test rather than a coerced one.
records: Vec::new(),
tiles: packed.tiles.clone(),
unpadded,
mb_count: 0,
});
for &id in &dxva.release_after_decode {
assert!(
slots.release(id),
"unit {i}: a deferred release named a picture holding no surface"
);
}
}
}
assert_eq!(out.len(), VENDORED_AV1_FRAMES);
out
}
/// The IVF frame walk — the same one `video_d3d11_native`'s parity module and every AV1 test in
/// this program use: a 32-byte file header, then a 12-byte header per packet carrying its size.
fn split_ivf(stream: &[u8]) -> Vec<&[u8]> {
let mut out = Vec::new();
let mut at = 32usize;
while at + 12 <= stream.len() {
let size = u32::from_le_bytes([stream[at], stream[at + 1], stream[at + 2], stream[at + 3]])
as usize;
at += 12;
if at + size > stream.len() {
break;
}
out.push(&stream[at..at + size]);
at += size;
}
out
}
// ---------------------------------------------------------------------------
// Offset → field name
// ---------------------------------------------------------------------------
@@ -543,9 +681,12 @@ fn our_hevc_submissions() -> Vec<OurSubmission> {
/// order, built from the field IDENTIFIERS so a name and the offset it reports cannot drift
/// apart — the whole point of the table is to turn a differing byte into a field name, and a
/// table with a copy-pasted mismatch would name the wrong one.
/// Nested paths (`tiles.cols`) are accepted as well as plain identifiers, and are named by
/// the whole path — AV1's picture parameters are eight nested blocks, and a table that could
/// only reach the outer members would report "segmentation differs" for a 140-byte struct.
macro_rules! field_table {
($ty:ty, $($field:ident),+ $(,)?) => {
&[$((stringify!($field), offset_of!($ty, $field))),+]
($ty:ty, $($($field:ident).+),+ $(,)?) => {
&[$((stringify!($($field).+), offset_of!($ty, $($field).+))),+]
};
}
@@ -670,6 +811,87 @@ const HEVC_QMATRIX_FIELDS: &[(&str, usize)] = field_table!(
ucScalingListDCCoefSizeID3,
);
/// Every field of `DXVA_PicParams_AV1`, same construction — but reaching INTO the eight nested
/// blocks, because they are where AV1 keeps almost the whole frame header. `tiles` alone is 260
/// bytes and `segmentation` 140; a table stopping at the outer members would turn every finding
/// in them into one useless name.
///
/// Three members stay whole on purpose. `frame_refs` is seven 36-byte entries which — unlike
/// the other two codecs' reference arrays — carry NO surface index and so compare byte for
/// byte: `Index` is `ref_frame_idx[name]`, an AV1 SLOT both sides read out of the same frame
/// header. `ref_frame_map_texture_index` is the surface array, which cannot be compared by
/// value at all ([`av1_reference_store`]). And `film_grain` is 158 bytes neither vendored
/// vector codes.
const AV1_FIELDS: &[(&str, usize)] = field_table!(
PicParamsAv1,
width,
height,
max_width,
max_height,
curr_pic_texture_index,
superres_denom,
bitdepth,
seq_profile,
tiles.cols,
tiles.rows,
tiles.context_update_id,
tiles.widths,
tiles.heights,
coding,
format,
primary_ref_frame,
order_hint,
order_hint_bits,
frame_refs,
ref_frame_map_texture_index,
loop_filter.filter_level,
loop_filter.filter_level_u,
loop_filter.filter_level_v,
loop_filter.sharpness_level,
loop_filter.control_flags,
loop_filter.ref_deltas,
loop_filter.mode_deltas,
loop_filter.delta_lf_res,
loop_filter.frame_restoration_type,
loop_filter.log2_restoration_unit_size,
loop_filter.reserved16,
quantization.control_flags,
quantization.base_qindex,
quantization.y_dc_delta_q,
quantization.u_dc_delta_q,
quantization.v_dc_delta_q,
quantization.u_ac_delta_q,
quantization.v_ac_delta_q,
quantization.qm_y,
quantization.qm_u,
quantization.qm_v,
quantization.reserved16,
cdef.control_flags,
cdef.y_strengths,
cdef.uv_strengths,
interp_filter,
segmentation.control_flags,
segmentation.reserved24,
segmentation.feature_mask,
segmentation.feature_data,
film_grain,
reserved32,
status_report_feedback_number,
);
/// `DXVA_Tile_AV1` — AV1's slice-control record, and the one this crate had to derive rather
/// than measure (`dxva.h` declares it; the SIZE is what the descriptor states).
const AV1_TILE_FIELDS: &[(&str, usize)] = field_table!(
TileAv1,
data_offset,
data_size,
row,
column,
reserved16,
anchor_frame,
reserved8,
);
/// Turn a field table into `(name, byte range)`, the last field running to `total`.
fn field_ranges(
fields: &[(&'static str, usize)],
@@ -1038,6 +1260,37 @@ fn hevc_ref_entries(pp: &[u8]) -> Vec<(u8, RefEntry)> {
.collect()
}
/// One side's AV1 reference store, read out of the SUBMITTED BYTES: `(CurrPicTextureIndex,
/// RefFrameMapTextureIndex[8], frame_refs[name].Index for the seven names)`.
///
/// AV1's reference numbering is two arrays that mean different things at once and the split is
/// exactly what a comparison has to respect. `frame_refs[i].Index` is an AV1 reference SLOT —
/// `ref_frame_idx[i]`, which both sides read out of the same frame header — so it is a VALUE
/// that must match libavcodec's exactly, and it is compared as part of the `frame_refs` field.
/// `RefFrameMapTextureIndex[slot]` and `CurrPicTextureIndex` are SURFACES, which come from each
/// side's own pool and are only ever a bijection.
///
/// So the store is compared as a SHAPE: which slots are occupied, and whether the decode target
/// collides with any of them.
fn av1_reference_store(pp: &[u8]) -> (u8, [u8; 8], [u8; 7]) {
let curr = pp[offset_of!(PicParamsAv1, curr_pic_texture_index)];
let mut store = [UNUSED_INDEX; 8];
let base = offset_of!(PicParamsAv1, ref_frame_map_texture_index);
store.copy_from_slice(&pp[base..base + 8]);
let mut names = [UNUSED_INDEX; 7];
// The stride and the member offset come from the TYPE, never from the two numbers
// `dxva_av1.rs` measured (36 and 33). Those are pinned there as compile-time assertions
// against the Windows SDK's own header, and re-typing them here would be a second copy that
// can drift from the first — which for a reader of this array is the difference between a
// reference slot and a warp coefficient.
for (name, slot) in names.iter_mut().enumerate() {
*slot = pp[offset_of!(PicParamsAv1, frame_refs)
+ name * size_of::<PicEntryAv1>()
+ offset_of!(PicEntryAv1, index)];
}
(curr, store, names)
}
/// The surface mapping between the two sides, tracked per PICTURE.
///
/// A global index-to-index bijection over a whole stream is the wrong model: both sides reuse a
@@ -1287,10 +1540,17 @@ fn preflight(capture: &Capture, ours: usize, codec: &str, reserved16: Option<usi
}
// A `PFCFG` line is optional, but a wrong one voids the slice-control comparison: it means
// the driver read the other slice-control struct entirely.
let want = pf_dxvadec::short_slice_config(match codec {
"h264" => Codec::H264,
_ => Codec::H265,
});
//
// ⚠ AV1 is exempt, and not because the check is inconvenient: `ConfigBitstreamRaw`'s short
// format is a property of the two SHORT SLICE-CONTROL structs, and AV1's slice-control
// record is `DXVA_Tile_AV1`, which has no short/long pair for a config to select between.
// Comparing an AV1 capture's number against HEVC's 1 — which a `_ =>` arm would do — is a
// check of nothing that fails on anything.
let want = match codec {
"h264" => pf_dxvadec::short_slice_config(Codec::H264),
"hevc" => pf_dxvadec::short_slice_config(Codec::H265),
_ => return,
};
for (au, &raw) in &capture.config_bitstream_raw {
assert_eq!(
raw, want,
@@ -1568,6 +1828,108 @@ fn hevc_rps_pictures(pp: &[u8], array: usize, entries: &[(u8, RefEntry)]) -> Vec
.collect()
}
/// The whole AV1 picture-parameter comparison.
///
/// Structurally simpler than the other two and the reason is worth stating: AV1 puts NO surface
/// index in its reference entries. `frame_refs[i].Index` is `ref_frame_idx[i]`, an AV1 SLOT both
/// sides read out of the same frame header, so the seven 36-byte entries — sizes, warp
/// parameters, warp type and slot alike — compare byte for byte with no re-indexing, no set
/// comparison and no allowance. Only two members carry surfaces, and they are handled as the
/// SHAPE of the store rather than by value ([`av1_reference_store`]).
///
/// There is no POC base to derive either: AV1's `order_hint` is a coded field, not a decoder's
/// running count, so libavcodec has nothing to seed it with.
///
/// ⚠ **`width`/`height` is a divergence waiting to be measured, and this comparison will
/// report it rather than absorb it.** libavcodec sends `avctx->width`, which
/// `update_context_with_frame_header` sets from `frame_width_minus_1 + 1` — FrameWidth, the
/// PRE-superres coded width — and the same for `frame_refs[i].width` off the reference's
/// `AVFrame`; this crate sends `UpscaledWidth`. With superres off the two are equal by
/// definition (7.20), which is every frame of the vendored vector and every frame a punktfunk
/// host emits, so a capture made from this vector cannot tell them apart. Deliberately given no
/// allowance: if a superres capture ever reaches this harness, the difference must be a finding
/// somebody reads, not a line somebody already excused. See `pic_av1.rs`'s note at `pp.width`.
fn compare_av1_picparams(ours: &[OurSubmission], capture: &Capture) -> Findings {
let ranges = field_ranges(AV1_FIELDS, size_of::<PicParamsAv1>());
// The two surface arrays, and nothing else: every other byte of this struct is a fact about
// the bitstream that both sides derive from the same frame header.
let structural = ["curr_pic_texture_index", "ref_frame_map_texture_index"];
let mut findings = Findings::default();
for (au, sub) in ours.iter().enumerate() {
let Some(theirs) = capture.pic_params.get(&au) else {
findings.note(
"<no capture>",
au,
"the capture holds no PFPP line for this frame",
);
continue;
};
if theirs.len() != sub.pic_params.len() {
findings.note(
"<struct size>",
au,
format!(
"the capture's picture parameters are {} bytes and ours are {}",
theirs.len(),
sub.pic_params.len()
),
);
continue;
}
compare_scalars(
au,
&sub.pic_params,
theirs,
&ranges,
&structural,
no_allowance,
&mut findings,
);
// The store, as a shape. Which SLOTS hold a picture is a fact about the bitstream and
// must agree; which SURFACE each holds is each side's own pool and never can.
let (our_curr, our_store, _) = av1_reference_store(&sub.pic_params);
let (their_curr, their_store, _) = av1_reference_store(theirs);
for slot in 0..8 {
let ours_occupied = our_store[slot] != UNUSED_INDEX;
let theirs_occupied = their_store[slot] != UNUSED_INDEX;
if ours_occupied != theirs_occupied {
findings.note(
format!("ref_frame_map_texture_index[{slot}][occupied]"),
au,
format!("ours {ours_occupied}, libav {theirs_occupied}"),
);
}
}
// The decode target must hold no store entry's surface. libavcodec cannot produce a
// collision — it fills the store from `h->ref[i]`, which the reference update has not
// run on yet, and takes `CurrPicTextureIndex` from `h->cur_frame.f` — so a collision on
// our side is a defect however the surfaces are numbered. This is the check that names
// the 2026-08-07 defect.
//
// ⚠ Note what is deliberately NOT checked: that two slots hold different surfaces. One
// picture in several reference slots is ordinary AV1 and this very vector does it —
// the key frame sits in BWDREF and ALTREF2 for the stream's whole length — so a
// "duplicate surface" check would fire on 273 of 274 frames of a correct conversion.
for (label, curr, store) in [
("ours", our_curr, our_store),
("libav", their_curr, their_store),
] {
if store.iter().any(|surface| *surface == curr) {
findings.note(
"curr_pic_texture_index[aliases the store]",
au,
format!(
"{label}: surface {curr} is both the decode target and a reference \
store entry the frame decodes into a picture it predicts from"
),
);
}
}
}
findings
}
/// The whole HEVC picture-parameter comparison.
fn compare_hevc_picparams(ours: &[OurSubmission], capture: &Capture) -> Findings {
let ranges = field_ranges(HEVC_FIELDS, size_of::<PicParamsHevc>());
@@ -1970,6 +2332,13 @@ fn every_hand_declared_dxva_struct_is_tiled_exactly_by_its_fields() {
HEVC_SLICE_FIELDS,
size_of::<SliceHevcShort>(),
),
// AV1's two. `PicParamsAv1` is the one struct in this crate whose offsets were
// MEASURED rather than mirrored — `layout-probe-av1.c` compiled with MSVC against the
// Windows SDK's own `dxva.h` — and `dxva_av1.rs` pins every one at compile time. What
// this adds is the other half: that the TABLE above reaches all 912 bytes, so a
// capture comparison can name every one of them.
("PicParamsAv1", AV1_FIELDS, size_of::<PicParamsAv1>()),
("TileAv1", AV1_TILE_FIELDS, size_of::<TileAv1>()),
] {
assert_eq!(fields[0].1, 0, "{what}: the first field must start at 0");
let ranges = field_ranges(fields, total);
@@ -2006,6 +2375,196 @@ fn every_h264_au_submits_four_buffers_in_libavcodecs_order() {
}
}
/// **No AV1 submission names its decode surface anywhere in the reference store**, and every
/// reference NAME resolves through a slot that holds one.
///
/// This is the defect the Windows parity harness caught on 2026-08-07 and the one nothing on
/// the CPU could see, stated over the SUBMITTED BYTES — which is where a libavcodec capture
/// would see it too, and the reason it belongs in this file as well as in `pic_av1`'s own
/// tests. `plan_to_dxva_av1` released the picture this frame's own `refresh_frame_flags`
/// displaces before assigning the decode target a slot, and `SlotMap::assign` hands back the
/// slot just vacated — so `CurrPicTextureIndex` and one `RefFrameMapTextureIndex` entry were
/// the same surface on 268 of these 274 frames: decode into the picture you predict from.
/// Intel Arc followed the aliased surface and got 245 of 250 delivered frames wrong; NVIDIA
/// tolerated it for 63 frames and then lost one 16x24 luma block at the `order_hint` wrap.
///
/// libavcodec cannot produce this shape and that is the whole argument for calling it a defect
/// rather than a convention: `ff_dxva2_av1_fill_picture_parameters` fills
/// `RefFrameMapTextureIndex` from `h->ref[i]`, the pre-refresh store, and takes
/// `CurrPicTextureIndex` from `h->cur_frame.f`, a frame the reference-frame update has not run
/// on yet. The two cannot be one surface.
///
/// The counts are asserted, not printed. At zero references this test would pass against a
/// conversion that named nothing at all.
#[test]
fn no_av1_submission_names_its_decode_surface_in_the_reference_store() {
let subs = our_av1_submissions();
let (mut with_store, mut named_refs) = (0usize, 0usize);
for (frame, sub) in subs.iter().enumerate() {
let (curr, store, names) = av1_reference_store(&sub.pic_params);
assert!(
store.iter().all(|surface| *surface != curr),
"frame {frame}: surface {curr} is both CurrPicTextureIndex and a \
RefFrameMapTextureIndex entry"
);
if store.iter().any(|s| *s != UNUSED_INDEX) {
with_store += 1;
}
for (name, slot) in names.iter().enumerate() {
if *slot == UNUSED_INDEX {
continue;
}
named_refs += 1;
assert!(
usize::from(*slot) < store.len(),
"frame {frame}, reference name {name}: slot {slot} is outside the eight-entry \
store `Index` is an AV1 reference SLOT, not a surface"
);
assert_ne!(
store[usize::from(*slot)],
UNUSED_INDEX,
"frame {frame}, reference name {name}: slot {slot} holds no surface, so the \
driver would follow `Index` into an empty entry"
);
}
}
assert_eq!(
with_store, 273,
"every frame but the opening key frame carries a populated reference store"
);
assert!(
named_refs > 0,
"no frame named a reference, so every check above was skipped"
);
}
/// AV1 submits THREE buffers and never a quantization matrix, on every frame of the vector.
///
/// The codec asymmetry the H.264 and HEVC tests above are about, taken to its third case.
/// H.264 submits the matrix unconditionally, HEVC only under `scaling_list_enabled_flag`, and
/// AV1 has no matrix BUFFER at all: its quantiser matrices are SELECTED by index
/// (`qm_y`/`qm_u`/`qm_v`) out of tables the decoder already holds, and `dxva2_av1_end_frame`
/// passes `NULL, 0` for the qm pair so the generic layer submits nothing. A fourth descriptor
/// here would be a buffer the driver has no `DXVA_Qmatrix_AV1` to read it as.
#[test]
fn every_av1_frame_submits_three_buffers_and_never_a_quantization_matrix() {
for (frame, sub) in our_av1_submissions().iter().enumerate() {
assert_eq!(
sub.descriptors
.iter()
.map(|d| d.buffer_type)
.collect::<Vec<_>>(),
vec![
BUFFER_PICTURE_PARAMETERS,
BUFFER_BITSTREAM,
BUFFER_SLICE_CONTROL,
],
"frame {frame}"
);
assert!(sub.qmatrix.is_none(), "frame {frame}");
}
}
/// No AV1 descriptor carries a macroblock count. AV1 has no macroblocks and `dxva2_av1.c`
/// never touches the field — the same statement `no_hevc_descriptor_ever_carries_a_macroblock_count`
/// makes, and the same defect class review 13 found on the H.264 side in the other direction.
#[test]
fn no_av1_descriptor_ever_carries_a_macroblock_count() {
for (frame, sub) in our_av1_submissions().iter().enumerate() {
assert_eq!(sub.mb_count, 0, "frame {frame}");
for desc in &sub.descriptors {
assert_eq!(
desc.num_mbs_in_buffer,
0,
"frame {frame}, {}",
buffer_name(desc.buffer_type)
);
}
}
}
/// AV1's slice-control buffer is `16 * tile count`, its bitstream descriptor is the packer's
/// PADDED size, and the tile records tile the unpadded window exactly — in order, without gaps
/// and without overlaps.
///
/// The last part is what distinguishes AV1 from the other two codecs here and is the reason
/// this cannot reuse `the_bitstream_descriptor_is_the_packers_padded_size_and_the_slice_records_tile_it_exactly`:
/// a `DXVA_Tile_AV1` addresses a TILE PAYLOAD, which is the bytes after that tile's
/// `tile_size_minus_1` field — so consecutive records are separated by those size fields and do
/// NOT abut, unlike H.264/HEVC slice records which tile their buffer with no gaps. What must
/// hold is weaker and still exact: strictly increasing, non-overlapping, inside the unpadded
/// window, and never starting at a tile-group OBU's first byte (which would hand the driver an
/// OBU header as entropy-coded tile data).
///
/// ⚠ The padding is charged to NO record. H.264 and HEVC add the tail padding to their last
/// slice record's `SliceBytesInBuffer`; `pack_av1` does not, because a tile's size is the
/// tile's, and the descriptor is the only place AV1's padding is accounted at all.
#[test]
fn the_av1_bitstream_descriptor_is_padded_and_the_tile_records_tile_it_without_overlapping() {
let mut frames_with_padding = 0usize;
for (frame, sub) in our_av1_submissions().iter().enumerate() {
let bitstream = sub
.descriptors
.iter()
.find(|d| d.buffer_type == BUFFER_BITSTREAM)
.unwrap_or_else(|| panic!("frame {frame} submits no bitstream buffer"));
assert_eq!(
bitstream.data_size % 128,
0,
"frame {frame}: the bitstream descriptor states the PADDED size"
);
// ⚠ `1..=128`, not `0..128`. `pack_av1` writes libavcodec's expression verbatim —
// `BITSTREAM_ALIGN - (cursor % BITSTREAM_ALIGN)` — so data that is ALREADY on the
// granule gets a whole 128-byte block rather than none (`pack_av1`'s
// `data_already_on_the_granule_still_gets_a_full_padding_block`). This vector never
// lands on the granule, which is exactly why the bound has to come from the rule and
// not from the measurement.
let padding = bitstream.data_size - sub.unpadded;
assert!(
(1..=128).contains(&padding),
"frame {frame}: {padding} bytes of padding"
);
frames_with_padding += 1;
let slice_control = sub
.descriptors
.iter()
.find(|d| d.buffer_type == BUFFER_SLICE_CONTROL)
.unwrap_or_else(|| panic!("frame {frame} submits no slice-control buffer"));
assert_eq!(
slice_control.data_size as usize,
size_of::<TileAv1>() * sub.tiles.len(),
"frame {frame}: sixteen bytes per TILE"
);
assert!(!sub.tiles.is_empty(), "frame {frame}: a frame has tiles");
let mut previous_end = 0u32;
for (i, tile) in sub.tiles.iter().enumerate() {
// `#[repr(packed)]` — copy the fields out before using them.
let (offset, size) = (tile.data_offset, tile.data_size);
assert!(size > 0, "frame {frame}, tile {i}: an empty tile payload");
assert!(
offset >= previous_end,
"frame {frame}, tile {i}: starts at {offset}, inside the previous tile which \
ends at {previous_end}"
);
assert!(
offset + size <= sub.unpadded,
"frame {frame}, tile {i}: runs past the bytes the packer wrote"
);
previous_end = offset + size;
assert_eq!(
tile.anchor_frame, UNUSED_INDEX,
"frame {frame}, tile {i}: large-scale-tile anchors are libavcodec's 0xFF"
);
}
}
assert_eq!(
frames_with_padding, VENDORED_AV1_FRAMES,
"every frame is padded — the rule is unconditional"
);
}
#[test]
fn every_h264_bitstream_and_slice_control_descriptor_carries_mb_width_times_mb_height() {
// Review 13's defect, over the whole vector: `NumMBsInBuffer` was 0 where libavcodec's
@@ -2281,6 +2840,7 @@ fn the_picture_parameter_buffer_is_the_whole_hand_declared_struct_for_both_codec
for (codec, subs, size) in [
("h264", our_h264_submissions(), size_of::<PicParamsH264>()),
("hevc", our_hevc_submissions(), size_of::<PicParamsHevc>()),
("av1", our_av1_submissions(), size_of::<PicParamsAv1>()),
] {
for (au, sub) in subs.iter().enumerate() {
assert_eq!(sub.pic_params.len(), size, "{codec} AU {au}");
@@ -2347,6 +2907,7 @@ fn hevc_case(enabled: bool, sps_coded: Option<u8>, pps_coded: Option<u8>) -> Our
.map(|qm| pf_dxvadec::as_bytes(qm).to_vec()),
descriptors: pf_dxvadec::descriptors_h265(&dxva, &packed),
records: packed.records,
tiles: Vec::new(),
unpadded,
mb_count: 0,
}
@@ -2508,6 +3069,35 @@ fn the_dump_and_the_parser_agree_and_the_comparison_finds_nothing_against_oursel
// The HEVC matrices are `absent` on this vector, and the parser must carry that fact rather
// than losing it — the whole of review 13's defect is the difference between the two.
assert!(capture.qmatrix.values().all(Option::is_none));
// AV1, on all 274 FRAMES. Nothing here has ever been run against libavcodec's own bytes
// (module docs say why), so this self-comparison is the only thing standing between
// `compare_av1_picparams` and a first capture: it proves the 912-byte field table reaches
// every byte, that `av1_reference_store` reads `CurrPicTextureIndex` and the eight-entry
// store from the offsets it thinks it does — a wrong one would report a false alias on a
// correct submission — and that the comparison invents nothing on identical input.
let ours = our_av1_submissions();
let capture = parse_capture(&dump("av1", &ours), "av1");
preflight(&capture, ours.len(), "av1", None);
assert_eq!(capture.pic_params.len(), VENDORED_AV1_FRAMES);
for findings in [
compare_av1_picparams(&ours, &capture),
compare_descriptors(&ours, &capture),
] {
assert!(
findings.is_empty(),
"comparing our own AV1 bytes against themselves must find nothing, got {:?}",
findings.fields()
);
assert!(
findings.documented_fields().is_empty(),
"identical bytes documented a divergence: {:?}",
findings.documented_fields()
);
}
// AV1 reports `absent` on every frame — it has no matrix BUFFER at all, unlike HEVC where
// the same spelling is a per-sequence decision.
assert!(capture.qmatrix.values().all(Option::is_none));
}
/// A submission holding only what [`compare_descriptors`] reads, for the bitstream-size
@@ -2527,6 +3117,7 @@ fn descriptor_only_submission(unpadded: u32, padded: u32, slices: usize) -> OurS
OurSubmission {
pic_params: vec![0u8; size_of::<PicParamsH264>()],
qmatrix: None,
tiles: Vec::new(),
descriptors: vec![
BufferDescriptor {
buffer_type: BUFFER_BITSTREAM,
@@ -2698,6 +3289,7 @@ fn the_hevc_tiles_flag_allowance_is_exactly_bit_ten_with_tiles_disabled_and_noth
qmatrix: sub.qmatrix.clone(),
descriptors: sub.descriptors.clone(),
records: sub.records.clone(),
tiles: sub.tiles.clone(),
unpadded: sub.unpadded,
mb_count: sub.mb_count,
}
@@ -3080,6 +3672,18 @@ fn renumber_h264_surfaces(pp: &[u8], f: impl Fn(u8) -> u8) -> Vec<u8> {
/// Emit this crate's whole submission — both codecs — in the capture's own format, so the two
/// files can be diffed by any tool without a capture at all.
#[test]
#[ignore = "needs a libavcodec capture: PF_LIBAV_CAPTURE_AV1=<file> (see the module docs)"]
fn our_av1_picture_parameters_match_libavcodecs() {
let capture = capture_from_env("PF_LIBAV_CAPTURE_AV1", "av1")
.expect("PF_LIBAV_CAPTURE_AV1=<file> names a capture (see the module docs)");
let ours = our_av1_submissions();
// No `Reserved16Bits` preflight: both libavcodec workarounds are H.264-only
// (`dxva2_h264.c`), and `DXVA_PicParams_AV1` has no such field to test.
preflight(&capture, ours.len(), "av1", None);
compare_av1_picparams(&ours, &capture).verdict("AV1 picture parameters", ours.len());
}
#[test]
#[ignore = "writes a dump: PF_DXVA_DUMP=<path>"]
fn dump_our_submission_in_the_captures_own_format() {