Commit Graph
8 Commits
Author SHA1 Message Date
enricobuehler a9e7c033c3 test(client/vaapi): the last rung of the ladder, finally checked in pixels — 7 legs, all bit-identical
Every other decode rung earns `verified` with frame-hash parity against
libavcodec. VAAPI could not: it hands out a DRM-PRIME dmabuf whose memory the
driver tiles, so nothing could read its decoded pixels back, and all four of its
legs sat at "never frame-hash parity-checked".

That was never bookkeeping. The D3D11VA AV1 rung decoded 250 frames, streamed
4K60 through a clean five-minute soak, and produced WRONG PIXELS for 186 of 250
frames on NVIDIA and 245 of 250 on Intel. It looked perfect on glass; only the
goldens caught it, and the same defect turned out to be in H.264 on two other
rungs. VAAPI was the one rung where that class of bug could still be sitting
with nothing able to see it.

It is not. Measured on .25 (Radeon 780M, RDNA3, radeonsi, Mesa 26.0.3, VA-API
1.23) on 2026-08-08, against the SAME golden files the Vulkan and D3D11VA rungs
are held to, read across the crate boundary rather than copied:

  H.264 vendored vector            250/250 bit-identical  (7 from the flush)
  H.264 our host, low-delay 640x480 120/120 bit-identical  (3 from the flush)
  H.265 vendored vector            250/250 bit-identical  (2 from the flush)
  H.265 our host, low-delay 640x480 120/120 bit-identical  (0 from the flush)
  HEVC Main 10, P010                 50/50 bit-identical  (2 from the flush)
  AV1 vendored vector              250/250 delivered of 274 decoded, and
                                   display frame 0 byte-identical to
                                   libavcodec's own PIXELS
  AV1 our host, 4K two-tile          60/60 bit-identical

⚠ ONE vendor. AMD/radeonsi only; no Intel iHD box has run these legs.

The readback that made it possible:

* `pf-vaadec`'s `va` module gains `VAImage` and `VAImageFormat`, hand-declared
  with every size and offset measured off libva 2.23.0's real headers by
  `layout-probe.c` and pinned as compile-time assertions — the same discipline
  the decode buffers already keep. The trap: `VAImage::width`/`height` are
  16-bit, so `data_size` sits at 60 and not at the 64 counting 32-bit fields
  gives, and every field after them is two bytes earlier than it looks.
* `pack_two_plane` is the pure geometry — the crop to the picture, the padding
  columns dropped per row, and the chroma plane taken from the driver's OWN
  `offsets[1]` rather than from `pitch * display_height`, which is the 1088-row
  smear this program has already paid for once. It needs no device, so ten CPU
  tests cover it on macOS and in the container.
* `video_vaapi_native::parity` drives the seven streams above through the
  production entry point and hashes what the rung DELIVERS, in delivery order,
  tail included — so the delivery path is under test as well as the decode, and
  a frame's surface comes from its own release token rather than from an
  inference about which pool entry holds which picture.

THE READBACK CANNOT REACH THE PRODUCTION PATH, and that is structural rather
than a promise. `vaDeriveImage`, `vaCreateImage`, `vaGetImage`, `vaMapBuffer`
and the rest are resolved by a `#[cfg(test)]` type that dlopens libva itself;
the production `Libva` gains no field; `sha2` is a dev dependency. A CPU test
scans this file's own source and fails if any of those symbols is dlsym'd
outside the harness, so a refactor cannot quietly undo it.

Derive is not guaranteed, so both routes are implemented and neither is
optional: `vaDeriveImage` first, `vaCreateImage` + `vaGetImage` as the fallback
(which also detiles), and if neither yields the pool's own fourcc the leg FAILS
naming what the driver gave it. There is no skip path — a parity test that
passes because it could not read anything is the failure mode this program has
been bitten by three times. Both answer on radeonsi, the first frame of every
leg is read through BOTH and they must agree, and `PF_VAAPI_READBACK=getimage`
reproduces the H.264 leg's 250/250 through the copying route alone, so the
fallback is exercised rather than merely written.

And it can fail — proven, not asserted. Planting the real geometry defect this
driver's layout makes visible (rows read contiguously, ignoring the 512-byte
pitch behind a 320-wide picture) fails at display frame 0 with the full
localisation: 68312 luma and 14998 chroma samples differing, max |delta| 255,
luma bounding box (0,1)..(319,239) — and with the goldens forced through one
route, 250/250 diverging with "suspect the readback geometry". `compare` and
`localise` also have CPU counterfactuals, and a hardware leg proves the readback
reads real and DISTINCT pixels and localises a one-byte flip to the exact pixel.

⚠ One thing the hardware legs do NOT cover, found by planting the other defect
and watching it do nothing: radeonsi's decode surfaces for every fixture here
have no VERTICAL padding — `offsets[1]` is exactly `pitch * height` — so the
chroma-plane trap is untested on this driver, and `pf-vaadec`'s
`reading_chroma_at_the_display_height_would_have_been_caught` is the only place
it is checked at all. `probe_this_machines_readback_routes` now prints the
derived layout and says which of the two it is, so the next driver answers for
itself instead of being assumed.
2026-08-08 00:55:03 +02:00
enricobuehler e8a7a1e6af fix(client/vaapi): the third rung does NOT alias — and now it cannot start to
The D3D11VA and Vulkan rungs both decoded into a surface they were predicting
from, on 117 of 120 access units of our own host's low-delay H.264 (`1c54d099`
for AV1, `834b2443` for H.264). `pf-vaadec` feeds `reference_frames` from the
same `plan.dpb_refs` snapshot, releases its whole `removed` list inline exactly
as the two broken conversions did, and neither fix commit touched it. It is
still exempt — this is the evidence, and the thing that keeps it true.

**Measured on the CPU, no GPU needed.** `walk_for_aliasing` drives the planner
and `plan_to_va` over both streams and counts four shapes. On
`lowdelay-640x480.h264` the aliasing PRECONDITION is fully present: 117 of 120
access units remove a picture their own `dpb_refs` still names, and on the same
117 the setup picture is handed the slot of a picture that access unit READS —
the D3D11VA/Vulkan defect verbatim, in this conversion, today. On the vendored
conformance vector both counts are 0, which is why that vector proved nothing
on two other backends for two milestones. Aliased submissions: **0 on both**.

**Why.** A slot is not a surface here. `plan_to_va` never invents one — every
reference it can name is read out of the `surfaces` table it is handed — and
the decode target is a separate parameter the caller takes from OUTSIDE that
table. `setup_surface` reaches the submission at exactly one field per codec
(H.264/H.265 `curr_pic.picture_id`, AV1 `current_frame` and
`current_display_picture`); HEVC is doubly safe, because its per-slice
`RefPicList` stores an INDEX into `reference_frames` rather than a surface.
AV1's documented substitution fallback is the one place the target can be named
as a reference, and only where the store resolved nothing at all to prefer.

**The exemption was incidental; it is structural now.** It needs the reference
table and the decode target to come from ONE snapshot of the bindings, and the
rung had that only by writing `free_surface()` and `surface_table()` adjacently
at three call sites. Split them and this rung acquires the defect exactly: the
table must be the PRE-removal one (that is where the references are), while a
free list consulted after the removals offers precisely the displaced picture's
surface. `Session::acquire_target` now returns the index, the surface and the
table together from `&self`, so a later edit cannot move one call and not the
other. No behaviour change: same order, same values, same refusal message.

Tests. `no_submission_names_its_decode_target_as_one_of_its_own_references`
(both streams, 0) with
`taking_the_decode_target_from_the_slot_table_aliases_on_the_low_delay_stream`
as the counterfactual that reproduces the defect on 117 of 120 — so the walk
demonstrably CAN see it when it is there.
`the_low_delay_stream_reassigns_slots_whose_pictures_it_still_reads` pins 0/250
and 117/120 so neither can drift silently.
`the_decode_target_can_never_be_a_surface_the_reference_table_names` sweeps
every binding state a 4-surface/3-slot pool can hold, and
`taking_the_free_surface_after_the_removals_would_hand_out_a_referenced_surface`
is the ordering counterfactual.

⚠ One existing test lost a VACUOUS half.
`the_setup_picture_routinely_inherits_a_just_freed_slot` asserted the decode
target was never also a reference while handing every picture its own
never-reused surface id — distinct integers cannot collide, so that assertion
could not fail whatever the conversion did. Its real measurement (225 of 250
access units reuse a just-freed slot, which is why the target is a parameter)
is kept; the collision half is gone, and the doc says where the question is
actually answered and why a recycling pool is what it takes to answer it.

Gates, run on `.25` (Radeon 780M, radeonsi, Mesa 26.0.3, VA-API 1.23), this
rung being Linux-only: `cargo fmt --all -- --check`; `cargo clippy -p
pf-client-core -p pf-vaadec --all-targets --features sdl3/build-from-source --
-D warnings`; `cargo test -p pf-client-core --lib --features
sdl3/build-from-source` (171 passed); the same filtered to `video_vaapi_native`
with `--include-ignored` (18 passed); `cargo test -p pf-vaadec` (48 passed).
Plus the pf-lxcheck2 container for the cross-platform half — fmt, clippy and
`cargo test -p pf-vaadec`, all clean.

All four VAAPI legs still decode with the refactor in place, not one access
unit refused: H.264 225 of 250 access units delivering a frame, H.265 204 of
250, HEVC Main 10 45 of 50 (P010), AV1 250 of 250 — the same counts and the
same tiled modifier 0x200000010401b04 those legs recorded before it. ⚠ The
H.26x legs live on `fix/vaapi-h264-h265-hardware-proof`, not on this branch, so
they were run by overlaying that commit's test module onto the scratch tree;
only the AV1 leg and the libva probe are reachable from here. This is a decode
measurement, not frame-hash parity — the rung exports a driver-tiled DRM-PRIME
dmabuf, so there is no CPU-readable image to hash. The alias assertions above
are the real evidence and they need no device.

⚠ NOT taken: `finish`'s `outputs.last()`, which ships one frame per access unit
and drops the rest of what a bump displaces (225/204/45 against 250/250/50),
with no end-of-stream flush. It cannot bite punktfunk — hosts emit zero-reorder
output, so `outputs` never holds more than one picture — and fixing it changes
`decode()`'s one-frame-per-access-unit contract with the pump (it wants a
deliverable queue, which `video_vk_native` already keeps) plus an end-of-stream
flush and the `keyframe`-labels-the-access-unit defect in the same function.
It is recorded and asserted on that other branch, whose three delivered-count
assertions any fix has to move in the same commit; doing that from here, blind
to them, would be worse than leaving it.
2026-08-07 22:55:25 +02:00
enricobuehler a20cd44ed4 feat(client): native VAAPI AV1 — the third rung, and two failure-path defects
The libva AV1 layouts, the AuPlan conversion and the Linux rung's AV1 arm,
completing AV1 across all three hardware backends. Pin-only.

Layouts measured, not transcribed: the committed probe grew the AV1
structures and every size and offset it printed against libva 2.23.0 is a
compile-time assertion. Three that a hand-count gets wrong — the picture
buffer is align 8 because anchor_frames_list is a pointer, inserting seven
bytes of padding; seg_info and film_grain_info carry their own padding tails
inside the parent; and THREE of AV1's six bit-field unions are narrower than
a word (one uint8_t, two uint16_t), so a u32 packer over any of them writes
through its neighbour.

This is the fifth way this program has had to spell "which pictures does this
frame use", and it is unlike the other four: ref_frame_map is indexed by SLOT
and holds actual VASurfaceIDs rather than indices into anything, ref_frame_idx
is indexed by NAME and holds slots taken from the header — not from the
plan's refs, where a lost reference leaves a hole and a hole is not a slot —
global motion is picture-level, and there is no per-reference size field at
all. Established from va_dec_av1.h and libavcodec's vaapi_av1.c, and stated
in the module docs so the next reader does not re-derive it.

Review verified the whole happy path — every layout assertion re-measured,
every packer width and bit position, the reference convention, the
num_elements buffer shape — and found both defects on FAILURE paths, neither
reachable on the vendored vector.

A conversion refusal permanently desynced the ledger. The mutation block sat
after the tile walk, so any tile-shape refusal left the planner holding a
picture with no ledger slot — and the resulting UnresolvedReference fires
before that block too, so it never repaired. Every later access unit
hard-errored until a shown key frame: one lost packet costing a GOP. The
arm's own doc already warned that skipping conversion would desynchronise the
slot map; the refusal door did exactly what the skip door was written to
avoid. The block is hoisted, and a tile-shape refusal on an already-damaged
plan is now concealed rather than refused.

Fixing that exposed a sharper edge: the conversion can release a slot and
reassign it to the refused picture in one call, so the binding would still
hold the PREVIOUS picture's surface — a wrong reference rather than a missing
one, which nothing downstream could notice. The caller now clears the binding
unconditionally on the refusal path.

And a damaged frame's surface was never written yet was bound as a reference
and left in pending, so a later clean show_existing_frame would claim it with
damaged = false and ship uninitialised GPU memory to the presenter — on
several drivers another client's framebuffer. The justification quoted half
of va_dec_av1.h; its next sentence gives the remedy, which is to point the
problematic index at an alternative buffer. Damaged frames now submit as they
do on the other two arms, with live surfaces substituted for invalid entries
and reported as a bitmask — preferring a reference that really decoded over
the decode target, and keeping libavcodec's deliberate all-invalid map on a
shown key frame.

Film grain is refused rather than decoded wrong: libva wants two surfaces,
one ungrained for prediction and one grained for output, and libavcodec
allocates a second frame for exactly that. The gate now sits after the
mutation block so a grained frame costs itself rather than the GOP, and stays
per-AU rather than per-sequence because a stream that merely DECLARES the tool
decodes here perfectly.

⚠ Residual, flagged not fixed: a picture decoded from substituted references
can still be shown by a later show_existing_frame. It is decoded memory now
rather than uninitialised, and it is what the H.264/H.265 arms do, but
tracking "this was concealed" through to display needs new session state.

Gates: macOS fmt/clippy/125 tests/cargo-doc, container clippy -D warnings over
seven crates and 548 tests, workspace check. pf-bitstream's diff is
comment-only — verified — so the Vulkan rung's 250/250 stands untouched.

Nothing here has decoded a frame: no VAAPI hardware is reachable.
2026-08-07 04:22:23 +02:00
enricobuehler a6e51215fd feat(client): M6's rung is wired — libva, dlopen'd, no libavcodec
The native VAAPI decoder now runs end to end: pf-vaadec's plans go into
libva's buffers, the surface comes back as DRM-PRIME dmabufs, and the
presenter imports them exactly as it does the FFmpeg rung's. Pin-only —
`PUNKTFUNK_DECODER=native-vaapi` — for the reason M5's D3D11VA rung was:
`auto` admission is earned with hardware parity and a soak, and this rung
has decoded nothing yet.

libva is dlopen'd rather than linked, so the pf-lxcheck2 container compiles
and clippies the whole thing without libva-dev, and a machine without a
VAAPI runtime gets a clean refusal instead of a packaging dependency.

The surface pool is not the slot map. `SlotMap::assign` hands out the lowest
free slot, and a slot freed by an access unit's own removals is free by the
time that unit's picture takes it — measured at 225 of the vendored vector's
250 access units. A surface bound by slot index would therefore decode, on
nine frames in ten, into the surface still holding the picture on screen. So
`plan_to_va` now takes the decode target as a parameter, bound by the caller
at activation time the way pf-vkdecode binds a pool image, and a surface is
free only when no live picture is bound to it, no output is owed for it, and
no consumer holds it.

Measured rather than transcribed, as everywhere else here: layout-probe.c
grew the export descriptor (312 bytes, objects[4]/layers[4]), the buffer-type
enumerators — VASliceParameterBufferType is 4 and VASliceDataBufferType is 5,
not the 3 and 4 that counting off the header suggests — and the config,
attribute and generic-value layouts. All pinned as compile-time assertions,
which is how the 12-byte VAGenericValue in the first draft was caught: the C
union holds a pointer, so it is 8-aligned and 16 bytes.

The plane walk lives in pf-vaadec, pure and unit-tested on macOS, because it
is the one structure the DRIVER writes and we read: SEPARATE_LAYERS returns
NV12 as two layers, and taking layers[0] is the green screen this project has
already paid for. It also refuses what it cannot express rather than guessing
— a bogus object count, a plane naming an object that is not there, objects
disagreeing on tiling.

Own DecodedImage variant, same payload type. The physical hand-off is
identical to the FFmpeg rung's, so the presenter keeps ONE arm and one
demotion streak; the variant exists so the compiler asks which rung decoded
wherever that matters. Both D3D11VA rungs share a variant and `1573a987` had
to fix the consequence afterwards — a "native" soak that could silently have
been an FFmpeg soak. Here the four uncovered matches were compile errors.

Buffers are destroyed by us, not by vaEndPicture: va.h is explicit that the
user must call vaDestroyBuffer, and the libva 0.x behaviour is long gone.
Leaking two per picture at 60 fps exhausts the driver's store in minutes.

pf-vaadec's presenter headroom was 4, written against no consumer. The Vulkan
rung had already measured the client pipeline at four to seven held frames;
it is 8 now, pinned to that crate's constant so a re-measurement moves both.

Gates: macOS fmt/clippy/341 tests/cargo doc, and in the container clippy
-D warnings over six crates, 795 tests, workspace check.

Hardware legs are still owed — no AMD/Mesa or Intel box was reachable.
2026-08-06 16:56:38 +02:00
enricobuehler 61b96c3837 feat(client): M6's HEVC conversion — the CPU half is complete
The H.265 twin of plan_to_va, and with it pf-vaadec covers both codecs end to
end from an AuPlan to the buffers a vaRenderPicture call carries. What remains
for the rung is the Linux-only plumbing.

HEVC differs from H.264 in four ways that each had to be got right rather than
assumed, and they are why this is a separate module instead of a parameter:

ReferenceFrames is 15 entries, not 16.

The reference sets are FLAGS, not arrays. There is no RefPicSetStCurrBefore
here: membership is ORed into each DPB entry's own flags. Vulkan wants slot
indices in identically named arrays, DXVA wants list positions in them, and
VAAPI wants neither — three spellings of one idea, and confusing the first two
is what made HEVC unplayable on every driver.

The per-slice lists are INDICES into ReferenceFrames, not pictures and not
surfaces. So the DPB array is built first and every list entry resolved
through it; a picture a slice names that is not in the marked DPB is a refusal
rather than something to paper over, because there is nothing to fall back to.

The offset is in BYTES. slice_data() is byte-aligned by byte_alignment(), so
header_bit_size / 8 is exact — and a header that is not a whole number of
bytes is an error rather than a rounded offset, which would decode garbage
from the first inter picture.

Two conversions that are NOT copies, and would have been silently wrong as
copies: libva takes the derived ChromaOffsetLX (equation 7-56) where the
parser stores the coded delta, so putting the delta there would tint every
weighted-predicted block; and only 32x32 matrixIds 0 and 3 exist, where the
parser keeps six slots. The IQ matrix is Optional and gated on
scaling_list_enabled_flag for the reason review round 13 found on the DXVA
side — a driver MUST apply what it is handed, so a table of parser defaults
dequantises every residual to zero.

The weight table is only filled where 7.3.6.1 says one is coded, and the
chroma denominator is clamped into a legal shift so a malformed stream cannot
panic a decode thread.

Tests walk both HEVC vectors — the 250-frame 8-bit one and the 50-frame
Main 10 one, so a depth field wired to a constant would show — asserting per
slice that the start code was trimmed, the byte offset is inside the slice,
and every used list index points at a DPB entry that is actually valid. Per
picture it asserts that exactly the three current sets carry RPS flags and
nothing else does, and the walk fails if it never saw an RPS flag or a
reference at all, so it cannot pass vacuously.
2026-08-06 15:15:04 +02:00
enricobuehler 6c379f8fef feat(client): M6's HEVC layouts, and a third way to spell a reference set
The HEVC twin of pf-vaadec's H.264 buffer layouts, measured the same way: the
committed probe extended to cover va_dec_hevc.h, every size and offset read
off real libva 2.23.0 headers and pinned as const assertions —
VAPictureHEVC 28, VAPictureParameterBufferHEVC 604,
VASliceParameterBufferHEVC 264, VAIQMatrixBufferHEVC 1016 — and every
bit-field position read back out of a real header rather than counted by eye.

The finding worth carrying: HEVC's reference plumbing is a THIRD convention,
and this program has now been bitten by confusing two of them.

  Vulkan takes DPB SLOT indices in RefPicSetStCurrBefore/After/LtCurr.
  Writing reference-list positions there is what made HEVC unplayable on every
  driver until it was root-caused.

  DXVA takes positions into RefPicList[] in identically named arrays.

  VAAPI takes neither. It marks set membership as FLAGS on the DPB entries
  themselves — VA_PICTURE_HEVC_RPS_ST_CURR_BEFORE / _AFTER / _LT_CURR — and
  its per-slice RefPicList[2][15] holds INDICES INTO ReferenceFrames, not
  pictures and not surfaces.

Three spellings of one idea, identical names on two of them, and different
referents on all three. The conversion will say which it is writing, every
time, and the docs now hold all three side by side.

Two more asymmetries with the H.264 side, recorded where they will be read:
ReferenceFrames is 15 entries here, not 16; and the offset is
slice_data_byte_offset — BYTES, where H.264 wants bits — over the same
definition. slice_data() is byte-aligned by byte_alignment(), so the parser's
header_bit_size / 8 is exact rather than rounded, which the conversion will
assert rather than assume.

Tests cover the probe's measured bit patterns plus a disjointness sweep over
every field of pic_fields and slice_parsing_fields — two probe vectors per
word would not catch a shift typo that overlapped two neighbours, and these
words are 20 and 14 fields wide.
2026-08-06 15:09:07 +02:00
enricobuehler 7a31f1089e feat(client): M6's conversion half — one AuPlan into libva's buffers
The second half of pf-vaadec: picture parameters, inverse-quantization
matrices and one slice-parameter record per slice, over the same transaction
discipline pf-dxvadec uses — validate, resolve references against the
PRE-removal slot map, then apply removals and assign the setup slot last. A
half-applied DPB update is the shape of a corrupt reference, so nothing
mutates until every fallible step has passed.

Three things VAAPI wants that neither other backend does, all of which the
existing plan already carries:

A bit offset. slice_data_bit_offset is where slice_data() begins, counted from
and including the NAL header byte with emulation-prevention bytes removed —
DXVA takes a byte offset, Vulkan takes nothing. It costs no new parsing: the
vendored parser records exactly that as SliceHeader::header_bit_size, because
cros-codecs' own production backend is VAAPI.

The slice data without its start code, since that offset is relative to the
NAL header byte. SlicePlan::data is start-code-inclusive and the prefix is
three OR four bytes — the host emits four on every access unit — so it is
measured per slice rather than assumed. Assuming it is the defect that made
HEVC unplayable on every driver.

The per-slice reference lists. DXVA's short-format slice control expresses no
lists at all; VAAPI wants RefPicList0/1 in 8.2.4.2 order, which is what the
plan's derived lists already are.

And the distinction that cost M5 a defect, now written down in a third place:
reference_frames is documented "in DPB", the same statement DXVA's
RefFrameList makes and the opposite of Vulkan's pReferenceSlots. It is filled
from the marked-DPB snapshot; the per-slice lists come from the slice's own.
Getting that backwards loses a long-term reference no slice happens to name.

Weight tables follow 7.3.3's presence rule rather than being copied
unconditionally: flagged only where the PPS actually enables explicit
weighting for that slice type and list. Flagging them otherwise hands the
driver defaults as though the stream had coded them. The vendored
PredWeightTable stores luma_offset_l0 as [i8; 32] but luma_offset_l1 as
[i16; 32] — an upstream inconsistency, not a semantic one — so the narrow side
widens.

Envelope refusals are errors, never silent narrowings: slice groups, separate
colour planes, a capacity mismatch, a reference holding no slot, lists past
their array bounds, a slice range outside its access unit.

Tests: 15. The one that matters walks all 250 access units of the vendored
conformance vector through H264Planner and this conversion, asserting per
slice that the range lies inside its access unit, that the declared size
matches it, that the start code really was trimmed, and that the header
neither is zero bits nor outruns the slice — plus that reference_frames
carries exactly as many valid entries as the marked DPB and every entry past
it is invalidated. It also asserts it saw a multi-slice picture and a
non-empty reference set, so a splitter bug cannot make it vacuous. Gates:
rustfmt, clippy, cargo doc with no unresolved links, and the container's
clippy -D warnings, tests and workspace check.
2026-08-06 14:22:33 +02:00
enricobuehler d3e000768d feat(client): M6 begins — the libva layouts, measured rather than transcribed
The VAAPI rung's crate, in the shape the other two native rungs established:
everything that can be a pure decision or a pure conversion lives in a
cross-platform crate the ordinary gates run, and only the parts that genuinely
need a device stay behind a platform cfg. This lands the first half of that —
the buffer layouts and the decoder-creation decisions — with the conversion to
follow.

Route: minimal FFI rather than cros-libva. The plan of record permits either
("cros-libva (or minimal FFI)"), and hand-declaring keeps the crate building
and testing on macOS and in the Linux container, which is the property that
made pf-dxvadec's defects findable on a laptop instead of on a box.

The layouts are not eyeballed. A C probe compiled against real libva 2.23.0
headers printed sizeof/alignof/offsetof for every field and set individual
bit-fields to read the resulting word back; those numbers are pinned as const
assertions, so a transcription slip is a compile error rather than a driver
reading the wrong byte. The probe is committed beside them, with the command
that runs it, because evidence that cannot be re-run is a claim.

What the probe settled that a reader would otherwise get wrong: VAPictureH264
is 36 bytes and is embedded 81 times across the two buffers, so its size is
load-bearing for every later offset; the three DEPRECATED FMO fields still
occupy bytes 624..628, and dropping them would shift everything after; and C
bit-fields allocate from the least significant bit on this ABI — proven, since
that is ABI-defined rather than standardised.

Groundwork for the conversion, established here so the next work package
starts from facts:

slice_data_bit_offset needs no new parsing. VAAPI is the only backend that
wants a bit position — DXVA takes a byte offset, Vulkan takes none — and the
vendored parser already records exactly it as SliceHeader::header_bit_size,
computed as (nalu.size - epb) * 8 - bits_left: from and including the NAL
header byte, emulation-prevention bytes removed. That is the field's
definition verbatim, and it is there because cros-codecs' own production
backend is VAAPI.

The slice data buffer starts at the NAL header byte, so the start code is
skipped — SlicePlan::data is start-code-inclusive and the prefix is three OR
four bytes, the host emitting four on 100% of access units.

reference_frames is the marked DPB, the same statement DXVA's RefFrameList
makes, so it comes from the dpb_refs snapshot; Vulkan's pReferenceSlots is the
opposite and takes the access unit's own set. All three conventions now have a
written home, which is the distinction that cost M5 a defect.

Unlike DXVA short-format, VAAPI wants the per-slice reference lists and the
full prediction weight tables inline — hence the 3128-byte slice record. One
wrinkle recorded rather than left to be discovered: the vendored
PredWeightTable stores luma_offset_l0 as [i8; 32] but luma_offset_l1 as
[i16; 32], and libva wants i16 for both.

Profile selection resolves H.264 to High for every 8-bit 4:2:0 stream instead
of reading profile_idc, because High is a superset for the tools our hosts
emit and picking Main for a stream that turns out to use 8x8 transforms is a
mid-stream failure where picking High is not. 4:4:4 and 10-bit H.264 are
refused rather than narrowed to an 8-bit profile — that class of silent
narrowing decodes to garbage instead of failing.

11 tests: the probe's bit patterns, a disjointness check per bit-field word
(two probe vectors alone would not catch a shift typo that overlapped two
fields), and the envelope refusals. Gates: rustfmt, clippy, cargo doc with no
unresolved links, and the Linux container's clippy -D warnings, tests and
workspace check.
2026-08-06 13:04:07 +02:00