6cffe29b1325e4eaa424055429f8ab712a201943
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
834b244301 |
fix(client): the H.264 twin was real — every low-delay picture decoded into a surface it predicted from
The AV1 review round flagged the H.264 leg as "plausibly the same defect, traced in source, not reproduced" and deliberately did not touch it. It is reproduced now, and it is worse than the AV1 one: it fires on 297 of 300 access units of every stream a punktfunk host emits, at 720p, 1080p and 2160p alike, on BOTH the DXVA rung and the Vulkan one. **Decided on the CPU, no GPU needed.** `H264Planner` snapshots `dpb_refs` in `begin_picture`, BEFORE `finish_picture` runs 8.2.5's marking and C.4.5.3's bump, so a picture the sliding window unmarks and the bump then evicts lands in both `dpb_refs` (which `RefFrameList` is built from) and `dpb.removed`. The conversion released the whole `removed` list and then assigned the decode target a slot; `SlotMap::assign` takes the lowest free slot, which is the one just vacated. `CurrPic = N` and `RefFrameList[k] = N`, in one submission. The two conditions have to coincide in ONE access unit, and low-delay H.264 is exactly what makes them: `max_num_reorder_frames = 0` means the evicted picture has already been output, which is what makes it evictable at all. NVENC seals it by writing `max_num_ref_frames = 3` ALONGSIDE `max_dec_frame_buffering = 3` — a DPB exactly as deep as its reference count — so the window unmarks the oldest reference in the very unit whose bump drops it. The aliased picture is `ref_idx 2` of a three-entry `num_ref_idx_l0_active` list: addressable by any macroblock, not a spare. **Why two hardware-proven codecs and four GPUs never saw it.** `test-25fps.h264` is level 1.3 with no VUI `bitstream_restriction`, so `dpb_limit` falls back to A.3.1's level ceiling and gives a 7-frame DPB against 2 reference frames — the window unmarks two units before the bump can evict — and it REORDERS, which keeps an unmarked picture alive past the unit that unmarked it. Two independent reasons, both properties of that vector rather than of H.264. It measured zero and passed 250/250 throughout. `data/lowdelay-640x480.h264` is vendored to close exactly that: our own host's output, 120 pictures, goldens from libavcodec cross-checked bit-identical across two ffmpeg builds on two architectures. **The fix is the AV1 fix.** `DecodePlanDxva` and `DecodePlanVk` grow `release_after_decode`, the conversions hand the removals back instead of applying them, and the callers release them once the decode op is issued. It costs no slot the map does not have: `SlotMap::new` allocates `max_dpb_frames + 1` and the DPB never exceeds `max_dpb_frames`, so a free slot always exists with the whole `removed` list still held — measured, peak 4 of 4 on the stream that defers on 117 of 120 units. The Vulkan rung breaks on it in both DPB modes and neither loudly: DISTINCT hands the aliased reference the same array layer the setup writes; COINCIDE clears `slot_image[setup]` in the binding sync and the reference then resolves to no bound image, dropping out of `pReferenceSlots` with a `trace!`. Its deferred release runs on the FAILURE paths too — the fallible region's Result is held rather than `?`-ed, because seven exits sat between the conversion and the release and each would have leaked a slot. `a_full_dpb_bump_reuses_the_slot_but_the_pool_model_binds_a_fresh_image` asserted the aliasing as "the planner's normal behaviour": an authored depth-1 stream whose AU1 references the picture it evicts. It now asserts the opposite, which is the defect in two lines. New evidence, all of it runnable: the CPU proof pins BOTH numbers (0 on the vector, 117 of 120 on the low-delay stream) so neither can drift silently; the ledger-pressure test measures the peak; and a low-delay parity leg is added to `pf-vkdecode`'s `gpu_parity` and `pf-client-core`'s `video_d3d11_native::parity` so both rungs are held to what they stream rather than only to what they conform to. |
||
|
|
dee97e893c |
fix(vkdecode): the address the driver keeps is now the address we keep
The AV1 use-after-free fix (
|
||
|
|
a404830456 |
feat(client): wire AV1 into the native Vulkan rung, pin-only
The third codec arm in video_vk_native, AV1 admitted to native_codec and to native_vulkan_gate by pin only. It stays out of `auto` on the same rule M5's D3D11VA rung follows: `auto` admission is earned with hardware evidence, and this has decoded nothing on a device. is_integrity_warning_av1 did not exist, so the client could not have concealed AV1 damage at all. Added, exhaustive, no wildcard: all three AV1 warnings really are damage, because AV1 has no spec-legal-but-noisy signal to mis-classify — no reorder envelope to announce, no MMCO to rebase — and the exhaustive match is what stops a future variant defaulting to clean. The blocking defect review found was two safety mechanisms cancelling each other. After a failure the decoder skipped to the next key frame answering Ok(None), and because AV1's planner has no flush its store kept planning cleanly, so those AUs carried no warnings and the client read them as proof the rung works — clearing the demotion streak and resetting its clock on every one. The streak could then never reach the threshold, which made the never-delivered fall-through to FFmpeg-Vulkan unreachable, which is the documented backstop for exactly three things: a level above maxLevelIdc, a sequence header disagreeing with the Welcome, and film grain. Film grain is the probe's own admitted assumption, so a grain stream would have frozen the screen for the session while DecodeHealth reported run 0 — recovered. AV1 now answers the wait with an error, as H.264 and H.265 already do through AwaitingIdr, so all three codecs are indistinguishable to the demotion machinery. That matters more than the extra precision of a third state: only the H.26x paths have hardware evidence, and they are proven WITH that behaviour. The obvious form of that fix would have wedged the decoder. A key frame can sit behind a skipped frame inside the same temporal unit — the vendored vector has 24 two-frame units — so erroring out of the per-plan loop would never reach it and the wait would never end. Skips are therefore counted per frame and the error raised only when the whole unit was skipped, with the metadata-only unit staying a clean Ok(None). Also closed: a refused temporal unit left an already-decoded frame in the ready queue, which shipped on the next AU as a clean success — putting a picture from a refused AU on screen, clearing the streak again, and latching delivered so the fall-through was disabled for good. The error arm now drains and releases unshown. MAX_DELIVERABLE is derived rather than picked: HOLD_HEADROOM minus the pipeline's own hold, pinned to pf-vkdecode's constant so a hardcoded depth fails the build. At the previous 8 the queue plus the presenter's 4-7 stood against a headroom of 8, so it capped memory without preventing the exhaustion it named, and a frame waiting 8 AUs burned 16 of the 17 query slots — where a re-armed slot reads as Failed and becomes a fabricated driver-corruption verdict in the very counter the Ally X signal lives in. The trim now runs after this AU's frame is taken, or at the derived depth it would drop a two-output unit's first frame and invert display order inside one AU. Its justification was also wrong: the claim that a temporal unit may carry a show_existing_frame alongside a shown frame is disproved by this repo's own golden — 250 units, 250 shown, zero show_existing. The bound is kept as defence in depth against a non-conformant or multi-operating-point stream, and now says so. Gates: macOS fmt/clippy/392 tests, container clippy -D warnings over six crates, 851 tests, workspace check. No hardware: the rung is pin-only and has still never decoded a frame on a device. |
||
|
|
cab3aa1726 |
feat(vkdecode): M7's Vulkan AV1 rung — GPU half, and the review that saved it
caps_av1 / session_av1 / decoder_av1, over the CPU half already committed, sharing the picture pool, bitstream ring, op ring, DPB settling and frame delivery with H.264 and H.265 rather than forking them. AV1 session parameters carry exactly one sequence header — no PPS, no VPS — so the parameters ledger is two-state: current, or recreate. The GPU plumbing came through review clean. The damage was all in the conversion committed two rounds ago, which nothing tested against a reference, and none of it would have failed a gate: clippy was clean, the tests were green, and the rung would have decoded its own conformance vector wrong on essentially every frame on AMD, silently. Four blocking defects, each measured on the vendored vector rather than argued: Nine StdVideoDecodeAV1PictureInfo flags were never set. Four change reconstruction — allow_screen_content_tools on 274 frames of 274, allow_warped_motion on 273, is_filter_switchable on 172, force_integer_mv on 1 — and RADV reads three of them directly. The block already set allow_intrabc, which is only codeable when screen-content tools are on, so it contradicted itself. LoopRestorationSize sent the pixel size where the field is log2(size) - 5. cros-codecs stores 64/128/256; RADV names its destination log2_restoration_size_minus5 and reads 1/2/3. Nothing truncates, nothing errors, and every frame with loop restoration reconstructs against a nonsense unit size. Per-reference Std info answered questions about the wrong picture: every reference carried the CURRENT frame's type, and RefFrameSignBias was never set at all. Sign bias is what tells a decoder a reference lies in the future, and this vector is the hidden-ALTREF one, so all-zero meant every reference was treated as past. Fixed at the source: pf-bitstream now records a RefState when a picture is stored — its own frame type, sign-bias mask, saved order hints — and carries it on the slot, so all three backends get answers about the reference rather than about the frame reading it. Film grain's six chroma-scaling fields were zero, which defeats the profile machinery that exists to refuse devices unable to synthesise grain. The reference-name compaction is fixed in the PLANNER, once. AuPlan::refs is now name-indexed with holes preserved, so a lost reference can no longer renumber every later AV1 reference name — a class that was live in both conversions and armed for the VAAPI rung that does not exist yet. The DXVA twin had a second name-versus-slot confusion: it read global motion by DPB slot from an array the spec indexes by reference name, and slot 0's matrix is all-zero rather than identity, so 273 references were given a zero warp. Also closed: pTileOffsets/pTileSizes were sized to tileCount while RADV reads AV1_MAX_NUM_TILES entries unconditionally — a 4-byte allocation read a kilobyte deep — now fixed 256-entry arrays with zeroed tails. And the test guarding the lost-reference refusal re-implemented the predicate inline, so deleting the guard left it green; both now call one named function. The bitstream layout now matches libavcodec: raw tile payloads only, frameHeaderOffset 0. The review established the spec-literal layout was NOT wrong — AV1 has no start-code scanning, so the 3-versus-4-byte and slices-only scars do not transfer, and no driver in the fleet reads frameHeaderOffset — but matching the validated reference deletes code, uploads 5835 fewer bytes over the vector, and removes the untested-driver tail. Upstream, and the third of its kind: the vendored parser writes ref_frame_sign_bias[i] in the same loop body where it writes order_hints[LAST_FRAME + i], so its array is shifted one down and index 7 is never written. Corrected in RefState::of with the shift documented, the vendored tree untouched, and pinned by a test that recomputes the bias from order_hints through the parser's own get_relative_dist. Gates: macOS fmt/clippy/tests, container clippy -D warnings over six crates, 845 tests, workspace check. No hardware: nothing here has reached a driver. |
||
|
|
db15c2615d |
fix(vkdecode): HEVC decoded from the wrong references, and told drivers the
wrong slice offsets Two independent defects, both in this crate. pf-bitstream is untouched — its HEVC plans were sound all along, which the D3D11VA rung proves by rendering correctly from the same AuPlans. The corrupter: StdVideoDecodeH265PictureInfo's RefPicSetStCurrBefore, StCurrAfter and LtCurr carry DPB SLOT indices. We wrote positions in the reference list. libavcodec's Vulkan HEVC hwaccel — the implementation every driver is validated against — writes the index into its own DPB array and passes that same value as slotIndex, while packing pReferenceSlots densely over the used entries; the two numberings are provably different there, and the RPS arrays follow the slot. The two readings coincide on a freshly anchored stream, because the references then occupy slots 0..n in reference-list order. They first diverge at the vendored vector's first B picture, AU 3, where refs are slots 0, 2, 1 — so we named slots 1 and 2 where the picture wanted 2 and 1, and every later access unit inherited the error through its own references. That is why this shipped and why review could not see it: correct for the opening pictures, wrong from the first reordering onward. It also accounts for the measurements exactly. Display order maps to decode order as display 0 from AU 0, display 3 from AU 1, display 2 from AU 2, display 1 from AU 3 — so the three frames that matched on AMD are precisely the three access units where positions and slots agree, and 250 - 3 = 247 is the divergence count that was measured. From AU 6 the named slots stop being merely wrong and become unbindable by that operation, which is where NVIDIA stopped reporting a verdict at all. The diagnostic: every slice offset must point at a THREE-byte start code. libavcodec discards the stream's prefix and writes 00 00 01, so that is the only pattern drivers are validated on, and pf-dxvadec's packer already normalised for exactly this reason and said so in its docs. This path uploaded the prefix verbatim. 249 of 250 HEVC slice segments in the vendored vector carry a four-byte prefix; all 500 H.264 slices carry three. That derives the driver's own complaint bit for bit. A decoder reaching the slice header by a fixed skip lands on the NAL header's second byte, reads first_slice_segment_in_pic_flag as 0, and then takes six bits of the real slice header as the tail of a long ue(v): 0xd0 gives 115, 0xe0 gives 119. A P-slice header and a B-slice header — which is why exactly two bogus pps_id values ever appeared. ⚠ H.264 was NOT protected structurally, only by its encoder's convention, and the real host does not share that convention: every one of 1514 H.264 access units and 1133 HEVC access units captured from an NVENC host prefixes its slices with FOUR bytes. The vendored H.264 vector is therefore not representative of what ships, and its bit-exactness was passing on a prefix form the field never sends. The normalisation lives in the shared ring layer and covers both codecs for that reason. rebased_offsets is replaced by pack_slices, which trims the leading zero byte and computes the offsets from the trimmed lengths in one call, so the bytes and the offsets cannot drift apart; upload and the CPU test go through the same pack_into. Hardware, after the fix — H.264 AND H.265 both 250/250 bit-identical to libavcodec, all four smoke legs green: NVIDIA RTX 4090 610.88 Windows coincide AMD Adrenalin 25.10.30.02 Windows distinct NVIDIA RTX 5070 Ti 610.43.03 Linux coincide On glass on the 4090 against a real NVENC host, 2800x1260 HEVC through the auto ladder: 73 one-second windows all native-vulkan, fps avg 59.3 of 60, decode 1.1 ms, e2e 4.5 ms p50, and ZERO driver-reported status failures where the same session before the fix logged 1489 in 181 seconds and had dragged ABR down to a 5 Mb/s target. No refusals, demotions, PlanWarnings, concealment, DEVICE_LOSTs or panics. Both defects now have CPU tests that were confirmed FAILING before the fix: one walks every access unit of both vendored vectors and asserts each declared offset opens on a three-byte start code, its own NAL header and a first_slice_segment_in_pic_flag consistent with the segment index; the other resolves every RPS entry by slot and asserts that 247 access units disagree with the positional reading, so it cannot go vacuous on a stream where the two happen to agree. |
||
|
|
2a57ee36f8 |
feat(client): M4 — the decoder's own verdict reaches the session
This program exists because a field corruption was architecturally undetectable through FFmpeg: no decode-status read, no corrupt-frame flag, errors only as scraped log lines, and no recovery-point signal so intra-refresh healing was invisible. The native decoder has all of those. M4 is where they stop being internal. DecodeHealth counts, per session and without allocating per frame, what the three answers actually are: damaged (the stream arrived incomplete), refused (the rung would not decode it at all) and driver-failed (the hardware says it could not decode what arrived), plus the current and worst concealment run — the figures that separate one bad AU from a stream that never came back. They ride the stats line additively, so an FFmpeg session and a healthy native session emit byte-identical output to today. The status-query capability is reported too: without it a clean report cannot be told from an unmeasured one, which is the whole nb_queries=0 lesson. The headline is local recovery. Until now the pump could only learn that intra-refresh healing finished from wire flags the host sends; absent those it froze until the 500 ms backstop forced an IDR. The parsed recovery-point SEI now feeds the re-anchor gate directly, so a session lifts on the picture that is actually clean. Wire semantics are untouched for every client that never calls it. Detection now asks for recovery instead of erroring — an integrity warning ticking the error streak would demote the native rung on exactly the lossy links it exists to diagnose, where an FFmpeg rung conceals silently and keeps its job. Review round 12 found that trade had removed the escape hatch entirely. Concealment returning Ok(None) reset the demotion streak, and worse: the driver-verdict ledger is only populated when a frame ships, so under continuous concealment no verdict was ever read and the erroring arm could not fire at all. A host framing regression of the 0.23.0 slice-wire class — which does not self-heal, and which a keyframe does not clear — would have frozen indefinitely with no demotion and a clean integrity line, where before it demoted to FFmpeg-Vulkan and showed a picture. Now only an answer that proves the rung works clears the streak: a shipped frame, or a clean no-frame. Concealment neither ticks nor clears, so a lossy link still cannot demote a healthy rung while a driver failure interleaved with concealment reaches the threshold again. Two more honesty defects from the same round. A rung refusing every AU reported no integrity line at all — the founding failure mode, wearing the shape of a clean bill of health; refusals are now counted. And driver-failed could be non-zero on a device that cannot produce driver verdicts, because a degraded timeline read looked the same as one; the attribution is now withheld inside the counter rather than at call sites, so the self-contradictory line is unrepresentable. Local recovery also no longer trusts any recovery-point SEI: only one whose target advances past an outstanding wave counts as a new wave, so an encoder re-announcing the current wave with a decreasing count — legal, and what x264 intra-refresh does — cannot lift the freeze early onto a partially stale picture. Frames buffered across an arm are dropped by decode order for the same reason. Fault injection is a first-class tool now (PUNKTFUNK_AU_FAULT, inert unless set, env read once). Its test replays the vendored vectors through the real planners and asserts a negative the plan assumed away: truncation and bit flips are PROVABLY invisible to the parser — Annex-B carries no NALU length, so a cut slice is just a shorter slice and a flipped payload byte is syntactically perfect. Only dropped AUs are parser-detectable; the rest need the driver verdict, which is why the status query matters. The H.265 leg found a second: three of that vector's faulted AUs are sub-layer non-reference pictures, so dropping them damages nothing and silence is correct — the test asserts both verdicts and guards that neither half goes vacuous. Per-frame decode latency was deliberately NOT built. Polling answers only 'complete by now', and the pump polls once per AU, so every sample would quantise up by as much as a frame interval — 8.3 ms at 120 Hz against decodes of 0.1-2 ms. Sampling faster needs a spin or a second thread on a decoder that is deliberately not Sync. A blocking per-frame wait is the field scar that once capped a stream at 51 fps. The honest sampled stat stands. Also fixed, pre-existing: the re-anchor gate re-armed on every damaged AU, so sustained damage permanently zeroed the mark count — meaning the wire's two-mark rule could never complete on exactly the lossy links it was written for. Field note recorded while wiring this: intra_refresh_recovery is set by exactly one encoder backend (Linux libav-NVENC under PUNKTFUNK_INTRA_REFRESH). AMF and QSV run a wave with no wire mark, and AMF emits no recovery-point SEI either, so AMD/Windows intra-refresh sessions still have no clean recovery point by either route. Gates: fmt clean; container clippy -D warnings zero across pf-client-core + pf-presenter + pf-vkdecode + punktfunk-core; tests 69/131/129/354/41 plus 5 fault-detection green; cargo check --workspace clean. |
||
|
|
6d8f3b45b5 |
feat(pf-vkdecode): the GPU half of HEVC decode — session, pools, recording
M3 WP-2 complete. caps_h265.rs builds the profile the stream actually needs (profile idc + chroma + bit depths, all three stated on every Vulkan object) and resolves its picture format — Main to NV12, Main 10 to P010, RExt 4:4:4 to the two-plane 4:4:4 formats — validating it against the format list of every role the chosen arrangement creates images in. A Main 10 stream on an 8-bit-only device is refused BEFORE a session exists, never narrowed: decoding 10-bit into an 8-bit surface is the silent-wrongness class this crate exists to refuse. session_h265.rs adds the three-array parameters ledger; decoder_h265.rs adds VkH265Decoder, mirroring VkH264Decoder method-for-method so the client wiring is a two-arm dispatch away. H.264 and H.265 now SHARE the machinery instead of duplicating it: derive_arrangement (one coincide/distinct/layered decision table), ring::rebased_offsets (the slices-only rebase — non-VCL NALUs in the decode range hang VCN firmware), session::bind_session_memory, and a parameterised build_frame. A DecodeProfile enum replaces the bare profile idc that images.rs and ring.rs used to take: both codecs' idc types are c_uint, so handing an H.265 idc to the H.264 path COMPILED SILENTLY and built a mismatched profile chain. That is now unrepresentable. The VPS leg is the ledger's real work. The vendored parser attaches a VPS to an SPS only when it saw the NALU, and clients join live streams, so VpsSource is Parsed-or-FromSps and is stored BY VALUE: re-activating a VPS-less SPS is Current (no churn), but the real VPS arriving under the same id is a content change and RECREATES onto it, because Vulkan cannot replace a stored parameter set. Review round 10 (adversarial) confirmed the hardware-proven H.264 path is NOT regressed — derive_arrangement's check order and error identity are byte-for-byte the original, build_frame's call sites still pass the granularity-aligned extent (the 1088-row scar stays shut), and rebased_offsets reproduces the deleted inline loop for every input while moving the sum to u64 so overflow errors instead of wrapping. Also verified: the refs-order contract on every path, the RESULT_STATUS caps gate (each of reset/begin/end individually gated, no pool created when unsupported — recording one on RADV hangs its VCN), pNext lifetimes, and that no panic is reachable on stream input. Its 10 findings are fixed. The two that mattered: - A failed decode stranded a DPB slot. Once plan_to_vk_h265 had mutated the slot map, five later failure paths returned without restoring it, so planner and slot map both believed a picture was resident while no image held it — and every later AU referencing it failed, where H.264 soft-degrades and keeps delivering. Fail-closed is kept (substituting a reference silently is the corruption-hiding this program exists to end) but made RECOVERABLE: a latch flushes the planner to AwaitingIdr and resets the bindings on the next decode, which composes with the client already requesting a keyframe on every decode error. The fix deliberately covers pre-mutation failures too — those strand the picture the other way round and wedge identically. - DecodedVkFrame carried no picture format, so a Main 10 frame would decode correctly and be rendered with 8-bit transfer/range math. It now carries one, stamped from the pool so it is truthful for both decoders by construction. The presenter comment says depth 8 is because only H.264 is WIRED, not a decoder limit. Plus: bind_session_memory freed allocations before the session that may hold them was destroyed (an ordering regression from the extraction, with a SAFETY comment asserting the opposite) — the bind-stage exit now hands them back so Drop destroys first; max_level_idc is codec-tagged rather than an H.264 type carrying H.265 code points; and the decode family's videoCodecOperations is now checked, turning 'create an H.265 session on a device without the extension' from UB into a clean ladder demote. Deferred by design: no HEVC gpu_smoke/gpu_parity yet (its goldens are already in tests/data/test-25fps-h265.nv12.sha256), and no codec dispatch in the client — both later legs. Gates: fmt clean; mac clippy zero warnings, pf-vkdecode 106 + pf-bitstream 69 green; container clippy -D warnings zero for pf-client-core + pf-presenter + pf-vkdecode, tests 69/121/106 green. HARDWARE (.173, after the refactor — review saying the proven path is safe is not the GPU saying it): gpu_parity '250 frames bit-identical to libavcodec software decode' on BOTH the NVIDIA 4090 (610.88, coincide mode) and the AMD iGPU (Adrenalin 25.10.30.02, distinct mode), gpu_smoke green on both. Two independent drivers, both DPB modes, still bit-exact. The smoke trace also shows the new videoCodecOperations capture reading DECODE_H264 | DECODE_H265 | DECODE_AV1 off the real decode family. |
||
|
|
dc0766b2f3 |
fix(client): the native rung now follows the stream's colour and reports true decode latency
The round-4 residuals, closed after the WP-D hardware verdict: - VUI colour plumbing (the one silent-wrong): the picture's ACTIVE SPS's colour signalling (H.273 code points + range, with E.2.1's 'unspecified' inference where the VUI is silent — the vendored parser's defaults ARE the inferred values, verified) rides PicturePlan -> DecodedVkFrame -> NativeVkFrame per frame, never latched: the Windows host switches an HDR desktop to PQ/BT.2020 IN-BAND while the Welcome still says SDR. Before this, the native path would have painted PQ washed out, silently. - Native decode-latency stat: the deliberately-deferred NativeVk arm of the pump's sampled once-per-stats-window decode measurement now feeds - the frame's (semaphore, semaphore_value) is the decode-done signal, resolved through the shipped ledger before a bounded, pure-measurement vkWaitSemaphores (VkH264Decoder::wait_decoded). - The renegotiation-teardown window is settled as NO HOLE: rebuild_state now documents the full safety argument (graveyarded pools stay intact under presenter holds, tokens route strictly by generation, session objects die only post-drain with the generation gate INSIDE read_status), and the two backend comments that wrongly claimed stale pools were 'gone' are fixed. - VK_KHR_unified_image_layouts stays deferred (fleet drivers lack it). Adversarial review round 6: 3 minor findings (2 doc fixes applied; the SPS-replaced-without-PPS-resend divergence stays a documented envelope assumption - hosts re-send both at every keyframe, and a hardening PlanWarning could cost real frames on a false positive). Gates: fmt clean; clippy -D warnings zero (mac + pf-lxcheck2 container, incl. pf-client-core/pf-presenter); tests 45+30+53 mac, 30+121+53 container. |
||
|
|
e6d6498a49 |
test(pf-vkdecode): frame-hash parity vs libavcodec — bit-exact on the whole fleet
WP-D parity A/B. gpu_parity (ignored) decodes the conformance vector, reads every frame back through the presenter's exact contract (wait, layout round-trip, signal-back, release), crops at the copy so pitch can never leak, and compares SHA-256s in display order against goldens from ffmpeg software decode — cross-checked bit-identical between ffmpeg 8.0.1 (linux) and 8.1.1 (macOS), so the reference is the spec, not one build. PF_VKD_TEST_READBACK=1 is the one test-only hook (ORs TRANSFER_SRC into pool usage; production pools stay zero-copy-tight). Fleet verdict: 250/250 frames bit-identical to libavcodec on RADV (Mesa 26.0.3, distinct), AMD proprietary Windows (25.10.30.02, distinct) and NVIDIA Windows (610.88, coincide) — H.264 decode is exactly specified, and the native path meets the spec on every driver and both DPB arrangements. |
||
|
|
6331ae7fd9 |
fix(pf-vkdecode): zero-copy pool model + the two faults the first hardware run found
WP-D leg 1 (.25 RADV, distinct mode) root causes, both real: 1. Output starvation: the fixed 4-deep ring lost to a stream that keeps max_dpb_frames+1 = 8 pictures pending. Zero-copy fix (user requirement, no copies): one picture pool of required_slots + HOLD_HEADROOM(8) images decoupled from DPB slots — a re-activated slot binds a fresh free image, so a delivered picture is never a decode target; the WP-B pin layer became dead and is deleted. Per-image timeline semaphores carry the AVVkFrame contract: decode signals value+1, the presenter waits and signals back, later decodes wait the image's latest value — layout traffic ordered against reference reads with no copy anywhere. 2. RESULT_STATUS queries HANG RADV's VCN firmware (ring timeout, DEVICE_LOST): queryResultStatusSupport=false on the decode family. Queries are now caps-gated; without them poll/wait degrade to timeline-completion verdicts (FFmpeg parity — and the likely reason upstream never wired nb_queries). The Ally-X-class detection runs where drivers advertise the query; .173 probes NVIDIA/Windows-AMD. Also: slice-only bitstream feeding (the field-proven consumer shape), graveyarded pool retirement keyed by release tokens + generation, decode-current-AU-before-status attribution, take_ready drained, H264-bit gating, teardown short-circuit on disconnected channel. On-glass: 48 AUs green on .25 holding 4 frames like the real client. Gates: fmt clean, container clippy -D warnings zero, 27+121+52 green both platforms. |
||
|
|
d0659d2b61 |
feat(client): wire the native Vulkan decoder in behind PUNKTFUNK_DECODER=native-vulkan
M2 WP-C. video_vk_native.rs adapts the presenter's VulkanDecodeDevice to pf-vkdecode (queue lock shared only when the families actually collide — the one case the 2026-07-09 DEVICE_LOST race proved matters), and the presenter consumes DecodedImage::NativeVk on its own device: no handle import, no AVVkFrame co-authoring — wait the timeline, barrier to sampled, existing crop-aware CSC, barrier back, release after the fence. Frame lifetime is a token: presented, retired, displaced or dropped mid-demotion, the guard's drop sends it exactly once; the backend releases the decoder slot only after the status query resolves, so a recycled slot can never report a false Failed. Driver-reported decode failures and plan warnings ride the existing streak/reanchor machinery — the Ally X corruption class is now a visible error, not a silent frame. Opt-in only until WP-D's on-glass parity verdict; H.264 sessions only; failures demote to the existing ladder. Known WP-D items recorded in code: coincide-mode cross-queue reference overlap, renegotiation teardown window, VUI colour plumbing. Gates: fmt clean; container clippy -D warnings zero for pf-client-core + pf-presenter + pf-vkdecode; 121+53+27 tests green. |
||
|
|
540c0d3027 |
feat(pf-vkdecode): the GPU half — session, DPB pools, decode recording, status queries
M2 WP-B. VkVideoSessionKHR lifecycle with drain-before-destroy on parameters recreation, DPB pools in both coincide and distinct modes (caps-derived, usage/flags validated against the driver's format properties), an aligned bitstream ring, vkCmdDecodeVideoKHR recording with one-shot RESET re-armed on failed submits, timeline-semaphore completion, and the per-op RESULT_STATUS query ring — the signal FFmpeg's hwaccel never reads and the reason this program exists. Frame lifetime is two-phase by construction: release_frame pins a delivered frame's slot against reuse, closing the coincide-mode overwrite the adversarial review round proved (a full DPB handed a just-returned frame's image back as the same call's decode target). Nine review findings fixed pre-commit; a counterfactual test pins the collision. Generation-stamped frames, memory-type misses as errors, granularity-aligned extents, level gate. AuPlan now carries its activated SPS/PPS (Rc) so backends never re-parse. GPU smoke test (ignored) decodes 48 AUs past DPB-full with releases — the fleet runs it in WP-D. Gates: fmt clean, clippy -D warnings zero, 45+27+53 tests green on macOS and the linux/amd64 container. |