2fa40f16fa0a97cc18516bd5ec209b54f91f9124
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
db15c2615d |
fix(vkdecode): HEVC decoded from the wrong references, and told drivers the
wrong slice offsets Two independent defects, both in this crate. pf-bitstream is untouched — its HEVC plans were sound all along, which the D3D11VA rung proves by rendering correctly from the same AuPlans. The corrupter: StdVideoDecodeH265PictureInfo's RefPicSetStCurrBefore, StCurrAfter and LtCurr carry DPB SLOT indices. We wrote positions in the reference list. libavcodec's Vulkan HEVC hwaccel — the implementation every driver is validated against — writes the index into its own DPB array and passes that same value as slotIndex, while packing pReferenceSlots densely over the used entries; the two numberings are provably different there, and the RPS arrays follow the slot. The two readings coincide on a freshly anchored stream, because the references then occupy slots 0..n in reference-list order. They first diverge at the vendored vector's first B picture, AU 3, where refs are slots 0, 2, 1 — so we named slots 1 and 2 where the picture wanted 2 and 1, and every later access unit inherited the error through its own references. That is why this shipped and why review could not see it: correct for the opening pictures, wrong from the first reordering onward. It also accounts for the measurements exactly. Display order maps to decode order as display 0 from AU 0, display 3 from AU 1, display 2 from AU 2, display 1 from AU 3 — so the three frames that matched on AMD are precisely the three access units where positions and slots agree, and 250 - 3 = 247 is the divergence count that was measured. From AU 6 the named slots stop being merely wrong and become unbindable by that operation, which is where NVIDIA stopped reporting a verdict at all. The diagnostic: every slice offset must point at a THREE-byte start code. libavcodec discards the stream's prefix and writes 00 00 01, so that is the only pattern drivers are validated on, and pf-dxvadec's packer already normalised for exactly this reason and said so in its docs. This path uploaded the prefix verbatim. 249 of 250 HEVC slice segments in the vendored vector carry a four-byte prefix; all 500 H.264 slices carry three. That derives the driver's own complaint bit for bit. A decoder reaching the slice header by a fixed skip lands on the NAL header's second byte, reads first_slice_segment_in_pic_flag as 0, and then takes six bits of the real slice header as the tail of a long ue(v): 0xd0 gives 115, 0xe0 gives 119. A P-slice header and a B-slice header — which is why exactly two bogus pps_id values ever appeared. ⚠ H.264 was NOT protected structurally, only by its encoder's convention, and the real host does not share that convention: every one of 1514 H.264 access units and 1133 HEVC access units captured from an NVENC host prefixes its slices with FOUR bytes. The vendored H.264 vector is therefore not representative of what ships, and its bit-exactness was passing on a prefix form the field never sends. The normalisation lives in the shared ring layer and covers both codecs for that reason. rebased_offsets is replaced by pack_slices, which trims the leading zero byte and computes the offsets from the trimmed lengths in one call, so the bytes and the offsets cannot drift apart; upload and the CPU test go through the same pack_into. Hardware, after the fix — H.264 AND H.265 both 250/250 bit-identical to libavcodec, all four smoke legs green: NVIDIA RTX 4090 610.88 Windows coincide AMD Adrenalin 25.10.30.02 Windows distinct NVIDIA RTX 5070 Ti 610.43.03 Linux coincide On glass on the 4090 against a real NVENC host, 2800x1260 HEVC through the auto ladder: 73 one-second windows all native-vulkan, fps avg 59.3 of 60, decode 1.1 ms, e2e 4.5 ms p50, and ZERO driver-reported status failures where the same session before the fix logged 1489 in 181 seconds and had dragged ABR down to a 5 Mb/s target. No refusals, demotions, PlanWarnings, concealment, DEVICE_LOSTs or panics. Both defects now have CPU tests that were confirmed FAILING before the fix: one walks every access unit of both vendored vectors and asserts each declared offset opens on a three-byte start code, its own NAL header and a first_slice_segment_in_pic_flag consistent with the segment index; the other resolves every RPS entry by slot and asserts that 247 access units disagree with the positional reading, so it cannot go vacuous on a stream where the two happen to agree. |
||
|
|
2a57ee36f8 |
feat(client): M4 — the decoder's own verdict reaches the session
This program exists because a field corruption was architecturally undetectable through FFmpeg: no decode-status read, no corrupt-frame flag, errors only as scraped log lines, and no recovery-point signal so intra-refresh healing was invisible. The native decoder has all of those. M4 is where they stop being internal. DecodeHealth counts, per session and without allocating per frame, what the three answers actually are: damaged (the stream arrived incomplete), refused (the rung would not decode it at all) and driver-failed (the hardware says it could not decode what arrived), plus the current and worst concealment run — the figures that separate one bad AU from a stream that never came back. They ride the stats line additively, so an FFmpeg session and a healthy native session emit byte-identical output to today. The status-query capability is reported too: without it a clean report cannot be told from an unmeasured one, which is the whole nb_queries=0 lesson. The headline is local recovery. Until now the pump could only learn that intra-refresh healing finished from wire flags the host sends; absent those it froze until the 500 ms backstop forced an IDR. The parsed recovery-point SEI now feeds the re-anchor gate directly, so a session lifts on the picture that is actually clean. Wire semantics are untouched for every client that never calls it. Detection now asks for recovery instead of erroring — an integrity warning ticking the error streak would demote the native rung on exactly the lossy links it exists to diagnose, where an FFmpeg rung conceals silently and keeps its job. Review round 12 found that trade had removed the escape hatch entirely. Concealment returning Ok(None) reset the demotion streak, and worse: the driver-verdict ledger is only populated when a frame ships, so under continuous concealment no verdict was ever read and the erroring arm could not fire at all. A host framing regression of the 0.23.0 slice-wire class — which does not self-heal, and which a keyframe does not clear — would have frozen indefinitely with no demotion and a clean integrity line, where before it demoted to FFmpeg-Vulkan and showed a picture. Now only an answer that proves the rung works clears the streak: a shipped frame, or a clean no-frame. Concealment neither ticks nor clears, so a lossy link still cannot demote a healthy rung while a driver failure interleaved with concealment reaches the threshold again. Two more honesty defects from the same round. A rung refusing every AU reported no integrity line at all — the founding failure mode, wearing the shape of a clean bill of health; refusals are now counted. And driver-failed could be non-zero on a device that cannot produce driver verdicts, because a degraded timeline read looked the same as one; the attribution is now withheld inside the counter rather than at call sites, so the self-contradictory line is unrepresentable. Local recovery also no longer trusts any recovery-point SEI: only one whose target advances past an outstanding wave counts as a new wave, so an encoder re-announcing the current wave with a decreasing count — legal, and what x264 intra-refresh does — cannot lift the freeze early onto a partially stale picture. Frames buffered across an arm are dropped by decode order for the same reason. Fault injection is a first-class tool now (PUNKTFUNK_AU_FAULT, inert unless set, env read once). Its test replays the vendored vectors through the real planners and asserts a negative the plan assumed away: truncation and bit flips are PROVABLY invisible to the parser — Annex-B carries no NALU length, so a cut slice is just a shorter slice and a flipped payload byte is syntactically perfect. Only dropped AUs are parser-detectable; the rest need the driver verdict, which is why the status query matters. The H.265 leg found a second: three of that vector's faulted AUs are sub-layer non-reference pictures, so dropping them damages nothing and silence is correct — the test asserts both verdicts and guards that neither half goes vacuous. Per-frame decode latency was deliberately NOT built. Polling answers only 'complete by now', and the pump polls once per AU, so every sample would quantise up by as much as a frame interval — 8.3 ms at 120 Hz against decodes of 0.1-2 ms. Sampling faster needs a spin or a second thread on a decoder that is deliberately not Sync. A blocking per-frame wait is the field scar that once capped a stream at 51 fps. The honest sampled stat stands. Also fixed, pre-existing: the re-anchor gate re-armed on every damaged AU, so sustained damage permanently zeroed the mark count — meaning the wire's two-mark rule could never complete on exactly the lossy links it was written for. Field note recorded while wiring this: intra_refresh_recovery is set by exactly one encoder backend (Linux libav-NVENC under PUNKTFUNK_INTRA_REFRESH). AMF and QSV run a wave with no wire mark, and AMF emits no recovery-point SEI either, so AMD/Windows intra-refresh sessions still have no clean recovery point by either route. Gates: fmt clean; container clippy -D warnings zero across pf-client-core + pf-presenter + pf-vkdecode + punktfunk-core; tests 69/131/129/354/41 plus 5 fault-detection green; cargo check --workspace clean. |
||
|
|
e4d8573475 |
feat(client): the native rung now decodes HEVC as well as H.264
The last piece of M3 WP-2 — VkH265Decoder was built and hardware-gated but nothing drove it. video_vk_native.rs holds a two-arm codec enum and forwards to it; the ledger, release tokens, status-query settling and timeline waits are byte-for-byte what they were, since they were always codec-agnostic over one DecodedVkFrame contract. The forwarders are written out per arm rather than macro'd so the unchanged H.264 arm is visible to a reviewer. The picture's own format now reaches the presenter, which picks bit depth and MSB packing from it instead of assuming the H.264 envelope. That incidentally fixes a live bug on the SHIPPING FFmpeg-Vulkan path: it derived ten-bit-ness by comparing against the 10-bit 4:2:0 format alone, so a 10-bit two-plane 4:4:4 surface — which its own format table accepts, and which NVIDIA reports for HEVC RExt — got 8-bit range and transfer maths. Reachable today with Full chroma plus 10-bit: decoded correctly, displayed wrong. Review round 11 caught a regression this WP would otherwise have shipped. pf-vkdecode refuses a stream whose (chroma, depth) pair has no picture format on the device, but the session is built lazily from the first SPS, so the refusal arrived AFTER construction — past the point where a native init failure falls through to FFmpeg-Vulkan. It burned the error streak instead and demoted to VAAPI/D3D11VA, which on NVIDIA/Linux means software. Turning on Full chroma on any non-NVIDIA GPU was enough: a 4K HEVC session that ran on FFmpeg-Vulkan before this branch would have landed on software decode. Both halves are fixed. The negotiated chroma and bit depth — already at the call site, the PyroWave arm four lines up uses them — are threaded into the backend, which probes the same caps path ensure_state would run, so the whole class refuses at CONSTRUCTION where the fall-through already exists. For the legs no negotiation can carry (a level above maxLevelIdc, an SPS that disagrees with the Welcome) the decoder latches 'never delivered a frame' and routes that first streak to FFmpeg-Vulkan rather than down the hardware ladder. H.264 is deliberately not probed: its envelope is fixed, so a probe would only add a profile guess on the bit-exact path; it gets the latch as its backstop. Two more from the round. Planner warnings are typed again rather than Debug strings — pf-vkdecode simply lacked the h265 re-export its h264 twin already had — which restores the H.264 log rendering exactly and unblocks M4, whose job is counting concealment by kind. And concealment is now the integrity set only: NonZeroReorder is documented spec-legal and fully planned, but the client treated every warning as damage, so the opening IDR and every ABR renegotiation's IDR were released unshown and re-anchored — a visible hitch on a healthy stream. Also: a raw-format newtype so a neighbouring i32 field cannot be passed to the colour maths, the presenter's depth table now pinned against pf-vkdecode's actual output vocabulary rather than the FFmpeg lane's, a per-format warn latch, and four stale docs. Gates: fmt clean; container clippy -D warnings zero across pf-client-core + pf-presenter + pf-vkdecode; tests 69/125/108/40 green; cargo check --workspace clean. |
||
|
|
6d8f3b45b5 |
feat(pf-vkdecode): the GPU half of HEVC decode — session, pools, recording
M3 WP-2 complete. caps_h265.rs builds the profile the stream actually needs (profile idc + chroma + bit depths, all three stated on every Vulkan object) and resolves its picture format — Main to NV12, Main 10 to P010, RExt 4:4:4 to the two-plane 4:4:4 formats — validating it against the format list of every role the chosen arrangement creates images in. A Main 10 stream on an 8-bit-only device is refused BEFORE a session exists, never narrowed: decoding 10-bit into an 8-bit surface is the silent-wrongness class this crate exists to refuse. session_h265.rs adds the three-array parameters ledger; decoder_h265.rs adds VkH265Decoder, mirroring VkH264Decoder method-for-method so the client wiring is a two-arm dispatch away. H.264 and H.265 now SHARE the machinery instead of duplicating it: derive_arrangement (one coincide/distinct/layered decision table), ring::rebased_offsets (the slices-only rebase — non-VCL NALUs in the decode range hang VCN firmware), session::bind_session_memory, and a parameterised build_frame. A DecodeProfile enum replaces the bare profile idc that images.rs and ring.rs used to take: both codecs' idc types are c_uint, so handing an H.265 idc to the H.264 path COMPILED SILENTLY and built a mismatched profile chain. That is now unrepresentable. The VPS leg is the ledger's real work. The vendored parser attaches a VPS to an SPS only when it saw the NALU, and clients join live streams, so VpsSource is Parsed-or-FromSps and is stored BY VALUE: re-activating a VPS-less SPS is Current (no churn), but the real VPS arriving under the same id is a content change and RECREATES onto it, because Vulkan cannot replace a stored parameter set. Review round 10 (adversarial) confirmed the hardware-proven H.264 path is NOT regressed — derive_arrangement's check order and error identity are byte-for-byte the original, build_frame's call sites still pass the granularity-aligned extent (the 1088-row scar stays shut), and rebased_offsets reproduces the deleted inline loop for every input while moving the sum to u64 so overflow errors instead of wrapping. Also verified: the refs-order contract on every path, the RESULT_STATUS caps gate (each of reset/begin/end individually gated, no pool created when unsupported — recording one on RADV hangs its VCN), pNext lifetimes, and that no panic is reachable on stream input. Its 10 findings are fixed. The two that mattered: - A failed decode stranded a DPB slot. Once plan_to_vk_h265 had mutated the slot map, five later failure paths returned without restoring it, so planner and slot map both believed a picture was resident while no image held it — and every later AU referencing it failed, where H.264 soft-degrades and keeps delivering. Fail-closed is kept (substituting a reference silently is the corruption-hiding this program exists to end) but made RECOVERABLE: a latch flushes the planner to AwaitingIdr and resets the bindings on the next decode, which composes with the client already requesting a keyframe on every decode error. The fix deliberately covers pre-mutation failures too — those strand the picture the other way round and wedge identically. - DecodedVkFrame carried no picture format, so a Main 10 frame would decode correctly and be rendered with 8-bit transfer/range math. It now carries one, stamped from the pool so it is truthful for both decoders by construction. The presenter comment says depth 8 is because only H.264 is WIRED, not a decoder limit. Plus: bind_session_memory freed allocations before the session that may hold them was destroyed (an ordering regression from the extraction, with a SAFETY comment asserting the opposite) — the bind-stage exit now hands them back so Drop destroys first; max_level_idc is codec-tagged rather than an H.264 type carrying H.265 code points; and the decode family's videoCodecOperations is now checked, turning 'create an H.265 session on a device without the extension' from UB into a clean ladder demote. Deferred by design: no HEVC gpu_smoke/gpu_parity yet (its goldens are already in tests/data/test-25fps-h265.nv12.sha256), and no codec dispatch in the client — both later legs. Gates: fmt clean; mac clippy zero warnings, pf-vkdecode 106 + pf-bitstream 69 green; container clippy -D warnings zero for pf-client-core + pf-presenter + pf-vkdecode, tests 69/121/106 green. HARDWARE (.173, after the refactor — review saying the proven path is safe is not the GPU saying it): gpu_parity '250 frames bit-identical to libavcodec software decode' on BOTH the NVIDIA 4090 (610.88, coincide mode) and the AMD iGPU (Adrenalin 25.10.30.02, distinct mode), gpu_smoke green on both. Two independent drivers, both DPB modes, still bit-exact. The smoke trace also shows the new videoCodecOperations capture reading DECODE_H264 | DECODE_H265 | DECODE_AV1 off the real decode family. |