fix(encode): gate multi-slice frames on the client's decoder — the 0.17.0 Chromecast crash

Field report: since 0.17.0 a stream to a Chromecast with Google TV 4K
freezes on the first frame and ~80% of the time crashes + reboots the
DEVICE — with both the Punktfunk app and Moonlight, while an Xbox
Series S is fine. Root cause: LN1 Phase 3 (67b79810) defaulted Linux
direct-NVENC to 4 slices per frame for EVERY session. Amlogic HEVC
decoders wedge on multi-slice AUs — exactly why moonlight-android
requests slicesPerFrame=1 for every hardware decoder (4 only for
software slice-threading) — and our RTSP parser never read the request.
The Phase-3 commit recorded the untested leg ("a live Moonlight re-test
joins the standing owed Moonlight item"); this report is that re-test.

The slicing ceiling now belongs to the CLIENT, threaded as open_video's
new max_slices from both planes:
- GameStream: parse x-nv-video[0].videoEncoderSlicesPerFrame into
  StreamConfig and honor it; absent/out-of-range (pre-auth input) => 1.
- punktfunk/1: new Hello cap VIDEO_CAP_MULTI_SLICE (0x80 — the byte's
  LAST free bit; the next cap needs a second byte + ABI bump).
  SessionPlan.max_slices = 32 with the bit, 1 without, applied to every
  encoder the plan opens so rebuilds can't change the wire shape. The
  desktop session client advertises it (FFmpeg/D3D11VA/Vulkan decode
  stacks are fine); Android/Apple stay off until they can decide
  per-decoder like Moonlight does — the cap is embedder-set decoder
  truth, never OR'd in by the shared pump.
- Linux direct-NVENC clamps its Phase-3 default to the ceiling
  (resolve_slices(codec, 4.min(max_slices))) and logs slices/max_slices
  in the caps-probe line; PUNKTFUNK_NVENC_SLICES stays the explicit
  operator override in both directions. Windows keeps its single-slice
  default untouched.

Also repairs the nvenc_cuda #[ignore] hardware tests: d2c46eaf added
open()'s cursor_blend param without updating them, invisible because CI
never compiles tests with the nvenc feature.

Verified on .21 (RTX 5070 Ti): pf-encode + host check/clippy clean with
nvenc; host unit suite 263/0; rtsp announce tests 7/7 incl. the new
slicesPerFrame coverage; on-hardware smokes — default e2e still 4
chunks/frame, NEW single-slice client-ceiling test clamps + disarms
chunked poll with no env involved, env escape unchanged. rustfmt clean.
Windows leg (prepare_display param) is CI-only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 22:23:42 +02:00
co-authored by Claude Fable 5
parent c04c5be224
commit f87c1e6cec
15 changed files with 235 additions and 21 deletions
+16
View File
@@ -528,6 +528,22 @@
#define VIDEO_CAP_CHACHA20 64
#endif
#if defined(PUNKTFUNK_FEATURE_QUIC)
// [`Hello::video_caps`] bit: the client's decoder accepts **multi-slice access units** — H.264/
// HEVC frames carrying several slice NALs (latency plan §7 LN1: the encoder splits frames so
// sub-frame readback can ship early slices while the tail encodes). Decoder-level, so the
// EMBEDDER sets it from what its decode stack actually handles: the desktop clients' FFmpeg/
// D3D11VA/Vulkan-video decoders are fine, but mobile/TV MediaCodec is per-SoC — Amlogic HEVC
// decoders (Chromecast with Google TV, Fire TV) wedge the whole DEVICE on multi-slice frames
// (the 0.17.0 field regression: the 4-slice Linux default froze streams on first frame and
// watchdog-rebooted the CCwGTV), which is exactly why Moonlight requests 1 slice per frame for
// every hardware decoder. The host defaults to >1 slice ONLY toward a client that sets this
// bit (`PUNKTFUNK_NVENC_SLICES` stays the explicit operator override in both directions);
// every other client gets single-slice frames — the pre-0.17 wire shape. NOTE: this takes the
// video_caps byte's last free bit — the next video cap needs a second byte (ABI bump).
#define VIDEO_CAP_MULTI_SLICE 128
#endif
#if defined(PUNKTFUNK_FEATURE_QUIC)
// [`Welcome::host_caps`] bit: the host applies [`InputKind::GamepadState`]
// (crate::input::InputKind::GamepadState) snapshot events — full per-pad state with a reorder