fix(encode): gate multi-slice frames on the client's decoder — the 0.17.0 Chromecast crash
ci / docs-site (push) Successful in 1m9s
apple / swift (push) Successful in 1m15s
ci / web (push) Successful in 3m16s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 7s
deb / build-publish-client-arm64 (push) Successful in 2m13s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 5s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 7s
ci / rust-arm64 (push) Successful in 3m34s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 5s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 13s
deb / build-publish-host (push) Successful in 4m20s
docker / builders-arm64cross (push) Successful in 6s
docker / deploy-docs (push) Successful in 28s
deb / build-publish (push) Successful in 6m17s
arch / build-publish (push) Failing after 7m58s
flatpak / build-publish (push) Successful in 9m6s
ci / rust (push) Successful in 13m37s
release / apple (push) Successful in 16m32s
windows-host / package (push) Successful in 18m21s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 32s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m21s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m44s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m24s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m27s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 2m56s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 3m54s
apple / screenshots (push) Successful in 20m45s
android / android (push) Successful in 5m0s

Field report: since 0.17.0 a stream to a Chromecast with Google TV 4K
freezes on the first frame and ~80% of the time crashes + reboots the
DEVICE — with both the Punktfunk app and Moonlight, while an Xbox
Series S is fine. Root cause: LN1 Phase 3 (67b79810) defaulted Linux
direct-NVENC to 4 slices per frame for EVERY session. Amlogic HEVC
decoders wedge on multi-slice AUs — exactly why moonlight-android
requests slicesPerFrame=1 for every hardware decoder (4 only for
software slice-threading) — and our RTSP parser never read the request.
The Phase-3 commit recorded the untested leg ("a live Moonlight re-test
joins the standing owed Moonlight item"); this report is that re-test.

The slicing ceiling now belongs to the CLIENT, threaded as open_video's
new max_slices from both planes:
- GameStream: parse x-nv-video[0].videoEncoderSlicesPerFrame into
  StreamConfig and honor it; absent/out-of-range (pre-auth input) => 1.
- punktfunk/1: new Hello cap VIDEO_CAP_MULTI_SLICE (0x80 — the byte's
  LAST free bit; the next cap needs a second byte + ABI bump).
  SessionPlan.max_slices = 32 with the bit, 1 without, applied to every
  encoder the plan opens so rebuilds can't change the wire shape. The
  desktop session client advertises it (FFmpeg/D3D11VA/Vulkan decode
  stacks are fine); Android/Apple stay off until they can decide
  per-decoder like Moonlight does — the cap is embedder-set decoder
  truth, never OR'd in by the shared pump.
- Linux direct-NVENC clamps its Phase-3 default to the ceiling
  (resolve_slices(codec, 4.min(max_slices))) and logs slices/max_slices
  in the caps-probe line; PUNKTFUNK_NVENC_SLICES stays the explicit
  operator override in both directions. Windows keeps its single-slice
  default untouched.

Also repairs the nvenc_cuda #[ignore] hardware tests: d2c46eaf added
open()'s cursor_blend param without updating them, invisible because CI
never compiles tests with the nvenc feature.

Verified on .21 (RTX 5070 Ti): pf-encode + host check/clippy clean with
nvenc; host unit suite 263/0; rtsp announce tests 7/7 incl. the new
slicesPerFrame coverage; on-hardware smokes — default e2e still 4
chunks/frame, NEW single-slice client-ceiling test clamps + disarms
chunked poll with no env involved, env escape unchanged. rustfmt clean.
Windows leg (prepare_display param) is CI-only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-30 22:23:42 +02:00
co-authored by Claude Fable 5
parent c04c5be224
commit f87c1e6cec
15 changed files with 235 additions and 21 deletions
+77 -4
View File
@@ -795,10 +795,17 @@ pub struct NvencCudaEncoder {
/// [`stream_ordered_requested`]). The per-frame gate additionally requires `pending` empty.
stream_ordered: bool,
/// Slice count the live session was configured with ([`resolve_slices`] — env override,
/// else the Linux direct-NVENC default of 4 since Phase 3; 1 = the preset's single slice).
/// else the Linux direct-NVENC default of 4 since Phase 3 clamped to
/// [`max_slices`](Self::max_slices); 1 = the preset's single slice).
/// Chunked poll needs ≥ 2 to have boundaries to cut at. Latched at init, consumed by
/// `build_config` (so an in-place reconfigure presents the same slicing).
slices: u32,
/// Ceiling on the per-frame slice count the session's CLIENT decoder accepts (from
/// negotiation: `VIDEO_CAP_MULTI_SLICE`, or GameStream's `videoEncoderSlicesPerFrame`).
/// 1 = single-slice only — the safe shape toward decoders that never asked (Amlogic TV
/// SoCs wedge on multi-slice AUs). Clamps the Phase-3 default; the explicit
/// `PUNKTFUNK_NVENC_SLICES` env override still wins in both directions.
max_slices: u32,
/// `NV_ENC_CAPS_SUPPORT_SUBFRAME_READBACK` from the caps probe — gates the DEFAULT-on
/// sub-frame arming (an unsupported GPU must not have `enableSubFrameWrite` forced into its
/// init params, which could fail the session open). `PUNKTFUNK_NVENC_SUBFRAME=1` overrides.
@@ -847,6 +854,7 @@ impl NvencCudaEncoder {
bit_depth: u8,
chroma: ChromaFormat,
cursor_blend: bool,
max_slices: u32,
) -> Result<Self> {
// The runtime `.so` load is the real "is NVENC possible here" gate: fail the open with a
// clear reason instead of an opaque session error on the first frame.
@@ -895,6 +903,8 @@ impl NvencCudaEncoder {
io_stream: ptr::null_mut(),
stream_ordered: false,
slices: 1,
// A zero from a misbehaving caller must not zero the resolver's default arithmetic.
max_slices: max_slices.max(1),
subframe_cap: false,
subframe_on: false,
subframe_forced: false,
@@ -1094,9 +1104,14 @@ impl NvencCudaEncoder {
// every Linux direct-NVENC session, resolved HERE (before the session opens) so the
// config author, the init params and the chunked-poll latch all agree; the caps probe
// gates the sub-frame default so a GPU without SUBFRAME_READBACK never has it forced
// into its init params. PUNKTFUNK_NVENC_SLICES=1 / PUNKTFUNK_NVENC_SUBFRAME=0 are the
// escapes.
self.slices = resolve_slices(self.codec, 4);
// into its init params. The Phase-3 default is CLAMPED to the session's negotiated
// `max_slices` — the client-decoder ceiling (`VIDEO_CAP_MULTI_SLICE`, or GameStream's
// `videoEncoderSlicesPerFrame`): a client that never asked for multi-slice AUs gets
// single-slice frames, because TV-SoC decoders (Amlogic — Chromecast with Google TV)
// wedge the whole device on frames carrying several slice NALs (the 0.17.0 field
// regression). PUNKTFUNK_NVENC_SLICES / PUNKTFUNK_NVENC_SUBFRAME stay the explicit
// operator overrides in both directions.
self.slices = resolve_slices(self.codec, 4.min(self.max_slices));
self.subframe_on = resolve_subframe(self.subframe_cap);
self.subframe_forced = subframe_env_forced();
tracing::info!(
@@ -1105,6 +1120,8 @@ impl NvencCudaEncoder {
yuv444 = self.yuv444_supported,
subframe_readback = subframe != 0,
dynamic_slice = dyn_slice != 0,
slices = self.slices,
max_slices = self.max_slices,
max = %format!("{wmax}x{hmax}"),
"NVENC (Linux direct) capabilities probed"
);
@@ -2559,6 +2576,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA session");
@@ -2657,6 +2675,7 @@ mod tests {
10,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA session");
@@ -2737,6 +2756,7 @@ mod tests {
10,
ChromaFormat::Yuv420,
true, // cursor_blend: bring up the Vulkan slot ring + the 10-bit blend
4,
)
.expect("open NVENC CUDA session");
let cursor = |serial: u64, x: i32, y: i32| pf_frame::CursorOverlay {
@@ -2802,6 +2822,7 @@ mod tests {
8,
ChromaFormat::Yuv444,
false,
4,
)
.expect("open NVENC CUDA 4:4:4 session");
@@ -2848,6 +2869,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA session");
@@ -2906,6 +2928,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
) else {
eprintln!(
"skipping rfi_declines_impossible_ranges: NVENC unavailable (no NVIDIA driver)"
@@ -2933,6 +2956,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA encoder")
}
@@ -2968,6 +2992,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open");
for f in 0..4u32 {
@@ -3106,6 +3131,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA session");
let frame = nv12_frame(W, H, 0);
@@ -3149,6 +3175,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
true, // cursor_blend: bring up the Vulkan slot ring + blend
4,
)
.expect("open NVENC CUDA session");
let cursor = |serial: u64, x: i32, y: i32| pf_frame::CursorOverlay {
@@ -3227,6 +3254,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA session");
// Steady sync frames first (stream-ordered mode).
@@ -3307,6 +3335,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA session");
@@ -3413,6 +3442,7 @@ mod tests {
8,
ChromaFormat::Yuv420,
false,
4,
)
.expect("open NVENC CUDA session");
@@ -3483,6 +3513,49 @@ mod tests {
assert!(!au.data.is_empty());
}
/// ON-HARDWARE (RTX box `.21`): a session whose CLIENT ceiling is 1 slice (`max_slices` from
/// negotiation — no cap bit / Moonlight `videoEncoderSlicesPerFrame:1`, the Chromecast field
/// regression's fix) must encode single-slice frames with chunked poll disarmed, with NO env
/// knobs involved. Run with `--test-threads=1` (env vars are process-global).
#[test]
#[ignore = "requires an NVIDIA GPU + driver — run manually on the RTX box (.21)"]
fn nvenc_cuda_single_slice_client_ceiling() {
const W: u32 = 1920;
const H: u32 = 1080;
// The ceiling under test is the negotiated one, not the operator override.
std::env::remove_var("PUNKTFUNK_NVENC_SLICES");
std::env::remove_var("PUNKTFUNK_NVENC_SUBFRAME");
pf_zerocopy::cuda::make_current().expect("shared CUDA context current");
let mut enc = NvencCudaEncoder::open(
Codec::H265,
PixelFormat::Nv12,
W,
H,
60,
20_000_000,
true,
8,
ChromaFormat::Yuv420,
false,
1, // the client never advertised multi-slice tolerance
)
.expect("open NVENC CUDA session");
for i in 0..4u32 {
let frame = nv12_frame(W, H, i);
enc.submit_indexed(&frame, i).expect("submit");
assert_eq!(
enc.slices, 1,
"a 1-slice client ceiling must clamp the Phase-3 default"
);
assert!(
!enc.supports_chunked_poll(),
"single-slice sessions have no boundaries — chunked poll must stay disarmed"
);
let au = enc.poll().expect("poll").expect("one AU per sync frame");
assert!(!au.data.is_empty(), "frame {i} produced an empty AU");
}
}
/// ON-HARDWARE (RTX box `.21`): the Phase-3 default-on ESCAPES — `PUNKTFUNK_NVENC_SLICES=1`
/// must fully disarm chunked poll (and `poll_chunk` degrades
/// to exactly one self-closing whole-AU chunk (the default-path contract every non-chunked