feat(encode/nvenc-linux): LN1 phase-3 — multi-slice sub-frame encode DEFAULT-ON for Linux direct-NVENC
apple / swift (push) Successful in 1m16s
apple / screenshots (push) Successful in 6m15s
windows-host / package (push) Successful in 9m39s
ci / web (push) Successful in 1m10s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
apple / swift (push) Successful in 1m16s
apple / screenshots (push) Successful in 6m15s
windows-host / package (push) Successful in 9m39s
ci / web (push) Successful in 1m10s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
Gates recorded (nvenc-subframe-slice-output.md Phase 3): loss-harness curve identical streamed-vs-whole; adversarial security review clean; live .21 3-leg A/B — streamed 0 corrupted frames, p99 8527→5363 µs vs sliced-whole, first_slice_us ~500 µs of a ~1250 µs encode reaching the send thread; wire bytes rate-governed under CBR + 1-frame VBV (the ~1-2 % slice-header cost is absorbed by rate control, not added to the wire). GameStream: plane has zero code diff (own RTP packetizer; shares only untouched send_pacing); a live Moonlight re-test joins the standing owed Moonlight item. - resolve_slices(codec, default): env 1..=32 wins (1 = the NEW explicit single-slice escape), else the backend default — 4 on Linux direct-NVENC, 1 elsewhere. resolve_subframe(default_on): PUNKTFUNK_NVENC_SUBFRAME tri-state (0 = never, 1 = force, unset = backend default). - Linux nvenc_cuda: resolves both ONCE in query_caps — the sub-frame default is gated on the GPU's SUBFRAME_READBACK cap so an unsupporting GPU never has enableSubFrameWrite forced into its init params — and feeds the same resolved values to build_config, build_init_params (open AND in-place reconfigure) and the chunked-poll latch. - Windows: env-only as before (async path untouched, byte-identical config for unset env); LowLatencyConfig.slices + the explicit subframe param replace the shared env reads. - Tests: chunked e2e now runs at the DEFAULTS (no knobs); the fallback test becomes the escape test (SLICES=1 disarms; SUBFRAME=0 disarms while 4-slice encode continues on the plain poll path). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -44,13 +44,14 @@
|
||||
//! the retrieve thread the same way it would hang the encode thread today (Linux has no
|
||||
//! event-timeout escape) — no regression, just no new watchdog either.
|
||||
//!
|
||||
//! **Sub-frame chunked poll** (`PUNKTFUNK_NVENC_SLICES=N` + `PUNKTFUNK_NVENC_SUBFRAME=1` — latency
|
||||
//! plan §7 LN1 Phase 1): on a sync depth-1 session, [`Encoder::poll_chunk`] hands the in-flight AU
|
||||
//! out as slice-boundary chunks read through `doNotWait` sub-frame locks while the tail is still
|
||||
//! encoding; one final blocking lock closes the AU (the completion authority — `numSlices` alone
|
||||
//! is not trusted across driver branches). Mutually exclusive with the pipelined retrieve (the
|
||||
//! escalated rebuild drops it); composes with stream-ordered submit (both are sync depth-1
|
||||
//! features).
|
||||
//! **Sub-frame chunked poll** (latency plan §7 LN1 — **default-on since Phase 3**: 4 slices +
|
||||
//! sub-frame readback on every session whose GPU advertises `SUBFRAME_READBACK`; escapes are
|
||||
//! `PUNKTFUNK_NVENC_SLICES=1` and `PUNKTFUNK_NVENC_SUBFRAME=0`): on a sync depth-1 session,
|
||||
//! [`Encoder::poll_chunk`] hands the in-flight AU out as slice-boundary chunks read through
|
||||
//! `doNotWait` sub-frame locks while the tail is still encoding; one final blocking lock closes
|
||||
//! the AU (the completion authority — `numSlices` alone is not trusted across driver branches).
|
||||
//! Mutually exclusive with the pipelined retrieve (the escalated rebuild drops it); composes
|
||||
//! with stream-ordered submit (both are sync depth-1 features).
|
||||
//!
|
||||
//! Needs a real NVIDIA GPU at runtime (session creation fails otherwise); compiles GPU-less and
|
||||
//! starts driver-less (the `.so` resolves at runtime — on an AMD/Intel box [`try_api`] fails cleanly
|
||||
@@ -60,7 +61,7 @@
|
||||
#![deny(clippy::undocumented_unsafe_blocks)]
|
||||
|
||||
use super::nvenc_core::{
|
||||
apply_low_latency_config, build_init_params, codec_guid, slices_env, subframe_requested,
|
||||
apply_low_latency_config, build_init_params, codec_guid, resolve_slices, resolve_subframe,
|
||||
LowLatencyConfig, NvStatusExt, RFI_DPB,
|
||||
};
|
||||
use super::nvenc_status;
|
||||
@@ -559,9 +560,19 @@ pub struct NvencCudaEncoder {
|
||||
/// Stream-ordered submit armed for the live session (sync-retrieve mode only; see
|
||||
/// [`stream_ordered_requested`]). The per-frame gate additionally requires `pending` empty.
|
||||
stream_ordered: bool,
|
||||
/// Slice count the live session was configured with (`PUNKTFUNK_NVENC_SLICES`, latched at
|
||||
/// init; 1 = the preset's single slice). Chunked poll needs ≥ 2 to have boundaries to cut at.
|
||||
/// Slice count the live session was configured with ([`resolve_slices`] — env override,
|
||||
/// else the Linux direct-NVENC default of 4 since Phase 3; 1 = the preset's single slice).
|
||||
/// Chunked poll needs ≥ 2 to have boundaries to cut at. Latched at init, consumed by
|
||||
/// `build_config` (so an in-place reconfigure presents the same slicing).
|
||||
slices: u32,
|
||||
/// `NV_ENC_CAPS_SUPPORT_SUBFRAME_READBACK` from the caps probe — gates the DEFAULT-on
|
||||
/// sub-frame arming (an unsupported GPU must not have `enableSubFrameWrite` forced into its
|
||||
/// init params, which could fail the session open). `PUNKTFUNK_NVENC_SUBFRAME=1` overrides.
|
||||
subframe_cap: bool,
|
||||
/// Sub-frame readback resolved for the live session ([`resolve_subframe`] over
|
||||
/// [`subframe_cap`](Self::subframe_cap)); consumed by every `build_init_params` call so the
|
||||
/// open and the in-place reconfigure present identical init params.
|
||||
subframe_on: bool,
|
||||
/// Sub-frame chunked poll armed for the live session (§7 LN1 Phase 1): multi-slice +
|
||||
/// sub-frame readback configured AND sync retrieve at init. See [`Encoder::poll_chunk`].
|
||||
subframe_chunks: bool,
|
||||
@@ -650,6 +661,8 @@ impl NvencCudaEncoder {
|
||||
io_stream: ptr::null_mut(),
|
||||
stream_ordered: false,
|
||||
slices: 1,
|
||||
subframe_cap: false,
|
||||
subframe_on: false,
|
||||
subframe_chunks: false,
|
||||
chunk: None,
|
||||
})
|
||||
@@ -810,6 +823,15 @@ impl NvencCudaEncoder {
|
||||
}
|
||||
self.rfi_supported = rfi != 0;
|
||||
self.custom_vbv = custom_vbv != 0;
|
||||
self.subframe_cap = subframe != 0;
|
||||
// Phase-3 default-on (nvenc-subframe-slice-output.md): 4 slices + sub-frame readback on
|
||||
// every Linux direct-NVENC session, resolved HERE (before the session opens) so the
|
||||
// config author, the init params and the chunked-poll latch all agree; the caps probe
|
||||
// gates the sub-frame default so a GPU without SUBFRAME_READBACK never has it forced
|
||||
// into its init params. PUNKTFUNK_NVENC_SLICES=1 / PUNKTFUNK_NVENC_SUBFRAME=0 are the
|
||||
// escapes.
|
||||
self.slices = resolve_slices(self.codec, 4);
|
||||
self.subframe_on = resolve_subframe(self.subframe_cap);
|
||||
tracing::info!(
|
||||
rfi = self.rfi_supported,
|
||||
custom_vbv = self.custom_vbv,
|
||||
@@ -923,6 +945,7 @@ impl NvencCudaEncoder {
|
||||
av1_input_depth_minus8: 0,
|
||||
hdr: self.hdr,
|
||||
rfi_supported: self.rfi_supported,
|
||||
slices: self.slices,
|
||||
},
|
||||
);
|
||||
Ok(cfg)
|
||||
@@ -963,6 +986,7 @@ impl NvencCudaEncoder {
|
||||
&mut cfg,
|
||||
split_mode,
|
||||
false,
|
||||
self.subframe_on,
|
||||
);
|
||||
|
||||
match (api().initialize_encoder)(enc, &mut init).nv_ok() {
|
||||
@@ -1186,14 +1210,14 @@ impl NvencCudaEncoder {
|
||||
}
|
||||
}
|
||||
}
|
||||
// Sub-frame chunked poll (§7 LN1 Phase 1): armed iff this session was CONFIGURED
|
||||
// multi-slice + sub-frame readback (apply_low_latency_config / build_init_params
|
||||
// consume the same env parses, so the latch can't disagree with the session config)
|
||||
// and the retrieve is sync — chunked poll is a depth-1 sync feature; a pipelined
|
||||
// session's non-blocking poll owns the bitstream from the retrieve thread instead.
|
||||
self.slices = slices_env(self.codec).unwrap_or(1);
|
||||
self.subframe_chunks =
|
||||
self.slices >= 2 && subframe_requested() && self.async_rt.is_none();
|
||||
// Sub-frame chunked poll (§7 LN1 Phase 1; default-on since Phase 3): armed iff this
|
||||
// session was CONFIGURED multi-slice + sub-frame readback (`self.slices` /
|
||||
// `self.subframe_on` were resolved once in `query_caps` and consumed by
|
||||
// `build_config` / `build_init_params`, so the latch can't disagree with the session
|
||||
// config) and the retrieve is sync — chunked poll is a depth-1 sync feature; a
|
||||
// pipelined session's non-blocking poll owns the bitstream from the retrieve thread
|
||||
// instead (the sub-frame write itself stays armed there; it's harmless).
|
||||
self.subframe_chunks = self.slices >= 2 && self.subframe_on && self.async_rt.is_none();
|
||||
if self.subframe_chunks {
|
||||
tracing::info!(
|
||||
slices = self.slices,
|
||||
@@ -1762,7 +1786,7 @@ impl Encoder for NvencCudaEncoder {
|
||||
// (version set, doNotWait) and `offsets` are live stack locals across the
|
||||
// synchronous call; `reportSliceOffsets` was armed at init so the driver may write
|
||||
// up to `numSlices` ≤ 32 offsets (`sliceModeData` is clamped 2..=32 by
|
||||
// `slices_env`). On a successful sub-frame lock `bitstreamBufferPtr` holds
|
||||
// `resolve_slices`). On a successful sub-frame lock `bitstreamBufferPtr` holds
|
||||
// `bitstreamSizeInBytes` readable bytes of COMPLETED slices (enableSubFrameWrite
|
||||
// publishes them mid-encode; proven by the on-hw probe) valid until the matching
|
||||
// unlock; the emitted range is copied out BEFORE the unlock. Every successful lock
|
||||
@@ -1934,6 +1958,7 @@ impl Encoder for NvencCudaEncoder {
|
||||
&mut cfg,
|
||||
self.split_mode,
|
||||
false,
|
||||
self.subframe_on,
|
||||
),
|
||||
..Default::default()
|
||||
};
|
||||
@@ -2585,10 +2610,11 @@ mod tests {
|
||||
d.starts_with(&[0, 0, 0, 1]) || d.starts_with(&[0, 0, 1])
|
||||
}
|
||||
|
||||
/// ON-HARDWARE (RTX box `.21`): LN1 Phase 1 — the chunked poll end to end. With 4 slices +
|
||||
/// sub-frame readback armed, `poll_chunk` must (a) report the mode armed, (b) hand every AU
|
||||
/// out as chunks whose first chunk opens the AU with the right metadata and whose `last`
|
||||
/// closes it, (c) cut every chunk at an Annex-B start code, and (d) reassemble byte-
|
||||
/// ON-HARDWARE (RTX box `.21`): LN1 Phase 1 — the chunked poll end to end, at the Phase-3
|
||||
/// DEFAULTS (no env knobs: 4 slices + sub-frame readback arm on their own on a
|
||||
/// SUBFRAME_READBACK-capable GPU). `poll_chunk` must (a) report the mode armed, (b) hand
|
||||
/// every AU out as chunks whose first chunk opens the AU with the right metadata and whose
|
||||
/// `last` closes it, (c) cut every chunk at an Annex-B start code, and (d) reassemble byte-
|
||||
/// identically to the finishing blocking lock's AU (enforced by the debug-build shadow check
|
||||
/// inside `poll_chunk` — a mismatch errors the test). At least one frame must actually chunk
|
||||
/// (>1 chunk) — the 5070 Ti probe shows every frame does. Run single-threaded (env vars are
|
||||
@@ -2598,16 +2624,9 @@ mod tests {
|
||||
fn nvenc_cuda_chunked_poll_end_to_end() {
|
||||
const W: u32 = 1920;
|
||||
const H: u32 = 1080;
|
||||
struct EnvGuard;
|
||||
impl Drop for EnvGuard {
|
||||
fn drop(&mut self) {
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SLICES");
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SUBFRAME");
|
||||
}
|
||||
}
|
||||
std::env::set_var("PUNKTFUNK_NVENC_SLICES", "4");
|
||||
std::env::set_var("PUNKTFUNK_NVENC_SUBFRAME", "1");
|
||||
let _guard = EnvGuard;
|
||||
// Defaults under test — make sure another test's knobs aren't leaking in.
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SLICES");
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SUBFRAME");
|
||||
|
||||
pf_zerocopy::cuda::make_current().expect("shared CUDA context current");
|
||||
let mut enc = NvencCudaEncoder::open(
|
||||
@@ -2690,7 +2709,8 @@ mod tests {
|
||||
assert!(!au.data.is_empty());
|
||||
}
|
||||
|
||||
/// ON-HARDWARE (RTX box `.21`): without the slice/subframe knobs, `poll_chunk` must degrade
|
||||
/// ON-HARDWARE (RTX box `.21`): the Phase-3 default-on ESCAPES — `PUNKTFUNK_NVENC_SLICES=1`
|
||||
/// must fully disarm chunked poll (and `poll_chunk` degrades
|
||||
/// to exactly one self-closing whole-AU chunk (the default-path contract every non-chunked
|
||||
/// session shares). Run with `--test-threads=1` (env vars are process-global).
|
||||
#[test]
|
||||
@@ -2698,16 +2718,25 @@ mod tests {
|
||||
fn nvenc_cuda_chunked_poll_fallback_whole_au() {
|
||||
const W: u32 = 1280;
|
||||
const H: u32 = 720;
|
||||
// Belt-and-braces: another test in this process may have set the knobs (env is global).
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SLICES");
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SUBFRAME");
|
||||
struct EnvGuard;
|
||||
impl Drop for EnvGuard {
|
||||
fn drop(&mut self) {
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SLICES");
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SUBFRAME");
|
||||
}
|
||||
}
|
||||
let _guard = EnvGuard;
|
||||
pf_zerocopy::cuda::make_current().expect("shared CUDA context current");
|
||||
|
||||
// Escape 1: explicit single slice — no boundaries to cut, chunked poll disarmed.
|
||||
std::env::set_var("PUNKTFUNK_NVENC_SLICES", "1");
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SUBFRAME");
|
||||
let mut enc = open_h265();
|
||||
let frame = nv12_frame(W, H, 0);
|
||||
enc.submit_indexed(&frame, 0).expect("submit");
|
||||
assert!(
|
||||
!enc.supports_chunked_poll(),
|
||||
"no knobs → chunked poll must not arm"
|
||||
"PUNKTFUNK_NVENC_SLICES=1 → chunked poll must not arm"
|
||||
);
|
||||
let c = enc
|
||||
.poll_chunk()
|
||||
@@ -2720,5 +2749,20 @@ mod tests {
|
||||
enc.poll_chunk().expect("poll_chunk").is_none(),
|
||||
"nothing in flight → None"
|
||||
);
|
||||
drop(enc);
|
||||
|
||||
// Escape 2: sub-frame readback vetoed — slices stay (default 4) but chunked poll
|
||||
// disarms and the plain poll path carries the session.
|
||||
std::env::remove_var("PUNKTFUNK_NVENC_SLICES");
|
||||
std::env::set_var("PUNKTFUNK_NVENC_SUBFRAME", "0");
|
||||
let mut enc = open_h265();
|
||||
let frame = nv12_frame(W, H, 0);
|
||||
enc.submit_indexed(&frame, 0).expect("submit");
|
||||
assert!(
|
||||
!enc.supports_chunked_poll(),
|
||||
"PUNKTFUNK_NVENC_SUBFRAME=0 → chunked poll must not arm"
|
||||
);
|
||||
let au = enc.poll().expect("poll").expect("AU");
|
||||
assert!(au.keyframe && !au.data.is_empty());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -35,26 +35,36 @@ pub(super) fn codec_guid(codec: Codec) -> nv::GUID {
|
||||
}
|
||||
}
|
||||
|
||||
/// `PUNKTFUNK_NVENC_SLICES` — the per-frame slice count (2..=32) for multi-slice encode (latency
|
||||
/// plan §7 LN1), H.264/HEVC only (AV1 partitions via tiles, so the knob never applies there).
|
||||
/// `None` = the preset's default single slice. ONE parse shared by the config author
|
||||
/// ([`apply_low_latency_config`]) and the Linux backend's chunked-poll arming, so the two can
|
||||
/// Resolved per-frame slice count for a session (latency plan §7 LN1, Phase 3): the
|
||||
/// `PUNKTFUNK_NVENC_SLICES` env override wins (1..=32; **1 = the explicit single-slice
|
||||
/// escape**, needed now that a backend can default higher), else the backend's
|
||||
/// `default_slices` — 4 on Linux direct-NVENC since the Phase-3 default-on, 1 everywhere else
|
||||
/// (the Windows async path is deliberately untouched). H.264/HEVC only (AV1 partitions via
|
||||
/// tiles). ONE parse shared by the config author ([`apply_low_latency_config`] via
|
||||
/// [`LowLatencyConfig::slices`]) and the Linux backend's chunked-poll arming, so the two can
|
||||
/// never disagree about whether a session is multi-slice.
|
||||
pub(super) fn slices_env(codec: Codec) -> Option<u32> {
|
||||
pub(super) fn resolve_slices(codec: Codec, default_slices: u32) -> u32 {
|
||||
if !matches!(codec, Codec::H264 | Codec::H265) {
|
||||
return None;
|
||||
return 1;
|
||||
}
|
||||
std::env::var("PUNKTFUNK_NVENC_SLICES")
|
||||
.ok()
|
||||
.and_then(|s| s.parse::<u32>().ok())
|
||||
.filter(|n| (2..=32).contains(n))
|
||||
.filter(|n| (1..=32).contains(n))
|
||||
.unwrap_or(default_slices)
|
||||
}
|
||||
|
||||
/// `PUNKTFUNK_NVENC_SUBFRAME=1` — arm sub-frame readback (`enableSubFrameWrite` +
|
||||
/// `reportSliceOffsets`; sync sessions only, see [`build_init_params`]). Shared for the same
|
||||
/// reason as [`slices_env`].
|
||||
pub(super) fn subframe_requested() -> bool {
|
||||
std::env::var("PUNKTFUNK_NVENC_SUBFRAME").as_deref() == Ok("1")
|
||||
/// Resolved sub-frame readback (`enableSubFrameWrite` + `reportSliceOffsets`; sync sessions
|
||||
/// only, see [`build_init_params`]): `PUNKTFUNK_NVENC_SUBFRAME` tri-state — `0` = never (the
|
||||
/// default-on escape), `1` = force (even where the caps probe says unsupported — an operator
|
||||
/// explicitly testing), unset = the backend's `default_on` (Linux direct-NVENC passes its
|
||||
/// SUBFRAME_READBACK caps-probe result since Phase 3; Windows passes `false`).
|
||||
pub(super) fn resolve_subframe(default_on: bool) -> bool {
|
||||
match std::env::var("PUNKTFUNK_NVENC_SUBFRAME").as_deref() {
|
||||
Ok("0") => false,
|
||||
Ok("1") => true,
|
||||
_ => default_on,
|
||||
}
|
||||
}
|
||||
|
||||
/// Reference-frame DPB depth when RFI is supported (Apollo uses 5). A deeper DPB lets an invalidated
|
||||
@@ -88,6 +98,9 @@ pub(super) struct LowLatencyConfig {
|
||||
pub hdr: bool,
|
||||
/// This GPU supports reference-frame invalidation (a deeper DPB for graceful loss recovery).
|
||||
pub rfi_supported: bool,
|
||||
/// Resolved per-frame slice count ([`resolve_slices`] — env override, else the backend
|
||||
/// default). ≤ 1 leaves the preset's single slice untouched.
|
||||
pub slices: u32,
|
||||
}
|
||||
|
||||
/// Author the shared `NV_ENC_INITIALIZE_PARAMS` (P1/ULL preset, PTD, the session dimensions/rate)
|
||||
@@ -104,6 +117,7 @@ pub(super) fn build_init_params(
|
||||
cfg: &mut nv::NV_ENC_CONFIG,
|
||||
split_mode: u32,
|
||||
enable_async: bool,
|
||||
subframe: bool,
|
||||
) -> nv::NV_ENC_INITIALIZE_PARAMS {
|
||||
let mut init = nv::NV_ENC_INITIALIZE_PARAMS {
|
||||
version: nv::NV_ENC_INITIALIZE_PARAMS_VER,
|
||||
@@ -123,12 +137,13 @@ pub(super) fn build_init_params(
|
||||
};
|
||||
// splitEncodeMode is a C bitfield — set via the generated accessor, not a struct field.
|
||||
init.set_splitEncodeMode(split_mode);
|
||||
// Sub-frame readback (latency plan §7 LN1 groundwork — EXPERIMENTAL, default off): the driver
|
||||
// writes each slice into the output buffer as it completes and reports per-slice offsets, so a
|
||||
// sync-mode consumer can read slices out while the frame is still encoding. Pair with
|
||||
// `PUNKTFUNK_NVENC_SLICES` (a single-slice frame yields nothing to read early).
|
||||
// `reportSliceOffsets` requires `enableEncodeAsync = 0`, so async (Windows) sessions never arm.
|
||||
if !enable_async && subframe_requested() {
|
||||
// Sub-frame readback (latency plan §7 LN1; default-on for Linux direct-NVENC since Phase 3 —
|
||||
// the caller resolves `subframe` via [`resolve_subframe`] + its caps probe): the driver
|
||||
// writes each slice into the output buffer as it completes and reports per-slice offsets, so
|
||||
// a sync-mode consumer can read slices out while the frame is still encoding. Pair with
|
||||
// multi-slice (a single-slice frame yields nothing to read early). `reportSliceOffsets`
|
||||
// requires `enableEncodeAsync = 0`, so async (Windows) sessions never arm.
|
||||
if !enable_async && subframe {
|
||||
init.set_enableSubFrameWrite(1);
|
||||
init.set_reportSliceOffsets(1);
|
||||
}
|
||||
@@ -180,12 +195,13 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
|
||||
Codec::PyroWave => unreachable!("PyroWave never opens the direct-NVENC backend"),
|
||||
}
|
||||
|
||||
// Multi-slice frames (latency plan §7 LN1 groundwork — EXPERIMENTAL, default off = the preset's
|
||||
// single slice): `PUNKTFUNK_NVENC_SLICES=N` (2..=32) splits every frame into N slices
|
||||
// Multi-slice frames (latency plan §7 LN1): `c.slices` splits every frame into N slices
|
||||
// (sliceMode 3 = "N slices per frame"), the unit sub-frame readback ships early and loss
|
||||
// concealment can discard independently. Costs ~1-2 % bitrate in slice headers. H.264/HEVC
|
||||
// only — AV1 partitions via tiles, not slices.
|
||||
if let Some(n) = slices_env(c.codec) {
|
||||
// only — AV1 partitions via tiles, not slices (the resolver already returns 1 there).
|
||||
// Default 4 on Linux direct-NVENC (Phase 3), env-only elsewhere; ≤ 1 keeps the preset's
|
||||
// single slice.
|
||||
if let Some(n) = Some(c.slices).filter(|n| *n >= 2) {
|
||||
match c.codec {
|
||||
Codec::H264 => {
|
||||
cfg.encodeCodecConfig.h264Config.sliceMode = 3;
|
||||
|
||||
@@ -37,7 +37,8 @@
|
||||
#![deny(clippy::undocumented_unsafe_blocks)]
|
||||
|
||||
use super::nvenc_core::{
|
||||
apply_low_latency_config, build_init_params, codec_guid, LowLatencyConfig, NvStatusExt, RFI_DPB,
|
||||
apply_low_latency_config, build_init_params, codec_guid, resolve_slices, resolve_subframe,
|
||||
LowLatencyConfig, NvStatusExt, RFI_DPB,
|
||||
};
|
||||
use super::nvenc_status;
|
||||
use super::{ChromaFormat, Codec, EncodedFrame, Encoder, EncoderCaps};
|
||||
@@ -744,6 +745,10 @@ impl NvencD3d11Encoder {
|
||||
av1_input_depth_minus8: if ten_bit_in { 2 } else { 0 },
|
||||
hdr: self.hdr,
|
||||
rfi_supported: self.rfi_supported,
|
||||
// Env-only on Windows (default single slice): the Phase-3 default-on is a
|
||||
// Linux-direct-NVENC decision — the Windows async path stays untouched, and a
|
||||
// Windows operator opting in must choose slices+sync over async retrieve.
|
||||
slices: resolve_slices(self.codec, 1),
|
||||
},
|
||||
);
|
||||
Ok(cfg)
|
||||
@@ -791,6 +796,9 @@ impl NvencD3d11Encoder {
|
||||
&mut cfg,
|
||||
split_mode,
|
||||
enable_async,
|
||||
// Windows: env opt-in only ("1"), never a default — and build_init_params
|
||||
// additionally refuses it on an async session.
|
||||
resolve_subframe(false),
|
||||
);
|
||||
|
||||
match (api().initialize_encoder)(enc, &mut init).nv_ok() {
|
||||
@@ -1568,6 +1576,7 @@ impl Encoder for NvencD3d11Encoder {
|
||||
&mut cfg,
|
||||
self.split_mode,
|
||||
self.session_async,
|
||||
resolve_subframe(false),
|
||||
),
|
||||
..Default::default()
|
||||
};
|
||||
|
||||
Reference in New Issue
Block a user