refactor(pf-encode): WP4 — one split policy, shared with the libav path

The libav NVENC path carried its own inline copy of the split decision and had
already drifted from the direct-SDK selector: it hard-coded a 2-way split
regardless of engine count, and had no depth rule at all. That is the drift the
shared resolver was extracted to prevent, and the copy quietly reintroduced it.

Routing it through `resolve_split_mode` needed the policy to MOVE. `nvenc_core`
is gated on `feature = "nvenc"`, but the libav path is precisely the build where
that feature is OFF (`PUNKTFUNK_NVENC_DIRECT=0`, and the featureless packages --
the packaging gap this project has been bitten by before). So
resolve_split_mode / max_forced_split_mode / clamp_to_engines, plus a new
`forced_split_width`, now live in `codec.rs`, which is always compiled and
already owned SPLIT_FORCE_PIXEL_RATE.

That means the NV_ENC_SPLIT_ENCODE_MODE values had to be hand-written as plain
constants, since the SDK enum does not exist without the feature. They are
therefore pinned: `nvenc_split_constants_match_the_sdk` (feature-gated, the only
place both are visible at once) asserts all five against the real enum, so the
copies cannot rot.

⚠ Only the FORCED outcomes are actionable on the libav side -- libavcodec's
`split_encode_mode` AVOption is its own vocabulary and our DISABLE is the NVENC
enum's 15, which would be meaningless there. DISABLE/AUTO both map to "leave the
option unset", which is exactly today's behaviour (unset = the driver's auto).
`engines = 0` ("not probed") maps to 2-way, preserving what that site always did;
a 3-NVENC part gets the wider split only on the direct-SDK path, which is the one
that actually probes.

⚠⚠ VERIFICATION GAP: .133 went down mid-change (no ping), so the WINDOWS leg is
UNVERIFIED. This matters more than usual -- the Windows backend imported
resolve_split_mode from nvenc_core and that import had to move too, which a grep
caught rather than a compiler. Re-run before trusting it:
  cargo clippy -p pf-encode --features nvenc --all-targets -- -D warnings

Verified .21: clippy -D warnings clean BOTH with and without the nvenc feature
(the featureless build is the whole point of the move) and with
nvenc,vulkan-encode; 65 unit tests incl. the new constant-parity test; 25/25
NVENC on-hardware; punktfunk-host clippy clean. fmt clean.
This commit is contained in:
2026-08-07 08:16:35 +02:00
parent 1062aa780f
commit 01294e3a53
5 changed files with 222 additions and 150 deletions
+156 -1
View File
@@ -518,7 +518,7 @@ impl Codec {
}
/// Pixel rate (luma samples/s) at or above which NVENC split-frame encoding is FORCED 2-way —
/// one number shared by the direct-SDK selector (`nvenc_core::resolve_split_mode`) and the libav
/// one number shared by the direct-SDK selector ([`resolve_split_mode`]) and the libav
/// `split_encode_mode` option author (`linux::NvencEncoder`), so the two paths can never disagree
/// about which modes split. A single NVENC engine tops out ~1 Gpix/s on HEVC, and AUTO doesn't
/// engage below ~2112 px height, so the sessions that need the second engine must be forced. Set
@@ -528,6 +528,161 @@ impl Codec {
/// comfortably single-engine) on AUTO.
pub const SPLIT_FORCE_PIXEL_RATE: u64 = 950_000_000;
/// The `NV_ENC_SPLIT_ENCODE_MODE` values, as plain constants.
///
/// They live HERE, not in `nvenc_core`, because the split policy below has to be shared with the
/// **libav** NVENC path — which compiles with the `nvenc` feature OFF (that is the whole
/// `PUNKTFUNK_NVENC_DIRECT=0` / featureless-package build), where the SDK enum does not exist.
/// One policy, no drift, was the point of extracting it; gating it behind the feature would have
/// left the libav copy free to diverge again, which is exactly what it had already done.
///
/// `nvenc_split_constants_match_the_sdk` (feature-gated) pins these against the real enum, so the
/// hand-written values cannot rot.
pub(crate) const SPLIT_AUTO: u32 = 0;
pub(crate) const SPLIT_AUTO_FORCED: u32 = 1;
pub(crate) const SPLIT_TWO_FORCED: u32 = 2;
pub(crate) const SPLIT_THREE_FORCED: u32 = 3;
pub(crate) const SPLIT_DISABLE: u32 = 15;
/// Resolved NVENC split-frame encode mode for a session — ONE selector shared by the Windows and
/// Linux direct-SDK backends (they had drifted into byte-identical duplicates, one of which
/// logged and one didn't). Precedence:
/// 1. `PUNKTFUNK_SPLIT_ENCODE` = `0`/`disable` | `1`/`auto` (AUTO_FORCED) | `2` | `3` — operator
/// override, always wins, except that `2`/`3` are clamped to the GPU's real engine count (see
/// [`clamp_to_engines`]; the driver honours an over-ask and silently encodes narrower).
/// 2. Pixel rate ≥ [`SPLIT_FORCE_PIXEL_RATE`] → force the WIDEST split the GPU can deliver
/// ([`max_forced_split_mode`]), not a hard-coded 2 (AUTO never engages below ~2112 px height,
/// so 4K120 must be forced onto the other engines; and a 3-NVENC part left at 2-way wastes a
/// third of its encode silicon).
/// 3. **HEVC** Main10 below that bar → DISABLE: 2-way split measured SLOWER on Ada for Main10 — at
/// 5120×1440@240 forced-2 took 7.6 ms/frame (~131 fps) vs 2.8 ms (~357 fps) single-engine, the
/// "broken animations in HDR" cap. ⚠ This rule used to sit ABOVE the pixel-rate arm and take no
/// codec, so it (a) vetoed 10-bit **4K120** — the very case the pixel-rate arm exists for — and
/// (b) applied an HEVC-on-Ada result to **AV1 10-bit**, which has no such measurement. Both
/// fixed; what remains is a conservative default in the regime where a second engine buys
/// nothing anyway.
/// ⚠⚠ **UNVALIDATED CONSEQUENCE:** 5120×1440@240 Main10 (1.77 Gpix/s) now clears the pixel-rate
/// bar and WILL be forced to split — i.e. the exact configuration that measurement came from
/// flips behaviour. That is deliberate (the datapoint is one sample, at low bits/frame, and the
/// bits/frame hypothesis predicts it should not generalise) but it is **the first thing to
/// re-measure on Ada**; `PUNKTFUNK_SPLIT_ENCODE=0` is the escape if it regresses.
/// 4. Else AUTO — ⚠ whose behaviour is **conditional on sub-frame**, measured on `.21` at 4K:
/// - sub-frame **ON** (the fleet default): AUTO **does not split** — 5023/5157 µs against
/// DISABLE's 4979/5000. Split and sub-frame are mutually unsupported for HEVC, so the driver
/// resolves AUTO to no-split and this arm silently means DISABLE.
/// - sub-frame **OFF**: AUTO **does split** — 2401/2352 µs against TWO_FORCED's 2319/2378.
///
/// So AUTO is NOT dead in general and must not be retired: doing so would lose a real split on
/// every sub-frame-off session. It is dead only in the sub-frame-on combination, which
/// [`resolve_split_subframe`] logs rather than silently accepting.
///
/// The caller still owns the rejection fallback (retry split-disabled) — a codec/config that
/// rejects the chosen mode downgrades at open, not here.
///
/// `engines` is the GPU's `NV_ENC_CAPS_NUM_ENCODER_ENGINES`; pass `0` when it could not be probed
/// (treated as "unknown", which keeps the pre-probe behaviour of assuming a second engine exists
/// and letting the open-time rejection fallback sort it out).
pub(crate) fn resolve_split_mode(
codec: Codec,
bit_depth: u8,
pixel_rate: u64,
engines: u32,
) -> u32 {
let hw_max = max_forced_split_mode(engines);
let mode = match std::env::var("PUNKTFUNK_SPLIT_ENCODE").ok().as_deref() {
Some("0") | Some("disable") => SPLIT_DISABLE,
Some("1") | Some("auto") => SPLIT_AUTO_FORCED,
Some("3") => clamp_to_engines(SPLIT_THREE_FORCED, hw_max, engines),
Some("2") => clamp_to_engines(SPLIT_TWO_FORCED, hw_max, engines),
// Use every engine the card has, not a hard-coded two: on a 3-NVENC part (GB202, AD102
// workstation) forcing 2 leaves a third of the silicon idle.
//
// ⚠ This arm now comes FIRST, ahead of the 10-bit rule. That reordering is the D1 fix: a
// 10-bit 4K120 session (995.3 Mpix/s) used to be vetoed by the depth rule before ever
// reaching the pixel-rate arm written for exactly it.
_ if pixel_rate >= SPLIT_FORCE_PIXEL_RATE => hw_max,
// Below that bar, HEVC Main10 keeps the conservative single-engine default. The one Ada
// measurement we have says split can be *slower* for Main10, and nothing under this bar
// needs a second engine anyway — so the cost of being wrong here is ~nil, unlike above it.
//
// ⚠ Now codec-scoped (the D2 fix): the measurement behind this was HEVC Main10 on Ada, and
// it used to veto **AV1 10-bit** too, which has neither the sub-frame conflict nor any
// measurement against it.
_ if codec == Codec::H265 && bit_depth >= 10 => SPLIT_DISABLE,
_ => SPLIT_AUTO,
};
tracing::debug!(
split_mode = mode,
?codec,
bit_depth,
pixel_rate,
engines,
"NVENC split-encode mode selected"
);
mode
}
/// The strongest split mode this GPU's engine count can actually deliver.
///
/// ⚠ **The driver will NOT tell you when you over-ask.** Measured on `.21` (RTX 5070 Ti, 2 NVENC,
/// driver 610.57.04, 4K HEVC): requesting `THREE_FORCED` was **HONOURED** — session opened in mode
/// 3 — and ran at **2303 µs/frame, identical to `TWO_FORCED`'s 2308**. No rejection, no warning,
/// no third engine; just a log line claiming 3-way over a 2-way encode. So the rejection fallback
/// cannot be relied on to find the ceiling and the clamp has to happen here.
///
/// `NV_ENC_SPLIT_ENCODE_MODE` can only *name* counts up to three (SDK 0.4.0 / NVENCAPI 12.1;
/// values 4..14 are unallocated, so a future API may extend it). Above that we fall back to
/// `AUTO_FORCED` = "split, driver picks how many", which measurably does force a split (2.01× vs
/// disabled on the same box) and is the only way to express "use everything you have".
pub(crate) fn max_forced_split_mode(engines: u32) -> u32 {
match engines {
// Unknown (cap unreadable / not probed): keep the historical assumption of a second
// engine and let the open-time rejection fallback correct it.
0 => SPLIT_TWO_FORCED,
1 => SPLIT_DISABLE,
2 => SPLIT_TWO_FORCED,
3 => SPLIT_THREE_FORCED,
// More engines than the enum can name — let the driver use them all.
_ => SPLIT_AUTO_FORCED,
}
}
/// The N of an N-way FORCED split, or `None` for the modes that do not name a width
/// (`DISABLE`, plain `AUTO`, and `AUTO_FORCED` — the last forces a split but lets the driver
/// choose how wide).
///
/// For callers that can only express "split this many ways" and have no vocabulary for our other
/// modes — the libav path, whose `split_encode_mode` AVOption is libavcodec's own enum, not the
/// NVENC one (our `DISABLE` is `15`, which would be meaningless there).
pub(crate) fn forced_split_width(mode: u32) -> Option<u32> {
match mode {
m if m == SPLIT_TWO_FORCED => Some(2),
m if m == SPLIT_THREE_FORCED => Some(3),
_ => None,
}
}
/// Hold an operator's `PUNKTFUNK_SPLIT_ENCODE=2|3` to what the hardware can deliver, loudly.
/// Without this the knob silently lies (see [`max_forced_split_mode`]); an override that asks for
/// more engines than exist is a mistake worth surfacing, not honouring.
pub(crate) fn clamp_to_engines(requested: u32, hw_max: u32, engines: u32) -> u32 {
// Only the named N-way modes are ordered; `hw_max` may be AUTO_FORCED (1) on a >3-engine part,
// which is not "less than" TWO_FORCED and must not clamp a legitimate request down.
let named = |m: u32| (2..=3).contains(&m);
if engines != 0 && named(requested) && named(hw_max) && requested > hw_max {
tracing::warn!(
requested,
engines,
using = hw_max,
"PUNKTFUNK_SPLIT_ENCODE asks for more NVENC engines than this GPU has — clamping. \
(The driver would ACCEPT the over-ask and silently encode with fewer, so the log \
would otherwise claim a split width that never happened.)"
);
return hw_max;
}
requested
}
/// `PUNKTFUNK_VBV_FRAMES` — HRD/VBV size in frame intervals (default 1.0, the strict low-latency
/// shape every backend ships: each frame must fit its rate share, keeping frame sizes uniform for
/// the pacer). The AMF/VAAPI/QSV paths parse the same variable locally; this helper brings the
+27 -15
View File
@@ -476,13 +476,22 @@ impl NvencEncoder {
opts.set("profile", "main10");
}
// Split-frame encode across both NVENC engines (GB203 has 2) when the pixel rate exceeds
// a single engine's HEVC capacity; e.g. 5120x1440@240 = 1.77 Gpix/s needs it, @120
// (0.88 Gpix/s) does not. HEVC/AV1 only (not H.264). AUTO won't engage below ~2112px
// height, so we force `2`; below the threshold we leave it AUTO (split costs ~2% BD-rate).
// Threshold shared with the direct-SDK selector ([`super::SPLIT_FORCE_PIXEL_RATE`] — set
// so 4K120 = 995.3 Mpix/s forces, which `> 1e9` famously missed by 0.47%). Output is
// standard HEVC — transparent to the client. Override with PUNKTFUNK_SPLIT_ENCODE.
// Split-frame encode across the GPU's NVENC engines. WP4: the policy is no longer
// duplicated here — it comes from the SAME [`resolve_split_mode`] the two direct-SDK
// backends use, so the pixel-rate threshold, the codec scoping and the (dropped) 10-bit
// short circuit cannot drift between the libav path and the rest. This copy had already
// diverged: it hard-coded a 2-way split regardless of engine count and carried no depth
// rule at all.
//
// ⚠ Only the FORCED outcomes are actionable here. libavcodec's `split_encode_mode`
// AVOption is its own vocabulary, and our `DISABLE` is the NVENC enum's `15` — passing
// that through would be meaningless to it (or fail the open). `DISABLE`/`AUTO` therefore
// both mean "leave the option unset", which is exactly today's behaviour: unset = the
// driver's own auto.
//
// ⚠ `engines = 0` = "not probed": the libav path has no caps probe of its own, and
// [`max_forced_split_mode`] maps unknown to 2-way, preserving what this site always did.
// A 3-NVENC part gets the wider split only on the direct-SDK path.
let pix_rate = width as u64 * height as u64 * fps as u64;
let split = std::env::var("PUNKTFUNK_SPLIT_ENCODE").ok();
match split.as_deref() {
@@ -497,14 +506,17 @@ impl NvencEncoder {
"PUNKTFUNK_SPLIT_ENCODE ignored — split encoding is not applicable to H.264 \
(nvEncodeAPI.h)"
),
None if matches!(codec, Codec::H265 | Codec::Av1)
&& pix_rate >= super::SPLIT_FORCE_PIXEL_RATE =>
{
opts.set("split_encode_mode", "2");
tracing::info!(
pix_rate,
"NVENC: forcing 2-way split encode (high pixel rate)"
);
None if matches!(codec, Codec::H265 | Codec::Av1) => {
let resolved = super::resolve_split_mode(codec, bit_depth, pix_rate, 0);
if let Some(n) = super::forced_split_width(resolved) {
opts.set("split_encode_mode", &n.to_string());
tracing::info!(
pix_rate,
bit_depth,
split_encode_mode = n,
"NVENC (libav): forcing split encode (shared selector)"
);
}
}
None => {}
}
+5 -5
View File
@@ -68,12 +68,12 @@
use super::nvenc_core::{
apply_low_latency_config, build_init_params, cached_ceiling, cached_split_verdict, codec_guid,
max_forced_split_mode, plan_range_recovery, resolve_slices, resolve_split_mode,
resolve_split_subframe, resolve_subframe, store_ceiling, store_split_verdict,
subframe_env_forced, ArbAction, CeilingKey, LowLatencyConfig, NvStatusExt, RangePlan,
SplitArbiter, SplitKey,
plan_range_recovery, resolve_slices, resolve_split_subframe, resolve_subframe, store_ceiling,
store_split_verdict, subframe_env_forced, ArbAction, CeilingKey, LowLatencyConfig, NvStatusExt,
RangePlan, SplitArbiter, SplitKey,
};
use super::nvenc_status;
use super::{max_forced_split_mode, resolve_split_mode};
use super::{AuChunk, ChromaFormat, Codec, EncodedFrame, Encoder, EncoderCaps};
use anyhow::{anyhow, bail, Context, Result};
use pf_frame::{CapturedFrame, FramePayload};
@@ -826,7 +826,7 @@ pub struct NvencCudaEncoder {
/// `NV_ENC_CAPS_NUM_ENCODER_ENGINES` — how many NVENC engines this GPU has, probed in
/// [`query_caps`]. `0` = not probed / unreadable. The split-encode ceiling: the driver accepts
/// a split wider than the hardware and silently encodes narrower, so this is the only honest
/// source for how wide we may go (see `nvenc_core::max_forced_split_mode`).
/// source for how wide we may go (see `codec::max_forced_split_mode`).
encoder_engines: u32,
/// Submit stamp for the split arbiter's per-frame cost (sync depth-1 path only).
last_submit_at: Option<std::time::Instant>,
+28 -126
View File
@@ -80,132 +80,6 @@ pub(super) fn resolve_subframe(default_on: bool) -> bool {
}
}
/// Resolved NVENC split-frame encode mode for a session — ONE selector shared by the Windows and
/// Linux direct-SDK backends (they had drifted into byte-identical duplicates, one of which
/// logged and one didn't). Precedence:
/// 1. `PUNKTFUNK_SPLIT_ENCODE` = `0`/`disable` | `1`/`auto` (AUTO_FORCED) | `2` | `3` — operator
/// override, always wins, except that `2`/`3` are clamped to the GPU's real engine count (see
/// [`clamp_to_engines`]; the driver honours an over-ask and silently encodes narrower).
/// 2. Pixel rate ≥ [`super::SPLIT_FORCE_PIXEL_RATE`] → force the WIDEST split the GPU can deliver
/// ([`max_forced_split_mode`]), not a hard-coded 2 (AUTO never engages below ~2112 px height,
/// so 4K120 must be forced onto the other engines; and a 3-NVENC part left at 2-way wastes a
/// third of its encode silicon).
/// 3. **HEVC** Main10 below that bar → DISABLE: 2-way split measured SLOWER on Ada for Main10 — at
/// 5120×1440@240 forced-2 took 7.6 ms/frame (~131 fps) vs 2.8 ms (~357 fps) single-engine, the
/// "broken animations in HDR" cap. ⚠ This rule used to sit ABOVE the pixel-rate arm and take no
/// codec, so it (a) vetoed 10-bit **4K120** — the very case the pixel-rate arm exists for — and
/// (b) applied an HEVC-on-Ada result to **AV1 10-bit**, which has no such measurement. Both
/// fixed; what remains is a conservative default in the regime where a second engine buys
/// nothing anyway.
/// ⚠⚠ **UNVALIDATED CONSEQUENCE:** 5120×1440@240 Main10 (1.77 Gpix/s) now clears the pixel-rate
/// bar and WILL be forced to split — i.e. the exact configuration that measurement came from
/// flips behaviour. That is deliberate (the datapoint is one sample, at low bits/frame, and the
/// bits/frame hypothesis predicts it should not generalise) but it is **the first thing to
/// re-measure on Ada**; `PUNKTFUNK_SPLIT_ENCODE=0` is the escape if it regresses.
/// 4. Else AUTO — ⚠ whose behaviour is **conditional on sub-frame**, measured on `.21` at 4K:
/// - sub-frame **ON** (the fleet default): AUTO **does not split** — 5023/5157 µs against
/// DISABLE's 4979/5000. Split and sub-frame are mutually unsupported for HEVC, so the driver
/// resolves AUTO to no-split and this arm silently means DISABLE.
/// - sub-frame **OFF**: AUTO **does split** — 2401/2352 µs against TWO_FORCED's 2319/2378.
///
/// So AUTO is NOT dead in general and must not be retired: doing so would lose a real split on
/// every sub-frame-off session. It is dead only in the sub-frame-on combination, which
/// [`resolve_split_subframe`] logs rather than silently accepting.
///
/// The caller still owns the rejection fallback (retry split-disabled) — a codec/config that
/// rejects the chosen mode downgrades at open, not here.
///
/// `engines` is the GPU's `NV_ENC_CAPS_NUM_ENCODER_ENGINES`; pass `0` when it could not be probed
/// (treated as "unknown", which keeps the pre-probe behaviour of assuming a second engine exists
/// and letting the open-time rejection fallback sort it out).
pub(super) fn resolve_split_mode(
codec: Codec,
bit_depth: u8,
pixel_rate: u64,
engines: u32,
) -> u32 {
use nv::NV_ENC_SPLIT_ENCODE_MODE as M;
let hw_max = max_forced_split_mode(engines);
let mode = match std::env::var("PUNKTFUNK_SPLIT_ENCODE").ok().as_deref() {
Some("0") | Some("disable") => M::NV_ENC_SPLIT_DISABLE_MODE as u32,
Some("1") | Some("auto") => M::NV_ENC_SPLIT_AUTO_FORCED_MODE as u32,
Some("3") => clamp_to_engines(M::NV_ENC_SPLIT_THREE_FORCED_MODE as u32, hw_max, engines),
Some("2") => clamp_to_engines(M::NV_ENC_SPLIT_TWO_FORCED_MODE as u32, hw_max, engines),
// Use every engine the card has, not a hard-coded two: on a 3-NVENC part (GB202, AD102
// workstation) forcing 2 leaves a third of the silicon idle.
//
// ⚠ This arm now comes FIRST, ahead of the 10-bit rule. That reordering is the D1 fix: a
// 10-bit 4K120 session (995.3 Mpix/s) used to be vetoed by the depth rule before ever
// reaching the pixel-rate arm written for exactly it.
_ if pixel_rate >= super::SPLIT_FORCE_PIXEL_RATE => hw_max,
// Below that bar, HEVC Main10 keeps the conservative single-engine default. The one Ada
// measurement we have says split can be *slower* for Main10, and nothing under this bar
// needs a second engine anyway — so the cost of being wrong here is ~nil, unlike above it.
//
// ⚠ Now codec-scoped (the D2 fix): the measurement behind this was HEVC Main10 on Ada, and
// it used to veto **AV1 10-bit** too, which has neither the sub-frame conflict nor any
// measurement against it.
_ if codec == Codec::H265 && bit_depth >= 10 => M::NV_ENC_SPLIT_DISABLE_MODE as u32,
_ => M::NV_ENC_SPLIT_AUTO_MODE as u32,
};
tracing::debug!(
split_mode = mode,
?codec,
bit_depth,
pixel_rate,
engines,
"NVENC split-encode mode selected"
);
mode
}
/// The strongest split mode this GPU's engine count can actually deliver.
///
/// ⚠ **The driver will NOT tell you when you over-ask.** Measured on `.21` (RTX 5070 Ti, 2 NVENC,
/// driver 610.57.04, 4K HEVC): requesting `THREE_FORCED` was **HONOURED** — session opened in mode
/// 3 — and ran at **2303 µs/frame, identical to `TWO_FORCED`'s 2308**. No rejection, no warning,
/// no third engine; just a log line claiming 3-way over a 2-way encode. So the rejection fallback
/// cannot be relied on to find the ceiling and the clamp has to happen here.
///
/// `NV_ENC_SPLIT_ENCODE_MODE` can only *name* counts up to three (SDK 0.4.0 / NVENCAPI 12.1;
/// values 4..14 are unallocated, so a future API may extend it). Above that we fall back to
/// `AUTO_FORCED` = "split, driver picks how many", which measurably does force a split (2.01× vs
/// disabled on the same box) and is the only way to express "use everything you have".
pub(super) fn max_forced_split_mode(engines: u32) -> u32 {
use nv::NV_ENC_SPLIT_ENCODE_MODE as M;
match engines {
// Unknown (cap unreadable / not probed): keep the historical assumption of a second
// engine and let the open-time rejection fallback correct it.
0 => M::NV_ENC_SPLIT_TWO_FORCED_MODE as u32,
1 => M::NV_ENC_SPLIT_DISABLE_MODE as u32,
2 => M::NV_ENC_SPLIT_TWO_FORCED_MODE as u32,
3 => M::NV_ENC_SPLIT_THREE_FORCED_MODE as u32,
// More engines than the enum can name — let the driver use them all.
_ => M::NV_ENC_SPLIT_AUTO_FORCED_MODE as u32,
}
}
/// Hold an operator's `PUNKTFUNK_SPLIT_ENCODE=2|3` to what the hardware can deliver, loudly.
/// Without this the knob silently lies (see [`max_forced_split_mode`]); an override that asks for
/// more engines than exist is a mistake worth surfacing, not honouring.
fn clamp_to_engines(requested: u32, hw_max: u32, engines: u32) -> u32 {
// Only the named N-way modes are ordered; `hw_max` may be AUTO_FORCED (1) on a >3-engine part,
// which is not "less than" TWO_FORCED and must not clamp a legitimate request down.
let named = |m: u32| (2..=3).contains(&m);
if engines != 0 && named(requested) && named(hw_max) && requested > hw_max {
tracing::warn!(
requested,
engines,
using = hw_max,
"PUNKTFUNK_SPLIT_ENCODE asks for more NVENC engines than this GPU has — clamping. \
(The driver would ACCEPT the over-ask and silently encode with fewer, so the log \
would otherwise claim a split width that never happened.)"
);
return hw_max;
}
requested
}
/// Whether the operator EXPLICITLY forced sub-frame readback on (`PUNKTFUNK_NVENC_SUBFRAME=1`)
/// — the log-severity input to [`resolve_split_subframe`]: a forced knob being overridden
/// deserves a `warn`, a default being tuned an `info`. Callers LATCH this once next to their
@@ -633,6 +507,7 @@ pub(super) fn clear_split_verdicts() {
#[cfg(test)]
mod tests {
use super::*;
use crate::{clamp_to_engines, max_forced_split_mode, resolve_split_mode};
use nv::NV_ENC_SPLIT_ENCODE_MODE as M;
// These assume PUNKTFUNK_SPLIT_ENCODE is unset (CI); an operator override deliberately wins.
@@ -1462,3 +1337,30 @@ mod arbiter_tests {
);
}
}
/// The hand-written split constants in `codec.rs` MUST equal the SDK enum they mirror. They are
/// duplicated there so the libav path — which builds without the `nvenc` feature, where the enum
/// does not exist — can share one policy instead of keeping the copy that had already drifted.
/// This is the only place both are visible at once.
#[cfg(test)]
mod split_constant_parity {
use nvidia_video_codec_sdk::sys::nvEncodeAPI::NV_ENC_SPLIT_ENCODE_MODE as M;
#[test]
fn nvenc_split_constants_match_the_sdk() {
assert_eq!(crate::SPLIT_AUTO, M::NV_ENC_SPLIT_AUTO_MODE as u32);
assert_eq!(
crate::SPLIT_AUTO_FORCED,
M::NV_ENC_SPLIT_AUTO_FORCED_MODE as u32
);
assert_eq!(
crate::SPLIT_TWO_FORCED,
M::NV_ENC_SPLIT_TWO_FORCED_MODE as u32
);
assert_eq!(
crate::SPLIT_THREE_FORCED,
M::NV_ENC_SPLIT_THREE_FORCED_MODE as u32
);
assert_eq!(crate::SPLIT_DISABLE, M::NV_ENC_SPLIT_DISABLE_MODE as u32);
}
}
+6 -3
View File
@@ -44,10 +44,13 @@
use super::nvenc_core::{
apply_low_latency_config, build_init_params, cached_ceiling, codec_guid, plan_range_recovery,
resolve_slices, resolve_split_mode, resolve_split_subframe, resolve_subframe, store_ceiling,
subframe_env_forced, CeilingKey, LowLatencyConfig, NvStatusExt, RangePlan,
resolve_slices, resolve_split_subframe, resolve_subframe, store_ceiling, subframe_env_forced,
CeilingKey, LowLatencyConfig, NvStatusExt, RangePlan,
};
// Moved to `codec.rs` (WP4) so the libav path, which builds without the `nvenc` feature, can share
// one split policy instead of keeping the copy that had already drifted.
use super::nvenc_status;
use super::resolve_split_mode;
use super::{AuChunk, ChromaFormat, Codec, EncodedFrame, Encoder, EncoderCaps};
use anyhow::{anyhow, bail, Context, Result};
use pf_frame::{CapturedFrame, FramePayload, PixelFormat};
@@ -595,7 +598,7 @@ pub struct NvencD3d11Encoder {
/// `NV_ENC_CAPS_NUM_ENCODER_ENGINES` — how many NVENC engines this GPU has, probed in
/// [`query_caps`](Self::query_caps). `0` = not probed / unreadable. The split-encode ceiling:
/// the driver accepts a split wider than the hardware and silently encodes narrower, so this
/// is the only honest source for how wide we may go (see `nvenc_core::max_forced_split_mode`).
/// is the only honest source for how wide we may go (see `codec::max_forced_split_mode`).
encoder_engines: u32,
/// (bitstream, mapped input resource to unmap after retrieval, pts_ns, recovery-anchor) per
/// in-flight encode. The fourth field tags the first frame encoded after a successful