feat(pf-encode,host): price the HEVC sub-frame trade so arbitration can cover it
The named next step after WP3's first increment. That increment deliberately REFUSED to arbitrate HEVC-with-sub-frame -- the fleet default, and the reported field case -- because engaging split there gives up sub-frame readback, whose whole value is that the send overlaps the encode. An encoder measuring only encode time would see split as ~2x faster, take it, and make end-to-end latency worse while reporting a win. This supplies the missing number. The real comparison is encode_1eng + send_of_last_slice against encode_2eng + send_of_whole_AU, so the challenger owes roughly spread x (slices-1)/slices. Split across the two sides that can each see half: - Host: new `Encoder::set_send_spread_us` (defaulted, forwarded by TrackedEncoder -- same trap class as set_wire_chunking, and unforwarded it would fail SILENTLY IN THE SAFE DIRECTION, which is the hardest kind to notice). The send thread is the only place a paced send is observed and the encode loop the only place the encoder can be touched, so it goes over an AtomicU32 like encoder_ceiling_kbps, EWMA-smoothed 3:1 per completed AU: one content spike must not flip a verdict that then gets cached. - Encoder: turns the raw spread into the handicap, because only it knows `slices`. SplitArbiter::with_handicap charges it to the challenger before the comparison. A unit test runs identical encode numbers with a cheap and an expensive send and asserts the verdict REVERSES -- with an expensive send the arm that looks twice as fast is a loss end to end, and the incumbent must hold. That is precisely the regression an encode-only arbiter ships. Gate now opens for HEVC+sub-frame only when a spread has actually been reported (and slices >= 2); with no hint it still refuses, so behaviour is unchanged until the host feeds it. Two mechanics this needed: - apply_split_mode became a PAIR flip (split + sub-frame), routed through resolve_split_subframe and restoring from `subframe_opened_with` so a session that never had sub-frame can never gain it. It also recomputes `subframe_chunks`, which reconfigure_bitrate does NOT -- spike S1c's finding; leave it stale and supports_chunked_poll keeps saying yes while numSlices never advances, so poll_chunk busy-polls its whole budget every AU. - The arbiter is now fed from BOTH completion points. A sub-frame session finishes through poll_chunk, so the incumbent arm of an HEVC experiment would otherwise never deliver a sample -- only the challenger, with sub-frame dropped, comes through poll. Verified .21: clippy -D warnings clean for pf-encode AND punktfunk-host with nvenc, 63 unit tests (1 new), 23/23 NVENC on-hardware green. Verified .133: Windows clippy -D warnings clean, zero dead_code. fmt clean.
This commit is contained in:
@@ -443,6 +443,20 @@ pub trait Encoder: Send {
|
||||
/// flagged [`EncodedFrame::chunk_aligned`] and the session marks them on the wire.
|
||||
/// Default: no-op (the H.26x backends' bitstreams cannot be cut losslessly).
|
||||
fn set_wire_chunking(&mut self, _shard_payload: usize) {}
|
||||
/// How long a whole AU's packets currently take to leave the socket (µs, smoothed) — the
|
||||
/// host's paced-send `spread_us`.
|
||||
///
|
||||
/// Exists for ONE decision, and only the host can supply it. The Linux direct-NVENC split
|
||||
/// arbitration compares single-engine against split, but on HEVC engaging split costs
|
||||
/// sub-frame readback, and sub-frame's whole value is that the send overlaps the encode. So
|
||||
/// the real comparison is `encode_1eng + send_of_last_slice` against
|
||||
/// `encode_2eng + send_of_whole_AU`, and an encoder that measures only encode time would
|
||||
/// reliably pick split and make end-to-end latency WORSE. The backend turns this number into
|
||||
/// that handicap (it knows its own slice count); the host just reports what it observes.
|
||||
///
|
||||
/// Optional by design: a backend that ignores it simply never arbitrates the sub-frame trade,
|
||||
/// which is the safe direction. `0` = unknown / not reported yet.
|
||||
fn set_send_spread_us(&mut self, _us: u32) {}
|
||||
/// How many frames the CAPTURER guarantees the encoder may hold in flight before it starts
|
||||
/// reusing an input texture (`Capturer::pipeline_depth`). Backends that encode the capturer's
|
||||
/// textures IN PLACE — no `CopyResource` — must not pipeline deeper than this: the capturer
|
||||
|
||||
Reference in New Issue
Block a user