59b6d3796b
Latency plan §7 LN0 (Linux/NVIDIA encode follow-on): - sampled PUNKTFUNK_PERF submit split (copy/blend/map/pic) in nvenc_cuda — the host loop's submit_us folds all four together; the D2D input copy is the LN2 zero-copy target and now measurable on its own - sampled blocking lock_bitstream timing on pf-nvenc-out — in two-thread mode the host loop's wait_us wraps a non-blocking poll, so the real encode wait was measured by no timer - caps probe + log SUPPORT_SUBFRAME_READBACK / SUPPORT_DYNAMIC_SLICE_MODE (LN1 sub-frame slice-output prerequisites, fleet visibility) - explicit zeroReorderDelay=1 in the shared low-latency config (P-only + no lookahead has no reordering anyway; pins the bit against preset or driver drift; shared with the Windows backend) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>