perf(latency): T2.5b — NV12 compute CSC on the LINEAR/gamescope zero-copy path
apple / swift (push) Successful in 1m21s
apple / screenshots (push) Successful in 6m23s
ci / web (push) Successful in 54s
arch / build-publish (push) Successful in 10m52s
ci / docs-site (push) Successful in 1m7s
decky / build-publish (push) Successful in 35s
android / android (push) Successful in 13m2s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 16s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 16s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
ci / bench (push) Successful in 6m8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 7m56s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 9m44s
docker / deploy-docs (push) Successful in 24s
deb / build-publish (push) Successful in 11m52s
ci / rust (push) Successful in 27m4s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m26s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m20s

design/latency-reduction-2026-07.md T2.5's Linux half: the LINEAR dmabuf path
(gamescope's only offer) fed NVENC RGB, paying its internal RGB->YUV CSC on
the SM the game is saturating — the exact contention §5.A removed everywhere
else. The Vulkan bridge now carries a buffer-to-buffer RGB->NV12 compute
shader (rgb2nv12_buf.comp, BT.709 limited, coefficient-identical to
pf-encode's rgb2yuv.comp; whole-word writes so no 8-bit-storage feature is
needed): import dmabuf -> dispatch CSC into the exportable buffer -> CUDA
de-strides both planes into a pooled two-plane NV12 buffer. PUNKTFUNK_NV12
(default-on) now covers LINEAR; a CSC failure latches RGB for the stream
(mid-frame fallback, no dropped frame); 4:4:4 LINEAR sessions stay RGB (never
silently subsample). New ImportKind::LinearNv12 rides the existing worker IPC
(appended last per the wire-tag rule); cursor stays downstream (blend_nv12).

Validated: .21 clippy -D warnings (pf-zerocopy/pf-capture/host+nvenc) + 17
zero-copy tests. Owed: on-glass gamescope session (visual + dmon sm% check).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-17 20:14:20 +02:00
parent fbe1e62ef2
commit 5e1e64e50b
10 changed files with 500 additions and 19 deletions
+36
View File
@@ -504,6 +504,9 @@ pub struct EglImporter {
/// created lazily on the first LINEAR frame, + the destination pool.
vk: Option<super::vulkan::VkBridge>,
linear_pool: Option<cuda::BufferPool>,
/// NV12 twin of [`linear_pool`](Self::linear_pool) for the bridge's compute-CSC output
/// (T2.5b) — separate pools because a session may fall back RGB mid-stream.
linear_nv12_pool: Option<cuda::BufferPool>,
gbm: *mut c_void,
render_fd: c_int,
}
@@ -647,6 +650,7 @@ impl EglImporter {
yuv444_blit: None,
vk: None,
linear_pool: None,
linear_nv12_pool: None,
gbm,
render_fd,
})
@@ -677,6 +681,38 @@ impl EglImporter {
)
}
/// Like [`import_linear`](Self::import_linear), but the bridge's compute CSC converts to a
/// two-plane **NV12** buffer (latency plan T2.5b) — the gamescope/LINEAR analogue of
/// [`import_nv12`](Self::import_nv12), so NVENC encodes native YUV on the dedicated-session
/// path too instead of paying its internal RGB→YUV CSC on the contended SM.
pub fn import_linear_nv12(
&mut self,
plane: &DmabufPlane,
width: u32,
height: u32,
) -> Result<DeviceBuffer> {
cuda::make_current()?;
if self
.linear_nv12_pool
.as_ref()
.map(|p| (p.width(), p.height()))
!= Some((width, height))
{
self.linear_nv12_pool = Some(cuda::BufferPool::new_nv12(width, height)?);
}
if self.vk.is_none() {
self.vk = Some(super::vulkan::VkBridge::new()?);
}
self.vk.as_mut().unwrap().import_linear_nv12(
plane.fd,
plane.offset,
plane.stride,
width,
height,
self.linear_nv12_pool.as_ref().unwrap(),
)
}
/// Drop the Vulkan bridge's cached per-fd import (see [`super::vulkan::VkBridge::forget_fd`]).
/// No-op when the bridge hasn't been built (tiled-only captures).
pub fn forget_linear_fd(&mut self, fd: i32) {