fix(pyrowave): per-session raw-dmabuf zero-copy capture on the Linux NVIDIA host
ci / web (push) Successful in 52s
apple / swift (push) Successful in 1m13s
ci / docs-site (push) Successful in 1m11s
ci / bench (push) Successful in 5m35s
apple / screenshots (push) Successful in 6m28s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 29s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
android / android (push) Successful in 13m34s
deb / build-publish (push) Successful in 9m10s
deb / build-publish-host (push) Successful in 9m26s
arch / build-publish (push) Successful in 17m58s
windows-host / package (push) Successful in 16m40s
ci / rust (push) Successful in 29m45s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m12s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m8s
docker / deploy-docs (push) Successful in 28s

A PyroWave session on an NVIDIA-auto host was forced onto CPU-RGB capture
(session_plan flipped gpu=false): Mutter blits tiled->LINEAR, we mmap +
de-pad ~30 MB, the encoder re-uploads it - three full-frame CPU touches
per frame at 5120x1440 while an HEVC session on the same box rides the
tiled EGL/CUDA zero-copy. The dmabuf passthrough + Vulkan tiled import
were already validated (8dc5d672) but only reachable via the global
PUNKTFUNK_ENCODER=pyrowave lab policy.

ZeroCopyPolicy gains pyrowave_session (from OutputFormat.pyrowave, i.e.
the negotiated codec): the capturer skips the NVENC-only EGL->CUDA
importer, takes the raw-dmabuf passthrough, and advertises the wavelet
encoder's Vulkan-importable modifiers so Mutter+NVIDIA negotiates tiled
zero-copy. The forced-CPU flip in session_plan is gone.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
enricobuehler
2026-07-19 15:57:12 +02:00
committed by enricobuehler
parent 1d587a259e
commit d2daeacc60
4 changed files with 65 additions and 54 deletions
+6 -1
View File
@@ -244,8 +244,13 @@ pub struct ZeroCopyPolicy {
/// The resolved backend produces GPU-resident frames (everything but the software encoder) — /// The resolved backend produces GPU-resident frames (everything but the software encoder) —
/// used only to phrase the CPU-fallback warning (the host `encode::resolved_backend_is_gpu`). /// used only to phrase the CPU-fallback warning (the host `encode::resolved_backend_is_gpu`).
pub backend_is_gpu: bool, pub backend_is_gpu: bool,
/// THIS session encodes PyroWave: the frames' consumer is the wavelet encoder's own Vulkan
/// device, which imports raw dmabufs on ANY vendor — so the capturer takes the raw-dmabuf
/// passthrough (like the VAAPI backend) instead of the EGL→CUDA import whose payloads only
/// NVENC can consume. Per-session (the codec is negotiated), unlike `backend_is_vaapi`.
pub pyrowave_session: bool,
/// The PyroWave encoder's Vulkan-importable dmabuf modifiers for the capture's packed-RGB fourcc, /// The PyroWave encoder's Vulkan-importable dmabuf modifiers for the capture's packed-RGB fourcc,
/// resolved when the encoder pref is `pyrowave` (the passthrough advertises them so Mutter+NVIDIA, /// resolved when the session encodes PyroWave (the passthrough advertises them so Mutter+NVIDIA,
/// which allocates tiled-only, still negotiates zero-copy). Empty otherwise. /// which allocates tiled-only, still negotiates zero-copy). Empty otherwise.
pub pyrowave_modifiers: Vec<u64>, pub pyrowave_modifiers: Vec<u64>,
} }
+23 -16
View File
@@ -255,9 +255,11 @@ fn spawn_pipewire(
want_hdr want_hdr
}; };
// Mirror of the thread's `vaapi_passthrough` decision (deterministic from here: on a VAAPI // Mirror of the thread's `vaapi_passthrough` decision (deterministic from here: on a VAAPI
// backend the EGL→CUDA importer is never built) — kept on the capturer so `next_frame`'s // backend or a PyroWave session the EGL→CUDA importer is never built) — kept on the capturer
// negotiation-timeout branch knows a failed negotiation was the LINEAR-dmabuf offer. // so `next_frame`'s negotiation-timeout branch knows a failed negotiation was the raw-dmabuf
let vaapi_dmabuf = zerocopy && !force_shm && policy.backend_is_vaapi; // passthrough offer.
let vaapi_dmabuf =
zerocopy && !force_shm && (policy.backend_is_vaapi || policy.pyrowave_session);
let join = thread::Builder::new() let join = thread::Builder::new()
.name("punktfunk-pipewire".into()) .name("punktfunk-pipewire".into())
.spawn(move || { .spawn(move || {
@@ -1945,16 +1947,18 @@ mod pipewire {
// Build the GPU importer up front — normally the ISOLATED worker process // Build the GPU importer up front — normally the ISOLATED worker process
// (design/zerocopy-worker-isolation.md), so a driver fault on a dying compositor's // (design/zerocopy-worker-isolation.md), so a driver fault on a dying compositor's
// dmabuf kills the worker, not this host. If it fails, log and fall back to the CPU path // dmabuf kills the worker, not this host. If it fails, log and fall back to the CPU path
// (we simply won't request dmabuf below). Skipped entirely when the encode backend is // (we simply won't request dmabuf below). Skipped entirely when the frames go to the
// VAAPI: those frames go to the raw-dmabuf passthrough, and building the importer there // raw-dmabuf passthrough — the encode backend is VAAPI, or the SESSION encodes PyroWave
// would waste a CUDA probe — or worse, on an NVIDIA box forced to PUNKTFUNK_ENCODER=vaapi, // (its Vulkan device imports raw dmabufs on any vendor): building the importer there
// succeed and produce CUDA payloads the VAAPI encoder must reject. Also skipped once // would waste a CUDA probe — or worse, succeed and produce CUDA payloads only NVENC can
// repeated worker deaths latched the import off (a wedged GPU stack must not crash-loop). // consume. Also skipped once repeated worker deaths latched the import off (a wedged GPU
// stack must not crash-loop).
let backend_is_vaapi = policy.backend_is_vaapi; let backend_is_vaapi = policy.backend_is_vaapi;
let raw_passthrough = backend_is_vaapi || policy.pyrowave_session;
// HDR never builds the EGL→CUDA importer: its de-tile blit renders into 8-bit RGBA8, // HDR never builds the EGL→CUDA importer: its de-tile blit renders into 8-bit RGBA8,
// which would silently crush the 10-bit depth. The HDR consumers are the CPU mmap path // which would silently crush the 10-bit depth. The HDR consumers are the CPU mmap path
// (LINEAR de-pad → X2Rgb10 CPU frames) and the VAAPI raw-dmabuf passthrough. // (LINEAR de-pad → X2Rgb10 CPU frames) and the VAAPI raw-dmabuf passthrough.
let mut importer = if zerocopy && !backend_is_vaapi && !want_hdr { let mut importer = if zerocopy && !raw_passthrough && !want_hdr {
if pf_zerocopy::gpu_import_disabled() { if pf_zerocopy::gpu_import_disabled() {
tracing::warn!( tracing::warn!(
"zero-copy GPU import disabled after repeated import-worker deaths — using CPU path" "zero-copy GPU import disabled after repeated import-worker deaths — using CPU path"
@@ -1980,9 +1984,11 @@ mod pipewire {
// host. KWin/gamescope don't need it (they blit into the buffer, so no read-before-render // host. KWin/gamescope don't need it (they blit into the buffer, so no read-before-render
// race). // race).
let force_shm = std::env::var("PUNKTFUNK_FORCE_SHM").as_deref() == Ok("1"); let force_shm = std::env::var("PUNKTFUNK_FORCE_SHM").as_deref() == Ok("1");
// VAAPI zero-copy passthrough: zero-copy on, no EGL→CUDA importer (any non-NVIDIA host), and // Raw-dmabuf zero-copy passthrough: zero-copy on, no EGL→CUDA importer, and the frames'
// the encoder backend is VAAPI → hand the raw dmabuf to the encoder (it imports + GPU-CSCs). // consumer imports raw dmabufs itself — the VAAPI backend (libva import + GPU CSC) or a
let vaapi_passthrough = zerocopy && !force_shm && importer.is_none() && backend_is_vaapi; // PyroWave session (the wavelet encoder's own Vulkan device, any vendor) → hand the raw
// dmabuf straight to the encoder.
let vaapi_passthrough = zerocopy && !force_shm && importer.is_none() && raw_passthrough;
// Modifiers our import stack handles for BGRx: the EGL-importable (tiled) set, plus LINEAR // Modifiers our import stack handles for BGRx: the EGL-importable (tiled) set, plus LINEAR
// (0) — NVIDIA's EGL won't list it, but LINEAR dmabufs (gamescope's only offer) import via // (0) — NVIDIA's EGL won't list it, but LINEAR dmabufs (gamescope's only offer) import via
// CUDA external memory instead. For the VAAPI passthrough path we advertise LINEAR only: // CUDA external memory instead. For the VAAPI passthrough path we advertise LINEAR only:
@@ -1998,8 +2004,9 @@ mod pipewire {
// advertisement with every modifier its device samples from, so compositors that // advertisement with every modifier its device samples from, so compositors that
// never allocate LINEAR (Mutter+NVIDIA) still negotiate zero-copy dmabufs. The modifiers // never allocate LINEAR (Mutter+NVIDIA) still negotiate zero-copy dmabufs. The modifiers
// were resolved by the facade (`ZeroCopyPolicy::pyrowave_modifiers`) — non-empty only when // were resolved by the facade (`ZeroCopyPolicy::pyrowave_modifiers`) — non-empty only when
// the host's `pyrowave` feature is on AND the encoder pref is `pyrowave` — so capture never // the host's `pyrowave` feature is on AND the session (or the global encoder pref) is
// calls back into `encode` and needs no feature gate of its own (the emptiness check gates it). // PyroWave — so capture never calls back into `encode` and needs no feature gate of its
// own (the emptiness check gates it).
if vaapi_passthrough && !policy.pyrowave_modifiers.is_empty() { if vaapi_passthrough && !policy.pyrowave_modifiers.is_empty() {
for &m in &policy.pyrowave_modifiers { for &m in &policy.pyrowave_modifiers {
if !modifiers.contains(&m) { if !modifiers.contains(&m) {
@@ -2019,11 +2026,11 @@ mod pipewire {
); );
} else if zerocopy && !want_dmabuf { } else if zerocopy && !want_dmabuf {
tracing::warn!("zero-copy: no importable dmabuf modifiers — using CPU path"); tracing::warn!("zero-copy: no importable dmabuf modifiers — using CPU path");
} else if vaapi_passthrough { } else if vaapi_passthrough && policy.pyrowave_modifiers.is_empty() {
tracing::info!( tracing::info!(
"zero-copy: advertising LINEAR dmabuf for direct VAAPI import (GPU CSC)" "zero-copy: advertising LINEAR dmabuf for direct VAAPI import (GPU CSC)"
); );
} else if want_dmabuf { } else if want_dmabuf && !vaapi_passthrough {
tracing::info!( tracing::info!(
count = modifiers.len(), count = modifiers.len(),
sample = ?&modifiers[..modifiers.len().min(6)], sample = ?&modifiers[..modifiers.len().min(6)],
+26 -13
View File
@@ -26,24 +26,35 @@ pub use pf_capture::{dxgi, synthetic_nv12};
/// capture→encode cycle). Resolved here (the host facade) and threaded in, so the edge stays one-way /// capture→encode cycle). Resolved here (the host facade) and threaded in, so the edge stays one-way
/// (plan §2.4 / §W6). /// (plan §2.4 / §W6).
#[cfg(target_os = "linux")] #[cfg(target_os = "linux")]
fn zero_copy_policy() -> pf_capture::ZeroCopyPolicy { fn zero_copy_policy(pyrowave_session: bool) -> pf_capture::ZeroCopyPolicy {
let backend_is_vaapi = crate::encode::linux_zero_copy_is_vaapi(); let backend_is_vaapi = crate::encode::linux_zero_copy_is_vaapi();
// The raw-dmabuf passthrough serves a PyroWave session on ANY vendor (the wavelet encoder's
// own Vulkan device imports the dmabuf) — per-session from the negotiated codec, plus the
// global `PUNKTFUNK_ENCODER=pyrowave` lab lever (which also flips `backend_is_vaapi`).
#[cfg(feature = "pyrowave")] #[cfg(feature = "pyrowave")]
let pyrowave_modifiers = let pyrowave_session =
if backend_is_vaapi && pf_host_config::config().encoder_pref.as_str() == "pyrowave" { pyrowave_session || pf_host_config::config().encoder_pref.as_str() == "pyrowave";
// BGRx is the capture path's canonical packed-RGB format (the modifier advertisement keys #[cfg(not(feature = "pyrowave"))]
// on it). `drm_fourcc(Bgrx)` is always `Some`. let pyrowave_session = {
pf_frame::drm_fourcc(PixelFormat::Bgrx) let _ = pyrowave_session;
.map(crate::encode::pyrowave_capture_modifiers) false
.unwrap_or_default() };
} else { #[cfg(feature = "pyrowave")]
Vec::new() let pyrowave_modifiers = if pyrowave_session {
}; // BGRx is the capture path's canonical packed-RGB format (the modifier advertisement keys
// on it). `drm_fourcc(Bgrx)` is always `Some`.
pf_frame::drm_fourcc(PixelFormat::Bgrx)
.map(crate::encode::pyrowave_capture_modifiers)
.unwrap_or_default()
} else {
Vec::new()
};
#[cfg(not(feature = "pyrowave"))] #[cfg(not(feature = "pyrowave"))]
let pyrowave_modifiers = Vec::new(); let pyrowave_modifiers = Vec::new();
pf_capture::ZeroCopyPolicy { pf_capture::ZeroCopyPolicy {
backend_is_vaapi, backend_is_vaapi,
backend_is_gpu: crate::encode::resolved_backend_is_gpu(), backend_is_gpu: crate::encode::resolved_backend_is_gpu(),
pyrowave_session,
pyrowave_modifiers, pyrowave_modifiers,
} }
} }
@@ -57,7 +68,9 @@ pub fn open_portal_monitor(want_hdr: bool) -> Result<Box<dyn Capturer>> {
// session so it inherits that grant headlessly; wlroots/Sway has no RemoteDesktop portal, // session so it inherits that grant headlessly; wlroots/Sway has no RemoteDesktop portal,
// so use a plain ScreenCast session there. // so use a plain ScreenCast session there.
let anchored = crate::inject::default_backend() == crate::inject::Backend::Libei; let anchored = crate::inject::default_backend() == crate::inject::Backend::Libei;
pf_capture::open_portal_monitor(anchored, want_hdr, zero_copy_policy()) // Monitor mirrors never carry the native PyroWave plane (GameStream protocol) — per-session
// passthrough is virtual-output-only; the global encoder-pref lever still applies inside.
pf_capture::open_portal_monitor(anchored, want_hdr, zero_copy_policy(false))
} }
#[cfg(not(target_os = "linux"))] #[cfg(not(target_os = "linux"))]
@@ -88,7 +101,7 @@ pub fn capture_virtual_output(
vout.keepalive, vout.keepalive,
want.gpu, want.gpu,
want.chroma_444, want.chroma_444,
zero_copy_policy(), zero_copy_policy(want.pyrowave),
) )
} }
+10 -24
View File
@@ -155,36 +155,22 @@ impl SessionPlan {
} }
gpu && !force_cpu_for_nvenc_444 gpu && !force_cpu_for_nvenc_444
}; };
// PyroWave on an NVIDIA-auto host: the `gpu` capture path resolves to the EGL→CUDA // PyroWave on Linux keeps `gpu = true`: the capture facade sees `pyrowave` below and
// import that only NVENC can consume — the wavelet backend ingests raw dmabufs // routes the session onto the raw-dmabuf passthrough (the wavelet encoder's own Vulkan
// (the AMD/Intel path) or CPU RGB. Flip THIS session to CPU RGB capture; the // device imports the compositor's dmabuf on ANY vendor — `ZeroCopyPolicy::pyrowave_session`
// Phase-2 exit sessions ran exactly this shape at 60 fps (the encode itself stays // advertises its importable modifiers, so Mutter+NVIDIA negotiates tiled zero-copy instead
// sub-ms GPU compute). Per-session raw-dmabuf passthrough on NVIDIA (true // of the old forced CPU-RGB readback). The EGL→CUDA importer is skipped there — its
// zero-copy without the PUNKTFUNK_ENCODER=pyrowave capture policy) is the // payloads only NVENC consumes.
// follow-up; the AMD/Intel dmabuf path is untouched.
#[cfg(target_os = "linux")]
let gpu = {
let pyro_needs_cpu = self.codec == crate::encode::Codec::PyroWave
&& !crate::encode::linux_zero_copy_is_vaapi();
if gpu && pyro_needs_cpu {
tracing::info!(
"PyroWave session on the NVIDIA capture path: GPU (CUDA) capture disabled \
for this session — frames arrive as CPU RGB and upload to the wavelet \
encoder (raw-dmabuf zero-copy on NVIDIA is a follow-up)"
);
}
gpu && !pyro_needs_cpu
};
crate::capture::OutputFormat { crate::capture::OutputFormat {
gpu, gpu,
hdr: self.hdr, hdr: self.hdr,
// 4:4:4 needs a full-chroma source: on Windows this keeps the capturer on RGB (not the // 4:4:4 needs a full-chroma source: on Windows this keeps the capturer on RGB (not the
// default NV12/P010 video-engine output) so NVENC can CSC to 4:4:4. // default NV12/P010 video-engine output) so NVENC can CSC to 4:4:4.
chroma_444: self.chroma.is_444(), chroma_444: self.chroma.is_444(),
// PyroWave (Windows): the IDD-push capturer makes its NV12 out-ring shareable + signals a // PyroWave: on Windows the IDD-push capturer makes its NV12 out-ring shareable + signals
// shared fence so the wavelet encoder can zero-copy-import the texture into its own Vulkan // a shared fence so the wavelet encoder can zero-copy-import the texture into its own
// device. Inert on Linux (the wavelet backend ingests dmabufs / CPU RGB there — handled // Vulkan device; on Linux the capture facade flips the zero-copy policy to the
// by the `gpu` flips above, not this flag). // raw-dmabuf passthrough (see above).
pyrowave: self.codec == crate::encode::Codec::PyroWave, pyrowave: self.codec == crate::encode::Codec::PyroWave,
} }
} }