fix(pyrowave): per-session raw-dmabuf zero-copy capture on the Linux NVIDIA host
ci / web (push) Successful in 52s
apple / swift (push) Successful in 1m13s
ci / docs-site (push) Successful in 1m11s
ci / bench (push) Successful in 5m35s
apple / screenshots (push) Successful in 6m28s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 29s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
android / android (push) Successful in 13m34s
deb / build-publish (push) Successful in 9m10s
deb / build-publish-host (push) Successful in 9m26s
arch / build-publish (push) Successful in 17m58s
windows-host / package (push) Successful in 16m40s
ci / rust (push) Successful in 29m45s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m12s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m8s
docker / deploy-docs (push) Successful in 28s
ci / web (push) Successful in 52s
apple / swift (push) Successful in 1m13s
ci / docs-site (push) Successful in 1m11s
ci / bench (push) Successful in 5m35s
apple / screenshots (push) Successful in 6m28s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 29s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
android / android (push) Successful in 13m34s
deb / build-publish (push) Successful in 9m10s
deb / build-publish-host (push) Successful in 9m26s
arch / build-publish (push) Successful in 17m58s
windows-host / package (push) Successful in 16m40s
ci / rust (push) Successful in 29m45s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m12s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m8s
docker / deploy-docs (push) Successful in 28s
A PyroWave session on an NVIDIA-auto host was forced onto CPU-RGB capture
(session_plan flipped gpu=false): Mutter blits tiled->LINEAR, we mmap +
de-pad ~30 MB, the encoder re-uploads it - three full-frame CPU touches
per frame at 5120x1440 while an HEVC session on the same box rides the
tiled EGL/CUDA zero-copy. The dmabuf passthrough + Vulkan tiled import
were already validated (8dc5d672) but only reachable via the global
PUNKTFUNK_ENCODER=pyrowave lab policy.
ZeroCopyPolicy gains pyrowave_session (from OutputFormat.pyrowave, i.e.
the negotiated codec): the capturer skips the NVENC-only EGL->CUDA
importer, takes the raw-dmabuf passthrough, and advertises the wavelet
encoder's Vulkan-importable modifiers so Mutter+NVIDIA negotiates tiled
zero-copy. The forced-CPU flip in session_plan is gone.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
@@ -244,8 +244,13 @@ pub struct ZeroCopyPolicy {
|
|||||||
/// The resolved backend produces GPU-resident frames (everything but the software encoder) —
|
/// The resolved backend produces GPU-resident frames (everything but the software encoder) —
|
||||||
/// used only to phrase the CPU-fallback warning (the host `encode::resolved_backend_is_gpu`).
|
/// used only to phrase the CPU-fallback warning (the host `encode::resolved_backend_is_gpu`).
|
||||||
pub backend_is_gpu: bool,
|
pub backend_is_gpu: bool,
|
||||||
|
/// THIS session encodes PyroWave: the frames' consumer is the wavelet encoder's own Vulkan
|
||||||
|
/// device, which imports raw dmabufs on ANY vendor — so the capturer takes the raw-dmabuf
|
||||||
|
/// passthrough (like the VAAPI backend) instead of the EGL→CUDA import whose payloads only
|
||||||
|
/// NVENC can consume. Per-session (the codec is negotiated), unlike `backend_is_vaapi`.
|
||||||
|
pub pyrowave_session: bool,
|
||||||
/// The PyroWave encoder's Vulkan-importable dmabuf modifiers for the capture's packed-RGB fourcc,
|
/// The PyroWave encoder's Vulkan-importable dmabuf modifiers for the capture's packed-RGB fourcc,
|
||||||
/// resolved when the encoder pref is `pyrowave` (the passthrough advertises them so Mutter+NVIDIA,
|
/// resolved when the session encodes PyroWave (the passthrough advertises them so Mutter+NVIDIA,
|
||||||
/// which allocates tiled-only, still negotiates zero-copy). Empty otherwise.
|
/// which allocates tiled-only, still negotiates zero-copy). Empty otherwise.
|
||||||
pub pyrowave_modifiers: Vec<u64>,
|
pub pyrowave_modifiers: Vec<u64>,
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -255,9 +255,11 @@ fn spawn_pipewire(
|
|||||||
want_hdr
|
want_hdr
|
||||||
};
|
};
|
||||||
// Mirror of the thread's `vaapi_passthrough` decision (deterministic from here: on a VAAPI
|
// Mirror of the thread's `vaapi_passthrough` decision (deterministic from here: on a VAAPI
|
||||||
// backend the EGL→CUDA importer is never built) — kept on the capturer so `next_frame`'s
|
// backend or a PyroWave session the EGL→CUDA importer is never built) — kept on the capturer
|
||||||
// negotiation-timeout branch knows a failed negotiation was the LINEAR-dmabuf offer.
|
// so `next_frame`'s negotiation-timeout branch knows a failed negotiation was the raw-dmabuf
|
||||||
let vaapi_dmabuf = zerocopy && !force_shm && policy.backend_is_vaapi;
|
// passthrough offer.
|
||||||
|
let vaapi_dmabuf =
|
||||||
|
zerocopy && !force_shm && (policy.backend_is_vaapi || policy.pyrowave_session);
|
||||||
let join = thread::Builder::new()
|
let join = thread::Builder::new()
|
||||||
.name("punktfunk-pipewire".into())
|
.name("punktfunk-pipewire".into())
|
||||||
.spawn(move || {
|
.spawn(move || {
|
||||||
@@ -1945,16 +1947,18 @@ mod pipewire {
|
|||||||
// Build the GPU importer up front — normally the ISOLATED worker process
|
// Build the GPU importer up front — normally the ISOLATED worker process
|
||||||
// (design/zerocopy-worker-isolation.md), so a driver fault on a dying compositor's
|
// (design/zerocopy-worker-isolation.md), so a driver fault on a dying compositor's
|
||||||
// dmabuf kills the worker, not this host. If it fails, log and fall back to the CPU path
|
// dmabuf kills the worker, not this host. If it fails, log and fall back to the CPU path
|
||||||
// (we simply won't request dmabuf below). Skipped entirely when the encode backend is
|
// (we simply won't request dmabuf below). Skipped entirely when the frames go to the
|
||||||
// VAAPI: those frames go to the raw-dmabuf passthrough, and building the importer there
|
// raw-dmabuf passthrough — the encode backend is VAAPI, or the SESSION encodes PyroWave
|
||||||
// would waste a CUDA probe — or worse, on an NVIDIA box forced to PUNKTFUNK_ENCODER=vaapi,
|
// (its Vulkan device imports raw dmabufs on any vendor): building the importer there
|
||||||
// succeed and produce CUDA payloads the VAAPI encoder must reject. Also skipped once
|
// would waste a CUDA probe — or worse, succeed and produce CUDA payloads only NVENC can
|
||||||
// repeated worker deaths latched the import off (a wedged GPU stack must not crash-loop).
|
// consume. Also skipped once repeated worker deaths latched the import off (a wedged GPU
|
||||||
|
// stack must not crash-loop).
|
||||||
let backend_is_vaapi = policy.backend_is_vaapi;
|
let backend_is_vaapi = policy.backend_is_vaapi;
|
||||||
|
let raw_passthrough = backend_is_vaapi || policy.pyrowave_session;
|
||||||
// HDR never builds the EGL→CUDA importer: its de-tile blit renders into 8-bit RGBA8,
|
// HDR never builds the EGL→CUDA importer: its de-tile blit renders into 8-bit RGBA8,
|
||||||
// which would silently crush the 10-bit depth. The HDR consumers are the CPU mmap path
|
// which would silently crush the 10-bit depth. The HDR consumers are the CPU mmap path
|
||||||
// (LINEAR de-pad → X2Rgb10 CPU frames) and the VAAPI raw-dmabuf passthrough.
|
// (LINEAR de-pad → X2Rgb10 CPU frames) and the VAAPI raw-dmabuf passthrough.
|
||||||
let mut importer = if zerocopy && !backend_is_vaapi && !want_hdr {
|
let mut importer = if zerocopy && !raw_passthrough && !want_hdr {
|
||||||
if pf_zerocopy::gpu_import_disabled() {
|
if pf_zerocopy::gpu_import_disabled() {
|
||||||
tracing::warn!(
|
tracing::warn!(
|
||||||
"zero-copy GPU import disabled after repeated import-worker deaths — using CPU path"
|
"zero-copy GPU import disabled after repeated import-worker deaths — using CPU path"
|
||||||
@@ -1980,9 +1984,11 @@ mod pipewire {
|
|||||||
// host. KWin/gamescope don't need it (they blit into the buffer, so no read-before-render
|
// host. KWin/gamescope don't need it (they blit into the buffer, so no read-before-render
|
||||||
// race).
|
// race).
|
||||||
let force_shm = std::env::var("PUNKTFUNK_FORCE_SHM").as_deref() == Ok("1");
|
let force_shm = std::env::var("PUNKTFUNK_FORCE_SHM").as_deref() == Ok("1");
|
||||||
// VAAPI zero-copy passthrough: zero-copy on, no EGL→CUDA importer (any non-NVIDIA host), and
|
// Raw-dmabuf zero-copy passthrough: zero-copy on, no EGL→CUDA importer, and the frames'
|
||||||
// the encoder backend is VAAPI → hand the raw dmabuf to the encoder (it imports + GPU-CSCs).
|
// consumer imports raw dmabufs itself — the VAAPI backend (libva import + GPU CSC) or a
|
||||||
let vaapi_passthrough = zerocopy && !force_shm && importer.is_none() && backend_is_vaapi;
|
// PyroWave session (the wavelet encoder's own Vulkan device, any vendor) → hand the raw
|
||||||
|
// dmabuf straight to the encoder.
|
||||||
|
let vaapi_passthrough = zerocopy && !force_shm && importer.is_none() && raw_passthrough;
|
||||||
// Modifiers our import stack handles for BGRx: the EGL-importable (tiled) set, plus LINEAR
|
// Modifiers our import stack handles for BGRx: the EGL-importable (tiled) set, plus LINEAR
|
||||||
// (0) — NVIDIA's EGL won't list it, but LINEAR dmabufs (gamescope's only offer) import via
|
// (0) — NVIDIA's EGL won't list it, but LINEAR dmabufs (gamescope's only offer) import via
|
||||||
// CUDA external memory instead. For the VAAPI passthrough path we advertise LINEAR only:
|
// CUDA external memory instead. For the VAAPI passthrough path we advertise LINEAR only:
|
||||||
@@ -1998,8 +2004,9 @@ mod pipewire {
|
|||||||
// advertisement with every modifier its device samples from, so compositors that
|
// advertisement with every modifier its device samples from, so compositors that
|
||||||
// never allocate LINEAR (Mutter+NVIDIA) still negotiate zero-copy dmabufs. The modifiers
|
// never allocate LINEAR (Mutter+NVIDIA) still negotiate zero-copy dmabufs. The modifiers
|
||||||
// were resolved by the facade (`ZeroCopyPolicy::pyrowave_modifiers`) — non-empty only when
|
// were resolved by the facade (`ZeroCopyPolicy::pyrowave_modifiers`) — non-empty only when
|
||||||
// the host's `pyrowave` feature is on AND the encoder pref is `pyrowave` — so capture never
|
// the host's `pyrowave` feature is on AND the session (or the global encoder pref) is
|
||||||
// calls back into `encode` and needs no feature gate of its own (the emptiness check gates it).
|
// PyroWave — so capture never calls back into `encode` and needs no feature gate of its
|
||||||
|
// own (the emptiness check gates it).
|
||||||
if vaapi_passthrough && !policy.pyrowave_modifiers.is_empty() {
|
if vaapi_passthrough && !policy.pyrowave_modifiers.is_empty() {
|
||||||
for &m in &policy.pyrowave_modifiers {
|
for &m in &policy.pyrowave_modifiers {
|
||||||
if !modifiers.contains(&m) {
|
if !modifiers.contains(&m) {
|
||||||
@@ -2019,11 +2026,11 @@ mod pipewire {
|
|||||||
);
|
);
|
||||||
} else if zerocopy && !want_dmabuf {
|
} else if zerocopy && !want_dmabuf {
|
||||||
tracing::warn!("zero-copy: no importable dmabuf modifiers — using CPU path");
|
tracing::warn!("zero-copy: no importable dmabuf modifiers — using CPU path");
|
||||||
} else if vaapi_passthrough {
|
} else if vaapi_passthrough && policy.pyrowave_modifiers.is_empty() {
|
||||||
tracing::info!(
|
tracing::info!(
|
||||||
"zero-copy: advertising LINEAR dmabuf for direct VAAPI import (GPU CSC)"
|
"zero-copy: advertising LINEAR dmabuf for direct VAAPI import (GPU CSC)"
|
||||||
);
|
);
|
||||||
} else if want_dmabuf {
|
} else if want_dmabuf && !vaapi_passthrough {
|
||||||
tracing::info!(
|
tracing::info!(
|
||||||
count = modifiers.len(),
|
count = modifiers.len(),
|
||||||
sample = ?&modifiers[..modifiers.len().min(6)],
|
sample = ?&modifiers[..modifiers.len().min(6)],
|
||||||
|
|||||||
@@ -26,24 +26,35 @@ pub use pf_capture::{dxgi, synthetic_nv12};
|
|||||||
/// capture→encode cycle). Resolved here (the host facade) and threaded in, so the edge stays one-way
|
/// capture→encode cycle). Resolved here (the host facade) and threaded in, so the edge stays one-way
|
||||||
/// (plan §2.4 / §W6).
|
/// (plan §2.4 / §W6).
|
||||||
#[cfg(target_os = "linux")]
|
#[cfg(target_os = "linux")]
|
||||||
fn zero_copy_policy() -> pf_capture::ZeroCopyPolicy {
|
fn zero_copy_policy(pyrowave_session: bool) -> pf_capture::ZeroCopyPolicy {
|
||||||
let backend_is_vaapi = crate::encode::linux_zero_copy_is_vaapi();
|
let backend_is_vaapi = crate::encode::linux_zero_copy_is_vaapi();
|
||||||
|
// The raw-dmabuf passthrough serves a PyroWave session on ANY vendor (the wavelet encoder's
|
||||||
|
// own Vulkan device imports the dmabuf) — per-session from the negotiated codec, plus the
|
||||||
|
// global `PUNKTFUNK_ENCODER=pyrowave` lab lever (which also flips `backend_is_vaapi`).
|
||||||
#[cfg(feature = "pyrowave")]
|
#[cfg(feature = "pyrowave")]
|
||||||
let pyrowave_modifiers =
|
let pyrowave_session =
|
||||||
if backend_is_vaapi && pf_host_config::config().encoder_pref.as_str() == "pyrowave" {
|
pyrowave_session || pf_host_config::config().encoder_pref.as_str() == "pyrowave";
|
||||||
// BGRx is the capture path's canonical packed-RGB format (the modifier advertisement keys
|
#[cfg(not(feature = "pyrowave"))]
|
||||||
// on it). `drm_fourcc(Bgrx)` is always `Some`.
|
let pyrowave_session = {
|
||||||
pf_frame::drm_fourcc(PixelFormat::Bgrx)
|
let _ = pyrowave_session;
|
||||||
.map(crate::encode::pyrowave_capture_modifiers)
|
false
|
||||||
.unwrap_or_default()
|
};
|
||||||
} else {
|
#[cfg(feature = "pyrowave")]
|
||||||
Vec::new()
|
let pyrowave_modifiers = if pyrowave_session {
|
||||||
};
|
// BGRx is the capture path's canonical packed-RGB format (the modifier advertisement keys
|
||||||
|
// on it). `drm_fourcc(Bgrx)` is always `Some`.
|
||||||
|
pf_frame::drm_fourcc(PixelFormat::Bgrx)
|
||||||
|
.map(crate::encode::pyrowave_capture_modifiers)
|
||||||
|
.unwrap_or_default()
|
||||||
|
} else {
|
||||||
|
Vec::new()
|
||||||
|
};
|
||||||
#[cfg(not(feature = "pyrowave"))]
|
#[cfg(not(feature = "pyrowave"))]
|
||||||
let pyrowave_modifiers = Vec::new();
|
let pyrowave_modifiers = Vec::new();
|
||||||
pf_capture::ZeroCopyPolicy {
|
pf_capture::ZeroCopyPolicy {
|
||||||
backend_is_vaapi,
|
backend_is_vaapi,
|
||||||
backend_is_gpu: crate::encode::resolved_backend_is_gpu(),
|
backend_is_gpu: crate::encode::resolved_backend_is_gpu(),
|
||||||
|
pyrowave_session,
|
||||||
pyrowave_modifiers,
|
pyrowave_modifiers,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -57,7 +68,9 @@ pub fn open_portal_monitor(want_hdr: bool) -> Result<Box<dyn Capturer>> {
|
|||||||
// session so it inherits that grant headlessly; wlroots/Sway has no RemoteDesktop portal,
|
// session so it inherits that grant headlessly; wlroots/Sway has no RemoteDesktop portal,
|
||||||
// so use a plain ScreenCast session there.
|
// so use a plain ScreenCast session there.
|
||||||
let anchored = crate::inject::default_backend() == crate::inject::Backend::Libei;
|
let anchored = crate::inject::default_backend() == crate::inject::Backend::Libei;
|
||||||
pf_capture::open_portal_monitor(anchored, want_hdr, zero_copy_policy())
|
// Monitor mirrors never carry the native PyroWave plane (GameStream protocol) — per-session
|
||||||
|
// passthrough is virtual-output-only; the global encoder-pref lever still applies inside.
|
||||||
|
pf_capture::open_portal_monitor(anchored, want_hdr, zero_copy_policy(false))
|
||||||
}
|
}
|
||||||
|
|
||||||
#[cfg(not(target_os = "linux"))]
|
#[cfg(not(target_os = "linux"))]
|
||||||
@@ -88,7 +101,7 @@ pub fn capture_virtual_output(
|
|||||||
vout.keepalive,
|
vout.keepalive,
|
||||||
want.gpu,
|
want.gpu,
|
||||||
want.chroma_444,
|
want.chroma_444,
|
||||||
zero_copy_policy(),
|
zero_copy_policy(want.pyrowave),
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -155,36 +155,22 @@ impl SessionPlan {
|
|||||||
}
|
}
|
||||||
gpu && !force_cpu_for_nvenc_444
|
gpu && !force_cpu_for_nvenc_444
|
||||||
};
|
};
|
||||||
// PyroWave on an NVIDIA-auto host: the `gpu` capture path resolves to the EGL→CUDA
|
// PyroWave on Linux keeps `gpu = true`: the capture facade sees `pyrowave` below and
|
||||||
// import that only NVENC can consume — the wavelet backend ingests raw dmabufs
|
// routes the session onto the raw-dmabuf passthrough (the wavelet encoder's own Vulkan
|
||||||
// (the AMD/Intel path) or CPU RGB. Flip THIS session to CPU RGB capture; the
|
// device imports the compositor's dmabuf on ANY vendor — `ZeroCopyPolicy::pyrowave_session`
|
||||||
// Phase-2 exit sessions ran exactly this shape at 60 fps (the encode itself stays
|
// advertises its importable modifiers, so Mutter+NVIDIA negotiates tiled zero-copy instead
|
||||||
// sub-ms GPU compute). Per-session raw-dmabuf passthrough on NVIDIA (true
|
// of the old forced CPU-RGB readback). The EGL→CUDA importer is skipped there — its
|
||||||
// zero-copy without the PUNKTFUNK_ENCODER=pyrowave capture policy) is the
|
// payloads only NVENC consumes.
|
||||||
// follow-up; the AMD/Intel dmabuf path is untouched.
|
|
||||||
#[cfg(target_os = "linux")]
|
|
||||||
let gpu = {
|
|
||||||
let pyro_needs_cpu = self.codec == crate::encode::Codec::PyroWave
|
|
||||||
&& !crate::encode::linux_zero_copy_is_vaapi();
|
|
||||||
if gpu && pyro_needs_cpu {
|
|
||||||
tracing::info!(
|
|
||||||
"PyroWave session on the NVIDIA capture path: GPU (CUDA) capture disabled \
|
|
||||||
for this session — frames arrive as CPU RGB and upload to the wavelet \
|
|
||||||
encoder (raw-dmabuf zero-copy on NVIDIA is a follow-up)"
|
|
||||||
);
|
|
||||||
}
|
|
||||||
gpu && !pyro_needs_cpu
|
|
||||||
};
|
|
||||||
crate::capture::OutputFormat {
|
crate::capture::OutputFormat {
|
||||||
gpu,
|
gpu,
|
||||||
hdr: self.hdr,
|
hdr: self.hdr,
|
||||||
// 4:4:4 needs a full-chroma source: on Windows this keeps the capturer on RGB (not the
|
// 4:4:4 needs a full-chroma source: on Windows this keeps the capturer on RGB (not the
|
||||||
// default NV12/P010 video-engine output) so NVENC can CSC to 4:4:4.
|
// default NV12/P010 video-engine output) so NVENC can CSC to 4:4:4.
|
||||||
chroma_444: self.chroma.is_444(),
|
chroma_444: self.chroma.is_444(),
|
||||||
// PyroWave (Windows): the IDD-push capturer makes its NV12 out-ring shareable + signals a
|
// PyroWave: on Windows the IDD-push capturer makes its NV12 out-ring shareable + signals
|
||||||
// shared fence so the wavelet encoder can zero-copy-import the texture into its own Vulkan
|
// a shared fence so the wavelet encoder can zero-copy-import the texture into its own
|
||||||
// device. Inert on Linux (the wavelet backend ingests dmabufs / CPU RGB there — handled
|
// Vulkan device; on Linux the capture facade flips the zero-copy policy to the
|
||||||
// by the `gpu` flips above, not this flag).
|
// raw-dmabuf passthrough (see above).
|
||||||
pyrowave: self.codec == crate::encode::Codec::PyroWave,
|
pyrowave: self.codec == crate::encode::Codec::PyroWave,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user