PyroWave encodes on the same GPU shader cores the game saturates, and an elevated VK_KHR_global_priority queue is the compute-preemption lever for it — measured on .21 (RTX 5070 Ti, GRID 2 loop): encode p99 6.4 -> 4.4 ms. Every driver refuses every priority class without CAP_SYS_NICE, on NVIDIA and on RADV alike, so the lever is decoration on a packaged host. 0.26.0-1 granted that capability to punktfunk-host and killed desktop streaming on every KDE box: KWin identifies a client by resolving /proc/<pid>/exe and matching an installed .desktop's Exec=, the kernel refuses that readlink to a reader whose effective set is not a superset of the target's PERMITTED set (cap_ptrace_access_check), and KWin holds no capabilities. #136 revoked it everywhere. The capability therefore cannot live in the process that fronts KWin. It lives in a new, deliberately small binary — punktfunk-encode-worker — which owns the priority-elevated Vulkan device and talks to nothing but the socket its parent spawned it on: no Wayland, no D-Bus, no network, no plugins. It is a SEPARATE FILE and must stay one; a hardlink or a hidden host subcommand shares the inode, hence the capability, and silently re-creates the incident. That rule is written where someone would break it, in the worker crate's own Cargo.toml. `open_inner` is reused verbatim in the worker — the same REALTIME->HIGH->none ladder, the same refusal-never-fails-open invariant, the same PUNKTFUNK_PERF split — so the A/B stays comparable with PW1. The only in-process change is a flag for whether THIS process prints the INERT warn, plus an out-parameter reporting the class that was granted. Three things the design did not anticipate: * An AU cannot ride in the message body. MAX_MSG is 64 KiB and bodies are serde_json, which renders a Vec<u8> as one decimal per byte: a 1080p60 AU is ~333 KB of JSON and 4K ~3.3 MB, and the minimum per-frame budget is already 64 KiB. So the AU crosses on a memfd the worker creates once and pwrites each frame; the fd crosses once, in Ready. A test pins the arithmetic so nobody "simplifies" the memfd away. Cursor bitmaps take the same route, only when their serial changes. * set_wire_chunking has to cross the wire even though poll_chunk does not. Chunking changes the AU BYTES, not merely how they are handed out — it feeds rate_budget()'s deflation and build_au's windowed framing — so a proxy-local copy would have the host cutting dense AUs at boundaries that are not window boundaries. Forwarded and mirrored. poll_chunk itself needs no protocol: the identical AuChunker runs host-side on the whole AU the worker returns. * CPU-backed frames really do reach this encoder (force_cpu_for_nvenc_444, and the raw-dmabuf degrade latch), and a 1080p BGRA frame is ~8 MB. The first non-dmabuf frame pins the session in-process with one warn rather than putting 480 MB/s on a socket. Every rung falls back to the in-process encoder exactly as today with one warn and never a dead session: PUNKTFUNK_ENCODE_WORKER=off, binary missing, spawn failure, handshake timeout, proto or workspace-version skew (host and worker are different files now, so that check is load-bearing), InitErr, a refused frame, and socket EOF mid-session — which respawns once, then pins inline. Also: recv retries EINTR with the REMAINING deadline, not a fresh one. With SO_RCVTIMEO the kernel returns EINTR rather than restarting, so a signal would otherwise read as a dead worker; re-arming with the full budget would instead let a steady signal rate defer a real hang forever.
112 lines
5.1 KiB
TOML
112 lines
5.1 KiB
TOML
[workspace]
|
|
resolver = "2"
|
|
members = [
|
|
"crates/punktfunk-core",
|
|
"crates/punktfunk-host",
|
|
"crates/punktfunk-host/vendor/usbip-sim",
|
|
# The capability-carrying PyroWave encode worker. A SEPARATE binary by design — never a
|
|
# hardlink of, or a subcommand of, punktfunk-host (design/gpu-priority-capability-worker.md).
|
|
"crates/punktfunk-encode-worker",
|
|
"crates/punktfunk-tray",
|
|
"crates/pf-bitstream",
|
|
"crates/pf-bitstream/vendor/cros-codecs",
|
|
"crates/pf-client-core",
|
|
"crates/pf-clipboard",
|
|
"crates/pf-presenter",
|
|
"crates/pf-console-ui",
|
|
"crates/pf-driver-proto",
|
|
"crates/pf-paths",
|
|
"crates/pf-update",
|
|
"crates/pf-update-check",
|
|
"crates/pf-host-config",
|
|
"crates/pf-gpu",
|
|
"crates/pf-zerocopy",
|
|
"crates/pf-frame",
|
|
"crates/pf-win-display",
|
|
"crates/pf-encode",
|
|
"crates/pf-capture",
|
|
"crates/pf-inject",
|
|
"crates/pf-vdisplay",
|
|
"crates/pf-vkdecode",
|
|
"crates/pf-dxvadec",
|
|
"crates/pf-vaadec",
|
|
"crates/pyrowave-sys",
|
|
"crates/libvpl-sys",
|
|
"clients/probe",
|
|
"clients/cli",
|
|
"clients/linux",
|
|
"clients/session",
|
|
"clients/windows",
|
|
"clients/android/native",
|
|
"tools/cursor-probe",
|
|
"tools/display-disturb",
|
|
"tools/latency-probe",
|
|
"tools/loss-harness",
|
|
]
|
|
# Standalone PoC (built on its own; pulls usbip/tokio/libusb we don't want in the workspace).
|
|
# The vendored `ndk` is a [patch.crates-io] source, not a member: it only compiles for the
|
|
# `*-linux-android` targets, so workspace membership would break host `cargo build --workspace`.
|
|
exclude = [
|
|
"packaging/linux/steam-deck-gadget/usbip-poc",
|
|
"clients/android/native/vendor/ndk",
|
|
]
|
|
|
|
# ndk 0.9.0 verbatim from crates.io plus ONE visibility change (and two warning fixes — an
|
|
# unnecessary `std::` qualification and a feature-gated `Result` import): `MediaCodec::as_ptr` made public
|
|
# (upstream keeps it private and exposes no frame-rendered binding), so the Android client can
|
|
# call `AMediaCodec_setOnFrameRenderedCallback` via ndk-sys for the HUD's `display` stage
|
|
# (design/stats-unification.md). Drop the patch when upstream exposes the pointer or the callback.
|
|
[patch.crates-io]
|
|
ndk = { path = "clients/android/native/vendor/ndk" }
|
|
|
|
[workspace.package]
|
|
version = "0.26.0"
|
|
edition = "2021"
|
|
rust-version = "1.82"
|
|
license = "MIT OR Apache-2.0"
|
|
authors = ["unom"]
|
|
repository = "https://git.unom.io/unom/punktfunk"
|
|
|
|
# The `unsafe` discipline the `packaging/windows/drivers/*` crates already run, extended to the
|
|
# workspace. `unsafe fn` marks a CONTRACT the caller must uphold; it is not a licence for the whole
|
|
# body to skip checking. Without this lint an `unsafe fn` body is unchecked end to end, so a 600-line
|
|
# function hides which handful of lines are actually the unsafe ones — exactly the reviewer-hostile
|
|
# shape we are working down. (This is the Rust 2024 default; adopting it early also pays off the
|
|
# edition migration.)
|
|
#
|
|
# `deny`, not `warn`. `warn` was never actually a softer setting: CI runs `cargo clippy … -D
|
|
# warnings`, which promotes it to a hard error anyway — that is how adopting this lint turned main
|
|
# red on every platform for a day without the level in this file ever saying `deny`. A level that
|
|
# lies about its own severity is worse than a strict one, so this now states what CI already does,
|
|
# and the exemptions are written down per file instead of hiding in a 689-warning wall nobody reads.
|
|
#
|
|
# THE EXEMPTIONS. Fourteen GPU/FFI backend files carry `#![allow(unsafe_op_in_unsafe_fn)]` with a
|
|
# one-line reason each. They are not "not done yet" — they are where this lint stops paying:
|
|
# their bodies are ash/CUDA/AMF/libav calls almost line for line (measured: 64% of the sites are a
|
|
# single third-party FFI call, and of the 44 `unsafe fn`s in them only 4 have a body containing no
|
|
# unsafe operation at all). Narrowing them means one `unsafe {}` per line plus, since pf-encode also
|
|
# denies `clippy::undocumented_unsafe_blocks`, one hand-written SAFETY comment per line that could
|
|
# only ever restate "an ash call on a live device" — the precise noise that made `unsafe` stop
|
|
# meaning anything here before (see the header of `pf-win-display/src/win_display.rs`).
|
|
#
|
|
# Everything else in the workspace is at zero and enforced. Removing one of those allows, file by
|
|
# file, is real work with a real payoff; blanket-narrowing all fourteen is not. Prefer DELETING an
|
|
# `unsafe fn` marker over wrapping its body: keep the marker only where a caller can actually break
|
|
# something (a raw pointer, a borrowed HANDLE, a GPU object that must not be in flight).
|
|
[workspace.lints.rust]
|
|
unsafe_op_in_unsafe_fn = "deny"
|
|
|
|
[profile.release]
|
|
opt-level = 3
|
|
lto = "thin"
|
|
codegen-units = 1
|
|
# NOTE: deliberately NOT `panic = "abort"`. punktfunk-core ships as a cdylib/staticlib into
|
|
# third-party apps (Swift/Kotlin/C) and its C ABI catches panics at the boundary
|
|
# (`catch_unwind` → `PunktfunkStatus::Panic`). `panic = "abort"` would make that guard a
|
|
# no-op and let a stray panic abort the embedding application. Unwinding keeps the
|
|
# documented isolation guarantee real.
|
|
|
|
# The per-frame hot path must stay fast even in dev builds.
|
|
[profile.dev.package."*"]
|
|
opt-level = 2
|