Compare commits

..
Author SHA1 Message Date
enricobuehler 6a506a8fa9 fix(vdisplay/driver,pf-frame): no punktfunk process holds REALTIME GPU priority by default
windows-drivers / probe-and-proto (pull_request) Successful in 30s
ci / bun-nix (pull_request) Successful in 56s
ci / docs-site (pull_request) Successful in 1m16s
ci / web (pull_request) Successful in 1m17s
ci / rust-arm64 (pull_request) Successful in 1m31s
windows-drivers / driver-build (pull_request) Successful in 1m47s
android / android (pull_request) Successful in 4m40s
ci / rust (pull_request) Successful in 9m54s
apple / swift (pull_request) Failing after 13m27s
apple / screenshots (pull_request) Skipped
The RX 9070 XT field A/B (2026-08-11/12 logs) convicted BOTH of our REALTIME
GPU-scheduling levers of generating the metronomic capture-stall class the
stall program has chased for weeks — compose-silence holes of 150-800 ms in
which ETW shows NO process presenting while the GPU stays responsive:

- the vdisplay driver's IddCxSetRealtimeGPUPriority raise beat at ~1.75-1.78 s
  (PFVD_NO_RT_GPU=1 alone removed that metronome: ~0.35 stalls/s metronomic ->
  10 sparse aperiodic over 3.9 min);
- the host auto-gate's HIGH->REALTIME upgrade (pf-frame dxgi.rs, T2.3) beat at
  ~3.58 s in the AV1 sessions where it promoted (vram_pct=1, 12:59:26); pinning
  PUNKTFUNK_GPU_PRIORITY_CLASS=high removed that residual too (13:45 session:
  zero metronomic, stall rate at the clean-run baseline).

Neither period matches any punktfunk clock: the full periodic-actor census
(driver: event-paced drain + 16 ms E_PENDING wait, 33 ms cursor poll, 3 s
watchdog reap; host: 250 ms descriptor poll, 5/50/100 ms probes + ~2 s scanline
retarget, 2 s VRAM gate, 2 s exclusive re-assert, 3.33 s pinger, 1 s stats,
~1 Hz phase-lock, fps/2 LTR marks) has nothing in the 1.69-2.29 s band, and
every host-side actor ran unchanged in the A/B that killed the fast metronome.
The periodicity is emergent from holding an unreachable-priority queue against
the WDDM scheduler on this AMD family (the period even differs by which of our
processes holds REALTIME); it is not a punktfunk cadence being amplified, so
there is nothing punktfunk-periodic to fix - the fix is to stop holding
REALTIME by default, which is also canonical parity (no shipping IDD raises
it, and HIGH was the class that delivered the original Sunshine-parity encode
win).

- Driver: PFVD_NO_RT_GPU (default-ON, opt-OUT) becomes the PFVD_RT_GPU ladder,
  default OFF on every vendor: unset = no raise (canonical IDD behavior);
  =thread = SetGPUThreadPriority(+7), a graduated in-band middle rung for field
  A/B (not default: unmeasured here, and the host measured the same call as "no
  help" for its own starvation case); anything else = the old REALTIME DDI.
  PFVD_NO_RT_GPU stays recognized and WINS over the opt-in, so the field boxes
  that carry it through the default-ON era keep meaning OFF. Both directions
  remain A/B-able without a rebuild (machine env + device restart). The CPU
  half of the original branch-2 hardening (MMCSS / TIME_CRITICAL) is untouched
  - it addressed the delivery holes that were actually observed.
- Host: PUNKTFUNK_GPU_PRIORITY_CLASS default auto -> high. `auto` (the gated
  REALTIME upgrade) stays available as an explicit opt-in, `realtime` still
  pins; unrecognized values now land on the HIGH default instead of silently
  opting into the gate - a typo must not buy the hazard. The VRAM/HAGS gate
  machinery is unchanged for `auto`; it guards the NVENC-hang hazard but cannot
  see this one.
- stall.rs: the no-OS-event METRONOMIC warning now carries rt_gpu_driver /
  rt_gpu_host fields (the machine-env state of both levers) and names clearing
  them as the FIRST cure, ahead of the display-hardware suspects - a field log
  self-answers the triage question this program just spent a week on.

No console policy axis for the driver knob: the lever is default-safe now, the
driver reads config at WUDFHost scope where machine env already matches the
device-restart lifecycle, and a policy axis would need pf-driver-proto churn
(or a device-key registry write) for an experimental lever that only exists to
be A/B-ed. If the `thread` rung ever proves out as a default-worthy raise,
that is the moment to revisit.
2026-08-12 13:57:20 +02:00
enricobuehler 66ba61b12c Merge pull request 'Safety round 2: ASAN+LSAN over the C ABI boundary, two soundness fixes, and WP4's AvFrame RAII' (#172) from worktree-safety-round-2 into main
audit / bun-audit (plugin-kit) (push) Successful in 24s
audit / bun-audit (sdk) (push) Successful in 24s
audit / bun-audit (web) (push) Successful in 27s
audit / cargo-audit (push) Successful in 34s
audit / docs-site-audit (push) Successful in 26s
apple / swift (push) Successful in 1m46s
audit / pnpm-audit (push) Successful in 1m16s
ci / rust-arm64 (push) Successful in 1m35s
ci / web (push) Successful in 1m25s
ci / docs-site (push) Successful in 1m22s
ci / bun-nix (push) Successful in 23s
audit / miri (push) Successful in 5m10s
audit / license-gate (push) Successful in 7m0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 15s
android / android (push) Successful in 8m2s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 9s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 8s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
deb / build-publish-client-arm64 (push) Successful in 2m34s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
audit / c-abi-asan (push) Successful in 7m51s
arch / build-publish (push) Successful in 8m42s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 16s
deb / build-publish (push) Successful in 5m49s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m25s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m39s
docker / builders-arm64cross (push) Successful in 10s
docker / deploy-docs (push) Failing after 1m4s
deb / build-publish-host (push) Successful in 8m54s
release / apple (push) Successful in 12m10s
windows-host / package (push) Successful in 13m41s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 20s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m19s
apple / screenshots (push) Successful in 5m52s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 4m0s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m14s
ci / rust (push) Successful in 13m52s
flatpak / build-publish (push) Successful in 15m17s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m54s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 18m39s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m41s
Reviewed-on: #172
2026-08-12 05:57:16 +00:00
enricobuehler 5002849737 feat(pf-encode): WP4 — AvFrame/AvSwsContext RAII across all three libav backends
ci / bun-nix (pull_request) Successful in 31s
ci / docs-site (pull_request) Successful in 1m17s
ci / rust-arm64 (pull_request) Successful in 1m45s
apple / swift (pull_request) Successful in 1m49s
apple / screenshots (pull_request) Skipped
ci / web (pull_request) Successful in 2m7s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m56s
android / android (pull_request) Successful in 6m21s
ci / rust (pull_request) Successful in 7m42s
windows / build (x86_64-pc-windows-msvc) (pull_request) Failing after 12m28s
Two newtypes beside AvBuffer/AvFilterGraph (same house shape: alloc/from_raw
rejects the allocator's null once, as_ptr lends, Drop frees, no Clone; NonNull
inside so the Options get a niche). 8 av_frame_alloc + 3 sws_getContext sites
converted; all 22 hand-placed av_frame_free and 5 sws_freeContext calls are
gone, and the three hand-written Drop impls (CpuInner, SystemInner,
NvencEncoder) with them.

The live defect this closes: ZeroCopyInner::submit (ffmpeg_win) leaked the
frame AND one pooled hwframe surface on each of three ? exits between the
pool pull and the send — under a SAFETY comment asserting no leak — and with
POOL=8, eight such failures starved the pool and wedged the encoder with no
error naming the cause. Every exit now returns the surface.

Drop-order care (the hidden cost the survey flagged): NvencEncoder's sws_csc
moved to field #1 (its hand-Drop freed it before all fields; this path runs
on every stall-watchdog recovery via *self = fresh); CpuInner's nv12/sws
declaration order flipped to match its hand-Drop; SystemInner's already
agreed. Pinned by FIELD ORDER comments, not offset_of asserts — the survey's
assert suggestion is the wrong tool: offset_of measures repr(Rust) memory
layout, which the compiler may reorder independently of the declaration
order that drop order actually follows.

The dmabuf path keeps its early descriptor release via an explicit drop()
at the exact point the hand-written free sat.

Gates: .25 clippy -D warnings + tests green (nvenc,vulkan-encode,pyrowave);
.133 check --all-targets + clippy --release -D warnings + 80 tests green
(nvenc,amf-qsv,qsv; test step needs ffmpeg\bin on Path — 0xC0000135
otherwise). Owed on hardware: the #[ignore]d alloc/drop cycles on
.136/.116/.173/.47 and the pool-exhaustion assertion (9th submit succeeds
after 8 forced failures).
2026-08-12 00:31:14 +02:00
enricobuehler 9a59504ba4 fix(punktfunk-core): validate InputKind before forming &InputEvent in the C ABI
abi.rs's two send-input entry points built &InputEvent straight out of
caller memory with ev.as_ref(); InputKind is repr(u8) with 16 valid
discriminants, so a C embedder writing ev->kind = 42 was immediate UB the
moment the reference formed — in a file whose stated principle is that
failures become status codes. New read_input_event() checks null, reads
the tag as a raw byte, validates through the same InputKind::from_u8 the
wire path uses, and only then forms the reference; bad tags return
InvalidArg. Every other field is a plain integer, valid for any pattern.

Test stages the event in MaybeUninit storage so the test itself never
holds a reference to the invalid value. 380 lib tests + the C harness
round-trip + clippy -D warnings green on .25; header regenerated.
2026-08-12 00:12:01 +02:00
enricobuehler e8c306b9c0 fix(punktfunk-host): WP3c/3d — align the TOKEN_USER buffer, make EqualSid fail closed
3c: forming &TOKEN_USER (align 8) out of a bare [u8; 256] (align 1) was UB
by the validity rule whenever the stack slot landed misaligned — shipped
codegen happened to 8-align it, which is luck, not a contract. Fixed with
a repr(align(8)) wrapper that keeps the buffer at 256 BYTES; the comment
records why [u64; 32] is the wrong shape (len() would silently become 32
and misclassify every hand-run host as SYSTEM via ERROR_INSUFFICIENT_BUFFER,
invisibly to a SYSTEM-side test). Length arg now size_of_val.

3d: EqualSid().is_ok() read BOTH 'SIDs differ' and 'EqualSid failed' as
Err, so a genuine failure yielded 'not SYSTEM' — the fail-OPEN direction,
contradicting the documented fail-closed contract. Now split three ways on
the last-error code, with SetLastError(0) cleared first so a stale value
cannot misclassify.

Gate: cargo check -p punktfunk-host + cargo clippy --release -D warnings
both green on .133 (real MSVC, fresh extraction, sentinel-verified).
2026-08-12 00:12:00 +02:00
enricobuehler c3b57438e1 chore(ci): c-abi-asan job in audit.yml — the harness under ASAN+LSAN, weekly + on demand
Same shape as the miri job (dated nightly, own san- cache prefixes,
non-blocking day one via a step-level ::warning::, a proved-it-ran grep).
run.sh gains PF_SAN_TOOLCHAIN so CI can pin its dated nightly — bare
+nightly would ask for the rolling channel the job never installs. Both
the pinned and vanilla paths re-verified green on .25.
2026-08-12 00:02:26 +02:00
enricobuehler e20b614059 chore(safety): PF_SAN sanitizer gate for the C ABI harness
PF_SAN=address builds the punktfunk-core staticlib on nightly with
-Zsanitizer/-Zbuild-std and the C harness with clang -fsanitize, so ASAN
instruments both sides of the boundary at once and LSAN (detect_leaks=1)
becomes the first automated check on abi.rs's Box::into_raw/from_raw leak
contract. Verified on the .25 box: green run passes byte-exact; deleting
one punktfunk_session_free() in the harness makes LSAN report the 308
Rust-side allocations behind the handle and the script exit 1.

The harness binary moves from mktemp to target/ — a debug+ASAN static
binary can exceed a tmpfs /tmp (it did, on .25's 3.6G tmpfs).
2026-08-11 23:58:06 +02:00
enricobuehler 6eb89b3f34 Merge pull request 'The lint ratchets (WP2b + WP2c): crate-level gaps closed, the three-workspace hoist, three blocking grep gates' (#171) from worktree-lint-ratchets into main
apple / swift (push) Successful in 1m39s
windows-drivers / probe-and-proto (push) Successful in 26s
ci / rust-arm64 (push) Successful in 2m18s
windows-drivers / driver-build (push) Successful in 2m15s
ci / web (push) Successful in 1m27s
android / android (push) Successful in 6m53s
ci / bun-nix (push) Successful in 1m3s
ci / docs-site (push) Successful in 1m52s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m14s
release / apple (push) Successful in 10m10s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m9s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m20s
deb / build-publish-client-arm64 (push) Successful in 6m37s
apple / screenshots (push) Successful in 6m16s
ci / rust (push) Successful in 14m7s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m33s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 51s
decky / build-publish (push) Successful in 33s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 8s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m23s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 10s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 9s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
deb / build-publish (push) Successful in 11m15s
deb / build-publish-host (push) Successful in 11m36s
arch / build-publish (push) Successful in 19m0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m3s
docker / builders-arm64cross (push) Successful in 19s
docker / deploy-docs (push) Successful in 6m39s
flatpak / build-publish (push) Successful in 8m5s
windows-host / package (push) Successful in 15m24s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 21s
nix / flake (push) Successful in 14m13s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m18s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m5s
2026-08-11 21:57:26 +00:00
enricobuehler bbd26ea82c Merge pull request 'Miri interprets the FFI-free leaf crates, one of them at MSVC layout' (#169) from worktree-miri-ci into main
ci / rust (push) Canceled after 0s
ci / rust-arm64 (push) Canceled after 0s
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
audit / cargo-audit (push) Successful in 42s
audit / bun-audit (sdk) (push) Successful in 31s
audit / bun-audit (plugin-kit) (push) Successful in 43s
audit / bun-audit (web) (push) Successful in 29s
audit / pnpm-audit (push) Successful in 10s
audit / docs-site-audit (push) Successful in 28s
audit / license-gate (push) Successful in 6m9s
audit / miri (push) Successful in 7m20s
2026-08-11 21:57:24 +00:00
enricobuehler 549fdf238b Merge branch 'main' into worktree-lint-ratchets
apple / swift (pull_request) Successful in 1m45s
apple / screenshots (pull_request) Skipped
windows-drivers / driver-build (pull_request) Successful in 2m1s
ci / rust-arm64 (pull_request) Successful in 1m54s
ci / web (pull_request) Successful in 1m13s
ci / docs-site (pull_request) Successful in 1m17s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m58s
ci / bun-nix (pull_request) Successful in 1m15s
ci / rust (pull_request) Successful in 6m1s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m44s
windows-drivers / probe-and-proto (pull_request) Successful in 26s
android / android (pull_request) Successful in 7m1s
nix / flake (pull_request) Successful in 16m48s
2026-08-11 21:57:17 +00:00
enricobuehler 3ea411fa39 Merge branch 'main' into worktree-miri-ci
ci / docs-site (pull_request) Successful in 1m16s
ci / bun-nix (pull_request) Successful in 2m10s
ci / web (pull_request) Successful in 2m55s
ci / rust-arm64 (pull_request) Successful in 3m21s
ci / rust (pull_request) Successful in 11m18s
2026-08-11 21:57:15 +00:00
enricobuehler 3b2fcd076d Merge pull request 'chore(api): regenerate openapi.json — #164's unpair change rewrote the unpairClient description without regenerating' (#170) from worktree-openapi-regen into main
ci / rust (push) Canceled after 31s
ci / rust-arm64 (push) Canceled after 29s
ci / web (push) Canceled after 26s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 17s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 17s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 15s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 13s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
2026-08-11 21:57:07 +00:00
enricobuehler d67ab9ede4 chore(safety): two .133 gate findings — cfg the abi lock helper, re-anchor a layer proof
ci / bun-nix (pull_request) Successful in 40s
ci / web (pull_request) Successful in 1m24s
ci / docs-site (pull_request) Successful in 1m30s
android / android (pull_request) Canceled after 1m45s
apple / swift (pull_request) Canceled after 1m41s
apple / screenshots (pull_request) Canceled after 0s
ci / rust (pull_request) Canceled after 1m43s
ci / rust-arm64 (pull_request) Canceled after 1m43s
nix / flake (pull_request) Canceled after 1m29s
windows-drivers / probe-and-proto (pull_request) Canceled after 0s
windows-drivers / driver-build (pull_request) Canceled after 1m25s
windows / build (aarch64-pc-windows-msvc) (pull_request) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (pull_request) Canceled after 0s
The tray leg builds punktfunk-core with default-features off, where
lock_recover's only callers (the quic-gated punktfunk_connection_* entry
points) do not exist — dead code under -D warnings. The helper takes the
same feature gate.

In pf-vkhdr-layer, rustfmt had reflowed destroy_surface's lookup into a
multiline closure, leaving the SAFETY comment outside the closure that
contains its unsafe block — the box's clippy rightly stopped accepting the
adjacency. The comment moves inside, directly above the block.
2026-08-11 23:51:01 +02:00
enricobuehler abec2a1457 chore(safety): nvenc_core — plain union-arm writes are safe by language rule
The .25 gate corrected the carve-out: rustc flags `unsafe { u.arm.field = x }`
as unused_unsafe — plain assignment through a union projection is safe
(writing an arm cannot itself be UB; the hazard is the mismatched READ).
The 11 plain writes go back to bare statements under their codec matches.
What stays in per-op unsafe blocks with arm-guard proofs is the real unsafe
surface: union reads, borrows, and the bindgen bitfield-setter calls — which
is exactly the surface the shipped 4:4:4 bug lived on (set_chromaFormatIDC
stamped under a wrong codec).
2026-08-11 23:45:36 +02:00
enricobuehler 5f097d530d chore(safety): exempt the two bindings-only sys crates from the hoisted deny
Linux fallout from the hoist the mac could not see: bindgen emits unsafe
blocks (layout tests/accessors) into OUT_DIR, where nobody hand-writes
SAFETY proofs — pyrowave-sys failed clippy on .25 with 17 of them, and
libvpl-sys would do the same on the Windows leg. Both crates are
bindings-only by charter (the safe wrapper lives with the consumer), so the
allow is crate-wide with the rationale at the crate root; the hand-written
link-sanity tests keep their proofs by convention.
2026-08-11 23:41:39 +02:00
enricobuehler 2bfd1cd2d5 chore(safety): three unsafe-hygiene grep gates, blocking in ci.yml (WP2c gates)
scripts/ci/check-unsafe-hygiene.sh — textual gates for three classes no lint
covers:

A. unsafe fn markers carrying no contract. unsafe_op_in_unsafe_fn forces real
   ops into blocks, so an unsafe fn with no `unsafe` in its body is a marker
   with no contract (db659809 found two by hand). Contract-deferring fns
   (Vec::set_len shape) waive with `// unsafe-fn-no-op-ok: <reason>`; fenced
   files and `unsafe extern "ABI" fn` (signature-mandated markers) are
   skipped structurally.

B. unwrap/expect/panic! inside extern "C"/"system" bodies — an abort since
   Rust 1.81, not linted, not fuzzable (8b98d0b3). catch_unwind bodies are
   exempt; `// panic-in-extern-ok: <reason>` waives a deliberate abort.

C. Safe-but-process-global APIs (env::set_var/remove_var, sigaction,
   setlocale, set_current_dir) — the 972af299 environ race lived in a file
   with zero occurrences of the word `unsafe`. Per-file count ratchet with
   the baseline in the script; any increase or new file fails.

Making gate B clean on main surfaced 14 real instances of exactly its class —
`.lock().unwrap()` in unguarded extern fns, where a poisoned mutex aborts the
embedding process: six punktfunk-core abi.rs entry points (poll_frame,
next_au, next_audio, next_audio_pcm, next_cursor_shape, next_clipboard),
seven Android JNI entry points, and the Windows client's deeplink wnd_proc.
All fixed with poison-recovering locks (the slots are last-value caches,
valid whatever a poisoned writer left) and Option::insert for the
set-then-unwrap shape; punktfunk-core's 203 lib tests pass. Gate A's
findings were six genuine contract-deferring fns — waived with reasons, not
fixed, because the markers are correct.

Gate-of-the-gate: all three shown to FAIL on deliberately planted instances
(marker fn, panicking extern callback, env::set_var in an unlisted file) and
to run clean on the tree, before the ci.yml step made them blocking.
2026-08-11 23:36:17 +02:00
enricobuehler dfebb9dfbb chore(safety): hoist the unsafe lints into the workspace tables (WP2c hoist)
undocumented_unsafe_blocks joins unsafe_op_in_unsafe_fn in
[workspace.lints], and the ~100 scattered per-file #![deny(...)] attributes
(85 files) are deleted — a new crate, or a new module in an old one, is now
covered on creation rather than on remembering. The per-file form is how
pf-vkhdr-layer, wdk-probe and half of pf-clipboard stayed uncovered.

There are THREE workspaces, so the claim is made three times: the main
Cargo.toml, packaging/windows/drivers (workspace table + [lints]
workspace = true in all seven members), and packaging/windows/pf-vkhdr-layer
(its [lints] table, previous commit). pf-update now opts into workspace
lints; the two vendored member snapshots (cros-codecs, usbip-sim) stay out
deliberately and now both say so.

Newly-covered fallout was two link-sanity tests (pyrowave-sys, libvpl-sys)
— proofs written. Stale prose that claimed the workspace held
unsafe_op_in_unsafe_fn at "warn" (it has been deny) or pointed at the
deleted attributes is corrected.

nvenc_core.rs is carved OUT of the unsafe_op_in_unsafe_fn fence: its
exemption rationale ("raw entry-table calls almost line for line") was
false — the file makes zero FFI calls. Its unsafe surface is C-union writes
whose soundness hangs on which codec arm is active, and its own 4:4:4 note
records the shipped bug (hevcConfig bytes stamped onto an AV1 config) that
per-operation blocks make visible. It now runs the strictest discipline in
the crate: clippy::multiple_unsafe_ops_per_block at deny, one union access
per block, each naming its codec guard.

Verified here: cargo fmt clean in all three workspaces; native clippy
-D warnings clean for everything that compiles on macOS (the three
pre-existing mac-native failures — pf-client-core wol.rs, pf-encode
dead-code/closure-call, probe mic_burst — reproduce on the clean tree).
Linux/Windows legs ride the .25/.133 gate.
2026-08-11 23:26:28 +02:00
enricobuehler f675b3710e chore(api): regenerate openapi.json — fcf4c9fd rewrote the unpairClient description without regenerating
ci / web (pull_request) Successful in 1m48s
ci / bun-nix (pull_request) Successful in 20s
ci / docs-site (pull_request) Successful in 2m22s
ci / rust-arm64 (pull_request) Successful in 4m3s
ci / rust (pull_request) Successful in 6m58s
2026-08-11 23:23:08 +02:00
enricobuehler 23fa03b051 chore(safety): close the three crate-level lint gaps (WP2b)
pf-vkhdr-layer — the sharpest gap: an implicit layer injected into every
Vulkan game process, 32 unsafe usages, zero SAFETY comments, own workspace so
no lint table reached it, and an explicit missing_safety_doc allow. Now: a
[lints] table (unsafe_op_in_unsafe_fn + clippy::undocumented_unsafe_blocks,
both deny), the allow removed, every unsafe operation in an explicit block
with a real proof (loader layer protocol / Vulkan valid-usage), # Safety docs
on the contract-carrying fns, const layout asserts for the SurfaceFormat2Raw
mirror, the five helpers with no caller-facing contract demoted to safe fns,
and the two redundant `unsafe impl Send` deleted (fn pointers and vk handles
are Send intrinsically — the type-check proves it).

wdk-probe — 21 unsafe blocks, 12 proofs: the 9 missing SAFETY comments are
written (the iddcx_rt.rs DDI slot-dispatch ones are about table population
and PFN/index pairing, not pattern fill), the sibling denies added at the
crate root, missing_safety_doc allow dropped, # Safety on DriverEntry, and
the crate joins windows-drivers.yml's clippy list — it was the only driver
crate not in it.

pf-clipboard — the undocumented_unsafe_blocks deny moves from host/windows.rs
to the crate root so host/wayland.rs (4 blocks), host/mutter.rs (2) and any
future backend under host/ are covered on creation. All existing blocks
already carry proofs; free today, structural tomorrow.

Verified here: pf-vkhdr-layer cargo fmt --check + clippy --release
-D warnings at x86_64-pc-windows-msvc. wdk-probe and pf-clipboard compile
checks need the WDK/Linux boxes and ride the .133/.25 gate.
2026-08-11 23:10:00 +02:00
enricobuehler 6e4638dab5 ci(audit): interpret the FFI-free leaf crates under Miri, one at MSVC layout
ci / web (pull_request) Successful in 59s
ci / rust-arm64 (pull_request) Successful in 2m3s
ci / bun-nix (pull_request) Successful in 1m17s
ci / docs-site (pull_request) Successful in 2m48s
ci / rust (pull_request) Failing after 8m8s
Adds a non-blocking `miri` job to audit.yml, per rust-safety-programme.md §7.

What it buys is one narrow, real thing: pf-driver-proto interpreted CROSS-COMPILED to
x86_64-pc-windows-msvc, on a Linux runner, with no Windows box in the loop. That crate is
`#![forbid(unsafe_code)]` and path-dep'd by BOTH the main workspace and the driver
workspace, so it is the layout oracle for every frame and IOCTL crossing that boundary,
and nothing else in CI checks it at MSVC layout. It is NOT unsafe coverage — Miri can
execute on the order of 2% of the host's unsafe and cannot run ash, windows-rs, ffmpeg,
CUDA or the WDK — so no "Miri coverage" number is reported anywhere.

Three steps, every one of them measured on 192.168.1.25 with a cold target dir and cold
sysroot cache, on the dated toolchain the job installs, BEFORE being committed:

  step A  pf-driver-proto + pf-host-config + pf-gpu   21 + 12 + 4 pass   43 s
  step B  pf-driver-proto @ x86_64-pc-windows-msvc           21 pass     26 s
  step C  punktfunk-core fec::gf8 with +avx2,+ssse3            2 pass     63 s

Four corrections to the §7.3 job spec, found while doing this and folded into comments:

* `-p punktfunk-core fec packet crypto` does not parse — cargo rejects the extra
  positionals. Corrected (filters after `--`) it selects 63 tests and was killed at a
  25-minute cap with not one test complete, so the bulk step is dropped entirely and only
  the narrow `fec::gf8` selection is kept, timed at 63 s.
* `nightly-2026-08-10` resolves to rustc 1.99.0-nightly (969b803cb 2026-08-09), NOT the
  12c36e253 2026-08-10 the doc cites: `nightly-<date>` names the day rustup PUBLISHED the
  build, which is compiled from the previous day's commit. The doc's hash came from the
  ROLLING `nightly` channel and was mislabelled. All three steps were re-run and are green
  on the dated pin actually installed here.
* fec-rs dispatches its GF(2^8) multiply through RUNTIME `is_x86_feature_detected!`, so
  step C's RUSTFLAGS are load-bearing in both directions. Verified by probe: bare,
  avx2=false and the step would silently interpret the scalar fallback; with the flags,
  avx2=true and `_mm256_shuffle_epi8` genuinely executes under the interpreter. GFNI stays
  false either way, so that branch is simply not covered.
* `RUSTC_WRAPPER: ""` is a guard, not a fix, and the comment says so — audit.yml sets no
  sccache today, and cargo-miri warns "Ignoring `RUSTC_WRAPPER`, Miri does not support
  wrapping" and carries on regardless.

Non-blocking via a step-level `||`, not job-level continue-on-error, following the
precedent audit.yml already documents for docs-site-audit. Each step additionally asserts
a non-zero pass count, so a crate rename or a filter that stops matching surfaces as a
warning rather than as a green zero-test run. Both paths were exercised directly: a
failing run emits the annotation and still exits 0, and a zero-selection run trips the
guard, while a green run with empty bin/doctest targets does not false-positive.

Leak checking stays ON (no -Zmiri-ignore-leaks); the two deliberate leaks in the tree are
named in a comment so whoever expands coverage annotates those sites instead of blanket-
disabling the check. pf-bitstream and the FFI crates are excluded with the reasons inline
so they are not helpfully re-added. `paths:` is deliberately not widened to
crates/pf-driver-proto/** — that filter is workflow-level and would fire all six audit
jobs on every driver-proto edit.
2026-08-11 22:52:28 +02:00
enricobuehler 0c2ac333ae Merge pull request 'Chore/rust safety programme' (#164) from chore/rust-safety-programme into main
apple / swift (push) Successful in 1m43s
ci / web (push) Successful in 1m33s
windows-drivers / probe-and-proto (push) Successful in 27s
ci / docs-site (push) Successful in 1m44s
ci / bun-nix (push) Successful in 21s
windows-drivers / driver-build (push) Successful in 1m45s
ci / rust-arm64 (push) Successful in 7m52s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 21s
deb / build-publish-client-arm64 (push) Successful in 2m35s
arch / build-publish (push) Successful in 11m28s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 18s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 19s
release / apple (push) Successful in 9m41s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 14s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 14s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 15s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 52s
ci / rust (push) Failing after 11m41s
android / android (push) Successful in 14m57s
deb / build-publish-host (push) Successful in 7m17s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 3m13s
apple / screenshots (push) Successful in 6m20s
docker / builders-arm64cross (push) Successful in 18s
docker / deploy-docs (push) Successful in 44s
nix / flake (push) Successful in 14m52s
windows-host / package (push) Successful in 20m38s
windows-host / winget-source (push) Skipped
deb / build-publish (push) Successful in 13m0s
windows-host / canary-manifest (push) Successful in 49s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m28s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m13s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m24s
flatpak / build-publish (push) Successful in 24m54s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m33s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 39m36s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 38m33s
2026-08-11 20:47:15 +00:00
enricobuehler bc70a58fb1 Merge main into chore/rust-safety-programme
windows-drivers / probe-and-proto (pull_request) Successful in 22s
ci / bun-nix (pull_request) Successful in 36s
ci / web (pull_request) Successful in 1m17s
apple / swift (pull_request) Successful in 1m39s
apple / screenshots (pull_request) Skipped
windows-drivers / driver-build (pull_request) Successful in 1m58s
ci / rust-arm64 (pull_request) Successful in 3m18s
ci / docs-site (pull_request) Successful in 3m55s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m29s
android / android (pull_request) Successful in 4m46s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m30s
ci / rust (pull_request) Failing after 10m11s
nix / flake (pull_request) Successful in 15m6s
Two conflicts: the test-module import list in gamescope.rs (union — the branch's takeover-state
tests and main's WSI opt-out tests both stay), and next_frame_timed_out in pf-capture, where the
branch still carried the pre-#168 else-if chain — resolved to main's match-based refactor, which
already embeds the same arm semantics plus the provisional-budget latch gate.
2026-08-11 22:34:38 +02:00
enricobuehler ce25aca7bd Merge pull request 'Two black screens from the .41 field session — a NO_FOCUS window stole the composite, and one truncated timeout downgraded the host forever' (#168) from worktree-blackscreen-fixes into main
apple / swift (push) Successful in 1m42s
ci / web (push) Successful in 1m14s
ci / docs-site (push) Successful in 1m20s
ci / bun-nix (push) Successful in 19s
deb / build-publish-client-arm64 (push) Successful in 1m46s
deb / build-publish (push) Successful in 5m9s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 6s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 8s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 10s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 5s
apple / screenshots (push) Canceled after 1m18s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 42s
deb / build-publish-host (push) Successful in 7m18s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m28s
arch / build-publish (push) Successful in 11m28s
android / android (push) Canceled after 6m31s
ci / rust (push) Canceled after 2m21s
ci / rust-arm64 (push) Canceled after 1m13s
docker / builders-arm64cross (push) Successful in 11s
docker / deploy-docs (push) Canceled after 5s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 4m10s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 3m2s
windows-host / package (push) Canceled after 0s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
2026-08-11 20:32:16 +00:00
enricobuehler 5587699a85 Merge pull request 'Every pinned card gets a library, and it launches with that card's profile' (#167) from worktree-console-pinned-profile-library into main
apple / swift (push) Successful in 1m40s
android / android (push) Canceled after 0s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 0s
ci / rust (push) Canceled after 0s
ci / rust-arm64 (push) Canceled after 1m47s
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 19s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 56s
deb / build-publish (push) Canceled after 0s
deb / build-publish-host (push) Canceled after 0s
deb / build-publish-client-arm64 (push) Canceled after 1m55s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 28s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 55s
docker / builders-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 11s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 11s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m34s
release / apple (push) Successful in 9m29s
flatpak / build-publish (push) Canceled after 13m45s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m59s
windows / build (aarch64-pc-windows-msvc) (push) Canceled after 7s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 0s
2026-08-11 20:30:03 +00:00
enricobuehler c946fcdcb5 Merge pull request 'build(web): silence rollup's "use client" directive warnings in the nitro pass' (#166) from build/web-silence-rollup-directive-warnings into main
arch / build-publish (push) Canceled after 47s
ci / bun-nix (push) Successful in 23s
ci / rust (push) Canceled after 42s
ci / docs-site (push) Canceled after 47s
ci / rust-arm64 (push) Canceled after 59s
ci / web (push) Canceled after 58s
deb / build-publish (push) Canceled after 5s
deb / build-publish-host (push) Canceled after 53s
deb / build-publish-client-arm64 (push) Canceled after 43s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 15s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 2s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 5s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 1m28s
windows-host / package (push) Canceled after 3m49s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
2026-08-11 20:29:36 +00:00
enricobuehler fcf4c9fd63 fix(mgmt): unpair now revokes a LIVE session on both planes
ci / bun-nix (pull_request) Successful in 37s
ci / web (pull_request) Successful in 1m38s
apple / swift (pull_request) Successful in 1m50s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m38s
ci / docs-site (pull_request) Successful in 2m32s
windows-drivers / driver-build (pull_request) Successful in 1m50s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m29s
android / android (pull_request) Successful in 6m31s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m54s
ci / rust (pull_request) Failing after 10m42s
nix / flake (pull_request) Canceled after 8m26s
windows-drivers / probe-and-proto (pull_request) Canceled after 0s
An unpair removed the certificate but left the revoked client's running
session streaming until the client chose to leave. Now it is a complete
revocation:

- GameStream: when the removed certificate owns the active launch, the
  session is quit_session'd — the ENet control thread's ended-session arm
  gives the client the standard TERMINATION+disconnect. (An owner-less
  launch cannot be attributed and is left to the WP0 port teardown when the
  last pairing goes.) The endpoint docstring's long-standing caveat
  ('removes the client from the listing without severing its ability to
  reconnect') is retired: TLS handshakes complete by design, authorization
  is per-request, and a live session no longer survives its own revocation.
- Native: session_status::stop_by_fingerprint signals the unpaired
  client's live session(s) to tear down deliberately (quit+stop), matched
  by the registry's client label — the fingerprint's 12-hex-char prefix for
  every pairable client; anonymous/TOFU sessions carry IP labels and are
  never touched (they have no pairing to revoke).

(The unpair-didn't-PERSIST half of 'unpairing was broken' was already fixed
in 13d57210 — save_paired was never called; this closes the other half.)

Gates: Linux amd64 both flavors clippy --all-targets -D warnings clean;
session_status 2/2 (new revocation test), the extended paired-clients test
green in both flavors, native_pairing test green.
2026-08-11 22:17:41 +02:00
enricobuehler 6ca192b9ab fix(packaging/gamescope): +pfhdr6 — a GAMESCOPE_NO_FOCUS window can no longer steal the composite
ci / web (pull_request) Successful in 1m3s
ci / rust-arm64 (pull_request) Successful in 1m35s
apple / swift (pull_request) Successful in 1m40s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m16s
ci / bun-nix (pull_request) Successful in 1m30s
ci / rust (pull_request) Successful in 4m52s
android / android (pull_request) Successful in 5m36s
Patch 0008: honor GAMESCOPE_NO_FOCUS in steamcompmgr's focus selection. hhd (Handheld Daemon)
sets the atom once at init on its hhd-ui overlay window and never clears it; MangoHud sets it
too; show/hide for these clients runs over the STEAM_OVERLAY protocol. NOTHING consumed the atom
— not upstream gamescope, not Bazzite's fork (checked ba148 by strings) — so a
mapped-but-unpainted hhd-ui window (it crash-loops under a headless punktfunk takeover and remaps
on every respawn, stamping Steam's appid 769) was an ordinary focus candidate, and steamcompmgr
picked it over Big Picture. The composite, and the stream fed from it, went black while every
health signal stayed green: on .41 the client sat decoding 60 fps at 0.1 Mb/s of black,
GAMESCOPE_FOCUSED_WINDOW named the hhd-ui window with GAMESCOPE_NO_FOCUS(CARDINAL)=1 on it, and
killing hhd-ui brought the picture back the same second.

The patch wires the atom exactly like GAMESCOPE_EXTERNAL_OVERLAY — read at map,
PropertyNotify-tracked with MakeFocusDirty, skipped by both focus-candidate collectors (X11 and
XDG) — and touches neither compositing nor appID, so a NO_FOCUS window still paints if the
baselayer protocol brings it into view; it is only barred from being CHOSEN. Applies cleanly on
the full 0001..0008 series from the bare 5fb8dce4 pin (verified with git am).

Banner +pfhdr5 → +pfhdr6, PKGBUILD 3.16.25.pfhdr6-1; README gains the 0008 row, the missing
+pfhdr5 ledger row, and the reconciled bump rule (a bugfix bumps the level only when field triage
must read the difference off a box's banner — 0007's crash-loop, 0008's lost composite).
2026-08-11 22:07:24 +02:00
enricobuehler 022ede651f fix(pf-capture): the truncated first attempt no longer latches the sticky downgrades
The pipeline retry loop deliberately shortens its first attempt's first-frame wait to 2.5s so a
stream bound during a gamescope re-init fails over quickly. But the portal capturer's timeout
diagnosis treated EVERY expiry as a verdict: it latched whichever offer it implicated — HDR
capture off for the source, the raw-dmabuf offer off, the EGL→CUDA offer off — process-wide and
permanently, when the attempt was truncated by design and a gamescope cold start routinely
delivers nothing inside that window while accepting every offer a few seconds later (observed on
.41: pid 1962 hit the expiry at connect and every later session in that process ran silently
degraded). This is bug #6 from the pf-capture sweep, verified then and unfixed until now.

The truncated attempt is now declared PROVISIONAL end to end: a new
`Capturer::next_frame_within_provisional` (default: delegates) lets the retry loop say "this
budget is the schedule, not a verdict", and the portal capturer's timeout classification — split
out as the pure `classify_first_frame_timeout` + `timeout_convicts`, with tests — names the same
suspect in the error text but latches nothing unless the expired budget was full-length.
2026-08-11 22:06:31 +02:00
enricobuehler 9c6e06d3b9 feat(host): GameStream is now a cargo feature — WP19, compile-time isolation
A new 'gamestream' feature (default ON — every stock package is behaviorally
identical, and GameStream stays runtime-opt-in via --gamestream /
PUNKTFUNK_GAMESTREAM) gates the whole Moonlight-protocol surface: control
(the ENet plane), rtsp, nvhttp, pairing, serverinfo, the _nvstream mDNS
advert, the compat media path (stream/video/audio), pen/gamepad/input
decode, apps, crypto, cert (the RSA identity), and tls's
Moonlight-client-cert leniency. AppState keeps the shared vocabulary
unconditional and cfg-gates the Moonlight-only fields; the mgmt API's PIN
endpoints (routes, handlers, OpenAPI entries, lane classifications, tests)
exist only under the feature.

Building --no-default-features --features pyrowave yields the hardened
NATIVE-ONLY host: no rusty_enet (the c2rust-transpiled C ENet stack, 158
unsafe sites) and no rsa (the identity split's legacy fallback became a
pem-only read — rustls/ring serves an existing RSA cert without the crate —
so the accepted Marvin advisory no longer applies to native-only builds).
Both claims are ASSERTED, not assumed: a new CI leg keeps the native-only
flavor clippy-clean and fails if cargo tree finds either crate in its graph.
serve --gamestream (or the env knob) against such a binary refuses to start
with a clear error rather than serving less than the operator configured.

En route: the logs-paging test assumed a quiet process-global log ring
between its cursors and raced other tests' legitimate log lines (the
identity tests added new emitters) — it now asserts on its own markers
within the page.

Gates: Linux amd64 — BOTH flavors clippy --all-targets -D warnings clean;
default tests identity 3/3, mgmt 37/37, gamestream 59/59; native-only tests
identity 3/3, mgmt 35/35, residue 4/4; rusty_enet+rsa absent native-only,
present default. .133 Windows — both flavors clippy clean (clean-first,
sentinel-checked), tree claims hold, and the WP0 port-lifecycle functional
gate PASSES on the default build.
2026-08-11 22:05:30 +02:00
enricobuehler 1009e14a44 build(web): silence rollup's "use client" directive warnings in the nitro pass
ci / bun-nix (pull_request) Successful in 31s
ci / docs-site (pull_request) Successful in 1m17s
ci / web (pull_request) Successful in 1m19s
ci / rust-arm64 (pull_request) Successful in 2m52s
ci / rust (pull_request) Successful in 7m9s
The nitro server build re-bundles the whole dep tree (`noExternals: true`), so
every React package shipping a `"use client"` banner earns a MODULE_LEVEL_DIRECTIVE
warning — ~150 locally, ~800 in CI — which buries the warnings worth reading.

Ignoring the banner is correct rather than papered over: this bundle is the
Bun/Nitro server, not an RSC module graph, and TanStack Start splits client from
server with its own transform, so nothing downstream consults it.

Supplying `onwarn` replaces nitro's own handler, so its three filters
(CIRCULAR_DEPENDENCY, EVAL, "Unsupported source map comment") are restated.

Verified: `bun run build` drops from 148 such lines to 0 with no other log
delta; `tsc --noEmit` and `biome check` clean.
2026-08-11 21:00:19 +02:00
enricobuehler e658ad726b feat(host): the identity split — the native planes get their own P-256 identity
One RSA-2048 identity served every plane, because Moonlight mandates RSA and
the planes grew out of the GameStream host. The native punktfunk/1 QUIC plane
and the management API now share a separate ECDSA P-256 identity
(native-cert.pem/native-key.pem, src/identity.rs): ring-generated via rcgen
(no rsa crate on the native path — the accepted Marvin advisory stops
applying once WP19 gates the compat planes), real SANs (localhost, loopback,
machine hostname — the legacy cert had none), and browser-compatible on
purpose: Ed25519 was rejected because no mainstream browser accepts an
Ed25519 server cert and /api/docs is opened in one. GameStream keeps the RSA
identity untouched (Moonlight pins it; its pairing hashes bind its X.509
signature bytes).

Migration is pin-preserving by construction. Clients TOFU-pin ONE leaf-DER
SHA-256 for both QUIC and the mgmt/library API, so the identity is resolved
ONCE in serve (the planes cannot race the first-run mint) under the rule:
identity files exist → use them; else the native trust store is EMPTY →
mint P-256 (fresh installs); else keep presenting the legacy RSA cert the
paired clients pinned, and log the migration path (unpair all, restart,
re-pair). Fingerprint pinning is algorithm-agnostic — existing shipped
clients pair against P-256 hosts unchanged.

Followers updated: the tray's loopback pin and the plugin SDK's mgmt CA
prefer native-cert.pem → cert.pem; the Windows runner ACL grant lists both
(the grant loop tolerates absent files). The in-process native tests now run
on an EPHEMERAL identity — they previously read, and would newly have
MINTED, identity files in the real config dir, which on a dev box that is
also a live host would have switched its identity and stranded every pinned
client.

Gates: Linux amd64 clippy --all-targets -D warnings clean (host+tray);
identity 2/2, mgmt 37/37, control 6/6, native 68/68 (C-ABI roundtrips over
the ephemeral identity). .133 Windows clippy clean; the port-lifecycle gate
re-run PASSES with the split live — the fresh host minted P-256 and served
mgmt over it (curl 200/204), ports tracked the paired list as before.
2026-08-11 20:52:01 +02:00
enricobuehler 23d0452157 feat(host): GameStream opt-in on every route; the native plane is deny(unsafe_code)-enforced
The user direction after WP0: ENet exists only for Moonlight, so the native
plane must be provably safe and the compat planes a deliberate choice.

Opt-in, everywhere. Windows already was (unchecked installer task). The three
opt-out surfaces are flipped: the shipped systemd user unit (deb/RPM/Arch/
sysext) no longer bakes --gamestream into ExecStart — a new
PUNKTFUNK_GAMESTREAM=1 host.env knob (pf-host-config, OR-ed with the CLI
flag) is the packaged opt-in; the NixOS module default goes true→false, with
a module-check assertion that unset = native-only; the Deck installer takes
--gamestream to opt in (--no-gamestream kept as explicit-off). Docs
(quickstart, running-as-a-service, moonlight, ubuntu/fedora/arch firewall
sections, gnome/sway, how-it-works) rewritten to the opt-in shape; the
CHANGELOG carries the upgrade note.

Enforced-safe. punktfunk-core is #![deny(unsafe_code)] crate-wide — every
module that parses network bytes is safe Rust as a compile error, not a
census result. Carve-outs are exactly two documented classes, neither of
which interprets attacker bytes: the client surface (abi, client) and the
transport syscall-batching shims (udp/{apple,linux,windows}, qos_windows).
In punktfunk-host, the modules a secure-default host exposes — native
(cfg-not-test: its tests exercise the client C ABI on purpose),
native_pairing, mgmt, mgmt_token, discovery, wol — are #[forbid(unsafe_code)].

Gates: Linux amd64 container clippy --all-targets -D warnings clean over
core+host-config+host; core 204 tests green under the deny; mgmt 46/46,
control 6/6. .133 Windows clippy (shipped features, clean-first,
sentinel-checked) clean — covers the qos_windows/udp-windows carve-outs.
macOS + iOS cargo check green (the apple.rs carve-out compiles for real).
2026-08-11 20:21:16 +02:00
enricobuehler 13d5721049 feat(gamestream): the ENet control port exists only while a pairing does (WP0)
rusty_enet — a c2rust-style transpile of C ENet, 158 unsafe sites — parsed
unauthenticated UDP on 47999 from GameStream startup, before any client had
ever paired: the host's entire pre-auth-reachable unsafe surface. Pairing
itself is HTTPS on nvhttp and never touches the port, so it now binds only
while the paired-client list is non-empty: a Gate in control.rs reconciles
the port to the list (armed only under --gamestream), pairing phase 4 brings
it up before the new client can /launch, and removing the last pairing tears
it down — a live client gets the same termination+disconnect farewell as a
host-side session end. A never-paired host on a hostile LAN exposes no ENet.

En route: the management API's unpair never called save_paired, so a restart
resurrected the client — and would now have silently re-opened the port; it
persists (the test now runs against a throwaway PUNKTFUNK_CONFIG_DIR so it
can't clobber a real paired.json). rusty_enet is pinned =0.4.0 per the WP,
left to the cargo-audit job to flag advisories against it.

Gate (amd64 container): clippy --all-targets -D warnings clean;
gamestream::control 6/6; mgmt::tests 37/37 incl. the regenerated
api/openapi.json. On-box .133 verification (ports/pair/stream) still owed.
2026-08-11 19:29:50 +02:00
enricobuehler d4366e7464 fix(pf-encode): the Vulkan extension probe walked a driver-filled array with no bound
ci / web (pull_request) Successful in 1m20s
apple / swift (pull_request) Successful in 1m38s
apple / screenshots (pull_request) Skipped
windows-drivers / driver-build (pull_request) Successful in 1m49s
windows-drivers / probe-and-proto (pull_request) Successful in 27s
ci / docs-site (pull_request) Successful in 1m22s
ci / rust-arm64 (pull_request) Successful in 2m50s
ci / bun-nix (pull_request) Successful in 24s
android / android (pull_request) Successful in 4m37s
ci / rust (pull_request) Successful in 10m25s
`ext_advertised` did `CStr::from_ptr(e.extension_name.as_ptr())` over a
driver-filled `[c_char; VK_MAX_EXTENSION_NAME_SIZE]`, and `vk_build.rs` open-coded
the identical call a second time. Neither had an in-Rust bound: a driver that
fills all 256 bytes without a NUL runs the walk into the NEXT
`ExtensionProperties`, and on the LAST element past the allocation.

The SAFETY comment asserted the spec guarantee ("a spec-guaranteed NUL-terminated
byte array") instead of enforcing it. That is the defect class this programme
keeps finding: a proof that restates what the other side promised rather than
checking it. Vulkan drivers are exactly the other side.

The bounded answer already shipped in the same crate — `pyrowave.rs:210` uses
ash's `extension_name_as_c_str()` for the identical job. It stops at
VK_MAX_EXTENSION_NAME_SIZE and returns Err when there is no terminator, so a
malformed entry is a non-match instead of an overrun. Both sites now route
through the one helper, which is no longer unsafe at all.

Deletes 2 unsafe operations and one duplicated walk.

⚠ The pre-existing test could not have caught this: it only ever built
well-formed, NUL-terminated entries. Added a case whose LAST element is 256
non-NUL bytes — the exact shape that used to leave the array — and a
prefix-match case, so the bound is now asserted rather than assumed.

Verified on 192.168.1.25 (Ubuntu, cargo 1.96.0 — the pinned toolchain):
  cargo check  -p pf-encode --features vulkan-encode,pyrowave --locked      ok
  cargo test   -p pf-encode --features vulkan-encode,pyrowave ext_advertised
                                                              2 passed / 0 failed
  cargo clippy -p pf-encode --all-targets --locked
        --features vulkan-encode,pyrowave -- -D warnings                    clean
Linux-only code (`enc/linux/`), so the Windows leg is unaffected.
2026-08-11 16:34:34 +02:00
enricobuehler cd72f77a3c fix(pf-encode): the AMF layout guards broke Windows clippy — 0*SLOT and 1*SLOT
windows-drivers / probe-and-proto (pull_request) Successful in 29s
apple / swift (pull_request) Successful in 1m39s
apple / screenshots (pull_request) Skipped
windows-drivers / driver-build (pull_request) Successful in 2m10s
ci / rust-arm64 (pull_request) Successful in 2m53s
ci / web (pull_request) Successful in 1m15s
android / android (pull_request) Successful in 4m44s
ci / bun-nix (pull_request) Successful in 26s
ci / rust (pull_request) Canceled after 5m25s
ci / docs-site (pull_request) Canceled after 1m7s
`27f08340` wrote every vtable offset assertion as `offset_of!(T, f) == N * SLOT`
so the slot INDEX stays visible in the assertion. For N=0 and N=1 that is
`0 * SLOT` and `1 * SLOT`, which clippy rejects as `erasing_op` and
`identity_op` — six errors, and windows-host.yml runs clippy with `-D warnings`,
so the branch as pushed would have turned the Windows leg red.

This is the blind spot the programme document names in §1.5, demonstrated on the
programme's own first code commit: 44% of the host's unsafe is `#[cfg(windows)]`,
no Linux or macOS check compiles it, and `cargo fmt`/`cargo check` on a Mac are
all clean. Only the .133 gate sees it.

Fixed with a `const fn slot(i: usize) -> usize` rather than by writing the two
offending cases as bare `0` and `SLOT`: that would have made those two the only
assertions where the slot index is invisible, and the index is the entire point.

Also records the cheap local gate that would have caught this without a Windows
round-trip: `amf_sys.rs` depends on nothing but `c_void`, so copying it into a
throwaway one-file crate and running `cargo clippy -- -D warnings` reproduces the
exact error on any host. Verified by reintroducing `0 * SLOT` and watching the
harness fail with the same message the runner gave.

Verified on 192.168.1.133 (Windows CI runner, the box with the WDK), after a
`cargo clean -p pf-encode` that reported `Removed 47 files, 135.5MiB` so the
recompile is real and not a cached green:
  cargo check -p pf-encode                                                 ok
  cargo check -p pf-encode --all-targets --features nvenc,amf-qsv,qsv      ok
  cargo clippy -p pf-encode --all-targets --features nvenc,amf-qsv,qsv
        -- -D warnings                                          exit 0 (was 101)
  cargo clippy -p punktfunk-host --features nvenc,amf-qsv,qsv -- -D warnings
                                                                exit 0 (was 101)
The gate also greps the extracted tree for the assertions before building, so a
stale upload cannot produce a passing run.
2026-08-11 16:28:59 +02:00
enricobuehler cd3f5474bf fix(pf-driver-proto): a layout test read an align-8 struct out of an align-1 buffer
ci / web (pull_request) Successful in 1m11s
ci / docs-site (pull_request) Successful in 1m18s
apple / swift (pull_request) Successful in 1m48s
apple / screenshots (pull_request) Skipped
ci / bun-nix (pull_request) Successful in 20s
windows-drivers / driver-build (pull_request) Successful in 2m14s
windows-drivers / probe-and-proto (pull_request) Successful in 40s
android / android (pull_request) Successful in 4m11s
ci / rust-arm64 (pull_request) Successful in 3m7s
ci / rust (pull_request) Successful in 7m4s
`control_structs_roundtrip_through_bytes` built the legacy-size wire form in a
stack `let mut legacy = [0u8; 40]` (align 1) and then called
`bytemuck::from_bytes::<control::AddRequest>`. `AddRequest` opens with
`session_id: u64`, so it is align 8, and `from_bytes` hands back a REFERENCE
into the buffer — it panics unless the buffer happens to be 8-aligned.

A stack `[u8; 40]` usually is, which is why this passed on every machine and
every CI leg since it was written. Under Miri it fails outright: Miri does not
let an accidentally-favourable stack slot stand in for a guarantee.

Switched to `pod_read_unaligned`, which reads by value and has no alignment
precondition. That is not a new idea here — `ChannelProof::parse` at lib.rs:1013
already carries a comment saying "`pod_read_unaligned`, NOT `from_bytes`" for
exactly this reason. This site is the only other one in the crate that reads a
POD out of a stack byte array; every other `from_bytes` call in the tests reads
from `bytes_of(&x)`, which is aligned by construction.

Test-only, so no shipped defect — but the crate is `#![forbid(unsafe_code)]` and
is path-dep'd by BOTH the main workspace and the driver workspace, so it is the
layout oracle for every frame and IOCTL that crosses that boundary. A test that
cannot be trusted to fail is worth fixing there more than anywhere else.

Found by the first Miri run ever performed against this repo.

Verified on 192.168.1.25 (Ubuntu, cargo 1.96.0):
  cargo +nightly miri test -p pf-driver-proto                              21/21
  cargo +nightly miri test -p pf-driver-proto --target x86_64-pc-windows-msvc
                                                                          21/21
  cargo test -p pf-driver-proto --locked                                     ok
  cargo clippy -p pf-driver-proto --all-targets --locked -- -D warnings    clean

The cross-target run is the interesting one: it interprets the crate at MSVC
layout on a Linux box with no Windows anywhere. Nothing else in CI does that.
2026-08-11 13:57:34 +02:00
enricobuehler 972af2992f fix(pf-capture): the gamescope cursor fallback rewrote environ under a live multithreaded host
`connect_via_env_swap` did set_var("XAUTHORITY", …) / connect / restore, guarded
by a mutex that serialised this source against itself and against nothing else.
`getenv` takes no lock. setenv/unsetenv rewrite the process-global `environ`, and
glibc REALLOCATES that array when a variable is added — while, at that exact
moment, the PipeWire thread is inside pw_init()'s dlopen making bare getenv()
calls and EGL/CUDA init is running alongside. The file's own doc already called
the pattern "unsound from a live multithreaded host"; it stayed as a fallback.

Three things made it worse than the comment suggested:

- The damaging branch is the one where XAUTHORITY is ABSENT and therefore gets
  ADDED (the realloc case). scripts/punktfunk-host.service deliberately does not
  import the login shell's environment, so absent is the DOCUMENTED NORMAL
  configuration for the shipped unit, not an edge case.
- `rediscover` re-runs this every 2 s for the whole session. A display whose
  connect fails is never pushed into `displays`, so the dead-display skip never
  covers it — the race is not once at startup, it repeats forever.
- It is unfixable in place. Sharing pf_vdisplay's ENV_LOCK is the wrong layer: it
  cannot make C `getenv` take a lock.

The fix is to stop writing `environ` at all. Connecting with an explicitly empty
auth token is what the swap actually achieved: we only reach the fallback when
our own lookup found no usable MIT-MAGIC-COOKIE-1 entry, and x11rb's internal
lookup reads the same file with a STRICTER matcher (it matches family/address
too, which we deliberately do not), so where we find nothing it finds nothing
either and connects unauthenticated. That is exactly why the swap "worked"
against a nested Xwayland started without -auth.

Gives up one case: an .Xauthority using an auth family we decline to guess at but
x11rb would have handled. A gamescope Xwayland writes a single-entry
MIT-MAGIC-COOKIE-1 file, so it is not reachable here, and declining to attach a
cursor overlay beats tearing `environ` out from under a live session.

Also removes XAUTH_LOCK, whose only user this was.

Verified on 192.168.1.25 (Ubuntu, cargo 1.96.0 — the pinned toolchain, pipewire
dev headers present): `cargo check -p pf-capture --locked` and
`cargo clippy -p pf-capture --all-targets --locked -- -D warnings` both clean.
Not verified on glass: the fallback is only reached when the cookie parse fails,
so a normal gamescope session does not enter it. Forcing it needs a nested
Xwayland started without -auth, or a mangled cookie file, on .181/.136.
2026-08-11 13:54:56 +02:00
enricobuehler df6f270e7b chore(safety): forbid unsafe on the crates that are already at zero
Five permanent ratchets, all free today — the point is that they cannot regress
tomorrow. Each crate was re-measured at the commit, not taken from a survey.

`forbid(unsafe_code)`:

  punktfunk-encode-worker  the binary that carries cap_sys_nice. Its header
                           claims "no Wayland, no D-Bus, no network, no
                           plugins"; this makes the memory-safety half of that
                           claim mechanical. `forbid`, not `deny`, so it cannot
                           be re-opened by an #[allow] further down.
  pf-update-check          parses a signed, network-fetched manifest and its own
                           header says it "owns the part where being wrong is a
                           security bug". Signature checking is worthless if the
                           parser around it can be walked out of bounds.
  pf-vaadec                its header states the design constraint outright — it
                           links no libva and compiles on macOS, "which is the
                           point". The crate is full of hand-declared libva
                           repr(C) mirrors; one raw deref and it stops being the
                           CPU-testable half.
  tools/cursor-probe       free, and a probe is where "just deref it to see" is
                           most tempting.

`deny(unsafe_code)` + one localized allow:

  pf-update                root runs this. Its single unsafe operation, a bare
                           geteuid, moves into a named `effective_uid()` helper
                           carrying the crate's one #[allow(unsafe_code)].

Deliberately NOT rewritten to rustix, contrary to the programme document's first
draft: pf-update's Cargo.toml states that its zero-dependency posture IS a
security invariant of a root helper ("no HTTP client, no TLS, no argument
parsing"), and the extern block says the same. Pulling a general-purpose syscall
crate into a root helper to delete one `unsafe` would trade a real property for
a cosmetic one. The localized allow keeps the ratchet: any NEW unsafe anywhere
in the crate is a build error.

Verified: `cargo check -p pf-vaadec -p pf-update-check` and
`cargo check -p pf-update -p cursor-probe` clean on macOS, plus
`cargo check -p pf-update --target x86_64-unknown-linux-gnu` — pf-update's whole
body is behind `cfg(target_os = "linux")`, so the macOS check does not reach the
line that changed. punktfunk-encode-worker is not built here (pf-encode's C
dependencies do not cross-compile from macOS) and needs the Linux CI leg.
2026-08-11 13:49:41 +02:00
enricobuehler 27f0834025 fix(pf-encode): const-assert the AMF vtable and POD layouts
amf_sys.rs mirrors five AMF COM vtables by hand and amf.rs dispatches through
them BY SLOT POSITION — 18 distinct slots across the five tables. The mirrors
carried 118 `Slot` placeholders whose only job is to hold the following slots at
their C offsets, and not one layout assertion of any kind. A slot inserted,
removed or reordered in an AMF header bump calls an arbitrary function pointer
through a mismatched signature: no compile error, no runtime signal.

`AMF_MIN_VERSION` does not defend against this. It checks a version NUMBER, not
a layout, and it is a floor with no ceiling.

The three POD checks that did exist (`AmfVariant`, `AmfGuid`, `AmfHdrMetadata`)
lived in amf.rs's `#[cfg(test)]` module, so they were verified only when someone
ran pf-encode's tests, on Windows, with AMF enabled — and NEVER in a release
build, which is exactly where a mis-mirrored `AMFVariantStruct` does its damage:
it crosses the FFI BY VALUE on every SetProperty. This is the same hole
`a8dd348b` closed for the cuda.h mirrors and missed here.

Adds ~40 `const _: () = assert!(...)` guards next to the mirrors: size of each
of the five vtables, the byte offset of every slot amf.rs actually calls, the
three POD layouts promoted out of the test module, and the AMFData/AMFBuffer
shared-prefix agreement that `create_surface_from_dx11_native`'s
AMFSurface-through-AMFData reinterpretation silently depends on.

Verified by compiling amf_sys.rs standalone (it needs only `c_void`, and a
repr(C) struct of code pointers has the same layout on any 64-bit target, so a
macOS const-eval proves the Windows arithmetic), and by deliberately breaking one
offset to confirm the guard actually fires rather than silently passing.

That check earned its keep immediately: `alloc_buffer` sits at slot 43, not 42.
Counting AMFInterface(3) + AMFPropertyStorage(10) + the AMFContext block by hand
is exactly the error these assertions exist to catch.

Zero runtime behaviour change. The `AMF_MIN_VERSION` ceiling is deliberately NOT
part of this commit: a ceiling would make the next AMF driver release refuse
encode on every AMD box, so it needs a warn-and-continue policy plus an env
override and a real AMF session to gate it.
2026-08-11 13:49:25 +02:00
enricobuehler db6683a585 chore(safety): commit the unsafe census, fix its two bugs, record the baseline
Founding commit for a host-focused Rust safety programme. Adds the census tool
that measures the programme, the 2026-08-11 baseline it produces, and the
programme document itself.

The metric is SHIPPED NON-FFI UNSAFE OPERATIONS: 713. Raw `unsafe {}` block
count is the wrong target and the workspace manifest already says why — 63.3%
of unsafe operations in host scope (1542 of 2435) are a single third-party FFI
call that ash/windows-rs/ffmpeg mark unsafe on our behalf. A block count also
rewards merging blocks, ignores SAFETY comments, and IMPROVES when code moves
from Linux to Windows, because no local check can see the Windows half.

The tool shipped here had two defects, both fixed:

- `in_test_mod` cached parsed `#[cfg(test)]` spans in a dict keyed on `id(src)`,
  the memory ADDRESS of the source string. CPython recycles addresses, so once
  one file's source was collected the next file's string could be allocated at
  the same address and silently inherit the previous file's test spans. Ten
  consecutive runs over an unchanged tree produced 694, 695, 696, 701, 703,
  709, 710, 713, 714 and 721. Fixed by holding a strong reference to the string
  beside its spans, which makes the address un-recyclable while the entry is
  live. Five consecutive runs now agree exactly.

- The layout-assertion regex matched `const _: () = assert!(...)` but not the
  `const _: () = { ... };` block form, which 18 files use — including abi.rs,
  pf-inject/linux/gamepad.rs and pf-capture/.../idd_push/probes.rs. It reported
  102 unguarded repr(C) declarations across 25 files where the true figure is
  60 across 22, defaming three well-guarded files.

A metric that is not reproducible is not a ratchet. The acceptance gate for
this commit is therefore five consecutive identical runs, not one.

Baseline: 713 shipped non-FFI unsafe operations; 60 unguarded repr(C)
declarations across 22 files; unsafe reachable pre-authentication by an
unpaired peer = 0 first-party.
2026-08-11 13:42:05 +02:00
enricobuehler 7ffafb5ef3 chore(api): regenerate openapi.json after merging main
`main` gained the launcher brand tokens (`f62a48d4`) while this branch was open, and both sides
touch the generated document — so it was regenerated from the MERGED source rather than
text-merged. Verified to carry both: the 18 launcher-token entries from main, and this branch's
corrected schema descriptions. No `required` array changed, so no client regeneration is needed.
2026-08-11 11:03:42 +02:00
enricobuehler 4b686f026a Merge branch 'main' into worktree-vd-sweep-2
# Conflicts:
#	api/openapi.json
2026-08-11 10:57:35 +02:00
enricobuehler d6132f7523 chore(api): regenerate openapi.json for the pf-vdisplay policy doc corrections
The sweep rewrote doc comments on `ToSchema` types (`KeepAlive`, `Topology`, `ModeConflict`,
`Identity`, `LayoutMode`, `Layout`, `DisplayPolicy`, `EffectivePolicy`), and utoipa emits those
verbatim as schema descriptions — so the checked-in snapshot went stale and
`mgmt::tests::openapi_document_is_complete_and_checked_in` would have failed.

Several of the corrected descriptions were shipping outright falsehoods to API consumers. The worst:
`KeepAlive::Forever` documented itself as "**Not honored until the display-lifecycle stage**" while
the mgmt handler honors it end-to-end and the `gaming-rig` preset selects it (sweep item 11.7).

Diff is descriptions only — the `required` arrays are unchanged, so no SDK or client regeneration is
needed. Generated with `cargo run -p punktfunk-host -- openapi` in `ci/rust-ci.Dockerfile` under
`--platform linux/amd64`, and confirmed by running the host's own drift test there (37 mgmt tests).

`docs-site/public/openapi.json` is deliberately untouched: it is already ~34 KB behind `api/` from
earlier work, and refreshing it here would sweep in unrelated changes.
2026-08-11 10:55:27 +02:00
enricobuehler dc4d8d6832 fix(pf-vdisplay): correct the regressions this sweep introduced
An adversarial review of the sweep's own diff raised 39 claims; 23 survived independent
verification. This commit fixes them. Several are cases where the sweep traded one bug for another.

**The display budget was enforced in the wrong place.** The new Linux `max_displays` ceiling sat in
`registry::acquire` — which runs again on every mid-stream rebuild. All three create-before-drop
paths hold the old lease while acquiring the new display, and only the mode-switch path passes
`supersedes`, so a session at the ceiling counted itself against the budget and could never recover
from capture loss or a Game↔Desktop switch. At `max_displays = 1` that is a single streaming client.
Moved to `admission::admit`, which is where Windows has always applied it and which is reached once
per connect — so a rebuild cannot hit it.

**"Cannot tell" was collapsed into "wrong mode".** `unanimous_output_size` returning `None` for two
disagreeing gamescopes was compared with `== Some(target)`, so ambiguity took the destructive branch:
a nested per-title gamescope — the normal Game Mode shape — made every connect restart the box's
session and kill the running game. Now a three-state `BoxOutputSize`, where `Ambiguous` mirrors the
live node instead of re-moding, and the post-restart wait asks "did what we asked for come up"
rather than demanding unanimity.

**Decide-then-act lost its mutual exclusion.** Re-scoping the `MANAGED_SESSION` guard fixed the
shutdown restore but let two concurrent creates at the same mode both relaunch, the second stopping
the unit the first was polling. A separate `MANAGED_LAUNCH` mutex restores the exclusion without
putting launch progress back into the lock the restore samples.

**Per-axis policy salvage was applied to a selector.** `preset` chooses the other axes, so salvaging
it to the default silently re-pointed the whole document; it now refuses the document instead. A file
whose every axis is unreadable also reported `configured() == Some(default)` — flipping Linux
identity from Shared to PerClient — and now correctly reports unconfigured.

**The six `#[serde(default)]` on `EffectivePolicy` are reverted**: they loosened `POST/PUT
/display/presets` (an omitted axis defaulted where it used to 400), which nobody asked for. The
catalog salvage they were added for now lives in a private Deserialize-only mirror type, so the read
path stays lenient and the wire contract stays strict.

Also: the Windows create path stored the OS-committed refresh in the field `acquire` uses as its
resize discriminator, so a same-mode re-acquire looked like a hotplug — the requested and committed
modes are now separate fields; `output_within`'s timeout arm detached both reader threads (now
bounded by a drain grace, capped at 16 MiB, and logged honestly — a `systemd-run --pipe` unit escapes
the process group and cannot be reached); `reenable_outputs_kscreen` abandoned the mode restore
whenever kscreen-doctor hit its budget even though the enable may have landed (now tri-state);
`write_atomic` replaced a symlinked portal config with a regular file, severing dotfiles management;
several new budgets were too short for the helper they bound (`steam -shutdown` was being killed
before it could deliver the request; `linger_enabled` read a 300 ms timeout as "not lingering" and
hard-failed a correctly configured box); and a restore logged an operator-facing error for a
`systemctl` call that had merely outlived its budget while systemd still owned the queued job.

Verified: 107 tests on macOS, 202 on Linux (executed in a container, not merely type-checked),
Linux and Windows clippy clean at `-D warnings`, fmt clean.
2026-08-11 10:06:16 +02:00
enricobuehler 8b98d0b3ec fix(pf-capture): a sweep found nine real defects behind comments that asserted the opposite
Reviewed the whole crate (15.6 kloc) for bugs, safety, structure and comment truth.
Both compile gates are green: `scripts/xcheck.sh windows clippy` and
`cargo clippy -p pf-capture --all-targets --locked -- -D warnings` in the amd64 CI
image (the Linux half needs libpipewire, so it cannot ride xcheck).

Code defects, each one contradicted by a comment sitting next to it:

* `pipeline_depth` clamped to `OUT_RING` (3) while both `repeat_last` and `OUT_RING`
  state the safe maximum is 2. `d` frames in flight need `d + 1` textures, so
  `PUNKTFUNK_IDD_DEPTH=3` rotated onto the slot NVENC was still reading and the convert
  overwrote it in place — torn frames, silently. Now `OUT_RING - 1`.
* The GDI cursor poller published `visible: true` for a NULL `hCursor` carrying
  `CURSOR_SHOWING` — how an app hides the pointer for its own window. The last
  rasterised arrow was then blended into a game that had hidden its cursor. Every
  rasterise gate already tested `handle != 0`; the published verdict now agrees.
* The ETW event callback did `RING.lock().unwrap()`. That is an `extern "system"` fn, so
  a poisoned lock panicked across an FFI boundary and ABORTED the host — a diagnostic
  taking down capture. Poison-tolerant now, which also makes the poison unreachable.
* `ChannelBroker::send` bounded the ring with `debug_assert`, so a release build instead
  panicked mid-`duplicate_and_deliver`, unwinding past the reap and leaking every handle
  already planted in the driver's WUDFHost. Refuses before the first duplication.
* `set_active(false)` did not clear `stall_since`, so a pooled capturer carried a stale
  stall clock into its next stream and reported capture loss microseconds in.
* `attach_gamescope_cursor` evaluated `spawn` before dropping the old source: two readers
  published into one slot, and a failed spawn destroyed a working reader. Idempotent now.
* `PUNKTFUNK_FORCE_SHM` used a bare `== "1"` compare, silently ignoring `=true`/`=on`.
* `spa_meta_bitmap.offset == 0` is SPA's "no image data" signal, distinct from the
  `bitmap_offset == 0` position-only case. Unhandled, it decoded the header's own words
  as cursor pixels and cached them.
* A `VideoInfoRaw::parse` failure was swallowed, so a malformed Format pod surfaced as
  the generic "no acceptable format" timeout. It is logged, and parsed once, not twice.

Comment corrections, all verified against the code they describe: four claims that a
failed open falls back to DDA (removed — the caller drops the keepalive under
"no fallback"); three comparisons to the removed WGC path; "we do NOT gate HDR on the
client's VIDEO_CAP_10BIT" (it does, in three places); the P010 sampler's "4 explicit
taps / 2x2 box" (two taps, left-cosited — the box was the bug it replaced); the cursor
meta cap quoted as 256x256 (1024, and 256 is the value that cost the whole Linux cursor
channel on-glass); the poller's "~60 Hz" (4 ms, ~250 Hz); "several minutes of coverage"
(~26 s); "8 frames in 400 ms >= 20 fps" (7 intervals, so 17.5); three "process-wide" HDR
latch claims (per-source, which is why HdrSource exists); a SAFETY proof claiming a view
is "unmapped never" (Drop unmaps it); the Linux module header describing a bounded
channel and BGRx-only frames (one-deep overwriting slot, several formats); and a doc
line stranded on `DisplayDescriptor` by an earlier split, restored to `IddPushCapturer`,
which had none.
2026-08-11 10:04:58 +02:00
enricobuehler 6b33750edc fix(pf-vdisplay): one non-UTF-8 byte in a portal config destroyed the whole file — in the module written to prevent exactly that
`portal_config::ensure_key` folded EVERY read failure into an empty string
(`read_to_string(path).unwrap_or_default()`). `upsert("", …)` then produced a file containing only
our block, the one-time backup was skipped because `!existing.is_empty()` was false, and the write
replaced the user's config — returning `Ok(true)`.

So a single Latin-1 character in a comment in `~/.config/hypr/xdph.conf` or
`~/.config/xdg-desktop-portal-wlr/config` destroyed the operator's entire portal configuration, with
no backup and no warning. The module doc says flat-writing these files "destroyed [everything else]
on first connect, silently and permanently" and that this module exists so it cannot happen; that
one line re-opened the door. The same shape hit a transient EIO on an NFS or overlay config dir.

Now: bytes are read with an explicit match, only `NotFound` may mean "empty", a non-UTF-8 config is
refused by name rather than replaced, the backup is taken by BYTES, and the write is atomic
(temp + `sync_all` + rename in the same directory, permissions carried over). Five new tests, all
running on macOS — `a_non_utf8_config_is_refused_not_replaced` fails against the old code.

Also in the wlr/Mutter family:

* **Mutter's `Primary` rebuilt kept physicals from scratch** — scale forced to 1.0, transform to 0,
  disabled heads re-enabled — so a rotated, 2x-scaled or deliberately-disabled monitor came back
  wrong, while the code went to real trouble to preserve refresh. Each head now carries its
  pre-connect scale and transform, and x advances by the LOGICAL width.

* Three availability probes read session env (`SWAYSOCK`, `XDG_CURRENT_DESKTOP`,
  `HYPRLAND_INSTANCE_SIGNATURE`) with no `ENV_LOCK` while `apply_session_env` `set_var`s the same
  keys from another thread — the glibc setenv/getenv race this crate's own lib.rs documents as UB.

* `wlroots::create_output` ran a statement before its `OutputGuard` existed, so a raced
  `wait_new_output` orphaned the output permanently — hyprland takes the guard first. The
  before/after name diff also ran outside any lock, so two concurrent creates could adopt each
  other's output. Both now run under a create lock, with a stray sweep on the failure path.

* `select_and_cast`'s timeout arm dropped the portal thread's `stop` flag un-set — the same leak
  Mutter was already fixed for. The guard is now built before the wait, in both copies.

* The xdpw chooser file was written per session and never removed, permanently shadowing the
  config's fallback with the name of an already-unplugged output. Its lifetime is now the handshake,
  not the session — scoped deliberately, because tying removal to the keepalive would let one
  session delete another's selection hours later.

* Hyprland's headless outputs are now named `PF-<pid>-<n>` and reconciled at startup, so a crashed
  host's leftovers are reclaimed while a live sibling host's outputs cannot be pulled out from under
  it. `set_monitor_rule` no longer discards hyprctl's rejection text and then hard-codes a
  GBM/dmabuf diagnosis it never verified.

* Both wlr backends silently dropped the `topology` policy axis: `Primary`/`Exclusive` was accepted,
  echoed by the mgmt API, applied on three backends and a no-op on two. They now say so.

Item 8.1: `swaymsg`, `hyprctl` and the portal `systemctl --user try-restart` calls are bounded
through `proc` with named budgets.
2026-08-11 09:22:32 +02:00
enricobuehler ef72d102b6 fix(pf-vdisplay): KWin's re-enable reported success when it matched no outputs at all, leaving a physical monitor dark
* **`reenable_outputs` returned `true` when it matched NONE of the requested outputs.** Unresolvable
  outputs were `continue`d and the return was the apply verdict alone — but an empty
  `kde_output_configuration_v2` still gets an `applied` event. So a total no-op suppressed the
  `reenable_outputs_kscreen` backstop and the operator's physical monitor stayed dark. Now counts
  staged outputs and returns `ok && matched == outputs.len()`, and refuses to apply an empty
  configuration at all.

* **The kscreen restore logged "restored the physical/bootstrap outputs" unconditionally**, with
  both call results discarded — including when `kscreen_ok` returned false on its 5 s budget, which
  is exactly the wedged state that fallback exists for.

* **`Session::open` swallowed every failure reason** — connect error, barrier timeout, missing global
  — and three of four callers degraded to kscreen-doctor with zero log. This is the class that hid
  the KWin >= 6.7 registry regression: a shipped fallback firing silently on every machine. It now
  logs at warn with the reason and the caller's operation name.

* `last_name` was seeded with a name kscreen-doctor can never resolve (KWin's address is
  `Virtual-punktfunk…`), so the intended default was guarded by an `is_none()` that could never hold
  and `apply_position` ran against no output. `our_uuid` was never reset per `create` and only
  assigned under `outcome.handled`, so a supersede positioned the *previous* output and never fell
  back.

* `probe()`'s `roundtrip` was the only unbudgeted compositor wait in the crate — every sibling path
  is budgeted — and it is reached from an async mgmt handler. Now bounded at 3 s. The pre-`created`
  dispatch loops gained deadlines and now set `stop` on the timeout arm.

* Every `wl_output` global was bound for the session's life with no `GlobalRemove` arm and no
  `release()`, on the virtual-output path too, which never reads them: unbounded growth on a
  hotplugging session.

* `monitors::list` was the one KWin call site with no kscreen fallback at all, despite `list_monitors`
  failing on exactly the condition the other four fall back for. It has one now.

* `CVT_H_GRANULARITY` and `MANAGED_PREFIX` existed as two literals under prose asserting they match;
  the second copy now imports the first.

The wider facade extraction (item 9.1) is deliberately not in this commit, but its two prerequisites
are — a comment at the restore seam records why they had to come first: a fallback arm that returns a
value the helper never checked re-introduces the silent success, behind a seam whose selling point is
one honest log per decline.

Also corrects the `PhysicalMonitor` type doc, which claimed "logical geometry throughout" while
`width`/`height` are the mode's PIXELS and `x`/`y` are logical, and adds the `logical_size()` helper
that is the only correct way to compare an extent against a position.
2026-08-11 09:22:11 +02:00
enricobuehler b2c03f1904 fix(pf-vdisplay): a managed launch blocked the shutdown restore that was meant to rescue it, and re-moding could flip the operator's own screen
The gamescope subsystem — the crate's largest and fastest-churning area, and the one the 2026-07-28
sweep predates most of.

* **`MANAGED_SESSION` was held across the ~90 s managed launch**, and the shutdown/idle restore
  blocks on that same lock — *after* it has already stopped our unit. So the display-manager restore
  never ran and the box was left with no session at all. `create_managed_session` now decides under
  the guard and acts outside it, re-acquiring only to store the result; `do_restore_tv_session`
  consumes the record in a short scope at the top. Same shape the SteamOS twin already used.

* **The physical-display guard was bypassed whenever no gamescope node happened to be published.**
  `if physical_display_connected() { if let Some(node) = find_gamescope_node() { … } }` fell through
  to `set-environment SCREEN_WIDTH/HEIGHT/CUSTOM_REFRESH_RATES` + `restart` when the node was
  momentarily absent — gamescope restarting between titles, or built without PipeWire — flipping the
  operator's own screen to the client's resolution and bouncing a DM-driven login session. The guard
  now refuses instead of falling through, and the forced `SCREEN_*` values (which were never unset,
  so every later session on the box inherited them) are tracked and `unset-environment`ed on restore.

* **`current_gamescope_output_size()` reported an arbitrary gamescope's `-W`/`-H`** — whichever
  `/proc` enumerated first — and four consumers treated it as this session's output size. It now
  answers only when every gamescope on the box agrees, and `None` ("cannot tell") when they differ.
  `heads.rs` no longer takes it at all: it reads the size off the DRM-backed argv it already
  selected. Its test previously passed `None`, which is why the hazard was invisible.

Resource and honesty fixes: the ATTACH path armed the box's own session-unit bind drop-in and no
in-process path ever removed it (now tracked and disarmed on both restore arms); `wait_for_node`
never called `try_wait`, so a gamescope that died at `vkCreateDevice` was polled for the full 15 s
and the error then blamed headless capture support; `do_restore_tv_session` deleted its crash-recovery
state *before* the unbounded work that state records, so a grace-period expiry in that window left
the DM down with nothing on disk to heal it; the SteamOS takeover's two failure arms never armed the
TV restore though the session-plus twin does; the TV-session restore logged success with the
`systemctl` status discarded; the `steam -shutdown` child was dropped un-reaped; and a managed
session that took nothing over was never persisted, so a host crash orphaned the transient unit.

Item 8.1: the unbounded `pw-dump`, `systemctl`, `loginctl` and `pkexec` calls in this subsystem now
go through `proc::{status_within, output_within}` with per-call budgets. `pw-dump` is polled from
three separate 45 s loops against the very daemon this file documents gamescope as head-blocking,
and until now a hang there pinned the session's stream thread forever.
2026-08-11 09:22:09 +02:00
enricobuehler db65980979 fix(pf-vdisplay): the ghost-monitor reap fed live devices to pnputil, and two unsafe fns had no unsafe in them
Windows half of the sweep — the reap bug, a panic that poisons two locks, and a round of unsafe
reduction.

* **The ghost reap selected the wrong devices.** It filtered `Status -ne 'OK'`, a HEALTH field: that
  matches devices that are PRESENT but in Error/Degraded/Unknown, not the ABSENT ones the reap is
  for — and it handed them to `pnputil /remove-device`, contradicting its own documented contract.
  It runs from `add_monitor`'s mid-session slot-exhaustion recovery, so the blast radius is a live
  session. Now filters on `-not $_.Present`.

* **`ensure_pinger` still used the panicking `thread::spawn` while holding two locks**, poisoning
  both — the un-fixed twin of a fix that already landed for `ensure_exclusive_watch`. Same shape
  applied.

Unsafe reduction, continuing the program that made pf-win-display's CCD helpers safe fns:

* `resolve_target_gdi` and `reisolate_after_swap` were `unsafe fn`s containing zero unsafe
  operations, and the three call-site SAFETY proofs described FFI they no longer perform. Both are
  now safe fns and those blocks are gone.
* `VdisplayDriver::open`'s `# Safety` section named no caller obligation — the same empty shape an
  earlier phase already removed from `open_device`.
* `(*detail).DevicePath.as_ptr()` derived a pointer from a `[u16; 1]` field and handed it to
  `CreateFileW`, which reads the whole flexible-array path beyond it. Now taken with `&raw const`
  from the full struct, so the pointer carries the provenance of the bytes actually read — the same
  correction already made for `MONITORINFOEXW` in ddc.rs.

Comment fixes, all verified against the code: three intra-doc links to a type this crate does not
have; a doc-comment run merged so that `shrink_action` — the gate that keeps a `Primary` group's
physical panels lit — read as undocumented while its rationale sat on an unrelated polling helper;
and the backend module header, which documented itself against a `sudovda` module that does not
exist and a fallback the crate says was removed.

Adds the first tests for `knobs.rs`, `instance.rs` and `driver.rs` — including `is_privileged_sid`,
the security-relevant predicate that decides whether an existing single-instance name is another
host or a squat, which had no coverage on any platform.
2026-08-11 08:49:06 +02:00
enricobuehler a1ff0dde0c fix(pf-vdisplay): the host promised HDR and cursor forwarding for gamescope sessions it did not start
`gamescope_ours_and` answered "did WE spawn this gamescope?" by reading `PUNKTFUNK_GAMESCOPE_NODE`.
Phase 2.3 deleted the code that published that key — routing.rs's own doc says "Nothing is written
back to the two knobs" — but this consumer was never migrated, so the read now returns "not
attaching" for every attach.

Both consumers then answer for a session this host has no flags on. On a plain box with a foreign
gamescope already running, `pick_gamescope_mode` resolves Attach at its fifth rung while the env key
stays unset, and the probe half only inspects the resolved BINARY, which is our patched build:

* `gamescope_composites_cursor()` returns true, so the host attaches no XFixes reader and blends
  nothing — while the stock gamescope actually running was never given
  `--pipewire-composite-cursor`, so the stream carries no pointer at all.
* `gamescope_hdr_available()` returns true, so the Welcome fixes `bit_depth` at 10 and the session
  negotiates BT.2020/PQ over an 8-bit SDR composite. The Welcome cannot take that back.

The same two failures hit the `capture_monitor` mirror route on any Bazzite or SteamOS box, where
the running Game Mode gamescope is by definition not one this host spawned.

The question is now asked of the resolved route rather than the environment, via a pure
`session_is_a_foreign_gamescope` that runs — and is tested — on every platform. The residual gap is
named in the doc rather than papered over: `create_managed_session`'s create-time degrade to a
foreign attach is still invisible to a ladder re-run.

Also in this commit:

* Two unguarded session-env reads now take `ENV_LOCK` (`detect()`'s `XDG_CURRENT_DESKTOP` fallback
  and `effective_topology()`'s legacy pins). `apply_session_env` `set_var`s those same keys from
  another thread, which is the glibc setenv/getenv race this crate's own lib.rs documents as UB.
* `mirror.rs`'s `names_ours_conclusively` was a `matches!` whose omitted default was the UNSAFE
  direction — a new backend would silently get its own virtual displays mirrored. Now exhaustive, so
  adding a `Compositor` is a compile error at the one site where the answer is a safety decision.
* `MirrorDisplay` overrides `poolable_now() -> false`; its `create` always reports `External`, so
  the trait's `true` default was a pre-create claim contradicting the post-create fact. The trait
  doc now says plainly that the default is a default and not a fact.
* The crate front-door doc listed 3 of 7 backends and quoted line counts half the size of the
  current crate; `routing.rs`'s summary was attached to the wrong item and described a published env
  channel that no longer exists; `available()` is no longer documented as cheap when it forks
  `gamescope --version` and does an unbudgeted Wayland roundtrip per call.
2026-08-11 08:48:49 +02:00
enricobuehler 9d58f4c170 fix(pf-vdisplay): one unreadable byte reverted the host to built-in display defaults, and one bad preset dropped the whole catalog
The policy layer folded every failure into "unconfigured", then wrote that emptiness back.

* **Any parse error reverted the WHOLE policy.** An unknown enum variant, a mistyped scalar, an
  EACCES or EIO — all became `Err(_) => None`, i.e. the host silently ran on built-in defaults with
  the operator's `display-settings.json` still sitting on disk. Parsing is now layered: strict
  first, then per-axis salvage so one unreadable axis costs only that axis, and only `NotFound` is
  quiet — EACCES/EIO warn loudly that the host is on defaults. `version` is read instead of being
  blindly rewritten to 1.

* **One malformed entry dropped the entire custom-preset catalog**, and the next CRUD atomically
  renamed the empty vector over the file. Entries are parsed one at a time now; a lossy load is
  flagged and refuses to overwrite.

* `sanitized()` clamped `max_displays` but never `KeepAlive::Duration.seconds`, so a PUT could pin a
  display for ~136 years — a deadline the reaper never reaches and a nonsense `expires_in_ms` in
  `/display/state`. Clamped to a day, in both `sanitized()` and `sanitize_preset_fields`, and
  sanitization now runs on LOAD as well as on write.

* The two stores' temp files had fixed names and no write lock, so concurrent saves could interleave
  serialize -> rename -> in-memory update. Unique suffixes, a lock, and the in-memory update ordered
  after the rename.

* `new_preset_id` never consulted the loaded entries for collisions.

* **Manual layout could place an unpinned display exactly on top of a pinned one**: the fallback was
  the unconditional auto-row prefix sum, blind to where prior members were pinned. Unpinned members
  now pack clear of the pins. Layout keys are canonicalized and unusable ones dropped at write time
  rather than persisted-and-ignored.

Adds 20 tests, all running on macOS: a 20k-round randomized property test asserting no unpinned
member ever overlaps a sibling (verified to fail against the pre-fix `arrange_manual`), the salvage
and quarantine paths, the clamps, and a field-count guard that fails the moment a 13th policy axis
appears without being wired into the merge path.

Note: `partial_json_fills_defaults` was renamed to `serde_defaults_fill_a_partial_document` with no
assertion weakened — it pins the FILE contract (an old settings file must still load), which is not
the mgmt PUT contract that sweep item 11.1 is about.
2026-08-11 08:48:47 +02:00
enricobuehler 61ff543acc fix(pf-vdisplay): a new client could be handed a streaming client's display, and a blind /proc scan tore every backend down
Five defects in the registry/identity half, plus the restructure that finally makes them testable.

* **A new client could be assigned a LIVE client's identity slot.** `DisplayIdentityMap::resolve`
  LRU-evicted purely on its `seen` stamp, with no knowledge of which ids are streaming. On Windows
  that id keys the manager's slot map, so the newcomer took the plain-JOIN branch and inherited the
  other client's monitor, capture target and stop flag. `resolve` now takes the live set, never
  evicts a live id, and REFUSES rather than hand one over — degrading to the shared/auto identity.

* **A transient `ActiveKind::None` invalidated every backend entry, including live streaming ones.**
  A `read_dir("/proc")` that happened to fail satisfied the change test and bumped the session
  epoch. A `None` observation is no longer evidence a desktop went away, and no longer overwrites
  the baseline (which would have bumped the epoch on the next poll anyway).

* **The Linux pool had no display ceiling at all** — `max_displays` was enforced only on Windows,
  while the pool keys on the CLIENT-SUPPLIED mode, so each distinct requested resolution minted a
  new display. Now capped in `linux::acquire`, gated on `poolable_now` so a gamescope attach or
  managed session (which consumes no pool slot) is not refused.

* **Two different definitions of "display group"** — `group_key` and a bare backend-name compare —
  and only one separated gamescope spawns. Unified as `pool::in_group`. The `position_for_new`
  collection also lacked the supersede exclusion the topology check 70 lines earlier had, so a
  mid-stream resize auto-rowed the replacement past its own dying predecessor, walking the display
  one width to the right on every mode switch.

* **Lifecycle events were wrong in both directions**: `Created` fired on keep-alive reuse, and
  `Released` fired only from the mgmt endpoint — never from a lease drop, the linger reaper,
  `mark_failed`, `retire` or `invalidate_backend`. All six now emit.

Also: `Release::Noop` no longer runs a full teardown (the one outcome the state machine defines as
"do nothing"); a failed linger-reaper spawn logs and retries instead of consuming its `Once` and
never tearing a kept display down again; group ids are a monotonic per-key counter instead of an
index into the currently-live sorted set, so an unrelated group appearing no longer renumbers a
display; and a corrupt `display-identity.json` is renamed to `.bad` with a warning rather than
silently overwritten, which used to reset every client's stable id and its saved DPI.

The pure half of the pool (`Entry`, `group_key`, `epoch_matches`, `take_expired`, `at_display_budget`,
`position_for_new`, `assign_group_ids`, `assemble_displays`) is now a non-cfg'd `mod pool`, so the
registry's decisions are exercised on every platform's CI instead of only on a Linux box. Crate test
count 53 -> 94.
2026-08-11 08:48:15 +02:00
enricobuehler dd9bbaf1c5 fix(pf-vdisplay): a helper that outran the pipe buffer had its output thrown away as a timeout
`output_within` read stdout/stderr only after the child exited, and its doc justified that with
"these helpers emit at most a few hundred KiB, well under any real pipe pressure". A pipe holds
64 KiB. Anything past that blocks the helper in `write()`, so it never exits, the budget kills it,
and a successful query is reported to the caller as `TimedOut` with its answer discarded.

The busiest caller is the one that trips it: `pw-dump` on a populated PipeWire graph clears 64 KiB
routinely and is polled from the 45 s gamescope loops. Confirmed empirically — a child writing
1 MiB into an undrained pipe never exits.

Both pipes are now drained on their own threads, concurrently with the wait.

That makes the joins load-bearing, which exposed the second half: the Unix `tree::Guard` was an
empty stub whose doc claimed `Child::kill` "already ends the only process there is". It never did
for this crate's Linux helpers — `pkexec`, `systemd-run`, `systemctl --user` and the `sh -c`
wrappers all fork — and a surviving grandchild holds the pipes' write ends, so a reader would wait
for an EOF that never arrives. The child is now the leader of its own process group and the guard
`killpg`s it, which is the Unix shape of the Job object the Windows half already used.

Also gates `pf-frame`, `pf-gpu` and `pf-encode` to Windows: every use site of all three is
`cfg(windows)`, and between them they dragged FFmpeg, ash and openh264 into the Linux build for
nothing (sweep item 13.19).
2026-08-11 08:22:16 +02:00
200 changed files with 11155 additions and 2768 deletions
+269
View File
@@ -15,6 +15,17 @@
# fails if any crate carries a license outside the allowlist — the regression
# guard about.toml always promised. (The Android Gradle tree has no lockfile, so
# nothing scans it — see the CRA roadmap.)
# * miri → NON-BLOCKING interpretation of the few FFI-free leaf crates, one of them
# cross-compiled to MSVC layout. Not a supply-chain scan; it lives here because
# audit.yml already has exactly the shape it needs (weekly cron,
# workflow_dispatch, the rust-ci container, the same cache pattern) and because
# ci.yml runs on every push against a fleet where 37 of 46 jobs contend for
# ubuntu-24.04. See the `miri:` job below for what it does and does not buy.
# * c-abi-asan → NON-BLOCKING ASAN+LSAN run of the C ABI harness (tests/c/run.sh under
# PF_SAN=address): both sides of the abi.rs boundary instrumented at once, and
# the only automated check on its Box::into_raw/from_raw leak contract. Same
# here-not-ci.yml reasoning as miri — plus -Zbuild-std defeats sccache, so it
# must not ride the per-push leg.
# Triggers: weekly (catch newly-disclosed CVEs in pinned deps), on every lockfile/allowlist
# change, and on demand.
# To silence a known-unfixable Rust advisory, add it to `.cargo/audit.toml` ([advisories] ignore=[…]).
@@ -44,6 +55,13 @@ on:
- 'about.toml'
- '.gitea/workflows/audit.yml'
workflow_dispatch:
# NOTE on the `paths:` list above and the `miri:` job: `crates/pf-driver-proto/**` is deliberately
# NOT listed, even though that crate is what the Miri job exists to watch. `paths:` is a
# WORKFLOW-level filter — adding it would fire all six jobs (three bun trees, pnpm, cargo-audit,
# the license gate) on every driver-proto edit, onto a fleet where 37 of 46 jobs contend for
# ubuntu-24.04, to run one 2-minute job. Weekly cron + workflow_dispatch is the day-one cadence;
# revisit once the job has a green history, and if you do, prefer moving miri to its own workflow
# file over widening this filter.
jobs:
cargo-audit:
@@ -177,3 +195,254 @@ jobs:
command -v cargo-about >/dev/null 2>&1 || cargo install --locked cargo-about --version 0.9.1 --features cli
cargo about generate about.hbs --fail -o /dev/null
cargo about generate -m packaging/windows/drivers/Cargo.toml -c about.toml about.hbs --fail -o /dev/null
# ── Miri ─────────────────────────────────────────────────────────────────────────────────────
# WHAT THIS BUYS, precisely — one thing, and it is worth having:
# It interprets `pf-driver-proto` CROSS-COMPILED TO `x86_64-pc-windows-msvc`, on a Linux
# runner, with no Windows box anywhere in the loop. That crate is `#![forbid(unsafe_code)]`
# and is path-dep'd by BOTH the main workspace and the driver workspace, so it is the layout
# oracle for every frame and IOCTL crossing that boundary — and drift there is silent
# corruption, not a compile error. Nothing else in CI checks it at MSVC layout.
# On the first run ever performed against this repo it found a real defect: a layout test
# reading an align-8 struct out of an align-1 stack buffer, which had passed on every machine
# and every CI leg since it was written because a stack `[u8; 40]` usually lands 8-aligned.
#
# WHAT IT DOES NOT BUY — do not let anyone report this as unsafe coverage, and do not publish a
# "Miri coverage" percentage; it would be noise. Miri can execute on the order of 2% of the
# host's unsafe. It cannot run ash, windows-rs, ffmpeg, CUDA or the WDK, and in those crates
# the unsafe *is* the foreign call, so there is nothing for an interpreter to execute. This
# job is a targeted instrument for three leaf surfaces, not a safety net.
#
# NON-BLOCKING, deliberately, and via a step-level `||` — NOT job-level `continue-on-error`,
# which act_runner does not reliably honor (same reasoning as docs-site-audit above; a red job
# here would take the whole run red). Flip to blocking only after several weeks of green
# establish the nightly-drift rate.
#
# Do NOT add crates here because they merely compile under Miri. Add them because they contain
# pure-Rust unsafe or a layout contract worth interpreting. Explicitly excluded:
# * pf-bitstream — its compile did not finish in 27 min at 2.1 GB RSS, and it is
# `forbid(unsafe_code)`, so there is nothing to find. Do not re-add it.
# * pf-update-check — ring; every FFI crate — dies on the first foreign call. Structural.
# * punktfunk-core in bulk — `-- fec packet crypto` selects 63 tests and was killed at a
# 25-minute cap with not one test reported complete. Only the narrow
# `fec::gf8` selection below is affordable, and it was timed before it
# was committed. Do not widen this filter without timing the result.
#
# MEASURED, not estimated — 192.168.1.25 (Ubuntu, 8 cores), on the DATED toolchain this job
# actually installs, with a COLD target dir and a COLD sysroot cache (so each step's figure
# includes building the Miri sysroot it needs) and a warm cargo registry. Every step below has
# been run start to finish; nothing here is extrapolated:
# step A 21 + 12 + 4 pass 43 s
# step B 21 pass 26 s
# step C 2 pass 63 s
# TOTAL 132 s cold. Interpretation itself is ~10 s of that; the rest is compiling, plus ~38 s
# of one-time sysroot builds (21 s host + 17 s MSVC) that the cache below then carries.
# Warm, the three steps are ~6 s / ~3 s / ~10 s. `timeout-minutes: 30` is therefore vast
# headroom, kept deliberately so a first fully-uncached run — which additionally downloads a
# ~400 MB toolchain and the registry — cannot trip it.
# If you add a step, MEASURE IT FIRST. The estimate this job replaced said "under 15 s across
# all four steps" and was extrapolated from a partial run; the real punktfunk-core figure was
# >25 min. Extrapolation is exactly how that happened.
miri:
runs-on: ubuntu-24.04
container:
image: 192.168.1.58:5010/punktfunk-rust-ci:latest
timeout-minutes: 30
env:
# A DATED nightly, bumped deliberately — exactly like rust-toolchain.toml, and for the same
# reason. The cache keys below carry this value, so bumping it self-invalidates them.
# ⚠ `nightly-<date>` names the day rustup PUBLISHED the build, and that build is compiled
# from the PREVIOUS day's commit. This pin therefore resolves to
# `rustc 1.99.0-nightly (969b803cb 2026-08-09)` [verified by installing it], NOT the
# `12c36e253 2026-08-10` that the rust-safety programme doc's §7 table cites — that figure
# came from the ROLLING `nightly` channel and was mislabelled as the dated one. Harmless,
# but do not "fix" the date to chase that hash: all three steps below were re-run and are
# green on the dated toolchain this job actually installs.
MIRI_TOOLCHAIN: nightly-2026-08-10
# A GUARD, not a fix for a present problem: audit.yml sets no sccache — only ci.yml does, at
# workflow level (ci.yml:27). `cargo-miri` REPLACES rustc and cannot be wrapped; it prints
# "Ignoring `RUSTC_WRAPPER` environment variable, Miri does not support wrapping" and
# carries on [verified]. This keeps a future workflow-level sccache from becoming a puzzle.
RUSTC_WRAPPER: ""
# -Zmiri-disable-isolation: pf-gpu's tests mkdir, and Miri aborts them without it [verified].
# -Zmiri-symbolic-alignment-check: the whole point — it refuses to let an accidentally
# favourable stack slot stand in for an alignment guarantee. This is the flag that caught
# the pf-driver-proto defect.
# NOTE the absence of -Zmiri-ignore-leaks. Miri leak-checks by DEFAULT, and that is the one
# leak-detection capability it offers here. None of the crates below leaks, so the job is
# green. The tree does contain DELIBERATE leaks (pf-umdf-util/src/section.rs `ViewCell`,
# gamepad_raii.rs leak-on-timeout) — when coverage ever reaches them, annotate those two
# sites; do not blanket-disable the check.
MIRIFLAGS: -Zmiri-disable-isolation -Zmiri-symbolic-alignment-check
steps:
- uses: actions/checkout@v4
# Two caches, split on purpose so a Cargo.lock change does not re-download a ~400 MB
# toolchain. Both use their OWN `miri-` key prefix — never a shared one.
# The Miri sysroot is per-toolchain and per-target (two are built here: host + MSVC), so it
# belongs with the toolchain, not with the lockfile.
- name: cache the nightly toolchain + Miri sysroots
uses: actions/cache@v4
with:
path: |
/usr/local/rustup/toolchains/${{ env.MIRI_TOOLCHAIN }}-x86_64-unknown-linux-gnu
~/.cache/miri
key: miri-toolchain-v1-${{ env.MIRI_TOOLCHAIN }}
- name: cache the cargo registry
uses: actions/cache@v4
with:
path: /usr/local/cargo/registry
key: miri-registry-v1-${{ hashFiles('Cargo.lock') }}
restore-keys: miri-registry-v1-
# The image needs no change for this: ci/rust-ci.Dockerfile:51-54 installs via rustup and
# `chmod -R a+w`s both RUSTUP_HOME and CARGO_HOME, so a job can add a toolchain at runtime.
# `rust-src` is required — cargo-miri builds its sysroot from source, per target.
#
# This does NOT disturb the 1.96.0 pin: `cargo +<toolchain>` overrides rust-toolchain.toml
# for that single invocation only, so `cargo fmt` / `clippy` keep resolving 1.96.0 and the
# fmt-parity contract in CLAUDE.md is untouched. The two echo lines below keep that claim
# honest in the log. They are deliberately NOT `rustup show active-toolchain`: that command
# RESOLVES the toolchain file and would install the whole 1.96.0 toolchain just to print a
# line, in a job where every cargo call is `+$MIRI_TOOLCHAIN` and 1.96.0 is never needed.
# Deliberately NOT `rustup override set` — that writes persistent per-directory state into
# the runner's rustup config, which leaks into unrelated later jobs on a self-hosted fleet.
# Deliberately NOT a second rust-toolchain.toml in a subdirectory — that would apply to
# every cargo invocation under that subtree including fmt, which is the drift the root pin
# exists to prevent.
- name: install the pinned nightly + miri
run: |
git config --global --add safe.directory "$PWD"
rustup toolchain install "$MIRI_TOOLCHAIN" \
--profile minimal \
--component miri,rust-src \
--target x86_64-pc-windows-msvc
echo "root pin, untouched by this job: $(grep -E '^channel' rust-toolchain.toml)"
cargo +"$MIRI_TOOLCHAIN" --version
# A run that reports `0 passed` is a selection that matched nothing, not a success — that
# exact mistake has already cost one round-trip here. So each step below checks a zero exit
# AND that at least one target reported a non-zero pass count, which is what catches a
# crate rename or a `--` filter that stops matching. (Each step legitimately prints several
# `0 passed` lines too — the empty bin/doctest targets — so the check is "at least one
# non-zero", not "no zeroes".) Expected counts at the time of writing: 21 + 12 + 4.
- name: miri — FFI-free leaf crates (native)
run: |
set -o pipefail
ok=1
cargo +"$MIRI_TOOLCHAIN" miri test \
-p pf-driver-proto -p pf-host-config -p pf-gpu 2>&1 | tee /tmp/miri-native.log || ok=0
grep -qE 'test result: ok\. [1-9][0-9]* passed' /tmp/miri-native.log || ok=0
[ "$ok" = 1 ] || echo "::warning::miri (FFI-free leaf crates, native) did not pass — non-blocking; see punktfunk-planning design/rust-safety-programme.md §7"
# THE step that justifies the job: pf-driver-proto at MSVC layout, on Linux, no Windows box.
# Expected: 21 passed. If this one ever goes red, treat it as a layout-contract break
# between the host and driver workspaces until proven otherwise.
- name: miri — pf-driver-proto at x86_64-pc-windows-msvc layout
run: |
set -o pipefail
ok=1
cargo +"$MIRI_TOOLCHAIN" miri test \
-p pf-driver-proto --target x86_64-pc-windows-msvc 2>&1 | tee /tmp/miri-msvc.log || ok=0
grep -qE 'test result: ok\. [1-9][0-9]* passed' /tmp/miri-msvc.log || ok=0
[ "$ok" = 1 ] || echo "::warning::miri (pf-driver-proto @ MSVC layout) did not pass — non-blocking, but this is the layout oracle for every frame and IOCTL; see design/rust-safety-programme.md §7"
# fec-rs dispatches its GF(2^8) multiply through RUNTIME `is_x86_feature_detected!`. Under
# Miri that detection reports the COMPILE-TIME target features, so WITHOUT these RUSTFLAGS
# the step silently interprets the scalar fallback and is worthless. Verified both ways on
# 192.168.1.25: bare, `avx2=false ssse3=false`; with the flags, `avx2=true ssse3=true` and
# `_mm256_shuffle_epi8` genuinely executes under the interpreter. GFNI stays false either
# way — Miri does not implement it — so the gfni branch is simply not covered here.
#
# ⚠ x86_64 ONLY, and it must stay that way. A RUSTFLAGS env var OVERRIDES config rustflags
# ENTIRELY (.cargo/config.toml:11-13 says so), and that config carries `--cfg aes_armv8` /
# `--cfg polyval_armv8` for aarch64 — worth a measured ~3x decrypt-throughput cliff if
# dropped. Harmless here because this job pins ubuntu-24.04/x86_64; fatal on mac-mini-1.
# Narrow selection is mandatory, not an optimisation: see the punktfunk-core note above.
- name: miri — punktfunk-core fec::gf8, taking the real AVX2/SSSE3 branches
env:
RUSTFLAGS: -C target-feature=+avx2,+ssse3
run: |
set -o pipefail
ok=1
cargo +"$MIRI_TOOLCHAIN" miri test \
-p punktfunk-core --lib -- fec::gf8 2>&1 | tee /tmp/miri-gf8.log || ok=0
grep -qE 'test result: ok\. [1-9][0-9]* passed' /tmp/miri-gf8.log || ok=0
[ "$ok" = 1 ] || echo "::warning::miri (punktfunk-core fec::gf8, AVX2/SSSE3) did not pass — non-blocking; see design/rust-safety-programme.md §7"
# ASAN + LSAN over the C ABI harness — §6.1 of design/rust-safety-programme.md, its rank-1
# tooling item. crates/punktfunk-core/tests/c/run.sh already proves the staticlib links and
# round-trips 4 frames byte-exact from C on every push (ci.yml); PF_SAN=address rebuilds BOTH
# sides instrumented — the staticlib on nightly with -Zsanitizer/-Zbuild-std (std itself
# included), the harness with clang -fsanitize — so ASAN sees the seam a Rust-only tool cannot,
# and LSAN (detect_leaks=1, the script's default) becomes the one automated check on abi.rs's
# Box::into_raw/from_raw leak contract.
# Proven to fail on 192.168.1.25: deleting a single punktfunk_session_free() from harness.c
# makes LSAN report the ~308 Rust-side allocations behind the handle and run.sh exit 1.
# What it does NOT see: the invalid-InputKind-discriminant UB at abi.rs (that needs the
# validator, tracked in §5 of the programme doc), and nothing GPU/Windows — this is the
# default-feature (quic-less, opus-less) core only.
c-abi-asan:
runs-on: ubuntu-24.04
container:
image: 192.168.1.58:5010/punktfunk-rust-ci:latest
timeout-minutes: 30
env:
# The SAME dated pin as the miri job above, deliberately — one nightly date to bump for
# both jobs (they have no toolchain interaction; sharing the date just halves the chores).
SAN_TOOLCHAIN: nightly-2026-08-10
# Same guard as the miri job: audit.yml sets no sccache today, and -Zbuild-std could not
# use it anyway. Keeps a future workflow-level sccache from becoming a puzzle.
RUSTC_WRAPPER: ""
steps:
- uses: actions/checkout@v4
# Own `san-` key prefixes — never shared with the miri caches, per the cache-poisoning
# note there (and so an incomplete save from one job can never starve the other).
- name: cache the nightly toolchain
uses: actions/cache@v4
with:
path: /usr/local/rustup/toolchains/${{ env.SAN_TOOLCHAIN }}-x86_64-unknown-linux-gnu
key: san-toolchain-v1-${{ env.SAN_TOOLCHAIN }}
- name: cache the cargo registry
uses: actions/cache@v4
with:
path: /usr/local/cargo/registry
key: san-registry-v1-${{ hashFiles('Cargo.lock') }}
restore-keys: san-registry-v1-
# rust-src is required: -Zbuild-std compiles std from source so it is instrumented too —
# without that, LSAN cannot attribute allocations made inside std (Vec, Box, HashMap).
- name: install the pinned nightly + rust-src
run: |
git config --global --add safe.directory "$PWD"
rustup toolchain install "$SAN_TOOLCHAIN" --profile minimal --component rust-src
echo "root pin, untouched by this job: $(grep -E '^channel' rust-toolchain.toml)"
cargo +"$SAN_TOOLCHAIN" --version
# The image installs clang but Ubuntu does not always pull the compiler-rt sanitizer
# runtime with it (verified absent on a stock 26.04 box). Probe with an actual ASAN link
# and self-heal via apt if it fails — container jobs on this fleet run as root (the
# bun-audit job's apt-get above relies on the same fact).
- name: ensure clang's ASAN runtime
run: |
if ! echo 'int main(void){return 0;}' | clang -fsanitize=address -x c - -o /tmp/asan-probe 2>/dev/null; then
apt-get update && apt-get install -y --no-install-recommends "libclang-rt-$(clang -dumpversion | cut -d. -f1)-dev"
echo 'int main(void){return 0;}' | clang -fsanitize=address -x c - -o /tmp/asan-probe
fi
# run.sh handles everything behind PF_SAN (nightly build, target path, clang flags,
# ASAN_OPTIONS=detect_leaks=1) and exits non-zero on any report. The grep is the
# proved-it-ran guard, same reasoning as the miri steps: a script change that silently
# skips the harness must not read as green. run.sh expects bash and PATH cargo — both true
# in this container. PF_SAN_TOOLCHAIN pins the script's `cargo +<toolchain>` to the dated
# nightly installed above — without it the script would ask for the ROLLING `nightly`
# channel, which this job deliberately does not install.
- name: C ABI harness under ASAN+LSAN
run: |
set -o pipefail
ok=1
PF_SAN=address PF_SAN_TOOLCHAIN="$SAN_TOOLCHAIN" \
bash crates/punktfunk-core/tests/c/run.sh 2>&1 | tee /tmp/asan-harness.log || ok=0
grep -q 'PASS: 4 frames round-tripped byte-exact' /tmp/asan-harness.log || ok=0
[ "$ok" = 1 ] || echo "::warning::c-abi-asan did not pass — non-blocking on day one; see design/rust-safety-programme.md §6.1. An LSAN report here means the abi.rs into_raw/from_raw contract broke."
+24 -2
View File
@@ -111,9 +111,31 @@ jobs:
- name: Format
run: cargo fmt --all --check
# rust-safety WP2c: three textual gates for classes no lint covers — unsafe fn markers
# carrying no contract, panic across an extern boundary (an abort since 1.81), and
# process-global safe APIs (env::set_var & co, count-ratcheted). Pure grep/awk, no cargo.
# Both failure modes were demonstrated before this became blocking (planted instances).
- name: Unsafe-hygiene grep gates
run: sh scripts/ci/check-unsafe-hygiene.sh
- name: Clippy (deny warnings)
run: cargo clippy --workspace --all-targets --locked -- -D warnings
# WP19 (rust-safety): the hardened NATIVE-ONLY host — no Moonlight-compat planes, no
# `rusty_enet` (transpiled C ENet), no `rsa`. Kept compiling here so the cfg boundary can't
# rot, and the dependency claim is ASSERTED, not assumed: `cargo tree -i` must find neither
# crate in the native-only graph (it exits non-zero with "nothing depends on" — inverted).
- name: Clippy + tree (native-only host, no gamestream feature)
run: |
cargo clippy -p punktfunk-host --no-default-features --features pyrowave \
--all-targets --locked -- -D warnings
if cargo tree -p punktfunk-host --no-default-features --features pyrowave \
--locked -i rusty_enet 2>/dev/null | grep -q rusty_enet; then
echo "native-only build still depends on rusty_enet"; exit 1; fi
if cargo tree -p punktfunk-host --no-default-features --features pyrowave \
--locked -i rsa 2>/dev/null | grep -q "^rsa"; then
echo "native-only build still depends on rsa"; exit 1; fi
- name: Build
run: cargo build --workspace --locked
@@ -124,8 +146,8 @@ jobs:
# `nvenc` gates enc/linux/nvenc_cuda.rs (+ nvenc_core/nvenc_status) and `vulkan-encode` gates
# enc/linux/vulkan_video.rs (+ the vendored vk_av1_encode/vk_valve_rgb bindings) — ~8,150
# lines carrying ~70 `unsafe` blocks. Their ONLY prior CI coverage was deb.yml's
# `cargo build`, where warnings are not errors, so pf-encode's own
# `#![deny(clippy::undocumented_unsafe_blocks)]` — the crate's stated unsafe-proof gate —
# `cargo build`, where warnings are not errors, so the `undocumented_unsafe_blocks` deny
# (now hoisted into [workspace.lints]) — pf-encode's stated unsafe-proof gate —
# was never actually enforced on them. (`pyrowave` needs no extra step: punktfunk-host has
# `default = ["pyrowave"]`, so the steps above already cover it.)
#
+4 -3
View File
@@ -159,9 +159,10 @@ jobs:
# The gamepad drivers' business logic is 100% safe (it moved onto pf-umdf-util, the audited
# unsafe layer); pf-vdisplay + wdk-iddcx are inherently FFI-bound but every `unsafe {}` carries a
# `// SAFETY:` proof. Both invariants are lint-gated (`unsafe_op_in_unsafe_fn` +
# `undocumented_unsafe_blocks`); this step keeps them from regressing. (wdk-probe is a
# toolchain-only probe crate and is excluded.)
run: cargo clippy -p pf-umdf-util -p pf-xusb -p pf-gamepad -p pf-mouse -p wdk-iddcx -p pf-vdisplay --all-targets -- -D warnings
# `undocumented_unsafe_blocks`); this step keeps them from regressing. wdk-probe is a
# toolchain-only probe crate, but it holds real DDI slot-dispatch unsafe (iddcx_rt.rs), so it
# runs the same gates.
run: cargo clippy -p pf-umdf-util -p pf-xusb -p pf-gamepad -p pf-mouse -p wdk-iddcx -p pf-vdisplay -p wdk-probe --all-targets -- -D warnings
- name: cargo fmt --check the safe-layer + gamepad/mouse drivers
run: cargo fmt -p pf-umdf-util -p pf-xusb -p pf-gamepad -p pf-mouse --check
- name: Inspect /INTEGRITYCHECK (before) — expect FORCE_INTEGRITY set by wdk-build
+121
View File
@@ -14,6 +14,98 @@ with the version table of the release you are moving to, then read **Breaking ch
## v0.27.1 — in development
### GameStream is now opt-in on EVERY route (⚠ packager-visible default change)
The secure native-only host is the default everywhere; the Moonlight-compat planes (plain-HTTP
pairing + the legacy GCM path, security-review #5/#9) are enabled only by an explicit choice:
- **The shipped systemd user unit** (`scripts/punktfunk-host.service`, installed by deb/RPM/Arch/
sysext) runs bare `serve``--gamestream` is no longer baked into `ExecStart`. Opt in via the
new **`PUNKTFUNK_GAMESTREAM=1`** knob in `host.env` (pf-host-config; equivalent to the flag —
either source enables), so no unit editing survives-upgrades dance is needed.
**Upgrade note:** a packaged host that served Moonlight by default becomes native-only until
the operator sets the knob (a hand-made `ExecStart` drop-in keeps winning as before).
- **NixOS module**: `services.punktfunk.host.gamestream` default flipped `true``false`
(module-check gained a "default is native-only" assertion); enabling it still opens the
GameStream firewall ports.
- **Steam Deck installer**: `--gamestream` opts in (was on-by-default with `--no-gamestream`;
the old flag is still accepted as explicit-off).
- Windows was already opt-in (unchecked installer task) and is unchanged.
### The ENet control port now exists only while a pairing does (rust-safety WP0)
`rusty_enet` — a c2rust-style transpile of C ENet, and the host's only pre-auth-reachable unsafe
surface — no longer listens unconditionally: UDP 47999 binds when the paired-client list becomes
non-empty and is torn down when the last pairing is removed (a live client gets the same
TERMINATION+disconnect farewell as a host-side session end). Pairing itself is HTTPS on nvhttp and
never touches the port, so a never-paired `--gamestream` host exposes no ENet at all. En route:
the management API's unpair endpoint never persisted (`save_paired` was missing), so an unpair
lasted only until the next restart — fixed. `rusty_enet` is now pinned `=0.4.0`.
**Unpair is now a complete revocation, on both planes.** Beyond the persistence fix above, an
unpair used to leave the revoked client's LIVE session streaming until the client chose to
leave. Now: unpairing a GameStream client whose certificate owns the active launch ends that
session (the client gets the standard TERMINATION+disconnect, and unpair-all still closes the
ENet port); unpairing a native client deliberately stops its live punktfunk/1 session(s)
(matched by certificate fingerprint — anonymous/TOFU sessions are unaffected, they have no
pairing to revoke). The unpair endpoint's long-standing docstring caveat ("removes the client
from the listing without severing its ability to reconnect") is retired: TLS-level handshakes
still complete by design, but authorization is per-request and a live session no longer
survives its own revocation.
### GameStream is now a cargo feature (compile-time isolation — packager-visible)
The Moonlight-compat planes (nvhttp pairing, RTSP, the ENet control stream, `_nvstream` mDNS,
the compat media path) are gated behind a new **`gamestream` cargo feature — default ON**, so
every stock package is behaviorally identical (GameStream stays runtime-opt-in via
`--gamestream` / `PUNKTFUNK_GAMESTREAM`). Building with
`--no-default-features --features pyrowave` produces the **hardened native-only host**:
- **no `rusty_enet`** — the c2rust-transpiled C ENet stack (158 unsafe sites) is absent from
the binary, provably (`cargo tree -i rusty_enet` finds nothing; CI asserts it);
- **no `rsa`** — the native planes run on the P-256 identity (above), and the legacy-identity
fallback is a pem-only read (rustls/ring serves an existing RSA cert without the crate), so
the accepted Marvin advisory (RUSTSEC-2023-0071) no longer applies to native-only builds;
- ~6,700 lines of Moonlight protocol code gone; `serve --gamestream` (or the env knob) against
such a binary **refuses to start** with a clear error rather than serving less than asked;
- the native-only management API (and its OpenAPI document) has no GameStream PIN endpoints
(`/api/v1/pair`, `/api/v1/pair/pin`); everything else — including the paired-client list and
unpair — is identical, so consoles work unchanged.
The checked-in `api/openapi.json` remains the default-features document.
### The identity split — the native planes get their own (P-256) host identity
One RSA-2048 identity historically served every plane, because Moonlight mandates RSA and the
planes grew out of the GameStream host. The native punktfunk/1 QUIC plane and the management API
now share a separate **ECDSA P-256** identity (`native-cert.pem`/`native-key.pem`): generated by
ring via rcgen, browser-compatible (Ed25519 server certs are not), carrying real SANs
(localhost, loopback, the machine hostname — the legacy cert had none), and free of the accepted
`rsa`-crate Marvin advisory. The GameStream plane keeps the RSA identity untouched.
**Migration is pin-preserving by construction**: clients TOFU-pin the leaf-cert SHA-256 at
pairing and use that one pin for both QUIC and the mgmt/library API, so the new identity is
adopted **only when the native trust store is empty** (fresh installs, or after an explicit
unpair-all + restart). An upgraded host with live native pairings keeps presenting the legacy
RSA cert those clients pinned, and logs the migration path. Fingerprint pinning is
algorithm-agnostic, so existing shipped clients pair against P-256 hosts unchanged.
Follow-the-identity consumers updated in-tree: the tray's loopback pin and the plugin SDK's
mgmt CA now prefer `native-cert.pem` (falling back to `cert.pem`), and the Windows runner ACL
grant covers both. ⚠ A plugin bundling an **older** `@punktfunk/host` SDK on a **fresh**
(P-256) host trusts the wrong cert — set `PUNKTFUNK_MGMT_CA=<config>/native-cert.pem` in its
environment or rebuild against the current SDK.
### Memory-safety, compiler-enforced (embedder-visible lint tightening)
`punktfunk-core` now carries `#![deny(unsafe_code)]` crate-wide: everything that parses network
bytes is safe Rust by compiler-enforced invariant. The documented `#![allow]` carve-outs are the
client surface (`abi`, `client`) and the platform syscall-batching shims under `transport`
(`udp/{apple,linux,windows}`, `qos_windows`) — none of which interpret attacker bytes. In
`punktfunk-host`, the modules a secure-default host exposes (`native`, `native_pairing`, `mgmt`,
`mgmt_token`, `discovery`, `wol`) are `#[forbid(unsafe_code)]`. If you embed `punktfunk-core` and
patch it, new unsafe outside the carve-outs is now a compile error.
### NixOS + KDE — session detection, the other half
🛑 **v0.27.0's NixOS session-detection fix did not reach a stock NixOS + Plasma 6 box.** It resolved
@@ -63,6 +155,35 @@ and `disable_environment` is then consulted last and wins on **presence alone**,
session script never mentions that second variable, so it is the one that survives. Both spellings
go out, on the transient unit and on the box's own session drop-in.
### punktfunk-gamescope `+pfhdr6` — a NO_FOCUS window can no longer steal the composite
🛑 **A mapped-but-unpainted window carrying `GAMESCOPE_NO_FOCUS=1` could win gamescope's focus
selection and turn the composite — and the stream fed from it — black while every health signal
stayed green.** Bazzite's hhd-ui (Handheld Daemon overlay) sets that atom once at init, stamps
Steam's appid, and crash-loops under a headless takeover; each respawn remapped a fullscreen black
window that steamcompmgr then chose over Big Picture (observed on a Bazzite box: client stats
happily decoding 60 fps at 0.1 Mb/s of black; killing hhd-ui restored the picture instantly). No
gamescope — upstream or Bazzite's fork — ever consumed the atom; its setters (hhd-ui, MangoHud)
show and hide via the `STEAM_OVERLAY` protocol and rely on never being focusable. Patch 0008 wires
`GAMESCOPE_NO_FOCUS` exactly like `GAMESCOPE_EXTERNAL_OVERLAY` (read at map, PropertyNotify-tracked,
skipped by both focus-candidate collectors) without touching compositing or `appID`. Banner
`+pfhdr5``+pfhdr6`; no new capability — the bump is so a field box's banner tells the two
behaviors apart.
### Linux capture — the truncated first attempt no longer latches sticky downgrades
🛑 **The pipeline retry loop's deliberately short (2.5 s) first-frame attempt could permanently
downgrade the whole host process.** On expiry, the portal capturer's timeout diagnosis latched
whichever offer it implicated — HDR capture off (per source), the raw-dmabuf offer off, the
EGL→CUDA offer off — as if the compositor had refused it, when the budget was truncated by design
and a gamescope cold start routinely needs longer before delivering anything. One lost race at
connect then pinned every later session to SDR and/or CPU capture until the host restarted. The
truncated attempt is now declared provisional end to end
(`Capturer::next_frame_within_provisional`): its expiry names the same suspect in the error text
but latches nothing; only the full-length attempts that follow hand down negotiation verdicts. The
classification is a pure function with tests
(`pf_capture::linux::first_frame_timeout_tests`).
## v0.27.0
87 commits since v0.26.0.
+11
View File
@@ -101,6 +101,17 @@ repository = "https://git.unom.io/unom/punktfunk"
[workspace.lints.rust]
unsafe_op_in_unsafe_fn = "deny"
# The companion lint: every `unsafe {}` / `unsafe impl` carries a `// SAFETY:` proof. Hoisted here
# from ~85 per-file `#![deny(...)]` attributes so a NEW crate (or a new module in an old one) is
# covered on creation rather than on remembering — the per-file form left pf-vkhdr-layer,
# wdk-probe, and half of pf-clipboard uncovered for months. NOTE: this table reaches only crates
# with `[lints] workspace = true`; `packaging/windows/drivers` and `packaging/windows/pf-vkhdr-layer`
# are SEPARATE workspaces and restate it (any "workspace-wide" claim must be made three times or it
# is false). Of the members, only the two vendored snapshots (pf-bitstream/vendor/cros-codecs,
# punktfunk-host/vendor/usbip-sim) stay out, deliberately — upstream code stays pristine.
[workspace.lints.clippy]
undocumented_unsafe_blocks = "deny"
[profile.release]
opt-level = 3
lto = "thin"
+10 -9
View File
@@ -10,7 +10,7 @@
"name": "MIT OR Apache-2.0",
"identifier": "MIT OR Apache-2.0"
},
"version": "0.26.0"
"version": "0.27.0"
},
"paths": {
"/api/v1/clients": {
@@ -53,7 +53,7 @@
"clients"
],
"summary": "Unpair a client",
"description": "Removes the client's certificate from the pairing store. Caveat: the nvhttp TLS layer\ndoes not yet reject unlisted certificates (`gamestream/tls.rs` accepts any well-formed\nclient cert — a planned hardening step), so until that lands this removes the client\nfrom the listing without severing its ability to reconnect.",
"description": "Removes the client's certificate from the pairing store (persisted — the removal survives a\nhost restart). Revocation is complete: a LIVE GameStream session owned by this certificate is\nended (the client gets the standard TERMINATION+disconnect), and removing the last pairing\nalso closes the ENet control port (UDP 47999), which is only bound while at least one pairing\nexists. The nvhttp TLS layer still completes a handshake with any well-formed client cert BY\nDESIGN (authorization is per-request via the paired-fingerprint check) — an unpaired client\nthat reconnects is rejected at every post-pair endpoint.",
"operationId": "unpairClient",
"parameters": [
{
@@ -4788,7 +4788,7 @@
"version": {
"type": "integer",
"format": "int32",
"description": "Schema version (currently 1) — lets a future field addition migrate rather than reject.",
"description": "Schema version (currently 1) — lets a future field addition migrate rather than reject. Read\nat load time ([`DisplayPolicyStore::load_from`] warns when a file claims a version this host\ndoes not know, then reads it best-effort) and pinned back to the current version on write.",
"minimum": 0
}
}
@@ -4857,7 +4857,7 @@
},
"EffectivePolicy": {
"type": "object",
"description": "The six resolved fields after preset expansion — what the lifecycle/registry and the Stage-0 call\nsites read, and what the mgmt API echoes as the \"currently in force\" policy. Pure output of\n[`DisplayPolicy::effective`].",
"description": "The six resolved fields after preset expansion — what the lifecycle/registry and the policy call\nsites read, and what the mgmt API echoes as the \"currently in force\" policy. Pure output of\n[`DisplayPolicy::effective`].\n\n**Every field is required on the wire, deliberately.** Unlike [`DisplayPolicy`] — which is only\never a *file* — this shape is also the `fields` member of [`CustomPresetInput`], i.e. the request\nbody of `POST /display/presets` and `PUT /display/presets/{id}`, and a *response* member three\ntimes over (`DisplaySettingsState.effective`, `PresetInfo.fields`, `CustomPreset.fields`).\n`#[serde(default)]` here would (a) turn `{\"name\":\"Kiosk\",\"fields\":{}}` — or any camelCase typo —\nfrom a serde rejection into a 201 storing a preset that expands to six axes nobody chose, and\n(b) make all six OPTIONAL in the generated OpenAPI schema, so every codegen'd client has to\nnull-check them. The *persisted* catalog's tolerance for an entry written before an axis existed\nis bought where it belongs, on the read path only: see [`StoredEffectivePolicy`].",
"required": [
"keep_alive",
"topology",
@@ -5915,7 +5915,7 @@
},
"Identity": {
"type": "string",
"description": "Stable display identity, so desktop environments persist per-display config (KDE scaling). Stored\nat Stage 0; carriers wired from the identity stage.",
"description": "Stable display identity, so desktop environments persist per-display config (KDE scaling). The\nslot this resolves to is carried per backend: the Windows EDID serial + IddCx connector index,\nKWin's per-slot output name, and the host-persisted Mutter scale map.",
"enum": [
"shared",
"per-client",
@@ -6132,14 +6132,14 @@
"seconds": {
"type": "integer",
"format": "int32",
"description": "Linger window in seconds.",
"description": "Linger window in seconds, clamped to `0..=86400` on write (see\n[`DisplayPolicy::sanitized`]): a window longer than a day is `forever` by any honest\nreading, and `u32` seconds is ~136 years — a deadline the reaper would never reach and a\nnonsense `expires_in_ms` in `/display/state`.",
"minimum": 0
}
}
},
{
"type": "object",
"description": "Keep the display until host shutdown or an explicit release (the `Pinned` lifecycle state).\n**Not honored until the display-lifecycle stage** — rejected by the mgmt PUT at Stage 0.",
"description": "Keep the display until host shutdown or an explicit release (the `Pinned` lifecycle state).\nHonored end-to-end: the registry resolves it to `Release::Pin`, so the display survives every\ndisconnect — free it with `POST /display/release` (which force-releases `Pinned` exactly like\na `Lingering` display). This is what the `gaming-rig` preset selects.",
"required": [
"mode"
],
@@ -6183,6 +6183,7 @@
},
"positions": {
"type": "object",
"description": "Keys are the **canonical decimal** identity-slot id (`\"1\"`..`\"15\"`) — the exact string\n`arrange` looks a member up by. [`DisplayPolicy::sanitized`] re-canonicalizes them on write\n(`\"01\"` → `\"1\"`) and drops anything that is not a slot id, because a key that never matches is\na pin the operator can see in the console and in `GET /display/settings` while every session\nsilently auto-rows past it.",
"additionalProperties": {
"$ref": "#/components/schemas/Position"
},
@@ -6194,7 +6195,7 @@
},
"LayoutMode": {
"type": "string",
"description": "How group members are arranged in the desktop coordinate space. Stored at Stage 0; applied from\nthe multi-monitor stage.",
"description": "How group members are arranged in the desktop coordinate space, resolved by `layout::arrange` —\nwhich both the `/display/state` readout and (on Linux, KWin only) the per-backend position apply\nconsume, so the answer is computed in exactly one place.",
"enum": [
"auto-row",
"manual"
@@ -6354,7 +6355,7 @@
},
"ModeConflict": {
"type": "string",
"description": "Admission when a *different* client connects while a display/session is already live and asks for\na different mode. Stored at Stage 0; enforced from the mode-conflict admission stage.",
"description": "Admission when a *different* client connects while a display/session is already live and asks for\na different mode. Enforced by [`super::admission`] before the Welcome is sent, so a `reject` is a\nclean handshake error rather than a half-built session.",
"enum": [
"separate",
"steal",
@@ -9,7 +9,7 @@ use punktfunk_core::config::{CompositorPref, GamepadPref, Mode};
use std::sync::{Arc, Mutex};
use std::time::Duration;
use super::{hex32, jni_guard, parse_hex32, SessionHandle};
use super::{hex32, jni_guard, lock_recover, parse_hex32, SessionHandle};
/// Machine token of the most recent `nativeConnect`/`nativePair` failure, taken (and cleared)
/// by `nativeTakeLastError` so Kotlin can render a cause-specific message instead of the old
@@ -41,7 +41,7 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeTakeLastErr
env: JNIEnv<'local>,
_this: JObject<'local>,
) -> jni::sys::jstring {
let token = std::mem::take(&mut *LAST_ERROR.lock().unwrap());
let token = std::mem::take(&mut *lock_recover(&LAST_ERROR));
match env.new_string(token) {
Ok(s) => s.into_raw(),
Err(_) => JObject::null().into_raw(),
@@ -45,6 +45,15 @@ pub(crate) fn jni_guard<T>(default: T, f: impl FnOnce() -> T) -> T {
})
}
/// Poison-recovering lock for the JNI entry points that are NOT behind [`jni_guard`]: a
/// `.lock().unwrap()` there turns a poisoned mutex into a panic across the `extern "system"`
/// boundary — an abort of the whole app on Rust ≥ 1.81 (the panic-in-extern grep gate's class).
/// The slots behind these mutexes are plane-thread handles and last-value caches; whatever a
/// poisoned writer left is still valid to inspect or replace.
pub(crate) fn lock_recover<T>(m: &Mutex<T>) -> std::sync::MutexGuard<'_, T> {
m.lock().unwrap_or_else(std::sync::PoisonError::into_inner)
}
/// A live session behind the `jlong` handle: the connector + the decode thread it feeds.
pub(crate) struct SessionHandle {
// Read only by the android decode path (`nativeStartVideo` → `crate::decode`); on the host
+7 -7
View File
@@ -8,7 +8,7 @@ use jni::objects::JString;
use jni::sys::{jboolean, jdoubleArray, jintArray, jlong, jsize, jstring};
use jni::JNIEnv;
use super::{jni_guard, SessionHandle};
use super::{jni_guard, lock_recover, SessionHandle};
/// `NativeBridge.nativeStartVideo(handle, surface, decoderName, lowLatencyMode, lowLatencyFeature,
/// isTv, presentPriority, smoothBuffer)` — wrap the SurfaceView's `Surface` as an `ANativeWindow`
@@ -48,7 +48,7 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeStartVideo(
.filter(|s| !s.is_empty());
// SAFETY: live handle per the nativeConnect/nativeClose contract.
let h = unsafe { &*(handle as *const SessionHandle) };
let mut guard = h.video.lock().unwrap();
let mut guard = lock_recover(&h.video);
if guard.is_some() {
return; // already streaming
}
@@ -222,7 +222,7 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeVideoStats(
}
// SAFETY: live handle per the nativeConnect/nativeClose contract.
let h = unsafe { &*(handle as *const SessionHandle) };
if h.video.lock().unwrap().is_none() {
if lock_recover(&h.video).is_none() {
return std::ptr::null_mut(); // not streaming → no stats
}
let snap = h
@@ -385,7 +385,7 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeStartAudio(
}
// SAFETY: live handle per the nativeConnect/nativeClose contract.
let h = unsafe { &*(handle as *const SessionHandle) };
let mut guard = h.audio.lock().unwrap();
let mut guard = lock_recover(&h.audio);
if guard.is_some() {
return; // already playing
}
@@ -434,7 +434,7 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeStartMic(
}
// SAFETY: live handle per the nativeConnect/nativeClose contract.
let h = unsafe { &*(handle as *const SessionHandle) };
let mut guard = h.mic.lock().unwrap();
let mut guard = lock_recover(&h.mic);
if let Some(m) = guard.as_ref() {
return m.session_id(); // already capturing — same stream, same session
}
@@ -516,7 +516,7 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeStartPadAud
speaker != 0,
) {
Some(p) => {
*h.pad_audio.lock().unwrap() = Some(p);
*lock_recover(&h.pad_audio) = Some(p);
1
}
None => 0,
@@ -629,6 +629,6 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeMicActive(
}
// SAFETY: live handle per the nativeConnect/nativeClose contract.
let h = unsafe { &*(handle as *const SessionHandle) };
jboolean::from(h.mic.lock().unwrap().is_some())
jboolean::from(lock_recover(&h.mic).is_some())
})
}
+7 -1
View File
@@ -176,7 +176,13 @@ unsafe extern "system" fn wnd_proc(
let slice = unsafe { std::slice::from_raw_parts(cds.lpData as *const u16, len) };
let url = String::from_utf16_lossy(slice);
tracing::debug!(%url, "link from another instance");
INBOX.lock().unwrap().push(url);
// Poison-recover, never unwrap: a panic out of a window procedure is an abort since
// Rust 1.81, and the inbox is a plain Vec that stays valid whatever a poisoned
// writer left behind.
INBOX
.lock()
.unwrap_or_else(std::sync::PoisonError::into_inner)
.push(url);
return LRESULT(1);
}
}
-1
View File
@@ -15,7 +15,6 @@
//! (measure the path: probe burst → goodput / loss / recommended bitrate)
// Unsafe-proof program: every `unsafe {}` in this client carries a `// SAFETY:` proof.
#![deny(clippy::undocumented_unsafe_blocks)]
// Link as a GUI (windows) subsystem binary so the default windowed launch (MSIX / double-click)
// does NOT pop a console window. The CLI paths (--headless/--discover) reattach to the launching
// terminal's console at startup (see main), so their output is still visible when run from a shell.
+7
View File
@@ -10,6 +10,11 @@
#![allow(non_snake_case)]
// Bindgen output for a C API: u128 layout warnings and the like are upstream's concern.
#![allow(improper_ctypes)]
// The workspace-wide undocumented_unsafe_blocks deny cannot apply to GENERATED code: bindgen
// emits `unsafe {}` in layout tests/accessors and nobody hand-writes proofs into OUT_DIR. This
// crate is bindings-only by charter (the safe wrapper lives with the consumer), so the allow is
// crate-wide; the hand-written link-sanity test below still carries its proof by convention.
#![allow(clippy::undocumented_unsafe_blocks)]
// Generated code — clippy findings in it (missing safety docs on generated unsafe fns, style
// nits across 14k lines) are bindgen's shape, not ours; the safe wrapper in pf-encode is the
// linted surface.
@@ -27,6 +32,8 @@ mod tests {
/// implementations — that's fine, MFXLoad itself must still succeed).
#[test]
fn dispatcher_links_and_loads() {
// SAFETY: MFXLoad allocates the dispatcher's loader context (documented to work with no
// driver present) and MFXUnload frees that same non-null handle; nothing else is touched.
unsafe {
let loader = MFXLoad();
assert!(!loader.is_null(), "MFXLoad returned NULL");
+15 -7
View File
@@ -7,13 +7,6 @@
//! [`FrameChannelSender`] closure, so this crate reaches neither the encoder nor the host
//! orchestrator).
// Every unsafe block in this crate carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
// …and that program only covers a whole `unsafe fn` body once the body needs its own block: in
// edition 2021 `unsafe_op_in_unsafe_fn` is allow-by-default, which exempted the crate's hardest FFI
// (the ring/slot construction, the channel broker, every D3D converter ctor) from the deny above.
#![deny(unsafe_op_in_unsafe_fn)]
use anyhow::Result;
use pf_frame::{CapturedFrame, FramePayload, PixelFormat};
// The Linux capturer reaches `DmabufFrame` through `super::`; `CursorOverlay` it names directly as
@@ -43,6 +36,21 @@ pub trait Capturer: Send {
self.next_frame()
}
/// [`next_frame_within`](Self::next_frame_within), but the caller declares the budget
/// PROVISIONAL: its expiry is the retry schedule firing (the deliberately truncated first
/// attempt), not a verdict on anything this capture offered. The portal backend must NOT
/// latch its sticky process-wide downgrades (HDR capture, either dmabuf-only offer) from a
/// provisional expiry — a gamescope cold start routinely outlives the short window while it
/// would have accepted every offer, and one latched race used to pin the whole host process
/// to SDR/CPU capture. The full-length attempt that follows delivers the honest verdict.
/// Backends that latch nothing from a timeout just delegate.
fn next_frame_within_provisional(
&mut self,
budget: std::time::Duration,
) -> Result<CapturedFrame> {
self.next_frame_within(budget)
}
/// Non-blocking: the freshest frame available since the last call, or `None` if none has
/// arrived (the caller reuses its last frame to hold a steady output rate). The default
/// just produces a frame each call — fine for instant synthetic sources; the portal
+283 -73
View File
@@ -1,4 +1,4 @@
//! Live capture: xdg ScreenCast portal (`ashpd`) → PipeWire (`pipewire`), CPU-copy path.
//! Live capture: xdg ScreenCast portal (`ashpd`) → PipeWire (`pipewire`).
//!
//! Two dedicated threads, because both stacks are tied to their thread:
//! * **portal thread** drives the async ashpd handshake on a multi-thread tokio runtime
@@ -7,9 +7,13 @@
//! drops; ashpd's `Session` has no `Drop`);
//! * **pipewire thread** owns the (`!Send`) MainLoop/Stream and pumps frames.
//!
//! The portal hands the PipeWire remote fd + node id to the pipewire thread; decoded BGRx
//! frames leave the pipewire thread over a bounded channel. The authoritative frame size
//! comes from the negotiated PipeWire format, not the portal's size hint.
//! The portal hands the PipeWire remote fd + node id to the pipewire thread; frames leave that
//! thread through a ONE-DEEP OVERWRITING slot (`FrameSlot`) plus a wakeup edge — not the bounded
//! `sync_channel(8)` this once used, which was drop-NEWEST and so handed a stalled consumer stale
//! frames (see `FrameSlot`'s own note). The payload is not necessarily BGRx either: the negotiation
//! can settle on packed RGB, NV12, YUV444 or 10-bit PQ, and on a dmabuf passthrough it never touches
//! the CPU. The authoritative frame size comes from the negotiated PipeWire format, not the portal's
//! size hint.
//!
//! Cleanup: BOTH threads are stopped deterministically — [`PortalCapturer`]'s `Drop` sends a
//! pipewire `channel` quit and joins that thread (releasing its EGL importer / CUDA context
@@ -18,8 +22,9 @@
//! connection and so ENDS the compositor's ScreenCast session. Dropping a capturer (session end,
//! or a retried/failed pipeline build) therefore leaves nothing behind on either side.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
// Every `unsafe` block in this module TREE carries a `// SAFETY:` proof; enforce it (unsafe-proof
// program). This file itself has none — the FFI lives in the child modules declared at the bottom
// (`pipewire`, `pw_cursor`, `pw_pods`, `portal`, `xfixes_cursor`), which this inner attribute covers.
use super::{CapturedFrame, Capturer, DmabufFrame, FramePayload, PixelFormat, ZeroCopyPolicy};
use anyhow::{anyhow, Context, Result};
@@ -173,8 +178,9 @@ pub struct PortalCapturer {
/// capture, not per frame.
negotiation_confirmed: bool,
/// This capture ran the HDR (10-bit PQ/BT.2020 dmabuf) offer — see [`Self::open`]'s
/// `want_hdr`. Read by the negotiation-timeout diagnosis (a failed HDR offer latches the
/// process-wide SDR downgrade) and by [`hdr_meta`](Capturer::hdr_meta).
/// `want_hdr`. Read by the negotiation-timeout diagnosis (a failed HDR offer latches the SDR
/// downgrade for THIS [`Self::hdr_source`] only, not process-wide) and by
/// [`hdr_meta`](Capturer::hdr_meta).
hdr_offer: bool,
/// Which HDR source this capturer is — the latch a failed [`hdr_offer`](Self::hdr_offer)
/// belongs to. See [`super::HdrSource`] for why the latch is not one process-wide flag.
@@ -463,7 +469,10 @@ fn spawn_pipewire(
let zerocopy = allow_zerocopy && pf_zerocopy::enabled();
// HDR cannot ride the SHM path (see `want_hdr` above): under PUNKTFUNK_FORCE_SHM the HDR
// offer is dropped — SDR capture, loudly.
let force_shm = std::env::var("PUNKTFUNK_FORCE_SHM").as_deref() == Ok("1");
// The shared parser, not a bare `== "1"` compare — matching `PUNKTFUNK_PIPEWIRE_NV12` below.
// A bare compare silently ignored `PUNKTFUNK_FORCE_SHM=true`/`=on`/`=yes`, so the knob looked
// set and did nothing.
let force_shm = pf_host_config::env_on("PUNKTFUNK_FORCE_SHM").unwrap_or(false);
let want_hdr = if want_hdr && force_shm {
tracing::warn!(
"HDR capture requested but PUNKTFUNK_FORCE_SHM=1 — the SHM path is 8-bit only; \
@@ -533,7 +542,7 @@ fn spawn_pipewire(
impl Capturer for PortalCapturer {
fn next_frame(&mut self) -> Result<CapturedFrame> {
self.frame_within(Duration::from_secs(10))
self.frame_within(Duration::from_secs(10), TimeoutVerdict::Conclusive)
}
fn cursor(&mut self) -> Option<pf_frame::CursorOverlay> {
@@ -555,6 +564,14 @@ impl Capturer for PortalCapturer {
// every nested Xwayland the provider reports, RE-RUNS the provider so a game's Xwayland
// that appears later is adopted, and follows whichever one gamescope draws the pointer on.
// `frame_size` lets it map root-space coordinates into frame space.
//
// Idempotent by construction. The contract says "called once", but nothing enforced it, and a
// second call evaluated `spawn` BEFORE dropping the old source: two readers then published
// into the same slot for the construction window, and a `spawn` that returned `None` destroyed
// a perfectly good reader outright.
if self._gs_cursor.is_some() {
return;
}
self._gs_cursor = xfixes_cursor::XFixesCursorSource::spawn(
targets,
Arc::clone(&self.signals.cursor_live),
@@ -563,7 +580,13 @@ impl Capturer for PortalCapturer {
}
fn next_frame_within(&mut self, budget: Duration) -> Result<CapturedFrame> {
self.frame_within(budget)
self.frame_within(budget, TimeoutVerdict::Conclusive)
}
fn next_frame_within_provisional(&mut self, budget: Duration) -> Result<CapturedFrame> {
// The retry loop's truncated first attempt: its expiry re-runs the schedule, it does not
// convict an offer — see `TimeoutVerdict` and the latch arms in `next_frame_timed_out`.
self.frame_within(budget, TimeoutVerdict::Provisional)
}
fn supports_arrival_wait(&self) -> bool {
@@ -655,6 +678,11 @@ impl Capturer for PortalCapturer {
if let Ok(mut slot) = self.slot.lock() {
*slot = None;
}
// Clear the stall clock for the same reason the mailbox is flushed: a pooled capturer
// whose previous stream ended mid-stall carried that `Instant` into the next one, so the
// first `try_latest` that saw `!streaming` found the 1500 ms grace already expired and
// reported capture loss on a stream that had been running for microseconds.
self.stall_since = None;
}
}
@@ -699,12 +727,73 @@ impl Capturer for PortalCapturer {
}
}
/// Whether an expired first-frame budget is allowed to CONVICT an offer. The retry loop's
/// deliberately truncated first attempt passes `Provisional`: its expiry means the schedule
/// moved on, not that the compositor refused anything — a gamescope cold start regularly needs
/// longer than that window to accept every offer it would have accepted. Latching from it pinned
/// the whole host process to SDR + CPU capture off a race the attempt lost by design; only a
/// full-length wait carries a verdict.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum TimeoutVerdict {
Conclusive,
Provisional,
}
/// Which offer a first-frame timeout implicates — the diagnosis behind
/// [`PortalCapturer::next_frame_timed_out`], split out pure so the latch policy is testable.
/// Mirrors the negotiation state exactly: a negotiated format clears every offer (the compositor
/// accepted, it just produced nothing), and a forced `PUNKTFUNK_ZEROCOPY=1` keeps both dmabuf
/// arms erroring loudly instead of implicating them (the operator asked for exactly that path).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum TimeoutOffer {
/// Format negotiated; no offer implicated — the compositor produced no buffers.
NoBuffers,
/// The 10-bit PQ/BT.2020 (HDR) dmabuf offer was never accepted.
Hdr,
/// The dmabuf-only raw-passthrough offer was never accepted.
RawDmabuf,
/// The dmabuf-only EGL→CUDA offer was never accepted.
GpuDmabuf,
/// Nothing negotiated and no offer implicated — format/modifier mismatch.
NoFormat,
}
fn classify_first_frame_timeout(
negotiated: bool,
hdr_offer: bool,
vaapi_dmabuf: bool,
gpu_dmabuf_offer: bool,
zerocopy_forced: bool,
) -> TimeoutOffer {
if negotiated {
TimeoutOffer::NoBuffers
} else if hdr_offer {
TimeoutOffer::Hdr
} else if vaapi_dmabuf && !zerocopy_forced {
TimeoutOffer::RawDmabuf
} else if gpu_dmabuf_offer && !zerocopy_forced {
TimeoutOffer::GpuDmabuf
} else {
TimeoutOffer::NoFormat
}
}
/// The latch policy: only a conclusive expiry of an offer-implicating timeout fires the offer's
/// sticky process-wide downgrade.
fn timeout_convicts(offer: TimeoutOffer, verdict: TimeoutVerdict) -> bool {
verdict == TimeoutVerdict::Conclusive
&& matches!(
offer,
TimeoutOffer::Hdr | TimeoutOffer::RawDmabuf | TimeoutOffer::GpuDmabuf
)
}
impl PortalCapturer {
/// The blocking first-frame wait behind [`Capturer::next_frame`] /
/// [`Capturer::next_frame_within`]. First frame can lag behind format negotiation; later
/// frames arrive at ~fps. Wait in short slices so a GPU-import poison (worker death) fails
/// the capture within ~0.5 s instead of sitting out the full first-frame budget.
fn frame_within(&mut self, budget: Duration) -> Result<CapturedFrame> {
fn frame_within(&mut self, budget: Duration, verdict: TimeoutVerdict) -> Result<CapturedFrame> {
let deadline = std::time::Instant::now() + budget;
loop {
if self.signals.broken.load(Ordering::Relaxed) {
@@ -730,7 +819,7 @@ impl PortalCapturer {
if let Some(f) = self.take_frame() {
return Ok(f);
}
return self.next_frame_timed_out(e, budget);
return self.next_frame_timed_out(e, budget, verdict);
}
}
}
@@ -752,83 +841,118 @@ impl PortalCapturer {
}
/// The [`frame_within`](Self::frame_within) budget expired (or the thread ended) — turn it
/// into the diagnosis-bearing error. Split out of the slicing loop above; behavior unchanged.
/// into the diagnosis-bearing error, and fire the offer's sticky downgrade latch when — and
/// only when — the expiry convicts the offer (see [`timeout_convicts`]).
fn next_frame_timed_out(
&self,
err: RecvTimeoutError,
budget: Duration,
verdict: TimeoutVerdict,
) -> Result<CapturedFrame> {
let within = budget.as_secs_f32();
match err {
RecvTimeoutError::Timeout => {
// Split the two black-screen root causes apart so the operator gets a cause, not
// just a symptom: did the format negotiate (compositor produced no buffers) or
// not (no acceptable format / node never emitted a param)?
if self.signals.negotiated.load(Ordering::Relaxed) {
Err(anyhow!(
let offer = classify_first_frame_timeout(
self.signals.negotiated.load(Ordering::Relaxed),
self.hdr_offer,
self.vaapi_dmabuf,
self.signals.gpu_dmabuf_offer.load(Ordering::Relaxed),
pf_zerocopy::zerocopy_forced(),
);
let convicted = timeout_convicts(offer, verdict);
// A provisional expiry names the same suspect but hands down no sentence — the
// full-length retry that follows is the one whose timeout latches.
let sentence = if convicted {
"" // each arm below states its own downgrade
} else {
" (short first-attempt window — nothing is latched; the full-length retry \
decides)"
};
match offer {
TimeoutOffer::NoBuffers => Err(anyhow!(
"no PipeWire frame within {within}s (node {}): format negotiated but no \
buffers arrived the compositor produced no frames (virtual output \
idle/unmapped, capture never started, or a stream bound during a \
compositor (re)start that will never deliver a reconnect fixes that)",
self.node_id
))
} else if self.hdr_offer {
// The HDR (10-bit PQ dmabuf) offer was never accepted — the monitor left HDR
// mode between the probe and the negotiation, the compositor pre-dates the
// GNOME 50 HDR formats, or its allocator can't do LINEAR for XR30/XB30.
// Latch the process-wide SDR downgrade so the next session (Moonlight
// auto-reconnects) negotiates SDR instead of re-running this same timeout.
super::note_hdr_capture_failed(self.hdr_source);
Err(anyhow!(
"no PipeWire frame within {within}s (node {}): the compositor never \
accepted the HDR (10-bit PQ/BT.2020 dmabuf) offer is the mirrored \
monitor in HDR mode on GNOME 50+? Downgrading this host to SDR capture; \
reconnect to stream SDR",
self.node_id
))
} else if self.vaapi_dmabuf && !pf_zerocopy::zerocopy_forced() {
// The dmabuf-only raw-passthrough offer was never accepted. Latch the
// downgrade so the encode loop's pipeline rebuild retries on the CPU offer
// instead of failing this same negotiation forever. The latch is SCOPED to the
// raw-passthrough decision: it used to be `note_vaapi_dmabuf_failed`, which fed
// `pf_zerocopy::enabled()` and therefore dropped every later session on this
// host — NVENC's EGL→CUDA path included — to CPU capture. Since this offer is
// also the PyroWave one (any vendor), a single PyroWave negotiation timeout was
// enough to do that.
pf_zerocopy::note_raw_dmabuf_negotiation_failed();
Err(anyhow!(
"no PipeWire frame within {within}s (node {}): the compositor never \
accepted the dmabuf-only offer (raw-dmabuf passthrough) downgrading \
THIS path to CPU capture for the rest of the process; the pipeline \
rebuild will renegotiate without dmabuf",
self.node_id
))
} else if self.signals.gpu_dmabuf_offer.load(Ordering::Relaxed)
&& !pf_zerocopy::zerocopy_forced()
{
// The EGL→CUDA dmabuf-only offer was never accepted — the twin of the raw-
// passthrough arm above (the offer the thread ACTUALLY made, per the signal
// it set — see `CaptureSignals::gpu_dmabuf_offer`). One timeout is conclusive:
// a compositor that allocates none of the importer's modifiers refuses them
// identically on every retry, so latch the offer off and let the pipeline
// rebuild renegotiate the CPU path instead of re-running this same 10 s
// timeout on every reconnect. A forced PUNKTFUNK_ZEROCOPY=1 keeps erroring
// loudly instead (same rule as the raw arm).
pf_zerocopy::note_gpu_dmabuf_negotiation_failed();
Err(anyhow!(
"no PipeWire frame within {within}s (node {}): the compositor never \
accepted the dmabuf-only offer (EGLCUDA GPU import) downgrading THIS \
offer to the CPU path for the rest of the process; the pipeline rebuild \
will renegotiate without dmabuf",
self.node_id
))
} else {
Err(anyhow!(
)),
TimeoutOffer::Hdr => {
// The HDR (10-bit PQ dmabuf) offer was never accepted — the monitor left HDR
// mode between the probe and the negotiation, the compositor pre-dates the
// GNOME 50 HDR formats, or its allocator can't do LINEAR for XR30/XB30.
// Latch the SDR downgrade for THIS source (`HdrSource`, not process-wide — one
// shared flag let either Linux HDR source disable the other) so the next session
// (Moonlight auto-reconnects) negotiates SDR instead of re-running this timeout.
if convicted {
super::note_hdr_capture_failed(self.hdr_source);
}
Err(anyhow!(
"no PipeWire frame within {within}s (node {}): the compositor never \
accepted the HDR (10-bit PQ/BT.2020 dmabuf) offer is the mirrored \
monitor in HDR mode on GNOME 50+?{}",
self.node_id,
if convicted {
" Downgrading this host to SDR capture; reconnect to stream SDR"
} else {
sentence
}
))
}
TimeoutOffer::RawDmabuf => {
// The dmabuf-only raw-passthrough offer was never accepted. Latch the
// downgrade so the encode loop's pipeline rebuild retries on the CPU offer
// instead of failing this same negotiation forever. The latch is SCOPED to the
// raw-passthrough decision: it used to be `note_vaapi_dmabuf_failed`, which fed
// `pf_zerocopy::enabled()` and therefore dropped every later session on this
// host — NVENC's EGL→CUDA path included — to CPU capture. Since this offer is
// also the PyroWave one (any vendor), a single PyroWave negotiation timeout was
// enough to do that.
if convicted {
pf_zerocopy::note_raw_dmabuf_negotiation_failed();
}
Err(anyhow!(
"no PipeWire frame within {within}s (node {}): the compositor never \
accepted the dmabuf-only offer (raw-dmabuf passthrough){}",
self.node_id,
if convicted {
" — downgrading THIS path to CPU capture for the rest of the \
process; the pipeline rebuild will renegotiate without dmabuf"
} else {
sentence
}
))
}
TimeoutOffer::GpuDmabuf => {
// The EGL→CUDA dmabuf-only offer was never accepted — the twin of the raw-
// passthrough arm above (the offer the thread ACTUALLY made, per the signal
// it set — see `CaptureSignals::gpu_dmabuf_offer`). One FULL-LENGTH timeout
// is conclusive: a compositor that allocates none of the importer's
// modifiers refuses them identically on every retry, so latch the offer off
// and let the pipeline rebuild renegotiate the CPU path instead of
// re-running this same 10 s timeout on every reconnect. A forced
// PUNKTFUNK_ZEROCOPY=1 keeps erroring loudly instead (same rule as the raw
// arm).
if convicted {
pf_zerocopy::note_gpu_dmabuf_negotiation_failed();
}
Err(anyhow!(
"no PipeWire frame within {within}s (node {}): the compositor never \
accepted the dmabuf-only offer (EGLCUDA GPU import){}",
self.node_id,
if convicted {
" — downgrading THIS offer to the CPU path for the rest of the \
process; the pipeline rebuild will renegotiate without dmabuf"
} else {
sentence
}
))
}
TimeoutOffer::NoFormat => Err(anyhow!(
"no PipeWire frame within {within}s (node {}): format negotiation never \
completed the compositor offered no format this consumer accepts \
(pixel-format/modifier mismatch) or the node never emitted a Format param",
self.node_id
))
)),
}
}
RecvTimeoutError::Disconnected => Err(anyhow!(
@@ -874,3 +998,89 @@ mod pipewire;
// unit-test without a compositor, which is the point.
mod pw_cursor;
mod pw_pods;
#[cfg(test)]
mod first_frame_timeout_tests {
use super::{classify_first_frame_timeout, timeout_convicts, TimeoutOffer, TimeoutVerdict};
#[test]
fn a_provisional_expiry_convicts_no_offer_whatever_was_on_the_table() {
// The bug this pins down: the retry loop's truncated 2.5 s first attempt latched all
// three sticky process-wide downgrades as if the compositor had refused the offers — a
// gamescope HDR cold start then streamed SDR (and CPU-copied) for the process lifetime.
for offer in [
TimeoutOffer::NoBuffers,
TimeoutOffer::Hdr,
TimeoutOffer::RawDmabuf,
TimeoutOffer::GpuDmabuf,
TimeoutOffer::NoFormat,
] {
assert!(
!timeout_convicts(offer, TimeoutVerdict::Provisional),
"provisional expiry must not latch {offer:?}"
);
}
}
#[test]
fn a_conclusive_expiry_convicts_exactly_the_offer_bearing_diagnoses() {
assert!(timeout_convicts(
TimeoutOffer::Hdr,
TimeoutVerdict::Conclusive
));
assert!(timeout_convicts(
TimeoutOffer::RawDmabuf,
TimeoutVerdict::Conclusive
));
assert!(timeout_convicts(
TimeoutOffer::GpuDmabuf,
TimeoutVerdict::Conclusive
));
// A negotiated-but-idle stream and a plain format mismatch implicate no offer — nothing
// to latch even on a full-length wait.
assert!(!timeout_convicts(
TimeoutOffer::NoBuffers,
TimeoutVerdict::Conclusive
));
assert!(!timeout_convicts(
TimeoutOffer::NoFormat,
TimeoutVerdict::Conclusive
));
}
#[test]
fn classification_mirrors_the_negotiation_state_precedence() {
// A negotiated format clears every offer, whatever else was on the table.
assert_eq!(
classify_first_frame_timeout(true, true, true, true, false),
TimeoutOffer::NoBuffers
);
// The HDR offer outranks the dmabuf arms (it is the offer that failed to negotiate).
assert_eq!(
classify_first_frame_timeout(false, true, true, true, false),
TimeoutOffer::Hdr
);
assert_eq!(
classify_first_frame_timeout(false, false, true, true, false),
TimeoutOffer::RawDmabuf
);
assert_eq!(
classify_first_frame_timeout(false, false, false, true, false),
TimeoutOffer::GpuDmabuf
);
assert_eq!(
classify_first_frame_timeout(false, false, false, false, false),
TimeoutOffer::NoFormat
);
}
#[test]
fn a_forced_zerocopy_keeps_both_dmabuf_arms_erroring_loudly_instead_of_implicated() {
// PUNKTFUNK_ZEROCOPY=1 is the operator insisting on the path — the timeout falls through
// to the generic diagnosis (and so never latches), exactly as the old else-if chain did.
assert_eq!(
classify_first_frame_timeout(false, false, true, true, true),
TimeoutOffer::NoFormat
);
}
}
+14 -1
View File
@@ -1506,7 +1506,20 @@ pub fn pipewire_thread(
{
return;
}
if ud.info.parse(param).is_ok() {
// Parse ONCE — `parse` takes `&mut self` — and report a failure instead of swallowing it.
// On `Err`, `negotiated` stays false and `format`/`modifier`/`frame_size` keep their
// previous values, so the capture dies on the generic "the compositor offered no format
// this consumer accepts" timeout — sending the operator hunting a format mismatch when
// the real fault was a malformed Format pod we DID accept.
let parsed = ud.info.parse(param);
if let Err(e) = &parsed {
tracing::error!(
error = %e,
"pipewire: failed to parse the negotiated Format pod — capture will time out \
with no usable format"
);
}
if parsed.is_ok() {
ud.signals.negotiated.store(true, Ordering::Relaxed);
// A (re)negotiation replaces the buffer pool: every cached per-buffer import
// (stored fds in the worker, the Vulkan bridge's per-fd sources) keys on
+11 -1
View File
@@ -197,6 +197,15 @@ pub(super) fn update_cursor_meta(cursor: &mut CursorState, spa_buf: *mut spa::sy
if bw == 0 || bh == 0 || bw > 1024 || bh > 1024 {
return;
}
// SPA's second "no image data" signal, distinct from the `bitmap_offset == 0` position-only
// case above: `spa_meta_bitmap.offset` is the offset of the PIXELS within the bitmap struct,
// and 0 means there are none. Without this, `pix_off == 0` made the pixel extent start at the
// `spa_meta_bitmap` header itself, so a producer signalling an invisible pointer got its own
// header words (format/size/stride/offset) decoded and cached as the cursor bitmap. In bounds,
// so not unsound — just garbage pixels blitted into every later frame.
if pix_off == 0 {
return;
}
let row = bw as usize * 4;
let stride = if stride < row { row } else { stride };
let Some(extent) = bitmap_extent(bmp_off, pix_off, stride, row, bh as usize, region_size)
@@ -327,7 +336,8 @@ pub(super) fn composite_cursor_rgb10(
}
/// Alpha-blend the cached cursor bitmap into the tightly-packed CPU frame at its latched
/// position. Cheap: a straight-alpha blit over at most ~256×256 pixels, clipped to the frame —
/// position. Cheap: a straight-alpha blit over at most 1024×1024 pixels (the accepted cap; real
/// cursors are ≤96 px), clipped to the frame —
/// the whole point of cursor-as-metadata (no forced full-frame composite on the producer).
pub(super) fn composite_cursor(
tight: &mut [u8],
+2 -1
View File
@@ -377,7 +377,8 @@ pub(super) fn build_dmabuf_buffers() -> Result<Vec<u8>> {
/// Request the compositor attach `SPA_META_Cursor` to each buffer, so the pointer travels as
/// metadata (position + an occasional bitmap) instead of being burned into the frame. Paired
/// with the portal's `CursorMode::Metadata`; producers that don't support it simply don't
/// attach it (harmless). Size is a range up to a 256×256 bitmap — bigger than any real cursor.
/// attach it (harmless). Size is a range up to a 1024×1024 bitmap — see the note on `max` below for
/// why this is not the "bigger than any real cursor" 256² it used to be.
pub(super) fn build_cursor_meta_param() -> Result<Vec<u8>> {
fn meta_size(w: u32, h: u32) -> i32 {
(std::mem::size_of::<spa::sys::spa_meta_cursor>()
+38 -35
View File
@@ -55,16 +55,6 @@ use x11rb::rust_connection::{DefaultStream, RustConnection};
use crate::GamescopeCursorTargets;
/// Serializes the `XAUTHORITY` env swap of the LEGACY connect fallback (the var is process-global).
///
/// The fallback is a last resort now — see [`connect_conn`]. It serialises this source against
/// itself and nothing else: `getenv` needs no lock to be racy, so every OTHER thread's read (libspa
/// plugin load, EGL/CUDA init — concurrent by construction, since `attach_gamescope_cursor` runs
/// while the PipeWire thread is starting) could still observe the swapped value or a torn
/// environ. That is why the primary path parses the cookie itself and never touches the
/// environment.
static XAUTH_LOCK: Mutex<()> = Mutex::new(());
/// The `MIT-MAGIC-COOKIE-1` auth-protocol name, as it appears in an `.Xauthority` entry.
const MIT_MAGIC_COOKIE_1: &[u8] = b"MIT-MAGIC-COOKIE-1";
@@ -267,17 +257,18 @@ fn connect(dpy: &str, xauthority: Option<&str>) -> Result<Connected, String> {
/// environment.
///
/// `RustConnection::connect` reads `XAUTHORITY` from the env, so the original implementation
/// `set_var`'d it around each connect under [`XAUTH_LOCK`]. That is unsound from a live
/// multithreaded host: the lock serialises this source against itself, but `getenv` takes no lock,
/// so any concurrent reader (libspa's plugin load, EGL/CUDA init — running at exactly this moment,
/// since the PipeWire thread is starting up) could read the swapped value or race the environ
/// rewrite outright. The project already has a process-wide env-lock discipline elsewhere, but
/// sharing it would be the wrong layer AND would still not fix `getenv`.
/// `set_var`'d it around each connect under a mutex. That is unsound from a live multithreaded
/// host: the lock serialised this source against itself, but `getenv` takes no lock, so any
/// concurrent reader (libspa's plugin load, EGL/CUDA init — running at exactly this moment, since
/// the PipeWire thread is starting up) could read the swapped value or race the environ rewrite
/// outright. The project already has a process-wide env-lock discipline elsewhere, but sharing it
/// would be the wrong layer AND would still not fix `getenv`.
///
/// So: parse the MIT-MAGIC-COOKIE-1 entry out of the file ourselves and hand it to
/// `connect_to_stream_with_auth_info`, which is what `RustConnection::connect` does internally with
/// the cookie IT found. The env swap survives only as a fallback for a file we cannot parse (an
/// unexpected layout, or an auth family whose entry we decline to guess at).
/// the cookie IT found. Where that finds nothing usable we connect with an explicitly empty token
/// ([`connect_unauthenticated`]) rather than swapping the environment — this process no longer
/// writes `environ` at all.
fn connect_conn(dpy: &str, xauthority: Option<&str>) -> Result<(RustConnection, usize), String> {
let Some(path) = xauthority else {
// No per-display cookie file to inject: the ambient environment is already what this
@@ -289,16 +280,16 @@ fn connect_conn(dpy: &str, xauthority: Option<&str>) -> Result<(RustConnection,
Ok(v) => return Ok(v),
Err(e) => tracing::debug!(
dpy = %dpy, xauthority = %path, error = %e,
"gamescope cursor: cookie connect failed — falling back to the XAUTHORITY env swap"
"gamescope cursor: cookie connect failed — retrying unauthenticated"
),
},
None => tracing::debug!(
dpy = %dpy, xauthority = %path,
"gamescope cursor: no MIT-MAGIC-COOKIE-1 entry for this display — falling back to the \
XAUTHORITY env swap"
"gamescope cursor: no MIT-MAGIC-COOKIE-1 entry for this display — connecting \
unauthenticated"
),
}
connect_via_env_swap(dpy, path)
connect_unauthenticated(dpy)
}
/// Connect to `dpy` and complete the setup handshake with an explicit cookie — the same two steps
@@ -331,19 +322,31 @@ fn connect_with_cookie(
.map_err(|e| format!("setup: {e}"))
}
/// LEGACY fallback (see [`connect_conn`]): swap `XAUTHORITY`, connect, restore. Serialised against
/// this source's own concurrent connects, but NOT against other threads' `getenv` — which is why it
/// is a fallback and not the path taken.
fn connect_via_env_swap(dpy: &str, xauthority: &str) -> Result<(RustConnection, usize), String> {
let _g = XAUTH_LOCK.lock().unwrap_or_else(|e| e.into_inner());
let prev = std::env::var_os("XAUTHORITY");
std::env::set_var("XAUTHORITY", xauthority);
let out = RustConnection::connect(Some(dpy));
match prev {
Some(p) => std::env::set_var("XAUTHORITY", p),
None => std::env::remove_var("XAUTHORITY"),
}
out.map_err(|e| format!("connect: {e}"))
/// Last-resort fallback (see [`connect_conn`]): connect with an EXPLICITLY EMPTY auth token.
///
/// This replaces a `set_var("XAUTHORITY", …)` / connect / restore dance, which was unsound and is
/// not fixable in place. `setenv`/`unsetenv` rewrite the process-global `environ`; glibc
/// *reallocates* that array when a variable is added, and the host is emphatically multithreaded
/// at this moment — `attach_gamescope_cursor` runs while the PipeWire thread is inside `pw_init`'s
/// `dlopen` and a dozen bare `getenv()` calls, with EGL/CUDA init alongside. A mutex here
/// serialised this source against itself and against nothing else, because `getenv` takes no lock.
/// The damaging branch is the one where `XAUTHORITY` is ABSENT and therefore gets *added* — which
/// `scripts/punktfunk-host.service` makes the normal configuration, since the unit deliberately
/// does not import the login shell's environment. And `rediscover` re-runs this every 2 s for the
/// whole session, because a display whose connect fails is never recorded and so is never skipped.
///
/// Connecting with an empty token is what the swap actually achieved. We only reach here when our
/// own lookup found no usable `MIT-MAGIC-COOKIE-1` entry, and x11rb's internal lookup reads the
/// same file with a STRICTER matcher (it also matches family/address, which we deliberately do
/// not) — so where we find nothing, it finds nothing too, and connects unauthenticated. That is
/// precisely why the swap "worked" against a nested Xwayland started without `-auth`.
///
/// The one case this gives up is an `.Xauthority` whose entry uses an auth family we decline to
/// guess at but x11rb would have handled. A gamescope Xwayland writes a single-entry
/// MIT-MAGIC-COOKIE-1 file, so that case is not reachable here — and a cursor overlay that
/// declines to attach is the correct outcome anyway, against a torn `environ` in a live session.
fn connect_unauthenticated(dpy: &str) -> Result<(RustConnection, usize), String> {
connect_with_cookie(dpy, Vec::new(), Vec::new())
}
/// The `MIT-MAGIC-COOKIE-1` `(name, data)` for `dpy` from the `.Xauthority`-format file at `path`.
+5 -6
View File
@@ -9,9 +9,6 @@
//! `crate::dxgi::*` path keeps resolving. DXGI Desktop Duplication has been removed; this
//! module contains no capturer.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
pub use pf_frame::dxgi::{make_device, pack_luid, D3d11Frame, PyroFrameShare, WinCaptureTarget};
// The P010 colour self-test (sweep Phase 5.5) — the `hdr-p010-selftest` subcommand, its f64
@@ -554,9 +551,11 @@ impl HdrP010Converter {
let mut ps_uv = None;
device.CreatePixelShader(&uvb, None, Some(&mut ps_uv))?;
let sd = D3D11_SAMPLER_DESC {
// POINT: the Y pass samples a single texel centre exactly, and the UV pass does its OWN
// 2x2 box average via 4 explicit taps at texel centres (offset half a texel). Point
// sampling keeps each tap exact; the averaging is in the shader, not the sampler.
// POINT: the Y pass samples a single texel centre exactly, and the UV pass takes its OWN
// two explicit taps on the 2x2 block's LEFT column (left-cositing) and averages them.
// Point sampling keeps each tap exact; the averaging is in the shader, not the sampler.
// (It was a 4-tap CENTER-sited 2x2 box until that was found to shift chroma by half a
// luma pixel — see `HDR_P010_UV_PS`.)
Filter: D3D11_FILTER_MIN_MAG_MIP_POINT,
AddressU: D3D11_TEXTURE_ADDRESS_CLAMP,
AddressV: D3D11_TEXTURE_ADDRESS_CLAMP,
+17 -12
View File
@@ -16,9 +16,6 @@
//! [`pf_driver_proto`] (which OWNS the contract, with `const` size asserts) — both sides `use` it, so
//! drift is a compile error rather than a "must match" comment.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::dxgi::{
make_device, BgraToYuvPlanes, D3d11Frame, HdrP010Converter, HdrRgb10Converter, PyroFrameShare,
VideoConverter, WinCaptureTarget,
@@ -337,6 +334,7 @@ use channel::ChannelBroker;
use descriptor::{DescriptorPoller, DisplayDescriptor};
use stall::{StallEvidence, StallWatch};
/// Creates + owns the shared ring; yields the driver's frames as [`FramePayload::D3d11`].
pub struct IddPushCapturer {
device: ID3D11Device,
context: ID3D11DeviceContext,
@@ -652,14 +650,18 @@ impl IddPushCapturer {
}
/// The output texture format + the [`PixelFormat`] NVENC encodes, driven by the DISPLAY's HDR
/// state (like the WGC path) plus the session's 4:4:4 negotiation: HDR → `P010` (BT.2020 PQ
/// state plus the session's 4:4:4 negotiation: HDR → `P010` (BT.2020 PQ
/// 10-bit limited) → NVENC Main10, and the client auto-detects PQ from the HEVC VUI; SDR →
/// `Nv12` (BT.709 8-bit limited), or full-chroma `Bgra` passthrough on a 4:4:4 session (NVENC
/// CSCs RGB→YUV444 itself, following the BT.709 VUI — the one path that deliberately pays the
/// SM-side CSC, because the video processor can only produce subsampled output). We do NOT
/// gate HDR on the client's advertised `VIDEO_CAP_10BIT` — clients under-report it (e.g. the
/// Mac advertises 10-bit only when its OWN display is HDR), yet all decode Main10 +
/// auto-switch, exactly as on the WGC path. HDR and 4:4:4 now COMPOSE: an HDR display that
/// SM-side CSC, because the video processor can only produce subsampled output). The
/// composition depth DOES follow the session's negotiated `client_10bit` — pinned at open
/// (`open.rs`, the `!client_10bit` force-off and the 10-bit enable) and re-pinned every sample
/// by [`Self::poll_display_hdr`], because a PQ stream sent to a client that advertised SDR-only
/// lands on an SDR desktop and blows out. (The older note here claimed the opposite — that the
/// advertised `VIDEO_CAP_10BIT` was ignored because clients under-report it. That reasoning
/// survives only in the CODEC choice: an HDR-negotiated H.26x session still follows a host
/// "Use HDR" flip in either direction.) HDR and 4:4:4 now COMPOSE: an HDR display that
/// negotiated full chroma emits packed 10-bit BT.2020 PQ RGB (`Rgb10a2`) for NVENC to CSC to
/// YUV 4:4:4 — HEVC Main 4:4:4 10. (Before, HDR won and the stream silently downgraded to
/// 4:2:0 *after* the Welcome had already promised 4:4:4.)
@@ -969,7 +971,7 @@ impl IddPushCapturer {
},
Usage: D3D11_USAGE_DEFAULT,
// RENDER_TARGET: the VIDEO processor (NV12) and the P010 shader passes both write here, and
// NVENC registers it as encode input — matching the WGC YUV ring. (PyroWave uses its own
// NVENC registers it as encode input. (PyroWave uses its own
// shareable two-plane `pyro_ring` instead, so this NVENC/AMF/QSV ring stays unshared.)
BindFlags: D3D11_BIND_RENDER_TARGET.0 as u32,
CPUAccessFlags: 0,
@@ -1970,9 +1972,12 @@ impl Capturer for IddPushCapturer {
fn pipeline_depth(&self) -> usize {
// 2 = one frame deferred: submit N+1 (capture + convert/copy into a fresh out-ring texture) while
// NVENC encodes N on the ASIC. We hand a rotating `OUT_RING` of output textures, so this is safe.
// `PUNKTFUNK_IDD_DEPTH` overrides (1 disables pipelining; clamp to ≤ OUT_RING so a frame in flight
// always has its own texture).
pf_host_config::config().idd_depth.clamp(1, OUT_RING)
// `PUNKTFUNK_IDD_DEPTH` overrides (1 disables pipelining). The ceiling is `OUT_RING - 1`, NOT
// `OUT_RING`: `d` frames in flight need `d + 1` textures, because the rotation has to hand out a
// slot that is not one of the `d` still being encoded. Clamping to `OUT_RING` admitted depth 3 on
// a 3-slot ring, where `repeat_last`'s rotation lands back on the slot NVENC is reading and the
// convert overwrites it in place — torn frames, silently, with no error anywhere.
pf_host_config::config().idd_depth.clamp(1, OUT_RING - 1)
}
fn capture_target_id(&self) -> Option<u32> {
@@ -2,9 +2,6 @@
//! capturer): duplicates the unnamed shared header / ring / event handles into the driver's WUDFHost
//! and delivers them as bare handle values over the SYSTEM-only control device.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::*;
/// The sealed channel's handle-duplication broker (`design/idd-push-security.md`): the frame objects
@@ -160,7 +157,18 @@ impl ChannelBroker {
event: HANDLE,
slots: &[HostSlot],
) -> Result<()> {
debug_assert!(slots.len() <= control::RING_LEN_USIZE);
// An ERROR, not a `debug_assert`: in a release build the assert is compiled out and the
// over-long slice instead panics on `req.texture_handles[k]` in the middle of
// `duplicate_and_deliver` — after handles have already been planted in WUDFHost. That panic
// unwinds straight past the reap below, leaking every duplicate made so far into the driver
// process. Refuse before the first duplication, while there is nothing to reap.
if slots.len() > control::RING_LEN_USIZE {
anyhow::bail!(
"frame channel: {} ring slots exceeds the wire limit of {}",
slots.len(),
control::RING_LEN_USIZE
);
}
let mut req = control::SetFrameChannelRequest {
target_id,
generation,
@@ -5,9 +5,6 @@
//! [`pf_frame::CursorOverlay`] the Linux portal path produces — everything downstream (the
//! cursor forwarder, the wire, the client renderer) is shared.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use super::*;
use pf_driver_proto::cursor::{
CursorShm, CURSOR_MAGIC, CURSOR_SHAPE_BYTES, CURSOR_SHAPE_MAX, CURSOR_SHAPE_OFFSET,
@@ -42,7 +39,9 @@ impl CursorShared {
/// the section itself (owned by `self`); the caller duplicates it into the WUDFHost.
pub(super) fn create(target_id: u32) -> Result<CursorShared> {
// SAFETY: plain FFI. Unnamed pagefile-backed section, host-lifetime owned; the view is
// mapped once and unmapped never (the capturer's life = the session's life).
// mapped once here and unmapped exactly once by `MappedSection::drop` (which unmaps before
// closing the mapping handle). No borrow into the view outlives the `MappedSection`: every
// access goes through `&self` accessors on the owner.
let section = unsafe {
let map = CreateFileMappingW(
INVALID_HANDLE_VALUE,
@@ -10,9 +10,6 @@
//! alpha-blended quad (the GDI poller's full-fidelity shape at its polled position), entirely
//! GPU-side on the capture device, before the normal conversion runs from the scratch.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::*;
use windows::core::s;
use windows::Win32::Graphics::Direct3D::D3D_PRIMITIVE_TOPOLOGY_TRIANGLELIST;
@@ -20,9 +20,6 @@
//! `winsta0\default` (the service supervisor retargets the token — `windows/service.rs`
//! `spawn_host`), so the poller thread sees the session's cursor directly; no helper process.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::*;
use windows::Win32::Graphics::Gdi::{
DeleteObject, GetDC, GetDIBits, GetObjectW, ReleaseDC, BITMAP, BITMAPINFO, BITMAPINFOHEADER,
@@ -55,8 +52,10 @@ struct Shape {
serial: u64,
}
/// Off-thread GDI cursor poller. Samples `GetCursorInfo` at ~60 Hz, rasterises the `HCURSOR` only
/// when its handle value changes, and publishes a ready [`pf_frame::CursorOverlay`] snapshot; the
/// Off-thread GDI cursor poller. Samples `GetCursorInfo` every [`Self::INTERVAL`] (4 ms, ~250 Hz —
/// see that constant for why 16 ms was the bug), rasterises the `HCURSOR` when its handle value
/// changes and when [`Self::EXTENT_PROBE`] catches a resize under a STABLE handle, and publishes a
/// ready [`pf_frame::CursorOverlay`] snapshot; the
/// capture thread's per-tick cost is one uncontended mutex read + an `Arc` clone
/// (same split as [`DescriptorPoller`], and for the same reason: user32/gdi32 calls have no place
/// on the capture/encode thread).
@@ -186,7 +185,6 @@ fn run(
// against, and this poller outlives all of them. `None` keeps the last good value — a
// transient CCD failure must not park the pointer at a `(0, 0, 0, 0)` rect, which would
// report every position invisible.
//
let fresh = pf_win_display::win_display::source_desktop_rect(target_id);
if let Some(fresh) = fresh {
if fresh != rect {
@@ -302,7 +300,14 @@ fn run(
serial: s.serial,
hot_x: s.hot_x,
hot_y: s.hot_y,
visible: showing && in_rect,
// `handle != 0` is part of "visible", not just of "worth rasterising": `SetCursor(NULL)`
// — how a game or a video player hides the pointer for its own window — leaves
// `CURSOR_SHOWING` set with a NULL `hCursor`. Judging on the flags alone published
// `visible: true` carrying the last shape we rasterised, so the composite path blended a
// ghost arrow into a game that had hidden its cursor, and the forward path told the
// client to draw one too. Every rasterise gate below already tests this; the published
// verdict has to agree with them.
visible: showing && in_rect && handle != 0,
}
});
*slot.lock().unwrap_or_else(|p| p.into_inner()) = overlay;
@@ -1,12 +1,8 @@
//! Off-thread display-descriptor polling (plan §W4, carved out of the IDD-push capturer): the
//! live HDR state + active resolution of the virtual target, sampled off the capture loop via CCD.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::*;
/// Creates + owns the shared ring; yields the driver's frames as [`FramePayload::D3d11`].
/// The display descriptor the capture loop follows: live HDR state + active resolution of the
/// virtual target.
#[derive(Clone, Copy, PartialEq, Eq)]
@@ -33,9 +33,6 @@
//! The session's `FlushTimer` is 1 s, so a bracket from the trailing second of a gap can land
//! AFTER that stall's report line — the next report (and the metronomic tally) still carries it.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use std::collections::VecDeque;
use std::sync::{Arc, Mutex, OnceLock, Weak};
use std::time::{Duration, Instant};
@@ -142,7 +139,12 @@ unsafe extern "system" fn on_event(record: *mut EVENT_RECORD) {
(*record).EventHeader.ProcessId,
)
};
let mut ring = RING.lock().unwrap();
// Poison-tolerant, and that is load-bearing rather than tidy: this is an `extern "system"`
// callback invoked from an OS thread, so a panic here unwinds across an FFI boundary and
// ABORTS the host process. `unwrap()` made a single poisoned lock turn every subsequent event
// delivery into a hard abort — a diagnostic taking down capture. Nothing else under this lock
// can panic, so recovering the guard also makes the poison unreachable in the first place.
let mut ring = RING.lock().unwrap_or_else(|e| e.into_inner());
if ring.len() == RING_CAP {
ring.pop_front();
}
@@ -152,8 +152,11 @@ impl IddPushCapturer {
}
/// Open the IDD-push capturer. On success the caller's `keepalive` is attached (the capturer owns the
/// virtual display); on FAILURE the keepalive is handed BACK so the caller can fall back to DDA
/// instead of tearing the display down (audit §5.1 — no more 20 s black bail). "Failure" includes the
/// virtual display); on FAILURE the keepalive is handed BACK so the caller decides the display's fate
/// itself — retire it, or reuse the monitor for a retry — instead of this function tearing it down
/// (audit §5.1 — no more 20 s black bail). There is no second capture path to fall back TO: DDA was
/// removed (see `lib.rs`), and `punktfunk-host`'s caller drops the returned keepalive under
/// `.context("IDD-push capture open (no fallback)")`. "Failure" includes the
/// driver not attaching to the ring within a few seconds (e.g. a hybrid-GPU render mismatch).
#[allow(clippy::too_many_arguments)]
pub fn open(
@@ -666,7 +669,7 @@ impl IddPushCapturer {
// wait for the first compose) until the capturer drops with the session.
_display_wake: pf_frame::session_tuning::DisplayWakeRequest::new(),
// Placeholder; `open()` attaches the real keepalive on success, so a FAILED open can hand
// it back to the caller for the DDA fallback (audit §5.1).
// it back to the caller to retire or reuse the display (audit §5.1).
_keepalive: Box::new(()),
};
// The HDR SDR-white reference for the composited cursor, queried ONCE here rather than
@@ -675,15 +678,15 @@ impl IddPushCapturer {
me.refresh_sdr_white_scale();
// Bounded wait for the driver to ATTACH to the ring AND publish a first frame. An attach
// failure (DRV_STATUS_TEX_FAIL) or an attach-but-no-frames (a game left the display in a
// format/size the ring can't match) becomes an open failure the caller falls back from (→ DDA),
// instead of next_frame's 20 s black-then-bail.
// format/size the ring can't match) becomes an open failure the caller handles by retiring the
// display, instead of next_frame's 20 s black-then-bail.
me.wait_for_attach()?;
Ok(me)
}
}
/// Block (bounded) until the driver has ATTACHED to the host ring (`DRV_STATUS_OPENED`) **and published
/// a first frame**, else fail so the caller can fall back to DDA (audit §5.1 +
/// a first frame**, else fail so the caller can retire the display and rebuild (audit §5.1 +
/// `design/windows-host-rewrite.md` §2.5 — the GB1 game-capture fix).
///
/// Requiring the first frame — not just the attach — catches the *reconnect-into-a-broken-state* case:
@@ -25,9 +25,6 @@
//! ([`acquire`]), refcounted across parallel capturers; probes sample at 20 Hz or slower and cost
//! microseconds each, so the engine is invisible next to a streaming session.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use std::collections::VecDeque;
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{Arc, Mutex, Weak};
@@ -53,7 +50,9 @@ use super::stall::ProbeWindow;
/// One probe's sample ring: `(completed_at, span, value_us)` — `value` is the measurement (a call
/// latency or a frozen-span/overshoot), `span` the wall interval it describes ending at
/// `completed_at`. Capped; ~20 Hz per probe → several minutes of coverage.
/// `completed_at`. Capped at 512 samples: at the fastest producer's ~20 Hz that is ~26 s of
/// coverage, ~51 s for the 100 ms loops — comfortably longer than the seconds-old windows a stall
/// report asks for, but NOT the "several minutes" this used to claim.
struct Ring {
samples: Mutex<VecDeque<(Instant, Duration, u64)>>,
}
@@ -1,9 +1,6 @@
//! Capture-stall detection (plan §W4, carved out of the IDD-push capturer): flags multi-hundred-ms
//! holes in DWM frame delivery that open while the desktop was actively composing.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::*;
/// A detected capture stall: a multi-hundred-ms hole in DWM's frame delivery that opened while the
@@ -317,7 +314,8 @@ impl StallWatch {
/// Frames of pre-gap history that must be tight for flow to count as active. Stalls are thus
/// naturally spaced ≥ RECENT frame times apart — no extra log rate limit needed.
const RECENT: usize = 8;
/// The RECENT pre-gap frames must all fit in this span (8 frames in 400 ms ≈ ≥ 20 fps flow —
/// The RECENT pre-gap frames must all fit in this span (8 frames spanning 400 ms is 7 intervals,
/// so the real bar is ≈ ≥ 17.5 fps flow —
/// loose enough for a 30 fps-capped game, tight enough to reject idle-desktop damage).
const ACTIVE_SPAN: Duration = Duration::from_millis(400);
/// The smallest hole that counts as a stall (~9 missed frames at 60 Hz) — well below the
@@ -535,14 +533,47 @@ impl StallWatch {
suspects)"
);
} else {
// The two REALTIME GPU-priority opt-ins, as configured in THIS process's
// environment (machine env; the WUDFHost driver process resolves the PFVD pair
// the same way, so this read mirrors what the driver decided — modulo a machine
// env edited after either process started, which a restart heals). The RX 9070
// XT field A/B (2026-08-12) convicted EXACTLY this warning's signature twice
// over: the driver's swap-chain REALTIME raise beat at ~1.8 s, the host
// auto-gate's REALTIME upgrade at ~3.6 s — so a log carrying this warning must
// say whether either lever is engaged before anyone chases display hardware.
let rt_gpu_driver = if std::env::var_os("PFVD_NO_RT_GPU").is_some() {
"off (PFVD_NO_RT_GPU)"
} else {
match std::env::var_os("PFVD_RT_GPU") {
None => "off (default)",
Some(v) if v.eq_ignore_ascii_case("thread") => "gpu-thread (+7)",
Some(_) => "REALTIME (PFVD_RT_GPU)",
}
};
let rt_gpu_host = match std::env::var("PUNKTFUNK_GPU_PRIORITY_CLASS")
.ok()
.as_deref()
{
Some("off") => "off",
Some("normal") => "normal",
Some("realtime") => "REALTIME (pinned)",
Some("auto") => "auto (gated REALTIME upgrade)",
_ => "high (default)",
};
tracing::warn!(
period_s = format!("{:.2}", period.as_secs_f64()),
os_correlated = correlated,
connected_inactive = %suspects,
rt_gpu_driver,
rt_gpu_host,
verdicts = %verdict_tally,
classes = %class_tally,
"capture stalls are METRONOMIC with NO coinciding OS display event — \
the disturbance is BELOW Windows: the GPU driver servicing a \
the disturbance is BELOW Windows. FIRST: if rt_gpu_driver or \
rt_gpu_host shows a REALTIME opt-in, clear it (unset PFVD_RT_GPU / \
set PUNKTFUNK_GPU_PRIORITY_CLASS=high) a punktfunk process holding \
REALTIME GPU priority is the field-proven amplifier of exactly this \
signature on AMD. Otherwise: the GPU driver servicing a \
connected-but-asleep sink (standby HPD/DDC/link probing), \
display-poller software (the SteelSeries-GG/SignalRGB class \
correlate 'slow display-descriptor poll' lines), or the DWM present \
-1
View File
@@ -18,7 +18,6 @@
// proof of why it is sound. This crate held ~91 unsafe items with NO enforcement while every
// other subsystem crate denied it — the decoders' `unsafe impl Send`s had a one-line aside
// instead of an argument precisely because nothing required one.
#![deny(clippy::undocumented_unsafe_blocks)]
#[cfg(any(target_os = "linux", windows))]
mod au_dump;
+2 -2
View File
@@ -17,8 +17,8 @@
//! (`PostMessage` is the documented thread-safe way to poke a message loop). Per-window state hangs
//! off `GWLP_USERDATA`, so multiple concurrent sessions each get their own window + state.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
// Every `unsafe` block in this file carries a `// SAFETY:` proof; the deny enforcing it sits at
// the crate root (lib.rs), covering every backend.
use std::cell::RefCell;
use std::sync::{Arc, Mutex};
+4
View File
@@ -10,6 +10,10 @@
//! [`spawn_decline_loop`] — so its control loop compiles unchanged on every host platform; the
//! platform split lives entirely behind [`start`].
// Unsafe-proof program: every `unsafe` block in any backend carries a `// SAFETY:` proof,
// enforced workspace-wide by `[workspace.lints]` — a new backend under `host/` is covered on
// creation.
use std::sync::atomic::AtomicBool;
use std::sync::Arc;
-1
View File
@@ -11,7 +11,6 @@
//! capture hint, start banner.
// Unsafe-proof program: every `unsafe {}` in the Skia/Vulkan overlay carries a `// SAFETY:` proof.
#![deny(clippy::undocumented_unsafe_blocks)]
#[cfg(any(target_os = "linux", windows))]
mod anim;
+7 -1
View File
@@ -1712,7 +1712,13 @@ mod tests {
let mut legacy = [0u8; 40];
legacy[..control::ADD_REQUEST_LEGACY_SIZE]
.copy_from_slice(&bytes[..control::ADD_REQUEST_LEGACY_SIZE]);
let old = *bytemuck::from_bytes::<control::AddRequest>(&legacy);
// `pod_read_unaligned`, NOT `from_bytes` — same rule as `ChannelProof::parse` above, and
// for the same reason. `legacy` is a `[u8; 40]` (align 1) but `AddRequest` opens with a
// `u64`, so it is align 8; `from_bytes` takes a REFERENCE into the buffer and panics
// unless the buffer happens to be 8-aligned. A stack `[u8; 40]` usually is, which is why
// this passed everywhere for so long — Miri caught it because Miri does not let an
// accidentally-favourable stack slot stand in for a guarantee.
let old = bytemuck::pod_read_unaligned::<control::AddRequest>(&legacy);
assert_eq!(old.preferred_monitor_id, 7);
assert_eq!(
(
-1
View File
@@ -52,7 +52,6 @@
//! ([`dxva::as_bytes`] / [`dxva::slice_bytes`]), fenced behind a sealed trait
//! that only this crate's `#[repr(C)]` PODs implement, and carrying a written
//! proof — enforced:
#![deny(clippy::undocumented_unsafe_blocks)]
pub mod config;
pub mod descriptors;
+84
View File
@@ -48,6 +48,8 @@ impl AvBuffer {
/// allocator returns on failure (so the `is_null` check every caller used to open-code happens
/// once, here).
///
// unsafe-fn-no-op-ok: contract-deferring constructor (`Vec::set_len` shape) — the body is
// safe; the ownership transfer promised here is what Drop/as_ptr later rely on.
/// # Safety
/// `p` must be null, or a live `AVBufferRef` whose ownership passes to the returned value —
/// nothing else may unref it.
@@ -117,6 +119,88 @@ impl Drop for AvFilterGraph {
}
}
/// An owned `AVFrame`, freed exactly once when it drops.
///
/// The house pattern (`AvBuffer` above): `alloc` rejects the allocator's null once, `as_ptr`
/// lends, `Drop` frees, no `Clone`. Before this type existed the crate held 8 `av_frame_alloc`
/// sites matched by 22 hand-placed `av_frame_free`s — an ownership contract upheld by nobody,
/// and broken in practice: the Windows zero-copy submit path leaked the frame AND a pooled
/// hwframe surface on three `?` exits, under a comment asserting the opposite (fixed in the
/// same change that introduced this type).
///
/// Why not ffmpeg-next's own RAII frame (`frame::Video::empty()`, already used as `VideoFrame`
/// in the Linux NVENC path): `Frame::empty()` does not null-check — on allocator failure it
/// wraps null and the next field write through it is UB — whereas every open-coded site here
/// null-checked. This type keeps that: `alloc` returns `Option`, mirroring
/// `AvFilterGraph::alloc`.
pub(crate) struct AvFrame(std::ptr::NonNull<ffi::AVFrame>);
impl AvFrame {
/// Allocate a frame, rejecting the null `av_frame_alloc` returns on OOM.
///
/// Safe: the call takes no arguments and has no precondition a caller could violate — the
/// only contract is what happens to the result, and that is exactly what this type owns.
pub(crate) fn alloc() -> Option<Self> {
// SAFETY: parameterless allocator; it returns either a fresh, uniquely-owned frame whose
// ownership passes to the value returned here, or null (rejected by NonNull::new).
std::ptr::NonNull::new(unsafe { ffi::av_frame_alloc() }).map(AvFrame)
}
/// The borrowed pointer, for the ffmpeg calls that fill or read the frame without taking
/// ownership of it. Borrowed only — the `AvFrame` stays the owner, so callers must not free
/// or move-from what this returns.
pub(crate) fn as_ptr(&self) -> *mut ffi::AVFrame {
self.0.as_ptr()
}
}
impl Drop for AvFrame {
fn drop(&mut self) {
let mut p = self.0.as_ptr();
// SAFETY: `p` is the non-null frame `alloc` took ownership of, and this type is its
// sole owner (neither `Clone` nor `Copy`; `as_ptr` only lends), so this runs exactly
// once. `av_frame_free` unrefs any buffers the frame holds (returning pooled hwframe
// surfaces to their pool) and frees the frame; it nulls only the local copy.
unsafe { ffi::av_frame_free(&mut p) };
}
}
/// An owned swscale context, freed exactly once when it drops.
///
/// Same ownership question as the frame above — `sws_getContext` at 3 sites was matched by 5
/// hand-placed `sws_freeContext`s, two of them inside hand-written `Drop` impls whose real job
/// this type absorbs.
pub(crate) struct AvSwsContext(std::ptr::NonNull<ffi::SwsContext>);
impl AvSwsContext {
/// Take ownership of a freshly-created `SwsContext`, rejecting the null `sws_getContext`
/// returns on failure (unsupported conversion or OOM).
///
// unsafe-fn-no-op-ok: contract-deferring constructor (`Vec::set_len` shape) — the body is
// safe; the ownership transfer promised here is what Drop/as_ptr later rely on.
/// # Safety
/// `p` must be null, or a live `SwsContext` whose ownership passes to the returned value —
/// nothing else may free it.
pub(crate) unsafe fn from_raw(p: *mut ffi::SwsContext) -> Option<Self> {
std::ptr::NonNull::new(p).map(AvSwsContext)
}
/// The borrowed pointer, for `sws_scale` calls. Borrowed only — the `AvSwsContext` stays
/// the owner.
pub(crate) fn as_ptr(&self) -> *mut ffi::SwsContext {
self.0.as_ptr()
}
}
impl Drop for AvSwsContext {
fn drop(&mut self) {
// SAFETY: `self.0` is the non-null context `from_raw` took ownership of, and this type
// is its sole owner (neither `Clone` nor `Copy`; `as_ptr` only lends), so this runs
// exactly once.
unsafe { ffi::sws_freeContext(self.0.as_ptr()) };
}
}
/// One `receive_packet` attempt, with the not-ready states kept distinct so a blocking drain can
/// tell "still encoding" (retry) from "stream over" (stop). The Linux NVENC/VAAPI polls collapse
/// `Again`/`Eof` to `None`; the Windows AMF/QSV path keeps them apart for its deadline-driven loop.
+66 -71
View File
@@ -12,8 +12,6 @@
//! does *not* accept — we expand it to `rgb0` (one padding byte/pixel, no colour math).
//! The encoder is opened *without* a global header so VPS/SPS/PPS are emitted in-band on
//! every IDR — the output is both a playable raw Annex-B stream and self-contained AUs.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{ChromaFormat, Codec, EncodedFrame, Encoder};
use anyhow::{anyhow, bail, Context, Result};
@@ -26,8 +24,8 @@ use std::os::raw::c_int;
use std::ptr;
use super::libav::{
apply_low_latency_rc, pixel_to_av, poll_encoder, AvBuffer, PollOutcome, SWS_CS_ITU709,
SWS_POINT,
apply_low_latency_rc, pixel_to_av, poll_encoder, AvBuffer, AvFrame, AvSwsContext, PollOutcome,
SWS_CS_ITU709, SWS_POINT,
};
use ffmpeg::ffi; // = ffmpeg_sys_next
@@ -193,6 +191,17 @@ struct OpenArgs {
}
pub struct NvencEncoder {
// FIELD ORDER IS LOAD-BEARING: the hand-written `Drop` this replaced ran before any field
// drop, freeing `sws_csc` ahead of `enc`/`frame`/`cuda` — and this path runs on every
// stall-watchdog recovery via `*self = fresh` in `reset`. Declaration order is what
// preserves that sequence now (drop order follows declaration; an offset_of assert cannot
// pin it — repr(Rust) may lay memory out in any order).
/// CPU CSC paths only: swscale context converting the captured packed source into
/// [`Self::frame`] — RGB/BGR → planar YUV444P for a 4:4:4 session (`hevc_nvenc` only emits
/// 4:4:4 from a YUV444 *input*; RGB-in is always 4:2:0), or X2RGB10/X2BGR10 → P010 (BT.2020
/// limited) for an HDR session. `None` on the plain RGB paths AND on the zero-copy paths (the
/// worker's GPU convert delivers ready CUDA frames).
sws_csc: Option<AvSwsContext>,
enc: encoder::video::Encoder,
/// Reusable 4-bpp CPU input frame (CPU path only; `None` for the zero-copy/CUDA path).
/// Mutating it in place across frames is sound only because the encoder is opened with
@@ -201,12 +210,6 @@ pub struct NvencEncoder {
frame: Option<VideoFrame>,
/// Zero-copy path: CUDA hwdevice/hwframes contexts (the encoder takes `AV_PIX_FMT_CUDA`).
cuda: Option<CudaHw>,
/// CPU CSC paths only: swscale context converting the captured packed source into
/// [`Self::frame`] — RGB/BGR → planar YUV444P for a 4:4:4 session (`hevc_nvenc` only emits
/// 4:4:4 from a YUV444 *input*; RGB-in is always 4:2:0), or X2RGB10/X2BGR10 → P010 (BT.2020
/// limited) for an HDR session. `None` on the plain RGB paths AND on the zero-copy paths (the
/// worker's GPU convert delivers ready CUDA frames). Freed in `Drop`.
sws_csc: Option<*mut ffi::SwsContext>,
/// This session opened as full-chroma 4:4:4 (FREXT) — via either input path.
want_444: bool,
src_format: PixelFormat,
@@ -228,7 +231,7 @@ pub struct NvencEncoder {
args: OpenArgs,
}
// `CudaHw` holds raw `AVBufferRef`s and `sws_csc` a raw `SwsContext`; the encoder lives on a single
// `CudaHw` holds raw `AVBufferRef`s and `sws_csc` an owned `SwsContext`; the encoder lives on a single
// thread. The CPU encoder is already `Send` via ffmpeg-next; assert it for the raw fields too.
// SAFETY: `NvencEncoder` owns an ffmpeg-next `Encoder`/`VideoFrame` (already `Send`) plus a `CudaHw`
// holding raw `AVBufferRef`s and an optional raw `SwsContext`, none of which are `Send` by default.
@@ -610,14 +613,13 @@ impl NvencEncoder {
);
}
// Built HERE, below the fallible encoder open, NOT above it. `sws_getContext` returns a raw
// pointer whose only free is `Drop for NvencEncoder` — and `Drop` needs a CONSTRUCTED
// `Self`, which does not exist on `open`'s early returns (the intra-refresh-unsupported
// retry, which recurses into `Self::open`, and the plain error return). Creating the
// context above them leaked one per failed attempt, and `open_nvenc_probed`'s EINVAL
// bitrate ladder calls `open` up to ~10 times, so a host stepping its bitrate down leaked a
// context per step. Nothing between here and the `Ok(NvencEncoder { … })` below can return,
// so this placement makes the leak unrepresentable rather than merely unlikely.
// Built HERE, below the fallible encoder open, NOT above it — historically because the
// context's only free was `Drop for NvencEncoder`, which needs a CONSTRUCTED `Self` that
// does not exist on `open`'s early returns; creating it above them leaked one per failed
// attempt, and `open_nvenc_probed`'s EINVAL bitrate ladder calls `open` up to ~10 times.
// The owned `AvSwsContext` now frees itself on any exit, but the placement stays: it
// documents the dependency on the post-open `nvenc_pixel`, and there is no reason to
// build a context an early return would just throw away.
// CPU CSC paths: build the packed-RGB → planar swscale (no rescale) into the encoder's
// input frame. THREE users: 4:4:4 (RGB→YUV444P, BT.709, range per the flag), HDR
// (X2RGB10/X2BGR10→P010, BT.2020 limited — the PQ transfer is per-channel and rides
@@ -642,10 +644,10 @@ impl NvencEncoder {
// formats. Both dims are the encoder's positive `width`/`height` as `c_int`; `src_av` is a
// valid `AVPixelFormat` (from the `sws_src_pixel`-validated packed-RGB source), the dst is
// YUV444P (4:4:4) or P010LE (HDR). The trailing filter/param pointers are null = "use
// defaults" (documented as accepted). No Rust memory is borrowed; the returned pointer is
// null-checked below.
// defaults" (documented as accepted). No Rust memory is borrowed; ownership of the
// returned context passes to the `AvSwsContext` (null rejected by `from_raw`).
let sws = unsafe {
ffi::sws_getContext(
AvSwsContext::from_raw(ffi::sws_getContext(
width as c_int,
height as c_int,
src_av,
@@ -656,11 +658,11 @@ impl NvencEncoder {
ptr::null_mut(),
ptr::null_mut(),
ptr::null(),
)
))
};
if sws.is_null() {
let Some(sws) = sws else {
bail!("sws_getContext(RGB→{nvenc_pixel:?}) failed");
}
};
// Colour math applies to the CSC users ONLY. The expand is a pure byte shuffle —
// packed 3-bpp RGB/BGR to the same channels in 4 bytes, `nvenc_pixel` being `rgb0`/
// `bgr0` — and NVENC does the RGB→YUV itself downstream. Handing it a matrix + range
@@ -680,7 +682,16 @@ impl NvencEncoder {
SWS_CS_ITU709
});
let dst_range = i32::from(full_range_444);
ffi::sws_setColorspaceDetails(sws, cs, 1, cs, dst_range, 0, 1 << 16, 1 << 16);
ffi::sws_setColorspaceDetails(
sws.as_ptr(),
cs,
1,
cs,
dst_range,
0,
1 << 16,
1 << 16,
);
}
}
Some(sws)
@@ -694,10 +705,10 @@ impl NvencEncoder {
Some(VideoFrame::new(nvenc_pixel, width, height))
};
Ok(NvencEncoder {
sws_csc,
enc,
frame,
cuda: cuda_hw,
sws_csc,
want_444,
src_format: format,
width,
@@ -840,7 +851,7 @@ impl NvencEncoder {
// three CSC users (see `open`): 4:4:4 → planar YUV444P, HDR → P010, and the packed 3-bpp
// expand → `rgb0`/`bgr0`. The remaining branch below is the 4-bpp source, which needs no
// conversion at all — just a row copy honouring the destination stride.
if let Some(sws) = self.sws_csc {
if let Some(sws) = self.sws_csc.as_ref().map(AvSwsContext::as_ptr) {
let frame = self
.frame
.as_mut()
@@ -929,27 +940,23 @@ impl NvencEncoder {
// SAFETY: `frames_ref` is the non-null CUDA frames ctx from `self.cuda` (unwrapped via
// `.context(..)?` above), and the shared CUDA context was just made current on THIS thread
// (`make_current()?`), the precondition for the device-pointer copies below.
// * `av_frame_alloc` → `f` (null-checked). `av_hwframe_get_buffer(frames_ref, f, 0)` fills `f`
// with a pooled CUDA surface (sets `data[]`/`linesize[]`/`buf[0]`/`hw_frames_ctx`); on
// failure we free `f` and bail.
// * For NV12 we read `(*f).data[0..2]` / `linesize[0..2]` (Y + interleaved UV), else
// `data[0]`/`linesize[0]` — in-struct fields of the non-null `f`, valid for the surface dims
// ffmpeg allocated — and pass them to the cuda copy helpers, which device→device copy `buf`
// (the imported `DeviceBuffer`, owned by the caller and live for this call) into the surface.
// * On copy error we free `f` and return. Otherwise we write `pts`/`pict_type` through `f` and
// `avcodec_send_frame` it into the live owned `self.enc` context (which takes its own ref of
// the pooled surface), then free our `f` ref exactly once. Single-threaded encoder → no race.
// * `f` is an owned `AvFrame` — every exit below (bail, copy error, success) drops it
// exactly once, releasing its ref on the pooled surface. `av_hwframe_get_buffer` fills
// it with a pooled CUDA surface (sets `data[]`/`linesize[]`/`buf[0]`/`hw_frames_ctx`).
// * For NV12 we read `data[0..2]` / `linesize[0..2]` (Y + interleaved UV), else
// `data[0]`/`linesize[0]` — in-struct fields of the live frame, valid for the surface
// dims ffmpeg allocated — and pass them to the cuda copy helpers, which device→device
// copy `buf` (the imported `DeviceBuffer`, owned by the caller and live for this call)
// into the surface.
// * `avcodec_send_frame` takes its own ref of the pooled surface, so the drop afterwards
// is the sole owning free. Single-threaded encoder → no race.
unsafe {
let mut f = ffi::av_frame_alloc();
if f.is_null() {
bail!("av_frame_alloc failed");
}
let f = AvFrame::alloc().context("av_frame_alloc failed")?;
// Pooled CUDA surface: sets format, width/height, data[0]/linesize[0], buf[0] and
// hw_frames_ctx. Reused across frames (the pool recycles), keeping NVENC's
// registration cache warm.
let r = ffi::av_hwframe_get_buffer(frames_ref, f, 0);
let r = ffi::av_hwframe_get_buffer(frames_ref, f.as_ptr(), 0);
if r < 0 {
ffi::av_frame_free(&mut f);
bail!("av_hwframe_get_buffer(CUDA) failed ({r})");
}
// NV12 surfaces are two-plane (Y in data[0], interleaved UV in data[1]); YUV444
@@ -960,41 +967,36 @@ impl NvencEncoder {
let copy_res = if buf.yuv444 {
let dsts = core::array::from_fn(|i| {
(
(*f).data[i] as pf_zerocopy::cuda::CUdeviceptr,
(*f).linesize[i] as usize,
(*f.as_ptr()).data[i] as pf_zerocopy::cuda::CUdeviceptr,
(*f.as_ptr()).linesize[i] as usize,
)
});
pf_zerocopy::cuda::copy_yuv444_to_device(buf, dsts, true)
} else if self.want_444 {
ffi::av_frame_free(&mut f);
bail!(
"4:4:4 session but the zero-copy frame is not YUV444 (LINEAR/gamescope \
capture has no GPU 4:4:4 convert) unset PUNKTFUNK_ZEROCOPY to use the \
CPU 4:4:4 path on this compositor"
);
} else if buf.is_nv12() {
let y_ptr = (*f).data[0] as pf_zerocopy::cuda::CUdeviceptr;
let y_pitch = (*f).linesize[0] as usize;
let uv_ptr = (*f).data[1] as pf_zerocopy::cuda::CUdeviceptr;
let uv_pitch = (*f).linesize[1] as usize;
let y_ptr = (*f.as_ptr()).data[0] as pf_zerocopy::cuda::CUdeviceptr;
let y_pitch = (*f.as_ptr()).linesize[0] as usize;
let uv_ptr = (*f.as_ptr()).data[1] as pf_zerocopy::cuda::CUdeviceptr;
let uv_pitch = (*f.as_ptr()).linesize[1] as usize;
pf_zerocopy::cuda::copy_nv12_to_device(buf, y_ptr, y_pitch, uv_ptr, uv_pitch, true)
} else {
let dst_ptr = (*f).data[0] as pf_zerocopy::cuda::CUdeviceptr;
let dst_pitch = (*f).linesize[0] as usize;
let dst_ptr = (*f.as_ptr()).data[0] as pf_zerocopy::cuda::CUdeviceptr;
let dst_pitch = (*f.as_ptr()).linesize[0] as usize;
pf_zerocopy::cuda::copy_device_to_device(buf, dst_ptr, dst_pitch, true)
};
if let Err(e) = copy_res {
ffi::av_frame_free(&mut f);
return Err(e).context("copy imported buffer into NVENC surface");
}
(*f).pts = pts;
(*f).pict_type = if idr {
copy_res.context("copy imported buffer into NVENC surface")?;
(*f.as_ptr()).pts = pts;
(*f.as_ptr()).pict_type = if idr {
ffi::AVPictureType::AV_PICTURE_TYPE_I
} else {
ffi::AVPictureType::AV_PICTURE_TYPE_NONE
};
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), f);
ffi::av_frame_free(&mut f);
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), f.as_ptr());
if r < 0 {
bail!("avcodec_send_frame(CUDA) failed ({r})");
}
@@ -1003,16 +1005,9 @@ impl NvencEncoder {
}
}
impl Drop for NvencEncoder {
fn drop(&mut self) {
if let Some(sws) = self.sws_csc.take() {
// SAFETY: `sws` is the non-null `SwsContext` allocated by `sws_getContext` in `open` and
// owned exclusively by this encoder (taken out of the field so it can't be freed twice).
// `sws_freeContext` frees it; nothing else references it after this single-threaded drop.
unsafe { ffi::sws_freeContext(sws) };
}
}
}
// No `Drop` for `NvencEncoder`: `sws_csc` (`Option<AvSwsContext>`) frees itself, and as field #1
// it does so ahead of `enc`/`frame`/`cuda` — the same sequence the hand-written `Drop` performed
// (see the field-order note on the struct).
/// Serialises the save → `AV_LOG_FATAL` → restore window that every capability probe opens around
/// an encoder open it *expects* to fail.
@@ -63,8 +63,6 @@
// the signature. Clearing this file means DELETING the markers that carry no caller contract, not
// wrapping the calls — until then the lint is off HERE and enforced everywhere else.
#![allow(unsafe_op_in_unsafe_fn)]
// Every `unsafe` block / impl in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use super::nvenc_core::{
apply_low_latency_config, build_init_params, cached_ceiling, cached_split_verdict, codec_guid,
+86 -118
View File
@@ -19,8 +19,6 @@
//! hwdevice/hwframes/buffersrc/buffersink calls go through `ffmpeg::ffi` (= `ffmpeg_sys_next`),
//! as the CUDA encode path and the clients' decode paths already do. The encoder is opened
//! *without* a global header, so VPS/SPS/PPS are in-band on every IDR.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{Codec, EncodedFrame, Encoder};
use anyhow::{anyhow, bail, Context, Result};
@@ -36,8 +34,8 @@ use std::ptr;
use std::sync::{Mutex, OnceLock};
use super::libav::{
apply_low_latency_rc, pixel_to_av, poll_encoder, AvBuffer, AvFilterGraph, PollOutcome,
SWS_CS_ITU709, SWS_POINT,
apply_low_latency_rc, pixel_to_av, poll_encoder, AvBuffer, AvFilterGraph, AvFrame,
AvSwsContext, PollOutcome, SWS_CS_ITU709, SWS_POINT,
};
use ffmpeg::ffi; // = ffmpeg_sys_next
@@ -546,8 +544,13 @@ impl VaapiHw {
struct CpuInner {
enc: encoder::video::Encoder,
hw: VaapiHw,
sws: *mut ffi::SwsContext,
nv12: *mut ffi::AVFrame, // reusable software NV12 staging frame (swscale dst → upload src)
// FIELD ORDER IS LOAD-BEARING: the hand-written `Drop` this replaced freed `nv12` BEFORE
// `sws` — the reverse of the old declaration order — and field-DECLARATION order is what
// preserves that now (drop order follows declaration; an offset_of assert cannot pin it,
// repr(Rust) may lay memory out in any order).
/// Reusable software NV12/P010 staging frame (swscale dst → upload src).
nv12: AvFrame,
sws: AvSwsContext,
src_format: PixelFormat,
width: u32,
height: u32,
@@ -602,10 +605,10 @@ impl CpuInner {
// `src_av` is a valid `AVPixelFormat` (from `pixel_to_av` of the `vaapi_sws_src`-validated
// `src_pixel`), the dst is NV12/P010. The three trailing pointers (srcFilter, dstFilter,
// param) are explicitly null = "use defaults", which the API documents as accepted. No Rust
// memory is borrowed — only by-value ints/enums — and the returned pointer is null-checked
// just below.
// memory is borrowed — only by-value ints/enums — and ownership of the returned context
// passes to the `AvSwsContext` (null rejected by `from_raw`).
let sws = unsafe {
ffi::sws_getContext(
AvSwsContext::from_raw(ffi::sws_getContext(
width as c_int,
height as c_int,
src_av,
@@ -616,16 +619,15 @@ impl CpuInner {
ptr::null_mut(),
ptr::null_mut(),
ptr::null(),
)
))
};
if sws.is_null() {
let Some(sws) = sws else {
bail!(
"sws_getContext(RGB→{})",
if ten_bit { "P010" } else { "NV12" }
);
}
// SAFETY: `sws` is the non-null `SwsContext` from `sws_getContext` above (the `is_null()`
// check immediately preceding returned false). The coefficient table from
};
// SAFETY: `sws` is the live owned context from above. The coefficient table from
// `sws_getCoefficients` (ITU-709, or BT.2020 NCL for the HDR path — matching the VUI) is a
// libswscale static const valid for the whole process, reused here for both the inverse
// (src) and forward (dst) matrices. `sws_setColorspaceDetails` only reads those tables and
@@ -637,32 +639,22 @@ impl CpuInner {
} else {
SWS_CS_ITU709
});
ffi::sws_setColorspaceDetails(sws, cs, 1, cs, 0, 0, 1 << 16, 1 << 16);
ffi::sws_setColorspaceDetails(sws.as_ptr(), cs, 1, cs, 0, 0, 1 << 16, 1 << 16);
}
// SAFETY: `av_frame_alloc` returns a fresh, uniquely-owned heap `AVFrame` (null-checked — on
// null we free the already-built `sws` and bail). We then write the plain `format`/`width`/
// `height` fields through the non-null, properly-aligned `f` (sole owner, not yet shared).
// `av_frame_get_buffer(f, 0)` allocates backing storage for those dims/format; on failure we
// free `f` and `sws` (unwinding the half-built state) and bail. On success `f` is a fully-owned
// NV12/P010 frame stored in `CpuInner.nv12` and freed once in `CpuInner::drop`. `f` is a
// unique fresh pointer, so none of these writes alias anything.
let nv12 = unsafe {
let f = ffi::av_frame_alloc();
if f.is_null() {
ffi::sws_freeContext(sws);
bail!("av_frame_alloc(staging) failed");
}
(*f).format = staging_av as c_int;
(*f).width = width as c_int;
(*f).height = height as c_int;
if ffi::av_frame_get_buffer(f, 0) < 0 {
let mut f = f;
ffi::av_frame_free(&mut f);
ffi::sws_freeContext(sws);
let nv12 = AvFrame::alloc().context("av_frame_alloc(staging) failed")?;
// SAFETY: writing the plain `format`/`width`/`height` fields through the owned frame's
// pointer stays inside its allocation (sole owner, not yet shared).
// `av_frame_get_buffer` allocates backing storage for those dims/format; on failure the
// owned `nv12` (and the `sws` above it) simply drop — the hand-written unwind this
// replaced had to free both by hand on every branch.
unsafe {
(*nv12.as_ptr()).format = staging_av as c_int;
(*nv12.as_ptr()).width = width as c_int;
(*nv12.as_ptr()).height = height as c_int;
if ffi::av_frame_get_buffer(nv12.as_ptr(), 0) < 0 {
bail!("av_frame_get_buffer(staging) failed");
}
f
};
}
tracing::info!(
encoder = codec.vaapi_name(),
"VAAPI encode active ({width}x{height}@{fps}, CPU→{} upload path)",
@@ -671,8 +663,8 @@ impl CpuInner {
Ok(CpuInner {
enc,
hw,
sws,
nv12,
sws,
src_format: format,
width,
height,
@@ -693,49 +685,43 @@ impl CpuInner {
// `bytes.len() >= src_row * h`. `sws_scale` reads `h` rows of `src_row` bytes from
// `src_data[0] = bytes.as_ptr()` (the other planes null/0 — packed RGB is single-plane), all
// in bounds; `bytes`, `src_data`, `src_stride` are live locals for this synchronous call.
// `self.sws` is the non-null context built in `open`; it writes into `self.nv12` (a non-null
// owned frame whose `data`/`linesize` in-struct arrays were sized by `av_frame_get_buffer`).
// `av_frame_alloc` (null-checked) yields a fresh `hwf`; `av_hwframe_get_buffer` pulls a pooled
// VAAPI surface from the live non-null `self.hw.frames_ref`; `av_hwframe_transfer_data` uploads
// the staged NV12 into it — both frames live, failures free `hwf` and bail. We then write
// `pts`/`pict_type` through the non-null `hwf` and `avcodec_send_frame` it into the live
// owned `self.enc` context (which takes its own ref), then free our `hwf` ref exactly once.
// The encoder runs only on this thread (see `unsafe impl Send`), so no aliasing/data race.
// `self.sws` is the owned context built in `open`; it writes into `self.nv12` (an owned
// frame whose `data`/`linesize` in-struct arrays were sized by `av_frame_get_buffer`).
// `hwf` is an owned `AvFrame` — every exit below drops it exactly once, releasing its ref
// on the pooled VAAPI surface. `av_hwframe_get_buffer` pulls that surface from the live
// non-null `self.hw.frames_ref`; `av_hwframe_transfer_data` uploads the staged NV12 into
// it. `avcodec_send_frame` takes its own ref, so the drop afterwards is the sole owning
// free. The encoder runs only on this thread (see `unsafe impl Send`), so no
// aliasing/data race.
unsafe {
let src_data: [*const u8; 4] = [bytes.as_ptr(), ptr::null(), ptr::null(), ptr::null()];
let src_stride: [c_int; 4] = [src_row as c_int, 0, 0, 0];
if ffi::sws_scale(
self.sws,
self.sws.as_ptr(),
src_data.as_ptr(),
src_stride.as_ptr(),
0,
h as c_int,
(*self.nv12).data.as_ptr(),
(*self.nv12).linesize.as_ptr(),
(*self.nv12.as_ptr()).data.as_ptr(),
(*self.nv12.as_ptr()).linesize.as_ptr(),
) < 0
{
bail!("sws_scale RGB→NV12 failed");
}
let mut hwf = ffi::av_frame_alloc();
if hwf.is_null() {
bail!("av_frame_alloc(hw) failed");
}
if ffi::av_hwframe_get_buffer(self.hw.frames_ref.as_ptr(), hwf, 0) < 0 {
ffi::av_frame_free(&mut hwf);
let hwf = AvFrame::alloc().context("av_frame_alloc(hw) failed")?;
if ffi::av_hwframe_get_buffer(self.hw.frames_ref.as_ptr(), hwf.as_ptr(), 0) < 0 {
bail!("av_hwframe_get_buffer(VAAPI) failed");
}
if ffi::av_hwframe_transfer_data(hwf, self.nv12, 0) < 0 {
ffi::av_frame_free(&mut hwf);
if ffi::av_hwframe_transfer_data(hwf.as_ptr(), self.nv12.as_ptr(), 0) < 0 {
bail!("av_hwframe_transfer_data(→VAAPI) failed");
}
(*hwf).pts = pts;
(*hwf).pict_type = if idr {
(*hwf.as_ptr()).pts = pts;
(*hwf.as_ptr()).pict_type = if idr {
ffi::AVPictureType::AV_PICTURE_TYPE_I
} else {
ffi::AVPictureType::AV_PICTURE_TYPE_NONE
};
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), hwf);
ffi::av_frame_free(&mut hwf);
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), hwf.as_ptr());
if r < 0 {
bail!("avcodec_send_frame(VAAPI) failed ({r})");
}
@@ -744,24 +730,10 @@ impl CpuInner {
}
}
impl Drop for CpuInner {
fn drop(&mut self) {
// SAFETY: `self.nv12` (an owned `AVFrame`) and `self.sws` (an owned `SwsContext`) are each
// freed exactly once here, guarded by `is_null()` so a never-set pointer is skipped (no double
// free). `CpuInner` owns both exclusively and `Drop` runs once. `av_frame_free` takes `&mut`
// and nulls the pointer. `self.enc`/`self.hw` are freed afterward by their own `Drop` impls;
// the encoder holds its own `av_buffer_ref`'d device/frames copies, so field-drop order is
// irrelevant to soundness.
unsafe {
if !self.nv12.is_null() {
ffi::av_frame_free(&mut self.nv12);
}
if !self.sws.is_null() {
ffi::sws_freeContext(self.sws);
}
}
}
}
// No `Drop` for `CpuInner`: `nv12` (`AvFrame`) and `sws` (`AvSwsContext`) free themselves, in
// field-declaration order — the same nv12-then-sws sequence the hand-written `Drop` performed
// (see the field-order note on the struct). The encoder holds its own `av_buffer_ref`'d
// device/frames copies, so their order against `enc`/`hw` is irrelevant to soundness.
// ---------------------------------------------------------------------------------------------
// Zero-copy dmabuf path: DRM-PRIME → hwmap(vaapi) → scale_vaapi(nv12) filter graph → encode.
@@ -1043,16 +1015,20 @@ impl DmabufInner {
// whole synchronous `submit`; we describe one object/layer/plane from its
// fourcc/modifier/offset/stride and its `lseek`-queried size. `libc::lseek` on that live
// fd only reads the description's size and returns it (or -1); it touches no Rust memory.
// * `av_frame_alloc` → `drm` (null-checked); we set its scalar fields and
// `hw_frames_ctx = av_buffer_ref(self.drm_frames)` (new ref of the live owned ctx).
// * `drm`/`nv12` are owned `AvFrame`s — every exit drops each exactly once (the
// hand-placed frees this replaced were branch-clean, but only by inspection). We set
// `drm`'s scalar fields and `hw_frames_ctx = av_buffer_ref(self.drm_frames)` (new ref
// of the live owned ctx).
// * `data[0] = Box::into_raw(desc)` transfers the box into the frame; `buf[0] =
// av_buffer_create(.., free_desc, ..)` registers a destructor that reclaims it exactly once
// when the buffer's refcount hits zero — matched alloc/free, no leak/double-free.
// * `av_buffersrc_add_frame_flags(self.src, drm, KEEP_REF)` pushes a ref into the live
// buffersrc; KEEP_REF keeps our own `drm` ref, which we then `av_frame_free`. We pull the
// converted surface with `av_buffersink_get_frame(self.sink, nv12)` BEFORE returning, so the
// dmabuf (owned by the caller) is read while still valid. `nv12` is sent into the live owned
// `self.enc` (takes its own ref) and our ref freed once. Single-threaded encoder → no race.
// buffersrc; KEEP_REF keeps our own `drm` ref, dropped explicitly right after the push
// (the same point the hand-written free sat, kept so the descriptor's release timing
// across the pull does not change). We pull the converted surface with
// `av_buffersink_get_frame(self.sink, nv12)` BEFORE returning, so the dmabuf (owned by
// the caller) is read while still valid. `nv12` is sent into the live owned `self.enc`
// (takes its own ref) and dropped. Single-threaded encoder → no race.
unsafe {
// Build a DRM-PRIME AVFrame describing the dmabuf (one object/fd, one layer/plane).
let mut desc: Box<ffi::AVDRMFrameDescriptor> = Box::new(std::mem::zeroed());
@@ -1077,21 +1053,18 @@ impl DmabufInner {
desc.layers[0].planes[0].offset = dmabuf.offset as isize;
desc.layers[0].planes[0].pitch = dmabuf.stride as isize;
let mut drm = ffi::av_frame_alloc();
if drm.is_null() {
bail!("av_frame_alloc(drm) failed");
}
(*drm).format = ffi::AVPixelFormat::AV_PIX_FMT_DRM_PRIME as c_int;
(*drm).width = self.width as c_int;
(*drm).height = self.height as c_int;
let drm = AvFrame::alloc().context("av_frame_alloc(drm) failed")?;
(*drm.as_ptr()).format = ffi::AVPixelFormat::AV_PIX_FMT_DRM_PRIME as c_int;
(*drm.as_ptr()).width = self.width as c_int;
(*drm.as_ptr()).height = self.height as c_int;
// The dmabuf is the compositor's rendered desktop: full-range RGB. Tag the frame so
// the VPP's colour negotiation sees the real input instead of "unspecified" (an
// untagged input lets the driver pick its own default for the RGB→NV12 conversion —
// Mesa's is BT.601, contradicting the BT.709-limited VUI the encoder signals).
(*drm).color_range = ffi::AVColorRange::AVCOL_RANGE_JPEG;
(*drm).colorspace = ffi::AVColorSpace::AVCOL_SPC_RGB;
(*drm).hw_frames_ctx = ffi::av_buffer_ref(self.drm_frames.as_ptr());
(*drm).data[0] = Box::into_raw(desc) as *mut u8;
(*drm.as_ptr()).color_range = ffi::AVColorRange::AVCOL_RANGE_JPEG;
(*drm.as_ptr()).colorspace = ffi::AVColorSpace::AVCOL_SPC_RGB;
(*drm.as_ptr()).hw_frames_ctx = ffi::av_buffer_ref(self.drm_frames.as_ptr());
(*drm.as_ptr()).data[0] = Box::into_raw(desc) as *mut u8;
// Own the descriptor so it frees with the frame (the fd is owned by the DmabufFrame,
// which outlives this call — the graph reads the surface before submit returns).
extern "C" fn free_desc(_opaque: *mut std::ffi::c_void, data: *mut u8) {
@@ -1102,8 +1075,8 @@ impl DmabufInner {
// reclaims it exactly once — no double-free. `_opaque` is unused (we passed null).
unsafe { drop(Box::from_raw(data as *mut ffi::AVDRMFrameDescriptor)) };
}
(*drm).buf[0] = ffi::av_buffer_create(
(*drm).data[0],
(*drm.as_ptr()).buf[0] = ffi::av_buffer_create(
(*drm.as_ptr()).data[0],
std::mem::size_of::<ffi::AVDRMFrameDescriptor>(),
Some(free_desc),
ptr::null_mut(),
@@ -1113,45 +1086,40 @@ impl DmabufInner {
// Push through hwmap → scale_vaapi; pull the NV12 surface back out.
let r = ffi::av_buffersrc_add_frame_flags(
self.src,
drm,
drm.as_ptr(),
ffi::AV_BUFFERSRC_FLAG_KEEP_REF as c_int,
);
ffi::av_frame_free(&mut drm);
// These two stages ARE the import: the push hands libav our DRM-PRIME descriptor, and
// the pull is where `hwmap` actually maps it into a VA surface (and `scale_vaapi` runs
// the CSC). A failure here means this driver would not take this compositor's dmabuf —
// which no encoder rebuild can fix — so tell the process-wide latch, and capture
// negotiates CPU frames from the next session on. `avcodec_send_frame` below is
// deliberately NOT counted: that one is the encoder stalling, which the in-place
// rebuild above us exists to recover, and disabling zero-copy over it would be a
// permanent penalty for a transient fault.
drop(drm); // release our ref where the hand-written free sat (see the SAFETY note)
// These two stages ARE the import: the push hands libav our DRM-PRIME descriptor, and
// the pull is where `hwmap` actually maps it into a VA surface (and `scale_vaapi` runs
// the CSC). A failure here means this driver would not take this compositor's dmabuf —
// which no encoder rebuild can fix — so tell the process-wide latch, and capture
// negotiates CPU frames from the next session on. `avcodec_send_frame` below is
// deliberately NOT counted: that one is the encoder stalling, which the in-place
// rebuild above us exists to recover, and disabling zero-copy over it would be a
// permanent penalty for a transient fault.
if r < 0 {
let e = format!("av_buffersrc_add_frame failed ({r})");
pf_zerocopy::note_raw_dmabuf_import_failure(&e);
bail!("{e}");
}
t_push = t0.elapsed();
let mut nv12 = ffi::av_frame_alloc();
if nv12.is_null() {
bail!("av_frame_alloc(nv12) failed");
}
let r = ffi::av_buffersink_get_frame(self.sink, nv12);
let nv12 = AvFrame::alloc().context("av_frame_alloc(nv12) failed")?;
let r = ffi::av_buffersink_get_frame(self.sink, nv12.as_ptr());
if r < 0 {
ffi::av_frame_free(&mut nv12);
let e = format!("av_buffersink_get_frame failed ({r})");
pf_zerocopy::note_raw_dmabuf_import_failure(&e);
bail!("{e}");
}
pf_zerocopy::note_raw_dmabuf_import_ok();
t_pull = t0.elapsed() - t_push;
(*nv12).pts = pts;
(*nv12).pict_type = if idr {
(*nv12.as_ptr()).pts = pts;
(*nv12.as_ptr()).pict_type = if idr {
ffi::AVPictureType::AV_PICTURE_TYPE_I
} else {
ffi::AVPictureType::AV_PICTURE_TYPE_NONE
};
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), nv12);
ffi::av_frame_free(&mut nv12);
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), nv12.as_ptr());
if r < 0 {
bail!("avcodec_send_frame(VAAPI) failed ({r})");
}
+5 -5
View File
@@ -17,7 +17,7 @@
// child-module shape. External imports are this file's own; `vk_util` is a crate-root sibling,
// so the path is `crate::`, not the parent-relative `super::` the parent uses.
use super::*;
use crate::vk_util::{find_mem, make_plain_image, make_view};
use crate::vk_util::{ext_advertised, find_mem, make_plain_image, make_view};
use anyhow::{bail, Result};
use ash::vk;
use std::ffi::c_void;
@@ -53,10 +53,10 @@ pub(super) unsafe fn probe_rgb_direct(
let Ok(exts) = instance.enumerate_device_extension_properties(pd) else {
return Err("probe-failed(ext-enum)");
};
if !exts
.iter()
.any(|e| std::ffi::CStr::from_ptr(e.extension_name.as_ptr()) == vrgb::EXTENSION_NAME)
{
// Route through `vk_util::ext_advertised` rather than open-coding the walk a second time:
// this copy used the same unbounded `CStr::from_ptr` and had the same read-past-the-array
// hazard on a driver that fills all VK_MAX_EXTENSION_NAME_SIZE bytes without a NUL.
if !ext_advertised(&exts, vrgb::EXTENSION_NAME) {
return Err("no-ext(mesa<26.0-or-no-efc)");
}
// 2. Feature bit.
+32 -5
View File
@@ -19,11 +19,15 @@ use pf_frame::PixelFormat;
/// barriers were used without the extension ever being enabled; `pf-presenter/dmabuf.rs` is the
/// in-repo precedent that enables it).
pub(super) fn ext_advertised(exts: &[vk::ExtensionProperties], name: &std::ffi::CStr) -> bool {
exts.iter().any(|e| {
// SAFETY: `extension_name` is a spec-guaranteed NUL-terminated UTF-8 byte array inside
// the driver-filled `VkExtensionProperties` (VK_MAX_EXTENSION_NAME_SIZE bound).
unsafe { std::ffi::CStr::from_ptr(e.extension_name.as_ptr()) == name }
})
// `extension_name_as_c_str()` is ash's BOUNDED accessor: it stops at
// `VK_MAX_EXTENSION_NAME_SIZE` and returns `Err` when the array holds no NUL, so a
// malformed driver entry is a non-match rather than a read past the array. The previous
// `CStr::from_ptr(e.extension_name.as_ptr())` had no in-Rust bound at all — its SAFETY
// comment asserted the spec guarantee instead of enforcing it, so a driver that filled all
// 256 bytes without a terminator ran the walk into the NEXT `ExtensionProperties` and, on
// the last element, past the allocation. Same accessor `pyrowave.rs` already uses for the
// identical job. No unsafe, no unchecked read, same answer on every well-formed driver.
exts.iter().any(|e| e.extension_name_as_c_str() == Ok(name))
}
pub(crate) fn color_range(layer: u32) -> vk::ImageSubresourceRange {
@@ -453,6 +457,29 @@ mod tests {
));
}
/// A driver entry with NO terminator anywhere in `extension_name` must be a non-match, not a
/// read past the array.
///
/// This is the case the old `CStr::from_ptr(e.extension_name.as_ptr())` could not survive:
/// with every one of VK_MAX_EXTENSION_NAME_SIZE bytes non-NUL it walked into the NEXT
/// `ExtensionProperties`, and on the LAST element past the allocation entirely. The old test
/// only ever built well-formed, NUL-terminated entries, so it proved nothing about the bound
/// — which is why the hazard survived a SAFETY comment that asserted the spec guarantee
/// rather than enforcing it.
#[test]
fn ext_advertised_rejects_unterminated_name_without_overrunning() {
let mut bad = ash::vk::ExtensionProperties::default();
bad.extension_name.fill(b'A' as std::ffi::c_char);
// Deliberately LAST, so an unbounded walk would leave the whole array.
let exts = [ash::vk::ExtensionProperties::default(), bad];
assert!(!super::ext_advertised(
&exts,
ash::ext::queue_family_foreign::NAME
));
// And a name that is a prefix of the garbage still must not match.
assert!(!super::ext_advertised(&exts, c"AAAA"));
}
use super::*;
/// CSC mode (`bgra_target = false`): the 3→4 expand is a pure byte shuffle — no channel
-3
View File
@@ -41,9 +41,6 @@
//! worker caches it, so the steady state passes **zero** descriptors (the PipeWire pool recycles a
//! small buffer set).
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use anyhow::{Context, Result};
use pf_frame::{CapturedFrame, CursorOverlay, DmabufFrame, FramePayload, PixelFormat};
use pf_zerocopy::ipc;
+39 -20
View File
@@ -5,12 +5,16 @@
//! `libloading`), the device binding (D3D11 vs CUDA), input-surface registration, and the
//! Windows-only async retrieve — stay in their backends. Sibling of [`super::nvenc_status`].
// UNSAFE-LINT EXEMPTION (rationale + exit criteria: `unsafe_op_in_unsafe_fn` in the workspace
// Cargo.toml). This body is raw `nvEncodeAPI` entry-table calls almost line for line; narrowing it
// would add one `unsafe {}` plus one SAFETY comment per call that could only restate the signature.
// Clearing this file means DELETING the markers that carry no caller contract, not wrapping the
// calls — until then the lint is off HERE and enforced everywhere else.
#![allow(unsafe_op_in_unsafe_fn)]
// UNSAFE-LINT EXEMPTION REMOVED — the old fence rationale ("raw nvEncodeAPI entry-table calls
// almost line for line") was false for this file: it makes ZERO FFI calls. Its unsafe surface is
// C-union access whose soundness hangs entirely on which codec arm is active, and the 4:4:4 note
// below records the shipped bug (hevcConfig bytes stamped onto an AV1 config) that per-operation
// visibility makes findable. So this file runs the strictest discipline in the crate: every
// union READ, borrow, or bitfield-setter call sits in its own `unsafe {}` block naming the codec
// guard it relies on. (Plain union-arm field WRITES are safe by language rule — writing an arm
// cannot itself be UB; the hazard is the mismatched read — so those stay bare, guarded by the
// same codec matches.)
#![deny(clippy::multiple_unsafe_ops_per_block)]
use super::Codec;
use nvidia_video_codec_sdk::sys::nvEncodeAPI as nv;
@@ -694,10 +698,9 @@ mod tests {
};
assert_eq!(cfg.profileGUID, nv::NV_ENC_HEVC_PROFILE_FREXT_GUID);
// SAFETY: an HEVC session's union arm is `hevcConfig` — the one this path wrote.
unsafe {
assert_eq!(cfg.encodeCodecConfig.hevcConfig.chromaFormatIDC(), 3);
assert_eq!(cfg.encodeCodecConfig.hevcConfig.pixelBitDepthMinus8(), 2);
}
unsafe { assert_eq!(cfg.encodeCodecConfig.hevcConfig.chromaFormatIDC(), 3) };
// SAFETY: same HEVC arm as above.
unsafe { assert_eq!(cfg.encodeCodecConfig.hevcConfig.pixelBitDepthMinus8(), 2) };
}
#[test]
@@ -1210,6 +1213,8 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
// are the only accepted config). H.264 has no tier. Level 0 = autoselect for HEVC.
match c.codec {
Codec::H265 => {
// Plain union-arm writes are safe by language rule (the hazard is a mismatched
// READ later); the match on `c.codec` keeps the arm honest.
cfg.encodeCodecConfig.hevcConfig.tier = 1;
cfg.encodeCodecConfig.hevcConfig.level = 0;
}
@@ -1264,21 +1269,29 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
}
if want_444 && c.codec == Codec::H265 {
cfg.profileGUID = nv::NV_ENC_HEVC_PROFILE_FREXT_GUID;
cfg.encodeCodecConfig.hevcConfig.set_chromaFormatIDC(3);
// SAFETY: HEVC session (guarded by `c.codec == Codec::H265` on this branch), so
// `hevcConfig` is the active arm.
unsafe { cfg.encodeCodecConfig.hevcConfig.set_chromaFormatIDC(3) };
if c.bit_depth == 10 {
cfg.encodeCodecConfig.hevcConfig.set_pixelBitDepthMinus8(2); // Main 4:4:4 10
// SAFETY: same HEVC arm, same branch guard. (Main 4:4:4 10)
unsafe { cfg.encodeCodecConfig.hevcConfig.set_pixelBitDepthMinus8(2) };
}
} else if c.bit_depth == 10 {
match c.codec {
Codec::H265 => {
cfg.profileGUID = nv::NV_ENC_HEVC_PROFILE_MAIN10_GUID;
cfg.encodeCodecConfig.hevcConfig.set_pixelBitDepthMinus8(2);
// SAFETY: HEVC session (matched on `c.codec`), so `hevcConfig` is the active arm.
unsafe { cfg.encodeCodecConfig.hevcConfig.set_pixelBitDepthMinus8(2) };
}
Codec::Av1 => {
cfg.encodeCodecConfig.av1Config.set_pixelBitDepthMinus8(2);
cfg.encodeCodecConfig
.av1Config
.set_inputPixelBitDepthMinus8(c.av1_input_depth_minus8);
// SAFETY: AV1 session (matched on `c.codec`), so `av1Config` is the active arm.
unsafe { cfg.encodeCodecConfig.av1Config.set_pixelBitDepthMinus8(2) };
// SAFETY: same AV1 arm, same match guard.
unsafe {
cfg.encodeCodecConfig
.av1Config
.set_inputPixelBitDepthMinus8(c.av1_input_depth_minus8)
};
}
Codec::H264 => {} // no 10-bit H.264 encode on NVENC — negotiation never asks
Codec::PyroWave => unreachable!("PyroWave never opens the direct-NVENC backend"),
@@ -1306,7 +1319,9 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
};
match c.codec {
Codec::H265 => {
let vui = &mut cfg.encodeCodecConfig.hevcConfig.hevcVUIParameters;
// SAFETY: HEVC session (matched on `c.codec`), so `hevcConfig` is the active
// arm; the borrow is dropped before any other union access.
let vui = unsafe { &mut cfg.encodeCodecConfig.hevcConfig.hevcVUIParameters };
vui.videoSignalTypePresentFlag = 1;
vui.videoFullRangeFlag = 0;
vui.colourDescriptionPresentFlag = 1;
@@ -1315,7 +1330,9 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
vui.colourMatrix = mat;
}
Codec::H264 => {
let vui = &mut cfg.encodeCodecConfig.h264Config.h264VUIParameters;
// SAFETY: H.264 session (matched on `c.codec`), so `h264Config` is the active
// arm; the borrow is dropped before any other union access.
let vui = unsafe { &mut cfg.encodeCodecConfig.h264Config.h264VUIParameters };
vui.videoSignalTypePresentFlag = 1;
vui.videoFullRangeFlag = 0;
vui.colourDescriptionPresentFlag = 1;
@@ -1324,7 +1341,9 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
vui.colourMatrix = mat;
}
Codec::Av1 => {
let av1 = &mut cfg.encodeCodecConfig.av1Config;
// SAFETY: AV1 session (matched on `c.codec`), so `av1Config` is the active arm;
// the borrow is dropped before any other union access.
let av1 = unsafe { &mut cfg.encodeCodecConfig.av1Config };
av1.colorPrimaries = prim;
av1.transferCharacteristics = trc;
av1.matrixCoefficients = mat;
-2
View File
@@ -12,8 +12,6 @@
//! defaulting to BT.709 limited — true of every punktfunk client (`csc_rows` falls back to 709 on
//! "unspecified"), but NOT of vendor TV decoders, which guess colorimetry from RESOLUTION: an LG
//! webOS panel reads a 4K SDR stream as BT.2020 and renders it visibly washed out.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{EncodedFrame, Encoder};
use anyhow::{bail, ensure, Context, Result};
+10 -28
View File
@@ -49,8 +49,6 @@
// restate the signature. Clearing this file means DELETING the markers that carry no caller
// contract, not wrapping the calls — until then the lint is off HERE and enforced everywhere else.
#![allow(unsafe_op_in_unsafe_fn)]
// Every `unsafe` block / impl in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{ChromaFormat, Codec, EncodedFrame, Encoder, EncoderCaps};
use anyhow::{anyhow, bail, Context, Result};
@@ -2239,38 +2237,22 @@ impl Encoder for AmfEncoder {
mod tests {
use super::*;
/// The mirrored `AMFVariantStruct` must match the C layout: 4-byte tag + 4 padding + 16-byte
/// union = 24 bytes, align 8, payload at offset 8 (it is passed BY VALUE across the FFI).
// The LAYOUT of `AmfVariant`, `AmfGuid` and `AmfHdrMetadata` is no longer asserted here.
// Those checks moved to `const _: ()` assertions beside the mirrors themselves in
// `amf_sys.rs`, together with per-slot offset guards for the five vtables. As `#[test]`s
// they only ran when someone ran pf-encode's tests, on Windows, with AMF enabled — never in
// a release build, which is precisely where a mis-mirrored `AMFVariantStruct` would do its
// damage. As const assertions they hold on EVERY build that compiles the module.
//
// What stays here is the part a layout check cannot express: that the little-endian packing
// of the union payload matches what the C side will read out of those bytes.
#[test]
fn variant_layout_matches_c() {
assert_eq!(std::mem::size_of::<AmfVariant>(), 24);
assert_eq!(std::mem::align_of::<AmfVariant>(), 8);
assert_eq!(std::mem::offset_of!(AmfVariant, payload), 8);
fn variant_payload_packing_matches_c() {
let v = AmfVariant::from_rate(60, 1);
assert_eq!(v.payload[0], 60u64 | (1u64 << 32));
assert_eq!(AmfVariant::from_i64(-1).payload[0], u64::MAX);
}
/// `AMFGuid` is the flattened Win32-GUID layout (16 bytes).
#[test]
fn guid_layout_matches_c() {
assert_eq!(std::mem::size_of::<sys::AmfGuid>(), 16);
}
/// `AMFHDRMetadata` (components/ColorSpace.h): 8×u16 + 2×u32 + 2×u16 = 28 bytes, no padding.
#[test]
fn hdr_metadata_layout_matches_c() {
assert_eq!(std::mem::size_of::<sys::AmfHdrMetadata>(), 28);
assert_eq!(
std::mem::offset_of!(sys::AmfHdrMetadata, max_mastering_luminance),
16
);
assert_eq!(
std::mem::offset_of!(sys::AmfHdrMetadata, max_content_light_level),
24
);
}
/// A representative HDR10 grade for the live tests (BT.2020 primaries, 1000-nit mastering)
/// in [`HdrMeta`]'s ST.2086 wire units/order (primaries G, B, R).
fn sample_hdr_meta() -> punktfunk_core::quic::HdrMeta {
+120
View File
@@ -409,6 +409,126 @@ pub struct AmfBufferVtbl {
pub remove_observer_buffer: Slot,
}
// -- Layout guards ---------------------------------------------------------------------------
//
// THE CONTRACT, STATED ONCE. Everything above is a hand-written mirror of a C type this crate
// does not own and cannot include. Two classes of drift are possible and NEITHER fails to
// compile on its own:
//
// 1. A POD passed by value (`AmfVariant`, `AmfGuid`, `AmfHdrMetadata`) whose field offsets
// disagree with the C struct. The runtime then reads a tag or a payload out of the wrong
// bytes — `AmfVariant` crosses the FFI by value on EVERY `SetProperty`.
// 2. A vtable slot inserted, removed or reordered. `amf.rs` dispatches BY POSITION through
// these mirrors, so a shifted slot calls an arbitrary function pointer through a
// mismatched signature. There is no compile error, no runtime signal, and the failure is
// whatever the neighbouring AMF entry point happens to do with our arguments.
//
// `AMF_MIN_VERSION` does not defend against either: it checks a version NUMBER, not a layout,
// and it is a floor with no ceiling. The assertions below are the actual defence. They are
// `const _: ()` rather than `#[cfg(test)]` deliberately — the three POD checks below used to
// live only in `amf.rs`'s test module, which means they were verified only when someone ran
// pf-encode's tests, on Windows, with AMF enabled, and NEVER in a release build. This is the
// same hole `a8dd348b` closed for the cuda.h mirrors; it was missed here.
//
// Every slot index below was counted against the vtable declarations above. A slot is asserted
// when `amf.rs` calls it — those are the ones whose displacement is directly exploitable — plus
// the total size of each table, which catches an insertion PAST the last called slot (invisible
// to a per-slot check, but still a sign the mirror has drifted from the header).
/// One vtable slot. Every mirrored table is a flat array of these, so an offset in bytes is
/// always `index * SLOT`.
const SLOT: usize = core::mem::size_of::<Slot>();
/// Byte offset of vtable slot `i`. A `const fn` rather than a bare `i * SLOT` expression because
/// clippy's `erasing_op`/`identity_op` reject `0 * SLOT` and `1 * SLOT` under the `-D warnings`
/// the Windows CI leg runs with — and writing those two as bare `0` and `SLOT` would be the one
/// place the slot INDEX stops being visible, which is the entire readability of these assertions.
const fn slot(i: usize) -> usize {
i * SLOT
}
// Every slot is a plain code pointer, so all five tables are pointer-sized-array-shaped. If this
// ever fails, the tables are not flat arrays any more and every offset below is meaningless.
const _: () = assert!(SLOT == core::mem::size_of::<usize>());
const _: () = assert!(core::mem::align_of::<Slot>() == core::mem::align_of::<usize>());
// -- PODs crossing the FFI by value --
// `AMFVariantStruct`: 4-byte tag + 4 padding + 16-byte union = 24 bytes, payload at 8.
const _: () = assert!(core::mem::size_of::<AmfVariant>() == 24);
const _: () = assert!(core::mem::align_of::<AmfVariant>() == 8);
const _: () = assert!(core::mem::offset_of!(AmfVariant, payload) == 8);
// `AMFGuid`: the flattened Win32 GUID.
const _: () = assert!(core::mem::size_of::<AmfGuid>() == 16);
const _: () = assert!(core::mem::align_of::<AmfGuid>() == 4);
// `AMFHDRMetadata` (components/ColorSpace.h): 8×u16 + 2×u32 + 2×u16 = 28 bytes, no padding.
const _: () = assert!(core::mem::size_of::<AmfHdrMetadata>() == 28);
const _: () = assert!(core::mem::offset_of!(AmfHdrMetadata, max_mastering_luminance) == 16);
const _: () = assert!(core::mem::offset_of!(AmfHdrMetadata, max_content_light_level) == 24);
// -- AMFFactory (7 slots) — `create_context` 0, `create_component` 1 --
const _: () = assert!(core::mem::size_of::<AmfFactoryVtbl>() == slot(7));
const _: () = assert!(core::mem::offset_of!(AmfFactoryVtbl, create_context) == slot(0));
const _: () = assert!(core::mem::offset_of!(AmfFactoryVtbl, create_component) == slot(1));
// -- AMFContext (55 slots) = AMFInterface(3) + AMFPropertyStorage(10) + AMFContext(42) --
const _: () = assert!(core::mem::size_of::<AmfContextVtbl>() == slot(55));
const _: () = assert!(core::mem::offset_of!(AmfContextVtbl, release) == slot(1));
const _: () = assert!(core::mem::offset_of!(AmfContextVtbl, terminate) == slot(13));
const _: () = assert!(core::mem::offset_of!(AmfContextVtbl, init_dx11) == slot(18));
const _: () = assert!(core::mem::offset_of!(AmfContextVtbl, alloc_buffer) == slot(43));
const _: () =
assert!(core::mem::offset_of!(AmfContextVtbl, create_surface_from_dx11_native) == slot(49));
// -- AMFComponent (28 slots) = AMFInterface(3) + PropertyStorage(10) + StorageEx(4) + Component(11) --
const _: () = assert!(core::mem::size_of::<AmfComponentVtbl>() == slot(28));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, release) == slot(1));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, set_property) == slot(3));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, init) == slot(17));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, terminate) == slot(19));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, drain) == slot(20));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, flush) == slot(21));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, submit_input) == slot(22));
const _: () = assert!(core::mem::offset_of!(AmfComponentVtbl, query_output) == slot(23));
// -- AMFData (23 slots) = AMFInterface(3) + AMFPropertyStorage(10) + AMFData(10) --
const _: () = assert!(core::mem::size_of::<AmfDataVtbl>() == slot(23));
const _: () = assert!(core::mem::offset_of!(AmfDataVtbl, release) == slot(1));
const _: () = assert!(core::mem::offset_of!(AmfDataVtbl, query_interface) == slot(2));
const _: () = assert!(core::mem::offset_of!(AmfDataVtbl, set_property) == slot(3));
const _: () = assert!(core::mem::offset_of!(AmfDataVtbl, get_property) == slot(4));
const _: () = assert!(core::mem::offset_of!(AmfDataVtbl, set_pts) == slot(19));
// -- AMFBuffer (28 slots) = the AMFData prefix (23) + AMFBuffer(5) --
const _: () = assert!(core::mem::size_of::<AmfBufferVtbl>() == slot(28));
const _: () = assert!(core::mem::offset_of!(AmfBufferVtbl, release) == slot(1));
const _: () = assert!(core::mem::offset_of!(AmfBufferVtbl, get_size) == slot(24));
const _: () = assert!(core::mem::offset_of!(AmfBufferVtbl, get_native) == slot(25));
// -- The shared-prefix agreement --
// `AMFBuffer` derives from `AMFData`, and `create_surface_from_dx11_native` hands back an
// `AMFSurface*` that this module drives through the `AmfData` mirror on the strength of that
// single-inheritance prefix (see the comment on that slot). If the two mirrors ever disagree
// about where a shared slot lives, that reinterpretation is silently wrong — so assert the
// agreement rather than restating it in prose.
const _: () = assert!(
core::mem::offset_of!(AmfDataVtbl, release) == core::mem::offset_of!(AmfBufferVtbl, release)
);
const _: () = assert!(
core::mem::offset_of!(AmfDataVtbl, set_property)
== core::mem::offset_of!(AmfBufferVtbl, set_property)
);
const _: () = assert!(
core::mem::offset_of!(AmfDataVtbl, get_property)
== core::mem::offset_of!(AmfBufferVtbl, get_property)
);
const _: () = assert!(
core::mem::offset_of!(AmfDataVtbl, set_pts) == core::mem::offset_of!(AmfBufferVtbl, set_pts)
);
const _: () = assert!(
core::mem::offset_of!(AmfDataVtbl, get_duration)
== core::mem::offset_of!(AmfBufferVtbl, get_duration)
);
// -- DLL entry points (core/Factory.h; AMF_CDECL_CALL) --------------------------------------
pub type AmfQueryVersionFn = unsafe extern "C" fn(*mut u64) -> AmfResult;
pub type AmfInitFn = unsafe extern "C" fn(u64, *mut *mut AmfFactory) -> AmfResult;
+89 -119
View File
@@ -37,8 +37,6 @@
//! through `ffmpeg::ffi` (= `ffmpeg_sys_next`), exactly as the Linux CUDA/VAAPI paths do. The
//! `AVD3D11VADeviceContext`/`AVD3D11VAFramesContext` layouts are mirrored (the bindings don't
//! allowlist `hwcontext_d3d11va.h`), as [`super::linux`] mirrors `AVCUDADeviceContext`.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{ChromaFormat, Codec, EncodedFrame, Encoder};
use anyhow::{anyhow, bail, Context, Result};
@@ -61,8 +59,8 @@ use windows::Win32::Graphics::Dxgi::Common::{
};
use super::libav::{
apply_low_latency_rc, pixel_to_av, poll_encoder, AvBuffer, PollOutcome, SWS_CS_BT2020,
SWS_CS_ITU709, SWS_POINT,
apply_low_latency_rc, pixel_to_av, poll_encoder, AvBuffer, AvFrame, AvSwsContext, PollOutcome,
SWS_CS_BT2020, SWS_CS_ITU709, SWS_POINT,
};
use ffmpeg::ffi; // = ffmpeg_sys_next
@@ -499,10 +497,14 @@ fn immediate_context(device: &ID3D11Device) -> ID3D11DeviceContext {
struct SystemInner {
enc: encoder::video::Encoder,
// FIELD ORDER IS LOAD-BEARING: the hand-written `Drop` this replaced freed `sw_frame`
// before `sws`, and field-DECLARATION order is what preserves that now (an offset_of assert
// cannot pin this — repr(Rust) may reorder memory independently of declaration order, and
// drop order follows declaration).
/// Reusable software NV12/P010 frame: swscale dst / readback dst, and the `send_frame` src.
sw_frame: *mut ffi::AVFrame,
/// swscale ctx for the BGRA→NV12 fallback (built lazily; null for the YUV-readback path).
sws: *mut ffi::SwsContext,
sw_frame: AvFrame,
/// swscale ctx for the BGRA→NV12 fallback (built lazily; `None` for the YUV-readback path).
sws: Option<AvSwsContext>,
/// CPU-readable staging texture for the D3D11 readback (built lazily on the captured device).
staging: Option<ID3D11Texture2D>,
ctx: Option<ID3D11DeviceContext>,
@@ -549,26 +551,18 @@ impl SystemInner {
ptr::null_mut(),
)?
};
// SAFETY: `av_frame_alloc` returns a freshly-allocated, uniquely-owned `AVFrame` (null-checked
// before any deref); writing `format`/`width`/`height` through `*f` stays inside that
// allocation. `av_frame_get_buffer(f, 0)` allocates the backing planes — on failure we
// `av_frame_free` the sole owner (no double-free) and bail; on success the raw `f` is moved into
// `self.sw_frame` and freed exactly once in `Drop`.
let sw_frame = unsafe {
let f = ffi::av_frame_alloc();
if f.is_null() {
bail!("av_frame_alloc(sw) failed");
}
(*f).format = sw_av as c_int;
(*f).width = width as c_int;
(*f).height = height as c_int;
if ffi::av_frame_get_buffer(f, 0) < 0 {
let mut f = f;
ffi::av_frame_free(&mut f);
let sw_frame = AvFrame::alloc().context("av_frame_alloc(sw) failed")?;
// SAFETY: writing `format`/`width`/`height` through the owned frame's pointer stays inside
// its allocation. `av_frame_get_buffer` allocates the backing planes — on failure the
// owned `sw_frame` simply drops (freed once, by the wrapper).
unsafe {
(*sw_frame.as_ptr()).format = sw_av as c_int;
(*sw_frame.as_ptr()).width = width as c_int;
(*sw_frame.as_ptr()).height = height as c_int;
if ffi::av_frame_get_buffer(sw_frame.as_ptr(), 0) < 0 {
bail!("av_frame_get_buffer(sw) failed");
}
f
};
}
tracing::info!(
encoder = vendor.encoder_name(codec),
"{} encode active ({width}x{height}@{fps}, system-memory {} path)",
@@ -578,7 +572,7 @@ impl SystemInner {
Ok(SystemInner {
enc,
sw_frame,
sws: ptr::null_mut(),
sws: None,
staging: None,
ctx: None,
format,
@@ -634,13 +628,13 @@ impl SystemInner {
// frame and `self.enc`'s own context, both live for the call and neither retained by libav
// (it references the frame's buffers itself).
unsafe {
(*self.sw_frame).pts = pts;
(*self.sw_frame).pict_type = if idr {
(*self.sw_frame.as_ptr()).pts = pts;
(*self.sw_frame.as_ptr()).pict_type = if idr {
ffi::AVPictureType::AV_PICTURE_TYPE_I
} else {
ffi::AVPictureType::AV_PICTURE_TYPE_NONE
};
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), self.sw_frame);
let r = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), self.sw_frame.as_ptr());
if r < 0 {
bail!("avcodec_send_frame({} system) failed ({r})", "ffmpeg_win");
}
@@ -705,10 +699,10 @@ impl SystemInner {
let total = pitch.saturating_mul(h + h.div_ceil(2));
let mapped = std::slice::from_raw_parts(base, total);
let chroma_off = pitch * h;
let y_dst = (*self.sw_frame).data[0];
let y_stride = (*self.sw_frame).linesize[0] as usize;
let uv_dst = (*self.sw_frame).data[1];
let uv_stride = (*self.sw_frame).linesize[1] as usize;
let y_dst = (*self.sw_frame.as_ptr()).data[0];
let y_stride = (*self.sw_frame.as_ptr()).linesize[0] as usize;
let uv_dst = (*self.sw_frame.as_ptr()).data[1];
let uv_stride = (*self.sw_frame.as_ptr()).linesize[1] as usize;
for y in 0..h {
let s = &mapped[y * pitch..y * pitch + row_bytes];
ptr::copy_nonoverlapping(s.as_ptr(), y_dst.add(y * y_stride), row_bytes);
@@ -748,7 +742,7 @@ impl SystemInner {
let pitch = map.RowPitch as usize;
let h = self.height as usize;
let base = map.pData as *const u8;
self.ensure_sws(
let sws = self.ensure_sws(
pixel_to_av(Pixel::BGRA),
ffi::AVPixelFormat::AV_PIX_FMT_NV12,
SWS_CS_ITU709,
@@ -756,13 +750,13 @@ impl SystemInner {
let src_data: [*const u8; 4] = [base, ptr::null(), ptr::null(), ptr::null()];
let src_stride: [c_int; 4] = [pitch as c_int, 0, 0, 0];
let r = ffi::sws_scale(
self.sws,
sws,
src_data.as_ptr(),
src_stride.as_ptr(),
0,
h as c_int,
(*self.sw_frame).data.as_ptr(),
(*self.sw_frame).linesize.as_ptr(),
(*self.sw_frame.as_ptr()).data.as_ptr(),
(*self.sw_frame.as_ptr()).linesize.as_ptr(),
);
ctx.Unmap(&staging, 0);
if r < 0 {
@@ -798,7 +792,7 @@ impl SystemInner {
let h = self.height as usize;
let base = map.pData as *const u8;
// RGB(BT.2020 PQ) → YUV(BT.2020 PQ): a matrix-only repack (same PQ transfer), full→limited.
self.ensure_sws(
let sws = self.ensure_sws(
ffi::AVPixelFormat::AV_PIX_FMT_X2BGR10LE,
ffi::AVPixelFormat::AV_PIX_FMT_P010LE,
SWS_CS_BT2020,
@@ -806,13 +800,13 @@ impl SystemInner {
let src_data: [*const u8; 4] = [base, ptr::null(), ptr::null(), ptr::null()];
let src_stride: [c_int; 4] = [pitch as c_int, 0, 0, 0];
let r = ffi::sws_scale(
self.sws,
sws,
src_data.as_ptr(),
src_stride.as_ptr(),
0,
h as c_int,
(*self.sw_frame).data.as_ptr(),
(*self.sw_frame).linesize.as_ptr(),
(*self.sw_frame.as_ptr()).data.as_ptr(),
(*self.sw_frame.as_ptr()).linesize.as_ptr(),
);
ctx.Unmap(&staging, 0);
if r < 0 {
@@ -844,7 +838,7 @@ impl SystemInner {
// `width`×`height`). `bytes` is borrowed for the call only and never aliases the owned
// `sw_frame`. `send` then hands `sw_frame` to the encoder.
unsafe {
self.ensure_sws(
let sws = self.ensure_sws(
pixel_to_av(sws_src(format)?),
ffi::AVPixelFormat::AV_PIX_FMT_NV12,
SWS_CS_ITU709,
@@ -852,13 +846,13 @@ impl SystemInner {
let src_data: [*const u8; 4] = [bytes.as_ptr(), ptr::null(), ptr::null(), ptr::null()];
let src_stride: [c_int; 4] = [src_row as c_int, 0, 0, 0];
if ffi::sws_scale(
self.sws,
sws,
src_data.as_ptr(),
src_stride.as_ptr(),
0,
h as c_int,
(*self.sw_frame).data.as_ptr(),
(*self.sw_frame).linesize.as_ptr(),
(*self.sw_frame.as_ptr()).data.as_ptr(),
(*self.sw_frame.as_ptr()).linesize.as_ptr(),
) < 0
{
bail!("sws_scale RGB→NV12 failed");
@@ -872,23 +866,24 @@ impl SystemInner {
/// 10-bit RGB10→P010 BT.2020), so caching a single context is sound.
///
/// Safe: every argument is a plain libav enum/int, and the context it caches belongs to `self`
/// (freed once in `Drop`).
/// (an owned `AvSwsContext`, freed by its own drop). Returns the borrowed pointer for the
/// caller's `sws_scale` — borrowed only, `self.sws` stays the owner.
fn ensure_sws(
&mut self,
src_av: ffi::AVPixelFormat,
dst_av: ffi::AVPixelFormat,
cs: c_int,
) -> Result<()> {
if !self.sws.is_null() {
return Ok(());
) -> Result<*mut ffi::SwsContext> {
if let Some(sws) = &self.sws {
return Ok(sws.as_ptr());
}
// SAFETY: `sws_getContext` takes only scalars plus the documented "no filters, no params"
// null trio, and returns an owned context or null — which is checked before use, so
// `sws_setColorspaceDetails` and the store below only ever see a live one.
// `sws_getCoefficients` returns a pointer into libav's own static tables, valid for the
// process, and the call only reads it.
// null trio, and returns an owned context or null — `from_raw` rejects the null, so
// `sws_setColorspaceDetails` only ever sees a live one, and ownership passes to the
// `AvSwsContext`. `sws_getCoefficients` returns a pointer into libav's own static tables,
// valid for the process, and the call only reads it.
let sws = unsafe {
let sws = ffi::sws_getContext(
let raw = ffi::sws_getContext(
self.width as c_int,
self.height as c_int,
src_av,
@@ -900,36 +895,22 @@ impl SystemInner {
ptr::null_mut(),
ptr::null(),
);
if sws.is_null() {
let Some(owned) = AvSwsContext::from_raw(raw) else {
bail!("sws_getContext(RGB→YUV) failed");
}
};
// Source full-range RGB → destination limited-range YUV (matches the limited-range VUI
// we signal). For RGB input the src coefficient table is unused; pass dst for both.
let coeff = ffi::sws_getCoefficients(cs);
ffi::sws_setColorspaceDetails(sws, coeff, 1, coeff, 0, 0, 1 << 16, 1 << 16);
sws
ffi::sws_setColorspaceDetails(owned.as_ptr(), coeff, 1, coeff, 0, 0, 1 << 16, 1 << 16);
owned
};
self.sws = sws;
Ok(())
Ok(self.sws.insert(sws).as_ptr())
}
}
impl Drop for SystemInner {
fn drop(&mut self) {
// SAFETY: `sw_frame` is the `AVFrame` allocated in `open` (or null) — `av_frame_free` drops it
// once and nulls the pointer through the `&mut`; `sws` is the cached `SwsContext` (or null) —
// `sws_freeContext` frees it once. This `Drop` runs exactly once and `SystemInner` owns both
// exclusively, so there is no double-free or use-after-free.
unsafe {
if !self.sw_frame.is_null() {
ffi::av_frame_free(&mut self.sw_frame);
}
if !self.sws.is_null() {
ffi::sws_freeContext(self.sws);
}
}
}
}
// No `Drop` for `SystemInner`: `sw_frame` (`AvFrame`) and `sws` (`Option<AvSwsContext>`) free
// themselves, in field-declaration order — the same sw_frame-then-sws sequence the hand-written
// `Drop` performed, pinned by the offset_of assert at the struct.
// ---------------------------------------------------------------------------------------------
// Zero-copy D3D11 path (the AMF default; QSV opt-in — see `zerocopy_enabled`): share the capture
@@ -1214,32 +1195,29 @@ impl ZeroCopyInner {
}
fn submit(&mut self, frame: &D3d11Frame, pts: i64, idr: bool) -> Result<()> {
// SAFETY: `d3d = av_frame_alloc()` is a fresh owned frame (null-checked) and is `av_frame_free`d
// exactly once on every path below. `av_hwframe_get_buffer` fills it from the pool — on failure
// we free it and bail. `(*d3d).data[0]` is the pool's texture-array and `data[1]` the array
// index; `from_raw_borrowed` borrows that `ID3D11Texture2D` WITHOUT taking ownership (no Release
// — the frame owns it) and is null-checked. `src` (the captured texture) and `dst` (the pooled
// slice) live on the SAME D3D11 device wrapped by `self.hw`, and the caller guarantees
// `captured.format == pool_format` before calling, so `CopySubresourceRegion(dst, dst_index, ..,
// src, 0, ..)` on the single-threaded immediate context `self.ctx` is a valid same-format GPU
// copy. For QSV the mapped `qsv` frame is a fresh owned frame whose `hw_frames_ctx` takes an
// `av_buffer_ref` of `self.qsv_frames`; it is `av_frame_free`d (releasing that ref) on both the
// map-failure and success paths. `avcodec_send_frame` only internally refs the input frame, so
// the `av_frame_free(d3d)`/`av_frame_free(qsv)` afterwards are the sole owning frees — no leak,
// no double-free, no use-after-free.
// SAFETY: `d3d`/`qsv` are owned `AvFrame`s, so EVERY exit — including the three `?` exits
// between the pool pull and the send, which as hand-placed frees previously leaked the
// frame plus one of the POOL-sized hwframe surfaces per failure (eight failures wedged
// the encoder permanently) — unrefs the pooled surface back to the pool. `(*d3d).data[0]`
// is the pool's texture-array and `data[1]` the array index; `from_raw_borrowed` borrows
// that `ID3D11Texture2D` WITHOUT taking ownership (no Release — the frame owns it) and is
// null-checked. `src` (the captured texture) and `dst` (the pooled slice) live on the
// SAME D3D11 device wrapped by `self.hw`, and the caller guarantees `captured.format ==
// pool_format` before calling, so `CopySubresourceRegion(dst, dst_index, .., src, 0, ..)`
// on the single-threaded immediate context `self.ctx` is a valid same-format GPU copy.
// For QSV the mapped `qsv` frame's `hw_frames_ctx` takes an `av_buffer_ref` of
// `self.qsv_frames`; its drop at the end of the arm releases that ref at the same point
// the hand-written free did. `avcodec_send_frame` only internally refs the input frame,
// so the drops are the sole owning frees — no leak, no double-free, no use-after-free.
unsafe {
// Pull a pooled D3D11 surface; its data[0] is the pool's texture-ARRAY, data[1] the slice.
let mut d3d = ffi::av_frame_alloc();
if d3d.is_null() {
bail!("av_frame_alloc(d3d11) failed");
}
let r = ffi::av_hwframe_get_buffer(self.hw.frames_ref.as_ptr(), d3d, 0);
let d3d = AvFrame::alloc().context("av_frame_alloc(d3d11) failed")?;
let r = ffi::av_hwframe_get_buffer(self.hw.frames_ref.as_ptr(), d3d.as_ptr(), 0);
if r < 0 {
ffi::av_frame_free(&mut d3d);
bail!("av_hwframe_get_buffer(D3D11) failed ({r})");
}
let dst_ptr = (*d3d).data[0] as *mut c_void;
let dst_index = (*d3d).data[1] as usize as u32;
let dst_ptr = (*d3d.as_ptr()).data[0] as *mut c_void;
let dst_index = (*d3d.as_ptr()).data[1] as usize as u32;
let dst_tex = ID3D11Texture2D::from_raw_borrowed(&dst_ptr)
.ok_or_else(|| anyhow!("pooled D3D11 frame has null texture"))?;
// GPU-local copy of the captured slice into the pooled array slice (like NVENC's CUDA
@@ -1249,58 +1227,50 @@ impl ZeroCopyInner {
self.ctx
.CopySubresourceRegion(&dst, dst_index, 0, 0, 0, &src, 0, None);
(*d3d).pts = pts;
(*d3d).pict_type = if idr {
(*d3d.as_ptr()).pts = pts;
(*d3d.as_ptr()).pict_type = if idr {
ffi::AVPictureType::AV_PICTURE_TYPE_I
} else {
ffi::AVPictureType::AV_PICTURE_TYPE_NONE
};
let send = match self.vendor {
WinVendor::Amf => ffi::avcodec_send_frame(self.enc.as_mut_ptr(), d3d),
WinVendor::Amf => ffi::avcodec_send_frame(self.enc.as_mut_ptr(), d3d.as_ptr()),
WinVendor::Qsv => {
// Map the D3D11 frame to a QSV surface (1:1, no copy), then send the mapped frame.
let mut qsv = ffi::av_frame_alloc();
if qsv.is_null() {
ffi::av_frame_free(&mut d3d);
bail!("av_frame_alloc(qsv) failed");
}
let qsv = AvFrame::alloc().context("av_frame_alloc(qsv) failed")?;
// Always `Some` on this arm — `open` fills the pair for `WinVendor::Qsv` and
// leaves it `None` only for AMF — but say so with a bail rather than an unwrap,
// matching the null check above it. The `Option` is what the raw pointer's
// "null means AMF" convention was already encoding.
let Some(qsv_frames) = self.qsv_frames.as_ref() else {
ffi::av_frame_free(&mut qsv);
ffi::av_frame_free(&mut d3d);
bail!("QSV send path without a derived QSV frames context");
};
(*qsv).format = ffi::AVPixelFormat::AV_PIX_FMT_QSV as c_int;
(*qsv).hw_frames_ctx = ffi::av_buffer_ref(qsv_frames.as_ptr());
(*qsv.as_ptr()).format = ffi::AVPixelFormat::AV_PIX_FMT_QSV as c_int;
(*qsv.as_ptr()).hw_frames_ctx = ffi::av_buffer_ref(qsv_frames.as_ptr());
// The map flags are a bindgen enum (no BitOr) — cast each to int before OR-ing.
let r = ffi::av_hwframe_map(
qsv,
d3d,
qsv.as_ptr(),
d3d.as_ptr(),
ffi::AV_HWFRAME_MAP_DIRECT as c_int | ffi::AV_HWFRAME_MAP_READ as c_int,
);
if r < 0 {
ffi::av_frame_free(&mut qsv);
ffi::av_frame_free(&mut d3d);
bail!("av_hwframe_map(D3D11→QSV) failed ({r})");
}
(*qsv).pts = pts;
(*qsv).pict_type = (*d3d).pict_type;
let s = ffi::avcodec_send_frame(self.enc.as_mut_ptr(), qsv);
ffi::av_frame_free(&mut qsv);
s
(*qsv.as_ptr()).pts = pts;
(*qsv.as_ptr()).pict_type = (*d3d.as_ptr()).pict_type;
ffi::avcodec_send_frame(self.enc.as_mut_ptr(), qsv.as_ptr())
// `qsv` drops here — releasing the mapped frame and its frames-ctx ref at the
// same point the hand-written `av_frame_free(&mut qsv)` did.
}
};
ffi::av_frame_free(&mut d3d);
if send < 0 {
bail!(
"avcodec_send_frame({}) failed ({send})",
self.vendor.label()
);
}
// `d3d` drops here (and on every early exit above), returning the pooled surface.
}
Ok(())
}
@@ -39,8 +39,6 @@
// the signature. Clearing this file means DELETING the markers that carry no caller contract, not
// wrapping the calls — until then the lint is off HERE and enforced everywhere else.
#![allow(unsafe_op_in_unsafe_fn)]
// Every `unsafe` block / impl in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use super::nvenc_core::{
apply_low_latency_config, build_init_params, cached_ceiling, codec_guid, plan_range_recovery,
-3
View File
@@ -37,9 +37,6 @@
//! it stays behind the same gate and falls back to IDR wherever the driver declines. 4:4:4 stays
//! `false` until probed on real hardware (design §8.6).
// Every `unsafe` block / impl in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{ChromaFormat, Codec, EncodedFrame, Encoder, EncoderCaps};
use anyhow::{anyhow, bail, Context, Result};
use libvpl_sys as vpl;
-1
View File
@@ -12,7 +12,6 @@
// `#[cfg(test)]` instead.
// Every unsafe block in this module tree carries a `// SAFETY:` proof; enforce it (unsafe-proof
// program). As a parent module this also covers the child modules (windows/linux backends).
#![deny(clippy::undocumented_unsafe_blocks)]
use anyhow::Result;
use pf_frame::{CapturedFrame, PixelFormat};
+42 -27
View File
@@ -7,9 +7,6 @@
//! The win32u GPU-preference hook, the HDR/video-engine converters, and the self-tests stay in the
//! capture crate — they are capture mechanics, not shared identity.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use anyhow::{Context, Result};
use windows::core::Interface;
use windows::Win32::Foundation::{HMODULE, LUID};
@@ -158,18 +155,26 @@ enum PrioMode {
Off,
/// A fixed class the operator pinned (`normal`=2 / `high`=4 / `realtime`=5).
Static(i32),
/// The default: HIGH immediately, then upgrade to REALTIME when it is safe — HAGS off, or
/// Opt-in (`auto`): HIGH immediately, then upgrade to REALTIME when it is safe — HAGS off, or
/// HAGS on with comfortable VRAM headroom (with a monitor that downgrades the moment VRAM
/// tightens). REALTIME is the proven ceiling-raiser (it is how our brief encode preempts a
/// saturating game), but REALTIME + NVIDIA + HAGS + near-full VRAM is a documented NVENC
/// hang the gate takes the win everywhere it cannot hit the hazard.
/// tightens). REALTIME is the T2.3 ceiling-raiser (a higher-priority context preempts at
/// pixel granularity), but it carries TWO field-proven hazards: REALTIME + NVIDIA + HAGS +
/// near-full VRAM is a documented NVENC hang (the VRAM gate covers that one), and on AMD the
/// upgrade itself produced a metronomic content-starving stall class (~3.6 s period, RX 9070
/// XT, 2026-08-12 A/B: pinning `high` removed it) that no VRAM gate can see — which is why
/// `auto` is no longer the default.
Auto,
}
/// Resolve `PUNKTFUNK_GPU_PRIORITY_CLASS` (`off|normal|high|realtime|auto`, default **auto**).
/// Resolve `PUNKTFUNK_GPU_PRIORITY_CLASS` (`off|normal|high|realtime|auto`, default **high**).
/// D3DKMT_SCHEDULINGPRIORITYCLASS: IDLE 0, BELOW_NORMAL 1, NORMAL 2, ABOVE_NORMAL 3, HIGH 4,
/// REALTIME 5. `realtime` pins REALTIME statically (no gate — the operator owns the hazard);
/// `high` restores the pre-T2.3 static default.
/// `auto` is the T2.3 gated-REALTIME mode, opt-in since the 2026-08-12 field A/B convicted the
/// REALTIME upgrade of its own metronomic stall class on AMD (see [`PrioMode::Auto`]) — HIGH is
/// the Sunshine/Apollo-parity lever that delivered the original decisive win, and the default
/// must not hold REALTIME anywhere (the same inversion as the vdisplay driver's `PFVD_RT_GPU`
/// ladder, which fixed the faster ~1.8 s metronome the same day). Unrecognized values read as
/// the default, not as `auto` — a typo must not opt a box into the hazard.
fn configured_gpu_priority_mode() -> PrioMode {
match std::env::var("PUNKTFUNK_GPU_PRIORITY_CLASS")
.ok()
@@ -177,9 +182,10 @@ fn configured_gpu_priority_mode() -> PrioMode {
{
Some("off") => PrioMode::Off,
Some("normal") => PrioMode::Static(2),
Some("high") => PrioMode::Static(4),
Some("realtime") => PrioMode::Static(5),
_ => PrioMode::Auto,
Some("auto") => PrioMode::Auto,
// `high`, unset, and anything unrecognized all land on the HIGH default.
_ => PrioMode::Static(4),
}
}
@@ -278,14 +284,17 @@ unsafe fn d3dkmt_set_scheduling_priority_class(
/// GPU-saturated game our capture+encode process is starved of GPU time slices — NVENC sits ~idle but
/// `lock_bitstream` waits ~20 ms for our context to be scheduled. Elevating the PROCESS GPU scheduling
/// priority class (the strong cross-process lever — far more effective than `SetGPUThreadPriority`
/// alone, which we measured as no help) lets our brief encode preempt the game. Default is the
/// T2.3 `auto` mode: HIGH immediately here, then [`auto_priority_gate`] upgrades to REALTIME
/// where the NVIDIA+HAGS+full-VRAM NVENC-hang hazard cannot bite (and a monitor downgrades when
/// it could). Runs once per process; best-effort.
/// `PUNKTFUNK_GPU_PRIORITY_CLASS = off|normal|high|realtime|auto` (default auto; `high` = the
/// pre-gate static behavior; `realtime` = pinned, operator owns the hazard). Best-effort:
/// silently no-ops under a UAC-filtered token (the process will not hold SE_INC_BASE_PRIORITY,
/// so the D3DKMT call is a no-op).
/// alone, which we measured as no help) lets our brief encode preempt the game. Default is a
/// static HIGH — the class that delivered that win. The T2.3 `auto` mode (HIGH here, then
/// [`auto_priority_gate`] upgrades to REALTIME behind the NVENC-hang VRAM gate) is opt-in since
/// the 2026-08-12 field A/B: on AMD the REALTIME upgrade generated its own metronomic
/// content-starving stall class (~3.6 s period) that the VRAM gate cannot see, and pinning HIGH
/// removed it. Runs once per process; best-effort.
/// `PUNKTFUNK_GPU_PRIORITY_CLASS = off|normal|high|realtime|auto` (default high; `auto` = the
/// gated-REALTIME upgrade, operator opts into the AMD stall hazard for the extra ceiling;
/// `realtime` = pinned, operator owns every hazard). Best-effort: silently no-ops under a
/// UAC-filtered token (the process will not hold SE_INC_BASE_PRIORITY, so the D3DKMT call is a
/// no-op).
fn elevate_process_gpu_priority() {
use std::sync::Once;
static ONCE: Once = Once::new();
@@ -319,17 +328,23 @@ fn elevate_process_gpu_priority() {
});
}
// --- REALTIME auto-gate (gpu-contention §5.C / latency plan T2.3) --------------------------------
// --- REALTIME auto-gate (gpu-contention §5.C / latency plan T2.3) — OPT-IN since 2026-08-12 ------
//
// REALTIME GPU scheduling priority is the genuine cross-process ceiling-raiser under a saturating
// game (a higher-priority context preempts at pixel granularity — the Async-TimeWarp mechanism),
// and our SYSTEM service uniquely holds the SE_INC_BASE_PRIORITY it needs. The one documented
// hazard: REALTIME + NVIDIA + HAGS-on + near-full VRAM can hang NVENC. So: probe HAGS once via
// D3DKMT; HAGS off ⇒ REALTIME unconditionally; HAGS on ⇒ REALTIME gated on LOCAL-segment VRAM
// headroom, with a monitor thread that downgrades to HIGH the moment usage crosses
// [`VRAM_DOWNGRADE_PCT`] of the OS budget and restores REALTIME after it has stayed under
// [`VRAM_RESTORE_PCT`] for [`VRAM_RESTORE_TICKS`] consecutive polls (hysteresis against flapping
// on the boundary of the hazard window).
// and our SYSTEM service uniquely holds the SE_INC_BASE_PRIORITY it needs. Two field-proven
// hazards bound it. (1) REALTIME + NVIDIA + HAGS-on + near-full VRAM can hang NVENC — the VRAM
// gate below exists for that one: probe HAGS once via D3DKMT; HAGS off ⇒ REALTIME
// unconditionally; HAGS on ⇒ REALTIME gated on LOCAL-segment VRAM headroom, with a monitor
// thread that downgrades to HIGH the moment usage crosses [`VRAM_DOWNGRADE_PCT`] of the OS
// budget and restores REALTIME after it has stayed under [`VRAM_RESTORE_PCT`] for
// [`VRAM_RESTORE_TICKS`] consecutive polls (hysteresis against flapping on the boundary of the
// hazard window). (2) On AMD (RX 9070 XT A/B), a punktfunk process holding REALTIME generated a
// metronomic content-starving stall class — every ~3.6 s ALL processes' presents paused
// 150800 ms with the GPU responsive — that no VRAM gate can see, and the vdisplay driver's
// REALTIME swap-chain raise produced the same pathology on its own ~1.8 s beat. That second
// hazard is why the whole gate now runs only under an explicit `auto`, and the default stays a
// static HIGH.
/// Downgrade REALTIME→HIGH when local VRAM usage exceeds this share of the OS budget.
const VRAM_DOWNGRADE_PCT: u64 = 92;
-1
View File
@@ -10,7 +10,6 @@
//! tuning), and — on Windows — [`dxgi`] (the capture identity + D3D11 device creation).
// Unsafe-proof program: every `unsafe {}` / `unsafe impl` must carry a `// SAFETY:` proof.
#![deny(clippy::undocumented_unsafe_blocks)]
pub mod hdr;
pub mod metronome;
-3
View File
@@ -11,9 +11,6 @@
//! state) auto-revert at thread exit (= session end); the process-wide bits revert at process exit.
//! See `design/host-latency-plan.md` Tier 3A.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
#[cfg(target_os = "windows")]
mod imp {
#![allow(non_snake_case)]
-3
View File
@@ -3,9 +3,6 @@
//! can't deschedule them; the native, GameStream, and direct-NVENC send threads all reach this the
//! same way (`pf_frame::thread_qos::boost_thread_priority`).
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
/// Raise the current thread's OS scheduling priority so a CPU-heavy game can't deschedule our
/// capture/encode/send threads. This matters even though our GPU work is already HIGH priority: the
/// GPU scheduler can only favour commands we've actually SUBMITTED, so if a normal-priority thread is
-1
View File
@@ -23,7 +23,6 @@
//! live session actually encodes on, for the console's "in use" display.
// Unsafe-proof program: every `unsafe {}` in this leaf carries a `// SAFETY:` proof.
#![deny(clippy::undocumented_unsafe_blocks)]
use anyhow::Result;
use serde::{Deserialize, Serialize};
+10
View File
@@ -144,6 +144,13 @@ pub struct HostConfig {
/// text ("Living Room PC"); the DNS-level `<label>.local.` target keeps using a sanitized
/// machine-safe label, so a spacey display name can't produce an invalid mDNS record.
pub host_name: Option<String>,
/// `PUNKTFUNK_GAMESTREAM` — enable the GameStream/Moonlight-compat planes (nvhttp pairing,
/// RTSP, ENet control, `_nvstream` mDNS) from `host.env`, equivalent to the `--gamestream`
/// CLI flag (either source turns it on). **Default OFF** — the secure native-only host: the
/// compat planes carry plain-HTTP pairing + the legacy GCM-nonce path (security-review
/// #5/#9), so stock-Moonlight support is opt-in on every route, and the packaged units ship
/// without the flag so this knob is how a package user opts in.
pub gamestream: bool,
/// `PUNKTFUNK_ENCODER` — explicit encoder-backend override (lowercased; empty = auto-detect by GPU vendor).
pub encoder_pref: String,
/// `PUNKTFUNK_RENDER_ADAPTER` — discrete render-GPU pin by description substring (`Some` even when empty:
@@ -356,6 +363,9 @@ impl HostConfig {
host_name: val("PUNKTFUNK_HOST_NAME")
.map(|s| s.trim().to_string())
.filter(|s| !s.is_empty()),
// Default OFF, explicit-on grammar: the Moonlight-compat planes are opt-in
// everywhere (see the field doc); `--gamestream` on the CLI also turns them on.
gamestream: env_on("PUNKTFUNK_GAMESTREAM").unwrap_or(false),
encoder_pref: std::env::var("PUNKTFUNK_ENCODER")
.unwrap_or_default()
.to_ascii_lowercase(),
@@ -15,9 +15,6 @@
//! `<linux/uinput.h>` on x86_64. `/dev/uinput` needs a udev rule + `input` group membership
//! (see `scripts/60-punktfunk.rules`); creation fails with a clear error otherwise.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use crate::pad_slots::PadSlots;
use anyhow::{bail, Result};
use punktfunk_core::input::{gamepad, GamepadFrame, MAX_PADS};
@@ -17,8 +17,6 @@
//! output's logical rectangle — the same shape the libei backend uses with its EI region.
#![allow(clippy::all, dead_code, non_camel_case_types, non_snake_case, unused)]
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{gs_button_to_evdev, vk_to_evdev, InputEvent, InputInjector};
use anyhow::{Context, Result};
-3
View File
@@ -6,9 +6,6 @@
//! to evdev/US), and translate events into virtual pointer/keyboard requests, tracking modifier
//! state so the compositor resolves shifted keysyms correctly.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{gs_button_to_evdev, vk_to_evdev, InputEvent, InputInjector};
use anyhow::{bail, Context, Result};
use punktfunk_core::input::InputKind;
@@ -15,9 +15,6 @@
//! with its position (never at a stale point), tip edges get their own DOWN/UP frames, and a
//! range-leave is a final frame without `INRANGE`.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use anyhow::{Context, Result};
use punktfunk_core::input::{InputEvent, InputKind};
use punktfunk_core::quic::{
@@ -14,9 +14,6 @@
//! user's, and any layout re-reads a *position* as a *character* — on a German host that is
//! exactly the y↔z swap / ü-on-ö scramble.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use anyhow::Result;
use punktfunk_core::input::{InputEvent, InputKind};
use std::mem::size_of;
-7
View File
@@ -14,13 +14,6 @@
// Scaffold: trait methods + per-OS backends are defined ahead of the target that uses them.
#![allow(dead_code)]
// Every unsafe block in this crate carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
// …and its companion: without this, an `unsafe fn` body needs no blocks, so an unproven FFI call
// could hide inside one and still satisfy the deny above. The workspace keeps
// `unsafe_op_in_unsafe_fn` at `warn` while the encoder backends are cleared; this crate is at zero.
#![deny(unsafe_op_in_unsafe_fn)]
use anyhow::Result;
use punktfunk_core::input::{InputEvent, InputKind};
-1
View File
@@ -17,7 +17,6 @@
//! the decode chain there is Vulkan → D3D11VA → software.
// Unsafe-proof program: every `unsafe {}` in this crate carries a `// SAFETY:` proof.
#![deny(clippy::undocumented_unsafe_blocks)]
// THE VULKAN CONTRACT, stated once - most `// SAFETY:` proofs in this crate are an instance of it.
//
+6
View File
@@ -14,6 +14,12 @@
//! and per-platform; it lives with the product that does it (`punktfunk-host::update`,
//! `pf-client-core::update`, and the root helper in `pf-update`).
// This crate parses a SIGNED, NETWORK-FETCHED manifest and, per the header above, "owns the part
// where being wrong is a security bug". Signature verification is worthless if the parser around
// it can be made to read out of bounds, so the absence of unsafe here is a security property and
// is now enforced rather than merely true today.
#![forbid(unsafe_code)]
/// The Ed25519 public keys trusted for update manifests — two slots, so a key rotation is
/// "sign with the new one, ship builds trusting both, retire the old" (the plugin-store
/// `OFFICIAL_KEYS` drill) rather than a flag day. The private half is the
+3
View File
@@ -18,3 +18,6 @@ path = "src/main.rs"
[target.'cfg(target_os = "linux")'.dependencies]
serde = { version = "1", features = ["derive"] }
serde_json = "1"
[lints]
workspace = true
+21 -2
View File
@@ -25,6 +25,12 @@
//! (root-written, world-readable) for the unprivileged caller to read; stdout/stderr land in
//! the unit's journal.
// ROOT RUNS THIS. `deny` rather than `forbid` only because of the single `geteuid` call in
// `linux_main::effective_uid`, which carries the one localized `#[allow(unsafe_code)]` in the
// crate and explains there why it is not worth a dependency to remove. Any NEW unsafe anywhere
// in this helper is a build error.
#![deny(unsafe_code)]
#[cfg(target_os = "linux")]
mod linux_main {
use serde::Serialize;
@@ -313,8 +319,7 @@ mod linux_main {
};
// Effective root is required for every leg; refuse early with a clear message
// rather than half-running.
// SAFETY: geteuid has no preconditions.
if unsafe { libc_geteuid() } != 0 {
if effective_uid() != 0 {
eprintln!("pf-update: must run as root (start punktfunk-update.service)");
std::process::exit(1);
}
@@ -397,6 +402,20 @@ mod linux_main {
#[link_name = "geteuid"]
fn libc_geteuid() -> u32;
}
/// The crate's ONLY unsafe operation, isolated so the crate-level `deny(unsafe_code)` can
/// stand and the exemption is one named function rather than a whole call site.
///
/// Deliberately NOT rewritten to `rustix::process::geteuid()`: this crate's Cargo.toml states
/// that the zero-dependency posture *is* a security invariant of a root helper ("no HTTP
/// client, no TLS, no argument parsing"), so pulling in a general-purpose syscall crate to
/// delete one `unsafe` would trade a real property for a cosmetic one.
#[allow(unsafe_code)]
fn effective_uid() -> u32 {
// SAFETY: `geteuid` is a POSIX syscall wrapper that takes no arguments, reads no memory
// through a pointer, cannot fail, and has no preconditions whatsoever.
unsafe { libc_geteuid() }
}
}
#[cfg(target_os = "linux")]
+7
View File
@@ -78,6 +78,13 @@
//! a `VASurfaceID` rather than an index — so the conversion will take that table as
//! a parameter and stay pure.
// The header above states the crate's whole design constraint: it is the CPU-testable half, it
// links no libva, and it compiles on macOS — "which is the point". That constraint is exactly
// what `forbid(unsafe_code)` encodes. The crate is full of hand-declared libva `repr(C)` mirrors,
// and the moment one of them gets dereferenced through a raw pointer here, the crate has quietly
// become the other half and stops being testable off a Linux box with a GPU.
#![forbid(unsafe_code)]
pub mod config;
pub mod drm;
pub mod pic;
+14 -5
View File
@@ -16,13 +16,9 @@ publish = false
[dependencies]
punktfunk-core = { path = "../punktfunk-core", features = ["quic"] }
pf-frame = { path = "../pf-frame" }
pf-gpu = { path = "../pf-gpu" }
pf-host-config = { path = "../pf-host-config" }
pf-paths = { path = "../pf-paths" }
pf-win-display = { path = "../pf-win-display" }
# The Windows admission gate consults NVENC's session budget (can_open_another_session).
pf-encode = { path = "../pf-encode" }
anyhow = "1"
tracing = "0.1"
# The platform-neutral policy/identity/custom-preset state is serde-serialized (persisted + the mgmt
@@ -41,8 +37,12 @@ hex = "0.4"
# the shipped host's dependency closure through this crate is unchanged.
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
[target.'cfg(target_os = "linux")'.dependencies]
# `proc`'s process-group tree guard is Unix-wide, not Linux-only: the module is compiled on every
# platform and its tests run on whatever the developer is sitting at (macOS, here).
[target.'cfg(unix)'.dependencies]
libc = "0.2"
[target.'cfg(target_os = "linux")'.dependencies]
# The Mutter backend drives D-Bus RemoteDesktop + ScreenCast.RecordVirtual via ashpd on a tokio
# runtime; the gamescope restore worker + portal handshakes use tokio too.
ashpd = { version = "0.13", features = ["screencast", "remote_desktop"] }
@@ -61,6 +61,15 @@ bitflags = "2"
x11rb = { version = "0.13", default-features = false }
[target.'cfg(target_os = "windows")'.dependencies]
# Windows-only, all three, and gated here rather than unconditionally so the LINUX build does not
# drag their closures in for nothing: `pf-frame` for the DXGI capture identity + the CTA-861.3 HDR
# luminance fields, `pf-gpu` for the render-adapter LUID, and `pf-encode` for the admission gate's
# NVENC session budget (`can_open_another_session`, admission.rs, itself `#[cfg(windows)]`). Every
# use site of all three is Windows-gated — verified by grep — and between them they pull FFmpeg,
# ash and openh264, none of which a Linux host reaches through this crate.
pf-frame = { path = "../pf-frame" }
pf-gpu = { path = "../pf-gpu" }
pf-encode = { path = "../pf-encode" }
# The host<->driver wire contract for the pf-vdisplay IddCx backend (control IOCTLs + Pod structs).
pf-driver-proto = { path = "../pf-driver-proto" }
bytemuck = { version = "1.19", features = ["derive"] }
+173 -29
View File
@@ -8,25 +8,37 @@
//! * **KWin** — privileged `zkde_screencast_unstable_v1::stream_virtual_output` ([`kwin`]).
//! * **wlroots/Sway** — `swaymsg create_output` + `output mode --custom` ([`wlroots`]).
//! * **Mutter/GNOME** — D-Bus `RemoteDesktop` + `ScreenCast.RecordVirtual` ([`mutter`]).
//! * **Hyprland** — `hyprctl output create headless` + the xdg-desktop-portal-hyprland ScreenCast
//! portal. Its own backend, not a wlroots dialect (`design/hyprland-support.md` D1).
//! * **gamescope** — three sub-modes behind one backend ([`GamescopeRoute`]): bare
//! **spawn** of a nested headless session, host-**managed** `gamescope-session-plus`/SteamOS
//! takeover, and **attach** to a session somebody else started. By far the largest backend here,
//! because it owns session lifecycle rather than just minting an output.
//! * **monitor mirror** — no virtual display at all: stream a PHYSICAL head the compositor already
//! has (the `PUNKTFUNK_CAPTURE_MONITOR` pin), reporting [`DisplayOwnership::External`] so none of
//! the lifecycle policy is applied to someone else's screen.
//! * **Windows pf-vdisplay** — the all-Rust IddCx driver + its `manager`, the sole Windows backend.
//!
//! No list of file sizes here: it rots. The rule instead — the Linux backends plus the Windows
//! manager are the bulk of this crate, and the platform-neutral half (`policy`, `registry`,
//! `lifecycle`, `layout`, `identity`, `admission`, `monitors`, `session`, `routing`, `proc`,
//! `portal_config`) is the minority that every platform's CI actually compiles and tests.
//!
//! [`VirtualDisplay::create`] returns a [`VirtualOutput`]: the PipeWire node to capture plus an
//! owned keepalive whose `Drop` releases the output (RAII — no explicit `destroy`). Capture
//! consumes the node via the host `capture::capture_virtual_output`.
// `dead_code` is ENFORCED on Linux, where ~10k of this crate's ~17k lines live. Off elsewhere for
// one structural reason: `proc`, `session`, `routing`, `monitors` and `lifecycle` are declared
// unconditionally but exist to serve the Linux backends, so on Windows/macOS most of their surface
// is legitimately unreferenced. Scoping it this way rather than crate-wide keeps the platform that
// owns the code honest. (Was a bare crate-wide allow whose "scaffold, defined ahead of the target
// that uses them" rationale had stopped being true.)
// `dead_code` is ENFORCED on Linux, where the clear majority of this crate lives — every compositor
// backend under `vdisplay/linux/` plus everything only they consume, which is roughly half the crate
// on its own and the half that carries the session-lifecycle risk. Off elsewhere for one structural
// reason: `proc`, `session`, `routing`, `monitors` and `lifecycle` are declared unconditionally but
// exist to serve the Linux backends, so on Windows/macOS most of their surface is legitimately
// unreferenced. Note what that waives: the Windows backend (`vdisplay/windows/`, itself thousands of
// lines) gets NO dead-code enforcement, so an orphaned Windows path has to be found by review.
// Scoping it this way rather than crate-wide still keeps the platform that owns most of the code
// honest. (Was a bare crate-wide allow whose "scaffold, defined ahead of the target that uses them"
// rationale had stopped being true.)
#![cfg_attr(not(target_os = "linux"), allow(dead_code))]
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
// …and that program only covers a whole `unsafe fn` body once the body needs its own block: in
// edition 2021 `unsafe_op_in_unsafe_fn` is allow-by-default, which exempted this crate's hardest
// FFI from the deny above — every IOCTL wrapper, and `restore_displays_ccd`, the call the whole
// Windows teardown path depends on to give the operator their physical panels back.
#![deny(unsafe_op_in_unsafe_fn)]
use anyhow::Result;
pub use punktfunk_core::Mode;
@@ -200,9 +212,16 @@ impl Compositor {
/// The compositor backends usable on this host *right now*: gamescope wherever its binary is
/// installed (it spawns a nested session — independent of the running desktop), plus the live
/// session's own compositor (KWin / Mutter / wlroots / Hyprland) when the host runs inside it.
/// Cheap, side-effect-free probes — safe to call per management request. A concrete client
/// preference is validated against this set before it's honored (see the punktfunk/1 handshake's
/// resolution).
/// Side-effect-free, but **not cheap, and not memoized**: every call re-walks `/proc`
/// ([`detect_active_session`]), and each backend probe that the live/pinned short-circuit below does
/// not exempt does real work — `gamescope::is_available` FORKS `gamescope --version`,
/// `kwin::is_available` does a Wayland registry roundtrip, `wlroots`/`hyprland` read a socket path
/// and `mutter` a D-Bus name. So a console polling `/host/compositors` on a KDE box still forks a
/// gamescope per poll, on a thread the caller must therefore not assume is cheap to block (mgmt
/// calls it inline on the async runtime). Callers wanting a hot path should cache the answer;
/// treating this as free is what the "cheap, safe per management request" claim this doc used to
/// make invited. A concrete client preference is validated against this set before it's honored
/// (see the punktfunk/1 handshake's resolution).
///
/// The **live session is the primary signal**, ahead of each backend's own probe. Those probes read
/// the process env (`XDG_CURRENT_DESKTOP` for Mutter, `WAYLAND_DISPLAY` for KWin's registry
@@ -311,7 +330,12 @@ pub fn detect() -> Result<Compositor> {
if let Some(c) = compositor_for_kind(detect_active_session().kind) {
return Ok(c);
}
let desktop = std::env::var("XDG_CURRENT_DESKTOP")
// Under [`ENV_LOCK`]: `apply_session_env` `set_var`s — and, for a dead session,
// `remove_var`s — this very key from another session's `spawn_blocking`, and a glibc
// `getenv` concurrent with a `setenv` is the `environ` realloc data race ENV_LOCK exists
// for (it is UB regardless of which key each side touches, so "different variable" is no
// defence). Read-then-drop: only the read needs serializing.
let desktop = with_env_lock(|| std::env::var("XDG_CURRENT_DESKTOP"))
.unwrap_or_default()
.to_ascii_uppercase();
if desktop.contains("KDE") {
@@ -559,13 +583,18 @@ pub fn effective_topology() -> policy::Topology {
return resolve_topology(e.topology);
}
// Unconfigured: honor a legacy operator env if present (a host runs one desktop backend, so at
// most one of these is set), else the Auto default.
let legacy = [
"PUNKTFUNK_KWIN_VIRTUAL_PRIMARY",
"PUNKTFUNK_MUTTER_VIRTUAL_PRIMARY",
]
.iter()
.find_map(|k| std::env::var(k).ok());
// most one of these is set), else the Auto default. Read under [`ENV_LOCK`] like every other
// env read on the session-setup path: this runs inside `create`, concurrent with another
// session's `apply_session_env` `set_var`s, and glibc's `environ` realloc makes a racing
// `getenv` UB no matter that these particular keys are ones nobody writes.
let legacy = with_env_lock(|| {
[
"PUNKTFUNK_KWIN_VIRTUAL_PRIMARY",
"PUNKTFUNK_MUTTER_VIRTUAL_PRIMARY",
]
.iter()
.find_map(|k| std::env::var(k).ok())
});
match legacy.as_deref().map(str::trim) {
Some("1" | "true" | "yes" | "on") => policy::Topology::Exclusive,
Some("0" | "false" | "no" | "off") => policy::Topology::Extend,
@@ -637,19 +666,92 @@ pub fn gamescope_composites_cursor() -> bool {
///
/// A host-managed `gamescope-session-plus` / SteamOS session counts as a spawn: we own its
/// `GAMESCOPE_BIN` wrapper (or PATH shim), so the flags are ours.
///
/// **Ask the resolved ROUTE, never the env.** This used to test the spawn-vs-attach term by reading
/// `PUNKTFUNK_GAMESCOPE_NODE`, which worked only while `apply_input_env` PUBLISHED its decision into
/// that key. Phase 2.3 deleted the publication (routing.rs: "Nothing is written back to the two
/// knobs") and left the key as an operator override — rung 2 of a 6-rung ladder — so the session
/// that reaches [`GamescopeRoute::Attach`] at the ladder's rung 5 instead (a foreign gamescope on an
/// infra-less box), and the monitor-pin mirror that never consults the ladder at all, both answered
/// "ours". The two consequences were silent and unrecoverable: the punktfunk/1 Welcome fixed the
/// session at 10-bit BT.2020/PQ against a foreign 8-bit SDR composite, and the host skipped the
/// XFixes cursor reconstruction for a session whose gamescope was never given
/// `--pipewire-composite-cursor` — a stream with no pointer in it at all.
///
/// **Two residual gaps**, both of which need a route this crate cannot see from here:
///
/// * the ladder is re-run with `dedicated_launch = false`, since a capability query carries no
/// session context — so it cannot see the one input that would move a session from
/// Managed/Attach to Spawn. On a box with no session infrastructure AND a foreign gamescope
/// running, a `game_session=dedicated` launch really takes rung 3 (`Spawn`) while this re-run
/// takes rung 5 (`Attach`) and answers "foreign";
/// * `create_managed_session` can degrade a resolved `Managed` to an ATTACH at create time (a
/// mask-fragile DM it may not stop — it then mirrors the box's own game-mode session). That
/// happens after this answer is due, and the ladder re-run here still says `Managed`, so such a
/// session is still credited with flags it does not own.
///
/// The second over-promises. The first UNDER-promises, and `false` is the deliberate choice for an
/// input we cannot see, because the two directions do not cost the same: over-promising fixes the
/// punktfunk/1 Welcome at 10-bit PQ against an 8-bit SDR composite and leaves a stream with **no
/// pointer at all**, while under-promising costs HDR and draws the pointer twice. But do not read
/// that as "fails closed": it is not, for the cursor. `gamescope::cursor_args` adds
/// `--pipewire-composite-cursor` from the BINARY probe alone, ungated by this answer, so on the
/// bare spawn above gamescope paints the pointer into the node while the host's
/// `session_plan::gamescope_needs_host_cursor` (`gamescope && !gamescope_composites_cursor()`) also
/// blends the XFixes pointer on top — two pointers, plus the encoder pushed off its zero-copy arm.
/// Do not "fix" that by re-running the ladder with a guessed `dedicated_launch = true`: that trades
/// the mild failure for the severe one on every non-launching session. Both gaps close the same
/// way, and only that way: give these two functions the session's own [`GamescopeRoute`] (which
/// `SessionContext` already carries) and have the backend report the degrade — a change to two
/// public signatures and every host call site, i.e. work outside this crate.
fn gamescope_ours_and(#[cfg(target_os = "linux")] probe: fn() -> bool) -> bool {
#[cfg(target_os = "linux")]
{
let attaching = with_env_lock(|| std::env::var_os("PUNKTFUNK_GAMESCOPE_NODE").is_some());
!attaching && probe()
// `probe` first: it is memoized (the `--version` banner is parsed once per process), while
// the route resolution walks `/proc` for a foreign gamescope. On a box with a stock
// gamescope the answer is already `false` and the walk never happens.
probe()
&& !session_is_a_foreign_gamescope(
capture_monitor().is_some(),
resolve_gamescope_route(Compositor::Gamescope, false).as_ref(),
)
}
#[cfg(not(target_os = "linux"))]
false
}
// Platform-neutral per-client stable display-id map (Stage 3): Windows seeds the monitor EDID +
// ConnectorIndex from the id; KWin names its output from it. `allow(dead_code)` because only Windows
// consumes it in non-test code today — the KWin wiring is the next Stage-3 step.
/// Pure predicate behind [`gamescope_ours_and`]: is the gamescope this session will use one
/// SOMEBODY ELSE started, whose spawn flags we therefore cannot vouch for?
///
/// Two ways to land on a foreign session, and both must count:
///
/// * `mirror_pinned` — a `PUNKTFUNK_CAPTURE_MONITOR` pin routes [`open`] to the mirror backend,
/// whose gamescope arm attaches to the node the RUNNING session already publishes without
/// consulting the sub-mode ladder at all. On a Bazzite/SteamOS box that session is Game Mode's,
/// i.e. by definition not ours.
/// * a [`GamescopeRoute::Attach`] verdict — however the ladder reached it (operator override,
/// or the foreign-gamescope rung).
///
/// [`GamescopeRoute::Managed`] is NOT foreign: the managed takeover starts the session through our
/// own `GAMESCOPE_BIN` wrapper / PATH shim, so its flags are the ones we chose.
///
/// `mirror_pinned` is judged from the pin alone, not from whether the mirror actually took: [`open`]
/// degrades a pin to the virtual-display path when the session reports no physical heads, and a
/// pinned box that lands there is called foreign here although it will bare-spawn. That is the
/// fail-closed direction — a capability withheld from a session that could have had it — and the
/// alternative (enumerating heads from a capability query) would put a compositor roundtrip on a
/// path that must answer before anything exists to ask.
fn session_is_a_foreign_gamescope(mirror_pinned: bool, route: Option<&GamescopeRoute>) -> bool {
mirror_pinned || matches!(route, Some(GamescopeRoute::Attach { .. }))
}
// Platform-neutral per-client stable display-id map: Windows seeds the monitor EDID serial +
// IddCx ConnectorIndex from the id; KWin names its output `Virtual-punktfunk-<id>` (kwin.rs's
// `resolve_slot` call); Mutter cannot carry the id into its virtual monitor at all, so it keys the
// host-persisted `ScaleMap` on the same identity key. All three are production call sites, so the
// `allow(dead_code)` below no longer stands for "unwired yet" (it did when only Windows consumed the
// map); it now covers whatever helpers no CURRENT backend reaches. Worth re-testing without it —
// that has to happen on a Linux build, since this is the platform where dead_code is enforced.
#[allow(dead_code)]
#[path = "vdisplay/identity.rs"]
pub(crate) mod identity;
@@ -735,6 +837,48 @@ mod tests {
assert_eq!(compositor_for_kind(ActiveKind::None), None);
}
/// The spawn-vs-attach term behind [`gamescope_hdr_available`] /
/// [`gamescope_composites_cursor`]. Both answers are IRREVOCABLE once the punktfunk/1 Welcome
/// has gone out (bit depth is fixed there; the session plan's cursor decision feeds the encoder
/// open), so an over-promise here is not recoverable at runtime — which is why the regression
/// this pins mattered: the term used to be read off `PUNKTFUNK_GAMESCOPE_NODE`, a key nothing
/// writes any more, so every foreign session answered "ours".
#[test]
fn only_a_session_we_start_can_promise_gamescope_capabilities() {
// Attach — however the ladder got there — is somebody else's session: unknown spawn flags.
assert!(session_is_a_foreign_gamescope(
false,
Some(&GamescopeRoute::Attach {
node: "auto".into()
})
));
assert!(session_is_a_foreign_gamescope(
false,
Some(&GamescopeRoute::Attach { node: "42".into() })
));
// A bare spawn is ours by definition; so is the managed takeover (it starts gamescope
// through our own GAMESCOPE_BIN wrapper / PATH shim, so the flags are the ones we chose).
assert!(!session_is_a_foreign_gamescope(
false,
Some(&GamescopeRoute::Spawn)
));
assert!(!session_is_a_foreign_gamescope(
false,
Some(&GamescopeRoute::Managed {
client: "steam".into()
})
));
// No route at all = not a gamescope session; the binary probe alone then decides.
assert!(!session_is_a_foreign_gamescope(false, None));
// A monitor pin bypasses the ladder entirely (mirror backend → attach to the node the
// RUNNING session publishes), so it is foreign whatever the ladder would have said.
assert!(session_is_a_foreign_gamescope(true, None));
assert!(session_is_a_foreign_gamescope(
true,
Some(&GamescopeRoute::Spawn)
));
}
#[test]
fn detect_active_session_is_side_effect_free_and_terminates() {
// A pure probe of /proc + the runtime dir: it must not panic and must return promptly on
+20 -1
View File
@@ -136,7 +136,26 @@ pub fn admit(req_identity: Option<[u8; 32]>) -> Admission {
!live.is_empty(),
)
};
let _ = any_live; // read only by the Windows budget block below
let _ = any_live; // read only by the budget blocks below
// The operator's `max_displays` ceiling (design §5.3). Applied HERE, once per connecting
// session, and deliberately NOT in the display create path: `acquire` runs again on every
// mid-stream rebuild (capture loss, a Game↔Desktop switch), and those rebuild before dropping
// the old display — so a ceiling enforced there counts the session against itself and refuses
// the recovery. Admission is reached once per connect, so it cannot.
#[cfg(target_os = "linux")]
if matches!(decision, Admission::Separate) && any_live {
// The Linux pool had no ceiling at all: its reuse key includes the CLIENT-SUPPLIED mode, so
// a client reconnecting at a different resolution misses reuse and mints a fresh display,
// and a handful of reconnects could row out an unbounded number of compositor outputs.
let max = policy::prefs().get().effective().max_displays;
let live = super::registry::live_display_count();
if live >= max {
return Admission::Reject(format!(
"host display budget exhausted: {live} display(s) live/kept, max_displays = {max}"
));
}
}
#[cfg(windows)]
if matches!(decision, Admission::Separate) && any_live {
let max = policy::prefs().get().effective().max_displays;
+14 -3
View File
@@ -225,9 +225,20 @@ pub trait VirtualDisplay: Send {
/// ([`DisplayOwnership::Owned`], keep-alive-able) display? The registry consults this **before**
/// its keep-alive reuse lookup, so it never hands a kept display of one flavor to a request of
/// another — specifically a gamescope managed/attach acquire must not reuse a kept **bare-spawn**
/// (they share the backend name `"gamescope"`). Default `true`; only gamescope overrides it,
/// returning `false` when the env selects attach/managed (consistent with the `ownership` its
/// `create` will report). See `design/gamemode-and-dedicated-sessions.md` A1.
/// (they share the backend name `"gamescope"`). Overridden by gamescope, which reads the
/// resolved [`GamescopeRoute`](crate::GamescopeRoute) carried on the instance (`self.route`, NOT
/// env — the sub-mode stopped travelling through `PUNKTFUNK_GAMESCOPE_NODE`/`_SESSION` in Phase
/// 2.3): `false` for `Managed` and `Attach`, `true` for `Spawn` **and for no route at all**,
/// since `create`'s own `None` arm falls through to the bare spawn — so an instance nobody
/// called `set_gamescope_route` on (the operator-pinned `PUNKTFUNK_COMPOSITOR` path) is
/// poolable, and takes both the reuse lookup and the `max_displays` ceiling. Also overridden by
/// the mirror backend (`false` always). See `design/gamemode-and-dedicated-sessions.md` A1.
///
/// The default `true` is a DEFAULT, not a fact: it happens to be right for every backend that
/// creates a display it owns, and it is wrong for any backend whose `create` reports something
/// other than [`DisplayOwnership::Owned`] — this answer and that one must agree, and nothing
/// enforces it. A required method would; making it one costs an impl in each of the five
/// per-compositor backends plus Windows.
fn poolable_now(&self) -> bool {
true
}
+176 -36
View File
@@ -21,6 +21,7 @@
//! Persisted to `<config>/display-identity.json` (migrated from the legacy Windows
//! `pf-vdisplay-identity.json`) so ids — and the client→config association — survive host restarts.
use std::collections::BTreeSet;
use std::path::PathBuf;
use std::sync::{Mutex, OnceLock};
@@ -78,12 +79,38 @@ impl DisplayIdentityMap {
pub(crate) fn load() -> Self {
let dir = pf_paths::config_dir();
let path = dir.join(FILE);
let bytes = std::fs::read(&path)
.or_else(|_| std::fs::read(dir.join(LEGACY_FILE)))
.ok();
let mut store = bytes
.and_then(|b| serde_json::from_slice::<Store>(&b).ok())
.unwrap_or_default();
let (from, bytes) = match std::fs::read(&path) {
Ok(b) => (path.clone(), Some(b)),
Err(_) => {
let legacy = dir.join(LEGACY_FILE);
match std::fs::read(&legacy) {
Ok(b) => (legacy, Some(b)),
// No file at all is the ordinary first-run case — not worth a word.
Err(_) => (path.clone(), None),
}
}
};
let mut store = match bytes {
Some(b) => match serde_json::from_slice::<Store>(&b) {
Ok(s) => s,
Err(e) => {
// An UNPARSEABLE map used to be swallowed into `Default::default()`, and the very
// next `resolve` persisted that empty store OVER the file — silently discarding
// every client's Windows EDID serial / KWin `Virtual-punktfunk-<id>` and the
// per-display DPI the OS keyed to them. Say so, and move the file aside so the
// damage is recoverable by hand (same treatment `display-presets.json` gets).
tracing::warn!(
path = %from.display(),
error = %e,
"display-identity map is unreadable — starting a fresh one; \
the old file is kept as .bad (every client re-derives its display id once)"
);
let _ = std::fs::rename(&from, from.with_extension("json.bad"));
Store::default()
}
},
None => Store::default(),
};
// SANITIZE a hand-edited / corrupt / cross-version file before trusting it: resolve()'s
// found-entry branch returns the stored id verbatim, so an out-of-range id (0 = the "auto"
// sentinel, or > MAX_ID) or a duplicate id/key would flow straight into the display identity.
@@ -100,7 +127,17 @@ impl DisplayIdentityMap {
/// The stable id (`1..=15`) for the client `key` ([`identity_key`]): its remembered id, or a
/// freshly assigned one (lowest free, else LRU-evict at the cap). Bumps the entry to MRU and persists.
pub(crate) fn resolve(&mut self, key: &str) -> u32 {
///
/// `live` is the set of ids that currently drive a REAL display (the Windows manager's slot keys
/// / the Linux pool's `identity_slot`s). An id in it is never evicted, and when every eviction
/// candidate is live this **refuses** (`None`) rather than handing the newcomer an id that is
/// already someone else's monitor. That is not hypothetical: the id keys the Windows manager's
/// slot map, whose plain-JOIN branch attaches an arriving session to whatever monitor the slot
/// already holds — so evicting a live id handed client B client A's streaming monitor, capture
/// target and all. Refusing costs the newcomer its stable identity (upstream falls back to the
/// shared/auto slot: `resolve_slot` → `None`, `slot_id_for` → `0`); evicting cost a live client
/// its session.
pub(crate) fn resolve(&mut self, key: &str, live: &BTreeSet<u32>) -> Option<u32> {
self.store.tick = self.store.tick.wrapping_add(1);
let now = self.store.tick;
@@ -108,32 +145,43 @@ impl DisplayIdentityMap {
e.seen = now;
let id = e.id;
self.persist();
return id;
return Some(id);
}
// New client: prefer the lowest free id in 1..=MAX_ID; if all are taken, evict the LRU entry and
// reuse its id (the evicted client re-establishes its scaling once on its next connect).
let id = (1..=MAX_ID)
.find(|i| !self.store.entries.iter().any(|e| e.id == *i))
.unwrap_or_else(|| {
// New client: prefer the lowest free id in 1..=MAX_ID; if all are taken, evict the
// least-recently-seen entry that is NOT live and reuse its id (that client re-establishes its
// scaling once on its next connect).
let id = match (1..=MAX_ID).find(|i| !self.store.entries.iter().any(|e| e.id == *i)) {
Some(free) => free,
None => {
let lru = self
.store
.entries
.iter()
.enumerate()
.filter(|(_, e)| !live.contains(&e.id))
.min_by_key(|(_, e)| e.seen)
.map(|(i, _)| i)
.expect("entries are non-empty whenever every id 1..=MAX_ID is taken");
let evicted = self.store.entries.remove(lru);
evicted.id
});
.map(|(i, _)| i);
let Some(lru) = lru else {
tracing::warn!(
cap = MAX_ID,
live = live.len(),
"display identity map is full and every id is driving a live display — \
this client gets the shared/auto display identity (no persisted per-client \
scaling) rather than displacing a live one"
);
return None;
};
self.store.entries.remove(lru).id
}
};
self.store.entries.push(Entry {
key: key.to_string(),
id,
seen: now,
});
self.persist();
id
Some(id)
}
/// Persist atomically (temp file + rename). Best-effort: a write failure just means a restart may
@@ -168,7 +216,8 @@ pub(crate) fn global() -> &'static Mutex<DisplayIdentityMap> {
/// Resolve the connecting client's stable slot id per the `identity` policy. When no policy is
/// configured, `default` applies — **PerClient on Windows / Shared on Linux**, preserving each
/// platform's historical behavior (Windows always keyed monitors per-client; Linux used one shared
/// output name). `None` ⇒ shared / anonymous the backend uses its base name / auto slot.
/// output name). `None` ⇒ shared / anonymous (or the map [refused](DisplayIdentityMap::resolve) an
/// id because every one is live) → the backend uses its base name / auto slot.
pub(crate) fn resolve_slot(
fp: Option<[u8; 32]>,
mode: (u32, u32),
@@ -185,12 +234,40 @@ pub(crate) fn resolve_slot(
Identity::PerClientMode => true,
};
let fp = fp?;
Some(
global()
.lock()
.unwrap()
.resolve(&identity_key(fp, mode, per_client_mode)),
)
// Sample the live ids BEFORE taking the map lock, never under it: the sources below take the
// Windows manager's `state` lock / the Linux pool lock, and this map is reached from inside a
// backend `create` — a lock order of (display owner → identity map) in both directions would be
// a deadlock. One direction only, and the map lock stays a leaf.
let live = live_slot_ids();
global()
.lock()
.unwrap()
.resolve(&identity_key(fp, mode, per_client_mode), &live)
}
/// The identity slots currently driving a REAL display — the eviction guard for
/// [`DisplayIdentityMap::resolve`]. Windows reads the manager's slot map (the key IS the identity
/// slot); Linux reads the registry pool's per-entry `identity_slot`. Both include KEPT
/// (lingering/pinned) displays on purpose: a kept display is a live compositor/driver resource whose
/// owner is expected back, and the whole point of the id is that the reconnect finds it again.
/// Anonymous (`0`) is not an identity and never blocks an assignment.
fn live_slot_ids() -> BTreeSet<u32> {
#[cfg(target_os = "windows")]
{
crate::manager::snapshot()
.into_iter()
.map(|i| i.slot_id)
.filter(|s| *s != 0)
.collect()
}
#[cfg(target_os = "linux")]
{
crate::registry::live_identity_slots()
}
#[cfg(not(any(target_os = "windows", target_os = "linux")))]
{
BTreeSet::new()
}
}
// ---------------------------------------------------------------------------------------
@@ -306,24 +383,31 @@ mod tests {
}
}
/// Nothing is streaming — the ordinary case, where the live set never constrains anything.
fn nothing_live() -> BTreeSet<u32> {
BTreeSet::new()
}
#[test]
fn stable_across_calls_and_distinct_per_client() {
let mut m = temp_map("stable");
let a1 = m.resolve(&identity_key(fp(1), (1920, 1080), false));
let b = m.resolve(&identity_key(fp(2), (1920, 1080), false));
let a2 = m.resolve(&identity_key(fp(1), (1280, 720), false)); // per-client: mode ignored
let a1 = m.resolve(&identity_key(fp(1), (1920, 1080), false), &nothing_live());
let b = m.resolve(&identity_key(fp(2), (1920, 1080), false), &nothing_live());
// per-client: mode ignored
let a2 = m.resolve(&identity_key(fp(1), (1280, 720), false), &nothing_live());
assert_eq!(a1, a2, "same client → same id (per-client ignores mode)");
assert_ne!(a1, b, "distinct clients → distinct ids");
assert!((1..=MAX_ID).contains(&a1) && (1..=MAX_ID).contains(&b));
assert!(a1.is_some_and(|i| (1..=MAX_ID).contains(&i)));
assert!(b.is_some_and(|i| (1..=MAX_ID).contains(&i)));
let _ = std::fs::remove_file(&m.path);
}
#[test]
fn per_client_mode_splits_by_resolution() {
let mut m = temp_map("permode");
let hd = m.resolve(&identity_key(fp(1), (1920, 1080), true));
let uhd = m.resolve(&identity_key(fp(1), (3840, 2160), true));
let hd2 = m.resolve(&identity_key(fp(1), (1920, 1080), true));
let hd = m.resolve(&identity_key(fp(1), (1920, 1080), true), &nothing_live());
let uhd = m.resolve(&identity_key(fp(1), (3840, 2160), true), &nothing_live());
let hd2 = m.resolve(&identity_key(fp(1), (1920, 1080), true), &nothing_live());
assert_ne!(hd, uhd, "same client, different resolution → different id");
assert_eq!(hd, hd2, "same client + resolution → same id");
let _ = std::fs::remove_file(&m.path);
@@ -333,16 +417,72 @@ mod tests {
fn lru_eviction_reuses_an_id_at_the_cap() {
let mut m = temp_map("lru");
for n in 1..=15u8 {
m.resolve(&identity_key(fp(n), (1920, 1080), false));
m.resolve(&identity_key(fp(n), (1920, 1080), false), &nothing_live());
}
let _ = m.resolve(&identity_key(fp(2), (1920, 1080), false)); // touch 2 so 1 is LRU
let id16 = m.resolve(&identity_key(fp(16), (1920, 1080), false));
// touch 2 so 1 is LRU
let _ = m.resolve(&identity_key(fp(2), (1920, 1080), false), &nothing_live());
let id16 = m
.resolve(&identity_key(fp(16), (1920, 1080), false), &nothing_live())
.expect("nothing is live → the LRU id is free to take");
assert!((1..=MAX_ID).contains(&id16));
assert_eq!(m.store.entries.len(), 15, "cap holds at 15 entries");
assert!(m.store.entries.iter().all(|e| (1..=MAX_ID).contains(&e.id)));
let _ = std::fs::remove_file(&m.path);
}
/// 10.2: the LRU victim is chosen among ids that are NOT driving a display. Handing the LRU id
/// to a newcomer while its owner streams is what let the Windows manager's plain-JOIN branch
/// attach the newcomer to the live client's monitor.
#[test]
fn lru_eviction_never_takes_a_live_id() {
let mut m = temp_map("lru-live");
let mut ids = Vec::new();
for n in 1..=15u8 {
ids.push(
m.resolve(&identity_key(fp(n), (1920, 1080), false), &nothing_live())
.unwrap(),
);
}
// fp(1) is the least-recently-seen — and it is the one that is streaming.
let lru_id = ids[0];
let live: BTreeSet<u32> = [lru_id].into_iter().collect();
let id16 = m
.resolve(&identity_key(fp(16), (1920, 1080), false), &live)
.expect("14 idle ids remain — one of them is the victim");
assert_ne!(id16, lru_id, "must not take the id of a live display");
assert_eq!(id16, ids[1], "the next-least-recently-seen IDLE id instead");
// The live client's mapping is untouched, so its reconnect still finds its own display.
assert_eq!(
m.resolve(&identity_key(fp(1), (1920, 1080), false), &live),
Some(lru_id)
);
let _ = std::fs::remove_file(&m.path);
}
/// Fail-closed at the extreme: every id live ⇒ refuse, rather than displace a streaming client.
/// The caller degrades to the shared/auto identity (`resolve_slot` → `None`, `slot_id_for` → 0).
#[test]
fn refuses_rather_than_evicting_when_every_id_is_live() {
let mut m = temp_map("lru-all-live");
let mut live = BTreeSet::new();
for n in 1..=15u8 {
live.insert(
m.resolve(&identity_key(fp(n), (1920, 1080), false), &BTreeSet::new())
.unwrap(),
);
}
assert_eq!(
m.resolve(&identity_key(fp(16), (1920, 1080), false), &live),
None
);
assert_eq!(m.store.entries.len(), 15, "nothing was evicted");
// A KNOWN client is still resolved even when everything is live — it owns that id already.
assert!(m
.resolve(&identity_key(fp(3), (1920, 1080), false), &live)
.is_some());
let _ = std::fs::remove_file(&m.path);
}
#[test]
fn key_composition() {
assert_eq!(identity_key(fp(0xab), (1920, 1080), false).len(), 64); // hex fp only
+228 -22
View File
@@ -10,8 +10,16 @@
//! deterministic.
//! * **manual** — per-identity-slot offsets from [`Layout::positions`] (console-arranged): a member
//! whose stable identity slot has a stored position sits there; a member with no pin (no stored
//! position, or a shared/anonymous identity that has no slot) falls back to its auto-row origin, so
//! a half-arranged group never collapses everything onto the origin.
//! position, or a shared/anonymous identity that has no slot) is **packed clear of the pins** —
//! rowed left-to-right starting past the rightmost pinned edge — so a half-arranged group neither
//! collapses everything onto the origin nor drops an unpinned display exactly on top of a pinned
//! one. The pins themselves are reproduced verbatim: where two of them overlap, that is the
//! operator's own arrangement and not ours to second-guess.
//!
//! Members carry no height, so "clear of the pins" is decided on the x axis alone and every pin
//! counts regardless of its `y` — a vertically-stacked arrangement therefore packs further right than
//! it strictly needs to. That is the conservative direction: a gap is a cosmetic waste of desktop
//! coordinate space, an overlap is two desktops fighting over the same pixels.
//!
//! Group membership + acquire order live in the registry ([`super::registry`]); this file only turns
//! that ordered member list into positions.
@@ -24,8 +32,18 @@ pub struct Member {
/// Stable per-client identity slot — the manual-layout key. `None` for a shared/anonymous
/// identity (no per-client slot), which can't carry a manual pin and therefore always auto-rows.
pub identity_slot: Option<u32>,
/// Pixel width, for auto-row `x` accumulation. Clamped at 0 (a bogus negative never shifts a
/// sibling left).
/// The member's width **in the same coordinate space the resulting [`Placement`] is expressed
/// in**, for row `x` accumulation. Clamped at 0 (a bogus negative never shifts a sibling left).
///
/// ⚠ Every fill site currently uses the requested *mode* width, i.e. pixels. On Windows
/// that is also the desktop space (CCD geometry is pixels), so the two agree; on KWin the
/// placement is handed to `config.position()`, which is the compositor's **logical** space — the
/// two coincide only at scale 1.0, and a per-output scale is exactly what the identity machinery
/// exists to make KDE reapply. A 150 %-scaled 2560-wide output occupies 1707 logical px, so
/// auto-rowing past it by 2560 leaves an 853-px dead band. Fixing that means dividing by the
/// output's applied scale at the KWin fill site (`kwin_output_mgmt` already reads `scale` into
/// its device state); this type stays unit-agnostic, and the contract is that whoever fills it
/// speaks the consumer's space.
pub width: i32,
}
@@ -37,30 +55,79 @@ pub struct Placement {
}
/// The auto-row origin of member `i`: the summed width of every prior member, top-aligned.
/// `saturating_add` because the widths are client-supplied through the requested mode — an absurd
/// one must produce an absurd coordinate, not a debug-build panic inside the state readout.
fn auto_row_x(members: &[Member], i: usize) -> i32 {
members[..i].iter().map(|m| m.width.max(0)).sum()
members[..i]
.iter()
.fold(0i32, |x, m| x.saturating_add(m.width.max(0)))
}
/// The manual pin for `m`, if its identity slot carries one. The lookup is an exact string match on
/// the canonical decimal slot id — `DisplayPolicy::sanitized` re-keys the table to that form on
/// write, so a `"01"` typed into a hand-edited settings file still resolves here.
fn pin_of(m: &Member, layout: &Layout) -> Option<Placement> {
m.identity_slot
.and_then(|slot| layout.positions.get(&slot.to_string()))
.map(|p| Placement { x: p.x, y: p.y })
}
/// Arrange `members` (in acquire order) per `layout`, returning one [`Placement`] per member in the
/// same order. Pure — the single source of truth for auto-row / manual placement, shared by the
/// state readout and (KWin) the per-backend position apply.
pub fn arrange(members: &[Member], layout: &Layout) -> Vec<Placement> {
members
.iter()
.enumerate()
.map(|(i, m)| {
let auto = Placement {
match layout.mode {
LayoutMode::AutoRow => (0..members.len())
.map(|i| Placement {
x: auto_row_x(members, i),
y: 0,
};
match layout.mode {
LayoutMode::AutoRow => auto,
// A pinned member sits at its stored offset; an unpinned one falls back to auto-row.
LayoutMode::Manual => m
.identity_slot
.and_then(|slot| layout.positions.get(&slot.to_string()))
.map(|p| Placement { x: p.x, y: p.y })
.unwrap_or(auto),
})
.collect(),
LayoutMode::Manual => arrange_manual(members, layout),
}
}
/// Manual placement: pins verbatim, everything else rowed out past them.
///
/// The unpinned fallback used to be the unconditional auto-row prefix sum — computed as if the pins
/// did not exist — so an unpinned display could land exactly on top of a pinned sibling with nothing
/// downstream noticing (the arrangement is only ever *reported* and *applied*, never validated). One
/// number in this crate's own fixture separated the tested case from that collision. Rowing the
/// unpinned members from the rightmost pinned edge instead makes the overlap unrepresentable within
/// one call, and keeps three of the fallback's properties: deterministic, acquire-ordered, and
/// identical to plain auto-row when nothing is pinned.
///
/// ⚠ **The fourth property is gone, knowingly: incremental stability.** The prefix sum could not
/// move member `i` when member `i+1` joined; this cursor is seeded from the pins of *all* members,
/// so an already-placed unpinned member's computed `x` shifts the moment a pinned sibling arrives
/// later in acquire order. Nothing re-applies it — `registry::position_for_new` takes only the
/// `.last()` placement and the registry moves the newly-acquired display alone — so in that ordering
/// `GET /display/state` reports a position the desktop never received (the pre-existing shape of
/// this: an auto-row teardown already shifts every survivor's reported `x` with no re-apply; the
/// packing widens the class to joins under `Manual`). It is not fixable here: the honest fix is for
/// the registry to re-apply the WHOLE group's arrangement on any membership change under
/// `LayoutMode::Manual`, the way `windows/manager.rs`'s `arrange_slots` already does, at which point
/// this function is right in every ordering. Seeding the cursor from preceding pins only would buy
/// incremental stability back by reintroducing the collision this exists to prevent — the wrong
/// trade, since the common ordering (the pin exists, an unpinned client joins) does reach the apply
/// path and is placed correctly.
fn arrange_manual(members: &[Member], layout: &Layout) -> Vec<Placement> {
let pins: Vec<Option<Placement>> = members.iter().map(|m| pin_of(m, layout)).collect();
// Start the unpinned row at the desktop origin, or past the rightmost pinned edge when there is
// one. `max(0)` on the width keeps a bogus negative from pulling the cursor back over a pin.
let mut cursor = pins
.iter()
.zip(members)
.filter_map(|(pin, m)| pin.map(|p| p.x.saturating_add(m.width.max(0))))
.fold(0i32, i32::max);
pins.iter()
.zip(members)
.map(|(pin, m)| match pin {
Some(p) => *p,
None => {
let at = Placement { x: cursor, y: 0 };
cursor = cursor.saturating_add(m.width.max(0));
at
}
})
.collect()
@@ -115,14 +182,153 @@ mod tests {
}
#[test]
fn manual_unpinned_and_slotless_fall_back_to_auto_row() {
fn manual_unpinned_and_slotless_pack_clear_of_the_pins() {
let members = [m(Some(1), 2560), m(Some(9), 1920), m(None, 1280)];
// Only slot 1 is pinned; slot 9 has no stored pin; the third has no slot at all.
let layout = manual(&[("1", 100, 50)]);
let out = arrange(&members, &layout);
assert_eq!(out[0], Placement { x: 100, y: 50 }, "pinned");
assert_eq!(out[1], Placement { x: 2560, y: 0 }, "unpinned → auto-row");
assert_eq!(out[2], Placement { x: 4480, y: 0 }, "slotless → auto-row");
// The pin occupies [100, 2660); the unpinned members row out from its right edge in acquire
// order, NOT from the pin-blind prefix sum (which would have put the first one at 2560 —
// inside the pin).
assert_eq!(
out[1],
Placement { x: 2660, y: 0 },
"unpinned → past the pin"
);
assert_eq!(out[2], Placement { x: 4580, y: 0 }, "slotless → past both");
}
#[test]
fn manual_with_no_pins_at_all_is_plain_auto_row() {
// The fallback must not drift from auto-row when the manual table happens to be empty (the
// state a group is in the instant `manual` is selected and nothing has been arranged yet).
let members = [m(Some(1), 2560), m(Some(2), 1920), m(None, 1280)];
let out = arrange(&members, &manual(&[]));
assert_eq!(out, arrange(&members, &Layout::default()));
}
#[test]
fn a_manual_pin_that_would_collide_with_an_auto_row_sibling_is_packed_clear() {
// The exact geometry §13 11.8 names: a pin sitting where the pin-blind auto-row would have
// put the unpinned sibling. Two displays on one origin = two desktops on the same pixels.
let members = [m(Some(1), 2560), m(Some(9), 1920)];
let layout = manual(&[("1", 2560, 0)]);
let out = arrange(&members, &layout);
assert_eq!(out[0], Placement { x: 2560, y: 0 }, "pin honored verbatim");
assert_ne!(
out[1], out[0],
"the unpinned sibling must not land on the pin"
);
assert_eq!(
out[1],
Placement { x: 5120, y: 0 },
"past the pin's right edge"
);
}
#[test]
fn a_pin_left_of_the_origin_still_leaves_the_unpinned_row_at_zero() {
// A negative pin is legal (KWin's global space extends left of 0). Its right edge is what
// matters: at -3000+2560 = -440 it constrains nothing, so the row still starts at the origin.
let members = [m(Some(1), 2560), m(Some(9), 1920)];
let out = arrange(&members, &manual(&[("1", -3000, 0)]));
assert_eq!(out[0], Placement { x: -3000, y: 0 });
assert_eq!(out[1], Placement { x: 0, y: 0 });
}
#[test]
fn absurd_widths_saturate_instead_of_panicking() {
// Widths originate in the client-requested mode; a hostile or corrupt one must produce an
// absurd coordinate, not an overflow panic inside the `/display/state` readout.
let members = [m(Some(1), i32::MAX), m(Some(2), i32::MAX), m(None, 4096)];
let out = arrange(&members, &Layout::default());
assert_eq!(out[2], Placement { x: i32::MAX, y: 0 });
let out = arrange(&members, &manual(&[("1", i32::MAX, 0)]));
assert_eq!(out[2], Placement { x: i32::MAX, y: 0 });
}
/// Property (deterministic seeded walk): across arbitrary member widths, slot assignments and pin
/// tables, **no unpinned member may share desktop space with any sibling**. Overlap between two
/// *pins* is excluded from the invariant — that is the operator's own arrangement, faithfully
/// reproduced. Members carry no height, so "share space" is decided on the x interval alone,
/// which is the strictest reading available here.
#[test]
fn no_unpinned_member_overlaps_a_sibling_under_any_layout() {
// Tiny deterministic LCG (Numerical Recipes) — reproducible, no dependency. Same shape as
// `lifecycle`'s property walk.
let mut rng: u64 = 0x0bad_f00d_dead_beef;
let mut next = || {
rng = rng
.wrapping_mul(6364136223846793005)
.wrapping_add(1442695040888963407);
(rng >> 33) as u32
};
for _ in 0..20_000 {
let count = (next() % 6) as usize;
let members: Vec<Member> = (0..count)
.map(|_| {
// A slot only sometimes, and from a small pool so collisions with the pin table
// are frequent; widths include 0 and the odd negative.
let slot = match next() % 4 {
0 => None,
_ => Some(next() % 6 + 1),
};
let width = match next() % 8 {
0 => 0,
1 => -((next() % 4000) as i32),
_ => (next() % 4000) as i32,
};
m(slot, width)
})
.collect();
let mut pairs: Vec<(String, i32, i32)> = Vec::new();
for slot in 1..=6u32 {
if next() % 2 == 0 {
let x = (next() % 8000) as i32 - 2000;
let y = ((next() % 3) * 1440) as i32;
pairs.push((slot.to_string(), x, y));
}
}
let borrowed: Vec<(&str, i32, i32)> =
pairs.iter().map(|(k, x, y)| (k.as_str(), *x, *y)).collect();
for layout in [Layout::default(), manual(&borrowed)] {
let out = arrange(&members, &layout);
assert_eq!(out.len(), members.len());
let pinned: Vec<bool> = members
.iter()
.map(|mem| pin_of(mem, &layout).is_some())
.collect();
for i in 0..out.len() {
for j in (i + 1)..out.len() {
if pinned[i] && pinned[j] {
continue; // the operator's own arrangement
}
let span = |k: usize| {
let x = out[k].x as i64;
(x, x + members[k].width.max(0) as i64)
};
let (ai, bi) = span(i);
let (aj, bj) = span(j);
// Empty spans (a zero/negative width) can't collide with anything.
if ai >= bi || aj >= bj {
continue;
}
assert!(
bi <= aj || bj <= ai,
"members {i} {:?} and {j} {:?} overlap under {layout:?} \
(widths {} / {})",
out[i],
out[j],
members[i].width,
members[j].width
);
}
}
}
}
}
#[test]
File diff suppressed because it is too large Load Diff
@@ -5,10 +5,18 @@
use super::*;
/// Wait for gamescope to report its PipeWire node. Authoritative source: gamescope's own log
/// line `stream available on node ID: N` (its node carries `node.name=gamescope` on TWO objects
/// — the adapter and the inner stream — and only the advertised id is the correct capture
/// target). Falls back to `pw-dump` discovery if the log line doesn't show.
/// Budget for a `pw-dump` snapshot. Two facts make an unbounded one the worst call in this file:
/// it is polled every 300500 ms from three separate 45 s loops, and it talks to the very daemon
/// this module documents gamescope as head-blocking below [`MIN_GAMESCOPE`] — so the failure mode
/// is not "slow", it is "never returns", on the session's own stream thread. Two seconds is far
/// above a populated graph's real cost; every caller already has a "couldn't ask" path.
const PW_DUMP_BUDGET: Duration = Duration::from_secs(2);
/// Budget for a `gamescope --version` probe. It loads the binary and prints a banner — no Vulkan
/// device, no daemon — so anything approaching this bound is a binary that cannot run at all,
/// which is exactly what a `None`/`false` answer means to each caller.
const VERSION_PROBE_BUDGET: Duration = Duration::from_secs(2);
/// B2 (game-exit detection): confirm a **dedicated** gamescope session's game has exited. gamescope is
/// a single-app compositor — it exits when its nested app exits — so once capture is lost, THIS
/// session's `node_id` not reappearing within a short confirmation window means the game quit (vs. a
@@ -159,16 +167,51 @@ pub(super) fn poll_managed_node(timeout: Duration) -> Option<u32> {
}
}
/// Wait for a freshly spawned gamescope to report its PipeWire node. Authoritative source:
/// gamescope's own log line `stream available on node ID: N` (its node carries
/// `node.name=gamescope` on TWO objects — the adapter and the inner stream — and only the
/// advertised id is the correct capture target). Falls back, at the deadline, to `pw-dump`
/// discovery SCOPED to this spawn's process tree (`child`'s pid, A5), so a coexisting gamescope's
/// node is never mistaken for ours.
///
/// Takes the `Child` rather than a bare pid so it can **stop early when gamescope is already
/// dead**. A gamescope that fails `vkCreateDevice` exits in under a second, and polling its corpse
/// for the full 15 s bought nothing except a caller error that blamed the wrong thing ("headless
/// capture is unsupported on this GPU/driver"). `try_wait` turns that into an immediate `None`
/// while the log — which the caller names in the same error — still holds the real reason.
pub(super) fn wait_for_node(
timeout: Duration,
log: &std::path::Path,
child_pid: u32,
child: &mut Child,
) -> Option<u32> {
let child_pid = child.id();
let deadline = Instant::now() + timeout;
loop {
if let Some(id) = node_from_log(log) {
return Some(id);
}
// Check for a node FIRST, then for death: a gamescope that published its node and then
// exited in the same tick still gives us the id, and the caller's own liveness handling
// (the keepalive `Child`, `kept_display_alive`) owns what happens next.
match child.try_wait() {
// Still running — keep waiting.
Ok(None) => {}
// Exited. One last scoped look (the node line may have been written between the two
// reads above), then give up rather than poll a corpse to the deadline.
Ok(Some(status)) => {
tracing::warn!(
pid = child_pid,
%status,
log = %log.display(),
"gamescope: the spawned process exited before publishing a PipeWire node — \
not waiting out the rest of the budget"
);
return node_from_log(log).or_else(|| find_gamescope_node_scoped(Some(child_pid)));
}
// `try_wait` itself failed (the child was reaped elsewhere, ECHILD): fall back to the
// old behaviour rather than inventing a death.
Err(_) => {}
}
if Instant::now() >= deadline {
// Last-resort fallback scoped to THIS spawn's process tree (A5), so a coexisting gamescope's
// node isn't picked by mistake.
@@ -197,7 +240,10 @@ fn node_from_log(log: &std::path::Path) -> Option<u32> {
/// keep-alive reuse liveness probe ([`GamescopeDisplay::kept_display_alive`]): a kept gamescope node
/// vanishes when its nested game exits, so a missing id means "recreate, don't reuse the corpse".
pub(super) fn gamescope_node_present(node_id: u32) -> bool {
let Ok(out) = Command::new("pw-dump").arg(node_id.to_string()).output() else {
let Ok(out) = crate::proc::output_within(
Command::new("pw-dump").arg(node_id.to_string()),
PW_DUMP_BUDGET,
) else {
// pw-dump unavailable → don't block reuse (mark_failed is the backstop on a genuinely dead node).
return true;
};
@@ -229,7 +275,7 @@ pub(super) fn find_gamescope_node() -> Option<u32> {
/// belong to OUR gamescope's process tree, so a coexisting foreign / other-session gamescope node is
/// never mistaken for ours). `None` = any gamescope node (the managed/attach paths, single-session).
fn find_gamescope_node_scoped(scope: Option<u32>) -> Option<u32> {
let out = Command::new("pw-dump").output().ok()?;
let out = crate::proc::output_within(&mut Command::new("pw-dump"), PW_DUMP_BUDGET).ok()?;
let dump: serde_json::Value = serde_json::from_slice(&out.stdout).ok()?;
let nodes = dump.as_array()?;
let node_props = |obj: &serde_json::Value| -> Option<(u32, String, String, Option<u32>)> {
@@ -302,7 +348,12 @@ fn find_gamescope_node_scoped(scope: Option<u32>) -> Option<u32> {
/// most recently created (the live session). Returns the bare socket *name* (the injector
/// resolves it against `XDG_RUNTIME_DIR`, matching libei's own `LIBEI_SOCKET` semantics).
pub(super) fn find_gamescope_eis_socket() -> Option<String> {
let runtime = std::env::var("XDG_RUNTIME_DIR").ok()?;
// Under the shared env lock: `session::apply_session_env` `set_var`s XDG_RUNTIME_DIR from the
// connect thread, and glibc's setenv/getenv pair is a data race the crate's own `lib.rs`
// documents as UB. The lock is not reentrant, so this must stay a read taken HERE and not
// hoisted into a caller — the only caller, `point_injector_at_eis`, holds nothing (its
// `ei_socket_file()` takes and releases the same lock separately).
let runtime = crate::with_env_lock(|| std::env::var("XDG_RUNTIME_DIR").ok())?;
let mut live: Vec<(std::time::SystemTime, String)> = Vec::new();
for entry in std::fs::read_dir(&runtime).ok()?.flatten() {
let name = entry.file_name().to_string_lossy().into_owned();
@@ -328,11 +379,12 @@ pub(super) fn find_gamescope_eis_socket() -> Option<String> {
/// not require any particular desktop to be running. Quiet (no version warning — that's for the
/// create path); just checks the binary executes.
pub(crate) fn is_available() -> bool {
std::process::Command::new(gamescope_bin())
.arg("--version")
.output()
.map(|o| o.status.success())
.unwrap_or(false)
crate::proc::output_within(
Command::new(gamescope_bin()).arg("--version"),
VERSION_PROBE_BUDGET,
)
.map(|o| o.status.success())
.unwrap_or(false)
}
/// The gamescope binary this host spawns, resolved ONCE per process:
@@ -400,14 +452,20 @@ fn which_in_path(name: &str) -> Option<String> {
///
/// Monotonic, so one probe answers every capability:
/// * `1` — 10-bit BT.2020/PQ capture formats ([`gamescope_hdr_capable`]);
/// * `2` — …and `--pipewire-composite-cursor` ([`gamescope_can_composite_cursor`]).
/// * `2` — …and `--pipewire-composite-cursor` ([`gamescope_can_composite_cursor`]);
/// * `3` — …and `--custom-refresh-rates` ([`gamescope_can_offer_refresh_rates`]);
/// * `4` — …and `--pipewire-composite-external-overlay`
/// ([`gamescope_can_composite_external_overlay`]).
///
/// When upstream takes the functional patches this becomes a plain version floor, exactly like
/// [`MIN_GAMESCOPE_OVERLAY`].
fn gamescope_patch_level() -> u32 {
static LEVEL: std::sync::OnceLock<u32> = std::sync::OnceLock::new();
*LEVEL.get_or_init(|| {
let Ok(out) = Command::new(gamescope_bin()).arg("--version").output() else {
let Ok(out) = crate::proc::output_within(
Command::new(gamescope_bin()).arg("--version"),
VERSION_PROBE_BUDGET,
) else {
return 0;
};
// The banner goes to stderr on some builds, stdout on others (same as the version gate).
@@ -530,7 +588,8 @@ fn parse_patch_level(banner: &str) -> u32 {
/// WSI-layer check has to compare TWO binaries — ours and the distro's — and a `None` there means
/// "leave the layer alone", not "assume old".
pub(super) fn gamescope_version_of(bin: &std::path::Path) -> Option<(u32, u32, u32)> {
let out = Command::new(bin).arg("--version").output().ok()?;
let out = crate::proc::output_within(Command::new(bin).arg("--version"), VERSION_PROBE_BUDGET)
.ok()?;
// Same stdout/stderr split as the version gate: builds disagree on where the banner goes.
let text = format!(
"{}{}",
@@ -549,8 +608,15 @@ const MIN_GAMESCOPE: (u32, u32, u32) = (3, 16, 22);
/// the overlay-window paint (gated on the consumer negotiating `gamescope_focus_appid == 0`, which
/// we do by never advertising that property — see the capturer's EnumFormat builders) first ships
/// in 3.16.23 (gamescope commits `ccd62074` + `f8b33d38`). Below this the overlay is *never* in the
/// node, so it cannot appear in the stream no matter what the host does. The cursor and
/// external-overlay / notification layers are excluded on *every* version (handled host-side).
/// node, so it cannot appear in the stream no matter what the host does.
///
/// On a **stock** gamescope the cursor and external-overlay / notification layers are excluded from
/// `paint_pipewire` on every version, and the host handles the cursor itself. punktfunk's own build
/// puts both back: `--pipewire-composite-cursor` at patch level 2+
/// ([`gamescope_can_composite_cursor`], which is what suppresses the host-side blend) and
/// `--pipewire-composite-external-overlay` at 4+ ([`gamescope_can_composite_external_overlay`]) —
/// see [`gamescope_patch_level`]. So "the overlay is missing from the stream" is a question about
/// which flags reached the running compositor, not about host-side compositing.
const MIN_GAMESCOPE_OVERLAY: (u32, u32, u32) = (3, 16, 23);
/// Best-effort: warn if the installed gamescope is older than [`MIN_GAMESCOPE`] (capture is
@@ -558,10 +624,11 @@ const MIN_GAMESCOPE_OVERLAY: (u32, u32, u32) = (3, 16, 23);
/// the stream). Parsing failures are silent (don't block a possibly-fine custom build) — this is a
/// diagnostic, not a gate. Returns the parsed version when it could read one.
pub(super) fn check_gamescope_version() -> Option<(u32, u32, u32)> {
let out = Command::new(gamescope_bin())
.arg("--version")
.output()
.ok()?;
let out = crate::proc::output_within(
Command::new(gamescope_bin()).arg("--version"),
VERSION_PROBE_BUDGET,
)
.ok()?;
// gamescope prints the version banner to stderr on some builds, stdout on others.
let text = format!(
"{}{}",
@@ -34,24 +34,28 @@ pub(crate) fn list_monitors() -> anyhow::Result<Vec<PhysicalMonitor>> {
Ok(heads_under(
Path::new("/sys/class/drm"),
&super::gamescope_argvs(),
super::current_gamescope_output_size(),
))
}
/// [`list_monitors`] against an arbitrary sysfs root and a supplied argv set — the unit-testable
/// core. `output_size` is gamescope's own `-W`/`-H`, which OUTRANKS the EDID's preferred timing
/// because it is the size the capture node actually produces.
fn heads_under(
base: &Path,
argvs: &[Vec<String>],
output_size: Option<(u32, u32)>,
) -> Vec<PhysicalMonitor> {
/// core.
///
/// The head's size comes from the `-W`/`-H` of the argv selected HERE, which OUTRANKS the EDID's
/// preferred timing because it is the size the capture node actually produces. It used to arrive as
/// a parameter filled by a scan over ALL gamescopes on the box — including the nested child this
/// function had just deliberately rejected, and any headless one the crate spawned itself. On a
/// Deck driving eDP-1 at 1280x800 with a game nested at `-W 1920 -H 1080`, the panel was listed as
/// 1920x1080, and `mirror::create` publishes that row verbatim as the `preferred_mode` the stream
/// negotiates against — a mode the composited node never produces, and one `check_mirrorable` waves
/// through because it only rejects `0x0`.
fn heads_under(base: &Path, argvs: &[Vec<String>]) -> Vec<PhysicalMonitor> {
// A gamescope that isn't on DRM has no head of its own. Any DRM-backed one qualifies the box:
// a Deck streaming from Game Mode often has a second, nested gamescope running the game inside
// the session one, and that child must not disqualify its parent.
let Some(argv) = argvs.iter().find(|a| drives_drm(a)) else {
return Vec::new();
};
let output_size = super::gamescope_output_size(argv);
let connected = connected_connectors(base);
if connected.is_empty() {
return Vec::new();
@@ -342,7 +346,6 @@ mod tests {
let heads = heads_under(
&base,
&[argv("/usr/bin/gamescope --prefer-output HDMI-A-1 --steam")],
None,
);
assert_eq!(heads.len(), 1);
assert_eq!(heads[0].connector, "HDMI-A-1");
@@ -366,12 +369,12 @@ mod tests {
"gamescope --backend sdl",
] {
assert!(
heads_under(&base, &[argv(a)], None).is_empty(),
heads_under(&base, &[argv(a)]).is_empty(),
"expected no heads for {a:?}"
);
}
// No gamescope at all is the same answer, not an error.
assert!(heads_under(&base, &[], None).is_empty());
assert!(heads_under(&base, &[]).is_empty());
std::fs::remove_dir_all(&base).unwrap();
}
@@ -384,12 +387,15 @@ mod tests {
&base,
&[
argv("gamescope --backend wayland -W 1280 -H 800"),
argv("/usr/bin/gamescope --prefer-output *,eDP-1 --steam"),
argv("/usr/bin/gamescope --prefer-output *,eDP-1 -W 2560 -H 1440 --steam"),
],
None,
);
assert_eq!(heads.len(), 1);
assert_eq!(heads[0].connector, "eDP-1");
// …and the size comes from the DRM PARENT, not from the nested child listed first. Reading
// it off any-gamescope-on-the-box is what published a 1280x800 panel as the mirror's
// preferred mode on a box where the game happened to be nested at a different size.
assert_eq!((heads[0].width, heads[0].height), (2560, 1440));
std::fs::remove_dir_all(&base).unwrap();
}
@@ -404,7 +410,7 @@ mod tests {
("card1-HDMI-A-1", "connected\n", "enabled\n"),
],
);
let heads = heads_under(&base, &[argv("gamescope --prefer-output *,eDP-1")], None);
let heads = heads_under(&base, &[argv("gamescope --prefer-output *,eDP-1")]);
assert_eq!(heads.len(), 1);
assert_eq!(heads[0].connector, "eDP-1");
std::fs::remove_dir_all(&base).unwrap();
@@ -421,7 +427,7 @@ mod tests {
("card1-HDMI-A-1", "connected\n", "enabled\n"),
],
);
let heads = heads_under(&base, &[argv("gamescope --steam")], None);
let heads = heads_under(&base, &[argv("gamescope --steam")]);
assert_eq!(
heads
.iter()
@@ -440,7 +446,7 @@ mod tests {
"unplugged",
&[("card1-HDMI-A-1", "disconnected\n", "disabled\n")],
);
assert!(heads_under(&base, &[argv("gamescope --steam")], None).is_empty());
assert!(heads_under(&base, &[argv("gamescope --steam")]).is_empty());
std::fs::remove_dir_all(&base).unwrap();
}
@@ -452,7 +458,6 @@ mod tests {
let heads = heads_under(
&base,
&[argv("gamescope -W 2560 -H 1440 --prefer-output HDMI-A-1")],
Some((2560, 1440)),
);
assert_eq!((heads[0].width, heads[0].height), (2560, 1440));
std::fs::remove_dir_all(&base).unwrap();
@@ -511,7 +516,6 @@ mod tests {
&[argv(
"gamescope --nested-refresh 30 --prefer-output HDMI-A-1",
)],
None,
);
assert_eq!(heads[0].refresh_mhz, 60_000);
assert_eq!(heads[0].mode_label(), "1920x1080@60");
@@ -539,7 +543,7 @@ mod tests {
"3840x2160\n1920x1080\n",
)
.unwrap();
let heads = heads_under(&base, &[argv("gamescope --steam")], None);
let heads = heads_under(&base, &[argv("gamescope --steam")]);
assert_eq!((heads[0].width, heads[0].height), (3840, 2160));
std::fs::remove_dir_all(&base).unwrap();
}
@@ -143,17 +143,57 @@ pub(crate) fn run() -> Result<()> {
}
}
/// How long the splash waits for the session's X server before giving up.
const CONNECT_BUDGET: Duration = Duration::from_secs(10);
/// Connect to the session's `DISPLAY`, retrying briefly — gamescope sets the variable before
/// exec'ing the nested command, but a slow Xwayland under cold driver init gets a grace window.
///
/// The retry runs on a worker thread and the budget is enforced by `recv_timeout` rather than by
/// re-checking a deadline between attempts. The difference is the whole point: `x11rb::connect`
/// has no timeout of its own, so against an Xwayland that ACCEPTED the socket and then never
/// answered the setup handshake it blocks indefinitely — and a deadline consulted only in the
/// `Err` arm is never reached at all. That is the failure this module exists to prevent, from the
/// inside: no painting client, no composite, no PipeWire buffers, and the capture dies on its 10 s
/// first-frame timeout having never logged "gamescope splash: mapped", so the diagnosis points
/// anywhere but here.
///
/// A worker still stuck in `connect` is abandoned rather than joined; it is one thread in a
/// process whose whole job is this window, and the alternative is the hang.
fn connect_with_retry() -> Result<(RustConnection, usize)> {
let deadline = std::time::Instant::now() + Duration::from_secs(10);
loop {
match x11rb::connect(None) {
Ok(ok) => return Ok(ok),
Err(e) if std::time::Instant::now() >= deadline => {
return Err(e).context("gamescope splash: could not connect to the session DISPLAY")
let (tx, rx) = std::sync::mpsc::channel();
std::thread::Builder::new()
.name("pf-splash-x11-connect".into())
.spawn(move || {
let deadline = std::time::Instant::now() + CONNECT_BUDGET;
loop {
match x11rb::connect(None) {
Ok(ok) => {
let _ = tx.send(Ok(ok));
return;
}
Err(e) if std::time::Instant::now() >= deadline => {
let _ = tx.send(Err(e));
return;
}
Err(_) => std::thread::sleep(Duration::from_millis(200)),
}
}
Err(_) => std::thread::sleep(Duration::from_millis(200)),
})
.context("gamescope splash: could not start the X connect thread")?;
// A little past the worker's own deadline, so a connect that merely finished slowly still wins
// and only a genuinely blocked one trips this.
match rx.recv_timeout(CONNECT_BUDGET + Duration::from_secs(1)) {
Ok(Ok(conn)) => Ok(conn),
Ok(Err(e)) => Err(e).context("gamescope splash: could not connect to the session DISPLAY"),
Err(_) => {
tracing::warn!(
secs = CONNECT_BUDGET.as_secs(),
"gamescope splash: the session's X server accepted no connection and never \
answered giving up. Nothing will paint in this gamescope, so it will composite \
nothing and the capture will starve; the gamescope log is where the reason is."
);
anyhow::bail!("gamescope splash: connecting to the session DISPLAY did not return")
}
}
}
+239 -23
View File
@@ -5,9 +5,10 @@
//! protocols, so it shares the wlr virtual-input path with sway — but it needs its own IPC and
//! portal, so it is a **distinct backend** from [`super::wlroots`], not a branch inside it (D1):
//!
//! 1. `hyprctl output create headless PF-<n>` adds a named headless output — Hyprland supports
//! 1. `hyprctl output create headless PF-<pid>-<n>` adds a named headless output — Hyprland supports
//! **explicit names**, so no before/after diffing like sway's `HEADLESS-N` (D6). We poll
//! `hyprctl -j monitors` until the name shows up.
//! `hyprctl -j monitors` until the name shows up. The creator's pid rides in the name so a
//! crashed host's leftovers are attributable, and only those (see [`reclaim_leftovers_once`]).
//! 2. A monitor rule sets the client's exact mode. [`set_monitor_rule`] uses `hyprctl keyword
//! monitor NAME,WxH@Hz,auto,1` (the hyprlang path — the default config manager on every current
//! release, ≥0.55 included) and falls back to the Lua `hyprctl eval 'hl.monitor{…}'` only for a
@@ -69,12 +70,46 @@ fn picker_selection_line(name: &str) -> String {
format!("[SELECTION]screen:{name}\n")
}
/// Monotonic per-process counter for headless output names (`PF-1`, `PF-2`, …). Named outputs kill
/// the before/after diff race sway needs (D6).
/// Monotonic per-process counter for headless output names (`PF-<pid>-1`, `PF-<pid>-2`, …). Named
/// outputs kill the before/after diff race sway needs (D6).
static OUTPUT_SEQ: AtomicU32 = AtomicU32::new(0);
/// The name for our next headless output: `PF-<pid>-<n>`.
///
/// The pid is not decoration. `OutputGuard::drop` is the only thing that removes an output, so a
/// host that was SIGKILLed leaves its outputs in the compositor — and a bare `PF-<n>` counter starts
/// again at `PF-1` in the next process, colliding with the corpses it just inherited. Stamping the
/// creator's pid into the name makes a leftover both recognisable and *attributable*, which is what
/// lets [`reclaim_leftovers_once`] remove only the ones whose owner is gone.
fn next_output_name() -> String {
format!("PF-{}", OUTPUT_SEQ.fetch_add(1, Ordering::Relaxed) + 1)
format!(
"PF-{}-{}",
std::process::id(),
OUTPUT_SEQ.fetch_add(1, Ordering::Relaxed) + 1
)
}
/// Is `name` an output some punktfunk host created (`PF-<pid>-<n>`, or a legacy `PF-<n>`)? Pure —
/// this is what [`list_monitors`] reports as `managed`, so a user's own monitor called `PF-office`
/// must not qualify.
fn is_managed_output(name: &str) -> bool {
let Some(rest) = name.strip_prefix("PF-") else {
return false;
};
!rest.is_empty()
&& rest
.split('-')
.all(|part| !part.is_empty() && part.bytes().all(|b| b.is_ascii_digit()))
}
/// The pid of the host that created `name`, for `PF-<pid>-<n>` only. `None` for anything else —
/// including a legacy `PF-<n>` from a host older than this naming scheme, which carries no owner and
/// therefore may not be reclaimed on a guess.
fn output_owner_pid(name: &str) -> Option<u32> {
let rest = name.strip_prefix("PF-")?;
let (pid, seq) = rest.split_once('-')?;
seq.parse::<u32>().ok()?;
pid.parse::<u32>().ok()
}
/// The Hyprland virtual-display driver. Stateless — each [`create`](VirtualDisplay::create) adds one
@@ -100,11 +135,24 @@ impl HyprlandDisplay {
/// under `$XDG_RUNTIME_DIR/hypr/*/.socket.sock` (so the systemd `--user` host works without env
/// import, unlike sway's `SWAYSOCK`; the signature is then exported by `apply_session_env`). Cheap,
/// side-effect-free — safe on the enumeration path.
///
/// Both env reads take [`crate::with_env_lock`] — in ONE scope, so the pair is sampled from a single
/// consistent view. This runs on a management worker (`/host/compositors` → [`crate::available`])
/// concurrently with another connect's `apply_session_env`, which `set_var`s the signature for a
/// live Hyprland session and `remove_var`s it for anything else; a glibc `getenv` racing that
/// `setenv`/`unsetenv` is the `environ` realloc data race ENV_LOCK exists for. No caller holds the
/// lock (it is not reentrant), and the `read_dir` below deliberately runs outside it.
pub fn is_available() -> bool {
if std::env::var_os("HYPRLAND_INSTANCE_SIGNATURE").is_some() {
let (sig, runtime) = crate::with_env_lock(|| {
(
std::env::var_os("HYPRLAND_INSTANCE_SIGNATURE"),
std::env::var_os("XDG_RUNTIME_DIR"),
)
});
if sig.is_some() {
return true;
}
let dir = match std::env::var_os("XDG_RUNTIME_DIR") {
let dir = match runtime {
Some(d) => std::path::PathBuf::from(d).join("hypr"),
None => return false,
};
@@ -147,6 +195,9 @@ impl VirtualDisplay for HyprlandDisplay {
fn create(&mut self, mode: Mode) -> Result<VirtualOutput> {
// Log the permission-system caveat once per process (silent black frames otherwise).
preflight_once();
// Remove any output a PREVIOUS host left in this compositor, before we mint our first.
reclaim_leftovers_once();
warn_topology_is_extend_only();
let name = next_output_name();
hyprctl_dispatch(&["output", "create", "headless", &name]).with_context(|| {
@@ -181,7 +232,7 @@ impl VirtualDisplay for HyprlandDisplay {
remote_fd: Some(fd),
preferred_mode: Some((mode.width, mode.height, mode.refresh_hz)),
keepalive: Box::new(Keepalive {
_stop: StopGuard(stop),
_stop: stop,
_output: output,
}),
// Owned (the compositor output is ours to tear down), but not registry-poolable: the
@@ -212,6 +263,62 @@ impl Drop for StopGuard {
}
}
/// Remove the `PF-<pid>-<n>` outputs left behind by host processes that are **gone**, once per
/// process before we create our first.
///
/// [`OutputGuard::drop`] is the only unplug path there is, so a host that was SIGKILLed, OOM-killed
/// or crashed leaves its headless outputs in the compositor for as long as the Hyprland session
/// lives — a dead `PF-…` head in the operator's layout, forever, with the next host start happily
/// adding more beside it. Reclaim is keyed on the OWNER pid in the name and only removes an output
/// whose creator no longer exists, so a second live host on the same session (or this very process)
/// can never have its output pulled out from under it. `Once` puts the sweep strictly before this
/// process owns anything, and blocks a concurrent first `create` until it is done.
fn reclaim_leftovers_once() {
static RECLAIMED: Once = Once::new();
RECLAIMED.call_once(|| {
let Ok(names) = monitor_names() else { return };
for name in names {
let Some(pid) = output_owner_pid(&name) else {
// Either not ours, or a legacy `PF-<n>` with no owner recorded — which we must not
// remove on a guess, because a still-running older host may be streaming it.
if is_managed_output(&name) {
tracing::debug!(output = %name, "a managed headless output with no owner pid in \
its name (an older host build) left alone");
}
continue;
};
if pid == std::process::id() || std::path::Path::new(&format!("/proc/{pid}")).exists() {
continue;
}
match hyprctl_dispatch(&["output", "remove", &name]) {
Ok(()) => tracing::info!(output = %name, owner_pid = pid, "removed a headless \
output left behind by a host that is no longer running"),
Err(e) => tracing::warn!(output = %name, owner_pid = pid, error = %format!("{e:#}"),
"could not remove a leftover headless output"),
}
}
});
}
/// The configured [`crate::policy::Topology`] is not implemented on this backend — say so once per
/// create instead of leaving the management API's echo as the only signal that the pin was dropped
/// (sweep 13.18). The Hyprland headless output is always an EXTENSION: nothing here promotes it to
/// primary or disables the operator's heads.
fn warn_topology_is_extend_only() {
let topology = crate::effective_topology();
if !matches!(
topology,
crate::policy::Topology::Extend | crate::policy::Topology::Auto
) {
tracing::warn!(
?topology,
"hyprland: this backend implements EXTEND only — the headless output is added beside \
the operator's heads and nothing is promoted or disabled. Configure `topology: extend` \
to stop the console promising otherwise."
);
}
}
/// Owns the created headless output; dropping it removes it from Hyprland.
struct OutputGuard(String);
@@ -226,14 +333,25 @@ impl Drop for OutputGuard {
}
}
/// Budget for one `hyprctl` call ([`crate::proc`]).
///
/// `hyprctl` is a client of the compositor it drives — it connects to the instance socket and waits
/// for a reply, so against a wedged Hyprland it never returns. These calls run on the session's
/// stream thread, whose only way to end a session is to return, so one hung query used to wedge the
/// session for good. Generous next to a healthy call (single-digit milliseconds), and every call
/// site already has a failed-query path.
const HYPRCTL_BUDGET: Duration = Duration::from_secs(5);
/// Budget for the one-shot xdph restart. `systemctl --user try-restart` waits for the user manager's
/// job to settle, so it is the slowest helper on this path — and its result is already ignored.
const PORTAL_RESTART_BUDGET: Duration = Duration::from_secs(10);
/// Run `hyprctl <args>`, returning stdout. `hyprctl` reads `HYPRLAND_INSTANCE_SIGNATURE` from the
/// env (exported by `apply_session_env`) to reach the right instance socket. It exits non-zero on a
/// hard failure, but for dispatch commands it can print an error with status 0 — see
/// [`hyprctl_dispatch`].
fn hyprctl(args: &[&str]) -> Result<String> {
let out = Command::new("hyprctl")
.args(args)
.output()
let out = crate::proc::output_within(Command::new("hyprctl").args(args), HYPRCTL_BUDGET)
.context("run hyprctl (is Hyprland installed?)")?;
if !out.status.success() {
bail!(
@@ -251,12 +369,36 @@ fn hyprctl(args: &[&str]) -> Result<String> {
/// write between ours and xdph's read would silently steer capture at the other session's output.
static SELECTION_LOCK: std::sync::Mutex<()> = std::sync::Mutex::new(());
/// The per-session selection file, removed when the handshake it steers is over.
///
/// Its lifetime is the HANDSHAKE, not the session: the shim cats it once, inside
/// [`select_and_cast`]'s critical section, and everything after that is the cast's own business.
/// Left behind (as it was) the stale `[SELECTION]screen:PF-…` outlives the output `Drop` has since
/// removed, and it permanently shadows xdph's documented empty-read fallback — every later capture
/// that reaches the picker without a session of ours is steered at an output that is gone. Tying
/// removal to the CAST instead would be worse: the file is one per user, so a session ending hours
/// later would delete a *sibling's* selection out from under its picker.
struct SelectionFile(String);
impl Drop for SelectionFile {
fn drop(&mut self) {
if let Err(e) = std::fs::remove_file(&self.0) {
if e.kind() != std::io::ErrorKind::NotFound {
tracing::debug!(path = %self.0, error = %e, "could not remove the xdph selection file");
}
}
}
}
/// Point xdph's custom picker at `output` and run the ScreenCast handshake, returning the portal fd
/// + node id and the guard that stops the cast. The caller must hold [`SELECTION_LOCK`].
fn select_and_cast(output: &str, hw_cursor: bool) -> Result<(OwnedFd, u32, Arc<AtomicBool>)> {
fn select_and_cast(output: &str, hw_cursor: bool) -> Result<(OwnedFd, u32, StopGuard)> {
ensure_xdph_config()?;
let sel = selection_file();
std::fs::write(&sel, picker_selection_line(output)).with_context(|| format!("write {sel}"))?;
// Owned from the write on: every arm below (and every `?`) leaves the handshake, which is the
// only thing that reads it.
let _sel_file = SelectionFile(sel);
let (setup_tx, setup_rx) = std::sync::mpsc::channel::<Result<(OwnedFd, u32), String>>();
let stop = Arc::new(AtomicBool::new(false));
let stop_thread = stop.clone();
@@ -264,8 +406,16 @@ fn select_and_cast(output: &str, hw_cursor: bool) -> Result<(OwnedFd, u32, Arc<A
.name("punktfunk-hypr-cast".into())
.spawn(move || portal_thread(setup_tx, stop_thread, hw_cursor))
.context("spawn hyprland portal thread")?;
// Built BEFORE the wait so EVERY error arm below sets the flag on its way out — as Mutter's
// `create` does. Returning the bare `Arc` and letting the CALLER wrap it left the two failure
// arms dropping an un-set flag: the thread's `send` can still LAND in the queue in the window
// between `recv_timeout` giving up and `setup_rx` being dropped, so it reports success and then
// parks forever on `while !stop`, holding a live ScreenCast session, its zbus connection, an
// `OwnedFd` and a 2-worker tokio runtime — one more set per slow-portal connect, for the host's
// lifetime, against an output that no longer exists.
let guard = StopGuard(stop);
match setup_rx.recv_timeout(Duration::from_secs(20)) {
Ok(Ok((fd, node_id))) => Ok((fd, node_id, stop)),
Ok(Ok((fd, node_id))) => Ok((fd, node_id, guard)),
Ok(Err(e)) => bail!("ScreenCast portal on {output} failed: {e}"),
Err(_) => bail!("timed out waiting for the ScreenCast portal on {output}"),
}
@@ -285,7 +435,7 @@ pub(crate) fn stream_existing_output(
Ok(crate::mirror::MirrorStream {
node_id,
remote_fd: Some(fd),
keepalive: Box::new(StopGuard(stop)),
keepalive: Box::new(stop),
})
}
@@ -330,11 +480,12 @@ pub(crate) fn list_monitors() -> Result<Vec<crate::monitors::PhysicalMonitor>> {
.unwrap_or(1.0),
primary: m.get("focused").and_then(|v| v.as_bool()).unwrap_or(false),
enabled: !m.get("disabled").and_then(|v| v.as_bool()).unwrap_or(false),
// Our headless outputs are named `PF-<n>` (see `next_output_name`).
// Our headless outputs are named `PF-<pid>-<n>` (see `next_output_name`); the shape
// is checked, not just the prefix, so a user's own `PF-office` stays theirs.
managed: m
.get("name")
.and_then(|v| v.as_str())
.is_some_and(|n| n.starts_with("PF-")),
.is_some_and(is_managed_output),
})
})
.collect();
@@ -382,6 +533,23 @@ fn wait_monitor_ready(name: &str, timeout: Duration) -> Result<()> {
}
}
/// Every monitor name Hyprland reports, **disabled ones included** (`-j monitors all`) — a leftover
/// output from a dead host may well have ended up disabled, and [`reclaim_leftovers_once`] must see
/// it anyway.
fn monitor_names() -> Result<Vec<String>> {
let out = hyprctl(&["-j", "monitors", "all"])?;
let monitors: serde_json::Value =
serde_json::from_str(&out).context("parse hyprctl -j monitors all")?;
Ok(monitors
.as_array()
.map(|a| {
a.iter()
.filter_map(|m| m.get("name").and_then(|n| n.as_str()).map(str::to_owned))
.collect()
})
.unwrap_or_default())
}
/// Is a monitor named `name` present in `hyprctl -j monitors` (JSON)?
fn monitor_exists(name: &str) -> Result<bool> {
let out = hyprctl(&["-j", "monitors"])?;
@@ -417,17 +585,33 @@ fn set_monitor_rule(name: &str, mode: Mode) -> Result<()> {
);
let keyword: Vec<&str> = vec!["keyword", "monitor", &spec];
let eval: Vec<&str> = vec!["eval", &lua];
// What each form actually said. hyprctl reports a rejection in its OUTPUT TEXT ("eval is only
// supported with the lua config manager", "invalid monitor rule", a permission denial), and
// dropping it on the floor with `.is_err()` is what left the failure below guessing at GBM when
// the compositor had already named the real cause.
let mut attempts: Vec<String> = Vec::new();
for a in [&keyword, &eval] {
// A wrong-era command errors (`keyword` gone under Lua, or `eval` under hyprlang) — skip to
// the other form. A command that's accepted then has up to the timeout to take effect.
if hyprctl_dispatch(a).is_err() {
if let Err(e) = hyprctl_dispatch(a) {
let said = format!("{e:#}");
tracing::debug!(output = %name, cmd = ?a, error = %said, "hyprctl rejected this monitor-rule form — trying the other config era");
attempts.push(said);
continue;
}
if wait_exact_mode(name, mode, Duration::from_millis(1500)) {
tracing::debug!(output = %name, cmd = ?a, w = mode.width, h = mode.height, "monitor adopted the requested mode");
return Ok(());
}
attempts.push(format!(
"hyprctl {a:?} was accepted but the mode never took effect"
));
}
let said = if attempts.is_empty() {
"nothing (no form was attempted)".to_string()
} else {
attempts.join("; ")
};
// Neither form produced the exact mode. Distinguish "usable but different size" (proceed with a
// warning — a working stream beats none) from "0×0 / gone" (the output has no framebuffer at all).
match monitor_size(name)? {
@@ -436,14 +620,20 @@ fn set_monitor_rule(name: &str, mode: Mode) -> Result<()> {
output = %name,
requested = %format!("{}x{}", mode.width, mode.height),
got = %format!("{w}x{h}"),
hyprctl = %said,
"Hyprland did not adopt the exact requested mode — streaming at the output's current size"
);
Ok(())
}
// The output has no framebuffer at all. Lead with what hyprctl SAID: if every form was
// rejected the cause is named right there (wrong config era, a permission denial, a bad
// rule) and no allocation was ever attempted; only a form that was accepted and still left
// the output at 0×0 points at the compositor failing to back the mode.
_ => bail!(
"headless output {name} never got a framebuffer (stayed 0x0) after the monitor rule for \
{}x{}@{hz} the compositor could not back the mode, likely a headless GBM/dmabuf \
allocation failure (GPU driver; cf. Sunshine#4197). Check the Hyprland log.",
{}x{}@{hz}. hyprctl said: {said}. If a form was accepted, the compositor could not back \
the mode likely a headless GBM/dmabuf allocation failure (GPU driver; cf. \
Sunshine#4197). Check the Hyprland log.",
mode.width,
mode.height
),
@@ -574,13 +764,17 @@ fn ensure_xdph_config() -> Result<()> {
return Ok(());
}
tracing::info!(path = %path.display(), "pointed xdg-desktop-portal-hyprland at the managed picker shim");
let _ = Command::new("systemctl")
.args([
// Bounded: `systemctl --user` blocks on the user manager's job queue, and this runs on the
// session's stream thread. Its result was already ignored — a timeout just means xdph picks the
// new config up whenever it next starts.
let _ = crate::proc::status_within(
Command::new("systemctl").args([
"--user",
"try-restart",
"xdg-desktop-portal-hyprland.service",
])
.status();
]),
PORTAL_RESTART_BUDGET,
);
Ok(())
}
@@ -702,6 +896,28 @@ mod tests {
assert_ne!(a, b);
}
/// The name carries the creating host's pid, which is what makes a leftover attributable — a
/// reclaim that could not tell whose output it was would have to remove a LIVE sibling's or
/// nothing at all.
#[test]
fn a_name_carries_its_owner_pid_and_only_ours_does() {
let mine = next_output_name();
assert_eq!(output_owner_pid(&mine), Some(std::process::id()));
assert!(is_managed_output(&mine));
// A legacy `PF-<n>` from an older host: recognisably managed, but with no owner recorded —
// so it may be reported, never reclaimed on a guess.
assert!(is_managed_output("PF-1"));
assert_eq!(output_owner_pid("PF-1"), None);
// Not ours: a user's own monitor name that happens to start with the prefix, and the
// connectors every wlr-family compositor mints.
for theirs in ["PF-office", "PF-", "PF-12-abc", "HEADLESS-1", "DP-1", ""] {
assert!(!is_managed_output(theirs), "{theirs:?} is not ours");
assert_eq!(output_owner_pid(theirs), None, "{theirs:?} has no owner");
}
}
#[test]
fn picker_line_carries_the_selection_marker() {
// xdph requires the `[SELECTION]` prefix; a bare `screen:NAME` is rejected as strange output.
+633 -130
View File
@@ -23,9 +23,6 @@
//! "Could not find output". We talk raw Wayland on `$WAYLAND_DISPLAY`, so the host must run inside
//! the KWin session's environment.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{Mode, VirtualDisplay, VirtualOutput};
use anyhow::{anyhow, bail, Context, Result};
use std::os::fd::{AsFd, AsRawFd};
@@ -33,7 +30,8 @@ use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::mpsc::Sender;
use std::sync::Arc;
use std::thread;
use std::time::Duration;
use std::time::{Duration, Instant};
use wayland_client::protocol::wl_callback::{self, WlCallback};
use wayland_client::protocol::wl_output::{self, WlOutput};
use wayland_client::protocol::wl_registry::{self, WlRegistry};
use wayland_client::{Connection, Dispatch, Proxy, QueueHandle};
@@ -237,7 +235,20 @@ impl VirtualDisplay for KwinDisplay {
Some(id) => format!("{VOUT_NAME}-{id}"),
None => VOUT_NAME.to_string(),
};
self.last_name = Some(name.clone()); // for apply_position (registry-driven §6.2 layout)
// `apply_position`'s kscreen-doctor fallback (the registry-driven §6.2 layout) addresses
// `last_name`, so seed it with `Virtual-<name>`: the address KWin exposes our output under
// and the ONLY spelling kscreen-doctor can resolve. The bare `name` we ask KWin for
// (`punktfunk`) matches no output at all, so seeding it with that left every position apply
// shelling out against an address that can never exist — and the `is_none()` guard that was
// supposed to correct it later could never fire, because this write is never `None`.
let our_prefix = format!("Virtual-{name}");
self.last_name = Some(our_prefix.clone());
// Every `create` re-resolves its own output, so the PREVIOUS one's UUID must not survive
// into this one. A supersede keeps this `KwinDisplay` and creates the replacement while the
// predecessor is still alive, so a stale UUID still RESOLVES: `set_position` would find the
// old output, position it, report success, and never reach the fallback — the new display
// silently stays where it was born. Re-set below only if the in-process path handles us.
self.our_uuid = None;
let (width, height) = (mode.width, mode.height);
let pointer_mode = if self.hw_cursor {
POINTER_METADATA
@@ -255,10 +266,18 @@ impl VirtualDisplay for KwinDisplay {
virtual_output_thread(w, h, name_thread, pointer_mode, setup_tx, stop_thread)
})
.context("spawn KWin virtual-output thread")?;
match setup_rx.recv_timeout(Duration::from_secs(20)) {
match setup_rx.recv_timeout(OPENER_BUDGET) {
Ok(Ok(v)) => Ok((v, stop)),
Ok(Err(e)) => bail!("KWin virtual output failed: {e}"),
Err(_) => bail!("timed out creating the KWin virtual output"),
Err(_) => {
// Nothing else will ever flip this `stop`: it is dropped with the error, and
// the `StopGuard` that normally owns it is only built on the success path. So
// the worker — which is by construction still inside `await_created` — would
// sit out its own budget holding a half-built output whose Wayland connection
// KWin keeps the output alive for. Release it here.
stop.store(true, Ordering::Relaxed);
bail!("timed out creating the KWin virtual output")
}
}
};
// KWin creates virtual outputs at a hardcoded 60 Hz, `stream_virtual_output` has no
@@ -293,8 +312,8 @@ impl VirtualDisplay for KwinDisplay {
);
// Topology + positioning address OUR output by its kde_output_management UUID (resolved
// in-process in `apply_topology`, supersede-robust) — no early kscreen-doctor resolve, so
// the path never shells out. `Virtual-<name>` is the name KWin exposes our output as.
let our_prefix = format!("Virtual-{name}");
// the path never shells out. `our_prefix` (computed above with `last_name`) is the name
// KWin exposes our output as.
let mut expect_exact_dims = false;
// The size the output actually ENDS UP at — the request, unless KWin's CVT generator had to
// shrink the width to the cell grain (see `CVT_H_GRANULARITY`). Reported as the output's
@@ -372,12 +391,11 @@ impl VirtualDisplay for KwinDisplay {
// kscreen-doctor backend; see `apply_topology`), with a kscreen-doctor fallback. `disabled`
// is the physical/bootstrap outputs, each `(name, "WxH@Hz")`, to restore on teardown.
let disabled = self.apply_topology(&name, &our_prefix, final_dims);
// A plain managed name is enough for apply_position's kscreen-doctor fallback when the
// in-process UUID path isn't set (single-output sessions are unambiguous; a supersede uses
// the UUID path instead). `want_high` already set `last_name` to the resolved kscreen id.
if self.last_name.is_none() {
self.last_name = Some(our_prefix);
}
// `last_name` is already the best address we have: `Virtual-<name>` from the top of this
// function, upgraded in place to the RESOLVED numeric kscreen id by whichever of the
// `want_high` fallback or `apply_topology`'s fallback actually ran a resolve. Nothing to
// fill in here — the guard that used to sit at this spot could never fire (`last_name` is
// written unconditionally above) and only made the plain-name case look handled.
// Per-group restore (§6.1): DON'T bind the re-enable to this session's keepalive (a per-session
// `StopGuard` restore would re-enable the physical the moment the FIRST of several exclusive
// sessions drops — under a still-live sibling). Instead stash it as a closure the registry lifts
@@ -385,7 +403,16 @@ impl VirtualDisplay for KwinDisplay {
// that display's output is reclaimed, so KWin never sees zero outputs). Empty ⇒ nothing to restore.
self.pending_restore = (!disabled.is_empty()).then(|| {
let disabled = disabled.clone();
// In-process first; fall back to kscreen-doctor if the compositor doesn't answer in budget.
// In-process first; fall back to kscreen-doctor if the compositor doesn't answer in
// budget. **Both halves now return honest verdicts** — `reenable_outputs` reports
// `false` unless every requested output was actually staged (an empty configuration
// used to ack as `applied` and suppress this backstop), and
// `reenable_outputs_kscreen` branches on its own exit status instead of logging
// success unconditionally. Any future extraction of these hand-rolled
// in-process-then-kscreen ladders into one facade must keep that property: a fallback
// arm that returns a value the helper never checked would re-introduce exactly the
// silent-success this pair was fixed for, behind a seam that claims to have one log
// site for every decline.
Box::new(move || {
if !crate::kwin_output_mgmt::reenable_outputs(&disabled) {
reenable_outputs_kscreen(&disabled);
@@ -409,6 +436,17 @@ impl VirtualDisplay for KwinDisplay {
/// closure only when the in-process path reports the compositor didn't answer. Called by the registry
/// when the display group's last member is torn down (design §6.1), BEFORE that member's output is
/// reclaimed — so KWin is never momentarily left with zero enabled outputs.
///
/// **This is the last line of defence for a physical monitor**, so it reports what actually
/// happened. It used to discard both `kscreen_ok` verdicts and log restored-everything
/// unconditionally — including when the call had been killed at [`KSCREEN_BUDGET`], i.e. exactly
/// the wedged compositor this fallback exists for, with a screen left dark and a green line in the
/// log saying otherwise.
///
/// Reporting honestly is not the same as *stopping* on a bad verdict, and the difference is
/// [`kscreen_verdict`]'s third state: a helper killed at its budget has told us nothing, and this
/// path must go on to the mode re-assert and the settle in that case exactly as the pre-verdict
/// code did — see the `None` arm below for what skipping them costs.
fn reenable_outputs_kscreen(outputs: &[(String, String)]) {
if outputs.is_empty() {
return;
@@ -420,20 +458,65 @@ fn reenable_outputs_kscreen(outputs: &[(String, String)]) {
.iter()
.map(|(name, _)| format!("output.{name}.enable"))
.collect();
let _ = kscreen_ok(&enable_args);
let enable_verdict = kscreen_verdict(&enable_args);
match enable_verdict {
// It ran and it refused (or could not be run at all). Nothing further to try: both the
// in-process path and this one have now declined, so the outputs stay as `exclusive` left
// them. Say so loudly — a dark monitor with no line in the log is what this whole restore
// chain exists to prevent.
Some(false) => {
tracing::error!(
outputs = ?outputs,
args = ?enable_args,
"KWin: could NOT re-enable the physical/bootstrap outputs (kscreen-doctor refused \
the config, or could not be run, after the in-process restore already declined) \
a monitor may be left dark"
);
return;
}
// Killed at [`KSCREEN_BUDGET`] — which is NOT the same as a refusal, and treating it as one
// is a regression this path already had once. kscreen-doctor applies the config and only
// THEN waits on the compositor before exiting, so a loaded KWin routinely lands the enable
// and still gets killed: the output is lit, and returning here would skip both halves of
// the rest of the restore — the mode re-assert (a 120 Hz panel comes back at the
// EDID-preferred ~60 Hz without it) and the 200 ms settle that keeps KWin from seeing zero
// enabled outputs when the caller reclaims the virtual one right after us (§6.1). The
// second budget this costs on the stream thread is deliberate and bounded, and is what the
// pre-`match` code spent unconditionally.
None => tracing::warn!(
outputs = ?outputs,
args = ?enable_args,
"KWin: kscreen-doctor was killed at its budget re-enabling the physical/bootstrap \
outputs the apply may well have landed, so continuing with the mode restore"
),
Some(true) => {}
}
// THEN re-assert each captured mode, best-effort — a bare re-enable lets KWin fall back to the
// EDID-preferred mode (a 120 Hz panel returns at ~60 Hz); this restores the exact refresh. The
// output is enabled now, so the mode set is valid; a rejected mode just leaves KWin's default.
// output is enabled now, so the mode set is valid; a rejected mode just leaves KWin's default
// a wrong refresh, not a dark screen, which is why only this half degrades to a warn.
let mode_args: Vec<String> = outputs
.iter()
.filter(|(_, mode)| !mode.is_empty())
.map(|(name, mode)| format!("output.{name}.mode.{mode}"))
.collect();
if !mode_args.is_empty() {
let _ = kscreen_ok(&mode_args);
}
let modes_restored = mode_args.is_empty() || kscreen_ok(&mode_args);
std::thread::sleep(Duration::from_millis(200));
tracing::info!(reenabled = ?outputs, "KWin: restored the physical/bootstrap outputs at their captured modes (group empty)");
// `enable_confirmed` rides along on both lines: after a budget kill the enable is *probable*,
// not established, and a log that cannot tell the operator which of the two it is put us here
// in the first place.
let enable_confirmed = enable_verdict == Some(true);
if modes_restored {
tracing::info!(reenabled = ?outputs, enable_confirmed, "KWin: restored the physical/bootstrap outputs at their captured modes (group empty)");
} else {
tracing::warn!(
reenabled = ?outputs,
args = ?mode_args,
enable_confirmed,
"KWin: re-enabled the physical/bootstrap outputs but could not re-assert their captured \
modes they are back at KWin's preferred refresh, not the one they were streaming at"
);
}
}
/// Resolve the kscreen address of the virtual output the host JUST created: the managed-prefix
@@ -488,12 +571,27 @@ const KSCREEN_BUDGET: Duration = Duration::from_secs(5);
/// `kscreen-doctor <args>` run for its exit status, bounded by [`KSCREEN_BUDGET`]. A timeout reads
/// as a failed apply — the same best-effort path a rejected argument already takes.
fn kscreen_ok(args: &[String]) -> bool {
crate::proc::status_within(
kscreen_verdict(args) == Some(true)
}
/// The same call, keeping the outcome that [`kscreen_ok`]'s `bool` throws away.
///
/// `Some(true)`/`Some(false)`: kscreen-doctor ran to completion and accepted / refused (a helper
/// that cannot be spawned at all counts as a refusal — there is nothing to wait for and no reason
/// to retry the next invocation). `None`: it was **killed at [`KSCREEN_BUDGET`]**, which is a
/// different fact entirely. kscreen-doctor applies the config and then waits on the compositor
/// before exiting, so a slow-but-working KWin gives us a kill on a request that already landed;
/// any caller that treats `None` as "it failed" is asserting something it does not know, and for
/// the restore path that assertion costs a monitor its refresh rate.
fn kscreen_verdict(args: &[String]) -> Option<bool> {
match crate::proc::status_within(
std::process::Command::new("kscreen-doctor").args(args),
KSCREEN_BUDGET,
)
.map(|s| s.success())
.unwrap_or(false)
) {
Ok(status) => Some(status.success()),
Err(e) if e.kind() == std::io::ErrorKind::TimedOut => None,
Err(_) => Some(false),
}
}
/// `kscreen-doctor -j` stdout, bounded by [`KSCREEN_BUDGET`]; `None` on any failure.
@@ -511,24 +609,111 @@ fn kscreen_json() -> Option<serde_json::Value> {
serde_json::from_slice(&kscreen_json_bytes()?).ok()
}
/// The `(width, height)` of an output's CURRENT mode from its `kscreen-doctor -j` entry.
fn output_active_size(o: &serde_json::Value) -> Option<(u32, u32)> {
let as_id = |v: &serde_json::Value| -> Option<String> {
v.as_str()
.map(|s| s.to_string())
.or_else(|| v.as_u64().map(|n| n.to_string()))
};
let current = o.get("currentModeId").and_then(as_id)?;
/// The CURRENT mode of an output from its `kscreen-doctor -j` entry, as `(width, height,
/// refresh_mHz)`. `None` if the entry names no current mode or that mode carries no size; a mode
/// with no `refreshRate` reports 0 mHz, which is the "unknown" the monitor type documents.
fn output_active_mode(o: &serde_json::Value) -> Option<(u32, u32, u32)> {
let current = o.get("currentModeId").and_then(json_id)?;
let mode = o
.get("modes")?
.as_array()?
.iter()
.find(|m| m.get("id").and_then(as_id).as_deref() == Some(current.as_str()))?;
.find(|m| m.get("id").and_then(json_id).as_deref() == Some(current.as_str()))?;
let size = mode.get("size")?;
Some((
size.get("width").and_then(|v| v.as_u64())? as u32,
size.get("height").and_then(|v| v.as_u64())? as u32,
))
let w = size.get("width").and_then(|v| v.as_u64())? as u32;
let h = size.get("height").and_then(|v| v.as_u64())? as u32;
// Hz → mHz without an intermediate round: `refreshRate` is a float (59.94, 119.92) and whole
// Hz would throw away exactly the distinction `PhysicalMonitor::refresh_mhz` exists to keep.
let mhz = mode
.get("refreshRate")
.and_then(|r| r.as_f64())
.map(|hz| (hz * 1000.0).round().max(0.0) as u32)
.unwrap_or(0);
Some((w, h, mhz))
}
/// The `(width, height)` of an output's CURRENT mode from its `kscreen-doctor -j` entry.
fn output_active_size(o: &serde_json::Value) -> Option<(u32, u32)> {
output_active_mode(o).map(|(w, h, _)| (w, h))
}
/// Every head KWin reports, for [`crate::monitors::list`] — the in-process enumerate
/// ([`crate::kwin_output_mgmt::list_monitors`]) with a `kscreen-doctor -j` fallback.
///
/// This was the ONE KWin call site with no fallback at all, while the in-process session it depends
/// on declines for exactly the reasons the other five fall back for: management global absent
/// (pre-6.x KWin), or a compositor that does not answer in budget. The console's monitor picker and
/// `PUNKTFUNK_CAPTURE_MONITOR`'s resolve then failed outright on a box whose `kscreen-doctor` was
/// perfectly able to answer — and a failed `list` is not "no monitors", it is a session that
/// refuses to start (`monitors::resolve` treats a miss as a hard error, deliberately).
pub(crate) fn list_monitors() -> Result<Vec<crate::monitors::PhysicalMonitor>> {
let declined = match crate::kwin_output_mgmt::list_monitors() {
Ok(monitors) => return Ok(monitors),
Err(e) => e,
};
let Some(doc) = kscreen_json() else {
return Err(declined.context(
"kscreen-doctor -j did not answer either (not installed, or killed at its budget)",
));
};
let monitors = monitors_from_kscreen_json(&doc);
tracing::info!(
count = monitors.len(),
reason = %declined,
"KWin: enumerated monitors via kscreen-doctor (in-process output management declined)"
);
Ok(monitors)
}
/// Parse `kscreen-doctor -j` into the shared monitor type. Split from the process call so it can be
/// tested against captured JSON — the mapping is where a picker's identity keys come from, and
/// `x`/`y` are what make two same-sized heads distinguishable at all.
///
/// Deliberately mirrors the in-process reader's contract: a disabled output has no current mode and
/// reports zeroed geometry rather than an invented one, `primary` accepts either the modern
/// `priority: 1` or the older `primary: true`, and the list is sorted by desktop position so it
/// reads left-to-right the way the desk looks.
fn monitors_from_kscreen_json(doc: &serde_json::Value) -> Vec<crate::monitors::PhysicalMonitor> {
let Some(outputs) = doc.get("outputs").and_then(|o| o.as_array()) else {
return Vec::new();
};
let mut out: Vec<crate::monitors::PhysicalMonitor> = outputs
.iter()
.filter_map(|o| {
let connector = o.get("name").and_then(|n| n.as_str())?.to_string();
let mode = output_active_mode(o);
let coord = |k: &str| {
o.get("pos")
.and_then(|p| p.get(k))
.and_then(|v| v.as_i64())
.unwrap_or(0) as i32
};
Some(crate::monitors::PhysicalMonitor {
managed: connector.starts_with(MANAGED_PREFIX),
description: crate::monitors::describe(
o.get("vendor").and_then(|v| v.as_str()).unwrap_or(""),
o.get("model").and_then(|v| v.as_str()).unwrap_or(""),
&connector,
),
width: mode.map(|m| m.0).unwrap_or(0),
height: mode.map(|m| m.1).unwrap_or(0),
refresh_mhz: mode.map(|m| m.2).unwrap_or(0),
x: coord("x"),
y: coord("y"),
scale: o
.get("scale")
.and_then(|v| v.as_f64())
.filter(|s| *s > 0.0)
.unwrap_or(1.0),
primary: o.get("primary").and_then(|p| p.as_bool()).unwrap_or(false)
|| o.get("priority").and_then(|p| p.as_u64()) == Some(1),
enabled: o.get("enabled").and_then(|e| e.as_bool()).unwrap_or(false),
connector,
})
})
.collect();
out.sort_by_key(|m| (m.x, m.y, m.connector.clone()));
out
}
/// CVT's horizontal cell granularity. KWin generates every custom mode's timing with **libxcvt**,
@@ -542,7 +727,11 @@ fn output_active_size(o: &serde_json::Value) -> Option<(u32, u32)> {
/// birth mode, and the caller falls back to 60 Hz — while KDE's display list shows the perfectly
/// good 2864x1320@119.92 mode sitting there unselected. Widths like 1920/2560/3840 are all
/// multiples of 8, which is why only phone-shaped clients ever hit it.
const CVT_H_GRANULARITY: u32 = 8;
///
/// Shared with [`crate::kwin_output_mgmt`], which matches the generated mode back the same way —
/// it used to keep its own copy under a comment claiming the two "match", which is a claim no
/// compiler was checking.
pub(crate) const CVT_H_GRANULARITY: u32 = 8;
/// One row of an output's mode list, as parsed from `kscreen-doctor -j`.
#[derive(Clone, Debug, PartialEq)]
@@ -762,7 +951,11 @@ fn read_active_mode(output: &str) -> Option<(u32, u32, u32)> {
/// The prefix EVERY managed KWin output shares — Stage 3 names them `punktfunk` / `punktfunk-<id>`,
/// which KWin exposes as `Virtual-punktfunk` / `Virtual-punktfunk-<id>`. Group membership (§6.1) is
/// recognised by this prefix, so we never have to thread the live set through the backend.
const MANAGED_PREFIX: &str = "Virtual-punktfunk";
///
/// Shared with [`crate::kwin_output_mgmt`] rather than copied: both halves of the ladder decide
/// "is this output one of OURS?" with it, and a drift between two copies would make the in-process
/// path disable a sibling session's output that the kscreen path deliberately spares.
pub(crate) const MANAGED_PREFIX: &str = "Virtual-punktfunk";
/// The current mode of an output as a kscreen-doctor mode setter, from its `-j` entry — preferring
/// the human `WxH@Hz` form (survives a mode-id re-enumeration across disable→enable) and falling back
@@ -875,14 +1068,27 @@ fn apply_virtual_primary(ours: &str) -> Vec<(String, String)> {
// the group is unambiguously the desktop — never a sibling session's output (group-aware filter).
// Each is captured WITH its current mode so teardown restores its real refresh, not KWin's default.
let others = other_enabled_outputs();
if !others.is_empty() {
let args: Vec<String> = others
.iter()
.map(|(o, _mode)| format!("output.{o}.disable"))
.collect();
let _ = kscreen(&args);
if others.is_empty() {
tracing::info!("KWin: streamed output set as the sole desktop (nothing else was enabled)");
return others;
}
let args: Vec<String> = others
.iter()
.map(|(o, _mode)| format!("output.{o}.disable"))
.collect();
if kscreen(&args) {
tracing::info!(also_disabled = ?others, "KWin: streamed output set as the sole desktop");
} else {
// Report the request, not a success: the outputs are still enabled, so the client sees the
// shell wherever KWin left it. They are returned for the restore regardless — re-enabling an
// output that was never disabled is a harmless no-op, and dropping them here would strand a
// physical dark if the disable actually landed and only the ack was lost to the budget.
tracing::warn!(
attempted_disable = ?others,
"KWin: could not disable the other outputs for the exclusive topology (kscreen-doctor \
failed or hit its budget) the streamed output is not the sole desktop"
);
}
tracing::info!(also_disabled = ?others, "KWin: streamed output set as the sole desktop");
others
}
@@ -919,10 +1125,42 @@ struct State {
node_id: Option<u32>,
failed: Option<String>,
closed: bool,
/// Every `wl_output` KWin advertises, keyed by the proxy, with its connector name once the
/// `name` event arrives. Only the monitor-mirror path ([`stream_existing_output`]) needs these
/// — `stream_output` takes a `wl_output` object, so the connector has to be resolved to one.
outputs: Vec<(WlOutput, Option<String>)>,
/// Highest `wl_display.sync` serial whose `done` has arrived the barrier [`roundtrip_within`]
/// waits on, so a compositor that accepted the connection and then stopped serving costs a
/// budget instead of the thread.
sync_done: u32,
/// Whether this connection needs `wl_output` objects at all — true ONLY on the monitor-mirror
/// path. `stream_virtual_output` names its output by string, so the virtual-output path never
/// reads [`State::outputs`]; binding them there was pure accumulation on a connection that
/// lives for the whole session, and every managed display this host creates is itself another
/// `wl_output` global.
want_outputs: bool,
/// Every `wl_output` KWin advertises, as (registry global name, proxy, connector once the
/// `name` event arrives). Only the monitor-mirror path ([`stream_existing_output`]) needs these —
/// `stream_output` takes a `wl_output` object, so the connector has to be resolved to one. The
/// global name is carried so `global_remove` can find the entry again ([`State::forget_output`]).
outputs: Vec<(u32, WlOutput, Option<String>)>,
}
impl State {
/// Drop the `wl_output` whose registry global just went away.
///
/// Both halves matter. The proxy must be `release`d — wayland-rs sends no destructor when a
/// proxy is merely dropped, so an unreleased binding is a server-side object leaked for the
/// life of a connection that lasts as long as the session. And the ENTRY must go, because
/// [`run_existing`]'s connector resolve scans this vector: a stale row for an unplugged head
/// would shadow the live output that took its connector name.
fn forget_output(&mut self, global: u32) {
let Some(pos) = self.outputs.iter().position(|(n, _, _)| *n == global) else {
return;
};
let (_, out, connector) = self.outputs.remove(pos);
// `wl_output.release` is `since 3`; below that the object simply has no destructor.
if out.version() >= 3 {
out.release();
}
tracing::debug!(?connector, "KWin: a wl_output went away — released it");
}
}
impl Dispatch<WlRegistry, ()> for State {
@@ -934,23 +1172,45 @@ impl Dispatch<WlRegistry, ()> for State {
_: &Connection,
qh: &QueueHandle<Self>,
) {
if let wl_registry::Event::Global {
name,
interface,
version,
} = event
{
if interface == Screencast::interface().name {
let v = version.min(MAX_VERSION);
state.screencast = Some(registry.bind::<Screencast, _, _>(name, v, qh, ()));
} else if interface == WlOutput::interface().name {
// v4 is where `wl_output.name` (the connector) arrives; bind at least that when the
// compositor offers it, else bind what it has and let the resolve fail loudly
// rather than mirroring an unidentifiable head.
let v = version.min(WL_OUTPUT_MAX_VERSION);
let out = registry.bind::<WlOutput, _, _>(name, v, qh, ());
state.outputs.push((out, None));
match event {
wl_registry::Event::Global {
name,
interface,
version,
} => {
if interface == Screencast::interface().name {
let v = version.min(MAX_VERSION);
state.screencast = Some(registry.bind::<Screencast, _, _>(name, v, qh, ()));
} else if state.want_outputs && interface == WlOutput::interface().name {
// v4 is where `wl_output.name` (the connector) arrives; bind at least that when
// the compositor offers it, else bind what it has and let the resolve fail
// loudly rather than mirroring an unidentifiable head.
let v = version.min(WL_OUTPUT_MAX_VERSION);
let out = registry.bind::<WlOutput, _, _>(name, v, qh, ());
state.outputs.push((name, out, None));
}
}
wl_registry::Event::GlobalRemove { name } => state.forget_output(name),
_ => {}
}
}
}
/// The `wl_display.sync` callback: `done` releases whichever [`roundtrip_within`] is waiting on
/// this serial. A plain `roundtrip()` would do the same job in one call, but it blocks on the
/// socket with no ceiling — against a compositor that accepted the connection and then stopped
/// answering, that is the session's stream thread pinned forever.
impl Dispatch<WlCallback, u32> for State {
fn event(
state: &mut Self,
_: &WlCallback,
event: wl_callback::Event,
serial: &u32,
_: &Connection,
_: &QueueHandle<Self>,
) {
if let wl_callback::Event::Done { .. } = event {
state.sync_done = state.sync_done.max(*serial);
}
}
}
@@ -969,8 +1229,8 @@ impl Dispatch<WlOutput, ()> for State {
_: &QueueHandle<Self>,
) {
if let wl_output::Event::Name { name } = event {
if let Some(slot) = state.outputs.iter_mut().find(|(o, _)| o == output) {
slot.1 = Some(name);
if let Some(slot) = state.outputs.iter_mut().find(|(_, o, _)| o == output) {
slot.2 = Some(name);
}
}
}
@@ -1052,10 +1312,16 @@ pub(crate) fn stream_existing_output(
}
})
.context("spawn KWin monitor-mirror thread")?;
let node_id = match setup_rx.recv_timeout(Duration::from_secs(20)) {
let node_id = match setup_rx.recv_timeout(OPENER_BUDGET) {
Ok(Ok(v)) => v,
Ok(Err(e)) => bail!("KWin monitor mirror failed: {e}"),
Err(_) => bail!("timed out recording the KWin output {connector:?}"),
Err(_) => {
// Same leak as the virtual-output opener: `StopOnDrop` only takes ownership of `stop`
// on the success path, so without this the mirror thread keeps recording a monitor
// nobody is watching until its own budget runs out.
stop.store(true, Ordering::Relaxed);
bail!("timed out recording the KWin output {connector:?}")
}
};
Ok(crate::mirror::MirrorStream {
node_id,
@@ -1187,7 +1453,16 @@ pub fn probe() -> Result<()> {
let qh = queue.handle();
let _registry = conn.display().get_registry(&qh, ());
let mut state = State::default();
queue.roundtrip(&mut state).context("registry roundtrip")?;
// Nothing to interrupt a probe: it is a one-shot question, bounded by the roundtrip budget.
let never = AtomicBool::new(false);
roundtrip_within(
&conn,
&mut queue,
&mut state,
&never,
1,
"registry roundtrip",
)?;
if state.screencast.is_none() {
bail!(
"KWin is up but does not expose zkde_screencast_unstable_v1 to this client — KWin gates \
@@ -1221,19 +1496,32 @@ fn run_existing(
setup_tx: &Sender<Result<u32, String>>,
stop: &AtomicBool,
) -> Result<()> {
// The opener started its own clock a moment ago; everything this worker spends before
// `await_created` comes out of the same 20 s (see [`CREATE_BUDGET`] — this path has two
// barriers, which is exactly why the create wait cannot be a fixed 15 s here).
let started = Instant::now();
let conn = Connection::connect_to_env()
.context("connect to KWin Wayland (is WAYLAND_DISPLAY set to the KWin socket?)")?;
let mut queue = conn.new_event_queue();
let qh = queue.handle();
let _registry = conn.display().get_registry(&qh, ());
let mut state = State::default();
// The one path that resolves a connector to a `wl_output`, so the only one that binds them.
let mut state = State {
want_outputs: true,
..State::default()
};
// Two roundtrips: the first processes the globals (binding screencast + every wl_output), the
// second drains each output's property burst — the `name` event we resolve the connector by.
queue.roundtrip(&mut state).context("registry roundtrip")?;
queue
.roundtrip(&mut state)
.context("wl_output property roundtrip")?;
roundtrip_within(&conn, &mut queue, &mut state, stop, 1, "registry roundtrip")?;
roundtrip_within(
&conn,
&mut queue,
&mut state,
stop,
2,
"wl_output property roundtrip",
)?;
let screencast = state.screencast.clone().ok_or_else(|| {
anyhow!(
@@ -1251,19 +1539,19 @@ fn run_existing(
let named: Vec<&str> = state
.outputs
.iter()
.filter_map(|(_, n)| n.as_deref())
.filter_map(|(_, _, n)| n.as_deref())
.collect();
let output = state
.outputs
.iter()
.find(|(_, n)| n.as_deref() == Some(connector))
.find(|(_, _, n)| n.as_deref() == Some(connector))
.or_else(|| {
state.outputs.iter().find(|(_, n)| {
state.outputs.iter().find(|(_, _, n)| {
n.as_deref()
.is_some_and(|n| n.eq_ignore_ascii_case(connector))
})
})
.map(|(o, _)| o.clone())
.map(|(_, o, _)| o.clone())
.ok_or_else(|| {
if named.is_empty() {
anyhow!(
@@ -1285,20 +1573,14 @@ fn run_existing(
"KWin: recording an existing output; awaiting PipeWire node"
);
let node_id = loop {
queue
.blocking_dispatch(&mut state)
.context("wayland dispatch (awaiting created)")?;
if let Some(node) = state.node_id {
break node;
}
if let Some(e) = state.failed.take() {
bail!("stream_output failed: {e}");
}
if state.closed {
bail!("KWin closed the stream before it was created");
}
};
let node_id = await_created(
&conn,
&mut queue,
&mut state,
stop,
"stream_output",
started,
)?;
setup_tx
.send(Ok(node_id))
.map_err(|_| anyhow!("monitor-mirror opener went away"))?;
@@ -1317,14 +1599,19 @@ fn run(
setup_tx: &Sender<Result<u32, String>>,
stop: &AtomicBool,
) -> Result<()> {
// Same clock as the mirror path: one barrier here rather than two, but the create wait is
// bounded against the opener either way (see [`CREATE_BUDGET`]).
let started = Instant::now();
let conn = Connection::connect_to_env()
.context("connect to KWin Wayland (is WAYLAND_DISPLAY set to the KWin socket?)")?;
let mut queue = conn.new_event_queue();
let qh = queue.handle();
let _registry = conn.display().get_registry(&qh, ());
// `want_outputs` stays false: `stream_virtual_output` names its output by string, so this
// connection never needs a `wl_output` — and it lives for the whole session (see `State`).
let mut state = State::default();
queue.roundtrip(&mut state).context("registry roundtrip")?;
roundtrip_within(&conn, &mut queue, &mut state, stop, 1, "registry roundtrip")?;
let screencast = state.screencast.clone().ok_or_else(|| {
anyhow!(
@@ -1353,21 +1640,15 @@ fn run(
"KWin: requested virtual output; awaiting PipeWire node"
);
// Pump events until KWin reports the node id (or an error).
let node_id = loop {
queue
.blocking_dispatch(&mut state)
.context("wayland dispatch (awaiting created)")?;
if let Some(node) = state.node_id {
break node;
}
if let Some(e) = state.failed.take() {
bail!("stream_virtual_output failed: {e}");
}
if state.closed {
bail!("KWin closed the stream before it was created");
}
};
// Pump events until KWin reports the node id (or an error, or the budget).
let node_id = await_created(
&conn,
&mut queue,
&mut state,
stop,
"stream_virtual_output",
started,
)?;
setup_tx
.send(Ok(node_id))
.map_err(|_| anyhow!("virtual-output opener went away"))?;
@@ -1380,25 +1661,83 @@ fn run(
Ok(())
}
/// Keep the connection (and thus the stream) alive until told to stop, observing `closed`.
/// `blocking_dispatch` can't be interrupted, so poll the connection fd with a short timeout and
/// honor `stop` within ~200 ms. Shared by the virtual-output and monitor-mirror paths — for a
/// virtual output this connection IS the output's lifetime; for a mirror it is only the
/// recording's, and the monitor itself is untouched either way.
fn park_until_stopped(
/// Poll slice while waiting on the Wayland fd — the granularity at which `stop` and a deadline are
/// observed (matches `kwin_output_mgmt`'s `POLL_MS`).
const POLL_MS: i32 = 200;
/// Budget for one compositor roundtrip. Generous next to a healthy one (a few ms); it exists only
/// so a KWin that accepted the connection and then stopped serving cannot pin the calling thread —
/// which for [`probe`] is whatever thread the mgmt API answered a `/display/compositors` on, and
/// for [`run`] is the session's own bring-up.
const ROUNDTRIP_BUDGET: Duration = Duration::from_secs(3);
/// How long an opener ([`spawn_vout`](VirtualDisplay::create), [`stream_existing_output`]) waits
/// for the worker's first word before giving up on it.
const OPENER_BUDGET: Duration = Duration::from_secs(20);
/// Slack subtracted from [`OPENER_BUDGET`] to get the worker's own ceiling: enough for its error to
/// travel one `mpsc` send while the opener is still listening.
const WORKER_MARGIN: Duration = Duration::from_millis(500);
/// Budget for the `created` handshake (the PipeWire node id) — but only as a ceiling, because
/// what actually matters is that the WORKER gives up before its opener does, so the failure the
/// client sees is a REASON ("KWin never created the output") rather than a bare timeout with the
/// worker still parked behind it.
///
/// That is a property of the whole worker, not of this one step, and it cannot be had by comparing
/// this constant with [`OPENER_BUDGET`]: the two workers do a different amount of work before they
/// get here. [`run`] spends one [`ROUNDTRIP_BUDGET`] barrier, so 3 + 15 < 20 ✓ — but [`run_existing`]
/// needs TWO (the registry globals, then the `wl_output` property burst that carries the connector
/// name), so 3 + 3 + 15 = 21 s and the mirror path lost the property that the doc here once claimed
/// for both. Hence [`await_created`] takes the worker's start instant and bounds itself by whichever
/// comes first, this budget or the opener's deadline; adding a third barrier to some future worker
/// cannot silently break it again.
const CREATE_BUDGET: Duration = Duration::from_secs(15);
/// How a bounded pump ended.
enum Pumped {
/// The predicate held.
Done,
/// `stop` was set — the caller's output/recording was released while we waited.
Stopped,
/// The deadline passed first.
Expired,
}
/// Bounded manual event loop: dispatch what's queued, then poll the connection fd for up to
/// [`POLL_MS`] and read, until `done(&state)` holds, `stop` is set, or `deadline` passes.
///
/// This is the only way to wait on this connection. `blocking_dispatch` and `roundtrip` cannot be
/// interrupted and have no ceiling, so a compositor that stops answering turns any wait into a
/// permanently stuck thread — and on the host that thread is the session's, whose only way to end a
/// session is to return. `deadline: None` means "no ceiling", which is correct for exactly one
/// caller: [`park_until_stopped`], where the wait IS the output's lifetime.
fn pump_until(
conn: &Connection,
queue: &mut wayland_client::EventQueue<State>,
state: &mut State,
deadline: Option<Instant>,
stop: &AtomicBool,
output: &str,
node_id: u32,
) -> Result<()> {
while !stop.load(Ordering::Relaxed) {
done: impl Fn(&State) -> bool,
) -> Result<Pumped> {
loop {
queue.dispatch_pending(state).context("dispatch_pending")?;
if state.closed {
tracing::warn!(output = %output, node_id, "KWin closed the screencast stream");
break;
if done(state) {
return Ok(Pumped::Done);
}
if stop.load(Ordering::Relaxed) {
return Ok(Pumped::Stopped);
}
let timeout = match deadline {
Some(d) => {
let remaining = d.saturating_duration_since(Instant::now());
if remaining.is_zero() {
return Ok(Pumped::Expired);
}
(remaining.as_millis() as i64).clamp(0, i64::from(POLL_MS)) as i32
}
None => POLL_MS,
};
conn.flush().context("wayland flush")?;
let Some(guard) = conn.prepare_read() else {
continue; // events already queued — loop dispatches them
@@ -1411,19 +1750,109 @@ fn park_until_stopped(
// SAFETY: `&mut pfd` points at a single live, fully-initialized `libc::pollfd` on the stack, and
// the count `1` matches that one-element array, so `poll` reads `fd`/`events` and writes `revents`
// strictly within `pfd`. `pfd.fd` is the Wayland connection's fd, valid because `conn` (and the
// `prepare_read` guard) are alive across the call. `poll` blocks up to 200 ms and writes only
// `revents`; `pfd` outlives the synchronous call and aliases nothing (a fresh local).
let r = unsafe { libc::poll(&mut pfd, 1, 200) };
// `prepare_read` guard) are alive across the call. `poll` blocks up to `timeout` ms and writes
// only `revents`; `pfd` outlives the synchronous call and aliases nothing (a fresh local).
let r = unsafe { libc::poll(&mut pfd, 1, timeout) };
if r > 0 && (pfd.revents & libc::POLLIN) != 0 {
let _ = guard.read();
} // else: timeout or signal — drop the guard, re-check `stop`
} // else: timeout or signal — drop the guard, re-check `stop` and the deadline
}
}
/// A `wl_display.sync` barrier bounded by [`ROUNDTRIP_BUDGET`] — the replacement for
/// `EventQueue::roundtrip`, which waits on the socket with no ceiling. `serial` must be unique per
/// connection (callers number theirs from 1); `what` names the wait in the error.
fn roundtrip_within(
conn: &Connection,
queue: &mut wayland_client::EventQueue<State>,
state: &mut State,
stop: &AtomicBool,
serial: u32,
what: &str,
) -> Result<()> {
let qh = queue.handle();
let _cb = conn.display().sync(&qh, serial);
let deadline = Instant::now() + ROUNDTRIP_BUDGET;
match pump_until(conn, queue, state, Some(deadline), stop, |st| {
st.sync_done >= serial
})? {
Pumped::Done => Ok(()),
Pumped::Stopped => bail!("{what} abandoned — the stream was released while we waited"),
Pumped::Expired => bail!(
"KWin accepted the Wayland connection but did not answer the {what} within \
{ROUNDTRIP_BUDGET:?} the compositor is not serving this client"
),
}
}
/// Keep the connection (and thus the stream) alive until told to stop, observing `closed`.
/// Shared by the virtual-output and monitor-mirror paths — for a virtual output this connection IS
/// the output's lifetime; for a mirror it is only the recording's, and the monitor itself is
/// untouched either way. The only deadline-free [`pump_until`] in the file, for that reason.
fn park_until_stopped(
conn: &Connection,
queue: &mut wayland_client::EventQueue<State>,
state: &mut State,
stop: &AtomicBool,
output: &str,
node_id: u32,
) -> Result<()> {
match pump_until(conn, queue, state, None, stop, |st| st.closed)? {
Pumped::Done => {
tracing::warn!(output = %output, node_id, "KWin closed the screencast stream");
}
// `Expired` cannot happen without a deadline; `Stopped` is the ordinary teardown.
Pumped::Stopped | Pumped::Expired => {}
}
Ok(())
}
/// Wait for the `created` event carrying the PipeWire node id, bounded and interruptible by `stop`.
///
/// The loop this replaced was a bare `blocking_dispatch` with no deadline that never read `stop`:
/// a KWin that acknowledged `stream_virtual_output` and then never answered parked the worker
/// thread for good, and the opener's `recv_timeout` arm — which did not set `stop` either — left it
/// there holding a half-built output. `request` names the request in the error.
///
/// `started` is when the WORKER began, not when this wait did: the bound is the earlier of
/// [`CREATE_BUDGET`] and the opener's own deadline, so whatever the barriers before us consumed
/// comes out of this wait rather than out of the opener's patience (see [`CREATE_BUDGET`] for the
/// arithmetic that made a fixed budget wrong on the mirror path).
fn await_created(
conn: &Connection,
queue: &mut wayland_client::EventQueue<State>,
state: &mut State,
stop: &AtomicBool,
request: &str,
started: Instant,
) -> Result<u32> {
let began = Instant::now();
let deadline = (began + CREATE_BUDGET).min(started + OPENER_BUDGET - WORKER_MARGIN);
let settled = |st: &State| st.node_id.is_some() || st.failed.is_some() || st.closed;
match pump_until(conn, queue, state, Some(deadline), stop, settled)? {
// Node id first: a `closed` that arrives in the same burst as `created` is a stream that
// was made and then torn down, not a failure to make one.
Pumped::Done => match (state.node_id, state.failed.take()) {
(Some(node), _) => Ok(node),
(None, Some(e)) => bail!("{request} failed: {e}"),
(None, None) => bail!("KWin closed the stream before it was created"),
},
Pumped::Stopped => bail!("{request} abandoned — released before KWin created the stream"),
// Report the wait we actually got, not the budget we asked for — they differ whenever the
// opener's deadline was the tighter of the two, and a message naming 15 s after 11 s is the
// kind of thing that sends the next person hunting for a stall that never happened.
Pumped::Expired => bail!(
"KWin acknowledged {request} but never sent the PipeWire node within {:?}",
began.elapsed()
),
}
}
#[cfg(test)]
mod tests {
use super::{modes_from_json, pick_custom_mode, KModeRow, MANAGED_PREFIX};
use super::{
modes_from_json, monitors_from_kscreen_json, pick_custom_mode, KModeRow, MANAGED_PREFIX,
};
fn row(id: &str, w: u32, h: u32, hz: f64) -> KModeRow {
KModeRow {
@@ -1506,6 +1935,80 @@ mod tests {
assert!(modes_from_json(&doc, "Virtual-nope").is_empty());
}
/// The kscreen fallback for `monitors::list` must produce the same contract the in-process
/// reader promises: geometry from `pos` (the identity key), the mode in PIXELS with refresh in
/// mHz precise enough to keep 59.94 apart from 60, a DISABLED head still listed but zeroed
/// rather than invented, our own managed output flagged, and the list sorted by position.
#[test]
fn parses_a_kscreen_monitor_list() {
let doc: serde_json::Value = serde_json::from_str(
r#"{"outputs":[
{"id":2,"name":"HDMI-A-1","enabled":true,"priority":2,"scale":1,
"pos":{"x":1920,"y":0},"vendor":"ACME","model":"U2720Q",
"currentModeId":"m9","modes":[
{"id":"m9","size":{"width":1920,"height":1080},"refreshRate":59.94}]},
{"id":1,"name":"eDP-1","enabled":true,"priority":1,"scale":1.5,
"pos":{"x":0,"y":0},
"currentModeId":7,"modes":[
{"id":7,"size":{"width":3840,"height":2160},"refreshRate":120.0}]},
{"id":3,"name":"DP-3","enabled":false,"scale":1,"pos":{"x":0,"y":0},
"modes":[{"id":"z","size":{"width":2560,"height":1440},"refreshRate":60.0}]},
{"id":4,"name":"Virtual-punktfunk-7","enabled":true,"scale":1,
"pos":{"x":5760,"y":0},"currentModeId":"v1","modes":[
{"id":"v1","size":{"width":2560,"height":1440},"refreshRate":119.98}]}
]}"#,
)
.expect("fixture parses");
let mons = monitors_from_kscreen_json(&doc);
let by = |c: &str| {
mons.iter()
.find(|m| m.connector == c)
.unwrap_or_else(|| panic!("{c} missing"))
.clone()
};
// Sorted by desktop position, not by kscreen's own order.
let order: Vec<&str> = mons.iter().map(|m| m.connector.as_str()).collect();
assert_eq!(
order,
vec!["DP-3", "eDP-1", "HDMI-A-1", "Virtual-punktfunk-7"]
);
let edp = by("eDP-1");
// PIXELS, at the scale the desk actually runs — the whole point of `logical_size`.
assert_eq!((edp.width, edp.height), (3840, 2160));
assert_eq!(edp.scale, 1.5);
assert_eq!(edp.logical_size(), (2560.0, 1440.0));
assert!(edp.primary, "priority 1 is KWin's primary");
assert_eq!(edp.refresh_mhz, 120_000);
// 59.94 must survive as mHz; rounding to whole Hz here is the bug this guards.
assert_eq!(by("HDMI-A-1").refresh_mhz, 59_940);
assert_eq!(by("HDMI-A-1").description, "ACME U2720Q");
assert!(!by("HDMI-A-1").primary);
// Disabled: listed (so "why can't I pick it?" has an answer) with no invented mode.
let dark = by("DP-3");
assert!(!dark.enabled);
assert_eq!((dark.width, dark.height, dark.refresh_mhz), (0, 0, 0));
// Ours, and labelled by connector when the entry carries no make/model.
let ours = by("Virtual-punktfunk-7");
assert!(ours.managed);
assert_eq!(ours.description, "Virtual-punktfunk-7");
assert!(!by("eDP-1").managed);
}
/// A document with no `outputs` array (an error object, or a kscreen-doctor whose schema
/// changed) is an empty list, never a panic — the caller's own error path already covers "the
/// tool did not answer".
#[test]
fn a_malformed_kscreen_document_yields_no_monitors() {
assert!(monitors_from_kscreen_json(&serde_json::json!({})).is_empty());
assert!(monitors_from_kscreen_json(&serde_json::json!({"outputs": 7})).is_empty());
// An output with no name cannot be pinned or resolved, so it is dropped rather than
// reported under an empty connector.
assert!(
monitors_from_kscreen_json(&serde_json::json!({"outputs": [{"enabled": true}]}))
.is_empty()
);
}
/// Group-aware exclusive (§6.1): with two managed group members + a physical panel enabled,
/// exclusive disables ONLY the non-managed panel — never a sibling session's per-slot output
/// (the Stage-3 naming would otherwise make a 2nd exclusive session black out the 1st).
@@ -20,8 +20,6 @@
//! each output's name / enabled / priority / current-mode size, then build a
//! `kde_output_configuration_v2` and `apply()` it, waiting for `applied` / `failed`.
#![deny(clippy::undocumented_unsafe_blocks)]
use std::collections::HashMap;
use std::os::fd::{AsFd, AsRawFd};
use std::time::{Duration, Instant};
@@ -107,10 +105,13 @@ const OP_BUDGET: Duration = Duration::from_secs(3);
/// Poll slice while waiting on the Wayland fd (matches the keepalive loop's cadence in `kwin.rs`).
const POLL_MS: i32 = 100;
/// KWin's CVT generator aligns a custom mode's width DOWN to a multiple of this (libxcvt's cell
/// grain), so the mode it builds for a `set_custom_modes` request may be a few px narrower than
/// asked — matches `kwin::CVT_H_GRANULARITY`. Used when matching the generated mode back.
const CVT_H_GRANULARITY: u32 = 8;
// KWin's CVT generator aligns a custom mode's width DOWN to a multiple of `CVT_H_GRANULARITY`
// (libxcvt's cell grain), so the mode it builds for a `set_custom_modes` request may be a few px
// narrower than asked — used below when matching the generated mode back. IMPORTED, not re-declared:
// this and `MANAGED_PREFIX` used to be second copies of `kwin.rs`'s literals, each under prose
// asserting the two "match" — an assertion no compiler was checking, on the two values that decide
// which output is OURS and which mode is the one we asked for.
use crate::kwin::{CVT_H_GRANULARITY, MANAGED_PREFIX};
/// `kde_output_management_v2.set_replication_source` (and the device's `replication_source` event)
/// arrived in v13. wayland-rs does not range-check requests, so sending one to a lower-version bind
@@ -166,9 +167,19 @@ pub(crate) struct TopologyOutcome {
/// One output as read from `kde_output_device_v2`.
#[derive(Default, Clone)]
struct DeviceState {
/// The global `name` number (higher = more recently advertised) — used to pick the newest of two
/// same-named outputs during a supersede.
/// The global `name` number (higher = more recently advertised) — the primary newest-wins
/// tie-break between two same-named outputs during a supersede. **Zero for every device on
/// KWin ≥ 6.7**, which hands outputs out through `kde_output_device_registry_v2` instead of one
/// global per output: those carry no global name at all (see [`seq`](DeviceState::seq)).
global: u32,
/// Order in which THIS connection first saw the device, from 1. The tie-break of last resort
/// behind `global`: on the registry model every `global` is 0, so without this the `max_by_key`
/// below degrades to "whichever entry `HashMap` iteration happened to reach last" — and `HashMap`
/// is seeded per process, so the supersede resolve was a coin flip that could pick the
/// PREDECESSOR (same name, same size) and configure the output that is about to disappear.
/// Announce order is not a proof of newness — it is the compositor's own enumeration order — but
/// it is deterministic, which the hash order was not.
seq: u32,
name: Option<String>,
uuid: Option<String>,
enabled: bool,
@@ -204,6 +215,8 @@ struct State {
/// the life of the session (dropping it would end the announcements).
device_registry: Option<DeviceRegistry>,
devices: HashMap<ObjectId, DeviceState>,
/// Highest [`DeviceState::seq`] handed out so far — the announce counter.
next_device_seq: u32,
/// mode object id → `(width, height, refresh_mHz)`.
mode_dims: HashMap<ObjectId, (u32, u32, u32)>,
/// Highest `wl_callback` serial whose `done` has arrived — the barrier the pump waits on.
@@ -213,6 +226,46 @@ struct State {
failure_reason: Option<String>,
}
impl State {
/// The entry for a device, stamping its announce order ([`DeviceState::seq`]) the first time we
/// see it. Every path that creates a device entry goes through here — the two announce models
/// (per-output global, and the ≥ 6.7 registry) plus the event handler, which can race ahead of
/// both — so the counter really does reflect the order the devices arrived in.
fn device_entry(&mut self, id: ObjectId) -> &mut DeviceState {
// Disjoint field borrows: `entry` holds `devices`, the closure holds only the counter.
let next = &mut self.next_device_seq;
self.devices.entry(id).or_insert_with(|| {
*next += 1;
DeviceState {
seq: *next,
..Default::default()
}
})
}
/// Forget a `kde_output_device_mode_v2` the compositor has destroyed.
///
/// The protocol's `removed` event says the compositor destroys the object *immediately after*
/// sending it — and the event is NOT marked `type="destructor"`, so wayland-rs happily keeps the
/// proxy alive locally. Anything still holding that id would later hand it back to KWin
/// (`kde_output_configuration_v2.mode`) as a request against a dead object, which is a protocol
/// error: KWin kills the connection, the apply "fails", and a >60 Hz session degrades to the
/// kscreen-doctor path with a log indistinguishable from "this KWin is too old". Reachable
/// precisely because `set_custom_modes` REPLACES the persisted custom list, so the mode a
/// previous session left behind is destroyed the moment this session installs its own.
fn forget_mode(&mut self, id: &ObjectId) {
self.mode_dims.remove(id);
for dev in self.devices.values_mut() {
dev.modes.retain(|(mid, _)| mid != id);
if dev.current_mode.as_ref() == Some(id) {
// Don't invent a size for a destroyed mode: a resolve keyed on current dims must
// miss (and fall back) rather than match on a mode that no longer exists.
dev.current_mode = None;
}
}
}
}
impl Dispatch<WlRegistry, ()> for State {
fn event(
state: &mut Self,
@@ -239,7 +292,7 @@ impl Dispatch<WlRegistry, ()> for State {
// handler can record it (newest-wins tie-break during a supersede).
let dev = registry.bind::<OutputDevice, _, _>(name, v, qh, name);
let id = dev.id();
state.devices.entry(id).or_default().proxy = Some(dev);
state.device_entry(id).proxy = Some(dev);
} else if interface == DeviceRegistry::interface().name {
// KWin ≥ 6.7 (Plasma 6.7.3 verified) no longer advertises ONE
// `kde_output_device_v2` global per output — it advertises this registry and
@@ -260,9 +313,14 @@ impl Dispatch<WlRegistry, ()> for State {
}
/// The device registry hands out one `kde_output_device_v2` per output via its `output` event
/// (a `new_id`, so the child is created by the `event_created_child!` binding below). Devices that
/// arrive this way have no global `name` number — the newest-wins supersede tie-break uses 0 for
/// them, which is fine: that tie-break only matters for the per-output-global model.
/// (a `new_id`, so the child is created by the `event_created_child!` binding below).
///
/// Devices that arrive this way have no global `name` number — the `0u32` UserData below is stamped
/// on every one of them, so [`DeviceState::global`] is 0 across the board. That is **not** harmless,
/// and an earlier comment here claimed it was: the registry model is what CURRENT KWin uses, so the
/// newest-wins supersede tie-break is unavailable exactly where it is needed (two same-named,
/// same-sized outputs, predecessor still alive). [`DeviceState::seq`] is the deterministic
/// fallback the tie-break actually lands on there.
impl Dispatch<DeviceRegistry, ()> for State {
fn event(
state: &mut Self,
@@ -274,7 +332,7 @@ impl Dispatch<DeviceRegistry, ()> for State {
) {
if let RegistryEvent::Output { output } = event {
let id = output.id();
state.devices.entry(id).or_default().proxy = Some(output);
state.device_entry(id).proxy = Some(output);
}
}
@@ -319,7 +377,23 @@ impl Dispatch<OutputDevice, u32> for State {
_: &Connection,
_: &QueueHandle<Self>,
) {
let entry = state.devices.entry(device.id()).or_default();
// Before anything re-creates the entry: `removed` (device ≥ v21, and we bind up to 24) means
// this output is gone for good and no further update will arrive. Dropping it keeps a
// hot-unplugged head from being resolved, disabled or "restored" minutes later, and the XML
// asks the client to `release` the object — the only way the server-side one is ever freed,
// since wayland-rs sends no destructor when a proxy is merely dropped.
if matches!(event, DeviceEvent::Removed) {
if let Some(dead) = state.devices.remove(&device.id()) {
for (mid, _) in &dead.modes {
state.mode_dims.remove(mid);
}
}
if device.version() >= 21 {
device.release();
}
return;
}
let entry = state.device_entry(device.id());
entry.global = *global;
if entry.proxy.is_none() {
entry.proxy = Some(device.clone());
@@ -363,6 +437,12 @@ impl Dispatch<DeviceMode, ()> for State {
_: &Connection,
_: &QueueHandle<Self>,
) {
// `removed` first, and NOT through the entry below: re-inserting a destroyed mode is exactly
// the stale row a later `config.mode(...)` would send back to KWin (see [`State::forget_mode`]).
if matches!(event, ModeEvent::Removed) {
state.forget_mode(&mode.id());
return;
}
let entry = state.mode_dims.entry(mode.id()).or_insert((0, 0, 0));
match event {
ModeEvent::Size { width, height } => {
@@ -370,6 +450,7 @@ impl Dispatch<DeviceMode, ()> for State {
entry.1 = height.max(0) as u32;
}
ModeEvent::Refresh { refresh } => entry.2 = refresh.max(0) as u32,
// `preferred` / `flags` / `cvt` carry nothing we drive an apply from.
_ => {}
}
}
@@ -419,13 +500,76 @@ struct Session {
next_sync: u32,
}
/// Why [`Session::open`] declined, i.e. why this operation degraded to the `kscreen-doctor`
/// shell-out.
///
/// The distinction is the whole value of the type: a bare `None` made every one of these read as
/// "not a KDE box", which is how a genuine regression — KWin ≥ 6.7 no longer advertising per-output
/// `kde_output_device_v2` globals, so the device list came back EMPTY — shipped as a fallback that
/// fired on every current KDE machine with nothing in the log to say so.
enum OpenFailure {
/// No Wayland connection at all (`WAYLAND_DISPLAY` unset/stale) — not a session we can drive.
Connect(String),
/// The compositor accepted the connection but did not answer the registry barrier in budget:
/// the wedge case this whole module exists for.
RegistryBarrier,
/// Connected and answering, but `kde_output_management_v2` is not advertised to this client
/// (too old a KWin, or not KWin at all).
NoManagementGlobal,
/// Management is there, but the outputs' own property bursts never completed in budget.
DeviceBarrier,
}
impl std::fmt::Display for OpenFailure {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
OpenFailure::Connect(e) => write!(f, "no Wayland connection ({e})"),
OpenFailure::RegistryBarrier => {
write!(
f,
"the compositor did not answer the registry roundtrip in budget"
)
}
OpenFailure::NoManagementGlobal => {
write!(
f,
"kde_output_management_v2 is not advertised to this client"
)
}
OpenFailure::DeviceBarrier => {
write!(
f,
"the outputs never finished announcing their state in budget"
)
}
}
}
}
impl Session {
/// [`Session::connect`] for the operation named by `op`, logging the reason on the way out.
///
/// One log site for all six callers: every one of them silently degraded to `kscreen-doctor`
/// before, so on a box where the in-process path never worked the only symptom was that
/// topology took ~26 s and nothing said why.
fn open(op: &'static str) -> Result<Session, OpenFailure> {
let opened = Session::connect();
if let Err(reason) = &opened {
tracing::warn!(
op,
%reason,
"KWin in-process output management unavailable — falling back to kscreen-doctor"
);
}
opened
}
/// Connect to the KWin Wayland socket, bind `kde_output_management_v2` + every
/// `kde_output_device_v2`, and read each output's state — all bounded by `OP_BUDGET`. `None` if
/// we can't connect, the management global isn't advertised, or the compositor doesn't answer in
/// budget (the wedge case — the caller then falls back to `kscreen-doctor`).
fn open() -> Option<Session> {
let conn = Connection::connect_to_env().ok()?;
/// `kde_output_device_v2`, and read each output's state — all bounded by `OP_BUDGET`. The
/// [`OpenFailure`] says which rung declined; every one of them sends the caller to
/// `kscreen-doctor`.
fn connect() -> Result<Session, OpenFailure> {
let conn = Connection::connect_to_env().map_err(|e| OpenFailure::Connect(e.to_string()))?;
let queue = conn.new_event_queue();
let qh = queue.handle();
let _registry = conn.display().get_registry(&qh, ());
@@ -438,19 +582,15 @@ impl Session {
let deadline = Instant::now() + OP_BUDGET;
// Phase 1: process the registry globals (binds management + every device in the handler).
if !s.sync_barrier(deadline) {
return None;
return Err(OpenFailure::RegistryBarrier);
}
if s.state.management.is_none() {
tracing::debug!(
"KWin does not advertise kde_output_management_v2 to this client — kscreen-doctor \
fallback"
);
return None;
return Err(OpenFailure::NoManagementGlobal);
}
// Phase 2: flush the device binds issued in phase 1 and drain each output's state burst
// (name / enabled / priority / current_mode / mode sizes / done).
if !s.sync_barrier(deadline) {
return None;
return Err(OpenFailure::DeviceBarrier);
}
// Phase 3 (KWin ≥ 6.7, the registry model): the devices themselves only arrive as the
// registry's `output` events during phase 2, so their property bursts are one round further
@@ -460,9 +600,9 @@ impl Session {
&& s.state.devices.values().any(|d| !d.seen_done)
&& !s.sync_barrier(deadline)
{
return None;
return Err(OpenFailure::DeviceBarrier);
}
Some(s)
Ok(s)
}
/// Send a `wl_display.sync` and pump the queue until its `done` arrives or `deadline` passes.
@@ -554,6 +694,26 @@ impl Session {
let id = dev.current_mode.as_ref()?;
self.state.mode_dims.get(id).copied()
}
/// Resolve OUR just-created virtual output: a managed-prefix name AND a current size equal to
/// the size we created it at — only the just-created output sits there during a supersede,
/// because the replacement deliberately reuses the per-slot name while the predecessor is still
/// alive. Newest wins the remaining tie: the global `name` number where there is one, else
/// announce order (see [`DeviceState::seq`] — on KWin ≥ 6.7 that is every device).
///
/// One resolve for all three operations (topology / de-mirror / custom mode). They had drifted
/// into three copies of the same filter, which is how a tie-break fix lands in two of them.
fn resolve_ours(&self, our_prefix: &str, our_w: u32, our_h: u32) -> Option<DeviceState> {
self.state
.devices
.values()
.filter(|d| {
d.name.as_deref().is_some_and(|n| n.starts_with(our_prefix))
&& self.current_dims(d).map(|(w, h, _)| (w, h)) == Some((our_w, our_h))
})
.max_by_key(|d| (d.global, d.seq))
.cloned()
}
}
/// `(width, height, "WxH@Hz")` capture of a device's current mode, Hz rounded — the same shape the
@@ -563,10 +723,10 @@ fn mode_spec(dims: (u32, u32, u32)) -> String {
format!("{}x{}@{}", dims.0, dims.1, hz)
}
/// Prefix EVERY managed KWin output shares (mirrors `kwin::MANAGED_PREFIX`) — the streamed outputs
/// are `Virtual-punktfunk` / `Virtual-punktfunk-<id>`, so a same-family sibling session is never
/// treated as a physical to disable, and its primary is never stolen (first-slot-wins).
const MANAGED_PREFIX: &str = "Virtual-punktfunk";
// `MANAGED_PREFIX` — the prefix EVERY managed KWin output shares (`Virtual-punktfunk` /
// `Virtual-punktfunk-<id>`), so a same-family sibling session is never treated as a physical to
// disable and its primary is never stolen (first-slot-wins) — is imported at the top of this file
// from `kwin.rs`, which owns the naming.
/// Every head KWin reports, for [`crate::monitors::list`].
///
@@ -575,12 +735,8 @@ const MANAGED_PREFIX: &str = "Virtual-punktfunk";
/// burst is skipped rather than reported half-read (its geometry would be a guess, and geometry is
/// exactly what callers key on).
pub(crate) fn list_monitors() -> anyhow::Result<Vec<crate::monitors::PhysicalMonitor>> {
let session = Session::open().ok_or_else(|| {
anyhow::anyhow!(
"KWin did not answer kde_output_management_v2 (not a KWin session, the protocol is \
not advertised to this client, or the compositor is wedged)"
)
})?;
let session = Session::open("list_monitors")
.map_err(|e| anyhow::anyhow!("KWin did not answer kde_output_management_v2: {e}"))?;
let mut out: Vec<_> = session
.state
.devices
@@ -630,24 +786,12 @@ pub(crate) fn apply_topology(
disabled: Vec::new(),
handled: false,
};
let Some(mut sess) = Session::open() else {
let Ok(mut sess) = Session::open("topology") else {
return miss();
};
let deadline = Instant::now() + OP_BUDGET;
// Resolve OUR output: managed-prefix name AND current size == the birth size (only the
// just-created output sits there during a supersede); newest global wins the tie.
let ours = sess
.state
.devices
.values()
.filter(|d| {
d.name.as_deref().is_some_and(|n| n.starts_with(our_prefix))
&& sess.current_dims(d).map(|(w, h, _)| (w, h)) == Some((our_w, our_h))
})
.max_by_key(|d| d.global)
.cloned();
let Some(ours) = ours else {
let Some(ours) = sess.resolve_ours(our_prefix, our_w, our_h) else {
tracing::warn!(
our_prefix,
our_w,
@@ -846,7 +990,7 @@ pub(crate) fn apply_topology(
/// which is broken under every topology equally. So this reads the state and applies **only** when
/// our output really is mirroring; the ordinary session pays one bounded enumerate and no apply.
pub(crate) fn clear_replication_source(our_prefix: &str, our_w: u32, our_h: u32) {
let Some(mut sess) = Session::open() else {
let Ok(mut sess) = Session::open("clear_replication_source") else {
return;
};
let deadline = Instant::now() + OP_BUDGET;
@@ -858,18 +1002,7 @@ pub(crate) fn clear_replication_source(our_prefix: &str, our_w: u32, our_h: u32)
if mgmt_version < REPLICATION_SOURCE_SINCE {
return;
}
// Same resolve as `apply_topology`: managed-prefix name AND the birth size, newest global wins.
let Some(ours) = sess
.state
.devices
.values()
.filter(|d| {
d.name.as_deref().is_some_and(|n| n.starts_with(our_prefix))
&& sess.current_dims(d).map(|(w, h, _)| (w, h)) == Some((our_w, our_h))
})
.max_by_key(|d| d.global)
.cloned()
else {
let Some(ours) = sess.resolve_ours(our_prefix, our_w, our_h) else {
return;
};
if !is_mirroring(ours.replication_source.as_deref()) {
@@ -918,7 +1051,7 @@ pub(crate) fn set_custom_mode(
want_h: u32,
want_hz: u32,
) -> Option<(u32, u32, u32)> {
let mut sess = Session::open()?;
let mut sess = Session::open("custom_mode").ok()?;
let deadline = Instant::now() + OP_BUDGET;
// `set_custom_modes` is `since 18`; calling it on an older bound management object is a protocol
@@ -928,16 +1061,9 @@ pub(crate) fn set_custom_mode(
return None;
}
// Resolve our output at its birth size (newest global wins a supersede).
// Resolve our output at its birth size (newest wins a supersede — see `resolve_ours`).
let our_proxy = sess
.state
.devices
.values()
.filter(|d| {
d.name.as_deref().is_some_and(|n| n.starts_with(our_prefix))
&& sess.current_dims(d).map(|(w, h, _)| (w, h)) == Some((birth_w, birth_h))
})
.max_by_key(|d| d.global)
.resolve_ours(our_prefix, birth_w, birth_h)
.and_then(|d| d.proxy.clone())?;
let our_key = our_proxy.id();
@@ -985,10 +1111,15 @@ pub(crate) fn set_custom_mode(
}
// Grab the generated mode's proxy, then select it (this is what changes the size).
// Newest match wins: `modes` is in announce order, and the entry we just had KWin generate is
// the last one. An earlier session's identical custom mode may still be listed here — KWin only
// destroys it (`kde_output_device_mode_v2.removed`) when it processes our `set_custom_modes`,
// and that removal may not have been dispatched yet.
let mode_proxy = {
let dev = sess.state.devices.get(&our_key)?;
dev.modes
.iter()
.rev()
.find(|(mid, _)| mode_matches(&sess.state, mid))
.map(|(_, p)| p.clone())?
};
@@ -1031,19 +1162,31 @@ pub(crate) fn set_custom_mode(
}
/// Re-enable outputs by name at their captured `WxH@Hz` modes (teardown), in-process. Returns
/// `true` if the config applied; `false` (compositor unresponsive / management absent) tells the
/// caller to fall back to `kscreen-doctor`.
/// `true` only if EVERY requested output was staged and the config applied; `false` (compositor
/// unresponsive, management absent, or an output we could not address) tells the caller to fall
/// back to `kscreen-doctor`.
///
/// The "every requested output" half is load-bearing, not pedantry. The names in `outputs` were
/// captured on a DIFFERENT connection during [`apply_topology`] and this restore opens a fresh
/// session minutes later, when the display group's last member drops — so a name that no longer
/// resolves is a live possibility. An empty `kde_output_configuration_v2` still gets an `applied`
/// event, so returning the apply verdict alone reported SUCCESS for a total no-op, suppressed the
/// `reenable_outputs_kscreen` backstop, and left a physical monitor dark.
pub(crate) fn reenable_outputs(outputs: &[(String, String)]) -> bool {
if outputs.is_empty() {
return true;
}
let Some(mut sess) = Session::open() else {
let Ok(mut sess) = Session::open("restore_outputs") else {
return false;
};
let deadline = Instant::now() + OP_BUDGET;
let config = sess.new_config();
let mut matched = 0usize;
for (name, spec) in outputs {
// Find the device by name (physical names are stable across a session).
// Find the device by name (physical names are stable across a session). BOTH misses below
// leave `matched` un-incremented, the proxy one included: a `DeviceState` can be created by
// the event handler ([`State::device_entry`]) and carry a name before the announce that
// records its proxy has been dispatched, and a name with no proxy is not addressable.
let Some(dev) = sess
.state
.devices
@@ -1056,6 +1199,7 @@ pub(crate) fn reenable_outputs(outputs: &[(String, String)]) -> bool {
let Some(proxy) = dev.proxy.as_ref() else {
continue;
};
matched += 1;
// Enable first — a bare enable always succeeds, so a physical is never left dark.
config.enable(proxy, 1);
// Then re-assert the captured mode so a 120 Hz panel doesn't return at KWin's ~60 Hz default.
@@ -1063,18 +1207,39 @@ pub(crate) fn reenable_outputs(outputs: &[(String, String)]) -> bool {
config.mode(proxy, &mode);
}
}
if matched == 0 {
// Nothing staged: applying would ack an empty config and read as success. Hand the whole
// restore to kscreen-doctor, which addresses outputs by name and needs no live proxy.
config.destroy();
tracing::warn!(
requested = ?outputs,
"KWin output management: none of the outputs to restore are addressable on this \
connection kscreen-doctor fallback"
);
return false;
}
let ok = sess.apply(&config, deadline);
config.destroy();
if ok {
let complete = ok && matched == outputs.len();
if complete {
tracing::info!(reenabled = ?outputs, "KWin output management: restored outputs (in-process)");
} else {
tracing::warn!(
requested = ?outputs,
matched,
applied = ok,
reason = ?sess.state.failure_reason,
"KWin output management: restore incomplete — kscreen-doctor backstop takes the rest \
(an output left disabled is a physical left dark)"
);
}
ok
complete
}
/// Position the output identified by `uuid` at `(x, y)` in the desktop layout, in-process. Returns
/// `true` if applied; `false` tells the caller to fall back to `kscreen-doctor`.
pub(crate) fn set_position(uuid: &str, x: i32, y: i32) -> bool {
let Some(mut sess) = Session::open() else {
let Ok(mut sess) = Session::open("position") else {
return false;
};
let deadline = Instant::now() + OP_BUDGET;
+160 -33
View File
@@ -122,8 +122,14 @@ impl MutterDisplay {
/// `XDG_SESSION_DESKTOP` alongside would resurrect the bug that scrub exists to prevent — a stale
/// `gnome` there after a gnome-shell crash reports Mutter usable and routes the next client into a
/// dead session (45 s create timeouts instead of a crisp handshake error).
///
/// The read takes [`crate::with_env_lock`]: this runs on a management worker (`/host/compositors` →
/// [`crate::available`]) concurrently with another connect's `apply_session_env`, which `set_var`s
/// this key for a live session and `remove_var`s it when nothing is — and a glibc `getenv` racing
/// that is the `environ` realloc data race ENV_LOCK exists for, torn answer at best and a host
/// segfault mid-connect at worst. Read-then-drop; no caller holds the lock (it is not reentrant).
pub fn is_available() -> bool {
std::env::var("XDG_CURRENT_DESKTOP")
crate::with_env_lock(|| std::env::var("XDG_CURRENT_DESKTOP"))
.map(|d| d.to_ascii_uppercase().contains("GNOME"))
.unwrap_or(false)
}
@@ -718,13 +724,24 @@ async fn connect(
}
// ---------------------------------------------------------------------------------------------
// Optional: make the per-session virtual output the PRIMARY monitor (PUNKTFUNK_MUTTER_VIRTUAL_PRIMARY).
// Optional: make the per-session virtual output the PRIMARY monitor.
//
// `RecordVirtual` adds the virtual monitor as an *extended* desktop. On a headless host that's the
// only display, so the shell + windows live there. But when a physical monitor is attached, GNOME
// keeps it primary and the virtual output is an empty extension — the stream shows only the
// wallpaper. We fix that by promoting the virtual output to primary (physical kept on, secondary)
// via `org.gnome.Mutter.DisplayConfig.ApplyMonitorsConfig`, and restore on teardown.
// wallpaper. We fix that by promoting the virtual output via
// `org.gnome.Mutter.DisplayConfig.ApplyMonitorsConfig`.
//
// Which shape is `crate::effective_topology()`'s call, not this module's: the console policy first,
// then the legacy `PUNKTFUNK_{KWIN,MUTTER}_VIRTUAL_PRIMARY` env, then the Auto default. `Primary`
// keeps the physicals on as secondaries; `Exclusive` omits them, so Mutter disables them for the
// session; `Extend` skips this block entirely.
//
// Applied at APPLY_TEMPORARY, and **MUTTER ITSELF REVERTS IT** when the virtual monitor disappears
// and our DisplayConfig connection closes. We must never re-assert the layout on teardown: the
// banner used to promise a "restore on teardown" that the teardown deliberately does not do, and
// issuing that ApplyMonitorsConfig is what SIGSEGVed gnome-shell on Mutter 50 + NVIDIA and wedged a
// box at the GDM greeter (see the teardown comment in `session_thread`).
// ---------------------------------------------------------------------------------------------
/// `org.gnome.Mutter.DisplayConfig.GetCurrentState` reply shapes (see the interface XML):
@@ -811,7 +828,9 @@ fn current_mode(state: &CurrentState, connector: &str) -> Option<(String, i32, i
/// Pure mode-pick for a KEPT physical (unit-tested). Given the physical's PRE-connect mode
/// (`pre_mode = (id, w, h, refresh)`; `None` when the connector is new since the snapshot) and the
/// mode list Mutter reports for it in the POST-virtual state
/// (`(id, w, h, refresh, is_current, is_preferred)`), return the `(mode_id, width)` to re-apply.
/// (`(id, w, h, refresh, is_current, is_preferred)`), return the `(mode_id, width, height)` to
/// re-apply. The height is not decoration: a head rotated 90°/270° is as wide on the desktop as its
/// mode is tall, and the caller lays the kept heads out side by side.
///
/// Mutter re-derives its layout when the `RecordVirtual` output appears and can silently drop a
/// 120 Hz panel to its EDID-preferred 60 Hz — so the post-virtual `is-current` is *already* 60 Hz.
@@ -821,40 +840,40 @@ fn current_mode(state: &CurrentState, connector: &str) -> Option<(String, i32, i
fn pick_keep_mode(
pre_mode: Option<(String, i32, i32, f64)>,
state_modes: &[(String, i32, i32, f64, bool, bool)],
) -> Option<(String, i32)> {
) -> Option<(String, i32, i32)> {
let state_current = || {
state_modes
.iter()
.find(|m| m.4)
.or_else(|| state_modes.iter().find(|m| m.5))
.or_else(|| state_modes.first())
.map(|m| (m.0.clone(), m.1))
.map(|m| (m.0.clone(), m.1, m.2))
};
let Some((pre_id, w, h, hz)) = pre_mode else {
return state_current();
};
// The exact pre mode id, if the connector still offers it (same session ⇒ usually true).
if state_modes.iter().any(|m| m.0 == pre_id) {
return Some((pre_id, w));
return Some((pre_id, w, h));
}
// Else a re-keyed id with the same geometry + refresh (still the real 120 Hz).
if let Some(m) = state_modes
.iter()
.find(|m| m.1 == w && m.2 == h && (m.3 - hz).abs() < 0.5)
{
return Some((m.0.clone(), m.1));
return Some((m.0.clone(), m.1, m.2));
}
// The physical genuinely no longer offers that mode — use whatever is valid now.
state_current()
}
/// The `(mode_id, width)` a kept physical should be RE-APPLIED at — its PRE-connect mode preserved
/// across Mutter's virtual-output layout re-derive. See [`pick_keep_mode`].
/// The `(mode_id, width, height)` a kept physical should be RE-APPLIED at — its PRE-connect mode
/// preserved across Mutter's virtual-output layout re-derive. See [`pick_keep_mode`].
fn physical_keep_mode(
pre: &CurrentState,
state: &CurrentState,
conn: &str,
) -> Option<(String, i32)> {
) -> Option<(String, i32, i32)> {
let pre_mode = current_mode_full(pre, conn);
let state_modes: Vec<(String, i32, i32, f64, bool, bool)> = state
.1
@@ -1044,13 +1063,57 @@ fn snap_integral_scale(want: f64, width: u32, height: u32) -> f64 {
.unwrap_or(want)
}
/// The scale of the logical monitor carrying `connector`, if present.
fn logical_scale(state: &CurrentState, connector: &str) -> Option<f64> {
/// The `(scale, transform)` of the logical monitor carrying `connector`. `None` means **no logical
/// monitor carries it** — which is how Mutter reports a head the operator has DISABLED, and is the
/// distinction [`keep_head_layout`] turns into "leave it off".
fn logical_placement(state: &CurrentState, connector: &str) -> Option<(f64, u32)> {
state
.2
.iter()
.find(|l| l.5.iter().any(|spec| spec.0 == connector))
.map(|l| l.2)
.map(|l| (l.2, l.3))
}
/// The scale of the logical monitor carrying `connector`, if present.
fn logical_scale(state: &CurrentState, connector: &str) -> Option<f64> {
logical_placement(state, connector).map(|(scale, _)| scale)
}
/// Whether a kept physical should be re-applied at all, and with what `(scale, transform)`. Pure —
/// unit-tested, because getting it wrong is invisible on a headless lab box and very visible on the
/// operator's desk.
///
/// The rebuild used to hardcode `scale = 1.0`, `transform = 0` and to list every connector Mutter
/// reported, so one connect un-rotated a portrait panel, dropped a 2×-scaled 4K head to native
/// pixels, and switched a deliberately-dark monitor back on. All three facts are in the PRE-connect
/// snapshot: `pre_logical` is the head's logical-monitor entry there, and Mutter reports a disabled
/// head by omitting it from `logical_monitors` entirely. So: carry the pre values when the head was
/// on; leave it out when the connector existed pre-connect and carried no logical monitor (disabled
/// on purpose); and for a connector that was not in the snapshot at all — a hotplug inside our
/// window — keep it on at whatever Mutter has just derived for it, which is the friendlier reading
/// of "the operator plugged this in while we were connecting".
fn keep_head_layout(
existed_pre: bool,
pre_logical: Option<(f64, u32)>,
state_logical: Option<(f64, u32)>,
) -> Option<(f64, u32)> {
// A non-finite or non-positive scale would fail the whole ApplyMonitorsConfig, taking the
// primary switch down with it.
let sane = |(scale, transform): (f64, u32)| {
(
if scale.is_finite() && scale > 0.0 {
scale
} else {
1.0
},
transform,
)
};
match (pre_logical, existed_pre) {
(Some(l), _) => Some(sane(l)),
(None, true) => None,
(None, false) => Some(sane(state_logical.unwrap_or((1.0, 0)))),
}
}
/// Every head Mutter reports, for [`crate::monitors::list`].
@@ -1142,16 +1205,20 @@ fn build_exclusive_config(vconn: &str, vmode: &str, scale: f64) -> Vec<ApplyLogi
)]
}
/// **Primary** — the virtual output primary at `(0, 0)`, with every currently-active physical
/// monitor KEPT as a secondary (laid left-to-right past the virtual, each at its **pre-connect**
/// mode). So the shell + new windows land on the streamed surface, but the operator's physical
/// screen stays on **at its real refresh**. On a headless host (no physicals) this is identical to
/// [`build_exclusive_config`].
/// **Primary** — the virtual output primary at `(0, 0)`, with every physical monitor the operator
/// had ENABLED kept as a secondary (laid left-to-right past the virtual, each at its **pre-connect**
/// mode, scale and transform). So the shell + new windows land on the streamed surface, but the
/// operator's physical screen stays exactly as they left it. On a headless host (no physicals) this
/// is identical to [`build_exclusive_config`].
///
/// `pre` is the snapshot taken *before* the virtual output existed (physical still at its true
/// refresh); `state` is the post-virtual state. We read each physical's mode from `pre` because
/// Mutter can knock a 120 Hz panel down to 60 Hz when it re-derives the layout for the virtual
/// monitor — reading `state` would cement that 60 Hz (`physical_keep_mode`).
/// refresh); `state` is the post-virtual state. Everything about a kept head is read from `pre`,
/// because the post-virtual state is already contaminated: Mutter re-derives the layout when the
/// `RecordVirtual` output appears and can knock a 120 Hz panel down to 60 Hz, so reading `state`
/// would cement that 60 Hz (`physical_keep_mode`). Scale, transform and enabled-ness come from the
/// same snapshot for the same reason — and because rebuilding them from scratch is what used to
/// un-rotate portrait panels, flatten a 2× scale and re-light a head the operator had switched off
/// ([`keep_head_layout`]).
///
/// *Physical-keep is unvalidated on-glass* — the lab boxes are headless (no attached display to keep
/// on); the layout math is conservative (append to the right) but wants a display-attached box.
@@ -1190,16 +1257,42 @@ fn build_primary_keeping_physicals(
if conn == vconn {
continue;
}
if let Some((mode_id, w)) = physical_keep_mode(pre, state, conn) {
let existed_pre = pre.1.iter().any(|m| m.0 .0 == *conn);
let Some((head_scale, transform)) = keep_head_layout(
existed_pre,
logical_placement(pre, conn),
logical_placement(state, conn),
) else {
// Omitted from the config ⇒ Mutter leaves it disabled, which is what the operator asked
// for. Listing it would switch their dark head on for the length of the session.
tracing::debug!(
connector = %conn,
"mutter: this head was disabled before the session — leaving it disabled"
);
continue;
};
if let Some((mode_id, w, h)) = physical_keep_mode(pre, state, conn) {
logicals.push((
x,
0,
1.0,
0,
head_scale,
transform,
false,
vec![(conn.clone(), mode_id, HashMap::new())],
));
x += w.max(0);
// Advance by the head's own LOGICAL footprint, in the layout's coordinate space — the
// same space the virtual's advance above uses. A 3840-wide panel at scale 2 occupies
// 1920, and a head rotated 90°/270° (transform 1/3, or their flipped twins 5/7) is as
// wide as its mode is TALL. Advancing by raw mode width was only ever *consistent* with
// the forced scale of 1.0 this rebuild used to apply; preserving the real scale without
// this would just trade one wrong layout for another (overlapping or gapped heads).
let rotated = matches!(transform, 1 | 3 | 5 | 7);
let footprint = if rotated { h } else { w };
x += if physical_layout {
footprint.max(0)
} else {
((footprint as f64 / head_scale).round() as i32).max(0)
};
}
}
logicals
@@ -1207,7 +1300,10 @@ fn build_primary_keeping_physicals(
#[cfg(test)]
mod tests {
use super::{pick_keep_mode, pick_virtual, snap_integral_scale, HashMap, Mode, MonitorInfo};
use super::{
keep_head_layout, pick_keep_mode, pick_virtual, snap_integral_scale, HashMap, Mode,
MonitorInfo,
};
// (id, w, h, refresh, is_current, is_preferred)
fn m(
@@ -1232,7 +1328,7 @@ mod tests {
];
assert_eq!(
pick_keep_mode(pre, &state),
Some(("M120".to_string(), 2560))
Some(("M120".to_string(), 2560, 1440))
);
}
@@ -1247,7 +1343,7 @@ mod tests {
];
assert_eq!(
pick_keep_mode(pre, &state),
Some(("new-120".to_string(), 2560))
Some(("new-120".to_string(), 2560, 1440))
);
}
@@ -1262,7 +1358,7 @@ mod tests {
];
assert_eq!(
pick_keep_mode(pre, &state),
Some(("s-100".to_string(), 3440))
Some(("s-100".to_string(), 3440, 1440))
);
}
@@ -1289,7 +1385,10 @@ mod tests {
m("A", 1920, 1080, 60.0, true, false),
m("B", 1920, 1080, 144.0, false, true),
];
assert_eq!(pick_keep_mode(None, &state), Some(("A".to_string(), 1920)));
assert_eq!(
pick_keep_mode(None, &state),
Some(("A".to_string(), 1920, 1080))
);
let no_current = vec![
m("A", 1920, 1080, 60.0, false, false),
@@ -1297,7 +1396,35 @@ mod tests {
];
assert_eq!(
pick_keep_mode(None, &no_current),
Some(("B".to_string(), 1920))
Some(("B".to_string(), 1920, 1080))
);
}
/// A kept physical must come back exactly as the operator had it. Rebuilding the layout from
/// scratch (`scale = 1.0`, `transform = 0`, every connector listed) un-rotated portrait panels,
/// flattened a 2× scale, and switched a deliberately-dark head back on the moment a client
/// connected — while the code went to real trouble to preserve the refresh.
#[test]
fn a_kept_head_carries_its_pre_connect_scale_and_transform() {
// Rotated + 2×-scaled, exactly as it was before the virtual output appeared.
assert_eq!(
keep_head_layout(true, Some((2.0, 1)), Some((1.0, 0))),
Some((2.0, 1))
);
// Disabled on purpose (present pre-connect, carried by no logical monitor) — stays off.
assert_eq!(keep_head_layout(true, None, Some((1.0, 0))), None);
// Hotplugged inside our window: not in the snapshot at all, so keep it on at whatever
// Mutter derived rather than disabling a monitor the operator just plugged in.
assert_eq!(
keep_head_layout(false, None, Some((1.5, 2))),
Some((1.5, 2))
);
assert_eq!(keep_head_layout(false, None, None), Some((1.0, 0)));
// A junk scale would fail the WHOLE ApplyMonitorsConfig, taking the primary switch with it.
assert_eq!(keep_head_layout(true, Some((0.0, 3)), None), Some((1.0, 3)));
assert_eq!(
keep_head_layout(true, Some((f64::NAN, 0)), None),
Some((1.0, 0))
);
}
@@ -99,8 +99,42 @@ pub(crate) fn upsert(existing: &str, block: Block<'_>, key: &str, value: &str) -
/// Read `path`, set `key` in `block`, write it back — and back the original up ONCE, the first time
/// we touch a file we did not write. Returns `true` when the file changed (the caller restarts the
/// portal only then).
///
/// The read is matched EXPLICITLY, and only [`ErrorKind::NotFound`](std::io::ErrorKind::NotFound)
/// may mean "empty". This used to be `read_to_string(path).unwrap_or_default()`, which folded every
/// read failure into an empty string — and an empty string is the one input for which this function
/// destroys data: `upsert("")` yields a file holding ONLY our block, the backup below is skipped
/// because there is nothing to back up, and the write replaces the user's config. One non-UTF-8 byte
/// in a comment (a Latin-1 character, an 8-bit paste) or a transient EIO on an NFS/overlay config
/// dir was enough, and the result was exactly the silent, permanent loss this module exists to
/// prevent. A config we cannot read is a config we refuse to rewrite.
pub(crate) fn ensure_key(path: &Path, block: Block<'_>, key: &str, value: &str) -> Result<bool> {
let existing = std::fs::read_to_string(path).unwrap_or_default();
// Read BYTES: whether a backup is owed is a question about what is on disk, not about what
// decoded — and the decode failure below is itself one of the cases that must not be silent.
let raw = match std::fs::read(path) {
Ok(b) => Some(b),
Err(e) if e.kind() == std::io::ErrorKind::NotFound => None,
Err(e) => {
return Err(e).with_context(|| {
format!(
"read {} (refusing to rewrite a portal config we could not read)",
path.display()
)
})
}
};
let existing = match &raw {
Some(bytes) => std::str::from_utf8(bytes)
.with_context(|| {
format!(
"{} is not UTF-8 — refusing to rewrite it (the one key we own is not worth \
losing the rest of the file for; fix or move the file and reconnect)",
path.display()
)
})?
.to_string(),
None => String::new(),
};
let updated = upsert(&existing, block, key, value);
if updated == existing {
return Ok(false);
@@ -108,9 +142,9 @@ pub(crate) fn ensure_key(path: &Path, block: Block<'_>, key: &str, value: &str)
if let Some(dir) = path.parent() {
std::fs::create_dir_all(dir).with_context(|| format!("mkdir {}", dir.display()))?;
}
// One-time backup. `create_new` makes this genuinely once: a later edit must not overwrite the
// user's ORIGINAL with our own previous output.
if !existing.is_empty() {
// One-time backup, of the bytes we actually read. `create_new` makes this genuinely once: a
// later edit must not overwrite the user's ORIGINAL with our own previous output.
if let Some(bytes) = raw.as_deref().filter(|b| !b.is_empty()) {
let backup = path.with_extension("punktfunk-backup");
match std::fs::OpenOptions::new()
.write(true)
@@ -119,7 +153,7 @@ pub(crate) fn ensure_key(path: &Path, block: Block<'_>, key: &str, value: &str)
{
Ok(mut f) => {
use std::io::Write;
let _ = f.write_all(existing.as_bytes());
let _ = f.write_all(bytes);
tracing::info!(
backup = %backup.display(),
"backed up the existing portal config before editing it"
@@ -133,10 +167,92 @@ pub(crate) fn ensure_key(path: &Path, block: Block<'_>, key: &str, value: &str)
),
}
}
std::fs::write(path, &updated).with_context(|| format!("write {}", path.display()))?;
write_atomic(path, updated.as_bytes())?;
Ok(true)
}
/// Replace `path`'s contents with `bytes` **atomically**: fill a temp file beside it, then rename
/// over it. `fs::write` truncates first and fills afterwards, so a crash, a full disk or a killed
/// host between the two leaves the user's config truncated — the same loss this module exists to
/// prevent, arrived at from the other side. The temp file goes in the SAME directory because a
/// rename is only atomic within one filesystem, and it inherits the original's permission bits so
/// an operator's 0600 config does not come back at the umask default.
///
/// A **symlinked** config is followed first, and that is not a nicety: `fs::write` opens the path
/// and therefore writes through the link, while `rename(2)` replaces the link itself. Individual
/// files under `~/.config` are symlinks into a dotfiles repo on every stow / chezmoi / home-manager
/// setup, so renaming over `~/.config/hypr/xdph.conf` would detach the user's repo — their next
/// `stow` reports a conflict or quietly reverts our key, and the connect after that writes it
/// again, forever. Following the link keeps this write byte-for-byte equivalent to the `fs::write`
/// it replaced, atomicity aside; it also makes the permission copy below sample the file the
/// rename actually lands on rather than one it was about to orphan.
///
/// The case this deliberately does NOT paper over: a link into a read-only target (home-manager
/// pointing at `/nix/store`). Following it fails the write, and the caller fails the connect with
/// the store path in the error — exactly as the pre-atomic `fs::write` did. Renaming over the link
/// instead would "work" by quietly detaching a declaratively managed file, which the user's next
/// `home-manager switch` refuses or reverts; a nix-managed config has to gain our key in the
/// user's flake, and a legible error is the only thing that tells them so.
fn write_atomic(path: &Path, bytes: &[u8]) -> Result<()> {
use std::io::Write;
let resolved = follow_link(path);
let path = resolved.as_path();
let dir = path.parent().unwrap_or_else(|| Path::new("."));
let stem = path
.file_name()
.map(|n| n.to_string_lossy().into_owned())
.unwrap_or_else(|| "config".to_string());
// Per-process name: two hosts editing the same config must not fill one another's temp file.
let tmp = dir.join(format!(".{stem}.punktfunk-{}.tmp", std::process::id()));
let write = || -> Result<()> {
{
let mut f =
std::fs::File::create(&tmp).with_context(|| format!("create {}", tmp.display()))?;
f.write_all(bytes)
.with_context(|| format!("write {}", tmp.display()))?;
// The rename must not publish a name whose contents are still in the page cache only.
f.sync_all()
.with_context(|| format!("sync {}", tmp.display()))?;
} // closed before the rename — Windows is far happier renaming a file nobody holds open.
if let Ok(md) = std::fs::metadata(path) {
let _ = std::fs::set_permissions(&tmp, md.permissions());
}
std::fs::rename(&tmp, path)
.with_context(|| format!("rename {} -> {}", tmp.display(), path.display()))
};
let r = write();
if r.is_err() {
// Never leave a half-written dotfile beside the user's config.
let _ = std::fs::remove_file(&tmp);
}
r
}
/// `path` with a symlink chain followed to the file it names, or `path` itself when it is not a
/// link (including when it does not exist yet — the ordinary first-connect case).
///
/// `symlink_metadata` rather than `metadata`, because the question is what `path` IS, not what it
/// points at. A **dangling** link is resolved by hand from its target text: `canonicalize` refuses
/// a target that does not exist, but `fs::write` through such a link creates it, and this write
/// stands in for that one.
fn follow_link(path: &Path) -> std::path::PathBuf {
match std::fs::symlink_metadata(path) {
Ok(md) if md.file_type().is_symlink() => std::fs::canonicalize(path)
.or_else(|_| {
std::fs::read_link(path).map(|target| {
if target.is_absolute() {
target
} else {
// A relative link is relative to the DIRECTORY holding it.
path.parent().unwrap_or_else(|| Path::new(".")).join(target)
}
})
})
.unwrap_or_else(|_| path.to_path_buf()),
_ => path.to_path_buf(),
}
}
#[cfg(test)]
mod tests {
use super::*;
@@ -237,3 +353,240 @@ mod tests {
);
}
}
/// [`ensure_key`] itself — the half that touches the user's disk.
///
/// The merge above was pinned by seven cases while the I/O wrapper around it, which is where the
/// destructive behaviour lives (the read, the once-only backup, the replacing write), had none. That
/// is backwards: `upsert` can at worst return a wrong string, `ensure_key` can delete a config.
/// Filesystem-only — no compositor, no portal — so these run on every platform, like the merge tests.
#[cfg(test)]
mod io_tests {
use super::*;
/// A scratch directory removed on drop. `tempfile` is deliberately not a dependency of this
/// crate; the temp-dir + pid + counter convention is the one `proc.rs`'s fixtures already use.
struct Scratch(std::path::PathBuf);
impl Scratch {
fn new(tag: &str) -> Self {
static N: std::sync::atomic::AtomicU32 = std::sync::atomic::AtomicU32::new(0);
let n = N.fetch_add(1, std::sync::atomic::Ordering::Relaxed);
let dir = std::env::temp_dir()
.join(format!("pf-vd-portalcfg-{tag}-{}-{n}", std::process::id()));
let _ = std::fs::remove_dir_all(&dir);
std::fs::create_dir_all(&dir).expect("scratch dir");
Self(dir)
}
fn path(&self, name: &str) -> std::path::PathBuf {
self.0.join(name)
}
}
impl Drop for Scratch {
fn drop(&mut self) {
let _ = std::fs::remove_dir_all(&self.0);
}
}
fn backup_of(p: &Path) -> std::path::PathBuf {
p.with_extension("punktfunk-backup")
}
/// The data-loss case. A config that cannot be decoded must be left EXACTLY as it is: the old
/// `unwrap_or_default()` turned it into an empty string, wrote a file holding only our block,
/// skipped the backup (nothing to back up, as far as it could tell) and returned `Ok(true)`.
#[test]
fn a_non_utf8_config_is_refused_not_replaced() {
let s = Scratch::new("nonutf8");
let p = s.path("config");
// A Latin-1 'ÿ' in a comment — the whole file is otherwise perfectly ordinary.
let raw: &[u8] = b"[screencast]\n# r\xffgler\nchooser_type=simple\noutput_name=DP-1\n";
std::fs::write(&p, raw).expect("seed");
let err = ensure_key(&p, Block::Ini("screencast"), "chooser_cmd", "cat x")
.expect_err("an unreadable config must not be rewritten");
assert!(
format!("{err:#}").contains("not UTF-8"),
"the error must name the real cause: {err:#}"
);
assert_eq!(
std::fs::read(&p).expect("still there"),
raw,
"byte-identical"
);
assert!(
!backup_of(&p).exists(),
"nothing was edited, so nothing is owed a backup"
);
}
/// The ordinary first-connect path: no file yet, so one is created — and there is no original
/// to preserve, so no backup is left lying beside it.
#[test]
fn a_missing_file_is_created_without_a_backup() {
let s = Scratch::new("missing");
let p = s.path("nested").join("config");
assert!(ensure_key(&p, Block::Ini("screencast"), "chooser_cmd", "cat x").expect("write"));
assert_eq!(
std::fs::read_to_string(&p).expect("created"),
"[screencast]\nchooser_cmd=cat x\n"
);
assert!(!backup_of(&p).exists());
}
/// `create_new` is what makes the backup once-only, and this is the invariant it buys: after a
/// second edit (a new `$XDG_RUNTIME_DIR`, so a new value) the backup must still hold the user's
/// PRISTINE file — not our own previous output.
#[test]
fn the_backup_holds_the_original_across_two_edits() {
let s = Scratch::new("backup");
let p = s.path("config");
let pristine = "[screencast]\nchooser_type=simple\noutput_name=DP-1\n";
std::fs::write(&p, pristine).expect("seed");
assert!(
ensure_key(&p, Block::Ini("screencast"), "chooser_cmd", "cat /run/a").expect("1st")
);
assert!(
ensure_key(&p, Block::Ini("screencast"), "chooser_cmd", "cat /run/b").expect("2nd")
);
assert_eq!(
std::fs::read_to_string(backup_of(&p)).expect("backup"),
pristine
);
let now = std::fs::read_to_string(&p).expect("edited");
assert!(
now.contains("chooser_cmd=cat /run/b"),
"the second value won"
);
assert!(
now.contains("output_name=DP-1"),
"the user's other keys survived"
);
}
/// Idempotence at the I/O level: an already-correct file is not rewritten and reports `false`,
/// because the caller RESTARTS the portal on `true` — a spurious `true` restarts xdpw/xdph on
/// every connect.
#[test]
fn an_unchanged_file_returns_false_and_does_not_rewrite() {
let s = Scratch::new("unchanged");
let p = s.path("config");
assert!(ensure_key(&p, Block::Ini("screencast"), "chooser_cmd", "cat x").expect("1st"));
let after_first = std::fs::read_to_string(&p).expect("written");
let mtime = std::fs::metadata(&p)
.and_then(|m| m.modified())
.expect("mtime");
assert!(
!ensure_key(&p, Block::Ini("screencast"), "chooser_cmd", "cat x").expect("2nd"),
"an unchanged config must report no change"
);
assert_eq!(
std::fs::read_to_string(&p).expect("still there"),
after_first
);
assert_eq!(
std::fs::metadata(&p)
.and_then(|m| m.modified())
.expect("mtime"),
mtime,
"the file must not have been touched at all"
);
}
/// The write publishes the WHOLE new file or nothing (temp + rename), and it leaves no debris
/// beside the config — a stray dotfile in `~/.config/hypr` is the kind of thing that outlives
/// several releases.
#[test]
fn the_write_is_atomic_and_leaves_no_temp_behind() {
let s = Scratch::new("atomic");
let p = s.path("config");
std::fs::write(&p, "[other]\nkeep=me\n").expect("seed");
assert!(ensure_key(&p, Block::Ini("screencast"), "chooser_cmd", "cat x").expect("write"));
let names: Vec<String> = std::fs::read_dir(&s.0)
.expect("dir")
.flatten()
.map(|e| e.file_name().to_string_lossy().into_owned())
.collect();
assert!(
!names.iter().any(|n| n.ends_with(".tmp")),
"temp file left behind: {names:?}"
);
assert!(std::fs::read_to_string(&p)
.expect("edited")
.contains("keep=me"));
}
/// A user who manages dotfiles (stow, chezmoi, home-manager) has `~/.config/hypr/xdph.conf` as
/// a SYMLINK into their repo. The edit has to land in the repo file with the link intact:
/// `fs::write` followed the link, the temp-file + `rename` that replaced it does not, and a
/// detached link is a config the user's tooling then fights us over on every connect.
#[cfg(unix)]
#[test]
fn a_symlinked_config_is_edited_through_the_link() {
let s = Scratch::new("symlink");
let repo = s.path("dotfiles");
std::fs::create_dir_all(&repo).expect("repo dir");
let real = repo.join("xdph.conf");
std::fs::write(
&real,
"screencopy {\n allow_token_by_default = true\n}\n",
)
.expect("seed");
let link = s.path("xdph.conf");
std::os::unix::fs::symlink(&real, &link).expect("symlink");
assert!(ensure_key(
&link,
Block::Hyprlang("screencopy"),
"custom_picker_binary",
"/run/user/1000/shim.sh",
)
.expect("write"));
assert!(
std::fs::symlink_metadata(&link)
.expect("still there")
.file_type()
.is_symlink(),
"the dotfiles link was replaced by a detached regular file"
);
let target = std::fs::read_to_string(&real).expect("the repo file");
assert!(
target.contains("custom_picker_binary = /run/user/1000/shim.sh"),
"the edit never reached the repo file: {target}"
);
assert!(
target.contains("allow_token_by_default = true"),
"the user's own keys survived"
);
}
/// The link may point at a file that does not exist yet (a repo checkout that has not been
/// populated). `fs::write` created the target through it, so this must too — replacing the
/// link would again detach it.
#[cfg(unix)]
#[test]
fn a_dangling_symlink_is_written_through_to_its_target() {
let s = Scratch::new("dangling");
let repo = s.path("dotfiles");
std::fs::create_dir_all(&repo).expect("repo dir");
let real = repo.join("config");
let link = s.path("config");
std::os::unix::fs::symlink(&real, &link).expect("symlink");
assert!(
ensure_key(&link, Block::Ini("screencast"), "chooser_cmd", "cat x").expect("write")
);
assert!(
std::fs::symlink_metadata(&link)
.expect("still there")
.file_type()
.is_symlink(),
"the link was replaced instead of written through"
);
assert_eq!(
std::fs::read_to_string(&real).expect("target created"),
"[screencast]\nchooser_cmd=cat x\n"
);
}
}
+149 -25
View File
@@ -40,7 +40,11 @@ fn chooser_file() -> String {
}
/// The chooser command xdpw runs via `/bin/sh -c`, reading stdout. The `|| echo` fallback keeps
/// plain portal capture (`--source portal`) working when no session has written the chooser file.
/// plain portal capture (`--source portal`) working when no session of ours is mid-handshake — it
/// is a GUESS at sway's own first headless output, right on a box whose sway loads the headless
/// backend with one output of its own and wrong (a cast of nothing) otherwise. It is reachable
/// again: the per-session file is removed with the handshake it steers ([`ChooserFile`]), so it no
/// longer sits there naming an output we have since unplugged.
fn chooser_cmd() -> String {
format!(
"cat {} 2>/dev/null || echo 'Monitor: HEADLESS-1'",
@@ -68,8 +72,14 @@ impl WlrootsDisplay {
/// wlroots/Sway is usable when the host runs inside a Sway session — signalled by `SWAYSOCK`
/// (the IPC socket `swaymsg create_output` needs). Cheap env check for the enumeration path.
///
/// Under [`crate::with_env_lock`]: this runs on a management worker (`/host/compositors` →
/// [`crate::available`]) concurrently with another connect's `apply_session_env`, which `set_var`s
/// — and, when no sway session is live, `remove_var`s — this very key. A glibc `getenv` racing a
/// `setenv` is the `environ` realloc data race ENV_LOCK exists for, and it is UB whichever key each
/// side names. No caller holds the lock (the mutex is not reentrant).
pub fn is_available() -> bool {
std::env::var_os("SWAYSOCK").is_some()
crate::with_env_lock(|| std::env::var_os("SWAYSOCK")).is_some()
}
impl VirtualDisplay for WlrootsDisplay {
@@ -86,13 +96,33 @@ impl VirtualDisplay for WlrootsDisplay {
}
fn create(&mut self, mode: Mode) -> Result<VirtualOutput> {
let before = output_names()
.context("swaymsg get_outputs (is the host inside the sway session env — SWAYSOCK?)")?;
swaymsg(&["create_output"])
.context("swaymsg create_output (sway needs the headless backend loaded)")?;
// The output appears synchronously in practice; poll briefly to be safe, and own it
// from here on so error unwinding unplugs it.
let output = OutputGuard(wait_new_output(&before, Duration::from_secs(5))?);
warn_topology_is_extend_only();
// Snapshot → create → identify, all under CREATE_LOCK. sway names the headless output
// itself (`HEADLESS-N`), so the only way to know which one is ours is "the name that was not
// there before" — and two concurrent creates each picking the other's output is a silent
// mis-capture, not a failure (mutter's TOPOLOGY_LOCK exists for exactly this class). The
// lock also gives the failure path somewhere safe to unplug from: the output already exists
// by the time `wait_new_output` can fail, and nothing else may have created one meanwhile.
let output = {
let _create = CREATE_LOCK.lock().unwrap_or_else(|e| e.into_inner());
let before = output_names().context(
"swaymsg get_outputs (is the host inside the sway session env — SWAYSOCK?)",
)?;
swaymsg(&["create_output"])
.context("swaymsg create_output (sway needs the headless backend loaded)")?;
// The output appears synchronously in practice; poll briefly to be safe, and own it
// from here on so error unwinding unplugs it.
match wait_new_output(&before, Duration::from_secs(5)) {
Ok(name) => OutputGuard(name),
Err(e) => {
// `create_output` reported success, so an output very probably exists — it just
// never showed up in time (or showed up a moment after we gave up). Unowned, it
// would sit in the operator's sway layout forever.
unplug_strays(&before);
return Err(e);
}
}
};
let name = output.0.clone();
// The client's exact mode (also the refresh clock that makes the output produce frames).
@@ -128,7 +158,7 @@ impl VirtualDisplay for WlrootsDisplay {
remote_fd: Some(fd),
preferred_mode: Some((mode.width, mode.height, mode.refresh_hz)),
keepalive: Box::new(Keepalive {
_stop: StopGuard(stop),
_stop: stop,
_output: output,
}),
// Owned (the compositor output is ours to tear down), but not registry-poolable: the
@@ -159,6 +189,52 @@ impl Drop for StopGuard {
}
}
/// Serializes **snapshot → `create_output` → identify-the-new-name**, process-wide. sway names its
/// headless outputs itself, so ownership is established by a before/after diff and two concurrent
/// creates would each adopt the other's output — which does not fail, it silently streams the wrong
/// one. Mutter's `TOPOLOGY_LOCK` is the same guard for the same reason; Hyprland needs none because
/// it lets us NAME the output (D6).
static CREATE_LOCK: std::sync::Mutex<()> = std::sync::Mutex::new(());
/// Unplug any headless output that appeared since `before` and that nothing owns — the cleanup for a
/// `create_output` whose output we could not identify in time. Only `HEADLESS-*` is touched: a
/// physical hotplug in the same window is the operator's, not ours, and `unplug` on a real connector
/// would take their screen away. Best-effort by construction, and it runs with [`CREATE_LOCK`] held
/// so nothing else in this process can have created the strays it sees.
fn unplug_strays(before: &[String]) {
let Ok(now) = output_names() else { return };
for name in now
.into_iter()
.filter(|n| n.starts_with("HEADLESS-") && !before.iter().any(|b| b == n))
{
match swaymsg(&["output", &name, "unplug"]) {
Ok(_) => tracing::warn!(output = %name, "unplugged a headless output we created but \
could not identify in time"),
Err(e) => tracing::warn!(output = %name, error = %format!("{e:#}"), "could not unplug \
the headless output left behind by a failed create"),
}
}
}
/// The configured [`crate::policy::Topology`] is not implemented on this backend — say so once per
/// create instead of leaving the management API's echo as the only signal that the pin was dropped
/// (sweep 13.18). sway's virtual output is always an EXTENSION: nothing here promotes it to primary
/// or disables the operator's heads.
fn warn_topology_is_extend_only() {
let topology = crate::effective_topology();
if !matches!(
topology,
crate::policy::Topology::Extend | crate::policy::Topology::Auto
) {
tracing::warn!(
?topology,
"wlroots: this backend implements EXTEND only — the headless output is added beside the \
operator's heads and nothing is promoted or disabled. Configure `topology: extend` to \
stop the console promising otherwise."
);
}
}
/// Owns the created headless output; dropping it unplugs it from sway.
struct OutputGuard(String);
@@ -171,15 +247,26 @@ impl Drop for OutputGuard {
}
}
/// Budget for one `swaymsg` call ([`crate::proc`]).
///
/// swaymsg is a CLIENT of the compositor it drives: against a wedged sway it blocks in its own
/// connect to the IPC socket and never returns — and these calls run on the session's stream thread,
/// whose only way to end a session is to return, so one hung query used to wedge the session
/// permanently. Generous next to a healthy call (single-digit milliseconds), and every call site
/// here already has a failed-query path, so a timeout lands on behaviour that already exists.
const SWAYMSG_BUDGET: Duration = Duration::from_secs(5);
/// Budget for the one-shot xdpw restart. `systemctl --user try-restart` waits for the unit's job to
/// settle, so it is the slowest helper on this path — and its result is already ignored.
const PORTAL_RESTART_BUDGET: Duration = Duration::from_secs(10);
/// Run `swaymsg -- <args>`, returning stdout (`--` so command tokens like `--custom` reach
/// sway instead of swaymsg's own getopt). swaymsg exits non-zero (with the error on stderr/
/// stdout) when the command fails, so checking the status covers `{"success": false}` too.
fn swaymsg(args: &[&str]) -> Result<String> {
let out = Command::new("swaymsg")
.arg("--")
.args(args)
.output()
.context("run swaymsg (is sway installed?)")?;
let out =
crate::proc::output_within(Command::new("swaymsg").arg("--").args(args), SWAYMSG_BUDGET)
.context("run swaymsg (is sway installed?)")?;
if !out.status.success() {
bail!(
"swaymsg {:?} failed: {}{}",
@@ -197,10 +284,11 @@ fn swaymsg(args: &[&str]) -> Result<String> {
/// *command*, which is right for `create_output` and wrong for a query — `-t` after `--` comes back
/// as `Unknown/invalid command '-t'` (caught on-glass writing the monitor enumeration).
fn swaymsg_query(kind: &str) -> Result<serde_json::Value> {
let out = Command::new("swaymsg")
.args(["-t", kind, "--raw"])
.output()
.context("run swaymsg (is sway installed?)")?;
let out = crate::proc::output_within(
Command::new("swaymsg").args(["-t", kind, "--raw"]),
SWAYMSG_BUDGET,
)
.context("run swaymsg (is sway installed?)")?;
if !out.status.success() {
bail!(
"swaymsg -t {kind} failed: {}",
@@ -230,13 +318,37 @@ fn output_names() -> Result<Vec<String>> {
/// handshake, not just the write, because the read happens inside it.
static SELECTION_LOCK: std::sync::Mutex<()> = std::sync::Mutex::new(());
/// The per-session chooser file, removed when the handshake it steers is over.
///
/// Its lifetime is the HANDSHAKE, not the session: xdpw reads it once, inside
/// [`select_and_cast`]'s critical section, and everything after that is the cast's own business.
/// Left behind (as it was) the stale `Monitor: HEADLESS-3` outlives the output `Drop` has since
/// unplugged, and it permanently shadows [`chooser_cmd`]'s `|| echo` fallback — so a later
/// `--source portal` capture with no session of ours running steers at a connector that is gone.
/// Tying removal to the CAST instead would be worse still: the file is one per user, so a session
/// ending hours later would delete a *sibling's* selection out from under its picker.
struct ChooserFile(String);
impl Drop for ChooserFile {
fn drop(&mut self) {
if let Err(e) = std::fs::remove_file(&self.0) {
if e.kind() != std::io::ErrorKind::NotFound {
tracing::debug!(path = %self.0, error = %e, "could not remove the xdpw chooser file");
}
}
}
}
/// Point xdpw's chooser at `output` and run the ScreenCast handshake, returning the portal fd +
/// node id and the guard that stops the cast. The caller must hold [`SELECTION_LOCK`].
fn select_and_cast(output: &str, hw_cursor: bool) -> Result<(OwnedFd, u32, Arc<AtomicBool>)> {
fn select_and_cast(output: &str, hw_cursor: bool) -> Result<(OwnedFd, u32, StopGuard)> {
ensure_xdpw_config()?;
let chooser = chooser_file();
std::fs::write(&chooser, format!("Monitor: {output}\n"))
.with_context(|| format!("write {chooser}"))?;
// Owned from the write on: every arm below (and every `?`) leaves the handshake, which is the
// only thing that reads it.
let _chooser = ChooserFile(chooser);
let (setup_tx, setup_rx) = std::sync::mpsc::channel::<Result<(OwnedFd, u32), String>>();
let stop = Arc::new(AtomicBool::new(false));
let stop_thread = stop.clone();
@@ -244,8 +356,16 @@ fn select_and_cast(output: &str, hw_cursor: bool) -> Result<(OwnedFd, u32, Arc<A
.name("punktfunk-wlr-cast".into())
.spawn(move || portal_thread(setup_tx, stop_thread, hw_cursor))
.context("spawn wlroots portal thread")?;
// Built BEFORE the wait so EVERY error arm below sets the flag on its way out — as Mutter's
// `create` does. Returning the bare `Arc` and letting the CALLER wrap it left the two failure
// arms dropping an un-set flag: the thread's `send` can still LAND in the queue in the window
// between `recv_timeout` giving up and `setup_rx` being dropped, so it reports success and then
// parks forever on `while !stop`, holding a live ScreenCast session, its zbus connection, an
// `OwnedFd` and a 2-worker tokio runtime — one more set per slow-portal connect, for the host's
// lifetime, against an output that no longer exists.
let guard = StopGuard(stop);
match setup_rx.recv_timeout(Duration::from_secs(20)) {
Ok(Ok((fd, node_id))) => Ok((fd, node_id, stop)),
Ok(Ok((fd, node_id))) => Ok((fd, node_id, guard)),
Ok(Err(e)) => bail!("ScreenCast portal on {output} failed: {e}"),
Err(_) => bail!("timed out waiting for the ScreenCast portal on {output}"),
}
@@ -266,7 +386,7 @@ pub(crate) fn stream_existing_output(
Ok(crate::mirror::MirrorStream {
node_id,
remote_fd: Some(fd),
keepalive: Box::new(StopGuard(stop)),
keepalive: Box::new(stop),
})
}
@@ -374,9 +494,13 @@ fn ensure_xdpw_config() -> Result<()> {
return Ok(());
}
tracing::info!(path = %path.display(), "pointed xdg-desktop-portal-wlr at the managed output chooser");
let _ = Command::new("systemctl")
.args(["--user", "try-restart", "xdg-desktop-portal-wlr.service"])
.status();
// Bounded: `systemctl --user` blocks on the user manager's job queue, and this runs on the
// session's stream thread. Its result was already ignored — a timeout just means the portal
// picks the new config up whenever it next starts.
let _ = crate::proc::status_within(
Command::new("systemctl").args(["--user", "try-restart", "xdg-desktop-portal-wlr.service"]),
PORTAL_RESTART_BUDGET,
);
Ok(())
}
+61 -5
View File
@@ -65,6 +65,16 @@ impl VirtualDisplay for MirrorDisplay {
self.hw_cursor
}
fn poolable_now(&self) -> bool {
// Never. `create` below always reports `DisplayOwnership::External` — we did not make this
// head and must not keep it — so the registry never pools a mirror, and the trait's `true`
// default was a claim this backend cannot honour on any request. It costs nothing today
// (the reuse lookup can only miss: no `"mirror"` entry ever enters the pool), but it is the
// answer the registry consults BEFORE `create` gets to declare ownership, so leaving it
// optimistic means the one pre-create statement of intent contradicts the post-create fact.
false
}
fn create(&mut self, _mode: Mode) -> Result<VirtualOutput> {
// Resolve the pin against the live head list FIRST: it yields the geometry the input anchor
// needs, and it turns "that monitor is gone" into one clear error before any compositor
@@ -101,7 +111,14 @@ impl VirtualDisplay for MirrorDisplay {
Compositor::Gamescope => {
crate::gamescope::stream_existing_output(&target.connector, self.hw_cursor)?
}
#[allow(unreachable_patterns)]
// Gated to non-Linux (`monitors::list`'s shape), NOT the bare `#[allow(unreachable_
// patterns)] other =>` this replaced: with it, a newly added `Compositor` variant fell
// through to a runtime bail on the very platform that would define it, silently, in the
// one place that decides which backends can mirror a head. Cfg'd out on Linux, the match
// is exhaustive and the new variant is a compile error here instead. The arm exists at
// all only because every arm above is itself `cfg(target_os = "linux")` — this module is
// Linux-only today, so it is a placeholder that keeps the shape honest if that changes.
#[cfg(not(target_os = "linux"))]
other => bail!(
"mirroring an existing monitor is not supported on the {} backend",
other.id()
@@ -172,11 +189,28 @@ fn check_mirrorable(target: &monitors::PhysicalMonitor, compositor: Compositor)
Ok(())
}
/// Does this compositor's `managed` flag mean "ours, for certain"? KWin outputs carry the
/// `Virtual-punktfunk` prefix we chose, and Hyprland's are `PF-N` — both ours by construction.
/// Sway's `HEADLESS-N` is sway's own generic naming, so it is a hint, not proof.
/// Does this compositor's `managed` flag mean "ours, for certain"?
///
/// EXHAUSTIVE on purpose, unlike the `matches!` it used to be. This is the one table in the crate
/// whose un-listed default is the UNSAFE direction: a `false` sends [`check_mirrorable`] down the
/// warn-and-proceed branch, which for a backend that DOES name its managed outputs by construction
/// (the KWin/Hyprland shape — i.e. both backends that have the property today) means streaming
/// punktfunk's own virtual display back to the client, the capture loop
/// `one_of_our_own_virtual_displays_is_refused` exists to forbid. Adding a `Compositor` variant must
/// therefore be a compile error here rather than a silent opt-out. (Contrast
/// [`Compositor::needs_live_session`], also a `matches!` — its omitted default is the safe one.)
fn names_ours_conclusively(compositor: Compositor) -> bool {
matches!(compositor, Compositor::Kwin | Compositor::Hyprland)
match compositor {
// Ours by construction: KWin outputs carry the `Virtual-punktfunk-<id>` name the identity
// module hands the backend, Hyprland's are `PF-N`. Nothing else mints those names.
Compositor::Kwin | Compositor::Hyprland => true,
// Sway names EVERY headless output `HEADLESS-N`, its own included; Mutter's virtual monitors
// carry no distinguishing name at all (it won't take one from us); and gamescope's
// `list_monitors` only ever reports the real DRM head a Game Mode session drives, so
// `managed` is never even set there. A hint at most — refusing would break the legitimate
// headless-sway setup this feature serves.
Compositor::Wlroots | Compositor::Mutter | Compositor::Gamescope => false,
}
}
/// mHz → whole Hz for [`VirtualOutput::preferred_mode`], never 0 (the negotiation treats 0 as
@@ -252,6 +286,28 @@ mod tests {
assert!(check_mirrorable(&m, Compositor::Hyprland).is_err());
}
/// Pin the conclusive-naming table per variant. The answer is a safety decision whose wrong
/// direction is the SILENT one: a backend that mints punktfunk-named outputs but is missing
/// from the `true` arm takes the warn-and-proceed branch and streams our own virtual display
/// back to the client. Exhaustive `match` + this test = the new variant has to be considered.
#[test]
fn the_conclusive_naming_table_is_pinned_per_backend() {
assert!(names_ours_conclusively(Compositor::Kwin));
assert!(names_ours_conclusively(Compositor::Hyprland));
assert!(!names_ours_conclusively(Compositor::Wlroots));
assert!(!names_ours_conclusively(Compositor::Mutter));
assert!(!names_ours_conclusively(Compositor::Gamescope));
}
/// The registry asks `poolable_now` BEFORE `create` gets to report ownership, so the two must
/// agree: a mirror's `create` always reports `External` (we did not make this head), therefore
/// no mirror request is ever poolable.
#[test]
fn a_mirrored_head_is_never_registry_poolable() {
let vd = MirrorDisplay::new(Compositor::Kwin, "DP-2".into()).unwrap();
assert!(!vd.poolable_now());
}
/// A head listed but not driving a mode (enabled yet modeless) would negotiate a 0x0 stream.
#[test]
fn a_head_with_no_current_mode_is_refused() {
+147 -14
View File
@@ -18,15 +18,22 @@
use crate::Compositor;
use anyhow::{bail, Result};
/// One head as the compositor currently reports it. Logical (post-scale) geometry throughout —
/// the same coordinate space libei regions and compositor layout use, *not* pixels.
/// One head as the compositor currently reports it.
///
/// **The two halves live in different spaces, and that is not an accident.** `x`/`y` are LOGICAL —
/// the compositor's global layout coordinates, the same space libei regions use — while
/// `width`/`height` are the current mode in PIXELS, because that is what every backend actually
/// reports (KWin's `current_mode` size, `hyprctl`'s mode, the CCD path's source mode) and what a
/// capturer has to open against. `scale` is the factor between them: see [`Self::logical_size`],
/// which is the only correct way to compare a size against `x`/`y`. An earlier version of this doc
/// claimed logical geometry "throughout", which is a trap for exactly the consumer that mixes them.
#[derive(Clone, Debug, PartialEq)]
pub struct PhysicalMonitor {
/// Connector name — `DP-1`, `HDMI-A-2`, `eDP-1`. The id `PUNKTFUNK_CAPTURE_MONITOR` names.
pub connector: String,
/// Human label for a picker (`make model`, else the connector). Never used for matching.
pub description: String,
/// Current mode, in pixels.
/// Current mode, in PIXELS (not the logical size — see the type doc and [`Self::logical_size`]).
pub width: u32,
pub height: u32,
/// Refresh in mHz (60000 = 60 Hz). 0 when the backend doesn't report it.
@@ -71,6 +78,24 @@ pub(crate) fn describe(make: &str, model: &str, connector: &str) -> String {
}
impl PhysicalMonitor {
/// The head's extent in the SAME space as `x`/`y` — mode pixels divided by `scale`.
///
/// The bridge between the two spaces this type carries, and the only correct way to ask "does
/// this head's box contain that layout coordinate?". A consumer that compares `width`/`height`
/// against `x`/`y` directly is right only at scale 1.0 and silently wrong on every fractional
/// KDE/GNOME desk (a 3840-px panel at 150 % occupies 2560 logical units, so a naive
/// `x + width` overlaps the head to its right by 1280).
///
/// A non-positive scale can only come from a backend that reported nonsense; it is treated as
/// 1.0 rather than dividing by zero.
pub fn logical_size(&self) -> (f64, f64) {
let scale = if self.scale > 0.0 { self.scale } else { 1.0 };
(
f64::from(self.width) / scale,
f64::from(self.height) / scale,
)
}
/// `1920x1080@60` — for logs and pickers.
pub fn mode_label(&self) -> String {
if self.refresh_mhz == 0 {
@@ -94,8 +119,11 @@ impl PhysicalMonitor {
/// callers resolving a pinned monitor must not (see [`resolve`]).
pub fn list(compositor: Compositor) -> Result<Vec<PhysicalMonitor>> {
match compositor {
// Via the `kwin` backend rather than `kwin_output_mgmt` directly: it owns the
// in-process-then-`kscreen-doctor` ladder, so this read degrades the same way every other
// KWin operation does instead of being the one that hard-fails on a wedged/old compositor.
#[cfg(target_os = "linux")]
Compositor::Kwin => crate::kwin_output_mgmt::list_monitors(),
Compositor::Kwin => crate::kwin::list_monitors(),
#[cfg(target_os = "linux")]
Compositor::Mutter => crate::mutter::list_monitors(),
#[cfg(target_os = "linux")]
@@ -133,15 +161,21 @@ pub fn list(compositor: Compositor) -> Result<Vec<PhysicalMonitor>> {
/// * `refresh_mhz` comes from the path's own rational rate, which keeps 59.94 distinct from 60.
#[cfg(windows)]
pub fn list_windows() -> Result<Vec<PhysicalMonitor>> {
let inv = pf_win_display::win_display::target_inventory();
if inv.is_empty() {
// Distinguish "reached it, nothing there" from a failure, exactly as [`list`] promises:
// an empty CCD database is a real state (every panel off — measured on .173 with the TV
// powered down), not an error.
return Ok(Vec::new());
}
Ok(inv
.into_iter()
// `Ok` even when the inventory is empty, exactly as [`list`] promises: an empty CCD database is
// a real state (every panel off — measured on .173 with the TV powered down), not a failure.
// Everything past the OS call is the pure mapping, so it lives where a test can reach it.
Ok(from_inventory(
pf_win_display::win_display::target_inventory(),
))
}
/// The CCD inventory → [`PhysicalMonitor`] mapping, split from the OS call so the Windows test leg
/// can exercise it (`list_windows` touches the display database on its first line, which left the
/// only mapping that decides what an operator can PIN with no coverage on the one platform that
/// runs it).
#[cfg(windows)]
fn from_inventory(inv: Vec<pf_win_display::win_display::TargetInventory>) -> Vec<PhysicalMonitor> {
inv.into_iter()
.map(|t| {
// The GDI name is what an operator recognises and what capture pins on; an inactive
// path has none, so fall back to the stable target id rather than an empty string —
@@ -167,7 +201,7 @@ pub fn list_windows() -> Result<Vec<PhysicalMonitor>> {
managed: t.ours,
}
})
.collect())
.collect()
}
/// Resolve a configured monitor name against `monitors`, exactly then case-insensitively.
@@ -257,6 +291,24 @@ mod tests {
assert_eq!(describe(" ", "unknown", "DP-2"), "DP-2");
}
/// The two spaces this type carries: the mode is pixels, `x`/`y` are logical, and `scale` is
/// the only thing that relates them. A 4K panel at KDE's 150 % really does occupy 2560x1440
/// logical units, which is what a consumer comparing against `x`/`y` must use.
#[test]
fn logical_size_divides_the_mode_by_the_scale() {
let mut m = mon("DP-1");
m.width = 3840;
m.height = 2160;
m.scale = 1.5;
assert_eq!(m.logical_size(), (2560.0, 1440.0));
// Unscaled: the two spaces coincide, which is why the trap goes unnoticed on most desks.
m.scale = 1.0;
assert_eq!(m.logical_size(), (3840.0, 2160.0));
// A backend that reported nonsense must not produce an infinity or a NaN.
m.scale = 0.0;
assert_eq!(m.logical_size(), (3840.0, 2160.0));
}
#[test]
fn mode_label_drops_an_unknown_refresh() {
let mut m = mon("DP-1");
@@ -265,3 +317,84 @@ mod tests {
assert_eq!(m.mode_label(), "1920x1080");
}
}
/// The Windows inventory mapping. Windows-only because it maps a Windows-only type — the CI leg
/// that runs it (`windows-host.yml`, `cargo test --release -p pf-vdisplay`) already exists; until
/// [`from_inventory`] was split out of the OS call there was simply nothing there to run.
#[cfg(all(test, windows))]
mod windows_tests {
use super::*;
use pf_win_display::win_display::TargetInventory;
/// One inventory row. Built through a single helper so a field rename shows up in one place —
/// the struct is another crate's and carries no `Default`.
fn target(target_id: u32, gdi_name: &str, active: bool) -> TargetInventory {
TargetInventory {
target_id,
active,
external_physical: true,
internal_panel: false,
tech: "HDMI",
friendly: "ACME TV".into(),
monitor_device_path: r"\\?\DISPLAY#ACM1234#".into(),
ours: false,
gdi_name: gdi_name.into(),
x: 0,
y: 0,
width: 1920,
height: 1080,
refresh_mhz: 59940,
primary: active,
}
}
/// An INACTIVE path has no source and therefore no GDI name. It must still be listed (the
/// "why can't I pick it?" contract) under an id that can actually be pinned — a blank connector
/// could never be resolved, and an operator would have no way to name the head at all.
#[test]
fn an_inactive_path_gets_a_target_id_connector_and_enabled_false() {
let mons = from_inventory(vec![target(4352, "", false)]);
assert_eq!(mons.len(), 1);
assert_eq!(mons[0].connector, "target-4352");
assert!(!mons[0].enabled);
// Windows applies DPI per application rather than a compositor-global logical scale, so
// the geometry above is pixels and the factor is honestly 1.0 — see the fn doc.
assert_eq!(mons[0].scale, 1.0);
}
/// The two halves must agree: whatever connector this mapping synthesizes has to be a name
/// [`resolve`] can find, because that pair is the whole pin round-trip the console offers.
#[test]
fn resolve_can_find_a_synthesized_target_name() {
let mons = from_inventory(vec![
target(4352, "", false),
target(1, r"\\.\DISPLAY1", true),
]);
assert_eq!(
resolve(&mons, "target-4352")
.expect("synthesized name")
.width,
1920
);
// An active path keeps its GDI name — the id an operator recognises.
assert_eq!(
resolve(&mons, r"\\.\DISPLAY1").expect("gdi name").connector,
r"\\.\DISPLAY1"
);
assert!(
resolve(&mons, r"\\.\display1").is_ok(),
"and case-insensitively, as `resolve` promises"
);
}
/// Our own IddCx display is flagged, so a picker can grey it out — the one thing Windows can
/// answer reliably and the Linux backends cannot.
#[test]
fn our_own_idd_is_marked_managed() {
let mut ours = target(257, r"\\.\DISPLAY2", true);
ours.ours = true;
let mons = from_inventory(vec![ours]);
assert!(mons[0].managed);
assert!(!from_inventory(vec![target(1, r"\\.\DISPLAY1", true)])[0].managed);
}
}
File diff suppressed because it is too large Load Diff
+233 -24
View File
@@ -12,21 +12,51 @@
//! take their existing failure path instead of hanging.
//!
//! What the budget bounds is the whole **process tree**, not just the process we spawned — see
//! [`tree`] for why that distinction is the entire difference on Windows.
//! [`tree`] for why that distinction is the entire difference on Windows, and for the one Unix
//! case (a unit the *user manager* forks for us) that even a process group cannot reach.
use std::io::{Error, ErrorKind, Result};
// `Read` is in scope for `Take::read_to_end` below — a `Take<R>` is a concrete type, so the
// generic bound alone does not bring the trait's methods with it.
use std::io::{Error, ErrorKind, Read, Result};
use std::process::{Command, ExitStatus, Output};
use std::sync::mpsc::{self, Receiver, RecvTimeoutError};
use std::time::{Duration, Instant};
/// Poll interval while waiting for a child to exit. Short enough that a fast helper (the normal
/// case — `kscreen-doctor` answers in tens of ms) isn't measurably delayed.
const POLL: Duration = Duration::from_millis(20);
/// Ceiling on how long [`output_within`] waits for its two reader threads once the child **and the
/// process group under it** are dead.
///
/// This is not a working budget — with every write end we can reach closed, the readers hit EOF
/// within a scheduler slice — it is the bound on the one case we cannot reach. [`tree`] ends a
/// *group*, so a descendant that deliberately left it keeps the write end open: `systemd-run
/// --pipe` (the gamescope bind probe) hands our pipes to a transient unit the **user manager**
/// forks, in its own group and session, and `killpg` by construction cannot touch it. Waiting on
/// that reader would pin the caller — on the host, the session's stream thread — for as long as
/// the unit lives, which is exactly the unbounded wait this module exists to prevent. So the
/// *call* is bounded here and the reader thread, not the call, is what gets left behind. The
/// price, paid only in that case, is that a call can return up to this much after its own budget —
/// still a bound, which an unreachable EOF is not.
const DRAIN_GRACE: Duration = Duration::from_secs(2);
/// Ceiling on what one drained pipe may buffer.
///
/// `read_to_end` is unbounded in memory, and a reader thread that outlived its call (see
/// [`DRAIN_GRACE`]) has nobody left to stop it — the cap is what keeps such a thread finite in
/// both memory and lifetime, and closing its read end is also what finally gives the escaped
/// writer an EPIPE. 16 MiB is an order of magnitude above the largest `pw-dump` a populated
/// PipeWire graph produces, so hitting it means a helper that ran away rather than one that was
/// busy; it is logged instead of being returned as quietly short output.
const DRAIN_CAP: u64 = 16 * 1024 * 1024;
/// Run `cmd` to completion, killing it if it outlives `budget`.
///
/// Stdout/stderr are left as the caller configured them (inherited by default), so this is for
/// commands run for their exit status alone — see [`output_within`] when the output is read.
pub(crate) fn status_within(cmd: &mut Command, budget: Duration) -> Result<ExitStatus> {
tree::prepare(cmd);
let mut child = cmd.spawn()?;
let tree = tree::Guard::attach(&child);
let deadline = Instant::now() + budget;
@@ -51,37 +81,132 @@ pub(crate) fn status_within(cmd: &mut Command, budget: Duration) -> Result<ExitS
/// Run `cmd` to completion and capture its stdout/stderr, killing it if it outlives `budget`.
///
/// The output is read only after the child has exited, so a helper that fills the pipe buffer and
/// stalls is caught by the budget rather than deadlocking the reader (these helpers emit at most a
/// few hundred KiB, well under any real pipe pressure).
/// Both pipes are drained **concurrently with the wait**, on their own threads. Reading them only
/// after exit — the obvious shape, and what this did originally — deadlocks on any helper that
/// outtalks the pipe buffer: a pipe holds **64 KiB** on Linux (`/proc/sys/fs/pipe-max-size`'s page
/// default), not the "few hundred KiB" the old comment claimed, so a chatty helper blocks in
/// `write()`, never reaches exit, is killed at the budget, and its output is discarded as a
/// timeout. `pw-dump` on a populated PipeWire graph clears 64 KiB routinely, and it is polled from
/// the 45 s gamescope loops — so the failure was not hypothetical, it was the busiest caller.
pub(crate) fn output_within(cmd: &mut Command, budget: Duration) -> Result<Output> {
tree::prepare(cmd);
let mut child = cmd
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.spawn()?;
let tree = tree::Guard::attach(&child);
// Taken off the `Child` so the reader threads own them outright: `wait_with_output` must not
// also be reading these, and `try_wait` below needs `&mut child` while they run.
let (stdout, stderr) = (child.stdout.take(), child.stderr.take());
let (out_rx, err_rx) = (drain(stdout), drain(stderr));
let deadline = Instant::now() + budget;
loop {
let status = loop {
match child.try_wait()? {
Some(_) => {
// Exited: `wait_with_output` now only drains already-buffered pipes — but only if
// nothing else still holds their WRITE end. A grandchild that outlived the helper
// does, and `wait_with_output` reads to an EOF that would then never arrive, which
// is the one way this "bounded" helper could still hang forever. End the tree first.
Some(status) => {
// The helper is gone, but a grandchild it left behind still holds the pipes' WRITE
// ends, so the readers below would wait for an EOF that never arrives. Ending the
// tree closes them for every descendant that stayed in the group — which is all of
// them for a direct exec, but NOT for one that left it (see [`DRAIN_GRACE`]), so
// the collection below is bounded rather than a plain join.
tree.terminate();
return child.wait_with_output();
break status;
}
None if Instant::now() >= deadline => {
tree.terminate();
let _ = child.kill();
let _ = child.wait();
let _ = child.wait(); // reap it — never leave a zombie behind
// Reap the READERS too. This arm used to just drop their handles, i.e. detach two
// threads still blocked in `read_to_end` and still owning the pipes' read ends —
// so a writer that escaped the group (a `systemd-run --pipe` unit) never even got
// the EPIPE the pre-drain implementation gave it by closing those fds with the
// `Child`. Joining unconditionally instead would be worse: it would hand the
// escaped writer the caller's thread, forever, which is the failure this whole
// module exists to prevent. So: a bounded collection, and an honest log when one
// of them cannot be reclaimed.
let until = Instant::now() + DRAIN_GRACE;
let (out, err) = (collect(&out_rx, until), collect(&err_rx, until));
if out.is_none() || err.is_none() {
stuck_reader(cmd, "killed at its budget");
}
return Err(timed_out(cmd, budget));
}
None => std::thread::sleep(POLL),
}
};
// Both halves of the output, or none: a caller parsing half a `pw-dump` is a caller being lied
// to, and its failure path is the one it already has for a helper that did not answer.
let until = Instant::now() + DRAIN_GRACE;
let (Some(stdout), Some(stderr)) = (collect(&out_rx, until), collect(&err_rx, until)) else {
stuck_reader(cmd, "exited");
let program = cmd.get_program().to_string_lossy().to_string();
return Err(Error::new(
ErrorKind::TimedOut,
format!(
"`{program}` exited but its output could not be drained within {DRAIN_GRACE:?}"
),
));
};
Ok(Output {
status,
stdout,
stderr,
})
}
/// Read one of a child's pipes on its own thread, so the child never blocks in `write()` waiting
/// for us to catch up, and hand the result back over a channel — not a `JoinHandle`, because the
/// caller must be able to give up on a reader it cannot unblock (see [`DRAIN_GRACE`]) and a
/// `join` offers no way to. Returns whatever was read; a read error yields the partial buffer,
/// because the caller's failure signal is the budget, not a short pipe.
fn drain<R: std::io::Read + Send + 'static>(pipe: Option<R>) -> Receiver<Vec<u8>> {
let (tx, rx) = mpsc::channel();
std::thread::spawn(move || {
let mut buf = Vec::new();
if let Some(r) = pipe {
let mut r = r.take(DRAIN_CAP);
let _ = r.read_to_end(&mut buf);
if buf.len() as u64 >= DRAIN_CAP {
tracing::warn!(
cap_bytes = DRAIN_CAP,
"a helper outran the drain cap — its output is truncated here, which the \
caller sees as an unparseable answer (i.e. a failed query)"
);
}
}
// The receiver is gone whenever the call has already returned — a timeout, or a grace that
// ran out. That is the only way this send fails, and it is a case we chose.
let _ = tx.send(buf);
});
rx
}
/// Take one drained pipe, waiting no longer than `until`. `None` means the reader is still parked
/// on a write end nothing we can signal is holding open.
fn collect(rx: &Receiver<Vec<u8>>, until: Instant) -> Option<Vec<u8>> {
match rx.recv_timeout(until.saturating_duration_since(Instant::now())) {
Ok(buf) => Some(buf),
// The reader panicked: that loses its half of the output, never the call.
Err(RecvTimeoutError::Disconnected) => Some(Vec::new()),
Err(RecvTimeoutError::Timeout) => None,
}
}
/// Say plainly what a stuck reader costs, because the thread is genuinely leaked and there is no
/// portable way to unblock a thread already inside `read()` on a pipe (a `dup2` over the fd does
/// not re-target a read in flight, and closing it under the thread is a use-after-free waiting for
/// an fd number to be reused). It ends when the escaped writer closes or [`DRAIN_CAP`] is reached.
fn stuck_reader(cmd: &Command, what: &str) {
tracing::warn!(
program = %cmd.get_program().to_string_lossy(),
grace_ms = DRAIN_GRACE.as_millis() as u64,
"helper {what} but its pipes never reached EOF — something it started is outside our \
process group and still holds the write end (`systemd-run --pipe` is the known case). \
The call is bounded; the reader thread is detached until that writer closes."
);
}
fn timed_out(cmd: &Command, budget: Duration) -> Error {
let program = cmd.get_program().to_string_lossy().to_string();
tracing::warn!(
@@ -240,11 +365,10 @@ fn undecorate(name: &str) -> &str {
/// Ending the *tree* the helper started, not just the process we spawned.
///
/// [`std::process::Child::kill`] is one `TerminateProcess` / one `SIGKILL`: it ends exactly the
/// process we launched. On Unix that is the whole story here — `kscreen-doctor`, `systemctl`,
/// `pw-dump` and friends are single processes we exec directly, and none of them forks a worker
/// that outlives it.
/// process we launched. That is never the whole story — see the Unix twin below for why it is not
/// enough there either — but Windows is where it fails hardest.
///
/// On Windows it is not, because there is no direct exec: every helper is reached through a shell
/// On Windows there is no direct exec: every helper is reached through a shell
/// (`cmd /c …`, `powershell -Command "… | pnputil …"`), so the process that actually hangs is a
/// **grandchild**. Killing the shell leaves it running — holding the stdio handles and the working
/// directory it inherited from us — and a budget that leaves that behind has not bounded anything.
@@ -273,6 +397,10 @@ mod tree {
JOB_OBJECT_LIMIT_KILL_ON_JOB_CLOSE,
};
/// Nothing to arrange before the spawn: job membership is assigned to the live process, so
/// [`Guard::attach`] does all of it. The Unix twin has to act here instead.
pub(super) fn prepare(_cmd: &mut std::process::Command) {}
/// Owns a Job object holding the spawned helper and everything it spawns. `None` when the job
/// could not be set up (see the module doc: degrade, don't fail).
pub(super) struct Guard(Option<HANDLE>);
@@ -352,18 +480,62 @@ mod tree {
}
}
/// The Unix half: `Child::kill` already ends the only process there is (see the Windows module doc
/// for why that is not true there). Kept as a real type rather than `cfg`ing the call sites, so the
/// two platforms read as one flow.
/// The Unix half — a **process group**, which is what Unix offers in place of a Job object.
///
/// This used to be an empty stub whose doc said `Child::kill` "already ends the only process there
/// is". Most Linux helpers here really are a single exec — `kscreen-doctor`, `pw-dump`, `hyprctl`,
/// `swaymsg` — but not all of them: `systemd-run --user` and `systemctl --user` do their work
/// through the user manager, which forks the actual process, so what hangs is routinely something
/// `Child::kill` cannot reach. With the reader threads in [`output_within`] waiting until every
/// write end of a pipe closes, one surviving relative is all it takes to keep a "bounded" call
/// going, which is why the group exists here too.
///
/// [`prepare`] puts the child in a new process group (it becomes the leader, so the group id is its
/// pid) and [`Guard::terminate`] `killpg`s that group, reaching every descendant that has not
/// deliberately left it. `process_group` changes only the group — not the session — so the helper
/// keeps its controlling terminal and login session, which anything doing a logind/polkit session
/// lookup depends on. Note the limit that follows from this and is NOT closed here: a process the
/// **user manager** forks on our behalf (`systemd-run --pipe`, whose transient unit inherits our
/// pipe write ends) is in another group and session by construction, so `killpg` misses it — see
/// [`DRAIN_GRACE`] for how the reader side is bounded in spite of that. The crate's one privileged
/// path, `pkexec` for the DM helper, deliberately does not come through this module at all: it
/// calls `Command::output()` directly and is documented as unbounded, because a `stop`/`restore`
/// verb legitimately takes seconds and killing it mid-flight is worse than waiting.
///
/// Best-effort in the same way as the Windows half: a failed `killpg` is ignored, and the
/// single-process `Child::kill` on the timeout path still runs.
#[cfg(not(windows))]
mod tree {
pub(super) struct Guard;
use std::os::unix::process::CommandExt;
/// The child's process-group id, captured while the child is still ours to reap.
pub(super) struct Guard(Option<i32>);
/// Make the child the leader of its own process group, so its descendants are reachable as one.
pub(super) fn prepare(cmd: &mut std::process::Command) {
cmd.process_group(0);
}
impl Guard {
pub(super) fn attach(_child: &std::process::Child) -> Self {
Self
pub(super) fn attach(child: &std::process::Child) -> Self {
// `prepare` asked for `process_group(0)`, so the group id IS the child's pid.
Self(i32::try_from(child.id()).ok())
}
/// End every process still in the group. A no-op once they have all exited, so this is safe
/// to call on the success path as well as the timeout one.
pub(super) fn terminate(&self) {
let Some(pgid) = self.0 else { return };
// `killpg` is a signal to a group we created and whose leader is the child we spawned;
// it cannot name a process we did not start. The one theoretical hazard is pid reuse
// between the leader's reap and this call, which needs a brand-new process to land on
// exactly that pid AND be a group leader — Linux hands out pids sequentially to
// `pid_max`, so there is no window to speak of, and the alternative (not killing) is
// the unbounded wait this module exists to prevent.
// SAFETY: a plain signal send by group id. No pointer is passed, nothing is aliased,
// and the result is deliberately ignored — ESRCH just means the group is already gone.
unsafe { libc::killpg(pgid, libc::SIGKILL) };
}
pub(super) fn terminate(&self) {}
}
}
@@ -390,6 +562,43 @@ mod tests {
);
}
/// A helper whose output exceeds one pipe buffer must still be captured IN FULL.
///
/// This is the case that fails against a `wait_with_output`-after-exit implementation: the
/// child blocks in `write()` with the pipe full, never exits, and the budget turns a perfectly
/// successful query into a `TimedOut` with its output thrown away. 1 MiB is ~16× a Linux pipe
/// (64 KiB) and ~64× the smallest macOS one, so it cannot be absorbed by a buffer on either.
#[test]
fn a_child_that_outruns_the_pipe_buffer_is_captured_in_full() {
const BYTES: usize = 1024 * 1024;
let mut cmd = Command::new("sh");
cmd.arg("-c")
.arg(format!("yes punktfunk | head -c {BYTES}; echo done >&2"));
let out = output_within(&mut cmd, Duration::from_secs(20)).expect("must not time out");
assert!(out.status.success(), "helper failed: {:?}", out.status);
assert_eq!(out.stdout.len(), BYTES, "stdout was truncated");
assert_eq!(String::from_utf8_lossy(&out.stderr).trim(), "done");
}
/// A helper that exits while a background child of its own still holds the pipe must not park
/// the caller: the reader waits for EOF on ALL write ends, so the grandchild's copy is what
/// would keep it there. Ending the process group is what closes it — and the collection is
/// bounded ([`DRAIN_GRACE`]) so that even the one relative a `killpg` cannot reach (a unit the
/// user manager forked for us) costs a detached thread rather than the calling thread.
#[test]
fn a_grandchild_holding_the_pipe_does_not_park_the_caller() {
let started = Instant::now();
let mut cmd = Command::new("sh");
cmd.arg("-c").arg("sleep 30 & echo punktfunk");
let out = output_within(&mut cmd, Duration::from_secs(10)).expect("the helper exited");
assert_eq!(String::from_utf8_lossy(&out.stdout).trim(), "punktfunk");
assert!(
started.elapsed() < Duration::from_secs(5),
"the call waited on the grandchild's EOF (took {:?})",
started.elapsed()
);
}
/// The normal path is unaffected: a quick command still yields its status and its output.
#[test]
fn a_quick_child_returns_normally() {
File diff suppressed because it is too large Load Diff
+25 -21
View File
@@ -77,29 +77,22 @@ fn pick_gamescope_mode(
}
}
/// Route input to match the chosen video backend (they must not diverge), via the highest-priority
/// `PUNKTFUNK_INPUT_BACKEND` knob the injector honors. For gamescope the sub-mode ladder
/// ([`pick_gamescope_mode`]) selects **managed** (a host-managed session at the client's mode —
/// tears the TV's autologin down on connect, restored on a debounced idle; only where
/// session-plus/SteamOS actually exists), **attach** (mirror a running gamescope at its own mode;
/// explicit via `PUNKTFUNK_GAMESCOPE_ATTACH`/`PUNKTFUNK_GAMESCOPE_NODE`, or the fallback for a
/// foreign gamescope on an infra-less box), or **bare spawn** (a per-session headless gamescope
/// nesting the session's launch command — the plain-distro default). `PUNKTFUNK_GAMESCOPE_MANAGED`
/// forces managed over all of it.
/// The operator's gamescope overrides, sampled ONCE — before this module has written anything.
/// The operator's gamescope overrides, sampled ONCE — at first use, and never written back.
///
/// [`apply_input_env`] both WRITES `PUNKTFUNK_GAMESCOPE_NODE`/`_SESSION` (to publish the sub-mode it
/// chose) and READS them as operator overrides. Reading them live therefore fed the ladder its own
/// previous output: the Attach arm sets `_NODE=auto`, and `node_env` sits at rung 2 of
/// `apply_input_env` used to both WRITE `PUNKTFUNK_GAMESCOPE_NODE`/`_SESSION` (to publish the
/// sub-mode it chose) and READ them as operator overrides. Reading them live therefore fed the
/// ladder its own previous output: the Attach arm set `_NODE=auto`, and `node_env` sits at rung 2 of
/// [`pick_gamescope_mode`] — ABOVE `dedicated_launch` at rung 3 — so one Attach decision latched
/// Attach for the rest of the host's life and silently overrode `game_session=dedicated`. Only rung
/// 1 (`_MANAGED`) could escape, because the Spawn arm that would clear the keys sits below the rung
/// that by then always fired.
///
/// Sampling at first use keeps the override's actual meaning — "the operator set this before we
/// ran" — and makes it immune to our own writes. The live reads that remain
/// ([`launch_is_nested`], gamescope's `poolable_now`) are deliberate: those consume the PUBLISHED
/// decision, which is what the keys carry after this function has run.
/// ran". Nothing publishes these keys any more (see [`resolve_gamescope_route`]): the resolved
/// decision travels as a [`GamescopeRoute`] VALUE carried on the backend instance, and every
/// consumer takes it that way — [`launch_is_nested`] by parameter, gamescope's `poolable_now` off
/// `self.route`, `crate::gamescope_hdr_available` by re-resolving the ladder. A change that
/// "restores" the write to serve some reader would restore the latch with it.
#[cfg(target_os = "linux")]
static OPERATOR_GAMESCOPE: std::sync::OnceLock<OperatorGamescope> = std::sync::OnceLock::new();
@@ -138,6 +131,16 @@ fn operator_gamescope() -> &'static OperatorGamescope {
})
}
/// Route input to match the chosen video backend (they must not diverge), via the highest-priority
/// `PUNKTFUNK_INPUT_BACKEND` knob the injector honors.
///
/// For gamescope the sub-mode ladder ([`pick_gamescope_mode`]) selects **managed** (a host-managed
/// session at the client's mode — tears the TV's autologin down on connect, restored on a debounced
/// idle; only where session-plus/SteamOS actually exists), **attach** (mirror a running gamescope at
/// its own mode; explicit via `PUNKTFUNK_GAMESCOPE_ATTACH`/`PUNKTFUNK_GAMESCOPE_NODE`, or the
/// fallback for a foreign gamescope on an infra-less box), or **bare spawn** (a per-session headless
/// gamescope nesting the session's launch command — the plain-distro default).
/// `PUNKTFUNK_GAMESCOPE_MANAGED` forces managed over all of it.
///
/// Returns the resolved [`GamescopeRoute`] when `chosen` is gamescope — the caller must carry it to
/// the backend instance via `VirtualDisplay::set_gamescope_route`. It is a RETURN VALUE and no
@@ -449,11 +452,12 @@ mod tests {
assert_eq!(pick(true, false, false, true, false, false, false), Attach);
}
/// The ladder must not be able to read back its own output. `apply_input_env`'s Attach arm
/// writes `PUNKTFUNK_GAMESCOPE_NODE=auto`, and `node_env` outranks `dedicated_launch` — so when
/// the override was read live, one Attach latched Attach for the host's lifetime and silently
/// overrode `game_session=dedicated`. Sampling once is what breaks the loop; this pins that the
/// sample does not move when the key is written afterwards.
/// The ladder must not be able to read back its own output. `apply_input_env`'s Attach arm used
/// to write `PUNKTFUNK_GAMESCOPE_NODE=auto`, and `node_env` outranks `dedicated_launch` — so
/// while the override was read live, one Attach latched Attach for the host's lifetime and
/// silently overrode `game_session=dedicated`. Sampling once is what breaks the loop, and it is
/// what makes restoring the write a non-event rather than a relapse; this pins that the sample
/// does not move when the key is written afterwards.
#[test]
#[cfg(target_os = "linux")]
fn operator_overrides_do_not_see_our_own_writes() {
+143 -12
View File
@@ -58,22 +58,21 @@ pub fn observe_session_instance(active: &ActiveSession) {
let changed = {
let mut last = LAST_INSTANCE.lock().unwrap_or_else(|e| e.into_inner());
let prev = *last;
*last = Some(cur);
// A `None` scan result is NOT an observation (see [`classify_instance_change`]), so it must
// not become the baseline either: recording it would make the NEXT poll — the one that sees
// the still-running desktop again — read as `None → DesktopKde`, i.e. a fresh instance, and
// bump the epoch out from under every pooled display. Leave the baseline on the last REAL
// instance and a transient miss is fully inert, in both directions.
if cur.0 != ActiveKind::None {
*last = Some(cur);
}
prev
};
if let Some(prev) = changed {
// Only a **desktop** compositor (KWin / Mutter / wlroots) instance change bumps the epoch +
// invalidates its kept displays — its PipeWire node dies with the compositor. A **gamescope**
// session (`ActiveKind::Gaming`) is NOT the epoch's subject: the box's game-mode / managed
// gamescope isn't pooled, and dedicated **spawns** are independent nested sessions whose nodes
// outlive any active-session change. So a game-mode gamescope restart, a Gaming↔Gaming winning-PID
// flap (e.g. B1 stopping the autologin before a dedicated spawn), or a coexisting-gamescope set
// change must NOT bump/invalidate — that would tear down a live/kept dedicated session (review
// findings #6/#7/#10). Gate the whole action on a desktop kind being involved.
if prev != cur && (is_desktop_kind(prev.0) || is_desktop_kind(cur.0)) {
if let InstanceChange::NewInstance { invalidate } = classify_instance_change(prev, cur) {
// Invalidate only the OLD backend, and only if it was a desktop compositor (never gamescope).
if is_desktop_kind(prev.0) {
if let Some(old) = compositor_for_kind(prev.0) {
if let Some(old_kind) = invalidate {
if let Some(old) = compositor_for_kind(old_kind) {
registry::invalidate_backend(old.id());
}
// The dead desktop's socket vars may still sit in the systemd --user manager env
@@ -95,6 +94,54 @@ pub fn observe_session_instance(active: &ActiveSession) {
}
}
/// What a `prev` → `cur` observation means for the session epoch — the pure core of
/// [`observe_session_instance`], so the (surprisingly load-bearing) rules below are unit-tested
/// without the process-global baseline.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
enum InstanceChange {
/// The same instance, or a change the epoch does not track — do nothing.
Nothing,
/// A new compositor instance: bump the epoch. `invalidate` names the OUTGOING desktop
/// compositor whose kept displays must be dropped (its PipeWire nodes died with it); `None`
/// when the outgoing session was gamescope / nothing, which owns no pooled displays.
NewInstance { invalidate: Option<ActiveKind> },
}
/// The epoch's rules, in one place:
///
/// * A `cur` of [`ActiveKind::None`] is **never** a change. `detect_active_session` answers `None`
/// both for "no graphical session is running" and for a scan that simply saw nothing — its whole
/// probe hangs off `if let Ok(entries) = std::fs::read_dir("/proc")`, and every per-PID rung
/// (`metadata`, `match_name`) can lose a race with a re-exec. Treating that as "the desktop
/// changed" ran `registry::invalidate_backend`, which removes pool entries in ANY lifecycle state
/// — Active ones included — so one unlucky `/proc` read tore down displays that were mid-stream,
/// and scrubbed the live session's socket vars out of the systemd `--user` manager on the way.
/// A real logout is picked up by the NEXT real observation (a different kind, or the same kind at
/// a new PID), which is the evidence-carrying end of the same transition.
/// * Only a **desktop** compositor (KWin / Mutter / wlroots) instance change counts. A **gamescope**
/// session ([`ActiveKind::Gaming`]) is not the epoch's subject: the box's game-mode / managed
/// gamescope isn't pooled, and dedicated **spawns** are independent nested sessions whose nodes
/// outlive any active-session change. So a game-mode gamescope restart, a Gaming↔Gaming
/// winning-PID flap (e.g. B1 stopping the autologin before a dedicated spawn), or a
/// coexisting-gamescope set change must NOT bump/invalidate — that would tear down a live/kept
/// dedicated session (review findings #6/#7/#10).
/// * A same-kind PID change IS a change: a fresh KWin's node-id space is unrelated to the dead
/// one's (A4).
fn classify_instance_change(
prev: (ActiveKind, Option<u32>),
cur: (ActiveKind, Option<u32>),
) -> InstanceChange {
if cur.0 == ActiveKind::None
|| prev == cur
|| !(is_desktop_kind(prev.0) || is_desktop_kind(cur.0))
{
return InstanceChange::Nothing;
}
InstanceChange::NewInstance {
invalidate: is_desktop_kind(prev.0).then_some(prev.0),
}
}
/// Counterpart to [`settle_desktop_portal`]'s `import-environment`: drop the desktop session's
/// socket vars from the systemd `--user` manager env once that desktop instance is GONE. They
/// persist in the manager otherwise, and every later user unit inherits them — including
@@ -702,6 +749,90 @@ pub fn settle_desktop_portal(chosen: Compositor) {
#[cfg(not(target_os = "linux"))]
pub fn settle_desktop_portal(_chosen: Compositor) {}
/// The epoch rules are platform-neutral (they are pure over [`ActiveKind`] + PID), so — unlike the
/// `/proc`-and-socket tests below — these run on every host this crate builds on.
#[cfg(test)]
mod instance_change_tests {
use super::*;
/// The 10.9 regression: a scan that answered `None` while KDE was in fact still up used to
/// satisfy `is_desktop_kind(prev)` and run the full invalidate — which drops pool entries in
/// ANY state, live streaming ones included.
#[test]
fn a_none_observation_is_never_a_change() {
for prev in [
(ActiveKind::DesktopKde, Some(42)),
(ActiveKind::DesktopGnome, Some(7)),
(ActiveKind::Gaming, Some(9)),
(ActiveKind::None, None),
] {
assert_eq!(
classify_instance_change(prev, (ActiveKind::None, None)),
InstanceChange::Nothing,
"a None scan result must not invalidate {prev:?}"
);
}
}
#[test]
fn a_desktop_swap_invalidates_the_outgoing_desktop() {
assert_eq!(
classify_instance_change(
(ActiveKind::DesktopKde, Some(1)),
(ActiveKind::DesktopGnome, Some(2))
),
InstanceChange::NewInstance {
invalidate: Some(ActiveKind::DesktopKde)
}
);
// Desktop → gamescope (Game Mode): the dead KWin's kept displays go with it.
assert_eq!(
classify_instance_change(
(ActiveKind::DesktopKde, Some(1)),
(ActiveKind::Gaming, Some(2))
),
InstanceChange::NewInstance {
invalidate: Some(ActiveKind::DesktopKde)
}
);
// gamescope → desktop: a new epoch, but gamescope owns no pooled entries to invalidate.
assert_eq!(
classify_instance_change(
(ActiveKind::Gaming, Some(1)),
(ActiveKind::DesktopKde, Some(2))
),
InstanceChange::NewInstance { invalidate: None }
);
}
#[test]
fn a_same_kind_restart_is_a_new_instance_but_a_gamescope_flap_is_not() {
// A fresh KWin (new PID) has an unrelated node-id space — A4.
assert_eq!(
classify_instance_change(
(ActiveKind::DesktopKde, Some(1)),
(ActiveKind::DesktopKde, Some(2))
),
InstanceChange::NewInstance {
invalidate: Some(ActiveKind::DesktopKde)
}
);
// The same instance re-detected: inert.
assert_eq!(
classify_instance_change(
(ActiveKind::DesktopKde, Some(1)),
(ActiveKind::DesktopKde, Some(1))
),
InstanceChange::Nothing
);
// Gaming↔Gaming winning-PID flap: never the epoch's business (findings #6/#7/#10).
assert_eq!(
classify_instance_change((ActiveKind::Gaming, Some(1)), (ActiveKind::Gaming, Some(2))),
InstanceChange::Nothing
);
}
}
#[cfg(all(test, target_os = "linux"))]
mod tests {
use super::*;
@@ -1,21 +1,19 @@
//! Host-lifetime virtual-display **ownership model** (Goal-1 §2.5). One reference-counted monitor
//! lifecycle, shared by both Windows backends (SudoVDA + pf-vdisplay) instead of the two verbatim-
//! duplicated `MGR: Mutex<Mgr>` globals each backend used to carry.
//! lifecycle, born as the shared half of two Windows backends (SudoVDA + pf-vdisplay) so the two
//! verbatim-duplicated `MGR: Mutex<Mgr>` globals could go; the SudoVDA backend has since been
//! removed, so pf-vdisplay is the sole driver behind the seam.
//!
//! [`VirtualDisplayManager`] owns the earned Idle/Active/Lingering refcount machine + the linger timer +
//! a **typed** [`OwnedHandle`] control device (no more raw `isize` smuggled across the pinger/linger
//! threads). The backend differences — the IOCTL protocol and the per-monitor REMOVE key — are the only
//! threads). The driver-specific part — the IOCTL protocol and the per-monitor REMOVE key — is the only
//! thing behind the [`VdisplayDriver`] seam; the state machine, the render-adapter pin decision, the
//! GDI/CCD glue (`pf_win_display::win_display`), and the generation-stamped [`MonitorLease`] are backend-neutral.
//! GDI/CCD glue (`pf_win_display::win_display`), and the generation-stamped [`MonitorLease`] are driver-neutral.
//!
//! It's a process-wide singleton ([`vdm`]) initialised once with the chosen backend's driver — the
//! host runs exactly one virtual-display backend per process. The session holds a [`MonitorLease`];
//! It's a process-wide singleton ([`vdm`]) initialised once with the driver — the host runs exactly
//! one virtual-display backend per process. The session holds a [`MonitorLease`];
//! its `Drop` releases the refcount (a *stale* lease — its monitor was preempted + recreated under it —
//! is a no-op, so it can never tear down the live monitor).
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use std::collections::BTreeMap;
use std::os::windows::io::{AsRawHandle, FromRawHandle, OwnedHandle};
use std::sync::atomic::{AtomicBool, AtomicU32, AtomicU64, Ordering};
@@ -86,7 +84,28 @@ struct Monitor {
/// is why WUDFHost death is ALL-slot shared fate.
wudf_pid: u32,
gdi_name: Option<String>,
/// The mode the OS actually COMMITTED for this monitor, not the one the client asked for — all
/// three paths that write it (create, re-arrival, in-place resize) read it back through
/// [`committed_mode_or`]. It is what `output_for` hands the capturer as `preferred_mode` and
/// what `/display/state` reports, so a requested-but-never-committed refresh here mis-paces the
/// encoder. It is NOT, on its own, the resize discriminator — see `requested_mode` below.
mode: Mode,
/// The mode the monitor was ASKED for at its last ADD/mode-set — the client's negotiated mode,
/// verbatim.
///
/// Kept beside the committed one because [`needs_resize`] is a two-sided question and `mode`
/// alone cannot answer it. `set_active_mode` deliberately commits the highest advertised refresh
/// <= the requested one rather than lose the client's resolution, so on a box that will not
/// advertise the negotiated rate (5120x1440@240 = 1.77 Gpix/s is the documented example) the two
/// fields PERMANENTLY disagree. If `acquire` diffed the incoming request against `mode` only,
/// every later acquire at the very mode the session already negotiated would read as a
/// mid-stream resize — and since `slot_id_for` keys on resolution alone, a refresh-only
/// divergence stays in the same slot and the divergence re-records itself on each pass. That is
/// not academic: `build_pipeline_with_retry` takes a retry-hold lease and then EVERY build
/// attempt re-`create`s the identical mode expecting a refcount++ join, so the slot would take
/// an in-place-resize attempt and then a full REMOVE→ADD hotplug per attempt — the exact churn
/// that exhausts the IddCx monitor-slot pool and wedges ADD at 0x80070490.
requested_mode: Mode,
/// The monitor id the driver actually resolved (the EDID serial / ConnectorIndex) — equals the
/// slot key when the per-client preference was honored, or the auto-allocated id (diagnostics).
resolved_monitor_id: u32,
@@ -165,7 +184,7 @@ struct GroupState {
ccd_exclusive: bool,
}
/// How a mid-stream re-arrival ([`ManagerInner::re_add`]) ended.
/// How a mid-stream re-arrival ([`VirtualDisplayManager::re_add`]) ended.
///
/// Three-way on purpose. `re_add` REMOVEs the old driver monitor before it ADDs the new one, so
/// once the ADD fails the old monitor is GONE — and the caller used to answer that by putting its
@@ -188,7 +207,7 @@ enum ReAdd {
/// What a NON-LAST-member teardown owes the group's topology.
///
/// Split out of [`ManagerInner::teardown_removed`] so the gate is testable without a driver, a CCD
/// Split out of [`VirtualDisplayManager::teardown_removed`] so the gate is testable without a driver, a CCD
/// device or a desktop — the Windows half of this crate has no other way to pin a decision.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
enum ShrinkAction {
@@ -200,11 +219,7 @@ enum ShrinkAction {
Nothing,
}
/// `ccd_exclusive` is the discriminator, NOT `ccd_saved.is_some()`: `Topology::Primary` stores a
/// snapshot too (from `set_virtual_primary_ccd`), so keying on the snapshot ran the EXCLUSIVE
/// isolate on a Primary group — clearing `DISPLAYCONFIG_PATH_ACTIVE` on every non-kept path, i.e.
/// blanking the very physical displays `Primary` exists to keep lit.
/// One stage of [`ManagerInner::resolve_target_gdi`]'s ladder: poll for the target's GDI name until
/// One stage of [`VirtualDisplayManager::resolve_target_gdi`]'s ladder: poll for the target's GDI name until
/// the 3 s ceiling. 50 ms sampling (latency plan P0.5) — a typical activation resolves on an early
/// poll, so finer sampling shaves ~150 ms off every stage crossing.
///
@@ -257,6 +272,15 @@ fn isolate_displays_ccd_seam(keep_target_ids: &[u32]) -> Option<SavedConfig> {
isolate_displays_ccd(keep_target_ids)
}
/// Decide the [`ShrinkAction`] a NON-LAST-member teardown owes the group.
///
/// `ccd_exclusive` is the discriminator, NOT `has_saved`: `Topology::Primary` stores a `ccd_saved`
/// snapshot too (from `set_virtual_primary_ccd`), so keying on the snapshot ran the EXCLUSIVE
/// isolate on a Primary group — clearing `DISPLAYCONFIG_PATH_ACTIVE` on every non-kept path, i.e.
/// blanking the very physical displays `Primary` exists to keep lit. (This paragraph had been
/// concatenated onto `poll_gdi_name`'s doc with no blank line between them, so the only written
/// record of the Phase-3.3 gate documented an unrelated polling helper and this fn read as
/// undocumented — a maintainer's invitation to "simplify" it back to the broken predicate.)
fn shrink_action(ccd_exclusive: bool, has_saved: bool) -> ShrinkAction {
if ccd_exclusive {
ShrinkAction::Reisolate
@@ -267,6 +291,69 @@ fn shrink_action(ccd_exclusive: bool, has_saved: bool) -> ShrinkAction {
}
}
/// The mode `target_id` is ACTUALLY running, for a caller about to RECORD it, with `requested` as
/// the fallback whenever the read-back cannot be trusted.
///
/// Every path that stores a `Monitor.mode` owes this call: `set_active_mode` deliberately commits
/// the highest advertised refresh <= the requested one rather than lose the client's resolution, and
/// the `wait_mode_settled` that precedes every store verifies the RESOLUTION only — so a `true`
/// settle is no evidence at all about the refresh. `active_mode`'s own doc states the contract:
/// "Callers that RECORD a mode must record this, or they claim a refresh the display is not
/// running."
///
/// Deliberately narrowed to the REFRESH. A read-back that FAILS, or that reports a different
/// RESOLUTION, keeps `requested`: the create path proceeds even when its settle timed out, so the
/// OS may still be sitting on its own default there, and recording that would hand the capturer +
/// the client a size nobody negotiated. The capturer already re-resolves the live size on its own
/// (`active_resolution` poll, game-capture GB1); the refresh is the field only this read-back can
/// answer.
fn committed_mode_or(target_id: u32, requested: Mode) -> Mode {
let Some((width, height, refresh_hz)) = pf_win_display::win_display::active_mode(target_id)
else {
return requested;
};
if (width, height) != (requested.width, requested.height) {
tracing::warn!(
target_id,
requested = format!("{}x{}", requested.width, requested.height),
active = format!("{width}x{height}"),
"the OS is not running the requested resolution after the settle — recording the \
requested mode (the capturer re-resolves the live size itself)"
);
return requested;
}
if refresh_hz != requested.refresh_hz {
tracing::info!(
target_id,
requested_hz = requested.refresh_hz,
committed_hz = refresh_hz,
"the OS committed a different refresh than requested (the driver does not advertise \
it) recording what the display actually runs"
);
}
Mode {
width,
height,
refresh_hz,
}
}
/// Does an acquire for `want` on a live monitor need a mid-stream resize, or can it JOIN?
///
/// Two modes describe one monitor and the caller may legitimately name either: `requested` is what
/// the session negotiated and re-asks for on every rebuild attempt, `committed` is what the OS
/// actually runs (they differ exactly when the driver would not advertise the negotiated refresh —
/// see [`committed_mode_or`]). Matching EITHER means there is nothing to do: asking again for the
/// negotiated mode cannot get a better result than the ADD already got, and asking for the mode the
/// display is already running is satisfied by definition. Only a genuinely NEW mode is a resize.
///
/// Keying on `committed` alone was the trap: it turned every same-mode re-acquire on a
/// refresh-clamping box into an in-place-resize attempt followed by a REMOVE→ADD hotplug, forever
/// (nothing ever re-records the negotiated rate, so the mismatch is self-perpetuating).
fn needs_resize(requested: Mode, committed: Mode, want: Mode) -> bool {
want != requested && want != committed
}
/// The manager's guarded state: the slot map + the (single) group record. One lock for both — every
/// group mutation happens on a slot transition, so splitting them would only invite lock-order bugs.
#[derive(Default)]
@@ -528,11 +615,11 @@ impl VirtualDisplayManager {
}
let reap = !slot.opened_once;
claim_instance()?;
// SAFETY: `VdisplayDriver::open` is `unsafe` only because it issues SetupAPI + `DeviceIoControl`
// FFI in the caller's apartment; the `device` mutex (held here) serializes it, so there is no
// concurrent open. `open` has no handle precondition to uphold, and the `OwnedHandle` it
// returns is the sole owner of the device.
let (handle, watchdog_s, driver_proto) = unsafe { self.driver.open(reap)? };
// `open` is a SAFE fn: it discharges every FFI precondition inside its own body (it opens the
// handle it then IOCTLs) and returns an `OwnedHandle` that is the sole owner of the device.
// The `device` mutex held here serializes racing opens — a *serialization* requirement, not
// a soundness one, which is exactly why it is not expressed as `unsafe`.
let (handle, watchdog_s, driver_proto) = self.driver.open(reap)?;
slot.opened_once = true;
self.watchdog_s.store(watchdog_s, Ordering::Relaxed);
self.driver_proto.store(driver_proto, Ordering::Relaxed);
@@ -690,11 +777,17 @@ impl VirtualDisplayManager {
// advertised mode list at ADD time, so we can't reach an arbitrary new mode in place — RE-
// ARRIVE the monitor at the exact mode instead (Fix 1). Own the slot for the swap: `re_add`
// needs `&mut inner` for the topology re-isolate, which the borrowed `mon` would block.
let cur_mode = match inner.slots.get(&slot) {
Some(SlotState::Active { mon, .. }) => mon.mode,
// Diff against BOTH of the slot's modes ([`needs_resize`]): the negotiated one the
// session keeps re-asking for and the one the OS actually committed. `mon.mode` alone is
// not the discriminator — on a box that clamps the negotiated refresh the two disagree
// for the monitor's whole life, and a request for the negotiated mode would then look
// like a resize on every acquire (`build_pipeline_with_retry` makes exactly that request
// once per build attempt, expecting a refcount++ join).
let (req_mode, cur_mode) = match inner.slots.get(&slot) {
Some(SlotState::Active { mon, .. }) => (mon.requested_mode, mon.mode),
_ => unreachable!("just matched Active"),
};
if cur_mode != mode {
if needs_resize(req_mode, cur_mode, mode) {
// IN-PLACE mode set first (latency plan P2): an already-advertised resolution
// (arrival list + the driver's same-id mode history) is CCD-forced on the SAME
// monitor — no REMOVE→ADD, so the monitor's OS identity (saved per-monitor DPI),
@@ -909,36 +1002,57 @@ impl VirtualDisplayManager {
let interval =
Duration::from_millis(self.watchdog_s.load(Ordering::Relaxed) as u64 * 1000 / 3);
let stop_t = stop.clone();
let thread = thread::spawn(move || {
let mut warned = false;
while !stop_t.load(Ordering::Relaxed) {
if let Some(h) = vdm().device_handle() {
// SAFETY: `ping` requires `dev` to be a valid control handle. The `h` Arc from
// `device_handle()` is held across this call, so the handle stays open even if
// it is retired concurrently — at worst the IOCTL fails (the retire drops only
// the manager's reference; see `DeviceSlot`). The pinger thread only spins
// while the `&'static` manager singleton lives.
match unsafe { vdm().driver.ping(dev_raw(&h)) } {
Ok(()) => warned = false,
Err(e) if is_device_gone(&e) => {
// The device itself is gone (driver upgrade / WUDFHost restart) — pings
// can only keep failing on this handle. Retire it so the next session's
// `ensure_device` reopens; the monitors are already dead driver-side.
vdm().invalidate_device(&e);
}
Err(e) => {
if !warned {
tracing::warn!(
"virtual-display keepalive PING failed (control handle lost?): {e:#}"
);
warned = true;
let thread = thread::Builder::new()
.name("vdisplay-pinger".into())
.spawn(move || {
let mut warned = false;
while !stop_t.load(Ordering::Relaxed) {
if let Some(h) = vdm().device_handle() {
// SAFETY: `ping` requires `dev` to be a valid control handle. The `h` Arc
// from `device_handle()` is held across this call, so the handle stays open
// even if it is retired concurrently — at worst the IOCTL fails (the retire
// drops only the manager's reference; see `DeviceSlot`). The pinger thread
// only spins while the `&'static` manager singleton lives.
match unsafe { vdm().driver.ping(dev_raw(&h)) } {
Ok(()) => warned = false,
Err(e) if is_device_gone(&e) => {
// The device itself is gone (driver upgrade / WUDFHost restart) —
// pings can only keep failing on this handle. Retire it so the next
// session's `ensure_device` reopens; the monitors are already dead
// driver-side.
vdm().invalidate_device(&e);
}
Err(e) => {
if !warned {
tracing::warn!(
"virtual-display keepalive PING failed (control handle lost?): {e:#}"
);
warned = true;
}
}
}
}
thread::sleep(interval);
}
thread::sleep(interval);
});
// NOT `thread::spawn` (which PANICS when the OS refuses the thread), for the same reason
// `ensure_exclusive_watch` was moved off it: this runs holding `pinger` and — via
// `create_monitor` ← `acquire` — the manager `state` guard, so an unwind here poisons the
// two locks the whole manager runs on, and every later `acquire`/`release`/`snapshot`
// `.lock().unwrap()` panics for the rest of the process. A missing pinger degrades to the
// driver's watchdog tearing the displays down (recoverable, and loud); a poisoned manager is
// neither. It also gains the thread name its two siblings already have.
let thread = match thread {
Ok(t) => t,
Err(e) => {
tracing::error!(
error = %e,
"could not spawn the virtual-display keepalive pinger — the driver's host-gone \
watchdog will tear this monitor down when it expires"
);
return;
}
});
};
*guard = Some(Pinger { stop, thread });
}
@@ -1167,9 +1281,14 @@ impl VirtualDisplayManager {
/// commits the target's path directly (supplied-config apply, the same thing display Settings
/// does), which doesn't consult the lid policy at all.
///
/// # Safety
/// Runs the CCD (QueryDisplayConfig / SetDisplayConfig) FFI; call under the `state` lock.
unsafe fn resolve_target_gdi(&self, target_id: u32) -> Option<String> {
/// Call under the `state` lock: this mutates the LIVE CCD topology (force-EXTEND, explicit path
/// activation), and the manager's sole-topology-mutator contract is what keeps two acquires from
/// interleaving path commits. A *serialization* requirement, not a soundness one — every CCD
/// helper it calls is a safe fn in `pf_win_display::win_display`, so this function performs no
/// unsafe operation at all. It was an `unsafe fn` back when the FFI was inline here, and stayed
/// one after the FFI moved out: three call sites then carried `unsafe {}` blocks whose SAFETY
/// proofs asserted things about FFI that is no longer in the body.
fn resolve_target_gdi(&self, target_id: u32) -> Option<String> {
// 50 ms sampling (latency plan P0.5): the SAME 3 s per-stage ceilings — the 3-stage ladder
// structure encodes real failure modes (headless auto-activate, integrated-panel clone,
// lid-closed path activation) and is untouched — but a typical activation resolves on an
@@ -1194,12 +1313,16 @@ impl VirtualDisplayManager {
/// (first member isolates and captures the restore; a later member re-issues the isolate with
/// the grown managed set — a sibling slot is never deactivated).
///
/// The returned `Monitor.mode` is what the OS COMMITTED, which need not be `mode` — see the
/// read-back after the settle. `Monitor.requested_mode` keeps `mode` verbatim, because that is
/// what the session re-asks for and `acquire`'s join/resize gate has to recognise.
///
/// # Safety
/// `dev` must be the live control handle.
unsafe fn create_monitor(
&'static self,
dev: HANDLE,
mode: Mode,
mut mode: Mode,
slot: u32,
client_hdr: Option<punktfunk_core::quic::HdrMeta>,
hw_cursor: bool,
@@ -1209,6 +1332,11 @@ impl VirtualDisplayManager {
// Windows reapplies the client's saved per-monitor config (DPI scaling) on reconnect;
// `0` (anonymous) = the driver auto-allocates the lowest-free id.
let preferred_id = slot;
// The client's negotiated mode, before the post-settle read-back below overwrites `mode`
// with what the OS committed. Both end up on the `Monitor`: the session re-asks for THIS one
// on every rebuild attempt, so it — not the committed one — is what `acquire`'s join/resize
// gate must recognise (see `Monitor::requested_mode`).
let requested_mode = mode;
let render_pin = resolve_render_pin();
// Hardware cursor only against a driver that implements the v5 channel: an older driver
// ignores the AddRequest field anyway (composited cursor), but gating here keeps the
@@ -1230,9 +1358,8 @@ impl VirtualDisplayManager {
// Resolve the capture target — wait for Windows to auto-activate the freshly-ADDed IDD into its
// OWN display path, with the integrated-screen clone fallback (shared by the re-arrival path).
// SAFETY: `resolve_target_gdi` runs the CCD FFI (a `Copy` `u32` target by value, owned return),
// under the `state` lock.
let gdi_name = unsafe { self.resolve_target_gdi(added.target_id) };
// Its `state`-lock discipline is satisfied: `acquire` holds the lock across this whole call.
let gdi_name = self.resolve_target_gdi(added.target_id);
match &gdi_name {
Some(n) => {
tracing::info!(
@@ -1368,6 +1495,17 @@ impl VirtualDisplayManager {
verified = settled,
"topology settle (verified-state wait)"
);
// Record what actually COMMITTED, not what was asked for — the same read-back
// `resize_in_place` does, for the same reason. `set_active_mode` deliberately falls
// back to the highest advertised refresh <= requested rather than lose the client's
// resolution, and `wait_mode_settled` verifies the RESOLUTION only, so `settled`
// says nothing about the refresh. Storing the request would make `mon.mode` claim a
// rate the display is not running: `output_for` hands the capturer that as
// `preferred_mode` (the encoder then paces to a rate the output never reaches),
// `/display/state` reports it, and the next Reconfigure diffs against it — a client
// re-requesting the rate it actually has would pay a needless resize, while one
// re-requesting the phantom rate takes the plain JOIN branch and never tries again.
mode = committed_mode_or(added.target_id, mode);
// EXPERIMENTAL `pnp_disable_monitors`, second selector (ANY topology): monitors
// that are connected but NOT part of the desktop — the standby TV/monitor the
@@ -1412,6 +1550,7 @@ impl VirtualDisplayManager {
wudf_pid: added.wudf_pid,
gdi_name,
mode,
requested_mode,
resolved_monitor_id: added.resolved_monitor_id,
position: (0, 0),
gen: self.gen.fetch_add(1, Ordering::Relaxed),
@@ -1507,29 +1646,10 @@ impl VirtualDisplayManager {
"in-place mode set did not commit within 1.5s (advertised after {advertised_ms} ms)"
);
}
// Record what actually COMMITTED, not what was asked for. `set_active_mode` deliberately
// falls back to the highest advertised refresh <= requested rather than lose the client's
// resolution, so `mon.mode = mode` claimed a rate the display might not be running — and
// `mon.mode` is what the next resize diffs against and what `/display/state` reports.
let committed = pf_win_display::win_display::active_mode(mon.target_id);
let landed = match committed {
Some((w, h, hz)) => Mode {
width: w,
height: h,
refresh_hz: hz,
},
// The settle above already verified the resolution; if the read-back races we still
// know the size took, so trust the request rather than leaving `mon.mode` stale.
None => mode,
};
if landed.refresh_hz != mode.refresh_hz {
tracing::info!(
requested_hz = mode.refresh_hz,
committed_hz = landed.refresh_hz,
"in-place resize: the OS committed a different refresh than requested (the driver \
does not advertise it) recording what it actually runs"
);
}
// Record what actually COMMITTED, not what was asked for — see [`committed_mode_or`], which
// the fresh-create and re-arrival paths share with this one so all three store the same
// truth: `mon.mode` is what `/display/state` reports and what the capturer paces to.
let landed = committed_mode_or(mon.target_id, mode);
tracing::info!(
advertised_ms,
settle_ms = settle_start.elapsed().as_millis() as u64,
@@ -1537,6 +1657,12 @@ impl VirtualDisplayManager {
"in-place resize committed (verified-state wait)"
);
mon.mode = landed;
// …and separately what was ASKED for, because that is what the session will re-request on
// its next acquire (a build retry, a build-then-drop overlap). Dropping it here would leave
// the slot only knowing a clamped refresh, and every such re-acquire would re-enter this
// function — or, once `wait_mode_advertised` refuses the un-advertised rate, the re-arrival
// hotplug below it.
mon.requested_mode = mode;
Ok(())
}
@@ -1607,7 +1733,7 @@ impl VirtualDisplayManager {
// values passed by value — no borrow crosses the call.
// SAFETY (both ADDs): `dev` is the live control handle; `render_pin`/`client_hdr` are owned
// `Copy`/`Option` values passed by value — no borrow crosses the call.
let (added, mode, rollback_err) = match unsafe {
let (added, mut mode, rollback_err) = match unsafe {
self.driver
.add_monitor(dev, mode, render_pin, slot, client_hdr, old.hw_cursor)
} {
@@ -1616,6 +1742,12 @@ impl VirtualDisplayManager {
// The old monitor is already REMOVEd, so there is nothing to "keep". Re-ADD it at
// the mode it had: the resize fails, but the session keeps streaming instead of
// being handed a slot whose driver monitor does not exist.
//
// At its REQUESTED mode, not its committed one — those differ exactly when the OS
// clamped the negotiated refresh, and this ADD is meant to replay the original one
// (same advertised mode list, same identity). Re-ADDing at the clamped rate would
// also make the clamp the slot's new negotiated mode, so the session's next acquire
// at the rate it still believes it has would read as yet another resize.
let e = e.context("re-arrival ADD at the new mode");
tracing::warn!(
slot,
@@ -1626,14 +1758,14 @@ impl VirtualDisplayManager {
match unsafe {
self.driver.add_monitor(
dev,
old.mode,
old.requested_mode,
render_pin,
slot,
client_hdr,
old.hw_cursor,
)
} {
Ok(a) => (a, old.mode, Some(e)),
Ok(a) => (a, old.requested_mode, Some(e)),
Err(e2) => {
tracing::error!(
slot,
@@ -1645,10 +1777,15 @@ impl VirtualDisplayManager {
}
}
};
// What the surviving ADD actually asked for (the new mode, or the old monitor's on a
// rollback) — pinned before the post-settle read-back below overwrites `mode` with the
// committed one. The session re-asks for THIS on its next acquire, so it is the join/resize
// gate's side of the pair (see `Monitor::requested_mode`).
let requested_mode = mode;
self.ensure_pinger();
// 3. Resolve the NEW target's GDI name (target_id changes across a re-arrival).
// SAFETY: CCD FFI over a `Copy` target id, under the `state` lock.
let gdi_name = unsafe { self.resolve_target_gdi(added.target_id) };
// 3. Resolve the NEW target's GDI name (target_id changes across a re-arrival). Under the
// `state` lock, as its topology-mutator discipline requires.
let gdi_name = self.resolve_target_gdi(added.target_id);
match &gdi_name {
Some(n) => {
tracing::info!(
@@ -1659,9 +1796,9 @@ impl VirtualDisplayManager {
// ADD only advertises the mode; force it active so DXGI/IDD captures the new size.
set_active_mode(n, mode);
// 4. Re-isolate the composited set with the NEW target replacing the old — preserving
// the group's first-member restore snapshot.
// SAFETY: CCD FFI over borrowed Copy target ids, under the `state` lock.
unsafe { self.reisolate_after_swap(inner, added.target_id) };
// the group's first-member restore snapshot. Under the `state` lock (the caller
// holds it and lent us `inner`), as its topology-mutator discipline requires.
self.reisolate_after_swap(inner, added.target_id);
// Topology settle before capture reopens: verified-state wait, ceiling = the old
// fixed 1500 ms sleep (latency plan P0.2 — the re-arrival twin).
let settle_start = std::time::Instant::now();
@@ -1671,6 +1808,12 @@ impl VirtualDisplayManager {
verified = settled,
"re-arrival topology settle (verified-state wait)"
);
// Store what COMMITTED, not what was asked for — the settle above verifies the
// resolution only, so it is no evidence about the refresh (see
// [`committed_mode_or`]). Doing this here rather than at the `Monitor` construction
// below keeps it on the arm where a path actually exists: with no GDI name there is
// no committed mode to read, and the request stands.
mode = committed_mode_or(added.target_id, mode);
}
None => tracing::warn!(
"re-arrival target {} not yet an active display path (auto-activate, EXTEND preset \
@@ -1688,6 +1831,7 @@ impl VirtualDisplayManager {
wudf_pid: added.wudf_pid,
gdi_name,
mode,
requested_mode,
resolved_monitor_id: added.resolved_monitor_id,
position: old.position,
gen: old.gen,
@@ -1708,9 +1852,11 @@ impl VirtualDisplayManager {
/// old slot has already been removed from the map by the caller, so `inner.target_ids()` is the
/// surviving siblings; the new target joins them.
///
/// # Safety
/// Drives the CCD topology FFI; call under the `state` lock.
unsafe fn reisolate_after_swap(&self, inner: &mut MgrInner, new_target: u32) {
/// Call under the `state` lock — it commits a new CCD topology, so it must not interleave with
/// another slot transition's commit. A *serialization* requirement, not a soundness one: every
/// helper it reaches (`isolate_displays_ccd_seam`, `set_virtual_primary_ccd`) is a safe fn, so
/// this body performs no unsafe operation. (`&mut MgrInner` already proves the lock is held.)
fn reisolate_after_swap(&self, inner: &mut MgrInner, new_target: u32) {
use crate::policy::Topology;
match topology_action() {
Topology::Exclusive => {
@@ -2271,7 +2417,15 @@ pub(crate) fn force_release(slot: Option<u64>) -> usize {
#[cfg(test)]
mod tests {
use super::{shrink_action, ShrinkAction};
use super::{needs_resize, shrink_action, Mode, ShrinkAction};
const fn m(width: u32, height: u32, refresh_hz: u32) -> Mode {
Mode {
width,
height,
refresh_hz,
}
}
/// The gate a non-last-member teardown keys off. It used to be `ccd_saved.is_some()`, which is
/// true for BOTH topologies — so a `Primary` group shrinking ran the exclusive isolate and
@@ -2297,4 +2451,35 @@ mod tests {
fn exclusivity_decides_without_a_snapshot() {
assert_eq!(shrink_action(true, false), ShrinkAction::Reisolate);
}
/// The join/resize gate on a box that CLAMPED the negotiated refresh (the driver would not
/// advertise 240 Hz at that pixel rate, so the OS committed 120). The session still re-requests
/// its negotiated mode on every build attempt — that must JOIN. Keying the gate on the committed
/// mode alone (which is what recording the read-back into `Monitor.mode` without keeping the
/// request amounts to) makes each of those a resize, i.e. an in-place attempt that fails on an
/// un-advertised rate and then a REMOVE→ADD hotplug, once per attempt, forever.
#[test]
fn a_reacquire_at_the_negotiated_mode_joins_even_when_the_os_clamped_the_refresh() {
let requested = m(5120, 1440, 240);
let committed = m(5120, 1440, 120);
assert!(
!needs_resize(requested, committed, requested),
"re-asking for the negotiated mode must JOIN, not hotplug the monitor"
);
// The other side of the pair: a client that re-asks for the rate the display actually runs
// has nothing to change either.
assert!(!needs_resize(requested, committed, committed));
// A genuinely new mode is still a resize — that is the branch's whole reason to exist.
assert!(needs_resize(requested, committed, m(3840, 2160, 120)));
assert!(needs_resize(requested, committed, m(5120, 1440, 60)));
}
/// The ordinary box (the OS advertises and commits exactly what was asked): both fields agree,
/// so the gate behaves exactly as the single-field one did.
#[test]
fn without_a_clamp_the_gate_is_plain_mode_equality() {
let mode = m(1920, 1080, 60);
assert!(!needs_resize(mode, mode, mode));
assert!(needs_resize(mode, mode, m(2560, 1440, 60)));
}
}
@@ -1,12 +1,20 @@
//! The backend-specific virtual-display **seam** (SudoVDA vs pf-vdisplay), carved out of the manager
//! (plan §W3): the REMOVE-key type, the `add_monitor` reply, and the IOCTL trait. This is the ONLY
//! thing that differs between the two Windows backends — the refcount machine, linger, pinger, and
//! CCD/GDI glue are all backend-neutral in [`super::VirtualDisplayManager`].
//! The virtual-display driver **seam**, carved out of the manager (plan §W3): the REMOVE-key type,
//! the `add_monitor` reply, and the IOCTL trait. It isolates the DRIVER's wire protocol from the
//! lifecycle — the refcount machine, linger, pinger and CCD/GDI glue are all driver-neutral in
//! [`super::VirtualDisplayManager`]. It was born as a two-backend seam (SudoVDA vs pf-vdisplay) and
//! has exactly one implementor since SudoVDA was removed: `crate::driver::PfVdisplayDriver` (the
//! flattened module name of `vdisplay/windows/pf_vdisplay.rs`). Kept as a trait because it is also
//! the only place the IOCTL surface can be faked, not because a second backend is expected.
use super::*;
/// The per-backend REMOVE key the driver stamps on ADD and consumes on REMOVE. SudoVDA keys monitors by
/// a fresh `GUID`; pf-vdisplay keys them by a monotonic `u64` session id.
/// The per-driver REMOVE key stamped on ADD and consumed on REMOVE. pf-vdisplay keys monitors by a
/// monotonic `u64` session id.
///
/// `Guid` is a RETAINED, UNUSED variant: it keyed SudoVDA's monitors (a fresh `GUID` per monitor) and
/// nothing constructs it since that backend was removed — the `else` arms in `pf_vdisplay`'s
/// `update_modes`/`remove_monitor` that reject it are therefore dead today. Left in place so the
/// enum still documents that the key is a per-driver choice rather than a `u64` by nature.
#[derive(Clone, Copy)]
pub(crate) enum MonitorKey {
Guid(windows::core::GUID),
@@ -29,10 +37,10 @@ pub(crate) struct AddedMonitor {
pub cursor_excluded: bool,
}
/// The backend-specific IOCTL surface — the *only* thing that differs between SudoVDA and pf-vdisplay.
/// Everything else (the refcount machine, the linger, the pinger, the CCD/GDI glue) is shared in
/// [`VirtualDisplayManager`]. `Send + Sync` because the manager (and so the boxed driver) is a
/// `&'static` singleton reached from the pinger + linger threads.
/// The driver's IOCTL surface — everything else (the refcount machine, the linger, the pinger, the
/// CCD/GDI glue) is driver-neutral and shared in [`VirtualDisplayManager`]. `Send + Sync` because the
/// manager (and so the boxed driver) is a `&'static` singleton reached from the pinger + linger
/// threads.
pub(crate) trait VdisplayDriver: Send + Sync {
fn name(&self) -> &'static str;
/// Find + open the control device, validate it (version handshake), and read the watchdog
@@ -42,9 +50,14 @@ pub(crate) trait VdisplayDriver: Send + Sync {
/// owned handle + watchdog seconds + the driver's reported protocol version (the in-place
/// resize gates on it).
///
/// # Safety
/// Issues setup-API + `DeviceIoControl` calls; runs in the caller's apartment.
unsafe fn open(&self, reap_orphans: bool) -> Result<(OwnedHandle, u32, u32)>;
/// SAFE, and owning — unlike every other method here, which takes the raw `dev` handle. It has
/// no caller obligation: it takes only a `bool`, opens the handle it then IOCTLs, and hands back
/// an `OwnedHandle` that closes on drop. It used to be an `unsafe fn` whose `# Safety` section
/// ("issues setup-API + `DeviceIoControl` calls; runs in the caller's apartment") restated what
/// the body does rather than naming anything a caller could uphold — an un-checkable proof
/// obligation at the one call site, which trains a reviewer to wave through the neighbouring
/// blocks where the `dev` precondition is real.
fn open(&self, reap_orphans: bool) -> Result<(OwnedHandle, u32, u32)>;
/// ADD a virtual monitor at `mode`, pinning the IDD render GPU to `render_luid` first if `Some`, and
/// requesting `preferred_monitor_id` (the host's per-client stable id; `0` = auto). `client_hdr`
/// is the CLIENT display's HDR volume for the monitor's EDID CTA HDR block (`None` = the
@@ -68,6 +81,8 @@ pub(crate) trait VdisplayDriver: Send + Sync {
/// The monitor is NOT departed; the caller CCD-forces the freshly-advertised mode afterwards.
/// The default errs so a backend without support routes to the re-arrival fallback.
///
// unsafe-fn-no-op-ok: trait method — the "dev is live" contract binds every impl; this
// default body is a stub that bails.
/// # Safety
/// `dev` must be the live control handle.
unsafe fn update_modes(&self, dev: HANDLE, key: &MonitorKey, mode: Mode) -> Result<()> {
@@ -85,3 +100,65 @@ pub(crate) trait VdisplayDriver: Send + Sync {
/// `dev` must be the live control handle.
unsafe fn ping(&self, dev: HANDLE) -> Result<()>;
}
#[cfg(test)]
mod tests {
use super::*;
/// A driver that implements nothing but the required methods — so the DEFAULTED `update_modes`
/// is what gets called.
struct FakeDriver;
impl VdisplayDriver for FakeDriver {
fn name(&self) -> &'static str {
"fake"
}
fn open(&self, _reap_orphans: bool) -> Result<(OwnedHandle, u32, u32)> {
anyhow::bail!("fake driver has no control device")
}
// unsafe-fn-no-op-ok: signature mandated by the trait; test stub.
unsafe fn add_monitor(
&self,
_dev: HANDLE,
_mode: Mode,
_render_luid: Option<LUID>,
_preferred_monitor_id: u32,
_client_hdr: Option<punktfunk_core::quic::HdrMeta>,
_hw_cursor: bool,
) -> Result<AddedMonitor> {
anyhow::bail!("fake driver adds no monitors")
}
// unsafe-fn-no-op-ok: signature mandated by the trait; test stub.
unsafe fn remove_monitor(&self, _dev: HANDLE, _key: &MonitorKey) -> Result<()> {
Ok(())
}
// unsafe-fn-no-op-ok: signature mandated by the trait; test stub.
unsafe fn ping(&self, _dev: HANDLE) -> Result<()> {
Ok(())
}
}
/// The `update_modes` default must ERR, not silently succeed: `resize_in_place` treats `Ok(())`
/// as "the driver refreshed the monitor's advertised mode list" and goes straight on to the CCD
/// force-set + settle — so a default that returned `Ok` would burn the full 1.5 s settle against
/// a mode list nobody updated, on every mid-stream resize, before falling back to the
/// re-arrival it should have taken immediately.
#[test]
fn the_defaulted_update_modes_reports_not_supported() {
let d = FakeDriver;
let mode = Mode {
width: 1920,
height: 1080,
refresh_hz: 60,
};
// SAFETY: the defaulted `update_modes` discharges its `dev` obligation by never using it —
// the body discards all three arguments and errs — so the null handle is never touched.
let err = unsafe { d.update_modes(HANDLE::default(), &MonitorKey::Session(1), mode) }
.expect_err("the default must not report success");
assert!(
err.to_string()
.contains("does not support in-place mode updates"),
"unexpected error text: {err:#}"
);
}
}
@@ -64,16 +64,27 @@ fn acquire_single_instance() -> Result<OwnedHandle> {
unsafe {
let h = match CreateMutexW(Some(&sa), false, w!("Global\\punktfunk-vdisplay-manager")) {
Ok(h) => h,
// The name exists but its creator's DACL denies this token the implicit OPEN (the SCM
// service creates it as SYSTEM; a second elevated-admin host lands here instead of in
// the ALREADY_EXISTS branch — validated on-glass). Legitimately that means an instance
// is live; it is ALSO exactly what a squat looks like, so say both.
// ACCESS_DENIED has THREE causes here and the handle alone cannot tell them apart, so
// name all three rather than assert one. (1) The name exists but its creator's DACL
// denies this token the implicit OPEN — the SCM service creates it as SYSTEM, so a
// second elevated-admin host lands here instead of in the ALREADY_EXISTS branch
// (validated on-glass); that is a live instance. (2) The same shape is exactly what a
// SQUAT looks like. (3) `CreateMutexW` also fails ACCESS_DENIED when the caller holds no
// SeCreateGlobalPrivilege at all — granted by default to Administrators, SYSTEM and the
// SERVICE groups but NOT to an ordinary interactive user, so an un-elevated
// `punktfunk-host serve` reaches this arm with no such object existing anywhere. Naming
// only (1)+(2) sent that operator hunting a process that does not exist and a
// `handle.exe` that finds nothing — the same misdiagnosis family as 2026-08-05 L-16,
// which this block exists to remove.
Err(e) if e.code().0 == 0x8007_0005u32 as i32 => anyhow::bail!(
"{IN_USE}\n\nIf no other punktfunk-host is running, the name \
`Global\\punktfunk-vdisplay-manager` has been SQUATTED by another process any \
account with SeCreateGlobalPrivilege can create it first and deny us access, \
which disables virtual-display streaming until that process exits. Find the \
holder with Sysinternals `handle.exe -a punktfunk-vdisplay-manager`."
"{IN_USE}\n\nIf no other punktfunk-host is running, either this process cannot \
create a `Global\\` kernel object at all (it needs SeCreateGlobalPrivilege run \
the host ELEVATED or as the installed service account; an ordinary interactive \
user does not hold it), or the name `Global\\punktfunk-vdisplay-manager` has been \
SQUATTED by another process any account with that privilege can create it first \
and deny us access, which disables virtual-display streaming until that process \
exits. Sysinternals `handle.exe -a punktfunk-vdisplay-manager` tells the two \
apart: a holder means a squat, NOTHING means the privilege."
),
Err(e) => {
return Err(e).context("CreateMutexW(punktfunk-vdisplay single-instance guard)");
@@ -190,6 +201,48 @@ fn object_owner_sid(h: HANDLE) -> Option<String> {
/// SYSTEM, BUILTIN\Administrators, or a member of the Administrators-owned set — the principals a
/// legitimate pf-vdisplay manager runs as.
///
/// Deliberately NARROW, and the narrowness is the security property: this predicate is what decides
/// whether an existing single-instance name is reported as "another punktfunk-host" (benign, wait it
/// out) or as a SQUAT (an attack on virtual-display availability). Widening it — `S-1-5-32-` as a
/// prefix, or any `S-1-5-21-…` domain account — silently reclassifies a non-administrative squatter
/// as one of ours and restores the exact misdiagnosis the 2026-08-05 L-16 fix removed. LocalService
/// (`S-1-5-19`) and NetworkService (`S-1-5-20`) are excluded ON PURPOSE: the plugin runner is forced
/// to LocalService, so a name owned by it is a plugin, not a host.
fn is_privileged_sid(sid: &str) -> bool {
matches!(sid, "S-1-5-18" | "S-1-5-32-544") || sid.starts_with("S-1-5-80-") // service SIDs
}
#[cfg(test)]
mod tests {
use super::is_privileged_sid;
/// Pins the classification above — the only pure decision in this module, and the one whose
/// widening is silent (nothing fails; a squat merely starts reading as a sibling host).
#[test]
fn is_privileged_sid_accepts_system_admins_and_service_sids_only() {
assert!(is_privileged_sid("S-1-5-18"), "SYSTEM");
assert!(is_privileged_sid("S-1-5-32-544"), "BUILTIN\\Administrators");
assert!(
is_privileged_sid("S-1-5-80-3139157870-2983391045-3678747466-658725712-1809340420"),
"an NT SERVICE\\… per-service SID"
);
assert!(!is_privileged_sid("S-1-5-32-545"), "BUILTIN\\Users");
assert!(
!is_privileged_sid("S-1-5-21-1004336348-1177238915-682003330-1001"),
"a local/domain user account"
);
// LocalService / NetworkService: the plugin runner's accounts, deliberately NOT ours.
assert!(!is_privileged_sid("S-1-5-19"), "LocalService");
assert!(!is_privileged_sid("S-1-5-20"), "NetworkService");
assert!(
!is_privileged_sid(""),
"an unreadable owner is never 'fine'"
);
// Prefix discipline: `S-1-5-80` without the trailing dash is a different SID string, and
// `S-1-5-8` (Proxy) must not slip in under a loosened prefix.
assert!(!is_privileged_sid("S-1-5-8"), "Proxy");
assert!(!is_privileged_sid("S-1-5-800-1"), "not a service SID");
}
}
@@ -2,25 +2,44 @@
//! carved out of the manager (plan §W3): the linger window, the keep-alive-forever pin, and the
//! per-monitor topology action. Pure readers of [`crate::policy`] + env — no manager state.
/// The historical Windows linger window, and the fallback for every rung that cannot answer.
const DEFAULT_LINGER_MS: u64 = 10_000;
/// Linger window before a session-less monitor is torn down. The console display-management policy
/// wins when configured (`keep_alive`); otherwise the legacy `PUNKTFUNK_MONITOR_LINGER_MS` env knob,
/// else the 10 s default.
pub(super) fn linger_ms() -> u64 {
use crate::policy::{prefs, Linger};
if let Some(eff) = prefs().configured_effective() {
return match eff.keep_alive.linger() {
Linger::Immediate => 0,
Linger::For(d) => d.as_millis() as u64,
// `forever` is handled BEFORE this by `keep_alive_forever()` in `release` (→ `Pinned`), so
// this arm is only reached defensively (e.g. a caller that resolves ms without the pin
// check) — fall back to the default rather than a huge linger.
Linger::Forever => 10_000,
};
resolve_linger_ms(
crate::policy::prefs()
.configured_effective()
.map(|eff| eff.keep_alive.linger()),
std::env::var("PUNKTFUNK_MONITOR_LINGER_MS")
.ok()
.and_then(|s| s.parse().ok()),
)
}
/// The precedence itself, lifted out of the readers so it is pinnable without a settings file, an
/// environment or a manager (this module's decisions are the ONLY ones on the Windows lifecycle path
/// that need neither a driver nor a desktop, and they had no tests at all).
///
/// `configured` is the console policy's resolved [`Linger`](crate::policy::Linger) (`None` = the
/// host was never configured), `env_ms` the parsed legacy knob. The configured policy outranks the
/// env knob entirely — an operator who set the console must not have it silently overridden by a
/// leftover variable.
fn resolve_linger_ms(configured: Option<crate::policy::Linger>, env_ms: Option<u64>) -> u64 {
use crate::policy::Linger;
match configured {
Some(Linger::Immediate) => 0,
Some(Linger::For(d)) => d.as_millis() as u64,
// `forever` is handled BEFORE this by `keep_alive_forever()` in `release` (→ `Pinned`), so
// this arm is only reached defensively (e.g. a caller that resolves ms without the pin
// check) — fall back to the default rather than a huge linger.
Some(Linger::Forever) => DEFAULT_LINGER_MS,
// Unconfigured: the legacy env knob, else the historical default. An unparseable value
// arrives here as `None` (the caller's `parse().ok()`), i.e. it reads as unset.
None => env_ms.unwrap_or(DEFAULT_LINGER_MS),
}
std::env::var("PUNKTFUNK_MONITOR_LINGER_MS")
.ok()
.and_then(|s| s.parse().ok())
.unwrap_or(10_000)
}
/// Whether the configured console policy's `keep_alive` resolves to **forever** (`Pinned`) — the
@@ -50,13 +69,79 @@ pub(super) fn exclusive_reassert_ms() -> u64 {
/// extended; `Primary` makes it primary while keeping the physical(s) active; `Exclusive` disables the
/// physical(s) so the IDD is the sole composited desktop.
pub(super) fn topology_action() -> crate::policy::Topology {
let configured = crate::policy::prefs()
.configured_effective()
.map(|_| crate::effective_topology());
resolve_topology_action(configured, std::env::var("PUNKTFUNK_NO_ISOLATE").is_ok())
}
/// The precedence for [`topology_action`], lifted out for the same reason as [`resolve_linger_ms`].
/// `configured` is [`crate::effective_topology`]'s answer when the console configured anything at
/// all (that fn is the rung responsible for never returning `Auto`); `no_isolate_env` is the legacy
/// `PUNKTFUNK_NO_ISOLATE` opt-out, which an unconfigured host still honors.
fn resolve_topology_action(
configured: Option<crate::policy::Topology>,
no_isolate_env: bool,
) -> crate::policy::Topology {
use crate::policy::Topology;
if crate::policy::prefs().configured_effective().is_some() {
return crate::effective_topology();
match configured {
Some(t) => t,
None if no_isolate_env => Topology::Extend,
None => Topology::Exclusive,
}
if std::env::var("PUNKTFUNK_NO_ISOLATE").is_ok() {
Topology::Extend
} else {
Topology::Exclusive
}
#[cfg(test)]
mod tests {
use super::{resolve_linger_ms, resolve_topology_action, DEFAULT_LINGER_MS};
use crate::policy::{Linger, Topology};
use std::time::Duration;
/// The console policy is the top rung: a host that configured `keep_alive` must not have it
/// silently overridden by a leftover `PUNKTFUNK_MONITOR_LINGER_MS`.
#[test]
fn configured_policy_beats_the_legacy_env_knob() {
assert_eq!(
resolve_linger_ms(Some(Linger::For(Duration::from_secs(3))), Some(60_000)),
3_000
);
assert_eq!(resolve_linger_ms(Some(Linger::Immediate), Some(60_000)), 0);
}
/// Unconfigured hosts keep the historical behavior: the env knob, else the 10 s default. An
/// unparseable value reaches this fn as `None` (the reader's `parse().ok()`), so it reads as
/// unset rather than as zero — a `linger_ms = 0` would tear the monitor down on every
/// disconnect.
#[test]
fn an_unconfigured_host_honours_the_env_knob_then_the_default() {
assert_eq!(resolve_linger_ms(None, Some(250)), 250);
assert_eq!(resolve_linger_ms(None, None), DEFAULT_LINGER_MS);
}
/// `Forever` is the `Pinned` lifecycle, resolved by `keep_alive_forever()` before any ms are
/// asked for; reaching this fn with it means a caller skipped the pin check, and the answer is
/// the default window — NOT an effectively infinite linger that would keep the physical panels
/// dark with nothing to release them.
#[test]
fn forever_resolves_to_the_default_not_a_huge_linger() {
assert_eq!(
resolve_linger_ms(Some(Linger::Forever), None),
DEFAULT_LINGER_MS
);
}
/// The unconfigured rungs are `Exclusive` by default, `Extend` under the legacy opt-out — and
/// neither is `Auto`, which the manager's `match` would treat as plain extend without ever
/// saying so.
#[test]
fn the_unconfigured_topology_rungs_never_yield_auto() {
assert_eq!(resolve_topology_action(None, false), Topology::Exclusive);
assert_eq!(resolve_topology_action(None, true), Topology::Extend);
// A configured host's answer is whatever `effective_topology()` resolved — passed through
// verbatim, env knob or not.
assert_eq!(
resolve_topology_action(Some(Topology::Primary), true),
Topology::Primary
);
}
}
@@ -8,14 +8,13 @@
//! the wire contract OWNED by [`pf_driver_proto::control`] (versioned + `#[repr(C)] Pod` structs,
//! NOT the SudoVDA ABI). No DLL, no named pipe. See `design/windows-host-rewrite.md`.
//!
//! This is a faithful clone of [`super::sudovda`] (the shipping fallback) repointed at the new driver:
//! same reference-counted/lingering monitor lifecycle, same CCD isolation + active-mode forcing — those
//! backend-NEUTRAL helpers are REUSED from `sudovda` (a pf-vdisplay monitor's `target_id` is a real OS
//! target id, so the CCD/DXGI code works unchanged). Only the driver-specific bits (GUID, IOCTL codes,
//! request/reply structs, the version handshake) differ, per `pf_driver_proto`.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
//! punktfunk's IddCx driver is the SOLE Windows backend — the legacy SudoVDA fallback was removed and
//! its driver is no longer shipped (`lib.rs`), so nothing here is a "clone of the fallback" any more.
//! The backend-NEUTRAL half — the reference-counted/lingering monitor lifecycle, the CCD isolation and
//! the active-mode forcing — lives in [`super::manager`] and `pf_win_display::win_display` (a
//! pf-vdisplay monitor's `target_id` is a real OS target id, so that CCD/DXGI code applies unchanged).
//! Only the driver-specific bits (GUID, IOCTL codes, request/reply structs, the version handshake) are
//! here, per `pf_driver_proto`.
use std::ffi::c_void;
use std::mem::size_of;
@@ -97,9 +96,10 @@ unsafe fn ioctl(h: HANDLE, code: u32, input: &[u8], output: &mut [u8]) -> Result
/// pinning an OS VidPN target against the IddCx adapter's fixed monitor-slot budget; once ~16 accumulate,
/// `IOCTL_ADD` wedges at 0x80070490 (`ERROR_NOT_FOUND`) and every session black-screens until a manual
/// reset/reboot. Removing the not-present PDOs frees the slots — the in-process equivalent of
/// `reset-pf-vdisplay.ps1` step 2 (proven on-box). Best-effort + idempotent: only NOT-present nodes
/// (`Status != OK`) are removed, so the LIVE session's monitor (`Status OK`) is never touched; any
/// failure is logged and swallowed. Returns the number removed.
/// `reset-pf-vdisplay.ps1` step 2 (proven on-box). Best-effort + idempotent: only ABSENT nodes
/// (`Present` false AND `Status` `Unknown`) are removed, so a LIVE session's monitor is never
/// touched — not even while it is in a transient problem state; any failure is logged and
/// swallowed. Returns the number removed.
///
/// The outcome is logged UNCONDITIONALLY, as found + removed: the old script counted only removals
/// and the host spoke only when that count was positive, so a reap whose pnputil never launched and
@@ -108,8 +108,17 @@ unsafe fn ioctl(h: HANDLE, code: u32, input: &[u8], output: &mut [u8]) -> Result
/// wedge with every sleep cycle.
fn reap_ghost_monitors() -> u32 {
// Mirrors reset-pf-vdisplay.ps1 step 2. powershell is always present for the SYSTEM service; the
// matched tokens ('OK', 'punktfunk', the InstanceId) are locale-invariant, so this is safe on a
// non-English box (unlike a .ps1 *file* read in the machine codepage).
// matched tokens ('Unknown', 'punktfunk', the InstanceId) are locale-invariant, so this is safe
// on a non-English box (unlike a .ps1 *file* read in the machine codepage).
//
// The selector asks about PRESENCE, not health — the exact complement of the liveness predicate
// the adapter reload below uses (`$_.Present -or $_.Status -ne 'Unknown'`). It used to read
// `Status -ne 'OK'`, which is a HEALTH field: `Error`, `Degraded` and `Unknown` all satisfy it,
// so a PRESENT virtual monitor in a transient problem state was handed to `pnputil
// /remove-device` — and this runs mid-session from `add_monitor`'s 0x80070490 recovery, i.e.
// while sibling sessions are live, so it could rip out a live client's monitor. `Present` is the
// authoritative bit; the `Status -eq 'Unknown'` conjunct is the guard for `Present` reading null
// (`-not $null` is TRUE, which alone would select every device on the box).
//
// pnputil is resolved by full path and `$LASTEXITCODE` pre-seeded to failure before every
// launch, exactly like the reload path below: a LocalSystem service's PATH need not include
@@ -117,7 +126,7 @@ fn reap_ghost_monitors() -> u32 {
// elevated), and the old bare-name call failed INVISIBLY there — `SilentlyContinue` swallowed
// the miss, no exit code was written, and the ghosts stayed to wedge `IOCTL_ADD` at 0x80070490.
const REAP_PS: &str = "$ErrorActionPreference='SilentlyContinue'; \
$g = @(Get-PnpDevice -Class Monitor | Where-Object { $_.Status -ne 'OK' -and $_.FriendlyName -match 'punktfunk' }); \
$g = @(Get-PnpDevice -Class Monitor | Where-Object { -not $_.Present -and $_.Status -eq 'Unknown' -and $_.FriendlyName -match 'punktfunk' }); \
$pnp = ($env:SystemRoot + '\\System32\\pnputil.exe'); \
$n = 0; foreach ($d in $g) { $LASTEXITCODE = 1; if (Test-Path $pnp) { & $pnp /remove-device $d.InstanceId *> $null }; if ($LASTEXITCODE -eq 0) { $n++ } }; \
Write-Output ($g.Count.ToString() + ' ' + $n)";
@@ -593,14 +602,20 @@ fn probe_device() -> Probe {
// SAFETY: `buf` is at least `required` bytes and aligned to 8 (so also to the struct's 4),
// so stamping `cbSize` and letting the API fill up to `required` bytes stays in bounds;
// `detail` aliases `buf` only within this iteration, and the `DevicePath` pointer is read
// before `buf` is dropped.
// before `buf` is dropped. That path pointer is taken as a RAW place projection off
// `detail`, so it keeps the whole `buf` allocation's provenance: `DevicePath` is declared
// `[u16; 1]` (a flexible-array-member stub), so `.as_ptr()` would auto-ref it and hand
// `CreateFileW` a pointer tagged for TWO bytes while the API reads the full NUL-terminated
// path (100+ bytes) — everything past `DevicePath[0]` out of bounds for that tag, and a
// compiler entitled to fold the zero-init back in and pass an EMPTY device name. Same
// defect class (and same fix) as the `MONITORINFOEXW` retag in `vdisplay/ddc.rs`.
let opened = unsafe {
(*detail).cbSize = size_of::<SP_DEVICE_INTERFACE_DETAIL_DATA_W>() as u32;
SetupDiGetDeviceInterfaceDetailW(hdev.0, &idata, Some(detail), required, None, None)
.context("SetupDiGetDeviceInterfaceDetailW(pf-vdisplay)")
.and_then(|()| {
CreateFileW(
PCWSTR((*detail).DevicePath.as_ptr()),
PCWSTR((&raw const (*detail).DevicePath).cast::<u16>()),
0xC000_0000, // GENERIC_READ | GENERIC_WRITE
FILE_SHARE_READ | FILE_SHARE_WRITE,
None,
@@ -635,7 +650,7 @@ impl VdisplayDriver for PfVdisplayDriver {
"pf-vdisplay"
}
unsafe fn open(&self, reap_orphans: bool) -> Result<(OwnedHandle, u32, u32)> {
fn open(&self, reap_orphans: bool) -> Result<(OwnedHandle, u32, u32)> {
// A short re-probe, and deliberately NO adapter reload — this replaces the second, impatient
// copy of the recovery that used to live here. Session bring-up already ran the full
// `ensure_available` before constructing the backend, so anything left for this open to
@@ -686,18 +701,42 @@ impl VdisplayDriver for PfVdisplayDriver {
);
}
let watchdog_s = info.watchdog_timeout_s.max(1);
if info.protocol_version < pf_driver_proto::PROTOCOL_VERSION {
// UNCONDITIONAL: this line is the only place the negotiated watchdog is reported, and the
// pinger's cadence (`watchdog/3`) is derived from it — yet it used to sit in the `else` of
// the version warning, so exactly the hosts where the number is worth having (anything but
// an exact-version pair) logged nothing at all.
tracing::info!(
"pf-vdisplay protocol {} (host drives {}..={}, watchdog timeout {}s)",
info.protocol_version,
pf_driver_proto::MIN_DRIVER_PROTOCOL_VERSION,
pf_driver_proto::PROTOCOL_VERSION,
watchdog_s
);
// Version-SPECIFIC capability gaps, reported independently. Every bump since v3 is ADDITIVE,
// so the old blanket `< PROTOCOL_VERSION` test named the WRONG gap: it told a v4 or v5
// driver it "lacks the in-place resize" — added IN v4 — purely because it was not v6. Each
// rung below names the capability the host actually gates on that version.
if info.protocol_version < 4 {
tracing::warn!(
"pf-vdisplay protocol {} (host supports {}): driver lacks the in-place resize \
mid-stream resizes use the monitor re-arrival path until the driver is updated",
info.protocol_version,
pf_driver_proto::PROTOCOL_VERSION
"pf-vdisplay protocol {}: driver lacks the in-place mid-stream resize \
(IOCTL_UPDATE_MODES, added in v4) every mid-stream resize costs a monitor \
re-arrival (one hotplug per switch) until the driver is updated",
info.protocol_version
);
} else {
}
if info.protocol_version < 5 {
tracing::warn!(
"pf-vdisplay protocol {}: driver lacks the IddCx hardware-cursor channel (added in \
v5) the pointer stays composited into the captured frame",
info.protocol_version
);
}
if info.protocol_version < 6 {
tracing::info!(
"pf-vdisplay protocol {} (watchdog timeout {}s)",
info.protocol_version,
watchdog_s
"pf-vdisplay protocol {}: driver lacks the mid-stream cursor-forward flip \
(IOCTL_SET_CURSOR_FORWARD, added in v6) the cursor model declared at monitor ADD \
stands for the whole session",
info.protocol_version
);
}
// Reap monitors orphaned by a crashed previous host — a FIRST-CLASS op (driver returns
+3 -4
View File
@@ -112,10 +112,9 @@
//! Unsafe posture: unlike pf-bitstream (which forbids unsafe outright), this crate
//! cannot — the `ash::vk::native` bindgen structs are zero-initialized the way the
//! encode side does it (`pf-encode/src/enc/linux/vk_build.rs`), and the GPU half is
//! Vulkan FFI. Every unsafe block therefore carries a written `// SAFETY:` proof,
//! enforced (and unlike the encoder there is NO file-level
//! `unsafe_op_in_unsafe_fn` exemption every operation is individually fenced):
#![deny(clippy::undocumented_unsafe_blocks)]
//! Vulkan FFI. Every unsafe block therefore carries a written `// SAFETY:` proof — enforced by
//! the workspace `[workspace.lints]` tables, and (unlike the encoder) with NO file-level
//! `unsafe_op_in_unsafe_fn` exemption: every operation is individually fenced.
pub mod caps;
pub mod caps_av1;
-2
View File
@@ -88,8 +88,6 @@
//! the readback geometry (row pitch / crop) or intra decode; mismatches that
//! only appear on later frames point at inter prediction / DPB management.
#![deny(clippy::undocumented_unsafe_blocks)]
mod common;
use ash::vk;
-2
View File
@@ -39,8 +39,6 @@
//! so releases pass `false`), soak, and both vendors' DPB arrangements at once
//! (each box exercises only its own).
#![deny(clippy::undocumented_unsafe_blocks)]
mod common;
use ash::vk;
@@ -28,9 +28,6 @@
//! suspects — without ever touching the CCD lock itself (the display-config lock is exactly what
//! stalls during churn; the capture thread must never block on it).
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it (unsafe-proof program).
#![deny(clippy::undocumented_unsafe_blocks)]
use std::collections::VecDeque;
use std::sync::{Mutex, Once, OnceLock};
use std::time::Instant;
-2
View File
@@ -12,8 +12,6 @@
// `win_display` has denied both unsafe-proof lints since its CCD helpers stopped being `unsafe fn`;
// hoist that to the crate root so the smaller modules (`input_desktop`, `monitor_devnode`,
// `display_events`) and any future one are covered by default rather than by remembering to opt in.
#![deny(clippy::undocumented_unsafe_blocks)]
#![deny(unsafe_op_in_unsafe_fn)]
#[cfg(target_os = "windows")]
pub mod display_events;

Some files were not shown because too many files have changed in this diff Show More