Compare commits

..
104 Commits
Author SHA1 Message Date
enricobuehler 1317901122 Merge pull request 'Uninstalling the Windows host left every audio device it minted behind forever — and the installer script documented that as a decision' (#145) from worktree-win-audio-uninstall-cleanup into main
android / android (push) Failing after 1m32s
ci / rust-arm64 (push) Successful in 1m53s
apple / swift (push) Successful in 1m34s
ci / bun-nix (push) Successful in 22s
ci / web (push) Successful in 1m48s
ci / docs-site (push) Successful in 1m44s
deb / build-publish-client-arm64 (push) Successful in 1m44s
deb / build-publish (push) Successful in 4m5s
ci / rust (push) Successful in 7m24s
apple / screenshots (push) Successful in 5m54s
arch / build-publish (push) Successful in 9m40s
deb / build-publish-host (push) Successful in 7m28s
windows-host / package (push) Successful in 13m52s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 20s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m53s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 18m0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 11s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 7s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 7s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 6s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 6s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / builders-arm64cross (push) Successful in 7s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m4s
docker / deploy-docs (push) Failing after 6m14s
Reviewed-on: #145
2026-08-09 19:30:31 +00:00
enricobuehler 7a9fa4501c Merge pull request 'Nobara could never use the patched gamescope — and the RPM it was told to install was unsigned' (#144) from worktree-gamescope-pin-bump-nobara into main
apple / swift (push) Successful in 1m35s
ci / rust-arm64 (push) Successful in 1m59s
android / android (push) Failing after 2m37s
ci / bun-nix (push) Successful in 17s
ci / docs-site (push) Successful in 1m19s
ci / web (push) Successful in 2m27s
deb / build-publish-client-arm64 (push) Successful in 1m54s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 8s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 10s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 13s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
deb / build-publish (push) Successful in 3m37s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
apple / screenshots (push) Successful in 5m43s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Failing after 41s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m2s
docker / deploy-docs (push) Skipped
arch / build-publish (push) Successful in 9m55s
deb / build-publish-host (push) Successful in 6m53s
docker / builders-arm64cross (push) Successful in 12s
ci / rust (push) Successful in 11m23s
flatpak / build-publish (push) Successful in 9m37s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 8m57s
windows-host / package (push) Successful in 18m53s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 16s
nix / flake (push) Failing after 17m1s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 24m26s
Reviewed-on: #144
2026-08-09 18:53:59 +00:00
enricobuehler d87a8df28d fix(windows): uninstall removes the audio devices the host mints
ci / bun-nix (pull_request) Successful in 25s
ci / web (pull_request) Successful in 1m21s
ci / docs-site (pull_request) Successful in 1m23s
apple / swift (pull_request) Successful in 1m42s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m37s
android / android (pull_request) Successful in 6m12s
ci / rust (pull_request) Successful in 7m37s
The field report: uninstalling punktfunk left "Punktfunk Speakers",
"Punktfunk Microphone" and the per-pad "Wireless Controller" endpoints
sitting in Windows' Sound settings forever.

They have no installer payload behind them, which is why nothing in the
uninstall touched them. The host mints them at RUNTIME as extra devnodes
on Valve's streaming-audio drivers, and both providers deliberately
re-resolve their devnode across restarts instead of re-minting it — so
they persist by design. Persistent across restarts must not mean
permanent: the .iss even documented leaving them behind as a decision.

New `driver uninstall --audio` leg (a third Inno [UninstallRun] entry,
after the two driver legs and well after `service uninstall`, since a
live host re-mints on its next wiring pass):

* restores the default playback device first, if a host that died
  mid-stream left it parked on our loopback sink — otherwise Windows
  re-picks by its own ranking rather than giving the operator back the
  device they had;
* removes every MEDIA-class devnode carrying one of our three durable
  owner markers (pad slot, minted role, probe), phantoms included;
* deletes each endpoint's MMDevices record, resolved through the
  devnode link BEFORE the devnode goes.

Marker-matched, never name-matched: our instances are name-identical to
Steam's own, and Steam's devnodes, its drivers, and a VB-CABLE from the
era when we bundled one carry no marker and stay untouched. A ROOT\
enumeration guard means a marker-shaped value on a real sound card can
never cost the user their hardware.

The registry half is best-effort: those keys are SYSTEM-owned and the
uninstaller runs elevated but as a user, so on a stock box the record
survives as an inert NOTPRESENT entry that Sound settings only shows
behind "Show Disconnected Devices". The device itself is gone either
way, and seizing ownership of SYSTEM registry keys from an uninstaller
is a worse thing to ship than that scrap.
2026-08-09 20:53:22 +02:00
enricobuehler 46390739d8 fix(pf-vdisplay): the box's OWN session unit needs the gamescope bind too
apple / swift (pull_request) Successful in 1m39s
apple / screenshots (pull_request) Skipped
android / android (pull_request) Successful in 5m40s
nix / flake (pull_request) Failing after 19m15s
ci / bun-nix (pull_request) Successful in 20s
ci / docs-site (pull_request) Successful in 1m1s
ci / web (pull_request) Successful in 1m11s
ci / rust-arm64 (pull_request) Successful in 1m28s
ci / rust (pull_request) Successful in 6m26s
`launch_session` spawns a transient unit and can hand `systemd-run` the
`BindReadOnlyPaths` directly, but a box that owns an autologin
`gamescope-session-plus@<client>.service` is RESTARTED IN PLACE instead — no `systemd-run`,
so that path kept running Nobara's hardcoded `/usr/bin/gamescope` and the previous commit
fixed only half the problem. Found on the box: after a reboot the host took the
`ensure_box_gamescope_mode` path (the autologin unit was live) rather than the managed one.

Deliver the same two fixes as a drop-in on that unit — the bind, and the WSI opt-out when the
box's layer was built for a different gamescope — plus `PF_HZ`/`PF_HDR_ARGS`, which the
wrapper reads and would otherwise default to 60 Hz. `daemon-reload` before the restart or
systemd runs the old unit. Best-effort: a failure to write it must not block a restart that
would otherwise work, and it is a no-op on a box already resolving to `/usr/bin/gamescope`.

⚠ REMOVED on restore, deliberately. Leaving it would put the patched build — and our HDR and
cursor flags — under the user's ORDINARY game mode, which is exactly what
`packaging/gamescope/README.md`'s "sits BESIDE the distro package" rule exists to prevent. The
bind is ours only for as long as we are driving the session.

`ensure_box_gamescope_mode` grows an `hdr` param to build those args; both call sites already
had it in scope (`self.hdr`, and `create_managed_session`'s parameter).

Gate: `scripts/xcheck.sh linux clippy` clean (0 warning/error lines), `cargo fmt` clean.
2026-08-09 20:28:00 +02:00
enricobuehler 3500e95660 fix(pf-vdisplay): make Nobara's session run the patched gamescope, and stop the WSI layer killing every client
Two independent reasons a Nobara box could never stream from a gamescope session,
both found on glass (VM 123, Nobara 44, RTX 5070 Ti).

**1. The session ran a stock gamescope, so the host refused it.**

Nobara's `gamescope-session-plus` builds its command as

    GAMESCOPECMD="/usr/bin/gamescope \

and reads `GAMESCOPE_BIN` NOWHERE. All three of our spawn levers miss at once: the env
var is ignored, and an absolute path cannot be redirected by a PATH shim. So the session
ran stock gamescope, the capability probe rejected it, and every session died with
"pipeline build failed (out of retries) … it ignored GAMESCOPE_BIN / the PATH shim".
`~/.gamescope-cmd.log` — which the script writes with the exact command it ran — settles
that in one line, and is the first thing to read on any such report.

Fixed by binding our wrapper over `/usr/bin/gamescope` inside the transient unit's mount
namespace (`BindReadOnlyPaths`). Deliberately a bind, not a replacement: punktfunk-gamescope
ships under its own name precisely so it sits BESIDE the distro package, and the bind is
scoped to the session — nothing outside it sees the redirect and nothing is written to
`/usr`. Skipped when the resolved binary already IS `/usr/bin/gamescope`.

**2. With the patched gamescope finally running, every Vulkan client died — black screen.**

The box's `VkLayer_FROG_gamescope_wsi` ships with the DISTRO's gamescope and speaks its
`gamescope_swapchain` protocol. Ours disagrees, so the compositor rejects the client's
`swapchain_feedback` ("message too short") and drops it. Steam never paints; there is no
other symptom, which is what makes it expensive to find.

Measured with `vkcube` under each build, layer on:

    ours 3.16.25-17  ON  -> 1 rejected client
    ours 3.16.25-17  OFF -> 0
    OLD pin 3.16.25-4 ON -> 1 rejected client
    stock 3.16.23.2  ON  -> 0

 The upstream protocol XML is BYTE-IDENTICAL between the distro's commit (5cdb5b0) and
our pin — same interface version, same `uuuuuus` signature — so this is the distro patching
gamescope, not a version bump. Hence the gate is "do the upstream triples differ", not a
floor, and an unreadable version on either side leaves the layer alone rather than degrading
a box that works (Bazzite/SteamOS, where it has always been fine).

⚠⚠ The old pin fails identically, so REVERTING the pin bump fixes nothing here — this is
pre-existing, not a regression from 5fb8dce4.

Verified against the UNPATCHED distro script, reproducing exactly what this code emits:
the session's own log reports `punktfunk-gamescope version 3.16.25-17-ga87390d+pfhdr4`,
with 0 swapchain_feedback errors, 0 client-communication errors and 0 aborts.

Gate: `scripts/xcheck.sh linux clippy` clean (0 warning/error lines), `cargo fmt` clean.
Non-vacuity re-verified per the xcheck note — a planted type error in the new function
produced 3 errors, and removing it went back to Finished.

Still open, deliberately NOT addressed here: a 10-bit HDR stream aborts gamescope in
`destroy_buffer` (upstream `pipewire.cpp:88`), which is a separate defect.
2026-08-09 20:15:52 +02:00
enricobuehler bc9201d136 fix(packaging/gamescope): bump the pin past upstream's capture-format probe, and sign the RPM
Three things, one delivery path — a Fedora/Nobara box getting the patched gamescope.

**The pin moves 8c676c39 -> 5fb8dce4** (3.16.25-1 -> 3.16.25-11). The commit that matters
is ff6b924, `rendervulkan: fall back to XBGR2101010 when XRGB2101010 is unsupported`: it
probes `linearTilingFeatures` for STORAGE+SAMPLED and captures as XBGR2101010 where
A2R10G10B10 linear storage is unavailable — which is every NVIDIA. That covers the paths
that are upstream's rather than ours: the RGB intermediate `paint_pipewire()` acquires when
the stream is YCbCr, and AVIF screenshots. #143 fixed our own node host-side; this is the
other half, and its commit message asked for exactly this bump.

All six patches rebased. Only 0006 conflicted: upstream's f8be7ee added
`vulkan_has_drm_modifiers_for_features()` immediately above the `g_device` declaration our
patch turns into a reference — both kept. 0003 and 0005 come out byte-identical; 0006 also
picks up the `--zero-commit --no-signature` form 0001-0005 already used.

**Patch 0001 now offers `xBGR_210LE` BEFORE `xRGB_210LE`**, mirroring the host-side
`HDR_FORMAT_ORDER` rationale on the producer end. A consumer takes the first pod it can use,
and we were handing third-party consumers (OBS and friends) the one format NVIDIA fills
byte-reversed under a correct-looking label. Deliberately NOT done by calling upstream's
`vulkan_get_rgb10_capture_format()`, which is what pw_pods.rs proposes: that symbol landed
after 3.16.25, so it would break `packaging/nix/gamescope.nix` — which applies these patches
to whatever gamescope nixpkgs pins — with an opaque C++ error instead of a patch conflict.
The reorder gets the same outcome on any base. Note added there so the next reader does not
"fix" it.

**And the RPM was never signed.** `Sign RPMs` runs right after `Build RPM`; the gamescope
RPM is built ~90 steps later, behind its own ~10-minute cache, so it missed the signing pass
entirely — every punktfunk-gamescope RPM ever published went out unsigned. The repo file we
tell users to install carries `gpgcheck=1`, so `dnf install punktfunk-gamescope` failed with
"The package is not signed" on every Fedora and Nobara box. The package was in the channel
the whole time and could not be installed from it, which is worse than absent: the notes and
the docs-site both say it is there. `sign-rpms.sh` now takes explicit paths (defaulting to
`dist/*.rpm` as before) and a second pass signs this one before publish, fail-closed on a tag
like the first.

Verified on Nobara 44 (VM 123, RTX 5070 Ti passthrough), canary 0.27.0-0.ci12611.g516a2954:

* Builds clean in the fc44 CI image; banner `3.16.25-17-ga87390d+pfhdr4` (11 upstream + our
  6), so the marker the host probes still reads 4 — no capability moved, hence pkgrel 3 and
  `.pfhdrN` staying put.
* `pw-cli enum-params` on the live node: BGRx, NV12, **xBGR_210LE (81), xRGB_210LE (80)** —
  8-bit consumers still negotiate bit-for-bit, 10-bit now leads with the safe one.
* All four patched flags present, `--pipewire-composite-external-overlay` included.
* Patch 0006 confirmed working by comparison, which is the only way to see it: the new build
  exits 0 where both the pre-0006 `+pfhdr2` build and the stock 3.16.23.2 abort with 134.
* Signing fix proven with a throwaway key: `Signature: (none)` -> `digests signatures OK`.
* Host health on the canary: synthetic spike 300/300 encoded, loopback 300 recovered, 0
  mismatches.

One unexplained one-off: the very first headless run after install segfaulted at exit
(SIGSEGV, after "Primary child shut down!"). Not reproduced in 11 subsequent runs across
every flag combination, so it is recorded rather than diagnosed — the binary is stripped and
there is no symbolised core.
2026-08-09 18:00:28 +02:00
enricobuehler 003ce8bea7 Merge pull request 'Every NVIDIA gamescope HDR stream had red and blue swapped — and a sysext step added in a release was unreachable forever' (#143) from worktree-hdr-rb-swap-nvidia into main
arch / build-publish (push) Failing after 3s
ci / rust (push) Failing after 2s
ci / rust-arm64 (push) Failing after 2s
deb / build-publish (push) Failing after 3s
deb / build-publish-host (push) Failing after 0s
deb / build-publish-client-arm64 (push) Failing after 1s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 12s
ci / bun-nix (push) Successful in 31s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 22s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 23s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 7s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 18s
apple / swift (push) Successful in 1m40s
ci / web (push) Successful in 1m8s
ci / docs-site (push) Successful in 1m16s
docker / builders-arm64cross (push) Successful in 16s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m10s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m28s
docker / deploy-docs (push) Failing after 1m41s
android / android (push) Successful in 5m38s
apple / screenshots (push) Successful in 5m53s
windows-host / package (push) Successful in 16m59s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 13s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m54s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 18m1s
Reviewed-on: #143
2026-08-09 15:34:02 +00:00
enricobuehler 235b8e55d4 Merge pull request 'chore(web): console onto @unom/ui 0.9.2' (#142) from worktree-console-unom-092 into main
audit / license-gate (push) Failing after 2s
audit / cargo-audit (push) Failing after 2s
audit / docs-site-audit (push) Successful in 22s
audit / bun-audit (web) (push) Failing after 22s
audit / bun-audit (sdk) (push) Successful in 27s
audit / bun-audit (plugin-kit) (push) Successful in 29s
audit / pnpm-audit (push) Successful in 20s
ci / rust-arm64 (push) Failing after 2s
deb / build-publish-client-arm64 (push) Failing after 2s
ci / rust (push) Failing after 2s
ci / bun-nix (push) Successful in 31s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 15s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 19s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
arch / build-publish (push) Canceled after 1m29s
ci / web (push) Canceled after 1m12s
ci / docs-site (push) Canceled after 1m9s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
deb / build-publish (push) Canceled after 1m14s
deb / build-publish-host (push) Canceled after 1m13s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 1s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 1s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 11s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 1m3s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 28s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 25s
docker / deploy-docs (push) Canceled after 0s
windows-host / package (push) Canceled after 1m1s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
nix / flake (push) Failing after 4m16s
Reviewed-on: #142
2026-08-09 15:32:35 +00:00
enricobuehler 0b252403cd fix(web): fix the card inset at the root, not at the call sites
ci / bun-nix (pull_request) Successful in 51s
ci / docs-site (pull_request) Successful in 1m35s
ci / web (pull_request) Successful in 2m30s
ci / rust-arm64 (pull_request) Successful in 3m16s
ci / rust (pull_request) Failing after 9m12s
nix / flake (pull_request) Failing after 19m50s
The broken inset on the Displays configuration card was the symptom. The cause is
structural, and it had already been diagnosed at least twice in-tree without being fixed.

Two faults, both in components/ui/card.tsx:

1. The padding was a RESPONSIVE COMPOUND: `p-4 pt-0 sm:p-6 sm:pt-0`. tailwind-merge
   resolves conflicts only within a variant, so any call-site override won at the base
   and lost at `sm:` — correct on a phone, wrong on every desktop. Measured on the
   Displays card before this change: padding-top 24px at 500px, 0px at 1440px.

2. `pt-0` encoded an assumption about a SIBLING that nothing enforced — "a CardHeader is
   above me and supplies the top inset". Delete the header, which is exactly what tabbing
   a page does since the tab label replaces the card title, and the top inset silently
   vanishes at ≥640px.

Fix:

- One single-variant utility, `p-padding-card` — the same `--spacing-padding-card` token
  @unom/ui's own Card uses, so nested cards finally agree on their inset. A single
  variant cannot half-lose an override.
- Top inset is now self-correcting: `[&:not(:first-child)]:pt-0`. Ask the DOM instead of
  the author. A headerless CardContent keeps its inset with nothing to remember.

Seven call sites had grown their own compensation in five dialects — `p-6`,
`p-card pt-card sm:pt-card` (×3), `p-4 sm:pt-6` (×3), `pt-4 sm:pt-6`, and my own `pt-6`
from the tabs commit. All removed; they are the symptom-fixes this replaces. LogsCard
even carried a six-line comment correctly describing the trap and working around it
locally — that comment is now three lines saying it no longer needs saying.

`flush` stays: full-bleed content is a real intent, expressed as a prop the component
honours rather than a utility that has to out-argue the one already there.

Guarded by UI/Card → "Inset with and without header", a headered/headerless pair that has
to look identical on every side. It must be checked at BOTH widths — a single width
cannot show this class of bug, which is why it kept surviving.

Verified by measuring computed padding at 500px and 1440px: first child 20px on all four
sides, after-a-header 0px top and 20px elsewhere, identical at both widths. tsc clean,
biome clean on every touched file, 9/9 server tests, build + i18n clean, 32/32 screenshots.
2026-08-09 17:23:09 +02:00
enricobuehler 0ab17ee81d fix(packaging): a post_merge step added in a release was unreachable forever
ci / bun-nix (pull_request) Successful in 48s
ci / docs-site (pull_request) Successful in 1m20s
ci / web (pull_request) Successful in 1m16s
apple / swift (pull_request) Successful in 1m45s
ci / rust-arm64 (pull_request) Successful in 1m43s
apple / screenshots (pull_request) Skipped
android / android (pull_request) Successful in 3m52s
ci / rust (pull_request) Failing after 8m7s
A sysext upgrade is driven by the script from the OLD image -- /usr/bin/punktfunk-sysext
is replaced by the very `systemd-sysext refresh` that runs mid-upgrade -- so a
post_merge step ADDED in the new release is executed by nobody. The old script
does not have it, and the new script never gets a turn: from then on `update`
matches the "already on $cur" branch and returns before post_merge. The step is
permanently unreachable on exactly the installs that need it, and nothing says so.

Field-proven on the Bazzite host that took 0.25.0 -> 0.26.0 (2026-08-09). The
casualty was the `punktfunk` group, which post_merge learned to create in 0.26.0
(62a6fa9f): 0.25.0's script ran the upgrade, so the group was never created, and
every `punktfunk-sysext update` since has said "nothing to do". `pf-dm-helper`
gates on membership in that group, so it refused every caller -- pkexec authorised
it and the helper then declined itself -- and every managed gamescope takeover fell
back to "stopping the display manager needs privilege", leaving sddm's autologin
Relogin loop churning logind sessions for the whole stream.

Re-run post_merge when already current. Everything in it is idempotent (guarded
getent/groupadd, `install` of /etc mirrors, udevadm reload/trigger, sysctl,
modprobe), so convergence is the honest behaviour and "nothing to do" was a lie
about host state. Add an explicit `reapply` verb too, so the steps a sysext image
cannot carry can be re-applied without reinstalling the image.

Also print the membership hint. Creating the group is necessary but NOT sufficient
and the difference is invisible until a stream fails: joining stays opt-in by
design (writing vhci `attach` materialises an arbitrary emulated USB device), so
post_merge now names the exact usermod when SUDO_USER is not a member. Matched with
`grep -qx` so `punktfunk-update` does not read as `punktfunk`.

bash -n clean; shellcheck clean apart from the pre-existing SC1091 on
`. /etc/os-release`, which fires on the unmodified file too.
2026-08-09 17:08:24 +02:00
enricobuehler 31aef4b09f feat(web): tab the Virtual displays page
ci / rust-arm64 (pull_request) Successful in 2m6s
ci / docs-site (pull_request) Successful in 3m59s
ci / bun-nix (pull_request) Successful in 4m39s
ci / web (pull_request) Successful in 5m1s
ci / rust (pull_request) Failing after 13m18s
nix / flake (pull_request) Failing after 23m11s
Same pill strip the plugin UIs use, via @unom/ui's Tabs: Configuration | Live displays.

The page was two stacked cards, and the configuration card ALONE is taller than the
viewport — the existing comment on the unsaved badge says as much, because that height
is how pending edits went unnoticed. The live-display list sat below all of it, so in
practice it was off screen.

Two details that are not cosmetic:

- The dirty marker moved from the card header onto the Configuration TRIGGER. Behind a
  tab the old badge would vanish entirely while Live was open — a strictly worse version
  of the problem it was added to solve. On the trigger it survives both tabs, and the
  Custom block keeps its own inline badge for when the tab IS open.
- The strip is extracted as a presentational `DisplayTabs` rather than inlined in
  `DisplaySection`. The container calls `useBlocker`, which needs a router, so it cannot
  render in Storybook — and this page's story exists specifically to pin the MOTION
  NESTING of the preset grid (a card sets no delayChildren, so tiles nested one level
  deeper stop staggering). Inserting tabs changes that ancestor chain, so the story has
  to render the real one or it passes for the wrong reason.

Adds Pages/Displays → "Unsaved on other tab", which switches to Live with a dirty draft:
if the marker ever goes silent there, the warning is gone exactly when it matters.

Verified: tsc clean, biome clean, `bun test server/` 9/9, vite build + i18n check clean,
Storybook builds, 32/32 screenshots.
2026-08-09 16:58:46 +02:00
enricobuehler 97928516a0 fix(pf-capture): every NVIDIA HDR stream had red and blue swapped
gamescope's capture textures are mappable, hence linear-tiled, and NVIDIA does
not implement linear-tiled STORAGE for A2R10G10B10_UNORM_PACK32. Upstream says
it plainly in rendervulkan.cpp: "imageStore lands in XBGR order there, swapping
R/B". So the composite writes XBGR bytes into a buffer still LABELLED
XRGB2101010, and our patch's spa_format_to_drm() derives that label from the
negotiated SPA format alone, never asking the hardware what it can actually
write.

The host then believed the label, correctly at every step:
xRGB_210LE -> PixelFormat::X2Rgb10 -> NV_ENC_BUFFER_FORMAT_ARGB10. DRM
XRGB2101010 really is "B in the low 10 bits" and NVENC ARGB10 really is "B in
the lowest 10 bits"; the Windows twin (R10G10B10A2 -> ABGR10) is correct by the
same rule. Every mapping audits clean because the label was right and only the
CONTENT was wrong -- which is why this survived a full trace of both ends.

Fix the preference host-side: offer xBGR_210LE FIRST. The first compatible
consumer pod wins, so that is what a gamescope session lands on, and an
XBGR2101010 texture is one NVIDIA writes in its own order -- label and content
agree. It costs nothing elsewhere: A2B10G10R10_UNORM_PACK32 is the universally
supported packed-10 format, it is what upstream's own fallback picks, and
X2Bgr10 has a first-class encoder path (NVENC ABGR10, VAAPI X2BGR10LE).
xRGB_210LE stays as the second pod so a producer offering only it can still
negotiate HDR instead of dropping to the SDR downgrade.

Doing it here rather than in the patch set is deliberate: the real fix is for
spa_format_to_drm() to offer only what vulkan_get_rgb10_capture_format()
reports, but that function landed after 3.16.25 and the pin is
3.16.25-7-g60561e2+pfhdr4 (0 "2101010" strings in the shipped binary), so the
deployed gamescope cannot self-correct. This ships in the host binary with no
gamescope rebuild.

Field-confirmed on the RTX 5070 Ti Bazzite host with 0.26.0, and confirmed
host-side rather than client-side by reproducing the identical swap from two
unrelated clients (16" MacBook Pro and Mac Studio). SDR was never affected --
it takes no packed-10 path.

Gate (pf-lxcheck2, linux/amd64): fmt clean, clippy --all-targets -D warnings
clean, cargo test -p pf-capture 60 passed / 0 failed incl. the new
hdr_offers_xbgr_before_xrgb order pin.
2026-08-09 16:58:15 +02:00
enricobuehler d13d253c2f chore(web): @unom/ui 0.8.16 → 0.9.2
ci / rust-arm64 (pull_request) Failing after 31s
ci / docs-site (pull_request) Successful in 3m1s
ci / bun-nix (pull_request) Successful in 3m27s
ci / web (pull_request) Successful in 4m29s
ci / rust (pull_request) Failing after 13m16s
nix / flake (pull_request) Failing after 19m47s
Brings the console onto the current design system. 0.9.x adds the Badge, Spinner,
Skeleton, Switch, Table, EmptyState and CodeBlock primitives, and 0.9.2 carries the
form fixes found while overhauling the rom-manager plugin UI:

- Select's border and focus ring resolved to `--main`, which is the FOREGROUND here
  (`--main: var(--foreground)` in web/src/styles.css), so the trigger wore a near-white
  border and a 3px near-white focus ring. Its chevron and placeholder were painted
  `--secondary`, a SURFACE colour, and all but vanished. Now on `--input`/`--ring`, the
  same tokens InputText already used.
- InputNumber declares a color-scheme, so the browser-drawn spinner arrows stop being
  near-black on a near-black field.

Both defects were live in this console too — the console palette is what exposes them.

Verified: codegen + vite build clean, `tsc --noEmit` clean, `bun test server/` 9/9,
Storybook builds, 31/31 screenshots. A probe over all 61 stories reports ZERO page
errors, and the two stories containing a Select now render it at h-input-height with
`border: rgb(42, 33, 72)` (the input token) and a muted-foreground chevron.

Note: the console's components/ui/ wrapper layer is unchanged and still required —
@unom/ui's DialogContent remains a surface with no Portal or placement, which is
exactly what web/src/components/ui/dialog.tsx supplies.
2026-08-09 16:27:08 +02:00
enricobuehler 516a295432 Merge pull request 'My gamescope gate withheld the host .deb it was meant to protect — the release still ships the KDE-breaking one' (#140) from worktree-gamescope-gate-placement into main
ci / web (push) Successful in 1m5s
ci / bun-nix (push) Successful in 17s
ci / rust-arm64 (push) Successful in 1m38s
ci / docs-site (push) Successful in 2m36s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 7s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 8s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 7s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 5s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 4s
deb / build-publish-client-arm64 (push) Successful in 1m38s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 16s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 14s
deb / build-publish (push) Successful in 3m45s
docker / builders-arm64cross (push) Successful in 8s
docker / deploy-docs (push) Successful in 31s
deb / build-publish-host (push) Successful in 6m39s
ci / rust (push) Successful in 10m37s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m48s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m37s
Reviewed-on: #140
2026-08-09 09:12:55 +00:00
enricobuehler 2c190b27b4 Merge pull request 'Switching audio device mid-stream killed the sound for the rest of the session — AVAudioEngine stops itself, and nothing ever restarted it' (#141) from worktree-audio-device-switch-silence into main
ci / rust (push) Canceled after 0s
ci / rust-arm64 (push) Canceled after 0s
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
apple / swift (push) Successful in 1m41s
release / apple (push) Successful in 10m9s
apple / screenshots (push) Successful in 6m10s
Reviewed-on: #141
2026-08-09 09:10:52 +00:00
enricobuehler 3cfa5ca194 Merge pull request 'The capability-hint test asserted the environment, not the code — main is red on a machine where nothing is wrong' (#139) from worktree-kwin-capability-test-env into main
apple / swift (push) Canceled after 0s
apple / screenshots (push) Canceled after 0s
ci / rust (push) Canceled after 0s
ci / rust-arm64 (push) Canceled after 0s
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
deb / build-publish (push) Canceled after 1m55s
deb / build-publish-host (push) Canceled after 1m10s
deb / build-publish-client-arm64 (push) Canceled after 50s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 2s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 0s
android / android (push) Successful in 6m20s
windows-host / package (push) Successful in 11m23s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 18s
arch / build-publish (push) Successful in 11m50s
Reviewed-on: #139
2026-08-09 09:10:33 +00:00
enricobuehler bf913c5706 fix(apple): switching audio device mid-stream killed the sound for the rest of the session
ci / bun-nix (pull_request) Successful in 36s
ci / web (pull_request) Successful in 1m22s
apple / swift (pull_request) Successful in 1m39s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m40s
ci / rust-arm64 (pull_request) Successful in 2m53s
ci / rust (pull_request) Failing after 9m43s
Field report, macOS client, host-independent: start a stream with AirPods in, take them
out — nothing on the speakers; put them back in — nothing in the AirPods either. Only
restarting the whole stream brought audio back.

An AVAudioEngine does not follow the audio hardware. When the output device changes under
a running engine, its IO unit sees the new hardware, THE ENGINE STOPS ITSELF, and it posts
AVAudioEngineConfigurationChange. It stays stopped until somebody starts it again, and
nothing here ever did — no error, no log line, just a session rendering silence from that
moment on. Putting the AirPods back in is a second stop, not a recovery, which is exactly
why that half of the report looked so strange.

Measured on the client's own playback topology (source node -> main mixer, 48 kHz stereo)
by moving the default output device programmatically: render callbacks go from ~94/s to
zero the instant the device changes, and both restarting the same engine and building a
fresh one resume them.

The fix watches the hardware and rebuilds the topology the session was started with, on
whatever device is there now. Three triggers, because no single one covers the ground:

  - the engine's own configuration-change notification, every platform — the direct
    signal, but it can only be posted BY an engine, so it cannot report a rebuild that
    failed to start;
  - a CoreAudio HAL default-output-device listener on macOS — independent of any engine
    and of the engine's topology. This is what makes the recovery work for the
    voice-processing engine, which is the DEFAULT macOS configuration (mic and echo
    cancellation both default on) and whose notification behaviour could not be verified:
    no Mac in the fleet can initialize VPIO at all;
  - route-change and media-services-reset on iOS/tvOS, where the session rather than the
    device is what moves. The route observer is now installed for mic-off (.playback)
    sessions and on tvOS too — it used to be iOS-and-mic-only, for the earpiece steer,
    but every platform has engines a route change can stop.

They collapse into one debounced rebuild (one switch produces a burst), with a floor
between rebuilds so a device that renegotiates in a loop cannot spin the session, and a
short retry ladder for a device caught mid-transition — a rebuild that fails leaves no
engine to post the next notification, so that path must not simply give up. The ring is
deliberately carried across: the drain thread keeps decoding through the switch, and the
ring's overflow policy has already dropped whatever went stale while the engine was down.

A rebuild is only ever done when it concerns us. A healthy engine that followed the change
on its own is left alone, and somebody changing the system default while this session is
pinned to a named speaker is none of our business — rebuilding for that would cost an
audible gap for nothing.

The trigger wiring is split into AudioDeviceWatcher for one reason: an end-to-end test of
the recovery needs a live session, which needs a host, and punktfunk-host does not build
on macOS — so the part where a silent failure costs the session ALL of its audio would
otherwise ship unverified. On its own the watcher is pointed at the real hardware from a
unit test: a real default-output-device move must reach the owner, our engine's
notification must get through, a foreign engine's must not. Neutralizing the wiring fails
both positive tests and neither negative one.

AudioDeviceSwitchTests drives the real SessionAudio through the out-and-back switch
against the loopback host; it skips wherever that fixture cannot run (which is every Mac,
today) and the open host's frame budget is raised so it outlives the switch.
2026-08-09 11:03:10 +02:00
enricobuehler 5bd92dac5d fix(ci): my gamescope gate withheld the host .deb it was supposed to protect
ci / bun-nix (pull_request) Successful in 25s
apple / swift (pull_request) Successful in 1m33s
apple / screenshots (pull_request) Skipped
ci / web (pull_request) Successful in 1m40s
ci / rust-arm64 (pull_request) Successful in 1m40s
ci / docs-site (pull_request) Successful in 2m3s
android / android (pull_request) Successful in 4m10s
ci / rust (pull_request) Successful in 8m9s
The gate #135 added fails the job at the gamescope BUILD step. In deb.yml that
step runs before "Publish to the Gitea apt registry" and "Attach the host .deb
to the Gitea release", so failing it skipped both.

Consequence on the v0.26.0 tag, and it is the worst thing in this release so
far: the host .deb on the release is from 00:17 — re-point #1, BEFORE #136
revoked CAP_SYS_NICE. Every other .deb is from 08:29-08:31. So the published
Debian host still runs `setcap cap_sys_nice=ep` in its postinst, which is
exactly what makes the host unidentifiable to KWin and kills every KDE desktop
session. A gate meant to protect the release withheld the fix for it and left
the broken artifact in place.

rpm.yml has the identical latent bug and only escaped it because Fedora went
green: a gamescope failure there would skip the sysext image, the feed publish
and the release attach, withholding the punktfunk RPMs and .raw images too.

Both now warn at the build/package steps and gate as the LAST step of the job,
after everything has published. A missing EXTRA must never stop a good artifact
shipping — go red afterwards instead.

Also: name noble's dependencies outright. `apt-get build-dep gamescope` gives it
almost nothing (the distro has no comparable package), which is why this peeled
one dep per CI cycle — wayland-protocols, then xdamage. The full set is derived
from the Arch package's depends+makedepends, which is the build that demonstrably
works, plus wlroots' own (it is a forced fallback subproject).

One `apt-get` per name on purpose: a single transaction aborts wholesale on one
unknown package, installing NOTHING and hiding the real gap behind a name typo.
Per-package, best-effort, with the missing name echoed; the end-of-job gate is
what actually decides.

⚠ Verification: both YAML files parse; every gamescope-touching `run:` block is
`bash -n` clean with matrix placeholders substituted (9 blocks); the .deb glob
matches build-gamescope-deb.sh's documented output
(`dist/punktfunk-gamescope_<version>_<arch>.deb`) and the RPM glob excludes
debuginfo/debugsource exactly as the attach loop above it does. The noble dep
NAMES cannot be proven from macOS — that is what the next tag run decides, and
it now decides it without holding the host .deb hostage.
2026-08-09 10:52:42 +02:00
enricobuehler e8a4f54c07 fix(pf-vdisplay): the capability-hint test asserted the environment, not the code
apple / swift (pull_request) Successful in 1m33s
apple / screenshots (pull_request) Skipped
android / android (pull_request) Successful in 5m0s
ci / bun-nix (pull_request) Successful in 42s
ci / docs-site (pull_request) Successful in 1m21s
ci / web (pull_request) Successful in 1m48s
ci / rust-arm64 (pull_request) Successful in 2m47s
ci / rust (pull_request) Successful in 6m57s
`silent_without_capabilities` called the real `capability_denial_hint()` and
asserted it returns "", on the strength of a doc comment that read "The test
process has no capabilities."

That is true on a dev box and false in CI, where the runner container is root
with a full permitted set. main went red on 0f79587d with:

    left: " — NOTE: this process carries capabilities (CapPrm=0x000001ffffffffff) …"
   right: ""

Nothing was wrong: the hint fired correctly, on a process that really did hold
every capability. The test was reading the ambient environment and calling it a
property of the code.

`permitted_caps_from_status` had already been split out for exactly this reason
— "so that shape is testable without a capability-carrying process to point at"
— but only the PARSE half. The message half still went to /proc/self/status.
This finishes the split: `capability_denial_hint_for(Option<u64>)` holds the
formatting and takes the mask, `capability_denial_hint()` reads /proc and
delegates. Both keep their callers, so neither is dead code.

Also adds `names_the_mask_and_the_repair_when_capped`. Without it the silent
case passes just as well against a function that returns "" unconditionally —
which is the failure mode this repo has been bitten by before, and the reason
every decode fix carries a counterfactual.

No behaviour change: the three error paths call the same function and get the
same string.

⚠ Verification is CI. `kwin.rs` is `#[cfg(target_os = "linux")]`, so it does not
compile on the macOS host this was written from; `cargo fmt --all --check` is
clean and a Linux container check was attempted but the stock rust image has no
cmake for audiopus_sys, so it never reached the test. ci.yml going green on main
is the proof — and unlike the case it replaces, this test now fails or passes
for reasons that have nothing to do with the machine running it.

Does not touch the v0.26.0 tag: ci.yml runs on `push: branches: [main]` and
`pull_request` only, and no tag leg runs cargo test.
2026-08-09 10:39:41 +02:00
enricobuehler f80636f901 Merge pull request 'The release notes advertise a privilege 0.26.0 deliberately does not grant' (#138) from worktree-notes-capsysnice-correction into main
android-screenshots / screenshots (push) Successful in 1m29s
release / apple (push) Successful in 12m13s
decky / build-publish (push) Successful in 37s
windows-host / package (push) Successful in 11m48s
windows-host / canary-manifest (push) Skipped
deb / build-publish-client-arm64 (push) Successful in 1m27s
deb / build-publish (push) Successful in 4m18s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m41s
linux-client-screenshots / screenshots (push) Successful in 2m54s
sbom / sbom (push) Successful in 20s
deb / build-publish-host (push) Failing after 5m18s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m54s
windows-host / winget-source (push) Successful in 21s
docker / builders-arm64cross (push) Successful in 9s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 6s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 18s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 21s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 13s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 18s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 38s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m5s
docker / deploy-docs (push) Successful in 28s
android / android (push) Successful in 10m26s
arch / build-publish (push) Successful in 11m24s
web-screenshots / screenshots (push) Successful in 5m5s
flatpak / build-publish (push) Successful in 16m37s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m18s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m6s
ci / rust-arm64 (push) Successful in 1m40s
ci / web (push) Successful in 2m5s
ci / docs-site (push) Successful in 1m12s
ci / bun-nix (push) Successful in 38s
ci / rust (push) Canceled after 1m31s
2026-08-09 08:13:45 +00:00
enricobuehler 0f79587dd6 docs(release): the notes claimed a privilege 0.26.0 deliberately does not grant
ci / rust-arm64 (pull_request) Successful in 1m55s
ci / web (pull_request) Successful in 1m51s
ci / bun-nix (pull_request) Successful in 35s
ci / docs-site (pull_request) Successful in 1m18s
ci / rust (pull_request) Failing after 8m47s
The user-facing v0.26.0 notes said, of the PyroWave GPU-priority lever:

    "it is now, and the package grants the host the permission that switch needs"

That was true of 0.26.0-1 and is now the opposite of true. Granting CAP_SYS_NICE
made the host unidentifiable to KWin and killed desktop streaming on every KDE
box across all five Linux channels, so 0.26.0-2 revokes it everywhere and must
keep doing so. The lever is wired natively on Linux for the first time — that
part stands — but it is dormant on an ordinary install, and the notes have to
say so rather than advertise a speed-up nobody gets.

CHANGELOG.md was already corrected in #136 (the 0.26.0-2 note under PW1 and the
qualifier on the owed A/B). This is the user-facing half, which #136 did not
touch:

  * the PyroWave bullet now leads with what DID land (two encoder handles, the
    capture buffer headroom) and describes the priority switch as present but
    dormant, with the reason.
  * a new Fixed entry for the KDE breakage itself. Worth telling users even
    though the release was never announced: 0.26.0-1 packages did reach the
    registries, and anyone who pulled one has a desktop session that fails with
    a missing-screencast error surviving a clean reinstall. It also explains the
    dormancy the bullet above now refers to.

Deliberately NOT written as a "Before you update" action: upgrading strips the
capability by itself on every channel, so there is nothing for a reader to do.

Commit count 47 -> 52.

Voice check clean (0 internal-vocabulary hits above "## For developers"); notes
67 lines.
2026-08-09 10:12:58 +02:00
enricobuehler 651a7a82a1 Merge pull request '0.26.0-1 gave the host CAP_SYS_NICE, which made it invisible to KWin — every KDE desktop session died, on five packaging channels' (#136) from worktree-kwin-capability-identification into main
ci / web (push) Successful in 1m4s
apple / swift (push) Successful in 1m37s
ci / rust-arm64 (push) Failing after 2m20s
ci / rust (push) Failing after 2m21s
ci / docs-site (push) Successful in 1m18s
ci / bun-nix (push) Successful in 26s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 37s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 30s
deb / build-publish-client-arm64 (push) Successful in 1m56s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 12s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 33s
android / android (push) Successful in 6m21s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m21s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 11s
apple / screenshots (push) Successful in 5m58s
deb / build-publish (push) Successful in 5m32s
deb / build-publish-host (push) Successful in 6m11s
docker / builders-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
arch / build-publish (push) Successful in 11m40s
windows-host / package (push) Successful in 12m30s
windows-host / winget-source (push) Skipped
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m2s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m58s
nix / flake (push) Successful in 18m34s
windows-host / canary-manifest (push) Successful in 14s
Reviewed-on: #136
2026-08-09 08:04:43 +00:00
enricobuehler 4d383811c0 fix(packaging): the same CAP_SYS_NICE broke KDE on FIVE channels, not one — Bazzite included
ci / bun-nix (pull_request) Successful in 17s
ci / web (pull_request) Successful in 1m7s
apple / swift (pull_request) Successful in 1m38s
ci / rust-arm64 (pull_request) Successful in 1m38s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m46s
android / android (pull_request) Successful in 5m31s
ci / rust (pull_request) Failing after 9m2s
nix / flake (pull_request) Successful in 12m24s
The Arch fix in the previous commit was incomplete. 0.26.0-1 granted the host CAP_SYS_NICE through
every Linux channel we ship, and each one breaks KWin identification the same way:

  * packaging/rpm/punktfunk.spec .......... %caps(cap_sys_nice=ep) in %files  <- Fedora AND Bazzite
                                            via rpm-ostree layering
  * packaging/bazzite/build-sysext.sh ..... setcap on the staging tree, recorded by mksquashfs
  * packaging/debian/build-deb.sh ......... setcap in the postinst
  * packaging/nix/nixos-module.nix ........ security.wrappers with capabilities = "cap_sys_nice=ep"
  * scripts/steamdeck/install.sh .......... setcap on $BIN, six lines after writing the .desktop
                                            whose Exec= it thereby voids

Bazzite was NOT a separate fault, as first reported here — it is this one. Verified by mounting the
published punktfunk-0.26.0-1-x86-64.raw: `getcap usr/bin/punktfunk-host` reports cap_sys_nice=ep,
stored as security.capability in the squashfs. The claim in packaging/arch/build-sysext.sh that
"file capabilities don't survive this squashfs path" is false and is corrected here; mksquashfs
records them, which is exactly why the image shipped one.

NixOS deserves its own note: a security.wrappers entry does not dodge the problem. The wrapper
raises the capability into its AMBIENT set before exec'ing the store binary, precisely so it
survives — which lands CAP_SYS_NICE in the exec'd process's permitted set and fails the readlink
identically to a file capability. ExecStart now points at the store path directly, which is also the
path packages.nix substitutes into the .desktop's Exec=, so the two finally agree.

Measured blast radius of holding a capability, same-uid reader, CachyOS kernel 7.1.6:

    /proc/PID/exe ....... EPERM   <- KWin's identification. Desktop sessions die.
    /proc/PID/root/* .... EPERM   <- xdg-desktop-portal reads .flatpak-info here to resolve an
                                     app id; the wlroots and Hyprland backends go through it
    /proc/PID/environ ... EPERM
    /proc/PID/cgroup .... OK
    /proc/PID/status .... OK
    /proc/PID/cmdline ... OK

Compositor backends, by exposure: KWin is broken outright (proven, field-confirmed). gamescope has
no identity gate and was never affected, which matches the field — only Desktop mode was reported.
Mutter drives Mutter's own D-Bus API, not the portal, and looks unaffected. wlroots and Hyprland go
through the ScreenCast portal, whose app-id resolution reads a path the capability blocks — a real
exposure, not something I reproduced end to end.

The sysext build now HARD-FAILS if a capability is staged, rather than trusting that the RPM payload
never carries one: a merged sysext's /usr is read-only squashfs, so a bad image cannot be repaired
on the box, and the spec was one %caps() away from baking one in again.

Docs corrected, because they advertised the capability as a feature:
  * docs-site running-as-a-service "GPU scheduling priority" — rewritten: the host carries no
    capability, why it must not, and how to clear a 0.26.0-1 install (Bazzite needs a new image)
  * docs-site configuration.md — the PYROWAVE_QUEUE_PRIORITY row no longer claims the packages grant it
  * packaging/bazzite/README.md — §6.5 still described the kde-desktop-setup.sh behaviour from
    before it stopped writing KWIN_WAYLAND_NO_PERMISSION_CHECKS and started REMOVING it; plus a
    note that 0.26.0-1 Desktop mode cannot be repaired in place
  * packaging/arch/README.md — the false "capabilities don't survive the sysext" line
  * CHANGELOG v0.26.0 PW1 — annotated with the 0.26.0-2 correction rather than rewritten, and the
    owed PyroWave-under-load A/B now says it needs a gamescope-only box

Verified: bash -n on all five changed shell files; nix-instantiate --parse on nixos-module.nix and
packages.nix; the published 0.26.0-1 sysext mounted and its capability read; getcap on an uncapped
file exits 0 with empty output, so the new build assertion cannot false-positive.
2026-08-09 09:56:36 +02:00
enricobuehler 42ee6c5628 fix(packaging): the host's CAP_SYS_NICE made it invisible to KWin, killing every KDE session
0.26.0-1 setcap'd `cap_sys_nice=ep` on /usr/bin/punktfunk-host so the encoder could open an
elevated global-priority Vulkan queue. On every KDE box that ended desktop streaming outright:

    KWin virtual output failed: KWin does not expose zkde_screencast_unstable_v1 to this client

reported from CachyOS on NVIDIA and on AMD, surviving a clean reinstall of host and client, and
worked around only by KWIN_WAYLAND_NO_PERMISSION_CHECKS=1.

The two cannot coexist. KWin hands out its restricted protocols — zkde_screencast_unstable_v1,
which mints our virtual output, and org_kde_kwin_fake_input, which injects input — only to a client
it can IDENTIFY, by resolving that client's /proc/<pid>/exe and matching it against an installed
.desktop's Exec=. The kernel refuses that readlink to any reader whose effective set is not a
superset of the target's PERMITTED set (cap_ptrace_access_check), and KWin holds no capabilities.
So the instant the binary carries one, KWin's executablePath() is empty, nothing matches, and the
global is never advertised — presenting exactly as a missing or mis-installed .desktop file.

Measured on CachyOS (kernel 7.1.6), same-uid reader, cap_sys_nice=ep on the target:

    no capability .............................. readlink /proc/<pid>/exe OK
    capability ................................. EPERM
    capability + prctl(PR_SET_DUMPABLE, 1) ..... EPERM   <- dumpable is NOT the gate
    capability dropped + PR_SET_DUMPABLE(1) .... OK      <- only an uncapped process works

The third row also rules out the reflex fix of moving the grant to systemd AmbientCapabilities=,
which lands CAP_SYS_NICE in the very same permitted set. Nothing short of not holding the
capability restores identification, so the host does not get one.

The cost is pacing only. pf-zerocopy's device create already walks REALTIME -> HIGH -> default when
a priority class is refused, and pf-frame's thread nice is a documented best-effort no-op without
the capability — so this is 0.25.0's behaviour exactly, which is the behaviour that worked.

  * packaging/arch/punktfunk-host.install: grant -> revoke. post_upgrade strips the capability from
    boxes that already ran 0.26.0-1's scriptlet. A pacman upgrade writes a new inode and file
    capabilities do not survive that, so this is belt-and-braces for reinstall/downgrade paths.
  * pf-vdisplay kwin.rs: all three "KWin does not expose zkde_screencast" errors now read
    /proc/self/status and, if this process holds ANY capability, name it with its CapPrm mask and
    the `setcap -r` that repairs it. The failure stays impossible to diagnose from the Wayland side
    otherwise, and it is not unique to our own packaging — a hand-rolled setcap does it too.

Verified on 192.168.1.21 (CachyOS): the capability/dumpable matrix above; cargo check and
cargo clippy --all-targets -- -D warnings clean for pf-vdisplay; both new unit tests pass; and the
hint itself exercised end-to-end, silent uncapped and firing with CapPrm=0x0000000000800000 under
cap_sys_nice=ep. The shipped punktfunk-host-0.26.0-1-x86_64.pkg.tar.zst was unpacked to confirm its
.INSTALL carries the setcap on both post_install and post_upgrade.

Ships as 0.26.0-2 — packaging plus one crate, no version bump.
2026-08-09 09:37:30 +02:00
enricobuehler 08eaf337e8 Merge pull request 'v0.26.0 promised a Fedora and an apt gamescope that were never built' (#135) from worktree-gamescope-rpm-deb-builddeps into main
ci / web (push) Successful in 1m6s
ci / rust-arm64 (push) Successful in 1m35s
ci / bun-nix (push) Successful in 27s
ci / docs-site (push) Successful in 1m12s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 16s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 12s
deb / build-publish-client-arm64 (push) Successful in 1m30s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 19s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 16s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m12s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m14s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 9s
deb / build-publish-host (push) Successful in 5m53s
deb / build-publish (push) Successful in 6m9s
docker / builders-arm64cross (push) Successful in 7s
ci / rust (push) Successful in 11m22s
docker / deploy-docs (push) Failing after 6m12s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 18m32s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 18m20s
2026-08-09 07:22:59 +00:00
enricobuehler 39869031be fix(ci): the gamescope RPM and .deb never built, and a warning let the tag ship anyway
ci / bun-nix (pull_request) Successful in 23s
ci / docs-site (pull_request) Successful in 1m18s
ci / web (pull_request) Successful in 1m32s
ci / rust-arm64 (pull_request) Successful in 3m22s
ci / rust (pull_request) Successful in 11m20s
v0.26.0's notes and docs-site say the patched gamescope is now installable on
Fedora and on Debian/Ubuntu. Neither package exists on the release. Both builds
failed inside best-effort steps that emit `::warning::` and return 0, so every
job stayed green and the only evidence was a warning nobody reads. Arch built
fine, which is why it is the sole gamescope package attached.

Two distinct missing build deps, same root cause: `dnf builddep gamescope` /
`apt-get build-dep gamescope` resolve the DISTRO'S OLDER PACKAGED gamescope,
which does not need what the pinned master tree needs.

  Fedora (f43 AND f44)
    /usr/sbin/ld: cannot find -lstdc++
    have you installed the static version of the stdc++ library ?
    ERROR: Compiler sccache c++ cannot compile programs.

  build-punktfunk-gamescope.sh appends `-static-libstdc++ -static-libgcc` to
  LDFLAGS deliberately, so the binary still starts on SteamOS's older libstdc++.
  Without libstdc++-static that trips meson's very FIRST sanity check, so
  nothing builds at all.

  Debian/Ubuntu noble
    protocol/meson.build:7:17: ERROR: Neither a subproject directory nor a
    wayland-protocols.wrap file was found.

  The tree carries no wrap fallback for wayland-protocols.

Both proven deps are installed WITHOUT `|| true` so a rename is loud. The
remaining Arch makedepends the older packaged gamescope may not pull (glm,
cmake, libXcursor, wayland-protocols-devel on Fedora) stay best-effort, since
meson finds fallbacks and a name that moves between releases should not fail
the job.

And the part that actually matters: on `refs/tags/v*` a missing gamescope is
now an ERROR, not a warning. A release must not be able to make a claim its own
CI silently dropped. Gated in two places per platform — the build step, and the
packaging step that is authoritative and also covers the cache path (the build
step is skipped entirely on a cache hit, so a stale cache would otherwise reach
packaging and skip in silence). Canary keeps the old best-effort behaviour.

Deliberately NOT gated: the sysext leg. The notes make no claim about gamescope
inside the sysext, and with the build fixed gs-cache is populated so it gets the
binary anyway — gating it would add release-blocking risk with no matching
promise.

⚠ Verification is CI itself: both YAML files parse, and every gamescope-touching
`run:` block is `bash -n` clean with the matrix placeholders substituted. The
dep names cannot be proven from macOS; the rpm and deb legs on the next tag are
the proof, and they are now hard-gated, so a wrong name fails loudly instead of
shipping another empty promise.
2026-08-09 09:22:20 +02:00
enricobuehler 55f361cb92 Merge pull request 'The v0.26.0 tag went red on Windows — a Linux-only reader tripped dead_code' (#134) from worktree-pyrowave-wire-dead-code into main
apple / swift (push) Successful in 1m33s
ci / rust-arm64 (push) Successful in 4m54s
ci / web (push) Successful in 1m47s
release / apple (push) Successful in 10m42s
ci / rust (push) Successful in 8m51s
ci / docs-site (push) Successful in 1m44s
ci / bun-nix (push) Successful in 32s
apple / screenshots (push) Successful in 5m55s
android-screenshots / screenshots (push) Successful in 2m16s
deb / build-publish (push) Successful in 3m56s
decky / build-publish (push) Successful in 23s
deb / build-publish-client-arm64 (push) Successful in 2m39s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m32s
deb / build-publish-host (push) Successful in 6m49s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m0s
linux-client-screenshots / screenshots (push) Successful in 2m55s
android / android (push) Successful in 11m36s
arch / build-publish (push) Successful in 13m24s
flatpak / build-publish (push) Successful in 8m4s
windows-host / winget-source (push) Successful in 35s
windows-host / package (push) Successful in 11m45s
windows-host / canary-manifest (push) Skipped
sbom / sbom (push) Successful in 35s
docker / deploy-docs (push) Failing after 6m10s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 11s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 9s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 11s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 1m5s
docker / builders-arm64cross (push) Successful in 16s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 57s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 3m57s
web-screenshots / screenshots (push) Successful in 4m48s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 18m13s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 19m38s
2026-08-08 23:44:44 +00:00
enricobuehler 2079411f4f fix(pf-encode): the Windows host could not compile — a Linux-only reader tripped dead_code
apple / swift (pull_request) Successful in 1m35s
apple / screenshots (pull_request) Skipped
ci / web (pull_request) Successful in 1m24s
android / android (pull_request) Successful in 5m55s
ci / rust-arm64 (pull_request) Successful in 4m5s
ci / bun-nix (pull_request) Successful in 33s
ci / docs-site (pull_request) Successful in 1m32s
ci / rust (pull_request) Successful in 21m44s
The v0.26.0 tag went red on windows-host at the clippy step, after a clean
build:

  error: function `wire_sequence` is never used
    --> crates\pf-encode\src\enc\pyrowave_wire.rs:68:15
     = note: `-D dead-code` implied by `-D warnings`

`pyrowave_wire` is cfg'd for linux OR windows and is genuinely shared —
`packet_boundary` and `stamp_color_bits` each have callers on both backends.
`wire_sequence` does not: every call site is in `enc/linux/pyrowave.rs`, which
is `#[cfg(all(target_os = "linux", feature = "pyrowave"))]`. Alternating
encoder handles are a Linux-side concern (PW5); the Windows backend drives
pyrowave's compat device with a single handle and never needs the counter. The
module's own `#[cfg(test)]` block does not reference it either, so on Windows
the item has zero callers in every target and dead_code is correct — it is the
`-D warnings` promotion to a hard error that stops the lib compiling.

Scoped to the one item rather than the file, and expressed as
`cfg_attr(not(target_os = "linux"), ...)` rather than a bare `allow`, so
dead_code stays LIVE on Linux — where the caller lives, and where this function
quietly losing its last caller would be a real finding rather than noise.

⚠ Not reproducible off a Windows box: cross-compiling to
x86_64-pc-windows-msvc from macOS dies in openh264-sys2's build script
(clang++ rejects `-fPIC` for that target) long before the lint stage. The
mechanism is nonetheless exact — one item, one cfg, zero callers behind it —
and the windows-host and windows-msix legs are the proof.

No behaviour change on any platform: this adds a lint attribute and eight
lines of comment.
2026-08-09 01:43:54 +02:00
enricobuehler 4d1a1348c0 Merge pull request 'chore(release): bump workspace version to 0.26.0' (#133) from worktree-release-0260 into main
apple / swift (push) Successful in 1m41s
audit / bun-audit (plugin-kit) (push) Successful in 1m1s
audit / bun-audit (sdk) (push) Successful in 33s
audit / bun-audit (web) (push) Failing after 36s
audit / docs-site-audit (push) Successful in 27s
audit / pnpm-audit (push) Successful in 30s
ci / web (push) Successful in 2m25s
audit / license-gate (push) Successful in 4m21s
ci / bun-nix (push) Successful in 24s
ci / docs-site (push) Successful in 1m26s
apple / screenshots (push) Canceled after 0s
ci / rust (push) Canceled after 0s
ci / rust-arm64 (push) Canceled after 6m20s
android-screenshots / screenshots (push) Canceled after 0s
android / android (push) Canceled after 0s
arch / build-publish (push) Canceled after 0s
deb / build-publish (push) Canceled after 0s
deb / build-publish-host (push) Canceled after 24s
deb / build-publish-client-arm64 (push) Canceled after 17s
decky / build-publish (push) Canceled after 5s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
linux-client-screenshots / screenshots (push) Canceled after 0s
release / apple (push) Canceled after 1m6s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 0s
sbom / sbom (push) Canceled after 0s
web-screenshots / screenshots (push) Canceled after 1s
windows-host / package (push) Canceled after 0s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m21s
audit / cargo-audit (push) Successful in 2m17s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m48s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m6s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m13s
nix / flake (push) Successful in 20m20s
flatpak / build-publish (push) Successful in 21m23s
2026-08-08 23:29:38 +00:00
enricobuehler e5180a5b7d chore(release): bump workspace version to 0.26.0
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m13s
apple / swift (pull_request) Successful in 1m35s
apple / screenshots (pull_request) Skipped
ci / web (pull_request) Successful in 2m39s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m15s
ci / bun-nix (pull_request) Successful in 56s
ci / docs-site (pull_request) Successful in 3m39s
ci / rust-arm64 (pull_request) Successful in 7m12s
android / android (pull_request) Successful in 9m5s
ci / rust (pull_request) Successful in 20m42s
nix / flake (pull_request) Failing after 20m28s
47 commits since v0.25.0, most of them from field reports on 0.25.0 itself,
plus Wave 2 of the PyroWave Linux host-performance program.

Nothing breaks: the wire protocol stays at 2 and the C ABI stays at 17, so
this release adds no call, no message and no capability bit. pf-driver-proto
is byte-for-byte identical to v0.25.0 and to v0.24.0.

Four new environment variables (PUNKTFUNK_OVERLAY_MASK,
PUNKTFUNK_GAMESCOPE_REFRESH_RATES, PUNKTFUNK_PYROWAVE_CHUNK_KIB,
PUNKTFUNK_PYROWAVE_STREAMED_AU), verified new by git grep at the v0.25.0 tag
rather than assumed. plugin-kit goes 0.3.2 -> 0.4.0 for the `plugin` launch
kind; the SDK goes 0.1.2 -> 0.1.4; gamescope patch level +pfhdr2 -> +pfhdr4.

Two behaviour changes make a client advertise LESS than it used to, both
deliberate: VIDEO_CAP_444 is now probed against the driver rather than ridden
off the setting alone (every Steam Deck with "Full chroma" on was losing HEVC
entirely, not crispness — no AMD silicon decodes HEVC 4:4:4), and the Decky
client-update check now reports a failure instead of dressing it up as
"up to date".

Bump is the same four files as 0.25.0: Cargo.toml, Cargo.lock,
docs/releases/v0.26.0.md, docs/releases/whatsnew/v0.26.0.txt — plus the
CHANGELOG.md section, which the split at 0.25.0 made part of the ritual.

Gates run locally, all green:
  * cargo fmt --all --check          clean
  * cargo metadata --locked          resolves
  * Cargo.lock diff                  versions-only, 70/70 changed lines, 35 crates
  * Play whatsnew gate               398/500 chars, not byte-identical to any other release
  * notes voice check                0 internal-vocabulary hits above "## For developers"

Notes are 66 lines against 0.25.0's 83, covering 47 commits.

Still owed on glass and recorded in the CHANGELOG's verification table:
iPhone + Bluetooth listen, Apple TV stats overlay, MacBook audio listen, the
Deck HEVC/4:4:4 retest, a Windows wake-from-sleep cycle, and the
PyroWave-under-game-load A/B with CAP_SYS_NICE actually granted.
2026-08-09 01:28:39 +02:00
enricobuehler 4070d043d6 Merge pull request 'PyroWave on Linux: the GPU-priority knob never fired, a dmabuf timeout condemned the host, and the jumbo grow was dead code' (#132) from worktree-wave2-pyrowave into main
apple / swift (push) Successful in 1m35s
android / android (push) Canceled after 3m13s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 3m9s
ci / rust (push) Canceled after 2m42s
ci / rust-arm64 (push) Canceled after 1m10s
ci / web (push) Canceled after 3s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
deb / build-publish (push) Canceled after 2s
deb / build-publish-host (push) Canceled after 0s
deb / build-publish-client-arm64 (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 8s
nix / flake (push) Canceled after 7s
release / apple (push) Canceled after 5m6s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 15s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 14s
windows-host / package (push) Canceled after 3m27s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Canceled after 0s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Canceled after 0s
windows / build (aarch64-pc-windows-msvc) (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 0s
Reviewed-on: #132
2026-08-08 23:23:12 +00:00
enricobuehler ebf61cb448 Merge branch 'worktree-wave2-pw5-encode-overlap' into worktree-wave2-pyrowave
apple / swift (pull_request) Successful in 1m34s
apple / screenshots (pull_request) Skipped
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m9s
ci / bun-nix (pull_request) Successful in 1m19s
ci / docs-site (pull_request) Successful in 1m57s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m13s
ci / web (pull_request) Successful in 3m25s
android / android (pull_request) Successful in 5m4s
ci / rust-arm64 (pull_request) Successful in 6m29s
ci / rust (pull_request) Successful in 11m37s
nix / flake (pull_request) Successful in 18m17s
# Conflicts:
#	crates/pf-encode/src/enc/linux/pyrowave.rs
2026-08-09 01:14:59 +02:00
enricobuehler 4a92c64144 Merge branch 'worktree-wave2-pw7a-jumbo-shard' into worktree-wave2-pyrowave 2026-08-09 01:13:30 +02:00
enricobuehler 2426056465 Merge branch 'worktree-wave2-pw6-streamed-au' into worktree-wave2-pyrowave 2026-08-09 01:13:25 +02:00
enricobuehler d3aaa16a7d Merge branch 'worktree-wave2-pw3-dmabuf-latch' into worktree-wave2-pyrowave
# Conflicts:
#	packaging/arch/punktfunk-host.install
#	scripts/steamdeck/install.sh
2026-08-09 01:13:23 +02:00
enricobuehler 2dd65bdd41 Merge pull request 'Opening the Steam menu on the Deck moved the game too — the pad is now held neutral while an overlay owns it' (#131) from worktree-deck-overlay-input-mask into main
apple / swift (push) Successful in 1m38s
audit / cargo-audit (push) Successful in 1m45s
audit / bun-audit (plugin-kit) (push) Successful in 37s
audit / bun-audit (sdk) (push) Successful in 30s
audit / bun-audit (web) (push) Failing after 26s
audit / docs-site-audit (push) Successful in 26s
audit / pnpm-audit (push) Successful in 11s
arch / build-publish (push) Successful in 9m2s
ci / rust-arm64 (push) Successful in 3m6s
android / android (push) Successful in 10m27s
audit / license-gate (push) Successful in 5m59s
ci / web (push) Successful in 1m30s
ci / bun-nix (push) Successful in 58s
ci / docs-site (push) Successful in 1m56s
release / apple (push) Successful in 9m51s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 1m5s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 1m52s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 12s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 31s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
windows-host / package (push) Successful in 12m31s
windows-host / winget-source (push) Skipped
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 13s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 19s
apple / screenshots (push) Successful in 5m48s
deb / build-publish-client-arm64 (push) Successful in 6m51s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m35s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m45s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m13s
deb / build-publish-host (push) Successful in 11m34s
ci / rust (push) Canceled after 18m48s
deb / build-publish (push) Canceled after 12m54s
docker / builders-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 8m0s
nix / flake (push) Canceled after 8m1s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 8m2s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 5m37s
windows-host / canary-manifest (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 1m30s
Reviewed-on: #131
2026-08-08 22:59:30 +00:00
enricobuehler 5cbaca7789 feat(client/pads): stop forwarding the pad while the Steam overlay owns it
ci / bun-nix (pull_request) Successful in 32s
ci / docs-site (pull_request) Successful in 1m21s
ci / web (pull_request) Successful in 1m37s
apple / swift (pull_request) Successful in 1m41s
apple / screenshots (pull_request) Skipped
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m19s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m37s
ci / rust-arm64 (pull_request) Successful in 5m50s
android / android (pull_request) Successful in 7m56s
ci / rust (pull_request) Successful in 15m55s
nix / flake (pull_request) Failing after 16m33s
On a Deck in Gaming Mode the Steam menu and the QAM are driven by the SAME
physical controller the client forwards, so opening either one moved the game
on the host as well as Steam's UI — a second, invisible player. Steam Input
masks a normal game here; it cannot mask us, because masking happens on Steam
Input's virtual pad and we deliberately forward the REAL one (28DE:1205 — the
virtual pad has no gyro, trackpads or paddles).

SDL ships the exact behaviour we want and it is on by default: presses are
dropped while the process has windows but no keyboard focus, releases still get
through. It CANNOT fire on a Deck. gamescope resolves focus per Xwayland ctx
and the client sits alone in its own, so the Steam overlay — which lives in the
root ctx — never takes our X focus away and no FocusOut is ever generated.
Measured on glass: with the QAM open, X input focus inside the client's ctx
stayed on its window for the whole 4 s, while GAMESCOPE_FOCUSED_APP flipped to
769 (Steam) and GAMESCOPE_FOCUSED_APP_GFX stayed on the app.

So the signal is explicit. `overlay_focus` watches those two atoms on the
gamescope root ctx — which is NOT our own $DISPLAY under `--xwayland-count 2`,
hence the socket-directory walk and the flatpak filesystem line — and the
presenter ORs it with window focus into one `set_masked`.

Masking is deliberately not `set_forwarding`: that closes the slot and sends
GamepadRemove, so the game would see a controller UNPLUG every time somebody
opened the QAM. This keeps every slot open and only stops the transitions,
after flushing what the host believes is held so a stick deflected at
overlay-open stops steering instead of freezing at its last value. On the way
back, held buttons are adopted rather than replayed — the A that picked a QAM
row must not fire in the game as it closes — while axes are re-sent, since a
stick has no press to ghost and SDL only speaks on change.

Fails open throughout: no gamescope, no X, or an unreadable signal all leave
forwarding exactly as it was. `PUNKTFUNK_OVERLAY_MASK=0` opts out.
2026-08-09 00:55:16 +02:00
enricobuehler 9c854893bc Merge pull request 'plugin-kit 0.4.0 — publish the launch surface #129 added, because rom-manager's main is red until it exists' (#130) from worktree-plugin-kit-040 into main
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
ci / rust (push) Canceled after 7s
ci / rust-arm64 (push) Canceled after 2s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 1s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
nix / flake (push) Canceled after 2s
plugin-kit-publish / publish (push) Successful in 39s
Reviewed-on: #130
2026-08-08 22:53:39 +00:00
enricobuehler 6b7997cace chore(plugin-kit): 0.4.0 — the launch surface a plugin needs to publish a tile the host cannot name
ci / docs-site (pull_request) Canceled after 30s
ci / web (pull_request) Canceled after 45s
ci / rust (pull_request) Canceled after 3m3s
ci / rust-arm64 (pull_request) Canceled after 1m52s
ci / bun-nix (pull_request) Canceled after 0s
nix / flake (pull_request) Canceled after 5s
`serveUi({launch})`, `PluginLaunchTarget` and `makeLaunchHandler` (#129) are new API, so this is a
minor bump rather than a patch. It also carries `SyncError.message`, without which a host refusal
reaches a plugin's own UI as the bare tag `SyncError` and nothing else.

Unblocks rom-manager, whose main is currently RED: it merged the consuming change while still
pinning `^0.2.0`, so `bun install --frozen-lockfile` there resolves a kit without these exports and
the typecheck fails on all three. Publishing this and then bumping that pin is the fix — in that
order, because the lockfile cannot resolve 0.4.0 until it exists on the registry.

Tag `plugin-kit-v0.4.0` to publish; the workflow asserts the tag matches this version.
2026-08-09 00:52:37 +02:00
enricobuehler 78ba2342b5 Merge pull request 'rom-manager has been putting 0 games in the library since 08-05 — a plugin launch kind, so a scanner can publish tiles the host cannot name' (#129) from worktree-rom-manager-plugin-launch into main
apple / swift (push) Successful in 1m39s
ci / web (push) Successful in 1m45s
ci / docs-site (push) Successful in 1m35s
ci / bun-nix (push) Successful in 42s
ci / rust-arm64 (push) Successful in 2m45s
deb / build-publish-client-arm64 (push) Successful in 1m4s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 7s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 6s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 9s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 7s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 7s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
deb / build-publish (push) Successful in 3m55s
android / android (push) Successful in 6m31s
docker / builders-arm64cross (push) Successful in 6s
docker / deploy-docs (push) Successful in 30s
apple / screenshots (push) Successful in 6m8s
ci / rust (push) Canceled after 8m47s
deb / build-publish-host (push) Successful in 8m5s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 5m54s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 4m57s
windows-host / package (push) Successful in 11m52s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 17s
arch / build-publish (push) Successful in 13m1s
Reviewed-on: #129
2026-08-08 22:44:44 +00:00
enricobuehler 7d37fe450d test(pf-encode): run PyroWave at depth 2 on real hardware — without shipping depth 2
Wave-2 PW5, the stage-6 experiment. Shipped behaviour is UNCHANGED: `max_inflight` is still 1.

Stage 6 is the frame-corruption stage, and its gate is an on-glass tear-hunt with a live compositor,
a real client and ten minutes of moving content. That is not runnable from here. But the depth-2
risk has two halves, and one of them lives entirely in this crate — the per-slot resources
(`cmd`/`fence`/`csc_set`/y/uv/cursor) and the alternating encoder handles — so that half can be
answered now, on the GPU, and the answer is worth having before anyone attempts the other.

The experiment drives the backend with two frames genuinely in flight (submit N+1, then poll N) and
compares the result against the encoder's OWN synchronous output over the same 16 moving frames.
Its own depth-1 decode is the honest reference: pyrowave's raw AU bytes are not reproducible
run-to-run (see the stage-3 commit), but its decoded planes are.

RESULT, .21 / RTX 5070 Ti (GPU idle at 180 MHz of 3090 — the slow-clock worst case on this card):

  depth-2 vs depth-1 over 16 frames: worst-case PSNR identical (inf)

Bit-identical luma, every frame, in order. So stages 4 and 5 between them are sufficient for the
encoder side: doubling the six single-slot resources and alternating two `pyrowave_encoder` handles
under one monotonic wire sequence really does make overlap invisible to the decoder.

The test is built to fail rather than to pass. Content MOVES every frame (flat fills are the
documented false-green trap — a torn frame stitched from two halves of a static card is invisible),
it asserts two frames were ACTUALLY in flight rather than silently proving nothing, it asserts the
AU count is unchanged, and it carries an off-by-one discriminator that raw PSNR would miss: each
overlapped frame must match its own reference BETTER than it matches the previous one, so a
pipeline delivering frames one position late fails even though every individual PSNR looks fine.

It reaches `max_inflight` directly instead of through a shipped knob, precisely so the shipped
value stays 1.

⚠ WHAT THIS DOES NOT COVER, stated here so the next person does not read it as a green light for
stage 6: the CAPTURE side. `.process` hands the SPA buffer back to the compositor at callback
return while the encode thread holds only a dup of its dmabuf fd, so a second frame in flight
widens the window in which the producer may overwrite a buffer we are still reading by a full frame
period. Nothing in this crate can test that — it needs a live producer. Stages 1 and 2 are what
make it answerable (the pool census says how deep the producer's ring is; the Choice range asks for
headroom), and the on-glass hunt is what would settle it.

Gates green at CI parity.
2026-08-09 00:30:25 +02:00
enricobuehler 077db416ec feat(pf-encode): two PyroWave encoder handles, and the 3-bit landmine that makes them work
Wave-2 PW5 stage 5. Depth is STILL 1 — the handles alternate per frame, one in flight.

PyroWave's `Encoder` cannot hold two frames. Not "probably not" — structurally not. `Encoder::Impl`
owns ONE each of `wavelet_img_high_res`, `bucket_buffer`, `meta_buffer`, `block_stat_buffer`,
`payload_data` and `quant_buffer`, and `Impl::encode` OPENS by discarding them: an image barrier
with `VK_IMAGE_LAYOUT_UNDEFINED` as the old layout — a written promise that nothing else is reading
it — plus three `fill_buffer` clears. Two encodes recorded into two command buffers and submitted
to one queue have no execution dependency in Vulkan (submission order orders the START, not the
completion), so N+1's DWT would overwrite the wavelet bands and zero the RDO buckets while N's
block packing still reads them. Content-dependent, silent.

So overlap means TWO handles on one device, alternated — one per slot. Every resource above is
then private per handle, and within a handle the encodes stay strictly serialized (a slot's next
frame is recorded only after that slot's previous one retired), which leaves patch 0004's
scratch-pool invariant intact without touching it.

THE LANDMINE, and it is the reason this stage is its own commit: `sequence_count` ALSO lives on
`Impl`, and it is the 3-bit counter stamped into every block header. Two handles each count
1,2,3... alone, so the wire sees 1,1,2,2,3,3.... The decoder restarts a frame only when the value
CHANGES (`diff = (hdr.sequence - last_seq) & 0x7; restart = diff != 0`), so a repeat reads as MORE
BLOCKS OF THE SAME FRAME: `clear()` never runs, `decoded_frame_for_current_sequence` stays true,
and the second frame of each pair is swallowed. Half frame rate, occasional mixed-frame blocks, no
error anywhere — on every client, since pf-client-core and the Apple Metal hand-port parse the same
field.

`patches/0007-encoder-sequence-override.patch` (new, ~38 lines) exposes
`Encoder::set_next_sequence` + a `pyrowave_encoder_set_next_sequence` C entry + a
`PYROWAVE_SEQUENCE_MASK` define, so ONE monotonic counter on the Rust side is stamped regardless of
which handle encodes. The setter stores `(seq - 1) & mask` because `Impl::encode` pre-increments —
its contract is about the next ENCODE, not the next store. Inert when unused, so the whole Windows
backend is untouched. No `.def` change: the C API is a static archive.

PREDICTED, THEN OBSERVED. A negative control on .21 (the override call removed, nothing else) reads
the wire out at exactly:

  [1, 1, 2, 2, 3, 3, 4, 4, 5, 5, 6, 6, 7, 7, 0, 0, 1, 1, 2, 2]

which is the analysis's prediction character for character, and with the override:

  +1 mod 8, all 20 frames, through the 3-bit wrap.

THE GATE, `wire_sequence_increments_across_alternating_handles`, checks three things over 20 frames
because any one alone could pass while the stream is broken: the wire counter advances by 1 mod 8;
ONE persistent decoder (its `last_seq` carried across every push, exactly like a client's) reports
every AU decodable; and consecutive decoded pictures DIFFER. Content moves every frame — and the
first run caught a trap in the harness itself rather than the encoder: `test_card` starts its LCG
at `seed | 1`, so seeds 2 and 3 build a byte-identical card and the test faked the very repeat it
hunts. Odd seeds only now, with the reason written down.

A runtime self-check backs the test up where the test cannot reach: after packetize, the stamped
sequence is compared against what we asked for, and a mismatch logs once per process naming patch
0007. A re-vendor that loses the patch would not fail to build — it would fail on glass, subtly,
and this makes it loud instead. Two byte reads per frame.

`reset()` rebuilds both handles and `Drop` destroys both, each with the same null-immediately
discipline the single handle had (`pyrowave_encoder_destroy` is a bare `delete` with no null
check, so a stale pointer left in the field is a double free).

Vendored-patch discipline: patch 0007 re-applies clean to a pristine vendor checkout (verified by
stashing the vendor tree and re-applying), and `git diff crates/pyrowave-sys/vendor/` touches
exactly the four intended files.

VERIFIED ON GLASS (.21, RTX 5070 Ti, GPU idle at 180 MHz of 3090): all 8 `#[ignore]`d GPU tests
pass, including the new gate and the 4:2:0 / 4:4:4 / 24-bpp PSNR smokes.

Gates green at CI parity.
2026-08-09 00:27:28 +02:00
enricobuehler fd98406868 Merge pull request 'The Decky plugin's "update the client" has never once detected an update' (#128) from worktree-decky-client-update into main
ci / web (push) Successful in 1m54s
ci / rust-arm64 (push) Successful in 3m56s
ci / bun-nix (push) Successful in 29s
ci / docs-site (push) Successful in 1m32s
decky / build-publish (push) Successful in 30s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 12s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 8s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 10s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 17s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 15s
ci / rust (push) Successful in 8m12s
docker / builders-arm64cross (push) Successful in 20s
docker / deploy-docs (push) Successful in 46s
Reviewed-on: #128
2026-08-08 22:21:08 +00:00
enricobuehler 5c70a90358 Merge pull request 'A Steam Deck could lose HEVC entirely to a chroma switch nothing checked — and the tool you'd triage it with denied the queue it was decoding on' (#127) from worktree-deck-hevc-shape-gates into main
ci / rust (push) Canceled after 25s
ci / rust-arm64 (push) Canceled after 25s
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
apple / swift (push) Successful in 1m36s
deb / build-publish-client-arm64 (push) Successful in 2m8s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m44s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m51s
deb / build-publish (push) Successful in 5m35s
deb / build-publish-host (push) Successful in 6m5s
apple / screenshots (push) Successful in 6m15s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m20s
android / android (push) Successful in 9m35s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m23s
arch / build-publish (push) Successful in 12m29s
flatpak / build-publish (push) Successful in 10m3s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m52s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 18m7s
Reviewed-on: #127
2026-08-08 22:20:39 +00:00
enricobuehler 29248dcab9 feat(pf-encode): PyroWave had six single-slot resources, not the two the plan named
Wave-2 PW5 stage 4. Pure capacity — `max_inflight` is STILL 1, nothing overlaps yet.

The plan named the y/uv images as the thing to double. Reading the backend found five more, and
each is a correctness problem under overlap rather than a performance one:

  * `csc_set` — ONE descriptor set, rewritten every frame by `bind_rgb`. Updating a set still bound
    by a PENDING command buffer violates VUID-vkUpdateDescriptorSets-None-03047, and on most
    drivers that is a wrong picture rather than an error.
  * `y_img`/`uv_img` — the CSC of N+1 storage-writes exactly the images pyrowave is still sampling
    for N. The barrier comment ("the previous frame's encode already completed under our
    synchronous fence") was load-bearing and said so.
  * `cursor_img` + `cursor_stage` — the struct comment stated the assumption outright: *"Single
    (not ring) because PyroWave encodes one frame synchronously — no in-flight overlap to race."*
  * `cmd` + `fence` — you cannot record into a PENDING command buffer at all.
  * `cpu_img`/`cpu_stage` (software capture / tests) — the host writes staging while the previous
    frame's copy is still pending.

All of it moves into a `Slot`, and the encoder now owns `SLOTS` of them. Two, because Granite caps
the overlap at two for us: the pyrowave device defaults to `init_frame_contexts(2)` and
`next_frame_context()` — called at the top of every `encode_gpu_synchronous` — waits the context it
rotates into. A third slot would need a vendored `init_frame_contexts(3)` that is not exposed.

`bitstream` and `import_cache` are deliberately NOT per-slot, and the `Slot` doc says why so a
later sweep does not "fix" it: `bitstream` is only touched during packetize, i.e. only on the poll
side one frame at a time, and `import_cache` retaining the VkImage/VkDeviceMemory per dmabuf inode
is precisely what makes it safe for two slots to sample the same imported buffer. `cpu_expand` is
shared for the same reason — it is copied into staging before `submit_frame` returns, so no GPU
work ever reads it.

Each frame carries its slot index in `InFlight` rather than recomputing it, so `wait_and_packetize`
cannot wait the wrong fence — the failure that would look like corruption rather than an error.
`reset()` now waits EVERY in-flight fence, not just one, which matters the moment depth rises.

WHAT IT COSTS, measured from the driver's own memory requirements rather than estimated (.21,
RTX 5070 Ti, and there is now an `#[ignore]`d test that prints it on any GPU):

  1080p 4:2:0   3872 KiB per slot    7744 KiB for both
  4K    4:2:0  12992 KiB per slot   25984 KiB for both
  4K    4:4:4  24992 KiB per slot   49984 KiB for both

So the extra slot costs ~3.8 MiB at 1080p and ~24 MiB at 4K 4:4:4 — an order of magnitude under
the plan's ~25-35 MB / 100-150 MB estimate, because that estimate included pyrowave's internal
wavelet and scratch buffers, which stage 5's second encoder handle will add and this stage does
not. Affordable on an iGPU. The open line now logs `slots`, `slot_kib` and `slots_kib` so this is
visible per session and not only in a test.

VERIFIED ON GLASS (.21, GPU idle at 195 MHz of 3090 — slow-clock, the worst case on this card):
all 6 `#[ignore]`d GPU tests pass, and all NINE decoded-plane hashes (`ref-dense-{y,cb,cr}`,
`ref-chunked-*`, `ref-dense444-*`) are bit-identical to the pre-PW5 base. Decode identity is the
meaningful gate here — the raw AU bytes are not reproducible run-to-run even from an unmodified
binary, which stage 3's message documents.

Gates green at CI parity.
2026-08-09 00:18:38 +02:00
enricobuehler 95962f55d0 refactor(pf-encode): PyroWave waited its fence inside submit — the one backend that did
Wave-2 PW5 stage 3. Depth is STILL 1; this is the shape change alone.

`encode_frame` recorded CSC+encode, queue-submitted, waited the fence and packetized, all inside
`Encoder::submit`. Every other backend in this crate puts the wait on the POLL side. That
difference is the whole reason the host loop's cadence folds around this encoder: with the wait
inline, `submit` returns only after the GPU is done, so the arrival-anchored floor absorbs the
encode only while it stays under 0.9x the frame interval.

Split into `submit_frame` (ingest -> CSC -> pyrowave encode -> queue-submit -> return) and
`wait_and_packetize` (fence wait -> packetize -> AU), with an `InFlight` deque between them capped
by `max_inflight`, which is 1. **One is the only value the resources can support today** — `cmd`,
`fence`, `csc_set` and the y/uv images are one each, so a second concurrent frame would record into
a PENDING command buffer and storage-write images pyrowave is still sampling. `submit` therefore
drains to `max_inflight - 1` before recording, which states that invariant in one place instead of
leaving it implicit in "the encode is synchronous".

The subtle part is the command-buffer state machine, and it is unchanged: the record-and-submit
closure still resets `cmd` on every PRE-submit failure (RECORDING/INVALID/EXECUTABLE, never
PENDING), and the fence wait still does NOT reset on failure, because a timeout leaves the buffer
PENDING where a reset violates VUID-vkResetCommandBuffer-commandBuffer-00045. What changed is that
a failed wait now also leaves the entry IN FLIGHT — which is precisely what tells `reset()` there
is live GPU work to re-wait before the pyrowave encoder object may be destroyed. `gpu_pending` is
gone; `!inflight.is_empty()` is the same fact, and cannot drift from it.

The split opened two windows that did not exist when everything ran inline, both closed here:
`reconfigure_bitrate` and `set_wire_chunking` can now land BETWEEN a submit and its poll, so the
packetize boundary and the bitstream cap are snapshotted into `InFlight` at submit time. Reading
the live fields would have let a mid-flight bitrate drop turn a perfectly good frame into
"unexpected packet count", and a mid-flight chunking change into an AU with the wrong
`chunk_aligned` flag.

`flush()` is no longer a no-op — it drains the in-flight frame, so the trait's poll-until-None
contract still returns every AU (the `spike` subcommand and the hardware smoke tests are the real
users).

The perf instrument still measures submit->AU, stamped at submit and taken when the AU becomes
readable, so `92326312`'s numbers stay directly comparable; the log line now carries `depth` and
says plainly that above depth 1 the number legitimately grows by about one loop period.

VERIFIED ON GLASS (.21, RTX 5070 Ti, GPU idle at 180 MHz of 3090 — so these are slow-clock runs,
which is the worst case on this card, not the best): all 6 `#[ignore]`d GPU tests pass — the
4:2:0, 4:4:4 and 24-bpp PSNR smokes, the mode-mismatch refusal, the fd-leak check and the golden
dump.

Byte-identity, honestly: the AU bytes are NOT reproducible, and were not before this commit
either. Three runs of the SAME unmodified binary produced three different `au-dense.bin` hashes
(ab7ecaf6 / 8735700e / 933b3d40) — the vendored 4:2:0 encoder emits run-varying bytes that the
decoder ignores. So the meaningful gate is DECODE identity, and that holds exactly: every decoded
plane (`ref-dense-{y,cb,cr}`, `ref-chunked-{y,cb,cr}`, `ref-dense444-{y,cb,cr}`) is bit-identical
between the pre-split base and this commit, across four runs. 4:4:4 AUs are additionally
bit-stable and match the checked-in Apple fixture exactly.

Gates green at CI parity.
2026-08-09 00:11:22 +02:00
enricobuehler 9e598f8595 fix(client): the 4:4:4 switch could cost a Deck its whole codec, and --probe-decode denied the queue it was decoding on
ci / bun-nix (pull_request) Successful in 41s
ci / web (pull_request) Successful in 1m17s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m21s
apple / swift (pull_request) Successful in 1m35s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m37s
ci / rust-arm64 (pull_request) Successful in 3m39s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m22s
android / android (pull_request) Successful in 3m56s
ci / rust (pull_request) Successful in 13m6s
Two Steam Deck findings from a field report of "the decoder was not found, it
fell back to H.264 — but sometimes HEVC worked".

**The 4:4:4 advertisement was a promise nothing checked.** `VIDEO_CAP_444` rode
the "Full chroma" setting alone. That was safe while a software HEVC decoder
existed underneath it; M8 removed one (there is no permissively licensed HEVC
CPU decoder, so `software_decodable_codecs()` is H.264|AV1). The host grants
4:4:4 on HEVC ONLY, and answers the resolved chroma in the Welcome before the
client builds a decoder — so on a device with no 4:4:4 decode the toggle did not
cost crispness, it cost the entire codec: the Vulkan rung refuses the shape at
construction, VAAPI refuses it too, there is no CPU rung, and the session
reconnects on H.264. AMD has no HEVC 4:4:4 decode on any silicon, so every Deck
with that switch on lost HEVC. It is per-profile and default-off, which is
exactly why it looked intermittent — a "Work" profile lost HEVC where "Game"
kept it, same box, same host.

Gated on `hevc_444_hardware_decodable`, which asks the driver through the SAME
code the rung uses at construction (`VkH265Decoder::probe_stream_support`), so
the advertisement and the rung that must honour it cannot disagree. Both depths
are required, not either: with HDR on the host may resolve 4:4:4 10-bit, and a
device offering YUV444_8 but not YUV444_10 would land in the same hole.

Answering from the Vulkan rung alone is exact rather than approximate — it is
the only rung in this build that implements 4:4:4 at all (`pf_vaadec::profile_for`
errors on chroma_format_idc 3, pf-dxvadec refuses anything but 4:2:0, the CPU
rung is 8-bit 4:2:0). Deliberately NOT extended to VIDEO_CAP_10BIT/HDR: all
three rungs implement 10-bit 4:2:0, so a Vulkan-only probe there would withdraw
HDR from boxes whose VAAPI/DXVA rung decodes it perfectly — a real regression
against a case never observed.

The bit arithmetic moves into `video::video_caps_for` so the part that was
wrong is testable without a GPU, a host or a Hello; the test is verified
non-vacuous against the planted original defect.

**`--probe-decode` described a different device from the one that streams.** The
RADV video-decode opt-in sat AFTER the --list-adapters/--probe-decode/--list-audio
/--pair early exits, so the triage tool never had it. Measured on a Deck
(canary e22af40f), same binary back to back: bare `--probe-decode` printed
"vulkan video decode: no", "driver decode ops: none (0x0)", "no queue family
advertises VIDEO_DECODE"; with RADV_PERFTEST=video_decode in the environment,
"YES" and "H.264, H.265, AV1, VP9". Any Deck triage that consulted it reached
the opposite of the truth. Hoisted to the top of `run`, ahead of every early
exit — nothing touches Vulkan before it (`main` calls `run` directly).

Gates, in the Linux container: fmt, plain `cargo build` (not only
--all-targets), `clippy --all-targets -D warnings`, and 185 tests.
2026-08-09 00:10:46 +02:00
enricobuehler bd86598d97 fix(decky): the client update the plugin offers was never once detected
ci / bun-nix (pull_request) Successful in 21s
ci / web (pull_request) Successful in 1m19s
ci / docs-site (pull_request) Successful in 1m22s
ci / rust-arm64 (pull_request) Successful in 2m22s
ci / rust (pull_request) Successful in 5m23s
The QAM has offered to update the client since 0.24, and on every Deck it has
answered "up to date" — including right now, with a client a day out of date.

The check asks flatpak for the remote's commit and compares it to the installed
one, and it named the app id with no branch: `flatpak remote-info punktfunk-origin
io.unom.Punktfunk`. The punktfunk remote publishes `stable` AND `canary`, so that
ref is ambiguous and flatpak refuses it — "Multiple branches available" — rather
than picking one. One branch INSTALLED does not help; the ambiguity is on the
remote. The call failed on every box, every time, and the failure returned
`available=False`, which the panel renders as good news. Hence: the plugin
appeared to update only itself.

Every query now names the ref in full, resolved once by `_flatpak_ref()` off the
exported tree (no subprocess — `_client_argv` is on the path of every headless
call). That resolution also carries the SCOPE, so a system-wide install is no
longer invisible to a check that hardcoded `--user`, and the launcher pins the
same `--branch=`, so the client we start is the client we check and update.

A check that cannot run now says so instead of reporting up-to-date: the flatpak
leg reports `client_error` exactly as the native leg already did. Dressing that
failure up as good news is the whole reason this went a week unnoticed.

Also: the button no longer promises "+ client" when the client is manual-only and
the tap can only print a command.

Verified on the Deck (192.168.1.253, canary, user scope) by running both code
paths against the real install, minutes apart:

  pre-fix   available=False  remote=''
  post-fix  available=True   remote=ca010668  (installed e22af40f)

and `flatpak {info,remote-info,update}` all accept the `id//branch` form there.
37 backend checks pass, 6 of them new and about exactly this.
2026-08-09 00:08:24 +02:00
enricobuehler c3ecc29117 feat(pf-capture): the zero-copy path never asked the compositor for buffer headroom
Wave-2 PW5 stage 2, on the number stage 1 just made visible.

`build_dmabuf_buffers` set `SPA_PARAM_BUFFERS_dataType` and stopped there — no
`SPA_PARAM_BUFFERS_buffers` at all, so the pool depth the whole zero-copy safety argument rests on
was entirely the producer's choice, and we never even expressed a preference. This asks for 8
(min 2, max 16).

A **Choice Range**, deliberately, not a fixed count. SPA intersects the consumer's and producer's
Buffers params, so a fixed 8 against a producer that can only afford 4 empties the intersection and
the link stalls in "negotiating" with no error anywhere — the exact trap that cost this codebase
the entire Linux cursor channel once, when a 256^2 cursor-meta max failed to intersect Mutter's
fixed 384^2 offer. With a range the producer clamps into it and negotiation still succeeds; the
min stays at 2 so nothing that works today stops working.

The numbers, and what they are not: 8 buffers is ~133 ms of pool at 60 Hz and ~33 ms at 240 Hz,
well past the ~3-4 ms capture-to-fence latency PW3/PW4 measured, with room for a second frame in
flight. 16 is a ceiling rather than a request — a 4K 4:4:4 buffer is ~25 MB, so 16 of them is
~400 MB of compositor allocation. These are the values we ASK for; what a producer actually
allocates is what stage 1's census line reports, and that line is the one to trust.

Scoped to the dmabuf pod only. The mappable and SHM-only builders are untouched: their consumers
copy out of the buffer inside `.process`, so pool depth is not part of their correctness argument.

A test pins the pod SHAPE — Choice, Range, Int children, values default-first — so a later
simplification cannot quietly turn the range back into a number and take the negotiation down with
it.

Gates green at CI parity; on-glass negotiation on each producer is stage 2's own gate and is
reported with the stage-1 census numbers.
2026-08-09 00:01:24 +02:00
enricobuehler 6d550530fe feat(pf-capture): nothing had ever counted the compositor's buffer pool — the number every zero-copy safety argument rests on
Wave-2 PW5 stage 1, and the one stage with no risk at all.

The zero-copy capture path dups the dmabuf fd, publishes the frame, and hands the SPA buffer
straight back to the producer at `.process` return — while the encode thread has not yet imported
it, let alone read it. The code says so itself ("content stability across the brief import/encode
window relies on the compositor's buffer-pool depth, like any zero-copy capture"). That depth is
therefore load-bearing: it is the ONLY thing standing between us and the producer overwriting a
buffer mid-read.

And it had never been measured. Not logged, not asserted, not even requested — `build_dmabuf_buffers`
set `SPA_PARAM_BUFFERS_dataType` and nothing else, so whatever the producer picked is what we got,
silently.

This adds the `add_buffer`/`remove_buffer` stream callbacks PipeWire has always offered and logs the
count once per distinct depth: `pool_depth`, `high_water`, and the latest-frame-only `drained`
count beside it. One line per session on a stable pool (`.process` runs at the capture rate — an
unconditional log would be 240 lines a second of the same number), a second line if a
renegotiation changes the depth.

`high_water` is tracked separately from `live` because a renegotiation frees the pool before
re-allocating it: any decision keyed on the live count would read that dip as "the pool shrank".
`remove` saturates at zero rather than wrapping, so an unmatched remove cannot report `u32::MAX`
buffers.

Measurement only — no behaviour change, and no consumer of the number yet. PW5's later stages need
it (a deeper encode pipeline widens the overwrite window by a full frame period), but the number is
worth having regardless of whether those stages ever land: it is the answer to "is our zero-copy
capture actually safe on this compositor", and until now the honest answer was "nobody knows".

3 tests pin the once-per-depth logging, the renegotiation dip, and the saturating remove.

Gates green at CI parity.
2026-08-09 00:01:08 +02:00
enricobuehler e22082ac2a Merge pull request 'No audio over Bluetooth on iOS: .defaultToSpeaker is an override that outranks A2DP' (#126) from worktree-ios-bluetooth-audio-route into main
ci / bun-nix (push) Successful in 37s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 17s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 6s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 16s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 12s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 19s
docker / builders-arm64cross (push) Successful in 14s
ci / web (push) Successful in 2m17s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m32s
ci / docs-site (push) Successful in 2m37s
ci / rust-arm64 (push) Successful in 3m44s
apple / swift (push) Successful in 1m37s
ci / rust (push) Successful in 6m29s
docker / deploy-docs (push) Failing after 6m50s
release / apple (push) Successful in 9m45s
apple / screenshots (push) Successful in 5m52s
Reviewed-on: #126
2026-08-08 21:54:23 +00:00
enricobuehler 4f5ca5f9bc docs(pf-capture): KWin/RADV was the last untested producer — it has no fence either
Closes PW4's one remaining gap. The Steam Deck switched to Desktop Mode gives KWin on RADV, the
combination none of the earlier legs covered, and it reports no implicit fence like every other:

  gamescope + NVIDIA (RTX 5070 Ti)   NoFence
  Mutter    + NVIDIA (RTX 5070 Ti)   NoFence
  gamescope + RADV   (Deck VANGOGH)  300/300 NoFence, mean 23us, p99 <=100us
  KWin      + RADV   (Deck desktop)  no fence  (older build's wording: waited=false)

That is every compositor x vendor this fleet has. PW4 retires with no outstanding doubt rather
than "probably fine except one box we never tried".

Measured with the Deck's OWN already-authorized binary rather than a scratch build, because KWin
grants zkde_screencast_unstable_v1 per EXECUTABLE PATH: it resolves /proc/<pid>/exe against a
.desktop's Exec= and caches the grant on first connect, so an unregistered path is refused outright
and registering one needs a re-login. The fence probe is pre-existing capture-path code, so a build
from July answers the outcome question perfectly well — and nothing of the user's was modified to
get it.

Comment-only; no behaviour change. fmt + pf-capture clippy -D warnings green.
2026-08-08 23:52:54 +02:00
enricobuehler 07f6d6f324 fix(apple): .defaultToSpeaker outranks Bluetooth, so every headset lost the stream
ci / bun-nix (pull_request) Successful in 57s
ci / docs-site (pull_request) Successful in 1m21s
ci / web (pull_request) Successful in 1m40s
ci / rust-arm64 (pull_request) Successful in 2m42s
apple / swift (pull_request) Successful in 1m55s
apple / screenshots (pull_request) Skipped
ci / rust (pull_request) Successful in 12m44s
Field report on 0.25, iOS: "no audio over Bluetooth ... plays through speakers
if Mic input is enabled".

Both halves are one bug. `micEnabled` and `echoCancel` both default to true
(EffectiveSettings.swift), so the DEFAULT iOS session is `.playAndRecord` — and
that branch set `.defaultToSpeaker`. That option is not the polite preference it
reads as: it is an output OVERRIDE, and it outranks an A2DP route. Wired
headphones beat it, Bluetooth does not, so a cable is the one way to test it and
get the right answer — which is what the comment sitting on it asserted
("headphones/BT still win"). Every Bluetooth listener on the default settings got
the phone's own speaker instead. Turning the mic off was the accidental
workaround the reporter found: that path takes `.playback`, which routes to A2DP
happily and always did.

The earpiece problem `.defaultToSpeaker` was reaching for is real —
`.playAndRecord` really does park the built-in output on the receiver. So solve
it against the route we were ACTUALLY given rather than pre-emptively: after
activation, if the current output is `.builtInReceiver`, override to the speaker;
anything external (Bluetooth, wired, CarPlay, AirPlay) is left strictly alone.

That override is a property of the current route — iOS drops it whenever the
route changes, which is exactly what lets a newly-connected headset win — so it
has to be re-applied per route. Hence the route-change observer: without it,
dropping Bluetooth mid-stream would hand the game to the earpiece. Registered
only for a `.playAndRecord` session (a `.playback` one needs no steering),
removed in stop() before the session deactivate, with deinit as a backstop.

Deliberately NOT adding `.allowBluetooth`: it would make a headset's mic usable,
but buys that by dragging the whole route onto HFP/SCO and collapsing game audio
to narrowband. High-quality A2DP output plus the built-in mic is the better trade
for a game-streaming client.

Verified: builds clean on arm64-apple-ios17.0 (the triple that actually compiles
these `#if os(iOS)` blocks — a plain `swift build` is macOS and skips them),
arm64-apple-tvos17.0, and macOS; 257 Swift tests pass, 0 failures.
On-glass iPhone + Bluetooth listen still owed.
2026-08-08 23:49:38 +02:00
enricobuehler 54666e66da Merge pull request 'The Apple TV client had no way to show its statistics overlay' (#125) from worktree-tvos-stats-shortcut into main
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 13s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 15s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 14s
ci / bun-nix (push) Successful in 53s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 14s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
ci / web (push) Successful in 1m17s
ci / docs-site (push) Successful in 1m27s
apple / swift (push) Successful in 1m34s
docker / builders-arm64cross (push) Successful in 15s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 27s
ci / rust-arm64 (push) Successful in 2m24s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m12s
docker / deploy-docs (push) Failing after 1m49s
ci / rust (push) Successful in 5m6s
apple / screenshots (push) Canceled after 0s
release / apple (push) Canceled after 4m52s
Reviewed-on: #125
2026-08-08 21:48:19 +00:00
enricobuehler 5872dfc649 feat(library): a plugin launch kind, so a scanner can publish tiles the host cannot name
apple / swift (pull_request) Successful in 1m40s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m34s
ci / web (pull_request) Successful in 3m46s
ci / bun-nix (pull_request) Successful in 54s
ci / rust-arm64 (pull_request) Successful in 5m54s
android / android (pull_request) Successful in 7m29s
ci / rust (pull_request) Successful in 21m2s
The 2026-08-05 review made `launch.kind = "command"` operator-only, and a reconcile refuses
on the FIRST offending entry — so rom-manager, whose every ROM is `<emulator> <args> <rom>`,
stopped putting anything in the library at all. Playnite hit the same wall and was rescued
with a typed kind the host resolves itself; there is no fixed scheme for "whichever emulator
the operator configured, with the core and flags they chose", so that trick does not
generalise.

So the entry now carries an opaque key and nothing executable, and the host asks the plugin
that owns it what to run — at launch time, over the loopback UI port and per-boot secret it
already registered. A stolen plugin token stops being command execution: planting an entry is
not enough, because the live plugin answers 404 for a key it never published. Nothing
executable is persisted or served to a client, and an emulator that moved is picked up on the
next launch instead of leaving a dead tile (the same reasoning as `xbox` resolving its AUMID
at launch time).

The host still SPAWNS it, because only the host can put the process where the stream can see
it: on Linux the line is either gamescope's own argv or a spawn carrying the session's
compositor env, and the returned child is what session-game-lifetime tracks to know the game
exited. A plugin spawning the emulator itself would land it outside both.

- library/plugin_launch.rs — the ask: blocking ureq, bounded body, absolute cwd, no control
  characters, and a log line for every way it can come back empty
- library/launch.rs — `plugin_recipe` tried before both per-OS resolvers, plus
  `launch_is_resolvable` so the async handshake probe never makes the blocking call
- native.rs — the session's `resolve_launch` moves onto `spawn_blocking`
- plugin-kit — `serveUi({launch})` serves `POST /__launch`; and `SyncError` finally renders
  its cause, which is why a host refusal with a fully explanatory 403 could reach a plugin's
  own UI as nothing but "Decode error"
2026-08-08 23:46:05 +02:00
enricobuehler bed58b75b6 feat(apple): the statistics overlay is reachable on tvOS
ci / web (pull_request) Successful in 1m0s
ci / bun-nix (pull_request) Successful in 1m11s
ci / docs-site (pull_request) Successful in 1m18s
apple / swift (pull_request) Successful in 1m34s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 4m5s
ci / rust (pull_request) Successful in 6m27s
An Apple TV session had no way to the stats overlay at all. Every other client
cycles it in-stream — Ctrl+Alt+Shift+S on the desktops, a three-finger tap on
touch — and tvOS has neither a keyboard nor a screen to tap, so the only route
was Settings before connecting (or a profile). The docs' own "cycle with" table
simply had no row for it.

Two surfaces, because an Apple TV may have a controller in the room or only the
remote:

- Select + X on a controller, cycling one tier per completion. Built like
  Android's mic chord (Select + Y) and deliberately disjoint from the escape
  chord — X is none of its four buttons, so reaching for one can never trip the
  other. Read off the wire mask like the escape chord, so a Select the
  hold-Select gesture has turned into a guide can't cycle the overlay on its way
  past. Available on every Apple platform: a controller in both hands is exactly
  the case the keyboard combo and the three-finger tap can't serve.

- Hold Play/Pause on the Siri Remote. Its right-click is therefore deferred until
  the press resolves — a tap still right-clicks, delivered on release with the
  release trailing by TAP_PRESS — because a right button held for half a second
  is a context menu on every desktop this streams.

A non-forwarding slot now claims the stats chord's elements too, alongside the
escape chord's: on tvOS an unclaimed button's press stays the system's and the
chord would silently never complete.

Tests pin both chords' masks against their GameController alias lists, that the
two overlap only on Select, and that the claim list covers both without
duplicates — the failure mode is nothing happening, with nothing logged.
2026-08-08 23:24:09 +02:00
enricobuehler ce31a9ddfd Merge pull request 'The Steam Deck updater kept sabotaging its own next update, and hand-deleting a file was the only way through' (#122) from worktree-steamdeck-update-bunnix-dirt into main
ci / docs-site (push) Successful in 1m8s
ci / web (push) Successful in 1m17s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 16s
ci / bun-nix (push) Successful in 20s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 9s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 8s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 6s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 5s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
docker / builders-arm64cross (push) Successful in 7s
ci / rust-arm64 (push) Successful in 2m36s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 48s
docker / deploy-docs (push) Successful in 28s
ci / rust (push) Successful in 12m23s
Reviewed-on: #122
2026-08-08 21:17:27 +00:00
enricobuehler 8fe834c89b Merge pull request 'Every gamescope session ended in a SIGSEGV at exit — the Vulkan device was being destroyed after the driver had gone' (#124) from worktree-gamescope-exit-segfault into main
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 13s
ci / bun-nix (push) Successful in 25s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 11s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 40s
ci / web (push) Successful in 1m1s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
ci / docs-site (push) Successful in 1m45s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m0s
docker / builders-arm64cross (push) Successful in 11s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m19s
docker / deploy-docs (push) Successful in 29s
ci / rust (push) Canceled after 3m41s
ci / rust-arm64 (push) Canceled after 3m42s
arch / build-publish (push) Successful in 13m16s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 22m19s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 23m3s
Reviewed-on: #124
2026-08-08 21:13:43 +00:00
enricobuehler 744bcb468b feat(host/wire): a jumbo path can now be PROVEN — and the shipped grow never could
Wave-2 PW7a: a PyroWave session on a proven-jumbo LAN should START at the big shard, because it
is the one codec that can never be re-keyed mid-stream (its client parses chunk-aligned AUs in
windows of the `Welcome` value, read once over the C ABI). At an 8908-byte shard that is ~6×
fewer datagrams per frame — ~49k → ~8k pps at 550 Mb/s — and proportionally less window-tail
padding.

THE BLOCKER FOUND FIRST: the whole jumbo leg was dead code, not just the missing half. quinn
caps a peer's MTU-discovery search at `min(MtuDiscoveryConfig::upper_bound, the OTHER side's
advertised max_udp_payload_size)` (`quinn_proto::connection::mtud::SearchState::new`), and
`EndpointConfig::max_udp_payload_size` defaults to 1472. Nothing in the repo had ever touched
`EndpointConfig`, so raising the host's PROBE ceiling — all `stream_transport_idle` did — could
never make discovery settle above 1472, and the shipped mid-session grow's
`settled >= sealed_datagram_bytes(target)` gate was unreachable on every path that has ever
existed. Two smaller contributors, fixed here too: the watcher stopped sampling the moment
`settled >= 1472`, discarding the very climb the proof needs, and a session sealed ABOVE the
1500-byte default was never checked against the path at all.

The advertisement is raised on the CLIENT endpoint, under the same `jumbo_wire_mtu()` opt-in as
the probe ceiling, because it is not free: quinn sizes its endpoint receive buffer
`max_udp_payload_size × max_receive_segments × BATCH_SIZE`, so on a GRO-capable Linux/Android
client that is ~2.9 MiB at the default and ~18 MiB at jumbo (47 KiB → 288 KiB on Apple/Windows).
Consequence: jumbo now needs the opt-in on BOTH ends. Without it, every byte on the wire and
every byte of buffer is exactly what it was.

WHY THE GROW IS AS SAFE AS THE CLAMP, which is not obvious — the failure modes are opposite. A
stale clamp only makes datagrams smaller than they had to be; a stale grow seals an oversized
datagram onto a 1500-byte path, where it is silently dropped, and a PyroWave session cannot
recover from that for its whole life. Mirroring the clamp's keying is therefore NOT sufficient.
So the memory is demoted: the persisted verdict only decides whether it is worth WAITING for a
proof, and what authorises the grow is a LIVE re-proof on the very connection being welcomed —
`conn.stats().path.current_mtu` ≥ the sealed target, i.e. a datagram of exactly that size acked
by this client, on this connection, seconds ago. The moved laptop cannot inherit anything: its
new path's live MTU is 1472 and the grow does not happen, whatever the memory says.

The remembered half is keyed strictly anyway — `(local_ip, peer_ip)`, so a verdict earned over
the host's 10 GbE NIC does not apply to the same peer over Wi-Fi or a VPN — and carries the
operator target it was proven under plus a 6 h TTL. It is erased by any contrary evidence: a
lower settle, a session that ended before the window closed (what a client staring at black
does), a changed opt-in, or a constrained-path clamp that disagrees.

The proof-wait is on the bring-up critical path (`handshake.rs` sends the `Welcome` and only
then kicks the display prep), so it is bounded at 300 ms, exits the instant the proof lands, and
is entered ONLY for a path a previous session already proved. Its worst case is the moved
laptop, and that is self-limiting: that session's watcher erases the verdict.

MEASURED, NOT ARGUED: `mtu_discovery_climbs_only_as_high_as_the_peer_advertises` (`#[ignore]`d,
loopback — whose own MTU is 64 KiB, so configuration is the only thing that can stop the search),
on .21:

  leg A (server opted in, client NOT): settled at 1472 B UDP payload   <- the dead-code proof
  leg B (both opted in):               reached 8972 B in 5 ms          <- the fix, and its speed

Leg A is the finding restated as an experiment. Leg B says the climb costs ~5 ms once both sides
advertise it, so the 300 ms proof-wait is ~60x the loopback convergence time — enough headroom
for a real LAN's RTT and per-probe ack delay across the ~11 probes the search takes.

Still owed: the A/B on a real jumbo LAN segment (9000-MTU NIC + switch on both ends) — pps per
frame, wire/pin ratio, and a PyroWave session observed starting at 8908. Not runnable without
the hardware.
2026-08-08 21:53:02 +02:00
enricobuehler 20f4d23f2d test(pw6): the streamed-AU trap is real — and at 2 % loss it costs exactly nothing
PW6 shipped behind a knob because one pre-registered risk was unmeasured: a
streamed frame whose FINAL block is lost has no totals, so where the whole-AU
path hands the consumer a usable blurred partial, a streamed frame may deliver
nothing. PyroWave clients opt into partial delivery unconditionally, so this
would have been a live behaviour change for every one of them. Measured now,
three ways, instead of reasoned about.

`tools/loss-harness` gains a partial-delivery leg: FEC pinned OFF, chunk-aligned
AUs, deliver_partial ON, realistic 1408/200 geometry, and AU sizes swept across
the whole 1..=200-shard range of FINAL-block sizes — because the final block's
size is what bounds the exposure. Loss is injected per packet from a seeded
xorshift rather than through `loopback_drop_period`, whose deterministic 1-in-N
would systematically always-or-never hit the final block, which is the entire
question. `tc netem` on `lo` was deliberately not used: the in-process model
gives exact per-frame attribution, needs no sudo, cannot disturb a box running a
live desktop session, and — decisively — can drop precisely the final block.

Leg 1, deterministic (drop exactly the last block, 200 frames): whole-AU
delivers 200 partials and 0 losses; streamed delivers 0 partials and 200 total
losses. The trap is real and, when it fires, total.

Leg 2, random loss, 20 000 frames per cell, same seed and sizes for both shapes.
At 2 % the two are indistinguishable — 20000/20000 partials and ZERO vanished
frames on both, matching the analytic bound E[loss^k] over final-block sizes k
(~1e-4). The gap only appears at 30 % (99.94 % vs 100 % rescue) and 50 %
(99.79 %). `complete` is 0 throughout by construction: with FEC off and ~500
packets per AU, essentially every frame is damaged — which is the regime the
partial path exists for.

The spike gains `--wire-chunk` and a streamed loopback path, so the wire shape
is reachable end to end outside a real client: `poll_chunk` drains the AU,
`begin_streamed_frame_at`/`seal_streamed_chunk`/`seal_streamed_finish` seal each
piece, and the client byte-compares the reassembly. On 120 real PyroWave AUs the
streamed legs (56.5 and 2.0 chunks/AU) and the whole-AU control emit a
byte-identical 47 373 568-byte stream with 0 mismatches — the cut changes the
wire shape and not one byte of content, and with the knob unset it does not
engage at all.

A new `#[ignore]`d GPU test closes the picture question on real hardware with a
BUSY card (gradients + checker + noise), never a flat fill: chunks are whole
windows, exactly one `first` and one `last`, the AU decodes through the client's
own window walk, and luma PSNR lands at 40.2 dB. Unset the knob and the test
refuses to run, which is the default-off claim verified rather than asserted.

Verdict recorded in the plan: KEEP IT OFF. The 2 % tie is an argument about
typical loss, but the failure is not graceful when it fires and the measured win
is host send-side pipelining that nobody has yet put a millisecond number on.
2026-08-08 21:26:25 +02:00
enricobuehler 49f5c815ea feat(pf-encode): PyroWave can stream its AU to the wire — and newest-wins was never in the way
PW6 was gated on one question: what happens to the client's newest-wins
draining when a PyroWave AU arrives in pieces, given that
`Session::set_deliver_frame_parts` refuses to combine with an all-intra
stream. The answer is that the doc and the plan conflated two different
axes, and the question never applied to this package.

Host STREAMED_AU chunks change only the WIRE shape. The reassembler
completes such a frame exactly like a whole one (`block_count != 0 &&
blocks_ok == block_count`) and hands up ONE Frame, so the frame channel
still sees one entry per AU and the drain is untouched.

What newest-wins genuinely cannot survive is the client's SEPARATE prefix
delivery, and the mechanism is sharper than "assumes whole AUs" said:
`FrameChannel::pop` counts QUEUE ENTRIES and takes one entry to be one AU.
With parts on, one AU pushes several, so `len > 1` stops meaning "the
consumer is behind" — the drain fires mid-AU, returns a SUFFIX and clears
that same AU's prefixes. For PyroWave that is fatal rather than lossy: the
sequence header lives in window 0 of every AU (`au_dims` reads it there), so
every frame would arrive headerless, and `FramePart`'s own orphan contract
would have a correct consumer abandon essentially all of them. Written into
`pop`, `set_deliver_frame_parts` and the handshake, together with what a fix
would take (skip whole SUPERSEDED AUs, never split one).

That answer shrinks what this package may claim, so the code says so
plainly. `encode_frame` is synchronous: the whole AU exists before the first
chunk can be polled, so `poll_chunk` is not "emit as produced" and there is
no encode/send overlap here (PW6 ⟂ PW5, confirmed). And with the client
still receiving one whole Frame there is no decode-while-arriving either —
the "~7 ms, decouple e2e latency from AU size" framing needs client work
this commit does not do. What IS left is real and host-side: the whole-AU
path FEC-protects, packetizes and seals the entire ~830 KB AU before its
first datagram may leave the socket, while the streamed path seals and paces
each FEC block as it completes.

All of the cutting lives in the shared `pyrowave_wire` helper, which
compiles and unit-tests on every platform, so both backends' `poll_chunk` /
`supports_chunked_poll` are thin delegations — the Windows backend cannot be
compiled from a Linux box, and logic written into it directly would ship
unverified. Chunks are whole numbers of framing windows because `build_au`
gives each window exactly ONE kind; that also makes them shard-aligned for
free, which is what the sealer's sentinel bases require. Dense mode never
streams (no window framing to cut on). `poll()` now errors while a chunk
cursor is live — the trait's one-drain-method-per-AU contract, where
double-emitting would put the same bytes on the wire twice under one frame
index — and `reset()` drops the cursor so a rebuild cannot splice a dead
AU's tail onto a fresh one. No new Encoder trait method, so neither the
TrackedEncoder forwarding trap nor the EncoderCaps default trap is in play.

Shipped OFF: `PUNKTFUNK_PYROWAVE_STREAMED_AU=1` arms it,
`PUNKTFUNK_PYROWAVE_CHUNK_KIB` tunes the 256 KiB target. The pre-registered
partial-delivery trap is real and now has a named cost — an unpinned
streamed frame (final block lost) is excluded from partial delivery, where
the whole-AU path still hands the consumer a usable blur, and PyroWave
clients opt into partials unconditionally. The netem loss-harness leg is the
prerequisite for default-on and has not been run.
2026-08-08 19:23:05 +02:00
enricobuehler 9e7713eecf fix(gamescope): every session ended in a SIGSEGV at exit
ci / bun-nix (pull_request) Successful in 27s
ci / web (pull_request) Successful in 1m1s
ci / docs-site (pull_request) Successful in 1m11s
ci / rust-arm64 (pull_request) Successful in 2m30s
ci / rust (pull_request) Successful in 4m47s
Each gamescope-backed session left a coredump behind. It happened after the
compositor had finished its work — "Primary child shut down!", then the crash —
so the stream itself looked fine and it surfaced only as a steady drip of
coredumps and a non-zero exit from the spawn.

It is a static-destruction-order bug, not a race and not anything gamescope does
wrong at runtime. `g_device` (CVulkanDevice) and `g_output` (VulkanOutput_t) were
plain globals, so glibc ran their destructors from `__run_exit_handlers` once
main() returned. Those destructors call back into the driver —
`~CVulkanCmdBuffer` -> `vk.FreeCommandBuffers`, `~CVulkanTexture` -> `vk.Destroy*`
— but the Vulkan ICD has already been torn down and unloaded by then, so each
call jumps through a function pointer into an unmapped page. The faulting address
equalling the instruction pointer is the signature:

  #0  0x00007fe8fd1d1070 in ?? ()
  #1  CVulkanCmdBuffer::~CVulkanCmdBuffer   at rendervulkan.cpp:1543
  #9  std::vector<unique_ptr<CVulkanCmdBuffer>>::~vector  (g_device+1792)
  #10 CVulkanDevice::~CVulkanDevice         at rendervulkan.hpp:768
  #11 __run_exit_handlers / exit()

Patch 0006 gives both globals storage that is constructed exactly as before but
never destroyed; a union member is destroyed only if the union's destructor says
so, and ours deliberately does not. Nothing needs freeing there — the process is
exiting and the kernel reclaims the device, its command buffers and every GPU
allocation. Both objects are needed: pinning only the device relocated the fault
into ~VulkanOutput_t.

The `.pfhdrN` level deliberately stays at 4. It is a capability tier the host
probes before it spawns, and this patch adds no capability — bumping it would
advertise a tier that does not exist. Per the PKGBUILD's own rule this ships as a
`pkgrel` bump instead.

Verified on an NVIDIA box, all six patches `git am`-ing onto the pinned upstream
commit and then a RELEASE build (the shipped configuration):

  version banner   3.16.25-7-gea635c1+pfhdr4   (marker intact)
  patched, real spawn shape (2752x2064@120 --steam --xwayland-count 1)
                   6/6 exit 0
  distro control, same shape
                   SIGSEGV

Not filed upstream, though it is not punktfunk-specific and would apply as-is.

Unrelated and left alone: `--xwayland-count 0` dies much earlier, in main() at
wlserver.cpp:3215, dereferencing a null `gamescope_xwayland_server_t`. Punktfunk
always spawns with `--xwayland-count 1`, so that path is never taken here.
2026-08-08 19:21:39 +02:00
enricobuehler 9232631299 feat(pf-encode): PyroWave had no encode split — so the one cost this program protects was unmeasurable
Wave-2 PW1's exit criterion, and the instrument it needed.

VAAPI and direct NVENC both log a PUNKTFUNK_PERF submit split. PyroWave did not — which meant the
single encoder the GPU-priority work exists to defend was the one you could not put a number on.
Adds per-frame timing of the synchronous encode (whole `submit`: CSC + encode + fence wait +
packetize, which for this backend IS the encode), summarised every 2 s as mean/p50/p99/max.

p99 rather than mean-only on purpose. The failure patch 0005 describes is a TAIL event — frames
going ~2 ms to 15-18 ms at 95 % game load while the mean barely moves — so a mean-only readout
would report "fine" straight through the thing being measured.

WHAT IT MEASURED — .21, RTX 5070 Ti (610.57.04), GRID 2 benchmark loop saturating the GPU at
54-87 %, PyroWave 1080p, same binary both arms (only CAP_SYS_NICE differs), 30-frame windows with
the warm-up window dropped:

  arm                        p50        p99        worst frame
  default priority (refused) ~2.6 ms    ~6.4 ms    9.5 ms
  REALTIME granted           ~3.2 ms    ~4.4 ms    5.4 ms
  REALTIME granted (repeat)  ~3.35 ms   ~4.8 ms    5.1 ms

p99 down ~30 %, worst frame roughly halved, for ~0.6 ms on the median. For a streaming encoder
that is the right side of the trade — the tail is what becomes a visible hitch.

This CONTRADICTS the patch's only prior datum (RTX 4090 / Windows / WDDM: "did not reduce the
spikes"), so patch 0005's header now records the Linux/NVIDIA result beside it, with an explicit
"do NOT delete this patch on the strength of the WDDM result — the two stacks disagree". Header
prose only; the diff hunks stay byte-identical and `git diff crates/pyrowave-sys/vendor/` is
untouched by this commit.

Caveats recorded rather than buried: the arms were not interleaved and the game load drifted
between them, capture was frame-starved (~2.5 fps) so this is encode latency under contention and
not a full-rate stream, and it is two granted runs against one refused run. The direction held
across all 25 windows.

Also worth knowing for anyone repeating this: `encode_fps` is a VACUOUS metric on this rig. A
headless gamescope with no real content emits ~12 fps, so both arms simply report the capture rate.
Measure latency, not throughput.

Gates green at CI parity.
2026-08-08 18:40:52 +02:00
enricobuehler 0cd946acb5 Merge pull request 'The virtual Steam Deck pad's udev rule names a group that four of six install paths never create' (#123) from worktree-steamdeck-group-and-secret-gaps into main
ci / web (push) Successful in 2m26s
ci / docs-site (push) Successful in 1m49s
ci / bun-nix (push) Successful in 42s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 11s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 29s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 14s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 18s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 19s
ci / rust-arm64 (push) Successful in 4m56s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m32s
ci / rust (push) Successful in 5m51s
docker / builders-arm64cross (push) Successful in 17s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 2m12s
docker / deploy-docs (push) Successful in 6m40s
arch / build-publish (push) Successful in 13m48s
nix / flake (push) Successful in 13m56s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m42s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 28m30s
Reviewed-on: #123
2026-08-08 16:00:47 +00:00
enricobuehler 62a6fa9fac fix(packaging): create the punktfunk group everywhere the udev rule needs it
ci / bun-nix (pull_request) Successful in 29s
ci / docs-site (pull_request) Successful in 1m37s
ci / web (pull_request) Successful in 2m39s
ci / rust-arm64 (pull_request) Successful in 4m10s
ci / rust (pull_request) Successful in 6m50s
nix / flake (pull_request) Failing after 23m28s
60-punktfunk.rules chgrp's the usbip vhci attach/detach nodes to a dedicated
`punktfunk` group (security-review 2026-08-05 M-4: writing `attach` materialises
an arbitrary emulated USB device, so it must not ride on `input`). Four of the
six install paths shipped that rule in 0.25.0 without ever creating the group.
chgrp then failed, the nodes stayed root:root 0644, and the virtual Steam Deck
pad silently never attached — while `usermod -aG punktfunk` failed outright with
"group 'punktfunk' does not exist".

Affected and fixed:

  * arch  — post_upgrade() called only _ensure_update_group, so every box that
            reached 0.25.0 by `pacman -Syu` missed it; post_install was correct.
  * nix   — no users.groups.punktfunk at all, though host.users' own description
            already promised the usbip/vhci pad. Declares it now and adds
            host.users to both groups.
  * bazzite sysext — a group is host state and cannot ride an image, and the
            deb/rpm scriptlets that would create it never run there.
  * steamdeck install.sh/update.sh — handled `input` only. Both now create the
            group and join it: running that script IS the statement "make my
            Deck a host with native pad passthrough".

deb and rpm were correct throughout (one postinst/%post for install + upgrade).

Also on the Deck path: web.env secret hygiene. install.sh's `chmod 600` sat
inside the create-only branch despite a comment calling it "the idempotent belt
for a pre-existing file", and update.sh never touched the config dir at all — so
an install set up once and only updated since kept web.env world-readable
(0644) with the console password and session secret in it. Both scripts now
harden ~/.config/punktfunk to 0700 and web.env to 0600 on every run, and say so
loudly, because a chmod does not un-leak an already-readable secret: the
password still needs rotating.

Both group blocks are `if ensure_group ...` rather than `ensure_group || true`:
a failed groupadd must not fall through to a usermod against a nonexistent
group, which under `set -e` aborted install.sh after the long build and
update.sh before the service restart (verified: exit 6, no restart).

Docs: the group is now documented where people actually look — the per-distro
guides, install.md, steamos-host.md, a new troubleshooting entry for "pad
arrives as an Xbox 360 controller", and the uninstall pages. The 0.25.0 notes
gain the "group does not exist" caveat and turn the password bullet from
"consider rotating" into a real instruction, and CHANGELOG records the known
issue against the breaking change that introduced it.

Verified: bash -n on all four scripts; the arch scriptlet's post_upgrade driven
in a container (creates the group, idempotent on re-run); the ensure_group
helper and both membership branches, including a control that reproduces the
original bug (chgrp to a missing group leaves the node root:root 0644); the
find -perm /0077 probe across 0644/0640/0604/0600/0400 on GNU findutils;
`nix flake check --no-build` (the exact CI gate) and a NixOS eval showing
alice.extraGroups == ["input","punktfunk"]; docs-site build + typecheck.
2026-08-08 17:53:36 +02:00
enricobuehler d402e9b996 Merge pull request 'A compositor pin silently vetoed dedicated game sessions — and a mid-bring-up mode switch was killing GNOME outright' (#121) from worktree-dedicated-session-pin-and-recovery into main
apple / swift (push) Successful in 1m41s
ci / web (push) Successful in 1m13s
ci / docs-site (push) Successful in 1m53s
ci / bun-nix (push) Successful in 34s
ci / rust (push) Successful in 5m31s
ci / rust-arm64 (push) Successful in 5m16s
android / android (push) Successful in 6m30s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 19s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 7s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 39s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 14s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 16s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 25s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 17s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 41s
deb / build-publish-client-arm64 (push) Successful in 1m31s
docker / builders-arm64cross (push) Successful in 9s
docker / deploy-docs (push) Successful in 47s
apple / screenshots (push) Successful in 6m18s
windows-host / package (push) Successful in 11m16s
windows-host / winget-source (push) Skipped
deb / build-publish (push) Successful in 6m0s
windows-host / canary-manifest (push) Successful in 30s
deb / build-publish-host (push) Successful in 6m46s
arch / build-publish (push) Successful in 12m33s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 14m41s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 14m37s
Reviewed-on: #121
2026-08-08 15:38:39 +00:00
enricobuehler f23e0df64c fix(host): a compositor pin silently vetoed dedicated game sessions
ci / bun-nix (pull_request) Successful in 26s
ci / web (pull_request) Successful in 1m4s
apple / swift (pull_request) Successful in 1m46s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m59s
ci / rust-arm64 (pull_request) Successful in 2m31s
android / android (pull_request) Successful in 7m23s
ci / rust (pull_request) Successful in 8m5s
`PUNKTFUNK_COMPOSITOR` is documented as "which backend to drive", but it also
quietly discarded `game_session=dedicated`: `resolve_compositor` gated the
dedicated route on `!overridden` and logged nothing either way. A host whose pin
was a forgotten validation leftover therefore went on displaying "dedicated" in
the console while every launch landed in the desktop instead — for 30 days on the
box that surfaced this, the only evidence being the ABSENCE of a log line.

The pin still wins, since it is the operator's explicit hand-set knob, but it now
says so and names itself.

Two further holes the same triage turned up:

- The pin put its backend into `available()` unconditionally AND skipped
  `apply_session_env`'s `XDG_CURRENT_DESKTOP` scrub, so `pick_compositor` could
  never return `None` — the one place `try_recover_session()` is called from. A
  pinned host whose gnome-shell had segfaulted therefore spent every connect on 8
  doomed `RemoteDesktop.CreateSession: ServiceUnknown` retries while the
  operator's configured `PUNKTFUNK_RECOVER_SESSION_CMD` sat unreachable behind
  that arm. Liveness is now read on both paths, and a pin aimed at a dead session
  takes the recovery exit with an error naming the pin. `needs_live_session()`
  exempts gamescope, which stands its own session up — pinning it on a headless
  box stays supported.

- A mode switch accepted before the pipeline existed was served the long way
  round: build at the now-stale mode, then immediately rebuild at the new one in
  the stream loop. That burns a display create, capture attach and encoder open
  on every such connect, and because the rebuild is deliberately
  create-before-drop it stands up two Mutter `RecordVirtual` monitors ~400 ms
  apart — which segfaults mutter 50.4 inside `meta_monitor_manager_rebuild` and
  takes down the whole desktop session, along with the game just launched into
  it (so the GAME looks like what crashed). Bring-up now adopts the newest queued
  mode and builds once, carrying over the H2/H3 correction ack that the replaced
  rebuild would have sent.

Verified on a real Linux host (192.168.1.21, x86_64): `cargo clippy --workspace
--all-targets --locked -- -D warnings`, `cargo fmt --all --check` and the
punktfunk-host + pf-vdisplay test suites all clean. The gate was proved
non-vacuous against a planted `compile_error!`.
2026-08-08 17:37:02 +02:00
enricobuehler 12f39e1967 fix(steamdeck): the updater stops dirtying the checkout it just pulled into
ci / web (pull_request) Successful in 57s
ci / bun-nix (pull_request) Successful in 25s
ci / rust-arm64 (pull_request) Successful in 2m56s
ci / docs-site (pull_request) Successful in 2m13s
ci / rust (pull_request) Successful in 6m8s
`update.sh --pull` could abort with "Your local changes to the following files
would be overwritten by merge: web/bun.nix" — before a single service was
restarted — and the only way past it was to delete the file by hand.

The updater did it to itself. web/bun.nix is generated (bun2nix, a pure
function of web/bun.lock) but committed, because the Nix build fetches
node_modules only from it. The web step ran `bun install --frozen-lockfile`
without --ignore-scripts, so web's `postinstall` (`bun2nix -o bun.nix`)
rewrote that tracked file on every update. Harmless while the committed file
is in sync — but main carried a stale web/bun.nix from 1db8f763 to b79d90b4,
so any Deck updated in that window had it rewritten to the *correct* content
and has been sitting dirty ever since. The SDK step has always passed
--ignore-scripts, which is why only web/bun.nix ever went dirty.

Two changes, both in install.sh and update.sh:

  * the web install now passes --ignore-scripts and runs `bun run codegen`
    explicitly. web has two install lifecycle scripts and we want exactly one:
    `prepare` IS `bun run codegen` (orval + paraglide + the i18n check) and is
    required, since src/api/gen, src/paraglide and src/routeTree.gen.ts are
    gitignored and `prebuild` only re-runs orval; `postinstall` is the one that
    writes a committed file. Equivalent to the old behaviour minus bun2nix.

  * --pull restores web/bun.nix and sdk/bun.nix before pulling, which unsticks
    the installs already broken out there. Deliberately NOT `git reset --hard`:
    $SRC is the operator's own checkout and may carry real local work, so a
    still-dirty tree now fails with a message that names the files and the way
    out instead of git's raw abort. Discarding these two is provably lossless —
    regenerating them from the lockfiles is exactly what bun2nix does.

CI already gates the drift that made this visible (scripts/ci/check-bun-nix.sh,
ci.yml), so main cannot ship a stale bun.nix again.
2026-08-08 17:15:49 +02:00
enricobuehler 8387e48ac6 docs(pf-capture): the fence wait is already free — PW4 retires into this comment
Wave-2 PW4's outcome. The package proposed moving the producer-fence wait off the PipeWire loop
thread, and was pre-registered to be ABANDONED if the wait turned out to already be free. It is,
on every producer and vendor measured — including the one where implicit sync actually exists.

Steam Deck, RADV VANGOGH, gamescope producer (built in distrobox pf2, run on the host):

  samples=300  mean_us=23  max_us=48  p50=<=100us  p99=<=100us
  signaled=0   no_fence=300  timed_out=0  failed=0

p99 in the first bucket is the plan's own abandonment condition, and the outcome split explains
why: 300 of 300 buffers reported NoFence. Same on both NVIDIA producers (gamescope and Mutter's
virtual output — the exact no-explicit-sync case the comment cites as the reason the wait exists).

So `wait_read_ready` here is one ioctl and a return, not a block. Moving it to the consumer side
would buy nothing measurable and would take on the hazard the package itself names — a slot holding
a not-yet-ready dmabuf, and `repeat_last` re-waiting a fence it already consumed. Not a trade worth
making for 23 microseconds.

The 100 ms budget stays: it guards a producer that DOES fence, which is a real thing even if
nothing in this fleet does it. KWin/AMD is the one combination still unmeasured, and the histogram
from the previous commit is deliberately kept as the way to re-check — run with PUNKTFUNK_PERF=1
and read the p99 bucket.

Comment-only; no behaviour change. Gates green at CI parity.
2026-08-08 16:19:34 +02:00
enricobuehler fb60bf653e feat(pf-capture): instrument the fence wait PW4 wants to move, before moving it
Wave-2 PW4, step one of one-so-far. The package's own first line is "investigation step first
(measure, then decide)", and it is pre-registered to be ABANDONED if the wait's p99 is ~0 — so the
instrument ships before the change, not after.

A per-session histogram of `wait_read_ready`, taken on the PipeWire loop thread, which is exactly
where the wait is expensive: that thread is the compositor's consumer, so time blocked there delays
buffer recycling for the NEXT frame. Logged under PUNKTFUNK_PERF at the same cadence and gate the
encode backends use for their submit splits, so a perf run reads as one instrument: samples, mean,
max, p50/p99 bucket, and the Signaled/NoFence/TimedOut/failed split.

Buckets are coarse on purpose (100us -> 10ms, plus overflow). The decision this feeds is binary —
a p99 in the first bucket means the wait is already free and PW4 becomes a comment correction; a
p99 past 1ms is a real stall against a 16.6ms frame budget. Edges are placed so those two worlds
cannot be confused, and anything past the last edge reports as overflow rather than clamping into
the top bucket ("worse than 10ms" is a distinct finding).

The outcome split sits next to the timings because "the wait is short" and "there is nothing to
wait for" are different results with different consequences, and one data point already shows the
second: on gamescope/NVIDIA the probe reports NoFence, i.e. that producer attaches no implicit
fence at all.

5 tests pin the arithmetic, including that an empty histogram reports "no answer" rather than a
decisive-looking zero — the failure mode that would retire the package on no evidence.

No behaviour change: the wait still happens where it always did. Gates green at CI parity.
2026-08-08 15:47:25 +02:00
enricobuehler 2a1c968a0e Merge pull request 'A gamescope session told every game its display was 60 Hz — and Fedora had no way to install the build that knows better' (#120) from worktree-gamescope-virtual-display into main
apple / swift (push) Successful in 1m37s
ci / web (push) Successful in 1m42s
ci / bun-nix (push) Successful in 29s
ci / docs-site (push) Successful in 1m37s
ci / rust-arm64 (push) Successful in 5m17s
deb / build-publish-client-arm64 (push) Successful in 1m48s
apple / screenshots (push) Successful in 6m27s
android / android (push) Successful in 9m56s
windows-host / package (push) Successful in 10m52s
windows-host / winget-source (push) Skipped
deb / build-publish (push) Successful in 11m20s
arch / build-publish (push) Successful in 16m0s
windows-host / canary-manifest (push) Failing after 32s
deb / build-publish-host (push) Successful in 12m7s
ci / rust (push) Successful in 17m46s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 12s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 10s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 11s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / builders-arm64cross (push) Successful in 23s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m15s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 2m8s
docker / deploy-docs (push) Successful in 33s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 13m51s
nix / flake (push) Failing after 14m39s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m35s
Reviewed-on: #120
2026-08-08 13:41:30 +00:00
enricobuehler b815e00a87 fix(pf-zerocopy): one dmabuf timeout condemned every later capture on the host, forever
Wave-2 PW3.

The raw-dmabuf passthrough has two very different reasons to switch itself off, and they shared one
`AtomicBool`:

  * the encoder repeatedly failed to import what this compositor allocates — unrecoverable, a
    driver fact, and the reason this latch was written (it stops the encode-stall recovery
    rebuilding the same doomed encoder five times and then ending the session, on every connection,
    forever);
  * the dmabuf-only capture offer never negotiated — which can simply mean the compositor was
    mid-restart.

Sharing the flag made the second as permanent as the first. One timeout, and EVERY later session on
the host captured CPU frames until the process was restarted — including sessions against a
completely different compositor and a different node, which had never failed at anything. Nothing
said so; the arm line PW2 added would have shown `cpu` with no explanation.

Now the two causes have the lifetimes they should have, in a `RawDmabufLatch` that owns both:

  * Import failures stay sticky. Unchanged threshold (3 consecutive), unchanged hazard coverage.
  * Negotiation timeouts get a retry budget of 2 — one retry, deliberately small: each failure
    costs a ~10 s stall, so a larger budget is paid by the user in dead air. One retry survives the
    mid-restart transient; a compositor that genuinely never accepts keeps the same identity, so it
    latches on the second try, one extra stall per host lifetime versus the old behaviour.
  * A capture that negotiates credits the budget back, so an evening of reconnects against a
    compositor that failed once cannot accumulate its way into a latch.
  * BOTH are keyed to a capture identity (node id + portal bit). A new node — fresh virtual output,
    compositor restart, the Bazzite Gaming↔Desktop switch — is a genuinely different question and
    earns a fresh dmabuf attempt instead of inheriting a verdict about something else. The SAME
    capture keeps its verdict, which is what preserves the 10 s-stall protection the latch exists
    for.

The session-open line now carries the latch state, so `cpu` is no longer ambiguous between "this
host was never going to do dmabuf" and "something failed earlier and we are still living with the
verdict" — only the second is a bug worth chasing, and only the second is now visible as one.

Atomics rather than a lock because `note_import_ok` is on the per-frame import path; everything
else runs at pipeline build or on failure. The state machine is tested against a local instance
rather than the process-wide static — seven tests covering both lifetimes, the identity clear, the
same-identity hold, the budget credit, and the cause naming.

One honest note on the identity: it is the PipeWire node id, not the "(compositor-id, modifier
list)" pair the design sketched. Node id is what capture actually has at that point, and it changes
on exactly the events that matter here (new virtual output, compositor restart, session switch).
Keying on the modifier list too would need the list before the importer is built, which is the
wrong order.
2026-08-08 15:40:55 +02:00
enricobuehler eb8c943572 feat(packaging): ship punktfunk-gamescope on RPM and apt too
ci / web (pull_request) Successful in 1m4s
ci / bun-nix (pull_request) Successful in 26s
apple / swift (pull_request) Successful in 1m40s
apple / screenshots (pull_request) Skipped
ci / rust (pull_request) Failing after 2m3s
ci / docs-site (pull_request) Successful in 1m41s
ci / rust-arm64 (pull_request) Successful in 2m44s
android / android (pull_request) Successful in 4m50s
nix / flake (pull_request) Failing after 16m12s
Until now the patched gamescope reached exactly four kinds of box: the
Bazzite/Fedora-Atomic sysext, the Arch package, the SteamOS installer and a NixOS
option. Everyone else was told to build gamescope from source. A traditional
Fedora-family box — Nobara, plain Fedora, the HTPCs people actually stream from —
therefore ran stock gamescope by default, which streams SDR, cursorless, and
tells every game its display is 60 Hz. That is not a user error; there was no
package to install.

Both new packages REPACK the binary CI already builds rather than building
gamescope again: it is a ~10-minute meson compile of an unrelated tree, cached
per distro base because the binary is soname-coupled to it. The Arch PKGBUILD
stays the one recipe that builds from source, because that is what makepkg is
for.

- packaging/gamescope/punktfunk-gamescope.spec + build-gamescope-rpm.sh. Version
  is derived from the binary's own banner (3.16.25.pfhdr4) — the only source that
  cannot drift from what is in the package. rpmbuild's automatic ELF Requires are
  what stop an f43 build installing on f44.
- packaging/debian/build-gamescope-deb.sh, same shape, with dpkg-shlibdeps for
  Depends.
- rpm.yml packages and publishes it beside the host RPMs; deb.yml gains a cached
  gamescope build (keyed on packaging/gamescope/** alone) and packages it into the
  existing publish loop. Both legs are best-effort, matching the sysext's existing
  rule: no binary, no package, and the host stays on its current SDR path.

Neither package Provides or Conflicts with gamescope — it installs as
/usr/bin/punktfunk-gamescope and only the sessions the host starts itself resolve
it, so a box's own Game Mode keeps using the distro binary.

Both refuse to package a binary without the +pfhdr marker. That marker is the
host's entire capability probe, so a build that lost the patches would install
fine and then silently stream SDR with no cursor.

Verified: build-gamescope-deb.sh produces an installable .deb from a stand-in
binary (correct version derived from the banner, 0755 tree, control fields) and
exits 1 on an unmarked one. The .spec is not yet exercised — no rpm tooling on
the box I had; CI's Fedora leg is its first run.
2026-08-08 15:36:19 +02:00
enricobuehler 102f550bba feat(host): use the new gamescope capabilities, and say so when the mode is lost
Pass --custom-refresh-rates (patch level 3+) and
--pipewire-composite-external-overlay (level 4+) on both spawn paths, with the
same probe-then-pass shape the HDR and cursor flags already use. A stock
gamescope has neither flag and gets neither, which is exactly today's behaviour.

New knob PUNKTFUNK_GAMESCOPE_REFRESH_RATES=60,90,120 widens the set a session
offers in Steam's in-session display settings. The rate the session actually runs
at is always included, so it can only add options; junk entries are skipped
rather than failing the host, because the worst a typo can cost is the extra
option the operator wanted.

And the part that would have turned a week of field triage into one log line:
warn_if_mode_lost(). --nested-refresh is the ONLY refresh a headless gamescope
has, and it reaches a gamescope-session-plus solely through the GAMESCOPE_BIN
wrapper, which the session script is free to lose — a sessions.d file sourced
with `set -a` can reassign GAMESCOPE_BIN, and one that sets GAMESCOPECMD outright
skips the whole builder. When that happens the stream still runs, still looks
right, and the client's own fps counter still reads the negotiated rate (the
encode loop repeats the held frame), while the game underneath is capped to 60.
Nothing anywhere said so.

It warns rather than refusing, deliberately: verify_managed_spawn_flags refuses
because its retry resolves a different plan, but a relaunch here would hand the
session the same environment and lose the mode the same way, so refusing would
only loop. Fails open on the same rule as the flag check — nothing to compare
against says nothing.

Also corrects the comment above the launch env, which claimed
CUSTOM_REFRESH_RATES "generates the mode the session ADVERTISES … what makes
games see the real refresh". It never did: no upstream gamescope has
--custom-refresh-rates, so gamescope_has_option gated it off and the variable was
inert. That belief is why the real lever went unexamined.

configuration.md gains the new knob and a warning on PUNKTFUNK_MAX_FPS, which
also lowers the refresh the session REPORTS on gamescope — the docs said it does
not cap the stream, which is true of the wire and not of what games are told.

Linux-verified on Ubuntu: cargo check --all-targets, clippy -D warnings, 133
tests (2 new), cargo fmt --check.
2026-08-08 15:35:58 +02:00
enricobuehler 818531a26e feat(gamescope): a headless session now reports its own mode, and the perf overlay reaches the stream
Two new patches on the pinned upstream, and the marker patch moves last so the
banner is stamped after the capabilities it advertises.

0003 — headless: advertise the virtual display's mode and refresh rates.
A headless gamescope is how we give a game a display: we pass the client's exact
mode and the session runs at it. It never told anyone. CHeadlessConnector
returned empty spans from GetModes() and GetValidDynamicRefreshRates() and
reported GAMESCOPE_SCREEN_TYPE_INTERNAL, so update_mode_atoms DELETED the
mode-list atom (no resolution list) and wlserver fell through to a one-entry
refresh list built from g_nOutputRefresh (no refresh list). With --nested-refresh
absent that entry is Init()'s 60 Hz default — which is why a field report on a
1920x1080@120 client saw "gamescope only shows 60hz, and there's no other
option", and why Overwatch capped itself to 60 while the stream ran at 120.
Populate both from the resolved mode, report EXTERNAL, and add
--custom-refresh-rates so the offered set can be widened. gamescope-session-plus
has probed for that flag for years; upstream never had it, so the
CUSTOM_REFRESH_RATES env it plumbs was a no-op everywhere.

0004 — pipewire: optionally composite the external overlay into the capture
stream. That layer is mangoapp: the fps/frametime readout the Deck UI turns on.
paint_pipewire has never referenced it on any version, so a consumer whose only
view of the session is the node sees the overlay it just enabled not appear, with
nothing to configure. Behind --pipewire-composite-external-overlay, off by
default, same argument as the cursor flag. Its commit id joins the repaint test —
the numbers change while the picture behind them is static, exactly the case the
existing test skips.

Verified: the series git-am's cleanly onto the pinned 8c676c39, and both new
functions were extracted verbatim and compiled with -Wall -Wextra under C++23
against stubs, with unit assertions for the parser and the mode/rate publication
(sorting, dedup, the running rate always present, re-entrancy, zero rejected).
A full gamescope build was not run — no box here has its dependency set; CI's
per-Fedora-major leg is the first real compile.
2026-08-08 15:35:37 +02:00
enricobuehler 608baf63be Merge pull request 'Post-sleep sessions still failed on 0.25.0 — the host was holding open the very device its recovery asks PnP to cycle' (#119) from worktree-vdisplay-reap-pnputil into main
apple / swift (push) Successful in 1m40s
ci / web (push) Successful in 1m21s
ci / rust-arm64 (push) Successful in 2m46s
ci / docs-site (push) Successful in 1m20s
ci / bun-nix (push) Successful in 26s
android / android (push) Successful in 6m37s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 12s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 7s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 7s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
apple / screenshots (push) Successful in 6m8s
deb / build-publish-host (push) Successful in 4m27s
deb / build-publish-client-arm64 (push) Successful in 2m0s
deb / build-publish (push) Successful in 5m36s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m1s
docker / builders-arm64cross (push) Successful in 8s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m23s
ci / rust (push) Successful in 9m44s
docker / deploy-docs (push) Successful in 36s
arch / build-publish (push) Successful in 12m2s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 4m14s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 4m14s
windows-host / package (push) Canceled after 11m59s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
Reviewed-on: #119
2026-08-08 13:29:09 +00:00
enricobuehler 767e67caf4 feat(packaging): grant the host CAP_SYS_NICE, without which the GPU-priority lever does nothing
Wave-2 PW1, second half. The companion commit wires `PYROWAVE_QUEUE_PRIORITY` into the Linux
PyroWave device; this is what makes it work on a packaged host.

Measured on .21 (RTX 5070 Ti, NVIDIA 610.43.02), same binary in both arms:

  as packaged (no capability)     every class refused, REALTIME *and* HIGH -> default priority
  same binary, cap_sys_nice+ep    granted REALTIME on the FIRST attempt, no downgrade

RADV behaves the same way. So this is not the RADV-specific "expect one downgrade to HIGH" the
plan predicted — without the capability there is no elevated priority at all, on any vendor, and
the knob is decoration.

Worth being precise about what is being granted, because it is a network-facing daemon.
CAP_SYS_NICE permits raising scheduling priority (nice, ioprio, affinity, RT class) and nothing
else: no filesystem access, no network privilege, no user switching, and it is NOT setuid. The
repo already ships exactly this capability on its gamescope binary for the same reason. Two side
effects that will otherwise confuse someone debugging: a capability-carrying binary is AT_SECURE,
so the loader ignores LD_LIBRARY_PATH/LD_PRELOAD for it (note this box was propped up by exactly
such a shim during the ffmpeg-9 soname break — that workaround would now be silently ignored), and
core dumps are suppressed by default.

Per packaging path, because none of them are the same:

- Arch: a `_grant_sched_capability` in the scriptlet, called from post_install AND post_upgrade —
  a replaced binary is a new inode, so the capability does not survive an upgrade by itself.
- Debian: the same setcap in the postinst `configure` branch.
- RPM: `%caps(cap_sys_nice=ep)` on the binary in `%files`, which is the rpm-native form — rpm then
  applies it on install, restores it on upgrade, and verifies it. A `%post setcap` does none of
  those.
- NixOS: `security.wrappers`, because a store path is read-only and shared and cannot be setcap'd.
  The unit's ExecStart moves to `config.security.wrapperDir` — without that the wrapper exists and
  the service still runs the uncapped store path, which is the whole failure this fixes.
- Steam Deck: setcap in the installer's sudo block. That box needs it most (one small Van Gogh GPU
  shared between the game and the encode). The binary lives under $HOME, so unlike the /etc
  drop-ins it survives a SteamOS A/B update on its own and needs no atomic-keep entry — but it
  does need re-applying after each rebuild, which re-running the installer does.
- Bazzite sysext: at IMAGE BUILD time, before mksquashfs. It cannot be done in the merge hook (a
  merged sysext's /usr is read-only squashfs) and it cannot ride in from the RPM either — rpm keeps
  capabilities in its own header and `rpm2cpio | cpio` carries only the payload, so the staged file
  arrives with none. mksquashfs does record security.capability (only security.selinux is
  excluded), so a setcap on the staging tree is what lands in the image. Needs root/CAP_SETFCAP;
  a plain-user CI build warns and ships without it rather than failing a release over a
  performance lever.

Every one of them is best-effort and cannot fail an install: a box without libcap, or a filesystem
that cannot store capabilities, simply runs at default priority exactly as it does today.

Documented in the same PR — the configuration row now says the packages grant it, and
running-as-a-service gets a section explaining what it is, how to check it (`getcap`), and how to
remove it (`setcap -r`, or just `PYROWAVE_QUEUE_PRIORITY=off`), including the two debugging side
effects.

Verified: the Arch scriptlet grants the capability from a fake package root exactly as pacman
would invoke it, and the resulting binary reaches REALTIME end to end on the RTX 5070 Ti; the RPM
spec's %caps line parses under rpmspec in a Fedora 41 container; the NixOS module parses under
nix-instantiate; all five edited shell scripts pass `bash -n`. No Rust file changed in this
commit, so the CI-parity Rust gates from the companion commit still stand.
2026-08-08 15:24:54 +02:00
enricobuehler 8f9c72877e Merge pull request 'An SDK fix could never reach an installed plugin — the runner now carries it' (#117) from worktree-runner-sdk-reconcile into main
ci / bun-nix (push) Successful in 24s
ci / web (push) Successful in 1m9s
ci / docs-site (push) Successful in 1m16s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 15s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 12s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 12s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 18s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 41s
ci / rust-arm64 (push) Successful in 2m28s
deb / build-publish-client-arm64 (push) Successful in 1m35s
deb / build-publish-host (push) Successful in 4m13s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 2m59s
sdk-publish / publish (push) Successful in 1m52s
docker / builders-arm64cross (push) Successful in 15s
deb / build-publish (push) Successful in 7m16s
ci / rust (push) Successful in 6m52s
arch / build-publish (push) Successful in 7m29s
docker / deploy-docs (push) Failing after 3m54s
windows-host / package (push) Successful in 15m55s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 25s
nix / flake (push) Successful in 14m35s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m52s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 21m11s
Reviewed-on: #117
2026-08-08 12:41:58 +00:00
enricobuehler 8f32976349 Merge pull request 'Library source settings never opened — the drawer still called the old plugin origin' (#118) from worktree-library-settings-origin-split into main
ci / bun-nix (push) Successful in 31s
arch / build-publish (push) Canceled after 1m8s
ci / rust (push) Canceled after 52s
ci / docs-site (push) Canceled after 1m12s
ci / rust-arm64 (push) Canceled after 1m15s
ci / web (push) Canceled after 1m15s
deb / build-publish (push) Canceled after 55s
deb / build-publish-host (push) Canceled after 45s
deb / build-publish-client-arm64 (push) Canceled after 8s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 14s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 19s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 4s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 4s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 17s
docker / builders-arm64cross (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 13s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 26s
windows-host / package (push) Canceled after 1m54s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 53s
Reviewed-on: #118
2026-08-08 12:40:44 +00:00
enricobuehler 2bf571a5ad feat(pf-encode): PyroWave's Linux encode device never asked for the priority its own patch requests
Wave-2 PW1, first half = Wave-1 WP14 step 4, executed as specced.

PyroWave encodes on the same GPU shader cores a game saturates, and that is measured to hurt:
patch 0005's header records `encode_gpu_synchronous` going from ~2 ms to 15-18 ms at 95 % game
load, with the stream frame rate collapsing. NVENC is immune because it has its own ASIC. The
lever for a compute workload is an elevated global-priority QUEUE — a process-priority raise only
reorders submission, not hardware preemption.

The vendored patch requests exactly that. It is gated `if (!inherit_info)`, and only Windows
leaves `inherit_info` null (`pyrowave_create_device_by_compat`, where Granite builds the device
itself). Linux passes its own create-infos into `pyrowave_device_create_info`, Granite's
`get_existing_create_info()` hands them back, `create_device` takes the inherit branch — and the
whole block is skipped. On Linux the knob has never done anything at all. Meanwhile pf-zerocopy's
VkBridge has shipped the identical ladder on Linux for some time and calls it "the actual NVIDIA
compute-preemption lever"; the encoder that needs it most did not have it.

This wires it natively in `open_inner`'s `DeviceHold`:

- The extension probe reuses the `dev_ext_props` already fetched for queue_family_foreign, and
  takes KHR or the EXT alias — the same spelling pf-zerocopy probes, so the two cannot disagree.
- `queue_priority_candidates` is a pure fn with the grammar copied from the C patch: unset →
  realtime, ASCII-lowercased, `off` alone disables, `high` asks for HIGH only, junk falls back to
  the ladder rather than to off. One env var must not mean two things on two platforms — that is
  the documentation trap this package exists to close — so the grammar is unit-tested against the
  patch's, including where they are both deliberately un-clever (neither trims).
- The create ladder is REALTIME → HIGH → no-priority, stepping only on a refusal. A refused class
  can never fail the open, which matters more here than on Windows: this path is reached only by a
  NEGOTIATED PyroWave session, so a hard error is a dead stream, not a fallback to another encoder.

The subtle part is the write-back. `pyrowave_create_device` RETAINS `device_create_info` for the
device's lifetime and Granite reads the chain back. If the ladder ends on the no-priority attempt
while `_queue_ci[0].p_next` still points at the global-priority struct, Granite is handed a chain
the device was not created with. The `None` arm therefore nulls `p_next` before the final create,
and the field's doc says why. The enabled extension deliberately STAYS in the list: it really is
enabled on the device, it just carries no request.

One deviation from the plan, stated because it is a deviation: the ladder also steps down on
`ERROR_INITIALIZATION_FAILED`, not only `ERROR_NOT_PERMITTED_KHR`. The plan and the C patch handle
only the latter; pf-zerocopy's shipped ladder accepts both. Given a hard error here kills a
negotiated session, treating one extra driver-specific refusal as a downgrade is the cheap side of
that asymmetry.

Also corrects the two vendored notes, which claimed a Linux behaviour the gate made impossible,
and records that patch 0005's negative RTX-4090 result is Windows/WDDM and does not transfer to a
different driver stack. Patch hunks are byte-identical (header prose only) and
`git diff crates/pyrowave-sys/vendor/` is PUNKTFUNK-VENDOR.txt alone.

`PYROWAVE_QUEUE_PRIORITY` is now reachable on Linux, so it is documented in the same PR.

MEASURED ON GLASS, and it changes what this package is worth on its own — .21, RTX 5070 Ti,
NVIDIA 610.43.02, same binary in both arms:

  as packaged (no capability)     every class refused, REALTIME *and* HIGH -> default priority
  same binary, cap_sys_nice+ep    granted REALTIME on the FIRST attempt, no downgrade

So the lever is INERT on an unprivileged host, and that is not the RADV-specific downgrade the
plan predicted — on NVIDIA it is a downgrade to nothing at all. The ladder itself is proven good
across all three legs (unset / high / off): a refused class never fails the open, and `off`
enables no extension and logs nothing. It simply has nothing to grant yet.

The privilege needed is CAP_SYS_NICE on the host binary, which is NOT what Wave-1 WP3 ships
(RLIMIT_NICE, PAM limits, CPUWeight — all different things). That grant is a security-posture
change on a network-facing daemon, so it is deliberately NOT in this commit; the warn line now
names the capability so an operator is not left guessing, and the docs row says the setting has no
effect on most hosts today rather than implying it works.

The loaded-GPU encode_us p99 A/B is therefore not run: it needs a GPU-saturating game (hence a
desktop session the box does not currently have) and it is pointless before the capability lands,
since the unprivileged arm has no priority to measure.

NO unit test is possible for the device-create ladder itself — it needs a real Vulkan device. Its
coverage is the clippy pass, the grammar tests, and the on-glass log line. Stated here rather than
left for a reviewer to wonder about.
2026-08-08 14:37:41 +02:00
enricobuehler deef5e4382 fix(console): library source settings 404'd — the drawer still called the old plugin origin
ci / bun-nix (pull_request) Successful in 35s
ci / web (pull_request) Successful in 1m19s
ci / docs-site (pull_request) Successful in 1m18s
ci / rust-arm64 (pull_request) Successful in 3m22s
ci / rust (pull_request) Successful in 4m30s
Opening a library source's settings did nothing, for every library plugin. Confirmed on
`.21` against the running console:

    console origin :47992  /plugin-ui/lutris/__config -> 404
    plugin  origin :47993  /plugin-ui/lutris/__config -> 401

The drawer fetches a RELATIVE `/plugin-ui/<id>/__config`, so it resolves against the
console's own origin — where `middleware/auth.ts` answers 404 for `/plugin-ui/**`
unconditionally and by design. That refusal is the 2026-08-05 review's origin split
(H-3): plugin UIs moved to their own listener, and neither origin may serve the other's
paths. The drawer is the only consumer of `/plugin-ui` that is NOT an iframe — every
other caller builds an absolute URL from `pluginOriginFrom(uiConfig)` — so it was the
one thing the split broke, and nothing failed loudly enough to notice.

The fix is deliberately not to point the drawer at the plugin origin. That needs CORS
plus cross-site cookies, and it would put a plugin-controlled response inside a
credentialed cross-origin fetch — reopening exactly the hole the split closed. What
this drawer needs is DATA, not an embedded UI: `/api/plugin-config/<id>` reads the
plugin's `__config` server-side over loopback and returns the JSON same-origin, so no
plugin markup or script is ever served from the console origin and the per-boot secret
stays on the server, as with the `/plugin-ui` proxy.

`/api/**` is always session-gated (`isPublicPath`), so the new route inherits the gate
and answers 401 as JSON rather than redirecting to /login — which is what a `fetch`
needs and what the old path could never give it. It forwards only GET and PUT, reads
the body BEFORE the stale-credential retry (`readRawBody` drains the stream, so a
retried PUT would have saved `{}` over the operator's config), and passes the plugin's
own body through untouched so a 400's decode issue still reaches the operator.

Verified against the real built server: `/api/plugin-config/lutris` answers 401 — the
route resolves and is gated, and the BFF catch-all at `api/[...]` does not swallow it —
while `/plugin-ui/lutris/__config` still answers 404 on the console origin, i.e. the
split is intact. `/api/v1/status` still reaches the BFF. tsc clean, production build
clean, i18n 633 messages across en+de, biome clean on both touched files (the one
warning in SourceSettings.tsx pre-dates this change).
2026-08-08 14:17:58 +02:00
enricobuehler 9c24569db6 fix(spike): --codec pyrowave encoded PyroWave off a capture negotiated for somebody else
Found while taking PW2's on-glass measurement, and it is what made the measurement possible.

`spike` built its capture request from `OutputFormat::resolve`, the constructor shared with the
GameStream path, which hard-codes `pyrowave: false` ("GameStream never negotiates PyroWave").
On Linux that flag is not cosmetic: `capture_virtual_output` feeds it to `zero_copy_policy` as
`ZeroCopyPolicy::pyrowave_session`, which is what puts the capture on the raw-dmabuf passthrough.
So `--codec pyrowave` opened a PyroWave encoder over a capture negotiated for a different
consumer, and the only way to exercise the real path was the host-global
`PUNKTFUNK_ENCODER=pyrowave` lever.

That lever cannot stand in for the per-session flag, which is the part that matters here: it
resolves the backend to `Pyrowave`, and `linux_zero_copy_is_vaapi_for` returns true for that —
so it ALSO flips `backend_is_vaapi` on. A per-session PyroWave negotiation on an auto/NVENC host,
where `backend_is_vaapi` is false, was therefore unreachable from the CLI — and that is exactly
the configuration whose CPU downgrade logged nothing at all.

The spike now sets the flag from its own codec, the same comparison `session_plan::output_format`
makes for a real session. With it, the before/after on .21 is unambiguous: origin/main logs zero
capture-path lines on that configuration, this branch logs two (the resolved arm, and the named
downgrade with its cause and fix).
2026-08-08 14:08:27 +02:00
enricobuehler 32cc8dd529 fix(runner): an SDK fix could never reach an installed plugin
ci / bun-nix (pull_request) Successful in 21s
ci / web (pull_request) Successful in 1m17s
ci / docs-site (pull_request) Successful in 1m29s
ci / rust-arm64 (pull_request) Successful in 3m5s
ci / rust (pull_request) Successful in 4m31s
nix / flake (pull_request) Failing after 11m38s
Publishing `@punktfunk/host@0.1.3` — the release that lets a library scanner register
`category`, so Lutris and Heroic stay out of the console nav — reached **no existing
install**. Measured on `.21`: the only thing that moved it was deleting `bun.lock` by
hand over ssh. A fix that needs an ssh session is not a fix.

**Why nothing reached it.** Every plugin resolves the SDK from the plugins tree, and
`bun.lock` pins it to an exact version with an integrity hash. Nothing in any
user-facing flow re-resolves that pin: installing a plugin, reinstalling it, and even
updating it to a newer release all leave the SDK alone, because the plugin's `^0.1.x`
range is already satisfied by what is locked. `bun update` does not help either — the
plugins are pinned exactly in the root manifest, so there is no direct dependency to
update through.

**Where the fix belongs.** The runner. It is bundled from this same `sdk/` at the
host's release commit (`packaging/arch/PKGBUILD` builds `src/runner-cli.ts` into the
punktfunk-scripting package), so `SDK_VERSION` is by construction the SDK matching the
host now on disk. A host upgrade is therefore the one moment that can carry an SDK fix
to already-installed plugins, and now it does — before any plugin loads, and with no
operator action at all.

**Why it re-resolves the whole lockfile** rather than pinning the SDK at the root: a
targeted `bun add @punktfunk/host@<v>` does NOT work while plugins declare the SDK in
their own `dependencies` (all six scanners do, though none import it). bun honours
their locked resolution and gives each a private nested copy that then SHADOWS the
root — measured, 5 nested copies, which is how I first "fixed" the box while leaving
every plugin still importing 0.1.2. A lockless resolve hoists one copy for everyone.
Once the plugins drop that spurious dependency this can become the targeted form.

Safety, because this runs unattended at boot on a tree the operator's plugins load
from: plugin versions are pinned exactly in the root manifest so a re-resolve cannot
move them (verified — lutris stays 0.1.0); the lockfile is backed up and restored if
the install fails or fails to deliver; and every failure is logged and swallowed, so a
dependency refresh can never stop working plugins from starting. The no-op path is the
one that runs on every healthy box, so it is tested first: same version, or no SDK at
all, touches nothing and logs nothing.

The SDK is bumped to 0.1.4 because its published content changed. Republishing 0.1.3
is impossible, and letting source drift from a published version is precisely the
defect that produced this whole chain — 0.1.2 was published before it forwarded
`category`, then the source changed underneath it without a bump. `version.test.ts`
fails if `SDK_VERSION` and `package.json` ever disagree.

Verified end to end on `.21` against a tree seeded from the operator's real pre-fix
backup: 0.1.2 → 0.1.3 automatically, one hoisted copy, no nested copies, plugin
versions preserved, and a second run is a silent no-op. SDK 79 tests pass (5 new),
typecheck clean.
2026-08-08 14:07:11 +02:00
enricobuehler fba22c6c64 fix(host/vdisplay): the host no longer vetoes its own wake-from-sleep recovery — control handles close on retire
ci / bun-nix (pull_request) Successful in 21s
ci / docs-site (pull_request) Successful in 1m23s
ci / web (pull_request) Successful in 1m29s
apple / swift (pull_request) Successful in 1m39s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m19s
android / android (pull_request) Successful in 8m16s
ci / rust (pull_request) Successful in 9m59s
The control-device sharing contract was 'bare HANDLE copies, never
closed for the process lifetime': retired handles were deliberately kept
alive because the pinger/linger threads and the capture delivery
closures held raw copies whose soundness depended on no-close. The cost
surfaced in the 2026-08-08 field log: after a wake left the driver
hostless, every adapter reload came back REFUSED (Generic failure) —
and an open control handle is exactly what vetoes the PnP disable (and
can wedge the pnputil restart) the recovery leans on.
reset-pf-vdisplay.ps1 stops the whole host service precisely to get
those handles closed; the in-process recovery could not, because the
process could never close them.

Ownership is now Arc all the way out: ensure_device/device_handle/
control_device_handle hand out Arc<OwnedHandle> clones, every consumer
holds its clone across its IOCTLs (the capture closures each own one —
Arc<OwnedHandle> is Send+Sync, ending the isize smuggling), and
retiring drops only the manager's reference, so the handle CLOSES when
the last in-flight user drains. DeviceSlot::retired is gone. The
recovery path now releases the manager's reference at the first absent
sighting — the 3 s ABSENT_SETTLE doubles as the drain window — and
again before a not-ready-deadline reload, so the PnP cycle finally runs
against a device the host is no longer holding open.

The driver attaches no meaning to the control file closing (host-gone
is the IOCTL-liveness watchdog, EvtFileClose deliberately unhooked), so
the close has no driver-side side effects. Lock order note: RECOVERY →
device is now taken (the release hooks); the forbidden inverse still
never occurs — VdisplayDriver::open never reloads.
2026-08-08 13:37:05 +02:00
enricobuehler 2aa763ce70 feat(pf-capture): a PyroWave session could drop to CPU capture and log nothing at all
Wave-2 PW2 (design/linux-host-performance-wave2-pyrowave.md). Observability only — no
behaviour change to any capture decision — and it lands first because every later package
in the program is measured by an A/B whose "before" is currently unreadable.

The defect: the capture path's CPU-fallback warning was gated on `backend_is_vaapi`, which
reads the HOST-GLOBAL encoder pref. A PyroWave session is negotiated PER SESSION, so on an
NVIDIA/auto host that gate is false — and the session then fell out of every arm of the
negotiation log chain, emitting nothing whatsoever while paying a full-resolution CPU pixel
touch on every frame. A degraded host and a healthy one produced identical logs.

Four sites, matching PW2.1-2.4:

1. The CPU-path warning now asks the per-session question (`consumer_kind`) instead of the
   pref, and names the consumer. Its gate widened to every GPU consumer and excludes only
   the software encoder, whose native input IS CPU frames — an NVENC session silently on
   the CPU path is the same defect, not a different one. `pyrowave_session` deliberately
   outranks `backend_is_vaapi`, because a PyroWave pref flips `backend_is_vaapi` on too
   (`linux_zero_copy_is_vaapi_for`'s `Pyrowave` arm), so testing vaapi first would swallow
   every PyroWave session.

2. The raw-passthrough block in `consume_frame` had four silent exits — no format, an
   SHM/MemFd buffer, no DRM fourcc, a failed `F_DUPFD_CLOEXEC` — each falling out of three
   nested `if`s into the CPU de-pad path. It is now a labeled block that breaks with a named
   `PassthroughFallback`, logged once per distinct reason per session with a running count,
   so a persistent downgrade is distinguishable from a hiccup at renegotiation. `.process`
   runs per frame, so the rate limit is the shippable part and is what the tests pin.

   Note `NoFormat` does NOT fall back — the CPU path needs `ud.format` too and returns — so
   the line says DROPPED for that one. Three of four downgrade; one loses the frame.

3. `force_cpu_for_nvenc_444` told a 4:4:4 PyroWave session it was "on the NVENC path", which
   is false in every particular: the wavelet encoder never touches NVENC, never swscales to
   YUV444P, and what it actually loses is the raw-dmabuf passthrough its design assumes.

4. One INFO line at pipeline build states the resolved arm and consumer
   (`capture pipeline resolved: dmabuf-passthrough → pyrowave`). Nothing stated it before;
   the 2026-08-08 triage reconstructed it from four files, and for the arm that matters most
   there was no detail line to reconstruct it from.

Also: `spike --codec pyrowave`, so a PyroWave capture→encode pass can be driven without a
client. That is the harness the rest of this program measures on, and it did not exist.

Gates on .21 at CI parity: fmt, workspace clippy -D warnings, pf-encode clippy with
nvenc,vulkan-encode,pyrowave and without, workspace tests.
2026-08-08 13:36:51 +02:00
enricobuehler 4b514cc07c Merge pull request 'An OLED palette, and split WHETHER the gamepad UI is offered from WHEN it appears' (#116) from worktree-oled-theme-gamepad-ui-split into main
apple / swift (push) Successful in 1m33s
ci / rust-arm64 (push) Successful in 2m51s
ci / web (push) Successful in 3m13s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m45s
ci / bun-nix (push) Successful in 45s
ci / docs-site (push) Successful in 1m30s
ci / rust (push) Successful in 4m55s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 23s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 40s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 1m3s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m51s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 1m2s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 40s
deb / build-publish-client-arm64 (push) Successful in 2m42s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m17s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 47s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m18s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m17s
deb / build-publish-host (push) Successful in 6m16s
docker / builders-arm64cross (push) Successful in 11s
release / apple (push) Successful in 9m52s
docker / deploy-docs (push) Successful in 36s
android / android (push) Successful in 13m40s
deb / build-publish (push) Successful in 9m17s
arch / build-publish (push) Successful in 14m36s
flatpak / build-publish (push) Successful in 7m20s
apple / screenshots (push) Successful in 6m0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m15s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 21m21s
Reviewed-on: #116
2026-08-08 11:22:34 +00:00
enricobuehler f242b2d2fc fix(host/vdisplay): a refused adapter reload now says WHY, and never targets a phantom devnode
Field log 2026-08-08 (0.25.0, wake from sleep): every session died on
'the adapter devnode could not be reloaded (Generic failure)' — the WMI
catch-all — because the REFUSED branch reported only the Disable
exception and threw away everything that would identify the failure
mode: why the pnputil /restart-device fallback ALSO failed (its exit
code — 3010 'needs a reboot' is its own diagnosis), what state the
devnode was in, and whether the right devnode was even targeted.

That last one is a real trap, not just missing telemetry: Get-PnpDevice
lists not-present PHANTOM devnodes (upgrade/reinstall leftovers), and
Select-Object -First 1 could hand every recovery attempt a phantom —
whose disable and restart both fail exactly like the field log — while
a live node sat unexamined. The selector now prefers present nodes (OK
before problem-state), and a phantom-only state gets a truthful
refusal: no reload can revive a devnode record whose device is gone;
only reinstalling re-creates it.

The REFUSED line now carries devnode counts, the chosen node's PnP
Status + ConfigManager problem code, and the restart exit code, so the
next field log decides between handle-veto, phantom, and problem-state
instead of reading 'Generic failure'. Decode pinned by test.
2026-08-08 13:22:05 +02:00
enricobuehler 27ceab2f6c Merge pull request 'The SDK could not be published at all — bun publish runs prepare, and prepare needs bun2nix' (#115) from worktree-sdk-publish-prepare-hook into main
ci / bun-nix (push) Successful in 43s
ci / web (push) Successful in 1m23s
ci / docs-site (push) Successful in 2m13s
ci / rust (push) Successful in 5m9s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 9s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 7s
ci / rust-arm64 (push) Successful in 5m35s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 7s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 6s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 24s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 29s
deb / build-publish-client-arm64 (push) Successful in 3m17s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 18s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 14s
arch / build-publish (push) Successful in 7m39s
deb / build-publish-host (push) Successful in 6m34s
docker / builders-arm64cross (push) Successful in 7s
sdk-publish / publish (push) Successful in 1m10s
deb / build-publish (push) Successful in 7m36s
docker / deploy-docs (push) Failing after 1m45s
plugin-kit-publish / publish (push) Successful in 1m5s
windows-host / package (push) Successful in 15m11s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 28s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 11m41s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 11m27s
nix / flake (push) Failing after 16m48s
Reviewed-on: #115
2026-08-08 11:04:33 +00:00
enricobuehler 30bd10e301 feat(clients): an OLED palette, and split WHETHER the gamepad UI is offered from WHEN it appears
apple / swift (pull_request) Successful in 1m33s
apple / screenshots (pull_request) Skipped
ci / web (pull_request) Successful in 1m39s
ci / rust-arm64 (pull_request) Successful in 4m7s
android / android (pull_request) Successful in 5m1s
ci / docs-site (pull_request) Successful in 1m47s
ci / bun-nix (pull_request) Successful in 42s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m19s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m23s
ci / rust (pull_request) Successful in 14m32s
Four changes to the client interface, kept together because two of them touch the same rows
and the last is a bug the first would have made far more visible.

A thirteenth `ui_palette` entry, `oled`. The palette table is hand-mirrored in three languages
(`pf-console-ui`'s `library.rs`, `GamepadPalette.swift`, `GamepadPalette.kt`), so it goes into
all three at index 1, directly after the brand default — which keeps `PALETTES[0]` the unknown-id
fallback and keeps the dark-to-pale cycling order intact. What earns the name is arithmetic, not
a darker shade of violet: the ramp's first two stops are literally (0,0,0) and the ground is pure
black, so the shaded half of the field is pixels switched off rather than "very dark grey", and
the calm mix the form screens sit under lifts toward nothing at all. Mean cell luminance is 0.019
against Violet's 0.254. The bright corner keeps a faint indigo-to-violet ember so the backdrop is
still a field with somewhere to go, and that ember carries enough chroma at that luminance
(60 degrees of hue travel across 13 of the 16 cells) to satisfy the existing multi-tone assertion
without adding `oled` to the near-neutral exemption Graphite and Opal take. Each port gains an
`oled_is_actually_black` test that measures the claim — pure-black corner cells, a mean under half
the darkest other field's — rather than restating the table.

A new device key, `gamepad_ui_mode`. The gamepad-UI switch had been deciding two things at once:
whether to offer the controller-optimized interface at all, and that it appears only while a pad
is attached. A user asked for the second half to stop applying. `"connected"` (the default, and
exactly what the lone Bool meant) and `"always"` separate them, surfaced as a "Show it" row
directly under the switch on all five settings surfaces and built only while that switch is on —
a picker whose every option decides nothing is worse than no picker. `GamepadUIEnvironment.isActive`
takes the mode with NO default argument on purpose: a call site that forgot it would silently
strand everyone who chose Always back on "only with a controller", which is the one bug this
parameter exists to make impossible. An unrecognized value waits for a controller, so a mode a
newer client wrote can never trap an older one in a layout it has no way back out of. It stays a
device preference on both platforms, never part of a profile: which interface this device wears
has nothing to do with how a host streams to it.

The smoothness buffer is hidden under Lowest latency, not dimmed. Everywhere else already hid it
— the GTK and WinUI shells, the Apple touch and tvOS screens, the Android touch screen — because
under that intent it names a quantity that does not exist. Two surfaces disagreed: Apple's gamepad
settings screen left the row live and steppable, and the desktop console dimmed it, having no way
to drop a row from a fixed list. That list is now rebuilt each frame through a `row_applies`
filter. The concern about a vanishing row moving everything under the cursor does not apply here
and the new test says why: the row it drops sits directly BELOW the row that drops it, so the only
cursor that can be present when the list shrinks is the one on the intent row, which does not
move. Two latent hazards went with it — `apply_row` had been indexing the row list on the
assumption the cursor is always in range, and nothing re-clamped that cursor when another writer
changed the intent behind the screen's back.

Pale palettes were unreadable on tvOS, reported from the field. `GamepadInk` was never the
problem: it flips correctly for a pale field, it is not platform-gated, and every tvOS gamepad
entry point already published it. The cause is that this app sets `preferredColorScheme` nowhere
and declares no `UIUserInterfaceStyle`, so every SYSTEM-derived colour landing on those screens —
a `.secondary` placeholder, a `.bordered` button's chrome, a NavigationStack title, a material's
frost — resolved against the DEVICE appearance, which the palette cannot reach. On iPhone, iPad
and Mac a great many users sit in Light mode, so under a pale palette those colours came out dark
and the theme looked correct by accident; an Apple TV is Dark essentially always, so every one of
them rendered white on a light field. The mirror image was broken too and had simply never been
reported: a dark palette on a Light-mode iPhone was already drawing dark on dark. The scheme is
now published beside the ink, once, in `GamepadInkModifier`, because the two are halves of one
decision and publishing only the ink silently loses every colour the frameworks draw on the app's
behalf. Two structural amplifiers went with it: `ConsoleGlass` had been scoping the scheme to the
fill inside its `.background {}` on the tvOS and pre-26 branches while the 26 branch put it on the
content, so no console row's own content ever saw it on tvOS; and `LibraryView`'s navigation
chrome and its loading, error and empty states sit above `LibraryCoverflowView` and so were never
inked at all on tvOS and macOS, where that view is presented directly rather than through the
iOS-only `GamepadLibraryScreen` wrapper.

That last one exposed a second tvOS gap worth closing in the same breath: `ui_palette` had no row
in tvOS's ordinary Settings, and the gamepad settings screen that owns it everywhere else needs an
extended-profile controller to open on tvOS. An Apple TV driven by the Siri Remote alone could not
reach the palettes at all, which would now include the OLED one. `SettingsView.tvBody` carries a
Background row.

Verified: pf-console-ui builds, passes `clippy --all-targets -D warnings` and runs 74 tests clean
under linux/amd64 (a Mac `cargo check` of that crate is vacuous — every module is cfg'd to
linux/windows); `cargo fmt --check` clean for it and pf-client-core. Android `:app` runs 80 tests
with 0 failures, including four new `gamepadUiActive` cases and the palette parity table. The
Apple package builds for macOS AND tvOS and its 9 palette/gamepad-UI tests pass — the tvOS
typecheck is possible because the checked-in xcframework already carries a `tvos-arm64` slice. The
tvOS RENDERING fix is compile-verified only; an on-glass Apple TV check under a pale palette is
still owed, and is the one thing here that a build cannot answer.
2026-08-08 12:57:12 +02:00
enricobuehler 1df39d9617 fix(ci/sdk): the SDK could not be published at all — bun publish runs prepare
ci / web (pull_request) Successful in 1m10s
ci / rust-arm64 (pull_request) Successful in 2m28s
ci / docs-site (pull_request) Successful in 1m23s
ci / bun-nix (pull_request) Successful in 26s
ci / rust (pull_request) Successful in 4m45s
nix / flake (pull_request) Successful in 15m6s
`sdk-v0.1.3` failed at the publish step with `bun2nix: command not found`, exit 127.
Nothing was published, so 0.1.3 is still free.

`bun publish` runs the `prepare` lifecycle script, and sdk's `prepare` is
`bun2nix -o bun.nix` — regenerating the nix dependency file. That tool is a
devDependency of the repo, not something the `oven/bun:1` publish container has, and
the workflow's own install is `--ignore-scripts`, so nothing put it on PATH either.

This was latent, not new. `prepare` gained the bun2nix call on 2026-07-27 (1db8f763,
"move the bun packages to bun2nix"), while the last SDK publish was 0.1.2, bumped
2026-07-20. So the hook has been broken for every SDK release since it landed, and
0.1.3 is simply the first one to try. `@punktfunk/plugin-kit` has no `prepare` and was
never affected, which is why kit 0.3.2 published fine in that window and hid this.

The fix is NOT to copy `web/package.json`, which does the same job from `postinstall`.
That is right for web — it is never published — and would be worse here: a published
package's `postinstall` runs in every CONSUMER's install, so every plugin depending on
`@punktfunk/host` would try to run bun2nix and fail. `prepare` is the correct hook for
a published package (it does not run for consumers); it just must not assume a
repo-maintenance tool exists wherever a publish happens.

So the script skips when bun2nix is absent — and ONLY then. A present-but-failing
bun2nix still fails the script, because swallowing that would publish with a silently
stale bun.nix, which is the exact hand-maintained-hash problem 1db8f763 set out to end.
Both directions measured against the same `sh -e` bun and the Gitea runner use:
absent → exit 0, present-and-failing → exit 3.

`bun publish --dry-run` now completes and reports `+ @punktfunk/host@0.1.3`.
2026-08-08 12:55:37 +02:00
enricobuehler e4f8c64b9f Merge pull request 'Library scanners sat in the nav and could not sync local art — and you can now hide one game' (#113) from worktree-plugin-nav-category-and-art into main
audit / bun-audit (plugin-kit) (push) Successful in 19s
apple / swift (push) Successful in 1m38s
audit / bun-audit (sdk) (push) Successful in 48s
audit / pnpm-audit (push) Successful in 11s
audit / docs-site-audit (push) Successful in 1m8s
audit / bun-audit (web) (push) Failing after 1m14s
apple / screenshots (push) Successful in 5m46s
ci / rust-arm64 (push) Successful in 4m32s
audit / license-gate (push) Successful in 5m12s
ci / bun-nix (push) Successful in 38s
arch / build-publish (push) Successful in 8m1s
ci / docs-site (push) Successful in 1m12s
ci / web (push) Successful in 1m28s
android / android (push) Successful in 9m3s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 33s
deb / build-publish-client-arm64 (push) Successful in 1m25s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 27s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 28s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
audit / cargo-audit (push) Failing after 10m5s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 27s
ci / rust (push) Successful in 7m55s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m31s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m22s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 8s
sdk-publish / publish (push) Failing after 31s
docker / builders-arm64cross (push) Successful in 11s
deb / build-publish-host (push) Successful in 4m20s
docker / deploy-docs (push) Successful in 35s
windows-host / package (push) Successful in 16m9s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 31s
deb / build-publish (push) Successful in 12m39s
nix / flake (push) Canceled after 14m7s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 14m17s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 13m18s
Reviewed-on: #113
2026-08-08 10:39:25 +00:00
enricobuehler 690ff7016b Merge pull request 'The config page missed the whole 0.25 env-var wave — jumbo frames and ten other knobs documented' (#114) from worktree-docs-config-page-0250-vars into main
ci / rust (push) Canceled after 11s
ci / rust-arm64 (push) Canceled after 0s
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 15s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 9s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
Reviewed-on: #114
2026-08-08 10:37:14 +00:00
enricobuehler 6cffe29b13 feat(host,console): hide individual library titles
ci / web (pull_request) Successful in 1m13s
ci / docs-site (pull_request) Successful in 1m23s
apple / swift (pull_request) Successful in 1m40s
ci / bun-nix (pull_request) Successful in 21s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m28s
android / android (pull_request) Successful in 4m25s
ci / rust (pull_request) Successful in 6m26s
nix / flake (pull_request) Successful in 15m40s
The library had one visibility control and it was all-or-nothing: turn a SOURCE off
and every one of its games goes. There was no way to drop a single title — a Proton
tool the filter missed, a demo, a game someone doesn't want on the TV — short of
hiding the whole launcher it came from.

**Where the setting lives.** Not on the entry. Only manual custom entries are stored;
a scanner's and a plugin's titles are rebuilt from scratch on every scan and every
reconcile, so a flag written onto one would be erased by the next sync — silently, and
minutes later, which is the worst possible shape for a setting. So `library-hidden.json`
holds the ids, mirroring how `library-scanners.json` holds disabled sources. The id is
stable by construction (D2: a claimed store's entries keep `<store>:<external_id>`
across reconciles), so a hide survives a re-scan, a plugin restart, and a store's
built-in→plugin migration.

**Where it takes effect.** In `all_games`, which is the one place every play surface
already funnels through — the grid on a client, native clients, the GameStream app
list, and launch resolution. Putting it there rather than at each call site is
deliberate: a per-surface filter is a rule someone has to remember, and forgetting one
is precisely the class of bug the `file://` art asymmetry in the previous commit was.
Hiding is curation, not access control — nothing is deleted, and un-hiding is instant.

**The console is the one surface that still sees them**, or a hidden title could never
be brought back. That exception is a TYPE, not a flag: `GET /library` answers
`Vec<GameEntry>` on every lane but the operator's and `Vec<OperatorGameEntry>` on
theirs, so a hidden entry cannot reach a paired streaming client by someone forgetting
a filter — there is no field there to leak. `hidden` is skipped when false, so the
response is byte-identical to today's for a library with nothing hidden.

`PUT /library/hidden/{id}` is operator-only — neither the plugin lane nor a paired cert,
unlike the scanner toggle. A plugin has no business deciding what its operator sees, and
a client must not be able to hide a game on the host it is streaming from. The id is not
validated against the current library on purpose: a title can be legitimately absent at
that moment (launcher closed, plugin mid-sync, drive unmounted), and refusing the
operator's choice in that window is worse than storing an id that matches nothing today.

On the card, the poster dims and a Hidden badge says why — a faded tile with no label
reads as a broken cover. Its controls stay at full contrast and, unlike an ordinary
card's, are not hover-revealed: the un-hide button is the only way out of the state, and
hiding it behind a hover would strand anyone on a touch screen.

Verified on .21 (Linux): 469 host tests pass (5 new), clippy clean under `-D warnings`,
`cargo fmt --all --check` clean. The routing test is the one that earns its keep — every
library id contains a colon and Heroic's contain two, so a router that split on it would
404 the console against ids the host itself produced. Console: tsc clean, production
build clean, i18n 633 messages across en+de, biome clean on the touched files.
2026-08-08 12:33:54 +02:00
enricobuehler 44c87d7ac1 docs(site): configuration page catches up to 0.25 — jumbo frames and seven other missing knobs
ci / web (pull_request) Successful in 1m8s
ci / rust-arm64 (pull_request) Successful in 2m27s
ci / bun-nix (pull_request) Successful in 21s
ci / docs-site (pull_request) Successful in 1m17s
ci / rust (pull_request) Successful in 8m15s
The env-var reference had fallen behind the v0.25.0 CHANGELOG table. Added, with
the semantics taken from the code rather than the changelog one-liners:

- PUNKTFUNK_JUMBO / PUNKTFUNK_WIRE_MTU (Network & discovery), with a note
  explaining the ack-gated mid-session grow, the start-at-1500 behavior, the
  NIC/switch prerequisites, and the sub-1500 shrink direction of WIRE_MTU
- PUNKTFUNK_AUDIO_QUALITY / AUDIO_REDUNDANCY / AUDIO_OUTPUT_MODE — the legacy
  HOST_AUDIO / KEEP_DEFAULT rows are folded into the OUTPUT_MODE row as the
  aliases they now are (follow_default wins when both are set)
- PUNKTFUNK_NO_AUDIO_MINT (Windows minted-endpoint opt-out)
- PUNKTFUNK_PAD_AUDIO / PAD_AUDIO_SLOTS (Gamepads — DualSense speaker+haptics)
- PUNKTFUNK_NVENC_SPLIT_ARBITRATE (Advanced performance tuning)
- PUNKTFUNK_UI_PLUGIN_PORT / PUNKTFUNK_LIBRARY_ART_ROOTS (Auth, API & paths)
- PUNKTFUNK_VAAPI_DEVICE (client-side table)

Verified against the actual read sites (pf-host-config, wire_mtu.rs,
config.rs jumbo_wire_mtu, pad_audio.rs, minted.rs, art.rs, bun-https.mjs);
the page's remaining vars all still exist in code. MDX-compiles clean with GFM.
2026-08-08 12:31:25 +02:00
enricobuehler 975fef2048 fix(host/vdisplay): the ghost-monitor reap can no longer fail in silence
The reap that keeps departed virtual monitors from exhausting the IddCx
monitor-slot budget launched pnputil by BARE NAME — under the LocalSystem
service's PATH that can miss System32, SilentlyContinue swallowed the
miss, and the Rust side logged only when the count was positive: a reap
that removed nothing and a box with no ghosts were byte-identical
(silence). Ghosts then ratcheted up with every sleep cycle until
IOCTL_ADD wedged at 0x80070490 and every session black-screened — and
the wedge self-heal shipped in 0.25.0 retried an ADD behind a reap that
could never remove anything, which is exactly a persistent post-sleep
"no connection" surviving the b6acbd09 probe fix.

Same family and same cure as the adapter-reload path one function down:
resolve pnputil via $env:SystemRoot (a SYSTEM process must not trust
PATH anyway — a planted pnputil.exe would run elevated), pre-seed
$LASTEXITCODE to failure before every launch, and report found AND
removed unconditionally so "no ghosts" and "removed nothing" are
finally distinguishable in a field log. The report parse is split out
and pinned by tests like classify_reload_output.
2026-08-08 12:19:48 +02:00
enricobuehler 9089651406 Merge pull request 'The jitter ring only ever learned from clicks — it now grows on near-misses, un-does refused shrinks, and cashes growth on the click it already paid' (#111) from worktree-audio-jitter-lowwater into main
apple / swift (push) Successful in 1m36s
ci / web (push) Successful in 1m14s
ci / docs-site (push) Successful in 1m20s
ci / bun-nix (push) Successful in 1m59s
ci / rust-arm64 (push) Successful in 2m31s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 12s
ci / rust (push) Failing after 3m6s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 21s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 12s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 5s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 27s
deb / build-publish-client-arm64 (push) Successful in 1m55s
deb / build-publish-host (push) Successful in 4m38s
docker / builders-arm64cross (push) Successful in 9s
deb / build-publish (push) Successful in 5m5s
docker / deploy-docs (push) Successful in 32s
android / android (push) Successful in 10m35s
flatpak / build-publish (push) Successful in 7m16s
release / apple (push) Successful in 10m45s
windows-host / package (push) Successful in 12m21s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 25s
arch / build-publish (push) Successful in 13m42s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m35s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m50s
apple / screenshots (push) Successful in 6m9s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m19s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m15s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 23m55s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 24m41s
Reviewed-on: #111
2026-08-08 10:11:10 +00:00
enricobuehler 2a2427afc8 Merge pull request '"Native resolution" streamed the compositor's points, not the panel's pixels — and the window was never high-DPI either' (#112) from worktree-wayland-native-pixel-density into main
ci / bun-nix (push) Successful in 26s
ci / docs-site (push) Successful in 1m4s
android / android (push) Canceled after 1m19s
ci / web (push) Successful in 1m12s
apple / swift (push) Canceled after 1m24s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 1m31s
ci / rust (push) Canceled after 1m35s
ci / rust-arm64 (push) Canceled after 1m34s
deb / build-publish (push) Canceled after 1m15s
deb / build-publish-host (push) Canceled after 37s
deb / build-publish-client-arm64 (push) Canceled after 25s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 17s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 4s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 4s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 7s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 1s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Canceled after 1m45s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Canceled after 0s
windows / build (aarch64-pc-windows-msvc) (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 0s
Reviewed-on: #112
2026-08-08 10:09:55 +00:00
enricobuehler d237646c66 fix(host,sdk,kit): library scanners sat in the nav, could not sync local art, and so never got their settings
Three symptoms on .21, two defects. Lutris and Heroic appeared in the console sidebar
they explicitly opt out of; Lutris's settings were unreachable from the Library
screen; and Lutris and Steam logged `sync (startup) failed: HostRequestError`.

**The sidebar is a publish gap.** The console is correct — it keeps
`category: "library"` plugins out of the nav (`uiPlugins`, app-shell.tsx) — but the
host reports no category for them at all. `defineLibraryPlugin` sets it and
`sdk/src/ui.ts` forwards it; what SHIPS does not. `@punktfunk/host` was bumped to
0.1.2 on 2026-07-20 and `category` landed 2026-08-05 without a bump, so the registry's
0.1.2 is the pre-category build and every installed scanner registers without one.
Bumps the SDK to 0.1.3 — **inert until it is published**.

Because the field rides the untyped `pf.request` seam so an older host ignores it
rather than rejecting the registration, dropping it is silent by design. `serveUi` now
reads its own directory entry back and warns once when a requested category did not
land, the same way `defineLibraryPlugin` already warns when a store claim did not take.
That is what turns the next occurrence into a log line instead of a bug report.

**The missing settings and the failed sync are ONE defect: a write/read disagreement
about `file://`.** `local_art_bytes` decodes a `file://` value before testing
containment; `validate_art_paths` handed the raw value to `Path::new`, where
`file:///home/u/c.jpg` is a RELATIVE path whose first component is `file:`. It
canonicalized against the cwd, failed, and read as "outside every art root". So the
host refused every cover the kit's own `fileUrl` helper emits — the documented way for
a plugin to publish local art — while the read path would have served those same files.

That the two symptoms share a cause is not obvious and is why this is one commit: the
Library screen's settings control renders only for `origin: "plugin"`, and a source
becomes `plugin` only once it holds a store CLAIM, which is taken during a successful
reconcile. Lutris failed at entry 0 and Steam at entry 3, so neither ever claimed its
store, both stayed `origin: "builtin"`, and neither got a settings button. Heroic
reconciled (its art is http(s)) and has had its settings all along; rom-manager was
never affected because zero entries meant it never applied.

`art_path_is_servable` now decodes first, so both halves of the confinement judge the
same string. Confinement itself is unchanged: an out-of-root path is still refused in
`file://` clothing, which the test asserts alongside the accept case.

Diagnosing this took the HOST's journal, because both surfaces that should have
explained it lied. `HostRequestError` stringified to its bare tag, so the sync engine's
`${e.cause}` logged `HostRequestError` and discarded the method, the path and the
host's own message; it now renders all three, including an object-shaped cause that
used to print `[object Object]`. And the host logged "payload carries a field this lane
may not set" for BOTH refusals in `check_entry_fields`, so a 400 about an art path read
as an auth problem — it now logs the real reason and the entry title.

Verified on .21 (Linux): 463 host tests pass, clippy clean under `-D warnings`,
`cargo fmt --all --check` clean. The new art test fails without the fix and passes with
it. plugin-kit 71 and SDK 72 tests pass, both typecheck clean, biome clean.
2026-08-08 11:43:30 +02:00
enricobuehler 69728b6f4e fix(pf-presenter): "Native resolution" streamed the compositor's POINTS, not the panel's pixels
ci / bun-nix (pull_request) Successful in 23s
ci / web (pull_request) Successful in 1m11s
ci / docs-site (pull_request) Successful in 1m19s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m33s
apple / swift (pull_request) Successful in 1m33s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 3m25s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m22s
android / android (pull_request) Successful in 4m41s
ci / rust (pull_request) Successful in 6m52s
A CachyOS / KDE Plasma 6.7.4 Wayland client with its 2560x1600@165 laptop panel at
150 % scaling negotiated 1706x1066 for "Native resolution" and streamed a visibly
blurry image. Two independent defects, and they stack — which is why forcing the mode
to 2560x1600 by hand did not fully fix it either.

1. `SDL_GetDesktopDisplayMode` reports a mode in SCREEN COORDINATES and hands the
   pixels-per-point ratio back separately as `pixel_density`. We read `m.w`/`m.h` raw.
   KDE advertises that panel as 1707x1067 points with a density of ~1.4997,
   `render_scale::apply` even-floors both odd axes, and 1706x1066 goes on the wire —
   exactly the mode in the reporter's handshake log. Multiplying by the density
   recovers 2560x1600 to the pixel, because SDL derives it as the output's exact
   pixels/points ratio. On X11 and Windows SDL never sets a density and `SDL_video.c`
   normalizes the unset 0.0 to 1.0, so this is inert there: the bug needed a
   compositor doing FRACTIONAL scaling.

2. The SDL window was created without `HIGH_PIXEL_DENSITY`, so the Wayland surface
   stayed at buffer scale 1 — the Vulkan swapchain was built at 1707x1067 and KWin
   upscaled it to the glass. Even a correct 2560x1600 stream was resampled down and
   then back up. The same flaw silently shrank "Match window", which asks the host for
   `size_in_pixels()`. The reporter's `SDL_VIDEO_WAYLAND_SCALE_TO_DISPLAY=1` workaround
   is this same fix applied from outside SDL, which is why it helped.

The surrounding code was already written for pixels != points — the swapchain,
match-window and pointer mapping all read `size_in_pixels()` while window-size
persistence reads logical `size()` — so the flag only makes those two stop being the
same number. `display_scale()` starts reporting 1.5 into a swapchain that is 1.5x
larger, leaving the OSD the size it already was.

Also closes a smaller hole on the way past: only an `Err` from SDL reached the
1920x1080 fallback, so a display that reported a 0x0 mode sent a 0x0 request.

Verified on home-worker-5 (CachyOS — the reporter's distro, real SDL 3.4.14):
`cargo clippy --all-targets -p pf-presenter -- -D warnings` clean and 18/18
pf-presenter tests pass, three of them new and pinned to the field-reported numbers.
2026-08-08 11:15:58 +02:00
enricobuehler 3bb87d260e fix(audio): detect jitter before it is audible, and stop re-probing a depth the link just refused
apple / swift (pull_request) Successful in 1m38s
apple / screenshots (pull_request) Skipped
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 2m1s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m44s
ci / web (pull_request) Successful in 1m38s
android / android (pull_request) Successful in 4m52s
ci / docs-site (pull_request) Successful in 1m33s
ci / rust-arm64 (pull_request) Successful in 4m10s
ci / bun-nix (pull_request) Successful in 28s
ci / rust (pull_request) Successful in 9m26s
The 0.25.0 MacBook field report — audio jitter 'at certain points' — is the
jitter policy learning exclusively from audible failures, on both of its
sides. Growth needed THREE audible underruns before deepening the ring; the
A/V sync loop re-tested a shallower ring every five quiet seconds and paid an
audible starvation event every time it was wrong, forever; and a grown target
was never re-banked — growth raises a threshold, only a re-prime deepens the
ring — so a bunching link rode the knife edge, clicking once per bunching
period with the 'grown' target sitting inert. A ten-minute simulation of the
Wi-Fi power-save pattern (25 ms gaps / 300 ms, −50 ppm skew) measured ~2000
audible events under the shipped policy.

Three mechanisms, in JitterPolicy (Linux/Windows/Android) and mirrored in the
Swift AudioRing:

- NEAR-MISS: a read served with less than one protocol frame left over is the
  same evidence as an underrun, heard by no one. It grows the target one step
  per window, BEFORE the click — waiting for the third audible underrun means
  the user heard two.
- SHRINK PROBES: every shrink is armed for five seconds; answered by an
  underrun or near-miss it is undone on the spot, and a failed sync-driven
  shrink is not retried for a doubling backoff (60 s → 8 min). A probe that
  survives resets the backoff. Continuity outranks sync, now with a memory.
- HOLLOW RE-PRIME: an underrun while the depth AVERAGE runs more than a step
  below the target re-primes immediately, spending the click it already cost
  on the whole refill instead of limping. The average, not the instant, is
  what separates a hollow ring from one late packet, and it is seeded on
  prime so a fresh ring is never spuriously hollow.

Same simulation after: 9 audible events, tail clean but for the clock-skew
re-anchor (a genuinely slow host must re-bank every few minutes; only rate
adaptation would remove that, and no client has it). Neutralising the three
constants reproduces the ~2000 — the convergence tests fail against the old
behaviour.

Verified: 203 punktfunk-core tests, 254 Swift tests (5 skipped), clippy -D
warnings on punktfunk-core --all-features, cargo fmt --all --check.
2026-08-08 11:05:53 +02:00
enricobuehler be57587572 Merge pull request 'The release-rebuild prune called a helper that cannot exist in a release rebuild' (#110) from worktree-arch-rebuild-prune into main
apple / swift (push) Successful in 1m30s
ci / bun-nix (push) Successful in 19s
ci / web (push) Successful in 1m40s
ci / docs-site (push) Successful in 3m12s
ci / rust-arm64 (push) Successful in 6m8s
apple / screenshots (push) Successful in 6m19s
android / android (push) Successful in 10m12s
arch / build-publish (push) Successful in 9m5s
decky / build-publish (push) Successful in 55s
deb / build-publish-client-arm64 (push) Successful in 1m58s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 18s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 15s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 52s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 32s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 28s
deb / build-publish (push) Successful in 9m25s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m28s
deb / build-publish-host (push) Successful in 7m54s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 3m3s
docker / builders-arm64cross (push) Successful in 14s
ci / rust (push) Successful in 17m40s
docker / deploy-docs (push) Failing after 3m46s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 19m52s
Reviewed-on: #110
2026-08-08 09:03:28 +00:00
enricobuehler 8f1c34c6bf fix(ci/arch): the release-rebuild prune called a helper that cannot exist there
apple / swift (pull_request) Successful in 1m32s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Failing after 1m51s
ci / web (pull_request) Successful in 1m25s
ci / bun-nix (pull_request) Successful in 52s
ci / docs-site (pull_request) Successful in 1m53s
android / android (pull_request) Successful in 6m47s
ci / rust (pull_request) Successful in 28m33s
The v0.25.0 rebuild published perfectly — registry has punktfunk-host 0.25.0-2 with
libavcodec.so=63-64, and it resolves on a real ffmpeg-9 box — then failed its last
step with

    prune_release_assets: command not found

`. scripts/ci/gitea-release.sh` sources from the CHECKED-OUT TREE, and a release
rebuild checks out the OLD TAG. So the step could only ever see the helpers that
existed when that tag was cut, and the prune is gated on exactly that path: the
helper was guaranteed absent in the only case that calls it. Adding it to a shared
script made it look available at review time while being unreachable at run time.

Only the workflow file is read from the dispatched ref, so the logic moves there,
inline. Same reasoning documented at both ends, including the corollary worth knowing
before the next rebuild: a PKGBUILD fix made after a tag does NOT reach a rebuild of
that tag either — the packaging comes from the tag too.

Verified by executing the one-liner's exact bytes out of arch.yml under /bin/sh (the
shell Gitea actually uses): keeps the new -2 set and gamescope, drops the superseded
-1 packages and their .sha256 sidecars, leaves other legs' .dmg/.deb untouched. The
`'\n'` survives the shell quoting, which was the part worth proving.

Also drops the now-dead helper from gitea-release.sh rather than leaving a function
no caller can reach, and leaves a warning there against the next one.
2026-08-08 10:57:49 +02:00
enricobuehler 1ef212a78d Merge pull request 'v0.25.0 shipped an Arch host no up-to-date box can install — and the pipeline had no way to tell' (#109) from worktree-arch-ffmpeg9-repackage into main
apple / swift (push) Successful in 1m39s
ci / rust-arm64 (push) Successful in 3m1s
ci / web (push) Successful in 1m32s
ci / bun-nix (push) Successful in 24s
ci / docs-site (push) Successful in 1m16s
android / android (push) Successful in 6m21s
decky / build-publish (push) Successful in 25s
deb / build-publish-client-arm64 (push) Successful in 57s
apple / screenshots (push) Successful in 6m0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 3m7s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 3m27s
arch / build-publish (push) Successful in 10m32s
deb / build-publish-host (push) Successful in 6m13s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 2m9s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m2s
deb / build-publish (push) Successful in 8m38s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 40s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 4m19s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 3m28s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 2m1s
docker / deploy-docs (push) Successful in 32s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 5m38s
docker / builders-arm64cross (push) Successful in 3m38s
ci / rust (push) Successful in 21m50s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 9m36s
Reviewed-on: #109
2026-08-08 08:39:32 +00:00
enricobuehler e044f68500 fix(ci/arch): v0.25.0 shipped a host no Arch box can install, and nothing could tell
apple / swift (pull_request) Successful in 1m38s
apple / screenshots (pull_request) Skipped
android / android (pull_request) Successful in 5m47s
ci / rust-arm64 (pull_request) Successful in 2m35s
ci / web (pull_request) Successful in 1m53s
ci / docs-site (pull_request) Successful in 1m24s
ci / bun-nix (pull_request) Successful in 26s
ci / rust (pull_request) Successful in 7m26s
Arch moved FFmpeg 8 -> 9 (every libav soname +1) hours before the release. PR #108
fixed the real bug — packaging/arch/PKGBUILD now binds punktfunk-host to the sonames
it actually linked, so pacman refuses an upgrade instead of bricking the install — and
re-keyed ci/arch-ci.Dockerfile so the builder would carry FFmpeg 9.

The tag was pushed four minutes later. arch.yml and docker.yml have no `needs:` between
them, and arch.yml deliberately runs no -Syu ("the image's snapshot IS the build
environment"), so the release build pulled the still-FFmpeg-8 `:latest` and published

    punktfunk-host 0.25.0-1  depends: libavcodec.so=62-64, libavutil.so=60-64,
                                      libavfilter.so=11-64, libavdevice.so=62-64,
                                      libswscale.so=9-64

against a world that had moved to 63/61/12/63/10. It fails safely — pacman refuses,
nothing bricks — but it fails broadly: pacman prepares one transaction, so an
unsatisfiable dependency of OURS stopped affected users' entire `pacman -Syu`.

Nothing in the pipeline could have caught it. The existing assert proves the dep is
VERSIONED; it cannot prove the version EXISTS. So two guards, plus the lever to repair
a release that has already shipped:

* Preflight parity — compare the builder's libav `provides` against the live repos and
  `-Syu` the container if they differ. The image is a cache and may lag; on this one
  axis it may not. Syncs into a throwaway --dbpath so the container never sits in the
  partial-upgrade state a bare `pacman -Sy` leaves.

* Publish gate — resolve every built package with `pacman -U --print` against a
  PRISTINE --dbpath. Empty db means "nothing is installed", so every dependency must
  come from the repos exactly as on a user's box. Resolving against the builder's own
  installed set is what would hide this: a stale ffmpeg satisfies a stale bound.
  gamescope stays best-effort (dropped from the upload with a warning, never fatal).

* workflow_dispatch(release_tag, pkgrel) — a published release cannot be repaired by
  re-running its tag: pkgrel would stay 1, which is invisible to a box that already
  recorded the broken build, and the workflow file at the tag can never carry inputs
  added after it. Dispatched from main it takes the WORKFLOW from main and the SOURCE
  from the tag, publishes to the stable repo at a higher pkgrel, and replaces the
  release-page assets (prune_release_assets: upsert replaces by NAME, and a rebuild's
  filenames differ, so the superseded package would otherwise stay one click away).

Verified on a real ffmpeg-9 box (.21, CachyOS) rather than reasoned about: the gate
rejects the published 0.25.0-1 host with the user-visible error verbatim, and passes
client, web, scripting and gamescope — 0 false positives across all five artifacts.
The parity snippet reads today's `provides` correctly (`-Si --dbpath` on an empty db
works; pacman does not wrap fields when piped). Version logic exercised on all four
paths: rebuild -> 0.25.0-2 stable, tag push and canary unchanged, pkgrel=1 refused.

Ships as punktfunk-host 0.25.0-2. README gains the pacman error and what to do about
it; CHANGELOG says plainly that 0.25.0's Arch packages were wrong.
2026-08-08 10:34:11 +02:00
194 changed files with 14386 additions and 1361 deletions
+181 -8
View File
@@ -48,7 +48,26 @@ on:
# `punktfunk-canary` pacman repo as X.Y.Z-0.<run#> (sorts below the eventual X.Y.Z-1),
# tags to `punktfunk` — separate repos, so neither channel can shadow the other.
tags: ['v*']
# REBUILDING A PUBLISHED RELEASE, because on a rolling distro the ground moves under one.
# Arch went FFmpeg 8 -> 9 (every libav soname +1) four minutes before v0.25.0 was tagged, so
# the release's punktfunk-host was linked in a builder image that still had 8 and shipped
# `libavcodec.so=62-64`. No up-to-date Arch box can satisfy that — and pacman prepares the
# whole transaction at once, so it did not merely block our package, it blocked those users'
# entire `pacman -Syu`. The repair is a rebuild of the SAME upstream version at a HIGHER
# pkgrel; nothing else reaches a box that already has the broken build recorded in its db.
# The workflow file at the tag can never carry inputs added after it was tagged, so dispatch
# this from `main`: it checks the tag's SOURCE out, publishes to the STABLE repo, and
# replaces the release-page assets. Same lever for any future "the distro moved" rebuild.
workflow_dispatch:
inputs:
release_tag:
description: 'Rebuild this published release (e.g. v0.25.0) into the stable `punktfunk` repo. Empty = ordinary canary build of the dispatched ref.'
required: false
default: ''
pkgrel:
description: 'pkgrel for that rebuild — MUST be above the published one (2, 3, …); a same-pkgrel republish is invisible to pacman. Ignored without release_tag.'
required: false
default: '2'
env:
REGISTRY: git.unom.io
@@ -94,7 +113,52 @@ jobs:
}
bun --version
# THE BUILDER'S FFmpeg IS PART OF THE PACKAGE CONTRACT, not merely a build detail.
# packaging/arch/PKGBUILD binds punktfunk-host to the exact libav sonames it linked
# (`libavcodec.so=63-64` …), so a builder one FFmpeg major behind Arch emits a package
# that NOBODY can install — and takes the user's whole `pacman -Syu` down with it, since
# pacman prepares the transaction as a unit. That is exactly how v0.25.0 shipped: PR #108
# re-keyed this image for FFmpeg 9, the release tag fired four minutes later, and the job
# still got the FFmpeg-8 `:latest`. The image is a cache and is allowed to lag — but never
# on this one axis. So heal it in-job and shout, instead of building a dead package.
# (Runs BEFORE checkout: a stale image should be repaired before anything depends on it.)
- name: FFmpeg soname parity with today's Arch (heals a stale builder image)
run: |
export LC_ALL=C # `Provides` is a localized field name
# Piped (never a TTY here) pacman prints each field on ONE line, unwrapped.
sonames() { sed -n 's/^Provides *: *//p' | tr ' ' '\n' | grep -E '^lib(av|sw)[a-z]*\.so=' | sort | tr '\n' ' '; }
# A SEPARATE --dbpath: this refreshes only a throwaway view of the repos, so the
# container's own db never enters the partial-upgrade state a bare `pacman -Sy` leaves.
mkdir -p /tmp/pf-archsync
if ! pacman -Sy --dbpath /tmp/pf-archsync --logfile /dev/null >/dev/null 2>&1; then
echo "::warning::could not refresh the Arch db — skipping the FFmpeg parity check"
exit 0
fi
HAVE="$(pacman -Qi ffmpeg | sonames)"
WANT="$(pacman -Si --dbpath /tmp/pf-archsync ffmpeg | sonames)"
echo "builder ffmpeg $(pacman -Q ffmpeg | cut -d' ' -f2): $HAVE"
echo "arch ffmpeg $(pacman -Si --dbpath /tmp/pf-archsync ffmpeg | sed -n 's/^Version *: *//p'): $WANT"
if [ "$HAVE" = "$WANT" ]; then
echo "OK: the builder links the FFmpeg every up-to-date Arch box already has"
exit 0
fi
echo "::warning::arch-ci is stale ACROSS AN FFMPEG SONAME BUMP — upgrading it for this run."
echo "::warning::Bump the 'refreshed:' date in ci/arch-ci.Dockerfile so the IMAGE carries it."
pacman -Syu --noconfirm || true
HAVE="$(pacman -Qi ffmpeg | sonames)"
if [ "$HAVE" != "$WANT" ]; then
echo "::error::builder still links $HAVE while Arch ships $WANT."
echo "::error::Building on would publish a package no Arch box can install."
exit 1
fi
echo "healed: builder now links $HAVE"
- uses: actions/checkout@v4
with:
# A dispatched release rebuild takes its WORKFLOW from the ref you dispatch (the only
# way it can carry inputs the tag predates) and its SOURCE from the tag. Empty string
# = checkout's own default, i.e. the triggering ref, for every other trigger.
ref: ${{ github.event.inputs.release_tag }}
# Cache cargo's git dir too, not just the registry: the workspace includes
# clients/windows, whose windows-reactor/windows deps are git-pinned — cargo must CLONE
@@ -127,12 +191,30 @@ jobs:
# Keep the leading `0.` — it is what sorts a canary BELOW the eventual `X.Y.Z-1` stable
# release. (A pkgrel is digits+dots only, so `0.` is the only prefix available; raising
# it to `1.` would sort canaries ABOVE the release and is not an option.)
env:
RELEASE_TAG: ${{ github.event.inputs.release_tag }}
REBUILD_PKGREL: ${{ github.event.inputs.pkgrel }}
run: |
eval "$(bash scripts/ci/pf-version.sh)" # -> PF_BASE (one minor ahead of latest stable)
case "$GITHUB_REF" in
refs/tags/v*) V="${GITHUB_REF_NAME#v}"; R="1"; REPO=punktfunk ;;
*) V="$PF_BASE"; R="0.$(printf '%08d' "$GITHUB_RUN_NUMBER")"; REPO=punktfunk-canary ;;
esac
if [ -n "${RELEASE_TAG:-}" ]; then
# Dispatched rebuild of a published release (see the workflow_dispatch note at the
# top): same upstream version, higher pkgrel, straight into the stable repo.
# ⚠ Keep that pkgrel SINGLE-DIGIT. Gitea's Arch registry picks the version its .db
# advertises by STRING order (the same trap the canary zero-padding below exists for),
# so "0.25.0-10" sorts BELOW "0.25.0-2" and the rebuild would never be advertised.
V="${RELEASE_TAG#v}"
R="${REBUILD_PKGREL:-2}"
REPO=punktfunk
case "$R" in
''|*[!0-9.]*) echo "::error::pkgrel '$R' is not digits+dots"; exit 1 ;;
1) echo "::error::pkgrel 1 is the published build — a rebuild MUST go up (2, 3, …)"; exit 1 ;;
esac
else
case "$GITHUB_REF" in
refs/tags/v*) V="${GITHUB_REF_NAME#v}"; R="1"; REPO=punktfunk ;;
*) V="$PF_BASE"; R="0.$(printf '%08d' "$GITHUB_RUN_NUMBER")"; REPO=punktfunk-canary ;;
esac
fi
echo "PF_PKGVER=$V" >> "$GITHUB_ENV"
echo "PF_PKGREL=$R" >> "$GITHUB_ENV"
echo "REPO=$REPO" >> "$GITHUB_ENV"
@@ -235,6 +317,63 @@ jobs:
rm -rf dist-gamescope # never cache a failed build (an empty path is not saved)
fi
# THE GATE THIS PIPELINE WAS MISSING. The soname assert above proves the libav dep is
# VERSIONED; it cannot prove the version is one that EXISTS. v0.25.0 passed it and still
# shipped `libavcodec.so=62-64` to a world that had moved to 63 — every affected user got
# "unable to satisfy dependency … required by punktfunk-host", and because pacman prepares
# one transaction, their whole system upgrade stopped there. So ask the only question that
# matters before publishing: would a real, up-to-date Arch box install this?
#
# An empty --dbpath is what makes the answer honest. It means "nothing is installed", so
# pacman must satisfy every dependency FROM THE REPOS exactly as a user's box does. Checking
# against the builder's own installed set instead would let a stale ffmpeg satisfy the stale
# bound and hide the break completely — the very illusion that shipped v0.25.0. `--print`
# resolves and prints; it downloads nothing and installs nothing. Verified against the real
# broken artifact on an ffmpeg-9 box: it reproduces the user-visible failure verbatim.
- name: Assert every package installs on an up-to-date Arch box
run: |
export LC_ALL=C
mkdir -p /tmp/pf-instcheck
if ! pacman -Sy --dbpath /tmp/pf-instcheck --logfile /dev/null >/dev/null 2>&1; then
echo "::error::could not sync the Arch db — cannot prove these packages install"
exit 1
fi
check() { # check FILE -> 0 installable, 1 not (reason on stdout)
pacman -U --print --noconfirm --dbpath /tmp/pf-instcheck --logfile /dev/null "$1" 2>&1
}
ls dist/*.pkg.tar.zst >/dev/null 2>&1 || { echo "::error::nothing in dist/ to check"; exit 1; }
rc=0
for pkg in dist/*.pkg.tar.zst; do
if out="$(check "$pkg")"; then
echo "OK $(basename "$pkg") ($(echo "$out" | wc -l) targets resolve)"
else
rc=1
echo "::error::$(basename "$pkg") CANNOT be installed on an up-to-date Arch box:"
echo "$out" | sed 's/^/ /'
fi
done
# gamescope stays best-effort, exactly as its build step is: a companion that cannot
# install is dropped from the upload with a warning, never a reason to withhold the
# packages this workflow exists to publish. (It is also the one package that can be
# restored from a cache older than the current Arch snapshot.)
for pkg in dist-gamescope/*.pkg.tar.zst; do
[ -e "$pkg" ] || continue
if out="$(check "$pkg")"; then
echo "OK $(basename "$pkg") ($(echo "$out" | wc -l) targets resolve)"
else
echo "::warning::$(basename "$pkg") is not installable on current Arch — NOT publishing it"
echo "$out" | sed 's/^/ /'
rm -f "$pkg"
fi
done
if [ "$rc" != 0 ]; then
echo "::error::refusing to publish: pacman would reject this on a current box, and a"
echo "::error::rejected dependency blocks the user's ENTIRE upgrade, not just punktfunk."
echo "::error::Usual cause: the arch-ci builder image lags Arch across a soname bump —"
echo "::error::bump 'refreshed:' in ci/arch-ci.Dockerfile, let docker.yml republish it, re-run."
exit 1
fi
# NOTE deliberately NO sysext image is built or published here: a prebuilt HOST binary on
# SteamOS breaks on the next A/B soname bump (and /var — where sysexts live — is
# per-partition-set), which is the standing packaging verdict behind the on-device
@@ -262,14 +401,48 @@ jobs:
done
echo "published to $OWNER/arch/$REPO"
# On a real release, also attach the packages to the unified Gitea Release.
- name: Attach packages to the Gitea release (stable tags only)
if: startsWith(gitea.ref, 'refs/tags/v')
# On a real release, also attach the packages to the unified Gitea Release. A dispatched
# rebuild attaches to that SAME release object: the release page is a distribution surface
# too, and leaving the superseded .pkg.tar.zst sitting on it is one click away from handing
# someone the exact break the rebuild exists to fix.
- name: Attach packages to the Gitea release (stable tags + release rebuilds)
if: startsWith(gitea.ref, 'refs/tags/v') || github.event.inputs.release_tag != ''
env:
GITEA_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
RELEASE_TAG: ${{ github.event.inputs.release_tag }}
run: |
. scripts/ci/gitea-release.sh
RID=$(ensure_release "$GITHUB_REF_NAME" "$GITHUB_REF_NAME" auto)
TAG="${RELEASE_TAG:-$GITHUB_REF_NAME}"
RID=$(ensure_release "$TAG" "$TAG" auto)
for pkg in dist/*.pkg.tar.zst; do
upsert_asset "$RID" "$pkg"
done
# A rebuild bumps pkgrel, so its FILENAMES differ from the ones already attached, and
# upsert_asset only replaces by name — the superseded set would survive untouched.
# Drop every pacman asset (and .sha256 sidecar) this upload did not just write.
#
# ⚠⚠ THIS MUST LIVE IN THE WORKFLOW, NOT IN scripts/ci/gitea-release.sh. The sourced
# script comes from the CHECKED-OUT TREE, which on a release rebuild is the OLD TAG —
# so it can only ever offer the helpers that existed when that tag was cut. A helper
# added for this feature is therefore guaranteed ABSENT in the one code path that
# calls it: the first attempt failed with `prune_release_assets: command not found`
# after publishing perfectly. Only the workflow file itself is taken from the ref you
# dispatch. Same reason a packaging fix made after a tag does NOT reach a rebuild of
# that tag — the PKGBUILD is the tag's too.
if [ -n "${RELEASE_TAG:-}" ]; then
KEEP="$(cd dist && printf '%s ' *.pkg.tar.zst)"
# An UNMATCHED glob would come through literally and match nothing in the keep set —
# i.e. "delete every pacman asset on the release". Skip entirely instead.
case "$KEEP" in *'*'*) KEEP="" ;; esac
API="$GITHUB_SERVER_URL/api/v1/repos/$GITHUB_REPOSITORY"
if [ -n "$KEEP" ]; then
curl -fsS "$API/releases/$RID/assets" -H "Authorization: token $GITEA_TOKEN" \
| python3 -c "import json,sys;k=set(sys.argv[1].split());k|={n+'.sha256' for n in k};print('\n'.join('%s %s'%(a['id'],a['name']) for a in json.load(sys.stdin) if a.get('name','').endswith(('.pkg.tar.zst','.pkg.tar.zst.sha256')) and a['name'] not in k))" "$KEEP" \
| while read -r id name; do
[ -n "$id" ] || continue
echo "dropping superseded release asset: $name"
curl -fsS -o /dev/null -X DELETE "$API/releases/$RID/assets/$id" \
-H "Authorization: token $GITEA_TOKEN" || true
done
fi
fi
+96
View File
@@ -320,6 +320,82 @@ jobs:
run: |
VERSION="$VERSION" BUNDLE_FFMPEG=1 bash packaging/debian/build-deb.sh
# punktfunk-gamescope for apt. Same reasoning as the RPM leg in rpm.yml: without a packaged
# build, a Debian/Ubuntu box has no route to the patched gamescope except compiling it, and a
# stock gamescope streams SDR, cursorless, and tells every game its display is 60 Hz.
#
# CACHED on packaging/gamescope/** alone — it depends on nothing else in this repo, so a
# normal push restores a binary instead of spending ~10 minutes on someone else's tree.
- uses: actions/cache@v4
id: gamescope
with:
path: gs-cache
key: punktfunk-gamescope-noble-${{ hashFiles('packaging/gamescope/**') }}
- name: Build the patched gamescope
if: steps.gamescope.outputs.cache-hit != 'true'
# Best-effort, exactly like rpm.yml: the host packages above are the primary delivery and
# work without this binary, so a hiccup building an unrelated tree must not fail the job.
# `build-dep gamescope` resolves the distro's much older packaged version, so it can come up
# short — that is what the `|| true`s absorb, and the marker check downstream is what makes
# a half-built result impossible to ship.
run: |
set -x
apt-get update
apt-get install -y --no-install-recommends meson ninja-build glslc git || true
apt-get build-dep -y gamescope || true
# NOT best-effort. `build-dep gamescope` resolves the distro's much older packaged
# gamescope — where noble has one at all — so it misses what the master tree needs, and
# wayland-protocols is the gap that actually stops the build: meson dies in
# protocol/meson.build with "Neither a subproject directory nor a wayland-protocols.wrap
# file was found", because the tree has no wrap fallback for it. That is what happened on
# the v0.26.0 tag: the step warned and skipped, the job stayed green, and the release
# shipped with no gamescope .deb while the notes said it had one.
apt-get install -y --no-install-recommends wayland-protocols
# The remaining Arch makedepends the older packaged gamescope does not necessarily pull.
# Best-effort: meson falls back or does without, and a name that moves between Ubuntu
# releases should not fail the job. (No libstdc++ static package is needed here — g++
# ships libstdc++.a, which is why only Fedora tripped the sanity check.)
# `build-dep gamescope` gives noble almost nothing — the distro has no comparable package
# — so the tree's real dependency set has to be named outright. One `apt-get` per name on
# purpose: a single transaction aborts wholesale on one unknown package, which would
# install NOTHING and hide the real gap behind a name typo. Best-effort per package, with
# the missing one named; the end-of-job gate below is what actually decides.
for p in libxdamage-dev libxcomposite-dev libxrender-dev libxext-dev libxxf86vm-dev \
libxtst-dev libx11-dev libxres-dev libxmu-dev libxcursor-dev libxi-dev \
libxfixes-dev libxkbcommon-dev libxkbcommon-x11-dev libcap-dev libdrm-dev \
libinput-dev libudev-dev libpipewire-0.3-dev libseat-dev libsdl2-dev \
libluajit-5.1-dev libavif-dev libdecor-0-dev hwdata libglm-dev libbenchmark-dev \
glslang-tools libvulkan-dev libwayland-dev libxcb1-dev libxcb-composite0-dev \
libxcb-xfixes0-dev libxcb-res0-dev libxcb-ewmh-dev libxcb-icccm4-dev \
libxcb-errors-dev libpixman-1-dev libdisplay-info-dev libgbm-dev libegl-dev \
cmake xwayland; do
apt-get install -y --no-install-recommends "$p" \
|| echo "::warning::no such noble package: $p (gamescope may still build without it)"
done
if bash packaging/gamescope/build-punktfunk-gamescope.sh \
--destdir "$PWD/gs-stage" --prefix /usr --jobs "$(nproc)"; then
install -Dm0755 gs-stage/usr/bin/punktfunk-gamescope gs-cache/punktfunk-gamescope
else
# Warn only, even on a tag. The hard gate moved to the END of this job: failing HERE
# skips the host .deb's own publish + release-attach steps below, which is how the
# v0.26.0 release ended up still carrying the pre-CAP_SYS_NICE host .deb from an
# earlier tag commit — a KDE-breaking artifact withheld from replacement by a gate
# meant to protect the release. Never let a missing EXTRA stop a good artifact
# shipping; go red afterwards instead.
echo "::warning::punktfunk-gamescope failed to build on noble — no .deb this run (gamescope sessions stay SDR)"
fi
- name: Build punktfunk-gamescope .deb
# Picked up by the publish loop below, which globs dist/*.deb.
run: |
if [ -x gs-cache/punktfunk-gamescope ] && gs-cache/punktfunk-gamescope --version >/dev/null 2>&1; then
bash packaging/debian/build-gamescope-deb.sh --binary gs-cache/punktfunk-gamescope
else
# Warn only — see the note on the build step. The gate is the last step of this job.
echo "::warning::no usable punktfunk-gamescope — skipping its .deb"
fi
- name: Publish to the Gitea apt registry
env:
TOKEN: ${{ secrets.REGISTRY_TOKEN }}
@@ -347,6 +423,26 @@ jobs:
upsert_asset "$RID" "$DEB"
done
# A release must not be able to make a claim its own CI silently dropped: v0.26.0's notes and
# docs-site said the patched gamescope was apt-installable while no .deb had ever been built,
# because every failure on this path was a `::warning::` that returned 0.
#
# ⚠ LAST step on purpose. The first version of this gate failed at the build step instead, and
# that skipped the host .deb's own publish + attach below — so the release kept the PREVIOUS
# tag commit's host .deb, which still carried the CAP_SYS_NICE postinst that breaks KDE. A
# gate protecting the release withheld the fix for it. Everything good ships first; the job
# goes red afterwards.
- name: A stable tag must ship the gamescope .deb
if: startsWith(gitea.ref, 'refs/tags/v')
run: |
shopt -s nullglob
built=(dist/punktfunk-gamescope_*.deb)
if [ ${#built[@]} -eq 0 ]; then
echo "::error::no punktfunk-gamescope .deb was built — a stable tag must not ship without it (the release notes and docs-site say it is apt-installable). Everything else in this job published normally; see the gamescope build step above for the meson error."
exit 1
fi
echo "gamescope .deb present: ${built[*]}"
# ---------------------------------------------------------------------------------------------
# The aarch64 CLIENT .deb. Cross-compiled on the ordinary amd64 runner in the
# punktfunk-rust-ci-arm64cross image (the rust-ci toolchain + an arm64 multiarch sysroot — see
+99
View File
@@ -206,13 +206,89 @@ jobs:
dnf -y install dnf-plugins-core meson ninja-build glslc || true
dnf builddep -y gamescope || true
dnf -y install xorg-x11-server-Xwayland-devel || true
# NOT best-effort: build-punktfunk-gamescope.sh appends `-static-libstdc++` to LDFLAGS
# (so the binary still starts on SteamOS's older libstdc++ — see its comment), and
# without the static library meson's very FIRST sanity check dies with
# "cannot find -lstdc++ / have you installed the static version", so nothing builds at
# all. That is what happened on the v0.26.0 tag: both Fedora bases warned and skipped,
# the job stayed green, and the release shipped with no gamescope RPM while the notes
# said it had one. A rename here must be LOUD, hence no `|| true`.
dnf -y install libstdc++-static
# The rest of the Arch package's makedepends that Fedora's older packaged gamescope does
# not necessarily pull. Best-effort: unlike the static runtime, meson finds fallbacks or
# does without, and a name that moves between Fedora releases should not fail the job.
dnf -y install wayland-protocols-devel glm-devel cmake libXcursor-devel || true
if bash packaging/gamescope/build-punktfunk-gamescope.sh \
--destdir "$PWD/gs-stage" --prefix /usr --jobs "$(nproc)"; then
install -Dm0755 gs-stage/usr/bin/punktfunk-gamescope gs-cache/punktfunk-gamescope
else
# Warn only, even on a tag — the hard gate is the LAST step of this job. Failing here
# would skip the sysext build, the sysext feed, AND the release attach below, so a
# missing gamescope would also withhold the punktfunk RPMs and the .raw images that
# built perfectly well. deb.yml learned that the expensive way on v0.26.0.
echo "::warning::punktfunk-gamescope failed to build for f${{ matrix.fedver }} — the sysext ships without it (gamescope sessions stay SDR)"
fi
# The same binary, as an ordinary RPM. The sysext below is the Atomic/Bazzite delivery; this
# is the one a traditional Fedora-family box (Nobara, plain Fedora) can actually install —
# until it existed those users had no packaged route to the patched build at all, and a stock
# gamescope tells every game its display is 60 Hz whatever the client negotiated.
#
# Same best-effort rule as the build above: no binary, no package, and the host stays on its
# existing SDR/host-composited path. The spec re-checks the +pfhdr marker itself.
- name: Package punktfunk-gamescope as an RPM
run: |
if [ -x gs-cache/punktfunk-gamescope ] && gs-cache/punktfunk-gamescope --version >/dev/null 2>&1; then
bash packaging/gamescope/build-gamescope-rpm.sh \
--binary gs-cache/punktfunk-gamescope \
--release "$PF_RELEASE"
else
# Warn only — see the note on the build step. The gate is the last step of this job.
echo "::warning::no usable punktfunk-gamescope for f${{ matrix.fedver }} — skipping its RPM"
fi
# A SECOND signing pass, for this package only. The main "Sign RPMs" step ran back at build
# time, long before this RPM existed — the gamescope build sits behind its own ~10-minute
# cache and deliberately runs after the host RPMs are already published. So every
# punktfunk-gamescope RPM went to the registry UNSIGNED, and the repo file we tell users to
# install carries gpgcheck=1: `dnf install punktfunk-gamescope` failed with "The package is
# not signed" on every Fedora and Nobara box. The package was in the channel the whole time
# and could not be installed from it — which is worse than absent, because the release notes
# and the docs-site both say it is there.
#
# Same fail-closed rule as the first pass: sign-rpms.sh hard-fails on refs/tags/v* if the org
# secret is missing, rather than republishing something a user's dnf will reject.
- name: Sign punktfunk-gamescope
env:
RPM_GPG_PRIVATE_KEY: ${{ secrets.RPM_GPG_PRIVATE_KEY }}
RPM_GPG_PASSPHRASE: ${{ secrets.RPM_GPG_PASSPHRASE }}
run: |
shopt -s nullglob
rpms=(dist/punktfunk-gamescope-*.rpm)
# No RPM here is the best-effort skip above, already warned about — not a signing failure.
if [ "${#rpms[@]}" -eq 0 ]; then
echo "no punktfunk-gamescope RPM to sign (see the packaging step above)"
exit 0
fi
bash packaging/rpm/sign-rpms.sh "${rpms[@]}"
- name: Publish punktfunk-gamescope to the Gitea RPM registry
env:
TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
shopt -s nullglob
for rpm in dist/punktfunk-gamescope-*.rpm; do
case "$rpm" in *debuginfo*|*debugsource*) continue;; esac
NAME=$(rpm -qp --qf '%{NAME}' "$rpm" 2>/dev/null)
VR=$(rpm -qp --qf '%{VERSION}-%{RELEASE}' "$rpm" 2>/dev/null)
ARCH=$(rpm -qp --qf '%{ARCH}' "$rpm" 2>/dev/null)
echo "uploading $rpm"
curl -fsS -o /dev/null --user "enricobuehler:$TOKEN" -X DELETE \
"https://$REGISTRY/api/packages/$OWNER/rpm/$GROUP/package/$NAME/$VR/$ARCH" || true
curl -fsS --user "enricobuehler:$TOKEN" --upload-file "$rpm" \
"https://$REGISTRY/api/packages/$OWNER/rpm/$GROUP/upload"
done
# The no-layering Bazzite path: wrap the just-built host + web RPMs into a systemd-sysext
# image and publish it to the per-Fedora-major feed (punktfunk-sysext/f43[-canary], …) that
# `punktfunk-sysext install|update` reads. Same RPMs, same channels — just no rpm-ostree.
@@ -276,3 +352,26 @@ jobs:
for raw in dist-sysext/*.raw; do
upsert_asset "$RID" "$raw" "$(basename "$raw" .raw).f${{ matrix.fedver }}.raw"
done
# A release must not be able to make a claim its own CI silently dropped — v0.26.0's notes
# said the patched gamescope was dnf-installable while both Fedora bases had skipped it on a
# `::warning::` (missing libstdc++-static, which the -static-libstdc++ link needs).
#
# ⚠ LAST step on purpose, matching deb.yml: failing at the build step instead would skip the
# sysext image, the feed publish AND the attach above, withholding the punktfunk RPMs and
# .raw images that built perfectly well. Everything good ships first; the job goes red after.
- name: A stable tag must ship the gamescope RPM
if: startsWith(gitea.ref, 'refs/tags/v')
run: |
shopt -s nullglob
built=(dist/punktfunk-gamescope-*.rpm)
keep=()
for r in "${built[@]}"; do
case "$r" in *debuginfo*|*debugsource*) continue;; esac
keep+=("$r")
done
if [ ${#keep[@]} -eq 0 ]; then
echo "::error::no punktfunk-gamescope RPM was built for f${{ matrix.fedver }} — a stable tag must not ship without it (the release notes and docs-site say it is installable). Everything else in this job published normally; see the gamescope build step above for the meson error."
exit 1
fi
echo "gamescope RPM present: ${keep[*]}"
+452
View File
@@ -12,6 +12,426 @@ with the version table of the release you are moving to, then read **Breaking ch
---
## v0.26.0
52 commits since v0.25.0.
### Versions
| | v0.25.0 | v0.26.0 | Notes |
|---|---|---|---|
| Wire protocol | 2 | **2** | unchanged |
| C ABI | 17 | **17** | unchanged — no symbol added, removed or changed |
| Workspace crate dirs | 26 | **26** | unchanged (40 workspace members) |
| Virtual-display driver protocol | 6 | **6** | unchanged (minimum accepted still 3) |
| Windows virtual-gamepad channel | 3 | **3** | unchanged |
| Plugin index schema | 1 | **1** | unchanged |
| `api/openapi.json` | 0.24.0 | **0.25.0** | tracks API edits, lags one release by convention |
| gamescope patch level (`+pfhdrN`) | 2 | **4** | 3 patches → 6; `pkgrel` 1 → 2 |
| `@punktfunk/host` (SDK) | 0.1.2 | **0.1.4** | |
| `@punktfunk/plugin-kit` | 0.3.2 | **0.4.0** | the `plugin` launch kind |
`crates/pf-driver-proto` is byte-for-byte identical to v0.25.0 and to v0.24.0 — if you ship the
virtual-display driver or the gamepad channel, the last two releases have not touched you.
### ⚠ Breaking changes
**None.** This is a fixes release. Every embedder, packager and plugin that works against v0.25.0
works against v0.26.0 unchanged. Two behaviour changes are worth knowing about anyway, because both
make a client advertise *less* than it used to — see **Capability advertisement** below.
### Capability advertisement
- **`VIDEO_CAP_444` is now probed, not asserted.** It rode the "Full chroma" setting alone. That was
safe while a software HEVC decoder sat underneath it; M8 removed one (there is no permissively
licensed HEVC CPU decoder, so `software_decodable_codecs()` is `H264|AV1`). The host grants 4:4:4
on HEVC **only** and answers the resolved chroma in the `Welcome` *before* the client builds a
decoder — so on a device with no 4:4:4 decode the toggle did not cost crispness, it cost the whole
codec: the Vulkan rung refuses the shape at construction, VAAPI refuses it too, there is no CPU
rung, and the session reconnects on H.264. No AMD silicon has HEVC 4:4:4 decode, so every Steam
Deck with that switch on lost HEVC. Per-profile and default-off, which is why it read as
intermittent.
Now gated on `hevc_444_hardware_decodable`, which asks the driver through the same code the rung
uses at construction (`VkH265Decoder::probe_stream_support`). **Both depths are required**, not
either: with HDR the host may resolve 4:4:4 10-bit, and a device offering `YUV444_8` but not
`YUV444_10` lands in the same hole. Answering from the Vulkan rung alone is exact rather than
approximate — it is the only rung in this build that implements 4:4:4 at all
(`pf_vaadec::profile_for` errors on `chroma_format_idc 3`, pf-dxvadec refuses anything but 4:2:0,
the CPU rung is 8-bit 4:2:0).
⚠ Deliberately **not** extended to `VIDEO_CAP_10BIT`/HDR: all three rungs implement 10-bit 4:2:0,
so a Vulkan-only probe there would withdraw HDR from boxes whose VAAPI/DXVA rung decodes it
perfectly — a regression against a case never observed.
The bit arithmetic moved into `video::video_caps_for` so the part that was wrong is testable
without a GPU, a host or a `Hello`; the test is verified non-vacuous against the planted defect.
### Host and client environment variables
Four new, one clarified. Verified new by `git grep` at the v0.25.0 tag, not assumed —
`PUNKTFUNK_JUMBO`, `PUNKTFUNK_WIRE_MTU`, `PUNKTFUNK_STREAMED_AU`, `PUNKTFUNK_LIBRARY_ART_ROOTS`,
`PUNKTFUNK_RECOVER_SESSION_CMD`, `PUNKTFUNK_GAMESCOPE_SDR_NITS`, `PUNKTFUNK_MAX_FPS` and
`PUNKTFUNK_ON_CONNECT_CMD` all already existed.
- **`PUNKTFUNK_OVERLAY_MASK`** *(new, client)* — controls the Steam-overlay input mask below.
- **`PUNKTFUNK_PYROWAVE_CHUNK_KIB`** *(new)* and **`PUNKTFUNK_PYROWAVE_STREAMED_AU`** *(new)*
PyroWave AU chunking and the streamed-AU path.
- **`PYROWAVE_QUEUE_PRIORITY`** *(existed, but was inert on Linux — see below)* — grammar: unset →
realtime, ASCII-lowercased, `off` alone disables, `high` asks for HIGH only, junk falls back to
the ladder rather than to off. ⚠ **One env var must not mean two things on two platforms**, so
the Rust grammar is unit-tested against the C patch's, including where both are deliberately
un-clever (neither trims).
- **`PUNKTFUNK_GAMESCOPE_REFRESH_RATES=60,90,120`** *(new)* — widens the set a gamescope session
offers in Steam's in-session display settings. The rate the session actually runs at is always
included, so it can only add options; junk entries are skipped rather than failing the host.
Requires gamescope patch level 3+.
- **`PUNKTFUNK_COMPOSITOR`** *(behaviour clarified, not changed)* — documented as "which backend to
drive", it also silently discarded `game_session=dedicated`: `resolve_compositor` gated the
dedicated route on `!overridden` and logged nothing either way. The pin still wins — it is the
operator's explicit knob — but it now says so and names itself. Two further holes closed with it:
the pin put its backend into `available()` unconditionally *and* skipped `apply_session_env`'s
`XDG_CURRENT_DESKTOP` scrub, so `pick_compositor` could never return `None` — the one call site of
`try_recover_session()`, which left `PUNKTFUNK_RECOVER_SESSION_CMD` unreachable behind that arm.
Liveness is now read on both paths. `needs_live_session()` exempts gamescope, which stands up its
own session, so pinning it on a headless box stays supported.
### Client settings keys
All additive; an older client ignores what it does not know, and a newer value can never trap an
older client.
- **`gamepad_ui_mode`** — `"connected"` (default, and exactly what the previous lone Bool meant) or
`"always"`. Splits *whether* the controller UI is offered from *when* it appears.
`GamepadUIEnvironment.isActive` takes the mode with **no default argument** on purpose: a call
site that forgot it would silently strand everyone who chose Always. An unrecognized value waits
for a controller.
- **`ui_palette`** gains `oled` at **index 1**, directly after the brand default — keeping
`PALETTES[0]` the unknown-id fallback and the dark-to-pale cycling order intact. Hand-mirrored in
three languages (`pf-console-ui`'s `library.rs`, `GamepadPalette.swift`, `GamepadPalette.kt`); each
port carries an `oled_is_actually_black` test that measures the claim (mean cell luminance 0.019
against Violet's 0.254) rather than restating the table.
- **`library-hidden.json`** — per-title hide list, mirroring how `library-scanners.json` holds
disabled sources. Deliberately **not** stored on the entry: a scanner's and a plugin's titles are
rebuilt from scratch on every scan and reconcile, so a flag written onto one would be erased
minutes later. Applied in `all_games`, the single funnel every play surface already goes through
(client grid, native clients, the GameStream app list, launch resolution).
### gamescope patches
Three → six, and the marker patch moves last so the banner is stamped after the capabilities it
advertises.
- **0003 — headless: advertise the virtual display's mode and refresh rates.** `CHeadlessConnector`
returned empty spans from `GetModes()` and `GetValidDynamicRefreshRates()` and reported
`GAMESCOPE_SCREEN_TYPE_INTERNAL`, so `update_mode_atoms` **deleted** the mode-list atom and
wlserver fell through to a one-entry refresh list built from `g_nOutputRefresh` — which, with
`--nested-refresh` absent, is `Init()`'s 60 Hz default. That is why a 1920x1080@120 client saw
"gamescope only shows 60hz" and Overwatch capped itself to 60 while the stream ran at 120. Now
populates both from the resolved mode, reports `EXTERNAL`, and adds `--custom-refresh-rates`.
gamescope-session-plus has probed for that flag for years; upstream never had it, so the
`CUSTOM_REFRESH_RATES` env it plumbs was a no-op everywhere.
- **0004 — pipewire: optionally composite the external overlay into the capture stream.** That layer
is mangoapp. `paint_pipewire` has never referenced it on any version. Behind
`--pipewire-composite-external-overlay`, off by default.
- **0006 — never destroy the Vulkan device or output.** `g_device` (`CVulkanDevice`) and `g_output`
(`VulkanOutput_t`) were plain globals, so glibc ran their destructors from `__run_exit_handlers`
once `main()` returned — calling back into an ICD that had already been torn down and unloaded.
Faulting address equalling the instruction pointer is the signature. Reproducible with
`gamescope --backend headless -W 1280 -H 720 -r 60 --xwayland-count 1 -- true` (exit 139, every
time). Both globals get storage constructed exactly as before but never destroyed; pinning only
the device relocated the fault into `~VulkanOutput_t`, hence a shared `CNoDestroy<T>`.
**`+pfhdrN` deliberately does not move for 0006.** The marker is a capability tier the host
probes via `gamescope_patch_level()` *before* it spawns; this patch adds no capability, so bumping
it would advertise a tier that does not exist. Ships as a `pkgrel` bump instead.
⚠ gamescope CI legs are best-effort — a broken patch is a **missing package**, not a red run.
### Virtual-display handle ownership (Windows)
The control-device sharing contract was "bare `HANDLE` copies, never closed for the process
lifetime": retired handles were kept alive because pinger/linger threads and capture closures held
raw copies whose soundness depended on no-close. An open control handle is exactly what vetoes the
PnP disable — and can wedge the `pnputil` restart — that wake-from-sleep recovery leans on, so every
post-wake adapter reload came back REFUSED. `reset-pf-vdisplay.ps1` stops the whole host service
precisely to get those handles closed; the in-process recovery could not.
Ownership is now `Arc` all the way out: `ensure_device` / `device_handle` / `control_device_handle`
hand out `Arc<OwnedHandle>` clones, every consumer holds its clone across its IOCTLs (ending the
`isize` smuggling — `Arc<OwnedHandle>` is `Send + Sync`), and retiring drops only the manager's
reference. `DeviceSlot::retired` is gone.
**Nothing may store a bare control `HANDLE` again.** The whole fix is that the handle closes when
the last in-flight user drains.
### Presenter — points are not pixels
`SDL_GetDesktopDisplayMode` reports a mode in **screen coordinates** and hands the pixels-per-point
ratio back separately as `pixel_density`; `m.w`/`m.h` were read raw. KDE advertises a 2560x1600 panel
at 150 % as 1707x1067 points with a density of ~1.4997, `render_scale::apply` even-floors both odd
axes, and 1706x1066 went on the wire. Multiplying by the density recovers 2560x1600 to the pixel.
⚠ Inert on X11 and Windows: SDL never sets a density there and `SDL_video.c` normalizes the unset
0.0 to 1.0. **This bug needed a compositor doing fractional scaling.**
Second, independent defect: the SDL window was created without `HIGH_PIXEL_DENSITY`, so the Wayland
surface stayed at buffer scale 1 and the swapchain was built at 1707x1067 for KWin to upscale. That
one also silently shrank "Match window", which asks the host for `size_in_pixels()`.
### Apple audio session
`micEnabled` and `echoCancel` both default to `true`, so the **default** iOS session is
`.playAndRecord` — and that branch set `.defaultToSpeaker`. That option is an output **override**,
not a preference, and it outranks an A2DP route. ⚠ **Wired headphones beat it, Bluetooth does not**,
so testing with a cable returns the wrong answer — which is what the comment sitting on it asserted.
Now solved against the route actually given: after activation, if the current output is
`.builtInReceiver`, override to speaker; anything external (Bluetooth, wired, CarPlay, AirPlay) is
left strictly alone. The override is a property of the current route — iOS drops it on every route
change, which is what lets a newly-connected headset win — so it is re-applied per route via an
observer, registered only for `.playAndRecord`, removed in `stop()` before deactivate, `deinit` as
backstop. Without it, dropping Bluetooth mid-stream lands on the earpiece.
⚠ Deliberately **not** adding `.allowBluetooth`: it would make a headset's mic usable but drag the
whole route onto HFP/SCO and collapse game audio to narrowband.
### Audio jitter policy
`JitterPolicy` (`punktfunk-core/src/audio.rs`, used by Linux/Windows/Android) and its mirror in
Swift `AudioRing`. The policy learned exclusively from audible failures on both sides: growth needed
**three** audible underruns; the A/V sync loop re-tested a shallower ring every five quiet seconds
and paid an audible starvation event every time it was wrong, forever; and a grown target was never
re-banked (growth raises a threshold — only a re-prime deepens the ring), so a bunching link rode
the knife edge with the "grown" target sitting inert.
Three mechanisms: **near-miss** (a read served with less than one protocol frame left over is the
same evidence as an underrun, heard by no one — grows one step per window, *before* the click);
**shrink probes** (every shrink armed for 5 s, undone on the spot if answered by an underrun or
near-miss, with a doubling backoff 60 s → 8 min on a failed sync-driven shrink; a surviving probe
resets it); **hollow re-prime** (an underrun while the depth *average* runs more than a step below
target re-primes immediately — the average, not the instant, separates a hollow ring from one late
packet, and it is seeded on prime so a fresh ring is never spuriously hollow).
Measured on a ten-minute simulation of the Wi-Fi power-save pattern (25 ms gaps / 300 ms, 50 ppm
skew): **~2000 audible events → 9.**
### Plugins, SDK and the runner
- **`category` never shipped.** The console correctly keeps `category: "library"` plugins out of the
nav; the host reported no category for them at all. `defineLibraryPlugin` sets it and
`sdk/src/ui.ts` forwards it — what shipped did not: `@punktfunk/host` was bumped to 0.1.2 on
2026-07-20 and `category` landed 2026-08-05 without a bump, so the registry's 0.1.2 is the
pre-category build. ⚠ **Inert until published.** `serveUi` now reads its own directory entry back
and warns once when a requested category did not land.
- **Local art sync failed on a `file://` disagreement.** `local_art_bytes` decodes a `file://` value
before testing containment; `validate_art_paths` handed the raw value to `Path::new`. Same defect
produced both the unreachable settings and `sync (startup) failed: HostRequestError`.
- **The runner now carries SDK updates.** The copy each installed plugin runs was pinned at install
time, so an SDK fix could never reach it.
- **`bun publish` runs `prepare`, and `prepare` needs bun2nix** — the SDK could not be published at
all. Also fixed: a corrupt committed `bun.lock` in plugin-kit.
- **Decky client update.** `flatpak remote-info punktfunk-origin io.unom.Punktfunk` names no branch;
the remote publishes `stable` **and** `canary`, so the ref is ambiguous and flatpak refuses it —
⚠ one branch being *installed* does not disambiguate, the ambiguity is on the remote. The call
failed on every box, every time, and returned `available=False`, which the panel rendered as good
news. Every query now names the ref in full via `_flatpak_ref()` (no subprocess), carrying the
**scope** too, so a system-wide install is no longer invisible to a check that hardcoded `--user`.
A check that cannot run now reports `client_error`.
### Packaging
- **The `punktfunk` group is created everywhere the udev rule needs it.** `60-punktfunk.rules`
chgrp's the usbip vhci attach/detach nodes to a dedicated group (security review 2026-08-05 M-4:
writing `attach` materialises an arbitrary emulated USB device, so it must not ride on `input`).
**Four of six install paths shipped that rule in 0.25.0 without creating the group** — chgrp
failed, nodes stayed `root:root 0644`, the virtual Deck pad silently never attached, and
`usermod -aG punktfunk` failed outright. Fixed in arch `post_upgrade()` (only `post_install` was
correct, so every box that reached 0.25.0 by `pacman -Syu` missed it), nix (`users.groups.punktfunk`
did not exist), the bazzite sysext (a group is host state and cannot ride an image), and the Steam
Deck scripts. deb and rpm were correct throughout.
- **`punktfunk-gamescope` now builds for RPM and apt**, not Arch only.
- **Arch release-rebuild prune** called a helper that cannot exist in a release rebuild. Together
with the FFmpeg 9 repackage this closes the 0.25.0-1 → 0.25.0-2 episode in the pipeline rather
than by hand.
- **Steam Deck `update.sh` / `install.sh`.** The web step ran `bun install --frozen-lockfile` with
no `--ignore-scripts`, so web's `postinstall` (`bun2nix -o bun.nix`) rewrote a **tracked** file on
every update; the SDK step below it had always passed `--ignore-scripts`, and that asymmetry is
the whole bug. Now `--ignore-scripts` plus an explicit `bun run codegen` — provably equivalent,
since web's `prepare` is literally `"bun run codegen"` and `src/api/gen`, `src/paraglide` and
`src/routeTree.gen.ts` are gitignored. `--pull` restores `web/bun.nix` and `sdk/bun.nix` before
pulling, which is lossless by construction. ⚠ Deliberately **not** `git reset --hard`: `$SRC`
defaults to the operator's own checkout. Also: `web.env` secret hygiene — `chmod 600` sat inside
the create-only branch, so an install set up once and only updated since kept it world-readable.
`packaging/debian/build-web-deb.sh`, `packaging/arch/PKGBUILD` and `packaging/rpm/punktfunk.spec`
still lack `--ignore-scripts` for web — harmless (throwaway build trees), left as follow-up.
### Triage tooling
**`--probe-decode` described a different device from the one that streams.** The RADV
video-decode opt-in sat *after* the `--list-adapters` / `--probe-decode` / `--list-audio` / `--pair`
early exits, so the triage tool never had it. Measured on a Deck, same binary back to back: bare
`--probe-decode` printed "vulkan video decode: no", "driver decode ops: none (0x0)", "no queue
family advertises VIDEO_DECODE"; with `RADV_PERFTEST=video_decode` in the environment, "YES" and
"H.264, H.265, AV1, VP9". ⚠ **Any Deck triage that consulted it reached the opposite of the truth.**
Hoisted to the top of `run`, ahead of every early exit.
### PyroWave on Linux — Wave 2
The program's own measurement, from patch 0005's header: `encode_gpu_synchronous` goes from ~2 ms
to **1518 ms at 95 % game load**, with the stream frame rate collapsing. PyroWave encodes on the
same shader cores a game saturates; NVENC is immune because it has its own ASIC.
- **PW1 — the GPU-priority lever had never fired on Linux.** The vendored patch requests an elevated
global-priority queue, gated `if (!inherit_info)` — and **only Windows leaves `inherit_info` null**
(`pyrowave_create_device_by_compat`, where Granite builds the device itself). Linux passes its own
create-infos, Granite's `get_existing_create_info()` hands them back, `create_device` takes the
inherit branch, and the whole block is skipped. Now wired natively in `open_inner`'s `DeviceHold`,
ladder REALTIME → HIGH → no-priority, stepping only on refusal; a refused class can never fail the
open. The extension probe reuses the `dev_ext_props` already fetched for `queue_family_foreign` and
takes KHR or the EXT alias — the same spelling pf-zerocopy probes, so the two cannot disagree.
**Needs `CAP_SYS_NICE`**, which the packaging granted in `0.26.0-1`; without it the lever does
nothing.
🛑 **Corrected in `0.26.0-2`: the packaging no longer grants it, and must not.** Every channel that
did (Arch `.install`, RPM `%caps()`, the Bazzite sysext image, the deb postinst, the NixOS
`security.wrappers` entry) broke desktop streaming on KDE outright — field-reported on CachyOS and
Bazzite as `KWin does not expose zkde_screencast_unstable_v1 to this client`. KWin identifies a
client by resolving its `/proc/<pid>/exe` against an installed `.desktop`, and the kernel refuses
that readlink to any reader whose effective set is not a superset of the target's **permitted**
set (`cap_ptrace_access_check`) — KWin has no capabilities, so a capability-carrying host is
unidentifiable and the restricted globals are never advertised. Neither `prctl(PR_SET_DUMPABLE, 1)`
nor systemd `AmbientCapabilities=` rescues it; only an uncapped process is identifiable. The lever
therefore stays wired but unexercised on a stock install (the ladder degrades to default priority),
and is opt-in for gamescope-only hosts, which have no such identity check.
- **PW5 — two encoder handles.** `Encoder::Impl` owns exactly one each of `wavelet_img_high_res`,
`bucket_buffer`, `meta_buffer`, `block_stat_buffer`, `payload_data`, `quant_buffer`, and
`Impl::encode` *opens* by discarding them (an image barrier with `VK_IMAGE_LAYOUT_UNDEFINED` as the
old layout, plus three `fill_buffer` clears). Two encodes submitted to one queue have **no**
execution dependency in Vulkan — submission order orders the start, not the completion — so N+1's
DWT would overwrite N's wavelet bands while N's block packing still reads them. Content-dependent
and silent. Overlap therefore means two handles alternated, one per slot. ⚠⚠ **The landmine:**
`sequence_count` also lives on `Impl`, and it is the **3-bit** counter stamped into every block
header. Two handles each counting 1,2,3… put 1,1,2,2,3,3… on the wire, and the decoder restarts a
frame only when the value *changes* — so a repeat reads as more blocks of the same frame. Depth is
**still 1**; the handles alternate with one in flight.
- **PW3 — the fence wait moved out of submit.** PyroWave was the one backend waiting its fence inside
`submit`.
- **PW7a — the jumbo leg was dead code.** quinn caps a peer's MTU-discovery search at
`min(MtuDiscoveryConfig::upper_bound, the other side's advertised max_udp_payload_size)`, and
`EndpointConfig::max_udp_payload_size` **defaults to 1472**. Nothing in the repo had ever touched
`EndpointConfig`, so raising the host's probe ceiling could never make discovery settle above 1472
— and the shipped mid-session grow's `settled >= sealed_datagram_bytes(target)` gate was
unreachable on **every path that has ever existed**. Two smaller contributors fixed with it: the
watcher stopped sampling the moment `settled >= 1472`, discarding the very climb the proof needs;
and a session sealed above the 1500-byte default was never checked against the path at all.
The advertisement is raised on the **client** endpoint under the same `jumbo_wire_mtu()` opt-in,
because it is not free: quinn sizes its endpoint receive buffer
`max_udp_payload_size × max_receive_segments × BATCH_SIZE` — on a GRO-capable Linux/Android client
that is ~2.9 MiB at the default and **~18 MiB at jumbo** (47 KiB → 288 KiB on Apple/Windows).
PyroWave is the codec that most wants this: it can never be re-keyed mid-stream (its client parses
chunk-aligned AUs in windows of the `Welcome` value, read once over the C ABI), so it should
*start* at the big shard. At an 8908-byte shard that is ~6× fewer datagrams per frame — **~49k → ~8k
pps at 550 Mb/s**.
### Zero-copy capture
- **The dmabuf latch conflated two causes with different lifetimes.** One `AtomicBool` served both
"the encoder repeatedly failed to import what this compositor allocates" (unrecoverable, a driver
fact) and "the dmabuf-only capture offer never negotiated" (which can just mean the compositor was
mid-restart). Sharing it made the second as permanent as the first: **one timeout, and every later
session on that host captured CPU frames until the process restarted** — including sessions against
a different compositor and a different node that had never failed at anything, with nothing said.
Now a `RawDmabufLatch` owning both: import failures stay sticky (unchanged 3-consecutive threshold);
negotiation timeouts get a retry budget of **2** — deliberately small, since each failure costs a
~10 s stall the user pays in dead air; a capture that negotiates credits the budget back; and both
are keyed to a capture identity (node id + portal bit).
- **The zero-copy path never asked for buffer headroom.** `build_dmabuf_buffers` set
`SPA_PARAM_BUFFERS_dataType` and stopped — no `SPA_PARAM_BUFFERS_buffers` at all, so the pool depth
every zero-copy safety argument rests on was entirely the producer's choice and we never expressed
a preference. Now asks for 8 (min 2, max 16) as a **Choice Range, deliberately not a fixed count**:
SPA intersects consumer and producer params, so a fixed 8 against a producer that can only afford 4
empties the intersection and the link stalls in "negotiating" with no error anywhere — ⚠ the exact
trap that once cost this codebase the entire Linux cursor channel, when a 256² cursor-meta max
failed to intersect Mutter's fixed 384². 8 buffers is ~133 ms of pool at 60 Hz and ~33 ms at 240 Hz;
16 is a ceiling, not a request (a 4K 4:4:4 buffer is ~25 MB).
- **A PyroWave session could drop to CPU capture and log nothing.** The CPU-fallback warning was gated
on `backend_is_vaapi`, which reads the **host-global** encoder pref — but a PyroWave session is
negotiated **per session**, so on an NVIDIA/auto host that gate is false and the session fell out of
every arm of the negotiation log chain while paying a full-resolution CPU pixel touch every frame.
A degraded host and a healthy one produced identical logs. Now asks the per-session question
(`consumer_kind`), widened to every GPU consumer and excluding only the software encoder, whose
native input *is* CPU frames. ⚠ `pyrowave_session` must outrank `backend_is_vaapi`, because a
PyroWave pref flips `backend_is_vaapi` on too.
### Steam-overlay input masking (Steam Deck)
On a Deck in Gaming Mode the Steam menu and the QAM are driven by the **same physical controller** the
client forwards, so opening either moved the game on the host as well — a second, invisible player.
Steam Input masks a normal game here; it cannot mask us, because masking happens on Steam Input's
virtual pad and we deliberately forward the **real** one (the virtual pad has no gyro, trackpads or
paddles).
**SDL's own gate cannot fire on a Deck.** SDL drops presses while a process has windows but no
keyboard focus, and it is on by default — but gamescope resolves focus per Xwayland ctx and the client
sits alone in its own, so the Steam overlay (which lives in the root ctx) never takes our X focus and
no `FocusOut` is ever generated. Measured on glass: with the QAM open, X input focus inside the
client's ctx stayed on its window for the whole 4 s while `GAMESCOPE_FOCUSED_APP` flipped to 769
(Steam) and `GAMESCOPE_FOCUSED_APP_GFX` stayed on the app. **That pair of atoms is the signal.**
`overlay_focus` watches them on the gamescope **root** ctx, which is *not* our own `$DISPLAY` under
`--xwayland-count 2` — hence the socket-directory walk and the flatpak filesystem line.
⚠⚠ Masking is deliberately **not** `set_forwarding`: that closes the slot and sends `GamepadRemove`,
so the game would see a controller **unplug** every time somebody opened the QAM. Every slot stays
open and only transitions stop, after flushing what the host believes is held (so a stick deflected at
overlay-open stops steering instead of freezing at its last value). On the way back, held buttons are
**adopted rather than replayed** — the A that picked a QAM row must not fire in the game as it closes
— while axes *are* re-sent, since a stick has no press to ghost and SDL only speaks on change.
### The `plugin` launch kind
The 2026-08-05 review made `launch.kind = "command"` operator-only, and a reconcile refuses on the
**first** offending entry — so rom-manager, whose every ROM is `<emulator> <args> <rom>`, stopped
putting anything in the library at all. Playnite hit the same wall and was rescued with a typed kind
the host resolves itself; there is no fixed scheme for "whichever emulator the operator configured,
with the core and flags they chose", so that trick does not generalise.
The entry now carries an **opaque key and nothing executable**, and the host asks the owning plugin
what to run at launch time, over the loopback UI port and per-boot secret it already registered.
**A stolen plugin token stops being command execution:** planting an entry is not enough, because
the live plugin answers 404 for a key it never published. Nothing executable is persisted or served to
a client, and an emulator that moved is picked up on the next launch rather than leaving a dead tile
(same reasoning as `xbox` resolving its AUMID at launch time).
**The host still spawns it**, because only the host can put the process where the stream can see it:
on Linux that is either gamescope's own argv or a spawn carrying the session's compositor env, and the
returned child is what session-game-lifetime tracks to know the game exited. A plugin spawning the
emulator itself would land it outside both.
### Verification status
| | |
|---|---|
| gamescope 0006 | 6/6 exit 0 on a release build at the real spawn shape (`2752x2064@120 --steam --xwayland-count 1`); distro control SIGSEGVs |
| Decky client update | on the Deck against the real install — pre-fix `available=False remote=''`, post-fix `available=True remote=ca010668` |
| `--probe-decode` | on a Deck, same binary back to back, with and without the RADV opt-in |
| Apple audio | builds on arm64-apple-ios17.0 (the triple that compiles the `#if os(iOS)` blocks — a plain `swift build` is macOS and skips them), arm64-apple-tvos17.0, macOS; 257 Swift tests |
| Audio jitter | 10-minute Wi-Fi power-save simulation, ~2000 → 9 audible events |
| 4:4:4 gate | test verified non-vacuous against the planted original defect |
| Steam Deck scripts | `bash -n` + shellcheck 0.11.0 clean at `-S warning`; exec bits preserved |
| Steam-overlay masking | on glass on a Deck — atom flip and X-focus non-flip both measured over a 4 s QAM open |
| PyroWave depth 2 | exercised on real hardware **without shipping depth 2** (dedicated test, shipped depth stays 1) |
| PW6 streamed AU | the trap is real, and at 2 % loss it costs exactly nothing |
**Owed on glass:** iPhone + Bluetooth listen, Apple TV stats overlay, MacBook audio listen, the
Deck HEVC/4:4:4 retest, a Windows wake-from-sleep cycle, and the PyroWave-under-game-load A/B on a
Linux host with `CAP_SYS_NICE` actually granted — the number this whole wave is aimed at. ⚠ That
last one now needs a **gamescope-only** host, or a hand-granted capability on a box you are not
streaming the KDE desktop from: see the `0.26.0-2` correction under PW1 above.
---
## v0.25.0
407 commits since v0.24.0.
@@ -79,6 +499,20 @@ capability rode on `input`, which every gamepad guide tells users to join — bu
arbitrary USB hardware. Operators must `usermod -aG punktfunk "$USER"` and re-login or the pad stops
attaching. Ordinary virtual gamepads are unaffected.
> **Known issue in 0.25.0, fixed after it.** Four of the six install paths shipped
> `60-punktfunk.rules` — whose `RUN+=` does `chgrp punktfunk` on the vhci `attach`/`detach` nodes —
> without ever creating the group, so the `chgrp` failed, the nodes stayed root-only, and the pad
> silently never attached. The `usermod` above also fails outright on those boxes with *group
> 'punktfunk' does not exist*. Affected: **Arch/CachyOS upgraded** rather than freshly installed
> (`post_upgrade` called only `_ensure_update_group`), the **NixOS module** (no
> `users.groups.punktfunk`), the **Bazzite sysext** (a group is host state and cannot ride an
> image), and **Steam Deck source installs** (`scripts/steamdeck/install.sh`/`update.sh` handled
> only `input`). The deb and rpm scriptlets were correct throughout — they run one `%post`/`postinst`
> on install and upgrade alike. All four now create the group, and the two that know which user
> runs the host (the Deck scripts and the NixOS module's `host.users`) add that user to it as well.
> Workaround on an unpatched box:
> `sudo groupadd --system punktfunk`, then the `usermod`, then re-login.
**3. Plugins may no longer set `launch.command` or the pre-launch command.** Both run through a
shell and are now operator-token only; a plugin that sets them is refused. Third-party plugins that
populated them need updating — use the `launcher_ui` / `xbox` launch kinds instead.
@@ -437,6 +871,24 @@ refuses the upgrade instead of bricking the install. All seven libs are listed e
`--as-needed` currently drops two: an unlinked soname is left bare by makepkg and satisfied by any
ffmpeg, so listing it costs nothing and a future link picks up the bound automatically.
🛑 **The v0.25.0 Arch packages shipped with that bound pointing at the WRONG FFmpeg — install
`punktfunk-host 0.25.0-2` or newer.** The soname fix and the FFmpeg-9 build landed as one merge;
the release tag was pushed four minutes later, while the CI builder image was still being
rebuilt. arch.yml deliberately runs no `-Syu` ("the image's snapshot IS the build environment"),
so the release was linked against FFmpeg 8 and published `libavcodec.so=62-64` — a bound no
up-to-date Arch box can satisfy. It fails *safely* (pacman refuses; nothing bricks), but it fails
**loudly and broadly**: pacman prepares one transaction, so an unsatisfiable dependency of ours
stopped affected users' entire `pacman -Syu`. `0.25.0-2` is the identical source rebuilt against
FFmpeg 9. Only Arch was exposed — every other format derives its dependency from the ELF at build
time and could not disagree with itself this way.
Two guards now stand where only a convention did. arch.yml compares the builder's libav
`provides` against the live repos before building and `-Syu`s itself if they differ; and no
package is published until a **pristine-`--dbpath`** `pacman -U --print` resolves it, which asks
"would a real, up-to-date Arch box install this?" instead of "does the builder happen to satisfy
it?" — the distinction that let this ship. Keeping `ci/arch-ci.Dockerfile` current is still the
cheap path; the guards are the backstop.
### Linux playback filled the buffer ceiling
The PipeWire playback callback sized its writes from the mapped buffer's **capacity** — PipeWire's
Generated
+36 -35
View File
@@ -994,7 +994,7 @@ dependencies = [
[[package]]
name = "cursor-probe"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"pf-capture",
@@ -1114,7 +1114,7 @@ dependencies = [
[[package]]
name = "display-disturb"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"windows 0.62.2 (registry+https://github.com/rust-lang/crates.io-index)",
]
@@ -2358,7 +2358,7 @@ dependencies = [
[[package]]
name = "latency-probe"
version = "0.25.0"
version = "0.26.0"
[[package]]
name = "lazy_static"
@@ -2463,7 +2463,7 @@ dependencies = [
[[package]]
name = "libvpl-sys"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"bindgen",
"cmake",
@@ -2498,7 +2498,7 @@ checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad"
[[package]]
name = "loss-harness"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"punktfunk-core",
]
@@ -2988,7 +2988,7 @@ checksum = "9b4f627cb1b25917193a259e49bdad08f671f8d9708acfd5fe0a8c1455d87220"
[[package]]
name = "pf-bitstream"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"cros-codecs",
"tracing",
@@ -2996,7 +2996,7 @@ dependencies = [
[[package]]
name = "pf-capture"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ashpd",
@@ -3017,7 +3017,7 @@ dependencies = [
[[package]]
name = "pf-client-core"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ash",
@@ -3047,11 +3047,12 @@ dependencies = [
"wasapi",
"windows 0.62.2 (git+https://github.com/microsoft/windows-rs?rev=acb5a1a7441033d9312b16842af02eb0c2b403dc)",
"winreg",
"x11rb",
]
[[package]]
name = "pf-clipboard"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ashpd",
@@ -3069,7 +3070,7 @@ dependencies = [
[[package]]
name = "pf-console-ui"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ash",
@@ -3090,7 +3091,7 @@ dependencies = [
[[package]]
name = "pf-dxvadec"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"cros-codecs",
"pf-bitstream",
@@ -3100,7 +3101,7 @@ dependencies = [
[[package]]
name = "pf-encode"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ash",
@@ -3124,7 +3125,7 @@ dependencies = [
[[package]]
name = "pf-frame"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"libc",
@@ -3136,7 +3137,7 @@ dependencies = [
[[package]]
name = "pf-gpu"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"pf-host-config",
@@ -3150,11 +3151,11 @@ dependencies = [
[[package]]
name = "pf-host-config"
version = "0.25.0"
version = "0.26.0"
[[package]]
name = "pf-inject"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ashpd",
@@ -3183,14 +3184,14 @@ dependencies = [
[[package]]
name = "pf-paths"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"tracing",
]
[[package]]
name = "pf-presenter"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ash",
@@ -3205,7 +3206,7 @@ dependencies = [
[[package]]
name = "pf-update"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"serde",
"serde_json",
@@ -3213,7 +3214,7 @@ dependencies = [
[[package]]
name = "pf-update-check"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"base64",
@@ -3225,7 +3226,7 @@ dependencies = [
[[package]]
name = "pf-vaadec"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"cros-codecs",
"pf-bitstream",
@@ -3234,7 +3235,7 @@ dependencies = [
[[package]]
name = "pf-vdisplay"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ashpd",
@@ -3267,7 +3268,7 @@ dependencies = [
[[package]]
name = "pf-vkdecode"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"ash",
"cros-codecs",
@@ -3278,7 +3279,7 @@ dependencies = [
[[package]]
name = "pf-win-display"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"pf-paths",
@@ -3290,7 +3291,7 @@ dependencies = [
[[package]]
name = "pf-zerocopy"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ash",
@@ -3513,7 +3514,7 @@ dependencies = [
[[package]]
name = "punktfunk-cli"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"pf-client-core",
"punktfunk-core",
@@ -3524,7 +3525,7 @@ dependencies = [
[[package]]
name = "punktfunk-client-android"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"android_logger",
"jni",
@@ -3542,7 +3543,7 @@ dependencies = [
[[package]]
name = "punktfunk-client-linux"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"async-channel",
@@ -3559,7 +3560,7 @@ dependencies = [
[[package]]
name = "punktfunk-client-session"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"pf-client-core",
@@ -3574,7 +3575,7 @@ dependencies = [
[[package]]
name = "punktfunk-client-windows"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"async-channel",
"mdns-sd",
@@ -3593,7 +3594,7 @@ dependencies = [
[[package]]
name = "punktfunk-core"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"aes-gcm",
"bytes",
@@ -3625,7 +3626,7 @@ dependencies = [
[[package]]
name = "punktfunk-host"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"aes",
"aes-gcm",
@@ -3710,7 +3711,7 @@ dependencies = [
[[package]]
name = "punktfunk-probe"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"mdns-sd",
@@ -3724,7 +3725,7 @@ dependencies = [
[[package]]
name = "punktfunk-tray"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"anyhow",
"ksni",
@@ -3747,7 +3748,7 @@ checksum = "d55d956fa96f5ec02be2e13af0e20391a5aa83d6a074e3ad368959d0fab299ea"
[[package]]
name = "pyrowave-sys"
version = "0.25.0"
version = "0.26.0"
dependencies = [
"bindgen",
"cmake",
+1 -1
View File
@@ -57,7 +57,7 @@ exclude = [
ndk = { path = "clients/android/native/vendor/ndk" }
[workspace.package]
version = "0.25.0"
version = "0.26.0"
edition = "2021"
rust-version = "1.82"
license = "MIT OR Apache-2.0"
+125 -4
View File
@@ -10,7 +10,7 @@
"name": "MIT OR Apache-2.0",
"identifier": "MIT OR Apache-2.0"
},
"version": "0.24.0"
"version": "0.25.0"
},
"paths": {
"/api/v1/clients": {
@@ -997,7 +997,7 @@
"library"
],
"summary": "List the game library",
"description": "Every installed-store title (Steam, read from the host's local files — no Steam API key)\nmerged with the user's custom entries, sorted by title. Artwork fields are URLs the client\nfetches directly (the public Steam CDN for Steam titles). `?provider=` narrows to the\nentries a given external provider owns; `?platform=` to one platform (case-insensitive —\ninstalled-store titles are `PC`, custom/provider entries carry whatever was authored).",
"description": "Every installed-store title (Steam, read from the host's local files — no Steam API key)\nmerged with the user's custom entries, sorted by title. Artwork fields are URLs the client\nfetches directly (the public Steam CDN for Steam titles). `?provider=` narrows to the\nentries a given external provider owns; `?platform=` to one platform (case-insensitive —\ninstalled-store titles are `PC`, custom/provider entries carry whatever was authored).\n\n**The operator's own lane additionally sees the titles they have HIDDEN**, each carrying\n`hidden: true`; every other lane gets them filtered out upstream and cannot tell they exist. The\nconsole needs them to offer \"un-hide\", and it is the only surface that does.",
"operationId": "getLibrary",
"parameters": [
{
@@ -1021,13 +1021,13 @@
],
"responses": {
"200": {
"description": "Unified library across all stores",
"description": "Unified library across all stores (the operator's lane also gets hidden entries, flagged)",
"content": {
"application/json": {
"schema": {
"type": "array",
"items": {
"$ref": "#/components/schemas/GameEntry"
"$ref": "#/components/schemas/OperatorGameEntry"
}
}
}
@@ -1301,6 +1301,79 @@
}
}
},
"/api/v1/library/hidden/{id}": {
"put": {
"tags": [
"library"
],
"summary": "Hide or un-hide one library title",
"description": "Curation, not access control: a hidden title disappears from every play surface — the console\ngrid on a client, native clients, the GameStream app list, and launch resolution — while nothing\nis deleted and un-hiding restores it immediately. The operator's own console still lists it\n(flagged `hidden`) so it can be brought back.\n\nKeyed by the entry's stable `<store>:<external_id>` id, which survives re-scans and reconciles by\nconstruction (D2). The id is **not** validated against the current library on purpose: a title\ncan be legitimately absent at this moment (launcher closed, plugin mid-sync, drive unmounted),\nand refusing the operator's choice in that window would be worse than storing an id that\ncurrently matches nothing. Emits `library.changed` (source = the store) only on a real change.",
"operationId": "setLibraryEntryHidden",
"parameters": [
{
"name": "id",
"in": "path",
"description": "The library entry id (e.g. `steam:70`)",
"required": true,
"schema": {
"type": "string"
}
}
],
"requestBody": {
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/HiddenToggle"
}
}
},
"required": true
},
"responses": {
"200": {
"description": "Stored; the entry's visibility after the call",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/HiddenState"
}
}
}
},
"400": {
"description": "Empty entry id",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ApiError"
}
}
}
},
"401": {
"description": "Missing or invalid bearer token",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ApiError"
}
}
}
},
"500": {
"description": "Could not persist the settings",
"content": {
"application/json": {
"schema": {
"$ref": "#/components/schemas/ApiError"
}
}
}
}
}
}
},
"/api/v1/library/provider/{provider}": {
"put": {
"tags": [
@@ -5553,6 +5626,37 @@
}
}
},
"HiddenState": {
"type": "object",
"description": "What `setLibraryEntryHidden` echoes back.",
"required": [
"id",
"hidden"
],
"properties": {
"hidden": {
"type": "boolean",
"description": "Its visibility after the call."
},
"id": {
"type": "string",
"description": "The entry id the call addressed."
}
}
},
"HiddenToggle": {
"type": "object",
"description": "Request body for `setLibraryEntryHidden`.",
"required": [
"hidden"
],
"properties": {
"hidden": {
"type": "boolean",
"description": "Whether this title should be hidden from every play surface."
}
}
},
"HookEntry": {
"type": "object",
"description": "One hook: fire `run` and/or `webhook` when an event matching `on` (+ `filter`) occurs.",
@@ -6339,6 +6443,23 @@
}
}
},
"OperatorGameEntry": {
"allOf": [
{
"$ref": "#/components/schemas/GameEntry"
},
{
"type": "object",
"properties": {
"hidden": {
"type": "boolean",
"description": "The operator hid this title ([`set_entry_hidden`]) — omitted when false, so the shape only\ngrows for entries that actually are hidden."
}
}
}
],
"description": "A library entry plus the operator's own view of it — today, whether they hid it.\n\nA separate type rather than a field on [`GameEntry`] for two reasons. It keeps the visibility\nanswer out of the providers entirely: a store parser has no opinion on what the operator hid, and\nadding `hidden: false` to all eight construction sites would imply it does. More importantly it\nmakes the lane rule a TYPE guarantee instead of a discipline — `GET /library` answers\n`Vec<GameEntry>` on every lane but the operator's, so a hidden entry cannot leak to a paired\nclient by someone forgetting a filter; there is no field there to leak.\n\n`flatten` keeps the wire shape identical to a plain entry with one extra key, so the console\nparses one model either way."
},
"PairedClient": {
"type": "object",
"description": "A paired (certificate-pinned) Moonlight client.",
+10
View File
@@ -19,6 +19,16 @@
# 63-64), so it would simply refuse to install rather than start. Re-keying this image is the step
# that makes the ffmpeg-9 bump actually reach the package — a Cargo.toml bump alone does nothing
# here. Whenever Arch moves to an FFmpeg major, bump the date in the same commit.
#
# ⚠ AND KNOW WHY THAT WAS NOT ENOUGH: bumping this date only helps once docker.yml has actually
# republished the image, and nothing sequences the two workflows. v0.25.0 was tagged four minutes
# after the ffmpeg-9 merge, so the release build still pulled the FFmpeg-8 `:latest` and published
# a punktfunk-host that no up-to-date Arch box could install — which blocks the user's ENTIRE
# `pacman -Syu`, not just our package. arch.yml therefore no longer trusts this image on that one
# axis: it compares the builder's libav sonames against the repos before building (and `-Syu`s
# itself if they differ), and refuses to publish anything a pristine-db `pacman -U --print` says
# is unsatisfiable. This file staying current is still the CHEAP path — those guards are the
# backstop, not the plan.
FROM docker.io/library/archlinux:base-devel
# One transaction: the main build/runtime deps (first list) + the gamescope companion's
@@ -69,11 +69,14 @@ fun App(forceGamepadUi: Boolean = false) {
// later manual Back out of the library is not undone by a stale value.
var reopenLibraryHostId by remember { mutableStateOf<String?>(null) }
// Console (gamepad) mode mirrors the Apple client: the setting AND (a pad is attached OR this is
// a TV OR the dev force flag). Flips live as controllers connect/disconnect.
// Console (gamepad) mode mirrors the Apple client: the setting AND (its mode says Always OR a
// pad is attached OR this is a TV OR the dev force flag). Flips live as controllers
// connect/disconnect — unless the mode is Always, where it simply stays.
val tv = remember { isTvDevice(context) }
val controllerConnected by rememberControllerConnected()
val gamepadUi = gamepadUiActive(settings.gamepadUiEnabled, controllerConnected, tv, forceGamepadUi)
val gamepadUi = gamepadUiActive(
settings.gamepadUiEnabled, settings.gamepadUiMode, controllerConnected, tv, forceGamepadUi,
)
// Publish the live session process-wide, so a `punktfunk://` link that arrives as a SECOND
// activity instance (the normal case under `launchMode = standard`) can refuse it before that
@@ -67,7 +67,7 @@ class GamepadPalette(
)
/**
* The twelve shipped palettes: the brand default, five more dark fields, then six pale
* The thirteen shipped palettes: the brand default, six more dark fields, then six pale
* ones. Cycling order runs dark → light, so stepping the row walks the range one way.
*/
val ALL = listOf(
@@ -77,6 +77,22 @@ class GamepadPalette(
ground = Triple(0.075, 0.060, 0.160),
accent = Triple(0.525, 0.471, 0.961), light = false,
),
GamepadPalette(
// For OLED and AMOLED panels, where a black pixel is a pixel switched off — no
// glow, no power. The first two stops are literally (0,0,0), so the shaded half
// of the field is genuinely off rather than "very dark grey", and the ground is
// pure black too: the calm mix on the form screens lifts toward nothing. What is
// left is a faint indigo→violet ember in the bright corner. The accent stays the
// brand violet — focus has to be findable on black.
"oled", "OLED",
listOf(
Triple(0.000, 0.000, 0.000), Triple(0.000, 0.000, 0.000),
Triple(0.010, 0.020, 0.100), Triple(0.045, 0.016, 0.115),
Triple(0.120, 0.024, 0.130),
),
ground = Triple(0.0, 0.0, 0.0),
accent = Triple(0.525, 0.471, 0.961), light = false,
),
GamepadPalette(
// Deep indigo climbing through violet into a hot magenta.
"nebula", "Nebula",
@@ -665,6 +665,21 @@ internal fun buildSettingsRows(
"Turn off to use the touch interface even with a controller connected.",
s.gamepadUiEnabled,
) { update(s.copy(gamepadUiEnabled = it)) },
) + listOfNotNull(
// WHEN the switch above takes over. Built only while it is ON: turn the switch off from
// this very screen and the row under the cursor would otherwise be one deciding nothing,
// on a screen that is itself about to disappear.
if (s.gamepadUiEnabled) {
choice(
"gamepadUIMode", GpTab.INTERFACE, null, "Show it",
"With a controller: the touch interface comes back when the last one " +
"disconnects. Always keeps this layout either way — for a device that lives " +
"docked to a TV. A TV itself is always in this mode regardless.",
GAMEPAD_UI_MODE_OPTIONS, s.gamepadUiMode,
) { update(s.copy(gamepadUiMode = it)) }
} else {
null
},
)
}
@@ -16,15 +16,35 @@ import androidx.compose.runtime.remember
import androidx.compose.ui.platform.LocalContext
import io.unom.punktfunk.kit.Gamepad
/**
* [Settings.gamepadUiMode]: take over only while a controller is attached. The default, and what
* the switch meant when it was a lone Boolean.
*/
const val GAMEPAD_UI_WHEN_CONNECTED = "connected"
/**
* [Settings.gamepadUiMode]: take over whenever the switch is on, pad or no pad — for a phone or
* tablet that lives docked to a TV, where the console layout is the one wanted and the pad is not
* always awake.
*/
const val GAMEPAD_UI_ALWAYS = "always"
/**
* Whether the controller-optimized "console" home (the host carousel + gamepad chrome) should
* replace the touch UI — the Android mirror of the Apple client's `GamepadUIEnvironment.isActive`:
* the user's [enabled] setting AND (a controller is attached OR this is a TV OR the dev [forced]
* flag). A TV counts unconditionally — its remote/gamepad is the only input, so it's always the
* console UI (as long as the setting is on).
* the user's [enabled] setting AND (the [mode] is [GAMEPAD_UI_ALWAYS] OR a controller is attached
* OR this is a TV OR the dev [forced] flag). A TV counts unconditionally — its remote/gamepad is
* the only input, so it's always the console UI (as long as the setting is on), which is why the
* mode row means nothing there. An unrecognized [mode] waits for a controller, so a value a newer
* client wrote can never strand this one in a layout it has no way back out of.
*/
fun gamepadUiActive(enabled: Boolean, controllerConnected: Boolean, tv: Boolean, forced: Boolean): Boolean =
enabled && (controllerConnected || tv || forced)
fun gamepadUiActive(
enabled: Boolean,
mode: String,
controllerConnected: Boolean,
tv: Boolean,
forced: Boolean,
): Boolean = enabled && (mode == GAMEPAD_UI_ALWAYS || controllerConnected || tv || forced)
/** True on a TV: the leanback/television feature or the TELEVISION ui-mode. */
fun isTvDevice(context: Context): Boolean {
@@ -94,11 +94,20 @@ data class Settings(
val touchMode: TouchMode = TouchMode.TRACKPAD,
/**
* Swap the whole home screen for the controller-optimized "console" UI (the host carousel +
* gamepad chrome) whenever a controller is connected — mirrors the Apple client's
* `gamepadUIEnabled`. On by default; turn it off to keep the touch UI even with a pad attached.
* gamepad chrome) — mirrors the Apple client's `gamepadUIEnabled`. On by default; turn it off
* to keep the touch UI even with a pad attached. WHEN it takes over is [gamepadUiMode].
* A TV (leanback) is always in this mode regardless (its remote/pad is the only input).
*/
val gamepadUiEnabled: Boolean = true,
/**
* When [gamepadUiEnabled] actually takes over — the cross-client `gamepad_ui_mode` pair,
* mirroring the Apple client's `gamepadUIMode`: `"connected"` (default, and what the switch
* has always meant) waits for a controller; `"always"` keeps the console UI with no pad in
* reach, for a phone or tablet that lives docked to a TV. Read only while [gamepadUiEnabled]
* is on, which is why both settings screens hide the row when the switch is off. Anything
* unrecognized resolves to `"connected"`. A TV ignores it — it is always in console mode.
*/
val gamepadUiMode: String = GAMEPAD_UI_WHEN_CONNECTED,
/**
* Show the experimental game-library browser (the coverflow reached with Y from a saved host).
* Fetched from the host's management API over mTLS; needs a paired host. Mirrors the Apple
@@ -107,9 +116,10 @@ data class Settings(
val libraryEnabled: Boolean = true,
/**
* Which colour family the console (gamepad) UI's living backdrop drifts through — the
* cross-client `ui_palette` key: `"violet"` (the brand default), `"tide"`, `"forest"`,
* `"ember"`, `"rose"`, `"graphite"`. See [GamepadPalette], whose table and maths mirror the
* desktop console's and the Apple client's under the same names. Presentation only: nothing
* cross-client `ui_palette` key: `"violet"` (the brand default), then `"oled"`, `"nebula"`,
* `"abyss"`, `"ember"`, `"moss"`, `"graphite"`, then the six pale fields. See
* [GamepadPalette], whose table and maths mirror the desktop console's and the Apple
* client's under the same names. Presentation only: nothing
* about a stream depends on it, so it is a device preference and never part of a profile.
* An unknown value reads as the default rather than failing — a newer client may have shipped
* a palette this build doesn't know.
@@ -303,6 +313,8 @@ class SettingsStore(context: Context) {
// Migration: the pre-enum Boolean "trackpad_mode" (true = trackpad, false = direct).
?: if (prefs.getBoolean(K_TRACKPAD, true)) TouchMode.TRACKPAD else TouchMode.POINTER,
gamepadUiEnabled = prefs.getBoolean(K_GAMEPAD_UI, true),
gamepadUiMode = prefs.getString(K_GAMEPAD_UI_MODE, GAMEPAD_UI_WHEN_CONNECTED)
?: GAMEPAD_UI_WHEN_CONNECTED,
libraryEnabled = prefs.getBoolean(K_LIBRARY, true),
uiPalette = prefs.getString(K_UI_PALETTE, "violet") ?: "violet",
lowLatencyMode = prefs.getBoolean(K_LOW_LATENCY, true),
@@ -344,6 +356,7 @@ class SettingsStore(context: Context) {
.putString(K_STATS_VERBOSITY, s.statsVerbosity.name)
.putString(K_TOUCH_MODE, s.touchMode.name)
.putBoolean(K_GAMEPAD_UI, s.gamepadUiEnabled)
.putString(K_GAMEPAD_UI_MODE, s.gamepadUiMode)
.putBoolean(K_LIBRARY, s.libraryEnabled)
.putString(K_UI_PALETTE, s.uiPalette)
.putBoolean(K_LOW_LATENCY, s.lowLatencyMode)
@@ -384,6 +397,7 @@ class SettingsStore(context: Context) {
const val K_HUD = "stats_hud_enabled"
const val K_TOUCH_MODE = "touch_mode"
const val K_GAMEPAD_UI = "gamepad_ui_enabled"
const val K_GAMEPAD_UI_MODE = "gamepad_ui_mode"
const val K_LIBRARY = "library_enabled"
const val K_UI_PALETTE = "ui_palette"
@@ -778,6 +792,13 @@ fun smoothBufferOptions(hz: Int): List<Pair<Int, String>> {
)
}
/** (stored value, label) for when the console UI takes over — the Apple client's table verbatim.
* Only offered while [Settings.gamepadUiEnabled] is on; a TV is in console mode either way. */
val GAMEPAD_UI_MODE_OPTIONS = listOf(
GAMEPAD_UI_WHEN_CONNECTED to "With a controller",
GAMEPAD_UI_ALWAYS to "Always",
)
/** (mode, label) for the touch-input model. */
val TOUCH_MODE_OPTIONS = listOf(
TouchMode.TRACKPAD to "Trackpad",
@@ -592,11 +592,24 @@ private fun GeneralSettings(s: Settings, update: (Settings) -> Unit) {
SettingsGroup("Interface") {
ToggleRow(
title = "Controller-optimized UI",
subtitle = "Switch to the console home when a controller is connected. A TV " +
"always uses it.",
subtitle = "Swap the touch home for the console home — the host carousel and " +
"gamepad chrome. A TV always uses it.",
checked = s.gamepadUiEnabled,
onCheckedChange = { on -> update(s.copy(gamepadUiEnabled = on)) },
)
// Only decides anything while the switch above is on, so it is HIDDEN rather than
// dimmed when it isn't — a picker whose every option changes nothing is worse than
// no picker, and this group is short enough that nothing jumps far.
if (s.gamepadUiEnabled) {
SettingDropdown(
label = "Show it",
options = GAMEPAD_UI_MODE_OPTIONS,
selected = s.gamepadUiMode,
caption = "With a controller: the touch home comes back when the last one " +
"disconnects. Always keeps the console home either way — for a device " +
"that lives docked to a TV.",
) { v -> update(s.copy(gamepadUiMode = v)) }
}
}
}
}
@@ -33,14 +33,14 @@ class GamepadPaletteTest {
fun tableMatchesTheOtherClients() {
assertEquals(
listOf(
"violet", "nebula", "abyss", "ember", "moss", "graphite",
"violet", "oled", "nebula", "abyss", "ember", "moss", "graphite",
"holo", "sunset", "bloom", "dawn", "mint", "opal",
),
GamepadPalette.ALL.map { it.id },
)
// Dark fields lead, pale ones follow, so stepping the row walks one direction.
val firstLight = GamepadPalette.ALL.indexOfFirst { it.light }
assertEquals(6, firstLight)
assertEquals(7, firstLight)
assertTrue(GamepadPalette.ALL.drop(firstLight).all { it.light })
// An unknown name is a newer client's palette, not an error.
assertEquals("violet", GamepadPalette.named("chartreuse").id)
@@ -72,6 +72,25 @@ class GamepadPaletteTest {
}
}
/**
* OLED is the one palette whose selling point is measurable: it has to be genuinely black,
* not merely the darkest of the dark fields. The blob field this client draws samples the
* ramp at 0.15/0.40/0.65/0.90, so its darkest blob lands in the all-black head of the ramp.
*/
@Test
fun oledIsActuallyBlack() {
val oled = GamepadPalette.named("oled")
assertEquals(Triple(0.0, 0.0, 0.0), oled.ground)
assertEquals(0f, oled.blobColors[0].red, 1e-6f)
assertEquals(0f, oled.blobColors[0].green, 1e-6f)
assertEquals(0f, oled.blobColors[0].blue, 1e-6f)
val mean = oled.stops.sumOf { luma(it) } / oled.stops.size
val darkestOther = GamepadPalette.ALL
.filter { it.id != "oled" && it.stops.isNotEmpty() }
.minOf { p -> p.stops.sumOf { luma(it) } / p.stops.size }
assertTrue("oled means $mean, barely under $darkestOther", mean < darkestOther / 2)
}
/** A pale palette really is pale — its ink flips, so a mislabelled one is unreadable. */
@Test
fun palettesAreHonestAboutLightness() {
@@ -95,4 +95,47 @@ class GamepadSettingsRowsTest {
// Drawn as a switch, and reading the persisted default.
assertEquals(true, row(on, "dsCapture").toggled)
}
/**
* The activation-mode row is a sub-setting of the Controller-optimized UI switch, so it is
* OFFERED only while that switch is on hidden rather than dimmed, because with the switch
* off this whole screen is about to be replaced by the touch UI and a dimmed row there would
* be one last thing to step past on the way out.
*/
@Test
fun `the activation-mode row follows the switch it belongs to`() {
fun ids(enabled: Boolean) = buildSettingsRows(
Settings(gamepadUiEnabled = enabled),
hasBodyVibrator = false, hasGyroscope = false, av1Capable = false,
) {}.map { it.id }
val on = ids(enabled = true)
assertTrue("the mode row is missing", "gamepadUIMode" in on)
assertEquals(
"the mode belongs directly under the switch it qualifies",
on.indexOf("gamepadUI") + 1,
on.indexOf("gamepadUIMode"),
)
val off = ids(enabled = false)
assertFalse("the mode row must not outlive its switch", "gamepadUIMode" in off)
assertTrue("the switch itself stays, or it could never be turned back on", "gamepadUI" in off)
}
/** Stepping the mode row writes the shared `gamepad_ui_mode` value, and wraps on A. */
@Test
fun `the activation-mode row steps the shared key`() {
var s = Settings()
fun mode() = buildSettingsRows(
s, hasBodyVibrator = false, hasGyroscope = false, av1Capable = false,
) { s = it }.first { it.id == "gamepadUIMode" }
assertEquals(GAMEPAD_UI_WHEN_CONNECTED, s.gamepadUiMode)
assertEquals("With a controller", mode().value)
assertFalse("already the first = thud", mode().adjust(-1))
assertTrue(mode().adjust(1))
assertEquals(GAMEPAD_UI_ALWAYS, s.gamepadUiMode)
// A from the last entry wraps home.
mode().activate()
assertEquals(GAMEPAD_UI_WHEN_CONNECTED, s.gamepadUiMode)
}
}
@@ -0,0 +1,53 @@
package io.unom.punktfunk
import org.junit.Assert.assertEquals
import org.junit.Assert.assertFalse
import org.junit.Assert.assertTrue
import org.junit.Test
/**
* [gamepadUiActive] is pure table-tested over its inputs, and the mirror of the Apple client's
* `GamepadUIEnvironmentTests`. The two clients share the stored `gamepad_ui_mode` values, so a
* disagreement here is a device that behaves differently from the same setting.
*/
class GamepadUiTest {
/** The default mode is what the switch meant when it was a lone Boolean. */
@Test
fun whenConnectedWaitsForAPad() {
assertTrue(gamepadUiActive(true, GAMEPAD_UI_WHEN_CONNECTED, true, tv = false, forced = false))
assertFalse(gamepadUiActive(true, GAMEPAD_UI_WHEN_CONNECTED, false, tv = false, forced = false))
assertFalse(gamepadUiActive(false, GAMEPAD_UI_WHEN_CONNECTED, true, tv = false, forced = false))
assertFalse(gamepadUiActive(false, GAMEPAD_UI_WHEN_CONNECTED, false, tv = false, forced = false))
// A TV is in console mode whatever the mode says — its remote is the only input.
assertTrue(gamepadUiActive(true, GAMEPAD_UI_WHEN_CONNECTED, false, tv = true, forced = false))
}
/** Always drops the controller from the decision but never the switch, which is the one
* way back to the touch UI. */
@Test
fun alwaysIgnoresThePadButNotTheSwitch() {
assertTrue(gamepadUiActive(true, GAMEPAD_UI_ALWAYS, false, tv = false, forced = false))
assertTrue(gamepadUiActive(true, GAMEPAD_UI_ALWAYS, true, tv = false, forced = false))
assertFalse(gamepadUiActive(false, GAMEPAD_UI_ALWAYS, false, tv = false, forced = false))
assertFalse(gamepadUiActive(false, GAMEPAD_UI_ALWAYS, true, tv = false, forced = false))
}
/** A value a newer client wrote waits for a pad rather than stranding this build in a
* layout it has no way back out of. */
@Test
fun anUnknownModeWaitsForAPad() {
assertFalse(gamepadUiActive(true, "whenever-i-say-so", false, tv = false, forced = false))
assertTrue(gamepadUiActive(true, "whenever-i-say-so", true, tv = false, forced = false))
assertFalse(gamepadUiActive(true, "", false, tv = false, forced = false))
}
/** The shipped default: the console UI still waits for a controller. */
@Test
fun theDefaultIsUnchangedBehaviour() {
val s = Settings()
assertTrue(s.gamepadUiEnabled)
assertEquals(GAMEPAD_UI_WHEN_CONNECTED, s.gamepadUiMode)
assertFalse(gamepadUiActive(s.gamepadUiEnabled, s.gamepadUiMode, false, tv = false, forced = false))
}
}
@@ -77,6 +77,7 @@ class ProfilesTest {
// Device-scope settings are not in the overlay at all, so no profile can move them.
assertEquals(base.gamepadUiEnabled, out.gamepadUiEnabled)
assertEquals(base.gamepadUiMode, out.gamepadUiMode)
assertEquals(base.libraryEnabled, out.libraryEnabled)
assertEquals(base.autoWakeEnabled, out.autoWakeEnabled)
assertEquals(base.sc2Capture, out.sc2Capture)
@@ -99,6 +99,10 @@ struct ContentView: View {
// with no (extended) controller attached tvOS falls back to HomeView as before.
@ObservedObject private var gamepadManager = GamepadManager.shared
@AppStorage(DefaultsKey.gamepadUIEnabled) private var gamepadUIEnabled = true
/// When the switch above takes over "connected" (default) or "always". See
/// `GamepadUIEnvironment`.
@AppStorage(DefaultsKey.gamepadUIMode) private var gamepadUIMode =
GamepadUIEnvironment.modeWhenConnected
/// Auto-wake on connect (Settings General). On (default): a dial to an offline saved host
/// fires Wake-on-LAN up front and falls into the "Waking" wait if the dial fails. Off: connects
/// go straight through with no wake. The explicit "Wake Host" action is unaffected either way.
@@ -113,7 +117,8 @@ struct ContentView: View {
@Environment(\.scenePhase) private var scenePhase
private var gamepadUIActive: Bool {
GamepadUIEnvironment.isActive(
gamepadConnected: gamepadManager.active != nil, enabledSetting: gamepadUIEnabled)
gamepadConnected: gamepadManager.active != nil, enabledSetting: gamepadUIEnabled,
mode: gamepadUIMode)
}
// The body is split in two `driven` (the screen plus its lifecycle drivers and sheets) and
@@ -901,6 +906,7 @@ struct ContentView: View {
private var shortcutHintText: String {
"Hold the remote's Back button — or L1+R1+Start+Select on a controller — to disconnect"
+ " · Touch surface moves the pointer · press clicks · Play/Pause right-clicks"
+ " · Hold Play/Pause, or Select+X on a controller, for statistics"
}
private static let shortcutHintFont: CGFloat = 22 // read from the couch
#endif
@@ -85,16 +85,40 @@ extension EnvironmentValues {
}
extension View {
/// Resolve the stored `ui_palette` and publish its ink to everything below. Applied by the
/// gamepad screens' common root so no individual view has to read the setting.
func gamepadPaletteInk() -> some View { modifier(GamepadInkModifier()) }
/// Resolve the stored `ui_palette` and publish its ink AND the matching colour scheme to
/// everything below. Applied by the gamepad screens' common root so no individual view has to
/// read the setting.
///
/// `active` exists for the one surface that is the same view in both worlds: `LibraryView`
/// renders the coverflow under the gamepad UI and a plain grid without it. Passing `false`
/// publishes nothing, because the touch/desktop layouts sit on the SYSTEM background, where a
/// palette's scheme would invert their own system colours instead of matching them.
func gamepadPaletteInk(_ active: Bool = true) -> some View {
modifier(GamepadInkModifier(active: active))
}
}
private struct GamepadInkModifier: ViewModifier {
var active = true
@AppStorage(DefaultsKey.uiPalette) private var paletteID = "violet"
/// The ambient scheme from ABOVE this modifier what gets republished unchanged when the
/// gamepad UI isn't the one drawing, so `active: false` is a true no-op rather than a branch
/// that would change this view's identity.
@Environment(\.colorScheme) private var systemScheme
func body(content: Content) -> some View {
content.environment(\.gamepadInk, GamepadInk.of(GamepadPalette.named(paletteID)))
let palette = GamepadPalette.named(paletteID)
return content
.environment(\.gamepadInk, active ? GamepadInk.of(palette) : .dark)
// The ink alone was never enough. Every SYSTEM-derived colour that lands on these
// screens `.secondary` in a placeholder, a `.bordered` button's chrome, a
// NavigationStack's title, a material's frost resolves against the DEVICE's
// appearance, which no part of this app had ever set. On iPhone and Mac that is often
// Light, so the pale palettes looked correct by accident; an Apple TV is Dark
// essentially always, so on tvOS every one of them came out WHITE on a pale field and
// the interface was unreadable. Publishing the scheme here once, beside the ink it
// has to agree with is what makes a pale palette mean "light" to UIKit too.
.environment(\.colorScheme, active ? (palette.light ? .light : .dark) : systemScheme)
}
}
@@ -32,9 +32,12 @@ struct LibraryView: View {
// setting off) every platform keeps the plain-grid presentation of this same view.
@ObservedObject private var gamepadManager = GamepadManager.shared
@AppStorage(DefaultsKey.gamepadUIEnabled) private var gamepadUIEnabled = true
@AppStorage(DefaultsKey.gamepadUIMode) private var gamepadUIMode =
GamepadUIEnvironment.modeWhenConnected
private var gamepadUIActive: Bool {
GamepadUIEnvironment.isActive(
gamepadConnected: gamepadManager.active != nil, enabledSetting: gamepadUIEnabled)
gamepadConnected: gamepadManager.active != nil, enabledSetting: gamepadUIEnabled,
mode: gamepadUIMode)
}
#endif
@@ -78,6 +81,16 @@ struct LibraryView: View {
}
}
#endif
#if os(iOS) || os(macOS) || os(tvOS)
// Published HERE, not just inside the coverflow, because the coverflow is only one of
// four things this view renders: the loading spinner, the error state and the empty
// state sit above it, as do the navigation title and toolbar. On iOS those are wrapped
// by GamepadLibraryScreen, which inks the whole thing; tvOS and macOS present this view
// directly in a NavigationStack, so under a pale palette every one of them kept the
// system's own (dark, on an Apple TV) chrome over a light field. Off when the gamepad
// UI isn't drawing the plain grid belongs to the system background.
.gamepadPaletteInk(gamepadUIActive)
#endif
}
@ViewBuilder private var content: some View {
@@ -81,6 +81,9 @@ struct GamepadSettingsView: View {
@AppStorage(DefaultsKey.hudPlacement) private var hudPlacement = HUDPlacement.topTrailing.rawValue
@AppStorage(DefaultsKey.libraryEnabled) private var libraryEnabled = true
@AppStorage(DefaultsKey.gamepadUIEnabled) private var gamepadUIEnabled = true
/// When the switch above takes over the row is only built while it is on.
@AppStorage(DefaultsKey.gamepadUIMode) private var gamepadUIMode =
GamepadUIEnvironment.modeWhenConnected
/// The gamepad UI's background colour family the backdrop BEHIND this screen re-colours as
/// the row steps, which is why the picker lives here and not in a sheet.
@AppStorage(DefaultsKey.uiPalette) private var paletteID = "violet"
@@ -659,6 +662,21 @@ struct GamepadSettingsView: View {
detail: "Turn off to use the touch interface even with a controller connected.",
value: $gamepadUIEnabled),
]
// WHEN the switch above takes over. Built only while it is on: with the switch off this
// screen is unreachable in the first place (no gamepad UI to open it from), so a row
// that decides nothing would exist purely to be found in a screenshot.
if gamepadUIEnabled, let at = list.firstIndex(where: { $0.id == "gamepadUI" }) {
list.insert(
choiceRow(
id: "gamepadUIMode", tab: .interface, icon: "gamecontroller.circle",
label: "Show it",
detail: "With a controller: the touch interface comes back when the last one "
+ "disconnects. Always keeps this layout either way — for a device that "
+ "lives on a TV.",
options: SettingsOptions.gamepadUIModes, current: gamepadUIMode
) { gamepadUIMode = $0 },
at: at + 1)
}
#if os(macOS)
// The windowed safe-present toggle slots in after "Smoothness buffer" (staying inside
// the Video tab) macOS only, mirroring the touch SettingsView's Presentation row
@@ -707,6 +725,14 @@ struct GamepadSettingsView: View {
at: anchor + 1)
}
#endif
// The smoothness buffer only decides anything under Smoothness. Every other settings
// surface touch, tvOS, the GTK and WinUI shells hides it under Lowest latency; this
// screen alone left it live and steppable, which is a row that thuds or silently stores
// a value nothing reads. Removed here rather than omitted from the literal above so the
// macOS safe-present insertion can still anchor on it.
if presentPriority != "smooth" {
list.removeAll { $0.id == "smoothBuffer" }
}
return list + profileRows
}
@@ -53,6 +53,14 @@ enum SettingsOptions {
static let hudPlacements: [(label: String, tag: String)] =
HUDPlacement.allCases.map { ($0.label, $0.rawValue) }
/// When the gamepad UI takes over (`DefaultsKey.gamepadUIMode`) only meaningful while
/// `gamepadUIEnabled` is on, so every surface that offers it hides the row when the switch
/// is off rather than showing a picker that decides nothing.
static let gamepadUIModes: [(label: String, tag: String)] = [
("With a controller", GamepadUIEnvironment.modeWhenConnected),
("Always", GamepadUIEnvironment.modeAlways),
]
/// Presentation intent (`DefaultsKey.presentPriority` the 2026-07 rebuild that replaced
/// the visible stage picker with intent; see SessionPresenter's PresentPriority and
/// design/apple-presentation-rebuild.md). The stage ladder survives only as the hidden
@@ -724,11 +724,24 @@ extension SettingsView {
#endif
#if !os(tvOS)
if !inProfileScope {
described("With a controller connected, the host list and library switch to a "
+ "controller-friendly layout — larger focus targets, a swipeable cover "
+ "browser.") {
described("The host list and library switch to a controller-friendly layout — "
+ "larger focus targets, a swipeable cover browser.") {
Toggle("Gamepad-optimized browsing", isOn: $gamepadUIEnabled)
}
// Only meaningful while the switch above is on, so it is HIDDEN rather than
// disabled when it isn't: a picker whose every option decides nothing is worse
// than no picker, and this Section is short enough that nothing jumps far.
if gamepadUIEnabled {
described("With a controller: the touch interface comes back when the last "
+ "one disconnects. Always keeps the controller-friendly layout either "
+ "way — for a device that lives on a TV.") {
Picker("Show it", selection: $gamepadUIMode) {
ForEach(SettingsOptions.gamepadUIModes, id: \.tag) { option in
Text(option.label).tag(option.tag)
}
}
}
}
}
#endif
#if DEBUG && !os(tvOS)
@@ -75,6 +75,13 @@ struct SettingsView: View {
@AppStorage(DefaultsKey.hudPlacement) var hudPlacement = HUDPlacement.topTrailing.rawValue
@ObservedObject var gamepads = GamepadManager.shared
@AppStorage(DefaultsKey.gamepadUIEnabled) var gamepadUIEnabled = true
/// When the switch above takes over read (and shown) only while it is on.
@AppStorage(DefaultsKey.gamepadUIMode) var gamepadUIMode =
GamepadUIEnvironment.modeWhenConnected
/// The gamepad UI's background palette. Edited here on tvOS only (see `tvBody`) every other
/// platform reaches it through the gamepad settings screen, which an Apple TV without a
/// controller cannot open.
@AppStorage(DefaultsKey.uiPalette) var uiPalette = "violet"
@AppStorage(DefaultsKey.autoWake) var autoWakeEnabled = true
@AppStorage(DefaultsKey.backgroundKeepAlive) var backgroundKeepAlive = false
@AppStorage(DefaultsKey.backgroundTimeoutMinutes) var backgroundTimeoutMinutes = 10
@@ -488,6 +495,22 @@ struct SettingsView: View {
TVSelectionRow(
title: "Gamepad-optimized browsing",
options: [("On", "on"), ("Off", "off")], selection: gamepadUIEnabledTag)
// Hidden while the switch above is off see the touch settings' identical gate.
if gamepadUIEnabled {
TVSelectionRow(
title: "Show it",
options: SettingsOptions.gamepadUIModes, selection: $gamepadUIMode)
// The Apple TV's ONLY route to the shared `ui_palette`. Everywhere else the
// Background row lives on the gamepad settings screen, which is reached from
// the gamepad launcher and on tvOS that launcher needs an extended-profile
// controller, so an Apple TV driven by the Siri Remote alone could not reach
// the palettes at all. It belongs beside "Show it" because both describe the
// same interface: this row is what that interface looks like once it is up.
TVSelectionRow(
title: "Background",
options: GamepadPalette.all.map { (label: $0.name, tag: $0.id) },
selection: $uiPalette)
}
tvCaption(Self.controllersFooter)
NavigationLink("About") { AboutView() }
.padding(.top, 8)
@@ -95,24 +95,19 @@ private struct ConsoleGlass<S: Shape>: ViewModifier {
private var materialWash: Color { ink.glass(ink.isLight ? 0.55 : 0.40) }
func body(content: Content) -> some View {
// The scheme goes on the WHOLE modified view, not just the fill inside `.background {}`.
// Scoped to the fill it frosts the material correctly and stops there, so a system colour
// in the row's own content (a `.secondary` label, a `.bordered` button) still resolved
// against the device appearance which is how the pale palettes came out light-on-light
// on tvOS, whose appearance is always Dark. The 26 branch had it right all along; the
// tvOS and pre-26 branches were the odd ones out.
#if os(tvOS)
// ALWAYS the material fallback on tvOS: the gamepad settings list is 15+ of these
// surfaces, and live Liquid Glass per row made the whole screen visibly laggy on the
// Apple TV's GPU (same class of call GlassProminentButton already makes glass fights
// the 10-foot platform). The wash and tint ride overlays two flat fills, no GPU cost.
content.background {
shape.fill(.ultraThinMaterial)
.environment(\.colorScheme, scheme)
.overlay { shape.fill(materialWash) }
.overlay {
if let tint { shape.fill(tint) }
}
}
#else
if #available(iOS 26, macOS 26, *) {
content.glassEffect(glass, in: shape).environment(\.colorScheme, scheme)
} else {
content.background {
content
.background {
shape.fill(.ultraThinMaterial)
.environment(\.colorScheme, scheme)
.overlay { shape.fill(materialWash) }
@@ -120,6 +115,21 @@ private struct ConsoleGlass<S: Shape>: ViewModifier {
if let tint { shape.fill(tint) }
}
}
.environment(\.colorScheme, scheme)
#else
if #available(iOS 26, macOS 26, *) {
content.glassEffect(glass, in: shape).environment(\.colorScheme, scheme)
} else {
content
.background {
shape.fill(.ultraThinMaterial)
.environment(\.colorScheme, scheme)
.overlay { shape.fill(materialWash) }
.overlay {
if let tint { shape.fill(tint) }
}
}
.environment(\.colorScheme, scheme)
}
#endif
}
@@ -173,11 +183,14 @@ private struct ConsoleGlassBackground<S: Shape>: ViewModifier {
in: shape)
.environment(\.colorScheme, scheme)
} else {
content.background {
shape.fill(.regularMaterial)
.environment(\.colorScheme, scheme)
.overlay { shape.fill(ink.glass(ink.isLight ? 0.55 : 0.40)) }
}
// Same hoist as ConsoleGlass: the content needs the scheme too, not only the frost.
content
.background {
shape.fill(.regularMaterial)
.environment(\.colorScheme, scheme)
.overlay { shape.fill(ink.glass(ink.isLight ? 0.55 : 0.40)) }
}
.environment(\.colorScheme, scheme)
}
}
}
@@ -0,0 +1,129 @@
// "The audio output moved under us" the one signal `SessionAudio` needs to survive a device
// change, and the one piece of it that can be tested without a stream.
//
// Split out of SessionAudio deliberately. An end-to-end test of the recovery needs a live session,
// which needs a host, and punktfunk-host does not build on macOS so the wiring that matters most
// (is the observer actually installed? does the identity check let the notification through?) would
// otherwise ship unverified, and a silent failure in it costs the session ALL of its audio. On its
// own this can be pointed at the real hardware from a unit test: see AudioDeviceWatcherTests.
//
// What it does NOT own: anything with session semantics. The iOS route-change steer and the
// media-services-reset re-activation stay in SessionAudio, next to the AVAudioSession they act on.
import AVFoundation
import os
#if os(macOS)
import CoreAudio
#endif
private let log = Logger(subsystem: "io.unom.punktfunk", category: "audio")
final class AudioDeviceWatcher {
/// Why the owner is being told. Only for the log line every reason leads to the same
/// question, "is playback still on the device it should be on".
enum Reason: String {
/// An engine stopped itself because its IO hardware changed underneath it.
case engineConfiguration = "the audio hardware configuration changed"
/// The system's default output device moved (macOS).
case defaultOutputDevice = "the default output device changed"
}
/// Does this configuration change belong to an engine the session still owns? A retired engine
/// posts one last change as it is torn down, and other AVAudioEngines in the process are not
/// ours to restart.
private let isOurs: (AnyObject?) -> Bool
/// Delivered on the main queue.
private let onChange: (Reason) -> Void
private let lock = NSLock()
private var configObserver: NSObjectProtocol?
#if os(macOS)
private var defaultOutputListener: AudioObjectPropertyListenerBlock?
#endif
init(isOurs: @escaping (AnyObject?) -> Bool, onChange: @escaping (Reason) -> Void) {
self.isOurs = isOurs
self.onChange = onChange
}
deinit { stop() }
/// Idempotent.
func start() {
lock.lock()
let already = configObserver != nil
lock.unlock()
guard !already else { return }
let token = NotificationCenter.default.addObserver(
forName: .AVAudioEngineConfigurationChange, object: nil, queue: nil
) { [weak self] note in
// Posted from whatever thread the IO unit noticed on. The engine is the notification's
// object; it is only ever compared by identity, never resurrected.
let posted = note.object as AnyObject?
DispatchQueue.main.async {
guard let self, self.isOurs(posted) else { return }
self.onChange(.engineConfiguration)
}
}
lock.lock()
configObserver = token
lock.unlock()
#if os(macOS)
// The engine notification is the direct signal, but it is delivered BY an engine useless
// in the two places it is needed most: after a rebuild that could not start (no engine left
// to notify anyone) and on an engine topology whose notification behaviour is unverified
// (the voice-processing engine, which is the DEFAULT macOS configuration and which no Mac
// here can even initialize). The HAL is told either way.
let block: AudioObjectPropertyListenerBlock = { [weak self] _, _ in
self?.onChange(.defaultOutputDevice) // on the main queue registered against it below
}
var address = Self.defaultOutputAddress()
let status = AudioObjectAddPropertyListenerBlock(
AudioObjectID(kAudioObjectSystemObject), &address, DispatchQueue.main, block)
guard status == noErr else {
log.warning("""
could not watch the default output device (\(status)) an output device change \
mid-stream may need a reconnect
""")
return
}
lock.lock()
defaultOutputListener = block
lock.unlock()
#endif
}
/// Idempotent, and safe from any thread. After it returns, no further `onChange` is delivered
/// except one already in flight on the main queue which the owner's own stopped-flag catches.
func stop() {
lock.lock()
let token = configObserver
configObserver = nil
#if os(macOS)
let listener = defaultOutputListener
defaultOutputListener = nil
#endif
lock.unlock()
if let token { NotificationCenter.default.removeObserver(token) }
#if os(macOS)
guard let listener else { return }
var address = Self.defaultOutputAddress()
AudioObjectRemovePropertyListenerBlock(
AudioObjectID(kAudioObjectSystemObject), &address, DispatchQueue.main, listener)
#endif
}
#if os(macOS)
/// Freshly built per call rather than held in a mutable static: the HAL takes the address
/// `inout` and copies it, so there is nothing to share and a shared one would only be a
/// mutable global.
private static func defaultOutputAddress() -> AudioObjectPropertyAddress {
AudioObjectPropertyAddress(
mSelector: kAudioHardwarePropertyDefaultOutputDevice,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
}
#endif
}
@@ -43,8 +43,21 @@ public enum AudioDevices {
}
private static func defaultInputDevice() -> AudioDeviceID? {
systemDevice(kAudioHardwarePropertyDefaultInputDevice)
}
/// The device the system is currently playing to what an engine with no pinned speaker UID
/// follows, and so what `SessionAudio` compares its live output device against when the
/// default moves (AirPods in or out, a headset unplugged).
static func defaultOutputDevice() -> AudioDeviceID? {
systemDevice(kAudioHardwarePropertyDefaultOutputDevice)
}
private static func systemDevice(
_ selector: AudioObjectPropertySelector
) -> AudioDeviceID? {
var address = AudioObjectPropertyAddress(
mSelector: kAudioHardwarePropertyDefaultInputDevice,
mSelector: selector,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
var dev = AudioDeviceID(0)
@@ -16,11 +16,17 @@ import os
/// (`punktfunk_core::audio::JitterPolicy`): a slow depth average that sits above target for a
/// sustained window sheds ONE 5 ms frame with a crossfade, and the hard cap is only a backstop.
///
/// **Adaptive depth.** The target is a floor, not a constant: repeated genuine underruns grow it
/// a step at a time (`noteRead`, mirroring `JitterPolicy::note_read`) up to `maxTargetMS`, and a
/// long quiet spell relaxes it back toward the base so a session on Wi-Fi that bunches arrivals
/// deepens until it stops crackling, while a clean LAN keeps the tight base latency. Keep the
/// constants here in step with `JitterTuning.COREAUDIO`.
/// **Adaptive depth.** The target is a floor, not a constant: a NEAR-MISS a read served with
/// less than one frame left over grows it a step BEFORE anything was audible, repeated genuine
/// underruns grow it too (`noteRead`, mirroring `JitterPolicy::note_read`) up to `maxTargetMS`,
/// and a long quiet spell relaxes it back toward the base so a session on Wi-Fi that bunches
/// arrivals deepens until it stops crackling, while a clean LAN keeps the tight base latency.
/// Growth only raises a promise; the one thing that re-banks real depth is a re-prime, so an
/// underrun while the ring is HOLLOW (depth average far below the target) re-primes at once,
/// spending the click it already cost on the whole refill. Every shrink is armed as a PROBE:
/// answered by an underrun or near-miss within its window, it is undone on the spot, and a
/// failed sync-driven shrink is not retried for a growing backoff. Keep the constants here in
/// step with `JitterTuning.COREAUDIO`.
///
/// **A/V sync.** On top of all that the depth can be STEERED, by `setSyncTarget` from the drain
/// thread's `AvSync` because a ring that is the right depth for the link is not thereby the
@@ -58,9 +64,29 @@ final class AudioRing: @unchecked Sendable {
/// target normally relaxes only after a long spell because, absent other evidence, the only
/// thing that can justify giving up hard-won slack is time; a sync request IS that evidence
/// a measurement saying the extra depth is costing alignment right now so a smaller target
/// gets tested sooner. Wrong guesses are cheap and self-correcting (one underrun and the
/// growth path takes it straight back). Mirrors `SHRINK_QUIET_SYNC_MS`.
/// gets tested sooner. Mirrors `SHRINK_QUIET_SYNC_MS`.
private static let shrinkQuietSyncMS = 5_000
/// Post-read depth below which a served callback counts as a NEAR-MISS: the device got its
/// samples, but with less than one protocol frame left in hand the same evidence as an
/// underrun, except nobody heard it yet, so the target grows BEFORE the click instead of
/// after the third one. Mirrors `NEAR_MISS_MARGIN_MS`.
private static let nearMissMarginMS = frameMS
/// How long a shrink remains a PROBE, in consumed audio: an underrun or near-miss inside
/// this window means the shrink was wrong, and the previous target is restored at once.
/// Mirrors `SHRINK_PROBE_MS`.
private static let shrinkProbeMS = 5_000
/// How long a failed probe keeps the sync loop from driving another shrink without it the
/// loop pays an audible starvation event every `shrinkQuietSyncMS` on any link whose jitter
/// genuinely needs the depth, forever. Doubles per consecutive failure, capped; a probe that
/// survives its window resets it. Mirror `SYNC_BACKOFF_MS` / `SYNC_BACKOFF_MAX_MS`.
private static let syncBackoffMS = 60_000
private static let syncBackoffMaxMS = 480_000
/// A ring is HOLLOW when its depth AVERAGE sits this far below the target: growth only ever
/// raises the promise, and the one thing that re-banks real depth is a re-prime so an
/// underrun in a hollow ring re-primes AT ONCE, spending the click it already cost on the
/// whole refill instead of riding the knife edge one click per bunching period. Mirrors
/// `DEPRIME_DEBT_MS`.
private static let deprimeDebtMS = growStepMS
private var buf: [Float]
private var readIdx = 0
@@ -87,6 +113,24 @@ final class AudioRing: @unchecked Sendable {
/// `nil` the default, and what an un-wired session keeps reproduces the pre-sync
/// behaviour exactly, so this ring could adopt sync without the other three diverging.
private var syncTarget: Int?
/// This read was served with less than `nearMissMarginMS` left over (set in `read`,
/// consumed by `noteRead`).
private var nearMiss = false
/// A near-miss already grew the target this window one step per window, so a bunching
/// episode (a RUN of consecutive near-misses while the ring refills) buys one measured
/// step, not a sprint to the ceiling.
private var nearMissGrown = false
/// The depth average runs a `deprimeDebtMS` debt against the target (set in `read`): an
/// underrun should re-prime at once instead of waiting out the hysteresis.
private var hollow = false
/// Interleaved samples left in the current shrink-probe window (0 = no probe outstanding).
private var probeRun = 0
/// The live target before the probed shrink, restored if the probe fails.
private var probePrevTarget = 0
/// Interleaved samples before the sync loop may drive another shrink (0 = allowed now).
private var syncBackoffRun = 0
/// Length of the NEXT backoff, in ms doubles per consecutive failed probe, capped.
private var syncBackoffLenMS = AudioRing.syncBackoffMS
/// The sync loop's smoothed offset in ms, STORED not computed: the ring owns the depth but has
/// no timestamps, so the drain thread (which has both a packet's `pts_ns` and the video leg)
/// hands the number back for reporting. Mirrors `NativeClient::audio_av_offset_ms`.
@@ -121,8 +165,15 @@ final class AudioRing: @unchecked Sendable {
/// then return the CAP i.e. quietly below the continuity floor, inverting the very ordering
/// this exists to guarantee, on exactly the awkward hardware it exists to survive. (Rust's
/// `Ord::clamp` announces the same condition by panicking; Swift would just get it wrong.)
private var target: Int {
let floor = max(targetLive, renderQuantum + Self.frameMS * perMS)
private var target: Int { target(lift: renderQuantum) }
/// The effective target with an explicit quantum lift. The property above uses the high-water
/// `renderQuantum` (priming must survive the biggest callback seen); the hollow check in
/// `read` passes the CURRENT callback instead, mirroring the Rust side's `want` a one-off
/// oversized read would otherwise inflate the debt threshold forever and turn the very next
/// late packet into a full re-prime.
private func target(lift quantum: Int) -> Int {
let floor = max(targetLive, quantum + Self.frameMS * perMS)
guard let want = syncTarget else { return floor }
let cap = max(Self.hardCapMS * perMS, floor)
return min(max(want, floor), cap)
@@ -211,12 +262,24 @@ final class AudioRing: @unchecked Sendable {
if available >= target {
primed = true
emptyReads = 0
// The refill just banked this much: seed the average with it rather than letting
// it climb from wherever the drought left it a freshly-primed ring would
// otherwise read as hollow for the EWMA's whole settling time, and the FIRST
// late packet would re-prime a ring that is actually full.
depthAvg = Double(available)
} else {
for i in 0..<count { out[i] = 0 }
return
}
}
// Hollow: the depth AVERAGE runs a debt against the target the promise has been raised
// but the depth was never re-banked (see `deprimeDebtMS`). Judged on the average, not
// this instant: a single late packet empties the ring for a callback without making it
// hollow, and must keep the consecutive-empties hysteresis. Lifted by THIS callback's
// size, not the high-water quantum see `target(lift:)`.
hollow = depthAvg + Double(Self.deprimeDebtMS * perMS) < Double(target(lift: count))
// Drift correction: shed exactly one frame, crossfaded, once the AVERAGE has sat above
// the threshold for the sustain window. Anything shorter is jitter and must be left alone.
if depthAvg > Double(target + Self.shedExcessMS * perMS) {
@@ -240,6 +303,9 @@ final class AudioRing: @unchecked Sendable {
if n < count {
for i in n..<count { out[i] = 0 }
}
// Near-miss: served in full, but with less than one frame left over the next callback
// starves unless a packet lands within one frame time.
nearMiss = n == count && writeIdx - readIdx < Self.nearMissMarginMS * perMS
noteRead(ranShort: n < count, count: count)
}
@@ -254,32 +320,84 @@ final class AudioRing: @unchecked Sendable {
if windowRun >= Self.growWindowMS * perMS {
windowRun = 0
underrunsInWindow = 0
nearMissGrown = false
}
syncBackoffRun = max(0, syncBackoffRun - count)
var restored = false
if probeRun > 0 {
probeRun = max(0, probeRun - count)
if ranShort || nearMiss {
// The probe FAILED: the link answered a shrink with (nearly) starving the ring.
// Take the depth straight back re-learning it three audible underruns at a
// time is what made the sync-vs-growth tug-of-war audible and keep the sync
// loop from probing again for a while, doubling per consecutive failure. The
// residual A/V offset is reported instead; continuity outranks sync. The
// restore CONSUMES this event as growth evidence: it answered a depth the ring
// is no longer at, so growing past the proven target on top would overshoot.
probeRun = 0
targetLive = max(targetLive, probePrevTarget)
syncBackoffRun = syncBackoffLenMS * perMS
syncBackoffLenMS = min(syncBackoffLenMS * 2, Self.syncBackoffMaxMS)
restored = true
} else if probeRun == 0 {
// Survived the whole window: the shallower depth is genuinely safe here, so the
// next probe starts from a clean slate.
syncBackoffLenMS = Self.syncBackoffMS
}
}
if ranShort {
quietRun = 0
emptyReads += 1
underrunCount += 1
if emptyReads >= Self.deprimeAfter {
if emptyReads >= Self.deprimeAfter || hollow {
// The consecutive-empties hysteresis protects a FULL ring from one late packet.
// A hollow ring is the opposite case: the target has been raised but the depth
// never re-banked (growth is a promise; only a re-prime cashes it), and riding
// that out is a click per bunching period, forever. The click just heard has
// already paid for the refill take it now.
primed = false
emptyReads = 0
}
underrunsInWindow += 1
if !restored {
underrunsInWindow += 1
}
if underrunsInWindow >= Self.growUnderruns {
underrunsInWindow = 0
windowRun = 0
targetLive = min(targetLive + Self.growStepMS * perMS, Self.maxTargetMS * perMS)
}
} else if nearMiss {
// Came within one frame of an underrun the same evidence as one, heard by no one.
// Growing here, BEFORE the click, is what "no audible jitter" means: waiting for
// the third audible underrun means the user heard two. One step per window (a
// bunching episode is a RUN of near-misses while the ring refills, and must buy one
// measured step, not a sprint to the ceiling); if it worsens into real underruns
// the path above takes over. A near-miss is pressure, not quiet.
quietRun = 0
emptyReads = 0
if !nearMissGrown, !restored {
nearMissGrown = true
targetLive = min(targetLive + Self.growStepMS * perMS, Self.maxTargetMS * perMS)
}
} else {
emptyReads = 0
quietRun += count
// Without a sync request, time is the only evidence that hard-won slack is no longer
// needed, so a grown target waits out the long window. A request for less IS evidence,
// and without this branch a ring that ratcheted to the ceiling during a transient would
// hold audio a ceiling's worth late for minutes after the cause had gone.
let quietNeeded = syncWantsLess ? Self.shrinkQuietSyncMS : Self.shrinkQuietMS
// hold audio a ceiling's worth late for minutes after the cause had gone. Every shrink
// is armed as a PROBE answered by an underrun or near-miss it is undone at once (see
// above), and a failed sync-driven guess is not retried for a backoff.
let syncShrink = syncWantsLess && syncBackoffRun == 0
let quietNeeded = syncShrink ? Self.shrinkQuietSyncMS : Self.shrinkQuietMS
if quietRun >= quietNeeded * perMS {
quietRun = 0
let prev = targetLive
targetLive = max(targetLive - Self.growStepMS * perMS, Self.targetMS * perMS)
if targetLive < prev {
probeRun = Self.shrinkProbeMS * perMS
probePrevTarget = prev
}
}
}
}
@@ -21,6 +21,10 @@
//
// Devices are chosen by UID ("" = system default: the engine is then never pinned to a
// concrete device and follows default-device changes).
//
// Surviving the hardware. An AVAudioEngine does NOT follow the audio hardware: when the output
// device changes underneath a running engine, the engine stops itself and stays stopped. The
// session therefore watches for that and rebuilds its engines see "Device changes" below.
import AVFoundation
import os
@@ -79,6 +83,48 @@ public final class SessionAudio {
/// session's activate.
private static let sessionQueue = DispatchQueue(label: "io.unom.punktfunk.audio.session")
#endif
#if !os(macOS)
/// Token for the route-change observer: it revives an engine the route change stopped, and on
/// iOS re-applies the earpiece steer (see `installRouteObserver`). Guarded by `stateLock`.
private var routeObserver: NSObjectProtocol?
/// Token for the media-services-reset observer the audio server restarting takes the
/// session's configuration and every engine with it. Guarded by `stateLock`.
private var mediaResetObserver: NSObjectProtocol?
#endif
// MARK: - Device changes (see `installDeviceChangeRecovery`)
/// What `start()` was asked for, so a rebuild can put back the SAME topology the session was
/// started with. Main-thread confined, like the start paths that read it.
private var startConfig: StartConfig?
private struct StartConfig {
let speakerUID: String
let micUID: String
let micChannel: Int
let micEnabled: Bool
let echoCancel: Bool
}
/// Watches the hardware for us (see `AudioDeviceWatcher`). Guarded by `stateLock`.
private var deviceWatcher: AudioDeviceWatcher?
/// Whether the engines have been built at least once. Distinguishes "not started yet" (iOS
/// starts asynchronously) from "started and dead", which is what the recovery may act on.
/// Main-thread confined.
private var enginesAttempted = false
/// A rebuild is already on the main queue one device switch produces a burst of triggers
/// and they must collapse into one restart. Main-thread confined.
private var rebuildQueued = false
/// `systemUptime` of the last rebuild, so a device that renegotiates in a loop cannot spin
/// the session. Main-thread confined.
private var lastRebuildAt: TimeInterval = 0
/// Let the burst of triggers from one switch land before rebuilding.
private static let rebuildDebounce: TimeInterval = 0.15
/// Floor between two rebuilds.
private static let rebuildFloor: TimeInterval = 0.5
/// Retries when a rebuild's `start()` loses the race with a device that is still going away
/// (0.3 s, 0.6 s, 1.2 s). A failed rebuild leaves no engine to post the next notification,
/// so this ladder and, on macOS, the HAL listener is all that stands between a mistimed
/// switch and a silent session.
private static let rebuildAttempts = 3
public init(connection: PunktfunkConnection) {
self.connection = connection
@@ -89,6 +135,15 @@ public final class SessionAudio {
/// Engine teardown still belongs to stop().
deinit {
flag.stop()
// The observers only hold self weakly, so we can be deinited with them still registered;
// drop them here too rather than leaking them when an owner skips stop().
deviceWatcher?.stop()
#if !os(macOS)
if let routeObserver { NotificationCenter.default.removeObserver(routeObserver) }
if let mediaResetObserver {
NotificationCenter.default.removeObserver(mediaResetObserver)
}
#endif
}
/// Start playback (and, if enabled+authorized, the mic uplink). Empty UIDs = system default
@@ -108,6 +163,12 @@ public final class SessionAudio {
videoLatency: LatencyMeter? = nil
) {
self.videoLatency = videoLatency
// Before any engine exists: the recovery watches the hardware, not the engines, and the
// config it rebuilds from has to be recorded whether or not this start succeeds.
startConfig = StartConfig(
speakerUID: speakerUID, micUID: micUID, micChannel: micChannel,
micEnabled: micEnabled, echoCancel: echoCancel)
installDeviceChangeRecovery(micEnabled: micEnabled)
#if os(macOS)
// No AVAudioSession on macOS start the engines directly (caller's thread, as before).
startEngines(
@@ -138,11 +199,29 @@ public final class SessionAudio {
do {
#if os(iOS)
if micEnabled {
// .defaultToSpeaker: .playAndRecord otherwise routes to the iPhone EARPIECE; only
// affects the built-in route (headphones/BT still win).
// NO .defaultToSpeaker here, deliberately. It reads like "prefer the speaker over
// the earpiece", and the comment that used to sit here claimed headphones and
// Bluetooth still won. That is true of WIRED headphones and false of Bluetooth
// a cable is the one way to test this and see the right answer. It is an
// OVERRIDE, and it outranks an A2DP route: with it set, every Bluetooth headset
// lost the stream to the phone's own speaker. That is the 0.25 field report ("no
// audio over Bluetooth ... plays through speakers if Mic input is enabled") mic
// and echo cancellation both default to ON, so this branch is the DEFAULT path
// and every Bluetooth listener hit it; turning the mic off was the accidental
// workaround, because that lands on `.playback` below, which routes to A2DP
// happily.
//
// The earpiece problem it was reaching for is real, so it is solved after
// activation instead, against the route we were ACTUALLY given
// see `steerBuiltInOutputToSpeaker`.
//
// `.allowBluetoothA2DP` alone, also deliberately: adding `.allowBluetooth` would
// make a headset's MIC usable, but it buys that by dragging the whole route onto
// HFP/SCO and collapsing game audio to narrowband. High-quality A2DP output plus
// the built-in mic is the better trade for a game-streaming client.
try session.setCategory(
.playAndRecord, mode: .default,
options: [.allowBluetoothA2DP, .defaultToSpeaker])
options: [.allowBluetoothA2DP])
// Uplink latency: ask for 5 ms IO quanta at the wire rate (the default ~10-23 ms
// quantum is most of the mic path's burst latency). Best-effort the hardware
// has the final word (a Bluetooth route will ignore both), and whatever quantum
@@ -156,18 +235,85 @@ public final class SessionAudio {
try session.setCategory(.playback, mode: .default)
#endif
try session.setActive(true)
#if os(iOS)
// Only the `.playAndRecord` session can land on the earpiece, and only it accepts an
// output override so the mic-off (`.playback`) path deliberately does neither.
// (The route OBSERVER that re-applies this per route is installed by
// `installDeviceChangeRecovery`, for every session a `.playback` session steers
// nothing but still has engines a route change can stop.)
if micEnabled { steerBuiltInOutputToSpeaker(session) }
#endif
} catch {
log.warning("AVAudioSession setup failed: \(error.localizedDescription)")
}
}
#endif
#if os(iOS)
/// `.playAndRecord` parks the BUILT-IN output on the earpiece right for a phone call,
/// useless for a game. Move it to the speaker, but ONLY when the route we were actually given
/// is the receiver: anything external (Bluetooth, wired, CarPlay, AirPlay) is left strictly
/// alone. That "look first" is the whole difference between this and the `.defaultToSpeaker`
/// option it replaced, which forced the speaker unconditionally and so beat Bluetooth.
///
/// Idempotent and cheap, so the route observer can simply call it again.
private func steerBuiltInOutputToSpeaker(_ session: AVAudioSession) {
// An override already in force shows up as `.builtInSpeaker`, not `.builtInReceiver`, so
// re-running this never fights its own previous result.
guard session.currentRoute.outputs.contains(where: { $0.portType == .builtInReceiver })
else { return }
do {
try session.overrideOutputAudioPort(.speaker)
} catch {
log.warning("could not move audio off the earpiece: \(error.localizedDescription)")
}
}
#endif
#if !os(macOS)
/// Routes change under a live session: a headset connects mid-stream, or disconnects and hands
/// the stream back to the built-in output. Two things follow from that.
///
/// iOS drops an output override whenever the route changes which is what lets a newly-
/// connected headset win so the earpiece steer is a property of the CURRENT route and has to
/// be re-applied per route. Without it, dropping Bluetooth mid-stream lands the game on the
/// earpiece.
///
/// And on every platform a route change can take the engines down with it (see
/// `installDeviceChangeRecovery`), which is why this is installed for `.playback` sessions and
/// on tvOS too, where there is no earpiece to steer away from.
private func installRouteObserver() {
let observer = NotificationCenter.default.addObserver(
forName: AVAudioSession.routeChangeNotification,
object: AVAudioSession.sharedInstance(), queue: nil
) { [weak self] _ in
// Arrives on whatever thread AVFoundation posts it from, and the session API blocks
// on the audio server so do the work on the shared session queue, like every
// other call into it.
SessionAudio.sessionQueue.async {
guard let self, !self.flag.isStopped else { return }
#if os(iOS)
self.steerBuiltInOutputToSpeaker(AVAudioSession.sharedInstance())
#endif
DispatchQueue.main.async { self.reviveStoppedEngines("the audio route changed") }
}
}
stateLock.lock()
let stale = routeObserver
routeObserver = observer
stateLock.unlock()
if let stale { NotificationCenter.default.removeObserver(stale) }
}
#endif
/// Build + start the engines combined (voice-processed) or split, per `wantsCombined`
/// with the mic uplink only when enabled + authorized. Main thread (engine setup); on
/// iOS/tvOS the session is already active by the time this runs.
private func startEngines(
speakerUID: String, micUID: String, micChannel: Int, micEnabled: Bool, echoCancel: Bool
) {
enginesAttempted = true // even if every path below fails see `reviveStoppedEngines`
#if os(tvOS)
// No app-accessible microphone input on tvOS playback only.
startPlayback(speakerUID: speakerUID)
@@ -241,24 +387,27 @@ public final class SessionAudio {
public func stop() {
flag.stop() // before taking the engines see stateLock's comment
stateLock.lock()
let capture = captureEngine
captureEngine = nil
let playback = playbackEngine
playbackEngine = nil
let combined = combinedEngine
combinedEngine = nil
let wasDraining = drainStarted
drainStarted = false
let watcher = deviceWatcher
deviceWatcher = nil
#if !os(macOS)
let route = routeObserver
routeObserver = nil
let mediaReset = mediaResetObserver
mediaResetObserver = nil
#endif
stateLock.unlock()
if let capture {
capture.inputNode.removeTap(onBus: 0)
capture.stop()
}
playback?.stop()
if let combined {
combined.inputNode.removeTap(onBus: 0)
combined.stop()
}
// Every watcher goes before the engines do: a device change landing during teardown must
// not schedule a rebuild of a session we are in the middle of releasing. (`flag` already
// guards that, but not arming the trigger is better than catching it.) On iOS this is
// also ahead of the deactivate below, so a route change cannot re-steer a dying session.
watcher?.stop()
#if !os(macOS)
if let route { NotificationCenter.default.removeObserver(route) }
if let mediaReset { NotificationCenter.default.removeObserver(mediaReset) }
#endif
tearDownEngines()
#if !os(macOS)
// Release the session so audio we interrupted (Music, podcasts) gets its resume cue. Like
// activation, setActive is synchronous/blocking run it on the shared serial session queue
@@ -279,6 +428,234 @@ public final class SessionAudio {
}
}
/// Stop and release every engine we own, leaving the ring, the drain thread, the observers and
/// the audio session alone the teardown half shared by `stop()` and a rebuild. Safe from any
/// thread; the engines are taken under the lock before any of them is touched.
private func tearDownEngines() {
stateLock.lock()
let capture = captureEngine
captureEngine = nil
let playback = playbackEngine
playbackEngine = nil
let combined = combinedEngine
combinedEngine = nil
stateLock.unlock()
if let capture {
capture.inputNode.removeTap(onBus: 0)
capture.stop()
}
playback?.stop()
if let combined {
combined.inputNode.removeTap(onBus: 0)
combined.stop()
}
}
// MARK: - Device changes
/// An AVAudioEngine does not follow the audio hardware. When the output device changes under a
/// running engine AirPods taken out of an ear, a headset unplugged, the default switched in
/// System Settings the engine's IO unit sees the new hardware, THE ENGINE STOPS ITSELF, and
/// it posts `AVAudioEngineConfigurationChange`. It stays stopped until somebody starts it
/// again. Nothing here ever did, so from that moment the session rendered silence: no audio on
/// the speakers the stream had just moved to, and none in the AirPods when they went back in
/// (that is a second stop, not a recovery), until the whole stream was restarted. Measured on
/// this exact topology: render callbacks go from ~94/s to zero the instant the default output
/// device changes, and both restarting the same engine and building a fresh one resume them.
///
/// Three triggers feed one rebuild, because no single one of them covers the ground:
///
/// - the engine notification, everywhere the direct signal, but only an engine that still
/// EXISTS can post it, so it cannot report a rebuild that failed to start;
/// - the HAL default-output-device listener, macOS independent of any engine and of the
/// engine's topology. It is what makes the recovery work for the voice-processing engine
/// (mic + echo cancellation, the DEFAULT macOS configuration) without having to assume that
/// a VPIO engine posts the notification the plain one demonstrably does;
/// - the route-change and media-services-reset notifications, iOS/tvOS, where the session and
/// not the device is what moves.
///
/// `micEnabled` only decides whether the mic-bearing session observers are worth installing.
/// Main thread.
private func installDeviceChangeRecovery(micEnabled: Bool) {
stateLock.lock()
let already = deviceWatcher != nil
stateLock.unlock()
guard !already else { return } // a second start() on one SessionAudio: keep the first set
let watcher = AudioDeviceWatcher(
isOurs: { [weak self] posted in self?.ownsEngine(posted) ?? false },
onChange: { [weak self] reason in self?.hardwareMoved(reason) })
stateLock.lock()
deviceWatcher = watcher
stateLock.unlock()
watcher.start()
#if !os(macOS)
installRouteObserver()
installMediaResetObserver(micEnabled: micEnabled)
#endif
}
/// Is `posted` one of the engines this session currently owns? A retired engine posts one last
/// configuration change as it is torn down, and another AVAudioEngine in the process is none of
/// our business identity only, the object is never resurrected.
private func ownsEngine(_ posted: AnyObject?) -> Bool {
stateLock.lock()
defer { stateLock.unlock() }
return posted === playbackEngine || posted === captureEngine || posted === combinedEngine
}
/// The hardware moved (main queue, from `AudioDeviceWatcher`). Both reasons ask the same
/// question is playback still where it should be but they answer it differently: an engine
/// that told us it stopped is definitive, while the default device moving might not concern us
/// at all.
private func hardwareMoved(_ reason: AudioDeviceWatcher.Reason) {
guard !flag.isStopped else { return }
switch reason {
case .engineConfiguration:
scheduleEngineRebuild(reason: reason.rawValue)
case .defaultOutputDevice:
#if os(macOS)
defaultOutputChanged()
#else
break // the watcher only raises this one on macOS
#endif
}
}
/// Restart the engines if and only if playback is down. The conservative trigger: it is
/// what a route change (iOS/tvOS) and the macOS backstop get to do, since a HEALTHY engine
/// that followed the change on its own must not be interrupted for it.
///
/// Gated on a start having been ATTEMPTED rather than on an engine existing, which is the
/// difference between recovering a session whose very first `startPlayback` failed no
/// output device at the moment it connected and leaving it silent for good. On iOS the same
/// flag keeps this from racing the asynchronous start, where no engine yet is normal.
private func reviveStoppedEngines(_ reason: String) {
guard !flag.isStopped, enginesAttempted, !playbackIsLive else { return }
scheduleEngineRebuild(reason: "playback is stopped and \(reason)")
}
/// Is the render side actually running? Both engines can carry it (`combinedEngine` when the
/// voice processor is engaged, `playbackEngine` otherwise). Taken out from under `stateLock`
/// before asking AVAudioEngine anything the lock guards our handles, not the framework.
private var playbackIsLive: Bool {
stateLock.lock()
let playback = playbackEngine
let combined = combinedEngine
stateLock.unlock()
return (playback?.isRunning ?? false) || (combined?.isRunning ?? false)
}
/// Coalesce: one device switch produces a burst the old device leaving, the default moving,
/// the new device settling, and each engine we own posting its own change and one rebuild
/// serves all of it. The floor between rebuilds keeps a device that renegotiates in a loop
/// from spinning the session. Main thread.
private func scheduleEngineRebuild(reason: String) {
guard !rebuildQueued else { return }
rebuildQueued = true
let since = ProcessInfo.processInfo.systemUptime - lastRebuildAt
let delay = max(Self.rebuildDebounce, Self.rebuildFloor - since)
log.info("\(reason) — restarting the audio engines in \(Int(delay * 1000)) ms")
DispatchQueue.main.asyncAfter(deadline: .now() + delay) { [weak self] in
self?.rebuildEngines(attempt: 0)
}
}
/// Put back the topology this session was started with, on whatever hardware is there now.
///
/// A full rebuild rather than a `start()` on the stopped engine, because the mic side has to
/// follow too: `installMicTap` reads the input's live format, and the voice processor
/// renegotiates its own. The RING is deliberately not touched it is the one thing carried
/// across (`makePlaybackChain` reuses it, `startDrain` is idempotent), so the drain thread
/// keeps decoding right through the switch and its overflow policy has already dropped
/// everything that went stale while the engine was down.
private func rebuildEngines(attempt: Int) {
rebuildQueued = false
guard !flag.isStopped, let config = startConfig else { return }
lastRebuildAt = ProcessInfo.processInfo.systemUptime
tearDownEngines()
startEngines(
speakerUID: config.speakerUID, micUID: config.micUID, micChannel: config.micChannel,
micEnabled: config.micEnabled, echoCancel: config.echoCancel)
// Did playback actually come back? A device caught mid-transition can refuse to start, and
// a rebuild that fails leaves no engine to post the next notification so this is the one
// path that must not just give up. (`startEngines` has logged the reason already.)
if playbackIsLive {
log.info("audio engines restarted on the current device")
return
}
guard attempt < Self.rebuildAttempts else {
#if os(macOS)
log.error("""
audio did not come back after the device change the default-output watcher will \
try again when a device appears
""")
#else
log.error("audio did not come back after the route change")
#endif
return
}
rebuildQueued = true // holds off a trigger that would only race this ladder
let delay = Self.rebuildDebounce * Double(1 << (attempt + 1))
DispatchQueue.main.asyncAfter(deadline: .now() + delay) { [weak self] in
self?.rebuildEngines(attempt: attempt + 1)
}
}
#if os(macOS)
/// The system's output device moved. Rebuild only when it actually concerns this session: the
/// engine is gone or stopped, or it is playing to a device that is no longer the one we should
/// be on. Somebody changing the default while we are pinned to a named speaker is none of our
/// business, and rebuilding for it would cost an audible gap for nothing. Main queue (the
/// listener block is registered against it).
private func defaultOutputChanged() {
guard !flag.isStopped, let config = startConfig else { return }
stateLock.lock()
let engine = combinedEngine ?? playbackEngine
stateLock.unlock()
guard let engine, engine.isRunning, let unit = engine.outputNode.audioUnit,
let playingOn = Self.currentDevice(of: unit)
else {
// Nothing is playing. If an engine was expected at all, this is the backstop firing.
reviveStoppedEngines("the default output device moved")
return
}
// Empty UID = follow the system default; a pinned UID only moves if that device itself
// came or went, which `deviceID(forUID:)` reports by resolving to a different ID or none.
let shouldBeOn = config.speakerUID.isEmpty
? AudioDevices.defaultOutputDevice()
: AudioDevices.deviceID(forUID: config.speakerUID)
guard let shouldBeOn, shouldBeOn != playingOn else { return }
scheduleEngineRebuild(reason: "the output device changed under the session")
}
#endif
#if !os(macOS)
/// The audio server can die and restart. It takes the session's configuration and every engine
/// with it, and the documented recovery is to build all of it again the same rebuild a route
/// change uses, with the session activation back in front of it.
private func installMediaResetObserver(micEnabled: Bool) {
let observer = NotificationCenter.default.addObserver(
forName: AVAudioSession.mediaServicesWereResetNotification, object: nil, queue: nil
) { [weak self] _ in
SessionAudio.sessionQueue.async {
guard let self, !self.flag.isStopped else { return }
self.activateAudioSession(micEnabled: micEnabled)
DispatchQueue.main.async {
self.scheduleEngineRebuild(reason: "the audio services were reset")
}
}
}
stateLock.lock()
let stale = mediaResetObserver
mediaResetObserver = observer
stateLock.unlock()
if let stale { NotificationCenter.default.removeObserver(stale) }
}
#endif
/// Silence the mic uplink (no room audio leaves the device) or restore it. THE one muting
/// mechanism: the owner composes its reasons the user's in-stream mute and the background
/// keep-alive's privacy mute into one effective state and passes that here, so neither can
@@ -344,6 +721,21 @@ public final class SessionAudio {
return Stats(bufferMS: s.bufferedMS, avOffsetMS: s.avOffsetMS)
}
#if os(macOS)
/// Whether playback is rendering, and the device it is rendering to. The device-change
/// recovery has exactly one observable signature from outside "running again, on the device
/// the system just moved to" and nothing else here could tell the two halves apart: a
/// stopped engine can still name the old device, and a retargeted one can still be stopped.
/// Used by `AudioDeviceSwitchTests`.
var playbackState: (running: Bool, device: AudioDeviceID?) {
stateLock.lock()
let engine = combinedEngine ?? playbackEngine
stateLock.unlock()
guard let engine else { return (false, nil) }
return (engine.isRunning, engine.outputNode.audioUnit.flatMap(Self.currentDevice(of:)))
}
#endif
// MARK: - Playback (host speaker)
/// The playback jitter ring + the source node draining it shared by the plain playback
@@ -112,6 +112,24 @@ public final class GamepadCapture {
static let escapeChordElements = [
GCInputLeftShoulder, GCInputRightShoulder, GCInputButtonMenu, GCInputButtonOptions,
]
/// The stats-overlay chord: Select + X, one tier per completion (off compact normal
/// detailed off). It exists because a controller in both hands has no other way to the
/// numbers the S combo needs a keyboard and the three-finger tap needs a free screen
/// and on tvOS there is no other way AT ALL, which is what this fixes.
///
/// Built like Android's mic chord (`GamepadRouter.MIC_CHORD`, Select + Y) and deliberately
/// not overlapping `escapeChord`: X is none of its four buttons, so no way of reaching the
/// exit chord passes through this one on the way, and vice versa. Select is a menu button
/// rather than a twitch action, which keeps the pair out of real play. Y is left free so the
/// mic chord can be ported onto it later without moving this one.
static let statsChord: UInt32 = GamepadWire.back | GamepadWire.x
/// `statsChord`'s elements by GameController alias same mirror-the-mask rule (and same
/// invisible failure) as `escapeChordElements`; the same test pins both.
static let statsChordElements = [GCInputButtonOptions, GCInputButtonX]
/// Every element some chord reads what a NON-forwarding slot claims (see `openSlot`). The
/// escape chord's four plus the stats chord's X; Select is shared, so it appears once.
static let chordElements: [String] =
escapeChordElements + statsChordElements.filter { !escapeChordElements.contains($0) }
/// pf-client-core's `DISCONNECT_HOLD` the same 1.5 s on every client.
private static let disconnectHold: TimeInterval = 1.5
/// pf-client-core's `GUIDE_HOLD`: hold Select alone this long the HOST's guide goes
@@ -288,14 +306,15 @@ public final class GamepadCapture {
// the PS button must open the host's Steam overlay. Restored to .enabled on close.
//
// With forwarding OFF none of that applies no press reaches the host, so taking the
// user's screenshot gesture away buys nothing. NARROWED, not skipped: the escape chord
// is still read off this slot, and on tvOS it is the only controller way out of a
// stream, so the chord's own four elements keep their claim. (Menu especially: leave
// its gesture attached on tvOS and the press is the system's the chord would never
// complete and the session would have no controller exit at all.)
// user's screenshot gesture away buys nothing. NARROWED, not skipped: the CHORDS are
// still read off this slot on tvOS the escape chord is the only controller way out of
// a stream, and the stats chord the only way to the overlay so their own elements keep
// their claim. (Menu especially: leave its gesture attached on tvOS and the press is the
// system's the chord would never complete and the session would have no controller
// exit at all.)
let claimed = forwarding
? Array(c.physicalInputProfile.elements.values)
: Self.escapeChordElements.compactMap { c.physicalInputProfile.elements[$0] }
: Self.chordElements.compactMap { c.physicalInputProfile.elements[$0] }
for element in claimed {
element.preferredSystemGestureState = .disabled
}
@@ -437,10 +456,24 @@ public final class GamepadCapture {
let newButtons = raw | (slot.buttons & GamepadWire.guide)
let changed = newButtons ^ slot.buttons
if changed != 0 {
let was = slot.buttons
for bit in GamepadWire.allButtons where changed & bit != 0 {
wire?.send(.gamepadButton(bit, down: newButtons & bit != 0, pad: slot.pad))
}
slot.buttons = newButtons
// The stats chord, edge-triggered on the press that COMPLETES it: one cycle per
// chord rather than one per press, since a third button pressed on top finds the
// mask already complete and can't re-fire it. Read off the wire mask like the escape
// chord, which means a Select the hold-Select gesture has turned into a guide is not
// in it a guide hold can't cycle the overlay on its way past. The buttons still
// forward (the chord is a local overlay change, not an input the host must not see).
if was & Self.statsChord != Self.statsChord,
newButtons & Self.statsChord == Self.statsChord {
// Straight to the shared tier default, like TouchMouse's three-finger tap: every
// reader (the HUD, the Settings pickers, the live session) observes it through
// @AppStorage, so no wiring back to the app is needed.
StatsVerbosity.cycle()
}
}
let newAxes: [Int32] = [
Int32(g.leftThumbstick.xAxis.value * 32767),
@@ -3,20 +3,40 @@
// layouts). A pure function, not a singleton: the reactivity comes from callers already observing
// `GamepadManager.shared` and the `DefaultsKey.gamepadUIEnabled` @AppStorage themselves (the same
// local-read pattern SettingsView already uses for GamepadManager), so this stays the single place
// the two combine without adding a second ObservableObject or an environment key nobody else needs.
// the inputs combine without adding a second ObservableObject or an environment key nobody else needs.
import Foundation
import PunktfunkShared
public enum GamepadUIEnvironment {
/// `enabledSetting` is the user's Settings toggle (`DefaultsKey.gamepadUIEnabled`);
/// `DefaultsKey.gamepadUIMode`: take over only while a controller is attached. The default,
/// and what the switch meant when it was a lone Bool.
public static let modeWhenConnected = "connected"
/// `DefaultsKey.gamepadUIMode`: take over whenever the switch is on, pad or no pad asked
/// for by people driving a TV-connected iPad or a couch Mac, where the console layout is the
/// one they want and the pad is not always awake.
public static let modeAlways = "always"
/// `enabledSetting` is the user's Settings switch (`DefaultsKey.gamepadUIEnabled`) off means
/// the touch/desktop UI, full stop. `mode` is `DefaultsKey.gamepadUIMode`, and only matters
/// once the switch is on: `modeAlways` takes over unconditionally, anything else (including a
/// value a newer client wrote) waits for a controller.
///
/// `gamepadConnected` is `GamepadManager.shared.active != nil` active only once a usable
/// controller is actually attached (a non-extended-profile device leaves `active` nil, which
/// keeps the touch UI). A `Bool` rather than the `DiscoveredController` itself: this function's
/// whole job is the AND, so there's nothing else to inspect, and it keeps the helper testable
/// without a real `GCController` (which XCTest can't construct).
public static func isActive(gamepadConnected: Bool, enabledSetting: Bool) -> Bool {
enabledSetting && (gamepadConnected || forced)
/// keeps the touch UI). A `Bool` rather than the `DiscoveredController` itself: this function
/// has nothing else to inspect, and it keeps the helper testable without a real `GCController`
/// (which XCTest can't construct).
/// `mode` carries no default on purpose: a call site that forgot it would silently strand
/// everyone who picked Always back on "only with a controller", which is exactly the bug
/// this parameter exists to make impossible.
public static func isActive(
gamepadConnected: Bool,
enabledSetting: Bool,
mode: String
) -> Bool {
guard enabledSetting else { return false }
return mode == modeAlways || gamepadConnected || forced
}
/// Dev-only escape hatch (like ContentView's `PUNKTFUNK_AUTOCONNECT`): pretend a controller is
@@ -34,10 +34,26 @@ public final class SiriRemotePointer {
private var heldButtons: Set<UInt32> = []
/// When Back/Menu went down; a release after `disconnectHold` fires the exit.
private var menuDownAt: Date?
/// Counts a held Play/Pause down to `statsHold`; nil when the button is up or already
/// resolved. See `playPauseChanged`.
private var playPauseTimer: Timer?
/// The held Play/Pause has already been spent on a stats cycle, so its release must not also
/// right-click.
private var statsHoldFired = false
/// Trails a delivered right-click tap by `tapPress` to release it see `deliverRightClick`.
private var rightReleaseTimer: Timer?
/// Hold Back/Menu at least this long (then release) to end the session. Shorter than the
/// controller chord's 1.5 s the remote has no way to trip this during gameplay.
private static let disconnectHold: TimeInterval = 1.0
/// Hold Play/Pause this long to cycle the stats overlay instead of right-clicking. It is the
/// remote's only spare button, and on an Apple TV with no controller in the room this is the
/// ONLY route to the numbers (S wants a keyboard, the three-finger tap a touchscreen).
/// Shorter than `disconnectHold`: nothing destructive rides on it.
private static let statsHold: TimeInterval = 0.5
/// pf-client-core's `TAP_PRESS`, borrowed for the deferred right-click: its release trails
/// the press by this much, so the two transitions can't fold into nothing downstream.
private static let tapPress: TimeInterval = 0.05
/// A full edge-to-edge swipe moves the host cursor about this many pixels. The surface is
/// small; two comfortable swipes should cross a 1080p desktop.
private static let pointerScale: Float = 1100
@@ -95,6 +111,9 @@ public final class SiriRemotePointer {
old.buttonX.pressedChangedHandler = nil
old.buttonMenu.pressedChangedHandler = nil
}
// Timers first, then the lift: a tap whose release is still owed is held state, so
// `releaseHeld` below is what sends its button-up.
cancelPlayPause()
releaseHeld()
lastTouch = nil
menuDownAt = nil
@@ -109,12 +128,13 @@ public final class SiriRemotePointer {
micro.dpad.valueChangedHandler = { [weak self] _, x, y in
MainActor.assumeIsolated { self?.touchMoved(x: x, y: y) }
}
// Surface click = left button; Play/Pause = right (the remote's only spare face button).
// Surface click = left button; Play/Pause = right (the remote's only spare face button),
// or held the stats-overlay cycle. See `playPauseChanged`.
micro.buttonA.pressedChangedHandler = { [weak self] _, _, pressed in
MainActor.assumeIsolated { self?.setButton(1, down: pressed) }
}
micro.buttonX.pressedChangedHandler = { [weak self] _, _, pressed in
MainActor.assumeIsolated { self?.setButton(3, down: pressed) }
MainActor.assumeIsolated { self?.playPauseChanged(pressed: pressed) }
}
micro.buttonMenu.pressedChangedHandler = { [weak self] _, _, pressed in
MainActor.assumeIsolated { self?.menuChanged(pressed: pressed) }
@@ -149,6 +169,76 @@ public final class SiriRemotePointer {
connection.send(.mouseButton(button, down: down))
}
/// Play/Pause: a TAP right-clicks, a HOLD (`statsHold`) cycles the stats overlay instead.
///
/// The right button is therefore DEFERRED until the press resolves, rather than going down on
/// contact: once the host has seen a button-down there is no taking it back, and a right
/// button held for half a second is a context menu on every desktop this streams. The shape
/// is the hold-Select gesture's (`GamepadCapture.gestureFiltered`) suppress, then deliver a
/// tap on release or the gesture past the threshold so the two behave alike.
private func playPauseChanged(pressed: Bool) {
if pressed {
statsHoldFired = false
let timer = Timer(timeInterval: Self.statsHold, repeats: false) { [weak self] _ in
Task { @MainActor in self?.statsHoldElapsed() }
}
RunLoop.main.add(timer, forMode: .common)
playPauseTimer?.invalidate()
playPauseTimer = timer
return
}
playPauseTimer?.invalidate()
playPauseTimer = nil
// The hold already spent this press on a cycle its release clicks nothing.
guard !statsHoldFired else {
statsHoldFired = false
return
}
deliverRightClick()
}
/// The threshold passed with Play/Pause still down cycle the overlay and consume the press.
/// Writes the shared `statsVerbosity` default every reader observes through @AppStorage the
/// same cycle as S, the three-finger tap and the controller's Select + X.
private func statsHoldElapsed() {
playPauseTimer = nil
statsHoldFired = true
StatsVerbosity.cycle()
}
/// A Play/Pause tap, delivered now that it resolved as one: the right button down, its
/// release `tapPress` behind so the pair can't collapse into nothing downstream.
private func deliverRightClick() {
// A previous tap's owed release goes out FIRST two taps inside `tapPress` would
// otherwise send the host two downs in a row (the rule GamepadCapture's held-back Select
// tap follows for the same reason).
finishRightClick()
setButton(3, down: true)
let timer = Timer(timeInterval: Self.tapPress, repeats: false) { [weak self] _ in
Task { @MainActor in self?.finishRightClick() }
}
RunLoop.main.add(timer, forMode: .common)
rightReleaseTimer = timer
}
/// Release a tap's right button if one is still owed; nothing otherwise.
private func finishRightClick() {
guard rightReleaseTimer != nil else { return }
rightReleaseTimer?.invalidate()
rightReleaseTimer = nil
setButton(3, down: false)
}
/// Drop any in-flight Play/Pause state (unbind / stop). Timers only a right button already
/// sent down is held state, and `releaseHeld` is what lifts it.
private func cancelPlayPause() {
playPauseTimer?.invalidate()
playPauseTimer = nil
rightReleaseTimer?.invalidate()
rightReleaseTimer = nil
statsHoldFired = false
}
private func menuChanged(pressed: Bool) {
if pressed {
menuDownAt = Date()
@@ -176,16 +176,23 @@ public enum DefaultsKey {
/// ("topLeading"/"topTrailing"/"bottomLeading"/"bottomTrailing"). Default top-trailing.
public static let hudPlacement = "punktfunk.hudPlacement"
/// iOS/iPadOS/macOS: switch the host list, settings and game library to a controller-friendly
/// layout (the console launcher, gamepad-navigable settings, a coverflow-style library)
/// whenever a gamepad is connected. On by default; see `GamepadUIEnvironment.isActive`.
/// layout (the console launcher, gamepad-navigable settings, a coverflow-style library).
/// On by default; WHEN it takes over is `gamepadUIMode`. See `GamepadUIEnvironment.isActive`.
public static let gamepadUIEnabled = "punktfunk.gamepadUIEnabled"
/// When `gamepadUIEnabled` actually takes over: `"connected"` (the default only while a
/// usable controller is attached, the behaviour this switch has always had) or `"always"`,
/// for someone who prefers the console layout with no pad in reach (a TV-connected iPad, a
/// Mac driven from the couch). Read only while `gamepadUIEnabled` is on, which is why the
/// settings rows hide it when the switch is off. Anything unrecognized reads as
/// `"connected"`. A device preference, never part of a stream profile.
public static let gamepadUIMode = "punktfunk.gamepadUIMode"
/// Which colour family the gamepad UI's living backdrop drifts through a
/// `GamepadPalette` id ("violet" = the brand default, then "tide"/"forest"/"ember"/
/// "rose"/"graphite"). The cross-client `ui_palette` key: the desktop console and the
/// Android client carry the same table under the same names. Presentation only, so it is
/// a device preference and never part of a stream profile. An unknown value reads as the
/// default rather than failing a newer client may have shipped a palette this build
/// doesn't know.
/// `GamepadPalette` id ("violet" = the brand default, then "oled"/"nebula"/"abyss"/"ember"/
/// "moss"/"graphite", then the pale ones). The cross-client `ui_palette` key: the desktop
/// console and the Android client carry the same table under the same names. Presentation
/// only, so it is a device preference and never part of a stream profile. An unknown value
/// reads as the default rather than failing a newer client may have shipped a palette this
/// build doesn't know.
public static let uiPalette = "punktfunk.uiPalette"
/// iPhone: ALSO play the rumble the host addresses to controller 1 (wire pad 0) on this
/// device's own Taptic Engine for phone-clip pads that ship without rumble motors, where
@@ -65,13 +65,25 @@ public struct GamepadPalette: Identifiable, Equatable, Sendable {
SIMD3(0.22, 0.38, 0.86), SIMD3(0.53, 0.47, 0.96),
]
/// The twelve shipped palettes: the brand default, five more dark fields, then six pale
/// The thirteen shipped palettes: the brand default, six more dark fields, then six pale
/// ones. Cycling order runs dark light, so stepping the row walks the whole range one way.
public static let all: [GamepadPalette] = [
// --- dark fields (white ink) ---
GamepadPalette(
id: "violet", name: "Violet", stops: [],
ground: SIMD3(0.075, 0.060, 0.160), accent: SIMD3(0.525, 0.471, 0.961), light: false),
GamepadPalette(
// For OLED and AMOLED panels, where a black pixel is a pixel switched off no glow,
// no power. The first two stops are literally (0,0,0), so the shaded half of the
// field is genuinely off rather than "very dark grey", and the ground is pure black
// too: the calm mix on the form screens lifts toward nothing. What is left is a
// faint indigoviolet ember in the bright corner. The accent stays the brand violet
// focus has to be findable on black.
id: "oled", name: "OLED",
stops: [SIMD3(0.000, 0.000, 0.000), SIMD3(0.000, 0.000, 0.000),
SIMD3(0.010, 0.020, 0.100), SIMD3(0.045, 0.016, 0.115),
SIMD3(0.120, 0.024, 0.130)],
ground: SIMD3(0, 0, 0), accent: SIMD3(0.525, 0.471, 0.961), light: false),
GamepadPalette(
// Deep indigo climbing through violet into a hot magenta.
id: "nebula", name: "Nebula",
@@ -0,0 +1,102 @@
// The device-switch regression, end to end against a real session.
//
// An AVAudioEngine does not follow the audio hardware: when the output device changes under a
// running engine it STOPS ITSELF and stays stopped. Nothing restarted it, so a stream whose
// output moved mid-session AirPods taken out of an ear, a headset unplugged, the default
// changed in System Settings played silence from that moment on: nothing on the speakers the
// system had just moved to, and nothing in the AirPods when they went back in, since that is a
// second stop rather than a recovery. Only restarting the whole stream brought audio back.
//
// This drives the real `SessionAudio` against the loopback host and moves the system's default
// output device out from under it, twice out and back, the exact shape of the field report.
// Playback-only (mic off): it is the render side that died, and a mic would drag the microphone
// permission and the voice processor into a test that is about neither.
//
// Driven by clients/apple/test-loopback.sh, like its LoopbackIntegrationTests siblings.
#if os(macOS)
import AVFoundation
import CoreAudio
import XCTest
@testable import PunktfunkKit
final class AudioDeviceSwitchTests: XCTestCase {
/// Set the system default output device. Test-local on purpose: nothing in the app ever
/// changes the user's device, it only follows it.
private func setDefaultOutput(_ id: AudioDeviceID) -> OSStatus {
var address = AudioObjectPropertyAddress(
mSelector: kAudioHardwarePropertyDefaultOutputDevice,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
var dev = id
return AudioObjectSetPropertyData(
AudioObjectID(kAudioObjectSystemObject), &address, 0, nil,
UInt32(MemoryLayout<AudioDeviceID>.size), &dev)
}
/// Pump the MAIN runloop until playback is running on `device`, or the deadline passes. The
/// recovery lands on the main queue (a debounced hop, then possibly a retry ladder), so a
/// sleeping test would block the very thing it is waiting for.
private func waitForPlayback(
_ audio: SessionAudio, on device: AudioDeviceID, timeout: TimeInterval
) -> Bool {
let deadline = Date().addingTimeInterval(timeout)
while Date() < deadline {
RunLoop.current.run(until: Date().addingTimeInterval(0.05))
let state = audio.playbackState
if state.running, state.device == device { return true }
}
return false
}
func testPlaybackFollowsAnOutputDeviceChange() throws {
guard let portStr = ProcessInfo.processInfo.environment["PUNKTFUNK_LOOPBACK_PORT"],
let port = UInt16(portStr)
else {
throw XCTSkip("needs a running punktfunk1-host — use clients/apple/test-loopback.sh")
}
guard let original = AudioDevices.defaultOutputDevice() else {
throw XCTSkip("no default output device")
}
let others = AudioDevices.outputs()
.compactMap { AudioDevices.deviceID(forUID: $0.uid) }
.filter { $0 != original }
guard let target = others.first else {
throw XCTSkip("needs a second output device to switch to")
}
let conn = try PunktfunkConnection(
host: "127.0.0.1", port: port, width: 1280, height: 720, refreshHz: 60,
bitrateKbps: 50_000)
let audio = SessionAudio(connection: conn)
// "" speaker UID = follow the system default, which is what the report was running and
// the only configuration a default-device change is supposed to move.
audio.start(
speakerUID: "", micUID: "", micChannel: 0, micEnabled: false, echoCancel: false)
defer {
audio.stop()
_ = setDefaultOutput(original)
}
XCTAssertTrue(
waitForPlayback(audio, on: original, timeout: 5),
"playback never started on the current default output device")
// Out: the device the stream was playing to goes away underneath it.
XCTAssertEqual(setDefaultOutput(target), noErr)
XCTAssertTrue(
waitForPlayback(audio, on: target, timeout: 10),
"playback did not come back after the output device changed — this is the field "
+ "report: no sound on the device the system moved to, until the stream is "
+ "restarted")
// And back: the second half of the report, where putting the AirPods back in produced a
// second stop rather than a recovery.
XCTAssertEqual(setDefaultOutput(original), noErr)
XCTAssertTrue(
waitForPlayback(audio, on: original, timeout: 10),
"playback did not come back after the output device changed back")
}
}
#endif
@@ -0,0 +1,121 @@
// The trigger half of surviving a device change: does the session actually get TOLD?
//
// An AVAudioEngine stops itself when its output hardware changes and never restarts on its own, so
// everything downstream of these notifications is dead code if the notification never arrives. The
// rebuild itself needs a live session to exercise (and so a host, which does not build on macOS),
// but the wiring does not and the wiring is where a silent failure costs a session all of its
// audio, which is exactly the shape of the bug this watcher exists to fix.
import AVFoundation
import XCTest
#if os(macOS)
import CoreAudio
#endif
@testable import PunktfunkKit
final class AudioDeviceWatcherTests: XCTestCase {
/// The callbacks land on the main queue, so a test that slept would block the thing it waits
/// for. Pumps until `predicate` holds or the deadline passes.
private func pump(until predicate: () -> Bool, timeout: TimeInterval = 2) -> Bool {
let deadline = Date().addingTimeInterval(timeout)
while Date() < deadline {
if predicate() { return true }
RunLoop.current.run(until: Date().addingTimeInterval(0.02))
}
return predicate()
}
/// The identity gate is the one line that could swallow every notification silently: get it
/// wrong and the recovery compiles, installs, runs and never fires.
func testAConfigurationChangeFromOurEngineReachesTheOwner() {
let engine = AVAudioEngine()
var reasons: [AudioDeviceWatcher.Reason] = []
let watcher = AudioDeviceWatcher(
isOurs: { $0 === engine }, onChange: { reasons.append($0) })
watcher.start()
defer { watcher.stop() }
NotificationCenter.default.post(
name: .AVAudioEngineConfigurationChange, object: engine)
XCTAssertTrue(
pump(until: { reasons.contains(.engineConfiguration) }),
"the session was never told its engine's configuration changed")
}
/// A retired engine posts one last change as it is torn down, and other AVAudioEngines in the
/// process are not ours to restart rebuilding for either would interrupt healthy playback.
func testAConfigurationChangeFromAForeignEngineIsIgnored() {
let ours = AVAudioEngine()
let stranger = AVAudioEngine()
var reasons: [AudioDeviceWatcher.Reason] = []
let watcher = AudioDeviceWatcher(
isOurs: { $0 === ours }, onChange: { reasons.append($0) })
watcher.start()
defer { watcher.stop() }
NotificationCenter.default.post(
name: .AVAudioEngineConfigurationChange, object: stranger)
// Give it the same grace the positive case gets, then require silence.
_ = pump(until: { !reasons.isEmpty }, timeout: 0.5)
XCTAssertTrue(reasons.isEmpty, "a foreign engine's change was taken for ours")
}
func testStopSilencesTheWatcher() {
let engine = AVAudioEngine()
var reasons: [AudioDeviceWatcher.Reason] = []
let watcher = AudioDeviceWatcher(
isOurs: { $0 === engine }, onChange: { reasons.append($0) })
watcher.start()
watcher.stop()
NotificationCenter.default.post(
name: .AVAudioEngineConfigurationChange, object: engine)
_ = pump(until: { !reasons.isEmpty }, timeout: 0.5)
XCTAssertTrue(reasons.isEmpty, "a stopped watcher still reported")
}
#if os(macOS)
/// The backstop, against the real HAL: move the system's default output device the thing that
/// happens when AirPods come out of an ear and require that the session hears about it. This
/// is the trigger the recovery leans on for the voice-processing engine, whose own notification
/// behaviour cannot be verified here (no Mac in this project's fleet can initialize VPIO).
func testTheDefaultOutputDeviceMovingReachesTheOwner() throws {
guard let original = AudioDevices.defaultOutputDevice() else {
throw XCTSkip("no default output device")
}
let others = AudioDevices.outputs()
.compactMap { AudioDevices.deviceID(forUID: $0.uid) }
.filter { $0 != original }
guard let target = others.first else {
throw XCTSkip("needs a second output device to switch to")
}
var reasons: [AudioDeviceWatcher.Reason] = []
let watcher = AudioDeviceWatcher(isOurs: { _ in false }, onChange: { reasons.append($0) })
watcher.start()
defer {
_ = Self.setDefaultOutput(original)
watcher.stop()
}
XCTAssertEqual(Self.setDefaultOutput(target), noErr)
XCTAssertTrue(
pump(until: { reasons.contains(.defaultOutputDevice) }, timeout: 5),
"the session was never told the default output device moved")
}
/// Test-local on purpose: nothing in the app ever changes the user's device, it only follows it.
private static func setDefaultOutput(_ id: AudioDeviceID) -> OSStatus {
var address = AudioObjectPropertyAddress(
mSelector: kAudioHardwarePropertyDefaultOutputDevice,
mScope: kAudioObjectPropertyScopeGlobal,
mElement: kAudioObjectPropertyElementMain)
var dev = id
return AudioObjectSetPropertyData(
AudioObjectID(kAudioObjectSystemObject), &address, 0, nil,
UInt32(MemoryLayout<AudioDeviceID>.size), &dev)
}
#endif
}
@@ -53,11 +53,22 @@ final class AudioRingDriftTests: XCTestCase {
XCTAssertEqual(silent, 0, "drift correction must never starve the callback")
}
/// The mirror case: a host clock running SLOW must keep audio flowing rather than being
/// "corrected" into a stutter.
func testNegativeDriftKeepsPlaying() {
/// The mirror case: a host clock running SLOW is a genuine deficit no depth is ever deep
/// enough forever so the ring must spend it on RARE, clean re-banks (a hollow ring
/// re-primes on its first click and refills the whole target) rather than riding the knife
/// edge in permanent sub-frame chatter, which is what "silence-free" used to hide: every
/// callback a fraction of a frame short, none of them fully silent, all of them audible.
/// 200 ppm is an exaggeration of real DAC skew (tens of ppm); even so, two minutes may
/// cost at most a couple of refills' worth of silent callbacks.
func testNegativeDriftBanksRarelyInsteadOfChattering() {
let (_, _, silent) = simulate(ms: 2 * 60 * 1_000, quantumMS: 5, driftPPM: -200)
XCTAssertEqual(silent, 0, "a draining ring must re-prime, not chatter")
XCTAssertLessThanOrEqual(
silent, 24,
"a draining ring re-banks a few times; a silent-callback stream means it is thrashing")
XCTAssertGreaterThan(
silent, 0,
"a persistent deficit cannot be ridden out silence-free — if this is zero the ring "
+ "is back to sub-frame chatter, which is audible without ever being silent")
}
/// A device that pulls a large quantum cannot sustain a target below it the ring must lift
@@ -79,11 +90,14 @@ final class AudioRingDriftTests: XCTestCase {
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: want) }
XCTAssertTrue(scratch.contains { $0 != 0 }, "should be playing after priming")
// Drain it dry with one oversized read, then feed a normal quantum again. The length comes
// off the buffer pointer, not off `huge`: touching the array inside the closure that is
// already holding it exclusively is an exclusivity violation.
var huge = [Float](repeating: 0, count: 200 * perMS)
huge.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: $0.count) }
// Drain it dry at the device's own quantum an oversized read would count as ITS OWN
// huge callback and legitimately read as hollow then starve one callback and feed a
// normal quantum again. The ring is freshly primed, so its depth average is nowhere near
// hollow, and one short read must ride on the hysteresis.
while ring.bufferedMS > 0 {
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: want) }
}
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: want) }
let feed = [Float](repeating: 0.5, count: want)
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: want) }
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: want) }
@@ -92,14 +106,17 @@ final class AudioRingDriftTests: XCTestCase {
"a single short read must not force a full re-prime")
}
/// Mirror of the Rust `target_grows_on_underruns_and_relaxes_when_quiet`: clustered genuine
/// underruns raise the target floor (that session needs the slack), a long quiet spell gives
/// it back and the floor never dips below the base.
/// Mirror of the Rust `target_grows_on_underruns_and_relaxes_when_quiet`, updated for
/// near-miss growth: the drain's LAST full read (less than a frame left over) already grows
/// the floor before anything was audible, clustered genuine underruns raise it further, and
/// a long genuinely quiet spell gives it back, never below the base. The quiet refill
/// runs DEEP: a knife-edge refill (exactly what each read takes) leaves the ring within a
/// frame of empty every callback, which now correctly reads as pressure, not quiet.
func testTargetGrowsOnUnderrunsAndRelaxesWhenQuiet() {
let ring = AudioRing(capacity: 48_000 * channels, channels: channels)
let want = 5 * perMS
var scratch = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 25 * perMS)
let feed = [Float](repeating: 0.5, count: 60 * perMS)
func write(ms: Int) {
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: ms * perMS) }
}
@@ -108,20 +125,27 @@ final class AudioRingDriftTests: XCTestCase {
}
XCTAssertEqual(ring.stats.targetMS, 20, "base target must match JitterTuning.COREAUDIO")
// Prime, drain dry, then alternate starve/refill: each dry read is a genuine underrun,
// each full read in between keeps the de-prime hysteresis from tripping.
// Prime, then drain: the 5th read is still served in full but leaves nothing over a
// near-miss, and the floor grows BEFORE any click.
write(ms: 25)
for _ in 0..<5 { read() } // drains to zero
for _ in 0..<5 { read() }
XCTAssertEqual(ring.stats.targetMS, 30, "a near-miss must grow the floor pre-click")
XCTAssertEqual(ring.stats.underruns, 0, "nothing was audible yet")
// Then alternate starve/refill: each dry read is a genuine underrun, each full read in
// between keeps the de-prime hysteresis from tripping. (The refills land as further
// near-misses, but growth is one step per window the cluster is what grows it again.)
read() // short underrun 1
write(ms: 5); read() // full hysteresis reset
read() // short underrun 2
write(ms: 5); read() // full
read() // short underrun 3 the floor grows one step
XCTAssertEqual(ring.stats.targetMS, 30, "3 clustered underruns must grow the target 10 ms")
XCTAssertEqual(ring.stats.targetMS, 40, "3 clustered underruns must grow the target 10 ms")
XCTAssertEqual(ring.stats.underruns, 3)
// A long clean run (30 s of consumed audio) relaxes the growth back to the base
for _ in 0..<(30_000 / 5 + 10) {
// A long clean run at a healthy depth relaxes the growth back to the base
write(ms: 60)
for _ in 0..<(90_000 / 5 + 10) {
write(ms: 5)
read()
}
@@ -435,10 +459,14 @@ final class AudioRingDriftTests: XCTestCase {
write(ms: 5); read()
read()
}
/// Quiet (full) reads needed before the grown target relaxes one step.
/// Quiet (full) reads needed before the grown target relaxes one step. The ring is
/// refilled DEEP first: a knife-edge refill (exactly what each read takes) leaves less
/// than a frame over every callback, which now correctly reads as pressure near-misses
/// and pressure never relaxes anything.
func quietToRelax(_ ring: AudioRing) -> Int {
var scratch = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 5 * perMS)
let feed = [Float](repeating: 0.5, count: 60 * perMS)
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: 60 * perMS) }
let start = ring.stats.targetMS
var reads = 0
while ring.stats.targetMS == start, reads < 200_000 {
@@ -466,6 +494,118 @@ final class AudioRingDriftTests: XCTestCase {
"sync pressure should relax sooner: \(fastReads) vs \(slowReads) quiet reads")
}
/// A shrink answered by an underrun or near-miss inside its probe window is undone AT ONCE,
/// and the sync loop is backed off mirrors the Rust `a_failed_shrink_probe_is_undone_at_once`
/// and `a_failed_probe_backs_the_sync_shrink_off`. Before this, the loop re-probed a proven
/// depth every five quiet seconds and paid an audible starvation event each time it was wrong,
/// forever the 0.25.0 MacBook field report.
func testAFailedShrinkProbeIsUndoneAtOnceAndBacksTheSyncLoopOff() {
let ring = AudioRing(capacity: 48_000 * channels, channels: channels)
let want = 5 * perMS
var scratch = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 60 * perMS)
func write(ms: Int) {
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: ms * perMS) }
}
func read() {
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: want) }
}
// Grow the floor (near-miss + a cluster of genuine underruns), as the usual pattern does.
write(ms: 25)
for _ in 0..<5 { read() }
read()
write(ms: 5); read()
read()
write(ms: 5); read()
read()
let grown = ring.stats.targetMS
XCTAssertGreaterThan(grown, 20, "the test needs a GROWN floor to probe")
// Sync asks for less; a deep, genuinely quiet spell later the shrink probes.
ring.setSyncTarget(perMS)
write(ms: 60)
var reads = 0
while ring.stats.targetMS == grown, reads < 10_000 {
write(ms: 5)
read()
reads += 1
}
XCTAssertEqual(ring.stats.targetMS, grown - 10, "the sync-driven shrink must have probed")
// Drain to the knife edge: the last full read leaves nothing over a near-miss, nobody
// heard anything and the probe must be undone on the spot.
while ring.bufferedMS > 5 { read() }
read()
XCTAssertEqual(
ring.stats.targetMS, grown,
"a failed probe must restore the target on the first near-miss")
XCTAssertEqual(ring.stats.underruns, 3, "and nothing audible may have paid for it")
// Backed off: two accelerated windows of clean, deep audio must NOT shrink again
write(ms: 60)
for _ in 0..<(2 * 5_000 / 5) {
write(ms: 5)
read()
}
XCTAssertEqual(
ring.stats.targetMS, grown,
"the five-second cadence must be suspended after a failure")
// while the slow, pre-sync window eventually still tests one backoff is not a freeze.
for _ in 0..<(2 * 30_000 / 5) {
write(ms: 5)
read()
}
XCTAssertLessThan(
ring.stats.targetMS, grown,
"the slow window must still be allowed to test a shrink")
}
/// Growth raises a promise; only a re-prime banks real depth. An underrun while the ring is
/// HOLLOW its depth AVERAGE far below the target re-primes immediately, spending the click
/// it already cost on the whole refill, instead of riding the knife edge and clicking once per
/// bunching period indefinitely. The average, not the instant, is what separates a hollow ring
/// from one late packet (`testSingleShortReadDoesNotDeprime` pins that side).
func testAHollowRingReprimesOnItsFirstClick() {
let ring = AudioRing(capacity: 48_000 * channels, channels: channels)
let want = 5 * perMS
var scratch = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 60 * perMS)
func write(ms: Int) {
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: ms * perMS) }
}
func read() {
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: want) }
}
// Grow the floor to 40 the usual way
write(ms: 25)
for _ in 0..<5 { read() }
read()
write(ms: 5); read()
read()
write(ms: 5); read()
read()
XCTAssertEqual(ring.stats.targetMS, 40)
// then ride the knife edge for ~2 s of audio, so the depth average genuinely sinks far
// below the promised 40 ms.
for _ in 0..<400 {
write(ms: 5)
read()
}
// One dry read the click. The ring is hollow, so this single click must re-prime.
read()
// A packet arrives, but the ring stays SILENT: it is re-priming toward the full target
// rather than playing the packet and clicking again at the next bunch.
write(ms: 10)
read()
XCTAssertTrue(
scratch.allSatisfy { $0 == 0 },
"a hollow ring must spend its click on the whole refill, not keep limping")
// And once the refill reaches the target, it plays again.
write(ms: 40)
read()
XCTAssertTrue(scratch.contains { $0 != 0 }, "refilled to target — playback resumes")
}
/// The four client rings adopt sync one at a time; an un-wired one must behave exactly as it
/// did. `nil` is the default, so this pins the initializer too and every other test in this
/// file runs without a sync target, which is the real guard that nothing moved underneath them.
@@ -5,8 +5,9 @@ import XCTest
/// The escape chord's mask and its GameController alias list have to describe the same four
/// buttons. `GamepadCapture.openSlot` claims the system gesture of every element while forwarding
/// is on, but only of `escapeChordElements` while it is off so if the alias list ever stops
/// covering the mask, the missing button's press stays the system's and the chord never completes.
/// is on, but only of `chordElements` `escapeChordElements` plus the stats chord's while it is
/// off, so if this alias list ever stops covering the mask, the missing button's press stays the
/// system's and the chord never completes. (`GamepadStatsChordTests` pins the claim list itself.)
///
/// That matters most on tvOS, where this chord is the only controller way out of a stream: the
/// symptom is a session nobody can leave with the pad in their hands, and nothing logs or crashes.
@@ -46,12 +46,29 @@ final class GamepadPaletteTests: XCTestCase {
func testTableMatchesTheOtherClients() {
XCTAssertEqual(
GamepadPalette.all.map(\.id),
["violet", "nebula", "abyss", "ember", "moss", "graphite",
["violet", "oled", "nebula", "abyss", "ember", "moss", "graphite",
"holo", "sunset", "bloom", "dawn", "mint", "opal"])
// Dark fields lead, pale ones follow, so stepping the row walks one direction.
let firstLight = GamepadPalette.all.firstIndex { $0.light }
XCTAssertEqual(firstLight, 6)
XCTAssertTrue(GamepadPalette.all.dropFirst(6).allSatisfy(\.light))
XCTAssertEqual(firstLight, 7)
XCTAssertTrue(GamepadPalette.all.dropFirst(7).allSatisfy(\.light))
}
/// OLED is the one palette whose selling point is measurable: it has to be genuinely black,
/// not merely the darkest of the dark fields.
func testOLEDIsActuallyBlack() {
let oled = GamepadPalette.named("oled")
XCTAssertEqual(oled.ground, SIMD3(0, 0, 0), "the calm lift must be nothing")
let cells = oled.meshColors
XCTAssertGreaterThanOrEqual(
cells.filter { luma($0) == 0 }.count, 3,
"the shaded corner has to be switched off, not dimmed")
let mean = cells.map(luma).reduce(0, +) / Double(cells.count)
let darkestOther = GamepadPalette.all
.filter { $0.id != "oled" }
.map { p in p.meshColors.map(luma).reduce(0, +) / Double(p.meshColors.count) }
.min() ?? 0
XCTAssertLessThan(mean, darkestOther / 2, "oled is barely darker than \(darkestOther)")
}
/// A palette must read as SEVERAL hues, not one hue at several brightnesses that was
@@ -0,0 +1,79 @@
import GameController
import XCTest
@testable import PunktfunkKit
/// The stats chord (Select + X) has the same drift hazard as the escape chord it sits beside: its
/// mask and its GameController alias list must describe the same buttons, and every element some
/// chord reads has to appear in the list a NON-forwarding slot claims otherwise that button's
/// press stays the system's and the chord silently never completes.
///
/// It matters most on tvOS, where this is the only way to the statistics overlay at all (no
/// keyboard for S, no touchscreen for the three-finger tap). The failure looks like nothing
/// happening, so it is pinned here rather than left to the comments.
@MainActor
final class GamepadStatsChordTests: XCTestCase {
/// The intended aliasbit pairing, spelled out independently of the implementation.
private let pairing: [(alias: String, bit: UInt32)] = [
(GCInputButtonOptions, GamepadWire.back),
(GCInputButtonX, GamepadWire.x),
]
func testChordMaskIsExactlyTheTwoPairedButtons() {
XCTAssertEqual(
pairing.reduce(UInt32(0)) { $0 | $1.bit },
GamepadCapture.statsChord,
"the chord mask and the alias pairing describe different buttons")
}
func testAliasListMirrorsTheMask() {
XCTAssertEqual(
GamepadCapture.statsChordElements.count,
GamepadCapture.statsChord.nonzeroBitCount,
"alias list and chord mask differ in size")
XCTAssertEqual(GamepadCapture.statsChordElements, pairing.map(\.alias))
}
/// The two chords must not be reachable through one another: pressing toward the exit chord
/// may not cycle the overlay on the way, and holding the stats chord may not arm a disconnect.
/// Select is the one button they share by design everything else has to be disjoint.
func testChordsOverlapOnlyOnSelect() {
XCTAssertEqual(
GamepadCapture.statsChord & GamepadCapture.escapeChord,
GamepadWire.back,
"the stats and escape chords share a button other than Select")
// Neither is a subset of the other, so completing one can never complete the other.
XCTAssertNotEqual(
GamepadCapture.statsChord & GamepadCapture.escapeChord, GamepadCapture.statsChord)
XCTAssertNotEqual(
GamepadCapture.statsChord & GamepadCapture.escapeChord, GamepadCapture.escapeChord)
}
/// `chordElements` is what `openSlot` claims when forwarding is OFF. It must cover BOTH
/// chords' aliases and repeat none of them (a duplicate would mean a bit with no element).
func testClaimListCoversBothChordsWithoutDuplicates() {
let claim = GamepadCapture.chordElements
for alias in GamepadCapture.escapeChordElements + GamepadCapture.statsChordElements {
XCTAssertTrue(claim.contains(alias), "\(alias) is read by a chord but never claimed")
}
XCTAssertEqual(Set(claim).count, claim.count, "a repeated alias in the claim list")
// Shared Select means the union is one shorter than the two lists laid end to end.
XCTAssertEqual(
claim.count,
GamepadCapture.escapeChordElements.count + GamepadCapture.statsChordElements.count - 1)
}
/// A cycle is a pure rotation through the four tiers the chord fires `StatsVerbosity.cycle`,
/// and a tier that dead-ended would strand a tvOS user with no other way back.
func testCycleReachesEveryTierAndReturns() {
var tier = StatsVerbosity.off
var seen: [StatsVerbosity] = []
for _ in 0..<StatsVerbosity.allCases.count {
seen.append(tier)
tier = tier.next()
}
XCTAssertEqual(Set(seen).count, StatsVerbosity.allCases.count, "a tier is unreachable")
XCTAssertEqual(tier, .off, "the cycle does not return to where it started")
}
}
@@ -1,14 +1,58 @@
// GamepadUIEnvironment.isActive is a pure AND table-tested exhaustively over its 2x2 inputs.
// GamepadUIEnvironment.isActive is pure table-tested exhaustively over its inputs.
import XCTest
@testable import PunktfunkKit
final class GamepadUIEnvironmentTests: XCTestCase {
func testActiveOnlyWhenEnabledAndConnected() {
XCTAssertTrue(GamepadUIEnvironment.isActive(gamepadConnected: true, enabledSetting: true))
XCTAssertFalse(GamepadUIEnvironment.isActive(gamepadConnected: true, enabledSetting: false))
XCTAssertFalse(GamepadUIEnvironment.isActive(gamepadConnected: false, enabledSetting: true))
XCTAssertFalse(GamepadUIEnvironment.isActive(gamepadConnected: false, enabledSetting: false))
private let connected = GamepadUIEnvironment.modeWhenConnected
private let always = GamepadUIEnvironment.modeAlways
/// The default mode is the behaviour the switch had when it was a lone Bool, so an install
/// that never sees the new row is exactly where it was.
func testWhenConnectedIsAPlainAnd() {
XCTAssertTrue(
GamepadUIEnvironment.isActive(
gamepadConnected: true, enabledSetting: true, mode: connected))
XCTAssertFalse(
GamepadUIEnvironment.isActive(
gamepadConnected: true, enabledSetting: false, mode: connected))
XCTAssertFalse(
GamepadUIEnvironment.isActive(
gamepadConnected: false, enabledSetting: true, mode: connected))
XCTAssertFalse(
GamepadUIEnvironment.isActive(
gamepadConnected: false, enabledSetting: false, mode: connected))
}
/// Always drops the controller from the decision entirely but NOT the switch, which stays
/// the one way back to the touch UI.
func testAlwaysIgnoresTheControllerButNotTheSwitch() {
XCTAssertTrue(
GamepadUIEnvironment.isActive(
gamepadConnected: false, enabledSetting: true, mode: always))
XCTAssertTrue(
GamepadUIEnvironment.isActive(
gamepadConnected: true, enabledSetting: true, mode: always))
XCTAssertFalse(
GamepadUIEnvironment.isActive(
gamepadConnected: false, enabledSetting: false, mode: always))
XCTAssertFalse(
GamepadUIEnvironment.isActive(
gamepadConnected: true, enabledSetting: false, mode: always))
}
/// A value a newer client wrote must wait for a controller, never strand this build in a
/// layout it has no way back out of.
func testUnknownModeWaitsForAController() {
XCTAssertFalse(
GamepadUIEnvironment.isActive(
gamepadConnected: false, enabledSetting: true, mode: "whenever-i-say-so"))
XCTAssertTrue(
GamepadUIEnvironment.isActive(
gamepadConnected: true, enabledSetting: true, mode: "whenever-i-say-so"))
XCTAssertFalse(
GamepadUIEnvironment.isActive(
gamepadConnected: false, enabledSetting: true, mode: ""))
}
}
+5 -2
View File
@@ -26,8 +26,11 @@ mkdir -p "$CFG/open" "$CFG/paired" "$CFG/guess"
trap 'kill "${HOST_PID:-}" "${PAIR_PID:-}" "${GUESS_PID:-}" 2>/dev/null || true' EXIT
# The open host also scripts a feedback burst (rumble + DualSense hidout) right after the
# handshake, so the Swift test can assert the host→client feedback planes end to end.
# The open host outlives the others on purpose: AudioDeviceSwitchTests connects to it and then
# spends tens of seconds moving the system's output device around, long after the 300 frames the
# round-trip test needs.
HOME="$CFG/open" XDG_CONFIG_HOME="$CFG/open/.config" PUNKTFUNK_TEST_FEEDBACK=1 \
target/release/punktfunk-host punktfunk1-host --port "$PORT" --source synthetic --frames 300 \
target/release/punktfunk-host punktfunk1-host --port "$PORT" --source synthetic --frames 12000 \
--allow-tofu &
HOST_PID=$!
HOME="$CFG/paired" XDG_CONFIG_HOME="$CFG/paired/.config" \
@@ -61,4 +64,4 @@ cd clients/apple
PUNKTFUNK_LOOPBACK_PORT="$PORT" PUNKTFUNK_PAIRING_PORT="$PAIR_PORT" PUNKTFUNK_PAIRING_PIN="$PIN" \
PUNKTFUNK_GUESS_PORT="$GUESS_PORT" PUNKTFUNK_GUESS_PIN="$GUESS_PIN" \
PUNKTFUNK_TEST_FEEDBACK=1 \
swift test --filter LoopbackIntegrationTests
swift test --filter 'LoopbackIntegrationTests|AudioDeviceSwitchTests'
+111 -26
View File
@@ -303,6 +303,58 @@ def _native_client() -> str | None:
return None
# The one architecture the flatpak client is built for.
_FLATPAK_ARCH = "x86_64"
def _flatpak_ref() -> dict | None:
"""The INSTALLED client flatpak resolved to a SCOPE and a BRANCH, or None when there is none.
``{"scope": "--user"|"--system", "branch": "canary", "ref": "io.unom.Punktfunk//canary"}``.
**Naming no branch is not a shorthand for "the only one".** flatpak refuses an ambiguous
ref rather than guessing at one, and the ambiguity does not need two branches *installed*:
the punktfunk remote publishes `stable` AND `canary`, so an unqualified
``flatpak remote-info <origin> io.unom.Punktfunk`` errors with "Multiple branches available"
on a Deck that has exactly one. That error is why the client update check silently answered
"up to date" on every Deck so every query downstream now names the ref in full.
Read off the exported tree rather than by shelling out to ``flatpak list``, because
:func:`_client_argv` is on the path of every headless call and a subprocess per call would be
absurd (the same reason :func:`_flatpak_installed` reads the filesystem). ``active`` is the
symlink flatpak points at the deployed commit its presence is what makes a branch directory
an INSTALL rather than the leftovers of one.
With more than one branch installed, `stable` wins, because that is the branch a plain
``flatpak run`` resolves to: the check has to describe the client the launcher really starts,
or a stale `stable` silently beats a current `canary` in both places at once.
"""
if not _flatpak():
return None
for root, scope in (
(Path(decky.DECKY_USER_HOME) / ".local" / "share" / "flatpak", "--user"),
(Path("/var/lib/flatpak"), "--system"),
):
try:
branches = sorted(
p.name for p in (root / "app" / APP_ID / _FLATPAK_ARCH).iterdir()
if (p / "active").exists()
)
except OSError:
continue # not installed in this scope
if not branches:
continue
branch = "stable" if "stable" in branches else branches[0]
if len(branches) > 1:
decky.logger.warning(
"%s is installed on %d branches (%s) — using %s, the one `flatpak run` resolves "
"to; uninstall the others so the client you launch is the client we update",
APP_ID, len(branches), ", ".join(branches), branch,
)
return {"scope": scope, "branch": branch, "ref": f"{APP_ID}//{branch}"}
return None
def _flatpak_installed() -> bool:
"""True when the flatpak APP is actually installed — not merely that `flatpak` exists.
@@ -310,10 +362,7 @@ def _flatpak_installed() -> bool:
because this is on the path of every headless call and a subprocess per call would be absurd.
Both scopes count: the Deck installs --user, a distro image may ship it system-wide.
"""
if not _flatpak():
return False
user = Path(decky.DECKY_USER_HOME) / ".local" / "share" / "flatpak" / "app" / APP_ID
return user.exists() or Path("/var/lib/flatpak/app", APP_ID).exists()
return _flatpak_ref() is not None
def _client_argv() -> list[str] | None:
@@ -323,15 +372,21 @@ def _client_argv() -> list[str] | None:
behaving exactly as it did. A native binary is the fallback and on a machine with no
flatpak client, the thing that makes the plugin work at all. `PF_DECKY_CLIENT=native|flatpak`
forces one when a machine has both.
The branch is PINNED (`--branch=`, which keeps the app id last :func:`_cli_argv` appends
`--command=` and flatpak treats everything after the id as the app's own argv), so the client
this launches is the exact ref :func:`_client_update_state` checks and :meth:`Plugin.
update_client` updates.
"""
forced = os.environ.get("PF_DECKY_CLIENT", "").strip().lower()
native = _native_client()
if forced == "native":
return [native] if native else None
if forced != "flatpak" and not _flatpak_installed() and native:
ref = _flatpak_ref()
if forced != "flatpak" and not ref and native:
return [native]
if _flatpak_installed():
return [_flatpak(), "run", "--arch=x86_64", APP_ID]
if ref:
return [_flatpak(), "run", f"--arch={_FLATPAK_ARCH}", f"--branch={ref['branch']}", APP_ID]
return [native] if native else None
@@ -575,27 +630,44 @@ def _looks_outdated(stderr: str) -> bool:
async def _client_update_state() -> dict:
"""Is a newer commit of the flatpak client available in the remote it tracks? The client is a
**per-user** install (so ``sudo flatpak update``, which is system-scope, never touches it), and
it versions independently of this plugin so we compare the installed commit against the
remote's here and let the QAM offer a user-scope update. Best-effort; all-``False`` on any error
(not installed, no flatpak, offline).
"""Is a newer commit of the flatpak client available in the remote it tracks? The client
versions independently of this plugin, so we compare the installed commit against the
remote's here and let the QAM offer an update in the scope the client is actually installed
in a per-user install is one ``sudo flatpak update`` (system-scope) never reaches.
Flatpak keeps its OWN comparison (commits, not versions) because it is the exact one: a
flatpak built from main between releases carries the release's crate version, so the
signed-manifest comparison the native path uses would call it up to date when it isn't.
Native installs have no commit to compare and go through :func:`_native_update_state`."""
state = {"available": False, "installed": "", "remote": ""}
rc, info = await _flatpak_capture(["info", "--user", APP_ID], timeout=10.0)
Native installs have no commit to compare and go through :func:`_native_update_state`.
Every query names the ref IN FULL (see :func:`_flatpak_ref`) the remote publishes both
`stable` and `canary`, and an unqualified one is an error, not a default."""
state = {"available": False, "installed": "", "remote": "", "error": ""}
ref = _flatpak_ref()
if not ref:
return state # no flatpak client in either scope
scope, full = ref["scope"], ref["ref"]
rc, info = await _flatpak_capture(["info", scope, full], timeout=10.0)
if rc != 0:
return state # client not installed as a user app / no flatpak
decky.logger.warning("flatpak info %s %s failed (rc=%s): %s", scope, full, rc, info[-200:])
state["error"] = "client-unavailable"
return state
state["installed"] = _field_from(info, "Commit")
origin = _field_from(info, "Origin")
if not origin:
state["error"] = "no-origin" # a sideloaded bundle tracks no remote to compare against
return state
rc, rinfo = await _flatpak_capture(["remote-info", "--user", origin, APP_ID], timeout=25.0)
rc, rinfo = await _flatpak_capture(["remote-info", scope, origin, full], timeout=25.0)
if rc != 0:
return state # remote unreachable — treat as "up to date", retry next check
# ⭐ NOT "up to date". Silently swallowing this is precisely how the whole leg stayed
# broken in the field: an unqualified ref made every one of these calls fail, and
# returning `available=False` dressed the failure up as good news. A check that could
# not run says so, and the panel says so too.
decky.logger.warning(
"flatpak remote-info %s %s failed (rc=%s): %s", origin, full, rc, rinfo.strip()[-200:]
)
state["error"] = "fetch-failed"
return state
state["remote"] = _field_from(rinfo, "Commit")
state["available"] = bool(
state["installed"] and state["remote"] and state["installed"] != state["remote"]
@@ -946,8 +1018,9 @@ class Plugin:
async def update_client(self) -> dict:
"""Update the **client**, by whichever route this box's install actually supports.
* **flatpak** ``flatpak update --user`` in the USER installation, the scope a Steam
Deck install lives in and which ``sudo flatpak update`` (system-scope) never reaches.
* **flatpak** ``flatpak update`` against the FULL ref, in the scope the client is
installed in (a per-user install is one ``sudo flatpak update`` never reaches, and an
unqualified ref is an error on a remote publishing more than one branch).
* **native, one-tap capable** (.deb / .rpm / pacman with the packaged root helper and
the operator's group opt-in) — ``punktfunk-client --apply-update``, which starts the
fixed, parameterless ``punktfunk-client-update.service`` through polkit. This backend
@@ -960,18 +1033,22 @@ class Plugin:
"""
if not _client_is_flatpak():
return await self._update_native_client()
_, before = await _flatpak_capture(["info", "--user", APP_ID], timeout=10.0)
ref = _flatpak_ref()
if not ref:
return {"ok": False, "updated": False, "error": "client-unavailable"}
scope, full = ref["scope"], ref["ref"]
_, before = await _flatpak_capture(["info", scope, full], timeout=10.0)
before_commit = _field_from(before, "Commit")
rc, out = await _flatpak_capture(["update", "--user", "-y", APP_ID], timeout=300.0)
rc, out = await _flatpak_capture(["update", scope, "-y", full], timeout=300.0)
if rc != 0:
decky.logger.warning("flatpak client update failed (rc=%s): %s", rc, out[-400:])
return {"ok": False, "updated": False, "error": "update-failed"}
_, after = await _flatpak_capture(["info", "--user", APP_ID], timeout=10.0)
_, after = await _flatpak_capture(["info", scope, full], timeout=10.0)
after_commit = _field_from(after, "Commit")
updated = bool(before_commit and after_commit and before_commit != after_commit)
decky.logger.info(
"flatpak client update: %s -> %s (updated=%s)",
before_commit[:10], after_commit[:10], updated,
"flatpak client update (%s %s): %s -> %s (updated=%s)",
scope, full, before_commit[:10], after_commit[:10], updated,
)
_update_cache["data"] = None # invalidate the cached "update available" snapshot
return {"ok": True, "updated": updated}
@@ -1018,12 +1095,20 @@ class Plugin:
try:
if _client_is_flatpak():
cu = await _client_update_state()
ref = _flatpak_ref()
result["client_update_available"] = bool(cu["available"])
result["client_current"] = (cu["installed"] or "")[:10]
result["client_latest"] = (cu["remote"] or "")[:10]
result["client_install"] = "flatpak"
result["client_applier"] = "flatpak"
result["client_command"] = f"flatpak update --user {APP_ID}"
# The line a user could actually run — same scope, same full ref we use. The old
# unqualified one errored out ("Multiple branches available") when pasted, too.
result["client_command"] = (
f"flatpak update {ref['scope']} -y {ref['ref']}" if ref else ""
)
if cu["error"]:
# Same contract as the native leg: "couldn't tell" is never "up to date".
result["client_error"] = cu["error"]
else:
nu = await _native_update_state()
result["client_update_available"] = bool(nu.get("update_available"))
+58
View File
@@ -30,6 +30,10 @@ sys.modules["decky"] = decky
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
import main # noqa: E402 (the plugin backend)
# The argv fixtures below monkey-patch `_client_argv` to pin one install shape; the
# _flatpak_ref block wants the REAL resolver back, so keep a handle on it.
_real_client_argv = main._client_argv
failures = 0
@@ -75,6 +79,60 @@ check("cli argv: native without a sibling CLI is None", main._cli_argv() is None
(tmp / "punktfunk").write_text("")
check("cli argv: native sibling found", main._cli_argv() == [str(tmp / "punktfunk")])
# ---- _flatpak_ref: the branch must be NAMED, always ---------------------------------------
#
# The bug this exists to prevent: every client-update query used to name no branch, and the
# punktfunk remote publishes `stable` AND `canary` — so `flatpak remote-info <origin>
# io.unom.Punktfunk` failed with "Multiple branches available", the check swallowed the failure,
# and the panel reported the client up to date forever. One branch INSTALLED is not enough to
# make the query unambiguous; the ambiguity lives on the remote.
shutil.rmtree("/tmp/pf-test-home", ignore_errors=True)
_fp_root = Path("/tmp/pf-test-home/.local/share/flatpak/app/io.unom.Punktfunk/x86_64")
main._flatpak = lambda: "/usr/bin/flatpak"
main._client_argv = _real_client_argv # undo the fixture patches above
check("ref: nothing installed => None", main._flatpak_ref() is None)
def _install_branch(name: str):
"""A deployed branch: the `active` symlink is what distinguishes an install from leftovers."""
commit = _fp_root / name / "deadbeef"
commit.mkdir(parents=True, exist_ok=True)
(_fp_root / name / "active").symlink_to("deadbeef")
(_fp_root / "canary").mkdir(parents=True, exist_ok=True)
check("ref: a branch dir without `active` is leftovers, not an install", main._flatpak_ref() is None)
_install_branch("canary")
ref = main._flatpak_ref()
check("ref: the single installed branch is used", ref == {
"scope": "--user", "branch": "canary", "ref": "io.unom.Punktfunk//canary",
})
check(
"ref: the launcher pins that branch, app id still LAST",
main._client_argv() == [
"/usr/bin/flatpak", "run", "--arch=x86_64", "--branch=canary", "io.unom.Punktfunk",
],
)
# The pin must survive _cli_argv's rewrite, or the CLI runs a different build than the GUI.
check(
"ref: --command= is inserted before the app id, keeping the pin",
main._cli_argv() == [
"/usr/bin/flatpak", "run", "--arch=x86_64", "--branch=canary",
"--command=punktfunk", "io.unom.Punktfunk",
],
)
# Two installed: `stable` is what a plain `flatpak run` resolves to, so it must be what we
# check and update too — otherwise a leftover stale `stable` wins the launch while `canary`
# gets the update, and the two halves disagree about which client is even running.
_install_branch("stable")
check("ref: with both installed, stable wins (what `flatpak run` picks)",
main._flatpak_ref()["branch"] == "stable")
shutil.rmtree("/tmp/pf-test-home", ignore_errors=True)
# ---- _cli_error: the CLI's exit-code contract -------------------------------------------
#
# Exit 5 + `unknown command` is how a client too old for a verb announces itself — the ONE
+3 -1
View File
@@ -120,7 +120,9 @@ export interface UpdateInfo {
client_applier: string;
client_command: string; // one copy-pastable line that updates this install by hand
client_opt_in: string; // set when one-tap WOULD work after `usermod -aG punktfunk-update`
client_error?: string; // the client check couldn't complete (e.g. "client-outdated")
// The client check couldn't complete — NEVER rendered as "up to date". "client-outdated" |
// "client-unavailable" | "no-origin" | "fetch-failed" (flatpak: the remote was unreachable).
client_error?: string;
error?: string; // "update-channel-unknown" (dev build) | "fetch-failed"
}
+6 -2
View File
@@ -33,6 +33,7 @@ import {
applyUpdate,
checkForUpdatesNow,
clientUpdateIsManualOnly,
clientUpdateIsOneTap,
hasUpdate,
HostView,
needsPair,
@@ -183,8 +184,11 @@ const QamPanel: FC = () => {
onClick={() => applyUpdate(update!, check)}
label={
update!.update_available
? `Plugin v${update!.current} → v${update!.latest}${
update!.client_update_available ? " + client" : ""
? // "+ client" only when this tap will really install it. A manual-only
// client rides along as a toast with the command, and promising it in the
// label would make that read as a failure.
`Plugin v${update!.current} → v${update!.latest}${
clientUpdateIsOneTap(update) ? " + client" : ""
}`
: "New client version"
}
+46 -24
View File
@@ -325,6 +325,22 @@ mod session_main {
};
// Before the struct literal — `vulkan` moves into it below.
let phase_lock = vulkan.as_ref().is_some_and(|v| v.present_timing);
// …and the 4:4:4 promise, for the same reason: asked while the device bundle is
// still borrowable. `&&` short-circuits, so a box that never enabled Full chroma
// pays no capability queries for a feature it does not want.
let want_444 = settings.enable_444
&& pf_client_core::video::hevc_444_hardware_decodable(vulkan.as_ref());
if settings.enable_444 && !want_444 {
// Loud, because the user turned a switch on and is not getting it. The
// alternative is what this replaces: the host grants 4:4:4, the decode ladder
// has no rung that can take it, and the session drops HEVC entirely.
tracing::warn!(
"Full chroma (4:4:4) requested but this device has no 4:4:4 HEVC decode — \
asking for 4:2:0 instead. Advertising it would cost the whole codec: 4:4:4 \
is granted on HEVC only, and there is no software HEVC decoder to fall back \
to (PyroWave carries 4:4:4 on any GPU, if the link can take it)."
);
}
SessionParams {
host: addr,
port,
@@ -356,30 +372,16 @@ mod session_main {
// slice NALs, so the host may keep its multi-slice low-latency default (§7 LN1).
// The mobile/TV embedders must NOT copy this blindly — Amlogic MediaCodec wedges
// on multi-slice AUs (see `VIDEO_CAP_MULTI_SLICE`), so they advertise per-decoder.
// 4:4:4 is opt-in and off by default (Settings "Full chroma"): the bit only says
// 4:4:4 is opt-in and off by default (Settings "Full chroma"): the bit says
// "upgrade me if you can" — the host still gates on its own policy, its capturer,
// HEVC, and a real GPU 4:4:4 encode probe, and answers the resolved chroma in the
// Welcome BEFORE we build a decoder. Advertised whenever the user asks because
// every path can DISPLAY it: the Vulkan presenter samples the 2-plane 4:4:4 pool
// formats (hardware RExt decode where the driver offers it — NVIDIA today),
// with the decoder ladder demoting on its own. No capability probe gates the
// bit — but note (M8) that the software rung below it is 4:2:0 8-bit ONLY and
// refuses anything else rather than mis-scaling it, so on a box whose hardware
// 4:4:4 decode fails the floor is a codec fallback, not a converted picture.
// Welcome BEFORE we build a decoder. It is now ALSO gated on this device being
// able to decode 4:4:4 (`want_444`, computed above); the rule and its reasoning
// live in `video::video_caps_for`, which is where they get tested.
// The cost stays VISIBLE, not silent: the Detailed stats overlay prints the
// resolved chroma ("4:4:4→4:2:0" when the host declined) and the decode path
// frames actually took.
video_caps: punktfunk_core::quic::VIDEO_CAP_MULTI_SLICE
| if settings.hdr_enabled {
punktfunk_core::quic::VIDEO_CAP_10BIT | punktfunk_core::quic::VIDEO_CAP_HDR
} else {
0
}
| if settings.enable_444 {
punktfunk_core::quic::VIDEO_CAP_444
} else {
0
},
video_caps: pf_client_core::video::video_caps_for(settings.hdr_enabled, want_444),
// This panel's HDR colour volume → the host's virtual-display EDID, so host
// apps tone-map to the real glass. Windows reads it from DXGI (the
// `--window-pos` monitor; advanced-color outputs only) — gated on the HDR
@@ -496,6 +498,12 @@ mod session_main {
/// decode is already the default just no-ops. Append rather than clobber so a user's own
/// `RADV_PERFTEST` survives; `PUNKTFUNK_DECODER=native-vaapi` still overrides the decoder
/// choice (the pre-M10 `vaapi` spelling reaches the same rung — it migrates, loudly).
///
/// ⚠⚠ Called from the TOP of [`run`], ahead of the `--list-adapters` / `--probe-decode`
/// early exits — not merely "before `run_session` creates the instance". Those flags
/// create Vulkan instances of their own and RADV latches `RADV_PERFTEST` when its ICD
/// initialises, so a call placed after them leaves the triage tool describing a device
/// that cannot decode while the streaming path decodes on it.
#[cfg(target_os = "linux")]
fn enable_radv_video_decode() {
const TOKEN: &str = "video_decode";
@@ -579,6 +587,23 @@ mod session_main {
)
.init();
// Before ANY Vulkan call — and that includes the two probe flags below, which is the
// whole reason this sits at the top of `run` instead of beside the session setup it
// was written for. Make RADV expose its video-decode queue + extensions so the
// decoder's `auto` path prefers Vulkan Video over VAAPI (Steam Deck, and any gated
// RADV). Windows drivers (NVIDIA/AMD Adrenalin) expose theirs unconditionally.
//
// ⚠⚠ It USED to sit after the `--list-adapters` / `--probe-decode` / `--list-audio` /
// `--pair` early exits, which meant the triage tool answered a DIFFERENT question from
// the one the streaming path asks. Measured on a Steam Deck (2026-08-08, canary
// `e22af40f`), same binary, back to back: bare `--probe-decode` printed `vulkan video
// decode: no`, `driver decode ops: none (0x0)`, `no queue family advertises
// VIDEO_DECODE`; the same call with `RADV_PERFTEST=video_decode` in the environment
// printed `YES` and `H.264, H.265, AV1, VP9`. The tool exists to be believed, so any
// Deck triage that consulted it reached the opposite of the truth.
#[cfg(target_os = "linux")]
enable_radv_video_decode();
// `--list-adapters`: print the Vulkan physical devices' marketing names (one per
// line, discrete first) for the desktop shells' GPU picker, then exit.
if arg_flag("--list-adapters") {
@@ -753,11 +778,8 @@ mod session_main {
return headless_pair(&pin);
}
// Before any Vulkan call: make RADV expose its video-decode queue + extensions so the
// decoder's `auto` path prefers Vulkan Video over VAAPI (Steam Deck, and any gated RADV).
// Windows drivers (NVIDIA/AMD Adrenalin) expose theirs unconditionally.
#[cfg(target_os = "linux")]
enable_radv_video_decode();
// (The RADV video-decode opt-in that used to live here now runs at the very top of
// `run` — it has to precede the probe flags too, not just the session.)
// The Settings device picks → env, unless the user already forced one by hand:
// the GPU (the shells' pickers store the adapter's marketing name) for the
+23
View File
@@ -168,6 +168,10 @@ pub struct PortalCapturer {
/// downgrade ([`pf_zerocopy::note_raw_dmabuf_negotiation_failed`]) so the pipeline rebuild
/// retries on the CPU offer instead of failing identically forever.
vaapi_dmabuf: bool,
/// PW3: this capture's dmabuf offer has been confirmed to negotiate (a frame arrived), so the
/// negotiation retry budget has already been credited back. One-shot — the credit is per
/// capture, not per frame.
negotiation_confirmed: bool,
/// This capture ran the HDR (10-bit PQ/BT.2020 dmabuf) offer — see [`Self::open`]'s
/// `want_hdr`. Read by the negotiation-timeout diagnosis (a failed HDR offer latches the
/// process-wide SDR downgrade) and by [`hdr_meta`](Capturer::hdr_meta).
@@ -412,6 +416,7 @@ impl PwHandles {
signals: self.signals,
stall_since: None,
vaapi_dmabuf: self.vaapi_dmabuf,
negotiation_confirmed: false,
hdr_offer: self.hdr_offer,
hdr_source,
node_id,
@@ -468,6 +473,13 @@ fn spawn_pipewire(
} else {
want_hdr
};
// PW3: tell the raw-dmabuf latch which capture this is BEFORE reading its verdict below. A
// different node id is a different question — a fresh virtual output, a compositor restart,
// the Bazzite Gaming↔Desktop switch — and inheriting "dmabuf does not work here" from an
// unrelated capture is how one transient timeout used to cost a host CPU capture until it was
// restarted. The portal bit is in the key because a portal-fd capture and a virtual-output
// capture with the same node number are genuinely different sources.
pf_zerocopy::note_raw_dmabuf_capture(u64::from(node_id) | (u64::from(fd.is_some()) << 32));
// THE negotiation decision, resolved once here and handed to the thread — no mirror (L3/F1).
// Every environment/latch read the decision depends on happens at this single point.
let plan = pipewire::negotiation_plan(pipewire::NegotiationInputs {
@@ -705,6 +717,7 @@ impl PortalCapturer {
// The slot before the wakeup: a publish that coalesced its edge (or landed while we were
// not waiting) is still visible here.
if let Some(f) = self.take_frame() {
self.note_negotiation_confirmed();
return Ok(f);
}
let slice = Duration::from_millis(500)
@@ -728,6 +741,16 @@ impl PortalCapturer {
self.slot.lock().ok().and_then(|mut s| s.take())
}
/// PW3: a frame arrived, so this capture's dmabuf-only offer DID negotiate — credit the
/// negotiation retry budget back. Only meaningful for a capture that actually made that offer,
/// and only once per capture (the budget counts consecutive failed BUILDS, not frames).
fn note_negotiation_confirmed(&mut self) {
if self.vaapi_dmabuf && !self.negotiation_confirmed {
self.negotiation_confirmed = true;
pf_zerocopy::note_raw_dmabuf_negotiation_ok();
}
}
/// The [`frame_within`](Self::frame_within) budget expired (or the thread ended) — turn it
/// into the diagnosis-bearing error. Split out of the slicing loop above; behavior unchanged.
fn next_frame_timed_out(
File diff suppressed because it is too large Load Diff
+155 -6
View File
@@ -121,6 +121,38 @@ pub(super) fn build_dmabuf_format(
/// SDR — the same outcome as not offering HDR.
const SPA_VIDEO_TRANSFER_SMPTE2084: u32 = 14;
/// The two 10-bit PQ formats an HDR session offers, **in negotiation order**. The order is not a
/// style choice — on NVIDIA it is the difference between correct colour and red/blue swapped.
///
/// `xBGR_210LE` (DRM `XBGR2101010`, Vulkan `A2B10G10R10_UNORM_PACK32`) comes FIRST because the
/// first compatible consumer pod wins, and it is the only one gamescope fills correctly on every
/// vendor:
///
/// * `A2R10G10B10_UNORM_PACK32` **linear-tiled storage** is an optional Vulkan feature that
/// NVIDIA does not implement. gamescope's capture textures are mappable, hence linear, so on
/// NVIDIA its composite `imageStore` into that image lands in XBGR order — the bytes come out
/// byte-reversed while the buffer is still LABELLED `XRGB2101010`.
/// * The host believes the label: `xRGB_210LE → PixelFormat::X2Rgb10 →`
/// `NV_ENC_BUFFER_FORMAT_ARGB10`. Every mapping in that chain is individually correct, which is
/// exactly why the bug is invisible from this side — the *content* is what's wrong.
/// * Upstream gamescope hit the same wall and fixed it with `vulkan_get_rgb10_capture_format()`,
/// which probes `linearTilingFeatures` for STORAGE+SAMPLED and falls back to `XBGR2101010`.
/// That landed AFTER 3.16.25, so the pinned `punktfunk-gamescope` (3.16.25-7-g60561e2 +pfhdr4)
/// predates it and cannot self-correct — hence fixing the preference host-side, where it ships
/// in the host binary with no gamescope rebuild.
///
/// Preferring xBGR costs nothing anywhere else: `A2B10G10R10_UNORM_PACK32` is the universally
/// supported packed-10 format (it is the standard HDR10 swapchain format), it is what upstream
/// falls back to, and `X2Bgr10` has a first-class encoder path (NVENC `ABGR10`, VAAPI
/// `X2BGR10LE`). `xRGB_210LE` stays as the second pod so a producer that somehow offers only it
/// can still negotiate HDR rather than falling off to the SDR downgrade.
///
/// ⚠ The real fix belongs upstream in the patch set: `spa_format_to_drm()` should offer only the
/// format `vulkan_get_rgb10_capture_format()` reports. Until the gamescope pin moves past that
/// commit, THIS ORDER is what keeps NVIDIA HDR sessions correct — do not "tidy" it.
pub(super) const HDR_FORMAT_ORDER: [VideoFormat; 2] =
[VideoFormat::xBGR_210LE, VideoFormat::xRGB_210LE];
pub(super) fn build_hdr_dmabuf_format(
format: VideoFormat,
preferred: Option<(u32, u32, u32)>,
@@ -288,16 +320,57 @@ pub(super) fn build_shm_only_buffers() -> Result<Vec<u8>> {
})
}
/// Build a Buffers param requesting dmabuf-only buffers.
/// PW5 stage 2: the buffer-pool depth we ASK for on the zero-copy path, as a Choice range.
///
/// The zero-copy path hands the SPA buffer back to the producer at `.process` return, while the
/// encode thread still holds a dup of its dmabuf fd and has not yet imported, let alone read, the
/// contents. Nothing bounds that window — see the `queue_raw_buffer` comment in `pipewire.rs` — so
/// the only thing that keeps capture untorn is the producer round-robining a pool deeper than our
/// import+encode latency. Until PW5 stage 1 nobody had ever counted what that pool was; we never
/// even asked for a size (`build_dmabuf_buffers` set `dataType` and nothing else).
///
/// A **range**, deliberately, not a fixed count: SPA intersects the consumer's and producer's
/// Buffers params, so a fixed 8 against a producer that can only afford 4 empties the intersection
/// and the link silently stalls in "negotiating" — the exact failure mode the cursor-meta `size`
/// property already cost this codebase once (see `build_cursor_meta_param`). With a range the
/// producer clamps into it and negotiation still succeeds.
///
/// The numbers: `min` stays at 2 so nothing that works today stops working; `default` 8 is ~133 ms
/// of buffer at 60 Hz and ~33 ms at 240 Hz, comfortably past the ~3-4 ms capture→fence latency
/// measured in PW3/PW4 even with a second frame in flight; `max` 16 is a ceiling, not a request
/// (a 4K 4:4:4 buffer is ~25 MB, so 16 is ~400 MB of compositor allocation and worth capping).
/// **What the producer actually picks is logged by the stage-1 census — trust that line, not
/// these constants.**
const POOL_MIN: i32 = 2;
const POOL_DEFAULT: i32 = 8;
const POOL_MAX: i32 = 16;
/// Build a Buffers param requesting dmabuf-only buffers, with pool headroom (see [`POOL_DEFAULT`]).
pub(super) fn build_dmabuf_buffers() -> Result<Vec<u8>> {
serialize_pod(pw::spa::pod::Object {
type_: pw::spa::utils::SpaTypes::ObjectParamBuffers.as_raw(),
id: pw::spa::param::ParamType::Buffers.as_raw(),
properties: vec![pw::spa::pod::Property {
key: pw::spa::sys::SPA_PARAM_BUFFERS_dataType,
flags: pw::spa::pod::PropertyFlags::empty(),
value: pw::spa::pod::Value::Int(1i32 << pw::spa::sys::SPA_DATA_DmaBuf),
}],
properties: vec![
pw::spa::pod::Property {
key: pw::spa::sys::SPA_PARAM_BUFFERS_dataType,
flags: pw::spa::pod::PropertyFlags::empty(),
value: pw::spa::pod::Value::Int(1i32 << pw::spa::sys::SPA_DATA_DmaBuf),
},
pw::spa::pod::Property {
key: pw::spa::sys::SPA_PARAM_BUFFERS_buffers,
flags: pw::spa::pod::PropertyFlags::empty(),
value: pw::spa::pod::Value::Choice(pw::spa::pod::ChoiceValue::Int(
pw::spa::utils::Choice(
pw::spa::utils::ChoiceFlags::empty(),
pw::spa::utils::ChoiceEnum::Range {
default: POOL_DEFAULT,
min: POOL_MIN,
max: POOL_MAX,
},
),
)),
},
],
})
}
@@ -512,4 +585,80 @@ mod tests {
"libspa renumbered spa_video_transfer_function — update the hardcoded PQ id"
);
}
/// PW5 stage 2: the pool request must be a **Choice Range**, never a fixed Int.
///
/// This is the whole safety argument for asking at all: SPA intersects the two sides' Buffers
/// params, so a fixed count a producer cannot afford empties the intersection and the link
/// stalls in "negotiating" with no error anywhere — the same trap that cost this codebase the
/// entire Linux cursor channel once (see `build_cursor_meta_param`). Asserting the pod shape
/// is what keeps a later "simplify" from turning the range back into a number.
#[test]
fn the_dmabuf_pool_request_is_a_range_not_a_fixed_count() {
let pod = build_dmabuf_buffers().unwrap();
let key = spa::sys::SPA_PARAM_BUFFERS_buffers.to_ne_bytes();
let at = pod
.windows(4)
.position(|w| w == key)
.expect("the dmabuf Buffers pod must carry a buffers count");
let word = |off: usize| u32::from_ne_bytes(pod[off..off + 4].try_into().unwrap());
// Property = { key, flags, value_pod }; value_pod = { size, type, body }. A Choice body
// is { type: u32, flags: u32, child_size: u32, child_type: u32, values… }.
assert_eq!(
word(at + 12),
spa::sys::SPA_TYPE_Choice,
"the buffers count must be a Choice, not a bare Int — a fixed count can fail \
negotiation outright"
);
assert_eq!(
word(at + 16),
spa::sys::SPA_CHOICE_Range,
"the Choice must be a Range (default, min, max)"
);
assert_eq!(word(at + 24), 4, "Choice child pods are 4-byte Ints");
assert_eq!(word(at + 28), spa::sys::SPA_TYPE_Int, "…of type Int");
let vals: Vec<i32> = (0..3)
.map(|i| i32::from_ne_bytes(pod[at + 32 + i * 4..at + 36 + i * 4].try_into().unwrap()))
.collect();
assert_eq!(
vals,
vec![POOL_DEFAULT, POOL_MIN, POOL_MAX],
"Range values are serialized default-first"
);
// The minimum must not exceed what producers already serve, or the ask becomes a demand.
const { assert!(POOL_MIN <= 2) };
}
/// xBGR_210LE must be offered FIRST, and this is a correctness test, not a style one.
///
/// The first compatible consumer pod wins the negotiation. Leading with `xRGB_210LE` makes an
/// NVIDIA gamescope session land on `XRGB2101010`, whose linear-tiled `A2R10G10B10` storage
/// NVIDIA does not support — gamescope's composite `imageStore` writes XBGR bytes under an
/// XRGB label and the whole stream comes out with red and blue swapped. Every format mapping
/// on the host side is individually correct, so nothing downstream can detect it.
///
/// Field-confirmed 2026-08-09 on the RTX 5070 Ti Bazzite host with 0.26.0. See the
/// [`HDR_FORMAT_ORDER`] docs for the upstream fix this predates.
#[test]
fn hdr_offers_xbgr_before_xrgb() {
assert_eq!(
HDR_FORMAT_ORDER[0],
VideoFormat::xBGR_210LE,
"xBGR_210LE must be offered first — leading with xRGB_210LE swaps red and blue on \
every NVIDIA gamescope HDR session"
);
assert_eq!(
HDR_FORMAT_ORDER[1],
VideoFormat::xRGB_210LE,
"xRGB_210LE stays as the fallback pod so a producer offering only it can still \
negotiate HDR instead of dropping to the SDR downgrade"
);
// Both must still build: the order is a preference, never a removal.
for fmt in HDR_FORMAT_ORDER {
assert!(
!build_hdr_dmabuf_format(fmt, None).unwrap().is_empty(),
"{fmt:?} must still produce a format pod"
);
}
}
}
+2 -1
View File
@@ -547,7 +547,8 @@ pub struct IddPushCapturer {
_keepalive: Box<dyn Send>,
}
// SAFETY: `IddPushCapturer` is `!Send` only because of its `*mut SharedHeader` raw pointer (and the
// COM interfaces / the broker's bare control `HANDLE`, which is process-global and never closed). It is
// COM interfaces; the frame/cursor delivery closures own `Arc` clones of the control device and are
// `Send + Sync` on their own). It is
// created, used, and dropped by a SINGLE thread — the owning capture/encode thread — never shared: the
// `ID3D11DeviceContext` is the device's IMMEDIATE context (single-threaded by D3D11 contract) and is
// only ever touched from that thread, and the header pointer (into the mapping this struct owns) is
+6
View File
@@ -134,6 +134,12 @@ pf-vaadec = { path = "../pf-vaadec" }
# container can then compile and clippy the whole rung without `libva-dev`, and a machine
# without a VAAPI runtime gets a clean refusal instead of a packaging dependency.
libloading = "0.8"
# The gamescope overlay watcher (`overlay_focus`): read two CARDINAL properties off a
# gamescope root window and block on PropertyNotify. `default-features = false` keeps the
# pure-Rust `RustConnection` — no libxcb link, so no new C dependency on any client package
# — the same stance pf-capture and pf-vdisplay already take on this crate. No extension
# features: root-window properties and an event mask are core X11.
x11rb = { version = "0.13", default-features = false }
[target.'cfg(windows)'.dependencies]
wasapi = "0.23"
+164 -1
View File
@@ -381,6 +381,7 @@ enum Ctl {
PadAudioPrefs(u8),
MenuMode(bool),
MenuRumble(MenuPulse),
Mask(bool),
}
#[derive(Clone)]
@@ -548,6 +549,31 @@ impl GamepadService {
let _ = self.ctl.send(Ctl::Forwarding(on));
}
/// A system overlay owns the controller right now — hold every forwarded pad NEUTRAL
/// until it closes. This is the Steam Input behaviour a streaming client has to
/// reproduce by hand: while the Deck's Steam menu or QAM is up, the same physical
/// sticks and buttons drive Steam's UI, and anything we keep forwarding lands in the
/// game underneath as a second, invisible player.
///
/// **Masking is not [`set_forwarding`](Self::set_forwarding).** Forwarding-off closes the
/// slot and sends the host a [`GamepadRemove`](InputKind::GamepadRemove) — the game sees a
/// controller *unplug*, which is a hardware event with real in-game consequences (pause
/// menus, "reconnect your controller", player-slot churn). Opening the QAM must not look
/// like that. Masking keeps every slot open and merely stops the transitions, after
/// flushing what the host believes is held so a stick held at overlay-open stops steering
/// instead of freezing at its last value.
///
/// SDL has this gate of its own — it drops presses while the process has windows but no
/// keyboard focus — and on a desktop it fires. It CANNOT fire on a Deck in Gaming Mode:
/// gamescope resolves focus per Xwayland ctx, and the client sits alone in its own ctx, so
/// its X input focus never moves when the overlay takes over (measured). That is why this
/// exists as an explicit lever rather than something inherited for free.
///
/// Held state is adopted, not replayed, on the way back — see [`Ctl::Mask`]'s handling.
pub fn set_masked(&self, on: bool) {
let _ = self.ctl.send(Ctl::Mask(on));
}
/// The session's system-button policy, resolved from
/// [`Settings::system_buttons_forward`] × [`Settings::guide_gesture_enabled`]:
/// `forward_raw` gates the physical guide/QAM presses onto the wire (off = they stay
@@ -1069,6 +1095,9 @@ struct Worker {
menu_mode: bool,
menu_nav: MenuNav,
menu_tx: async_channel::Sender<MenuEvent>,
/// A system overlay owns input ([`GamepadService::set_masked`]): forwarded pads are held
/// neutral and menu translation is paused, with every slot still OPEN.
masked: bool,
}
impl Worker {
@@ -1519,6 +1548,87 @@ impl Worker {
}
}
/// Re-adopt what the pads are physically holding when an overlay mask lifts.
///
/// Buttons are taken back into `held_buttons` **without** a wire press: a button pressed
/// inside the overlay (the A that picked a QAM row) must not fire in the game the instant it
/// closes — releasing it and pressing again is what arms it. Same rule menu mode already
/// applies across a screen handoff ([`MenuNav::reset`]), for the same reason.
///
/// Axes ARE re-sent, because a stick has no press semantics to ghost — it is deflected or it
/// is not. The mask flushed them to zero, and SDL only speaks on *change*, so a stick still
/// held when the overlay closes would stay dead host-side until the user happened to move it.
///
/// Neither half can run against a pad that is gone: this only walks open slots, and every SDL
/// read here is a state query on a handle the slot owns.
fn readopt_held(&mut self) {
use sdl3::gamepad::{Axis, Button};
// Every button `button_bit` maps — the same surface the press path forwards.
const BUTTONS: [Button; 21] = [
Button::South,
Button::East,
Button::West,
Button::North,
Button::Back,
Button::Start,
Button::Guide,
Button::LeftStick,
Button::RightStick,
Button::LeftShoulder,
Button::RightShoulder,
Button::DPadUp,
Button::DPadDown,
Button::DPadLeft,
Button::DPadRight,
Button::Touchpad,
Button::RightPaddle1,
Button::LeftPaddle1,
Button::RightPaddle2,
Button::LeftPaddle2,
Button::Misc1,
];
const AXES: [Axis; 6] = [
Axis::LeftX,
Axis::LeftY,
Axis::RightX,
Axis::RightY,
Axis::TriggerLeft,
Axis::TriggerRight,
];
// Copied out: the slot walk below borrows `self` mutably.
let system_forward = self.system_forward;
let attached = self.attached.clone();
for slot in &mut self.slots {
slot.held_buttons.clear();
for b in BUTTONS {
let Some(bit) = button_bit(b) else {
continue;
};
// The press path returns before `held_buttons` for un-forwarded system
// buttons; tracking them here would invent state it never keeps.
if !system_forward && matches!(bit, wire::BTN_GUIDE | wire::BTN_MISC1) {
continue;
}
if slot.pad.button(b) {
slot.held_buttons.push(bit);
}
}
let Some(c) = &attached else {
continue;
};
for a in AXES {
let (id, v) = axis_value(a, slot.pad.axis(a));
if slot.last_axis[id as usize] != v {
slot.last_axis[id as usize] = v;
send(c, InputKind::GamepadAxis, id, v, slot.index);
}
}
}
// The chord latch was cleared on the way in; drop it again if what we just adopted
// doesn't actually hold it.
self.rearm_escape();
}
/// True when any one forwarded pad holds the entire escape chord (any player can leave).
fn chord_held(&self) -> bool {
self.slots
@@ -1785,6 +1895,34 @@ impl Worker {
.push((pad, bit, Instant::now() + TAP_PRESS));
}
}
Ok(Ctl::Mask(on)) => {
if self.masked == on {
continue;
}
self.masked = on;
if on {
// Neutral NOW, and while the slots stay open: a stick held when the
// overlay opened must stop steering, but the host must not see the pad
// unplug (that is `close_slot_at`'s job, and a game reacts to it).
if let Some(c) = self.attached.clone() {
for slot in &mut self.slots {
Self::flush_slot(&c, slot);
}
}
// Nothing can be mid-chord across the flip: the transitions that would
// complete or break it are about to be dropped.
self.reset_chord();
} else {
// Coming back. Whatever is still physically held was never delivered —
// adopt it silently rather than replay it as a fresh press, the same
// rule menu mode uses across a screen handoff (`MenuNav::reset`). A
// button you pressed *inside* the overlay must not fire in the game the
// instant it closes; releasing and pressing again is what arms it.
self.readopt_held();
self.menu_nav.reset();
}
tracing::info!(masked = on, "overlay input mask");
}
Ok(Ctl::Forwarding(on)) => {
if self.forwarding == on {
continue;
@@ -1846,6 +1984,28 @@ impl Worker {
/// "is a session live".
fn handle_event(&mut self, event: sdl3::event::Event) {
use sdl3::event::Event;
// A system overlay owns the controller ([`GamepadService::set_masked`]): drop every
// input transition. The pads were flushed neutral when the mask went on, so dropping
// the ups as well as the downs is what keeps the two in agreement — `readopt_held`
// rebuilds the held set from the hardware when it lifts.
//
// Device add/remove deliberately still count: a controller genuinely plugged in or
// pulled out behind an overlay is a fact about the world, not an input, and losing it
// would leave the slot table lying about what exists.
if self.masked
&& matches!(
event,
Event::ControllerButtonDown { .. }
| Event::ControllerButtonUp { .. }
| Event::ControllerAxisMotion { .. }
| Event::ControllerTouchpadDown { .. }
| Event::ControllerTouchpadMotion { .. }
| Event::ControllerTouchpadUp { .. }
| Event::ControllerSensorUpdated { .. }
)
{
return;
}
match event {
Event::ControllerDeviceAdded { which, .. } => {
if !self.order.contains(&which) {
@@ -2074,7 +2234,9 @@ impl Worker {
/// on and no session is attached (attach supersedes; SDL events merely wake the loop,
/// so a press is translated the iteration it arrives).
fn menu_poll(&mut self) {
if !self.menu_mode || self.attached.is_some() {
// Masked covers the launcher too: with the Deck's Steam menu up over our console, the
// same stick that scrolls Steam's UI would otherwise also be scrolling ours behind it.
if !self.menu_mode || self.attached.is_some() || self.masked {
return;
}
let Some((_, pad)) = self.menu_open.as_ref() else {
@@ -2301,6 +2463,7 @@ impl Worker {
menu_mode: false,
menu_nav: MenuNav::new(),
menu_tx,
masked: false,
}
}
}
+4
View File
@@ -46,6 +46,10 @@ pub mod orchestrate;
// The host's OS-identity chain (mDNS `os=` TXT): sanitize + icon-walk order. Pure string
// logic, built everywhere (the Apple/Android ports mirror it rather than link it).
pub mod os;
// "A system overlay owns the controller" for gamescope Gaming Mode — the signal behind the
// gamepad input mask, which SDL's own focus gate structurally cannot provide there.
#[cfg(target_os = "linux")]
pub mod overlay_focus;
// Client settings profiles: the override catalog + the one connect-time resolver
// (design/client-settings-profiles.md §4). Sits beside `trust`, which owns the host records
// the bindings live on.
+284
View File
@@ -0,0 +1,284 @@
//! "A system overlay owns the controller right now" — the gamescope half of the input mask.
//!
//! On a Steam Deck in Gaming Mode the Steam menu and the QAM are drawn by Steam and driven by
//! the *same physical controller* the client is forwarding. Steam does not mask us the way it
//! masks a normal game: masking happens on Steam Input's virtual pad, and the client
//! deliberately forwards the REAL pad instead (28DE:1205 — the virtual one has no gyro,
//! trackpads or paddles). So while the QAM is up, one thumbstick drives Steam's UI *and* the
//! game on the host. This watcher is what tells [`crate::gamepad::GamepadService::set_masked`]
//! to stop that.
//!
//! **Why the free mechanism can't do it.** SDL already drops gamepad presses while the process
//! has windows but no keyboard focus (`SDL_PrivateJoystickShouldIgnoreEvent`, on by default —
//! we never set `SDL_JOYSTICK_ALLOW_BACKGROUND_EVENTS`), and on a desktop that fires. It cannot
//! fire here: gamescope resolves focus **per Xwayland ctx** (`determine_and_apply_focus` scans
//! only that ctx's window list), the Steam overlay lives in the root ctx, and the client sits
//! alone in its own. Measured on a Deck 2026-08-08: with the QAM open, X input focus inside the
//! client's ctx never moved off its window, so no `FocusOut` is ever generated. Hence an
//! explicit signal.
//!
//! **The signal.** gamescope publishes two CARDINALs on the ROOT ctx's root window (Steam mode
//! only, i.e. `gamescope -e` — which is what Gaming Mode runs):
//!
//! * `GAMESCOPE_FOCUSED_APP` — appid of the window holding **input** focus
//! * `GAMESCOPE_FOCUSED_APP_GFX` — appid of the window being **displayed**
//!
//! They are equal in normal play and diverge exactly while something else has taken input over
//! the running app. Measured, both for the Steam menu and for the QAM:
//!
//! ```text
//! app=3856846079 gfx=3856846079 ← streaming, we own input
//! app=769 gfx=3856846079 ← overlay open (769 = Steam)
//! ```
//!
//! Note `app != gfx` rather than "app is Steam": anything that takes input away from the
//! displayed app is a thing we should stop forwarding through, and comparing to our own appid
//! would need us to know it (a non-Steam shortcut's appid is assigned by Steam at creation).
//!
//! **Which display.** Not necessarily ours. Gaming Mode runs `gamescope --xwayland-count 2`:
//! Steam and the atoms live on the first server, the app is given the second, and the client's
//! own `$DISPLAY` therefore has none of these properties. So discovery walks candidates — our
//! `$DISPLAY` first (correct for a single-server gamescope), then every socket in
//! `/tmp/.X11-unix` — and keeps the first whose root actually carries both atoms. gamescope's
//! Xwayland accepts unauthenticated local connections (verified: `xprop` against it succeeds
//! with no `.Xauthority` at all), so no cookie plumbing is needed.
//!
//! Everything here is best-effort by construction: no gamescope, no X, a sandbox that cannot
//! see the other socket, or a session that restarts underneath us all end in "no signal", which
//! degrades to exactly the behaviour that shipped before this module existed.
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;
use std::time::Duration;
use x11rb::connection::Connection;
use x11rb::protocol::xproto::{
Atom, AtomEnum, ChangeWindowAttributesAux, ConnectionExt, EventMask, Window,
};
use x11rb::protocol::Event;
use x11rb::rust_connection::RustConnection;
/// How long to wait before rebuilding everything after the X connection drops. Gaming Mode
/// recreates its Xwayland servers across a session restart, so "gone" is not permanent — but it
/// is also not worth a hot retry loop.
const RECONNECT_DELAY: Duration = Duration::from_secs(3);
/// Live "an overlay owns input" flag, updated by a background thread.
///
/// Cheap to poll (one relaxed atomic load), which is what the presenter's event loop wants — it
/// checks once per iteration and only talks to the gamepad service on an edge.
pub struct OverlayFocus {
open: Arc<AtomicBool>,
}
impl OverlayFocus {
/// Start watching, or return `None` when this isn't a gamescope Steam session (the common
/// case — every desktop client) or the user opted out with `PUNKTFUNK_OVERLAY_MASK=0`.
///
/// Returning `None` is not a failure: the caller keeps its window-focus path, which is the
/// right signal everywhere the compositor actually moves focus.
pub fn start() -> Option<OverlayFocus> {
if std::env::var("PUNKTFUNK_OVERLAY_MASK").is_ok_and(|v| v == "0" || v == "false") {
tracing::info!("overlay input mask disabled by PUNKTFUNK_OVERLAY_MASK");
return None;
}
if !gamescope_session() {
return None;
}
let open = Arc::new(AtomicBool::new(false));
let flag = open.clone();
std::thread::Builder::new()
.name("punktfunk-overlay-focus".into())
.spawn(move || watch(&flag))
.map_err(|e| tracing::warn!(error = %e, "overlay focus watcher failed to start"))
.ok()?;
Some(OverlayFocus { open })
}
/// Does something other than the displayed app own input right now?
pub fn is_open(&self) -> bool {
self.open.load(Ordering::Relaxed)
}
}
/// Gaming Mode / any gamescope session — the only place this signal exists. Mirrors the same
/// env checks the shells already use to detect Gaming Mode.
fn gamescope_session() -> bool {
std::env::var_os("GAMESCOPE_WAYLAND_DISPLAY").is_some()
|| std::env::var_os("SteamDeck").is_some()
|| std::env::var("XDG_CURRENT_DESKTOP").is_ok_and(|d| d.eq_ignore_ascii_case("gamescope"))
}
/// Displays worth trying, in order: ours first (a single-server gamescope publishes the atoms on
/// the display the app is already on), then every other socket present. `/tmp/.X11-unix` is
/// listed rather than probing `:0..:N` blindly so we never connect to a display that isn't there.
fn candidate_displays() -> Vec<String> {
let mut out = Vec::new();
if let Ok(d) = std::env::var("DISPLAY") {
if !d.is_empty() {
out.push(d);
}
}
if let Ok(entries) = std::fs::read_dir("/tmp/.X11-unix") {
let mut found: Vec<String> = entries
.flatten()
.filter_map(|e| {
let name = e.file_name().into_string().ok()?;
let n = name.strip_prefix('X')?;
n.parse::<u32>().ok().map(|n| format!(":{n}"))
})
.collect();
found.sort();
for d in found {
if !out.contains(&d) {
out.push(d);
}
}
}
out
}
/// The two atoms on a root that carries them, or `None` for a display that isn't gamescope's
/// root ctx. `only_if_exists` keeps this from interning atoms into unrelated X servers.
fn gamescope_atoms(conn: &RustConnection) -> Option<(Atom, Atom)> {
let app = conn
.intern_atom(true, b"GAMESCOPE_FOCUSED_APP")
.ok()?
.reply()
.ok()?
.atom;
let gfx = conn
.intern_atom(true, b"GAMESCOPE_FOCUSED_APP_GFX")
.ok()?
.reply()
.ok()?
.atom;
(app != 0 && gfx != 0).then_some((app, gfx))
}
/// Read one CARDINAL appid. gamescope writes these with a length of ZERO when the appid is 0
/// (`focusedAppId != 0 ? 1 : 0`), so "present but empty" is a real state meaning "no app" — it
/// must read as `None`, not as `Some(0)` that would then compare unequal to everything.
fn read_appid(conn: &RustConnection, root: Window, atom: Atom) -> Option<u32> {
let reply = conn
.get_property(false, root, atom, AtomEnum::CARDINAL, 0, 1)
.ok()?
.reply()
.ok()?;
// Bound rather than returned inline: the iterator borrows `reply`, and as a tail
// expression its temporary would outlive it.
let id = reply.value32()?.next();
id
}
/// The whole decision, separated from X so it can be tested: an overlay is up exactly when
/// input focus and the displayed app are both known and DIFFER.
///
/// Absence is never an overlay. A missing value means "no app focused" (gamescope's zero-length
/// write) or "this display stopped answering" — and a mask that latched on when the signal went
/// away would silently kill the controller for the rest of the session, which is a far worse
/// failure than not masking at all.
fn overlay_open_from(app: Option<u32>, gfx: Option<u32>) -> bool {
matches!((app, gfx), (Some(a), Some(g)) if a != g)
}
/// True when input focus and the displayed app have diverged — an overlay is up.
fn overlay_open(conn: &RustConnection, root: Window, app: Atom, gfx: Atom) -> bool {
overlay_open_from(read_appid(conn, root, app), read_appid(conn, root, gfx))
}
/// Connect, find the root ctx, then block on PropertyNotify for the two atoms. Returns on any X
/// error so the outer loop can rebuild after a session restart.
fn watch(flag: &Arc<AtomicBool>) {
loop {
if let Some((conn, root, app, gfx)) = connect() {
// Seed before the first event: the overlay may already be up when we start.
flag.store(overlay_open(&conn, root, app, gfx), Ordering::Relaxed);
loop {
match conn.wait_for_event() {
Ok(Event::PropertyNotify(e)) if e.atom == app || e.atom == gfx => {
let open = overlay_open(&conn, root, app, gfx);
if flag.swap(open, Ordering::Relaxed) != open {
tracing::debug!(open, "gamescope overlay focus changed");
}
}
Ok(_) => {}
Err(e) => {
tracing::info!(error = %e, "gamescope focus watcher disconnected");
break;
}
}
}
// A dropped connection tells us nothing about the controller — unmask, or a
// gamescope restart mid-overlay would leave the pad dead with nothing to revive it.
flag.store(false, Ordering::Relaxed);
}
std::thread::sleep(RECONNECT_DELAY);
}
}
/// The first candidate display whose root carries both atoms, with PropertyNotify selected.
fn connect() -> Option<(RustConnection, Window, Atom, Atom)> {
for dpy in candidate_displays() {
// `dpy`, not `display`: `display` is one of tracing's own value helpers, and a field
// named after it resolves to the helper inside the macro rather than to this string.
let Ok((conn, screen_num)) = RustConnection::connect(Some(&dpy)) else {
continue;
};
let Some((app, gfx)) = gamescope_atoms(&conn) else {
continue;
};
let root = conn.setup().roots[screen_num].root;
// Both atoms must actually be PRESENT on this root, not merely interned: a second
// gamescope Xwayland knows the atom names (they are per-server strings) but only the
// root ctx publishes the values.
if read_appid(&conn, root, gfx).is_none() {
continue;
}
// Checked rather than fire-and-forget: an event mask that silently failed to apply
// would leave the watcher blocked forever on a display that never speaks to it.
let selected = match conn.change_window_attributes(
root,
&ChangeWindowAttributesAux::new().event_mask(EventMask::PROPERTY_CHANGE),
) {
Ok(cookie) => cookie.check().is_ok(),
Err(_) => false,
};
if !selected {
continue;
}
tracing::info!(dpy, "watching gamescope focus for overlay input masking");
return Some((conn, root, app, gfx));
}
None
}
#[cfg(test)]
mod tests {
use super::*;
/// The measured Deck states, both directions (2026-08-08, Steam menu and QAM alike):
/// equal appids while we own input, divergent while the overlay does.
#[test]
fn divergent_appids_are_an_overlay() {
assert!(!overlay_open_from(Some(3856846079), Some(3856846079)));
assert!(overlay_open_from(Some(769), Some(3856846079)));
}
/// gamescope writes these properties with a length of ZERO when the appid is 0, so "no app"
/// arrives as a missing value rather than `Some(0)`. Reading it as `Some(0)` would make it
/// differ from every real appid and mask the pad on an empty Gaming Mode home screen.
#[test]
fn a_missing_appid_is_never_an_overlay() {
assert!(!overlay_open_from(None, Some(3856846079)));
assert!(!overlay_open_from(Some(769), None));
assert!(!overlay_open_from(None, None));
}
/// The safety property that outranks the feature: if the signal is unreadable we forward as
/// before. A latched mask would leave a streaming session with a dead controller and no way
/// back short of restarting it.
#[test]
fn absence_fails_open_not_closed() {
assert!(!overlay_open_from(None, None));
}
}
+6 -6
View File
@@ -1174,12 +1174,12 @@ pub struct Settings {
/// mirrors the Apple client's "Show game library" toggle, default off.
pub library_enabled: bool,
/// Which colour family the gamepad UI's living backdrop drifts through — the shared
/// `ui_palette` key (`"violet"` = the brand default, then `tide`/`forest`/`ember`/
/// `rose`/`graphite`; see `pf-console-ui`'s palette table, and the Apple/Android
/// clients' twins). Presentation only: nothing about a stream depends on it, which is
/// why it is a device preference and never part of a settings profile. An unknown
/// name reads as the default rather than erroring — a newer client may have shipped a
/// palette this binary doesn't know.
/// `ui_palette` key (`"violet"` = the brand default, then `oled`/`nebula`/`abyss`/
/// `ember`/`moss`/`graphite`, then the six pale fields; see `pf-console-ui`'s palette
/// table, and the Apple/Android clients' twins). Presentation only: nothing about a
/// stream depends on it, which is why it is a device preference and never part of a
/// settings profile. An unknown name reads as the default rather than erroring — a
/// newer client may have shipped a palette this binary doesn't know.
#[serde(default = "default_ui_palette")]
pub ui_palette: String,
/// Send Wake-on-LAN before connecting to a saved host and wait for it to boot (the
+126
View File
@@ -1531,6 +1531,85 @@ pub fn av1_hardware_decodable(vk: Option<&VulkanDecodeDevice>) -> bool {
d3d11
}
/// Can this client actually DECODE 4:4:4 HEVC — the question `VIDEO_CAP_444` is a promise
/// about, and the one nothing asked until a Steam Deck lost HEVC over it.
///
/// The bit used to ride the "Full chroma" toggle alone, with a comment saying the software
/// rung was the floor underneath it. M8 removed that floor: there is no CPU HEVC decoder at
/// all ([`software_decodable_codecs`]), and the host grants 4:4:4 only on HEVC. So on a
/// device with no 4:4:4 decode the toggle did not cost crispness — it cost the whole codec.
/// The Welcome resolves the chroma before a decoder exists, the native Vulkan constructor
/// then refuses the shape, VAAPI refuses it too, and the session reconnects on H.264 with
/// "HEVC decoding failed on this device" (field report 2026-08-08, Deck / VanGogh).
///
/// ⭐ Answered from the VULKAN rung alone, and that is exact rather than approximate: it is
/// the only rung in this build that implements 4:4:4 at all. `pf_vaadec::profile_for` maps
/// only `chroma_format_idc == 1` and errors `UnsupportedShape` on 3; `pf_dxvadec`'s config
/// refuses "anything but 4:2:0" by construction; the CPU rung is 8-bit 4:2:0 only. So a
/// device whose Vulkan driver offers no 4:4:4 decode profile has no 4:4:4 path in this
/// client, whatever its silicon can do. (That is why an Intel box — whose hardware HAS done
/// HEVC 4:4:4 since Ice Lake — is still a `false` here: our DXVA/VAAPI rungs do not
/// implement it, so advertising it would be a lie about US, not about the GPU.)
///
/// ⚠ Both depths are required, not either: with HDR on, the host may resolve 4:4:4 **10-bit**,
/// and a device offering `YUV444_8` but not `YUV444_10` would land in exactly the hole this
/// closes. Asking for both costs one extra capability query and removes the case entirely.
///
/// ⚠ Deliberately NOT extended to `VIDEO_CAP_10BIT`/`VIDEO_CAP_HDR`, which are advertised
/// unprobed for the same reason this one was. The asymmetry is real: all three hardware
/// rungs implement 10-bit 4:2:0 (`profile_for` maps `(H265, 1, 10)` and `(Av1, 1, 10)`;
/// pf-dxvadec carries P010), so a Vulkan-only probe there would answer `false` on boxes
/// whose VAAPI/DXVA rung decodes 10-bit perfectly and would silently withdraw HDR from
/// them — a visible regression bought against a case that has never been observed. Gating
/// 10-bit honestly needs a libva/D3D11 probe, which this path cannot afford (same reason
/// [`av1_hardware_decodable`] does not consult VAAPI).
pub fn hevc_444_hardware_decodable(vk: Option<&VulkanDecodeDevice>) -> bool {
#[cfg(any(target_os = "linux", windows))]
{
vk.is_some_and(|v| {
crate::video_vk_native::hevc_shape_supported(v, CHROMA_444, 0)
&& crate::video_vk_native::hevc_shape_supported(v, CHROMA_444, 2)
})
}
// No native Vulkan rung is compiled in off the two desktop OSes, so nothing here can
// decode 4:4:4 and the honest answer is a constant.
#[cfg(not(any(target_os = "linux", windows)))]
{
let _ = vk;
false
}
}
/// `chroma_format_idc` for 4:4:4 (H.265 7.4.3.2) — spelled once so the two depth probes
/// above and any future caller cannot disagree about the magic number.
const CHROMA_444: u8 = 3;
/// The desktop session's `video_caps` bitfield, as a pure function of the two user
/// switches that move it — so the rule can be tested without a GPU, a host or a Hello.
///
/// `want_444` is the "Full chroma" setting **already ANDed with this device's ability to
/// decode it** ([`hevc_444_hardware_decodable`]). Split that way on purpose: the caller
/// owns the expensive driver question and can log its own refusal with the user's setting
/// in hand, while the bit arithmetic — the part that was wrong — stays testable.
///
/// `MULTI_SLICE` is unconditional and is decoder truth for THIS embedder: every desktop
/// decode stack (Vulkan Video, D3D11VA, VAAPI, openh264/rav1d) handles AUs carrying
/// several slice NALs, so the host may keep its multi-slice low-latency default (§7 LN1).
/// ⚠ The mobile/TV embedders must NOT copy this blindly — Amlogic MediaCodec wedges on
/// multi-slice AUs (see `VIDEO_CAP_MULTI_SLICE`), so they advertise per-decoder.
///
/// HDR off means 10-bit is not advertised either, so the host never upgrades depth.
pub fn video_caps_for(hdr_enabled: bool, want_444: bool) -> u8 {
let mut caps = punktfunk_core::quic::VIDEO_CAP_MULTI_SLICE;
if hdr_enabled {
caps |= punktfunk_core::quic::VIDEO_CAP_10BIT | punktfunk_core::quic::VIDEO_CAP_HDR;
}
if want_444 {
caps |= punktfunk_core::quic::VIDEO_CAP_444;
}
caps
}
/// [`decodable_codecs`] plus the PyroWave bit when the presenter's device passed the
/// compute-feature probe, minus the codecs `decoder_pref` makes unreachable.
/// Advertisement-only: `resolve_codec` never auto-picks PyroWave — the session must also
@@ -2701,6 +2780,53 @@ mod tests {
use super::*;
use punktfunk_core::quic::{CODEC_AV1, CODEC_H264, CODEC_HEVC, CODEC_PYROWAVE};
/// The 4:4:4 advertisement is a PROMISE, and M8 removed the floor that used to make a
/// broken one survivable: there is no CPU HEVC decoder, and the host grants 4:4:4 on
/// HEVC only, so advertising it on a device that cannot decode it costs the entire
/// codec (field 2026-08-08, Steam Deck / VanGogh — HEVC fell back to H.264).
///
/// The device question needs a GPU; THIS is the half that does not, and it is the half
/// that was wrong — the bit used to ride `enable_444` alone.
#[test]
fn the_444_bit_needs_the_setting_and_a_device_that_can_decode_it() {
const V444: u8 = punktfunk_core::quic::VIDEO_CAP_444;
// The regression itself: setting on, device can't → the bit must NOT go out.
assert_eq!(
video_caps_for(true, false) & V444,
0,
"a 4:4:4 promise this device cannot keep costs HEVC entirely"
);
// ...and the feature still works where it can be honoured.
assert_ne!(video_caps_for(true, true) & V444, 0);
// Never advertised unasked, whatever the device can do.
assert_eq!(video_caps_for(true, false) & V444, 0);
assert_eq!(video_caps_for(false, false) & V444, 0);
// The 4:4:4 gate must not disturb the other two bits (10-bit/HDR is deliberately
// NOT probe-gated — see `hevc_444_hardware_decodable`'s docs for why).
const HDR_BITS: u8 =
punktfunk_core::quic::VIDEO_CAP_10BIT | punktfunk_core::quic::VIDEO_CAP_HDR;
for want_444 in [false, true] {
assert_eq!(video_caps_for(true, want_444) & HDR_BITS, HDR_BITS);
assert_eq!(video_caps_for(false, want_444) & HDR_BITS, 0);
assert_ne!(
video_caps_for(false, want_444) & punktfunk_core::quic::VIDEO_CAP_MULTI_SLICE,
0,
"MULTI_SLICE is unconditional for this embedder"
);
}
}
/// No presenter Vulkan device ⇒ no 4:4:4, and that is an ANSWER rather than a missing
/// one: the native Vulkan rung is the only one in this build that implements 4:4:4 at
/// all (`pf_vaadec::profile_for` errors on `chroma_format_idc == 3`, pf-dxvadec refuses
/// anything but 4:2:0, the CPU rung is 8-bit 4:2:0). The `Some` arm needs real hardware
/// and lives in the GPU suites.
#[test]
fn no_vulkan_device_means_no_444_promise() {
assert!(!hevc_444_hardware_decodable(None));
}
/// The reconnect rule, as the invariant it is: an exhausted codec must come back as
/// one this client can decode ALL THE WAY DOWN, and must never come back as itself.
///
+66 -15
View File
@@ -216,6 +216,70 @@ fn submit_queues_collide(graphics_qf: u32, decode_qf: u32) -> bool {
graphics_qf == decode_qf
}
/// The queue lock this device's decode lane submits under. One function so the
/// pre-session shape probe ([`hevc_shape_supported`]) and the real decoder cannot pick
/// different serialization for the same device.
fn queue_lock_for(vk: &VulkanDecodeDevice) -> Box<dyn pf_vkdecode::QueueLock> {
if submit_queues_collide(vk.graphics_qf, vk.decode_qf) {
Box::new(NativeQueueLock::Shared(vk.queue_lock.clone()))
} else {
Box::new(NativeQueueLock::Uncontended)
}
}
/// The presenter's handles in pf-vkdecode's shape. Same reason as [`queue_lock_for`]:
/// the probe must ask about the DEVICE THE SESSION WOULD USE, not a re-derived one.
fn device_handles(vk: &VulkanDecodeDevice) -> DeviceHandles {
DeviceHandles {
get_instance_proc_addr: vk.get_instance_proc_addr,
instance: vk.instance,
physical_device: vk.physical_device,
device: vk.device,
decode_qf: vk.decode_qf,
decode_queue_index: DECODE_QUEUE_INDEX,
graphics_qf: vk.graphics_qf,
}
}
/// Can this device hardware-decode HEVC at the given picture shape? Asked BEFORE the
/// Hello, so the client never advertises a shape it would have to refuse a session over.
///
/// This is the same question, through the same code, that
/// [`NativeVulkanDecoder::new`]'s H.265 arm asks at construction — `VkH265Decoder::new`
/// then `probe_stream_support` — deliberately, so an advertisement and the rung that has
/// to honour it cannot disagree. It creates and drops a decoder object; that costs a
/// handful of driver capability queries and no session, no images and no submits.
///
/// `false` when the presenter has no Vulkan Video decode at all, which for 4:4:4 is the
/// right answer rather than a missing one — see
/// [`crate::video::hevc_444_hardware_decodable`] for why no other rung can be asked.
pub(crate) fn hevc_shape_supported(
vk: &VulkanDecodeDevice,
chroma_format_idc: u8,
bit_depth_luma_minus8: u8,
) -> bool {
if !vk.video_decode {
return false;
}
// The device-independent half first: a shape pf-vkdecode has no picture format for
// needs no driver to refuse it (and `probe_stream_support` would only re-derive it).
if pf_vkdecode::output_format_for(chroma_format_idc, bit_depth_luma_minus8).is_none() {
return false;
}
// SAFETY: the `DeviceHandles` contract exactly as `NativeVulkanDecoder::new` states
// it — these are the presenter's live instance/device, which outlive this call by
// construction (the presenter owns them for the whole process, and this runs on its
// thread while building the session's Hello). The decoder is dropped before return,
// so nothing outlives the borrow.
let dec = unsafe { pf_vkdecode::VkH265Decoder::new(&device_handles(vk), queue_lock_for(vk)) };
match dec {
Ok(d) => d
.probe_stream_support(chroma_format_idc, bit_depth_luma_minus8)
.is_ok(),
Err(_) => false,
}
}
/// [`pf_vkdecode::QueueLock`] over the device's shared [`crate::video::QueueLock`] —
/// or over nothing, when the decode queue provably has no other submitter (see the
/// module doc's queue-lock section).
@@ -934,21 +998,8 @@ impl NativeVulkanDecoder {
if !vk.video_decode {
bail!("presenter device lacks Vulkan Video decode");
}
let lock: Box<dyn pf_vkdecode::QueueLock> =
if submit_queues_collide(vk.graphics_qf, vk.decode_qf) {
Box::new(NativeQueueLock::Shared(vk.queue_lock.clone()))
} else {
Box::new(NativeQueueLock::Uncontended)
};
let handles = DeviceHandles {
get_instance_proc_addr: vk.get_instance_proc_addr,
instance: vk.instance,
physical_device: vk.physical_device,
device: vk.device,
decode_qf: vk.decode_qf,
decode_queue_index: DECODE_QUEUE_INDEX,
graphics_qf: vk.graphics_qf,
};
let lock = queue_lock_for(vk);
let handles = device_handles(vk);
// The `DeviceHandles` caller contract, held for the decoder's whole lifetime
// and identical for both arms (it is the HANDLES' contract, not the codec's):
// the handles are the presenter's live instance/device, which outlives every
+50 -4
View File
@@ -246,17 +246,34 @@ const CELL_RAMP: [f64; 16] = [
-0.10, 0.08, -0.06, 0.12,
];
/// The twelve shipped palettes: the brand default, five more dark fields, then six pale ones.
/// The thirteen shipped palettes: the brand default, six more dark fields, then six pale ones.
/// Cycling order runs dark → light, so stepping the row walks the whole range in one direction.
/// Adding one here adds it to every console settings screen; the Apple and Android tables must
/// gain the same entry to keep the `ui_palette` key portable.
#[rustfmt::skip]
pub const PALETTES: [Palette; 12] = [
pub const PALETTES: [Palette; 13] = [
// --- dark fields (white ink) ---
Palette {
id: "violet", name: "Violet", stops: None,
ground: (0.075, 0.060, 0.160), accent: (0.525, 0.471, 0.961), light: false,
},
Palette {
// For OLED and AMOLED panels, where a black pixel is a pixel switched off — no glow,
// no power. The ramp's first two stops are literally (0,0,0), so the whole shaded half
// of the field is genuinely off rather than "very dark grey", and the ground is pure
// black too: the calm mix on the form screens lifts toward nothing, so settings and
// pairing sit on an unlit panel. What is left is a faint indigo→violet ember in the
// bright corner, dim enough to stay under a tenth of the other dark fields' mean
// luminance while keeping the backdrop a field with somewhere to go rather than a
// dead rectangle. The accent stays the brand violet — focus has to be findable on
// black.
id: "oled", name: "OLED",
stops: Some(&[
(0.000, 0.000, 0.000), (0.000, 0.000, 0.000), (0.010, 0.020, 0.100),
(0.045, 0.016, 0.115), (0.120, 0.024, 0.130),
]),
ground: (0.0, 0.0, 0.0), accent: (0.525, 0.471, 0.961), light: false,
},
Palette {
// Deep indigo climbing through violet into a hot magenta.
id: "nebula", name: "Nebula",
@@ -857,7 +874,7 @@ mod tests {
assert_eq!(
ids,
[
"violet", "nebula", "abyss", "ember", "moss", "graphite", "holo", "sunset",
"violet", "oled", "nebula", "abyss", "ember", "moss", "graphite", "holo", "sunset",
"bloom", "dawn", "mint", "opal",
]
);
@@ -867,7 +884,36 @@ mod tests {
.position(|p| p.light)
.expect("some are light");
assert!(PALETTES[first_light..].iter().all(|p| p.light));
assert_eq!(first_light, 6);
assert_eq!(first_light, 7);
}
/// OLED is the one palette whose selling point is measurable: it has to be genuinely
/// black, not merely the darkest of the dark fields. Pure black corners, a mean well
/// under every other field's, and a ground that lifts to nothing on the form screens.
#[test]
fn oled_is_actually_black() {
let luma = |c: (f64, f64, f64)| 0.2126 * c.0 + 0.7152 * c.1 + 0.0722 * c.2;
let oled = palette("oled");
assert_eq!(
oled.ground,
(0.0, 0.0, 0.0),
"the calm lift must be nothing"
);
let cells = oled.mesh_colors();
assert!(
cells.iter().filter(|c| luma(**c) == 0.0).count() >= 3,
"the shaded corner has to be switched off, not dimmed"
);
let mean = cells.iter().map(|c| luma(*c)).sum::<f64>() / 16.0;
let darkest_other = PALETTES
.iter()
.filter(|p| p.id != "oled")
.map(|p| p.mesh_colors().iter().map(|c| luma(*c)).sum::<f64>() / 16.0)
.fold(f64::MAX, f64::min);
assert!(
mean < darkest_other / 2.0,
"oled means {mean:.3}, only half a stop under {darkest_other:.3}"
);
}
/// Every colour a palette produces stays in gamut, and a pale palette really is pale —
+158 -44
View File
@@ -258,11 +258,17 @@ impl SettingsScreen {
}
}
/// The rows of the CURRENT tab. Profiles is built from the catalog: one row per
/// profile, or the explainer placeholder while there are none.
fn row_ids(&self) -> Vec<RowId> {
/// The rows of the CURRENT tab, minus any whose setting has nothing to act on (see
/// [`row_applies`]). Profiles is built from the catalog: one row per profile, or the
/// explainer placeholder while there are none.
fn row_ids(&self, ctx: &Ctx) -> Vec<RowId> {
if self.tab != PROFILES_TAB {
return TABS[self.tab].1.to_vec();
return TABS[self.tab]
.1
.iter()
.copied()
.filter(|id| row_applies(*id, ctx.settings))
.collect();
}
if self.profiles.is_empty() {
vec![RowId::NoProfiles]
@@ -271,6 +277,16 @@ impl SettingsScreen {
}
}
/// Pull the cursor back onto the list. Every tab but Profiles used to be a fixed length,
/// so this only mattered on entry ([`show_tab`]); the smoothness buffer's row now comes
/// and goes, and another writer (a desktop shell, a session's match-window persist) can
/// take it away between frames while this screen is open.
fn clamp_cursor(&mut self, len: usize) {
if self.list.cursor >= len {
self.list.jump_to(len.saturating_sub(1));
}
}
#[cfg(test)]
pub(crate) fn tab_for_test(&self) -> usize {
self.tab
@@ -278,21 +294,22 @@ impl SettingsScreen {
/// L1/R1 (and Tab/PgUp/PgDn) — move one tab, wrapping (the strip is a ring, like A's
/// value cycle), keeping each tab's own cursor.
fn switch_tab(&mut self, delta: i32) -> Option<MenuPulse> {
fn switch_tab(&mut self, delta: i32, ctx: &Ctx) -> Option<MenuPulse> {
let n = TABS.len() as i32;
self.show_tab((self.tab as i32 + delta).rem_euclid(n) as usize)
self.show_tab((self.tab as i32 + delta).rem_euclid(n) as usize, ctx)
}
/// Show `tab`, parking the cursor the outgoing tab was on. Also the pointer's path in:
/// a press on a pill names a tab outright rather than a direction to step in.
fn show_tab(&mut self, tab: usize) -> Option<MenuPulse> {
fn show_tab(&mut self, tab: usize, ctx: &Ctx) -> Option<MenuPulse> {
if tab >= TABS.len() {
return None;
}
self.tab_cursors[self.tab] = self.list.cursor;
self.tab = tab;
// Clamp the remembered cursor: the Profiles tab's length follows the catalog.
let len = self.row_ids().len();
// Clamp the remembered cursor: the Profiles tab's length follows the catalog, and
// Video's follows whether the smoothness buffer is offered.
let len = self.row_ids(ctx).len();
self.list
.jump_to(self.tab_cursors[self.tab].min(len.saturating_sub(1)));
Some(MenuPulse::Move)
@@ -302,10 +319,11 @@ impl SettingsScreen {
/// there is never meant for a row.
pub(crate) fn pointer(&mut self, p: Pointer, ctx: &mut Ctx, fx: &mut Outbox) -> bool {
if let Some(tab) = self.strip.pointer(p) {
self.show_tab(tab);
self.show_tab(tab, ctx);
return true;
}
let ids = self.row_ids();
let ids = self.row_ids(ctx);
self.clamp_cursor(ids.len());
let (msg, pulse) = self.list.pointer(p, ids.len());
if matches!(msg, ListMsg::None) && pulse.is_none() {
return false;
@@ -325,11 +343,12 @@ impl SettingsScreen {
fx.pop();
return None;
}
MenuEvent::JumpBack => return self.switch_tab(-1),
MenuEvent::JumpForward => return self.switch_tab(1),
MenuEvent::JumpBack => return self.switch_tab(-1, ctx),
MenuEvent::JumpForward => return self.switch_tab(1, ctx),
_ => {}
}
let ids = self.row_ids();
let ids = self.row_ids(ctx);
self.clamp_cursor(ids.len());
let (msg, pulse) = self.list.menu(ev, ids.len());
self.apply_row(msg, pulse, &ids, ctx, fx)
}
@@ -344,8 +363,14 @@ impl SettingsScreen {
ctx: &mut Ctx,
fx: &mut Outbox,
) -> Option<MenuPulse> {
// A cursor with no row under it can only mean the list shrank between the clamp above
// and here, which nothing does today — but indexing on the assumption would turn that
// into a panic in a shipping console rather than a dropped keypress.
let Some(&focused) = ids.get(self.list.cursor) else {
return pulse;
};
// The Profiles rows navigate instead of editing the settings file.
match ids[self.list.cursor] {
match focused {
RowId::Profile(i) => {
return match msg {
ListMsg::Activate => {
@@ -378,7 +403,7 @@ impl SettingsScreen {
}
match msg {
ListMsg::Adjust(delta) => {
let changed = adjust(ids[self.list.cursor], delta, false, ctx);
let changed = adjust(focused, delta, false, ctx);
if changed {
ctx.settings.save();
Some(MenuPulse::Move)
@@ -388,7 +413,7 @@ impl SettingsScreen {
}
ListMsg::Activate => {
// A cycles forward WRAPPING, so every option is reachable one-handed.
if adjust(ids[self.list.cursor], 1, true, ctx) {
if adjust(focused, 1, true, ctx) {
ctx.settings.save();
}
pulse
@@ -397,8 +422,8 @@ impl SettingsScreen {
}
}
pub(crate) fn hints(&self, _ctx: &Ctx) -> Vec<Hint> {
let ids = self.row_ids();
pub(crate) fn hints(&self, ctx: &Ctx) -> Vec<Hint> {
let ids = self.row_ids(ctx);
// The shoulders always change section, so that hint leads on every row.
let mut hints = vec![Hint::new(HintKey::Shoulders, "Section")];
hints.extend(match ids.get(self.list.cursor) {
@@ -445,7 +470,8 @@ impl SettingsScreen {
rect.right,
rect.bottom - detail_h as f32,
);
let ids = self.row_ids();
let ids = self.row_ids(ctx);
self.clamp_cursor(ids.len());
let rows: Vec<RowSpec> = ids
.iter()
.map(|id| row_spec(*id, ctx, &self.profiles))
@@ -466,6 +492,24 @@ impl SettingsScreen {
}
}
/// Whether a row is OFFERED at all, as opposed to offered-but-inert.
///
/// The two are a real distinction. Echo cancellation and the pad rows follow a switch the user
/// can see a line or two above them, so dimming them shows the relationship — dropping them
/// would just make settings appear and disappear as the switch flips. The smoothness buffer is
/// different: it is not a sub-setting of a switch, it is a knob on ONE of two intents, and
/// under Lowest latency it names a quantity that doesn't exist. Every other settings surface —
/// the GTK and WinUI shells, the Apple touch/tvOS screens, the Android touch screen — hides it
/// there. This screen was the lone exception because its row list was fixed; it is rebuilt from
/// this filter each frame now, and the row it drops sits directly BELOW the row that drops it,
/// so the cursor is never under anything that moves.
fn row_applies(id: RowId, s: &pf_client_core::trust::Settings) -> bool {
match id {
RowId::SmoothBuffer => s.present_priority == "smooth",
_ => true,
}
}
fn row_spec(id: RowId, ctx: &Ctx, profiles: &[(String, String)]) -> RowSpec {
// The Profiles section: name + how many hosts pin it (counted from the live rows, so
// it reflects what the carousel shows). Read-only here beyond opening the pin screen.
@@ -497,18 +541,17 @@ fn row_spec(id: RowId, ctx: &Ctx, profiles: &[(String, String)]) -> RowSpec {
_ => {}
}
let s = &ctx.settings;
// Several rows follow another: echo cancellation only means anything while the mic
// streams, the pad rows only while any controller is forwarded at all, and the
// smoothness buffer only while that intent is chosen. All go dim and inert otherwise
// — the same relationship the desktop shells draw by greying a row out (they hide the
// buffer row entirely; a fixed row list can't, and a row that vanished mid-list would
// move everything under the cursor).
// Two rows follow a switch a line or two above them: echo cancellation only means
// anything while the mic streams, and the pad rows only while any controller is
// forwarded at all. Both go dim and inert otherwise — the same relationship the desktop
// shells draw by greying a row out, and dimming (not dropping) is what shows the
// relationship. The smoothness buffer used to be listed here too; it is dropped from the
// list instead now — see [`row_applies`] for why that one is different.
let enabled = match id {
RowId::EchoCancel => s.mic_enabled,
RowId::Pad | RowId::PadType | RowId::SystemButtons | RowId::GuideGesture => {
s.gamepad_forwarding
}
RowId::SmoothBuffer => s.present_priority == "smooth",
_ => true,
};
let (header, label, value): (Option<&'static str>, &str, String) = match id {
@@ -848,7 +891,10 @@ fn adjust(id: RowId, delta: i32, wrap: bool, ctx: &mut Ctx) -> bool {
step_option(cur, PRESENT_PRIORITIES.len(), delta, wrap)
.map(|i| s.present_priority = PRESENT_PRIORITIES[i].0.to_string())
}
// Inert unless smoothness is chosen — a boundary thud, matching the dimmed row.
// Under Lowest latency the row isn't offered at all ([`row_applies`]), so this branch
// is only reachable if another writer flipped the intent between the frame that built
// the list and the keypress that lands here — a boundary thud, not a stored value
// nothing will read.
RowId::SmoothBuffer => {
if s.present_priority == "smooth" {
let cur = SMOOTH_BUFFERS
@@ -1093,9 +1139,6 @@ mod tests {
fake_home();
let mut s = SettingsScreen::with_profiles(Vec::new());
rendered(&mut s);
// Row 0 of the leading tab is Resolution, whose first step is Native → Match
// window: one field, one unambiguous effect to assert on.
assert_eq!(s.row_ids()[0], RowId::Resolution);
let first = s.list.row_rect(0).expect("the list drew its rows");
let (mut settings, pads) = ctx_parts();
settings.save(); // seat the fake HOME's file — `apply_row` rebases on it
@@ -1109,6 +1152,9 @@ mod tests {
device_name: "t",
t: 0.0,
};
// Row 0 of the leading tab is Resolution, whose first step is Native → Match
// window: one field, one unambiguous effect to assert on.
assert_eq!(s.row_ids(&ctx)[0], RowId::Resolution);
let mut fx = Outbox::default();
assert!(!ctx.settings.match_window);
assert!(s.pointer(press(first), &mut ctx, &mut fx));
@@ -1232,13 +1278,12 @@ mod tests {
assert!(ctx.settings.echo_cancel);
}
/// The smoothness buffer follows the presentation intent, exactly as echo cancellation
/// follows the mic: dimmed and inert under Lowest latency (where holding frames means
/// nothing), live under Smoothness. The desktop shells hide the row instead; a fixed
/// row list dims it, because a row vanishing mid-list would shift everything under the
/// cursor.
/// The smoothness buffer is OFFERED only under Smoothness — under Lowest latency it names
/// a quantity that doesn't exist, so the row is gone from the Video tab rather than sitting
/// there dimmed. This is what the GTK and WinUI shells and the Apple/Android screens have
/// always done; this screen was the exception until its row list stopped being fixed.
#[test]
fn smoothness_buffer_follows_the_intent() {
fn smoothness_buffer_is_offered_only_under_smoothness() {
let (mut settings, pads) = ctx_parts();
assert_eq!(settings.present_priority, "latency", "the shipped default");
let library = crate::library::LibraryShared::default();
@@ -1251,24 +1296,93 @@ mod tests {
device_name: "t",
t: 0.0,
};
assert!(!row_spec(RowId::SmoothBuffer, &ctx, &[]).enabled);
let mut s = SettingsScreen::with_profiles(Vec::new());
s.tab = TABS
.iter()
.position(|(name, _)| *name == "Video")
.expect("the Video tab");
let video = s.row_ids(&ctx);
assert!(
!video.contains(&RowId::SmoothBuffer),
"latency hides the buffer row: {video:?}"
);
assert!(video.contains(&RowId::PresentPriority), "the intent stays");
// Even reached out of band it writes nothing — the list it came from is a frame old.
assert!(
!adjust(RowId::SmoothBuffer, 1, false, &mut ctx),
"latency intent = thud"
);
assert_eq!(ctx.settings.smooth_buffer, 0, "and nothing was written");
// Stepping the intent to Smoothness brings the buffer row to life.
// Stepping the intent to Smoothness brings the row into the list, directly under it.
assert!(adjust(RowId::PresentPriority, 1, false, &mut ctx));
assert_eq!(ctx.settings.present_priority, "smooth");
assert!(row_spec(RowId::SmoothBuffer, &ctx, &[]).enabled);
let video = s.row_ids(&ctx);
let intent = video
.iter()
.position(|id| *id == RowId::PresentPriority)
.expect("the intent row");
assert_eq!(
video.get(intent + 1),
Some(&RowId::SmoothBuffer),
"the row that comes and goes sits BELOW the row that decides it, so the cursor \
never has anything move out from under it"
);
assert!(adjust(RowId::SmoothBuffer, 1, false, &mut ctx));
assert_eq!(ctx.settings.smooth_buffer, 1);
// The intent wraps back and the row goes inert again.
// The intent wraps back and the row leaves again — with the cursor parked on the
// intent row, which is where a user who just stepped it necessarily is.
s.list.cursor = intent;
assert!(adjust(RowId::PresentPriority, -1, false, &mut ctx));
assert_eq!(ctx.settings.present_priority, "latency");
assert!(!row_spec(RowId::SmoothBuffer, &ctx, &[]).enabled);
let video = s.row_ids(&ctx);
assert!(!video.contains(&RowId::SmoothBuffer));
assert_eq!(
video.get(s.list.cursor),
Some(&RowId::PresentPriority),
"the cursor is still on the row the user was stepping"
);
}
/// A cursor parked past the end of a list that shrank underneath it is pulled back rather
/// than indexed with — the console must not panic because another writer changed the
/// presentation intent while its settings screen was open.
#[test]
fn a_shrinking_list_pulls_the_cursor_back() {
// `apply_row` rebases on the FILE before acting, so this has to be seated — and
// seated with the SHRUNKEN list's intent, which is the state being tested.
fake_home();
let (mut settings, pads) = ctx_parts();
settings.present_priority = "latency".into();
settings.save();
settings.present_priority = "smooth".into();
let library = crate::library::LibraryShared::default();
let mut ctx = Ctx {
hosts: &[],
library: &library,
settings: &mut settings,
pads: &pads,
deck: false,
device_name: "t",
t: 0.0,
};
let mut s = SettingsScreen::with_profiles(Vec::new());
s.tab = TABS
.iter()
.position(|(name, _)| *name == "Video")
.expect("the Video tab");
// Park on the last row while the buffer row is still there…
s.list.cursor = s.row_ids(&ctx).len() - 1;
let parked = s.list.cursor;
// …then take it away behind the screen's back, as a desktop shell would.
ctx.settings.present_priority = "latency".into();
let mut fx = Outbox::default();
let pulse = s.menu(MenuEvent::Confirm, &mut ctx, &mut fx);
assert!(pulse.is_some(), "the press was routed, not dropped");
assert!(s.list.cursor < parked, "the cursor came back onto the list");
assert!(fx.nav.is_none());
}
#[test]
@@ -1392,7 +1506,7 @@ mod tests {
("p2".into(), "Game".into()),
]);
s.tab = PROFILES_TAB;
let ids = s.row_ids();
let ids = s.row_ids(&ctx);
assert_eq!(ids, vec![RowId::Profile(0), RowId::Profile(1)]);
let spec = row_spec(RowId::Profile(0), &ctx, &s.profiles);
@@ -1438,7 +1552,7 @@ mod tests {
};
let mut s = SettingsScreen::with_profiles(Vec::new());
s.tab = PROFILES_TAB;
let ids = s.row_ids();
let ids = s.row_ids(&ctx);
assert_eq!(ids, vec![RowId::NoProfiles]);
let spec = row_spec(RowId::NoProfiles, &ctx, &s.profiles);
assert!(!spec.enabled);
+1 -1
View File
@@ -329,7 +329,7 @@ fn dump_console_screens() {
for _ in 0..5 {
s.handle_menu(MenuEvent::JumpForward);
}
for id in ["violet", "ember", "abyss", "holo", "sunset", "mint"] {
for id in ["violet", "oled", "ember", "abyss", "holo", "sunset", "mint"] {
s.settings.ui_palette = id.to_string();
dump(&mut s, 40, 8, &format!("03-settings-{id}"), true);
}
File diff suppressed because it is too large Load Diff
+328
View File
@@ -53,6 +53,32 @@ pub(crate) fn stamp_color_bits(bitstream: &mut [u8], seq_offset: usize, bt2020_p
}
}
/// Read the 3-bit wire sequence counter out of a pyrowave block header.
///
/// Every block header is `{ u16 ballot; u16 payload_words:12, sequence:3, extended:1; u32 ... }`
/// (`pyrowave_common.hpp`, `static_assert(sizeof == 8)`), so the counter is bits 12..14 of the
/// little-endian half-word at `packet_offset + 2` — the same word `stamp_color_bits` reaches into
/// from the other end.
///
/// This field is the entire frame-boundary signal on the wire: the decoder restarts a frame only
/// when the value CHANGES (`diff = (hdr.sequence - last_seq) & 0x7; restart = diff != 0`), so a
/// repeated value is read as more blocks of the same frame. That is why PW5's alternating encoder
/// handles need `pyrowave_encoder_set_next_sequence`, and why a test asserts this reader sees
/// +1 mod 8 across the pair.
///
/// Its only caller is the Linux backend — alternating encoder handles are a Linux-side concern, and
/// the Windows backend drives pyrowave's compat device with a single handle. The rest of this module
/// really is shared (`packet_boundary` and `stamp_color_bits` have callers on both), so the exemption
/// is scoped to this one item rather than the file: `dead_code` stays live on Linux, where the caller
/// lives and where its disappearing would be a real finding. Windows builds with `-D warnings`, so
/// without this the host and tray clippy legs fail to compile the lib at all.
#[cfg_attr(not(target_os = "linux"), allow(dead_code))]
pub(crate) fn wire_sequence(bitstream: &[u8], packet_offset: usize) -> Option<u8> {
let lo = *bitstream.get(packet_offset + 2)?;
let hi = *bitstream.get(packet_offset + 3)?;
Some(((u16::from_le_bytes([lo, hi]) >> 12) & 0x7) as u8)
}
/// The wavelet block space's total 32x32-block count for a mode — the exact counting walk of
/// upstream `WaveletBuffers::init_block_meta` (also ported to the Apple `WaveletLayout`, whose
/// golden tests pin it against real host AUs). Needed because the vendored RDO pass packs the
@@ -201,6 +227,193 @@ pub(crate) fn build_au(
au
}
// ---------------------------------------------------------------------------
// Streamed-AU chunk cutting (PW6 — latency plan §T3.4, wave-2 plan PW6)
// ---------------------------------------------------------------------------
/// Default per-chunk target — ~34 chunks for a 400 Mb/s 60 fps AU (~833 KB). Deliberately coarse,
/// because the SEALER, not this size, sets how early bytes actually leave:
///
/// * Toward a plain `VIDEO_CAP_STREAMED_AU` client, `Packetizer::push_streamed` flushes only when
/// its pending buffer exceeds one FEC block — `fec.max_data_per_block × shard_payload`, which is
/// 200 × 1408 = 281 600 B on the shipped 1500-MTU IPv4 geometry. Anything smaller than that is
/// simply buffered. (256 KiB sits just under one block, so the first flush lands on the SECOND
/// chunk; the win is intact either way — the whole-AU path seals all ~3 blocks before its first
/// datagram may leave.) Only a client that ALSO negotiated `VIDEO_CAP_MULTI_SLICE` gets the
/// finer `MIN_STREAM_BLOCK_SHARDS` floor (16 shards ≈ 22 KB), where the chunk size does set the
/// flush granularity directly. pf-encode is not told the session's FEC geometry, so this is a
/// fixed byte target rather than a block-derived one.
/// * Chunks are not free: the send thread paces each sealed batch on its own
/// (`stream.rs::pace_sealed`), and every call grants a fresh `max(bytes/4, 128 KiB)` microburst
/// allowance. Cutting an AU into dozens of chunks therefore erodes the pacing this host does to
/// stop line-rate bursts from overrunning the NIC — the failure mode the pacer exists for.
const STREAM_CHUNK_TARGET_BYTES: usize = 256 * 1024;
/// Clamp on the `PUNKTFUNK_PYROWAVE_CHUNK_KIB` override (see [`stream_chunk_step`]).
const STREAM_CHUNK_MIN_KIB: usize = 4;
const STREAM_CHUNK_MAX_KIB: usize = 8192;
/// Whether streamed-AU output is armed for this host process.
///
/// **Default OFF, and deliberately so.** The streamed wire shape costs one PyroWave-specific
/// regression that has not been measured: an UNPINNED streamed frame (its final block never
/// arrived, so `frame_bytes` is still the 0 sentinel) is excluded from partial delivery
/// (`reassemble.rs`, 2026-07 security-review finding 10) — where today's whole-AU path hands the
/// consumer a usable blurred partial, a streamed frame that loses its final block delivers
/// NOTHING. PyroWave clients opt into partial delivery unconditionally
/// (`client/pump/handshake.rs`), so this is a live behaviour change for every one of them. The
/// netem loss-harness leg (2 % on `lo`, FEC pinned off — the Phase-4 recipe) comparing
/// partial-delivery rates streamed vs whole-AU is the prerequisite for flipping the default;
/// until it has run, `PUNKTFUNK_PYROWAVE_STREAMED_AU=1` is how you get it.
///
/// The client's `VIDEO_CAP_STREAMED_AU` and the host's `PUNKTFUNK_STREAMED_AU` remain the outer
/// gates (`stream.rs`) — this only decides whether the ENCODER offers chunks at all.
fn stream_armed() -> bool {
static ARMED: std::sync::OnceLock<bool> = std::sync::OnceLock::new();
// Latched once: `supports_chunked_poll` is re-queried per AU, and a knob that could change
// mid-session would flip the wire shape under an open `StreamedAu`.
*ARMED.get_or_init(|| {
matches!(
std::env::var("PUNKTFUNK_PYROWAVE_STREAMED_AU").as_deref(),
Ok("1")
)
})
}
/// Bytes per streamed chunk, rounded DOWN to a whole number of `window`-sized windows (never
/// below one). The rounding is the whole point — see [`AuChunker`].
fn chunk_step(window: usize, target: usize) -> usize {
(target / window.max(1)).max(1) * window.max(1)
}
/// The streamed-AU chunk size for a backend whose wire chunking is `wire_chunk`, or `None` when
/// this session must stay on the whole-AU path — which is the answer whenever the feature is not
/// armed ([`stream_armed`]) or the encoder is in DENSE mode.
///
/// Dense mode is excluded on purpose: there the AU is ONE atomic pyrowave packet with no window
/// framing, so a cut is neither shard-aligned nor a framing boundary. Every real PyroWave session
/// runs datagram-aligned (`stream.rs` sets `plan.wire_chunk = Some(session.shard_payload())`), so
/// nothing is lost — but the invariant this file promises stays true instead of nearly true.
///
/// `PUNKTFUNK_PYROWAVE_CHUNK_KIB` overrides the target (clamped to
/// [`STREAM_CHUNK_MIN_KIB`]..=[`STREAM_CHUNK_MAX_KIB`]); garbage falls back to the default.
pub(crate) fn stream_chunk_step(wire_chunk: Option<usize>) -> Option<usize> {
let window = wire_chunk.filter(|&w| w > 0)?;
if !stream_armed() {
return None;
}
static TARGET: std::sync::OnceLock<usize> = std::sync::OnceLock::new();
let target = *TARGET.get_or_init(|| {
std::env::var("PUNKTFUNK_PYROWAVE_CHUNK_KIB")
.ok()
.and_then(|v| v.trim().parse::<usize>().ok())
.filter(|k| (STREAM_CHUNK_MIN_KIB..=STREAM_CHUNK_MAX_KIB).contains(k))
.map(|k| k * 1024)
.unwrap_or(STREAM_CHUNK_TARGET_BYTES)
});
Some(chunk_step(window, target))
}
/// Hands a **finished** datagram-aligned AU out in window-aligned pieces for the streamed-AU wire
/// ([`crate::Encoder::poll_chunk`], `punktfunk_core::quic::VIDEO_CAP_STREAMED_AU`). Shared by both
/// pyrowave backends so the cut rule cannot drift between Linux and Windows — the Windows backend
/// cannot even be compiled from a Linux/macOS dev box, so logic written into it directly ships
/// unverified.
///
/// ## What this does NOT buy (read before quoting PW6 as a latency win)
///
/// pyrowave's `encode_frame` is **synchronous**: `submit` returns only once the whole AU sits in
/// `pending`, so by the time the host can poll a chunk the encode is over. `poll_chunk` is
/// therefore NOT "emit slices as the encoder produces them" — it is "hand the finished AU out in
/// pieces so the wire work pipelines with itself". Concretely, what moves:
///
/// * whole-AU path: `Session::seal_frame_at` FEC-protects, packetizes and AEAD-seals the ENTIRE
/// ~830 KB AU before its first datagram may leave the socket;
/// * streamed path: each FEC block seals and paces as it completes, so the first byte reaches the
/// wire after one block's seal, and the remaining seal work overlaps its own transmission.
///
/// There is NO encode/send overlap here — unlike the H.26x sub-frame slice path, where chunks
/// genuinely appear while the encoder is still working. PW6 and PW5 (encode overlap) are
/// independent packages, not sequential ones.
///
/// It also does **not** give the client decode-while-arriving: the reassembler completes a
/// streamed AU exactly like a whole one (`reassemble.rs` — `block_count != 0 && blocks_ok ==
/// block_count`) and hands up ONE `Frame`. Client-side prefix decode is the separate
/// `Session::set_deliver_frame_parts` opt-in, which PyroWave's newest-wins frame channel cannot
/// take — see the PW6 section of `design/linux-host-performance-wave2-pyrowave.md`.
///
/// ## The cut rule
///
/// A chunk is a whole number of `chunk`-sized WINDOWS. [`build_au`] gives every window exactly ONE
/// `kind` in its 4-byte prefix (`WIN_PACKED` or one link of a `WIN_FRAG_*` chain), so a cut inside
/// a window would split a unit the clients parse atomically. Whole windows are `shard_payload`
/// multiples by construction, which is what makes the sealer's sentinel block bases shard-aligned
/// for free (plan §4.4) — the streamed path's placement contract.
pub(crate) struct AuChunker {
au: Vec<u8>,
/// Bytes already handed out.
cursor: usize,
/// Bytes per chunk — a whole number of windows ([`chunk_step`]).
step: usize,
pts_ns: u64,
keyframe: bool,
recovery_anchor: bool,
chunk_aligned: bool,
/// Set once anything has been emitted, so the degenerate EMPTY AU still owes exactly one
/// chunk and not an infinite stream of them.
emitted: bool,
}
impl AuChunker {
pub(crate) fn new(frame: crate::EncodedFrame, step: usize) -> AuChunker {
AuChunker {
au: frame.data,
cursor: 0,
step: step.max(1),
pts_ns: frame.pts_ns,
keyframe: frame.keyframe,
recovery_anchor: frame.recovery_anchor,
chunk_aligned: frame.chunk_aligned,
emitted: false,
}
}
/// The next piece, or `None` once the AU is spent. The pieces concatenate to exactly the bytes
/// [`crate::Encoder::poll`] would have returned; `first` opens the wire frame and `last` closes
/// it (the host's `handle_chunk` keys its `begin`/`finish` off precisely those two).
pub(crate) fn next(&mut self) -> Option<crate::AuChunk> {
if self.cursor >= self.au.len() {
// A zero-byte AU is not reachable through `build_au` (it always emits at least one
// window), but the host would leak its open `StreamedAu` if a chunked poll returned
// nothing at all — so the degenerate case still owes one self-closing chunk.
if self.emitted {
return None;
}
self.emitted = true;
return Some(self.chunk(Vec::new(), true, true));
}
let first = self.cursor == 0;
let end = (self.cursor + self.step).min(self.au.len());
let data = self.au[self.cursor..end].to_vec();
self.cursor = end;
self.emitted = true;
Some(self.chunk(data, first, end == self.au.len()))
}
/// AU-level metadata rides every chunk (the `AuChunk` contract only makes it authoritative on
/// `first`, but a truthful copy on each one costs nothing and keeps a mid-AU log honest).
fn chunk(&self, data: Vec<u8>, first: bool, last: bool) -> crate::AuChunk {
crate::AuChunk {
data,
pts_ns: self.pts_ns,
keyframe: self.keyframe,
recovery_anchor: self.recovery_anchor,
chunk_aligned: self.chunk_aligned,
first,
last,
}
}
}
#[cfg(test)]
mod tests {
use super::*;
@@ -362,4 +575,119 @@ mod tests {
stamp_color_bits(&mut bs, 0, true);
assert_eq!(bs[7], 0x78);
}
// --- streamed-AU chunk cutting (PW6) ------------------------------------
// Appended at module END per the wave plan's ownership rule.
fn frame(data: Vec<u8>) -> crate::EncodedFrame {
crate::EncodedFrame {
data,
pts_ns: 1_234_567,
keyframe: true,
recovery_anchor: false,
chunk_aligned: true,
}
}
/// Drain a chunker into `(concatenated bytes, per-chunk lengths, first flags, last flags)`.
fn drain(mut c: AuChunker) -> (Vec<u8>, Vec<usize>, Vec<bool>, Vec<bool>) {
let (mut bytes, mut lens, mut firsts, mut lasts) = (Vec::new(), Vec::new(), vec![], vec![]);
while let Some(ch) = c.next() {
lens.push(ch.data.len());
firsts.push(ch.first);
lasts.push(ch.last);
bytes.extend_from_slice(&ch.data);
assert_eq!(ch.pts_ns, 1_234_567, "AU metadata rides every chunk");
assert!(ch.keyframe && ch.chunk_aligned && !ch.recovery_anchor);
}
(bytes, lens, firsts, lasts)
}
/// The invariant PW6 rests on: chunks concatenate to EXACTLY the AU, every cut lands on a
/// whole-window boundary (so no window's single `kind` is split across two wire frames), and
/// the reassembled stream still walks back to the same codec packets. A cut inside a window
/// would hand the client a 4-byte prefix whose body arrives in a different chunk — the
/// framing is one-kind-per-window, so there is no way to express that.
#[test]
fn stream_chunks_tile_the_au_on_window_boundaries() {
let bs: Vec<u8> = (0..4000u32).map(|i| (i % 251) as u8).collect();
let packets = [(0, 20), (20, 300), (320, 55), (375, 900), (1275, 40)];
let chunk = 64;
let au = build_au(&packets, &bs, Some(chunk));
assert!(au.len() / chunk > 4, "need several windows to cut between");
let step = chunk_step(chunk, 3 * chunk);
assert_eq!(step, 3 * chunk);
let (bytes, lens, firsts, lasts) = drain(AuChunker::new(frame(au.clone()), step));
assert_eq!(bytes, au, "chunks concatenate to exactly the AU");
assert!(
lens.iter().all(|l| l % chunk == 0),
"every chunk is a whole number of windows: {lens:?}"
);
assert!(
lens[..lens.len() - 1].iter().all(|&l| l == step),
"only the tail chunk may be short: {lens:?}"
);
assert_eq!(
firsts,
(0..lens.len()).map(|i| i == 0).collect::<Vec<_>>(),
"exactly one opening chunk"
);
assert_eq!(
lasts,
(0..lens.len())
.map(|i| i + 1 == lens.len())
.collect::<Vec<_>>(),
"exactly one closing chunk"
);
// And the client's parse is unchanged by the cutting.
let mut expect = Vec::new();
for &(o, s) in &packets {
expect.extend_from_slice(&bs[o..o + s]);
}
assert_eq!(walk(&bytes, chunk), expect);
}
/// The step always rounds DOWN to whole windows and never to zero — a target below one window
/// degenerates to one window per chunk rather than an empty chunk (which would spin forever).
#[test]
fn chunk_step_rounds_down_to_whole_windows() {
// 262144 / 1408 = 186.2 → 186 whole windows (261 888 B), never the 262 144 asked for.
assert_eq!(chunk_step(1408, 256 * 1024), 186 * 1408);
assert_eq!(chunk_step(1408, 1408), 1408);
assert_eq!(chunk_step(1408, 1407), 1408); // below one window → one window
assert_eq!(chunk_step(1408, 0), 1408);
assert_eq!(chunk_step(0, 4096), 4096); // defensive: never divides by zero
}
/// An AU that fits one chunk is a single `first && last` piece — the shape the host's
/// `handle_chunk` turns into begin+finish on one message, and byte-identical on the wire to
/// what the whole-AU path would have sealed.
#[test]
fn single_chunk_au_opens_and_closes_itself() {
let au = vec![7u8; 512];
let (bytes, lens, firsts, lasts) = drain(AuChunker::new(frame(au.clone()), 4096));
assert_eq!(bytes, au);
assert_eq!(lens, vec![512]);
assert_eq!(firsts, vec![true]);
assert_eq!(lasts, vec![true]);
}
/// The degenerate empty AU still owes exactly ONE self-closing chunk: a chunked poll that
/// returned nothing would leave the host's `StreamedAu` open forever (its `begin` fires on
/// `first`, its `finish` on `last`).
#[test]
fn empty_au_still_emits_one_self_closing_chunk() {
let mut c = AuChunker::new(frame(Vec::new()), 4096);
let ch = c.next().expect("one chunk");
assert!(ch.first && ch.last && ch.data.is_empty());
assert!(c.next().is_none(), "and never a second one");
}
/// Dense (non-windowed) AUs never stream: there is no window framing to cut on, so a chunk
/// boundary would be neither shard-aligned nor a parse boundary.
#[test]
fn dense_mode_never_streams() {
assert!(stream_chunk_step(None).is_none());
assert!(stream_chunk_step(Some(0)).is_none());
}
}
@@ -128,6 +128,11 @@ pub struct PyroWaveEncoder {
wire_budget: pyrowave_wire::WireBudget,
bitstream: Vec<u8>,
pending: VecDeque<EncodedFrame>,
/// The AU currently being handed out in streamed chunks (PW6 — `Some` strictly between a
/// `first` chunk and its `last`). See [`pyrowave_wire::AuChunker`]: this backend's encode is
/// synchronous, so the AU is COMPLETE before the first chunk leaves — the split is for the
/// send side, never an encode/send overlap.
chunker: Option<pyrowave_wire::AuChunker>,
}
// SAFETY: used only from the single encode thread; the pyrowave handles are owned and only touched
@@ -255,6 +260,7 @@ impl PyroWaveEncoder {
wire_budget: pyrowave_wire::WireBudget::new(),
bitstream: Vec::new(),
pending: VecDeque::new(),
chunker: None,
})
}
}
@@ -676,10 +682,55 @@ impl Encoder for PyroWaveEncoder {
}
fn poll(&mut self) -> Result<Option<EncodedFrame>> {
// Trait contract: each AU is drained through ONE method. Erroring beats double-emitting
// the bytes the chunk cursor already handed out (which would reach the wire twice, under
// the same frame index, and fail the receiver's retro-validation).
if self.chunker.is_some() {
bail!("pyrowave: poll() on an AU already being drained through poll_chunk");
}
Ok(self.pending.pop_front())
}
// --- streamed AU (PW6) — see `pyrowave_wire::AuChunker` for what this does and does NOT buy.
// Byte-identical to the Linux twin BY CONSTRUCTION: all of the cutting lives in the shared
// helper, which compiles and unit-tests on every platform. This file cannot be compiled from
// a Linux/macOS dev box, so anything written here directly would ship unverified.
fn supports_chunked_poll(&self) -> bool {
pyrowave_wire::stream_chunk_step(self.wire_chunk).is_some()
}
fn poll_chunk(&mut self) -> Result<Option<crate::AuChunk>> {
// Finish the AU already in flight before opening the next one — the host's `handle_chunk`
// keys begin/finish off `first`/`last` and cannot interleave two AUs.
if let Some(c) = self.chunker.as_mut() {
if let Some(chunk) = c.next() {
return Ok(Some(chunk));
}
self.chunker = None;
}
let Some(f) = self.pending.pop_front() else {
return Ok(None);
};
// No blocking wait here (the trait allows one): `submit` already ran the whole encode
// synchronously, so an AU in `pending` is complete by construction.
match pyrowave_wire::stream_chunk_step(self.wire_chunk) {
Some(step) => Ok(self
.chunker
.insert(pyrowave_wire::AuChunker::new(f, step))
.next()),
// Unarmed / dense: the trait's own default shape, so a host that polls chunks anyway
// still gets whole AUs.
None => Ok(Some(crate::AuChunk::whole(f))),
}
}
fn reset(&mut self) -> bool {
// A rebuild forfeits every in-flight frame — including an AU only half-handed-out through
// `poll_chunk`. Dropping the cursor here (ahead of every `pending.clear()` arm below) is
// what keeps the next `poll_chunk` from splicing the tail of a dead AU onto a fresh one;
// the host sees a `first` without the previous `last`, logs "streamed AU abandoned
// mid-flight" and lets the client age that frame out.
self.chunker = None;
// Cheap in-place rebuild: recreate only the pyrowave encoder object (no rate-control /
// reference state to preserve). The device, imported textures and fence survive.
// SAFETY: encode is synchronous (no work in flight); the device outlives the swapped encoder.
+50
View File
@@ -260,6 +260,18 @@ pub struct HostConfig {
/// encode, so this is the knob that decides how bright "white" looks on the client's panel.
/// `None` = leave gamescope's own default.
pub gamescope_sdr_nits: Option<u32>,
/// `PUNKTFUNK_GAMESCOPE_REFRESH_RATES` — extra refresh rates (Hz, comma-separated) a gamescope
/// session offers its clients on top of the one it runs at, e.g. `60,90,120`.
///
/// A headless gamescope has no EDID, so it cannot work out what else its display could run at:
/// on a stock build it advertises exactly ONE rate and Steam's in-session display settings show
/// a single entry. Our `+pfhdr3` build takes this list (`--custom-refresh-rates`) and publishes
/// it, which is what puts real choices in that menu. The session's own rate is always included
/// whatever is set here, so this can only ever ADD options.
///
/// Empty (the default) = advertise only the negotiated rate. Ignored on a stock gamescope,
/// which has no flag to take it.
pub gamescope_refresh_rates: Vec<u32>,
/// `PUNKTFUNK_RECOVER_SESSION_CMD` — operator hook fired (debounced) when a client connects while NO
/// graphical session is live for this uid: the state a compositor crash leaves behind (gnome-shell
/// SIGSEGV → GDM greeter, whose auto-login is once-per-boot, so the box would otherwise need a walk-up
@@ -379,6 +391,12 @@ impl HostConfig {
gamescope_sdr_nits: val("PUNKTFUNK_GAMESCOPE_SDR_NITS")
.and_then(|s| s.trim().parse::<u32>().ok())
.filter(|n| (1..=10_000).contains(n)),
// Unparseable entries are DROPPED rather than failing the host: this only ever widens a
// menu, and the session's own rate is added back unconditionally, so the worst a typo
// can cost is the extra option the operator wanted — never the session.
gamescope_refresh_rates: parse_refresh_rates(
val("PUNKTFUNK_GAMESCOPE_REFRESH_RATES").as_deref(),
),
recover_session_cmd: val("PUNKTFUNK_RECOVER_SESSION_CMD")
.filter(|s| !s.trim().is_empty()),
on_connect_cmd: val("PUNKTFUNK_ON_CONNECT_CMD").filter(|s| !s.trim().is_empty()),
@@ -397,6 +415,20 @@ impl HostConfig {
}
}
/// `"60, 90,120"` → `[60, 90, 120]`, sorted and deduped. Junk entries and out-of-range rates are
/// skipped rather than rejected wholesale — see the call site for why. Pure + unit-tested.
fn parse_refresh_rates(raw: Option<&str>) -> Vec<u32> {
let mut out: Vec<u32> = raw
.unwrap_or_default()
.split(',')
.filter_map(|s| s.trim().parse::<u32>().ok())
.filter(|&hz| (1..=1000).contains(&hz))
.collect();
out.sort_unstable();
out.dedup();
out
}
impl HostConfig {
/// The rate to hand the compositor as the GAME's refresh: the session's rate, capped by
/// [`Self::max_fps`]. Only the compositor's game-facing rate goes through here — the session's
@@ -446,6 +478,24 @@ mod tests {
assert_eq!(c.game_fps(0), 0);
}
#[test]
fn refresh_rate_list_parses_and_tolerates_junk() {
assert_eq!(parse_refresh_rates(Some("60,90,120")), vec![60, 90, 120]);
// Spaces, unsorted input and duplicates all normalise.
assert_eq!(
parse_refresh_rates(Some(" 120, 60 ,90, 60")),
vec![60, 90, 120]
);
// Unset and empty are the default: advertise only the session's own rate.
assert!(parse_refresh_rates(None).is_empty());
assert!(parse_refresh_rates(Some("")).is_empty());
assert!(parse_refresh_rates(Some(" ")).is_empty());
// A typo costs its own entry, never the whole list — the knob only widens a menu.
assert_eq!(parse_refresh_rates(Some("60,abc,120")), vec![60, 120]);
// Out of range in both directions (0 is not a refresh rate; 1920 is a width).
assert_eq!(parse_refresh_rates(Some("0,60,1920")), vec![60]);
}
#[test]
fn audio_output_mode_parses_its_spellings() {
for (s, want) in [
+124 -6
View File
@@ -550,7 +550,15 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
Some((x, y)) => b.position(x, y),
None => b.position_centered(),
};
b.resizable().vulkan();
// HIGH_PIXEL_DENSITY: give us a backbuffer in the panel's REAL pixels. Without it
// SDL leaves the Wayland surface at buffer scale 1, so on a fractionally scaled
// output (KDE at 150 %: a 2560×1600 panel reported as 1707×1067 points) the
// swapchain is built at 1707×1067 and the compositor upscales it to the glass —
// a 2560×1600 stream is resampled DOWN and back UP, and looks it. The flag only
// widens `size_in_pixels()`; `size()` stays logical, which is what the persisted
// window size and SDL's own mouse coordinates are in, and both callers already
// use the right one.
b.resizable().vulkan().high_pixel_density();
if opts.fullscreen {
b.fullscreen();
}
@@ -638,17 +646,29 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
// translation automatically — the GTK launcher never turned it off either).
gamepad.set_menu_mode(true);
}
// Gaming Mode's Steam menu / QAM drive the SAME physical pad we forward, and gamescope
// never takes our X focus away (it resolves focus per Xwayland ctx, and we are alone in
// ours), so SDL's own background-input gate cannot fire there. `None` everywhere else,
// where window focus IS the signal — see the FocusLost/FocusGained arms below.
#[cfg(target_os = "linux")]
let overlay_focus = pf_client_core::overlay_focus::OverlayFocus::start();
// Two independent reasons the pad is not ours — window focus and the gamescope overlay —
// OR'd into ONE value that is pushed to the service on an edge. Kept as separate inputs
// rather than one flag each source writes: either would otherwise clear the other's mask
// (a focus-loss mask undone by the next overlay poll saying "no overlay", and vice versa).
let mut focus_lost = false;
let mut mask_applied = false;
// The native display mode — the `0 = native` fallback for the requested stream mode
// (the GTK client reads the monitor under its window; same idea).
let native = window
.get_display()
.and_then(|d| d.get_mode())
.map(|m| Mode {
width: m.w.max(0) as u32,
height: m.h.max(0) as u32,
refresh_hz: m.refresh_rate.round().max(0.0) as u32,
})
.map(|m| native_mode(m.w, m.h, m.pixel_density, m.refresh_rate))
.ok()
// A zero-sized mode is as useless as no mode at all — only `Err` used to reach
// the fallback, so a display that reported 0×0 streamed a 0×0 request.
.filter(|m: &Mode| m.width > 0 && m.height > 0)
.unwrap_or(Mode {
width: 1920,
height: 1080,
@@ -750,8 +770,17 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
tracing::info!("focus lost — input released");
}
}
// Controllers go with the keyboard and mouse. SDL already stops
// delivering their PRESSES here, but nothing zeroed what the host
// still believes is held — so a stick deflected at the moment focus
// went away kept steering. Masking flushes it neutral.
focus_lost = true;
}
WindowEvent::FocusGained => {
// Unlike capture, the controller mask has no "the user meant it"
// variant to respect — it exists only to mirror who owns the pad —
// so regaining focus always lifts its half.
focus_lost = false;
// An auto-release (Alt-Tab) undoes itself; a chord release
// stays released until the user opts back in.
if let Some(cap) = stream.as_mut().and_then(|s| s.capture.as_mut()) {
@@ -1062,6 +1091,18 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
other => pump.handle_event(other),
}
}
// Who owns the pad right now: window focus, plus Gaming Mode's overlay signal where it
// exists (one relaxed atomic load; `None` off gamescope). Edge-triggered — the service
// hears only about CHANGES, so an open QAM doesn't re-flush the pads every iteration.
#[cfg(target_os = "linux")]
let overlay_now = overlay_focus.as_ref().is_some_and(|of| of.is_open());
#[cfg(not(target_os = "linux"))]
let overlay_now = false;
let want_mask = focus_lost || overlay_now;
if want_mask != mask_applied {
mask_applied = want_mask;
gamepad.set_masked(want_mask);
}
pump.tick();
// One coalesced MouseMove per iteration — pure motion must reach the host
// without waiting for a click/key to flush it.
@@ -2136,6 +2177,37 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
Ok(outcome)
}
/// An `SDL_DisplayMode` as the panel's REAL pixels — the `0 = native` stream mode.
///
/// SDL3 reports a display mode in SCREEN COORDINATES, not pixels, and hands you the ratio
/// between the two separately as `pixel_density`. On X11 and Windows that ratio is always
/// 1.0 (SDL never sets it there, and `SDL_video.c` normalizes the unset 0.0 up to 1.0), so
/// this is a no-op — but under a Wayland compositor doing FRACTIONAL scaling it is the
/// whole ballgame: KDE at 150 % advertises a 2560×1600 panel as 1707×1067 points with
/// `pixel_density` ≈ 1.4997, and taking `m.w`/`m.h` raw is what made "Native resolution"
/// negotiate 1706×1066 (1707×1067 even-floored by `render_scale::apply`) and stream a
/// blurry two-thirds-size image. `SDL_VIDEO_WAYLAND_SCALE_TO_DISPLAY=1` is the same fix
/// from the outside — it makes SDL report the native mode itself — which is why setting it
/// was a workaround.
///
/// The density is the exact `pixels / points` ratio SDL derived from the output, so the
/// multiplication recovers the panel size to the pixel rather than approximating it.
fn native_mode(w: i32, h: i32, pixel_density: f32, refresh_rate: f32) -> Mode {
// A non-finite or non-positive density is SDL telling us nothing useful; 1× at least
// preserves the pre-fix behaviour instead of collapsing the mode to zero.
let density = if pixel_density.is_finite() && pixel_density > 0.0 {
pixel_density
} else {
1.0
};
let px = |v: i32| (v.max(0) as f32 * density).round().max(0.0) as u32;
Mode {
width: px(w),
height: px(h),
refresh_hz: refresh_rate.round().max(0.0) as u32,
}
}
/// Match-window (D1): replace the params' requested w/h with the window's physical pixel
/// size — even-floored (the host's `validate_dimensions` rejects odd) and clamped to a
/// sane minimum — keeping the resolved refresh. Under `--fullscreen` the window IS the
@@ -2908,6 +2980,52 @@ fn stats_text(
mod tests {
use super::*;
/// The field report this exists for: CachyOS/KDE Plasma 6.7.4 Wayland, a 2560×1600@165
/// laptop panel at 150 % scaling. KDE advertises the output as 1707×1067 points, SDL
/// hands that back as the desktop mode with `pixel_density` = 2560/1707, and "Native
/// resolution" streamed 1706×1066 — the points, even-floored by `render_scale::apply`.
#[test]
fn native_is_the_panels_pixels_under_fractional_wayland_scaling() {
// SDL derives the density as the exact pixels-per-point ratio of the output.
let density = 2560.0 / 1707.0;
let m = native_mode(1707, 1067, density, 165.0);
assert_eq!((m.width, m.height, m.refresh_hz), (2560, 1600, 165));
// …and it survives the even-floor the host's `validate_dimensions` forces, which is
// where 1707×1067 lost its odd pixel and became the reported 1706×1066.
assert_eq!(
punktfunk_core::render_scale::apply(m.width, m.height, 1.0, 8192),
(2560, 1600)
);
assert_eq!(
punktfunk_core::render_scale::apply(1707, 1067, 1.0, 8192),
(1706, 1066),
"the pre-fix mode, kept here so the regression is legible"
);
}
#[test]
fn native_is_unchanged_where_the_density_is_one() {
// X11, Windows, and Wayland at 100 % all report 1.0 — the fix must be inert there.
let m = native_mode(2560, 1600, 1.0, 165.0);
assert_eq!((m.width, m.height, m.refresh_hz), (2560, 1600, 165));
// Integer scaling (a 200 % 4K panel reported as 1920×1080 points) doubles cleanly.
let m = native_mode(1920, 1080, 2.0, 60.0);
assert_eq!((m.width, m.height), (3840, 2160));
}
#[test]
fn a_nonsense_density_falls_back_to_one_rather_than_zeroing_the_mode() {
// SDL normalizes an unset density to 1.0, but this must not be the one place a
// driver quirk can hand the host a 0×0 mode request.
for bogus in [0.0, -1.0, f32::NAN, f32::INFINITY] {
let m = native_mode(2560, 1600, bogus, 60.0);
assert_eq!((m.width, m.height), (2560, 1600), "density {bogus}");
}
// A negative mode size is clamped, not wrapped into a huge u32.
let m = native_mode(-1, -1, 1.5, 60.0);
assert_eq!((m.width, m.height), (0, 0));
}
#[test]
fn overlay_scale_follows_dpi_and_survives_a_bogus_display() {
// 100 % / 96 dpi is the identity — the chrome keeps the size it always had.
+16
View File
@@ -130,6 +130,22 @@ impl Compositor {
}
}
/// Does this backend need a compositor that is ALREADY RUNNING for this uid?
///
/// Every desktop backend attaches to a live session — it asks Mutter/KWin/sway/Hyprland to mint
/// a virtual output over their IPC, so with nothing running there is no one to ask and `create`
/// can only fail (on GNOME: `RemoteDesktop.CreateSession:
/// org.freedesktop.DBus.Error.ServiceUnknown`). [`Compositor::Gamescope`] is the exception: it
/// stands its own session up from nothing (bare headless spawn / managed takeover), which is
/// exactly why a headless box pins to it.
///
/// Callers use this to tell "the session is up" from "the session is a corpse" BEFORE marching a
/// client into a doomed bring-up — the state a compositor crash leaves behind (gnome-shell
/// SIGSEGV → GDM greeter, whose auto-login is once-per-boot, so it never returns on its own).
pub fn needs_live_session(self) -> bool {
!matches!(self, Compositor::Gamescope)
}
/// Human label for UIs.
pub fn label(self) -> &'static str {
match self {
@@ -27,6 +27,7 @@ mod heads;
mod splash;
use discovery::{
check_gamescope_version, find_gamescope_eis_socket, find_gamescope_node, gamescope_bin,
gamescope_can_composite_external_overlay, gamescope_can_offer_refresh_rates,
gamescope_node_present, poll_managed_node, wait_for_node,
};
pub(crate) use discovery::{
@@ -327,7 +328,7 @@ impl VirtualDisplay for GamescopeDisplay {
// client's resolution (the box is headless, so its game-mode mode is ours to set).
// Reuse if it already matches (fast, no restart); otherwise relaunch the box's own
// session at the client mode. Without this the client gets the box's default mode.
ensure_box_gamescope_mode(mode)?
ensure_box_gamescope_mode(mode, self.hdr)?
} else {
id.parse()
.context("PUNKTFUNK_GAMESCOPE_NODE must be a node id or 'auto'")?
@@ -493,7 +494,7 @@ fn create_managed_session(client: &str, mode: Mode, hdr: bool) -> Result<Virtual
"gamescope: managed takeover unavailable — degrading to ATTACH (mirroring the box's \
own game-mode session)"
);
let node_id = ensure_box_gamescope_mode(mode)?;
let node_id = ensure_box_gamescope_mode(mode, hdr)?;
point_injector_at_eis();
return Ok(VirtualOutput {
node_id,
@@ -922,6 +923,68 @@ fn remove_steamos_dropin() {
let _ = std::fs::remove_file(steamos_dropin_path());
}
/// Drop-in for the box's OWN autologin `gamescope-session-plus@*.service`.
///
/// The transient-unit path ([`launch_session`]) can pass `BindReadOnlyPaths` straight to
/// `systemd-run`, but a box that owns an autologin session is RESTARTED in place instead — no
/// `systemd-run`, so the bind has to arrive as a drop-in or Nobara's hardcoded
/// `/usr/bin/gamescope` wins there too (see [`DISTRO_GAMESCOPE_PATH`]).
///
/// `zz-` so it sorts last, matching the SteamOS drop-in convention above.
fn session_plus_dropin_path() -> std::path::PathBuf {
let home = std::env::var("HOME").unwrap_or_else(|_| "/home/deck".to_string());
std::path::Path::new(&home)
.join(".config/systemd/user/gamescope-session-plus@.service.d/zz-punktfunk-bind.conf")
}
/// Write the box-session drop-in carrying the same two fixes the transient path gets: the bind, and
/// the WSI opt-out when the box's layer was built for a different gamescope. `PF_HZ`/`PF_HDR_ARGS`
/// ride along because the wrapper reads them (without `PF_HZ` it falls back to 60).
///
/// A no-op returning `Ok(false)` when there is nothing to redirect, so a box already running our
/// binary keeps a clean unit.
fn write_session_plus_dropin(
wrapper: &std::path::Path,
mode: Mode,
hdr: bool,
wsi_ok: bool,
) -> Result<bool> {
if gamescope_bin() == DISTRO_GAMESCOPE_PATH {
return Ok(false);
}
let path = session_plus_dropin_path();
if let Some(parent) = path.parent() {
std::fs::create_dir_all(parent).with_context(|| format!("mkdir {}", parent.display()))?;
}
let body = format!(
"[Service]\n\
BindReadOnlyPaths={wrapper}:{DISTRO_GAMESCOPE_PATH}\n\
Environment=PF_HZ={hz}\n\
Environment=\"PF_HDR_ARGS={hdr_args}\"\n\
{wsi}",
wrapper = wrapper.display(),
hz = game_hz(mode.refresh_hz),
hdr_args = hdr_args(hdr)
.into_iter()
.chain(cursor_args())
.collect::<Vec<_>>()
.join(" "),
wsi = if wsi_ok {
String::new()
} else {
"Environment=ENABLE_GAMESCOPE_WSI=0\n".to_string()
},
);
std::fs::write(&path, body).with_context(|| format!("write drop-in {}", path.display()))?;
Ok(true)
}
/// Remove the box-session drop-in (restore-on-disconnect). Best-effort, mirroring
/// [`remove_steamos_dropin`].
fn remove_session_plus_dropin() {
let _ = std::fs::remove_file(session_plus_dropin_path());
}
/// Take over SteamOS's `gamescope-session.target` headless at the CLIENT's mode: write the shim + a
/// drop-in carrying the mode, `daemon-reload`, then RESTART the target so `steam-launcher.service`
/// brings Steam up in the fresh headless gamescope — and attach to its node. A same-mode reconnect
@@ -1011,7 +1074,7 @@ fn create_managed_session_steamos(mode: Mode, hdr: bool) -> Result<VirtualOutput
/// box's own unit (rather than spawning a competing one) avoids the autologin-respawn fight the old
/// MANAGED path hit. A headless box has no physical panel, so its game-mode resolution is ours to set;
/// Steam restarts only on an actual resolution CHANGE.
fn ensure_box_gamescope_mode(mode: Mode) -> Result<u32> {
fn ensure_box_gamescope_mode(mode: Mode, hdr: bool) -> Result<u32> {
let target = (mode.width, mode.height);
// Fast path: already at the client's resolution — just attach to the live node.
if current_gamescope_output_size() == Some(target) {
@@ -1083,6 +1146,26 @@ fn ensure_box_gamescope_mode(mode: Mode) -> Result<u32> {
&format!("SCREEN_HEIGHT={}", mode.height),
&format!("CUSTOM_REFRESH_RATES={}", mode.refresh_hz.max(1)),
]);
// Same two fixes the transient path gets, but this unit is the BOX's own — they have to arrive
// as a drop-in, and `daemon-reload` before the restart or systemd runs the old unit.
match write_gamescope_bin_wrapper()
.and_then(|w| write_session_plus_dropin(&w, mode, hdr, wsi_layer_matches_our_gamescope()))
{
Ok(true) => {
tracing::info!(
bin = %gamescope_bin(),
%unit,
"gamescope: dropped in a bind over {DISTRO_GAMESCOPE_PATH} for the box's own \
session unit a session script that hardcodes that path (Nobara) gets the \
patched build on this restart too"
);
systemctl_user(&["daemon-reload"]);
}
Ok(false) => {}
// Best-effort: a box whose session already runs our binary loses nothing, and a failure
// here must not block a restart that would otherwise work.
Err(e) => tracing::warn!(error = %e, "gamescope: could not write the box-session drop-in"),
}
systemctl_user(&["restart", &unit]);
// Wait for the relaunched session to come up at the new size and publish its capture node. The
// node appears when gamescope is up (well before Steam finishes booting); the caller's
@@ -1153,17 +1236,9 @@ fn gamescope_argvs() -> Vec<Vec<String>> {
/// also the final filter that separates a compositor from anything else [`gamescope_argvs`] let by.
fn current_gamescope_output_size() -> Option<(u32, u32)> {
gamescope_argvs().into_iter().find_map(|args| {
let flag = |names: &[&str]| -> Option<u32> {
args.iter().enumerate().find_map(|(i, a)| {
names
.contains(&a.as_str())
.then(|| args.get(i + 1).and_then(|v| v.parse().ok()))
.flatten()
})
};
match (
flag(&["-W", "--output-width"]),
flag(&["-H", "--output-height"]),
argv_u32(&args, &["-W", "--output-width"]),
argv_u32(&args, &["-H", "--output-height"]),
) {
(Some(w), Some(h)) => Some((w, h)),
_ => None,
@@ -1171,6 +1246,104 @@ fn current_gamescope_output_size() -> Option<(u32, u32)> {
})
}
/// The numeric value following the first of `names` present in `argv`. Pure + unit-tested — it is
/// the shared reader behind both the output-size probe above and the mode verification below.
fn argv_u32(argv: &[String], names: &[&str]) -> Option<u32> {
argv.iter().enumerate().find_map(|(i, a)| {
names
.contains(&a.as_str())
.then(|| argv.get(i + 1).and_then(|v| v.parse().ok()))
.flatten()
})
}
/// Did the MODE we asked an indirectly-spawned session for actually reach its gamescope?
///
/// [`verify_managed_spawn_flags`] answers the same question for the capability flags and REFUSES
/// the session when they are missing, because the retry then resolves a different (correct) plan.
/// The mode has no such recovery: relaunching would hand the session the exact same environment and
/// lose it the same way, so refusing would only loop. It is not silent either, though — and it used
/// to be, in the way that costs the most:
///
/// `--nested-refresh` is the ONLY refresh a headless gamescope has. `CHeadlessBackend::Init`
/// assigns `g_nOutputRefresh = g_nNestedRefresh`, defaulting to **60 Hz** when the flag is absent,
/// and that one number is what the session composites at, what `vblankmanager` paces to, and what
/// Steam and every game are told the display runs at. It reaches a `gamescope-session-plus` only
/// through the `GAMESCOPE_BIN` wrapper — which the session script is free to lose (a `sessions.d`
/// file sourced with `set -a` can reassign `GAMESCOPE_BIN`; one that sets `GAMESCOPECMD` outright
/// skips the whole builder). When that happened the stream still ran, still looked right, and still
/// showed the client's own fps counter at the negotiated rate — because the encode loop repeats the
/// held frame — while the game underneath was capped to 60. Field report 2026-08-08.
///
/// So: warn, name the numbers, and carry on. Same "any running gamescope carrying it" rule as the
/// flag check, and the same silence when `/proc` cannot be read.
fn warn_if_mode_lost(mode: Mode, want_hz: u32) {
let argvs = gamescope_argvs();
let lost = mode_mismatch(mode.width, mode.height, want_hz, &argvs);
if lost.is_empty() {
return;
}
tracing::warn!(
lost = %lost.join(", "),
"gamescope: the session did not start at the mode we asked for — the session script \
dropped GAMESCOPE_BIN / SCREEN_WIDTH / SCREEN_HEIGHT. A headless gamescope reports \
`--nested-refresh` as its ONE refresh rate (60 Hz when the flag never arrives), so games \
and Steam will believe the display runs at that rate however fast the stream is. Install \
punktfunk-gamescope, or check /etc/gamescope-session-plus/sessions.d/ for a file that \
overrides GAMESCOPE_BIN or sets GAMESCOPECMD"
);
}
/// Which parts of the requested mode no running gamescope was started with, as human-readable
/// `asked=…, got=…` fragments. Empty when it matches — or when there is nothing to compare against,
/// which is the same fail-open rule [`missing_flags`] has and for the same reason. Pure +
/// unit-tested.
fn mode_mismatch(want_w: u32, want_h: u32, want_hz: u32, argvs: &[Vec<String>]) -> Vec<String> {
if argvs.is_empty() {
return Vec::new();
}
let mut lost = Vec::new();
let sizes: Vec<(u32, u32)> = argvs
.iter()
.filter_map(|a| {
Some((
argv_u32(a, &["-W", "--output-width"])?,
argv_u32(a, &["-H", "--output-height"])?,
))
})
.collect();
// No gamescope carries an output size at all → we cannot tell ours apart from a nested one;
// stay quiet rather than warn on every box that runs a second gamescope.
if !sizes.is_empty() && !sizes.contains(&(want_w, want_h)) {
lost.push(format!(
"resolution asked={want_w}x{want_h}, got={}",
sizes
.iter()
.map(|(w, h)| format!("{w}x{h}"))
.collect::<Vec<_>>()
.join("/")
));
}
let rates: Vec<u32> = argvs
.iter()
.filter_map(|a| argv_u32(a, &["-r", "--nested-refresh"]))
.collect();
if !rates.contains(&want_hz) {
lost.push(match rates.as_slice() {
// The flag is absent everywhere — the exact shape that silently yields 60 Hz.
[] => format!(
"refresh asked={want_hz}Hz, got=no --nested-refresh at all (gamescope defaults to \
60Hz headless)"
),
got => format!(
"refresh asked={want_hz}Hz, got={}Hz",
got.iter().map(u32::to_string).collect::<Vec<_>>().join("/")
),
});
}
lost
}
/// Did the flags we passed an INDIRECTLY-spawned session actually reach its gamescope?
///
/// The bare spawn builds argv itself and cannot lose them. The two managed modes can: a
@@ -2166,6 +2339,12 @@ fn do_restore_tv_session() {
}
return;
}
// Hand the box back its OWN gamescope before restarting its session: our bind drop-in exists
// to serve a punktfunk stream, and leaving it would silently put the patched build (plus our
// HDR/cursor flags) under the user's ordinary game mode — exactly the "sits beside the distro
// package" rule this whole design rests on.
remove_session_plus_dropin();
systemctl_user(&["daemon-reload"]);
for unit in units {
let _ = Command::new("systemctl")
.args(["--user", "start", &unit])
@@ -2322,6 +2501,60 @@ fn write_gamescope_bin_wrapper() -> Result<std::path::PathBuf> {
Ok(path)
}
/// The absolute path a session script may hardcode instead of honouring `GAMESCOPE_BIN`.
///
/// Nobara's `gamescope-session-plus` builds its command as `GAMESCOPECMD="/usr/bin/gamescope …"`
/// and reads `GAMESCOPE_BIN` NOWHERE, so all three of our spawn levers miss at once: the env var
/// is ignored, and an absolute path cannot be redirected by a PATH shim. The session then runs a
/// stock gamescope, the capability probe rejects it, and every session dies with
/// "pipeline build failed (out of retries)".
const DISTRO_GAMESCOPE_PATH: &str = "/usr/bin/gamescope";
/// Bind our wrapper over [`DISTRO_GAMESCOPE_PATH`] **inside the session unit's mount namespace**,
/// so a script that hardcodes that path still gets the patched build.
///
/// Deliberately a bind rather than replacing the distro's binary: `punktfunk-gamescope` ships under
/// its own name precisely so it sits BESIDE the distro package (a Steam gaming session keeps using
/// its own gamescope — see packaging/gamescope/README.md). The bind is scoped to this transient
/// unit, so nothing outside the session sees it and nothing is written to `/usr`.
///
/// Skipped when the resolved binary IS the distro path (nothing to redirect) — binding a file over
/// itself is pointless, and on a box with no `punktfunk-gamescope` we must not pretend otherwise.
fn session_gamescope_bind(wrapper: &std::path::Path) -> Option<String> {
if gamescope_bin() == DISTRO_GAMESCOPE_PATH {
return None;
}
Some(format!(
"--property=BindReadOnlyPaths={}:{DISTRO_GAMESCOPE_PATH}",
wrapper.display()
))
}
/// Whether the box's `VkLayer_FROG_gamescope_wsi` can be trusted against the gamescope we run.
///
/// The layer ships with the DISTRO's gamescope and speaks its `gamescope_swapchain` protocol; we
/// run our own build. When the two disagree the compositor rejects the client's
/// `swapchain_feedback` ("message too short") and **kills every Vulkan client** — Steam never
/// paints and the stream is a black screen with no error anywhere else.
///
/// Measured on Nobara 44 (`vkcube` under each build, layer on):
/// distro 3.16.23.2 → 0 errors; our 3.16.25 → 1 rejected client. The upstream protocol XML is
/// byte-identical between those commits, so this is the distro PATCHING gamescope, not a version
/// bump — which is why the check is "do the version triples differ", not a floor.
///
/// `ENABLE_GAMESCOPE_WSI=0` is gamescope's own opt-out and costs only the layer's extras
/// (present-mode control, client HDR metadata) — far cheaper than a client that cannot start.
fn wsi_layer_matches_our_gamescope() -> bool {
let ours = discovery::gamescope_version_of(std::path::Path::new(gamescope_bin()));
let distro = discovery::gamescope_version_of(std::path::Path::new(DISTRO_GAMESCOPE_PATH));
match (ours, distro) {
// Same upstream triple ⇒ the layer was built from the same protocol. Keep it.
(Some(a), Some(b)) => a == b,
// Either side unreadable: leave the layer alone rather than degrade a box that works.
_ => true,
}
}
/// Launch `gamescope-session-plus <client>` headless at `mode` as a transient `systemd --user`
/// unit (clean cgroup teardown of the whole Steam tree on stop). Injects `--nested-refresh` (via
/// the wrapper) + `--generate-drm-mode cvt` so games see exactly `mode` (resolution + refresh) and
@@ -2337,15 +2570,61 @@ fn launch_session(client: &str, unit_name: &str, mode: Mode, hdr: bool) -> Resul
let wrapper = write_gamescope_bin_wrapper()?;
stop_session(unit_name); // clear any stale unit + relay so a relaunch is clean
let hz = mode.refresh_hz.max(1);
// The two rates are deliberately different when the frame limiter is set. CUSTOM_REFRESH_RATES
// generates the mode the session ADVERTISES, which must stay the client's — that is what makes
// games see the real refresh instead of the box's EDID. PF_HZ becomes `--nested-refresh`, the
// rate the game is clamped to, and is the only one the limiter touches. Identical when it's
// unset, which is the default.
// ONE rate reaches gamescope, and it is `--nested-refresh` (via the wrapper's `PF_HZ`). On the
// headless backend that flag IS the output refresh — `CHeadlessBackend::Init` assigns
// `g_nOutputRefresh = g_nNestedRefresh` — so it is simultaneously the rate the session
// composites at, the rate `vblankmanager` paces to, and the rate Steam and every game are told
// the display runs at. When the frame limiter (`PUNKTFUNK_MAX_FPS`) is set they all drop
// together; that is the trade the knob is, and it is off by default.
//
// `CUSTOM_REFRESH_RATES` below does NOT do this, whatever its name suggests: it is the *set* of
// rates the session may offer, and `gamescope-session-plus` gates it on the binary having
// `--custom-refresh-rates`, which no upstream gamescope has ever had. On a stock gamescope it
// is inert (it was a silent no-op for years); on our `+pfhdr3` build it is what puts more than
// one entry in Steam's refresh menu. Either way it cannot fix a wrong `--nested-refresh`.
let game = game_hz(mode.refresh_hz);
// The advertised SET, which always contains the rate we actually run at.
let offered = {
let mut r = pf_host_config::config().gamescope_refresh_rates.clone();
if !r.contains(&hz) {
r.push(hz);
}
r.sort_unstable();
r.dedup();
r.iter().map(u32::to_string).collect::<Vec<_>>().join(",")
};
// Redirect a hardcoded `/usr/bin/gamescope` at our wrapper, for session scripts that never
// read `GAMESCOPE_BIN` (Nobara). Computed once so the log line below reflects what we did.
let bind = session_gamescope_bind(&wrapper);
if bind.is_some() {
tracing::info!(
bin = %gamescope_bin(),
"gamescope: binding the patched build over {DISTRO_GAMESCOPE_PATH} inside the session \
unit a session script that hardcodes that path (Nobara) gets the patched build \
instead of the distro's stock one. Nothing outside this unit is affected."
);
}
// The distro's Vulkan WSI layer speaks the distro gamescope's protocol; ours may differ, and a
// mismatch kills every Vulkan client (Steam included) with no error but a black screen.
let wsi_ok = wsi_layer_matches_our_gamescope();
if !wsi_ok {
tracing::warn!(
"gamescope: this box's VkLayer_FROG_gamescope_wsi was built for a different gamescope \
than the one we run disabling it for this session (ENABLE_GAMESCOPE_WSI=0). Left \
enabled it rejects the client's swapchain_feedback and every Vulkan client dies, \
which shows up as a black screen with no other symptom."
);
}
let start_unit = || -> Result<()> {
let status = Command::new("systemd-run")
.args(["--user", "--collect", &format!("--unit={unit_name}")])
let mut cmd = Command::new("systemd-run");
cmd.args(["--user", "--collect", &format!("--unit={unit_name}")]);
if let Some(b) = bind.as_deref() {
cmd.arg(b);
}
if !wsi_ok {
cmd.arg("--setenv=ENABLE_GAMESCOPE_WSI=0");
}
let status = cmd
// Same headless-must-not-attach rule as [`spawn`]: the transient unit inherits the
// user manager env, which can carry a (possibly stale) desktop DISPLAY/WAYLAND_DISPLAY
// that would abort gamescope at startup.
@@ -2366,7 +2645,7 @@ fn launch_session(client: &str, unit_name: &str, mode: Mode, hdr: bool) -> Resul
))
.arg(format!("--setenv=GAMESCOPE_BIN={}", wrapper.display()))
.arg("--setenv=DRM_MODE=cvt")
.arg(format!("--setenv=CUSTOM_REFRESH_RATES={hz}"))
.arg(format!("--setenv=CUSTOM_REFRESH_RATES={offered}"))
.arg("--")
.arg(SESSION_PLUS_BIN)
.arg(client)
@@ -2394,6 +2673,9 @@ fn launch_session(client: &str, unit_name: &str, mode: Mode, hdr: bool) -> Resul
stop_session(unit_name);
return Err(e);
}
// Loud, but not fatal — see [`warn_if_mode_lost`] for why this one warns where the
// capability flags above refuse.
warn_if_mode_lost(mode, game);
return Ok(id);
}
if Instant::now() >= deadline {
@@ -2526,7 +2808,15 @@ fn add_bare_gamescope_args(
if grab_cursor {
command.arg("--force-grab-cursor");
}
for arg in hdr_args(hdr).into_iter().chain(cursor_args()) {
// `-r` above is what this headless session will REPORT as its refresh (the headless backend
// assigns `g_nOutputRefresh = g_nNestedRefresh`), so it is already correct here. This adds the
// rest of the SET the in-session UI may offer — the bare spawn passes it directly, with none of
// the session-script indirection the managed path has to route it through.
for arg in hdr_args(hdr)
.into_iter()
.chain(cursor_args())
.chain(refresh_rate_args(hz))
{
command.arg(arg);
}
command.args(["--xwayland-count", "1", "--"]);
@@ -2571,11 +2861,48 @@ fn hdr_args(hdr: bool) -> Vec<String> {
/// host-side (it costs the host a full-frame pass, and on the zero-CSC encode source it cannot be
/// done at all). Empty on a stock gamescope, which is exactly the old behaviour.
fn cursor_args() -> Vec<String> {
let mut args = Vec::new();
if gamescope_can_composite_cursor() {
vec!["--pipewire-composite-cursor".to_string()]
} else {
Vec::new()
args.push("--pipewire-composite-cursor".to_string());
}
// The external overlay (mangoapp — the Deck UI's fps/frametime readout, patch level 4+). Unlike
// the cursor there is no host-side fallback: the host cannot reconstruct another process's
// overlay window, so without this the layer is simply absent from every gamescope stream.
if gamescope_can_composite_external_overlay() {
args.push("--pipewire-composite-external-overlay".to_string());
}
args
}
/// `--custom-refresh-rates <list>` when the resolved gamescope has it (patch level 3+): the rates a
/// HEADLESS session may offer its clients.
///
/// Without it a headless connector advertises exactly one rate, so Steam's in-session display
/// settings show a single entry and a game reads the display as that one number. `session_hz` is
/// always in the list — it is the mode the session actually runs at, and an advertised set that
/// excluded it would be a lie in the other direction.
///
/// The operator can widen the set (`PUNKTFUNK_GAMESCOPE_REFRESH_RATES=60,90,120`) so the in-session
/// UI offers real choices; unset, we advertise the one rate we run at, which is what the client
/// asked for.
fn refresh_rate_args(session_hz: u32) -> Vec<String> {
if !gamescope_can_offer_refresh_rates() {
return Vec::new();
}
let mut rates = pf_host_config::config().gamescope_refresh_rates.clone();
if !rates.contains(&session_hz) {
rates.push(session_hz);
}
rates.sort_unstable();
rates.dedup();
vec![
"--custom-refresh-rates".to_string(),
rates
.iter()
.map(u32::to_string)
.collect::<Vec<_>>()
.join(","),
]
}
/// Spawn `gamescope --backend headless -W w -H h -r hz -- <app>`. The app comes from
@@ -2717,7 +3044,7 @@ mod tests {
use super::{
cgroup_is_punktfunk_owned, cgroup_under_user_manager, connected_connector_under,
display_manager_unit_under, dm_plan, dm_survives_masked_unit, game_hz, hdr_args,
is_steam_launch, missing_flags, nested_wrapper_script, sentinel_advanced,
is_steam_launch, missing_flags, mode_mismatch, nested_wrapper_script, sentinel_advanced,
shape_dedicated_command,
};
@@ -2949,6 +3276,63 @@ mod tests {
assert!(!cgroup_is_punktfunk_owned(""));
}
/// The silent-60Hz guard. A headless gamescope reports `--nested-refresh` as its ONE refresh
/// rate and falls back to 60 Hz when the flag never arrives, so a session that lost the
/// `GAMESCOPE_BIN` wrapper streams at the client's rate while telling every game it is 60 —
/// the exact shape of the 2026-08-08 field report, and invisible without this.
#[test]
fn mode_mismatch_names_what_the_session_actually_got() {
let argv = |s: &str| -> Vec<String> { s.split(' ').map(str::to_string).collect() };
// The good case: our own managed spawn, carrying everything we asked for.
let ok = vec![argv(
"/usr/bin/gamescope --backend headless -W 1920 -H 1080 --nested-refresh 120 --steam",
)];
assert!(mode_mismatch(1920, 1080, 120, &ok).is_empty());
// THE field case: the wrapper was dropped, so there is no `--nested-refresh` anywhere and
// gamescope silently ran its 60 Hz default. Size still landed (SCREEN_WIDTH survived).
let lost = vec![argv(
"/usr/bin/gamescope --backend headless -W 1920 -H 1080 --steam",
)];
let got = mode_mismatch(1920, 1080, 120, &lost);
assert_eq!(got.len(), 1, "only the refresh is wrong: {got:?}");
assert!(got[0].contains("asked=120Hz"), "{got:?}");
assert!(got[0].contains("no --nested-refresh at all"), "{got:?}");
// A wrong rate is reported with the number it actually got, not just "missing".
let wrong = vec![argv("gamescope -W 1920 -H 1080 --nested-refresh 60")];
let got = mode_mismatch(1920, 1080, 120, &wrong);
assert_eq!(got.len(), 1);
assert!(got[0].contains("got=60Hz"), "{got:?}");
// Resolution lost too (SCREEN_WIDTH/HEIGHT dropped as well) — both are named.
let both = vec![argv("gamescope -W 1280 -H 720")];
assert_eq!(mode_mismatch(1920, 1080, 120, &both).len(), 2);
// Fail OPEN, exactly like `missing_flags`: nothing to compare against says nothing. A box
// with a second gamescope that carries no output size must not produce a false alarm.
assert!(mode_mismatch(1920, 1080, 120, &[]).is_empty());
// ANY running gamescope carrying the mode satisfies it — a Deck commonly runs a nested one
// beside the session, and demanding that every gamescope match would reject a good session.
let two = vec![
argv("gamescope -W 1280 -H 800 --nested-refresh 60"),
argv("gamescope -W 1920 -H 1080 --nested-refresh 120"),
];
assert!(mode_mismatch(1920, 1080, 120, &two).is_empty());
// The long spellings are read too.
let long = vec![argv(
"gamescope --output-width 1920 --output-height 1080 --nested-refresh 120",
)];
assert!(mode_mismatch(1920, 1080, 120, &long).is_empty());
// A flag with no value after it must not panic or read past the end.
let truncated = vec![argv("gamescope -W 1920 -H 1080 --nested-refresh")];
assert_eq!(mode_mismatch(1920, 1080, 120, &truncated).len(), 1);
}
/// The silent-cursor guard: a managed session that ignored `GAMESCOPE_BIN` / the PATH shim runs
/// a stock gamescope, and the host — already told the compositor would paint the pointer —
/// paints none either. Only a compositor we can SEE, missing a flag we can NAME, may fail.
@@ -449,6 +449,34 @@ pub(crate) fn gamescope_can_composite_cursor() -> bool {
gamescope_patch_level() >= 2 && !flags_lost()
}
/// Does the resolved gamescope let us hand a headless session the list of refresh rates it may
/// offer (`--custom-refresh-rates`)?
///
/// Below this level a headless gamescope advertises **one** rate — whatever `--nested-refresh`
/// resolved to, or its own 60 Hz default — and no resolution list at all, because its connector
/// returns empty spans from `GetModes()`/`GetValidDynamicRefreshRates()` and reports an INTERNAL
/// screen (which makes `update_mode_atoms` delete the mode-list atom outright). So on a stock
/// gamescope, Steam's in-session display settings show exactly one refresh rate and no
/// resolutions, and games read the display as 60 Hz whatever the client negotiated.
///
/// `gamescope-session-plus` has probed for this flag for years (`CUSTOM_REFRESH_RATES` is gated on
/// `gamescope --help` mentioning it) — upstream simply never had it, so the env var it plumbs was
/// a no-op everywhere.
pub(crate) fn gamescope_can_offer_refresh_rates() -> bool {
gamescope_patch_level() >= 3 && !flags_lost()
}
/// Can the resolved gamescope paint the EXTERNAL OVERLAY — mangoapp, the Deck-UI fps/frametime
/// readout — into its PipeWire node (`--pipewire-composite-external-overlay`)?
///
/// `paint_pipewire` has never referenced that layer on any upstream version, so a client whose
/// only view of the session is the node sees the overlay it just enabled simply not appear.
/// Unlike the cursor there is no host-side substitute: the host cannot reconstruct someone else's
/// overlay window.
pub(crate) fn gamescope_can_composite_external_overlay() -> bool {
gamescope_patch_level() >= 4 && !flags_lost()
}
/// Has a spawn been observed where our flags did NOT reach the gamescope process?
///
/// The binary probe above answers "can it", which is all the bare spawn needs — there we build
@@ -496,6 +524,22 @@ fn parse_patch_level(banner: &str) -> u32 {
.unwrap_or(0)
}
/// The upstream `X.Y.Z` a specific gamescope binary reports, or `None` if it cannot be run/parsed.
///
/// Split from [`check_gamescope_version`] (which only ever probes the RESOLVED binary) because the
/// WSI-layer check has to compare TWO binaries — ours and the distro's — and a `None` there means
/// "leave the layer alone", not "assume old".
pub(super) fn gamescope_version_of(bin: &std::path::Path) -> Option<(u32, u32, u32)> {
let out = Command::new(bin).arg("--version").output().ok()?;
// Same stdout/stderr split as the version gate: builds disagree on where the banner goes.
let text = format!(
"{}{}",
String::from_utf8_lossy(&out.stdout),
String::from_utf8_lossy(&out.stderr)
);
parse_version(&text)
}
/// Minimum gamescope that captures reliably: below 3.16.22, headless PipeWire capture deadlocks
/// against PipeWire ≥ 1.6 (a loop-lock bug) and a stuck link head-blocks the whole daemon.
const MIN_GAMESCOPE: (u32, u32, u32) = (3, 16, 22);
+114 -5
View File
@@ -13,8 +13,11 @@
//! So an interactive Plasma session does NOT hand it to a bare client — the host packages ship
//! `io.unom.Punktfunk.Host.desktop` (`Exec=/usr/bin/punktfunk-host`,
//! `X-KDE-Wayland-Interfaces=zkde_screencast_unstable_v1,…`) so it is present before the host first
//! connects. The headless test path instead exposes it to bare clients via
//! `KWIN_WAYLAND_NO_PERMISSION_CHECKS=1`. The compositor backend must implement
//! connects. That identification is also why **the host binary must carry no file capability**: a
//! process holding capabilities KWin lacks is one the kernel will not let KWin resolve
//! `/proc/<pid>/exe` for, so it can never be matched to a `.desktop` no matter how correctly the
//! file is installed (see [`capability_denial_hint`]). The headless test path instead exposes it to
//! bare clients via `KWIN_WAYLAND_NO_PERMISSION_CHECKS=1`. The compositor backend must implement
//! `createVirtualOutput`: the **DRM backend** (any version) or the **VirtualBackend since KWin
//! 6.5.6** (`kwin_wayland --virtual`); on `--virtual` < 6.5.6 the request fails with
//! "Could not find output". We talk raw Wayland on `$WAYLAND_DISPLAY`, so the host must run inside
@@ -1071,6 +1074,107 @@ impl Drop for StopOnDrop {
}
}
/// Extra sentence appended to every "KWin never advertised the screencast global" error when this
/// process carries capabilities — the one cause that is completely invisible from the Wayland side.
///
/// KWin authorizes a restricted interface by resolving the *client's* `/proc/<pid>/exe` and
/// matching it against an installed `.desktop`. The kernel refuses that readlink to any reader
/// whose effective set is not a superset of the target's **permitted** set
/// (`cap_ptrace_access_check`), and KWin has no capabilities at all. So a host binary carrying any
/// file capability is simply unidentifiable: `executablePath()` comes back empty, no `.desktop` can
/// match, and the global is never advertised — indistinguishable, from here, from a missing
/// `.desktop`. Neither half of the obvious workaround helps: `prctl(PR_SET_DUMPABLE, 1)` leaves the
/// permitted-set check failing, and moving the grant to systemd `AmbientCapabilities=` lands the
/// capability in the same permitted set. Only an uncapped binary is identifiable.
///
/// This is not hypothetical: 0.26.0-1 setcap'd `cap_sys_nice` on the host for the GPU-priority
/// lever and took out desktop streaming on every KDE box until the capability was removed again.
fn capability_denial_hint() -> String {
let permitted = std::fs::read_to_string("/proc/self/status")
.ok()
.and_then(|status| permitted_caps_from_status(&status));
capability_denial_hint_for(permitted)
}
/// The message half of [`capability_denial_hint`], split from the `/proc/self/status` read so it is
/// testable against a *given* mask instead of whatever the test process happens to hold.
///
/// That distinction is not academic: the first version of this asserted the empty case by calling
/// the real thing and trusting the test process to be uncapped. That holds on a dev box and is
/// false in CI, where the runner container is root with a full permitted set
/// (`CapPrm=0x000001ffffffffff`) — so the hint fired, correctly, and the test failed on a machine
/// where nothing was wrong. A check whose answer depends on the ambient environment tests the
/// environment, not the code.
fn capability_denial_hint_for(permitted: Option<u64>) -> String {
match permitted {
Some(caps) if caps != 0 => format!(
" — NOTE: this process carries capabilities (CapPrm={caps:#018x}), which is enough on \
its own to cause this: the kernel then refuses KWin the /proc/<pid>/exe read it \
identifies clients by, so no .desktop can match however correctly it is installed. \
Clear them with `sudo setcap -r /usr/bin/punktfunk-host` and restart the host"
),
_ => String::new(),
}
}
/// The permitted-capability mask out of a `/proc/<pid>/status` body, or `None` if the field is
/// absent/unparseable. The kernel prints it as a tab-separated 16-digit hex word with no `0x`
/// (`CapPrm:\t0000000000800000` = CAP_SYS_NICE), which is what the split-and-radix-16 parse below
/// expects — split out from [`capability_denial_hint`] purely so that shape is testable without a
/// capability-carrying process to point at.
fn permitted_caps_from_status(status: &str) -> Option<u64> {
let field = status.lines().find(|l| l.starts_with("CapPrm:"))?;
u64::from_str_radix(field.split_whitespace().nth(1)?, 16).ok()
}
#[cfg(test)]
mod capability_hint_tests {
use super::*;
/// Verbatim from a `cap_sys_nice=ep` process on CachyOS — the case that broke 0.26.0-1.
const CAPPED: &str = "Name:\tpunktfunk-host\nUid:\t1000\t1000\t1000\t1000\nCapPrm:\t0000000000800000\nCapEff:\t0000000000800000\n";
/// ...and from the same binary with no capability, where the hint must stay silent.
const CLEAN: &str = "Name:\tpunktfunk-host\nUid:\t1000\t1000\t1000\t1000\nCapPrm:\t0000000000000000\nCapEff:\t0000000000000000\n";
#[test]
fn parses_the_kernels_permitted_mask() {
assert_eq!(permitted_caps_from_status(CAPPED), Some(0x0080_0000));
assert_eq!(permitted_caps_from_status(CLEAN), Some(0));
// CapPrm is not guaranteed present (older/again-different kernels): stay quiet, never panic.
assert_eq!(permitted_caps_from_status("Name:\tx\n"), None);
assert_eq!(permitted_caps_from_status("CapPrm:\tzzzz\n"), None);
assert_eq!(permitted_caps_from_status("CapPrm:\n"), None);
}
/// A capability-free host must not append the hint — the message it decorates is also printed
/// on genuinely missing `.desktop` files, and a spurious "you have capabilities" line would
/// send the reader chasing a setcap that was never there.
///
/// Driven off an explicit mask rather than the test process's own: see
/// [`capability_denial_hint_for`] for why calling the real reader here fails in CI.
#[test]
fn silent_without_capabilities() {
assert_eq!(
capability_denial_hint_for(permitted_caps_from_status(CLEAN)),
""
);
// Absent or unparseable field: also silent, never a panic and never a spurious hint.
assert_eq!(capability_denial_hint_for(None), "");
}
/// ...and the case that matters actually speaks, naming the mask and the repair. Without this
/// the test above passes just as well against a function that returns `""` unconditionally.
#[test]
fn names_the_mask_and_the_repair_when_capped() {
let hint = capability_denial_hint_for(permitted_caps_from_status(CAPPED));
assert!(
hint.contains("0x0000000000800000"),
"names the mask: {hint}"
);
assert!(hint.contains("setcap -r"), "names the repair: {hint}");
}
}
/// Readiness probe: connect to the KWin Wayland socket, roundtrip the registry, and confirm
/// the privileged `zkde_screencast` global is actually advertised. This is exactly what
/// [`run`] needs before it can create a virtual output, so a session-bringup script can poll
@@ -1090,7 +1194,8 @@ pub fn probe() -> Result<()> {
it on the host's .desktop X-KDE-Wayland-Interfaces (install \
io.unom.Punktfunk.Host.desktop with Exec=/usr/bin/punktfunk-host, then re-login so KWin \
re-reads it the grant is cached per-exe on first connect), or set \
KWIN_WAYLAND_NO_PERMISSION_CHECKS=1 for the headless test; needs KWin 6.5.6"
KWIN_WAYLAND_NO_PERMISSION_CHECKS=1 for the headless test; needs KWin 6.5.6{}",
capability_denial_hint()
);
}
Ok(())
@@ -1134,7 +1239,9 @@ fn run_existing(
anyhow!(
"KWin does not expose zkde_screencast_unstable_v1 to this client — install the host's \
.desktop (io.unom.Punktfunk.Host.desktop, X-KDE-Wayland-Interfaces) and re-login so \
KWin authorizes it, or run KWin with KWIN_WAYLAND_NO_PERMISSION_CHECKS=1 (headless test)"
KWin authorizes it, or run KWin with KWIN_WAYLAND_NO_PERMISSION_CHECKS=1 (headless \
test){}",
capability_denial_hint()
)
})?;
@@ -1223,7 +1330,9 @@ fn run(
anyhow!(
"KWin does not expose zkde_screencast_unstable_v1 to this client — install the host's \
.desktop (io.unom.Punktfunk.Host.desktop, X-KDE-Wayland-Interfaces) and re-login so \
KWin authorizes it, or run KWin with KWIN_WAYLAND_NO_PERMISSION_CHECKS=1 (headless test)"
KWin authorizes it, or run KWin with KWIN_WAYLAND_NO_PERMISSION_CHECKS=1 (headless \
test){}",
capability_denial_hint()
)
})?;
@@ -299,16 +299,21 @@ struct Pinger {
/// The manager's control-device cache. Reopenable: a driver upgrade / WUDFHost restart kills the
/// cached handle (every IOCTL fails with a gone-class code forever), so such a failure RETIRES it and
/// the next [`VirtualDisplayManager::ensure_device`] reopens the (new) device interface, re-running
/// the version handshake. Retired handles are deliberately kept alive — never closed — for the
/// process lifetime: the pinger/linger threads and every capturer's `ChannelBroker` hold BARE
/// `HANDLE` copies whose soundness contract is "never closed"; a retired handle only ever FAILS
/// IOCTLs, which every holder already tolerates. Reopens are rare (a driver restart), so the retained
/// list is bounded in practice.
/// the version handshake.
///
/// Ownership is `Arc` all the way out: every consumer — `acquire`'s IOCTL runs, the pinger/linger
/// threads, the capture layer's delivery closures — holds its OWN clone across its use, so retiring
/// here merely drops the manager's reference and the handle CLOSES when the last in-flight user
/// drains. That close is load-bearing, not housekeeping: an open control handle is exactly what
/// vetoes the PnP disable/restart the wake-from-sleep recovery leans on (field 2026-08-08 — every
/// reload REFUSED `Generic failure`; `reset-pf-vdisplay.ps1` stops the whole host service precisely
/// to get its handles closed, and Arc ownership buys the same release without dying). The previous
/// contract kept retired handles open for the process lifetime because bare `HANDLE` copies were
/// smuggled into threads and closures; those copies are gone, and nothing may rely on a dead
/// handle staying open again.
#[derive(Default)]
struct DeviceSlot {
current: Option<Arc<OwnedHandle>>,
/// Never dropped — see the type doc (bare-`HANDLE` holders rely on no-close).
retired: Vec<Arc<OwnedHandle>>,
/// `CLEAR_ALL` (crashed-host orphan reap) runs only on the FIRST open of the process; a reopen
/// races sessions this process still considers live and must not raze them.
opened_once: bool,
@@ -397,11 +402,6 @@ pub fn vdm() -> &'static VirtualDisplayManager {
.expect("VirtualDisplayManager used before a backend initialised it")
}
/// The live pf-vdisplay control-device handle, for the IDD-push capturer's sealed-channel delivery
/// (`IOCTL_SET_FRAME_CHANNEL`). Safe to hand out as a bare `HANDLE`: cached handles are never closed
/// for the process lifetime — a dead one is RETIRED (kept alive, see [`DeviceSlot`]), so a stale copy
/// can only fail IOCTLs, never dangle. `None` before the first backend open — impossible for a
/// capturer, which only exists on a monitor the manager created.
/// Can this host's pf-vdisplay driver run the v5 hardware-cursor channel? Reads the
/// handshake-latched protocol version, opening the control device once if no session has
/// opened it yet this service run (the same open every session performs anyway) — so the
@@ -421,7 +421,13 @@ pub fn hw_cursor_capable() -> bool {
m.driver_proto.load(Ordering::Relaxed) >= 5
}
pub fn control_device_handle() -> Option<HANDLE> {
/// The live pf-vdisplay control device, for the IDD-push capturer's sealed-channel delivery
/// (`IOCTL_SET_FRAME_CHANNEL`) — an `Arc` clone the caller (and every closure it builds) holds for
/// as long as it may issue IOCTLs: the handle stays open while any holder lives and closes when the
/// last drains, which is what lets the wake-from-sleep recovery's PnP disable proceed once the
/// manager retires it (see [`DeviceSlot`]). `None` before the first backend open — impossible for a
/// capturer, which only exists on a monitor the manager created.
pub fn control_device_handle() -> Option<Arc<OwnedHandle>> {
VDM.get().and_then(VirtualDisplayManager::device_handle)
}
@@ -497,17 +503,28 @@ fn is_device_gone(e: &anyhow::Error) -> bool {
GONE.contains(&w.code().0)
}
/// The transient raw `HANDLE` view of an Arc-held control device, for the backend IOCTL surface.
/// Sound only while the `Arc` it borrows from is held — which the borrow makes structural: every
/// use site necessarily has the owning clone alive across the call, so a concurrent retire (which
/// now really closes the handle once its users drain — see [`DeviceSlot`]) can never close it
/// mid-IOCTL.
fn dev_raw(dev: &OwnedHandle) -> HANDLE {
HANDLE(dev.as_raw_handle())
}
impl VirtualDisplayManager {
pub(crate) fn backend_name(&self) -> &'static str {
self.driver.name()
}
/// Open + cache the control device; REOPEN when a gone-classified failure retired the cached one
/// (driver upgrade / WUDFHost restart). The `device` mutex serializes racing opens.
fn ensure_device(&self) -> Result<HANDLE> {
/// (driver upgrade / WUDFHost restart). The `device` mutex serializes racing opens. Returns an
/// `Arc` clone the caller holds across every IOCTL it derives from it — a concurrent retire then
/// drops only the manager's reference and closes nothing under the caller (see [`DeviceSlot`]).
fn ensure_device(&self) -> Result<Arc<OwnedHandle>> {
let mut slot = self.device.lock().unwrap();
if let Some(d) = &slot.current {
return Ok(HANDLE(d.as_raw_handle()));
return Ok(d.clone());
}
let reap = !slot.opened_once;
claim_instance()?;
@@ -519,35 +536,33 @@ impl VirtualDisplayManager {
slot.opened_once = true;
self.watchdog_s.store(watchdog_s, Ordering::Relaxed);
self.driver_proto.store(driver_proto, Ordering::Relaxed);
let raw = HANDLE(handle.as_raw_handle());
slot.current = Some(Arc::new(handle));
let dev = Arc::new(handle);
slot.current = Some(dev.clone());
if !reap {
tracing::info!("virtual-display control device reopened (retired handle replaced)");
}
Ok(raw)
Ok(dev)
}
/// The live control handle for the pinger/linger threads. `None` before the first acquire opened
/// it, or between a retire and the next reopen.
fn device_handle(&self) -> Option<HANDLE> {
self.device
.lock()
.unwrap()
.current
.as_ref()
.map(|d| HANDLE(d.as_raw_handle()))
/// The live control device for the pinger/linger threads — an `Arc` clone the caller holds
/// across its IOCTLs. `None` before the first acquire opened it, or between a retire and the
/// next reopen.
fn device_handle(&self) -> Option<Arc<OwnedHandle>> {
self.device.lock().unwrap().current.clone()
}
/// Retire the cached control handle after a gone-classified IOCTL failure. The handle is retained
/// un-closed (see [`DeviceSlot`]); the next [`ensure_device`](Self::ensure_device) reopens the
/// (new) device interface and re-runs the version handshake.
/// Retire the cached control handle after a gone-classified IOCTL failure: drop the manager's
/// reference, so the handle CLOSES once the last in-flight user drains (see [`DeviceSlot`]) —
/// the release the wake-from-sleep recovery needs before it can cycle the adapter devnode. The
/// next [`ensure_device`](Self::ensure_device) reopens the (new) device interface and re-runs
/// the version handshake.
fn invalidate_device(&self, why: &anyhow::Error) {
let mut slot = self.device.lock().unwrap();
if let Some(cur) = slot.current.take() {
if slot.current.take().is_some() {
tracing::warn!(
"virtual-display control device retired — reopening on next use (cause: {why:#})"
"virtual-display control device retired — closes when its last user drains, \
reopening on next use (cause: {why:#})"
);
slot.retired.push(cur);
}
}
@@ -620,11 +635,11 @@ impl VirtualDisplayManager {
old_target,
"IDD-push reconnect — preempting the kept (lingering/pinned) monitor, recreating a fresh one"
);
// SAFETY: `teardown_removed` requires `dev` to be a valid control handle; `dev` is the
// value `ensure_device()` returned above (cached handles are never closed — a dead one
// is retired, kept alive; see `DeviceSlot`). `mon` was just removed from the map, so it
// SAFETY: `teardown_removed` requires `dev` to be a valid control handle; the `dev`
// Arc `ensure_device()` returned above is held across this call, so the handle stays
// open even against a concurrent retire. `mon` was just removed from the map, so it
// is exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev, &mut inner, mon) };
unsafe { self.teardown_removed(dev_raw(&dev), &mut inner, mon) };
// Let the OS finish the ASYNC monitor departure before the next ADD; a back-to-back
// REMOVE→ADD races the teardown and the ADD IOCTL is rejected under reconnect churn.
// Verified-state wait, ceiling = the old fixed 400 ms settle (latency plan P0.3).
@@ -657,11 +672,11 @@ impl VirtualDisplayManager {
wudf_pid = mon.wudf_pid,
"virtual monitor's WUDFHost is gone — preempting the dead monitor, recreating"
);
// SAFETY: `teardown_removed` requires a valid control handle; `dev` is the value
// `ensure_device()` returned above (cached handles are never closed — a dead one is
// retired, kept alive; see `DeviceSlot`). `mon` was just removed from the map, so it
// SAFETY: `teardown_removed` requires a valid control handle; the `dev` Arc
// `ensure_device()` returned above is held across this call, so the handle stays
// open even against a concurrent retire. `mon` was just removed from the map, so it
// is exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev, &mut inner, mon) };
unsafe { self.teardown_removed(dev_raw(&dev), &mut inner, mon) };
// Same async-departure settle as the reconnect preempt above (verified wait, P0.3).
let _ = wait_target_departed(old_target, Duration::from_millis(400));
}
@@ -693,9 +708,10 @@ impl VirtualDisplayManager {
else {
unreachable!("just matched Active");
};
// SAFETY: `dev` is the handle `ensure_device()` returned above; the CCD
// waits inside run under the held `state` lock (this fn's discipline).
match unsafe { self.resize_in_place(dev, mon, mode) } {
// SAFETY: the `dev` Arc `ensure_device()` returned above is held across
// this call (so the handle stays open); the CCD waits inside run under
// the held `state` lock (this fn's discipline).
match unsafe { self.resize_in_place(dev_raw(&dev), mon, mode) } {
Ok(()) => {
// Same join semantics as the re-arrival: +1 ref for the new
// (build-then-drop overlap) lease; `gen` untouched, so the old
@@ -734,10 +750,11 @@ impl VirtualDisplayManager {
let Some(SlotState::Active { mon, refs }) = inner.slots.remove(&slot) else {
unreachable!("just matched Active");
};
// SAFETY: `dev` is the handle `ensure_device()` returned above; `re_add` touches the
// live topology under the held `state` lock. `mon` is owned here (removed from the map).
// SAFETY: the `dev` Arc `ensure_device()` returned above is held across this call
// (so the handle stays open); `re_add` touches the live topology under the held
// `state` lock. `mon` is owned here (removed from the map).
let new_mon = match unsafe {
self.re_add(dev, &mut inner, slot, &mon, mode, client_hdr)
self.re_add(dev_raw(&dev), &mut inner, slot, &mon, mode, client_hdr)
} {
ReAdd::Arrived(m) => *m,
ReAdd::RolledBack {
@@ -815,11 +832,11 @@ impl VirtualDisplayManager {
}
// The slot is empty: create a fresh monitor for it.
// SAFETY: `create_monitor` requires `dev` to be a valid control handle; `dev` is the handle
// `ensure_device()` returned above (cached handles are never closed — a dead one is retired,
// kept alive; see `DeviceSlot`), and we hold the `state` lock.
// SAFETY: `create_monitor` requires `dev` to be a valid control handle; the `dev` Arc
// `ensure_device()` returned above is held across this call (so the handle stays open even
// against a concurrent retire), and we hold the `state` lock.
let mon = match unsafe {
self.create_monitor(dev, mode, slot, client_hdr, hw_cursor, &mut inner)
self.create_monitor(dev_raw(&dev), mode, slot, client_hdr, hw_cursor, &mut inner)
} {
// The cached device died under us (driver upgrade / WUDFHost restart, detected only
// now — e.g. the host sat idle past the pinger-less window). Retire it, reopen, and
@@ -831,9 +848,18 @@ impl VirtualDisplayManager {
tracing::info!(
"virtual-display control device reopened — retrying the monitor create"
);
// SAFETY: as above — `dev` is the handle the reopening `ensure_device` just
// returned, and the `state` lock is still held.
unsafe { self.create_monitor(dev, mode, slot, client_hdr, hw_cursor, &mut inner)? }
// SAFETY: as above — the `dev` Arc the reopening `ensure_device` just returned is
// held across this call, and the `state` lock is still held.
unsafe {
self.create_monitor(
dev_raw(&dev),
mode,
slot,
client_hdr,
hw_cursor,
&mut inner,
)?
}
}
r => r?,
};
@@ -887,13 +913,12 @@ impl VirtualDisplayManager {
let mut warned = false;
while !stop_t.load(Ordering::Relaxed) {
if let Some(h) = vdm().device_handle() {
// SAFETY: `ping` requires `dev` to be a valid control handle. `h` is from
// `device_handle()` (the `Some` branch) — cached handles are NEVER closed for the
// process lifetime (a dead one is retired, kept alive; see `DeviceSlot`), so the
// handle stays valid for this call even if it was retired concurrently — at worst
// the IOCTL fails. The pinger thread only spins while the `&'static` manager
// singleton lives.
match unsafe { vdm().driver.ping(h) } {
// SAFETY: `ping` requires `dev` to be a valid control handle. The `h` Arc from
// `device_handle()` is held across this call, so the handle stays open even if
// it is retired concurrently — at worst the IOCTL fails (the retire drops only
// the manager's reference; see `DeviceSlot`). The pinger thread only spins
// while the `&'static` manager singleton lives.
match unsafe { vdm().driver.ping(dev_raw(&h)) } {
Ok(()) => warned = false,
Err(e) if is_device_gone(&e) => {
// The device itself is gone (driver upgrade / WUDFHost restart) — pings
@@ -1897,12 +1922,11 @@ impl VirtualDisplayManager {
slot,
"virtual-display: last session left (deliberate quit) — tearing down now, linger skipped"
);
// SAFETY: `teardown_removed` requires `dev` to be the live control handle; `dev`
// is the cached process-lifetime `OwnedHandle` from `device_handle()` (the `Some`
// checked above; cached handles are never closed — a dead one is retired, kept
// alive). `mon` was moved out of the map under the `state` lock, so it is
// exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev, &mut inner, mon) };
// SAFETY: `teardown_removed` requires `dev` to be the live control handle; the
// `dev` Arc from `device_handle()` (the `Some` checked above) is held across
// this call, so the handle stays open. `mon` was moved out of the map under the
// `state` lock, so it is exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev_raw(&dev), &mut inner, mon) };
}
None => {
inner.slots.insert(
@@ -1980,10 +2004,10 @@ impl VirtualDisplayManager {
"IDD-push setup: force-preempting the stuck-Active prior monitor (its IddCx swap-chain is dead)"
);
// SAFETY: `teardown_removed` requires `dev` to be the live control handle;
// `dev` is the cached process-lifetime `OwnedHandle` from `device_handle()`
// (the `Some` checked above). `mon` was moved out of the map under the
// `state` lock, so it is exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev, &mut inner, mon) };
// the `dev` Arc from `device_handle()` (the `Some` checked above) is held
// across this call, so the handle stays open. `mon` was moved out of the
// map under the `state` lock, so it is exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev_raw(&dev), &mut inner, mon) };
// Let the OS finish the ASYNC departure before the next ADD (mirrors the
// acquire() Lingering-preempt settle).
thread::sleep(Duration::from_millis(400));
@@ -2051,11 +2075,12 @@ impl VirtualDisplayManager {
// its session. Lock order stays state → device (teardown's invalidate
// path), same as every other holder; the pinger takes only the device
// lock — no inversion.
// SAFETY: `teardown_removed` requires a valid control handle; `dev` is
// from `self.device_handle()` (cached handles are never closed — a dead
// one is retired, kept alive; see `DeviceSlot`). `mon` was moved out of
// the map under the lock, so it is exclusively owned here.
unsafe { self.teardown_removed(dev, &mut g, mon) };
// SAFETY: `teardown_removed` requires a valid control handle; the `dev`
// Arc from `self.device_handle()` is held across this call, so the
// handle stays open (a concurrent retire drops only the manager's
// reference; see `DeviceSlot`). `mon` was moved out of the map under
// the lock, so it is exclusively owned here.
unsafe { self.teardown_removed(dev_raw(&dev), &mut g, mon) };
}
}
})
@@ -2218,11 +2243,11 @@ impl VirtualDisplayManager {
if let Some(SlotState::Lingering { mon, .. } | SlotState::Pinned { mon }) =
inner.slots.remove(&k)
{
// SAFETY: `teardown_removed` needs a live control handle; `dev` is from
// `device_handle()` (cached handles are never closed — a dead one is retired, kept
// alive; see `DeviceSlot`). `mon` was moved out of the map under the `state` lock,
// so it is exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev, &mut inner, mon) };
// SAFETY: `teardown_removed` needs a live control handle; the `dev` Arc from
// `device_handle()` is held across this call, so the handle stays open (see
// `DeviceSlot`). `mon` was moved out of the map under the `state` lock, so it is
// exclusively owned here — no aliasing.
unsafe { self.teardown_removed(dev_raw(&dev), &mut inner, mon) };
released += 1;
}
}
@@ -100,14 +100,27 @@ unsafe fn ioctl(h: HANDLE, code: u32, input: &[u8], output: &mut [u8]) -> Result
/// `reset-pf-vdisplay.ps1` step 2 (proven on-box). Best-effort + idempotent: only NOT-present nodes
/// (`Status != OK`) are removed, so the LIVE session's monitor (`Status OK`) is never touched; any
/// failure is logged and swallowed. Returns the number removed.
///
/// The outcome is logged UNCONDITIONALLY, as found + removed: the old script counted only removals
/// and the host spoke only when that count was positive, so a reap whose pnputil never launched and
/// a box with no ghosts produced byte-identical logs (silence) — the same vacuous-signal family as
/// the `status=OK` trap [`reload_vdisplay_adapter`] answers — while ghosts ratcheted toward the
/// wedge with every sleep cycle.
fn reap_ghost_monitors() -> u32 {
// Mirrors reset-pf-vdisplay.ps1 step 2. powershell is always present for the SYSTEM service; the
// matched tokens ('OK', 'punktfunk', the InstanceId) are locale-invariant, so this is safe on a
// non-English box (unlike a .ps1 *file* read in the machine codepage).
//
// pnputil is resolved by full path and `$LASTEXITCODE` pre-seeded to failure before every
// launch, exactly like the reload path below: a LocalSystem service's PATH need not include
// System32 (and a SYSTEM process must not trust PATH anyway — a planted `pnputil.exe` would run
// elevated), and the old bare-name call failed INVISIBLY there — `SilentlyContinue` swallowed
// the miss, no exit code was written, and the ghosts stayed to wedge `IOCTL_ADD` at 0x80070490.
const REAP_PS: &str = "$ErrorActionPreference='SilentlyContinue'; \
$g = Get-PnpDevice -Class Monitor | Where-Object { $_.Status -ne 'OK' -and $_.FriendlyName -match 'punktfunk' }; \
$n = 0; foreach ($d in $g) { pnputil /remove-device $d.InstanceId *> $null; if ($LASTEXITCODE -eq 0) { $n++ } }; \
Write-Output $n";
$g = @(Get-PnpDevice -Class Monitor | Where-Object { $_.Status -ne 'OK' -and $_.FriendlyName -match 'punktfunk' }); \
$pnp = ($env:SystemRoot + '\\System32\\pnputil.exe'); \
$n = 0; foreach ($d in $g) { $LASTEXITCODE = 1; if (Test-Path $pnp) { & $pnp /remove-device $d.InstanceId *> $null }; if ($LASTEXITCODE -eq 0) { $n++ } }; \
Write-Output ($g.Count.ToString() + ' ' + $n)";
// Resolve powershell by full path — the LocalSystem service's PATH is not guaranteed to include
// System32 — with a bare-name fallback.
let ps = std::env::var("SystemRoot")
@@ -125,17 +138,29 @@ fn reap_ghost_monitors() -> u32 {
.output()
{
Ok(o) => {
let n = String::from_utf8_lossy(&o.stdout)
.trim()
.parse::<u32>()
.unwrap_or(0);
if n > 0 {
let raw = String::from_utf8_lossy(&o.stdout);
let Some((found, removed)) = parse_reap_output(&raw) else {
tracing::warn!(
reaped = n,
output = %raw.trim(),
"pf-vdisplay: ghost-monitor reap died before reporting — ghost nodes (if any) still pin IddCx monitor slots"
);
return 0;
};
if found == 0 {
tracing::info!("pf-vdisplay: no ghost (not-present) virtual-monitor nodes to reap");
} else if removed < found {
tracing::warn!(
found,
removed,
"pf-vdisplay: ghost-monitor reap could NOT remove every ghost node — the leftovers keep pinning IddCx monitor slots toward the 0x80070490 wedge"
);
} else {
tracing::warn!(
reaped = removed,
"pf-vdisplay: reaped ghost (not-present) virtual-monitor nodes — IddCx slot-exhaustion prevention"
);
}
n
removed
}
Err(e) => {
tracing::warn!(error = %e, "pf-vdisplay: ghost-monitor reap could not spawn powershell");
@@ -144,6 +169,18 @@ fn reap_ghost_monitors() -> u32 {
}
}
/// Parse [`reap_ghost_monitors`]'s script output — `"<found> <removed>"`. Split out to be testable
/// without a box, like [`classify_reload_output`]: the field failure this answers was a reap whose
/// outcome could not be decoded from the log at all, so the decoding is worth pinning down. `None`
/// = the script died before reporting (callers treat that as "removed nothing", loudly).
fn parse_reap_output(out: &str) -> Option<(u32, u32)> {
let mut it = out.split_whitespace().map(str::parse::<u32>);
match (it.next(), it.next()) {
(Some(Ok(found)), Some(Ok(removed))) => Some((found, removed)),
_ => None,
}
}
/// What an adapter-cycle attempt actually DID — deliberately NOT the devnode's PnP status afterwards.
/// The old script reported that status, and a device it had failed to touch at all still reads `OK`,
/// so a no-op cycle was indistinguishable from a real one in the log (field report 2026-08-02: a
@@ -178,6 +215,14 @@ fn reload_vdisplay_adapter() -> AdapterCycle {
// device description — locale-invariant). Same spawn shape as `reap_ghost_monitors` above; the
// reported tokens are ours, so parsing them is locale-invariant too.
//
// The selector prefers LIVE devnodes: `Get-PnpDevice` also lists not-present PHANTOMS (an
// upgrade/reinstall leftover), and the old `Select-Object -First 1` could hand every recovery
// attempt a phantom — whose disable AND restart both fail — while a live node sat unexamined.
// A phantom-only state gets its own truthful refusal: no reload lever can revive a devnode
// record whose device is GONE; only re-creating the node (reinstall) can. `Present` is the
// authoritative bit, with `Status -ne 'Unknown'` as the fallback should it read null; live
// `OK` nodes sort ahead of problem-state ones.
//
// Every step that can fail is `-ErrorAction Stop` inside a `try` — the old script ran the whole
// cycle under `SilentlyContinue` and then reported `(Get-PnpDevice …).Status`, which reports the
// DEVICE, not the cycle: a disable that was refused left the device untouched, started, and
@@ -188,10 +233,19 @@ fn reload_vdisplay_adapter() -> AdapterCycle {
// let "never ran" read as "returned 0". Pre-seeding a failure means only a real exit 0 reports a
// reload. pnputil is resolved by full path — a LocalSystem service's PATH need not include
// System32.
//
// The REFUSED line carries the evidence a field log needs to tell the failure modes apart
// (2026-08-08: a woken box logged only `REFUSED Generic failure` — the WMI catch-all — leaving
// handle-veto vs phantom vs problem-state undecidable): how many devnodes matched and how many
// are live, the chosen node's PnP Status + ConfigManager problem code, and the pnputil
// /restart-device exit code the old script threw away (3010 = needs a reboot, which is its own
// diagnosis).
const CYCLE_PS: &str = "$ErrorActionPreference='SilentlyContinue'; \
$ad = Get-PnpDevice -Class Display | Where-Object { $_.FriendlyName -match 'punktfunk Virtual Display' } | Select-Object -First 1; \
if (-not $ad) { Write-Output 'ABSENT'; exit }; \
$id = $ad.InstanceId; $err = ''; \
$all = @(Get-PnpDevice -Class Display | Where-Object { $_.FriendlyName -match 'punktfunk Virtual Display' }); \
if ($all.Count -eq 0) { Write-Output 'ABSENT'; exit }; \
$live = @($all | Where-Object { $_.Present -or $_.Status -ne 'Unknown' } | Sort-Object { $_.Status -ne 'OK' }); \
if ($live.Count -eq 0) { Write-Output ('REFUSED only phantom (not-present) adapter devnodes remain (' + $all.Count + ') - the device node itself is gone and no reload can revive it; reinstalling the host re-creates it'); exit }; \
$ad = $live[0]; $id = $ad.InstanceId; $err = ''; \
try { \
Disable-PnpDevice -InstanceId $id -Confirm:$false -ErrorAction Stop; Start-Sleep -Seconds 2; \
try { Enable-PnpDevice -InstanceId $id -Confirm:$false -ErrorAction Stop } \
@@ -201,9 +255,11 @@ fn reload_vdisplay_adapter() -> AdapterCycle {
} catch { $err = ($_.Exception.Message -replace '\\s+', ' ') }; \
$pnp = ($env:SystemRoot + '\\System32\\pnputil.exe'); $LASTEXITCODE = 1; \
if (Test-Path $pnp) { & $pnp /restart-device $id *> $null }; \
if ($LASTEXITCODE -eq 0) { Start-Sleep -Seconds 2; \
$rx = $LASTEXITCODE; \
if ($rx -eq 0) { Start-Sleep -Seconds 2; \
Write-Output ('RELOADED restart ' + (Get-PnpDevice -InstanceId $id).Status) } \
else { Enable-PnpDevice -InstanceId $id -Confirm:$false; Write-Output ('REFUSED ' + $err) }";
else { Enable-PnpDevice -InstanceId $id -Confirm:$false; \
Write-Output ('REFUSED devnodes=' + $all.Count + ' live=' + $live.Count + ' status=' + $ad.Status + ' problem=' + $ad.ConfigManagerErrorCode + ' restart_exit=' + $rx + ' ' + $err) }";
let ps = std::env::var("SystemRoot")
.map(|r| format!(r"{r}\System32\WindowsPowerShell\v1.0\powershell.exe"))
.unwrap_or_else(|_| "powershell.exe".to_string());
@@ -1050,10 +1106,12 @@ const BRIEF_RETRY: Duration = Duration::from_secs(3);
/// them rather than N interleaved ones — each of which tears down the stack the others are waiting
/// on. The second caller through typically finds the interface already up and returns at once.
///
/// Taken ONLY by [`ensure_available`], which holds no manager lock, and released before the retire
/// hook below takes the manager's `device` mutex. That is what keeps the lock order one-way:
/// [`VdisplayDriver::open`] runs *inside* that same `device` mutex, so if it could also take this
/// lock the two orders would invert and deadlock. It cannot — it never reloads.
/// Taken ONLY by [`ensure_available`], which holds no manager lock. The lock order is one-way —
/// `RECOVERY` → `device`: the recovery's handle-release hooks (`invalidate_cached_device`, which
/// drops the manager's reference so the control handle can CLOSE before the PnP cycle) take the
/// `device` mutex while this is held. It must stay one-way: [`VdisplayDriver::open`] runs *inside*
/// that same `device` mutex, so if it could also take this lock the two orders would invert and
/// deadlock. It cannot — it never reloads.
static RECOVERY: std::sync::Mutex<()> = std::sync::Mutex::new(());
/// [`is_available`], with self-heal — and with PATIENCE, which is the part that matters after a
@@ -1069,10 +1127,11 @@ pub fn ensure_available() -> Result<()> {
let _serialize = RECOVERY.lock().unwrap_or_else(|e| e.into_inner());
wait_for_interface(NOT_READY_GRACE, true)
};
// OUTSIDE the recovery lock, by the ordering contract on `RECOVERY`. A reload tore the driver
// stack down and back up, so any control handle a previous session cached is dead by
// construction — retire it while we know that for certain, rather than leaving the next session
// to discover it by having an IOCTL fail. No-op before any backend opened the device.
// A reload tore the driver stack down and back up, so any control handle cached MEANWHILE (a
// racing open during the arrival window) is dead by construction — retire it while we know
// that for certain, rather than leaving the next session to discover it by having an IOCTL
// fail. Usually a no-op now: the recovery path already released the manager's reference
// before the reload (the handle-drain that lets the PnP cycle proceed at all).
if reloaded {
super::manager::invalidate_cached_device(
"the pf-vdisplay adapter was reloaded (hostless-zombie recovery)",
@@ -1119,12 +1178,33 @@ fn wait_for_interface(not_ready_grace: Duration, reload: bool) -> (Result<OwnedH
// Track how long we have seen NOTHING. Reset by any sighting, so a device that flickers
// between absent and not-ready is treated as the transition it is.
if probe.is_absent() {
if absent_since.is_none() && reload {
// First absent sighting on the recovery path: drop the manager's reference to the
// (dead) control device NOW, so the ABSENT_SETTLE below doubles as the drain window
// for every outstanding `Arc` clone — the handle then actually CLOSES before the
// reload runs. An open control handle is exactly what vetoes the PnP disable (and
// can wedge the pnputil restart) that the reload leans on; reset-pf-vdisplay.ps1
// stops the whole host service to get the same release (field 2026-08-08: every
// reload on a woken box came back REFUSED `Generic failure`). Gated on `reload`:
// the BRIEF_RETRY caller runs inside the manager's `device` mutex, where taking it
// again would deadlock — and that caller never reloads anyway.
super::manager::invalidate_cached_device(
"control interface absent — releasing the host's own device handle ahead of a \
possible adapter reload",
);
}
absent_since.get_or_insert_with(Instant::now);
} else {
absent_since = None;
}
let absent_long_enough = absent_since.is_some_and(|t| t.elapsed() >= ABSENT_SETTLE);
if reload && !reloaded && (absent_long_enough || Instant::now() >= deadline) {
// The not-ready path reaches here without the absent-sighting release above — drop the
// manager's reference now for the same reason (idempotent: a second call is a no-op).
super::manager::invalidate_cached_device(
"adapter reload imminent — releasing the host's own device handle (open handles \
veto the PnP cycle)",
);
match reload_vdisplay_adapter() {
// No devnode at all — waiting cannot conjure a driver. Fail immediately rather than
// burning the arrival window on a box that simply does not have it installed.
@@ -1195,6 +1275,32 @@ mod tests {
}
}
/// A refusal must carry evidence, not just a verdict. The 2026-08-08 field log showed only
/// `REFUSED Generic failure` — the WMI catch-all — leaving handle-veto vs phantom vs
/// problem-state undecidable from the log. The enriched line's tokens (devnode counts, PnP
/// status, problem code, the pnputil restart exit code the old script discarded) must survive
/// decoding verbatim, and the phantom-only state must decode as a refusal too — a reload
/// cannot revive a devnode record whose device is gone.
#[test]
fn a_refusal_keeps_its_evidence() {
let why = match classify_reload_output(
"REFUSED devnodes=2 live=1 status=OK problem=0 restart_exit=3010 Generic failure",
) {
AdapterCycle::Refused(why) => why,
other => panic!("expected Refused, got {}", variant(&other)),
};
for token in ["devnodes=2", "live=1", "status=OK", "restart_exit=3010"] {
assert!(why.contains(token), "{token} must survive: {why:?}");
}
assert!(matches!(
classify_reload_output(
"REFUSED only phantom (not-present) adapter devnodes remain (2) - the device node \
itself is gone and no reload can revive it; reinstalling the host re-creates it"
),
AdapterCycle::Refused(why) if why.contains("phantom")
));
}
/// The outcomes callers branch on: `NotInstalled` fails a session fast, `Reloaded` earns the
/// arrival window, and the lever that worked stays visible in the log (`restart` means the
/// disable was refused and something still holds the device open).
@@ -1226,6 +1332,29 @@ mod tests {
));
}
/// The reap's outcome must decode losslessly — the field ratchet (0.23→0.25) was a reap whose
/// bare-named pnputil never launched under the LocalSystem PATH while the host stayed silent:
/// "no ghosts" and "removed nothing" were byte-identical. Found and removed now travel
/// separately so a leftover ghost is loud, and the old single-number output (or a powershell
/// that died before reporting) must not decode as anything.
#[test]
fn reap_output_decodes_found_and_removed() {
assert_eq!(parse_reap_output("3 3\r\n"), Some((3, 3)));
assert_eq!(
parse_reap_output("4 0"),
Some((4, 0)),
"pnputil unlaunchable"
);
assert_eq!(parse_reap_output("0 0"), Some((0, 0)), "clean box");
for dead in ["5", "", " ", "garbage", "OK"] {
assert_eq!(
parse_reap_output(dead),
None,
"{dead:?} is not a reap report"
);
}
}
/// `is_absent` is what decides between WAITING and performing device surgery, so the two states
/// it separates are pinned here. An interface that is registered but not yet ACTIVE is a devnode
/// mid-transition — the wake-from-sleep case — and reloading the adapter under it only lengthens
+322 -28
View File
@@ -19,7 +19,7 @@ pub mod vkslot;
pub mod vulkan;
pub mod worker;
use std::sync::atomic::{AtomicBool, AtomicU32, Ordering};
use std::sync::atomic::{AtomicBool, AtomicU32, AtomicU64, Ordering};
pub use cuda::DeviceBuffer;
pub use egl::{DmabufPlane, EglImporter};
@@ -261,56 +261,223 @@ pub fn gpu_import_disabled() -> bool {
/// operator found `PUNKTFUNK_ZEROCOPY=0` by hand. The host already knows how to encode that
/// machine — capture just has to stop handing it dmabufs. Latching here is what makes the next
/// session negotiate CPU frames on its own.
static RAW_DMABUF_FAILURE_STREAK: AtomicU32 = AtomicU32::new(0);
static RAW_DMABUF_DISABLED: AtomicBool = AtomicBool::new(false);
/// Below the encoder's own rebuild budget, so the latch is set before the session it doomed ends.
const RAW_DMABUF_FAILURE_LATCH: u32 = 3;
/// Record an encoder-side raw-dmabuf import failure. Latches the process-wide disable after
/// `RAW_DMABUF_FAILURE_LATCH` consecutive failures.
/// Consecutive capture rebuilds whose dmabuf-only offer never negotiated before the passthrough is
/// latched off. **2 = one retry**, deliberately: each failed negotiation costs a ~10 s stall, so a
/// larger budget is paid by the user in dead air. One retry is enough to survive a compositor
/// caught mid-restart, which is the transient this exists for; a compositor that genuinely never
/// accepts keeps the same capture identity, so its streak accumulates and it latches on the second
/// try — one extra stall versus the old behaviour, once per host lifetime.
const RAW_DMABUF_NEGOTIATION_LATCH: u32 = 2;
/// The raw-dmabuf passthrough's off-switch — **two causes with two different lifetimes**, which is
/// the whole point of this type.
///
/// They used to share one `AtomicBool`, so the cheap recoverable cause (a negotiation that timed
/// out, possibly because the compositor was mid-restart) was as permanent as the expensive
/// unrecoverable one (an encoder that cannot import what this compositor allocates). Once either
/// fired, EVERY later session on the host captured CPU frames until the process was restarted —
/// including sessions against a completely different compositor and node, which had never failed
/// at anything.
///
/// * **Import failures stay sticky.** A driver that will not take what the compositor allocates
/// refuses identically on every retry, and the encode-stall recovery above cannot tell that from
/// a transient — it rebuilt the same failing encoder five times and then ended the session, on
/// every connection, forever. That is what this latch was born to stop, and it must keep
/// stopping it.
/// * **Negotiation timeouts get a retry budget** ([`RAW_DMABUF_NEGOTIATION_LATCH`]).
/// * **Both are keyed to a capture identity.** A new node id — a fresh virtual output, the
/// Bazzite Gaming↔Desktop switch, a compositor restart — is a genuinely different question, so
/// it earns a fresh dmabuf attempt instead of inheriting a verdict about something else.
///
/// Atomics rather than a lock because [`note_import_ok`](Self::note_import_ok) is on the per-frame
/// import path; everything else here runs at pipeline build or on failure.
#[derive(Debug)]
pub struct RawDmabufLatch {
import_streak: AtomicU32,
import_latched: AtomicBool,
negotiation_streak: AtomicU32,
negotiation_latched: AtomicBool,
/// The capture identity the counters above describe. `u64::MAX` = nothing observed yet (a real
/// identity is a node id, so it can never collide with the sentinel).
identity: AtomicU64,
}
/// Nothing observed yet — distinct from any real capture identity.
const NO_IDENTITY: u64 = u64::MAX;
impl RawDmabufLatch {
pub const fn new() -> Self {
RawDmabufLatch {
import_streak: AtomicU32::new(0),
import_latched: AtomicBool::new(false),
negotiation_streak: AtomicU32::new(0),
negotiation_latched: AtomicBool::new(false),
identity: AtomicU64::new(NO_IDENTITY),
}
}
/// Whether the raw-dmabuf passthrough is currently off, for either cause.
pub fn disabled(&self) -> bool {
self.import_latched.load(Ordering::Relaxed)
|| self.negotiation_latched.load(Ordering::Relaxed)
}
/// Tell the latch which capture is about to be built. A DIFFERENT capture from the one the
/// current verdict was formed against clears every counter and both latches, so the new
/// pipeline earns a fresh dmabuf attempt.
///
/// Returns `true` only when that clear actually **re-armed something** — i.e. the identity
/// changed *and* a latch was set. Deliberately not "the identity changed": every session on a
/// fresh virtual output changes it, and a caller that logged on that would print a re-arm line
/// on every healthy session open, which is noise. `true` means "this capture would have been
/// forced to CPU by an earlier capture's verdict, and no longer is".
///
/// Call this BEFORE reading [`disabled`](Self::disabled) for a negotiation decision, or the
/// decision is made against the previous capture's verdict.
pub fn observe_capture(&self, identity: u64) -> bool {
if self.identity.swap(identity, Ordering::Relaxed) == identity {
return false;
}
let was_latched = self.disabled();
self.import_streak.store(0, Ordering::Relaxed);
self.import_latched.store(false, Ordering::Relaxed);
self.negotiation_streak.store(0, Ordering::Relaxed);
self.negotiation_latched.store(false, Ordering::Relaxed);
was_latched
}
/// Record an encoder-side raw-dmabuf import failure. Returns `true` if this failure is the one
/// that latched the passthrough off.
pub fn note_import_failure(&self) -> Option<u32> {
let streak = self.import_streak.fetch_add(1, Ordering::Relaxed) + 1;
(streak >= RAW_DMABUF_FAILURE_LATCH && !self.import_latched.swap(true, Ordering::Relaxed))
.then_some(streak)
}
/// Record a raw dmabuf that imported and encoded — resets the failure streak. The per-frame
/// hot path, hence a single relaxed store.
///
/// Deliberately does NOT clear `import_latched`: once the latch fires, capture has already
/// moved to CPU frames, so there are no more dmabuf imports to succeed. Only a new capture
/// identity clears it.
pub fn note_import_ok(&self) {
self.import_streak.store(0, Ordering::Relaxed);
}
/// Record a capture rebuild whose dmabuf-only offer never negotiated. Returns `Some(streak)`
/// if this is the failure that latched the passthrough off, `None` while retries remain.
pub fn note_negotiation_timeout(&self) -> Option<u32> {
let streak = self.negotiation_streak.fetch_add(1, Ordering::Relaxed) + 1;
(streak >= RAW_DMABUF_NEGOTIATION_LATCH
&& !self.negotiation_latched.swap(true, Ordering::Relaxed))
.then_some(streak)
}
/// Record a capture whose dmabuf offer DID negotiate — the retry budget is per consecutive
/// run of failures, so a success spends none of it.
pub fn note_negotiation_ok(&self) {
self.negotiation_streak.store(0, Ordering::Relaxed);
}
/// Diagnostic for the session-open line: which cause (if any) currently holds it off.
pub fn state(&self) -> &'static str {
match (
self.import_latched.load(Ordering::Relaxed),
self.negotiation_latched.load(Ordering::Relaxed),
) {
(true, true) => "latched: encoder-import + negotiation",
(true, false) => "latched: encoder-import failures",
(false, true) => "latched: negotiation timeouts",
(false, false) => "live",
}
}
}
impl Default for RawDmabufLatch {
fn default() -> Self {
Self::new()
}
}
static RAW_DMABUF: RawDmabufLatch = RawDmabufLatch::new();
/// Record an encoder-side raw-dmabuf import failure. Latches the passthrough off after
/// `RAW_DMABUF_FAILURE_LATCH` consecutive failures, until the capture identity changes.
pub fn note_raw_dmabuf_import_failure(reason: &str) {
let streak = RAW_DMABUF_FAILURE_STREAK.fetch_add(1, Ordering::Relaxed) + 1;
if streak >= RAW_DMABUF_FAILURE_LATCH && !RAW_DMABUF_DISABLED.swap(true, Ordering::Relaxed) {
if let Some(streak) = RAW_DMABUF.note_import_failure() {
tracing::error!(
streak,
reason,
"zero-copy raw-dmabuf passthrough disabled for this host process: the encoder failed \
to import the compositor's dmabuf {streak} times in a row captures fall back to the \
CPU path (slower, but this host could not stream at all otherwise)"
"zero-copy raw-dmabuf passthrough disabled: the encoder failed to import the \
compositor's dmabuf {streak} times in a row captures fall back to the CPU path \
(slower, but this host could not stream at all otherwise). A new capture (different \
node / compositor) clears this."
);
}
}
/// Record a raw dmabuf that imported and encoded — resets the failure streak.
pub fn note_raw_dmabuf_import_ok() {
RAW_DMABUF_FAILURE_STREAK.store(0, Ordering::Relaxed);
RAW_DMABUF.note_import_ok();
}
/// Latch the raw-dmabuf passthrough off because its dmabuf-only *offer never negotiated* — the
/// CAPTURE-side counterpart to [`note_raw_dmabuf_import_failure`]'s encoder-side streak. One
/// timeout is conclusive for this offer (a compositor that cannot allocate the requested
/// LINEAR/modifier BGRx dmabuf refuses it identically on every retry), so there is no streak to
/// count: the next capture skips the passthrough and negotiates SHM/CPU instead of re-running the
/// same 10 s timeout on every reconnect.
/// CAPTURE-side counterpart to [`note_raw_dmabuf_import_failure`]'s encoder-side streak.
///
/// Unlike the import streak this gets a retry budget: the offer can time out because the
/// compositor was mid-restart rather than because it will never accept, and the old behaviour
/// (one timeout = CPU capture for the rest of the host's life, for every compositor and every
/// node) turned a transient into a permanent downgrade nobody could see.
///
/// Scoped deliberately. This used to be `note_vaapi_dmabuf_failed`, which fed [`enabled`] and so
/// disabled ALL zero-copy host-wide — see [`enabled`]. `RAW_DMABUF_DISABLED` gates only the
/// raw-passthrough decision, so the EGL→CUDA importer that a later NVENC session builds is
/// untouched.
/// disabled ALL zero-copy host-wide — see [`enabled`]. It gates only the raw-passthrough decision,
/// so the EGL→CUDA importer that a later NVENC session builds is untouched.
pub fn note_raw_dmabuf_negotiation_failed() {
if !RAW_DMABUF_DISABLED.swap(true, Ordering::Relaxed) {
tracing::warn!(
"zero-copy raw-dmabuf passthrough disabled for this host process: the compositor never \
accepted the dmabuf-only capture offer, so later captures negotiate the CPU path \
instead of repeating that timeout (the EGLCUDA import path is NOT affected)"
);
match RAW_DMABUF.note_negotiation_timeout() {
Some(streak) => tracing::warn!(
streak,
"zero-copy raw-dmabuf passthrough disabled: the compositor did not accept the \
dmabuf-only capture offer {streak} builds in a row, so later captures negotiate the \
CPU path instead of repeating that timeout (the EGLCUDA import path is NOT \
affected). A new capture (different node / compositor) clears this."
),
None => tracing::warn!(
"the compositor did not accept the dmabuf-only capture offer — retrying dmabuf on the \
next capture build before giving up on it"
),
}
}
/// True once repeated encoder import failures latched the raw-dmabuf passthrough off (see
/// [`note_raw_dmabuf_import_failure`]).
/// Record a capture whose dmabuf offer negotiated — spends none of the retry budget.
pub fn note_raw_dmabuf_negotiation_ok() {
RAW_DMABUF.note_negotiation_ok();
}
/// Tell the latch which capture is about to be built, so a verdict formed against a DIFFERENT
/// compositor/node is not inherited. Returns `true` if a latch was cleared by the change.
pub fn note_raw_dmabuf_capture(identity: u64) -> bool {
let cleared = RAW_DMABUF.observe_capture(identity);
if cleared {
tracing::info!(
identity,
"zero-copy raw-dmabuf passthrough re-armed: this is a different capture from the one \
that failed, so it gets a fresh dmabuf attempt"
);
}
cleared
}
/// True while either cause holds the raw-dmabuf passthrough off (see [`RawDmabufLatch`]).
pub fn raw_dmabuf_import_disabled() -> bool {
RAW_DMABUF_DISABLED.load(Ordering::Relaxed)
RAW_DMABUF.disabled()
}
/// Which cause holds the passthrough off, for the session-open diagnostic line.
pub fn raw_dmabuf_latch_state() -> &'static str {
RAW_DMABUF.state()
}
/// The EGL→CUDA twin of the raw-passthrough negotiation latch: the capture advertised the GPU
@@ -564,4 +731,131 @@ mod tests {
note_gpu_import_death(); // third consecutive death
assert!(gpu_import_disabled());
}
// ---- PW3: the raw-dmabuf latch's two lifetimes ------------------------------------------
//
// Against a LOCAL `RawDmabufLatch`, never the process-wide static: these assertions are about
// the state machine, and sharing one global across a test binary's threads is how a latch test
// becomes order-dependent.
/// The expensive cause stays sticky. A driver that cannot import what this compositor
/// allocates refuses identically every time, and the encode-stall recovery cannot tell that
/// from a transient — this latch is what stops it rebuilding the same doomed encoder forever.
#[test]
fn import_failures_latch_and_stay_latched() {
let l = RawDmabufLatch::new();
assert!(!l.disabled());
assert_eq!(l.note_import_failure(), None); // 1
assert_eq!(l.note_import_failure(), None); // 2
assert!(!l.disabled(), "must not latch before the streak completes");
assert_eq!(l.note_import_failure(), Some(3));
assert!(l.disabled());
// Only the FIRST crossing reports, so the error line cannot repeat per frame.
assert_eq!(l.note_import_failure(), None);
// A success resets the streak but must NOT unlatch: once capture moved to CPU frames there
// are no more dmabuf imports, so an "ok" here would be about something else entirely.
l.note_import_ok();
assert!(l.disabled());
}
/// A run of failures broken by a success spends none of the budget — the streak is
/// consecutive-only, which is what makes an occasional failure survivable.
#[test]
fn a_success_breaks_the_import_streak() {
let l = RawDmabufLatch::new();
l.note_import_failure();
l.note_import_failure();
l.note_import_ok();
assert_eq!(l.note_import_failure(), None, "streak restarted at 1");
assert_eq!(l.note_import_failure(), None);
assert!(!l.disabled());
assert_eq!(l.note_import_failure(), Some(3));
}
/// The cheap cause gets a retry. This is the behaviour change PW3 exists for: one timeout used
/// to mean CPU capture for the rest of the host's life, on every compositor and every node.
#[test]
fn a_negotiation_timeout_is_retried_before_it_latches() {
let l = RawDmabufLatch::new();
assert_eq!(l.note_negotiation_timeout(), None, "first one retries");
assert!(
!l.disabled(),
"the next capture build must still be allowed to try dmabuf"
);
assert_eq!(l.note_negotiation_timeout(), Some(2));
assert!(l.disabled());
assert_eq!(l.note_negotiation_timeout(), None, "reports once");
}
/// A capture that negotiates credits the budget back, so a compositor that fails once and then
/// works never accumulates its way to a latch across an evening of reconnects.
#[test]
fn a_negotiated_capture_credits_the_retry_budget() {
let l = RawDmabufLatch::new();
for _ in 0..10 {
assert_eq!(l.note_negotiation_timeout(), None);
l.note_negotiation_ok();
}
assert!(!l.disabled());
}
/// A different capture is a different question. New node id (fresh virtual output, compositor
/// restart, the Bazzite Gaming↔Desktop switch) clears BOTH causes — the same capture does not.
#[test]
fn a_new_capture_identity_clears_the_latch_and_the_same_one_does_not() {
let l = RawDmabufLatch::new();
// Nothing is latched yet, so observing a new capture re-arms NOTHING — that is what the
// return value means, and it is why a healthy session open logs no re-arm line.
assert!(
!l.observe_capture(7),
"nothing was latched, nothing re-armed"
);
assert!(!l.observe_capture(7), "same capture, no clear");
for _ in 0..RAW_DMABUF_FAILURE_LATCH {
l.note_import_failure();
}
assert!(l.disabled());
assert!(
!l.observe_capture(7),
"the SAME capture must keep its verdict — this is the 10s-stall hazard the latch exists for"
);
assert!(l.disabled());
assert!(l.observe_capture(9), "a different node re-arms it");
assert!(!l.disabled());
// ...and the streaks reset with it, so the fresh attempt gets a full budget.
assert_eq!(l.note_import_failure(), None);
}
/// The negotiation latch is keyed the same way — a compositor restart must not inherit the
/// previous one's timeout verdict.
#[test]
fn a_new_capture_identity_clears_the_negotiation_latch_too() {
let l = RawDmabufLatch::new();
l.observe_capture(1);
l.note_negotiation_timeout();
l.note_negotiation_timeout();
assert!(l.disabled());
assert!(l.observe_capture(2));
assert!(!l.disabled());
}
/// The session-open line has to name WHICH cause holds it off — "cpu because nothing here
/// does dmabuf" and "cpu because something failed earlier" are different bugs.
#[test]
fn latch_state_names_the_cause() {
let l = RawDmabufLatch::new();
assert_eq!(l.state(), "live");
l.note_negotiation_timeout();
l.note_negotiation_timeout();
assert_eq!(l.state(), "latched: negotiation timeouts");
let l = RawDmabufLatch::new();
for _ in 0..RAW_DMABUF_FAILURE_LATCH {
l.note_import_failure();
}
assert_eq!(l.state(), "latched: encoder-import failures");
for _ in 0..RAW_DMABUF_NEGOTIATION_LATCH {
l.note_negotiation_timeout();
}
assert_eq!(l.state(), "latched: encoder-import + negotiation");
}
}
+368 -7
View File
@@ -497,6 +497,31 @@ const SHRINK_QUIET_MS: u32 = 30_000;
/// The same, while the A/V sync loop is actively asking for a shallower ring — see the branch in
/// [`JitterPolicy::note_read`] that selects between them.
const SHRINK_QUIET_SYNC_MS: u32 = 5_000;
/// Post-read depth below which a served callback counts as a NEAR-MISS: the device got its
/// samples, but less than one protocol frame was left in hand, so the next callback starves
/// unless a packet lands inside one frame time. On a healthy link the post-read depth hovers a
/// whole target above this, which is what makes a near-miss evidence of real delivery jitter —
/// the same evidence as an underrun, except nobody heard it yet.
const NEAR_MISS_MARGIN_MS: u32 = FRAME_MS;
/// How long a shrink remains a PROBE, in consumed audio: an underrun or near-miss inside this
/// window means the shrink was wrong, and the previous target is restored at once instead of
/// being re-learned three audible underruns at a time.
const SHRINK_PROBE_MS: u32 = 5_000;
/// A ring is HOLLOW when its depth AVERAGE sits this far below the target: the target promises a
/// depth the ring does not actually hold. Growth only ever raises the promise — the one thing
/// that re-banks real depth is a re-prime — so an underrun in a hollow ring re-primes AT ONCE:
/// the click has already happened, and spending it on the whole refill is strictly better than
/// riding the knife edge and paying a click per bunching period indefinitely, which is what the
/// consecutive-empties hysteresis alone converges to. A full ring's underrun (one packet a few
/// ms late) is nowhere near hollow and keeps the hysteresis.
const DEPRIME_DEBT_MS: u32 = GROW_STEP_MS;
/// How long a failed probe keeps the sync loop from driving another shrink. Without this the
/// loop pays an audible starvation event every [`SHRINK_QUIET_SYNC_MS`] on any link whose jitter
/// genuinely needs the depth — sync asks for less, the ring shrinks, the link answers, the ring
/// grows back, five quiet seconds later sync asks again, forever. Doubles per consecutive
/// failure up to [`SYNC_BACKOFF_MAX_MS`]; a probe that survives its window resets it.
const SYNC_BACKOFF_MS: u32 = 60_000;
const SYNC_BACKOFF_MAX_MS: u32 = 480_000;
/// The playback de-jitter state machine shared by every client's audio ring.
///
@@ -539,6 +564,24 @@ pub struct JitterPolicy {
/// behaviour exactly, which is what lets the four client rings adopt this one at a time
/// without diverging in the meantime.
sync_target: Option<usize>,
/// Set by [`step`](Self::step) when the read it authorised leaves less than
/// [`NEAR_MISS_MARGIN_MS`] buffered; consumed by [`note_read`](Self::note_read).
near_miss: bool,
/// A near-miss already grew the target this window — one step per window, so a single
/// bunching episode (which lands as a RUN of consecutive near-misses while the ring refills)
/// buys one measured step, not a sprint to the ceiling.
near_miss_grown: bool,
/// Set by [`step`](Self::step): the depth average sits more than [`DEPRIME_DEBT_MS`] below
/// the target, so an underrun should re-prime at once instead of waiting out the hysteresis.
hollow: bool,
/// Consumed samples left in the current shrink-probe window (0 = no probe outstanding).
probe_run: usize,
/// The live target before the probed shrink, restored if the probe fails.
probe_prev_target: usize,
/// Consumed samples before the sync loop may drive another shrink (0 = allowed now).
sync_backoff_run: usize,
/// Length of the NEXT backoff, in ms — doubles per consecutive failed probe, capped.
sync_backoff_ms: u32,
}
impl JitterPolicy {
@@ -558,6 +601,13 @@ impl JitterPolicy {
quiet_run: 0,
last_want: 0,
sync_target: None,
near_miss: false,
near_miss_grown: false,
hollow: false,
probe_run: 0,
probe_prev_target: 0,
sync_backoff_run: 0,
sync_backoff_ms: SYNC_BACKOFF_MS,
}
}
@@ -667,8 +717,26 @@ impl JitterPolicy {
if !self.primed && depth.saturating_sub(out.drop_front) >= target {
self.primed = true;
self.empties = 0;
// The refill just banked this much: seed the average with it rather than letting it
// climb from wherever the drought left it — a freshly-primed ring would otherwise
// read as hollow for the EWMA's whole settling time, and the FIRST late packet
// would re-prime a ring that is actually full.
self.depth_avg = depth.saturating_sub(out.drop_front) as f32;
}
out.silence = !self.primed;
// Near-miss: this read will be served, but with less than one frame left over — the
// next callback starves unless a packet lands within one frame time. Unconditional
// assignment, so a stale flag can never survive a de-prime into the next primed read.
let after = depth.saturating_sub(out.drop_front);
self.near_miss = self.primed
&& after >= want
&& after - want < NEAR_MISS_MARGIN_MS as usize * self.per_ms;
// Hollow: the depth AVERAGE runs a debt against the target — the promise has been raised
// but the depth was never re-banked (see `DEPRIME_DEBT_MS`). Judged on the average, not
// this instant: a single late packet empties the ring for a callback without making it
// hollow, and must keep the consecutive-empties hysteresis.
self.hollow = self.primed
&& (self.depth_avg as usize + DEPRIME_DEBT_MS as usize * self.per_ms) < target;
out
}
@@ -683,19 +751,51 @@ impl JitterPolicy {
return;
}
let want = self.last_want.max(1);
let near_miss = std::mem::take(&mut self.near_miss);
self.window_run += want;
if self.window_run >= GROW_WINDOW_MS as usize * self.per_ms {
self.window_run = 0;
self.underruns = 0;
self.near_miss_grown = false;
}
self.sync_backoff_run = self.sync_backoff_run.saturating_sub(want);
let mut restored = false;
if self.probe_run > 0 {
self.probe_run = self.probe_run.saturating_sub(want);
if ran_short || near_miss {
// The probe FAILED: the link answered a shrink with (nearly) starving the ring.
// Take the depth straight back — re-learning it three audible underruns at a
// time is what made the sync-vs-growth tug-of-war audible — and keep the sync
// loop from probing again for a while, doubling per consecutive failure. The
// residual A/V offset is reported instead; continuity outranks sync. The
// restore CONSUMES this event as growth evidence: it answered a depth the ring
// is no longer at, so growing past the proven target on top would overshoot.
self.probe_run = 0;
self.target = self.target.max(self.probe_prev_target);
self.sync_backoff_run = self.sync_backoff_ms as usize * self.per_ms;
self.sync_backoff_ms = (self.sync_backoff_ms * 2).min(SYNC_BACKOFF_MAX_MS);
restored = true;
} else if self.probe_run == 0 {
// Survived the whole window: the shallower depth is genuinely safe here, so the
// next probe starts from a clean slate.
self.sync_backoff_ms = SYNC_BACKOFF_MS;
}
}
if ran_short {
self.quiet_run = 0;
self.empties += 1;
if self.empties >= self.tuning.deprime_after {
if self.empties >= self.tuning.deprime_after || self.hollow {
// The consecutive-empties hysteresis protects a FULL ring from one late packet.
// A hollow ring is the opposite case: the target has been raised but the depth
// never re-banked (growth is a promise; only a re-prime cashes it), and riding
// that out is a click per bunching period, forever. The click just heard has
// already paid for the refill — take it now.
self.primed = false;
self.empties = 0;
}
self.underruns += 1;
if !restored {
self.underruns += 1;
}
if self.underruns >= GROW_UNDERRUNS {
// This device genuinely needs more slack than the base target. Grow ONCE per
// window, capped — the alternative (every device pre-paying the worst device's
@@ -705,17 +805,33 @@ impl JitterPolicy {
let grown = self.target + GROW_STEP_MS as usize * self.per_ms;
self.target = grown.min(self.tuning.max_target_ms as usize * self.per_ms);
}
} else if near_miss {
// Came within one frame of an underrun — the same evidence as one, heard by no one.
// Growing here, BEFORE the click, is what "no audible jitter" means: waiting for
// the third audible underrun means the user heard two. One step per window (a
// bunching episode is a RUN of near-misses while the ring refills, and must buy one
// measured step, not a sprint to the ceiling); if it worsens into real underruns
// the path above takes over. A near-miss is pressure, not quiet.
self.quiet_run = 0;
self.empties = 0;
if !self.near_miss_grown && !restored {
self.near_miss_grown = true;
let grown = self.target + GROW_STEP_MS as usize * self.per_ms;
self.target = grown.min(self.tuning.max_target_ms as usize * self.per_ms);
}
} else {
self.empties = 0;
self.quiet_run += want;
// A grown target normally relaxes only after a long quiet spell, because without other
// evidence the only thing that can justify giving up hard-won slack is time. When the
// sync loop is asking to run shallower it IS that evidence — a measurement saying the
// extra depth is costing alignment right now — so test a smaller target sooner. Wrong
// guesses are cheap and self-correcting: one underrun and the growth path takes it
// straight back. Without this a ring that ratcheted to the ceiling during a transient
// would hold the audio a ceiling's worth late for minutes after the cause had gone.
let quiet_needed = if self.sync_wants_less() {
// extra depth is costing alignment right now — so test a smaller target sooner. Every
// shrink is armed as a PROBE: answered by an underrun or near-miss it is undone at
// once (see above), and a failed sync-driven guess is not retried for a backoff —
// without that, a link whose jitter genuinely needs the depth pays an audible
// starvation event every five seconds, forever.
let sync_shrink = self.sync_wants_less() && self.sync_backoff_run == 0;
let quiet_needed = if sync_shrink {
SHRINK_QUIET_SYNC_MS
} else {
SHRINK_QUIET_MS
@@ -725,10 +841,15 @@ impl JitterPolicy {
// doesn't cost latency for the rest of the session.
self.quiet_run = 0;
let base = self.tuning.base_target_ms as usize * self.per_ms;
let prev = self.target;
self.target = self
.target
.saturating_sub(GROW_STEP_MS as usize * self.per_ms)
.max(base);
if self.target < prev {
self.probe_run = SHRINK_PROBE_MS as usize * self.per_ms;
self.probe_prev_target = prev;
}
}
}
}
@@ -1937,4 +2058,244 @@ mod tests {
"sync pressure should relax sooner: {fast_reads} vs {slow_reads} quiet reads"
);
}
// ---- near-miss growth and shrink probes (the audible-limit-cycle fixes) ---------------
/// A primed read that is served but leaves less than one frame buffered is a NEAR-MISS —
/// the same evidence as an underrun, heard by no one — and must grow the target BEFORE the
/// click, not after the third one. One step per window: a bunching episode lands as a run of
/// consecutive near-misses while the ring refills, and must not sprint to the ceiling.
#[test]
fn a_near_miss_grows_the_target_without_an_underrun() {
let t = JitterTuning::COREAUDIO;
let pm = per_ms(2);
let want = 5 * pm;
let mut p = JitterPolicy::new(t, 2);
p.step(t.base_target_ms as usize * pm, want); // primes exactly at target
assert!(p.is_primed());
let base = p.target_ms();
// Serve the callback with less than one frame left over: depth = want + (margin 1).
p.step(want + NEAR_MISS_MARGIN_MS as usize * pm - 1, want);
p.note_read(false); // NOT short — the device got its samples
assert_eq!(
p.target_ms(),
base + GROW_STEP_MS,
"a near-miss must buy one step"
);
// A second near-miss in the same window is the same episode: no further growth.
p.step(want + pm, want);
p.note_read(false);
assert_eq!(p.target_ms(), base + GROW_STEP_MS, "one step per window");
// A healthy read does not grow anything.
let grown = p.target_ms();
p.step(grown as usize * pm + want, want);
p.note_read(false);
assert_eq!(p.target_ms(), grown);
}
/// A healthy steady depth must never read as a near-miss: the margin is one frame, and a
/// ring hovering at target sits a whole target above it.
#[test]
fn steady_depth_never_grows_the_target() {
let t = JitterTuning::PIPEWIRE;
let pm = per_ms(2);
let want = 5 * pm;
let mut p = JitterPolicy::new(t, 2);
for _ in 0..(60_000 / 5) {
// one minute of clean callbacks
p.step(t.base_target_ms as usize * pm + want, want);
p.note_read(false);
}
assert_eq!(p.target_ms(), t.base_target_ms);
}
/// A shrink answered by an underrun (or near-miss) inside its probe window is undone AT
/// ONCE — re-learning the depth three audible underruns at a time is what made the
/// sync-vs-growth tug-of-war audible in the field.
#[test]
fn a_failed_shrink_probe_is_undone_at_once() {
let t = JitterTuning::COREAUDIO;
let pm = per_ms(2);
let want = 5 * pm;
let mut p = JitterPolicy::new(t, 2);
// Grow the floor two steps the audible way.
for _ in 0..(2 * GROW_UNDERRUNS) {
while !p.is_primed() {
p.step(200 * pm, want);
}
p.step(200 * pm, want);
p.note_read(true);
}
let grown = p.target_ms();
assert!(grown > t.base_target_ms);
// Sync asks for less; five quiet seconds later the shrink probes.
p.set_sync_target(Some(pm));
let depth = grown as usize * pm + want;
while p.target_ms() == grown {
p.step(depth, want);
p.note_read(false);
}
assert_eq!(p.target_ms(), grown - GROW_STEP_MS);
// ONE near-miss — nobody heard anything yet — and the depth is back.
p.step(want + pm, want);
p.note_read(false);
assert_eq!(
p.target_ms(),
grown,
"a failed probe must restore the target on the first near-miss"
);
}
/// After a failed probe the sync loop may not drive another shrink at the accelerated
/// cadence — the slow, pre-sync window still applies, the five-second one does not.
#[test]
fn a_failed_probe_backs_the_sync_shrink_off() {
let t = JitterTuning::COREAUDIO;
let pm = per_ms(2);
let want = 5 * pm;
let mut p = JitterPolicy::new(t, 2);
for _ in 0..(2 * GROW_UNDERRUNS) {
while !p.is_primed() {
p.step(200 * pm, want);
}
p.step(200 * pm, want);
p.note_read(true);
}
let grown = p.target_ms();
p.set_sync_target(Some(pm));
let depth = grown as usize * pm + want;
// First sync-driven shrink, then fail its probe.
while p.target_ms() == grown {
p.step(depth, want);
p.note_read(false);
}
p.step(want + pm, want);
p.note_read(false);
assert_eq!(p.target_ms(), grown, "restored");
// Twice the accelerated window of clean audio: the backed-off loop must NOT have
// shrunk again (before the fix this was exactly one audible failure per five seconds).
for _ in 0..(2 * SHRINK_QUIET_SYNC_MS / 5) {
p.step(depth, want);
p.note_read(false);
}
assert_eq!(
p.target_ms(),
grown,
"the accelerated cadence must be suspended after a failure"
);
// The slow pre-sync window still relaxes it eventually — backoff is not a freeze.
for _ in 0..(2 * SHRINK_QUIET_MS / 5) {
p.step(depth, want);
p.note_read(false);
}
assert!(
p.target_ms() < grown,
"the slow window must still be allowed to test a shrink"
);
}
/// One simulated bunching run's outcome.
#[derive(Debug, Default)]
struct BunchSim {
/// Reads that actually starved the device — each one is audible.
audible: u32,
/// Audible reads in the second half of the run: non-zero means the policy never
/// converged and the user hears it forever.
audible_tail: u32,
}
/// Drive a policy over a link that BUNCHES: delivery pauses for `gap_ms` every `period_ms`,
/// then the withheld audio arrives at once — the Wi-Fi power-save pattern from the field
/// reports, where the total rate is fine and only the spacing is wrong. `drift_ppm` is the
/// host-vs-DAC clock skew; a slightly slow host (negative) erodes the depth over minutes,
/// which is what keeps re-testing whatever target the policy has settled on — without it a
/// simulated ring freezes wherever priming left it and a wrong target is never punished.
fn simulate_bunching(
tuning: JitterTuning,
sync_target: Option<usize>,
ms: u32,
gap_ms: u32,
period_ms: u32,
drift_ppm: i64,
) -> BunchSim {
let pm = per_ms(2);
let want = 5 * pm;
let mut p = JitterPolicy::new(tuning, 2);
p.set_sync_target(sync_target);
let mut depth = 0usize;
let mut withheld = 0usize;
let mut carry: i64 = 0;
let mut out = BunchSim::default();
for cb in 0..(ms / 5) {
// The host keeps producing (want ± drift per callback); the link decides delivery.
carry += want as i64 * drift_ppm;
let extra = carry / 1_000_000;
carry -= extra * 1_000_000;
let produced = (want as i64 + extra).max(0) as usize;
let in_gap = (cb * 5) % period_ms < gap_ms;
if in_gap {
withheld += produced;
} else {
depth += produced + std::mem::take(&mut withheld);
}
let s = p.step(depth, want);
depth -= s.drop_front.min(depth);
if s.silence {
p.note_read(false);
continue;
}
let short = depth < want;
depth -= want.min(depth);
if short {
out.audible += 1;
if cb >= ms / 10 {
out.audible_tail += 1;
}
}
p.note_read(short);
}
out
}
/// THE field regression this whole change is for. A link that bunches needs ~30 ms of ring;
/// the sync loop wants less. Before this change the policy paid an audible event nearly
/// every bunching period, indefinitely — this exact simulation measured ~2000 over ten
/// minutes: the sync loop re-probed a proven depth every five quiet seconds, growth needed
/// three audible underruns to answer, and a grown target was never re-banked (growth raises
/// a threshold; only a re-prime deepens the ring), so the depth rode the knife edge. Now
/// near-misses grow the target before the first click, a failed shrink probe is undone at
/// once and backs the sync loop off, and a hollow ring cashes the whole refill on the click
/// it already paid. What remains is the clock-skew re-anchor — a slightly slow host
/// genuinely starves the ring every few minutes, and only rate adaptation (which no client
/// has) could remove that — so the bound is "a handful over ten minutes", not zero.
#[test]
fn sync_pressure_on_a_bunching_link_converges_instead_of_clicking_forever() {
// 25 ms gaps every 300 ms, a slightly slow host, ten minutes, sync permanently asking
// for a 5 ms ring.
let s = simulate_bunching(
JitterTuning::COREAUDIO,
Some(per_ms(2) * 5),
600_000,
25,
300,
-50,
);
assert!(
s.audible_tail <= 4,
"the tug-of-war must converge to the skew floor: {s:?}"
);
assert!(
s.audible <= 12,
"learning the link may cost a handful of audible events, not a stream of them: {s:?}"
);
}
/// The same link without sync pressure — the plain adaptive-growth behaviour — must land in
/// the same place: sync steering may not add a persistent audible cost over not steering.
#[test]
fn a_bunching_link_without_sync_stays_clean_after_growing() {
let s = simulate_bunching(JitterTuning::COREAUDIO, None, 600_000, 25, 300, -50);
assert!(s.audible_tail <= 4, "{s:?}");
assert!(s.audible <= 12, "{s:?}");
}
}
@@ -353,6 +353,18 @@ impl FrameChannel {
/// all-intra stream ([`Self::set_all_intra`]) a multi-deep queue drains to the NEWEST AU
/// instead — the skipped ones are already superseded and decode independently, so showing
/// them only adds latency.
///
/// ⚠ **The all-intra drain counts QUEUE ENTRIES and assumes one entry == one AU.** That holds
/// today only because slice-progressive delivery is refused on PyroWave
/// (`client/pump/handshake.rs`; see [`crate::session::Session::set_deliver_frame_parts`]).
/// Turn parts on for an all-intra stream and one AU pushes several entries, at which point
/// `len > 1` no longer means "the consumer is behind": this fires mid-AU, hands back a SUFFIX
/// and `clear()`s that AU's own prefixes — a headerless frame, every frame. Anyone making the
/// two composable must skip whole SUPERSEDED AUs (drop up to the newest entry whose
/// `part.first` is set, never split an AU), give `push`'s `FRAME_QUEUE_HARD_CAP` eviction the
/// same rule, and count `skipped_total` in AUs. Host-side streamed AUs
/// ([`crate::quic::VIDEO_CAP_STREAMED_AU`]) are NOT affected — they still arrive as one
/// completed `Frame` per AU.
pub(crate) fn pop(&self, timeout: Duration) -> FramePop {
let mut st = self.inner.lock().unwrap();
if st.q.is_empty() && !st.closed {
@@ -229,7 +229,10 @@ pub(super) async fn connect_and_handshake(args: &WorkerArgs) -> Result<Handshake
}
// Slice-progressive delivery (the embedder's opt-in): AU prefixes hand up as
// `Frame::part` pieces while the tail is still on the wire. Never on PyroWave — its
// all-intra frame channel drains newest-wins, which assumes whole AUs.
// all-intra frame channel drains newest-wins per QUEUE ENTRY, so parts of one AU read as
// separate AUs and the drain shreds the AU it is mid-way through (`FrameChannel::pop`
// spells out the mechanism and what a fix would take). Unrelated to the host's streamed-AU
// wire (`VIDEO_CAP_STREAMED_AU`), which still completes one whole `Frame` per AU.
if args.frame_parts && welcome.codec != crate::quic::CODEC_PYROWAVE {
session.set_deliver_frame_parts(true);
}
+128 -1
View File
@@ -82,6 +82,49 @@ fn stream_transport_idle(idle: std::time::Duration) -> Arc<quinn::TransportConfi
Arc::new(t)
}
/// Endpoint config for the CLIENT endpoint — the half of the jumbo opt-in that lives on the
/// receiving side, and without which the whole jumbo leg is unreachable.
///
/// `EndpointConfig::max_udp_payload_size` is the QUIC transport parameter this endpoint
/// advertises: "the largest UDP payload I accept". quinn defaults it to **1472** (a 1500-byte
/// Ethernet MTU), and a peer's MTU-discovery search is upper-bounded by
/// `min(MtuDiscoveryConfig::upper_bound, the value the OTHER side advertised)`
/// (`quinn_proto::connection::mtud::SearchState::new`). So raising the host's probe ceiling
/// alone — which is all [`stream_transport_idle`] did — can never make a host's discovery
/// settle above 1472: the *client's* default advertisement caps it, and the host's
/// settled-at-jumbo proof (`native/wire_mtu.rs`, both the mid-session grow and the
/// session-start one) could never fire. This raises the advertisement to the sealed jumbo
/// datagram size so the proof is obtainable at all.
///
/// Gated on the SAME operator opt-in as the probe ceiling ([`crate::config::jumbo_wire_mtu`],
/// i.e. `PUNKTFUNK_JUMBO=1` / `PUNKTFUNK_WIRE_MTU` > 1500) because it is not free: quinn sizes
/// its endpoint receive buffer as `max_udp_payload_size × max_receive_segments × BATCH_SIZE`,
/// which on a GRO-capable Linux/Android client is 64 × 32 segments — ~2.9 MiB at the 1472
/// default, ~18 MiB at jumbo. A jumbo LAN is a deliberate deployment; every other client keeps
/// today's buffer to the byte. Without the opt-in this returns the stock config, so the
/// advertisement, the wire, and the memory are all unchanged.
fn endpoint_config() -> quinn::EndpointConfig {
let mut cfg = quinn::EndpointConfig::default();
if let Some(mtu) = crate::config::jumbo_wire_mtu() {
// Derived exactly like the probe ceiling above (IPv4 overhead — a v6 peer's sealed
// target is smaller, so this covers it), and clamped into quinn's accepted range.
let shard = crate::config::jumbo_shard_payload_for(
mtu,
std::net::IpAddr::V4(std::net::Ipv4Addr::UNSPECIFIED),
);
let accept = crate::config::sealed_datagram_bytes(shard).clamp(1200, 65_527) as u16;
if cfg.max_udp_payload_size(accept).is_ok() {
tracing::info!(
max_udp_payload_size = accept,
wire_mtu = mtu,
"jumbo opt-in: this endpoint advertises a jumbo QUIC receive ceiling, so the \
peer's MTU discovery can prove a jumbo path (it is capped by this value)"
);
}
}
cfg
}
/// Server endpoint with a fresh self-signed certificate (tests/dev — production hosts
/// persist an identity and use [`server_with_identity`] so clients can pin it).
pub fn server(addr: std::net::SocketAddr) -> anyhow_result::Result<quinn::Endpoint> {
@@ -238,7 +281,15 @@ pub fn client_pinned_with_identity(
.map_err(|e| anyhow_result::Error::msg(format!("quic client config: {e}")))?;
let mut client_cfg = quinn::ClientConfig::new(Arc::new(quic_cfg));
client_cfg.transport_config(stream_transport()); // keep-alive — see stream_transport
let mut ep = quinn::Endpoint::client("0.0.0.0:0".parse().unwrap())?;
// `Endpoint::client` hardcodes `EndpointConfig::default()`, whose 1472-byte
// `max_udp_payload_size` caps the HOST's MTU discovery (see `endpoint_config`), so the
// endpoint is built by hand to carry the jumbo opt-in. Same bind as before
// (`0.0.0.0:0`, v4 — no dual-stack flag to reproduce) and the same default runtime.
let socket = std::net::UdpSocket::bind("0.0.0.0:0")?;
let runtime = quinn::default_runtime()
.ok_or_else(|| anyhow_result::Error::msg("no async runtime found".into()))?;
let mut ep = quinn::Endpoint::new(endpoint_config(), None, socket, runtime)?;
ep.set_default_client_config(client_cfg);
Ok(ep)
})();
@@ -348,4 +399,80 @@ mod tests {
let _ = super::stream_transport_idle(std::time::Duration::MAX);
let _ = super::stream_transport_idle(std::time::Duration::ZERO);
}
/// Where a connection's MTU discovery is allowed to climb to, measured rather than argued
/// (PW7a). Loopback's own MTU is 64 KiB, so the ONLY thing that can stop the search here is
/// configuration — which makes this a clean instrument for the two ceilings:
///
/// * **leg A** — server opted in, client NOT: the search stalls at the client's default
/// `max_udp_payload_size` advertisement (1472) no matter how high the server's probe
/// ceiling is. This is why the shipped jumbo grow could never fire: `wire_mtu.rs` waits
/// for a settle at the sealed jumbo size and the peer's transport parameter forbids it.
/// * **leg B** — both opted in: the search reaches the sealed jumbo datagram, and the
/// elapsed time is what the `Welcome`'s bounded proof-wait has to cover.
///
/// `#[ignore]`d: it sets process-wide env (each endpoint reads the opt-in at construction,
/// which is exactly how the two legs are built) and spends seconds of wall clock.
/// Run it alone: `cargo test -p punktfunk-core --features quic mtu_discovery -- --ignored
/// --nocapture --test-threads=1`.
#[tokio::test]
#[ignore = "measurement: sets process env and takes ~15 s of wall clock"]
async fn mtu_discovery_climbs_only_as_high_as_the_peer_advertises() {
async fn climb(server_jumbo: bool, client_jumbo: bool) -> (u16, u128) {
let set = |on: bool| {
if on {
std::env::set_var("PUNKTFUNK_JUMBO", "1");
} else {
std::env::remove_var("PUNKTFUNK_JUMBO");
}
};
set(server_jumbo);
let server = endpoint::server("127.0.0.1:0".parse().unwrap()).unwrap();
let addr = server.local_addr().unwrap();
set(client_jumbo);
let client = endpoint::client_insecure().unwrap();
set(false);
let accept = tokio::spawn(async move {
let incoming = server.accept().await.expect("incoming");
let conn = incoming.await.expect("host side connects");
(server, conn)
});
let client_conn = client.connect(addr, "punktfunk").unwrap().await.unwrap();
let (_server_ep, host_conn) = accept.await.unwrap();
// A stream write gives the driver something to transmit, which is what starts the
// search (probes ride `poll_transmit`); after that each probe's ack drives the next.
let mut s = host_conn.open_uni().await.unwrap();
s.write_all(b"go").await.unwrap();
let want = crate::config::sealed_datagram_bytes(crate::config::jumbo_shard_payload_for(
9000,
std::net::IpAddr::V4(std::net::Ipv4Addr::UNSPECIFIED),
)) as u16;
let t0 = std::time::Instant::now();
let mut mtu = host_conn.stats().path.current_mtu;
while t0.elapsed() < std::time::Duration::from_secs(6) && mtu < want {
tokio::time::sleep(std::time::Duration::from_millis(5)).await;
mtu = host_conn.stats().path.current_mtu;
}
let elapsed = t0.elapsed().as_millis();
drop(client_conn);
drop(client);
(mtu, elapsed)
}
let (capped, _) = climb(true, false).await;
println!("leg A (server opted in, client not): settled at {capped} B UDP payload");
assert_eq!(
capped, 1472,
"a peer that advertises the stock max_udp_payload_size caps the search at 1472 — \
the whole point of raising it on the client endpoint"
);
let (grown, ms) = climb(true, true).await;
println!("leg B (both opted in): reached {grown} B UDP payload in {ms} ms");
assert!(
grown >= 8972,
"both sides opted in, loopback MTU is 64 KiB — discovery should reach the sealed \
jumbo datagram, got {grown}"
);
}
}
+15 -2
View File
@@ -677,8 +677,21 @@ impl Session {
/// [`Frame::part`]` = Some` while the rest is still on the wire, instead of one whole-AU
/// delivery (the slice-progressive decode path — [`crate::packet::USER_FLAG_SLICE_STREAM`]).
/// With it on, EVERY video frame delivery carries `part: Some` (a frame with no early
/// parts arrives as the degenerate `{offset: 0, first, last}` whole). Do not combine with
/// an all-intra (PyroWave) stream: its newest-wins draining assumes whole AUs.
/// parts arrives as the degenerate `{offset: 0, first, last}` whole).
///
/// **Do not combine with an all-intra (PyroWave) stream**, and the reason is sharper than
/// "newest-wins draining assumes whole AUs" (2026-08-08, PW6): the drain
/// (`client::frame_channel::FrameChannel::pop`) counts QUEUE ENTRIES and takes one entry to be
/// one AU. With parts on, a single AU pushes K entries, so `len > 1` stops meaning "the consumer
/// is behind" — the drain fires mid-AU, returns the newest entry (a SUFFIX) and clears that
/// same AU's prefixes. For PyroWave that is unrecoverable rather than lossy: the sequence
/// header lives in window 0 of every AU, so every frame would arrive headerless. Making the
/// two composable means teaching the drain to skip whole superseded AUs (never to split one)
/// — see the PW6 section of `design/linux-host-performance-wave2-pyrowave.md`.
///
/// Note this is a DIFFERENT axis from the host's streamed-AU wire
/// ([`crate::quic::VIDEO_CAP_STREAMED_AU`]): a streamed AU still completes as ONE `Frame`
/// here, so it is unaffected by any of the above.
pub fn set_deliver_frame_parts(&mut self, on: bool) {
self.reassembler.set_deliver_parts(on);
}
+5
View File
@@ -199,6 +199,11 @@ pub(crate) mod audio_probe;
#[cfg(target_os = "windows")]
#[path = "audio/windows/minted.rs"]
pub(crate) mod minted;
// The uninstall sweep over every audio devnode the two providers above (and the probe) mint —
// pub(crate) for `driver uninstall --audio`, the installer's [UninstallRun] leg.
#[cfg(target_os = "windows")]
#[path = "audio/windows/devnode_cleanup.rs"]
pub(crate) mod devnode_cleanup;
#[cfg(target_os = "windows")]
#[path = "audio/windows/wasapi_cap.rs"]
mod wasapi_cap;
@@ -406,6 +406,34 @@ fn recover_orphaned_default() {
});
}
/// [`recover_orphaned_default`]'s uninstall-time twin: same "put the operator's device back if
/// the default is still parked on ours" rule, minus the `Once` gate (the uninstaller is a fresh
/// process that runs it exactly once) — and it always drops the marker file, because there is no
/// next host run to consume it.
///
/// Why the uninstaller needs this at all: the devnode sweep that follows deletes the endpoint the
/// default may still point at. Windows would then re-pick something on its own, but it re-picks by
/// its OWN ranking, not the device the operator had before we parked it. Restoring first means
/// uninstalling gives the box back exactly the default it came with.
///
/// Returns whether a device was actually put back — the caller only logs it.
pub(crate) fn unpark_default_for_uninstall() -> bool {
let path = park_marker_path();
let Ok(s) = std::fs::read_to_string(&path) else {
return false;
};
let _ = std::fs::remove_file(&path);
let mut lines = s.lines();
let (Some(prev), Some(set)) = (lines.next(), lines.next()) else {
return false;
};
// A default the operator changed by hand since the park wins, exactly as on the recovery path.
if default_render_id().as_deref() != Some(set) {
return false;
}
set_default_endpoint(prev).is_ok()
}
/// Make `id` the default playback device for the duration of the desktop-audio capture,
/// remembering the operator's current default (in memory + the crash marker) the FIRST time so
/// [`restore_default_playback`] can put it back. Nothing is remembered when `id` already is the
@@ -42,8 +42,10 @@ use windows::Win32::System::Registry::{
};
/// Marker value in a probe devnode's `Device Parameters` key — how `cleanup` finds what this
/// devtest minted (and nothing else).
const PROBE_MARKER: &str = "PunktfunkAudioProbe";
/// devtest minted (and nothing else). pub(crate): the uninstall sweep
/// ([`devnode_cleanup`](super::devnode_cleanup)) sweeps this family too, so a devtest run on an
/// operator's box cannot outlive the product.
pub(crate) const PROBE_MARKER: &str = "PunktfunkAudioProbe";
/// DeviceDesc for probe devnodes (visible in Device Manager until the INF install renames it).
const PROBE_DESC: &str = "Punktfunk Audio Probe";
/// How long to wait for audiosrv to register a minted endpoint.
@@ -0,0 +1,217 @@
//! Uninstall-time removal of every audio device punktfunk minted on this box — the
//! `punktfunk-host driver uninstall --audio` leg the installer's Inno `[UninstallRun]` calls.
//!
//! The field report this exists for: uninstalling punktfunk left "Punktfunk Speakers",
//! "Punktfunk Microphone" and the per-pad "Wireless Controller" endpoints sitting in Windows'
//! Sound settings forever. They are not files and no uninstaller deletes them by walking a
//! payload list — they are DEVNODES this host created at runtime, and they persist exactly
//! because they are designed to ([`minted`](super::minted) and
//! [`pad_endpoint`](super::pad_endpoint) both re-resolve their devnodes across host restarts
//! rather than re-minting them). Persistent across restarts must not mean permanent.
//!
//! What gets swept: every MEDIA-class devnode carrying one of the three durable owner markers
//! this product writes into `Device Parameters`, whatever minted it —
//!
//! * [`pad_endpoint::PAD_INDEX_VALUE`](super::pad_endpoint::PAD_INDEX_VALUE) — the per-pad
//! DualSense speaker endpoints,
//! * [`minted::ROLE_MARKER`](super::minted::ROLE_MARKER) — the Speakers/Microphone substrate,
//! * [`audio_probe::PROBE_MARKER`](super::audio_probe::PROBE_MARKER) — devtest leftovers, so a
//! probe run on an operator's box cannot outlive the product either.
//!
//! Marker-matched, never name-matched: our devnodes are instances of VALVE's streaming-audio
//! drivers and are name-identical to Steam's own (the same reason the wiring plan works by
//! recorded id). Steam's devnodes, its driver packages, and a VB-CABLE from the era when we
//! bundled one all carry no marker and are therefore untouchable here — uninstalling punktfunk
//! removes what punktfunk created, and nothing else.
// Every `unsafe` block in this file carries a `// SAFETY:` proof; enforce it.
#![deny(clippy::undocumented_unsafe_blocks)]
use super::{audio_control, audio_probe, minted, pad_endpoint as pe};
use anyhow::Result;
use windows::Win32::Devices::DeviceAndDriverInstallation::SetupDiEnumDeviceInfo;
/// The `Device Parameters` REG_DWORD each punktfunk-minted devnode family stamps on itself. The
/// VALUE is what differs per family; presence of the NAME is "this one is ours", which is all a
/// sweep needs.
const OWNER_MARKERS: [&str; 3] = [
pe::PAD_INDEX_VALUE,
minted::ROLE_MARKER,
audio_probe::PROBE_MARKER,
];
/// What one sweep removed. `endpoint_records` is counted separately from `devnodes` because the
/// registry half is best-effort by design — see [`delete_endpoint_record`].
#[derive(Debug, Default, Clone, Copy, PartialEq, Eq)]
pub(crate) struct Removed {
pub devnodes: usize,
pub devnode_failures: usize,
pub endpoint_records: usize,
}
/// Restore the default playback device if we left it parked, then remove every audio devnode
/// this product minted, newest registry record and all.
///
/// Best-effort throughout, like the rest of the (un)install path: a devnode that refuses to go
/// is counted and reported, never fatal — a non-zero exit here would abort the whole uninstaller
/// over a virtual speaker.
pub(crate) fn purge() -> Result<Removed> {
// FIRST, before the sweep deletes the endpoint the default may still point at. A host that
// died mid-stream leaves the box's default playback parked on our loopback sink; Windows
// would re-pick on its own once the device vanishes, but by its own ranking rather than by
// what the operator had. Putting it back is the difference between "the box works again"
// and "the box works again, on the device it started with".
if audio_control::unpark_default_for_uninstall() {
println!("restored the default playback device this host had parked");
}
let mut out = Removed::default();
for inst in owned_devnodes()? {
// Resolve the endpoint records BEFORE the devnode goes. An endpoint's MMDevices key is
// tied to us only through its `{1}.<instance id>` devnode link — once the devnode is
// removed, nothing left in the store says the record was ever ours, and a sweep that
// guessed by NAME is exactly the mistake this module refuses to make.
let records: Vec<(&str, String)> = [
(pe::MMDEV_RENDER_PATH, pe::find_endpoint_for_devnode(&inst)),
(
pe::MMDEV_CAPTURE_PATH,
pe::find_capture_endpoint_for_devnode(&inst),
),
]
.into_iter()
.filter_map(|(path, found)| Some((path, found.ok().flatten()?)))
.collect();
if !remove_devnode(&inst) {
out.devnode_failures += 1;
// The device is still there, so its record still belongs to a live endpoint.
continue;
}
out.devnodes += 1;
for (path, endpoint) in records {
if delete_endpoint_record(path, &endpoint) {
out.endpoint_records += 1;
}
}
}
Ok(out)
}
/// Every MEDIA-class devnode carrying one of [`OWNER_MARKERS`]. Enumerated WITHOUT `DIGCF_PRESENT`
/// (that is what [`pe::media_class_devs`] gives us), so a phantom left by a crashed host is swept
/// too — the same "ghost in Device Manager forever" complaint the pad and vdisplay legs fixed.
fn owned_devnodes() -> Result<Vec<String>> {
let set = pe::media_class_devs()?;
let mut out = Vec::new();
for i in 0.. {
let mut did = pe::devinfo_data();
// SAFETY: live set; `did` is a live out-param with cbSize set.
if unsafe { SetupDiEnumDeviceInfo(set.0, i, &mut did) }.is_err() {
break; // ERROR_NO_MORE_ITEMS
}
let Some(inst) = pe::instance_id(&set, &did) else {
continue;
};
if !is_removable_instance(&inst) {
continue;
}
if OWNER_MARKERS
.iter()
.any(|m| pe::read_devparam_dword(&set, &did, m).is_some())
{
out.push(inst);
}
}
Ok(out)
}
/// A devnode this sweep is allowed to remove: ROOT-enumerated, i.e. software-created.
///
/// Every devnode we mint comes from `SetupDiCreateDeviceInfoW(… DICD_GENERATE_ID)` on the MEDIA
/// class, which always yields `ROOT\MEDIA\NNNN`. Nothing else can be ours — so if a marker name
/// we own ever collides with a value some vendor writes under a REAL sound card's `Device
/// Parameters`, this guard is what stops an uninstall from taking the user's hardware with it.
fn is_removable_instance(instance_id: &str) -> bool {
instance_id.to_ascii_uppercase().starts_with("ROOT\\")
}
/// `pnputil /remove-device` — the same teardown `audio-probe cleanup` and the driver legs use.
/// Called by absolute path: an uninstaller must not depend on the invoking shell's `%PATH%`.
fn remove_devnode(instance_id: &str) -> bool {
let windir = std::env::var("WINDIR").unwrap_or_else(|_| r"C:\Windows".into());
match std::process::Command::new(format!(r"{windir}\System32\pnputil.exe"))
.args(["/remove-device", instance_id])
.output()
{
Ok(o) if o.status.success() => {
println!("removed audio devnode {instance_id}");
true
}
Ok(o) => {
eprintln!(
"warning: pnputil could not remove {instance_id} (status {:?}): {}",
o.status.code(),
String::from_utf8_lossy(&o.stderr).trim()
);
false
}
Err(e) => {
eprintln!("warning: could not run pnputil for {instance_id}: {e}");
false
}
}
}
/// Delete one endpoint's MMDevices record — the `{guid}` subkey holding its name, its stamped
/// formats and its per-endpoint volume/settings.
///
/// BEST-EFFORT ON PURPOSE, and quiet when it fails. These keys are owned by SYSTEM and grant
/// Administrators read only (the same ACL that forces the stamping path through
/// `grant_system_full_control`), while the uninstaller runs elevated but as a USER — so on a
/// stock box this is denied and the record stays. What stays is inert: with the devnode gone the
/// endpoint is NOTPRESENT, which Sound settings surface only behind "Show Disconnected Devices",
/// and nothing re-animates it without a devnode to link to. Buying that last cosmetic scrap would
/// mean an uninstaller seizing ownership of SYSTEM-owned registry keys — a worse thing to ship
/// than the leftover. The DEVICE, which is what the field report was about, is gone either way.
fn delete_endpoint_record(reg_path: &str, endpoint_id: &str) -> bool {
use winreg::enums::{HKEY_LOCAL_MACHINE, KEY_ALL_ACCESS};
use winreg::RegKey;
let Ok(guid) = pe::endpoint_guid_part(endpoint_id) else {
return false;
};
let Ok(store) =
RegKey::predef(HKEY_LOCAL_MACHINE).open_subkey_with_flags(reg_path, KEY_ALL_ACCESS)
else {
return false;
};
store.delete_subkey_all(guid).is_ok()
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn only_root_enumerated_devnodes_are_ours() {
assert!(is_removable_instance(r"ROOT\MEDIA\0003"));
// PnP casing is not guaranteed.
assert!(is_removable_instance(r"root\media\0004"));
// A real sound card, however it got a marker-shaped value written under it.
assert!(!is_removable_instance(
r"HDAUDIO\FUNC_01&VEN_10EC&DEV_0900\4&1c4a4e5&0&0001"
));
assert!(!is_removable_instance(r"USB\VID_046D&PID_0A38\ABCDEF"));
// Not a prefix match on the string "ROOT" appearing anywhere.
assert!(!is_removable_instance(r"SWD\ROOT\MEDIA\0003"));
}
#[test]
fn every_minted_family_is_swept() {
// The sweep is only as complete as this list — a new minted-devnode family that forgets
// to register here would ship the same leak again.
assert!(OWNER_MARKERS.contains(&"PunktfunkPadIndex"));
assert!(OWNER_MARKERS.contains(&"PunktfunkAudioRole"));
assert!(OWNER_MARKERS.contains(&"PunktfunkAudioProbe"));
}
}
@@ -32,8 +32,9 @@ use std::sync::{Arc, Mutex, OnceLock};
use std::thread;
use std::time::{Duration, Instant};
/// Durable role marker in a minted devnode's `Device Parameters` key.
const ROLE_MARKER: &str = "PunktfunkAudioRole";
/// Durable role marker in a minted devnode's `Device Parameters` key. pub(crate): the uninstall
/// sweep ([`devnode_cleanup`](super::devnode_cleanup)) matches devnodes on it.
pub(crate) const ROLE_MARKER: &str = "PunktfunkAudioRole";
/// How long to wait for audiosrv to register a freshly minted endpoint.
const ENDPOINT_WAIT: Duration = Duration::from_secs(15);
/// Minimum spacing between provisioning retries once the startup attempt failed
@@ -87,13 +87,15 @@ const DEVNODE_DESC: &str = "Punktfunk Pad Audio";
/// The multi-instancing Steam Remote Play render driver we ride on.
const SSS_HWID: &str = "ROOT\\SteamStreamingSpeakers";
/// Registry value under the devnode's `Device Parameters` key persisting which pad slot the
/// devnode serves (REG_DWORD).
const PAD_INDEX_VALUE: &str = "PunktfunkPadIndex";
/// devnode serves (REG_DWORD). pub(crate): the uninstall sweep
/// ([`devnode_cleanup`](super::devnode_cleanup)) matches devnodes on it.
pub(crate) const PAD_INDEX_VALUE: &str = "PunktfunkPadIndex";
/// The endpoint store for render endpoints (each subkey = one endpoint GUID).
const MMDEV_RENDER_PATH: &str = r"SOFTWARE\Microsoft\Windows\CurrentVersion\MMDevices\Audio\Render";
pub(crate) const MMDEV_RENDER_PATH: &str =
r"SOFTWARE\Microsoft\Windows\CurrentVersion\MMDevices\Audio\Render";
/// The capture-direction sibling of [`MMDEV_RENDER_PATH`] — where a paired device's microphone
/// half registers (the `audio-probe` devtest's S3 lookup).
const MMDEV_CAPTURE_PATH: &str =
pub(crate) const MMDEV_CAPTURE_PATH: &str =
r"SOFTWARE\Microsoft\Windows\CurrentVersion\MMDevices\Audio\Capture";
/// WASAPI endpoint-id prefix for render endpoints (`{0.0.0.00000000}.{guid}`).
const ENDPOINT_ID_PREFIX: &str = "{0.0.0.00000000}.";
@@ -355,8 +357,9 @@ fn reg_registry_value(v: &StampValue) -> winreg::RegValue<'static> {
}
/// The per-endpoint GUID portion of a WASAPI endpoint id (`{0.0.0.00000000}.{guid}` →
/// `{guid}`) — the endpoint's MMDevices registry key name.
fn endpoint_guid_part(endpoint_id: &str) -> Result<&str> {
/// `{guid}`) — the endpoint's MMDevices registry key name. pub(crate): the uninstall sweep
/// deletes those keys by name.
pub(crate) fn endpoint_guid_part(endpoint_id: &str) -> Result<&str> {
endpoint_id
.rfind('{')
.map(|i| &endpoint_id[i..])
+25 -18
View File
@@ -194,27 +194,29 @@ pub fn capture_virtual_output(
crate::inject::set_stream_target(Some(target.target_id));
let pref = vout.preferred_mode;
let keep = vout.keepalive;
// The sealed-channel delivery seam: resolve the pf-vdisplay control device ONCE (it is
// process-global — a dead one is retired, kept alive — so the raw value is stable for the
// process) and wrap `send_frame_channel` in a `Send + Sync` closure the IDD-push capturer calls
// at ring attach. This is the ONE reach into `crate::vdisplay` the capturer would otherwise make;
// building it here keeps the capture→vdisplay dependency out of pf-capture (plan §W6).
// The sealed-channel delivery seam: resolve the pf-vdisplay control device ONCE and wrap
// `send_frame_channel` in a `Send + Sync` closure the IDD-push capturer calls at ring attach.
// This is the ONE reach into `crate::vdisplay` the capturer would otherwise make; building it
// here keeps the capture→vdisplay dependency out of pf-capture (plan §W6).
let control = crate::vdisplay::manager::control_device_handle().ok_or_else(|| {
anyhow::anyhow!(
"pf-vdisplay control device not open (monitor not created via the manager?)"
)
})?;
// `HANDLE` is not `Send`; capture the raw value and rebuild it inside the closure (the control
// device is never closed for the process lifetime, so the value stays valid).
let control_raw = control.0 as isize;
// Each closure keeps its own `Arc<OwnedHandle>` clone (`Send + Sync`), so the handle is open
// for exactly as long as any delivery closure lives — and CLOSES once the manager retires it
// and the last session drops, which is what lets the wake-from-sleep recovery's PnP device
// cycle proceed (an open control handle vetoes it).
let control_frame = control.clone();
let sender: pf_capture::FrameChannelSender = std::sync::Arc::new(
move |req: &pf_driver_proto::control::SetFrameChannelRequest| {
// SAFETY: `control_raw` is the pf-vdisplay control handle resolved above; it is never
// closed for the process lifetime, so reconstructing the `HANDLE` and issuing the
// `IOCTL_SET_FRAME_CHANNEL` is sound (`send_frame_channel`'s precondition).
// SAFETY: the captured `control_frame` Arc keeps the control handle open across this
// call — `send_frame_channel`'s precondition.
unsafe {
crate::vdisplay::driver::send_frame_channel(
windows::Win32::Foundation::HANDLE(control_raw as *mut core::ffi::c_void),
windows::Win32::Foundation::HANDLE(
std::os::windows::io::AsRawHandle::as_raw_handle(&*control_frame),
),
req,
)
}
@@ -231,14 +233,17 @@ pub fn capture_virtual_output(
// Cursor-forward sessions (M2c): hand the capturer the v5 cursor-channel delivery closure —
// its presence opts the session in (the capturer creates + delivers the CursorShm section,
// the driver declares the IddCx hardware cursor). Built exactly like `sender` above.
let control_cursor = control.clone();
let cursor_sender: Option<pf_capture::CursorChannelSender> = want.hw_cursor.then(|| {
std::sync::Arc::new(
move |req: &pf_driver_proto::control::SetCursorChannelRequest| {
// SAFETY: `control_raw` is the pf-vdisplay control handle resolved above; it is
// never closed for the process lifetime (`send_cursor_channel`'s precondition).
// SAFETY: the captured `control_cursor` Arc keeps the control handle open across
// this call (`send_cursor_channel`'s precondition).
unsafe {
crate::vdisplay::driver::send_cursor_channel(
windows::Win32::Foundation::HANDLE(control_raw as *mut core::ffi::c_void),
windows::Win32::Foundation::HANDLE(
std::os::windows::io::AsRawHandle::as_raw_handle(&*control_cursor),
),
req,
)
}
@@ -261,11 +266,13 @@ pub fn capture_virtual_output(
target_id,
enable: enable as u32,
};
// SAFETY: `control_raw` is the pf-vdisplay control handle resolved above; it is
// never closed for the process lifetime (`send_cursor_forward`'s precondition).
// SAFETY: the captured `control` Arc keeps the control handle open across this call
// (`send_cursor_forward`'s precondition).
unsafe {
crate::vdisplay::driver::send_cursor_forward(
windows::Win32::Foundation::HANDLE(control_raw as *mut core::ffi::c_void),
windows::Win32::Foundation::HANDLE(
std::os::windows::io::AsRawHandle::as_raw_handle(&*control),
),
&req,
)?;
}
+149
View File
@@ -29,9 +29,11 @@ mod epic;
mod gog;
#[cfg(target_os = "linux")]
mod heroic;
mod hidden;
mod launch;
#[cfg(target_os = "linux")]
mod lutris;
mod plugin_launch;
mod scanners;
mod steam;
#[cfg(windows)]
@@ -46,9 +48,11 @@ pub use epic::*;
pub use gog::*;
#[cfg(target_os = "linux")]
pub use heroic::*;
pub use hidden::*;
pub use launch::*;
#[cfg(target_os = "linux")]
pub use lutris::*;
pub use plugin_launch::*;
pub use scanners::*;
pub use steam::*;
#[cfg(windows)]
@@ -195,6 +199,32 @@ pub struct GameEntry {
pub meta: GameMeta,
}
/// A library entry plus the operator's own view of it — today, whether they hid it.
///
/// A separate type rather than a field on [`GameEntry`] for two reasons. It keeps the visibility
/// answer out of the providers entirely: a store parser has no opinion on what the operator hid, and
/// adding `hidden: false` to all eight construction sites would imply it does. More importantly it
/// makes the lane rule a TYPE guarantee instead of a discipline — `GET /library` answers
/// `Vec<GameEntry>` on every lane but the operator's, so a hidden entry cannot leak to a paired
/// client by someone forgetting a filter; there is no field there to leak.
///
/// `flatten` keeps the wire shape identical to a plain entry with one extra key, so the console
/// parses one model either way.
#[derive(Clone, Debug, Serialize, ToSchema)]
pub struct OperatorGameEntry {
#[serde(flatten)]
pub entry: GameEntry,
/// The operator hid this title ([`set_entry_hidden`]) — omitted when false, so the shape only
/// grows for entries that actually are hidden.
#[serde(skip_serializing_if = "is_not_hidden")]
pub hidden: bool,
}
/// `skip_serializing_if` predicate for [`OperatorGameEntry::hidden`] — `&bool` as serde requires.
fn is_not_hidden(hidden: &bool) -> bool {
!*hidden
}
/// A store that contributes titles to the library. The trait is the extension point for future
/// launchers; today only [`SteamProvider`] implements it.
pub trait LibraryProvider {
@@ -268,7 +298,39 @@ impl ArtKind {
/// Removing the plugin releases the claim and the built-in comes straight back.
///
/// The user-curated custom store is not a source and always contributes.
///
/// A **third** gate rides on top of these two: the operator's per-entry hides (`hidden.rs`). It is
/// applied here rather than at each call site so a hidden title is gone from every surface by
/// construction — the grid, native clients, `/applist`, and launch resolution — exactly as a
/// disabled source's titles are. [`all_games_for_operator`] is the single deliberate exception.
pub fn all_games() -> Vec<GameEntry> {
let hidden = hidden_ids();
let mut games = collect_games();
games.retain(|g| !hidden.contains(&g.id));
games
}
/// The library **including** the operator's hidden titles, each flagged.
///
/// The console's list is the only caller, and only on the operator's own lane (`GET /library`
/// branches on it): a hidden entry has to be visible SOMEWHERE or it could never be brought back.
/// Everything else — every paired client, the GameStream app list, launch resolution — goes through
/// [`all_games`] and never sees them.
pub fn all_games_for_operator() -> Vec<OperatorGameEntry> {
let hidden = hidden_ids();
collect_games()
.into_iter()
.map(|entry| OperatorGameEntry {
hidden: hidden.contains(&entry.id),
entry,
})
.collect()
}
/// Merge every enabled source + the custom entries, sorted by title — with no visibility gate of its
/// own. Split out so the two public views above cannot drift: they differ only in what they do with
/// the hidden set, never in what they collect.
fn collect_games() -> Vec<GameEntry> {
let off = disabled_scanners();
let claimed = claimed_stores();
// A built-in scanner runs when the operator hasn't disabled it AND no plugin has claimed its
@@ -314,3 +376,90 @@ pub fn all_games() -> Vec<GameEntry> {
games.sort_by_key(|g| g.title.to_lowercase());
games
}
#[cfg(test)]
mod tests {
use super::*;
fn entry(id: &str, title: &str) -> GameEntry {
GameEntry {
id: id.into(),
store: id.split_once(':').map_or("custom", |(s, _)| s).into(),
title: title.into(),
art: Artwork::default(),
role: GameRole::default(),
launch: None,
provider: None,
detect: DetectSpec::default(),
meta: GameMeta::default(),
}
}
/// The console codes against this shape, so pin it: the operator view must be a normal entry
/// with ONE extra key, and that key must vanish when the title is visible.
///
/// The skip matters beyond tidiness — it is what keeps this response byte-identical to the old
/// one for a library with nothing hidden, so shipping the feature cannot change what an existing
/// console renders until someone actually hides something.
#[test]
fn operator_entry_flattens_and_omits_hidden_when_false() {
let visible = OperatorGameEntry {
entry: entry("steam:70", "Half-Life"),
hidden: false,
};
let v = serde_json::to_value(&visible).expect("serializes");
assert_eq!(v["id"], "steam:70", "the entry's fields stay at top level");
assert_eq!(v["title"], "Half-Life");
assert!(
v.get("hidden").is_none(),
"a visible entry must not carry the key at all: {v}"
);
let hidden = OperatorGameEntry {
entry: entry("steam:70", "Half-Life"),
hidden: true,
};
let v = serde_json::to_value(&hidden).expect("serializes");
assert_eq!(v["hidden"], true);
assert_eq!(v["id"], "steam:70", "flatten still applies when hidden");
}
/// `all_games` and `all_games_for_operator` must agree on WHICH entries exist and differ only in
/// visibility — they share `collect_games` for exactly that reason. This pins the shared-source
/// property the same way the art test pins write/read symmetry: both views of an id-set built
/// from one collector, so a future edit that inlines one of them is caught.
#[test]
fn hidden_filter_is_the_only_difference_between_the_two_views() {
let games = vec![
entry("steam:70", "Half-Life"),
entry("lutris:4", "Syndicate"),
entry("custom:abc", "Chrono Trigger"),
];
let hidden: HashSet<String> = ["lutris:4".to_string()].into_iter().collect();
let operator: Vec<OperatorGameEntry> = games
.iter()
.cloned()
.map(|entry| OperatorGameEntry {
hidden: hidden.contains(&entry.id),
entry,
})
.collect();
let played: Vec<GameEntry> = games
.into_iter()
.filter(|g| !hidden.contains(&g.id))
.collect();
assert_eq!(operator.len(), 3, "the operator sees every title");
assert_eq!(played.len(), 2, "a player does not see the hidden one");
assert!(
!played.iter().any(|g| g.id == "lutris:4"),
"the hidden id must be absent, not merely flagged"
);
assert_eq!(
operator.iter().filter(|r| r.hidden).count(),
1,
"exactly the hidden one is flagged"
);
}
}
+99 -1
View File
@@ -350,8 +350,21 @@ fn sniff_image_type(bytes: &[u8]) -> Option<&'static str> {
/// write-time half of the art confinement — [`validate_art_paths`] refuses to persist a value this
/// rejects, so an out-of-root path never reaches the catalog in the first place, and
/// [`local_art_bytes`] re-checks at read time so an entry written before this existed is still safe.
///
/// A `file://` value is decoded to a plain path FIRST, exactly as [`local_art_bytes`] does. Both
/// halves of the confinement must judge the *same* string or they disagree: `Path::new` on a raw
/// `file:///home/u/c.jpg` yields a RELATIVE path whose first component is `file:`, which
/// canonicalizes against the cwd, fails, and reads as "outside every root". That is not a
/// conservative failure — it rejected every `file://` cover the plugin kit emits (`fileUrl`, the
/// documented way for a library plugin to publish local art), so the Lutris and Steam scanners
/// could not reconcile a single entry while the read path would have served those same files
/// happily.
pub fn art_path_is_servable(value: &str) -> bool {
let p = Path::new(value);
// Idempotent for the already-decoded caller: the decoded form no longer carries the prefix,
// so `local_art_bytes` passing its own output back through here is a no-op, not a second
// percent-decode of a path that legitimately contains `%`.
let value = file_url_to_path(value);
let p = Path::new(&*value);
let ext_ok = p
.extension()
.and_then(|e| e.to_str())
@@ -699,11 +712,23 @@ mod tests {
const PNG: &[u8] = &[0x89, b'P', b'N', b'G', 0x0D, 0x0A, 0x1A, 0x0A, 0, 0, 0, 13];
/// `PUNKTFUNK_LIBRARY_ART_ROOTS` is process-global while cargo runs tests as threads, so the
/// tests that repoint it must not overlap — one clearing the variable mid-flight makes the
/// other's temp root stop being a root, which fails as a confinement bug that isn't there.
/// Poisoning is recovered rather than propagated: a panic in one test should report ITS
/// failure, not cascade into an unrelated `PoisonError`.
static ART_ROOTS_LOCK: std::sync::Mutex<()> = std::sync::Mutex::new(());
fn lock_art_roots() -> std::sync::MutexGuard<'static, ()> {
ART_ROOTS_LOCK.lock().unwrap_or_else(|e| e.into_inner())
}
/// The art proxy reads bytes in the HOST process (LocalSystem on Windows) from a path the
/// plugin lane can write — so what it will and will not read IS the security boundary
/// (2026-08-05 review H-2). Confinement, extension, and content are all load-bearing.
#[test]
fn local_art_bytes_is_confined_and_image_only() {
let _guard = lock_art_roots();
let dir = std::env::temp_dir().join(format!("pf-art-test-{}", std::process::id()));
let outside = std::env::temp_dir().join(format!("pf-art-out-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
@@ -837,6 +862,79 @@ mod tests {
);
}
/// The write gate and the read gate must judge the SAME string.
///
/// Regression for 2026-08-08: `validate_art_paths` handed the raw value to `Path::new`, so a
/// `file:///…` cover became a *relative* path starting with a `file:` component, canonicalized
/// against the cwd, failed, and was refused as "outside every art root" — while
/// `local_art_bytes` decoded the very same value and served the file. Every Lutris and Steam
/// entry carrying local art was rejected with a 400 the plugin could only report as
/// `HostRequestError`, so neither scanner could sync a single game. Asserting servable and
/// readable together is the point: either alone passes with the bug present.
#[test]
fn file_url_art_is_accepted_at_write_time_exactly_as_at_read_time() {
let _guard = lock_art_roots();
let dir = std::env::temp_dir().join(format!("pf-art-wr-{}", std::process::id()));
let outside = std::env::temp_dir().join(format!("pf-art-wr-out-{}", std::process::id()));
std::fs::create_dir_all(&dir).unwrap();
std::fs::create_dir_all(&outside).unwrap();
std::env::set_var("PUNKTFUNK_LIBRARY_ART_ROOTS", &dir);
let cover = dir.join("cover.png");
std::fs::write(&cover, PNG).unwrap();
// What the kit's `fileUrl` actually emits for a Lutris/Steam cover.
let url = file_url(&cover);
assert!(
is_local_art_path(&url),
"a file:// value is local art, so the confinement applies to it"
);
assert!(
art_path_is_servable(&url),
"write time must accept the file:// form of a servable cover"
);
assert!(
validate_art_paths(&Artwork {
portrait: Some(url.clone()),
header: Some(url),
..Default::default()
})
.is_ok(),
"a real Lutris-shaped payload must reconcile"
);
// A percent-encoded name (the reason the decode exists at all) survives the round trip.
let spaced = dir.join("My Cover.png");
std::fs::write(&spaced, PNG).unwrap();
let spaced_url = file_url(&spaced).replace(' ', "%20");
assert!(
art_path_is_servable(&spaced_url),
"percent-encoded names must decode before the containment test: {spaced_url}"
);
assert!(local_art_bytes(&spaced_url).is_some(), "read time agrees");
// Loosening the write gate must not loosen the confinement: outside the root is still
// refused in file:// clothing, which is what the raw-string bug was accidentally doing.
let elsewhere = outside.join("cover.png");
std::fs::write(&elsewhere, PNG).unwrap();
assert!(
!art_path_is_servable(&file_url(&elsewhere)),
"file:// must not escape the art roots at write time either"
);
assert!(
validate_art_paths(&Artwork {
portrait: Some(file_url(&elsewhere)),
..Default::default()
})
.is_err(),
"an out-of-root file:// cover is still refused"
);
std::env::remove_var("PUNKTFUNK_LIBRARY_ART_ROOTS");
let _ = std::fs::remove_dir_all(&dir);
let _ = std::fs::remove_dir_all(&outside);
}
#[test]
fn sniff_image_type_recognizes_containers_and_rejects_secrets() {
assert_eq!(sniff_image_type(PNG), Some("image/png"));
@@ -476,6 +476,15 @@ pub fn validate_provider_payload(inputs: &[ProviderEntryInput]) -> Result<(), St
"entries[{i}]: `launch.value` for kind `xbox` must be `<Identity>!<AppId>`"
));
}
// `plugin`: the value is an opaque key in the OWNING plugin's own namespace, handed back
// to it at launch time (see `library::ask_plugin_launch`). The host never parses it, so
// the only checks are the ones that keep it loggable and bounded.
if launch.kind == "plugin" && !valid_plugin_entry_key(&launch.value) {
return Err(format!(
"entries[{i}]: `launch.value` for kind `plugin` must be 1512 chars with no \
control characters"
));
}
}
if let Some(marker) = &e.detect.env_marker {
if !valid_env_key(&marker.key) {
+141
View File
@@ -0,0 +1,141 @@
//! Per-entry visibility: the operator hides one *title*, where `scanners.rs` hides a whole source.
//!
//! **Why this is a side table and not a field on the entry.** Only manual custom entries are stored;
//! a scanner's and a plugin's titles are regenerated from scratch on every scan and every reconcile.
//! A `hidden` flag written onto one of those would be erased by the next sync — silently, and
//! minutes later, which is the worst possible shape for a setting. So the operator's choice lives
//! here, keyed by the entry's stable `<store>:<external_id>` id, and the entries stay disposable.
//!
//! That id is stable *by construction* (design D2): a claimed store's entries keep
//! `<store>:<external_id>` across reconciles no matter what the host-assigned id does, which is the
//! same property GameStream app ids and client art caches already depend on. Hiding therefore
//! survives a re-scan, a plugin restart, and the built-in→plugin migration for a store.
//!
//! Hiding is **curation, not access control** — it declutters a grid. It is applied in
//! [`all_games`](crate::library::all_games), so a hidden title is gone from every play surface
//! *including* launch resolution (the same reach a disabled scanner has), but nothing is deleted and
//! un-hiding is immediate. The console is the one surface that still sees hidden titles — otherwise
//! there would be no way to un-hide one — and only on the operator's own lane.
use super::*;
/// Persisted shape (`library-hidden.json`): the ids the operator hid. Absent file = nothing hidden.
///
/// Mirrors `library-scanners.json`'s disabled-set rather than sharing it: that file answers "which
/// SOURCES run", this one answers "which TITLES show", and a source id (`steam`) and an entry id
/// (`steam:70`) are different namespaces. Keeping them apart means neither migration can corrupt the
/// other, and an operator reading either file sees one idea.
#[derive(Debug, Default, Serialize, Deserialize)]
struct HiddenSettings {
#[serde(default)]
hidden: Vec<String>,
}
fn settings_path() -> PathBuf {
// Same hardened config dir as library.json / library-scanners.json.
pf_paths::config_dir().join("library-hidden.json")
}
/// Load the hidden set (default + non-fatal if the file is absent or malformed).
///
/// A malformed file means "nothing hidden", never "hide everything": the failure mode of a bad parse
/// must be a library that shows too much, not one that looks empty and reads as data loss.
fn load_settings() -> HiddenSettings {
match std::fs::read_to_string(settings_path()) {
Ok(raw) => serde_json::from_str(&raw).unwrap_or_else(|e| {
tracing::warn!(error = %e, "library-hidden.json malformed — nothing hidden");
HiddenSettings::default()
}),
Err(_) => HiddenSettings::default(),
}
}
fn save_settings(settings: &HiddenSettings) -> Result<()> {
let dir = pf_paths::config_dir();
pf_paths::create_private_dir(&dir).with_context(|| format!("create {}", dir.display()))?;
let json = serde_json::to_string_pretty(settings)?;
// Write-then-rename like the catalog, so a crash mid-write never truncates the settings.
let tmp = settings_path().with_extension("json.tmp");
pf_paths::write_secret_file(&tmp, json.as_bytes())
.with_context(|| format!("write {}", tmp.display()))?;
std::fs::rename(&tmp, settings_path()).context("rename library-hidden.json")?;
Ok(())
}
/// The hidden entry ids, loaded once per library read.
pub(crate) fn hidden_ids() -> HashSet<String> {
load_settings().hidden.into_iter().collect()
}
/// The store half of a library id (`steam:70` → `steam`), for the `library.changed` source.
///
/// Falls back to the whole id rather than an empty string: an id without a `:` is not a shape this
/// host produces, and naming it in the event beats emitting a blank source that matches no cache key.
fn store_of(id: &str) -> &str {
id.split_once(':').map_or(id, |(store, _)| store)
}
/// Hide or un-hide one entry. Returns whether the entry is hidden **after** the call.
///
/// Idempotent, and deliberately not validated against the current library: an entry can be absent
/// right now for reasons that have nothing to do with the operator's intent — the launcher is closed,
/// a plugin has not finished its first sync, a disk is unmounted. Refusing to hide a title that is
/// temporarily missing, or silently dropping the choice when it comes back, would both be worse than
/// storing an id that currently matches nothing. Persists and emits `library.changed` only when the
/// state actually changed, so a repeated PUT is a cheap no-op.
pub fn set_entry_hidden(id: &str, hidden: bool) -> Result<bool> {
let mut settings = load_settings();
let was_hidden = settings.hidden.iter().any(|h| h == id);
if was_hidden == hidden {
return Ok(hidden);
}
if hidden {
settings.hidden.push(id.to_string());
settings.hidden.sort();
settings.hidden.dedup();
} else {
settings.hidden.retain(|h| h != id);
}
save_settings(&settings)?;
crate::events::emit(crate::events::EventKind::LibraryChanged {
source: store_of(id).to_string(),
});
Ok(hidden)
}
#[cfg(test)]
mod tests {
use super::*;
/// The event source is the STORE, not the whole id — that is the key every client cache and the
/// console's query invalidation is grouped by.
#[test]
fn store_of_takes_the_prefix_and_tolerates_a_bare_id() {
assert_eq!(store_of("steam:70"), "steam");
assert_eq!(store_of("custom:abc"), "custom");
// An external id may itself contain a colon (Heroic's `legendary:<hash>`): split on the
// FIRST one, or the store would come back wrong for exactly the store that does this.
assert_eq!(store_of("heroic:legendary:fc0b13b7"), "heroic");
assert_eq!(store_of("weird-no-colon"), "weird-no-colon");
}
/// A malformed settings file must read as "nothing hidden". The inverse — treating a parse
/// failure as "hide everything" — would present as a library that lost its games.
#[test]
fn malformed_settings_hide_nothing() {
let s: HiddenSettings = serde_json::from_str("{ not json").unwrap_or_default();
assert!(s.hidden.is_empty());
let s: HiddenSettings = serde_json::from_str("{}").expect("an empty object is valid");
assert!(s.hidden.is_empty(), "absent key means nothing hidden");
}
/// The persisted shape is the contract an operator may hand-edit — pin it.
#[test]
fn settings_roundtrip_the_documented_shape() {
let s: HiddenSettings =
serde_json::from_str(r#"{"hidden":["steam:70","lutris:4"]}"#).expect("parses");
assert_eq!(s.hidden, vec!["steam:70", "lutris:4"]);
let json = serde_json::to_string(&s).expect("serializes");
assert_eq!(json, r#"{"hidden":["steam:70","lutris:4"]}"#);
}
}
+78 -9
View File
@@ -52,7 +52,9 @@ pub fn resolve_launch(id: &str) -> Option<LaunchTarget> {
{
// Linux runs the command itself, so a title without one has nothing to launch — same answer
// (and same warning path) as before this resolution existed.
let command = entry.launch.as_ref().and_then(command_for)?;
let command = plugin_recipe(&entry)
.map(|l| l.command)
.or_else(|| entry.launch.as_ref().and_then(command_for))?;
Some(LaunchTarget {
game,
launcher: entry.role == GameRole::Launcher,
@@ -74,9 +76,66 @@ pub fn resolve_launch(id: &str) -> Option<LaunchTarget> {
}
}
/// The recipe for a `plugin`-kind entry, asked of the plugin that owns it. `None` for every other
/// kind (without doing any I/O), so both per-OS resolvers can simply try this first.
///
/// This lives beside [`resolve_launch`] / [`launch_title`] rather than inside `command_for` /
/// `windows_launch_for` because it needs the entry's **`provider`** — and that field is the whole
/// authorization story. `provider` is stamped by the host from the reconcile URL
/// (`PUT /library/provider/{provider}`), never taken from the payload, so it is what decides which
/// plugin gets asked. A plugin that plants an entry under someone else's provider only causes that
/// *other* plugin to be asked about a key it never published — which is a 404, not a launch.
///
/// **Blocking**: see [`ask_plugin_launch`]. `resolve_launch`'s async callers hop through
/// `spawn_blocking`; the handshake probe uses [`launch_is_resolvable`], which never asks.
fn plugin_recipe(entry: &GameEntry) -> Option<PluginLaunch> {
let spec = entry.launch.as_ref()?;
if spec.kind != "plugin" {
return None;
}
let Some(provider) = entry.provider.as_deref() else {
// Only a provider reconcile can author this kind, so this is unreachable short of a
// hand-edited library.json — say so rather than silently doing nothing.
tracing::warn!(
id = %entry.id,
"plugin launch: entry carries no provider, so no plugin can answer for it"
);
return None;
};
ask_plugin_launch(provider, &spec.value)
}
/// Whether `id` will actually launch something — **without asking a plugin**.
///
/// The handshake needs this one bit to decide dedicated-session routing, and it runs on the async
/// path, so it must not make a blocking call out to a plugin. For a `plugin`-kind entry the cheap
/// answer is "a live plugin is registered under its provider, and the key is well formed"; if that
/// plugin later refuses the ask, the launch fails the same way any unresolvable entry does and the
/// player is left on the session.
#[cfg(not(windows))]
pub fn launch_is_resolvable(id: &str) -> bool {
let Some(entry) = all_games().into_iter().find(|g| g.id == id) else {
return false;
};
let Some(spec) = entry.launch.as_ref() else {
return false;
};
if spec.kind == "plugin" {
return valid_plugin_entry_key(&spec.value)
&& entry
.provider
.as_deref()
.is_some_and(|p| crate::mgmt::ui_credential(p).is_some());
}
command_for(spec).is_some()
}
/// Map a resolved [`LaunchSpec`] to its shell command (pure — the unit-testable core of
/// [`resolve_launch`], split out so the appid-validation can be tested without a Steam install).
///
/// The `plugin` kind is deliberately absent: its answer comes from another process, so it is
/// resolved by [`plugin_recipe`] before this is reached.
///
/// - `steam_appid` → `steam steam://rungameid/<appid>` (appid validated as digits).
/// - `command` → the stored command verbatim. This string comes from the host's own custom store
/// (added by the host operator via the admin UI), never from the client, so it is trusted.
@@ -126,17 +185,24 @@ fn command_for(spec: &LaunchSpec) -> Option<String> {
/// desktop and grabs foreground.
#[cfg(windows)]
pub fn launch_title(id: &str) -> Result<()> {
let spec = all_games()
let entry = all_games()
.into_iter()
.find(|g| g.id == id)
.and_then(|g| g.launch)
.filter(|g| g.launch.is_some())
.ok_or_else(|| anyhow::anyhow!("no launchable library entry '{id}'"))?;
let (cmdline, workdir) = windows_launch_for(&spec).ok_or_else(|| {
anyhow::anyhow!(
"library entry '{id}' has no Windows launch recipe (kind '{}')",
spec.kind
)
})?;
let spec = entry.launch.clone().expect("filtered to Some above");
// A `plugin` entry's recipe comes from the plugin that owns it, and arrives in the same
// (command line, working dir) shape this path already spawns. `windows_launch_for` has no arm
// for the kind, so a failed ask falls through to the "no recipe" error below.
let (cmdline, workdir) = plugin_recipe(&entry)
.map(|l| (l.command, l.cwd))
.or_else(|| windows_launch_for(&spec))
.ok_or_else(|| {
anyhow::anyhow!(
"library entry '{id}' has no Windows launch recipe (kind '{}')",
spec.kind
)
})?;
let pid = crate::interactive::spawn_in_active_session(&cmdline, workdir.as_deref())
.with_context(|| format!("launch '{id}' in the interactive session"))?;
tracing::info!(launch_id = id, %cmdline, pid, "launched library title in the interactive session");
@@ -148,6 +214,9 @@ pub fn launch_title(id: &str) -> Result<()> {
///
/// CreateProcessAsUserW does NO shell or protocol resolution, so the URI/flags are handed to a
/// concrete EXE as plain arguments — a (host-derived) URI string can never reach a command interpreter.
///
/// The `plugin` kind is deliberately absent: its answer comes from another process, so it is
/// resolved by [`plugin_recipe`] before this is reached.
#[cfg(windows)]
fn windows_launch_for(spec: &LaunchSpec) -> Option<(String, Option<std::path::PathBuf>)> {
match spec.kind.as_str() {
@@ -0,0 +1,364 @@
//! The `plugin` launch kind's transport: ask a library plugin what to run for one of **its own**
//! entries, at launch time, over the loopback UI surface it already registered.
//!
//! ## Why the host asks instead of storing a command
//!
//! A ROM tile is `<emulator> <args> <rom>` — an operator-configured command line, and the one shape
//! [`super::privileged_field`] refuses from the plugin lane (2026-08-05 review H-1). The Playnite
//! plugin hit the same wall and was rescued with a typed `playnite` kind the host resolves itself
//! (see `command_for`), but that only works because a Playnite launch is a fixed URI scheme. There
//! is no fixed scheme for "some emulator the operator installed, with the core and flags they chose"
//! — the knowledge lives in the plugin, and it is the plugin that owns the hardened quoting seam for
//! it (ROM filenames are untrusted input).
//!
//! So the entry carries an **opaque key** and nothing executable, and the command is fetched from
//! the owning plugin at the moment of an actual launch. What that buys over letting the plugin write
//! `kind = "command"` straight into the library:
//!
//! * **A stolen plugin token is no longer command execution.** Planting an entry is not enough — the
//! host asks the *live registered plugin* what to run, authenticated with the per-boot secret only
//! that process knows. A plugin asked about an entry it never published answers 404 (this is why
//! the ask names the entry rather than trusting the payload), so a forged entry launches nothing.
//! * **Nothing executable is ever persisted or served.** No command lands in `library.json`, and
//! `GET /library` has none to redact for a paired client.
//! * **No stale recipes.** The same reasoning as the `xbox` kind resolving its AUMID at launch time:
//! an emulator that moved, or a config the operator has since edited, is picked up on the next
//! launch instead of leaving an unlaunchable tile behind.
//!
//! The host still *runs* the command, because only the host can put the process where the stream can
//! see it: on Linux the line is either gamescope's own argv (a bare-spawn session nests it) or a
//! spawn carrying the session's compositor env, and the returned child is what
//! `design/session-game-lifetime.md` tracks to know the game exited. A plugin spawning the emulator
//! itself would land it outside the captured session and outside that lifetime.
use super::*;
use std::io::Read;
use std::time::Duration;
/// The whole ask, end to end. A plugin resolving one of its own entries is a local lookup against
/// state it already holds, so this is generous for a healthy plugin and short enough that a wedged
/// one cannot hold a launch — or, on the GameStream plane, the data-plane thread that calls this —
/// for longer than a player would keep staring at a tile that did nothing.
const ASK_TIMEOUT: Duration = Duration::from_secs(3);
/// A command LINE, not a script. Generous for `flatpak run … --core=… "/very/long/rom path"`,
/// bounded so a malformed answer cannot land a megabyte in the logs or in a shell argument.
const MAX_COMMAND: usize = 4096;
/// Cap the whole response body — the shape is two short strings.
const MAX_BODY: usize = 64 * 1024;
/// What a plugin answered: the command line to run, and optionally the directory to run it in
/// (emulators that resolve cores or configs relative to their install dir need one).
pub struct PluginLaunch {
pub command: String,
pub cwd: Option<PathBuf>,
}
/// The wire shape of `POST /__launch`'s response.
#[derive(Deserialize)]
struct LaunchReply {
command: String,
#[serde(default)]
cwd: Option<String>,
}
/// The opaque per-entry key a `plugin` launch carries. It is echoed to the owning plugin as JSON and
/// lands in log lines, so bound it and keep control characters out; everything else is the plugin's
/// own namespace (rom-manager uses its `<platform>/<relpath>` external id).
pub fn valid_plugin_entry_key(v: &str) -> bool {
!v.is_empty() && v.len() <= 512 && !v.chars().any(char::is_control)
}
/// Ask `plugin` what to run for its entry `key`.
///
/// `None` — the plugin is not registered/live, has no UI surface, disowns the entry, or answered
/// something unusable. Every arm logs, because from a player's seat all of them look like "the tile
/// did nothing", and the difference is exactly what an operator needs to fix it.
///
/// **Blocking** (`ureq`, the host's existing off-runtime HTTP client): callers run on a blocking
/// thread. `resolve_launch`'s async callers hop through `spawn_blocking`, and the handshake's
/// "is this launchable at all" probe uses [`super::launch_is_resolvable`], which never asks.
pub fn ask_plugin_launch(plugin: &str, key: &str) -> Option<PluginLaunch> {
if !valid_plugin_entry_key(key) {
tracing::warn!(
plugin,
"plugin launch: entry key failed validation — ignoring"
);
return None;
}
let Some(cred) = crate::mgmt::ui_credential(plugin) else {
tracing::warn!(
plugin,
entry = key,
"plugin launch: no live plugin registered under that provider id (is it running?) — \
nothing to launch"
);
return None;
};
let agent = ureq::AgentBuilder::new().timeout(ASK_TIMEOUT).build();
// Loopback + the plugin's own per-boot secret, exactly what the console proxy presents. The
// registration stores a PORT, never an address (mgmt::plugins D5), so this can only ever dial
// this machine.
// `send_string` + an explicit content type rather than `send_json`: that one needs ureq's `json`
// feature, and the body is one field.
let body = serde_json::json!({ "entry": key }).to_string();
let resp = match agent
.post(&format!("http://127.0.0.1:{}/__launch", cred.port))
.set("Authorization", &format!("Bearer {}", cred.secret))
.set("Content-Type", "application/json")
.send_string(&body)
{
Ok(r) => r,
// A plugin that does not know the entry says so with a 404 — the answer a FORGED entry gets,
// and the reason planting one is not enough to make the host run anything.
Err(ureq::Error::Status(404, _)) => {
tracing::warn!(
plugin,
entry = key,
"plugin launch: the plugin does not own an entry with that key — nothing to launch"
);
return None;
}
Err(ureq::Error::Status(code, _)) => {
tracing::warn!(
plugin,
entry = key,
code,
"plugin launch: the plugin refused to resolve the entry"
);
return None;
}
Err(e) => {
tracing::warn!(
plugin,
entry = key,
error = %e,
"plugin launch: could not reach the plugin's launch surface"
);
return None;
}
};
let mut buf = Vec::new();
if let Err(e) = resp
.into_reader()
.take((MAX_BODY + 1) as u64)
.read_to_end(&mut buf)
{
tracing::warn!(plugin, entry = key, error = %e, "plugin launch: reading the answer failed");
return None;
}
if buf.len() > MAX_BODY {
tracing::warn!(
plugin,
entry = key,
"plugin launch: answer exceeds the {MAX_BODY}-byte cap"
);
return None;
}
let reply: LaunchReply = match serde_json::from_slice(&buf) {
Ok(r) => r,
Err(e) => {
tracing::warn!(plugin, entry = key, error = %e, "plugin launch: answer was not {{command, cwd}}");
return None;
}
};
validate_reply(plugin, key, reply)
}
/// The checks on what came back, split out so they can be tested without a plugin on a port.
fn validate_reply(plugin: &str, key: &str, reply: LaunchReply) -> Option<PluginLaunch> {
let command = reply.command.trim().to_string();
if command.is_empty() {
tracing::warn!(
plugin,
entry = key,
"plugin launch: answered an empty command"
);
return None;
}
if command.len() > MAX_COMMAND {
tracing::warn!(
plugin,
entry = key,
"plugin launch: command exceeds the {MAX_COMMAND}-byte cap"
);
return None;
}
// Hygiene rather than a security boundary — a plugin that wanted two commands could always write
// `a; b`, and composing the line is its job. But a launch command is ONE line: keeping control
// characters out is what makes the logged line the line that ran, and what stops a stray `\r`
// from mangling the Windows `cmd.exe /c` form.
if command.chars().any(char::is_control) {
tracing::warn!(
plugin,
entry = key,
"plugin launch: command contains control characters — refusing it"
);
return None;
}
let cwd = match reply
.cwd
.as_deref()
.map(str::trim)
.filter(|c| !c.is_empty())
{
None => None,
Some(dir) => {
let path = PathBuf::from(dir);
// Relative to WHAT? The host's cwd is not the plugin's, and a launch that silently ran
// somewhere unintended is worse than one that says why it did not.
if !path.is_absolute() {
tracing::warn!(
plugin,
entry = key,
cwd = dir,
"plugin launch: working directory must be absolute — refusing it"
);
return None;
}
Some(path)
}
};
Some(PluginLaunch { command, cwd })
}
#[cfg(test)]
mod tests {
use super::*;
use std::io::Write;
/// A one-shot HTTP/1.1 stub on an ephemeral loopback port. Returns the port and a handle that
/// yields the raw request text — so the assertions about what the HOST sent (method, path,
/// bearer, body) live in the test thread, where a failure reads as a failure.
fn stub_plugin(status: u16, body: &'static str) -> (u16, std::thread::JoinHandle<String>) {
let listener = std::net::TcpListener::bind("127.0.0.1:0").expect("bind loopback");
let port = listener.local_addr().expect("local addr").port();
let handle = std::thread::spawn(move || {
let (mut sock, _) = listener.accept().expect("accept");
let mut buf = Vec::new();
let mut chunk = [0u8; 1024];
// Read until the body named by Content-Length has arrived (ureq always sends one here).
loop {
let n = sock.read(&mut chunk).expect("read request");
if n == 0 {
break;
}
buf.extend_from_slice(&chunk[..n]);
let text = String::from_utf8_lossy(&buf).to_string();
if let Some(end) = text.find("\r\n\r\n") {
let len = text[..end]
.lines()
.find_map(|l| {
let (k, v) = l.split_once(':')?;
k.eq_ignore_ascii_case("content-length")
.then(|| v.trim().parse::<usize>().ok())?
})
.unwrap_or(0);
if buf.len() >= end + 4 + len {
break;
}
}
}
let resp = format!(
"HTTP/1.1 {status} STATUS\r\nContent-Type: application/json\r\n\
Content-Length: {}\r\nConnection: close\r\n\r\n{body}",
body.len()
);
sock.write_all(resp.as_bytes()).expect("write response");
let _ = sock.flush();
String::from_utf8_lossy(&buf).to_string()
});
(port, handle)
}
#[test]
fn asks_the_registered_plugin_and_takes_its_answer() {
let (port, server) =
stub_plugin(200, r#"{"command":"retroarch 'smw.sfc'","cwd":"/opt/emu"}"#);
crate::mgmt::register_ui_for_test("stub-launcher", port, "s3cr3t");
let got = ask_plugin_launch("stub-launcher", "snes/smw.sfc").expect("a recipe");
assert_eq!(got.command, "retroarch 'smw.sfc'");
assert_eq!(got.cwd.as_deref(), Some(std::path::Path::new("/opt/emu")));
let req = server.join().expect("stub thread");
assert!(req.starts_with("POST /__launch "), "request was {req:?}");
// The plugin's own per-boot secret, the same credential the console proxy presents.
assert!(
req.contains("Bearer s3cr3t"),
"the ask must authenticate: {req:?}"
);
// The entry key is what the plugin resolves against its own state — it must be on the wire.
assert!(
req.contains(r#""entry":"snes/smw.sfc""#),
"body was {req:?}"
);
}
#[test]
fn a_404_means_the_plugin_disowns_the_entry() {
// The forged-entry case: planting a library row is not enough, because the plugin that would
// have to answer for it never published one.
let (port, server) = stub_plugin(404, r#"{"error":"no launchable entry \"forged\""}"#);
crate::mgmt::register_ui_for_test("stub-disowner", port, "s");
assert!(ask_plugin_launch("stub-disowner", "forged").is_none());
server.join().expect("stub thread");
}
#[test]
fn an_unregistered_provider_resolves_to_nothing() {
// No live plugin, no port to dial, no launch — and no panic.
assert!(ask_plugin_launch("no-such-plugin-is-registered", "k").is_none());
}
fn reply(command: &str, cwd: Option<&str>) -> LaunchReply {
LaunchReply {
command: command.into(),
cwd: cwd.map(str::to_string),
}
}
#[test]
fn entry_keys_are_bounded_and_printable() {
assert!(valid_plugin_entry_key("snes/Super Mario World.sfc"));
assert!(!valid_plugin_entry_key(""));
assert!(!valid_plugin_entry_key("with\nnewline"));
assert!(!valid_plugin_entry_key("with\0nul"));
assert!(!valid_plugin_entry_key(&"x".repeat(513)));
}
#[test]
fn a_usable_answer_passes_through_trimmed() {
let got = validate_reply(
"rom-manager",
"snes/smw",
reply(" retroarch 'smw.sfc' \n", None),
)
.expect("usable");
assert_eq!(got.command, "retroarch 'smw.sfc'");
assert!(got.cwd.is_none());
}
#[test]
fn empty_and_oversized_and_control_char_commands_are_refused() {
assert!(validate_reply("p", "k", reply(" ", None)).is_none());
assert!(validate_reply("p", "k", reply(&"x".repeat(MAX_COMMAND + 1), None)).is_none());
// The interesting one: a second line smuggled into what the host logs as a single command.
assert!(validate_reply("p", "k", reply("retroarch rom\nrm -rf ~", None)).is_none());
}
#[test]
fn a_working_directory_must_be_absolute() {
let abs = if cfg!(windows) { r"C:\emu" } else { "/opt/emu" };
let got = validate_reply("p", "k", reply("run", Some(abs))).expect("absolute cwd is fine");
assert_eq!(got.cwd.as_deref(), Some(std::path::Path::new(abs)));
assert!(validate_reply("p", "k", reply("run", Some("emu/cores"))).is_none());
// An empty/whitespace cwd is "no preference", not a refusal.
assert!(validate_reply("p", "k", reply("run", Some(" ")))
.expect("blank cwd is tolerated")
.cwd
.is_none());
}
}
+23 -2
View File
@@ -844,6 +844,7 @@ fn parse_spike(args: &[String]) -> Result<Options> {
let mut bitrate_mbps = 20u64;
let mut out: Option<PathBuf> = None;
let mut loopback = true;
let mut wire_chunk: Option<usize> = None;
let mut i = 0;
while i < args.len() {
@@ -890,7 +891,13 @@ fn parse_spike(args: &[String]) -> Result<Options> {
"h264" => Codec::H264,
"h265" | "hevc" => Codec::H265,
"av1" => Codec::Av1,
other => bail!("unknown --codec '{other}' (h264|h265|av1)"),
// The spike is the only way to drive a PyroWave capture→encode pass without
// a client, which is what the Linux-host PyroWave work measures against.
// Needs the `pyrowave` feature (default-on) and pairs with
// `PUNKTFUNK_ENCODER=pyrowave`, which is what puts the CAPTURE side on the
// raw-dmabuf passthrough.
"pyrowave" => Codec::PyroWave,
other => bail!("unknown --codec '{other}' (h264|h265|av1|pyrowave)"),
}
}
"--bitrate" => {
@@ -900,6 +907,12 @@ fn parse_spike(args: &[String]) -> Result<Options> {
}
"--out" => out = Some(PathBuf::from(next()?)),
"--no-loopback" => loopback = false,
"--wire-chunk" => {
let v: usize = next()?
.parse()
.map_err(|_| anyhow::anyhow!("bad --wire-chunk (bytes)"))?;
wire_chunk = (v > 0).then_some(v);
}
"-h" | "--help" => {
print_usage();
std::process::exit(0);
@@ -934,6 +947,7 @@ fn parse_spike(args: &[String]) -> Result<Options> {
bitrate_bps: bitrate_mbps.saturating_mul(1_000_000),
out,
loopback,
wire_chunk,
})
}
@@ -1007,11 +1021,18 @@ SPIKE OPTIONS:
KWin virtual output at --width x --height and captures it
--seconds <N> capture duration in seconds (default: 5)
--fps <N> target frame rate (default: 60)
--codec <h264|h265|av1> NVENC codec (default: h265)
--codec <h264|h265|av1|pyrowave>
encode codec (default: h265). 'pyrowave' also wants
PUNKTFUNK_ENCODER=pyrowave so capture takes the passthrough
--bitrate <MBPS> target bitrate in Mbps (default: 20)
--width <W> --height <H> synthetic source size (default: 1920x1080)
--out <PATH> raw Annex-B output (default: /tmp/punktfunk-spike.<ext>)
--no-loopback skip the punktfunk_core round-trip verification
--wire-chunk <BYTES> PyroWave datagram-aligned packetization at this shard payload
(a real session passes its negotiated shard_payload, e.g. 1408).
With PUNKTFUNK_PYROWAVE_STREAMED_AU=1 also armed, the AU is
drained through poll_chunk and sealed as a STREAMED wire frame
(VIDEO_CAP_STREAMED_AU), then byte-verified by the loopback
-h, --help this help
NOTES:
+9
View File
@@ -47,6 +47,14 @@ mod store;
mod tests;
mod update;
/// Lets `library::plugin_launch`'s tests put a stub plugin in the registry (test-only).
#[cfg(test)]
pub(crate) use plugins::register_ui_for_test;
/// The launch path asks a library plugin what to run for its own entries, and needs the loopback
/// credential this process already holds for it. Re-exported (rather than opening the whole
/// `plugins` module crate-wide) so these two are the ONLY things `mgmt` lends to the library side.
pub(crate) use plugins::ui_credential;
/// Default management port — adjacent to the GameStream block (47984…48010), and the same
/// number Sunshine users already associate with "the config UI".
pub const DEFAULT_PORT: u16 = 47990;
@@ -228,6 +236,7 @@ fn api_router_parts() -> (Router<Arc<MgmtState>>, utoipa::openapi::OpenApi) {
.routes(routes!(library::get_library))
.routes(routes!(library::list_library_scanners))
.routes(routes!(library::set_library_scanner))
.routes(routes!(library::set_library_entry_hidden))
.routes(routes!(library::create_custom_game))
.routes(routes!(
library::update_custom_game,
+12
View File
@@ -51,6 +51,18 @@ impl AuthLane {
pub(crate) fn may_set_privileged_fields(self) -> bool {
matches!(self, AuthLane::Admin)
}
/// Whether this is the operator's own lane — the console, as opposed to a paired client or a
/// plugin.
///
/// Same arm as [`may_set_privileged_fields`](Self::may_set_privileged_fields) today, and
/// deliberately a separate question: that one asks "may this caller cause command execution",
/// this one asks "is this caller the person curating the library". A read-only view the operator
/// alone should see (their hidden titles) is not a privilege escalation, and collapsing the two
/// would leave whichever one changes first silently answering for the other.
pub(crate) fn is_operator(self) -> bool {
matches!(self, AuthLane::Admin)
}
}
/// Auth gate on the `/api/v1` routes: a paired client cert (mTLS, from anywhere) or the bearer token
+140 -37
View File
@@ -14,33 +14,42 @@ use axum::Extension;
/// scanner plugin — while `prep` / `launch.kind = "command"` inside that payload are the operator's
/// authority alone. Route reachability and field authority are separate questions.
///
/// `Some(response)` is the refusal to return; `None` means the payload may proceed. Deliberately
/// not `Result<(), Response>`: the "error" here IS the response the handler sends, so there is no
/// error value to propagate, and a 128-byte `Response` in an `Err` variant is what
/// `Some((reason, response))` is the refusal to return; `None` means the payload may proceed.
/// Deliberately not `Result<(), Response>`: the "error" here IS the response the handler sends, so
/// there is no error value to propagate, and a 128-byte `Response` in an `Err` variant is what
/// `clippy::result_large_err` objects to.
///
/// `reason` is the caller's log line. It exists because these are TWO different refusals — an
/// operator-privileged field (403) and an unservable art path (400) — and logging both as "carries
/// a field this lane may not set" sent the Lutris/Steam `file://` art rejection looking like an
/// auth problem. The plugin only ever sees `HostRequestError`, so this log line is the sole
/// diagnosis surface for whoever has to explain why a scanner syncs nothing.
fn check_entry_fields(
lane: AuthLane,
art: &crate::library::Artwork,
launch: Option<&crate::library::LaunchSpec>,
prep: &[crate::hooks::PrepCmd],
) -> Option<Response> {
) -> Option<(String, Response)> {
if !lane.may_set_privileged_fields() {
if let Some(field) = crate::library::privileged_field(launch, prep) {
return Some(api_error(
StatusCode::FORBIDDEN,
&format!(
"`{field}` is executed as the host user and may only be set with the \
operator's admin token a plugin may publish entries with any host-resolved \
launch kind (steam_appid, steam_ui, launcher_ui, epic, gog, aumid, xbox, lutris_id, \
heroic, playnite) \
instead"
return Some((
format!("payload carries `{field}`, which this lane may not set"),
api_error(
StatusCode::FORBIDDEN,
&format!(
"`{field}` is executed as the host user and may only be set with the \
operator's admin token a plugin may publish entries with any host-resolved \
launch kind (steam_appid, steam_ui, launcher_ui, epic, gog, aumid, xbox, lutris_id, \
heroic, playnite) \
instead"
),
),
));
}
}
crate::library::validate_art_paths(art)
.err()
.map(|e| api_error(StatusCode::BAD_REQUEST, &e))
.map(|e| (e.clone(), api_error(StatusCode::BAD_REQUEST, &e)))
}
#[derive(Deserialize)]
@@ -58,6 +67,10 @@ pub(crate) struct LibraryQuery {
/// fetches directly (the public Steam CDN for Steam titles). `?provider=` narrows to the
/// entries a given external provider owns; `?platform=` to one platform (case-insensitive —
/// installed-store titles are `PC`, custom/provider entries carry whatever was authored).
///
/// **The operator's own lane additionally sees the titles they have HIDDEN**, each carrying
/// `hidden: true`; every other lane gets them filtered out upstream and cannot tell they exist. The
/// console needs them to offer "un-hide", and it is the only surface that does.
#[utoipa::path(
get,
path = "/library",
@@ -68,26 +81,28 @@ pub(crate) struct LibraryQuery {
("platform" = Option<String>, Query, description = "Only entries on this platform (case-insensitive, e.g. `PS2`)"),
),
responses(
(status = OK, description = "Unified library across all stores", body = [crate::library::GameEntry]),
(status = OK, description = "Unified library across all stores (the operator's lane also gets hidden entries, flagged)", body = [crate::library::OperatorGameEntry]),
(status = UNAUTHORIZED, description = "Missing or invalid bearer token", body = ApiError),
)
)]
pub(crate) async fn get_library(
Extension(lane): Extension<AuthLane>,
Query(q): Query<LibraryQuery>,
) -> Json<Vec<crate::library::GameEntry>> {
) -> Response {
// The operator's list is a DIFFERENT TYPE, not the same one with a flag set — which is what
// makes "a hidden title never reaches a paired client" structural rather than a filter someone
// has to remember. The redaction below is skipped here because this arm is the operator's own
// token: the command line being redacted is the one they typed.
if lane.is_operator() {
let mut rows = crate::library::all_games_for_operator();
rows.retain(|r| matches_query(&r.entry, &q));
for r in &mut rows {
crate::library::proxy_local_art(&r.entry.id, &mut r.entry.art);
}
return Json(rows).into_response();
}
let mut games = crate::library::all_games();
if let Some(provider) = q.provider.filter(|p| !p.is_empty()) {
games.retain(|g| g.provider.as_deref() == Some(provider.as_str()));
}
if let Some(platform) = q.platform.filter(|p| !p.is_empty()) {
games.retain(|g| {
g.meta
.platform
.as_deref()
.is_some_and(|p| p.eq_ignore_ascii_case(&platform))
});
}
games.retain(|g| matches_query(g, &q));
// Rewrite provider entries' local-file art into host art-proxy URLs so a client fetches covers
// from the host (a provider like Playnite stores on-host paths; the payload stays tiny at any
// library size, and the client never sees an unreachable `C:\…`).
@@ -103,16 +118,97 @@ pub(crate) async fn get_library(
// a client picks a title by ID and the host resolves the recipe itself (`resolve_launch`),
// which is the invariant that stops a client injecting a command in the first place. The
// `kind` stays, so "this is launchable, and how" still renders.
if !lane.may_set_privileged_fields() {
for g in &mut games {
if let Some(l) = g.launch.as_mut() {
if l.kind == "command" {
l.value.clear();
}
//
// Unconditional now: the operator's lane returned above, so reaching here IS "some lane but
// theirs". Leaving the old `if !lane.may_set_privileged_fields()` would read as though an
// unredacted path still existed here, and would quietly stop redacting if that early return
// ever moved.
for g in &mut games {
if let Some(l) = g.launch.as_mut() {
if l.kind == "command" {
l.value.clear();
}
}
}
Json(games)
Json(games).into_response()
}
/// The `?provider=` / `?platform=` narrowing, shared by both lane arms so they cannot drift.
fn matches_query(g: &crate::library::GameEntry, q: &LibraryQuery) -> bool {
if let Some(provider) = q.provider.as_deref().filter(|p| !p.is_empty()) {
if g.provider.as_deref() != Some(provider) {
return false;
}
}
if let Some(platform) = q.platform.as_deref().filter(|p| !p.is_empty()) {
if !g
.meta
.platform
.as_deref()
.is_some_and(|p| p.eq_ignore_ascii_case(platform))
{
return false;
}
}
true
}
/// Request body for `setLibraryEntryHidden`.
#[derive(Deserialize, ToSchema)]
pub(crate) struct HiddenToggle {
/// Whether this title should be hidden from every play surface.
hidden: bool,
}
/// What `setLibraryEntryHidden` echoes back.
#[derive(Serialize, ToSchema)]
pub(crate) struct HiddenState {
/// The entry id the call addressed.
id: String,
/// Its visibility after the call.
hidden: bool,
}
/// Hide or un-hide one library title
///
/// Curation, not access control: a hidden title disappears from every play surface — the console
/// grid on a client, native clients, the GameStream app list, and launch resolution — while nothing
/// is deleted and un-hiding restores it immediately. The operator's own console still lists it
/// (flagged `hidden`) so it can be brought back.
///
/// Keyed by the entry's stable `<store>:<external_id>` id, which survives re-scans and reconciles by
/// construction (D2). The id is **not** validated against the current library on purpose: a title
/// can be legitimately absent at this moment (launcher closed, plugin mid-sync, drive unmounted),
/// and refusing the operator's choice in that window would be worse than storing an id that
/// currently matches nothing. Emits `library.changed` (source = the store) only on a real change.
#[utoipa::path(
put,
path = "/library/hidden/{id}",
tag = "library",
operation_id = "setLibraryEntryHidden",
params(("id" = String, Path, description = "The library entry id (e.g. `steam:70`)")),
request_body = HiddenToggle,
responses(
(status = OK, description = "Stored; the entry's visibility after the call", body = HiddenState),
(status = BAD_REQUEST, description = "Empty entry id", body = ApiError),
(status = UNAUTHORIZED, description = "Missing or invalid bearer token", body = ApiError),
(status = INTERNAL_SERVER_ERROR, description = "Could not persist the settings", body = ApiError),
)
)]
pub(crate) async fn set_library_entry_hidden(
Path(id): Path<String>,
ApiJson(toggle): ApiJson<HiddenToggle>,
) -> Response {
if id.trim().is_empty() {
return api_error(StatusCode::BAD_REQUEST, "entry id must not be empty");
}
match crate::library::set_entry_hidden(&id, toggle.hidden) {
Ok(hidden) => {
tracing::info!(entry = %id, hidden, "management API: library entry visibility set");
Json(HiddenState { id, hidden }).into_response()
}
Err(e) => api_error(StatusCode::INTERNAL_SERVER_ERROR, &e.to_string()),
}
}
/// Request body for `setLibraryScanner`.
@@ -205,7 +301,9 @@ pub(crate) async fn create_custom_game(
if input.title.trim().is_empty() {
return api_error(StatusCode::BAD_REQUEST, "title must not be empty");
}
if let Some(denied) = check_entry_fields(lane, &input.art, input.launch.as_ref(), &input.prep) {
if let Some((_, denied)) =
check_entry_fields(lane, &input.art, input.launch.as_ref(), &input.prep)
{
return denied;
}
match crate::library::add_custom(input) {
@@ -238,7 +336,9 @@ pub(crate) async fn update_custom_game(
if input.title.trim().is_empty() {
return api_error(StatusCode::BAD_REQUEST, "title must not be empty");
}
if let Some(denied) = check_entry_fields(lane, &input.art, input.launch.as_ref(), &input.prep) {
if let Some((_, denied)) =
check_entry_fields(lane, &input.art, input.launch.as_ref(), &input.prep)
{
return denied;
}
use crate::library::MutateOutcome;
@@ -364,11 +464,14 @@ pub(crate) async fn reconcile_provider_entries(
// Every entry in the payload, not just the first — a reconcile replaces a whole entry set, so
// one privileged field anywhere in it is one command execution.
for (i, e) in inputs.iter().enumerate() {
if let Some(denied) = check_entry_fields(lane, &e.art, e.launch.as_ref(), &e.prep) {
if let Some((reason, denied)) = check_entry_fields(lane, &e.art, e.launch.as_ref(), &e.prep)
{
tracing::warn!(
provider,
index = i,
"library reconcile refused: payload carries a field this lane may not set"
title = %e.title,
reason = %reason,
"library reconcile refused"
);
return denied;
}
+32
View File
@@ -286,6 +286,38 @@ pub(crate) fn live_plugin_ids() -> Vec<String> {
registry().live_ids()
}
/// The loopback `{port, secret}` a live plugin serves its UI on — the credential the **host itself**
/// presents when it asks a library plugin what to run for one of its `plugin`-kind launch entries
/// ([`crate::library::ask_plugin_launch`]).
///
/// The same lookup the console proxy gets from `GET /plugins/{id}/ui-credential`, exposed in-process
/// so the launch path never round-trips through the management API to reach a port this process
/// already holds. `None` for an unknown, expired, or UI-less plugin — which the launch path reports
/// as "no recipe", exactly like any other unresolvable entry.
pub(crate) fn ui_credential(id: &str) -> Option<UiCredential> {
registry().credential(id)
}
/// Put a live UI registration in the registry directly — **test only**, so the launch path
/// ([`crate::library::ask_plugin_launch`]) can be driven against a stub server without standing up
/// the whole management router just to reach `PUT /plugins/{id}`.
#[cfg(test)]
pub(crate) fn register_ui_for_test(id: &str, port: u16, secret: &str) {
registry().upsert(
id,
Valid {
title: id.to_string(),
version: None,
ui: Some(StoredUi {
port,
secret: secret.to_string(),
icon: None,
}),
category: None,
},
);
}
// ---------------------------------------------------------------- validation
/// A plugin id: `definePlugin`'s kebab-case name (`^[a-z][a-z0-9-]*$`, ≤64) — the same regex the SDK
+38
View File
@@ -1198,6 +1198,10 @@ fn every_route_is_classified_for_the_plugin_and_cert_lanes() {
("GET", "/api/v1/library/art/{id}/{kind}", true, true),
("GET", "/api/v1/library/scanners", true, false),
("PUT", "/api/v1/library/scanners/{id}", true, false),
// Hiding a title is the OPERATOR curating their own library: a plugin has no business
// deciding what the operator sees, and a paired client must not be able to hide a game on
// the host it is streaming from. Neither lane, unlike the scanner toggle above.
("PUT", "/api/v1/library/hidden/{id}", false, false),
("POST", "/api/v1/library/custom", true, false),
("PUT", "/api/v1/library/custom/{id}", true, false),
("DELETE", "/api/v1/library/custom/{id}", true, false),
@@ -2048,6 +2052,40 @@ async fn library_scanner_list_and_unknown_toggle() {
);
}
/// A library id is `<store>:<external_id>`, so the hide route's path segment CONTAINS A COLON —
/// and for Heroic (`heroic:legendary:<hash>`) it contains two.
///
/// This is the one thing about the endpoint that could be silently wrong: if the router did not
/// match, or split on the colon, the console's hide button would 404 against an id the host itself
/// produced. Asserting "not 404" is the whole point, so the body is deliberately INVALID — that
/// stops at the JSON layer with a 4xx and never reaches the handler, which would otherwise write
/// `library-hidden.json` into the developer's real config dir (the same reason the toggle test
/// above only exercises its rejection path).
#[tokio::test]
async fn hide_route_matches_ids_containing_colons() {
let app = test_app(test_state(), None);
let put = |id: &str| {
axum::http::Request::put(format!("/api/v1/library/hidden/{id}"))
.header(axum::http::header::CONTENT_TYPE, "application/json")
// Not a `HiddenToggle` — rejected before the handler runs.
.body(Body::from(serde_json::json!({"nope": 1}).to_string()))
.unwrap()
};
for id in ["steam:70", "custom:abc", "heroic:legendary:fc0b13b7"] {
let (s, json) = send(&app, put(id)).await;
assert_ne!(
s,
StatusCode::NOT_FOUND,
"`{id}` must ROUTE — a colon is a legal path character and every library id has one: {json}"
);
assert!(
s.is_client_error(),
"a body that is not a HiddenToggle must be refused, not accepted: {s} {json}"
);
}
}
// ------------------------------------------------------------------ library providers
/// Provider reconcile validation (the write path itself is unit-tested in `library::custom`
+20 -7
View File
@@ -1148,7 +1148,12 @@ async fn serve_session(
// path verdict (WARN + learned clamp for the next session on a constrained path; clears
// a stale clamp on a healthy one) — and, with the driver above, heal or grow THIS
// session mid-stream. Bounded ~10 s task unless a jumbo grow leaves it as revert guard.
wire_mtu::spawn_watch(conn.clone(), welcome.shard_payload as usize, shard_reneg);
wire_mtu::spawn_watch(
conn.clone(),
welcome.shard_payload as usize,
hello.max_shard_payload,
shard_reneg,
);
// Negotiated cursor forwarding: the HOST_CAP_CURSOR bit the Welcome advertised, read back
// rather than recomputed (`handshake::cursor_forward` computed it once, with the encoder
// blend-capability gate — re-running it here could drift, and would re-probe).
@@ -1507,11 +1512,17 @@ async fn serve_session(
// launcher's on-disk metadata, and the data plane needs three things out of it — what to run, what
// to call the title, and how to recognize its process once a launcher has handed off
// (design/session-game-lifetime.md §4).
let launch_target =
hello
.launch
.as_deref()
.and_then(|id| match crate::library::resolve_launch(id) {
//
// On a blocking thread: a `plugin`-kind entry resolves by asking the plugin that owns it over
// loopback (`library::ask_plugin_launch`), and this is an async context.
let launch_target = match hello.launch.as_deref() {
None => None,
Some(id) => {
let owned = id.to_string();
match tokio::task::spawn_blocking(move || crate::library::resolve_launch(&owned))
.await
.context("resolve the session's library launch")?
{
Some(t) => {
tracing::info!(
launch_id = id,
@@ -1528,7 +1539,9 @@ async fn serve_session(
);
None
}
});
}
}
};
#[cfg(target_os = "windows")]
let launch_for_dp = launch_target.as_ref().and(hello.launch.clone());
#[cfg(not(target_os = "windows"))]
+134 -25
View File
@@ -33,6 +33,52 @@ fn pick_compositor(
}
}
/// Is this connect pinned at a compositor that is not actually running?
///
/// Pure (the I/O shell passes in the observed liveness) so the interaction is unit-tested, because
/// it is invisible from the outside: an operator pin puts its backend into
/// [`crate::vdisplay::available`] unconditionally AND skips `apply_session_env`'s
/// `XDG_CURRENT_DESKTOP` scrub, so [`pick_compositor`] hands back a compositor that may be a corpse
/// and its `None` (recover) arm can never fire. [`Compositor::Gamescope`] is exempt — it stands its
/// own session up, which is the whole reason a headless box pins it.
#[cfg(not(target_os = "windows"))]
fn pinned_at_a_dead_session(
overridden: bool,
chosen: crate::vdisplay::Compositor,
live: crate::vdisplay::ActiveKind,
) -> bool {
overridden && chosen.needs_live_session() && live == crate::vdisplay::ActiveKind::None
}
/// The handshake error for "no graphical session is live for this uid" — the state a compositor
/// crash leaves behind (gnome-shell SIGSEGV → GDM greeter, whose auto-login is once-per-boot, so the
/// box would otherwise need a walk-up or a reboot).
///
/// Fires the operator's recovery hook (debounced) on the way out when one is configured, so the
/// client's retry a few seconds later lands in a recovered desktop. `pinned` names the
/// `PUNKTFUNK_COMPOSITOR` value when the pin is what got us here, so the message can say which knob
/// to change rather than the generic advice to *set* the knob that caused it.
#[cfg(not(target_os = "windows"))]
fn no_live_session(pinned: Option<&str>) -> anyhow::Error {
if crate::vdisplay::try_recover_session() {
return anyhow::anyhow!(
"no live graphical session for this uid — host session recovery launched \
(PUNKTFUNK_RECOVER_SESSION_CMD); retry in a few seconds"
);
}
match pinned {
Some(pin) => anyhow::anyhow!(
"PUNKTFUNK_COMPOSITOR={pin} pins this host to a backend that can only attach to an \
already-running compositor, and no graphical session is live for this uid start a \
session, pin `gamescope` (it stands its own up), or set PUNKTFUNK_RECOVER_SESSION_CMD"
),
None => anyhow::anyhow!(
"no usable compositor (no live graphical session for this uid; set \
PUNKTFUNK_COMPOSITOR or start a desktop/gaming session)"
),
}
}
/// Resolve the client's compositor preference to a concrete backend (the I/O shell around
/// [`pick_compositor`]): enumerate what's available, auto-detect the default, pick, and log
/// whether the explicit request was honored or fell back. Runs blocking probes — call off the
@@ -61,13 +107,21 @@ pub(super) fn resolve_compositor(
// Explicit operator override (legacy / CI / forcing a backend for a test) wins and is assumed
// to come with a hand-set env — don't retarget the process env in that case.
let overridden = pf_host_config::config().compositor.is_some();
// Liveness is read on BOTH paths. The auto path retargets the process env at the live
// session (below); the PINNED path needs it too, because a pin names a BACKEND, not a
// running session — and a pin whose compositor has died used to be indistinguishable from a
// healthy one here (it skips `apply_session_env`'s `XDG_CURRENT_DESKTOP` scrub and lands
// itself in `available()`, so `pick_compositor` could never return `None`). That combination
// marched every client through 8 doomed `create` retries and left the operator's
// `PUNKTFUNK_RECOVER_SESSION_CMD` unreachable — see the `needs_live_session` gate below.
let active = crate::vdisplay::detect_active_session();
let detected = if overridden {
crate::vdisplay::detect().ok()
} else {
// Auto: detect the LIVE session (Gaming vs Desktop) and retarget the process env at it so
// every backend (video capture + input) this connect opens against the active session —
// this is the state machine that lets one host follow a Bazzite box across Gaming↔Desktop.
let active = crate::vdisplay::detect_active_session();
//
// A4: if the compositor instance changed since the last connect (an idle-time Game↔Desktop
// switch), bump the epoch + invalidate the old backend's kept displays so this connect never
// reuses a node id from the dead instance.
@@ -84,14 +138,36 @@ pub(super) fn resolve_compositor(
// under `game_session=dedicated` (gamescope confirmed available) forces its OWN headless
// gamescope spawn at the client's mode, overriding the detected desktop/game-mode backend. The
// env was already retargeted above (for XDG_RUNTIME_DIR / the PipeWire daemon); we just pin the
// backend + input to the spawn sub-mode. Skipped under an explicit operator compositor pin.
if dedicated_launch && !overridden {
let route = crate::vdisplay::apply_input_env(Compositor::Gamescope, true);
tracing::info!(
?route,
"dedicated game session — routing to a headless gamescope spawn at the client mode"
);
return Ok((Compositor::Gamescope, route));
// backend + input to the spawn sub-mode. An explicit operator compositor pin still outranks
// it — but says so out loud (below), because a silent veto is indistinguishable from the
// feature being broken.
if dedicated_launch {
if overridden {
// The pin still wins (it is the operator's explicit, hand-configured knob), but it
// must NEVER win silently: the console goes on displaying `game_session=dedicated`
// while every launch lands in the pinned session instead, and nothing in the log
// connects the two. That cost a full triage on a box whose `PUNKTFUNK_COMPOSITOR`
// was a forgotten validation leftover — the setting had never once taken effect and
// the only evidence was the ABSENCE of the info! line below.
tracing::warn!(
pin = pf_host_config::config()
.compositor
.as_deref()
.unwrap_or("-"),
"game_session=dedicated asked for this launch's OWN headless gamescope, but \
PUNKTFUNK_COMPOSITOR pins this host to a backend the operator pin wins and \
the game launches into the pinned session instead. Unset PUNKTFUNK_COMPOSITOR \
to get dedicated game sessions."
);
} else {
let route = crate::vdisplay::apply_input_env(Compositor::Gamescope, true);
tracing::info!(
?route,
"dedicated game session — routing to a headless gamescope spawn at the client \
mode"
);
return Ok((Compositor::Gamescope, route));
}
}
let available = crate::vdisplay::available();
let chosen = match pick_compositor(pref, &available, detected) {
@@ -112,23 +188,18 @@ pub(super) fn resolve_compositor(
);
Compositor::Gamescope
}
None => {
// The state a compositor crash leaves behind (gnome-shell
// SIGSEGV → GDM greeter, whose auto-login is once-per-boot). If the operator
// configured a recovery hook, fire it (debounced) and tell the client to retry:
// its next knock lands in the recovered desktop.
if crate::vdisplay::try_recover_session() {
anyhow::bail!(
"no live graphical session for this uid — host session recovery launched \
(PUNKTFUNK_RECOVER_SESSION_CMD); retry in a few seconds"
);
}
anyhow::bail!(
"no usable compositor (no live graphical session for this uid; set \
PUNKTFUNK_COMPOSITOR or start a desktop/gaming session)"
);
}
None => return Err(no_live_session(None)),
};
// Same dead-session exit, reached the other way: a pin puts its backend in `available()`
// unconditionally, so `pick_compositor` above can hand back a compositor that is not
// actually running and the `None` arm never fires. Check the backend's own requirement
// against observed liveness instead of trusting the pin. Gamescope is exempt — it stands
// its own session up, which is the whole point of pinning it on a headless box.
if pinned_at_a_dead_session(overridden, chosen, active.kind) {
return Err(no_live_session(
pf_host_config::config().compositor.as_deref(),
));
}
// Point input at the same backend and resolve the gamescope sub-mode (managed where the
// session infra exists, attach to a foreign gamescope, else per-session bare spawn). The
// route travels back to the caller as a VALUE and is carried on the backend instance — an
@@ -170,6 +241,44 @@ mod tests {
use super::pick_compositor;
use punktfunk_core::config::CompositorPref;
/// A pin at a compositor that ISN'T RUNNING must take the recovery exit rather than march the
/// client into a bring-up that can only fail.
///
/// The regression this pins down: `PUNKTFUNK_COMPOSITOR=mutter` on a box whose gnome-shell had
/// segfaulted. The pin put Mutter in `available()` and suppressed the `XDG_CURRENT_DESKTOP`
/// scrub, so every connect "resolved" happily and then spent 8 retries on
/// `RemoteDesktop.CreateSession: ServiceUnknown` — while the operator's
/// `PUNKTFUNK_RECOVER_SESSION_CMD` sat unreachable behind a `None` arm that could never fire.
#[cfg(not(target_os = "windows"))]
#[test]
fn a_pin_at_a_dead_session_recovers_instead_of_retrying() {
use super::pinned_at_a_dead_session as dead;
use crate::vdisplay::{ActiveKind, Compositor::*};
// The bug: pinned to a desktop backend with nothing live for this uid.
assert!(dead(true, Mutter, ActiveKind::None));
assert!(dead(true, Kwin, ActiveKind::None));
assert!(dead(true, Wlroots, ActiveKind::None));
assert!(dead(true, Hyprland, ActiveKind::None));
// Pinned but the session IS up — the ordinary case, must not bail.
assert!(!dead(true, Mutter, ActiveKind::DesktopGnome));
// Gamescope stands its own session up from nothing: pinning it on a headless box is a
// SUPPORTED setup, not a dead session. (This is the .21 no-login workaround — never break it.)
assert!(!dead(true, Gamescope, ActiveKind::None));
// Unpinned is untouched: the auto path already reaches `pick_compositor`'s `None` arm via
// `compositor_for_kind(ActiveKind::None)`, and it owns the managed-takeover case.
assert!(!dead(false, Mutter, ActiveKind::None));
}
/// gamescope is the ONLY backend that can serve a connect with no session already running.
#[test]
fn only_gamescope_survives_a_dead_session() {
use crate::vdisplay::Compositor::*;
assert!(!Gamescope.needs_live_session());
for c in [Mutter, Kwin, Wlroots, Hyprland] {
assert!(c.needs_live_session(), "{c:?} needs a live compositor");
}
}
#[test]
fn compositor_resolution_precedence() {
use crate::vdisplay::Compositor::*;

Some files were not shown because too many files have changed in this diff Show More