662df795b2aa8306ea427ee369700a6be6a754ed
365
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
55dbb14cf4 |
docs(windows): the signing pages still told users to import a certificate we no longer publish
ci / bun-nix (pull_request) Successful in 20s
ci / web (pull_request) Successful in 1m2s
ci / rust-arm64 (pull_request) Successful in 1m22s
windows-client / client (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (pull_request) Failing after 1m47s
windows-client / client (x64, , x86_64-pc-windows-msvc, C:\t) (pull_request) Failing after 1m53s
ci / docs-site (pull_request) Successful in 8m3s
ci / rust (pull_request) Successful in 27m30s
Azure Artifact Signing chains to a public root, so neither the host installer nor the client MSIX ships a .cer any more and users import nothing. install.md, install-client.md and windows-host.md still walked through importing one — and because the pre-Azure files are still sitting in the package registry those URLs return 200 rather than 404, so following the docs didn't fail loudly, it quietly planted a retired self-signed cert in machine Root and TrustedPublisher. Verified on glass against the 0.29 canary while checking the release: both artifacts verify Valid, timestamped, and publicly trusted, and the MSIX installs with nothing imported. Also corrected while here: the MSIX publisher change makes a different package identity, so installs from 0.28.1 or earlier need an uninstall rather than an upgrade (and a packaged app's settings go with it, so the client pairs again); a silent UPGRADE reuses the task selection the previous install recorded instead of the wizard defaults, so a once-declined installgamepad silently keeps stale gamepad drivers; installaudiocable stopped being a task name in 4a621de6; and Add-AppxPackage from a non-interactive session can fail 0x80070005 when the Windows App Runtime it depends on is in use. |
||
|
|
daabb85373 |
Merge pull request 'The Linux data-plane renice was a silent no-op on every install — RealtimeKit fallback, audio threads boosted at all, nice-limit headroom on every channel' (#232) from worktree-thread-qos-rtkit into main
audit / bun-audit (sdk) (push) Successful in 36s
audit / bun-audit (web) (push) Successful in 14s
audit / docs-site-audit (push) Successful in 26s
audit / pnpm-audit (push) Successful in 18s
audit / cargo-audit (push) Successful in 2m26s
apple / swift (push) Successful in 1m56s
audit / bun-audit (plugin-kit) (push) Successful in 3m33s
ci / rust-arm64 (push) Failing after 2m12s
audit / miri (push) Successful in 5m14s
android / android (push) Successful in 8m30s
ci / bun-nix (push) Successful in 26s
ci / docs-site (push) Successful in 1m35s
audit / license-gate (push) Successful in 8m7s
audit / c-abi-asan (push) Successful in 8m29s
ci / web (push) Successful in 6m5s
apple / distribute (push) Successful in 10m58s
apple / screenshots (push) Successful in 8m44s
windows-host / package (push) Successful in 13m15s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 31s
ci / rust (push) Canceled after 29m44s
windows-client / client (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Failing after 1m49s
windows-client / client (x64, , x86_64-pc-windows-msvc, C:\t) (push) Failing after 1m49s
flatpak / build-publish (push) Successful in 33m23s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 8s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / builders (ci/gamescope-trixie.Dockerfile, punktfunk-gamescope-trixie) (push) Successful in 8s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 27s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 29s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 31s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 31s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 25s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 22s
docker / builders-arm64cross (push) Successful in 11s
docker / deploy-docs (push) Successful in 38s
deb / build-publish-gamescope (push) Successful in 1m3s
deb / build-publish (push) Successful in 4m9s
deb / build-publish-client-arm64 (push) Successful in 6m0s
arch / build-publish (push) Successful in 8m3s
deb / build-publish-host (push) Successful in 8m35s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 8m57s
deb / smoke-install (push) Canceled after 1m17s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 18m5s
nix / flake (push) Failing after 26m33s
|
||
|
|
52df9c59af |
pkg(linux): nice-limit headroom on every channel, so the renice also works without rtkit
A new shared drop-in, packaging/linux/50-punktfunk-nice.conf (user@.service.d, LimitNICE=-15), raises the user-session nice hard limit so the direct setpriority() path works on rtkit-less boxes — a limit, not a grant, effective from the next login. Shipped by rpm (%files + install, flows into the Bazzite sysext via rpm2cpio), Arch, and deb; the Steam Deck installer writes it to /etc/systemd/system/user@.service.d instead (SteamOS /usr is read-only), following its existing sudo-to-/etc pattern. rpm and deb gain a weak Recommends: rtkit and Arch an optdepends hint — with rtkit the fix needs no relogin at all. The NixOS module instead sets security.rtkit.enable = mkDefault true (rtkit is not a given there; mkDefault keeps it operator-overridable). It remains true on every channel that the host binary must never carry a file capability — the spec's no-caps note now names the two fallback rungs instead of calling the thread nice a best-effort no-op. |
||
|
|
bb78117504 |
feat(host): moving the management port off 47990 now survives, and the console follows
47990 is the management API's port and also Sunshine's (and Apollo's, and Vibeshine's) web UI port. With the GameStream planes off it is the ONLY port the two still share, so moving it is the whole of what "run both on one box" needs — except moving it was barely possible: * `--mgmt-bind` was the sole route, and it lives in a unit file / service registration that a package upgrade rewrites. There was no `host.env` key, so the change did not survive. * The literal 47990 appeared in SIX places — mgmt::DEFAULT_PORT, the Windows service's console launch, scripts/punktfunk-web.service, the NixOS module, web/web-run.cmd, and the console's own default. Nothing downstream could learn a different port, so moving the listener silently left the console proxying to a port nothing was listening on. Now there is one source of truth. `PUNKTFUNK_MGMT_BIND` joins `host.env` (the `--gamestream` / PUNKTFUNK_GAMESTREAM shape: either source works, the CLI flag wins), and `serve` publishes the port it ACTUALLY bound to ~/.config/punktfunk/mgmt-endpoint, in the same KEY=VALUE form mgmt-token already uses so it is sourceable as a systemd EnvironmentFile and readable by the Windows service's existing read_env_file_value. Every consumer derives from that; the 47990 literals survive only as the fallback that keeps an OLD host working with a NEW console. The two unit files drop their hardcoded `Environment=PUNKTFUNK_MGMT_URL=` rather than layering a default beneath the file: whether Environment= or EnvironmentFile= wins is a directive-ordering question, and the hand-written unit and the Nix-generated one do not order the same way. No default, no precedence puzzle — the server's own built-in fallback covers a host that never wrote the file. Two robustness details worth naming, because both fail in the same direction: * mgmt-endpoint is written write-then-rename. A torn read would set PUNKTFUNK_MGMT_URL to EMPTY, which is worse than a missing file — a built-in default only rescues an *unset* variable. * mgmtUrl() now treats blank as unset, which `??` alone does not. The publish happens in parse_serve next to the token persistence, so both files appear together; the console's unit gates on mgmt-token, and its Restart=always picks up a lost race anyway. What this does NOT change: a lost 47990 bind is still fatal to the whole host (the bind sits in tokio::try_join! with the native plane), and running two Moonlight-compatible hosts at once is still unsupported — on Windows the exclusive display topology is a second, independent conflict. Both are documented rather than altered. Verified on Linux in punktfunk-rust-ci (amd64): cargo check --all-targets clean for punktfunk-host and pf-host-config with the "Checking punktfunk-host" marker confirmed present (a first run exited 0 having compiled nothing — the warm shared target dir judged it fresh), 40/40 mgmt tests pass including the new one pinning the published line against both parsers that consume it. Console: tsc --noEmit clean, bun test server/ 9/9. cargo fmt --all --check clean. |
||
|
|
b79ff45bd1 |
feat(windows): sign via Azure Artifact Signing — a 3-day leaf makes timestamping mandatory
Releases move from the self-signed CN=unom cert to Azure Artifact Signing (formerly Trusted Signing): account `unomsigning`, profile `unom-io`, signed by the `punktfunk-ci-signing` service principal, which holds only the Artifact Signing Certificate Profile Signer role scoped to that one profile. Both pack scripts gain the backend ahead of the existing .pfx and ephemeral fallbacks, so canary and fork builds are unaffected. Three things that are easy to get wrong, and are handled here rather than discovered in the field: Azure mints a leaf certificate per signing request that expires in about three days. Both scripts previously retried WITHOUT a timestamp when a timestamped sign failed — under Azure that ships an artifact which verifies on the runner and goes untrusted days later, on every user's machine at once. The retry is now gated on the mode: still lenient for a .pfx whose cert outlives the release, a hard failure for Azure. The MSIX manifest Publisher must equal the signer subject byte-for-byte, because package identity is Name + Publisher. The default is now the profile's verified subject, written with `[char]0xFC` escapes rather than literal umlauts so this UTF-8-without-BOM file cannot silently mojibake the DN into one that no longer matches. pack-msix.ps1 now also reads the signature back off the packed .msix and fails on drift — asymmetric on purpose: a subject that disagrees is fatal, a subject that cannot be read is only a warning, since Get-AuthenticodeSignature's .msix support varies by Windows version and signtool has already reported success by then. NOTE this changes package identity, so existing installs need an uninstall, not an upgrade. The updater's leaf-pinning note was wrong and is corrected: update/windows.rs claimed the AUTHENTICODE_SHA256 field made Trusted Signing "a manifest edit", but a per-request leaf is exactly what a leaf pin cannot track — a pin would go stale within days and reject every release after it. Drivers are deliberately untouched: their catalogs keep the DRIVER_CERT_* cert and the installer still plants it as a machine root. The two signatures were always independent (SmartScreen/UAC vs PnP), which is why the installer could move without them. Whether a publicly-trusted catalog would let us drop that root plant is recorded as an unverified follow-up, not assumed. Verified: both scripts parse under the PowerShell 7 AST parser, both workflows are valid YAML, the evaluated Publisher default matches the subject Azure reports for the profile (86 chars, ordinal), rustfmt clean. NOT verified on Windows — the sign path itself needs an on-glass run on .133. |
||
|
|
4499313749 |
fix(nix): the module started a second host in root's systemd, stealing the ports from the real one
ci / docs-site (pull_request) Successful in 1m17s
ci / bun-nix (pull_request) Successful in 1m33s
nix / flake (pull_request) Failing after 1m23s
ci / rust (pull_request) Canceled after 1m54s
ci / rust-arm64 (pull_request) Canceled after 2m3s
ci / web (pull_request) Successful in 1m8s
`systemd.user.*` has no per-user form in NixOS — it installs units into every user's manager. With `host.autoStart` adding them to `default.target`, that included root, whose `user@0.service` exists the moment anybody SSHes in as root. Root's host won the race for the fixed ports and the desktop user's copy crash-looped forever on `bind RTSP 48010: Address already in use`. Every other listener binds first and logs success, so the log reads like a clash with an unrelated program; a second copy of itself running as root is the last thing you look for. `host.users` did not help — it only granted input/punktfunk group membership and never scoped the units. Render `ConditionUser=` on all four user units from `host.users`. Entries are written `|user`: the pipe makes each a triggering condition, which systemd ORs, where plain repeated `ConditionUser=` lines are ANDed and would match nobody. With `host.users` empty, fall back to `!@system` — still keeps root out while leaving the manual `systemctl --user enable --now` route working for a login. module-check.nix gains three assertions covering both branches and web-init keeping its non-triggering ConditionPathExists alongside the new condition. They run in nix.yml's eval leg, and were confirmed to fail against the unfixed module (2 of 23) before being committed. Verified on the box that found this: root force-starting the host now yields ConditionResult=no. |
||
|
|
4caf2b76e8 |
Merge pull request 'gamescope +pfhdr7 — linger no longer dies of its own capture teardown (luxus's fix, overlay#9)' (#212) from worktree-gamescope-linger-pw-destroy-race into main
ci / web (push) Successful in 1m15s
ci / docs-site (push) Successful in 1m29s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 6s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 7s
ci / rust (push) Successful in 7m33s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / builders (ci/gamescope-trixie.Dockerfile, punktfunk-gamescope-trixie) (push) Successful in 8s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 12s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 1m34s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 55s
ci / rust-arm64 (push) Successful in 4m37s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m23s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 1m29s
ci / bun-nix (push) Successful in 4m9s
docker / deploy-docs (push) Successful in 33s
arch / build-publish (push) Successful in 10m38s
docker / builders-arm64cross (push) Successful in 1m45s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m28s
Reviewed-on: #212 |
||
|
|
0f64551c56 |
fix(packaging/gamescope): +pfhdr7 — linger no longer dies of its own capture teardown
Patch 0009, reported, written and proven live by luxus (punktfunk-overlay#9): when the capture consumer leaves, stream_handle_remove_buffer — and the stale-push path in dispatch_nudge — destroyed idle buffers on the PipeWire thread. Dropping the last CVulkanTexture reference there calls into the Vulkan driver (vkDestroyImage / FreeMemory / dmabuf fds) while steamcompmgr can still be inside vulkan_screenshot on another buffer of the same 4-buffer pool; on NVIDIA the race lands as a SIGSEGV in CVulkanCmdBuffer::insertBarrier. The timing is what made it selectively lethal: it fires at stream END — exactly the window where the host keeps the headless display lingering for a reconnect. So the kept display was already dead (journal: linger line → coredump → "kept display was dead — recreating") and the "resumed" session was a fresh compositor with the game lost. The fix queues the corpses (bury_buffer, mutex-guarded) and steamcompmgr reaps them on every vblank, including while the stream is only paused — the linger state itself. Field-proven on the reporter's NVIDIA host: 4 coredumps in one evening of BG3 at 4K60 HDR with --pipewire-composite- cursor (the heaviest paint path we ship), zero after; disconnect/reconnect confirmed live to reuse the lingered session (2026-08-13). Three of the four stacks are this race; the fourth (~CVulkanDevice during exit) is patch 0006's already-fixed static-destruction bug — do not re-diagnose it as part of this. Ours differs from the overlay's original only by the meson.build banner hunk: +pfhdr6 → +pfhdr7, PKGBUILD 3.16.25.pfhdr7-1. No new capability — same rule as pfhdr5/6: "reconnect lost my game" triage has to read a box's exposure off its banner, and every probe is >=. Known residual, deliberately untouched: add_buffer's error path still deletes on the PW thread. By the later `goto error`s a texture may be attached, so the race is reachable there in theory — but only when an add FAILS mid-renegotiation, which no field coredump shows; the patch stays byte-identical with what was proven on-glass. Verified: the full 0001..0009 series applies onto the bare 5fb8dce4 pin with plain `git am` (the build script's own invocation, no -3, no fuzz) and with `git am -3`; the fc44 CI image (punktfunk-fedora44-rpm) builds the result with rpm.yml's exact dep recipe to a binary whose banner reads `3.16.25-20-g40fe8b5+pfhdr7 (gcc 16.1.1)`. After 0009, destroy_buffer has exactly two callers left — pipewire_reap_dead_buffers (steamcompmgr vblank) and pipewire_destroy_buffer (steamcompmgr's copy-completion path) — both on the compositor thread. Nix, deb, sysext and rpm all glob patches/*.patch and read the level off the banner, so no other packaging file moves. |
||
|
|
082c65755f |
fix(packaging): every channel ships the WSI layer, so in-game HDR works off a stock install
ci / web (pull_request) Successful in 1m0s
ci / bun-nix (pull_request) Successful in 2m6s
ci / rust-arm64 (pull_request) Successful in 3m29s
ci / docs-site (pull_request) Successful in 3m53s
android / android (pull_request) Successful in 5m4s
ci / rust (pull_request) Successful in 5m9s
apple / swift (pull_request) Successful in 2m0s
apple / distribute (pull_request) Skipped
apple / screenshots (pull_request) Skipped
nix / flake (pull_request) Failing after 6m17s
The previous commit built the layer and taught the host to use it, but only the
Arch PKGBUILD carried the files, so every other channel still landed on the
no-game-HDR fallback. This finishes the job.
The packaging scripts now take `--stage`, the DESTDIR the gamescope build script
wrote, instead of a path to one binary. That is the part worth keeping: the next
file this package needs will not require a new flag in four scripts and two
workflows. CI caches the whole staged tree for the same reason. The gs-cache key
already hashes packaging/gamescope/**, which this commit changes, so stale caches
in the old single-file shape cannot be restored into the new layout.
Channels, all of them:
rpm spec gains Source1/Source2 and %files entries
deb build-gamescope-deb.sh copies the layer into the package root
Arch PKGBUILD (previous commit); the sysext extracts the whole usr tree
sysext bazzite takes --gamescope-stage; arch asserts the layer arrived
nix the derivation keeps, renames and rewrites the layer rather than
deleting it with everything else
A missing layer is fatal in every one of them, not best-effort. A package that
carries the compositor without it looks completely healthy and then silently
denies every game an HDR10 swapchain -- the exact failure this whole change
exists to end, so it must not be possible to ship it again by accident.
Two things needed care:
The layer manifest carries an ABSOLUTE library_path baked in at build time, so
every channel has to install the .so at exactly that path. That means literal
/usr/lib/punktfunk, not %{_libdir} (which is /usr/lib64 on Fedora) and not a
Debian multiarch triplet. Nothing links the .so by soname -- the loader dlopens
it by that path -- so multilib has no claim here. The rpm and nix install checks
now read the path back out of the manifest and fail if it names a file the
package does not install, because a manifest pointing at nothing is the silent
shape of this bug.
NixOS has no /usr, so the layer lives inside the gamescope derivation and the
host's path is overridable via PUNKTFUNK_GAMESCOPE_WSI_LAYER_DIR, which the
module sets -- the same posture as PUNKTFUNK_GAMESCOPE_BIN, and documented.
The manifest rewrite moved out of a heredoc into
packaging/gamescope/rewrite-wsi-layer-manifest.py because the FHS builds and the
Nix store both need it and must rename the layer identically; two copies would
drift into a host looking for a name only one of them produces.
Verified: 214 pf-vdisplay tests pass in a linux container, clippy -D warnings and
rustfmt clean, bash -n on all five changed shell scripts, both workflow YAMLs
parse, and the rewrite script was run against a synthetic FROG manifest to
confirm it renames/repoints/regates while preserving the `functions` block --
which is the field that decides whether the layer loads at all.
NOT verified: no nix on this machine, so gamescope.nix, flake.nix and the module
are unevaluated; no gamescope build, no package build of any kind, and no game
has taken an HDR swapchain on glass.
|
||
|
|
3ac4548cf8 |
fix(gamescope): ship the WSI layer built beside our compositor, instead of guessing at the distro's
A game nested under gamescope gets an HDR10 swapchain from the FROG WSI layer and from nothing else -- gamescope advertises no runtime colour-management protocol a Mesa/NVIDIA WSI could negotiate through. That layer talks `gamescope_swapchain` to the compositor, and when the two disagree the compositor rejects the client's swapchain_feedback and every Vulkan client dies on a black screen with sound and input and no error anywhere. We ship our own compositor and did NOT ship a layer, on the recorded grounds that the layer is "version-independent of the compositor binary". It is not, and wsi_layer_matches_our_gamescope() exists because it is not. So the host was left guessing from version triples, and that guess is wrong in both directions: a distro at the same upstream tag that patched the protocol compares EQUAL and keeps a layer that will kill every game, while a distro at a different tag with a byte-identical protocol compares unequal and loses HDR for nothing. Since we pin a rev, the second case is the normal one -- on essentially every box with a distro gamescope, the layer was disabled and no game could render HDR. Ship the layer instead. It is built from the same tree at the same rev as the compositor, so the two cannot drift, and the guess stops being load-bearing. It is installed under our own name (VK_LAYER_PUNKTFUNK_gamescope_wsi) at our own path with our own enable/disable variables, so it coexists with the distro's rather than colliding -- the Vulkan loader keys implicit layers on that name -- and the host switches the two independently in one session. WsiPlan makes the three states explicit and resolves them once per launch, since the fallback spawns `--version` probes: Ours our layer is installed: enable it, force the distro's off DistroKept no layer of ours, distro's looks compatible: touch nothing DistroDisabled no layer of ours, distro's untrusted: today's behaviour That last arm is the fail-safe. A host newer than its gamescope package behaves exactly as it does today rather than enabling a layer that is not there, so this can roll out one packaging surface at a time without a flag day. Only the Arch PKGBUILD carries the new files so far. The rpm path takes a CI-cached binary rather than the build script's stage dir, so it needs the cache, build-gamescope-rpm.sh and the spec moved together; the deb, both sysexts and gamescope.nix need the same two files added. Until each lands, those boxes take the DistroDisabled arm and are no worse off than before. Verified: 214 pf-vdisplay tests pass in a linux container (including a new one pinning that the Ours arm enables ours AND forces the distro's off together -- either half alone is a bug), clippy -D warnings and rustfmt clean, both shell files pass bash -n, and the manifest rewrite was run against a synthetic FROG manifest to confirm it renames/repoints/regates while preserving the `functions` block. NOT verified: an actual gamescope build, any package build, or a game taking an HDR swapchain on glass. |
||
|
|
4ebe7d1185 |
fix(host): the plugin runner was located by FHS path only, so NixOS never found it
On NixOS every plugin PACKAGE op failed with "the plugin runner isn't installed" on a box where the runner was installed, enabled and running. `runner_command()` checked FHS locations exclusively — /usr/bin, the /usr/lib + /usr/share pair behind it, and the ~/.local mirror the SteamOS installer lays down. Nix ships punktfunk-scripting as a derivation of its OWN, so its wrapper is neither beside the host binary nor anywhere under /usr, and no rung could ever match. Service ops go through systemd and were unaffected, which is what made it read as arbitrary: `plugins status` said running/enabled while `plugins add` said not installed. Resolution now matches punktfunk-encode-worker's: PUNKTFUNK_SCRIPTING -> beside the host binary -> PATH -> /usr -> ~/.local. PATH is the rung Nix lands on. The /usr rungs stay AFTER it rather than being dropped, because a systemd unit's PATH need not include /usr/bin. As with the encode worker the env override is deliberately not existence-checked — a named path that is wrong should fail naming itself, not fall through to some other runner. Lifted into a pure injected function so the whole table is testable, which is also how the regression is pinned: removing the PATH rung fails the NixOS row specifically. Second half, and the reason the Rust change alone would not have fixed the console: the NixOS module now puts the runner on the HOST UNIT's `path`. The console installs plugins from inside the host service, whose PATH is exactly that unit list — `environment.systemPackages` only ever covered an operator's interactive shell. Without it the CLI would have been fixed and the console would not. module-check.nix gains both the positive and the negative assertion, so CI's `nix flake check --no-build` holds the property. The error text named only apt and SteamOS; it now names NixOS and the override. The ~/.local/bin symlink workaround is no longer needed. |
||
|
|
1679275272 |
fix(flatpak): pin the skia-binaries archive to 0.99.0 — #193 bumped the crate and left the tarball at 0.87
ci / bun-nix (pull_request) Successful in 26s
ci / docs-site (pull_request) Successful in 1m19s
ci / rust-arm64 (pull_request) Successful in 1m47s
ci / web (pull_request) Successful in 3m23s
android / android (pull_request) Successful in 5m16s
windows-client / client (x64, , x86_64-pc-windows-msvc, C:\t) (pull_request) Successful in 7m3s
windows-client / client (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (pull_request) Successful in 2m57s
ci / rust (pull_request) Successful in 8m51s
The dependency currency wave took skia-safe/skia-bindings 0.87.0 -> 0.99.0 in
crates/pf-console-ui/Cargo.toml, but packaging/flatpak/io.unom.Punktfunk.yml still
pinned the 0.87.0 prebuilt archive, so every flatpak leg since the merge dies with
error[E0599]: no variant, associated function, or constant named `Default`
found for enum `SkPathFillType` (and `SkPathDirection`)
--> cargo/vendor/skia-bindings-0.99.0/src/defaults.rs:57
Nothing about that message points at the manifest, so it reads like a crate bug. It
isn't. `SKIA_BINARIES_URL: file://…` makes skia-bindings unpack the pinned tarball
verbatim into target/…/build/skia-bindings-*/out/skia/ — *including the bindings.rs
it was generated with*. Those two `Default`s are associated consts emitted INTO
bindings.rs, so they travel with the archive, not with the crate: 0.99.0's
src/defaults.rs was compiling against 0.87.0-era bindings. Verified directly — the
0.99.0 archive carries `impl SkPathFillType { pub const Default = Winding }` and
`impl SkPathDirection { pub const Default = CW }` on both x86_64 and aarch64.
Because the URL is file://, the fetch can never fail, so there is no download error
to notice — the only symptom is a compile error deep in a vendored crate.
The asset name changed across the bump: `jpeg` entered skia-safe's defaults at 0.99,
so the resolved-feature key went `pdf-textlayout-vulkan` -> `jpegd-jpege-pdf-textlayout-vulkan`.
Confirmed against each archive's own key.txt/tag.txt (tag 0.99.0, key
a25a0fdb7d90429aa2d1-<target>-jpegd-jpege-pdf-textlayout-vulkan), and libskparagraph.a
plus the Vulkan backend symbols are present, so the feature set still matches what
pf-console-ui resolves.
Everything else in the offline chain (Cargo.lock, cargo-sources.json) is regenerated
from the lock and self-corrects; this tarball is the single hand-maintained pin, which
is exactly why it was the thing left behind. Both bump sites now carry a pointer to
the other so the next one can't split-brain the same way.
|
||
|
|
c95db8eebc |
Merge pull request 'Move TLS to aws-lc-rs with post-quantum key exchange, drop ring via ureq 3, and fix the dependency defects behind it' (#192) from worktree-aws-lc-rs-migration into main
apple / swift (push) Successful in 2m18s
audit / bun-audit (plugin-kit) (push) Successful in 20s
audit / bun-audit (web) (push) Successful in 14s
audit / docs-site-audit (push) Successful in 16s
audit / pnpm-audit (push) Successful in 12s
audit / cargo-audit (push) Successful in 2m44s
audit / bun-audit (sdk) (push) Successful in 1m33s
audit / license-gate (push) Successful in 5m13s
ci / rust (push) Failing after 1m48s
ci / web (push) Successful in 1m45s
audit / miri (push) Successful in 10m5s
ci / docs-site (push) Successful in 1m29s
audit / c-abi-asan (push) Successful in 10m24s
ci / bun-nix (push) Successful in 1m6s
apple / distribute (push) Successful in 14m0s
deb / build-publish-gamescope (push) Failing after 56s
windows-drivers / driver-build (push) Successful in 1m59s
android / android (push) Successful in 19m12s
arch / build-publish (push) Successful in 19m57s
windows-drivers / probe-and-proto (push) Successful in 25s
ci / rust-arm64 (push) Successful in 13m8s
apple / screenshots (push) Successful in 7m34s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 2m14s
deb / build-publish (push) Successful in 10m28s
docker / builders (ci/gamescope-trixie.Dockerfile, punktfunk-gamescope-trixie) (push) Successful in 10s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 7m38s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 2m52s
deb / build-publish-client-arm64 (push) Successful in 12m9s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 7m56s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 4m49s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m29s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 4m27s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 3m12s
deb / build-publish-host (push) Successful in 19m29s
docker / deploy-docs (push) Failing after 1m58s
docker / builders-arm64cross (push) Successful in 6m28s
deb / smoke-install (push) Failing after 4m16s
flatpak / build-publish (push) Successful in 12m18s
nix / flake (push) Failing after 20m59s
windows-host / package (push) Successful in 19m52s
windows-host / winget-source (push) Skipped
windows-client / client (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 7m36s
windows-host / canary-manifest (push) Successful in 27s
windows-client / client (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 9m53s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 22m36s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 2m50s
Reviewed-on: #192 |
||
|
|
6202543b21 |
Merge pull request 'ci: cache the C/C++ half, link with mold, fix the debug/release cache collision, consolidate the Apple and Windows-client workflows' (#191) from worktree-ci-optimization into main
android / android (push) Canceled after 34s
apple / swift (push) Canceled after 1m49s
apple / distribute (push) Canceled after 0s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 0s
ci / rust (push) Canceled after 0s
ci / rust-arm64 (push) Canceled after 55s
ci / web (push) Canceled after 49s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 42s
deb / build-publish (push) Canceled after 15s
deb / build-publish-host (push) Canceled after 0s
deb / build-publish-gamescope (push) Canceled after 0s
deb / build-publish-client-arm64 (push) Canceled after 0s
deb / smoke-install (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/gamescope-trixie.Dockerfile, punktfunk-gamescope-trixie) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 0s
windows-client / client (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Canceled after 0s
windows-client / client (x64, , x86_64-pc-windows-msvc, C:\t) (push) Canceled after 0s
windows-drivers / probe-and-proto (push) Canceled after 0s
windows-drivers / driver-build (push) Canceled after 0s
windows-host / package (push) Canceled after 0s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
decky / build-publish (push) Successful in 1m23s
Reviewed-on: #191 |
||
|
|
51a005dd43 |
fix(deps): close the audit gaps, drop unused declarations, declare what is used
Acting on the 2026-08-13 dependency sweep. Every claim below was re-verified against
the tree before acting on it (greps carry a positive control; the advisories were
re-checked with cargo audit 0.22.2).
SECURITY
- event-listener 5.4.1 -> 5.4.2 (RUSTSEC-2026-0221, unsound Send/Sync on StackSlot;
reaches the tray via zbus and the host via ashpd). This sat unnoticed because
`cargo audit` reports unsoundness as a WARNING and the job fails only on
vulnerabilities — audit.toml now says so out loud.
- spin 0.9.8 -> 0.9.9. 0.9.8 is YANKED and was genuinely compiled (flume via mdns-sd
and relm4, plus lazy_static).
- wayland-scanner 0.31.10 -> 0.31.11, which moves quick-xml 0.39 -> 0.41. That is the
exact trigger audit.toml documented for RUSTSEC-2026-0194/0195, so both ignores are
deleted rather than left as permanent exceptions. Only RUSTSEC-2023-0071 (rsa
Marvin, still unfixed upstream) remains.
- Corrected audit.toml's claim that `paste` arrives "via utoipa-axum": rav1d pulls it
too, so every client has it through the decode path and dropping utoipa-axum would
not have cleared it.
TWO CI GATES THAT SCANNED NOTHING
- `cargo audit` only ever reads the ROOT Cargo.lock. The drivers lock was already in
this job's `paths:` filter, so edits to it triggered a run that then ignored them.
All four secondary workspaces now get an explicit `--file` (verified: clean, bar the
known `paste` warning in drivers).
- packaging/windows/pf-vkhdr-layer had NO lockfile at all while shipping as a DLL in
the host installer, so every build resolved fresh and neither cargo-audit nor
cargo-about ever saw it. Lockfile generated and committed, and added to `paths:`.
UNUSED / DUPLICATE DECLARATIONS
- punktfunk-host: removed 13 dependencies it never references — the Wayland stack
(client, protocols{,-wlr,-misc}, scanner, backend), xkbcommon, reis, khronos-egl,
ash, usbip-sim, parking_lot, bytemuck. The code moved to pf-inject and pf-zerocopy
in the subsystem extraction and those crates declare them; only the manifest entries
and their now-false comments stayed. Also dropped four redundant re-declarations
(tokio/serde_json/futures-util in the Linux block, tower in dev-deps).
- Removed genuinely unused: bytes (punktfunk-core), anyhow (pf-win-display),
tracing (clients/cli), anyhow (clients/session), serde (clients/windows).
- Removed the high-level `wdk` crate from all five driver crates and the drivers
workspace: none of them ever referenced `wdk::` (62 `wdk_sys::` uses; pf-umdf-util
is a full WDF crate that never declared it). `tracing`/`tracing-subscriber` remain
in that lock afterwards but ONLY as wdk-sys build-dependencies, not in the DLLs.
- pf-win-display took punktfunk-core with `quic` for one type (`Mode`) that lives in
the ungated `config` module; now `default-features = false`, which keeps
quinn/tokio/rcgen/opus out of a leaf crate's declared closure.
- pf-encode declared the windows-rs feature `Wdk_Graphics_Direct3D` for a call that
lives in pf-frame and is resolved via GetProcAddress on gdi32.
LATENT BREAKAGE (compiled only by feature unification)
- pf-inject uses `tokio::select!` without declaring `macros` (borrowed from
punktfunk-core's quic feature); pf-capture uses `tokio::sync::oneshot` without
declaring `sync` (borrowed from ashpd->zbus); pf-client-core uses the `minwindef`
and `winnt` windows-rs headers without declaring them (borrowed from
clients/windows). Each now declares what it uses, so an unrelated crate changing its
features cannot break them.
- pf-console-ui took pf-client-core WITHOUT `default-features = false`, unlike every
other consumer. That default is `pyrowave`, which compiles the vendored PyroWave C++
— "fatal on Windows ARM64". Only safe today because the ARM64 leg passes
--no-default-features (which also drops `ui`).
CORRECTED A FALSE INVARIANT
- clients/windows claimed "the workspace builds ONE windows-rs". It does not: wasapi
pulls the crates.io windows 0.62.2 beside the git-rev copy. The invariant that DOES
hold is narrower (reactor and that crate share one rev, which is what makes the
IDXGISwapChain1 hand-off type-check). Comment rewritten, with a warning against
"fixing" it via a blanket [patch.crates-io] — this rev uses header-named features
while a dozen other manifests use the old Win32_* namespace ones.
Plus the safe in-compat `cargo update` sweep (no manifest edits).
Verified on macOS: punktfunk-core 385, pf-update-check 32, c_abi 1 (with
LIBRARY_PATH=/opt/homebrew/opt/opus/lib), cargo audit clean bar the two known
unmaintained warnings. Linux and Windows legs follow.
|
||
|
|
5fbf04f56d |
Ship punktfunk-gamescope on apt, support Debian 13, and state the real host floor (#190)
ci / docs-site (push) Successful in 1m19s
windows-host / package (push) Failing after 55s
ci / rust-arm64 (push) Successful in 1m40s
windows-host / canary-manifest (push) Skipped
windows-host / winget-source (push) Skipped
ci / bun-nix (push) Successful in 20s
apple / swift (push) Successful in 1m39s
deb / build-publish-gamescope (push) Failing after 2s
android / android (push) Successful in 6m22s
deb / build-publish-host (push) Successful in 4m38s
deb / build-publish (push) Successful in 5m31s
ci / web (push) Successful in 7m54s
apple / screenshots (push) Successful in 5m57s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 3m30s
deb / build-publish-client-arm64 (push) Successful in 8m26s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 6m4s
docker / builders (ci/gamescope-trixie.Dockerfile, punktfunk-gamescope-trixie) (push) Successful in 1m45s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 5m41s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m27s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 3m19s
decky / build-publish (push) Successful in 51s
arch / build-publish (push) Successful in 10m45s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10m24s
ci / rust (push) Canceled after 6m35s
deb / smoke-install (push) Canceled after 2m50s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 2m34s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 1m2s
docker / builders-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 1m33s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 6m22s
|
||
|
|
3f7fbf1061 |
ci: build the web console once per push instead of once per packaging job
The Nitro console bundle is a pure function of web/ and sdk/, and it was being built
six times on every push: ci.yml, deb, both RPM legs (f43 + f44), arch, and the docker
app image, at roughly 2.5 min each. windows-host.yml has cached it on exactly this
shape for a while — this extends the same arrangement to the Linux packaging legs,
sharing one key family so whichever job builds it first warms the others.
The bun version is part of the key. Each builder image runs the bun.sh installer at
image-build time, so rust-ci, fedora-rpm and arch-ci can drift apart; keying on it
means they share while they agree and simply stop sharing when they do not, rather
than one image's bun silently producing the bundle another image ships.
Each packaging path needed a different hand-off:
* deb — build-web-deb.sh already builds only if web/.output is missing, so the
restore alone is enough; the workflow's build+smoke step is now gated on
the miss.
* arch — makepkg builds with PF_SRCDIR pointing at the workspace, so a restored
bundle is already where it needs to be. PKGBUILD gains the same
build-if-missing guard the deb script has.
* rpm — neither direction works by default. build-rpm.sh packages a `git archive`
tarball and web/.output is gitignored, so a bundle in the workspace is
invisible to rpmbuild; and the spec's own build lands in rpmbuild's
%{_topdir}, which build-rpm.sh mktemps and removes on EXIT, so a console
built there is gone before the cache's post step and the cache would never
populate — every run a miss that quietly rebuilt. So the workflow builds it,
and hands it over by absolute path through a new optional `pf_prebuilt_web`
macro. Undefined (plain rpmbuild, COPR) takes the original build path.
Every path asserts the bundle exists and carries the Bun.serve marker, on cache hits
too. A cache is one more place a wrong artifact can come from, and the packaging
scripts' build-if-missing behaviour — correct for a local build — would otherwise turn
a broken restore into either a silent rebuild or, with the build step skipped, a
package with no console in it. That is not hypothetical: windows-host.yml shipped
0.22.1 and 0.22.2 with no console because an unset path variable was handled by a
single Write-Host, which is why its equivalent step throws.
|
||
|
|
346385bad8 |
fix(deb): ship punktfunk-gamescope on apt at last, and support Debian 13
`punktfunk-gamescope` had never been published to the apt registry — not in any release. It was built inside the host job's Ubuntu 24.04 image, where it cannot build: our pin vendors wlroots 0.19.3, which floors `wayland-server` at 1.23.1, and noble ships 1.22.0 (it also lacks libxcb-errors-dev and has only libdisplay-info 0.1.1). Every rung of that path was a `::warning::` returning 0 and the one hard gate ran last by design, so v0.26.0 and v0.27.0 both released with the package missing while docs-site told apt users to install it. The same tags shipped it fine for Arch, Fedora 44 and Bazzite. It now builds in its own job on Debian 13 (ci/gamescope-trixie.Dockerfile), the oldest apt base the tree configures on. One package serves Debian 13 AND Ubuntu 26.04 — measured by installing and running it on both — because the build also vendors libdisplay-info via the new `--extra-fallback` option: linked against the distro copy it demands `libdisplay-info2` on trixie, which Ubuntu 26.04 does not have (it carries libdisplay-info3). The option is opt-in, so the Arch/Fedora/nix outputs are byte-for-byte unchanged. Ubuntu 24.04 gets no gamescope package and cannot — its wayland is too old to run one however built. Debian 13 is now a documented host target. That needed no packaging change at all: the host .deb's glibc-2.39 floor and bundled FFmpeg already made it installable, and it had been working for a long time while docs-site said Debian was unsupported and unverified. Verified by installing: host, web console and plugin runner install, resolve every soname and run. The desktop client stays Ubuntu-26.04-only (built there, floors at `libc6 >= 2.43`; Debian 13 has 2.41). Compositor detection now answers Cinnamon (Mint, LMDE) with the route that works instead of advice that cannot help. Muffin forked from Mutter 3.36: `org.cinnamon.Muffin.ScreenCast` has only RecordMonitor/RecordWindow, never RecordVirtual, and xdg-desktop-portal-xapp implements no ScreenCast — so no value of PUNKTFUNK_COMPOSITOR makes a Cinnamon desktop host a virtual display. The error names headless gamescope, which needs no desktop compositor. The XDG sniff moved into a pure function so those branches are testable; Cinnamon is matched before GNOME, since it is a GNOME derivative and the generic arm would otherwise hand it the Mutter backend (caught by the new test). New `smoke-install` job installs every published package from the registry in pristine ubuntu:24.04, ubuntu:26.04 and debian:trixie images, asserts each binary resolves its libraries and runs, and insists the version served is the one this run built. Nothing in deb.yml had ever installed a package it produced, which is how both of the above survived unnoticed. ⚠ Bootstrap: seed `punktfunk-gamescope-trixie:latest` into the LAN registry once (docker.yml builds it thereafter) or the new job cannot start. |
||
|
|
ba16237c35 |
fix(bazzite): the shipped template pinned ATTACH, so Game Mode mirrored the box's screen instead of giving the client its own display
apple / swift (pull_request) Successful in 1m46s
apple / screenshots (pull_request) Skipped
android / android (pull_request) Successful in 8m43s
ci / bun-nix (pull_request) Successful in 22s
ci / web (pull_request) Successful in 1m7s
ci / docs-site (pull_request) Successful in 3m13s
ci / rust-arm64 (pull_request) Successful in 3m18s
ci / rust (pull_request) Successful in 15m34s
Field report: "on Bazzite when using gaming mode it is mirroring the main display instead of giving the client its own." It is our own template that does it. `packaging/bazzite/host.env` set `PUNKTFUNK_GAMESCOPE_ATTACH=1`, and every install path — rpm, deb, Arch, nix — ships that file as `/usr/share/punktfunk/host.env.bazzite` with the docs telling people to copy it verbatim. So the recommended Bazzite setup turned the attach override ON for everyone. That override is rung 2 of `pick_gamescope_mode`, ABOVE `dedicated_launch` at rung 3. The rung comment calls the operator overrides a debug/CI escape hatch, which is right — but we were shipping one as a distro default, so on a Bazzite box the managed takeover and the dedicated game session were both unreachable. A game launched from a client's library could not get a session of its own either, which is the case the dedicated route exists for. With a physical display connected, attach then takes the `physical_display_connected()` arm and streams the box's own head at the box's own mode: the mirror the reporter saw. The template now forces nothing and lets the per-connect detection answer, which on a box with `gamescope-session-plus` is MANAGED. Attach stays available, documented as the opt-in it is, with the mirror and the dedicated-session cost stated. Because managed depends on the `punktfunk` group to stop the display manager, the template now says so where someone choosing a model will read it, rather than only in the distro guide. Also fixes the off-switch. Both overrides were read with `var_os(..).is_some()`, so `PUNKTFUNK_GAMESCOPE_ATTACH=0` meant ATTACH ON — the opposite of what the line says, and of every other knob on this host. They now use the shared `env_on` grammar, so `0|false|off|no` disable and a bare `=1` keeps working. Anyone who "turned attach off" in an older host.env had it on the whole time. Note an upgrade never rewrites an existing `~/.config/punktfunk/host.env`, so boxes set up from an older template keep the pin until the line is deleted by hand; the Bazzite and HDR pages now say that. Verified: `scripts/xcheck.sh linux` check + clippy `-D warnings` clean, pf-vdisplay 206/0 under rust:1.96, `cargo fmt --all --check` clean. Gate proved non-vacuous against a planted `compile_error!` in routing.rs. |
||
|
|
c68e0be688 |
Merge remote-tracking branch 'origin/main' into worktree-edition-2024
ci / bun-nix (pull_request) Successful in 24s
windows-drivers / probe-and-proto (pull_request) Successful in 34s
ci / docs-site (pull_request) Successful in 1m23s
apple / swift (pull_request) Successful in 1m45s
apple / screenshots (pull_request) Skipped
ci / web (pull_request) Successful in 3m4s
windows-drivers / driver-build (pull_request) Successful in 2m23s
ci / rust (pull_request) Failing after 4m6s
ci / rust-arm64 (pull_request) Failing after 5m5s
windows / build (x86_64-pc-windows-msvc) (pull_request) Failing after 1m54s
nix / flake (pull_request) Successful in 13m36s
windows / build (aarch64-pc-windows-msvc) (pull_request) Failing after 54s
android / android (pull_request) Successful in 15m6s
|
||
|
|
f373dffb5e |
chore: migrate the main workspace and pf-vkhdr-layer to edition 2024 (WP20)
The safety half of the rust-safety programme's §8.4: `std::env::set_var`/`remove_var` are
`unsafe fn` in edition 2024, converting the class of bug the programme found the hard way
(the
|
||
|
|
6a506a8fa9 |
fix(vdisplay/driver,pf-frame): no punktfunk process holds REALTIME GPU priority by default
windows-drivers / probe-and-proto (pull_request) Successful in 30s
ci / bun-nix (pull_request) Successful in 56s
ci / docs-site (pull_request) Successful in 1m16s
ci / web (pull_request) Successful in 1m17s
ci / rust-arm64 (pull_request) Successful in 1m31s
windows-drivers / driver-build (pull_request) Successful in 1m47s
android / android (pull_request) Successful in 4m40s
ci / rust (pull_request) Successful in 9m54s
apple / swift (pull_request) Failing after 13m27s
apple / screenshots (pull_request) Skipped
The RX 9070 XT field A/B (2026-08-11/12 logs) convicted BOTH of our REALTIME GPU-scheduling levers of generating the metronomic capture-stall class the stall program has chased for weeks — compose-silence holes of 150-800 ms in which ETW shows NO process presenting while the GPU stays responsive: - the vdisplay driver's IddCxSetRealtimeGPUPriority raise beat at ~1.75-1.78 s (PFVD_NO_RT_GPU=1 alone removed that metronome: ~0.35 stalls/s metronomic -> 10 sparse aperiodic over 3.9 min); - the host auto-gate's HIGH->REALTIME upgrade (pf-frame dxgi.rs, T2.3) beat at ~3.58 s in the AV1 sessions where it promoted (vram_pct=1, 12:59:26); pinning PUNKTFUNK_GPU_PRIORITY_CLASS=high removed that residual too (13:45 session: zero metronomic, stall rate at the clean-run baseline). Neither period matches any punktfunk clock: the full periodic-actor census (driver: event-paced drain + 16 ms E_PENDING wait, 33 ms cursor poll, 3 s watchdog reap; host: 250 ms descriptor poll, 5/50/100 ms probes + ~2 s scanline retarget, 2 s VRAM gate, 2 s exclusive re-assert, 3.33 s pinger, 1 s stats, ~1 Hz phase-lock, fps/2 LTR marks) has nothing in the 1.69-2.29 s band, and every host-side actor ran unchanged in the A/B that killed the fast metronome. The periodicity is emergent from holding an unreachable-priority queue against the WDDM scheduler on this AMD family (the period even differs by which of our processes holds REALTIME); it is not a punktfunk cadence being amplified, so there is nothing punktfunk-periodic to fix - the fix is to stop holding REALTIME by default, which is also canonical parity (no shipping IDD raises it, and HIGH was the class that delivered the original Sunshine-parity encode win). - Driver: PFVD_NO_RT_GPU (default-ON, opt-OUT) becomes the PFVD_RT_GPU ladder, default OFF on every vendor: unset = no raise (canonical IDD behavior); =thread = SetGPUThreadPriority(+7), a graduated in-band middle rung for field A/B (not default: unmeasured here, and the host measured the same call as "no help" for its own starvation case); anything else = the old REALTIME DDI. PFVD_NO_RT_GPU stays recognized and WINS over the opt-in, so the field boxes that carry it through the default-ON era keep meaning OFF. Both directions remain A/B-able without a rebuild (machine env + device restart). The CPU half of the original branch-2 hardening (MMCSS / TIME_CRITICAL) is untouched - it addressed the delivery holes that were actually observed. - Host: PUNKTFUNK_GPU_PRIORITY_CLASS default auto -> high. `auto` (the gated REALTIME upgrade) stays available as an explicit opt-in, `realtime` still pins; unrecognized values now land on the HIGH default instead of silently opting into the gate - a typo must not buy the hazard. The VRAM/HAGS gate machinery is unchanged for `auto`; it guards the NVENC-hang hazard but cannot see this one. - stall.rs: the no-OS-event METRONOMIC warning now carries rt_gpu_driver / rt_gpu_host fields (the machine-env state of both levers) and names clearing them as the FIRST cure, ahead of the display-hardware suspects - a field log self-answers the triage question this program just spent a week on. No console policy axis for the driver knob: the lever is default-safe now, the driver reads config at WUDFHost scope where machine env already matches the device-restart lifecycle, and a policy axis would need pf-driver-proto churn (or a device-key registry write) for an experimental lever that only exists to be A/B-ed. If the `thread` rung ever proves out as a default-worthy raise, that is the moment to revisit. |
||
|
|
d67ab9ede4 |
chore(safety): two .133 gate findings — cfg the abi lock helper, re-anchor a layer proof
ci / bun-nix (pull_request) Successful in 40s
ci / web (pull_request) Successful in 1m24s
ci / docs-site (pull_request) Successful in 1m30s
android / android (pull_request) Canceled after 1m45s
apple / swift (pull_request) Canceled after 1m41s
apple / screenshots (pull_request) Canceled after 0s
ci / rust (pull_request) Canceled after 1m43s
ci / rust-arm64 (pull_request) Canceled after 1m43s
nix / flake (pull_request) Canceled after 1m29s
windows-drivers / probe-and-proto (pull_request) Canceled after 0s
windows-drivers / driver-build (pull_request) Canceled after 1m25s
windows / build (aarch64-pc-windows-msvc) (pull_request) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (pull_request) Canceled after 0s
The tray leg builds punktfunk-core with default-features off, where lock_recover's only callers (the quic-gated punktfunk_connection_* entry points) do not exist — dead code under -D warnings. The helper takes the same feature gate. In pf-vkhdr-layer, rustfmt had reflowed destroy_surface's lookup into a multiline closure, leaving the SAFETY comment outside the closure that contains its unsafe block — the box's clippy rightly stopped accepting the adjacency. The comment moves inside, directly above the block. |
||
|
|
2bfd1cd2d5 |
chore(safety): three unsafe-hygiene grep gates, blocking in ci.yml (WP2c gates)
scripts/ci/check-unsafe-hygiene.sh — textual gates for three classes no lint covers: A. unsafe fn markers carrying no contract. unsafe_op_in_unsafe_fn forces real ops into blocks, so an unsafe fn with no `unsafe` in its body is a marker with no contract ( |
||
|
|
dfebb9dfbb |
chore(safety): hoist the unsafe lints into the workspace tables (WP2c hoist)
undocumented_unsafe_blocks joins unsafe_op_in_unsafe_fn in
[workspace.lints], and the ~100 scattered per-file #![deny(...)] attributes
(85 files) are deleted — a new crate, or a new module in an old one, is now
covered on creation rather than on remembering. The per-file form is how
pf-vkhdr-layer, wdk-probe and half of pf-clipboard stayed uncovered.
There are THREE workspaces, so the claim is made three times: the main
Cargo.toml, packaging/windows/drivers (workspace table + [lints]
workspace = true in all seven members), and packaging/windows/pf-vkhdr-layer
(its [lints] table, previous commit). pf-update now opts into workspace
lints; the two vendored member snapshots (cros-codecs, usbip-sim) stay out
deliberately and now both say so.
Newly-covered fallout was two link-sanity tests (pyrowave-sys, libvpl-sys)
— proofs written. Stale prose that claimed the workspace held
unsafe_op_in_unsafe_fn at "warn" (it has been deny) or pointed at the
deleted attributes is corrected.
nvenc_core.rs is carved OUT of the unsafe_op_in_unsafe_fn fence: its
exemption rationale ("raw entry-table calls almost line for line") was
false — the file makes zero FFI calls. Its unsafe surface is C-union writes
whose soundness hangs on which codec arm is active, and its own 4:4:4 note
records the shipped bug (hevcConfig bytes stamped onto an AV1 config) that
per-operation blocks make visible. It now runs the strictest discipline in
the crate: clippy::multiple_unsafe_ops_per_block at deny, one union access
per block, each naming its codec guard.
Verified here: cargo fmt clean in all three workspaces; native clippy
-D warnings clean for everything that compiles on macOS (the three
pre-existing mac-native failures — pf-client-core wol.rs, pf-encode
dead-code/closure-call, probe mic_burst — reproduce on the clean tree).
Linux/Windows legs ride the .25/.133 gate.
|
||
|
|
23fa03b051 |
chore(safety): close the three crate-level lint gaps (WP2b)
pf-vkhdr-layer — the sharpest gap: an implicit layer injected into every Vulkan game process, 32 unsafe usages, zero SAFETY comments, own workspace so no lint table reached it, and an explicit missing_safety_doc allow. Now: a [lints] table (unsafe_op_in_unsafe_fn + clippy::undocumented_unsafe_blocks, both deny), the allow removed, every unsafe operation in an explicit block with a real proof (loader layer protocol / Vulkan valid-usage), # Safety docs on the contract-carrying fns, const layout asserts for the SurfaceFormat2Raw mirror, the five helpers with no caller-facing contract demoted to safe fns, and the two redundant `unsafe impl Send` deleted (fn pointers and vk handles are Send intrinsically — the type-check proves it). wdk-probe — 21 unsafe blocks, 12 proofs: the 9 missing SAFETY comments are written (the iddcx_rt.rs DDI slot-dispatch ones are about table population and PFN/index pairing, not pattern fill), the sibling denies added at the crate root, missing_safety_doc allow dropped, # Safety on DriverEntry, and the crate joins windows-drivers.yml's clippy list — it was the only driver crate not in it. pf-clipboard — the undocumented_unsafe_blocks deny moves from host/windows.rs to the crate root so host/wayland.rs (4 blocks), host/mutter.rs (2) and any future backend under host/ are covered on creation. All existing blocks already carry proofs; free today, structural tomorrow. Verified here: pf-vkhdr-layer cargo fmt --check + clippy --release -D warnings at x86_64-pc-windows-msvc. wdk-probe and pf-clipboard compile checks need the WDK/Linux boxes and ride the .133/.25 gate. |
||
|
|
bc70a58fb1 |
Merge main into chore/rust-safety-programme
windows-drivers / probe-and-proto (pull_request) Successful in 22s
ci / bun-nix (pull_request) Successful in 36s
ci / web (pull_request) Successful in 1m17s
apple / swift (pull_request) Successful in 1m39s
apple / screenshots (pull_request) Skipped
windows-drivers / driver-build (pull_request) Successful in 1m58s
ci / rust-arm64 (pull_request) Successful in 3m18s
ci / docs-site (pull_request) Successful in 3m55s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m29s
android / android (pull_request) Successful in 4m46s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m30s
ci / rust (pull_request) Failing after 10m11s
nix / flake (pull_request) Successful in 15m6s
Two conflicts: the test-module import list in gamescope.rs (union — the branch's takeover-state tests and main's WSI opt-out tests both stay), and next_frame_timed_out in pf-capture, where the branch still carried the pre-#168 else-if chain — resolved to main's match-based refactor, which already embeds the same arm semantics plus the provisional-budget latch gate. |
||
|
|
6ca192b9ab |
fix(packaging/gamescope): +pfhdr6 — a GAMESCOPE_NO_FOCUS window can no longer steal the composite
ci / web (pull_request) Successful in 1m3s
ci / rust-arm64 (pull_request) Successful in 1m35s
apple / swift (pull_request) Successful in 1m40s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m16s
ci / bun-nix (pull_request) Successful in 1m30s
ci / rust (pull_request) Successful in 4m52s
android / android (pull_request) Successful in 5m36s
Patch 0008: honor GAMESCOPE_NO_FOCUS in steamcompmgr's focus selection. hhd (Handheld Daemon) sets the atom once at init on its hhd-ui overlay window and never clears it; MangoHud sets it too; show/hide for these clients runs over the STEAM_OVERLAY protocol. NOTHING consumed the atom — not upstream gamescope, not Bazzite's fork (checked ba148 by strings) — so a mapped-but-unpainted hhd-ui window (it crash-loops under a headless punktfunk takeover and remaps on every respawn, stamping Steam's appid 769) was an ordinary focus candidate, and steamcompmgr picked it over Big Picture. The composite, and the stream fed from it, went black while every health signal stayed green: on .41 the client sat decoding 60 fps at 0.1 Mb/s of black, GAMESCOPE_FOCUSED_WINDOW named the hhd-ui window with GAMESCOPE_NO_FOCUS(CARDINAL)=1 on it, and killing hhd-ui brought the picture back the same second. The patch wires the atom exactly like GAMESCOPE_EXTERNAL_OVERLAY — read at map, PropertyNotify-tracked with MakeFocusDirty, skipped by both focus-candidate collectors (X11 and XDG) — and touches neither compositing nor appID, so a NO_FOCUS window still paints if the baselayer protocol brings it into view; it is only barred from being CHOSEN. Applies cleanly on the full 0001..0008 series from the bare 5fb8dce4 pin (verified with git am). Banner +pfhdr5 → +pfhdr6, PKGBUILD 3.16.25.pfhdr6-1; README gains the 0008 row, the missing +pfhdr5 ledger row, and the reconciled bump rule (a bugfix bumps the level only when field triage must read the difference off a box's banner — 0007's crash-loop, 0008's lost composite). |
||
|
|
23d0452157 |
feat(host): GameStream opt-in on every route; the native plane is deny(unsafe_code)-enforced
The user direction after WP0: ENet exists only for Moonlight, so the native
plane must be provably safe and the compat planes a deliberate choice.
Opt-in, everywhere. Windows already was (unchecked installer task). The three
opt-out surfaces are flipped: the shipped systemd user unit (deb/RPM/Arch/
sysext) no longer bakes --gamestream into ExecStart — a new
PUNKTFUNK_GAMESTREAM=1 host.env knob (pf-host-config, OR-ed with the CLI
flag) is the packaged opt-in; the NixOS module default goes true→false, with
a module-check assertion that unset = native-only; the Deck installer takes
--gamestream to opt in (--no-gamestream kept as explicit-off). Docs
(quickstart, running-as-a-service, moonlight, ubuntu/fedora/arch firewall
sections, gnome/sway, how-it-works) rewritten to the opt-in shape; the
CHANGELOG carries the upgrade note.
Enforced-safe. punktfunk-core is #![deny(unsafe_code)] crate-wide — every
module that parses network bytes is safe Rust as a compile error, not a
census result. Carve-outs are exactly two documented classes, neither of
which interprets attacker bytes: the client surface (abi, client) and the
transport syscall-batching shims (udp/{apple,linux,windows}, qos_windows).
In punktfunk-host, the modules a secure-default host exposes — native
(cfg-not-test: its tests exercise the client C ABI on purpose),
native_pairing, mgmt, mgmt_token, discovery, wol — are #[forbid(unsafe_code)].
Gates: Linux amd64 container clippy --all-targets -D warnings clean over
core+host-config+host; core 204 tests green under the deny; mgmt 46/46,
control 6/6. .133 Windows clippy (shipped features, clean-first,
sentinel-checked) clean — covers the qos_windows/udp-windows carve-outs.
macOS + iOS cargo check green (the apple.rs carve-out compiles for real).
|
||
|
|
1befa8a2c4 |
docs(nix): bring the Nix docs in line with the module, and fix a stale claim they shared
There are three places Nix is documented — the public docs-site, packaging/nix/
README.md, and packaging/README.md — plus the changelog. All had drifted.
STALE CLAIM, and not only for Nix. install.md said the plugin runner's "user unit
ships **disabled** — enable it once you have" something to run. That is true only
of Arch and source installs: the deb postinst and RPM %post both
`systemctl --global enable punktfunk-scripting.service`, and the Bazzite sysext
bakes in a default.target.wants symlink (build-sysext.sh:113). bazzite.md carried
the same claim about its own image. Both corrected, per channel, with the reason
the default flipped — the library scanners are plugins, so a host without the
runner can come up with an empty library — and the `mask`-not-`disable` opt-out
the sysext's own comment documents.
docs-site:
* install.md NixOS — `desktopSession` in the example and explained, the runner
no longer needs enabling, and the host/console line says what autoStart does.
* running-as-a-service.md — "Restart the host with your desktop" documented the
drop-in for packaged installs only; NixOS gets its one-liner beside it.
* bazzite.md — the runner is started for you, not "isn't started".
packaging/nix/README.md:
* option tables gain `desktopSession`, `gamescopeHdr`, `gamescopePackage`, and
the `punktfunk` group next to `input` (both are required — the udev rule
chgrp's the vhci nodes and fails outright if the group was never created).
* "what the module configures" gains the security.wrappers entry, and a note on
why the capability sits on the encode worker and never on the host: a wrapper
raises it into the ambient set, which lands it in the permitted set and fails
KWin's /proc/<pid>/exe readlink identically to a file capability.
* the appliance snippet no longer tells you to put pkgs.gamescope on PATH —
gamescopeHdr does that with the patched build, and desktopSession is called
out as the thing to leave off there.
* a caveat recording that `nix flake check` does not check the module, and the
two rules for editing module-check.nix (assertions stay pure Nix; assert
list-valued unit fields on the lists, not the rendered text).
packaging/README.md: the flake ships five packages, not "host + client".
CHANGELOG.md v0.27.0: a NixOS section covering the comm/session-detection fix, the
module changes including the scripting default flip as an explicit behaviour
change, and the flake-check gap — plus the documentation bullets above.
|
||
|
|
f8cde0adaf |
feat(nix): actually check the NixOS module in CI, and close the sweep's open issues
THE CI GAP. `nix flake check` does not check `nixosModules`. It forces the value
and asserts it is a lambda taking an open attribute set — nothing more; nix's own
source carries `// FIXME: if we have a 'nixpkgs' input, use it to check the
module.` Measured: a flake whose module sets a nonexistent OPTION, references a
nonexistent `pkgs` attribute AND calls a nonexistent `lib` function passes clean,
printing `checking NixOS module 'nixosModules.default'... all checks passed!`.
nix.yml's header claimed that leg covered the module; it never did, for the
module's whole life — on a flake whose history is Nix regressions reaching main
invisibly.
Closed with `checks.<system>.nixos-module` (packaging/nix/module-check.nix): it
evaluates the module against real nixpkgs in four scenarios (desktop, appliance,
native-only, client-only) and asserts on the rendered systemd units. The
assertions are PURE NIX so instantiating the check runs them — which means the
eval-only `--no-build` leg CI already runs is sufficient, and no Rust is built.
Stub fake-derivation packages keep it independent of punktfunk-host/-client and
the from-source gamescope; crane and bun2nix are provably not needed (they are
`throw`s in the wiring test and it still instantiates).
17 checks, including regression guards for every divergence the sweep found and
for the KWin identification trap (host ExecStart must stay on the plain store
path, never a capability wrapper, while the encode worker points AT the wrapper).
Mutation-tested: 8 mutants, each re-introducing one real defect, all 8 rejected,
baseline green. The suite already earned it once — its first run failed a correct
module because systemd renders `After=` as one space-separated line, so those
assertions now read the evaluated lists instead of the text.
Also closed from the sweep:
* services.punktfunk.host.desktopSession (new, default false) — binds the host
to graphical-session.target, the declarative form of the
punktfunk-host-desktop-session.conf drop-in. Without it a Plasma/GNOME
restart leaves the host holding a Wayland socket and portal D-Bus connection
that died with the old compositor: it still listens, still answers, and every
session it then serves fails at capture. Off by default because an appliance
may never reach that target and would be left permanently stopped.
* scripting.autoStart now defaults ON, matching the deb postinst and RPM %post,
which both `systemctl --global enable` the runner. It was opt-in here on the
reasoning that the runner is inert until you add automation — which stopped
being true when the game-library scanners became plugins. A NixOS host came up
with an empty library and no obvious reason why. The module and README carried
the superseded rationale verbatim; both updated.
* A warning when the host is enabled and xdg.portal is not. A warning rather
than `xdg.portal.enable = mkDefault true`, because enabling the portal service
with no `extraPortals` backend is its own broken state and only the operator
knows which backend their compositor needs.
* punktfunk-gamescope gets a `build-gamescope` dispatch input. It is on the
critical path of every host build (`gamescopeHdr` defaults true) yet nothing
compiled it; it tracks nixpkgs' gamescope, so a flake.lock bump — not a change
of ours — is what breaks it, and the first to find out would be an operator
whose system rebuild fails.
All .nix files reformatted with the flake's own declared formatter
(nixfmt-rfc-style from the PINNED nixpkgs, not a channel's).
|
||
|
|
159bbdbfc2 |
fix(nix): port three NixOS-module divergences from the shipped systemd units
A sweep of the Nix packaging against the units the deb/rpm actually install found three decisions that were made, documented and deliberate everywhere else, and simply not carried into packaging/nix/nixos-module.nix. punktfunk-web — StartLimitIntervalSec=0. The unit's EnvironmentFile for the mgmt token is mandatory ON PURPOSE, so the console genuinely fails until the host's first `serve` writes it. systemd's default rate limit (5 starts / 10 s) against RestartSec=2 then gives up permanently after ~10 s — which on an appliance is exactly the window before the host is ready, so a console enabled before the host's first run stayed dead until someone restarted it by hand. scripts/punktfunk-web.service has carried the override since that defect was found; the Nix module omitted it while its own comment went on promising "Restart retries until the host has created it". punktfunk-web — Restart=always, not on-failure. A console that exits 0 has still stopped serving, and on-failure leaves it down. Matches the shipped unit and web-run.cmd on Windows, both of which relaunch bun on ANY exit. An explicit `systemctl --user stop` is unaffected. punktfunk-scripting — the sandbox was missing entirely. The shipped unit confines the runner with NoNewPrivileges, ProtectSystem= strict, ReadWritePaths=%h /tmp and an AF_UNIX/AF_INET/AF_INET6 address-family restriction, plus PrivateTmp=no (a field report: a private /tmp hides /tmp/vhclient and /tmp/.X11-unix, so a plugin launches its vendor binary and then cannot reach the daemon behind it). The NixOS unit had none of it — so the one unit here that executes arbitrary operator TypeScript by design ran strictly LESS confined on NixOS than on every other channel. Verified by evaluating the module against the pinned nixpkgs and rendering the units: assertions clean, cap_sys_nice=ep on the encode-worker wrapper, firewall 47984/47989/47990/47992/47993/48010, and each unit carrying exactly the directives above. That evaluation is NOT something CI does — measured: `nix flake check` passes a nixosModule containing a nonexistent option, a nonexistent pkgs attribute and a nonexistent lib function, printing "checking NixOS module ... all checks passed!" while never evaluating it against nixpkgs. nix.yml's header claims that leg covers the module. It does not; tracked separately. |
||
|
|
35b5ee6a36 |
Merge pull request 'punktfunk-encode-worker: GPU priority via a capability-carrying worker, with WP3 on-glass complete' (#153) from worktree-worktree-encode-worker into main
audit / bun-audit (plugin-kit) (push) Successful in 20s
audit / bun-audit (web) (push) Failing after 20s
audit / bun-audit (sdk) (push) Successful in 20s
audit / pnpm-audit (push) Successful in 9s
audit / docs-site-audit (push) Successful in 20s
audit / cargo-audit (push) Successful in 1m9s
apple / swift (push) Successful in 1m42s
ci / web (push) Successful in 1m21s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m48s
ci / docs-site (push) Successful in 1m19s
ci / bun-nix (push) Successful in 17s
android / android (push) Canceled after 5m0s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 5m1s
ci / rust (push) Canceled after 4m17s
ci / rust-arm64 (push) Canceled after 4m8s
deb / build-publish (push) Canceled after 54s
deb / build-publish-host (push) Canceled after 0s
deb / build-publish-client-arm64 (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 3s
release / apple (push) Canceled after 3m58s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 1s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 0s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Canceled after 2m10s
windows / build (aarch64-pc-windows-msvc) (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 0s
decky / build-publish (push) Successful in 26s
audit / license-gate (push) Successful in 6m39s
windows-host / package (push) Successful in 13m21s
windows-host / winget-source (push) Skipped
nix / flake (push) Successful in 15m53s
windows-host / canary-manifest (push) Successful in 25s
Reviewed-on: #153 |
||
|
|
fdef4c90ce |
fix(gamescope): stop the PipeWire use-after-free that aborted a session on every connect
A managed gamescope session on Nobara 44 (VM 123) died on essentially every
client connect. The visible symptom was a black screen; underneath,
`punktfunk-gamescope` was SIGABRT crash-looping — 11 coredumps in three minutes
— until `gamescope-session-plus` ran out of retries and came up on the *stock*
`/usr/bin/gamescope` at its default 1920x1080, which looks like a working game
mode and carries none of our capture patches.
punktfunk-gamescope: ../src/pipewire.cpp:88: void destroy_buffer(
pipewire_buffer*): Assertion `false' failed.
#4 __assert_fail
#5 destroy_buffer(pipewire_buffer*).cold
The abort is a use-after-free wearing an `assert(false); // unreachable`.
`pw_buffer->user_data` is associated with its `pipewire_buffer` in exactly one
place, at the bottom of `stream_handle_add_buffer` — after all four `goto error`
paths, whose label is a bare `delete buffer`. And `stream_handle_remove_buffer`
clears `buffer->buffer`, the only route back to the `pw_buffer`, while a still-
`copying` buffer is deleted later on the steamcompmgr thread with no way to
reach the slot. PipeWire recycles `pw_buffer` slots across renegotiations, so
the next remove reads `buffer->type` out of freed memory, falls off the end of
the switch and aborts.
The host sets the session to the client's mode on connect, and that mode change
is what renegotiates the stream — which is why "every connect" was the trigger.
Patch 0007 fixes the association rather than the symptom: set `user_data` at
allocation so it is valid on every path out of `add_buffer` and clear it on the
error path; clear it in `remove_buffer`, the last point both halves are known;
null-check the two consumers. The `default:` arm then logs instead of aborting.
Offered upstream — nothing about it is punktfunk-specific.
Two traps this cost time on, both now written down in the README:
* It is NOT HDR-specific. The abort was first seen right after a 10-bit
stream negotiated, so `PUNKTFUNK_GAMESCOPE_HDR=0` looked like a workaround.
The failing argv carries no `--hdr-enabled` at all.
* `gamescope-session-plus` hides it by falling back to stock gamescope, so a
session existing proves nothing — read the banner.
`.pfhdrN` moves to 5 even though no capability moved: every deployed pfhdr4
binary crash-loops, so an operator needs to be able to tell them apart. All
`>=` thresholds in the host's probe are unaffected.
Also documents `libstdc++-static` as a build dependency — it is punktfunk's
requirement (the script links the C++ runtime statically on purpose), so no
`dnf builddep` will ever pull it, and without it meson fails with a message
naming neither the flag nor the package.
Verified on VM 123 with the patched binary installed: 5 rapid connect/
disconnect cycles plus 3 further sessions, zero new gamescope coredumps (43
before, 43 after), Steam game mode streaming real content at 5120x1440, and
`/tmp/chimeraos-short-session-tracker` never created — the short-session latch
that used to strand the box in plasma was downstream of this crash.
|
||
|
|
c7df7b45af |
fix(drivers/pf-gamepad): the three Xbox identities as a range — the driver clippy gate is red on main
ci / web (pull_request) Successful in 1m12s
ci / docs-site (pull_request) Successful in 1m46s
ci / bun-nix (pull_request) Successful in 37s
ci / rust-arm64 (pull_request) Successful in 3m49s
ci / rust (pull_request) Successful in 11m37s
windows-drivers / probe-and-proto (pull_request) Successful in 23s
windows-drivers / driver-build (pull_request) Successful in 1m43s
`cargo clippy --all-targets -- -D warnings` over the shipped drivers (the step that enforces the unsafe-audit gates) fails on main since #149 landed: clippy 1.96's `manual_range_patterns` fires on all five `4 | 5 | 6` device-type arms, and `-D warnings` turns each into an error, so `pf-gamepad` fails to compile as both lib and lib-test and the whole step never reaches the other five crates. Device types 4/5/6 are the Xbox Wireless / One S / Elite Series 2 identities added by #149 — contiguous by construction, so `4..=6` is the same set. Purely a lint fix: no arm gains or loses a device type, and the comments that already record *why* the three share one report shape, one descriptor and one vendor string are untouched. |
||
|
|
5d7091bf87 |
Merge pull request 'The Windows Xbox pad: make games actually see it' (#149) from worktree-xbox-pad-wgi-visibility into main
apple / swift (push) Successful in 1m35s
windows-drivers / probe-and-proto (push) Successful in 26s
windows-drivers / driver-build (push) Failing after 1m40s
android / android (push) Successful in 6m12s
ci / rust-arm64 (push) Successful in 5m5s
ci / bun-nix (push) Successful in 21s
ci / web (push) Successful in 5m31s
arch / build-publish (push) Successful in 8m28s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 13s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 9s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 9s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
ci / docs-site (push) Successful in 4m14s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
release / apple (push) Successful in 10m1s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Failing after 45s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 3m4s
docker / deploy-docs (push) Skipped
ci / rust (push) Successful in 12m22s
deb / build-publish-client-arm64 (push) Successful in 4m27s
deb / build-publish-host (push) Successful in 7m16s
docker / builders-arm64cross (push) Successful in 12s
deb / build-publish (push) Successful in 9m58s
apple / screenshots (push) Successful in 6m11s
windows-host / package (push) Canceled after 12m58s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
flatpak / build-publish (push) Successful in 9m46s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m48s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m50s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m23s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m31s
nix / flake (push) Failing after 21m34s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 24m43s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 25m26s
Reviewed-on: #149 |
||
|
|
fb309e0262 |
fix(pf-vdisplay): the takeover blamed polkit for a group it never named, and offered two remedies that cannot work
ci / bun-nix (pull_request) Successful in 33s
ci / web (pull_request) Successful in 1m11s
apple / swift (pull_request) Successful in 1m38s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 2m9s
ci / rust-arm64 (pull_request) Successful in 3m25s
android / android (pull_request) Successful in 4m28s
ci / rust (pull_request) Successful in 19m29s
Field triage on Nobara, 2026-08-09. Every connect degraded to ATTACH — which on that box mirrors a
game-mode session the host never configured, and looked like a black screen on every connect. The
host said:
the packaged pf-dm-helper polkit action is missing or was denied (reinstall the punktfunk
package, or install the display-manager polkit rule from the docs)
Every clause of that was wrong. The action was installed, `allow_any`, and its exec.path annotation
matched the installed helper; pkexec authorized it and RAN the helper. The helper refused, and said
exactly why:
pf-dm-helper: user 'nobara-user' is not in the 'punktfunk' group — refusing.
Grant it with: sudo usermod -aG punktfunk nobara-user (then re-login)
That text never reached the log, because `dm_helper` ran the helper with `.status()` — which
discards stderr and collapses the exit code to a bool. The one thing that would have ended the
investigation in seconds was thrown away at the call site, and the caller then guessed. Neither
suggested remedy adds anyone to a group, so a reader who followed both stayed broken and learned the
docs were useless. It fails soft, with no error and no failed unit, so nobody finds it on purpose.
Now: `.output()`, and four failure modes that stay distinguishable because they need different
fixes — helper not installed, pkexec could not run it, polkit denied it (pkexec's own 126/127), and
the helper ran and refused, whose stderr rides through VERBATIM rather than being re-described. Null
stdin too, so a pkexec that decides to prompt gets EOF instead of parking a stream thread on a tty
read.
The same gate gates the `linger` verb, so on a sessionless host an unjoined user fails there first —
carrying the reason there as well, or the misdiagnosis just moves one message earlier.
A new startup preflight says it before a stream is being built rather than during one, gated so it
cannot nag a box that would never attempt a takeover: not root, a display-manager alias exists, a
managed session launcher exists, a packaged helper exists, and the user is not in the group. It reads
membership from the user database rather than this process's groups, deliberately: that is what the
helper reads (it runs as root and resolves the caller from the database), so `usermod -aG` satisfies
the DM gate immediately and the warning stops. Using `getgroups()` would keep warning on a box where
the takeover already works.
Packaging said the group was for "the virtual Steam Deck pad (usbip)" — so anyone without a Deck pad
correctly skipped it and landed here by following instructions properly. All three scriptlets now
lead with Game Mode, name both grants, and record that creating the group is necessary and NOT
sufficient. Docs get the same treatment: the group is an admonition above the DM-flavor list in
gamescope.md, a black-screen entry in troubleshooting.md that tells the reader to read the quoted
reason FIRST, and the per-distro install pages no longer frame it as pad-only.
|
||
|
|
2b1843ed1c |
fix(drivers/pf-gamepad): the right stick is Z/Rz — as declared, it was dead
Found on glass, first real streaming session: everything worked except the right stick, and Steam correctly showed "Xbox One S Controller". `XBOX_RDESC` declared the right stick as `Rx`/`Ry`. `xinputhid`, which translates our HID collection into XUSB, maps `Z`/`Rz` to the right stick and does not treat `Rx`/`Ry` as one, so those two axes reached nothing. Two usage bytes. Left and right were declared identically here — same collection, same globals, same size and count — so the usages were the entire difference, which is what makes the diagnosis airtight rather than plausible. Note `DUALSENSE_RDESC`, a real capture, also uses `Z`/`Rz` for its right stick and puts the TRIGGERS on `Rx`/`Ry`; that is most likely where the original mistake came from. ⚠️ Byte offsets are unchanged — still 16×2 at bit 5.0 — so `xbox_proto`'s layout tests and the host-side packing are untouched. This is a pure relabelling. 🛑 THE REAL LESSON IS THE HARNESS, AND IT IS FIXED HERE TOO. This survived every bench measurement because `dualsense-windows-test` drove LS-X and the A button and left the other five analogue axes at zero. `XInputGetState` read `RX [0..0]`, which I read as "the devtest doesn't move it" — true, and useless: a harness that exercises one axis cannot tell "this axis is not mapped" from "nothing is driving it", and the two are indistinguishable in every consumer. The devtest now sweeps all six axes on distinct phases and ramps both triggers, so one run shows which axes arrive AND that they are not crosstalking onto each other's bytes. MEASURED ON .173, same run shape before and after, devtest sweeping all six axes: before: LX [-11264..24576] LY [-32768..31744] RX [0..0] RY [-1..-1] LT [0..248] RT [7..255] after: LX [-8192..26624] LY [-32768..31744] RX [-32768..31744] RY [-24576..10240] LT [0..248] RT [7..255] VERIFIED * `cargo test -p pf-inject --lib` 104/104 on Windows; `xbox` subset 11/11 on macOS — the layout tests still pass because nothing moved. * Driver rebuilds and signs; the descriptor is still 223 bytes so the `wReportLength` const assert is undisturbed. * `cargo fmt --all --check` clean. NOT VERIFIED * Not yet re-tested in a real streaming session — that is the next on-glass run. * ⚠️ A leftover finding from the same session, unrelated to this fix and NOT investigated: the session's pad devnode SURVIVES client disconnect and keeps the `Global\pfds-boot-0` bootstrap mailbox, so a devtest run afterwards fails with `Zugriff verweigert (0x80070005)` and silently measures the stale pad instead. Restarting the service releases it. Worth its own look. |
||
|
|
4f9071b980 |
feat(pads/windows): three Xbox identities — Wireless, One S and Elite Series 2
Until now there was one Xbox identity, `device_type = 4` / `045E:0B13`, and Windows folded a client's `XboxOne` request onto it because the only Windows Xbox backend was the XUSB companion, which presents one fixed 360 identity and cannot vary it. The HID backend can, so the fold goes and two identities join it: devtype 4 045E:0B13 pf_xboxwireless Xbox Wireless Controller devtype 5 045E:02FD pf_xboxones Xbox Wireless Controller (One S) devtype 6 045E:0B22 pf_xboxelite Xbox Elite Wireless Controller Series 2 `GamepadPref::XboxElite` takes wire byte 11 — the first unassigned one, and the round-trip test previously asserted `from_u8(11) == Auto` with a comment saying assigning it must update that; the sentinel moved to 12. The C ABI mirror and the generated header moved with it. ⭐ ALL THREE SHARE ONE REPORT DESCRIPTOR, deliberately. In HID terms they are the same pad; the descriptor is the report shape, not the identity. §3 of the handoff records that our single hand-written descriptor already cost three separate bugs, and inventing two more would multiply that debt for no measured gain. They differ in VID/PID, product string, hardware id and Device Manager description only. ⚠️ All three install `pfGamepadXbox`, the section that attaches the `xinputhid` bus filter. That was the open risk: Microsoft's `xinputhid.inf` promotes by an explicit hardware-id allow-list containing `02D1, 02DD, 02E3, 02EA, 0B00, 0B0A, 0B13, 02FF` — and NEITHER `02FD` NOR `0B22` is on it. Measured on .173: promotion does not care, because it comes from our own AddReg rather than from matching Microsoft's ids. All three gain `IG_00`, register an XUSB interface, and are read live by classic XInput. Had this gone the other way the two new identities would have been strictly worse than the one they joined. The XUSB escape hatch needed a runtime degrade to stay honest. `pick_gamepad` is compile-time only, so with `PUNKTFUNK_XBOX_BACKEND=xusb` the host would have resolved and echoed `xboxelite` in its `Welcome` while actually building a 360 pad. `degrade_xbox_identity` folds the identity back at runtime, mirroring `degrade_if_no_uhid`. VERIFIED ON WINDOWS (.173 — none of this compiles on macOS; the driver needs the WDK and the rest is `cfg(windows)`): * `cargo test -p pf-inject --lib` 104/104 — including `hwid_matches_inf`, `hwid_devtype_table_matches_the_driver` and `only_the_xbox_identity_installs_the_xinputhid_section`, all now sweeping the whole identity set and asserting the section split in both directions. * `cargo test -p punktfunk-core --lib gamepad` 7/7; `cargo check -p punktfunk-host` clean. * Driver builds and signs; the descriptor/`wReportLength` const asserts still hold with the descriptor shared three ways. * ON GLASS, per identity, via the new `--xboxones` / `--xboxelite` devtest legs: each gets its own devnode (`PF_XBOX_0` / `PF_XBOX_ONES_0` / `PF_XBOX_ELITE_0`), each HID child gains `IG_00`, each registers an XUSB interface, and XInput reads each live (packets advancing, `buttons=0x1000`). * macOS: `cargo fmt --all --check` clean in both workspaces. NOT VERIFIED / NOT DONE * **Elite paddles are NOT implemented.** `BTN_PADDLE1..4` would need descriptor buttons, and once `xinputhid` promotes the pad it claims the HID collection exclusively — XInput has no paddle fields and the HID consumers that do may be locked out, so the buttons would likely reach nobody. The decisive measurement is cheap and named in the code: hold a paddle bit set and see whether a user-mode HID reader still gets reports. Until then the Edge remains the only virtual pad with native back-button slots and nothing should be advertised otherwise. * **No client picker offers the Elite**, and none can auto-detect it — SDL3's `GamepadType` has no Elite variant. It is reachable today only via `PUNKTFUNK_GAMEPAD=xboxelite` or a hand-edited client setting. All five clients ship the same curated six options by deliberate parity, so adding one is a cross-client UX change, not part of this. * Nothing here has run in a real streaming session; every measurement came from the devtest. |
||
|
|
1317901122 |
Merge pull request 'Uninstalling the Windows host left every audio device it minted behind forever — and the installer script documented that as a decision' (#145) from worktree-win-audio-uninstall-cleanup into main
android / android (push) Failing after 1m32s
ci / rust-arm64 (push) Successful in 1m53s
apple / swift (push) Successful in 1m34s
ci / bun-nix (push) Successful in 22s
ci / web (push) Successful in 1m48s
ci / docs-site (push) Successful in 1m44s
deb / build-publish-client-arm64 (push) Successful in 1m44s
deb / build-publish (push) Successful in 4m5s
ci / rust (push) Successful in 7m24s
apple / screenshots (push) Successful in 5m54s
arch / build-publish (push) Successful in 9m40s
deb / build-publish-host (push) Successful in 7m28s
windows-host / package (push) Successful in 13m52s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 20s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m53s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 18m0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 11s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 7s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 7s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 6s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 6s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / builders-arm64cross (push) Successful in 7s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m4s
docker / deploy-docs (push) Failing after 6m14s
Reviewed-on: #145 |
||
|
|
d87a8df28d |
fix(windows): uninstall removes the audio devices the host mints
ci / bun-nix (pull_request) Successful in 25s
ci / web (pull_request) Successful in 1m21s
ci / docs-site (pull_request) Successful in 1m23s
apple / swift (pull_request) Successful in 1m42s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m37s
android / android (pull_request) Successful in 6m12s
ci / rust (pull_request) Successful in 7m37s
The field report: uninstalling punktfunk left "Punktfunk Speakers", "Punktfunk Microphone" and the per-pad "Wireless Controller" endpoints sitting in Windows' Sound settings forever. They have no installer payload behind them, which is why nothing in the uninstall touched them. The host mints them at RUNTIME as extra devnodes on Valve's streaming-audio drivers, and both providers deliberately re-resolve their devnode across restarts instead of re-minting it — so they persist by design. Persistent across restarts must not mean permanent: the .iss even documented leaving them behind as a decision. New `driver uninstall --audio` leg (a third Inno [UninstallRun] entry, after the two driver legs and well after `service uninstall`, since a live host re-mints on its next wiring pass): * restores the default playback device first, if a host that died mid-stream left it parked on our loopback sink — otherwise Windows re-picks by its own ranking rather than giving the operator back the device they had; * removes every MEDIA-class devnode carrying one of our three durable owner markers (pad slot, minted role, probe), phantoms included; * deletes each endpoint's MMDevices record, resolved through the devnode link BEFORE the devnode goes. Marker-matched, never name-matched: our instances are name-identical to Steam's own, and Steam's devnodes, its drivers, and a VB-CABLE from the era when we bundled one carry no marker and stay untouched. A ROOT\ enumeration guard means a marker-shaped value on a real sound card can never cost the user their hardware. The registry half is best-effort: those keys are SYSTEM-owned and the uninstaller runs elevated but as a user, so on a stock box the record survives as an inert NOTPRESENT entry that Sound settings only shows behind "Show Disconnected Devices". The device itself is gone either way, and seizing ownership of SYSTEM registry keys from an uninstaller is a worse thing to ship than that scrap. |
||
|
|
77f0a25d18 |
feat(drivers/pf-gamepad): ship the xinputhid bus filter, so Windows finally promotes our Xbox pad
The field report that started this work was an Xbox controller that no game could see on a Windows host for two weeks. Root cause was that our Xbox pad reaches no Windows input API a modern title uses. This is the fix, and it is two registry values. Windows promotes Xbox pads with `xinputhid`, whose INF is an explicit hardware-id ALLOW-LIST — its own comment says "we can not use a Compatability ID for the loading of this driver, and so rely on individual hardware IDs". A software-enumerated devnode can never match those ids, so we write what the matching install sections would have written. `045E:0B13`, the PID this identity already claimed, is on that allow-list twice, so the identity choice turned out to be exactly right. 🛑 THE PAIRING IS THE WHOLE FINDING, AND THE TWO VALUES GO IN DIFFERENT KEYS. `UpperFilters` is a `.HW` AddReg (hardware key); `DevicePropertyFlags` is a DDInstall AddReg (software key). A live A/B on .173: removing `DevicePropertyFlags` alone reverts EVERYTHING — no `IG_00`, no XUSB interface, no XInput, no WGI entry — while `UpperFilters` alone is completely inert. `1` = `BusDevice`, which Microsoft glosses as "a focused bus filter driver for the IG_ problem". It is not a description of the device, it is the switch. An earlier session installed the filter WITHOUT it, measured a device that produced nothing, and recorded "never ship it". The filter was never broken; it had never been switched on. That conclusion is now retracted. ⚠️ The Xbox line gets its OWN DDInstall section, `pfGamepadXbox`. All five identities previously shared `pfGamepad`, so an AddReg there would have handed a DualSense, DualShock 4, Edge and Steam Deck to Microsoft's Xbox translator. The regression check below exists for exactly that. MEASURED ON .173 (Win11 26200), INF-SHIPPED — no hand-written registry values: * `UpperFilters=xinputhid` lands on the hardware key and `DevicePropertyFlags=1` on the software key, applied by the INF at install. * The HID child gains the `IG_00` token: `HID\PUNKTFUNK&IG_00\...`. * An XUSB interface appears: `\\?\hid#punktfunk&ig_00#...#{ec87f1e3-...}`. * classic XInput reads it live — packets ADVANCING, `buttons=0x1000` (the devtest's A), and the stick sweeping. XInput had NEVER seen this backend before. * `XInputSetState` rumble round-trips: `rumble from game: pad=0 low=65535 high=32767`. * REGRESSION CHECK PASSED: with the DualSense identity up, its devnode has an EMPTY `UpperFilters` and no `DevicePropertyFlags`. The PlayStation pads are untouched. WGI `Gamepad` lists the pad but reads `ts=0`. That is NOT ours: a real Xbox Elite Series 2, promoted by Microsoft's own driver on the same box, reads `ts=0` in WGI at the very moment classic XInput is reading live data from it (`buttons=0x1000 LY=-32768`). Our pad is behaviourally indistinguishable from real hardware here; the row is a property of the non-interactive session. NOT VERIFIED * On-glass in a console session. Everything above ran over ssh, which is what makes the WGI row unreadable; the real-Elite control is what settles it, not a clean WGI reading. * GameInput — no binding in the `windows` crate, still unmeasured for this backend. * `PUNKTFUNK_XBOX_BACKEND` still defaults to XUSB. This changes what the HID backend CAN do; it does not change which backend is chosen. That is WP-E and it is a separate decision. * Trigger-actuator enable bits, still conjecture — `XINPUT_VIBRATION` has two members and cannot exercise them. |
||
|
|
f9fe496dbc |
feat(drivers/pf-gamepad): declare the rumble output report, and the Xbox pad gets rumble at all
`XBOX_RDESC` declared no OUTPUT item — zero `0x91` bytes. hidclass routes an output report only if
the descriptor declares one, so `on_output_report` never fired, `publish_output` never wrote the
out-ring, and `parse_xbox_output` in `inject/windows/xbox_windows.rs` was unreachable code. The
entire host-side rumble plane was already built, wired and tested, and was simply never fed. The
HID Xbox pad therefore had NO rumble whatsoever, not merely no trigger rumble.
This appends the PID-page `Set Effect Report` collection, report id `0x03`, 8 payload bytes, sized
to exactly the layout `parse_xbox_output` and `design/trigger-rumble-plane.md` §2.1 already
specify. It is declared AFTER the final Input item and re-states every global it uses, so the
16-byte input layout `xbox_proto`'s tests pin is untouched.
⚠️ PROVENANCE: hand-written, and it could not be otherwise. The Elite capture taken for WP-A
reports `OUTPUT items: 0` — Windows exposes no literal descriptor bytes and the reconstruction
carries no output collection for that pad — so there was nothing to copy. The comment says so and
asks for a Linux hidraw capture to replace it.
Also adds a compile-time assert pairing every descriptor with its HID-descriptor `wReportLength`.
Those are two copies of one length, edited in different places, and a mismatch fails SILENTLY:
hidclass asks for `wReportLength` bytes, parses whatever it got, and the pad either enumerates
truncated or not at all with nothing naming the cause. It now cannot build out of step. This
caught nothing today because I updated both by hand, but it is exactly the trap this descriptor
has already sprung twice in other forms.
MEASURED ON .173 (Win11 26200), with the pad promoted via the WP-B0 xinputhid bus-filter config:
* `XInputSetState(0xFFFF, 0x8000)` produced, on the host side,
`rumble from game: pad=0 low=65535 high=32767`
`rumble from game: pad=0 low=0 high=0`
i.e. XInputSetState -> xinputhid -> HID output report 0x03 -> on_output_report -> out-ring ->
parse_xbox_output -> PadFeedback. First rumble this backend has ever delivered.
* The round-trip values confirm the descriptor's `Logical Maximum (100)` percent domain is
right: 0x8000 -> 50% -> 32767. A 0..255 domain would have produced different numbers.
* This also answers `trigger-rumble-plane.md`'s WP0 gate — YES, Windows writes output reports
to a synthesized 045E:0B13 — which was blocking the whole trigger plane.
* classic XInput reads the pad fully: packets advancing, `buttons=0x1000` (the devtest's A), and
`LX [-32768..31744]`, the complete sweep. LY/RX/RY frozen is correct; the devtest drives only
LS-X and A.
VERIFIED
* `cargo test -p pf-inject --lib xbox` 11/11 — the input layout is byte-identical, as intended.
* `hid-descriptor-dump --rust-source ... --symbol XBOX_RDESC` decodes it clean: input report
0x01 unchanged at 16 bytes and the same offsets, new output report 0x03 at 9 bytes on the
wire, feature 0x85 unchanged, `structure: OK`.
* Driver builds and signs on .173 with the WDK; the new const asserts compile, so all five
descriptor/wReportLength pairs agree.
* fmt clean on both tools; .173 fully reverted afterwards.
NOT VERIFIED
* The enable-mask bit assignments for the two TRIGGER actuators. `XINPUT_VIBRATION` has only two
members, so XInput can never drive them and this run could not exercise them. Still open, as
trigger-rumble-plane.md WP0 says.
* That this equals the real pad's output collection, byte for byte. Needs Linux hidraw.
* Nothing about the INF is changed: `pf_gamepad.inx` still has no AddReg, so none of the
promotion config ships. The rumble descriptor is inert until something drives it.
|
||
|
|
ae35e8b4d7 |
test(tools): capture the real Xbox descriptor, because ours was invented and disagrees with it
`XBOX_RDESC` is the only report descriptor in `pf-gamepad` that was hand-written rather than
captured off hardware, and its own provenance warning has now come true three times. The fix for
that class of bug is not another careful reading — it is a tool that goes and asks the device.
`tools/hid-descriptor-dump` does that: it dumps a real HID device's report descriptor, decodes it
into an annotated item listing plus a bit-offset LAYOUT TABLE, and can decode a blob we already
ship through the same decoder (`--rust-source <file> --symbol <NAME>`) so the two are diffable
line for line. `--read N` pulls live wire bytes, which is the only ground truth a reconstructed
descriptor cannot give you.
Deliberately NOT a workspace member — it pulls `hidapi`, a C library wanting libudev on Linux,
which has no business in `cargo build --workspace` or on a CI leg with no pad attached. It is a
bring-your-own-hardware tool and it is excluded in the root manifest, so CI never sees it.
The captured Elite disagrees with our blob in four ways, and the dangerous one is field ORDER:
the real pad reports sticks, ONE combined 16-bit Z trigger, then BUTTONS, then the hat, in an
UNNUMBERED 15-byte report; ours declares Report ID 1, two Simulation-page trigger axes, then the
hat, then 15 buttons. Since we claim a genuine Microsoft VID/PID and SDL/Steam/Windows all apply
stock mappings keyed on it, that ordering difference is exactly how every control silently lands
on the wrong action. The driver comment now records the diff and the two blockers that stop the
capture from simply being pasted in.
VERIFIED
* `cargo fmt --check` clean, `cargo clippy --all-targets -- -D warnings` clean (macOS).
* The tool builds and runs on macOS and on .173 (Windows 11 26200, cargo 1.96, MSVC, no WDK).
* TOOL VALIDATED AGAINST A KNOWN-GOOD CONTROL: pointed at the live DualSense on .173, it
reproduces the real `DUALSENSE_RDESC` layout exactly (input 0x01, 64 B, X,Y,Z,Rz,Rx,Ry at
bytes 1..6, hat 8.0, 15 buttons 8.4, output 0x02, the full feature ladder), and `--read`
returned live len=64 reports with sticks centred at 80 80 80 80 and the counter incrementing.
* `cargo metadata` on the root workspace still resolves and does NOT list this crate.
* The Elite capture is reproducible: `--vid 045E --pid 0B22`.
NOT VERIFIED
* That the capture equals the pad's NATIVE report map. Windows exposes no API for a device's
literal descriptor bytes, so hidapi reconstructs from `HidD_GetPreparsedData` — faithful in
structure, item order and bit offsets, not byte-exact (measured: the DualSense's real 273-byte
descriptor reconstructs to 467). `xinputhid` also filters that pad, and the captured shape is
the legacy DirectInput view. A byte-exact answer needs Linux hidraw.
* Why the Elite returned ZERO input reports across two runs (72 s and 90 s) while the DualSense
streamed fine on the same code path — untouched pad, or exclusive claim by the XInput
translator. Unresolved.
* Nothing here was built on Windows as a driver: `XBOX_RDESC` itself is UNCHANGED, so no
behaviour changes. The only edit to the driver is its provenance comment.
|
||
|
|
f34acf1d73 |
fix(drivers/pf-gamepad): the Xbox descriptor never declared the channel-proof report, so the pad served neutral forever
`XBOX_RDESC` declared only Input report 1. The sealed pad channel delivers its DATA section over a vendor Feature report `0x85` (`ProofTransport::HidFeatureReport`), and the proof handler's own comment records the assumption that made this invisible — "0x85 is already declared as a Feature report in all three captured descriptors". True of the captured PlayStation blobs; false of this hand-constructed one. So hidclass rejected the host's `HidD_GetFeature` before the driver ever saw it, the host refused to hand over the section, and the pad answered every read with its neutral report. The HID Xbox pad had never delivered a single input report since it was written. Declaring `0x85` with a 63-byte payload (1 id + 63 = 64 = FeatureReportByteLength) fixes it. Verified on glass on .173: `gamepad driver attached to the shared section proto=3 late=false`, and WGI's RawGameController path then reads the pad live — advancing timestamps, the devtest's left-stick sweep, buttons toggling. Before the fix: 12 consecutive samples, one frozen timestamp, every axis at dead centre. This is the descriptor-provenance warning in this file coming true. It is still CONSTRUCTED rather than captured, and that remains the open risk — `xinputhid` appears to validate the descriptor and refuses ours, and a real Elite is a multi-collection device where ours has one. Codec layout tests still 11/11; fmt clean. Only device_type 4 is affected, which nothing shipping uses yet. |
||
|
|
bc9201d136 |
fix(packaging/gamescope): bump the pin past upstream's capture-format probe, and sign the RPM
Three things, one delivery path — a Fedora/Nobara box getting the patched gamescope. **The pin moves 8c676c39 -> 5fb8dce4** (3.16.25-1 -> 3.16.25-11). The commit that matters is ff6b924, `rendervulkan: fall back to XBGR2101010 when XRGB2101010 is unsupported`: it probes `linearTilingFeatures` for STORAGE+SAMPLED and captures as XBGR2101010 where A2R10G10B10 linear storage is unavailable — which is every NVIDIA. That covers the paths that are upstream's rather than ours: the RGB intermediate `paint_pipewire()` acquires when the stream is YCbCr, and AVIF screenshots. #143 fixed our own node host-side; this is the other half, and its commit message asked for exactly this bump. All six patches rebased. Only 0006 conflicted: upstream's f8be7ee added `vulkan_has_drm_modifiers_for_features()` immediately above the `g_device` declaration our patch turns into a reference — both kept. 0003 and 0005 come out byte-identical; 0006 also picks up the `--zero-commit --no-signature` form 0001-0005 already used. **Patch 0001 now offers `xBGR_210LE` BEFORE `xRGB_210LE`**, mirroring the host-side `HDR_FORMAT_ORDER` rationale on the producer end. A consumer takes the first pod it can use, and we were handing third-party consumers (OBS and friends) the one format NVIDIA fills byte-reversed under a correct-looking label. Deliberately NOT done by calling upstream's `vulkan_get_rgb10_capture_format()`, which is what pw_pods.rs proposes: that symbol landed after 3.16.25, so it would break `packaging/nix/gamescope.nix` — which applies these patches to whatever gamescope nixpkgs pins — with an opaque C++ error instead of a patch conflict. The reorder gets the same outcome on any base. Note added there so the next reader does not "fix" it. **And the RPM was never signed.** `Sign RPMs` runs right after `Build RPM`; the gamescope RPM is built ~90 steps later, behind its own ~10-minute cache, so it missed the signing pass entirely — every punktfunk-gamescope RPM ever published went out unsigned. The repo file we tell users to install carries `gpgcheck=1`, so `dnf install punktfunk-gamescope` failed with "The package is not signed" on every Fedora and Nobara box. The package was in the channel the whole time and could not be installed from it, which is worse than absent: the notes and the docs-site both say it is there. `sign-rpms.sh` now takes explicit paths (defaulting to `dist/*.rpm` as before) and a second pass signs this one before publish, fail-closed on a tag like the first. Verified on Nobara 44 (VM 123, RTX 5070 Ti passthrough), canary 0.27.0-0.ci12611.g516a2954: * Builds clean in the fc44 CI image; banner `3.16.25-17-ga87390d+pfhdr4` (11 upstream + our 6), so the marker the host probes still reads 4 — no capability moved, hence pkgrel 3 and `.pfhdrN` staying put. * `pw-cli enum-params` on the live node: BGRx, NV12, **xBGR_210LE (81), xRGB_210LE (80)** — 8-bit consumers still negotiate bit-for-bit, 10-bit now leads with the safe one. * All four patched flags present, `--pipewire-composite-external-overlay` included. * Patch 0006 confirmed working by comparison, which is the only way to see it: the new build exits 0 where both the pre-0006 `+pfhdr2` build and the stock 3.16.23.2 abort with 134. * Signing fix proven with a throwaway key: `Signature: (none)` -> `digests signatures OK`. * Host health on the canary: synthetic spike 300/300 encoded, loopback 300 recovered, 0 mismatches. One unexplained one-off: the very first headless run after install segfaulted at exit (SIGSEGV, after "Primary child shut down!"). Not reproduced in 11 subsequent runs across every flag combination, so it is recorded rather than diagnosed — the binary is stripped and there is no symbolised core. |
||
|
|
d2a2bcc25d |
feat(drivers/pf-xusb): answer the async input wait, and put xinputhid on the stack
The two things this driver's README has always listed as the missing WGI/GameInput work, both user-mode, neither needing a bus driver: `IOCTL_XUSB_WAIT_FOR_INPUT` is now pended on a manual queue and completed by the periodic timer on a dwPacketNumber edge, answering with the same 29-byte GET_STATE payload the synchronous path serves. Declining it was enough for classic xinput1_4, which just falls back to sync GET_STATE polling — that is why the pad has always worked there. It is not enough for WGI/GameInput, which poll asynchronously: to them a decline is a refusal, not a fallback. Completion is edge-gated because releasing a waiter on an unchanged packet spins its caller at timer rate. WAIT_GUIDE_BUTTON stays declined — we have no state to signal on. The INF adds UpperFilters=xinputhid on the XUSB devnode. Note the earlier attempt put that filter on the HID child of the *other* backend, which was simply the wrong devnode: XInput does not read HID at all, it enumerates GUID_DEVINTERFACE_XUSB, which is what this driver registers. Verified on .173: build + sign + catalog exit 0; infverif "INF is VALID"; the devnode starts Status OK with UpperFilters=xinputhid readable back from its enum key; and XInput still sees the pad (slot 1 live alongside the box's real Elite in slot 0), so the async queue is no regression to the path that already worked. NOT yet measured: whether WGI/GameInput now admit the pad. `IG_` is the wrong probe for this driver — it is a HID-path artifact and pf-xusb is System-class with no HID child, so its absence says nothing either way. That needs a real WinRT/GameInput enumeration test. |
||
|
|
0ab17ee81d |
fix(packaging): a post_merge step added in a release was unreachable forever
ci / bun-nix (pull_request) Successful in 48s
ci / docs-site (pull_request) Successful in 1m20s
ci / web (pull_request) Successful in 1m16s
apple / swift (pull_request) Successful in 1m45s
ci / rust-arm64 (pull_request) Successful in 1m43s
apple / screenshots (pull_request) Skipped
android / android (pull_request) Successful in 3m52s
ci / rust (pull_request) Failing after 8m7s
A sysext upgrade is driven by the script from the OLD image -- /usr/bin/punktfunk-sysext
is replaced by the very `systemd-sysext refresh` that runs mid-upgrade -- so a
post_merge step ADDED in the new release is executed by nobody. The old script
does not have it, and the new script never gets a turn: from then on `update`
matches the "already on $cur" branch and returns before post_merge. The step is
permanently unreachable on exactly the installs that need it, and nothing says so.
Field-proven on the Bazzite host that took 0.25.0 -> 0.26.0 (2026-08-09). The
casualty was the `punktfunk` group, which post_merge learned to create in 0.26.0
(
|
||
|
|
d498ff4a60 |
test(drivers): give the Xbox identity a root-enumerated id, and verify the whole thing on Windows
`root\pf_xboxwireless` alongside the plain id, mirroring the DualSense model line — the INF already documents that variant as the one devgen/devcon tests bind, and without it the Xbox identity could only be exercised through a running host. Verified end to end on .173 (Windows 11 26200, WDK 10.0.26100.0): - build-gamepad-drivers.ps1 builds + signs + catalogs the driver, exit 0 - infverif /v /w on the generated pf_gamepad.inf: "INF is VALID" - pnputil stages the package; devgen creates the devnode; it starts clean: Status OK, Class HIDClass, "Punktfunk Virtual Xbox Wireless Controller" - it enumerates a HID child, Status OK, carrying HID_DEVICE_SYSTEM_GAME and HID_DEVICE_UP:0001_U:0005 — Windows parsed the constructed report descriptor and classified the pad as a Game Pad (usage page 0x01, usage 0x05), which is precisely what pf-xusb could never do Test devnode, phantom child, driver package and both certs were removed afterwards. Two build gotchas worth knowing, both already handled inside build-gamepad-drivers.ps1 and both of which cost a cycle here: CARGO_TARGET_DIR pointing outside the workspace breaks wdk-sys (wdk-build walks up from OUT_DIR looking for a Cargo.lock and finds none), and the WDK version must be pinned via Version_Number=10.0.26100.0 or bindgen picks SDK 10.0.28000.0, which ships no km/crt headers. Still open: the SwDeviceCreate USB identity (HID\VID_045E&PID_0B13) cannot be checked through a devgen node, which has no USB hardware ids — that needs the host path. So the WGI-promotion question is still unanswered, and host routing is still unwritten. |
||
|
|
4f8cce6751 |
feat(packaging): grant CAP_SYS_NICE to the encode worker on all six channels, and assert the host never gets it
767e67ca's per-channel mechanics were correct; they were aimed at the wrong binary. Each one is restored here pointed at punktfunk-encode-worker, and every host-side removal from #136 stays verbatim. All grants remain best-effort — an uncapped worker still encodes, at default priority, so a failed setcap must never fail an install. * Arch: setcap in post_install AND post_upgrade (a replaced binary is a new inode). * RPM: %caps(cap_sys_nice=ep) in %files, never a %post setcap — %caps applies, restores and verifies, and covers Fedora as well as Bazzite via rpm-ostree layering. * Bazzite + Arch sysext: setcap on the staging tree before mksquashfs, which does record security.capability. The assertion is amended, not removed: host EMPTY is still a hard fail, and the worker must carry exactly cap_sys_nice=ep — missing is fine, anything else is not. * deb: setcap in postinst. * NixOS: security.wrappers for the WORKER plus PUNKTFUNK_ENCODE_WORKER in the unit. A file capability cannot live on a store path, and an ambient grant is right here precisely because nothing ever identifies the worker. The host's ExecStart stays on the store path. * Steam Deck: setcap the worker; the .desktop the script writes stays valid this time. Four things the plan's channel table missed: * packaging/arch/build-sysext.sh had no capability handling at all, and a sysext can never run a pacman scriptlet — the SteamOS image would have shipped the lever permanently inert. * scripts/steamdeck/update.sh had none either. It rebuilds both binaries, so a new inode drops the grant, and it is the documented steady-state path: the lever would have died on the first update. It also never healed a Deck already capped by 0.26.0-1. * A capped worker is AT_SECURE, and glibc drops $ORIGIN-expanded RPATH entries for secure binaries unless they normalise into a trusted system dir. Copying the host's rpath under BUNDLE_FFMPEG=1 would have left the capped worker unable to find libavcodec on exactly the channel that bundles it. Absolute DT_RPATH instead. * Nix crane scopes by -p, so the worker would not have been built at all, and it needs its own addDriverRunpath. scripts/ci/assert-cap-matrix.sh mechanizes the lesson from 0.26.0-1 — verify the PACKAGE, never the board. It unpacks the built Arch package, the deb, the rpm and the mounted sysext raw and asserts one matrix: the host carries NOTHING (hard fail), the worker exactly cap_sys_nice=ep. The sysext reader first proves it can round-trip a capability through mksquashfs/unsquashfs at all, so an unreadable artifact fails rather than issuing a blind PASS, and --self-test red-teams the assertions themselves. Red-teaming the leg found a real bug: setcap originally ran BEFORE the assertion, so "the worker arrived carrying something unexpected" was unreachable and a stray %caps would have been silently overwritten. Both sysext scripts now assert, then grant, then assert again. |