forked from unom/punktfunk
790db5edbbd26088bb755684f32e00a08fc71278
223
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
790db5edbb |
Merge pull request 'The cache signing key is installed, and its DNS was never a dashboard click' (#318) from worktree-nix-binary-cache into main
Reviewed-on: unom/punktfunk#318 |
||
|
|
7e4fe80793 |
feat(nix): install the cache signing key and correct how its ingress is provisioned
Two corrections and one thing actually done. DNS here is not a dashboard click. unom/infra owns the unom.io zone in OpenTofu (terraform/cloudflare/records.tf, applied by dns-cutover.yml), and that file's `local.hostnames` set carries its own invariant: "a name here with no vhost 404s, a vhost with no name here never cuts over." A record added by hand in Cloudflare is out-of-band and risks the duplicate-record round-robin the file documents a few lines further down — the same class of trap as hand-editing ~/caddy/Caddyfile on the box. The setup steps said "in the unom.io Cloudflare zone" as though it were a manual change; they now name both files, the workflow that applies them, and the one-added-record check to expect from `plan`. unom/infra#20 makes the change. The signing key is generated and `NIX_CACHE_SIGNING_KEY` is installed as a repo Actions secret, so its public half is no longer a placeholder: punktfunk-cache-1:yhOJmHxzg6tzXpxSFzlYn6Pc6r0jHprsWqt8MZC654o= pinned in both docs. The publish step still writes the same value to /punktfunk-cache.pub, so the docs can always be checked against the cache itself — and the wizard now compares the two and warns on a mismatch, because docs that disagree with the cache mean users reject everything it serves. The wizard drops to four stages. DNS and the vhost were separate stages when they looked like separate manual steps; they are one PR against one repo, so they are one stage. The key stage now detects the installed key, prints it, and refuses to casually regenerate — a new key invalidates every signature already published and breaks every user pinning the old one. Verified: shellcheck + `bash -n` clean, 4 stages against TOTAL_STAGES=4, and the already-installed path's key extraction tested against the real README. |
||
|
|
c4cf53c1fc |
Merge pull request 'The Nix cache's setup steps pointed at a home-lab proxy that no longer exists' (#316) from worktree-nix-binary-cache into main
Reviewed-on: unom/punktfunk#316 |
||
|
|
7c411f7ef4 |
fix(nix): the cache's setup steps described a topology that no longer exists
The bring-up instructions were copied from packaging/flatpak/README.md, which
still describes an edge proxy on `home-reverse-proxy-1` forwarding to
192.168.50.50. That home-lab topology is gone. packaging/winget/server/
compose.production.yml — the newest of the three and the only one written since
the move — says so outright: "the sibling docs/flatpak compose files still carry
stale comments … the public hostnames resolve straight to the hcloud box and are
served by Caddy there — no local proxy is involved." flatpak.unom.io resolves to
167.233.145.172, which is unom-1 itself, confirming it.
So the steps now match how docs and winget were actually stood up:
* DNS in the unom.io Cloudflare zone, DNS-only, straight at the hcloud box.
* The vhost in unom/infra `caddy/Caddyfile`, proxying to localhost:3250 —
NOT 192.168.50.50, and NOT hand-edited on the box. ~/caddy/Caddyfile there
looks like the config but is an rsynced copy with no .git to warn you; a
vhost added only on the box lasts until the next deploy. That is how the
winget source vanished on 2026-07-26, and it is now called out here too.
* `caddy_target_ports` + terraform is dropped. It was the home-lab firewall
allowlist; winget's setup, written post-move, has no such step.
Also adds the SNI diagnostic winget's README hard-won: Caddy 308s every Host on
:80 to https, including names it has never heard of, so probing port 80 proves
nothing — check the certificate by SNI instead.
scripts/setup-nix-cache.sh walks the five steps interactively (built from the
/wizard template): it opens each page, says exactly what to click, and verifies
each stage before moving on, because the failure signatures are easy to confuse
— a TLS handshake failure means the vhost is missing, a 502 means the container
is down, and a 404 means the cache is healthy and empty.
It also closes the loop the first version left open: it generates the signing
key locally (a local nix, or the nixos/nix image — MEASURED: both produce the
`name:base64` line, and convert-secret-to-public round-trips), then writes the
PUBLIC half straight into the two docs that carried a `<fill-in>` placeholder.
Nobody has to wait an hour for the first publish to print a value we can derive
up front. The secret half is shown once for pasting into Gitea and never
touches disk. Re-running detects an installed key and refuses to silently
replace it, since that would invalidate every signature already published.
Verified: shellcheck clean, `bash -n` clean, 5 stages against TOTAL_STAGES=5,
and the doc substitution tested against a real generated key — public keys are
base64 and contain `/`, so the sed uses `|` as its delimiter.
|
||
|
|
13f8a1c5cd | Merge remote-tracking branch 'origin/main' into worktree-dualsense-handoff | ||
|
|
37813199b5 |
Merge origin/main — the WirePlumber DualSense policy and the UCM drop-in are complements
Three packaging conflicts, all the same shape: #307 added a `60-punktfunk-dualsense.conf` install at the exact line this branch added the ALSA UCM install to. Both sides kept — they act on different layers and neither subsumes the other: * the WirePlumber rules govern how the pad's nodes BEHAVE once they exist (`node.always-process` so GE-Proton's raw open cannot race itself, `priority.driver = 0` so a pad never clocks somebody else's graph); * the UCM drop-in governs WHICH nodes exist at all (a `SpeakerHaptic` device at priority 200, so the 1-channel sink games overrun is never minted). Checked rather than assumed: the drop-in's node-name matchers (`~alsa_output.usb-Sony_Interactive_Entertainment_DualSense.*`) still match the sink the UCM change introduces — `…DualSense_Wireless_Controller-00.HiFi__ SpeakerHaptic__sink` — so the policy follows the pad onto the new profile. And neither touches volume, so the 0 dB pin on this branch is untouched by both. The Android side of #301 deleted the Compose gamepad mirror, not `SettingsScreen.kt`, so the "Controller speaker" subtitle survives; the Skia console that replaced it carries no speaker row of its own (it opens Android's connected-controllers view instead), so there is no second place to say it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c0dcac7fa2 |
fix(sdk,tray): follow the mgmt port the host actually bound — a moved PUNKTFUNK_MGMT_BIND left every plugin and the tray dialing 47990
Field report 2026-08-18, confirmed: the operator had moved the management API off 47990 (`PUNKTFUNK_MGMT_BIND` in host.env — the supported way to share a box with Sunshine/Apollo). The web console followed, because it reads `<config_dir>/mgmt-endpoint`, the one line the host publishes on every start with the port it REALLY bound. Nothing else did: - The plugin runner / SDK resolved `PUNKTFUNK_MGMT_URL` → literal `https://127.0.0.1:47990`. The runner is a scheduled task (Windows) / systemd unit that inherits nothing from host.env — on Windows it cannot even read it — so every plugin, and the runner's own log shipper, dialed a dead port forever. Task Running, plugins never registering, empty library, and "no logs at all". - The tray defaulted `--mgmt-port` to 47990 and told the operator to edit the autostart command line if they moved the bind. Nobody knows to do that; the tray reports a running host as unreachable. One source, two readers, no new file: - `sdk/src/config.ts::publishedMgmtUrl` reads `mgmt-endpoint`; `resolveConfig` uses it after the env override and before the 47990 default. Every plugin `connect()` follows, on every platform, with no unit/task changes. `runner-cli.ts` additionally exports it into `PUNKTFUNK_MGMT_URL` before any plugin loads, so a plugin still carrying an older vendored `@punktfunk/host` follows too (on Windows `reconcileSharedSdk` cannot refresh the read-only tree, so old copies can outlive several host upgrades). An explicit PUNKTFUNK_MGMT_URL still wins. - `pf_paths::published_mgmt_port` (std-only leaf; the tray now depends on it) parses the same line. The tray's `mgmt_port` becomes `Option<u16>`: `--mgmt-port` pins, `None` re-reads the file on every poll tick, so a host restarted on a new port is picked up without relaunching the tray. Swept the rest: the web console (`windows::service::spawn_web`, the systemd unit, NixOS module) already sourced the file; the host CLI, plugin-kit (goes through the SDK), gaming-mode console and native clients derive the port from discovery / the Welcome — no other literal remained on a loopback path. The console's web port (47992) is not operator-configurable, so the tray's literal there is not the same bug. Verified: SDK 83 tests pass (4 new: absent file → default, published line followed, env wins, blank = unset), `tsc` clean, biome clean; `pf-paths` unit test; `cargo fmt --check` clean; `cargo clippy -p pf-paths -D warnings` clean; `cargo check -p punktfunk-tray -p pf-paths` on Linux (docker rust:1.96) — the tray is cfg-gated off macOS. Not built on Windows from here. |
||
|
|
ab88a8fb40 |
fix(pad-audio): the DualSense's only playback route was a mono sink, and games overran it
A wired DualSense on Fedora 44 / Bazzite / Arch presents exactly one playback
sink: the 1-channel `…Default__Speaker__sink`. GE-Proton mints its synthetic
"Sony controller speaker" endpoint from that lone mono sink, and Marvel's
Spider-Man Remastered overruns it — reliably, ~74 s in:
73.846 render_GetBuffer (…)->(5034, …) <- GE's mono endpoint
73.846 EXCEPTION_ACCESS_VIOLATION info[0]=1 (WRITE) info[1]=5CB9A000
Not a format mismatch: `GetMixFormat` and the game's `Initialize` both agree on
mono float32 `nBlockAlign 4`, and pulse sized `maxlength: 20136` = 5034 x 4
correctly. At the fault `rsi=rbp=0x13aa` (5034, the frame count) while
`rcx`/`rdx` are 5206/5207 — the copy loop had already run past the count. It is
a game/GE bug on a code path that ONLY EXISTS WHEN THE MONO SINK DOES.
So delete the mono sink rather than chase the overrun. `alsa-ucm-conf` describes
the pad as Speaker / Headphones / Mic / Headset and has never carried a
`SpeakerHaptic` device — the DualSense profile arrived upstream in 1.2.15
(36a111a) already without it, and the Deck's is a Valve downstream patch they
still carry on their own 1.2.16.1. With `SpeakerHaptic` at `PlaybackPriority
200` against `Speaker`'s 100 the card takes `HiFi (Mic, SpeakerHaptic)`, the
sink is the 4-channel one, and the mono sink — with the crash path — never
exists. The voice coils reach their own channels as a bonus.
Shipped WITHOUT replacing a file `alsa-ucm-conf` owns, which is what made this
awkward to package. `USB-Audio/USB-Audio.conf` ends with an unconditional,
optional include of `USB-Audio/conf.d/{vid}-{pid}.conf`, placed after its device
table has chosen `${var:ProfileName}` and before it includes the profile that
name resolves to — so a two-line drop-in keyed by 054c:0ce6 / 054c:0df2 swaps
the profile with no diversion, no `Conflicts`, and no `%config` fight. Verified
against alsa-lib rather than assumed: `ucm_cond.c` makes `Condition` optional
for a syntax-v8 `If` carrying `Append`, and `uc_mgr_evaluate_include` evaluates
each included subtree in place before moving to the next include, so the
`Define` lands before the profile include substitutes the variable. The hook and
the DualSense profile shipped in the SAME release (1.2.15), so every tree that
has the bug has the hook.
Host packages only (rpm — and therefore the Bazzite sysext, which unpacks the
RPMs — deb, Arch). The client already has a working fallback in
`ensure_pro_audio`, and a shared file in two co-installable packages is a file
conflict for a nicety. NixOS is not covered: it has no /usr/share/alsa/ucm2 to
drop into and needs a package override instead.
`scripts/ci/check-dualsense-ucm.sh` runs the whole chain on a real distro tree
with no hardware, via UCM's card-less `conf.virt.d` path with only the four card
built-ins stubbed. Against pristine Fedora 44 alsa-ucm-conf 1.2.16.1: baseline
`Headphones/Headset/Mic/Speaker`; with the drop-in, `SpeakerHaptic` and
`HeadphonesHaptic` too, `PlaybackPriority/SpeakerHaptic=200` over `Speaker`'s
100, `PlaybackPCM/SpeakerHaptic=…dualsense_haptic_out:…,1,1,2,3`. It exists
because this fix hooks another project's dispatcher: an upstream rename would
neuter it silently, and what comes back is the crash, not a quieter pad.
Negative-tested both ways (typo'd ProfileName, hook deleted).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c7c9500e89 |
fix(windows/scripting): the runner task writes a log file, so a runner that can't reach the host is no longer silent
Field report 2026-08-18, Windows host on 0.30: PunktfunkScripting task Running, Playnite and Steam plugins installed, library empty, and "no logs at all for plugins" — nowhere on the box. That is by construction, not by accident. The runner's only log door is the log shipper, which tees console output to `POST /plugins/logs` over the mgmt API; the scheduled task itself had no console and no file. So every failure that stops the runner reaching the host — LocalService lost its read grant on plugin-token / native-cert.pem, a moved mgmt bind, a TLS pin miss, a 401 — is exactly the failure the shipper cannot report, and it leaves the same picture: task Running, plugins never registering, an empty grid, and nothing to send when asked for logs. `scripting-run.cmd` now redirects the runner's stdout+stderr to `%ProgramData%\punktfunk\plugin-state\runner.log`, keeping the previous run as `runner.log.1`. plugin-state is the one directory `plugins enable` makes writable for LocalService, and it inherits Users-read from the config dir, so the operator can `type` it from any prompt. Writability is probed with `copy /y nul` first; if the dir is not writable (the task was started by the installer before `plugins enable` ever ran) the runner starts unlogged as before rather than not at all. No `goto`: the file is stored LF and cmd's label scan is unreliable there. The console's empty-Plugins hint (en/de) and the plugin docs now name the file; the log-ship header no longer claims the task writes no file. Verified by reading only — no Windows box reachable from here; the cmd semantics used (`copy nul` as a write probe, `if defined` blocks, leading redirect on `echo`) are the boring ones. |
||
|
|
f24eb02692 |
fix(packaging): the DualSense driver-priority guard has to be 0, and cover the capture node
The shipped WirePlumber policy sets `priority.driver = 1` on a DS5's ALSA sink and says it "keeps the pad from ever driving the graph". Read against PipeWire's own recalc, it does not: `priority_driver` is unsigned and `pw_context_recalc_graph` skips a driver only when it is `<= 0`. At 1 the pad is merely LAST in the ordering — and last is still elected whenever nothing above it qualifies, which on a punktfunk host is the ordinary in-session state, because claiming our own sink as the default output leaves the box's real card idle. With `node.always-process` on the same node it is also permanently runnable, i.e. permanently eligible. Zero is the value that means excluded. The pad keeps driving the streams actually linked to it — a driver always drives its own group, priority orders the election and nothing else — so GE-Proton's haptics are unaffected. The second rule covers the capture side of the same cards. That node is what clocked a reporter's desktop audio for a whole session: in the Pro Audio profile it carries `priority.driver = 2600`, never suspends, and had nothing linked to it at all — its only function on that machine was to clock other people's graphs. The `alsa_output` matches never touched it. Only the priority is set there; holding a device open is about the playback node GE opens raw, and an always-processing microphone is not something this host should ask for. Both of these are belt to the braces of the host-side fix — a capture group that carries its own driver cannot be handed one — but they are worth having on their own: they are what stops a pad from clocking anything else on the box, including a build that predates it. |
||
|
|
0a6a49a9aa |
feat(packaging): a WirePlumber policy holds a DualSense's sound card for GE-Proton
GE-Proton's DS5 haptic router opens the pad sink's backing hw: device RAW whenever it is free — then its own path re-probe EBUSYs against its own handle, invalidates the stream, and spins a 100 Hz "device generation" refresh loop: haptics dead, speaker dead, and in one game a buffer race in the same machinery crashed the title outright. On SteamOS, where that code was developed, PipeWire always holds the device, so GE lands on its well-tested Pulse-routing fallback immediately and none of this fires. Ship the SteamOS-shaped environment: node.always-process + no suspend keeps PipeWire holding the device from the moment the card appears, and priority.driver = 1 keeps the pad — whose USB audio clock (virtual or physical) is nobody's idea of a house clock — from ever driving the graph. Installed by rpm/deb/arch/nix into /usr/share/wireplumber/wireplumber.conf.d/. Matches both DS5 product-string spellings; covers physically plugged pads on a headless host identically. |
||
|
|
a519491928 |
fix(pad): the pad's sound card was root-only, so PipeWire never even saw it
The usbip DualSense's ALSA card is minted mid-session-bringup while no seat session is
active, so logind's uaccess ACL never materialises and /dev/snd/controlC*/pcmC* stay
root:audio 0660 with the user in neither. WirePlumber's probe fails EACCES ("spa.alsa:
can't open control for card hw:2: Permission denied"), the card never appears in PipeWire,
no pad sink exists for winepulse to route to — which is why every mmdevapi endpoint in the
GE logs was a punktfunk-speaker and the audio ContainerIds were all GUID_NULL — and
GE-Proton's direct ALSA haptics leg (find_dualsense_haptic_alsa_path) cannot open the PCM
either. Same mechanism the hidraw rules in this file already handle, one subsystem over.
Verified live on .41 (Bazzite f44): installing the rule + udevadm trigger made WirePlumber
adopt the card mid-session — device, Default__Speaker__sink and Mic source all appeared,
under exactly the alsa_output.usb-Sony_Interactive_Entertainment_ name prefix GE matches.
|
||
|
|
faa00ed142 |
fix(usbip): an OUT reply said 0 bytes accepted, so every hidraw write on the pad "failed"
GE-Proton's `hidraw_enable_dualsense_usb_haptics` never enabled the DualSense's USB haptics
mode against our usbip pad — `err:hid:hidraw_device_set_output_report id 2 write failed
error: 2 No such file or directory`, then feature report 0x08 retried forever with
EINVAL/EAGAIN. Adaptive triggers, voice-coil haptics and the speaker are all gated behind
that one enable, so nothing downstream could ever show a result. Five theories were ruled
out by log inspection; the sixth was measured on the live pad on .41 today:
write(hidraw, output 0x02, 48 B) -> 0
ioctl(HIDIOCSFEATURE 0x08, 48 B) -> 0
ioctl(HIDIOCGFEATURE 0x05, 41 B) -> 41
The vendored simulator answered every non-isochronous OUT URB through the IN constructor
with an empty buffer, i.e. `actual_length = 0` — and a debug_assert pinned that as the
rule ("OUT nothing"). vhci_hcd copies the field into `urb->actual_length` verbatim
(`usbip_pack_pdu(pdu, urb, USBIP_RET_SUBMIT, 0)` in `vhci_recv_ret_submit()`) and has no
other source for it, so `usbhid_output_report()` returned 0 as `write()`'s byte count and
`usb_control_msg()` returned 0 for the SET_REPORT data stage. winebus checks `count > 0`,
takes 0 as failure, and prints the thread's *stale* errno — the ENOENT/EINVAL/EAGAIN in the
log were never kernel verdicts. The earlier `/tmp/hidwrite.py` "150/150 ok" was the same
illusion: `os.write` returning 0 does not raise. A real usbip stub reports the real URB's
`actual_length`, which on OUT is the bytes sent.
Fix: `UsbIpResponse::usbip_ret_submit_out_success(header, accepted)` acknowledges the bytes
taken (`data.len()`, which `read_from_socket` sized from `transfer_buffer_length`) with no
payload back; the handler uses it for OUT; the assertion now pins "OUT carries no buffer",
not "OUT claims 0". Two wire-byte tests pin both directions. The Steam Controller 2 shares
this handler, so its OUT writes were being reported as 0 bytes too.
`scripts/usbip-trace-analyse.py` flagged ANY nonzero OUT actual_length as a desync — the
wrong rule (its own framing never reads a payload back on OUT) and one that would have hid
this bug and flagged the fix. It now flags an OUT reply claiming more than it was sent, or
0 against a non-empty write.
|
||
|
|
8e8cc84d1a |
fix(pad): the usbip DualSense died because its calibration report was one byte too long
`PUNKTFUNK_DUALSENSE_USBIP=1` enumerated the pad and then lost it ~400 ms later,
taking the controller with it (the usbip transport replaces uhid, so there was
nothing to fall back to). Three sessions blamed the ISO stream, the link speed and
`actual_length` in turn. It was none of them.
`DS_FEATURE_CALIBRATION` is 42 bytes. `hid-playstation` asks for 41
(`DS_FEATURE_REPORT_CALIBRATION_SIZE`), and on a USB backend an over-long reply is
not truncated, it is fatal to the transport:
size = urb->actual_length; /* 42, what we declared */
if (size > urb->transfer_buffer_length) /* 42 > 41 */
goto error; /* "probably malicious packet" */
error:
dev_err(&urb->dev->dev, "recv xbuf, %d\n", ret); /* ret still 0 */
usbip_event_add(ud, VDEV_EVENT_ERROR_TCP);
`VDEV_EVENT_ERROR_TCP` tears down the whole connection, not the one URB — hence
`recv xbuf, 0` (that 0 is the untouched initialiser, not a byte count), then
-EPROTO on the calibration read, `Failed to create dualsense`, and the disconnect.
The dmesg order made the teardown look like the cause; it was the consequence.
The blob had been wrong since it was written, and a FIXME said so. It stayed
invisible because every other backend truncates: hidraw for the uhid pad, hidclass
on Windows. USB/IP is the first transport that checks.
Three changes, because one of them alone would leave the same trap set:
- Trim the constant to 41 and pin all three feature-report sizes in a test.
- Clamp every reply to the requested length in the transport (`clamp_reply`), and
drop any payload a handler returns on an OUT transfer — the kernel never reads
one, so those bytes would misframe every PDU after them. A handler bug now costs
one wrong reply instead of the device.
- `DualSenseUsbip::open` waits for the kernel to actually bind a HID driver before
reporting success. A `vhci_hcd` attach succeeds immediately and enumerates
asynchronously, so bringup faults were being reported as working pads; now they
return Err and the caller's existing uhid fallback catches them.
Also adds `PUNKTFUNK_USBIP_TRACE` (both socket directions to disk) and
`scripts/usbip-trace-analyse.py`, which walks a capture and names the first frame
whose declared length disagrees with what the kernel will consume. The handler's
Err arm is no longer discarded either — it was the only signal distinguishing "we
dropped the connection" from "the kernel did", and both read identically in dmesg.
Verified on .21 (CachyOS, kernel 7.1.8): `Registered DualSense controller
hw_version=0x01000208 fw_version=0x01000036`, the device stays enumerated, and
snd-usb-audio mints a real ALSA card. Audio over the isochronous endpoint now runs
for the first time — a 300 Hz tone on the coil pair reads back channel-exact
(peak_coils=0.5000, peak_speaker=0.0000) for the whole run. A 4957-frame capture
analyses clean.
|
||
|
|
be8183caab |
feat(console): host tiles get their OS mark, the chip gets a battery, the strip gets Rescan
`HostRow.os` has been plumbed since the model landed, with a comment saying the drawing was a follow-up because "the Skia glyph set doesn't exist yet". It does exist: assets/os-icons ships thirteen licensed masters, and `pf_client_core::os::os_icon_tokens` already resolves a chain to them - walking most-specific-first and applying the brand aliases (`macos` -> `apple`, `steamos` -> `steam`). Every other front-end walks that same list. So the console takes the shared resolver rather than inventing one, and gets its table GENERATED from the masters (`scripts/gen_os_mark_table.py`, hooked into the existing `gen-os-icons.sh`) rather than hand-transcribed. Thirteen paths of up to 3.5 kB where one mangled character is a silently wrong logo is not work for a human, which is precisely the reasoning the launcher-icon tables already carry. A new master now reaches the console for free; the script's closing note says so. Two corrections to the plan this implements, both found in the code: - The chain is SLASH-separated and resolves most-specific-FIRST, not "the first known token of a `;`-chain". A `linux/fedora/bazzite` host draws Bazzite, and falls back through Fedora to Tux - so the console is right about thirteen distros rather than the four the plan scoped. - The hint bar was already a glass pill, not "ink on the field". What it was missing is that it mixed its OWN glass (a flat wash and a hand-rolled stroke), making it the one floating surface that ignored the palette; it now goes through `theme::panel` like the chip and the toast, and picks up the lit edge. `draw_monogram` becomes `draw_badge`: the OS mark when the chain resolves, the initial when it doesn't. A substitution, not an addition - a badge showing both a Tux and an "L" says the same thing twice - and an older host that advertises no `os` keeps its monogram pixel for pixel. The controller chip gains a pad silhouette and a battery pip. `PadInfo` gets an additive `battery: Option<PadBattery>`; nothing crosses the wire, this is local SDL state. The plan expected to poll "on the existing pad-refresh cadence" - there isn't one, `publish()` is entirely event-driven (hotplug, pin change). And `pad_info` is deliberately open-free because an open GRABS the hardware, while SDL only reports power for an OPEN device. So the level is read from the ONE pad the service already holds open - `menu_open`, the nav pad, which is open exactly while a console is on screen and is the only pad any UI asks about - on a 15 s poll inside the loop that already wakes every 10 ms. Every other pad publishes `None`, which is the honest answer. `None` renders as no battery at all, never 0 %: a wired pad, a Steam virtual pad and SDL's `-1` "powered, level unknown" are all the same non-answer, and 0 % is the one reading that sends someone hunting for a charger. Charging outranks the low-charge red, because a pad at 4 % on the cable is not the problem a pad at 4 % off it is. Finally, Rescan: a second sentinel tile trailing Add Host, sending the `ConsoleCmd::Probe` that has existed unsent by any screen since it was written. A controller surface has no pull-to-refresh, so the affordance has to be a tile. The two trailing tiles are actions rather than hosts, so `hosts.get(i)` answering both "which host" and "which action" with one `None` became a `Slot` enum - with a second action tile that ambiguity is a bug waiting, and the test that matters is that an accidental A on the end of the strip can never start a session. Verified in the pf-gtkflow container: fmt, clippy --all-targets -D warnings, plain build, 104 tests green. |
||
|
|
165a42fcc9 |
fix(host/udev): the virtual Steam Controller 2's hidraw node was root-only, so Steam never saw it
Field report: the Android client captures a wired SC2 and says so ("captured — streams as-is"),
the host attaches the virtual pad over usbip cleanly, and Steam's Settings → Controller →
Connected Controllers is empty. Nothing works in game mode.
The host's own log names the culprit by what is missing from it. Every kernel control transfer is
there — device/config/string descriptors, SET_CONFIGURATION, SET_IDLE, and a GET of the 372-byte
report descriptor — and there is not one SET_REPORT and not one `answering feature GET`. The
kernel enumerated the controller; Steam never opened it.
60-punktfunk.rules grants hidraw access per product id, and it lists what the host used to mint:
the Sony pads, the Switch Pro, the Deck (28DE:1205) and the classic Steam Controller (28DE:1102).
The SC2 identities `steam_backend_product` mints — wired 28DE:1302 and Puck 28DE:1304 — were never
added, so their nodes stayed root-only while the host runs as a user service.
For any other pad that would be a degradation. For this one it is total: no kernel driver claims
the PID (mainline hid-steam stops at the Deck) and the state reports ride a vendor collection, so
there is no evdev node either. Steam is the only consumer there is, and a hidraw node it cannot
open is not a degraded controller, it is no controller. Leaning on the distro's steam-devices
rules doesn't save it — those lists are per-PID and the SC2 shipped in 2026, so a host whose copy
predates it grants nothing.
Add both identities in the same two forms the Deck uses (KERNELS for the UHID shape, ATTRS for the
usbip/gadget one), and document the symptom in troubleshooting with the hand-rollable rule for
hosts on an older package, the one-line-plus-one-absence log signature, and a note that the
trackpads needing Steam to act as a mouse is lizard mode being off on purpose, not a fault.
Single source of truth: every distro installs this file (the NixOS module takes it from the
package via services.udev.packages), so no per-packaging change is needed.
|
||
|
|
daabb85373 | Merge pull request 'The Linux data-plane renice was a silent no-op on every install — RealtimeKit fallback, audio threads boosted at all, nice-limit headroom on every channel' (#232) from worktree-thread-qos-rtkit into main | ||
|
|
52df9c59af |
pkg(linux): nice-limit headroom on every channel, so the renice also works without rtkit
A new shared drop-in, packaging/linux/50-punktfunk-nice.conf (user@.service.d, LimitNICE=-15), raises the user-session nice hard limit so the direct setpriority() path works on rtkit-less boxes — a limit, not a grant, effective from the next login. Shipped by rpm (%files + install, flows into the Bazzite sysext via rpm2cpio), Arch, and deb; the Steam Deck installer writes it to /etc/systemd/system/user@.service.d instead (SteamOS /usr is read-only), following its existing sudo-to-/etc pattern. rpm and deb gain a weak Recommends: rtkit and Arch an optdepends hint — with rtkit the fix needs no relogin at all. The NixOS module instead sets security.rtkit.enable = mkDefault true (rtkit is not a given there; mkDefault keeps it operator-overridable). It remains true on every channel that the host binary must never carry a file capability — the spec's no-caps note now names the two fallback rungs instead of calling the thread nice a best-effort no-op. |
||
|
|
99eb679c07 |
feat(clients): a moved mgmt port now outlives the advert that announced it
Moving the mgmt port off 47990 (the fix for sharing a box with a Sunshine fork, whose web UI owns
that port) only ever worked for as long as mDNS did. The real port lived in the advert and nowhere
else: every client read it live and threw it away, so on a VPN, a routed subnet, or any
multicast-dead network the library silently fell back to a port nothing was listening on.
`KnownHost` gains `mgmt_port: Option<u16>` + `effective_mgmt_port()` + `learn_mgmt_port()`, exactly
the shape `mac` and `os` already use ("learned from the advert while online, persisted so it
survives the host going to sleep") — except this one is load-bearing rather than cosmetic, so
`upsert` states the preserve rule explicitly instead of relying on the does-not-mention-it accident
that `clipboard_sync` survives by, and `upsert_trusted` carries it across a re-key.
Wired through all four client families, each of which was wrong in its own way:
* CLI / Windows / Linux reached for `DEFAULT_MGMT_PORT` at the call site — the constant is the
FALLBACK, not the answer. Windows also needed the port on `Target`, which the library screen has
instead of a `KnownHost`.
* The session console read `advert.and_then(mgmt_port)` with NO saved fallback, two lines above an
`os` that gets the three-rung treatment right. It now matches, and learns on every tick.
* Linux's `mgmt_port_for` consulted live adverts only; it now falls back to the store.
* Android never carried the port at all — its native discovery record stopped at 8 fields. Added
`mgmt` as the 9th (the record's own documented "new fields append, never reorder" rule), then
through `DiscoveredHost` -> `KnownHost` -> `LibraryScreen`.
* Apple LOOKED done and was not: `StoredHost.mgmtPort` and `effectiveMgmtPort` have existed all
along, but nothing anywhere wrote the field and the `mgmt` TXT was never parsed — so it was
permanently nil and every Apple client resolved to 47990 regardless. That is worse than the
honest omissions above, because it reads as finished. Now parsed, carried on `DiscoveredHost`,
and written by `HostStore.updateMgmtPort` at the same site that learns MACs and the OS chain.
Also `PUNKTFUNK_NATIVE_PORT` in host.env, finishing the pair with PUNKTFUNK_MGMT_BIND: `--native-port`
was likewise CLI-only and died on a package upgrade. A bad value is a startup ERROR rather than the
silent fall back to 9777 that `PUNKTFUNK_DATA_PORT` still does — the failure that reads as "I moved
the port and the client still can't reach me". The client side of the native port already worked
(`KnownHost.port` is persisted, `--connect HOST:PORT` names it).
Adding the field broke three `KnownHost` literals in tests, which is the `Default` impl's stated
purpose working ("adding a field here can't silently produce records that lack it"). All three now
carry 47991 — deliberately NOT the default, so the assertions cannot pass vacuously against a
hardcode. New coverage: forward-compat decode of a store predating the field, the resolver
fallback, re-key carry-forward, and on Android the 9th-field parse plus 0/non-numeric/out-of-range
all reading as unknown.
What this does NOT fix: a host that moved its mgmt port and has NEVER been seen over mDNS. Nothing
tells the client where to look, and the honest fix is for the host to announce it in-band — the
`Welcome` message has an established "append a trailing field, older peer decodes to the default"
pattern for exactly this, at the cost of a C ABI accessor and a bump. Left for a separate change.
Verified: Linux (punktfunk-rust-ci/pf-lxcheck2, amd64) `cargo check --all-targets` clean for
pf-host-config, punktfunk-host, pf-client-core, punktfunk-cli, punktfunk-client-linux and
punktfunk-client-session — the last confirmed non-vacuous by planting a compile_error! and watching
the gate fail (cargo prints "Compiling", not "Checking", for bin-only packages, so the usual marker
grep lies about it). Android: :kit + :app compileDebugKotlin clean, ParseRecordTest 12/12 with both
new cases named in the XML. Apple: xcframework built, `swift build` complete, SharedFoundationTests
pass. cargo fmt --all --check clean. NOT verified: the Windows client (192.168.1.133 unreachable).
|
||
|
|
bb78117504 |
feat(host): moving the management port off 47990 now survives, and the console follows
47990 is the management API's port and also Sunshine's (and Apollo's, and Vibeshine's) web UI port. With the GameStream planes off it is the ONLY port the two still share, so moving it is the whole of what "run both on one box" needs — except moving it was barely possible: * `--mgmt-bind` was the sole route, and it lives in a unit file / service registration that a package upgrade rewrites. There was no `host.env` key, so the change did not survive. * The literal 47990 appeared in SIX places — mgmt::DEFAULT_PORT, the Windows service's console launch, scripts/punktfunk-web.service, the NixOS module, web/web-run.cmd, and the console's own default. Nothing downstream could learn a different port, so moving the listener silently left the console proxying to a port nothing was listening on. Now there is one source of truth. `PUNKTFUNK_MGMT_BIND` joins `host.env` (the `--gamestream` / PUNKTFUNK_GAMESTREAM shape: either source works, the CLI flag wins), and `serve` publishes the port it ACTUALLY bound to ~/.config/punktfunk/mgmt-endpoint, in the same KEY=VALUE form mgmt-token already uses so it is sourceable as a systemd EnvironmentFile and readable by the Windows service's existing read_env_file_value. Every consumer derives from that; the 47990 literals survive only as the fallback that keeps an OLD host working with a NEW console. The two unit files drop their hardcoded `Environment=PUNKTFUNK_MGMT_URL=` rather than layering a default beneath the file: whether Environment= or EnvironmentFile= wins is a directive-ordering question, and the hand-written unit and the Nix-generated one do not order the same way. No default, no precedence puzzle — the server's own built-in fallback covers a host that never wrote the file. Two robustness details worth naming, because both fail in the same direction: * mgmt-endpoint is written write-then-rename. A torn read would set PUNKTFUNK_MGMT_URL to EMPTY, which is worse than a missing file — a built-in default only rescues an *unset* variable. * mgmtUrl() now treats blank as unset, which `??` alone does not. The publish happens in parse_serve next to the token persistence, so both files appear together; the console's unit gates on mgmt-token, and its Restart=always picks up a lost race anyway. What this does NOT change: a lost 47990 bind is still fatal to the whole host (the bind sits in tokio::try_join! with the native plane), and running two Moonlight-compatible hosts at once is still unsupported — on Windows the exclusive display topology is a second, independent conflict. Both are documented rather than altered. Verified on Linux in punktfunk-rust-ci (amd64): cargo check --all-targets clean for punktfunk-host and pf-host-config with the "Checking punktfunk-host" marker confirmed present (a first run exited 0 having compiled nothing — the warm shared target dir judged it fresh), 40/40 mgmt tests pass including the new one pinning the published line against both parsers that consume it. Console: tsc --noEmit clean, bun test server/ 9/9. cargo fmt --all --check clean. |
||
|
|
2d15548e38 |
ci(windows): provision the signing toolchain — no .NET runtime meant signtool exited 3 in silence
Verified the whole Azure signing path on the runner (.133) today and it failed twice, for two reasons that neither error message named. Both are now provisioned here so a rebuild from the unom/infra Packer template cannot silently un-fix them. Azure.CodeSigning.Dlib.dll is a mixed-mode C++/CLI assembly: it ships Ijwhost.dll and a runtimeconfig.json pinning Microsoft.NETCore.App 8.0.0. The runner had NO .NET runtime at all — pwsh 7 is a self-contained install and brings no shared runtime — so signtool exited 3 having printed absolutely nothing. Installing the .NET 8 runtime turned that into a clean sign. The client itself installs machine-wide under C:\trusted-signing rather than a user's .nuget, because act_runner runs as SYSTEM, whose USERPROFILE is C:\Windows\System32\config\systemprofile. A per-user install under Administrator is invisible to every job that actually builds. Confirmed by resolving Find-AzureDlib from a SYSTEM scheduled task, which is also how the earlier SSH-only attempts misled: over a network logon New-SelfSignedCertificate hits NTE_PERM, so a control test that "fails" there proves nothing about how CI will behave. Both downloads are SHA-256 pinned against version-immutable URLs (nuget.org flat-container and the dotnet builds CDN), so they fail closed on tampering rather than on every Microsoft patch release — unlike the BtbN `latest` pin above, which re-rolls. The .NET install uses Start-Process -Wait because the bundle is a GUI PE that returns instantly under `&`, leaving $LASTEXITCODE unset and racing the completion check (cost one false failure here). End-to-end result on .133, as SYSTEM: sign rc=0, verify rc=0, chain Microsoft Identity Verification Root CA 2020 -> ID Verified CS EOC CA 04 -> "unom - Enrico Buhler", leaf thumbprint DD6A610F242CB5B2078C2A5D628699B6AB0CAC07 (matches the profile Azure reports), timestamped, leaf expires in 3 days as expected. Signing an unsigned binary and reading the subject back reproduces pack-msix.ps1's Publisher assertion exactly (match=True) — checked against a NON-catalog-signed binary on purpose, because Get-AuthenticodeSignature on a catalog-signed system exe returns the catalog signer and would have read as a false mismatch. |
||
|
|
0f9ccfa8b6 |
fix(ci): funnel the art tests' env overrides through one RAII guard
CI gate C (unsafe hygiene) failed on the previous commit: `library/art.rs` went from 4 process-global-API mentions to 10, because the two new tests each hand-rolled a set/restore pair the way the two existing ones already did. The gate says fix the call sites rather than raise the baseline, and it is right to here — the hand-rolled pattern was also leaking. Each test set `PUNKTFUNK_LIBRARY_ART_ROOTS` and unset it at the end, so any assertion firing between the two halves left the override installed for every later test in the process, turning one real failure into a cascade. `ArtRootsEnv` now holds the lock and the saved values and restores them on drop, which runs on an unwind too. `write_env` is the single write point, so the gate has exactly one pair of call sites to judge: the count drops to 2, below the old baseline of 4, and stays flat however many tests are added. Baseline lowered to 2 in the same commit, as the ratchet's policy requires. ⚠ The gate greps for the API names in COMMENTS as well as code, so the SAFETY comments here deliberately describe the calls instead of naming them. Re-verified after the refactor: .25 493/493 + clippy clean, .133 12/12 art tests + clippy clean, `check-unsafe-hygiene.sh` clean locally. |
||
|
|
a4af1ee8bd |
chore(deps): regenerate third-party notices for the currency wave
Covers all five generated files, not just the root one: the four per-client copies are scoped to the binaries their package installs, so they move independently of the workspace-wide file. Root: 571 -> 575 crates, reflecting this wave (skia-safe 0.99, the RustCrypto digest-0.11 family, jni 0.22, x11rb 0.14, reis 0.7, xkbcommon 0.9, wasapi 0.24, windows-service 0.8.1, x509-parser 0.18, rand 0.9, base64 0.23, libloading 0.9, mdns-sd 0.21 + if-addrs 0.15, rcgen 0.14, criterion 0.8, android_logger 0.15). The per-client diffs are much larger than the wave alone explains, because they were never regenerated after #192: all four still attributed `ring` and named no aws-lc-rs at all. Since #192 removed ring from the tree entirely, the shipped Acknowledgements screens have been crediting a crypto library the clients do not carry while omitting the one they do. They now catch up on both changes at once. (`ring` still appears via the generator's deliberate `--all-features` over-approximation, which sees quinn-proto's wasm-only edge; that is by design — listing an unlinked crate is untidy, omitting a linked one is the failure the file exists to prevent.) Also stops gen-third-party-notices.sh preferring `cargo about` for the root file. That preference was silently destructive: cargo-about only sees CARGO dependencies, so it drops every VENDORED_TREES entry -- pyrowave, the Granite subset, volk, Vulkan-Headers, the Font Awesome brand icons, Simple Icons -- which are third-party sources shipped inside first-party crates under their own licences. Measured today: cargo-about emitted 7,274 lines / ~514 crates with zero mentions of volk, Vulkan-Headers or Font Awesome, against the python generator's 17,324 / 575 with all of them. Merely having cargo-about on PATH was enough to degrade the file, so anyone regenerating after this commit would have undone it. cargo-about remains what the CI licence gate runs -- that job asks a different question (is every licence in the about.toml allowlist) and writes to /dev/null. Both licence-gate legs pass: `cargo about generate about.hbs --fail` and the drivers-workspace leg, RC=0. |
||
|
|
6202543b21 |
Merge pull request 'ci: cache the C/C++ half, link with mold, fix the debug/release cache collision, consolidate the Apple and Windows-client workflows' (#191) from worktree-ci-optimization into main
Reviewed-on: unom/punktfunk#191 |
||
|
|
0bfc7fe913 |
ci: fold release.yml into apple.yml and the two Windows client workflows into one
Two merges, both of which exist to express an ordering Gitea cannot express across files, and both of which delete a duplicated build. release.yml -> apple.yml (as the `distribute` job) The name described neither what it did (Apple only — every other platform's release is its own packaging workflow attaching to the same Gitea release on a v* tag, with announce.yml as the manual "go") nor anything a reader would guess. The name was the smaller problem. Gitea has no cross-workflow `needs`, so nothing sequenced it against apple.yml's tests: a canary main push uploaded iOS, macOS and tvOS builds to TestFlight even when `swift test` had just failed on that same commit. It is now `needs: swift`, which is only expressible in one file. The two files' paths: filters had also drifted — apple.yml watched crates/**, release.yml watched crates/punktfunk-core/**. The merged filter takes the NARROW one, because that is the correct one: everything on this runner is built from punktfunk-core via build-xcframework.sh, and punktfunk-core's only path dependency is its own vendored fec-rs. That is checkable in one command, and the header says so, and says to widen it if that ever stops being true. Net effect on the shared mac mini: pushes that touch host-side crates no longer build or upload anything Apple. windows.yml + windows-msix.yml -> windows-client.yml The pair built the same three crates FOUR times per client push on ONE runner: debug x64 + arm64 for lint/test, release x64 + arm64 for packaging. windows-host.yml already records why a second (debug) dep tree on this machine is a liability rather than a cost — it re-runs openh264-sys2's vendored C++ through cc-rs's cl.exe fan-out and tips the runner into C1069, which is disk exhaustion wearing a compiler error's clothes. So there is one release build per arch now and clippy/fmt/test run against it, exactly as windows-host.yml does. The paths list went from three copies to one; PRs get the build/lint/test signal and stop before packaging. The rename is safe, and this is worth recording because the GitHub instinct is wrong here: `github.run_number` is REPO-WIDE in Gitea, not per-workflow — consecutive runs of DIFFERENT workflows get consecutive numbers (verified against the API: android 13226, apple 13227, arch 13228, ci 13229, deb 13230). The canary MSIX version <minor>.<run>.0 and Apple's CURRENT_PROJECT_VERSION therefore keep climbing across a rename. On GitHub the same rename would reset both to 1, sorting every new canary below the published ones and getting the TestFlight uploads rejected outright. 25 workflows, down from 27, and every `name:` now matches its filename. Cross-references in windows-host.yml, windows-drivers.yml, android.yml, flatpak.yml, sbom.yml, the provisioning scripts, gitea-release.sh and clients/windows/packaging/README.md updated. |
||
|
|
f1dc6c9f94 |
ci: cache the C/C++ half, link with mold, and split the debug/release target caches
Three independent reasons Rust CI stayed slow despite sccache, fixed together because
they share the same measurement.
1. sccache only ever covered RUSTC. Every C/C++ dependency in the tree — aws-lc-sys,
openh264-sys2's vendored C++, the CMake-built libopus behind audiopus_sys — was
compiled from scratch on every job of every workflow. CMAKE_{C,CXX}_COMPILER_LAUNCHER
plus CC_/CXX_x86_64_unknown_linux_gnu route both build-script styles (cc-rs and
cmake-rs) through the same shared cache.
The CC_* vars are JOB-scoped in ci.yml and deb.yml, never workflow-scoped: the
arm64 cross image sets its own CC_x86_64_unknown_linux_gnu=pf-host-cc, the wrapper
that keeps ffmpeg-sys-next's host probe off the arm64 include dirs. Overwriting it
would surface as a header mismatch rather than as a CI config error.
2. Linking is cacheable by nothing, and these jobs relink the host, client, session,
cli, worker and tray on every run — twice per push for rpm (f43 + f44). The four
Linux builder images now install mold and carry a $CARGO_HOME/config.toml that uses
it for x86_64. aarch64 is deliberately left alone (cross driver, already-fast legs).
Each image asserts `mold --version` in its build, so an image can never ship the
flag without the linker: docker.yml goes red and :latest stays on the last good one.
3. THE EXPENSIVE ONE. ci.yml (debug) and deb.yml (release) named a byte-identical
target-cache key, under a comment claiming the release build reused ci.yml's
artifacts. It never could. actions/cache is first-saver-wins on an exact key and
ci.yml is the faster job, so the shared key always held a debug-only target/ — and,
worse, deb.yml could then never save its own, because the key was taken. Every
canary .deb has been a from-scratch release build for as long as both keys existed.
Same collision on the arm64 pair, and a third participant in
linux-client-screenshots.yml. Split into -debug-/-release- key families; that job
reads deb's tree via restore-keys but keeps its own exact key so it can never win
the save race and replace a full tree with its single-crate one.
Also: one scripts/ci/ensure-sccache.sh replaces ten copy-pasted bootstrap blocks that
had already drifted into two dialects (GNU tar --wildcards vs bsdtar), every Rust job
now ends with --show-stats so a cache regression is visible instead of just "CI got
slower", and deb.yml's web install joins every other CI install on --ignore-scripts.
No behaviour change to any artifact: same compilers, same flags, same outputs.
|
||
|
|
346385bad8 |
fix(deb): ship punktfunk-gamescope on apt at last, and support Debian 13
`punktfunk-gamescope` had never been published to the apt registry — not in any release. It was built inside the host job's Ubuntu 24.04 image, where it cannot build: our pin vendors wlroots 0.19.3, which floors `wayland-server` at 1.23.1, and noble ships 1.22.0 (it also lacks libxcb-errors-dev and has only libdisplay-info 0.1.1). Every rung of that path was a `::warning::` returning 0 and the one hard gate ran last by design, so v0.26.0 and v0.27.0 both released with the package missing while docs-site told apt users to install it. The same tags shipped it fine for Arch, Fedora 44 and Bazzite. It now builds in its own job on Debian 13 (ci/gamescope-trixie.Dockerfile), the oldest apt base the tree configures on. One package serves Debian 13 AND Ubuntu 26.04 — measured by installing and running it on both — because the build also vendors libdisplay-info via the new `--extra-fallback` option: linked against the distro copy it demands `libdisplay-info2` on trixie, which Ubuntu 26.04 does not have (it carries libdisplay-info3). The option is opt-in, so the Arch/Fedora/nix outputs are byte-for-byte unchanged. Ubuntu 24.04 gets no gamescope package and cannot — its wayland is too old to run one however built. Debian 13 is now a documented host target. That needed no packaging change at all: the host .deb's glibc-2.39 floor and bundled FFmpeg already made it installable, and it had been working for a long time while docs-site said Debian was unsupported and unverified. Verified by installing: host, web console and plugin runner install, resolve every soname and run. The desktop client stays Ubuntu-26.04-only (built there, floors at `libc6 >= 2.43`; Debian 13 has 2.41). Compositor detection now answers Cinnamon (Mint, LMDE) with the route that works instead of advice that cannot help. Muffin forked from Mutter 3.36: `org.cinnamon.Muffin.ScreenCast` has only RecordMonitor/RecordWindow, never RecordVirtual, and xdg-desktop-portal-xapp implements no ScreenCast — so no value of PUNKTFUNK_COMPOSITOR makes a Cinnamon desktop host a virtual display. The error names headless gamescope, which needs no desktop compositor. The XDG sniff moved into a pure function so those branches are testable; Cinnamon is matched before GNOME, since it is a GNOME derivative and the generic arm would otherwise hand it the Mutter backend (caught by the new test). New `smoke-install` job installs every published package from the registry in pristine ubuntu:24.04, ubuntu:26.04 and debian:trixie images, asserts each binary resolves its libraries and runs, and insists the version served is the one this run built. Nothing in deb.yml had ever installed a package it produced, which is how both of the above survived unnoticed. ⚠ Bootstrap: seed `punktfunk-gamescope-trixie:latest` into the LAN registry once (docker.yml builds it thereafter) or the new job cannot start. |
||
|
|
f373dffb5e |
chore: migrate the main workspace and pf-vkhdr-layer to edition 2024 (WP20)
The safety half of the rust-safety programme's §8.4: `std::env::set_var`/`remove_var` are
`unsafe fn` in edition 2024, converting the class of bug the programme found the hard way
(the
|
||
|
|
2bfd1cd2d5 |
chore(safety): three unsafe-hygiene grep gates, blocking in ci.yml (WP2c gates)
scripts/ci/check-unsafe-hygiene.sh — textual gates for three classes no lint covers: A. unsafe fn markers carrying no contract. unsafe_op_in_unsafe_fn forces real ops into blocks, so an unsafe fn with no `unsafe` in its body is a marker with no contract ( |
||
|
|
23d0452157 |
feat(host): GameStream opt-in on every route; the native plane is deny(unsafe_code)-enforced
The user direction after WP0: ENet exists only for Moonlight, so the native
plane must be provably safe and the compat planes a deliberate choice.
Opt-in, everywhere. Windows already was (unchecked installer task). The three
opt-out surfaces are flipped: the shipped systemd user unit (deb/RPM/Arch/
sysext) no longer bakes --gamestream into ExecStart — a new
PUNKTFUNK_GAMESTREAM=1 host.env knob (pf-host-config, OR-ed with the CLI
flag) is the packaged opt-in; the NixOS module default goes true→false, with
a module-check assertion that unset = native-only; the Deck installer takes
--gamestream to opt in (--no-gamestream kept as explicit-off). Docs
(quickstart, running-as-a-service, moonlight, ubuntu/fedora/arch firewall
sections, gnome/sway, how-it-works) rewritten to the opt-in shape; the
CHANGELOG carries the upgrade note.
Enforced-safe. punktfunk-core is #![deny(unsafe_code)] crate-wide — every
module that parses network bytes is safe Rust as a compile error, not a
census result. Carve-outs are exactly two documented classes, neither of
which interprets attacker bytes: the client surface (abi, client) and the
transport syscall-batching shims (udp/{apple,linux,windows}, qos_windows).
In punktfunk-host, the modules a secure-default host exposes — native
(cfg-not-test: its tests exercise the client C ABI on purpose),
native_pairing, mgmt, mgmt_token, discovery, wol — are #[forbid(unsafe_code)].
Gates: Linux amd64 container clippy --all-targets -D warnings clean over
core+host-config+host; core 204 tests green under the deny; mgmt 46/46,
control 6/6. .133 Windows clippy (shipped features, clean-first,
sentinel-checked) clean — covers the qos_windows/udp-windows carve-outs.
macOS + iOS cargo check green (the apple.rs carve-out compiles for real).
|
||
|
|
db6683a585 |
chore(safety): commit the unsafe census, fix its two bugs, record the baseline
Founding commit for a host-focused Rust safety programme. Adds the census tool
that measures the programme, the 2026-08-11 baseline it produces, and the
programme document itself.
The metric is SHIPPED NON-FFI UNSAFE OPERATIONS: 713. Raw `unsafe {}` block
count is the wrong target and the workspace manifest already says why — 63.3%
of unsafe operations in host scope (1542 of 2435) are a single third-party FFI
call that ash/windows-rs/ffmpeg mark unsafe on our behalf. A block count also
rewards merging blocks, ignores SAFETY comments, and IMPROVES when code moves
from Linux to Windows, because no local check can see the Windows half.
The tool shipped here had two defects, both fixed:
- `in_test_mod` cached parsed `#[cfg(test)]` spans in a dict keyed on `id(src)`,
the memory ADDRESS of the source string. CPython recycles addresses, so once
one file's source was collected the next file's string could be allocated at
the same address and silently inherit the previous file's test spans. Ten
consecutive runs over an unchanged tree produced 694, 695, 696, 701, 703,
709, 710, 713, 714 and 721. Fixed by holding a strong reference to the string
beside its spans, which makes the address un-recyclable while the entry is
live. Five consecutive runs now agree exactly.
- The layout-assertion regex matched `const _: () = assert!(...)` but not the
`const _: () = { ... };` block form, which 18 files use — including abi.rs,
pf-inject/linux/gamepad.rs and pf-capture/.../idd_push/probes.rs. It reported
102 unguarded repr(C) declarations across 25 files where the true figure is
60 across 22, defaming three well-guarded files.
A metric that is not reproducible is not a ratchet. The acceptance gate for
this commit is therefore five consecutive identical runs, not one.
Baseline: 713 shipped non-FFI unsafe operations; 60 unguarded repr(C)
declarations across 22 files; unsafe reachable pre-authentication by an
unpaired peer = 0 first-party.
|
||
|
|
f62a48d4a9 |
feat(library): launcher tiles get their launcher's logo — a brand token on the wire, the vector in every client
A launcher tile (role: "launcher", design D4) shipped no art on purpose:
a launcher's own icon is square, every client cover-crops a 2:3 poster,
and the crop turns a mark into a strip. So the tiles were the launcher's
name on a flat accent face — legible, and the blandest thing in the grid.
Entries now carry an optional `icon`: the NAME of a brand mark, never
image bytes and never a URL. `[a-z][a-z0-9-]{0,31}`, shape-validated by
the host on every lane (a client interpolates the value into a resource
name or an asset lookup, so the guard belongs upstream of all of them,
and each client re-checks rather than trusting the peer).
A token rather than art because the alternative is closed by
construction, and deliberately: the art proxy serves what the bytes ARE
(sniff_image_type) and SVG is not on that list — it is script-capable
XML and the console renders library art in a browser. Widening that
sniff would trade a rendering nicety for a stored-XSS surface. Naming
the mark keeps the refusal intact, keeps the glyph vector at whatever
size a tile happens to be, lets it take the tile's ink, and adds nothing
to a reconcile payload that is already body-limited. The cost is that a
third-party plugin cannot ship a mark no client bundles; its tile falls
back to the launcher's name, exactly as before, and the fix is a PR
adding the master.
assets/launcher-icons/ holds seven monochrome masters with per-mark
provenance and licensing (Simple Icons CC0: lutris, heroic, epic, gog;
Font Awesome CC BY: steam, xbox; Playnite's own logo, MIT). steam is
generated FROM assets/os-icons/steam.svg so the SteamOS host badge and
the Steam launcher tile can never drift.
scripts/gen-launcher-icons.sh bakes the three derivatives that cannot
consume a master (GTK symbolic SVG, Windows PNG, Apple template PDF)
and — unlike gen-os-icons.sh, which prints path data for a human to
paste — GENERATES the three inline registries (web console, Android
ImageVector, pf-console-ui Skia). Three clients x seven paths of up to
3 kB is a transcription error waiting to happen, and a mangled character
is a silently wrong logo rather than a build failure. The generated Rust
goes through rustfmt, since `cargo fmt --all --check` is a CI gate and a
generated file that fails it would fail every regeneration.
All six renderers draw the mark CONTAINED, never cover-cropped: the
masters' viewports are not square (steam 496x512, playnite 1024x1024)
and filling a 2:3 frame would reproduce the strip this exists to avoid.
Every one keeps its old fallback for a token it has no art for.
Epic, GOG and Xbox marks ship dormant. Those plugins' launcher switches
are off by default and emit nothing, because the host has no verified
launcher_ui activation for them yet — shipping the art now keeps turning
one on the one-line plugin change those plugins promise, instead of also
needing a release of all six clients.
api/openapi.json and the SDK are regenerated (the spec's version field
was stale at 0.25.0 and now reads 0.26.0, which is the crate's actual
version — an unrelated line that regeneration necessarily corrects).
Verified: host cargo check, clippy -D warnings across pf-client-core /
pf-console-ui / punktfunk-client-session / punktfunk-client-linux, plain
build, pf-console-ui tests (77, including a new one asserting all seven
masters parse under Skia and one asserting the letterbox stays inside
its box), pf-client-core tests (188), cargo fmt --all --check, Apple
swift build, Android compileDebugKotlin, web tsc + vite build,
plugin-kit tsc, biome. The Windows client is NOT compile-verified — it
cannot be built from a Mac (scripts/xcheck.sh covers only the capture
stack by design) and CI does not build it either; its tile change needs
a real box before it ships.
|
||
|
|
890b67a863 |
fix(validation): v5 called jitter a leak — a leak is a trend, not a spread
v5's verdict was `max - min` over the sampled fd counts with a default tolerance
of 0. An encode worker's fd count legitimately moves by one when a dmabuf fd is
in flight at the sampling instant, so the spread was permanently 1 and the leg
failed on a perfectly healthy box — reported, like every red leg here, as "a
shipping blocker, not a flake".
Measured on home-nobara-1 (KDE, RTX 5070 Ti), 33 samples over 480 s:
54 54 54 54 54 54 54 54 55 55 54 54 54 55 54 54 55 54 54 54 55 54 …54
It oscillates and ENDS on 54, exactly where it started. Nothing accumulates.
The replacement is median-of-thirds: median(last third) - median(first third).
That is strictly MORE sensitive to what R2 is actually about — a steady leak
moves the trend just as much as it moves the spread, while bounded jitter moves
only the spread — so this is not the tolerance being widened to get a green.
The spread is still printed, now labelled as jitter when the trend is flat. The
warm-up window already covers the one-off first-sight-of-each-buffer cost, so a
plateau inside it is by design not a leak; a step that never comes back still
trends and still fails.
The self-test grows the cases that force this to be a real assertion: the
measured oscillation must trend to zero, a synthetic leak must still trend up, a
flat series must be flat, and a step that never returns must be caught. Writing
them is what caught my own arithmetic — the first draft asserted a leak trend of
12 where the reader correctly says 10.
Also records what the v5 log now makes obvious: `--minutes` does NOT set the wall
clock. `spike` is frame-count bounded (`seconds * fps`), and a KWin virtual
output being driven hard delivers ~197 fps against a `--fps 60` budget, so a
"10 minute" run ended after 182 s. Ask for more minutes than you want.
|
||
|
|
9cdbfabd4d |
fix(validation): v4.e demanded a rung the spike vehicle cannot reach
v4.e killed the worker mid-session and then required "the encode worker died
mid-session" in the spike's log. That line, and the respawn that follows it, are
emitted by `RemotePyroWave::reset` — and the only caller of `Encoder::reset` is
the real session's `reset_stalled_encoder` loop in native/stream.rs. `spike` is
a dev tool with no recovery loop at all: it does
encoder.submit(&frame).context("encoder submit")?
and exits. So a worker killed under the spike can never reach reset, the line
can never appear, and the leg reported
FAILED — a red leg here is a shipping blocker, not a flake.
for a ladder rung the product implements correctly. A false negative in the one
place that must not have one: this kit exists to refuse false PASSes, and a
false FAIL spends exactly the same credibility.
Verified on glass first, so the rung is not being excused on a reading of the
source. home-nobara-1 (KDE, RTX 5070 Ti), real client session, worker pid 44249
killed with -9: `video_streaming` stayed true across the kill, and the host
logged
pyrowave: respawned the encode worker after a mid-session death
worker=/usr/bin/punktfunk-encode-worker priority=Granted(Realtime)
encoder submit failed — encoder rebuilt in place, forcing an IDR
error=... Broken pipe (os error 32) reset=1 max=5
v4.e now asserts the half the spike can actually observe — the death surfaces as
an ATTRIBUTABLE worker-IPC error naming the worker, after real encode windows,
and the host process does not die with it. A hang, an unexplained failure, or a
dead host still fails. The respawn half is printed as the human follow-up, in
the same idiom v1 already uses for its on-glass half, and written into `recipe`
with the two commands that close it.
|
||
|
|
620f017d9a |
fix(validation): tell the CAPTURE it is a PyroWave session, or NVIDIA hands it a CUDA buffer
`--codec pyrowave` selects the ENCODER. The capture pipeline picks its consumer from
`ZeroCopyPolicy::pyrowave_session`, which on the spike path is fed only by the global
`PUNKTFUNK_ENCODER=pyrowave` lab lever (punktfunk-host/src/capture.rs). Without it, .21 resolved
capture pipeline resolved: cuda-import -> nvenc capture_arm="cuda-import" consumer="nvenc"
zero-copy: dmabuf imported to CUDA (no CPU copy) nv12=true
and the wavelet encoder refused the payload on its first submit: "unsupported FramePayload (need
Dmabuf or Cpu RGB)". That is not a worker bug — the arm that failed was the pure in-process one.
It reproduces only where the A/B actually lives. An AMD box has no CUDA arm to pick, so .25 resolved
straight to dmabuf-passthrough and the kit looked correct there. With the lever set, .21 resolves
`dmabuf-passthrough -> pyrowave` and both arms encode 2700/2700 frames.
V3b then passes on .21 (RTX 5070 Ti, GRID 2 at ~100% GPU, 5120x1440 — the portal captures the real
monitor, --width/--height being synthetic-only):
in-process, refused p50 2.85 ms p99 8.39 ms (10 windows)
capped worker, granted p50 2.65 ms p99 4.10 ms (11 windows)
p99 delta -4.29 ms
The worker reports `priority=Granted(Realtime)` with `ext=VK_KHR_global_priority` on the FIRST
attempt and logs no fallback line; the refused arm logs "every global queue priority class was
refused". So the capability still buys the lever from a SEPARATE process, with the IPC hop in the
loop — 8.39 -> 4.10 ms is a 51% p99 cut, against PW1's in-host 6.4 -> 4.4 at 1080p. Different
resolution and a harder load, so treat the class as confirmed and the absolute numbers as not
comparable to PW1's.
|
||
|
|
854b14a52e |
fix(validation): a leg that never ran must say so, not blame the arm
The V3b run on .21 died with `open portal capturer: timed out waiting for the ScreenCast portal` — a GNOME consent dialog nobody answered — and the kit reported "arm A is not the in-process arm". That is false: the arm was constructed correctly (`PUNKTFUNK_ENCODE_WORKER=off` is right there in the captured env header), it simply never reached encoder-open, so the line the assert looks for could not exist. A red that points at the wrong thing costs the same debugging time as a green that hides a real one. `spike_failure_reason` now runs BEFORE any arm-identity assert in v2, v3a and v3b, and names the actual cause: the portal timeout gets its own message saying the dialog appears on the HOST's own screen and cannot be answered from inside a stream — which is precisely the situation that produced this failure, since the operator was watching the box through a game session at the time. Falls back to the first ERROR line, then to "no PUNKTFUNK_PERF window at all", so a spike that dies some other way still reports that rather than a misattribution. |
||
|
|
bbc0513f0c |
fix(validation): the kit could not read a real log — tracing wraps field names in ANSI
Found by running it. The first V3a run on .25 encoded 2700 frames in BOTH arms, at 59.6 fps, with 22
perf windows each — and the kit reported "fewer than 3 usable perf windows", because `tracing`'s fmt
layer wraps field NAMES in SGR escapes. The bytes on disk are `p99_us\e[0m\e[2m=\e[0m4601`, so
`s/.*p99_us=\([0-9][0-9]*\).*/\1/p` never matched. The message text is plain, which is why the
window COUNT was right and only the numbers vanished — and why the fixtures never caught it: they
were hand-written, and cleaner than reality.
Anything matching a field breaks the same way, so this was not only V3a: v2's `priority=Realtime`,
the demotion `reason=`, and v4's rungs all read fields. Every log read now goes through one
`log_cat` that strips SGR, and the spike is launched with NO_COLOR=1 so fresh logs are plain at the
source too — a human grepping a red leg by hand is defeated by those escapes exactly as the parser
was.
The self-test gains the same four perf windows a second time, ANSI-wrapped, asserting an identical
result: same numbers, same expectation, so a failure there can only mean the stripping broke. That
fixture caught its own first draft, which built the line in one printf with 27 placeholders against
23 arguments and emitted empty escapes — hence the field-at-a-time helper.
With this, V3a self-reports on .25 (sway headless, real dmabuf capture, AMD 780M/RADV, 2700 frames
per arm, both arms at default GPU priority):
in-process p50 2.08 ms p99 4.18 ms (21 windows)
uncapped worker p50 2.07 ms p99 3.52 ms (21 windows)
p99 delta -0.66 ms -> PASS
R1's pre-registered abandonment gate does not fire: the process boundary is not merely under the
+1.0 ms ceiling, it is measurably FASTER at the tail, while p50 is unchanged (2.08 vs 2.07). An
earlier hand-extraction of the same logs gave -0.43 ms, so the direction reproduces across runs.
Caveat for whoever reads this later: idle iGPU in a KVM guest, RADV, no GPU-bound load. This bounds
the IPC hop; it says nothing about V3b, which still needs .21 under GRID 2.
|
||
|
|
2de604ecab |
test(validation): the on-glass kit for the encode worker, including the 0.26.0-1 regression test
WP3 of design/gpu-priority-capability-worker-implementation-plan.md. Five legs, the first of which is
the test that would have caught the field incident: in a KDE session with the worker installed and
capped, `getcap` on the host must be EMPTY, its CapPrm all zeroes, `readlink /proc/<pid>/exe` must
resolve, and `punktfunk-host probe-compositor` must exit 0 — which on KWin succeeds only when the
privileged zkde_screencast_unstable_v1 global was actually advertised to this client.
Read-only by default; the one mutating rung (kill -9) is behind --allow-mutate and kills only a
worker that is a child of the spike the script itself started. It NEVER calls setcap: the uncapped
arms use a plain copy of the worker, which does not carry security.capability, verified uncapped
before use. So no leg needs root and none restores state. A skip is never a pass — exit 2 means
incomplete, distinct from 1 (failure).
V3 is split, which the plan did not do. Its stated form compares against PW1's in-process-capped
baselines, and those exist only on .21 under GRID 2:
* V3a is the pre-registered abandonment gate and needs no capability at all — in-process versus an
UNCAPPED worker, both at default priority, so the only difference is the process boundary. Fails
if the worker's p99 exceeds inline by more than --gate-ms (1.0). This runs on any box with a GPU.
* V3b is the lever itself, capped worker versus the refused in-process arm, and says plainly that
an idle GPU makes it meaningless.
The false PASS this kit exists to refuse: a CPU-backed frame makes the proxy pin itself in-process
for the session, so a synthetic source would quietly turn the "worker" arm into a second in-process
arm and pass the gate for the wrong reason. The worker arm is only accepted with a dmabuf-passthrough
capture, a capability-carrying-worker line, and no fallback line anywhere in the log.
Also asserts the host and worker are different inodes — a hardlink shares the file capability, which
is the same incident by another route.
|
||
|
|
4f8cce6751 |
feat(packaging): grant CAP_SYS_NICE to the encode worker on all six channels, and assert the host never gets it
767e67ca's per-channel mechanics were correct; they were aimed at the wrong binary. Each one is restored here pointed at punktfunk-encode-worker, and every host-side removal from #136 stays verbatim. All grants remain best-effort — an uncapped worker still encodes, at default priority, so a failed setcap must never fail an install. * Arch: setcap in post_install AND post_upgrade (a replaced binary is a new inode). * RPM: %caps(cap_sys_nice=ep) in %files, never a %post setcap — %caps applies, restores and verifies, and covers Fedora as well as Bazzite via rpm-ostree layering. * Bazzite + Arch sysext: setcap on the staging tree before mksquashfs, which does record security.capability. The assertion is amended, not removed: host EMPTY is still a hard fail, and the worker must carry exactly cap_sys_nice=ep — missing is fine, anything else is not. * deb: setcap in postinst. * NixOS: security.wrappers for the WORKER plus PUNKTFUNK_ENCODE_WORKER in the unit. A file capability cannot live on a store path, and an ambient grant is right here precisely because nothing ever identifies the worker. The host's ExecStart stays on the store path. * Steam Deck: setcap the worker; the .desktop the script writes stays valid this time. Four things the plan's channel table missed: * packaging/arch/build-sysext.sh had no capability handling at all, and a sysext can never run a pacman scriptlet — the SteamOS image would have shipped the lever permanently inert. * scripts/steamdeck/update.sh had none either. It rebuilds both binaries, so a new inode drops the grant, and it is the documented steady-state path: the lever would have died on the first update. It also never healed a Deck already capped by 0.26.0-1. * A capped worker is AT_SECURE, and glibc drops $ORIGIN-expanded RPATH entries for secure binaries unless they normalise into a trusted system dir. Copying the host's rpath under BUNDLE_FFMPEG=1 would have left the capped worker unable to find libavcodec on exactly the channel that bundles it. Absolute DT_RPATH instead. * Nix crane scopes by -p, so the worker would not have been built at all, and it needs its own addDriverRunpath. scripts/ci/assert-cap-matrix.sh mechanizes the lesson from 0.26.0-1 — verify the PACKAGE, never the board. It unpacks the built Arch package, the deb, the rpm and the mounted sysext raw and asserts one matrix: the host carries NOTHING (hard fail), the worker exactly cap_sys_nice=ep. The sysext reader first proves it can round-trip a capability through mksquashfs/unsquashfs at all, so an unreadable artifact fails rather than issuing a blind PASS, and --self-test red-teams the assertions themselves. Red-teaming the leg found a real bug: setcap originally ran BEFORE the assertion, so "the worker arrived carrying something unexpected" was unreachable and a stray %caps would have been silently overwritten. Both sysext scripts now assert, then grant, then assert again. |
||
|
|
4d383811c0 |
fix(packaging): the same CAP_SYS_NICE broke KDE on FIVE channels, not one — Bazzite included
The Arch fix in the previous commit was incomplete. 0.26.0-1 granted the host CAP_SYS_NICE through
every Linux channel we ship, and each one breaks KWin identification the same way:
* packaging/rpm/punktfunk.spec .......... %caps(cap_sys_nice=ep) in %files <- Fedora AND Bazzite
via rpm-ostree layering
* packaging/bazzite/build-sysext.sh ..... setcap on the staging tree, recorded by mksquashfs
* packaging/debian/build-deb.sh ......... setcap in the postinst
* packaging/nix/nixos-module.nix ........ security.wrappers with capabilities = "cap_sys_nice=ep"
* scripts/steamdeck/install.sh .......... setcap on $BIN, six lines after writing the .desktop
whose Exec= it thereby voids
Bazzite was NOT a separate fault, as first reported here — it is this one. Verified by mounting the
published punktfunk-0.26.0-1-x86-64.raw: `getcap usr/bin/punktfunk-host` reports cap_sys_nice=ep,
stored as security.capability in the squashfs. The claim in packaging/arch/build-sysext.sh that
"file capabilities don't survive this squashfs path" is false and is corrected here; mksquashfs
records them, which is exactly why the image shipped one.
NixOS deserves its own note: a security.wrappers entry does not dodge the problem. The wrapper
raises the capability into its AMBIENT set before exec'ing the store binary, precisely so it
survives — which lands CAP_SYS_NICE in the exec'd process's permitted set and fails the readlink
identically to a file capability. ExecStart now points at the store path directly, which is also the
path packages.nix substitutes into the .desktop's Exec=, so the two finally agree.
Measured blast radius of holding a capability, same-uid reader, CachyOS kernel 7.1.6:
/proc/PID/exe ....... EPERM <- KWin's identification. Desktop sessions die.
/proc/PID/root/* .... EPERM <- xdg-desktop-portal reads .flatpak-info here to resolve an
app id; the wlroots and Hyprland backends go through it
/proc/PID/environ ... EPERM
/proc/PID/cgroup .... OK
/proc/PID/status .... OK
/proc/PID/cmdline ... OK
Compositor backends, by exposure: KWin is broken outright (proven, field-confirmed). gamescope has
no identity gate and was never affected, which matches the field — only Desktop mode was reported.
Mutter drives Mutter's own D-Bus API, not the portal, and looks unaffected. wlroots and Hyprland go
through the ScreenCast portal, whose app-id resolution reads a path the capability blocks — a real
exposure, not something I reproduced end to end.
The sysext build now HARD-FAILS if a capability is staged, rather than trusting that the RPM payload
never carries one: a merged sysext's /usr is read-only squashfs, so a bad image cannot be repaired
on the box, and the spec was one %caps() away from baking one in again.
Docs corrected, because they advertised the capability as a feature:
* docs-site running-as-a-service "GPU scheduling priority" — rewritten: the host carries no
capability, why it must not, and how to clear a 0.26.0-1 install (Bazzite needs a new image)
* docs-site configuration.md — the PYROWAVE_QUEUE_PRIORITY row no longer claims the packages grant it
* packaging/bazzite/README.md — §6.5 still described the kde-desktop-setup.sh behaviour from
before it stopped writing KWIN_WAYLAND_NO_PERMISSION_CHECKS and started REMOVING it; plus a
note that 0.26.0-1 Desktop mode cannot be repaired in place
* packaging/arch/README.md — the false "capabilities don't survive the sysext" line
* CHANGELOG v0.26.0 PW1 — annotated with the 0.26.0-2 correction rather than rewritten, and the
owed PyroWave-under-load A/B now says it needs a gamescope-only box
Verified: bash -n on all five changed shell files; nix-instantiate --parse on nixos-module.nix and
packages.nix; the published 0.26.0-1 sysext mounted and its capability read; getcap on an uncapped
file exits 0 with empty output, so the new build assertion cannot false-positive.
|
||
|
|
d3aaa16a7d |
Merge branch 'worktree-wave2-pw3-dmabuf-latch' into worktree-wave2-pyrowave
# Conflicts: # packaging/arch/punktfunk-host.install # scripts/steamdeck/install.sh |
||
|
|
ce31a9ddfd |
Merge pull request 'The Steam Deck updater kept sabotaging its own next update, and hand-deleting a file was the only way through' (#122) from worktree-steamdeck-update-bunnix-dirt into main
Reviewed-on: unom/punktfunk#122 |
||
|
|
62a6fa9fac |
fix(packaging): create the punktfunk group everywhere the udev rule needs it
60-punktfunk.rules chgrp's the usbip vhci attach/detach nodes to a dedicated
`punktfunk` group (security-review 2026-08-05 M-4: writing `attach` materialises
an arbitrary emulated USB device, so it must not ride on `input`). Four of the
six install paths shipped that rule in 0.25.0 without ever creating the group.
chgrp then failed, the nodes stayed root:root 0644, and the virtual Steam Deck
pad silently never attached — while `usermod -aG punktfunk` failed outright with
"group 'punktfunk' does not exist".
Affected and fixed:
* arch — post_upgrade() called only _ensure_update_group, so every box that
reached 0.25.0 by `pacman -Syu` missed it; post_install was correct.
* nix — no users.groups.punktfunk at all, though host.users' own description
already promised the usbip/vhci pad. Declares it now and adds
host.users to both groups.
* bazzite sysext — a group is host state and cannot ride an image, and the
deb/rpm scriptlets that would create it never run there.
* steamdeck install.sh/update.sh — handled `input` only. Both now create the
group and join it: running that script IS the statement "make my
Deck a host with native pad passthrough".
deb and rpm were correct throughout (one postinst/%post for install + upgrade).
Also on the Deck path: web.env secret hygiene. install.sh's `chmod 600` sat
inside the create-only branch despite a comment calling it "the idempotent belt
for a pre-existing file", and update.sh never touched the config dir at all — so
an install set up once and only updated since kept web.env world-readable
(0644) with the console password and session secret in it. Both scripts now
harden ~/.config/punktfunk to 0700 and web.env to 0600 on every run, and say so
loudly, because a chmod does not un-leak an already-readable secret: the
password still needs rotating.
Both group blocks are `if ensure_group ...` rather than `ensure_group || true`:
a failed groupadd must not fall through to a usermod against a nonexistent
group, which under `set -e` aborted install.sh after the long build and
update.sh before the service restart (verified: exit 6, no restart).
Docs: the group is now documented where people actually look — the per-distro
guides, install.md, steamos-host.md, a new troubleshooting entry for "pad
arrives as an Xbox 360 controller", and the uninstall pages. The 0.25.0 notes
gain the "group does not exist" caveat and turn the password bullet from
"consider rotating" into a real instruction, and CHANGELOG records the known
issue against the breaking change that introduced it.
Verified: bash -n on all four scripts; the arch scriptlet's post_upgrade driven
in a container (creates the group, idempotent on re-run); the ensure_group
helper and both membership branches, including a control that reproduces the
original bug (chgrp to a missing group leaves the node root:root 0644); the
find -perm /0077 probe across 0644/0640/0604/0600/0400 on GNU findutils;
`nix flake check --no-build` (the exact CI gate) and a NixOS eval showing
alice.extraGroups == ["input","punktfunk"]; docs-site build + typecheck.
|
||
|
|
12f39e1967 |
fix(steamdeck): the updater stops dirtying the checkout it just pulled into
`update.sh --pull` could abort with "Your local changes to the following files would be overwritten by merge: web/bun.nix" — before a single service was restarted — and the only way past it was to delete the file by hand. The updater did it to itself. web/bun.nix is generated (bun2nix, a pure function of web/bun.lock) but committed, because the Nix build fetches node_modules only from it. The web step ran `bun install --frozen-lockfile` without --ignore-scripts, so web's `postinstall` (`bun2nix -o bun.nix`) rewrote that tracked file on every update. Harmless while the committed file is in sync — but main carried a stale web/bun.nix from |
||
|
|
767e67caf4 |
feat(packaging): grant the host CAP_SYS_NICE, without which the GPU-priority lever does nothing
Wave-2 PW1, second half. The companion commit wires `PYROWAVE_QUEUE_PRIORITY` into the Linux PyroWave device; this is what makes it work on a packaged host. Measured on .21 (RTX 5070 Ti, NVIDIA 610.43.02), same binary in both arms: as packaged (no capability) every class refused, REALTIME *and* HIGH -> default priority same binary, cap_sys_nice+ep granted REALTIME on the FIRST attempt, no downgrade RADV behaves the same way. So this is not the RADV-specific "expect one downgrade to HIGH" the plan predicted — without the capability there is no elevated priority at all, on any vendor, and the knob is decoration. Worth being precise about what is being granted, because it is a network-facing daemon. CAP_SYS_NICE permits raising scheduling priority (nice, ioprio, affinity, RT class) and nothing else: no filesystem access, no network privilege, no user switching, and it is NOT setuid. The repo already ships exactly this capability on its gamescope binary for the same reason. Two side effects that will otherwise confuse someone debugging: a capability-carrying binary is AT_SECURE, so the loader ignores LD_LIBRARY_PATH/LD_PRELOAD for it (note this box was propped up by exactly such a shim during the ffmpeg-9 soname break — that workaround would now be silently ignored), and core dumps are suppressed by default. Per packaging path, because none of them are the same: - Arch: a `_grant_sched_capability` in the scriptlet, called from post_install AND post_upgrade — a replaced binary is a new inode, so the capability does not survive an upgrade by itself. - Debian: the same setcap in the postinst `configure` branch. - RPM: `%caps(cap_sys_nice=ep)` on the binary in `%files`, which is the rpm-native form — rpm then applies it on install, restores it on upgrade, and verifies it. A `%post setcap` does none of those. - NixOS: `security.wrappers`, because a store path is read-only and shared and cannot be setcap'd. The unit's ExecStart moves to `config.security.wrapperDir` — without that the wrapper exists and the service still runs the uncapped store path, which is the whole failure this fixes. - Steam Deck: setcap in the installer's sudo block. That box needs it most (one small Van Gogh GPU shared between the game and the encode). The binary lives under $HOME, so unlike the /etc drop-ins it survives a SteamOS A/B update on its own and needs no atomic-keep entry — but it does need re-applying after each rebuild, which re-running the installer does. - Bazzite sysext: at IMAGE BUILD time, before mksquashfs. It cannot be done in the merge hook (a merged sysext's /usr is read-only squashfs) and it cannot ride in from the RPM either — rpm keeps capabilities in its own header and `rpm2cpio | cpio` carries only the payload, so the staged file arrives with none. mksquashfs does record security.capability (only security.selinux is excluded), so a setcap on the staging tree is what lands in the image. Needs root/CAP_SETFCAP; a plain-user CI build warns and ships without it rather than failing a release over a performance lever. Every one of them is best-effort and cannot fail an install: a box without libcap, or a filesystem that cannot store capabilities, simply runs at default priority exactly as it does today. Documented in the same PR — the configuration row now says the packages grant it, and running-as-a-service gets a section explaining what it is, how to check it (`getcap`), and how to remove it (`setcap -r`, or just `PYROWAVE_QUEUE_PRIORITY=off`), including the two debugging side effects. Verified: the Arch scriptlet grants the capability from a fake package root exactly as pacman would invoke it, and the resulting binary reaches REALTIME end to end on the RTX 5070 Ti; the RPM spec's %caps line parses under rpmspec in a Fedora 41 container; the NixOS module parses under nix-instantiate; all five edited shell scripts pass `bash -n`. No Rust file changed in this commit, so the CI-parity Rust gates from the companion commit still stand. |
||
|
|
8f1c34c6bf |
fix(ci/arch): the release-rebuild prune called a helper that cannot exist there
The v0.25.0 rebuild published perfectly — registry has punktfunk-host 0.25.0-2 with
libavcodec.so=63-64, and it resolves on a real ffmpeg-9 box — then failed its last
step with
prune_release_assets: command not found
`. scripts/ci/gitea-release.sh` sources from the CHECKED-OUT TREE, and a release
rebuild checks out the OLD TAG. So the step could only ever see the helpers that
existed when that tag was cut, and the prune is gated on exactly that path: the
helper was guaranteed absent in the only case that calls it. Adding it to a shared
script made it look available at review time while being unreachable at run time.
Only the workflow file is read from the dispatched ref, so the logic moves there,
inline. Same reasoning documented at both ends, including the corollary worth knowing
before the next rebuild: a PKGBUILD fix made after a tag does NOT reach a rebuild of
that tag either — the packaging comes from the tag too.
Verified by executing the one-liner's exact bytes out of arch.yml under /bin/sh (the
shell Gitea actually uses): keeps the new -2 set and gamescope, drops the superseded
-1 packages and their .sha256 sidecars, leaves other legs' .dmg/.deb untouched. The
`'\n'` survives the shell quoting, which was the part worth proving.
Also drops the now-dead helper from gitea-release.sh rather than leaving a function
no caller can reach, and leaves a warning there against the next one.
|
||
|
|
e044f68500 |
fix(ci/arch): v0.25.0 shipped a host no Arch box can install, and nothing could tell
Arch moved FFmpeg 8 -> 9 (every libav soname +1) hours before the release. PR #108 fixed the real bug — packaging/arch/PKGBUILD now binds punktfunk-host to the sonames it actually linked, so pacman refuses an upgrade instead of bricking the install — and re-keyed ci/arch-ci.Dockerfile so the builder would carry FFmpeg 9. The tag was pushed four minutes later. arch.yml and docker.yml have no `needs:` between them, and arch.yml deliberately runs no -Syu ("the image's snapshot IS the build environment"), so the release build pulled the still-FFmpeg-8 `:latest` and published punktfunk-host 0.25.0-1 depends: libavcodec.so=62-64, libavutil.so=60-64, libavfilter.so=11-64, libavdevice.so=62-64, libswscale.so=9-64 against a world that had moved to 63/61/12/63/10. It fails safely — pacman refuses, nothing bricks — but it fails broadly: pacman prepares one transaction, so an unsatisfiable dependency of OURS stopped affected users' entire `pacman -Syu`. Nothing in the pipeline could have caught it. The existing assert proves the dep is VERSIONED; it cannot prove the version EXISTS. So two guards, plus the lever to repair a release that has already shipped: * Preflight parity — compare the builder's libav `provides` against the live repos and `-Syu` the container if they differ. The image is a cache and may lag; on this one axis it may not. Syncs into a throwaway --dbpath so the container never sits in the partial-upgrade state a bare `pacman -Sy` leaves. * Publish gate — resolve every built package with `pacman -U --print` against a PRISTINE --dbpath. Empty db means "nothing is installed", so every dependency must come from the repos exactly as on a user's box. Resolving against the builder's own installed set is what would hide this: a stale ffmpeg satisfies a stale bound. gamescope stays best-effort (dropped from the upload with a warning, never fatal). * workflow_dispatch(release_tag, pkgrel) — a published release cannot be repaired by re-running its tag: pkgrel would stay 1, which is invisible to a box that already recorded the broken build, and the workflow file at the tag can never carry inputs added after it. Dispatched from main it takes the WORKFLOW from main and the SOURCE from the tag, publishes to the stable repo at a higher pkgrel, and replaces the release-page assets (prune_release_assets: upsert replaces by NAME, and a rebuild's filenames differ, so the superseded package would otherwise stay one click away). Verified on a real ffmpeg-9 box (.21, CachyOS) rather than reasoned about: the gate rejects the published 0.25.0-1 host with the user-visible error verbatim, and passes client, web, scripting and gamescope — 0 false positives across all five artifacts. The parity snippet reads today's `provides` correctly (`-Si --dbpath` on an empty db works; pacman does not wrap fields when piped). Version logic exercised on all four paths: rebuild -> 0.25.0-2 stable, tag push and canary unchanged, pkgrel=1 refused. Ships as punktfunk-host 0.25.0-2. README gains the pacman error and what to do about it; CHANGELOG says plainly that 0.25.0's Arch packages were wrong. |
||
|
|
deeb8b6700 |
feat(pf-encode): build against FFmpeg 9
ffmpeg-next 8.1.0 could not accept FFmpeg 9 at all: ffmpeg-sys-next's version probe
covered avcodec majors 56..62 (the range is exclusive of its end), so libavcodec 63 fell
outside what it knew how to bind. 9.0.0 widens that to 56..63, which is what actually
unblocks Arch. Bump both pins — the unconditional Linux dep and the optional Windows
amf-qsv one — and the lock with them.
No API drift to fix. The crate major is a CEILING, not a target: one source tree still
spans FFmpeg 7.x/libavcodec 61, 8.x/62 and 9.x/63 via per-version cfgs, and every wrapper
symbol the NVENC-libav, VAAPI and amf-qsv backends name survives 8.1.0 -> 9.0.0
unchanged. The three hand-written #[repr(C)] hwcontext mirrors are the parts no compiler
checks, so they were re-read against the real headers rather than trusted:
AVCUDADeviceContext and AVD3D11VAFramesContext are byte-identical across 7.1/8/9, and
AVD3D11VADeviceContext gained two trailing UINTs in 8 that 7.1 lacks — which is why that
mirror deliberately stops at the common prefix, and why its assertions now say what they
do and do not buy you. They pin our layout, not libav's; a green build is not evidence.
The CI image is the step that makes this reach users. arch.yml deliberately runs no -Syu
("the image's snapshot IS the build environment"), so the builder stayed frozen on ffmpeg
8 no matter what Arch shipped, and a canary built from that snapshot could not satisfy the
soname dep the PKGBUILD now derives. Re-keying ci/ rebuilds it against ffmpeg 9.
Ubuntu and Windows deliberately stay put: the noble .deb bundles its own FFmpeg 8 behind
an rpath and strips the libav sonames from its Depends, and Windows bundles BtbN DLLs into
the signed installer — neither is exposed to the break, BtbN publishes no FFmpeg 9 build,
and moving either would re-qualify an encode stack to buy nothing.
Verified end to end on 192.168.1.21 (CachyOS, system ffmpeg 2:9.0-5, RTX 5070 Ti): host
builds clean and links libavcodec.so.63/libavutil.so.61/libavfilter.so.12/libswscale.so.10
with no unresolved sonames; the ffmpeg-8 compat shim is gone and the service runs with
NRestarts=0 and answers 401 on :47990; pf-encode's 67 tests pass; and a live synthetic
encode drives real NVENC hardware through FFmpeg 9's libavcodec to a decodable 1080p HEVC
stream (180/180 frames, FEC loopback 0 mismatches) with libavcodec.so.63 and
libnvidia-encode both mapped into the encoding process.
|
||
|
|
8551e88fcb |
merge: bring current main into the audio-substrate branch
Two conflicts, both unions of independent removals/fixes: main fixed the same three install.rs SAFETY comments this branch fixed (main's phrasing kept), and the runner provisioning drops BOTH env lines — main removed PF_FFVK_VULKAN_INCLUDE (pf-ffvk is gone since the FFmpeg replacement), this branch removed VBCABLE_DIR (the retirement). |
||
|
|
4a621de6b1 |
chore(packaging): retire VB-Cable — audio's substrate is Steam's drivers
The other half of the audio-substrate decision (spikes S2+S3 green, minted
endpoints landed in the previous commit): stop bundling a third-party
kernel driver the host no longer needs.
installer the VB-CABLE task, payload, silent-install run and the
donationware notice are gone; a suppressible notice tells
a Steam-less box that audio needs Steam INSTALLED (never
running) and that installing it later just works. A cable
from an older install is still deliberately not removed.
packer + CI -VbCableDir/VBCABLE_DIR, the staged-payload check and the
runner provisioning download are gone; SBOM drops the
redistributed-driver component.
winget the VB-Audio bundling-grant agreement becomes the honest
Steam requirement (surfaced on the unattended path where
no wizard is on screen).
docs windows-host/uninstall/security/echo say what actually
ships: no kernel-mode driver of our own, endpoints minted
from Valve's vendor-signed drivers, VB-CABLE mentioned
only as the historical fallback that keeps working.
host wording the mic-open guidance and module headers lead with Steam;
the NAME ladder itself is untouched — demoting 'cable
input' was considered and rejected (on a box where minting
transiently fails, the SSM would outrank an installed
cable, steal the silent sink, and make audio host-audible).
|