Commit Graph
2318 Commits
Author SHA1 Message Date
enricobuehlerandClaude Opus 5 6b3c582eb1 feat(client/present): use the driver's queue-free vblank mode where it exists
`VK_PRESENT_MODE_FIFO_LATEST_READY_EXT` is FIFO's tear-free vblank pacing that
presents the LATEST READY image at each refresh and retires the older ones,
instead of draining a queue. That is precisely what the software glass gate
emulates — so where the driver offers it, the driver does the job, and it does
it exactly where the gate matters most: a surface with no MAILBOX gets
newest-wins behaviour back without the app holding frames.

Found by asking the surface what it actually offers rather than trusting a
comment: the previous commit's `surface present modes` line read back
`[MAILBOX, 1000361000, FIFO]` on NVIDIA/Wayland, and 1000361000 is this mode.

The extension postdates the Vulkan headers ash 0.38 is generated from (1.3.281),
so there is no binding — hence the bare number in the log. It is hand-declared
here: mode value, extension name, and
`VkPhysicalDevicePresentModeFifoLatestReadyFeaturesEXT` spliced into the device
pNext chain. One trap worth naming: the SURFACE advertises the mode even with
the extension disabled, and using it on that basis is undefined — so the ladder
only offers it when the device feature actually came back true and we enabled it.

The gate/probe predicate had to split in two, and the distinction is the point:

* `needs_glass_gate()` — FIFO and FIFO_RELAXED only. NOT this mode: gating on
  top of a driver that already retires stale images would hold frames back to
  emulate something the presentation engine is doing, paying the serialisation
  twice, which is the ~27 ms the last commit measured.
* `vblank_locked()` — the whole FIFO family INCLUDING this mode, because it
  still presents on the refresh boundary, so the VRR cadence probe's premise
  ("with VRR off, a present waits for vblank") still holds.

Ranking: MAILBOX first (measured good at 1.4 ms), then LATEST_READY, then plain
FIFO — so a MAILBOX-less surface reaches newest-wins in the driver rather than
in our gate.

MEASURED ON GLASS (.21, NVIDIA 610.43.03, GNOME/Wayland): the extension probe,
feature enable and swapchain creation all succeed with a mode ash has no binding
for. Default ladder selects MAILBOX with `fifo_latest_ready=true`; the VRR ladder
selects `present_mode=1000361000` and measures `display 2.6 ms (pace 0.6 + latch
2.0)` — against 13-28 ms for plain FIFO + gate on the same box. The vblank-locked
path is now MAILBOX-class.

That changes the previous commit's reversal. The VRR ladder was reverted to
opt-in because it led with plain FIFO and cost ~27 ms; led with LATEST_READY it
costs 0.6 ms over MAILBOX. So `allow_vrr` is automatic again WHERE THE DEVICE
OFFERS THE MODE, and stays behind `PUNKTFUNK_VRR_FIFO=1` where it does not — on
those drivers the ladder would fall back to plain FIFO and the regression
returns. Both branches are pinned by tests. This also retires a dead switch: the
"Follow variable refresh rate" row did nothing at all after the reversal, and now
does something real on any driver with the extension.

⚠ Still unverified off this box: whether Windows and Intel drivers expose the
mode at all. Nothing measured here carries over — Windows Vulkan WSI goes through
DXGI, so exposing the enum and mapping it usefully onto flip-model semantics are
separate questions, and Intel is a different vendor stack again. Both facts are
logged unconditionally now (`surface present modes` + `fifo_latest_ready=`), so
one run on any box settles it. The code is safe either way: the mode is only
requested where the device feature enabled, and `allow_vrr` only goes automatic
there — everywhere else the shipped MAILBOX-first behaviour is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:01:56 +02:00
enricobuehlerandClaude Opus 5 e08474d96d fix(client/present): log the surface's actual present modes, and document the VRR opt-in
"AMD's Windows driver offers no MAILBOX" is the premise the FIFO glass gate is
built on, and it has been carried in a code comment rather than measured. Present
modes are a property of the (surface, device) pair — they vary by platform
surface, driver version and fullscreen state — so the only way to settle it is to
read them back from real machines. One unconditional log line makes every field
log answer the question.

First reading, .21 (NVIDIA 610.43.03, GNOME/Wayland):
  surface present modes available=[MAILBOX, 1000361000, FIFO]

Two things fall out. No IMMEDIATE and no FIFO_RELAXED on this surface, which is
why a PUNKTFUNK_PRESENT_MODE=immediate run reported mode=fifo — the pin was not
offered and the ladder fell through; previously that looked like a puzzling
result and is now evidence. And 1000361000 is
VK_PRESENT_MODE_FIFO_LATEST_READY_EXT: FIFO's tear-free vblank pacing that
presents the LATEST READY image instead of draining a queue — the driver-native
version of what the glass gate emulates in software, and a candidate to replace
it wherever the driver exposes it (needs VK_EXT_present_mode_fifo_latest_ready
enabled at device creation, so a work package rather than a tweak).

Also documents PUNKTFUNK_VRR_FIFO, which the previous commit introduced without
a docs entry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:01:56 +02:00
enricobuehlerandClaude Opus 5 f422ae3e38 fix(client/present): what the first on-glass session found, including a reversed default
WP6 ran against .21 (CachyOS, RTX 5070 Ti, NVIDIA 610.43.03, GNOME/Wayland,
1080p60 HDMI, VRR provably disabled — `org.gnome.mutter experimental-features`
is empty), host and client on the same box, `VK_KHR_present_wait` available.

Five defects that unit tests and both CI gates had passed over:

1. The latch learner and the VRR probe observed NOTHING. Both derived spacings
   with `windows(2)` inside a single batch, but the run loop drains present-wait
   samples every pass, so a batch is normally ONE stamp. `period_us` read back
   exactly the mode fallback — correct by luck on a 60 Hz panel, wrong the moment
   a mode lies, which is the entire reason PanelGrid exists. The tests fed
   40-stamp batches, a shape the live loop never produces. Spacings are now
   measured against the previous stamp across calls.

2. The VRR reference was circular. It compared spacings against the LEARNED
   period, but the grid cannot be learned from our own presents when the stream
   runs below panel rate — we only ever observe multiples ≥ our frame interval,
   so the learner adopts our own cadence and every delta is on-grid by
   construction. It learned 18-22 ms from a 40-50 fps stream and reported VRR on
   a display with VRR off. The reference is now the DISPLAY MODE's period, which
   is the vblank grid presents actually quantize to.

3. The probe is meaningless outside FIFO. MAILBOX deliberately decouples presents
   from scanout, so its stamps are never grid-quantized: same panel, same minute,
   FIFO read `no` (correct, period 16.4 ms) and MAILBOX read `yes` (wrong).
   Outside a FIFO-family mode the honest answer is Unknown, and that is now what
   it reports.

4. Round evaluation was per-CALL rather than per-sample, so the verdict depended
   on how the caller batched its stamps. Closed inside the sample loop now, with
   a test pinning bulk-vs-one-at-a-time equivalence — the same invariant (1)
   violated, in a second place.

5. `force_latency` was dead code without the `pyrowave` feature: a warning in the
   `--no-default-features` build CI actually ships (the Windows ARM64 leg). The
   gate only ever tested default features; it now tests both.

DESIGN REVERSAL — the VRR FIFO-first ladder is opt-in (`PUNKTFUNK_VRR_FIFO=1`),
no longer default. It shipped default-on for `allow_vrr` + fullscreen, which is
the default configuration. Measured A/B, same box, back to back, reproduced
across three runs: FIFO+engine `display 28.4 ms (pace 11.8 + latch 16.6)` versus
MAILBOX `1.4 ms (0.2 + 1.2)`. Under a compositor the FIFO present's on-glass
confirmation arrives a whole refresh later and the presenter serialises behind
it. The VRR upside is real in principle but UNMEASURED — no VRR panel was
available — and a default that is measurably ~27 ms worse on the hardware we
could test, bought against an unproven win on hardware we could not, is the
wrong way round. A test pins the default to MAILBOX; flip it back when a VRR
panel confirms the win.

NOT measured, and not claimed: the FIFO glass gate's own headline. The standing
queue only forms when the stream rate approaches the panel rate, and an idle
GNOME desktop is damage-driven at 40-50 fps on a 60 Hz panel, so `gated`/`forced`
read 0 in every mode and the mechanism never engaged. The 11-13 ms figure is
still the code's inherited documentation, not a fresh measurement. It needs its
actual target: AMD-on-Windows (no MAILBOX, direct scanout) under load.

Rig caveats recorded rather than smoothed over: host and client shared one GPU,
so absolute latencies are contended and run-to-run variance was large, and it
could not be visually confirmed what the physical screen showed. Mode selection,
the fallback ladder, the VRR verdict and the counter plumbing are robust to
that; absolute numbers are not.

Gates: fmt, clippy -D warnings over the five client crates AND the
`--no-default-features` build (added because defect 5 hid there), 160 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:01:56 +02:00
enricobuehlerandClaude Opus 5 e38e3c44c9 feat(client/present): V-Sync and VRR become real settings, and VRR is measured
WP3 of design/desktop-presentation-rebuild.md. The `vsync` and `allow_vrr`
settings have existed since WP1 but nothing consumed them — the swapchain picked
MAILBOX-or-FIFO once, from an env var, and froze. This makes them mean
something, which is also what unblocks their settings rows (deliberately
withheld from WP5 rather than shipped as dead switches).

Present-mode selection is now a preference ladder, not a constant:

* V-Sync off — IMMEDIATE, then FIFO_RELAXED, then the tear-free modes. Asking
  to tear and silently getting vsync is a lie, so the mode that actually took is
  named in the stats line and a refused preference is logged requested-vs-active.
* V-Sync on + VRR allowed + fullscreen — FIFO first. On a variable-refresh panel
  with direct scanout the FIFO present IS the flip, so the panel follows the
  stream's cadence instead of a fixed grid; MAILBOX would decouple presents from
  scanout and re-quantize to the compositor's clock. This is only safe because
  WP2's glass gate bounds the standing queue that historically made FIFO costly.
* Otherwise — MAILBOX then FIFO, the shipped default, unchanged.

`PUNKTFUNK_PRESENT_MODE` still pins a mode outright and now falls back to the
settings (rather than to mailbox) when the name is unknown.

VRR detection is MEASURED, never queried. No portable query exists — SDL exposes
none, Wayland does not report adaptive-sync state, Windows surfaces nothing
through Vulkan — and the platforms that do answer have been caught lying (see
the Android per-uid refresh-rate finding). The discriminator is quantization: on
a fixed-refresh panel every on-glass instant lands on the vblank grid, so the
spacing between presents is ~k×period for whole k even when the stream runs
slower than the panel (it just picks a larger k); under real VRR the panel
refreshes when we present, so the spacing follows our own cadence and sits off
the grid. `CadenceProbe` folds each delta to its distance from the nearest
multiple of the learned period and takes the median. Tri-state: it stays Unknown
below 24 deltas and after a display change, so `vrr` is reported only when it
has been measured — never inferred from what the display claims.

Also fixes the read-once refresh rate: `native.refresh_hz` was sampled at
startup and never revisited, so dragging the window to another monitor left a
60 Hz-seeded clock pacing a 144 Hz panel. `WindowEvent::DisplayChanged` now
relearns the latch grid, resets the cadence verdict, and clears the served-slot
latch.

Settings rows for both, on all three surfaces (GTK, WinUI, console). The
console's V-Sync row is reachable in Gaming Mode, which is the only editor a
Deck user has.

Gates: punktfunk-rust-ci linux/amd64 — fmt, clippy -D warnings over
pf-client-core, pf-presenter, pf-console-ui, the session binary and the GTK
client, 160 tests (the two new ones cover every ladder and both cadence
regimes, including the case that matters most: a stream slower than a FIXED
panel must still read as fixed). WinUI leg on the Windows runner .133:
clippy=0 tests=0, against a tree proven by content to contain the edit.

⚠ On-glass validation is still owed and is NOT claimed here: every box with a
real display was powered off when this landed, so the VRR ladder and the
detector have been exercised only against synthetic stamps in unit tests.

Rebase follow-up: `20de58a7` landed the same "panel grid can be wrong in both
directions" defect fix on Android and extracted the corrected learner into
`punktfunk_core::phase::PanelGrid` for the iOS and desktop presenters to share.
This clock had the identical bug — it capped the learned period at the display
mode's refresh, and the mode is only a CLAIM, so a display really running slower
than it advertises pinned a grid whose instants never arrive, for the session,
with no way back. Adopted the shared learner rather than carrying a second,
buggier copy; still fed the window's MIN spacing, which preserves the k×period
resistance the cap was actually aimed at while the streak requirement lets a
genuinely slower panel be discovered. New test: seed 120 Hz, real panel 60 Hz,
the clock must climb back out.

Took the same commit's third lesson too: the adaptive margin widened on a
latch over 1.5×period (a number picked here), and now widens on the latch
exceeding one period plus the lead already applied — the slot actually aimed at.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:01:56 +02:00
enricobuehlerandClaude Opus 5 b1ac4d02de feat(client/present): the display stat splits, and the intent reaches the settings UI
WP4 + WP5 of design/desktop-presentation-rebuild.md, on top of the WP1/WP2
engine. The engine shipped with no way to choose it and no way to see what it
cost; this closes both.

WP4 — the display stage splits into `pace` (decoded → present-submit, our own
pipeline) + `latch` (submit → on-glass, the presentation queue and the vblank
wait), off the `submitted_ns` stamp WP2 already carried. That split is what
makes a high `display` self-diagnosing: latch dominating is the vsync floor or
a standing queue, pace dominating is us. A `present:` line joins the Detailed
tier naming the live swapchain mode — the answer to most "why is my latch a
whole refresh" questions, since a MAILBOX request silently lands on FIFO
wherever the driver has no mailbox — plus the engine's counters, rendered only
when they are non-zero so a healthy latency session shows just the mode.

Deviation from the plan: the planned `display_adj` twin is NOT here. It was
specified as `display − latch_p50` for parity with the Apple HUD's shaved
figure, but with a real per-sample `pace` percentile that twin is the same
quantity derived worse (subtracting percentiles). `pace` IS the
Apple-comparable number — Apple subtracts its OS present floor, the latch is
ours — and the user docs now say exactly that.

WP5 — Prioritize + Smoothness buffer on all three surfaces: the GTK dialog (a
new Presentation group on the Display page), the WinUI settings page, and the
console settings screen, which is the ONLY editor reachable in Gaming Mode and
so the one that decides whether Deck users can reach this at all. The buffer
control follows the intent the way echo cancellation follows the mic: hidden on
the desktop shells, dimmed and inert on the console, where a row that vanished
mid-list would shift everything under the cursor.

The V-Sync and VRR rows are deliberately NOT here. Their settings exist and are
profile-routed, but the swapchain does not honour them until WP3, and a toggle
that does nothing is exactly how "Full chroma (4:4:4)" shipped inert on desktop
for three releases after being announced.

Buffer labels carry no millisecond hints (Apple/Android derive them from the
session refresh): under a Native mode the shells do not know the refresh at
settings time, so the captions state the cost as one refresh per frame rather
than a confident wrong number.

Docs: the stats page documents the split and the `present:` line, and stops
claiming Linux/Windows measure to the present instant (untrue since
present_wait); client-settings documents both new rows and drops the stale
claim that the desktop 4:4:4 toggle has no effect (it was wired to
VIDEO_CAP_444); configuration documents PUNKTFUNK_PRESENTER and
PUNKTFUNK_PRESENT_DEBUG.

Gates: punktfunk-rust-ci linux/amd64 — fmt, clippy -D warnings over
pf-client-core, pf-presenter, pf-console-ui, the session binary and the GTK
client, 158 tests. The WinUI leg cannot be reached by any Linux or macOS check,
so it was compiled on the Windows runner .133: clippy -D warnings and tests
both exit 0, against a tree proven by content to contain the edit. ⚠ The first
run there reported a false pass — the script printed its done-marker while the
log carried a test failure (a STATUS_DLL_NOT_FOUND launch failure, ffmpeg's
DLLs missing from PATH); the harness now echoes each phase's exit code so the
verdict is a fact in the log rather than an inference from a marker.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:01:38 +02:00
enricobuehlerandClaude Opus 5 5f55fa874a feat(client/present): the desktop presenter gains the Apple/Android intent model
WP1+WP2 of design/desktop-presentation-rebuild.md. The shared Linux/Windows
session client presented arrival-paced with no pacing layer at all: two depth-2
newest-wins hops into a drain-to-newest and an immediate present. That IS the
lowest-latency intent, but it was unnamed, unselectable, and had no alternative
— and on a surface without MAILBOX (AMD's Windows driver offers none, and any
compositor holding images does the same) the swapchain's own FIFO becomes a
standing queue worth a measured 11-13 ms at 60 Hz.

WP1 — the settings cluster, under the keys the Apple client already writes into
the shared profile catalog (present_priority / smooth_buffer / vsync /
allow_vrr): mismatched names would ride SettingsOverlay::extra, carried but
never applied. PresentPriority::resolve mirrors the Android reference exactly
(anything but an explicit "smooth" is latency; a buffer outside 1..=3 becomes
2), so a profile authored on any client means the same thing on all of them.
Only the first two are consumed here; vsync/allow_vrr land in WP3.

WP2 — the engine (present_pace.rs, pure state + arithmetic, 6 tests):
- FrameStore: newest-wins slot, or the smoothing FIFO with preroll-to-capacity,
  drop-oldest overflow, and an underflow that re-arms the preroll (repeat by
  omission) — the Apple/Android semantics, with qDrop/qDry counters.
- LatchClock: the panel grid learned from VK_KHR_present_wait glass stamps,
  min positive spacing capped by the mode refresh (measured, never queried —
  VRR and Android's per-uid refresh lie both punish trusting a reported rate).
  It now also publishes the host-facing LatchGrid, so the phase-lock report and
  the local scheduler cannot disagree about the grid.
- PresentGate: one undisplayed present in flight on FIFO surfaces, with the
  100 ms stale force-open. This is the standing-queue killer, and it is inert
  on MAILBOX/IMMEDIATE and without present timing — where behaviour stays
  byte-for-byte the shipped arrival pacing.

Wiring: glass samples drain every pass (a 1 Hz batch would starve clock and
gate) and the waiter pushes an SDL wake, so a gate reopen never waits out the
event timeout; smoothness serves one frame per latch slot and tightens the
loop's wait to that deadline; the adaptive slot margin starts at 0 and widens
+500 us per missed window toward 2.5 ms (a fixed lead was measured to be pure
display tax). PUNKTFUNK_PRESENTER=arrival disables the whole engine for field
A/B without a rebuild.

PyroWave collapses smoothness to latency for the stream: its plane-ring
retirement accounting assumes the depth-2 newest-wins hand-off, and all-intra
frames make buffering moot anyway.

Gates (punktfunk-rust-ci, linux/amd64, sources touched first so a warm target
cannot print a vacuous Finished): clippy -D warnings across pf-client-core,
pf-presenter and punktfunk-client-session; 80 + 32 tests pass; rustfmt clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 00:00:50 +02:00
enricobuehlerandClaude Fable 5 d839f4c2b6 fix(client/windows): settings stop going stale behind your back, and the log has a door
A field reporter's codec setting "changed by itself" between sessions. Nothing writes
the negotiated codec back — what they saw was a stale snapshot. `AppCtx.settings` is
loaded ONCE at process start and the page renders from it, but this process is not the
file's only writer (the spawned session persists its match-window size, the console UI
and Decky save too), so the page showed values another process had already replaced —
until a row was touched and `commit`'s rebase pulled the file in, at which point the
value visibly jumped. The 2026-07-31 rebase fix covered the whole-file writers and
missed two spots: nothing re-based on page ENTRY, and the profile-scope commit arm
cloned the snapshot without reloading, so overlay absorption diffed against stale
globals. Both now re-base on the file.

Two more ways a setting could vanish or cost time:

* An older binary's whole-file save DROPPED a newer client's keys — `Settings` had no
  unknown-key passthrough, unlike `SettingsOverlay`, whose `extra` map already gives
  profiles exactly that contract. Extended to the globals: additive, empty on every
  existing store, and an empty map serializes to nothing so no file churns. (`save()`
  was already temp+rename, so the torn-file → silent-Default reset was closed.)
* "Check the client log" never said WHERE. Settings ▸ About grows an Open log folder
  row (%LOCALAPPDATA%\punktfunk\logs, folder not file so the rotated .old generation
  is in reach), and the failed-spawn banner now names the path.

The 4:4:4 caption said "HEVC only, and only where the host can encode it", which sends
people hunting: the host gate is PyroWave or an NVENC backend. It says so now.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 23:49:45 +02:00
enricobuehler 0de161e29b Merge pull request 'feat(clients/input): controllers can stop being forwarded, for couches that hand the pad over another way' (#22) from worktree-gamepad-passthrough-toggle into main
Reviewed-on: unom/punktfunk#22
2026-08-02 21:38:01 +00:00
enricobuehlerandClaude Opus 5 b297542c4d feat(clients/input): controllers can stop being forwarded, for couches that hand the pad over another way
A controller that reaches the host by USB passthrough — VirtualHere and friends, or simply a
pad plugged into the host — arrived there twice: once as the real device, once as the virtual
pad this client built from the same hands. Games read both, so a stick drifts against the
centred second pad and menus take every input twice.

New per-client setting, "Forward controllers", default on (today's behaviour). It is tier-P,
so a profile can decline what another profile forwards.

On Linux and Windows it is deliberately stronger than "send nothing". Opening a controller is
what CLAIMS it — SDL's HIDAPI drivers take the device node — and a claimed device is one a
passthrough tool cannot bind, so with this off the session opens no slot at all and never
enables the Valve HIDAPI drivers. Menu navigation is untouched: the launcher still opens the
active pad, and a session supersedes menu mode whether it forwards or not, so the pad is free
for the whole time a stream is up. The consequence, documented at both the setting and the
chord: the controller escape chord is read off forwarded pads, so it is unavailable there.

The Apple and Android input stacks claim nothing, so those clients keep their slots and their
chords and only gate the wire sends — losing tvOS's only controller way out of a stream would
have been the worse bug. Android does stop its DualSense and Steam Controller 2 USB captures,
which do claim the device.

Surfaces: GTK, WinUI, the console settings screen, Apple's touch and gamepad settings, the
Android touch and gamepad settings, and Decky (which also hides the rows that now have nothing
to act on). Everywhere the "which pad" and "pad type" rows grey out while it is off.

Verified: cargo clippy --all-targets -D warnings + 79 tests on pf-client-core, pf-console-ui,
punktfunk-client-session and punktfunk-client-linux (linux/amd64 container, gate proven
non-vacuous with a planted error); swift build for the Apple clients; gradle compile + 49 unit
tests for Android (likewise proven); tsc for Decky. clients/windows is UNCOMPILED — both
Windows boxes were offline; its edits were reviewed against the helper signatures by hand.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 22:21:26 +02:00
enricobuehlerandClaude Fable 5 98e040fd01 fix(host/stream): the wire holds the session rate when the display outruns it
PUNKTFUNK_VDISPLAY_HZ_MULT promises extra display refreshes without one
extra frame on the wire, but the frame-driven trigger enforced its pace only
as a per-gap floor: sleep to 0.9×interval, then wake on arrival. A source
that always has a frame pending — the overdriven display under uncapped
content — settled at 0.9-interval spacing, 1.11× the negotiated rate. That
is the field report's 132 fps on a 120 fps session: ten percent more
bitrate, encode and decode for frames a 120 Hz panel can only drop.

A credit bucket (PaceBudget) now pins the long-run average at the pacing
rate: credit accrues at one frame per interval of real elapsed time, capped
at 1.25 frames of post-stall burst, and every submitted frame spends one. A
grab may run early only against banked credit, so the 0.9 floor keeps its
per-gap jitter headroom while the average cannot exceed the rate — and a
source at or below it banks faster than it spends and is never delayed.
Anchoring to real elapsed time also keeps the synchronous-encode overlap the
arrival-anchored floor bought (the owed fraction absorbs a constant encode
tail instead of stacking on top of it), and it cannot fight the phase lock's
submit grid: both agree the period is the interval.

The charge lives under the same guard as the gate — the legacy fixed tick
paces by its own grid, and charging it without ever accruing would bank
unbounded debt that stalls the loop if a rebuild later flips the capturer to
arrival-wait.

Verified on .25: native::stream tests 15/15 (three new PaceBudget tests),
punktfunk-host 369/369, clippy -D warnings clean, fmt clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 21:54:09 +02:00
enricobuehlerandClaude Fable 5 5174a59832 fix(capture/kwin): a hidden cursor leaves the stream — KWin's id-0 meta is the hide
Since the 0.22.0 cursor work (the seat-pointer park + the metadata
composite), a KWin capture-model stream always has a cursor — and it never
went away again: not in game, not in Big Picture, not with a controller in
hand (field report, 2026-08-01). The host blended the arrow forever because
pf-capture deliberately ignores SPA_META_Cursor id 0, and once `visible`
latched true nothing on Linux ever cleared it.

Two producer contracts meet on id 0, and one flag now carries which one a
stream follows. KWin rewrites the cursor meta on EVERY enqueued buffer and
writes id 0 whenever Cursor::isOnOutput says the pointer is not in this
stream — which covers both a globally hidden cursor and a client null-cursor
surface (empty geometry intersects nothing). There id 0 IS the hide, and
honoring it is what lets a game hide the pointer mid-stream. Mutter only
rewrites a buffer's meta when the cursor changed, so recycled buffers carry
stale id-0 regions between damage frames — honoring those flickered the
cursor off between hovers (on-glass round 5), and that path keeps its
last-known-state behavior.

The flag rides from the backend that created the output (correct for
registry-pooled reuse too — a kept display only ever matches its own
backend) through capture_virtual_output into the parser's CursorState. The
portal-monitor path stays on the stale-meta contract: the only thing routed
through it today is Mutter's HDR mirror.

Verified on .25: pf-capture 45/45, punktfunk-host 369/369, clippy
-D warnings clean (pf-capture, punktfunk-host, cursor-probe), fmt clean.
On-glass KDE validation still owed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 21:54:09 +02:00
enricobuehlerandClaude Opus 5 c2a6d30d7b fix(android/decode): a codec input slot the feeder can't fill goes back, and so does the AU
`AMediaCodec_getInputBuffer` returning null for an index the input-available
callback had just handed us dropped both the slot and the access unit on the
floor. Every sibling path in this loop recycles the slot — the orphan-part
discard and the oversize drop both say so in as many words — because nothing was
written and nothing was queued, so it is still ours. Forgetting it leaks one of
the codec's input buffers per occurrence: we never use it again and the codec
never frees what it never received, so the pipeline runs out of input slots,
`pending_aus` overflows into its drop-oldest arm, and the resulting keyframe storm
reads as a decode fault rather than a bookkeeping one.

The AU went with it, silently — no keyframe request, no freeze gate, unlike every
other loss path here — leaving a hole in the reference chain whose concealment
was free to reach the screen.

Both go back now. `break` rather than `continue`, because a codec that cannot
hand out an input buffer it has just advertised is in no state to be fed the rest
of the parked queue on this pass, and retrying the same index against every
parked AU would burn the whole backlog for nothing; the loop comes round again on
the housekeeping wake within 5 ms if it was transient.

Gates: cargo ndk check green on arm64 and armv7, fmt clean, Android clippy at the
same 4 pre-existing warnings as the base commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 20:15:45 +02:00
enricobuehlerandClaude Opus 5 20de58a78a fix(android/present): the panel grid can be wrong in both directions, and the margin listens to the latch
Three defects in the 0.23.0 timeline presenter, all found while root-causing the
field report that turned out to be the slice wire. None of them is that bug; all
three are real, and the first is the one that would still bite once it is fixed.

The panel-period learner could only ever narrow. It is seeded from the display
mode Kotlin asked for — and `preferredDisplayModeId` is a REQUEST the system may
refuse (Smooth Display off, battery saver, thermal, an OEM governor). Ask for
120 Hz on a panel that stays at 60 and the presenter pins an 8.33 ms grid on a
16.67 ms display with no way back, for the rest of the session: it then aims at
instants that never arrive and releases faster than the panel scans. The learner
moves both ways now, and lives in `punktfunk_core::phase::PanelGrid` where it is
host-testable and where the iOS and desktop presenters can share it. The
asymmetry is kept and made explicit — narrowing is immediate (a finer real grid
is always safe to subdivide onto, and it is the per-uid down-rate case the seed
most often gets wrong), widening needs eight consecutive agreeing observations
and then takes the narrowest of them, because one wide sample is a missed
callback and eight in a row is a display that really did slow down.

The glass budget was a prediction with nothing underneath it. `OnFrameRendered`
already reports what actually reached glass, but the budget never consulted it,
so a wrong grid could hand SurfaceFlinger frames indefinitely: BufferQueue fills,
MediaCodec runs out of output buffers, the decoder stalls, and the no-output
backstop starts begging for keyframes. Releases are now counted against their
confirms and the presenter holds back past six outstanding — loose on purpose,
since the callbacks are allowed to arrive batched and a held frame in the
newest-wins slot is a dropped one. It self-clears when the confirms catch up, and
writes the ledger off after the same 100 ms the stale reopen uses, so a platform
that stops confirming can never wedge the stream. `qWait` and `unconfirmed` join
the 1 Hz pf.present line, which is what would have made this visible from a log.

The adaptive latch margin widened on `paced_drops` — the newest-wins store's own
policy evictions, which happen whenever the stream out-runs the panel and say
nothing about SurfaceFlinger's latch lead. On a healthy device that walked the
margin to its 2.5 ms ceiling and re-imposed the display latency the P2e sweep had
just measured away. It now widens on the measured latch exceeding one panel
period plus the live margin, which is what a missed vsync actually looks like.

Also corrects two doc comments that named `display.refreshRate` as the panel_hz
source; it has been the mode table since the A024 down-rate fix.

Gates: 278 punktfunk-core lib tests (7 new PanelGrid cases incl. the refused-mode
regression), clippy -D warnings and fmt clean, cargo ndk check green on arm64 and
armv7. Android clippy reports the same 4 warnings as the base commit and no new
ones. NOT yet confirmed on glass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 20:15:45 +02:00
enricobuehlerandClaude Opus 5 97b2c01ac1 fix(core/packet): a slice-streamed frame costs its own size, not the whole frame ceiling
The 0.23.0 slice wire flushes a block every MIN_STREAM_BLOCK_SHARDS, so every
ordinary access unit is now opened by a SENTINEL — a header with no totals. The
reassembler sized those frames at `max_frame_bytes`, which the QUIC handshake
clamps to 8-64 MiB. That was survivable while sentinels were rare (the streamed
path emitted one only for an AU exceeding a whole FEC block, ~281 KB); it is not
survivable now that every frame is one.

Two consequences, both measured: each access unit allocated and ZEROED a
multi-megabyte buffer, and the in-flight budget (IN_FLIGHT_BUF_FACTOR x
max_frame_bytes) was spent after ~3 concurrent frames — with production geometry,
12 ordinary AUs in flight lost 9 of them outright, every packet dropped before it
could be placed. On a link with normal reorder that is a permanent loss storm:
frames never complete, the re-anchor gate freezes the picture, and the client begs
for keyframes. Only clients advertising VIDEO_CAP_MULTI_SLICE reach this path —
Android and the Linux/Windows session client; Apple and the Windows in-process
client never did, which is why it read as a platform-specific "video pipeline"
fault in the field.

A sentinel carries no total but does pin its own block's extent: a slice sentinel
by its wire base, a legacy one by its full-K position. Size the buffer to that and
grow as later blocks (or the final block's totals) reveal more. The budget is
re-checked on growth for the same reason it is checked at open.

The same flush also drained `pending` to empty whenever the AU's length was an
exact multiple of the shard payload, leaving `finish_streamed` to seal a final
block of one zero-padded FILLER shard. Its derived base overlapped the block
flushed a moment earlier, retro-validation correctly read that as a lying header,
and the whole AU died — one frame in every 1408 on a 1500-MTU link, ~12 s apart at
120 fps, each costing a freeze and a recovery keyframe. A flush now keeps one
whole shard back, restoring the invariant `StreamedAu::pending` already documented.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 20:15:45 +02:00
enricobuehler 29473d6280 Merge pull request 'fix(client/ios): Escape keeps the pointer captured instead of handing it back to iPadOS' (#19) from worktree-ipad-esc-pointer-relock into main
Reviewed-on: unom/punktfunk#19
2026-08-02 17:27:54 +00:00
enricobuehlerandClaude Opus 5 b6acbd096e fix(host/vdisplay): waking the PC stops failing the first session
A woken Windows host refused every connection with "pf-vdisplay driver
interface not found", on a box where the driver was installed and running.

Resuming re-enters D0 and re-registers the IddCx control interface while
the rest of the resume storm is still going. A client reconnecting a
second after wake lands inside that gap. `ensure_available` probed
exactly ONCE, so it read the gap as a dead driver and answered a device
that was seconds from ready by disabling and re-enabling it — then gave
the interface 4 s to come back, which a contended post-resume PnP does
not meet. The session failed, and the log blamed a missing install.

The recovery also could not tell whether it had recovered anything. It
ran the whole cycle under `SilentlyContinue` and reported
`(Get-PnpDevice).Status` — the DEVICE's status, not the cycle's outcome —
so a disable that was REFUSED left the adapter untouched, started, and
reading `OK`. That is the reporter's `cycled the adapter device …
status=OK` line: a recovery that never happened, announcing success. And
a refusal is the expected case here, not the exotic one:
reset-pf-vdisplay.ps1 stops the host service first precisely because the
host holds the driver's control device open, a step an in-process cycle
structurally cannot take.

- Distinguish a devnode MID-TRANSITION (interface registered, not started
  yet, or the open refused) from one genuinely ABSENT. Wait the first
  out; only the second earns a reload. `Probe` carries the counts.
- Report what the reload DID, not what the device looks like afterwards:
  every failable step is `-ErrorAction Stop` in a `try`, and
  `pnputil /restart-device` is the fallback for the in-use device that
  `Disable-PnpDevice` refuses. Failure paths re-enable, so a half-cycle
  can never strand the adapter DISABLED.
- Give the interface 15 s to arrive after a reload, not 4 — under a 30 s
  hard ceiling so a permanently wedged devnode still fails predictably.
- Serialize recovery: N sessions racing in after a wake perform ONE
  reload, not N interleaved ones. The lock is taken only where no manager
  lock is held, so the order stays one-way.
- Retire the manager's cached control handle when a reload runs, instead
  of letting the next session discover it via a failed IOCTL.
- Surface the real reason. `ensure_available` returns `Result`, so the
  log names how long it waited, whether a reload ran, and how many
  interface instances were seen in what state — the detail that would
  have identified this from the field report's log alone.

`VdisplayDriver::open` now shares the wait (brief, no reload) instead of
carrying a second, drifted copy of it — that path is also reached by
`hw_cursor_capable` mid-handshake, where a reload would be the wrong
trade for one capability bool.

Windows-gated, so verified with scripts/xcheck.sh (check + clippy -D
warnings, --all-targets) and cargo fmt; on-glass wake test still owed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 13:12:54 +02:00
enricobuehlerandClaude Opus 5 d63e913f52 fix(client/ios): Escape keeps the pointer captured instead of handing it back to iPadOS
iPadOS releases the scene's pointer lock by itself when Escape is pressed — the platform's
built-in "let me out", mirroring the web Pointer Lock API's default unlock gesture. Nothing in
our code does it: a bare Esc never touches `captured`, and it keeps forwarding to the host as
the game key it is. But the lock going away flips the mouse onto the absolute UIKit path and
un-hides the iPadOS cursor, so pressing Esc for an in-game menu silently cost the capture until
the user clicked into the video to win it back.

Esc is a GAME key in a stream, not a request to hand the pointer back to iPadOS, so an unwanted
drop is now re-requested. `syncPointerLock` arms a short, bounded burst (3 attempts over ~0.6 s,
no restart inside 2 s) whenever the lock is wanted, was previously HELD, and is now gone; the
first attempt re-asserts `prefersPointerLocked`, later ones present a real false→true transition
and re-anchor the PointerLockChain. Every deliberate release (⌘⎋, ⌃⌥⇧Q, the Stream menu,
resigning active) clears `captured` first, so `wantsPointerLock` is already false when their drop
is observed and none of them are fought.

The "previously held" half of the condition keeps a scene that never qualifies (Stage Manager,
Split View) from paying for a lock that isn't coming — there, a first grant is still driven by
the chain engage in setCaptured/viewDidAppear exactly as before.

While a re-lock is in flight the local cursor stays hidden and absolute pointer MOTION stays
muted, so the couple of frames it takes read as "Esc did nothing to my mouse" rather than a
cursor that blinks in and out and a host cursor that teleports to the pointer's absolute
position. Buttons still forward (they carry no position), so a click mid-relock isn't swallowed.
The burst clears itself on give-up, so the cursor can never stay hidden on a lock the system
won't grant.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 12:08:31 +02:00
enricobuehlerandClaude Opus 5 362595b20f fix(vdisplay/kwin): the streamed output declares it mirrors nothing, so a stored replicationSource can't clone a panel into the stream
KWin stores output configuration per *setup* — the exact set of connected
outputs, matched by EDID/connector — in `kwinoutputconfig.json`, and
`replicationSource` is one of the fields it saves and restores
(`OutputConfigurationStore::storeConfig` / `setupToConfig`). Our virtual output
carries a STABLE name on purpose, so once any setup has an entry making
`Virtual-punktfunk` a mirror of a physical head, KWin re-applies it to OUR
output on every session that reproduces that same monitor set — and only that
set, which is why the failure looks environment-dependent: a field report has
the stream cloning the panel whenever exactly one monitor is live, and behaving
normally the moment the others come back (a different setup key, a different
stored entry).

A mirroring output is not a desktop. KWin's `applyMirroring` overrides its scale
and render offset to the source's, so the stream carries the physical screen's
viewport at the physical screen's size instead of the mode the client
negotiated. The protocol says the rest out loud on `priority`: "an output may
not be in the output order if it's disabled or mirroring another screen" — so
the primary assertion this module works so hard to verify silently stops meaning
anything too.

Nothing we sent ever contradicted the stored value. The topology config enabled
our output, took priority 1 and disabled the others, but never stated the one
property that decides whether the thing is its own screen. Now it does:
`set_replication_source(ours, "")` rides along in the config we already build
(free, idempotent — an empty source is exactly what KWin resolves to "mirrors
nothing"), gated on management v13 where the request appeared, since wayland-rs
does not range-check requests and an out-of-range opcode would kill the
connection.

`extend`/`auto` issue no topology calls by design — the streamed output is meant
to join the desk without rearranging it — but a mirror is not an arrangement, it
is a broken source under every topology. So they get `clear_replication_source`,
which enumerates and applies ONLY when our output really is mirroring.

The device's `replication_source` event is now read, so the state is visible: a
mirrored streamed output names its source in a warn instead of leaving "the
stream just shows my monitor" as something only the reporter can see.

Verified on 192.168.1.25 (Ubuntu, cargo 1.96): `cargo test -p pf-vdisplay` 128
pass (7 in `kwin_output_mgmt`), `cargo clippy -p pf-vdisplay --all-targets
--locked -D warnings` clean, `scripts/xcheck.sh linux` clean, fmt clean. NOT yet
on-glass — no KDE box here reproduces a stored mirror; the reporter's setup is
the real test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 11:55:34 +02:00
enricobuehlerandClaude Fable 5 caa47e28e6 fix(client/decode): AV1 hardware decode stops silently opening libdav1d
avcodec_find_decoder(id) returns the registry's FIRST decoder for the id, and
upstream orders the native av1 decoder LAST on purpose ("hwaccel hooks only,
so prefer external decoders" — allcodecs.c). All three hardware backends
selected by id, so every AV1 session opened libdav1d: a software decoder that
silently ignores hw_device_ctx and never calls get_format. Each frame then
failed the backend's hw-format guard and the session burned the demotion
ladder MID-STREAM — field-logged as 68 Vulkan fails → D3D11VA → 102 fails →
software, ~3 s of black — with "hardware decode active" already printed and
the D3D11 profile/pool probes all green. H.264/HEVC never hit this only
because their native decoders happen to be registered first.

Selection is now by capability: find_hw_decoder walks av_codec_iterate and
takes the first decoder whose avcodec_get_hw_config advertises the backend's
surface via HW_DEVICE_CTX, so a build without a usable hw decoder fails at
OPEN in milliseconds and the ladder runs there — the idiom the D3D11 probes
already follow. Registry order still wins among capable decoders, so
H.264/HEVC select exactly what they always did. The software path keeps the
id lookup on purpose: libdav1d is the fastest CPU AV1, and the native av1
decoder has no software path at all.

Every decode log now carries the selected decoder's name — decoder="av1" vs
decoder="libdav1d" is the whole diagnosis, and no log line said it. The
session log names the WIRE codec and drops the FFmpeg id for PyroWave
(ffmpeg_codec_id's fallthrough claimed codec_id=HEVC for wavelet sessions
that never touch FFmpeg).

The CPU lane also stops passing raw PQ off as a tone-map: software-decoded
frames deliberately never take the HDR10 swapchain, but a PQ stream there was
then shown UNtonemapped (washed out) with no warning — the pq-downgrade warn
keys off the swapchain answer — while the Detailed OSD badge claimed the
"HDR→SDR" tone-map that only the hardware lane's CSC runs. The presenter now
warns once when a PQ CpuFrame arrives, and the badge distinguishes
"HDR→SDR (raw)" (no tone-map pass) from the hardware lane's real "HDR→SDR".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 09:54:43 +02:00
enricobuehlerandClaude Fable 5 652abeb397 fix(host/audio): an unsatisfiable wiring plan waits for an endpoint change instead of hammering
Field case 2026-08: the display isolate invalidated the only real render
endpoint; the mic held the Steam Streaming Microphone, the Speakers were
blacklisted, and the capture loop re-ran the full wiring pass — three
IPolicyConfig SetDefaultEndpoint writes included — every 2 s for 8+
minutes, retrying a verdict that could never change.

- wiring_plan: a plan with no loopback is a typed structural verdict
  (Wiring::loopback_unsatisfiable + an endpoint-set fingerprint); the
  dead leftover() tier (byte-identical to real_hw()) becomes a real last
  resort that accepts ONLY the Steam Streaming Speakers, flagged
  loopback_last_resort — a known-silent-loopback QUALITY risk, never the
  cable/VoiceMeeter echo CORRECTNESS risks. excluded_from_loopback stays
  untouched (judge_default's mid-stream snap-back semantics).
- wasapi_cap: an unsatisfiable plan errors ONCE per topology with the
  render inventory, per-endpoint rejection reasons and only the remedies
  not already taken, then parks on a cheap enumerate-and-hash poll and
  re-plans the instant the set changes; transient failures get a real
  capped exponential backoff (2 s → 60 s, reset on success or set
  change); a last-resort capture re-plans on any set change and promotes
  the 30 s zero-packet breadcrumb to warn.
- audio_control: the recording default is asserted only when the plan
  changed or the default drifted — set_default_endpoint fires all three
  SetDefaultEndpoint roles unconditionally, so the 2 s loop silently
  stomped any operator recording-device change; the "attach one, or let
  the host install the Steam Streaming pair" warn (already satisfied in
  the field case) is replaced by the same diagnosis.

Verified: 19/19 wiring_plan tests (native rustc --test and the Linux CI
image via docker); both Windows files type-check and clippy clean
against wasapi 0.23.0 / windows 0.62.2 for x86_64-pc-windows-msvc via an
xcheck-style stub harness (the in-tree target check dies in
openh264-sys2's build script on macOS, as scripts/xcheck.sh documents).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 09:53:36 +02:00
enricobuehlerandClaude Fable 5 48511d1267 fix(abr): probe throughput is measured over the client's receive interval
The capacity probe divided client-side bytes by the HOST's burst
duration — a window wrong on both edges (base snapshotted before the
burst reached the host, frozen only when the ProbeResult landed, while
the host's clock stops the moment ITS send window closes, before the
switch/kernel queue finishes draining toward the client). On a 1 GbE
link a 2 Gbps burst target "measured" 1266 Mbps and set an 886 Mbps
climb ceiling the link could never carry — permanent for the session,
because set_ceiling never lowers.

The reassembler now stamps probe-scoped counters (bytes, packets,
first/last arrival, monotonic ns) at its FLAG_PROBE routing, so video
around the burst contaminates neither numerator nor denominator; the
throughput divisor is the client's first→last arrival interval, with
the host duration kept as the fallback when fewer than two probe
packets arrived. The user-facing speed test shares the corrected
computation (ProbeOutcome/PunktfunkProbeResult layouts unchanged;
elapsed_ms docs updated to the new semantics).

Two guards ride along:
- PUNKTFUNK_ABR_MAX_MBPS clamps inside set_ceiling — the one funnel
  every learned ceiling passes through — so a user cap binds no matter
  what any probe concludes.
- The controller latches decode_cap_kbps when two CONSECUTIVE backoffs
  carry decode-severe evidence (deep decode excursion or jump-to-live
  flush) at a similar pre-backoff rate, mirroring host_cap_kbps for the
  client decoder: a knee below the link ceiling was a permanent 30-60 s
  sawtooth costing a flush + dropped-frame burst per cycle (1440p120
  HEVC field case, knee ~490 Mbps). One spurious flush never latches;
  the cap re-probes on the CAP_REPROBE_WINDOWS clock, so it lifts when
  the decoder recovers.

Also rights the three stale "3 Gbps" probe-clamp comments (the host
constant has been 10 Gbps since MAX_PROBE_KBPS moved).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 09:53:36 +02:00
enricobuehlerandClaude Fable 5 3e649d372e fix(capture/stall): the ETW witness testifies in QPC, so compose-silence stops convicting content
The stall classifier's present witness never worked: the consumer was opened
without PROCESS_TRACE_MODE_RAW_TIMESTAMP, so ProcessTrace converted every
event's TimeStamp to FILETIME regardless of the session's ClientContext=1 —
FILETIME ticks (100 ns since 1601) are ~4 orders of magnitude above QPC, so
every ts <= to_q comparison was false. summary() always printed etw=none,
window_counts() always returned presents=0/queue_adds=0 while present_history
was still true (satisfied by the unfiltered ring), and classify() therefore
convicted EVERY compose-silence hole as CONTENT-SILENCE; FRAME-GENERATION —
the class the program exists to catch — was unreachable. Two comments
asserted the wrong contract ("TimeStamp IS a QPC value"); both now state the
real one: ClientContext selects the session clock, RAW_TIMESTAMP is what
stops the FILETIME conversion on delivery.

Three adjacent defects fixed with it:

- summary() and window_counts() each took their own ring lock and their own
  (Instant::now(), qpc_now()) anchor, with OpenProcess syscalls between the
  two calls — the prose and the verdict could disagree about the same hole.
  Merged into window_report(): one snapshot, one anchor, both halves; the
  summary keeps its 300 ms lead-in, the counts keep the gap-only window, and
  the etw=/etw_presents=/etw_queue_adds= log fields are unchanged.

- present_history/queue_history meant "an event EVER sat in the ring" —
  satisfied by events arriving after the hole, or by a previous session's
  leftovers in the never-cleared static RING. Both flags now mean witness
  LIVENESS: at least one event inside a 5 s LOOKBACK ending at the hole's
  start, i.e. the witness demonstrably worked before the hole opened. The
  ring is cleared when a new session starts, so a dead session's events can
  never pose as the next one's history.

- window_counts() accepted only BltQueueAddEntry (1071) as queue history
  while summary() also took BltQueueCompleteIndirectPresent (1068); either
  proves the queue witness works, so the merged reader takes both.

The windowing math is factored into a pure count_window() (plain i64 tick
arithmetic) with unit tests, and the classify() matrix gains the live-witness
zero-presents case. Conviction thresholds are untouched.

Verified: scripts/xcheck.sh windows clippy (-D warnings, --all-targets) green
for pf-frame/pf-win-display/pf-capture/pf-vdisplay; native cargo check green.
The new Windows-gated tests type-check but need a Windows box to run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 09:53:36 +02:00
enricobuehlerandClaude Fable 5 0d004c4680 fix(host/handshake): the 4:4:4 gate names the encoder backend, not the capturer
capture_supports_444 was an encoder-backend fact (direct NVENC or PyroWave)
logged under a capture-ish name — a field report burned real time hunting a
capture problem because of it. The 'encode chroma' line now says
ingest_chain_supports_444, a requested-but-declined session logs WHICH gate
lost, and the console UI's Full chroma explainer names the real requirement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 09:53:29 +02:00
enricobuehlerandClaude Fable 5 8d7e273a96 feat(client/present): PUNKTFUNK_PRESENT_MODE gains explicit mailbox and fifo_relaxed arms, and the docs stop guessing
The env knob silently folded 'mailbox' and every typo into the default arm,
FIFO_RELAXED was not reachable at all, and clients/session/README.md claimed
the default is FIFO (it is MAILBOX with a FIFO fallback). An AMD-on-Windows
client always lands on FIFO because that driver offers no MAILBOX — now
documented at the picker and in the docs-site client table, next to the ABR
probe/ceiling knobs a field report went looking for and couldn't find.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 09:53:28 +02:00
enricobuehlerandClaude Opus 5 e726542f96 docs(winget): the vhost belongs in unom/infra, not on the box
Step 2 told you to add the Caddy vhost by hand on unom-1. That instruction is
what broke the source: ~/caddy/Caddyfile is a copy that deploy-all.sh rsyncs
over from unom/infra, with no .git there to warn you, so the hand-added block
survived until the 2026-07-31 hardening commit rewrote the file from the repo's
own copy and deleted it.

Point step 2 at unom/infra and record how the failure presents, since it does
not look like an ingress problem from the client side: no vhost means no
certificate for that SNI, so Caddy answers with TLS internal_error (alert 80)
before sending one, and winget surfaces that as
WINHTTP_CALLBACK_STATUS_FLAG_SECURITY_CHANNEL_ERROR / 0x8a15003b.

Also note that port 80 is useless for diagnosing it — Caddy 308s every Host to
https including names it has never heard of — and give the SNI probe that does
work.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 10:43:39 +02:00
enricobuehlerandClaude Opus 5 e0427a3bb6 feat(ci/android): a release tag without Play notes fails before it builds
Play does not show an empty "What's new" when the file is missing — it carries
the PREVIOUS release's text onto the new version. So the store listing ends up
describing a build nobody is getting, and nothing surfaces it except reading
the listing. That is the shape of the v0.22.3 notes announcing a feature the
tag never contained, and a soft warning in a log nobody reads does not prevent
it.

The gate runs FIRST in the job, before the ten-minute build: a miss costs a
second and leaves nothing half-published — no build, no assets on the Gitea
release, nothing on Play. It rejects three things: a missing file, a file
byte-identical to another release's (the same bug reached by copy-paste rather
than omission), and an empty or over-500-char one.

Length is checked here as well as in play-upload.py on purpose. The uploader
stays the last line of defence and is the only check android-promote.yml gets,
but it runs at step 9; this catches an unedited TEMPLATE copy at step 1. It
counts CHARACTERS, not bytes — Play's cap is 500 chars and `•` is three bytes
in UTF-8, so a `wc -c` check would have called the 356-char v0.23.0 notes 365
and can reject a legal file.

whatsnew/TEMPLATE.txt gives the file a starting point and says what the gate
does and does not enforce: it cannot tell whether the prose was ever edited, so
a copy that still reads "<The headline change>" ships exactly as written.

Canary stays exempt — no curated notes, and Play reusing text for internal
testers costs nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 10:34:32 +02:00
enricobuehlerandClaude Fable 5 02a5bdb965 fix(gamepad): the virtual DualSense stops demanding a firmware update it cannot take
The emulated pad's firmware-info feature report (0x20) advertised update
version 0x0154 — a 2021-era number. PlayStation Accessories compares it
against Sony's latest (0x0630 as of 2026-08) and offers an Update that can
only end in "can't complete the update", since the virtual pad speaks no DFU;
libScePad titles (Stellar Blade) surface the same nag in-game. A real pad
plugged in directly reads up to date, which made the prompt look like
punktfunk corrupting the controller.

The old value was chosen to keep the kernel and SDL on the flag0
COMPATIBLE_VIBRATION convention, but parse_ds_output has since learned the
firmware-≥2.24 COMPATIBLE_VIBRATION2 flag as well, so nothing depends on
looking old anymore. Advertise 0x0999 — above anything Sony has shipped and
comfortably ahead of their ~yearly cadence — instead of chasing their exact
latest, which would resurrect the prompt on every Sony release. Writers that
read the version now use the v2 flag; both conventions land in the same
rumble plane. Bumped in both copies of the blob (host uhid + Windows driver);
the DualSense Edge shares them, and its own versioning (0x0217 latest) sits
below the new value too.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 10:24:20 +02:00
enricobuehlerandClaude Opus 5 09b9ee8f53 feat(ci/android): a release tag publishes to Play production, not alpha
Production access came through on 2026-08-01. Until now a `vX.Y.Z` tag could
only reach `alpha` and someone had to promote it by hand in the Console; it now
goes to `production` at 100% (`completed`). Canary is unchanged on `internal`,
and its run-number versionCodes always outrank production, so testers keep
getting the newer build.

A tag therefore reaches real users with no further click. What keeps that
honest: the tag is only pushed once every platform is green, and Play reviews
each production release before it ships. Ramping instead is `--status
inProgress --user-fraction 0.2` on the upload step.

Play's "What's new" gets its own file, docs/releases/whatsnew/vX.Y.Z.txt — the
vX.Y.Z.md body is ~34 KB against a 500-char cap, so it cannot be reused. Only
tags have one; canary is a moving target and Play carrying the previous text
over is fine for internal testers. Same freeze rule as the notes: once the tag
exists, the file describes what that versionCode shipped.

android-promote.yml is the lever for everything that is not a fresh tag —
promote a tested build, halt a rollout, or roll production back onto an older
versionCode. It is separate from android.yml because promotion must not
rebuild, and an `if:` on all ten build steps is worse than one small workflow.
dry_run defaults to true, so a mis-typed versionCode validates and deletes the
edit instead of publishing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 09:29:20 +02:00
enricobuehlerandClaude Opus 5 43e3c7b69f feat(ci/android): play-upload can attach release notes and promote a build
Two things it could not do, both needed now that a tag ships to production.

Release notes: it never sent `releaseNotes`, so Play's "What's new" was whatever
the previous release said. It now takes --release-notes-file, and refuses text
over Play's 500-char-per-language cap with the actual count — that check has to
happen before the upload, because the API only rejects it at commit, by which
point the AAB is already on Play.

Promotion: --promote assigns a versionCode that is already on Play instead of
uploading, so what reaches production is the byte-identical artifact the testers
ran. Rebuilding would mint a fresh versionCode from possibly-newer sources and
ship something nobody tested. --promote-from asserts the code really is on that
track (a typo'd versionCode now fails before it touches production) and clears
that track in the SAME edit, so the build is never active on both at once.

--user-fraction comes along because --status inProgress is an API error without
it; it is validated as strictly between 0 and 1 rather than left to Google.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 09:29:04 +02:00
enricobuehler ea469162f9 fix(docs): three doc comments start a markdown list they never meant to
Windows clippy on the v0.23.0 tag: `doc_lazy_continuation` in
crates/pf-client-core/src/audio_wasapi.rs:37. The cause is one line break —
`+ wire cost.` begins a line, so the markdown parser reads `+` as a bullet
marker and the following line becomes a lazy continuation of that list item.

Fixed by reflowing so the `+` is mid-line rather than by taking clippy's
suggested indent: indenting would keep the accidental bullet in the rendered
docs, which is the actual defect. Same treatment for the two siblings a sweep
of every `///` line found, both invisible to the Linux gate for their own
reasons:

- gamestream/audio.rs:237 — `+ libopus;` at line start, on the
  cfg(not(linux/windows)) stub, so only a macOS clippy would ever see it.
- mgmt/tests.rs:1653 — `404.` at line start IS an ordered-list marker
  (CommonMark: 1-9 digits + `.`), and it is behind cfg(test), so only an
  --all-targets run sees it.

This is the [[Windows clippy sees what the Linux gate structurally cannot]]
shape again: audio_wasapi.rs is cfg(windows), so no amount of Linux CI would
have caught it.

Verified: a scanner over every .rs doc comment in the tree now reports zero
line-initial list markers with an unindented continuation; rustfmt clean (it
does not reflow doc comments, so these edits are stable).
2026-08-01 01:30:45 +02:00
enricobuehler 213b353dad docs(release): the 0.23.0 notes link to the docs site that exists
Both links pointed at https://punktfunk.unom.io/docs/... — the marketing host,
which 404s (verified live; the docs site serves from docs.punktfunk.unom.io, as
README.md has used throughout). The Updating link was the one a reader following
the one-click update section would actually click.

The echo link additionally pointed at the docs ROOT rather than the page it
names; both now resolve to their real slugs, confirmed against
docs-site/content/docs/{echo,updating}.md and by fetching them (200/200, vs 404
for the old host).

Release body re-synced by body-only PATCH.
2026-08-01 01:25:01 +02:00
enricobuehler 23ec0822d8 docs(release): the 0.23.0 notes stop calling gamescope HDR a handheld feature
Two wording fixes to the shipped v0.23.0 notes, which docs/releases/README.md
explicitly allows after the tag — the file stays authoritative and the announce
step re-syncs from it.

"HDR turns itself on for Steam Deck and other gamescope handhelds" framed a
LINUX HOST change as a device one. What 19392918 actually did is default
PUNKTFUNK_GAMESCOPE_HDR on and get the patched gamescope onto the Linux install
routes (Bazzite + Arch sysext, the nix module, the Deck's on-device build
script) — which covers any host running its games through gamescope: a Bazzite
or SteamOS box in Game Mode, an HTPC, a desktop, not only a handheld. The
lead-in carried the same framing and mattered more, since everything above the
first `##` is what the Discord embed shows.

Also rewords the mic-mute lead-in ("stop the room being heard" read oddly).

No claim changes: same features, same scope, same release. The live release body
needs a re-sync — done via a body-only PATCH rather than an announce dispatch,
since announcing also posts to Discord and publishes the stable update manifest,
neither of which should fire before the fleet is green.
2026-08-01 01:21:54 +02:00
enricobuehler 49bbdcf4ef chore(release): bump workspace version to 0.23.0
A minor bump: 133 commits since v0.22.3 across 400-odd files. The wire grew two
negotiated abilities (slice-streamed access units, phase-locked capture); the
Android presenter was rebuilt; the microphone path was rebuilt end to end on
every client; the web console moved from polling to the host's event stream;
gamescope HDR is on by default; HDR and 4:4:4 stopped being mutually exclusive
on Windows; and the one-click update apply that missed the 0.22.3 cut ships here.
The canary base is already 0.23 — release.yml derives it as one minor ahead of
the latest stable tag — so this is the version the canary channel has been
publishing against all along.

Lock touched for the 32 workspace members only, via `cargo update --workspace`
rather than a sed: `wasapi` is itself at 0.23.0 and `rustls` at 0.23.41, so the
version space we are moving into is occupied by third-party crates this time.
Diff against origin/main is versions-only, 32 insertions and 32 deletions;
`cargo metadata --locked` resolves; `cargo fmt --all --check` clean in both the
main and the packaging/windows/drivers workspaces.

Notes at docs/releases/v0.23.0.md, per docs/releases/README.md — authored with
the bump so CI's ensure_release seeds the body at tag creation.
2026-08-01 01:01:16 +02:00
enricobuehler 3c509d48c9 feat(core/abi): report_phase earns its version — C ABI 13 -> 14
`punktfunk_connection_report_phase` (fa822744, coherence tail 1d31e4c5) and the
`PUNKTFUNK_CLIENT_CAP_PHASE_LOCK` mirror const (7cf71dd2) grew the embeddable C
surface without moving ABI_VERSION. Every prior additive entry in that doc list
bumped it — v3's wake_on_lan, v5's next_rumble2, v8's clipboard block, v13's
send_pen — precisely so an embedder can ask punktfunk_abi_version() whether the
function it wants to link is there. Left at 13, the one number that answers that
question said "no report_phase" about a core that has one.

The header was already regenerated with both symbols, so it carried the new
surface under the old number; the regen here changes exactly the #define and its
doc block and nothing else, which also confirms the committed header was
otherwise current.

Additive and capability-gated: the host arms on report receipt, the wire grows
only PhaseReport (0x32) — a control message an old host never reads — and a
strict-prefix append on the 0xCF host-timing tail, so WIRE_VERSION stays 2. No
in-tree caller compares ABI_VERSION against a literal; mgmt/tests.rs asserts
against the symbol.

Verified: cargo build -p punktfunk-core regenerates include/punktfunk_core.h to
exactly this diff; punktfunk-core lib suite 134/134; rustfmt clean. The C ABI
harness cannot run on this Mac (`ld: library 'opus' not found`, the documented
pre-existing local linker gap) — CI's Linux leg is the gate for it.
2026-08-01 00:58:11 +02:00
enricobuehler b8d987b145 docs(release): the v0.22.3 notes stop claiming one-click updating it never shipped
The v0.22.3 tag is `1c836afc`, cut 14:36. The one-click apply work landed on
main between 15:01 and 16:28, and this file was then edited at 17:23 (5790a3e3)
to announce it — three "New"/"Under the hood" claims about a build that does not
contain them. The live release body is still the pre-edit text, so nothing wrong
has been published yet; but `announce.yml` re-asserts this file over the release
on every announce, and 0.22.3 has not been announced. Announcing it would have
published the false version and put its lead-in ("can install it for you where
the platform allows") into the Discord embed.

Verified against the tag rather than the commit graph: `update.available` is in
`1c836afc`, `update.applied` is not, and `mgmt/tests.rs` there literally probes
`/api/v1/update/apply-does-not-exist-yet` while `auth.rs` carries the comment
"today it is only a check". The Updates card itself (b275e6d3, cc015626) IS in
the tag, so that bullet stays; only the apply half goes.

The removed material is not lost — it is in docs/releases/v0.23.0.md, which is
the release that actually ships it.
2026-08-01 00:58:11 +02:00
enricobuehler 2d3f9f8690 Merge remote-tracking branch 'origin/main' into audio/mic-latency-echo 2026-08-01 00:54:20 +02:00
enricobuehler 951bcec650 Merge remote-tracking branch 'origin/main' into audio/mic-latency-echo 2026-08-01 00:47:50 +02:00
enricobuehlerandClaude Fable 5 badda070ef docs(android): the stats-array KDoc counts the doubles it actually returns
nativeVideoStats grew to 33 with the decode split and the overflow counter, but
its own KDoc still promised 30 and StatsOverlay still said 26 — a count that
was already two extensions stale before this one. Both now list the full index
set, with the JNI KDoc named as the authoritative one so the next extension has
a single place to update.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:47:04 +02:00
enricobuehlerandClaude Fable 5 46bcfc3041 fix(scripts/windows): the installer-run scripts go back to pure ASCII
CI's guard fired on build-web.ps1: an em-dash in the header comment. The
rule exists because PowerShell 5.1 mis-parses non-UTF-8-locale files, and
the check covers every script the installer can run, comments included.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:41:07 +02:00
enricobuehlerandClaude Fable 5 f3a39df7b3 test(host/audio): prove a reopen recovered with a live uplink, not one frame
`reopens_after_push_death` failed about one run in nine, and widening its
timeout did not help — the earlier commit blamed the backoff and was wrong.

The pump drops whatever queued while it was down: audio from before the
device came back is stale, so a fresh instance drains the channel right
after opening. The harness counts `opens` from the START of the open, so
the moment the test sees the counter move, the pump has not reached that
drain yet. The single frame it then sent landed inside the drain window and
was discarded exactly as designed, leaving the test waiting for audio that
was never going to arrive.

So the test now keeps feeding, which is what a real uplink does and what
the drain assumes. The sequence advances each time or the de-jitter reads
the repeats as duplicates and drops them for a second, correct reason.

Production behaviour is unchanged: this was the test asserting something
the pump never promised.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:41:07 +02:00
enricobuehlerandClaude Fable 5 f9c56eaf5c feat(android): Automatic prefers AV1 where the silicon says it should
The P3 format A/B (NP3 ↔ RTX 4090, identical conditions) measured AV1 ~1.2 ms
faster end-to-end than HEVC with slightly better codec-pure decode time. Under
"Automatic" the client now sends AV1 as its soft preference when this device
hardware-decodes it (the advertised AV1 bit is already gated on a real,
non-blocked hardware decoder) AND it lacks FEATURE_PartialFrame — a
partial-frame device keeps HEVC, whose slice-progressive overlap AV1 cannot
ride (no slices, the chunked poll never arms). The host honors the preference
only inside its probed shared codec set, so an AV1-less encoder still resolves
HEVC, and an explicit user choice wins unchanged. The codec picker caption
mirrors the same rule so "Automatic" says what it does on this device.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:40:58 +02:00
enricobuehlerandClaude Fable 5 b69ef02f4d fix(android): a cold-start connect no longer loses HDR or the native mode
A punktfunk:// deep link can reach the connect before the activity is attached
to its display; context.display then throws and the display probes silently
fell to their worst answers — displaySupportsHdr advertised SDR (the whole
session pinned to 8-bit BT.709) and nativeDisplayMode fell back to 1080p60.
Seen live on the NP3: one cold connect advertised hdr=false, the warm retry
true, nothing in the log either way.

Both probes now share probeDisplay: the context display when attached, else
DisplayManager DEFAULT_DISPLAY — which IS the panel on phones and TVs; the
activity-display distinction only matters on multi-display setups, where the
attached path still wins whenever available. Each fallback leg logs itself, so
a downgraded session can never again be silent about why.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:40:58 +02:00
enricobuehlerandClaude Fable 5 4d45a96ff9 feat(encode/windows): sub-frame readback defaults on where the GPU supports it
Linux parity, validated by the .173 on-glass A/B (no regression; the win goes
to clients that actually consume slice-progressive parts): the caps probe now
reads NV_ENC_CAPS_SUPPORT_SUBFRAME_READBACK and seeds resolve_subframe with it
instead of a hard false, so PUNKTFUNK_NVENC_SUBFRAME becomes the tri-state
escape it already is on Linux, and the split×sub-frame arbitration hears the
real forced flag for its log severity.

The A/B also caught the default path opening every session with a WARN: the
submit-time idr_hint missed that NVENC emits the session-opening frame as an
IDR regardless of pic flags, so frame 1's early chunks went out unflagged and
the divergence check fired at every start. The hint now carries the Linux
twin's `opening` term.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:40:58 +02:00
enricobuehlerandClaude Fable 5 849baea881 feat(android): the stream re-votes its refresh rate and touches keep their curvature
surfaceChanged re-asserts the frame-rate vote (FIXED_SOURCE; ALWAYS only on the
TV low-latency path, mirroring the native hint) — a buffer-geometry change on
some OEM builds silently drops the 120 Hz pin mid-stream. Touch passthrough and
direct-pointer moves forward the MotionEvent historical samples before the
current point, so a fast swipe lands with its real shape; the trackpad path
keeps summed deltas on purpose — its acceleration curve is tuned for per-frame
dt and historicals would change the feel, not the sum.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:40:58 +02:00
enricobuehlerandClaude Fable 5 c767a904d2 feat(android): the decode stage answers where its time goes, HUD on or off
P3 decode science: every AU is stamped as its last piece enters the codec, so
the decode stage splits into feed (received→queued: hand-off + input-slot wait)
and codec (queued→decoded: the decoder alone — a slice head start would show
here). The split + an always-on capture→decoded e2e ride the 1 Hz pf-present
line, so a wireless HUD-off A/B reads everything from logcat; the HUD equation
gains the split (indices 30/31), the skipped counter tells benign newest-wins
pacing from parked-AU overflow (32), and a −2-refresh Apple-HUD-equivalent twin
makes iPhone comparisons honest (Apple shaves its OS floor; Android shows raw).

Connect now logs the per-mime decoder picks + FEATURE_PartialFrame verdicts
(tag pf.caps) — on the NP3 all three c2.qti low-latency decoders say no, so
parts delivery never arms and P2d is inert there; a debug.punktfunk.force_parts
sysprop overrides the probe for the on-glass question the API cannot answer.
Forced on glass: c2.qti accepts PARTIAL_FRAME pieces without erroring but only
assembles them — codec time unchanged, so the overlap is dead on SM8735 either
way.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-01 00:40:58 +02:00
enricobuehlerandClaude Opus 5 0985726415 fix(host,web): an empty update channel stops looking like a broken host
A channel nobody has published to answers `manifest.json` with a 404, and the
check reported that the same way it reports a dead registry or a bad signature:
"Last check failed: feed returned HTTP 404". Every host on the stable channel
shows it today, because the stable manifest only publishes when someone
dispatches `announce` for a release tag — so the first thing an operator sees
from the new Updates card is a red failure caused by nothing being wrong.

The shared checker now distinguishes the two. `feed::fetch_manifest_blocking`
returns a typed `FeedError` instead of a string, and only a 404 on the manifest
ITSELF becomes `NotPublished` — a 404 on the detached signature still fails
loudly, because that is the half-published pair the manifest-then-signature
upload order can produce, and it must stay fail-closed.

The host carries that through as `UpdateStatus.not_published`, mutually
exclusive with `last_error`. It is benign only while no manifest has ever been
seen for the channel: once a check has succeeded, the same 404 means the feed
LOST a document it used to serve, which stays an error. The console then shows a
plain sentence naming the channel instead of the failure banner, and "None
published yet" rather than "Not checked yet".

The Linux client makes the same distinction but deliberately NOT the same
choice: `--check-update` keeps exiting 1 and keeps `error` set, because its
consumer is a shell script and an empty channel is the absence of evidence that
this build is current — not a confirmation that it is.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:37:56 +02:00
enricobuehler 23f1debe69 Merge remote-tracking branch 'origin/main' into audio/mic-latency-echo 2026-08-01 00:21:24 +02:00
enricobuehlerandClaude Opus 5 ff190e9825 fix(web): the logs filter row stops touching the card's edge on desktop
Found on glass. The logs card has no CardHeader, so it puts the top padding back
itself — but only at one breakpoint. `CardContent` is `p-4 pt-0 sm:p-6 sm:pt-0`,
and tailwind-merge resolves conflicts only within the same variant: a bare
`pt-6` cancels `pt-0` and leaves `sm:pt-0` standing. Measured on .173: 24px of
top padding at 420px wide, 0px at 1280px, with the level filters and the search
box sitting flush against the card border.

It is the same trap `components/ui/card.tsx` documents for `p-0` — a variant
cancelling its unprefixed counterpart and nothing else — just in the other
direction, so the note now points both ways. This card was the only offender.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:20:10 +02:00
enricobuehlerandClaude Opus 5 f66de3eba4 fix(web): close the last two high-severity items from the closeout audit
**The streamed-screen pin could still be clobbered by three other write paths.**
Deferring it to the server's value on Save fixed the reported sequence but not
the general case: the draft is only re-seeded while it is CLEAN, so once there is
an unsaved edit its `capture_monitor` is frozen at whatever it was before the
operator used the picker — and `applyAxis` (which spreads the last saved policy),
the built-in preset switch and the custom-preset apply all put that stale value
back. Every write path reads `serverCaptureMonitor()` now; no path spreads the
draft's copy.

**The session⇄game grace input had no accessible name.** That card has its own
`Field` and only DisplayCard's was fixed, so the number input was still announced
as an unnamed spin button. Same treatment: `htmlFor`/`id` for the single control,
`fieldset`/`legend` for the two button groups. Verified in a browser — zero
inputs without an accessible name across the Displays page.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:20:10 +02:00
enricobuehlerandClaude Opus 5 f2e1b9872c fix(web): four "fixes" from this branch that did not actually fix anything
A verification pass re-read every finding from the original sweep against the
code on this branch rather than against the commit messages. It found that four
of them were still broken, two because the edit I made was inert. Commit
messages claim; code decides.

- **The Storybook typecheck was never on.** `tsconfig.json` listed `.storybook`
  as a bare directory name, and tsc silently skips dot-prefixed directories in
  that form — so the entry typechecked nothing at all. Proved it by planting
  `export const __probe: string = 1` in `.storybook/preview.tsx` and watching
  `bun run lint` pass. `.storybook/**/*` is what actually pulls it in; the same
  probe now fails as it should.

- **The Moonlight stale-PIN reset was a no-op.** `submit.reset()` sat at the top
  of `onSubmit`, immediately before `submit.mutate(...)` — which moves the status
  to pending in the same update, so it cleared a flag that was already changing.
  The green "PIN sent" note therefore still greeted the next pairing attempt over
  an empty PIN box. It now resets on the transition that actually matters:
  `pin_pending` going false → true.

- **The session⇄game controls had the enforcement flag inverted**, and I never
  touched it. `enforced.length === 0 || …` reads an EMPTY list as "this build
  enforces everything", when the contract says the opposite in as many words:
  "Empty on a platform with no launch path (macOS), so the console can say so
  instead of offering a switch that does nothing". On exactly the platform the
  flag exists for, every control stayed live and reported success for an axis the
  host would never act on. Absent still means "assume it acts" — that is the
  compatible reading for an older host, and a different case from present-empty.

- **Logout stopped revoking after a restart.** The epoch was a module-level
  counter starting at 1, so it revoked within one process run and then reset —
  and since the seal key derives from the stable mgmt token, a cookie captured
  before a restart unsealed fine and was accepted again for the rest of its
  7-day TTL. One service restart undid the whole fix. It persists next to the
  host's config now. Verified: log out, restart the console, the captured cookie
  still 401s, a fresh login still works.

Two more the pass rated as partial, both worth closing:

- The plugin-UI response filter was a denylist of four header names, so
  `Clear-Site-Data` sailed through — a plugin error page could wipe `pf_session`
  and sign the operator out of the console, on our own origin, because the iframe
  is same-origin by design. It is an allowlist now; a plugin-supplied CSP,
  `X-Frame-Options` or CORS header no longer speaks for us either.

- A half-configured TLS setup now refuses to start instead of logging a warning
  and serving anyway. Neither shape can work — one path missing puts the login
  password on the LAN in the clear, and PUNKTFUNK_UI_SECURE without TLS marks the
  cookie Secure so the browser drops it and login can never stick. Exiting with a
  reason beats a console that looks fine and is not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 00:20:10 +02:00