df74dd5aee96ef6f96008c9319fa95cf88b9d659
1556
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
df74dd5aee |
fix(windows/cursor): a device restart is NOT free — gate it again, and stop swallowing pnputil's reason
Two corrections, both from the box. 1. The previous commit made the between-session clean unconditional on the grounds that /restart-device is idempotent and costs 0.07 s. It is not free: after ~6 restarts in one afternoon .173 put the devnode into RESTART PENDING, and pnputil then refuses every further attempt -- 'a system restart is pending for this device to complete a previous operation' -- until an actual reboot. Restarting speculatively on every capture-mode connect would burn the only lever we have. So gate it on an outstanding declare again. 2. That failure was invisible. The script discarded pnputil's output and reported a bare exit code, so three separate runs looked like a wiring bug (the clean 'not firing') when in fact it ran every time and the restart failed. Keep the text, and surface a restart_pending field so the one failure that no retry can fix is named in the log. Also corrects the previous commit's message: CURSOR_DECLARED was never the problem. Both host processes run as SYSTEM and serve/handshake/capture share one process, so the flag was set and read correctly all along. |
||
|
|
b2020396c9 |
fix(windows/cursor): stop gating the between-session clean on host-process state
The reconnect case still failed after the keep-alive guard fix, and the logs said why by omission: neither outcome line appeared, so clean_cursor_for_next_session was returning at its first early-out — CURSOR_DECLARED was false even though the previous session had logged 'driver declares the hardware cursor'. The flag is written in capture_virtual_output and read in the handshake; one of those is not on the path that actually runs. Rather than chase that, delete the dependency on it. The operation being guarded is idempotent and costs 0.07 s: restarting an already-clean adapter wastes 70 ms once per capture-mode connect, while failing to restart a dirty one costs that session a full-frame copy for every frame with a visible pointer, for its entire life. The flag stays only as a log field, so the next run still shows whether the hint was set. The real guard remains the one that matters: no session is streaming. |
||
|
|
b20184462d |
fix(windows/cursor): gate the between-session clean on STREAMING sessions, not keep-alive
Measured on .173: the guard never fired in the one case it exists for. After a
desktop-mode session disconnects its monitor LINGERS, so `no_live_displays()`
— which counted keep-alive slots as held — refused the restart on exactly the
reconnect that needed it, and the capture session went back to
`cursor_excluded=true` and forced compositing.
Only a streaming session (SlotState::Active) can be damaged by the restart. A
lingering/pinned monitor has no session attached and a reconnect preempts and
recreates it anyway ("a reused IddCx swap-chain is dead"), so the restart
destroys nothing that was going to survive.
|
||
|
|
537a1852ed |
feat(windows/cursor): give the pointer back after a desktop session — clear the declare between sessions
The start-up clean (
|
||
|
|
5819cf054b |
feat(windows/cursor): clear a sticky hardware-cursor declare at start-up — capture sessions get the OS's own pointer back
The goal this serves: a session that never engages desktop mode should get the LOSSLESS cursor — composited by Windows itself, for free, with true XOR — and the host should only pay for compositing when the user actually asks for the desktop mouse model. That is already what happens on a clean adapter. The problem is that adapters do not stay clean. A hardware-cursor declare is irrevocable and ADAPTER-WIDE (pf-driver-proto v6): once any desktop-mode session declares, DWM stops compositing the pointer into every later frame on that adapter, so every capture-latched session afterwards has to self-composite — a full-frame copy per visible-pointer frame, and our straight-alpha approximation of an XOR cursor instead of the real thing — for the rest of the adapter's life. "Until the next reboot" turned out to be far longer than it sounds. With Fast Startup on (the Windows default) a shutdown plus power-on is a HIBERBOOT: it restores session 0 and its drivers, so the declare survives what the operator calls a reboot. Measured on .173 — Kernel-Boot event id 27 reporting `0x1` where a cold boot reports `0x0`, with lsass/services/wininit keeping their pre-"reboot" start times, while `LastBootUpTime` reports the older cold boot and makes uptime checks lie. On such a box the lossless path can be gone for weeks. So clear it explicitly at host start, where no session holds a display yet. `pnputil /restart-device` recycles the WUDFHost process the driver's `DECLARED_TARGETS` lives in, which is all it takes. Measured at **0.07 s** against ~6 s of sleeps for the existing Disable+Enable cycle, and unlike that cycle it is designed for a device in use, so it does not hit the refusal `reload_vdisplay_adapter` documents as "the expected case here". In the same call it also repaired an adapter found in CM_PROB_FAILED_POST_START (Code 43). Best-effort throughout: a failure just leaves the adapter as it was and sessions self-composite exactly as before. `PUNKTFUNK_CURSOR_CLEAN_START=0` opts out. Verified: `scripts/xcheck.sh windows` green; `cargo fmt --all` clean; full `cargo build -p punktfunk-host --release --features nvenc` on .173 (xcheck cannot reach punktfunk-host — it needs ffmpeg). |
||
|
|
7f6d1622ee |
test(pf-capture): pin the composite-cursor regen key and the blend retry escalation
The two pieces of logic the previous commit added had no tests, and both are the kind that fail silently in opposite directions. `blend_key_of` is now ONE definition shared by the regen test and the blend itself, rather than the same tuple built at two call sites. That drift is the actual bug shape: a key that reports "changed" while the drawn frame is identical re-encodes for nothing, and a key that reports "unchanged" while the pointer moved freezes it on screen. Both directions are asserted — a hidden pointer keys identically wherever it moves, and every visible change (position, shape serial, and the visible→hidden transition that must strip the pointer from the frame) moves the key. `next_blend_backoff` is extracted for the same reason `mono_planes_to_rgba` was: the arithmetic a bug hides in does not need a live D3D11 device around it to be checked. The test walks the escalation well past its ceiling and asserts it PARKS there — an unbounded doubling would mean a device that comes back after a long stall never gets picked up. Verified: `scripts/xcheck.sh windows` green; `cargo fmt --all` clean. The tests themselves compile and run only on Windows (the module is `cfg(windows)`), so they are pending a run on a box. |
||
|
|
5c1db4662f |
fix(windows/cursor): the composite model paid a full-frame copy to draw nothing, and its failures were terminal
Three defects in the capture-model cursor path, all of them in the state a session spends most of its life in. 1. The blend copied the whole frame even when nothing would be drawn. `prepare_blend_scratch` ran the scratch build + `CopyResource` unconditionally whenever compositing was on, and only THEN skipped the quad for a hidden pointer. A 4K FP16 ring slot is 66 MB, so at 120 fps that is ~8 GB/s of write bandwidth bought for nothing — and a game that grabbed the pointer hides it, which is exactly when the composite model is engaged. The overlay is now resolved FIRST and a hidden or unknown pointer returns `None` before any allocation or copy; the conversion reads the slot directly, which is the same frame it would have got. 2. A hidden pointer moving forced frame regeneration on an idle desktop. The regen key was `(serial, x, y, visible)` — raw cursor state — so a pointer a game had hidden re-encoded the last slot every time it moved, despite the frame being pixel-identical. The key is now what the blend would DRAW: `Some((serial, x, y))` visible, `None` hidden. The visible⇄hidden transitions still change it, so the frame that must gain or lose the pointer is still regenerated. 3. A blend failure was permanent. `cursor_blend_failed` was set once, warned once, and the session then streamed a pointer-less desktop for the rest of its life — including for a device-loss that heals a frame later. It is now a backoff (250 ms doubling to 4 s) that suppresses the blend, drops the pass so it rebuilds, and clears on the first success. Every escalation logs, so a permanently broken session is distinguishable from one transient hiccup at startup, and the recovery says so. Also: the GDI poller now carries a heartbeat. `alive()` only asks whether the thread exited, so a poller wedged on an input desktop it can no longer read (`GetCursorInfo` failing every tick `continue`s before the publish) froze the pointer in every frame at its last sampled shape while looking perfectly healthy in the log. The capturer samples the publish count and warns once per stall — this poller is the ONLY full-fidelity shape source, since the driver's IddCx query is alpha-only, so its silence is the difference between a correct pointer and a frozen one. None of this depends on the open iPad diagnosis (planning-repo `windows-cursor-model-determinism.md` §2): these are defects on their own terms, and they are the prerequisites that make the composite path cheap and observable enough to reason about. Verified: `scripts/xcheck.sh windows` (clippy -D warnings for pf-frame, pf-win-display, pf-capture, pf-vdisplay against x86_64-pc-windows-msvc) green; `cargo fmt --all` clean. NOT run on a Windows box. |
||
|
|
4b514cc07c |
Merge pull request 'An OLED palette, and split WHETHER the gamepad UI is offered from WHEN it appears' (#116) from worktree-oled-theme-gamepad-ui-split into main
apple / swift (push) Successful in 1m33s
ci / rust-arm64 (push) Successful in 2m51s
ci / web (push) Successful in 3m13s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m45s
ci / bun-nix (push) Successful in 45s
ci / docs-site (push) Successful in 1m30s
ci / rust (push) Successful in 4m55s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 23s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 40s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 1m3s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m51s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 1m2s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 40s
deb / build-publish-client-arm64 (push) Successful in 2m42s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m17s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 47s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m18s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m17s
deb / build-publish-host (push) Successful in 6m16s
docker / builders-arm64cross (push) Successful in 11s
release / apple (push) Successful in 9m52s
docker / deploy-docs (push) Successful in 36s
android / android (push) Successful in 13m40s
deb / build-publish (push) Successful in 9m17s
arch / build-publish (push) Successful in 14m36s
flatpak / build-publish (push) Successful in 7m20s
apple / screenshots (push) Successful in 6m0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m15s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 21m21s
Reviewed-on: #116 |
||
|
|
30bd10e301 |
feat(clients): an OLED palette, and split WHETHER the gamepad UI is offered from WHEN it appears
apple / swift (pull_request) Successful in 1m33s
apple / screenshots (pull_request) Skipped
ci / web (pull_request) Successful in 1m39s
ci / rust-arm64 (pull_request) Successful in 4m7s
android / android (pull_request) Successful in 5m1s
ci / docs-site (pull_request) Successful in 1m47s
ci / bun-nix (pull_request) Successful in 42s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m19s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m23s
ci / rust (pull_request) Successful in 14m32s
Four changes to the client interface, kept together because two of them touch the same rows
and the last is a bug the first would have made far more visible.
A thirteenth `ui_palette` entry, `oled`. The palette table is hand-mirrored in three languages
(`pf-console-ui`'s `library.rs`, `GamepadPalette.swift`, `GamepadPalette.kt`), so it goes into
all three at index 1, directly after the brand default — which keeps `PALETTES[0]` the unknown-id
fallback and keeps the dark-to-pale cycling order intact. What earns the name is arithmetic, not
a darker shade of violet: the ramp's first two stops are literally (0,0,0) and the ground is pure
black, so the shaded half of the field is pixels switched off rather than "very dark grey", and
the calm mix the form screens sit under lifts toward nothing at all. Mean cell luminance is 0.019
against Violet's 0.254. The bright corner keeps a faint indigo-to-violet ember so the backdrop is
still a field with somewhere to go, and that ember carries enough chroma at that luminance
(60 degrees of hue travel across 13 of the 16 cells) to satisfy the existing multi-tone assertion
without adding `oled` to the near-neutral exemption Graphite and Opal take. Each port gains an
`oled_is_actually_black` test that measures the claim — pure-black corner cells, a mean under half
the darkest other field's — rather than restating the table.
A new device key, `gamepad_ui_mode`. The gamepad-UI switch had been deciding two things at once:
whether to offer the controller-optimized interface at all, and that it appears only while a pad
is attached. A user asked for the second half to stop applying. `"connected"` (the default, and
exactly what the lone Bool meant) and `"always"` separate them, surfaced as a "Show it" row
directly under the switch on all five settings surfaces and built only while that switch is on —
a picker whose every option decides nothing is worse than no picker. `GamepadUIEnvironment.isActive`
takes the mode with NO default argument on purpose: a call site that forgot it would silently
strand everyone who chose Always back on "only with a controller", which is the one bug this
parameter exists to make impossible. An unrecognized value waits for a controller, so a mode a
newer client wrote can never trap an older one in a layout it has no way back out of. It stays a
device preference on both platforms, never part of a profile: which interface this device wears
has nothing to do with how a host streams to it.
The smoothness buffer is hidden under Lowest latency, not dimmed. Everywhere else already hid it
— the GTK and WinUI shells, the Apple touch and tvOS screens, the Android touch screen — because
under that intent it names a quantity that does not exist. Two surfaces disagreed: Apple's gamepad
settings screen left the row live and steppable, and the desktop console dimmed it, having no way
to drop a row from a fixed list. That list is now rebuilt each frame through a `row_applies`
filter. The concern about a vanishing row moving everything under the cursor does not apply here
and the new test says why: the row it drops sits directly BELOW the row that drops it, so the only
cursor that can be present when the list shrinks is the one on the intent row, which does not
move. Two latent hazards went with it — `apply_row` had been indexing the row list on the
assumption the cursor is always in range, and nothing re-clamped that cursor when another writer
changed the intent behind the screen's back.
Pale palettes were unreadable on tvOS, reported from the field. `GamepadInk` was never the
problem: it flips correctly for a pale field, it is not platform-gated, and every tvOS gamepad
entry point already published it. The cause is that this app sets `preferredColorScheme` nowhere
and declares no `UIUserInterfaceStyle`, so every SYSTEM-derived colour landing on those screens —
a `.secondary` placeholder, a `.bordered` button's chrome, a NavigationStack title, a material's
frost — resolved against the DEVICE appearance, which the palette cannot reach. On iPhone, iPad
and Mac a great many users sit in Light mode, so under a pale palette those colours came out dark
and the theme looked correct by accident; an Apple TV is Dark essentially always, so every one of
them rendered white on a light field. The mirror image was broken too and had simply never been
reported: a dark palette on a Light-mode iPhone was already drawing dark on dark. The scheme is
now published beside the ink, once, in `GamepadInkModifier`, because the two are halves of one
decision and publishing only the ink silently loses every colour the frameworks draw on the app's
behalf. Two structural amplifiers went with it: `ConsoleGlass` had been scoping the scheme to the
fill inside its `.background {}` on the tvOS and pre-26 branches while the 26 branch put it on the
content, so no console row's own content ever saw it on tvOS; and `LibraryView`'s navigation
chrome and its loading, error and empty states sit above `LibraryCoverflowView` and so were never
inked at all on tvOS and macOS, where that view is presented directly rather than through the
iOS-only `GamepadLibraryScreen` wrapper.
That last one exposed a second tvOS gap worth closing in the same breath: `ui_palette` had no row
in tvOS's ordinary Settings, and the gamepad settings screen that owns it everywhere else needs an
extended-profile controller to open on tvOS. An Apple TV driven by the Siri Remote alone could not
reach the palettes at all, which would now include the OLED one. `SettingsView.tvBody` carries a
Background row.
Verified: pf-console-ui builds, passes `clippy --all-targets -D warnings` and runs 74 tests clean
under linux/amd64 (a Mac `cargo check` of that crate is vacuous — every module is cfg'd to
linux/windows); `cargo fmt --check` clean for it and pf-client-core. Android `:app` runs 80 tests
with 0 failures, including four new `gamepadUiActive` cases and the palette parity table. The
Apple package builds for macOS AND tvOS and its 9 palette/gamepad-UI tests pass — the tvOS
typecheck is possible because the checked-in xcframework already carries a `tvos-arm64` slice. The
tvOS RENDERING fix is compile-verified only; an on-glass Apple TV check under a pale palette is
still owed, and is the one thing here that a build cannot answer.
|
||
|
|
e4f8c64b9f |
Merge pull request 'Library scanners sat in the nav and could not sync local art — and you can now hide one game' (#113) from worktree-plugin-nav-category-and-art into main
audit / bun-audit (plugin-kit) (push) Successful in 19s
apple / swift (push) Successful in 1m38s
audit / bun-audit (sdk) (push) Successful in 48s
audit / pnpm-audit (push) Successful in 11s
audit / docs-site-audit (push) Successful in 1m8s
audit / bun-audit (web) (push) Failing after 1m14s
apple / screenshots (push) Successful in 5m46s
ci / rust-arm64 (push) Successful in 4m32s
audit / license-gate (push) Successful in 5m12s
ci / bun-nix (push) Successful in 38s
arch / build-publish (push) Successful in 8m1s
ci / docs-site (push) Successful in 1m12s
ci / web (push) Successful in 1m28s
android / android (push) Successful in 9m3s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 33s
deb / build-publish-client-arm64 (push) Successful in 1m25s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 27s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 28s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
audit / cargo-audit (push) Failing after 10m5s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 27s
ci / rust (push) Successful in 7m55s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m31s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m22s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 8s
sdk-publish / publish (push) Failing after 31s
docker / builders-arm64cross (push) Successful in 11s
deb / build-publish-host (push) Successful in 4m20s
docker / deploy-docs (push) Successful in 35s
windows-host / package (push) Successful in 16m9s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 31s
deb / build-publish (push) Successful in 12m39s
nix / flake (push) Canceled after 14m7s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 14m17s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 13m18s
Reviewed-on: #113 |
||
|
|
6cffe29b13 |
feat(host,console): hide individual library titles
ci / web (pull_request) Successful in 1m13s
ci / docs-site (pull_request) Successful in 1m23s
apple / swift (pull_request) Successful in 1m40s
ci / bun-nix (pull_request) Successful in 21s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m28s
android / android (pull_request) Successful in 4m25s
ci / rust (pull_request) Successful in 6m26s
nix / flake (pull_request) Successful in 15m40s
The library had one visibility control and it was all-or-nothing: turn a SOURCE off
and every one of its games goes. There was no way to drop a single title — a Proton
tool the filter missed, a demo, a game someone doesn't want on the TV — short of
hiding the whole launcher it came from.
**Where the setting lives.** Not on the entry. Only manual custom entries are stored;
a scanner's and a plugin's titles are rebuilt from scratch on every scan and every
reconcile, so a flag written onto one would be erased by the next sync — silently, and
minutes later, which is the worst possible shape for a setting. So `library-hidden.json`
holds the ids, mirroring how `library-scanners.json` holds disabled sources. The id is
stable by construction (D2: a claimed store's entries keep `<store>:<external_id>`
across reconciles), so a hide survives a re-scan, a plugin restart, and a store's
built-in→plugin migration.
**Where it takes effect.** In `all_games`, which is the one place every play surface
already funnels through — the grid on a client, native clients, the GameStream app
list, and launch resolution. Putting it there rather than at each call site is
deliberate: a per-surface filter is a rule someone has to remember, and forgetting one
is precisely the class of bug the `file://` art asymmetry in the previous commit was.
Hiding is curation, not access control — nothing is deleted, and un-hiding is instant.
**The console is the one surface that still sees them**, or a hidden title could never
be brought back. That exception is a TYPE, not a flag: `GET /library` answers
`Vec<GameEntry>` on every lane but the operator's and `Vec<OperatorGameEntry>` on
theirs, so a hidden entry cannot reach a paired streaming client by someone forgetting
a filter — there is no field there to leak. `hidden` is skipped when false, so the
response is byte-identical to today's for a library with nothing hidden.
`PUT /library/hidden/{id}` is operator-only — neither the plugin lane nor a paired cert,
unlike the scanner toggle. A plugin has no business deciding what its operator sees, and
a client must not be able to hide a game on the host it is streaming from. The id is not
validated against the current library on purpose: a title can be legitimately absent at
that moment (launcher closed, plugin mid-sync, drive unmounted), and refusing the
operator's choice in that window is worse than storing an id that matches nothing today.
On the card, the poster dims and a Hidden badge says why — a faded tile with no label
reads as a broken cover. Its controls stay at full contrast and, unlike an ordinary
card's, are not hover-revealed: the un-hide button is the only way out of the state, and
hiding it behind a hover would strand anyone on a touch screen.
Verified on .21 (Linux): 469 host tests pass (5 new), clippy clean under `-D warnings`,
`cargo fmt --all --check` clean. The routing test is the one that earns its keep — every
library id contains a colon and Heroic's contain two, so a router that split on it would
404 the console against ids the host itself produced. Console: tsc clean, production
build clean, i18n 633 messages across en+de, biome clean on the touched files.
|
||
|
|
9089651406 |
Merge pull request 'The jitter ring only ever learned from clicks — it now grows on near-misses, un-does refused shrinks, and cashes growth on the click it already paid' (#111) from worktree-audio-jitter-lowwater into main
apple / swift (push) Successful in 1m36s
ci / web (push) Successful in 1m14s
ci / docs-site (push) Successful in 1m20s
ci / bun-nix (push) Successful in 1m59s
ci / rust-arm64 (push) Successful in 2m31s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 12s
ci / rust (push) Failing after 3m6s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 21s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 12s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 5s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 27s
deb / build-publish-client-arm64 (push) Successful in 1m55s
deb / build-publish-host (push) Successful in 4m38s
docker / builders-arm64cross (push) Successful in 9s
deb / build-publish (push) Successful in 5m5s
docker / deploy-docs (push) Successful in 32s
android / android (push) Successful in 10m35s
flatpak / build-publish (push) Successful in 7m16s
release / apple (push) Successful in 10m45s
windows-host / package (push) Successful in 12m21s
windows-host / winget-source (push) Skipped
windows-host / canary-manifest (push) Successful in 25s
arch / build-publish (push) Successful in 13m42s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m35s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m50s
apple / screenshots (push) Successful in 6m9s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m19s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m15s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 23m55s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 24m41s
Reviewed-on: #111 |
||
|
|
d237646c66 |
fix(host,sdk,kit): library scanners sat in the nav, could not sync local art, and so never got their settings
Three symptoms on .21, two defects. Lutris and Heroic appeared in the console sidebar
they explicitly opt out of; Lutris's settings were unreachable from the Library
screen; and Lutris and Steam logged `sync (startup) failed: HostRequestError`.
**The sidebar is a publish gap.** The console is correct — it keeps
`category: "library"` plugins out of the nav (`uiPlugins`, app-shell.tsx) — but the
host reports no category for them at all. `defineLibraryPlugin` sets it and
`sdk/src/ui.ts` forwards it; what SHIPS does not. `@punktfunk/host` was bumped to
0.1.2 on 2026-07-20 and `category` landed 2026-08-05 without a bump, so the registry's
0.1.2 is the pre-category build and every installed scanner registers without one.
Bumps the SDK to 0.1.3 — **inert until it is published**.
Because the field rides the untyped `pf.request` seam so an older host ignores it
rather than rejecting the registration, dropping it is silent by design. `serveUi` now
reads its own directory entry back and warns once when a requested category did not
land, the same way `defineLibraryPlugin` already warns when a store claim did not take.
That is what turns the next occurrence into a log line instead of a bug report.
**The missing settings and the failed sync are ONE defect: a write/read disagreement
about `file://`.** `local_art_bytes` decodes a `file://` value before testing
containment; `validate_art_paths` handed the raw value to `Path::new`, where
`file:///home/u/c.jpg` is a RELATIVE path whose first component is `file:`. It
canonicalized against the cwd, failed, and read as "outside every art root". So the
host refused every cover the kit's own `fileUrl` helper emits — the documented way for
a plugin to publish local art — while the read path would have served those same files.
That the two symptoms share a cause is not obvious and is why this is one commit: the
Library screen's settings control renders only for `origin: "plugin"`, and a source
becomes `plugin` only once it holds a store CLAIM, which is taken during a successful
reconcile. Lutris failed at entry 0 and Steam at entry 3, so neither ever claimed its
store, both stayed `origin: "builtin"`, and neither got a settings button. Heroic
reconciled (its art is http(s)) and has had its settings all along; rom-manager was
never affected because zero entries meant it never applied.
`art_path_is_servable` now decodes first, so both halves of the confinement judge the
same string. Confinement itself is unchanged: an out-of-root path is still refused in
`file://` clothing, which the test asserts alongside the accept case.
Diagnosing this took the HOST's journal, because both surfaces that should have
explained it lied. `HostRequestError` stringified to its bare tag, so the sync engine's
`${e.cause}` logged `HostRequestError` and discarded the method, the path and the
host's own message; it now renders all three, including an object-shaped cause that
used to print `[object Object]`. And the host logged "payload carries a field this lane
may not set" for BOTH refusals in `check_entry_fields`, so a 400 about an art path read
as an auth problem — it now logs the real reason and the entry title.
Verified on .21 (Linux): 463 host tests pass, clippy clean under `-D warnings`,
`cargo fmt --all --check` clean. The new art test fails without the fix and passes with
it. plugin-kit 71 and SDK 72 tests pass, both typecheck clean, biome clean.
|
||
|
|
69728b6f4e |
fix(pf-presenter): "Native resolution" streamed the compositor's POINTS, not the panel's pixels
ci / bun-nix (pull_request) Successful in 23s
ci / web (pull_request) Successful in 1m11s
ci / docs-site (pull_request) Successful in 1m19s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m33s
apple / swift (pull_request) Successful in 1m33s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 3m25s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m22s
android / android (pull_request) Successful in 4m41s
ci / rust (pull_request) Successful in 6m52s
A CachyOS / KDE Plasma 6.7.4 Wayland client with its 2560x1600@165 laptop panel at 150 % scaling negotiated 1706x1066 for "Native resolution" and streamed a visibly blurry image. Two independent defects, and they stack — which is why forcing the mode to 2560x1600 by hand did not fully fix it either. 1. `SDL_GetDesktopDisplayMode` reports a mode in SCREEN COORDINATES and hands the pixels-per-point ratio back separately as `pixel_density`. We read `m.w`/`m.h` raw. KDE advertises that panel as 1707x1067 points with a density of ~1.4997, `render_scale::apply` even-floors both odd axes, and 1706x1066 goes on the wire — exactly the mode in the reporter's handshake log. Multiplying by the density recovers 2560x1600 to the pixel, because SDL derives it as the output's exact pixels/points ratio. On X11 and Windows SDL never sets a density and `SDL_video.c` normalizes the unset 0.0 to 1.0, so this is inert there: the bug needed a compositor doing FRACTIONAL scaling. 2. The SDL window was created without `HIGH_PIXEL_DENSITY`, so the Wayland surface stayed at buffer scale 1 — the Vulkan swapchain was built at 1707x1067 and KWin upscaled it to the glass. Even a correct 2560x1600 stream was resampled down and then back up. The same flaw silently shrank "Match window", which asks the host for `size_in_pixels()`. The reporter's `SDL_VIDEO_WAYLAND_SCALE_TO_DISPLAY=1` workaround is this same fix applied from outside SDL, which is why it helped. The surrounding code was already written for pixels != points — the swapchain, match-window and pointer mapping all read `size_in_pixels()` while window-size persistence reads logical `size()` — so the flag only makes those two stop being the same number. `display_scale()` starts reporting 1.5 into a swapchain that is 1.5x larger, leaving the OSD the size it already was. Also closes a smaller hole on the way past: only an `Err` from SDL reached the 1920x1080 fallback, so a display that reported a 0x0 mode sent a 0x0 request. Verified on home-worker-5 (CachyOS — the reporter's distro, real SDL 3.4.14): `cargo clippy --all-targets -p pf-presenter -- -D warnings` clean and 18/18 pf-presenter tests pass, three of them new and pinned to the field-reported numbers. |
||
|
|
3bb87d260e |
fix(audio): detect jitter before it is audible, and stop re-probing a depth the link just refused
apple / swift (pull_request) Successful in 1m38s
apple / screenshots (pull_request) Skipped
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 2m1s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m44s
ci / web (pull_request) Successful in 1m38s
android / android (pull_request) Successful in 4m52s
ci / docs-site (pull_request) Successful in 1m33s
ci / rust-arm64 (pull_request) Successful in 4m10s
ci / bun-nix (pull_request) Successful in 28s
ci / rust (pull_request) Successful in 9m26s
The 0.25.0 MacBook field report — audio jitter 'at certain points' — is the jitter policy learning exclusively from audible failures, on both of its sides. Growth needed THREE audible underruns before deepening the ring; the A/V sync loop re-tested a shallower ring every five quiet seconds and paid an audible starvation event every time it was wrong, forever; and a grown target was never re-banked — growth raises a threshold, only a re-prime deepens the ring — so a bunching link rode the knife edge, clicking once per bunching period with the 'grown' target sitting inert. A ten-minute simulation of the Wi-Fi power-save pattern (25 ms gaps / 300 ms, −50 ppm skew) measured ~2000 audible events under the shipped policy. Three mechanisms, in JitterPolicy (Linux/Windows/Android) and mirrored in the Swift AudioRing: - NEAR-MISS: a read served with less than one protocol frame left over is the same evidence as an underrun, heard by no one. It grows the target one step per window, BEFORE the click — waiting for the third audible underrun means the user heard two. - SHRINK PROBES: every shrink is armed for five seconds; answered by an underrun or near-miss it is undone on the spot, and a failed sync-driven shrink is not retried for a doubling backoff (60 s → 8 min). A probe that survives resets the backoff. Continuity outranks sync, now with a memory. - HOLLOW RE-PRIME: an underrun while the depth AVERAGE runs more than a step below the target re-primes immediately, spending the click it already cost on the whole refill instead of limping. The average, not the instant, is what separates a hollow ring from one late packet, and it is seeded on prime so a fresh ring is never spuriously hollow. Same simulation after: 9 audible events, tail clean but for the clock-skew re-anchor (a genuinely slow host must re-bank every few minutes; only rate adaptation would remove that, and no client has it). Neutralising the three constants reproduces the ~2000 — the convergence tests fail against the old behaviour. Verified: 203 punktfunk-core tests, 254 Swift tests (5 skipped), clippy -D warnings on punktfunk-core --all-features, cargo fmt --all --check. |
||
|
|
86bb09e2cf |
Merge pull request 'Arch could upgrade FFmpeg out from under the host and brick it — and the host now builds against FFmpeg 9' (#108) from worktree-ffmpeg9-support into main
apple / swift (push) Successful in 1m28s
android / android (push) Canceled after 0s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 0s
audit / cargo-audit (push) Canceled after 0s
audit / bun-audit (plugin-kit) (push) Canceled after 0s
audit / bun-audit (sdk) (push) Canceled after 0s
audit / bun-audit (web) (push) Canceled after 0s
audit / docs-site-audit (push) Canceled after 0s
audit / pnpm-audit (push) Canceled after 0s
audit / license-gate (push) Canceled after 0s
ci / rust (push) Canceled after 0s
ci / rust-arm64 (push) Canceled after 0s
ci / web (push) Canceled after 0s
ci / docs-site (push) Canceled after 0s
ci / bun-nix (push) Canceled after 0s
deb / build-publish (push) Canceled after 0s
deb / build-publish-host (push) Canceled after 0s
deb / build-publish-client-arm64 (push) Canceled after 0s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Canceled after 0s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Canceled after 0s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Canceled after 0s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / builders-arm64cross (push) Canceled after 0s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 0s
nix / flake (push) Canceled after 0s
release / apple (push) Canceled after 2m25s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 0s
windows-host / package (push) Canceled after 0s
windows-host / canary-manifest (push) Canceled after 0s
windows-host / winget-source (push) Canceled after 0s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Canceled after 4s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Canceled after 0s
windows / build (aarch64-pc-windows-msvc) (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 0s
decky / build-publish (push) Successful in 1m6s
Reviewed-on: #108 |
||
|
|
deeb8b6700 |
feat(pf-encode): build against FFmpeg 9
apple / swift (pull_request) Successful in 1m53s
apple / screenshots (pull_request) Skipped
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 2m34s
ci / web (pull_request) Successful in 2m32s
ci / docs-site (pull_request) Successful in 1m25s
ci / bun-nix (pull_request) Successful in 26s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 3m23s
android / android (pull_request) Successful in 6m47s
ci / rust-arm64 (pull_request) Successful in 8m49s
nix / flake (pull_request) Failing after 16m7s
ci / rust (pull_request) Successful in 23m39s
ffmpeg-next 8.1.0 could not accept FFmpeg 9 at all: ffmpeg-sys-next's version probe
covered avcodec majors 56..62 (the range is exclusive of its end), so libavcodec 63 fell
outside what it knew how to bind. 9.0.0 widens that to 56..63, which is what actually
unblocks Arch. Bump both pins — the unconditional Linux dep and the optional Windows
amf-qsv one — and the lock with them.
No API drift to fix. The crate major is a CEILING, not a target: one source tree still
spans FFmpeg 7.x/libavcodec 61, 8.x/62 and 9.x/63 via per-version cfgs, and every wrapper
symbol the NVENC-libav, VAAPI and amf-qsv backends name survives 8.1.0 -> 9.0.0
unchanged. The three hand-written #[repr(C)] hwcontext mirrors are the parts no compiler
checks, so they were re-read against the real headers rather than trusted:
AVCUDADeviceContext and AVD3D11VAFramesContext are byte-identical across 7.1/8/9, and
AVD3D11VADeviceContext gained two trailing UINTs in 8 that 7.1 lacks — which is why that
mirror deliberately stops at the common prefix, and why its assertions now say what they
do and do not buy you. They pin our layout, not libav's; a green build is not evidence.
The CI image is the step that makes this reach users. arch.yml deliberately runs no -Syu
("the image's snapshot IS the build environment"), so the builder stayed frozen on ffmpeg
8 no matter what Arch shipped, and a canary built from that snapshot could not satisfy the
soname dep the PKGBUILD now derives. Re-keying ci/ rebuilds it against ffmpeg 9.
Ubuntu and Windows deliberately stay put: the noble .deb bundles its own FFmpeg 8 behind
an rpath and strips the libav sonames from its Depends, and Windows bundles BtbN DLLs into
the signed installer — neither is exposed to the break, BtbN publishes no FFmpeg 9 build,
and moving either would re-qualify an encode stack to buy nothing.
Verified end to end on 192.168.1.21 (CachyOS, system ffmpeg 2:9.0-5, RTX 5070 Ti): host
builds clean and links libavcodec.so.63/libavutil.so.61/libavfilter.so.12/libswscale.so.10
with no unresolved sonames; the ffmpeg-8 compat shim is gone and the service runs with
NRestarts=0 and answers 401 on :47990; pf-encode's 67 tests pass; and a live synthetic
encode drives real NVENC hardware through FFmpeg 9's libavcodec to a decodable 1080p HEVC
stream (180/180 frames, FEC loopback 0 mismatches) with libavcodec.so.63 and
libnvidia-encode both mapped into the encoding process.
|
||
|
|
cabd011f1d |
Merge pull request 'The 272 ms audio buffer was legal: the PipeWire callback filled the buffer ceiling, not the graph's request' (#106) from fix/pw-playback-requested into main
apple / swift (push) Successful in 1m34s
ci / web (push) Successful in 1m3s
ci / bun-nix (push) Successful in 17s
ci / docs-site (push) Successful in 1m15s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m37s
ci / rust-arm64 (push) Successful in 3m1s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 8s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 9s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 13s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 12s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 27s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 23s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 17s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 17s
android / android (push) Successful in 5m6s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 2m43s
apple / screenshots (push) Successful in 6m5s
deb / build-publish (push) Successful in 4m8s
deb / build-publish-host (push) Successful in 3m54s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m9s
deb / build-publish-client-arm64 (push) Successful in 4m36s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m11s
ci / rust (push) Canceled after 2m31s
docker / builders-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
arch / build-publish (push) Successful in 8m34s
flatpak / build-publish (push) Successful in 18m15s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m28s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 22m4s
Reviewed-on: #106 |
||
|
|
be86cfcdc0 |
fix(client/audio): the PipeWire callback stops filling the buffer ceiling every cycle
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m12s
apple / swift (pull_request) Successful in 1m38s
apple / screenshots (pull_request) Skipped
ci / bun-nix (pull_request) Successful in 23s
ci / web (pull_request) Successful in 1m0s
ci / docs-site (pull_request) Successful in 1m8s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m18s
ci / rust-arm64 (pull_request) Successful in 1m40s
android / android (pull_request) Successful in 5m24s
ci / rust (pull_request) Successful in 12m51s
The playback process callback sized its writes from the mapped buffer's capacity — PipeWire's quantum-limit, 8192 frames ≈ 170 ms — instead of the graph's per-cycle ask (pw_buffer.requested). Every cycle therefore queued up to 170 ms of PCM downstream of the ring, and, worse, taught JitterPolicy that the device drains 170 ms per callback: the underrun floor (want + one frame) rose above any depth the A/V sync loop may request, so sync measured audio ~280 ms late and was forbidden — by its own continuity rule — from draining it. The first on-glass run of the latency overhaul showed exactly that: audio buffer 272 ms, a/v +284 ms, stable. Honor requested (capacity remains both the ceiling and the fallback for requested == 0), and log requested-vs-capacity once per stream in the shape of the host's per-capture-open quantum line, so the next on-glass report can say which one is sizing the writes. Needs libpipewire >= 0.3.49 (2022-03) for the requested field; every ship target clears that. Verified on .21: cargo clippy -p pf-client-core --all-targets -D warnings clean, 167 tests pass, fmt clean. |
||
|
|
e9e1ec7dc5 |
fix(pf-inject): the DualShock 4 Windows backend never imported OFF_INPUT
ci / web (pull_request) Successful in 1m22s
apple / swift (pull_request) Successful in 1m32s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m58s
ci / docs-site (pull_request) Successful in 1m25s
ci / bun-nix (pull_request) Successful in 30s
android / android (pull_request) Successful in 5m2s
ci / rust (pull_request) Successful in 8m2s
The Windows host does not build:
error[E0425]: cannot find value `OFF_INPUT` in this scope
--> crates\pf-inject\src\inject\windows\dualshock4_windows.rs:65:48
error: could not compile `pf-inject` (lib) due to 1 previous error
`dualshock4_windows.rs` writes the neutral report straight to `OFF_INPUT` in its
bootstrap path — correctly, and exactly as the DualSense and Steam Deck backends
do: the devnode does not exist yet at that point, so there is no reader to race
and no seqlock to take. Its steady-state path already goes through
`publish_input`, which is the v2.3 seqlock.
But the import list only names `publish_input`. `steam_deck_windows.rs` imports
`OFF_INPUT` explicitly for the same bootstrap write; this one was missed when the
list was edited to add `publish_input`.
One word in a `use`. No behaviour.
WHY CI DID NOT CATCH IT: `pf-inject`'s Windows backends compile only for
`*-pc-windows-msvc`, and the crate is host-side, so the client Windows workflow
never touches it. A cargo check from a Mac cannot stand in either — pf-inject
pulls punktfunk-core and therefore ring, whose C build wants MSVC headers, so the
cross-check dies in cc-rs long before it reaches this file.
FOUND BY: running windows-host.yml's own build line on the CI runner (.133)
against the v0.25.0 release tree before tagging —
`cargo build --release -p punktfunk-host --features nvenc,amf-qsv,qsv`. It fails
at `pf-inject`, which is step 1 of the host job, so a v0.25.0 tag would have
produced no Windows host binary, no installer, and no host asset on the release.
|
||
|
|
7f82bca9c0 |
fix(pf-dxvadec): a wrapped sentence turned "6." into an ordered list
ci / bun-nix (pull_request) Successful in 23s
ci / web (pull_request) Successful in 1m10s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m17s
ci / docs-site (pull_request) Successful in 2m3s
ci / rust-arm64 (pull_request) Successful in 2m16s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m10s
apple / swift (pull_request) Successful in 1m34s
apple / screenshots (pull_request) Skipped
android / android (pull_request) Successful in 4m59s
ci / rust (pull_request) Successful in 6m3s
A doc paragraph in `pic_av1.rs` wrapped so that "first at frame / 6. Releasing…" put `6.` at the start of a line. rustdoc reads that as an ordered-list item starting at 6, which makes the following unindented `///` line a lazy continuation — `clippy::doc_lazy_continuation`, denied by `-D warnings`. Reflowed so the number cannot begin a line. Prose is byte-identical in content; only the wrap points move. No code, no behaviour. WHY THIS MATTERS FOR THE TAG. `pf-dxvadec` is Windows-only, and no Windows leg runs on a push to main — so main being green proves nothing about this. The failure surfaces for the first time in a release tag's fan-out, which is exactly what happened to the FIRST v0.23.0 tag: it went red on Windows clippy for this same lint, and the cure was a tag re-point. Caught pre-tag by re-running the lazy-continuation scanner over the tree while preparing v0.25.0 (0 hits before this commit's parent merged the new decode crates, 1 after). Cannot be verified by compiling here — the crate does not build on macOS — so the evidence is the scanner plus the lint's own rule, not a clippy run. |
||
|
|
3d20f2c0e5 |
Merge pull request 'Three decode rungs were decoding into a surface they were predicting from' (#102) from integration/decode-aliasing-program into main
ci / bun-nix (push) Successful in 30s
ci / web (push) Successful in 1m22s
ci / docs-site (push) Successful in 1m24s
apple / swift (push) Successful in 1m35s
ci / rust-arm64 (push) Successful in 1m56s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 16s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 16s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 16s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 14s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
deb / build-publish-client-arm64 (push) Successful in 3m6s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m11s
android / android (push) Successful in 5m57s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m21s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m37s
apple / screenshots (push) Successful in 5m52s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 2m24s
deb / build-publish-host (push) Successful in 6m20s
ci / rust (push) Canceled after 6m56s
docker / builders-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
windows-host / package (push) Failing after 1m37s
windows-host / canary-manifest (push) Skipped
windows-host / winget-source (push) Skipped
deb / build-publish (push) Successful in 5m39s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m47s
arch / build-publish (push) Successful in 11m19s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m19s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 6m37s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 10m20s
flatpak / build-publish (push) Successful in 8m54s
Reviewed-on: #102 |
||
|
|
2b167595aa |
docs(client): the VAAPI rung has parity now — say what is actually left
ci / bun-nix (pull_request) Successful in 28s
windows / build (aarch64-pc-windows-msvc) (pull_request) Failing after 30s
apple / swift (pull_request) Successful in 1m36s
apple / screenshots (pull_request) Skipped
ci / docs-site (pull_request) Successful in 1m19s
ci / web (pull_request) Successful in 1m30s
ci / rust-arm64 (pull_request) Successful in 2m39s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m26s
android / android (pull_request) Successful in 4m31s
ci / rust (pull_request) Successful in 10m52s
Its rows still read "never frame-hash parity-checked: the rung exports a tiled dmabuf with no CPU-readable image, so parity needs a readback path that does not exist yet". That readback now exists, and all SEVEN legs came back bit-identical to libavcodec on RDNA3: vendored H.264 250/250, our host's low-delay H.264 120/120, vendored H.265 250/250, host low-delay H.265 120/120, HEVC Main 10 50/50 as P010, vendored AV1 250/250 of 274 decoded, and host low-delay 4K two-tile AV1 60/60. The two arms collapse into one, because the thing that split them — AV1 having evidence the other legs lacked — is gone. Every leg now has the same evidence. It stays `verified = false`, and the note says why in the words the unproven-rung test requires: it has NEVER run on a second vendor and has never been soaked. That is a real limit rather than a formality — every other verified pair in this table earned it on more than one part, and the D3D11VA AV1 row two entries up is a rung that passed on one vendor's driver while failing on another's. The second reason is not about evidence at all, and it belongs in the record rather than in a commit nobody reads later: flipping this flag is a ROUTING change. `native_rung_admitted` is `verified || !below.verified`, so a verified VAAPI outranks Vulkan Video on every Linux AMD and Intel client — the Steam Deck included. The parity result justifies that change; it should still be made on purpose, by someone who wants it, rather than arriving as a side effect of writing down a test result. |
||
|
|
a9e7c033c3 |
test(client/vaapi): the last rung of the ladder, finally checked in pixels — 7 legs, all bit-identical
Every other decode rung earns `verified` with frame-hash parity against
libavcodec. VAAPI could not: it hands out a DRM-PRIME dmabuf whose memory the
driver tiles, so nothing could read its decoded pixels back, and all four of its
legs sat at "never frame-hash parity-checked".
That was never bookkeeping. The D3D11VA AV1 rung decoded 250 frames, streamed
4K60 through a clean five-minute soak, and produced WRONG PIXELS for 186 of 250
frames on NVIDIA and 245 of 250 on Intel. It looked perfect on glass; only the
goldens caught it, and the same defect turned out to be in H.264 on two other
rungs. VAAPI was the one rung where that class of bug could still be sitting
with nothing able to see it.
It is not. Measured on .25 (Radeon 780M, RDNA3, radeonsi, Mesa 26.0.3, VA-API
1.23) on 2026-08-08, against the SAME golden files the Vulkan and D3D11VA rungs
are held to, read across the crate boundary rather than copied:
H.264 vendored vector 250/250 bit-identical (7 from the flush)
H.264 our host, low-delay 640x480 120/120 bit-identical (3 from the flush)
H.265 vendored vector 250/250 bit-identical (2 from the flush)
H.265 our host, low-delay 640x480 120/120 bit-identical (0 from the flush)
HEVC Main 10, P010 50/50 bit-identical (2 from the flush)
AV1 vendored vector 250/250 delivered of 274 decoded, and
display frame 0 byte-identical to
libavcodec's own PIXELS
AV1 our host, 4K two-tile 60/60 bit-identical
⚠ ONE vendor. AMD/radeonsi only; no Intel iHD box has run these legs.
The readback that made it possible:
* `pf-vaadec`'s `va` module gains `VAImage` and `VAImageFormat`, hand-declared
with every size and offset measured off libva 2.23.0's real headers by
`layout-probe.c` and pinned as compile-time assertions — the same discipline
the decode buffers already keep. The trap: `VAImage::width`/`height` are
16-bit, so `data_size` sits at 60 and not at the 64 counting 32-bit fields
gives, and every field after them is two bytes earlier than it looks.
* `pack_two_plane` is the pure geometry — the crop to the picture, the padding
columns dropped per row, and the chroma plane taken from the driver's OWN
`offsets[1]` rather than from `pitch * display_height`, which is the 1088-row
smear this program has already paid for once. It needs no device, so ten CPU
tests cover it on macOS and in the container.
* `video_vaapi_native::parity` drives the seven streams above through the
production entry point and hashes what the rung DELIVERS, in delivery order,
tail included — so the delivery path is under test as well as the decode, and
a frame's surface comes from its own release token rather than from an
inference about which pool entry holds which picture.
THE READBACK CANNOT REACH THE PRODUCTION PATH, and that is structural rather
than a promise. `vaDeriveImage`, `vaCreateImage`, `vaGetImage`, `vaMapBuffer`
and the rest are resolved by a `#[cfg(test)]` type that dlopens libva itself;
the production `Libva` gains no field; `sha2` is a dev dependency. A CPU test
scans this file's own source and fails if any of those symbols is dlsym'd
outside the harness, so a refactor cannot quietly undo it.
Derive is not guaranteed, so both routes are implemented and neither is
optional: `vaDeriveImage` first, `vaCreateImage` + `vaGetImage` as the fallback
(which also detiles), and if neither yields the pool's own fourcc the leg FAILS
naming what the driver gave it. There is no skip path — a parity test that
passes because it could not read anything is the failure mode this program has
been bitten by three times. Both answer on radeonsi, the first frame of every
leg is read through BOTH and they must agree, and `PF_VAAPI_READBACK=getimage`
reproduces the H.264 leg's 250/250 through the copying route alone, so the
fallback is exercised rather than merely written.
And it can fail — proven, not asserted. Planting the real geometry defect this
driver's layout makes visible (rows read contiguously, ignoring the 512-byte
pitch behind a 320-wide picture) fails at display frame 0 with the full
localisation: 68312 luma and 14998 chroma samples differing, max |delta| 255,
luma bounding box (0,1)..(319,239) — and with the goldens forced through one
route, 250/250 diverging with "suspect the readback geometry". `compare` and
`localise` also have CPU counterfactuals, and a hardware leg proves the readback
reads real and DISTINCT pixels and localises a one-byte flip to the exact pixel.
⚠ One thing the hardware legs do NOT cover, found by planting the other defect
and watching it do nothing: radeonsi's decode surfaces for every fixture here
have no VERTICAL padding — `offsets[1]` is exactly `pitch * height` — so the
chroma-plane trap is untested on this driver, and `pf-vaadec`'s
`reading_chroma_at_the_display_height_would_have_been_caught` is the only place
it is checked at all. `probe_this_machines_readback_routes` now prints the
derived layout and says which of the two it is, so the next driver answers for
itself instead of being assumed.
|
||
|
|
bfed711921 |
Merge remote-tracking branch 'origin/main' into audio/latency-overhaul
ci / bun-nix (pull_request) Successful in 33s
ci / web (pull_request) Successful in 1m9s
ci / docs-site (pull_request) Successful in 1m21s
apple / swift (pull_request) Successful in 1m33s
apple / screenshots (pull_request) Skipped
ci / rust-arm64 (pull_request) Successful in 2m13s
windows / build (x86_64-pc-windows-msvc) (pull_request) Successful in 2m36s
windows / build (aarch64-pc-windows-msvc) (pull_request) Successful in 1m9s
android / android (pull_request) Successful in 7m53s
ci / rust (pull_request) Successful in 12m23s
|
||
|
|
f926bab9f7 |
fix(client): the native VAAPI rung stopped dropping decoded frames on the floor
`finish` showed `outputs.last()` and retired every other picture an access unit bumped out of the DPB without ever displaying it, and nothing flushed the DPB at end of stream. Measured on .25 against the vendored vectors: 225 of 250 frames for H.264, 204 of 250 for H.265, 45 of 50 for HEVC Main 10. D3D11VA and Vulkan deliver every frame, so this was the rung's alone. All four legs now deliver 250 / 250 / 50 / 250. The same function carried a second defect. `DmabufFrame::keyframe` was stamped with the CURRENT access unit's `is_idr`, not the flag of the picture it was about to display, and on a reordering stream those are different pictures: the IDR is bumped out several units after it decodes and arrived flagged `false` on all three legs' first frame, while a later AU draining the DPB flagged some old trailing picture as a keyframe. That field is `DecodedImage::is_keyframe`, the pump's post-loss re-anchor signal, so a mislabel re-anchors on the wrong frame. Three changes, all inside this rung: * **A deliverable queue**, the same shape as `video_vk_native`'s — extend, ship the front, trim the oldest past the bound, count and rate-limit the drops into `DecodeHealth::dropped`. Its DEPTH is derived differently and the divergence is documented: the Vulkan rung's bound is `HOLD_HEADROOM - PIPELINE_HOLD` = 1 because a queued frame there counts against the pool ON TOP of the DPB's own residency. Here the three claims are disjoint and a bumped picture MOVES from `pending`/slot to `held`, so the queue inherits the claim rather than adding one. The bound is the DPB's depth — the deepest carry-over a bump can leave — and the measured cost is at most one surface (zero on H.264, whose three seven-picture IDR drains are the deepest bursts these vectors have). A bound of 1 would have left 235 of 250 on H.264, most of the defect still in place. * **An end-of-stream flush.** This rung has no EOS signal and cannot have one: the pump feeds access units until the session ends and then drops the decoder. So `flush` has the two honest callers — `Drop`, where nothing can be presented and the job is to release the queue's surfaces and the DPB's before the pool goes, and a caller that KNOWS the stream ended, which today is the conformance harness. One walk, not a production path and an untested teardown path. AV1 needs none: it shows at most one frame per temporal unit and buffers nothing, which its 250/250 says out loud. * **`PictureFacts` recorded when a picture decodes**, and read back when it is displayed. `keyframe` was the defect; `color` and `display` are the same mistake one field along — an in-band HDR switch changes the VUI mid-stream and AV1's render region is per-frame, so a queued frame shown two units later would have been drawn with the newest picture's signalling. Concealment answers `Ok(None)` and deliberately does NOT drain the queue, which is the Vulkan rung's order and is load-bearing: `clears_demotion_streak` is `delivered || !concealed`, so shipping a queued frame on a concealed AU would zero the streak and take away the escape hatch that stops a rung concealing forever from holding a frozen picture. The three delivered-count assertions moved with the fix, and so did the CPU derivation that reproduces them without a GPU — it now simulates the whole delivery model (ledger, queue, one-per-AU hand-off, flush) in the order `decode` does it, and carries the old behaviour beside the new one as a counterfactual: a queue bound of 0 with no flush still reproduces 225/204/45 exactly, and the test fails if it ever stops being SHORT. `settle` was split out as the pure half of `finish` so the claim walk, the display ordering and the picture facts are all assertable with no device; `the_queue_never_needs_a_surface_the_pool_does_not_have` runs the surface-lifetime arithmetic over the real vectors and pins the peak claims (9 of a 16-surface pool on H.264, 8 of 14 on both HEVC vectors), with an unbounded queue as the counterfactual that shows the bound doing its job. Gates run: `cargo fmt --all -- --check`, `cargo clippy -p pf-client-core -p pf-vaadec --all-targets --features sdl3/build-from-source -- -D warnings`, `cargo test -p pf-client-core --lib --features sdl3/build-from-source` (176 pass), the same filtered to `video_vaapi_native -- --include-ignored` (23 pass, 0 ignored) and `cargo test -p pf-vaadec` (48 pass) — all on .25 (Radeon 780M, RDNA3, radeonsi, Mesa 26.0.3, VA-API 1.23); plus `cargo fmt --all -- --check` and `cargo clippy --workspace --all-targets -- -D warnings` in pf-lxcheck2. |
||
|
|
12a5318397 |
fix(audio): place audio with the picture instead of wherever the ring settles
The host stamps `pts_ns` on every audio datagram and the client decoded it
into `AudioPacket` — and then never read it. Video's `pts_ns` is used end to
end (the presenter computes a true glass-to-glass `displayed + clock_offset −
pts`), so audio free-ran at whatever depth its jitter ring happened to reach,
video was presented on an independent path, and nothing ever compared them.
The A/V offset was an accident of buffer depths: it moved whenever the ring
ratcheted under underrun pressure, and it got WORSE every time video got
faster, because a quicker decoder lowers the video leg and leaves audio's
exactly where it was. That is what a field report on the Steam Deck heard as
"the audio delay is way too high", and it is why shaving milliseconds off the
audio budget had not helped.
Video is the master. In a game streamer the video leg is the input-feel budget
and must never be inflated to satisfy the audio clock, while audio tolerates
small crossfaded corrections that are inaudible — and `crossfade_drop` already
applies them. So audio moves:
audio_e2e = (now + buffered_ahead + clock_offset) − pts_ns
av_offset = audio_e2e − video_e2e (> 0 ⇒ audio behind the picture)
`AvSync` smooths that with an EWMA, ignores what sits inside a deadband no
listener can detect, refuses the implausible outright rather than clamping it
(a wall-clock step must not steer the ring), and proposes a depth.
Continuity outranks sync, always. `JitterPolicy::set_sync_target` only ever
takes a REQUEST, clamped between the existing underrun-driven floor and the
hard cap. A link whose jitter genuinely needs more buffer than the picture is
away keeps its buffer and the residual is reported — sync can never starve the
ring into dropouts. `None` is the default and reproduces the previous behaviour
exactly, so the four client rings can adopt this one at a time without
diverging.
Two upstream defects found on the way, both prerequisites:
* The host stamped `pts_ns` at ENCODE time, inside the loop draining an
already-accumulated chunk, so every frame of a chunk carried near-identical
timestamps describing when we got round to encoding. Harmless while nothing
consumed it; a sync loop regulating against it would regulate against a
fiction. It now comes off the capture clock.
* The host did not pace. One capture callback hands over a whole quantum — 5 ms
when the graph honours our ask, 21.3 ms on a VM, where stock PipeWire raises
`min-quantum` to 1024 — and the loop drained all of it into back-to-back
`send_datagram` calls. The wire carried a 4-5 frame burst then ~21 ms of
nothing, and a ring can only absorb that by standing a burst period deep.
Frames now leave on the audio clock, which costs no average latency.
And the reason none of this was visible: `buffer_ms`/`target_ms` existed only
as a `tracing::debug!` line, absent from `Stats`. On a Deck the client runs
under Steam's `reaper` with stdout on a pipe nobody can read, so the one number
identifying a deep ring was unobtainable on the device reporting the latency.
The HUD now carries `audio buffer N ms · a/v ±N ms` — both, because a deep ring
on a jittery link is correct and only the offset separates that from audio held
late. The host also reports its negotiated quantum against the one it asked
for, per capture open rather than once per process.
Verified: 364 core + 40 presenter tests on Linux, clippy -D warnings clean on
punktfunk-{core,host} + pf-{client-core,presenter}, fmt clean. New tests pin
the safety invariant (sync cannot pull the target below the continuity floor on
any preset), that `None` leaves the policy bit-identical, and that a device
quantum exceeding the hard cap does not panic `Ord::clamp` inside a realtime
callback.
Android and Apple keep today's behaviour (the `None` default) until their
presenters publish a video figure to align against; design/audio-latency-
overhaul.md carries the plan.
|
||
|
|
1482e6b373 |
docs(client): all four VAAPI legs have decoded — the evidence table said two never had
The H.264 and H.265 rows still read "NEVER decoded a frame on any hardware". That stopped being true on 2026-08-07, in the same session that proved AV1: every access unit of the vendored H.264 (250), H.265 (250) and HEVC Main 10 (50) vectors was accepted on .25 (Radeon 780M, RDNA3, Mesa 26.0.3) with no decode error — NV12 for the 8-bit legs, P010 for Main 10, all on the same tiled AMD modifier — and probe_this_machines_libva reports VLD decode for all three profiles. The row records the delivered counts honestly rather than rounding them up: 225/204/45 against 250/250/50 access units, because `finish` shows `outputs.last()` and drops the other pictures an AU bumps, and nothing flushes the DPB at end of stream. That is this rung's own behaviour — D3D11VA delivers all 250 — and it is invisible on punktfunk's zero-reorder host output. It is recorded and asserted rather than fixed: changing the one-frame-per-AU contract touches the pump's deliverable queue, an end-of-stream flush, and the `keyframe`-labels-the-access-unit defect in the same function, so it belongs in a commit that moves all three. Still `verified = false` for all four, and the note says why in the words the unproven-rung test requires: never frame-hash parity-checked. That is not pedantry — the D3D11VA AV1 row two lines above is a rung that decoded 250 frames and produced wrong pixels for every one of them. Parity is what distinguishes them, and this rung exports a tiled dmabuf with no CPU-readable image, so it needs a readback path nothing has written yet. |
||
|
|
79afa9ce79 | Merge branch 'fix/hevc-lowdelay-parity-gate' into integration/decode-aliasing-program | ||
|
|
bb9f482b2c | Merge branch 'fix/vaapi-decode-target-aliasing' into integration/decode-aliasing-program | ||
|
|
dc116d28ca | Merge branch 'fix/vaapi-h264-h265-hardware-proof' into integration/decode-aliasing-program | ||
|
|
d25a20a233 |
feat(vkdecode): the AV1 rungs meet a second tile for the first time
Every AV1 frame either decode rung has ever been measured against is `tile_cols = tile_rows = 1`. The vendored vector is single-tile on all 274 of its frames, so every tile array the conversions fill — `tiles.widths`, `tiles.heights`, the per-tile records — had only ever been written at index 0, and a conversion that wrote tile 0 and left the rest zero would pass the whole suite. Our encoder splits 4K into TWO TILE ROWS. **The fixture.** `lowdelay-3840x2160.ivf.av1`, 261 KB, 60 frames — `punktfunk-host spike --source synthetic --codec av1 --width 3840 --height 2160 --fps 60 --seconds 1 --bitrate 1` on .21 (NVENC, RTX 5070 Ti), wrapped to IVF with `ffmpeg -f obu … -c copy` so `common::split_av1_aus` (the vendored parser's own `IvfIterator`) frames it exactly as it frames the vector, with no second splitter that could disagree. **4K is not a size choice, it is the only shape with the property.** Measured on the same box with the same command: 1280x720, 1920x1080 and 2560x1440 all give `tile_cols = tile_rows = 1`; 3840x2160 gives `tile_cols = 1, tile_rows = 2` with `width_in_sbs_minus_1 = [59]`, `height_in_sbs_minus_1 = [16, 16]`, and both tiles in ONE Tile Group OBU. 60 frames instead of 120 pays for the resolution: 261 KB, under both the 282 KB H.264 and 270 KB H.265 low-delay fixtures. Goldens are libavcodec's software decode, cross-checked between ffmpeg n8.1.2 (Arch x86_64, libdav1d) and 8.1.1 (Homebrew, macOS arm64, libdav1d) whose 746,496,000-byte raw outputs are BYTE-IDENTICAL, not merely equal per frame. 60 of 60 digests distinct. **AV1's frame accounting is asserted, never derived.** The vendored vector is 250 temporal units carrying 274 coded frames of which 24 are hidden; this stream is 60 units, 60 coded, 60 shown, 0 hidden, 0 `show_existing_frame`, 1 key frame. Neither is the general case, so both parity harnesses now take units / decoded / shown as three independent parameters instead of computing one from another, and the CPU guard states all six numbers. **A CPU gate that needed no hardware at all.** `pic_av1`'s new `a_two_tile_frame_fills_both_row_entries_and_leaves_the_rest_zero` pins the second row entry against its OWN `height_in_sbs_minus_1`, requires the two rows to tile the frame exactly, and requires TWO tile RECORDS out of ONE tile group with rows (0,0) and (1,0) — the transposition a square grid could never reveal — each spanning real bytes. The existing one-tile test asserts index 0 is right and `1..` are zero, which a broken multi-tile conversion also satisfies. ⚠⚠ **This is a file, and on AV1 that distinction has already cost a release.** "250/250 delivered frames bit-identical to libavcodec" was true for the entire period the host was shipping only the FIRST TILE of every 4K frame: the verification ran against a vendored file while the truncation lived in packetisation, and the suite stayed green throughout. This fixture closes the multi-tile gap on the DECODE rungs and closes nothing about fragmentation, reassembly, loss or AU boundaries — the golden header, both module docs and the leg docs all say so, at length, so the next reader does not inherit the same false confidence. Legs: `low_delay_host_av1_every_frame_hashes_bit_identical_to_libavcodec` on the Vulkan rung (11 ignored legs now) and on the D3D11VA rung, plus two non-ignored CPU tests. Verified: 11/11 Vulkan parity legs on .21 (RTX 5070 Ti, 610.57.04), the new one 60/60 bit-identical; workspace clippy `-D warnings` and `cargo fmt --all --check` clean on .21. |
||
|
|
e8a7a1e6af |
fix(client/vaapi): the third rung does NOT alias — and now it cannot start to
The D3D11VA and Vulkan rungs both decoded into a surface they were predicting from, on 117 of 120 access units of our own host's low-delay H.264 (`1c54d099` for AV1, `834b2443` for H.264). `pf-vaadec` feeds `reference_frames` from the same `plan.dpb_refs` snapshot, releases its whole `removed` list inline exactly as the two broken conversions did, and neither fix commit touched it. It is still exempt — this is the evidence, and the thing that keeps it true. **Measured on the CPU, no GPU needed.** `walk_for_aliasing` drives the planner and `plan_to_va` over both streams and counts four shapes. On `lowdelay-640x480.h264` the aliasing PRECONDITION is fully present: 117 of 120 access units remove a picture their own `dpb_refs` still names, and on the same 117 the setup picture is handed the slot of a picture that access unit READS — the D3D11VA/Vulkan defect verbatim, in this conversion, today. On the vendored conformance vector both counts are 0, which is why that vector proved nothing on two other backends for two milestones. Aliased submissions: **0 on both**. **Why.** A slot is not a surface here. `plan_to_va` never invents one — every reference it can name is read out of the `surfaces` table it is handed — and the decode target is a separate parameter the caller takes from OUTSIDE that table. `setup_surface` reaches the submission at exactly one field per codec (H.264/H.265 `curr_pic.picture_id`, AV1 `current_frame` and `current_display_picture`); HEVC is doubly safe, because its per-slice `RefPicList` stores an INDEX into `reference_frames` rather than a surface. AV1's documented substitution fallback is the one place the target can be named as a reference, and only where the store resolved nothing at all to prefer. **The exemption was incidental; it is structural now.** It needs the reference table and the decode target to come from ONE snapshot of the bindings, and the rung had that only by writing `free_surface()` and `surface_table()` adjacently at three call sites. Split them and this rung acquires the defect exactly: the table must be the PRE-removal one (that is where the references are), while a free list consulted after the removals offers precisely the displaced picture's surface. `Session::acquire_target` now returns the index, the surface and the table together from `&self`, so a later edit cannot move one call and not the other. No behaviour change: same order, same values, same refusal message. Tests. `no_submission_names_its_decode_target_as_one_of_its_own_references` (both streams, 0) with `taking_the_decode_target_from_the_slot_table_aliases_on_the_low_delay_stream` as the counterfactual that reproduces the defect on 117 of 120 — so the walk demonstrably CAN see it when it is there. `the_low_delay_stream_reassigns_slots_whose_pictures_it_still_reads` pins 0/250 and 117/120 so neither can drift silently. `the_decode_target_can_never_be_a_surface_the_reference_table_names` sweeps every binding state a 4-surface/3-slot pool can hold, and `taking_the_free_surface_after_the_removals_would_hand_out_a_referenced_surface` is the ordering counterfactual. ⚠ One existing test lost a VACUOUS half. `the_setup_picture_routinely_inherits_a_just_freed_slot` asserted the decode target was never also a reference while handing every picture its own never-reused surface id — distinct integers cannot collide, so that assertion could not fail whatever the conversion did. Its real measurement (225 of 250 access units reuse a just-freed slot, which is why the target is a parameter) is kept; the collision half is gone, and the doc says where the question is actually answered and why a recycling pool is what it takes to answer it. Gates, run on `.25` (Radeon 780M, radeonsi, Mesa 26.0.3, VA-API 1.23), this rung being Linux-only: `cargo fmt --all -- --check`; `cargo clippy -p pf-client-core -p pf-vaadec --all-targets --features sdl3/build-from-source -- -D warnings`; `cargo test -p pf-client-core --lib --features sdl3/build-from-source` (171 passed); the same filtered to `video_vaapi_native` with `--include-ignored` (18 passed); `cargo test -p pf-vaadec` (48 passed). Plus the pf-lxcheck2 container for the cross-platform half — fmt, clippy and `cargo test -p pf-vaadec`, all clean. All four VAAPI legs still decode with the refactor in place, not one access unit refused: H.264 225 of 250 access units delivering a frame, H.265 204 of 250, HEVC Main 10 45 of 50 (P010), AV1 250 of 250 — the same counts and the same tiled modifier 0x200000010401b04 those legs recorded before it. ⚠ The H.26x legs live on `fix/vaapi-h264-h265-hardware-proof`, not on this branch, so they were run by overlaying that commit's test module onto the scratch tree; only the AV1 leg and the libva probe are reachable from here. This is a decode measurement, not frame-hash parity — the rung exports a driver-tiled DRM-PRIME dmabuf, so there is no CPU-readable image to hash. The alias assertions above are the real evidence and they need no device. ⚠ NOT taken: `finish`'s `outputs.last()`, which ships one frame per access unit and drops the rest of what a bump displaces (225/204/45 against 250/250/50), with no end-of-stream flush. It cannot bite punktfunk — hosts emit zero-reorder output, so `outputs` never holds more than one picture — and fixing it changes `decode()`'s one-frame-per-access-unit contract with the pump (it wants a deliverable queue, which `video_vk_native` already keeps) plus an end-of-stream flush and the `keyframe`-labels-the-access-unit defect in the same function. It is recorded and asserted on that other branch, whose three delivered-count assertions any fix has to move in the same commit; doing that from here, blind to them, would be worse than leaving it. |
||
|
|
f0702f3e06 |
feat(vkdecode): HEVC's exemption stops being an argument and becomes a vendored stream
`fd6241a2` made HEVC's freedom from the release-ordering defect falsifiable on CPU and
recorded what was still missing: no low-delay HEVC stream was vendored, so the exemption
rested on a structural argument plus one throwaway measurement. This vendors the stream,
and the exemption HELD.
**The fixture.** `lowdelay-640x480.h265`, 270 KB, 120 pictures — `punktfunk-host spike
--source synthetic --codec h265 --width 640 --height 480 --fps 60 --seconds 2 --bitrate 1`
on .21 (NVENC, RTX 5070 Ti, driver 610.57.04). Deliberately the H.264 sibling's resolution
and frame count: the two are then directly comparable, 640 and 480 are both multiples of
MinCbSizeY so there is no conformance window and a hash mismatch can only be decode rather
than readback geometry, and 270 KB sits alongside the 282 KB already accepted for H.264.
Goldens are libavcodec's software decode, cross-checked BIT-IDENTICAL across ffmpeg n8.1.2
(Arch, x86_64) and 8.1.1 (Homebrew, macOS arm64), 120 of 120 digests distinct.
**The exemption held, measured rather than argued.** `sps_max_dec_pic_buffering_minus1 = 4`
against the four pictures 8.3.2 keeps marked in steady state, `sps_max_num_reorder_pics = 0`,
`numRefL0 = 1` — a five-picture DPB filled exactly by four references plus the current
picture. 115 of the 120 access units retire a picture, and `removed ∩ dpb_refs` is **0 of
120**. A 300-picture 1080p stream from the same host reports the same shape: 295
retirements, 0 intersections. It is the encoder and not the resolution, exactly as for
H.264.
**A zero proves nothing on its own, so the fixture is pinned by its counterfactual.**
`test-25fps.h264` reported zero for two milestones while every stream we ship aliased on
99% of its frames. So the guarantee here is not "we looked and it was fine": hand
`plan_to_dxva_h265` the marked DPB as it stood BEFORE `decode_rps` — the mutation a
snapshot move would cause, reconstructed exactly as `dpb_refs(N-1) ∪ {stored(N-1)}` — and
the alias appears on **115 of 120** access units, driven through the real conversion rather
than through planner arithmetic. If a regeneration ever produced a stream that reordered,
or a DPB deeper than its reference count, that 115 collapses to 0 and the tests say so
instead of continuing to pass.
**The two rungs are exempt for different reasons, and the asymmetry is now a gate.** DXVA
binds the whole marked DPB — `RefPicList` is spec-defined that way, and an RFI long-term
anchor has to survive in it — so its exemption really is `H265Planner`'s snapshot ordering,
one call away from being untrue. `plan_to_vk_h265` never reads `dpb_refs` at all:
`pReferenceSlots` is the slots the operation uses, so it binds the current RPS sets, which
`decode_rps` itself derives and which therefore cannot name a picture that same RPS just
dropped. A new test feeds that conversion the identical widened snapshot and asserts
nothing changes, so a future change making the Vulkan rung bind the marked DPB — a
legitimate thing to want, since a *Foll* anchor invisible to the hardware is the RFI
failure shape — fails loudly instead of silently acquiring the defect.
What the Vulkan pixel leg adds is therefore NOT aliasing coverage, and its docs say so:
it is the first HEVC frame either rung has decoded from our own encoder, under a DPB that
retires and reissues a slot on 115 of 120 access units back to back, where the vendored
vector's reordering keeps that eviction slack.
Legs: `low_delay_host_h265_every_frame_hashes_bit_identical_to_libavcodec` on the Vulkan
rung (10 ignored legs now, up from 9) and on the D3D11VA rung, plus three non-ignored CPU
guards that run in ordinary CI.
Verified: 10/10 Vulkan parity legs on .21 (RTX 5070 Ti, 610.57.04), the new one 120/120
bit-identical; workspace clippy `-D warnings` and `cargo fmt --all --check` clean on .21.
|
||
|
|
fd6241a24f |
fix(dxvadec): the review round — a doc that had become false, a warn-storm on renegotiation, and HEVC's exemption made falsifiable
Four findings, all real. **`SlotMap`'s own docs had become false.** "feed it every `DpbUpdate` in decode order (via `Self::apply` or `plan_to_vk`, which applies internally)" — `plan_to_vk` no longer applies internally, which is the entire point of the change, and `release`'s docs named it as one of the two things that may free a slot. A reader following those docs would build the next caller wrong in exactly the way this commit's parent fixed. Both now say which conversions defer, which one does not, and why H.265 is the one that does not. **The deferred release warned on a legitimate event.** `release_deferred` warned per id when a deferred release found no slot — but a renegotiation replaces the whole `Session`, and with it the slot map, INSIDE `plan`, while the planner's own drain reports every drained picture in that same access unit's `removed`. Every one of those ids then misses, and nothing is wrong. `debug!`, with the legitimate cause named so the illegitimate one stays diagnosable. **HEVC's exemption was asserted only in its consequence.** `the_current_picture_is_ named_by_curr_pic_and_never_aliases_a_reference` checked that no reference shares the decode target's slot — which on the vendored vector holds whether or not the reasoning behind it does. That is precisely how the H.264 leg passed for two milestones. The test now also asserts the PLANNER property the exemption rests on (`removed ∩ dpb_refs = ∅`, falsified by moving `dpb_snapshot()` above `decode_rps`), and records that the low-delay measurement was 0 of 300 against H.264's 297 of 300 from the same host and the same run. It also records what is still missing: no low-delay HEVC stream is vendored, so HEVC's freedom is a re-derivable argument plus one measurement, not a standing hardware leg. **Two stale cross-references.** Both AV1 conversions told the reader the H.264/H.265 zero was "measured on reordering vectors and not a proof" — the open question this commit's parent closed. They now say what the answer was. |
||
|
|
834b244301 |
fix(client): the H.264 twin was real — every low-delay picture decoded into a surface it predicted from
The AV1 review round flagged the H.264 leg as "plausibly the same defect, traced in source, not reproduced" and deliberately did not touch it. It is reproduced now, and it is worse than the AV1 one: it fires on 297 of 300 access units of every stream a punktfunk host emits, at 720p, 1080p and 2160p alike, on BOTH the DXVA rung and the Vulkan one. **Decided on the CPU, no GPU needed.** `H264Planner` snapshots `dpb_refs` in `begin_picture`, BEFORE `finish_picture` runs 8.2.5's marking and C.4.5.3's bump, so a picture the sliding window unmarks and the bump then evicts lands in both `dpb_refs` (which `RefFrameList` is built from) and `dpb.removed`. The conversion released the whole `removed` list and then assigned the decode target a slot; `SlotMap::assign` takes the lowest free slot, which is the one just vacated. `CurrPic = N` and `RefFrameList[k] = N`, in one submission. The two conditions have to coincide in ONE access unit, and low-delay H.264 is exactly what makes them: `max_num_reorder_frames = 0` means the evicted picture has already been output, which is what makes it evictable at all. NVENC seals it by writing `max_num_ref_frames = 3` ALONGSIDE `max_dec_frame_buffering = 3` — a DPB exactly as deep as its reference count — so the window unmarks the oldest reference in the very unit whose bump drops it. The aliased picture is `ref_idx 2` of a three-entry `num_ref_idx_l0_active` list: addressable by any macroblock, not a spare. **Why two hardware-proven codecs and four GPUs never saw it.** `test-25fps.h264` is level 1.3 with no VUI `bitstream_restriction`, so `dpb_limit` falls back to A.3.1's level ceiling and gives a 7-frame DPB against 2 reference frames — the window unmarks two units before the bump can evict — and it REORDERS, which keeps an unmarked picture alive past the unit that unmarked it. Two independent reasons, both properties of that vector rather than of H.264. It measured zero and passed 250/250 throughout. `data/lowdelay-640x480.h264` is vendored to close exactly that: our own host's output, 120 pictures, goldens from libavcodec cross-checked bit-identical across two ffmpeg builds on two architectures. **The fix is the AV1 fix.** `DecodePlanDxva` and `DecodePlanVk` grow `release_after_decode`, the conversions hand the removals back instead of applying them, and the callers release them once the decode op is issued. It costs no slot the map does not have: `SlotMap::new` allocates `max_dpb_frames + 1` and the DPB never exceeds `max_dpb_frames`, so a free slot always exists with the whole `removed` list still held — measured, peak 4 of 4 on the stream that defers on 117 of 120 units. The Vulkan rung breaks on it in both DPB modes and neither loudly: DISTINCT hands the aliased reference the same array layer the setup writes; COINCIDE clears `slot_image[setup]` in the binding sync and the reference then resolves to no bound image, dropping out of `pReferenceSlots` with a `trace!`. Its deferred release runs on the FAILURE paths too — the fallible region's Result is held rather than `?`-ed, because seven exits sat between the conversion and the release and each would have leaked a slot. `a_full_dpb_bump_reuses_the_slot_but_the_pool_model_binds_a_fresh_image` asserted the aliasing as "the planner's normal behaviour": an authored depth-1 stream whose AU1 references the picture it evicts. It now asserts the opposite, which is the defect in two lines. New evidence, all of it runnable: the CPU proof pins BOTH numbers (0 on the vector, 117 of 120 on the low-delay stream) so neither can drift silently; the ledger-pressure test measures the peak; and a low-delay parity leg is added to `pf-vkdecode`'s `gpu_parity` and `pf-client-core`'s `video_d3d11_native::parity` so both rungs are held to what they stream rather than only to what they conform to. |
||
|
|
5aeb8d2552 |
fix(client): a failed AV1 decode left the surface's facts saying it holds the last picture
The `damaged` path has cleared `Session::held[setup_slot]` since M7, for a reason that now applies to the failure path too: the slot map says the surface holds THIS picture while the surface still carries whatever the previous occupant decoded, so a later `show_existing_frame` naming it blits the old picture's pixels with the old picture's geometry and colour. The failure path never reached that far before — `decode_into`'s error returned straight out of `frame_av1` — and the previous commit made it continue so the slot releases could run. |
||
|
|
3a4c94ad79 |
fix(dxvadec): the review round — a vacuous predicate, an overstated claim, and the H.264 twin of this defect
Five findings from the adversarial pass, all real. **The deferral predicate was vacuous.** `plan.dpb.removed` is ALWAYS a subset of `plan.dpb_refs`: `Av1Planner::plan_frame` snapshots `dpb_refs` before any mutation and `refresh_slots` can only report a picture that was in `self.slots` at that moment. So `filter(|id| dpb_refs.contains(id))` was a condition that is never false, the eager-release loop beside it could never release anything, and the test assertion "only a picture the submission points at earns the reprieve" could never fire. Now: defer every removal, say why in terms of the planner, and assert the PLANNER's property (`removed ⊆ dpb_refs`) — which is falsifiable, and whose failure would mean the conversion is releasing a surface `ref_frame_map` points at. **The failure-path claim was overstated.** Holding the decode's `Result` closes this frame's leak, not the unit's: `decode_av1` returns on the first failing frame and abandons the rest of the temporal unit's plans, so their removals are never released. 24 of 250 units carry a second frame. Named rather than fixed — what to do with the frames after a failure is the pump's question. **⚠⚠ The H.264 leg plausibly has the same defect, and the comment this change added said it could not.** `pic.rs` builds `RefFrameList` from `plan.dpb_refs`, and `H264Planner` snapshots that in `begin_picture` — BEFORE 8.2.5 marking and the DPB bump. The vendored bump drops a picture the sliding window just unmarked once it has been output, so a picture can land in both `RefFrameList` and `dpb.removed`: the AV1 aliasing shape exactly. Measured zero on the vendored vector — but that vector REORDERS, which is precisely what keeps an unmarked picture alive past the AU that unmarked it. A punktfunk host emits LOW-DELAY H.264, where output happens as each picture is decoded, which is the condition that makes eviction and unmarking land in the same access unit. Traced end to end in source, not reproduced (no low-delay vector). NOT fixed: changing a hardware-proven codec on an unreproduced suspicion is the worse risk two commits before a release. Instead `no_au_removes_a_picture_its_own_reference_list_names` makes the assumption falsifiable, and its message says what to do when it fires. HEVC is structurally safe and now says why: `H265Planner` snapshots `dpb_refs` AFTER `decode_rps`. **Four more stale promotion sites**, past the four already fixed: `Backend:: NativeD3d11va`'s variant doc, `Decoder::new`'s Windows rung comment, `lib.rs`'s module note and `clients/session/README.md`. Two sites that used the AV1 leg as the live EXAMPLE of an unproven rung are marked as expired rather than deleted — the reasoning is what the next bad-evidence leg will need. **The AV1 dump was missing.** `PF_DXVA_DUMP` wrote h264 and hevc only, for the one codec whose libavcodec capture has never been taken and where the dump is therefore the only tool. |
||
|
|
af4d265168 |
fix(client): the fourth site that swore the DXVA AV1 leg fails parity, and a clippy lint
`the_evidence_table_says_exactly_which_rungs_have_run_on_hardware` asserts the same fact a third way — a proven list and a NOT-proven list, both spelled out — so promoting the rung in the three places the handoff named still left a test saying "the DXVA AV1 leg FAILS parity on two GPUs — claiming otherwise is the dishonesty this program must not ship". It was right to fail; the pair moves lists here. Three prose sites that still described the leg as decoding wrong pixels move with it: `native_supports_av1`'s device-facts note, `log_rung`'s honesty-surface docs, and the OPEN question in the Windows Intel arm of `pick_native` — that last one is marked CLOSED rather than deleted, because the question it raised (the evidence filter asks "any evidence", and has no answer for BAD evidence) is a real gap in the rule that outlived this particular leg. |
||
|
|
a29e366b3e |
feat(vaapi): VAAPI decodes H.264, H.265 and Main 10 — their first frames on any hardware
The evidence table said these legs "have still never decoded a frame anywhere", and VAAPI is the rung every Linux AMD/Intel client lands on. They have now decoded, on `.25` (Radeon 780M / Phoenix1, RDNA3, radeonsi, Mesa 26.0.3, VA-API 1.23, /dev/dri/renderD128): H.264 250/250 access units accepted, 225 frames delivered, NV12 H.265 250/250 accepted, 204 delivered, NV12 HEVC Main 10 50/50 accepted, 45 delivered, P010 (AV1, unchanged: 250/250 accepted, 250 delivered, NV12) all on the same tiled AMD modifier (0x200000010401b04). Not one access unit of any vector was refused. Three `#[ignore]`d legs modelled on the AV1 one, plus the Annex-B access-unit splitters they need — ported verbatim from `video_d3d11_native`'s test module so the two platform rungs are driven over the same access units rather than over two splitters free to disagree. Main 10 earns a third leg rather than a variation on the second: ten bits is a different VAAPI profile, a different render-target format and a different surface fourcc, and that leg's fourcc assertion is the only thing that would catch a driver quietly handing back NV12 for a ten-bit stream. This is NOT frame-hash parity, and the doc comments say so rather than letting the test names imply it. The Vulkan and D3D11VA legs hash every frame against libavcodec because both can read their decoded surface back; this rung exports a DRM-PRIME dmabuf whose memory the driver tiles, so there is no CPU-readable image to hash without a `vaDeriveImage`/`vaGetImage` path production neither uses nor wants. What these legs prove is that every access unit is accepted, that the expected number of frames comes back, and that each one is a real exported surface of the right shape and fourcc — enough to turn "never decoded a frame anywhere" into a measurement, not enough to promote the rung to `verified`. Two findings the run surfaced, neither of which bites punktfunk's own streams: * The delivered counts are 225/204/45, not 250/250/50, and that is the RUNG, not the driver. `finish` shows `outputs.last()` and never more, so an access unit whose plan bumps several pictures out of the DPB displays the last and drops the rest — 18 dropped at the H.264 vector's three draining IDRs, 45 on the H.265 vector's 45 two-picture bumps — and there is no end-of-stream flush. Hosts emit zero-reorder low-delay output with no B pictures, so `outputs` never holds more than one picture in the field. A CPU-only test derives all three counts from the planner alone, on any Linux box with no GPU, so they stay explanations rather than recordings. * `DmabufFrame::keyframe` labels the ACCESS UNIT, not the picture delivered: `finish` is handed the current AU's `is_idr`. On a reordering stream the IDR is bumped out several access units after it decoded and arrives flagged `false`, while the access unit that drains the DPB at a later IDR flags whichever old picture it displays as a keyframe. That flag is `DecodedImage::is_keyframe`, the pump's post-loss re-anchor signal. Asserted so that fixing it is noticed, not so that it is preserved. Gates, all run on `.25` (this rung only compiles on Linux): `cargo fmt --all -- --check`; `cargo clippy -p pf-client-core --all-targets --features sdl3/build-from-source -- -D warnings`; `cargo test -p pf-client-core --lib --features sdl3/build-from-source` (169 passed); the same filtered to video_vaapi_native with `--include-ignored` (16 passed). Plus the pf-lxcheck2 container's workspace-wide `cargo fmt --all -- --check` and `cargo clippy --workspace --all-targets -- -D warnings`, both clean. The evidence table in `video.rs` still says these legs have never decoded a frame. It is being edited concurrently, so its replacement row is handed over rather than raced for here. |
||
|
|
f4dda9074b |
feat(dxvadec): the AV1 picparams harness AV1 forgot, and the D3D11VA AV1 rung is promoted
Two halves. **The harness.** `libav_picparams_parity` covered H.264 and HEVC only, which is exactly the gap that let a wrong AV1 submission ship. It now plans, converts and packs all 274 frames of the vendored AV1 vector and checks what needs no capture: the three-buffer descriptor set with no quantization matrix (AV1's matrices are selected by index, so `dxva2_av1_end_frame` passes NULL/0 and there is no buffer to submit), no macroblock count anywhere, the 912-byte picture-parameter buffer, and the tile records — which unlike H.264/HEVC slice records do NOT abut, because a `DXVA_Tile_AV1` addresses a tile PAYLOAD and consecutive payloads are separated by their `tile_size_minus_1` fields. The one that matters most is `no_av1_submission_names_its_decode_surface_in_the_ reference_store`: the invariant the previous commit fixed, over the submitted BYTES rather than over the plan. libavcodec cannot produce that shape — it fills `RefFrameMapTextureIndex` from the pre-refresh store and takes `CurrPicTextureIndex` from a frame the reference update has not run on — which is the argument for calling it a defect rather than a convention. `AV1_FIELDS` reaches into the eight nested blocks (`tiles.widths`, `segmentation.feature_data`, …) so a future capture reports a field and not "260 bytes of tiles differ"; `field_table!` grew nested-path support for it. The `#[ignore]`d `our_av1_picture_parameters_match_libavcodecs` and the capture recipe are in place, and `the_dump_and_the_parser_agree…` now self-compares AV1 too. ⚠ NO libavcodec AV1 capture was taken and the module docs say so rather than leaving an absent result to be read as a pass: `.221` has no MSYS2, no gcc and no make, so a patched FFmpeg there is a toolchain bring-up, not a build. Everything this file claims about libavcodec's AV1 side is READ out of `dxva2_av1.c` (n8.1). That reading did turn up one live divergence, recorded at `pic_av1.rs`'s `pp.width` and deliberately NOT changed: libavcodec sends `avctx->width`, which is FrameWidth (pre-superres), where this crate sends UpscaledWidth. The two are equal whenever superres is off, which is every stream that exists here, so the 250/250 result says nothing either way and a blind change would be unmeasured. **The promotion.** `(D3d11va, CODEC_AV1)` is `verified` — 250/250 delivered frames bit-identical to libavcodec on an RTX 3500 Ada AND an Intel Arc. All three places move together: the evidence arm, the module table and `every_rung_runs_and_the_unproven_ones_are_named`, whose `unproven` array loses the pair and whose proven list gains it. ⚠ This changes rung SELECTION, not just a label. `verified` is what lets `auto` pick D3D11VA ahead of Vulkan Video, so Windows Intel and unknown-vendor boxes — where the ladder is `native-d3d11va → native-vk → sw` — now decode AV1 on D3D11VA where they previously fell to Vulkan. Taken deliberately: ~10x the Vulkan leg's speed, and the parity that promoted it was measured on an Intel Arc, which is the vendor family the change moves. Still no soak on the goldens, and the notes say so. Also: `frame_av1` holds the decode's `Result` instead of `?`-ing it, so both slot releases run on the failure path. `decode_av1` notes an error and keeps the session rather than rebuilding the slot map, so an early return leaked a surface per failed frame and hit `SlotError::Full` after nine. |
||
|
|
1c54d0999b |
fix(client): the D3D11VA AV1 rung decoded every inter frame into a surface it was predicting from
AV1 applies `refresh_frame_flags` AFTER the frame is decoded (7.20), so a frame that reads a reference slot and then overwrites it is the ORDINARY case, not an exotic one: 268 of the vendored vector's 274 frames do it, first at frame 6. `plan_to_dxva_av1` released every displaced picture inside the conversion — which is what the H.264 and H.265 siblings do with their whole `removed` list — and then assigned the decode target a slot. `SlotMap::assign` takes the lowest free slot, and the lowest free slot is the one just vacated. So the submission said `CurrPicTextureIndex = N` and `RefFrameMapTextureIndex[k] = N` in the same breath, on 268 of 274 frames: decode into the surface you predict from. Neither vendored H.264 nor H.265 vector ever produces that shape (measured: zero on 250 AUs), which is why an eager release survived two hardware-proven codecs and opened on the first AV1 frame past the key frame's neighbourhood. HEVC even has the invariant under test already — `the_current_picture_is_named_by_curr_pic_and_ never_aliases_a_reference` — and AV1 had nothing. The Vulkan rung already carries the fix; this is the same contract, and the DXVA constraint is the STRICTER of the two: Vulkan binds only the references a frame names, while `RefFrameMapTextureIndex` declares the whole store, so every picture the store still names has to survive the conversion. `DecodePlanDxvaAv1` grows `release_after_decode` and `frame_av1` applies it once the decode op is issued — next to the `refresh_frame_flags == 0` release that already waits for the same reason. Peak surfaces held goes 7 of the 9 the pool allocates, so the spare slot `SlotMap::new` adds is doing exactly the job it exists for. Measured on hardware before the fix: Intel Arc got 245 of 250 delivered frames wrong — 47% of luma at the first bad frame, max |delta| 242, chroma wrong too, a frame predicted from the wrong picture — and the only late frame it got right was the one intra frame, which names no reference and so could not alias. That reads as a `primary_ref_frame` defect and is not one: PRIMARY_REF_NONE and "has no references to alias" are the same frames. |
||
|
|
6d0a389dd2 |
fix(client): the D3D11VA AV1 rung decodes wrong pixels — the parity harness existed all along
The follow-up was framed as "build the frame-hash parity harness the D3D11VA AV1 rung is missing, then flip hardware_verified to true". Both halves were wrong. The harness was never missing. `video_d3d11_native`'s `parity` module has carried `av1_every_delivered_frame_hashes_bit_identical_to_libavcodec` since M7 wired the rung — wired to the SAME libavcodec goldens the Vulkan AV1 leg passes against, with the display-order model that handles the vector's 24 hidden frames, sitting `#[ignore]`d beside the H.264/H.265/Main10 legs. It had simply never been run on a device; .173 was powered off the day it was written. What the old evidence note called a missing harness is real about pf-dxvadec the CRATE, which cannot host one — it links no D3D11 — but the device half lives here and was already done. Run on .221, it FAILS, on both GPUs, deterministically (three runs each, identical first-divergent frame and identical hashes): 186/250 diverging display frames on an RTX 3500 Ada, 245/250 on an Intel Arc. It is the decode that is wrong, not the measurement, and three independent checks say so. H.264 and H.265 pass 250/250 and HEVC Main 10 50/50 through the same harness, the same readback geometry, the same crop and the same slot map on those same two GPUs. pf-vkdecode's Vulkan AV1 leg reproduces the same golden file 250/250 on the same box. And the goldens regenerate byte-for-byte from the ffmpeg build their own header names. Two signatures, and they are not one defect wearing two faces. NVIDIA is bit-exact for display frames 0..=63 and then loses ONE 16x24 luma block — 174 pixels, max |delta| 8, chroma untouched — on the frame whose order_hint first reaches 64, after which every remaining frame is downstream of it through prediction. The stream parks the key frame (order_hint 0) in BWDREF and ALTREF2 for its whole length, so 64 is where the distance to it reaches the edge of what get_relative_dist can represent at OrderHintBits = 7. Intel is structurally wrong from display frame 4 — 47% of luma, max |delta| 242, chroma wrong too, a frame predicted from the wrong picture — and the only later frame it gets right is the one whose primary_ref_frame is PRIMARY_REF_NONE. None of this is visible on glass, which is the whole argument for goldens: the rung streams 4K60 on both parts with a clean five-minute soak at roughly ten times the Vulkan leg's decode time. The 2026-08-07 field sessions that looked clean were looking at wrong pixels. So hardware_verified stays false, and the note now says why in the strongest available terms — it prints at warn on every session that lands here, and "decodes AV1 to wrong pixels" is what a support engineer needs to read. The pair stays in `every_rung_runs_and_the_unproven_ones_are_named`'s unproven array; its note still contains NEVER, because the pair has never PASSED parity, which is now a measured statement rather than an absence. Left deliberately unchanged: `auto` on Windows can still reach this rung for AV1, and on Intel it is the arm that fires, because that vendor advertises no SAMPLED usage on any decode profile so zero-copy Vulkan Video cannot run there. Barring it trades visibly-wrong AV1 for the software rung, which cannot keep up at 4K and is itself unproven. Which way that trade goes is a product call, so it is recorded at the admission site rather than made silently here. `av1_divergence_map` is kept, cleaned up and documented: it is what turned "186 frames differ" into a lead — one line per display frame, its verdict beside the plan facts that could explain it, and an opt-in raw-NV12 dump. At a frame where one vendor hashes correctly, that vendor's bytes ARE libavcodec's bytes and so a valid reference for the other's, which is how "how badly" was answered without new goldens. The tool that would localise the rest does not exist: pf-dxvadec's libav_picparams_parity covers H.264 and HEVC only, so the AV1 conversion has never been compared against libavcodec at the picture-parameter level either. That is the next step, not another session. Also in this file, since it is the same table and the same day: the VAAPI rung's AV1 leg has now decoded 250/250 of the vendored vector on RDNA3 and its arm is split from the H.264/H.265 ones, which genuinely have still never decoded anything. It is unverified for the same reason as ever — no parity — and the D3D11VA row above is exactly why that distinction is worth keeping: a rung can decode 250 frames and still be wrong. |
||
|
|
f351eb01e9 |
feat(vaapi): VAAPI decodes AV1 — the rung's first frame on any hardware
The evidence table has said "native VAAPI: has never decoded a frame anywhere
(M6/M7)" since the rung was written. That is no longer true. Measured on `.25`
(Radeon 780M / Phoenix1 RDNA3, radeonsi, Mesa 26.0.3, VA-API 1.23, Ubuntu
26.04 — headless, no display server needed):
VAAPI AV1 rung constructed: native-vaapi av1
VAAPI AV1: 250 frames delivered, first 320x240 fourcc="NV12"
modifier=0x200000010401b04
250 of 250 displayed frames, first try, on the same vendored vector the Vulkan
and D3D11VA AV1 legs walk. The count matters as more than a smoke test: the
vector carries 274 coded frames in 250 temporal units — 24 units carry two, and
those extras are HIDDEN (decoded, referenced, never shown) — so 250 delivered is
this rung agreeing with the other two about which frames are output. A tiled AMD
DRM modifier rather than a linear one says the surface is a real decode target,
not a fallback.
Two changes, both in the rung's own file.
**The probe never asked about AV1.** `probe_this_machines_libva` walked H.264
High, HEVC Main and HEVC Main 10 and stopped there, which is part of why "never
decoded a frame" could stand so long without anyone noticing what had not been
asked. It now covers both AV1 profiles, and this box answers:
H.264 High: VLD decode AV1 Profile 0: VLD decode
HEVC Main: VLD decode AV1 Profile 1: no (VAProfile not supported)
Profile 1 being refused is correct — 4:4:4 AV1, which radeonsi does not do — and
it is the negative case that proves the probe reports rather than assumes.
**`av1_decodes_the_vendored_vector_on_this_machines_vaapi`** is the decode
itself, `#[ignore]`d beside the probe.
It is deliberately WEAKER than the Vulkan and D3D11VA AV1 legs, and the docs say
so rather than letting the name imply parity: those two hash every frame against
libavcodec's goldens because both can read their decoded surface back. This rung
hands out a DRM-PRIME dmabuf whose memory the driver tiles, so there is no
CPU-readable image to hash without adding a vaDeriveImage/vaGetImage path that
production neither uses nor wants. So it asserts what can be asserted honestly —
every temporal unit accepted, the right number of frames back, each a real
exported surface of the right shape, the first flagged as a keyframe — and it is
NOT frame-hash parity. Promoting this rung to `verified` still wants parity, and
parity wants a readback path first.
It fails loudly rather than skipping when the device has no AV1 entry point. It
is `#[ignore]`d, so it only runs when someone points it at a box that is supposed
to have one, and a silent pass there is exactly the invisible-failure mode this
program exists to end.
Gates: on `.25`, fmt clean, `clippy -p pf-client-core --all-targets -D warnings`
green under the Linux cfg where this rung actually compiles, the whole lib suite
167/167, and all 11 VAAPI tests green with `--include-ignored`. Workspace fmt +
clippy + lib suite also green in the Linux container.
⚠ Not touched here on purpose: the evidence table in `video.rs`. Its VAAPI row
still reads "never decoded a frame anywhere" and now understates what is known —
but a parallel agent is editing that same file for the D3D11VA AV1 row, so the
row is left for whoever lands second to update once, rather than conflicting.
Note for anyone reproducing on `.25`: it has no system SDL3 and no passwordless
sudo, so the test binary links only with `--features sdl3/build-from-source`
(SDL3 is gamepads, irrelevant to decode; production Linux still links the system
one). Its disk sits at ~99% full, and the tree there is a `git archive` export
with no `.git`, so `git apply`/`git checkout --` silently do nothing.
|
||
|
|
5a4305c072 |
merge: bring current main into the gyro correctness branch
main moved ~60 commits while this branch was in progress, and one of them matters here: PR #88 (the phone-gyro mirror) landed, touching the same motion path. One conflicted file, `GamepadCapture.swift`, in three places — all of them the two changes meeting rather than disagreeing: - **Slot fields.** #88 added `motionSent` + `lastAccel` for its flush-parks-motion fix; this branch removed `lastMotionNs` with the 4 ms drop-throttle. Kept both decisions: the parking state stays, the throttle field goes. - **forwardMotion's head.** #88 added the mirror stand-down (`pad 0` yields while the phone speaks for it); this branch deleted the throttle guard. Kept the stand-down, dropped the guard. - **The send.** This branch converts into the DualSense report frame; #88 records what went out so `flush` can replay it beside a zero gyro. Both, with the recording placed AFTER the conversion — `flush` replays `lastAccel`, so it has to be the vector that actually went on the wire, or a still pad's gravity gets parked in the wrong axis. The two features compose exactly, which is worth stating because it is not luck: this branch gates motion capture on `hasRotationRate`, and #88 engages the phone mirror when `hasRotationRate != true`. They are complements — a pad either drives its own gyro or the phone mirrors for it, never both and never neither. Everything else auto-merged. Note `DeviceGyroRemapTests` is `#if os(iOS)`, so the macOS suite reports the same 215 as before the merge rather than gaining #88's six — checked, not assumed. Gates re-run against the merged tree rather than trusting either side's: Linux fmt + build + `clippy --locked --all-targets -D warnings` + punktfunk-core and pf-inject suites; Apple 215 tests and the iOS-triple typecheck; Android kit + app compile and tests. All green. |
||
|
|
19c9165d4b |
docs(client): the D3D11VA AV1 rung has two vendors and a soak now — and still no parity
Re-measured against a host carrying #95, from .21 (RTX 5070 Ti, av1_nvenc) to .221, on glass: Intel Arc, auto -> native-d3d11va 4K60, decode 1.4 ms, e2e 16.7 ms p50 RTX 3500 Ada, pinned native-d3d11va 4K60, decode 1.0 ms RTX 3500 Ada, pinned native-vulkan 4K60, decode 11.6-16.7 ms Plus a 5-minute Arc soak: 297 stats lines, 60 fps, decode 1.3 ms, e2e 10.9/14.8 ms p50, and exactly one WARN in the whole run — the hardware_verified=false notice itself. No refusals, no demotions, no concealed runs. Three things that follow. The rung is no longer a one-session curiosity: it decodes 4K60 AV1 on TWO vendors and survives a soak. The Arc leg matters twice over, because the Arc advertises no SAMPLED usage on any decode profile — zero-copy Vulkan Video cannot work there — so `auto` demoting to D3D11VA and then decoding is the whole demotion path working as designed. It is roughly 10x faster than the Vulkan AV1 leg on the SAME NVIDIA GPU. That is the strongest argument yet for eventually letting `auto` pick it ahead of Vulkan Video, which is exactly what `verified` gates. And it stays `verified = false` anyway, because the missing piece is specific: there is no frame-hash parity against libavcodec. Every other verified pair in that table earned it with one, and pf-dxvadec has no harness that could produce one — `libav_picparams_parity` compares picture parameters on the CPU and never decodes a frame. Building that harness is the work that promotes this rung; a fourth session is not. The evidence string now says so, so the next reader does not have to rediscover which half is missing. The VAAPI row is corrected in the same spirit rather than left as a bare "NO": the reachable VAAPI box (.25, RDNA3) reports VAProfileAV1Profile0 / VAEntrypointVLD and advertises no Vulkan AV1 decode at all, which makes it the right box to prove that rung on and an unambiguous oracle when it happens. What stopped it is recorded too — no punktfunk checkout there and 4 GB of usable RAM. Documentation only — no behaviour change, and no flag flipped. |
||
|
|
c64cdc4ef7 |
docs(encode): close out the tile-aware AV1 sub-frame reader — measured, not worth it
#95 disarmed sub-frame readback for AV1, which means AV1 forgoes the latency win HEVC gets from shipping slice 1 while slice 2 encodes. The follow-up was to teach the reader AV1's units: cut on OBU boundaries rather than byte counts and arm from the driver's reported unit count. Measured on .21 (RTX 5070 Ti, av1_nvenc) before writing any of it, and the measurement closes it rather than scoping it. Reading the frame headers av1_nvenc actually emits at 4K: width_in_sbs_minus_1[0] = 59 one tile column, the full 3840 height_in_sbs_minus_1[0..1] = 16, 16 two tile rows tile_start_and_end_present_flag = 0 BOTH TILES IN ONE TILE GROUP OBU That last flag is the finding. "Cut on OBU boundaries" presumes the tiles are separate OBUs and they are not — there is no boundary between them to cut on. Shipping tile 1 early would need the HOST to re-author AV1 syntax per chunk, synthesising a fresh Tile Group OBU header with tile_start_and_end_present_flag = 1 and its own tg_start/tg_end. That is bitstream surgery on the encode path, not the reader change it was assumed to be. And the prize would be small even then, because split encode already spent it. The two tile rows go to two split-encode engines that run CONCURRENTLY, so they complete at nearly the same moment — the win is bounded by the skew between engines, not by half a frame. Whole-frame encode measures 3.3-3.6 ms at 4K60 against a 16.7 ms p50 end-to-end, so even the sequential-tiles fantasy caps near 1.7 ms and the real number is a fraction of it. HEVC's win is bigger for a structural reason that does not transfer: forced split and sub-frame are mutually unsupported, so HEVC's slices genuinely are produced one after another. 1080p settles it further: tile_cols_log2 = tile_rows_log2 = 0, a single tile, so there is nothing to pipeline at the commonest streaming resolution at all. Recorded next to the disarm with the reopen condition named — NVENC emitting one OBU per tile, or setting tile_start_and_end_present_flag = 1 — so this is closed on evidence rather than left as an open maybe. Documentation only — no behaviour change. |
||
|
|
6b4be28d24 |
docs(client): write down why the CPU rung is not process-isolated
#97's frame-context floor closes the one rav1d abort we hit and can prove. It does not make the rung panic-proof and nothing at that call site can, because rav1d's public surface is dav1d's C ABI: any reachable panic crosses `extern "C"` as `panic_cannot_unwind` and becomes `abort()`, past every `catch_unwind`, rung demotion and typed refusal we have. Counted across rav1d 1.1.0's 60 source files: 285 `unwrap()`, 214 `assert!`, 19 `unreachable!`, 11 `expect()`, 10 `panic!`. 539 sites that end the client if a stream can reach them. #97 fixed one of them. Process isolation is the only defence that actually works, and this records the decision NOT to build it, with the reasoning, so it is not re-argued from scratch each time someone reads that number: * the defect is upstream's and is one line (memorysafety/rav1d#1497, filed 2026-08-07 with the fix and a reproducer; still open, no PR, as of today); * 539 is an unbounded number, not a risk estimate — none of those sites is known reachable from a punktfunk stream, and the honest next step is to fuzz the rung and find out, which is cheap, rather than buy insurance, which is not; * the cost lands on the video path across Linux, Windows and Android (the Apple clients decode through VideoToolbox and never reach this code), each needing its own shared-memory frame transport, child lifecycle and backpressure, and it adds a scheduling boundary to the slowest rung on the ladder while zero-copy is a hard requirement; * an abort here costs a session that was already degraded — this rung exists because the GPU rungs failed first. The trigger to revisit is named as an event rather than a feeling: a SECOND distinct abort in the field, or a fuzzer finding a reachable panic. Either makes it a class of bugs instead of one, and a class is what would justify the architecture. Documentation only — no behaviour change. |
||
|
|
669176982d |
fix(h264): name the DPB cliff #96 left standing in the other codec
H.264 derives its DPB size the same way HEVC did before #96 — from a level ceiling that says what a stream MAY use, not what it needs — and the ceiling saturates at 16 frames, which is 17 hardware slots with the picture in flight. That is the exact arithmetic that cost 720p and 1080p their HEVC. Measured on real encoders (2026-08-07) rather than assumed: H.264 escapes it twice over, and both escapes belong to the encoders, not to the format. encoder level picked VUI restriction NVENC (RTX 5070 Ti, 610.57.04) 3.2/4.2/5.1/5.2 present, buffering 3 VAAPI via libavcodec (RDNA3, 26.0.3) 4.1/4.2/5.1/5.2 present, buffering 1 openh264 (the software rung) 3.2/4.2/5.1/5.2 present, buffering 1 Every one picks a level proportionate to the picture AND states its real need in the VUI bitstream restriction, so the ceiling is never reached and never consulted. Nothing is broken today, and clamping would be wrong: with the restriction present the number IS the stream's own statement, and a stream that genuinely asked for a deep DPB would decode wrong if we shrank it. So this does not change what any stream decodes. It gives the arithmetic one named home (`dpb_limit`, the twin of `h265::dpb_limit`) carrying the evidence and the reasoning, and it adds the signal that was missing: when an SPS carries no restriction AND its level ceiling would demand more slots than mainstream hardware provides, the plan now says so with `PlanWarning::LevelDerivedDpb` instead of a user silently losing the codec the way #96's users silently lost HEVC. It is not an integrity warning — the picture is intact; what fails is opening a session — so `is_integrity_warning` classifies it false. One thing the sweep corrects about how the follow-up was framed: it is SMALL pictures that saturate the ceiling most easily, not 720p specifically. 640x360 at level 3.1 computes 16 as readily as 720p at level 5.0, because the ceiling is MaxDpbMbs divided by the picture's macroblocks. The authored 64x64 test fixtures land there too, which is why they now assert through `picture_warnings`. Guards, as the missing consumer-end half of pf-encode's `rfi_dpb_fits_a_mainstream_vulkan_decoder`: * every_reachable_h264_stream_fits_a_mainstream_slot_pool — the measured (picture, level, declaration) pairs, asserting slots <= 16 * the_level_ceiling_alone_would_reproduce_96_and_is_warned_about — the same resolutions at levels that saturate, pinned WITH the warning * a_proportionate_level_fits_even_without_a_vui_restriction — so neither escape looks like it is doing all the work alone Gates: fmt + clippy -D warnings clean; pf-client-core 167/167; pf-bitstream 84/84; and gpu_parity 8/8 bit-identical to libavcodec on the RTX 5070 Ti, which is the gate that matters for anything touching the bitstream layer. |
||
|
|
d996449a82 |
fix(host/pads): a virtual pad at rest said it was in free fall
G14, unblocked by the frame measurement in
|