Compare commits

...

97 Commits

Author SHA1 Message Date
enricobuehler aacd610b68 fix(steamdeck/install): set up controller passthrough + surface the required reboot
ci / docs-site (push) Successful in 1m12s
ci / web (push) Successful in 1m13s
apple / swift (push) Successful in 1m18s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 9s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 14s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 51s
ci / bench (push) Successful in 6m35s
apple / screenshots (push) Successful in 6m13s
deb / build-publish (push) Successful in 10m21s
docker / deploy-docs (push) Successful in 34s
deb / build-publish-host (push) Successful in 9m58s
android / android (push) Successful in 17m7s
arch / build-publish (push) Successful in 17m47s
ci / rust (push) Successful in 19m28s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m59s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m2s
A fresh SteamOS host install failed on-glass two ways: KWin denied Desktop-mode
capture (zkde_screencast_unstable_v1) so every session died in the pipeline build,
and the native Steam Deck pad degraded to an Xbox 360 controller. Both stem from
setup that only a fresh login applies — the host user-service keeps the
pre-install group set and KWin reads its .desktop grant at session start — plus a
missing vhci-hcd autoload the pad's USB/IP transport needs (the deb/rpm/arch
packages ship it; the Deck installer did not).

- install.sh: install punktfunk-modules.conf (+ modprobe vhci-hcd now) for the
  native Deck controller transport, seed the KDE RemoteDesktop grant
  (kde-authorized) for Desktop-mode input, and print a loud "reboot before
  streaming" notice when a relogin is actually required (input group added or the
  .desktop grant first installed).
- update.sh: retrofit the udev rule, vhci-hcd autoload, input group, and grant so
  existing installs pick them up on the next update without a reinstall.
- steamos-host.md: document the first-install reboot and the native controller
  passthrough requirements (input group + vhci-hcd + client-side Steam Input Off).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 19:22:08 +02:00
enricobuehler 081ff64087 docs: explain how to actually run the VirtualHere passthrough example
apple / swift (push) Successful in 1m22s
ci / web (push) Successful in 1m11s
ci / docs-site (push) Successful in 1m11s
ci / bench (push) Successful in 5m39s
apple / screenshots (push) Successful in 6m38s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 9s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 57s
deb / build-publish (push) Successful in 9m0s
deb / build-publish-host (push) Successful in 9m20s
windows-host / package (push) Successful in 9m46s
android / android (push) Successful in 18m37s
arch / build-publish (push) Successful in 18m48s
docker / deploy-docs (push) Successful in 28s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m44s
ci / rust (push) Successful in 26m30s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m2s
The VirtualHere DualSense recipe linked the `virtualhere-dualsense.ts` SDK
example as a deployable recipe but skipped everything needed to run it: that
VirtualHere is a server (couch) + client (host) pair, and that the example —
like every example — imports `../src/index.js`, which only resolves inside the
SDK repo. A user copying it out had no way to know they must
`bun add @punktfunk/host` and swap that import.

- automation.md: rewrite the recipe into a full walkthrough — the two-sided
  VirtualHere setup, the `-t` verbs, the zero-code hooks version, and a new
  scripted section with exact deploy steps, env vars, and running it as a
  service (its own SIGTERM handler makes `systemctl stop` release the pad).
- sdk/README.md: say how to run an example in-repo vs deployed on a host.
- virtualhere-dualsense.ts: header note on the import swap + service setup,
  since that file is where the doc link lands.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 17:28:23 +02:00
enricobuehler 675fd24cce chore(release): bump workspace version to 0.17.1
audit / cargo-audit (push) Successful in 3m22s
audit / bun-audit (push) Failing after 13s
apple / swift (push) Successful in 1m23s
ci / web (push) Successful in 55s
ci / docs-site (push) Successful in 1m11s
ci / bench (push) Successful in 5m12s
apple / screenshots (push) Successful in 6m55s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 4m55s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 5m39s
ci / rust (push) Successful in 24m5s
decky / build-publish (push) Successful in 30s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 19s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 17s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 40s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
android-screenshots / screenshots (push) Successful in 3m34s
linux-client-screenshots / screenshots (push) Successful in 8m32s
release / apple (push) Successful in 11m16s
deb / build-publish-host (push) Successful in 12m51s
deb / build-publish (push) Successful in 12m57s
docker / deploy-docs (push) Successful in 27s
web-screenshots / screenshots (push) Successful in 3m13s
windows-host / package (push) Successful in 14m6s
android / android (push) Successful in 15m33s
arch / build-publish (push) Successful in 16m3s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 2m43s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m40s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m10s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 23m16s
flatpak / build-publish (push) Successful in 7m9s
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 16:37:32 +02:00
enricobuehler 0e60e58cab perf(apple): windowed macOS presents — off-main transactional commit, surface prototype, safe-present setting
ci / docs-site (push) Successful in 52s
ci / web (push) Successful in 56s
decky / build-publish (push) Successful in 21s
apple / swift (push) Successful in 1m24s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
ci / bench (push) Successful in 5m53s
release / apple (push) Successful in 9m19s
deb / build-publish (push) Successful in 12m27s
arch / build-publish (push) Successful in 12m52s
docker / deploy-docs (push) Successful in 26s
deb / build-publish-host (push) Successful in 13m18s
apple / screenshots (push) Successful in 6m20s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m51s
android / android (push) Successful in 18m40s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m15s
ci / rust (push) Successful in 27m38s
Task 1 (latency): the transactional present now commits from the RENDER
thread inside an explicit CATransaction + flush() instead of hopping to
main. The present harness (2026-07-21, 240 Hz Studio, full-size window)
measured the main hop batching presents at runloop-iteration rate when
main is busy — an active implicit transaction NESTS the explicit one —
which is the field's presents=55 @ fps=240 / display_p50 18.6 ms.
Off-main commits measured immune to main churn: glass p50 ~10 ms,
cadence a clean 4.17 ms at 240. flush() is load-bearing: without it the
render thread's own implicit transaction (drawableSize/colour writes)
swallows the explicit commit and NOTHING reaches glass — the harness
reproduced the exact freeze (every present dropped). Legacy main-hop
kept as PUNKTFUNK_TXN_PRESENT=main for field A/B. New pf-windowed
Console line (subsystem io.unom.punktfunk, category presenter)
decomposes the issue side per second.

Task 1 lever D: the f407f418 IOSurface contents-swap path is
resurrected format-aware as WindowedPresentMode.surface — bgra8 SDR /
rgba16Float+PQ-tagged HDR pool, swap committed off-main from the
completion handler; EDR anchored by the metal layer underneath.
Env-only prototype (PUNKTFUNK_WINDOWED_PRESENT=surface) until the HDR
composite is eyeballed on glass.

Task 2 (setting): punktfunk.windowedSafePresent (default ON =
transactional mitigation; OFF = the fast async path whose panic the
saga is about — the caption says so in plain words). Rows in the touch
Settings (Presentation) and gamepad settings, macOS-only.
PUNKTFUNK_WINDOWED_PRESENT=async|transaction|surface overrides for dev
A/B. maximumDrawableCount stays 3 — the harness measured 2 starving the
render loop (172/s) and worsening glass p50.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 15:58:47 +02:00
enricobuehler 71faf38a85 fix(apple): stats os_log used %d for 64-bit Swift Int — macOS 26 drops the line
Swift `Int` is 64-bit (Foundation infers %lld); the stats mirror's format
string used %d (32-bit C int) for fps/presents/lost. macOS 26's strict
String(format:) validator rejects the mismatch and drops the whole line
(the error also cascades onto the float args, mis-blaming floor_p50). Use
%lld for the three Int fields. Pre-existing, unrelated to the DCP work;
surfaced while reading presenter logs during the panic fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 15:14:41 +02:00
enricobuehler f49d38826b fix(apple): windowed macOS DCP panic via presentsWithTransaction — keep real HDR, drop the SDR IOSurface path
The windowed-macOS "mismatched swapID's" @UnifiedPipeline.cpp kernel
panic is the CAMetalLayer ASYNC image queue (commandBuffer.present,
mandatory with displaySyncEnabled=false) diverging from WindowServer's
compositor on a high-refresh composited session — it survived glass
pacing (PyroWave, 2026-07-18) and hit windowed HEVC on the same 240 Hz
Mac Studio (2026-07-21), so it's the async queue itself, at any codec
or pacing.

f407f418's mitigation routed windowed PyroWave through a BGRA8 IOSurface
layer.contents pool, which cannot carry HDR — extending it to HEVC would
have tone-mapped windowed HDR down to SDR. Instead, present the drawable
through a Core Animation transaction (CAMetalLayer.presentsWithTransaction)
for every windowed session: commit, waitUntilScheduled, then drawable.present()
inside a CATransaction, so the swap commits with the layer tree and stays
in lockstep with the compositor instead of racing it. The full
rgba16Float / PQ / EDR render path is untouched — windowed HDR is real
HDR again. Fullscreen keeps the async path (direct scanout, lowest latency,
no panic there).

Removes the IOSurface surface pool, its surfaceLayer sibling, and the
macOS windowed PQ->SDR tone-map (tvOS keeps its no-headroom tone-map).
Net -150 lines. swift build + 169 tests green. Needs on-glass soak on
the 240 Hz Mac Studio (kernel race — only real hardware confirms).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 14:45:11 +02:00
enricobuehler 42198eb1b6 feat(linux/vulkan-encode): default-on native NV12 capture + RGB true-extent
ci / web (push) Successful in 48s
ci / docs-site (push) Successful in 1m5s
apple / swift (push) Successful in 1m20s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 25s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 42s
ci / bench (push) Successful in 5m59s
apple / screenshots (push) Successful in 6m17s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 6m29s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 5m12s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m16s
deb / build-publish (push) Successful in 12m58s
deb / build-publish-host (push) Successful in 13m40s
windows-host / package (push) Successful in 10m46s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m52s
android / android (push) Successful in 16m18s
arch / build-publish (push) Successful in 16m20s
docker / deploy-docs (push) Successful in 30s
ci / rust (push) Successful in 19m25s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 14m19s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m19s
Flip both zero-copy levers from opt-in to default:

- The capture negotiation now PREFERS gamescope's producer-side NV12 pod by
  default (PUNKTFUNK_PIPEWIRE_NV12=0 restores the packed-RGB negotiation).
  Codec-aware gating rides a new OutputFormat::nv12_native ->
  ZeroCopyPolicy::native_nv12_session edge resolved by the host facade from
  the session plan (pf_encode::linux_native_nv12_ok): only H265/AV1 sessions
  whose backend can open the raw Vulkan Video encoder ever see the NV12 pod
  -- an H264/Moonlight session (libav VAAPI, which would misread the
  two-plane buffer) keeps today's BGRx negotiation, as do the
  GameStream-resolve and portal-mirror paths, PyroWave, and NVENC prefs.
- RGB-direct's unaligned modes default to the true-extent direct import
  (PUNKTFUNK_VULKAN_RGB_TRUE_EXTENT=0 restores the padded-copy staging).
  Guarded-tested on Van Gogh with the kernel journal watched: clean, and at
  5.38 ms p50 the fastest 1080p encode path measured on that hardware. The
  EFC only exists on Mesa >= 26, where the codedExtent-driven session_init
  padding is guaranteed (verified back to Mesa 24.2).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 14:01:28 +02:00
enricobuehler b4fbc94e36 feat(linux/vulkan-encode): PUNKTFUNK_VULKAN_RGB_TRUE_EXTENT bring-up gate
Opt-in alternative to RGB-direct's padded-copy staging at unaligned modes:
direct-import the visible-size capture and declare the TRUE-SIZE source
codedExtent. RADV has derived the VCN session_init firmware padding from
srcPictureResource.codedExtent since Mesa 24.2, so the EFC is told the source
lacks the alignment rows and the hardware edge-extends them internally --
which also reframes the 2026-07-20 field GPU reset: that crash passed the
ALIGNED extent as the source codedExtent (firmware padding zero) over a
visible-size buffer. Off by default until a guarded live test proves the EFC
front-end honors the padding like the YUV fetch path does; the padded-copy
staging remains the shipped behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 14:01:28 +02:00
enricobuehler 004f0cacb8 perf(linux/vulkan-encode): true-size headers make native NV12 zero-copy at every mode
Drop the padded-copy staging for producer-native NV12 and direct-import the
visible-size buffer at unaligned modes (1080p) too. Safe by the driver's own
contract rather than by allocation luck: native sessions now author the H265
SPS / AV1 sequence header at the RENDER size and pass the matching codedExtent
on every picture resource. RADV rounds the bitstream SPS up itself (with a
conformance window -- radv_video_patch_encode_session_parameters, per the
VK_KHR_video_encode_h265 proposal's "implementations may override" clause) and
programs the VCN session_init with the true extent plus nonzero firmware
padding, so the hardware edge-extends the alignment rows internally and never
fetches past the source's real extent -- the exact mechanism VAAPI/radeonsi has
always used to encode 1080p. The driver-emitted header NALs already carry the
patched SPS, so the wire format is unchanged.

The CSC and RGB-direct paths keep the app-aligned convention (their sources
genuinely cover the aligned extent); RGB padded-copy staging is untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 14:01:28 +02:00
enricobuehler 49c6dd00c9 fix(presenter): CPU frames never take the HDR10 passthrough surface
audit / bun-audit (push) Failing after 15s
ci / docs-site (push) Successful in 56s
ci / web (push) Successful in 58s
apple / swift (push) Successful in 1m22s
audit / cargo-audit (push) Successful in 2m20s
decky / build-publish (push) Successful in 32s
ci / bench (push) Successful in 7m0s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 15s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 14s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 7m18s
release / apple (push) Successful in 9m28s
deb / build-publish-host (push) Successful in 10m32s
deb / build-publish (push) Successful in 11m48s
docker / deploy-docs (push) Successful in 23s
android / android (push) Successful in 14m36s
arch / build-publish (push) Successful in 15m4s
windows-host / package (push) Successful in 16m11s
flatpak / build-publish (push) Failing after 8m19s
apple / screenshots (push) Successful in 6m40s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m11s
ci / rust (push) Successful in 19m38s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m22s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m29s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m39s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 4m39s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 5m36s
Software decode uploads swscale RGBA with no CSC/tonemap pass. On the new
desktop-compositor HDR10 surfaces (GNOME 48 / Plasma 6 + Mesa >= 25.1 offer
ST2084 even on SDR desktops) that sRGB-encoded content was composed as PQ —
the field-reported psychedelic picture (Fedora clients decode HEVC in
software because stock Mesa strips it; reproduced + fix live-validated on
GNOME Wayland: mode-0 soup -> HDR→SDR washed-out-but-correct).

The CPU lane keeps the SDR swapchain and its known untonemapped-PQ gap; a
real PQ→sRGB pass for CPU frames is the follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:30:35 +02:00
enricobuehler e7c6c96e5f fix(presenter): set the Wayland app_id so the session window gets its icon
SDL_APP_ID = io.unom.Punktfunk — compositors match it against the shipped
.desktop for the window icon; without it KWin/mutter show the generic Wayland
icon (field-noticed on the Deck). Linux analog of win32 AppUserModelID.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:30:35 +02:00
enricobuehler da134ad045 feat(presenter): PUNKTFUNK_HDR10=0 refuses the HDR10 swapchain (SDR tonemap pin)
GNOME 48 / Plasma 6 with Mesa >= 25.1 now offer HDR10_ST2084 surfaces even on
SDR desktops, silently flipping PQ streams into the mode-0 passthrough lane —
previously unreachable on desktop compositors and never validated there. The
hatch pins the shader tonemap for diagnosis and as the field escape.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:30:35 +02:00
enricobuehler 428bd2519c test(qsv): converter→ring-profile-P010→encoder live e2e (the RTV-written seam)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:30:35 +02:00
enricobuehler 87d7eceabf fix(encode/qsv): write the colour VUI unconditionally + Intel-lane colour diagnostics
The native QSV encoder only attached mfxExtVideoSignalInfo on HDR sessions —
an SDR stream carried no colour description at all (NVENC writes 709-limited
unconditionally; qsv.rs now mirrors it).

Diagnostics for the field-reported blue/magenta Intel-host colours:
- hdr-p010-selftest takes an optional WxH + GPU vendor (intel|nvidia|amd) and
  prints which adapter it tested — 64x64 on the default GPU proved nothing on
  dual-GPU boxes, and the field heights (1080/1400) are not 16-aligned.
- pf-capture: ignored live test running the converter at 1920x1080 pinned to
  the Intel adapter.
- pf-encode: qsv_live_p010_1080_colorbars_dump — known P010 bars through the
  UNALIGNED-height ingest copy (1080 src -> align16 1088 pool, the seam no
  640x480 test exercises) to Main10 HEVC, dumped for off-box decode checks.
- dxgi.rs: drop the stale 'falls back to the R10 path' claim (no such path).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 13:30:35 +02:00
enricobuehler 1d4795666e feat(linux/vulkan-encode): opt-in gamescope producer-native NV12 encode source
ci / web (push) Successful in 54s
ci / docs-site (push) Successful in 57s
decky / build-publish (push) Successful in 19s
apple / swift (push) Successful in 1m21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 6m53s
apple / screenshots (push) Successful in 6m16s
windows-host / package (push) Successful in 9m38s
deb / build-publish (push) Successful in 9m55s
docker / deploy-docs (push) Successful in 26s
arch / build-publish (push) Successful in 12m52s
android / android (push) Successful in 12m58s
deb / build-publish-host (push) Successful in 14m9s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 23m0s
ci / rust (push) Successful in 28m32s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 21m41s
With PUNKTFUNK_PIPEWIRE_NV12=1 (bring-up gate), the PipeWire negotiation
offers an NV12 LINEAR DMA-BUF pod (BT.709 limited pinned MANDATORY) ahead
of the BGRx one, and gamescope's producer-side RGB->YUV pass replaces the
host CSC entirely: the encoder imports the two-plane buffer as a profiled
VIDEO_ENCODE_SRC image and the VCN encodes it directly. Contributed
measurements: encode p99 2.9 ms (from ~4-4.5 ms via EFC RGB-direct),
60 fps capture, 0 send drops.

Hardening on top of the contributed patch:
- unaligned modes (1080p!) stage through a padded aligned NV12 copy (edge
  rows/columns duplicated, transfer-only) instead of direct-importing the
  visible-size buffer -- a direct import would make the VCN read past the
  producer allocation, the exact OOB class behind the 2026-07-20 field GPU
  reset; the encode extents return to the aligned coded extent everywhere
- the UV plane layout honors the producer's plane-1 chunk (offset/stride)
  when the SPA buffer carries one (same-BO verified by inode), with the
  contiguous-plane contract as fallback
- PyroWave sessions are excluded from the gate (their Vulkan compute CSC
  ingests packed RGB), and a native-NV12 session that resolves to libav
  VAAPI (H264 codec, PUNKTFUNK_VULKAN_ENCODE=0, feature off, or a failed
  Vulkan open) refuses at open instead of streaming garbage chroma
- pad staging images carry TRANSFER_SRC (the width-padding pass self-copies
  the staging image -- previously missing on 1366-wide modes)
- metadata-cursor one-shot warn (parity with RGB-direct) and a padded-NV12
  PUNKTFUNK_PERF split label; the padded RGB copy gets its timestamps too

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 11:57:42 +02:00
enricobuehler 58b27acbd5 fix(vdisplay/session): scrub the dead desktop's WAYLAND_DISPLAY from the user-manager env
ci / web (push) Successful in 1m2s
ci / docs-site (push) Successful in 1m14s
apple / swift (push) Successful in 1m22s
decky / build-publish (push) Successful in 27s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 19s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 17s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
ci / bench (push) Successful in 5m44s
apple / screenshots (push) Successful in 6m18s
android / android (push) Successful in 13m0s
docker / deploy-docs (push) Successful in 27s
deb / build-publish (push) Successful in 13m36s
deb / build-publish-host (push) Successful in 14m44s
windows-host / package (push) Successful in 16m32s
arch / build-publish (push) Successful in 17m49s
ci / rust (push) Successful in 19m58s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m35s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m42s
settle_desktop_portal pushes the live desktop's WAYLAND_DISPLAY into the
systemd --user manager (import-environment) so a re-activated portal inherits
the right session — but the import PERSISTS after that desktop dies. Every
later user unit then inherits the stale socket, including
gamescope-session.target: gamescope sees WAYLAND_DISPLAY, runs NESTED, and
aborts with "Failed to connect to wayland socket: wayland-0" — observed live
on a Deck, where Game Mode could not start at all (autologin crash-looped)
until the var was unset by hand.

Two-layer fix, mirroring the protection launch_session already gives its
transient unit:
- observe_session_instance: when a desktop compositor instance goes away,
  unset WAYLAND_DISPLAY/DISPLAY from the manager env (a desktop bounce is
  harmless — the next portal settle re-imports).
- write_steamos_dropin: UnsetEnvironment=DISPLAY WAYLAND_DISPLAY on the
  takeover drop-in, unit-scoped belt-and-suspenders for host-initiated
  restarts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 11:06:55 +02:00
enricobuehler 333c8979e9 fix(vdisplay/gamescope): capture-loss rebuild probes must not steal the box session
Observed live on a Deck host: the user switched Game Mode → Desktop, KDE came
up in 0.8 s — and the capture-loss rebuild's FIRST detection ran inside that
gap, still said Gaming, and its gamescope re-acquire restarted
gamescope-session.target. On SteamOS that steals the seat: the user was
yanked out of the KDE session they had just chosen, the two session managers
fought (job canceled / dependency failed), no gamescope node appeared within
30 s, and the stream died at the 40 s rebuild budget.

A rebuild probe acting on possibly-stale detection must never stop, relaunch,
or take over box sessions. New pf_vdisplay::rebuild_probe_scope (RAII,
counted): while held, the gamescope managed/takeover create paths only attach
to a live node and fail fast otherwise — no stop_autologin_sessions, no
session-plus relaunch, no gamescope-session.target restart. The capture-loss
loop holds it for the first 4 s after a loss (the detection-ambiguity
window); after that, destructive rebuilds are allowed again, so a genuine
switch INTO Game Mode still gets its headless takeover. Probe failures are
instant, so the loop now paces itself at ~2 Hz instead of spinning.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 11:00:02 +02:00
enricobuehler 46d09ed973 fix(host/linux): stop burning retry budget on doomed pipeline builds
Two session-transition stalls found live on a SteamOS Deck host, one cause:
the pipeline retry loop couldn't tell a wait-it-out failure from a
retry-now-and-it-works one.

- Gamescope first-frame race: a PipeWire stream connected while gamescope
  re-initializes its headless takeover negotiates a format, reaches
  Streaming, and never receives a buffer — while a fresh connect delivers
  within ~0.5 s. Every gamescope bring-up ate the full 10 s first-frame
  budget on attempt 1 (17 s bring-ups; KWin: 1.2 s). The retry loop's FIRST
  attempt now waits 2.5 s (Capturer::next_frame_within), so the reconnect
  that fixes the race happens at ~3 s. Later attempts keep the patient 10 s
  — the documented 30-60 s Big Picture cold start still fits the budget.

- Capture-loss rebuild vs session switch: the rebuild loop re-detects the
  active session between build_pipeline_with_retry calls, but each call
  burned 8 attempts (~13 s) against a compositor that no longer exists (a
  Desktop→Gaming switch spent 13 of its 27 s retrying gone-KWin). The
  capture-loss path now passes max_attempts=2, turning the outer loop into
  ~1 s detect-and-retry cycles that follow the box to the new session.

The resize path's direct build_pipeline call keeps the default budget (no
retry wrapper there to absorb an early bail).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 10:34:49 +02:00
enricobuehler 946f1b3d69 fix(steamdeck): build the host with vulkan-encode, and wire up what it needs to actually engage
ci / web (push) Successful in 54s
ci / docs-site (push) Successful in 57s
apple / swift (push) Successful in 1m20s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 6m43s
apple / screenshots (push) Successful in 6m32s
deb / build-publish (push) Successful in 9m57s
docker / deploy-docs (push) Successful in 26s
arch / build-publish (push) Successful in 12m46s
android / android (push) Successful in 13m40s
deb / build-publish-host (push) Successful in 13m3s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m42s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 18m59s
ci / rust (push) Successful in 26m41s
The SteamOS scripts built the host featureless, so PlatformAuto could never
pick the Vulkan Video backend (default-on since 84329205) and every session
silently ran libav VAAPI — unlike the deb/arch/rpm/sysext builds, which all
pass the feature. Three gaps, found live on a Deck host:

- install.sh/update.sh: cargo build with --features punktfunk-host/vulkan-encode
  (matches the packaged builds; pure-Rust ash, no new system dep).
- host.env: RADV_PERFTEST=video_encode — Van Gogh RADV still gates
  VK_KHR_video_encode_* behind it, so without it the backend can't open and
  falls back to VAAPI. Harmless where encode is exposed by default.
- io.unom.Punktfunk.Host.desktop (Exec rewritten to this install's binary):
  KWin only grants zkde_screencast_unstable_v1/org_kde_kwin_fake_input to an
  exe a .desktop authorizes — without it Desktop-Mode capture fails outright
  and a mid-stream Game↔Desktop switch strands the stream on the old backend.

update.sh retrofits the last two idempotently so existing installs get them
without a reinstall.

Validated on a SteamOS 3.8.14 Deck (Van Gogh): Vulkan HEVC opened on both the
KWin and gamescope paths at 5120x1440 — in-place ABR reconfigure (no IDR, no
encoder rebuild) vs the VAAPI path's full teardown per ABR step, and Desktop
capture came up without a re-login.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 10:14:31 +02:00
enricobuehler 7c0313a7cb fix(web): flush card content, so a full-bleed table stops doubling its inset
ci / web (push) Successful in 55s
ci / docs-site (push) Successful in 58s
apple / swift (push) Successful in 1m18s
ci / bench (push) Successful in 6m42s
apple / screenshots (push) Successful in 6m21s
ci / rust (push) Successful in 24m18s
android-screenshots / screenshots (push) Successful in 3m2s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
android / android (push) Successful in 15m7s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 37s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 12s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
windows-host / package (push) Successful in 10m7s
release / apple (push) Successful in 10m12s
flatpak / build-publish (push) Successful in 6m55s
deb / build-publish (push) Successful in 11m44s
linux-client-screenshots / screenshots (push) Successful in 7m40s
docker / deploy-docs (push) Successful in 28s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m22s
web-screenshots / screenshots (push) Successful in 3m27s
deb / build-publish-host (push) Successful in 12m25s
arch / build-publish (push) Successful in 15m42s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m32s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m48s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 14m41s
`<CardContent className="p-0">` never removed the padding it looked like it
removed: tailwind-merge resolves conflicts within a variant, so `p-0` cancels
`p-4` but leaves `sm:p-6` standing. Every call site that used it nests a
CardHeader inside, which brings its own `sm:p-6` — so at >=640px the title sat
24px further in than its neighbours'. Visible on the Plugins page as 'Installed
plugins' not lining up with 'Plugin runner'.

Give CardContent a `flush` prop that omits the padding outright, and use it at
the three call sites (Store, Pairing, Stats) rather than fixing the symptom
three times and leaving the footgun in place.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 00:25:31 +02:00
enricobuehler 04c975e9cb fix(store): uninstall must refuse a plugin's framework, not just odd names
apple / swift (push) Successful in 1m26s
apple / screenshots (push) Successful in 6m47s
ci / web (push) Successful in 1m9s
ci / docs-site (push) Successful in 1m5s
windows-host / package (push) Successful in 9m45s
android / android (push) Successful in 12m4s
arch / build-publish (push) Successful in 12m4s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 40s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
ci / bench (push) Successful in 5m57s
docker / deploy-docs (push) Successful in 24s
deb / build-publish (push) Successful in 9m0s
deb / build-publish-host (push) Successful in 9m32s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m15s
ci / rust (push) Successful in 24m40s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 19m55s
The guard checked the package NAME shape, and `@punktfunk/plugin-kit` — the
framework every kit-built plugin depends on — matches `@scope/plugin-*`
exactly. Windows on-glass accepted an uninstall of it; bun no-ops removing a
non-dependency so nothing broke, but the store was offering to pull a library
out from under the plugins using it.

Require membership in the top-level dependency list, the same authority that
already keeps plugin-kit out of the installed listing. Same blind spot, two
call sites — the listing was fixed during implementation, this one wasn't.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 00:10:49 +02:00
enricobuehler 402ec90edc fix(web/store): restore the plugin cards' top padding, and capitalise platform names
ci / web (push) Successful in 1m5s
ci / docs-site (push) Successful in 1m5s
apple / swift (push) Successful in 1m24s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 8s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 31s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
ci / bench (push) Successful in 5m34s
apple / screenshots (push) Successful in 6m33s
windows-host / package (push) Successful in 9m46s
deb / build-publish (push) Successful in 9m19s
docker / deploy-docs (push) Successful in 26s
deb / build-publish-host (push) Successful in 9m48s
android / android (push) Successful in 17m6s
arch / build-publish (push) Successful in 18m10s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m30s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 21m38s
ci / rust (push) Successful in 27m42s
Two things looked wrong in the store:

The cards had no space above the icon. `CardContent` zeroes its own top padding
(`pt-0`/`sm:pt-0`) because it normally sits under a `CardHeader` that supplies
it — but these cards have no header, so the top padding simply vanished while
the other three sides kept `p-card`. `p-card` does not undo it: `card` is a
custom `--spacing-*` token, which tailwind-merge does not recognise as a
spacing value and so never dedupes against `pt-0`, leaving the longhand to win.
Both header-less cards in this section now restore it explicitly, at both
breakpoints.

Platform chips read "windows" / "linux". Those are the catalog's platform
IDENTIFIERS (the index validator pins `linux | windows | macos`), which is right
for the wire and wrong on screen; the card now maps them to display names —
Linux, Windows, macOS. Proper nouns, so deliberately not routed through i18n.

Adds Store/StoreCard stories covering the padding, the platform labels, and the
installed / update / incompatible / blocked / external states, so the card can
be eyeballed without a host or a catalog.
2026-07-21 00:00:37 +02:00
enricobuehler d783d5445c chore(release): bump workspace version to 0.17.0
apple / swift (push) Successful in 1m20s
android / android (push) Has been cancelled
apple / screenshots (push) Has been cancelled
arch / build-publish (push) Has been cancelled
audit / cargo-audit (push) Has been cancelled
audit / bun-audit (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
deb / build-publish (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
flatpak / build-publish (push) Has been cancelled
release / apple (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
windows-host / package (push) Has been cancelled
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Has been cancelled
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Has been cancelled
windows / build (aarch64-pc-windows-msvc) (push) Has been cancelled
windows / build (x86_64-pc-windows-msvc) (push) Has been cancelled
2026-07-20 23:55:51 +02:00
enricobuehler 67b7981096 feat(encode/nvenc-linux): LN1 phase-3 — multi-slice sub-frame encode DEFAULT-ON for Linux direct-NVENC
apple / swift (push) Successful in 1m16s
apple / screenshots (push) Successful in 6m15s
windows-host / package (push) Successful in 9m39s
ci / web (push) Successful in 1m10s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
Gates recorded (nvenc-subframe-slice-output.md Phase 3): loss-harness curve
identical streamed-vs-whole; adversarial security review clean; live .21
3-leg A/B — streamed 0 corrupted frames, p99 8527→5363 µs vs sliced-whole,
first_slice_us ~500 µs of a ~1250 µs encode reaching the send thread; wire
bytes rate-governed under CBR + 1-frame VBV (the ~1-2 % slice-header cost is
absorbed by rate control, not added to the wire). GameStream: plane has zero
code diff (own RTP packetizer; shares only untouched send_pacing); a live
Moonlight re-test joins the standing owed Moonlight item.

- resolve_slices(codec, default): env 1..=32 wins (1 = the NEW explicit
  single-slice escape), else the backend default — 4 on Linux direct-NVENC,
  1 elsewhere. resolve_subframe(default_on): PUNKTFUNK_NVENC_SUBFRAME
  tri-state (0 = never, 1 = force, unset = backend default).
- Linux nvenc_cuda: resolves both ONCE in query_caps — the sub-frame default
  is gated on the GPU's SUBFRAME_READBACK cap so an unsupporting GPU never
  has enableSubFrameWrite forced into its init params — and feeds the same
  resolved values to build_config, build_init_params (open AND in-place
  reconfigure) and the chunked-poll latch.
- Windows: env-only as before (async path untouched, byte-identical config
  for unset env); LowLatencyConfig.slices + the explicit subframe param
  replace the shared env reads.
- Tests: chunked e2e now runs at the DEFAULTS (no knobs); the fallback test
  becomes the escape test (SLICES=1 disarms; SUBFRAME=0 disarms while
  4-slice encode continues on the plain poll path).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 23:44:41 +02:00
enricobuehler e6cdd79f71 fix(loss-harness): drop the needless mut on send_wires — workspace clippy is -D warnings
apple / swift (push) Successful in 1m17s
apple / screenshots (push) Has been cancelled
android / android (push) Has been cancelled
deb / build-publish-host (push) Failing after 5s
ci / web (push) Successful in 1m14s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
ci / docs-site (push) Successful in 1m1s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
ci / bench (push) Successful in 5m9s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 10m44s
deb / build-publish (push) Successful in 12m34s
arch / build-publish (push) Has been cancelled
ci / rust (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
Slipped through 421e8112: the on-box clippy runs covered punktfunk-core and
punktfunk-host but not the tools crates; CI's workspace-wide pass caught it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 23:44:37 +02:00
enricobuehler 7e5ec7eba9 chore(store): repin the official plugin-index signing key
apple / swift (push) Successful in 1m17s
apple / screenshots (push) Successful in 6m43s
ci / web (push) Successful in 1m8s
ci / docs-site (push) Successful in 1m13s
arch / build-publish (push) Successful in 12m43s
deb / build-publish-host (push) Failing after 25s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
ci / rust (push) Failing after 9m38s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
ci / bench (push) Successful in 6m58s
android / android (push) Successful in 16m41s
docker / deploy-docs (push) Successful in 11s
windows-host / package (push) Successful in 10m9s
deb / build-publish (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
The bootstrap key generated during implementation was lost, and its
replacement was pasted into the wrong repository's secret store — Gitea
secrets neither cross repos nor read back, so neither is usable. The private
half of this one lives in punktfunk-plugin-index's own INDEX_SIGNING_KEY,
which is the only thing that signs the catalog.

Nothing has shipped pinning any of these, so this is a straight replacement
of slot 0; the second slot stays reserved for a genuine rotation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 23:25:16 +02:00
enricobuehler c6caeca5bc feat(encode): RGB-direct (EFC) is now the DEFAULT on capable hosts — gated by a session cursor-blend hint
apple / screenshots (push) Has been cancelled
apple / swift (push) Has been cancelled
android / android (push) Has been cancelled
ci / web (push) Successful in 55s
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
ci / docs-site (push) Successful in 55s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 20s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 56s
ci / bench (push) Successful in 7m8s
ci / rust (push) Failing after 9m14s
windows-host / package (push) Successful in 10m24s
arch / build-publish (push) Successful in 12m58s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m25s
docker / deploy-docs (push) Successful in 28s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 20m25s
B2 default-on (design/vulkan-rgb-direct-encode.md): wherever the probe
passes, Vulkan Video sessions now take the EFC RGB source by default —
direct import on aligned modes, padded-copy staging on unaligned ones
(1080p). PUNKTFUNK_VULKAN_RGB_DIRECT becomes the override: =0 disables,
=1 forces, unset = the default below.

The one session class that must NOT default on: cursor-as-metadata
captures (every non-gamescope compositor), where the CSC shader's blend
IS the visible pointer — the EFC cannot composite, and defaulting there
would silently drop the cursor from the stream. The hint rides the
existing plumbing:

- SessionPlan gains cursor_blend, resolved once where the compositor is
  known (gamescope embeds the pointer itself → false; kwin/mutter/
  wlroots/hyprland → true), and shows up in the logged plan line.
- open_video/open_video_backend thread it through (native pump: all
  three encoder-open sites read plan.cursor_blend; GameStream monitor
  capture: true — it negotiates metadata cursor; spike: false).
- VulkanVideoEncoder::open resolves: env override, else ON iff the
  session never hands us cursor bitmaps. The warn-once for a cursor on
  an RGB session (forced via =1) stays.

Verified on-hw box (Linux): pf-encode + punktfunk-host compile, clippy
clean, unit suite green. The GPU paths themselves are unchanged from the
smoke-validated 96e19986 — this commit only changes which sessions
select them by default.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 23:24:46 +02:00
enricobuehler 3736bbdc73 fix(core): streamed-AU hardening from the security-review pass
apple / swift (push) Successful in 1m19s
docker / deploy-docs (push) Has been cancelled
android / android (push) Has been cancelled
apple / screenshots (push) Has been cancelled
arch / build-publish (push) Has been cancelled
release / apple (push) Successful in 8m41s
ci / rust (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
flatpak / build-publish (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
windows-host / package (push) Has been cancelled
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Has been cancelled
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Has been cancelled
windows / build (aarch64-pc-windows-msvc) (push) Has been cancelled
windows / build (x86_64-pc-windows-msvc) (push) Has been cancelled
Review verdict: no exploitable memory-safety / amplification / splice /
resurrection issue; nonce order and two-lane seal byte-identical for the
legacy path. Actionable findings applied:
- explicit per-block data_shards cross-check next to the recovery_shards
  one (defense-in-depth — the geometry invariants make it unreachable
  today, but have_data indexing assumes it)
- partial delivery skips an UNPINNED streamed frame (frame_bytes still 0)
  instead of truncating its max-sized buffer to an empty "partial"
- regression tests for the two untested load-bearing guards: the
  final-first out-of-range-sentinel drop (the exact-sized-buffer guard)
  and second-final-with-different-totals rejection

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 23:13:59 +02:00
enricobuehler 421e811225 feat(loss-harness): sweep the streamed-AU wire shape alongside whole-AU
The Phase-2 gate: more, smaller wire units must not regress FEC recovery.
Adds a GF16 streamed column (three chunk pushes + finish per AU — sentinel
blocks then real totals). Measured: identical to whole-AU at every loss
rate (50/50 through 1/6, the same 25%-budget cliff).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 22:59:09 +02:00
enricobuehler 79913a430e feat(probe): advertise VIDEO_CAP_STREAMED_AU
The probe's reassembler is the shared core one, and the probe is the tool
that measures the streamed-AU overlap win — advertise the cap so a host
with chunked encode streams to it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 22:55:34 +02:00
enricobuehler c5402cb1f7 feat(core+host): LN1 phase-2 — VIDEO_CAP_STREAMED_AU streamed access units
The wire half of sub-frame slice output (latency §7 LN1, planning
design/nvenc-subframe-slice-output.md Phase 2): toward a client that
advertises the new cap bit (0x20), a chunked-poll encoder session ships each
AU's completed FEC blocks while the tail of the frame is still encoding —
the AU's last packet leaves the host as the encode finishes instead of
after it.

Wire semantics (negotiated; zero change for anyone else):
- Non-final blocks ride SENTINEL headers: block_count = 0 (a value no legacy
  sender emits) + frame_bytes = 0 + exactly max_data_per_block data shards,
  so the receiver's shard-offset formula needs no total.
- The final block's headers carry the real frame_bytes/block_count (+
  FLAG_EOF) and RETRO-VALIDATE the whole frame: totals under which a
  received sentinel block is out of range or not full-K kill the frame
  wholesale (no spliced delivery) and the index can't be resurrected.
- Firewall: sentinels are bounded by the negotiated limits (full-K exactly,
  never the last block the limits allow, no total to lie about); the exact
  derived-geometry check runs unchanged on every non-sentinel packet and
  retroactively at pinning. Sentinel opens commit a max_frame_bytes buffer,
  bounded by the existing IN_FLIGHT_BUF_FACTOR budget (amplification test).
- Order-agnostic like legacy: a reversed frame (final block first) opens
  legacy-shaped and still accepts its sentinels against the pinned totals.
- Small/empty streamed AUs degenerate to byte-identical legacy headers.

Host: Packetizer::{begin,push,finish}_streamed seal full-K blocks (data +
parity per block) as chunks arrive; Session::seal_streamed_* share the
pooled-wire + two-lane seal machinery via the new seal_run; the send thread
paces each flush under the frame's existing deadline (pace_sealed split out
of paced_submit) and runs the whole-AU accounting at the last chunk; the
encode pump forwards poll_chunk output as ChunkMsg when the client has the
cap AND the encoder chunks (re-queried per AU — an escalation falls back
seamlessly). Probes never run mid-AU. PUNKTFUNK_STREAMED_AU=0 = host escape
hatch. Client core ORs the cap into Hello (the shared reassembler carries
the support). Sampled first_slice_us vs encode_us PERF log measures the
overlap; the 0xCF stage-field extension stays a follow-up.

Core tests: streamed round-trip (clean/loss/reorder/duplicate, both orders),
sentinel firewall bounds, lying-final wholesale kill + no-resurrect,
open-amplification budget, header-shape pins. Gates still owed before
default-on: security review pass, loss-harness curve, GameStream smoke
(plane untouched structurally), bitrate A/B.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 22:53:53 +02:00
enricobuehler 9d3b114fd6 feat(encode/nvenc): LN1 phase-1 — slice-boundary chunked encoder poll (poll_chunk/AuChunk)
apple / swift (push) Successful in 1m20s
apple / screenshots (push) Successful in 6m24s
windows-host / package (push) Successful in 10m45s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / web (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
The encoder half of sub-frame slice output (latency §7 LN1, planning
design/nvenc-subframe-slice-output.md Phase 1): with PUNKTFUNK_NVENC_SLICES=N +
PUNKTFUNK_NVENC_SUBFRAME=1 on a sync depth-1 session, Encoder::poll_chunk hands
the in-flight AU out as slice-boundary chunks read through doNotWait sub-frame
locks while the tail is still encoding — the readback loop the on-hw probe
validated (~200 µs slice spacing on the 5070 Ti), productionized.

- codec.rs: AuChunk (AU metadata on the first chunk, `last` closes the AU;
  chunks concatenate to exactly the poll() bytes) + supports_chunked_poll /
  poll_chunk trait surface. Default impl wraps poll() as one self-closing
  chunk, so a chunk-driven consumer works against every backend.
- nvenc_cuda: chunked readback cut at slice boundaries only (bitstream size at
  n reported slices = end of slice n, Annex-B contiguous); completion is NEVER
  numSlices alone — one finishing BLOCKING lock is the authority and the wedge
  watchdog, so the final chunk blocks exactly like sync poll (the depth-1 pump
  contract; 6dc195f9 bug class). Keyframe on early chunks is the submit-time
  IDR prediction (exact under P-only/infinite-GOP), cross-checked at finish.
  Debug builds shadow-check emitted chunks against the finished AU prefix.
  Mutually exclusive with pipelined retrieve (gated off when async_rt exists,
  dropped by the escalation rebuild); composes with stream-ordered submit.
- nvenc_core: slices_env/subframe_requested shared parses so the config author
  and the chunked-poll arming can't disagree.
- TrackedEncoder forwards both new methods (the set_wire_chunking trap class).

Host loop untouched — Phase 2 (VIDEO_CAP_STREAMED_AU seal/send) consumes this.
On-hw: nvenc_cuda_chunked_poll_end_to_end + nvenc_cuda_chunked_poll_fallback_whole_au.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 20:53:32 +02:00
enricobuehler 411c6dc06a fix(encode): forward set_pipelined through TrackedEncoder — the LN3 escalation no-oped through the wrapper
Every session encoder is boxed in TrackedEncoder (open_video), and the
wrapper never forwarded Encoder::set_pipelined — so the host loop's
contention escalation (ae673158) hit the trait default, which returns
false ("backend can't pipeline, stop asking"), and the pipelined-retrieve
stage could never engage adaptively. Only the explicit
PUNKTFUNK_NVENC_ASYNC=1 open-time path worked. The exact defaulted-method
trap the wrapper's own comment documents for set_wire_chunking.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 20:53:32 +02:00
enricobuehler 833f3348a0 feat(store): console plugin store, index repo, and the fixes on-glass found
apple / swift (push) Successful in 1m22s
release / apple (push) Successful in 9m59s
windows-host / package (push) Successful in 9m52s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m20s
apple / screenshots (push) Successful in 6m28s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m21s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 4m39s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 5m45s
audit / cargo-audit (push) Successful in 2m58s
audit / bun-audit (push) Successful in 13s
ci / web (push) Successful in 52s
ci / docs-site (push) Successful in 59s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / bench (push) Has been cancelled
ci / rust (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
decky / build-publish (push) Successful in 23s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 35s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 49s
flatpak / build-publish (push) Successful in 7m53s
docker / deploy-docs (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
deb / build-publish (push) Successful in 11m41s
deb / build-publish-host (push) Successful in 13m32s
Console: a Plugins section (Browse / Installed / Sources) on a static nav
entry, with install friction proportional to trust — a plain confirm for a
verified entry, a warning naming the curator for an external one, and a
danger dialog that makes you retype the spec for a raw package. Tier badges
are permanent and follow the plugin onto its own UI page.

Index: unom/punktfunk-plugin-index published and served from Gitea's raw
endpoint (real HTTPS, byte-exact, no vhost to stand up) — merge to main is
publish, which resolves the design's open hosting question.

Four things only running it could find:
- runner discovery matched @punktfunk/plugin-* only, so a third-party
  scoped plugin (which D8 requires) would install and never run
- ...and that convention also matches @punktfunk/plugin-kit, a plugin's
  own framework: it listed as installed and would have been imported as a
  unit. Both now key off the plugins dir's top-level dependencies, with an
  emptied dependency list meaning 'nothing installed' rather than falling
  back to the naming convention
- the store must not pass new flags to the runner: the scripting package
  ships separately and an older one reads an unknown flag's value as a
  package name. The host writes the bunfig scope mapping itself
- ureq reports only >= 400 as Err, so a conditional request's 304 arrives
  as Ok with an empty body — handled as an error it made every refresh
  after the first verify a signature over zero bytes and sit stale

Also: the console's first Tabs use exposed an @unom/ui theme gap that
rendered inactive tabs invisible (caught in a browser pass, not by types).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 20:48:02 +02:00
enricobuehler 45c3b96907 feat(store): plugin store host module — signed catalogs, tiered trust, install jobs
The index is the verification gate: a catalog entry pins one exact version
plus that version's tarball integrity hash, so 'verified on every release'
is a property of the data rather than a promise about process. Nothing can
express 'track latest' for a catalogued plugin.

- store/index.rs   signed index parse + ed25519 verify (ring), validate-and-drop
                   per entry, semver minHost/advisory ranges
- store/sources.rs built-in unom source (compiled-in URL + two key slots for
                   rotation) + operator sources in plugin-sources.json
- store/catalog.rs https fetch with size/timeout/redirect caps, signature before
                   parse, last-good disk cache (stale-but-usable when offline)
- store/jobs.rs    single-flight install/uninstall: registry-integrity preflight
                   against the pin, spawn the runner CLI with live log capture,
                   post-install version check with rollback, provenance record,
                   runner restart (discovery is startup-only)
- store/manifest.rs install provenance; absence means CLI-installed
- mgmt/store.rs    12 routes under /api/v1/store, denied to the plugin token
                   (a plugin that can install plugins is an escalation primitive)

Also generalizes runner discovery and listInstalled from @punktfunk/plugin-*
to ANY scope's plugin-*: catalog entries must be scoped so the scope can map
to that entry's registry, so a third-party plugin necessarily arrives under
its own scope and would otherwise install but never run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 20:48:02 +02:00
enricobuehler 96e19986bc feat(encode): RGB-direct padded-copy — unaligned modes (1080p) get the EFC path safely
apple / swift (push) Successful in 1m20s
apple / screenshots (push) Successful in 6m42s
windows-host / package (push) Successful in 9m24s
android / android (push) Failing after 18s
ci / web (push) Successful in 1m8s
ci / docs-site (push) Successful in 1m43s
arch / build-publish (push) Successful in 11m48s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 12s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
ci / bench (push) Successful in 7m33s
deb / build-publish (push) Successful in 9m1s
deb / build-publish-host (push) Successful in 13m6s
ci / rust (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
Lifts 3aacec53's alignment gate into a mode select: an aligned mode keeps
the true zero-copy direct import; an unaligned one (1080p!) now blits the
visible frame into a per-slot ALIGNED BGRA staging image and duplicates
the edge rows/columns into the 64x16 padding — transfer-only regions in
one vkCmdCopyImage (1080p: visible + 8 row regions; width padding adds a
second self-copy pass in GENERAL), no compute shader. The encode reads
the staging image, never the capture buffer, so the EFC can never read
past a producer allocation (the field GPU hang). Still one ~8 MB copy vs
the CSC path's ~17 MB + dispatch + plane copies. Verdict line:
active(padded-copy). The staging import drops the video profile entirely
(TRANSFER_SRC only).

CPU-payload paths made honest on the way (they were the smoke baseline
AND the software-capture fallback):
- rgb mode: the staging upload is padded CPU-side (edge duplication) so
  the aligned encode-src is fully defined;
- CSC mode: the sampled image is now SOURCE-sized with a matching copy
  extent — the old aligned-size image + tightly-packed buffer sheared
  rows and left garbage rows at unaligned modes (black-bar artifacts;
  YMIN=16 in every smoke frame), which also masked as a 24 dB PSNR
  'regression' against the (correct) padded output.

On-glass (780M, host Mesa 26.0.4): all four smokes pass at 256x256
(direct) AND 250x250 (padded); padded-EFC frames decode perfectly
uniform (YMIN==YMAX) at the exact 709-narrow values (79/148/60 for the
first three fills); CSC-vs-padded PSNR 49.9 dB avg after the baseline
fix. clippy clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 20:38:55 +02:00
enricobuehler 3aacec53d8 fix(encode): RGB-direct must not engage on unaligned modes — the EFC reads past the capture buffer (GPU hang)
apple / swift (push) Successful in 1m23s
apple / screenshots (push) Successful in 6m27s
windows-host / package (push) Successful in 9m29s
ci / web (push) Successful in 1m1s
ci / docs-site (push) Successful in 1m15s
ci / bench (push) Successful in 6m8s
android / android (push) Successful in 15m47s
arch / build-publish (push) Successful in 14m49s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 13s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 12s
decky / build-publish (push) Successful in 25s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13s
deb / build-publish (push) Successful in 9m32s
ci / rust (push) Successful in 19m0s
deb / build-publish-host (push) Successful in 13m3s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m8s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 22m27s
docker / deploy-docs (push) Successful in 13s
Field report (commit 25624443, RDNA2/VCN3 host, 1080p): repeated VCN0 VM
protection faults from punktfunk-host, vcn_enc_0.0 ring timeouts, two ring
resets, then a failed reset escalating to a full MODE1 GPU reset with VRAM
loss. Root cause is B1's RGB-direct source: the session's coded extent is
64x16-aligned (1920x1088) but the captured dmabuf is only the real mode
size (1920x1080) — the VCN EFC reads the 8 padding rows past the end of
the buffer (8 x 7680 B = 61 KB, matching the faulting page span exactly).
The CSC shader absorbs alignment by clamping reads and duplicating edges;
RGB-direct has no such stage. Every prior validation surface happened to
be aligned (256x256 smoke, 1440p live box) — 1080p is the common mode
that is not. The stall watchdog's rebuild-and-refault loop then turned a
per-session fault into the GPU death spiral.

Two layers:
- Open-time gate: RGB-direct engages only when the real mode equals the
  aligned coded extent (720p/1440p/4K yes, 1080p no); otherwise the CSC
  path runs and the verdict line says why (unaligned-mode). The padded-
  copy variant that re-enables 1080p is design-doc B2 work.
- Submit-time check: an RGB-direct frame that does not cover the coded
  extent errors out (encoder-rebuild path) instead of importing an
  undersized source.

On-glass (780M, host Mesa 26.0.4): all four smokes pass aligned; at
PF_SMOKE_W/H=250 the gate refuses RGB-direct and the rgb smokes soft-skip.
clippy clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 20:23:29 +02:00
enricobuehler 84e1e5aeb4 fix(web): restore the Dashboard status tiles' top padding
apple / swift (push) Successful in 1m30s
apple / screenshots (push) Successful in 6m28s
windows-host / package (push) Successful in 9m27s
ci / web (push) Successful in 1m10s
ci / docs-site (push) Successful in 1m11s
android / android (push) Successful in 13m1s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
decky / build-publish (push) Successful in 20s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 37s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 22s
ci / bench (push) Successful in 6m9s
arch / build-publish (push) Successful in 15m32s
deb / build-publish (push) Successful in 9m10s
deb / build-publish-host (push) Successful in 10m45s
ci / rust (push) Successful in 24m42s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m6s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 17m24s
docker / deploy-docs (push) Successful in 12s
CardContent is shadcn's header-adjacent variant (`p-4 pt-0 sm:p-6 sm:pt-0`),
but the four Live-status tiles use it with no CardHeader. tailwind-merge lets
the call site's unprefixed `p-4` cancel the base `pt-0`, yet nothing cancels
`sm:pt-0` — so the tiles lost their top padding at >=640px only, and the
content sat against the top edge. Restores the top inset at the sm breakpoint
and lets the content fill the stretched grid cell so it stays centred when a
sibling tile is taller.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 20:13:34 +02:00
enricobuehler 1262fec2ae perf(zerocopy/vulkan): LN4 — global-priority VkBridge queue for the LINEAR/gamescope CSC
apple / swift (push) Successful in 1m22s
apple / screenshots (push) Successful in 6m27s
ci / web (push) Successful in 1m0s
ci / docs-site (push) Successful in 1m6s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
ci / bench (push) Successful in 7m51s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
deb / build-publish (push) Successful in 8m47s
arch / build-publish (push) Successful in 15m47s
android / android (push) Successful in 16m2s
deb / build-publish-host (push) Successful in 12m13s
ci / rust (push) Successful in 26m31s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m54s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m3s
docker / deploy-docs (push) Successful in 16s
The RGB→NV12 compute CSC shares the SMs with the game; request an
elevated global-priority queue (VK_KHR/EXT_global_priority) so the
driver can schedule it ahead. PUNKTFUNK_VK_QUEUE_PRIORITY =
off|high|realtime (default realtime), downgrading REALTIME→HIGH→none
on NOT_PERMITTED/INITIALIZATION_FAILED — a refused class never fails
the bridge. Mirrors PyroWave's ac0e7332 (same honest caveat: WDDM
didn't move; the Linux driver may — contended-tail lever, correct and
harmless either way).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 20:12:42 +02:00
enricobuehler ae67315804 perf(encode/nvenc-linux): LN3 — pipelined-retrieve escalation replaces the depth-1 async foot-gun
apple / swift (push) Successful in 1m24s
apple / screenshots (push) Successful in 5m16s
windows-host / package (push) Successful in 10m11s
ci / web (push) Successful in 59s
ci / docs-site (push) Successful in 1m40s
ci / bench (push) Successful in 5m13s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
decky / build-publish (push) Successful in 19s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
android / android (push) Successful in 17m6s
arch / build-publish (push) Successful in 17m27s
deb / build-publish (push) Successful in 12m39s
deb / build-publish-host (push) Successful in 9m33s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 7m27s
ci / rust (push) Successful in 27m10s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m1s
docker / deploy-docs (push) Successful in 17s
PUNKTFUNK_NVENC_ASYNC gains a real tri-state: 1 = always (as before, now
documented as ~+1 tick at depth-1), 0 = never (vetoes escalation), unset =
ADAPTIVE — off until the session loop's cadence-overrun detector escalates.

The host loop's adaptive-depth leaky bucket grows a second stage: once the
capturer's depth is maxed (Linux portal is permanently depth-1), it asks
the encoder for pipelined retrieve via the new Encoder::set_pipelined hook
(asked exactly once; default impl declines, Windows untouched).

nvenc_cuda engages at a safe point via a clean session rebuild WITHOUT the
IO-stream binding: with input==output stream bound, later stream work
waits on prior encode completions and would serialize a pipelined session
— stream-ordered submit and two-thread retrieve are mutually exclusive.
The ordered gate now also requires async_rt absence (belt-and-braces for
the runtime switch). Re-open's first frame is the standard session IDR.

On-hardware test: escalate mid-session → retrieve thread live, binding
gone, all AUs deliver, first post-escalation AU is the IDR.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 20:10:24 +02:00
enricobuehler 256244430b feat(encode): Vulkan Video B1 — RGB-direct encode source via VCN EFC (PUNKTFUNK_VULKAN_RGB_DIRECT=1)
apple / swift (push) Successful in 1m25s
windows-host / package (push) Successful in 11m6s
apple / screenshots (push) Successful in 6m55s
ci / web (push) Successful in 1m8s
ci / docs-site (push) Successful in 1m4s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 15s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
ci / rust (push) Failing after 10m50s
arch / build-publish (push) Successful in 14m36s
ci / bench (push) Successful in 6m3s
android / android (push) Successful in 15m19s
deb / build-publish-host (push) Successful in 9m39s
deb / build-publish (push) Successful in 10m24s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m6s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m53s
docker / deploy-docs (push) Successful in 18s
design/vulkan-rgb-direct-encode.md B1: when the probe passes AND
PUNKTFUNK_VULKAN_RGB_DIRECT=1 (default OFF until the B2 on-glass A/B),
the session opens with pictureFormat=B8G8R8A8 and the captured RGB frame
becomes the encode source — the VCN EFC front-end does the 709-narrow CSC
inline during encode. Deleted from the per-frame path: the compute CSC
dispatch, both NV12 plane copies, the semaphore hop and one queue submit
(dmabuf frames are ONE encode-queue submit; ~17 MB/frame of GPU traffic
gone). The DPB stays NV12; RFI/RC/quality-level machinery is untouched.

Shape:
- probe_rgb_direct now returns the chroma-siting bits to create the
  session with (midpoint preferred, else cosited-even per axis); the open
  verdict line gains "active" / "available(off; ...)" states.
- RgbProfileStack: the rgb-chained video profile rebuilt on the stack per
  profiled-image creation after open (profile identity is by value) —
  dmabuf imports become profiled VIDEO_ENCODE_SRC images (import cache
  unchanged), the CPU staging image likewise (concurrent encode+compute).
- record_submit_rgb: steps 2–4 twin (dmabuf: single submit; CPU: staging
  copy on the compute queue, semaphore-ordered); shared step-1/bookkeeping.
- begin_encode_cmd takes a SrcAcquire (CSC general / fresh-import FOREIGN
  QFOT / cached visibility-only / staging TRANSFER_DST) so both paths
  share the encode recording.
- RGB-direct frames skip the CSC per-slot resources entirely (make_frame
  split into csc/common halves); cursor bitmaps warn once (EFC cannot
  composite; gamescope — the flagship — embeds the cursor itself).

On-glass (780M RADV PHOENIX, host Mesa 26.0.4): all four smokes pass
(vulkan_smoke{,_av1,_rgb,_rgb_av1}); the EFC-encoded H265 stream decodes
clean and matches the CSC-encoded stream at 49.9 dB average PSNR (min
48.9) — within the design's ±1-code-value tolerance. (Raw-OBU ffmpeg
probing of the tiny AV1 dumps fails identically for BOTH paths —
pre-existing dump quirk, not a stream defect.) check+clippy+full unit
suite green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 20:05:56 +02:00
enricobuehler 589d48a975 fix(plugin-kit): SSE keepalive must outpace Bun's idle timeout (0.1.4)
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
decky / build-publish (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
arch / build-publish (push) Has been cancelled
apple / swift (push) Has been cancelled
apple / screenshots (push) Has been cancelled
android / android (push) Has been cancelled
servePluginUi's Bun.serve leaves idleTimeout at Bun's 10s default, so the old
15s keepalive could never arrive — an idle status feed was closed before its
first ping, which is exactly what a live host showed. Default is now 5s.

Also adds the regression suite that should have caught it: the original test
only covered a finite, self-driving stream. The new one exercises the
production shape (PubSub-backed, published after the response opens) plus the
keepalive — with a reader that does not re-enter read(), which is what made the
first version of this test lie.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:59:04 +02:00
enricobuehler 424bab1c3a fix(plugin-kit): soften @unom/ui's card ring for the brand palette (0.1.3)
apple / swift (push) Successful in 1m20s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
decky / build-publish (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
arch / build-publish (push) Has been cancelled
apple / screenshots (push) Has been cancelled
android / android (push) Successful in 14m52s
@unom/ui rings cards ring-2 ring-accent, sized for its own subtle accent; in
this palette --accent is the brand violet, so every card shouted. The console
already softens the same ring in its Card wrapper — plugin UIs now inherit that
instead of each wrapping AnimatedCard themselves.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:47:51 +02:00
enricobuehler 5a27c02738 feat(encode/nvenc): LN1 phase-0 — slice-count + sub-frame-readback knobs and the slice-timing probe
apple / swift (push) Successful in 1m16s
apple / screenshots (push) Successful in 6m55s
windows-host / package (push) Successful in 11m18s
ci / web (push) Successful in 54s
ci / docs-site (push) Successful in 56s
ci / bench (push) Successful in 5m19s
android / android (push) Successful in 12m34s
decky / build-publish (push) Successful in 28s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 12s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 14s
arch / build-publish (push) Successful in 15m25s
deb / build-publish (push) Successful in 11m48s
deb / build-publish-host (push) Successful in 9m33s
ci / rust (push) Successful in 24m38s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m29s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m36s
docker / deploy-docs (push) Successful in 10s
PUNKTFUNK_NVENC_SLICES=N (2..=32, default off) splits H.264/HEVC frames
into N slices (sliceMode 3); PUNKTFUNK_NVENC_SUBFRAME=1 (default off)
arms enableSubFrameWrite + reportSliceOffsets on sync sessions only.
Both experimental groundwork for sub-frame slice output (plan §7 LN1).

nvenc_cuda_subframe_slice_probe (on-hardware, ignored) answers the LN1
go/no-go: spins lock_bitstream(doNotWait) against an in-flight frame and
prints the (t_us, status, numSlices, bytes) timeline — incremental slice
availability and its spacing, or all-at-completion.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:44:54 +02:00
enricobuehler d5a0f5012b fix(encode): RGB-direct probe — accept cosited-even chroma siting (what VCN EFC actually offers)
apple / swift (push) Successful in 1m17s
windows-host / package (push) Successful in 9m23s
apple / screenshots (push) Successful in 6m43s
ci / web (push) Successful in 1m8s
ci / docs-site (push) Successful in 1m8s
ci / bench (push) Successful in 6m28s
decky / build-publish (push) Successful in 21s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
deb / build-publish (push) Successful in 9m33s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
android / android (push) Successful in 17m10s
arch / build-publish (push) Successful in 17m21s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Failing after 7m30s
deb / build-publish-host (push) Successful in 12m39s
ci / rust (push) Successful in 26m58s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m26s
docker / deploy-docs (push) Successful in 30s
On-glass B0 result from the 780M (RADV PHOENIX, host Mesa 26.0.4): the
probe returned no-709-narrow-midpoint because RADV advertises
xChromaOffsets = COSITED_EVEN only — the VCN EFC does canonical H.26x
left-cosited x-chroma, while the requirement was written to bit-match our
2x2-average (midpoint) shader. The delta is a half-pel chroma-x phase,
imperceptible and unsignalled in our bitstream; EFC's siting is arguably
the more correct one. Accept either siting bit per axis (model + range
stay exact: 709 narrow); B1 will pass the preferred available bit.

With this the 780M probes rgb_direct=available; both GPU smoke tests
(H265 + AV1) pass on that box against host Mesa 26.0.4 — the first
real-RADV validation of the poll fix, the quality-level control, and B0.
(Container mesa alone can't run the H265 smoke: Fedora ships RADV with
H264/H265 encode compiled out — AV1 only. Use the host ICD via
VK_ICD_FILENAMES on Atomic boxes.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 19:44:33 +02:00
enricobuehler d833cff470 fix(plugin-kit): resolve the --secondary text/surface collision in theme.css (0.1.2)
apple / swift (push) Successful in 1m19s
apple / screenshots (push) Successful in 6m34s
ci / web (push) Successful in 51s
ci / docs-site (push) Successful in 1m2s
ci / bench (push) Successful in 5m24s
deb / build-publish (push) Successful in 9m25s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 27s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
deb / build-publish-host (push) Successful in 9m13s
android / android (push) Successful in 18m26s
arch / build-publish (push) Successful in 18m56s
ci / rust (push) Successful in 19m2s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m25s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m43s
docker / deploy-docs (push) Successful in 23s
@unom/ui reads --secondary as muted TEXT; this palette (like the console's) uses
it as a dark SURFACE, so EmptyState captions, badge text and table headers
rendered dark-on-dark. Tailwind derives both utilities from one token, so the
text lane is re-pointed at --muted-foreground here — once, for every plugin,
instead of in each plugin's stylesheet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:44:01 +02:00
enricobuehler 34ec57cb52 test(encode/nvenc-linux): on-hardware assertion that stream-ordered submit arms
apple / swift (push) Successful in 1m20s
apple / screenshots (push) Successful in 6m19s
windows-host / package (push) Successful in 11m33s
ci / web (push) Successful in 55s
ci / docs-site (push) Successful in 1m2s
android / android (push) Successful in 12m27s
ci / bench (push) Successful in 5m21s
decky / build-publish (push) Successful in 1m56s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
arch / build-publish (push) Successful in 16m24s
deb / build-publish (push) Successful in 9m21s
deb / build-publish-host (push) Successful in 9m24s
ci / rust (push) Successful in 19m1s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 23m11s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 22m26s
docker / deploy-docs (push) Successful in 24s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:38:22 +02:00
enricobuehler cd36c46388 perf(encode/nvenc-linux): LN2 v1 — stream-ordered submit via NvEncSetIOCudaStreams
Bind the session's IO streams to the encode thread's high-priority copy
stream: in sync-retrieve depth-1 use the per-frame input copy and cursor
blend now enqueue with NO cuStreamSynchronize and encode_picture orders
after them on the stream. Same stream both directions, so the encode's
completion is inserted into the stream and later work (the next frame's
copy into a reused ring slot) waits for it.

Soundness gate: the fast path engages only when pending is empty (true
depth-1 usage) — every prior encode was drained by a blocking poll, and
the caller holds the frame payload across the matching poll (contract now
documented on Encoder::submit; both host loops already comply). Pipelined
callers and PUNKTFUNK_NVENC_ASYNC mode keep the blocking copies.

True zero-copy input registration (registering the worker-owned IPC
buffer directly) stays the LN2 v2 follow-up — it needs a contiguous
worker-pool NV12 layout and a registration<->IPC-mapping lifetime tie.

PUNKTFUNK_NVENC_STREAM_ORDERED=0 restores the old blocking behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:38:22 +02:00
enricobuehler 59b6d3796b perf(encode/nvenc-linux): LN0 — sampled submit/retrieve-lock split, sub-frame caps, zeroReorderDelay
Latency plan §7 LN0 (Linux/NVIDIA encode follow-on):
- sampled PUNKTFUNK_PERF submit split (copy/blend/map/pic) in nvenc_cuda —
  the host loop's submit_us folds all four together; the D2D input copy is
  the LN2 zero-copy target and now measurable on its own
- sampled blocking lock_bitstream timing on pf-nvenc-out — in two-thread
  mode the host loop's wait_us wraps a non-blocking poll, so the real
  encode wait was measured by no timer
- caps probe + log SUPPORT_SUBFRAME_READBACK / SUPPORT_DYNAMIC_SLICE_MODE
  (LN1 sub-frame slice-output prerequisites, fleet visibility)
- explicit zeroReorderDelay=1 in the shared low-latency config (P-only +
  no lookahead has no reordering anyway; pins the bit against preset or
  driver drift; shared with the Windows backend)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 19:38:22 +02:00
enricobuehler 78dba293a8 feat(encode): Vulkan Video B0 — vendored VK_VALVE_video_encode_rgb_conversion + RGB-direct probe telemetry
apple / swift (push) Successful in 1m16s
apple / screenshots (push) Successful in 6m16s
windows-host / package (push) Successful in 10m2s
ci / web (push) Successful in 1m13s
ci / docs-site (push) Successful in 1m18s
android / android (push) Successful in 13m46s
arch / build-publish (push) Successful in 12m58s
ci / bench (push) Successful in 7m23s
decky / build-publish (push) Successful in 38s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 21s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 14s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 17s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 16s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 14s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
deb / build-publish (push) Successful in 9m35s
deb / build-publish-host (push) Successful in 9m43s
ci / rust (push) Successful in 28m29s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 21m33s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 21m42s
docker / deploy-docs (push) Successful in 25s
First phase of design/vulkan-rgb-direct-encode.md (punktfunk-planning): make
the captured BGRx dmabuf the direct encode source with the VCN EFC front-end
doing the 709-narrow CSC — deleting the per-frame compute CSC, both plane
copies, the semaphore hop and one queue submit.

B0 changes nothing about the encode path. It vendors the extension surface
(vk_valve_rgb.rs — ash 0.38 predates it; same rationale and style as the
vk_av1_encode module) and probes at open whether this host qualifies:
extension present (Mesa >= 26.0 + EFC hardware) → feature bit → conversion
caps cover the compute shader's exact math (709 / narrow / midpoint both
axes) → encode-src format set offers B8G8R8A8 with DRM-modifier tiling.
The verdict is one INFO line (rgb_direct=available | first missing
requirement) — the field telemetry that decides where B1 can default on.

Verified: cargo check + clippy clean + unit tests green (Linux box); the
probe short-circuits to no-ext on non-RADV drivers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 19:29:22 +02:00
enricobuehler cfdec2729d perf(encode): Vulkan Video — explicit VCN preset, persistent bitstream map, CSC split timing
apple / swift (push) Successful in 1m19s
apple / screenshots (push) Successful in 6m25s
windows-host / package (push) Successful in 9m18s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / web (push) Successful in 46s
ci / docs-site (push) Successful in 51s
ci / bench (push) Successful in 6m15s
decky / build-publish (push) Successful in 34s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 15s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 16s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 15s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 14s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 15s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m36s
deb / build-publish (push) Successful in 12m42s
deb / build-publish-host (push) Successful in 13m40s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m53s
ci / rust (push) Successful in 28m29s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 22m1s
docker / deploy-docs (push) Successful in 28s
Field follow-up to 6dc195f9 (the AMD host whose H265 encode sits at ~5 ms):
the remaining time is the VCN ASIC preset, and we never chose one.

1. Explicit encode quality level. RADV only emits a VCN preset op when the
   app issues ENCODE_QUALITY_LEVEL — we never did, so every session ran the
   firmware's default preset. Install the resolved level on the first
   frame's RESET control (chained ahead of the CBR state, which satisfies
   the spec's quality-change-carries-rate-control rule) and bake the same
   level into the session-parameters object (the spec requires the match).
   Default 0 = the fastest tier the driver exposes (RADV: SPEED, except
   H265 on pre-RDNA4 which the driver pins to BALANCE); clamped to the
   profile's maxQualityLevels. PUNKTFUNK_VULKAN_QUALITY=0..3 overrides for
   quality-biased setups. Logged at open so field logs show the tier.

2. Persistent bitstream mapping. read_slot vkMapMemory/vkUnmapMemory'd the
   host-visible bitstream buffer every frame; map each ring slot once at
   build instead (coherent memory; vkFreeMemory implicitly unmaps).

3. PUNKTFUNK_PERF CSC split. GPU timestamps bracket the compute batch
   (import barriers + cursor prep + CSC dispatch + plane copies) and a
   sampled csc_us line logs every ~2 s — separating shader/copy time from
   the ASIC encode inside the pump's wait_us, the VAAPI submit-split
   analog for this backend. Off (and unrecorded) unless PUNKTFUNK_PERF.

Verified: cargo check + clippy + DPB/RPS unit tests green (Linux box).
The GPU smoke tests still need a RADV host (NVIDIA driver fails session
open on the build box, pre-existing).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 19:17:34 +02:00
enricobuehler 6dc195f982 fix(encode): Vulkan Video poll() must block per the depth-1 pump contract
apple / swift (push) Successful in 1m20s
windows-host / package (push) Successful in 9m30s
apple / screenshots (push) Successful in 6m39s
ci / web (push) Successful in 50s
ci / docs-site (push) Successful in 1m2s
ci / bench (push) Successful in 8m45s
decky / build-publish (push) Successful in 20s
android / android (push) Successful in 20m36s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 44s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 5m4s
ci / rust (push) Successful in 18m50s
deb / build-publish (push) Successful in 12m45s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
arch / build-publish (push) Successful in 23m15s
deb / build-publish-host (push) Successful in 11m37s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11m4s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 15m29s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 13m4s
docker / deploy-docs (push) Successful in 12s
The session pump's depth-1 loop (capture → submit → poll) treats a poll()
None as "the backend holds the frame internally, re-poll next tick" —
true only of the libav AMF/QSV wrappers, whose codecs genuinely hold ~2
frames. The Vulkan Video backend probed its slot fence with
get_fence_status instead, so the AU of every frame — finished on the
ASIC in ~5 ms — sat unharvested until the tick AFTER the next frame was
submitted: encode_us read as one full frame period (~17 ms at 60 Hz vs
VAAPI's ~5.3 ms on the same VCN, the AMD field report) and every frame
shipped a frame period late — one frame of avoidable glass-to-glass
latency on the path that exists to beat libav VAAPI.

Wait the oldest in-flight slot's fence in poll() instead, matching the
sync NVENC backend's blocking lock_bitstream and the documented
Capturer::pipeline_depth contract ("capture → submit → poll-blocks").
Bounded by ENCODE_FENCE_TIMEOUT_NS like enqueue()'s backpressure wait,
for the same reason: poll runs on the thread the stall recovery runs on,
so a wedged GPU must surface as an error, not park it. None now only
means "nothing submitted". Covers H265 and AV1 (shared poll).

Verified: cargo check + DPB/RPS unit tests green (home-worker-5); the
GPU smoke tests fail at session open on that box's NVIDIA driver
identically on unmodified origin/main — pre-existing, needs the RADV
box for on-glass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 18:53:21 +02:00
enricobuehler da259d5b79 feat(plugin-kit): browser-safe ./wire subpath for provider schemas (0.1.1)
apple / swift (push) Successful in 1m22s
apple / screenshots (push) Successful in 6m35s
ci / web (push) Successful in 1m13s
ci / docs-site (push) Successful in 1m14s
android / android (push) Successful in 17m3s
decky / build-publish (push) Successful in 18s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 8s
arch / build-publish (push) Successful in 16m33s
ci / bench (push) Successful in 14m48s
deb / build-publish (push) Successful in 14m48s
ci / rust (push) Successful in 22m36s
deb / build-publish-host (push) Successful in 18m16s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 18m30s
docker / deploy-docs (push) Successful in 23s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 28m25s
Plugin contracts share ProviderEntry/Artwork/LaunchSpec/PrepStep with their
UIs; the kit root stays node-only (reconcile keeps the HostClient transport).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:49:13 +02:00
enricobuehler 4094f6208d fix(web): load the inlang plugin from node_modules, not a CDN
audit / bun-audit (push) Successful in 13s
ci / docs-site (push) Successful in 54s
ci / web (push) Successful in 56s
apple / swift (push) Successful in 1m19s
audit / cargo-audit (push) Successful in 3m6s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 55s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
apple / screenshots (push) Successful in 6m41s
deb / build-publish (push) Successful in 10m14s
ci / bench (push) Successful in 10m59s
windows-host / package (push) Successful in 11m11s
deb / build-publish-host (push) Successful in 13m54s
android / android (push) Successful in 17m22s
arch / build-publish (push) Successful in 17m37s
ci / rust (push) Successful in 22m53s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m22s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 18m48s
docker / deploy-docs (push) Successful in 29s
`web/project.inlang/settings.json` pointed `modules` at
https://cdn.jsdelivr.net/npm/@inlang/plugin-message-format@4/dist/index.js —
inlang's stock `paraglide-js init` scaffolding, which nobody had a reason to
keep. It made `bun run codegen` need the network at build time, so in a
sandbox that has none (the Nix build; any air-gapped CI) the plugin import
failed and paraglide compiled ZERO messages.

That failure is silent by design: the inlang SDK logs a PluginImportError as a
WARNING, paraglide still prints "Successfully compiled" and exits 0, vite
bundles the empty message set, and the console only dies at request time inside
renderToReadableStream because every `m.foo()` is undefined. Reported from a
NixOS host running punktfunk-web-server.

The plugin is a normal npm package, so depend on it like one: add
@inlang/plugin-message-format as a devDependency and reference it by relative
path. `loadProjectFromDirectory` (which both `paraglide-js compile` and the
vite plugin use) imports non-`http` modules straight off disk, and the shipped
dist/index.js is a self-contained bundle with no imports of its own, so it
loads from the data: URL exactly as the CDN copy did. Compiles 268 messages
offline, with no project.inlang/cache round-trip.

Add tools/check-i18n.mjs, run at the end of `codegen`, so this can never ship
silently again: it fails the build on a remote `modules` entry, a module that
does not resolve, or a compiled message count below what messages/en.json
defines. Verified it catches all three (remote URL, missing plugin, corrupt
plugin that compiles to 0 messages) and passes a healthy build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 18:43:39 +02:00
enricobuehler 0ed8ca90f2 fix(plugin-kit): regenerate bun.lock (duplicate package path after react devDep add)
android / android (push) Has been cancelled
apple / swift (push) Has been cancelled
apple / screenshots (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:41:34 +02:00
enricobuehler 8d72e1d27e feat(plugin-kit): @punktfunk/plugin-kit 0.1.0 — the Effect plugin framework
ci / rust (push) Has been cancelled
android / android (push) Has been cancelled
apple / swift (push) Has been cancelled
apple / screenshots (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / web (push) Has been cancelled
ci / bench (push) Has been cancelled
ci / docs-site (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
deb / build-publish (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
windows-host / package (push) Has been cancelled
The ~80% of every plugin that was copy-pasted (rom-manager ↔ playnite),
extracted as one Effect-v4 package:

- runtime: definePluginKit — async-main boundary hiding a ManagedRuntime
  (two-effect-instances discipline), signal-driven interruption with a
  bounded shutdown grace
- config: Schema-driven raw round-trip (defaults only in the Schema via
  withDecodingDefaultKey + encodingStrategy omit; file stays authored-shape),
  atomic writes, world-writable refusal, changes stream
- cache-store: disposable derived state, corrupt→empty, write-through
- reconcile: kit-owned provider wire schemas + typed client over the
  skew-safe untyped pf.request seam
- sync-engine: generic poll/watch/debounce/single-flight-coalesce/
  fingerprint-skip engine with a status PubSub (the SSE feed)
- ui-server + sse: effect/unstable/httpapi behind servePluginUi with
  core-only env layers (validated by the phase-0 spikes); raw SSE route
  (httpapi has no event-stream media type)
- cli: minimal command dispatcher reusing the plugin layer graph
  (deliberately not effect/unstable/cli — it needs platform packages)
- react subpath: plugin router (path→hash→fallback deep-link restore +
  pf-ui:navigate bridge), ResultGate, sseAtom, resolvePluginBase
- theme.css: the console's violet identity packaged for plugin UIs

18 bun tests incl. the two phase-0 spikes; publish workflow mirrors
sdk-publish (tag plugin-kit-v*; 0.1.0 published manually).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:41:11 +02:00
enricobuehler c15c80718a chore(sdk): bump @punktfunk/host to 0.1.2
Manual publish (no sdk-v tag — the version was published directly so the
plugin-kit and rewritten rom-manager can consume pluginStateDir/pluginIngestDir
from the registry; the previously published 0.1.1 predates the security-tiers
merge and lacks those exports).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:41:11 +02:00
enricobuehler bb58fde369 feat(plugin-kit): scaffold @punktfunk/plugin-kit + phase-0 spikes
Two spikes validating the riskiest seams of the plugin-kit design before
implementation:

- HttpApi (effect/unstable/httpapi) served on Bun behind the SDK's
  servePluginUi using ONLY effect-core layers (Etag.layerWeak, Path.layer,
  HttpPlatform.layer + FileSystem.layerNoop) — no platform package. Auth,
  schema validation, and static fallthrough all verified end-to-end.
- Browser client prefix strategy for the console proxy: both
  transformClient+prependUrl AND baseUrl preserve the /plugin-ui/<id>
  path prefix. Nested Schema.withDecodingDefaultKey defaults confirmed
  (Effect-valued defaults; encodingStrategy "omit" restores raw shape
  on encode).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:41:11 +02:00
enricobuehler de7b54d4e0 fix(core): remaining 15 sweep lows — client, clipboard, C ABI (v10), FEC, GSO, GTK
ci / docs-site (push) Successful in 58s
ci / web (push) Successful in 1m11s
apple / swift (push) Successful in 1m20s
decky / build-publish (push) Successful in 29s
android / android (push) Has been cancelled
apple / screenshots (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / rust (push) Has been cancelled
ci / bench (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
flatpak / build-publish (push) Has been cancelled
release / apple (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
windows-host / package (push) Has been cancelled
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Has been cancelled
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Has been cancelled
windows / build (aarch64-pc-windows-msvc) (push) Has been cancelled
windows / build (x86_64-pc-windows-msvc) (push) Has been cancelled
Second and final batch from the 2026-07-20 core-sweep lows:

client:
- connect() timeout now sets `quit` before shutdown, so a handshake that
  completes after the deadline closes with QUIT_CLOSE_CODE instead of
  leaving the host lingering (virtual display up) for a reconnect that
  never comes.
- probe_result: saturating_add on the wire-supplied wire_packets +
  send_dropped counters (debug-build overflow panic / release wrap).
- the standing-latency bleed no longer sets flush_in_window: that flag
  is the ABR's SEVERE (×0.7) verdict, and the bleed fires only on
  provably loss-free windows the controller itself scores as fine.

clipboard:
- fetch_cancels pruned on every new fetch (was: one dead oneshot per
  paste for the session).
- serve chunks gated on a parked waiter + capped at CLIP_FETCH_CAP with
  an Error event (was: unbounded silent accumulation under any req_id).
- serve_inbound park bounded by FETCH_STALL_SECS + send.stopped() (was:
  an unanswered FetchRequest parked the task, waiter, and bi-stream
  forever; ~100 of them exhaust the connection's bidi budget).

C ABI (ABI_VERSION 9 → 10, header regenerated, additive only):
- new punktfunk_connection_clock_offset_now_ns — the LIVE re-synced
  offset (Swift/Kotlin latency math read the frozen connect-time value
  ~40ms wrong after a wall-clock step).
- to_config: checked u64→usize narrowing of max_frame_bytes (32-bit
  armeabi truncated >4GiB to a plausible residue).
- host_poll_input: no &mut held across the embedder callback (re-entry
  aliased it — UB under noalias); mid-drain callback clears now stick.
- next_audio_pcm: DTX (empty) payloads skipped — decode synthesized
  120ms of concealment per 5ms slot and grew the playout ring forever.
- next_clipboard releases the parked payload on an empty poll (a one-off
  50MiB paste stayed resident all session).
- frames_dropped / wants_decode_latency write their documented 0/false
  defaults before the NULL-handle check.
- gamepad constant docs match pick_gamepad() reality (DualSenseEdge/
  SwitchPro landed; DualSense/DS4 honored on Windows UMDF too).

FEC:
- gf8 reconstruct/reconstruct_into reuse the (k,m) codec cache like
  encode_into (was: fresh 230×200 generator + decode inversion per
  lossy block on the pump thread).
- vendored fec-rs: the 8 safe wrap_mul_slice shims assert equal lengths
  (x86 SIMD callees bound stores on input.len() — safe-code OOB write);
  ReconstructShard's safety contract gains the len()==get().len()
  clause and reconstruct_internal sizes raw slices from the slice.

transport/GTK:
- Linux GSO super-buffer capped at the real UDP payload ceiling (65487)
  not 65535 — seg sizes ≥1024 could EMSGSIZE and latch GSO off
  process-wide, blamed on the network.
- GTK settings dialog no longer rewrites an unlisted-but-valid stored
  gamepad preference to "auto" on close.

Already fixed on main (verified stale, skipped): the reassembler FEC
ceiling, wants_decode_latency's third term, request_probe rollback +
the generalized probe watchdog.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:26:43 +02:00
enricobuehler 76ac3cf867 fix(core): five sweep lows — seal-lane fallback, flush partial, codec echo, idle clamp, pair-name UTF-8
From the 2026-07-20 punktfunk-core quality sweep's adjudicated lows
(~/punktfunk-sweeps/punktfunk-core-2026-07-20.json), each with a
regression test:

- session: the two-lane seal's dead-worker fallback dropped the frame's
  back half and returned Ok with half an access unit; the corpse lane
  was also never respawned. The send-failure arm now reclaims the tail
  (single-lane seals the WHOLE frame) and both failure arms drop the
  lane so the next large frame respawns it. A recv-side death now
  surfaces as an error instead of silent truncation.
- reassemble: Reassembler::reset() left pending_partial parked, so a
  pre-flush stale partial survived Session::flush_backlog and was
  delivered as the first "frame" after a jump-to-live.
- caps: resolve_codec echoed a non-conformant multi-bit preferred byte
  verbatim (downstream from_wire folds it to HEVC — possibly outside
  the shared set). It now isolates one bit of the intersection.
- endpoint: stream_transport_idle only floored the value; an absurd
  operator-supplied idle timeout blew past QUIC's VarInt ms ceiling and
  panicked host startup through the expect. Clamped to 1s..1h.
- pairing: PairRequest::encode cut the device name at a raw byte-64
  boundary, splitting multi-byte UTF-8 (host showed U+FFFD forever);
  it now shares Hello's char-boundary truncate_to.

The frozen-FEC-ceiling finding (reassemble.rs:109) was already fixed on
main by the sweep's high-severity commit — skipped as stale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:16:02 +02:00
enricobuehler f4ec18d528 chore(core): fix build.rs + crate-doc drift left by the W7/W8 splits
- build.rs: drop the stale `rerun-if-changed=src/client.rs` (the file
  became client/ in W7; the path never fires).
- lib.rs: the crate doc's module list named 6 of ~19 modules — add the
  missing ones (config/error/stats/input/reject/reanchor/render_scale/
  audio/wol + the quic-gated control-plane group), without intra-doc
  links to feature-gated modules.
- quic/mod.rs: the module doc still described the deleted `msgs`
  module; list the actual post-split module set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:08:39 +02:00
enricobuehler ce51a2ba74 refactor(core): decompose the client run_pump monolith into pump/ submodules
client/pump.rs was a single 1192-line async fn stacking six concerns.
It is now a 190-line orchestrator (same spawn order, same channel
wiring) over five focused modules:

- pump/handshake.rs   — connect + Hello/Welcome/Start + skew handshake +
                        hole punch + Session construction (HandshakeOut)
- pump/input_task.rs  — the gamepad snapshot/removal/arrival re-send
                        state machine + passthrough input forwarding
- pump/control_task.rs— the control-stream select loop (renegotiation,
                        probes, bitrate acks, clock re-sync, clip
                        metadata), bundled as ControlTask
- pump/datagram_task.rs — the datagram tag demux + rumble reorder gate
- pump/data.rs        — the blocking data-plane pump (loss reports, ABR
                        + capacity probe, jump-to-live detectors,
                        standing-latency bleed), bundled as DataPump

Every body moved verbatim (dedent + arg-struct destructures aliasing
the original binding names); no per-frame indirection added — the pump
loop, its locals, and the hot-path shape are byte-equivalent. The
Ctrl/DataPump arg structs exist only to stay under too_many_arguments.

196 lib tests pass; clippy clean on --features quic AND
--no-default-features (--all-targets); include/punktfunk_core.h
byte-identical; fmt clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 18:07:14 +02:00
enricobuehler eacaaa5cd8 refactor(core): peel session.rs anti-replay, perf telemetry, and seal lane into submodules
session.rs (1203) sharpens down to the two hot-path state machines +
lifecycle (887): ReplayWindow + seq_of + their tests -> session/replay.rs;
the PUNKTFUNK_PERF trio PumpPerf/SealPerf/TimedCoder -> session/perf.rs;
the Phase-1.5 lane machinery SealLane/SealJob/seal_wire_slice/
TWO_LANE_MIN_PACKETS -> session/seal.rs. Facade pattern (session.rs stays
the parent file); pub use keeps session::{PumpPerf,SealPerf} stable and
lib.rs re-exports are untouched. Pure code motion + pub(super) bumps —
seal_frame_inner/poll_frame/poll_input bodies unchanged; the
wire-equivalence tests stay co-located with the seal path they pin.

196 lib tests pass, clippy --features quic --all-targets clean, fmt clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:56:26 +02:00
enricobuehler e0a1edb9d4 refactor(core): co-locate the quic wire tests with their modules
quic/tests.rs (1813 lines, 43 tests) was the W7 split's leftover: the
source moved into handshake/caps/control/clock/pairing/pake/datagram/
endpoint/clipstream/io but every test stayed in one monolithic file.
Each test now lives in a #[cfg(test)] mod tests at the foot of the
module it exercises, verbatim. The two CompositorPref/GamepadPref
wire/name tests moved to config.rs (where those enums live), so they
now also run under --no-default-features. The clip_loopback and
ctrl_framing integration mods share connect_pair via a cfg(test)-only
quic/test_util.rs.

Test-only motion: 196 lib tests pass unchanged on macOS, clippy
--features quic --all-targets clean, include/punktfunk_core.h
byte-identical.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:53:36 +02:00
enricobuehler a53369bf21 fix(ci): green main again — clock_sync callers, two clippy denials, an unused import
ci / web (push) Successful in 56s
ci / docs-site (push) Successful in 1m11s
apple / swift (push) Successful in 1m20s
decky / build-publish (push) Successful in 36s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 7m0s
flatpak / build-publish (push) Successful in 6m25s
deb / build-publish (push) Successful in 9m53s
docker / deploy-docs (push) Successful in 28s
release / apple (push) Successful in 9m25s
deb / build-publish-host (push) Successful in 10m0s
windows-host / package (push) Successful in 15m34s
apple / screenshots (push) Successful in 6m30s
arch / build-publish (push) Successful in 18m15s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m6s
android / android (push) Successful in 19m13s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m14s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 14m6s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m3s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 4m37s
ci / rust (push) Successful in 26m59s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 5m31s
`ci.yml`'s clippy gate has been failing on main since the pre-0.16.0 sweep landed, and
because clippy stops at the first crate it can't compile, the visible error was only
ever the first of four. `ci.yml` is not tag-triggered, so v0.16.0 cut and shipped over
a red main; none of this reaches the release artifacts (the two E0308s are in a test
binary and a dev tool, and the rest are lints) but the gate has been blind since.

Fixed, in the order clippy surfaced them:

- clients/probe failed to COMPILE (E0308, x2). `810d918d` moved `clock_sync` onto the
  resumable `io::MsgReader` but only updated the client pump, leaving the probe passing
  a bare `&mut RecvStream`. The probe now wraps the control stream in a `MsgReader` at
  `open_bi` and threads that everywhere — Welcome, the --remode and --bitrate watchers,
  and the speed-test result read — which is also what the refactor was for: those reads
  sit behind timeouts and `select!`, exactly where a straddling frame would desync the
  stream for the rest of the run.

- pf-capture: `SPA_META_Cursor as u32` is a `u32 -> u32` no-op
  (`clippy::unnecessary_cast`). Line 1274 already passes the same constant uncast, so
  the type is not in question.

- pf-inject: `noop as usize` on the SIGUSR1 wake handler added in `986402f7` trips
  `clippy::function_casts_as_integer`; goes via `*const ()` as the lint asks. Same value
  in the `usize`-typed `sa_sigaction` slot.

- punktfunk-core: the `ctrl_framing` test module's `use super::*` is unused, which
  `-D warnings` promotes to an error.

Verified with CI's own commands on a Linux box (this is all Linux-gated code, so a Mac
cannot check it): `cargo clippy --workspace --all-targets --locked -- -D warnings`
finishes clean, `cargo build --workspace --locked` succeeds, and
`cargo test --workspace --locked` exits 0 with no failures — including punktfunk-core's
196-test suite, which covers the control-stream framing the probe change touches.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 11:49:49 +02:00
enricobuehler 3213c1b61c ci(flatpak): resolve over TCP so the flathub bootstrap stops failing on tag pushes
apple / swift (push) Successful in 1m22s
apple / screenshots (push) Successful in 6m43s
ci / web (push) Successful in 50s
ci / docs-site (push) Successful in 1m17s
ci / rust (push) Failing after 7m10s
ci / bench (push) Successful in 6m57s
decky / build-publish (push) Successful in 26s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 11s
arch / build-publish (push) Successful in 12m9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 13s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
deb / build-publish (push) Successful in 8m55s
android / android (push) Successful in 17m5s
deb / build-publish-host (push) Successful in 9m25s
docker / deploy-docs (push) Successful in 23s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m5s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 14m38s
flatpak / build-publish (push) Failing after 8m1s
The flathub remote-add has now burned its entire retry budget and failed the job on
three consecutive tag builds — v0.15.0, the v0.15.0 re-point, and v0.16.0 — each one
needing a manual re-run to land the bundle.

The cause was already root-caused here on 2026-07-11 (see the Tooling step's comment):
home-runner-1's Docker embedded resolver at 127.0.0.11 drops UDP lookups while the
shared multi-org fleet is saturated. Nothing retransmits a dropped datagram, so the
lookup simply times out. The fix then was to widen retry.sh to 10 attempts (~9 min),
which comfortably outlasts a main push's ~8-workflow fan-out — but a TAG push starts
13 workflows at once, and that burst now outlives the whole budget. Retrying harder
is chasing the symptom; the transport is the problem.

`options use-vc` moves resolution to TCP, where the kernel retransmits and a query
cannot be silently lost under load. Same resolver and search path — only the
transport changes — so internal names still resolve exactly as before. Deliberately
no additional nameservers: a public resolver in the list could answer an internal
name (git.unom.io) with NXDOMAIN. retry.sh stays as the backstop for real upstream
blips.

Applied non-fatally, because Docker bind-mounts /etc/resolv.conf and may present it
read-only: a DNS tuning that cannot be applied must not be the thing that fails a
release build. If the write is refused we warn and stay on UDP, i.e. exactly today's
behaviour.

Scoped to the flatpak workflow, where the failure is actually evidenced, rather than
applied fleet-wide on suspicion.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 08:33:34 +02:00
enricobuehler b3adfb3a56 ci(release): pin Xcode DerivedData so it stops filling the mac runner's disk
apple / swift (push) Successful in 1m23s
release / apple (push) Successful in 9m19s
ci / web (push) Successful in 52s
ci / docs-site (push) Successful in 59s
apple / screenshots (push) Successful in 6m31s
ci / rust (push) Failing after 6m18s
ci / bench (push) Successful in 5m7s
deb / build-publish (push) Successful in 8m58s
decky / build-publish (push) Successful in 30s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 21s
arch / build-publish (push) Failing after 16m42s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 46s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 15s
android / android (push) Successful in 18m40s
deb / build-publish-host (push) Successful in 9m25s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m5s
docker / deploy-docs (push) Successful in 25s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m38s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 19m36s
v0.16.0's Apple leg failed at the xcframework build with "No space left on
device (os error 28)": the runner's boot volume was at 100% with 152 MB free.

None of the four `xcodebuild archive` calls passed -derivedDataPath, so each one
used the default ~/Library/Developer/Xcode/DerivedData/<name>-<hash> — and that
hash is derived from the PROJECT'S ABSOLUTE PATH. act_runner rotates its
workspace (~/.cache/act/<hash>/hostexecutor), so every rotation looked like a new
project to Xcode and minted a fresh ~760 MB tree. Nothing ever collects those:
they live outside the workspace, so act's own cleanup never sees them. 31 had
accumulated in three days (17 -> 20 July), which with the 12 GB shared
ModuleCache came to 32 GB — on a 228 GB volume already 95% full.

Pin all four archives to one path. The tree is now REUSED rather than multiplied,
which also keeps the module cache warm instead of rebuilding it per run. The new
step additionally prunes anything week-stale left in the default root, covering
both the legacy per-path trees and any other job that lands there.

Cleared by hand on home-mac-mini-1 to unblock the release (152 MB -> 31 GB free);
this is the change that stops it coming back. Note the same class of problem bit
the Windows runner the same night from the other direction — its disk-cleanup task
purges C:\t and deleted files out from under ISCC mid-pack — which is worth its
own look and is NOT addressed here.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 08:17:33 +02:00
enricobuehler 191c9a4e18 chore(release): bump workspace version to 0.16.0
audit / bun-audit (push) Successful in 18s
apple / swift (push) Successful in 1m34s
audit / cargo-audit (push) Successful in 2m10s
ci / web (push) Successful in 54s
ci / docs-site (push) Successful in 1m3s
apple / screenshots (push) Failing after 1m56s
ci / rust (push) Failing after 6m33s
ci / bench (push) Successful in 7m17s
android-screenshots / screenshots (push) Successful in 2m58s
decky / build-publish (push) Successful in 19s
release / apple (push) Successful in 11m22s
deb / build-publish (push) Successful in 11m14s
deb / build-publish-host (push) Successful in 11m17s
android / android (push) Successful in 13m10s
arch / build-publish (push) Successful in 14m6s
web-screenshots / screenshots (push) Successful in 2m57s
linux-client-screenshots / screenshots (push) Successful in 7m33s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 22m38s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 22m47s
flatpak / build-publish (push) Successful in 8m9s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 8m7s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 9m3s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 13s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 13s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 15s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 14s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 12s
docker / deploy-docs (push) Successful in 9s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 6m17s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 4m3s
windows-host / package (push) Successful in 12m35s
MINOR, not the 0.15.1 originally planned: the 24 commits since v0.15.0 carry 7
features, and they are not incidental — they change the plugin runner's privilege
model. It stops holding the console's full-admin token (scoped plugin-token lane),
the Windows runner drops from SYSTEM to LocalService, public-registry plugin
installs are gated behind an explicit flag, and the systemd user unit is sandboxed.
Operators have workflow-visible changes to absorb (a plugin genuinely needing the
admin surface must now set PUNKTFUNK_MGMT_TOKEN; installing outside the @punktfunk
scope needs --allow-public-registry). A patch label would understate all of that,
and sdk-publish pushes this version straight to consumers.

The other 12 fixes are largely the pf-encode/zerocopy/capture/inject quality sweep
landing: teardown deadlocks, error-path leaks, a GL->CUDA copy race, an odd-height
NV12 overrun, a cursor-meta OOB read, plus the two loss-recovery gaps 0.15.0 left
open (Vulkan Video never got the LTR taint sweep; QSV's sweep skipped its modal
single-swept-slot case).

`WIRE_VERSION` stays 2 and the driver protocol stays v4 — pairings, clients and
installed drivers keep working. The C ABI moves 8 -> 9 (PunktfunkFrame grows
`received_ns`, so receipt is stamped at the session boundary rather than at the
embedder's pull); the gitignored Apple xcframework needs a rebuild for that to
reach the Apple clients. No OpenAPI path changes (45 before and after) — the new
plugin-token lane is auth-layer only — so api/openapi.json moves info.version only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 07:47:10 +02:00
enricobuehler dd558be55b fix(core/client): fix the four ways a speed-test probe corrupts ABR and the report tick
ci / web (push) Successful in 1m3s
ci / docs-site (push) Successful in 1m2s
apple / swift (push) Successful in 1m19s
decky / build-publish (push) Successful in 53s
deb / build-publish (push) Successful in 12m22s
ci / bench (push) Successful in 6m13s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 44s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 6m32s
ci / rust (push) Failing after 9m2s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 11s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m22s
release / apple (push) Successful in 9m50s
windows-host / package (push) Failing after 13m7s
deb / build-publish-host (push) Successful in 13m1s
arch / build-publish (push) Failing after 14m11s
apple / screenshots (push) Failing after 3m37s
android / android (push) Successful in 14m56s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 6m8s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8m32s
flatpak / build-publish (push) Failing after 8m13s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 6m2s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 6m33s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 7m44s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m32s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m13s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 9m0s
docker / deploy-docs (push) Successful in 26s
There is one `ProbeState` per session and no correlation id, and two independent
requesters share it: the pump's startup link-capacity probe and the embedder's
`NativeClient::request_probe` (the shipped "Test connection" in the Windows GUI
and the Linux GTK app). The startup path had a busy check, a watchdog and a
window rebase; the embedder path had none of them, and the two collide by
default — both shipped speed tests call `request_probe` on the statement right
after `connect()`, and the pump's probe fires 2 s later, inside that burst.

Four defects, one cluster:

- The pump overwrote the shared slot unconditionally. Mid-burst it drops
  `base_packets`/`base_bytes`, the pump re-snapshots them against the host's
  full-burst denominator, and the user's speed test reports roughly an order of
  magnitude low; if the result lands first instead, the reset wipes `done` and
  the embedder's poll loop never sees its own measurement. Now it defers and
  retries rather than stealing the slot.

- An unanswered embedder probe never timed out. `probe_active` gates the entire
  report tick — LossReport, the ABR window feed, the standing-latency ladder and
  a pending ClockResync all live inside it. A host that ignores ProbeRequest is
  an anticipated configuration (the startup path was given a 6 s timeout for
  exactly that) so the embedder path could latch `active` forever and silently
  switch off every adaptation mechanism for the rest of the session. Now a
  watchdog covers a probe of either origin.

- `request_probe` latched `active` before `try_send`, so a full or closed ctrl
  channel returned `Closed` to the caller while leaving the session wedged in
  the state above. It now rolls back, mirroring the startup path.

- Only the startup path rebased the ABR window past the burst. Probe filler is
  counted into `bytes_received` for every accepted datagram but never reaches
  the decoder, and the tick is suppressed for the whole burst, so the first
  post-burst window read the burst rate as `actual_kbps`. That poisons
  `proven_kbps`, a monotone high-water mark that is never decayed, which
  disables the x1.5 evidence-gated climb guard for the session and authorizes a
  climb into a rate the decoder has never carried. The rebase now happens on the
  falling edge of any probe, and covers the loss/packet anchors too so the first
  post-burst LossReport isn't divided by a filler-inflated packet count.

Also: `wants_decode_latency` advertised on two of the three terms the pump
actually requires to arm ABR, omitting `resolved_bitrate_kbps > 0`, so against
an old host that reports no rate an embedder fed decode latency to a controller
that never runs.

No wire or ABI change.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 07:42:37 +02:00
enricobuehler 810d918d36 fix(core/quic): make the control-stream read cancel-safe
`io::read_msg` frames a message with two `quinn::RecvStream::read_exact` calls,
and quinn documents `read_exact` as explicitly NOT cancel-safe: the bytes it has
already taken out of the stream live only in the future's own buffer and nothing
puts them back on drop. Both long-lived control loops drive that read from a
`tokio::select!` arm — the client pump alongside `ctrl_rx.recv()` and the resync
tick, the host alongside probe/reconfig/clip-offer channels — and neither uses
`biased;`, so any sibling that becomes ready ends the iteration and drops a
partially-progressed read. `clock_sync` has the same shape via
`tokio::time::timeout`, which can fire mid-frame before the session even starts.

A control frame only has to straddle two wakeups for this to bite: a ClipOffer
carries up to 16 kinds x 128 bytes of MIME, ~2 KB, which exceeds one QUIC packet
and is subject to the pacer; any frame whose second half is lost or reordered
does it too. Losing the consumed length prefix misaligns the stream permanently
— the next read takes two payload bytes as a length, so Reconfigured,
ProbeResult, BitrateChanged, ClockEcho and ClipState all decode as garbage and
are silently dropped, and a bogus length up to 64 KiB parks the read forever.
Mode switches, adaptive bitrate, mid-stream clock resync and clipboard are dead
for the rest of the session; only a reconnect recovers, and the log shows at
most one `warn!`.

Add `io::MsgReader`, which keeps the frame in progress in the reader rather than
the future and reads via quinn's cancel-safe `read`, and switch the three
cancelling sites to it (client control loop, host control loop, clock_sync).
The sequential handshake/pairing callers keep the plain `read_msg`, whose doc
comment now states the constraint. No wire bytes and no ABI change — only how
the same length-prefixed frames are assembled.

Tests: a frame split across two wakeups with the read cancelled in between must
resume and leave the following frame correctly framed (confirmed to fail — it
hangs on the desynced stream — against the old behavior), plus a zero-length
frame round-trip.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 07:42:37 +02:00
enricobuehler 7b2cdf5a7a fix(core/net): bind the data plane to the authenticated peer + stop adaptive FEC wedging large frames
Four defects from the punktfunk-core quality sweep, all in the data plane.

transport/udp: the hole-punch adopted the source address of ANY datagram whose
first 8 bytes matched PUNCH_MAGIC — a fixed public constant with no key, nonce
or session id — and the authenticated QUIC peer was passed only as the
no-punch fallback, so it was never used to validate. Hole-punch is the default
bring-up path (it is skipped only for a fixed --data-port), and the data socket
is an OS ephemeral, so spraying the ephemeral range during the 2500 ms punch
wait let anyone steal the video plane: the legitimate client is then filtered
out by the connect() and gets nothing, while QUIC stays healthy so no reconnect
fires. With a spoofed source the same 8 bytes aim a full-rate stream at a third
party. Take the authenticated peer IP and require the punch to match it — only
the PORT is in question (that is what a NAT remaps and what the punch exists to
discover); the client dials the same host IP as its QUIC connection, so a NAT
presents one source IP for both planes. Also budget each read from the REMAINING
window, so off-peer datagrams cannot stretch the wait past punch_timeout.

transport/udp: the punch keepalive treated every send error as fatal and broke
out of its loop permanently and silently. It holds a clone of the connected,
non-blocking data socket — exactly the socket whose transient conditions this
module defines and documents in is_transient_io. One ENOBUFS from a full wlan tx
queue or an ENETUNREACH during an AP roam killed the only thing holding the NAT
mapping open; the path recovers, video keeps flowing, and the stream dies later
when the idle timer expires the mapping during a static scene. Route it through
is_transient_io like every other send site in the file.

packet: adaptive FEC moved fec_percent live (host bands it 1..=50 while Welcome
advertises 10) but the receiver's per-block acceptance ceiling was computed once
from the negotiated percentage and never re-derived. Once FEC ramped, every
packet of a maximal block failed `total > max_total_shards`, the block never
accumulated a shard, the frame aged out, and the resulting loss drove FEC higher
still — large frames wedged at 100% loss exactly when FEC was meant to rescue
the link. Fixed on both sides, because hosts and clients update independently:
the sender clamps per-block parity to the ceiling the peer negotiated, and the
receiver sizes that ceiling from the whole clamp range rather than a stale
snapshot of it.

No wire bytes and no C ABI signature change; WIRE_VERSION and ABI_VERSION are
unchanged. Regression tests cover all three (the punch tests were confirmed to
fail without the fix).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 07:42:37 +02:00
enricobuehler 15d51bc0ff Merge remote-tracking branch 'origin/main' into land/sweep-all
apple / swift (push) Successful in 1m13s
apple / screenshots (push) Successful in 6m28s
ci / web (push) Successful in 55s
ci / docs-site (push) Successful in 54s
ci / rust (push) Failing after 6m24s
decky / build-publish (push) Successful in 20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 8s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 7m12s
docker / deploy-docs (push) Successful in 15s
deb / build-publish (push) Successful in 9m0s
deb / build-publish-host (push) Successful in 9m22s
arch / build-publish (push) Successful in 17m34s
android / android (push) Successful in 18m52s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 13m56s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 17m12s
windows-host / package (push) Failing after 14m42s
2026-07-20 00:46:41 +02:00
enricobuehler e9a4c4a601 feat(security): add a user-writable plugin ingest inbox for cross-account data
ci / web (push) Successful in 48s
ci / docs-site (push) Successful in 1m1s
apple / swift (push) Successful in 1m18s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 12s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 12s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 53s
ci / bench (push) Successful in 7m12s
apple / screenshots (push) Successful in 6m21s
deb / build-publish (push) Successful in 11m27s
deb / build-publish-host (push) Successful in 12m0s
arch / build-publish (push) Successful in 13m2s
android / android (push) Successful in 15m50s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m38s
ci / rust (push) Successful in 20m35s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 20m2s
docker / deploy-docs (push) Successful in 27s
windows-host / package (push) Successful in 15m19s
The LocalService runner can't traverse the interactive user's profile the way
the old SYSTEM runner could — so a plugin can no longer read a file an app
running as *you* produced (the acute case: the Playnite exporter's library
JSON under %APPDATA%). Confirmed on-glass: as LocalService the glob finds
nothing and the profile is un-traversable.

Add the inverse of plugin-state: <config_dir>\ingest, granted BUILTIN\Users
Modify by plugins enable (disable reverts to inherited Users:RX). An
interactive-user app drops ingest\<plugin>\… and the de-privileged runner
reads it there — the one Users-writable carve-out in the otherwise
Users-read-only tree. SDK exports pluginIngestDir(name) to resolve it; on
Linux the systemd --user runner owns the config dir so same-user producers
write there with no grant.

Accepted tradeoff: the inbox is writable by any local user (trusted-single-
user model; it feeds only a LocalService runner). Consumers must treat ingest
data as lower trust than their own state.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler fdbfa37e43 fix(cli): silence host capture/DPI/GPU startup noise on management subcommands
A plain `punktfunk-host plugins add …` printed the host's startup banner and
— on Windows — ran the win32u GPU-preference hook, whose DPI-awareness probe
emits a scary `SetProcessDpiAwarenessContext … "access denied"` WARN. None of
it is relevant to a package-manager command, and the WARN reads like a failure.

Gate both behind is_management_cli(): plugins / openapi / library /
detect-conflicts / driver / web / service-management verbs skip the banner and
the DXGI hook. `service run` (the SCM-launched host) is explicitly NOT
lightweight, so the hybrid-GPU ACCESS_LOST fix it depends on stays intact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler 73ec6ed9ef feat(security): give the de-privileged runner a writable plugin-state dir
The LocalService runner cannot write anywhere under %ProgramData%\punktfunk
(the config dir is Users-read-only), so a state-writing plugin's saveCache /
config-edit / first-run mkdir all fail EPERM — proven on-glass (rom-manager
only looked fine because its state dir was pre-created by an admin run and a
0-title reconcile skipped the write).

Add the one writable grant the model was missing, keeping the split crisp —
code dirs RX+WA, secrets R, and now a dedicated state root RW:

- plugins enable / build-scripting.ps1: create %ProgramData%\punktfunk  plugin-state and grant LocalService (OI)(CI)(M); disable revokes. Users stay
  read-only, so another non-admin still can't tamper with a plugin's launch
  templates.
- SDK: export pluginStateDir(name) -> <config_dir>/plugin-state/<name>. Same
  path on Linux (the systemd --user runner owns the config dir, writable with
  no grant), so plugins use one branch-free helper.

Plugins must persist under pluginStateDir(), not straight under the config dir.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler 5cd6e8f572 docs: reflect the plugin-runner security tiers
- plugins: the Windows runner task runs as LocalService now, and
  public-registry names need --allow-public-registry
- automation + SDK README: connect()'s zero-config credential is the
  scoped plugin token; pairing administration and hook registration
  need an explicit PUNKTFUNK_MGMT_TOKEN opt-in

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler 7b7231fdbe feat(security): gate public-registry plugin installs + sandbox the runner unit
Supply chain: resolvePackage() used to pass any punktfunk-plugin-* or
foreign-scope name straight to bun add — a typo or a squatted look-alike
on npmjs.org would install operator-privileged code. Only the @punktfunk
scope (pinned to the Gitea registry by the bunfig scope map) resolves by
default now; anything else throws with an explanation unless
--allow-public-registry is passed, and even then installs print a loud
warning. Removal never gates — uninstalling stays safe regardless of
origin.

systemd (user unit): free hardening for well-behaved plugins —
NoNewPrivileges, PrivateTmp, ProtectSystem=strict with ReadWritePaths=%h
(plugin state and download dirs under $HOME keep working), and
RestrictAddressFamilies=AF_UNIX AF_INET AF_INET6. Plugins writing
outside $HOME, and distros that restrict unprivileged user namespaces
for user units, are handled via a documented systemctl --user edit
drop-in.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler 7fe07d0fa5 feat(security): enforce the unit-file trust rule on Windows too
fileIsSafe() returned true unconditionally on win32 ('ACL check is a
follow-up'), leaving the best-effort directory DACL as the only guard on
what the runner imports and executes. The follow-up: before importing a
unit, read its SDDL (Get-Acl via the full-System32-path powershell) and
refuse loudly — exactly like the Unix mode-bits path — unless the owner
is SYSTEM/Administrators/TrustedInstaller (or the account the runner
itself runs as, mirroring Unix's 'your own file is fine') and no other
principal holds a write-capable allow ACE (write/append data, EA/attrs,
DELETE, WRITE_DAC, WRITE_OWNER, generic write/all).

The SDDL verdict is a pure exported function with platform-independent
tests; unknown ACE shapes and unreadable ACLs fail closed. Inherit-only
ACEs (templates for children) and deny ACEs don't trip it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler 84c47cd0a7 feat(security): run the Windows plugin runner as LocalService, not SYSTEM
The PunktfunkScripting scheduled task ran operator-installed plugin code
as SYSTEM with -RunLevel Highest — any plugin defect was a full
compromise of the box. The principal is now NT AUTHORITY\LocalService
(minimal privileges, no password to manage, exists at boot, loopback
networking works), registered without RunLevel:

- installer: New-ScheduledTaskPrincipal -UserId 'LocalService'
- plugins enable: converges the principal idempotently (migrating tasks
  an older installer or a dev box registered as SYSTEM) BEFORE starting,
  then grants LocalService read — via icacls, full-System32-path — on
  exactly the two SYSTEM/Admins-DACL'd files the runner's connect()
  needs: the scoped plugin-token and the TLS-pin cert.pem. Never
  mgmt-token. plugins disable revokes the grants; plugins status now
  prints the task principal so the migration is verifiable.
- build-scripting.ps1 mirrors the convergence + grants on dev deploys.

The usage text also mentions the new --allow-public-registry gate that
lands with the supply-chain commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler 9b5ca0eff3 feat(security): scope the plugin runner to a capability-limited bearer token
The scripting runner used to hold the console's full-admin mgmt-token — a
plugin defect could rewrite hooks.json (arbitrary command execution) or
administer pairing. The host now mints a second persisted secret,
plugin-token (PUNKTFUNK_PLUGIN_TOKEN), and require_auth grows a third lane
for it: loopback-confined like the admin token, but plugin_may_access
carves out /hooks (read AND write), everything under /pair, /native/pair,
/native/pending, client unpair DELETEs, and other plugins' ui-credential.
Everything a plugin legitimately does (status/library/events/sessions,
its own UI lease) is untouched.

The SDK's zero-config connect() now prefers PUNKTFUNK_PLUGIN_TOKEN /
plugin-token over mgmt-token, so plugins pick the scoped credential up
automatically; a script that genuinely needs the admin surface sets
PUNKTFUNK_MGMT_TOKEN explicitly. Old hosts without a plugin-token fall
back to mgmt-token unchanged. No OpenAPI change: the lane is auth-layer
only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler ecec7cc062 feat(deploy/windows): always deploy the plugin/script runner
`punktfunk-host plugins add/remove/list` forward to the bun runner, which on
Windows is resolved relative to the running exe (<exe-dir>\bun\bun.exe +
<exe-dir>\scripting\runner-cli.js). Only the installer ever laid that payload
down, so on a deploy-host.ps1 dev box — where the service runs out of
target\release — both paths are absent and the CLI bails with "the plugin
runner isn't installed". deploy-all.ps1 was host + web console only.

Add build-scripting.ps1: it mirrors CI (bun install --frozen-lockfile
--ignore-scripts + bun build src/runner-cli.ts --target=bun, gated on the same
`attempt=` sentinel that proves the dynamic plugin import stayed a runtime
import), then lays the bundle + scripting-run.cmd + the shared bun next to every
host exe it finds — the built one and whatever the PunktfunkHost service runs —
so it is correct on a dev checkout or an installed {app}. It stops a running
PunktfunkScripting first (a live runner holds bun.exe open) and does not
silently enable the opt-in task (-EnableTask to do so).

deploy-all.ps1 is now host -> web console -> runner, always, so the host binary
and the runner bundle never drift apart.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 00:45:27 +02:00
enricobuehler 1e8b267d93 Merge branch 'fix/vulkan-open-leak' into land/sweep-all 2026-07-20 00:43:19 +02:00
enricobuehler dafab58943 Merge branch 'fix/round1-highs' into land/sweep-all 2026-07-20 00:42:38 +02:00
enricobuehler 1dcba4dffa Merge branch 'fix/encode-medium-tier' into land/sweep-all 2026-07-20 00:42:32 +02:00
enricobuehler a784682d4c fix(client/net): split the receipt stamp from the pull + bleed standing latency
ci / docs-site (push) Successful in 56s
apple / swift (push) Successful in 1m14s
ci / web (push) Successful in 2m30s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 9s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 9s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 7s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 9s
ci / bench (push) Successful in 5m57s
release / apple (push) Successful in 9m8s
deb / build-publish (push) Successful in 9m33s
arch / build-publish (push) Successful in 13m14s
docker / deploy-docs (push) Successful in 23s
flatpak / build-publish (push) Failing after 8m1s
deb / build-publish-host (push) Successful in 10m24s
android / android (push) Successful in 14m56s
apple / screenshots (push) Successful in 6m43s
ci / rust (push) Successful in 19m29s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m35s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m29s
windows-host / package (push) Successful in 15m13s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 5m48s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 6m38s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 7m45s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 8m51s
The two-pair investigation (wired Mac clients stuck at a rock-steady
~18-19 ms "network" that survived the load ending and cleared only on
reconnect) exposed two structural gaps, one of measurement and one of
recovery:

- Receipt was stamped at the hand-off PULL (Swift nextAU, pf-client-core,
  Android decode loops), not at reassembly completion — so any client-side
  standing state between the reassembler and the pull read as NETWORK
  latency, undiagnosable from the HUD. ABI v9: `PunktfunkFrame`/`Frame`
  grow `received_ns`, stamped by `Session::poll_frame` as the AU crosses
  the session boundary. Every embedder now uses the core stamp; the Apple
  client keeps the pull instant as `AccessUnit.pulledNs` and shows the
  receipt→pull wait as its own "client queue" term (detailed HUD tier from
  2 ms + a `queue_p50` stats-log field). Decode stages keep their pull
  anchor on all platforms, so no historical stage shifts meaning.

- The jump-to-live detectors deliberately ignore anything under 6 queued
  frames / 400 ms behind — so a small, constant, loss-free elevation (a
  sub-frame standing backlog, or a stale clock offset after a wall-clock
  step/slew) is carried for the rest of the session. New third detector
  (`StandingLatency`, unit-tested ladder): window-MIN one-way delay
  ≥ 10 ms above the session floor with zero loss for ~4.5 s escalates
  gently — a free clock re-sync first (an applied re-sync re-bases the
  floor), then at most 3 flush+keyframe bleeds sharing the jump-to-live
  cooldown, then a loud disarm naming what it means. Loss windows reset
  the run: congestion belongs to FEC/ABR, not this detector.

Also: mid-stream re-sync apply/discard logs debug→info — they are the
forensic trail for the stale-offset case and were invisible in the field.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 23:38:11 +02:00
enricobuehler 134874b3c8 fix(client/windows): repair the v0.15.0 MSIX build — ARM64 feature leak + illegal XML comment
ci / docs-site (push) Successful in 52s
ci / web (push) Successful in 54s
decky / build-publish (push) Successful in 24s
apple / swift (push) Successful in 1m20s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 12s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 15s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 9s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 8s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 6m21s
ci / bench (push) Successful in 6m57s
flatpak / build-publish (push) Successful in 6m21s
deb / build-publish (push) Successful in 9m54s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 6m58s
android / android (push) Successful in 13m35s
apple / screenshots (push) Has been cancelled
arch / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
ci / rust (push) Failing after 20m38s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 7m56s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m9s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 20m26s
docker / deploy-docs (push) Successful in 25s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 9m18s
Both MSIX matrix jobs failed at the v0.15.0 tag, for two unrelated reasons, both
introduced by this release. 0.13.0 and 0.14.0 shipped both packages; 0.15.0 shipped
neither.

ARM64 failed to BUILD. The PyroWave Windows un-gating (1ef0229b) moved `pyrowave-sys`
in pf-client-core out of `cfg(target_os = "linux")` into a general optional dependency
that pf-client-core's own `default = ["pyrowave"]` turns on. The ARM64 leg builds
`--no-default-features` precisely to skip it, but that only disables defaults for the
two packages named with `-p` — every internal edge kept inheriting them:
clients/windows -> pf-client-core, clients/session -> pf-client-core, clients/session
-> pf-presenter, and pf-presenter's own `default = ["pyrowave"]` -> pf-client-core.
So the vendored Granite C++ compiled on ARM64 and stopped where it always does:
muglm.cpp reaching for _MM_TRANSPOSE4_PS/_mm_storeu_ps, then
`simd.hpp:103 #error "Implement me."`.

Those three internal edges now pass `default-features = false`; the feature is enabled
only through an explicit chain (clients/session's `pyrowave` -> pf-client-core/pyrowave
+ pf-presenter/pyrowave, and pf-presenter's `pyrowave` -> pf-client-core/pyrowave), so
`--no-default-features` finally means what it says. Verified by feature resolution
rather than assertion — `cargo tree -i pyrowave-sys`:

  aarch64-pc-windows-msvc, --no-default-features -> not in the graph (was: present)
  x86_64-pc-windows-msvc,  default features      -> present
  x86_64-unknown-linux-gnu, default features     -> present

PyroWave is NOT dropped from the Windows client. Decode has always run in the spawned
punktfunk-session binary, whose own default keeps the feature on, and unification turns
it back on for the shared pf-client-core wherever that binary is in the build (x64,
Linux). The shell only ever offered "pyrowave" as a codec preference string — it holds
no decoder and never called one, so linking the C++ into it was dead weight even on x64.

x64 built fine and failed at PACKAGING: `makepri new` rejected the manifest with
PRI191 "Appx manifest not found or is invalid" / malformed comment syntax. The console
tile added in 2508b720 documented its flag inside an XML comment, and XML forbids `--`
anywhere in a comment body, so the literal flag spelling made the whole manifest
unparseable. Reworded, with a note to keep the next person from reintroducing it.
Confirmed by parsing both revisions after template substitution: the tagged manifest
fails at line 63 col 9, this one parses.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 23:20:29 +02:00
enricobuehler 134fba1424 fix(inject): heap the SwDeviceCreate callback context; stop latching a pad slot on a failed create
Two medium findings from the round-1 sweep, each applied to both siblings.

- create_swdevice stack-allocated the SwCreateCtx that the async PnP completion
  callback writes through (result + up to 127 u16 of instance id) and then
  SetEvents. The wait is bounded at 10s, so on a wedged-PnP timeout the callback
  can still be PENDING: the frame is popped, the input thread reuses that stack,
  and a late callback corrupts it and SetEvents an already-closed (possibly
  recycled) handle. The context is now heap-allocated and reclaimed only where
  the callback provably ran; on the timeout path the box is deliberately leaked
  and the event left open, so a late write always targets live memory. Costs a
  one-off ~264 B + one HANDLE on that rare path. Applied to the DualSense path
  and its XUSB sibling in gamepad_windows.rs.

- Ds4WinPad::open swallowed a create_swdevice failure into a WARN and returned
  Ok with no devnode. PadSlots::ensure then stored Some(pad) AND called
  gate.on_success(), so the slot short-circuited on is_some() forever and the
  capped-backoff retry that exists precisely to self-heal a transient PnP failure
  never ran — the game saw no controller for the rest of the session unless the
  client unplugged the pad. Now propagates, matching the XUSB sibling. Same fix
  applied to steam_deck_windows.rs.

Windows .173: pf-inject 53/0. Linux .21: pf-inject 74/0 (8 ignored).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 23:19:19 +02:00
enricobuehler fe4af1761e fix(capture): stop stranding a PipeWire buffer on a caught panic; invalidate the PyroWave CSC on ring recreate
Two medium findings from the round-1 sweep.

- The `.process` callback dequeued a buffer INSIDE its `catch_unwind`, and every
  requeue site was inside too. `newest` is a raw pointer with no Drop, so any
  caught panic (update_cursor_meta / consume_frame) unwound past all three
  requeues and permanently stranded that buffer. Once the stream's fixed pool
  drained, `dequeue_raw_buffer` returned null every call and capture silently
  wedged while still reporting negotiated/active — defeating the very
  catch_unwind that was meant to keep a panic survivable. The drain loop now runs
  OUTSIDE the catch (dequeue/queue are non-panicking C FFI pointer ops) and
  `newest` is requeued exactly once after it, on every path: normal,
  corrupted-skip, or caught panic.

- `recreate_ring` invalidated `video_conv` and `hdr_p010_conv` but not
  `pyro_conv`. That converter is mode-baked — BgraToYuvPlanes selects entirely
  different shaders and output formats for SDR (8-bit BT.709 → R8/R8G8) vs HDR
  (scRGB→PQ BT.2020 → R16/R16G16) — and `ensure_pyro_conv` only builds when None,
  so a display_hdr flip reused the stale SDR converter against a freshly
  HDR-formatted pyro ring, corrupting every frame. Reachable at the documented
  "Downgrade point D": a PyroWave session with client_10bit=true that opens on a
  box where HDR can't enable, then flips once the display comes up.

Linux .21: pf-capture 1/0. Windows .173: `cargo check -p pf-capture` clean.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 23:16:25 +02:00
enricobuehler 26ff005e70 fix(zerocopy): error-path leaks, exception-safe fence wait, odd-height NV12 overrun, worker reaper deadline
Seven medium findings from the round-1 sweep, all in pf-zerocopy. Adjudicated
against source; all zero-copy-preserving (no GPU→CPU→GPU roundtrip introduced).

- import_src (vulkan.rs) leaked a VkBuffer + VkDeviceMemory + dup'd fd on every
  fallible step before the success-only src_cache.insert. A failed import is
  survived by the worker and RETRIED by the caller every frame, so this leaked
  per frame for the worker's whole lifetime. Now owns the dup fd via OwnedFd
  (closes on early return until Vulkan consumes it at allocate_memory) and
  destroys the buffer + frees memory on each error path.

- ensure_dst destroyed the old exportable buffer and nulled self.dst BEFORE
  building the replacement, so a failed rebuild both dropped the working buffer
  and leaked every object the partial rebuild created (raw ash handles, no Drop;
  VkBridge::drop only frees the live self.dst). Now builds fully, unwinds each
  error locally, and swaps only on success.

- The submit→wait→reset sequence in import_linear_nv12 / import_linear `?`-ed out
  of wait_for_fences BEFORE reset_fences on a TIMEOUT/DEVICE_LOST, leaving the
  shared self.cmd PENDING and self.fence IN-USE while the caller retries on the
  same bridge (and ensure_dst later destroys dst.buffer assuming nothing is in
  flight). Now drains the GPU (device_wait_idle) and resets the fence before
  propagating — fixing the reported ensure_dst UAF at its source.

- copy_pitched_nv12_to_buffer writes height.div_ceil(2) chroma rows into a UV
  plane sized at height/2, so an odd height overruns by one uv_pitch row (OOB
  device write / CUDA_ERROR_ILLEGAL_ADDRESS). Added the even-dimension guard to
  import_linear_nv12, matching Nv12Blit/Yuv444Blit.

- sweep_reaper only reaped already-exited workers, so a worker wedged inside a
  driver call lingered forever holding its CUcontext + BufferPool (hundreds of MB
  VRAM). Now force-kills a child parked past REAPER_KILL_DEADLINE (20s).

Compile + tests green on Linux .21 (RTX 5070 Ti): pf-zerocopy 17/0. The error
paths themselves are not fault-injected; the fixes are structural.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 23:10:06 +02:00
enricobuehler 986402f731 fix(inject,zerocopy,capture): teardown deadlock, GL→CUDA copy race, cursor-meta OOB read
The three high-severity defects from the round-1 sweep of pf-inject /
pf-zerocopy / pf-capture (adjudicated against source — all 7 reported criticals
in these crates downgraded; these were the real highs).

- pf-inject steam_gadget: `SteamDeckGadget::drop` set `running=false` then joined
  the control thread, which spends steady state parked in a blocking, no-timeout
  `EVENT_FETCH` ioctl that only tests `running` at its loop top. The flag never
  reaches it, closing the fd can't wake an in-flight ioctl (the syscall holds a
  file reference, and the fd is shared via Arc by the very threads being joined),
  so the join hung — and it runs on the session input thread via
  `PadSlots::sweep`, driven by the client's `active_mask`, so a remote peer
  clearing its pad bit could freeze all session input. Now wakes the parked
  threads with SIGUSR1 (no-op, non-SA_RESTART handler → the ioctl returns EINTR
  and the loop exits), retried until each reports done and bounded (~1s).

- pf-zerocopy cuda: the GL→CUDA "sync point" was never established for the copy.
  `cuGraphicsMapResources`/`UnmapResources` were issued on the NULL stream, but
  the D2D copy runs on `copy_stream()`, a `CU_STREAM_NON_BLOCKING` stream exempt
  from implicit NULL-stream ordering — and the GL de-tile/CSC that produced the
  texture ends with only `glFlush` (no fence). So the copy could race ahead of
  the not-yet-retired GL draw: intermittent stale/torn frames under GPU load, on
  the default NVIDIA capture→encode path. Map, copy, and unmap now share
  `copy_stream()`, so map's device-side guarantee orders the GL work before the
  copy. Zero-copy preserved (no GPU→CPU→GPU roundtrip).

- pf-capture cursor meta: `update_cursor_meta` trusted three producer-written
  fields (bitmap_offset, pixel offset, stride) with no bound against the metadata
  region, driving OOB pointer arithmetic and an oversized `from_raw_parts` — an
  OOB read that SIGSEGVs inside the PipeWire `.process` callback (uncatchable by
  the surrounding `catch_unwind`). Switched to `spa_buffer_find_meta` to obtain
  the region's real `size` and validate every offset with checked arithmetic
  before each deref/slice, mirroring the fd-length guard the main frame path
  already applies.

Compile + existing tests green on Linux .21 (real RTX 5070 Ti): pf-inject 74/0,
pf-zerocopy 17/0, pf-capture 1/0. The gadget deadlock path only executes on a
SteamOS host with raw_gadget/dummy_hcd (not reproducible on the CachyOS box), so
that fix is reasoned + compile-verified, not runtime-exercised.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 22:53:15 +02:00
enricobuehler 28491acf2a fix(encode): unwind Vulkan Video open failure instead of leaking every prior object
VulkanVideoEncoder::open_inner creates ~20 Vulkan objects across ~15
fallible steps, but all cleanup lived in the encoder's Drop — which only
runs once the value exists at the final Ok(Self). Any earlier ?/bail!
leaked everything built so far (a VkDevice + GPU memory per retried
open, and this backend is the default encode path on AMD/Intel Linux
hosts where open can fail transiently).

Factor the entire teardown sequence — unchanged — into a VkTeardown
guard whose Drop destroys any prefix of the build (vkDestroy*/vkFree*
are defined no-ops on VK_NULL_HANDLE): open_inner mirrors each object
into the guard as it is created and disarms it only at Ok(Self); the
encoder's own Drop rebuilds one from its fields, so both paths share
one sequence and cannot drift. make_frame now builds in place into a
guard-parked null Frame so a mid-build failure unwinds its partial
handles too, and make_video_image / vk_util::make_plain_image (also
used by the PyroWave backend) / build_parameters_h265 no longer leak
their own partially-created objects on failure.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 22:15:44 +02:00
enricobuehler b0fbb80fd5 fix(encode): key the PyroWave plane-import cache on the capturer's ring generation
Completes the partial fix from the previous commit. The Windows PyroWave backend
caches its imported plane images keyed on the D3D11 texture's COM address, and
holds no reference on that texture — so once the capturer recreates its ring,
those addresses can be handed straight back out by the allocator and a
pointer-keyed cache hit returns an image bound to a texture that no longer
exists. Adding the extent to the key ruled out same-address-different-size
aliasing, but a recycle at identical dimensions still aliased.

The capturer already tracks exactly the value needed: `generation`, bumped on
every ring recreate. Plumbed it onto `PyroFrameShare` and the encoder now
flushes every cached import when it changes, which makes cache identity
independent of allocator behaviour rather than a bet against pointer reuse.

Validated on the RTX box: `pyrowave_win_smoke` (forced with `--ignored`, the
only test that actually exercises this path on real hardware) passes all ten
configurations — 1024²/720p/1080p/1440p across SDR/HDR and 4:2:0/4:4:4 — with
correct decoded chroma means, so the steady-state cache-hit path still works.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 21:45:14 +02:00
enricobuehler 6f52397342 fix(encode): bound NVENC async pipelining by the capturer's texture ring
The Windows direct-NVENC backend registers and encodes the capturer's textures
IN PLACE (no CopyResource), so how deep it may pipeline is a property of the
CAPTURER, not of the encoder. It was bounded only by `async_inflight_cap()` —
`PUNKTFUNK_NVENC_ASYNC_DEPTH`, default 4, clamped to the output-bitstream pool
— which consults nothing about the capturer, while the comment at the
backpressure loop claimed it "keep[s] in-flight depth within the capturer's
texture ring". It never did.

The IDD-push capturer rotates `OUT_RING = 3` per delivered frame with no regard
for encode completion (its own invariant note says OUT_RING(3) > max
pipeline_depth(2)). With the default async depth of 4 the encoder can therefore
still be reading a texture the capturer has already handed out again and
overwritten: torn or mixed frames. It is visual corruption rather than UB, so it
fails silently and intermittently — the worst shape to diagnose from a field
report.

Adds `Encoder::set_input_ring_depth`, reported from `Capturer::pipeline_depth`,
and bounds the async backpressure loop by `min(async_inflight_cap(), depth)`.
For IDD-push that yields 2, matching the capturer's stated contract; backends
that copy their input, or are synchronous, ignore it.

Wired at ALL THREE encoder-creation sites (initial open, stall/resize rebuild,
ABR rebuild) and forwarded through `TrackedEncoder` — this crate has a
documented trap where an unforwarded defaulted trait method silently no-ops
through that wrapper, which has already bitten the direct-NVENC work once and
the wire-chunking probe once.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 21:42:53 +02:00
enricobuehler a38adad943 fix(encode): bound GPU waits, validate encode status, repair command-buffer and cache invariants
Five of the nine medium findings from the pf-encode sweep. The remaining four
need cross-crate plumbing or an unwind refactor and are deliberately left out.

- vulkan_video `enqueue` waited on the backpressure fence with `u64::MAX`. That
  wait runs ON the host encode thread — the same thread the stall watchdog's
  `reset()` would run on — so a wedged GPU parked the one thread that could
  recover the session: no error, no reset, and teardown blocking on the join.
  This is the DEFAULT encode path for AMD/Intel Linux hosts (both shipped build
  recipes enable `vulkan-encode` and `vulkan_encode_enabled()` defaults true).
  Now bounded by ENCODE_FENCE_TIMEOUT_NS with expiry surfaced as an error.

- vulkan_video `import_cached` evicted a cached dmabuf import and destroyed its
  image/view/memory with no fence wait, while up to `ring_depth - 1` submitted
  frames may still reference it — a GPU-side use-after-free. `Drop` and `reset`
  both idle first; this was the one unguarded destroy. Now idles before the
  eviction loop, guarded on the length test so the steady state pays nothing.

- vulkan_video `read_slot` never asked for the encode's operation status, so a
  FAILED encode was indistinguishable from a successful one and its feedback
  was read as if it described real bitstream. Now requests WITH_STATUS_KHR and
  refuses anything that is not COMPLETE.

- linux/pyrowave `encode_frame` opens its recording window early and has six
  fallible steps inside it; every one returned with `cmd` still RECORDING, and
  nothing repaired it (one `begin_command_buffer` in the file, and neither
  `reset()` nor `Drop` touches `cmd`), so the next frame called `begin` on a
  recording buffer — invalid usage. `submit` now resets the buffer on error;
  legal on all these paths since the pool carries RESET_COMMAND_BUFFER and the
  buffer is not pending.

- windows/pyrowave `encode_frame` ignored `frame.width`/`height` and imported
  planes at the encoder's configured extent, so a ring recreate at a new mode
  (the IDD capturer does this autonomously on a confirmed descriptor change)
  read the planes under a stale VkImageCreateInfo. Added the size guard its QSV
  and AMF siblings already carry, and keyed the plane cache on
  (address, width, height) so a recycled COM address cannot resurrect an import
  of a different size. NOTE: a recycle at the SAME size is still theoretically
  possible; the complete fix keys on the capturer's ring generation and needs
  that plumbed onto `PyroFrameShare`.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 21:34:16 +02:00
enricobuehler b22d0da75b fix(encode): port the RFI taint sweep to Vulkan Video, close the QSV sweep hole, bounds-check encode feedback
apple / swift (push) Successful in 1m19s
apple / screenshots (push) Successful in 6m24s
ci / web (push) Successful in 49s
ci / docs-site (push) Successful in 51s
ci / bench (push) Successful in 6m55s
deb / build-publish (push) Successful in 9m12s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 34s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
arch / build-publish (push) Successful in 17m2s
docker / deploy-docs (push) Successful in 14s
android / android (push) Successful in 17m59s
deb / build-publish-host (push) Successful in 11m19s
ci / rust (push) Successful in 25m3s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 15m52s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 15m54s
windows-host / package (push) Failing after 13m20s
Three defects from the pf-encode sweep, each adjudicated against source.

- Vulkan Video never received fecbec2d's taint sweep (it was carved out one
  commit later). `pick_recovery_slot` accepts any resident slot whose wire is
  below the CURRENT loss start, but "resident and older than this loss" is not
  "the client decoded it": after an earlier loss [a,b] recovered at wire r,
  everything in [a, r-1] is undecodable at the client — the lost frames plus
  every frame that predicted through the gap — and those wires stay eligible
  until the 8-slot ring rolls them out. A later loss could therefore anchor on
  one and ship it tagged `recovery_anchor`, which is the client's definitive
  re-anchor signal (punktfunk-core/src/reanchor.rs): the host lifts the client's
  post-loss freeze onto a picture built from a reference it never had. Swept
  before anchor selection, matching AMF/QSV.
  `slot_wire` is blanked and `slot_poc` deliberately is NOT: `slot_poc` feeds
  `build_h265_rps_s0`, which must keep naming every physically-resident DPB
  picture or a conforming decoder evicts them and the anchor then references a
  picture the client already dropped.

- QSV's sweep was incomplete, and in its MODAL case. `ltr_slots` mirrors the
  hardware DPB, but nulling an entry issues no VPL call — the frame stays marked
  long-term until that LongTermIdx is re-marked or an IDR flushes it (amf.rs
  states this verbatim). The rejection loop iterates the post-sweep mirror and
  only rejects `Some` slots, so it silently skipped the single entry the sweep
  exists to distrust, leaving the recovery frame free to predict from it. With
  NUM_LTR_SLOTS=2 the "exactly one slot swept" case is the common one, and the
  two existing tests cover only the both-survive and both-swept cases. Taint is
  now recorded in `ltr_tainted` with the FrameOrder left in place, so anchor
  selection and the queued-force guard skip it while the rejection list still
  names it.

- `read_slot` built a slice from the driver-reported (offset, bytes-written)
  encode feedback with no validation against `bs_size`, so a driver reporting a
  range outside the bitstream buffer produced an out-of-bounds read shipped
  straight onto the wire. Checked in u64 (so the add cannot wrap) before
  `map_memory`, so the error path has no unmap to unwind.

Adds `taint_sweep_excludes_slots_from_an_earlier_loss` covering the two-loss
case the existing single-loss test does not reach.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 21:18:17 +02:00
enricobuehler d398e2296f chore(release): bump workspace version to 0.15.0
apple / swift (push) Successful in 1m16s
audit / bun-audit (push) Successful in 16s
audit / cargo-audit (push) Successful in 2m12s
ci / web (push) Successful in 54s
ci / docs-site (push) Successful in 1m1s
ci / bench (push) Successful in 4m54s
release / apple (push) Successful in 10m19s
apple / screenshots (push) Successful in 5m19s
android-screenshots / screenshots (push) Successful in 3m37s
ci / rust (push) Successful in 24m3s
decky / build-publish (push) Successful in 20s
windows / build (aarch64-pc-windows-msvc) (push) Failing after 5m31s
android / android (push) Successful in 15m3s
deb / build-publish (push) Successful in 10m56s
arch / build-publish (push) Successful in 14m15s
deb / build-publish-host (push) Successful in 10m49s
web-screenshots / screenshots (push) Successful in 3m9s
linux-client-screenshots / screenshots (push) Successful in 7m19s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 9m6s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 16m19s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 16m20s
docker / deploy-docs (push) Successful in 24s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 12s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 11s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 11s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 10s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 11s
windows-host / package (push) Successful in 16m50s
flatpak / build-publish (push) Successful in 7m18s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Failing after 5m2s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Failing after 6m10s
MINOR, not patch: 44 commits since v0.14.0 split 21 features / 19 fixes, and
/api/v1 gained the per-scanner library-toggle endpoints (additive, already
regenerated into api/openapi.json by c2bba134). sdk-publish pushes this version
to consumers, so a patch bump would understate a new API surface to anyone
pinning ~0.14.

Headline work is the desktop clients drawing level with the Apple revamp: the
Windows client picked up settings parity, a findable console UI in the header,
the shared clipboard (with a per-host toggle), PyroWave decode in the codec
picker, and D3D11VA-first decode + HDR pass-through on Intel; the Linux GTK4
client got the same category-map settings rebuild. Apple landed the intent-based
presenter rebuild and the Dynamic Island redesign. Also: the punktfunk-host
plugins CLI, per-scanner library toggles in the console, and PyroWave raw-dmabuf
zero-copy capture on the Linux NVIDIA host.

Notable fixes: LTR-RFI loss recovery under sustained loss, two encode teardown
memory-safety holes, the audio first-open retry that was leaving sessions
silent, and the GameStream stream-marker announcement on the compat plane.

Every workspace crate is on version.workspace = true, so this stayed a one-line
bump plus the lock sync. (fec-rs, pf-driver-proto, usbip-sim, the Windows driver
crates and pf-vkhdr-layer are deliberately versioned independently and stay put.)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 20:57:58 +02:00
enricobuehler f9668b16a1 fix(encode): NVENC partial-init session leak + three backend-parity gaps
apple / swift (push) Successful in 1m18s
apple / screenshots (push) Successful in 4m33s
android / android (push) Has been cancelled
arch / build-publish (push) Has been cancelled
ci / web (push) Has been cancelled
ci / docs-site (push) Has been cancelled
ci / bench (push) Has been cancelled
ci / rust (push) Has been cancelled
deb / build-publish (push) Has been cancelled
deb / build-publish-host (push) Has been cancelled
decky / build-publish (push) Has been cancelled
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Has been cancelled
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Has been cancelled
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Has been cancelled
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Has been cancelled
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Has been cancelled
docker / deploy-docs (push) Has been cancelled
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Has been cancelled
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Has been cancelled
windows-host / package (push) Has been cancelled
All four found in the pf-encode quality sweep and verified against source.

- NVENC partial-init leak (BOTH platforms, high): `init_session` publishes
  `self.encoder` — and on Windows charges LIVE_SESSION_UNITS — *before* its
  remaining fallible steps (bitstream buffers; on Linux also the input-surface
  alloc and `register_resource`). A failure there left a live session with
  `inited == false`, and every guard on the re-init path keys off `inited`, so
  the next submit skipped teardown and overwrote `self.encoder`: the session
  leaked permanently toward the driver's per-process cap, and its budget units
  never returned, progressively starving parallel-display admission. `teardown`
  already keys off `encoder.is_null()` rather than `inited`, so it cleans up
  exactly this half-built state — it just was never called. Now invoked on the
  `init_session` error path on both platforms.

- `can_encode_10bit` asked the wrong backend (medium): it resolved via
  `linux_auto_is_vaapi`, which ignores `encoder_pref`, while `can_encode_444`
  and `open_video` honour it. On a host that forces a backend (e.g.
  `encoder_pref = "vaapi"` on an NVIDIA box) the probe answered for NVENC while
  the session opened VAAPI, so the negotiated bit depth — and the HDR/SDR colour
  label derived from it — described a backend that never ran. Now uses the same
  `linux_zero_copy_is_vaapi` mirror, and `linux_auto_is_vaapi` carries a warning
  that it resolves the `auto` case only and is not a dispatch mirror.

- Linux software arm ignored SW_BITRATE_CEIL (low): the Windows arm clamped
  openh264 to 100 Mbps, the Linux arm passed the full negotiated rate. The
  constant is now module-scope so both arms share one value.

- QSV/AMF env-parity (low): `PUNKTFUNK_IR_PERIOD_FRAMES` was a no-op on QSV
  despite the comment claiming parity with AMF, and `PUNKTFUNK_NO_QSV_LTR` /
  `PUNKTFUNK_INTRA_REFRESH` had dropped AMF's `trim()` and `yes`/`on` spellings,
  so a value with stray whitespace silently did nothing on Intel while the same
  value worked on AMD.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 20:47:55 +02:00
205 changed files with 23908 additions and 4383 deletions
+28 -2
View File
@@ -73,8 +73,34 @@ jobs:
# sufficient — the Tooling step's dnf install pulls a systemd package upgrade whose RPM # sufficient — the Tooling step's dnf install pulls a systemd package upgrade whose RPM
# trigger re-runs authselect and regenerates this file, undoing the fix. It's reapplied # trigger re-runs authselect and regenerates this file, undoing the fix. It's reapplied
# there, right before the first `flatpak` network call. # there, right before the first `flatpak` network call.
- name: Fix container DNS (drop nss-resolve — no systemd-resolved in CI) - name: Fix container DNS (drop nss-resolve, resolve over TCP)
run: sed -i 's/resolve \[!UNAVAIL=return\] //' /etc/nsswitch.conf run: |
sed -i 's/resolve \[!UNAVAIL=return\] //' /etc/nsswitch.conf
# Resolve over TCP instead of UDP. The documented root cause of the flathub
# bootstrap failures (investigated 2026-07-11, see the Tooling step) is this box's
# Docker embedded resolver at 127.0.0.11 DROPPING UDP lookups while the shared
# runner fleet is saturated — a datagram nobody retransmits, so the lookup just
# times out. The answer then was to widen retry.sh's budget to 10 attempts (~9 min),
# which is enough to outlast a main push's ~8-workflow fan-out but NOT a TAG push's
# 13: v0.15.0 (twice) and v0.16.0 each burned all 10 attempts and failed the job,
# each needing a manual re-run.
#
# `use-vc` makes glibc use TCP, where the kernel retransmits and the query cannot be
# silently lost under load. Same resolver, same search path — only the transport
# changes, so internal names (git.unom.io) resolve exactly as before; deliberately
# NO extra nameservers, which would risk answering an internal name from a public
# resolver. Docker's embedded DNS serves TCP on 127.0.0.11:53 as well as UDP.
# retry.sh stays as the backstop for genuine upstream blips.
#
# Non-fatal: Docker bind-mounts /etc/resolv.conf and can present it read-only, and a
# DNS tuning that cannot be applied must not be what fails the release build — that
# would trade an occasional re-run for a hard stop. Falling back to UDP just restores
# today's behaviour, which retry.sh already covers.
if ! grep -q '^options .*use-vc' /etc/resolv.conf 2>/dev/null; then
echo 'options use-vc timeout:3 attempts:3' >> /etc/resolv.conf \
|| echo "::warning::could not set use-vc (read-only resolv.conf?); staying on UDP"
fi
cat /etc/resolv.conf || true
# fedora:43 has no node, but actions/checkout (a JS action) needs it. A plain `run:` step # fedora:43 has no node, but actions/checkout (a JS action) needs it. A plain `run:` step
# executes via the container shell (no node needed), so install node BEFORE checkout. # executes via the container shell (no node needed), so install node BEFORE checkout.
+69
View File
@@ -0,0 +1,69 @@
# Publish the plugin framework (@punktfunk/plugin-kit) to the Gitea npm registry
# (https://git.unom.io/api/packages/unom/npm/).
#
# Trigger: push a tag `plugin-kit-vX.Y.Z` (must equal plugin-kit/package.json "version"),
# or run manually. Versions independently of the app's `v*` and the SDK's `sdk-v*` tags.
#
# The kit's devDependency on @punktfunk/host is `file:../sdk`, so the SDK's dist must be
# built BEFORE the kit's `bun install` copies it.
#
# Auth: REGISTRY_TOKEN — the same repo Actions secret sdk-publish.yml uses.
name: plugin-kit-publish
on:
push:
tags: ['plugin-kit-v*']
workflow_dispatch:
jobs:
publish:
runs-on: ubuntu-24.04
container:
image: oven/bun:1
timeout-minutes: 15
steps:
# oven/bun's slim base ships neither git, a CA bundle, nor node — actions/checkout's HTTPS
# fetch needs git + ca-certificates, and the version-guard step below uses node.
- name: Install git + node + CA certs
run: apt-get update && apt-get install -y --no-install-recommends ca-certificates git nodejs
- uses: actions/checkout@v4
- name: Build the SDK (file:../sdk dependency source)
working-directory: sdk
run: |
bun install --frozen-lockfile --ignore-scripts
bun run build
- name: Install dependencies
working-directory: plugin-kit
run: bun install --frozen-lockfile --ignore-scripts
- name: Typecheck
working-directory: plugin-kit
run: bun run typecheck
- name: Test
working-directory: plugin-kit
run: bun test
- name: Build (dist/ JS + .d.ts + theme.css)
working-directory: plugin-kit
run: bun run build
- name: Tag matches package version
if: startsWith(github.ref, 'refs/tags/')
working-directory: plugin-kit
run: |
TAG="${GITHUB_REF_NAME#plugin-kit-v}"
PKG="$(node -p "require('./package.json').version")"
test "$TAG" = "$PKG" || { echo "tag $GITHUB_REF_NAME does not match package version $PKG"; exit 1; }
- name: Publish to Gitea registry
working-directory: plugin-kit
env:
NODE_AUTH_TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
test -n "$NODE_AUTH_TOKEN" || { echo "REGISTRY_TOKEN secret is empty"; exit 1; }
printf '//git.unom.io/api/packages/unom/npm/:_authToken=%s\n' "$NODE_AUTH_TOKEN" >> .npmrc
bun publish
+24
View File
@@ -149,6 +149,26 @@ jobs:
# inherits this from the env during the xcframework build). # inherits this from the env during the xcframework build).
echo "CMAKE_POLICY_VERSION_MINIMUM=3.5" >> "$GITHUB_ENV" echo "CMAKE_POLICY_VERSION_MINIMUM=3.5" >> "$GITHUB_ENV"
- name: Pin + prune Xcode DerivedData
# Without -derivedDataPath, xcodebuild derives its DerivedData directory name from the
# PROJECT'S ABSOLUTE PATH — and act_runner rotates its workspace
# (~/.cache/act/<hash>/hostexecutor), so each rotation minted a brand new ~760 MB tree
# under ~/Library that nothing ever collected. 31 of them piled up in three days
# (~32 GB with the shared ModuleCache), filled the runner's boot volume, and failed
# v0.16.0's xcframework build with "No space left on device". Pinning one path makes the
# tree REUSED instead of multiplied — it also keeps the module cache warm between runs.
run: |
DD="$HOME/ci/derived-data/release"
mkdir -p "$DD"
echo "DERIVED_DATA=$DD" >> "$GITHUB_ENV"
# Safety net for trees the pin does not own: the legacy per-path ones from before this
# change, and anything another job leaves in the default root. Untouched for a week ⇒ gone.
if [ -d "$HOME/Library/Developer/Xcode/DerivedData" ]; then
find "$HOME/Library/Developer/Xcode/DerivedData" -mindepth 1 -maxdepth 1 \
-mtime +7 -exec rm -rf {} + 2>/dev/null || true
fi
echo "disk after prune:"; df -h /System/Volumes/Data | tail -1
- name: Build PunktfunkCore.xcframework (mac + iOS + tvOS) - name: Build PunktfunkCore.xcframework (mac + iOS + tvOS)
# tvOS is a tier-3 target (nightly -Zbuild-std): slow on the first build, then cached on # tvOS is a tier-3 target (nightly -Zbuild-std): slow on the first build, then cached on
# the self-hosted runner. Built on canary too so the tvOS archive/upload below runs on the # the self-hosted runner. Built on canary too so the tvOS archive/upload below runs on the
@@ -176,6 +196,7 @@ jobs:
-project "$PROJECT" -scheme Punktfunk \ -project "$PROJECT" -scheme Punktfunk \
-destination 'generic/platform=macOS' \ -destination 'generic/platform=macOS' \
-archivePath "$RUNNER_TEMP/Punktfunk-macos.xcarchive" \ -archivePath "$RUNNER_TEMP/Punktfunk-macos.xcarchive" \
-derivedDataPath "$DERIVED_DATA" \
-skipMacroValidation -skipPackagePluginValidation \ -skipMacroValidation -skipPackagePluginValidation \
MARKETING_VERSION="$VERSION" CURRENT_PROJECT_VERSION="$BUILD_NUM" \ MARKETING_VERSION="$VERSION" CURRENT_PROJECT_VERSION="$BUILD_NUM" \
CODE_SIGNING_ALLOWED=NO CODE_SIGNING_ALLOWED=NO
@@ -273,6 +294,7 @@ jobs:
-project "$PROJECT" -scheme Punktfunk \ -project "$PROJECT" -scheme Punktfunk \
-destination 'generic/platform=macOS' \ -destination 'generic/platform=macOS' \
-archivePath "$RUNNER_TEMP/Punktfunk-macos-appstore.xcarchive" \ -archivePath "$RUNNER_TEMP/Punktfunk-macos-appstore.xcarchive" \
-derivedDataPath "$DERIVED_DATA" \
-skipMacroValidation -skipPackagePluginValidation \ -skipMacroValidation -skipPackagePluginValidation \
-allowProvisioningUpdates \ -allowProvisioningUpdates \
-authenticationKeyPath "$RUNNER_TEMP/asc.p8" \ -authenticationKeyPath "$RUNNER_TEMP/asc.p8" \
@@ -336,6 +358,7 @@ jobs:
-project "$PROJECT" -scheme Punktfunk-iOS \ -project "$PROJECT" -scheme Punktfunk-iOS \
-destination 'generic/platform=iOS' \ -destination 'generic/platform=iOS' \
-archivePath "$RUNNER_TEMP/Punktfunk-ios.xcarchive" \ -archivePath "$RUNNER_TEMP/Punktfunk-ios.xcarchive" \
-derivedDataPath "$DERIVED_DATA" \
-skipMacroValidation -skipPackagePluginValidation \ -skipMacroValidation -skipPackagePluginValidation \
-allowProvisioningUpdates \ -allowProvisioningUpdates \
-authenticationKeyPath "$RUNNER_TEMP/asc.p8" \ -authenticationKeyPath "$RUNNER_TEMP/asc.p8" \
@@ -394,6 +417,7 @@ jobs:
-project "$PROJECT" -scheme Punktfunk-tvOS \ -project "$PROJECT" -scheme Punktfunk-tvOS \
-destination 'generic/platform=tvOS' \ -destination 'generic/platform=tvOS' \
-archivePath "$RUNNER_TEMP/Punktfunk-tvos.xcarchive" \ -archivePath "$RUNNER_TEMP/Punktfunk-tvos.xcarchive" \
-derivedDataPath "$DERIVED_DATA" \
-skipMacroValidation -skipPackagePluginValidation \ -skipMacroValidation -skipPackagePluginValidation \
-allowProvisioningUpdates \ -allowProvisioningUpdates \
-authenticationKeyPath "$RUNNER_TEMP/asc.p8" \ -authenticationKeyPath "$RUNNER_TEMP/asc.p8" \
Generated
+30 -27
View File
@@ -2159,7 +2159,7 @@ dependencies = [
[[package]] [[package]]
name = "latency-probe" name = "latency-probe"
version = "0.14.0" version = "0.17.1"
[[package]] [[package]]
name = "lazy_static" name = "lazy_static"
@@ -2264,7 +2264,7 @@ dependencies = [
[[package]] [[package]]
name = "libvpl-sys" name = "libvpl-sys"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"bindgen", "bindgen",
"cmake", "cmake",
@@ -2299,7 +2299,7 @@ checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad"
[[package]] [[package]]
name = "loss-harness" name = "loss-harness"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"punktfunk-core", "punktfunk-core",
] ]
@@ -2788,7 +2788,7 @@ checksum = "9b4f627cb1b25917193a259e49bdad08f671f8d9708acfd5fe0a8c1455d87220"
[[package]] [[package]]
name = "pf-capture" name = "pf-capture"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ashpd", "ashpd",
@@ -2808,7 +2808,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-client-core" name = "pf-client-core"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ash", "ash",
@@ -2832,7 +2832,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-clipboard" name = "pf-clipboard"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ashpd", "ashpd",
@@ -2850,7 +2850,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-console-ui" name = "pf-console-ui"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ash", "ash",
@@ -2871,7 +2871,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-encode" name = "pf-encode"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ash", "ash",
@@ -2881,6 +2881,7 @@ dependencies = [
"libvpl-sys", "libvpl-sys",
"nvidia-video-codec-sdk", "nvidia-video-codec-sdk",
"openh264", "openh264",
"pf-capture",
"pf-frame", "pf-frame",
"pf-gpu", "pf-gpu",
"pf-host-config", "pf-host-config",
@@ -2894,7 +2895,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-ffvk" name = "pf-ffvk"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"ash", "ash",
"bindgen", "bindgen",
@@ -2903,7 +2904,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-frame" name = "pf-frame"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"libc", "libc",
@@ -2915,7 +2916,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-gpu" name = "pf-gpu"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"pf-host-config", "pf-host-config",
@@ -2929,11 +2930,11 @@ dependencies = [
[[package]] [[package]]
name = "pf-host-config" name = "pf-host-config"
version = "0.14.0" version = "0.17.1"
[[package]] [[package]]
name = "pf-inject" name = "pf-inject"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ashpd", "ashpd",
@@ -2961,14 +2962,14 @@ dependencies = [
[[package]] [[package]]
name = "pf-paths" name = "pf-paths"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"tracing", "tracing",
] ]
[[package]] [[package]]
name = "pf-presenter" name = "pf-presenter"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ash", "ash",
@@ -2983,7 +2984,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-vdisplay" name = "pf-vdisplay"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ashpd", "ashpd",
@@ -3013,7 +3014,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-win-display" name = "pf-win-display"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"pf-paths", "pf-paths",
@@ -3025,7 +3026,7 @@ dependencies = [
[[package]] [[package]]
name = "pf-zerocopy" name = "pf-zerocopy"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ash", "ash",
@@ -3221,7 +3222,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-client-android" name = "punktfunk-client-android"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"android_logger", "android_logger",
"jni", "jni",
@@ -3237,7 +3238,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-client-linux" name = "punktfunk-client-linux"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"async-channel", "async-channel",
@@ -3253,7 +3254,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-client-session" name = "punktfunk-client-session"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"pf-client-core", "pf-client-core",
@@ -3268,7 +3269,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-client-windows" name = "punktfunk-client-windows"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"async-channel", "async-channel",
"ffmpeg-next", "ffmpeg-next",
@@ -3287,7 +3288,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-core" name = "punktfunk-core"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"aes-gcm", "aes-gcm",
"bytes", "bytes",
@@ -3318,7 +3319,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-host" name = "punktfunk-host"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"aes", "aes",
"aes-gcm", "aes-gcm",
@@ -3363,11 +3364,13 @@ dependencies = [
"rand 0.8.6", "rand 0.8.6",
"rcgen", "rcgen",
"reis", "reis",
"ring",
"roxmltree", "roxmltree",
"rsa", "rsa",
"rusqlite", "rusqlite",
"rustls", "rustls",
"rusty_enet", "rusty_enet",
"semver",
"serde", "serde",
"serde_json", "serde_json",
"sha2", "sha2",
@@ -3400,7 +3403,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-probe" name = "punktfunk-probe"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"mdns-sd", "mdns-sd",
@@ -3414,7 +3417,7 @@ dependencies = [
[[package]] [[package]]
name = "punktfunk-tray" name = "punktfunk-tray"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"ksni", "ksni",
@@ -3437,7 +3440,7 @@ checksum = "d55d956fa96f5ec02be2e13af0e20391a5aa83d6a074e3ad368959d0fab299ea"
[[package]] [[package]]
name = "pyrowave-sys" name = "pyrowave-sys"
version = "0.14.0" version = "0.17.1"
dependencies = [ dependencies = [
"bindgen", "bindgen",
"cmake", "cmake",
+1 -1
View File
@@ -48,7 +48,7 @@ exclude = [
ndk = { path = "clients/android/native/vendor/ndk" } ndk = { path = "clients/android/native/vendor/ndk" }
[workspace.package] [workspace.package]
version = "0.14.0" version = "0.17.1"
edition = "2021" edition = "2021"
rust-version = "1.82" rust-version = "1.82"
license = "MIT OR Apache-2.0" license = "MIT OR Apache-2.0"
+1133 -1
View File
File diff suppressed because it is too large Load Diff
@@ -404,7 +404,14 @@ fn feeder_loop(
// stage is consumed: the HUD, or the ABR decode signal (`measure_decode`). The // stage is consumed: the HUD, or the ABR decode signal (`measure_decode`). The
// HUD-only `received` point + host/network split stay gated on the overlay. // HUD-only `received` point + host/network split stay gated on the overlay.
if stats.enabled() || measure_decode { if stats.enabled() || measure_decode {
let received_ns = now_realtime_ns(); // Core reassembly-completion stamp (ABI v9), NOT the pull instant: stamping
// here would fold the hand-off queue wait into the network latency figure
// (a client-side standing backlog masquerading as network). 0 = older core.
let received_ns = if frame.received_ns > 0 {
frame.received_ns as i128
} else {
now_realtime_ns()
};
{ {
let mut g = in_flight let mut g = in_flight
.lock() .lock()
@@ -221,7 +221,13 @@ pub(super) fn run_sync(
// samplers (`received` point, host/network split) stay gated on the overlay so // samplers (`received` point, host/network split) stay gated on the overlay so
// the hidden steady state adds only a wall-clock read + the receipt push. // the hidden steady state adds only a wall-clock read + the receipt push.
if stats.enabled() || measure_decode { if stats.enabled() || measure_decode {
let received_ns = now_realtime_ns(); // Core reassembly-completion stamp (ABI v9), not the pull instant — see
// async_loop: a pull stamp folds hand-off queue wait into "network".
let received_ns = if frame.received_ns > 0 {
frame.received_ns as i128
} else {
now_realtime_ns()
};
in_flight.push_back((frame.pts_ns / 1000, received_ns)); in_flight.push_back((frame.pts_ns / 1000, received_ns));
if in_flight.len() > IN_FLIGHT_CAP { if in_flight.len() > IN_FLIGHT_CAP {
in_flight.pop_front(); // stale — codec never echoed it back in_flight.pop_front(); // stale — codec never echoed it back
@@ -579,13 +579,21 @@ struct ContentView: View {
model?.disconnect() // the captured-state D combo model?.disconnect() // the captured-state D combo
}, },
onFrame: { [meter = model.meter, latency = model.latency, onFrame: { [meter = model.meter, latency = model.latency,
split = model.latencySplit, offset = conn.clockOffsetNs] au in split = model.latencySplit, queue = model.clientQueue,
offset = conn.clockOffsetNs] au in
meter.note(byteCount: au.data.count) meter.note(byteCount: au.data.count)
latency.record(ptsNs: au.ptsNs, offsetNs: offset) latency.record(ptsNs: au.ptsNs, offsetNs: offset)
// The same receipt, keyed by pts, awaiting its 0xCF host timing (the // The same receipt, keyed by pts, awaiting its 0xCF host timing (the
// host/network split drained by the 1 s stats tick). // host/network split drained by the 1 s stats tick). receivedNs is
// the core's reassembly stamp (ABI v9), so the split's network term no
// longer contains the client-queue wait...
split.recordReceipt( split.recordReceipt(
ptsNs: au.ptsNs, receivedNs: au.receivedNs, offsetNs: offset) ptsNs: au.ptsNs, receivedNs: au.receivedNs, offsetNs: offset)
// ...which is measured as its own term instead (receiptpull, both
// client-local).
queue.record(
ptsNs: UInt64(bitPattern: au.receivedNs), atNs: au.pulledNs,
offsetNs: 0)
}, },
onSessionEnd: { [weak model] in onSessionEnd: { [weak model] in
Task { @MainActor in model?.sessionEnded() } Task { @MainActor in model?.sessionEnded() }
@@ -102,6 +102,12 @@ final class SessionModel: ObservableObject {
@Published var decodeValid = false @Published var decodeValid = false
@Published var displayP50Ms = 0.0 @Published var displayP50Ms = 0.0
@Published var displayValid = false @Published var displayValid = false
/// Client-queue wait: core reassembly receipt the pump's pull (`AccessUnit.pulledNs
/// receivedNs`, ABI v9 receipt split the 2026-07 two-pair investigation). ~0 on a healthy
/// stream; a persistent value is a client-side standing backlog that used to hide inside
/// "network". Shown in the detailed tier only when it says something ( ~2 ms).
@Published var clientQueueP50Ms = 0.0
@Published var clientQueueValid = false
/// The measured OS present floor (design/apple-presentation-rebuild.md): the deadline /// The measured OS present floor (design/apple-presentation-rebuild.md): the deadline
/// engine's vendglass pipeline depth an OS property no client can pace under (~2 refresh /// engine's vendglass pipeline depth an OS property no client can pace under (~2 refresh
/// intervals composited; would read ~1 under direct-to-display). The HUD subtracts it from /// intervals composited; would read ~1 under direct-to-display). The HUD subtracts it from
@@ -147,6 +153,9 @@ final class SessionModel: ObservableObject {
let endToEnd = LatencyMeter() let endToEnd = LatencyMeter()
let decodeStage = LatencyMeter() let decodeStage = LatencyMeter()
let displayStage = LatencyMeter() let displayStage = LatencyMeter()
/// Client-queue sampler (see `clientQueueP50Ms`) fed per AU by the stream view's onFrame,
/// drained by the same 1 s tick as the stage meters.
let clientQueue = LatencyMeter()
/// The OS present floor sampler (see `osFloorP50Ms`) fed one sample per display-link /// The OS present floor sampler (see `osFloorP50Ms`) fed one sample per display-link
/// update by the deadline engine, drained by the same 1 s tick as the stage meters. /// update by the deadline engine, drained by the same 1 s tick as the stage meters.
let presentFloor = LatencyMeter() let presentFloor = LatencyMeter()
@@ -489,6 +498,7 @@ final class SessionModel: ObservableObject {
endToEndValid = false endToEndValid = false
decodeValid = false decodeValid = false
displayValid = false displayValid = false
clientQueueValid = false
osFloorValid = false osFloorValid = false
lostFrames = 0 lostFrames = 0
lostPct = 0 lostPct = 0
@@ -679,6 +689,12 @@ final class SessionModel: ObservableObject {
} else { } else {
self.osFloorValid = false self.osFloorValid = false
} }
if let q = self.clientQueue.drain() {
self.clientQueueP50Ms = q.p50Ms
self.clientQueueValid = true
} else {
self.clientQueueValid = false
}
// Mirror the window to the unified log (see statsLog) one line per second, // Mirror the window to the unified log (see statsLog) one line per second,
// stages in ms, only while frames actually flowed. `fps` counts RECEIVED AUs; // stages in ms, only while frames actually flowed. `fps` counts RECEIVED AUs;
// `presents` counts frames that reached glass (the display meter's sample count) // `presents` counts frames that reached glass (the display meter's sample count)
@@ -689,9 +705,12 @@ final class SessionModel: ObservableObject {
// captured before the 2026-07 floor policy); the appended trio carries the // captured before the 2026-07 floor policy); the appended trio carries the
// measured OS present floor and the floor-shaved values the HUD displays. // measured OS present floor and the floor-shaved values the HUD displays.
let line = String( let line = String(
format: "fps=%d presents=%d e2e_p50=%.1f e2e_p95=%.1f hostnet_p50=%.1f " // Swift Int is 64-bit %lld, NOT %d (which is a 32-bit C int); macOS 26's
+ "decode_p50=%.1f display_p50=%.1f lost=%d " // strict String(format:) validator rejects the %d/Int mismatch and drops
+ "floor_p50=%.1f display_adj=%.1f e2e_adj=%.1f", // the whole line (a cascade error that also mis-blames the float args).
format: "fps=%lld presents=%lld e2e_p50=%.1f e2e_p95=%.1f hostnet_p50=%.1f "
+ "decode_p50=%.1f display_p50=%.1f lost=%lld "
+ "floor_p50=%.1f display_adj=%.1f e2e_adj=%.1f queue_p50=%.1f",
frames, frames,
displayWindow?.count ?? 0, displayWindow?.count ?? 0,
self.endToEndValid ? self.endToEndP50Ms : -1, self.endToEndValid ? self.endToEndP50Ms : -1,
@@ -702,7 +721,8 @@ final class SessionModel: ObservableObject {
lost, lost,
self.osFloorValid ? self.osFloorP50Ms : -1, self.osFloorValid ? self.osFloorP50Ms : -1,
self.displayValid ? self.displayAdjP50Ms : -1, self.displayValid ? self.displayAdjP50Ms : -1,
self.endToEndValid ? self.endToEndAdjP50Ms : -1) self.endToEndValid ? self.endToEndAdjP50Ms : -1,
self.clientQueueValid ? self.clientQueueP50Ms : -1)
statsLog.info("\(line, privacy: .public)") statsLog.info("\(line, privacy: .public)")
} }
} }
@@ -118,6 +118,16 @@ struct StreamHUDView: View {
.font(.system(.caption2, design: .monospaced)) .font(.system(.caption2, design: .monospaced))
.foregroundStyle(.tertiary) .foregroundStyle(.tertiary)
} }
// Client-queue wait (reassembly receipt decode pull, ABI v9 split): ~0 on
// a healthy stream and hidden as noise; shown from 2 ms a persistent value
// is a client-side standing backlog that pre-split builds displayed as
// "network" (the 2026-07 two-pair plateau). The core's standing-latency
// bleed logs alongside when it acts on the same state.
if model.clientQueueValid && model.clientQueueP50Ms >= 2 {
Text("client queue +\(model.clientQueueP50Ms, specifier: "%.1f") (receive backlog — standing if it persists)")
.font(.system(.caption2, design: .monospaced))
.foregroundStyle(.tertiary)
}
} }
} else if model.hostNetworkValid { } else if model.hostNetworkValid {
// Stage-1 fallback presenter: the layer decodes + presents internally with no // Stage-1 fallback presenter: the layer decodes + presents internally with no
@@ -43,6 +43,9 @@ struct GamepadSettingsView: View {
@AppStorage(DefaultsKey.presentPriority) private var presentPriority = @AppStorage(DefaultsKey.presentPriority) private var presentPriority =
SettingsOptions.presentPriorityDefault SettingsOptions.presentPriorityDefault
@AppStorage(DefaultsKey.smoothBuffer) private var smoothBuffer = 0 @AppStorage(DefaultsKey.smoothBuffer) private var smoothBuffer = 0
#if os(macOS)
@AppStorage(DefaultsKey.windowedSafePresent) private var windowedSafePresent = true
#endif
#if os(iOS) #if os(iOS)
@AppStorage(DefaultsKey.rumbleOnDevice) private var rumbleOnDevice = false @AppStorage(DefaultsKey.rumbleOnDevice) private var rumbleOnDevice = false
#endif #endif
@@ -345,6 +348,22 @@ struct GamepadSettingsView: View {
detail: "Turn off to use the touch interface even with a controller connected.", detail: "Turn off to use the touch interface even with a controller connected.",
value: $gamepadUIEnabled), value: $gamepadUIEnabled),
] ]
#if os(macOS)
// The windowed safe-present toggle slots in after "Smoothness buffer" (staying inside
// the Video group) macOS only, mirroring the touch SettingsView's Presentation row
// (the DCP swapID-panic mitigation; see DefaultsKey.windowedSafePresent).
if let at = list.firstIndex(where: { $0.id == "smoothBuffer" }) {
list.insert(
toggleRow(
id: "windowedSafePresent", icon: "macwindow.badge.plus",
label: "Safe windowed presentation",
detail: "Windowed streams present in step with the compositor — avoids a "
+ "macOS display-driver crash on high-refresh displays, at a small "
+ "latency cost. Fullscreen always uses the fastest path.",
value: $windowedSafePresent),
at: at + 1)
}
#endif
#if os(iOS) #if os(iOS)
// The device-rumble mirror slots in after "Controller type" (staying inside the // The device-rumble mirror slots in after "Controller type" (staying inside the
// Controller group the next row carries the "Interface" header). iPhone only in // Controller group the next row carries the "Interface" header). iPhone only in
@@ -300,6 +300,18 @@ extension SettingsView {
+ "of added latency. Off shows frames as soon as they're ready.") { + "of added latency. Off shows frames as soon as they're ready.") {
Toggle("V-Sync", isOn: $vsync) Toggle("V-Sync", isOn: $vsync)
} }
// The DCP swapID-panic mitigation's user handle (see DefaultsKey.windowedSafePresent
// for the saga). Default ON: turning it off re-arms a WHOLE-MACHINE kernel panic on
// affected setups, so the caption says so in plain words.
described(windowedSafePresent
? "Windowed streams present in step with the system compositor — avoids a macOS "
+ "display-driver crash seen on high-refresh displays, at a small latency "
+ "cost. Fullscreen always uses the fastest path."
: "Windowed streams use the fastest present path. On some high-refresh setups "
+ "this can crash macOS itself (kernel panic) — turn back on if your Mac "
+ "restarts during windowed streaming.") {
Toggle("Safe windowed presentation", isOn: $windowedSafePresent)
}
#endif #endif
} }
} }
@@ -40,6 +40,7 @@ struct SettingsView: View {
@AppStorage(DefaultsKey.smoothBuffer) var smoothBuffer = 0 @AppStorage(DefaultsKey.smoothBuffer) var smoothBuffer = 0
#if os(macOS) #if os(macOS)
@AppStorage(DefaultsKey.vsync) var vsync = false @AppStorage(DefaultsKey.vsync) var vsync = false
@AppStorage(DefaultsKey.windowedSafePresent) var windowedSafePresent = true
#endif #endif
#if !os(tvOS) #if !os(tvOS)
@AppStorage(DefaultsKey.allowVRR) var allowVRR = true @AppStorage(DefaultsKey.allowVRR) var allowVRR = true
@@ -35,10 +35,31 @@ public struct AccessUnit: Sendable {
public let ptsNs: UInt64 public let ptsNs: UInt64
public let frameIndex: UInt32 public let frameIndex: UInt32
public let flags: UInt32 public let flags: UInt32
/// Client `CLOCK_REALTIME` instant the AU was handed over by the core (post-FEC, decrypted) /// Client `CLOCK_REALTIME` instant the AU finished reassembly in the core (post-FEC,
/// the **received** measurement point of design/stats-unification.md. The decode stage is /// decrypted `PunktfunkFrame.received_ns`, ABI v9) the **received** measurement point of
/// `decodedNs - receivedNs`, both client-local (no skew offset applies). /// design/stats-unification.md. NOT the pull instant: stamping at the pull folded the
/// pre-decode hand-off wait into the network term, which is how the 2026-07 two-pair
/// standing-latency plateau hid as "network". The decode stage is `decodedNs - receivedNs`,
/// both client-local (no skew offset applies).
public let receivedNs: Int64 public let receivedNs: Int64
/// Client `CLOCK_REALTIME` instant this pull returned. `pulledNs - receivedNs` is the
/// client-queue wait (kernel hand-off + FrameChannel dwell) the term the HUD splits out
/// so a client-side standing backlog can never masquerade as network latency again.
public let pulledNs: Int64
/// `pulledNs` defaults to `receivedNs` (zero queue wait) for callers with no pull instant
/// the synthetic probe AUs and decode tests, where the split is meaningless.
public init(
data: Data, ptsNs: UInt64, frameIndex: UInt32, flags: UInt32,
receivedNs: Int64, pulledNs: Int64? = nil
) {
self.data = data
self.ptsNs = ptsNs
self.frameIndex = frameIndex
self.flags = flags
self.receivedNs = receivedNs
self.pulledNs = pulledNs ?? receivedNs
}
} }
/// One Opus audio packet (48 kHz stereo, 5 ms frames) decode with AVAudioConverter /// One Opus audio packet (48 kHz stereo, 5 ms frames) decode with AVAudioConverter
@@ -662,11 +683,16 @@ public final class PunktfunkConnection {
let data = Data(bytes: base, count: Int(frame.len)) // copy: ptr valid only until next call let data = Data(bytes: base, count: Int(frame.len)) // copy: ptr valid only until next call
var ts = timespec() var ts = timespec()
clock_gettime(CLOCK_REALTIME, &ts) clock_gettime(CLOCK_REALTIME, &ts)
let receivedNs = Int64(ts.tv_sec) * 1_000_000_000 + Int64(ts.tv_nsec) let pulledNs = Int64(ts.tv_sec) * 1_000_000_000 + Int64(ts.tv_nsec)
// Receipt = the core's reassembly-completion stamp (ABI v9); the pull instant is
// kept separately so the client-queue wait is its own measured term. 0 would mean a
// pre-v9 core impossible here (core and Kit ship in one binary), but fall back to
// the pull instant rather than record a 1970 receipt.
let receivedNs = frame.received_ns > 0 ? Int64(frame.received_ns) : pulledNs
return AccessUnit( return AccessUnit(
data: data, ptsNs: frame.pts_ns, data: data, ptsNs: frame.pts_ns,
frameIndex: frame.frame_index, flags: frame.flags, frameIndex: frame.frame_index, flags: frame.flags,
receivedNs: receivedNs) receivedNs: receivedNs, pulledNs: pulledNs)
case statusNoFrame: case statusNoFrame:
return nil return nil
case statusClosed: case statusClosed:
@@ -23,6 +23,34 @@ import os
private let presenterLog = Logger(subsystem: "io.unom.punktfunk", category: "presenter") private let presenterLog = Logger(subsystem: "io.unom.punktfunk", category: "presenter")
#if os(macOS)
/// HOW a windowed (composited) macOS session pushes finished frames to glass the DCP
/// "mismatched swapID's" kernel-panic saga's mechanism picker. Fullscreen always presents
/// `async` (direct-scanout promotion, lowest latency, no panic reports there); the windowed
/// mechanism is resolved per session by SessionPresenter (user setting +
/// PUNKTFUNK_WINDOWED_PRESENT env override) and routed here via `setWindowedPresent`.
///
/// - `async`: the CAMetalLayer image queue (`commandBuffer.present`) the fastest composited
/// path and the PANIC TRIGGER on high-refresh displays (the out-of-band swaps race
/// WindowServer's compositor; it survived glass pacing and every codec).
/// - `transaction`: `CAMetalLayer.presentsWithTransaction` the swap commits WITH the layer
/// tree, in lockstep with the compositor (Apple's documented remedy; validated no-panic on
/// the 240 Hz repro machine). The present is committed from the RENDER thread inside an
/// explicit CATransaction + flush see `encodePresent` for why that beats the original
/// main-thread hop.
/// - `surface`: no image queue at all render into a pooled IOSurface and swap it into a plain
/// CALayer's `contents` (the f407f418 PyroWave mitigation, resurrected format-aware:
/// rgba16Float + PQ tagging keeps HDR). WindowServer treats it as ordinary layer damage on
/// its own composite cadence. PROTOTYPE: whether the compositor honors PQ/EDR for plain-layer
/// IOSurface contents still needs an on-glass eyeball the metal layer stays underneath with
/// `wantsExtendedDynamicRangeContent` as the EDR anchor.
enum WindowedPresentMode: String, Sendable {
case async
case transaction
case surface
}
#endif
/// HDR reference white (BT.2408 "HDR Reference White"): the absolute luminance, in nits, that the /// HDR reference white (BT.2408 "HDR Reference White"): the absolute luminance, in nits, that the
/// PQ signal's diffuse white sits at. Passed to `CAEDRMetadata.hdr10(opticalOutputScale:)`, it anchors /// PQ signal's diffuse white sits at. Passed to `CAEDRMetadata.hdr10(opticalOutputScale:)`, it anchors
/// 203-nit diffuse white at EDR 1.0 (the display's SDR-white level) and lets the system tone-map the /// 203-nit diffuse white at EDR 1.0 (the display's SDR-white level) and lets the system tone-map the
@@ -198,8 +226,8 @@ fragment float4 pf_frag_hdr(VOut in [[stage_in]],
// in a genuine HDR10 output, PQ passthrough is the correct emission and the TV tone-maps.) // in a genuine HDR10 output, PQ passthrough is the correct emission and the TV tone-maps.)
// The shared PQdisplay-referred-SDR tail (see pf_frag_hdr_tv's rationale above): ST 2084 // The shared PQdisplay-referred-SDR tail (see pf_frag_hdr_tv's rationale above): ST 2084
// EOTF 203-nit-anchored scene light BT.2020709 primaries extended-Reinhard rolloff // EOTF 203-nit-anchored scene light BT.2020709 primaries extended-Reinhard rolloff
// BT.709 OETF. Used by the tvOS biplanar tone-map and the planar (PyroWave) tone-map the // BT.709 OETF. Used by the tvOS biplanar tone-map and the tvOS planar (PyroWave) tone-map (the
// latter also on macOS windowed sessions, whose IOSurface present path is BGRA8-only. // no-HDR-headroom fallback). macOS keeps real HDR windowed now see `WindowedPresentMode`.
static inline float3 pqToSdr(float3 pq) { static inline float3 pqToSdr(float3 pq) {
const float m1 = 2610.0/16384.0; const float m1 = 2610.0/16384.0;
const float m2 = 78.84375; const float m2 = 78.84375;
@@ -230,8 +258,7 @@ fragment float4 pf_frag_hdr_tv(VOut in [[stage_in]],
// PyroWave planar HDR tone-map: three separate R16 planes (P010-style studio codes; the rows // PyroWave planar HDR tone-map: three separate R16 planes (P010-style studio codes; the rows
// fold in depth-10 MSB packing) PQ RGB the shared SDR tail. Used when a PQ pyrowave // fold in depth-10 MSB packing) PQ RGB the shared SDR tail. Used when a PQ pyrowave
// stream must land on an 8-bit surface: tvOS without HDR headroom, and macOS WINDOWED sessions // stream must land on an 8-bit surface: tvOS without HDR headroom. The passthrough planar
// (the IOSurface present path the DCP-panic mitigation is BGRA8). The passthrough planar
// HDR pipeline reuses pf_frag_planar itself on an rgba16Float drawable (identical math the // HDR pipeline reuses pf_frag_planar itself on an rgba16Float drawable (identical math the
// layer's itur_2100_PQ colour space + EDR metadata do the interpretation). // layer's itur_2100_PQ colour space + EDR metadata do the interpretation).
fragment float4 pf_frag_planar_tm(VOut in [[stage_in]], fragment float4 pf_frag_planar_tm(VOut in [[stage_in]],
@@ -259,21 +286,51 @@ public final class MetalVideoPresenter {
public let layer: CAMetalLayer public let layer: CAMetalLayer
#if os(macOS) #if os(macOS)
/// The WINDOWED-mode PyroWave present target: a plain CALayer sized like `layer` (installed /// WINDOWED-mode present coordination the macOS DCP KERNEL PANIC mitigation.
/// as a sibling ABOVE it), fed IOSurfaces via `contents` inside ordinary CATransactions.
/// ///
/// Why this exists the macOS DCP KERNEL PANIC ("mismatched swapID's" @UnifiedPipeline.cpp, /// The panic ("mismatched swapID's" @UnifiedPipeline.cpp, WindowServer dies, machine reboots):
/// WindowServer dies, machine reboots): out-of-band CAMetalLayer image-queue swaps into a /// the CAMetalLayer's ASYNCHRONOUS image queue (`commandBuffer.present(drawable)` an
/// COMPOSITED (windowed) session race WindowServer's own swap submissions on high-refresh /// out-of-band flip, mandatory with `displaySyncEnabled=false`) diverges from WindowServer's
/// displays, and the race survives glass pacing a fully serialized one-in-flight present /// compositor on a high-refresh COMPOSITED (windowed) session the compositor's notion of the
/// stream still panicked a 240 Hz Mac Studio (2026-07-18, twice). So in windowed mode we stop /// current swap and the layer's queued swap disagree, and the DCP asserts. It survived glass
/// using the image queue entirely and present the way video players do: render the planar CSC /// pacing: a fully serialized one-in-flight present stream still panicked a 240 Hz Mac Studio
/// into an IOSurface pool and swap `contents` on main WindowServer treats it as ordinary /// (2026-07-18, PyroWave), and a windowed HEVC session panicked the same machine 2026-07-21
/// damage on its own composite cadence, coalescing faster-than-refresh updates instead of /// so it is the async image queue itself, at any pacing or codec, not a present rate.
/// latching queue swaps mid-cycle. Fullscreen keeps the CAMetalLayer path (direct-scanout ///
/// promotion, no compositing, no panic reports). Contents updates are transparent to the /// The fix keeps the full render path (rgba16Float / PQ / EDR real HDR is preserved) and
/// layer below when nil, so flipping modes just covers/uncovers the metal layer. /// only changes HOW the drawable is presented: `CAMetalLayer.presentsWithTransaction`. With it
public let surfaceLayer: CALayer = { /// set, we don't hand the drawable to the command buffer; we commit, wait until scheduled, then
/// call `drawable.present()` INSIDE a CATransaction the present is enrolled in Core
/// Animation's transaction and committed together with the layer tree, so the swap stays in
/// lockstep with the compositor instead of racing it (Apple's documented remedy for Metal
/// presentation drifting out of sync with CA). Fullscreen keeps the async path (direct-scanout
/// promotion, lowest latency, no compositor and no panic reports there).
///
/// 2026-07-21 latency rework: the mitigation MECHANISM is now a three-way pick
/// (`WindowedPresentMode`) and the transactional present commits from the RENDER thread
/// see `encodePresent`. Staged under `stagingLock` (main pushes it via
/// `setComposited``setWindowedPresent`); the render thread drains it and toggles the layer
/// property + present style. `Active` is the render-thread copy so the layer property flips
/// exactly once per mode change.
private var windowedPresentStaged: WindowedPresentMode = .async
private var windowedPresentActive: WindowedPresentMode = .async
/// PUNKTFUNK_TXN_PRESENT=main the ORIGINAL transactional present (commit
/// waitUntilScheduled hop to the MAIN thread and present inside its CATransaction), kept
/// as a field A/B lever. The default is the render-thread commit: the present harness
/// (2026-07-21, this saga) measured the main hop landing a runloop turn late on a busy main
/// thread, and an ACTIVE implicit transaction there NESTS the explicit one presents batch
/// at runloop-iteration rate (the field's presents=55 @ fps=240, display_p50 18.6 ms).
/// Off-main commits measured immune to main-thread churn (~10 ms glass p50 at 240 Hz
/// full-size vs 14+ ms under a churned main hop).
private let txnPresentOnMain =
ProcessInfo.processInfo.environment["PUNKTFUNK_TXN_PRESENT"] == "main"
/// The WINDOWED-mode `surface` present target: a plain CALayer sized like `layer` (installed
/// as a sibling ABOVE it by SessionPresenter), fed IOSurfaces via `contents` inside explicit
/// CATransactions. Transparent (nil contents) whenever surface mode is off, so the metal
/// layer below shows through. See `WindowedPresentMode.surface`.
let surfaceLayer: CALayer = {
let l = CALayer() let l = CALayer()
l.contentsGravity = .resize // frame is already aspect-fit + pixel-snapped by layout l.contentsGravity = .resize // frame is already aspect-fit + pixel-snapped by layout
l.isOpaque = true l.isOpaque = true
@@ -281,8 +338,8 @@ public final class MetalVideoPresenter {
return l return l
}() }()
/// One IOSurface-backed render target of the windowed present pool. All pool state is /// One IOSurface-backed render target of the windowed surface-present pool. All pool state
/// RENDER-THREAD confined; only the immutable surface refs cross to main (contents swap). /// is RENDER-THREAD confined; only the immutable surface refs cross threads (contents swap).
private struct SurfaceSlot { private struct SurfaceSlot {
let surface: IOSurfaceRef let surface: IOSurfaceRef
let texture: MTLTexture let texture: MTLTexture
@@ -292,15 +349,52 @@ public final class MetalVideoPresenter {
private var surfacePool: [SurfaceSlot] = [] private var surfacePool: [SurfaceSlot] = []
private var surfacePoolSize: CGSize = .zero private var surfacePoolSize: CGSize = .zero
private var surfacePoolHDR = false
private var surfaceSeq: UInt64 = 0 private var surfaceSeq: UInt64 = 0
/// Index of the slot most recently handed to the layer never rewritten next, even if its /// Index of the slot most recently handed to the layer never rewritten next, even if its
/// use count already dropped (the compositor may still be scanning out the previous frame). /// use count already dropped (the compositor may still be scanning out the previous frame).
private var lastHandedOff: Int? private var lastHandedOff: Int?
/// Staged (under `stagingLock`, like every cross-thread input): the hosting view's windowed
/// vs fullscreen state, pushed from main via `setSurfacePresents`. Drained in `renderPlanar`. /// Once-per-second decomposition of the ACTIVE windowed present path (the field-diagnosis
private var surfacePresentsStaged = false /// half of the DCP-latency work): scheduled/completed wait + commit/flush cost per present,
/// Render-thread copy, so pool teardown happens exactly once on a mode flip. /// and how many presents/swaps were issued. The pf-present line shows the GLASS side
private var surfacePresentsActive = false /// (latchMs / dropped); this shows the ISSUE side. Logged via `presenterLog` only while a
/// windowed mechanism is active (zero cost fullscreen). Lock-guarded: transaction mode
/// records from the render thread, surface mode from Metal completion threads.
private final class WindowedPresentDiag: @unchecked Sendable {
private let lock = NSLock()
private var presents = 0
private var schedMs: [Double] = []
private var commitMs: [Double] = []
private var last = CACurrentMediaTime()
func record(schedMs sched: Double, commitMs commit: Double, mode: WindowedPresentMode) {
lock.lock()
presents += 1
schedMs.append(sched)
commitMs.append(commit)
let now = CACurrentMediaTime()
guard now - last >= 1 else {
lock.unlock()
return
}
last = now
let sSched = schedMs.sorted()
let sCommit = commitMs.sorted()
let line = String(
format: "pf-windowed mode=%@ presents=%d schedMs p50=%.2f max=%.2f "
+ "commitMs p50=%.2f max=%.2f",
mode.rawValue, presents, sSched[sSched.count / 2], sSched.last ?? 0,
sCommit[sCommit.count / 2], sCommit.last ?? 0)
presents = 0
schedMs.removeAll(keepingCapacity: true)
commitMs.removeAll(keepingCapacity: true)
lock.unlock()
presenterLog.info("\(line, privacy: .public)")
}
}
private let windowedDiag = WindowedPresentDiag()
#endif #endif
private let device: MTLDevice private let device: MTLDevice
@@ -316,7 +410,7 @@ public final class MetalVideoPresenter {
private let pipelinePlanar: MTLRenderPipelineState private let pipelinePlanar: MTLRenderPipelineState
/// PyroWave planar HDR passthrough (pf_frag_planar rgba16Float; the layer's PQ colour /// PyroWave planar HDR passthrough (pf_frag_planar rgba16Float; the layer's PQ colour
/// space + EDR interpret the samples) and the planar PQSDR tone-map (pf_frag_planar_tm /// space + EDR interpret the samples) and the planar PQSDR tone-map (pf_frag_planar_tm
/// bgra8; tvOS without headroom + macOS windowed IOSurface presents). /// bgra8; tvOS without HDR headroom).
private let pipelinePlanarHDR: MTLRenderPipelineState private let pipelinePlanarHDR: MTLRenderPipelineState
private let pipelinePlanarToneMap: MTLRenderPipelineState private let pipelinePlanarToneMap: MTLRenderPipelineState
private var textureCache: CVMetalTextureCache? private var textureCache: CVMetalTextureCache?
@@ -591,13 +685,14 @@ public final class MetalVideoPresenter {
} }
#if os(macOS) #if os(macOS)
/// Park the windowed-vs-fullscreen present routing (MAIN thread the hosting view pushes its /// Park the windowed present mechanism (MAIN thread the hosting view pushes its window
/// window state on every layout). true = PyroWave frames present via `surfaceLayer` contents /// state on every layout; SessionPresenter resolves the mechanism per session). `.async` =
/// (the DCP swapID-panic mitigation see `surfaceLayer`); false = the CAMetalLayer path. /// FULLSCREEN (or the user opted out of the mitigation): the image queue. `.transaction` /
/// `.surface` = COMPOSITED (windowed) mitigation mechanisms see `WindowedPresentMode`.
/// Applied by the render thread on the next frame, like every other staged value here. /// Applied by the render thread on the next frame, like every other staged value here.
public func setSurfacePresents(_ on: Bool) { func setWindowedPresent(_ mode: WindowedPresentMode) {
stagingLock.lock() stagingLock.lock()
surfacePresentsStaged = on windowedPresentStaged = mode
stagingLock.unlock() stagingLock.unlock()
} }
#endif #endif
@@ -734,36 +829,12 @@ public final class MetalVideoPresenter {
) -> Bool { ) -> Bool {
stagingLock.lock() stagingLock.lock()
let targetFromLayout = drawableTarget let targetFromLayout = drawableTarget
#if os(macOS)
let surfaceMode = surfacePresentsStaged
#endif
stagingLock.unlock() stagingLock.unlock()
// A PQ (HDR) pyrowave stream drives the same layer/EDR machinery as the biplanar path; // A PQ (HDR) pyrowave stream drives the same layer/EDR machinery as the biplanar path
// macOS WINDOWED sessions stay on the SDR layer (the IOSurface path tone-maps in-shader). // including macOS windowed sessions, which keep real HDR (the DCP mitigation is the
#if os(macOS) // transactional present in `encodePresent`, not a colour downgrade).
configure(hdr: planes.pq && !surfaceMode)
#else
configure(hdr: planes.pq) configure(hdr: planes.pq)
#endif
var csc = planes.csc var csc = planes.csc
#if os(macOS)
if surfaceMode != surfacePresentsActive {
surfacePresentsActive = surfaceMode
presenterLog.info(
"stage2: windowed surface presents \(surfaceMode ? "ON" : "OFF", privacy: .public) (PyroWave DCP-panic mitigation)")
if !surfaceMode {
// Back to the metal path (fullscreen): drop the pool at 5K it holds >100 MB,
// and re-entering windowed mode rebuilds it in one frame.
surfacePool.removeAll()
surfacePoolSize = .zero
lastHandedOff = nil
}
}
if surfaceMode {
return renderPlanarToSurface(
planes, targetFromLayout: targetFromLayout, csc: &csc, onPresented: onPresented)
}
#endif
// PQ passthrough needs the HDR drawable; a PQ frame while the drawable is (still) // PQ passthrough needs the HDR drawable; a PQ frame while the drawable is (still)
// 8-bit tvOS without display headroom, or a not-yet-flipped layer tone-maps // 8-bit tvOS without display headroom, or a not-yet-flipped layer tone-maps
// in-shader instead (the pipeline must match the drawable's pixel format). // in-shader instead (the pipeline must match the drawable's pixel format).
@@ -792,118 +863,6 @@ public final class MetalVideoPresenter {
} }
} }
#if os(macOS)
/// The windowed-mode present tail (see `surfaceLayer` for why this path exists): render the
/// planar CSC into a pooled IOSurface and hand it to `surfaceLayer.contents` on MAIN inside a
/// plain CATransaction an ordinary damaged-layer update on WindowServer's own composite
/// cadence, no CAMetalLayer image-queue swap anywhere. `presentAtMediaTime` doesn't apply
/// (the compositor paces); `onPresented` fires after the contents swap is committed, stamped
/// with CLOCK_REALTIME then the closest observable analogue of "reached glass" here (the
/// composite follows within a refresh, so the meters' display stage reads slightly optimistic).
private func renderPlanarToSurface(
_ planes: WaveletPlanes, targetFromLayout: CGSize, csc: inout CscUniform,
onPresented: ((Int64?) -> Void)?
) -> Bool {
let decodedSize = CGSize(width: planes.width, height: planes.height)
let targetSize = (targetFromLayout.width > 0 && targetFromLayout.height > 0)
? targetFromLayout : decodedSize
ensureSurfacePool(size: targetSize)
guard let slotIndex = takeSurfaceSlot(),
let commandBuffer = queue.makeCommandBuffer()
else { return false }
let slot = surfacePool[slotIndex]
let pass = MTLRenderPassDescriptor()
pass.colorAttachments[0].texture = slot.texture
pass.colorAttachments[0].loadAction = .clear
pass.colorAttachments[0].clearColor = MTLClearColor(red: 0, green: 0, blue: 0, alpha: 1)
pass.colorAttachments[0].storeAction = .store
guard let encoder = commandBuffer.makeRenderCommandEncoder(descriptor: pass) else {
return false
}
encoder.setRenderPipelineState(planes.pq ? pipelinePlanarToneMap : pipelinePlanar)
encoder.setFragmentTexture(planes.y, index: 0)
encoder.setFragmentTexture(planes.cb, index: 1)
encoder.setFragmentTexture(planes.cr, index: 2)
encoder.setFragmentBytes(&csc, length: MemoryLayout<CscUniform>.stride, index: 0)
encoder.drawPrimitives(type: .triangle, vertexStart: 0, vertexCount: 3)
encoder.endEncoding()
let surface = slot.surface
let surfaceLayer = surfaceLayer // captured directly the handler must not retain self
let keepAlive: [Any] = [planes.y, planes.cb, planes.cr]
commandBuffer.addCompletedHandler { _ in
_ = keepAlive // ring textures pinned until the GPU finished sampling
DispatchQueue.main.async {
CATransaction.begin()
CATransaction.setDisableActions(true)
surfaceLayer.contents = surface
CATransaction.commit()
onPresented?(
Stage2Pipeline.realtimeNs(forDisplayLinkTimestamp: CACurrentMediaTime()))
}
}
commandBuffer.commit()
lastHandedOff = slotIndex
return true
}
/// (Re)build the pool at `size` 4 BGRA8 IOSurface render targets (one on glass, one queued
/// in CA, one rendering, one spare). RENDER THREAD. A failed allocation leaves the pool empty;
/// the caller returns false and the ring's putBack + display-link retry take over.
private func ensureSurfacePool(size: CGSize) {
guard size != surfacePoolSize else { return }
surfacePool.removeAll()
surfacePoolSize = size
lastHandedOff = nil
let w = Int(size.width)
let h = Int(size.height)
guard w > 0, h > 0 else { return }
// 256-byte row alignment satisfies both IOSurface and Metal linear-texture rules.
let bytesPerRow = ((w * 4) + 255) & ~255
let props: [String: Any] = [
kIOSurfaceWidth as String: w,
kIOSurfaceHeight as String: h,
kIOSurfaceBytesPerElement as String: 4,
kIOSurfaceBytesPerRow as String: bytesPerRow,
kIOSurfacePixelFormat as String: kCVPixelFormatType_32BGRA,
]
let desc = MTLTextureDescriptor.texture2DDescriptor(
pixelFormat: .bgra8Unorm, width: w, height: h, mipmapped: false)
desc.usage = [.renderTarget]
desc.storageMode = .shared
for _ in 0..<4 {
guard let surface = IOSurfaceCreate(props as CFDictionary),
let texture = device.makeTexture(descriptor: desc, iosurface: surface, plane: 0)
else {
surfacePool.removeAll()
return
}
surfacePool.append(SurfaceSlot(surface: surface, texture: texture))
}
}
/// Pick the slot to render into: never the one just handed to the layer (the compositor may
/// still scan it), prefer surfaces the window server isn't holding (`IOSurfaceIsInUse`), and
/// among those the least recently rendered. Falls back to the LRU busy slot rather than
/// stalling a visible glitch at worst, never a queue-up. RENDER THREAD.
private func takeSurfaceSlot() -> Int? {
guard !surfacePool.isEmpty else { return nil }
var free: Int?
var busy: Int?
for i in surfacePool.indices where i != lastHandedOff {
if !IOSurfaceIsInUse(surfacePool[i].surface) {
if free == nil || surfacePool[i].seq < surfacePool[free!].seq { free = i }
} else {
if busy == nil || surfacePool[i].seq < surfacePool[busy!].seq { busy = i }
}
}
guard let pick = free ?? busy else { return nil }
surfaceSeq += 1
surfacePool[pick].seq = surfaceSeq
return pick
}
#endif
/// The shared present tail of `render`/`renderPlanar`: size the drawable, encode one /// The shared present tail of `render`/`renderPlanar`: size the drawable, encode one
/// fullscreen triangle with `pipeline` (`bind` supplies the fragment resources), schedule /// fullscreen triangle with `pipeline` (`bind` supplies the fragment resources), schedule
/// the present and the on-glass callback. /// the present and the on-glass callback.
@@ -936,6 +895,36 @@ public final class MetalVideoPresenter {
#if DEBUG #if DEBUG
logSizeIfChanged(decoded: decodedSize, drawable: targetSize) logSizeIfChanged(decoded: decodedSize, drawable: targetSize)
#endif #endif
#if os(macOS)
// Windowed (composited) the DCP swapID-panic mitigation mechanism (see
// `WindowedPresentMode`). Toggle the layer property BEFORE vending a drawable so the
// vend matches how it will be presented; drained here on the render thread, flipped
// exactly once per mode change.
stagingLock.lock()
let windowedMode = windowedPresentStaged
stagingLock.unlock()
if windowedMode != windowedPresentActive {
windowedPresentActive = windowedMode
layer.presentsWithTransaction = windowedMode == .transaction
if windowedMode != .surface, !surfacePool.isEmpty {
// Leaving surface mode (fullscreen entry / mechanism A/B): drop the pool at 5K
// it holds >100 MB, and re-entering rebuilds it in one frame. SessionPresenter
// clears the surface layer's contents on main.
surfacePool.removeAll()
surfacePoolSize = .zero
lastHandedOff = nil
}
presenterLog.info(
"stage2: windowed present mode \(windowedMode.rawValue, privacy: .public) (DCP swapID-panic mitigation)")
}
if windowedMode == .surface {
// No image queue at all: render into a pooled IOSurface and swap it into the
// sibling layer's contents. The drawable/queue tail below never runs.
return encodeToSurface(
targetSize: targetSize, pipeline: pipeline, onPresented: onPresented,
keepAlive: keepAlive, bind: bind)
}
#endif
if let providedDrawable, if let providedDrawable,
providedDrawable.texture.pixelFormat != layer.pixelFormat { providedDrawable.texture.pixelFormat != layer.pixelFormat {
return false // config outran the vend (HDR flip) next vend has the new format return false // config outran the vend (HDR flip) next vend has the new format
@@ -974,6 +963,61 @@ public final class MetalVideoPresenter {
} }
#endif #endif
} }
// Keep the bound sources alive until the GPU finishes sampling (see the callers).
commandBuffer.addCompletedHandler { _ in _ = keepAlive }
#if os(macOS)
if windowedPresentActive == .transaction {
// Windowed DCP mitigation: present the drawable THROUGH a Core Animation transaction
// (`presentsWithTransaction`, set above) instead of the async image queue, so the swap
// commits with the layer tree and stays in lockstep with the compositor (no out-of-band
// flip to race WindowServer's swaps). Wait until the GPU work is scheduled (contents
// will be ready p50 ~0.1 ms), then present inside an EXPLICIT CATransaction ON THIS
// RENDER THREAD and `flush()`. `presentAtMediaTime` does not apply the transaction
// paces.
//
// Threading history, because BOTH failure modes shipped or nearly shipped:
// A bare `present()` from this thread (no transaction) never flushes nothing
// commits a runloop-less thread's implicit transaction, so drawables are never
// released; after maximumDrawableCount vends `nextDrawable()` blocks forever and
// the stream FREEZES (the fullscreenwindowed switch did exactly this).
// The explicit begin/commit alone is NOT enough either: this thread has an ACTIVE
// implicit transaction (the layer mutations above drawableSize/colour created
// it), so the explicit transaction NESTS inside it and its commit defers to the
// implicit one that never comes. The harness reproduced the exact freeze: every
// present reported presentedTime=0, nothing reached glass. `CATransaction.flush()`
// pushes the implicit transaction (present included) to the render server NOW.
// The original fix hopped to MAIN and presented there correct, but slow in the
// field (presents=55 @ fps=240, display_p50 18.6 ms on the 240 Hz Studio): each
// present lands a runloop turn late, and main's own implicit transaction batches
// enrolled presents at runloop-iteration rate. Kept as PUNKTFUNK_TXN_PRESENT=main.
// The off-main commit measured immune to main-thread churn in the harness
// (2026-07-21: glass p50 ~10 ms at 240 Hz full-size, cadence a clean 4.17 ms).
commandBuffer.commit()
let schedStart = CACurrentMediaTime()
commandBuffer.waitUntilScheduled()
let schedMs = (CACurrentMediaTime() - schedStart) * 1000
let commitStart = CACurrentMediaTime()
if txnPresentOnMain {
let presentedDrawable = drawable
DispatchQueue.main.async {
CATransaction.begin()
CATransaction.setDisableActions(true)
presentedDrawable.present()
CATransaction.commit()
}
} else {
CATransaction.begin()
CATransaction.setDisableActions(true)
drawable.present()
CATransaction.commit()
CATransaction.flush()
}
windowedDiag.record(
schedMs: schedMs, commitMs: (CACurrentMediaTime() - commitStart) * 1000,
mode: .transaction)
return true
}
#endif
// Scheduled on the vsync when the pipeline gave us the link's target (see the doc comment); // Scheduled on the vsync when the pipeline gave us the link's target (see the doc comment);
// immediate otherwise. A target already in the past presents immediately same thing. // immediate otherwise. A target already in the past presents immediately same thing.
if let presentAtMediaTime { if let presentAtMediaTime {
@@ -981,12 +1025,146 @@ public final class MetalVideoPresenter {
} else { } else {
commandBuffer.present(drawable) commandBuffer.present(drawable)
} }
// Keep the bound sources alive until the GPU finishes sampling (see the callers).
commandBuffer.addCompletedHandler { _ in _ = keepAlive }
commandBuffer.commit() commandBuffer.commit()
return true return true
} }
#if os(macOS)
/// The WINDOWED `surface` present tail (see `WindowedPresentMode.surface`): render with the
/// same per-frame pipeline into a pooled IOSurface and hand it to `surfaceLayer.contents`
/// from the command buffer's COMPLETION handler, inside an explicit CATransaction + flush
/// (the same off-main commit discipline as the transactional present an ordinary
/// damaged-layer update on WindowServer's own composite cadence, no image queue anywhere).
/// RENDER THREAD. `onPresented` is stamped right after the contents swap commits the
/// closest observable analogue of "reached glass" here (the composite follows within a
/// refresh, so the display-stage meters read slightly OPTIMISTIC in this mode).
///
/// The pool tracks `hdrActive`: bgra8 for SDR, rgba16Float tagged BT.2100 PQ for HDR
/// `configure` already ran, so the caller's `pipeline` attachment format always matches.
/// HDR OPEN RISK (why this whole mode is a prototype): whether the compositor honors the
/// PQ tag + EDR for plain-CALayer IOSurface contents needs an on-glass eyeball; the metal
/// layer underneath keeps `wantsExtendedDynamicRangeContent` as the EDR anchor (the harness
/// measured the display's EDR headroom engaging with this arrangement).
private func encodeToSurface(
targetSize: CGSize, pipeline: MTLRenderPipelineState,
onPresented: ((Int64?) -> Void)?,
keepAlive: [Any], bind: (MTLRenderCommandEncoder) -> Void
) -> Bool {
ensureSurfacePool(size: targetSize, hdr: hdrActive)
guard let slotIndex = takeSurfaceSlot(),
let commandBuffer = queue.makeCommandBuffer()
else { return false }
let slot = surfacePool[slotIndex]
let pass = MTLRenderPassDescriptor()
pass.colorAttachments[0].texture = slot.texture
pass.colorAttachments[0].loadAction = .clear
pass.colorAttachments[0].clearColor = MTLClearColor(red: 0, green: 0, blue: 0, alpha: 1)
pass.colorAttachments[0].storeAction = .store
guard let encoder = commandBuffer.makeRenderCommandEncoder(descriptor: pass) else {
return false
}
encoder.setRenderPipelineState(pipeline)
bind(encoder)
encoder.drawPrimitives(type: .triangle, vertexStart: 0, vertexCount: 3)
encoder.endEncoding()
let surface = slot.surface
let surfaceLayer = surfaceLayer // captured directly the handler must not retain self
let diag = windowedDiag
let commitStamp = CACurrentMediaTime()
commandBuffer.addCompletedHandler { _ in
_ = keepAlive // sources pinned until the GPU finished sampling
let completedAt = CACurrentMediaTime()
// Swap on THIS Metal completion thread: explicit transaction + flush, so the commit
// reaches the render server now, independent of main (completion handlers for one
// queue fire in execution order, so swaps can't reorder).
CATransaction.begin()
CATransaction.setDisableActions(true)
surfaceLayer.contents = surface
CATransaction.commit()
CATransaction.flush()
diag.record(
schedMs: (completedAt - commitStamp) * 1000,
commitMs: (CACurrentMediaTime() - completedAt) * 1000, mode: .surface)
onPresented?(Stage2Pipeline.realtimeNs(forDisplayLinkTimestamp: CACurrentMediaTime()))
}
commandBuffer.commit()
lastHandedOff = slotIndex
return true
}
/// (Re)build the pool at `size`/`hdr` 4 IOSurface render targets (one on glass, one
/// committed in CA, one rendering, one spare). RENDER THREAD. A failed allocation leaves the
/// pool empty; the caller returns false and the ring's putBack + display-link retry take
/// over.
private func ensureSurfacePool(size: CGSize, hdr: Bool) {
guard size != surfacePoolSize || hdr != surfacePoolHDR else { return }
surfacePool.removeAll()
surfacePoolSize = size
surfacePoolHDR = hdr
lastHandedOff = nil
let w = Int(size.width)
let h = Int(size.height)
guard w > 0, h > 0 else { return }
// rgba16Float (8 B/px) carries the PQ-encoded HDR samples; bgra8 the SDR ones. 256-byte
// row alignment satisfies both IOSurface and Metal linear-texture rules.
let bytesPerElement = hdr ? 8 : 4
let bytesPerRow = ((w * bytesPerElement) + 255) & ~255
let props: [String: Any] = [
kIOSurfaceWidth as String: w,
kIOSurfaceHeight as String: h,
kIOSurfaceBytesPerElement as String: bytesPerElement,
kIOSurfaceBytesPerRow as String: bytesPerRow,
kIOSurfacePixelFormat as String: hdr
? kCVPixelFormatType_64RGBAHalf : kCVPixelFormatType_32BGRA,
]
let desc = MTLTextureDescriptor.texture2DDescriptor(
pixelFormat: hdr ? .rgba16Float : .bgra8Unorm, width: w, height: h, mipmapped: false)
desc.usage = [.renderTarget]
desc.storageMode = .shared
for _ in 0..<4 {
guard let surface = IOSurfaceCreate(props as CFDictionary),
let texture = device.makeTexture(descriptor: desc, iosurface: surface, plane: 0)
else {
surfacePool.removeAll()
return
}
if hdr, let name = CGColorSpace(name: CGColorSpace.itur_2100_PQ)?.name {
// Tag the surface BT.2100 PQ so the compositor interprets the half-float
// samples as PQ-encoded HDR (the CALayer-contents analogue of the metal
// layer's colorspace).
IOSurfaceSetValue(surface, "IOSurfaceColorSpace" as CFString, name)
}
surfacePool.append(SurfaceSlot(surface: surface, texture: texture))
}
// The EDR request rides the SURFACE layer too (its contents are what composite); the
// metal layer underneath keeps its own from configureColor as the anchor. Layer flags
// are committed by the next swap's transaction flush.
surfaceLayer.wantsExtendedDynamicRangeContent = hdr
}
/// Pick the slot to render into: never the one just handed to the layer (the compositor may
/// still scan it), prefer surfaces the window server isn't holding (`IOSurfaceIsInUse`), and
/// among those the least recently rendered. Falls back to the LRU busy slot rather than
/// stalling a visible glitch at worst, never a queue-up. RENDER THREAD.
private func takeSurfaceSlot() -> Int? {
guard !surfacePool.isEmpty else { return nil }
var free: Int?
var busy: Int?
for i in surfacePool.indices where i != lastHandedOff {
if !IOSurfaceIsInUse(surfacePool[i].surface) {
if free == nil || surfacePool[i].seq < surfacePool[free!].seq { free = i }
} else {
if busy == nil || surfacePool[i].seq < surfacePool[busy!].seq { busy = i }
}
}
guard let pick = free ?? busy else { return nil }
surfaceSeq += 1
surfacePool[pick].seq = surfaceSeq
return pick
}
#endif
/// Returns the CVMetalTexture (not just its MTLTexture) so the caller can keep it alive past the /// Returns the CVMetalTexture (not just its MTLTexture) so the caller can keep it alive past the
/// draw the MTLTexture is only valid while its CVMetalTexture is retained. /// draw the MTLTexture is only valid while its CVMetalTexture is retained.
private func makeTexture( private func makeTexture(
@@ -142,20 +142,17 @@ enum PresentPriority: Equatable {
final class SessionPresenter { final class SessionPresenter {
/// Present pacing for this session. Stage-3 always means glass gating; under the stage-2 /// Present pacing for this session. Stage-3 always means glass gating; under the stage-2
/// default, macOS PyroWave sessions ALSO get glass gating a kernel-panic mitigation, not a /// default, macOS PyroWave sessions ALSO get glass gating for SMOOTHNESS, not as the panic
/// latency tweak. macOS's DCP panics ("mismatched swapID's" @UnifiedPipeline.cpp, the whole /// fix (that is the windowed transactional present see `setComposited`). PyroWave's wavelet
/// machine dies) when WindowServer's swap submissions race, and the reliable trigger is /// decode is near-instant Metal compute, so a network clump presents within the same
/// out-of-band CAMetalLayer presents (displaySyncEnabled=false mandatory for us, see /// millisecond, and it is the codec that sustains stream rates above the panel's refresh; the
/// MetalVideoPresenter's init) arriving faster than the compositor latches them in a /// glass gate admits one presented-but-undisplayed swap at a time (serialized on the on-glass
/// COMPOSITED (windowed) session. Arrival pacing does exactly that with PyroWave: the wavelet /// callback, 100 ms stale backstop) so those bursts coalesce in the newest-wins ring instead
/// decode is near-instant Metal compute, so a network clump of frames presents within the /// of flooding the queue. (Glass pacing was ALSO the original DCP-panic mitigation attempt
/// same millisecond, and PyroWave is the codec that sustains stream rates above the panel's /// disproven: a fully serialized stream still panicked, which is why the real fix moved to the
/// refresh. The glass gate admits one presented-but-undisplayed swap at a time (serialized on /// present mechanism.) An explicit stage-2 pick (setting/env) still forces arrival pacing
/// the on-glass callback, 100 ms stale backstop), which removes the racing pattern outright; /// that A/B lever must stay honest. VideoToolbox codecs keep arrival pacing: decode latency
/// frames the panel couldn't have shown anyway coalesce in the newest-wins ring. An explicit /// spaces their presents.
/// stage-2 pick (setting/env) still forces arrival pacing that A/B lever must stay honest.
/// VideoToolbox codecs keep arrival pacing: decode latency spaces their presents, and years
/// of stage-2 defaults there predate any panic report.
static func pacing( static func pacing(
for choice: PresenterChoice, explicit: PresenterChoice?, codec: VideoCodec for choice: PresenterChoice, explicit: PresenterChoice?, codec: VideoCodec
) -> PresentPacing { ) -> PresentPacing {
@@ -178,10 +175,25 @@ final class SessionPresenter {
/// that doesn't exist after the first Wi-Fi clump. Sub-refresh display latency needs pacing /// that doesn't exist after the first Wi-Fi clump. Sub-refresh display latency needs pacing
/// that can't queue at all that's stage-4 (`PresentPacing.deadline`), not a deeper gate. /// that can't queue at all that's stage-4 (`PresentPacing.deadline`), not a deeper gate.
/// ///
#if os(macOS)
/// Resolve the windowed (composited) present MECHANISM for this session the DCP
/// swapID-panic mitigation picker (see `WindowedPresentMode`). The
/// `PUNKTFUNK_WINDOWED_PRESENT=async|transaction|surface` env lever wins (dev A/B);
/// otherwise the user's safe-present setting: ON/unset `.transaction` (the validated
/// mitigation), OFF `.async` (the fast pre-mitigation path the panic returns on
/// affected high-refresh setups; the Settings caption says so). `.surface` is currently
/// env-only (prototype HDR-composite verification owed). Fullscreen always presents
/// async regardless (`setComposited`). Internal (not private) for unit tests.
static func windowedPresentMode(setting: Bool?, env: String?) -> WindowedPresentMode {
if let env, let mode = WindowedPresentMode(rawValue: env) { return mode }
return (setting ?? true) ? .transaction : .async
}
#endif
/// `PUNKTFUNK_GATE_DEPTH` (13) still overrides on iOS/tvOS so the standing-queue ladder /// `PUNKTFUNK_GATE_DEPTH` (13) still overrides on iOS/tvOS so the standing-queue ladder
/// stays reproducible on-device; macOS is pinned to 1, env ignored glass pacing exists /// stays reproducible on-device; macOS is pinned to 1, env ignored a deeper gate only builds
/// there as the DCP swapID kernel-panic mitigation (see `pacing`), and STRICT present /// a standing queue (see above), and macOS glass pacing exists for PyroWave smoothness
/// serialization is its point. Internal (not private) for unit tests. /// (see `pacing`), where depth 1 is the point. Internal (not private) for unit tests.
static func gateDepth(env: String?) -> Int { static func gateDepth(env: String?) -> Int {
#if os(macOS) #if os(macOS)
return 1 return 1
@@ -196,10 +208,16 @@ final class SessionPresenter {
private var stage2Link: CADisplayLink? private var stage2Link: CADisplayLink?
private var metalLayer: CAMetalLayer? private var metalLayer: CAMetalLayer?
#if os(macOS) #if os(macOS)
/// The windowed-mode PyroWave present target (sibling above `metalLayer`) and the last /// The windowed present MECHANISM this session runs while composited (resolved once per
/// routing pushed to the pipeline see `setComposited`. Main-thread only, like all of this. /// session in `start` the user's safe-present setting + the PUNKTFUNK_WINDOWED_PRESENT
/// dev override) and the routing last pushed to the pipeline see `setComposited` (the DCP
/// swapID-panic mitigation). Main-thread only, like all of this.
private var windowedMode: WindowedPresentMode = .transaction
private var windowedPresentApplied: WindowedPresentMode = .async
/// The windowed `surface` present target (sibling above `metalLayer`, transparent while
/// unused) installed whenever stage-2 runs so a mechanism flip never has to mutate the
/// layer tree mid-session.
private var surfaceLayer: CALayer? private var surfaceLayer: CALayer?
private var surfacePresentsActive = false
#endif #endif
private var connection: PunktfunkConnection? private var connection: PunktfunkConnection?
/// The decoded frame's REAL pixel dimensions (ground truth, pushed by the view from the pump's /// The decoded frame's REAL pixel dimensions (ground truth, pushed by the view from the pump's
@@ -283,11 +301,17 @@ final class SessionPresenter {
baseLayer.addSublayer(metal) baseLayer.addSublayer(metal)
metalLayer = metal metalLayer = metal
#if os(macOS) #if os(macOS)
// The windowed-PyroWave present target sits ABOVE the metal layer: transparent (nil windowedPresentApplied = .async
// contents) while the metal path presents, covering it while surface presents run. // Resolve THIS session's windowed mechanism once (setting + dev env lever)
// `setComposited` routes between it and fullscreen-async from every layout.
windowedMode = Self.windowedPresentMode(
setting: UserDefaults.standard.object(
forKey: DefaultsKey.windowedSafePresent) as? Bool,
env: ProcessInfo.processInfo.environment["PUNKTFUNK_WINDOWED_PRESENT"])
// The surface present target sits ABOVE the metal layer: transparent (nil contents)
// unless the surface mechanism actually presents, covering it while it does.
baseLayer.addSublayer(pipeline.surfaceLayer) baseLayer.addSublayer(pipeline.surfaceLayer)
surfaceLayer = pipeline.surfaceLayer surfaceLayer = pipeline.surfaceLayer
surfacePresentsActive = false
#endif #endif
stage2 = pipeline stage2 = pipeline
// The link is the vsync CLOCK + putBack-retry nudge, not the presentation trigger // The link is the vsync CLOCK + putBack-retry nudge, not the presentation trigger
@@ -432,19 +456,23 @@ final class SessionPresenter {
#if os(macOS) #if os(macOS)
/// Route presents for the window's composited state (MAIN thread the view pushes it on /// Route presents for the window's composited state (MAIN thread the view pushes it on
/// every layout, which fullscreen transitions always trigger). PyroWave sessions in a /// every layout, which fullscreen transitions always trigger). A COMPOSITED (windowed)
/// COMPOSITED (windowed) session present via `surfaceLayer` contents instead of the /// session presents through this session's resolved mitigation mechanism (`windowedMode`
/// CAMetalLayer image queue the DCP "mismatched swapID's" kernel-panic mitigation (see /// transactional by default, see `windowedPresentMode`) instead of the async image queue
/// `MetalVideoPresenter.surfaceLayer`; the metal-swap race survives glass pacing, so pacing /// the DCP "mismatched swapID's" kernel-panic mitigation (see `MetalVideoPresenter`; the
/// alone was not enough). VT codecs keep the metal path: no panic reports there, and their /// async-swap race survives glass pacing, so pacing alone was not enough). ALL codecs:
/// HDR/EDR presentation has no surface-contents equivalent wired. /// PyroWave hit it 2026-07-18 and windowed HEVC hit the same 240 Hz Mac Studio 2026-07-21
/// it is the async image queue itself, not any codec or present rate. Fullscreen keeps the
/// async path (direct scanout, lowest latency, no panic there). The full HDR/EDR render
/// path is preserved in every mechanism.
func setComposited(_ composited: Bool) { func setComposited(_ composited: Bool) {
guard let stage2, let connection else { return } guard let stage2 else { return }
let wantsSurface = composited && connection.videoCodec == .pyrowave let mode: WindowedPresentMode = composited ? windowedMode : .async
guard wantsSurface != surfacePresentsActive else { return } guard mode != windowedPresentApplied else { return }
surfacePresentsActive = wantsSurface let wasSurface = windowedPresentApplied == .surface
stage2.setSurfacePresents(wantsSurface) windowedPresentApplied = mode
if !wantsSurface { stage2.setWindowedPresent(mode)
if wasSurface {
// Uncover the metal layer NOW (its last drawable is still attached, so fullscreen // Uncover the metal layer NOW (its last drawable is still attached, so fullscreen
// entry shows the previous frame until the next present no black flash). // entry shows the previous frame until the next present no black flash).
CATransaction.begin() CATransaction.begin()
@@ -471,7 +499,7 @@ final class SessionPresenter {
#if os(macOS) #if os(macOS)
surfaceLayer?.removeFromSuperlayer() surfaceLayer?.removeFromSuperlayer()
surfaceLayer = nil surfaceLayer = nil
surfacePresentsActive = false windowedPresentApplied = .async
#endif #endif
connection = nil connection = nil
} }
@@ -1114,15 +1114,15 @@ public final class Stage2Pipeline {
} }
#if os(macOS) #if os(macOS)
/// The windowed-mode PyroWave present target (see `MetalVideoPresenter.surfaceLayer` the /// Forward the windowed present mechanism (MAIN thread see
/// DCP swapID-panic mitigation). The hosting view installs it as a sibling above `layer`. /// `MetalVideoPresenter.setWindowedPresent`, the DCP swapID-panic mitigation).
public var surfaceLayer: CALayer { presenter.surfaceLayer } func setWindowedPresent(_ mode: WindowedPresentMode) {
presenter.setWindowedPresent(mode)
/// Forward the windowed-vs-fullscreen present routing (MAIN thread see
/// `MetalVideoPresenter.setSurfacePresents`).
public func setSurfacePresents(_ on: Bool) {
presenter.setSurfacePresents(on)
} }
/// The windowed `surface` present target the hosting SessionPresenter installs as a sibling
/// ABOVE `layer` (transparent while unused see `MetalVideoPresenter.surfaceLayer`).
var surfaceLayer: CALayer { presenter.surfaceLayer }
#endif #endif
/// Forward the display's current EDR headroom to the presenter (MAIN thread a `UIScreen` /// Forward the display's current EDR headroom to the presenter (MAIN thread a `UIScreen`
@@ -1213,7 +1213,9 @@ public final class Stage2Pipeline {
let chunkAligned = let chunkAligned =
au.flags & PunktfunkConnection.userFlagChunkAligned != 0 au.flags & PunktfunkConnection.userFlagChunkAligned != 0
let ptsNs = au.ptsNs let ptsNs = au.ptsNs
let receivedNs = au.receivedNs // Decode stage starts at the PULL (matching the VT path's FrameContext
// receiptpull is the HUD's separate client-queue term, ABI v9 split).
let receivedNs = au.pulledNs
let flags = au.flags let flags = au.flags
let submitted = decoder.decode( let submitted = decoder.decode(
au: au.data, chunkAligned: chunkAligned, windowSize: windowSize au: au.data, chunkAligned: chunkAligned, windowSize: windowSize
@@ -32,9 +32,12 @@ public enum ReadyImage: @unchecked Sendable {
public struct ReadyFrame: @unchecked Sendable { public struct ReadyFrame: @unchecked Sendable {
/// Host capture clock (the AU's pts), in nanoseconds. /// Host capture clock (the AU's pts), in nanoseconds.
public let ptsNs: UInt64 public let ptsNs: UInt64
/// Client `CLOCK_REALTIME` instant the AU was received (`AccessUnit.receivedNs`, threaded /// Client `CLOCK_REALTIME` instant the AU left `nextAU` (`AccessUnit.pulledNs`, threaded
/// through the decode via the frame refcon), in nanoseconds. 0 when unknown (a caller that /// through the decode via the frame refcon), in nanoseconds the decode stage's start
/// didn't stamp receipt) the decode-stage meter then drops the sample via its sanity guard. /// point. (Named for its historical role; since the ABI v9 receipt split the true
/// reassembly receipt lives on `AccessUnit.receivedNs`, and receiptpull is the HUD's own
/// client-queue term.) 0 when unknown (a caller that didn't stamp) the decode-stage meter
/// then drops the sample via its sanity guard.
public let receivedNs: Int64 public let receivedNs: Int64
/// Client `CLOCK_REALTIME` instant decode completed, in nanoseconds. /// Client `CLOCK_REALTIME` instant decode completed, in nanoseconds.
public let decodedNs: Int64 public let decodedNs: Int64
@@ -167,7 +170,11 @@ public final class VideoDecoder: @unchecked Sendable {
var infoOut = VTDecodeInfoFlags() var infoOut = VTDecodeInfoFlags()
// The AU's receipt instant + wire flags ride through as a retained context; the output // The AU's receipt instant + wire flags ride through as a retained context; the output
// callback reclaims it. Retain immediately before submit so no early return can leak it. // callback reclaims it. Retain immediately before submit so no early return can leak it.
let ctx = FrameContext(receivedNs: au.receivedNs, flags: au.flags) // The decode stage starts at the PULL (the AU leaving nextAU), not the reassembly
// receipt: both consumers the decode-stage meter and the ABR decode signal are
// specified from the pull, and the receiptpull wait is the HUD's separate client-queue
// term (see AccessUnit.pulledNs).
let ctx = FrameContext(receivedNs: au.pulledNs, flags: au.flags)
let refcon = Unmanaged.passRetained(ctx).toOpaque() let refcon = Unmanaged.passRetained(ctx).toOpaque()
let status = VTDecompressionSessionDecodeFrame( let status = VTDecompressionSessionDecodeFrame(
session, session,
@@ -700,9 +700,9 @@ public final class StreamLayerView: NSView {
private func layoutPresenter() { private func layoutPresenter() {
presenter.layout(in: bounds, contentsScale: window?.backingScaleFactor ?? 1) presenter.layout(in: bounds, contentsScale: window?.backingScaleFactor ?? 1)
// Present routing tracks the window's composited state (fullscreen transitions always // Present routing tracks the window's composited state (fullscreen transitions always
// re-layout, so this stays current): windowed PyroWave presents via surface contents // re-layout, so this stays current): a windowed session presents through a Core Animation
// the DCP swapID kernel-panic mitigation (see SessionPresenter.setComposited). A view // transaction the DCP swapID kernel-panic mitigation (see SessionPresenter.setComposited).
// not yet in a window counts as composited (the safe default). // A view not yet in a window counts as composited (the safe default).
presenter.setComposited(!(window?.styleMask.contains(.fullScreen) ?? false)) presenter.setComposited(!(window?.styleMask.contains(.fullScreen) ?? false))
// Feed the follower only once in a window (backing scale is real then) and with real // Feed the follower only once in a window (backing scale is real then) and with real
// bounds a pre-window layout would report point-sized dimensions. // bounds a pre-window layout would report point-sized dimensions.
@@ -70,6 +70,16 @@ public enum DefaultsKey {
/// (lowest latency the default, OFF). Resolved once per session; /// (lowest latency the default, OFF). Resolved once per session;
/// PUNKTFUNK_PRESENT_MODE=immediate|vsync overrides it for A/B. See Stage2Pipeline's header. /// PUNKTFUNK_PRESENT_MODE=immediate|vsync overrides it for A/B. See Stage2Pipeline's header.
public static let vsync = "punktfunk.vsync" public static let vsync = "punktfunk.vsync"
/// macOS: present WINDOWED sessions in lockstep with the system compositor (the DCP
/// "mismatched swapID's" kernel-panic mitigation see SessionPresenter.windowedPresentMode
/// and the MetalVideoPresenter saga notes). ON/unset (the default): windowed presents ride
/// a Core Animation transaction validated panic-free on the 240 Hz repro machine, at a
/// small display-latency cost vs the raw path. OFF: windowed sessions keep the fast async
/// image queue ON AFFECTED SETUPS (high-refresh displays) THAT PATH KERNEL-PANICS THE
/// WHOLE MAC, which is why the default is ON. Fullscreen always presents async (fast path)
/// regardless. Resolved once per session; PUNKTFUNK_WINDOWED_PRESENT=async|transaction|
/// surface overrides it for dev A/B.
public static let windowedSafePresent = "punktfunk.windowedSafePresent"
/// Allow variable refresh rate: hand the display link a wide frame-rate RANGE (low floor, /// Allow variable refresh rate: hand the display link a wide frame-rate RANGE (low floor,
/// preferred = stream rate) so a ProMotion / adaptive-sync display can vary its physical /// preferred = stream rate) so a ProMotion / adaptive-sync display can vary its physical
/// refresh to match the stream. On by default; a no-op on fixed-refresh displays. When off, /// refresh to match the stream. On by default; a no-op on fixed-refresh displays. When off,
@@ -316,6 +316,41 @@ final class PresentPacingTests: XCTestCase {
SessionPresenter.pacing(for: .stage4, explicit: .stage4, codec: .pyrowave), .deadline) SessionPresenter.pacing(for: .stage4, explicit: .stage4, codec: .pyrowave), .deadline)
} }
// MARK: - Windowed present mechanism (the macOS DCP swapID-panic mitigation picker)
#if os(macOS)
/// The safe-present setting: ON/unset the validated transactional mitigation; an explicit
/// OFF the fast async path (the user accepted the affected-setup panic risk). The
/// PUNKTFUNK_WINDOWED_PRESENT env lever overrides both ways, `surface` is env-only (the
/// prototype mechanism), and garbage/empty env values are "unset", not an override.
func testWindowedPresentModeResolution() {
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: nil, env: nil), .transaction,
"unset defaults to the panic mitigation")
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: true, env: nil), .transaction)
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: false, env: nil), .async,
"an explicit opt-out gets the fast async path")
// The dev env lever wins over the setting, both directions.
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: true, env: "async"), .async)
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: false, env: "transaction"),
.transaction)
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: true, env: "surface"), .surface,
"the surface prototype is reachable via env only")
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: false, env: "surface"), .surface)
// Garbage/empty env = unset.
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: nil, env: "garbage"), .transaction)
XCTAssertEqual(
SessionPresenter.windowedPresentMode(setting: false, env: ""), .async)
}
#endif
// MARK: - Glass-gate depth // MARK: - Glass-gate depth
/// The in-flight present budget is 1 EVERYWHERE: any deeper gate keeps a standing queue /// The in-flight present budget is 1 EVERYWHERE: any deeper gate keeps a standing queue
+9 -1
View File
@@ -856,7 +856,15 @@ pub fn show(
s.render_scale = s.render_scale =
RENDER_SCALES[(scale_row.selected() as usize).min(RENDER_SCALES.len() - 1)]; RENDER_SCALES[(scale_row.selected() as usize).min(RENDER_SCALES.len() - 1)];
s.bitrate_kbps = (bitrate_row.value() * 1000.0) as u32; s.bitrate_kbps = (bitrate_row.value() * 1000.0) as u32;
s.gamepad = GAMEPADS[(pad_row.selected() as usize).min(GAMEPADS.len() - 1)].to_string(); // Keep a stored preference this table doesn't list (e.g. "switchpro" — valid to the
// session, hand-edited or written by another client): it displays as "Automatic", and
// writing that back would silently erase it just by opening + closing the dialog.
// Persist the row only when the user picked a non-Auto entry or the stored value was
// a listed one to begin with.
let pad_sel = (pad_row.selected() as usize).min(GAMEPADS.len() - 1);
if pad_sel != 0 || GAMEPADS.contains(&s.gamepad.as_str()) {
s.gamepad = GAMEPADS[pad_sel].to_string();
}
s.touch_mode = s.touch_mode =
TOUCH_MODES[(touch_row.selected() as usize).min(TOUCH_MODES.len() - 1)].to_string(); TOUCH_MODES[(touch_row.selected() as usize).min(TOUCH_MODES.len() - 1)].to_string();
s.forward_pad = chosen_pin.borrow().clone(); s.forward_pad = chosen_pin.borrow().clone();
+14 -13
View File
@@ -458,7 +458,11 @@ async fn session(args: Args) -> Result<()> {
), ),
(None, None) => tracing::info!(%remote, "punktfunk/1 connected"), (None, None) => tracing::info!(%remote, "punktfunk/1 connected"),
} }
let (mut send, mut recv) = conn.open_bi().await.context("open control stream")?; let (mut send, recv) = conn.open_bi().await.context("open control stream")?;
// Frame every read on the control stream through the resumable reader, exactly as the client
// pump does: `clock_sync` bounds each read with a timeout, and a frame straddling two wakeups
// would otherwise leave the stream permanently misaligned for the rest of the run.
let mut recv = io::MsgReader::new(recv);
io::write_msg( io::write_msg(
&mut send, &mut send,
@@ -483,8 +487,11 @@ async fn session(args: Args) -> Result<()> {
// host/network split is exactly what it exists to report. Old hosts ignore the bit. // host/network split is exactly what it exists to report. Old hosts ignore the bit.
// PROBE_SEQ: the shared-core reassembler windows probe-space frames, so the probe // PROBE_SEQ: the shared-core reassembler windows probe-space frames, so the probe
// qualifies for `--speed-test` bursts; without the bit the host declines them. // qualifies for `--speed-test` bursts; without the bit the host declines them.
// STREAMED_AU: the same shared reassembler accepts sentinel-headed streamed
// blocks, and the probe is exactly the tool that measures the overlap win.
let mut caps = punktfunk_core::quic::VIDEO_CAP_HOST_TIMING let mut caps = punktfunk_core::quic::VIDEO_CAP_HOST_TIMING
| punktfunk_core::quic::VIDEO_CAP_PROBE_SEQ; | punktfunk_core::quic::VIDEO_CAP_PROBE_SEQ
| punktfunk_core::quic::VIDEO_CAP_STREAMED_AU;
if std::env::var_os("PUNKTFUNK_CLIENT_10BIT").is_some() { if std::env::var_os("PUNKTFUNK_CLIENT_10BIT").is_some() {
caps |= punktfunk_core::quic::VIDEO_CAP_10BIT; caps |= punktfunk_core::quic::VIDEO_CAP_10BIT;
} }
@@ -513,8 +520,8 @@ async fn session(args: Args) -> Result<()> {
.encode(), .encode(),
) )
.await?; .await?;
let welcome = Welcome::decode(&io::read_msg(&mut recv).await?) let welcome =
.map_err(|e| anyhow!("Welcome decode: {e:?}"))?; Welcome::decode(&recv.read_msg().await?).map_err(|e| anyhow!("Welcome decode: {e:?}"))?;
tracing::info!( tracing::info!(
mode = ?welcome.mode, mode = ?welcome.mode,
fec = ?welcome.fec, fec = ?welcome.fec,
@@ -629,10 +636,7 @@ async fn session(args: Args) -> Result<()> {
tracing::error!("Reconfigure write failed"); tracing::error!("Reconfigure write failed");
return; return;
} }
match io::read_msg(&mut rr) match rr.read_msg().await.map(|b| Reconfigured::decode(&b)) {
.await
.map(|b| Reconfigured::decode(&b))
{
Ok(Ok(ack)) if ack.accepted => { Ok(Ok(ack)) if ack.accepted => {
tracing::info!(mode = ?ack.mode, "mode switch ACCEPTED") tracing::info!(mode = ?ack.mode, "mode switch ACCEPTED")
} }
@@ -685,10 +689,7 @@ async fn session(args: Args) -> Result<()> {
tracing::error!("SetBitrate write failed"); tracing::error!("SetBitrate write failed");
return; return;
} }
match io::read_msg(&mut rr) match rr.read_msg().await.map(|b| BitrateChanged::decode(&b)) {
.await
.map(|b| BitrateChanged::decode(&b))
{
Ok(Ok(ack)) => tracing::info!( Ok(Ok(ack)) => tracing::info!(
applied_kbps = ack.bitrate_kbps, applied_kbps = ack.bitrate_kbps,
"BITRATE CHANGE acked by host" "BITRATE CHANGE acked by host"
@@ -750,7 +751,7 @@ async fn session(args: Args) -> Result<()> {
tracing::error!("ProbeRequest write failed"); tracing::error!("ProbeRequest write failed");
return; return;
} }
let res = match io::read_msg(&mut sr).await.map(|b| ProbeResult::decode(&b)) { let res = match sr.read_msg().await.map(|b| ProbeResult::decode(&b)) {
Ok(Ok(r)) => r, Ok(Ok(r)) => r,
other => { other => {
tracing::error!(?other, "bad ProbeResult"); tracing::error!(?other, "bad ProbeResult");
+6 -2
View File
@@ -27,9 +27,13 @@ ui = ["dep:pf-console-ui", "dep:serde_json"]
# Same Linux+Windows gating as the rest of the client stack; elsewhere this is a stub # Same Linux+Windows gating as the rest of the client stack; elsewhere this is a stub
# binary. # binary.
[target.'cfg(any(target_os = "linux", windows))'.dependencies] [target.'cfg(any(target_os = "linux", windows))'.dependencies]
pf-presenter = { path = "../../crates/pf-presenter" } # `default-features = false` on both: THIS crate's `pyrowave` feature (above) is the single
# switch that turns the wavelet codec on, and it enables it explicitly on each. Inheriting their
# defaults instead would make `--no-default-features` a lie — the Windows ARM64 leg builds that
# way precisely to skip the vendored PyroWave C++, which has no ARM64 SIMD path.
pf-presenter = { path = "../../crates/pf-presenter", default-features = false }
pf-console-ui = { path = "../../crates/pf-console-ui", optional = true } pf-console-ui = { path = "../../crates/pf-console-ui", optional = true }
pf-client-core = { path = "../../crates/pf-client-core" } pf-client-core = { path = "../../crates/pf-client-core", default-features = false }
punktfunk-core = { path = "../../crates/punktfunk-core", features = ["quic"] } punktfunk-core = { path = "../../crates/punktfunk-core", features = ["quic"] }
# The fake-library dev hook (`PUNKTFUNK_FAKE_LIBRARY`, browse mode) parses GameEntry JSON. # The fake-library dev hook (`PUNKTFUNK_FAKE_LIBRARY`, browse mode) parses GameEntry JSON.
serde_json = { version = "1", optional = true } serde_json = { version = "1", optional = true }
+11 -1
View File
@@ -29,7 +29,17 @@ punktfunk-core = { path = "../../crates/punktfunk-core", features = ["quic"] }
# The shared client service layer: the trust/settings stores (ONE `Settings` struct for the # The shared client service layer: the trust/settings stores (ONE `Settings` struct for the
# shell and the spawned session binary — src/trust.rs re-exports it) and the game-library # shell and the spawned session binary — src/trust.rs re-exports it) and the game-library
# data model (fetch + art pipeline) behind the library page. # data model (fetch + art pipeline) behind the library page.
pf-client-core = { path = "../../crates/pf-client-core" } #
# `default-features = false` drops pf-client-core's default `pyrowave`, which would otherwise
# build the vendored PyroWave C++ INTO THE SHELL — dead weight here (the shell never decodes;
# it only offers "pyrowave" as a codec preference string the session binary acts on) and fatal
# on ARM64, where Granite's math falls back to x86 SSE intrinsics and stops at
# `simd.hpp: #error "Implement me."`. This does NOT drop PyroWave from the Windows client:
# decode lives in the spawned punktfunk-session binary, whose own default enables the feature,
# and cargo's feature unification turns it back on for the shared pf-client-core whenever that
# binary is in the same build (x64). On the ARM64 leg both are built --no-default-features, so
# nothing enables it and the C++ is never compiled.
pf-client-core = { path = "../../crates/pf-client-core", default-features = false }
# WinUI 3 UI via windows-reactor (a declarative React-like framework backed by WinUI). Its # WinUI 3 UI via windows-reactor (a declarative React-like framework backed by WinUI). Its
# `build.rs` downloads the Windows App SDK NuGets and stages the bootstrap DLL + resources.pri # `build.rs` downloads the Windows App SDK NuGets and stages the bootstrap DLL + resources.pri
+8 -3
View File
@@ -59,9 +59,14 @@
</Application> </Application>
<!-- <!--
Second entry point: the couch/console UI, for an HTPC or a TV-attached box where the Second entry point: the couch/console UI, for an HTPC or a TV-attached box where the
desktop shell is the wrong first screen. Same full-trust executable, launched with desktop shell is the wrong first screen. Its own executable (punktfunk-console.exe)
`--console`, which hands straight off to the session binary's controller-driven because an MSIX Application entry cannot pass arguments to a full-trust exe; it hands
browse mode (host list, pairing, settings, library) fullscreen. straight off to the session binary's controller-driven browse mode (host list,
pairing, settings, library) fullscreen.
NOTE: never write a double hyphen in this file. XML forbids it inside a comment, and
makepri rejects the whole manifest ("Appx manifest not found or is invalid") — which
is exactly how the console flag spelled out here broke the v0.15.0 MSIX build.
--> -->
<Application Id="PunktfunkConsole" Executable="punktfunk-console.exe" <Application Id="PunktfunkConsole" Executable="punktfunk-console.exe"
EntryPoint="Windows.FullTrustApplication"> EntryPoint="Windows.FullTrustApplication">
+16
View File
@@ -24,6 +24,16 @@ use pf_frame::DmabufFrame;
pub trait Capturer: Send { pub trait Capturer: Send {
fn next_frame(&mut self) -> Result<CapturedFrame>; fn next_frame(&mut self) -> Result<CapturedFrame>;
/// [`next_frame`](Self::next_frame) with a caller-chosen first-frame budget instead of the
/// backend's default. The pipeline retry loop shortens its FIRST attempt's wait: a PipeWire
/// stream connected while gamescope re-inits its headless takeover can negotiate a format,
/// reach `Streaming`, and still never receive a buffer — a fresh connect then delivers within
/// ~0.5 s, so waiting out the full default budget on a doomed stream just delays the retry
/// that fixes it. Backends without an internal wait budget ignore it (the default delegates).
fn next_frame_within(&mut self, _budget: std::time::Duration) -> Result<CapturedFrame> {
self.next_frame()
}
/// Non-blocking: the freshest frame available since the last call, or `None` if none has /// Non-blocking: the freshest frame available since the last call, or `None` if none has
/// arrived (the caller reuses its last frame to hold a steady output rate). The default /// arrived (the caller reuses its last frame to hold a steady output rate). The default
/// just produces a frame each call — fine for instant synthetic sources; the portal /// just produces a frame each call — fine for instant synthetic sources; the portal
@@ -249,6 +259,12 @@ pub struct ZeroCopyPolicy {
/// passthrough (like the VAAPI backend) instead of the EGL→CUDA import whose payloads only /// passthrough (like the VAAPI backend) instead of the EGL→CUDA import whose payloads only
/// NVENC can consume. Per-session (the codec is negotiated), unlike `backend_is_vaapi`. /// NVENC can consume. Per-session (the codec is negotiated), unlike `backend_is_vaapi`.
pub pyrowave_session: bool, pub pyrowave_session: bool,
/// THIS session's encoder can ingest a producer-native NV12 capture (the Linux raw Vulkan
/// Video backend on an H265/AV1 session — resolved by the host facade via
/// `pf_encode::linux_native_nv12_ok`). Gates whether the negotiation PREFERS gamescope's
/// producer-side NV12 pod: libav VAAPI (H264's backend) would misread the two-plane buffer,
/// so H264/GameStream/PyroWave sessions must never see NV12 frames.
pub native_nv12_session: bool,
/// The PyroWave encoder's Vulkan-importable dmabuf modifiers for the capture's packed-RGB fourcc, /// The PyroWave encoder's Vulkan-importable dmabuf modifiers for the capture's packed-RGB fourcc,
/// resolved when the session encodes PyroWave (the passthrough advertises them so Mutter+NVIDIA, /// resolved when the session encodes PyroWave (the passthrough advertises them so Mutter+NVIDIA,
/// which allocates tiled-only, still negotiates zero-copy). Empty otherwise. /// which allocates tiled-only, still negotiates zero-copy). Empty otherwise.
+229 -97
View File
@@ -299,29 +299,11 @@ fn spawn_pipewire(
impl Capturer for PortalCapturer { impl Capturer for PortalCapturer {
fn next_frame(&mut self) -> Result<CapturedFrame> { fn next_frame(&mut self) -> Result<CapturedFrame> {
// First frame can lag behind format negotiation; later frames arrive at ~fps. Wait in self.frame_within(Duration::from_secs(10))
// short slices so a GPU-import poison (worker death) fails the capture within ~0.5 s
// instead of sitting out the full first-frame budget.
let deadline = std::time::Instant::now() + Duration::from_secs(10);
loop {
if self.broken.load(Ordering::Relaxed) {
return Err(anyhow!(
"zero-copy GPU import lost (node {}): the import worker died or tiled imports \
failed repeatedly rebuilding capture",
self.node_id
));
}
if let Some(f) = self.pending.take() {
return Ok(f); // a wait_arrival stash outranks the channel (it's older)
}
let slice = Duration::from_millis(500)
.min(deadline.saturating_duration_since(std::time::Instant::now()));
match self.frames.recv_timeout(slice) {
Ok(frame) => return Ok(frame),
Err(RecvTimeoutError::Timeout) if std::time::Instant::now() < deadline => continue,
Err(e) => return self.next_frame_timed_out(e),
}
} }
fn next_frame_within(&mut self, budget: Duration) -> Result<CapturedFrame> {
self.frame_within(budget)
} }
fn supports_arrival_wait(&self) -> bool { fn supports_arrival_wait(&self) -> bool {
@@ -417,9 +399,41 @@ impl Capturer for PortalCapturer {
} }
impl PortalCapturer { impl PortalCapturer {
/// The [`Capturer::next_frame`] budget expired (or the thread ended) — turn it into the /// The blocking first-frame wait behind [`Capturer::next_frame`] /
/// diagnosis-bearing error. Split out of the slicing loop above; behavior unchanged. /// [`Capturer::next_frame_within`]. First frame can lag behind format negotiation; later
fn next_frame_timed_out(&self, err: RecvTimeoutError) -> Result<CapturedFrame> { /// frames arrive at ~fps. Wait in short slices so a GPU-import poison (worker death) fails
/// the capture within ~0.5 s instead of sitting out the full first-frame budget.
fn frame_within(&mut self, budget: Duration) -> Result<CapturedFrame> {
let deadline = std::time::Instant::now() + budget;
loop {
if self.broken.load(Ordering::Relaxed) {
return Err(anyhow!(
"zero-copy GPU import lost (node {}): the import worker died or tiled imports \
failed repeatedly rebuilding capture",
self.node_id
));
}
if let Some(f) = self.pending.take() {
return Ok(f); // a wait_arrival stash outranks the channel (it's older)
}
let slice = Duration::from_millis(500)
.min(deadline.saturating_duration_since(std::time::Instant::now()));
match self.frames.recv_timeout(slice) {
Ok(frame) => return Ok(frame),
Err(RecvTimeoutError::Timeout) if std::time::Instant::now() < deadline => continue,
Err(e) => return self.next_frame_timed_out(e, budget),
}
}
}
/// The [`frame_within`](Self::frame_within) budget expired (or the thread ended) — turn it
/// into the diagnosis-bearing error. Split out of the slicing loop above; behavior unchanged.
fn next_frame_timed_out(
&self,
err: RecvTimeoutError,
budget: Duration,
) -> Result<CapturedFrame> {
let within = budget.as_secs_f32();
match err { match err {
RecvTimeoutError::Timeout => { RecvTimeoutError::Timeout => {
// Split the two black-screen root causes apart so the operator gets a cause, not // Split the two black-screen root causes apart so the operator gets a cause, not
@@ -427,9 +441,10 @@ impl PortalCapturer {
// not (no acceptable format / node never emitted a param)? // not (no acceptable format / node never emitted a param)?
if self.negotiated.load(Ordering::Relaxed) { if self.negotiated.load(Ordering::Relaxed) {
Err(anyhow!( Err(anyhow!(
"no PipeWire frame within 10s (node {}): format negotiated but no buffers \ "no PipeWire frame within {within}s (node {}): format negotiated but no \
arrived the compositor produced no frames (virtual output idle/unmapped, \ buffers arrived the compositor produced no frames (virtual output \
or capture never started)", idle/unmapped, capture never started, or a stream bound during a \
compositor (re)start that will never deliver a reconnect fixes that)",
self.node_id self.node_id
)) ))
} else if self.hdr_offer { } else if self.hdr_offer {
@@ -440,10 +455,10 @@ impl PortalCapturer {
// auto-reconnects) negotiates SDR instead of re-running this same timeout. // auto-reconnects) negotiates SDR instead of re-running this same timeout.
super::note_hdr_capture_failed(); super::note_hdr_capture_failed();
Err(anyhow!( Err(anyhow!(
"no PipeWire frame within 10s (node {}): the compositor never accepted \ "no PipeWire frame within {within}s (node {}): the compositor never \
the HDR (10-bit PQ/BT.2020 dmabuf) offer is the mirrored monitor in \ accepted the HDR (10-bit PQ/BT.2020 dmabuf) offer is the mirrored \
HDR mode on GNOME 50+? Downgrading this host to SDR capture; reconnect \ monitor in HDR mode on GNOME 50+? Downgrading this host to SDR capture; \
to stream SDR", reconnect to stream SDR",
self.node_id self.node_id
)) ))
} else if self.vaapi_dmabuf && !pf_zerocopy::vaapi_dmabuf_forced() { } else if self.vaapi_dmabuf && !pf_zerocopy::vaapi_dmabuf_forced() {
@@ -452,14 +467,15 @@ impl PortalCapturer {
// retries on the CPU offer instead of failing this same negotiation forever. // retries on the CPU offer instead of failing this same negotiation forever.
pf_zerocopy::note_vaapi_dmabuf_failed(); pf_zerocopy::note_vaapi_dmabuf_failed();
Err(anyhow!( Err(anyhow!(
"no PipeWire frame within 10s (node {}): the compositor never accepted \ "no PipeWire frame within {within}s (node {}): the compositor never \
the LINEAR-dmabuf offer (VAAPI zero-copy) downgrading this host to the \ accepted the LINEAR-dmabuf offer (VAAPI zero-copy) downgrading this \
CPU capture path; the pipeline rebuild will renegotiate without dmabuf", host to the CPU capture path; the pipeline rebuild will renegotiate \
without dmabuf",
self.node_id self.node_id
)) ))
} else { } else {
Err(anyhow!( Err(anyhow!(
"no PipeWire frame within 10s (node {}): format negotiation never \ "no PipeWire frame within {within}s (node {}): format negotiation never \
completed the compositor offered no format this consumer accepts \ completed the compositor offered no format this consumer accepts \
(pixel-format/modifier mismatch) or the node never emitted a Format param", (pixel-format/modifier mismatch) or the node never emitted a Format param",
self.node_id self.node_id
@@ -824,6 +840,7 @@ mod pipewire {
VideoFormat::RGBA => PixelFormat::Rgba, VideoFormat::RGBA => PixelFormat::Rgba,
VideoFormat::RGB => PixelFormat::Rgb, VideoFormat::RGB => PixelFormat::Rgb,
VideoFormat::BGR => PixelFormat::Bgr, VideoFormat::BGR => PixelFormat::Bgr,
VideoFormat::NV12 => PixelFormat::Nv12,
// The GNOME 50+ HDR screencast formats (packed 2:10:10:10; only ever negotiated by // The GNOME 50+ HDR screencast formats (packed 2:10:10:10; only ever negotiated by
// the `want_hdr` offer, whose MANDATORY colorimetry props pin them to PQ/BT.2020). // the `want_hdr` offer, whose MANDATORY colorimetry props pin them to PQ/BT.2020).
VideoFormat::xRGB_210LE => PixelFormat::X2Rgb10, VideoFormat::xRGB_210LE => PixelFormat::X2Rgb10,
@@ -989,10 +1006,10 @@ mod pipewire {
.into_inner()) .into_inner())
} }
/// Build a BGRx dmabuf `EnumFormat` pod advertising the EGL-importable `modifiers` as a /// Build a LINEAR/modifier DMA-BUF `EnumFormat` pod. Packed BGRx is the existing import path;
/// mandatory enum Choice; the compositor fixates to one of them that it can allocate, which /// NV12 is gamescope's producer-side RGB→YUV path (opt-in during bring-up).
/// we read back in `param_changed`.
fn build_dmabuf_format( fn build_dmabuf_format(
format: VideoFormat,
modifiers: &[u64], modifiers: &[u64],
preferred: Option<(u32, u32, u32)>, preferred: Option<(u32, u32, u32)>,
) -> Result<Vec<u8>> { ) -> Result<Vec<u8>> {
@@ -1003,7 +1020,7 @@ mod pipewire {
pw::spa::param::ParamType::EnumFormat, pw::spa::param::ParamType::EnumFormat,
pw::spa::pod::property!(FormatProperties::MediaType, Id, MediaType::Video), pw::spa::pod::property!(FormatProperties::MediaType, Id, MediaType::Video),
pw::spa::pod::property!(FormatProperties::MediaSubtype, Id, MediaSubtype::Raw), pw::spa::pod::property!(FormatProperties::MediaSubtype, Id, MediaSubtype::Raw),
pw::spa::pod::property!(FormatProperties::VideoFormat, Id, VideoFormat::BGRx), pw::spa::pod::property!(FormatProperties::VideoFormat, Id, format),
pw::spa::pod::property!( pw::spa::pod::property!(
FormatProperties::VideoSize, FormatProperties::VideoSize,
Choice, Choice,
@@ -1032,6 +1049,22 @@ mod pipewire {
pw::spa::utils::Fraction { num: 240, denom: 1 } pw::spa::utils::Fraction { num: 240, denom: 1 }
), ),
); );
if format == VideoFormat::NV12 {
obj.properties.push(pw::spa::pod::Property {
key: pw::spa::sys::SPA_FORMAT_VIDEO_colorMatrix,
flags: pw::spa::pod::PropertyFlags::MANDATORY,
value: pw::spa::pod::Value::Id(pw::spa::utils::Id(
pw::spa::sys::SPA_VIDEO_COLOR_MATRIX_BT709,
)),
});
obj.properties.push(pw::spa::pod::Property {
key: pw::spa::sys::SPA_FORMAT_VIDEO_colorRange,
flags: pw::spa::pod::PropertyFlags::MANDATORY,
value: pw::spa::pod::Value::Id(pw::spa::utils::Id(
pw::spa::sys::SPA_VIDEO_COLOR_RANGE_16_235,
)),
});
}
obj.properties.push(pw::spa::pod::Property { obj.properties.push(pw::spa::pod::Property {
key: pw::spa::sys::SPA_FORMAT_VIDEO_modifier, key: pw::spa::sys::SPA_FORMAT_VIDEO_modifier,
flags: pw::spa::pod::PropertyFlags::MANDATORY, flags: pw::spa::pod::PropertyFlags::MANDATORY,
@@ -1310,21 +1343,25 @@ mod pipewire {
/// (which Mutter delivers as metadata-only "corrupted" buffers) still refresh the position. /// (which Mutter delivers as metadata-only "corrupted" buffers) still refresh the position.
fn update_cursor_meta(cursor: &mut CursorState, spa_buf: *mut spa::sys::spa_buffer) { fn update_cursor_meta(cursor: &mut CursorState, spa_buf: *mut spa::sys::spa_buffer) {
// SAFETY: `spa_buf` is the live buffer we still hold (dequeued, not yet requeued). // SAFETY: `spa_buf` is the live buffer we still hold (dequeued, not yet requeued).
// `spa_buffer_find_meta_data` scans its metadata array for a `SPA_META_Cursor` of at least // `spa_buffer_find_meta` returns the `spa_meta` (type + byte `size` + `data` pointer) for
// `size_of::<spa_meta_cursor>()` bytes and returns a pointer into that buffer's metadata // `SPA_META_Cursor`, or null. We take `find_meta` rather than `find_meta_data` specifically
// (or null), valid until requeue. The size argument matches the struct the result is cast to. // to obtain the region's real `size`: the bitmap offset, pixel offset and stride read below
let cur = unsafe { // are ALL producer-written, and without a bound against the actual region they drive
spa::sys::spa_buffer_find_meta_data( // out-of-bounds pointer arithmetic and an oversized `slice::from_raw_parts` — an OOB read
spa_buf, // that SIGSEGVs inside the PipeWire `.process` callback (a segfault `catch_unwind` cannot
spa::sys::SPA_META_Cursor, // catch). Every offset below is validated against `region_size` with checked arithmetic,
std::mem::size_of::<spa::sys::spa_meta_cursor>(), // mirroring the fd-length guard the main frame path already applies to xdg-desktop-portal-wlr.
) as *const spa::sys::spa_meta_cursor let meta = unsafe { spa::sys::spa_buffer_find_meta(spa_buf, spa::sys::SPA_META_Cursor) };
}; if meta.is_null() {
if cur.is_null() {
return; return;
} }
// SAFETY: `cur` is non-null and points to a `spa_meta_cursor` of at least its own size // SAFETY: `meta` is non-null and points into the held buffer's metadata array.
// inside the held buffer (guaranteed by the size arg above), so every field read is in bounds. let (region_size, data) = unsafe { ((*meta).size as usize, (*meta).data as *const u8) };
if data.is_null() || region_size < std::mem::size_of::<spa::sys::spa_meta_cursor>() {
return;
}
let cur = data as *const spa::sys::spa_meta_cursor;
// SAFETY: `region_size >= size_of::<spa_meta_cursor>()` checked above, so every field is in bounds.
let (id, pos_x, pos_y, hot_x, hot_y, bmp_off) = unsafe { let (id, pos_x, pos_y, hot_x, hot_y, bmp_off) = unsafe {
( (
(*cur).id, (*cur).id,
@@ -1347,13 +1384,18 @@ mod pipewire {
// Position-only update — keep the cached bitmap. // Position-only update — keep the cached bitmap.
return; return;
} }
// SAFETY: `bitmap_offset` is a byte offset from `cur` to a `spa_meta_bitmap`, which the let bmp_off = bmp_off as usize;
// producer placed inside the same meta region it sized for this cursor (>= the size we // The `spa_meta_bitmap` header must fit entirely inside the region before we read it —
// requested). The resulting pointer is in bounds and aligned for `spa_meta_bitmap`. // `bitmap_offset` is producer-controlled and otherwise reads past the metadata.
let bmp = match bmp_off.checked_add(std::mem::size_of::<spa::sys::spa_meta_bitmap>()) {
unsafe { (cur as *const u8).add(bmp_off as usize) as *const spa::sys::spa_meta_bitmap }; Some(end) if end <= region_size => {}
// SAFETY: `bmp` is the in-bounds, aligned `spa_meta_bitmap` pointer computed just above; the _ => return,
// producer fully initialized this header, so reading its scalar fields is sound. }
// SAFETY: `bmp_off + size_of::<spa_meta_bitmap>() <= region_size` (checked directly above),
// so the header is fully in bounds; the producer places it aligned as before.
let bmp = unsafe { data.add(bmp_off) as *const spa::sys::spa_meta_bitmap };
// SAFETY: `bmp` is the in-bounds `spa_meta_bitmap` header validated just above; reading its
// scalar fields is sound.
let (vfmt, bw, bh, stride, pix_off) = unsafe { let (vfmt, bw, bh, stride, pix_off) = unsafe {
( (
(*bmp).format, (*bmp).format,
@@ -1369,10 +1411,27 @@ mod pipewire {
} }
let row = bw as usize * 4; let row = bw as usize * 4;
let stride = if stride < row { row } else { stride }; let stride = if stride < row { row } else { stride };
let span = stride * (bh as usize - 1) + row; // `span` is the exact byte extent the strided loop reads: `stride·(bh-1) + row`. Compute it
// SAFETY: the bitmap pixels live at `bmp + pix_off` for `span` bytes, within the // with checked arithmetic (a producer stride near `i32::MAX` would otherwise overflow) and
// producer-sized meta region. `span` is the exact extent the strided copy below reads. // require the whole pixel block `[bmp_off + pix_off, +span)` to lie inside the region before
let src = unsafe { std::slice::from_raw_parts((bmp as *const u8).add(pix_off), span) }; // fabricating the slice — this is the check whose absence made the read go out of bounds.
let span = match stride
.checked_mul(bh as usize - 1)
.and_then(|v| v.checked_add(row))
{
Some(s) => s,
None => return,
};
match bmp_off
.checked_add(pix_off)
.and_then(|v| v.checked_add(span))
{
Some(end) if end <= region_size => {}
_ => return,
}
// SAFETY: `bmp_off + pix_off + span <= region_size` (checked directly above), so the slice
// is fully within the producer's meta region; `span` is exactly the strided loop's extent.
let src = unsafe { std::slice::from_raw_parts(data.add(bmp_off + pix_off), span) };
let mut rgba = vec![0u8; bw as usize * bh as usize * 4]; let mut rgba = vec![0u8; bw as usize * bh as usize * 4];
for y in 0..bh as usize { for y in 0..bh as usize {
for x in 0..bw as usize { for x in 0..bw as usize {
@@ -1578,8 +1637,8 @@ mod pipewire {
} }
} }
// VAAPI zero-copy passthrough: hand the raw dmabuf straight to the encoder, which imports // Raw DMA-BUF passthrough: packed RGB is imported for GPU CSC; producer-native NV12 can
// it into a VA surface and does RGB→NV12 on the GPU video engine. No CUDA importer here. // be consumed by the Vulkan Video encoder without another color conversion.
if ud.vaapi_passthrough { if ud.vaapi_passthrough {
if let Some(fmt) = ud.format { if let Some(fmt) = ud.format {
if datas[0].type_() == pw::spa::buffer::DataType::DmaBuf { if datas[0].type_() == pw::spa::buffer::DataType::DmaBuf {
@@ -1587,9 +1646,41 @@ mod pipewire {
let chunk = datas[0].chunk(); let chunk = datas[0].chunk();
let offset = chunk.offset(); let offset = chunk.offset();
let stride = chunk.stride().max(0) as u32; let stride = chunk.stride().max(0) as u32;
// Native NV12 usually arrives as a two-plane SPA buffer over ONE buffer
// object; plane 1's chunk carries the REAL UV offset/stride (compositors
// may align the Y plane before UV). Pass it through instead of assuming
// contiguity. Each spa_data holds its own (dup'd) fd, so BO identity is
// by inode, not fd number; a genuinely two-BO frame cannot travel through
// the single-fd import — drop it with a diagnosis instead of streaming
// garbage chroma.
let plane1 =
if fmt == PixelFormat::Nv12 && datas.len() >= 2 && datas[1].fd() > 0 {
// SAFETY: zeroed `libc::stat` is a valid POD initializer; both fds are
// owned by the live PipeWire buffer for this callback, and `fstat`
// only writes the out-param structs, whose fields are read only after
// the `== 0` success checks.
let same_bo = unsafe {
let mut s0: libc::stat = std::mem::zeroed();
let mut s1: libc::stat = std::mem::zeroed();
libc::fstat(datas[0].fd() as i32, &mut s0) == 0
&& libc::fstat(datas[1].fd() as i32, &mut s1) == 0
&& (s0.st_dev, s0.st_ino) == (s1.st_dev, s1.st_ino)
};
if !same_bo {
warn_once(
"NV12 planes live in different buffer objects — frames \
dropped (single-fd import only)",
);
return;
}
let c1 = datas[1].chunk();
Some((c1.offset(), c1.stride().max(0) as u32))
} else {
None
};
// dup the fd so it survives the SPA buffer recycle — the encode thread // dup the fd so it survives the SPA buffer recycle — the encode thread
// imports it. (Content stability across the brief map+CSC window relies on // imports it. Content stability across the brief import/encode window relies
// the compositor's buffer-pool depth, like any zero-copy capture.) // on the compositor's buffer-pool depth, like any zero-copy capture.
// SAFETY: `datas[0].fd()` is the dmabuf fd owned by the live PipeWire buffer (valid // SAFETY: `datas[0].fd()` is the dmabuf fd owned by the live PipeWire buffer (valid
// for this callback). `fcntl(fd, F_DUPFD_CLOEXEC, 0)` reads only the integer fd, // for this callback). `fcntl(fd, F_DUPFD_CLOEXEC, 0)` reads only the integer fd,
// touches no Rust memory, and returns a fresh independent CLOEXEC duplicate (or -1). // touches no Rust memory, and returns a fresh independent CLOEXEC duplicate (or -1).
@@ -1616,9 +1707,10 @@ mod pipewire {
modifier: ud.modifier, modifier: ud.modifier,
offset, offset,
stride, stride,
plane1,
}), }),
// Cursor-as-metadata: the encoder blends this into its owned VA // Cursor-as-metadata is blended only by RGB→NV12 backends. Gamescope
// surface (raw dmabuf never touched). // embeds its pointer in the produced pixels, so native NV12 has none.
cursor: ud.cursor.overlay(), cursor: ud.cursor.overlay(),
}); });
static ONCE: std::sync::atomic::AtomicBool = static ONCE: std::sync::atomic::AtomicBool =
@@ -1629,7 +1721,12 @@ mod pipewire {
h, h,
modifier = ud.modifier, modifier = ud.modifier,
fourcc = format_args!("{:#010x}", fourcc), fourcc = format_args!("{:#010x}", fourcc),
"zero-copy: handing the raw dmabuf to the encoder (GPU import + CSC)" source = if fmt == PixelFormat::Nv12 {
"producer-native NV12"
} else {
"packed RGB (encoder GPU CSC)"
},
"zero-copy: handing the raw DMA-BUF to the encoder"
); );
} }
return; return;
@@ -1989,6 +2086,28 @@ mod pipewire {
// PyroWave session (the wavelet encoder's own Vulkan device, any vendor) → hand the raw // PyroWave session (the wavelet encoder's own Vulkan device, any vendor) → hand the raw
// dmabuf straight to the encoder. // dmabuf straight to the encoder.
let vaapi_passthrough = zerocopy && !force_shm && importer.is_none() && raw_passthrough; let vaapi_passthrough = zerocopy && !force_shm && importer.is_none() && raw_passthrough;
// Producer-side NV12 (default-on; PUNKTFUNK_PIPEWIRE_NV12=0 escapes): gamescope offers a
// one-fd LINEAR NV12 image when the consumer asks — its compositor pass does the RGB→YUV,
// and the Vulkan Video encoder imports the buffer as its encode source directly (no host
// CSC at all). `native_nv12_session` restricts this to sessions whose encoder can ingest
// it (Linux vulkan-encode H265/AV1 — never H264/libav-VAAPI, GameStream-resolve, or
// PyroWave, whose Vulkan compute CSC ingests packed RGB only). Raw passthrough is
// required because the CUDA importer expects packed RGB, and 4:4:4/HDR must not be
// silently subsampled/downconverted. Non-NV12 compositors (KWin/GNOME) simply match the
// packed-RGB fallback pod.
let prefer_native_nv12 = std::env::var("PUNKTFUNK_PIPEWIRE_NV12").as_deref() != Ok("0")
&& policy.native_nv12_session
&& backend_is_vaapi
&& vaapi_passthrough
&& !policy.pyrowave_session
&& !want_444
&& !want_hdr;
if prefer_native_nv12 {
tracing::info!(
"zero-copy: preferring gamescope producer-side NV12 LINEAR DMA-BUF (no host \
RGB CSC; PUNKTFUNK_PIPEWIRE_NV12=0 restores the packed-RGB negotiation)"
);
}
// Modifiers our import stack handles for BGRx: the EGL-importable (tiled) set, plus LINEAR // Modifiers our import stack handles for BGRx: the EGL-importable (tiled) set, plus LINEAR
// (0) — NVIDIA's EGL won't list it, but LINEAR dmabufs (gamescope's only offer) import via // (0) — NVIDIA's EGL won't list it, but LINEAR dmabufs (gamescope's only offer) import via
// CUDA external memory instead. For the VAAPI passthrough path we advertise LINEAR only: // CUDA external memory instead. For the VAAPI passthrough path we advertise LINEAR only:
@@ -2028,7 +2147,9 @@ mod pipewire {
tracing::warn!("zero-copy: no importable dmabuf modifiers — using CPU path"); tracing::warn!("zero-copy: no importable dmabuf modifiers — using CPU path");
} else if vaapi_passthrough && policy.pyrowave_modifiers.is_empty() { } else if vaapi_passthrough && policy.pyrowave_modifiers.is_empty() {
tracing::info!( tracing::info!(
"zero-copy: advertising LINEAR dmabuf for direct VAAPI import (GPU CSC)" native_nv12_preferred = prefer_native_nv12,
"zero-copy: advertising LINEAR DMA-BUF for encoder import (native NV12 first \
when enabled, packed RGB fallback)"
); );
} else if want_dmabuf && !vaapi_passthrough { } else if want_dmabuf && !vaapi_passthrough {
tracing::info!( tracing::info!(
@@ -2164,36 +2285,37 @@ mod pipewire {
} }
}) })
.process(|stream, ud| { .process(|stream, ud| {
// PipeWire dispatches this from a C trampoline with no catch_unwind; a // Latest-frame-only (OBS pattern): Mutter delivers buffers in bursts and recycles its
// panic crossing that FFI boundary would abort the whole host. Contain it. // pool; an older queued buffer carries a STALE frame. Drain all queued buffers, requeue
let outcome = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { // the older ones, keep only the newest. This dequeue/requeue runs OUTSIDE the
// Latest-frame-only (OBS pattern): Mutter delivers buffers in bursts and // `catch_unwind` below — they are non-panicking C FFI pointer ops, and `newest` is
// recycles its pool; an older queued buffer carries a STALE frame. Drain all // requeued exactly once AFTER the panic-containing region. Previously the whole thing was
// queued buffers, requeue the older ones, keep only the newest. // inside the catch, so a caught panic (in `update_cursor_meta`/`consume_frame`) stranded
// SAFETY: `stream` is the live stream PipeWire passes into this `.process` callback on // `newest` forever, permanently shrinking the stream's fixed pool until capture wedged.
// the loop thread, where `pw_stream_dequeue_buffer` is the documented call. It returns // SAFETY: `stream` is the live stream PipeWire passes into this `.process` callback on the
// a `*mut pw_buffer` owned by the stream (or null when the queue is drained), // loop thread; `dequeue_raw_buffer` returns a stream-owned `*mut pw_buffer` or null
// null-checked before any use. The loop is single-threaded, so no concurrent access. // (null-checked), single-threaded so no concurrent access.
let mut newest = unsafe { stream.dequeue_raw_buffer() }; let mut newest = unsafe { stream.dequeue_raw_buffer() };
if newest.is_null() { if newest.is_null() {
return; return;
} }
let mut drained = 1u32; let mut drained = 1u32;
loop { loop {
// SAFETY: same stream/loop-thread contract as the dequeue above; each call returns // SAFETY: same stream/loop-thread contract; returns the next stream-owned buffer or null.
// the next stream-owned `*mut pw_buffer` or null (null-checked before use).
let next = unsafe { stream.dequeue_raw_buffer() }; let next = unsafe { stream.dequeue_raw_buffer() };
if next.is_null() { if next.is_null() {
break; break;
} }
// SAFETY: `newest` is a non-null `*mut pw_buffer` previously dequeued from this same // SAFETY: `newest` was dequeued from this stream and not yet requeued; we immediately
// stream and not yet requeued; `pw_stream_queue_buffer` hands ownership back to the // overwrite it, so the requeued pointer is never touched again.
// stream. We immediately overwrite `newest = next`, so the requeued pointer is never
// touched again (no use-after-requeue). Loop thread, single-threaded.
unsafe { stream.queue_raw_buffer(newest) }; unsafe { stream.queue_raw_buffer(newest) };
newest = next; newest = next;
drained += 1; drained += 1;
} }
// PipeWire dispatches from a C trampoline with no catch_unwind; a panic crossing that FFI
// boundary would abort the whole host. Contain the inspect/consume work — the only Rust
// code here that can panic — and requeue `newest` unconditionally after it.
let outcome = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| {
// SAFETY: `newest` is the non-null buffer we still own (dequeued, not requeued); // SAFETY: `newest` is the non-null buffer we still own (dequeued, not requeued);
// `.buffer` is a `*mut spa_buffer` field libpipewire populated. This is a single field // `.buffer` is a `*mut spa_buffer` field libpipewire populated. This is a single field
// load through a valid pointer — no mutation or aliasing. // load through a valid pointer — no mutation or aliasing.
@@ -2272,19 +2394,18 @@ mod pipewire {
"capture: skipped a stale CORRUPTED/cursor buffer (GNOME)" "capture: skipped a stale CORRUPTED/cursor buffer (GNOME)"
); );
} }
// SAFETY: `newest` is the non-null buffer we own (dequeued, never requeued on this // Skip this stale/cursor buffer — `newest` is requeued unconditionally below.
// skip path); hand it back to the stream exactly once and return without touching it
// again. Loop thread inside `.process`.
unsafe { stream.queue_raw_buffer(newest) };
return; return;
} }
consume_frame(ud, spa_buf); consume_frame(ud, spa_buf);
// SAFETY: `consume_frame` has finished reading `spa_buf` (and the `datas` borrows derived
// from `newest`), so requeuing the owned `newest` exactly once here is sound — no
// use-after-requeue. Loop thread inside `.process`.
unsafe { stream.queue_raw_buffer(newest) };
})); }));
// Hand `newest` back to the stream exactly once, on EVERY path — normal, corrupted-skip,
// or a caught panic in the closure above. This single requeue is what keeps the fixed
// buffer pool from draining.
// SAFETY: all reads of `spa_buf`/`newest` (update_cursor_meta, consume_frame) completed
// inside the closure above; `newest` was dequeued from this stream and not yet requeued.
unsafe { stream.queue_raw_buffer(newest) };
if outcome.is_err() { if outcome.is_err() {
// In the per-frame `.process` callback: a deterministic panic (e.g. a bad // In the per-frame `.process` callback: a deterministic panic (e.g. a bad
// format) would fire this every frame, so power-of-two throttle it — enough to // format) would fire this every frame, so power-of-two throttle it — enough to
@@ -2368,7 +2489,18 @@ mod pipewire {
build_hdr_dmabuf_format(VideoFormat::xBGR_210LE, preferred)?, build_hdr_dmabuf_format(VideoFormat::xBGR_210LE, preferred)?,
] ]
} else if want_dmabuf { } else if want_dmabuf {
vec![build_dmabuf_format(&modifiers, preferred)?] let mut pods = Vec::with_capacity(if prefer_native_nv12 { 2 } else { 1 });
if prefer_native_nv12 {
// First compatible consumer pod wins. Gamescope advertises NV12 and BGRx; pinning
// BT.709 limited here selects its RGB→NV12 shader with our bitstream colorimetry.
pods.push(build_dmabuf_format(VideoFormat::NV12, &[0], preferred)?);
}
pods.push(build_dmabuf_format(
VideoFormat::BGRx,
&modifiers,
preferred,
)?);
pods
} else { } else {
vec![serialize_pod(obj)?] vec![serialize_pod(obj)?]
}; };
+209 -12
View File
@@ -308,8 +308,9 @@ float2 main(float4 pos : SV_POSITION, float2 uv : TEXCOORD0) : SV_TARGET {
/// Plane writes use per-plane render-target views of the single P010 texture: an `R16_UNORM` RTV /// Plane writes use per-plane render-target views of the single P010 texture: an `R16_UNORM` RTV
/// selects plane 0 (luma, full WxH), an `R16G16_UNORM` RTV selects plane 1 (chroma, W/2 x H/2). This /// selects plane 0 (luma, full WxH), an `R16G16_UNORM` RTV selects plane 1 (chroma, W/2 x H/2). This
/// planar-RTV mechanism needs a D3D11.3+ runtime + driver support; [`HdrP010Converter::convert`] /// planar-RTV mechanism needs a D3D11.3+ runtime + driver support; [`HdrP010Converter::convert`]
/// surfaces a clear error if `CreateRenderTargetView` rejects the plane format so the caller can fall /// surfaces a clear error if `CreateRenderTargetView` rejects the plane format. (There is no runtime
/// back to the existing R10 path. /// fallback — the error propagates through `try_consume` and ends the session; the "R10 path" the
/// original design referenced was never kept.)
pub(crate) struct HdrP010Converter { pub(crate) struct HdrP010Converter {
vs: ID3D11VertexShader, vs: ID3D11VertexShader,
ps_y: ID3D11PixelShader, ps_y: ID3D11PixelShader,
@@ -737,14 +738,157 @@ fn p010_reference(r: f64, g: f64, b: f64) -> (f64, f64, f64) {
/// Y ≤ 4 codes, U/V ≤ 5 codes (rounding + chroma averaging). Prints a per-colour table + PASS/FAIL. /// Y ≤ 4 codes, U/V ≤ 5 codes (rounding + chroma averaging). Prints a per-colour table + PASS/FAIL.
#[cfg(target_os = "windows")] #[cfg(target_os = "windows")]
pub fn hdr_p010_selftest() -> Result<()> { pub fn hdr_p010_selftest() -> Result<()> {
use windows::Win32::Graphics::Direct3D::D3D_DRIVER_TYPE_HARDWARE; hdr_p010_selftest_at(64, 64, None)
use windows::Win32::Graphics::Dxgi::IDXGIAdapter; }
// 64x64, even dims. A 4x4 grid of 16x16 flat scRGB blocks (each 2x2 chroma footprint uniform → /// [`hdr_p010_selftest`] at an arbitrary even size and (optionally) on a specific GPU vendor
// exact chroma comparison) covering pure R/G/B/white/black/gray at plausible HDR nit levels, plus /// (PCI vendor id, e.g. `0x8086` Intel / `0x10de` NVIDIA / `0x1002` AMD). The size matters on
// a couple of bright (>1.0 scRGB) colours, then the rest is a gradient (compared on Y only). /// top of the 64×64 default because the field sessions run at capture resolutions whose height
const W: u32 = 64; /// is NOT 16-aligned (1080 → the encoder's align16 pool seam) and a driver may treat the planar
const H: u32 = 64; /// RTVs differently at real sizes; the vendor pin matters on dual-GPU boxes where the default
/// adapter is not the one the session encodes on.
/// Test support (used by pf-encode's live e2e): the 8 sRGB colour bars (white/yellow/cyan/green/
/// magenta/red/blue/black, sRGB 1.0 = scRGB 1.0 = 80 nits) as a w×h FP16 scRGB texture on the
/// adapter with `luid`, converted through the REAL [`HdrP010Converter`] into a P010 texture with
/// **`BIND_RENDER_TARGET` only, `MiscFlags` 0 — the exact bind profile of the IDD out-ring** (the
/// CPU-upload encoder tests can't use that profile, so only this path exercises "RTV-written P010
/// → encoder ingest copy"). Returns `(device, p010)`; expected decoded codes per bar are the
/// bars_pq2020 fixture's: (490,512,512) (478,423,518) (464,525,473) (450,432,476) (350,584,585)
/// (325,448,598) (226,650,535) (64,512,512).
#[cfg(target_os = "windows")]
#[doc(hidden)]
pub fn hdr_p010_convert_bars_on_luid(
luid: [u8; 8],
w: u32,
h: u32,
) -> Result<(ID3D11Device, ID3D11Texture2D)> {
use windows::Win32::Graphics::Direct3D::D3D_DRIVER_TYPE_UNKNOWN;
use windows::Win32::Graphics::Dxgi::{CreateDXGIFactory1, IDXGIAdapter1, IDXGIFactory4};
if w == 0 || h == 0 || w % 2 != 0 || h % 2 != 0 {
bail!("bars pattern needs even non-zero dimensions, got {w}x{h}");
}
// sRGB primaries at full/zero channels: sRGB EOTF(1.0)=1.0, (0)=0 → the scRGB pattern is
// pure 0/1 floats and the PQ/BT.2020 reference codes above are exact.
const BARS: [(f32, f32, f32); 8] = [
(1.0, 1.0, 1.0),
(1.0, 1.0, 0.0),
(0.0, 1.0, 1.0),
(0.0, 1.0, 0.0),
(1.0, 0.0, 1.0),
(1.0, 0.0, 0.0),
(0.0, 0.0, 1.0),
(0.0, 0.0, 0.0),
];
let bar_w = (w / 8).max(1) as usize;
let mut fp16 = vec![0u16; (w * h * 4) as usize];
for y in 0..h as usize {
for x in 0..w as usize {
let (r, g, b) = BARS[(x / bar_w).min(7)];
let i = (y * w as usize + x) * 4;
fp16[i] = f32_to_f16(r);
fp16[i + 1] = f32_to_f16(g);
fp16[i + 2] = f32_to_f16(b);
fp16[i + 3] = f32_to_f16(1.0);
}
}
// SAFETY: same single-device/single-thread contract as `hdr_p010_selftest_at`; the FP16
// initial-data Vec outlives the synchronous CreateTexture2D; the returned COM handles own
// their references.
unsafe {
let luid = windows::Win32::Foundation::LUID {
LowPart: u32::from_le_bytes(luid[..4].try_into().unwrap()),
HighPart: i32::from_le_bytes(luid[4..].try_into().unwrap()),
};
let factory: IDXGIFactory4 = CreateDXGIFactory1().context("dxgi factory")?;
let adapter: IDXGIAdapter1 = factory.EnumAdapterByLuid(luid).context("adapter by luid")?;
let mut device: Option<ID3D11Device> = None;
let mut context: Option<ID3D11DeviceContext> = None;
D3D11CreateDevice(
&adapter,
D3D_DRIVER_TYPE_UNKNOWN,
HMODULE::default(),
D3D11_CREATE_DEVICE_BGRA_SUPPORT,
Some(&[D3D_FEATURE_LEVEL_11_0]),
D3D11_SDK_VERSION,
Some(&mut device),
None,
Some(&mut context),
)
.context("D3D11CreateDevice(luid) for bars convert")?;
let device = device.context("null device")?;
let context = context.context("null context")?;
let src_desc = D3D11_TEXTURE2D_DESC {
Width: w,
Height: h,
MipLevels: 1,
ArraySize: 1,
Format: DXGI_FORMAT_R16G16B16A16_FLOAT,
SampleDesc: DXGI_SAMPLE_DESC {
Count: 1,
Quality: 0,
},
Usage: D3D11_USAGE_DEFAULT,
BindFlags: D3D11_BIND_SHADER_RESOURCE.0 as u32,
..Default::default()
};
let init = D3D11_SUBRESOURCE_DATA {
pSysMem: fp16.as_ptr() as *const c_void,
SysMemPitch: w * 8,
SysMemSlicePitch: 0,
};
let mut src_tex: Option<ID3D11Texture2D> = None;
device
.CreateTexture2D(&src_desc, Some(&init), Some(&mut src_tex))
.context("CreateTexture2D(fp16 bars)")?;
let src_tex = src_tex.context("null src tex")?;
let mut src_srv: Option<ID3D11ShaderResourceView> = None;
device
.CreateShaderResourceView(&src_tex, None, Some(&mut src_srv))
.context("CreateShaderResourceView(fp16 bars)")?;
let src_srv = src_srv.context("null src srv")?;
// The IDD out-ring's exact profile: P010, RENDER_TARGET only, MiscFlags 0.
let p010_desc = D3D11_TEXTURE2D_DESC {
Width: w,
Height: h,
MipLevels: 1,
ArraySize: 1,
Format: DXGI_FORMAT_P010,
SampleDesc: DXGI_SAMPLE_DESC {
Count: 1,
Quality: 0,
},
Usage: D3D11_USAGE_DEFAULT,
BindFlags: D3D11_BIND_RENDER_TARGET.0 as u32,
..Default::default()
};
let mut p010: Option<ID3D11Texture2D> = None;
device
.CreateTexture2D(&p010_desc, None, Some(&mut p010))
.context("CreateTexture2D(P010 bars dst)")?;
let p010 = p010.context("null p010 tex")?;
let conv = HdrP010Converter::new(&device)?;
conv.convert(&device, &context, &src_srv, &p010, w, h)?;
Ok((device, p010))
}
}
#[cfg(target_os = "windows")]
pub fn hdr_p010_selftest_at(w: u32, h: u32, vendor: Option<u32>) -> Result<()> {
use windows::Win32::Graphics::Direct3D::{D3D_DRIVER_TYPE_HARDWARE, D3D_DRIVER_TYPE_UNKNOWN};
use windows::Win32::Graphics::Dxgi::{CreateDXGIFactory1, IDXGIAdapter, IDXGIFactory1};
if w == 0 || h == 0 || w % 2 != 0 || h % 2 != 0 {
bail!("hdr-p010-selftest needs even non-zero dimensions, got {w}x{h}");
}
// A grid of 16x16 flat scRGB blocks (each 2x2 chroma footprint uniform → exact chroma
// comparison) covering pure R/G/B/white/black/gray at plausible HDR nit levels, plus a couple
// of bright (>1.0 scRGB) colours, then the rest is a gradient (compared on Y only).
#[allow(non_snake_case)]
let (W, H) = (w, h);
const BLK: u32 = 16; const BLK: u32 = 16;
// (name, r, g, b) scRGB linear (1.0 = 80 nits). Mix of SDR-ish and HDR (>1.0) values. // (name, r, g, b) scRGB linear (1.0 = 80 nits). Mix of SDR-ish and HDR (>1.0) values.
let named: [(&str, f32, f32, f32); 8] = [ let named: [(&str, f32, f32, f32); 8] = [
@@ -797,12 +941,36 @@ pub fn hdr_p010_selftest() -> Result<()> {
// `fp16` outlives the synchronous `CreateTexture2D` that reads it. The mapped-pointer reads are // `fp16` outlives the synchronous `CreateTexture2D` that reads it. The mapped-pointer reads are
// proven individually at the `read_u16` closure below. // proven individually at the `read_u16` closure below.
unsafe { unsafe {
// Hardware D3D11 device (no adapter pin — the default GPU is fine for the self-test). // Device on the requested vendor's adapter (dual-GPU boxes encode on a specific one), else
// the default hardware GPU. Always says which adapter ran — a PASS is only meaningful for
// the GPU it actually tested.
let adapter: Option<IDXGIAdapter> = match vendor {
None => None,
Some(want) => {
let factory: IDXGIFactory1 = CreateDXGIFactory1().context("dxgi factory")?;
let mut found = None;
for i in 0.. {
let Ok(a) = factory.EnumAdapters(i) else {
break;
};
let desc = a.GetDesc().context("adapter desc")?;
if desc.VendorId == want {
found = Some(a);
break;
}
}
Some(found.with_context(|| format!("no adapter with vendor id {want:#x}"))?)
}
};
let mut device: Option<ID3D11Device> = None; let mut device: Option<ID3D11Device> = None;
let mut context: Option<ID3D11DeviceContext> = None; let mut context: Option<ID3D11DeviceContext> = None;
D3D11CreateDevice( D3D11CreateDevice(
None::<&IDXGIAdapter>, adapter.as_ref(),
D3D_DRIVER_TYPE_HARDWARE, if adapter.is_some() {
D3D_DRIVER_TYPE_UNKNOWN
} else {
D3D_DRIVER_TYPE_HARDWARE
},
HMODULE::default(), HMODULE::default(),
D3D11_CREATE_DEVICE_BGRA_SUPPORT, D3D11_CREATE_DEVICE_BGRA_SUPPORT,
Some(&[D3D_FEATURE_LEVEL_11_0]), Some(&[D3D_FEATURE_LEVEL_11_0]),
@@ -814,6 +982,22 @@ pub fn hdr_p010_selftest() -> Result<()> {
.context("D3D11CreateDevice(hardware) for hdr-p010-selftest")?; .context("D3D11CreateDevice(hardware) for hdr-p010-selftest")?;
let device = device.context("null device")?; let device = device.context("null device")?;
let context = context.context("null context")?; let context = context.context("null context")?;
{
let dxgi: windows::Win32::Graphics::Dxgi::IDXGIDevice =
device.cast().context("device -> IDXGIDevice")?;
let desc = dxgi.GetAdapter().context("GetAdapter")?.GetDesc()?;
let name = String::from_utf16_lossy(
&desc.Description[..desc
.Description
.iter()
.position(|&c| c == 0)
.unwrap_or(desc.Description.len())],
);
println!(
"adapter: {name} (vendor {:#06x}, luid {:08x}:{:08x})",
desc.VendorId, desc.AdapterLuid.HighPart, desc.AdapterLuid.LowPart
);
}
// Source FP16 texture (initialized) + SRV. // Source FP16 texture (initialized) + SRV.
let src_desc = D3D11_TEXTURE2D_DESC { let src_desc = D3D11_TEXTURE2D_DESC {
@@ -1175,3 +1359,16 @@ impl VideoConverter {
blt.context("VideoProcessorBlt") blt.context("VideoProcessorBlt")
} }
} }
#[cfg(test)]
mod hdr_selftests {
/// LIVE (needs the GPU): [`super::hdr_p010_selftest_at`] at the field capture size — 1080 is
/// NOT 16-aligned, and the planar-RTV write path is driver-specific per vendor. Pinned to the
/// Intel adapter (`0x8086`), so it runs on the Intel validation boxes and errors out cleanly
/// ("no adapter") elsewhere. `cargo test -p pf-capture -- --ignored hdr_p010 --nocapture`.
#[test]
#[ignore]
fn hdr_p010_selftest_intel_1080_live() {
super::hdr_p010_selftest_at(1920, 1080, Some(0x8086)).expect("hdr p010 selftest @1080");
}
}
@@ -1341,6 +1341,12 @@ impl IddPushCapturer {
self.out_ring.clear(); // the output format changed → rebuild lazily at the new format self.out_ring.clear(); // the output format changed → rebuild lazily at the new format
self.video_conv = None; // converters are sized + HDR-specific → rebuild at the new mode self.video_conv = None; // converters are sized + HDR-specific → rebuild at the new mode
self.hdr_p010_conv = None; self.hdr_p010_conv = None;
// The PyroWave CSC is mode-baked too (BgraToYuvPlanes picks different SDR vs HDR shaders
// and R8/R8G8 vs R16/R16G16 outputs). Without this, a display_hdr flip (Downgrade point D:
// client_10bit=true but HDR couldn't enable at open) reused the stale SDR converter against
// the freshly HDR-formatted pyro ring — every frame corrupted. `ensure_pyro_conv` only
// builds when None, so it must be reset here like its siblings.
self.pyro_conv = None;
self.pyro_ring.clear(); // PyroWave two-plane ring is sized → rebuild at the new mode self.pyro_ring.clear(); // PyroWave two-plane ring is sized → rebuild at the new mode
self.pyro_last = None; self.pyro_last = None;
self.out_idx = 0; self.out_idx = 0;
@@ -1861,6 +1867,7 @@ impl IddPushCapturer {
cbcr, cbcr,
fence_handle, fence_handle,
fence_value, fence_value,
ring_gen: self.generation,
}), }),
) )
} else { } else {
@@ -1919,6 +1926,7 @@ impl IddPushCapturer {
cbcr: dst_cbcr, cbcr: dst_cbcr,
fence_handle, fence_handle,
fence_value, fence_value,
ring_gen: self.generation,
}), }),
}), }),
cursor: None, cursor: None,
+10 -2
View File
@@ -447,8 +447,16 @@ fn pump(
// every ~816 ms at 60120 Hz anyway, so this rarely times out mid-stream). // every ~816 ms at 60120 Hz anyway, so this rarely times out mid-stream).
match connector.next_frame(Duration::from_millis(20)) { match connector.next_frame(Duration::from_millis(20)) {
Ok(frame) => { Ok(frame) => {
// The `received` point: AU fully reassembled, in hand, before decode. // The `received` point: reassembly COMPLETION, stamped by the core session as
let received_ns = now_ns(); // the AU crossed poll_frame (ABI v9). Stamping here at the hand-off pull instead
// would fold the pre-decode queue wait into `host+network` — a client-side
// standing backlog masquerading as network latency (the 2026-07 two-pair
// investigation). 0 = a core predating the stamp; fall back to the pull instant.
let received_ns = if frame.received_ns > 0 {
frame.received_ns
} else {
now_ns()
};
// fps / goodput count every received AU (spec), decoded or not. // fps / goodput count every received AU (spec), decoded or not.
frames_n += 1; frames_n += 1;
bytes_n += frame.data.len() as u64; bytes_n += frame.data.len() as u64;
+5
View File
@@ -25,6 +25,11 @@ tracing = "0.1"
# A test writer for the NVENC backend's unit tests (`with_test_writer().try_init()`). # A test writer for the NVENC backend's unit tests (`with_test_writer().try_init()`).
tracing-subscriber = { version = "0.3", features = ["env-filter"] } tracing-subscriber = { version = "0.3", features = ["env-filter"] }
[target.'cfg(target_os = "windows")'.dev-dependencies]
# The QSV live e2e drives the REAL HdrP010Converter output (an RTV-written, ring-profile P010
# texture) into the encoder — the one seam the CPU-upload tests can't reach.
pf-capture = { path = "../pf-capture" }
[target.'cfg(any(target_os = "linux", target_os = "windows"))'.dependencies] [target.'cfg(any(target_os = "linux", target_os = "windows"))'.dependencies]
# Software H.264 (openh264, BSD-2) — the GPU-less encode path on both platforms. # Software H.264 (openh264, BSD-2) — the GPU-less encode path on both platforms.
openh264 = "0.9" openh264 = "0.9"
+101
View File
@@ -32,6 +32,46 @@ pub struct EncodedFrame {
pub chunk_aligned: bool, pub chunk_aligned: bool,
} }
/// One slice-boundary chunk of an encoded AU, emitted by a chunked-poll backend
/// ([`Encoder::poll_chunk`], latency plan §7 LN1): the encoder hands out completed slices while
/// the rest of the frame is still encoding, so packetize/FEC/pacing can overlap the encode tail.
/// The chunks of one AU concatenate to exactly the bytes [`Encoder::poll`] would have returned,
/// and every cut lands on an Annex-B NAL boundary (slice starts). AU-level metadata
/// (`pts_ns`/`keyframe`/`recovery_anchor`/`chunk_aligned`) is authoritative on the FIRST chunk
/// (`first`) — the host opens the wire frame from it; `last` closes the AU. `keyframe` on a
/// non-final chunk is the encoder's own prediction (exact under the P-only/infinite-GOP config —
/// the driver only ever emits an IDR we asked for); the final chunk re-checks it against the
/// driver's reported picture type.
pub struct AuChunk {
pub data: Vec<u8>,
pub pts_ns: u64,
pub keyframe: bool,
/// See [`EncodedFrame::recovery_anchor`].
pub recovery_anchor: bool,
/// See [`EncodedFrame::chunk_aligned`].
pub chunk_aligned: bool,
/// Opens the AU (carries the authoritative AU metadata).
pub first: bool,
/// Closes the AU (the concatenation is complete; the encoder's in-flight slot is released).
pub last: bool,
}
impl AuChunk {
/// A whole AU as a single self-closing chunk — what every non-chunked backend's
/// [`Encoder::poll_chunk`] default emits, so a chunk consumer needs no per-backend fork.
pub fn whole(f: EncodedFrame) -> Self {
AuChunk {
data: f.data,
pts_ns: f.pts_ns,
keyframe: f.keyframe,
recovery_anchor: f.recovery_anchor,
chunk_aligned: f.chunk_aligned,
first: true,
last: true,
}
}
}
/// Codec selection negotiated with the client. /// Codec selection negotiated with the client.
#[derive(Clone, Copy, Debug, PartialEq, Eq)] #[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Codec { pub enum Codec {
@@ -219,6 +259,11 @@ pub struct EncoderCaps {
/// A hardware encoder. One per session; runs on the encode thread. /// A hardware encoder. One per session; runs on the encode thread.
pub trait Encoder: Send { pub trait Encoder: Send {
/// Submit one captured frame for encoding. Lifetime contract: the caller must keep `frame`
/// (and its GPU payload) alive until this frame's AU has been returned by
/// [`poll`](Self::poll) — a stream-ordered backend (Linux direct-NVENC's IO-stream binding)
/// may still be reading the payload asynchronously after `submit` returns. Both host encode
/// loops already hold the frame across their poll drain; new callers must do the same.
fn submit(&mut self, frame: &CapturedFrame) -> Result<()>; fn submit(&mut self, frame: &CapturedFrame) -> Result<()>;
/// [`submit`](Self::submit) with the **wire frame index** this frame's AU will carry — the /// [`submit`](Self::submit) with the **wire frame index** this frame's AU will carry — the
/// number the packetizer stamps on it and the client's loss reports/RFI requests name. The /// number the packetizer stamps on it and the client's loss reports/RFI requests name. The
@@ -264,8 +309,37 @@ pub trait Encoder: Send {
fn invalidate_ref_frames(&mut self, _first_frame: i64, _last_frame: i64) -> bool { fn invalidate_ref_frames(&mut self, _first_frame: i64, _last_frame: i64) -> bool {
false false
} }
/// Escalate into a pipelined (two-thread) retrieve mode under sustained GPU contention — the
/// encoder analog of the capturer depth escalation: AUs ride ~one loop tick behind (`poll`
/// may return `None` while an encode is in flight) in exchange for capture/submit no longer
/// serializing on the encode wait. Returns whether pipelined retrieve is (now) active; the
/// switch may be deferred to the next safe point internally. `false` from the default impl =
/// unsupported — the session loop stops asking. De-escalation is a v2 item everywhere.
fn set_pipelined(&mut self, _on: bool) -> bool {
false
}
/// Pull the next encoded AU if one is ready. /// Pull the next encoded AU if one is ready.
fn poll(&mut self) -> Result<Option<EncodedFrame>>; fn poll(&mut self) -> Result<Option<EncodedFrame>>;
/// Whether [`poll_chunk`](Self::poll_chunk) currently emits sub-AU chunks — i.e. the LIVE
/// session has slice-level readback armed (Linux direct-NVENC with the
/// `PUNKTFUNK_NVENC_SLICES` and `PUNKTFUNK_NVENC_SUBFRAME` knobs on a sync depth-1
/// retrieve). Dynamic, not static: a pipelined-retrieve escalation or a session rebuild can
/// turn it off — re-query per AU, never cache across frames. `false` (the default) means
/// `poll_chunk` degrades to one whole-AU chunk per frame.
fn supports_chunked_poll(&self) -> bool {
false
}
/// Pull the next slice-boundary chunk of the oldest in-flight AU (latency plan §7 LN1).
/// Semantics when chunking is live: BLOCKS until the next chunk is readable, and the final
/// (`last`) chunk blocks exactly like [`poll`](Self::poll) does — the depth-1 pump treats
/// `None` as re-poll-next-tick, so a non-blocking tail would ride the AU one tick late (the
/// `6dc195f9` Vulkan bug class). `Ok(None)` only when no AU is in flight. Each AU must be
/// drained through ONE method: calling `poll` on a partially-chunked AU is a caller bug (the
/// backend errors rather than double-emit bytes). Default: delegates to `poll`, wrapping the
/// whole AU as a single `first && last` chunk.
fn poll_chunk(&mut self) -> Result<Option<AuChunk>> {
Ok(self.poll()?.map(AuChunk::whole))
}
/// Tear the underlying hardware encoder down and rebuild it in place, keeping the session's /// Tear the underlying hardware encoder down and rebuild it in place, keeping the session's
/// negotiated parameters — the encode-stall watchdog's recovery lever (a wedged AMF/QSV /// negotiated parameters — the encode-stall watchdog's recovery lever (a wedged AMF/QSV
/// driver stops emitting AUs or accepting frames without ever returning an error). Returns /// driver stops emitting AUs or accepting frames without ever returning an error). Returns
@@ -293,6 +367,16 @@ pub trait Encoder: Send {
/// flagged [`EncodedFrame::chunk_aligned`] and the session marks them on the wire. /// flagged [`EncodedFrame::chunk_aligned`] and the session marks them on the wire.
/// Default: no-op (the H.26x backends' bitstreams cannot be cut losslessly). /// Default: no-op (the H.26x backends' bitstreams cannot be cut losslessly).
fn set_wire_chunking(&mut self, _shard_payload: usize) {} fn set_wire_chunking(&mut self, _shard_payload: usize) {}
/// How many frames the CAPTURER guarantees the encoder may hold in flight before it starts
/// reusing an input texture (`Capturer::pipeline_depth`). Backends that encode the capturer's
/// textures IN PLACE — no `CopyResource` — must not pipeline deeper than this: the capturer
/// rotates its output ring per delivered frame with no regard for encode completion, so a
/// deeper pipeline lets it overwrite a texture mid-encode. That is visual corruption (torn or
/// mixed frames), not UB, so it fails silently and intermittently.
///
/// Called once by the session glue after the capturer is known; a backend that copies its
/// input, or is synchronous, ignores it. Default: no-op.
fn set_input_ring_depth(&mut self, _depth: usize) {}
/// Signal end-of-stream. After this, drain the remaining AUs with [`poll`](Self::poll) /// Signal end-of-stream. After this, drain the remaining AUs with [`poll`](Self::poll)
/// until it returns `None` — NVENC buffers frames internally even at `delay=0`. /// until it returns `None` — NVENC buffers frames internally even at `delay=0`.
fn flush(&mut self) -> Result<()>; fn flush(&mut self) -> Result<()>;
@@ -412,6 +496,23 @@ mod tests {
} }
} }
/// The whole-AU chunk (every non-chunked backend's `poll_chunk` shape) must carry the AU's
/// metadata verbatim and be self-closing (`first && last`).
#[test]
fn whole_au_chunk_is_self_closing() {
let c = AuChunk::whole(EncodedFrame {
data: vec![0, 0, 0, 1, 0x40],
pts_ns: 42,
keyframe: true,
recovery_anchor: true,
chunk_aligned: false,
});
assert_eq!(c.data, vec![0, 0, 0, 1, 0x40]);
assert_eq!(c.pts_ns, 42);
assert!(c.keyframe && c.recovery_anchor && !c.chunk_aligned);
assert!(c.first && c.last);
}
/// Wire round-trip and the stats label stay in lockstep with the `quic::CODEC_*` bits. /// Wire round-trip and the stats label stay in lockstep with the `quic::CODEC_*` bits.
#[test] #[test]
fn codec_wire_roundtrip_and_label() { fn codec_wire_roundtrip_and_label() {
+3 -3
View File
@@ -826,7 +826,7 @@ impl NvencEncoder {
(*f).linesize[i] as usize, (*f).linesize[i] as usize,
) )
}); });
pf_zerocopy::cuda::copy_yuv444_to_device(buf, dsts) pf_zerocopy::cuda::copy_yuv444_to_device(buf, dsts, true)
} else if self.want_444 { } else if self.want_444 {
ffi::av_frame_free(&mut f); ffi::av_frame_free(&mut f);
bail!( bail!(
@@ -839,11 +839,11 @@ impl NvencEncoder {
let y_pitch = (*f).linesize[0] as usize; let y_pitch = (*f).linesize[0] as usize;
let uv_ptr = (*f).data[1] as pf_zerocopy::cuda::CUdeviceptr; let uv_ptr = (*f).data[1] as pf_zerocopy::cuda::CUdeviceptr;
let uv_pitch = (*f).linesize[1] as usize; let uv_pitch = (*f).linesize[1] as usize;
pf_zerocopy::cuda::copy_nv12_to_device(buf, y_ptr, y_pitch, uv_ptr, uv_pitch) pf_zerocopy::cuda::copy_nv12_to_device(buf, y_ptr, y_pitch, uv_ptr, uv_pitch, true)
} else { } else {
let dst_ptr = (*f).data[0] as pf_zerocopy::cuda::CUdeviceptr; let dst_ptr = (*f).data[0] as pf_zerocopy::cuda::CUdeviceptr;
let dst_pitch = (*f).linesize[0] as usize; let dst_pitch = (*f).linesize[0] as usize;
pf_zerocopy::cuda::copy_device_to_device(buf, dst_ptr, dst_pitch) pf_zerocopy::cuda::copy_device_to_device(buf, dst_ptr, dst_pitch, true)
}; };
if let Err(e) = copy_res { if let Err(e) = copy_res {
ffi::av_frame_free(&mut f); ffi::av_frame_free(&mut f);
File diff suppressed because it is too large Load Diff
+18 -1
View File
@@ -1176,7 +1176,24 @@ impl Encoder for PyroWaveEncoder {
fn submit(&mut self, frame: &CapturedFrame) -> Result<()> { fn submit(&mut self, frame: &CapturedFrame) -> Result<()> {
// SAFETY: single-threaded encoder; `encode_frame` records/submits on handles this // SAFETY: single-threaded encoder; `encode_frame` records/submits on handles this
// struct owns and waits its own fence before touching results. // struct owns and waits its own fence before touching results.
unsafe { self.encode_frame(frame) } let r = unsafe { self.encode_frame(frame) };
if r.is_err() {
// `encode_frame` opens the recording window early and has several fallible steps
// inside it (cursor prep, dmabuf import, format mapping, the CPU-RGB staging path,
// an unsupported-payload bail, and the encode call itself). Every one returns with
// `self.cmd` still RECORDING, and nothing downstream repairs it — there is exactly
// one `begin_command_buffer` in this file and `reset()`/`Drop` never touch `cmd` —
// so the NEXT frame would call `begin` on a recording buffer, which is invalid usage.
// Legal here on every path: the pool carries RESET_COMMAND_BUFFER and the buffer is
// not pending (we never reached the submit, or the submit itself failed).
// SAFETY: `self.cmd` is owned by this encoder and, on these paths, not in flight.
unsafe {
let _ = self
.device
.reset_command_buffer(self.cmd, vk::CommandBufferResetFlags::empty());
}
}
r
} }
fn caps(&self) -> EncoderCaps { fn caps(&self) -> EncoderCaps {
+76 -14
View File
@@ -37,9 +37,11 @@ pub(crate) fn fourcc_to_vk(fourcc: u32) -> Option<vk::Format> {
const AR24: u32 = 0x3432_5241; // ARGB8888 const AR24: u32 = 0x3432_5241; // ARGB8888
const XB24: u32 = 0x3432_4258; // XBGR8888 const XB24: u32 = 0x3432_4258; // XBGR8888
const AB24: u32 = 0x3432_4241; // ABGR8888 const AB24: u32 = 0x3432_4241; // ABGR8888
const NV12: u32 = 0x3231_564e; // DRM_FORMAT_NV12
match fourcc { match fourcc {
XR24 | AR24 => Some(vk::Format::B8G8R8A8_UNORM), XR24 | AR24 => Some(vk::Format::B8G8R8A8_UNORM),
XB24 | AB24 => Some(vk::Format::R8G8B8A8_UNORM), XB24 | AB24 => Some(vk::Format::R8G8B8A8_UNORM),
NV12 => Some(vk::Format::G8_B8R8_2PLANE_420_UNORM),
_ => None, _ => None,
} }
} }
@@ -77,21 +79,62 @@ pub(crate) unsafe fn import_rgb_dmabuf(
d: &pf_frame::DmabufFrame, d: &pf_frame::DmabufFrame,
cw: u32, cw: u32,
ch: u32, ch: u32,
) -> Result<(vk::Image, vk::DeviceMemory, vk::ImageView)> {
import_rgb_dmabuf_as(
device,
ext_fd,
mem_props,
d,
cw,
ch,
vk::ImageUsageFlags::SAMPLED,
None,
)
}
/// [`import_rgb_dmabuf`] with the image usage explicit and an optional video-profile list.
/// Despite the historical name, this also imports gamescope's one-fd LINEAR NV12: the UV
/// subresource layout comes from the producer's plane-1 chunk when it reported one, falling
/// back to the shared-stride contiguous-plane contract.
#[allow(clippy::too_many_arguments)]
pub(crate) unsafe fn import_rgb_dmabuf_as(
device: &ash::Device,
ext_fd: &ash::khr::external_memory_fd::Device,
mem_props: &vk::PhysicalDeviceMemoryProperties,
d: &pf_frame::DmabufFrame,
cw: u32,
ch: u32,
usage: vk::ImageUsageFlags,
profile_list: Option<&mut vk::VideoProfileListInfoKHR>,
) -> Result<(vk::Image, vk::DeviceMemory, vk::ImageView)> { ) -> Result<(vk::Image, vk::DeviceMemory, vk::ImageView)> {
use anyhow::Context; use anyhow::Context;
use std::os::fd::IntoRawFd; use std::os::fd::IntoRawFd;
let fmt = fourcc_to_vk(d.fourcc) let fmt = fourcc_to_vk(d.fourcc)
.with_context(|| format!("unsupported dmabuf fourcc {:#x}", d.fourcc))?; .with_context(|| format!("unsupported dmabuf fourcc {:#x}", d.fourcc))?;
let plane = [vk::SubresourceLayout::default() let planes: Vec<vk::SubresourceLayout> = if fmt == vk::Format::G8_B8R8_2PLANE_420_UNORM {
let (uv_offset, uv_stride) = d.plane1.map(|(o, s)| (o as u64, s as u64)).unwrap_or((
d.offset as u64 + d.stride as u64 * ch as u64,
d.stride as u64,
));
vec![
vk::SubresourceLayout::default()
.offset(d.offset as u64) .offset(d.offset as u64)
.row_pitch(d.stride as u64)]; .row_pitch(d.stride as u64),
vk::SubresourceLayout::default()
.offset(uv_offset)
.row_pitch(uv_stride),
]
} else {
vec![vk::SubresourceLayout::default()
.offset(d.offset as u64)
.row_pitch(d.stride as u64)]
};
let mut drm = vk::ImageDrmFormatModifierExplicitCreateInfoEXT::default() let mut drm = vk::ImageDrmFormatModifierExplicitCreateInfoEXT::default()
.drm_format_modifier(d.modifier) .drm_format_modifier(d.modifier)
.plane_layouts(&plane); .plane_layouts(&planes);
let mut ext = vk::ExternalMemoryImageCreateInfo::default() let mut ext = vk::ExternalMemoryImageCreateInfo::default()
.handle_types(vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT); .handle_types(vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT);
let img = device.create_image( let mut ci = vk::ImageCreateInfo::default()
&vk::ImageCreateInfo::default()
.image_type(vk::ImageType::TYPE_2D) .image_type(vk::ImageType::TYPE_2D)
.format(fmt) .format(fmt)
.extent(vk::Extent3D { .extent(vk::Extent3D {
@@ -103,13 +146,15 @@ pub(crate) unsafe fn import_rgb_dmabuf(
.array_layers(1) .array_layers(1)
.samples(vk::SampleCountFlags::TYPE_1) .samples(vk::SampleCountFlags::TYPE_1)
.tiling(vk::ImageTiling::DRM_FORMAT_MODIFIER_EXT) .tiling(vk::ImageTiling::DRM_FORMAT_MODIFIER_EXT)
.usage(vk::ImageUsageFlags::SAMPLED) .usage(usage)
.sharing_mode(vk::SharingMode::EXCLUSIVE) .sharing_mode(vk::SharingMode::EXCLUSIVE)
.initial_layout(vk::ImageLayout::UNDEFINED) .initial_layout(vk::ImageLayout::UNDEFINED)
.push_next(&mut ext) .push_next(&mut ext)
.push_next(&mut drm), .push_next(&mut drm);
None, if let Some(pl) = profile_list {
)?; ci = ci.push_next(pl);
}
let img = device.create_image(&ci, None)?;
// dup the fd; Vulkan takes ownership of the dup on a successful import. // dup the fd; Vulkan takes ownership of the dup on a successful import.
let dup = d.fd.try_clone().context("dup dmabuf fd")?.into_raw_fd(); let dup = d.fd.try_clone().context("dup dmabuf fd")?.into_raw_fd();
let fd_props = { let fd_props = {
@@ -183,7 +228,8 @@ pub(crate) unsafe fn make_plain_image(
None, None,
)?; )?;
let req = device.get_image_memory_requirements(img); let req = device.get_image_memory_requirements(img);
let mem = device.allocate_memory( // Unwind on failure: callers (the encoders' open paths) only ever see the completed triple.
let mem = match device.allocate_memory(
&vk::MemoryAllocateInfo::default() &vk::MemoryAllocateInfo::default()
.allocation_size(req.size) .allocation_size(req.size)
.memory_type_index(find_mem( .memory_type_index(find_mem(
@@ -192,8 +238,24 @@ pub(crate) unsafe fn make_plain_image(
vk::MemoryPropertyFlags::DEVICE_LOCAL, vk::MemoryPropertyFlags::DEVICE_LOCAL,
)), )),
None, None,
)?; ) {
device.bind_image_memory(img, mem, 0)?; Ok(m) => m,
let view = make_view(device, img, fmt, 0)?; Err(e) => {
Ok((img, mem, view)) device.destroy_image(img, None);
return Err(e.into());
}
};
if let Err(e) = device.bind_image_memory(img, mem, 0) {
device.destroy_image(img, None);
device.free_memory(mem, None);
return Err(e.into());
}
match make_view(device, img, fmt, 0) {
Ok(view) => Ok((img, mem, view)),
Err(e) => {
device.destroy_image(img, None);
device.free_memory(mem, None);
Err(e)
}
}
} }
@@ -0,0 +1,82 @@
//! Vendored `VK_VALVE_video_encode_rgb_conversion` bindings — the RGB→YCbCr encode-source
//! extension (Vulkan 1.4.327; RADV since Mesa 26.0, hardware-gated on the VCN EFC front-end
//! conversion block). Our pinned `ash 0.38.0+1.3.281` predates it entirely; same vendoring
//! rationale as [`vk_av1_encode`](super::vk_av1_encode) — definitions copied from the registry
//! so the layouts are correct-by-construction, chained via raw `p_next`. Consumed by
//! `vulkan_video.rs`: B0 probes + logs availability (design/vulkan-rgb-direct-encode.md);
//! B1 makes the captured BGRx dmabuf the direct encode source with EFC doing the 709-narrow CSC.
#![allow(dead_code)]
use ash::vk;
use std::ffi::{c_void, CStr};
pub const EXTENSION_NAME: &CStr = c"VK_VALVE_video_encode_rgb_conversion";
// ---------- struct-type (VkStructureType) values — construct via `stype` ----------
pub const ST_PHYSICAL_DEVICE_FEATURES: i32 = 1_000_390_000;
pub const ST_CAPABILITIES: i32 = 1_000_390_001;
pub const ST_PROFILE_INFO: i32 = 1_000_390_002;
pub const ST_SESSION_CREATE_INFO: i32 = 1_000_390_003;
// `VkVideoEncodeRgbModelConversionFlagBitsVALVE`
pub const MODEL_RGB_IDENTITY: u32 = 0x01;
pub const MODEL_YCBCR_IDENTITY: u32 = 0x02;
pub const MODEL_YCBCR_709: u32 = 0x04;
pub const MODEL_YCBCR_601: u32 = 0x08;
pub const MODEL_YCBCR_2020: u32 = 0x10;
// `VkVideoEncodeRgbRangeCompressionFlagBitsVALVE`
pub const RANGE_FULL: u32 = 0x01;
pub const RANGE_NARROW: u32 = 0x02;
// `VkVideoEncodeRgbChromaOffsetFlagBitsVALVE`
pub const CHROMA_OFFSET_COSITED_EVEN: u32 = 0x01;
pub const CHROMA_OFFSET_MIDPOINT: u32 = 0x02;
/// `VkPhysicalDeviceVideoEncodeRgbConversionFeaturesVALVE` — chain into
/// `VkPhysicalDeviceFeatures2` (query) / `VkDeviceCreateInfo` (enable).
#[repr(C)]
pub struct PhysicalDeviceVideoEncodeRgbConversionFeaturesVALVE {
pub s_type: vk::StructureType,
pub p_next: *mut c_void,
pub video_encode_rgb_conversion: vk::Bool32,
}
/// `VkVideoEncodeRgbConversionCapabilitiesVALVE` — chain into the
/// `vkGetPhysicalDeviceVideoCapabilitiesKHR` output when the queried profile carries
/// [`VideoEncodeProfileRgbConversionInfoVALVE`]; reports which conversions the HW does.
#[repr(C)]
pub struct VideoEncodeRgbConversionCapabilitiesVALVE {
pub s_type: vk::StructureType,
pub p_next: *mut c_void,
pub rgb_models: u32,
pub rgb_ranges: u32,
pub x_chroma_offsets: u32,
pub y_chroma_offsets: u32,
}
/// `VkVideoEncodeProfileRgbConversionInfoVALVE` — part of the video-profile *identity*: every
/// consumer of the profile (caps query, format query, session, image profile lists) must carry
/// the same chain.
#[repr(C)]
pub struct VideoEncodeProfileRgbConversionInfoVALVE {
pub s_type: vk::StructureType,
pub p_next: *const c_void,
pub perform_encode_rgb_conversion: vk::Bool32,
}
/// `VkVideoEncodeSessionRgbConversionCreateInfoVALVE` — chain into
/// `VkVideoSessionCreateInfoKHR`; single-bit selections of the conversion actually performed.
#[repr(C)]
pub struct VideoEncodeSessionRgbConversionCreateInfoVALVE {
pub s_type: vk::StructureType,
pub p_next: *const c_void,
pub rgb_model: u32,
pub rgb_range: u32,
pub x_chroma_offset: u32,
pub y_chroma_offset: u32,
}
/// `vk::StructureType` for a raw `ST_*` constant above.
#[inline]
pub fn stype(raw: i32) -> vk::StructureType {
vk::StructureType::from_raw(raw)
}
File diff suppressed because it is too large Load Diff
+69
View File
@@ -35,6 +35,38 @@ pub(super) fn codec_guid(codec: Codec) -> nv::GUID {
} }
} }
/// Resolved per-frame slice count for a session (latency plan §7 LN1, Phase 3): the
/// `PUNKTFUNK_NVENC_SLICES` env override wins (1..=32; **1 = the explicit single-slice
/// escape**, needed now that a backend can default higher), else the backend's
/// `default_slices` — 4 on Linux direct-NVENC since the Phase-3 default-on, 1 everywhere else
/// (the Windows async path is deliberately untouched). H.264/HEVC only (AV1 partitions via
/// tiles). ONE parse shared by the config author ([`apply_low_latency_config`] via
/// [`LowLatencyConfig::slices`]) and the Linux backend's chunked-poll arming, so the two can
/// never disagree about whether a session is multi-slice.
pub(super) fn resolve_slices(codec: Codec, default_slices: u32) -> u32 {
if !matches!(codec, Codec::H264 | Codec::H265) {
return 1;
}
std::env::var("PUNKTFUNK_NVENC_SLICES")
.ok()
.and_then(|s| s.parse::<u32>().ok())
.filter(|n| (1..=32).contains(n))
.unwrap_or(default_slices)
}
/// Resolved sub-frame readback (`enableSubFrameWrite` + `reportSliceOffsets`; sync sessions
/// only, see [`build_init_params`]): `PUNKTFUNK_NVENC_SUBFRAME` tri-state — `0` = never (the
/// default-on escape), `1` = force (even where the caps probe says unsupported — an operator
/// explicitly testing), unset = the backend's `default_on` (Linux direct-NVENC passes its
/// SUBFRAME_READBACK caps-probe result since Phase 3; Windows passes `false`).
pub(super) fn resolve_subframe(default_on: bool) -> bool {
match std::env::var("PUNKTFUNK_NVENC_SUBFRAME").as_deref() {
Ok("0") => false,
Ok("1") => true,
_ => default_on,
}
}
/// Reference-frame DPB depth when RFI is supported (Apollo uses 5). A deeper DPB lets an invalidated /// Reference-frame DPB depth when RFI is supported (Apollo uses 5). A deeper DPB lets an invalidated
/// reference fall back to an older still-valid frame instead of a full IDR; `numRefL0 = 1` keeps each /// reference fall back to an older still-valid frame instead of a full IDR; `numRefL0 = 1` keeps each
/// P-frame single-reference for low latency. Also the window the backends' `invalidate_ref_frames` /// P-frame single-reference for low latency. Also the window the backends' `invalidate_ref_frames`
@@ -66,6 +98,9 @@ pub(super) struct LowLatencyConfig {
pub hdr: bool, pub hdr: bool,
/// This GPU supports reference-frame invalidation (a deeper DPB for graceful loss recovery). /// This GPU supports reference-frame invalidation (a deeper DPB for graceful loss recovery).
pub rfi_supported: bool, pub rfi_supported: bool,
/// Resolved per-frame slice count ([`resolve_slices`] — env override, else the backend
/// default). ≤ 1 leaves the preset's single slice untouched.
pub slices: u32,
} }
/// Author the shared `NV_ENC_INITIALIZE_PARAMS` (P1/ULL preset, PTD, the session dimensions/rate) /// Author the shared `NV_ENC_INITIALIZE_PARAMS` (P1/ULL preset, PTD, the session dimensions/rate)
@@ -82,6 +117,7 @@ pub(super) fn build_init_params(
cfg: &mut nv::NV_ENC_CONFIG, cfg: &mut nv::NV_ENC_CONFIG,
split_mode: u32, split_mode: u32,
enable_async: bool, enable_async: bool,
subframe: bool,
) -> nv::NV_ENC_INITIALIZE_PARAMS { ) -> nv::NV_ENC_INITIALIZE_PARAMS {
let mut init = nv::NV_ENC_INITIALIZE_PARAMS { let mut init = nv::NV_ENC_INITIALIZE_PARAMS {
version: nv::NV_ENC_INITIALIZE_PARAMS_VER, version: nv::NV_ENC_INITIALIZE_PARAMS_VER,
@@ -101,6 +137,16 @@ pub(super) fn build_init_params(
}; };
// splitEncodeMode is a C bitfield — set via the generated accessor, not a struct field. // splitEncodeMode is a C bitfield — set via the generated accessor, not a struct field.
init.set_splitEncodeMode(split_mode); init.set_splitEncodeMode(split_mode);
// Sub-frame readback (latency plan §7 LN1; default-on for Linux direct-NVENC since Phase 3 —
// the caller resolves `subframe` via [`resolve_subframe`] + its caps probe): the driver
// writes each slice into the output buffer as it completes and reports per-slice offsets, so
// a sync-mode consumer can read slices out while the frame is still encoding. Pair with
// multi-slice (a single-slice frame yields nothing to read early). `reportSliceOffsets`
// requires `enableEncodeAsync = 0`, so async (Windows) sessions never arm.
if !enable_async && subframe {
init.set_enableSubFrameWrite(1);
init.set_reportSliceOffsets(1);
}
init init
} }
@@ -118,6 +164,9 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
cfg.gopLength = nv::NVENC_INFINITE_GOPLENGTH; cfg.gopLength = nv::NVENC_INFINITE_GOPLENGTH;
cfg.frameIntervalP = 1; cfg.frameIntervalP = 1;
cfg.rcParams.rateControlMode = nv::NV_ENC_PARAMS_RC_MODE::NV_ENC_PARAMS_RC_CBR; cfg.rcParams.rateControlMode = nv::NV_ENC_PARAMS_RC_MODE::NV_ENC_PARAMS_RC_CBR;
// Explicit zero reorder delay: with P-only + no lookahead there is no reordering to buffer,
// but pin the bit so no preset/driver default can ever slip a frame of reorder delay in.
cfg.rcParams.set_zeroReorderDelay(1);
let bps = c.bitrate.min(u32::MAX as u64) as u32; let bps = c.bitrate.min(u32::MAX as u64) as u32;
cfg.rcParams.averageBitRate = bps; cfg.rcParams.averageBitRate = bps;
cfg.rcParams.maxBitRate = bps; cfg.rcParams.maxBitRate = bps;
@@ -146,6 +195,26 @@ pub(super) unsafe fn apply_low_latency_config(cfg: &mut nv::NV_ENC_CONFIG, c: Lo
Codec::PyroWave => unreachable!("PyroWave never opens the direct-NVENC backend"), Codec::PyroWave => unreachable!("PyroWave never opens the direct-NVENC backend"),
} }
// Multi-slice frames (latency plan §7 LN1): `c.slices` splits every frame into N slices
// (sliceMode 3 = "N slices per frame"), the unit sub-frame readback ships early and loss
// concealment can discard independently. Costs ~1-2 % bitrate in slice headers. H.264/HEVC
// only — AV1 partitions via tiles, not slices (the resolver already returns 1 there).
// Default 4 on Linux direct-NVENC (Phase 3), env-only elsewhere; ≤ 1 keeps the preset's
// single slice.
if let Some(n) = Some(c.slices).filter(|n| *n >= 2) {
match c.codec {
Codec::H264 => {
cfg.encodeCodecConfig.h264Config.sliceMode = 3;
cfg.encodeCodecConfig.h264Config.sliceModeData = n;
}
Codec::H265 => {
cfg.encodeCodecConfig.hevcConfig.sliceMode = 3;
cfg.encodeCodecConfig.hevcConfig.sliceModeData = n;
}
Codec::Av1 | Codec::PyroWave => {}
}
}
// Chroma + bit depth. Full-chroma 4:4:4 (HEVC Range Extensions, chromaFormatIDC=3 under the FREXT // Chroma + bit depth. Full-chroma 4:4:4 (HEVC Range Extensions, chromaFormatIDC=3 under the FREXT
// profile) takes precedence and composes with 10-bit (Main 4:4:4 10); it needs a full-chroma- // profile) takes precedence and composes with 10-bit (Main 4:4:4 10); it needs a full-chroma-
// capable input. Otherwise 10-bit selects Main10 (HEVC) or the AV1 output depth — stamping the // capable input. Otherwise 10-bit selects Main10 (HEVC) or the AV1 output depth — stamping the
+56 -7
View File
@@ -37,7 +37,8 @@
#![deny(clippy::undocumented_unsafe_blocks)] #![deny(clippy::undocumented_unsafe_blocks)]
use super::nvenc_core::{ use super::nvenc_core::{
apply_low_latency_config, build_init_params, codec_guid, LowLatencyConfig, NvStatusExt, RFI_DPB, apply_low_latency_config, build_init_params, codec_guid, resolve_slices, resolve_subframe,
LowLatencyConfig, NvStatusExt, RFI_DPB,
}; };
use super::nvenc_status; use super::nvenc_status;
use super::{ChromaFormat, Codec, EncodedFrame, Encoder, EncoderCaps}; use super::{ChromaFormat, Codec, EncodedFrame, Encoder, EncoderCaps};
@@ -409,6 +410,12 @@ pub struct NvencD3d11Encoder {
events: Vec<usize>, events: Vec<usize>,
/// Async mode: the retrieve thread + its channels (`None` = classic same-thread sync retrieve). /// Async mode: the retrieve thread + its channels (`None` = classic same-thread sync retrieve).
async_rt: Option<AsyncRetrieve>, async_rt: Option<AsyncRetrieve>,
/// The capturer's `pipeline_depth` (`set_input_ring_depth`). This backend encodes the
/// capturer's textures IN PLACE, so it is a HARD ceiling on async in-flight depth: the
/// capturer rotates its ring per delivered frame regardless of encode completion, so
/// pipelining deeper lets it overwrite a texture mid-encode (torn frames). `None` until the
/// session glue reports it — treated as "unknown, don't pipeline past the env cap".
input_ring_depth: Option<usize>,
/// `NV_ENC_CAPS_ASYNC_ENCODE_SUPPORT` from the caps probe — gates the async retrieve mode. /// `NV_ENC_CAPS_ASYNC_ENCODE_SUPPORT` from the caps probe — gates the async retrieve mode.
async_supported: bool, async_supported: bool,
/// (bitstream, mapped input resource to unmap after retrieval, pts_ns, recovery-anchor) per /// (bitstream, mapped input resource to unmap after retrieval, pts_ns, recovery-anchor) per
@@ -505,6 +512,7 @@ impl NvencD3d11Encoder {
bitstreams: Vec::new(), bitstreams: Vec::new(),
events: Vec::new(), events: Vec::new(),
async_rt: None, async_rt: None,
input_ring_depth: None,
async_supported: false, async_supported: false,
pending: VecDeque::new(), pending: VecDeque::new(),
frame_idx: 0, frame_idx: 0,
@@ -737,6 +745,10 @@ impl NvencD3d11Encoder {
av1_input_depth_minus8: if ten_bit_in { 2 } else { 0 }, av1_input_depth_minus8: if ten_bit_in { 2 } else { 0 },
hdr: self.hdr, hdr: self.hdr,
rfi_supported: self.rfi_supported, rfi_supported: self.rfi_supported,
// Env-only on Windows (default single slice): the Phase-3 default-on is a
// Linux-direct-NVENC decision — the Windows async path stays untouched, and a
// Windows operator opting in must choose slices+sync over async retrieve.
slices: resolve_slices(self.codec, 1),
}, },
); );
Ok(cfg) Ok(cfg)
@@ -784,6 +796,9 @@ impl NvencD3d11Encoder {
&mut cfg, &mut cfg,
split_mode, split_mode,
enable_async, enable_async,
// Windows: env opt-in only ("1"), never a default — and build_init_params
// additionally refuses it on an async session.
resolve_subframe(false),
); );
match (api().initialize_encoder)(enc, &mut init).nv_ok() { match (api().initialize_encoder)(enc, &mut init).nv_ok() {
@@ -1135,7 +1150,19 @@ impl Encoder for NvencD3d11Encoder {
self.chroma_444 = false; self.chroma_444 = false;
} }
let device = frame.device.clone(); let device = frame.device.clone();
self.init_session(&device)?; // `init_session` publishes `self.encoder` (and charges LIVE_SESSION_UNITS) BEFORE its
// last fallible steps, so a failure there leaves a live session with `inited == false`.
// Every guard on the re-init path keys off `inited`, so without this the next submit
// would skip teardown and overwrite `self.encoder` — leaking the session permanently
// (toward the driver's per-process cap) along with its session-budget units.
// `teardown` keys off `encoder.is_null()`, not `inited`, so it cleans up exactly this
// half-built state and is a no-op when nothing was opened.
if let Err(e) = self.init_session(&device) {
// SAFETY: same contract as the teardown above — the encode thread owns the session,
// and a failed init leaves nothing mid-encode to race with.
unsafe { self.teardown() };
return Err(e);
}
self.init_device = dev_raw; self.init_device = dev_raw;
} }
// The session's opening frame — NVENC emits it as an IDR regardless of pic flags, so the // The session's opening frame — NVENC emits it as an IDR regardless of pic flags, so the
@@ -1144,11 +1171,21 @@ impl Encoder for NvencD3d11Encoder {
// index, which is non-zero on a mid-session encoder rebuild's first frame. // index, which is non-zero on a mid-session encoder rebuild's first frame.
let opening = self.next == 0; let opening = self.next == 0;
// Async backpressure: never hand NVENC an output bitstream that is still in flight, and // Async backpressure: never hand NVENC an output bitstream that is still in flight, and
// keep in-flight depth within the capturer's texture ring (see `async_inflight_cap`). At // keep in-flight depth within the capturer's texture ring. At the cap, block on the OLDEST
// the cap, block on the OLDEST completion (the retrieve thread is already waiting on its // completion (the retrieve thread is already waiting on its event) before submitting more —
// event) before submitting more — bounding depth exactly like the sync path's per-tick // bounding depth exactly like the sync path's per-tick blocking poll, just `cap` deep
// blocking poll, just `cap` deep instead of 1. // instead of 1.
while self.async_rt.is_some() && self.pending.len() >= async_inflight_cap() { //
// The ring term is the one that matters for correctness: `async_inflight_cap()` is only the
// output-bitstream-pool ceiling plus an env knob, and consults NOTHING about the capturer,
// despite this comment previously claiming otherwise. Since this backend encodes the
// capturer's textures in place, exceeding the capturer's declared `pipeline_depth` lets it
// rotate a texture out from under a live encode — torn frames, silently.
let cap = match self.input_ring_depth {
Some(d) => async_inflight_cap().min(d.max(1)),
None => async_inflight_cap(),
};
while self.async_rt.is_some() && self.pending.len() >= cap {
let done = { let done = {
let rt = self.async_rt.as_mut().expect("checked in loop condition"); let rt = self.async_rt.as_mut().expect("checked in loop condition");
rt.done_rx rt.done_rx
@@ -1324,6 +1361,17 @@ impl Encoder for NvencD3d11Encoder {
self.submit(frame) self.submit(frame)
} }
fn set_input_ring_depth(&mut self, depth: usize) {
// This backend registers and encodes the capturer's textures in place (no CopyResource),
// so the capturer's ring depth is a hard ceiling on how deep async may pipeline.
self.input_ring_depth = Some(depth);
tracing::debug!(
depth,
env_cap = async_inflight_cap(),
"NVENC: capturer input-ring depth reported — async in-flight bounded by the smaller"
);
}
fn request_keyframe(&mut self) { fn request_keyframe(&mut self) {
self.force_kf = true; self.force_kf = true;
} }
@@ -1528,6 +1576,7 @@ impl Encoder for NvencD3d11Encoder {
&mut cfg, &mut cfg,
self.split_mode, self.split_mode,
self.session_async, self.session_async,
resolve_subframe(false),
), ),
..Default::default() ..Default::default()
}; };
+57 -6
View File
@@ -43,6 +43,9 @@ const BS_SLACK: usize = 256 * 1024;
/// (a desktop-switch device recreate), in which case the stale imports are evicted + destroyed. /// (a desktop-switch device recreate), in which case the stale imports are evicted + destroyed.
const IMPORT_CACHE_CAP: usize = 8; const IMPORT_CACHE_CAP: usize = 8;
/// Plane-import cache key: the texture's COM address plus the extent it was imported at.
type PlaneKey = (isize, u32, u32);
// --- Vulkan enum values not surfaced by pyrowave-sys' bindgen (only enums *reachable* from the // --- Vulkan enum values not surfaced by pyrowave-sys' bindgen (only enums *reachable* from the
// pyrowave C API are generated; these plain #define / flags-typedef values are stable spec // pyrowave C API are generated; these plain #define / flags-typedef values are stable spec
// constants). bindgen renders every reachable Vulkan enum as a `u32` type alias, so these u32 // constants). bindgen renders every reachable Vulkan enum as a `u32` type alias, so these u32
@@ -136,8 +139,12 @@ pub struct PyroWaveEncoder {
// Imported plane textures, cached by the out-ring texture's raw pointer (stable per ring slot): // Imported plane textures, cached by the out-ring texture's raw pointer (stable per ring slot):
// the full-res R8 Y plane and the half-res R8G8 CbCr plane, imported SEPARATELY (a single planar // the full-res R8 Y plane and the half-res R8G8 CbCr plane, imported SEPARATELY (a single planar
// NV12 import is unreliable on NVIDIA at arbitrary sizes). // NV12 import is unreliable on NVIDIA at arbitrary sizes).
y_images: Vec<(isize, pw::pyrowave_image)>, /// The capturer ring generation the cached plane imports below belong to. A recreate bumps it,
cbcr_images: Vec<(isize, pw::pyrowave_image)>, /// and every cached import is destroyed — the COM addresses they are keyed on can be recycled
/// by the allocator after a recreate, so identity cannot rest on the pointer alone.
ring_gen: Option<u32>,
y_images: Vec<(PlaneKey, pw::pyrowave_image)>,
cbcr_images: Vec<(PlaneKey, pw::pyrowave_image)>,
width: u32, width: u32,
height: u32, height: u32,
@@ -268,6 +275,7 @@ impl PyroWaveEncoder {
pw_dev, pw_dev,
pw_enc, pw_enc,
sync: std::ptr::null_mut(), sync: std::ptr::null_mut(),
ring_gen: None,
y_images: Vec::new(), y_images: Vec::new(),
cbcr_images: Vec::new(), cbcr_images: Vec::new(),
width, width,
@@ -351,10 +359,16 @@ impl PyroWaveEncoder {
/// ///
/// # Safety /// # Safety
/// Same contract as [`import_plane`]. /// Same contract as [`import_plane`].
/// Keyed on `(texture address, width, height)` rather than the bare address: the COM pointer
/// carries no reference here, so a released texture's address can be recycled by a later
/// allocation and return an import describing the WRONG surface. Folding the extent in means a
/// recycled address at a different size can never alias. (A recycle at the SAME size is still
/// possible in principle — the complete fix is to key on the capturer's ring generation, which
/// needs that generation plumbed onto `PyroFrameShare`.)
unsafe fn cached_plane( unsafe fn cached_plane(
cache: &mut Vec<(isize, pw::pyrowave_image)>, cache: &mut Vec<(PlaneKey, pw::pyrowave_image)>,
make: impl FnOnce() -> Result<pw::pyrowave_image>, make: impl FnOnce() -> Result<pw::pyrowave_image>,
key: isize, key: PlaneKey,
) -> Result<pw::pyrowave_image> { ) -> Result<pw::pyrowave_image> {
if let Some((_, img)) = cache.iter().find(|(k, _)| *k == key) { if let Some((_, img)) = cache.iter().find(|(k, _)| *k == key) {
return Ok(*img); return Ok(*img);
@@ -423,6 +437,21 @@ impl PyroWaveEncoder {
!self.pw_enc.is_null(), !self.pw_enc.is_null(),
"pyrowave: encode after a failed reset (encoder was destroyed and not rebuilt)" "pyrowave: encode after a failed reset (encoder was destroyed and not rebuilt)"
); );
// The plane textures are imported at the encoder's CONFIGURED extent, not the frame's, so a
// capture that changed size would be read under a stale `VkImageCreateInfo`. This is
// reachable without any client Reconfigure: the IDD capturer autonomously recreates its ring
// on a confirmed display-descriptor change (e.g. a fullscreen game mode-setting the virtual
// display). Refuse instead — the session must reopen the encoder at the new mode. Mirrors
// the guard the QSV and AMF backends already carry.
anyhow::ensure!(
frame.width == self.width && frame.height == self.height,
"pyrowave: captured frame {}x{} != encoder {}x{} (the capturer recreated its ring at a \
new mode the encoder must be reopened)",
frame.width,
frame.height,
self.width,
self.height
);
let FramePayload::D3d11(d3d) = &frame.payload else { let FramePayload::D3d11(d3d) = &frame.payload else {
bail!("pyrowave (Windows) needs a D3D11 frame (the capturer must be in pyrowave mode)") bail!("pyrowave (Windows) needs a D3D11 frame (the capturer must be in pyrowave mode)")
}; };
@@ -431,6 +460,25 @@ impl PyroWaveEncoder {
in pyrowave mode (session_plan::output_format must set OutputFormat::pyrowave)", in pyrowave mode (session_plan::output_format must set OutputFormat::pyrowave)",
)?; )?;
// Ring recreate ⇒ every cached plane import belongs to textures that no longer exist. Their
// COM addresses can be handed back out by the allocator, so a pointer-keyed hit could return
// an image bound to freed memory. Flush on the generation change rather than relying on the
// address (or the FIFO cap) to notice.
if self.ring_gen != Some(share.ring_gen) {
if self.ring_gen.is_some() {
tracing::info!(
from = ?self.ring_gen,
to = share.ring_gen,
cached = self.y_images.len() + self.cbcr_images.len(),
"pyrowave: capturer recreated its ring — flushing stale plane imports"
);
}
for (_, img) in self.y_images.drain(..).chain(self.cbcr_images.drain(..)) {
pw::pyrowave_image_destroy(img);
}
self.ring_gen = Some(share.ring_gen);
}
// Import the fence whenever this encoder has no timeline yet — the first frame, OR a fresh // Import the fence whenever this encoder has no timeline yet — the first frame, OR a fresh
// encoder after a client mode-switch rebuild (the capturer passes the persistent handle on // encoder after a client mode-switch rebuild (the capturer passes the persistent handle on
// every frame precisely so a rebuilt encoder can re-import it). // every frame precisely so a rebuilt encoder can re-import it).
@@ -465,7 +513,7 @@ impl PyroWaveEncoder {
}; };
let pw_dev = self.pw_dev; let pw_dev = self.pw_dev;
let y_img = { let y_img = {
let key = d3d.texture.as_raw() as isize; let key = (d3d.texture.as_raw() as isize, w, h);
let tex = &d3d.texture; let tex = &d3d.texture;
Self::cached_plane( Self::cached_plane(
&mut self.y_images, &mut self.y_images,
@@ -474,7 +522,7 @@ impl PyroWaveEncoder {
)? )?
}; };
let cbcr_img = { let cbcr_img = {
let key = share.cbcr.as_raw() as isize; let key = (share.cbcr.as_raw() as isize, cw, ch);
let tex = &share.cbcr; let tex = &share.cbcr;
Self::cached_plane( Self::cached_plane(
&mut self.cbcr_images, &mut self.cbcr_images,
@@ -976,6 +1024,9 @@ mod tests {
cbcr: cbcr_tex, cbcr: cbcr_tex,
fence_handle: Some(fence_handle.0 as isize), fence_handle: Some(fence_handle.0 as isize),
fence_value: 1, fence_value: 1,
// One synthetic ring for the whole case: a constant generation exercises the
// steady-state cache-hit path (a changing one would flush every frame).
ring_gen: 1,
}), }),
}), }),
cursor: None, cursor: None,
+306 -12
View File
@@ -146,7 +146,12 @@ const NUM_LTR_SLOTS: usize = 2;
/// `PUNKTFUNK_NO_QSV_LTR` — defeat switch for the LTR-RFI path (parity with /// `PUNKTFUNK_NO_QSV_LTR` — defeat switch for the LTR-RFI path (parity with
/// `PUNKTFUNK_NO_AMF_LTR`); loss recovery then always falls back to IDR. /// `PUNKTFUNK_NO_AMF_LTR`); loss recovery then always falls back to IDR.
fn ltr_disabled() -> bool { fn ltr_disabled() -> bool {
std::env::var("PUNKTFUNK_NO_QSV_LTR").is_ok_and(|v| v == "1" || v.eq_ignore_ascii_case("true")) // Same accepted spellings as AMF's `ltr_disabled` — this had dropped the `trim()` and the
// `yes`/`on` forms, so a value with stray whitespace (easy to produce with `set VAR=1 `)
// silently left LTR enabled on Intel while the identical value worked on AMD.
std::env::var("PUNKTFUNK_NO_QSV_LTR")
.map(|v| matches!(v.trim(), "1" | "true" | "yes" | "on"))
.unwrap_or(false)
} }
/// Frames between LTR marks (`PUNKTFUNK_LTR_INTERVAL_FRAMES`, shared with AMF); default ~1/4 s /// Frames between LTR marks (`PUNKTFUNK_LTR_INTERVAL_FRAMES`, shared with AMF); default ~1/4 s
@@ -171,13 +176,22 @@ fn ltr_test_force_at() -> Option<i64> {
/// Mirrors [`super::amf`]'s `PUNKTFUNK_INTRA_REFRESH` opt-in: request the intra-refresh wave /// Mirrors [`super::amf`]'s `PUNKTFUNK_INTRA_REFRESH` opt-in: request the intra-refresh wave
/// instead of LTR (mutually exclusive — the wave sweeps the whole picture, LTR pins references). /// instead of LTR (mutually exclusive — the wave sweeps the whole picture, LTR pins references).
fn intra_refresh_requested() -> bool { fn intra_refresh_requested() -> bool {
// Spelling parity with AMF (see `ltr_disabled` above).
std::env::var("PUNKTFUNK_INTRA_REFRESH") std::env::var("PUNKTFUNK_INTRA_REFRESH")
.is_ok_and(|v| v == "1" || v.eq_ignore_ascii_case("true")) .map(|v| matches!(v.trim(), "1" | "true" | "yes" | "on"))
.unwrap_or(false)
} }
/// The wave period in frames (~0.5 s), the same shape as Linux NVENC / AMF. /// The wave period in frames (~0.5 s), `PUNKTFUNK_IR_PERIOD_FRAMES` overrides — the same knob and
/// default as AMF / Linux NVENC. (This claimed parity while ignoring the env var entirely, so the
/// knob silently did nothing on Intel; the clamp is kept because `mfxU16` bounds the field.)
fn intra_refresh_period(fps: u32) -> u16 { fn intra_refresh_period(fps: u32) -> u16 {
(fps / 2).clamp(8, 240) as u16 std::env::var("PUNKTFUNK_IR_PERIOD_FRAMES")
.ok()
.and_then(|s| s.trim().parse::<u32>().ok())
.filter(|v| *v >= 2)
.unwrap_or(fps / 2)
.clamp(8, 240) as u16
} }
// --------------------------------------------------------------------------------------------- // ---------------------------------------------------------------------------------------------
@@ -464,10 +478,14 @@ fn build_params(cfg: &EncodeConfig) -> ParamSet {
b b
}); });
// HDR signalling (10-bit sessions are the HDR path on Windows — same coupling as NVENC): // Colour signalling, written UNCONDITIONALLY (mirrors nvenc_core.rs): the input is already
// BT.2020/PQ colour description + the source's mastering/CLL grade at every IDR. // CSC'd to a specific matrix — BT.709 limited for SDR (the capture-side VideoConverter),
// BT.2020 PQ for HDR (HdrP010Converter) — so the stream must say so. An SDR stream without a
// colour description leaves the choice to the decoder's "unspecified" default, and
// Moonlight/third-party/Android-vendor decoders default to 601 at sub-HD → mis-rendered
// colours. (10-bit sessions are the HDR path on Windows — same coupling as NVENC.)
let hdr = cfg.ten_bit && cfg.codec != Codec::H264; let hdr = cfg.ten_bit && cfg.codec != Codec::H264;
let vsi = hdr.then(|| { let vsi = {
// SAFETY: all-zero is valid; header stamped below. // SAFETY: all-zero is valid; header stamped below.
let mut b: Box<vpl::mfxExtVideoSignalInfo> = Box::new(unsafe { std::mem::zeroed() }); let mut b: Box<vpl::mfxExtVideoSignalInfo> = Box::new(unsafe { std::mem::zeroed() });
b.Header.BufferId = vpl::MFX_EXTBUFF_VIDEO_SIGNAL_INFO as u32; b.Header.BufferId = vpl::MFX_EXTBUFF_VIDEO_SIGNAL_INFO as u32;
@@ -475,11 +493,17 @@ fn build_params(cfg: &EncodeConfig) -> ParamSet {
b.VideoFormat = 5; // unspecified b.VideoFormat = 5; // unspecified
b.VideoFullRange = 0; b.VideoFullRange = 0;
b.ColourDescriptionPresent = 1; b.ColourDescriptionPresent = 1;
if hdr {
b.ColourPrimaries = 9; // BT.2020 b.ColourPrimaries = 9; // BT.2020
b.TransferCharacteristics = 16; // SMPTE ST 2084 (PQ) b.TransferCharacteristics = 16; // SMPTE ST 2084 (PQ)
b.MatrixCoefficients = 9; // BT.2020 non-constant b.MatrixCoefficients = 9; // BT.2020 non-constant
b } else {
}); b.ColourPrimaries = 1; // BT.709
b.TransferCharacteristics = 1; // BT.709
b.MatrixCoefficients = 1; // BT.709
}
Some(b)
};
let mastering = cfg.hdr_meta.filter(|_| hdr).map(|m| { let mastering = cfg.hdr_meta.filter(|_| hdr).map(|m| {
// SAFETY: all-zero is valid; header stamped below. // SAFETY: all-zero is valid; header stamped below.
let mut b: Box<vpl::mfxExtMasteringDisplayColourVolume> = let mut b: Box<vpl::mfxExtMasteringDisplayColourVolume> =
@@ -700,7 +724,16 @@ pub struct QsvEncoder {
/// `EncoderCaps::supports_rfi` and all per-frame marking/forcing below. /// `EncoderCaps::supports_rfi` and all per-frame marking/forcing below.
ltr_active: bool, ltr_active: bool,
/// The wire frame index stored in each LTR slot (`None` = never marked). /// The wire frame index stored in each LTR slot (`None` = never marked).
///
/// This mirrors the HARDWARE DPB, so an entry must not be cleared merely because we distrust
/// it: nulling issues no VPL call, and the encoder keeps the frame marked long-term until that
/// `LongTermIdx` is re-marked or an IDR flushes it. Distrust is recorded in `ltr_tainted`
/// instead, so the rejection list can still NAME the entry the hardware is holding.
ltr_slots: [Option<i64>; NUM_LTR_SLOTS], ltr_slots: [Option<i64>; NUM_LTR_SLOTS],
/// Per-slot taint from `invalidate_ref_frames`' sweep: the mark is still live in the hardware
/// DPB but was encoded inside the client's corrupt window, so it may not anchor a recovery —
/// it must be REJECTED instead. Cleared wherever the slot is re-marked or the DPB is flushed.
ltr_tainted: [bool; NUM_LTR_SLOTS],
next_ltr_slot: usize, next_ltr_slot: usize,
ltr_mark_interval: i64, ltr_mark_interval: i64,
/// Set by `invalidate_ref_frames`: the slot the next submitted frame force-references. /// Set by `invalidate_ref_frames`: the slot the next submitted frame force-references.
@@ -775,6 +808,7 @@ impl QsvEncoder {
ir_active: false, ir_active: false,
ltr_active: false, ltr_active: false,
ltr_slots: [None; NUM_LTR_SLOTS], ltr_slots: [None; NUM_LTR_SLOTS],
ltr_tainted: [false; NUM_LTR_SLOTS],
next_ltr_slot: 0, next_ltr_slot: 0,
ltr_mark_interval: ltr_mark_interval(fps), ltr_mark_interval: ltr_mark_interval(fps),
pending_force: None, pending_force: None,
@@ -899,6 +933,7 @@ impl QsvEncoder {
self.ltr_active = ltr_active; self.ltr_active = ltr_active;
self.ir_active = ir_active; self.ir_active = ir_active;
self.ltr_slots = [None; NUM_LTR_SLOTS]; self.ltr_slots = [None; NUM_LTR_SLOTS];
self.ltr_tainted = [false; NUM_LTR_SLOTS];
self.next_ltr_slot = 0; self.next_ltr_slot = 0;
self.pending_force = None; self.pending_force = None;
self.hdr_applied = self.hdr_meta; self.hdr_applied = self.hdr_meta;
@@ -1024,6 +1059,7 @@ impl Encoder for QsvEncoder {
// An IDR voids the decoder's reference buffers — drop stale slots and any // An IDR voids the decoder's reference buffers — drop stale slots and any
// queued force; the mark cadence below re-anchors on the IDR itself. // queued force; the mark cadence below re-anchors on the IDR itself.
self.ltr_slots = [None; NUM_LTR_SLOTS]; self.ltr_slots = [None; NUM_LTR_SLOTS];
self.ltr_tainted = [false; NUM_LTR_SLOTS]; // the IDR flushed the DPB with them
self.next_ltr_slot = 0; self.next_ltr_slot = 0;
self.pending_force = None; self.pending_force = None;
} else if self.ltr_test_force_at == Some(cur_idx) { } else if self.ltr_test_force_at == Some(cur_idx) {
@@ -1039,7 +1075,9 @@ impl Encoder for QsvEncoder {
// emptied the slot since the force was queued. An empty slot means there is // emptied the slot since the force was queued. An empty slot means there is
// nothing clean to re-reference — the frame must ship as a plain P WITHOUT the // nothing clean to re-reference — the frame must ship as a plain P WITHOUT the
// `recovery_anchor` tag (the client lifts its post-loss freeze on that tag). // `recovery_anchor` tag (the client lifts its post-loss freeze on that tag).
if let Some(idx) = self.ltr_slots[slot] { // The slot is no longer emptied by the sweep, so test the taint flag too — a
// tainted slot is exactly the "nothing clean to re-reference" case.
if let Some(idx) = self.ltr_slots[slot].filter(|_| !self.ltr_tainted[slot]) {
force_ltr = Some((slot, idx)); force_ltr = Some((slot, idx));
recovery_anchor = true; recovery_anchor = true;
} }
@@ -1047,6 +1085,9 @@ impl Encoder for QsvEncoder {
if force_ltr.is_none() && (forced || cur_idx % self.ltr_mark_interval == 0) { if force_ltr.is_none() && (forced || cur_idx % self.ltr_mark_interval == 0) {
let slot = self.next_ltr_slot; let slot = self.next_ltr_slot;
self.ltr_slots[slot] = Some(cur_idx); self.ltr_slots[slot] = Some(cur_idx);
// Re-marking replaces the hardware's LongTermIdx: the tainted frame is gone from
// the DPB and this slot is clean again.
self.ltr_tainted[slot] = false;
self.next_ltr_slot = (self.next_ltr_slot + 1) % NUM_LTR_SLOTS; self.next_ltr_slot = (self.next_ltr_slot + 1) % NUM_LTR_SLOTS;
mark_slot = Some(slot); mark_slot = Some(slot);
} }
@@ -1309,13 +1350,23 @@ impl Encoder for QsvEncoder {
// loss ships corruption as the recovery anchor, and every subsequent mark re-samples // loss ships corruption as the recovery anchor, and every subsequent mark re-samples
// the soup — the sustained-loss field failure where the picture never healed. Dropped // the soup — the sustained-loss field failure where the picture never healed. Dropped
// slots stay dropped; the cadence re-marks a clean frame within ~1/4 s. // slots stay dropped; the cadence re-marks a clean frame within ~1/4 s.
for marked in self.ltr_slots.iter_mut() { //
// Mark tainted rather than clearing: `ltr_slots` mirrors the HARDWARE DPB, and nulling an
// entry issues no VPL call — the frame stays marked long-term in the encoder. Clearing it
// made the rejection list below (which iterates the post-sweep mirror and only names `Some`
// slots) silently SKIP the one entry the sweep exists to distrust, so the recovery frame
// could still predict from it. With two slots the "exactly one swept" case is the modal
// one, and it was the broken one.
for (slot, marked) in self.ltr_slots.iter().enumerate() {
if marked.is_some_and(|idx| idx >= first) { if marked.is_some_and(|idx| idx >= first) {
*marked = None; self.ltr_tainted[slot] = true;
} }
} }
let mut best: Option<(usize, i64)> = None; let mut best: Option<(usize, i64)> = None;
for (slot, marked) in self.ltr_slots.iter().enumerate() { for (slot, marked) in self.ltr_slots.iter().enumerate() {
if self.ltr_tainted[slot] {
continue; // still in the DPB, but encoded inside the corrupt window
}
if let Some(idx) = *marked { if let Some(idx) = *marked {
if idx < first && best.is_none_or(|(_, b)| idx > b) { if idx < first && best.is_none_or(|(_, b)| idx > b) {
best = Some((slot, idx)); best = Some((slot, idx));
@@ -1440,6 +1491,7 @@ impl Encoder for QsvEncoder {
self.ltr_active = ltr; self.ltr_active = ltr;
self.ir_active = ir; self.ir_active = ir;
self.ltr_slots = [None; NUM_LTR_SLOTS]; self.ltr_slots = [None; NUM_LTR_SLOTS];
self.ltr_tainted = [false; NUM_LTR_SLOTS];
self.next_ltr_slot = 0; self.next_ltr_slot = 0;
self.pending_force = None; self.pending_force = None;
if let Some(inner) = self.inner.as_mut() { if let Some(inner) = self.inner.as_mut() {
@@ -1952,4 +2004,246 @@ mod tests {
"the bitrate retarget emitted a keyframe (StartNewSequence leak)" "the bitrate retarget emitted a keyframe (StartNewSequence leak)"
); );
} }
/// FULL-CHAIN colour check at the field capture size: a known P010 colour-bar source at
/// 1920x1080 — whose height is NOT 16-aligned, so the ingest `CopySubresourceRegion` copies
/// into a 1920x1088 runtime pool surface whose chroma plane sits at a DIFFERENT row offset
/// than the source's (the seam no 640x480 test exercises) — encoded to Main10 HEVC and
/// dumped to `%TEMP%\pf_qsv_1080_bars.h265` for off-box decode verification against the
/// same codes. On-box this asserts stream shape; the pixel verdict needs a decoder.
#[test]
fn qsv_live_p010_1080_colorbars_dump() {
use windows::Win32::Graphics::Direct3D::D3D_DRIVER_TYPE_UNKNOWN;
use windows::Win32::Graphics::Direct3D11::{
D3D11CreateDevice, D3D11_BIND_RENDER_TARGET, D3D11_BIND_SHADER_RESOURCE,
D3D11_SDK_VERSION, D3D11_SUBRESOURCE_DATA, D3D11_USAGE_DEFAULT,
};
use windows::Win32::Graphics::Dxgi::Common::{DXGI_FORMAT_P010, DXGI_SAMPLE_DESC};
use windows::Win32::Graphics::Dxgi::{CreateDXGIFactory1, IDXGIAdapter1, IDXGIFactory4};
// (Y, Cb, Cr) 10-bit limited codes for the 8 sRGB bars white/yellow/cyan/green/magenta/
// red/blue/black at 80-nit SDR white under PQ/BT.2020 — the same math as pf-capture's
// `p010_reference` (and the bars_pq2020 client fixture). Stored MSB-aligned (`<<6`).
const BARS: [(u16, u16, u16); 8] = [
(490, 512, 512),
(478, 423, 518),
(464, 525, 473),
(450, 432, 476),
(350, 584, 585),
(325, 448, 598),
(226, 650, 535),
(64, 512, 512),
];
const W: u32 = 1920;
const H: u32 = 1080;
init_tracing();
let Ok((_l, impls)) = intel_loader() else {
eprintln!("skipping: no VPL loader");
return;
};
let Some(imp) = impls.iter().find(|i| i.luid_valid) else {
eprintln!("skipping: no Intel VPL implementation on this box");
return;
};
if !probe_can_encode_10bit(Codec::H265) {
eprintln!("skipping: this GPU declines 10-bit HEVC");
return;
}
// P010 initial data: plane 0 = H rows of W u16 luma; plane 1 = H/2 rows of W u16
// (interleaved Cb,Cr pairs), same pitch. Bars are vertical: bar index = x / (W/8).
let bar_w = (W / 8) as usize;
let mut init = vec![0u16; (W as usize) * (H as usize + H as usize / 2)];
for y in 0..H as usize {
for x in 0..W as usize {
init[y * W as usize + x] = BARS[(x / bar_w).min(7)].0 << 6;
}
}
let chroma_base = (W as usize) * (H as usize);
for cy in 0..(H as usize / 2) {
for cx in 0..(W as usize / 2) {
let (_, cb, cr) = BARS[((cx * 2) / bar_w).min(7)];
init[chroma_base + cy * W as usize + cx * 2] = cb << 6;
init[chroma_base + cy * W as usize + cx * 2 + 1] = cr << 6;
}
}
// SAFETY: self-contained harness on one thread/device (same contract as `drive_live`);
// the initial-data pointer outlives the synchronous CreateTexture2D that reads it.
let (device, tex) = unsafe {
let luid = windows::Win32::Foundation::LUID {
LowPart: u32::from_le_bytes(imp.luid[..4].try_into().unwrap()),
HighPart: i32::from_le_bytes(imp.luid[4..].try_into().unwrap()),
};
let factory: IDXGIFactory4 = CreateDXGIFactory1().expect("dxgi factory");
let adapter: IDXGIAdapter1 = factory.EnumAdapterByLuid(luid).expect("intel adapter");
let mut device = None;
D3D11CreateDevice(
&adapter,
D3D_DRIVER_TYPE_UNKNOWN,
windows::Win32::Foundation::HMODULE::default(),
Default::default(),
None,
D3D11_SDK_VERSION,
Some(&mut device),
None,
None,
)
.expect("d3d11 device on intel adapter");
let device: ID3D11Device = device.expect("device");
let desc = D3D11_TEXTURE2D_DESC {
Width: W,
Height: H,
MipLevels: 1,
ArraySize: 1,
Format: DXGI_FORMAT_P010,
SampleDesc: DXGI_SAMPLE_DESC {
Count: 1,
Quality: 0,
},
Usage: D3D11_USAGE_DEFAULT,
BindFlags: (D3D11_BIND_RENDER_TARGET.0 | D3D11_BIND_SHADER_RESOURCE.0) as u32,
CPUAccessFlags: 0,
MiscFlags: 0,
};
let data = D3D11_SUBRESOURCE_DATA {
pSysMem: init.as_ptr() as *const std::ffi::c_void,
SysMemPitch: W * 2,
SysMemSlicePitch: 0,
};
let mut t: Option<ID3D11Texture2D> = None;
device
.CreateTexture2D(&desc, Some(&data), Some(&mut t))
.expect("bar texture");
(device.clone(), t.expect("texture"))
};
let mut enc = QsvEncoder::open(
Codec::H265,
PixelFormat::P010,
W,
H,
30,
10_000_000,
10,
ChromaFormat::Yuv420,
)
.expect("open");
enc.set_hdr_meta(Some(test_hdr_meta()));
let mut stream = Vec::new();
let mut aus = 0usize;
let mut keyframes = 0usize;
for i in 0..12u32 {
let frame = CapturedFrame {
width: W,
height: H,
pts_ns: i as u64 * 33_333_333,
format: PixelFormat::P010,
payload: FramePayload::D3d11(pf_frame::dxgi::D3d11Frame {
texture: tex.clone(),
device: device.clone(),
pyro: None,
}),
cursor: None,
};
enc.submit_indexed(&frame, i).expect("submit");
if let Some(au) = enc.poll().expect("poll") {
aus += 1;
keyframes += au.keyframe as usize;
stream.extend_from_slice(&au.data);
}
}
enc.flush().expect("flush");
while let Some(au) = enc.poll().expect("drain") {
aus += 1;
keyframes += au.keyframe as usize;
stream.extend_from_slice(&au.data);
}
assert!(aus >= 10, "expected ≥10 AUs, got {aus}");
assert!(keyframes >= 1, "expected an IDR in the dump");
let path = std::env::temp_dir().join("pf_qsv_1080_bars.h265");
std::fs::write(&path, &stream).expect("write dump");
println!(
"wrote {} AUs ({} bytes, {keyframes} keyframes) to {}",
aus,
stream.len(),
path.display()
);
}
/// The PRODUCTION host chain minus the IDD ring: the REAL `HdrP010Converter` renders the 8
/// sRGB bars into a ring-profile P010 texture (`BIND_RENDER_TARGET` only — RTV-written, not
/// CPU-uploaded) on the VPL implementation's own adapter, and THAT texture goes through the
/// unaligned-height ingest copy into a Main10 encode. Dumped to
/// `%TEMP%\pf_qsv_conv_1080_bars.h265`; expected decode codes = the bars_pq2020 fixture set
/// (see `hdr_p010_convert_bars_on_luid`).
#[test]
fn qsv_live_hdr_converter_e2e_1080_dump() {
const W: u32 = 1920;
const H: u32 = 1080;
init_tracing();
let Ok((_l, impls)) = intel_loader() else {
eprintln!("skipping: no VPL loader");
return;
};
let Some(imp) = impls.iter().find(|i| i.luid_valid) else {
eprintln!("skipping: no Intel VPL implementation on this box");
return;
};
if !probe_can_encode_10bit(Codec::H265) {
eprintln!("skipping: this GPU declines 10-bit HEVC");
return;
}
let (device, tex) = pf_capture::dxgi::hdr_p010_convert_bars_on_luid(imp.luid, W, H)
.expect("converter bars");
let mut enc = QsvEncoder::open(
Codec::H265,
PixelFormat::P010,
W,
H,
30,
10_000_000,
10,
ChromaFormat::Yuv420,
)
.expect("open");
enc.set_hdr_meta(Some(test_hdr_meta()));
let mut stream = Vec::new();
let mut aus = 0usize;
for i in 0..12u32 {
let frame = CapturedFrame {
width: W,
height: H,
pts_ns: i as u64 * 33_333_333,
format: PixelFormat::P010,
payload: FramePayload::D3d11(pf_frame::dxgi::D3d11Frame {
texture: tex.clone(),
device: device.clone(),
pyro: None,
}),
cursor: None,
};
enc.submit_indexed(&frame, i).expect("submit");
if let Some(au) = enc.poll().expect("poll") {
aus += 1;
stream.extend_from_slice(&au.data);
}
}
enc.flush().expect("flush");
while let Some(au) = enc.poll().expect("drain") {
aus += 1;
stream.extend_from_slice(&au.data);
}
assert!(aus >= 10, "expected ≥10 AUs, got {aus}");
let path = std::env::temp_dir().join("pf_qsv_conv_1080_bars.h265");
std::fs::write(&path, &stream).expect("write dump");
println!(
"wrote {aus} AUs ({} bytes) to {}",
stream.len(),
path.display()
);
}
} }
+123 -7
View File
@@ -129,6 +129,9 @@ pub fn open_video(
cuda: bool, cuda: bool,
bit_depth: u8, bit_depth: u8,
chroma: ChromaFormat, chroma: ChromaFormat,
// The session may hand this encoder cursor bitmaps to composite (cursor-as-metadata
// captures). Backends whose fast path can't blend (Vulkan EFC RGB-direct) key off it.
cursor_blend: bool,
) -> Result<Box<dyn Encoder>> { ) -> Result<Box<dyn Encoder>> {
let (inner, backend) = open_video_backend( let (inner, backend) = open_video_backend(
codec, codec,
@@ -140,6 +143,7 @@ pub fn open_video(
cuda, cuda,
bit_depth, bit_depth,
chroma, chroma,
cursor_blend,
)?; )?;
// Record what this session encodes on (the mgmt API's "currently used GPU"): the backend label // Record what this session encodes on (the mgmt API's "currently used GPU"): the backend label
// is reported by `open_video_backend` from the branch that ACTUALLY opened — not re-derived by // is reported by `open_video_backend` from the branch that ACTUALLY opened — not re-derived by
@@ -203,15 +207,34 @@ impl Encoder for TrackedEncoder {
fn invalidate_ref_frames(&mut self, first_frame: i64, last_frame: i64) -> bool { fn invalidate_ref_frames(&mut self, first_frame: i64, last_frame: i64) -> bool {
self.inner.invalidate_ref_frames(first_frame, last_frame) self.inner.invalidate_ref_frames(first_frame, last_frame)
} }
// Forwarded for the same reason as `set_wire_chunking` below — the unforwarded default
// (`false` = "backend can't pipeline, stop asking") silently killed the §7 LN3 contention
// escalation for every session, since the host loop only ever holds the wrapped box.
fn set_pipelined(&mut self, on: bool) -> bool {
self.inner.set_pipelined(on)
}
// The classic TrackedEncoder trap: a defaulted trait method that isn't forwarded // The classic TrackedEncoder trap: a defaulted trait method that isn't forwarded
// silently no-ops through the wrapper (bit the direct-NVENC work, then THIS — the // silently no-ops through the wrapper (bit the direct-NVENC work, then THIS — the
// §4.4 chunking probe run hit the default while the plan said Some(1408)). // §4.4 chunking probe run hit the default while the plan said Some(1408)).
fn set_wire_chunking(&mut self, shard_payload: usize) { fn set_wire_chunking(&mut self, shard_payload: usize) {
self.inner.set_wire_chunking(shard_payload) self.inner.set_wire_chunking(shard_payload)
} }
// Forwarded for the same reason as `set_wire_chunking` above — an unforwarded default here
// would silently leave the in-place backends pipelining past the capturer's ring.
fn set_input_ring_depth(&mut self, depth: usize) {
self.inner.set_input_ring_depth(depth)
}
fn poll(&mut self) -> Result<Option<EncodedFrame>> { fn poll(&mut self) -> Result<Option<EncodedFrame>> {
self.inner.poll() self.inner.poll()
} }
// Both chunked-poll methods forwarded (the same trap class): the defaults would report
// "not chunked" and wrap whole AUs, silently discarding the sub-frame overlap.
fn supports_chunked_poll(&self) -> bool {
self.inner.supports_chunked_poll()
}
fn poll_chunk(&mut self) -> Result<Option<AuChunk>> {
self.inner.poll_chunk()
}
fn reset(&mut self) -> bool { fn reset(&mut self) -> bool {
self.inner.reset() self.inner.reset()
} }
@@ -223,6 +246,13 @@ impl Encoder for TrackedEncoder {
} }
} }
/// Ceiling applied to the negotiated bitrate before it reaches openh264: software H.264 realistically
/// caps far below the rates a hardware session negotiates, and handing it the full figure just
/// misconfigures its rate control. Module-scope so BOTH software arms share one value — the Linux
/// arm was missing the clamp the Windows arm applied.
#[cfg(any(target_os = "linux", target_os = "windows"))]
const SW_BITRATE_CEIL: u64 = 100_000_000;
/// Open the platform encoder backend. Returns the encoder together with the display label of the /// Open the platform encoder backend. Returns the encoder together with the display label of the
/// branch that ACTUALLY opened (`nvenc`/`vaapi`/`vulkan`/`amf`/`qsv`/`software`) — the label feeds /// branch that ACTUALLY opened (`nvenc`/`vaapi`/`vulkan`/`amf`/`qsv`/`software`) — the label feeds
/// the mgmt API's live-session record, and only the open site knows which internal fallback won /// the mgmt API's live-session record, and only the open site knows which internal fallback won
@@ -238,7 +268,9 @@ fn open_video_backend(
cuda: bool, cuda: bool,
bit_depth: u8, bit_depth: u8,
chroma: ChromaFormat, chroma: ChromaFormat,
cursor_blend: bool,
) -> Result<(Box<dyn Encoder>, &'static str)> { ) -> Result<(Box<dyn Encoder>, &'static str)> {
let _ = cursor_blend; // consumed only by the Linux vulkan-encode arm below
validate_dimensions(codec, width, height)?; validate_dimensions(codec, width, height)?;
// Refresh/fps must be positive and sane: fps feeds the encoder time_base (`Rational(1, fps)`) // Refresh/fps must be positive and sane: fps feeds the encoder time_base (`Rational(1, fps)`)
// and the pts→ns conversion (`pts * 1e9 / fps`), so 0 builds a 1/0 rational / divides by zero. // and the pts→ns conversion (`pts * 1e9 / fps`), so 0 builds a 1/0 rational / divides by zero.
@@ -296,8 +328,15 @@ fn open_video_backend(
&& vulkan_encode_enabled() && vulkan_encode_enabled()
&& !(bit_depth == 10 && format.is_hdr_rgb10()) && !(bit_depth == 10 && format.is_hdr_rgb10())
{ {
match vulkan_video::VulkanVideoEncoder::open(codec, width, height, fps, bitrate_bps) match vulkan_video::VulkanVideoEncoder::open(
{ codec,
format,
width,
height,
fps,
bitrate_bps,
cursor_blend,
) {
Ok(e) => { Ok(e) => {
tracing::info!( tracing::info!(
codec = ?codec, codec = ?codec,
@@ -306,12 +345,32 @@ fn open_video_backend(
); );
return Ok((Box::new(e) as Box<dyn Encoder>, "vulkan")); return Ok((Box::new(e) as Box<dyn Encoder>, "vulkan"));
} }
// Native NV12 (PUNKTFUNK_PIPEWIRE_NV12 capture) has no VAAPI fallback:
// libav's dmabuf lane would import the two-plane buffer as packed RGB
// (silent garbage) and its CPU lane bails per frame — die crisply instead.
Err(e) if format == PixelFormat::Nv12 => {
return Err(e.context(
"Vulkan Video open failed on a native-NV12 capture \
no VAAPI fallback exists; set PUNKTFUNK_PIPEWIRE_NV12=0 to \
restore the packed-RGB negotiation",
));
}
Err(e) => tracing::warn!( Err(e) => tracing::warn!(
error = %format!("{e:#}"), error = %format!("{e:#}"),
"Vulkan Video encode open failed — falling back to libav VAAPI" "Vulkan Video encode open failed — falling back to libav VAAPI"
), ),
} }
} }
// Same rule when the Vulkan backend was never eligible (H264 session,
// PUNKTFUNK_VULKAN_ENCODE=0, or a build without the feature).
if format == PixelFormat::Nv12 {
anyhow::bail!(
"native NV12 capture requires the Vulkan Video encoder (HEVC/AV1 \
session, --features vulkan-encode, PUNKTFUNK_VULKAN_ENCODE not 0) this \
session resolved to libav VAAPI; set PUNKTFUNK_PIPEWIRE_NV12=0 to restore \
the packed-RGB negotiation"
);
}
vaapi::VaapiEncoder::open( vaapi::VaapiEncoder::open(
codec, codec,
format, format,
@@ -350,7 +409,15 @@ fn open_video_backend(
"the Vulkan Video encoder supports HEVC + AV1; the session negotiated {codec:?}" "the Vulkan Video encoder supports HEVC + AV1; the session negotiated {codec:?}"
); );
} }
vulkan_video::VulkanVideoEncoder::open(codec, width, height, fps, bitrate_bps) vulkan_video::VulkanVideoEncoder::open(
codec,
format,
width,
height,
fps,
bitrate_bps,
cursor_blend,
)
.map(|e| (Box::new(e) as Box<dyn Encoder>, "vulkan")) .map(|e| (Box::new(e) as Box<dyn Encoder>, "vulkan"))
} }
#[cfg(not(feature = "vulkan-encode"))] #[cfg(not(feature = "vulkan-encode"))]
@@ -407,7 +474,13 @@ fn open_video_backend(
); );
} }
let _ = (cuda, bit_depth); // software path is CPU + 8-bit only let _ = (cuda, bit_depth); // software path is CPU + 8-bit only
sw::OpenH264Encoder::open(format, width, height, fps, bitrate_bps) sw::OpenH264Encoder::open(
format,
width,
height,
fps,
bitrate_bps.min(SW_BITRATE_CEIL),
)
.map(|e| (Box::new(e) as Box<dyn Encoder>, "software")) .map(|e| (Box::new(e) as Box<dyn Encoder>, "software"))
} }
"auto" | "" => { "auto" | "" => {
@@ -610,8 +683,6 @@ fn open_video_backend(
(build a GPU backend: --features nvenc or amf-qsv, or request H264)" (build a GPU backend: --features nvenc or amf-qsv, or request H264)"
); );
let _ = (bit_depth, chroma); // the software H.264 path is 8-bit 4:2:0 only let _ = (bit_depth, chroma); // the software H.264 path is 8-bit 4:2:0 only
// Software H.264 realistically caps far below the negotiated hardware rates.
const SW_BITRATE_CEIL: u64 = 100_000_000;
sw::OpenH264Encoder::open( sw::OpenH264Encoder::open(
format, format,
width, width,
@@ -770,6 +841,33 @@ fn vulkan_encode_enabled() -> bool {
.unwrap_or(true) .unwrap_or(true)
} }
/// Whether THIS session's encoder can ingest a producer-native NV12 capture: only the raw
/// Vulkan Video backend does (libav VAAPI would misread the two-plane buffer as packed RGB —
/// [`open_video`] refuses the combination), so the session's codec must be one it encodes and
/// the backend must be eligible to open. The host facade threads the verdict into the capture
/// negotiation (`OutputFormat::nv12_native` → `ZeroCopyPolicy::native_nv12_session`), which
/// then PREFERS gamescope's producer-side NV12 pod (default-on; `PUNKTFUNK_PIPEWIRE_NV12=0`
/// escapes at the capture gate).
#[cfg(target_os = "linux")]
pub fn linux_native_nv12_ok(codec: Codec) -> bool {
#[cfg(feature = "vulkan-encode")]
{
matches!(codec, Codec::H265 | Codec::Av1)
&& vulkan_encode_enabled()
// NVENC/PyroWave prefs never open the Vulkan Video backend; every other pref
// (auto/vaapi/amd/intel/vulkan) tries it first on AMD/Intel — see [`open_video`].
&& !matches!(
pf_host_config::config().encoder_pref.as_str(),
"nvenc" | "nvidia" | "cuda" | "pyrowave"
)
}
#[cfg(not(feature = "vulkan-encode"))]
{
let _ = codec;
false
}
}
/// Cheap, side-effect-free NVIDIA-presence probe for the `auto` backend selector: the NVIDIA /// Cheap, side-effect-free NVIDIA-presence probe for the `auto` backend selector: the NVIDIA
/// kernel driver exposes these device nodes, AMD/Intel boxes have neither. Deliberately does NOT /// kernel driver exposes these device nodes, AMD/Intel boxes have neither. Deliberately does NOT
/// create a CUDA context (that would allocate GPU state on every host that merely *might* be /// create a CUDA context (that would allocate GPU state on every host that merely *might* be
@@ -784,6 +882,12 @@ fn nvidia_present() -> bool {
/// picks its vendor's backend — AMD/Intel → VAAPI on that GPU's render node, NVIDIA → NVENC (still /// picks its vendor's backend — AMD/Intel → VAAPI on that GPU's render node, NVIDIA → NVENC (still
/// requiring the proprietary driver's device nodes; a nouveau NVIDIA GPU can't NVENC) — otherwise /// requiring the proprietary driver's device nodes; a nouveau NVIDIA GPU can't NVENC) — otherwise
/// today's NVIDIA-presence probe, unchanged. /// today's NVIDIA-presence probe, unchanged.
///
/// ⚠ This resolves the **`auto` case only** — it deliberately ignores `encoder_pref`. It is NOT a
/// mirror of [`open_video`]'s dispatch and must not be used to decide which backend a capability
/// probe should ask: use [`linux_zero_copy_is_vaapi`], which layers `encoder_pref` on top of this.
/// (`can_encode_10bit` used this directly and answered for the wrong backend whenever a host
/// forced one.)
#[cfg(target_os = "linux")] #[cfg(target_os = "linux")]
fn linux_auto_is_vaapi() -> bool { fn linux_auto_is_vaapi() -> bool {
if let Some(g) = pf_gpu::manual_selection() { if let Some(g) = pf_gpu::manual_selection() {
@@ -997,7 +1101,13 @@ pub fn can_encode_10bit(codec: Codec) -> bool {
// only half the Linux gate — the capture side (GNOME 50+ portal monitor in HDR mode) // only half the Linux gate — the capture side (GNOME 50+ portal monitor in HDR mode)
// is resolved separately by the host (`capturer_supports_hdr` / the GameStream RTSP // is resolved separately by the host (`capturer_supports_hdr` / the GameStream RTSP
// honor), since this probe can't know what the compositor will negotiate. // honor), since this probe can't know what the compositor will negotiate.
if linux_auto_is_vaapi() { // Resolve through the SAME helper `can_encode_444` uses (and which mirrors
// `open_video`'s dispatch): `linux_auto_is_vaapi` ignores `encoder_pref`, so on a box
// that forces a backend — e.g. `encoder_pref = "vaapi"` on an NVIDIA host — this probe
// would answer for NVENC while the session actually opens VAAPI, and the negotiated bit
// depth (plus the HDR/SDR colour label derived from it) would describe a backend that
// never runs. That is exactly the dishonesty this probe exists to prevent.
if linux_zero_copy_is_vaapi() {
vaapi::probe_can_encode_10bit(codec) vaapi::probe_can_encode_10bit(codec)
} else { } else {
linux::probe_can_encode_10bit(codec) linux::probe_can_encode_10bit(codec)
@@ -1315,6 +1425,12 @@ mod vulkan_video;
#[cfg(all(target_os = "linux", feature = "vulkan-encode"))] #[cfg(all(target_os = "linux", feature = "vulkan-encode"))]
#[path = "enc/linux/vk_av1_encode.rs"] #[path = "enc/linux/vk_av1_encode.rs"]
mod vk_av1_encode; mod vk_av1_encode;
// Vendored `VK_VALVE_video_encode_rgb_conversion` bindings (host-only) — RGB encode source with
// the VCN EFC front-end doing the CSC (design/vulkan-rgb-direct-encode.md). ash 0.38 predates
// the extension; same vendoring rationale as `vk_av1_encode`.
#[cfg(all(target_os = "linux", feature = "vulkan-encode"))]
#[path = "enc/linux/vk_valve_rgb.rs"]
mod vk_valve_rgb;
// Small ash leaf helpers shared by the Linux Vulkan encode backends (dmabuf import, image/memory // Small ash leaf helpers shared by the Linux Vulkan encode backends (dmabuf import, image/memory
// utilities) — extracted from `vulkan_video.rs` when the PyroWave backend arrived. // utilities) — extracted from `vulkan_video.rs` when the PyroWave backend arrived.
#[cfg(all( #[cfg(all(
+6
View File
@@ -52,6 +52,12 @@ pub struct PyroFrameShare {
/// The fence value the capturer signalled after THIS frame's convert. The encoder's Vulkan /// The fence value the capturer signalled after THIS frame's convert. The encoder's Vulkan
/// acquire waits on it, so the wavelet read is ordered after the D3D11 CSC. /// acquire waits on it, so the wavelet read is ordered after the D3D11 CSC.
pub fence_value: u64, pub fence_value: u64,
/// The capturer's ring generation, bumped every time it recreates its texture ring. The
/// PyroWave encoder caches its plane imports keyed on the texture's COM address, which carries
/// no reference — after a recreate those addresses can be recycled by the allocator, so a
/// cached import may describe a texture that no longer exists. The encoder flushes its import
/// cache whenever this changes, making cache identity independent of allocator behaviour.
pub ring_gen: u32,
} }
/// A GPU-resident captured texture (the Windows zero-copy path: NVENC/AMF/QSV encode it in place; /// A GPU-resident captured texture (the Windows zero-copy path: NVENC/AMF/QSV encode it in place;
+31 -10
View File
@@ -105,13 +105,15 @@ pub fn drm_fourcc(format: PixelFormat) -> Option<u32> {
Bgra => drm_fourcc_code(b"AR24"), // DRM_FORMAT_ARGB8888 Bgra => drm_fourcc_code(b"AR24"), // DRM_FORMAT_ARGB8888
Rgbx => drm_fourcc_code(b"XB24"), // DRM_FORMAT_XBGR8888 Rgbx => drm_fourcc_code(b"XB24"), // DRM_FORMAT_XBGR8888
Rgba => drm_fourcc_code(b"AB24"), // DRM_FORMAT_ABGR8888 Rgba => drm_fourcc_code(b"AB24"), // DRM_FORMAT_ABGR8888
// Linux native NV12 capture (gamescope PipeWire): one LINEAR dmabuf with contiguous Y then
// interleaved UV, exposed under DRM_FORMAT_NV12.
Nv12 => drm_fourcc_code(b"NV12"),
// The GNOME 50+ HDR screencast formats (packed 2:10:10:10, PQ/BT.2020). // The GNOME 50+ HDR screencast formats (packed 2:10:10:10, PQ/BT.2020).
X2Rgb10 => drm_fourcc_code(b"XR30"), // DRM_FORMAT_XRGB2101010 X2Rgb10 => drm_fourcc_code(b"XR30"), // DRM_FORMAT_XRGB2101010
X2Bgr10 => drm_fourcc_code(b"XB30"), // DRM_FORMAT_XBGR2101010 X2Bgr10 => drm_fourcc_code(b"XB30"), // DRM_FORMAT_XBGR2101010
// 24-bit packed RGB/BGR have no straightforward dmabuf import here; use the CPU path. // 24-bit packed RGB/BGR have no straightforward dmabuf import here; use the CPU path.
// Rgb10a2/Nv12/P010 are the Windows HDR / video-processor formats — never produced on // Rgb10a2/P010 are Windows formats; Yuv444 is OUR convert output, never a capture source.
// Linux; Yuv444 is OUR convert's OUTPUT, never a capture source format. Rgb | Bgr | Rgb10a2 | P010 | Yuv444 => return None,
Rgb | Bgr | Rgb10a2 | Nv12 | P010 | Yuv444 => return None,
}) })
} }
@@ -144,6 +146,12 @@ pub struct OutputFormat {
/// (never BGRA-passthrough / P010). `false` on every non-PyroWave session and on Linux (the /// (never BGRA-passthrough / P010). `false` on every non-PyroWave session and on Linux (the
/// wavelet encoder ingests dmabufs / CPU RGB there, not a D3D11 texture). /// wavelet encoder ingests dmabufs / CPU RGB there, not a D3D11 texture).
pub pyrowave: bool, pub pyrowave: bool,
/// THIS session's encoder can ingest a producer-native NV12 capture (Linux raw Vulkan Video
/// backend on an H265/AV1 session — see `pf_encode::linux_native_nv12_ok`). The Linux capture
/// negotiation only offers gamescope the NV12 pod when this is set: libav VAAPI (the H264
/// codec's backend, and the fallback family) would misread the two-plane buffer as packed
/// RGB. Always `false` on Windows (the IDD-push capturer owns its own formats).
pub nv12_native: bool,
} }
impl OutputFormat { impl OutputFormat {
@@ -161,6 +169,11 @@ impl OutputFormat {
chroma_444: false, chroma_444: false,
// GameStream never negotiates PyroWave (native punktfunk/1 only). // GameStream never negotiates PyroWave (native punktfunk/1 only).
pyrowave: false, pyrowave: false,
// Conservative: the GameStream + spike paths don't resolve the codec here, and a
// Moonlight client may negotiate H264 (whose VAAPI backend can't ingest NV12) — so
// they never prefer the producer-native NV12 pod. The punktfunk/1 plane opts in via
// `SessionPlan::output_format()`, which knows the codec.
nv12_native: false,
} }
} }
} }
@@ -201,18 +214,26 @@ pub struct CapturedFrame {
pub cursor: Option<CursorOverlay>, pub cursor: Option<CursorOverlay>,
} }
/// A captured frame still living in a single-plane packed-RGB dmabuf (the VAAPI zero-copy path). /// A captured frame still living in a DMA-BUF. Packed RGB uses one plane. Native Linux NV12
/// (gamescope PipeWire) travels in ONE fd: Y starts at `offset`, and the interleaved UV plane
/// lives at `plane1`'s offset/stride when the producer reported them — else at the contiguous
/// fallback `offset + stride * frame_height` with the shared `stride`.
///
/// Owns a *dup* of the PipeWire buffer's fd, so the frame can travel to the encode thread and be /// Owns a *dup* of the PipeWire buffer's fd, so the frame can travel to the encode thread and be
/// imported into a VA surface there without the compositor's buffer being closed underneath it. /// imported there without the compositor's buffer being closed underneath it. Content stability
/// (Content stability across the brief import window relies on the compositor's buffer pool depth, /// across the brief import window relies on the compositor's buffer pool depth, like any zero-copy
/// same as any zero-copy capture — the VAAPI importer copies into its own NV12 surface promptly.) /// capture.
#[cfg(target_os = "linux")] #[cfg(target_os = "linux")]
pub struct DmabufFrame { pub struct DmabufFrame {
pub fd: std::os::fd::OwnedFd, pub fd: std::os::fd::OwnedFd,
/// DRM FourCC of the packed-RGB plane (e.g. `XR24` for BGRx). /// DRM FourCC (`XR24` for BGRx, `NV12` for native 4:2:0).
pub fourcc: u32, pub fourcc: u32,
/// DRM format modifier the compositor allocated (0 = LINEAR). /// DRM format modifier the compositor allocated (0 = LINEAR).
pub modifier: u64, pub modifier: u64,
/// Second-plane `(offset, stride)` within the SAME fd, when the producer reported one (the
/// PipeWire buffer's plane-1 chunk — NV12's interleaved UV). `None` falls back to the
/// contiguous-plane contract above. Always `None` for single-plane packed RGB.
pub plane1: Option<(u32, u32)>,
pub offset: u32, pub offset: u32,
pub stride: u32, pub stride: u32,
} }
@@ -225,8 +246,8 @@ pub enum FramePayload {
/// The dmabuf has already been imported + copied into this owned device buffer. /// The dmabuf has already been imported + copied into this owned device buffer.
#[cfg(target_os = "linux")] #[cfg(target_os = "linux")]
Cuda(pf_zerocopy::DeviceBuffer), Cuda(pf_zerocopy::DeviceBuffer),
/// A raw packed-RGB dmabuf — the AMD/Intel (VAAPI) zero-copy path. The encoder imports it into /// A raw DMA-BUF: packed RGB for the existing GPU CSC paths, or native NV12 from a producer
/// a VA surface and does RGB→NV12 on the GPU video engine (no host CSC, no upload). /// such as gamescope. The encoder imports it without a host copy.
#[cfg(target_os = "linux")] #[cfg(target_os = "linux")]
Dmabuf(DmabufFrame), Dmabuf(DmabufFrame),
/// A GPU-resident D3D11 texture (Windows zero-copy path for NVENC). Owns the copied frame. /// A GPU-resident D3D11 texture (Windows zero-copy path for NVENC). Owns the copied frame.
@@ -17,7 +17,7 @@
use anyhow::{bail, Context, Result}; use anyhow::{bail, Context, Result};
use std::mem::size_of; use std::mem::size_of;
use std::os::fd::RawFd; use std::os::fd::RawFd;
use std::sync::atomic::{AtomicBool, Ordering}; use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use std::thread::JoinHandle; use std::thread::JoinHandle;
@@ -196,6 +196,45 @@ impl Drop for GadgetFd {
} }
} }
/// The signal used to break a worker thread out of a blocking raw_gadget ioctl at teardown.
/// `EVENT_FETCH`/`EP_WRITE` are `wait_event_interruptible` in the kernel with no timeout and no
/// `O_NONBLOCK` honouring, and closing the fd cannot wake a thread already inside the ioctl (the
/// in-flight syscall holds a reference to the struct file). A signal is the only reliable lever:
/// delivered with a no-op, non-`SA_RESTART` handler it forces the ioctl to return `EINTR`, after
/// which the loop's top-of-iteration `running` check exits. `SIGUSR1` is unused elsewhere in this
/// process; the handler is a no-op, so a stray `SIGUSR1` becomes harmless rather than fatal.
const WAKE_SIGNAL: libc::c_int = libc::SIGUSR1;
/// Install the no-op `WAKE_SIGNAL` handler exactly once. Crucially `sa_flags = 0` (no `SA_RESTART`)
/// so a delivered signal makes the interruptible ioctl return `EINTR` instead of auto-restarting.
fn install_wake_handler() {
static ONCE: std::sync::Once = std::sync::Once::new();
ONCE.call_once(|| {
extern "C" fn noop(_: libc::c_int) {}
// SAFETY: installing a well-formed `sigaction` with an empty mask and a valid no-op handler
// for a single signal; touches only this process's disposition for `WAKE_SIGNAL`.
unsafe {
let mut sa: libc::sigaction = std::mem::zeroed();
// Via `*const ()`: casting a function item straight to an integer is what
// `clippy::function_casts_as_integer` rejects, and the pointer hop is the documented
// way to spell it. `sa_sigaction` is a `usize`-typed handler slot, so the value is
// unchanged.
sa.sa_sigaction = noop as *const () as usize;
libc::sigemptyset(&mut sa.sa_mask);
sa.sa_flags = 0;
libc::sigaction(WAKE_SIGNAL, &sa, std::ptr::null_mut());
}
});
}
/// Lets `Drop` wake a specific worker thread parked in a blocking ioctl. `tid` is the thread's
/// `pthread_self()` (0 until it starts); `done` is set right before the thread returns, so `Drop`
/// stops signalling a thread that has already exited.
struct Waker {
tid: Arc<AtomicU64>,
done: Arc<AtomicBool>,
}
/// A virtual Steam Deck presented over the USB gadget subsystem. Dropping it stops the threads and /// A virtual Steam Deck presented over the USB gadget subsystem. Dropping it stops the threads and
/// closes the gadget (the kernel tears down the device). /// closes the gadget (the kernel tears down the device).
pub struct SteamDeckGadget { pub struct SteamDeckGadget {
@@ -203,6 +242,7 @@ pub struct SteamDeckGadget {
feedback: Arc<Mutex<super::steam_proto::SteamFeedback>>, feedback: Arc<Mutex<super::steam_proto::SteamFeedback>>,
running: Arc<AtomicBool>, running: Arc<AtomicBool>,
threads: Vec<JoinHandle<()>>, threads: Vec<JoinHandle<()>>,
wakers: Vec<Waker>,
_fd: Arc<GadgetFd>, _fd: Arc<GadgetFd>,
seq: u32, seq: u32,
} }
@@ -243,6 +283,18 @@ impl SteamDeckGadget {
let ctrl_ep = Arc::new(std::sync::atomic::AtomicI32::new(-1)); let ctrl_ep = Arc::new(std::sync::atomic::AtomicI32::new(-1));
let configured = Arc::new(AtomicBool::new(false)); let configured = Arc::new(AtomicBool::new(false));
// The teardown wake path (see `WAKE_SIGNAL`) needs the handler installed before any thread
// can park in a blocking ioctl.
install_wake_handler();
let ctrl_waker = Waker {
tid: Arc::new(AtomicU64::new(0)),
done: Arc::new(AtomicBool::new(false)),
};
let stream_waker = Waker {
tid: Arc::new(AtomicU64::new(0)),
done: Arc::new(AtomicBool::new(false)),
};
// Control thread: enumerate + answer every control transfer. // Control thread: enumerate + answer every control transfer.
let control = { let control = {
let fd = fd.clone(); let fd = fd.clone();
@@ -250,10 +302,15 @@ impl SteamDeckGadget {
let ctrl_ep = ctrl_ep.clone(); let ctrl_ep = ctrl_ep.clone();
let configured = configured.clone(); let configured = configured.clone();
let feedback = feedback.clone(); let feedback = feedback.clone();
let tid = ctrl_waker.tid.clone();
let done = ctrl_waker.done.clone();
std::thread::Builder::new() std::thread::Builder::new()
.name("pf-deck-gadget-ctrl".into()) .name("pf-deck-gadget-ctrl".into())
.spawn(move || { .spawn(move || {
control_loop(fd, running, ctrl_ep, configured, feedback, serial, unit_id) // SAFETY: `pthread_self` is always valid on the calling thread.
tid.store(unsafe { libc::pthread_self() } as u64, Ordering::SeqCst);
control_loop(fd, running, ctrl_ep, configured, feedback, serial, unit_id);
done.store(true, Ordering::SeqCst);
}) })
.context("spawn gadget control thread")? .context("spawn gadget control thread")?
}; };
@@ -264,9 +321,16 @@ impl SteamDeckGadget {
let ctrl_ep = ctrl_ep.clone(); let ctrl_ep = ctrl_ep.clone();
let configured = configured.clone(); let configured = configured.clone();
let report = report.clone(); let report = report.clone();
let tid = stream_waker.tid.clone();
let done = stream_waker.done.clone();
std::thread::Builder::new() std::thread::Builder::new()
.name("pf-deck-gadget-stream".into()) .name("pf-deck-gadget-stream".into())
.spawn(move || stream_loop(fd, running, ctrl_ep, configured, report)) .spawn(move || {
// SAFETY: `pthread_self` is always valid on the calling thread.
tid.store(unsafe { libc::pthread_self() } as u64, Ordering::SeqCst);
stream_loop(fd, running, ctrl_ep, configured, report);
done.store(true, Ordering::SeqCst);
})
.context("spawn gadget stream thread")? .context("spawn gadget stream thread")?
}; };
@@ -275,6 +339,7 @@ impl SteamDeckGadget {
feedback, feedback,
running, running,
threads: vec![control, stream], threads: vec![control, stream],
wakers: vec![ctrl_waker, stream_waker],
_fd: fd, _fd: fd,
seq: 0, seq: 0,
}) })
@@ -302,6 +367,32 @@ impl SteamDeckGadget {
impl Drop for SteamDeckGadget { impl Drop for SteamDeckGadget {
fn drop(&mut self) { fn drop(&mut self) {
self.running.store(false, Ordering::SeqCst); self.running.store(false, Ordering::SeqCst);
// The control thread spends steady state parked in a blocking `EVENT_FETCH` ioctl that only
// tests `running` at the top of its loop, so clearing the flag is not enough — it must be
// signalled out of the syscall (see `WAKE_SIGNAL`). Without this the join below can hang the
// caller (the session input thread, via `PadSlots::sweep`) indefinitely. Retry until each
// thread reports done, to cover the race where the signal lands just before the thread
// re-enters the ioctl; bounded (~1 s) so a genuinely stuck thread can't wedge teardown either.
for _ in 0..200 {
let mut all_done = true;
for w in &self.wakers {
if w.done.load(Ordering::SeqCst) {
continue;
}
all_done = false;
let tid = w.tid.load(Ordering::SeqCst);
if tid != 0 {
// SAFETY: the thread is joinable and not yet joined (join runs after this loop),
// so `tid` names a live pthread; `pthread_kill` on a finished-but-unjoined thread
// is defined (returns ESRCH), never UB.
unsafe { libc::pthread_kill(tid as libc::pthread_t, WAKE_SIGNAL) };
}
}
if all_done {
break;
}
std::thread::sleep(std::time::Duration::from_millis(5));
}
for t in self.threads.drain(..) { for t in self.threads.drain(..) {
let _ = t.join(); let _ = t.join();
} }
@@ -299,13 +299,20 @@ pub(super) fn create_swdevice(p: &SwDeviceProfile) -> Result<(HSWDEVICE, Option<
let event = unsafe { CreateEventW(None, true, false, PCWSTR::null())? }; let event = unsafe { CreateEventW(None, true, false, PCWSTR::null())? };
// `result` starts as E_FAIL, NOT S_OK: if the wait below times out, a zero-initialised HRESULT // `result` starts as E_FAIL, NOT S_OK: if the wait below times out, a zero-initialised HRESULT
// would read as success and mask the failure (found by the 2026-07 driver-health audit). // would read as success and mask the failure (found by the 2026-07 driver-health audit).
let mut ctx = SwCreateCtx { // HEAP-allocated, deliberately: `sw_create_cb` writes `result` + up to 127 u16 of instance id
// through this pointer and then `SetEvent`s. The wait below is bounded (10 s), so on a wedged-PnP
// timeout the callback may still be PENDING — a stack context would be popped and a late callback
// would corrupt whatever the input thread put there next, and SetEvent a closed/recycled handle.
// On the timeout path we therefore LEAK the box and leave the event open (a one-off ~264 B + one
// HANDLE, only on that rare path) so a late callback always writes to live memory.
let ctx = Box::into_raw(Box::new(SwCreateCtx {
event, event,
result: E_FAIL, result: E_FAIL,
instance_id: [0; 128], instance_id: [0; 128],
}; }));
// SAFETY: info + the buffers + ctx outlive the call (we wait on the event before returning); // SAFETY: info + the buffers outlive the call; `ctx` is a live heap allocation that outlives every
// windows-rs returns the HSWDEVICE (the C out-param) as the Result value. // path below (reclaimed only where the callback provably ran). windows-rs returns the HSWDEVICE
// (the C out-param) as the Result value.
let hsw = match unsafe { let hsw = match unsafe {
SwDeviceCreate( SwDeviceCreate(
w!("punktfunk"), w!("punktfunk"),
@@ -313,13 +320,15 @@ pub(super) fn create_swdevice(p: &SwDeviceProfile) -> Result<(HSWDEVICE, Option<
&info, &info,
None, None,
Some(sw_create_cb), Some(sw_create_cb),
Some(&mut ctx as *mut SwCreateCtx as *const c_void), Some(ctx as *const c_void),
) )
} { } {
Ok(h) => h, Ok(h) => h,
Err(e) => { Err(e) => {
// SAFETY: event is valid. // SAFETY: the call failed, so no callback was registered and `ctx` is ours to reclaim;
// `event` is valid and unreferenced.
unsafe { unsafe {
drop(Box::from_raw(ctx));
let _ = CloseHandle(event); let _ = CloseHandle(event);
} }
return Err(anyhow!("SwDeviceCreate failed: {e}")); return Err(anyhow!("SwDeviceCreate failed: {e}"));
@@ -328,17 +337,22 @@ pub(super) fn create_swdevice(p: &SwDeviceProfile) -> Result<(HSWDEVICE, Option<
// Block until PnP finishes enumerating (the callback signals), then check its result. // Block until PnP finishes enumerating (the callback signals), then check its result.
// SAFETY: event is valid. // SAFETY: event is valid.
let wait = unsafe { WaitForSingleObject(event, 10_000) }; let wait = unsafe { WaitForSingleObject(event, 10_000) };
// SAFETY: event is valid.
unsafe {
let _ = CloseHandle(event);
}
if wait != WAIT_OBJECT_0 { if wait != WAIT_OBJECT_0 {
// Timed out: the callback may still fire. Intentionally leak `ctx` AND leave `event` open so
// its eventual write + SetEvent target live memory/handle rather than freed ones.
// SAFETY: hsw is the handle SwDeviceCreate returned. // SAFETY: hsw is the handle SwDeviceCreate returned.
unsafe { SwDeviceClose(hsw) }; unsafe { SwDeviceClose(hsw) };
return Err(anyhow!( return Err(anyhow!(
"SwDeviceCreate enumeration callback never fired (10s) — PnP may be wedged" "SwDeviceCreate enumeration callback never fired (10s) — PnP may be wedged"
)); ));
} }
// The callback ran (it is what signalled the event), so nothing else will touch `ctx`/`event`.
// SAFETY: `ctx` came from `Box::into_raw` above and is reclaimed exactly once here; `event` is
// valid and no longer referenced by a pending callback.
let ctx = unsafe {
let _ = CloseHandle(event);
Box::from_raw(ctx)
};
if ctx.result.is_err() { if ctx.result.is_err() {
// SAFETY: hsw is the handle SwDeviceCreate returned. // SAFETY: hsw is the handle SwDeviceCreate returned.
unsafe { SwDeviceClose(hsw) }; unsafe { SwDeviceClose(hsw) };
@@ -62,7 +62,7 @@ impl Ds4WinPad {
std::ptr::write_unaligned(base as *mut u32, SHM_MAGIC); std::ptr::write_unaligned(base as *mut u32, SHM_MAGIC);
} }
let inst = format!("pf_ds4_{index}"); let inst = format!("pf_ds4_{index}");
let (hsw, instance_id) = match create_swdevice(&SwDeviceProfile { let (hsw, instance_id) = create_swdevice(&SwDeviceProfile {
instance: &inst, instance: &inst,
container_tag: 0x5046_4453, // "PFDS" container_tag: 0x5046_4453, // "PFDS"
container_index: index, container_index: index,
@@ -70,13 +70,13 @@ impl Ds4WinPad {
usb_vid_pid: "VID_054C&PID_09CC", usb_vid_pid: "VID_054C&PID_09CC",
usb_mi: None, usb_mi: None,
description: "punktfunk Virtual DualShock 4", description: "punktfunk Virtual DualShock 4",
}) { })?; // Propagate, do NOT swallow — see below.
Ok((h, id)) => (Some(h), id), let (hsw, instance_id) = (Some(hsw), instance_id);
Err(e) => { // Swallowing a create failure here (the previous behaviour) latched the pad slot to
tracing::warn!(error = %format!("{e:#}"), "SwDeviceCreate failed; DualShock 4 devnode unavailable"); // `Some(pad)` with no live devnode: `PadSlots::ensure` short-circuits on `is_some()` and
(None, None) // `gate.on_success()` cleared the backoff, so the create-gate that exists precisely to
} // self-heal a transient PnP failure never retried. The game saw no controller for the whole
}; // session unless the client unplugged the pad. Matches the XUSB sibling, which propagates.
let _sw = hsw.map(super::gamepad_raii::SwDevice::new); let _sw = hsw.map(super::gamepad_raii::SwDevice::new);
// Bounded eager delivery — for the DS4 this is what closes the identity race: the driver // Bounded eager delivery — for the DS4 this is what closes the identity race: the driver
// must read `device_type = 1` from the delivered DATA section before hidclass asks it for // must read `device_type = 1` from the delivered DATA section before hidclass asks it for
@@ -82,12 +82,16 @@ fn create_swdevice(index: u8) -> Result<(HSWDEVICE, Option<String>)> {
let event = unsafe { CreateEventW(None, true, false, PCWSTR::null())? }; let event = unsafe { CreateEventW(None, true, false, PCWSTR::null())? };
// `result` starts as E_FAIL, NOT S_OK: if the wait below times out, a zero-initialised HRESULT // `result` starts as E_FAIL, NOT S_OK: if the wait below times out, a zero-initialised HRESULT
// would read as success and mask the failure (found by the 2026-07 driver-health audit). // would read as success and mask the failure (found by the 2026-07 driver-health audit).
let mut ctx = SwCreateCtx { // HEAP-allocated for the same reason as the DualSense sibling: the callback writes through this
// pointer and SetEvents, and the wait below is bounded — a stack context would be popped while a
// late callback still holds it. On the timeout path the box is deliberately leaked and the event
// left open so a late write/SetEvent always targets live memory/handle.
let ctx = Box::into_raw(Box::new(SwCreateCtx {
event, event,
result: E_FAIL, result: E_FAIL,
instance_id: [0; 128], instance_id: [0; 128],
}; }));
// SAFETY: info + buffers + ctx outlive the call (we wait on the event before returning). // SAFETY: info + buffers outlive the call; `ctx` is a live heap allocation outliving every path.
let hsw = match unsafe { let hsw = match unsafe {
SwDeviceCreate( SwDeviceCreate(
w!("punktfunk"), w!("punktfunk"),
@@ -95,13 +99,14 @@ fn create_swdevice(index: u8) -> Result<(HSWDEVICE, Option<String>)> {
&info, &info,
None, None,
Some(sw_create_cb), Some(sw_create_cb),
Some(&mut ctx as *mut SwCreateCtx as *const c_void), Some(ctx as *const c_void),
) )
} { } {
Ok(h) => h, Ok(h) => h,
Err(e) => { Err(e) => {
// SAFETY: event is valid. // SAFETY: the call failed, so no callback is pending and `ctx` is ours to reclaim.
unsafe { unsafe {
drop(Box::from_raw(ctx));
let _ = CloseHandle(event); let _ = CloseHandle(event);
} }
return Err(anyhow!("SwDeviceCreate(pf_xusb) failed: {e}")); return Err(anyhow!("SwDeviceCreate(pf_xusb) failed: {e}"));
@@ -109,17 +114,20 @@ fn create_swdevice(index: u8) -> Result<(HSWDEVICE, Option<String>)> {
}; };
// SAFETY: event valid; block until PnP finishes enumerating, then check the callback result. // SAFETY: event valid; block until PnP finishes enumerating, then check the callback result.
let wait = unsafe { WaitForSingleObject(event, 10_000) }; let wait = unsafe { WaitForSingleObject(event, 10_000) };
// SAFETY: event is valid.
unsafe {
let _ = CloseHandle(event);
}
if wait != WAIT_OBJECT_0 { if wait != WAIT_OBJECT_0 {
// Timed out — intentionally leak `ctx` and leave `event` open (see above).
// SAFETY: hsw is the handle SwDeviceCreate returned. // SAFETY: hsw is the handle SwDeviceCreate returned.
unsafe { SwDeviceClose(hsw) }; unsafe { SwDeviceClose(hsw) };
return Err(anyhow!( return Err(anyhow!(
"SwDeviceCreate(pf_xusb) enumeration callback never fired (10s) — PnP may be wedged" "SwDeviceCreate(pf_xusb) enumeration callback never fired (10s) — PnP may be wedged"
)); ));
} }
// The callback ran (it signalled the event), so nothing else will touch `ctx`/`event`.
// SAFETY: `ctx` came from `Box::into_raw` and is reclaimed exactly once here.
let ctx = unsafe {
let _ = CloseHandle(event);
Box::from_raw(ctx)
};
if ctx.result.is_err() { if ctx.result.is_err() {
// SAFETY: hsw is the handle SwDeviceCreate returned. // SAFETY: hsw is the handle SwDeviceCreate returned.
unsafe { SwDeviceClose(hsw) }; unsafe { SwDeviceClose(hsw) };
@@ -66,7 +66,7 @@ impl DeckWinPad {
std::ptr::write_unaligned(base as *mut u32, SHM_MAGIC); std::ptr::write_unaligned(base as *mut u32, SHM_MAGIC);
} }
let inst = format!("pf_deck_{index}"); let inst = format!("pf_deck_{index}");
let (hsw, instance_id) = match create_swdevice(&SwDeviceProfile { let (hsw, instance_id) = create_swdevice(&SwDeviceProfile {
instance: &inst, instance: &inst,
container_tag: 0x5046_4453, // "PFDS" container_tag: 0x5046_4453, // "PFDS"
container_index: index, container_index: index,
@@ -77,13 +77,8 @@ impl DeckWinPad {
// spike's run-1 failure). // spike's run-1 failure).
usb_mi: Some(2), usb_mi: Some(2),
description: "punktfunk Virtual Steam Deck", description: "punktfunk Virtual Steam Deck",
}) { })?; // Propagate — swallowing latched the slot to a pad with no devnode (see the DS4 twin).
Ok((h, i)) => (Some(h), i), let (hsw, instance_id) = (Some(hsw), instance_id);
Err(e) => {
tracing::warn!(error = %format!("{e:#}"), "SwDeviceCreate failed; Steam Deck devnode unavailable");
(None, None)
}
};
let _sw = hsw.map(super::gamepad_raii::SwDevice::new); let _sw = hsw.map(super::gamepad_raii::SwDevice::new);
// Bounded eager delivery — the driver must read `device_type = 3` before hidclass asks // Bounded eager delivery — the driver must read `device_type = 3` before hidclass asks
// it for descriptors, or the pad would enumerate with the default DualSense identity. // it for descriptors, or the pad would enumerate with the default DualSense identity.
+5 -1
View File
@@ -11,7 +11,11 @@ repository.workspace = true
# Same Linux+Windows gating as the rest of the client stack (dmabuf import is the one # Same Linux+Windows gating as the rest of the client stack (dmabuf import is the one
# Linux-only module — see lib.rs). # Linux-only module — see lib.rs).
[target.'cfg(any(target_os = "linux", windows))'.dependencies] [target.'cfg(any(target_os = "linux", windows))'.dependencies]
pf-client-core = { path = "../pf-client-core" } # `default-features = false`: the PyroWave decode backend is turned on through THIS crate's own
# `pyrowave` feature (which re-exports it below), never by inheriting the dependency's default.
# Otherwise a consumer that deliberately builds us without `pyrowave` still drags the vendored
# C++ in — fatal on Windows ARM64, where Granite has no SIMD path.
pf-client-core = { path = "../pf-client-core", default-features = false }
# AVVkFrame access (Vulkan Video frames: live sync state under the frames lock). # AVVkFrame access (Vulkan Video frames: live sync state under the frames lock).
pf-ffvk = { path = "../pf-ffvk" } pf-ffvk = { path = "../pf-ffvk" }
punktfunk-core = { path = "../punktfunk-core", features = ["quic"] } punktfunk-core = { path = "../punktfunk-core", features = ["quic"] }
+5
View File
@@ -333,6 +333,11 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
// bottom-right corner (the reported bug). The menu/library is keyboard+gamepad-driven // bottom-right corner (the reported bug). The menu/library is keyboard+gamepad-driven
// and consumes no mouse, so nothing wanted these synthetic events anyway. // and consumes no mouse, so nothing wanted these synthetic events anyway.
sdl3::hint::set("SDL_TOUCH_MOUSE_EVENTS", "0"); sdl3::hint::set("SDL_TOUCH_MOUSE_EVENTS", "0");
// The Wayland `app_id` (and X11 WM_CLASS) — compositors match it against
// io.unom.Punktfunk.desktop for the window/taskbar icon. Without it SDL uses a generic
// identity and the session window gets the default-Wayland icon (the Linux analog of
// the AppUserModelID adoption above).
sdl3::hint::set("SDL_APP_ID", "io.unom.Punktfunk");
let sdl = sdl3::init().context("SDL init")?; let sdl = sdl3::init().context("SDL init")?;
let video = sdl.video().context("SDL video")?; let video = sdl.video().context("SDL video")?;
let events = sdl.event().context("SDL events")?; let events = sdl.event().context("SDL events")?;
+9 -1
View File
@@ -30,9 +30,17 @@ impl Presenter {
// switch modes before anything touches this frame. Only where the surface // switch modes before anything touches this frame. Only where the surface
// offers HDR10 — otherwise PQ stays on the SDR swapchain and the CSC shader // offers HDR10 — otherwise PQ stays on the SDR swapchain and the CSC shader
// tonemaps (mode 1). // tonemaps (mode 1).
//
// CPU frames NEVER take the HDR10 surface: software decode uploads swscale RGBA with
// no CSC/tonemap pass, so on a mode-0 swapchain that sRGB-encoded content would be
// composed as PQ — the field-reported psychedelic cyan/magenta picture (reproduced
// 2026-07-21: Fedora-class client, no hw HEVC decode, GNOME/Mesa offering HDR10 even
// on an SDR desktop). On the SDR swapchain the same frames are merely untonemapped
// (washed out) — wrong in the known, benign way until the CPU lane grows a real
// PQ→sRGB pass.
let frame_pq = match &input { let frame_pq = match &input {
FrameInput::Redraw => None, FrameInput::Redraw => None,
FrameInput::Cpu(f) => Some(f.color.is_pq()), FrameInput::Cpu(_) => Some(false),
#[cfg(target_os = "linux")] #[cfg(target_os = "linux")]
FrameInput::Dmabuf(d) => Some(d.color.is_pq()), FrameInput::Dmabuf(d) => Some(d.color.is_pq()),
FrameInput::VkFrame(v) => Some(v.color.is_pq()), FrameInput::VkFrame(v) => Some(v.color.is_pq()),
+9
View File
@@ -617,6 +617,15 @@ pub(super) fn pick_formats(
surface: vk::SurfaceKHR, surface: vk::SurfaceKHR,
colorspace_ext: bool, colorspace_ext: bool,
) -> Result<(vk::SurfaceFormatKHR, Option<vk::SurfaceFormatKHR>)> { ) -> Result<(vk::SurfaceFormatKHR, Option<vk::SurfaceFormatKHR>)> {
// `PUNKTFUNK_HDR10=0` (explicit-off grammar) refuses the HDR10/ST.2084 swapchain outright,
// pinning PQ streams to the shader tonemap on an SDR surface. Two reasons this exists:
// desktop compositors newly offer HDR10 even on SDR desktops (GNOME 48 / Plasma 6 with
// Mesa ≥ 25.1 — a lane that otherwise engages silently), and it is the A/B lever that
// splits "HDR10 passthrough composes wrong" from "the decoded planes are wrong" in the
// field without rebuilding anything.
let colorspace_ext = colorspace_ext
&& !std::env::var("PUNKTFUNK_HDR10")
.is_ok_and(|v| matches!(v.as_str(), "0" | "false" | "off" | "no"));
let formats = unsafe { surface_i.get_physical_device_surface_formats(pdev, surface) }?; let formats = unsafe { surface_i.get_physical_device_surface_formats(pdev, surface) }?;
let mut sdr = None; let mut sdr = None;
for want in [vk::Format::B8G8R8A8_UNORM, vk::Format::R8G8B8A8_UNORM] { for want in [vk::Format::B8G8R8A8_UNORM, vk::Format::R8G8B8A8_UNORM] {
+28
View File
@@ -254,6 +254,34 @@ pub fn detect() -> Result<Compositor> {
} }
} }
/// Attach-only probes: while any scope is held, backend `create` paths must not stop, relaunch,
/// or take over box sessions — they may only attach to an already-live output, and fail fast
/// otherwise. The capture-loss rebuild holds one for its first seconds: right after a capture
/// loss the active-session detection can be STALE (a Game→Desktop switch observed live: the
/// probe's gamescope re-acquire restarted `gamescope-session.target` and yanked the user out of
/// the KDE session they had just switched to). A counter, so overlapping scopes compose.
static REBUILD_PROBES: std::sync::atomic::AtomicU32 = std::sync::atomic::AtomicU32::new(0);
/// RAII scope marking pipeline builds as attach-only probes (see [`rebuild_probe_active`]).
pub struct RebuildProbeScope(());
pub fn rebuild_probe_scope() -> RebuildProbeScope {
REBUILD_PROBES.fetch_add(1, std::sync::atomic::Ordering::SeqCst);
RebuildProbeScope(())
}
impl Drop for RebuildProbeScope {
fn drop(&mut self) {
REBUILD_PROBES.fetch_sub(1, std::sync::atomic::Ordering::SeqCst);
}
}
/// Is any [`rebuild_probe_scope`] active? Destructive session operations (stop/relaunch/
/// takeover-restart) must be skipped while true.
pub fn rebuild_probe_active() -> bool {
REBUILD_PROBES.load(std::sync::atomic::Ordering::SeqCst) > 0
}
/// Open the virtual-display driver for `compositor`. /// Open the virtual-display driver for `compositor`.
pub fn open(compositor: Compositor) -> Result<Box<dyn VirtualDisplay>> { pub fn open(compositor: Compositor) -> Result<Box<dyn VirtualDisplay>> {
#[cfg(target_os = "linux")] #[cfg(target_os = "linux")]
@@ -325,6 +325,29 @@ fn create_managed_session(client: &str, mode: Mode) -> Result<VirtualOutput> {
if steamos_session_present() { if steamos_session_present() {
return create_managed_session_steamos(mode); return create_managed_session_steamos(mode);
} }
// Attach-only rebuild probe: reuse a live same-mode session, but NEVER stop/relaunch box
// sessions — right after a capture loss the caller's session detection can be stale, and a
// destructive rebuild here would fight the session the user just switched to.
if crate::rebuild_probe_active() {
let guard = MANAGED_SESSION.lock().unwrap_or_else(|e| e.into_inner());
let same_mode = guard.as_ref().is_some_and(|s| {
s.width == mode.width && s.height == mode.height && s.refresh_hz == mode.refresh_hz
});
if same_mode {
if let Some(node_id) = find_gamescope_node() {
point_injector_at_eis();
tracing::info!(
node_id,
"gamescope session: attach-only probe reusing live node"
);
return Ok(managed_output(node_id, mode));
}
}
return Err(anyhow!(
"gamescope session has no attachable live node — attach-only rebuild probe refuses \
to stop/relaunch box sessions (re-detection follows the live session)"
));
}
// Steam is single-instance: if the box autologged into gaming mode on a physical display (the // Steam is single-instance: if the box autologged into gaming mode on a physical display (the
// Bazzite default — `gamescope-session-plus@ogui-steam` on the TV), that session holds Steam and // Bazzite default — `gamescope-session-plus@ogui-steam` on the TV), that session holds Steam and
// renders to the TV's native mode, which we'd capture instead of the client's. Free Steam by // renders to the TV's native mode, which we'd capture instead of the client's. Free Steam by
@@ -607,12 +630,17 @@ fn write_steamos_dropin(shim_dir: &std::path::Path, mode: Mode) -> Result<()> {
if let Some(parent) = path.parent() { if let Some(parent) = path.parent() {
std::fs::create_dir_all(parent).with_context(|| format!("mkdir {}", parent.display()))?; std::fs::create_dir_all(parent).with_context(|| format!("mkdir {}", parent.display()))?;
} }
// UnsetEnvironment: the same headless-must-not-attach armor `launch_session` gives its
// transient unit — the manager env can carry a stale desktop DISPLAY/WAYLAND_DISPLAY (from a
// portal settle), and gamescope would abort trying to attach to it instead of becoming the
// display server. Unit-scoped belt-and-suspenders on top of the observe_session_instance scrub.
let body = format!( let body = format!(
"[Service]\n\ "[Service]\n\
Environment=PATH={shim}:/usr/bin:/bin:/usr/local/bin\n\ Environment=PATH={shim}:/usr/bin:/bin:/usr/local/bin\n\
Environment=PF_W={w}\n\ Environment=PF_W={w}\n\
Environment=PF_H={h}\n\ Environment=PF_H={h}\n\
Environment=PF_HZ={hz}\n", Environment=PF_HZ={hz}\n\
UnsetEnvironment=DISPLAY WAYLAND_DISPLAY\n",
shim = shim_dir.display(), shim = shim_dir.display(),
w = mode.width, w = mode.width,
h = mode.height, h = mode.height,
@@ -650,6 +678,16 @@ fn create_managed_session_steamos(mode: Mode) -> Result<VirtualOutput> {
} }
*guard = None; // tracked session lost its node — fall through to a clean restart *guard = None; // tracked session lost its node — fall through to a clean restart
} }
// Attach-only rebuild probe: the reuse path above may attach, but a restart of the session
// target is out of bounds — observed live on a Deck: a stale post-capture-loss detection made
// this restart steal the seat back from the KDE session the user had just switched to.
if crate::rebuild_probe_active() {
return Err(anyhow!(
"gamescope has no live node and this is an attach-only rebuild probe — refusing to \
restart {STEAMOS_SESSION_TARGET} (the box may be mid-switch to another session; \
re-detection follows it)"
));
}
let shim_dir = write_headless_shim()?; let shim_dir = write_headless_shim()?;
write_steamos_dropin(&shim_dir, mode)?; write_steamos_dropin(&shim_dir, mode)?;
systemctl_user(&["daemon-reload"]); systemctl_user(&["daemon-reload"]);
@@ -57,6 +57,13 @@ pub fn observe_session_instance(active: &ActiveSession) {
if let Some(old) = compositor_for_kind(prev.0) { if let Some(old) = compositor_for_kind(prev.0) {
registry::invalidate_backend(old.id()); registry::invalidate_backend(old.id());
} }
// The dead desktop's socket vars may still sit in the systemd --user manager env
// ([`settle_desktop_portal`]'s import-environment) — scrub them NOW, or the next
// `gamescope-session.target` start inherits a stale WAYLAND_DISPLAY and gamescope
// runs NESTED against the dead desktop socket instead of becoming the display
// server ("Failed to connect to wayland socket: wayland-0" — kept a Deck's Game
// Mode from starting at all, observed live 2026-07-21).
scrub_desktop_manager_env();
} }
let epoch = bump_session_epoch(); let epoch = bump_session_epoch();
tracing::info!( tracing::info!(
@@ -70,6 +77,23 @@ pub fn observe_session_instance(active: &ActiveSession) {
*last = Some(cur); *last = Some(cur);
} }
/// Counterpart to [`settle_desktop_portal`]'s `import-environment`: drop the desktop session's
/// socket vars from the systemd `--user` manager env once that desktop instance is GONE. They
/// persist in the manager otherwise, and every later user unit inherits them — including
/// `gamescope-session.target`, whose gamescope then aborts trying to attach to the dead desktop
/// socket. Best-effort; the D-Bus activation env has no unset op, but gamescope-session is
/// systemd-started, so the manager scrub is the one that matters. (A desktop restart re-imports
/// via the next [`settle_desktop_portal`], so scrubbing on a bounce is harmless.)
#[cfg(target_os = "linux")]
fn scrub_desktop_manager_env() {
let _ = std::process::Command::new("systemctl")
.args(["--user", "unset-environment", "WAYLAND_DISPLAY", "DISPLAY"])
.status();
}
#[cfg(not(target_os = "linux"))]
fn scrub_desktop_manager_env() {}
/// Is `kind` a **desktop** compositor (KWin / Mutter / wlroots) — one whose kept PipeWire outputs die /// Is `kind` a **desktop** compositor (KWin / Mutter / wlroots) — one whose kept PipeWire outputs die
/// with the compositor instance, so the session epoch tracks it? `Gaming` (gamescope) and `None` are /// with the compositor instance, so the session epoch tracks it? `Gaming` (gamescope) and `None` are
/// not (gamescope spawns are independent nested sessions — see [`observe_session_instance`]). /// not (gamescope spawns are independent nested sessions — see [`observe_session_instance`]).
+20 -4
View File
@@ -21,7 +21,7 @@ use std::path::{Path, PathBuf};
use std::process::{Child, Command}; use std::process::{Child, Command};
use std::sync::atomic::{AtomicBool, Ordering}; use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{Arc, Mutex, OnceLock}; use std::sync::{Arc, Mutex, OnceLock};
use std::time::Duration; use std::time::{Duration, Instant};
/// Handshake budget: EGL + CUDA bring-up is ~200 ms; a cold driver load can take seconds. /// Handshake budget: EGL + CUDA bring-up is ~200 ms; a cold driver load can take seconds.
const HANDSHAKE_TIMEOUT: Duration = Duration::from_secs(20); const HANDSHAKE_TIMEOUT: Duration = Duration::from_secs(20);
@@ -64,11 +64,27 @@ impl Drop for Shared {
/// Children whose worker hasn't exited yet at `RemoteImporter` drop time (it exits on socket /// Children whose worker hasn't exited yet at `RemoteImporter` drop time (it exits on socket
/// EOF, i.e. after the last in-flight frame drops). Swept on every spawn and every drop so /// EOF, i.e. after the last in-flight frame drops). Swept on every spawn and every drop so
/// workers don't linger as zombies for more than one capture generation. /// workers don't linger as zombies for more than one capture generation.
static REAPER: Mutex<Vec<Child>> = Mutex::new(Vec::new()); static REAPER: Mutex<Vec<(Child, Instant)>> = Mutex::new(Vec::new());
/// How long past `REPLY_TIMEOUT` a parked worker may linger before it is force-killed. A worker
/// wedged INSIDE a driver call never observes socket EOF, so `try_wait` alone would keep it (and
/// its CUcontext + BufferPool — order hundreds of MB of VRAM) forever.
const REAPER_KILL_DEADLINE: Duration = Duration::from_secs(20);
fn sweep_reaper() { fn sweep_reaper() {
let mut list = REAPER.lock().unwrap(); let mut list = REAPER.lock().unwrap();
list.retain_mut(|c| !matches!(c.try_wait(), Ok(Some(_)))); let now = Instant::now();
list.retain_mut(|(c, parked)| {
if matches!(c.try_wait(), Ok(Some(_))) {
return false; // exited on its own → reaped
}
if now.duration_since(*parked) > REAPER_KILL_DEADLINE {
let _ = c.kill();
let _ = c.wait();
return false; // wedged past the deadline → force-killed + reaped
}
true
});
} }
/// Fd pinned to this process's own executable image, opened (once, lazily) via the /// Fd pinned to this process's own executable image, opened (once, lazily) via the
@@ -455,7 +471,7 @@ impl Drop for RemoteImporter {
// gone; park the rest for the next sweep. // gone; park the rest for the next sweep.
if let Some(mut child) = self.child.take() { if let Some(mut child) = self.child.take() {
if !matches!(child.try_wait(), Ok(Some(_))) { if !matches!(child.try_wait(), Ok(Some(_))) {
REAPER.lock().unwrap().push(child); REAPER.lock().unwrap().push((child, Instant::now()));
} }
} }
sweep_reaper(); sweep_reaper();
+87 -28
View File
@@ -251,6 +251,32 @@ unsafe fn copy_blocking(copy: &CUDA_MEMCPY2D, what: &str) -> Result<()> {
ck(cuStreamSynchronize(stream), "cuStreamSynchronize") ck(cuStreamSynchronize(stream), "cuStreamSynchronize")
} }
/// Issue `copy` on this thread's priority stream WITHOUT waiting — for stream-ordered consumers
/// only (the direct-NVENC submit path with `NvEncSetIOCudaStreams` bound to this stream): the
/// stream, not the CPU, orders completion, so the SOURCE must stay valid until the downstream
/// stream work (the encode) has finished.
unsafe fn copy_async(copy: &CUDA_MEMCPY2D, what: &str) -> Result<()> {
ck(cuMemcpy2DAsync_v2(copy, copy_stream()), what)
}
/// `copy_blocking` when `sync`, else `copy_async` — the shared tail of the public `copy_*_to_device`
/// helpers, whose `sync: false` mode carries `copy_async`'s source-lifetime contract.
unsafe fn copy_issue(copy: &CUDA_MEMCPY2D, what: &str, sync: bool) -> Result<()> {
if sync {
copy_blocking(copy, what)
} else {
copy_async(copy, what)
}
}
/// The calling thread's copy/launch stream as a raw handle, for binding external stream-ordering
/// (the direct-NVENC `NvEncSetIOCudaStreams` hookup). Null = the NULL stream (priority-stream
/// creation failed) — callers should treat null as "stream-ordering unavailable" and keep their
/// blocking copies. The shared context must be current on this thread.
pub fn copy_stream_handle() -> *mut c_void {
copy_stream() // CUstream IS *mut c_void (opaque CUstream_st*)
}
/// Max cursor-overlay bitmap edge (px) uploaded to the device blend buffer — matches the Vulkan path. /// Max cursor-overlay bitmap edge (px) uploaded to the device blend buffer — matches the Vulkan path.
pub const CURSOR_MAX: u32 = 256; pub const CURSOR_MAX: u32 = 256;
@@ -354,6 +380,7 @@ impl CursorBlend {
ch: u32, ch: u32,
ox: i32, ox: i32,
oy: i32, oy: i32,
sync: bool,
) -> Result<()> { ) -> Result<()> {
let (mut a_surf, mut a_cur) = (surf, self.cur_buf); let (mut a_surf, mut a_cur) = (surf, self.cur_buf);
let (mut a_pitch, mut a_w, mut a_h) = (pitch as i32, w as i32, h as i32); let (mut a_pitch, mut a_w, mut a_h) = (pitch as i32, w as i32, h as i32);
@@ -370,7 +397,7 @@ impl CursorBlend {
&mut a_ox as *mut _ as *mut c_void, &mut a_ox as *mut _ as *mut c_void,
&mut a_oy as *mut _ as *mut c_void, &mut a_oy as *mut _ as *mut c_void,
]; ];
self.launch(self.f_argb, a_cw as u32, a_ch as u32, &mut args) self.launch(self.f_argb, a_cw as u32, a_ch as u32, &mut args, sync)
} }
/// Blend into an owned planar YUV444 surface (3 stacked full-res planes) at `(ox,oy)`. /// Blend into an owned planar YUV444 surface (3 stacked full-res planes) at `(ox,oy)`.
@@ -385,6 +412,7 @@ impl CursorBlend {
ch: u32, ch: u32,
ox: i32, ox: i32,
oy: i32, oy: i32,
sync: bool,
) -> Result<()> { ) -> Result<()> {
let (mut a_base, mut a_cur) = (base, self.cur_buf); let (mut a_base, mut a_cur) = (base, self.cur_buf);
let (mut a_pitch, mut a_w, mut a_h) = (pitch as i32, w as i32, h as i32); let (mut a_pitch, mut a_w, mut a_h) = (pitch as i32, w as i32, h as i32);
@@ -401,7 +429,7 @@ impl CursorBlend {
&mut a_ox as *mut _ as *mut c_void, &mut a_ox as *mut _ as *mut c_void,
&mut a_oy as *mut _ as *mut c_void, &mut a_oy as *mut _ as *mut c_void,
]; ];
self.launch(self.f_yuv444, a_cw as u32, a_ch as u32, &mut args) self.launch(self.f_yuv444, a_cw as u32, a_ch as u32, &mut args, sync)
} }
/// Blend into an owned NV12 surface (Y plane at `base`, interleaved UV at `base + pitch*h`). /// Blend into an owned NV12 surface (Y plane at `base`, interleaved UV at `base + pitch*h`).
@@ -416,6 +444,7 @@ impl CursorBlend {
ch: u32, ch: u32,
ox: i32, ox: i32,
oy: i32, oy: i32,
sync: bool,
) -> Result<()> { ) -> Result<()> {
let (mut a_yb, mut a_uvb, mut a_cur) = (base, base + pitch as u64 * h as u64, self.cur_buf); let (mut a_yb, mut a_uvb, mut a_cur) = (base, base + pitch as u64 * h as u64, self.cur_buf);
let (mut a_yp, mut a_uvp) = (pitch as i32, pitch as i32); let (mut a_yp, mut a_uvp) = (pitch as i32, pitch as i32);
@@ -441,16 +470,20 @@ impl CursorBlend {
(a_cw as u32).div_ceil(2), (a_cw as u32).div_ceil(2),
(a_ch as u32).div_ceil(2), (a_ch as u32).div_ceil(2),
&mut args, &mut args,
sync,
) )
} }
/// Launch `f` over a `work_w × work_h` grid (16×16 blocks) on the copy stream, then synchronize. /// Launch `f` over a `work_w × work_h` grid (16×16 blocks) on the copy stream; `sync` waits
/// for it, `!sync` leaves completion to the stream (stream-ordered consumers only — the
/// kernel PARAMETERS are copied at launch time, so the arg locals need not outlive the call).
fn launch( fn launch(
&self, &self,
f: CUfunction, f: CUfunction,
work_w: u32, work_w: u32,
work_h: u32, work_h: u32,
args: &mut [*mut c_void], args: &mut [*mut c_void],
sync: bool,
) -> Result<()> { ) -> Result<()> {
if work_w == 0 || work_h == 0 { if work_w == 0 || work_h == 0 {
return Ok(()); return Ok(());
@@ -458,9 +491,11 @@ impl CursorBlend {
const B: u32 = 16; const B: u32 = 16;
let stream = copy_stream(); let stream = copy_stream();
// SAFETY: `f` is a resolved kernel from our loaded module; `args` holds pointers to live // SAFETY: `f` is a resolved kernel from our loaded module; `args` holds pointers to live
// locals whose types match the kernel's C parameters (per the call site above); grid/block // locals whose types match the kernel's C parameters (per the call site above) — CUDA
// dims are non-zero. Launched on the copy stream (ordered after the input-surface copy that // copies the parameter values during `cuLaunchKernel` itself, so they need not outlive
// `copy_into_slot` already synchronized) then synchronized. Requires the context current. // the call. Grid/block dims are non-zero. Launched on the copy stream (ordered after the
// input-surface copy issued on the same stream); `sync` waits, `!sync` leaves ordering to
// the stream (the NVENC IO-stream binding). Requires the context current.
unsafe { unsafe {
ck( ck(
cuLaunchKernel( cuLaunchKernel(
@@ -478,7 +513,10 @@ impl CursorBlend {
), ),
"cuLaunchKernel(cursor)", "cuLaunchKernel(cursor)",
)?; )?;
ck(cuStreamSynchronize(stream), "cuStreamSynchronize(cursor)") if sync {
ck(cuStreamSynchronize(stream), "cuStreamSynchronize(cursor)")?;
}
Ok(())
} }
} }
} }
@@ -1003,7 +1041,13 @@ impl RegisteredTexture {
// SAFETY: `self.resource` is the valid `CUgraphicsResource` from a successful `register_gl` // SAFETY: `self.resource` is the valid `CUgraphicsResource` from a successful `register_gl`
// (its only constructor), so the wrappers forward to the live table; the caller holds the // (its only constructor), so the wrappers forward to the live table; the caller holds the
// GL+CUDA contexts current (the registration's contract). `cuGraphicsMapResources` maps // GL+CUDA contexts current (the registration's contract). `cuGraphicsMapResources` maps
// `count == 1` resource via `&mut self.resource` (a live field) on the default stream; // `count == 1` resource via `&mut self.resource` (a live field). It is issued on
// `copy_stream()` — NOT the NULL stream — because map's only ordering guarantee is that
// prior GL work completes before subsequent CUDA work issued IN THE STREAM PASSED TO IT;
// the copy below runs on `copy_stream()` (a `CU_STREAM_NON_BLOCKING` stream, exempt from
// implicit NULL-stream ordering), so mapping on NULL left the copy free to race the GL
// de-tile/CSC that produced this texture (glFlush only, no fence) — intermittent torn or
// stale frames under GPU load. Map, copy, and unmap now all share `copy_stream()`.
// `cuGraphicsSubResourceGetMappedArray` writes the mapped `CUarray` into the live local // `cuGraphicsSubResourceGetMappedArray` writes the mapped `CUarray` into the live local
// `array` (index 0, mip 0). On failure we unmap and bail (balanced). `&copy` is a live // `array` (index 0, mip 0). On failure we unmap and bail (balanced). `&copy` is a live
// local `CUDA_MEMCPY2D` outliving the synchronous `copy_blocking`: `srcArray` is valid // local `CUDA_MEMCPY2D` outliving the synchronous `copy_blocking`: `srcArray` is valid
@@ -1012,12 +1056,12 @@ impl RegisteredTexture {
// we always unmap afterward (even on error), keeping the map/unmap pair balanced. // we always unmap afterward (even on error), keeping the map/unmap pair balanced.
unsafe { unsafe {
ck( ck(
cuGraphicsMapResources(1, &mut self.resource, std::ptr::null_mut()), cuGraphicsMapResources(1, &mut self.resource, copy_stream()),
"cuGraphicsMapResources", "cuGraphicsMapResources",
)?; )?;
let mut array: CUarray = std::ptr::null_mut(); let mut array: CUarray = std::ptr::null_mut();
if cuGraphicsSubResourceGetMappedArray(&mut array, self.resource, 0, 0) != 0 { if cuGraphicsSubResourceGetMappedArray(&mut array, self.resource, 0, 0) != 0 {
let _ = cuGraphicsUnmapResources(1, &mut self.resource, std::ptr::null_mut()); let _ = cuGraphicsUnmapResources(1, &mut self.resource, copy_stream());
bail!("cuGraphicsSubResourceGetMappedArray failed"); bail!("cuGraphicsSubResourceGetMappedArray failed");
} }
let copy = CUDA_MEMCPY2D { let copy = CUDA_MEMCPY2D {
@@ -1031,7 +1075,7 @@ impl RegisteredTexture {
..Default::default() ..Default::default()
}; };
let res = copy_blocking(&copy, "cuMemcpy2DAsync_v2"); let res = copy_blocking(&copy, "cuMemcpy2DAsync_v2");
let _ = cuGraphicsUnmapResources(1, &mut self.resource, std::ptr::null_mut()); let _ = cuGraphicsUnmapResources(1, &mut self.resource, copy_stream());
res res
} }
} }
@@ -1058,12 +1102,12 @@ impl RegisteredTexture {
// so the map/unmap pair stays balanced and the array outlives the copy. // so the map/unmap pair stays balanced and the array outlives the copy.
unsafe { unsafe {
ck( ck(
cuGraphicsMapResources(1, &mut self.resource, std::ptr::null_mut()), cuGraphicsMapResources(1, &mut self.resource, copy_stream()),
"cuGraphicsMapResources", "cuGraphicsMapResources",
)?; )?;
let mut array: CUarray = std::ptr::null_mut(); let mut array: CUarray = std::ptr::null_mut();
if cuGraphicsSubResourceGetMappedArray(&mut array, self.resource, 0, 0) != 0 { if cuGraphicsSubResourceGetMappedArray(&mut array, self.resource, 0, 0) != 0 {
let _ = cuGraphicsUnmapResources(1, &mut self.resource, std::ptr::null_mut()); let _ = cuGraphicsUnmapResources(1, &mut self.resource, copy_stream());
bail!("cuGraphicsSubResourceGetMappedArray failed"); bail!("cuGraphicsSubResourceGetMappedArray failed");
} }
let copy = CUDA_MEMCPY2D { let copy = CUDA_MEMCPY2D {
@@ -1077,7 +1121,7 @@ impl RegisteredTexture {
..Default::default() ..Default::default()
}; };
let res = copy_blocking(&copy, "cuMemcpy2DAsync_v2(plane)"); let res = copy_blocking(&copy, "cuMemcpy2DAsync_v2(plane)");
let _ = cuGraphicsUnmapResources(1, &mut self.resource, std::ptr::null_mut()); let _ = cuGraphicsUnmapResources(1, &mut self.resource, copy_stream());
res res
} }
} }
@@ -1122,10 +1166,13 @@ pub fn copy_mapped_yuv444(
/// Copy a pitched device buffer into another device region (device→device), e.g. our imported /// Copy a pitched device buffer into another device region (device→device), e.g. our imported
/// [`DeviceBuffer`] into a pooled CUDA surface NVENC owns. Both are 4-byte (BGRx) pixels. /// [`DeviceBuffer`] into a pooled CUDA surface NVENC owns. Both are 4-byte (BGRx) pixels.
/// The caller must have the shared context current on this thread (see [`make_current`]). /// The caller must have the shared context current on this thread (see [`make_current`]).
/// `sync: false` enqueues without a CPU wait (stream-ordered consumers only — `src` must stay
/// valid until the downstream stream work completes; see [`copy_stream_handle`]).
pub fn copy_device_to_device( pub fn copy_device_to_device(
src: &DeviceBuffer, src: &DeviceBuffer,
dst_ptr: CUdeviceptr, dst_ptr: CUdeviceptr,
dst_pitch: usize, dst_pitch: usize,
sync: bool,
) -> Result<()> { ) -> Result<()> {
let copy = CUDA_MEMCPY2D { let copy = CUDA_MEMCPY2D {
srcMemoryType: CU_MEMORYTYPE_DEVICE, srcMemoryType: CU_MEMORYTYPE_DEVICE,
@@ -1138,23 +1185,27 @@ pub fn copy_device_to_device(
Height: src.height as usize, Height: src.height as usize,
..Default::default() ..Default::default()
}; };
// SAFETY: `copy_blocking` is unsafe (issues a CUDA copy); the caller must have the shared // SAFETY: `copy_issue` is unsafe (issues a CUDA copy); the caller must have the shared
// context current (documented). `&copy` is a live local device→device `CUDA_MEMCPY2D` outliving // context current (documented). `&copy` is a live local device→device `CUDA_MEMCPY2D` outliving
// the synchronous call: `srcDevice`/`srcPitch` are `src`'s live allocation, `dstDevice`/ // the enqueue: `srcDevice`/`srcPitch` are `src`'s live allocation, `dstDevice`/`dstPitch` the
// `dstPitch` the caller's live region, `width*4`×`height` within both. Wrapper → live table. // caller's live region, `width*4`×`height` within both; `sync: false` shifts the source-
unsafe { copy_blocking(&copy, "cuMemcpy2DAsync_v2(dev->dev)") } // lifetime obligation to the caller (documented above). Wrapper → live table.
unsafe { copy_issue(&copy, "cuMemcpy2DAsync_v2(dev->dev)", sync) }
} }
/// Copy our imported NV12 [`DeviceBuffer`] (Y + UV planes) into NVENC's two-plane CUDA surface /// Copy our imported NV12 [`DeviceBuffer`] (Y + UV planes) into NVENC's two-plane CUDA surface
/// `(y_dst, y_pitch)` / `(uv_dst, uv_pitch)` (`av_hwframe_get_buffer`'s `data[0]`/`data[1]` + /// `(y_dst, y_pitch)` / `(uv_dst, uv_pitch)` (`av_hwframe_get_buffer`'s `data[0]`/`data[1]` +
/// `linesize[0]`/`linesize[1]`). The Y plane is `width`×`height` bytes; the chroma plane is /// `linesize[0]`/`linesize[1]`). The Y plane is `width`×`height` bytes; the chroma plane is
/// `(width/2)·2` bytes × `height/2` rows. The caller must have the shared context current. /// `(width/2)·2` bytes × `height/2` rows. The caller must have the shared context current.
/// `sync: false` enqueues without a CPU wait (stream-ordered consumers only — `src` must stay
/// valid until the downstream stream work completes; see [`copy_stream_handle`]).
pub fn copy_nv12_to_device( pub fn copy_nv12_to_device(
src: &DeviceBuffer, src: &DeviceBuffer,
y_dst: CUdeviceptr, y_dst: CUdeviceptr,
y_pitch: usize, y_pitch: usize,
uv_dst: CUdeviceptr, uv_dst: CUdeviceptr,
uv_pitch: usize, uv_pitch: usize,
sync: bool,
) -> Result<()> { ) -> Result<()> {
let (src_uv_ptr, src_uv_pitch) = src let (src_uv_ptr, src_uv_pitch) = src
.uv .uv
@@ -1183,15 +1234,16 @@ pub fn copy_nv12_to_device(
Height: h / 2, Height: h / 2,
..Default::default() ..Default::default()
}; };
// SAFETY: two unsafe `copy_blocking` device→device copies; the caller must have the shared // SAFETY: two unsafe `copy_issue` device→device copies; the caller must have the shared
// context current (documented). `&y`/`&uv` are live local `CUDA_MEMCPY2D`s outliving each // context current (documented). `&y`/`&uv` are live local `CUDA_MEMCPY2D`s outliving each
// synchronous call. All four device pointers are valid: `src.ptr`/`src_uv_ptr` come from a live // enqueue. All four device pointers are valid: `src.ptr`/`src_uv_ptr` come from a live
// NV12 `DeviceBuffer` (its `.uv` presence was checked via `ok_or_else`), `y_dst`/`uv_dst` are // NV12 `DeviceBuffer` (its `.uv` presence was checked via `ok_or_else`), `y_dst`/`uv_dst` are
// the caller's live NVENC surface planes; the luma copy is `w`×`h`, the chroma copy // the caller's live NVENC surface planes; the luma copy is `w`×`h`, the chroma copy
// `(w/2)*2`×`h/2`, each within its planes. Wrappers → live table. // `(w/2)*2`×`h/2`, each within its planes; `sync: false` shifts the source-lifetime obligation
// to the caller (documented above). Wrappers → live table.
unsafe { unsafe {
copy_blocking(&y, "cuMemcpy2DAsync_v2(nv12 Y dev->dev)")?; copy_issue(&y, "cuMemcpy2DAsync_v2(nv12 Y dev->dev)", sync)?;
copy_blocking(&uv, "cuMemcpy2DAsync_v2(nv12 UV dev->dev)") copy_issue(&uv, "cuMemcpy2DAsync_v2(nv12 UV dev->dev)", sync)
} }
} }
@@ -1199,7 +1251,13 @@ pub fn copy_nv12_to_device(
/// (`av_hwframe_get_buffer`'s `data[0..3]` + `linesize[0..3]` for a `yuv444p` frames context). /// (`av_hwframe_get_buffer`'s `data[0..3]` + `linesize[0..3]` for a `yuv444p` frames context).
/// Each plane is `width`×`height` bytes; the source planes sit at row offsets `0/H/2H` of the /// Each plane is `width`×`height` bytes; the source planes sit at row offsets `0/H/2H` of the
/// single allocation. The caller must have the shared context current. /// single allocation. The caller must have the shared context current.
pub fn copy_yuv444_to_device(src: &DeviceBuffer, dsts: [(CUdeviceptr, usize); 3]) -> Result<()> { /// `sync: false` enqueues without a CPU wait (stream-ordered consumers only — `src` must stay
/// valid until the downstream stream work completes; see [`copy_stream_handle`]).
pub fn copy_yuv444_to_device(
src: &DeviceBuffer,
dsts: [(CUdeviceptr, usize); 3],
sync: bool,
) -> Result<()> {
anyhow::ensure!(src.yuv444, "copy_yuv444_to_device on a non-YUV444 buffer"); anyhow::ensure!(src.yuv444, "copy_yuv444_to_device on a non-YUV444 buffer");
let w = src.width as usize; let w = src.width as usize;
let h = src.height as usize; let h = src.height as usize;
@@ -1215,12 +1273,13 @@ pub fn copy_yuv444_to_device(src: &DeviceBuffer, dsts: [(CUdeviceptr, usize); 3]
Height: h, Height: h,
..Default::default() ..Default::default()
}; };
// SAFETY: unsafe `copy_blocking` device→device copy; the caller must have the shared // SAFETY: unsafe `copy_issue` device→device copy; the caller must have the shared
// context current (documented). `&copy` is a live local outliving the synchronous call; // context current (documented). `&copy` is a live local outliving the enqueue;
// `src.ptr + pitch·h·i` stays within the live 3·H-row stacked allocation (`yuv444` // `src.ptr + pitch·h·i` stays within the live 3·H-row stacked allocation (`yuv444`
// checked above), `dst_ptr`/`dst_pitch` is the caller's live NVENC plane; `w`×`h` fits // checked above), `dst_ptr`/`dst_pitch` is the caller's live NVENC plane; `w`×`h` fits
// both. Wrapper → live table. // both; `sync: false` shifts the source-lifetime obligation to the caller (documented
unsafe { copy_blocking(&copy, "cuMemcpy2DAsync_v2(yuv444 plane dev->dev)")? }; // above). Wrapper → live table.
unsafe { copy_issue(&copy, "cuMemcpy2DAsync_v2(yuv444 plane dev->dev)", sync)? };
} }
Ok(()) Ok(())
} }
+9
View File
@@ -691,6 +691,15 @@ impl EglImporter {
width: u32, width: u32,
height: u32, height: u32,
) -> Result<DeviceBuffer> { ) -> Result<DeviceBuffer> {
// Even dimensions only: the UV copy walks `height.div_ceil(2)` chroma rows (the correct NV12
// count), but the pooled UV plane is sized at `height/2` rows — for an odd height those
// disagree by one row and the copy writes a full `uv_pitch` past the allocation (OOB device
// write / CUDA_ERROR_ILLEGAL_ADDRESS that poisons the shared context). Reject here, matching
// the guards `Nv12Blit::new`/`Yuv444Blit::new` already carry.
anyhow::ensure!(
width % 2 == 0 && height % 2 == 0,
"LINEAR NV12 needs even dimensions (got {width}x{height})"
);
cuda::make_current()?; cuda::make_current()?;
if self if self
.linear_nv12_pool .linear_nv12_pool
+196 -55
View File
@@ -86,9 +86,11 @@ impl VkBridge {
// SAFETY: standard ash bring-up — every call is `unsafe` only because ash cannot statically // SAFETY: standard ash bring-up — every call is `unsafe` only because ash cannot statically
// verify Vulkan handle/CreateInfo validity. `ash::Entry::load` dlopens a real system // verify Vulkan handle/CreateInfo validity. `ash::Entry::load` dlopens a real system
// libvulkan. Each `*CreateInfo`/`AllocateInfo` is built by ash's builders from locals (`app`, // libvulkan. Each `*CreateInfo`/`AllocateInfo` is built by ash's builders from locals (`app`,
// `exts`, `prio`, `qci`, and the inline infos) that all live for the duration of the // `exts`, `prio`, `qci`, `gp_info`, and the inline infos) that all live for the duration of
// synchronous `create_*`/`enumerate_*` call that reads them — in particular the // the synchronous `create_*`/`enumerate_*` call that reads them — the ladder loop rebuilds
// `enabled_extension_names(&exts)` and `queue_priorities(&prio)` borrows outlive their calls. // `prio`/`gp_info`/`qci`/`exts` fresh per attempt, so every `enabled_extension_names(&exts)`
// / `queue_priorities(&prio)` / `push_next(&mut gp_info)` borrow outlives its own
// `create_device` call.
// Every handle passed (`instance`, `phys`, `device`, `qf`, `cmd_pool`) was just created and // Every handle passed (`instance`, `phys`, `device`, `qf`, `cmd_pool`) was just created and
// checked via `?`/`ok_or_else` in this same function, so no invalid handle is ever used. This // checked via `?`/`ok_or_else` in this same function, so no invalid handle is ever used. This
// constructor shares nothing across threads. // constructor shares nothing across threads.
@@ -122,23 +124,93 @@ impl VkBridge {
.ok_or_else(|| anyhow!("no compute-capable queue family"))? .ok_or_else(|| anyhow!("no compute-capable queue family"))?
as u32; as u32;
let exts = [ // Global-priority queue (latency plan §7 LN4, PyroWave's `ac0e7332` lever for the
// VkBridge): the LINEAR/gamescope CSC dispatch shares the SM/compute cores with the
// game, so ask for an elevated global priority to get scheduled ahead of it.
// `PUNKTFUNK_VK_QUEUE_PRIORITY` = off | high | realtime (default realtime); the
// create loop downgrades REALTIME→HIGH→none on NOT_PERMITTED (and retries a plain
// create on INITIALIZATION_FAILED) so a refused class never fails the bridge.
let gp_ext = std::env::var("PUNKTFUNK_VK_QUEUE_PRIORITY")
.ok()
.as_deref()
.map_or(Some(vk::QueueGlobalPriorityKHR::REALTIME), |v| match v {
"off" | "0" => None,
"high" => Some(vk::QueueGlobalPriorityKHR::HIGH),
_ => Some(vk::QueueGlobalPriorityKHR::REALTIME),
})
.and_then(|want| {
// Enable whichever alias the driver advertises (KHR = the promoted name).
let props = instance.enumerate_device_extension_properties(phys).ok()?;
let has = |name: &std::ffi::CStr| {
props
.iter()
.any(|p| p.extension_name_as_c_str() == Ok(name))
};
if has(vk::KHR_GLOBAL_PRIORITY_NAME) {
Some((vk::KHR_GLOBAL_PRIORITY_NAME, want))
} else if has(vk::EXT_GLOBAL_PRIORITY_NAME) {
Some((vk::EXT_GLOBAL_PRIORITY_NAME, want))
} else {
None
}
});
let base_exts = [
ash::khr::external_memory_fd::NAME.as_ptr(), ash::khr::external_memory_fd::NAME.as_ptr(),
ash::ext::external_memory_dma_buf::NAME.as_ptr(), ash::ext::external_memory_dma_buf::NAME.as_ptr(),
]; ];
let mut try_priority = gp_ext.map(|(_, want)| want);
let device = loop {
let prio = [1.0f32]; let prio = [1.0f32];
let qci = [vk::DeviceQueueCreateInfo::default() let mut gp_info = vk::DeviceQueueGlobalPriorityCreateInfoKHR::default()
.global_priority(try_priority.unwrap_or(vk::QueueGlobalPriorityKHR::MEDIUM));
let mut qci0 = vk::DeviceQueueCreateInfo::default()
.queue_family_index(qf) .queue_family_index(qf)
.queue_priorities(&prio)]; .queue_priorities(&prio);
let device = instance let mut exts: Vec<*const std::ffi::c_char> = base_exts.to_vec();
.create_device( if try_priority.is_some() {
qci0 = qci0.push_next(&mut gp_info);
exts.push(gp_ext.expect("try_priority implies gp_ext").0.as_ptr());
}
let qci = [qci0];
match instance.create_device(
phys, phys,
&vk::DeviceCreateInfo::default() &vk::DeviceCreateInfo::default()
.queue_create_infos(&qci) .queue_create_infos(&qci)
.enabled_extension_names(&exts), .enabled_extension_names(&exts),
None, None,
) ) {
.context("vkCreateDevice (external-memory extensions supported?)")?; Ok(d) => {
if let Some(p) = try_priority {
tracing::info!(
priority = ?p,
"VkBridge queue at elevated global priority (CSC schedules \
ahead of a GPU-bound game where the driver honors it)"
);
}
break d;
}
// A refused class must never fail the bridge — walk the ladder down.
Err(
vk::Result::ERROR_NOT_PERMITTED_KHR
| vk::Result::ERROR_INITIALIZATION_FAILED,
) if try_priority == Some(vk::QueueGlobalPriorityKHR::REALTIME) => {
try_priority = Some(vk::QueueGlobalPriorityKHR::HIGH);
}
Err(
vk::Result::ERROR_NOT_PERMITTED_KHR
| vk::Result::ERROR_INITIALIZATION_FAILED,
) if try_priority.is_some() => {
tracing::debug!(
"global-priority queue not permitted — VkBridge at default priority"
);
try_priority = None;
}
Err(e) => {
return Err(e)
.context("vkCreateDevice (external-memory extensions supported?)")
}
}
};
let ext_fd = ash::khr::external_memory_fd::Device::new(&instance, &device); let ext_fd = ash::khr::external_memory_fd::Device::new(&instance, &device);
let queue = device.get_device_queue(qf, 0); let queue = device.get_device_queue(qf, 0);
@@ -193,10 +265,17 @@ impl VkBridge {
/// Import `fd` (dup'd internally; Vulkan owns the dup) as a transfer-src buffer of `size`. /// Import `fd` (dup'd internally; Vulkan owns the dup) as a transfer-src buffer of `size`.
unsafe fn import_src(&mut self, fd: i32, size: u64) -> Result<()> { unsafe fn import_src(&mut self, fd: i32, size: u64) -> Result<()> {
use std::os::fd::{AsRawFd, FromRawFd, IntoRawFd, OwnedFd};
let dup = libc::dup(fd); let dup = libc::dup(fd);
if dup < 0 { if dup < 0 {
bail!("dup(dmabuf fd)"); bail!("dup(dmabuf fd)");
} }
// Own the dup so every early return BEFORE Vulkan consumes it (at `allocate_memory` success)
// closes it. `SrcBuf` holds raw handles with no Drop and is only populated on the success
// path, so each fallible step below must also destroy the buffer it created — otherwise a
// failed import (which the worker survives and the caller retries every frame) leaks a
// VkBuffer + VkDeviceMemory + fd per frame for the worker's whole lifetime.
let dup = OwnedFd::from_raw_fd(dup);
let mut ext_info = vk::ExternalMemoryBufferCreateInfo::default() let mut ext_info = vk::ExternalMemoryBufferCreateInfo::default()
.handle_types(vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT); .handle_types(vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT);
let buffer = self let buffer = self
@@ -212,41 +291,55 @@ impl VkBridge {
.push_next(&mut ext_info), .push_next(&mut ext_info),
None, None,
) )
.context("create import buffer")?; .context("create import buffer")?; // `dup` drops → closes on failure
let mut fd_props = vk::MemoryFdPropertiesKHR::default(); let mut fd_props = vk::MemoryFdPropertiesKHR::default();
self.ext_fd if let Err(e) = self.ext_fd.get_memory_fd_properties(
.get_memory_fd_properties(
vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT, vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT,
dup, dup.as_raw_fd(),
&mut fd_props, &mut fd_props,
) ) {
.context("vkGetMemoryFdPropertiesKHR")?; self.device.destroy_buffer(buffer, None);
return Err(e).context("vkGetMemoryFdPropertiesKHR");
}
let reqs = self.device.get_buffer_memory_requirements(buffer); let reqs = self.device.get_buffer_memory_requirements(buffer);
let mem_type = self.memory_type( let mem_type = match self.memory_type(
reqs.memory_type_bits & fd_props.memory_type_bits, reqs.memory_type_bits & fd_props.memory_type_bits,
vk::MemoryPropertyFlags::empty(), vk::MemoryPropertyFlags::empty(),
)?; ) {
Ok(t) => t,
Err(e) => {
self.device.destroy_buffer(buffer, None);
return Err(e);
}
};
// Vulkan takes ownership of the fd on a SUCCESSFUL import: hand over the raw fd now, and on
// failure close it ourselves (matching the original contract) plus destroy the buffer.
let raw = dup.into_raw_fd();
let mut import = vk::ImportMemoryFdInfoKHR::default() let mut import = vk::ImportMemoryFdInfoKHR::default()
.handle_type(vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT) .handle_type(vk::ExternalMemoryHandleTypeFlags::DMA_BUF_EXT)
.fd(dup); // Vulkan takes ownership of `dup` on success .fd(raw);
let mut dedicated = vk::MemoryDedicatedAllocateInfo::default().buffer(buffer); let mut dedicated = vk::MemoryDedicatedAllocateInfo::default().buffer(buffer);
let memory = self let memory = match self.device.allocate_memory(
.device
.allocate_memory(
&vk::MemoryAllocateInfo::default() &vk::MemoryAllocateInfo::default()
.allocation_size(reqs.size.max(size)) .allocation_size(reqs.size.max(size))
.memory_type_index(mem_type) .memory_type_index(mem_type)
.push_next(&mut import) .push_next(&mut import)
.push_next(&mut dedicated), .push_next(&mut dedicated),
None, None,
) ) {
.map_err(|e| { Ok(m) => m,
libc::close(dup); // failed import does not consume the fd Err(e) => {
anyhow!("import dmabuf memory: {e}") libc::close(raw); // failed import does not consume the fd
})?; self.device.destroy_buffer(buffer, None);
self.device return Err(anyhow!("import dmabuf memory: {e}"));
.bind_buffer_memory(buffer, memory, 0) }
.context("bind import memory")?; };
if let Err(e) = self.device.bind_buffer_memory(buffer, memory, 0) {
// `memory` owns the imported fd — freeing it releases the fd too.
self.device.free_memory(memory, None);
self.device.destroy_buffer(buffer, None);
return Err(e).context("bind import memory");
}
self.src_cache.insert( self.src_cache.insert(
fd, fd,
SrcBuf { SrcBuf {
@@ -263,11 +356,11 @@ impl VkBridge {
if self.dst.as_ref().is_some_and(|d| d.size >= size) { if self.dst.as_ref().is_some_and(|d| d.size >= size) {
return Ok(()); return Ok(());
} }
if let Some(old) = self.dst.take() { // Build the replacement FULLY before retiring the old one. Previously the old dst was
self.device.destroy_buffer(old.buffer, None); // destroyed and `self.dst` nulled up front, so a failed rebuild both dropped the working
self.device.free_memory(old.memory, None); // buffer AND leaked every object the partial rebuild created (`buffer`/`memory` are raw ash
// old.cuda drops its mapping with it // handles with no Drop, and `VkBridge::drop` only frees the live `self.dst`). Now every
} // fallible step unwinds locally, and the swap happens only on full success.
let mut ext_info = vk::ExternalMemoryBufferCreateInfo::default() let mut ext_info = vk::ExternalMemoryBufferCreateInfo::default()
.handle_types(vk::ExternalMemoryHandleTypeFlags::OPAQUE_FD); .handle_types(vk::ExternalMemoryHandleTypeFlags::OPAQUE_FD);
let buffer = self let buffer = self
@@ -285,35 +378,63 @@ impl VkBridge {
.context("create export buffer")?; .context("create export buffer")?;
let reqs = self.device.get_buffer_memory_requirements(buffer); let reqs = self.device.get_buffer_memory_requirements(buffer);
let mem_type = let mem_type =
self.memory_type(reqs.memory_type_bits, vk::MemoryPropertyFlags::DEVICE_LOCAL)?; match self.memory_type(reqs.memory_type_bits, vk::MemoryPropertyFlags::DEVICE_LOCAL) {
Ok(t) => t,
Err(e) => {
self.device.destroy_buffer(buffer, None);
return Err(e);
}
};
let mut export = vk::ExportMemoryAllocateInfo::default() let mut export = vk::ExportMemoryAllocateInfo::default()
.handle_types(vk::ExternalMemoryHandleTypeFlags::OPAQUE_FD); .handle_types(vk::ExternalMemoryHandleTypeFlags::OPAQUE_FD);
let mut dedicated = vk::MemoryDedicatedAllocateInfo::default().buffer(buffer); let mut dedicated = vk::MemoryDedicatedAllocateInfo::default().buffer(buffer);
let memory = self let memory = match self.device.allocate_memory(
.device
.allocate_memory(
&vk::MemoryAllocateInfo::default() &vk::MemoryAllocateInfo::default()
.allocation_size(reqs.size) .allocation_size(reqs.size)
.memory_type_index(mem_type) .memory_type_index(mem_type)
.push_next(&mut export) .push_next(&mut export)
.push_next(&mut dedicated), .push_next(&mut dedicated),
None, None,
) ) {
.context("allocate exportable memory")?; Ok(m) => m,
self.device Err(e) => {
.bind_buffer_memory(buffer, memory, 0) self.device.destroy_buffer(buffer, None);
.context("bind export memory")?; return Err(e).context("allocate exportable memory");
let opaque_fd = self }
.ext_fd };
.get_memory_fd( if let Err(e) = self.device.bind_buffer_memory(buffer, memory, 0) {
self.device.free_memory(memory, None);
self.device.destroy_buffer(buffer, None);
return Err(e).context("bind export memory");
}
let opaque_fd = match self.ext_fd.get_memory_fd(
&vk::MemoryGetFdInfoKHR::default() &vk::MemoryGetFdInfoKHR::default()
.memory(memory) .memory(memory)
.handle_type(vk::ExternalMemoryHandleTypeFlags::OPAQUE_FD), .handle_type(vk::ExternalMemoryHandleTypeFlags::OPAQUE_FD),
) ) {
.context("vkGetMemoryFdKHR")?; Ok(f) => f,
Err(e) => {
self.device.free_memory(memory, None);
self.device.destroy_buffer(buffer, None);
return Err(e).context("vkGetMemoryFdKHR");
}
};
// CUDA imports (and on success owns) the exported fd. Size must match the allocation. // CUDA imports (and on success owns) the exported fd. Size must match the allocation.
let cuda = cuda::ExternalDmabuf::import_owned_fd(opaque_fd, reqs.size) // `import_owned_fd` closes `opaque_fd` on its own failure, so only the Vulkan objects unwind.
.context("cuImportExternalMemory(OPAQUE_FD from Vulkan)")?; let cuda = match cuda::ExternalDmabuf::import_owned_fd(opaque_fd, reqs.size) {
Ok(c) => c,
Err(e) => {
self.device.free_memory(memory, None);
self.device.destroy_buffer(buffer, None);
return Err(e).context("cuImportExternalMemory(OPAQUE_FD from Vulkan)");
}
};
// Full success: retire the previous buffer now, then publish the new one.
if let Some(old) = self.dst.take() {
self.device.destroy_buffer(old.buffer, None);
self.device.free_memory(old.memory, None);
// old.cuda drops its mapping with it
}
tracing::info!(size, "Vulkan→CUDA exportable staging buffer ready"); tracing::info!(size, "Vulkan→CUDA exportable staging buffer ready");
self.dst = Some(DstBuf { self.dst = Some(DstBuf {
buffer, buffer,
@@ -544,9 +665,19 @@ impl VkBridge {
self.device self.device
.queue_submit(self.queue, &[submit], self.fence) .queue_submit(self.queue, &[submit], self.fence)
.context("queue submit")?; .context("queue submit")?;
self.device // Exception-safe wait: a TIMEOUT/DEVICE_LOST must not `?` out with the submission still
// executing — `self.cmd` and `self.fence` are reused every frame, and the caller retries
// on the SAME bridge (and `ensure_dst` later destroys `dst.buffer` assuming no in-flight
// work references it). Drain the GPU and reset the fence before propagating so the shared
// cmd/fence return clean.
if let Err(e) = self
.device
.wait_for_fences(&[self.fence], true, 1_000_000_000) .wait_for_fences(&[self.fence], true, 1_000_000_000)
.context("fence wait")?; {
let _ = self.device.device_wait_idle();
let _ = self.device.reset_fences(&[self.fence]);
return Err(e).context("fence wait");
}
self.device self.device
.reset_fences(&[self.fence]) .reset_fences(&[self.fence])
.context("reset fence")?; .context("reset fence")?;
@@ -639,9 +770,19 @@ impl VkBridge {
self.device self.device
.queue_submit(self.queue, &[submit], self.fence) .queue_submit(self.queue, &[submit], self.fence)
.context("queue submit")?; .context("queue submit")?;
self.device // Exception-safe wait: a TIMEOUT/DEVICE_LOST must not `?` out with the submission still
// executing — `self.cmd` and `self.fence` are reused every frame, and the caller retries
// on the SAME bridge (and `ensure_dst` later destroys `dst.buffer` assuming no in-flight
// work references it). Drain the GPU and reset the fence before propagating so the shared
// cmd/fence return clean.
if let Err(e) = self
.device
.wait_for_fences(&[self.fence], true, 1_000_000_000) .wait_for_fences(&[self.fence], true, 1_000_000_000)
.context("fence wait")?; {
let _ = self.device.device_wait_idle();
let _ = self.device.reset_fences(&[self.fence]);
return Err(e).context("fence wait");
}
self.device self.device
.reset_fences(&[self.fence]) .reset_fences(&[self.fence])
.context("reset fence")?; .context("reset fence")?;
-1
View File
@@ -10,7 +10,6 @@ fn main() {
println!("cargo:rerun-if-changed=src/abi.rs"); println!("cargo:rerun-if-changed=src/abi.rs");
println!("cargo:rerun-if-changed=src/config.rs"); println!("cargo:rerun-if-changed=src/config.rs");
println!("cargo:rerun-if-changed=src/input.rs"); println!("cargo:rerun-if-changed=src/input.rs");
println!("cargo:rerun-if-changed=src/client.rs");
println!("cargo:rerun-if-changed=src/error.rs"); println!("cargo:rerun-if-changed=src/error.rs");
println!("cargo:rerun-if-changed=cbindgen.toml"); println!("cargo:rerun-if-changed=cbindgen.toml");
+92 -16
View File
@@ -78,6 +78,11 @@ impl PunktfunkConfig {
u8::try_from(self.fec_percent).map_err(|_| PunktfunkStatus::InvalidArg)?; u8::try_from(self.fec_percent).map_err(|_| PunktfunkStatus::InvalidArg)?;
let max_data_per_block = let max_data_per_block =
u16::try_from(self.max_data_per_block).map_err(|_| PunktfunkStatus::InvalidArg)?; u16::try_from(self.max_data_per_block).map_err(|_| PunktfunkStatus::InvalidArg)?;
// The one narrowing here that differs by target width: on 32-bit (armeabi-v7a) an
// `as usize` silently truncates a >4 GiB value to a plausible-looking residue that
// passes validate() — reject it instead, like every narrowing above.
let max_frame_bytes =
usize::try_from(self.max_frame_bytes).map_err(|_| PunktfunkStatus::InvalidArg)?;
let cfg = Config { let cfg = Config {
role, role,
phase, phase,
@@ -87,7 +92,7 @@ impl PunktfunkConfig {
max_data_per_block, max_data_per_block,
}, },
shard_payload: self.shard_payload as usize, shard_payload: self.shard_payload as usize,
max_frame_bytes: self.max_frame_bytes as usize, max_frame_bytes,
encrypt: self.encrypt != 0, encrypt: self.encrypt != 0,
key: self.key, key: self.key,
salt: self.salt, salt: self.salt,
@@ -125,6 +130,12 @@ pub struct PunktfunkFrame {
pub frame_index: u32, pub frame_index: u32,
pub pts_ns: u64, pub pts_ns: u64,
pub flags: u32, pub flags: u32,
/// Wall-clock reassembly-completion instant (ns since the Unix epoch, CLOCK_REALTIME — the
/// clock `pts_ns` and the skew handshake use). THIS is the receipt stamp for latency math:
/// a stamp the embedder takes itself at the poll return additionally contains the
/// pre-decode hand-off queue wait, so a client-side standing backlog would masquerade as
/// network latency (ABI v9 — the 2026-07 two-pair standing-latency investigation).
pub received_ns: u64,
} }
/// Snapshot of session counters. /// Snapshot of session counters.
@@ -391,6 +402,7 @@ pub unsafe extern "C" fn punktfunk_client_poll_frame(
frame_index: f.frame_index, frame_index: f.frame_index,
pts_ns: f.pts_ns, pts_ns: f.pts_ns,
flags: f.flags, flags: f.flags,
received_ns: f.received_ns,
}; };
} }
PunktfunkStatus::Ok PunktfunkStatus::Ok
@@ -456,24 +468,32 @@ pub unsafe extern "C" fn punktfunk_set_input_callback(
#[no_mangle] #[no_mangle]
pub unsafe extern "C" fn punktfunk_host_poll_input(s: *mut PunktfunkSession) -> i32 { pub unsafe extern "C" fn punktfunk_host_poll_input(s: *mut PunktfunkSession) -> i32 {
let r = std::panic::catch_unwind(AssertUnwindSafe(|| { let r = std::panic::catch_unwind(AssertUnwindSafe(|| {
let mut count = 0i32;
loop {
// Narrow scope: re-derive the handle and pull ONE event, then drop the borrow
// before dispatching. The callback may legally re-enter `punktfunk_*` on this
// handle (get_stats, send_input, clearing the callback) — with a `&mut` held
// across the call that re-entry aliased it (UB under noalias). Re-reading
// `input_cb` per iteration also makes a mid-drain
// `punktfunk_set_input_callback(s, NULL, NULL)` take effect immediately instead
// of firing the cleared callback for the queued remainder. (Freeing the session
// from inside the callback remains forbidden, as on every entry point.)
let (ev, cb) = {
let s = match unsafe { s.as_mut() } { let s = match unsafe { s.as_mut() } {
Some(s) => s, Some(s) => s,
None => return PunktfunkStatus::NullPointer as i32, None => return PunktfunkStatus::NullPointer as i32,
}; };
let cb = s.input_cb;
let mut count = 0i32;
loop {
match s.inner.poll_input() { match s.inner.poll_input() {
Ok(Some(ev)) => { Ok(Some(ev)) => (ev, s.input_cb),
Ok(None) => break,
Err(e) => return e.status() as i32,
}
};
if let Some((cb, user)) = cb { if let Some((cb, user)) = cb {
cb(&ev as *const InputEvent, user); cb(&ev as *const InputEvent, user);
} }
count += 1; count += 1;
} }
Ok(None) => break,
Err(e) => return e.status() as i32,
}
}
count count
})); }));
r.unwrap_or(PunktfunkStatus::Panic as i32) r.unwrap_or(PunktfunkStatus::Panic as i32)
@@ -878,8 +898,8 @@ pub const PUNKTFUNK_GAMEPAD_AUTO: u32 = 0;
/// uinput X-Box 360 pad (the universal default — every game speaks XInput). /// uinput X-Box 360 pad (the universal default — every game speaks XInput).
pub const PUNKTFUNK_GAMEPAD_XBOX360: u32 = 1; pub const PUNKTFUNK_GAMEPAD_XBOX360: u32 = 1;
/// UHID DualSense (kernel `hid-playstation`): adaptive triggers, lightbar, touchpad, motion — /// UHID DualSense (kernel `hid-playstation`): adaptive triggers, lightbar, touchpad, motion —
/// feedback arrives on the HID-output plane ([`punktfunk_connection_next_hidout`]). Honored /// feedback arrives on the HID-output plane ([`punktfunk_connection_next_hidout`]). Honored on
/// only where available (Linux hosts); otherwise the host falls back to X-Box 360. /// Linux (UHID) and Windows (UMDF minidriver) hosts; otherwise the host falls back to X-Box 360.
pub const PUNKTFUNK_GAMEPAD_DUALSENSE: u32 = 2; pub const PUNKTFUNK_GAMEPAD_DUALSENSE: u32 = 2;
/// uinput X-Box One / Series pad — the X-Box 360 backend with the One/Series USB identity, so /// uinput X-Box One / Series pad — the X-Box 360 backend with the One/Series USB identity, so
/// games show One/Series glyphs. XInput-identical to `XBOX360` otherwise (no game-visible gain; /// games show One/Series glyphs. XInput-identical to `XBOX360` otherwise (no game-visible gain;
@@ -888,8 +908,8 @@ pub const PUNKTFUNK_GAMEPAD_DUALSENSE: u32 = 2;
pub const PUNKTFUNK_GAMEPAD_XBOXONE: u32 = 3; pub const PUNKTFUNK_GAMEPAD_XBOXONE: u32 = 3;
/// UHID DualShock 4 (kernel `hid-playstation` ≥ 6.2): lightbar, touchpad, motion, rumble — the /// UHID DualShock 4 (kernel `hid-playstation` ≥ 6.2): lightbar, touchpad, motion, rumble — the
/// touchpad/motion arrive over the rich-input plane and lightbar over the HID-output plane, like /// touchpad/motion arrive over the rich-input plane and lightbar over the HID-output plane, like
/// DualSense (minus adaptive triggers / player LEDs / mute). Honored only where available (Linux /// DualSense (minus adaptive triggers / player LEDs / mute). Honored on Linux (UHID) and Windows
/// hosts); otherwise the host falls back to X-Box 360. /// (UMDF minidriver) hosts; otherwise the host falls back to X-Box 360.
pub const PUNKTFUNK_GAMEPAD_DUALSHOCK4: u32 = 4; pub const PUNKTFUNK_GAMEPAD_DUALSHOCK4: u32 = 4;
/// UHID classic Steam Controller (Valve `28DE:1102`, kernel `hid-steam`): one stick + dual /// UHID classic Steam Controller (Valve `28DE:1102`, kernel `hid-steam`): one stick + dual
/// trackpads + two grip paddles. Honored only where available (Linux hosts); else Xbox 360. /// trackpads + two grip paddles. Honored only where available (Linux hosts); else Xbox 360.
@@ -899,10 +919,12 @@ pub const PUNKTFUNK_GAMEPAD_STEAMCONTROLLER: u32 = 5;
/// host. Honored on Linux AND Windows hosts; else folds to X-Box 360. /// host. Honored on Linux AND Windows hosts; else folds to X-Box 360.
pub const PUNKTFUNK_GAMEPAD_STEAMDECK: u32 = 6; pub const PUNKTFUNK_GAMEPAD_STEAMDECK: u32 = 6;
/// DualSense Edge (Sony `054C:0DF2`): the DualSense plus two back buttons + two Fn buttons, so a /// DualSense Edge (Sony `054C:0DF2`): the DualSense plus two back buttons + two Fn buttons, so a
/// client's back paddles land on native slots. Folds to `DUALSENSE` until its backend lands. /// client's back paddles land on native slots. Honored on Linux (UHID `hid-playstation`) and
/// Windows (UMDF) hosts; otherwise the host falls back to X-Box 360.
pub const PUNKTFUNK_GAMEPAD_DUALSENSEEDGE: u32 = 7; pub const PUNKTFUNK_GAMEPAD_DUALSENSEEDGE: u32 = 7;
/// Nintendo Switch Pro Controller (Nintendo `057E:2009`, kernel `hid-nintendo`): Nintendo glyphs + /// Nintendo Switch Pro Controller (Nintendo `057E:2009`, kernel `hid-nintendo`): Nintendo glyphs +
/// positional layout, gyro/accel, HD rumble. Folds to `XBOX360` until its backend lands. /// positional layout, gyro/accel, HD rumble. Honored only where available (Linux hosts, UHID
/// `hid-nintendo`); otherwise the host falls back to X-Box 360.
pub const PUNKTFUNK_GAMEPAD_SWITCHPRO: u32 = 8; pub const PUNKTFUNK_GAMEPAD_SWITCHPRO: u32 = 8;
/// New Steam Controller (2026, Valve `28DE:1302`) passed through AS-IS: the host mirrors the /// New Steam Controller (2026, Valve `28DE:1302`) passed through AS-IS: the host mirrors the
/// client's raw Triton input reports out of a virtual SC2 with the real identity, and Steam's /// client's raw Triton input reports out of a virtual SC2 with the real identity, and Steam's
@@ -1744,6 +1766,7 @@ pub unsafe extern "C" fn punktfunk_connection_next_au(
frame_index: f.frame_index, frame_index: f.frame_index,
pts_ns: f.pts_ns, pts_ns: f.pts_ns,
flags: f.flags, flags: f.flags,
received_ns: f.received_ns,
}; };
} }
PunktfunkStatus::Ok PunktfunkStatus::Ok
@@ -1906,6 +1929,13 @@ pub unsafe extern "C" fn punktfunk_connection_next_audio_pcm(
} }
let AudioPcmState { decoder, pcm } = &mut *state; let AudioPcmState { decoder, pcm } = &mut *state;
let dec = decoder.as_mut().unwrap(); let dec = decoder.as_mut().unwrap();
// A header-only datagram (DTX silence — a legal wire form) must be SKIPPED, not
// decoded: `decode_float` treats an empty payload as a loss and synthesizes a full
// 120 ms of concealment for a ~5 ms slot, growing the playout ring without bound.
// Mirrors the host mic pump's guard; the sink underruns to silence on its own.
if pkt.data.is_empty() {
return PunktfunkStatus::NoFrame;
}
// `decode_float` divides the output buffer length by the channel count to get the // `decode_float` divides the output buffer length by the channel count to get the
// per-channel capacity; an empty payload requests packet-loss concealment. // per-channel capacity; an empty payload requests packet-loss concealment.
match dec.decode_float(&pkt.data, pcm, false) { match dec.decode_float(&pkt.data, pcm, false) {
@@ -2956,7 +2986,14 @@ pub unsafe extern "C" fn punktfunk_connection_next_clipboard(
unsafe { *out = out_ev }; unsafe { *out = out_ev };
PunktfunkStatus::Ok PunktfunkStatus::Ok
} }
Err(e) => e.status(), Err(e) => {
// Release the parked payload once the embedder polls past it: clipboard
// traffic is sporadic, so without this a one-off 50 MiB paste stays resident
// for the rest of the session (there is no other release entry point). The
// borrow contract already says `out` data is valid only until the next call.
*c.last_clip.lock().unwrap() = None;
e.status()
}
} }
}) })
} }
@@ -3046,6 +3083,35 @@ pub unsafe extern "C" fn punktfunk_connection_clock_offset_ns(
}) })
} }
/// The **live** host↔client wall-clock offset (nanoseconds, host minus client): the
/// connect-time estimate of [`punktfunk_connection_clock_offset_ns`], updated by every applied
/// mid-stream clock re-sync. Ongoing latency math (per-frame `received pts` splits, the
/// glass-to-glass meter) must use this one — after a wall-clock step/slew the frozen
/// connect-time value reads tens of milliseconds wrong for the rest of the session, while the
/// core itself has already re-synced. Same clock contract as the connect-time getter.
///
/// # Safety
/// `c` is a valid connection handle; `offset_ns` is writable (NULL is skipped).
#[cfg(feature = "quic")]
#[no_mangle]
pub unsafe extern "C" fn punktfunk_connection_clock_offset_now_ns(
c: *const PunktfunkConnection,
offset_ns: *mut i64,
) -> PunktfunkStatus {
guard(|| {
let c = match unsafe { c.as_ref() } {
Some(c) => c,
None => return PunktfunkStatus::NullPointer,
};
unsafe {
if !offset_ns.is_null() {
*offset_ns = c.inner.clock_offset_now_ns();
}
}
PunktfunkStatus::Ok
})
}
/// Ask the host to switch the live session to `width`x`height`@`refresh_hz` without /// Ask the host to switch the live session to `width`x`height`@`refresh_hz` without
/// reconnecting (window resized, refresh changed). Non-blocking enqueue: on acceptance the /// reconnecting (window resized, refresh changed). Non-blocking enqueue: on acceptance the
/// stream continues at the new mode — the first new-mode access unit is an IDR with /// stream continues at the new mode — the first new-mode access unit is an IDR with
@@ -3185,6 +3251,11 @@ pub unsafe extern "C" fn punktfunk_connection_frames_dropped(
out: *mut u64, out: *mut u64,
) -> PunktfunkStatus { ) -> PunktfunkStatus {
guard(|| { guard(|| {
// The header promises "writes 0 on a NULL connection" — honor it BEFORE the handle
// check, so an embedder that skips the status never reads an uninitialized slot.
if !out.is_null() {
unsafe { *out = 0 };
}
let c = match unsafe { c.as_ref() } { let c = match unsafe { c.as_ref() } {
Some(c) => c, Some(c) => c,
None => return PunktfunkStatus::NullPointer, None => return PunktfunkStatus::NullPointer,
@@ -3242,6 +3313,11 @@ pub unsafe extern "C" fn punktfunk_connection_wants_decode_latency(
out: *mut bool, out: *mut bool,
) -> PunktfunkStatus { ) -> PunktfunkStatus {
guard(|| { guard(|| {
// The header promises "writes 0 on a NULL connection" — honor it BEFORE the handle
// check: an uninitialized byte is not even a valid C++/Swift bool to read.
if !out.is_null() {
unsafe { *out = false };
}
let c = match unsafe { c.as_ref() } { let c = match unsafe { c.as_ref() } {
Some(c) => c, Some(c) => c,
None => return PunktfunkStatus::NullPointer, None => return PunktfunkStatus::NullPointer,
@@ -73,6 +73,142 @@ pub(crate) const NOOP_CLOCK_FLUSHES_TO_DISARM: u32 = 2;
/// FIRST no-op clock flush — the moment a step is actually suspected. /// FIRST no-op clock flush — the moment a step is actually suspected.
pub(crate) const CLOCK_RESYNC_INTERVAL: Duration = Duration::from_secs(60); pub(crate) const CLOCK_RESYNC_INTERVAL: Duration = Duration::from_secs(60);
/// Standing-latency bleed (the 2026-07 two-pair investigation): how far above the session's own
/// one-way-delay floor a report window's MINIMUM must sit to count as a standing elevation. The
/// jump-to-live detectors above deliberately ignore anything below ~6 frames / 400 ms, so a
/// small standing state — a sub-frame kernel/reassembly backlog, or a stale clock offset after a
/// wall-clock step — is carried forever and reads as permanent extra "network" latency. 10 ms
/// sits above skew-handshake error + normal LAN jitter, and below a single 60 fps frame period,
/// so the observed one-frame plateau (~17 ms) trips it while a healthy stream cannot.
pub(crate) const STANDING_LAT_THRESH_NS: i128 = 10_000_000;
/// Consecutive elevated report windows (~750 ms each) before the bleed escalates — ~4.5 s of a
/// continuously standing, loss-free elevation. Windows with any loss reset the run: loss means
/// genuine congestion, which the FEC/ABR machinery owns, not this detector.
pub(crate) const STANDING_LAT_WINDOWS: u32 = 6;
/// Per-session cap on flush+keyframe bleeds. A standing state that survives a clock re-sync AND
/// this many local flushes is not local and not clock — the path latency itself changed; the
/// detector disarms with a warning instead of paying a recovery keyframe every few seconds.
pub(crate) const STANDING_LAT_MAX_BLEEDS: u32 = 3;
/// What the standing-latency detector asks the pump to do this window (see [`StandingLatency`]).
#[derive(Debug, PartialEq, Eq)]
pub(crate) enum StandingLatAction {
None,
/// First escalation: ask for a mid-stream clock re-sync — free, and a stale offset from a
/// stepped/slewed wall clock produces exactly this signature (an applied re-sync re-bases
/// the floor via the pump's `clock_gen` watch, clearing the elevation if that was the cause).
Resync {
above_ms: i64,
},
/// The elevation survived a re-sync attempt: flush the local receive backlog + request a
/// keyframe (the jump-to-live action), draining a real sub-threshold standing queue. The
/// pump reports execution back via [`StandingLatency::bled`]; an unexecuted action simply
/// re-arms next window.
Bleed {
above_ms: i64,
},
/// Bleed cap reached and the elevation is back: give up and say so.
Disarm {
above_ms: i64,
},
}
/// Detector for a small, constant, loss-free one-way-delay elevation — the standing state the
/// jump-to-live thresholds deliberately tolerate. Tracks the session's OWD floor (minimum of
/// report-window minimums since start / last re-base) and escalates when windows sit
/// persistently above it: re-sync first, then a bounded number of flush+keyframe bleeds, then
/// disarm. Pure state machine (no clocks, no I/O) so the escalation ladder is unit-testable.
pub(crate) struct StandingLatency {
/// Lowest window-minimum OWD seen since session start / last [`rebase`](Self::rebase).
floor_ns: Option<i128>,
/// Minimum per-frame OWD this report window; `None` = no frames yet.
window_min_ns: Option<i128>,
/// Consecutive elevated windows.
run: u32,
/// The current elevation already got its re-sync request — next escalation is a bleed.
resync_tried: bool,
bleeds: u32,
disarmed: bool,
}
impl StandingLatency {
pub(crate) fn new() -> Self {
StandingLatency {
floor_ns: None,
window_min_ns: None,
run: 0,
resync_tried: false,
bleeds: 0,
disarmed: false,
}
}
/// Feed one frame's skew-corrected OWD (capture→reassembly-complete, ns). Caller gates on a
/// live clock offset and plausibility (0 < owd < 10 s), like the ABR OWD signal.
pub(crate) fn note_frame(&mut self, owd_ns: i128) {
self.window_min_ns = Some(match self.window_min_ns {
Some(m) => m.min(owd_ns),
None => owd_ns,
});
}
/// Close a report window. `loss_free` = the window carried zero loss (loss resets the run —
/// congestion is the FEC/ABR machinery's problem, and queues under loss are not "standing").
pub(crate) fn on_window(&mut self, loss_free: bool) -> StandingLatAction {
let Some(wmin) = self.window_min_ns.take() else {
return StandingLatAction::None; // no frames this window — no evidence either way
};
let floor = *self.floor_ns.get_or_insert(wmin);
self.floor_ns = Some(floor.min(wmin));
let above_ns = wmin - floor;
if self.disarmed {
return StandingLatAction::None;
}
if !loss_free || above_ns < STANDING_LAT_THRESH_NS {
self.run = 0;
if above_ns < STANDING_LAT_THRESH_NS {
self.resync_tried = false; // elevation cleared — a future one re-syncs first again
}
return StandingLatAction::None;
}
self.run += 1;
if self.run < STANDING_LAT_WINDOWS {
return StandingLatAction::None;
}
self.run = 0; // each escalation gets a fresh observation run
let above_ms = (above_ns / 1_000_000) as i64;
if !self.resync_tried {
self.resync_tried = true;
StandingLatAction::Resync { above_ms }
} else if self.bleeds < STANDING_LAT_MAX_BLEEDS {
StandingLatAction::Bleed { above_ms }
} else {
self.disarmed = true;
StandingLatAction::Disarm { above_ms }
}
}
/// The pump executed a [`StandingLatAction::Bleed`] (flush + keyframe). The floor is KEPT: a
/// successful bleed brings OWD back down to it (elevation clears naturally); an unsuccessful
/// one leaves the elevation visible so the ladder continues toward the cap.
pub(crate) fn bled(&mut self) {
self.bleeds += 1;
self.window_min_ns = None;
}
/// A mid-stream clock re-sync was APPLIED (the pump's `clock_gen` watch): every OWD reading
/// shifted, so the floor and any elevation measured under the old offset are meaningless —
/// re-learn from scratch. The bleed budget survives (it caps keyframes per session).
pub(crate) fn rebase(&mut self) {
self.floor_ns = None;
self.window_min_ns = None;
self.run = 0;
self.resync_tried = false;
}
}
/// Client decode-stage latency accumulator for the adaptive-bitrate controller's decode signal. /// Client decode-stage latency accumulator for the adaptive-bitrate controller's decode signal.
/// The embedder adds one sample per decoded frame ([`NativeClient::report_decode_us`], µs from the /// The embedder adds one sample per decoded frame ([`NativeClient::report_decode_us`], µs from the
/// AU leaving [`NativeClient::next_frame`] to its decoded output) and the data-plane pump drains a /// AU leaving [`NativeClient::next_frame`] to its decoded output) and the data-plane pump drains a
@@ -191,6 +327,7 @@ mod frame_channel_tests {
pts_ns: i as u64, pts_ns: i as u64,
flags: 0, flags: 0,
complete: true, complete: true,
received_ns: 0,
} }
} }
@@ -258,3 +395,143 @@ mod frame_channel_tests {
assert_eq!(popped(&ch), Some(total - FRAME_QUEUE_HARD_CAP as u32)); assert_eq!(popped(&ch), Some(total - FRAME_QUEUE_HARD_CAP as u32));
} }
} }
#[cfg(test)]
mod standing_latency_tests {
use super::{
StandingLatAction, StandingLatency, STANDING_LAT_MAX_BLEEDS, STANDING_LAT_THRESH_NS,
STANDING_LAT_WINDOWS,
};
const FLOOR: i128 = 2_000_000; // a healthy 2 ms LAN OWD
const ELEVATED: i128 = FLOOR + STANDING_LAT_THRESH_NS + 7_000_000; // ~one 60fps frame above
/// Run `n` windows at `owd`, asserting every window but the last returns None; returns the
/// last window's action.
fn run_windows(d: &mut StandingLatency, owd: i128, n: u32) -> StandingLatAction {
for i in 0..n {
d.note_frame(owd);
let a = d.on_window(true);
if i + 1 < n {
assert_eq!(a, StandingLatAction::None, "window {i} escalated early");
} else {
return a;
}
}
unreachable!("n > 0 by construction");
}
/// Learn a clean floor: one window at the healthy OWD.
fn learned(d: &mut StandingLatency) {
d.note_frame(FLOOR);
assert_eq!(d.on_window(true), StandingLatAction::None);
}
#[test]
fn healthy_stream_never_escalates() {
let mut d = StandingLatency::new();
learned(&mut d);
// Jitter riding above the floor but under the threshold: never a run.
for _ in 0..(STANDING_LAT_WINDOWS * 4) {
d.note_frame(FLOOR + STANDING_LAT_THRESH_NS - 1);
assert_eq!(d.on_window(true), StandingLatAction::None);
}
}
#[test]
fn escalation_ladder_resync_then_bleeds_then_disarm() {
let mut d = StandingLatency::new();
learned(&mut d);
// First full elevated run asks for the free fix: a clock re-sync.
assert!(matches!(
run_windows(&mut d, ELEVATED, STANDING_LAT_WINDOWS),
StandingLatAction::Resync { .. }
));
// Re-sync didn't help (no rebase came) — each further run is a bleed, up to the cap...
for _ in 0..STANDING_LAT_MAX_BLEEDS {
assert!(matches!(
run_windows(&mut d, ELEVATED, STANDING_LAT_WINDOWS),
StandingLatAction::Bleed { .. }
));
d.bled();
}
// ...then the detector gives up loudly, once, and stays quiet.
assert!(matches!(
run_windows(&mut d, ELEVATED, STANDING_LAT_WINDOWS),
StandingLatAction::Disarm { .. }
));
d.note_frame(ELEVATED);
assert_eq!(d.on_window(true), StandingLatAction::None);
}
#[test]
fn loss_windows_reset_the_run() {
let mut d = StandingLatency::new();
learned(&mut d);
for _ in 0..(STANDING_LAT_WINDOWS - 1) {
d.note_frame(ELEVATED);
assert_eq!(d.on_window(true), StandingLatAction::None);
}
// A lossy window means congestion, not a standing state: run resets...
d.note_frame(ELEVATED);
assert_eq!(d.on_window(false), StandingLatAction::None);
// ...so the ladder needs the full run again before acting.
assert!(matches!(
run_windows(&mut d, ELEVATED, STANDING_LAT_WINDOWS),
StandingLatAction::Resync { .. }
));
}
#[test]
fn recovery_resets_the_ladder_to_resync_first() {
let mut d = StandingLatency::new();
learned(&mut d);
assert!(matches!(
run_windows(&mut d, ELEVATED, STANDING_LAT_WINDOWS),
StandingLatAction::Resync { .. }
));
// The elevation clears on its own (e.g. the successful bleed case, or transient): the
// next episode starts back at the free escalation, not at a bleed.
d.note_frame(FLOOR);
assert_eq!(d.on_window(true), StandingLatAction::None);
assert!(matches!(
run_windows(&mut d, ELEVATED, STANDING_LAT_WINDOWS),
StandingLatAction::Resync { .. }
));
}
#[test]
fn applied_resync_rebases_and_clears_a_stale_offset_elevation() {
let mut d = StandingLatency::new();
learned(&mut d);
assert!(matches!(
run_windows(&mut d, ELEVATED, STANDING_LAT_WINDOWS),
StandingLatAction::Resync { .. }
));
// The re-sync APPLIES (pump sees clock_gen move) → rebase. The corrected offset brings
// OWD readings back to truth; the floor re-learns and nothing ever escalates to a bleed.
d.rebase();
for _ in 0..(STANDING_LAT_WINDOWS * 2) {
d.note_frame(FLOOR);
assert_eq!(d.on_window(true), StandingLatAction::None);
}
}
#[test]
fn empty_windows_are_no_evidence() {
let mut d = StandingLatency::new();
learned(&mut d);
for _ in 0..(STANDING_LAT_WINDOWS - 1) {
d.note_frame(ELEVATED);
assert_eq!(d.on_window(true), StandingLatAction::None);
}
// A frameless window (paused stream) neither advances nor resets the run...
assert_eq!(d.on_window(true), StandingLatAction::None);
// ...so one more elevated window completes it.
d.note_frame(ELEVATED);
assert!(matches!(
d.on_window(true),
StandingLatAction::Resync { .. }
));
}
}
+28 -6
View File
@@ -425,6 +425,12 @@ impl NativeClient {
Ok(Ok(t)) => t, Ok(Ok(t)) => t,
Ok(Err(e)) => return Err(e), Ok(Err(e)) => return Err(e),
Err(_) => { Err(_) => {
// A connect we already reported as failed must not leave a lingering host
// session if the handshake lands late: mark it a deliberate QUIT (not a plain
// drop / close code 0) so the worker's close tells the host to tear down now
// instead of holding the session (and its virtual display) for a reconnect
// that will never come.
quit.store(true, Ordering::SeqCst);
shutdown.store(true, Ordering::SeqCst); shutdown.store(true, Ordering::SeqCst);
return Err(PunktfunkError::Timeout); return Err(PunktfunkError::Timeout);
} }
@@ -456,9 +462,14 @@ impl NativeClient {
hot_tids, hot_tids,
clock_offset, clock_offset,
decode_lat, decode_lat,
// The controller arms exactly when the pump does (see `abr::BitrateController::new` // The controller arms exactly when the pump does — all three terms, not two: Automatic
// below): Automatic (the user asked for bitrate 0) and not a rate-pinned PyroWave stream. // (the user asked for bitrate 0), not a rate-pinned PyroWave stream, AND the host
wants_decode: bitrate_kbps == 0 && negotiated.codec != crate::quic::CODEC_PYROWAVE, // echoed the rate it actually configured. Dropping the last term made this
// over-advertise against an old host that reports no rate, so an embedder fed decode
// latency to a controller that never runs.
wants_decode: bitrate_kbps == 0
&& negotiated.codec != crate::quic::CODEC_PYROWAVE
&& negotiated.bitrate_kbps > 0,
mode: mode_slot, mode: mode_slot,
host_fingerprint: negotiated.host_fingerprint, host_fingerprint: negotiated.host_fingerprint,
resolved_compositor: negotiated.compositor, resolved_compositor: negotiated.compositor,
@@ -703,14 +714,23 @@ impl NativeClient {
// Reset the accumulator so a fresh run doesn't blend into the previous one. // Reset the accumulator so a fresh run doesn't blend into the previous one.
*self.probe.lock().unwrap() = ProbeState { *self.probe.lock().unwrap() = ProbeState {
active: true, active: true,
duration_ms,
..Default::default() ..Default::default()
}; };
self.ctrl_tx let sent = self
.ctrl_tx
.try_send(CtrlRequest::Probe(ProbeRequest { .try_send(CtrlRequest::Probe(ProbeRequest {
target_kbps, target_kbps,
duration_ms, duration_ms,
})) }))
.map_err(|_| PunktfunkError::Closed) .map_err(|_| PunktfunkError::Closed);
if sent.is_err() {
// Nothing was asked of the host, so nothing will ever answer. Leaving `active` latched
// would suppress the pump's entire report tick for the rest of the session (the pump
// mirrors the startup path's rollback at the same point).
self.probe.lock().unwrap().active = false;
}
sent
} }
/// Read the current speed-test measurement (partial until `done`, final once the host's /// Read the current speed-test measurement (partial until `done`, final once the host's
@@ -745,7 +765,9 @@ impl NativeClient {
0.0 0.0
} as f32; } as f32;
// Host-side drop: what the send buffer couldn't even accept (the host-side ceiling). // Host-side drop: what the send buffer couldn't even accept (the host-side ceiling).
let offered_wire = p.host_wire_packets + p.host_send_dropped; // Saturating: both counters arrive verbatim off the wire (same discipline as the
// saturating_sub/mul above — a hostile sum must not overflow-panic a debug build).
let offered_wire = p.host_wire_packets.saturating_add(p.host_send_dropped);
let host_drop_pct = if offered_wire > 0 { let host_drop_pct = if offered_wire > 0 {
p.host_send_dropped as f64 / offered_wire as f64 * 100.0 p.host_send_dropped as f64 / offered_wire as f64 * 100.0
} else { } else {
@@ -32,6 +32,11 @@ pub(crate) struct ProbeState {
pub(crate) host_duration_ms: u32, pub(crate) host_duration_ms: u32,
/// The host's `ProbeResult` arrived → the measurement is final. /// The host's `ProbeResult` arrived → the measurement is final.
pub(crate) done: bool, pub(crate) done: bool,
/// The requested burst length, so the pump can arm a watchdog for a host that never answers.
/// Without one, an ignored `ProbeRequest` latches `active` forever and the pump's whole report
/// tick — loss reports, the ABR window feed, the standing-latency ladder and pending clock
/// re-syncs — stays suppressed for the rest of the session.
pub(crate) duration_ms: u32,
} }
/// A finished/partial speed-test measurement, returned by [`NativeClient::probe_result`]. /// A finished/partial speed-test measurement, returned by [`NativeClient::probe_result`].
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,181 @@
//! Control task: the handshake stream stays open for mid-stream renegotiation + speed tests.
//! Outbound requests (mode switch, probe) and inbound replies (Reconfigured, ProbeResult) are
//! multiplexed with `select!`; a single outbound channel (`ctrl_rx`) keeps one writer so the
//! two `&mut ctrl_send` borrows don't collide across branches.
use super::super::*;
use super::*;
pub(super) struct ControlTask {
pub(super) ctrl_rx: tokio::sync::mpsc::Receiver<CtrlRequest>,
pub(super) ctrl_send: quinn::SendStream,
pub(super) ctrl_recv: io::MsgReader,
/// `None` = no connect-time skew handshake (old host) — clock re-sync stays off.
pub(super) clock_rtt_ns: Option<u64>,
pub(super) mode_slot: Arc<Mutex<Mode>>,
pub(super) probe: Arc<Mutex<ProbeState>>,
/// The latest host `BitrateChanged` ack, drained by the pump's ABR on its report tick.
pub(super) bitrate_ack: Arc<Mutex<Option<u32>>>,
pub(super) clock_offset: Arc<std::sync::atomic::AtomicI64>,
pub(super) clock_gen: Arc<AtomicU32>,
/// Clipboard metadata events (ClipState/ClipOffer) feed the same event plane the
/// clipboard task uses for fetch data.
pub(super) clip_event_tx: std::sync::mpsc::SyncSender<ClipEventCore>,
}
impl ControlTask {
pub(super) async fn run(self) {
let ControlTask {
mut ctrl_rx,
mut ctrl_send,
mut ctrl_recv,
clock_rtt_ns,
mode_slot,
probe,
bitrate_ack,
clock_offset,
clock_gen,
clip_event_tx,
} = self;
// Mid-stream clock re-sync (see [`ClockResync`]): a batch runs every
// CLOCK_RESYNC_INTERVAL and whenever the pump asks (CtrlRequest::ClockResync after
// its first no-op clock flush). Echoes interleave with the other control replies in
// the read arm below; only when the host answered the connect-time handshake — an
// old host would just eat the probes.
let mut resync = ClockResync::new();
let mut resync_tick = tokio::time::interval_at(
tokio::time::Instant::now() + CLOCK_RESYNC_INTERVAL,
CLOCK_RESYNC_INTERVAL,
);
resync_tick.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay);
loop {
tokio::select! {
req = ctrl_rx.recv() => {
let Some(req) = req else { break }; // client dropped
let bytes = match req {
CtrlRequest::Mode(m) => Reconfigure { mode: m }.encode(),
CtrlRequest::Probe(p) => p.encode(),
CtrlRequest::Keyframe => RequestKeyframe.encode(),
CtrlRequest::Rfi(r) => r.encode(),
CtrlRequest::Loss(r) => r.encode(),
CtrlRequest::SetBitrate(k) => SetBitrate { bitrate_kbps: k }.encode(),
CtrlRequest::ClockResync => {
if clock_rtt_ns.is_none() {
continue; // no connect-time handshake — host can't answer
}
resync.begin(wall_clock_ns()).encode()
}
CtrlRequest::ClipControl(c) => c.encode(),
CtrlRequest::ClipOffer(o) => o.encode(),
};
if io::write_msg(&mut ctrl_send, &bytes).await.is_err() {
break;
}
}
_ = resync_tick.tick(), if clock_rtt_ns.is_some() => {
let probe = resync.begin(wall_clock_ns());
if io::write_msg(&mut ctrl_send, &probe.encode()).await.is_err() {
break;
}
}
msg = ctrl_recv.read_msg() => {
let Ok(msg) = msg else { break }; // stream closed
if let Ok(ack) = Reconfigured::decode(&msg) {
if ack.accepted {
*mode_slot.lock().unwrap() = ack.mode;
tracing::info!(mode = ?ack.mode, "host accepted mode switch");
} else {
tracing::warn!(active = ?ack.mode, "host rejected mode switch");
}
} else if let Ok(result) = ProbeResult::decode(&msg) {
let mut p = probe.lock().unwrap();
// Freeze the delivered figures now (the burst is done), before resumed
// video can inflate the packet counters.
let base_p = p.base_packets.unwrap_or(p.rx_packets_now);
let base_b = p.base_bytes.unwrap_or(p.rx_bytes_now);
p.delivered_packets = p.rx_packets_now.saturating_sub(base_p);
p.delivered_bytes = p.rx_bytes_now.saturating_sub(base_b);
p.host_goodput_bytes = result.bytes_sent;
p.host_au = result.packets_sent;
p.host_wire_packets = result.wire_packets_sent;
p.host_send_dropped = result.send_dropped;
p.host_duration_ms = result.duration_ms;
p.done = true;
p.active = false; // burst over — the pump stops mirroring counters
tracing::info!(
host_goodput_bytes = result.bytes_sent,
wire_packets_sent = result.wire_packets_sent,
send_dropped = result.send_dropped,
duration_ms = result.duration_ms,
delivered_packets = p.delivered_packets,
"speed-test probe result"
);
} else if let Ok(ack) = BitrateChanged::decode(&msg) {
// Adaptive bitrate: the host's clamp is authoritative — park it for
// the pump's controller (which also reads any ack as "this host
// renegotiates", arming further steps).
tracing::info!(
kbps = ack.bitrate_kbps,
"host re-targeted encoder bitrate"
);
*bitrate_ack.lock().unwrap() = Some(ack.bitrate_kbps);
} else if let Ok(echo) = ClockEcho::decode(&msg) {
match resync.on_echo(&echo, wall_clock_ns()) {
ResyncStep::Probe(p) => {
if io::write_msg(&mut ctrl_send, &p.encode()).await.is_err() {
break;
}
}
ResyncStep::Done { offset_ns, rtt_ns } => {
// Never let a congested window bias the offset (frames read
// late exactly then) — keep the old estimate and let the next
// periodic batch try again.
if accept_resync(rtt_ns, clock_rtt_ns.unwrap_or(0)) {
// info, not debug: ≤1/min, and it is THE forensic
// trail for a stale-offset (stepped/slewed wall clock)
// latency plateau — the 2026-07 two-pair investigation
// had to reconstruct this blind.
tracing::info!(
offset_ns,
rtt_us = rtt_ns / 1000,
"mid-stream clock re-sync applied"
);
clock_offset.store(offset_ns, Ordering::Relaxed);
clock_gen.fetch_add(1, Ordering::Relaxed);
} else {
tracing::info!(
rtt_us = rtt_ns / 1000,
"clock re-sync batch discarded — RTT above the \
connect-time baseline (congested window)"
);
}
}
ResyncStep::Idle => {}
}
} else if let Ok(state) = ClipState::decode(&msg) {
// Host ack / policy / backend update for the toggle UI (try_send: a
// lagging embedder drops the newest — a stale toggle heals on the next).
let _ = clip_event_tx.try_send(ClipEventCore::State {
enabled: state.enabled,
policy: state.policy,
reason: state.reason,
});
} else if let Ok(offer) = ClipOffer::decode(&msg) {
// The host copied something: surface the lazy format list; the embedder
// fetches only if a local app pastes.
let _ = clip_event_tx.try_send(ClipEventCore::RemoteOffer {
seq: offer.seq,
kinds: offer.kinds,
});
} else {
tracing::warn!(
tag = ?msg.first(),
len = msg.len(),
"unknown control message — ignoring"
);
}
}
}
}
}
}
@@ -0,0 +1,568 @@
//! The blocking data-plane pump: poll the session for access units, run the adaptive-FEC
//! loss reports, the ABR controller + startup capacity probe, the jump-to-live detectors,
//! and the standing-latency bleed, and hand frames to the embedder.
use super::super::*;
use super::*;
/// Data-plane pump on a blocking thread: poll the session, hand frames to the embedder.
/// try_send drops the newest frame when the embedder lags (freshness over completeness).
/// Speed-test filler ([`FLAG_PROBE`]) is folded into the probe accumulator instead of the
/// decoder queue — it isn't video.
pub(super) struct DataPump {
pub(super) session: Session,
pub(super) frames: Arc<FrameChannel>,
pub(super) ctrl_tx: tokio::sync::mpsc::Sender<CtrlRequest>,
pub(super) shutdown: Arc<std::sync::atomic::AtomicBool>,
pub(super) probe: Arc<Mutex<ProbeState>>,
pub(super) hot_tids: Arc<Mutex<Vec<i32>>>,
pub(super) clock_offset: Arc<std::sync::atomic::AtomicI64>,
pub(super) clock_gen: Arc<AtomicU32>,
pub(super) decode_lat: Arc<Mutex<DecodeLatAcc>>,
pub(super) frames_dropped: Arc<std::sync::atomic::AtomicU64>,
pub(super) fec_recovered: Arc<std::sync::atomic::AtomicU64>,
pub(super) bitrate_ack: Arc<Mutex<Option<u32>>>,
/// The embedder's REQUESTED rate (0 = Automatic — the only case the ABR arms).
pub(super) bitrate_kbps: u32,
/// The rate the host actually configured (echoed in Welcome).
pub(super) resolved_bitrate_kbps: u32,
pub(super) negotiated_codec: u8,
}
impl DataPump {
pub(super) fn run(self) {
let DataPump {
mut session,
frames,
ctrl_tx,
shutdown: pump_shutdown,
probe: pump_probe,
hot_tids: pump_hot_tids,
clock_offset: pump_clock_offset,
clock_gen: pump_clock_gen,
decode_lat: pump_decode_lat,
frames_dropped,
fec_recovered,
bitrate_ack,
bitrate_kbps,
resolved_bitrate_kbps,
negotiated_codec,
} = self;
pin_thread_user_interactive(); // feeds the frame channel → the user-interactive video pump
register_hot_tid(&pump_hot_tids); // this thread does UDP receive + FEC reassembly — hint it
// Adaptive-FEC loss reporting: every ADAPT_REPORT_INTERVAL, report the loss observed over the
// window (shards FEC recovered, plus a bump if any frame went unrecoverable) so the host can
// size FEC to the link. Suppressed during a speed test (its FLAG_PROBE filler would skew it).
const ADAPT_REPORT_INTERVAL: Duration = Duration::from_millis(750);
let mut last_report = Instant::now();
let (
mut last_recovered,
mut last_late,
mut last_received,
mut last_dropped,
mut last_bytes,
) = (0u64, 0u64, 0u64, 0u64, 0u64);
// PUNKTFUNK_PERF: per-window pump observability — the Session's receive stage split
// (recv / decrypt / reassemble+FEC, see `Session::take_pump_perf`) and completed-AU
// inter-arrival jitter. Smoothness has no metric otherwise: jump-to-live counters only
// fire after the stream is already seconds behind.
let pump_perf_on = std::env::var("PUNKTFUNK_PERF").is_ok_and(|v| v != "0");
let mut arrivals_us: Vec<u32> = Vec::new();
let mut last_arrival: Option<Instant> = None;
// Adaptive bitrate (see `crate::abr`): armed only when the embedder asked for Automatic
// (`bitrate_kbps == 0`) and the host echoed the rate it actually configured (an old host
// echoes 0 → controller stays permanently off). Fed once per report window with the same
// deltas the LossReport uses, plus the window's mean skew-corrected one-way delay, the
// actual delivered throughput (climb gate + proven-throughput mark), and whether a
// jump-to-live flush fired.
// PyroWave sessions PIN their rate (§4.6): AIMD descent turns wavelets to mush well
// above its floor, and the climb probe's VBV reasoning doesn't apply to hard
// per-frame CBR — controller and capacity probe stay off (0 = permanently off).
let rate_pinned = negotiated_codec == crate::quic::CODEC_PYROWAVE;
let mut abr = BitrateController::new(if bitrate_kbps == 0 && !rate_pinned {
resolved_bitrate_kbps
} else {
0
});
// Startup link-capacity probe (Automatic sessions): the controller's ceiling is the
// negotiated start rate — the conservative 20 Mbps default, historically a box Automatic
// could NEVER climb out of. One speed-test burst shortly after the stream settles
// measures what the link actually delivers; ×0.7 (headroom for FEC overhead + variance)
// becomes the climb ceiling and slow start does the rest. Old hosts decline (all-zero
// reply) or never answer (timeout clears the state so LossReports resume) — either way
// the ceiling stays negotiated, exactly the old behavior. PUNKTFUNK_ABR_PROBE=0 opts out.
const CAPACITY_PROBE_KBPS: u32 = 2_000_000;
const CAPACITY_PROBE_MS: u32 = 800;
const CAPACITY_PROBE_DELAY: Duration = Duration::from_secs(2);
const CAPACITY_PROBE_TIMEOUT: Duration = Duration::from_secs(6);
let mut capacity_probe_at: Option<Instant> = (bitrate_kbps == 0
&& !rate_pinned
&& resolved_bitrate_kbps > 0
&& std::env::var("PUNKTFUNK_ABR_PROBE").map_or(true, |v| v != "0"))
.then(|| Instant::now() + CAPACITY_PROBE_DELAY);
let mut capacity_probe_deadline: Option<Instant> = None;
// Edge detector + watchdog for a probe of EITHER origin (the startup capacity probe or an
// embedder speed test via `NativeClient::request_probe`). The startup path had both built
// in; the embedder path had neither, so an unanswered request wedged the report tick and a
// finished one left the ABR window anchored before the burst.
let mut was_probing = false;
let mut probe_watchdog: Option<Instant> = None;
let (mut owd_sum_ns, mut owd_frames) = (0i128, 0u32);
let mut flush_in_window = false;
// Jump-to-live state (see the guard in the loop below): when the clock-based over-bound
// run began (`stale_since`, armed only when the skew handshake succeeded so the clocks
// are comparable), when the clock-free non-draining-queue run began (`standing_since`),
// and the last-jump instant for the shared cooldown. Wall-clock runs (T1.4), not frame
// counts — the detectors' sensitivity must not scale with fps or repeat cadence.
let mut stale_since: Option<Instant> = None;
let mut standing_since: Option<Instant> = None;
let mut last_flush: Option<Instant> = None;
// Clock-detector health: consecutive clock-triggered flushes that found no local backlog
// (see NOOP_FLUSH_DATAGRAMS). Reaching NOOP_CLOCK_FLUSHES_TO_DISARM turns the clock-based
// detector off (a clock step / upstream queue it can't fix) — until a mid-stream clock
// re-sync lands and re-arms it (`pump_clock_gen` below). The FIRST no-op flush also asks
// the control task for an immediate re-sync (via the report tick): the flush finding no
// local backlog IS the "the wall clock stepped under me" signal.
let mut noop_clock_flushes: u32 = 0;
let mut clock_detector_armed = true;
let mut resync_wanted = false;
let mut seen_clock_gen = pump_clock_gen.load(Ordering::Relaxed);
// Standing-latency bleed (see StandingLatency): the third detector, for the small,
// constant, loss-free OWD elevation the two jump-to-live detectors deliberately
// tolerate (< QUEUE_HIGH frames, < FLUSH_LATENCY behind) — a sub-frame standing
// backlog, or a stale clock offset after a wall-clock step, either of which otherwise
// reads as permanent extra "network" latency for the rest of the session.
let mut standing_lat = StandingLatency::new();
while !pump_shutdown.load(Ordering::SeqCst) {
// The live host↔client offset: re-loaded every iteration so an applied mid-stream
// re-sync takes effect on the very next frame's latency math.
let clock_offset_ns = pump_clock_offset.load(Ordering::Relaxed);
// An applied re-sync invalidates the staleness run measured under the OLD offset:
// reset the counters and re-arm the clock-based detector if a step had disarmed it.
let gen = pump_clock_gen.load(Ordering::Relaxed);
if gen != seen_clock_gen {
seen_clock_gen = gen;
stale_since = None;
noop_clock_flushes = 0;
// Every OWD reading shifted with the offset — the standing-latency floor and
// any elevation measured under the old one are meaningless now. If a stale
// offset WAS the elevation, this is also the moment it gets fixed.
standing_lat.rebase();
if !clock_detector_armed {
clock_detector_armed = true;
tracing::info!("clock re-sync applied — clock-based jump-to-live re-armed");
}
}
// Mirror the reassembler's unrecoverable-drop count for the client's keyframe-recovery
// loop, and (during a speed test) the packet-level receive counters for the throughput
// measurement. Updated every iteration (not just on a produced frame) so they stay current
// through a total-loss drought where no AU completes. Cheap: a few relaxed atomic loads.
let st = session.stats();
frames_dropped.store(st.frames_dropped, Ordering::Relaxed);
fec_recovered.store(st.fec_recovered_shards, Ordering::Relaxed);
let probe_active = {
let mut p = pump_probe.lock().unwrap();
if p.active && !p.done {
p.rx_packets_now = st.packets_received;
p.rx_bytes_now = st.bytes_received;
p.base_packets.get_or_insert(st.packets_received);
p.base_bytes.get_or_insert(st.bytes_received);
}
p.active && !p.done
};
// A probe just ended (either kind): rebase EVERY window anchor past the burst. Its
// FLAG_PROBE filler landed in `bytes_received`/`packets_received` (session.rs counts
// every accepted datagram) but never reached the decoder, and the report tick was
// suppressed for the whole burst, so `last_*` still points before it. Without this the
// first post-burst window reads the burst rate as `actual_kbps` and poisons the ABR's
// monotone proven-throughput high-water mark — which never decays — and divides the
// window's loss by a packet count inflated with filler.
if was_probing && !probe_active {
last_recovered = st.fec_recovered_shards;
last_late = st.fec_late_shards;
last_received = st.packets_received;
last_dropped = st.frames_dropped;
last_bytes = st.bytes_received;
last_report = Instant::now();
}
// Arm a watchdog on the leading edge of ANY probe, so a host that silently ignores
// `ProbeRequest` (an old build — anticipated, see the capacity-probe timeout below)
// cannot latch `active` forever and suppress the report tick for the whole session.
if !was_probing && probe_active {
let burst = Duration::from_millis(pump_probe.lock().unwrap().duration_ms as u64);
probe_watchdog = Some(Instant::now() + burst + CAPACITY_PROBE_TIMEOUT);
}
if !probe_active {
probe_watchdog = None;
} else if let Some(deadline) = probe_watchdog {
if Instant::now() >= deadline {
probe_watchdog = None;
pump_probe.lock().unwrap().active = false;
tracing::warn!(
"speed-test probe unanswered — clearing it so loss reports and ABR resume"
);
}
}
was_probing = probe_active;
// Fire the startup link-capacity probe once the stream has settled (see the constants
// above), and fold its measurement into the ABR ceiling when the result lands.
// Never steal the slot from an embedder speed test in flight: there is one `ProbeState`
// and no correlation id, so a clobber both wrecks the user's "Test connection" figure
// (its base counters get re-snapshotted mid-burst against the full-burst denominator)
// and mis-scales our own ceiling. Retry once it finishes.
if capacity_probe_at.is_some_and(|at| Instant::now() >= at) && probe_active {
capacity_probe_at = Some(Instant::now() + CAPACITY_PROBE_DELAY);
} else if capacity_probe_at.is_some_and(|at| Instant::now() >= at) {
capacity_probe_at = None;
*pump_probe.lock().unwrap() = ProbeState {
active: true,
duration_ms: CAPACITY_PROBE_MS,
..Default::default()
};
if ctrl_tx
.try_send(CtrlRequest::Probe(ProbeRequest {
target_kbps: CAPACITY_PROBE_KBPS,
duration_ms: CAPACITY_PROBE_MS,
}))
.is_ok()
{
capacity_probe_deadline = Some(Instant::now() + CAPACITY_PROBE_TIMEOUT);
tracing::info!(
target_kbps = CAPACITY_PROBE_KBPS,
duration_ms = CAPACITY_PROBE_MS,
"adaptive bitrate: startup link-capacity probe"
);
} else {
pump_probe.lock().unwrap().active = false; // ctrl queue full — skip
}
}
if let Some(deadline) = capacity_probe_deadline {
let mut p = pump_probe.lock().unwrap();
if p.done {
capacity_probe_deadline = None;
// An all-zero reply is a decline (old host / probe-less build) — keep the
// negotiated ceiling. Otherwise: delivered wire kbps × 0.7.
if p.host_duration_ms > 0 && p.delivered_bytes > 0 {
let delivered_kbps = (p.delivered_bytes.saturating_mul(8)
/ p.host_duration_ms.max(1) as u64)
as u32;
let ceiling = delivered_kbps.saturating_mul(7) / 10;
abr.set_ceiling(ceiling);
tracing::info!(
delivered_kbps,
ceiling_kbps = ceiling,
"adaptive bitrate: link-capacity probe done — climb ceiling set"
);
} else {
tracing::info!(
"adaptive bitrate: capacity probe declined — keeping negotiated ceiling"
);
}
// The probe's FLAG_PROBE filler landed in `bytes_received` but never reached
// the decoder — rebase the ABR window's byte counter past it, or the next
// window's "actual throughput" reads as the burst rate and poisons the
// controller's proven-throughput high-water mark with the LINK rate.
last_bytes = st.bytes_received;
} else if Instant::now() >= deadline {
// The host never answered (a build that ignores ProbeRequest): clear the
// stuck-active state so LossReports resume, keep the negotiated ceiling.
p.active = false;
capacity_probe_deadline = None;
tracing::info!(
"adaptive bitrate: capacity probe timed out (old host?) — keeping negotiated ceiling"
);
}
}
if !probe_active && last_report.elapsed() >= ADAPT_REPORT_INTERVAL {
// A no-op clock flush earlier in this window suspected a wall-clock step: fire
// the mid-stream re-sync now (once — the 60 s periodic covers everything else).
if resync_wanted {
resync_wanted = false;
let _ = ctrl_tx.try_send(CtrlRequest::ClockResync);
}
let window_dropped = st.frames_dropped.wrapping_sub(last_dropped);
let loss_ppm = window_loss_ppm(
st.fec_recovered_shards.wrapping_sub(last_recovered),
st.fec_late_shards.wrapping_sub(last_late),
st.packets_received.wrapping_sub(last_received),
window_dropped,
);
let _ = ctrl_tx.try_send(CtrlRequest::Loss(LossReport { loss_ppm }));
// Standing-latency bleed: close the detector's window with this report's loss
// verdict and run its escalation ladder — re-sync first (free; a stale offset
// from a stepped wall clock produces exactly this signature and the applied
// re-sync rebases the floor), then a bounded flush+keyframe (drains a real
// sub-threshold standing backlog the jump-to-live thresholds tolerate), then a
// loud disarm (the path latency itself changed; nothing local fixes that).
match standing_lat.on_window(loss_ppm == 0 && window_dropped == 0) {
StandingLatAction::None => {}
StandingLatAction::Resync { above_ms } => {
tracing::info!(
above_ms,
"standing latency above the session floor with zero loss — \
requesting a clock re-sync first (a stale offset reads exactly \
like this)"
);
let _ = ctrl_tx.try_send(CtrlRequest::ClockResync);
}
StandingLatAction::Bleed { above_ms } => {
// Shares the jump-to-live cooldown: an unexecuted bleed simply re-arms
// over the next windows (the detector's run rebuilds).
if last_flush.is_none_or(|t| t.elapsed() >= FLUSH_COOLDOWN) {
last_flush = Some(Instant::now());
// Deliberately NOT `flush_in_window = true`: that flag is the ABR's
// SEVERE verdict (an immediate ×0.7 back-off), and the bleed fires
// only after ~6 provably loss-free windows with a sub-25ms elevation
// the controller itself scores as fine. The bleed's effect reaches
// the ABR through the window's own honest signals (OWD/loss/decode);
// the flag stays exclusive to the jump-to-live path below.
let flushed = session.flush_backlog().unwrap_or(0);
let dropped = frames.clear();
let _ = ctrl_tx.try_send(CtrlRequest::Keyframe);
standing_lat.bled();
tracing::warn!(
above_ms,
flushed_datagrams = flushed,
dropped_frames = dropped,
"standing latency survived a clock re-sync — bled the local \
backlog (flush + keyframe)"
);
}
}
StandingLatAction::Disarm { above_ms } => {
tracing::warn!(
above_ms,
"standing latency persists after a re-sync and every bleed — not \
local, not clock; the path latency changed. Leaving it be \
(reconnect re-baselines)"
);
}
}
// Adaptive bitrate: drain any host ack first (its clamp is authoritative), then
// feed the controller this window's congestion signals; a decision becomes a
// SetBitrate on the control stream.
if let Some(acked) = bitrate_ack.lock().unwrap().take() {
abr.on_ack(acked);
}
let owd_mean_us =
(owd_frames > 0).then(|| (owd_sum_ns / owd_frames as i128 / 1000) as i64);
(owd_sum_ns, owd_frames) = (0, 0);
// Drain the embedder's decode-latency window (always, so it stays bounded even when
// the controller is disabled) → the mean feeds the decode signal; `None` when the
// embedder reported nothing this window (old embedder / no decoded frames).
let decode_mean_us = {
let mut acc = pump_decode_lat.lock().unwrap();
let (sum, count) = (acc.sum_us, acc.count);
*acc = DecodeLatAcc::default();
(count > 0).then(|| (sum / count as u64) as i64)
};
// The window's ACTUAL delivered throughput — what the pipeline really carried, vs
// the target it was allowed. Wire bytes (headers + FEC) slightly overstate the
// media rate the decoder ingests; acceptable for the climb gate / proven-mark
// semantics (both compare against targets with their own headroom).
let window_ms = last_report.elapsed().as_millis().max(1) as u64;
let actual_kbps = (st.bytes_received.wrapping_sub(last_bytes).saturating_mul(8)
/ window_ms) as u32;
if let Some(kbps) = abr.on_window(
Instant::now(),
window_dropped,
loss_ppm,
owd_mean_us,
decode_mean_us,
actual_kbps,
flush_in_window,
) {
// Log the window's signals alongside the decision so an on-glass session can
// tell a decode-driven re-target (the new signal — decode_mean_us elevated with
// loss/OWD flat) from a network-driven one.
tracing::info!(
kbps,
loss_ppm,
owd_mean_us = owd_mean_us.unwrap_or(-1),
decode_mean_us = decode_mean_us.unwrap_or(-1),
actual_kbps,
flushed = flush_in_window,
"adaptive bitrate: requesting encoder re-target"
);
let _ = ctrl_tx.try_send(CtrlRequest::SetBitrate(kbps));
}
flush_in_window = false;
last_report = Instant::now();
last_recovered = st.fec_recovered_shards;
last_late = st.fec_late_shards;
last_received = st.packets_received;
last_dropped = st.frames_dropped;
last_bytes = st.bytes_received;
if pump_perf_on {
if let Some(p) = session.take_pump_perf() {
let per_pkt_ns = |ns: u64| ns.checked_div(p.packets).unwrap_or(0);
tracing::info!(
recv_ms = p.recv_ns / 1_000_000,
decrypt_ms = p.decrypt_ns / 1_000_000,
reasm_ms = p.reasm_ns / 1_000_000,
packets = p.packets,
batches = p.batches,
pkts_per_batch = p.packets.checked_div(p.batches).unwrap_or(0),
decrypt_ns_pkt = per_pkt_ns(p.decrypt_ns),
reasm_ns_pkt = per_pkt_ns(p.reasm_ns),
"pump stage split (window)"
);
}
// Inter-arrival jitter over the window's completed AUs. `late` counts gaps
// over 2× the window median — the "a frame arrived visibly off-beat" tally.
if arrivals_us.len() >= 8 {
arrivals_us.sort_unstable();
let pct = |q: usize| arrivals_us[(arrivals_us.len() - 1) * q / 100];
let (p50, p95) = (pct(50), pct(95));
let late = arrivals_us.iter().filter(|&&d| d > p50 * 2).count();
tracing::info!(
frames = arrivals_us.len() + 1,
arrival_p50_us = p50,
arrival_p95_us = p95,
arrival_max_us = arrivals_us.last().copied().unwrap_or(0),
late,
"frame inter-arrival jitter (window)"
);
}
arrivals_us.clear();
}
}
match session.poll_frame() {
Ok(frame) => {
if frame.flags & FLAG_PROBE as u32 != 0 {
continue; // speed-test filler, not video — measured via the counters above
}
if pump_perf_on {
let now = Instant::now();
if let Some(prev) = last_arrival.replace(now) {
// 4096 ≈ 17 s at 240 fps — a stuck window can't grow it unbounded.
if arrivals_us.len() < 4096 {
arrivals_us
.push((now - prev).as_micros().min(u32::MAX as u128) as u32);
}
}
}
// Jump-to-live guard. A standing receive/hand-off queue never drains by itself —
// the pump consumes strictly in order at the arrival rate, so once behind, the
// stream stays behind for good (observed live: stuck 67 s). Pre-decode AUs are
// reference-chained (infinite GOP), so we can NOT drop a frame mid-stream to catch
// up; the only safe recovery is to discard the whole backlog and re-anchor decode
// on a fresh keyframe. Two independent "we're behind" signals arm it, both gated by
// FLUSH_COOLDOWN, both suspended during a speed test (the probe MEASURES a saturated
// queue; flushing would corrupt its counters):
// * clock-based — completed frames sit > FLUSH_LATENCY behind the skew-corrected
// capture clock continuously for FLUSH_AFTER. Needs the skew handshake, and
// also catches kernel/reassembler backlog the hand-off queue hasn't reached yet.
// * clock-free — the pre-decode hand-off queue stopped draining: its depth stayed
// ≥ QUEUE_HIGH (never falling to QUEUE_LOW, still high at the trip) for
// STANDING_TIME. Works with no handshake / a same-clock session (where the
// clock path is disarmed), and is the direct signal that the embedder can't
// keep up. A transient Wi-Fi clump drains within ~100 ms and never trips it.
if probe_active {
// Keep both detectors disarmed across a speed test so its (deliberately)
// saturated queue doesn't leave a primed run that fires the moment it ends.
stale_since = None;
standing_since = None;
} else {
let lat_ns = if clock_offset_ns != 0 {
now_realtime_ns() + clock_offset_ns as i128 - frame.pts_ns as i128
} else {
0
};
// Feed the adaptive-bitrate controller's OWD window (mean capture→received
// delay): rising delay under zero loss is queue growth — the pre-loss
// congestion signal. Only meaningful with a clock handshake.
if clock_offset_ns != 0 && lat_ns > 0 {
owd_sum_ns += lat_ns;
owd_frames += 1;
// The standing-latency detector rides the same signal, but off the
// window MINIMUM (robust against jitter/burst spikes — a standing
// state elevates the floor itself). Same 10 s plausibility clamp as
// the hn stats use.
if lat_ns < 10_000_000_000 {
standing_lat.note_frame(lat_ns);
}
}
if clock_detector_armed
&& clock_offset_ns != 0
&& lat_ns > FLUSH_LATENCY.as_nanos() as i128
{
stale_since.get_or_insert_with(Instant::now);
} else {
stale_since = None;
}
let depth = frames.depth();
if depth >= QUEUE_HIGH {
standing_since.get_or_insert_with(Instant::now);
} else if depth <= QUEUE_LOW {
standing_since = None;
}
// The queue trip additionally requires the depth to still be high NOW, so
// a run that started ≥ high but is hovering in the hysteresis band (a
// clump mid-drain) never fires on elapsed time alone.
let clock_behind = stale_since.is_some_and(|t| t.elapsed() >= FLUSH_AFTER);
let queue_behind = depth >= QUEUE_HIGH
&& standing_since.is_some_and(|t| t.elapsed() >= STANDING_TIME);
if (clock_behind || queue_behind)
&& last_flush.is_none_or(|t| t.elapsed() >= FLUSH_COOLDOWN)
{
stale_since = None;
standing_since = None;
last_flush = Some(Instant::now());
flush_in_window = true; // strongest "link can't hold the rate" signal
let flushed = session.flush_backlog().unwrap_or(0);
let dropped = frames.clear();
let _ = ctrl_tx.try_send(CtrlRequest::Keyframe);
tracing::warn!(
behind_ms = if clock_behind { lat_ns / 1_000_000 } else { -1 },
queue_depth = depth,
flushed_datagrams = flushed,
dropped_frames = dropped,
"receive backlog stopped draining — jumped to live (flush + keyframe)"
);
// Clock-detector health check: a clock-only trigger whose flush found
// no local backlog is a false "behind" reading (a wall-clock step, or
// an upstream queue a local flush can't drain) — repeated, it would
// cost a recovery IDR every cooldown forever. Disarm after two in a
// row; the clock-free queue detector keeps covering real backlogs.
if clock_behind
&& !queue_behind
&& flushed < NOOP_FLUSH_DATAGRAMS
&& dropped == 0
{
noop_clock_flushes += 1;
if noop_clock_flushes == 1 {
// First no-op flush = a wall-clock step is the prime
// suspect: ask for an immediate re-sync (sent on the next
// report tick). Applied, it resets these counters and
// re-arms the detector before the disarm below triggers.
resync_wanted = true;
}
if noop_clock_flushes >= NOOP_CLOCK_FLUSHES_TO_DISARM {
clock_detector_armed = false;
tracing::warn!(
"clock-based jump-to-live disarmed — its flushes found no \
local backlog (clock step or upstream queueing suspected); \
the queue-depth detector stays armed"
);
}
} else {
noop_clock_flushes = 0;
}
continue; // this frame is part of the stale past — don't render it
}
}
frames.push(frame);
}
Err(PunktfunkError::NoFrame) => {
std::thread::sleep(Duration::from_micros(300));
}
Err(_) => break,
}
}
// The pump exited (shutdown / fatal session error) — wake any consumer blocked in
// `next_frame` with a Closed signal instead of a spurious timeout (the old mpsc did this
// implicitly when the sender dropped).
frames.close();
}
}
@@ -0,0 +1,79 @@
//! Datagram demux: host → client audio/rumble (try_send: a lagging embedder drops the
//! newest packet rather than backing up the QUIC receive path).
use super::*;
pub(super) async fn run(
conn: quinn::Connection,
audio_tx: std::sync::mpsc::SyncSender<AudioPacket>,
rumble_tx: std::sync::mpsc::SyncSender<RumbleUpdate>,
rumble_feed: super::super::rumble::RumbleFeed,
hidout_tx: std::sync::mpsc::SyncSender<crate::quic::HidOutput>,
hdr_meta_tx: std::sync::mpsc::SyncSender<crate::quic::HdrMeta>,
host_timing_tx: std::sync::mpsc::SyncSender<crate::quic::HostTiming>,
) {
// Per-pad reorder gate for v2 rumble envelopes (the seq analog of the host's gamepad-state
// gate): a datagram the network reordered must not roll a stopped motor back on. Legacy v1
// datagrams carry no seq and bypass it (an old host's own periodic re-send is the only heal).
let mut rumble_last_seq: [Option<u8>; crate::input::MAX_PADS] = [None; crate::input::MAX_PADS];
while let Ok(d) = conn.read_datagram().await {
match d.first() {
Some(&crate::quic::AUDIO_MAGIC) => {
if let Some((seq, pts_ns, opus)) = crate::quic::decode_audio_datagram(&d) {
let _ = audio_tx.try_send(AudioPacket {
seq,
pts_ns,
data: opus.to_vec(),
});
}
}
Some(&crate::quic::RUMBLE_MAGIC) => {
if let Some(u) = crate::quic::decode_rumble_envelope(&d) {
// Gate v2 envelopes on their per-pad seq; forward v1 (envelope: None) as-is.
let fresh = match u.envelope {
Some(env) => {
let idx = u.pad as usize;
if idx < crate::input::MAX_PADS {
if crate::input::GamepadSnapshot::seq_newer(
env.seq,
rumble_last_seq[idx],
) {
rumble_last_seq[idx] = Some(env.seq);
true
} else {
false // reordered/duplicate — drop, keep the newer state
}
} else {
true // out-of-range pad (host never sends these): no gate
}
}
None => true,
};
if fresh {
let ttl = u.envelope.map(|e| e.ttl_ms);
// Both consumers are fed; an embedder drains exactly one of them
// (the legacy queue, or the policy engine's command API).
let _ = rumble_tx.try_send((u.pad, u.low, u.high, ttl));
rumble_feed.wire_update(u.pad, u.low, u.high, ttl);
}
}
}
Some(&crate::quic::HIDOUT_MAGIC) => {
if let Some(h) = HidOutput::decode(&d) {
let _ = hidout_tx.try_send(h);
}
}
Some(&crate::quic::HDR_META_MAGIC) => {
if let Some(m) = crate::quic::decode_hdr_meta_datagram(&d) {
let _ = hdr_meta_tx.try_send(m);
}
}
Some(&crate::quic::HOST_TIMING_MAGIC) => {
if let Some(t) = crate::quic::decode_host_timing_datagram(&d) {
let _ = host_timing_tx.try_send(t);
}
}
_ => {} // unknown tag — a newer host; ignore
}
}
}
@@ -0,0 +1,204 @@
//! Connect + handshake: dial the host (cert-pinned), exchange Hello/Welcome/Start on the
//! control stream, run the wall-clock skew handshake, hole-punch the data port, and stand
//! up the data-plane [`Session`]. A typed application close from the host surfaces as
//! [`PunktfunkError::Rejected`] instead of the generic transport error.
use super::super::*;
use super::*;
/// Everything [`run_pump`](super::run_pump) needs from a successful connect + handshake.
pub(super) struct HandshakeOut {
pub(super) conn: quinn::Connection,
pub(super) session: Session,
pub(super) ctrl_send: quinn::SendStream,
pub(super) ctrl_recv: io::MsgReader,
pub(super) negotiated: Negotiated,
pub(super) host_caps: u8,
}
pub(super) async fn connect_and_handshake(args: &WorkerArgs) -> Result<HandshakeOut> {
let (host, port, pin) = (&args.host, args.port, args.pin);
let (mode, compositor, gamepad) = (args.mode, args.compositor, args.gamepad);
let (bitrate_kbps, video_caps, audio_channels) =
(args.bitrate_kbps, args.video_caps, args.audio_channels);
let (video_codecs, preferred_codec, display_hdr) =
(args.video_codecs, args.preferred_codec, args.display_hdr);
let (launch, identity, shutdown) = (&args.launch, &args.identity, &args.shutdown);
let remote: std::net::SocketAddr = join_host_port(host, port)
.parse()
.map_err(|_| PunktfunkError::InvalidArg("host:port"))?;
let (ep, observed) = endpoint::client_pinned_with_identity(
pin,
identity.as_ref().map(|(c, k)| (c.as_str(), k.as_str())),
);
let ep = ep.map_err(|e| PunktfunkError::Io(std::io::Error::other(e.to_string())))?;
let conn = ep
.connect(remote, "punktfunk")
.map_err(|_| PunktfunkError::InvalidArg("connect"))?
.await
.map_err(|e| {
// A pin mismatch surfaces as a TLS failure; report it as a crypto error so
// the embedder can distinguish "wrong host identity" from plain IO trouble.
let fp_mismatch =
pin.is_some() && observed.lock().unwrap().map(|fp| Some(fp) != pin) == Some(true);
if fp_mismatch {
PunktfunkError::Crypto
} else {
PunktfunkError::Io(std::io::Error::other(e.to_string()))
}
})?;
let fingerprint = observed.lock().unwrap().unwrap_or([0u8; 32]);
// The rest of the handshake runs in an inner future so a failure can consult
// `conn.close_reason()`: a host that turned us away with a typed application close
// (pairing not armed / denied / approval timeout / version mismatch / busy) surfaces
// as `PunktfunkError::Rejected` instead of the generic transport error the failed
// read produces — the difference between "not accepted" and the actual cause.
let handshake = async {
let (mut send, recv) = conn
.open_bi()
.await
.map_err(|e| PunktfunkError::Io(std::io::Error::other(e.to_string())))?;
// Frame every read on this stream through the resumable reader: the control loop
// below drives it from a `select!` arm and `clock_sync` wraps it in a timeout, and a
// partial frame lost to either would misalign the stream for the whole session.
let mut recv = io::MsgReader::new(recv);
io::write_msg(
&mut send,
&Hello {
abi_version: crate::WIRE_VERSION,
mode,
compositor,
gamepad,
bitrate_kbps,
// No device name yet: the connect ABI has no name parameter (pairing does). The
// host falls back to a fingerprint-derived label in its pending-approval list.
name: None,
// Library id to launch this session, if the embedder asked for one.
launch: launch.clone(),
// The embedder's decode/present caps (e.g. the Windows client advertises
// VIDEO_CAP_10BIT | VIDEO_CAP_HDR). The host only upgrades to a 10-bit / HDR encode
// when the matching bit is set, so `0` stays an 8-bit BT.709 stream. HOST_TIMING is
// OR'd in unconditionally: every NativeClient build demuxes the 0xCF plane, and the
// bit only asks the host for observability datagrams (never changes the encode).
// PROBE_SEQ likewise: the shared reassembler keeps probe filler in its own window
// (every embedder inherits it), so the host may burst speed tests without consuming
// video frame indexes. STREAMED_AU the same way: the shared reassembler accepts
// sentinel-headed streamed blocks (retro-validated at the final block), so the host
// may overlap a multi-slice encode's tail with packetize/FEC/pacing.
video_caps: video_caps
| crate::quic::VIDEO_CAP_HOST_TIMING
| crate::quic::VIDEO_CAP_PROBE_SEQ
| crate::quic::VIDEO_CAP_STREAMED_AU,
// Requested surround channel count; the host echoes the resolved value in Welcome.
audio_channels,
// The codecs this client can decode + its soft preference (0 = auto). The host
// resolves the emitted codec from these and reports it in `Welcome::codec`.
video_codecs,
preferred_codec,
// The client display's HDR volume → the host's virtual-display EDID (host apps
// tone-map to the client's real panel). `None` = unknown/SDR.
display_hdr,
}
.encode(),
)
.await?;
let welcome = Welcome::decode(&recv.read_msg().await?)?;
if welcome.compositor != CompositorPref::Auto {
tracing::info!(
compositor = welcome.compositor.as_str(),
"host resolved compositor"
);
}
if welcome.gamepad != GamepadPref::Auto {
tracing::info!(
gamepad = welcome.gamepad.as_str(),
"host resolved gamepad backend"
);
}
// Reserve our data-plane port, then start the host.
let probe = std::net::UdpSocket::bind("0.0.0.0:0")?;
let udp_port = probe.local_addr()?.port();
drop(probe);
io::write_msg(
&mut send,
&Start {
client_udp_port: udp_port,
}
.encode(),
)
.await?;
// Wall-clock skew handshake on the control stream (before the session's control task takes
// it): align our clock to the host's so the embedder can express receive/present instants in
// the host's capture clock (the AU `pts_ns`). 0 ⇒ an old host that didn't answer (shared-clock
// assumption, as before). This is the substrate for glass-to-glass present-time measurement.
let (clock_offset_ns, clock_rtt_ns) =
match crate::quic::clock_sync(&mut send, &mut recv).await {
Some(skew) => {
tracing::info!(
offset_ns = skew.offset_ns,
rtt_us = skew.rtt_ns / 1000,
rounds = skew.rounds,
"clock skew estimated (host-client)"
);
(skew.offset_ns, Some(skew.rtt_ns))
}
None => (0, None),
};
let host_udp = std::net::SocketAddr::new(remote.ip(), welcome.udp_port);
let transport =
UdpTransport::connect(&format!("0.0.0.0:{udp_port}"), &host_udp.to_string())?;
// Hole-punch the host's data port so video traverses a NAT / stateful inter-VLAN firewall
// (control + side planes ride the client-initiated QUIC; the raw video UDP needs the client
// to open the path first). Stops with the session via the shared shutdown flag.
if let Ok(sock) = transport.try_clone_socket() {
crate::transport::spawn_data_punch(sock, shutdown.clone());
}
let mut session = Session::new(welcome.session_config(Role::Client), Box::new(transport))?;
// PyroWave sessions opt into partial delivery (plan §4.4): an aged-out lossy
// frame arrives as blocks-with-holes instead of vanishing — the all-intra codec
// renders it as one frame of localized blur, strictly better than a freeze.
if welcome.codec == crate::quic::CODEC_PYROWAVE {
session.set_deliver_partial_frames(true);
}
Ok::<_, PunktfunkError>((
session,
send,
recv,
Negotiated {
mode: welcome.mode,
compositor: welcome.compositor,
gamepad: welcome.gamepad,
host_fingerprint: fingerprint,
bitrate_kbps: welcome.bitrate_kbps,
clock_offset_ns,
clock_rtt_ns,
bit_depth: welcome.bit_depth,
color: welcome.color,
chroma_format: welcome.chroma_format,
audio_channels: welcome.audio_channels,
codec: welcome.codec,
shard_payload: welcome.shard_payload,
host_caps: welcome.host_caps,
},
welcome.host_caps,
))
};
match handshake.await {
Ok((session, send, recv, negotiated, host_caps)) => Ok(HandshakeOut {
conn,
session,
ctrl_send: send,
ctrl_recv: recv,
negotiated,
host_caps,
}),
Err(e) => Err(match reject_from_close(&conn) {
Some(r) => PunktfunkError::Rejected(r),
None => e,
}),
}
}
@@ -0,0 +1,139 @@
//! Input task: embedder events → QUIC datagrams. Toward a host that advertised
//! HOST_CAP_GAMEPAD_STATE, the per-transition gamepad events every embedder still emits are
//! folded into idempotent, sequence-numbered full-state snapshots (`GamepadSnapshot`): the
//! datagram plane drops and reorders (and sheds oldest-first at the 4 KiB send cap), so a lost
//! per-transition event would corrupt held pad state until the *next* change — a held trigger
//! stuck wrong indefinitely. Snapshots heal on the next send, the seq lets the host drop stale
//! reorders, and a periodic refresh of every touched pad bounds any loss to one refresh
//! interval — the same idempotent-state discipline as the host's 500 ms rumble refresh.
//! Keyboard/mouse/touch events pass through unchanged; an older host (no caps bit) keeps
//! getting the legacy per-transition gamepad events.
use super::super::*;
use super::*;
pub(super) async fn run(
conn: quinn::Connection,
mut input_rx: tokio::sync::mpsc::UnboundedReceiver<InputEvent>,
gamepad_snapshots: bool,
) {
use crate::input::{GamepadSnapshot, InputKind, MAX_PADS};
// Touched pads only: an entry appears on the first gamepad event for that index, so the
// refresh never conjures a virtual pad the embedder didn't drive.
let mut pads: [Option<GamepadSnapshot>; MAX_PADS] = [None; MAX_PADS];
// Per-pad wrapping seq that PERSISTS across a pad's remove/re-add on the same index (the
// snapshot itself is cleared to `None` on removal). A removal takes `seq[idx] + 1` so it
// supersedes every prior snapshot; the re-added pad's first snapshot takes the next value
// after that, so the host's seq gate accepts it instead of rejecting a restarted-at-0 seq.
let mut seq: [u8; MAX_PADS] = [0; MAX_PADS];
// Re-sends of a removal still owed on refresh ticks (the removal rides the lossy datagram
// plane; a single lost one would silently strand a ghost pad on the host — the exact bug
// the removal fixes). Mirrors the host's rumble stop burst: a few time-spread re-sends,
// each with a fresh (higher) seq, and canceled the moment the pad is driven again.
const REMOVE_RESENDS: u8 = 2;
let mut remove_owed: [u8; MAX_PADS] = [0; MAX_PADS];
// Per-pad declared controller kind ([`GamepadArrival`]) + its owed re-sends: the host needs
// the kind before the pad's first frame to build a matching virtual device (mixed types), so
// like the removal it rides the lossy plane with a small time-spread re-send burst.
const ARRIVAL_RESENDS: u8 = 2;
let mut arrival: [Option<u8>; MAX_PADS] = [None; MAX_PADS];
let mut arrival_owed: [u8; MAX_PADS] = [0; MAX_PADS];
let mut refresh = tokio::time::interval(Duration::from_millis(100));
refresh.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay);
loop {
tokio::select! {
ev = input_rx.recv() => {
let Some(ev) = ev else { break };
let idx = ev.flags as usize;
if gamepad_snapshots
&& matches!(ev.kind, InputKind::GamepadButton | InputKind::GamepadAxis)
&& idx < MAX_PADS
{
// The pad is being driven — cancel any owed removal (a re-plug on this
// index; its fresh snapshot seq already supersedes the removal's).
remove_owed[idx] = 0;
let snap = pads[idx].get_or_insert(GamepadSnapshot {
pad: idx as u8,
..Default::default()
});
// Unknown axis ids don't send (the host's legacy fold drops them too).
if snap.fold(&ev) {
seq[idx] = seq[idx].wrapping_add(1);
snap.seq = seq[idx];
let _ = conn
.send_datagram(snap.to_event().encode().to_vec().into());
}
continue;
}
if gamepad_snapshots && ev.kind == InputKind::GamepadRemove && idx < MAX_PADS {
// Stop refreshing the pad and forward a seq-stamped removal (in the shared
// seq space) so the host tears its virtual device down and no reordered
// snapshot can resurrect it; arm the re-send burst against datagram loss.
// Drop any owed kind declaration too — a re-plug on this index sends its own.
pads[idx] = None;
arrival[idx] = None;
arrival_owed[idx] = 0;
seq[idx] = seq[idx].wrapping_add(1);
remove_owed[idx] = REMOVE_RESENDS;
let rem = crate::input::InputEvent {
flags: crate::input::encode_gamepad_remove(idx as u8, seq[idx]),
..ev
};
let _ = conn.send_datagram(rem.encode().to_vec().into());
continue;
}
if gamepad_snapshots && ev.kind == InputKind::GamepadArrival && idx < MAX_PADS {
// Remember the declared kind (`code`) and forward it, arming a re-send burst
// so the host learns it before the pad's first frame even under loss.
arrival[idx] = Some(ev.code as u8);
arrival_owed[idx] = ARRIVAL_RESENDS;
let _ = conn.send_datagram(ev.encode().to_vec().into());
continue;
}
let _ = conn.send_datagram(ev.encode().to_vec().into());
}
_ = refresh.tick() => {
for idx in 0..MAX_PADS {
// Re-send an owed kind declaration (independent of whether the pad has state
// yet — it may be idle-but-connected). Idempotent on the host.
if arrival_owed[idx] > 0 {
if let Some(kind) = arrival[idx] {
arrival_owed[idx] -= 1;
let arr = crate::input::InputEvent {
kind: InputKind::GamepadArrival,
_pad: [0; 3],
code: kind as u32,
x: 0,
y: 0,
flags: idx as u32,
};
let _ = conn.send_datagram(arr.encode().to_vec().into());
} else {
arrival_owed[idx] = 0;
}
}
if let Some(snap) = pads[idx].as_mut() {
seq[idx] = seq[idx].wrapping_add(1);
snap.seq = seq[idx];
let _ = conn.send_datagram(snap.to_event().encode().to_vec().into());
} else if remove_owed[idx] > 0 {
// Idempotent removal re-send with a fresh seq (the host drops it as a
// no-op once the pad is already gone, but a re-plug's later snapshot
// still wins by seq).
remove_owed[idx] -= 1;
seq[idx] = seq[idx].wrapping_add(1);
let rem = crate::input::InputEvent {
kind: InputKind::GamepadRemove,
_pad: [0; 3],
code: 0,
x: 0,
y: 0,
flags: crate::input::encode_gamepad_remove(idx as u8, seq[idx]),
};
let _ = conn.send_datagram(rem.encode().to_vec().into());
}
}
}
}
}
}
+115 -4
View File
@@ -145,6 +145,11 @@ pub async fn run(
let Some(cmd) = cmd else { break }; // NativeClient dropped let Some(cmd) = cmd else { break }; // NativeClient dropped
match cmd { match cmd {
ClipCommand::Fetch { xfer_id, seq, file_index, mime } => { ClipCommand::Fetch { xfer_id, seq, file_index, mime } => {
// Prune finished fetches first: a completed/failed/timed-out fetch task
// drops its cancel receiver, so its sender reads closed. Without this
// the map grew by one dead sender per paste for the whole session (only
// an explicit Cancel ever removed entries).
fetch_cancels.retain(|_, tx| !tx.is_closed());
let (cancel_tx, cancel_rx) = oneshot::channel(); let (cancel_tx, cancel_rx) = oneshot::channel();
fetch_cancels.insert(xfer_id, cancel_tx); fetch_cancels.insert(xfer_id, cancel_tx);
let conn = conn.clone(); let conn = conn.clone();
@@ -153,14 +158,39 @@ pub async fn run(
tokio::spawn(run_outbound_fetch(conn, xfer_id, req, events, cancel_rx)); tokio::spawn(run_outbound_fetch(conn, xfer_id, req, events, cancel_rx));
} }
ClipCommand::Serve { req_id, bytes, last } => { ClipCommand::Serve { req_id, bytes, last } => {
serve_bufs.entry(req_id).or_default().extend_from_slice(&bytes); // Gate on a genuinely parked fetch (the waiter registers before the
// FetchRequest event is emitted, so a live serve always finds it):
// bytes served under a stale/unknown/cancelled req_id would otherwise
// pool here for the whole session with Ok returned for every chunk.
if !serve_waiters.lock().unwrap().contains_key(&req_id) {
serve_bufs.remove(&req_id);
} else {
let buf = serve_bufs.entry(req_id).or_default();
if buf.len().saturating_add(bytes.len()) > CLIP_FETCH_CAP {
// The requester bounds its read at CLIP_FETCH_CAP — anything
// larger is memory burned toward a guaranteed peer rejection.
// Fail the transfer NOW: peer reads UNAVAILABLE, embedder gets
// an Error instead of silent Ok-per-chunk.
serve_bufs.remove(&req_id);
if let Some(tx) = serve_waiters.lock().unwrap().remove(&req_id) {
let _ = tx.send(None);
}
let _ = events.try_send(ClipEventCore::Error {
id: req_id,
code: PunktfunkStatus::InvalidArg as i32,
});
} else {
buf.extend_from_slice(&bytes);
if last { if last {
let full = serve_bufs.remove(&req_id).unwrap_or_default(); let full = serve_bufs.remove(&req_id).unwrap_or_default();
if let Some(tx) = serve_waiters.lock().unwrap().remove(&req_id) { if let Some(tx) = serve_waiters.lock().unwrap().remove(&req_id)
{
let _ = tx.send(Some(full)); let _ = tx.send(Some(full));
} }
} }
} }
}
}
ClipCommand::Cancel { id } => { ClipCommand::Cancel { id } => {
// Route to whichever table owns the id (they are disjoint by the high bit). // Route to whichever table owns the id (they are disjoint by the high bit).
if let Some(tx) = fetch_cancels.remove(&id) { if let Some(tx) = fetch_cancels.remove(&id) {
@@ -224,8 +254,27 @@ async fn serve_inbound(
return; return;
} }
match rx.await { // Overall stall bound (§3.4) — the inbound mirror of `run_outbound_fetch`'s: an embedder
Ok(Some(bytes)) => { // that never answers must not hold the waiter, this task, and the accepted bi-stream open
// for the rest of the session (~100 unanswered pastes exhaust the connection's bidi-stream
// budget and every later host paste stalls). `send.stopped()` additionally wakes us the
// moment the peer gives up on its side of the fetch.
let answer = tokio::select! {
r = rx => r.ok().flatten(),
_ = send.stopped() => {
let _ = events.try_send(ClipEventCore::Cancelled { id: req_id });
None
}
_ = tokio::time::sleep(std::time::Duration::from_secs(FETCH_STALL_SECS)) => {
let _ = events.try_send(ClipEventCore::Cancelled { id: req_id });
None
}
};
// Idempotent — the answering/denying paths already removed it; the stall paths must.
waiters.lock().unwrap().remove(&req_id);
match answer {
Some(bytes) => {
if clipstream::write_fetch_hdr( if clipstream::write_fetch_hdr(
&mut send, &mut send,
&ClipFetchHdr { &ClipFetchHdr {
@@ -302,3 +351,65 @@ async fn run_outbound_fetch(
} }
} }
} }
#[cfg(test)]
mod tests {
use super::*;
use crate::quic::test_util::connect_pair;
use crate::quic::CLIP_FILE_INDEX_NONE;
/// A serve chunk that alone breaches the requester-side CLIP_FETCH_CAP must fail the
/// transfer immediately — Error to the embedder, UNAVAILABLE to the peer — instead of
/// accumulating Ok-per-chunk toward a guaranteed peer-side rejection. Also pins the
/// waiter-membership gate: the serve only accumulated because the fetch was parked.
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn oversized_serve_fails_the_transfer_instead_of_buffering() {
let (_s, _c, host_conn, client_conn) = connect_pair().await;
let (ev_tx, ev_rx) = std::sync::mpsc::sync_channel(16);
let (cmd_tx, cmd_rx) = tokio::sync::mpsc::unbounded_channel();
tokio::spawn(run(client_conn, ev_tx, cmd_rx));
// Host pastes: it opens a fetch bi-stream toward the client.
let req = ClipFetch {
seq: 1,
file_index: CLIP_FILE_INDEX_NONE,
mime: "text/plain;charset=utf-8".into(),
};
let (_send, mut recv) = clipstream::open_fetch(&host_conn, &req).await.unwrap();
// The client core surfaces the FetchRequest (the waiter is parked now).
let (req_id, ev_rx) = tokio::task::spawn_blocking(move || {
match ev_rx
.recv_timeout(std::time::Duration::from_secs(5))
.unwrap()
{
ClipEventCore::FetchRequest { req_id, .. } => (req_id, ev_rx),
other => panic!("expected FetchRequest, got {other:?}"),
}
})
.await
.unwrap();
cmd_tx
.send(ClipCommand::Serve {
req_id,
bytes: vec![0u8; CLIP_FETCH_CAP + 1],
last: false,
})
.unwrap();
let ev = tokio::task::spawn_blocking(move || {
ev_rx
.recv_timeout(std::time::Duration::from_secs(5))
.unwrap()
})
.await
.unwrap();
match ev {
ClipEventCore::Error { id, .. } => assert_eq!(id, req_id),
other => panic!("expected Error, got {other:?}"),
}
// The peer's read side sees the transfer refused, not a hang.
let hdr = clipstream::read_fetch_hdr(&mut recv).await.unwrap();
assert_eq!(hdr.status, CLIP_FETCH_UNAVAILABLE);
}
}
+96
View File
@@ -548,4 +548,100 @@ mod tests {
Some(SteamController) Some(SteamController)
); );
} }
#[test]
fn compositor_pref_wire_and_names() {
for p in [
CompositorPref::Auto,
CompositorPref::Kwin,
CompositorPref::Wlroots,
CompositorPref::Mutter,
CompositorPref::Gamescope,
] {
assert_eq!(CompositorPref::from_u8(p.to_u8()), p);
assert_eq!(CompositorPref::from_name(p.as_str()), Some(p));
}
// Aliases + unknowns.
assert_eq!(CompositorPref::from_name("KDE"), Some(CompositorPref::Kwin));
assert_eq!(
CompositorPref::from_name("sway"),
Some(CompositorPref::Wlroots)
);
assert_eq!(CompositorPref::from_name("nope"), None);
// Unknown wire byte degrades to Auto (forward-compatible).
assert_eq!(CompositorPref::from_u8(200), CompositorPref::Auto);
}
#[test]
fn gamepad_pref_wire_and_names() {
for p in [
GamepadPref::Auto,
GamepadPref::Xbox360,
GamepadPref::DualSense,
GamepadPref::XboxOne,
GamepadPref::DualShock4,
GamepadPref::SteamController,
GamepadPref::SteamDeck,
GamepadPref::DualSenseEdge,
GamepadPref::SwitchPro,
GamepadPref::SteamController2,
GamepadPref::SteamController2Puck,
] {
assert_eq!(GamepadPref::from_u8(p.to_u8()), p);
assert_eq!(GamepadPref::from_name(p.as_str()), Some(p));
}
// Every wire byte 0..=10 is assigned, distinct, and pinned (forward-compat with peers
// that only know a prefix of the range).
for (v, p) in [
(0, GamepadPref::Auto),
(1, GamepadPref::Xbox360),
(2, GamepadPref::DualSense),
(3, GamepadPref::XboxOne),
(4, GamepadPref::DualShock4),
(5, GamepadPref::SteamController),
(6, GamepadPref::SteamDeck),
(7, GamepadPref::DualSenseEdge),
(8, GamepadPref::SwitchPro),
(9, GamepadPref::SteamController2),
(10, GamepadPref::SteamController2Puck),
] {
assert_eq!(p.to_u8(), v);
assert_eq!(GamepadPref::from_u8(v), p);
}
// The next unassigned byte degrades to Auto today; assigning it later must update this.
assert_eq!(GamepadPref::from_u8(11), GamepadPref::Auto);
// Aliases + unknowns.
assert_eq!(GamepadPref::from_name("PS5"), Some(GamepadPref::DualSense));
assert_eq!(GamepadPref::from_name("x360"), Some(GamepadPref::Xbox360));
assert_eq!(GamepadPref::from_name("ps4"), Some(GamepadPref::DualShock4));
assert_eq!(GamepadPref::from_name("DS4"), Some(GamepadPref::DualShock4));
assert_eq!(
GamepadPref::from_name("edge"),
Some(GamepadPref::DualSenseEdge)
);
assert_eq!(
GamepadPref::from_name("Switch-Pro"),
Some(GamepadPref::SwitchPro)
);
assert_eq!(
GamepadPref::from_name("ibex"),
Some(GamepadPref::SteamController2)
);
assert_eq!(
GamepadPref::from_name("sc2"),
Some(GamepadPref::SteamController2)
);
assert_eq!(
GamepadPref::from_name("sc2puck"),
Some(GamepadPref::SteamController2Puck)
);
assert_eq!(
GamepadPref::from_name("xbox-one"),
Some(GamepadPref::XboxOne)
);
assert_eq!(GamepadPref::from_name("series"), Some(GamepadPref::XboxOne));
assert_eq!(GamepadPref::from_name("nope"), None);
// Unknown wire byte degrades to Auto (forward-compatible).
assert_eq!(GamepadPref::from_u8(200), GamepadPref::Auto);
}
} }
+19
View File
@@ -86,8 +86,18 @@ impl ErasureCoder for Gf8Coder {
// No FEC: every original must already be present. // No FEC: every original must already be present.
return collect_originals(received, data_count); return collect_originals(received, data_count);
} }
// Same (k, m)-keyed cache as `encode_into`: a fresh ReedSolomon per lossy block costs a
// full generator build (k×2k Gauss-Jordan + total×k×k multiply) on the real-time pump
// thread AND forfeits the instance's decode-matrix cache for stable loss patterns.
let mut guard = self.rs.lock().unwrap_or_else(|p| p.into_inner());
let cached =
matches!(&*guard, Some((ck, cm, _)) if *ck == data_count && *cm == recovery_count);
if !cached {
let rs = ReedSolomon::new(data_count, recovery_count) let rs = ReedSolomon::new(data_count, recovery_count)
.map_err(|_| FecError::Config("invalid GF(2^8) shard counts"))?; .map_err(|_| FecError::Config("invalid GF(2^8) shard counts"))?;
*guard = Some((data_count, recovery_count, rs));
}
let rs = &guard.as_ref().expect("cache populated above").2;
rs.reconstruct_data(received) rs.reconstruct_data(received)
.map_err(|_| FecError::Backend("gf8 reconstruct"))?; .map_err(|_| FecError::Backend("gf8 reconstruct"))?;
collect_originals(received, data_count) collect_originals(received, data_count)
@@ -116,8 +126,17 @@ impl ErasureCoder for Gf8Coder {
for &(j, bytes) in recovery { for &(j, bytes) in recovery {
received[data_count + j] = Some(bytes.to_vec()); received[data_count + j] = Some(bytes.to_vec());
} }
// Cache the codec by (k, m) exactly as `encode_into`/`reconstruct` do (see the note
// there) — this path runs per lossy block on the pump thread.
let mut guard = self.rs.lock().unwrap_or_else(|p| p.into_inner());
let cached =
matches!(&*guard, Some((ck, cm, _)) if *ck == data_count && *cm == recovery_count);
if !cached {
let rs = ReedSolomon::new(data_count, recovery_count) let rs = ReedSolomon::new(data_count, recovery_count)
.map_err(|_| FecError::Config("invalid GF(2^8) shard counts"))?; .map_err(|_| FecError::Config("invalid GF(2^8) shard counts"))?;
*guard = Some((data_count, recovery_count, rs));
}
let rs = &guard.as_ref().expect("cache populated above").2;
rs.reconstruct_data(&mut received) rs.reconstruct_data(&mut received)
.map_err(|_| FecError::Backend("gf8 reconstruct"))?; .map_err(|_| FecError::Backend("gf8 reconstruct"))?;
for (i, h) in have.iter().enumerate() { for (i, h) in have.iter().enumerate() {
+20 -1
View File
@@ -16,6 +16,17 @@
//! (recv → open → reorder → FEC recover → reassemble) state machines. //! (recv → open → reorder → FEC recover → reassemble) state machines.
//! - [`transport`] — pluggable packet I/O (in-process loopback for tests; UDP for real). //! - [`transport`] — pluggable packet I/O (in-process loopback for tests; UDP for real).
//! - [`abi`] — the `extern "C"` surface and `cbindgen`-generated `punktfunk_core.h`. //! - [`abi`] — the `extern "C"` surface and `cbindgen`-generated `punktfunk_core.h`.
//! - [`config`] / [`error`] / [`stats`] — session configuration, the shared error/status
//! vocabulary, and the counters snapshot.
//! - [`input`] — the wire input-event vocabulary (keyboard/mouse/touch, gamepad snapshots).
//! - [`reject`] — typed application-close rejection codes · [`reanchor`] — the post-loss
//! freeze-until-reanchor client gate · [`render_scale`] — the shared render-scale setting ·
//! [`audio`] — Opus PCM decode for C-ABI embedders · [`wol`] — Wake-on-LAN.
//! - `quic` (feature `quic`) — the punktfunk/1 control plane: handshake, typed control
//! messages, pairing (SPAKE2), the datagram plane codecs, and clock sync. With it come
//! `client` (the embeddable NativeClient worker), `abr` (the adaptive-bitrate
//! controller), and `clipboard` (the shared-clipboard transport task). `tls`
//! (feature `tls`) — the pinned-fingerprint certificate verifier.
//! //!
//! ## Threading contract //! ## Threading contract
//! //!
@@ -83,7 +94,15 @@ pub use stats::Stats;
/// `punktfunk_connection_clipboard_{control,offer,fetch,serve,cancel}` + /// `punktfunk_connection_clipboard_{control,offer,fetch,serve,cancel}` +
/// `punktfunk_connection_next_clipboard`. Additive; the wire grows only backward-compatible control /// `punktfunk_connection_next_clipboard`. Additive; the wire grows only backward-compatible control
/// messages (0x40-0x44) and a new `Welcome::host_caps` bit, so [`WIRE_VERSION`] is unchanged. /// messages (0x40-0x44) and a new `Welcome::host_caps` bit, so [`WIRE_VERSION`] is unchanged.
pub const ABI_VERSION: u32 = 8; /// v9: `PunktfunkFrame` grew `received_ns` — the reassembly-completion receipt stamp, so
/// embedders stop stamping receipt at the hand-off pull (which folds the pre-decode queue wait
/// into apparent network latency). Struct-size change on the frame poll surface = a hard ABI
/// break for embedders reading `PunktfunkFrame`; nothing on the wire moved, so [`WIRE_VERSION`]
/// is unchanged.
/// v10: added `punktfunk_connection_clock_offset_now_ns` — the LIVE (mid-stream re-synced)
/// clock offset ongoing latency math must use; the connect-time getter stays frozen by
/// contract. Additive, client-local — no wire change, so [`WIRE_VERSION`] is unchanged.
pub const ABI_VERSION: u32 = 10;
/// The punktfunk/1 **wire** version — what `Hello`/`Welcome` carry and hosts equality-check. /// The punktfunk/1 **wire** version — what `Hello`/`Welcome` carry and hosts equality-check.
/// Deliberately its own constant: [`ABI_VERSION`] tracks the embeddable **C surface** /// Deliberately its own constant: [`ABI_VERSION`] tracks the embeddable **C surface**
+263 -4
View File
@@ -39,10 +39,58 @@ pub struct Packetizer {
/// DATA shards before any block's parity — all blocks' parity must stay alive until the /// DATA shards before any block's parity — all blocks' parity must stay alive until the
/// frame's second emission pass. /// frame's second emission pass.
recovery: Vec<Vec<Vec<u8>>>, recovery: Vec<Vec<Vec<u8>>>,
/// The peer's per-block `data + recovery` acceptance ceiling, frozen from the **negotiated**
/// config exactly as the far side derives it in [`ReassemblerLimits::from_config`]. Adaptive
/// FEC moves `fec.fec_percent` live ([`set_fec_percent`](Self::set_fec_percent)) but the
/// receiver's ceiling is computed once at session construction and never re-derived, so parity
/// must be clamped against this or a raised percentage puts blocks over the far side's bound —
/// where every packet of the block is dropped wholesale, the frame never completes, and the
/// resulting loss pushes adaptive FEC *higher*. See the `recovery_for` clamp in `packetize_each`.
max_total_shards: usize,
/// The peer's per-frame block ceiling, mirroring [`ReassemblerLimits::from_config`]'s
/// `max_blocks` — the streamed path's bound on how many sentinel blocks it may emit (a
/// streamed AU's size isn't known up front, so this is the only pre-emission guard against
/// producing a frame the receiver must reject).
max_blocks: usize,
}
/// One in-progress **streamed** access unit (design/nvenc-subframe-slice-output.md Phase 2):
/// the caller feeds encoder chunks in as they exist ([`Packetizer::push_streamed`]) and every
/// completed `max_data_per_block × shard_payload` block leaves under SENTINEL headers
/// (`frame_bytes = 0`, `block_count = 0` — "not final yet") before the AU's total size is
/// known; [`Packetizer::finish_streamed`] seals the tail block with the real totals and
/// `FLAG_EOF`. Only ever sent to a peer that advertised
/// [`crate::quic::VIDEO_CAP_STREAMED_AU`].
pub struct StreamedAu {
frame_index: u32,
pts_ns: u64,
user_flags: u32,
/// Bytes not yet sealed into a block. Kept ≤ one block by `push_streamed` (it flushes only
/// while STRICTLY more than a block is buffered, so the final block always has ≥ 1 byte and
/// a sentinel block is never retroactively the frame's last).
pending: Vec<u8>,
/// Sentinel blocks already emitted.
blocks_out: u16,
/// Total AU bytes accumulated so far (`pending` included).
total_bytes: u64,
/// The frame's first packet (block 0, shard 0 — carries `FLAG_SOF`) has been emitted.
opened: bool,
}
impl StreamedAu {
/// The wire frame index this AU is sealed with (the caller's RFI bookkeeping domain).
pub fn frame_index(&self) -> u32 {
self.frame_index
}
} }
impl Packetizer { impl Packetizer {
pub fn new(config: &Config) -> Self { pub fn new(config: &Config) -> Self {
let max_data = config.fec.max_data_per_block as usize;
let total_data_max = config
.max_frame_bytes
.div_ceil(config.shard_payload.max(1))
.max(1);
Packetizer { Packetizer {
next_frame_index: 0, next_frame_index: 0,
next_probe_index: 0, next_probe_index: 0,
@@ -52,6 +100,10 @@ impl Packetizer {
version: config.phase as u8, version: config.phase as u8,
tail: Vec::new(), tail: Vec::new(),
recovery: Vec::new(), recovery: Vec::new(),
// Mirrors `ReassemblerLimits::from_config` — keep the two in step.
max_total_shards: (max_data + config.fec.recovery_for(max_data))
.min(config.fec.scheme.max_total_shards()),
max_blocks: total_data_max.div_ceil(max_data).max(1),
} }
} }
@@ -173,6 +225,15 @@ impl Packetizer {
}; };
// Per-block shard geometry (deterministic — recomputed in both passes). // Per-block shard geometry (deterministic — recomputed in both passes).
let block_data_count = |b: usize| ((b + 1) * max_block).min(total_data) - b * max_block; let block_data_count = |b: usize| ((b + 1) * max_block).min(total_data) - b * max_block;
// Parity for a `k`-shard block: the configured percentage, clamped so the block's wire
// total never exceeds what the peer will accept (see `max_total_shards`). The clamp only
// binds on blocks near `max_data_per_block`; smaller blocks keep the full adaptive range,
// so raising FEC still buys real protection wherever there is headroom. Bound as locals,
// not as a `&self` method: `emit_one` below would otherwise capture all of `self` and
// collide with the `&mut self.recovery[b]` parity borrow.
let (fec, max_total_shards) = (self.fec, self.max_total_shards);
let recovery_for =
move |k: usize| fec.recovery_for(k).min(max_total_shards.saturating_sub(k));
// One parity pool per block, reused across frames (steady-state zero-alloc). // One parity pool per block, reused across frames (steady-state zero-alloc).
if self.recovery.len() < block_count { if self.recovery.len() < block_count {
@@ -183,7 +244,7 @@ impl Packetizer {
let mut total_recovery = 0usize; let mut total_recovery = 0usize;
for b in 0..block_count { for b in 0..block_count {
let k = block_data_count(b); let k = block_data_count(b);
let m = self.fec.recovery_for(k); let m = recovery_for(k);
if k + m > u16::MAX as usize { if k + m > u16::MAX as usize {
return Err(PunktfunkError::Unsupported("block shard count exceeds u16")); return Err(PunktfunkError::Unsupported("block shard count exceeds u16"));
} }
@@ -204,7 +265,7 @@ impl Packetizer {
block_index: b as u16, block_index: b as u16,
block_count: block_count as u16, block_count: block_count as u16,
data_shards: k as u16, data_shards: k as u16,
recovery_shards: self.fec.recovery_for(k) as u16, recovery_shards: recovery_for(k) as u16,
shard_index: shard_index as u16, shard_index: shard_index as u16,
shard_bytes: payload as u16, shard_bytes: payload as u16,
magic: PUNKTFUNK_MAGIC, magic: PUNKTFUNK_MAGIC,
@@ -223,7 +284,7 @@ impl Packetizer {
// This block's data shards: references into `frame` (plus the staged tail). // This block's data shards: references into `frame` (plus the staged tail).
let data_shards: Vec<&[u8]> = (first..first + k).map(shard_at).collect(); let data_shards: Vec<&[u8]> = (first..first + k).map(shard_at).collect();
let recovery_count = self.fec.recovery_for(k); let recovery_count = recovery_for(k);
coder.encode_into(&data_shards, recovery_count, &mut self.recovery[b])?; coder.encode_into(&data_shards, recovery_count, &mut self.recovery[b])?;
for (shard_index, body) in data_shards.iter().enumerate() { for (shard_index, body) in data_shards.iter().enumerate() {
@@ -242,7 +303,7 @@ impl Packetizer {
let mut parity_left = total_recovery; let mut parity_left = total_recovery;
for b in 0..block_count { for b in 0..block_count {
let k = block_data_count(b); let k = block_data_count(b);
let recovery_count = self.fec.recovery_for(k); let recovery_count = recovery_for(k);
for r in 0..recovery_count { for r in 0..recovery_count {
parity_left -= 1; parity_left -= 1;
let mut flags = FLAG_PIC; let mut flags = FLAG_PIC;
@@ -256,4 +317,202 @@ impl Packetizer {
self.next_seq = next_seq; self.next_seq = next_seq;
Ok(()) Ok(())
} }
/// Open a **streamed** access unit (see [`StreamedAu`]). `frame_index` follows the same
/// contract as [`packetize_each`](Self::packetize_each): `Some(i)` = the caller owns the
/// video numbering; `None` draws from the internal counter.
pub fn begin_streamed(
&mut self,
pts_ns: u64,
user_flags: u32,
frame_index: Option<u32>,
) -> StreamedAu {
let frame_index = frame_index.unwrap_or_else(|| {
let i = self.next_frame_index;
self.next_frame_index = i.wrapping_add(1);
i
});
StreamedAu {
frame_index,
pts_ns,
user_flags,
pending: Vec::new(),
blocks_out: 0,
total_bytes: 0,
opened: false,
}
}
/// Feed one encoder chunk into a streamed AU, emitting every block that COMPLETES under
/// sentinel headers (`frame_bytes = 0`, `block_count = 0`, exactly `max_data_per_block`
/// data shards — the receiver's offset formula needs no total). Flushes only while
/// STRICTLY more than one block is buffered, so the frame's last block — whose header must
/// carry the real totals — is never emitted here. Packets reach `emit` in wire order (the
/// caller's nonce order), data then parity per block.
pub fn push_streamed(
&mut self,
au: &mut StreamedAu,
chunk: &[u8],
coder: &dyn ErasureCoder,
mut emit: impl FnMut(&PacketHeader, &[u8]) -> Result<()>,
) -> Result<()> {
au.total_bytes += chunk.len() as u64;
au.pending.extend_from_slice(chunk);
let block_bytes = self.fec.max_data_per_block as usize * self.shard_payload;
while au.pending.len() > block_bytes {
// Room for this sentinel block AND the final block after it, within the peer's
// per-frame block ceiling and the u16 wire field.
if au.blocks_out as usize + 2 > self.max_blocks.min(u16::MAX as usize) {
return Err(PunktfunkError::Unsupported(
"streamed AU exceeds the negotiated max_frame_bytes",
));
}
let sof = !au.opened;
let (bi, pts, uf) = (au.blocks_out, au.pts_ns, au.user_flags);
let fi = au.frame_index;
self.emit_streamed_block(
fi,
pts,
uf,
bi,
&au.pending[..block_bytes],
0,
0,
sof,
false,
coder,
&mut emit,
)?;
au.pending.drain(..block_bytes);
au.blocks_out += 1;
au.opened = true;
}
Ok(())
}
/// Close a streamed AU: seal the final block — its headers carry the REAL
/// `frame_bytes`/`block_count`, which retro-validate the whole frame at the receiver — with
/// `FLAG_EOF` on the last emitted packet. An empty AU degenerates to today's single
/// zero-padded-shard frame (`block_count = 1`, never a sentinel).
pub fn finish_streamed(
&mut self,
au: StreamedAu,
coder: &dyn ErasureCoder,
mut emit: impl FnMut(&PacketHeader, &[u8]) -> Result<()>,
) -> Result<()> {
let frame_bytes = u32::try_from(au.total_bytes)
.map_err(|_| PunktfunkError::Unsupported("streamed AU exceeds u32 bytes"))?;
let block_count = au.blocks_out + 1;
self.emit_streamed_block(
au.frame_index,
au.pts_ns,
au.user_flags,
au.blocks_out,
&au.pending,
frame_bytes,
block_count,
!au.opened,
true,
coder,
&mut emit,
)
}
/// Seal ONE streamed block (data shards + its parity, in wire order). Sentinel blocks pass
/// `frame_bytes = 0` / `block_count = 0`; the final block passes the real totals. `sof`
/// marks the frame's very first packet, `eof` its very last.
#[allow(clippy::too_many_arguments)]
fn emit_streamed_block(
&mut self,
frame_index: u32,
pts_ns: u64,
user_flags: u32,
block_index: u16,
bytes: &[u8],
frame_bytes: u32,
block_count: u16,
sof: bool,
eof: bool,
coder: &dyn ErasureCoder,
emit: &mut impl FnMut(&PacketHeader, &[u8]) -> Result<()>,
) -> Result<()> {
let payload = self.shard_payload;
if payload > u16::MAX as usize {
return Err(PunktfunkError::InvalidArg("shard_payload exceeds u16"));
}
// At least one (zero-padded) data shard even for an empty final block (empty AU).
let k = bytes.len().div_ceil(payload).max(1);
let m = self
.fec
.recovery_for(k)
.min(self.max_total_shards.saturating_sub(k));
if k + m > u16::MAX as usize {
return Err(PunktfunkError::Unsupported("block shard count exceeds u16"));
}
// Stage the one possibly-partial (or empty-frame) shard in the zero-padded scratch.
let full_shards = bytes.len() / payload;
self.tail.clear();
self.tail.resize(payload, 0);
let rem = bytes.len() % payload;
if rem > 0 {
self.tail[..rem].copy_from_slice(&bytes[full_shards * payload..]);
}
let tail = &self.tail;
let shard_at = |s: usize| -> &[u8] {
if s < full_shards {
&bytes[s * payload..(s + 1) * payload]
} else {
tail.as_slice()
}
};
let data_shards: Vec<&[u8]> = (0..k).map(shard_at).collect();
if self.recovery.is_empty() {
self.recovery.push(Vec::new());
}
coder.encode_into(&data_shards, m, &mut self.recovery[0])?;
let mut next_seq = self.next_seq;
let mut emit_one = |next_seq: &mut u32, shard_index: usize, body: &[u8], flags: u8| {
let seq = *next_seq;
*next_seq = next_seq.wrapping_add(1);
let hdr = PacketHeader {
pts_ns,
frame_index,
stream_seq: seq,
frame_bytes,
user_flags,
block_index,
block_count,
data_shards: k as u16,
recovery_shards: m as u16,
shard_index: shard_index as u16,
shard_bytes: payload as u16,
magic: PUNKTFUNK_MAGIC,
version: self.version,
fec_scheme: coder.scheme() as u8,
flags,
};
emit(&hdr, body)
};
for (shard_index, body) in data_shards.iter().enumerate() {
let mut flags = FLAG_PIC;
if sof && shard_index == 0 {
flags |= FLAG_SOF;
}
if eof && m == 0 && shard_index + 1 == k {
flags |= FLAG_EOF;
}
emit_one(&mut next_seq, shard_index, body, flags)?;
}
for r in 0..m {
let mut flags = FLAG_PIC;
if eof && r + 1 == m {
flags |= FLAG_EOF;
}
let body: &[u8] = &self.recovery[0][r];
emit_one(&mut next_seq, k + r, body, flags)?;
}
self.next_seq = next_seq;
Ok(())
}
} }
+177 -18
View File
@@ -70,7 +70,12 @@ struct BlockState {
} }
struct FrameBuf { struct FrameBuf {
/// Exact AU size. 0 = unknown: the frame was opened by a streamed-AU SENTINEL packet
/// ([`crate::quic::VIDEO_CAP_STREAMED_AU`]) and the final block's real totals haven't
/// arrived yet — the frame can't complete before they do (and retro-validate).
frame_bytes: usize, frame_bytes: usize,
/// Block count; 0 = unknown (sentinel-opened, totals not yet pinned). A legacy-opened
/// frame always has ≥ 1 here, so 0 doubles as the "unpinned streamed" marker.
block_count: usize, block_count: usize,
pts_ns: u64, pts_ns: u64,
user_flags: u32, user_flags: u32,
@@ -106,8 +111,19 @@ pub struct ReassemblerLimits {
impl ReassemblerLimits { impl ReassemblerLimits {
pub fn from_config(c: &Config) -> Self { pub fn from_config(c: &Config) -> Self {
let max_data = c.fec.max_data_per_block as usize; let max_data = c.fec.max_data_per_block as usize;
// Size the ceiling from the whole range adaptive FEC may reach, NOT from the percentage
// negotiated at session start: the sender moves `fec_percent` live (`Packetizer::
// set_fec_percent`, clamped to ≤ 90) and the wire is self-describing, so it never
// renegotiates. Deriving this from the start value made every packet of a large block
// fail the `total > max_total_shards` check once FEC ramped up — the block never
// accumulated a shard, the frame aged out, and the resulting loss drove FEC *higher*,
// wedging large frames at 100% loss exactly when FEC was meant to rescue the link. A
// current sender also clamps its side (`Packetizer::recovery_for`); this keeps an
// already-deployed sender that doesn't from wedging a current receiver. Still a hard
// pre-allocation bound against hostile headers — just the sender's clamp, not a stale
// snapshot of it.
let max_total = let max_total =
(max_data + c.fec.recovery_for(max_data)).min(c.fec.scheme.max_total_shards()); (max_data + (max_data * 90).div_ceil(100)).min(c.fec.scheme.max_total_shards());
let total_data = c.max_frame_bytes.div_ceil(c.shard_payload.max(1)).max(1); let total_data = c.max_frame_bytes.div_ceil(c.shard_payload.max(1)).max(1);
ReassemblerLimits { ReassemblerLimits {
shard_bytes: c.shard_payload, shard_bytes: c.shard_payload,
@@ -266,25 +282,49 @@ impl Reassembler {
|| total == 0 || total == 0
|| total > lim.max_total_shards || total > lim.max_total_shards
|| shard_index >= total || shard_index >= total
|| block_count == 0
|| block_count > lim.max_blocks
|| hdr.block_index as usize >= block_count
|| frame_bytes > lim.max_frame_bytes || frame_bytes > lim.max_frame_bytes
{ {
drop(stats); drop(stats);
return Ok(None); return Ok(None);
} }
// Derived-geometry firewall: every sender (our Packetizer, any version) slices a frame // Streamed-AU sentinel ([`crate::quic::VIDEO_CAP_STREAMED_AU`]): `block_count == 0` — a
// into consecutive blocks of exactly `max_data_per_block` data shards with only the LAST // value no legacy sender ever emits — marks a NON-FINAL block of an AU whose total size
// block smaller, and stamps the exact `frame_bytes` in every header. That makes every // doesn't exist yet (the host is still encoding its tail). The exact derived-geometry
// data shard's final AU offset computable on arrival — // check below can't run without a total, so a sentinel is bounded by the negotiated
// limits instead — and by construction: full-K exactly (the offset formula needs it),
// never the last block the limits allow (the real final block must still fit after it),
// and no total to lie about. The frame-wide exact check runs retroactively the moment
// the final block's real totals arrive (the pinning arm below).
let sentinel = block_count == 0;
let block_idx = hdr.block_index as usize;
// For a sentinel-opened frame the buffer must hold ANY final geometry the totals may
// later pin — the maximum the negotiated limits allow (the design's "allocate at
// max_frame_bytes"; the existing in-flight budget bounds the amplification).
let total_data_max = lim.max_frame_bytes.div_ceil(shard_bytes).max(1);
let total_data = frame_bytes.div_ceil(shard_bytes).max(1);
if sentinel {
if frame_bytes != 0
|| data_shards != lim.max_data_shards
|| block_idx + 1 >= lim.max_blocks
{
drop(stats);
return Ok(None);
}
} else {
if block_count > lim.max_blocks || block_idx >= block_count {
drop(stats);
return Ok(None);
}
// Derived-geometry firewall: every sender (our Packetizer, any version) slices a
// frame into consecutive blocks of exactly `max_data_per_block` data shards with
// only the LAST block smaller, and stamps the exact `frame_bytes` in every
// non-sentinel header. That makes every data shard's final AU offset computable on
// arrival —
// offset = (block_index × max_data_per_block + shard_index) × shard_bytes // offset = (block_index × max_data_per_block + shard_index) × shard_bytes
// — which is what lets shards land straight in the frame buffer below. Enforce the // — which is what lets shards land straight in the frame buffer below. Enforce the
// invariant so a header lying about its geometry is dropped instead of scribbling into // invariant so a header lying about its geometry is dropped instead of scribbling
// another shard's range. // into another shard's range.
let total_data = frame_bytes.div_ceil(shard_bytes).max(1);
let expect_blocks = total_data.div_ceil(lim.max_data_shards).max(1); let expect_blocks = total_data.div_ceil(lim.max_data_shards).max(1);
let block_idx = hdr.block_index as usize;
let expect_data_shards = if block_idx + 1 == expect_blocks { let expect_data_shards = if block_idx + 1 == expect_blocks {
total_data - (expect_blocks - 1) * lim.max_data_shards total_data - (expect_blocks - 1) * lim.max_data_shards
} else { } else {
@@ -294,6 +334,7 @@ impl Reassembler {
drop(stats); drop(stats);
return Ok(None); return Ok(None);
} }
}
let body = &pkt[HEADER_LEN..HEADER_LEN + shard_bytes]; let body = &pkt[HEADER_LEN..HEADER_LEN + shard_bytes];
// Route by index space: speed-test probe filler (FLAG_PROBE in user_flags) reassembles in // Route by index space: speed-test probe filler (FLAG_PROBE in user_flags) reassembles in
@@ -339,8 +380,13 @@ impl Reassembler {
} }
// First packet of a frame allocates its whole (zeroed) buffer, budget-gated; later // First packet of a frame allocates its whole (zeroed) buffer, budget-gated; later
// packets must agree with its geometry. // packets must agree with its geometry. A sentinel-opened (streamed) frame allocates at
let buf_len = total_data * shard_bytes; // the limits' maximum — its real size doesn't exist yet.
let buf_len = if sentinel {
total_data_max * shard_bytes
} else {
total_data * shard_bytes
};
let frame = match win.frames.entry(hdr.frame_index) { let frame = match win.frames.entry(hdr.frame_index) {
std::collections::hash_map::Entry::Occupied(e) => e.into_mut(), std::collections::hash_map::Entry::Occupied(e) => e.into_mut(),
std::collections::hash_map::Entry::Vacant(e) => { std::collections::hash_map::Entry::Vacant(e) => {
@@ -363,7 +409,66 @@ impl Reassembler {
}) })
} }
}; };
if frame.block_count != block_count || frame.frame_bytes != frame_bytes { if sentinel {
// Sentinel packets carry no totals to cross-check: while the frame is unpinned
// (`block_count == 0`) they match by construction, and once the totals are pinned —
// whichever packet arrived first, sentinel or final (reorder is normal) — a
// sentinel block must still be NON-final under them. Its full-K shape was already
// enforced by the firewall, which is exactly what the pinned geometry demands of
// every non-final block, and its shard offsets are within the pinned range by
// `block_idx + 1 < block_count`.
if frame.block_count != 0 && block_idx + 1 >= frame.block_count {
drop(stats);
return Ok(None);
}
} else if frame.block_count == 0 {
// A streamed frame meets its FINAL block's real totals: retro-validate every block
// the sentinels created against the exact geometry these totals derive (the same
// invariant the firewall enforces per-packet for legacy frames), then PIN them. A
// header that lies — totals under which an already-received sentinel block is
// out of range or not full-K — kills the WHOLE frame: its landed shards can't be
// trusted to be at the offsets this geometry means, so delivering any of it would
// hand the decoder spliced bytes.
let expect_blocks = total_data.div_ceil(lim.max_data_shards).max(1);
let final_k = total_data - (expect_blocks - 1) * lim.max_data_shards;
let lied = frame.blocks.iter().any(|(&bi, b)| {
let bi = bi as usize;
bi >= expect_blocks
|| (bi + 1 < expect_blocks && b.data_shards != lim.max_data_shards)
|| (bi + 1 == expect_blocks && b.data_shards != final_k)
});
if lied {
let mut f = win
.frames
.remove(&hdr.frame_index)
.expect("frame entry exists");
*in_flight_bytes -= f.buf.len();
// Remember the index (with its late-shard memory, exactly like an aged-out
// frame) so stragglers can't resurrect it, reclaim the parity buffers, and
// count the loss — the client's recovery request is the right outcome for a
// frame that was destroyed by a lying header.
win.completed.insert(
hdr.frame_index,
reconstructed_shards(&f.blocks, lim.max_data_shards),
);
for block in f.blocks.values_mut() {
for slot in block.recovery.iter_mut() {
if let Some(rb) = slot.take() {
if recovery_pool.len() < RECOVERY_POOL_MAX {
recovery_pool.push(rb);
}
}
}
}
if !is_probe {
StatsCounters::add(&stats.frames_dropped, 1);
}
drop(stats);
return Ok(None);
}
frame.frame_bytes = frame_bytes;
frame.block_count = block_count;
} else if frame.block_count != block_count || frame.frame_bytes != frame_bytes {
drop(stats); drop(stats);
return Ok(None); return Ok(None);
} }
@@ -371,6 +476,7 @@ impl Reassembler {
buf, buf,
blocks, blocks,
blocks_ok, blocks_ok,
block_count: frame_block_count,
.. ..
} = frame; } = frame;
@@ -391,6 +497,14 @@ impl Reassembler {
drop(stats); drop(stats);
return Ok(None); return Ok(None);
} }
// Defense-in-depth (2026-07 security review): the geometry invariants above guarantee
// every packet of a block agrees on K, so this can't fire today — but `have_data`
// indexing and the recovery-slot math below assume it, and an explicit check keeps a
// future firewall refactor from turning that assumption into an OOB panic.
if block.data_shards != data_shards {
drop(stats);
return Ok(None);
}
if block.done { if block.done {
// A data shard the parity reconstruct already restored (`!have_data`) was late, not // A data shard the parity reconstruct already restored (`!have_data`) was late, not
// lost — net it out of the `fec_recovered_shards` it was counted into (see the // lost — net it out of the `fec_recovered_shards` it was counted into (see the
@@ -474,8 +588,12 @@ impl Reassembler {
} }
} }
// Whole frame ready? // Whole frame ready? Judged against the FRAME's pinned block count, not this packet's
if *blocks_ok == block_count { // header — a streamed frame can complete on a reordered sentinel packet (header says 0)
// after the final block already pinned the real totals, and can never complete before
// they're pinned (`0` never equals a non-zero `blocks_ok`).
let block_count = *frame_block_count;
if block_count != 0 && *blocks_ok == block_count {
let mut done = win.frames.remove(&hdr.frame_index).unwrap(); let mut done = win.frames.remove(&hdr.frame_index).unwrap();
win.completed.insert( win.completed.insert(
hdr.frame_index, hdr.frame_index,
@@ -489,6 +607,7 @@ impl Reassembler {
pts_ns: done.pts_ns, pts_ns: done.pts_ns,
flags: done.user_flags, flags: done.user_flags,
complete: true, complete: true,
received_ns: 0, // stamped by Session::poll_frame at the session boundary
})); }));
} }
Ok(None) Ok(None)
@@ -506,6 +625,10 @@ impl Reassembler {
// The dropped frames' buffers (and their parity bufs) go back to the allocator, not the // The dropped frames' buffers (and their parity bufs) go back to the allocator, not the
// pool — a flush is the rare path. The budget resets with them. // pool — a flush is the rare path. The budget resets with them.
self.in_flight_bytes = 0; self.in_flight_bytes = 0;
// An aged-out partial parked for delivery is from the discarded past too — without this
// it survives `flush_backlog` and gets handed up as the first "frame" after the
// jump-to-live, exactly the stale content the flush existed to discard.
self.pending_partial = None;
} }
} }
@@ -579,7 +702,10 @@ impl ReassemblyWindow {
// where shards are missing (the codec's block walk skips zero windows). // where shards are missing (the codec's block walk skips zero windows).
// Newest-wins if several age out in one prune. Still counted dropped below. // Newest-wins if several age out in one prune. Still counted dropped below.
if let Some(sink) = partial_sink.as_deref_mut() { if let Some(sink) = partial_sink.as_deref_mut() {
if f.user_flags & USER_FLAG_CHUNK_ALIGNED != 0 { // `frame_bytes > 0` also excludes an UNPINNED streamed frame (its total is
// still the 0 sentinel value): truncating its max-sized buffer to 0 would
// deliver an empty "partial" — worse than the plain drop it gets instead.
if f.user_flags & USER_FLAG_CHUNK_ALIGNED != 0 && f.frame_bytes > 0 {
let mut buf = std::mem::take(&mut f.buf); let mut buf = std::mem::take(&mut f.buf);
buf.truncate(f.frame_bytes); buf.truncate(f.frame_bytes);
let newer = sink let newer = sink
@@ -592,6 +718,7 @@ impl ReassemblyWindow {
pts_ns: f.pts_ns, pts_ns: f.pts_ns,
flags: f.user_flags, flags: f.user_flags,
complete: false, complete: false,
received_ns: 0, // stamped by Session::poll_frame at the session boundary
}); });
} }
} }
@@ -632,3 +759,35 @@ impl ReassemblyWindow {
} }
} }
} }
#[cfg(test)]
mod reset_tests {
use super::*;
/// `flush_backlog` discards the past wholesale — an aged-out partial parked for delivery is
/// part of that past and must not survive [`Reassembler::reset`] to be handed up as the
/// first "frame" after a jump-to-live.
#[test]
fn reset_drops_a_parked_partial() {
let mut r = Reassembler::new(ReassemblerLimits {
shard_bytes: 64,
max_data_shards: 8,
max_total_shards: 16,
max_blocks: 4,
max_frame_bytes: 4096,
});
r.pending_partial = Some(Frame {
data: vec![0u8; 64],
frame_index: 7,
pts_ns: 1,
flags: 0,
complete: false,
received_ns: 0,
});
r.reset();
assert!(
r.take_partial().is_none(),
"a pre-flush partial must not survive reset()"
);
}
}
+460
View File
@@ -579,3 +579,463 @@ fn rejects_wrong_shard_bytes_and_oversized_frame() {
.is_none()); .is_none());
assert_eq!(stats.snapshot().packets_dropped, 1); assert_eq!(stats.snapshot().packets_dropped, 1);
} }
/// Adaptive FEC raises `fec_percent` mid-session while the receiver's per-block acceptance
/// ceiling is frozen at session construction and never renegotiated. A maximal block must
/// therefore still land: the sender clamps its parity to the ceiling
/// (`Packetizer::recovery_for`), and the receiver sizes that ceiling from the whole clamp range
/// rather than the start percentage. Regression guard for the wedge this caused — every packet
/// of a large block failing `total > max_total_shards`, so the frame never completed and the
/// resulting loss drove adaptive FEC higher still.
#[test]
fn adaptive_fec_ramp_keeps_maximal_blocks_within_the_peers_ceiling() {
let cfg = e2e_config(FecScheme::Gf16, 10);
let coder = coder_for(FecScheme::Gf16);
let lim = ReassemblerLimits::from_config(&cfg);
let mut pk = Packetizer::new(&cfg);
// Ramp far past the negotiated 10% — exactly what `apply_fec_target` does under loss.
pk.set_fec_percent(50);
// A frame of full `max_data_per_block` blocks: where the ceiling actually binds.
let frame_len = cfg.shard_payload * cfg.fec.max_data_per_block as usize * 2;
let src: Vec<u8> = (0..frame_len).map(|i| (i * 131 + 7) as u8).collect();
let pkts = pk.packetize(&src, 1, 0, coder.as_ref()).unwrap();
let k = cfg.fec.max_data_per_block as usize;
let mut clamped = false;
for p in &pkts {
let hdr = PacketHeader::read_from_bytes(&p[..HEADER_LEN]).unwrap();
let total = hdr.data_shards as usize + hdr.recovery_shards as usize;
assert!(
total <= lim.max_total_shards,
"block total {total} exceeds the peer's ceiling {} — every packet of this block \
would be dropped",
lim.max_total_shards
);
// The unclamped 50% would put 2 parity on a full block; the negotiated 10% ceiling
// leaves room for 1. Proves the clamp actually bound rather than passing vacuously.
if hdr.data_shards as usize == k {
assert!(
(hdr.recovery_shards as usize) < cfg.fec.recovery_for(k).max(1) + 1,
"parity must be clamped to the peer's ceiling"
);
clamped = true;
}
}
assert!(clamped, "test must exercise a maximal block");
// And the frame still reassembles byte-identically.
let mut r = Reassembler::new(lim);
let stats = StatsCounters::default();
let mut got = None;
for p in &pkts {
if let Some(f) = r.push(p, coder.as_ref(), &stats).unwrap() {
got = Some(f);
}
}
assert_eq!(
got.expect("frame must complete after an adaptive-FEC ramp")
.data,
src
);
}
// ---------------------------------------------------------------------------
// Streamed access units (VIDEO_CAP_STREAMED_AU — nvenc-subframe-slice-output.md Phase 2)
// ---------------------------------------------------------------------------
/// Packetize one streamed AU from `chunks` via begin/push/finish, returning the emitted wire
/// packets (header ++ shard) and the concatenated source bytes.
fn streamed_packets(
scheme: FecScheme,
fec_percent: u8,
chunks: &[&[u8]],
) -> (Vec<Vec<u8>>, Vec<u8>) {
let cfg = e2e_config(scheme, fec_percent);
let coder = coder_for(scheme);
let mut pk = Packetizer::new(&cfg);
let mut au = pk.begin_streamed(12345, 0, Some(0));
let mut pkts: Vec<Vec<u8>> = Vec::new();
let mut src = Vec::new();
for c in chunks {
src.extend_from_slice(c);
pk.push_streamed(&mut au, c, coder.as_ref(), |h: &PacketHeader, b: &[u8]| {
let mut p = Vec::with_capacity(HEADER_LEN + b.len());
p.extend_from_slice(h.as_bytes());
p.extend_from_slice(b);
pkts.push(p);
Ok(())
})
.unwrap();
}
pk.finish_streamed(au, coder.as_ref(), |h: &PacketHeader, b: &[u8]| {
let mut p = Vec::with_capacity(HEADER_LEN + b.len());
p.extend_from_slice(h.as_bytes());
p.extend_from_slice(b);
pkts.push(p);
Ok(())
})
.unwrap();
(pkts, src)
}
/// Deliver a streamed AU's packets (optionally with kills within the FEC budget, reversed
/// order, and a duplicate) and assert byte-identical completion. Reversed order is the
/// critical case: the FINAL block's real-total headers arrive FIRST, the frame opens
/// legacy-shaped, and the sentinels must still be accepted against the pinned totals.
fn streamed_roundtrip(scheme: FecScheme, kill: &[usize], reverse: bool) {
let chunks: Vec<Vec<u8>> = (0..3)
.map(|c| (0..50).map(|i| (c * 57 + i * 131 + 7) as u8).collect())
.collect();
let chunk_refs: Vec<&[u8]> = chunks.iter().map(|c| c.as_slice()).collect();
let (pkts, src) = streamed_packets(scheme, 50, &chunk_refs);
// 150 B / 16 B shards / 4-shard blocks → sentinel blocks 0,1 (4 data + 2 rec each) +
// final block (2 data + 1 rec) = 15 packets.
assert_eq!(
pkts.len(),
15,
"expected geometry changed — update the kills"
);
let mut delivery: Vec<Vec<u8>> = pkts
.iter()
.enumerate()
.filter(|(i, _)| !kill.contains(i))
.map(|(_, p)| p.clone())
.collect();
if reverse {
delivery.reverse();
}
if let Some(dup) = delivery.first().cloned() {
delivery.push(dup);
}
let cfg = e2e_config(scheme, 50);
let mut r = Reassembler::new(ReassemblerLimits::from_config(&cfg));
let coder = coder_for(scheme);
let stats = StatsCounters::default();
let mut got = None;
for p in &delivery {
if let Some(f) = r.push(p, coder.as_ref(), &stats).unwrap() {
assert!(got.is_none(), "frame must complete exactly once");
got = Some(f);
}
}
let f = got.expect("streamed frame must complete within the FEC budget");
assert_eq!(
f.data, src,
"reassembled streamed AU must be byte-identical"
);
assert_eq!(f.pts_ns, 12345);
assert!(f.complete);
}
#[test]
fn streamed_roundtrip_clean_and_reversed() {
streamed_roundtrip(FecScheme::Gf16, &[], false);
streamed_roundtrip(FecScheme::Gf16, &[], true);
streamed_roundtrip(FecScheme::Gf8, &[], true);
}
/// Loss within each block's FEC budget: one data shard from a sentinel block, one from the
/// final block — in both delivery orders.
#[test]
fn streamed_roundtrip_survives_loss_and_reorder() {
// Wire order: blk0 = 0..4 data + 4..6 rec, blk1 = 6..10 data + 10..12 rec,
// final = 12..14 data + 14 rec.
streamed_roundtrip(FecScheme::Gf16, &[1, 12], false);
streamed_roundtrip(FecScheme::Gf16, &[1, 12], true);
}
/// The wire shape of a streamed AU: sentinel headers (block_count = 0, frame_bytes = 0,
/// full-K) on every non-final block, real totals + EOF on the final block, SOF on the very
/// first packet only.
#[test]
fn streamed_headers_sentinel_then_final() {
let chunks: Vec<Vec<u8>> = (0..3).map(|_| vec![0xA5u8; 50]).collect();
let chunk_refs: Vec<&[u8]> = chunks.iter().map(|c| c.as_slice()).collect();
let (pkts, src) = streamed_packets(FecScheme::Gf16, 50, &chunk_refs);
let mut saw_final = false;
for (i, p) in pkts.iter().enumerate() {
let h = PacketHeader::read_from_bytes(&p[..HEADER_LEN]).unwrap();
assert_eq!(
h.flags & FLAG_SOF != 0,
i == 0,
"SOF exactly on the first packet"
);
assert_eq!(
h.flags & FLAG_EOF != 0,
i + 1 == pkts.len(),
"EOF exactly on the last packet"
);
if h.block_index < 2 {
assert_eq!(
h.block_count, 0,
"non-final block must ride sentinel headers"
);
assert_eq!(h.frame_bytes, 0);
assert_eq!(h.data_shards, 4, "sentinel blocks are exactly full-K");
} else {
saw_final = true;
assert_eq!(h.block_count, 3, "final block carries the real block count");
assert_eq!(h.frame_bytes as usize, src.len(), "and the real AU size");
}
}
assert!(saw_final);
}
/// A streamed AU smaller than one block emits NO sentinels — its single (final) block is
/// byte-identical in shape to a legacy frame, so small frames pay zero streaming overhead
/// and any receiver accepts them.
#[test]
fn streamed_small_frame_degenerates_to_legacy() {
let (pkts, src) = streamed_packets(FecScheme::Gf16, 50, &[&[0x5Au8; 40]]);
for p in &pkts {
let h = PacketHeader::read_from_bytes(&p[..HEADER_LEN]).unwrap();
assert_eq!(
h.block_count, 1,
"single-block streamed AU must be legacy-shaped"
);
assert_eq!(h.frame_bytes as usize, src.len());
}
}
/// Sentinel firewall: a sentinel header that is not exactly full-K, or claims a non-zero
/// total, or sits where the final block could no longer follow, is dropped before any
/// allocation happens.
#[test]
fn streamed_sentinel_firewall_bounds() {
let mut r = Reassembler::new(limits());
let coder = coder_for(FecScheme::Gf8);
let stats = StatsCounters::default();
let sentinel = |f: fn(&mut PacketHeader)| {
let mut h = base_header();
h.block_count = 0;
h.frame_bytes = 0;
h.data_shards = 8; // limits().max_data_shards — the only legal sentinel K
h.recovery_shards = 0;
f(&mut h);
h
};
// Not full-K.
let h = sentinel(|h| h.data_shards = 7);
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
// Claims a total.
let h = sentinel(|h| h.frame_bytes = 64);
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
// Sits on the last block the limits allow (no room for the final block after it).
let h = sentinel(|h| h.block_index = 3); // limits().max_blocks == 4
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
assert_eq!(stats.snapshot().packets_dropped, 3);
// A conformant sentinel IS accepted (proves the rejections above weren't vacuous).
let h = sentinel(|_| {});
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
assert_eq!(
stats.snapshot().packets_dropped,
3,
"conformant sentinel accepted"
);
}
/// Retro-validation: final-block totals under which an already-received sentinel block is
/// out of range (or mis-sized) kill the WHOLE frame — no spliced delivery — and the killed
/// index cannot be resurrected by stragglers.
#[test]
fn streamed_lying_final_totals_kill_the_frame_wholesale() {
let mut r = Reassembler::new(limits());
let coder = coder_for(FecScheme::Gf8);
let stats = StatsCounters::default();
// Two sentinel blocks (indexes 0 and 1) open the frame and land shards.
for bi in 0..2u16 {
let mut h = base_header();
h.block_count = 0;
h.frame_bytes = 0;
h.data_shards = 8;
h.recovery_shards = 0;
h.block_index = bi;
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
}
// A "final" header claiming the whole AU is ONE 16-byte shard: geometry-valid on its own
// (expect_blocks = 1, K = 1), but it disowns both sentinel blocks → the frame dies.
let mut lying = base_header();
lying.block_count = 1;
lying.frame_bytes = 16;
lying.data_shards = 1;
lying.recovery_shards = 0;
assert!(r
.push(&packet(lying), coder.as_ref(), &stats)
.unwrap()
.is_none());
let snap = stats.snapshot();
assert_eq!(
snap.frames_dropped, 1,
"the lying frame must be counted lost"
);
// A straggler sentinel for the killed index must not resurrect it.
let mut h = base_header();
h.block_count = 0;
h.frame_bytes = 0;
h.data_shards = 8;
h.recovery_shards = 0;
let before = stats.snapshot().packets_dropped;
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
assert_eq!(
stats.snapshot().packets_dropped,
before + 1,
"straggler for a killed frame must be dropped, not re-open it"
);
}
/// A sentinel first-packet commits a MAX-sized frame buffer, so the in-flight budget must
/// bite after IN_FLIGHT_BUF_FACTOR frames — the amplification bound for one-datagram opens.
#[test]
fn streamed_open_amplification_is_budget_bounded() {
let mut r = Reassembler::new(limits());
let coder = coder_for(FecScheme::Gf8);
let stats = StatsCounters::default();
// limits(): max_frame_bytes 4096 → each sentinel open commits 4096 B; budget = 4×4096.
for fi in 0..5u32 {
let mut h = base_header();
h.block_count = 0;
h.frame_bytes = 0;
h.data_shards = 8;
h.recovery_shards = 0;
h.frame_index = fi;
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
}
assert_eq!(
stats.snapshot().packets_dropped,
1,
"the fifth max-sized open must be refused by the in-flight budget"
);
}
/// The final-first order's single buffer-safety guard (2026-07 security review finding 3): a
/// frame opened by its FINAL block allocates an EXACT-sized buffer; a sentinel aimed at (or
/// past) the pinned final slot must be dropped — without the guard its full-K write would land
/// outside that buffer. And the reject must not corrupt the frame: it still completes.
#[test]
fn streamed_out_of_range_sentinel_after_final_first_is_dropped() {
let mut r = Reassembler::new(limits());
let coder = coder_for(FecScheme::Gf8);
let stats = StatsCounters::default();
// Final block opens the frame: block_count = 1, frame_bytes = 32 → K = 2. Send shard 0
// only, so the frame stays in flight with the totals pinned.
let mut fin = base_header();
fin.block_count = 1;
fin.frame_bytes = 32;
fin.data_shards = 2;
fin.recovery_shards = 0;
fin.shard_index = 0;
assert!(r
.push(&packet(fin), coder.as_ref(), &stats)
.unwrap()
.is_none());
// Sentinels at the final slot (0) and past it (1): both non-final-impossible under the
// pinned block_count = 1 → dropped, never written into the 32-byte buffer.
for bi in 0..2u16 {
let mut h = base_header();
h.block_count = 0;
h.frame_bytes = 0;
h.data_shards = 8;
h.recovery_shards = 0;
h.block_index = bi;
assert!(r
.push(&packet(h), coder.as_ref(), &stats)
.unwrap()
.is_none());
}
assert_eq!(stats.snapshot().packets_dropped, 2);
// The frame is unharmed: its real second shard completes it at the exact pinned length.
let mut fin2 = fin;
fin2.shard_index = 1;
let got = r
.push(&packet(fin2), coder.as_ref(), &stats)
.unwrap()
.expect("frame must still complete after the rejected sentinels");
assert_eq!(got.data.len(), 32);
assert!(got.complete);
}
/// A second "final" header with DIFFERENT totals must be rejected once a streamed frame is
/// pinned (re-pinning would re-interpret already-landed shards), and the frame must still
/// complete under the first totals.
#[test]
fn streamed_second_final_with_different_totals_is_rejected() {
let mut r = Reassembler::new(limits());
let coder = coder_for(FecScheme::Gf8);
let stats = StatsCounters::default();
let sentinel_shard = |shard_index: u16| {
let mut h = base_header();
h.block_count = 0;
h.frame_bytes = 0;
h.data_shards = 8;
h.recovery_shards = 0;
h.block_index = 0;
h.shard_index = shard_index;
h
};
let final_shard = |frame_bytes: u32, data_shards: u16, shard_index: u16| {
let mut h = base_header();
h.block_count = 2;
h.frame_bytes = frame_bytes;
h.data_shards = data_shards;
h.recovery_shards = 0;
h.block_index = 1;
h.shard_index = shard_index;
h
};
// Sentinel opens block 0, then the real final pins totals: 10 shards = 160 bytes.
assert!(r
.push(&packet(sentinel_shard(0)), coder.as_ref(), &stats)
.unwrap()
.is_none());
assert!(r
.push(&packet(final_shard(160, 2, 0)), coder.as_ref(), &stats)
.unwrap()
.is_none());
// A second final claiming 144 bytes (K = 1): geometry-valid alone, but it contradicts the
// pinned totals → dropped.
let before = stats.snapshot().packets_dropped;
assert!(r
.push(&packet(final_shard(144, 1, 0)), coder.as_ref(), &stats)
.unwrap()
.is_none());
assert_eq!(stats.snapshot().packets_dropped, before + 1);
// The frame still completes under the FIRST totals: the rest of block 0 + the final tail.
let mut got = None;
for s in 1..8u16 {
assert!(got.is_none());
got = r
.push(&packet(sentinel_shard(s)), coder.as_ref(), &stats)
.unwrap();
}
assert!(got.is_none(), "block 1 still owes a shard");
let got = r
.push(&packet(final_shard(160, 2, 1)), coder.as_ref(), &stats)
.unwrap()
.expect("frame completes under the first pinned totals");
assert_eq!(got.data.len(), 160);
}
+90 -1
View File
@@ -32,6 +32,18 @@ pub const VIDEO_CAP_HOST_TIMING: u8 = 0x08;
/// older client gets a declined (zeroed) [`ProbeResult`] instead of a measurement its single-window /// older client gets a declined (zeroed) [`ProbeResult`] instead of a measurement its single-window
/// reassembler would silently drop as stale. /// reassembler would silently drop as stale.
pub const VIDEO_CAP_PROBE_SEQ: u8 = 0x10; pub const VIDEO_CAP_PROBE_SEQ: u8 = 0x10;
/// [`Hello::video_caps`] bit: the client's reassembler accepts **streamed access units**
/// (design/nvenc-subframe-slice-output.md Phase 2): the host may ship an AU's early FEC blocks
/// before the AU's total size exists — while the tail of the frame is still encoding — so the
/// AU's last packet leaves the host sooner (latency plan §7 LN1). Non-final blocks ride
/// SENTINEL headers (`block_count == 0` — a value no legacy sender emits — with
/// `frame_bytes == 0` and exactly `max_data_per_block` data shards, so the shard-offset
/// formula needs no total); the FINAL block's headers carry the real
/// `frame_bytes`/`block_count` (+ `FLAG_EOF`), which retro-validate the whole frame's geometry
/// — a mismatch drops the frame wholesale. The host streams ONLY to clients advertising this
/// bit; every other client gets today's whole-AU path (chunks concatenated before sealing), so
/// the fallback is zero-risk.
pub const VIDEO_CAP_STREAMED_AU: u8 = 0x20;
/// [`Welcome::host_caps`] bit: the host applies [`InputKind::GamepadState`] /// [`Welcome::host_caps`] bit: the host applies [`InputKind::GamepadState`]
/// (crate::input::InputKind::GamepadState) snapshot events — full per-pad state with a reorder /// (crate::input::InputKind::GamepadState) snapshot events — full per-pad state with a reorder
@@ -91,8 +103,13 @@ pub fn resolve_codec(client_codecs: u8, host_capable: u8, preferred: u8) -> Opti
return None; return None;
} }
// Honor the client's preference when the host can also emit it; else fall back to precedence. // Honor the client's preference when the host can also emit it; else fall back to precedence.
// `preferred` is a single-bit field by contract but arrives as a raw wire byte — isolate ONE
// bit of the intersection instead of echoing the request, so a non-conformant multi-bit
// value can never escape as a codec id (downstream `from_wire` folds unknown values to HEVC,
// which may not even be in the shared set).
if preferred != 0 && shared & preferred != 0 { if preferred != 0 && shared & preferred != 0 {
return Some(preferred); let want = shared & preferred;
return Some(want & want.wrapping_neg());
} }
// Precedence: HEVC > AV1 > H.264. // Precedence: HEVC > AV1 > H.264.
[CODEC_HEVC, CODEC_AV1, CODEC_H264] [CODEC_HEVC, CODEC_AV1, CODEC_H264]
@@ -171,3 +188,75 @@ impl Default for ColorInfo {
Self::SDR_BT709 Self::SDR_BT709
} }
} }
#[cfg(test)]
mod tests {
use crate::config::{CompositorPref, FecConfig, FecScheme, GamepadPref, Mode};
use crate::quic::*;
#[test]
fn host_cap_clipboard_bit_is_distinct_and_survives_welcome() {
// The new cap packs into the existing trailing host_caps byte with no layout change.
assert_ne!(HOST_CAP_CLIPBOARD, HOST_CAP_GAMEPAD_STATE);
let mut w = Welcome {
abi_version: 1,
udp_port: 1,
mode: Mode {
width: 1920,
height: 1080,
refresh_hz: 60,
},
fec: FecConfig {
scheme: FecScheme::Gf16,
fec_percent: 0,
max_data_per_block: 1024,
},
shard_payload: 1024,
encrypt: false,
key: [0; 16],
salt: [0; 4],
frames: 0,
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
bit_depth: 8,
color: ColorInfo::SDR_BT709,
chroma_format: CHROMA_IDC_420,
audio_channels: 2,
codec: CODEC_HEVC,
host_caps: HOST_CAP_GAMEPAD_STATE | HOST_CAP_CLIPBOARD,
};
let got = Welcome::decode(&w.encode()).unwrap();
assert_eq!(got.host_caps & HOST_CAP_CLIPBOARD, HOST_CAP_CLIPBOARD);
assert_eq!(
got.host_caps & HOST_CAP_GAMEPAD_STATE,
HOST_CAP_GAMEPAD_STATE
);
// Clipboard-off host: the bit is clear, gamepad bit still set.
w.host_caps = HOST_CAP_GAMEPAD_STATE;
assert_eq!(
Welcome::decode(&w.encode()).unwrap().host_caps & HOST_CAP_CLIPBOARD,
0
);
}
#[test]
fn resolve_codec_canonicalizes_a_multi_bit_preference() {
// A non-conformant peer may stuff its capability MASK into `preferred` — the result
// must still be a single bit of the shared set, never the raw multi-bit echo (which
// folds to HEVC downstream and can select a codec the client can't decode).
assert_eq!(
resolve_codec(CODEC_H264, CODEC_H264 | CODEC_AV1, CODEC_H264 | CODEC_AV1),
Some(CODEC_H264)
);
// Several shared preferred bits: still exactly one bit, and one of the preferred ones.
let got = resolve_codec(
CODEC_H264 | CODEC_HEVC | CODEC_AV1,
CODEC_H264 | CODEC_HEVC | CODEC_AV1,
CODEC_AV1 | CODEC_HEVC,
)
.unwrap();
assert_eq!(got.count_ones(), 1);
assert_ne!(got & (CODEC_AV1 | CODEC_HEVC), 0);
}
}
@@ -115,3 +115,134 @@ pub async fn read_data(recv: &mut quinn::RecvStream, max_bytes: usize) -> std::i
.await .await
.map_err(std::io::Error::other) .map_err(std::io::Error::other)
} }
// In-process QUIC loopback: the real clipstream fetch transport, both success and cancel.
#[cfg(test)]
mod tests {
use crate::quic::clipstream;
use crate::quic::test_util::connect_pair;
use crate::quic::*;
#[tokio::test]
async fn fetch_text_transfers_then_cancel_resets() {
let (_server_ep, _client_ep, host_conn, client_conn) = connect_pair().await;
let payload = b"hello clipboard \xf0\x9f\x93\x8b".to_vec(); // text + a 4-byte emoji
let holder_payload = payload.clone();
// Holder = the host side: accept two fetch streams. Serve the first; cancel the second.
let holder = tokio::spawn(async move {
// Fetch #1 — serve the payload.
let (mut send, mut recv) = host_conn.accept_bi().await.expect("accept fetch #1");
let kind = clipstream::read_stream_header(&mut recv)
.await
.expect("stream header #1");
assert_eq!(kind, clipstream::CLIP_STREAM_KIND_FETCH);
let req = clipstream::read_fetch(&mut recv)
.await
.expect("fetch req #1");
assert_eq!(req.seq, 1);
assert_eq!(req.file_index, CLIP_FILE_INDEX_NONE);
assert_eq!(req.mime, "text/plain;charset=utf-8");
clipstream::write_fetch_hdr(
&mut send,
&ClipFetchHdr {
status: CLIP_FETCH_OK,
total_size: holder_payload.len() as u64,
},
)
.await
.expect("write hdr #1");
clipstream::write_data(&mut send, &holder_payload)
.await
.expect("write data #1");
// Fetch #2 — read the request, then cancel mid-transfer with RESET_STREAM.
let (mut send2, mut recv2) = host_conn.accept_bi().await.expect("accept fetch #2");
clipstream::read_stream_header(&mut recv2)
.await
.expect("stream header #2");
let _ = clipstream::read_fetch(&mut recv2)
.await
.expect("fetch req #2");
send2.reset(clipstream::cancelled_code()).unwrap();
host_conn // keep alive until the requester side is done
});
// Requester = the client side.
// #1: full lazy fetch of the text payload.
let req = ClipFetch {
seq: 1,
file_index: CLIP_FILE_INDEX_NONE,
mime: "text/plain;charset=utf-8".into(),
};
let (_send, mut recv) = clipstream::open_fetch(&client_conn, &req)
.await
.expect("open fetch #1");
let hdr = clipstream::read_fetch_hdr(&mut recv)
.await
.expect("read hdr #1");
assert_eq!(hdr.status, CLIP_FETCH_OK);
assert_eq!(hdr.total_size as usize, payload.len());
let got = clipstream::read_data(&mut recv, 8 << 20)
.await
.expect("read data #1");
assert_eq!(got, payload);
// #2: the holder resets the stream — the requester surfaces an error rather than hanging.
let req2 = ClipFetch {
seq: 2,
file_index: CLIP_FILE_INDEX_NONE,
mime: "text/plain;charset=utf-8".into(),
};
let (_send2, mut recv2) = clipstream::open_fetch(&client_conn, &req2)
.await
.expect("open fetch #2");
assert!(
clipstream::read_fetch_hdr(&mut recv2).await.is_err(),
"a cancelled fetch must surface as an error, not a hang"
);
let _host_conn = holder.await.unwrap();
}
#[tokio::test]
async fn read_data_enforces_size_cap() {
let (_server_ep, _client_ep, host_conn, client_conn) = connect_pair().await;
let big = vec![0xABu8; 200_000]; // > the 64 KiB chunk, and > the cap we set below
let holder_payload = big.clone();
let holder = tokio::spawn(async move {
let (mut send, mut recv) = host_conn.accept_bi().await.expect("accept");
clipstream::read_stream_header(&mut recv).await.unwrap();
let _ = clipstream::read_fetch(&mut recv).await.unwrap();
clipstream::write_fetch_hdr(
&mut send,
&ClipFetchHdr {
status: CLIP_FETCH_OK,
total_size: holder_payload.len() as u64,
},
)
.await
.unwrap();
let _ = clipstream::write_data(&mut send, &holder_payload).await;
host_conn
});
let req = ClipFetch {
seq: 1,
file_index: CLIP_FILE_INDEX_NONE,
mime: "application/octet-stream".into(),
};
let (_send, mut recv) = clipstream::open_fetch(&client_conn, &req).await.unwrap();
assert_eq!(
clipstream::read_fetch_hdr(&mut recv).await.unwrap().status,
CLIP_FETCH_OK
);
// Cap below the payload size ⇒ read_data errors instead of buffering unboundedly.
assert!(clipstream::read_data(&mut recv, 64 * 1024).await.is_err());
let _host_conn = holder.await.unwrap();
}
}
+114 -2
View File
@@ -38,7 +38,7 @@ pub struct ClockSkew {
/// with, so the offset aligns a client receive instant to the host's capture clock. /// with, so the offset aligns a client receive instant to the host's capture clock.
pub async fn clock_sync( pub async fn clock_sync(
send: &mut quinn::SendStream, send: &mut quinn::SendStream,
recv: &mut quinn::RecvStream, recv: &mut io::MsgReader,
) -> Option<ClockSkew> { ) -> Option<ClockSkew> {
use std::time::Duration; use std::time::Duration;
const ROUNDS: usize = 8; const ROUNDS: usize = 8;
@@ -50,7 +50,7 @@ pub async fn clock_sync(
if io::write_msg(send, &probe).await.is_err() { if io::write_msg(send, &probe).await.is_err() {
break; break;
} }
let read = tokio::time::timeout(read_timeout, io::read_msg(recv)).await; let read = tokio::time::timeout(read_timeout, recv.read_msg()).await;
let echo = match read { let echo = match read {
Ok(Ok(b)) => match ClockEcho::decode(&b) { Ok(Ok(b)) => match ClockEcho::decode(&b) {
Ok(e) => e, Ok(e) => e,
@@ -154,3 +154,115 @@ impl Default for ClockResync {
pub fn accept_resync(batch_rtt_ns: u64, connect_rtt_ns: u64) -> bool { pub fn accept_resync(batch_rtt_ns: u64, connect_rtt_ns: u64) -> bool {
batch_rtt_ns <= (connect_rtt_ns + connect_rtt_ns / 2).max(2_000_000) batch_rtt_ns <= (connect_rtt_ns + connect_rtt_ns / 2).max(2_000_000)
} }
#[cfg(test)]
mod tests {
use crate::quic::*;
#[test]
fn clock_offset_picks_min_rtt_and_recovers_offset() {
// Host clock is +1_000_000 ns ahead of the client. Construct samples where a symmetric
// round-trip recovers exactly that offset, and a noisy (asymmetric, high-RTT) sample is
// present but must be ignored by the min-RTT selection.
const OFF: i64 = 1_000_000;
// Clean sample: client t1=0, one-way=200µs each way → t2 = t1 + 200_000 + OFF (host clock),
// t3 = t2 + 50_000 (host processing), t4 = t3 - OFF + 200_000 (back in client clock).
let t1 = 0u64;
let t2 = (t1 as i64 + 200_000 + OFF) as u64;
let t3 = t2 + 50_000;
let t4 = (t3 as i64 - OFF + 200_000) as u64;
// Noisy sample: same offset but a fat, asymmetric RTT (slow return path) — higher RTT.
let n1 = 1_000_000u64;
let n2 = (n1 as i64 + 200_000 + OFF) as u64;
let n3 = n2 + 50_000;
let n4 = (n3 as i64 - OFF + 5_000_000) as u64; // 5 ms return → big RTT
let (offset, rtt) =
clock_offset_ns(&[(n1, n2, n3, n4), (t1, t2, t3, t4)]).expect("non-empty");
// The min-RTT sample recovers the offset exactly; its RTT is 2x200us, and the noisy
// (asymmetric, 5 ms return) sample is ignored by the min-RTT selection.
assert_eq!(offset, OFF);
assert_eq!(rtt, 400_000);
assert!(clock_offset_ns(&[]).is_none());
}
/// The mid-stream re-sync state machine: 8 rounds collected via matched echoes, stale
/// echoes ignored, a restarted batch abandons the old one, and the batch result is the
/// min-RTT estimate — the exact behavior the connect-time `clock_sync` loop has.
#[test]
fn clock_resync_collects_rounds_and_ignores_stale_echoes() {
// Host clock +1 ms ahead; symmetric 100 µs one-way paths except one congested round.
const OFF: i64 = 1_000_000;
let echo_for = |t1: u64, one_way: u64| ClockEcho {
t1_ns: t1,
t2_ns: (t1 as i64 + one_way as i64 + OFF) as u64,
t3_ns: (t1 as i64 + one_way as i64 + OFF) as u64 + 10_000,
};
let t4_for = |e: &ClockEcho, one_way: u64| (e.t3_ns as i64 - OFF + one_way as i64) as u64;
let mut rs = ClockResync::new();
// An unsolicited echo before any batch is ignored.
assert_eq!(
rs.on_echo(&echo_for(42, 100_000), 500_000),
ResyncStep::Idle
);
let mut probe = rs.begin(1_000_000);
// A stale echo (wrong t1: the abandoned pre-begin probe) is ignored mid-batch.
assert_eq!(
rs.on_echo(&echo_for(42, 100_000), 500_000),
ResyncStep::Idle
);
for round in 0..ClockResync::ROUNDS {
// Round 3 is congested (5 ms one-way) — it must lose the min-RTT selection.
let one_way = if round == 3 { 5_000_000 } else { 100_000 };
let echo = echo_for(probe.t1_ns, one_way);
let t4 = t4_for(&echo, one_way);
match rs.on_echo(&echo, t4) {
ResyncStep::Probe(p) => {
assert!(round < ClockResync::ROUNDS - 1, "batch overran its rounds");
probe = p;
}
ResyncStep::Done { offset_ns, rtt_ns } => {
assert_eq!(round, ClockResync::ROUNDS - 1, "batch ended early");
assert_eq!(offset_ns, OFF, "min-RTT round recovers the offset exactly");
assert_eq!(rtt_ns, 200_000); // 2×100 µs; host processing (t3t2) excluded
}
ResyncStep::Idle => panic!("matched echo must advance the batch"),
}
}
// The batch is done: even a matching-t1 replay no longer advances anything.
assert_eq!(
rs.on_echo(&echo_for(probe.t1_ns, 100_000), probe.t1_ns + 300_000),
ResyncStep::Idle
);
// begin() mid-batch abandons the in-flight batch: its echo is stale afterwards.
let old = rs.begin(2_000_000);
let fresh = rs.begin(3_000_000);
assert_eq!(
rs.on_echo(&echo_for(old.t1_ns, 100_000), 2_300_000),
ResyncStep::Idle
);
assert!(matches!(
rs.on_echo(&echo_for(fresh.t1_ns, 100_000), 3_300_000),
ResyncStep::Probe(_)
));
}
/// The acceptance guard: a batch measured through a congested window (fat RTT) must not
/// replace the offset — its queueing delay biases the estimate exactly when frames
/// already read late. Floor of 2 ms so a near-zero connect RTT (same-host/LAN) doesn't
/// reject every later batch over normal jitter.
#[test]
fn clock_resync_acceptance_guard() {
// Generous connect RTT (10 ms): accept up to 1.5×.
assert!(accept_resync(14_000_000, 10_000_000));
assert!(!accept_resync(16_000_000, 10_000_000));
// Tiny connect RTT (200 µs, wired LAN): the 2 ms floor governs.
assert!(accept_resync(1_900_000, 200_000));
assert!(!accept_resync(2_100_000, 200_000));
// Boundary: exactly at the bound is accepted.
assert!(accept_resync(2_000_000, 0));
assert!(accept_resync(15_000_000, 10_000_000));
}
}
+366
View File
@@ -782,3 +782,369 @@ impl ClipFetchHdr {
}) })
} }
} }
#[cfg(test)]
mod tests {
use crate::config::Mode;
use crate::quic::*;
#[test]
fn reconfigure_roundtrip() {
let rq = Reconfigure {
mode: Mode {
width: 1920,
height: 1080,
refresh_hz: 144,
},
};
assert_eq!(Reconfigure::decode(&rq.encode()).unwrap(), rq);
for accepted in [true, false] {
let rs = Reconfigured {
accepted,
mode: rq.mode,
};
assert_eq!(Reconfigured::decode(&rs.encode()).unwrap(), rs);
}
// The type byte separates the post-handshake messages from each other.
assert!(Reconfigure::decode(
&Reconfigured {
accepted: true,
mode: rq.mode
}
.encode()
)
.is_err());
}
#[test]
fn request_keyframe_roundtrip() {
let bytes = RequestKeyframe.encode();
assert!(RequestKeyframe::decode(&bytes).is_ok());
// Distinct from the other control messages — its type byte must not collide.
let mode = Mode {
width: 1280,
height: 720,
refresh_hz: 60,
};
assert!(RequestKeyframe::decode(&Reconfigure { mode }.encode()).is_err());
assert!(Reconfigure::decode(&bytes).is_err());
// Length is exact (no trailing bytes accepted).
assert!(RequestKeyframe::decode(&[bytes.as_slice(), &[0]].concat()).is_err());
}
#[test]
fn rfi_request_roundtrip() {
for (first_frame, last_frame) in [(0u32, 0u32), (40, 47), (5, 5), (1_000_000, u32::MAX)] {
let r = RfiRequest {
first_frame,
last_frame,
};
assert_eq!(RfiRequest::decode(&r.encode()).unwrap(), r);
}
// Disjoint from the bare keyframe request (its loss-unaware sibling) and others: type byte + length.
assert!(RfiRequest::decode(&RequestKeyframe.encode()).is_err());
assert!(RequestKeyframe::decode(
&RfiRequest {
first_frame: 1,
last_frame: 2
}
.encode()
)
.is_err());
// Exact length — no trailing bytes.
let bytes = RfiRequest {
first_frame: 3,
last_frame: 9,
}
.encode();
assert!(RfiRequest::decode(&[bytes.as_slice(), &[0]].concat()).is_err());
assert!(RfiRequest::decode(&bytes[..bytes.len() - 1]).is_err());
}
#[test]
fn loss_report_roundtrip() {
for loss_ppm in [0u32, 1, 12_345, 50_000, 1_000_000] {
let r = LossReport { loss_ppm };
assert_eq!(LossReport::decode(&r.encode()).unwrap(), r);
}
// Disjoint from the other control messages (type byte + length).
assert!(LossReport::decode(&RequestKeyframe.encode()).is_err());
assert!(RequestKeyframe::decode(&LossReport { loss_ppm: 0 }.encode()).is_err());
assert!(LossReport::decode(
&[LossReport { loss_ppm: 0 }.encode().as_slice(), &[0]].concat()
)
.is_err());
}
#[test]
fn window_loss_ppm_estimates_and_caps() {
// No traffic → 0. A clean window (nothing recovered) → 0.
assert_eq!(window_loss_ppm(0, 0, 0, 0), 0);
assert_eq!(window_loss_ppm(0, 0, 1000, 0), 0);
// 50 recovered of 1000 total (950 received + 50 recovered) = 5%.
assert_eq!(window_loss_ppm(50, 0, 950, 0), 50_000);
// An unrecoverable frame adds the +5% bump (push FEC past the current cap).
assert_eq!(window_loss_ppm(50, 0, 950, 1), 100_000);
// A total-loss window with a drop but nothing received still reports the bump, capped at 1e6.
assert_eq!(window_loss_ppm(0, 0, 0, 3), 50_000);
assert!(window_loss_ppm(u64::MAX, 0, 1, 9) <= 1_000_000);
// Reordering: shards "recovered" early that then arrived are late, not lost — netted out, so
// a pure-reorder window reads 0. Partially late nets to the true loss (20 of 1000 = 2%).
assert_eq!(window_loss_ppm(50, 50, 1000, 0), 0);
assert_eq!(window_loss_ppm(50, 30, 980, 0), 20_000);
// `late` can outrun `recovered` across a window boundary (reorder straddling the report
// tick) or via a rare wire duplicate — saturate at a clean window, never underflow.
assert_eq!(window_loss_ppm(10, 25, 1000, 0), 0);
}
#[test]
fn bitrate_messages_roundtrip() {
let req = SetBitrate {
bitrate_kbps: 14_000,
};
assert_eq!(SetBitrate::decode(&req.encode()).unwrap(), req);
let ack = BitrateChanged {
bitrate_kbps: 14_000,
};
assert_eq!(BitrateChanged::decode(&ack.encode()).unwrap(), ack);
// Same payload shape as LossReport — the type byte alone must keep them disjoint.
assert!(LossReport::decode(&req.encode()).is_err());
assert!(SetBitrate::decode(&ack.encode()).is_err());
assert!(BitrateChanged::decode(&req.encode()).is_err());
assert!(SetBitrate::decode(&LossReport { loss_ppm: 7 }.encode()).is_err());
}
#[test]
fn probe_messages_roundtrip() {
let req = ProbeRequest {
target_kbps: 250_000,
duration_ms: 2000,
};
assert_eq!(ProbeRequest::decode(&req.encode()).unwrap(), req);
let res = ProbeResult {
bytes_sent: 62_500_000,
packets_sent: 480,
duration_ms: 2003,
wire_packets_sent: 41_000,
send_dropped: 1_200,
};
assert_eq!(ProbeResult::decode(&res.encode()).unwrap(), res);
assert_eq!(res.encode().len(), 29);
// A pre-wire-stats host's 21-byte ProbeResult still decodes, with the new fields zeroed.
let legacy = {
let full = res.encode();
full[..21].to_vec()
};
let decoded = ProbeResult::decode(&legacy).unwrap();
assert_eq!(decoded.wire_packets_sent, 0);
assert_eq!(decoded.send_dropped, 0);
assert_eq!(decoded.bytes_sent, res.bytes_sent);
// Type bytes keep the control messages disjoint from each other.
assert!(ProbeRequest::decode(&res.encode()).is_err());
assert!(Reconfigure::decode(&req.encode()).is_err());
assert!(ProbeResult::decode(&req.encode()).is_err());
}
#[test]
fn clock_messages_roundtrip() {
let probe = ClockProbe {
t1_ns: 1_700_000_000_123,
};
assert_eq!(ClockProbe::decode(&probe.encode()).unwrap(), probe);
let echo = ClockEcho {
t1_ns: 1_700_000_000_123,
t2_ns: 1_700_000_050_456,
t3_ns: 1_700_000_050_789,
};
assert_eq!(ClockEcho::decode(&echo.encode()).unwrap(), echo);
// Disjoint from the other control messages (distinct type bytes).
assert!(ClockProbe::decode(&echo.encode()).is_err());
assert!(ProbeRequest::decode(&probe.encode()).is_err());
assert!(ClockEcho::decode(&probe.encode()).is_err());
}
// ---- Shared clipboard control + fetch-stream message codecs (0x40-0x44) -----------------------
#[test]
fn clip_control_roundtrip() {
for (enabled, flags) in [
(true, 0u8),
(false, 0),
(true, CLIP_FLAG_FILES),
(false, 0xFF),
] {
let m = ClipControl { enabled, flags };
assert_eq!(ClipControl::decode(&m.encode()).unwrap(), m);
}
// Disjoint from its host→client sibling (type byte + length) and exact length.
assert!(ClipControl::decode(
&ClipState {
enabled: true,
policy: 0,
reason: 0
}
.encode()
)
.is_err());
let bytes = ClipControl {
enabled: true,
flags: 0,
}
.encode();
assert!(ClipControl::decode(&[bytes.as_slice(), &[0]].concat()).is_err());
assert!(ClipControl::decode(&bytes[..bytes.len() - 1]).is_err());
}
#[test]
fn clip_state_roundtrip() {
let cases = [
ClipState {
enabled: true,
policy: CLIP_POLICY_TEXT | CLIP_POLICY_FILES,
reason: CLIP_REASON_OK,
},
ClipState {
enabled: false,
policy: 0,
reason: CLIP_REASON_BACKEND_UNAVAILABLE,
},
ClipState {
enabled: true,
policy: CLIP_POLICY_TEXT,
reason: CLIP_REASON_NO_FILES,
},
];
for m in cases {
assert_eq!(ClipState::decode(&m.encode()).unwrap(), m);
}
// A ClipControl must not decode as a ClipState (type byte).
assert!(ClipState::decode(
&ClipControl {
enabled: true,
flags: 0
}
.encode()
)
.is_err());
let bytes = cases[0].encode();
assert!(ClipState::decode(&bytes[..bytes.len() - 1]).is_err());
}
#[test]
fn clip_offer_roundtrip() {
// Empty offer, one kind, and a full multi-format offer (text/rich/image/files).
let cases = [
ClipOffer {
seq: 0,
kinds: vec![],
},
ClipOffer {
seq: 1,
kinds: vec![ClipKind {
mime: "text/plain;charset=utf-8".into(),
size_hint: 12,
}],
},
ClipOffer {
seq: u32::MAX,
kinds: vec![
ClipKind {
mime: "text/plain;charset=utf-8".into(),
size_hint: 0,
},
ClipKind {
mime: "text/html".into(),
size_hint: 4096,
},
ClipKind {
mime: "image/png".into(),
size_hint: 1 << 30,
},
ClipKind {
mime: "application/x-punktfunk-files".into(),
size_hint: 5_000_000_000,
},
],
},
];
for m in &cases {
assert_eq!(&ClipOffer::decode(&m.encode()).unwrap(), m);
}
// Trailing bytes are rejected (get_clip_kind consumes exactly to the end).
let mut padded = cases[1].encode();
padded.push(0);
assert!(ClipOffer::decode(&padded).is_err());
// A count byte over the cap is rejected before allocating.
let mut over = cases[0].encode();
over[9] = (CLIP_MAX_KINDS + 1) as u8;
assert!(ClipOffer::decode(&over).is_err());
// Disjoint from a same-family control message.
assert!(ClipOffer::decode(
&ClipControl {
enabled: true,
flags: 0
}
.encode()
)
.is_err());
}
#[test]
fn clip_fetch_roundtrip() {
let cases = [
ClipFetch {
seq: 1,
file_index: CLIP_FILE_INDEX_NONE,
mime: "text/plain;charset=utf-8".into(),
},
ClipFetch {
seq: 7,
file_index: 0,
mime: "application/x-punktfunk-files".into(),
},
ClipFetch {
seq: u32::MAX,
file_index: 41,
mime: String::new(),
},
];
for m in &cases {
assert_eq!(&ClipFetch::decode(&m.encode()).unwrap(), m);
}
// Trailing + truncation both rejected (exact-length mime check).
let bytes = cases[0].encode();
assert!(ClipFetch::decode(&[bytes.as_slice(), &[0]].concat()).is_err());
assert!(ClipFetch::decode(&bytes[..bytes.len() - 1]).is_err());
// A fetch-stream message must not decode as a control-stream offer, and vice-versa.
assert!(ClipOffer::decode(&cases[0].encode()).is_err());
assert!(ClipFetch::decode(
&ClipOffer {
seq: 1,
kinds: vec![]
}
.encode()
)
.is_err());
}
#[test]
fn clip_fetch_hdr_roundtrip() {
for (status, total_size) in [
(CLIP_FETCH_OK, 15u64),
(CLIP_FETCH_STALE, 0),
(CLIP_FETCH_UNAVAILABLE, 0),
(CLIP_FETCH_DENIED, 0),
(CLIP_FETCH_OK, u64::MAX),
] {
let m = ClipFetchHdr { status, total_size };
assert_eq!(ClipFetchHdr::decode(&m.encode()).unwrap(), m);
}
let bytes = ClipFetchHdr {
status: CLIP_FETCH_OK,
total_size: 1,
}
.encode();
assert!(ClipFetchHdr::decode(&[bytes.as_slice(), &[0]].concat()).is_err());
assert!(ClipFetchHdr::decode(&bytes[..bytes.len() - 1]).is_err());
}
}
+318
View File
@@ -605,3 +605,321 @@ pub fn decode_host_timing_datagram(b: &[u8]) -> Option<HostTiming> {
stages, stages,
}) })
} }
#[cfg(test)]
mod tests {
use crate::quic::*;
#[test]
fn hdr_meta_datagram_roundtrip_and_truncation() {
let m = HdrMeta {
// BT.2020 display primaries in 1/50000 units (the DXGI/ST.2086 reference values).
display_primaries: [[8500, 39850], [6550, 2300], [35400, 14600]],
white_point: [15635, 16450], // D65
max_display_mastering_luminance: 10_000_000, // 1000 nits in 0.0001 cd/m²
min_display_mastering_luminance: 1, // 0.0001 nits
max_cll: 1000,
max_fall: 400,
};
let d = encode_hdr_meta_datagram(&m);
assert_eq!(d[0], HDR_META_MAGIC);
assert_eq!(decode_hdr_meta_datagram(&d), Some(m));
// Truncated buffers and a wrong tag are rejected (never partially read).
for n in 0..d.len() {
assert_eq!(decode_hdr_meta_datagram(&d[..n]), None);
}
let mut bad = d.clone();
bad[0] = HIDOUT_MAGIC;
assert_eq!(decode_hdr_meta_datagram(&bad), None);
}
#[test]
fn host_timing_datagram_roundtrip_and_truncation() {
let t = HostTiming {
pts_ns: 1_751_500_000_123_456_789, // a realistic 2026 CLOCK_REALTIME capture stamp
host_us: 4_321,
stages: None,
};
let d = encode_host_timing_datagram(&t);
assert_eq!(d[0], HOST_TIMING_MAGIC);
assert_eq!(d.len(), 13);
assert_eq!(decode_host_timing_datagram(&d), Some(t));
// Truncated buffers and a wrong tag are rejected (never partially read).
for n in 0..d.len() {
assert_eq!(decode_host_timing_datagram(&d[..n]), None);
}
let mut bad = d.clone();
bad[0] = HDR_META_MAGIC;
assert_eq!(decode_host_timing_datagram(&bad), None);
// Extended form (T0.1): the stage tail roundtrips; a truncated tail (an old host's 13-byte
// datagram, or anything short of the full 25) degrades to `stages: None`, never a partial
// read; the prefix fields stay identical in both forms (the append-extensibility contract).
let ts = HostTiming {
stages: Some(HostStages {
queue_us: 900,
encode_us: 3_100,
pace_us: 2_500,
}),
..t
};
let ds = encode_host_timing_datagram(&ts);
assert_eq!(ds.len(), 25);
assert_eq!(
&ds[..13],
&d[..13],
"prefix is byte-identical to the legacy form"
);
assert_eq!(decode_host_timing_datagram(&ds), Some(ts));
for n in 13..ds.len() {
assert_eq!(
decode_host_timing_datagram(&ds[..n]),
Some(t),
"partial stage tail ({n} B) must degrade to the legacy decode"
);
}
}
#[test]
fn audio_datagram_roundtrip() {
let opus = [0x42u8; 97];
let d = encode_audio_datagram(7, 1_000_000_123, &opus);
assert_eq!(d[0], AUDIO_MAGIC);
let (seq, pts, payload) = decode_audio_datagram(&d).unwrap();
assert_eq!((seq, pts), (7, 1_000_000_123));
assert_eq!(payload, opus);
assert!(decode_audio_datagram(&d[..12]).is_none()); // truncated header
assert!(decode_audio_datagram(&[0u8; 13]).is_none()); // bad magic
// Empty payload is legal (DTX) — header-only datagram.
let header_only = encode_audio_datagram(0, 0, &[]);
let (_, _, empty) = decode_audio_datagram(&header_only).unwrap();
assert!(empty.is_empty());
}
#[test]
fn rumble_datagram_roundtrip() {
let d = encode_rumble_datagram(1, 0x1234, 0xFFFF);
assert_eq!(d[0], RUMBLE_MAGIC);
assert_eq!(decode_rumble_datagram(&d), Some((1, 0x1234, 0xFFFF)));
assert!(decode_rumble_datagram(&d[..6]).is_none());
}
#[test]
fn rumble_envelope_roundtrip_and_legacy_tolerance() {
// v2 envelope round-trips seq + ttl.
let d = encode_rumble_datagram_v2(2, 0x4000, 0x8000, 7, 400);
assert_eq!(d[0], RUMBLE_MAGIC);
assert_eq!(d.len(), RUMBLE_V2_LEN);
assert_eq!(
decode_rumble_envelope(&d),
Some(RumbleUpdate {
pad: 2,
low: 0x4000,
high: 0x8000,
envelope: Some(RumbleEnvelope {
seq: 7,
ttl_ms: 400
}),
})
);
// The legacy level decoder reads a v2 datagram as a plain level — the tail is ignored, so an
// old client running against a new host still renders the right amplitudes.
assert_eq!(decode_rumble_datagram(&d), Some((2, 0x4000, 0x8000)));
// A legacy 7-byte datagram (old host) decodes as a level with no envelope — a new client then
// applies its own staleness policy.
let v1 = encode_rumble_datagram(3, 0x1111, 0x2222);
assert_eq!(
decode_rumble_envelope(&v1),
Some(RumbleUpdate {
pad: 3,
low: 0x1111,
high: 0x2222,
envelope: None,
})
);
// A torn/short tail (8 or 9 bytes) is not a valid envelope — degrade to a level, never panic
// or drop. (The host never emits these; a truncating middlebox might.)
assert_eq!(
decode_rumble_envelope(&d[..8]).map(|u| u.envelope),
Some(None)
);
assert_eq!(
decode_rumble_envelope(&d[..9]).map(|u| u.envelope),
Some(None)
);
// Bad tag / too short → None on both decoders.
assert!(decode_rumble_envelope(&d[..6]).is_none());
let mut wrong_tag = d;
wrong_tag[0] = AUDIO_MAGIC;
assert!(decode_rumble_envelope(&wrong_tag).is_none());
}
#[test]
fn rumble_envelope_seq_gate_drops_reordered_stale_start() {
use crate::input::GamepadSnapshot;
// The client-side reorder gate (reused verbatim from gamepad snapshots): a stale start
// arriving after a stop must not re-light the motors.
let stop = decode_rumble_envelope(&encode_rumble_datagram_v2(0, 0, 0, 10, 0)).unwrap();
let stale_start =
decode_rumble_envelope(&encode_rumble_datagram_v2(0, 0x8000, 0x8000, 9, 400)).unwrap();
let stop_seq = stop.envelope.unwrap().seq;
let stale_seq = stale_start.envelope.unwrap().seq;
// Nothing applied yet → the first update always passes.
assert!(GamepadSnapshot::seq_newer(stop_seq, None));
// The reordered older start does NOT supersede the stop.
assert!(!GamepadSnapshot::seq_newer(stale_seq, Some(stop_seq)));
// A genuine later renewal does.
assert!(GamepadSnapshot::seq_newer(11, Some(stop_seq)));
// Wraps: seq 1 supersedes 254.
assert!(GamepadSnapshot::seq_newer(1, Some(254)));
}
#[test]
fn mic_datagram_roundtrip_and_disjoint_from_audio() {
let opus = [0x5Au8; 80];
let d = encode_mic_datagram(42, 9_999, &opus);
assert_eq!(d[0], MIC_MAGIC);
let (seq, pts, payload) = decode_mic_datagram(&d).unwrap();
assert_eq!((seq, pts), (42, 9_999));
assert_eq!(payload, opus);
assert!(decode_mic_datagram(&d[..12]).is_none()); // truncated
// Tag separation: a mic datagram is not an audio datagram and vice-versa.
assert!(decode_audio_datagram(&d).is_none());
assert!(decode_mic_datagram(&encode_audio_datagram(1, 2, &opus)).is_none());
// Empty payload (DTX) is legal.
assert!(decode_mic_datagram(&encode_mic_datagram(0, 0, &[]))
.unwrap()
.2
.is_empty());
}
#[test]
fn rich_input_roundtrip() {
for ev in [
RichInput::Touchpad {
pad: 1,
finger: 0,
active: true,
x: 40000,
y: 12345,
},
RichInput::Motion {
pad: 0,
gyro: [-100, 200, -300],
accel: [16384, -8192, 1],
},
RichInput::TouchpadEx {
pad: 2,
surface: 1,
finger: 1,
touch: true,
click: false,
x: -12345,
y: 30000,
pressure: 4000,
},
] {
let d = ev.encode();
assert_eq!(d[0], RICH_INPUT_MAGIC);
assert_eq!(RichInput::decode(&d), Some(ev));
}
// A raw Triton state report rides the plane verbatim (as-is SC2 passthrough).
let mut data = [0u8; HID_REPORT_MAX];
data[0] = 0x42; // ID_TRITON_CONTROLLER_STATE
for (i, b) in data.iter_mut().enumerate().take(46).skip(1) {
*b = i as u8;
}
let raw = RichInput::HidReport {
pad: 3,
len: 46,
data,
};
let d = raw.encode();
assert_eq!(d.len(), 4 + 46); // tag + kind + pad + len + body — no fixed-array padding
assert_eq!(RichInput::decode(&d), Some(raw));
// A torn HidReport truncates to what arrived rather than over-reading (len clamps).
assert_eq!(
RichInput::decode(&d[..20]),
Some(RichInput::HidReport {
pad: 3,
len: 16,
data: {
let mut t = [0u8; HID_REPORT_MAX];
t[..16].copy_from_slice(&data[..16]);
t
},
})
);
// Disjoint from the fixed input datagram (0xC8); unknown kind + truncation → None.
assert!(RichInput::decode(&[crate::input::INPUT_MAGIC; 18]).is_none());
assert!(RichInput::decode(&[RICH_INPUT_MAGIC, 0x7F]).is_none()); // unknown kind
assert!(RichInput::decode(&[RICH_INPUT_MAGIC, RICH_TOUCHPAD, 0]).is_none()); // short
assert!(RichInput::decode(&[RICH_INPUT_MAGIC, RICH_TOUCHPAD_EX, 0, 0, 0, 0]).is_none());
// short
}
#[test]
fn hid_output_roundtrip() {
let cases = [
HidOutput::Led {
pad: 2,
r: 0xAA,
g: 0xBB,
b: 0xCC,
},
HidOutput::PlayerLeds {
pad: 0,
bits: 0b10101,
},
HidOutput::Trigger {
pad: 1,
which: 1,
effect: vec![0x26, 0x90, 0xA0, 0xFF, 0x00, 0x00],
},
HidOutput::TrackpadHaptic {
pad: 0,
side: 1,
amplitude: 0x1234,
period: 0x5678,
count: 9,
},
// A raw Triton rumble output report (as-is SC2 passthrough, host→client).
HidOutput::HidRaw {
pad: 1,
kind: HID_RAW_OUTPUT,
data: vec![0x80, 0, 0, 0, 0x34, 0x12, 0, 0x78, 0x56, 0],
},
// A raw 64-byte feature report (lizard-off / IMU-enable settings write).
HidOutput::HidRaw {
pad: 0,
kind: HID_RAW_FEATURE,
data: {
let mut f = vec![0u8; HID_REPORT_MAX];
f[0] = 1; // Triton feature reports ride report id 1
f[1] = 0x87; // ID_SET_SETTINGS_VALUES
f
},
},
];
for ev in &cases {
let d = ev.encode();
assert_eq!(d[0], HIDOUT_MAGIC);
assert_eq!(HidOutput::decode(&d).as_ref(), Some(ev));
}
assert!(HidOutput::decode(&[HIDOUT_MAGIC, 0x7F]).is_none()); // unknown kind
// A rich-input datagram is not a HID-output datagram.
assert!(HidOutput::decode(
&RichInput::Motion {
pad: 0,
gyro: [0; 3],
accel: [0; 3]
}
.encode()
)
.is_none());
}
}
+26 -2
View File
@@ -26,10 +26,12 @@ fn stream_transport() -> Arc<quinn::TransportConfig> {
/// path is PINGed at least twice per window and a single lost PING (wifi roam / brief blip) won't /// path is PINGed at least twice per window and a single lost PING (wifi roam / brief blip) won't
/// false-close. `idle` is clamped to a ≥1s floor so a misconfigured tiny value can't tear live /// false-close. `idle` is clamped to a ≥1s floor so a misconfigured tiny value can't tear live
/// sessions down. Active sessions are unaffected either way: video keeps the connection live and /// sessions down. Active sessions are unaffected either way: video keeps the connection live and
/// the keep-alive holds it open through quiet control periods. /// the keep-alive holds it open through quiet control periods. Clamped to a 1 s..1 h window:
/// the ceiling keeps an absurd operator-supplied value inside QUIC's VarInt millisecond range,
/// so the conversion below genuinely cannot fail (it used to panic host startup instead).
fn stream_transport_idle(idle: std::time::Duration) -> Arc<quinn::TransportConfig> { fn stream_transport_idle(idle: std::time::Duration) -> Arc<quinn::TransportConfig> {
use std::time::Duration; use std::time::Duration;
let idle = idle.max(Duration::from_secs(1)); let idle = idle.clamp(Duration::from_secs(1), Duration::from_secs(3600));
let keep_alive = (idle / 2).min(Duration::from_secs(4)); let keep_alive = (idle / 2).min(Duration::from_secs(4));
let mut t = quinn::TransportConfig::default(); let mut t = quinn::TransportConfig::default();
t.max_idle_timeout(Some( t.max_idle_timeout(Some(
@@ -293,3 +295,25 @@ impl rustls::server::danger::ClientCertVerifier for AcceptAnyClientCert {
.supported_schemes() .supported_schemes()
} }
} }
#[cfg(test)]
mod tests {
use crate::quic::endpoint;
#[test]
fn fingerprint_is_sha256_of_der() {
// Stable across calls, distinct for distinct certs.
let a = endpoint::cert_fingerprint(b"cert-a");
assert_eq!(a, endpoint::cert_fingerprint(b"cert-a"));
assert_ne!(a, endpoint::cert_fingerprint(b"cert-b"));
}
#[test]
fn absurd_idle_timeout_is_clamped_not_a_panic() {
// The conversion to quinn's IdleTimeout fails past the QUIC VarInt millisecond
// ceiling — an operator-supplied huge PUNKTFUNK_IDLE_TIMEOUT_MS used to panic host
// startup through the `expect`. Both extremes must construct.
let _ = super::stream_transport_idle(std::time::Duration::MAX);
let _ = super::stream_transport_idle(std::time::Duration::ZERO);
}
}
+601 -2
View File
@@ -182,8 +182,9 @@ pub struct Start {
} }
/// Truncate `s` to at most `max` bytes on a UTF-8 char boundary (so a multi-byte char straddling /// Truncate `s` to at most `max` bytes on a UTF-8 char boundary (so a multi-byte char straddling
/// the cap is dropped whole, never split). Shared by Hello's length-prefixed name/launch fields. /// the cap is dropped whole, never split). Shared by Hello's length-prefixed name/launch fields
fn truncate_to(s: &str, max: usize) -> &str { /// and [`PairRequest`](super::PairRequest)'s copy of the same name cap.
pub(super) fn truncate_to(s: &str, max: usize) -> &str {
if s.len() <= max { if s.len() <= max {
return s; return s;
} }
@@ -495,3 +496,601 @@ impl Start {
}) })
} }
} }
#[cfg(test)]
mod tests {
use crate::config::{CompositorPref, FecConfig, FecScheme, GamepadPref, Mode, Role};
use crate::quic::*;
#[test]
fn welcome_roundtrip() {
let w = Welcome {
abi_version: 1,
udp_port: 9999,
mode: Mode {
width: 2560,
height: 1440,
refresh_hz: 240,
},
fec: FecConfig {
scheme: FecScheme::Gf16,
fec_percent: 20,
max_data_per_block: 4096,
},
shard_payload: 1200,
encrypt: true,
key: [7u8; 16],
salt: [1, 2, 3, 4],
frames: 600,
compositor: CompositorPref::Gamescope,
gamepad: GamepadPref::DualSense,
bitrate_kbps: 50_000,
bit_depth: 10,
color: ColorInfo::HDR10_BT2020_PQ,
chroma_format: CHROMA_IDC_444,
audio_channels: 2,
codec: CODEC_H264, // exercise a non-default codec through the roundtrip
host_caps: HOST_CAP_GAMEPAD_STATE,
};
assert_eq!(Welcome::decode(&w.encode()).unwrap(), w);
// Client-side reassembler ceiling derives from the negotiated rate: 4x the average frame at
// 50 Mbps/240 Hz is ~104 KB, so the 8 MiB floor governs. The host keeps the p1_defaults
// bound (it never reassembles video), as does a client of a bitrate-0 (older) host.
let cc = w.session_config(Role::Client);
assert_eq!(cc.max_frame_bytes, 8 << 20);
cc.validate().expect("derived client config validates");
assert_eq!(w.session_config(Role::Host).max_frame_bytes, 64 << 20);
let old_host = Welcome {
bitrate_kbps: 0,
..w
};
assert_eq!(
old_host.session_config(Role::Client).max_frame_bytes,
64 << 20
);
// A high-rate mode scales past the floor: 1.5 Gbps at 60 Hz = 4 x 3.125 MB = 12.5 MB.
let fat = Welcome {
bitrate_kbps: 1_500_000,
mode: Mode {
width: 5120,
height: 1440,
refresh_hz: 60,
},
..w
};
let derived = fat.session_config(Role::Client).max_frame_bytes;
assert_eq!(derived, 4 * 1_500_000 * 125 / 60);
assert!(derived > (8 << 20) && derived < (64 << 20));
}
#[test]
fn codec_negotiation_and_back_compat() {
// resolve_codec precedence (HEVC > AV1 > H.264), no preference (0).
assert_eq!(
resolve_codec(CODEC_H264 | CODEC_HEVC, CODEC_HEVC | CODEC_AV1, 0),
Some(CODEC_HEVC)
);
assert_eq!(
resolve_codec(CODEC_H264 | CODEC_AV1, CODEC_AV1 | CODEC_H264, 0),
Some(CODEC_AV1)
);
assert_eq!(resolve_codec(CODEC_H264, CODEC_H264, 0), Some(CODEC_H264));
// A software host (H.264 only) + an HEVC-only client share nothing → refuse.
assert_eq!(resolve_codec(CODEC_HEVC, CODEC_H264, 0), None);
// An older client (0 = no codec byte) is treated as HEVC-only.
assert_eq!(
resolve_codec(0, CODEC_HEVC | CODEC_H264, 0),
Some(CODEC_HEVC)
);
assert_eq!(resolve_codec(0, CODEC_H264, 0), None);
// Soft preference: honored when the host can also emit it, overriding precedence...
assert_eq!(
resolve_codec(CODEC_H264 | CODEC_HEVC, CODEC_H264 | CODEC_HEVC, CODEC_H264),
Some(CODEC_H264)
);
assert_eq!(
resolve_codec(CODEC_HEVC | CODEC_AV1, CODEC_HEVC | CODEC_AV1, CODEC_AV1),
Some(CODEC_AV1)
);
// ...but falls back to precedence when the preferred codec isn't in the shared set.
assert_eq!(
resolve_codec(CODEC_HEVC | CODEC_H264, CODEC_HEVC | CODEC_H264, CODEC_AV1),
Some(CODEC_HEVC)
);
// A preference the host can't emit still can't rescue a no-shared-codec case.
assert_eq!(resolve_codec(CODEC_HEVC, CODEC_H264, CODEC_HEVC), None);
// PyroWave is opt-in ONLY (plan §3): mutual support NEVER auto-selects it — the ladder
// ignores it entirely...
assert_eq!(
resolve_codec(CODEC_HEVC | CODEC_PYROWAVE, CODEC_HEVC | CODEC_PYROWAVE, 0),
Some(CODEC_HEVC)
);
// ...even when it is the ONLY shared codec (an all-intra 200 Mbps stream must never be a
// silent fallback)...
assert_eq!(resolve_codec(CODEC_PYROWAVE, CODEC_PYROWAVE, 0), None);
// ...it is reachable exclusively through the client's explicit preference.
assert_eq!(
resolve_codec(
CODEC_HEVC | CODEC_PYROWAVE,
CODEC_HEVC | CODEC_PYROWAVE,
CODEC_PYROWAVE
),
Some(CODEC_PYROWAVE)
);
// A pyrowave preference against a host without the backend falls back to the ladder.
assert_eq!(
resolve_codec(CODEC_HEVC | CODEC_PYROWAVE, CODEC_HEVC, CODEC_PYROWAVE),
Some(CODEC_HEVC)
);
// And the negotiated bit SURVIVES the Welcome wire roundtrip — the decode whitelist
// once folded unknown codec bytes (incl. PyroWave) to HEVC, which sent wavelet AUs
// into an FFmpeg HEVC decoder on the first on-glass run.
let mut pw_w = Welcome::decode(
&Welcome {
abi_version: 2,
udp_port: 1,
mode: Mode {
width: 1280,
height: 720,
refresh_hz: 60,
},
fec: FecConfig {
scheme: FecScheme::Gf16,
fec_percent: 0,
max_data_per_block: 1024,
},
shard_payload: 1024,
encrypt: false,
key: [0; 16],
salt: [0; 4],
frames: 0,
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
bit_depth: 8,
color: ColorInfo::SDR_BT709,
chroma_format: CHROMA_IDC_420,
audio_channels: 2,
codec: CODEC_PYROWAVE,
host_caps: 0,
}
.encode(),
)
.unwrap();
assert_eq!(pw_w.codec, CODEC_PYROWAVE);
// A genuinely unknown future bit still folds to the HEVC default.
pw_w.codec = 0x40;
assert_eq!(Welcome::decode(&pw_w.encode()).unwrap().codec, CODEC_HEVC);
// A Hello advertising codecs roundtrips, and the wire form of a codec-only Hello decodes on
// a build that ignores the trailing byte (back-compat: extra bytes are skipped).
let h = Hello {
abi_version: 2,
mode: Mode {
width: 1280,
height: 720,
refresh_hz: 60,
},
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
name: None,
launch: None,
video_caps: 0,
audio_channels: 2, // stereo — forces the video_caps/audio_channels placeholders
video_codecs: CODEC_H264 | CODEC_HEVC,
preferred_codec: CODEC_H264,
display_hdr: None,
};
let enc = h.encode();
let dec = Hello::decode(&enc).unwrap();
assert_eq!(dec.video_codecs, CODEC_H264 | CODEC_HEVC);
assert_eq!(dec.preferred_codec, CODEC_H264);
// Drop the preferred_codec byte → still decodes, video_codecs intact, preference gone.
let no_pref = &enc[..enc.len() - 1];
assert_eq!(
Hello::decode(no_pref).unwrap().video_codecs,
CODEC_H264 | CODEC_HEVC
);
assert_eq!(Hello::decode(no_pref).unwrap().preferred_codec, 0);
// A pre-codec Hello (no video_codecs/preferred bytes) decodes to 0 → HEVC-only.
let legacy = &enc[..enc.len() - 2];
assert_eq!(Hello::decode(legacy).unwrap().video_codecs, 0);
assert_eq!(Hello::decode(legacy).unwrap().preferred_codec, 0);
// A pre-codec Welcome (no codec byte) decodes to HEVC.
let mut w = Welcome::decode(
&Welcome {
abi_version: 2,
udp_port: 1,
mode: h.mode,
fec: FecConfig {
scheme: FecScheme::Gf16,
fec_percent: 0,
max_data_per_block: 1024,
},
shard_payload: 1024,
encrypt: false,
key: [0; 16],
salt: [0; 4],
frames: 0,
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
bit_depth: 8,
color: ColorInfo::SDR_BT709,
chroma_format: CHROMA_IDC_420,
audio_channels: 2,
codec: CODEC_H264,
host_caps: 0,
}
.encode(),
)
.unwrap();
assert_eq!(w.codec, CODEC_H264);
w.codec = CODEC_HEVC;
let wenc = w.encode();
assert_eq!(
Welcome::decode(&wenc[..wenc.len() - 1]).unwrap().codec,
CODEC_HEVC
);
}
#[test]
fn hello_start_roundtrip() {
let h = Hello {
abi_version: 1,
mode: Mode {
width: 1280,
height: 720,
refresh_hz: 120,
},
compositor: CompositorPref::Kwin,
gamepad: GamepadPref::DualSense,
bitrate_kbps: 25_000,
name: Some("Test Device".into()),
launch: Some("steam:570".into()),
video_caps: VIDEO_CAP_10BIT,
audio_channels: 2,
video_codecs: CODEC_H264 | CODEC_HEVC, // exercise the codec bitfield roundtrip
preferred_codec: CODEC_HEVC,
display_hdr: None,
};
assert_eq!(Hello::decode(&h.encode()).unwrap(), h);
let s = Start {
client_udp_port: 1234,
};
assert_eq!(Start::decode(&s.encode()).unwrap(), s);
}
#[test]
fn hello_welcome_compositor_back_compat() {
// Trailing optional bytes (compositor at 20/53, gamepad at 21/54): a legacy peer's
// shorter message still decodes (missing fields = Auto), and a legacy peer reading a
// new message ignores the trailing bytes. Simulate both directions by truncation.
let h = Hello {
abi_version: 2,
mode: Mode {
width: 1920,
height: 1080,
refresh_hz: 60,
},
compositor: CompositorPref::Mutter,
gamepad: GamepadPref::DualSense,
bitrate_kbps: 80_000,
name: None,
launch: None,
video_caps: 0,
audio_channels: 2,
video_codecs: 0,
preferred_codec: 0,
display_hdr: None,
};
let enc = h.encode();
assert_eq!(enc.len(), 26);
// Legacy (20-byte) Hello → both Auto, no bitrate, mode intact.
let legacy = Hello::decode(&enc[..20]).unwrap();
assert_eq!(legacy.compositor, CompositorPref::Auto);
assert_eq!(legacy.gamepad, GamepadPref::Auto);
assert_eq!(legacy.bitrate_kbps, 0);
assert_eq!(legacy.mode, h.mode);
// Compositor-era (21-byte) Hello → compositor intact, gamepad Auto.
let mid = Hello::decode(&enc[..21]).unwrap();
assert_eq!(mid.compositor, CompositorPref::Mutter);
assert_eq!(mid.gamepad, GamepadPref::Auto);
// Gamepad-era (22-byte) Hello → compositor + gamepad intact, bitrate 0 (host default).
let pre_bitrate = Hello::decode(&enc[..22]).unwrap();
assert_eq!(pre_bitrate.gamepad, GamepadPref::DualSense);
assert_eq!(pre_bitrate.bitrate_kbps, 0);
// Full message → bitrate intact.
assert_eq!(Hello::decode(&enc).unwrap().bitrate_kbps, 80_000);
let w = Welcome {
abi_version: 2,
udp_port: 7000,
mode: h.mode,
fec: FecConfig {
scheme: FecScheme::Gf16,
fec_percent: 20,
max_data_per_block: 4096,
},
shard_payload: 1200,
encrypt: true,
key: [3u8; 16],
salt: [9, 8, 7, 6],
frames: 0,
compositor: CompositorPref::Kwin,
gamepad: GamepadPref::Xbox360,
bitrate_kbps: 120_000,
bit_depth: 10,
color: ColorInfo::HDR10_BT2020_PQ,
chroma_format: CHROMA_IDC_444,
audio_channels: 6, // 5.1 — exercises the non-default trailing byte
codec: CODEC_HEVC,
host_caps: HOST_CAP_GAMEPAD_STATE,
};
let wenc = w.encode();
assert_eq!(wenc.len(), 68); // 60 base + 4 colour + chroma + audio-channels + codec + host-caps
let legacy_w = Welcome::decode(&wenc[..53]).unwrap();
assert_eq!(legacy_w.compositor, CompositorPref::Auto);
assert_eq!(legacy_w.gamepad, GamepadPref::Auto);
assert_eq!(legacy_w.bitrate_kbps, 0);
assert_eq!(legacy_w.frames, 0);
assert_eq!(legacy_w.key, w.key);
let mid_w = Welcome::decode(&wenc[..54]).unwrap();
assert_eq!(mid_w.compositor, CompositorPref::Kwin);
assert_eq!(mid_w.gamepad, GamepadPref::Auto);
// Gamepad-era (55-byte) Welcome → gamepad intact, bitrate 0 (unknown).
let pre_bitrate_w = Welcome::decode(&wenc[..55]).unwrap();
assert_eq!(pre_bitrate_w.gamepad, GamepadPref::Xbox360);
assert_eq!(pre_bitrate_w.bitrate_kbps, 0);
assert_eq!(pre_bitrate_w.bit_depth, 8); // older host (no trailing byte) → 8-bit assumed
assert_eq!(legacy_w.bit_depth, 8);
// A pre-colour (60-byte) Welcome → SDR BT.709 (the only colour those hosts produced).
let pre_color_w = Welcome::decode(&wenc[..60]).unwrap();
assert_eq!(pre_color_w.bit_depth, 10);
assert_eq!(pre_color_w.color, ColorInfo::SDR_BT709);
assert_eq!(pre_color_w.chroma_format, CHROMA_IDC_420); // pre-chroma host → 4:2:0
assert_eq!(legacy_w.color, ColorInfo::SDR_BT709);
assert_eq!(legacy_w.chroma_format, CHROMA_IDC_420);
// A pre-chroma (64-byte) Welcome carries colour but no chroma/audio bytes → 4:2:0 + stereo.
let pre_chroma_w = Welcome::decode(&wenc[..64]).unwrap();
assert_eq!(pre_chroma_w.color, ColorInfo::HDR10_BT2020_PQ);
assert_eq!(pre_chroma_w.chroma_format, CHROMA_IDC_420);
assert_eq!(pre_chroma_w.audio_channels, 2); // audio byte (offset 65) absent → stereo
// A pre-audio (65-byte) Welcome carries chroma but no audio byte → 4:4:4 + stereo.
let pre_audio_w = Welcome::decode(&wenc[..65]).unwrap();
assert_eq!(pre_audio_w.chroma_format, CHROMA_IDC_444);
assert_eq!(pre_audio_w.audio_channels, 2);
assert_eq!(Welcome::decode(&wenc).unwrap().bitrate_kbps, 120_000);
assert_eq!(Welcome::decode(&wenc).unwrap().bit_depth, 10); // full form carries it
assert_eq!(
Welcome::decode(&wenc).unwrap().color,
ColorInfo::HDR10_BT2020_PQ
);
assert_eq!(
Welcome::decode(&wenc).unwrap().chroma_format,
CHROMA_IDC_444
); // full form carries 4:4:4
assert_eq!(Welcome::decode(&wenc).unwrap().audio_channels, 6); // ...and 5.1
// A pre-host-caps (67-byte) Welcome → 0 (legacy input only); the full form carries the bit.
assert_eq!(Welcome::decode(&wenc[..67]).unwrap().host_caps, 0);
assert_eq!(
Welcome::decode(&wenc).unwrap().host_caps,
HOST_CAP_GAMEPAD_STATE
);
}
#[test]
fn hello_name_roundtrip_and_back_compat() {
let base = Hello {
abi_version: 2,
mode: Mode {
width: 1280,
height: 720,
refresh_hz: 60,
},
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
name: Some("Enrico's MacBook".into()),
launch: None,
video_caps: 0,
audio_channels: 2,
video_codecs: 0,
preferred_codec: 0,
display_hdr: None,
};
let enc = base.encode();
assert_eq!(
Hello::decode(&enc).unwrap().name.as_deref(),
Some("Enrico's MacBook")
);
// A bitrate-era (26-byte) peer reading a named Hello ignores the trailing name; a named
// host reading a bitrate-era Hello decodes name = None.
assert_eq!(Hello::decode(&enc[..26]).unwrap().name, None);
// No name → wire form is byte-identical to the bitrate-era message (26 bytes).
let unnamed = Hello {
name: None,
..base.clone()
};
assert_eq!(unnamed.encode().len(), 26);
// Over-long names truncate to a char boundary within HELLO_NAME_MAX on encode.
let long = Hello {
name: Some(format!("{}ü", "x".repeat(HELLO_NAME_MAX - 1))), // ü straddles the cap
..base.clone()
};
let dec = Hello::decode(&long.encode()).unwrap();
let n = dec.name.expect("truncated name still present");
assert!(n.len() <= HELLO_NAME_MAX && n.starts_with('x'));
// A corrupt length byte (longer than the buffer) or bad UTF-8 degrades to None, never Err.
let mut bad_len = unnamed.encode();
bad_len.push(40); // claims 40 name bytes, none follow
assert_eq!(Hello::decode(&bad_len).unwrap().name, None);
let mut bad_utf8 = unnamed.encode();
bad_utf8.extend_from_slice(&[2, 0xFF, 0xFE]);
assert_eq!(Hello::decode(&bad_utf8).unwrap().name, None);
}
#[test]
fn hello_launch_roundtrip_and_back_compat() {
let base = Hello {
abi_version: 2,
mode: Mode {
width: 1920,
height: 1080,
refresh_hz: 60,
},
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
name: None,
launch: None,
video_caps: 0,
audio_channels: 2,
video_codecs: 0,
preferred_codec: 0,
display_hdr: None,
};
// launch alone (no name): a zero-length name placeholder keeps the offset deterministic.
let with_launch = Hello {
launch: Some("steam:570".into()),
..base.clone()
};
assert_eq!(Hello::decode(&with_launch.encode()).unwrap(), with_launch);
// launch + name together.
let both = Hello {
name: Some("Enrico's Mac".into()),
launch: Some("custom:abc123".into()),
..base.clone()
};
assert_eq!(Hello::decode(&both.encode()).unwrap(), both);
// name but no launch (a name-era client): launch decodes None.
let name_only = Hello {
name: Some("Enrico's Mac".into()),
..base.clone()
};
assert_eq!(Hello::decode(&name_only.encode()).unwrap().launch, None);
// Neither field → still the 26-byte bitrate-era form (no launch placeholder emitted).
assert_eq!(base.encode().len(), 26);
assert_eq!(Hello::decode(&base.encode()).unwrap().launch, None);
// A bitrate-era (26-byte) peer reading a launch-bearing Hello ignores it.
assert_eq!(
Hello::decode(&with_launch.encode()[..26]).unwrap().launch,
None
);
// Over-long ids truncate on a char boundary within HELLO_LAUNCH_MAX.
let long = Hello {
launch: Some(format!("{}ü", "x".repeat(HELLO_LAUNCH_MAX - 1))),
..base.clone()
};
let dec = Hello::decode(&long.encode())
.unwrap()
.launch
.expect("present");
assert!(dec.len() <= HELLO_LAUNCH_MAX && dec.starts_with('x'));
}
#[test]
fn hello_display_hdr_roundtrip_and_back_compat() {
let base = Hello {
abi_version: 2,
mode: Mode {
width: 3840,
height: 2160,
refresh_hz: 120,
},
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
name: None,
launch: None,
video_caps: VIDEO_CAP_10BIT | VIDEO_CAP_HDR,
audio_channels: 2,
video_codecs: 0,
preferred_codec: 0,
display_hdr: None,
};
// A real client-panel volume (P3 primaries, 800-nit peak, 0.05-nit floor, 400-nit FALL).
let vol = HdrMeta {
display_primaries: [[13250, 34500], [7500, 3000], [34000, 16000]], // G, B, R
white_point: [15635, 16450], // D65
max_display_mastering_luminance: 8_000_000, // 800 nits
min_display_mastering_luminance: 500, // 0.05 nits
max_cll: 0,
max_fall: 400,
};
let with_hdr = Hello {
display_hdr: Some(vol),
..base.clone()
};
// Full roundtrip, including the forced placeholders for the earlier trailing fields.
assert_eq!(Hello::decode(&with_hdr.encode()).unwrap(), with_hdr);
// display_hdr alone (every earlier optional at its default) still lands at a deterministic
// offset — the placeholder discipline holds through the whole tail.
let hdr_only = Hello {
video_caps: 0,
display_hdr: Some(vol),
..base.clone()
};
assert_eq!(Hello::decode(&hdr_only.encode()).unwrap(), hdr_only);
// An older host reading a display_hdr-bearing Hello ignores the trailing block (its decode
// stops at preferred_codec); a new host reading an older client's Hello gets None.
let enc = with_hdr.encode();
assert_eq!(
Hello::decode(&enc[..enc.len() - HDR_META_BODY_LEN]).unwrap(),
Hello {
display_hdr: None,
..with_hdr.clone()
}
);
assert_eq!(Hello::decode(&base.encode()).unwrap().display_hdr, None);
// A TRUNCATED trailing block (mid-datagram cut) degrades to None, never a partial read.
assert_eq!(
Hello::decode(&enc[..enc.len() - 1]).unwrap().display_hdr,
None
);
// Exact wire length: 26 bitrate-era bytes + the 6 forced single-byte placeholders
// (name len, launch len, video_caps, audio_channels, video_codecs, preferred_codec) + the body.
assert_eq!(hdr_only.encode().len(), 26 + 6 + HDR_META_BODY_LEN);
}
#[test]
fn control_messages_disjoint_from_hello() {
// A Hello uses MAGIC (PKF1); control messages use CTL_MAGIC (PKFc). No Hello — at
// any abi_version — can be misparsed as a control message, and vice-versa.
for abi in [1u32, 2, 16, 0x10, 0x0113, 0x1410] {
let h = Hello {
abi_version: abi,
mode: Mode {
width: 1280,
height: 720,
refresh_hz: 60,
},
compositor: CompositorPref::Auto,
gamepad: GamepadPref::Auto,
bitrate_kbps: 0,
name: None,
launch: None,
video_caps: 0,
audio_channels: 2,
video_codecs: 0,
preferred_codec: 0,
display_hdr: None,
}
.encode();
assert!(PairRequest::decode(&h).is_err(), "abi {abi} parsed as pair");
assert!(Reconfigure::decode(&h).is_err());
}
// And a PairRequest never parses as a Hello.
let pr = PairRequest {
name: "x".into(),
spake_a: vec![0u8; 33],
}
.encode();
assert!(Hello::decode(&pr).is_err());
}
}
+157
View File
@@ -1,6 +1,14 @@
//! Length-prefixed framing for QUIC control-stream messages: a `u16` length header followed by the //! Length-prefixed framing for QUIC control-stream messages: a `u16` length header followed by the
//! payload, bounded at 64 KiB (control messages are tiny). //! payload, bounded at 64 KiB (control messages are tiny).
/// Read one framed message (bounded at 64 KiB — control messages are tiny). /// Read one framed message (bounded at 64 KiB — control messages are tiny).
///
/// **Not cancel-safe**: it frames with two `quinn::RecvStream::read_exact` calls, and quinn
/// documents `read_exact` as not cancel-safe (the bytes it has already taken out of the stream
/// live only in the future's own buffer, and nothing puts them back on drop). Dropping a
/// partially-progressed future therefore destroys the bytes it consumed and misaligns every
/// subsequent read on that stream. Use it only where the read runs to completion — the sequential
/// handshake/pairing exchanges. Anything driving a read from a `select!` arm or a
/// `tokio::time::timeout` must use [`MsgReader`] instead.
pub async fn read_msg(recv: &mut quinn::RecvStream) -> std::io::Result<Vec<u8>> { pub async fn read_msg(recv: &mut quinn::RecvStream) -> std::io::Result<Vec<u8>> {
let mut len = [0u8; 2]; let mut len = [0u8; 2];
recv.read_exact(&mut len) recv.read_exact(&mut len)
@@ -14,9 +22,158 @@ pub async fn read_msg(recv: &mut quinn::RecvStream) -> std::io::Result<Vec<u8>>
Ok(buf) Ok(buf)
} }
/// Cancel-safe framed reader for a long-lived control stream.
///
/// Keeps the frame in progress in `buf` rather than inside the read future, so dropping the future
/// — which both control loops do on every iteration where a sibling `select!` arm wins, and which
/// [`clock_sync`](super::clock_sync) does on a read timeout — resumes instead of losing bytes.
/// With the plain [`read_msg`] a control frame that straddles two wakeups (a ~2 KB `ClipOffer`
/// exceeds one QUIC packet; so does any frame whose second half is lost or reordered) left the
/// stream permanently misaligned: the next read took two payload bytes as a length, every later
/// message decoded as garbage and was silently ignored, and a bogus 64 KiB length parked the read
/// forever — killing mode switches, adaptive bitrate, clock re-sync and clipboard for the rest of
/// the session with nothing but a `warn!` in the log.
pub struct MsgReader {
recv: quinn::RecvStream,
/// The frame in progress, length prefix included.
buf: Vec<u8>,
/// Bytes `buf` must reach: 2 while reading the prefix, then `2 + payload length`.
need: usize,
}
impl MsgReader {
pub fn new(recv: quinn::RecvStream) -> Self {
MsgReader {
recv,
buf: Vec::new(),
need: 2,
}
}
/// Read one framed message. Cancel-safe: dropping the future keeps the partial frame, so the
/// next call resumes where this one stopped.
pub async fn read_msg(&mut self) -> std::io::Result<Vec<u8>> {
loop {
while self.buf.len() < self.need {
let mut chunk = [0u8; 2048];
let want = (self.need - self.buf.len()).min(chunk.len());
// `read` IS cancel-safe: it only reports bytes it hands back, and they are
// committed to `self.buf` before the next await point.
match self
.recv
.read(&mut chunk[..want])
.await
.map_err(std::io::Error::other)?
{
Some(n) => self.buf.extend_from_slice(&chunk[..n]),
None => {
return Err(std::io::Error::new(
std::io::ErrorKind::UnexpectedEof,
"control stream finished mid-frame",
))
}
}
}
if self.need == 2 {
self.need = 2 + u16::from_le_bytes([self.buf[0], self.buf[1]]) as usize;
if self.need == 2 {
self.buf.clear();
return Ok(Vec::new()); // zero-length frame
}
} else {
let msg = self.buf.split_off(2);
self.buf.clear();
self.need = 2;
return Ok(msg);
}
}
}
}
/// Write one framed message. /// Write one framed message.
pub async fn write_msg(send: &mut quinn::SendStream, payload: &[u8]) -> std::io::Result<()> { pub async fn write_msg(send: &mut quinn::SendStream, payload: &[u8]) -> std::io::Result<()> {
send.write_all(&super::frame(payload)) send.write_all(&super::frame(payload))
.await .await
.map_err(std::io::Error::other) .map_err(std::io::Error::other)
} }
/// The control stream is read from a `select!` arm on both peers, so the read future is dropped
/// routinely — and quinn documents `read_exact` (what `io::read_msg` uses) as NOT cancel-safe.
/// [`io::MsgReader`] must survive that: the partial frame lives in the reader, not the future.
#[cfg(test)]
mod tests {
use crate::quic::io;
use crate::quic::test_util::connect_pair;
/// A frame whose halves land in different wakeups, with the read cancelled in between, must
/// still be delivered whole — and the NEXT frame must decode correctly too. Without a
/// resumable reader the consumed length prefix is lost, the following read takes two payload
/// bytes as a length, and every later control message is garbage for the rest of the session.
#[tokio::test]
async fn cancelled_mid_frame_read_resumes_without_desync() {
let (_server_ep, _client_ep, host_conn, client_conn) = connect_pair().await;
let first = b"the-frame-that-straddles-two-wakeups".to_vec();
let second = b"the-frame-after-it".to_vec();
let (f1, f2) = (first.clone(), second.clone());
let writer = tokio::spawn(async move {
let (mut send, _recv) = host_conn.open_bi().await.expect("open bi");
let framed = crate::quic::frame(&f1);
// Length prefix + only part of the payload, then a real pause: this is the ClipOffer
// -sized frame split across two QUIC packets that made the bug reachable.
let split = 2 + f1.len() / 3;
send.write_all(&framed[..split]).await.expect("write head");
tokio::time::sleep(std::time::Duration::from_millis(120)).await;
send.write_all(&framed[split..]).await.expect("write tail");
send.write_all(&crate::quic::frame(&f2))
.await
.expect("write second");
tokio::time::sleep(std::time::Duration::from_millis(200)).await;
host_conn
});
let (_send, recv) = client_conn.accept_bi().await.expect("accept bi");
let mut reader = io::MsgReader::new(recv);
// Cancel mid-frame — exactly what a sibling `select!` arm does.
let cancelled =
tokio::time::timeout(std::time::Duration::from_millis(30), reader.read_msg()).await;
assert!(
cancelled.is_err(),
"the head-only frame must not complete yet (test setup)"
);
let got = tokio::time::timeout(std::time::Duration::from_secs(5), reader.read_msg())
.await
.expect("first frame must arrive after resuming")
.expect("first frame reads cleanly");
assert_eq!(got, first, "the cancelled read must resume, not lose bytes");
let got2 = tokio::time::timeout(std::time::Duration::from_secs(5), reader.read_msg())
.await
.expect("second frame must arrive")
.expect("second frame reads cleanly");
assert_eq!(got2, second, "stream must still be framed correctly");
let _host_conn = writer.await.unwrap();
}
/// A zero-length frame is a legal encoding and must not stall the reader or eat the next one.
#[tokio::test]
async fn zero_length_frame_round_trips() {
let (_server_ep, _client_ep, host_conn, client_conn) = connect_pair().await;
let writer = tokio::spawn(async move {
let (mut send, _recv) = host_conn.open_bi().await.expect("open bi");
send.write_all(&crate::quic::frame(&[])).await.unwrap();
send.write_all(&crate::quic::frame(b"after")).await.unwrap();
tokio::time::sleep(std::time::Duration::from_millis(200)).await;
host_conn
});
let (_send, recv) = client_conn.accept_bi().await.expect("accept bi");
let mut reader = io::MsgReader::new(recv);
assert!(reader.read_msg().await.unwrap().is_empty());
assert_eq!(reader.read_msg().await.unwrap(), b"after");
let _host_conn = writer.await.unwrap();
}
}
+9 -6
View File
@@ -22,11 +22,14 @@
//! reported back for persisting). The data plane adds AES-GCM on top. //! reported back for persisting). The data plane adds AES-GCM on top.
//! All integers little-endian; every message is `u16 length || payload`. //! All integers little-endian; every message is `u16 length || payload`.
//! //!
//! Split by concern (networking-audit deferred plan §3 — a pure move): [`msgs`] the //! Split by concern (networking-audit deferred plan §3 — a pure move): `handshake` the
//! handshake + typed control messages, [`pake`] the pairing SPAKE2, [`datagram`] the //! positional Hello/Welcome/Start codecs, `caps` the capability/codec-negotiation
//! 0xC90xCF plane codecs, [`io`] framed stream IO, [`clock`] skew estimation + mid-stream //! vocabulary, `control` the typed control + clipboard messages, `pairing` the pairing
//! re-sync, [`endpoint`] the quinn constructors. Every item is re-exported here, so all //! message codecs with [`pake`] the SPAKE2 itself, `datagram` the 0xC90xCF plane codecs,
//! existing `crate::quic::X` paths compile unchanged. //! [`io`] framed stream IO, `clock` skew estimation + mid-stream re-sync, [`endpoint`] the
//! quinn constructors, [`clipstream`] the per-transfer clipboard fetch streams. Every item
//! is re-exported here, so all existing `crate::quic::X` paths compile unchanged; each
//! module's tests sit at its own foot.
/// Protocol magic + version, first bytes of the positional handshake (Hello/Welcome/Start). /// Protocol magic + version, first bytes of the positional handshake (Hello/Welcome/Start).
pub const MAGIC: &[u8; 4] = b"PKF1"; pub const MAGIC: &[u8; 4] = b"PKF1";
@@ -76,4 +79,4 @@ pub use pairing::*;
pub use crate::reject::*; pub use crate::reject::*;
#[cfg(test)] #[cfg(test)]
mod tests; pub(crate) mod test_util;
+53 -3
View File
@@ -78,13 +78,16 @@ fn get_bytes(b: &[u8], off: usize) -> Result<(&[u8], usize)> {
impl PairRequest { impl PairRequest {
pub fn encode(&self) -> Vec<u8> { pub fn encode(&self) -> Vec<u8> {
let name = self.name.as_bytes(); // Same cap, same rule as Hello's copy of this field: truncate on a char boundary —
let n = name.len().min(64); // a raw byte cut mid-sequence put invalid UTF-8 on the wire, and the host showed the
// name with a permanent replacement char in its paired-clients list.
let name = super::handshake::truncate_to(&self.name, HELLO_NAME_MAX).as_bytes();
let n = name.len();
let mut b = Vec::with_capacity(8 + n + self.spake_a.len()); let mut b = Vec::with_capacity(8 + n + self.spake_a.len());
b.extend_from_slice(CTL_MAGIC); b.extend_from_slice(CTL_MAGIC);
b.push(MSG_PAIR_REQUEST); b.push(MSG_PAIR_REQUEST);
b.push(n as u8); b.push(n as u8);
b.extend_from_slice(&name[..n]); b.extend_from_slice(name);
put_bytes(&mut b, &self.spake_a); put_bytes(&mut b, &self.spake_a);
b b
} }
@@ -171,3 +174,50 @@ impl PairResult {
Ok(PairResult { ok: b[5] != 0 }) Ok(PairResult { ok: b[5] != 0 })
} }
} }
#[cfg(test)]
mod tests {
use crate::quic::*;
#[test]
fn pair_messages_roundtrip() {
let pr = PairRequest {
name: "Enrico's Mac".into(),
spake_a: vec![1, 2, 3, 4, 5],
};
assert_eq!(PairRequest::decode(&pr.encode()).unwrap(), pr);
let pc = PairChallenge {
spake_b: vec![9; 33],
confirm: [7u8; 32],
};
assert_eq!(PairChallenge::decode(&pc.encode()).unwrap(), pc);
let pp = PairProof { confirm: [3u8; 32] };
assert_eq!(PairProof::decode(&pp.encode()).unwrap(), pp);
for ok in [true, false] {
assert_eq!(
PairResult::decode(&PairResult { ok }.encode()).unwrap().ok,
ok
);
}
// Length-exact: a truncated/padded PairProof is rejected.
let mut bad = pp.encode();
bad.push(0);
assert!(PairProof::decode(&bad).is_err());
}
#[test]
fn pair_request_name_cap_respects_char_boundaries() {
// A multi-byte char straddling the 64-byte cap must be dropped whole (Hello's rule),
// not split mid-sequence into invalid UTF-8 the host then renders as U+FFFD forever.
let pr = PairRequest {
name: format!("{}\u{00fc}", "x".repeat(HELLO_NAME_MAX - 1)),
spake_a: vec![1, 2, 3],
};
let dec = PairRequest::decode(&pr.encode()).unwrap();
assert!(dec.name.len() <= HELLO_NAME_MAX && dec.name.starts_with('x'));
assert!(
!dec.name.contains('\u{FFFD}'),
"name must never be split mid-char on the wire"
);
}
}
+33
View File
@@ -80,3 +80,36 @@ impl PairingPake {
pub fn verify(expected: &[u8; 32], got: &[u8; 32]) -> bool { pub fn verify(expected: &[u8; 32], got: &[u8; 32]) -> bool {
ct_eq(expected, got) ct_eq(expected, got)
} }
#[cfg(test)]
mod tests {
use crate::quic::pake;
#[test]
fn spake2_pairing_agrees_only_on_matching_pin_and_certs() {
let cfp = [0x11u8; 32];
let hfp = [0x22u8; 32];
// Right PIN, same fingerprint views on both sides → both confirmations agree.
let (ca, ma) = pake::start(true, "4321", &cfp, &hfp);
let (cb, mb) = pake::start(false, "4321", &cfp, &hfp);
let a = ca.finish(&mb).unwrap();
let b = cb.finish(&ma).unwrap();
assert!(pake::verify(&a.host, &b.host) && pake::verify(&a.client, &b.client));
// Wrong PIN → different keys → confirmations DON'T match (one online guess wasted).
let (ca, ma) = pake::start(true, "0000", &cfp, &hfp);
let (cb, mb) = pake::start(false, "4321", &cfp, &hfp);
let a = ca.finish(&mb).unwrap();
let b = cb.finish(&ma).unwrap();
assert!(!pake::verify(&a.client, &b.client));
// MITM: the two legs saw different host certs → no agreement even with the right PIN.
let attacker_hfp = [0x33u8; 32];
let (ca, ma) = pake::start(true, "4321", &cfp, &attacker_hfp);
let (cb, mb) = pake::start(false, "4321", &cfp, &hfp);
let a = ca.finish(&mb).unwrap();
let b = cb.finish(&ma).unwrap();
assert!(!pake::verify(&a.client, &b.client));
}
}
@@ -0,0 +1,28 @@
//! Shared in-process QUIC loopback plumbing for the quic submodule tests.
use super::endpoint;
/// Stand up two loopback quinn endpoints, connect, and return
/// `(server_ep, client_ep, host_conn, client_conn)`. Both endpoints are returned so the caller
/// keeps them in scope — dropping a `quinn::Endpoint` tears down its connections.
pub(crate) async fn connect_pair() -> (
quinn::Endpoint,
quinn::Endpoint,
quinn::Connection,
quinn::Connection,
) {
let server = endpoint::server("127.0.0.1:0".parse().unwrap()).unwrap();
let addr = server.local_addr().unwrap();
let client = endpoint::client_insecure().unwrap();
let accept = tokio::spawn(async move {
let incoming = server.accept().await.expect("incoming connection");
let conn = incoming.await.expect("host side connects");
(server, conn)
});
let client_conn = client
.connect(addr, "punktfunk")
.unwrap()
.await
.expect("client side connects");
let (server, host_conn) = accept.await.unwrap();
(server, client, host_conn, client_conn)
}
File diff suppressed because it is too large Load Diff
+163 -337
View File
@@ -13,7 +13,9 @@ use crate::crypto::SessionCrypto;
use crate::error::{PunktfunkError, Result}; use crate::error::{PunktfunkError, Result};
use crate::fec::{coder_for, ErasureCoder}; use crate::fec::{coder_for, ErasureCoder};
use crate::input::InputEvent; use crate::input::InputEvent;
use crate::packet::{Packetizer, Reassembler, ReassemblerLimits, MAX_DATAGRAM_BYTES}; use crate::packet::{
PacketHeader, Packetizer, Reassembler, ReassemblerLimits, StreamedAu, MAX_DATAGRAM_BYTES,
};
use crate::stats::{Stats, StatsCounters}; use crate::stats::{Stats, StatsCounters};
use crate::transport::Transport; use crate::transport::Transport;
use zerocopy::IntoBytes; use zerocopy::IntoBytes;
@@ -30,6 +32,14 @@ pub struct Frame {
/// ([`crate::packet::USER_FLAG_CHUNK_ALIGNED`]) are ever delivered partial; missing /// ([`crate::packet::USER_FLAG_CHUNK_ALIGNED`]) are ever delivered partial; missing
/// shard ranges are zero-filled at their exact offsets. /// shard ranges are zero-filled at their exact offsets.
pub complete: bool, pub complete: bool,
/// Wall-clock instant (ns since the Unix epoch, CLOCK_REALTIME basis — the same clock the
/// skew handshake compares and the host stamps `pts_ns` with) at which this AU finished
/// reassembly, stamped by [`Session::poll_frame`] as the frame leaves the session. Embedders
/// that previously stamped receipt themselves at the hand-off pull should use this instead:
/// the pull stamp additionally contains the pre-decode queue wait, silently folding any
/// client-side standing backlog into the apparent NETWORK latency. The reassembler itself
/// leaves this 0 (it owns no clock — the stamp is the session boundary's job).
pub received_ns: u64,
} }
/// One end of a stream. Constructed for a single [`Role`]; calling the other role's /// One end of a stream. Constructed for a single [`Role`]; calling the other role's
@@ -82,160 +92,27 @@ pub struct Session {
lane_scratch: Vec<Vec<u8>>, lane_scratch: Vec<Vec<u8>>,
} }
/// Wire-packet count at which a frame's sealing splits across two lanes (plan Phase 1.5): /// Stamp [`Frame::received_ns`] as the frame crosses the session boundary in
/// below it the channel rendezvous (~µs) isn't worth it; at it the halved AES-GCM span /// [`Session::poll_frame`] — completed frames return the moment their last shard lands, so
/// (≥ ~125 µs of ~1 µs/packet work) dwarfs the hand-off. ≈300 KB of wire, i.e. ≥150 Mbps /// stamping at return IS stamping at reassembly completion (µs apart). CLOCK_REALTIME to match
/// at 60 fps — small frames and the probe's ~17-packet AUs stay strictly single-lane. /// `pts_ns` / the skew handshake (deliberately not monotonic — cross-machine latency math).
const TWO_LANE_MIN_PACKETS: usize = 256; fn stamp_received(mut f: Frame) -> Frame {
f.received_ns = std::time::SystemTime::now()
/// One two-lane seal hand-off: the frame's back-half wire buffers, sealed by the worker with .duration_since(std::time::UNIX_EPOCH)
/// nonces `seq_base + i` (the nonce order is deterministic per shard index, which is what .map(|d| d.as_nanos() as u64)
/// makes the split sound). Round-trips through the channels so the buffers return to the pool. .unwrap_or(0);
struct SealJob { f
bufs: Vec<Vec<u8>>,
seq_base: u64,
timed: bool,
/// Worker-lane CPU ns (when `timed`) and the seal outcome, filled in by the worker.
ns: u64,
result: Result<()>,
} }
/// The persistent second seal lane: a worker thread that AES-GCM-seals the back half of a mod perf;
/// large frame's packets while the send thread seals the front half. Rendezvous channels mod replay;
/// (bound 1) — the send thread submits, seals its half, then waits; no per-frame spawn. mod seal;
/// Dropping the struct closes the channel and the worker exits.
struct SealLane {
to_worker: std::sync::mpsc::SyncSender<SealJob>,
from_worker: std::sync::mpsc::Receiver<SealJob>,
}
impl SealLane { pub use perf::{PumpPerf, SealPerf};
fn spawn(crypto: std::sync::Arc<SessionCrypto>) -> Option<SealLane> {
let (to_worker, jobs) = std::sync::mpsc::sync_channel::<SealJob>(1);
let (done_tx, from_worker) = std::sync::mpsc::sync_channel::<SealJob>(1);
std::thread::Builder::new()
.name("punktfunk-seal2".into())
.spawn(move || {
while let Ok(mut job) = jobs.recv() {
let t0 = job.timed.then(std::time::Instant::now);
job.result = seal_wire_slice(&crypto, &mut job.bufs, job.seq_base);
if let Some(t0) = t0 {
job.ns = t0.elapsed().as_nanos() as u64;
}
if done_tx.send(job).is_err() {
break; // session gone mid-frame — nothing left to seal for
}
}
})
.ok()?;
Some(SealLane {
to_worker,
from_worker,
})
}
}
/// Seal a run of pre-written wire buffers in place: buffer `i` is `seq(8) ‖ plaintext ‖ tag use perf::TimedCoder;
/// scratch` and seals over `[8..]` with sequence `seq_base + i` — the exact per-packet layout use replay::{seq_of, ReplayWindow};
/// and nonce order of the fused single-lane path. Shared by both lanes. use seal::{seal_wire_slice, SealJob, SealLane, TWO_LANE_MIN_PACKETS};
fn seal_wire_slice(c: &SessionCrypto, wires: &mut [Vec<u8>], seq_base: u64) -> Result<()> {
for (i, wire) in wires.iter_mut().enumerate() {
c.seal_in_place(seq_base.wrapping_add(i as u64), &mut wire[8..])?;
}
Ok(())
}
/// Accumulated client receive-path stage timings since the last [`Session::take_pump_perf`].
/// Answers "where does the pump core go" at line rate: kernel drain (`recv_ns`) vs AES-GCM
/// (`decrypt_ns`) vs reassembly+FEC (`reasm_ns`, the `Reassembler::push` round-trip including
/// shard copies and block reconstruction). 2026-07-14 sweep context: the pump pegs one core at
/// ~1.5 Gbps wire, ~85% of it userspace — this split is what Phase 2.1 (pooled reassembly) is
/// validated against.
#[derive(Debug, Default, Clone, Copy)]
pub struct PumpPerf {
/// ns inside `recv_batch` (recvmmsg / recvmsg_x), i.e. syscall + kernel copy.
pub recv_ns: u64,
/// ns inside `open_in_place` across all datagrams (AES-128-GCM + replay-window upkeep).
pub decrypt_ns: u64,
/// ns inside `Reassembler::push` (header parse, shard copy, FEC reconstruct, AU assembly).
pub reasm_ns: u64,
/// recv_batch calls (batches) and datagrams processed over the accumulation window.
pub batches: u64,
pub packets: u64,
}
/// Accumulated host send-path stage timings since the last [`Session::take_seal_perf`] (plan
/// Phase 0.4, host half). Answers "where does the send thread go" at rate: FEC parity
/// generation (`fec_ns`, inside [`ErasureCoder::encode_into`]) vs AES-GCM (`seal_ns`,
/// per-packet `seal_in_place`) vs the socket handoff (`sock_ns` — `send_gso`/`sendmmsg`
/// syscalls; the internal submit paths time it here, the paced video path folds its chunk
/// sends in via [`Session::note_sock_ns`]). The Phase 1.5 gate reads off this split: build
/// two-lane seal only if `seal_ns` exceeds ~15% of the send thread at 2 Gbps.
#[derive(Debug, Default, Clone, Copy)]
pub struct SealPerf {
/// ns inside `ErasureCoder::encode_into` (parity generation).
pub fec_ns: u64,
/// ns inside `seal_in_place` across all wire packets (AES-128-GCM).
pub seal_ns: u64,
/// ns inside `send_sealed` (socket syscalls), where the session can see it.
pub sock_ns: u64,
/// Frames sealed and wire packets sealed over the accumulation window.
pub frames: u64,
pub packets: u64,
}
/// [`ErasureCoder`] shim accumulating the time spent in `encode_into` (the send-path FEC
/// stage) — only constructed when `PUNKTFUNK_PERF` armed the session's [`SealPerf`]. The
/// counter is atomic purely to satisfy the trait's `Sync` bound; it lives on one thread.
struct TimedCoder<'a> {
inner: &'a dyn ErasureCoder,
ns: &'a std::sync::atomic::AtomicU64,
}
impl ErasureCoder for TimedCoder<'_> {
fn scheme(&self) -> crate::config::FecScheme {
self.inner.scheme()
}
fn encode(
&self,
data: &[&[u8]],
recovery_count: usize,
) -> std::result::Result<Vec<Vec<u8>>, crate::fec::FecError> {
self.inner.encode(data, recovery_count)
}
fn encode_into(
&self,
data: &[&[u8]],
recovery_count: usize,
out: &mut Vec<Vec<u8>>,
) -> std::result::Result<(), crate::fec::FecError> {
let t0 = std::time::Instant::now();
let r = self.inner.encode_into(data, recovery_count, out);
self.ns.fetch_add(
t0.elapsed().as_nanos() as u64,
std::sync::atomic::Ordering::Relaxed,
);
r
}
fn reconstruct(
&self,
data_count: usize,
recovery_count: usize,
received: &mut [Option<Vec<u8>>],
) -> std::result::Result<Vec<Vec<u8>>, crate::fec::FecError> {
self.inner.reconstruct(data_count, recovery_count, received)
}
fn reconstruct_into(
&self,
recovery_count: usize,
data: &mut [&mut [u8]],
have: &[bool],
recovery: &[(usize, &[u8])],
) -> std::result::Result<(), crate::fec::FecError> {
self.inner
.reconstruct_into(recovery_count, data, have, recovery)
}
}
/// Datagrams drained per `recvmmsg` syscall on the client (the reused ring's size). 128 keeps /// Datagrams drained per `recvmmsg` syscall on the client (the reused ring's size). 128 keeps
/// the syscall rate ≤ ~3.4k/s even at the ~430k pkt/s the post-2026-07-14 receive path delivers /// the syscall rate ≤ ~3.4k/s even at the ~430k pkt/s the post-2026-07-14 receive path delivers
@@ -399,6 +276,68 @@ impl Session {
pts_ns: u64, pts_ns: u64,
user_flags: u32, user_flags: u32,
frame_index: Option<u32>, frame_index: Option<u32>,
) -> Result<Vec<Vec<u8>>> {
self.seal_run(true, |p, coder, emit| {
p.packetize_each(data, pts_ns, user_flags, frame_index, coder, emit)
})
}
/// Host: open a **streamed** access unit ([`crate::quic::VIDEO_CAP_STREAMED_AU`] — only
/// toward a client that advertised it; anyone else must get the whole-AU
/// [`seal_frame_at`](Self::seal_frame_at) path). The AU's bytes are fed in with
/// [`seal_streamed_chunk`](Self::seal_streamed_chunk) as the encoder produces them and
/// closed with [`seal_streamed_finish`](Self::seal_streamed_finish); the three calls'
/// returned wire batches are ONE frame, and the nonce order is emission order across the
/// calls — send each batch as it is returned, before sealing the next.
pub fn begin_streamed_frame_at(
&mut self,
pts_ns: u64,
user_flags: u32,
frame_index: u32,
) -> Result<StreamedAu> {
if self.config.role != Role::Host {
return Err(PunktfunkError::InvalidArg(
"seal_frame called on a client session",
));
}
Ok(self
.packetizer
.begin_streamed(pts_ns, user_flags, Some(frame_index)))
}
/// Feed one encoder chunk into a streamed AU, sealing every FEC block it completes under
/// sentinel headers (see [`begin_streamed_frame_at`](Self::begin_streamed_frame_at)). The
/// returned batch is often EMPTY (the chunk is buffered until a block fills) — that's
/// normal, not an error.
pub fn seal_streamed_chunk(
&mut self,
au: &mut StreamedAu,
chunk: &[u8],
) -> Result<Vec<Vec<u8>>> {
self.seal_run(false, |p, coder, emit| {
p.push_streamed(au, chunk, coder, emit)
})
}
/// Close a streamed AU: seal the final block with the real totals (+ `FLAG_EOF`), which
/// retro-validate the frame at the receiver. Counts the frame as submitted.
pub fn seal_streamed_finish(&mut self, au: StreamedAu) -> Result<Vec<Vec<u8>>> {
self.seal_run(true, |p, coder, emit| p.finish_streamed(au, coder, emit))
}
/// The shared packetize → pooled-wire → seal machinery behind [`seal_frame`](Self::seal_frame)
/// and the streamed sealers: `run` drives the packetizer against an emit sink that writes
/// each packet's plaintext at its final wire offset; the seal pass (two-lane for large runs)
/// then encrypts in place. `count_frame` gates the per-frame stats — a streamed AU counts
/// once, at its finish call.
fn seal_run(
&mut self,
count_frame: bool,
run: impl FnOnce(
&mut Packetizer,
&dyn ErasureCoder,
&mut dyn FnMut(&PacketHeader, &[u8]) -> Result<()>,
) -> Result<()>,
) -> Result<Vec<Vec<u8>>> { ) -> Result<Vec<Vec<u8>>> {
if self.config.role != Role::Host { if self.config.role != Role::Host {
return Err(PunktfunkError::InvalidArg( return Err(PunktfunkError::InvalidArg(
@@ -445,10 +384,10 @@ impl Session {
// sealing itself is a separate pass so it can split across lanes. // sealing itself is a separate pass so it can split across lanes.
let seq_base = *next_seq; let seq_base = *next_seq;
let encrypting = crypto.is_some(); let encrypting = crypto.is_some();
let result = packetizer.packetize_each(data, pts_ns, user_flags, frame_index, coder_ref, { let result = {
let wires = &mut wires; let wires = &mut wires;
let used = &mut used; let used = &mut used;
move |hdr, body| { let mut emit = move |hdr: &PacketHeader, body: &[u8]| -> Result<()> {
if *used == wires.len() { if *used == wires.len() {
wires.push(Vec::new()); wires.push(Vec::new());
} }
@@ -467,8 +406,9 @@ impl Session {
wire.extend_from_slice(body); wire.extend_from_slice(body);
} }
Ok(()) Ok(())
} };
}); run(packetizer, coder_ref, &mut emit)
};
result?; result?;
// A smaller frame uses fewer buffers than the pool holds: drop the unused tail, same // A smaller frame uses fewer buffers than the pool holds: drop the unused tail, same
// as the previous `resize_with(packets.len(), ..)` did. (Before the seal phase, so a // as the previous `resize_with(packets.len(), ..)` did. (Before the seal phase, so a
@@ -483,7 +423,10 @@ impl Session {
} }
let mut split_done = false; let mut split_done = false;
if two_lane && used >= TWO_LANE_MIN_PACKETS { if two_lane && used >= TWO_LANE_MIN_PACKETS {
if let Some(lane) = seal_lane.as_ref() { // Take the lane for the frame: a healthy round-trip puts it back; either
// failure arm drops the corpse so the next large frame respawns a fresh one
// instead of retrying a dead channel forever.
if let Some(lane) = seal_lane.take() {
let half = used / 2; let half = used / 2;
let mut tail = std::mem::take(lane_scratch); let mut tail = std::mem::take(lane_scratch);
tail.extend(wires.drain(half..)); tail.extend(wires.drain(half..));
@@ -494,7 +437,8 @@ impl Session {
ns: 0, ns: 0,
result: Ok(()), result: Ok(()),
}; };
if lane.to_worker.send(job).is_ok() { match lane.to_worker.send(job) {
Ok(()) => {
// Seal the front half while the worker runs; collect BOTH results // Seal the front half while the worker runs; collect BOTH results
// before erroring so the lane is always drained and reusable. // before erroring so the lane is always drained and reusable.
let t0 = perf_armed.then(std::time::Instant::now); let t0 = perf_armed.then(std::time::Instant::now);
@@ -502,10 +446,9 @@ impl Session {
if let Some(t0) = t0 { if let Some(t0) = t0 {
seal_ns += t0.elapsed().as_nanos() as u64; seal_ns += t0.elapsed().as_nanos() as u64;
} }
let mut done = lane match lane.from_worker.recv() {
.from_worker Ok(mut done) => {
.recv() *seal_lane = Some(lane);
.map_err(|_| PunktfunkError::Unsupported("seal lane died"))?;
seal_ns += done.ns; seal_ns += done.ns;
wires.append(&mut done.bufs); wires.append(&mut done.bufs);
*lane_scratch = done.bufs; *lane_scratch = done.bufs;
@@ -513,7 +456,23 @@ impl Session {
done.result?; done.result?;
split_done = true; split_done = true;
} }
// A failed send means the worker is gone — fall through to single-lane. Err(_) => {
// The worker died holding the back half — the frame is
// unrecoverable (its packets are gone), but the error now
// SURFACES instead of `Ok` with half an access unit.
front?;
return Err(PunktfunkError::Unsupported("seal lane died"));
}
}
}
Err(std::sync::mpsc::SendError(job)) => {
// The worker is gone but the channel hands the job back: reclaim
// the back half so the single-lane pass below seals the WHOLE
// frame — previously this fall-through sealed and returned only
// the front half, silently, as `Ok`.
wires.extend(job.bufs);
}
}
} }
} }
if !split_done { if !split_done {
@@ -527,10 +486,12 @@ impl Session {
if let Some(p) = self.seal_perf.as_mut() { if let Some(p) = self.seal_perf.as_mut() {
p.fec_ns += fec_ns.load(std::sync::atomic::Ordering::Relaxed); p.fec_ns += fec_ns.load(std::sync::atomic::Ordering::Relaxed);
p.seal_ns += seal_ns; p.seal_ns += seal_ns;
p.frames += 1; p.frames += count_frame as u64;
p.packets += used as u64; p.packets += used as u64;
} }
if count_frame {
StatsCounters::add(&self.stats.frames_submitted, 1); StatsCounters::add(&self.stats.frames_submitted, 1);
}
let bytes: u64 = wires.iter().map(|w| w.len() as u64).sum(); let bytes: u64 = wires.iter().map(|w| w.len() as u64).sum();
StatsCounters::add(&self.stats.packets_sent, wires.len() as u64); StatsCounters::add(&self.stats.packets_sent, wires.len() as u64);
StatsCounters::add(&self.stats.bytes_sent, bytes); StatsCounters::add(&self.stats.bytes_sent, bytes);
@@ -684,7 +645,7 @@ impl Session {
// Nothing new on the wire — hand over an aged-out partial if one is // Nothing new on the wire — hand over an aged-out partial if one is
// waiting (it can only get staler). // waiting (it can only get staler).
if let Some(p) = self.reassembler.take_partial() { if let Some(p) = self.reassembler.take_partial() {
return Ok(p); return Ok(stamp_received(p));
} }
return Err(PunktfunkError::NoFrame); return Err(PunktfunkError::NoFrame);
} }
@@ -748,12 +709,12 @@ impl Session {
} }
if let Some(frame) = pushed { if let Some(frame) = pushed {
StatsCounters::add(&self.stats.frames_completed, 1); StatsCounters::add(&self.stats.frames_completed, 1);
return Ok(frame); return Ok(stamp_received(frame));
} }
// A push that completed nothing may still have aged a partial out — deliver it // A push that completed nothing may still have aged a partial out — deliver it
// ahead of further draining (its successors are already arriving). // ahead of further draining (its successors are already arriving).
if let Some(p) = self.reassembler.take_partial() { if let Some(p) = self.reassembler.take_partial() {
return Ok(p); return Ok(stamp_received(p));
} }
} }
} }
@@ -816,102 +777,6 @@ impl Session {
} }
} }
/// Extract the AEAD-authenticated 8-byte big-endian sequence prefix from a sealed wire datagram.
/// Only called on the encrypted receive path, where a preceding successful open has already
/// established `wire.len() >= 8`.
fn seq_of(wire: &[u8]) -> u64 {
u64::from_be_bytes(wire[..8].try_into().unwrap())
}
/// Depth of the anti-replay window, in sequences. The sender advances its sequence once per
/// datagram, so this must cover the reassembler's 120 ms loss window
/// ([`LOSS_WINDOW_NS`](crate::packet)) at line-rate packet rates — otherwise the replay filter
/// silently re-tightens the "late ≠ lost" fix: a Wi-Fi-retry-delayed shard the reassembler would
/// still use gets dropped here as "older than the window" first (4096 was only ~33 ms at the
/// ~125k pkt/s of a 1 Gbps stream; 32768 topped out around ~2 Gbps — which the client now
/// exceeds: the 2026-07-14 zero-copy + hardware-AES work measured ~4.8 Gbps wire ≈ 430k pkt/s
/// delivered). 131072 covers 120 ms up to ~1.09M pkt/s (≈12 Gbps wire) and is effectively
/// unbounded for the sparse input stream, while still bounding how far back a replay could
/// hide; the bitmap costs 16 KiB per session.
const REPLAY_WINDOW: u64 = 131072;
const REPLAY_WORDS: usize = (REPLAY_WINDOW / 64) as usize;
/// Sliding-window anti-replay filter over the AEAD-authenticated wire sequence. The sender counts
/// its datagrams from 0, and the protocol never legitimately re-sends a sequence (FEC recovery
/// shards get fresh ones), so a sequence seen twice is a replay. The AEAD tag already authenticates
/// the sequence — a forged one can't open — so this only has to reject *duplicates* of validly
/// sealed datagrams (and anything older than the window, which we can no longer prove is fresh).
/// Genuine reordering within the window is accepted. Bitmap-per-sequence, indexed `seq % WINDOW`.
struct ReplayWindow {
/// Highest sequence accepted so far; `seen` stays false until the first datagram.
highest: u64,
seen: bool,
/// One bit per in-window sequence in `(highest - WINDOW, highest]`.
bits: [u64; REPLAY_WORDS],
}
impl ReplayWindow {
fn new() -> ReplayWindow {
ReplayWindow {
highest: 0,
seen: false,
bits: [0; REPLAY_WORDS],
}
}
#[inline]
fn word_bit(seq: u64) -> (usize, u64) {
let idx = (seq % REPLAY_WINDOW) as usize;
(idx / 64, 1u64 << (idx % 64))
}
fn is_set(&self, seq: u64) -> bool {
let (w, b) = Self::word_bit(seq);
self.bits[w] & b != 0
}
fn set(&mut self, seq: u64) {
let (w, b) = Self::word_bit(seq);
self.bits[w] |= b;
}
fn unset(&mut self, seq: u64) {
let (w, b) = Self::word_bit(seq);
self.bits[w] &= !b;
}
/// Record `seq`, returning `true` if it's fresh (accept) or `false` if it's a replay / too old.
fn accept(&mut self, seq: u64) -> bool {
if !self.seen {
self.seen = true;
self.highest = seq;
self.set(seq);
return true;
}
if seq > self.highest {
// Advance the window. Sequences between the old and new high slide in unseen, so clear
// their (possibly stale, from a full window ago) slots — unless we jumped an entire
// window, in which case wipe the bitmap wholesale.
if seq - self.highest >= REPLAY_WINDOW {
self.bits = [0; REPLAY_WORDS];
} else {
let mut s = self.highest + 1;
while s < seq {
self.unset(s);
s += 1;
}
}
self.highest = seq;
self.set(seq);
true
} else if self.highest - seq >= REPLAY_WINDOW || self.is_set(seq) {
// Older than the window (can't prove it isn't a replay) or already seen (a duplicate) —
// either way, drop it.
false
} else {
self.set(seq); // in-window and not yet seen — a genuine reorder
true
}
}
}
#[cfg(test)] #[cfg(test)]
mod wire_equivalence_tests { mod wire_equivalence_tests {
use super::*; use super::*;
@@ -1011,6 +876,42 @@ mod wire_equivalence_tests {
(0..len).map(|i| (i * 31 + 7) as u8).collect() (0..len).map(|i| (i * 31 + 7) as u8).collect()
} }
/// A dead seal lane (worker gone, channels dangling) must degrade to a single-lane seal of
/// the WHOLE frame — the old fall-through sealed and returned only the front half as `Ok` —
/// and the corpse must be dropped so the next large frame respawns a fresh lane.
#[test]
fn dead_seal_lane_falls_back_to_single_lane_whole_frame() {
let mut opt = host_session(host_cfg(FecScheme::Gf16, 20, true));
let mut refr = host_session(host_cfg(FecScheme::Gf16, 20, true));
// A lane whose worker has already exited: both far ends dropped, so `send` fails
// immediately and hands the job (with the frame's back half) back.
let (to_worker, jobs) = std::sync::mpsc::sync_channel::<SealJob>(1);
let (done_tx, from_worker) = std::sync::mpsc::sync_channel::<SealJob>(1);
drop(jobs);
drop(done_tx);
opt.seal_lane = Some(SealLane {
to_worker,
from_worker,
});
let frame = pattern(20000); // > TWO_LANE_MIN_PACKETS wire packets → takes the split path
let got = opt.seal_frame(&frame, 7, 0).unwrap();
let want = seal_via_wrapper(&mut refr, &frame, 7, 0);
assert_eq!(got, want, "fallback must seal the whole frame, not half");
assert!(
opt.seal_lane.is_none(),
"the dead lane must be dropped, not retried forever"
);
// The next large frame respawns a fresh, working lane.
opt.reclaim_wires(got);
let got2 = opt.seal_frame(&frame, 8, 1).unwrap();
let want2 = seal_via_wrapper(&mut refr, &frame, 8, 1);
assert_eq!(got2, want2);
assert!(
opt.seal_lane.is_some(),
"a fresh lane respawns on the next large frame"
);
}
/// Partial delivery (plan §4.4): a chunk-aligned frame that loses shards past FEC's /// Partial delivery (plan §4.4): a chunk-aligned frame that loses shards past FEC's
/// reach is DELIVERED once it ages out — `complete: false`, received shards at their /// reach is DELIVERED once it ages out — `complete: false`, received shards at their
/// exact offsets, missing ranges zero-filled — instead of silently dropping. Plain /// exact offsets, missing ranges zero-filled — instead of silently dropping. Plain
@@ -1106,78 +1007,3 @@ mod wire_equivalence_tests {
); );
} }
} }
#[cfg(test)]
mod replay_tests {
use super::*;
#[test]
fn accepts_in_order_and_rejects_duplicates() {
let mut w = ReplayWindow::new();
for seq in 0..1000 {
assert!(w.accept(seq), "fresh in-order seq {seq} must be accepted");
}
// Every one of those is now a replay.
for seq in 0..1000 {
assert!(!w.accept(seq), "replayed seq {seq} must be rejected");
}
}
#[test]
fn accepts_reorder_within_window_once() {
let mut w = ReplayWindow::new();
assert!(w.accept(100));
// Earlier-but-in-window sequences (a genuine reorder) are accepted exactly once.
assert!(w.accept(80));
assert!(!w.accept(80), "second copy of a reordered seq is a replay");
assert!(w.accept(99));
assert!(
!w.accept(100),
"the high-water seq itself can't be replayed"
);
}
#[test]
fn rejects_older_than_window() {
let mut w = ReplayWindow::new();
assert!(w.accept(REPLAY_WINDOW * 2));
// Anything a full window or more behind the high-water mark is dropped (can't prove fresh).
assert!(!w.accept(REPLAY_WINDOW * 2 - REPLAY_WINDOW));
assert!(!w.accept(0));
// But just inside the window is still accepted.
assert!(w.accept(REPLAY_WINDOW * 2 - (REPLAY_WINDOW - 1)));
}
#[test]
fn large_forward_jump_wipes_stale_bits() {
let mut w = ReplayWindow::new();
assert!(w.accept(5));
// Jump far forward (more than a window). The slot for an old seq that aliases 5 mod WINDOW
// must read as unseen afterward, i.e. the jump cleared it — so a NEW seq there is accepted.
let far = 10 * REPLAY_WINDOW + 5;
assert!(w.accept(far));
assert!(
!w.accept(5),
"the pre-jump seq is now far older than the window"
);
// A fresh seq aliasing 5 (mod WINDOW) but inside the new window is accepted, proving the
// stale bit was cleared rather than mistaken for a replay.
assert!(w.accept(far - REPLAY_WINDOW + 1));
}
#[test]
fn first_seq_need_not_be_zero() {
// Startup loss can mean the first datagram we ever open isn't seq 0.
let mut w = ReplayWindow::new();
assert!(w.accept(42));
assert!(!w.accept(42));
assert!(w.accept(43));
}
#[test]
fn seq_of_reads_the_big_endian_prefix() {
let mut wire = 0x0102_0304_0506_0708u64.to_be_bytes().to_vec();
wire.extend_from_slice(b"ciphertext-and-tag");
assert_eq!(seq_of(&wire), 0x0102_0304_0506_0708);
}
}
+98
View File
@@ -0,0 +1,98 @@
//! `PUNKTFUNK_PERF` stage-timing telemetry for the two hot paths: where the client pump
//! and the host send thread actually spend their time, accumulated per report window and
//! drained via [`Session::take_pump_perf`](super::Session::take_pump_perf) /
//! [`Session::take_seal_perf`](super::Session::take_seal_perf).
use crate::fec::ErasureCoder;
/// Accumulated client receive-path stage timings since the last [`Session::take_pump_perf`](super::Session::take_pump_perf).
/// Answers "where does the pump core go" at line rate: kernel drain (`recv_ns`) vs AES-GCM
/// (`decrypt_ns`) vs reassembly+FEC (`reasm_ns`, the `Reassembler::push` round-trip including
/// shard copies and block reconstruction). 2026-07-14 sweep context: the pump pegs one core at
/// ~1.5 Gbps wire, ~85% of it userspace — this split is what Phase 2.1 (pooled reassembly) is
/// validated against.
#[derive(Debug, Default, Clone, Copy)]
pub struct PumpPerf {
/// ns inside `recv_batch` (recvmmsg / recvmsg_x), i.e. syscall + kernel copy.
pub recv_ns: u64,
/// ns inside `open_in_place` across all datagrams (AES-128-GCM + replay-window upkeep).
pub decrypt_ns: u64,
/// ns inside `Reassembler::push` (header parse, shard copy, FEC reconstruct, AU assembly).
pub reasm_ns: u64,
/// recv_batch calls (batches) and datagrams processed over the accumulation window.
pub batches: u64,
pub packets: u64,
}
/// Accumulated host send-path stage timings since the last [`Session::take_seal_perf`](super::Session::take_seal_perf) (plan
/// Phase 0.4, host half). Answers "where does the send thread go" at rate: FEC parity
/// generation (`fec_ns`, inside [`ErasureCoder::encode_into`]) vs AES-GCM (`seal_ns`,
/// per-packet `seal_in_place`) vs the socket handoff (`sock_ns` — `send_gso`/`sendmmsg`
/// syscalls; the internal submit paths time it here, the paced video path folds its chunk
/// sends in via [`Session::note_sock_ns`](super::Session::note_sock_ns)). The Phase 1.5 gate reads off this split: build
/// two-lane seal only if `seal_ns` exceeds ~15% of the send thread at 2 Gbps.
#[derive(Debug, Default, Clone, Copy)]
pub struct SealPerf {
/// ns inside `ErasureCoder::encode_into` (parity generation).
pub fec_ns: u64,
/// ns inside `seal_in_place` across all wire packets (AES-128-GCM).
pub seal_ns: u64,
/// ns inside `send_sealed` (socket syscalls), where the session can see it.
pub sock_ns: u64,
/// Frames sealed and wire packets sealed over the accumulation window.
pub frames: u64,
pub packets: u64,
}
/// [`ErasureCoder`] shim accumulating the time spent in `encode_into` (the send-path FEC
/// stage) — only constructed when `PUNKTFUNK_PERF` armed the session's [`SealPerf`]. The
/// counter is atomic purely to satisfy the trait's `Sync` bound; it lives on one thread.
pub(super) struct TimedCoder<'a> {
pub(super) inner: &'a dyn ErasureCoder,
pub(super) ns: &'a std::sync::atomic::AtomicU64,
}
impl ErasureCoder for TimedCoder<'_> {
fn scheme(&self) -> crate::config::FecScheme {
self.inner.scheme()
}
fn encode(
&self,
data: &[&[u8]],
recovery_count: usize,
) -> std::result::Result<Vec<Vec<u8>>, crate::fec::FecError> {
self.inner.encode(data, recovery_count)
}
fn encode_into(
&self,
data: &[&[u8]],
recovery_count: usize,
out: &mut Vec<Vec<u8>>,
) -> std::result::Result<(), crate::fec::FecError> {
let t0 = std::time::Instant::now();
let r = self.inner.encode_into(data, recovery_count, out);
self.ns.fetch_add(
t0.elapsed().as_nanos() as u64,
std::sync::atomic::Ordering::Relaxed,
);
r
}
fn reconstruct(
&self,
data_count: usize,
recovery_count: usize,
received: &mut [Option<Vec<u8>>],
) -> std::result::Result<Vec<Vec<u8>>, crate::fec::FecError> {
self.inner.reconstruct(data_count, recovery_count, received)
}
fn reconstruct_into(
&self,
recovery_count: usize,
data: &mut [&mut [u8]],
have: &[bool],
recovery: &[(usize, &[u8])],
) -> std::result::Result<(), crate::fec::FecError> {
self.inner
.reconstruct_into(recovery_count, data, have, recovery)
}
}
+175
View File
@@ -0,0 +1,175 @@
//! Sliding-window anti-replay filter over the AEAD-authenticated wire sequence
//! (plan §1). Applied on both encrypted receive paths —
//! [`Session::poll_frame`](super::Session::poll_frame) and
//! [`Session::poll_input`](super::Session::poll_input).
/// Extract the AEAD-authenticated 8-byte big-endian sequence prefix from a sealed wire datagram.
/// Only called on the encrypted receive path, where a preceding successful open has already
/// established `wire.len() >= 8`.
pub(super) fn seq_of(wire: &[u8]) -> u64 {
u64::from_be_bytes(wire[..8].try_into().unwrap())
}
/// Depth of the anti-replay window, in sequences. The sender advances its sequence once per
/// datagram, so this must cover the reassembler's 120 ms loss window
/// ([`LOSS_WINDOW_NS`](crate::packet)) at line-rate packet rates — otherwise the replay filter
/// silently re-tightens the "late ≠ lost" fix: a Wi-Fi-retry-delayed shard the reassembler would
/// still use gets dropped here as "older than the window" first (4096 was only ~33 ms at the
/// ~125k pkt/s of a 1 Gbps stream; 32768 topped out around ~2 Gbps — which the client now
/// exceeds: the 2026-07-14 zero-copy + hardware-AES work measured ~4.8 Gbps wire ≈ 430k pkt/s
/// delivered). 131072 covers 120 ms up to ~1.09M pkt/s (≈12 Gbps wire) and is effectively
/// unbounded for the sparse input stream, while still bounding how far back a replay could
/// hide; the bitmap costs 16 KiB per session.
const REPLAY_WINDOW: u64 = 131072;
const REPLAY_WORDS: usize = (REPLAY_WINDOW / 64) as usize;
/// Sliding-window anti-replay filter over the AEAD-authenticated wire sequence. The sender counts
/// its datagrams from 0, and the protocol never legitimately re-sends a sequence (FEC recovery
/// shards get fresh ones), so a sequence seen twice is a replay. The AEAD tag already authenticates
/// the sequence — a forged one can't open — so this only has to reject *duplicates* of validly
/// sealed datagrams (and anything older than the window, which we can no longer prove is fresh).
/// Genuine reordering within the window is accepted. Bitmap-per-sequence, indexed `seq % WINDOW`.
pub(super) struct ReplayWindow {
/// Highest sequence accepted so far; `seen` stays false until the first datagram.
highest: u64,
seen: bool,
/// One bit per in-window sequence in `(highest - WINDOW, highest]`.
bits: [u64; REPLAY_WORDS],
}
impl ReplayWindow {
pub(super) fn new() -> ReplayWindow {
ReplayWindow {
highest: 0,
seen: false,
bits: [0; REPLAY_WORDS],
}
}
#[inline]
fn word_bit(seq: u64) -> (usize, u64) {
let idx = (seq % REPLAY_WINDOW) as usize;
(idx / 64, 1u64 << (idx % 64))
}
fn is_set(&self, seq: u64) -> bool {
let (w, b) = Self::word_bit(seq);
self.bits[w] & b != 0
}
fn set(&mut self, seq: u64) {
let (w, b) = Self::word_bit(seq);
self.bits[w] |= b;
}
fn unset(&mut self, seq: u64) {
let (w, b) = Self::word_bit(seq);
self.bits[w] &= !b;
}
/// Record `seq`, returning `true` if it's fresh (accept) or `false` if it's a replay / too old.
pub(super) fn accept(&mut self, seq: u64) -> bool {
if !self.seen {
self.seen = true;
self.highest = seq;
self.set(seq);
return true;
}
if seq > self.highest {
// Advance the window. Sequences between the old and new high slide in unseen, so clear
// their (possibly stale, from a full window ago) slots — unless we jumped an entire
// window, in which case wipe the bitmap wholesale.
if seq - self.highest >= REPLAY_WINDOW {
self.bits = [0; REPLAY_WORDS];
} else {
let mut s = self.highest + 1;
while s < seq {
self.unset(s);
s += 1;
}
}
self.highest = seq;
self.set(seq);
true
} else if self.highest - seq >= REPLAY_WINDOW || self.is_set(seq) {
// Older than the window (can't prove it isn't a replay) or already seen (a duplicate) —
// either way, drop it.
false
} else {
self.set(seq); // in-window and not yet seen — a genuine reorder
true
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn accepts_in_order_and_rejects_duplicates() {
let mut w = ReplayWindow::new();
for seq in 0..1000 {
assert!(w.accept(seq), "fresh in-order seq {seq} must be accepted");
}
// Every one of those is now a replay.
for seq in 0..1000 {
assert!(!w.accept(seq), "replayed seq {seq} must be rejected");
}
}
#[test]
fn accepts_reorder_within_window_once() {
let mut w = ReplayWindow::new();
assert!(w.accept(100));
// Earlier-but-in-window sequences (a genuine reorder) are accepted exactly once.
assert!(w.accept(80));
assert!(!w.accept(80), "second copy of a reordered seq is a replay");
assert!(w.accept(99));
assert!(
!w.accept(100),
"the high-water seq itself can't be replayed"
);
}
#[test]
fn rejects_older_than_window() {
let mut w = ReplayWindow::new();
assert!(w.accept(REPLAY_WINDOW * 2));
// Anything a full window or more behind the high-water mark is dropped (can't prove fresh).
assert!(!w.accept(REPLAY_WINDOW * 2 - REPLAY_WINDOW));
assert!(!w.accept(0));
// But just inside the window is still accepted.
assert!(w.accept(REPLAY_WINDOW * 2 - (REPLAY_WINDOW - 1)));
}
#[test]
fn large_forward_jump_wipes_stale_bits() {
let mut w = ReplayWindow::new();
assert!(w.accept(5));
// Jump far forward (more than a window). The slot for an old seq that aliases 5 mod WINDOW
// must read as unseen afterward, i.e. the jump cleared it — so a NEW seq there is accepted.
let far = 10 * REPLAY_WINDOW + 5;
assert!(w.accept(far));
assert!(
!w.accept(5),
"the pre-jump seq is now far older than the window"
);
// A fresh seq aliasing 5 (mod WINDOW) but inside the new window is accepted, proving the
// stale bit was cleared rather than mistaken for a replay.
assert!(w.accept(far - REPLAY_WINDOW + 1));
}
#[test]
fn first_seq_need_not_be_zero() {
// Startup loss can mean the first datagram we ever open isn't seq 0.
let mut w = ReplayWindow::new();
assert!(w.accept(42));
assert!(!w.accept(42));
assert!(w.accept(43));
}
#[test]
fn seq_of_reads_the_big_endian_prefix() {
let mut wire = 0x0102_0304_0506_0708u64.to_be_bytes().to_vec();
wire.extend_from_slice(b"ciphertext-and-tag");
assert_eq!(seq_of(&wire), 0x0102_0304_0506_0708);
}
}
+74
View File
@@ -0,0 +1,74 @@
//! The second seal lane (plan Phase 1.5): a persistent worker thread that AES-GCM-seals
//! the back half of a large frame's wire packets while the send thread seals the front.
//! [`Session`](super::Session)`::seal_frame_inner` owns the split policy; this module owns
//! the lane machinery and the shared per-slice seal loop.
use crate::crypto::SessionCrypto;
use crate::error::Result;
/// Wire-packet count at which a frame's sealing splits across two lanes (plan Phase 1.5):
/// below it the channel rendezvous (~µs) isn't worth it; at it the halved AES-GCM span
/// (≥ ~125 µs of ~1 µs/packet work) dwarfs the hand-off. ≈300 KB of wire, i.e. ≥150 Mbps
/// at 60 fps — small frames and the probe's ~17-packet AUs stay strictly single-lane.
pub(super) const TWO_LANE_MIN_PACKETS: usize = 256;
/// One two-lane seal hand-off: the frame's back-half wire buffers, sealed by the worker with
/// nonces `seq_base + i` (the nonce order is deterministic per shard index, which is what
/// makes the split sound). Round-trips through the channels so the buffers return to the pool.
pub(super) struct SealJob {
pub(super) bufs: Vec<Vec<u8>>,
pub(super) seq_base: u64,
pub(super) timed: bool,
/// Worker-lane CPU ns (when `timed`) and the seal outcome, filled in by the worker.
pub(super) ns: u64,
pub(super) result: Result<()>,
}
/// The persistent second seal lane: a worker thread that AES-GCM-seals the back half of a
/// large frame's packets while the send thread seals the front half. Rendezvous channels
/// (bound 1) — the send thread submits, seals its half, then waits; no per-frame spawn.
/// Dropping the struct closes the channel and the worker exits.
pub(super) struct SealLane {
pub(super) to_worker: std::sync::mpsc::SyncSender<SealJob>,
pub(super) from_worker: std::sync::mpsc::Receiver<SealJob>,
}
impl SealLane {
pub(super) fn spawn(crypto: std::sync::Arc<SessionCrypto>) -> Option<SealLane> {
let (to_worker, jobs) = std::sync::mpsc::sync_channel::<SealJob>(1);
let (done_tx, from_worker) = std::sync::mpsc::sync_channel::<SealJob>(1);
std::thread::Builder::new()
.name("punktfunk-seal2".into())
.spawn(move || {
while let Ok(mut job) = jobs.recv() {
let t0 = job.timed.then(std::time::Instant::now);
job.result = seal_wire_slice(&crypto, &mut job.bufs, job.seq_base);
if let Some(t0) = t0 {
job.ns = t0.elapsed().as_nanos() as u64;
}
if done_tx.send(job).is_err() {
break; // session gone mid-frame — nothing left to seal for
}
}
})
.ok()?;
Some(SealLane {
to_worker,
from_worker,
})
}
}
/// Seal a run of pre-written wire buffers in place: buffer `i` is `seq(8) ‖ plaintext ‖ tag
/// scratch` and seals over `[8..]` with sequence `seq_base + i` — the exact per-packet layout
/// and nonce order of the fused single-lane path. Shared by both lanes.
pub(super) fn seal_wire_slice(
c: &SessionCrypto,
wires: &mut [Vec<u8>],
seq_base: u64,
) -> Result<()> {
for (i, wire) in wires.iter_mut().enumerate() {
c.seal_in_place(seq_base.wrapping_add(i as u64), &mut wire[8..])?;
}
Ok(())
}
@@ -198,8 +198,14 @@ pub(super) fn send_gso(t: &UdpTransport, packets: &[&[u8]]) -> std::io::Result<u
return send_batch(t, packets); return send_batch(t, packets);
} }
let fd = t.socket.as_raw_fd(); let fd = t.socket.as_raw_fd();
// A GSO super-buffer is capped at 64 segments AND 65535 payload bytes (kernel limits). // A GSO super-buffer is capped at 64 segments AND, in bytes, by the kernel's UDP payload
let max_seg = (65535 / seg).clamp(1, 64); // ceiling. 65535 is the IP-datagram cap — the payload is that minus the IP + UDP headers
// the cork accounts for (IPv4: 65507, IPv6: 65487). Use the tighter v6 figure: it costs at
// most one segment per train, while a super-buffer over the ceiling is bounced with
// EMSGSIZE — which gso_unsupported() reads as "no GSO on this path" and latches GSO off
// process-wide, silently forfeiting the multi-Gbps lever over a local arithmetic slip.
const GSO_MAX_PAYLOAD: usize = 65535 - 40 - 8;
let max_seg = (GSO_MAX_PAYLOAD / seg).clamp(1, 64);
let mut scratch: Vec<u8> = Vec::with_capacity(seg * max_seg); let mut scratch: Vec<u8> = Vec::with_capacity(seg * max_seg);
let mut sent = 0usize; let mut sent = 0usize;
for chunk in packets.chunks(max_seg) { for chunk in packets.chunks(max_seg) {
+131 -8
View File
@@ -110,9 +110,19 @@ pub fn spawn_data_punch(sock: UdpSocket, stop: std::sync::Arc<std::sync::atomic:
.spawn(move || { .spawn(move || {
let mut i = 0u32; let mut i = 0u32;
while !stop.load(std::sync::atomic::Ordering::Relaxed) { while !stop.load(std::sync::atomic::Ordering::Relaxed) {
if sock.send(PUNCH_MAGIC).is_err() { match sock.send(PUNCH_MAGIC) {
Ok(_) => {}
// Same contract as `Transport::send`: a momentarily full tx queue, a stale
// ICMP or a network-path blip is a lossy drop, not a reason to stop holding
// the NAT/firewall path open. Breaking here is silent and permanent — the
// path recovers, video keeps flowing, and the stream dies later when the
// idle timer expires the mapping during a static scene.
Err(e) if is_transient_io(&e) => {}
Err(e) => {
tracing::debug!(error = %e, "data-plane punch send failed — stopping keepalive");
break; break;
} }
}
let delay_ms = if i < 15 { 200 } else { 2000 }; let delay_ms = if i < 15 { 200 } else { 2000 };
i = i.saturating_add(1); i = i.saturating_add(1);
std::thread::sleep(std::time::Duration::from_millis(delay_ms)); std::thread::sleep(std::time::Duration::from_millis(delay_ms));
@@ -160,34 +170,65 @@ impl UdpTransport {
/// NAT-translated one, which can differ from the client-reported `fallback_peer`). If no punch /// NAT-translated one, which can differ from the client-reported `fallback_peer`). If no punch
/// arrives (a client that doesn't hole-punch), fall back to `fallback_peer` — the same flat-LAN /// arrives (a client that doesn't hole-punch), fall back to `fallback_peer` — the same flat-LAN
/// behaviour as [`connect`](Self::connect). Returns `(transport, punched)`. /// behaviour as [`connect`](Self::connect). Returns `(transport, punched)`.
///
/// `expect_ip` is the *authenticated* peer address (the QUIC connection's remote IP) — see
/// [`from_socket_punch`](Self::from_socket_punch) for why only punches from it are honoured.
pub fn connect_via_punch( pub fn connect_via_punch(
local: &str, local: &str,
fallback_peer: &str, fallback_peer: &str,
expect_ip: std::net::IpAddr,
punch_timeout: std::time::Duration, punch_timeout: std::time::Duration,
) -> std::io::Result<(Self, bool)> { ) -> std::io::Result<(Self, bool)> {
Self::from_socket_punch(UdpSocket::bind(local)?, fallback_peer, punch_timeout) Self::from_socket_punch(
UdpSocket::bind(local)?,
fallback_peer,
expect_ip,
punch_timeout,
)
} }
/// [`connect_via_punch`](Self::connect_via_punch) on an already-bound socket — see /// [`connect_via_punch`](Self::connect_via_punch) on an already-bound socket — see
/// [`from_socket`](Self::from_socket) for why the host binds the data port up front. /// [`from_socket`](Self::from_socket) for why the host binds the data port up front.
///
/// `expect_ip` binds the data plane to the peer the control plane already authenticated.
/// [`PUNCH_MAGIC`] is a fixed public constant carrying no key, nonce or session id, so without
/// this check *any* source that lands an 8-byte datagram on the (ephemeral, sprayable) data
/// port during the punch wait becomes the video destination — the legitimate client is then
/// filtered out by the `connect` below and receives nothing, while QUIC stays healthy so no
/// reconnect is triggered. Only the *port* is in question here (that is what a NAT remaps, and
/// what the punch exists to discover); the IP is known, because the client binds `0.0.0.0:0`
/// and dials the same host IP as its QUIC connection, so the kernel picks the same source IP
/// for both planes and any NAT on the path presents one source IP for both.
pub fn from_socket_punch( pub fn from_socket_punch(
socket: UdpSocket, socket: UdpSocket,
fallback_peer: &str, fallback_peer: &str,
expect_ip: std::net::IpAddr,
punch_timeout: std::time::Duration, punch_timeout: std::time::Duration,
) -> std::io::Result<(Self, bool)> { ) -> std::io::Result<(Self, bool)> {
socket.set_read_timeout(Some(punch_timeout))?;
let deadline = std::time::Instant::now() + punch_timeout; let deadline = std::time::Instant::now() + punch_timeout;
let mut buf = [0u8; 64]; let mut buf = [0u8; 64];
let mut observed: Option<std::net::SocketAddr> = None; let mut observed: Option<std::net::SocketAddr> = None;
loop { loop {
// Budget the read from what's LEFT, not the full window: off-peer datagrams are
// discarded below, and a full-window timeout per read would let a stray flood stretch
// the punch wait far past `punch_timeout`.
let remaining = deadline.saturating_duration_since(std::time::Instant::now());
if remaining.is_zero() {
break;
}
socket.set_read_timeout(Some(remaining))?;
match socket.recv_from(&mut buf) { match socket.recv_from(&mut buf) {
Ok((n, src)) Ok((n, src))
if n >= PUNCH_MAGIC.len() && &buf[..PUNCH_MAGIC.len()] == PUNCH_MAGIC => if src.ip() == expect_ip
&& n >= PUNCH_MAGIC.len()
&& &buf[..PUNCH_MAGIC.len()] == PUNCH_MAGIC =>
{ {
observed = Some(src); observed = Some(src);
break; break;
} }
Ok(_) => {} // stray datagram — keep waiting for a real punch // Stray, or a well-formed punch from someone who isn't the authenticated peer —
// keep waiting for a real one.
Ok(_) => {}
Err(e) Err(e)
if matches!( if matches!(
e.kind(), e.kind(),
@@ -198,9 +239,6 @@ impl UdpTransport {
} }
Err(e) => return Err(e), Err(e) => return Err(e),
} }
if std::time::Instant::now() >= deadline {
break;
}
} }
let punched = observed.is_some(); let punched = observed.is_some();
let target = observed.map(|s| s.to_string()); let target = observed.map(|s| s.to_string());
@@ -482,4 +520,89 @@ mod tests {
"every datagram should be drained via recv_batch" "every datagram should be drained via recv_batch"
); );
} }
/// The punch discovers the peer's NAT-remapped *port*, so a punch from the authenticated IP on
/// a port that differs from the client-reported one must still be adopted — that is the whole
/// reason hole-punching exists, and the source-IP check must not break it.
#[test]
fn punch_adopts_remapped_port_from_the_authenticated_peer() {
// Stands in for the client's post-NAT data socket: same IP as the "QUIC peer", new port.
let puncher = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
puncher
.set_read_timeout(Some(std::time::Duration::from_millis(500)))
.unwrap();
// The client-*reported* address, which the NAT remapped — video must NOT go here.
let reported = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
reported
.set_read_timeout(Some(std::time::Duration::from_millis(200)))
.unwrap();
let host_sock = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
let host_addr = host_sock.local_addr().unwrap();
puncher.send_to(PUNCH_MAGIC, host_addr).unwrap();
let (transport, punched) = UdpTransport::from_socket_punch(
host_sock,
&reported.local_addr().unwrap().to_string(),
std::net::IpAddr::from([127, 0, 0, 1]),
std::time::Duration::from_millis(500),
)
.unwrap();
assert!(punched, "a punch from the authenticated IP must be adopted");
transport.send(b"video").unwrap();
let mut buf = [0u8; 64];
let n = puncher
.recv(&mut buf)
.expect("video must follow the punched (NAT-remapped) port");
assert_eq!(&buf[..n], b"video");
assert!(
reported.recv(&mut buf).is_err(),
"video must not go to the stale reported port"
);
}
/// A punch from any source other than the QUIC-authenticated peer must be ignored: `PUNCH_MAGIC`
/// is a fixed public constant with no key or session id, so honouring an off-peer punch lets
/// anyone who lands an 8-byte datagram on the ephemeral data port steal (or redirect) the video
/// plane while the control plane stays healthy. Falling back to the reported address is correct.
#[test]
fn punch_from_an_unauthenticated_source_is_ignored() {
let attacker = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
attacker
.set_read_timeout(Some(std::time::Duration::from_millis(200)))
.unwrap();
let legit = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
legit
.set_read_timeout(Some(std::time::Duration::from_millis(500)))
.unwrap();
let host_sock = std::net::UdpSocket::bind("127.0.0.1:0").unwrap();
let host_addr = host_sock.local_addr().unwrap();
attacker.send_to(PUNCH_MAGIC, host_addr).unwrap();
// The authenticated peer is TEST-NET-1, so nothing arriving over loopback is the peer.
let (transport, punched) = UdpTransport::from_socket_punch(
host_sock,
&legit.local_addr().unwrap().to_string(),
std::net::IpAddr::from([192, 0, 2, 1]),
std::time::Duration::from_millis(300),
)
.unwrap();
assert!(
!punched,
"an off-peer punch must not be adopted as the video destination"
);
transport.send(b"video").unwrap();
let mut buf = [0u8; 64];
assert!(
attacker.recv(&mut buf).is_err(),
"video must never be redirected to the punch source"
);
let n = legit
.recv(&mut buf)
.expect("video falls back to the reported peer address");
assert_eq!(&buf[..n], b"video");
}
} }

Some files were not shown because too many files have changed in this diff Show More