Compare commits

...
Author SHA1 Message Date
enricobuehler c5eea458b4 fix(windows/cursor): gate the between-session recycle off — calling it from ADD deadlocks
The ADD path holds the manager 'device' mutex, and invalidate_cached_device
takes it. Its own doc says it must not be called from inside that mutex. Calling
it there deadlocked the session: on .173 the ADD stopped after
SET_RENDER_ADAPTER, no monitor was created, and the client reported 'no frames
received'.

The mechanism itself is sound and measured — recycling the driver's WUDFHost
clears the declare (pid 3872 -> 19932, adapter_luid 0x8ed607 -> 0x1a8f6ca,
cursor_excluded true -> false, next session streamed normally). What is missing
is a call site that runs OUTSIDE that mutex and is still on every session's
path; the handshake site tried earlier is not reached on this host.

Off unless PUNKTFUNK_CURSOR_RECYCLE=1. A session that self-composites is the
old behaviour; a deadlocked one is a regression, and the tree should not carry
that while the call site is unsolved.
2026-08-09 00:40:54 +02:00
enricobuehler 7f141eb9a4 fix(windows/cursor): record the declare where it actually happens, and retire the stale control handle
Two fixes that together make the reconnect case work.

1. The declare point is send_cursor_channel, not the ADD request's hw_cursor
   flag. The driver declares its IddCx hardware cursor when the cursor CHANNEL
   arrives -- that is what emits 'cursor channel delivered - driver declares the
   hardware cursor'. Recording it at ADD time (and before that in
   capture_virtual_output) left CURSOR_DECLARED false, so the between-session
   clean returned at its first gate and logged nothing, three runs in a row.

2. The cached control handle dies with the recycled WUDFHost. The first run that
   actually reached the recycle produced a session with NO FRAMES, because the
   ADD that followed used a stale handle. Retire it via
   manager::invalidate_cached_device so the next ensure_device reopens against
   the respawned host, and give WUDF a moment to come back.
2026-08-09 00:32:39 +02:00
enricobuehler 06086de328 fix(windows/cursor): do the clean on the ADD path — the handshake call site was never reached
Third and last placement bug in this chain. clean_cursor_for_next_session was
called from the handshake, inside the '#[cfg(windows)] let prep = match (source,
compositor) { (Virtual, Some(comp)) => ...' arm — which is not the path taken on
this host, so the function never ran. The logs said so by omission: no 'skipping'
line, no 'recycled' line, and cursor_excluded stayed true across the reconnect.

Move it next to where the declare is recorded: the ADD request. Every session
that creates a monitor goes through it, with hw_cursor in hand, before the
monitor exists — which is also the moment the recycle is safe.
2026-08-09 00:24:15 +02:00
enricobuehler 7b91afd721 fix(windows/cursor): record the declare from the ADD request — the only place every one passes
CURSOR_DECLARED was set in capture_virtual_output, which is not on the Windows
prep path, so it was never true and the between-session clean returned at its
first gate every time, silently. Both on-glass runs that appeared to 'not fire'
were this; the one run that did fire was a build with no gate at all.

Move it to the ADD request in the driver module, which every session that wants
a hardware cursor goes through by construction.
2026-08-09 00:16:44 +02:00
enricobuehler 0ed4b51104 feat(windows/cursor): un-declare by recycling the driver's WUDFHost — the reconnect case now works
The declare is irrevocable and adapter-wide, but its SCOPE is the WUDFHost
process (monitor.rs DECLARED_TARGETS). So it can be dropped by recycling that
process instead of restarting the device — and unlike /restart-device, which is
a once-per-boot operation whose single use the start-up clean already spends,
this can run whenever no session holds a display.

Proven on .173 2026-08-08 by killing the host by hand: pid 3872 -> 19932,
adapter_luid 0x8ed607 -> 0x1a8f6ca (a fresh adapter object), cursor_excluded
true -> FALSE, and the next capture session streamed normally. Not a zombie:
WUDF respawns the host on the next open.

The pid comes from the driver itself — every ADD reply carries wudf_pid, so
there is no guessing which of the machine's WUDFHost processes is ours.

This is what makes the requirement reachable: run a desktop session, disconnect,
reconnect in capture mode, and the pointer goes back to the OS's own
compositing. restart_device_for_clean_cursor stays for the start-up path, where
a full device restart is the stronger reset and the once-per-boot budget is
untouched.
2026-08-09 00:07:50 +02:00
enricobuehler 3197a4e887 fix(windows/cursor): name the once-per-boot limit — /restart-device cannot clean between sessions
Measured on a freshly cold-booted .173: the start-up clean succeeded (0.07 s,
adapter clean, first session came up cursor_excluded=false and then declared),
and the between-session clean then fired correctly — gate passed,
CURSOR_DECLARED set, no session streaming — and the restart FAILED with 'Das
System muss neu gestartet werden, damit Konfigurationsvorgaenge abgeschlossen
werden'.

So /restart-device is a ONCE-PER-BOOT lever. The first call after a cold boot
works; every later one in the same boot needs a reboot first, and repeated
attempts additionally drive the devnode into 'restart pending'. The earlier
'0.07 s, so just do it whenever' measurement was true only of the first call.

That settles what this mechanism can and cannot do: it cleans the adapter at
host start-up, and nowhere else. Giving a capture session back the lossless
pointer AFTER a desktop session in the same boot needs a different way to
recycle the driver's WUDFHost process, which is where the declare lives.

Also fixes the detector, which missed this wording entirely and logged
restart_pending=false on exactly the failure it exists to name.
2026-08-08 21:10:06 +02:00
enricobuehler df74dd5aee fix(windows/cursor): a device restart is NOT free — gate it again, and stop swallowing pnputil's reason
Two corrections, both from the box.

1. The previous commit made the between-session clean unconditional on the
   grounds that /restart-device is idempotent and costs 0.07 s. It is not free:
   after ~6 restarts in one afternoon .173 put the devnode into RESTART PENDING,
   and pnputil then refuses every further attempt -- 'a system restart is
   pending for this device to complete a previous operation' -- until an actual
   reboot. Restarting speculatively on every capture-mode connect would burn the
   only lever we have. So gate it on an outstanding declare again.

2. That failure was invisible. The script discarded pnputil's output and
   reported a bare exit code, so three separate runs looked like a wiring bug
   (the clean 'not firing') when in fact it ran every time and the restart
   failed. Keep the text, and surface a restart_pending field so the one
   failure that no retry can fix is named in the log.

Also corrects the previous commit's message: CURSOR_DECLARED was never the
problem. Both host processes run as SYSTEM and serve/handshake/capture share
one process, so the flag was set and read correctly all along.
2026-08-08 19:02:49 +02:00
enricobuehler b2020396c9 fix(windows/cursor): stop gating the between-session clean on host-process state
The reconnect case still failed after the keep-alive guard fix, and the logs
said why by omission: neither outcome line appeared, so
clean_cursor_for_next_session was returning at its first early-out —
CURSOR_DECLARED was false even though the previous session had logged
'driver declares the hardware cursor'. The flag is written in
capture_virtual_output and read in the handshake; one of those is not on the
path that actually runs.

Rather than chase that, delete the dependency on it. The operation being
guarded is idempotent and costs 0.07 s: restarting an already-clean adapter
wastes 70 ms once per capture-mode connect, while failing to restart a dirty
one costs that session a full-frame copy for every frame with a visible
pointer, for its entire life. The flag stays only as a log field, so the next
run still shows whether the hint was set.

The real guard remains the one that matters: no session is streaming.
2026-08-08 18:52:19 +02:00
enricobuehler b20184462d fix(windows/cursor): gate the between-session clean on STREAMING sessions, not keep-alive
Measured on .173: the guard never fired in the one case it exists for. After a
desktop-mode session disconnects its monitor LINGERS, so `no_live_displays()`
— which counted keep-alive slots as held — refused the restart on exactly the
reconnect that needed it, and the capture session went back to
`cursor_excluded=true` and forced compositing.

Only a streaming session (SlotState::Active) can be damaged by the restart. A
lingering/pinned monitor has no session attached and a reconnect preempts and
recreates it anyway ("a reused IddCx swap-chain is dead"), so the restart
destroys nothing that was going to survive.
2026-08-08 18:36:41 +02:00
enricobuehler 537a1852ed feat(windows/cursor): give the pointer back after a desktop session — clear the declare between sessions
The start-up clean (5819cf05) cannot reach the case that actually bites: run a
desktop-mode session, disconnect, reconnect with mouse capture. Same host
process, so the adapter is still carrying the first session's hardware-cursor
declare — and because that declare is irrevocable and adapter-wide, the capture
session self-composites the pointer for its entire life. It pays a full-frame
copy for every frame with a visible pointer and draws our straight-alpha
approximation of an XOR cursor, when the OS would do it natively, for free, and
exactly right.

So clear it at the start of the session that does not want it. The host records
a declare when it delivers a cursor channel (`note_cursor_declared`), and the
next session whose Welcome carries no HOST_CAP_CURSOR restarts the device to
drop it. `pnputil /restart-device` recycles the WUDFHost process the driver's
DECLARED_TARGETS lives in — 0.07 s, measured on .173.

Placement is load-bearing: it runs in the handshake right after the Welcome is
sent and BEFORE the display prep is kicked. That is the one moment this
session's display does not exist yet, and the restart takes every monitor on
the adapter with it.

The guard is deliberately conservative. `manager::no_live_displays()` counts
KEEP-ALIVE slots as held, not just streaming ones: a kept monitor belongs to a
client that is expected back, and the cost of being wrong is that someone
else's session dies, against a skip costing only the old self-composited
behaviour. That does mean the clean will not fire while the previous session's
monitor is still kept — telling a kept slot apart from a live one needs
SlotState internals, and is the obvious follow-up.

Verified: `scripts/xcheck.sh windows` green; `cargo fmt --all` clean; full
`cargo build -p punktfunk-host --release --features nvenc` on .173. The
desktop -> disconnect -> capture sequence is NOT yet exercised on glass.
2026-08-08 18:29:18 +02:00
enricobuehler 5819cf054b feat(windows/cursor): clear a sticky hardware-cursor declare at start-up — capture sessions get the OS's own pointer back
The goal this serves: a session that never engages desktop mode should get the
LOSSLESS cursor — composited by Windows itself, for free, with true XOR — and
the host should only pay for compositing when the user actually asks for the
desktop mouse model.

That is already what happens on a clean adapter. The problem is that adapters
do not stay clean. A hardware-cursor declare is irrevocable and ADAPTER-WIDE
(pf-driver-proto v6): once any desktop-mode session declares, DWM stops
compositing the pointer into every later frame on that adapter, so every
capture-latched session afterwards has to self-composite — a full-frame copy
per visible-pointer frame, and our straight-alpha approximation of an XOR
cursor instead of the real thing — for the rest of the adapter's life.

"Until the next reboot" turned out to be far longer than it sounds. With Fast
Startup on (the Windows default) a shutdown plus power-on is a HIBERBOOT: it
restores session 0 and its drivers, so the declare survives what the operator
calls a reboot. Measured on .173 — Kernel-Boot event id 27 reporting `0x1`
where a cold boot reports `0x0`, with lsass/services/wininit keeping their
pre-"reboot" start times, while `LastBootUpTime` reports the older cold boot
and makes uptime checks lie. On such a box the lossless path can be gone for
weeks.

So clear it explicitly at host start, where no session holds a display yet.
`pnputil /restart-device` recycles the WUDFHost process the driver's
`DECLARED_TARGETS` lives in, which is all it takes. Measured at **0.07 s**
against ~6 s of sleeps for the existing Disable+Enable cycle, and unlike that
cycle it is designed for a device in use, so it does not hit the refusal
`reload_vdisplay_adapter` documents as "the expected case here". In the same
call it also repaired an adapter found in CM_PROB_FAILED_POST_START (Code 43).

Best-effort throughout: a failure just leaves the adapter as it was and
sessions self-composite exactly as before. `PUNKTFUNK_CURSOR_CLEAN_START=0`
opts out.

Verified: `scripts/xcheck.sh windows` green; `cargo fmt --all` clean; full
`cargo build -p punktfunk-host --release --features nvenc` on .173 (xcheck
cannot reach punktfunk-host — it needs ffmpeg).
2026-08-08 18:16:06 +02:00
enricobuehler 62624c1daf feat(probe): --cursor-hold — stop the wiggle so the pointer can be parked on a target
The capture-model repro flags (`--cursor-capture` / `--cursor-nochannel`) drive
relative pointer motion in circles for the whole dump, which is right for
keeping a damage-driven desktop publishing frames but makes the pointer
impossible to aim: at radius 10 every 25 ms it walks several hundred pixels a
second, so a `SetCursorPos` on the host is undone before the next frame.

That mattered because the shape UNDER the pointer is the whole question. The
arrow is a colour cursor and proves nothing about the monochrome path — the
I-beam is the only common Windows system cursor with `hbmColor == null`, so it
is the only one that exercises `mono_planes_to_rgba`. Without being able to
park the pointer on a text field there is no way to photograph the case that
matters.

`--cursor-hold` primes the wiggle for ~3 s (enough to clear CURSOR_SUPPRESSED
and get metadata flowing) and then stops, leaving the pointer wherever the host
puts it for the rest of the dump.

Used it to settle the question on .173: a `--cursor-nochannel` session — the
byte-identical wire behaviour of an iPad/Android/tvOS client, which never
advertises CLIENT_CAP_CURSOR — receives a correctly rendered monochrome I-beam
in the video. Decoded from the dump with ffmpeg; the host log shows the session
took the forced-composite path with the GDI poller live.

Verified: `scripts/xcheck.sh windows` green; `cargo fmt --all` clean.
2026-08-08 17:03:33 +02:00
enricobuehler 7f6d1622ee test(pf-capture): pin the composite-cursor regen key and the blend retry escalation
The two pieces of logic the previous commit added had no tests, and both are
the kind that fail silently in opposite directions.

`blend_key_of` is now ONE definition shared by the regen test and the blend
itself, rather than the same tuple built at two call sites. That drift is the
actual bug shape: a key that reports "changed" while the drawn frame is
identical re-encodes for nothing, and a key that reports "unchanged" while the
pointer moved freezes it on screen. Both directions are asserted — a hidden
pointer keys identically wherever it moves, and every visible change (position,
shape serial, and the visible→hidden transition that must strip the pointer
from the frame) moves the key.

`next_blend_backoff` is extracted for the same reason `mono_planes_to_rgba`
was: the arithmetic a bug hides in does not need a live D3D11 device around it
to be checked. The test walks the escalation well past its ceiling and asserts
it PARKS there — an unbounded doubling would mean a device that comes back
after a long stall never gets picked up.

Verified: `scripts/xcheck.sh windows` green; `cargo fmt --all` clean. The tests
themselves compile and run only on Windows (the module is `cfg(windows)`), so
they are pending a run on a box.
2026-08-08 15:31:18 +02:00
enricobuehler 5c1db4662f fix(windows/cursor): the composite model paid a full-frame copy to draw nothing, and its failures were terminal
Three defects in the capture-model cursor path, all of them in the state a
session spends most of its life in.

1. The blend copied the whole frame even when nothing would be drawn.
   `prepare_blend_scratch` ran the scratch build + `CopyResource`
   unconditionally whenever compositing was on, and only THEN skipped the
   quad for a hidden pointer. A 4K FP16 ring slot is 66 MB, so at 120 fps
   that is ~8 GB/s of write bandwidth bought for nothing — and a game that
   grabbed the pointer hides it, which is exactly when the composite model
   is engaged. The overlay is now resolved FIRST and a hidden or unknown
   pointer returns `None` before any allocation or copy; the conversion
   reads the slot directly, which is the same frame it would have got.

2. A hidden pointer moving forced frame regeneration on an idle desktop.
   The regen key was `(serial, x, y, visible)` — raw cursor state — so a
   pointer a game had hidden re-encoded the last slot every time it moved,
   despite the frame being pixel-identical. The key is now what the blend
   would DRAW: `Some((serial, x, y))` visible, `None` hidden. The
   visible⇄hidden transitions still change it, so the frame that must gain
   or lose the pointer is still regenerated.

3. A blend failure was permanent. `cursor_blend_failed` was set once, warned
   once, and the session then streamed a pointer-less desktop for the rest
   of its life — including for a device-loss that heals a frame later. It is
   now a backoff (250 ms doubling to 4 s) that suppresses the blend, drops
   the pass so it rebuilds, and clears on the first success. Every
   escalation logs, so a permanently broken session is distinguishable from
   one transient hiccup at startup, and the recovery says so.

Also: the GDI poller now carries a heartbeat. `alive()` only asks whether the
thread exited, so a poller wedged on an input desktop it can no longer read
(`GetCursorInfo` failing every tick `continue`s before the publish) froze the
pointer in every frame at its last sampled shape while looking perfectly
healthy in the log. The capturer samples the publish count and warns once per
stall — this poller is the ONLY full-fidelity shape source, since the driver's
IddCx query is alpha-only, so its silence is the difference between a correct
pointer and a frozen one.

None of this depends on the open iPad diagnosis (planning-repo
`windows-cursor-model-determinism.md` §2): these are defects on their own
terms, and they are the prerequisites that make the composite path cheap and
observable enough to reason about.

Verified: `scripts/xcheck.sh windows` (clippy -D warnings for pf-frame,
pf-win-display, pf-capture, pf-vdisplay against x86_64-pc-windows-msvc) green;
`cargo fmt --all` clean. NOT run on a Windows box.
2026-08-08 14:25:27 +02:00
7 changed files with 616 additions and 61 deletions
+24 -3
View File
@@ -129,6 +129,16 @@ struct Args {
/// host must composite the metadata cursor on its own; decode the dump and look for the /// host must composite the metadata cursor on its own; decode the dump and look for the
/// pointer. /// pointer.
cursor_nochannel: bool, cursor_nochannel: bool,
/// `--cursor-hold` — with `--cursor-capture`/`--cursor-nochannel`, stop the relative wiggle
/// after a short priming burst instead of circling forever. The wiggle exists to keep a
/// damage-driven desktop publishing frames, but it also DRAGS the host pointer several hundred
/// pixels a second, which makes it impossible to hold the pointer over a chosen target — and
/// the shape under the pointer is the whole point when the question is "does the MONOCHROME
/// I-beam survive compositing?" (the arrow is a colour cursor and proves nothing about the
/// mono path). With this flag: prime for ~3 s so the pointer is un-suppressed and metadata is
/// flowing, then hold still so a `SetCursorPos` on the host can park it on a text field for
/// the rest of the dump.
cursor_hold: bool,
/// `--discover [SECS]` — browse the LAN for native (`_punktfunk._udp`) hosts for `SECS` /// `--discover [SECS]` — browse the LAN for native (`_punktfunk._udp`) hosts for `SECS`
/// seconds (default 4), print what's found, and exit. No connection is made. /// seconds (default 4), print what's found, and exit. No connection is made.
discover: Option<u64>, discover: Option<u64>,
@@ -309,6 +319,7 @@ fn parse_args() -> Args {
clock_resync: argv.iter().any(|a| a == "--clock-resync"), clock_resync: argv.iter().any(|a| a == "--clock-resync"),
cursor_capture: argv.iter().any(|a| a == "--cursor-capture"), cursor_capture: argv.iter().any(|a| a == "--cursor-capture"),
cursor_nochannel: argv.iter().any(|a| a == "--cursor-nochannel"), cursor_nochannel: argv.iter().any(|a| a == "--cursor-nochannel"),
cursor_hold: argv.iter().any(|a| a == "--cursor-hold"),
} }
} }
@@ -900,13 +911,23 @@ async fn session(args: Args) -> Result<()> {
} }
}); });
let wiggle_conn = conn.clone(); let wiggle_conn = conn.clone();
let hold = args.cursor_hold;
tokio::spawn(async move { tokio::spawn(async move {
// Relative circles, forever: keeps the host pointer moving (and, on metadata-cursor // Relative circles: keeps the host pointer moving (and, on metadata-cursor
// compositors, keeps cursor updates flowing) for the whole dump. // compositors, keeps cursor updates flowing) for the whole dump — unless
// `--cursor-hold`, which primes and then stops so the pointer can be parked.
tokio::time::sleep(std::time::Duration::from_secs(2)).await; tokio::time::sleep(std::time::Duration::from_secs(2)).await;
tracing::info!("cursor-capture: relative pointer wiggle running"); tracing::info!(hold, "cursor-capture: relative pointer wiggle running");
let prime_until = std::time::Instant::now() + std::time::Duration::from_secs(3);
let mut t = 0.0f64; let mut t = 0.0f64;
loop { loop {
if hold && std::time::Instant::now() >= prime_until {
tracing::info!(
"cursor-capture: wiggle primed and STOPPED (--cursor-hold) — the pointer \
now stays where the host puts it"
);
return;
}
let e = InputEvent { let e = InputEvent {
kind: InputKind::MouseMove, kind: InputKind::MouseMove,
_pad: [0; 3], _pad: [0; 3],
+270 -54
View File
@@ -212,6 +212,47 @@ struct KeyedMutexGuard<'a> {
/// (`frame_transport.rs`). /// (`frame_transport.rs`).
const WAIT_ABANDONED_HRESULT: i32 = 0x0000_0080; const WAIT_ABANDONED_HRESULT: i32 = 0x0000_0080;
/// First retry delay after a composite-blend failure — short enough that a transient device-loss
/// costs a few pointer-less frames rather than the rest of the session.
const BLEND_RETRY_MIN: Duration = Duration::from_millis(250);
/// Ceiling for the doubling retry: a genuinely broken device stops burning a frame-sized texture
/// allocation every quarter second, while still recovering within ~4 s if it ever comes back.
const BLEND_RETRY_MAX: Duration = Duration::from_secs(4);
/// How long the poller may publish NOTHING before the capturer calls it wedged. It polls at
/// `CursorPoller::INTERVAL` (4 ms), so this is ~250 missed publishes — far outside any scheduling
/// hiccup, and still fast enough to name the fault while a user is still looking at it.
const POLLER_STALL: Duration = Duration::from_secs(1);
/// The next retry delay after a composite-blend failure: [`BLEND_RETRY_MIN`] for the first, then
/// doubling per consecutive failure up to [`BLEND_RETRY_MAX`]. Free function so the escalation is
/// testable without a live D3D11 device (the `mono_planes_to_rgba` precedent — the arithmetic a
/// bug would hide in does not need the plumbing around it).
fn next_blend_backoff(prev: Option<Duration>) -> Duration {
prev.map_or(BLEND_RETRY_MIN, |b| (b * 2).min(BLEND_RETRY_MAX))
}
/// The composite-regen change key for an overlay: what a blend would DRAW — `(serial, x, y)` for a
/// visible pointer, `None` when nothing would be drawn. ONE definition, used by both the regen test
/// and the blend itself, because the two drifting apart is precisely the bug shape here: a key that
/// says "changed" while the drawn frame is identical re-encodes for nothing, and a key that says
/// "unchanged" while the pointer moved freezes it on screen.
fn blend_key_of(ov: Option<&pf_frame::CursorOverlay>) -> Option<(u64, i32, i32)> {
ov.filter(|o| o.visible).map(|o| (o.serial, o.x, o.y))
}
/// A composite-blend failure and its pending retry ([`IddPushCapturer::blend_fail`]).
struct BlendFail {
/// No blend is attempted before this instant.
retry_at: Instant,
/// The delay that produced `retry_at`; doubles per consecutive failure up to
/// [`BLEND_RETRY_MAX`].
backoff: Duration,
/// Consecutive failures without an intervening success — logged, so a session that is
/// permanently pointer-less is distinguishable from one that hiccupped once.
consecutive: u32,
}
impl<'a> KeyedMutexGuard<'a> { impl<'a> KeyedMutexGuard<'a> {
/// Acquire `mutex` at `key`, waiting up to `timeout_ms`. `None` if the acquire times out / errors /// Acquire `mutex` at `key`, waiting up to `timeout_ms`. `None` if the acquire times out / errors
/// (the caller skips the frame), so the guard is only ever held when the lock is genuinely held. /// (the caller skips the frame), so the guard is only ever held when the lock is genuinely held.
@@ -385,13 +426,26 @@ pub struct IddPushCapturer {
/// to a visible pointer is compositing here. Pins `composite_cursor` on — nothing may turn /// to a visible pointer is compositing here. Pins `composite_cursor` on — nothing may turn
/// it off (there is no channel to hand the pointer to). /// it off (there is no channel to hand the pointer to).
composite_forced: bool, composite_forced: bool,
/// The cursor-quad blend pass (lazy; per capture device). `None` after a build failure — /// The cursor-quad blend pass (lazy; per capture device). `None` before the first blend and
/// composite mode then degrades to pointer-less frames (warned once). /// after a failure dropped it; rebuilt on the next attempt that is not suppressed.
cursor_blend: Option<cursor_blend::CursorBlendPass>, cursor_blend: Option<cursor_blend::CursorBlendPass>,
cursor_blend_failed: bool, /// Composite-blend failure state. `None` = healthy. A failure used to be TERMINAL — one warn,
/// a sticky flag, and the session then streamed a pointer-less desktop for its whole life —
/// but the causes that actually occur (device loss, a transient allocation failure on the
/// frame-sized scratch) heal, and the pointer is the one thing a capture-model session cannot
/// do without. So a failure now only suppresses the blend until `retry_at`, doubling from
/// [`BLEND_RETRY_MIN`] to [`BLEND_RETRY_MAX`] while failures continue, and the first success
/// clears it.
blend_fail: Option<BlendFail>,
/// Sticky: [`Self::live_cursor`] has fallen back to the driver's shm section. The two sources /// Sticky: [`Self::live_cursor`] has fallen back to the driver's shm section. The two sources
/// keep independent serial namespaces, so once crossed we never go back (see there). /// keep independent serial namespaces, so once crossed we never go back (see there).
cursor_shm_latched: bool, cursor_shm_latched: bool,
/// Poller heartbeat watch: the last sampled publish count and when it last ADVANCED. A poller
/// that is `alive()` but wedged stops advancing it while never exiting — invisible before.
cursor_poll_watch: (u64, Instant),
/// Whether the wedged-poller warning has already been emitted for the CURRENT stall (cleared
/// when it resumes), so a permanently wedged poller warns once rather than every tick.
cursor_poll_stalled: bool,
/// The frame-sized blend scratch (slot copy + cursor quad): texture + SRV + (w, h, fmt) /// The frame-sized blend scratch (slot copy + cursor quad): texture + SRV + (w, h, fmt)
/// it was built for — rebuilt when the ring geometry changes. /// it was built for — rebuilt when the ring geometry changes.
blend_scratch: Option<( blend_scratch: Option<(
@@ -401,10 +455,12 @@ pub struct IddPushCapturer {
u32, u32,
DXGI_FORMAT, DXGI_FORMAT,
)>, )>,
/// The (serial, x, y, visible) of the LAST blended pointer — the composite-regen change /// What the LAST blend actually DREW — the composite-regen change key: pointer-only motion
/// key: pointer-only motion produces no driver publish (the declared hardware cursor /// produces no driver publish (the declared hardware cursor doesn't dirty frames), so
/// doesn't dirty frames), so `try_consume` regenerates from the last slot when this moves. /// `try_consume` regenerates from the last slot when this changes. `None` = the frame carries
last_blend_key: Option<(u64, i32, i32, bool)>, /// no pointer (hidden or no shape yet), which is why a HIDDEN pointer's position is not part
/// of the key — see [`Self::cursor_blend_key`].
last_blend_key: Option<(u64, i32, i32)>,
/// The ring slot of the last FRESH publish — the regen source. /// The ring slot of the last FRESH publish — the regen source.
last_slot: Option<usize>, last_slot: Option<usize>,
/// The target's SDR-white scale (vs 80 nits) for HDR cursor compositing — refreshed on /// The target's SDR-white scale (vs 80 nits) for HDR cursor compositing — refreshed on
@@ -1211,10 +1267,17 @@ impl IddPushCapturer {
/// poller meant pointer-less frames, not a degraded pointer. /// poller meant pointer-less frames, not a degraded pointer.
fn live_cursor(&mut self) -> Option<pf_frame::CursorOverlay> { fn live_cursor(&mut self) -> Option<pf_frame::CursorOverlay> {
if !self.cursor_shm_latched { if !self.cursor_shm_latched {
if let Some(p) = &self.cursor_poll { // Sample the heartbeat and the snapshot together, then drop the borrow so the watch
if p.alive() { // can take `&mut self`. `alive()` is liveness only — `watch_cursor_publishes` is what
return p.read(); // tells a working poller apart from a wedged one.
} let sampled = self
.cursor_poll
.as_ref()
.filter(|p| p.alive())
.map(|p| (p.publishes(), p.read()));
if let Some((n, overlay)) = sampled {
self.watch_cursor_publishes(n);
return overlay;
} }
// The poller is gone (or never started) and we are about to read the shm — latch, so a // The poller is gone (or never started) and we are about to read the shm — latch, so a
// poller that somehow reports alive again cannot re-cross the serial namespaces. // poller that somehow reports alive again cannot re-cross the serial namespaces.
@@ -1255,17 +1318,91 @@ impl IddPushCapturer {
); );
} }
/// The (serial, x, y, visible) of the CURRENT live cursor — the composite-regen change key. /// Watch the GDI poller's heartbeat and log the transitions. The poller is the ONLY
/// `None` while no source has a shape yet. /// full-fidelity shape source (the driver's query is alpha-only — `cursor_poll.rs`), so a
fn cursor_blend_key(&mut self) -> Option<(u64, i32, i32, bool)> { /// poller that is alive but no longer publishing freezes the pointer in every frame at its
self.live_cursor().map(|o| (o.serial, o.x, o.y, o.visible)) /// last sampled shape and position. That state used to be completely silent: `alive()` stays
/// true, the slot keeps returning its last snapshot, and nothing in the log distinguishes it
/// from a genuinely motionless pointer.
fn watch_cursor_publishes(&mut self, n: u64) {
let (last, since) = self.cursor_poll_watch;
if n != last {
self.cursor_poll_watch = (n, Instant::now());
if self.cursor_poll_stalled {
self.cursor_poll_stalled = false;
tracing::info!(
target_id = self.target_id,
"cursor poller resumed publishing — the pointer tracks again"
);
}
} else if !self.cursor_poll_stalled && since.elapsed() >= POLLER_STALL {
self.cursor_poll_stalled = true;
tracing::warn!(
target_id = self.target_id,
stalled_ms = since.elapsed().as_millis() as u64,
"cursor poller is ALIVE but has stopped publishing — the pointer is frozen at its \
last sampled shape/position (input-desktop reads failing every tick?)"
);
}
}
/// Is the composite blend currently suppressed by a failure's backoff?
fn blend_suppressed(&self) -> bool {
self.blend_fail
.as_ref()
.is_some_and(|f| Instant::now() < f.retry_at)
}
/// Record a composite-blend failure and arm the next retry (see [`BlendFail`]). Logs EVERY
/// escalation rather than only the first — a pointer-less capture-model session is a
/// user-visible fault, and the old warn-once left a permanently broken one indistinguishable
/// in the log from a single transient hiccup at startup.
fn note_blend_failure(&mut self, why: &str) {
let backoff = next_blend_backoff(self.blend_fail.as_ref().map(|f| f.backoff));
let consecutive = self.blend_fail.as_ref().map_or(1, |f| f.consecutive + 1);
self.blend_fail = Some(BlendFail {
retry_at: Instant::now() + backoff,
backoff,
consecutive,
});
tracing::warn!(
consecutive,
retry_in_ms = backoff.as_millis() as u64,
"cursor composite: {why} — frames stay pointer-less until the retry succeeds"
);
}
/// A blend succeeded: retire any failure record so the next one starts at the short backoff.
fn note_blend_success(&mut self) {
if let Some(f) = self.blend_fail.take() {
tracing::info!(
after_consecutive_failures = f.consecutive,
"cursor composite: blend recovered — the pointer is back in frames"
);
}
}
/// What a blend would DRAW this tick — `(serial, x, y)` for a visible pointer, `None` for a
/// hidden or not-yet-known one. Keyed on the drawn RESULT rather than on raw cursor state so
/// that a HIDDEN pointer moving — routine, because that is exactly what a game that grabbed
/// the pointer does — cannot force a frame regeneration on an otherwise idle desktop. The
/// visible⇄hidden transitions still change the key (`Some`⇄`None`), so the frame that must
/// gain or lose the pointer is still regenerated.
fn cursor_blend_key(&mut self) -> Option<(u64, i32, i32)> {
blend_key_of(self.live_cursor().as_ref())
} }
/// Composite the pointer for this convert: ensure the frame-sized blend scratch, copy the /// Composite the pointer for this convert: ensure the frame-sized blend scratch, copy the
/// slot into it, and alpha-blend the GDI poller's shape at its polled position. Returns the /// slot into it, and alpha-blend the GDI poller's shape at its polled position. Returns the
/// scratch (texture + SRV) the conversion should read INSTEAD of the slot; `None` degrades /// scratch (texture + SRV) the conversion should read INSTEAD of the slot; `None` degrades
/// to the pointer-less slot (scratch/pass creation failed — warned once). A hidden pointer /// to the pointer-less slot, which is the correct frame whenever nothing would be drawn.
/// blends nothing (the plain copy is the correct frame). ///
/// **There is NO scratch and NO copy when the pointer is hidden or unknown.** The full-frame
/// `CopyResource` below is the single largest cost of the composite model — a 4K FP16 ring
/// slot is 66 MB, so at 120 fps an unconditional copy is ~8 GB/s of write bandwidth — and it
/// buys nothing when the blend that follows draws nothing. A game that grabbed the pointer
/// hides it, so this early-out is what makes the capture model free in the state it spends
/// most of its life in.
/// ///
/// # Safety /// # Safety
/// D3D11 calls on the owning capture/encode thread's device + immediate context, called /// D3D11 calls on the owning capture/encode thread's device + immediate context, called
@@ -1274,6 +1411,18 @@ impl IddPushCapturer {
&mut self, &mut self,
slot_tex: &ID3D11Texture2D, slot_tex: &ID3D11Texture2D,
) -> Option<(ID3D11Texture2D, ID3D11ShaderResourceView)> { ) -> Option<(ID3D11Texture2D, ID3D11ShaderResourceView)> {
// Resolve WHAT WOULD BE DRAWN first, and record it as the applied key even when that is
// "nothing" — `try_consume`'s regen test compares against this, so an early-out must still
// leave the key describing the frame we are about to emit. Through `live_cursor`, so a
// dead poller degrades to the shm section here too.
let overlay = self.live_cursor();
self.last_blend_key = blend_key_of(overlay.as_ref());
let ov = overlay.filter(|o| o.visible)?;
// Blending is suppressed while a recent failure's backoff runs — skip the scratch and the
// copy too, not just the draw: with nothing to draw onto it, the copy is pure waste.
if self.blend_suppressed() {
return None;
}
// SAFETY: per the contract above, D3D11 calls on the owning thread's device + immediate // SAFETY: per the contract above, D3D11 calls on the owning thread's device + immediate
// context while the slot's keyed mutex is held. `CreateTexture2D`/`CreateShaderResourceView` // context while the slot's keyed mutex is held. `CreateTexture2D`/`CreateShaderResourceView`
// take a fully-initialized stack descriptor plus live out-params and are `.ok()`-checked before // take a fully-initialized stack descriptor plus live out-params and are `.ok()`-checked before
@@ -1325,13 +1474,7 @@ impl IddPushCapturer {
self.blend_scratch = Some((t, v, self.width, self.height, fmt)); self.blend_scratch = Some((t, v, self.width, self.height, fmt));
} }
None => { None => {
if !self.cursor_blend_failed { self.note_blend_failure("scratch creation failed");
self.cursor_blend_failed = true;
tracing::warn!(
"cursor blend scratch creation failed — capture-model frames stay \
pointer-less this session"
);
}
return None; return None;
} }
} }
@@ -1339,38 +1482,33 @@ impl IddPushCapturer {
let (tex, srv, ..) = self.blend_scratch.as_ref().expect("just ensured"); let (tex, srv, ..) = self.blend_scratch.as_ref().expect("just ensured");
let (tex, srv) = (tex.clone(), srv.clone()); let (tex, srv) = (tex.clone(), srv.clone());
self.context.CopyResource(&tex, slot_tex); self.context.CopyResource(&tex, slot_tex);
// Blend the pointer (visible shapes only; hidden = the copy alone is the frame). // Draw `ov` — resolved and keyed at the top, where a hidden pointer already took the
// Through `live_cursor`, so a dead poller degrades to the shm section HERE too — this // early-out, so reaching here means there IS something to blend.
// is the path that actually draws the pointer in the composite model, and the one that if self.cursor_blend.is_none() {
// used to read the poller unconditionally. match cursor_blend::CursorBlendPass::new(&self.device) {
let overlay = self.live_cursor(); Ok(p) => self.cursor_blend = Some(p),
self.last_blend_key = overlay.as_ref().map(|o| (o.serial, o.x, o.y, o.visible)); Err(e) => {
if let Some(ov) = overlay.filter(|o| o.visible) { self.note_blend_failure(&format!("blend pass build failed: {e:#}"));
if self.cursor_blend.is_none() && !self.cursor_blend_failed {
match cursor_blend::CursorBlendPass::new(&self.device) {
Ok(p) => self.cursor_blend = Some(p),
Err(e) => {
self.cursor_blend_failed = true;
tracing::warn!(
"cursor blend pass build failed — capture-model frames stay \
pointer-less this session: {e:#}"
);
}
} }
} }
if let Some(pass) = self.cursor_blend.as_mut() { }
// FP16 ring = scRGB linear composition (HDR): linearize the sRGB shape and if let Some(pass) = self.cursor_blend.as_mut() {
// scale it to the target's SDR white so it matches the desktop around it. // FP16 ring = scRGB linear composition (HDR): linearize the sRGB shape and
let scale = if self.display_hdr { // scale it to the target's SDR white so it matches the desktop around it.
self.sdr_white_scale let scale = if self.display_hdr {
} else { self.sdr_white_scale
0.0 } else {
}; 0.0
if let Err(e) = pass.blend(&self.device, &self.context, &tex, &ov, scale) { };
if !self.cursor_blend_failed { match pass.blend(&self.device, &self.context, &tex, &ov, scale) {
self.cursor_blend_failed = true; // One good draw retires the whole failure record: whatever broke has healed,
tracing::warn!("cursor blend draw failed — pointer-less frames: {e:#}"); // and the next failure should get the SHORT retry, not the escalated one.
} Ok(()) => self.note_blend_success(),
Err(e) => {
// Drop the pass so the block above rebuilds it: a device-loss failure is
// transient, but a pass built against the lost device never succeeds again.
self.cursor_blend = None;
self.note_blend_failure(&format!("blend draw failed: {e:#}"));
} }
} }
} }
@@ -2075,6 +2213,84 @@ mod tests {
use super::stall::Stall; use super::stall::Stall;
use super::*; use super::*;
/// A `CursorOverlay` at `(x, y)` with `serial`, visible or not. `rgba` is never read by the
/// key/backoff logic under test, so a 1×1 pixel keeps the fixtures honest about that.
fn overlay(serial: u64, x: i32, y: i32, visible: bool) -> pf_frame::CursorOverlay {
pf_frame::CursorOverlay {
x,
y,
w: 1,
h: 1,
rgba: std::sync::Arc::new(vec![0, 0, 0, 0]),
serial,
hot_x: 0,
hot_y: 0,
visible,
}
}
/// The regen key is what would be DRAWN, so a hidden pointer keys to `None` no matter where it
/// is. This is the whole point: a game that grabbed the pointer moves it constantly, and each
/// of those moves used to re-encode the last slot for a frame that is pixel-identical.
#[test]
fn a_hidden_pointer_has_no_blend_key_wherever_it_moves() {
assert_eq!(blend_key_of(None), None, "no overlay ⇒ nothing drawn");
assert_eq!(
blend_key_of(Some(&overlay(7, 10, 10, false))),
None,
"hidden ⇒ nothing drawn"
);
assert_eq!(
blend_key_of(Some(&overlay(7, 999, 999, false))),
blend_key_of(Some(&overlay(7, 10, 10, false))),
"a hidden pointer moving must NOT look like a change"
);
}
/// …but every transition that alters the drawn frame still changes the key, or the pointer
/// would freeze on screen (the failure mode opposite to the one above).
#[test]
fn every_visible_change_moves_the_blend_key() {
let shown = blend_key_of(Some(&overlay(7, 10, 10, true)));
assert_eq!(shown, Some((7, 10, 10)));
assert_ne!(
shown,
blend_key_of(Some(&overlay(7, 11, 10, true))),
"a visible pointer moving is a change"
);
assert_ne!(
shown,
blend_key_of(Some(&overlay(8, 10, 10, true))),
"a new shape at the same spot is a change"
);
assert_ne!(
shown,
blend_key_of(Some(&overlay(7, 10, 10, false))),
"visible → hidden must regenerate the frame that loses the pointer"
);
}
/// The retry escalates and then holds at the ceiling — it must never grow without bound (the
/// point of a ceiling is that a device which comes back is picked up within it).
#[test]
fn the_blend_retry_backoff_doubles_then_caps() {
let first = next_blend_backoff(None);
assert_eq!(first, BLEND_RETRY_MIN, "the first failure retries quickly");
assert_eq!(next_blend_backoff(Some(first)), first * 2, "then doubles");
// Walk it well past the cap and assert it PARKS there rather than overshooting.
let mut b = first;
for _ in 0..32 {
b = next_blend_backoff(Some(b));
}
assert_eq!(b, BLEND_RETRY_MAX, "escalation parks at the ceiling");
assert_eq!(
next_blend_backoff(Some(BLEND_RETRY_MAX)),
BLEND_RETRY_MAX,
"and stays there"
);
}
/// W14: the mint must stay inside the publish token's 24-bit generation field, and must skip 0. /// W14: the mint must stay inside the publish token's 24-bit generation field, and must skip 0.
/// ///
/// `IDD_GENERATION` is a full `u32` while `FrameToken` carries 24 bits and `unpack` MASKS what it /// `IDD_GENERATION` is a full `u32` while `FrameToken` carries 24 bits and `unpack` MASKS what it
@@ -68,6 +68,11 @@ pub(super) struct CursorPoller {
/// while the secure desktop needs the software-cursor path to render (see /// while the secure desktop needs the software-cursor path to render (see
/// `IddPushCapturer::poll_secure_desktop`). /// `IddPushCapturer::poll_secure_desktop`).
secure: Arc<AtomicBool>, secure: Arc<AtomicBool>,
/// Monotonic count of published snapshots — the poller's HEARTBEAT. It advances once per
/// successful poll (a failed `GetCursorInfo` `continue`s before the publish), so a thread that
/// is wedged on an input desktop it can no longer read stops advancing this while never
/// exiting. [`Self::alive`] cannot see that state: it only asks whether the thread finished.
ticks: Arc<AtomicU64>,
thread: Option<std::thread::JoinHandle<()>>, thread: Option<std::thread::JoinHandle<()>>,
} }
@@ -106,10 +111,12 @@ impl CursorPoller {
let slot: Arc<Mutex<Option<pf_frame::CursorOverlay>>> = Arc::new(Mutex::new(None)); let slot: Arc<Mutex<Option<pf_frame::CursorOverlay>>> = Arc::new(Mutex::new(None));
let stop = Arc::new(AtomicBool::new(false)); let stop = Arc::new(AtomicBool::new(false));
let secure = Arc::new(AtomicBool::new(false)); let secure = Arc::new(AtomicBool::new(false));
let (slot_t, stop_t, secure_t) = (slot.clone(), stop.clone(), secure.clone()); let ticks = Arc::new(AtomicU64::new(0));
let (slot_t, stop_t, secure_t, ticks_t) =
(slot.clone(), stop.clone(), secure.clone(), ticks.clone());
let thread = std::thread::Builder::new() let thread = std::thread::Builder::new()
.name("pf-cursor-poll".into()) .name("pf-cursor-poll".into())
.spawn(move || run(target_id, rect, &slot_t, &stop_t, &secure_t)) .spawn(move || run(target_id, rect, &slot_t, &stop_t, &secure_t, &ticks_t))
.ok(); .ok();
if thread.is_none() { if thread.is_none() {
tracing::warn!("cursor poller thread spawn failed — cursor falls back to driver shm"); tracing::warn!("cursor poller thread spawn failed — cursor falls back to driver shm");
@@ -118,6 +125,7 @@ impl CursorPoller {
slot, slot,
stop, stop,
secure, secure,
ticks,
thread, thread,
} }
} }
@@ -133,7 +141,14 @@ impl CursorPoller {
self.secure.load(Ordering::Relaxed) self.secure.load(Ordering::Relaxed)
} }
/// The heartbeat count (see [`Self::ticks`]). Compared against its own previous value by the
/// capturer — the ABSOLUTE value means nothing, only whether it is still moving.
pub(super) fn publishes(&self) -> u64 {
self.ticks.load(Ordering::Relaxed)
}
/// Whether the worker thread is (still) alive — `false` degrades the capturer to the shm read. /// Whether the worker thread is (still) alive — `false` degrades the capturer to the shm read.
/// Note this is liveness, NOT health: see [`Self::publishes`].
pub(super) fn alive(&self) -> bool { pub(super) fn alive(&self) -> bool {
self.thread.as_ref().is_some_and(|t| !t.is_finished()) self.thread.as_ref().is_some_and(|t| !t.is_finished())
} }
@@ -155,6 +170,7 @@ fn run(
slot: &Mutex<Option<pf_frame::CursorOverlay>>, slot: &Mutex<Option<pf_frame::CursorOverlay>>,
stop: &AtomicBool, stop: &AtomicBool,
secure: &AtomicBool, secure: &AtomicBool,
ticks: &AtomicU64,
) { ) {
// Physical-pixel coordinates on this thread regardless of the process's DPI awareness: // Physical-pixel coordinates on this thread regardless of the process's DPI awareness:
// `rect` comes from CCD (always physical), and a DPI-virtualized `GetCursorInfo` position // `rect` comes from CCD (always physical), and a DPI-virtualized `GetCursorInfo` position
@@ -306,6 +322,8 @@ fn run(
} }
}); });
*slot.lock().unwrap_or_else(|p| p.into_inner()) = overlay; *slot.lock().unwrap_or_else(|p| p.into_inner()) = overlay;
// Heartbeat AFTER the publish, so it counts snapshots the capturer can actually read.
ticks.fetch_add(1, Ordering::Relaxed);
} }
} }
@@ -656,7 +656,9 @@ impl IddPushCapturer {
composite_cursor: composite_forced, composite_cursor: composite_forced,
composite_forced, composite_forced,
cursor_blend: None, cursor_blend: None,
cursor_blend_failed: false, blend_fail: None,
cursor_poll_watch: (0, std::time::Instant::now()),
cursor_poll_stalled: false,
cursor_shm_latched: false, cursor_shm_latched: false,
blend_scratch: None, blend_scratch: None,
last_blend_key: None, last_blend_key: None,
@@ -421,6 +421,34 @@ pub fn hw_cursor_capable() -> bool {
m.driver_proto.load(Ordering::Relaxed) >= 5 m.driver_proto.load(Ordering::Relaxed) >= 5
} }
/// Is NO session currently streaming to a virtual display?
///
/// The safety question for anything that tears the adapter down — notably
/// [`crate::driver::clean_cursor_for_next_session`], whose `pnputil /restart-device` takes every
/// monitor on the adapter with it. Only [`SlotState::Active`] counts: that is a session with live
/// references, and destroying its monitor mid-stream is the cross-session damage worth refusing.
///
/// `Lingering`/`Pinned` slots deliberately do NOT count. They are keep-alive monitors with no
/// session attached, and a reconnect **already** preempts and recreates them — "a reused IddCx
/// swap-chain is dead" (see [`SlotState::Pinned`]) — so a device restart destroys nothing the
/// reconnect was not going to destroy anyway. Counting them was too conservative to be useful: the
/// case this gate exists for is exactly *disconnect from a desktop session, reconnect in capture
/// mode*, and the disconnected session's monitor is lingering at precisely that moment, so the
/// clean-up could never fire when it was most wanted (observed on `.173`, 2026-08-08).
pub fn no_active_sessions() -> bool {
match VDM.get() {
// Before the first backend open there is nothing to protect.
None => true,
Some(m) => !m
.state
.lock()
.unwrap_or_else(|e| e.into_inner())
.slots
.values()
.any(|s| matches!(s, SlotState::Active { .. })),
}
}
pub fn control_device_handle() -> Option<HANDLE> { pub fn control_device_handle() -> Option<HANDLE> {
VDM.get().and_then(VirtualDisplayManager::device_handle) VDM.get().and_then(VirtualDisplayManager::device_handle)
} }
@@ -158,6 +158,226 @@ enum AdapterCycle {
Refused(String), Refused(String),
} }
/// Restart the pf-vdisplay device to CLEAR a sticky IddCx hardware-cursor declare, so sessions that
/// do not want the host to own the pointer get the OS's own cursor compositing back (full fidelity,
/// zero host cost — no GDI poller, no per-frame blend, true XOR instead of our outline
/// approximation).
///
/// **Why this exists.** A hardware-cursor declare is irrevocable and ADAPTER-WIDE
/// (`pf-driver-proto` v6 note): once any desktop-mode session declares, DWM stops compositing the
/// pointer into EVERY later frame on that adapter, and every subsequent session — including
/// capture-latched ones that never asked for a cursor channel — has to self-composite. The state
/// lives in the driver's `DECLARED_TARGETS`, whose scope is the WUDFHost process, so recycling that
/// process clears it.
///
/// **Why `/restart-device` and not the [`reload_vdisplay_adapter`] cycle.** Measured on-glass
/// 2026-08-08 (`.173`): `pnputil /restart-device` returned in **0.07 s** with a NEW WUDFHost pid,
/// against ~6 s of sleeps for `Disable`+`Enable` — and, being designed for a device that is in use,
/// it does not hit the refusal that doc calls "the expected case here". It also repaired an adapter
/// found in `CM_PROB_FAILED_POST_START` (Code 43) in the same call.
///
/// ⚠⚠ **This is a ONCE-PER-BOOT lever, not a cheap one.** Measured on `.173` 2026-08-08: the first
/// `/restart-device` after a cold boot succeeds in 0.07 s; every later one in the same boot fails
/// with *"Das System muss neu gestartet werden, damit Konfigurationsvorgänge abgeschlossen
/// werden"*, and repeated attempts additionally push the devnode into `restart pending`. So this
/// can clean the adapter at host start-up and nowhere else — anything wanting to un-declare
/// mid-boot (e.g. giving a capture session back the lossless pointer after a desktop session) needs
/// a different mechanism to recycle the driver's WUDFHost process, which is where the declare
/// actually lives.
///
/// ⚠ It tears the adapter down, so it must run only when NO session holds a display — the host
/// start-up path. `PUNKTFUNK_CURSOR_CLEAN_START=0` disables it.
///
/// Returns `true` only when pnputil reported success. Best-effort: a failure just leaves the
/// adapter as it was (sessions then self-composite exactly as before).
/// The driver's WUDFHost pid, from the most recent ADD reply. `0` before any monitor was created.
static LAST_WUDF_PID: std::sync::atomic::AtomicU32 = std::sync::atomic::AtomicU32::new(0);
/// Clear a sticky hardware-cursor declare by recycling the driver's WUDFHost process.
///
/// The declare is irrevocable and adapter-wide, but its scope is the WUDFHost process
/// (`monitor.rs` `DECLARED_TARGETS`) — so killing that process drops it. WUDF respawns the host on
/// the next open, with a fresh adapter object.
///
/// **This is what makes un-declaring possible mid-boot.** `pnputil /restart-device` also works but
/// is a ONCE-PER-BOOT operation (see [`restart_device_for_clean_cursor`]); the start-up clean
/// spends it, leaving nothing for the desktop-session→reconnect case. Measured on `.173`
/// 2026-08-08: pid 3872 → 19932, `adapter_luid` 0x8ed607 → 0x1a8f6ca, `cursor_excluded` true →
/// **false**, next session streamed normally.
///
/// Same precondition as the device restart: no session may hold a display, because every monitor
/// on the adapter dies with the host.
fn recycle_wudfhost() -> bool {
let pid = LAST_WUDF_PID.load(std::sync::atomic::Ordering::Relaxed);
if pid == 0 {
tracing::info!("cursor: no driver host pid known yet — nothing to recycle");
return false;
}
// taskkill rather than OpenProcess/TerminateProcess: the host runs as SYSTEM, so it already has
// the rights, and shelling out keeps this off the unsafe-proof budget for a once-per-session
// maintenance action.
match std::process::Command::new(
std::env::var("SystemRoot")
.map(|r| format!(r"{r}\System32 askkill.exe"))
.unwrap_or_else(|_| "taskkill.exe".to_string()),
)
.args(["/PID", &pid.to_string(), "/F"])
.output()
{
Ok(o) if o.status.success() => {
tracing::info!(
pid,
"cursor: recycled the driver's WUDFHost — the hardware-cursor declare is gone"
);
LAST_WUDF_PID.store(0, std::sync::atomic::Ordering::Relaxed);
true
}
Ok(o) => {
tracing::warn!(
pid,
stderr = %String::from_utf8_lossy(&o.stderr).trim().replace('\n', " "),
"cursor: could not recycle the driver's WUDFHost — this session self-composites"
);
false
}
Err(e) => {
tracing::warn!(pid, error = %e, "cursor: taskkill spawn failed");
false
}
}
}
/// Has this host process DECLARED an IddCx hardware cursor since the adapter was last restarted?
/// Set by the ADD path when a session ASKS for a hardware cursor (the one place every declare
/// passes through); cleared when the declare is dropped. The host's own mirror of the
/// driver's `DECLARED_TARGETS` — cheaper than probing, and it only ever needs to be right about
/// "did WE dirty it", because a declare from an earlier BOOT is handled by the start-up clean.
static CURSOR_DECLARED: std::sync::atomic::AtomicBool = std::sync::atomic::AtomicBool::new(false);
/// Give the NEXT session back the lossless cursor: if an earlier session on this host declared the
/// hardware cursor and this one does not want it, restart the device to clear the sticky declare.
///
/// This is the case the start-up clean cannot reach — **run a desktop-mode session, disconnect,
/// reconnect in capture mode**. Same host process, so the adapter is still dirty from the first
/// session and the capture session would self-composite the pointer for its whole life. Declaring
/// is one-way and adapter-wide (`pf-driver-proto` v6), so the only way back is a device restart —
/// 0.07 s, measured.
///
/// Must be called BEFORE this session creates its display, and only when nothing else holds one:
/// the restart takes every monitor on the adapter with it.
///
/// Returns `true` only when it actually restarted.
pub fn clean_cursor_for_next_session(session_wants_declare: bool) -> bool {
use std::sync::atomic::Ordering;
if session_wants_declare || !CURSOR_DECLARED.load(Ordering::Relaxed) {
return false;
}
// Gated deliberately — a device restart is NOT free. Windows puts the devnode into
// "restart pending" after repeated cycles, and `/restart-device` then refuses with "a system
// restart is pending for this device" until an actual reboot (hit on .173 2026-08-08 after ~6
// restarts in one afternoon, which is also what made the earlier runs look like a wiring bug:
// the call ran, the restart failed, and nothing logged the failure). So restart only when a
// declare is actually outstanding, never speculatively.
let previously_declared = true;
// Refuse only while another session is STREAMING — a keep-alive (lingering/pinned) monitor has
// no session attached and a reconnect recreates it regardless, so restarting the adapter costs
// it nothing. Gating on keep-alive too made this dead code in the one case it exists for: after
// a desktop session disconnects its monitor LINGERS, which is exactly when the next
// capture-mode connect needs the declare gone (observed on .173).
if !super::manager::no_active_sessions() {
tracing::info!(
"cursor: this session wants no hardware cursor and an earlier one declared, but a display is still held (live or keep-alive) — skipping the adapter restart, so the pointer stays host-composited for this session"
);
return false;
}
if recycle_wudfhost() {
// The cached control handle died with the host process. Retire it so the next
// `ensure_device` reopens against the respawned WUDFHost — without this the ADD that
// follows runs on a stale handle and the session comes up with no frames at all.
super::manager::invalidate_cached_device("cursor clean: recycled the driver host");
std::thread::sleep(std::time::Duration::from_millis(1500));
CURSOR_DECLARED.store(false, Ordering::Relaxed);
tracing::info!(
previously_declared,
"cursor: restarted the adapter for this capture-mode session — any hardware-cursor \
declare is gone, so the OS composites the pointer itself (full fidelity, no host \
blend). previously_declared=false only means the host-side hint was unset; the \
restart is idempotent either way"
);
return true;
}
false
}
pub fn restart_device_for_clean_cursor() -> bool {
if std::env::var("PUNKTFUNK_CURSOR_CLEAN_START").is_ok_and(|v| v == "0") {
tracing::info!(
"pf-vdisplay: cursor clean-start disabled (PUNKTFUNK_CURSOR_CLEAN_START=0) — a sticky \
hardware-cursor declare from an earlier boot will keep sessions self-compositing"
);
return false;
}
// `$LASTEXITCODE` is pre-seeded to 1 for the same reason `reload_vdisplay_adapter` does it: if
// pnputil never launches, a stale value must not read as success.
const PS: &str = "$ErrorActionPreference='SilentlyContinue'; \
$ad = Get-PnpDevice -Class Display | Where-Object { $_.FriendlyName -match 'punktfunk Virtual Display' } | Select-Object -First 1; \
if (-not $ad) { Write-Output 'ABSENT'; exit }; \
$pnp = ($env:SystemRoot + '\\System32\\pnputil.exe'); $LASTEXITCODE = 1; \
if (Test-Path $pnp) { $out = (& $pnp /restart-device $ad.InstanceId 2>&1 | Out-String) }; \
if ($LASTEXITCODE -eq 0) { Write-Output 'RESTARTED' } \
else { Write-Output ('FAILED ' + ($out -replace '\\s+', ' ')) }";
let ps = std::env::var("SystemRoot")
.map(|r| format!(r"{r}\System32\WindowsPowerShell\v1.0\powershell.exe"))
.unwrap_or_else(|_| "powershell.exe".to_string());
let out = match std::process::Command::new(&ps)
.args([
"-NoProfile",
"-NonInteractive",
"-ExecutionPolicy",
"Bypass",
"-Command",
PS,
])
.output()
{
Ok(o) => String::from_utf8_lossy(&o.stdout).trim().to_string(),
Err(e) => {
tracing::warn!(error = %e, "pf-vdisplay: cursor clean-start could not spawn powershell");
return false;
}
};
match out.as_str() {
"RESTARTED" => {
tracing::info!(
"pf-vdisplay: restarted the adapter at start-up — any sticky hardware-cursor \
declare is cleared, so sessions without a cursor channel get the OS's own \
(full-fidelity, zero-cost) pointer compositing until one declares again"
);
true
}
"ABSENT" => false, // driver not installed — nothing to clean, and `open` reports that later
// Keep pnputil's own text. The failure that actually occurs is "a system restart is
// pending for this device" — no retry fixes it, and a bare exit code hid it for three runs.
other => {
tracing::warn!(
outcome = other,
// Two distinct wordings, both meaning "not until you reboot":
// "Für das Gerät steht ein Systemneustart aus" (device restart pending)
// "Das System muss neu gestartet werden, damit …" (config ops need a reboot)
// The second is what you actually hit, and it appears after the FIRST successful
// restart of a boot — see the doc on `restart_device_for_clean_cursor`.
needs_reboot = other.contains("Systemneustart")
|| other.contains("muss neu gestartet werden")
|| other.to_ascii_lowercase().contains("restart is pending")
|| other.to_ascii_lowercase().contains("must be restarted"),
"pf-vdisplay: cursor clean-start did not restart the adapter — sessions without a \
cursor channel will self-composite the pointer if an earlier declare is sticky"
);
false
}
}
}
/// Reload the pf-vdisplay ADAPTER device — the in-process equivalent of `reset-pf-vdisplay.ps1` /// Reload the pf-vdisplay ADAPTER device — the in-process equivalent of `reset-pf-vdisplay.ps1`
/// step 3. A crashed/killed WUDFHost can leave the devnode "started" yet HOSTLESS (PnP Status OK, no /// step 3. A crashed/killed WUDFHost can leave the devnode "started" yet HOSTLESS (PnP Status OK, no
/// WUDFHost process, zero device-interface instances) — a zombie no session can open until the stack /// WUDFHost process, zero device-interface instances) — a zombie no session can open until the stack
@@ -353,6 +573,12 @@ pub unsafe fn send_cursor_channel(
dev: HANDLE, dev: HANDLE,
req: &control::SetCursorChannelRequest, req: &control::SetCursorChannelRequest,
) -> Result<()> { ) -> Result<()> {
// THE declare point. The driver declares its IddCx hardware cursor when this channel arrives —
// not from the ADD request's `hw_cursor` flag, which is why recording the declare there (and,
// before that, in `capture_virtual_output`) left the flag false and the between-session clean
// silently inert. The log line that names this moment is "cursor channel delivered - driver
// declares the hardware cursor".
CURSOR_DECLARED.store(true, std::sync::atomic::Ordering::Relaxed);
let mut none: [u8; 0] = []; let mut none: [u8; 0] = [];
// SAFETY: per this fn's contract `dev` is the live control handle; `bytes_of(req)` borrows the // SAFETY: per this fn's contract `dev` is the live control handle; `bytes_of(req)` borrows the
// caller's request across this synchronous call; no output buffer. // caller's request across this synchronous call; no output buffer.
@@ -679,6 +905,26 @@ impl VdisplayDriver for PfVdisplayDriver {
client_hdr: Option<punktfunk_core::quic::HdrMeta>, client_hdr: Option<punktfunk_core::quic::HdrMeta>,
hw_cursor: bool, hw_cursor: bool,
) -> Result<AddedMonitor> { ) -> Result<AddedMonitor> {
// Give a capture-mode session the LOSSLESS pointer back: if an earlier session declared a
// hardware cursor and this one does not want it, recycle the driver's host process BEFORE
// this monitor is added. The ADD path is the only place guaranteed to see every session
// (the handshake call site this replaced sat in a `match (source, compositor)` arm that is
// not taken on this host, so it never ran).
// ⚠ DISABLED BY DEFAULT — opt in with PUNKTFUNK_CURSOR_RECYCLE=1.
//
// The MECHANISM is proven (recycling the driver host clears the declare: measured pid
// 3872→19932, adapter_luid 0x8ed607→0x1a8f6ca, cursor_excluded true→false, next session
// streamed fine). What is NOT solved is calling it from HERE: `invalidate_cached_device`
// takes the manager `device` mutex, which this ADD path already holds, so the session
// DEADLOCKS — observed on .173, the ADD stops after SET_RENDER_ADAPTER and the client gets
// "no frames received". Its own doc warns about exactly this.
//
// The fix is a call site that runs OUTSIDE the mutex and still on every session's path;
// the handshake site tried before is not reached on this host. Until then this stays off:
// a session that self-composites is the old behaviour, a deadlocked one is a regression.
if !hw_cursor && std::env::var("PUNKTFUNK_CURSOR_RECYCLE").is_ok_and(|v| v == "1") {
clean_cursor_for_next_session(false);
}
let session_id = next_session_id(); let session_id = next_session_id();
// The client display's volume rides into the monitor's EDID CTA HDR block; all-zero = // The client display's volume rides into the monitor's EDID CTA HDR block; all-zero =
// unknown → the driver keeps its built-in defaults (also what an un-upgraded driver, which // unknown → the driver keeps its built-in defaults (also what an un-upgraded driver, which
@@ -824,7 +1070,14 @@ impl VdisplayDriver for PfVdisplayDriver {
tracing::info!( tracing::info!(
target_id = reply.target_id, target_id = reply.target_id,
adapter_luid = %format_args!("{:#x}", luid.LowPart), adapter_luid = %format_args!("{:#x}", luid.LowPart),
wudf_pid = reply.wudf_pid, wudf_pid = {
// The declare lives in THIS process (monitor.rs `DECLARED_TARGETS`), so remember it:
// recycling it is the only way to un-declare that does not cost the once-per-boot
// device restart. Proven on .173 2026-08-08 — killing it gave a new host pid, a NEW
// adapter luid, and `cursor_excluded=false`, with the next session streaming fine.
LAST_WUDF_PID.store(reply.wudf_pid, std::sync::atomic::Ordering::Relaxed);
reply.wudf_pid
},
cursor_excluded = reply.cursor_excluded != 0, cursor_excluded = reply.cursor_excluded != 0,
"pf-vdisplay monitor created {}x{}@{}", "pf-vdisplay monitor created {}x{}@{}",
mode.width, mode.width,
+17
View File
@@ -382,6 +382,23 @@ fn real_main() -> Result<()> {
// driver to a stray second host started while the service sat idle. // driver to a stray second host started while the service sat idle.
#[cfg(target_os = "windows")] #[cfg(target_os = "windows")]
vdisplay::manager::claim_instance_eagerly(); vdisplay::manager::claim_instance_eagerly();
// Clean-cursor start (design/windows-cursor-model-determinism.md §4.3): clear any
// sticky IddCx hardware-cursor declare left on the adapter by an EARLIER boot's
// desktop-mode session. That declare is irrevocable and adapter-wide, so without this
// every capture-latched session on the box self-composites the pointer for the rest of
// the adapter's life — paying a full-frame copy per visible-pointer frame and drawing
// our straight-alpha approximation of an XOR cursor — when the OS would otherwise
// composite it natively, for free, at full fidelity.
//
// It is NOT enough to wait for a reboot: with Fast Startup on (the Windows default) a
// shutdown+power-on is a hiberboot that RESTORES session 0 and its drivers, so the
// declare survives what the operator calls a reboot (measured: Kernel-Boot event id 27
// `0x1`, and `lsass`/`services` keeping their pre-"reboot" start times). Only a cold
// boot or a device restart actually clears it — and the device restart costs 0.07 s.
//
// Runs HERE, before any session holds a display: the restart tears the adapter down.
#[cfg(target_os = "windows")]
vdisplay::driver::restart_device_for_clean_cursor();
// Crash recovery for the experimental `pnp_disable_monitors` axis: re-enable any // Crash recovery for the experimental `pnp_disable_monitors` axis: re-enable any
// monitor devnodes a previous host disabled for an Exclusive session and never // monitor devnodes a previous host disabled for an Exclusive session and never
// restored (crash/kill/power loss) — before any new session touches the topology. // restored (crash/kill/power loss) — before any new session touches the topology.