The hand-back never checked that the panel came back, and a crashed host left game mode asleep #375

Merged
enricobuehler merged 1 commits from worktree-gamescope-idled-handback into main 2026-08-22 23:38:26 +00:00
1 Commits
Author SHA1 Message Date
enricobuehler c63e8cee39 fix(gamescope): the hand-back never checked that the panel came back, and a crashed host left game mode asleep
ci / bun-nix (pull_request) Successful in 28s
ci / docs-drift (pull_request) Successful in 1m5s
ci / web (pull_request) Successful in 1m7s
ci / docs-site (pull_request) Successful in 1m50s
ci / rust-arm64 (pull_request) Successful in 2m7s
ci / rust (pull_request) Successful in 7m5s
android / android (pull_request) Successful in 7m35s
Field reports on 0.31.x, Bazzite and Nobara: after disconnecting, the box's own
physical screen stays black.

I could not reproduce it (PR #375 has the full negative write-up: five scenarios
across both distro families on the real VMs, all recovering cleanly, and the
mechanism I first proposed disproved on glass). So this does not guess at the
trigger. It closes the gap that lets ANY trigger end as a dark panel, and fixes
the one black-screen path I could prove.

## The restore never checked its own work

`do_restore_tv_session` issues a lifecycle verb and logs what systemd said about
the JOB. "The job succeeded" and "the box shows a picture" are different
questions, and nothing in this file has ever asked the second one — the restore
walks away the moment the verb returns, so every way the box can end up dark
looks identical to success in the log.

So measure it. After the hand-back a detached watcher polls
`detect_active_session()`, whose `None` means no compositor of our uid is running
at all — exactly the symptom. If the box is still dark 25 s later it climbs a
ladder of remedies, each measured on both images (Bazzite 44.20260818, Nobara
f44, 2026-08-22):

1. STOP the autologin unit. Its login session's script is parked on
   `systemctl --user --wait start <unit>` on both images, so a stop releases that
   wait, the session exits, and `Relogin=true` logs back in — starting the unit
   inside a session with a seat. `stop`, not `restart`: a restart does NOT
   release the parked waiter (measured), which is why it cannot rescue a box the
   ordinary restart already failed to bring back.
2. Restart the display manager — what the pre-0.31.0 takeover did on every
   disconnect, and proven on the Bazzite VM to return the box to game mode.
3. `PUNKTFUNK_RECOVER_SESSION_CMD`, then an ERROR naming the command a human has
   to run.

Detached, and that is load-bearing: the restore holds `RESTORE_FLIGHT`, which a
reconnecting client must take before it can re-take the box, so watching for up
to a minute while holding it would put that wait in front of every reconnect.
The watcher also stands down the instant `takeover_live()` says a new takeover
armed — the box belongs to that stream now, and a remedy fired into it would be
a fresh bug. It runs after `clear_takeover()` so that check means "a client
reconnected" and not "our own takeover has not been filed yet".

Skipped on the shutdown path: `restore_takeover_now` runs inside `native.rs`'s
20 s `SHUTDOWN_RESTORE_GRACE`, and spending that grace watching would cost the
hand-back rather than check it. What covers a shutdown that left the box dark is
the next host start — which this commit also makes true.

## A crashed host left the box's game mode asleep, provably

`restore_takeover_on_startup` sweeps a leftover idle drop-in off the box and logs
that the box's "own Game Mode session would have started and then done nothing".
Removing the FILE does not touch the unit RUNNING under it: its `ExecStart` is
still the sleep, so it sits `active` drawing nothing. Nothing below that sweep
restarts it either — the takeover file may be absent, unparseable, or fail
`takeover_state_is_live`, and all three exits leave the box on a dark panel with
its game mode "running". Any host killed mid-takeover (SIGKILL, OOM, a yanked
update) lands exactly there, and it survives until someone reboots.

`hand_back_idled_units_after_crash` restarts those units, gated on the box
actually being dark so a user already in game mode or on a desktop is never
bounced, and only for ACTIVE instances — under a just-removed idle drop-in,
active means "running the sleep".

## Not changed

The `restart` verb on the ordinary restore path. It works on both distros
(measured), and 0.31.0 chose it deliberately for the idled unit. The `stop` idea
survives only as escalation rung 1, where it runs after the proven path has
already failed.

`listed_autologin_units` is factored out of `stop_autologin_sessions` so both
callers share it, and its column parsing — which decides whether a live gaming
session can be told from a dead leftover — finally has a test against real
`--plain` output from both images.
2026-08-23 01:02:27 +02:00