fix(native/session): make the stop flag enforceable so a session can't outlive its client

Native sessions could survive long after the client was gone, in two
independent ways.

HOST: `stop` was only advisory. The one thing that ends a session is the
stream thread returning — every teardown (conn.close, the joins, and the
RAII drops of the session permit, admission entry and stream marker)
sits after that await, and nothing forced the issue. The encode loop
checks `stop` between iterations, so any unbounded call INSIDE one never
reaches the check, and one stuck syscall became a permanent zombie: it
held its semaphore slot (four of those and the host stops accepting
QUIC entirely), its admission entry (a later client gets "host busy"
forever), and not even the console's Stop button could clear it — that
button sets this same flag.

  * Bound the wait: once the session has been told to stop, the thread
    gets STREAM_STOP_GRACE (90 s, well past the 40 s capture-rebuild
    budget) to return, then teardown runs anyway. The thread is detached,
    not killed — Rust can't cancel a blocking thread — so it keeps its
    capturer/encoder until the stuck call returns, but the session's slot
    and admission entry come back and the host keeps serving. It logs at
    ERROR as the host wedge it is.
  * Bound the audio/input joins too — the last unbounded await in
    teardown.
  * Take the session permit AFTER the QUIC handshake instead of before
    `accept()`, so a host at its concurrency cap still accepts and the
    waiting client sees a live path instead of a silent dial timeout.
  * Bound the compositor helpers that caused the wedge in the first
    place: new pf-vdisplay `proc::{status_within, output_within}` kill a
    child that outlives its budget. `kscreen-doctor` is a Wayland client
    of the very compositor it configures, so against a wedged KWin it
    never returned; same for systemctl/dbus against a stuck session bus.

CLIENTS: the connection was never closed, so the host was right to keep
the session — it still had a live, keep-alive-answering peer.

  * Android: backgrounding did no teardown at all, and Android doesn't
    suspend the process, so the worker kept answering keep-alives until
    the OS reclaimed it (on a TV box, never). End the session on ON_STOP,
    via the existing onDispose path; a plain close, not a quit, so the
    host lingers the display for a fast return.
  * Apple: the .background arm was iOS-only AND gated on an opt-in that
    defaults off, so backgrounding did nothing — while the `audio`
    background mode kept the app (and its connection) alive indefinitely.
    Act unconditionally, and cover tvOS.
  * Core: `conn.close()` only queues the frame, and run_pump is the body
    of a block_on whose runtime is dropped the instant it returns, so the
    driver could never put it on the wire — a deliberate quit reached the
    host as silence (8 s idle timeout, no quit code, and the linger meant
    for an unwanted disconnect). Carry the endpoint out of the handshake
    and flush with wait_idle(), the same discipline the pairing and probe
    paths already use.

Linux check/clippy/tests green: 262 host, 71 pf-vdisplay (incl. new
bounded-process tests), 231 core.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-24 02:11:49 +02:00
co-authored by Claude Fable 5
parent 41fa25c440
commit ac3dc4323f
9 changed files with 338 additions and 84 deletions
+5
View File
@@ -58,6 +58,11 @@ pub(crate) fn emit_display_event(ev: DisplayEvent) {
pub(crate) mod backend;
pub use backend::{DisplayOwnership, VirtualDisplay, VirtualOutput};
/// Time-bounded child-process helpers — every compositor query shells out, and an unbounded one
/// can wedge the calling (session) thread forever.
#[path = "vdisplay/proc.rs"]
pub(crate) mod proc;
/// Live-session detection + session-epoch + env retargeting (plan §W3).
#[path = "vdisplay/session.rs"]
pub(crate) mod session;
+41 -47
View File
@@ -149,11 +149,7 @@ impl VirtualDisplay for KwinDisplay {
return;
};
// kscreen-doctor position syntax: `output.<name-or-id>.position.<x>,<y>`.
let ok = std::process::Command::new("kscreen-doctor")
.arg(format!("output.{output}.position.{x},{y}"))
.status()
.map(|s| s.success())
.unwrap_or(false);
let ok = kscreen_ok(&[format!("output.{output}.position.{x},{y}")]);
if ok {
tracing::info!(output, x, y, "KWin: placed output in the desktop layout");
} else {
@@ -342,9 +338,7 @@ fn reenable_outputs(outputs: &[(String, String)]) {
.iter()
.map(|(name, _)| format!("output.{name}.enable"))
.collect();
let _ = std::process::Command::new("kscreen-doctor")
.args(&enable_args)
.status();
let _ = kscreen_ok(&enable_args);
// THEN re-assert each captured mode, best-effort — a bare re-enable lets KWin fall back to the
// EDID-preferred mode (a 120 Hz panel returns at ~60 Hz); this restores the exact refresh. The
// output is enabled now, so the mode set is valid; a rejected mode just leaves KWin's default.
@@ -354,9 +348,7 @@ fn reenable_outputs(outputs: &[(String, String)]) {
.map(|(name, mode)| format!("output.{name}.mode.{mode}"))
.collect();
if !mode_args.is_empty() {
let _ = std::process::Command::new("kscreen-doctor")
.args(&mode_args)
.status();
let _ = kscreen_ok(&mode_args);
}
std::thread::sleep(Duration::from_millis(200));
tracing::info!(reenabled = ?outputs, "KWin: restored the physical/bootstrap outputs at their captured modes (group empty)");
@@ -404,13 +396,37 @@ fn resolve_kscreen_addr(name: &str, w: u32, h: u32) -> String {
fallback
}
/// Budget for one `kscreen-doctor` call.
///
/// It is a Wayland client of the very compositor it configures, so against a wedged KWin it blocks
/// in its own connect and never returns — and these calls run on the session's stream thread, whose
/// only way to end a session is to return. Generous next to a healthy call (tens of ms).
const KSCREEN_BUDGET: Duration = Duration::from_secs(5);
/// `kscreen-doctor <args>` run for its exit status, bounded by [`KSCREEN_BUDGET`]. A timeout reads
/// as a failed apply — the same best-effort path a rejected argument already takes.
fn kscreen_ok(args: &[String]) -> bool {
crate::proc::status_within(
std::process::Command::new("kscreen-doctor").args(args),
KSCREEN_BUDGET,
)
.map(|s| s.success())
.unwrap_or(false)
}
/// `kscreen-doctor -j` stdout, bounded by [`KSCREEN_BUDGET`]; `None` on any failure.
fn kscreen_json_bytes() -> Option<Vec<u8>> {
crate::proc::output_within(
std::process::Command::new("kscreen-doctor").arg("-j"),
KSCREEN_BUDGET,
)
.ok()
.map(|o| o.stdout)
}
/// `kscreen-doctor -j` parsed, `None` on any failure.
fn kscreen_json() -> Option<serde_json::Value> {
let out = std::process::Command::new("kscreen-doctor")
.arg("-j")
.output()
.ok()?;
serde_json::from_slice(&out.stdout).ok()
serde_json::from_slice(&kscreen_json_bytes()?).ok()
}
/// The `(width, height)` of an output's CURRENT mode from its `kscreen-doctor -j` entry.
@@ -545,13 +561,7 @@ fn pick_custom_mode<'a>(
fn set_custom_refresh(width: u32, height: u32, hz: u32, output: &str) -> Option<(u32, u32, u32)> {
let output = output.to_string();
let mhz = hz.saturating_mul(1000);
let run = |arg: String| {
std::process::Command::new("kscreen-doctor")
.arg(arg)
.status()
.map(|s| s.success())
.unwrap_or(false)
};
let run = |arg: String| kscreen_ok(&[arg]);
// Install the mode only if the output doesn't already carry a usable one: kscreen-doctor
// APPENDS to the output's custom-mode list and KWin PERSISTS that list per output name
// (`kwinoutputconfig.json`, which is why the same per-slot name is reused across sessions) — so
@@ -710,14 +720,11 @@ fn output_current_mode_spec(o: &serde_json::Value) -> Option<String> {
/// session's own output — so a 2nd `exclusive` session (with a distinct per-slot name) never disables
/// the 1st session's live output. Parsed from `kscreen-doctor -j` (same source as [`read_active_mode`]).
fn other_enabled_outputs() -> Vec<(String, String)> {
let out = match std::process::Command::new("kscreen-doctor")
.arg("-j")
.output()
{
Ok(o) => o,
Err(_) => return Vec::new(),
let out = match kscreen_json_bytes() {
Some(o) => o,
None => return Vec::new(),
};
let doc: serde_json::Value = match serde_json::from_slice(&out.stdout) {
let doc: serde_json::Value = match serde_json::from_slice(&out) {
Ok(d) => d,
Err(_) => return Vec::new(),
};
@@ -746,13 +753,10 @@ fn other_enabled_outputs() -> Vec<(String, String)> {
/// then sets itself primary — the pre-group behavior). Recent kscreen marks the primary with
/// `"priority": 1`; older builds used a `"primary": true` bool — accept either.
fn a_managed_output_is_primary() -> bool {
let Ok(out) = std::process::Command::new("kscreen-doctor")
.arg("-j")
.output()
else {
let Some(out) = kscreen_json_bytes() else {
return false;
};
let Ok(doc) = serde_json::from_slice::<serde_json::Value>(&out.stdout) else {
let Ok(doc) = serde_json::from_slice::<serde_json::Value>(&out) else {
return false;
};
doc.get("outputs")
@@ -777,13 +781,7 @@ fn a_managed_output_is_primary() -> bool {
/// the keepalive to re-enable on teardown. Best-effort: on failure, streaming continues (just possibly
/// showing only the wallpaper) rather than failing the session.
fn apply_virtual_primary(ours: &str) -> Vec<(String, String)> {
let kscreen = |args: &[String]| {
std::process::Command::new("kscreen-doctor")
.args(args)
.status()
.map(|s| s.success())
.unwrap_or(false)
};
let kscreen = |args: &[String]| kscreen_ok(args);
// First-slot-wins (§6.1): only grab primary if no managed group member is primary yet — so a 2nd
// exclusive session joins as a secondary monitor of the shared desktop instead of stealing the
// shell off the 1st session's output. KWin usually then re-homes the desktop + disables the
@@ -815,11 +813,7 @@ fn apply_virtual_primary(ours: &str) -> Vec<(String, String)> {
/// (don't disable the bootstrap/physical) — so the shell re-homes onto the streamed surface while a
/// physical screen stays usable. Nothing to restore on teardown (we disabled nothing).
fn apply_virtual_primary_only(ours: &str) {
let ok = std::process::Command::new("kscreen-doctor")
.arg(format!("output.{ours}.primary"))
.status()
.map(|s| s.success())
.unwrap_or(false);
let ok = kscreen_ok(&[format!("output.{ours}.primary")]);
if ok {
tracing::info!("KWin: streamed output set primary (physical outputs kept)");
} else {
+114
View File
@@ -0,0 +1,114 @@
//! Time-bounded child-process helpers.
//!
//! Every compositor query this crate makes shells out to a helper (`kscreen-doctor`, `systemctl`,
//! `pw-dump`, …), and most of them are *clients of the very thing being diagnosed*: `kscreen-doctor`
//! is a Wayland client, so against a wedged KWin it blocks in its own connect and **never returns**.
//! `Command::status()` / `Command::output()` have no timeout, so one hung helper pinned the calling
//! thread forever — and on the host that thread is the session's stream thread, whose only way to
//! end a session is to return. A stuck query therefore became a permanently stuck session.
//!
//! These wrappers bound the wait: poll for exit until the budget runs out, then kill the child and
//! report [`std::io::ErrorKind::TimedOut`], so callers see a plain "the helper failed" error and
//! take their existing failure path instead of hanging.
use std::io::{Error, ErrorKind, Result};
use std::process::{Command, ExitStatus, Output};
use std::time::{Duration, Instant};
/// Poll interval while waiting for a child to exit. Short enough that a fast helper (the normal
/// case — `kscreen-doctor` answers in tens of ms) isn't measurably delayed.
const POLL: Duration = Duration::from_millis(20);
/// Run `cmd` to completion, killing it if it outlives `budget`.
///
/// Stdout/stderr are left as the caller configured them (inherited by default), so this is for
/// commands run for their exit status alone — see [`output_within`] when the output is read.
pub(crate) fn status_within(cmd: &mut Command, budget: Duration) -> Result<ExitStatus> {
let mut child = cmd.spawn()?;
let deadline = Instant::now() + budget;
loop {
match child.try_wait()? {
Some(status) => return Ok(status),
None if Instant::now() >= deadline => {
let _ = child.kill();
let _ = child.wait(); // reap it — never leave a zombie behind
return Err(timed_out(cmd, budget));
}
None => std::thread::sleep(POLL),
}
}
}
/// Run `cmd` to completion and capture its stdout/stderr, killing it if it outlives `budget`.
///
/// The output is read only after the child has exited, so a helper that fills the pipe buffer and
/// stalls is caught by the budget rather than deadlocking the reader (these helpers emit at most a
/// few hundred KiB, well under any real pipe pressure).
pub(crate) fn output_within(cmd: &mut Command, budget: Duration) -> Result<Output> {
let mut child = cmd
.stdout(std::process::Stdio::piped())
.stderr(std::process::Stdio::piped())
.spawn()?;
let deadline = Instant::now() + budget;
loop {
match child.try_wait()? {
// Exited: `wait_with_output` now only drains already-buffered pipes.
Some(_) => return child.wait_with_output(),
None if Instant::now() >= deadline => {
let _ = child.kill();
let _ = child.wait();
return Err(timed_out(cmd, budget));
}
None => std::thread::sleep(POLL),
}
}
}
fn timed_out(cmd: &Command, budget: Duration) -> Error {
let program = cmd.get_program().to_string_lossy().to_string();
tracing::warn!(
program,
budget_ms = budget.as_millis() as u64,
"helper did not exit within its budget — killed it (a wedged compositor/session bus is the \
usual cause); treating it as a failed query"
);
Error::new(
ErrorKind::TimedOut,
format!("`{program}` did not exit within {budget:?}"),
)
}
#[cfg(test)]
mod tests {
use super::*;
/// A helper that never exits must be killed at the budget and reported as `TimedOut` — the
/// whole point of the module (an unbounded `status()` here is what wedged a whole session).
#[test]
fn a_hung_child_is_killed_at_the_budget() {
let started = Instant::now();
let err = status_within(Command::new("sleep").arg("30"), Duration::from_millis(150))
.expect_err("must time out");
assert_eq!(err.kind(), ErrorKind::TimedOut);
assert!(
started.elapsed() < Duration::from_secs(5),
"must return at its budget, not the child's lifetime (took {:?})",
started.elapsed()
);
}
/// The normal path is unaffected: a quick command still yields its status and its output.
#[test]
fn a_quick_child_returns_normally() {
let st = status_within(&mut Command::new("true"), Duration::from_secs(5)).expect("ran");
assert!(st.success());
let out = output_within(
Command::new("echo").arg("punktfunk"),
Duration::from_secs(5),
)
.expect("ran");
assert!(out.status.success());
assert_eq!(String::from_utf8_lossy(&out.stdout).trim(), "punktfunk");
}
}
+40 -19
View File
@@ -6,6 +6,15 @@
use super::*;
/// Budget for one `systemctl --user` / `dbus-update-activation-environment` call.
///
/// These talk to the session bus, and a bus that is itself restarting or wedged answers nothing —
/// unbounded, that pinned the caller (on the host, the session's stream thread) forever. A restart
/// of the portal units is the slowest legitimate case, hence the generous window; missing it just
/// means the portal env settles late, which the callers already treat as best-effort.
#[cfg(target_os = "linux")]
const SYSTEMD_BUDGET: std::time::Duration = std::time::Duration::from_secs(10);
/// The **session epoch** — bumped whenever session detection observes a different compositor
/// *instance*: an [`ActiveKind`] change, **or** a new compositor PID for the same kind (the
/// Desktop→Game→Desktop bounce that brings up a fresh KWin/gamescope with an unrelated node-id space).
@@ -86,9 +95,15 @@ pub fn observe_session_instance(active: &ActiveSession) {
/// via the next [`settle_desktop_portal`], so scrubbing on a bounce is harmless.)
#[cfg(target_os = "linux")]
fn scrub_desktop_manager_env() {
let _ = std::process::Command::new("systemctl")
.args(["--user", "unset-environment", "WAYLAND_DISPLAY", "DISPLAY"])
.status();
let _ = crate::proc::status_within(
std::process::Command::new("systemctl").args([
"--user",
"unset-environment",
"WAYLAND_DISPLAY",
"DISPLAY",
]),
SYSTEMD_BUDGET,
);
}
#[cfg(not(target_os = "linux"))]
@@ -499,40 +514,46 @@ pub fn settle_desktop_portal(chosen: Compositor) {
];
// Push our (correct) env into the systemd --user manager + the D-Bus activation environment so a
// re-activated portal/backend inherits the live session.
let _ = std::process::Command::new("systemctl")
.args(["--user", "import-environment"])
.args(VARS)
.status();
let _ = std::process::Command::new("dbus-update-activation-environment")
.arg("--systemd")
.args(VARS)
.status();
let _ = crate::proc::status_within(
std::process::Command::new("systemctl")
.args(["--user", "import-environment"])
.args(VARS),
SYSTEMD_BUDGET,
);
let _ = crate::proc::status_within(
std::process::Command::new("dbus-update-activation-environment")
.arg("--systemd")
.args(VARS),
SYSTEMD_BUDGET,
);
// KWin input goes through the xdg RemoteDesktop portal; the frontend routes RemoteDesktop to a
// backend by its OWN startup XDG_CURRENT_DESKTOP, so restart it (+ the KDE backend) to re-read
// the now-live session, then let it settle before the injector reopens against it.
if chosen == Compositor::Kwin {
let _ = std::process::Command::new("systemctl")
.args([
let _ = crate::proc::status_within(
std::process::Command::new("systemctl").args([
"--user",
"try-restart",
"xdg-desktop-portal-kde.service",
"xdg-desktop-portal.service",
])
.status();
]),
SYSTEMD_BUDGET,
);
std::thread::sleep(std::time::Duration::from_millis(600));
}
// Hyprland capture rides the xdg ScreenCast portal serviced by xdph (G5): on a mid-stream switch
// xdph may still hold the old session's Wayland/instance env, so restart it (+ the frontend) to
// re-read the now-live session, mirroring the KWin settling above.
if chosen == Compositor::Hyprland {
let _ = std::process::Command::new("systemctl")
.args([
let _ = crate::proc::status_within(
std::process::Command::new("systemctl").args([
"--user",
"try-restart",
"xdg-desktop-portal-hyprland.service",
"xdg-desktop-portal.service",
])
.status();
]),
SYSTEMD_BUDGET,
);
std::thread::sleep(std::time::Duration::from_millis(600));
}
tracing::info!(