fix(native/session): make the stop flag enforceable so a session can't outlive its client

Native sessions could survive long after the client was gone, in two
independent ways.

HOST: `stop` was only advisory. The one thing that ends a session is the
stream thread returning — every teardown (conn.close, the joins, and the
RAII drops of the session permit, admission entry and stream marker)
sits after that await, and nothing forced the issue. The encode loop
checks `stop` between iterations, so any unbounded call INSIDE one never
reaches the check, and one stuck syscall became a permanent zombie: it
held its semaphore slot (four of those and the host stops accepting
QUIC entirely), its admission entry (a later client gets "host busy"
forever), and not even the console's Stop button could clear it — that
button sets this same flag.

  * Bound the wait: once the session has been told to stop, the thread
    gets STREAM_STOP_GRACE (90 s, well past the 40 s capture-rebuild
    budget) to return, then teardown runs anyway. The thread is detached,
    not killed — Rust can't cancel a blocking thread — so it keeps its
    capturer/encoder until the stuck call returns, but the session's slot
    and admission entry come back and the host keeps serving. It logs at
    ERROR as the host wedge it is.
  * Bound the audio/input joins too — the last unbounded await in
    teardown.
  * Take the session permit AFTER the QUIC handshake instead of before
    `accept()`, so a host at its concurrency cap still accepts and the
    waiting client sees a live path instead of a silent dial timeout.
  * Bound the compositor helpers that caused the wedge in the first
    place: new pf-vdisplay `proc::{status_within, output_within}` kill a
    child that outlives its budget. `kscreen-doctor` is a Wayland client
    of the very compositor it configures, so against a wedged KWin it
    never returned; same for systemctl/dbus against a stuck session bus.

CLIENTS: the connection was never closed, so the host was right to keep
the session — it still had a live, keep-alive-answering peer.

  * Android: backgrounding did no teardown at all, and Android doesn't
    suspend the process, so the worker kept answering keep-alives until
    the OS reclaimed it (on a TV box, never). End the session on ON_STOP,
    via the existing onDispose path; a plain close, not a quit, so the
    host lingers the display for a fast return.
  * Apple: the .background arm was iOS-only AND gated on an opt-in that
    defaults off, so backgrounding did nothing — while the `audio`
    background mode kept the app (and its connection) alive indefinitely.
    Act unconditionally, and cover tvOS.
  * Core: `conn.close()` only queues the frame, and run_pump is the body
    of a block_on whose runtime is dropped the instant it returns, so the
    driver could never put it on the wire — a deliberate quit reached the
    host as silence (8 s idle timeout, no quit code, and the linger meant
    for an unwanted disconnect). Carry the endpoint out of the handshake
    and flush with wait_idle(), the same discipline the pairing and probe
    paths already use.

Linux check/clippy/tests green: 262 host, 71 pf-vdisplay (incl. new
bounded-process tests), 231 core.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-24 02:11:49 +02:00
co-authored by Claude Fable 5
parent 41fa25c440
commit ac3dc4323f
9 changed files with 338 additions and 84 deletions
+8
View File
@@ -36,6 +36,7 @@ pub(super) async fn run_pump(args: WorkerArgs) {
};
let handshake::HandshakeOut {
conn,
ep,
session,
ctrl_send,
ctrl_recv,
@@ -192,4 +193,11 @@ pub(super) async fn run_pump(args: WorkerArgs) {
0
};
conn.close(close_code.into(), b"client closed");
// Flush the CONNECTION_CLOSE before the runtime is dropped (the same discipline as the pairing
// + probe paths). `close` only queues the frame — the endpoint driver puts it on the wire, and
// this fn is the body of a `block_on` whose runtime is dropped the instant it returns, so
// without this the driver could simply never be polled again. The host then saw a deliberate
// quit as silence: no `QUIT_CLOSE_CODE`, an 8 s idle timeout, and the keep-alive linger meant
// for an UNWANTED disconnect. Bounded — a host already gone must not delay the client's exit.
let _ = tokio::time::timeout(std::time::Duration::from_millis(300), ep.wait_idle()).await;
}
@@ -9,6 +9,12 @@ use super::*;
/// Everything [`run_pump`](super::run_pump) needs from a successful connect + handshake.
pub(super) struct HandshakeOut {
pub(super) conn: quinn::Connection,
/// The dialing endpoint, kept alive for the session so the pump can FLUSH its
/// `CONNECTION_CLOSE` before the runtime is dropped ([`super::run_pump`]). Dropping it here
/// left nothing to drive the close onto the wire, so a deliberate quit reached the host as
/// silence — an 8 s idle timeout with no quit code, which the host reads as an unwanted
/// disconnect and lingers the display for.
pub(super) ep: quinn::Endpoint,
pub(super) session: Session,
pub(super) ctrl_send: quinn::SendStream,
pub(super) ctrl_recv: io::MsgReader,
@@ -238,6 +244,7 @@ pub(super) async fn connect_and_handshake(args: &WorkerArgs) -> Result<Handshake
match handshake.await {
Ok((session, send, recv, negotiated, host_caps)) => Ok(HandshakeOut {
conn,
ep,
session,
ctrl_send: send,
ctrl_recv: recv,