Merge pull request 'Audio was late because nothing ever placed it against the picture' (#101) from audio/latency-overhaul into main
ci / web (push) Successful in 1m17s
apple / swift (push) Successful in 1m33s
ci / rust-arm64 (push) Successful in 1m35s
ci / bun-nix (push) Successful in 26s
ci / docs-site (push) Successful in 1m21s
windows-msix / package (x64, , x86_64-pc-windows-msvc, C:\t) (push) Successful in 3m24s
deb / build-publish-client-arm64 (push) Successful in 2m52s
docker / builders (--build-arg FEDORA_VERSION=44, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm, -f44) (push) Successful in 17s
docker / builders (ci/android-ci.Dockerfile, punktfunk-android-ci) (push) Successful in 8s
docker / builders (ci/arch-ci.Dockerfile, punktfunk-arch-ci) (push) Successful in 9s
docker / builders (ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 10s
docker / builders (ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 13s
docker / builders (ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 10s
windows / build (aarch64-pc-windows-msvc) (push) Successful in 1m49s
deb / build-publish-host (push) Successful in 4m41s
apple / screenshots (push) Successful in 5m50s
docker / apps (., web/Dockerfile, punktfunk-web) (push) Successful in 1m36s
docker / apps (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 1m31s
docker / builders-arm64cross (push) Successful in 11s
arch / build-publish (push) Successful in 8m51s
windows / build (x86_64-pc-windows-msvc) (push) Successful in 3m33s
android / android (push) Successful in 11m39s
ci / rust (push) Successful in 11m32s
docker / deploy-docs (push) Failing after 1m43s
windows-host / package (push) Failing after 2m16s
windows-host / canary-manifest (push) Skipped
windows-host / winget-source (push) Skipped
deb / build-publish (push) Successful in 5m3s
windows-msix / package (arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Successful in 3m12s
flatpak / build-publish (push) Successful in 8m42s
release / apple (push) Successful in 9m37s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 19m15s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 19m12s

Reviewed-on: #101
This commit was merged in pull request #101.
This commit is contained in:
2026-08-07 22:32:08 +00:00
30 changed files with 2021 additions and 85 deletions
+16
View File
@@ -160,6 +160,22 @@ jobs:
key: gradle-${{ hashFiles('clients/android/**/*.gradle.kts', 'clients/android/gradle/wrapper/gradle-wrapper.properties') }}
restore-keys: gradle-
# Clippy for the ANDROID target. Like the kit tests below, this was running NOWHERE: ci.yml
# lints `--workspace` on the host, where `clients/android/native` and every
# `#[cfg(target_os = "android")]` module elsewhere compile out, and this workflow only ever
# built. Discovered in 2026-08 with five lints already resident — code no gate had ever read.
#
# Placed BEFORE assembleDebug deliberately: a lint failure should cost the ~10 s the lint
# takes, not the full three-ABI build first. It shares sccache and the target dir with the
# build that follows, so the compile is not paid twice.
#
# The task lints arm64-v8a AND armeabi-v7a, and reuses the build task's exact cargo-ndk
# environment — see the long note on `registerCargoNdkClippy` in kit/build.gradle.kts for why
# both pointer widths are load-bearing and why the environment must not be duplicated here.
- name: Clippy (Android target, deny warnings)
working-directory: clients/android
run: ./gradlew :kit:cargoNdkClippy --stacktrace
# The kit's JVM unit tests — the pure parsers, migrations and feedback policies. They were
# running nowhere: this workflow only assembled, and android-screenshots.yml runs the :app
# module's tests, so nothing enforced :kit's. Cheap (a couple of seconds against an already
@@ -18,12 +18,13 @@ import kotlin.math.roundToInt
* The live stats overlay — the unified HUD (`design/stats-unification.md`): headline is
* `capture→displayed` tiled by `host+network` + `decode` + `display` when the platform delivered
* OnFrameRendered render callbacks this window (`dispValid`), falling back to the v1
* `capture→decoded` headline without the `display` term when it didn't. Reads the 33-double
* `capture→decoded` headline without the `display` term when it didn't. Reads the 35-double
* layout from [NativeBridge.nativeVideoStats] (that KDoc is the authoritative index list):
* `[fps, mbps, e2eP50Ms, e2eP95Ms, latValid, skew, w, h, hz, lostTotal, bitDepth, colorPrimaries,
* colorTransfer, chromaFormatIdc, hostNetP50Ms, decodeP50Ms, hostP50Ms, netP50Ms, lost, skipped,
* fec, frames, dispValid, displayP50Ms, e2eDispP50Ms, e2eDispP95Ms, paceP50Ms, latchP50Ms,
* presentsWindow, presenterActive, feedP50Ms, codecP50Ms, skippedOverflowWindow]`. Every read
* presentsWindow, presenterActive, feedP50Ms, codecP50Ms, skippedOverflowWindow, audioBufferMs,
* audioAvOffsetMs]`. Every read
* is length-guarded, so an older native lib simply omits the lines it can't feed.
*
* The shown `display` and `end-to-end` numbers EXCLUDE the OS present floor (see [osFloorMs]) at
@@ -44,7 +45,7 @@ import kotlin.math.roundToInt
* reliability counters (1821) when nonzero.
* - [StatsVerbosity.DETAILED] — also the decoder label, the video-feed descriptor (1013), the
* stage equation (14/15, split into `host + network` when the Phase-2 terms at 16/17 are nonzero),
* and the excluded-floor line when one was measured.
* the excluded-floor line when one was measured, and the audio plane's own latency (33/34).
* [StatsVerbosity.OFF] renders nothing. Older native layouts simply omit the lines they lack (the
* counter line falls back to the cumulative `lostTotal` at index 9 on a pre-window lib).
*/
@@ -178,10 +179,42 @@ internal fun StatsOverlay(
}
}
}
if (detailed) {
audioLine(s)?.let { statLine(it, Color.White) }
}
counterLine(s, lost)?.let { statLine(it, Color(0xFFFFB0B0)) }
}
}
/**
* The audio plane's own latency from the live gauges at 33/34 — `audio buffer 42 ms · a/v +18 ms`,
* the same wording the desktop HUD uses. `buffer` is how much decoded audio is queued ahead of the
* speaker; `a/v` is where that PUTS it relative to the picture (positive = audio behind). `null`
* before any audio has been queued (buffer 0 — audio off, or the ring not yet primed) and on an
* older native layout.
*
* Both terms, not just the depth: a deep ring on a jittery link is correct behaviour — the
* underrun-driven floor earned that buffer — and only the offset distinguishes it from a ring that
* is simply holding audio late. The offset term is dropped at zero, which is both "aligned" and
* "no measurement yet"; the depth alone is still the triage number, and it is the one that did not
* exist at all before (the plane published nothing any surface could render, so a "the audio delay
* is way too high" report had no instrument behind it).
*
* NOT shaved by [osFloorMs], unlike every video figure above. That shave is a reporting policy —
* metrics report what Punktfunk controls — but sound has to reach the ear when the light reaches
* the eye, so the sync loop aligns against the RAW capture→displayed time (see the native
* `DisplayTracker`) and this offset is stated in those same terms. Subtracting the floor here would
* report an alignment the listener is not getting.
*/
private fun audioLine(s: DoubleArray): String? {
if (s.size < 35) return null
val bufferMs = s[33].roundToInt()
if (bufferMs <= 0) return null
val avOffset = s[34].roundToInt()
val avTerm = if (avOffset != 0) " · a/v ${if (avOffset > 0) "+" else ""}$avOffset ms" else ""
return "audio buffer $bufferMs ms$avTerm"
}
/** One monospace HUD line — the shared type ramp so every tier's rows line up. */
@Composable
private fun statLine(text: String, color: Color) {
@@ -0,0 +1,94 @@
package io.unom.punktfunk
import androidx.activity.ComponentActivity
import androidx.compose.ui.test.junit4.createAndroidComposeRule
import androidx.compose.ui.test.onNodeWithText
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith
import org.robolectric.RobolectricTestRunner
import org.robolectric.annotation.Config
/**
* The stats HUD's audio line — `audio buffer N ms · a/v ±N ms`, from the live gauges at indexes
* 33/34 (`design/audio-latency-overhaul.md`).
*
* Worth pinning because the whole point of the overhaul's stats half is that the audio plane became
* OBSERVABLE. Before it, ring depth and A/V offset existed only as a log line, and on a device
* launched by a game launcher that goes to a pipe nobody can read — so the single number that
* identifies a deep ring was unobtainable on the exact device reporting the latency, and a field
* investigation ran to its conclusion without it. A measurement that never reaches a surface is
* indistinguishable from no measurement, which is what this asserts.
*
* `sdk = [36]` for the same reason as the screenshot tests: Robolectric ships android-all jars only
* up to API 36 while the app's compileSdk is 37.
*/
@RunWith(RobolectricTestRunner::class)
@Config(sdk = [36])
class StatsOverlayAudioTest {
@get:Rule
val compose = createAndroidComposeRule<ComponentActivity>()
/**
* A plausible 35-double window with the audio gauges dialled in. Everything before 33 is the
* DETAILED-renderable shape the ShotScenes fixture uses; only the last two matter here.
*/
private fun stats(bufferMs: Double, avOffsetMs: Double, size: Int = 35): DoubleArray {
val full = doubleArrayOf(
238.0, 921.4, 1.3, 2.1, 1.0, 1.0, 5120.0, 1440.0, 240.0, 2.0,
10.0, 9.0, 16.0, 1.0, 0.9, 0.4, 0.6, 0.3,
2.0, 1.0, 5.0, 238.0,
1.0, 0.5, 1.8, 2.6,
0.2, 0.3, 236.0, 1.0,
0.1, 0.3, 0.0,
bufferMs, avOffsetMs,
)
return full.copyOf(size)
}
private fun show(s: DoubleArray, verbosity: StatsVerbosity = StatsVerbosity.DETAILED) {
compose.setContent { StatsOverlay(s, verbosity = verbosity) }
}
@Test
fun detailedShowsDepthAndOffset() {
show(stats(bufferMs = 42.0, avOffsetMs = 18.0))
// Positive = audio playing BEHIND the picture, and the sign is explicit so a glance tells
// which way the loop still has to move.
compose.onNodeWithText("audio buffer 42 ms · a/v +18 ms").assertExists()
}
@Test
fun audioAheadOfThePictureReadsNegative() {
show(stats(bufferMs = 42.0, avOffsetMs = -12.0))
compose.onNodeWithText("audio buffer 42 ms · a/v -12 ms").assertExists()
}
/** Aligned (or not yet measured) drops the offset term; the depth alone is still the triage number. */
@Test
fun alignedShowsDepthAlone() {
show(stats(bufferMs = 42.0, avOffsetMs = 0.0))
compose.onNodeWithText("audio buffer 42 ms").assertExists()
}
/** Nothing queued (audio off, or the ring not yet primed) — the line has nothing to say. */
@Test
fun silentPlaneRendersNoLine() {
show(stats(bufferMs = 0.0, avOffsetMs = 0.0))
compose.onNodeWithText("audio buffer", substring = true).assertDoesNotExist()
}
/** The line is DETAILED-only, like every other per-stage figure. */
@Test
fun normalTierOmitsTheLine() {
show(stats(bufferMs = 42.0, avOffsetMs = 18.0), verbosity = StatsVerbosity.NORMAL)
compose.onNodeWithText("audio buffer", substring = true).assertDoesNotExist()
}
/** An older native lib emits 33 doubles; the overlay must omit the line, not index past the end. */
@Test
fun olderNativeLayoutOmitsTheLine() {
show(stats(bufferMs = 42.0, avOffsetMs = 18.0, size = 33))
compose.onNodeWithText("audio buffer", substring = true).assertDoesNotExist()
}
}
@@ -355,10 +355,12 @@ internal fun StreamScene(verbosity: StatsVerbosity = StatsVerbosity.DETAILED) {
Brush.linearGradient(listOf(Color(0xFF2A1E5C), Color(0xFF0E1B3D), Color(0xFF06122B))),
),
) {
// The full 26-double unified layout (design/stats-unification.md): [fps, mbps, e2eP50,
// e2eP95, latValid, skew, w, h, hz, lostTotal, bitDepth, colorPrimaries, colorTransfer,
// chromaFormatIdc, hostNetP50, decodeP50, hostP50, netP50, lost, skipped, fec, frames,
// dispValid, displayP50, e2eDispP50, e2eDispP95].
// The full 35-double unified layout — NativeBridge.nativeVideoStats' KDoc is the
// authoritative index list: [fps, mbps, e2eP50, e2eP95, latValid, skew, w, h, hz,
// lostTotal, bitDepth, colorPrimaries, colorTransfer, chromaFormatIdc, hostNetP50,
// decodeP50, hostP50, netP50, lost, skipped, fec, frames, dispValid, displayP50,
// e2eDispP50, e2eDispP95, paceP50, latchP50, presents, presenterActive, feedP50, codecP50,
// skippedOverflow, audioBufferMs, audioAvOffsetMs].
// 10/9/16/1 = a 10-bit BT.2020 PQ (HDR) 4:2:0 feed so the DETAILED HUD renders its
// video-feed line; the display stage is valid (dispValid 1) so the headline is the
// directly-measured capture→displayed pair, less the excluded OS present floor (the 0.3
@@ -376,6 +378,12 @@ internal fun StreamScene(verbosity: StatsVerbosity = StatsVerbosity.DETAILED) {
1.0, 0.5, 1.8, 2.6,
// Timeline-presenter split: pace + latch tile the display term; presents ≈ fps.
0.2, 0.3, 236.0, 1.0,
// The decode term's own split (feed + codec = 0.4), and no overflow — the one
// `skipped` above is benign newest-wins pacing, not a decoder falling behind.
0.1, 0.3, 0.0,
// The audio plane: a 28 ms ring placed 4 ms behind the picture — a converged sync
// loop, i.e. inside the deadband it deliberately leaves alone.
28.0, 4.0,
),
verbosity = verbosity,
decoderLabel = "c2.qti.hevc.decoder · low-latency",
+61 -19
View File
@@ -67,30 +67,37 @@ fun androidSdkDir(): String {
return "${System.getProperty("user.home")}/Library/Android/sdk"
}
// Every cargo-ndk invocation needs the same discovery environment, and they must not drift apart:
// a lint that ran against a different toolchain/sysroot than the build is a lint about a different
// program. Applied by both `registerCargoNdk` (build) and `registerCargoNdkClippy` (lint).
fun Exec.cargoNdkEnvironment() {
val sdk = androidSdkDir()
// A GUI Android Studio launch does not source the login shell, so make cargo, the NDK, and
// cmake (libopus builds via the cmake crate) discoverable explicitly — same as a bare CLI.
val cmakeBin = "$sdk/cmake/3.22.1/bin"
environment(
"PATH",
cargoBin + File.pathSeparator + cmakeBin + File.pathSeparator + System.getenv("PATH"),
)
environment("ANDROID_HOME", sdk)
environment("ANDROID_NDK_HOME", "$sdk/ndk/$ndkVer")
// CMake's built-in Android support (used by the cmake crate for libopus) finds the NDK via
// these, and uses Ninja (bundled next to the SDK cmake) since there's no `make`.
environment("ANDROID_NDK_ROOT", "$sdk/ndk/$ndkVer")
environment("ANDROID_NDK", "$sdk/ndk/$ndkVer")
environment("CMAKE_GENERATOR", "Ninja")
// audiopus_sys picks static-vs-dynamic by HOST not target — force the bundled static libopus
// (pure C) so the android .so links it instead of looking for the host's libopus.so.
environment("LIBOPUS_STATIC", "1")
environment("LIBOPUS_NO_PKG", "1")
}
fun registerCargoNdk(taskName: String, release: Boolean) =
tasks.register<Exec>(taskName) {
group = "rust"
description = "cargo-ndk build of punktfunk-client-android (${if (release) "release" else "debug"})"
workingDir = repoRoot
val sdk = androidSdkDir()
// A GUI Android Studio launch does not source the login shell, so make cargo, the NDK, and
// cmake (libopus builds via the cmake crate) discoverable explicitly — same as a bare CLI.
val cmakeBin = "$sdk/cmake/3.22.1/bin"
environment(
"PATH",
cargoBin + File.pathSeparator + cmakeBin + File.pathSeparator + System.getenv("PATH"),
)
environment("ANDROID_HOME", sdk)
environment("ANDROID_NDK_HOME", "$sdk/ndk/$ndkVer")
// CMake's built-in Android support (used by the cmake crate for libopus) finds the NDK via
// these, and uses Ninja (bundled next to the SDK cmake) since there's no `make`.
environment("ANDROID_NDK_ROOT", "$sdk/ndk/$ndkVer")
environment("ANDROID_NDK", "$sdk/ndk/$ndkVer")
environment("CMAKE_GENERATOR", "Ninja")
// audiopus_sys picks static-vs-dynamic by HOST not target — force the bundled static libopus
// (pure C) so the android .so links it instead of looking for the host's libopus.so.
environment("LIBOPUS_STATIC", "1")
environment("LIBOPUS_NO_PKG", "1")
cargoNdkEnvironment()
// Resolve cargo by ABSOLUTE path: Gradle's Exec resolves command[0] via the JVM's
// inherited PATH, NOT the environment("PATH", …) set above (that only reaches the spawned
// child). A GUI Android Studio launch (and any daemon it started) has no ~/.cargo/bin on
@@ -113,6 +120,41 @@ fun registerCargoNdk(taskName: String, release: Boolean) =
commandLine(cmd)
}
// ------------------------------------------------------------------------------------------------
// Lint the ANDROID target. `punktfunk-client-android` and every `#[cfg(target_os = "android")]`
// module elsewhere in the workspace were, until this task existed, **completely unlinted**: ci.yml
// runs `cargo clippy --workspace` on the HOST, where all of that code is compiled out, and this
// workflow only ever ran `build`. The gap was found in 2026-08 with five lints sitting in
// clients/android/native (two of them `unnecessary_cast`, which is exactly the class that decides
// whether a cast is redundant BY POINTER WIDTH).
//
// Both widths are linted, and that is the load-bearing part: arm64-v8a is 64-bit and armeabi-v7a is
// 32-bit, so a cast that is redundant on one can be required on the other. Linting only the primary
// ABI would license "fixes" that break the 32-bit build — the shipping ABI for the many 32-bit
// Google TV / Android TV boxes this client targets. x86_64 is deliberately omitted: it is
// emulator-only and shares its pointer width with arm64, so it costs a third of the job's lint time
// for no signal these two do not already carry.
//
// `--all-targets` for the same reason ci.yml spells it out: without it the `#[cfg(test)]` modules
// are never compiled, and un-compiled test code drifts silently.
fun registerCargoNdkClippy(taskName: String) =
tasks.register<Exec>(taskName) {
group = "verification"
description = "clippy (deny warnings) for punktfunk-client-android on both Android widths"
workingDir = repoRoot
cargoNdkEnvironment()
commandLine(
// Absolute cargo path for the same reason as the build task above.
"$cargoBin/cargo", "ndk",
"-t", "arm64-v8a", "-t", "armeabi-v7a",
"--platform", "28",
"clippy", "-p", "punktfunk-client-android", "--all-targets",
"--", "-D", "warnings",
)
}
val cargoNdkClippy = registerCargoNdkClippy("cargoNdkClippy")
// Post-link floor check: every undefined symbol in the built .so must exist in the API-28 stubs,
// else System.loadLibrary fails on devices at the minSdk floor (see the script header for the
// 0.9.0 incident this guards against). Runs right after its cargo-ndk task; the APK build depends
@@ -264,12 +264,12 @@ object NativeBridge {
/**
* Drain ~1 s of live decode stats for the on-stream HUD, or `null` when no decode thread runs.
* Returns 33 doubles (unified stats spec, `design/stats-unification.md`):
* Returns 35 doubles (unified stats spec, `design/stats-unification.md`):
* `[fps, mbps, e2eP50Ms, e2eP95Ms, latValid, skewCorrected, width, height, refreshHz, framesLost,
* bitDepth, colorPrimaries, colorTransfer, chromaFormatIdc, hostNetP50Ms, decodeP50Ms, hostP50Ms,
* netP50Ms, lostWindow, skippedWindow, fecWindow, framesWindow, dispValid, displayP50Ms,
* e2eDispP50Ms, e2eDispP95Ms, paceP50Ms, latchP50Ms, presentsWindow, presenterActive,
* feedP50Ms, codecP50Ms, skippedOverflowWindow]`
* feedP50Ms, codecP50Ms, skippedOverflowWindow, audioBufferMs, audioAvOffsetMs]`
* (the flags are 1.0/0.0; indexes 2/3 are the end-to-end capture→decoded headline; 1013
* describe the negotiated video feed — bit depth 8/10, CICP primaries/transfer, and the HEVC
* chroma_format_idc 1=4:2:0 / 3=4:4:4; 14/15 are the stage p50s tiling the headline —
@@ -285,7 +285,10 @@ object NativeBridge {
* the window's on-glass confirm count, and whether the presenter is active at all; 30/31
* split `decode` (15) the same way — `feed` = received→queued (hand-off + input-slot wait),
* `codec` = queued→decoded, the decoder's own time; 32 is the parked-AU overflow subset of
* `skipped` (19), i.e. the decoder falling behind rather than benign newest-wins pacing).
* `skipped` (19), i.e. the decoder falling behind rather than benign newest-wins pacing;
* 33/34 are the AUDIO plane — the playback ring's live depth in ms and the A/V sync loop's
* smoothed offset in ms, positive meaning audio plays BEHIND the picture. Those two are live
* gauges, not windowed samples, and the offset reads 0 until the loop has a video reference).
* Poll ~1 Hz; each call resets the measurement window.
*/
external fun nativeVideoStats(handle: Long): DoubleArray?
+106 -12
View File
@@ -20,6 +20,16 @@
//! (2) is now the SHARED `punktfunk_core::audio::JitterPolicy` at `JitterTuning::AAUDIO`, which also
//! fixed what this ring was missing: it had a hard cap but nothing that walked the depth back down,
//! so drift and arrival bursts raised latency permanently and Android settled on its ceiling.
//!
//! It is also **A/V synchronised** (`design/audio-latency-overhaul.md`): the decode thread reads the
//! host capture `pts_ns` every `AudioPacket` has always carried, compares where this frame will
//! actually play against where the picture it belongs with reached glass
//! (`decode::DisplayTracker` publishes that), and asks the ring for a depth that closes the gap.
//! Only ASKS — `JitterPolicy` clamps the request between its own underrun-driven floor and the hard
//! cap, so continuity outranks sync and a link whose jitter genuinely needs more buffer than the
//! picture is away keeps its buffer, with the residual reported on the HUD instead of taken out of
//! the listener's stream. With no video reference (below API 33 there are no render callbacks, so
//! nothing confirms a present) the target stays `None` and the ring behaves exactly as it did.
use ndk::audio::{
AudioCallbackResult, AudioContentType, AudioDirection, AudioFormat, AudioPerformanceMode,
@@ -34,6 +44,14 @@ use std::sync::mpsc::{sync_channel, Receiver, SyncSender, TrySendError};
use std::sync::Arc;
use std::time::Duration;
/// What one playback open attempt yields: the stream, plus both halves of the PCM hand-off — the
/// sender the decode thread fills and the receiver that returns drained buffers for refill.
///
/// Named rather than written inline because the closure's return type trips
/// `clippy::type_complexity`, which the Android target is now linted for (`:kit:cargoNdkClippy`)
/// after years of nothing checking it.
type OpenedPlayback = ndk::audio::Result<(AudioStream, SyncSender<Vec<f32>>, Receiver<Vec<f32>>)>;
const SAMPLE_RATE: i32 = 48_000;
/// Decoded-chunk hand-off depth: 64 × 5 ms = 320 ms slack (matches the core's AUDIO_QUEUE).
const RING_CHUNKS: usize = 64;
@@ -94,15 +112,45 @@ impl AudioDec {
/// Diagnostics — written by the decode thread + the realtime callback, logged periodically. The
/// audio analogue of the video `fed`/`rendered` counters (we can't "screenshot" sound).
///
/// The ring's DEPTH is not here: the A/V sync loop needs the same number in the same units, so it
/// is published once through [`punktfunk_core::audio::AudioSyncCell`] and read from there by the
/// log line below. One publisher, one reading — a second copy is a second thing to go stale.
#[derive(Default)]
struct Counters {
opus_decoded: AtomicU64, // Opus packets decoded OK (~200/s at 5 ms frames)
pcm_written: AtomicU64, // PCM frames copied out to AAudio (device clock is pulling)
underruns: AtomicU64, // callbacks that emitted silence (ring not primed / drained)
ring_depth: AtomicU64, // ring sample count at the last callback
target_ms: AtomicU64, // the policy's LIVE target depth (it grows on this device's underruns)
}
/// Whether the A/V sync loop runs this session. `false` leaves `JitterPolicy`'s sync target at
/// `None`, which reproduces the pre-overhaul ring behaviour exactly — the point of the hatch.
///
/// Two levers because Android has neither of the other clients' launch surfaces. `PUNKTFUNK_NO_AV_SYNC`
/// keeps the contract the desktop clients document (and works when the client is driven from a
/// shell), but an app started from the launcher inherits no such environment, so the one a field
/// tester can actually reach is the sysprop — `adb shell setprop debug.punktfunk.no_av_sync 1`,
/// no rebuild, exactly like `debug.punktfunk.presenter`. A loop that steers PLAYBACK has to be
/// bisectable on the device that reports the regression, not only on the bench.
fn av_sync_enabled() -> bool {
if matches!(
std::env::var("PUNKTFUNK_NO_AV_SYNC").as_deref(),
Ok("1") | Ok("true")
) {
return false;
}
let mut buf = [0u8; 92]; // PROP_VALUE_MAX
// SAFETY: __system_property_get with a valid name + PROP_VALUE_MAX buffer is always safe.
let n = unsafe {
libc::__system_property_get(
c"debug.punktfunk.no_av_sync".as_ptr(),
buf.as_mut_ptr().cast(),
)
};
!(n > 0 && matches!(&buf[..n as usize], b"1" | b"true"))
}
/// Owned by [`crate::session::SessionHandle`]: the live AAudio stream + the decode thread.
pub struct AudioPlayback {
_stream: AudioStream, // dropping it stops + closes the AAudio stream
@@ -127,15 +175,15 @@ impl AudioPlayback {
// Worst transient the ring can hold before the policy trims it.
let hard_cap_max = tuning.hard_cap_ms as usize * ms;
let counters = Arc::new(Counters::default());
// The A/V sync hand-off: the realtime callback owns the ring (so it publishes the depth and
// consumes the target), the decode thread owns the timestamps (so it computes the target).
// Two atomics, because the callback must not block on the thread that decodes Opus.
let sync: Arc<punktfunk_core::audio::AudioSyncCell> = Arc::default();
// One open attempt at a given sharing mode. Everything the realtime callback captures
// (channels, ring, prime state) is rebuilt per attempt — `open_stream` consumes the builder
// AND the callback, so nothing survives a failed try to reuse.
let try_open = |sharing: AudioSharingMode| -> ndk::audio::Result<(
AudioStream,
SyncSender<Vec<f32>>,
Receiver<Vec<f32>>,
)> {
let try_open = |sharing: AudioSharingMode| -> OpenedPlayback {
let (tx, rx) = sync_channel::<Vec<f32>>(RING_CHUNKS);
// Recycle free-list: drained PCM buffers go BACK to the decode thread to be refilled, so
// the realtime callback never frees heap (Android's Scudo allocator has unbounded free()
@@ -146,6 +194,7 @@ impl AudioPlayback {
// Realtime consumer state, owned by the callback (FnMut) — no lock: AAudio calls it from
// a single high-priority thread, and the decode thread only touches `tx`/`free_rx`.
let cb_counters = counters.clone();
let cb_sync = sync.clone();
// Pre-reserve the ring so `extend` never reallocates on the realtime thread. Worst
// transient before the trim below = the hard cap plus one full channel of 5 ms (480-f32)
// frames — the punktfunk protocol always sends 5 ms Opus frames (host `audio_thread`); a
@@ -171,6 +220,13 @@ impl AudioPlayback {
ring.extend(chunk.drain(..));
let _ = free_tx.try_send(chunk);
}
// A/V sync: take whatever depth the decode thread's sync loop last asked for, and
// publish where the ring actually is so it can measure the result. The policy
// clamps the request between its own underrun floor and the hard cap — continuity
// outranks sync, always (see `JitterPolicy::set_sync_target`). Read AFTER the
// drain, so the depth is everything a frame queued right now must wait behind.
policy.set_sync_target(cb_sync.target());
cb_sync.publish_depth(ring.len());
// Jitter buffer: the shared policy decides prime/silence, trims a burst, and —
// new here — sheds ONE crossfaded 5 ms frame when the depth average has sat above
// target long enough to be drift rather than jitter. Without that shed this ring
@@ -201,9 +257,6 @@ impl AudioPlayback {
// No-op while un-primed, so a deliberate priming silence is never counted as an
// underrun (which would otherwise drive the adaptive floor up for no reason).
policy.note_read(ran_short);
cb_counters
.ring_depth
.store(ring.len() as u64, Ordering::Relaxed);
cb_counters
.target_ms
.store(policy.target_ms() as u64, Ordering::Relaxed);
@@ -303,7 +356,7 @@ impl AudioPlayback {
let sd = shutdown.clone();
let join = std::thread::Builder::new()
.name("pf-audio".into())
.spawn(move || decode_loop(client, tx, free_rx, sd, counters, channels))
.spawn(move || decode_loop(client, tx, free_rx, sd, counters, channels, sync))
.ok();
Some(AudioPlayback {
@@ -334,6 +387,7 @@ fn decode_loop(
shutdown: Arc<AtomicBool>,
counters: Arc<Counters>,
channels: usize,
sync: Arc<punktfunk_core::audio::AudioSyncCell>,
) {
// Fold this Opus→AAudio thread into the client's hot-thread set so the ADPF session the decode
// thread opens also keeps audio decode on a fast core (registered before the video pump's first
@@ -354,9 +408,44 @@ fn decode_loop(
let mut window_peak = 0f32; // loudest |sample| since the last log — tells a tone from silence
let mut gaps = punktfunk_core::audio::AudioGapTracker::new();
let mut frame_samples = 0usize; // per-channel samples of the last decoded frame — the PLC unit
// A/V sync (audio latency overhaul). This thread is the only place holding all three
// ingredients at once: the packet's host capture `pts_ns`, the ring depth (via the sync cell)
// and the video plane's end-to-end figure. `pts_ns` arrived in every `AudioPacket` and was
// dropped on the floor here for the plane's whole existence, which is why audio ran at whatever
// depth its jitter ring settled at with nothing ever placing it against the picture.
let av_sync_enabled = av_sync_enabled();
let mut av = punktfunk_core::audio::AvSync::new(channels as u8);
let video_e2e = client.video_e2e_shared();
let av_offset_out = client.audio_av_offset_shared();
let buffer_ms_out = client.audio_buffer_ms_shared();
if !av_sync_enabled {
log::info!("audio: A/V sync disabled (PUNKTFUNK_NO_AV_SYNC / debug.punktfunk.no_av_sync)");
}
'pump: while !shutdown.load(Ordering::Relaxed) {
match client.next_audio(Duration::from_millis(5)) {
Ok(pkt) => {
// Place this frame against the picture it belongs with, BEFORE it is queued:
// `buffered_ahead` is everything that must still play first, so the depth read here
// is exactly what delays it.
let depth = sync.depth();
// Published unconditionally — the ring's depth is worth seeing even with sync off,
// and it is what makes a "the audio delay is way too high" report triageable at all.
buffer_ms_out.store((depth / ms.max(1)) as u32, Ordering::Relaxed);
if av_sync_enabled {
let ve2e = video_e2e.load(Ordering::Relaxed);
av.observe(punktfunk_core::audio::AvSyncObservation {
pts_ns: pkt.pts_ns,
now_local_ns: punktfunk_core::client::now_realtime_ns(),
clock_offset_ns: client.clock_offset_now_ns(),
buffered_ahead: depth,
// 0 = nothing confirmed on the glass yet (no render callback below API 33,
// or the stream has not presented a frame); no reference, no correction.
video_e2e_ns: (ve2e > 0).then_some(ve2e),
});
sync.set_target(av.desired_depth(depth));
av_offset_out.store(av.offset_ms() as i64, Ordering::Relaxed);
}
// Conceal lost packets (a seq gap) with libopus PLC before decoding the one that
// arrived: empty input synthesizes `frame_samples` of interpolation per missing
// packet — an inaudible fade instead of the click a hard gap makes in the ring.
@@ -404,12 +493,17 @@ fn decode_loop(
Err(TrySendError::Disconnected(_)) => break,
}
if count % 600 == 0 {
// `av_ms` is the sync loop's smoothed placement error (+ = audio behind
// the picture); 0 with sync off, or before it has a video reference.
// Logged next to the depth because a deep ring on a jittery link is
// correct and only the offset separates that from audio held late.
log::info!(
"audio: opus={count} pcm_frames={} underruns={} buffer_ms={} target_ms={} peak={window_peak:.3}",
"audio: opus={count} pcm_frames={} underruns={} buffer_ms={} target_ms={} av_ms={} peak={window_peak:.3}",
counters.pcm_written.load(Ordering::Relaxed),
counters.underruns.load(Ordering::Relaxed),
counters.ring_depth.load(Ordering::Relaxed) / ms.max(1) as u64,
(depth / ms.max(1)) as u64,
counters.target_ms.load(Ordering::Relaxed),
av.offset_ms(),
);
window_peak = 0.0;
}
@@ -204,7 +204,15 @@ pub(super) fn run_async(
// SurfaceFlinger's render timestamp. `render_cb` is the callback's leaked Arc refcount,
// reclaimed after the codec is dropped below.
let meter = Arc::new(PresentMeter::new());
let tracker = DisplayTracker::new(stats.clone(), clock_offset.clone(), meter.clone());
// The tracker also publishes each confirmed present's end-to-end into the shared cell the audio
// plane steers its jitter ring by (`design/audio-latency-overhaul.md`) — video is the master,
// and this is the only point that knows when a frame actually reached glass.
let tracker = DisplayTracker::new(
stats.clone(),
clock_offset.clone(),
client.video_e2e_shared(),
meter.clone(),
);
let render_cb = install_render_callback(&codec, &tracker);
// The timeline presenter (see `presenter.rs`): newest-wins / smoothing store, one-in-flight
+45 -9
View File
@@ -5,7 +5,7 @@ use ndk::media::media_codec::MediaCodec;
use ndk::native_window::NativeWindow;
use std::collections::VecDeque;
use std::ffi::c_void;
use std::sync::atomic::{AtomicI64, Ordering};
use std::sync::atomic::{AtomicI64, AtomicU64, Ordering};
use std::sync::{Arc, Mutex};
use super::latency::now_realtime_ns;
@@ -35,6 +35,16 @@ pub(super) struct DisplayTracker {
/// loaded per callback so mid-stream re-syncs apply. Holding the handle (not the client)
/// keeps the leaked render-callback refcount from pinning the whole session alive.
clock_offset: Arc<AtomicI64>,
/// Where the AUDIO plane reads the video leg it has to land with (ns) — `displayed +
/// clock_offset pts`, published on every confirmed present. Written here, read by
/// [`crate::audio`]'s sync loop; the two planes never touch each other directly (the presenter
/// must not know about audio, and the audio thread cannot see the glass).
///
/// Published RAW. The HUD shaves the OS present floor off its shown display / end-to-end
/// numbers (`StatsOverlay.osFloorMs` — metrics report what Punktfunk controls), but sound has
/// to reach the ear when the light reaches the eye, and a floor-shaved reference would place
/// audio a whole latch period early on every device. Presentation policy, not physics.
video_e2e: Arc<AtomicU64>,
/// Always-on latch/display accumulator for the presenter's 1 Hz `pf-present` line —
/// independent of the HUD gate, so a HUD-off A/B stays measurable from logcat.
meter: Arc<super::presenter::PresentMeter>,
@@ -48,11 +58,13 @@ impl DisplayTracker {
pub(super) fn new(
stats: Arc<crate::stats::VideoStats>,
clock_offset: Arc<AtomicI64>,
video_e2e: Arc<AtomicU64>,
meter: Arc<super::presenter::PresentMeter>,
) -> Arc<DisplayTracker> {
Arc::new(DisplayTracker {
stats,
clock_offset,
video_e2e,
meter,
rendered: Mutex::new(VecDeque::new()),
})
@@ -105,7 +117,14 @@ pub(super) fn install_render_callback(
}
let sym = libc::dlsym(lib, c"AMediaCodec_setOnFrameRenderedCallback".as_ptr());
if sym.is_null() {
log::info!("decode: no render callback on this API level (<33) — no display stage");
// No confirmed present ⇒ no `display` stage AND no reference for the audio plane's A/V
// sync, which then stays inert and leaves the ring exactly as it was. The release
// instant is NOT substituted: releases target a future vsync, so it runs a whole latch
// period (8-21 ms measured) ahead of glass — well outside the loop's deadband, i.e. it
// would place audio early on every frame while looking like it was working.
log::info!(
"decode: no render callback on this API level (<33) — no display stage, no A/V sync"
);
return None;
}
std::mem::transmute::<*mut c_void, SetOnFrameRenderedFn>(sym)
@@ -145,8 +164,10 @@ pub(super) unsafe fn release_render_callback(ud: *const DisplayTracker) {
/// between the frame rendering and the (batchable) callback delivery — to subtract against the
/// receipt/decode stamps and the host capture pts. Records the HUD's `displayed` point:
/// `end-to-end` = capture→displayed (skew-corrected) and `display` = decoded→displayed
/// (single-clock local). Panic-free by construction (poison-proof lock, saturating math) — an
/// unwind out of an `extern "C"` fn would abort the process.
/// (single-clock local) — and publishes that end-to-end figure for the audio plane to align
/// against, which is the only place in the client that knows when a frame truly reached glass.
/// Panic-free by construction (poison-proof lock, saturating math) — an unwind out of an
/// `extern "C"` fn would abort the process.
unsafe extern "C" fn on_frame_rendered(
_codec: *mut ndk_sys::AMediaCodec,
userdata: *mut c_void,
@@ -186,13 +207,28 @@ unsafe extern "C" fn on_frame_rendered(
let latch_us = paired.and_then(|(_, r)| clamp(displayed_ns - r));
// Always-on half: the presenter's pf-present line reads these with the HUD off.
t.meter.note_latch(latch_us);
if !t.stats.enabled() {
return; // HUD hidden — skip the skew math + the stats lock
}
// The glass-to-glass figure, computed ABOVE the HUD gate: the audio plane steers its ring by it
// (see `video_e2e`), and a sync loop that only worked while the overlay was up would be off on
// the exact devices that report latency — on a Deck-class report the overlay is precisely what
// the field cannot reach. The cost is one relaxed load and some integer arithmetic per confirmed
// present (≤ the panel rate); the stats LOCK stays behind the gate, which is what that
// early-return was really protecting.
let e2e_ns =
displayed_ns + t.clock_offset.load(Ordering::Relaxed) as i128 - pts_us as i128 * 1000;
let e2e_us = (e2e_ns > 0 && e2e_ns < 10_000_000_000).then_some((e2e_ns / 1000) as u64);
t.stats.note_displayed(e2e_us, display_us, latch_us);
// Same (0, 10 s) clamp as every other e2e sample — a vendor's first render callbacks can carry
// a garbage `system_nano`, and here that would step the audio ring rather than just a p95.
let e2e_valid = e2e_ns > 0 && e2e_ns < 10_000_000_000;
if e2e_valid {
t.video_e2e.store(e2e_ns as u64, Ordering::Relaxed);
}
if !t.stats.enabled() {
return; // HUD hidden — skip the stats lock
}
t.stats.note_displayed(
e2e_valid.then_some((e2e_ns / 1000) as u64),
display_us,
latch_us,
);
}
/// React to an output-format change by signalling the stream's HDR dataspace on the Surface (SDR
@@ -185,9 +185,12 @@ pub(super) fn run_sync(
// render = true are parked in the tracker; the OnFrameRendered callback pairs them with
// SurfaceFlinger's render timestamp. `render_cb` is the callback's leaked Arc refcount,
// reclaimed after the codec is dropped below.
// The `video_e2e` cell is the audio plane's alignment reference (see `DisplayTracker`): this
// legacy loop feeds it too, so A/V sync works with "Low-latency mode" off as well.
let tracker = DisplayTracker::new(
stats.clone(),
clock_offset.clone(),
client.video_e2e_shared(),
std::sync::Arc::new(super::presenter::PresentMeter::new()),
);
let render_cb = install_render_callback(&codec, &tracker);
+15 -2
View File
@@ -33,8 +33,21 @@ pub(super) fn now_monotonic_ns() -> i64 {
};
// SAFETY: `clock_gettime` with a valid out-pointer is an always-safe syscall.
unsafe { libc::clock_gettime(libc::CLOCK_MONOTONIC, &mut ts) };
// Explicit widening: timespec's fields are 32-bit on armv7 (time_t/c_long).
ts.tv_sec as i64 * 1_000_000_000 + ts.tv_nsec as i64
// Explicit widening: `timespec`'s fields are 32-bit on armv7 (`time_t`/`c_long`) and 64-bit on
// arm64, so these casts are REQUIRED on one shipping ABI and redundant on the other.
//
// `:kit:cargoNdkClippy` lints both widths, so it sees the redundant half and flags it; taking
// its advice would break the 32-bit build, which is the ABI for the many 32-bit Google TV /
// Android TV boxes this client targets. `i64::from`/`.into()` do not escape it either — they
// just trade `unnecessary_cast` for `useless_conversion` on the 64-bit side. So the cast stays
// and the lint is answered here rather than in whichever build breaks first.
#[allow(
clippy::unnecessary_cast,
reason = "required on 32-bit ABIs; redundant only on 64-bit"
)]
{
ts.tv_sec as i64 * 1_000_000_000 + ts.tv_nsec as i64
}
}
/// One upcoming frame timeline (API 33+ payload): when SurfaceFlinger expects to present the
+10 -7
View File
@@ -26,6 +26,15 @@ use std::sync::mpsc::{sync_channel, Receiver, RecvTimeoutError, SyncSender, TryS
use std::sync::Arc;
use std::time::{Duration, SystemTime, UNIX_EPOCH};
/// What one capture open attempt yields: the stream, plus both halves of the PCM hand-off — the
/// receiver the encode worker drains and the sender that returns emptied buffers for reuse. Note
/// the pair is the mirror image of [`crate::audio::OpenedPlayback`]'s: here the callback produces
/// and the worker consumes.
///
/// Named rather than written inline for the same reason as that one — `clippy::type_complexity`,
/// now that the Android target is actually linted (`:kit:cargoNdkClippy`).
type OpenedCapture = ndk::audio::Result<(AudioStream, Receiver<Vec<f32>>, SyncSender<Vec<f32>>)>;
const CHANNELS: usize = 1;
const SAMPLE_RATE: i32 = 48_000;
/// 10 ms per channel @ 48 kHz — half the desktop clients' 20 ms frame, trading a little Opus
@@ -84,13 +93,7 @@ impl MicCapture {
// One open attempt at a given sharing mode (same pattern as [`crate::audio`]: `open_stream`
// consumes the builder AND the callback, so each try rebuilds the channels it captures).
let try_open = |sharing: AudioSharingMode,
voice: bool|
-> ndk::audio::Result<(
AudioStream,
Receiver<Vec<f32>>,
SyncSender<Vec<f32>>,
)> {
let try_open = |sharing: AudioSharingMode, voice: bool| -> OpenedCapture {
let (tx, rx) = sync_channel::<Vec<f32>>(RING_CHUNKS);
// Recycle free-list, mirroring the playback path: the realtime capture callback must
// not touch the allocator (Android's Scudo has unbounded malloc/free tail latency — an
+2 -2
View File
@@ -408,8 +408,8 @@ pub(crate) unsafe fn self_test(fd: i32, seconds: i32, hz: i32) -> i32 {
frame.fill(0);
// Channels 2 and 3 are the voice coils; the speaker pair stays silent so a pass is
// unambiguously FELT rather than merely audible.
for c in 2..channels {
frame[c] = sample;
for slot in frame.iter_mut().take(channels).skip(2) {
*slot = sample;
}
}
if let Err(e) = playback.write_interleaved(&chunk) {
+16 -4
View File
@@ -177,12 +177,12 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeStopVideo(
}
/// `NativeBridge.nativeVideoStats(handle): DoubleArray?` — drain ~1 s of decode stats for the HUD
/// (unified stats spec, `design/stats-unification.md`). Returns 33 doubles
/// (unified stats spec, `design/stats-unification.md`). Returns 35 doubles
/// `[fps, mbps, e2eP50Ms, e2eP95Ms, latValid, skewCorrected, width, height, refreshHz, framesLost,
/// bitDepth, colorPrimaries, colorTransfer, chromaFormatIdc, hostNetP50Ms, decodeP50Ms, hostP50Ms,
/// netP50Ms, lostWindow, skippedWindow, fecWindow, framesWindow, dispValid, displayP50Ms,
/// e2eDispP50Ms, e2eDispP95Ms, paceP50Ms, latchP50Ms, presentsWindow, presenterActive,
/// feedP50Ms, codecP50Ms, skippedOverflowWindow]`
/// feedP50Ms, codecP50Ms, skippedOverflowWindow, audioBufferMs, audioAvOffsetMs]`
/// (the flags are 1.0/0.0; indexes 021 match the previous 22-double layout — 013 the original
/// 14-double one with the latency pair re-based to the end-to-end capture→decoded headline, 14/15
/// the stage p50s tiling it: `host+network` = capture→received, `decode` = received→decoded; 16/17
@@ -203,7 +203,10 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeStopVideo(
/// received→queued (hand-off + input-slot wait) at 30 and `codec` = queued→decoded (codec-pure,
/// from the AU's last piece) at 31, both 0.0 when no sample landed (sync loop); 32 is the
/// parked-AU overflow subset of the window's `skipped` at 19 (decoder fell behind, vs benign
/// newest-wins pacing)), or `null` when no decode thread is running.
/// newest-wins pacing); 33/34 are the AUDIO plane's latency — the playback ring's live depth in ms
/// and the A/V sync loop's smoothed offset in ms (positive = audio behind the picture) — both live
/// gauges rather than windowed samples, like the cumulative drop total at 9), or `null` when no
/// decode thread is running.
/// Poll ~1 Hz from the UI; each call
/// resets the measurement window. Not android-gated — pure `jni` + connector reads, so it links on
/// the host build too (Kotlin only ever calls it on device).
@@ -227,7 +230,7 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeVideoStats(
.drain(h.client.frames_dropped(), h.client.fec_recovered_shards());
let mode = h.client.mode();
let color = h.client.color;
let buf: [f64; 33] = [
let buf: [f64; 35] = [
snap.fps,
snap.mbps,
snap.e2e_p50_ms,
@@ -281,6 +284,15 @@ pub extern "system" fn Java_io_unom_punktfunk_kit_NativeBridge_nativeVideoStats(
snap.feed_p50_ms,
snap.codec_p50_ms,
snap.skipped_overflow as f64,
// The audio plane's own latency (`design/audio-latency-overhaul.md`): how much decoded
// audio is queued ahead of the speaker, and where the A/V sync loop measures that
// PUTS it relative to the picture (+ = audio behind). Both, because a deep ring on a
// jittery link is correct behaviour and only the offset tells that apart from audio
// simply held late. Live gauges written by the audio thread — before this the whole
// plane published nothing any surface could render, so a "the audio delay is way too
// high" report had no instrument behind it at all.
h.client.audio_buffer_ms() as f64,
h.client.audio_av_offset_ms() as f64,
];
let arr = match env.new_double_array(buf.len() as jsize) {
Ok(a) => a,
@@ -132,6 +132,17 @@ final class SessionModel: ObservableObject {
/// and under stage-1.
@Published var osFloorP50Ms = 0.0
@Published var osFloorValid = false
/// The AUDIO plane's latency, from the playback ring (`SessionAudio.Stats`): how much decoded
/// audio is queued ahead of the speaker, and where that PUTS it relative to the picture
/// (positive = audio behind). `audioValid` is false until playback runs.
///
/// Both numbers, never just the depth a deep ring on a jittery link is the adaptive floor
/// doing its job, and only the offset separates that from audio simply being held late. They
/// existed nowhere a surface could render them until now, which is why a field report of "the
/// audio delay seems way too high" was triaged all the way to a conclusion without them.
@Published var audioBufferMs = 0
@Published var audioAvOffsetMs = 0
@Published var audioValid = false
/// The floor-shaved values every HUD tier displays (raw floor, never below 0). Identical
/// to the raw values whenever no floor is measured.
@@ -662,6 +673,7 @@ final class SessionModel: ObservableObject {
displayValid = false
clientQueueValid = false
osFloorValid = false
audioValid = false
lostFrames = 0
lostPct = 0
mouseCaptured = false
@@ -736,7 +748,14 @@ final class SessionModel: ObservableObject {
micUID: settings.micUID,
micChannel: settings.micChannel,
micEnabled: settings.micEnabled,
echoCancel: settings.echoCancel)
echoCancel: settings.echoCancel,
// The A/V sync reference: `endToEnd` is captureon-glass, the one figure that says
// where the picture actually IS, and the audio ring steers its depth to land with it.
// The same meter object the presenter writes per presented frame, so audio reads the
// video plane's own measurement rather than a second estimate of it and under the
// stage-1 fallback presenter, which stamps nothing, it stays empty and the loop
// correctly declines to correct.
videoLatency: endToEnd)
self.audio = audio
// Gamepads: forward every controller GamepadManager selected each on its own wire pad
// index (a pin forwards only one, Automatic forwards all) and render the host's feedback
@@ -897,6 +916,15 @@ final class SessionModel: ObservableObject {
} else {
self.clientQueueValid = false
}
// The audio plane is a LEVEL, not a window: the ring's depth and the sync loop's
// smoothed offset are both current values, so they are read rather than drained.
if let a = self.audio?.stats {
self.audioBufferMs = a.bufferMS
self.audioAvOffsetMs = a.avOffsetMS
self.audioValid = true
} else {
self.audioValid = false
}
// Mirror the window to the unified log (see statsLog) one line per second,
// stages in ms, only while frames actually flowed. `fps` counts RECEIVED AUs;
// `presents` counts frames that reached glass (the display meter's sample count)
@@ -912,7 +940,12 @@ final class SessionModel: ObservableObject {
// the whole line (a cascade error that also mis-blames the float args).
format: "fps=%lld presents=%lld e2e_p50=%.1f e2e_p95=%.1f hostnet_p50=%.1f "
+ "decode_p50=%.1f display_p50=%.1f lost=%lld "
+ "floor_p50=%.1f display_adj=%.1f e2e_adj=%.1f queue_p50=%.1f",
+ "floor_p50=%.1f display_adj=%.1f e2e_adj=%.1f queue_p50=%.1f "
// Appended LAST, so every existing parser of this line is unaffected.
// In the log as well as on the HUD because the overlay is only up when
// someone thought to turn it on, and the reports that need these
// numbers arrive after the fact.
+ "audio_buffer=%lld audio_av_offset=%lld",
frames,
displayWindow?.count ?? 0,
self.endToEndValid ? self.endToEndP50Ms : -1,
@@ -924,7 +957,9 @@ final class SessionModel: ObservableObject {
self.osFloorValid ? self.osFloorP50Ms : -1,
self.displayValid ? self.displayAdjP50Ms : -1,
self.endToEndValid ? self.endToEndAdjP50Ms : -1,
self.clientQueueValid ? self.clientQueueP50Ms : -1)
self.clientQueueValid ? self.clientQueueP50Ms : -1,
self.audioValid ? self.audioBufferMs : -1,
self.audioValid ? self.audioAvOffsetMs : 0)
statsLog.info("\(line, privacy: .public)")
}
}
@@ -154,6 +154,28 @@ struct StreamHUDView: View {
.foregroundStyle(.secondary)
}
}
// The AUDIO plane's own latency (detailed tier). Deliberately OUTSIDE the video branch
// above: it is not a term of that equation audio is steered to MEET the video total,
// never summed into it and the depth is exactly as worth seeing under the stage-1
// fallback presenter, which measures no end-to-end at all.
//
// `buffer` is how much decoded audio is queued ahead of the speaker; `a/v` is where
// that puts it relative to the picture (+ = audio behind). Both, not just the depth: a
// deep ring on a jittery link is the adaptive floor doing its job, and only the offset
// distinguishes that from a ring holding audio late. Neither number was renderable
// anywhere before they lived in a periodic log line which is how a report of "the
// audio delay seems way too high" got triaged to a conclusion with no instrument.
if verbosity == .detailed && model.audioValid && model.audioBufferMs > 0 {
// String(format:) for the signed offset: `%+d` has no specifier-interpolation
// equivalent, and Swift's Int is 64-bit (%lld, never the 32-bit %d).
Text(model.audioAvOffsetMs == 0
? "audio buffer \(model.audioBufferMs) ms"
: String(
format: "audio buffer %lld ms · a/v %+lld ms",
model.audioBufferMs, model.audioAvOffsetMs))
.font(.system(.caption2, design: .monospaced))
.foregroundStyle(.tertiary)
}
if model.lostFrames > 0 {
// Unrecoverable network drops this window; hidden while the link is clean.
// String(format:) rather than specifier interpolation: the literal % would
@@ -21,6 +21,12 @@ import os
/// long quiet spell relaxes it back toward the base so a session on Wi-Fi that bunches arrivals
/// deepens until it stops crackling, while a clean LAN keeps the tight base latency. Keep the
/// constants here in step with `JitterTuning.COREAUDIO`.
///
/// **A/V sync.** On top of all that the depth can be STEERED, by `setSyncTarget` from the drain
/// thread's `AvSync` because a ring that is the right depth for the link is not thereby the
/// right depth for the picture. Continuity still outranks sync: the request is clamped between
/// the underrun-driven floor above and the hard cap, so the loop can never buy alignment with a
/// dropout. `nil` (the default) is exactly the pre-sync behaviour.
final class AudioRing: @unchecked Sendable {
/// Mirrors `JitterTuning::COREAUDIO` see that type for the rationale.
private static let targetMS = 20
@@ -48,6 +54,13 @@ final class AudioRing: @unchecked Sendable {
private static let growWindowMS = 5_000
private static let growStepMS = 10
private static let shrinkQuietMS = 30_000
/// The same quiet span, while the A/V sync loop is actively asking to run shallower. A grown
/// target normally relaxes only after a long spell because, absent other evidence, the only
/// thing that can justify giving up hard-won slack is time; a sync request IS that evidence
/// a measurement saying the extra depth is costing alignment right now so a smaller target
/// gets tested sooner. Wrong guesses are cheap and self-correcting (one underrun and the
/// growth path takes it straight back). Mirrors `SHRINK_QUIET_SYNC_MS`.
private static let shrinkQuietSyncMS = 5_000
private var buf: [Float]
private var readIdx = 0
@@ -70,6 +83,14 @@ final class AudioRing: @unchecked Sendable {
/// which is a different problem from the depth being wrong.
private var underrunCount = 0
private var shedCount = 0
/// The depth the A/V sync loop would like, in interleaved samples (`AvSync.desiredDepth`).
/// `nil` the default, and what an un-wired session keeps reproduces the pre-sync
/// behaviour exactly, so this ring could adopt sync without the other three diverging.
private var syncTarget: Int?
/// The sync loop's smoothed offset in ms, STORED not computed: the ring owns the depth but has
/// no timestamps, so the drain thread (which has both a packet's `pts_ns` and the video leg)
/// hands the number back for reporting. Mirrors `NativeClient::audio_av_offset_ms`.
private var avOffsetMS = 0
private let channels: Int
private let perMS: Int
private let lock = OSAllocatedUnfairLock()
@@ -85,9 +106,64 @@ final class AudioRing: @unchecked Sendable {
/// Effective target depth in interleaved samples: the (adaptively grown) live target, lifted
/// so it can always serve one device quantum plus a packet (a large-buffer device cannot
/// sustain a target below its own quantum).
/// sustain a target below its own quantum) then, if the A/V sync loop has asked for a depth,
/// its request CLAMPED into that band. Mirrors `JitterPolicy::effective_target`.
///
/// The clamp order is the whole safety argument for steering playback depth off a network
/// measurement at all: sync may pull the ring shallower to catch the picture up, or push it
/// deeper when audio runs early, but never below what underrun pressure has proven this link
/// needs, and never past the hard cap that bounds added latency. A link whose jitter genuinely
/// demands more buffer than the picture is away keeps its buffer and the residual is REPORTED
/// (`Stats.avOffsetMS`) rather than taken out of the listener's stream.
///
/// The ceiling is raised to the floor rather than used as-is: a device whose callback quantum
/// alone exceeds `hardCapMS` makes `floor > cap`, and a plain `min(max(s, floor), cap)` would
/// then return the CAP i.e. quietly below the continuity floor, inverting the very ordering
/// this exists to guarantee, on exactly the awkward hardware it exists to survive. (Rust's
/// `Ord::clamp` announces the same condition by panicking; Swift would just get it wrong.)
private var target: Int {
max(targetLive, renderQuantum + Self.frameMS * perMS)
let floor = max(targetLive, renderQuantum + Self.frameMS * perMS)
guard let want = syncTarget else { return floor }
let cap = max(Self.hardCapMS * perMS, floor)
return min(max(want, floor), cap)
}
/// The sync loop is asking to run shallower than the adaptive target has grown to the
/// evidence `noteRead` relaxes a grown target on. Compared against the LIVE target, not the
/// effective one: it is the underrun-driven growth that a sync request is evidence against,
/// not the device-quantum lift, which no amount of measurement can argue with.
private var syncWantsLess: Bool {
guard let want = syncTarget else { return false }
return want < targetLive
}
/// Hand the ring the depth the A/V sync loop wants (`AvSync.desiredDepth`), in interleaved
/// samples, or `nil` to run unsynchronised. Called from the drain thread.
///
/// This is a REQUEST, not a command see `target` for what happens to it. `nil` is the
/// default and reproduces the pre-sync behaviour exactly.
func setSyncTarget(_ samples: Int?) {
lock.lock()
defer { lock.unlock() }
syncTarget = samples
}
/// Store the sync loop's smoothed A/V offset for reporting (positive = audio behind the
/// picture). The ring cannot compute this it has no timestamps but it is where the two
/// numbers a listener's complaint needs, depth and offset, can be read under one lock.
func noteAvOffset(_ ms: Int) {
lock.lock()
defer { lock.unlock() }
avOffsetMS = ms
}
/// Buffered depth in interleaved samples what the sync loop measures against (`bufferedMS`
/// is the same quantity rounded for humans). Everything queued here must play before the frame
/// the drain thread is about to write, which is exactly what delays it.
var bufferedSamples: Int {
lock.lock()
defer { lock.unlock() }
return writeIdx - readIdx
}
func write(_ samples: UnsafePointer<Float>, count: Int) {
@@ -196,7 +272,12 @@ final class AudioRing: @unchecked Sendable {
} else {
emptyReads = 0
quietRun += count
if quietRun >= Self.shrinkQuietMS * perMS {
// Without a sync request, time is the only evidence that hard-won slack is no longer
// needed, so a grown target waits out the long window. A request for less IS evidence,
// and without this branch a ring that ratcheted to the ceiling during a transient would
// hold audio a ceiling's worth late for minutes after the cause had gone.
let quietNeeded = syncWantsLess ? Self.shrinkQuietSyncMS : Self.shrinkQuietMS
if quietRun >= quietNeeded * perMS {
quietRun = 0
targetLive = max(targetLive - Self.growStepMS * perMS, Self.targetMS * perMS)
}
@@ -239,6 +320,12 @@ final class AudioRing: @unchecked Sendable {
let targetMS: Int
let underruns: Int
let sheds: Int
/// The A/V sync loop's smoothed offset (ms): **positive = audio playing BEHIND the
/// picture**, negative = ahead of it. `0` before the loop has evidence, or with sync off.
///
/// Reported next to the depth, never instead of it: a deep ring on a jittery link is
/// CORRECT behaviour, and only the offset separates that from a ring holding audio late.
let avOffsetMS: Int
}
var stats: Stats {
@@ -248,7 +335,154 @@ final class AudioRing: @unchecked Sendable {
bufferedMS: (writeIdx - readIdx) / max(perMS, 1),
targetMS: target / max(perMS, 1),
underruns: underrunCount,
sheds: shedCount)
sheds: shedCount,
avOffsetMS: avOffsetMS)
}
}
// MARK: - A/V sync
/// The A/V synchronisation controller: turns "when will this audio actually play" and "when did
/// the picture it belongs with reach the glass" into a ring depth `AudioRing` should aim for.
/// The Swift mirror of `punktfunk_core::audio::AvSync` keep the two in step.
///
/// **The defect it exists to fix.** The host stamps `pts_ns` on every audio datagram and the
/// client decoded it into `AudioPCM` and then never read it. Video's `pts_ns`, by contrast, is
/// used end to end (`LatencyMeter` computes a true glass-to-glass `displayed + clockOffset pts`
/// per presented frame). So audio free-ran at whatever depth its jitter ring happened to settle
/// at, video was presented on a wholly independent path, and nothing ever compared them: the A/V
/// offset was an accident of buffer depths. It moved whenever the ring ratcheted under underrun
/// pressure, and the way this surfaced in the field it got WORSE every time video got faster,
/// because a quicker decoder lowers the video leg while leaving the audio leg exactly where it was.
///
/// **Video is the master.** In a game streamer the video leg is the input-feel budget and must
/// never be inflated to satisfy the audio clock; audio tolerates small, crossfaded, rate-limited
/// corrections that are inaudible, and `AudioRing.shedOneFrame` already applies them. So audio
/// moves.
///
/// **Continuity outranks sync.** This type only ever PROPOSES a depth. `AudioRing` clamps the
/// proposal to its own underrun-driven floor (see `AudioRing.target`), so a link whose jitter
/// genuinely needs more buffer than the picture is away keeps its buffer and the residual is
/// reported instead of being taken out of the listener's stream.
///
/// Not a class and not locked: it is owned outright by the drain thread that observes packets.
struct AvSync {
/// Smoothing time constant for the measured offset, in ms of consumed audio. Long enough that
/// network jitter and a single late datagram do not move it; short enough to track real drift.
private static let ewmaTauMS = 2_000
/// Offsets inside this band are left alone. Correcting a few ms costs a (crossfaded, but real)
/// discontinuity and buys nothing a listener can perceive detectability for A/V misalignment
/// sits an order of magnitude above it. The deadband is what keeps the loop from hunting
/// forever around zero, which would be audible in a way the misalignment it chased was not.
private static let deadbandMS = 10
/// Observations folded before the first correction is offered. The offset is derived from a
/// clock skew estimate and a video figure that both need a moment to settle after connect;
/// acting on the first sample would chase the handshake, not the stream.
private static let minObservations = 100
/// An offset larger than this is not believed. A wall-clock step, a paused host, or a stale
/// video figure can all produce an enormous apparent misalignment, and steering the ring by it
/// would empty or overfill it outright. Beyond this the loop reports and waits rather than acts.
private static let saneLimitMS = 1_000
/// The protocol's frame, in ms the EWMA is weighted by it so the time constant means the
/// same thing however often the caller observes.
private static let frameMS = 5
/// Interleaved samples per millisecond at the negotiated layout (48 × channels).
private let perMS: Int
/// EWMA of the measured offset in ns. Positive = audio is scheduled to play LATE relative to
/// the picture it belongs with.
private var offsetAvgNs: Double = 0
private var observations = 0
/// Set once an observation lands outside `saneLimitMS`, for reporting.
private(set) var implausible = false
/// `channels` is the negotiated interleaved channel count (2/6/8).
init(channels: Int) {
perMS = 48 * max(channels, 1)
}
/// One measurement handed to `observe`. Every field is in the units its source already
/// produces, so no caller has to do clock arithmetic to use it correctly.
struct Observation {
/// The host capture timestamp carried by the audio frame being queued (host clock).
let ptsNs: UInt64
/// Local `CLOCK_REALTIME` now the same basis `LatencyMeter` stamps video in.
let nowLocalNs: Int64
/// Host clock minus client clock, from the skew handshake (`clockOffsetNs`).
///
/// It very nearly CANCELS: the video figure this is differenced against was computed with
/// the same offset and the same sign, so as long as both terms use one value the skew
/// drops out of the result entirely. That is what makes the connect-time offset good
/// enough here even though the absolute legs would prefer a re-synced one.
let clockOffsetNs: Int64
/// How much audio is already queued AHEAD of this frame, in interleaved samples
/// everything that must play before it does.
let bufferedAhead: Int
/// The video plane's current end-to-end figure in ns: `displayed + clockOffset pts`, as
/// `LatencyMeter` already computes it per presented frame. `nil` while nothing has reached
/// the glass recently no reference, no correction.
let videoE2eNs: Int64?
}
/// Fold one measurement. Returns the smoothed offset in ns once there is enough evidence to
/// believe it (positive = audio late), or `nil` while still settling.
///
/// Rejecting the implausible rather than clamping it is deliberate: a wall-clock step or a
/// stale video figure produces a huge apparent offset, and a clamped-but-wrong value would be
/// acted on as though it were a small real one.
@discardableResult
mutating func observe(_ o: Observation) -> Int64? {
// No frame on the glass yet no reference to align against, so nothing to say.
guard let videoE2eNs = o.videoE2eNs else { return nil }
// When this frame's samples will actually reach the speaker, expressed in the host's
// capture clock the same clock, and the same shape, as the video figure it is compared
// against.
let bufferedNs = Int64(o.bufferedAhead / max(perMS, 1)) * 1_000_000
// Overflow-reporting arithmetic, NOT the wrapping `&+`/`&-` the meters use. Every term is
// a nanosecond count on the same epoch (~1.8e18), so the DIFFERENCE is tiny while the
// operands sit within a factor of five of `Int64.max` and a garbage `pts_ns` would wrap
// a nonsense value round into a small, plausible-looking offset. This loop's entire
// defence is that it can tell nonsense from a real misalignment, so an overflow takes the
// same exit the sanity limit does rather than being silently believed.
let (playAtLocal, o1) = o.nowLocalNs.addingReportingOverflow(bufferedNs)
let (playAtHost, o2) = playAtLocal.addingReportingOverflow(o.clockOffsetNs)
let (audioE2eNs, o3) = playAtHost.subtractingReportingOverflow(Int64(bitPattern: o.ptsNs))
let (offsetNs, o4) = audioE2eNs.subtractingReportingOverflow(videoE2eNs)
guard !o1, !o2, !o3, !o4, abs(offsetNs) <= Int64(Self.saneLimitMS) * 1_000_000 else {
implausible = true
return nil
}
implausible = false
let alpha = min(1.0, Double(Self.frameMS) / Double(Self.ewmaTauMS))
if observations == 0 {
offsetAvgNs = Double(offsetNs)
} else {
offsetAvgNs += (Double(offsetNs) - offsetAvgNs) * alpha
}
observations += 1
return settled ? Int64(offsetAvgNs) : nil
}
/// Enough evidence folded to act on.
var settled: Bool { observations >= Self.minObservations }
/// The smoothed offset in ms (positive = audio late), for the HUD. Reported as soon as it is
/// measured, including while still settling a number the operator can watch converge is more
/// useful than a blank that hides whether the loop is working at all.
var offsetMS: Int { Int(offsetAvgNs / 1_000_000) }
/// The ring depth that would place audio with the picture, given where the ring is now.
/// `nil` while unsettled or inside the deadband the caller then leaves the ring alone.
///
/// Audio late (offset > 0) means there is too much queued: aim shallower. Audio early means
/// aim deeper.
func desiredDepth(currentDepth: Int) -> Int? {
guard settled else { return nil }
let offsetMs = offsetAvgNs / 1_000_000
guard abs(offsetMs) >= Double(Self.deadbandMS) else { return nil }
let delta = Int(offsetMs * Double(perMS))
return max(0, currentDepth - delta)
}
}
@@ -62,6 +62,13 @@ public final class SessionAudio {
/// not the ring, so the drain thread never has to be re-pointed). Main-thread confined,
/// like every start path.
private var ring: AudioRing?
/// The video plane's end-to-end meter (captureon-glass), if the owner wired one the
/// reference the A/V sync loop steers the ring against. `nil` leaves the loop inert and the
/// ring exactly as it was before sync existed, which is also what the stage-1 fallback
/// presenter gets: it decodes and presents inside the layer with no per-frame stamp, so it can
/// offer no reference, and a loop with no reference must not invent one. Main-thread confined,
/// like `ring`; the meter itself is internally locked and read from the drain thread.
private var videoLatency: LatencyMeter?
#if !os(macOS)
/// AVAudioSession `setCategory`/`setActive` are synchronous and block on the audio server, so
/// they must not run on the main thread (UI stall AVFoundation warns about it). PROCESS-WIDE
@@ -91,9 +98,16 @@ public final class SessionAudio {
/// a later main-queue hop (gated by `!flag.isStopped`) so playback is live shortly after, not
/// on return. The mic may start later still if the permission prompt is pending.
/// `echoCancel` picks the engine topology see the header note and `wantsCombined`.
///
/// `videoLatency` is the session's END-TO-END latency meter (captureon-glass). Pass it to arm
/// A/V sync: it is the only thing that tells the audio plane where the picture actually is, and
/// without it the ring keeps today's free-running behaviour. Omit it for a playback-only or
/// stage-1 session, where no such figure is measured.
public func start(
speakerUID: String, micUID: String, micChannel: Int, micEnabled: Bool, echoCancel: Bool
speakerUID: String, micUID: String, micChannel: Int, micEnabled: Bool, echoCancel: Bool,
videoLatency: LatencyMeter? = nil
) {
self.videoLatency = videoLatency
#if os(macOS)
// No AVAudioSession on macOS start the engines directly (caller's thread, as before).
startEngines(
@@ -305,6 +319,31 @@ public final class SessionAudio {
}
}
// MARK: - Stats
/// The playback plane's two latency numbers, for the stats overlay.
///
/// Both, never just the depth: a deep ring on a jittery link is CORRECT behaviour the
/// adaptive floor put it there because the link kept starving and only the offset separates
/// that from a ring that is simply holding audio late. Before this pair existed the plane
/// published nothing any surface could render (depth and target lived in a periodic log line),
/// and a field investigation into "the audio delay seems way too high" ran all the way to its
/// conclusion without either number.
public struct Stats: Sendable {
/// Decoded audio queued ahead of the speaker (ms).
public let bufferMS: Int
/// The A/V sync loop's smoothed offset (ms): positive = audio playing BEHIND the picture.
/// `0` before the loop has evidence, with sync unwired, or genuinely aligned.
public let avOffsetMS: Int
}
/// A snapshot of `Stats`, or nil before playback starts. Main thread (`ring` is main-confined;
/// the ring's own numbers are taken under its lock, so they describe one instant).
public var stats: Stats? {
guard let s = ring?.stats else { return nil }
return Stats(bufferMS: s.bufferedMS, avOffsetMS: s.avOffsetMS)
}
// MARK: - Playback (host speaker)
/// The playback jitter ring + the source node draining it shared by the plain playback
@@ -401,9 +440,25 @@ public final class SessionAudio {
}
drainStarted = true
stateLock.unlock()
// A/V sync. This thread is the only place that holds all three ingredients at once: the
// packet's host capture `ptsNs`, the ring depth, and the video plane's end-to-end figure.
// `ptsNs` was decoded into `AudioPCM` and then dropped on the floor right here for the
// plane's entire existence, which is why audio ran at whatever depth its jitter ring
// happened to settle at and nothing ever placed it against the picture.
//
// The escape hatch mirrors the Rust clients': a field regression in a loop that steers
// PLAYBACK should be bisectable without a rebuild. macOS honours it from the environment;
// elsewhere it simply never trips, which is the same as today's behaviour.
let syncEnabled = !["1", "true"].contains(
ProcessInfo.processInfo.environment["PUNKTFUNK_NO_AV_SYNC"] ?? "")
// nil disarms the loop entirely no reference, no correction (see `videoLatency`).
let videoLatency = syncEnabled ? self.videoLatency : nil
if !syncEnabled { log.info("A/V sync disabled by PUNKTFUNK_NO_AV_SYNC") }
let channels = Int(connection.resolvedAudioChannels)
let thread = Thread { [connection, flag, drainDone] in
defer { drainDone.signal() }
var drained = 0
var av = AvSync(channels: channels)
// Decode happens IN-CORE (libopus multistream) AudioToolbox's Opus path is
// stereo-only and is handed back as interleaved f32 PCM in wire channel order.
// Per-iteration autorelease pool: no runloop on this thread (see Stage2Pipeline).
@@ -417,6 +472,25 @@ public final class SessionAudio {
return false // session closed
}
guard let pcm, pcm.frameCount > 0 else { return true }
// Place this frame against the picture it belongs with BEFORE queueing it: the
// depth read here is everything that must still play first, which is exactly what
// delays it. Skipped wholesale when no meter was wired, so an un-armed session
// does not even read the ring.
if let videoLatency {
let depth = ring.bufferedSamples
var ts = timespec()
clock_gettime(CLOCK_REALTIME, &ts)
let nowNs = Int64(ts.tv_sec) * 1_000_000_000 + Int64(ts.tv_nsec)
// Half a second of tolerance on the reference: long enough to ride out a
// stalled or hitching present path, short enough that a backgrounded session
// (video decode dropped, audio still playing) stops steering almost at once.
av.observe(AvSync.Observation(
ptsNs: pcm.ptsNs, nowLocalNs: nowNs,
clockOffsetNs: connection.clockOffsetNs, bufferedAhead: depth,
videoE2eNs: videoLatency.latestSample(asOfNs: nowNs, maxAgeMs: 500)))
ring.setSyncTarget(av.desiredDepth(currentDepth: depth))
ring.noteAvOffset(av.offsetMS)
}
pcm.samples.withUnsafeBufferPointer { p in
if let base = p.baseAddress {
ring.write(base, count: pcm.frameCount * pcm.channels)
@@ -430,7 +504,7 @@ public final class SessionAudio {
if drained % 2_000 == 0 {
let s = ring.stats
log.info(
"audio: buffer_ms=\(s.bufferedMS) target_ms=\(s.targetMS) underruns=\(s.underruns) drift_sheds=\(s.sheds)"
"audio: buffer_ms=\(s.bufferedMS) target_ms=\(s.targetMS) underruns=\(s.underruns) drift_sheds=\(s.sheds) av_offset_ms=\(s.avOffsetMS)"
)
}
return true
@@ -24,6 +24,10 @@ public final class LatencyMeter: @unchecked Sendable {
private let lock = NSLock()
private var samplesUs: [Int64] = []
private var skewCorrected = false
/// The most recent sample and the instant it ended, for `latestSample(asOfNs:maxAgeMs:)`
/// a LEVEL, not a window, so `drain` deliberately leaves both alone.
private var latestNs: Int64 = 0
private var latestAtNs: Int64 = 0
public init() {}
@@ -49,10 +53,42 @@ public final class LatencyMeter: @unchecked Sendable {
guard latNs > 0, latNs < 10_000_000_000 else { return }
lock.lock()
samplesUs.append(latNs / 1000)
latestNs = latNs
latestAtNs = atNs
if offsetNs != 0 { skewCorrected = true }
lock.unlock()
}
/// The most recent single sample in ns, or `nil` if none has landed or the last one ended more
/// than `maxAgeMs` before `nowNs` (both `CLOCK_REALTIME`). Unlike `drain`, this reports a level
/// rather than a window, and reading it consumes nothing.
///
/// **What it is for.** Read off the END-TO-END meter, this is the video plane's live
/// glass-to-glass figure `displayed + clockOffset pts`, exactly the shape `AvSync` compares
/// audio against and it is the reference the A/V sync loop needs. It is published from
/// `record`, so BOTH present paths (arrival and deadline) feed it without either knowing that
/// audio exists.
///
/// **Why staleness is not optional.** The number is a level, so absent an age check it would
/// simply keep its last value forever. This client has a state where that matters: the
/// backgrounded keep-alive keeps audio playing and DROPS video decode entirely, so the loop
/// would go on steering the ring against a reference minutes old and frozen. Expiring it
/// returns `nil`, which is the same "no reference yet" case as session start the loop holds
/// its last correction and stops chasing. `nowNs` is caller-supplied rather than read fresh so
/// the audio side compares against exactly the instant it timestamped its own frame at.
///
/// Only the PAST is bounded. A present stamp can legitimately sit a hair ahead of the reader's
/// clock (the deadline presenter stamps at the link's target present time), and discarding the
/// only reference we have over a fraction of a refresh would make it flap in and out; a stamp
/// wildly in the future instead yields a huge offset, which `AvSync` refuses on its own terms.
public func latestSample(asOfNs nowNs: Int64, maxAgeMs: Int) -> Int64? {
lock.lock()
defer { lock.unlock() }
guard latestNs > 0 else { return nil }
guard (nowNs &- latestAtNs) <= Int64(maxAgeMs) * 1_000_000 else { return nil }
return latestNs
}
public struct Stats: Sendable {
public let p50Ms: Double
public let p95Ms: Double
@@ -198,5 +198,315 @@ final class AudioRingDriftTests: XCTestCase {
silentTail, 0,
"after adapting, the last 3 s must play through the bunching without a dropout")
}
// MARK: - A/V sync (audio latency overhaul, W6)
//
// The second half of the same story. Depth alone is not correctness: a ring can be exactly as
// deep as its link needs and still put audio in the wrong place, because nothing ever compared
// it to the picture. `AvSync` measures that comparison and asks the ring to move; the ring is
// free to refuse. These pin both halves that the loop DOES act (the previous pass in this
// area shipped a correction that was structurally unreachable and had a green test), and that
// it can never act far enough to starve the callback.
/// Build an observation whose measured offset is exactly `offsetMS` (positive = audio late).
/// Mirrors the Rust `obs` helper: pin now/skew/pts so the only free term is the buffered depth,
/// then choose the video figure so the difference lands where we want it.
private func obs(offsetMS: Int, depth: Int) -> AvSync.Observation {
let bufferedMS = depth / perMS
let audioE2eMS = bufferedMS + 40 // 40 ms of transport, arbitrary but fixed
let videoE2eMS = audioE2eMS - offsetMS
return AvSync.Observation(
ptsNs: 1_000_000_000,
nowLocalNs: 1_000_000_000 + 40 * 1_000_000,
clockOffsetNs: 0,
bufferedAhead: depth,
videoE2eNs: Int64(max(0, videoE2eMS)) * 1_000_000)
}
/// Fold `n` identical observations in.
private func settle(_ sync: inout AvSync, offsetMS: Int, depth: Int, count: Int = 100) {
for _ in 0..<count { sync.observe(obs(offsetMS: offsetMS, depth: depth)) }
}
func testAvSyncNeedsEvidenceBeforeActing() {
var s = AvSync(channels: channels)
// One sample is never enough the skew estimate and the video figure both settle after
// connect, and acting on the first would chase the handshake, not the stream.
XCTAssertNil(s.observe(obs(offsetMS: 50, depth: 30 * perMS)))
XCTAssertFalse(s.settled)
XCTAssertNil(s.desiredDepth(currentDepth: 30 * perMS))
settle(&s, offsetMS: 50, depth: 30 * perMS, count: 99) // 1 + 99 = 100
XCTAssertTrue(s.settled, "should act once the evidence is in")
}
/// No frame on the glass no reference the loop says nothing, however many observations
/// arrive. This is the state every session starts in, and the one the stage-1 fallback
/// presenter stays in for its whole life.
func testAvSyncWithoutAVideoReferenceNeverActs() {
var s = AvSync(channels: channels)
for _ in 0..<500 {
s.observe(AvSync.Observation(
ptsNs: 1_000_000_000, nowLocalNs: 1_040_000_000, clockOffsetNs: 0,
bufferedAhead: 30 * perMS, videoE2eNs: nil))
}
XCTAssertFalse(s.settled)
XCTAssertNil(s.desiredDepth(currentDepth: 30 * perMS))
}
func testAvSyncAimsShallowerWhenAudioIsLate() {
let depth = 60 * perMS
var s = AvSync(channels: channels)
settle(&s, offsetMS: 40, depth: depth, count: 400)
guard let want = s.desiredDepth(currentDepth: depth) else {
return XCTFail("a 40 ms offset is actionable")
}
XCTAssertLessThan(want, depth, "audio late must aim shallower")
// The correction is the offset, not a guess at it.
let shedMS = (depth - want) / perMS
XCTAssertTrue((35...45).contains(shedMS), "should aim to shed ~40 ms, got \(shedMS)")
XCTAssertEqual(s.offsetMS, 40, "and report it, sign and all")
}
func testAvSyncAimsDeeperWhenAudioIsEarly() {
let depth = 20 * perMS
var s = AvSync(channels: channels)
settle(&s, offsetMS: -30, depth: depth, count: 400)
guard let want = s.desiredDepth(currentDepth: depth) else {
return XCTFail("a 30 ms offset is actionable")
}
XCTAssertGreaterThan(want, depth, "audio early must aim deeper")
XCTAssertEqual(s.offsetMS, -30)
}
func testAvSyncDeadbandsWhatNoOneCanHear() {
let depth = 30 * perMS
var s = AvSync(channels: channels)
settle(&s, offsetMS: 8, depth: depth, count: 400) // inside the 10 ms deadband
XCTAssertNil(
s.desiredDepth(currentDepth: depth),
"an offset inside the deadband must not provoke a (real, if crossfaded) discontinuity")
XCTAssertEqual(s.offsetMS, 8, "…but it is still REPORTED — the HUD shows the residual")
}
/// A wall-clock step or a stale video figure produces an enormous apparent misalignment.
/// Clamping it would act on a wrong number as though it were a small real one, so it is
/// refused outright and the running average is left untouched.
func testAvSyncRejectsTheImplausibleInsteadOfClampingIt() {
let depth = 30 * perMS
var s = AvSync(channels: channels)
settle(&s, offsetMS: 30, depth: depth, count: 400)
let before = s.offsetMS
// Built directly rather than through `obs`: that helper floors the video figure at zero,
// which would cap the offset at a merely LARGE value and let this pass without ever
// exercising the rejection.
let wild = AvSync.Observation(
ptsNs: 0, nowLocalNs: 5_000_000_000, clockOffsetNs: 0,
bufferedAhead: depth, videoE2eNs: 40_000_000)
XCTAssertNil(s.observe(wild))
XCTAssertTrue(s.implausible, "a ~5 s offset must be refused, not folded")
XCTAssertEqual(before, s.offsetMS, "an implausible sample must be discarded, not folded in")
}
/// The same refusal for arithmetic that cannot even be CARRIED OUT, which is why the terms are
/// combined with overflow-reporting operators rather than the wrapping `&-` the latency meters
/// use.
///
/// This input is not arbitrary. `ptsNs = 1 << 63` reads as `Int64.min` in two's complement, so
/// the audio leg overflows and the difference lands on EXACTLY `Int64.min` and `abs()` of
/// `Int64.min` has no representable result, so in Swift it traps. Check the overflow flags
/// after the sanity limit instead of before and this observation does not mis-measure the
/// stream, it aborts the process, from the audio drain thread. The guard's short-circuit
/// ordering is what makes the sanity check itself safe to run.
func testAvSyncRefusesAnOffsetItCannotEvenCompute() {
var s = AvSync(channels: channels)
let wild = AvSync.Observation(
ptsNs: 1 << 63, nowLocalNs: 40_000_000, clockOffsetNs: 0,
bufferedAhead: 0, videoE2eNs: 40_000_000)
XCTAssertNil(s.observe(wild))
XCTAssertTrue(s.implausible)
XCTAssertFalse(s.settled, "a refused sample is not evidence")
XCTAssertEqual(s.offsetMS, 0, "and nothing of it was folded in")
}
// MARK: - and what the ring does with the proposal
/// Drive one read so the ring knows the device quantum (`renderQuantum` seeds the floor).
private func primeQuantum(_ ring: AudioRing, quantumMS: Int) {
var scratch = [Float](repeating: 0, count: quantumMS * perMS)
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: $0.count) }
}
/// The loop is NOT inert: a settled proposal inside the ring's legal band actually moves the
/// effective target. Without this the whole feature could ship as unreachable code with every
/// other test still green which is exactly how the previous drift correction shipped dead.
func testSyncActuallyMovesTheTarget() {
let ring = AudioRing(capacity: 48_000 * channels, channels: channels)
primeQuantum(ring, quantumMS: 5)
XCTAssertEqual(ring.stats.targetMS, 20, "base target (JitterTuning.COREAUDIO)")
// Audio 30 ms EARLY at a 20 ms depth aim 50 ms deep: above the floor, under the 90 ms
// cap, so the ring has no reason to refuse.
var s = AvSync(channels: channels)
settle(&s, offsetMS: -30, depth: 20 * perMS, count: 400)
ring.setSyncTarget(s.desiredDepth(currentDepth: 20 * perMS))
XCTAssertEqual(ring.stats.targetMS, 50, "the ring must adopt a legal request")
// And releasing it returns the ring to exactly where it was.
ring.setSyncTarget(nil)
XCTAssertEqual(ring.stats.targetMS, 20)
}
/// THE safety invariant: sync only ever proposes. Continuity the underrun-driven floor
/// outranks it, or a lossy link would be "synced" into dropouts. Pinned against a GROWN floor,
/// not just the base, because the floor sync is most likely to argue with is the one a bad link
/// earned.
func testSyncCanNeverStarveTheRing() {
let ring = AudioRing(capacity: 48_000 * channels, channels: channels)
let want = 5 * perMS
var scratch = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 25 * perMS)
func write(ms: Int) {
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: ms * perMS) }
}
func read() {
scratch.withUnsafeMutableBufferPointer { ring.read(into: $0.baseAddress!, count: want) }
}
// Grow the floor above the base with three clustered genuine underruns (same shape as
// testTargetGrowsOnUnderrunsAndRelaxesWhenQuiet).
write(ms: 25)
for _ in 0..<5 { read() }
read()
write(ms: 5); read()
read()
write(ms: 5); read()
read()
let floor = ring.stats.targetMS
XCTAssertGreaterThan(floor, 20, "the test needs a GROWN floor to be meaningful")
// Ask for an absurdly shallow ring zero.
ring.setSyncTarget(0)
XCTAssertEqual(
ring.stats.targetMS, floor,
"sync pulled the target below the continuity floor — a link that needs the buffer must "
+ "keep it, and the residual gets reported instead")
// One frame under the floor is still under the floor.
ring.setSyncTarget(floor * perMS - perMS)
XCTAssertEqual(ring.stats.targetMS, floor)
// And it may not blow past the hard cap either added latency stays bounded.
ring.setSyncTarget(Int.max / 2)
XCTAssertLessThanOrEqual(ring.stats.targetMS, 90, "sync pushed the target past the hard cap")
}
/// A device whose callback quantum alone exceeds the hard cap puts the continuity floor ABOVE
/// the ceiling. The floor must win: clamping naively (`min(max(s, floor), cap)`) would hand
/// back the cap quietly below the floor, inverting the whole ordering on exactly the
/// awkward hardware this code exists to survive.
func testAHugeDeviceQuantumDoesNotInvertTheClamp() {
let ring = AudioRing(capacity: 48_000 * channels * 2, channels: channels)
let quantumMS = 500 // absurd, but not a reason to starve the callback
primeQuantum(ring, quantumMS: quantumMS)
ring.setSyncTarget(0)
XCTAssertGreaterThanOrEqual(
ring.stats.targetMS, quantumMS,
"the target must still be able to serve one callback")
}
/// A ring that ratcheted during a transient must not hold audio late for minutes after the
/// cause is gone: with sync asking for less, the relax window is the short one.
func testSyncPressureRelaxesAGrownTargetSoonerThanTimeAlone() {
let want = 5 * perMS
func grow(_ ring: AudioRing) {
var scratch = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 25 * perMS)
func write(ms: Int) {
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: ms * perMS) }
}
func read() {
scratch.withUnsafeMutableBufferPointer {
ring.read(into: $0.baseAddress!, count: want)
}
}
write(ms: 25)
for _ in 0..<5 { read() }
read()
write(ms: 5); read()
read()
write(ms: 5); read()
read()
}
/// Quiet (full) reads needed before the grown target relaxes one step.
func quietToRelax(_ ring: AudioRing) -> Int {
var scratch = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 5 * perMS)
let start = ring.stats.targetMS
var reads = 0
while ring.stats.targetMS == start, reads < 200_000 {
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: 5 * perMS) }
scratch.withUnsafeMutableBufferPointer {
ring.read(into: $0.baseAddress!, count: want)
}
reads += 1
}
return reads
}
let slow = AudioRing(capacity: 48_000 * channels, channels: channels)
grow(slow)
slow.setSyncTarget(nil)
let slowReads = quietToRelax(slow)
let fast = AudioRing(capacity: 48_000 * channels, channels: channels)
grow(fast)
fast.setSyncTarget(perMS) // strictly shallower than the grown target
let fastReads = quietToRelax(fast)
XCTAssertLessThan(
fastReads, slowReads,
"sync pressure should relax sooner: \(fastReads) vs \(slowReads) quiet reads")
}
/// The four client rings adopt sync one at a time; an un-wired one must behave exactly as it
/// did. `nil` is the default, so this pins the initializer too and every other test in this
/// file runs without a sync target, which is the real guard that nothing moved underneath them.
func testNoSyncTargetLeavesTheRingExactlyAsItWas() {
let a = AudioRing(capacity: 48_000 * channels, channels: channels)
let b = AudioRing(capacity: 48_000 * channels, channels: channels)
b.setSyncTarget(nil)
let want = 5 * perMS
var sa = [Float](repeating: 0, count: want)
var sb = [Float](repeating: 0, count: want)
let feed = [Float](repeating: 0.5, count: 30 * perMS)
for step in 0..<4_000 {
// Uneven delivery so the depth actually moves around and the two rings have something
// to disagree about.
if step % 7 == 0 {
for r in [a, b] {
feed.withUnsafeBufferPointer { r.write($0.baseAddress!, count: 30 * perMS) }
}
}
sa.withUnsafeMutableBufferPointer { a.read(into: $0.baseAddress!, count: want) }
sb.withUnsafeMutableBufferPointer { b.read(into: $0.baseAddress!, count: want) }
XCTAssertEqual(sa, sb, "step \(step): an explicit nil diverged from the default")
XCTAssertEqual(a.stats.targetMS, b.stats.targetMS, "step \(step)")
}
}
/// The reporting half of §1.3: the offset must reach the same snapshot the depth does, because
/// a depth on its own cannot distinguish "deep because the link needs it" from "deep and
/// therefore late". This is the number the HUD and the 1 Hz log line read.
func testAvOffsetIsReportedAlongsideTheDepth() {
let ring = AudioRing(capacity: 48_000 * channels, channels: channels)
XCTAssertEqual(ring.stats.avOffsetMS, 0, "no evidence yet reads as zero, not as noise")
let feed = [Float](repeating: 0.5, count: 30 * perMS)
feed.withUnsafeBufferPointer { ring.write($0.baseAddress!, count: 30 * perMS) }
var s = AvSync(channels: channels)
settle(&s, offsetMS: 37, depth: 30 * perMS, count: 400)
ring.noteAvOffset(s.offsetMS)
let stats = ring.stats
XCTAssertEqual(stats.bufferedMS, 30)
XCTAssertEqual(stats.avOffsetMS, 37, "positive = audio behind the picture")
}
}
#endif
@@ -71,4 +71,65 @@ final class LatencyMeterTests: XCTestCase {
m.record(ptsNs: now - 20_000_000_000, offsetNs: 0)
XCTAssertNil(m.drain())
}
// MARK: - latestSample: the A/V sync loop's video reference
/// The end-to-end meter doubles as the reference the audio ring steers against, so its most
/// recent sample must be readable as a LEVEL without consuming it, and independently of the
/// 1 Hz percentile window the HUD drains.
func testLatestSampleSurvivesDrainAndIsNotAWindow() {
let m = LatencyMeter()
let atNs: Int64 = 1_000_000_000_000
m.record(ptsNs: UInt64(atNs - 12_000_000), atNs: atNs, offsetNs: 0) // 12 ms
XCTAssertEqual(m.latestSample(asOfNs: atNs, maxAgeMs: 500), 12_000_000)
_ = m.drain()
XCTAssertEqual(
m.latestSample(asOfNs: atNs, maxAgeMs: 500), 12_000_000,
"the reference is a level — draining the percentile window must not clear it")
// and it tracks the newest frame.
m.record(ptsNs: UInt64(atNs - 20_000_000), atNs: atNs, offsetNs: 0)
XCTAssertEqual(m.latestSample(asOfNs: atNs, maxAgeMs: 500), 20_000_000)
}
/// No frame yet no reference. This is what keeps the sync loop inert at session start and
/// under the stage-1 presenter, which stamps no present at all.
func testLatestSampleIsNilBeforeAnyFrame() {
XCTAssertNil(LatencyMeter().latestSample(asOfNs: 1_000_000_000_000, maxAgeMs: 500))
}
/// THE staleness gate: video can stop while audio keeps playing (the backgrounded keep-alive
/// drops decode entirely). A level with no expiry would go on offering a minutes-old figure as
/// though it were live, and the ring would be steered against a frozen reference.
func testLatestSampleExpires() {
let m = LatencyMeter()
let atNs: Int64 = 1_000_000_000_000
m.record(ptsNs: UInt64(atNs - 12_000_000), atNs: atNs, offsetNs: 0)
XCTAssertNotNil(m.latestSample(asOfNs: atNs + 499_000_000, maxAgeMs: 500))
XCTAssertNil(
m.latestSample(asOfNs: atNs + 501_000_000, maxAgeMs: 500),
"a stale reference must read as NO reference, not as a live one")
// A stamp marginally ahead of the reader's clock is normal (the deadline presenter stamps
// at the link's TARGET present time) and must not drop the only reference we have.
XCTAssertNotNil(m.latestSample(asOfNs: atNs - 8_000_000, maxAgeMs: 500))
}
/// A sample the meter refused must not become a reference either the sync loop would then be
/// steered by a value the percentile window itself judged absurd.
///
/// The ABSURDLY LARGE case is the load-bearing one: a negative interval would also be stopped
/// by `latestSample`'s own `> 0` check, so on its own it proves nothing about where the publish
/// sits relative to the guard.
func testRefusedSampleIsNotPublishedAsAReference() {
let m = LatencyMeter()
let atNs: Int64 = 1_000_000_000_000
m.record(ptsNs: UInt64(atNs - 20_000_000_000), atNs: atNs, offsetNs: 0) // 20 s refused
XCTAssertNil(
m.latestSample(asOfNs: atNs, maxAgeMs: 500),
"a sample too absurd for the window is too absurd to steer the ring")
m.record(ptsNs: UInt64(atNs + 1), atNs: atNs, offsetNs: 0) // negative interval
XCTAssertNil(m.latestSample(asOfNs: atNs, maxAgeMs: 500))
// and a good sample after them still lands, so the refusals cost nothing.
m.record(ptsNs: UInt64(atNs - 9_000_000), atNs: atNs, offsetNs: 0)
XCTAssertEqual(m.latestSample(asOfNs: atNs, maxAgeMs: 500), 9_000_000)
}
}
+24 -1
View File
@@ -107,6 +107,9 @@ pub struct AudioPlayer {
recycle_rx: Receiver<Vec<f32>>,
quit_tx: pipewire::channel::Sender<Terminate>,
thread: Option<std::thread::JoinHandle<()>>,
/// A/V sync hand-off with the PipeWire callback: it publishes the ring depth, the decode
/// thread posts the depth the sync loop wants. See [`punktfunk_core::audio::AudioSyncCell`].
sync: Arc<punktfunk_core::audio::AudioSyncCell>,
}
impl AudioPlayer {
@@ -121,10 +124,12 @@ impl AudioPlayer {
// as the data channel; a full pool just drops the Vec (plain deallocation).
let (recycle_tx, recycle_rx) = std::sync::mpsc::sync_channel::<Vec<f32>>(64);
let (quit_tx, quit_rx) = pipewire::channel::channel::<Terminate>();
let sync: Arc<punktfunk_core::audio::AudioSyncCell> = Arc::default();
let sync_cb = sync.clone();
let thread = std::thread::Builder::new()
.name("punktfunk-audio".into())
.spawn(move || {
if let Err(e) = pw_thread(pcm_rx, recycle_tx, quit_rx, channels as usize) {
if let Err(e) = pw_thread(pcm_rx, recycle_tx, quit_rx, channels as usize, sync_cb) {
tracing::warn!(error = %e, "audio playback thread ended");
}
})
@@ -134,9 +139,16 @@ impl AudioPlayer {
recycle_rx,
quit_tx,
thread: Some(thread),
sync,
})
}
/// The A/V sync hand-off cell — the decode thread reads the ring depth from it and posts the
/// depth the sync loop wants back through it.
pub fn sync_cell(&self) -> Arc<punktfunk_core::audio::AudioSyncCell> {
self.sync.clone()
}
/// A recycled chunk Vec from the pool, empty but with its capacity intact — fill it
/// and hand it back through [`push`](Self::push). Allocates only when the pool is dry
/// (startup, or after the PipeWire side dropped chunks).
@@ -180,6 +192,8 @@ struct PlayerData {
underruns: u64,
sheds: u64,
callbacks: u64,
/// A/V sync hand-off with the decode thread (depth out, target in).
sync: Arc<punktfunk_core::audio::AudioSyncCell>,
}
fn pw_thread(
@@ -187,6 +201,7 @@ fn pw_thread(
recycle_tx: SyncSender<Vec<f32>>,
quit_rx: pipewire::channel::Receiver<Terminate>,
channels: usize,
sync: Arc<punktfunk_core::audio::AudioSyncCell>,
) -> Result<()> {
use pipewire as pw;
use pw::{properties::properties, spa};
@@ -240,6 +255,7 @@ fn pw_thread(
underruns: 0,
sheds: 0,
callbacks: 0,
sync,
};
let _listener = stream
@@ -267,6 +283,13 @@ fn pw_thread(
let want_frames = data.data().map(|s| s.len() / stride).unwrap_or(0);
let want = want_frames * ud.channels;
// A/V sync: take whatever depth the decode thread's sync loop last asked for, and
// publish where the ring actually is so it can measure the result. The policy
// clamps the request between its own underrun floor and the hard cap — continuity
// outranks sync, always (see `JitterPolicy::set_sync_target`).
ud.policy.set_sync_target(ud.sync.target());
ud.sync.publish_depth(ud.ring.len());
// Shared de-jitter policy: prime depth in MILLISECONDS, smooth drift correction
// (a crossfaded 5 ms shed) so latency returns to target instead of ratcheting,
// and a hard cap as the backstop.
+21 -1
View File
@@ -163,6 +163,9 @@ pub struct AudioPlayer {
recycle_rx: Receiver<Vec<f32>>,
stop: Arc<AtomicBool>,
thread: Option<std::thread::JoinHandle<()>>,
/// A/V sync hand-off with the render thread: it publishes the ring depth, the decode thread
/// posts the depth the sync loop wants. See [`punktfunk_core::audio::AudioSyncCell`].
sync: Arc<punktfunk_core::audio::AudioSyncCell>,
}
impl AudioPlayer {
@@ -179,10 +182,13 @@ impl AudioPlayer {
let stop = Arc::new(AtomicBool::new(false));
let (ready_tx, ready_rx) = std::sync::mpsc::sync_channel::<Result<()>>(1);
let stop_t = stop.clone();
let sync: Arc<punktfunk_core::audio::AudioSyncCell> = Arc::default();
let sync_t = sync.clone();
let thread = std::thread::Builder::new()
.name("punktfunk-audio".into())
.spawn(move || {
if let Err(e) = render_thread(pcm_rx, recycle_tx, stop_t, ready_tx, channels as u8)
if let Err(e) =
render_thread(pcm_rx, recycle_tx, stop_t, ready_tx, channels as u8, sync_t)
{
tracing::warn!(error = %format!("{e:#}"), "audio playback thread ended");
}
@@ -197,6 +203,7 @@ impl AudioPlayer {
recycle_rx,
stop,
thread: Some(thread),
sync,
})
}
Ok(Err(e)) => Err(e),
@@ -213,6 +220,12 @@ impl AudioPlayer {
self.recycle_rx.try_recv().unwrap_or_default()
}
/// The A/V sync hand-off cell — the decode thread reads the ring depth from it and posts the
/// depth the sync loop wants back through it.
pub fn sync_cell(&self) -> Arc<punktfunk_core::audio::AudioSyncCell> {
self.sync.clone()
}
/// Queue one interleaved f32 chunk (in the session's channel layout). Drops the chunk if the
/// WASAPI side is wedged (the renderer conceals the gap; never block the session pump).
pub fn push(&self, pcm: Vec<f32>) {
@@ -237,6 +250,7 @@ fn render_thread(
stop: Arc<AtomicBool>,
ready: SyncSender<Result<()>>,
channels: u8,
sync: Arc<punktfunk_core::audio::AudioSyncCell>,
) -> Result<()> {
if let Err(e) = wasapi::initialize_mta()
.ok()
@@ -315,6 +329,12 @@ fn render_thread(
}
let want = avail_frames * channels as usize;
// A/V sync: same contract as the PipeWire ring — take the decode thread's request,
// publish where the ring actually is. The policy clamps the request against its own
// underrun floor, so continuity always outranks alignment.
policy.set_sync_target(sync.target());
sync.publish_depth(ring.len());
let step = policy.step(ring.len(), want);
if step.drop_front > 0 {
sheds += 1;
+62
View File
@@ -178,6 +178,22 @@ pub struct Stats {
/// is actually going out — the muted case has its own badge, which does not need stats on.
pub mic_sent: u32,
pub mic_dropped: u32,
/// How much decoded audio is queued ahead of the speaker right now (ms) — the playback
/// ring's depth.
///
/// The audio plane used to publish nothing any surface could render: depth and target existed
/// only as a `tracing::debug!` line, and on a Steam Deck the client runs under Steam's
/// `reaper` with its stdout on a pipe, so the one number that identifies a deep ring was
/// unobtainable on the device reporting the latency. A field investigation ran to its
/// conclusion without it. That is the gap this closes.
pub audio_buffer_ms: u32,
/// The A/V sync loop's smoothed offset (ms): **positive = audio playing BEHIND the picture**,
/// negative = ahead of it. `0` before the loop has evidence, or with sync disabled.
///
/// This is the figure that says whether audio is placed correctly, and it is the one the
/// overhaul is judged by — an absolute buffer depth cannot distinguish "deep because the link
/// needs it" from "deep and therefore late".
pub audio_av_offset_ms: i32,
/// The decode path frames actually took this window (`"vaapi"`/`"software"`, empty
/// until the first frame) — the OSD's trailing tag; tracks a mid-session fallback.
pub decoder: &'static str,
@@ -1407,6 +1423,8 @@ fn pump(
},
mic_sent,
mic_dropped,
audio_buffer_ms: connector.audio_buffer_ms(),
audio_av_offset_ms: connector.audio_av_offset_ms() as i32,
decoder: dec_path,
target_kbps: connector.current_bitrate_kbps(),
auto_rate,
@@ -1523,15 +1541,59 @@ fn spawn_audio(
let mut dec = AudioDec::new(channels)
.map_err(|e| tracing::warn!(error = %e, "opus decoder failed — audio disabled"))
.ok()?;
// A/V sync (audio latency overhaul). This thread is the only place that holds all three
// ingredients at once: the packet's host capture `pts_ns`, the ring depth (via the sync cell)
// and the video plane's end-to-end figure. `pts_ns` was decoded into `AudioPacket` and then
// dropped on the floor here for the plane's entire existence, which is why audio ran at
// whatever depth its jitter ring happened to settle at and nothing ever placed it against the
// picture.
//
// The escape hatch is deliberate: a field regression in a loop that steers PLAYBACK should be
// bisectable without a rebuild, the same way `PUNKTFUNK_MIC_LEGACY_BUFFER` covers the uplink.
let av_sync_enabled = !matches!(
std::env::var("PUNKTFUNK_NO_AV_SYNC").as_deref(),
Ok("1") | Ok("true")
);
let sync_cell = player.sync_cell();
let video_e2e = connector.video_e2e_shared();
let av_offset_out = connector.audio_av_offset_shared();
let buffer_ms_out = connector.audio_buffer_ms_shared();
// Interleaved samples per ms, to report the ring depth in the unit a human reads.
let per_ms = 48 * channels.max(1) as usize;
std::thread::Builder::new()
.name("punktfunk-audio-rx".into())
.spawn(move || {
let mut pcm = vec![0f32; 5760 * channels as usize]; // scratch: max Opus frame (120 ms) × channels
let mut gaps = punktfunk_core::audio::AudioGapTracker::new();
let mut frame_samples = 0usize; // per-channel samples of the last decoded frame — the PLC unit
let mut av = punktfunk_core::audio::AvSync::new(channels);
if !av_sync_enabled {
tracing::info!("A/V sync disabled by PUNKTFUNK_NO_AV_SYNC");
}
while !stop.load(Ordering::SeqCst) {
match connector.next_audio(Duration::from_millis(100)) {
Ok(pkt) => {
// Place this frame against the picture it belongs with, BEFORE it is
// queued: `buffered_ahead` is everything that must still play first, so
// the depth read here is exactly what delays it.
let depth = sync_cell.depth();
// Published unconditionally — the ring's depth is worth seeing even with
// sync off, and it is what makes a "too much latency" report triageable.
buffer_ms_out.store((depth / per_ms) as u32, Ordering::Relaxed);
if av_sync_enabled {
let ve2e = video_e2e.load(Ordering::Relaxed);
let o = punktfunk_core::audio::AvSyncObservation {
pts_ns: pkt.pts_ns,
now_local_ns: punktfunk_core::client::now_realtime_ns(),
clock_offset_ns: connector.clock_offset_now_ns(),
buffered_ahead: depth,
// 0 = nothing on the glass yet; no reference, no correction.
video_e2e_ns: (ve2e > 0).then_some(ve2e),
};
av.observe(o);
sync_cell.set_target(av.desired_depth(depth));
av_offset_out.store(av.offset_ms() as i64, Ordering::Relaxed);
}
// Conceal lost packets (a seq gap) with libopus PLC before decoding the one
// that arrived: empty input synthesizes `frame_samples` of interpolation per
// missing packet — an inaudible fade instead of the click a hard gap makes.
+34
View File
@@ -222,6 +222,10 @@ struct StreamState {
/// Live host↔client clock offset handle (None until Connected): loaded per present so
/// mid-stream re-syncs keep the end-to-end number honest after an NTP step / drift.
clock_offset: Option<Arc<std::sync::atomic::AtomicI64>>,
/// Where the audio plane reads the video leg it must land with (ns). Published on every
/// presented frame; see the two `e2e` sites. The presenter deliberately knows nothing about
/// audio beyond writing this number.
video_e2e: Option<Arc<std::sync::atomic::AtomicU64>>,
hdr: bool,
/// The presented lane shows a PQ stream RAW — no tone-map pass ran — so the OSD badge
/// reads `HDR→SDR (raw)` instead of claiming one that never did.
@@ -391,6 +395,7 @@ impl StreamState {
profile,
latch_grid,
clock_offset: None,
video_e2e: None,
hdr: false,
hdr_untonemapped: false,
win_e2e_us: Vec::with_capacity(256),
@@ -1295,6 +1300,7 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
.ok();
gamepad.attach(c.clone());
st.clock_offset = Some(c.clock_offset_shared());
st.video_e2e = Some(c.video_e2e_shared());
// gamescope's EIS grants only a relative pointer — absolute sends
// would be dropped, so the desktop model is pinned off there. Auto
// (an older host that didn't say) stays allowed: Windows hosts and
@@ -1619,6 +1625,11 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
.max(0) as u64;
if e2e > 0 && e2e < 10_000_000_000 {
st.win_e2e_us.push(e2e / 1000);
// Hand the audio plane the figure it has to hit. This is the TRUE
// on-glass branch, so it is the best reference we can offer.
if let Some(c) = st.video_e2e.as_ref() {
c.store(e2e, Ordering::Relaxed);
}
}
st.win_disp_us
.push(s.displayed_ns.saturating_sub(s.decoded_ns) / 1000);
@@ -1995,6 +2006,13 @@ fn run_inner(mut opts: SessionOpts, mut mode: ModeCtl) -> Result<Option<Outcome>
.max(0) as u64;
if e2e > 0 && e2e < 10_000_000_000 {
st.win_e2e_us.push(e2e / 1000);
// Same hand-off as the glass-stamped branch above. This one is anchored
// on the submit instant rather than a true latch, so it UNDERSTATES the
// video leg by up to a refresh period — the audio loop's deadband is
// wider than that, which is what keeps the approximation harmless.
if let Some(c) = st.video_e2e.as_ref() {
c.store(e2e, Ordering::Relaxed);
}
}
st.win_disp_us
.push(displayed_ns.saturating_sub(decoded_ns) / 1000);
@@ -2805,6 +2823,20 @@ fn stats_text(
text.push_str(&format!(" · dropped {}", s.mic_dropped));
}
}
// The audio plane's own latency, Detailed-only. `buffer` is how much decoded audio is queued
// ahead of the speaker; `a/v` is where that PUTS it relative to the picture (+ = audio behind).
//
// Both, not just the depth: a deep ring on a jittery link is correct behaviour, and only the
// offset distinguishes that from a ring that is simply holding audio late. Before this the
// plane published neither — they lived in a `tracing::debug!` line that, on the Steam Deck,
// goes to a pipe under Steam's reaper that nobody can read, so the device that reported the
// latency was the one device where the numbers could not be seen.
if detailed && s.audio_buffer_ms > 0 {
text.push_str(&format!("\naudio buffer {} ms", s.audio_buffer_ms));
if s.audio_av_offset_ms != 0 {
text.push_str(&format!(" · a/v {:+} ms", s.audio_av_offset_ms));
}
}
// Decode integrity (M4) — the native lane's answer to "was that stream actually
// clean?". Appended LAST and only when it has something to say, which keeps it
// additive for the stdout `stats:` line's parsers (a machine interface: every
@@ -3116,6 +3148,8 @@ mod tests {
lost_pct: 0.4,
mic_sent: 0,
mic_dropped: 0,
audio_buffer_ms: 0,
audio_av_offset_ms: 0,
// The decode-path tag as the session actually spells it since M10 — the
// ladder's rung names (`NativeRung::name`), not the deleted libavcodec
// ones. A fixture carrying a tag no client emits would let this test go on
+492 -2
View File
@@ -494,6 +494,9 @@ const GROW_WINDOW_MS: u32 = 5_000;
const GROW_STEP_MS: u32 = 10;
/// Quiet time (no underrun) before a grown target relaxes one step back toward the base.
const SHRINK_QUIET_MS: u32 = 30_000;
/// The same, while the A/V sync loop is actively asking for a shallower ring — see the branch in
/// [`JitterPolicy::note_read`] that selects between them.
const SHRINK_QUIET_SYNC_MS: u32 = 5_000;
/// The playback de-jitter state machine shared by every client's audio ring.
///
@@ -531,6 +534,11 @@ pub struct JitterPolicy {
/// `want` from the most recent [`step`](Self::step), so [`note_read`](Self::note_read) can
/// advance the sample-denominated timers without the caller repeating it.
last_want: usize,
/// Depth the A/V sync loop would like, in interleaved samples ([`AvSync::desired_depth`]).
/// `None` — the default, and what every un-wired ring keeps — reproduces the pre-sync
/// behaviour exactly, which is what lets the four client rings adopt this one at a time
/// without diverging in the meantime.
sync_target: Option<usize>,
}
impl JitterPolicy {
@@ -549,9 +557,27 @@ impl JitterPolicy {
window_run: 0,
quiet_run: 0,
last_want: 0,
sync_target: None,
}
}
/// Hand the ring the depth the A/V sync loop wants ([`AvSync::desired_depth`]), or `None` to
/// run unsynchronised.
///
/// This is a REQUEST, not a command. [`effective_target`](Self::effective_target) clamps it
/// between the underrun-driven adaptive floor and the hard cap, so sync can never starve the
/// ring: if the link's jitter needs more buffer than the picture is away, the floor wins and
/// the residual shows up on the HUD instead of as a dropout. That ordering is the whole safety
/// argument for steering playback depth from a network measurement at all.
pub fn set_sync_target(&mut self, target: Option<usize>) {
self.sync_target = target;
}
/// The sync loop is asking to run shallower than the adaptive target has grown to.
fn sync_wants_less(&self) -> bool {
self.sync_target.is_some_and(|s| s < self.target)
}
/// The live target depth in ms (grows under underrun pressure; never below the base).
pub fn target_ms(&self) -> u32 {
(self.target / self.per_ms) as u32
@@ -577,7 +603,25 @@ impl JitterPolicy {
/// quantum, a legacy AAudio path) lifts it to `want` plus one protocol frame rather than
/// oscillating prime → dropout → re-prime forever.
fn effective_target(&self, want: usize) -> usize {
self.target.max(want + FRAME_MS as usize * self.per_ms)
let floor = self.target.max(want + FRAME_MS as usize * self.per_ms);
match self.sync_target {
// Continuity outranks sync — see `set_sync_target`. The loop may pull the ring
// shallower to catch the picture up, or push it deeper when audio runs early, but
// never below what underrun pressure has proven this link needs, and never past the
// hard cap that bounds added latency.
//
// The ceiling is raised to the floor rather than passed to `clamp` as-is: a device
// whose callback quantum alone exceeds the preset's `hard_cap_ms` makes `floor > cap`,
// and `Ord::clamp` PANICS when min > max. That would be a panic in a realtime audio
// callback on exactly the awkward hardware this code exists to survive — and the same
// reasoning `step` already applies when it computes its own cap with `.max(target +
// want)`.
Some(s) => {
let cap = (self.tuning.hard_cap_ms as usize * self.per_ms).max(floor);
s.clamp(floor, cap)
}
None => floor,
}
}
/// Decide this callback: what to trim, and whether to play. Call BEFORE reading, with the
@@ -664,7 +708,19 @@ impl JitterPolicy {
} else {
self.empties = 0;
self.quiet_run += want;
if self.quiet_run >= SHRINK_QUIET_MS as usize * self.per_ms {
// A grown target normally relaxes only after a long quiet spell, because without other
// evidence the only thing that can justify giving up hard-won slack is time. When the
// sync loop is asking to run shallower it IS that evidence — a measurement saying the
// extra depth is costing alignment right now — so test a smaller target sooner. Wrong
// guesses are cheap and self-correcting: one underrun and the growth path takes it
// straight back. Without this a ring that ratcheted to the ceiling during a transient
// would hold the audio a ceiling's worth late for minutes after the cause had gone.
let quiet_needed = if self.sync_wants_less() {
SHRINK_QUIET_SYNC_MS
} else {
SHRINK_QUIET_MS
};
if self.quiet_run >= quiet_needed as usize * self.per_ms {
// Long quiet spell: give a grown target one step back, so a single bad minute
// doesn't cost latency for the rest of the session.
self.quiet_run = 0;
@@ -750,6 +806,205 @@ pub fn spa_positions(channels: u8) -> &'static [u32] {
}
}
/// The lock-free hand-off between the thread that knows the TIMESTAMPS (the decode/pull thread,
/// which sees each packet's `pts_ns`) and the one that knows the RING (the realtime audio
/// callback, which owns the depth and the [`JitterPolicy`]). Neither can do the job alone and the
/// callback must not block, so they trade two words.
///
/// `usize::MAX` encodes "no target" rather than `0`, because `0` is a perfectly ordinary depth to
/// ask for and conflating the two would silently mean "run the ring dry".
#[derive(Debug)]
pub struct AudioSyncCell {
depth: std::sync::atomic::AtomicUsize,
target: std::sync::atomic::AtomicUsize,
}
impl Default for AudioSyncCell {
fn default() -> Self {
AudioSyncCell {
depth: std::sync::atomic::AtomicUsize::new(0),
target: std::sync::atomic::AtomicUsize::new(usize::MAX),
}
}
}
impl AudioSyncCell {
/// Callback side: publish the ring's current depth in interleaved samples.
pub fn publish_depth(&self, depth: usize) {
self.depth
.store(depth, std::sync::atomic::Ordering::Relaxed);
}
/// Decode side: the ring depth as last seen by the audio callback.
pub fn depth(&self) -> usize {
self.depth.load(std::sync::atomic::Ordering::Relaxed)
}
/// Decode side: ask the ring to aim for this depth (`None` = run unsynchronised).
pub fn set_target(&self, target: Option<usize>) {
self.target.store(
target.unwrap_or(usize::MAX),
std::sync::atomic::Ordering::Relaxed,
);
}
/// Callback side: the depth the sync loop wants, if any.
pub fn target(&self) -> Option<usize> {
match self.target.load(std::sync::atomic::Ordering::Relaxed) {
usize::MAX => None,
t => Some(t),
}
}
}
/// Smoothing time constant for the measured A/V offset, in ms of consumed audio. Long enough that
/// network jitter and a single late datagram do not move it; short enough to track real drift.
const AV_EWMA_TAU_MS: u32 = 2_000;
/// Offsets inside this band are left alone. Correcting a few ms costs a (crossfaded, but real)
/// discontinuity and buys nothing a listener can perceive — detectability for A/V misalignment sits
/// an order of magnitude above it. The deadband is what keeps the loop from hunting forever around
/// zero, which would be audible in a way the misalignment it was chasing was not.
const AV_DEADBAND_MS: u32 = 10;
/// Observations folded before the first correction is offered. The offset is derived from a clock
/// skew estimate and a video figure that both need a moment to settle after connect; acting on the
/// first sample would chase the handshake, not the stream.
const AV_MIN_OBSERVATIONS: u32 = 100;
/// An offset larger than this is not believed. A wall-clock step, a paused host, or a stale video
/// figure can all produce an enormous apparent misalignment, and steering the ring by it would
/// empty or overfill it outright. Beyond this the loop reports and waits rather than acting.
const AV_SANE_LIMIT_MS: u32 = 1_000;
/// The A/V synchronisation controller: turns "when will this audio actually play" and "when did the
/// picture it belongs with reach the glass" into a ring depth the [`JitterPolicy`] should aim for.
///
/// **The defect it exists to fix.** The host stamps `pts_ns` on every audio datagram and the client
/// decoded it into `AudioPacket` — and then never read it. Video's `pts_ns`, by contrast, is used
/// end to end (the presenter computes a true glass-to-glass `displayed + clock_offset pts`). So
/// audio free-ran at whatever depth its jitter ring happened to settle at, video was presented on a
/// wholly independent path, and nothing ever compared them: the A/V offset was an accident of
/// buffer depths. It moved whenever the ring ratcheted under underrun pressure, and — the way this
/// surfaced in the field — it got WORSE every time video got faster, because a quicker decoder
/// lowers the video leg while leaving the audio leg exactly where it was.
///
/// **Video is the master.** In a game streamer the video leg is the input-feel budget and must
/// never be inflated to satisfy the audio clock; audio tolerates small, crossfaded, rate-limited
/// corrections that are inaudible, and [`crossfade_drop`] already applies them. So audio moves.
///
/// **Continuity outranks sync.** This type only ever proposes a depth. [`JitterPolicy`] clamps the
/// proposal to its own underrun-driven floor, so a link whose jitter genuinely needs more buffer
/// than the picture is away keeps its buffer and the residual is reported instead of being taken
/// out of the listener's stream. See [`JitterPolicy::set_sync_target`].
#[derive(Clone, Debug)]
pub struct AvSync {
/// Interleaved samples per millisecond at the negotiated layout (48 × channels).
per_ms: usize,
/// EWMA of the measured offset in ns. Positive = audio is scheduled to play LATE relative to
/// the picture it belongs with.
offset_avg_ns: f32,
observations: u32,
/// Set once an observation lands outside [`AV_SANE_LIMIT_MS`], for reporting.
implausible: bool,
}
/// One measurement handed to [`AvSync::observe`]. Every field is in the units its source already
/// produces, so no caller has to do clock arithmetic to use it correctly.
#[derive(Clone, Copy, Debug)]
pub struct AvSyncObservation {
/// The host capture timestamp carried by the audio frame being queued (host clock).
pub pts_ns: u64,
/// Local wall-clock now, same basis the client's video latency math uses (CLOCK_REALTIME).
pub now_local_ns: i128,
/// Host clock minus client clock, from the skew handshake (`clock_offset_now_ns`).
pub clock_offset_ns: i64,
/// How much audio is already queued AHEAD of this frame, in interleaved samples — everything
/// that must play before it does.
pub buffered_ahead: usize,
/// The video plane's current end-to-end figure in ns: `displayed + clock_offset pts`, as the
/// presenter already computes it. `None` while no frame has been presented yet.
pub video_e2e_ns: Option<u64>,
}
impl AvSync {
/// `channels` is the negotiated interleaved channel count (2/6/8).
pub fn new(channels: u8) -> AvSync {
AvSync {
per_ms: (SAMPLE_RATE_HZ / 1000) as usize * channels.max(1) as usize,
offset_avg_ns: 0.0,
observations: 0,
implausible: false,
}
}
/// Fold one measurement. Returns the smoothed offset in ns once there is enough evidence to
/// believe it (positive = audio late), or `None` while still settling.
///
/// Rejecting the implausible rather than clamping it is deliberate: a wall-clock step or a
/// stale video figure produces a huge apparent offset, and a clamped-but-wrong value would be
/// acted on as though it were a small real one.
pub fn observe(&mut self, o: AvSyncObservation) -> Option<i64> {
// No frame on the glass yet ⇒ no reference to align against, so nothing to say.
let video_e2e_ns = o.video_e2e_ns?;
// When this frame's samples will actually reach the speaker, expressed in the host's
// capture clock — the same clock, and the same shape, as the video figure it is compared
// against.
let buffered_ns = (o.buffered_ahead / self.per_ms.max(1)) as i128 * 1_000_000;
let play_at_host = o.now_local_ns + buffered_ns + o.clock_offset_ns as i128;
let audio_e2e_ns = play_at_host - o.pts_ns as i128;
let offset_ns = audio_e2e_ns - video_e2e_ns as i128;
if offset_ns.unsigned_abs() > (AV_SANE_LIMIT_MS as u128) * 1_000_000 {
self.implausible = true;
return None;
}
self.implausible = false;
// Weight by one protocol frame so the time constant means the same thing regardless of how
// often the caller observes.
let alpha = (FRAME_MS as f32 / AV_EWMA_TAU_MS as f32).clamp(0.0, 1.0);
if self.observations == 0 {
self.offset_avg_ns = offset_ns as f32;
} else {
self.offset_avg_ns += (offset_ns as f32 - self.offset_avg_ns) * alpha;
}
self.observations = self.observations.saturating_add(1);
self.settled().then_some(self.offset_avg_ns as i64)
}
/// Enough evidence folded to act on.
pub fn settled(&self) -> bool {
self.observations >= AV_MIN_OBSERVATIONS
}
/// The smoothed offset in ms (positive = audio late), for the HUD. Reported as soon as it is
/// measured, including while still settling — a number the operator can watch converge is more
/// useful than a blank that hides whether the loop is working at all.
pub fn offset_ms(&self) -> i32 {
(self.offset_avg_ns / 1_000_000.0) as i32
}
/// The last observation was outside the believable range and was discarded.
pub fn implausible(&self) -> bool {
self.implausible
}
/// The ring depth that would place audio with the picture, given where the ring is now.
/// `None` while unsettled or inside the deadband — the caller then leaves the policy alone.
///
/// Audio late (offset > 0) means there is too much queued: aim shallower. Audio early means
/// aim deeper.
pub fn desired_depth(&self, current_depth: usize) -> Option<usize> {
if !self.settled() {
return None;
}
let offset_ms = self.offset_avg_ns / 1_000_000.0;
if offset_ms.abs() < AV_DEADBAND_MS as f32 {
return None;
}
let delta = (offset_ms * self.per_ms as f32) as i64;
Some((current_depth as i64 - delta).max(0) as usize)
}
}
#[cfg(test)]
mod tests {
use super::*;
@@ -1447,4 +1702,239 @@ mod tests {
}
}
}
// ---- A/V sync (audio latency overhaul) ----------------------------------------------
/// Build an observation whose measured offset is exactly `offset_ms` (positive = audio late).
fn obs(offset_ms: i64, depth: usize, per_ms: usize) -> AvSyncObservation {
// audio_e2e = buffered + (now + skew - pts); pin now/skew/pts so the only free term is the
// buffered depth, then choose video_e2e so the difference lands on `offset_ms`.
let buffered_ms = (depth / per_ms) as i64;
let audio_e2e_ms = buffered_ms + 40; // 40 ms of transport, arbitrary but fixed
let video_e2e_ms = audio_e2e_ms - offset_ms;
AvSyncObservation {
pts_ns: 1_000_000_000,
now_local_ns: 1_000_000_000i128 + 40 * 1_000_000,
clock_offset_ns: 0,
buffered_ahead: depth,
video_e2e_ns: Some((video_e2e_ms.max(0) as u64) * 1_000_000),
}
}
fn settle(sync: &mut AvSync, offset_ms: i64, depth: usize, per_ms: usize, n: u32) {
for _ in 0..n {
sync.observe(obs(offset_ms, depth, per_ms));
}
}
#[test]
fn av_sync_needs_evidence_before_acting() {
let pm = per_ms(2);
let mut s = AvSync::new(2);
// One sample is never enough — the skew estimate and the video figure both settle after
// connect, and acting on the first would chase the handshake.
assert!(s.observe(obs(50, 30 * pm, pm)).is_none());
assert!(!s.settled());
assert!(s.desired_depth(30 * pm).is_none());
settle(&mut s, 50, 30 * pm, pm, AV_MIN_OBSERVATIONS);
assert!(s.settled(), "should act once the evidence is in");
}
#[test]
fn av_sync_aims_shallower_when_audio_is_late() {
let pm = per_ms(2);
let depth = 60 * pm;
let mut s = AvSync::new(2);
settle(&mut s, 40, depth, pm, AV_MIN_OBSERVATIONS * 4);
let want = s
.desired_depth(depth)
.expect("a 40 ms offset is actionable");
assert!(
want < depth,
"audio late must aim shallower: {want} vs {depth}"
);
// The correction is the offset, not a guess at it.
let shed_ms = (depth - want) / pm;
assert!(
(35..=45).contains(&shed_ms),
"should aim to shed ~40 ms, got {shed_ms}"
);
}
#[test]
fn av_sync_aims_deeper_when_audio_is_early() {
let pm = per_ms(2);
let depth = 20 * pm;
let mut s = AvSync::new(2);
settle(&mut s, -30, depth, pm, AV_MIN_OBSERVATIONS * 4);
let want = s
.desired_depth(depth)
.expect("a 30 ms offset is actionable");
assert!(
want > depth,
"audio early must aim deeper: {want} vs {depth}"
);
}
#[test]
fn av_sync_deadbands_what_no_one_can_hear() {
let pm = per_ms(2);
let depth = 30 * pm;
let mut s = AvSync::new(2);
settle(
&mut s,
(AV_DEADBAND_MS - 2) as i64,
depth,
pm,
AV_MIN_OBSERVATIONS * 4,
);
assert!(
s.desired_depth(depth).is_none(),
"an offset inside the deadband must not provoke a (real, if crossfaded) discontinuity"
);
}
#[test]
fn av_sync_rejects_the_implausible_instead_of_clamping_it() {
let pm = per_ms(2);
let depth = 30 * pm;
let mut s = AvSync::new(2);
settle(&mut s, 30, depth, pm, AV_MIN_OBSERVATIONS * 4);
let before = s.offset_ms();
// A wall-clock step / stale video figure. Built directly rather than through `obs`: that
// helper floors the video figure at zero, which would cap the offset at a merely LARGE
// value and let this test pass without ever exercising the rejection.
let wild = AvSyncObservation {
pts_ns: 0,
now_local_ns: 5_000_000_000,
clock_offset_ns: 0,
buffered_ahead: depth,
video_e2e_ns: Some(40_000_000),
};
assert!(s.observe(wild).is_none());
assert!(s.implausible(), "a ~5 s offset must be refused, not folded");
assert_eq!(
before,
s.offset_ms(),
"an implausible sample must be discarded, not folded in"
);
}
#[test]
fn sync_can_never_starve_the_ring() {
// THE safety invariant: sync only ever proposes. Continuity — the underrun-driven floor —
// outranks it on every preset, or a lossy link would be "synced" into dropouts.
for (name, t) in [
("PIPEWIRE", JitterTuning::PIPEWIRE),
("WASAPI", JitterTuning::WASAPI),
("COREAUDIO", JitterTuning::COREAUDIO),
("AAUDIO", JitterTuning::AAUDIO),
] {
let pm = per_ms(2);
let want = 5 * pm;
let mut p = JitterPolicy::new(t, 2);
let floor = p.effective_target(want);
// Ask for an absurdly shallow ring — zero.
p.set_sync_target(Some(0));
assert_eq!(
p.effective_target(want),
floor,
"{name}: sync pulled the target below the continuity floor"
);
// And it may not blow past the hard cap either.
p.set_sync_target(Some(usize::MAX / 2));
assert!(
p.effective_target(want) <= t.hard_cap_ms as usize * pm,
"{name}: sync pushed the target past the hard cap"
);
}
}
#[test]
fn a_huge_device_quantum_does_not_panic_the_clamp() {
// `Ord::clamp` panics when min > max. A device whose callback quantum alone exceeds the
// preset's hard cap pushes the continuity floor above the ceiling, and this runs inside a
// realtime audio callback — so the ceiling yields to the floor instead.
let t = JitterTuning::PIPEWIRE; // hard_cap 80 ms
let pm = per_ms(2);
let want = 500 * pm; // a 500 ms quantum: absurd, but not a reason to abort the process
let mut p = JitterPolicy::new(t, 2);
p.set_sync_target(Some(0));
let target = p.effective_target(want); // must not panic
assert!(
target >= want,
"the target must still be able to serve one callback"
);
}
#[test]
fn no_sync_target_leaves_the_policy_exactly_as_it_was() {
// The four rings adopt sync one at a time; an un-wired ring must behave bit-identically to
// before. `None` is the default, so this also pins the constructor.
let t = JitterTuning::PIPEWIRE;
let pm = per_ms(2);
let want = 5 * pm;
let mut a = JitterPolicy::new(t, 2);
let mut b = JitterPolicy::new(t, 2);
b.set_sync_target(None);
assert_eq!(a.effective_target(want), b.effective_target(want));
for depth_ms in [0usize, 5, 15, 30, 60, 90, 200] {
let sa = a.step(depth_ms * pm, want);
let sb = b.step(depth_ms * pm, want);
assert_eq!(sa, sb, "depth {depth_ms} ms diverged with an explicit None");
a.note_read(sa.silence);
b.note_read(sb.silence);
}
}
#[test]
fn sync_pressure_relaxes_a_grown_target_sooner_than_time_alone() {
// A ring that ratcheted during a transient must not hold audio late for minutes after the
// cause is gone. With sync asking for less, the relax window is the short one.
let t = JitterTuning::PIPEWIRE;
let pm = per_ms(2);
let want = 5 * pm;
let grow = |p: &mut JitterPolicy| {
// Drive underruns until the target has grown above the base. Each round hands `step` a
// DEEP ring first: `note_read` ignores everything while un-primed (a priming silence is
// not an underrun), and `deprime_after` short reads in a row un-prime the ring — so
// hammering a zero-depth ring would report nothing and grow nothing, forever.
for _ in 0..10_000 {
if p.target_ms() > t.base_target_ms {
return;
}
p.step(200 * pm, want); // (re-)prime
p.note_read(true); // then one genuine short read
}
panic!("the adaptive floor never grew — the test cannot measure a relax");
};
// Quiet reads needed to relax one step, with and without sync pressure.
let quiet_to_relax = |p: &mut JitterPolicy| -> usize {
let start = p.target_ms();
let mut reads = 0usize;
while p.target_ms() == start && reads < 200_000 {
p.step(60 * pm, want);
p.note_read(false);
reads += 1;
}
reads
};
let mut slow = JitterPolicy::new(t, 2);
grow(&mut slow);
slow.set_sync_target(None);
let slow_reads = quiet_to_relax(&mut slow);
let mut fast = JitterPolicy::new(t, 2);
grow(&mut fast);
// Ask for something strictly shallower than the grown target.
fast.set_sync_target(Some(pm));
let fast_reads = quiet_to_relax(&mut fast);
assert!(
fast_reads < slow_reads,
"sync pressure should relax sooner: {fast_reads} vs {slow_reads} quiet reads"
);
}
}
+62 -1
View File
@@ -297,6 +297,26 @@ pub struct NativeClient {
/// the pump's first no-op clock flush). Shared with the pump and, via
/// [`clock_offset_shared`](Self::clock_offset_shared), with embedder latency-math threads.
clock_offset: Arc<AtomicI64>,
/// The video plane's live end-to-end latency in ns — `displayed + clock_offset pts`, the
/// figure the presenter already computes per frame (with a TRUE on-glass stamp where
/// `VK_KHR_present_wait` is available, and the submit instant otherwise). `0` = nothing
/// presented yet.
///
/// Written by whoever puts frames on the glass; read by the audio plane, which steers its ring
/// depth to land audio WITH the picture ([`crate::audio::AvSync`]). It lives here, next to
/// `clock_offset`, because those two are exactly the pair a synchroniser needs and neither
/// plane owns the other: the presenter must not know about audio, and the audio thread cannot
/// see the glass.
video_e2e_ns: Arc<AtomicU64>,
/// The A/V sync loop's smoothed offset in ms — positive = audio playing LATE relative to the
/// picture. Written by the audio thread, read by the stats HUD. The audio plane used to
/// publish NOTHING a surface could render (its depth and target existed only as a
/// `tracing::debug!` line, which on a Deck goes into a pipe under Steam's reaper that nobody
/// can read), so a latency report had no instrument behind it at all.
audio_av_offset_ms: Arc<AtomicI64>,
/// Decoded audio queued ahead of the speaker (ms) — the playback ring's depth, as last seen by
/// the audio callback. Written by the audio thread, read by the stats HUD.
audio_buffer_ms: Arc<AtomicU32>,
/// Decode-stage latency samples from the embedder ([`report_decode_us`](Self::report_decode_us)),
/// drained per window by the data-plane pump to feed the adaptive-bitrate controller's decode
/// signal. Shared with the pump; see [`DecodeLatAcc`].
@@ -408,7 +428,13 @@ fn pin_thread_user_interactive() {}
/// Wall-clock now in nanoseconds (CLOCK_REALTIME basis), to compare against the host-stamped
/// capture `pts_ns` after the skew offset is applied — the same latency math the stats HUDs use.
fn now_realtime_ns() -> i128 {
///
/// Public because the A/V sync loop ([`crate::audio::AvSync`]) lives in an embedder crate but must
/// read the clock in EXACTLY this basis: its whole output is a difference between a local instant
/// and a host `pts_ns`, so a caller reaching for `Instant` or a monotonic clock instead would get a
/// plausible-looking number that is wrong by the machine's boot time. Exporting the one correct
/// clock is cheaper than documenting which clocks are incorrect.
pub fn now_realtime_ns() -> i128 {
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_nanos() as i128)
@@ -553,6 +579,9 @@ impl NativeClient {
let mic_stats = Arc::new(MicUplinkCounters::default());
let hot_tids = Arc::new(Mutex::new(Vec::new()));
let clock_offset = Arc::new(AtomicI64::new(0));
let video_e2e_ns = Arc::new(AtomicU64::new(0));
let audio_av_offset_ms = Arc::new(AtomicI64::new(0));
let audio_buffer_ms = Arc::new(AtomicU32::new(0));
let decode_lat = Arc::new(Mutex::new(DecodeLatAcc::default()));
// Seeded by the pump from the Welcome (before ready_tx), then follows every ack.
let live_bitrate = Arc::new(AtomicU32::new(0));
@@ -698,6 +727,9 @@ impl NativeClient {
rfi: Mutex::new(RfiRecovery::default()),
hot_tids,
clock_offset,
video_e2e_ns,
audio_av_offset_ms,
audio_buffer_ms,
decode_lat,
live_bitrate_kbps: live_bitrate,
// The controller arms exactly when the pump does — all three terms, not two: Automatic
@@ -971,6 +1003,35 @@ impl NativeClient {
self.clock_offset.clone()
}
/// The shared cell carrying the video plane's end-to-end latency (ns, `0` = nothing presented
/// yet). The presenter WRITES it once per presented frame; the audio plane READS it to place
/// its samples with the picture. See the field docs on `video_e2e_ns`.
pub fn video_e2e_shared(&self) -> Arc<AtomicU64> {
self.video_e2e_ns.clone()
}
/// The cell carrying the A/V sync loop's smoothed offset in ms (positive = audio late).
/// Written by the audio thread; read by the HUD.
pub fn audio_av_offset_shared(&self) -> Arc<AtomicI64> {
self.audio_av_offset_ms.clone()
}
/// The A/V sync offset the audio plane last measured, in ms. Positive = audio is playing
/// behind the picture. `0` before the loop has evidence, or when sync is off.
pub fn audio_av_offset_ms(&self) -> i64 {
self.audio_av_offset_ms.load(Ordering::Relaxed)
}
/// The cell carrying the playback ring's depth in ms. Written by the audio thread.
pub fn audio_buffer_ms_shared(&self) -> Arc<AtomicU32> {
self.audio_buffer_ms.clone()
}
/// Decoded audio queued ahead of the speaker, in ms.
pub fn audio_buffer_ms(&self) -> u32 {
self.audio_buffer_ms.load(Ordering::Relaxed)
}
/// Report one decoded frame's decode-stage latency, in microseconds: the wall-clock elapsed from
/// the access unit leaving [`next_frame`](Self::next_frame) to its decoded output becoming
/// available (dequeued from the decoder). This feeds the "Automatic" bitrate controller's decode
+43 -4
View File
@@ -352,6 +352,12 @@ struct MicUserData {
/// bursting out as stale audio when recording (re)starts.
const MIC_STALE: Duration = Duration::from_secs(1);
/// The graph quantum every punktfunk PipeWire stream asks for, in frames: 240 @ 48 kHz = 5 ms,
/// one protocol audio frame. Named so the `NODE_LATENCY` request and the code that CHECKS whether
/// the request was honoured cannot drift apart — the check is only meaningful while it compares
/// against the same number the ask used.
const CAPTURE_QUANTUM_FRAMES: u32 = 240;
fn mic_pw_thread(
pcm_rx: Receiver<(std::time::Instant, Vec<f32>)>,
quit_rx: pipewire::channel::Receiver<Terminate>,
@@ -722,12 +728,19 @@ fn pw_thread(
channels: u32,
stats: crate::audio::capture_policy::CaptureStats,
last_stats: std::time::Instant,
/// Whether this OPEN has reported its negotiated buffer size yet. Per-open, not the
/// process-wide `static AtomicBool` this replaces: a host runs for days across many
/// sessions, so the old form reported the very first capture of the process and then
/// never again — the one number that identifies a clamped quantum, invisible on every
/// subsequent open (including every reopen after a device change).
reported_quantum: bool,
}
let ud = CapUd {
tx,
channels,
stats: Default::default(),
last_stats: std::time::Instant::now(),
reported_quantum: false,
};
let _listener = stream
.add_local_listener_with_user_data(ud)
@@ -788,10 +801,36 @@ fn pw_thread(
let region = &buf[offset..(offset + size).min(buf.len())];
// Negotiated as F32LE; reinterpret the byte region as interleaved f32.
let n = region.len() / 4;
static FIRST: std::sync::atomic::AtomicBool =
std::sync::atomic::AtomicBool::new(true);
if FIRST.swap(false, std::sync::atomic::Ordering::Relaxed) {
tracing::info!(samples = n, "audio first capture buffer");
if !ud.reported_quantum {
ud.reported_quantum = true;
// What we ASKED for vs what PipeWire actually handed us. Stating only the
// result ("samples=2048") reads as a fact about the device; stating it
// next to the request is what makes a clamp legible. A VM is the common
// cause — stock `pipewire.conf` raises `default.clock.min-quantum` to
// 1024 whenever `cpu.vm.name` is set, so a 5 ms ask silently becomes
// 21.3 ms and the audio plane starts arriving in bursts. That cost a
// whole field investigation to find; it should cost one log line.
let frames = n / (ud.channels.max(1) as usize);
let want = CAPTURE_QUANTUM_FRAMES as usize;
if frames > want {
tracing::warn!(
requested_frames = want,
negotiated_frames = frames,
negotiated_ms =
format!("{:.1}", frames as f32 * 1000.0 / SAMPLE_RATE as f32),
"the audio graph refused our low-latency quantum — capture arrives \
in bursts this size, and the client must buffer at least that \
much to play them smoothly. On a VM this is PipeWire's \
`default.clock.min-quantum = 1024` rule; check \
`pw-metadata -n settings`"
);
} else {
tracing::info!(
requested_frames = want,
negotiated_frames = frames,
"audio capture quantum negotiated"
);
}
}
let mut samples = Vec::with_capacity(n);
for i in 0..n {
+59 -1
View File
@@ -89,6 +89,16 @@ pub(super) fn audio_thread(
use crate::audio::SAMPLE_RATE;
const FRAME_MS: usize = 5;
const SAMPLES_PER_FRAME: usize = SAMPLE_RATE as usize * FRAME_MS / 1000; // 240
/// One protocol frame of wall time — the cadence paced sends aim for.
const FRAME_INTERVAL: std::time::Duration = std::time::Duration::from_millis(FRAME_MS as u64);
/// Ceiling on a single pacing sleep. The capture channel is finite and `next_chunk` has to be
/// serviced; sleeping past a couple of frames would trade a burst on the wire for a drop at
/// the capturer, which is strictly worse (a drop is a click AND a permanent shift).
const PACE_MAX_SLEEP: std::time::Duration = std::time::Duration::from_millis(10);
/// How far behind schedule the pacer may fall before it stops trying to catch up and simply
/// re-anchors. Chasing an old schedule after a stall would send a burst — the exact thing
/// pacing exists to prevent — so past this point the debt is forgiven, not repaid.
const PACE_REANCHOR: std::time::Duration = std::time::Duration::from_millis(100);
let want = punktfunk_core::audio::normalize_channels(channels);
// Tier and redundancy are ONE decision, budgeted against the session's video bitrate — see
// `handshake::audio_budget`. An unparseable `audio.quality` was already warned about there
@@ -151,6 +161,27 @@ pub(super) fn audio_thread(
// continuity breaks (a capture reopen), so we never advertise a predecessor the client's
// sequence numbering does not agree with.
let mut prev_frame: Vec<u8> = Vec::new();
// W1.1/W1.2 — the audio SAMPLE clock, and the schedule frames leave on.
//
// `pts_ns` used to be `now_ns()` evaluated inside the drain loop below, which made it the
// instant we got round to ENCODING rather than the instant the samples were CAPTURED. Every
// frame carved out of one capture chunk therefore carried a near-identical timestamp, and the
// value drifted with encoder scheduling. That was harmless only for as long as nothing
// consumed it; a client-side A/V sync loop regulating against it would be regulating against
// a fiction, so this is a prerequisite for the whole overhaul, not a tidy-up.
//
// `pace_due` exists because a chunk is not a frame. A capture callback hands us a whole
// quantum (5 ms when the graph honours our ask, 21.3 ms on a VM that clamps it to 1024 —
// see `audio::linux`'s quantum warning), and the old loop drained all of it into
// back-to-back `send_datagram` calls. The wire then carried a 4-5 frame burst followed by
// ~21 ms of nothing, and a client ring can only absorb that by standing at least a burst
// period deep. Releasing frames on the audio clock instead costs no AVERAGE latency — the
// client was buffering those frames anyway — and removes the burst the ring was sized for.
// Uninitialised on purpose: every read is preceded by the re-anchor at the top of the chunk
// loop, and seeding it with a placeholder would just be a value the compiler correctly points
// out is never read.
let mut next_pts_ns: u64;
let mut pace_due: Option<std::time::Instant> = None;
if capturer.is_some() {
tracing::info!(
channels = want,
@@ -194,10 +225,37 @@ pub(super) fn audio_thread(
continue;
}
};
// Anchor the sample clock on THIS chunk's arrival. PipeWire hands us a buffer of already
// captured audio, so the newest sample in `acc` is ~now and the oldest is one whole
// buffer-occupancy earlier. Re-deriving the anchor every chunk (rather than free-running
// a counter) keeps the stamp tied to the capture device's own cadence, so a drifting or
// resampling graph corrects itself instead of accumulating error over a long session.
let arrival_ns = now_ns();
acc.extend_from_slice(&chunk);
let queued_frames = (acc.len() / want as usize) as u64;
next_pts_ns = arrival_ns.saturating_sub(queued_frames * 1_000_000_000 / SAMPLE_RATE as u64);
while acc.len() >= frame_len {
// Hold each frame until its slot on the audio clock. The FIRST frame of a chunk is
// already due (its samples are the oldest we hold), so this only ever delays the
// tail of a multi-frame chunk — exactly the burst we are trying not to send. A
// schedule that has fallen more than one frame behind is re-anchored rather than
// chased, so a scheduling hiccup cannot turn into a permanent send-time debt.
let now = std::time::Instant::now();
match pace_due {
Some(due) if due > now => {
let wait = due - now;
// Never sleep longer than the audio we are holding: `next_chunk` has to be
// serviced or the capture channel backs up and starts dropping.
std::thread::sleep(wait.min(PACE_MAX_SLEEP));
}
Some(due) if now.duration_since(due) > PACE_REANCHOR => pace_due = None,
_ => {}
}
pace_due = Some(pace_due.unwrap_or_else(std::time::Instant::now) + FRAME_INTERVAL);
let frame: Vec<f32> = acc.drain(..frame_len).collect();
let pts_ns = now_ns();
let pts_ns = next_pts_ns;
next_pts_ns += FRAME_MS as u64 * 1_000_000;
match enc.encode_float(&frame, &mut opus_buf) {
Ok(n) => {
let opus = &opus_buf[..n];
+12
View File
@@ -73,6 +73,7 @@ e2e 14.2/19.8 ms (p50/p95) · host 3.1 · net 6.7 · decode 2.1 · display 2.3 m
host: queue 0.6 · encode 1.8 · xfer 0.2 · pace 0.5 ms
present: mailbox
lost 3 (2.4%)
audio buffer 28 ms · a/v +4 ms
```
Android (headline and `display` both floor-shaved, like the Apple clients — the raw
@@ -85,6 +86,7 @@ HEVC · 10-bit · HDR (BT.2020 PQ) · 4:2:0
end-to-end 14.2 ms p50 · 19.8 p95 · capture→displayed
= host 3.1 + network 6.7 + decode 2.1 + display 2.3 · presents 119
os present +16.7 excluded (display pipeline minimum)
audio buffer 28 ms · a/v +4 ms
lost 3 (2.4%) · skipped 1 · FEC 12
```
@@ -185,6 +187,16 @@ lost 3 (2.4%)
(frames your client chose not to display because a newer one had already arrived) and
`FEC` (packet shards the error correction recovered this second — loss you *didn't*
feel) are reported by the **Android client only**; the other clients show `lost` alone.
- **The audio line** — Detailed only, on Linux · Windows · Steam Deck · Android, and shown
once sound is actually playing. `audio buffer` is how much decoded audio is queued ahead
of your speakers; `a/v` is where that *puts* it relative to the picture — **positive means
audio is playing behind the picture**, negative means ahead of it. The client steers the
buffer to drive `a/v` toward zero, but never below the depth your link's jitter needs, so
on a rough connection you may see the buffer hold and a small `a/v` remain: that is the
client choosing an unbroken stream over perfect lip-sync, and it is the honest reading
rather than a hidden compromise. The `a/v` term is omitted when it is zero — aligned, or
not yet measured (it needs a frame on screen to compare against, and a few seconds to
settle). The Apple clients do not report it yet.
All values refresh once per second over the last second of frames.