forked from unom/punktfunk
Phase 0.4 host half: PUNKTFUNK_PERF now splits the send thread per window into fec/seal/sock (SealPerf via Session::take_seal_perf; the paced video path folds its chunk-send time in through note_sock_ns), logged with per-packet ns in the send loop's perf line. Measured on .21 at 2.5 Gbps offered: fec ~100 ns/pkt (Phase 1.4 landed), seal ~1000 ns/pkt = 21.5% of a core, sock ~1400 ns/pkt — the Phase 1.5 gate (seal > ~15% of the thread at 2 Gbps) trips. Phase 1.5: seal_frame_inner is now write-then-seal — packetize writes every packet's plaintext at its final wire offset, then a frame of >= 256 wire packets (~300 KB) splits the AES-GCM pass across two lanes: a persistent punktfunk-seal2 worker (lazy-spawned, rendezvous channels, no per-frame spawn, zero steady-state allocs via a reused hand-off Vec) seals the back half under nonces seq_base+i while the send thread seals the front. Nonce order is deterministic per shard index, so the wire is byte-identical to the sequential pass — pinned by the wire-equivalence test, now including a 469-packet frame plus an assertion that the lane actually spawned. Small frames and the probe's ~17-packet AUs stay single-lane; PUNKTFUNK_SEAL_LANES=1 forces single-lane. Validated: 84 core tests + workspace suites + clippy -D warnings on .21. Halves the seal wall-clock on big frames — headroom for the 10G pair's ~4.8 Gbps ceiling (seal alone would be ~47% of a core there) and PyroWave 4K rates. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>