AMD/Intel needed no new encoder code for HDR — the VAAPI path already ingests
an XR30 dmabuf into `format=p010:out_color_matrix=bt2020`. NVIDIA did: the
10-bit formats were excluded from the GPU import outright, so an HDR session
fell back to a CPU readback plus swscale, which is the one thing the capture
path is not allowed to ship.
It turns out no CSC kernel is needed. NVENC ingests packed 10-bit RGB natively
as `ARGB10`/`ABGR10` and does the conversion itself following the configured
VUI matrix — which `apply_low_latency_config` already sets to BT.2020 NCL for
an HDR session. So the frame travels LINEAR dmabuf → Vulkan bridge → CUDA →
NVENC unconverted: no host CSC pass, no depth loss, no extra work on a
contended SM.
* invariant 1 is restated rather than dropped: HDR must never take the TILED
EGL de-tile blit (it renders into an 8-bit `GL_RGBA8` texture). The HDR pods
are LINEAR-only by construction, so the plan may build the importer; the
per-frame gate — which sees the negotiated modifier the plan cannot — is what
enforces the tiled half, and falls back to the CPU path if a producer ever
ignores our offer.
* …but only where the encoder can actually take the payload
(`linux_hdr_cuda_ok`). libav's HDR route builds a P010 hardware frames
context and swscales into it, so on a host without the direct-SDK backend a
packed-2:10:10:10 CUDA buffer would land in a P010 surface as garbage. Those
keep the CPU path.
* `nvenc_cuda` stops pinning 8-bit/SDR. Depth and HDR now follow the INPUT
format, like the Windows backend: a 10-bit session whose capture came back
8-bit encodes AND labels 8-bit rather than mislabelling.
* the cursor-blend compute shader gains two 10-bit modes, so the pointer
gamescope leaves out of its node survives the HDR path. Same display-referred
blend the CPU path's `composite_cursor_rgb10` already does — the samples are
PQ, and a real sRGB→PQ cursor LUT is polish, not correctness for a pointer.