perf(pyrowave): elevated GPU scheduling + global-priority encode queue
apple / screenshots (push) Successful in 6m18s
ci / web (push) Successful in 1m25s
ci / docs-site (push) Successful in 1m5s
android / android (push) Successful in 13m2s
arch / build-publish (push) Successful in 12m39s
decky / build-publish (push) Successful in 19s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 10s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 10s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 8s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 9s
ci / bench (push) Successful in 5m36s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 10s
windows-host / package (push) Successful in 16m23s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 6m8s
deb / build-publish-host (push) Successful in 10m16s
deb / build-publish (push) Successful in 12m30s
ci / rust (push) Successful in 19m26s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Successful in 22m20s
docker / deploy-docs (push) Successful in 23s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Successful in 24m32s
apple / swift (push) Successful in 1m17s

PyroWave's wavelet encode runs on the GPU's compute/shader cores, so a GPU-bound
game starves it: submit spikes from ~2 ms to ~15 ms under a 95%+ game load and
the stream fps collapses. NVENC is immune (separate encoder ASIC). Two levers to
let the encode get scheduled ahead of the game's rendering:

- Windows process GPU scheduling: D3DKMTSetProcessSchedulingPriorityClass, env
  PUNKTFUNK_GPU_PRIORITY = off|above-normal|high (default)|realtime. Best-effort,
  once per process, non-fatal on refusal (enc/windows/pyrowave.rs).
- Global-priority Vulkan encode queue (Granite patch 0005): request a
  VK_KHR_global_priority queue (PYROWAVE_QUEUE_PRIORITY = off|high|realtime,
  default realtime), downgrading REALTIME→HIGH→none on NOT_PERMITTED so a refused
  class never regresses the encoder to HEVC.

HONEST STATUS: on an RTX 4090 / Windows / WDDM neither moved the ~15 ms spikes —
the graphics-vs-compute preemption granularity is the wall, not the priority
level. Kept because both are correct, harmless (graceful fallback), and may help
other GPUs/drivers. For a GPU-saturated game the working levers are reducing the
encode's GPU cost (4:2:0/8-bit) or H.265; PyroWave holds full rate on the desktop
and in games that leave the GPU headroom.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
2026-07-19 01:02:01 +02:00
parent dc20e4452e
commit ac0e73321c
5 changed files with 273 additions and 2 deletions
@@ -2115,6 +2115,51 @@ bool Context::create_device(VkPhysicalDevice gpu_, VkSurfaceKHR surface,
vpGetProfileProperties(profile.profile, &props);
#endif
// PUNKTFUNK PATCH (patches/0005-global-priority-queue.patch, NOT upstream): request a high
// global-priority queue so PyroWave's compute-shader encode can PREEMPT a GPU-bound game on
// the shared shader cores. A WDDM process-priority raise does not help (it only orders packet
// submission, not hardware preemption); a global-priority queue is the actual NVIDIA compute-
// preemption lever. `PYROWAVE_QUEUE_PRIORITY` = off | high | realtime (default realtime). The
// device-create loop below downgrades on NOT_PERMITTED, so a refused class NEVER fails the
// encoder — it just runs at default priority. Only the fresh encoder device (no inherit_info)
// is affected; this Granite copy is vendored solely for the PyroWave codec.
Util::SmallVector<VkQueueGlobalPriorityKHR> pf_priority_candidates;
Util::SmallVector<VkDeviceQueueGlobalPriorityCreateInfoKHR> pf_global_priority_infos;
if (!inherit_info)
{
const char *pf_gp_ext = nullptr;
if (has_extension(VK_KHR_GLOBAL_PRIORITY_EXTENSION_NAME))
pf_gp_ext = VK_KHR_GLOBAL_PRIORITY_EXTENSION_NAME;
else if (has_extension(VK_EXT_GLOBAL_PRIORITY_EXTENSION_NAME))
pf_gp_ext = VK_EXT_GLOBAL_PRIORITY_EXTENSION_NAME;
std::string pf_prio;
if (!Util::get_environment("PYROWAVE_QUEUE_PRIORITY", pf_prio))
pf_prio = "realtime";
for (auto &pf_ch : pf_prio)
if (pf_ch >= 'A' && pf_ch <= 'Z')
pf_ch = char(pf_ch + 32);
if (pf_gp_ext && pf_prio != "off")
{
if (pf_prio == "high")
{
pf_priority_candidates.push_back(VK_QUEUE_GLOBAL_PRIORITY_HIGH_KHR);
}
else
{
pf_priority_candidates.push_back(VK_QUEUE_GLOBAL_PRIORITY_REALTIME_KHR);
pf_priority_candidates.push_back(VK_QUEUE_GLOBAL_PRIORITY_HIGH_KHR);
}
enabled_extensions.push_back(pf_gp_ext);
pf_global_priority_infos.resize(queue_infos.size());
for (size_t pf_i = 0; pf_i < queue_infos.size(); pf_i++)
{
pf_global_priority_infos[pf_i] = { VK_STRUCTURE_TYPE_DEVICE_QUEUE_GLOBAL_PRIORITY_CREATE_INFO_KHR };
pf_global_priority_infos[pf_i].globalPriority = pf_priority_candidates.front();
queue_infos[pf_i].pNext = &pf_global_priority_infos[pf_i];
}
}
}
if (inherit_info)
{
device_info.enabledExtensionCount = inherit_info->enabledExtensionCount;
@@ -2144,8 +2189,41 @@ bool Context::create_device(VkPhysicalDevice gpu_, VkSurfaceKHR surface,
if (device == VK_NULL_HANDLE)
return false;
}
else if (vkCreateDevice(gpu, &device_info, nullptr, &device) != VK_SUCCESS)
return false;
else
{
// PUNKTFUNK: try the requested global-priority class, downgrade through the
// candidate list on NOT_PERMITTED, then finally create with no global priority.
VkResult pf_res;
if (pf_priority_candidates.empty())
{
pf_res = vkCreateDevice(gpu, &device_info, nullptr, &device);
}
else
{
pf_res = VK_ERROR_NOT_PERMITTED_KHR;
for (size_t pf_a = 0; pf_a < pf_priority_candidates.size(); pf_a++)
{
for (auto &pf_gp : pf_global_priority_infos)
pf_gp.globalPriority = pf_priority_candidates[pf_a];
pf_res = vkCreateDevice(gpu, &device_info, nullptr, &device);
if (pf_res != VK_ERROR_NOT_PERMITTED_KHR)
break;
LOGW("PyroWave: global queue priority %u not permitted; downgrading.\n",
unsigned(pf_priority_candidates[pf_a]));
}
if (pf_res == VK_ERROR_NOT_PERMITTED_KHR)
{
for (auto &pf_qi : queue_infos)
pf_qi.pNext = nullptr;
pf_res = vkCreateDevice(gpu, &device_info, nullptr, &device);
LOGW("PyroWave: all global queue priorities refused; default priority.\n");
}
else
LOGI("PyroWave: encode device created with an elevated global queue priority.\n");
}
if (pf_res != VK_SUCCESS)
return false;
}
}
if (inherit_info)