2e6cd0e235c29f9da1f20c122c2041817749aa43
5
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
fcdb90c39b |
feat(ci): builder images move to the LAN registry, content-keyed; fan-out scoped by paths
The five builder images now live on home-ci-core's LAN registry (192.168.1.58:5010) under content keys — a hash of the ci/ tree (+ rust-toolchain.toml for the cross image). docker.yml builds one only when its key has no manifest yet, so a push that doesn't touch ci/ costs a curl per image instead of seven WAN pushes and a set of per-SHA tags that no plain prune could ever reclaim. Releases pin builders by copying the key manifest to a vX.Y.Z tag via the registry API — no rebuild, no bytes moved. Around that: deb/rpm/arch/android/apple/decky get path filters so docs-only pushes stop lighting up the whole fleet (branch pushes only — tag runs match tags:, as flatpak/windows-msix releases have proven for months); the report-only bench job moves to bench.yml (nightly + dispatch) and stops occupying a fleet slot per push; flatpak caches its Flathub runtimes and builder state instead of re-downloading multi-GB every run; rpm's cargo registry cache gets its own key namespace instead of sharing the Ubuntu jobs'; audit caches cargo bin+registry rather than the whole toolchain dir; docker-prune.sh loses the local act-cache cap/burst-clear (the cache is central now — deleting it under disk pressure was how runner-2 ended up cold-building everything) and gains a leaked-network prune. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ab6d7727d3 |
fix(ci/runner): the disk guard has to fire before the disk is tight, not after
ci / web (push) Successful in 57s
ci / docs-site (push) Successful in 2m23s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Successful in 21s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Successful in 14s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Successful in 15s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Successful in 19s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Successful in 57s
ci / bench (push) Successful in 6m45s
ci / rust-arm64 (push) Successful in 10m32s
ci / rust (push) Failing after 10m39s
decky / build-publish (push) Successful in 38s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Successful in 22s
docker / build-push-arm64cross (push) Successful in 23s
android / android (push) Successful in 12m39s
docker / deploy-docs (push) Successful in 1m18s
apple / swift (push) Successful in 4m56s
arch / build-publish (push) Successful in 20m44s
apple / screenshots (push) Successful in 22m0s
deb / build-publish-host (push) Successful in 10m47s
deb / build-publish-client-arm64 (push) Successful in 7m53s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Failing after 7m27s
deb / build-publish (push) Failing after 6m6s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 11m40s
Measured over six hours after the last change: zero burst clears fired, while deb still died of ENOSPC between polls. Both 30 min and 10 min lost the same race — three concurrent Rust builds fill the disk and drain it again inside the interval, so every poll landed on a healthy df and the guard concluded all was well. Two changes, both aimed at that gap rather than at the symptom. The interval goes to 2 min: the script is a few docker calls and no-ops in about a second, so sampling five times more often costs nothing worth counting. And MIN_FREE_GB goes 45 -> 60, because the clear only reclaims idle images (~18 G measured) while three jobs can eat the remainder inside one interval — a guard that waits for "tight" has already lost. It has to act while there is still room to act in. This narrows the window; it does not close it. With a 43 G image baseline on a 172 G disk, three concurrent heavy jobs are working inside ~129 G and can exceed it. Sub-interval spikes are a polling problem, and the honest fixes are fewer concurrent replicas or more disk. Deployed to home-runner-1; deployed md5 matches the repo. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5f29db3975 |
fix(ci/runner): grow the runner disk, and simplify the burst condition
ci / rust-arm64 (push) Canceled after 42s
android / android (push) Canceled after 39s
apple / swift (push) Canceled after 41s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 44s
ci / rust (push) Canceled after 44s
ci / web (push) Canceled after 41s
ci / docs-site (push) Canceled after 41s
ci / bench (push) Canceled after 40s
deb / build-publish (push) Canceled after 7s
deb / build-publish-host (push) Canceled after 1s
deb / build-publish-client-arm64 (push) Canceled after 0s
decky / build-publish (push) Canceled after 0s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Canceled after 0s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / build-push-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 1m17s
windows-host / package (push) Canceled after 2m16s
windows-host / winget-source (push) Canceled after 0s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Canceled after 0s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Canceled after 0s
windows / build (aarch64-pc-windows-msvc) (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 0s
The capacity lever, taken: home-runner-1's LXC rootfs went 123 G -> 175 G
(`pct resize 116 rootfs +50G` on home-node-1, online, no downtime, thin pool had
~607 G spare). Free space went 77 G -> 124 G, which is the part that actually
gives three concurrent Rust builds room; the 10-minute prune is now a backstop
rather than the only thing standing between a push-storm and ENOSPC.
MIN_FREE_GB stays at 45. It is deliberately an absolute floor, not a percentage:
what three concurrent target/ dirs need does not change when the disk is resized,
but a percentage threshold silently does — 80% meant ~25 G free before and ~35 G
now. That is exactly why the percentage alone was the wrong instrument.
Also replaces the multi-line `{ …; } || { …; }` burst condition with two flat
tests into a flag. Same semantics, and shellcheck parses it — the brace-group
form across a line break did not survive an edit to the comment above it.
Truth-tabled: quiet at 25%/124G, fires on either signal alone, and stays quiet
when df returns nothing rather than treating an empty reading as pressure.
Deployed to home-runner-1; deployed md5 matches the repo.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7fc387bcbf |
fix(ci/runner): the disk guard had been dead code, and polled too slowly to matter
arch has been failing with `as: BFD (GNU Binutils) 2.47 assertion fail`, which reads like a toolchain regression and is not one. The line above it is `can't write 10 bytes to section .text._ZN6Vulkan...: 'No space left on device'` — the assembler handling ENOSPC badly. Both recent arch failures are the runner filling its disk, and by the time anyone looks, df reports 37% used. Three things were wrong with the hygiene that was supposed to prevent this. The cache cap and the burst-clear were dead code. They looked up the runner as `docker ps -f name=gitea-runner-runner`, which matches zero containers now that the replicas are `gitea-runner-fleet-runner-N-1`, so $RUNNER was always empty and both branches were skipped. The store also moved: the fleet runs a standalone cache-server bind-mounting a HOST directory, so no docker exec is needed at all. The routine prune reclaimed nothing. `--filter until=6h` on a runner that rebuilds its CI images every push means every image is younger than the window — measured: 0B reclaimed while docker system df reported 22.76 GB reclaimable. Now until=2h. The burst guard never fired. It polled every 30 minutes for >=80% used, but three concurrent Rust builds fill the disk and drain it again well inside that window, so the poll kept landing on a healthy df. Now every 10 minutes, and it triggers on a free-space FLOOR too — 80% of 123 G still leaves only ~25 G, which three jobs swallow before the next poll. Deployed to home-runner-1 and exercised: shellcheck clean, timer active on the new interval, one run reclaimed 551 MB. Honest limit: this improves the odds, it does not fix the cause. The remaining 22 GB of idle images are the fedora-rpm bases the next run wants back, so there is no free headroom to reclaim — three replicas building this workspace share one 123 G disk. The lever is capacity (grow the LXC) or concurrency (drop to two replicas), and that is a judgement call, not a script. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
693eedd314 |
ci(runner): cap the act_runner cache + 30-min prune (fix recurring disk-full)
The hourly docker-prune could never reclaim the real disk filler: the act_runner cache server's blob store (cache.dir:"" -> /root/.cache/actcache/cache) lives in the long-running runner container's WRITABLE LAYER, which docker prune can't see. It grew to ~66 GB and filled the 125 GB disk on its own. - New docker-prune.sh holds the logic (inline ExecStart= broke under systemd's own $-expansion, which emptied $SZ/$(...) before sh ran them — silently no-oping the burst guard). The unit now just calls the script. - Caps the actcache: clears the blobs once they exceed ~20 GB (act_runner repopulates; keys are content-hashed, so only stale entries drop). - Burst guard lowered 85%->80% and now also clears the actcache. - Timer hourly -> every 30 min; image/cache `until` 12h -> 6h. Live: cleared 66 GB on home-runner-1 (93% -> 20%), deployed + verified. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |