Files
punktfunk/scripts/ci/docker-prune.sh
T
enricobuehlerandClaude Opus 5 5f29db3975
ci / rust-arm64 (push) Canceled after 42s
android / android (push) Canceled after 39s
apple / swift (push) Canceled after 41s
apple / screenshots (push) Canceled after 0s
arch / build-publish (push) Canceled after 44s
ci / rust (push) Canceled after 44s
ci / web (push) Canceled after 41s
ci / docs-site (push) Canceled after 41s
ci / bench (push) Canceled after 40s
deb / build-publish (push) Canceled after 7s
deb / build-publish-host (push) Canceled after 1s
deb / build-publish-client-arm64 (push) Canceled after 0s
decky / build-publish (push) Canceled after 0s
docker / build-push (--build-arg FEDORA_VERSION=44, ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora44-rpm) (push) Canceled after 0s
docker / build-push (., web/Dockerfile, punktfunk-web) (push) Canceled after 0s
docker / build-push (ci, ci/fedora-rpm.Dockerfile, punktfunk-fedora-rpm) (push) Canceled after 0s
docker / build-push (ci, ci/rust-ci-noble.Dockerfile, punktfunk-rust-ci-noble) (push) Canceled after 0s
docker / build-push (ci, ci/rust-ci.Dockerfile, punktfunk-rust-ci) (push) Canceled after 0s
docker / build-push (docs-site, docs-site/Dockerfile, punktfunk-docs) (push) Canceled after 0s
docker / build-push-arm64cross (push) Canceled after 0s
docker / deploy-docs (push) Canceled after 0s
rpm / build-publish (43, bazzite, punktfunk-fedora-rpm) (push) Canceled after 0s
rpm / build-publish (44, fedora-44, punktfunk-fedora44-rpm) (push) Canceled after 0s
flatpak / build-publish (push) Canceled after 1m17s
windows-host / package (push) Canceled after 2m16s
windows-host / winget-source (push) Canceled after 0s
windows-msix / package (arm64, C:\Users\Public\ffmpeg-arm64, --no-default-features, aarch64-pc-windows-msvc, C:\t-a64) (push) Canceled after 0s
windows-msix / package (x64, C:\Users\Public\ffmpeg, , x86_64-pc-windows-msvc, C:\t) (push) Canceled after 0s
windows / build (aarch64-pc-windows-msvc) (push) Canceled after 0s
windows / build (x86_64-pc-windows-msvc) (push) Canceled after 0s
fix(ci/runner): grow the runner disk, and simplify the burst condition
The capacity lever, taken: home-runner-1's LXC rootfs went 123 G -> 175 G
(`pct resize 116 rootfs +50G` on home-node-1, online, no downtime, thin pool had
~607 G spare). Free space went 77 G -> 124 G, which is the part that actually
gives three concurrent Rust builds room; the 10-minute prune is now a backstop
rather than the only thing standing between a push-storm and ENOSPC.

MIN_FREE_GB stays at 45. It is deliberately an absolute floor, not a percentage:
what three concurrent target/ dirs need does not change when the disk is resized,
but a percentage threshold silently does — 80% meant ~25 G free before and ~35 G
now. That is exactly why the percentage alone was the wrong instrument.

Also replaces the multi-line `{ …; } || { …; }` burst condition with two flat
tests into a flag. Same semantics, and shellcheck parses it — the brace-group
form across a line break did not survive an edit to the comment above it.
Truth-tabled: quiet at 25%/124G, fires on either signal alone, and stays quiet
when df returns nothing rather than treating an empty reading as pressure.

Deployed to home-runner-1; deployed md5 matches the repo.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 13:19:07 +02:00

64 lines
4.0 KiB
Bash

#!/usr/bin/env bash
# CI runner disk hygiene — invoked by docker-prune.service (every 30 min). Lives in a real script
# rather than inline ExecStart= lines because systemd does its OWN $-expansion on ExecStart and
# empties shell vars / $(...) before /bin/sh sees them (silently breaking the logic under `|| true`).
#
# See docker-prune.service for the full why. The headline: the act_runner cache server's blob store
# lives INSIDE the long-running runner container's writable layer, where `docker prune` can't reach
# it — left alone it grows to tens of GB and fills the disk on its own.
set -u
export PATH=/usr/bin:/bin:/usr/local/bin:$PATH
# The cache-server's blob store is a HOST directory: the fleet runs a standalone cache-server
# service (compose.yml) that bind-mounts this path to /data, and every replica points at it over
# HTTP (`external_server`). It used to live inside a runner container's writable layer, which is
# why this reached in with `docker exec` — that is no longer where it is, and the container name it
# looked for (`gitea-runner-runner`) does not exist either now that the replicas are named
# `gitea-runner-fleet-runner-N-1`. Both halves silently did nothing: the filter matched zero
# containers, so the cap and the burst-clear below were dead code. A plain host path needs neither.
CACHE_DIR=${CACHE_DIR:-/home/runner/gitea-runner-fleet/cache}
CAP_MB=${CAP_MB:-20000} # clear the cache store once it exceeds ~20 GB
BURST_PCT=${BURST_PCT:-80} # full clear once the disk is this % full
MIN_FREE_GB=${MIN_FREE_GB:-45} # ...or this little is left, whichever trips first
# 1) Routine: trim aged images / build cache / stopped containers. sha-<commit> tags aren't
# dangling, so -a is required. until=2h, not 6h: on a busy day every image is younger than six
# hours, so the filter matched nothing and a run reclaimed 0B while `docker system df` was
# reporting 20+ GB reclaimable. Two hours still protects a re-run of the push being worked on.
docker image prune -af --filter until=2h || true
docker builder prune -af --filter until=2h || true
docker buildx prune -af --filter until=2h || true
docker container prune -f --filter until=2h || true
# 2) Cap the cache-server store. Clearing the blobs is safe — act_runner repopulates it and cache
# keys are content-hashed, so this only drops stale entries.
if [ -d "$CACHE_DIR" ]; then
SZ=$(du -sm "$CACHE_DIR" 2>/dev/null | cut -f1)
if [ -n "${SZ:-}" ] && [ "$SZ" -ge "$CAP_MB" ]; then
rm -rf "${CACHE_DIR:?}"/* && echo "cache-server store cleared (was ${SZ} MB)"
fi
fi
# 3) Burst guard: a push-storm fills the disk WITHIN one interval — three concurrent Rust builds,
# each with a multi-GB target/, on top of a ~40 GB containerd image baseline. Trigger on a free
# -space FLOOR as well as a percentage: the percentage is the wrong instrument on its own, since
# what matters is absolute headroom for three concurrent target/ dirs, not a ratio — and the
# ratio moves whenever the disk is resized (it went 123 G -> 175 G on 2026-07-29) while the
# headroom three jobs need does not. In-use images are protected by the daemon, so a burst clear
# cannot pull the rug from a live job.
PCT=$(df --output=pcent / | tr -dc '0-9')
FREE_GB=$(df --output=avail -BG / | tr -dc '0-9')
# Two flat tests into a flag rather than one multi-line `{ …; } || { …; }` condition: the brace-group
# form is easy to get subtly wrong across a line break, and this reads as what it is — either signal
# alone is enough. Empty values (a df that failed) simply leave the flag unset, i.e. no clear.
BURST=0
[ -n "$PCT" ] && [ "$PCT" -ge "$BURST_PCT" ] && BURST=1
[ -n "$FREE_GB" ] && [ "$FREE_GB" -lt "$MIN_FREE_GB" ] && BURST=1
if [ "$BURST" = 1 ]; then
echo "disk ${PCT}% used, ${FREE_GB}G free (thresholds ${BURST_PCT}% / ${MIN_FREE_GB}G) — burst clear"
docker image prune -af || true
docker builder prune -af || true
docker buildx prune -af || true
[ -d "$CACHE_DIR" ] && rm -rf "${CACHE_DIR:?}"/* || true
fi