Files
punktfunk/.gitea
enricobuehler fd4f032d20 ci: retry bun install — a truncated tarball reads as a corrupt package
docs-site died on `error: Fail extracting tarball for
"@rolldown/binding-linux-x64-gnu"` (run 19630, 2026-08-20). The message
points at the package; the package is fine.

MEASURED, because the message invites the wrong fix:
  * The tarball's sha512 matches docs-site/bun.lock exactly, and it is
    an ordinary 3-entry npm tgz — same gzip framing, same modes, no pax
    headers — as the 1.2.0 one that installs fine. Only the payload
    differs in size (20.6 MB vs 19.0 MB of .node).
  * bun 1.3.13 AND 1.3.14 both extract that exact tarball from disk in
    under 80 ms. So it is not the bun bump the floating oven/bun:1 tag
    brought in, and not a format bun stopped accepting.
  * In the SAME run, the web job installed the same registry over the
    same network and passed — it was 25 s ahead of docs-site.
  * Run 19632, seven minutes later, installed the identical lockfile
    and passed.

So: a transient truncation, not a bad package. bun streams
download-and-extract, so a tarball cut off mid-stream surfaces at the
extract step and names the package it was reading — which is why this
looks like `@rolldown/binding-linux-x64-gnu` is broken and why the
obvious fixes (bump rolldown, pin bun, refresh the lockfile) would all
have "worked" by changing which bytes were in flight, and none of them
would have fixed anything.

scripts/ci/retry.sh already exists for precisely this and its header
already diagnosed it: "the runner box executes many jobs in parallel and
its network drops packets under that load … Wrap every single-shot
network command in CI with this instead." `bun install` is a single-shot
network command and was the one class still unwrapped, so it is wrapped
now at all nine Linux sites — ci.yml (web, docs-site), arch, deb, rpm,
web-screenshots, sdk-publish and plugin-kit-publish (both installs).

3 attempts, not retry.sh's usual 5: a genuinely stale lockfile fails
deterministically under --frozen-lockfile, and 10s+20s of backoff is
enough to outlive a load burst without making that honest failure wait
a minute and a half.

The two windows-host.yml installs are left alone: pwsh, and a Windows
box that is not the contended runner.

Verified: all seven workflows still parse; the helper resolves from
web/, docs-site/ and sdk/ (the three working-directory shapes used);
the wrapper recovers a command that fails once and succeeds on the
retry; and `bash ../scripts/ci/retry.sh 3 bun install --frozen-lockfile
--ignore-scripts` in docs-site installs all 1138 packages, so the
lockfile is sound and the wrapper does not change the command.

Not done, deliberately: docs-site's lockfile still pins rolldown 1.1.2
where web has 1.2.0. That difference is real but it is not this bug,
and refreshing a lockfile to chase a network flake would have buried it.
2026-08-20 09:34:50 +02:00
..