Builds and CI
The rules
- Every repository builds with Bazel, at a version pinned in
.bazelversion.bazel build //...andbazel test //...work the same way in every repo. - Tools are pinned by the repository, never assumed from the machine. Programs a build runs come from nixpkgs through
rules_nixpkgs. A program nixpkgs doesn’t have comes from anhttp_archivewith asha256. Upgrading a tool is then a one-line commit, and Bazel rebuilds exactly what depends on it. - CI uses one shared base image. It contains Nix, Bazelisk, git, and CA certificates, and nothing project-specific. No repository maintains its own CI image.
- Only trusted machines write to the remote cache. CI and designated workstations have write credentials. Everyone else reads.
- Secrets live in GitLab CI/CD variables, never in a repository.
How the cache decides what to rebuild
Bazel models a build as actions. An action is one command with declared inputs and outputs, such as “engrave psalm-119-he with MuseScore into an SVG.” Each action’s key is a hash of its command line, its environment, its tools, and the content of every input. Bazel keeps two stores:
- the action cache, which maps an action key to the hashes of its outputs;
- the content-addressable store, which maps a hash to the bytes.
If an action’s key is already in the cache, Bazel downloads its outputs instead of running it. Timestamps play no part, so a fresh clone on another machine gets the same hits.
Bazel also stops early. If an action re-runs but produces byte-identical output, nothing downstream of it re-runs. Editing a comment in a hymn source, for example, doesn’t re-run MIDI rendering or the manifest.
Caching is only correct if every input is declared. Bazel runs actions in a sandbox that hides undeclared files, so a missing input fails the build instead of producing a wrong cache hit.
Remote cache
- One cache for every project: a single S3 bucket behind a single
bazel-remote. Cache keys hash the exact command, tools and inputs, so projects can’t collide, and identical work is stored once. Give a project its own cache (--remote_instance_name, or a separate bucket) only if its trust level differs, such as a public repo whose outside contributors’ CI must not write. Wherebazel-remoteruns is undecided. - Retention:
bazel-remoteevicts the least recently used entries above a size cap (--max_size). Any S3 lifecycle rule on the bucket is only a long backstop (90 days or more). S3 expires objects by creation date, so it would delete the most-used artifacts on the same schedule as stale ones. - Trust: the server accepts writes only with credentials (
--htpasswd_file), which only CI and designated workstations hold. Everyone else reads anonymously (--allow_unauthenticated_reads) and sets--remote_upload_local_results=false. A client without credentials that tries to upload is refused (UNAUTHENTICATED) and its build still succeeds. This guards against cache poisoning: a machine that can write could store a bad output under a legitimate key, and every other machine would trust it.
CI
Every repository includes a shared template from this project rather than writing its own cache and runner setup. The template comes after the pilot.
GitLab.com’s shared runners start every job from a clean state, so each job downloads its toolchain from the Nix binary cache. If that proves slow, the fix is a generated warm image (built from the same pins with nix2container) or a Nix binary cache near the runners. A hand-written Dockerfile is never the fix.
Why this setup
The goal: never redo a build step whose inputs haven’t changed, whether on a workstation, in CI, or on another machine, while keeping local edit-and-rebuild fast.
- Why not Nix on its own? Nix caches whole derivations and decides whether to rebuild by inputs. It has no early stop on unchanged output (that needs content-addressed derivations, which are experimental), and every build pays for evaluation and copying the source into the store. That’s fine for CI and slow for an edit-and-see loop. We still use Nix, for what it does best: pinned, reproducible tools.
- Why not a Docker image per project? Images drift from the build’s idea of its inputs. Upgrade Chromium in an image, forget to tell the build, and the cache serves stale artifacts. They also add a pipeline stage and tag juggling to every repo.
- Why not hand-rolled input hashes? doreancon.org’s
build-stamps.jsonalready skips unchanged hymns. But it only hashes the inputs someone remembered to list (not tool versions), and it can’t share results between machines.
Pilot results
doreancon.org’s hymn engraving, measured 2026-10-01 on the bazel-pilot branch. There are 18 hymns. MuseScore 4.7.4, Xvfb and Node 24 come from one pinned nixpkgs commit through rules_nixpkgs_core 0.14.0, with Bazel 9.2.0.
| Case | Result |
|---|---|
| Engrave all 18 hymns, warm toolchain | 12.6 s |
| Edit one hymn | 1.3–1.4 s, one action (1.8 s before the fonts were pinned) |
| Revert that edit | 0.09 s, from the disk cache |
| Change a manifest field the engraver doesn’t read | 0.23 s: the per-hymn cuts re-run, all engravings skipped |
| Fresh checkout, shared cache | 36/36 cached; 12.7 s, all of it Bazel start-up |
| Cold container (base image, nothing cached) | 60 s, about 55 s of it toolchain download |
| Cold container, remote cache warmed by trusted CI | 36/36 remote hits, 0.46 s critical path, 42 s total |
| Same hymn engraved directly, warm MuseScore, no Bazel | 0.8 s |
What we learned:
- A single local edit is not faster. It’s about 0.5 s slower than running the engraver directly (Bazel’s sandbox and start-up). The gains are elsewhere: no work is repeated across machines, nothing re-runs when an output comes out unchanged, and the tool versions are part of every cache key.
- Nix’s fontconfig reads the host’s fonts. On a non-NixOS machine it falls back to
/etc/fontsand scans/usr/share/fonts, so text could set differently per machine with nothing in the cache key to show it. Pin afonts.confthat lists only Nix fonts and a prebuilt cache, and pass it asFONTCONFIG_FILE(doreancon.org’snix/fonts.nix). That fixed hermeticity and also removed most of the per-action cost: a freshHOMEno longer rescans fonts. - MuseScore’s output isn’t byte-stable, so nothing downstream of a fresh engraving can be skipped. The skip happens before MuseScore, at the per-hymn manifest cut.
xvfb-run -araces when actions run in parallel: most X servers die on a shared display number. StartXvfb -displayfddirectly instead.- GTK starts gvfs daemons that leave a FUSE mount in
$HOME. SetGIO_USE_VFS=local. - Bazel falls back to
processwrapper-sandboxon Ubuntu workstations (AppArmor restricts unprivileged user namespaces) and in unprivileged containers. It works, but it can’t hide absolute paths outside the sandbox.
Open questions
The doreancon.org pilot exists to answer these before other repositories convert:
- [x] Does
rules_nixpkgswork under Bzlmod with Bazel 9? Yes (0.14.0, 9.2.0). - [x] MuseScore from nixpkgs or the AppImage? nixpkgs: same version, and every hymn’s crop matches the Docker-built SVG exactly.
- [x] Does a build work in an unprivileged container? Yes, with
processwrapper-sandbox, in the base image under local Docker. - [ ] The same on GitLab.com’s shared runners, including writing to a workspace the runner cloned as root.
- [ ] Per-job toolchain download on GitLab.com (about 55 s locally).
- [ ] Where does
bazel-remoterun? - [x] Is a single local edit fast enough? 1.3–1.4 s against 0.8 s direct, after pinning the fonts. Judged acceptable for the pilot.