Skip to content

Builds and CI ​

The rules ​

  1. Every repository builds with Bazel, at a version pinned in .bazelversion. bazel build //... and bazel test //... work the same way in every repo.
  2. Tools are pinned by the repository, never assumed from the machine. Programs a build runs come from nixpkgs through rules_nixpkgs. A program nixpkgs doesn’t have comes from an http_archive with a sha256. Upgrading a tool is then a one-line commit, and Bazel rebuilds exactly what depends on it.
  3. CI uses one shared base image. It contains Nix, Bazelisk, git, and CA certificates, and nothing project-specific. No repository maintains its own CI image.
  4. Only trusted machines write to the remote cache. CI and designated workstations have write credentials. Everyone else reads.
  5. Secrets live in GitLab CI/CD variables, never in a repository.

How the cache decides what to rebuild ​

Bazel models a build as actions. An action is one command with declared inputs and outputs, such as “engrave psalm-119-he with MuseScore into an SVG.” Each action’s key is a hash of its command line, its environment, its tools, and the content of every input. Bazel keeps two stores:

  • the action cache, which maps an action key to the hashes of its outputs;
  • the content-addressable store, which maps a hash to the bytes.

If an action’s key is already in the cache, Bazel downloads its outputs instead of running it. Timestamps play no part, so a fresh clone on another machine gets the same hits.

Bazel also stops early. If an action re-runs but produces byte-identical output, nothing downstream of it re-runs. Editing a comment in a hymn source, for example, doesn’t re-run MIDI rendering or the manifest.

Caching is only correct if every input is declared. Bazel runs actions in a sandbox that hides undeclared files, so a missing input fails the build instead of producing a wrong cache hit.

Remote cache ​

  • One cache for every project: a single S3 bucket behind a single bazel-remote. Cache keys hash the exact command, tools and inputs, so projects can’t collide, and identical work is stored once. Give a project its own cache (--remote_instance_name, or a separate bucket) only if its trust level differs, such as a public repo whose outside contributors’ CI must not write. Where bazel-remote runs is undecided.
  • Retention: bazel-remote evicts the least recently used entries above a size cap (--max_size). Any S3 lifecycle rule on the bucket is only a long backstop (90 days or more). S3 expires objects by creation date, so it would delete the most-used artifacts on the same schedule as stale ones.
  • Trust: the server accepts writes only with credentials (--htpasswd_file), which only CI and designated workstations hold. Everyone else reads anonymously (--allow_unauthenticated_reads) and sets --remote_upload_local_results=false. A client without credentials that tries to upload is refused (UNAUTHENTICATED) and its build still succeeds. This guards against cache poisoning: a machine that can write could store a bad output under a legitimate key, and every other machine would trust it.

CI ​

Every repository includes a shared template from this project rather than writing its own cache and runner setup. The template comes after the pilot.

GitLab.com’s shared runners start every job from a clean state, so each job downloads its toolchain from the Nix binary cache. If that proves slow, the fix is a generated warm image (built from the same pins with nix2container) or a Nix binary cache near the runners. A hand-written Dockerfile is never the fix.

Why this setup ​

The goal: never redo a build step whose inputs haven’t changed, whether on a workstation, in CI, or on another machine, while keeping local edit-and-rebuild fast.

  • Why not Nix on its own? Nix caches whole derivations and decides whether to rebuild by inputs. It has no early stop on unchanged output (that needs content-addressed derivations, which are experimental), and every build pays for evaluation and copying the source into the store. That’s fine for CI and slow for an edit-and-see loop. We still use Nix, for what it does best: pinned, reproducible tools.
  • Why not a Docker image per project? Images drift from the build’s idea of its inputs. Upgrade Chromium in an image, forget to tell the build, and the cache serves stale artifacts. They also add a pipeline stage and tag juggling to every repo.
  • Why not hand-rolled input hashes? doreancon.org’s build-stamps.json already skips unchanged hymns. But it only hashes the inputs someone remembered to list (not tool versions), and it can’t share results between machines.

Pilot results ​

doreancon.org’s hymn engraving, measured 2026-10-01 on the bazel-pilot branch. There are 18 hymns. MuseScore 4.7.4, Xvfb and Node 24 come from one pinned nixpkgs commit through rules_nixpkgs_core 0.14.0, with Bazel 9.2.0.

CaseResult
Engrave all 18 hymns, warm toolchain12.6 s
Edit one hymn1.3–1.4 s, one action (1.8 s before the fonts were pinned)
Revert that edit0.09 s, from the disk cache
Change a manifest field the engraver doesn’t read0.23 s: the per-hymn cuts re-run, all engravings skipped
Fresh checkout, shared cache36/36 cached; 12.7 s, all of it Bazel start-up
Cold container (base image, nothing cached)60 s, about 55 s of it toolchain download
Cold container, remote cache warmed by trusted CI36/36 remote hits, 0.46 s critical path, 42 s total
Same hymn engraved directly, warm MuseScore, no Bazel0.8 s

What we learned:

  • A single local edit is not faster. It’s about 0.5 s slower than running the engraver directly (Bazel’s sandbox and start-up). The gains are elsewhere: no work is repeated across machines, nothing re-runs when an output comes out unchanged, and the tool versions are part of every cache key.
  • Nix’s fontconfig reads the host’s fonts. On a non-NixOS machine it falls back to /etc/fonts and scans /usr/share/fonts, so text could set differently per machine with nothing in the cache key to show it. Pin a fonts.conf that lists only Nix fonts and a prebuilt cache, and pass it as FONTCONFIG_FILE (doreancon.org’s nix/fonts.nix). That fixed hermeticity and also removed most of the per-action cost: a fresh HOME no longer rescans fonts.
  • MuseScore’s output isn’t byte-stable, so nothing downstream of a fresh engraving can be skipped. The skip happens before MuseScore, at the per-hymn manifest cut.
  • xvfb-run -a races when actions run in parallel: most X servers die on a shared display number. Start Xvfb -displayfd directly instead.
  • GTK starts gvfs daemons that leave a FUSE mount in $HOME. Set GIO_USE_VFS=local.
  • Bazel falls back to processwrapper-sandbox on Ubuntu workstations (AppArmor restricts unprivileged user namespaces) and in unprivileged containers. It works, but it can’t hide absolute paths outside the sandbox.

Open questions ​

The doreancon.org pilot exists to answer these before other repositories convert:

  • [x] Does rules_nixpkgs work under Bzlmod with Bazel 9? Yes (0.14.0, 9.2.0).
  • [x] MuseScore from nixpkgs or the AppImage? nixpkgs: same version, and every hymn’s crop matches the Docker-built SVG exactly.
  • [x] Does a build work in an unprivileged container? Yes, with processwrapper-sandbox, in the base image under local Docker.
  • [ ] The same on GitLab.com’s shared runners, including writing to a workspace the runner cloned as root.
  • [ ] Per-job toolchain download on GitLab.com (about 55 s locally).
  • [ ] Where does bazel-remote run?
  • [x] Is a single local edit fast enough? 1.3–1.4 s against 0.8 s direct, after pinning the fonts. Judged acceptable for the pilot.