AWS
SVRBC has one AWS account, named svrbc, in an AWS Organization. People sign
in through IAM Identity Center, not as IAM users, and never as the root user.
Signing in
Section titled “Signing in”Everyone signs in at the access portal:
https://d-906661105f.awsapps.com/start
The portal lists the svrbc account and the roles you have in it. Choose a
role (for example AdministratorAccess) to open the console, or Access keys
for short-lived CLI credentials.
The portal’s own Preferences page holds only language and display mode. Everything else about your user is changed by an administrator.
Regions
Section titled “Regions”Put new resources in us-west-2 (Oregon). SVRBC is in California, so a
West Coast region keeps latency low. Oregon is preferred over us-west-1
(N. California), which is a little closer but costs more, has fewer
availability zones, and gets new services later.
IAM Identity Center is the exception: it lives in us-east-1. It was
enabled there, and an instance can’t be moved. Opening Identity Center in any
other region doesn’t show the instance. Instead it shows either
“[InstanceGuard] unexpected state” or a page offering to Enable an instance
of IAM Identity Center. Don’t click that. Switch the console to
US East (N. Virginia).
Naming
Section titled “Naming”Every SVRBC resource in AWS is named so that its name alone says it is
SVRBC’s, what it is for, and (when it belongs to one site) which site. Names
use lowercase letters, digits and hyphens only: S3 bucket names are global
across all of AWS, and a dot in one breaks HTTPS on the bucket’s own hostname. (Inside the GitLab
group it’s the other way round: projects drop the svrbc- prefix;
see New projects.)
| Kind | Pattern | Examples |
|---|---|---|
| Shared by every project | svrbc-<purpose> |
svrbc-bazel-cache (bucket) |
| Belongs to one site | svrbc-<domain, dots as hyphens>-<purpose> |
svrbc-doreancon-org-assets (the bucket behind assets.doreancon.org) |
| IAM role | <the resource it acts on>-<what it does> |
svrbc-bazel-cache-writer |
| CloudFront function, cache policy, origin access control | <the resource it serves>[-<what it does>] |
svrbc-bazel-cache-paths, svrbc-bazel-cache |
- One resource per site, not one shared one, when its lifecycle is the site’s: a site’s assets go when the site does. Shared infrastructure (the build cache) is one resource for everyone.
- Tag everything with
purpose(what it is for, in a few words) anddocs(the handbook page that explains it). A tag’s value may hold only letters, digits, spaces and_ . : / = + - @: no apostrophe or comma (“svrbc.org’s site” fails with aValidationExceptionwhen the resource is created, which a plan doesn’t catch). - Region
us-west-2(above), except where AWS requiresus-east-1: IAM Identity Center, and certificates (ACM) for CloudFront. - Write down how it was made, in this repository’s
infra/<resource>/: as OpenTofu configuration (below), or, until a resource is imported, as the CLI commands and their JSON (as forinfra/bazel-cache/).
A resource’s public hostname is named for the service it provides, under the domain it serves (see Domains).
Users live in the Identity Center directory itself; nothing is synced from an
outside identity provider. Administrators manage them under
IAM Identity Center → Users (in us-east-1).
Changing a username
Section titled “Changing a username”The console shows the username as read-only, but the API will change it.
Verified 2026-10-01 by renaming xcco3x to cco3. From CloudShell in
us-east-1:
aws identitystore update-user \ --region us-east-1 \ --identity-store-id d-906661105f \ --user-id <user-id> \ --operations '[{"AttributePath":"userName","AttributeValue":"<new-name>"}]'The user ID is the UUID in the URL of the user’s page in the console. The ID
doesn’t change, so groups, account assignments and MFA devices, which hang off
the ID, should carry over. Confirmed 2026-10-01: cco3 signs in to the
portal and the CLI (aws sso login) with its AdministratorAccess
assignment intact.
Two traps:
--operationsmust be JSON. TheAttributePath=…,AttributeValue=…shorthand is rejected becauseAttributeValueis a document type.- Typing into CloudShell through browser automation corrupts input. It
doubled the hyphens inside UUIDs and turned straight quotes into curly ones.
Paste the command by hand. If it has to be typed, build the ID and the JSON
in shell variables without literal
-or"characters.
Infrastructure as code (OpenTofu)
Section titled “Infrastructure as code (OpenTofu)”Adopted 2026-10-05: every SVRBC resource in AWS, and the DNS records SVRBC made at Porkbun, were imported then.
SVRBC’s AWS resources and DNS records are described by
OpenTofu configuration in this repository’s infra/,
and nowhere else. A site that needs a bucket, a certificate or a DNS
record gets it by a merge request here. This is the one project whose CI may
change infrastructure and the one that holds the registrar’s keys, so no other
repository needs either. A site’s own repository still deploys its releases
(its code and files) through its narrow deployer role. A change is a merge request whose pipeline shows the
plan, and merging applies it. The console is for reading. A change made
there is drift, and the next apply undoes it.
OpenTofu rather than Terraform: it’s open source (MPL), has the same CLI and language, and since 1.10 locks S3 state without a DynamoDB table. What follows applies equally to Terraform.
Layout
Section titled “Layout”- One stack per
infra/<name>/directory, each with its own state. A failed apply stays small, and a plan shows only what that stack touches. Things every stack uses (the GitLab OIDC provider) live ininfra/aws-account/. - In each stack:
main.tf(the resources),imports.tf(import blocks, while any remain),versions.tf(terraform {}block, backend, provider versions), and the committed.terraform.lock.hcl. - No module until something is repeated, and then only once the copies are known to be the same. A module written in advance guesses at the variation.
- Policies and CloudFront functions stay in their own
.json/.jsfiles, read withfile()ortemplatefile(): they stay readable, diffable and testable on their own.
Versions
Section titled “Versions”- Pin OpenTofu with Nix:
infra/shell.nix, the same nixpkgs commit asdocker/base. Run every command through it (nix-shell --run 'tofu plan'). - Constrain providers in
required_providers(version = "~> 6.0") and commit.terraform.lock.hcl, with hashes for every platform that runs it:tofu providers lock -platform=linux_amd64 -platform=darwin_arm64. Upgrade withtofu init -upgradein its own merge request.
- State lives in S3, in the
svrbc-tofu-statebucket (us-west-2): versioned, encrypted, public access blocked, key<stack>/terraform.tfstate, locked withuse_lockfile = true. The bucket was made by hand, since it can’t hold its own state (infra/tofu-state/README.md). - Never commit state or plans.
.terraform/,*.tfstate*and*.tfplanare in.gitignore. - State holds secrets in plain text (any generated password, any secret
passed to a resource). Only the apply and plan roles and administrators may
read the bucket. Prefer what keeps a secret out of state: an
ephemeralresource or a write-only (_wo) argument where the provider has one. SSM Parameter Store keeps a secret out of state but not from the planner (below), which can read it like everything else. - Don’t edit state by hand. Rename with a
movedblock, stop managing something without destroying it with aremovedblock, adopt something with animportblock. Each is reviewed in a merge request like any other change, wheretofu state mv/rmandtofu importhappen on one person’s machine and leave no record. Once applied, delete the block.
Writing resources
Section titled “Writing resources”- Bring existing resources in with
importblocks, then change the configuration untiltofu planreports no changes. Only then does the configuration describe what exists. Don’t fix things in the same merge request as the import: make the empty plan first, then the change. The exception is what an import can’t carry: a provider-only setting (publishon a CloudFront function) or adefault_tagstag the resource lacked. Name each such change in the stack’s README. - Refer, don’t repeat. Use
aws_cloudfront_distribution.site.arn, not a pasted ARN, so that order and dependencies follow from the configuration. Use adatasource for something another stack owns. - Tag through the provider’s
default_tags:purposeanddocs(Naming). prevent_destroyon whatever can’t be recreated: buckets with data, the state bucket, certificates that DNS validates.- What something else manages, ignore explicitly, with a comment saying
what manages it:
lifecycle { ignore_changes = [...] }for a Lambda function’s code (which releases deploy), and noaws_s3_objectfor objects that CI writes. for_eachwith stable keys, notcount, for collections: removing the first of threecountitems replaces the other two.- No credentials in
.tffiles or variables files. This repository is public. CI gets AWS credentials through GitLab OIDC; people use theirsvrbcSSO profile.
Why these choices
Section titled “Why these choices”- A bucket of its own for the state (
svrbc-tofu-state), not a generalsvrbc-opsbucket with the state under a prefix (considered 2026-10-05). A bucket costs nothing, and sharing would put per-prefix rules into the policy of the bucket that holds secrets. Moving the state later istofu init -migrate-state. jianyuan/porkbunfor DNS, pinned exactly. Of the five community providers, whose source was read 2026-10-05, it is the most used and the only one tested against Porkbun’s API (its sandbox). Its faults are guarded: it logs both keys atTF_LOG=DEBUG(tofu-cirefuses to run withTF_LOGset and the keys present); it can register domains and move nameservers (tofu-ci checkallows onlyporkbun_dns_record); and it retries a failed create without an idempotency key (the drift check would show a duplicate).icco/porkbunis better engineered (idempotency keys, no key logging, only DNS and nameservers) but had never been run against the live API. Worth revisiting once it has.- Records Porkbun’s own features make stay Porkbun’s: nameservers,
email and URL forwarding, and its free SSL certificate’s
_acme-challengerecords, which it rewrites at every renewal. Each DNS stack’s README lists them.
Changing infrastructure
Section titled “Changing infrastructure”- On a branch, edit the
.tffiles.infra/tofu-ci checkrunsfmt -checkandvalidateover every stack. tofu planlocally (AWS_PROFILE=svrbc) and read it. A replacement (-/+) of anything that serves traffic or holds data needs a reason in the merge request.- Open the merge request. Its pipeline (
tofu-plan) checks and plans every stack assvrbc-tofu-planner, without taking the state lock. - Releasing applies it. A release’s pipeline (the
v*tag the merge bot makes from a version bump; Releases) runstofu-applyassvrbc-tofu-applier, one pipeline at a time, wheninfra/changed since the last release whose pipeline passed (SVRBC_PREVIOUS_RELEASE). A failed apply fails its release, so the next release still compares with the one before and applies again. Merging alone applies nothing. Bump the version in the same merge request as an infrastructure change, so the plan reviewed is the one applied; changes merged without one are applied together by the next release, whose own merge request (only a version bump) plans nothing.
Verified 2026-10-05: !18’s pipeline planned as the planner, master’s
tofu-apply applied every stack as the applier (after !19 let it read
what it may not change), and a scheduled tofu-drift found no drift. IAM’s
policy simulator confirms the boundary holds.
Avoid -target except to recover from a failed apply. It applies part of
a change and leaves the rest of the plan unexplained.
Removing something still in use
Section titled “Removing something still in use”Take a resource out of use in one merge request, and remove it in a later one. OpenTofu deletes a resource removed from the configuration before it updates the resources that used it. Removed together, the resource goes first, or AWS refuses because it’s in use, and the update that would have stopped using it never runs; the stack is left half-applied.
That took doreancon.org down on 2026-10-08. One plan removed the site’s
Lambda function, its URL, and the CloudFront function in front of it, while
updating the distribution to serve from S3 instead. The Lambda and its URL
were deleted, AWS refused to delete the CloudFront function (FunctionInUse),
and the distribution was never updated: it kept sending every request to a
function that no longer existed. The site came back through CI by putting the
CloudFront function back in the configuration, so the next apply had only the
updates to make, and removing it in a later merge request.
So tofu-ci refuses such a plan, in the merge request’s tofu-plan and again
in tofu-apply before anything is applied:
infra/plan-order-check
reads the plan and fails when a resource it only deletes is one that a
resource it only updates depends on (by the dependencies the state records,
which OpenTofu keeps transitively, so it can name more users than the direct
one).
A scheduled pipeline on master runs tofu-drift daily: tofu plan -detailed-exitcode over every stack, failing when AWS differs from infra/
as the latest release has it (master may be ahead, until it’s released).
(The schedule isn’t set up yet.)
The roles CI uses
Section titled “The roles CI uses”Both are assumed through GitLab OIDC, like the cache’s writer: no AWS key is
stored anywhere. Both live in infra/aws-account.
| Role | Who may assume it | What it may do |
|---|---|---|
svrbc-tofu-planner |
any branch of svrbc/forks/developers.svrbc.org, and master |
read everything (ReadOnlyAccess) |
svrbc-tofu-applier |
release tags (v*) of svrbc/developers.svrbc.org only |
change S3, CloudFront, Lambda, ACM, logs and SSM; manage svrbc-* roles that carry the boundary |
infra/aws-accountis applied by an administrator, never by CI. It holds the CI roles themselves, the OIDC provider and the boundary, and the applier is denied all of them. Otherwise a merged change could widen the applier’s own permissions. In a release,tofu-applyplans that stack and fails when it has changes waiting.- Every role a stack makes carries
svrbc-workload-boundary. The applier may create or change asvrbc-*role only with that permissions boundary attached. The boundary allows only what sites and caches use, and never IAM or the state bucket. Without it, the applier could make a role with administrator rights and assume it. - The Porkbun keys reach only
tofu-applyandtofu-drift. They are thesvrbcsubaccount’s keys, stored as protected, masked CI variables (PORKBUN_API_KEY,PORKBUN_SECRET_KEY) scoped to theinfraenvironment. A merge request’s pipeline skips the DNS stacks, so their reviewer plans them locally. The provider (jianyuan/porkbun, pinned) logs both keys atTF_LOG=DEBUG, sotofu-cirefuses to run withTF_LOGset while they are present, andtofu-ci checkallows only itsporkbun_dns_recordresource. - Whatever the planner can read, every member who can push to the fork
can read. A merge request’s configuration runs with that role, and
configuration can print what it reads. So nothing secret may be readable
with
ReadOnlyAccess, state included. Keep secrets where it can’t see them, or accept that they are shared with every member.