Skip to content

AWS

SVRBC has one AWS account, named svrbc, in an AWS Organization. People sign in through IAM Identity Center, not as IAM users, and never as the root user.

Everyone signs in at the access portal:

https://d-906661105f.awsapps.com/start

The portal lists the svrbc account and the roles you have in it. Choose a role (for example AdministratorAccess) to open the console, or Access keys for short-lived CLI credentials.

The portal’s own Preferences page holds only language and display mode. Everything else about your user is changed by an administrator.

Put new resources in us-west-2 (Oregon). SVRBC is in California, so a West Coast region keeps latency low. Oregon is preferred over us-west-1 (N. California), which is a little closer but costs more, has fewer availability zones, and gets new services later.

IAM Identity Center is the exception: it lives in us-east-1. It was enabled there, and an instance can’t be moved. Opening Identity Center in any other region doesn’t show the instance. Instead it shows either “[InstanceGuard] unexpected state” or a page offering to Enable an instance of IAM Identity Center. Don’t click that. Switch the console to US East (N. Virginia).

Every SVRBC resource in AWS is named so that its name alone says it is SVRBC’s, what it is for, and (when it belongs to one site) which site. Names use lowercase letters, digits and hyphens only: S3 bucket names are global across all of AWS, and a dot in one breaks HTTPS on the bucket’s own hostname. (Inside the GitLab group it’s the other way round: projects drop the svrbc- prefix; see New projects.)

Kind Pattern Examples
Shared by every project svrbc-<purpose> svrbc-bazel-cache (bucket)
Belongs to one site svrbc-<domain, dots as hyphens>-<purpose> svrbc-doreancon-org-assets (the bucket behind assets.doreancon.org)
IAM role <the resource it acts on>-<what it does> svrbc-bazel-cache-writer
CloudFront function, cache policy, origin access control <the resource it serves>[-<what it does>] svrbc-bazel-cache-paths, svrbc-bazel-cache
  • One resource per site, not one shared one, when its lifecycle is the site’s: a site’s assets go when the site does. Shared infrastructure (the build cache) is one resource for everyone.
  • Tag everything with purpose (what it is for, in a few words) and docs (the handbook page that explains it). A tag’s value may hold only letters, digits, spaces and _ . : / = + - @: no apostrophe or comma (“svrbc.org’s site” fails with a ValidationException when the resource is created, which a plan doesn’t catch).
  • Region us-west-2 (above), except where AWS requires us-east-1: IAM Identity Center, and certificates (ACM) for CloudFront.
  • Write down how it was made, in this repository’s infra/<resource>/: as OpenTofu configuration (below), or, until a resource is imported, as the CLI commands and their JSON (as for infra/bazel-cache/).

A resource’s public hostname is named for the service it provides, under the domain it serves (see Domains).

Users live in the Identity Center directory itself; nothing is synced from an outside identity provider. Administrators manage them under IAM Identity Center → Users (in us-east-1).

The console shows the username as read-only, but the API will change it. Verified 2026-10-01 by renaming xcco3x to cco3. From CloudShell in us-east-1:

Terminal window
aws identitystore update-user \
--region us-east-1 \
--identity-store-id d-906661105f \
--user-id <user-id> \
--operations '[{"AttributePath":"userName","AttributeValue":"<new-name>"}]'

The user ID is the UUID in the URL of the user’s page in the console. The ID doesn’t change, so groups, account assignments and MFA devices, which hang off the ID, should carry over. Confirmed 2026-10-01: cco3 signs in to the portal and the CLI (aws sso login) with its AdministratorAccess assignment intact.

Two traps:

  • --operations must be JSON. The AttributePath=…,AttributeValue=… shorthand is rejected because AttributeValue is a document type.
  • Typing into CloudShell through browser automation corrupts input. It doubled the hyphens inside UUIDs and turned straight quotes into curly ones. Paste the command by hand. If it has to be typed, build the ID and the JSON in shell variables without literal - or " characters.

Adopted 2026-10-05: every SVRBC resource in AWS, and the DNS records SVRBC made at Porkbun, were imported then.

SVRBC’s AWS resources and DNS records are described by OpenTofu configuration in this repository’s infra/, and nowhere else. A site that needs a bucket, a certificate or a DNS record gets it by a merge request here. This is the one project whose CI may change infrastructure and the one that holds the registrar’s keys, so no other repository needs either. A site’s own repository still deploys its releases (its code and files) through its narrow deployer role. A change is a merge request whose pipeline shows the plan, and merging applies it. The console is for reading. A change made there is drift, and the next apply undoes it.

OpenTofu rather than Terraform: it’s open source (MPL), has the same CLI and language, and since 1.10 locks S3 state without a DynamoDB table. What follows applies equally to Terraform.

  • One stack per infra/<name>/ directory, each with its own state. A failed apply stays small, and a plan shows only what that stack touches. Things every stack uses (the GitLab OIDC provider) live in infra/aws-account/.
  • In each stack: main.tf (the resources), imports.tf (import blocks, while any remain), versions.tf (terraform {} block, backend, provider versions), and the committed .terraform.lock.hcl.
  • No module until something is repeated, and then only once the copies are known to be the same. A module written in advance guesses at the variation.
  • Policies and CloudFront functions stay in their own .json/.js files, read with file() or templatefile(): they stay readable, diffable and testable on their own.
  • Pin OpenTofu with Nix: infra/shell.nix, the same nixpkgs commit as docker/base. Run every command through it (nix-shell --run 'tofu plan').
  • Constrain providers in required_providers (version = "~> 6.0") and commit .terraform.lock.hcl, with hashes for every platform that runs it: tofu providers lock -platform=linux_amd64 -platform=darwin_arm64. Upgrade with tofu init -upgrade in its own merge request.
  • State lives in S3, in the svrbc-tofu-state bucket (us-west-2): versioned, encrypted, public access blocked, key <stack>/terraform.tfstate, locked with use_lockfile = true. The bucket was made by hand, since it can’t hold its own state (infra/tofu-state/README.md).
  • Never commit state or plans. .terraform/, *.tfstate* and *.tfplan are in .gitignore.
  • State holds secrets in plain text (any generated password, any secret passed to a resource). Only the apply and plan roles and administrators may read the bucket. Prefer what keeps a secret out of state: an ephemeral resource or a write-only (_wo) argument where the provider has one. SSM Parameter Store keeps a secret out of state but not from the planner (below), which can read it like everything else.
  • Don’t edit state by hand. Rename with a moved block, stop managing something without destroying it with a removed block, adopt something with an import block. Each is reviewed in a merge request like any other change, where tofu state mv/rm and tofu import happen on one person’s machine and leave no record. Once applied, delete the block.
  • Bring existing resources in with import blocks, then change the configuration until tofu plan reports no changes. Only then does the configuration describe what exists. Don’t fix things in the same merge request as the import: make the empty plan first, then the change. The exception is what an import can’t carry: a provider-only setting (publish on a CloudFront function) or a default_tags tag the resource lacked. Name each such change in the stack’s README.
  • Refer, don’t repeat. Use aws_cloudfront_distribution.site.arn, not a pasted ARN, so that order and dependencies follow from the configuration. Use a data source for something another stack owns.
  • Tag through the provider’s default_tags: purpose and docs (Naming).
  • prevent_destroy on whatever can’t be recreated: buckets with data, the state bucket, certificates that DNS validates.
  • What something else manages, ignore explicitly, with a comment saying what manages it: lifecycle { ignore_changes = [...] } for a Lambda function’s code (which releases deploy), and no aws_s3_object for objects that CI writes.
  • for_each with stable keys, not count, for collections: removing the first of three count items replaces the other two.
  • No credentials in .tf files or variables files. This repository is public. CI gets AWS credentials through GitLab OIDC; people use their svrbc SSO profile.
  • A bucket of its own for the state (svrbc-tofu-state), not a general svrbc-ops bucket with the state under a prefix (considered 2026-10-05). A bucket costs nothing, and sharing would put per-prefix rules into the policy of the bucket that holds secrets. Moving the state later is tofu init -migrate-state.
  • jianyuan/porkbun for DNS, pinned exactly. Of the five community providers, whose source was read 2026-10-05, it is the most used and the only one tested against Porkbun’s API (its sandbox). Its faults are guarded: it logs both keys at TF_LOG=DEBUG (tofu-ci refuses to run with TF_LOG set and the keys present); it can register domains and move nameservers (tofu-ci check allows only porkbun_dns_record); and it retries a failed create without an idempotency key (the drift check would show a duplicate). icco/porkbun is better engineered (idempotency keys, no key logging, only DNS and nameservers) but had never been run against the live API. Worth revisiting once it has.
  • Records Porkbun’s own features make stay Porkbun’s: nameservers, email and URL forwarding, and its free SSL certificate’s _acme-challenge records, which it rewrites at every renewal. Each DNS stack’s README lists them.
  1. On a branch, edit the .tf files. infra/tofu-ci check runs fmt -check and validate over every stack.
  2. tofu plan locally (AWS_PROFILE=svrbc) and read it. A replacement (-/+) of anything that serves traffic or holds data needs a reason in the merge request.
  3. Open the merge request. Its pipeline (tofu-plan) checks and plans every stack as svrbc-tofu-planner, without taking the state lock.
  4. Releasing applies it. A release’s pipeline (the v* tag the merge bot makes from a version bump; Releases) runs tofu-apply as svrbc-tofu-applier, one pipeline at a time, when infra/ changed since the last release whose pipeline passed (SVRBC_PREVIOUS_RELEASE). A failed apply fails its release, so the next release still compares with the one before and applies again. Merging alone applies nothing. Bump the version in the same merge request as an infrastructure change, so the plan reviewed is the one applied; changes merged without one are applied together by the next release, whose own merge request (only a version bump) plans nothing.

Verified 2026-10-05: !18’s pipeline planned as the planner, master’s tofu-apply applied every stack as the applier (after !19 let it read what it may not change), and a scheduled tofu-drift found no drift. IAM’s policy simulator confirms the boundary holds.

Avoid -target except to recover from a failed apply. It applies part of a change and leaves the rest of the plan unexplained.

Take a resource out of use in one merge request, and remove it in a later one. OpenTofu deletes a resource removed from the configuration before it updates the resources that used it. Removed together, the resource goes first, or AWS refuses because it’s in use, and the update that would have stopped using it never runs; the stack is left half-applied.

That took doreancon.org down on 2026-10-08. One plan removed the site’s Lambda function, its URL, and the CloudFront function in front of it, while updating the distribution to serve from S3 instead. The Lambda and its URL were deleted, AWS refused to delete the CloudFront function (FunctionInUse), and the distribution was never updated: it kept sending every request to a function that no longer existed. The site came back through CI by putting the CloudFront function back in the configuration, so the next apply had only the updates to make, and removing it in a later merge request.

So tofu-ci refuses such a plan, in the merge request’s tofu-plan and again in tofu-apply before anything is applied: infra/plan-order-check reads the plan and fails when a resource it only deletes is one that a resource it only updates depends on (by the dependencies the state records, which OpenTofu keeps transitively, so it can name more users than the direct one).

A scheduled pipeline on master runs tofu-drift daily: tofu plan -detailed-exitcode over every stack, failing when AWS differs from infra/ as the latest release has it (master may be ahead, until it’s released). (The schedule isn’t set up yet.)

Both are assumed through GitLab OIDC, like the cache’s writer: no AWS key is stored anywhere. Both live in infra/aws-account.

Role Who may assume it What it may do
svrbc-tofu-planner any branch of svrbc/forks/developers.svrbc.org, and master read everything (ReadOnlyAccess)
svrbc-tofu-applier release tags (v*) of svrbc/developers.svrbc.org only change S3, CloudFront, Lambda, ACM, logs and SSM; manage svrbc-* roles that carry the boundary
  • infra/aws-account is applied by an administrator, never by CI. It holds the CI roles themselves, the OIDC provider and the boundary, and the applier is denied all of them. Otherwise a merged change could widen the applier’s own permissions. In a release, tofu-apply plans that stack and fails when it has changes waiting.
  • Every role a stack makes carries svrbc-workload-boundary. The applier may create or change a svrbc-* role only with that permissions boundary attached. The boundary allows only what sites and caches use, and never IAM or the state bucket. Without it, the applier could make a role with administrator rights and assume it.
  • The Porkbun keys reach only tofu-apply and tofu-drift. They are the svrbc subaccount’s keys, stored as protected, masked CI variables (PORKBUN_API_KEY, PORKBUN_SECRET_KEY) scoped to the infra environment. A merge request’s pipeline skips the DNS stacks, so their reviewer plans them locally. The provider (jianyuan/porkbun, pinned) logs both keys at TF_LOG=DEBUG, so tofu-ci refuses to run with TF_LOG set while they are present, and tofu-ci check allows only its porkbun_dns_record resource.
  • Whatever the planner can read, every member who can push to the fork can read. A merge request’s configuration runs with that role, and configuration can print what it reads. So nothing secret may be readable with ReadOnlyAccess, state included. Keep secrets where it can’t see them, or accept that they are shared with every member.