Files
RosterChief/DEPLOYMENT.md
Bernard Siebens d30b163122 Add maintenance mode: lock the platform down from the control panel
Closes every club subdomain with a 503 in that club's own colours, stands the
scheduled jobs down, and keeps open exactly what is needed to end it again.

The exemptions ARE the feature:

- /accounts/ stays open on the base domain. Close it too and you cannot sign in to
  turn maintenance off -- a lock-down with no key, fixable only from a shell.
- /healthz answers on every host. Close it and the load balancer decides the node
  is dead, stops routing to it, and takes the control panel down with everything
  else.
- migrate and collectstatic are NOT blocked. Maintenance is usually declared in
  order to run them; a blanket guard on BaseCommand would mean turning the mode off
  to do the work you turned it on for. Only the domain jobs (archive_overdue_clubs,
  extend_event_series, import_members_csv) refuse, and they exit non-zero so cron
  mails you -- a scheduled job that silently skips itself is how a month of billing
  goes missing.

The state is cached with a 10-second TTL, not for ever. Write-through makes the
flip instant for the shared Redis of a real deployment, and the TTL is the belt to
that braces: on a per-process cache -- a dev box with no Redis, or a misconfigured
deploy -- a lock-down that reached only one gunicorn worker would be worse than
useless. Live-verified: a club subdomain, its login page and the base domain all
503 while the control panel and the sign-in screens stay up.

Also adds the two deployment pieces asked for: compose.behind-proxy.yaml for a
dev/test box that already runs Caddy on :80 (app on the loopback, host Caddy proxies
to it -- and the host's Caddy still needs the DNS plugin, because the wildcard is
still a wildcard), and deploy/backup.sh + restore-check.sh with a cron schedule. The
backup writes to a .part file and only lands it once gzip -t says it is readable: a
truncated dump that looks like a backup is the failure you find on the day you need
it. The weekly restore rehearsal is the only line in that cron that proves the rest
work.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-14 10:11:15 +02:00

12 KiB

Deploying RosterChief

One server today, several later, with no code changes in between — only environment variables. This document is the runbook and, more usefully, the list of things that are specific to this app and will bite you if you treat it as a generic Django deploy.

The five things that make this deployment unusual

1. You need a wildcard TLS certificate, and that forces DNS-01. Tenancy is subdomain-based (ajax.rosterchief.app), so the certificate must cover *.rosterchief.app. Let's Encrypt will not issue a wildcard over HTTP-01 — only over DNS-01, which means the TLS terminator needs API access to your DNS zone. That is why deploy/caddy/Dockerfile builds Caddy with a DNS provider plugin, and why CLOUDFLARE_API_TOKEN is a required variable rather than a nicety. Swap the plugin (caddy-dns/route53, caddy-dns/digitalocean, …) if your DNS lives elsewhere.

DNS needs two records, both pointing at the server:

A   rosterchief.app    -> <server ip>
A   *.rosterchief.app  -> <server ip>

2. Redis is not optional, even on one server. waffle caches each feature flag's targeting in the Django cache, and LocMemCache is private to a single process. Under several gunicorn workers, toggling a feature in the control panel flushes one worker's cache while the others keep serving the stale flag — a feature that "sometimes doesn't turn on". A shared cache is the fix.

3. SECURE_PROXY_SSL_HEADER must be set, and Caddy must send the header. Caddy terminates TLS, so without it Django believes every request is plain HTTP: request.is_secure() goes false, WebAuthn disagrees with the browser about the origin, and SECURE_SSL_REDIRECT becomes a redirect loop. Both halves are already wired (settings + header_up X-Forwarded-Proto); don't remove either.

4. Uploads must move to object storage before the second app server. Club logos go to MEDIA_ROOT on local disk. On one box that is fine. On two, a logo uploaded to node A is a 404 on node B. Setting AWS_STORAGE_BUCKET_NAME switches the default storage to S3 — do it before you scale, not during.

5. PDF invoices need native libraries. WeasyPrint binds to pango/cairo. The image installs them; a bare-metal deploy would need them too, and a Mac needs Homebrew. This is the main reason to run the container even in development if you touch invoicing.

First deploy

# 1. Configure
cp .env.compose.example .env                 # read by docker compose
cp .env.production.example .env.production   # read by Django
python -c "import secrets; print(secrets.token_urlsafe(64))"   # -> DJANGO_SECRET_KEY

# 2. Build and start
docker compose build
docker compose up -d db redis
docker compose run --rm web python manage.py migrate
docker compose run --rm web python manage.py createsuperuser
docker compose up -d

# 3. Verify
curl -fsS https://rosterchief.app/healthz          # {"status": "ok", ...}
docker compose run --rm web python manage.py check --deploy

check --deploy is what catches an env file that forgot the HTTPS flags: they default to off in code, because defaulting them to not DEBUG would redirect every test request to https and break the suite anywhere DEBUG is unset.

The first docker compose up will take a minute or two: Caddy is provisioning the wildcard certificate over DNS-01, and DNS propagation is not instant. Watch it with docker compose logs -f caddy.

Migrations

Deliberately not run by the container's entrypoint. With more than one web container they would race, and a starting gunicorn worker is a bad place to discover a failed migration. Run them once, explicitly, as part of the deploy:

docker compose build
docker compose run --rm web python manage.py migrate
docker compose up -d --no-deps web

Scheduled jobs

Two commands need to run on a schedule. Put them on the host, not in a container, and on exactly one node when you have several — three nodes archiving the same club is three emails to the same club.

# Bill: archive clubs unpaid past their grace period.
# Run it WITHOUT --commit for the first week and read the output. The flag exists because
# this switches off paying customers: a bad clock or a bad cron should cost you an email,
# not a morning of angry clubs.
0 6 * * *  cd /srv/rosterchief && docker compose run --rm web python manage.py archive_overdue_clubs --commit

# Events: extend recurring series so the calendar never runs dry.
0 3 * * *  cd /srv/rosterchief && docker compose run --rm web python manage.py extend_event_series

Maintenance mode

Control panel → Features → Maintenance mode. While it is on:

  • every club subdomain serves a 503 maintenance page, in that club's own colours;
  • the control panel and the sign-in screens stay open, because closing them would leave you with no way to turn it back off;
  • /healthz keeps answering on every host, or the load balancer would take the node out of rotation and the control panel with it;
  • the scheduled jobs stand downarchive_overdue_clubs, extend_event_series and import_members_csv refuse to run.

migrate and collectstatic are deliberately not blocked. Maintenance is usually declared in order to run them, and a guard that stopped them would mean turning the mode off to do the work you turned it on for.

The scheduled jobs exit non-zero while the platform is closed, so cron will mail you. That is intended: a job that silently skips itself is how a month of billing goes missing. If you genuinely mean to run one during a window, pass --ignore-maintenance.

So a migration-heavy deploy looks like:

# 1. Close the platform in the control panel (or from a shell):
docker compose run --rm web python manage.py shell -c \
  "from features.models import Maintenance; Maintenance.start(message='Upgrading. Back by 21:00.')"

# 2. Do the work — migrate is not blocked.
docker compose build
docker compose run --rm web python manage.py migrate
docker compose up -d --no-deps web

# 3. Reopen from the control panel.

The state lives in Redis as well as the database, so it takes effect on every worker and every server at once — a per-process cache would leave some workers still serving clubs.

Behind an existing Caddy (dev / test server)

If the box already runs Caddy on :80 and :443 — a test server sharing a host with other sites — do not run ours: two Caddies cannot both hold port 80. Run the app only, publish it on the loopback, and add a site block to the Caddy that is already there.

docker compose -f compose.behind-proxy.yaml up -d          # web + db + redis, no caddy

web publishes on 127.0.0.1:8001 (override with WEB_PORT). Loopback, not 0.0.0.0 — bound to all interfaces, a test instance is reachable at http://<server-ip>:8001 with no TLS, bypassing the proxy and every security header with it.

Then, in the host's Caddyfile:

test.rosterchief.app, *.test.rosterchief.app {
	tls {
		dns cloudflare {env.CLOUDFLARE_API_TOKEN}
	}

	reverse_proxy 127.0.0.1:8001 {
		header_up X-Forwarded-Proto {scheme}
		header_up X-Real-IP {remote_host}
	}
}

Three things this needs, and each one is a way to lose an afternoon:

  1. The host's Caddy must have the DNS plugin too. The wildcard is still a wildcard: the stock caddy package cannot answer a DNS-01 challenge. caddy add-package github.com/caddy-dns/cloudflare on a package install, or run a Caddy built like deploy/caddy/Dockerfile.
  2. Give the test instance its own subdomain tree (*.test.rosterchief.app) and set ROSTERCHIEF_BASE_DOMAIN=test.rosterchief.app. It drives tenant resolution, the shared session cookie and the WebAuthn RP ID — point it at the production domain and test passkeys start colliding with real ones.
  3. header_up X-Forwarded-Proto is not optional, exactly as in the bundled Caddyfile. Without it Django believes the request is plain HTTP behind the proxy.

DNS still needs both records, pointing at the test box:

A   test.rosterchief.app    -> <server ip>
A   *.test.rosterchief.app  -> <server ip>

The compose project is named rosterchief-test, so its containers and volumes never collide with a production stack on the same host.

Automated backups

deploy/backup.sh dumps the database, tars the uploads while they are still on local disk, prunes anything older than KEEP_DAYS, and — if you set BACKUP_REMOTE — copies the lot off the box with rclone.

deploy/backup.sh /var/backups/rosterchief

It writes to a .part file and only moves it into place once gzip -t says the archive is readable and non-empty. A truncated dump that looks like a backup is the failure mode worth engineering against, because you only discover it on the day you need it.

Schedule it as root on the host (single server; on several, run it on the database node):

# Nightly at 02:30, before the billing and event jobs.
30 2 * * *  cd /srv/rosterchief && BACKUP_REMOTE=b2:rosterchief-backups KEEP_DAYS=14 deploy/backup.sh /var/backups/rosterchief

# Weekly restore rehearsal into a throwaway database. This is the only line here that proves
# the others work.
0 4 * * 0   cd /srv/rosterchief && deploy/restore-check.sh

Cron mails you on non-zero exit, and the script uses set -Eeuo pipefail so it does exit non-zero. A backup script that fails quietly is worse than none, because you will believe you have backups.

Offsite matters more than frequency. A dump sitting on the same disk as the database survives a bad migration but not the server. BACKUP_REMOTE takes any rclone remote (S3, Backblaze, a second box).

Once uploads move to S3 (AWS_STORAGE_BUCKET_NAME), the script skips the media tarball: the bucket's own versioning is the backup. Turn versioning on when you create it.

Restoring

gunzip -c /var/backups/rosterchief/db-2026-07-14-0230.sql.gz \
  | docker compose exec -T db psql -U rosterchief rosterchief

The dump is taken with --clean --if-exists, so it drops and recreates rather than colliding with what is there. Rehearse it once, now, against a scratch database — not the first time you need it.

Backups (manual)

Two things carry state: Postgres and the uploads.

# Database
docker compose exec -T db pg_dump -U rosterchief rosterchief | gzip > rosterchief-$(date +%F).sql.gz

# Uploads — until they are on S3, in which case the bucket's own versioning is the backup.
docker compose cp web:/app/media ./media-backup

Restore is gunzip -c dump.sql.gz | docker compose exec -T db psql -U rosterchief rosterchief. Test it once, now, rather than the first time you need it.

Going multi-server

Nothing in the code changes. What changes is where the services live:

one server several
Postgres db container DJANGO_DATABASE_URL → your central Postgres
Cache / flags redis container managed Redis (or your existing one)
Uploads local disk S3 bucket (AWS_STORAGE_BUCKET_NAME)
Static files WhiteNoise, in the image unchanged — that is why WhiteNoise is there
Cron host crontab one node only
TLS Caddy on the box load balancer, or Caddy on each node

Drop db and redis from compose.yaml, point the URLs at the central services, and run web on as many nodes as you like behind a load balancer pointed at /healthz.

The health check tests the database and a cache round trip, not just that the process is listening — a node that cannot reach Postgres, or whose cache silently swallows writes, is not healthy, and a load balancer must not keep feeding it traffic.

Rollback

Images are the unit of rollback. Tag on build, keep the last few, and:

docker compose up -d --no-deps web   # with the previous image tag

Migrations are the exception: they don't roll back with the image. Prefer additive migrations (add a column, deploy, backfill, then stop writing the old one) so that yesterday's image still runs against today's schema.