Add maintenance mode: lock the platform down from the control panel
Closes every club subdomain with a 503 in that club's own colours, stands the scheduled jobs down, and keeps open exactly what is needed to end it again. The exemptions ARE the feature: - /accounts/ stays open on the base domain. Close it too and you cannot sign in to turn maintenance off -- a lock-down with no key, fixable only from a shell. - /healthz answers on every host. Close it and the load balancer decides the node is dead, stops routing to it, and takes the control panel down with everything else. - migrate and collectstatic are NOT blocked. Maintenance is usually declared in order to run them; a blanket guard on BaseCommand would mean turning the mode off to do the work you turned it on for. Only the domain jobs (archive_overdue_clubs, extend_event_series, import_members_csv) refuse, and they exit non-zero so cron mails you -- a scheduled job that silently skips itself is how a month of billing goes missing. The state is cached with a 10-second TTL, not for ever. Write-through makes the flip instant for the shared Redis of a real deployment, and the TTL is the belt to that braces: on a per-process cache -- a dev box with no Redis, or a misconfigured deploy -- a lock-down that reached only one gunicorn worker would be worse than useless. Live-verified: a club subdomain, its login page and the base domain all 503 while the control panel and the sign-in screens stay up. Also adds the two deployment pieces asked for: compose.behind-proxy.yaml for a dev/test box that already runs Caddy on :80 (app on the loopback, host Caddy proxies to it -- and the host's Caddy still needs the DNS plugin, because the wildcard is still a wildcard), and deploy/backup.sh + restore-check.sh with a cron schedule. The backup writes to a .part file and only lands it once gzip -t says it is readable: a truncated dump that looks like a backup is the failure you find on the day you need it. The weekly restore rehearsal is the only line in that cron that proves the rest work. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
31
features/commands.py
Normal file
31
features/commands.py
Normal file
@@ -0,0 +1,31 @@
|
||||
"""Scheduled work refuses to run while the platform is locked down.
|
||||
|
||||
Deliberately opt-in, per command, rather than a blanket guard on BaseCommand: maintenance is
|
||||
usually declared IN ORDER to run `migrate` or `collectstatic`, and a guard that blocked those
|
||||
would make the mode useless — you would have to turn it off to do the work you turned it on
|
||||
for. Only the domain jobs (which write club data, archive clubs, or import members) stand
|
||||
down.
|
||||
"""
|
||||
|
||||
from django.core.management.base import BaseCommand, CommandError
|
||||
|
||||
from features.models import Maintenance
|
||||
|
||||
|
||||
class MaintenanceAwareCommand(BaseCommand):
|
||||
"""A command that must not run while the platform is closed."""
|
||||
|
||||
def execute(self, *args, **options):
|
||||
if Maintenance.is_on() and not options.get("ignore_maintenance"):
|
||||
raise CommandError("The platform is in maintenance mode; this command stands down. Pass --ignore-maintenance to override.")
|
||||
|
||||
return super().execute(*args, **options)
|
||||
|
||||
def create_parser(self, prog_name, subcommand, **kwargs):
|
||||
parser = super().create_parser(prog_name, subcommand, **kwargs)
|
||||
parser.add_argument(
|
||||
"--ignore-maintenance",
|
||||
action="store_true",
|
||||
help="Run even though the platform is in maintenance mode.",
|
||||
)
|
||||
return parser
|
||||
Reference in New Issue
Block a user