Reliability operations¶
These controls protect worker requests while keeping experiment views responsive. The backup units are opt-in templates, not installed by the deployment workflow. Opening a PR or installing these scripts does not enable off-server backups.
Dashboard isolation¶
Experiment summary calculations use one persistent spawned process, separate from the HTTP server's Python interpreter. It opens a query-only SQLite connection and does not initialize connectors, credential files, migrations or report jobs. Workers' pairing, heartbeat, budget, submission and stop decisions still use the authoritative service, never this display cache. Owner checks and correction invalidation remain in the request path.
The existing cache bounds the queue to 128 jobs and stored JSON to 32 MiB. There is one projection in flight, a 32 MiB reply limit and a 30-second timeout; a hung child is terminated, errors back off, and a later request can start a fresh child. These are work/IPC limits, not a hard process-memory quota. The child runs at lower CPU priority on Unix. Report generation already has a separate subprocess.
Only expensive cache-miss projections are serialized. Page navigations, public report reads and cached polls have no separate dashboard admission gate. Slow clients do not occupy projection capacity while receiving an HTTP response. The existing overall HTTP connection limits still apply. The UI's polling handles temporary summary unavailability without blocking page navigation.
Projection failures log dashboard.projection.failed with only the exception's
class name in the parent's trace; no message, traceback or request identifiers
cross the child error boundary. Configuration is startup-scoped, like the parent
service's configuration. Restart the service through the normal coordinated
release process after changing experiment.yaml; hot reload is not supported.
Availability and latency¶
/healthremains the cheap process liveness endpoint./readychecks that the ledger can be read, with a 500 ms SQLite busy timeout. It returns 503 if the ledger is unavailable, without exposing paths or errors./readyalso returns only a dashboard health boolean, not queue counts. A stuck job or three consecutive projection failures marksdashboard.healthyfalse. This does not change HTTP readiness for worker traffic or trigger a server restart.
The Production availability workflow probes the homepage, liveness and readiness
from a GitHub-hosted runner once an hour, at minute 17. Each check samples three times and
fails the job if all three samples fail content validation or exceed three seconds.
This detects slow HTTP 200s and an unhealthy projection worker, not just downtime.
Scheduled Actions can be delayed; this is a baseline monitor, not a paging SLA.
Configure GitHub Actions failure notifications for the operator. No new paging
service or secret is configured by this change. At one billed minute per invocation,
the baseline is at most 744 minutes in a 31-day month, not the roughly 4,464 minutes
of a ten-minute schedule. Longer jobs, retries and other workflows add usage;
this is not a billing cap or a guarantee that the account's included quota covers
it. See GitHub's Actions billing rules.
Use a separately budgeted external monitor for faster detection.
Run the same read-only probe manually:
python3 scripts/check_site_health.py --url https://staging.scorebench.dev
To exercise a known owner's populated experiment, add --experiment-id exp_...
and --cookie-file /private/path/session.cookie. That file must be owned by the
caller, mode 0600, non-symlink, and contain one harness_admin=... cookie. It is
sent only to the experiment check on the specified origin. Cross-origin redirects
are refused, expired-login redirects fail, and no response body or cookie is
printed. Supply an existing session securely, never via a command-line cookie.
The default scheduled workflow has no login credentials, so it does not prove
that a specific experiment's cached summaries are fresh. The authenticated probe
checks that all summaries are ready and error-free; it does not verify their scores.
Backups and recovery drills¶
scripts/scorebench_backup.py uses SQLite's online backup API, including committed
WAL data without stopping the server. It copies all pages in one backup step
(pages=-1), so concurrent heartbeats cannot repeatedly restart a chunked copy.
WAL-mode writers can still commit; rollback-journal databases can block writers
during the copy. The progress callback checks the 120-second deadline between
attempts and after the native step, not during that step. The systemd unit bounds
the entire backup job. See SQLite's backup API.
It creates a new private destination, checks
SQLite integrity and foreign keys, records SHA-256 and key-table counts, then
restores to a disposable new database and verifies those counts. It never restores
over a live ledger or overwrites an existing destination.
python3 scripts/scorebench_backup.py create \
--database /path/to/run/harness.sqlite \
--destination /private/backups/new-snapshot
python3 scripts/scorebench_backup.py restore-check \
--snapshot /private/backups/new-snapshot \
--destination /private/drills/new-restore
This is a ledger backup, not a complete service backup. Candidate artifacts, traces, uploads, run configuration, credentials, state/signing material and TLS configuration outside SQLite need a separately protected backup. Do not use a ledger-only restore as evidence that the whole service can recover.
For off-server copies, provision an encrypted restic repository on independent storage, and preserve its password in a separate secret store. The tool neither purchases storage nor initializes repositories. Put the repository URL and password in separate owner-only files. Keep cloud credentials in private environment/configuration files, not command arguments. The remote address is an operator responsibility; pointing it back at this host is not off-server protection.
python3 scripts/scorebench_backup.py offsite \
--database /path/to/run/harness.sqlite \
--destination /private/backups/new-snapshot \
--repository-file /private/config/restic-repository \
--password-file /private/config/restic-password
Success requires uploading the snapshot, downloading its exact manifest and ledger
from restic, verifying checksums and restoring the downloaded ledger. Upload errors,
timeouts or failed checks exit nonzero. Restic output can contain credentials, so
only the non-secret snapshot id and verification results are printed. After a
verified remote restore, an owner-only uploaded.json receipt authorizes local
retention. Keep the newest three verified local snapshots (including the current
one) by default; --keep-local N overrides this. Only private sibling directories
named ledger-YYYYMMDDTHHMMSSZ with exactly the ledger, manifest and matching upload
receipt are eligible. Unverified, failed, modified or manually named snapshots,
symlinks and the live ledger are never pruned. Remote restic retention is separate
and is not changed by this tool. A preflight requires space for three database copies plus 1 GiB;
otherwise it fails before writing. This is a headroom check, not reserved disk
space. Monitor disk usage and investigate retained failed backups; automatic
retention does not cap space used by failure evidence or unrelated files.
For scheduled backups, adapt the three infra/systemd/scorebench-ledger-backup.*
templates to the deployment's user and actual ledger path, particularly the
service's ReadWritePaths. SQLite may need to open its WAL sidecars even for a
read-only connection. Create the private backup root and restic cache directory
first. The private environment file supplies SCOREBENCH_DATABASE,
SCOREBENCH_BACKUP_ROOT, RESTIC_REPOSITORY_FILE, RESTIC_PASSWORD_FILE and any
storage-provider credentials. SCOREBENCH_BACKUP_KEEP_LOCAL optionally overrides
the three-copy local retention. Install the unit and timer only after an operator
has run a successful staging backup and remote restore check. Enable the timer
and configure failed-unit notifications; check systemctl list-timers and the
last successful journal entry daily. A timer alone is not proof of backup success.
The hourly timer has up to five minutes of jitter. Backups are serialized and each network command has a ten-minute deadline. Recovery-point age depends on the last successful verified backup; no RTO is promised by these templates. Practice a full service restore on an isolated host with submissions and paid helpers disabled.
No-inference load test¶
python3 -m pytest -q tests/test_reliability_load.py
SCOREBENCH_SOAK_ROUNDS=60 python3 -m pytest -q \
tests/test_reliability_load.py::test_32_worker_heartbeats_during_summary_refresh
These tests start a disposable loopback server and 32 simulated swarm workers. They check heartbeat latency, concurrent summary refresh and readiness, including a deliberately blocked projection, concurrent polls and six slow page/report readers. They never use a production URL, cloud runner, model API or real venue credential. The default two-round smoke runs in normal CI; the longer soak is bounded to 120 rounds. Run the soak on a staging-sized host before release and compare its printed timings. This is a server-load check, not proof of cloud provisioning or provider behavior.
Rollout¶
Validate exact-head CI and the no-inference soak, deploy to staging, then verify
/health, /ready, an authenticated experiment view, correction invalidation and
ordinary worker heartbeats. Enable backup automation only after its own restore
drill. Coordinate a production rollout around active experiments; do not restart,
re-pair, recover or replace workers to test this change. A normal release rollback
needs no schema migration. Reverting to a version without /ready also requires
disabling or adapting the new availability probe.