Paradigm Puzzlesexperiments by ScoreBench

Public ScoreBench Deployment

This operator guide describes the repository's deployment configuration. Service names and filesystem paths retain legacy harness identifiers; these are compatibility details, not the product name. Local experiment workers do not need to deploy a server.

The public host exposes harnessd through Nginx and keeps the daemon bound to 127.0.0.1.

Current public URL:

https://scorebench.dev

The deployment workflow targets 161.97.77.106 as user ant. DNS, registrar, and host access are infrastructure configuration; verify their live state before changing a deployment rather than inferring it from this page.

https://scorebench.dev/ redirects to the canonical experiment product at /puzzles/scorebench. Authentication, experiments, charts, runs, and exports remain on the same origin.

Service

harnessd runs under systemd with secure admin cookies (/etc/systemd/system/harnessd.service):

[Unit]
Description=ScoreBench daemon (scorebench.dev)
After=network.target

[Service]
Type=simple
User=ant
Group=ant
WorkingDirectory=/home/ant/harness
Environment=PYTHONUNBUFFERED=1
Environment=SCOREBENCH_SECRET_ENV_KEY_FILE=/etc/scorebench/harnessd-credentials.key
ExecStart=/usr/bin/python3 /home/ant/harness/harnessd /home/ant/harness/runs/highload_sum_of_prime_numbers --host 127.0.0.1 --port 8718 --secure-cookies
Restart=always
RestartSec=3

[Install]
WantedBy=multi-user.target

Before enabling the external key path on an existing deployment, copy the current runs/credentials/.harness_secret_key to the configured path with owner ant and mode 0600. Use a distinct key path for staging. Do not remove the old key until every encrypted credential has been read successfully after restart.

Useful commands:

sudo systemctl status harnessd
sudo journalctl -u harnessd -f
sudo systemctl restart harnessd

Runtime dependencies on the host (Ubuntu 24.04): python3-cryptography, python3-yaml, python3-pytest (apt) plus mkdocs==1.6.1 and mkdocs-material==9.7.6 (user-level pip) for the admin UI Docs page.

CI/CD

Every push to main or staging runs .github/workflows/deploy.yml. The same workflow validates pull requests without deploying:

  1. Select coverage against the last successful release on the target branch. Only the explicit UI allowlist (ui_theme.py, chart_runtime.py, export_view.py, and their browser/unit tests) uses the UI path: rendering, security, generated JavaScript, and desktop/mobile browser checks. Docs-only changes use documentation tests and a strict MkDocs build/link validation. Mixed or unknown paths, daemon.py, worker/accounting/server code, MPP, dependencies, and CI policy changes require the complete Python 3.10/3.12, browser, TypeScript, docs, packaging and pinned Docker lifecycle suites. Missing baseline/API evidence also requires full coverage. Comparing with a successful release, not merely the previous commit, includes earlier failed or not-yet-validated changes.
  2. Require ci-ready: all selected jobs must succeed and all unselected jobs must be intentionally skipped. Failed, cancelled or missing jobs block deployment. On main pushes only, check the private VLIW judge; staging never redeploys production judge infrastructure.
  3. rsync of the repository to /home/ant/harness or /home/ant/harness-staging over SSH using the DEPLOY_SSH_KEY repository secret. State paths (runs/, artifacts/, venue_homes/, .harness_credentials/, .env*, .harness_secret_key*) are excluded and protected from deletion.
  4. Check the production host-local SSH tunnel to the VLIW judge. Judge/tunnel installation is skipped only when a stored input fingerprint matches and the service is healthy. Fingerprints include deployment policy and the tunnel key where applicable. Missing receipts, changed inputs, failed health checks and manual reruns force installation. Receipts are removed before mutation and written only after successful health checks, so a failed deployment cannot certify partial installation.
  5. Restart the branch's daemon and confirm health plus the private access contract: HTML reports redirect to login and report JSON rejects anonymous requests. The daemon acknowledges featured-report warmup through its existing background queue; a multi-minute render no longer delays deployment.

The deployment installs deploy/scorebench-report-warmup.conf for the selected daemon only. Startup skips warmup when all configured artifacts exist with the current renderer digest. Changed renderers or missing artifacts queue a full refresh through the same serialized subprocess used by on-demand reports (4 GiB address-space cap by default; startup renders have at most 900 seconds each). There is no separate expctl writer competing with the daemon. Live-data freshness still follows the existing on-demand refresh and snapshot safety rules; renderer_ready is not a claim that no new submissions have arrived.

Warmup status is in <run_dir>/.report-warmup.json. It identifies the daemon PID and renderer digest and distinguishes queued, running, renderer_ready, failed, and interrupted. A normal deploy verifies that its own daemon acknowledged warmup, not that pending work is complete. A separate report-warmup job waits up to 31 minutes, outside the deployment lock and reserved runner, then verifies the same startup's random warmup identifier and renderer digest. It fails the workflow on a render failure, deadline or interruption. A changed startup identifier with the same renderer digest (daemon restart or docs-only deployment) is followed within the original deadline; a crash loop cannot reset that deadline or certify reports. A different identifier and renderer digest exits successfully with an explicit superseded notice, not a report certification. The newer deployment owns its own verification. A changed digest for the same identifier still fails. The site need not wait for that job to serve requests. The verifier uses a GitHub-hosted runner only while warmup is pending, consuming hosted Actions minutes without occupying the self-hosted test/deploy pool. It polls public HTTPS health only; no deploy SSH key or production credentials are available to that job. The renderer digest conservatively hashes every package Python source, so most server-code releases need warmup. Already-current artifacts skip that job. Superseding warmup jobs cancel the old poller, never the deployment. Failures also emit reports.warmup.failed. The new /health field reports.startup_warmup exposes only the warmup phase, action, random identifier and renderer digest, not PID, filenames or error details. Other pre-existing health fields are unchanged. To verify eventual completion:

PYTHONPATH=/home/ant/harness python3 /home/ant/harness/scripts/check_report_warmup.py \
  /home/ant/harness/runs/highload_sum_of_prime_numbers \
  --pid "$(systemctl show harnessd --property=MainPID --value)" --require-complete

Use the staging paths and service for staging. A stale receipt from an earlier PID or renderer cannot pass this check; completed status also revalidates the artifact manifest. Health and access checks never wait for rendering and are never skipped. Graceful service closure records interrupted; abrupt termination can leave a pending receipt, which cannot certify a new daemon. The next start rechecks artifacts. Warmup setup/status-write failures are logged without preventing the API from starting, but cannot pass the deployment verification.

On the first attempt of a release push, an immutable validation receipt can reuse a successful same-repository PR run created in the last 24 hours (rerunning an older run does not extend this conservative age window). The full Git tree, workflow digest, target branch, run/attempt, artifact checksum and required GitHub job conclusions must match. Scoped receipts additionally bind the validated baseline tree and coverage profile; a UI receipt cannot certify worker changes. Full validation can cover narrower changes, not vice versa. Missing, expired, inaccessible or inconsistent receipts run tests instead. Successful failed-job reruns are eligible: the receipt and independently fetched job results must belong to the latest successful attempt, including preserved successful jobs. The gate rechecks that the attempt did not change during verification. It first looks up the merged PR's exact head, with bounded retries for an empty listing, rather than trusting GitHub's cached success-only index. Reusing validation never creates another receipt. This reuses test evidence, not a mutable Docker tag, dependency cache, or a staging build with different source. The release still checks out and deploys its own exact tree.

Tests use the datager pool; planning, gating, diagnostics and deployment use the reserved datager-deploy runner. The scorebench-deploy and vliw-judge locks and both live-payment idle guards remain mandatory. A weekly scheduled run on the default branch and manual workflow dispatches run all suites without deploying. Rerunning the whole workflow disables validation reuse and scoped selection. Retrying only a failed job retains the already-completed checks. Do not add whole-workflow path filters: ci-ready must always report a verdict.

Full validation separates Python 3.10/3.12, experiment/chart browsers, wallet verification, packaging/docs/TypeScript, and Docker lifecycle checks. A failed browser job can be rerun without repeating the Python suite. Wallet bundles are built and verified once, then downloaded by dependent jobs from the same run; missing bundles fail before tests, so native helper coverage cannot silently skip. Pip/npm download caches speed installation but never authorize skipping tests. Pytest uses two workers per job rather than every visible CPU; matrix concurrency is capped at two and the three-runner pool bounds heavy jobs across workflows. Test scratch remains on job-managed disk, not the small /tmp tmpfs. Readiness assertions wait for loaded state; tests of actual deadlines retain their deliberately short timeouts.

Manual worker changes span this repository's CLI and the separate scorebench-skill repository. Release the compatible skill to staging before deploying the staging CLI, and to main before deploying the production CLI. The generated setup prompt selects that environment's skill branch. Pi/OpenCode OpenRouter support requires openrouter_run.py --check --expected-model --expected-effort plus openrouter_agents.py and openrouter_accounting.py; older helpers must fail preflight before run start. Do not deploy only one half of this change. Run the skill's opt-in native OpenCode/Pi integration tests and a small authenticated OpenRouter probe before advertising live-provider validation. Local provider fixtures prove routing and accounting behavior, not OpenRouter availability or credential health. Already-running workers do not acquire updated helpers.

VLIW Judge

The VLIW judge binds only to 127.0.0.1:8787 on its host. The public ScoreBench host reaches it through scorebench-vliw-tunnel.service, which forwards local port 8790 over key-only SSH. No judge HTTP port is open in the VPS firewall.

The deployment installs services/vliw_judge/ under /opt/scorebench-vliw-judge, downloads problem.py from a pinned upstream commit, verifies its SHA-256, and restarts vliw-judge.service. Queue state is stored in /var/lib/vliw-judge/judge.sqlite; queued and running payloads live under /var/lib/vliw-judge/jobs and are removed after a terminal result.

Useful checks:

sudo systemctl status scorebench-vliw-tunnel.service
curl -fsS http://127.0.0.1:8790/health

# On the judge host:
sudo systemctl status vliw-judge.service
sudo journalctl -u vliw-judge.service -f
curl -fsS http://127.0.0.1:8787/health

Staging

https://staging.scorebench.dev is a second instance on the same VPS used to review UI and behavioral changes before they reach production:

  • Pushes to the staging branch deploy to /home/ant/harness-staging and restart harnessd-staging (same deploy.yml, branch-parameterized).
  • The staging daemon runs from /home/ant/harness-staging/runs/highload_sum_of_prime_numbers on port 8719; nginx vhost staging.scorebench.dev proxies to it and sends X-Robots-Tag: noindex, nofollow plus a deny-all robots.txt so the site is never indexed.
  • One-time provisioning is automated by the Staging Setup workflow (.github/workflows/staging-setup.yml, workflow_dispatch). It bootstraps the code dir from the prod checkout and can refresh staging report data from an online production SQLite snapshot. A refresh preserves staging run tokens, arm identities, run-visibility preferences, candidate-source publications, run_state.json, credentials, and secret keys; rewrites production filesystem paths to the staging root; keeps a timestamped rollback snapshot; and atomically swaps the integrity-checked database while only harnessd-staging is stopped. Run the refresh with:
gh workflow run staging-setup.yml --ref staging \
  -f seed_data=true -f run_certbot=false

The same workflow also creates the initial staging run directory, installs the systemd unit and nginx vhost, and can run certbot once the staging DNS A record (Cloudflare, DNS-only) exists. - Rollout flow: land risky changes on staging, review live, then merge staging into main.

One-time index prebuild

The covering containment index (idx_candidate_code_containment_cover) replaces the old lookup index on the first writable startup. On a production-sized database that build takes about 17 seconds, longer than a deploy should keep the daemon down, so build it once beforehand during an approved quiet window. Do this on staging first, after the usual idle checks:

timeout --kill-after=10 300 python3 scripts/prebuild_containment_index.py \
  runs/highload_sum_of_prime_numbers/harness.sqlite

The script waits at most two seconds for the write lock, then creates only the new index and keeps the old one. While it builds, daemon writes wait, so submissions and pings can stall for the whole build. That is about 17 seconds at production size, and up to the timeout limit if it runs slowly; plan the window for that. It then verifies the table, the column order, and that the index is neither unique nor partial.

An interruption before the commit rolls back. After the commit, the new index may already exist even if the script then failed or printed nothing, which is harmless. Rerun it; it is idempotent and never touches the old index. Deploy only after it exits 0 and prints ... ready. The next normal restart then drops the old index.

Paradigm Puzzles Product Origin

The current product lives at https://scorebench.dev/puzzles/scorebench and uses ScoreBench authentication. The optional native integration embeds the composer in Paradigm's Next.js application and uses Paradigm's X session. It requires a separate rollout; its API is disabled without an integration token. To enable it, configure a distinct server-only secret per environment:

Environment=SCOREBENCH_PARADIGM_INTEGRATION_TOKEN_FILE=/etc/scorebench/paradigm-integration-token

The matching Paradigm server-only environment uses SCOREBENCH_URL and SCOREBENCH_PARADIGM_INTEGRATION_TOKEN. See Paradigm-native experiment composer for the contract and rollout checklist.

The token file must be owned by the ant service account with mode 0600; startup fails closed if group or other permissions are present.

The canonical ScoreBench product and API use:

https://staging.scorebench.dev  # staging
https://scorebench.dev          # production

The compatibility hosts paradigm-staging.scorebench.dev and paradigm.scorebench.dev render the same experiment product while the native Paradigm handoff is integrated. They are aliases, not separate daemon processes. Each environment has one service layer, encrypted credential store, run ledger, submission queue, and report worker.

Provision an origin with the idempotent workflow after creating a DNS-only A record to 161.97.77.106:

gh workflow run paradigm-site-setup.yml --ref staging \
  -f environment=staging -f run_certbot=true

Use environment=production from main only after the staging product surface has been reviewed. The workflow preserves an existing Certbot-managed Nginx file, verifies the host policy directly against the daemon, and refuses to request TLS unless DNS resolves to the ScoreBench VPS.

Diagnostics

Every HTTP response includes an X-Harness-Trace-Id header. If an agent or UI action fails, copy that trace ID and search the run logs:

TRACE_ID=trc_...
rg "$TRACE_ID" /home/ant/harness/runs/highload_sum_of_prime_numbers/logs

The daemon writes structured JSONL files under the active run:

runs/highload_sum_of_prime_numbers/logs/harness.jsonl  # all structured events
runs/highload_sum_of_prime_numbers/logs/trace.jsonl    # request-scoped events
runs/highload_sum_of_prime_numbers/logs/errors.jsonl   # warnings/errors and tracebacks

Logs redact credential values, cookies, bearer tokens, passwords, CSRF values, submitted bundle bodies, and common provider-key formats even when they appear inside an otherwise innocuous error string.

Public-Service Safeguards

  • User passwords are salted PBKDF2-SHA256 hashes; successful legacy logins are upgraded in place.
  • Failed login and registration attempts are throttled in memory per client. Proxy client-IP headers are accepted only from the loopback reverse proxy; Nginx's validated X-Real-IP takes precedence over forwarded chains.
  • Run bearer tokens are shown once, stored only as hashes, and rotated rather than revealed. Reissuing immediately revokes the old token.
  • Each run has candidate, compressed-artifact, working-token, and rolling submit limits. Effective values are returned by authenticated GET /context.
  • Run traces have per-upload, count, and aggregate-byte limits.
  • Request bodies have size checks and a 120-second socket timeout. Dynamic authenticated responses use Cache-Control: no-store; unexpected HTTP 500 responses expose only a trace ID, while details remain in owner-only logs.

Submission rate limiting is derived from persisted candidate timestamps. The general run window and the stricter same-bundle window therefore survive daemon restarts, as do candidate, artifact, token, and trace budgets.

Report pressure is visible without opening a chart:

curl -fsS http://127.0.0.1:8718/health | jq '.reports'
rg '"event":"reports.generate' \
  /home/ant/harness/runs/highload_sum_of_prime_numbers/logs/harness.jsonl | tail -n 20
jq '{updated_at, last_generation, scope_summaries}' \
  /home/ant/harness/runs/highload_sum_of_prime_numbers/reports/.report-manifest.json

reports.running shows the active renderer. pending_full, pending_reports, and pending_reasons expose coalesced catch-up work. last_generation.phases_ms identifies database scans, exercise-data building, rendering, and publication separately. A true dirty_during_generation means submissions changed while the snapshot was rendering; the worker should perform another pass.

Nginx

The public Nginx site proxies to the local daemon (/etc/nginx/sites-available/scorebench.dev; Certbot manages the TLS parts):

server {
    server_name scorebench.dev www.scorebench.dev;
    client_max_body_size 128m;

    location / {
        proxy_pass http://127.0.0.1:8718;
        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
        proxy_read_timeout 600s;
        proxy_send_timeout 600s;
    }
}

Open the firewall for Nginx:

sudo ufw allow 'Nginx Full'

Issue or renew HTTPS (renewal is automatic via certbot.timer):

sudo certbot --nginx -d scorebench.dev -d www.scorebench.dev --redirect

Smoke Checks

curl -sS http://scorebench.dev/ui/login -D - -o /dev/null
curl -fsS https://scorebench.dev/ui/login | rg 'ScoreBench'

After login, agent tokens should use the public URL:

export SCOREBENCH_URL=https://scorebench.dev
export SCOREBENCH_RUN_TOKEN=hrun_...
scorebench context