ScoreBench Architecture¶
ScoreBench coordinates controlled coding-agent experiments and records their results. Users design experiments in the browser; a local coordinator launches isolated workers; ScoreBench stores the protocol, scoped assignments, candidate bundles, official scores, measured usage, and audit trail.
The hosted composer offers Anthropic Take-Home and Prop AMM through the
paradigm_puzzles adapter. Other adapters remain available for configured
deployments; see Integrations and Adapters.
The core rule is: official submissions go through ScoreBench, which owns the venue credential and result ledger. Workers may use their model provider, local tools, and the official problem page as allowed by the run's constraints. ScoreBench is not a blanket network proxy or an enforced network-egress sandbox.
Repository Boundary¶
This repository contains the product and its worker distribution:
- HTTP daemon, experiment composer, and web UI
- pinned worker distribution, launcher, and supervisor integration
- agent-facing CLI
- connector layer for challenge platforms
- SQLite ledger and immutable candidate bundles
- credential storage and scoped run tokens
- deterministic reports and charts
The agent instructions live separately in the ScoreBench skill repository:
https://github.com/josusanmartin/scorebench-skill
The skill should teach an agent how to use the ScoreBench API and CLI. It should not contain connector secrets, server internals, or challenge platform API keys.
High-Level Shape¶
Browser -> ScoreBench: protocol, recipes, budget, launch prompt, results
Local coordinator -> isolated Docker workers: prepare, launch, monitor
Worker -> model provider and local tools: solve and test
Worker -> scorebench CLI -> ScoreBench: scoped submissions and usage
ScoreBench -> Paradigm Puzzles adapter: official evaluation
ScoreBench -> ledger + immutable artifacts -> Charts and Export Studio
flowchart TB
User["User browser"] --> Server["ScoreBench daemon and service"]
User -->|launch prompt| Coordinator["Local coordinator"]
Coordinator --> Worker["Isolated worker: coding harness, model, ScoreBench CLI"]
Worker -->|scoped API| Server
Worker --> Provider["Model provider"]
Worker --> Tools["Local tools and permitted problem reads"]
Server --> Adapter["Paradigm Puzzles adapter"]
Adapter --> Venue["Official evaluation"]
Server --> DB[("SQLite ledger: harness.sqlite")]
Server --> Credentials["Encrypted venue credentials"]
Server --> Bundles["Immutable candidate bundles and sanitized traces"]
DB --> Reports["Charts and Export Studio"]
Reports --> User
Venue credentials remain server-side. Each worker has a separate run token and isolated workspace. Model-provider credentials are a different boundary: the launcher copies only the required credential or uses the Claude access-only relay described in the launch runbook. Workers do not receive the owner's ScoreBench account session.
The public deployment presents the experiment workflow at scorebench.dev.
The root route redirects to /puzzles/scorebench; charts, runs, credentials,
and reports are private and scoped to the signed-in owner. Compatibility
hostnames for the Paradigm handoff render the same product against the same
service layer, queue, encrypted credentials, and authoritative ledger.
Main Components¶
Daemon And Web UI¶
challenge_harness.daemon runs the local or public HTTP service. It serves:
- user login
- immutable experiments, appended run batches, and owner-scoped monitoring
- credential management
- exercise/run token creation
- charts and reports
- agent API endpoints used by the
scorebenchCLI - static assets such as CSS, reports, and favicon
The daemon is intentionally thin. It parses HTTP requests, renders pages, and
delegates most business logic to HarnessService.
Experiment pages render metadata immediately, without waiting for worker history
projections. The owner-only status route with view=summary returns a display
snapshot and readiness counts. One bounded background queue fills missing or
changed per-worker summaries; unchanged completed runs reuse their summaries.
The in-memory cache is limited to 32 MiB, 1,024 entries and 128 queued jobs per
server process. Corrections, deletions and scope changes invalidate old results
before they can be displayed. Owner authorization is checked on every request.
The browser polls every two seconds while summaries load and every 15 seconds for active experiments, with one request in flight, a ten-second timeout and bounded retry backoff. Hidden pages stop polling. The Overview requests score trajectories separately from the initial HTML; full chart reports remain lazy. Readiness counts measure completed worker summaries, not estimated time left. Failures retain explicitly refreshing data and retry automatically.
These snapshots are non-authoritative display data, not budget or lifecycle
inputs. Worker APIs and the status route without view=summary retain fresh
ledger reads. No accounting, stop, recovery or cleanup decision trusts the
display cache.
HarnessService¶
challenge_harness.service.HarnessService is the orchestration layer. It owns
the rules for:
- experiment protocols, run recipes, launch assignments, and execution gates
- owner-scoped experiment status, stop, exclusion, and recovery controls
- admin authentication and session behavior
- credential profile lookup
- exercise/run token validation
- run start, run current, run progress, run ping, and usage accounting
- run-scoped post-run trace artifact validation and retrieval
- candidate submission validation
- idempotency and duplicate-content checks
- connector execution
- pending result refresh
- report generation triggers
If the behavior affects correctness, security, or accounting, it belongs in the service layer rather than in HTML templates or connector-specific code.
SQLite Ledger¶
challenge_harness.db.HarnessStore is the durable ledger. The database lives in
the run directory as harness.sqlite.
The ledger records the facts that reports should trust:
- experiment and run metadata
- credential profile bindings
- scoped exercise/run tokens
- candidate sequence numbers
- content hashes
- immutable bundle locations
- submission timestamps
- connector status and raw response metadata
- normalized scores
- token snapshots and usage deltas
- run heartbeats and elapsed-time evidence
- append-only passive timing evidence and dedicated observer-token scope
- per-user chart visibility and candidate-source publication allowlists
SQLite provides a transactional ledger that is straightforward to back up and inspect. WAL mode allows report readers to coexist with live writes.
Sanitized agent traces are intentionally not stored in SQLite. The skill
produces bounded gzip NDJSON only after a run; the service validates and stores
it under artifacts/run_traces/ using a hash-derived run scope and trace ID.
Trace upload/download is observational, never triggers report generation, and
does not affect run activity or candidate state.
Agent CLI¶
challenge_harness.cli provides the scorebench command (legacy alias: harness) used by solving agents.
It reads:
SCOREBENCH_URLSCOREBENCH_RUN_TOKEN
The legacy HARNESS_URL / HARNESS_RUN_TOKEN names remain supported.
Then it calls the daemon for:
scorebench contextscorebench exercisescorebench run startscorebench run currentscorebench run progressscorebench run pingscorebench run usagescorebench submitscorebench invalidatescorebench reinstatescorebench refreshscorebench bestscorebench historyscorebench solution
The CLI creates deterministic candidate bundles from files or directories and sends them to the server. It does not need platform credentials.
scorebench run progress is the stable supervisor-facing accounting contract. It
uses the same canonical timing and token normalization as reports, is restricted
to one run-token scope, and does not add a heartbeat when polled. Chart
markup and /best are presentation and best-candidate views, not progress
measurement APIs.
Timing v2 adds a separate local observer inside the worker environment from the ScoreBench skill. It tails only new structured Codex or Claude session events, emits bounded categorical spans, and authenticates with a least-privilege observer token. The service stores those spans append-only and projects active lower and upper bounds without changing the authoritative v1 clock. Connector blockage is attested only by service code after repeated structured dependency failures. See Active-Time Accounting.
Connector Layer¶
Connectors live under challenge_harness/connectors/. A connector translates a
ScoreBench candidate into a platform-specific action and normalizes the result.
The hosted experiment composer uses paradigm_puzzles for Anthropic Take-Home
(vliw) and Prop AMM (prop-amm). The adapter supports additional contracts for
configured use; that broader catalog is not the hosted picker.
Retained adapters include the private vliw judge, Tensara, HighLoad, CPU.mode,
GPU Mode, GitHub PR transport, and fake test adapter. Their commands and
credential requirements are documented in the adapter reference.
Scoped discovery commands are adapter-specific; workers must also respect the
run's restrictions on external information.
Connectors are deliberately small. They should know how to talk to a platform, but they should not decide experiment policy, token accounting, run ownership, or chart semantics.
Submission Flow¶
The supported hosted flow is:
- The user signs in, connects a Paradigm Puzzles key, and creates a reviewed experiment with explicit run recipes and a common per-run budget.
- ScoreBench records the immutable specification and issues one expiring, single-use launch prompt for the batch.
- The user's coordinator prepares the pinned Docker workers. Each worker redeems its assigned pairing and waits for its execution gate.
- The supervisor establishes the assigned run, native-session accounting, and
timing observer before releasing the model. The model inspects
scorebench context,exercise,run current, andrun progress. - The worker submits a candidate bundle and usage snapshot. ScoreBench validates scope, limits, idempotency, and metadata before invoking the adapter.
- ScoreBench records the official result or pending state with its raw evidence. Pending candidates are refreshed; unchanged candidates are not resubmitted merely to poll status.
- At exit, the supervisor reconciles usage, checks completion policy, records the lifecycle outcome, and uploads the sanitized trace. The coordinator verifies the results; retained worker artifacts remain available for recovery.
- Charts and Export Studio read the ledger. Invalidations and reinstatements append audit events without erasing the underlying evidence.
sequenceDiagram
participant User as User browser
participant Server as ScoreBench service
participant Coordinator as Local coordinator
participant Worker as Worker and supervisor
participant Venue as Paradigm Puzzles
participant Store as Ledger and artifacts
User->>Server: Create reviewed experiment
Server->>Store: Save immutable protocol and assignments
Server-->>User: Single-use launch prompt
User->>Coordinator: Paste prompt
Coordinator->>Worker: Prepare and launch isolated worker
Worker->>Server: Redeem assigned pairing; wait for execution gate
Server-->>Worker: Scoped contract and start permission
Worker->>Server: Candidate bundle and usage snapshot
Server->>Store: Save candidate and evidence
Server->>Venue: Official submission with server-held key
Venue-->>Server: Score or pending result
Server->>Store: Save result and raw evidence
Worker->>Server: Final usage, lifecycle outcome, sanitized trace
User->>Server: Monitor and compare results
Manual token creation and YAML launches remain available for operators. They are separate from the generated experiment launcher; see the manual CLI workflow.
Credential Model¶
The hosted flow stores a named paradigm_puzzles profile containing
PARADIGM_PUZZLES_API_KEY. The browser accepts the owner's key over HTTPS;
ScoreBench encrypts it and does not return it to workers. Other configured
adapters have their own credential schemas.
The optional Paradigm-native integration obtains that key server-to-server instead. It has a separate authentication contract and is disabled unless configured; see native integration.
The run token given to an agent is scoped to a specific user, connector, credential profile, and exercise. This keeps runs isolated:
agent token can submit to one exercise with one credential profile
agent token cannot list credentials
agent token cannot inspect sibling runs
agent token cannot switch to another connector
Run bearer tokens are shown only in the creation or rotation response. SQLite stores a SHA-256 token digest and prefix, not recoverable plaintext. Reissuing an active token creates a replacement with the same scope and immediately revokes the old token.
User passwords use salted PBKDF2-SHA256 hashes. A successful login against a legacy unsalted SHA-256 record upgrades that record in place; login and registration attempts are throttled per client.
Resource Boundaries¶
Run-scoped submission budgets are enforced under the candidate-allocation lock:
- candidate count per run
- cumulative compressed candidate-bundle bytes per run
- normalized working-token total per run
- distinct candidates per rolling time window
- repeated submissions of the same bundle per rolling time window
The defaults are conservative and can be overridden by the experiment's
submission_limits mapping. Submission windows are calculated from persisted
candidate timestamps and survive daemon restarts. Idempotent response recovery
happens before budget consumption, so replaying the original request key remains
safe even while a window is full. Post-run traces have independent per-upload, count, and aggregate
compressed-byte limits. HTTP request bodies are length-bounded and read with a
socket timeout; the request ceiling includes base64 expansion of the 64 MiB
candidate-bundle limit, and the threaded HTTP server has a fixed worker ceiling
plus a smaller semaphore for concurrently buffered bodies above 1 MiB.
Sensitive dynamic responses and trace downloads are marked Cache-Control:
no-store.
Cross-run code-similarity warnings combine an exact normalized-token digest,
64-bit simhash for structural edits, and a compact bottom-k shingle-containment
sketch for additive padding. Indexed simhash and containment anchors nominate a
small candidate set before scoring. Solutions with 16-79 normalized tokens get
exact-match coverage only; approximate checks start at 80 tokens to limit false
positives on short boilerplate. For exercises with a trusted public starter,
ScoreBench removes that starter's normalized shingles before comparison and
requires at least 32 residual shingles, including for exact matches. This keeps
tiny identical edits around the starter from becoming evidence of copied
solution code. Fuzzy checks additionally require at least 80 residual shingles
on both sides; the total token count cannot make a small customization look
substantial. Residual containment also uses a 0.90 minimum instead of the
raw-code 0.72 threshold and requires residual sizes to be within 2x. Exact and
SimHash matches keep their normal rules above the residual floor. The registry is
keyed by connector and
exercise, so only an explicit, versioned template can be discounted; learned
corpus-common code is never silently ignored. Existing raw fingerprints remain
compatible through residual containment, so enabling a template does not require
a database backfill for new submissions. Historical annotations can be rebuilt
from immutable, hash-verified bundles with
scripts/recalculate_code_similarity.py; the command is dry-run by default,
creates an online SQLite backup before applying, preserves unrelated warnings,
and is expected to produce a zero-change second dry-run. Production maintenance
uses --registered-template-scopes-only so unrelated binary challenge bundles
are not traversed. Similarity remains an audit warning, not an automatic
rejection.
The offline replay summary includes non-source metadata for every resulting flag: candidate and run IDs, score, match reason, distance, containment, and residual sizes. This makes threshold changes auditable before a guarded apply without exposing submitted code.
Connector Responsibilities¶
A connector should provide a narrow contract:
- declare its credential schema for the UI
- list or validate supported exercises when possible
- return an exercise statement or instructions
- submit or evaluate a candidate bundle
- refresh a pending result when the platform is asynchronous
- normalize score, status, direction, and remote IDs
- preserve useful raw response metadata for debugging
A connector should not:
- trust agent-provided scores
- read another credential profile
- decide A/B grouping
- mutate unrelated runs
- hide platform errors from the service layer
Conceptually, every connector sits behind the same boundary:
flowchart LR
Bundle["Candidate bundle<br/>source files + metadata"]
Scope["Validated token scope<br/>user + connector + credential + exercise + run"]
Service["HarnessService"]
Schema["Connector credential schema"]
Connector["Connector implementation"]
Credential["Server-held credential profile"]
Platform["Challenge platform"]
Result["Normalized ConnectorResult<br/>status + score + direction + remote id + raw evidence"]
Store[("SQLite ledger")]
Bundle --> Service
Scope --> Service
Service --> Schema
Service --> Credential
Service --> Connector
Credential --> Connector
Connector --> Platform
Platform --> Connector
Connector --> Result
Result --> Service
Service --> Store
Database And Evidence Model¶
ScoreBench stores both normalized data and evidence.
Normalized data is what the chart view uses:
candidate_id
run_name
credential_profile
connector
exercise
status
score
score_unit
direction
tokens_total
tokens_delta
active_seconds
remote_submission_id
Evidence is what lets us debug and audit:
server_received_at
server_completed_at
request hash
bundle hash
source hash
raw connector response
connector logs
trace id
remote URL or submission id
refresh attempts
warnings
The chart should be reproducible from the ledger and immutable bundles. It should not depend on agent-written progress files as authoritative data.
The current database has implementation-specific details, but the core ledger relationship is:
erDiagram
USERS ||--o{ CREDENTIAL_PROFILES : owns
USERS ||--o{ RUN_TOKENS : receives
CREDENTIAL_PROFILES ||--o{ RUN_TOKENS : scopes
CONNECTORS ||--o{ CREDENTIAL_PROFILES : defines
CONNECTORS ||--o{ EXERCISES : provides
EXERCISES ||--o{ RUN_TOKENS : scopes
RUN_TOKENS ||--o{ RUNS : starts
RUNS ||--o{ CANDIDATES : contains
CANDIDATES ||--|| BUNDLES : stores
CANDIDATES ||--o{ SUBMISSION_RESULTS : records
CANDIDATES ||--o{ USAGE_SNAPSHOTS : accounts
CANDIDATES ||--o{ TRACE_EVENTS : explains
SUBMISSION_RESULTS ||--o{ RAW_EVIDENCE : preserves
USERS {
string username
string role
}
CONNECTORS {
string name
string credential_schema
}
CREDENTIAL_PROFILES {
string name
string connector
string owner_user
string encrypted_secret_ref
}
EXERCISES {
string connector
string exercise_id
string title
}
RUN_TOKENS {
string token_hash
string user
string connector
string credential_profile
string exercise_id
}
RUNS {
string run_name
string credential_profile
string exercise_id
}
CANDIDATES {
string candidate_id
string content_sha256
string status
}
SUBMISSION_RESULTS {
string remote_submission_id
float score
string status
string direction
}
Timing, Token, And Cost Accounting¶
ScoreBench can trust server-observed facts:
- when the request arrived
- when the connector returned
- candidate sequence number
- bundle hash
- raw platform response
- normalized platform score
ScoreBench cannot inherently know true model token usage unless the runner or provider reports it. For that reason token data is recorded with provenance.
Agents are required to submit token snapshots when using the ScoreBench workflow. ScoreBench stores totals and computes deltas server-side. If token data is missing, the service should reject or warn according to the current policy rather than silently producing empty token charts.
Provider aggregates are not universally comparable. In particular, Grok's aggregate includes cached reads, so the service requires a disjoint input/output/cache-read breakdown for Grok models and recomputes the working total server-side.
Run pings and usage snapshots give the charts a better view of active work than wall-clock timestamps alone. This matters when an agent session expires, idles, or is interrupted for unrelated reasons.
Active time is explicitly an estimate. Reports retain its source, raw wall value, and cumulative idle time removed for each candidate. Candidates where at least half of clock time was excluded receive a distinct chart marker and tooltip explanation. If submitted timestamps are unavailable, the fallback caps each inter-candidate gap independently instead of collapsing all later candidates to one timestamp.
Cost uses a recorded provider cost when available, otherwise the report derives
an API-equivalent estimate from usage and the pricing catalog. Neither is a
verified provider invoice.
The report builder prices model token categories with the validated,
provider-specific YAML catalog under challenge_harness/pricing/; legacy
aggregate-only runs receive a visibly marked estimate. Unknown and composite
models remain unpriced. See API Cost Accounting.
Pending And Stale Results¶
Some platforms return a result immediately. Others take minutes and require polling.
The connector layer can return a pending status. ScoreBench records that state and uses refresh paths to complete the result later:
submitted -> scored
submitted -> rejected
submitted -> failed
Queued private VLIW jobs and asynchronous adapter submissions can finish after
the initial request. Background refresh or explicit scorebench refresh
reconciles their results. Keep the candidate identity and inspect the recorded
status instead of submitting the same bundle again.
Reports And Charts¶
challenge_harness.report builds deterministic report artifacts from the
SQLite ledger and candidate bundles.
Important outputs include:
- strategy comparison charts
- per-exercise report JSON
- run summaries
- TSV/CSV exports for analysis
- raw evidence links where available
The primary chart is the strategy comparison view. It compares runs by:
- connector
- exercise
- credential profile
- run name
- score trajectory
- active time
- wall time
- token spend
- API-equivalent model cost
- failures and rejected candidates
- promotions and best-so-far changes
Invalidated candidates are preserved for auditability but excluded from best, best-so-far curves, promotion decisions, and chart winner calculations.
The chart should answer:
Which strategy improved fastest under the same budget?
Which strategy reached the best final score?
Which strategy spent fewer tokens or less active time?
Which strategy reached a score under the lowest API-equivalent cost?
Which failures were platform errors versus rejected candidates?
The reporting pipeline is one-way: reports are derived from the ledger, not the other way around.
Report generation is isolated from the live submission path:
- writable service connections use SQLite WAL mode;
- report builders open every ledger with a true read-only, query-only connection and read each ledger within one transaction; they never run migrations;
- every output is written to a temporary file and atomically replaced;
.report-manifest.jsonrecords a source fingerprint for each artifact, so a targeted refresh cannot make unrelated stale pages appear fresh;- concurrent refresh requests are coalesced, and requests that arrive during a render are handled by a bounded catch-up loop;
- if the source changes while a report is rendering, the generated artifact remains stale and is eligible for another render after the cooldown;
- a normal stale page is served immediately while its replacement renders in the background.
Report heartbeat reads stream rows and remove only the duplicated top-level
original_prompt before retaining them. Run records still provide the displayed
prompt, and the ledger retains the full audit evidence. All usage and timing
metadata is preserved using Python JSON parsing, including legacy duplicate-key
and non-finite-number semantics; report reads do not change accounting writes.
On Linux, background report subprocesses and expctl report have a default
4 GiB virtual-address-space ceiling (RLIMIT_AS). Set the positive integer
SCOREBENCH_REPORT_MEMORY_MB in the renderer's environment to tune it to the
host's capacity; an existing tighter limit is never raised. This is a per-process
limit, not a total host budget: allow headroom for the daemon, staging, SQLite
page cache and other processes, and avoid overlapping manual builds. Exceeding
the ceiling fails the build instead of publishing a successful partial report.
The background renderer also retains its configured timeout (900 seconds by
default). These limits do not apply to the daemon or agent workers. Non-Linux
platforms do not apply this memory ceiling. A failed warmup after a deployment
restart does not roll back the running service; check health and artifact
freshness separately, and do not blindly rerun the deployment.
Ledger migrations install transactional report revision counters. Chart data changes increment these counters; authentication/session bookkeeping and token last-use timestamps do not. Each artifact records both its data revision and a safety revision. New candidates, usage, and activity can reuse the previous snapshot while a refresh runs. Deletions, invalidations, and corrections to recorded results advance the safety revision and withhold the old snapshot. Counters are per ledger, not per experiment, so a correction can conservatively withhold other charts backed by that ledger. Legacy ledgers without counters retain the conservative database/WAL fingerprint fallback.
Experiment JSON requests return the last safe, currently authorized snapshot
with HTTP 200 and report_freshness, rather than repeatedly returning 503 while
new writes arrive. Newly bound runs absent from the snapshot appear in
pending_run_ids. If no data artifact exists yet, an empty pending response
keeps controls usable without claiming there are no runs. Experiment ownership,
membership, and local visibility are checked before serving data; an empty
experiment never falls back to the owner's other runs.
Charts and Export Studio retry stale or pending data after 15 seconds, then at 30-second intervals for at most 20 attempts. They pause network work in hidden tabs, retain user controls, and replace data in place without reloading. GET polls respect the server cooldown instead of requesting urgent generation. Duplicate requests for an artifact already being rendered coalesce; requests for a different artifact remain queued for the bounded catch-up pass.
Full generation still writes the compatibility JSON, CSV, TSV, SVG, compare, and export artifacts. A stale exercise page uses incremental generation: exercise topology and run-count summaries come from the manifest, only the requested exercise data is rebuilt, and only the requested compare/data or export artifacts are atomically replaced.
Interactive compare pages contain the application shell rather than the full
candidate history. The shell fetches compact
strategy-compare-<connector>-<exercise>.json; a deterministic .json.gz
sidecar is retained as an artifact, while the daemon caches viewer-scoped gzip
responses so privacy filtering does not require repeated parsing or compression.
Chart query filters can be applied before JSON transfer. Export Studio also loads authorized report data. Downloaded HTML snapshots
are self-contained so presentations work offline.
Generated comparison payloads retain all runs. The HTTP layer requires an authenticated account and scopes each JSON response to that owner's locally visible runs. An optional experiment identifier narrows the payload to the exact run keys bound to that immutable experiment. Local hide state therefore affects only the owner and cannot expose a run to another viewer.
Candidate source has a separate allowlist. /ui/candidates/code reads bounded
text from the immutable candidate bundle and is owner-only by default. An owner
can publish or revoke one candidate without exposing its run charts;
anonymous access succeeds only while that exact candidate has a matching publication row.
The page highlights source server-side and its copy control re-fetches the same
authorized JSON representation, so raw source is never interpolated into client
script.
The admin web keeps chart viewing and report management on separate routes:
/ui/dashboardsopens Charts;/ui/reports/redirects to the deployment's default exercise report./ui/runsis the personal Runs index: searchable run history, isolated chart links, local hide/show, and permanent deletion. Run charts remain owner-scoped; candidate source publication is a separate control./ui/keyscreates and manages scoped exercise API keys under Account./ui/docs/serves the generated MkDocs site and its search index from the run directory.
flowchart LR
DB[("SQLite ledger")]
Bundles["Immutable bundles"]
Raw["Raw connector evidence"]
Builder["Report builder<br/>challenge_harness.report"]
JSON["report.json<br/>exercise-report.json<br/>compact chart JSON + gzip"]
TSV["TSV / CSV exports"]
HTML["strategy-compare.html<br/>application shell"]
Manifest[".report-manifest.json<br/>freshness + summaries + timings"]
UI["Web UI report pages"]
DB --> Builder
Bundles --> Builder
Raw --> Builder
Builder --> JSON
Builder --> TSV
Builder --> HTML
Builder --> Manifest
JSON --> UI
HTML --> UI
Logging And Tracing¶
ScoreBench should log enough to debug connector issues without exposing secrets.
Useful log layers are:
- request log: HTTP method, path, status, duration, trace id
- service log: run token scope, candidate id, validation decisions
- connector log: platform request attempts, remote IDs, status transitions
- trace log: structured per-submission timeline
- error log: exceptions, connector failures, refresh failures
Secrets must be redacted from logs. The right debugging handle is the trace id, candidate id, remote submission id, and credential profile name, not the underlying cookie or API key.
GET /health includes a reports object with the active worker state, queued
scope, and last generation metrics. reports.generate.completed trace events
include total time, per-phase time, selected scopes, candidate and artifact
counts, payload sizes, and whether the source changed during generation.
Adding A New Connector¶
To add a connector:
- Create
challenge_harness/connectors/<connector_name>.py. - Define the credential schema used by the web UI.
- Register the connector in
challenge_harness/connectors/__init__.py. - Implement exercise lookup or validation.
- Implement submit/evaluate.
- Implement refresh if results are asynchronous.
- Normalize status, score, direction, and remote submission ID.
- Preserve raw response metadata for debugging.
- Add tests with fake HTTP responses or a fake CLI.
- Update the ScoreBench skill only if the agent-facing workflow changes.
The target is a connector that is boring and auditable. Platform-specific quirks should stay inside the connector, while run policy and experiment accounting stay in the ScoreBench service and ledger.
Operational Modes¶
The hosted product runs behind an HTTPS reverse proxy. Operators can also run ScoreBench locally for development or private deployments.
Local mode is useful for private experiments:
browser -> http://127.0.0.1:<port>/ui
agent -> http://127.0.0.1:<port>
Public mode is useful when multiple users need access:
browser -> HTTPS host -> reverse proxy -> ScoreBench daemon
agent -> HTTPS host -> reverse proxy -> ScoreBench daemon
In public mode, per-user credential isolation matters. A signed-in user should only see and manage their own credential profiles unless explicit admin tooling is added for cross-user operations.
Design Principles¶
- Keep agents lightweight: one run token, simple CLI commands, no platform keys.
- Keep the ledger authoritative: reports come from SQLite and immutable bundles.
- Keep connectors narrow: submit, refresh, normalize, preserve evidence.
- Keep credentials server-side: workers never receive venue cookies or venue API keys.
- Keep reports deterministic: fixed ordering, stable schemas, reproducible HTML.
- Keep failures visible: connector errors should be recorded and returned.
- Keep experiments isolated: one credential profile and one run scope per token.