API Cost Accounting¶
ScoreBench shows an API-equivalent model cost for each priced run. This is the cost implied by token usage and the standard synchronous public API price catalog loaded by the server. It is useful for comparing strategies under a common budget, but it is not an invoice and does not attempt to value a Claude Max, ChatGPT, enterprise, or other subscription plan.
The token and cost views deliberately count cache reads differently. Working tokens exclude cached-input reads so repeatedly loading the same context does not look like new work. API-equivalent cost includes those reads at the model's cached-input price. Cache writes remain part of working tokens and use their separate price where the provider publishes one.
The chart view and Export Studio support Cost, Working tokens, Fresh input, Output tokens, and Processed tokens as chart axes. Lower-cost candidate snapshots appear earlier on the cost axis, so the score chart answers "what quality did this run reach for this API budget?" Processed tokens add known cache reads to working tokens; a run without a cache-read counter remains visible as a labeled lower bound instead of disappearing. Export Studio's Runs view can compare one best-scoring point per run by any of these resource measures.
Price Table¶
A dash means the provider does not publish a separately billed cache-write category for that model; ScoreBench charges cache misses at the regular input price rather than treating them as free.
The current catalog version is 2026-09-22. Prices are USD per million tokens.
| Provider | Reviewed | ScoreBench model | Input | Cache hit/read | Cache write (default rate) | Output |
|---|---|---|---|---|---|---|
| OpenAI | 2026-09-22 |
gpt-6-astra |
$10.00 | $1.00 | $12.50 | $50.00 |
| OpenAI | 2026-09-22 |
gpt-6-sol |
$2.00 | $0.20 | $2.50 | $10.00 |
| OpenAI | 2026-09-22 |
gpt-6-luna |
$0.10 | $0.01 | $0.125 | $0.50 |
| OpenAI | 2026-09-22 |
gpt-5.6-sol |
$4.00 | $0.40 | $5.00 | $20.00 |
| OpenAI | 2026-09-22 |
gpt-5.6-terra |
$2.00 | $0.20 | $2.50 | $12.00 |
| OpenAI | 2026-09-22 |
gpt-5.6-luna |
$0.20 | $0.02 | $0.25 | $1.20 |
| OpenAI | 2026-09-22 |
gpt-5.5 |
$5.00 | $0.50 | - | $30.00 |
| OpenAI | 2026-09-22 |
gpt-5-codex |
$1.25 | $0.125 | - | $10.00 |
| Anthropic | 2026-09-22 |
claude-fable-5-1 |
$10.00 | $0.25 | $20.00 | $50.00 |
| Anthropic | 2026-09-22 |
claude-fable-5 |
$10.00 | $1.00 | $20.00 | $50.00 |
| Anthropic | 2026-09-22 |
claude-opus-5-5 |
$4.00 | $0.20 | $8.00 | $20.00 |
| Anthropic | 2026-09-22 |
claude-opus-5 |
$5.00 | $0.50 | $10.00 | $25.00 |
| Anthropic | 2026-09-22 |
claude-sonnet-5 |
$2.00 | $0.20 | $4.00 | $10.00 |
| Anthropic | 2026-09-22 |
claude-haiku-4.5 |
$1.00 | $0.10 | $2.00 | $5.00 |
| Anthropic | 2026-09-22 |
claude-opus-4.8 |
$5.00 | $0.50 | $10.00 | $25.00 |
| xAI | 2026-09-21 |
grok-4.7 |
$2.00 | $0.50 | - | $6.00 |
| xAI | 2026-09-21 |
grok-4.6 |
$2.00 | $0.50 | - | $6.00 |
| xAI | 2026-09-21 |
grok-4.5 |
$2.00 | $0.30 | - | $6.00 |
2026-09-01 |
gemini-3.5-flash |
$1.50 | $0.15 | - | $9.00 | |
2026-09-01 |
gemini-3.1-pro-preview |
$2.00 | $0.20 | - | $12.00 | |
| Moonshot AI | 2026-08-28 |
kimi-k3 |
$3.00 | $0.30 | - | $15.00 |
| Alibaba Cloud | 2026-09-01 |
qwen3-coder-plus |
$1.00 | $0.20 | $1.25 | $5.00 |
| DeepSeek | 2026-09-01 |
deepseek-v4-pro |
$1.32 | $0.044 | - | $3.96 |
| Z.AI | 2026-08-28 |
glm-5.2 |
$1.40 | $0.26 | - | $4.40 |
Provider-specific notes:
- OpenAI: GPT-5.6 Sol uses the current promotional rate, published as available at least through November 21, 2026.
- OpenAI: GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6, and GPT-5.5 requests above their published long-context thresholds have per-request price modifiers that ScoreBench does not infer from cumulative counters.
- OpenAI: GPT-6 Sol and GPT-6 Luna (released 2026-09-22) bill prompts above 272K input tokens at 2x input and cache rates and 1.5x output.
- Anthropic: Anthropic publishes separate 5-minute and 1-hour cache-write rates. ScoreBench's generic cache-creation counter uses the conservative 1-hour rate so native Claude Code runs cannot be undercounted; provider-metered cost_usd overrides this fallback when available.
- xAI: Native Grok workers price each recorded inference separately, doubling all token rates when its prompt reaches 200k tokens. Older cumulative-only records cannot identify this tier. These are global public API-equivalent prices, not subscription charges or regional-endpoint invoices.
- Google: Gemini 3.1 Pro Preview charges higher rates for prompts above 200k tokens. ScoreBench uses the <=200k standard rate because cumulative usage counters do not preserve each request's context tier.
- Google: Gemini context-cache storage is time-based and is not represented by ScoreBench's token-only cache-creation field; provider-metered cost_usd remains authoritative when available.
- Alibaba Cloud: Qwen3-Coder-Plus uses context-length tiers. ScoreBench uses the international <=32k standard rate because cumulative usage counters do not preserve each request's context tier.
- Alibaba Cloud: Alibaba publishes both implicit-cache and lower explicit-cache-read prices. ScoreBench's generic cache-read counter uses the conservative implicit-cache rate unless provider-metered cost_usd is available.
- DeepSeek: DeepSeek publishes peak and off-peak rates. ScoreBench uses the peak rate so a cost budget is not understated; provider-metered cost_usd remains authoritative and off-peak API usage can cost half as much.
- Z.AI: Z.AI currently lists cached-input storage for GLM-5.2 as free; ScoreBench models token charges only.
Official references: OpenAI model catalog, OpenAI prompt-caching billing, Anthropic pricing, xAI pricing, Gemini Developer API pricing, Kimi K3 pricing, Qwen3-Coder-Plus pricing, DeepSeek models and pricing, Z.AI pricing.
The standard table deliberately excludes long-context surcharges, regional uplifts, batch discounts, priority or fast-mode charges, tools, infrastructure, taxes, and subscription economics. Provider-specific modifiers cannot be reconstructed reliably from cumulative run counters because their thresholds and modes apply per request.
Derivation Order¶
ScoreBench chooses the strongest available measurement for every candidate:
- A provider-reported
cost_usd, captured by the runner when present. - A cumulative input/output/cache token breakdown priced with the table above.
- A final run-level breakdown allocated to candidates in proportion to each candidate's cumulative working-token total.
- A legacy aggregate estimate using 80% input and 20% output tokens:
cost = working_tokens * (0.8 * input_rate + 0.2 * output_rate) / 1,000,000
Methods 3 and 4 are estimates. A partial token breakdown is also an estimate
when a separately priced cache category is missing. The UI prefixes every
estimated value with ~, shows the derivation in candidate details, and leaves
unknown or composite model names unpriced instead of guessing.
Aggregate estimates cannot recover historical cache reads and can materially understate cache-heavy agent runs. They are intended to preserve a useful historical comparison until new runs provide full token categories; they are not backfilled into the database.
For OpenRouter runs, the ScoreBench skill starts each configured coding harness
through a loopback-only, per-worker accounting proxy. OpenRouter includes native
token counts, cache details, and billed cost in every completed response, so the
worker submits that cumulative cost with usage_source=openrouter. Those values
are provider measurements and are not prefixed with ~. Detection is based on
the configured endpoint or Codex provider, not merely on an API key being present.
Explicit Pi/OpenCode OpenRouter models are also detected and routed using each
CLI's native provider configuration; see
Pi and OpenCode.
The capture must begin before the coding harness starts and parallel workers
must not share a usage log. The wrapper creates a pinned zero baseline,
publishes changed usage every 30 seconds, and reconciles again at exit.
The proxy also keeps a durable, private request journal, saving generation IDs
as soon as they arrive. Missing receipts can be reconciled using provider
generation metadata, with original errors retained and each generation counted
once. For stopped workers, scorebench run reconcile --check previews this repair
and scorebench run reconcile applies it without inference or lifecycle changes.
See receipt recovery.
No-ID requests remain unknown; an API key's shared spend cannot allocate those
charges to an individual run.
OpenRouter Infrastructure Failures¶
Scientific experiments use openrouter-infrastructure-v1 with updated wrappers.
Their exact experiment usage excludes requests that receive an explicit upstream
408, 429, 500, 502, 503, or 504 failure without model output. Canonical provider
error types take precedence over lossy status codes, so an authentication or
token-limit error is not mistaken for infrastructure overhead. This also covers
explicit gateway error events in an HTTP 200 response or stream before output.
These requests contribute neither cost nor tokens to the experiment metric;
this does not assert that OpenRouter charged nothing.
Every exclusion is appended to the original ledger with the request identity,
status/error code, response digest, and any provider usage received. The wrapper
reports accounting_basis: experiment, accounting_policy, excluded request
count, known excluded provider cost, and the count whose provider cost is unknown.
Experiment accounting remains exact, not approximate. The provider's invoice
may differ. Standalone/prepaid allocations do not adopt this exclusion policy.
Rate limits and retryable gateway responses get at most three proxy attempts,
with backoff and jitter or Retry-After. A wait longer than 60 seconds is passed
back to the native client with its retry header unchanged. Requests with model
output are not replayed by this mechanism. Repeated failures can still interrupt
the worker after the retry limit; they do not falsely exhaust its experiment budget.
Authentication/credit/permission failures and invalid requests need correction, not repeated inference. Output-length stops, reasoning-only output, and normal empty completions remain measured model work. Ambiguous TLS disconnects and unidentified missing receipts are not retroactively called gateway failures. Known generation IDs still use receipt reconciliation. Historical errors that lack response evidence remain unresolved until supported reconciliation or an explicit partial-recovery decision; deployment does not rewrite old totals.
See OpenRouter's error reference for transport-level and in-stream error formats.
Working tokens exclude cache reads, while OpenRouter's full billed usage.cost
includes the costs it charged for the request. Cost is summed across every
included response, including helper requests routed through the proxy. Reasoning tokens
are part of output tokens, not added a second time. Responses API cached input
is normalized just like Chat Completions cached input.
The last completed response is the measurement boundary. During a long stream, spend may be temporarily unchanged until final usage arrives. An in-flight request can overshoot a budget; this is not a prepaid hard spending cap. Missing usage outside the explicit infrastructure exclusions is an accounting error, not free inference. If an included response omits cost, the helper must not present a partial cost sum as the run total. Repeated publication failures stop the wrapper instead of allowing indefinite unmetered work. Preserve the ledger for diagnosis; never reset the baseline to clear an error.
Runner Contract¶
For the best cost accuracy, report cumulative run-relative categories on every
submission and in the final scorebench run usage event:
scorebench submit candidate.py \
--input-tokens 120000 \
--output-tokens 6000 \
--cache-creation-tokens 4000 \
--cache-read-tokens 900000 \
--total-tokens 130000 \
--tokens-total-source launcher_usage
For canonical disjoint counters, --input-tokens means fresh input and
--cache-read-tokens means cached input. OpenAI/Codex instead reports
input_tokens as an inclusive total. Pass that raw total as --input-tokens,
its cached subset as --cached-input-tokens, and any cache-write subset as
--cache-creation-tokens; the CLI subtracts both subsets before sending the
canonical breakdown. Do not pass both cache-read flags, pre-normalize an
inclusive count, or fabricate a split.
New native Grok workers pin ScoreBench's price manifest before inference and price each deduplicated request separately. At 200,000 prompt tokens or more, all token rates for that request double, including cached input and output. Live and final usage snapshots carry this calculated cost; working-token counts do not change. These are global public API-equivalent costs, not subscription charges or regional-endpoint invoices. Older cumulative-only snapshots cannot reconstruct the request tiers, and deployment does not rewrite retained runs. Cost-budget workers refuse to launch when model pricing is unavailable.
Grok's native inputTokens and totalTokens both include cached reads. Grok
runs must send disjoint fresh input, output, and cache-read counters; the server
rejects an aggregate-only Grok snapshot because it cannot recover working
tokens from that value. The ScoreBench skill's token_usage.py --grok-jsonl
path parses the active Grok session safely. For other providers, when only the
aggregate is available, continue sending the exact aggregate and let the
chart mark its cost estimate.
Maintenance¶
Prices live in one provider file under challenge_harness/pricing/. To update a
price, edit that provider's YAML and advance its reviewed_at date. To add a
model, add it under models; no Python registry change is required. The
directory's README.md documents the schema and cache-write convention. Run
python3 scripts/validate_model_pricing.py for a fast local validation before
tests or deployment.
ScoreBench validates every provider file at process startup. The report payload includes the catalog version, per-provider review dates, config filenames, and official source URLs. The documentation table above is generated from the same catalog during the MkDocs build, so it cannot drift from server calculations. Report regeneration uses the currently deployed catalog, making displayed figures a current-price comparison rather than historical provider billing at the candidate's submission date.