Paradigm Puzzlesexperiments by ScoreBench

API Cost Accounting

ScoreBench shows an API-equivalent model cost for each priced run. This is the cost implied by token usage and the standard synchronous public API price catalog loaded by the server. It is useful for comparing strategies under a common budget, but it is not an invoice and does not attempt to value a Claude Max, ChatGPT, enterprise, or other subscription plan.

The token and cost views deliberately count cache reads differently. Working tokens exclude cached-input reads so repeatedly loading the same context does not look like new work. API-equivalent cost includes those reads at the model's cached-input price. Cache writes remain part of working tokens and use their separate price where the provider publishes one.

The chart view and Export Studio support Cost, Working tokens, Fresh input, Output tokens, and Processed tokens as chart axes. Lower-cost candidate snapshots appear earlier on the cost axis, so the score chart answers "what quality did this run reach for this API budget?" Processed tokens add known cache reads to working tokens; a run without a cache-read counter remains visible as a labeled lower bound instead of disappearing. Export Studio's Runs view can compare one best-scoring point per run by any of these resource measures.

Price Table

A dash means the provider does not publish a separately billed cache-write category for that model; ScoreBench charges cache misses at the regular input price rather than treating them as free.

The current catalog version is 2026-09-22. Prices are USD per million tokens.

Provider Reviewed ScoreBench model Input Cache hit/read Cache write (default rate) Output
OpenAI 2026-09-22 gpt-6-astra $10.00 $1.00 $12.50 $50.00
OpenAI 2026-09-22 gpt-6-sol $2.00 $0.20 $2.50 $10.00
OpenAI 2026-09-22 gpt-6-luna $0.10 $0.01 $0.125 $0.50
OpenAI 2026-09-22 gpt-5.6-sol $4.00 $0.40 $5.00 $20.00
OpenAI 2026-09-22 gpt-5.6-terra $2.00 $0.20 $2.50 $12.00
OpenAI 2026-09-22 gpt-5.6-luna $0.20 $0.02 $0.25 $1.20
OpenAI 2026-09-22 gpt-5.5 $5.00 $0.50 - $30.00
OpenAI 2026-09-22 gpt-5-codex $1.25 $0.125 - $10.00
Anthropic 2026-09-22 claude-fable-5-1 $10.00 $0.25 $20.00 $50.00
Anthropic 2026-09-22 claude-fable-5 $10.00 $1.00 $20.00 $50.00
Anthropic 2026-09-22 claude-opus-5-5 $4.00 $0.20 $8.00 $20.00
Anthropic 2026-09-22 claude-opus-5 $5.00 $0.50 $10.00 $25.00
Anthropic 2026-09-22 claude-sonnet-5 $2.00 $0.20 $4.00 $10.00
Anthropic 2026-09-22 claude-haiku-4.5 $1.00 $0.10 $2.00 $5.00
Anthropic 2026-09-22 claude-opus-4.8 $5.00 $0.50 $10.00 $25.00
xAI 2026-09-21 grok-4.7 $2.00 $0.50 - $6.00
xAI 2026-09-21 grok-4.6 $2.00 $0.50 - $6.00
xAI 2026-09-21 grok-4.5 $2.00 $0.30 - $6.00
Google 2026-09-01 gemini-3.5-flash $1.50 $0.15 - $9.00
Google 2026-09-01 gemini-3.1-pro-preview $2.00 $0.20 - $12.00
Moonshot AI 2026-08-28 kimi-k3 $3.00 $0.30 - $15.00
Alibaba Cloud 2026-09-01 qwen3-coder-plus $1.00 $0.20 $1.25 $5.00
DeepSeek 2026-09-01 deepseek-v4-pro $1.32 $0.044 - $3.96
Z.AI 2026-08-28 glm-5.2 $1.40 $0.26 - $4.40

Provider-specific notes:

  • OpenAI: GPT-5.6 Sol uses the current promotional rate, published as available at least through November 21, 2026.
  • OpenAI: GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6, and GPT-5.5 requests above their published long-context thresholds have per-request price modifiers that ScoreBench does not infer from cumulative counters.
  • OpenAI: GPT-6 Sol and GPT-6 Luna (released 2026-09-22) bill prompts above 272K input tokens at 2x input and cache rates and 1.5x output.
  • Anthropic: Anthropic publishes separate 5-minute and 1-hour cache-write rates. ScoreBench's generic cache-creation counter uses the conservative 1-hour rate so native Claude Code runs cannot be undercounted; provider-metered cost_usd overrides this fallback when available.
  • xAI: Native Grok workers price each recorded inference separately, doubling all token rates when its prompt reaches 200k tokens. Older cumulative-only records cannot identify this tier. These are global public API-equivalent prices, not subscription charges or regional-endpoint invoices.
  • Google: Gemini 3.1 Pro Preview charges higher rates for prompts above 200k tokens. ScoreBench uses the <=200k standard rate because cumulative usage counters do not preserve each request's context tier.
  • Google: Gemini context-cache storage is time-based and is not represented by ScoreBench's token-only cache-creation field; provider-metered cost_usd remains authoritative when available.
  • Alibaba Cloud: Qwen3-Coder-Plus uses context-length tiers. ScoreBench uses the international <=32k standard rate because cumulative usage counters do not preserve each request's context tier.
  • Alibaba Cloud: Alibaba publishes both implicit-cache and lower explicit-cache-read prices. ScoreBench's generic cache-read counter uses the conservative implicit-cache rate unless provider-metered cost_usd is available.
  • DeepSeek: DeepSeek publishes peak and off-peak rates. ScoreBench uses the peak rate so a cost budget is not understated; provider-metered cost_usd remains authoritative and off-peak API usage can cost half as much.
  • Z.AI: Z.AI currently lists cached-input storage for GLM-5.2 as free; ScoreBench models token charges only.

Official references: OpenAI model catalog, OpenAI prompt-caching billing, Anthropic pricing, xAI pricing, Gemini Developer API pricing, Kimi K3 pricing, Qwen3-Coder-Plus pricing, DeepSeek models and pricing, Z.AI pricing.

The standard table deliberately excludes long-context surcharges, regional uplifts, batch discounts, priority or fast-mode charges, tools, infrastructure, taxes, and subscription economics. Provider-specific modifiers cannot be reconstructed reliably from cumulative run counters because their thresholds and modes apply per request.

Derivation Order

ScoreBench chooses the strongest available measurement for every candidate:

  1. A provider-reported cost_usd, captured by the runner when present.
  2. A cumulative input/output/cache token breakdown priced with the table above.
  3. A final run-level breakdown allocated to candidates in proportion to each candidate's cumulative working-token total.
  4. A legacy aggregate estimate using 80% input and 20% output tokens:
cost = working_tokens * (0.8 * input_rate + 0.2 * output_rate) / 1,000,000

Methods 3 and 4 are estimates. A partial token breakdown is also an estimate when a separately priced cache category is missing. The UI prefixes every estimated value with ~, shows the derivation in candidate details, and leaves unknown or composite model names unpriced instead of guessing.

Aggregate estimates cannot recover historical cache reads and can materially understate cache-heavy agent runs. They are intended to preserve a useful historical comparison until new runs provide full token categories; they are not backfilled into the database.

For OpenRouter runs, the ScoreBench skill starts each configured coding harness through a loopback-only, per-worker accounting proxy. OpenRouter includes native token counts, cache details, and billed cost in every completed response, so the worker submits that cumulative cost with usage_source=openrouter. Those values are provider measurements and are not prefixed with ~. Detection is based on the configured endpoint or Codex provider, not merely on an API key being present. Explicit Pi/OpenCode OpenRouter models are also detected and routed using each CLI's native provider configuration; see Pi and OpenCode. The capture must begin before the coding harness starts and parallel workers must not share a usage log. The wrapper creates a pinned zero baseline, publishes changed usage every 30 seconds, and reconciles again at exit.

The proxy also keeps a durable, private request journal, saving generation IDs as soon as they arrive. Missing receipts can be reconciled using provider generation metadata, with original errors retained and each generation counted once. For stopped workers, scorebench run reconcile --check previews this repair and scorebench run reconcile applies it without inference or lifecycle changes. See receipt recovery. No-ID requests remain unknown; an API key's shared spend cannot allocate those charges to an individual run.

OpenRouter Infrastructure Failures

Scientific experiments use openrouter-infrastructure-v1 with updated wrappers. Their exact experiment usage excludes requests that receive an explicit upstream 408, 429, 500, 502, 503, or 504 failure without model output. Canonical provider error types take precedence over lossy status codes, so an authentication or token-limit error is not mistaken for infrastructure overhead. This also covers explicit gateway error events in an HTTP 200 response or stream before output. These requests contribute neither cost nor tokens to the experiment metric; this does not assert that OpenRouter charged nothing.

Every exclusion is appended to the original ledger with the request identity, status/error code, response digest, and any provider usage received. The wrapper reports accounting_basis: experiment, accounting_policy, excluded request count, known excluded provider cost, and the count whose provider cost is unknown. Experiment accounting remains exact, not approximate. The provider's invoice may differ. Standalone/prepaid allocations do not adopt this exclusion policy.

Rate limits and retryable gateway responses get at most three proxy attempts, with backoff and jitter or Retry-After. A wait longer than 60 seconds is passed back to the native client with its retry header unchanged. Requests with model output are not replayed by this mechanism. Repeated failures can still interrupt the worker after the retry limit; they do not falsely exhaust its experiment budget.

Authentication/credit/permission failures and invalid requests need correction, not repeated inference. Output-length stops, reasoning-only output, and normal empty completions remain measured model work. Ambiguous TLS disconnects and unidentified missing receipts are not retroactively called gateway failures. Known generation IDs still use receipt reconciliation. Historical errors that lack response evidence remain unresolved until supported reconciliation or an explicit partial-recovery decision; deployment does not rewrite old totals.

See OpenRouter's error reference for transport-level and in-stream error formats.

Working tokens exclude cache reads, while OpenRouter's full billed usage.cost includes the costs it charged for the request. Cost is summed across every included response, including helper requests routed through the proxy. Reasoning tokens are part of output tokens, not added a second time. Responses API cached input is normalized just like Chat Completions cached input.

The last completed response is the measurement boundary. During a long stream, spend may be temporarily unchanged until final usage arrives. An in-flight request can overshoot a budget; this is not a prepaid hard spending cap. Missing usage outside the explicit infrastructure exclusions is an accounting error, not free inference. If an included response omits cost, the helper must not present a partial cost sum as the run total. Repeated publication failures stop the wrapper instead of allowing indefinite unmetered work. Preserve the ledger for diagnosis; never reset the baseline to clear an error.

Runner Contract

For the best cost accuracy, report cumulative run-relative categories on every submission and in the final scorebench run usage event:

scorebench submit candidate.py \
  --input-tokens 120000 \
  --output-tokens 6000 \
  --cache-creation-tokens 4000 \
  --cache-read-tokens 900000 \
  --total-tokens 130000 \
  --tokens-total-source launcher_usage

For canonical disjoint counters, --input-tokens means fresh input and --cache-read-tokens means cached input. OpenAI/Codex instead reports input_tokens as an inclusive total. Pass that raw total as --input-tokens, its cached subset as --cached-input-tokens, and any cache-write subset as --cache-creation-tokens; the CLI subtracts both subsets before sending the canonical breakdown. Do not pass both cache-read flags, pre-normalize an inclusive count, or fabricate a split.

New native Grok workers pin ScoreBench's price manifest before inference and price each deduplicated request separately. At 200,000 prompt tokens or more, all token rates for that request double, including cached input and output. Live and final usage snapshots carry this calculated cost; working-token counts do not change. These are global public API-equivalent costs, not subscription charges or regional-endpoint invoices. Older cumulative-only snapshots cannot reconstruct the request tiers, and deployment does not rewrite retained runs. Cost-budget workers refuse to launch when model pricing is unavailable.

Grok's native inputTokens and totalTokens both include cached reads. Grok runs must send disjoint fresh input, output, and cache-read counters; the server rejects an aggregate-only Grok snapshot because it cannot recover working tokens from that value. The ScoreBench skill's token_usage.py --grok-jsonl path parses the active Grok session safely. For other providers, when only the aggregate is available, continue sending the exact aggregate and let the chart mark its cost estimate.

Maintenance

Prices live in one provider file under challenge_harness/pricing/. To update a price, edit that provider's YAML and advance its reviewed_at date. To add a model, add it under models; no Python registry change is required. The directory's README.md documents the schema and cache-write convention. Run python3 scripts/validate_model_pricing.py for a fast local validation before tests or deployment.

ScoreBench validates every provider file at process startup. The report payload includes the catalog version, per-provider review dates, config filenames, and official source URLs. The documentation table above is generated from the same catalog during the MkDocs build, so it cannot drift from server calculations. Report regeneration uses the currently deployed catalog, making displayed figures a current-price comparison rather than historical provider billing at the candidate's submission date.