Active-Time Accounting¶
This is the technical timing reference. Supported Docker workers set up their observer inside the worker environment; users do not install or run a second host observer for that flow. Start with the launch runbook for normal monitoring and completion checks.
ScoreBench timing v2 is a low-overhead, evidence-based shadow measurement. It does not force every wall-clock interval into active or idle. Unsupported time stays unknown, and the API reports lower and upper bounds.
Shadow mode
progress.active_seconds and the existing report fields remain
authoritative. progress.timing_v2.authoritative is false. Do not stop a
run, create completion markers, backfill history, or change promotion logic
from v2 until the shadow comparison has been reviewed and an explicit
migration is approved.
Four States¶
| State | Meaning | Counts toward target |
|---|---|---|
confirmed_active |
Structured model, tool, or agent-operation span | Yes |
verified_blocked |
Bounded dependency wait attested by ScoreBench | No |
explicit_pause |
Explicit pause or terminal finish interval | No |
unknown |
No supported classification | Neither credited nor deducted |
An observer may submit block_claim, but a claim is not a fifth accounting
state. It remains unknown until independent service evidence verifies it.
The public bounds are:
active_lower_bound = credited_active
active_upper_bound = credited_active + unknown
Target decisions are conservative:
- lower bound at or above target:
definitely_complete - upper bound below target:
definitely_incomplete - target between the bounds:
ambiguous
For every projection:
credited_active + verified_blocked + explicit_pause + unknown = wall
Evidence Precedence¶
The classifier is deterministic:
- Explicit pause or finish wins over every overlapping span.
- Direct structured activity wins over a service blockage attestation. This avoids deducting time when useful work continued during an outage.
- Verified blockage wins over inferred transition allowance.
- Everything else remains unknown.
A published five-second transition allowance may bridge consecutive direct
activity spans. It is reported separately and never folded into
confirmed_active_seconds.
Passive Local Observer¶
The ScoreBench skill includes scripts/scorebench_observer.py. Register it once
after the run start or resume ping:
SCOREBENCH_SKILL_DIR="${CODEX_HOME:-$HOME/.codex}/skills/scorebench"
test -f "$SCOREBENCH_SKILL_DIR/scripts/scorebench_observer.py" || \
SCOREBENCH_SKILL_DIR="$HOME/.claude/skills/scorebench"
python3 "$SCOREBENCH_SKILL_DIR/scripts/scorebench_observer.py" \
register --provider auto --cwd "$PWD"
Pairing saves the observer credential in the owner-only CLI config. The web
handoff and admin launcher expose it once as
SCOREBENCH_TIMING_OBSERVER_TOKEN. The regular hrun_... token cannot write
timing evidence, and the hobs_... token cannot submit candidates or certify a
blocked or paused interval. Explicit pauses come from the run lifecycle.
Registration starts one singleton host daemon. It uses filesystem events on Linux with a bounded five-second fallback, parses only newly appended structured session records, batches completed spans into at most 60-second leases, and uploads no more than once per minute. It does not prompt the model or consume model tokens.
The parser uses only event timestamps, event categories, opaque operation IDs, and process liveness. Prompt text, messages, reasoning, source code, tool arguments, tool output, and secrets are neither persisted in observer state nor uploaded. The server accepts only an allowlist of short scalar metadata. Generated public report payloads contain the aggregate timing projection, not the underlying observer rows or host identifiers.
If the observer cannot prove a span, the interval remains unknown. It does not fall back to pane text, candidate cadence, or a longer idle timeout.
Observer Endpoint¶
POST /run/timing-evidence requires the dedicated hobs_... bearer token:
{
"observer_id": "host-0123456789abcdef",
"observer_boot_id": "boot-0123456789abcdef",
"evidence": [
{
"id": "obs-0123456789abcdef01234567",
"started_at": "2026-08-28T10:00:00Z",
"ended_at": "2026-08-28T10:01:00Z",
"classification": "confirmed_active",
"source": "claude_session",
"sequence": 1,
"metadata": {
"provider": "claude",
"operation_kind": "model_or_tool",
"event_count": 12,
"parser_version": "timing-observer.1",
"lease_seconds": 60
}
}
]
}
Writes are append-only and idempotent. Reusing an evidence ID with different content is rejected. A batch has at most 256 spans; observer spans are at most five minutes, while the bundled observer emits one-minute leases.
Verified Blockage¶
Only ScoreBench service code can store verified_blocked. Automatic connector
attestation currently requires all of these:
- The exact run was synchronously waiting inside a bounded connector call.
- The connector returned
failedwith structuredretryable: trueevidence. - The failure kind is an allowlisted upstream transport or HTTP 5xx failure.
- A circuit breaker confirms at least two consecutive failures and has reached its configured threshold.
- No overlapping direct activity takes precedence in the projection.
A first transient failure, a slow request, a queued result, an unstructured exception, a global outage without lane correlation, and every worker claim stay unknown. Stored blockage spans have bounded start and end timestamps plus the connector and request ID.
Progress And Reports¶
GET /run/progress keeps schema version 1 and adds a nested timing_v2 object:
{
"schema_version": 2,
"timing_policy_version": "timing-v2-shadow.1",
"authoritative": false,
"wall_seconds": 14400,
"confirmed_active_seconds": 12600,
"transition_allowance_seconds": 30,
"credited_active_seconds": 12630,
"verified_blocked_seconds": 300,
"explicit_pause_seconds": 600,
"unknown_seconds": 870,
"active_lower_bound_seconds": 12630,
"active_upper_bound_seconds": 13500,
"evidence_coverage_ratio": 0.9396,
"target_active_seconds": 14400,
"target_state": "definitely_incomplete",
"measured_at": "2026-08-28T14:00:00Z"
}
A progress read advances wall time only as unknown; polling never creates active evidence. A terminal finish closes the summary endpoint, and a later resume reopens it. Candidate timing is projected only through that candidate's submission timestamp. Run summaries stop at the terminal finish timestamp, so post-finish usage or telemetry cannot inflate the result. Generated comparison reports add v2 only after that run has passive evidence; historical runs are not implicitly backfilled with all-unknown projections.
Migration Gates¶
Before v2 can become authoritative:
- Replay representative Codex and Claude runs, including long model calls, long tools, restarts, pauses, outages, and deadlocks.
- Inspect every material unknown interval and compare v1 with both v2 bounds.
- Keep parser CPU, memory, request rate, and payload size within the published staging budget.
- Run the deployment's isolated staging canary without backfilling existing runs. The canary exercises the live evidence and progress endpoints, then removes its temporary arm, tokens, lifecycle rows, and evidence.
- Approve the policy version and watcher completion behavior explicitly.
Until then, v1 remains the run-control clock and v2 is diagnostic evidence only.