bench:model-cost

Lower cost per accepted proof-backed PR at fixed quality.

At a fixed quality bar, Boetica routing delivers a lower cost per accepted proof-backed PR or remediation than provider-native routing.

Lower cost per accepted proof-backed PR at fixed quality.ScorecardFresh · Jun 2026Signed scorecard: Collected recently and within validity window, dated Jun 2026.

Baselines
Provider-native routing · OpenRouter-style routing
Collected
2026-06-22
Expires
2026-09-22
Dataset hash
sha256:0f0f0f0f
Boetica commit
b7e4c19
App version
engine 2026.6
Sandbox fabric
n/a (routing harness)
Signature state
Signed scorecard — Fresh: Collected recently and within validity window.

Signed results

Every row reports Boetica against the strongest baseline, names the winner without relying on color, and ties the result to an inspectable artifact.

bench:model-cost · collected 2026-06-22 · expires 2026-09-22
MetricBoeticaBest baselineWinnerArtifact
Cost per accepted PRLowerBaselineBoeticaCost report
Quality bar heldYesYesTieEval report
Fallback coveragePresentPartialBoeticaRouting log

Representative end-state figures. Replaced by live signed scorecard data before procurement review.

Re-verify due2026-06-22. Inside the validity window but scheduled for re-verification.
Signature present and re-verifiable
Algorithm
ECDSA P-256 (cosign keyless)
KMS key
gcpkms://projects/boetica-prod/locations/global/keyRings/evidence/cryptoKeys/scorecards
Sigstore bundle
sigstore-bundle://rekor/boetica/scorecards
Signer
boetica-evidence-signer
Signed at
2026-06-24T08:00:00Z
Digest
sha256:a8a29e01
  1. Routing datasetsha256:0f0f0f0f
  2. Cost reportsha256:1e1e1e1e
  3. Signaturesha256:a8a29e01
Open evidence packetFresh · Jun 2026Evidence packet: Collected recently and within validity window, dated Jun 2026.

Plain-text summary: across 3 measured metrics, Boetica leads its baselines on the bench:model-cost benchmark, signed ECDSA P-256 (cosign keyless) on 2026-06-24T08:00:00Z and re-verifiable from the hash trail above.

bench:model-cost · methodology

How this benchmark is run

At a fixed quality bar, routing is scored on cost per accepted proof-backed PR or remediation, quality-bar adherence, and fallback coverage.

Fixtures
A routing dataset of accepted tasks replayed across routing strategies at a held-constant quality gate.
Baseline collection
Provider-native and OpenRouter-style routing are replayed on the same tasks; cost is measured from provider billing units, not list price.
Statistical method
Cost per accepted PR is total attributed model cost divided by accepted tasks at the fixed quality bar.
Reviewer
Independent routing reviewer
Last updated
2026-06-22

Inclusion rules

  • Only tasks that pass the fixed quality gate are counted.
  • Cost includes all model usage attributed to the accepted task.
  • Fallback events are logged and attributed to the strategy.

Exclusion rules

  • Tasks that fail the quality gate under any strategy (excluded from cost comparison).

Limitations

  • This domain is re-verify-due; the verifier re-run is scheduled within the validity window.

Other signed domains

Each domain runs through the same trust boundary and leaves its own signed scorecard.

Run it on your own work

Prove bench:model-cost on your repo, not ours.

Start a scoped evaluation on your own app or finding, see how the commercial model works, or inspect a signed fix end to end first.