bench:model-cost
Lower cost per accepted proof-backed PR at fixed quality.
At a fixed quality bar, Boetica routing delivers a lower cost per accepted proof-backed PR or remediation than provider-native routing.
Lower cost per accepted proof-backed PR at fixed quality.ScorecardFresh · Jun 2026Signed scorecard: Collected recently and within validity window, dated Jun 2026.
Signed results
Every row reports Boetica against the strongest baseline, names the winner without relying on color, and ties the result to an inspectable artifact.
| Metric | Boetica | Best baseline | Winner | Artifact |
|---|---|---|---|---|
| Cost per accepted PR | Lower | Baseline | Boetica | Cost report |
| Quality bar held | Yes | Yes | Tie | Eval report |
| Fallback coverage | Present | Partial | Boetica | Routing log |
Representative end-state figures. Replaced by live signed scorecard data before procurement review.
- Algorithm
- ECDSA P-256 (cosign keyless)
- KMS key
- gcpkms://projects/boetica-prod/locations/global/keyRings/evidence/cryptoKeys/scorecards
- Sigstore bundle
- sigstore-bundle://rekor/boetica/scorecards
- Signer
- boetica-evidence-signer
- Signed at
- 2026-06-24T08:00:00Z
- Digest
- sha256:a8a29e01
- Routing dataset
sha256:0f0f0f0f - Cost report
sha256:1e1e1e1e - Signature
sha256:a8a29e01
Plain-text summary: across 3 measured metrics, Boetica leads its baselines on the bench:model-cost benchmark, signed ECDSA P-256 (cosign keyless) on 2026-06-24T08:00:00Z and re-verifiable from the hash trail above.
bench:model-cost · methodology
How this benchmark is run
At a fixed quality bar, routing is scored on cost per accepted proof-backed PR or remediation, quality-bar adherence, and fallback coverage.
- Fixtures
- A routing dataset of accepted tasks replayed across routing strategies at a held-constant quality gate.
- Baseline collection
- Provider-native and OpenRouter-style routing are replayed on the same tasks; cost is measured from provider billing units, not list price.
- Statistical method
- Cost per accepted PR is total attributed model cost divided by accepted tasks at the fixed quality bar.
- Reviewer
- Independent routing reviewer
- Last updated
- 2026-06-22
Inclusion rules
- Only tasks that pass the fixed quality gate are counted.
- Cost includes all model usage attributed to the accepted task.
- Fallback events are logged and attributed to the strategy.
Exclusion rules
- Tasks that fail the quality gate under any strategy (excluded from cost comparison).
Limitations
- This domain is re-verify-due; the verifier re-run is scheduled within the validity window.
Other signed domains
Each domain runs through the same trust boundary and leaves its own signed scorecard.
- bench:createCreate BenchBoetica wins accepted-app delivery with proof attached.
- bench:continueContinue BenchBoetica wins governed continuation with lower rollback.
- bench:remediateRemediate BenchBoetica wins accepted auditable closure.
- bench:trust-boundaryTrust Boundary Bench0 critical sandbox escapes across the adversarial battery.
- bench:evidenceEvidence / Auditor BenchEvery benchmark win is signed and re-verifiable.
- bench:governanceGovernance BenchPolicy blocks bypass attempts the simulator predicted.
- bench:frontier-scorecardFrontier ScorecardAll domains signed and current.
Run it on your own work
Prove bench:model-cost on your repo, not ours.
Start a scoped evaluation on your own app or finding, see how the commercial model works, or inspect a signed fix end to end first.
