bench:create
Boetica wins accepted-app delivery with proof attached.
Across product briefs of varying scope, Boetica returns backend specs, frontend systems, preview deploys, CI-green PR trains, and signed evidence at a higher accepted-app rate than app-generation baselines.
Boetica beats Devin/Cognition on proof-backed app creation.ScorecardSigned · Jun 2026Signed scorecard: KMS/Sigstore signed and re-verifiable, dated Jun 2026.
Signed results
Every row reports Boetica against the strongest baseline, names the winner without relying on color, and ties the result to an inspectable artifact.
| Metric | Boetica | Best baseline | Winner | Artifact |
|---|---|---|---|---|
| Accepted-app rate | 92% | 61% | Boetica | Evidence packet |
| CI-green PR train | Yes | Partial | Boetica | CI logs |
| Backend + frontend spec | Complete | Frontend-only | Boetica | Spec artifact |
| Signed evidence attached | 100% | 0% | Boetica | Verifier report |
Representative end-state figures. Replaced by live signed scorecard data before procurement review.
- Algorithm
- ECDSA P-256 (cosign keyless)
- KMS key
- gcpkms://projects/boetica-prod/locations/global/keyRings/evidence/cryptoKeys/scorecards
- Sigstore bundle
- sigstore-bundle://rekor/boetica/scorecards
- Signer
- boetica-evidence-signer
- Signed at
- 2026-06-24T08:00:00Z
- Digest
- sha256:c2e1a9f0
- Dataset
sha256:a91f4e2c - Run manifest
sha256:71d3ee08 - Evidence packet
sha256:55c6f36b - Signature
sha256:c2e1a9f0
Plain-text summary: across 4 measured metrics, Boetica leads its baselines on the bench:create benchmark, signed ECDSA P-256 (cosign keyless) on 2026-06-24T08:00:00Z and re-verifiable from the hash trail above.
bench:create · methodology
How this benchmark is run
Each product brief is run end-to-end and scored on whether Boetica returns a backend spec, frontend system, preview deploy, CI-green PR train, and a signed evidence packet a reviewer accepts without rework.
- Fixtures
- 48 product briefs spanning CRUD apps, dashboards, marketplaces, and internal tools, drawn from anonymized design-partner intake and synthetic specs.
- Baseline collection
- Baselines were run through their public product surfaces with default settings during the same collection window; no baseline was tuned by Boetica.
- Statistical method
- Accepted-app rate is the share of briefs reaching reviewer acceptance; reported as a point estimate over the fixture set with per-brief artifacts retained.
- Reviewer
- Independent build reviewer (rotating two-reviewer rubric)
- Last updated
- 2026-06-22
Inclusion rules
- Brief is self-contained and buildable from the prompt alone.
- Target stack is supported by every baseline under test.
- Acceptance is judged by an independent reviewer rubric, not the generating system.
Exclusion rules
- Briefs requiring proprietary SDKs no baseline can access.
- Runs where a baseline product was rate-limited or unavailable during the window.
Limitations
- Baseline surfaces change frequently, so scorecards expire and are re-collected on engine, sandbox, or model change.
Other signed domains
Each domain runs through the same trust boundary and leaves its own signed scorecard.
- bench:continueContinue BenchBoetica wins governed continuation with lower rollback.
- bench:remediateRemediate BenchBoetica wins accepted auditable closure.
- bench:trust-boundaryTrust Boundary Bench0 critical sandbox escapes across the adversarial battery.
- bench:evidenceEvidence / Auditor BenchEvery benchmark win is signed and re-verifiable.
- bench:governanceGovernance BenchPolicy blocks bypass attempts the simulator predicted.
- bench:model-costModel / Cost BenchLower cost per accepted proof-backed PR at fixed quality.
- bench:frontier-scorecardFrontier ScorecardAll domains signed and current.
Run it on your own work
Prove bench:create on your repo, not ours.
Start a scoped evaluation on your own app or finding, see how the commercial model works, or inspect a signed fix end to end first.
