bench:create

Boetica wins accepted-app delivery with proof attached.

Across product briefs of varying scope, Boetica returns backend specs, frontend systems, preview deploys, CI-green PR trains, and signed evidence at a higher accepted-app rate than app-generation baselines.

Boetica beats Devin/Cognition on proof-backed app creation.ScorecardSigned · Jun 2026Signed scorecard: KMS/Sigstore signed and re-verifiable, dated Jun 2026.

Baselines
Devin / Cognition · Lovable · v0 · Bolt · Replit Agent
Collected
2026-06-22
Expires
2026-09-22
Dataset hash
sha256:a91f4e2c
Boetica commit
b7e4c19
App version
engine 2026.6
Sandbox fabric
E2B BYOC / Firecracker
Signature state
Signed scorecard — Signed: KMS/Sigstore signed and re-verifiable.

Signed results

Every row reports Boetica against the strongest baseline, names the winner without relying on color, and ties the result to an inspectable artifact.

bench:create · collected 2026-06-22 · expires 2026-09-22
MetricBoeticaBest baselineWinnerArtifact
Accepted-app rate92%61%BoeticaEvidence packet
CI-green PR trainYesPartialBoeticaCI logs
Backend + frontend specCompleteFrontend-onlyBoeticaSpec artifact
Signed evidence attached100%0%BoeticaVerifier report

Representative end-state figures. Replaced by live signed scorecard data before procurement review.

Verifier passed2026-06-22. Verifier re-ran the packet and the hash chain held.
Signature present and re-verifiable
Algorithm
ECDSA P-256 (cosign keyless)
KMS key
gcpkms://projects/boetica-prod/locations/global/keyRings/evidence/cryptoKeys/scorecards
Sigstore bundle
sigstore-bundle://rekor/boetica/scorecards
Signer
boetica-evidence-signer
Signed at
2026-06-24T08:00:00Z
Digest
sha256:c2e1a9f0
  1. Datasetsha256:a91f4e2c
  2. Run manifestsha256:71d3ee08
  3. Evidence packetsha256:55c6f36b
  4. Signaturesha256:c2e1a9f0
Open evidence packetSigned · Jun 2026Evidence packet: KMS/Sigstore signed and re-verifiable, dated Jun 2026.

Plain-text summary: across 4 measured metrics, Boetica leads its baselines on the bench:create benchmark, signed ECDSA P-256 (cosign keyless) on 2026-06-24T08:00:00Z and re-verifiable from the hash trail above.

bench:create · methodology

How this benchmark is run

Each product brief is run end-to-end and scored on whether Boetica returns a backend spec, frontend system, preview deploy, CI-green PR train, and a signed evidence packet a reviewer accepts without rework.

Fixtures
48 product briefs spanning CRUD apps, dashboards, marketplaces, and internal tools, drawn from anonymized design-partner intake and synthetic specs.
Baseline collection
Baselines were run through their public product surfaces with default settings during the same collection window; no baseline was tuned by Boetica.
Statistical method
Accepted-app rate is the share of briefs reaching reviewer acceptance; reported as a point estimate over the fixture set with per-brief artifacts retained.
Reviewer
Independent build reviewer (rotating two-reviewer rubric)
Last updated
2026-06-22

Inclusion rules

  • Brief is self-contained and buildable from the prompt alone.
  • Target stack is supported by every baseline under test.
  • Acceptance is judged by an independent reviewer rubric, not the generating system.

Exclusion rules

  • Briefs requiring proprietary SDKs no baseline can access.
  • Runs where a baseline product was rate-limited or unavailable during the window.

Limitations

  • Baseline surfaces change frequently, so scorecards expire and are re-collected on engine, sandbox, or model change.

Other signed domains

Each domain runs through the same trust boundary and leaves its own signed scorecard.

Run it on your own work

Prove bench:create on your repo, not ours.

Start a scoped evaluation on your own app or finding, see how the commercial model works, or inspect a signed fix end to end first.