bench:trust-boundary
0 critical sandbox escapes across the adversarial battery.
Boetica ran the isolation battery across token exfiltration, metadata, RFC1918, public-internet, cross-tenant, snapshot canary, env/file/process/memory, and prompt-injection-to-secret attempts with zero critical escapes.
0 critical sandbox escapes.Trust attestationSigned · Jun 2026Trust-boundary attestation: KMS/Sigstore signed and re-verifiable, dated Jun 2026.
Signed results
Every row reports Boetica against the strongest baseline, names the winner without relying on color, and ties the result to an inspectable artifact.
| Metric | Boetica | Best baseline | Winner | Artifact |
|---|---|---|---|---|
| Critical escapes | 0 | Not attested | Boetica | Trust attestation |
| Token exfiltration | Blocked | Not attested | Boetica | Red-team result |
| Cross-tenant access | Blocked | Not attested | Boetica | Red-team result |
| Prompt-injection-to-secret | Blocked | Not attested | Boetica | Red-team result |
Representative end-state figures. Replaced by live signed scorecard data before procurement review.
- Algorithm
- ECDSA P-256 (cosign keyless)
- KMS key
- gcpkms://projects/boetica-prod/locations/global/keyRings/evidence/cryptoKeys/scorecards
- Sigstore bundle
- sigstore-bundle://rekor/boetica/scorecards
- Signer
- boetica-evidence-signer
- Signed at
- 2026-06-21T08:00:00Z
- Digest
- sha256:f7d77401
- Battery manifest
sha256:1a2b3c4d - Red-team result
sha256:5e6f7a8b - Attestation
sha256:f7d77401
Plain-text summary: across 4 measured metrics, Boetica leads its baselines on the bench:trust-boundary benchmark, signed ECDSA P-256 (cosign keyless) on 2026-06-21T08:00:00Z and re-verifiable from the hash trail above.
bench:trust-boundary · methodology
How this benchmark is run
The sandbox is attacked with an adversarial battery and scored on critical escapes across token exfiltration, metadata, RFC1918, public-internet, cross-tenant, snapshot canary, env/file/process/memory, and prompt-injection-to-secret attempts.
- Fixtures
- An isolation battery of adversarial probes run inside the sandbox fabric, including an external red-team pass against the same image digest.
- Baseline collection
- Published sandbox claims and the internal adversarial fixture suite form the comparison; baseline columns read 'Not attested' where no signed evidence exists.
- Statistical method
- Result is a pass/fail count of critical escapes; the headline claim holds only at zero critical escapes for the attested digest.
- Reviewer
- External red team + internal isolation owner
- Last updated
- 2026-06-20
Inclusion rules
- Every probe targets a documented boundary control.
- Sandbox image digest and egress policy hash are pinned and recorded.
- A critical escape is any probe that reaches a usable secret, another tenant, or the public internet.
Exclusion rules
- Theoretical attacks with no executable probe in the battery.
- Probes against controls outside the attested boundary (tracked separately).
Limitations
- Attestation is bound to a specific image digest and expires on fabric change.
Other signed domains
Each domain runs through the same trust boundary and leaves its own signed scorecard.
- bench:createCreate BenchBoetica wins accepted-app delivery with proof attached.
- bench:continueContinue BenchBoetica wins governed continuation with lower rollback.
- bench:remediateRemediate BenchBoetica wins accepted auditable closure.
- bench:evidenceEvidence / Auditor BenchEvery benchmark win is signed and re-verifiable.
- bench:governanceGovernance BenchPolicy blocks bypass attempts the simulator predicted.
- bench:model-costModel / Cost BenchLower cost per accepted proof-backed PR at fixed quality.
- bench:frontier-scorecardFrontier ScorecardAll domains signed and current.
Run it on your own work
Prove bench:trust-boundary on your repo, not ours.
Start a scoped evaluation on your own app or finding, see how the commercial model works, or inspect a signed fix end to end first.
