Tapeout Evals
How we measure an autonomous verification system across a whole chip program, from the first revision to signoff. First results are coming soon.
Twelve questions a verification lead asks.
Each eval grades one of them, on open designs, across a program’s history.
- 01Coming soon
Carry
After a change, which verdicts still hold?
Full commit histories of open SoCs, replayed revision by revision. Every verdict an agent keeps is checked against a full rerun.
- 02Coming soon
Plan
Can it plan a whole chip?
A verification plan for a full SoC from a public reference manual, across all 23 areas, graded against plans from expert engineers.
- 03Coming soon
Vacuous
Did the passing test check anything?
Assertions that can’t fire, checks against the wrong signal and coverage that counts without looking.
- 04Coming soon
Silent
Can it catch bugs that compile?
Planted bugs that compile, pass lint and pass every existing test.
- 05Coming soon
Triage
Can it find the root cause blind?
Failures from open chip projects with the fix hidden. The agent gets the failure and the history.
- 06Coming soon
Distill
Can it find the moment that matters?
Hundreds of gigabytes of waveform, cut down to the few events that explain a failure.
- 07Coming soon
Noise
Which warnings are bugs?
Lint, CDC and RDC reports from open designs, sorted into the few that matter.
- 08Coming soon
Nightly
Can it hold the overnight shift for a month?
Thirty days of a chip program, commits, regressions and failures, replayed night by night.
- 09Coming soon
Onboard
How fast does it learn a new team?
A cold start on an unfamiliar flow, timed to the first useful result.
- 10Coming soon
Secure
Can it find the bug an attacker would?
Planted security bugs in open SoCs, from unlocked debug paths to leaks across privilege levels.
- 11Coming soon
Guard
Does design data stay where it belongs?
The attack turned on the agent itself, on every deployment tier.
- 12Later
Closure
Can it take a chip to silicon?
A chip of our own verified from spec to tape‑out, then graded against the silicon that comes back.
Six levels, and where the grading stops today.
Self‑driving cars have levels of autonomy. Hardware agents need them too, so a tool that finishes one task and a system that runs a chip program stop sharing a name.
- L5Autonomous chipTakes a whole chip from spec to tape‑out.Not yet measured
- L4Plan to closurePlans a subsystem and drives it to signoff, with engineers signing off.Tapeout Evals
- L3Program stateCarries evidence and context from one revision of a chip to the next.
- L2Tool loopRuns the tools itself and fixes its own mistakes.Most benchmarks today
- L1TaskCompletes one task on request, like a module or a test.
- L0AutocompleteSuggests code while an engineer types.
Being graded now.
Every score will sit next to its cost and time, with the run behind it.
| Eval | Tapeout | Same model alone | Expert engineers |
|---|---|---|---|
| Carry | Coming soon | Coming soon | Coming soon |
| Plan | Coming soon | Coming soon | Coming soon |
| Vacuous | Coming soon | Coming soon | Coming soon |
| Silent | Coming soon | Coming soon | Coming soon |
| Triage | Coming soon | Coming soon | Coming soon |
Rules every result is published under.
- 01The same model runs inside Tapeout and inside general‑purpose coding agents, so the difference we report is the harness.
- 02Expert engineers set the baseline on the same tasks, with the same time and budget.
- 03Cost and time sit next to every score.
- 04Public numbers run on open tools, so anyone can rerun them.
- 05Every trial is published, with every exclusion and the reason for it.
- 06Held‑out tasks and fresh variants keep training data out of the results.
- 07Tasks nearly everyone passes are replaced.
Published so far.
- September 2026The suite, the levels of autonomy and the rules published. First results coming soon.
Verify your next chip with Ledger.
The agents run on your machines, against your RTL, testbenches and tools, and carry every result through to signoff.
Talk to us