Grading the whole chip program
Hardware AI is graded one task at a time, and chip programs run for years. Here’s how we’re going to grade an agent across a whole program, the six levels of autonomy we’ll grade it against, and what we’ll publish soon.
Ask a frontier model to write a Verilog module from a description and it will usually get it right. The benchmarks that made that question famous are close to saturated. The top models pass nearly every task.
That was never the hard part of building a chip. A verification team works on a program that runs for years, across thousands of revisions, dozens of blocks and a regression that never really stops. The question its lead asks every morning is harder to grade. What changed overnight, what broke, and which of last week’s results can we still sign off on?
No public benchmark asks that. So we’re building the ones that do. We call them Tapeout Evals, and the first results are coming soon.
Six levels of autonomy
Self‑driving cars have levels of autonomy. Hardware agents don’t yet, so the field talks past itself. A tool that finishes one task and a system that runs a chip program both get called an agent. We think the field needs a shared scale, and this is the one we’ll grade against.
- L5Autonomous chipTakes a whole chip from spec to tape‑out.Not yet measured
- L4Plan to closurePlans a subsystem and drives it to signoff, with engineers signing off.Tapeout Evals
- L3Program stateCarries evidence and context from one revision of a chip to the next.
- L2Tool loopRuns the tools itself and fixes its own mistakes.Most benchmarks today
- L1TaskCompletes one task on request, like a module or a test.
- L0AutocompleteSuggests code while an engineer types.
Nearly every hardware benchmark published so far grades levels one and two, one task or one tool loop at a time. Tapeout Evals start at level three, where an agent has to carry what it knows from one revision of a chip to the next. Nobody grades level five yet. We intend to be the first.
What we’ll measure
Carry asks the question every verification lead asks after a change. Which verdicts still hold? We replay the full commit history of open SoCs, let the agent decide which verdicts to keep, and check every decision against a full rerun. The number that matters is how often it keeps a verdict it shouldn’t. A system that gets this wrong ships bugs, however fast it runs.
Vacuous grades the tests that pass. Some pass because they never checked anything. An assertion that can’t fire, a comparison against the wrong signal, coverage that counts a line without looking at the result. We plant them in open designs and measure how many an agent catches. A higher coverage number means very little if the tests behind it never check anything.
Silent plants bugs that compile, pass lint and pass every existing test, like a typo in a signal name that quietly creates a new wire. These are the bugs that reach silicon.
Plan hands the agent a public reference manual and asks for a verification plan for the whole chip, across all 23 areas an SoC team has to cover, from clocks and resets to secure boot. Expert engineers write plans for the same chips, and we grade one against the other.
Triage takes failures from open chip projects, hides the fix and asks for the root cause. Distill cuts hundreds of gigabytes of waveform down to the handful of events that explain a failure. Noise sorts lint, CDC and RDC reports into the few warnings that matter.
Nightly replays a month of a chip program, night by night, with an agent on the overnight shift. Onboard times how long it takes to become useful on a flow it has never seen.
Secure plants the bugs an attacker would look for, from debug paths left unlocked to data that leaks across privilege levels. Guard turns the attack on the agent itself and measures whether design data stays where it belongs.
And Closure takes a chip of our own from spec to tape‑out, then grades the verification against the silicon that comes back.
How we’ll run them
The same model runs inside Tapeout and inside general‑purpose coding agents, so the difference we report is the harness. Expert engineers set the human baseline on the same tasks, with the same time and budget. Cost and time sit next to every score.
Public numbers run on open tools, so anyone can rerun them. Every trial gets published, with every exclusion and the reason for it. Tasks nearly everyone passes get replaced, and held‑out variants keep training data out of the results.
Coming soon
The first results are being graded now. The suite, the levels and the rules are on our benchmarks page, and results will land there first.
If you lead a verification team and want to see how your flow would score, we’d like to talk.