Tapeout Evals

How we measure an autonomous verification system across a whole chip program, from the first revision to signoff. First results are coming soon.

The suite

Twelve questions a verification lead asks.

Each eval grades one of them, on open designs, across a program’s history.

  1. 01Coming soon

    Carry

    After a change, which verdicts still hold?

    Full commit histories of open SoCs, replayed revision by revision. Every verdict an agent keeps is checked against a full rerun.

  2. 02Coming soon

    Plan

    Can it plan a whole chip?

    A verification plan for a full SoC from a public reference manual, across all 23 areas, graded against plans from expert engineers.

  3. 03Coming soon

    Vacuous

    Did the passing test check anything?

    Assertions that can’t fire, checks against the wrong signal and coverage that counts without looking.

  4. 04Coming soon

    Silent

    Can it catch bugs that compile?

    Planted bugs that compile, pass lint and pass every existing test.

  5. 05Coming soon

    Triage

    Can it find the root cause blind?

    Failures from open chip projects with the fix hidden. The agent gets the failure and the history.

  6. 06Coming soon

    Distill

    Can it find the moment that matters?

    Hundreds of gigabytes of waveform, cut down to the few events that explain a failure.

  7. 07Coming soon

    Noise

    Which warnings are bugs?

    Lint, CDC and RDC reports from open designs, sorted into the few that matter.

  8. 08Coming soon

    Nightly

    Can it hold the overnight shift for a month?

    Thirty days of a chip program, commits, regressions and failures, replayed night by night.

  9. 09Coming soon

    Onboard

    How fast does it learn a new team?

    A cold start on an unfamiliar flow, timed to the first useful result.

  10. 10Coming soon

    Secure

    Can it find the bug an attacker would?

    Planted security bugs in open SoCs, from unlocked debug paths to leaks across privilege levels.

  11. 11Coming soon

    Guard

    Does design data stay where it belongs?

    The attack turned on the agent itself, on every deployment tier.

  12. 12Later

    Closure

    Can it take a chip to silicon?

    A chip of our own verified from spec to tape‑out, then graded against the silicon that comes back.

Levels of autonomy

Six levels, and where the grading stops today.

Self‑driving cars have levels of autonomy. Hardware agents need them too, so a tool that finishes one task and a system that runs a chip program stop sharing a name.

  1. L5Autonomous chipTakes a whole chip from spec to tape‑out.Not yet measured
  2. L4Plan to closurePlans a subsystem and drives it to signoff, with engineers signing off.Tapeout Evals
  3. L3Program stateCarries evidence and context from one revision of a chip to the next.
  4. L2Tool loopRuns the tools itself and fixes its own mistakes.Most benchmarks today
  5. L1TaskCompletes one task on request, like a module or a test.
  6. L0AutocompleteSuggests code while an engineer types.
Results

Being graded now.

Every score will sit next to its cost and time, with the run behind it.

EvalTapeoutSame model aloneExpert engineers
CarryComing soonComing soonComing soon
PlanComing soonComing soonComing soon
VacuousComing soonComing soonComing soon
SilentComing soonComing soonComing soon
TriageComing soonComing soonComing soon
How we run them

Rules every result is published under.

  1. 01The same model runs inside Tapeout and inside general‑purpose coding agents, so the difference we report is the harness.
  2. 02Expert engineers set the baseline on the same tasks, with the same time and budget.
  3. 03Cost and time sit next to every score.
  4. 04Public numbers run on open tools, so anyone can rerun them.
  5. 05Every trial is published, with every exclusion and the reason for it.
  6. 06Held‑out tasks and fresh variants keep training data out of the results.
  7. 07Tasks nearly everyone passes are replaced.
Changelog

Published so far.

  • September 2026The suite, the levels of autonomy and the rules published. First results coming soon.

Verify your next chip with Ledger.

The agents run on your machines, against your RTL, testbenches and tools, and carry every result through to signoff.

Talk to us