GPT-Synopsys and the models we can put to work
OpenAI and Synopsys are developing a model for chip design. What they announced, why we’re considering open models and fine-tuning, and what we might build around a chip as more of its engineering becomes autonomous.

OpenAI and Synopsys announced GPT-Synopsys on September 30. They’re developing a specialized model that can operate Synopsys tools, read their outputs and iterate on chip designs. Early customer engagements are underway. This is an agreement to develop and deliver the model, rather than a generally available release.
Imagine giving an agent a timing report and asking it to get a block to the target frequency. It opens the report, follows the slow path, changes the design and runs the tools again. The next report tells it whether the change helped, whether area grew and where the remaining problem sits. That is the kind of engineering loop the partnership aims to put inside a model’s working repertoire.
A model that knows the tools
Synopsys says OpenAI will license its EDA tools for the model’s development. The intended work includes power, performance and area optimization, timing and verification closure. The tools supply the results; the model has to decide how to act on them. At Tapeout, our agents keep timing, power and area targets in mind as they investigate failures and propose changes.
The distinction matters when a run fails. A log might point to a read returning the wrong value. The useful next move could be to inspect the waveform, reproduce the failure with a smaller test or check an assumption in the testbench. A model that understands the engineering environment has more to do than write a plausible patch. It has to choose an experiment and learn from what comes back.
Synopsys is also building Autopilot and long-horizon engineering agents, with persistent memory and workflows through verification closure. GPT-Synopsys is intended to integrate with that platform and interoperate with customer agent harnesses.
What we can actually run
The announced GPT-Synopsys offering will run on OpenAI-hosted infrastructure and bundle compute, the model and licenses. The announcement does not offer downloadable weights or establish access for Tapeout. We don’t know yet whether we’ll be able to evaluate it in our harness, or which interfaces and tools would be available to us.
That leaves practical questions for a verification team. Can the model run where the design data is allowed to go? Can it use the team’s existing tools and licenses? Can the team save enough of the run to reproduce a failure? Those answers help determine whether a model can join a live chip program.
Learning a team’s flow
We’re considering open models too. We want to understand which models can do useful verification work inside the environments teams already use. Tapeout Evals could help us make those choices, with the runs, failures, cost and time behind each result available for inspection.
Customer-specific model training is another possibility. A team may have years of failure logs, reviewed fixes and conventions for running its tools. Could training on material the customer authorizes help a model recognize a familiar failure or follow that team’s flow? We don’t know yet how that would scale. Each customer’s data, training costs, model updates and evaluation would need to be handled, and a useful result for one team might not transfer to another.
A small trial could help us find out. Compare the original and adapted model on failures held out from training, using the same tools and budget. Does it find the reset bug sooner? Does it choose a smaller reproducer before launching another full regression? Does it notice that its patch made an assertion stop firing? Any improvement would have to survive those checks before we knew whether customer-specific training was worth pursuing.
We’d like Tapeout Evals to help answer those questions for open-weight models and hosted models we can access. Give each one the same failing design, the same tools and a fixed budget. Keep its commands, edits, tool outputs and final result. Then an engineer can inspect where it found the bug, where it got stuck and how much work the answer cost.
Start with a failure it can finish
The whole-program ambition can be tested in smaller pieces. Start with a simulation failure whose fix is hidden. Can the model reproduce it, identify the cause and produce a repair that survives additional checks? Next, give it a passing assertion that never fires. Does it notice? Then introduce a new RTL revision and ask which earlier results need to be rerun.
Those trials could help us decide which models belong in which parts of Tapeout. A model might be useful at sorting lint warnings but struggle to plan verification for a subsystem. Another might diagnose failures accurately but spend too much compute pursuing a fix. We should be able to report those differences without compressing them into one winner.
Our existing Tapeout Evals include Triage, Vacuous and Carry for these kinds of questions, followed by longer trials such as Nightly and Closure. We’d also like to examine what changes when a model gets the program’s history, which experiment it chooses next and when it stops to ask an engineer for help. Those are directions to investigate, rather than results we have today.
A successful debugging trial would establish something useful and bounded. It would not prove that the same model can take a chip to signoff. If it fails, the commands and tool outputs should show which step needs work before we attempt a longer run.
The chip stays on the bench
A direction we’re interested in is letting a team change the agent working on its chip without losing the history of the work. A specialist model, an open-weight model and an engineer could all work from the same revisions and evidence, with their contributions recorded against the design. The model can change while the chip’s history stays with the program.
One page for the chip
That could lead to a fairly simple prototype. Open a page for a public chip such as OpenTitan Earl Grey. See its major blocks, the revision being examined, the runs that passed, the failures still open and the results that haven’t been established. Click Memory and follow an outstanding failure into the assertion, log and RTL behind it. Put that page in front of engineers and find out whether it answers questions they currently need several tools and a meeting to resolve.
A table might be the best first view. Block, owner, revision, open failures, latest run. Click a cell and see where its value came from. An architect and a verification lead could point to the same block while opening different details. We could test whether that shared view helps before attempting to automate the work underneath it.
What did those seventeen lines change
Another possibility is a diff that follows a source change into the chip. An engineer widens an interface. The review shows the blocks that consume it, the assertions referring to its old width and the firmware definition that needs another look. Any effect the system cannot establish stays an open question. The useful test is whether engineers find consequences they would otherwise have missed, and how often the system sends them down the wrong path.
Which run should come next
We could also explore choosing experiments under a budget. Give an agent a failure and an hour of compute. Does it spend the hour rerunning a broad regression, or find a short test that isolates the problem in minutes? Does it account for a formal license that won’t become available until later? An evaluation could measure the bugs found, the unresolved questions, the compute spent and the point where an engineer had to step in.
None of these ideas requires us to build another simulator, source-control system or place-and-route engine. Those tools already contain decades of engineering. We could connect their outputs to the chip’s revisions and use small prototypes and evaluations to learn which connections are useful. The larger ideas would have to earn their way out of those trials.
GPT-Synopsys puts a specialized model alongside the tools that already do chip engineering. We’re interested in what that model can do if we gain access, what open models and fine-tuning might make possible, and what a team needs around those models to keep understanding its chip. There’s plenty we don’t know yet. A public design, a recorded investigation and a result another engineer can check give us somewhere concrete to start.