GenAlpha

Introducing GenAlpha AnalogBench

18 September 2026

Analog is silicon's biggest bottleneck. GenAlpha builds AI Designers that accelerate it. Agents that do the design work in a real flow, drawing schematics & layout, running & debugging DRC, LVS, and parasitics.

Our AI Designers can be paired with any of the leading models, so we evaluate those models regularly — which ones do the best analog work, and how much output and time each one takes to get there. Our primary evaluation framework is GenAlpha AnalogBench: a curated set of real analog design and layout tasks, executed by our AI Designers paired with each model. We recently evaluated twelve leading open- and closed-weight models, and want to share the results.

Leaderboard

Higher scores are better.

1claude-fable-5.1796
2claude-opus-5771
3grok-4.6715
4gpt-5.6-sol709
5gemini-3.8-flash701
6muse-spark-1.3681
7gpt-6-astra665
8glm-5.3-flash624
9kimi-k3621
10qwen3.8-max616
11deepseek-v4.1-flash528
12inkling184

By task type

Highest score in each column in teal. Simple edits, such as deleting an instance or moving a net, are counted with layout.

ModelOverallSchematicLayoutQ&A
claude-fable-5.1796794765856
claude-opus-5771798695862
grok-4.6715666737757
gpt-5.6-sol709710666787
gemini-3.8-flash701749637737
muse-spark-1.3681694657702
gpt-6-astra665867430748
glm-5.3-flash624614585717
kimi-k3621524603819
qwen3.8-max616622532753
deepseek-v4.1-flash528374572717
inkling184125160329

Output tokens and time

One point per model. Higher and further left means a better score for less work. Tokens and time include the AI Designer and any sub-agents it starts.

Score vs output tokens

0 200 400 600 800 1000 0k 20k 40k 60k 80k 100k 120k Average output tokens per task Score claude-fable-5.1: 796, 36.5k output tokens/task claude-fable-5.1 claude-opus-5: 771, 74.2k output tokens/task claude-opus-5 grok-4.6: 715, 65.2k output tokens/task gpt-5.6-sol: 709, 28.2k output tokens/task gpt-5.6-sol gemini-3.8-flash: 701, 52.2k output tokens/task muse-spark-1.3: 681, 85.5k output tokens/task gpt-6-astra: 665, 24.7k output tokens/task gpt-6-astra glm-5.3-flash: 624, 75.1k output tokens/task kimi-k3: 621, 62.8k output tokens/task qwen3.8-max: 616, 66.7k output tokens/task deepseek-v4.1-flash: 528, 107.9k output tokens/task deepseek-v4.1-flash inkling: 184, 27.4k output tokens/task inkling

Score vs time

0 200 400 600 800 1000 0m 5m 10m 15m 20m 25m Average time per task (minutes) Score claude-fable-5.1: 796, 7m 47s/task claude-fable-5.1 claude-opus-5: 771, 13m 57s/task claude-opus-5 grok-4.6: 715, 15m 35s/task gpt-5.6-sol: 709, 6m 13s/task gpt-5.6-sol gemini-3.8-flash: 701, 9m 21s/task muse-spark-1.3: 681, 10m 50s/task gpt-6-astra: 665, 6m 20s/task glm-5.3-flash: 624, 17m 50s/task glm-5.3-flash kimi-k3: 621, 22m 01s/task qwen3.8-max: 616, 22m 00s/task deepseek-v4.1-flash: 528, 7m 58s/task inkling: 184, 23m 04s/task inkling

Benchmark makeup

Schematic: 33 tasks Layout: 35 tasks Q&A: 19 tasks 87tasks
  • Schematic3338%
  • Layout3540%
  • Q&A1922%

Learn more →

← All updates


Generation Alpha Transistor · San Francisco, CA