Razavi-Bench

Tasks50
Rubric0-4
Models26
Judges2

An expert-curated benchmark that tests whether frontier AI
models can reason about analog circuit design.

Razavi-Bench covers 26 frontier models as of August 2026 -- including Claude Fable 5, Claude Opus 5, GPT 5.6, Qwen 3.8 Max Preview, Kimi K3, and Gemini 3.1 Pro -- on 50 expert-written analog design questions, with all scores recalculated using a corrected rubric after a careful audit of Q15.

Direct QA Leaderboard

#ModelScore RangeToken Cost

Score vs Efficiency

Token Composition

Score Distribution

Task Examples

Part 1 -- Fundamentals (first 6 of 30)

Loading canonical task data...
View all 30 Part 1 tasks →

Part 2 -- Design & Simulation (first 6 of 20)

Loading canonical task data...
View all 20 Part 2 tasks →

Evaluation Rubric

ScoreMeaningCriteria
4CorrectCorrect conclusion and reasoning; topology, device roles, dominant mechanism, trend, and key assumptions are right.
3Mostly correctMain conclusion is right, with a minor omission, imprecision, or modeling flaw that does not change the result.
2Partially correctIdentifies some relevant mechanism, but misses an important circuit detail, trend, or design consequence.
1Mostly incorrectMain conclusion is wrong, but the answer contains a small amount of relevant circuit understanding.
0Incorrect / unusableFundamentally wrong, internally inconsistent, or based on a mistaken topology/device/connection.

Judges evaluate against golden solutions prioritizing analog-circuit reasoning over surface similarity. See GitHub for full details.

Citation

@misc{zhang2026razavibench,
  title        = {Razavi-Bench: An Expert-Curated Benchmark for Analog-Design Reasoning},
  author       = {Zhishuai Zhang and Behzad Razavi},
  year         = {2026},
  howpublished = {\url{https://github.com/Arcadia-1/razavi-bench}},
  url          = {https://razavi-bench.tokenzhang.com/},
  note         = {Benchmark repository}
}