Razavi-Bench

Tasks50
Models32
MultimodalYes

An expert-curated benchmark that tests whether frontier AI
models can reason about analog circuit design.

Direct QA Leaderboard

#ModelScore RangeToken Cost

Score vs Efficiency

Token Composition

Score Distribution

Task Examples

Part 1 -- Fundamentals (first 6 of 30)

Loading canonical task data...
View all 30 Part 1 tasks →

Part 2 -- Design & Simulation (first 6 of 20)

Loading canonical task data...
View all 20 Part 2 tasks →

Evaluation Rubric

ScoreMeaningCriteria
4CorrectCorrect conclusion and reasoning; topology, device roles, dominant mechanism, trend, and key assumptions are right.
3Mostly correctMain conclusion is right, with a minor omission, imprecision, or modeling flaw that does not change the result.
2Partially correctIdentifies some relevant mechanism, but misses an important circuit detail, trend, or design consequence.
1Mostly incorrectMain conclusion is wrong, but the answer contains a small amount of relevant circuit understanding.
0Incorrect / unusableFundamentally wrong, internally inconsistent, or based on a mistaken topology/device/connection.

Judges evaluate against golden solutions prioritizing analog-circuit reasoning over surface similarity. See GitHub for full details.

Citation

@misc{zhang2026razavibench,
  title        = {Razavi-Bench: An Expert-Curated Benchmark for Analog-Design Reasoning},
  author       = {Zhishuai Zhang and Behzad Razavi},
  year         = {2026},
  howpublished = {\url{https://github.com/Arcadia-1/razavi-bench}},
  url          = {https://razavi-bench.tokenzhang.com/},
  note         = {Benchmark repository}
}