An expert-curated benchmark that tests whether frontier AI
models can reason about analog circuit design.
Razavi-Bench covers 26 frontier models as of August 2026 -- including Claude Fable 5, Claude Opus 5, GPT 5.6, Qwen 3.8 Max Preview, Kimi K3, and Gemini 3.1 Pro -- on 50 expert-written analog design questions, with all scores recalculated using a corrected rubric after a careful audit of Q15.
| # | Model | Score | Range | Token | Cost |
|---|
| Score | Meaning | Criteria |
|---|---|---|
| 4 | Correct | Correct conclusion and reasoning; topology, device roles, dominant mechanism, trend, and key assumptions are right. |
| 3 | Mostly correct | Main conclusion is right, with a minor omission, imprecision, or modeling flaw that does not change the result. |
| 2 | Partially correct | Identifies some relevant mechanism, but misses an important circuit detail, trend, or design consequence. |
| 1 | Mostly incorrect | Main conclusion is wrong, but the answer contains a small amount of relevant circuit understanding. |
| 0 | Incorrect / unusable | Fundamentally wrong, internally inconsistent, or based on a mistaken topology/device/connection. |
Judges evaluate against golden solutions prioritizing analog-circuit reasoning over surface similarity. See GitHub for full details.