GLM-5.3 scores 49.4 on the Respan Index (#24), ahead of Gemini 3.7 Flash at 48.6 (#26). Gemini 3.7 Flash leads on 13 of the 22 benchmarks both report. Gemini 3.7 Flash is 1.4x cheaper per token at list price. Gemini 3.7 Flash has the larger context window (1.0M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 71.6% |
| 68.8% |
| FrontierSWE V2 (Proximal Labs) | 20.3% | 30.2% |
| GPQA Diamond | 94.8% | 90.9% |
| LiveBench Agentic Coding (LiveBench) | 58.3% | 60.9% |
| LiveBench Coding (LiveBench) | 78.9% | 79% |
| LiveBench Data Analysis (LiveBench) | 68% | 70.2% |
| LiveBench Instruction Following (LiveBench) | 79.9% | 69.3% |
| LiveBench Language (LiveBench) | 85.5% | 79.9% |
| LiveBench Mathematics (LiveBench) | 93.5% | 87.9% |
| LiveBench Reasoning (LiveBench) | 87.8% | 85.8% |
| Mystery Game Puzzles (Epoch AI) | 37% | 33% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 97.2% | 91.1% |
| SimpleQA Verified | 69.2% | 41% |
| Terminal-Bench 4.0 (Vals AI) | 12.1% | 38.9% |
| LMArena Agent (LMArena) | -0.0149 | 0.0236 |
| LMArena Elo | 1488 | 1479 |
| LMArena WebDev (LMArena) | 1592 | 1623 |
| DeepSWE v1.1 | 65.3% | 66.9% |
Higher is better. Each score links to where it was published.