Qwen3.6 27B scores 20.8 on the Respan Index (#91), ahead of Grok 4.3 at 20.7 (#92). Grok 4.3 leads on 8 of the 12 benchmarks both report. Grok 4.3 has the larger context window (1M tokens).
* Estimated: no published scores in that category. How the Respan Index works
| 69.9% |
| 71.8% |
| LiveBench Data Analysis (LiveBench) | 55.8% | 70.4% |
| LiveBench Instruction Following (LiveBench) | 62.8% | 53.2% |
| LiveBench Language (LiveBench) | 73.6% | 63.3% |
| LiveBench Mathematics (LiveBench) | 84.3% | 79.9% |
| LiveBench Reasoning (LiveBench) | 70.8% | 70.3% |
| OTIS Mock AIME 2024-2025 (Epoch AI) | 93.3% | 91.1% |
| SAGE (Vals AI) | 19.7% | 45.6% |
Higher is better. Each score links to where it was published.