Latest Results
Overall Quality Index
Mean of all benchmark scores per model, latest run.
Latest Scores by Benchmark
Score Trends Over Time
Overall Quality Index
Average Latency Trend (s)
Estimated Daily Benchmark Cost (USD)
Model Comparison
Latest run vs. previous run (Δ shows change).
Radar
Run Details
About the Benchmarks
Each benchmark uses a fixed, frozen subset of questions so daily scores stay comparable over time. All scoring is fully programmatic (no LLM judge): exact-match, code execution, or rule checks.