The Two MMLU Scores: What a Benchmark Name Does Not Fix
- Date: September 6, 2026 · Author: Dmitrii Zatona TL;DR Two MMLU accuracies, 0.781 and 0.79, for two builds of one model family under the same benchmark name; for a score-delta query the verifier returns incomparable (Sections 1 and 5).
- The split, the implementation, the prompt format, the grader and the runner’s network access stay open, and where published measurements exist for them the differences are points of accuracy, not thousandths (Section 2).
Unverified
- Date: September 6, 2026 · Author: Dmitrii Zatona TL;DR Two MMLU accuracies, 0.781 and 0.79, for two builds of one model family under the same benchmark name; for a score-delta query the verifier returns incomparable (Sections 1 and 5).
- The split, the implementation, the prompt format, the grader and the runner’s network access stay open, and where published measurements exist for them the differences are points of accuracy, not thousandths (Section 2).
Sources: Zatona