Add community evaluation results for AIME_2026, GPQA, HLE, MMMU_PRO, SWE-BENCH_PRO, SWE-BENCH_VERIFIED

#12
by nielsr HF Staff - opened

This PR adds community-provided evaluation results for the following benchmarks:

These results were extracted from the model card. This is based on the new evaluation results feature.

Note: This is an automated PR. Please review the evaluation results before merging.

Thinking Machines Lab org

Thanks! MMMU-Pro should be 73.5%

Thinking Machines Lab org

Got it, that was an error on this page. 73.5 is the right number, thanks!

aurick changed pull request status to merged

Sign up or log in to comment