菜单

关于 🐙 GitHub
arXiv 提交日期: 2026-06-04
📄 Abstract - Benchmarks in Leipzig

Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3-day workshop *Benchmarks in Leipzig* with 35 participants at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. We present the resulting collection of 100 questions. We evaluated these questions in three stages: a single attempt by five state-of-the-art LLMs, followed by a 20-runs-per-model evaluation with three of these models, and finally a 3-run attempt with two heavy-thinking models. After Stage 1, 41 questions remained completely unsolved; after Stage 2, this count dropped to 16; and we concluded Stage 3 with only 2 unsolved questions. This demonstrates that the mathematical reasoning capabilities of LLMs are becoming impressive.

顶级标签: llm benchmark machine learning
详细标签: math reasoning evaluation dataset question answering 或 搜索:

莱比锡基准测试 / Benchmarks in Leipzig


1️⃣ 一句话总结

本文介绍了一个由49位数学家合作创建的高难度数学问答数据集,包含100个研究级问题,并通过三轮逐步加强的测试(从单次尝试到深度思考模型多次尝试)评估了最先进的大语言模型,结果显示模型能力惊人,最终仅剩2个问题未被解决。

源自 arXiv: 2606.05818