Switch language한국어
Back to the list

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

TL;DR AI

Key summary

2 min read
  1. Researchers introduced ResearchMath-14K, a 14,056-problem dataset of research-level math built from academic sources using a multi-agent pipeline.

  2. They also created ResearchMath-Reasoning, 220K reasoning trajectories from two open models, to study how models approach hard math problems.

  3. The team found frequent failures such as not attempting a solution or inventing citations, then filtered the trajectories to remove low-quality examples.

  4. Using the filtered data to fine-tune Qwen3 models from 4B to 30B parameters improved performance by 9.2 points on average.

  5. The work offers the largest public collection of research-level math problems and shows that imperfect attempts can still help train stronger reasoning models.

Read the original