Researchers from Texas A&M University, Harvard University, Vanderbilt University, Stanford University, the University of North Carolina at Chapel Hill, and DARPA have introduced the concept of a "Mathematical Primitive" to assess the structural mathematical understanding of large language models (LLMs). Their work, detailed in a paper titled "The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models," suggests that an LLM's accuracy on mathematical problems can mask a lack of deeper comprehension.
The research team, including Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, and Zhengzhong Tu, developed a new benchmark called PRIM. This benchmark evaluates mathematical reasoning across four distinct dimensions: Discovery, Generation, Digestion, and Execution. Discovery measures a model's ability to independently identify the core conceptual structure of a problem. Generation assesses standard zero-shot chain-of-thought performance. Digestion determines if a model can recognize the primitive when a correct solution is provided, and Execution evaluates problem-solving when the correct primitive is explicitly given.
The study's diagnosis indicates that similar solution accuracy scores among LLMs can hide vastly different underlying capability profiles. Many models can perform calculations or generate derivations that appear correct without identifying the compact idea that makes the problem solvable. Conversely, when a model is provided with the correct primitive, it often demonstrates a substantial latent capacity to complete the remaining steps. The researchers identified Discovery, the ability to independently spot the mathematical primitive from the problem statement, as the dominant bottleneck in mathematical reasoning for LLMs.
The PRIM benchmark utilizes a curated set of challenging mathematical problems from Humanity's Last Exam (HLE), a benchmark designed to push the boundaries of AI capabilities. HLE includes 2,500 to 3,000 questions across various subjects, requiring graduate-level expertise and resisting quick internet retrieval. The primitives for the PRIM benchmark were initially drafted by a strong model (GPT-5.4-High) and then reviewed and verified by three human experts with graduate-level mathematical training.
Building on these findings, the researchers introduced Absorb, a primitive-privileged self-distillation framework. Instead of training models to generate primitives, which the authors found could lead to performance regressions, Absorb uses the primitive as privileged information to guide a teacher model during training. This approach allows the student model to internalize structural reasoning without requiring explicit primitives during inference. Experiments on various model scales and benchmarks, including HLE, HMMT25, and Omni-MATH, showed that Absorb consistently improved mathematical reasoning compared to baselines. The researchers noted that failures primarily caused by a lack of Discovery were more amenable to repair through post-training methods.
