When OpenAI released hundreds of claimed solutions to some of the world’s hardest math problems this week, the frontier lab stated that it had consulted an advisory group of elite mathematicians to avoid the controversy that arose during its previous model's solution of a long-standing problem in the field. However, OpenAI fell short of the standards set by this advisory group, particularly regarding the need for human understanding of mathematical results.
Advisory Group's Concerns
A new paper has highlighted gaps between the natural language and formally expressed solutions to a million-dollar problem that was ostensibly solved by OpenAI’s models. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton University’s Institute for Advanced Studies, consists of nine prominent researchers from institutions worldwide. They released guidelines for frontier labs solving math problems at the end of September.
In a statement regarding the latest proofs, AGMAI noted, "it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully." The organization’s first request was to stop testing advanced mathematical problems on proprietary models. OpenAI’s release explicitly states that it is evaluating its proprietary models using open research problems in mathematics.
Evaluation and Transparency Issues
The advisory group did not respond when asked by TechCrunch for a more thorough evaluation of OpenAI’s latest proof release. The lab did follow some of AGMAI's principles, including releasing results promptly and providing information about how the models reached their conclusions. However, only ten of the 719 manuscripts included releases of the model’s chain of thought. For papers that are difficult to understand, the mathematicians suggested that proofs should be formalized, yet only 42% of the proofs released by OpenAI had undergone this process.












