TIDEZINE.

OpenAI Releases 722 Math Manuscripts, Single-Prompt Problem Solving Sparks Verification Controversy

OpenAI has publicly released 722 manuscripts on GitHub, claiming progress on 372 hard math problems. Most results came from a single prompt handed to a single AI agent, but the model itself remains unreleased, and specific prompts and compute times were not disclosed.

OpenAI Releases 722 Math Manuscripts, Single-Prompt Problem Solving Sparks Verification Controversy

722 manuscripts, 372 hard math problems, with each result averaging roughly three hours of ChatGPT Pro usage—that's the scale of what OpenAI has just disclosed. But what's really intriguing is the method: according to what OpenAI told Scientific American, nearly every single paper was produced using one prompt handed to one AI agent.

All of the manuscripts have been uploaded to GitHub, verified by OpenAI's newly formed advisory group, the Advisory Group on Mathematics and Artificial Intelligence (AGMAI). Last month, OpenAI had already signaled its intentions, claiming its model "solved more than 100 long-standing open problems spanning most areas of mathematics." This release essentially lays out the concrete list. The company stated that the results were produced by a ChatGPT pilot model not yet released to the public.

Among the headline results OpenAI claims are a solution to the four-dimensional Kakeya conjecture, improvements to several key computer algorithms, and progress on the Riemann hypothesis—any one of which, if confirmed, would be a major event in the math world. In its announcement, the company wrote that it is starting with a GitHub repository release, complete with mechanisms for paper revisions and citations, "while still seeking other community-hosted options that meet the advisory group's guidelines," and promised that future releases will be more complete in terms of citation standards, mathematical exposition, and presentation.

But OpenAI only followed through on half of AGMAI's recommended transparency. The advisory group had originally recommended disclosing the compute time and prompts used for each individual problem, but OpenAI did not do this, instead only providing an overview of the overall reasoning process, compute estimates, and the number of problems attempted. In other words, the public can see the list of results, but cannot see the specific path that would let anyone verify how those results were actually produced.

This is exactly where the controversy lies. Prior to this, OpenAI's claims related to Navier-Stokes had already stirred up a wave of backlash in the math community, leaving many scholars wary of "one-shot AI problem solving" claims. MIT mathematician Andrew Sutherland told Scientific American: "Unless they release the model so results can be reproduced, any claim of a single agent solving a problem in one shot should be considered unverified."

The batch of manuscripts OpenAI has released still needs to be reviewed and verified by mathematicians one by one, and the real impact won't be clear until that review process concludes. As for whether the model itself will ever be made public, or when the specific prompts will be filled in, OpenAI has not given a timeline.

Related

Florida AG Files Emergency Injunction Against OpenAI, Demanding Halt to New Model Training and Block on Minor Users
Tech

Florida AG Files Emergency Injunction Against OpenAI, Demanding Halt to New Model Training and Block on Minor Users

Florida Attorney General Uthmeier filed an emergency motion on Monday asking the court to bar OpenAI from training new models without independent safety oversight and to cut off minors' access to ChatGPT, in a case stemming from a criminal investigation earlier this year into the FSU shooting.

$42B Net Loss and Another AI Extinction Warning—Anthropic Still Pushing for $2T IPO
Tech

$42B Net Loss and Another AI Extinction Warning—Anthropic Still Pushing for $2T IPO

Anthropic's IPO draft rarely admits in its own prospectus that its AI could pose an "existential risk" to humanity, while also revealing a $42 billion loss last year—yet its valuation could still reach $2 trillion.

Out of Battery? Ring's New Smart Lock Comes With a Hand-Crank Dial
Tech

Out of Battery? Ring's New Smart Lock Comes With a Hand-Crank Dial

Ring founder Jamie Siminoff believes dead-battery anxiety is the key reason smart locks haven't gone mainstream. The new Smart Lock features a built-in hand-crank generator—just turn the handle for emergency access. It's slated to launch in early 2027 at $249.

Type to Build a Game: Google Launches Experimental Platform Playground, But Your Google One Tier Decides What You Get
Tech

Type to Build a Game: Google Launches Experimental Platform Playground, But Your Google One Tier Decides What You Get

Google has launched Playground, an AI-powered game creation platform that lets you generate 2D or 3D games with just a few prompts—no coding, no drawing required. For now, though, it's limited to users in the US who are 18 or older, and access is tiered based on your Google One membership level.

After Skydance's $76B Merger, HBO Max and Paramount+ Set to Combine Into One Streaming Service
Tech

After Skydance's $76B Merger, HBO Max and Paramount+ Set to Combine Into One Streaming Service

Skydance has announced that its three streaming platforms—HBO Max, Paramount+, and Discovery+—will gradually merge into one. The news comes alongside the completion of the $110 billion Warner Bros. Discovery acquisition, though the name and pricing of the new service have yet to be determined.

$11 Billion Merger Sealed: Paramount and Warner Bros. Officially Fold into Skydance
Tech

$11 Billion Merger Sealed: Paramount and Warner Bros. Officially Fold into Skydance

After nearly a year of negotiations and regulatory review, the $11 billion merger between Paramount and Warner Bros. Discovery is officially complete. The new company will be named Skydance — and it comes with $80 billion in debt attached.