deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Math community sets release norms for AI-generated proofs and results

A statement on agmai.org, shaped by more than 600 survey replies, asks frontier labs to stop testing math on private models and details how AI-generated results should be released and funded.

Math community sets release norms for AI-generated proofs and results

A call to stop testing math on private models

A statement published on September 29, 2026, on agmai.org, which surfaced on Hacker News' front page, opens with a blunt request: frontier AI labs should stop running advanced mathematical problems on proprietary models that the wider scientific community cannot examine. The authors say their recommendations were written with that practice as the practical context, but they would prefer labs simply not do it, and they explicitly withhold endorsement of it.

Why mathematicians are worried

The document grounds its argument in long-standing scholarly norms: a paper's authors are expected to understand the argument, verify it themselves, and accept responsibility for it, and researchers behind major advances routinely explain their work in seminars and answer questions from colleagues. AI upends this because models can now produce proofs that nobody in the loop can follow, check, or stand behind. The authors argue that human understanding of mathematics remains paramount, and the statement is an attempt to define what responsible scholarly output looks like when that understanding does not yet exist.

A survey with over 600 replies

Before drafting the recommendations, the group asked the mathematical community what responsible release would mean and received more than 600 replies. According to the statement, the survey concerned a specific episode in which OpenAI announced the existence of many results without providing details. The final recommendations, supported by a clear plurality of respondents, target any AI lab whose models are likely to significantly affect mathematics.

Three principles underpin the document. Significant results should be responsibly released as soon as possible. Labs that release substantial output without immediate human understanding must take responsibility for ensuring that understanding follows, including through funding. And the development of that understanding must remain organic and community-led rather than directed by the labs, even when the labs produced the results.

Two tracks depending on who understands the result

The recommendations split on whether a human understands the work. If a mathematician fully understands the content and takes responsibility for it, ordinary academic norms apply: post a preprint, submit to a journal for peer review, and give talks explaining the work.

For output that nobody yet understands, the statement proposes technical release norms followed by a support obligation.

What an initial release should include

The technical norms are unusually concrete:

  • Before release, labs should use LLMs to scour the literature for ideas related to the proofs and cite the papers that introduced them, even if the model discovered those ideas independently.
  • A model should produce each proof in the style of a traditional mathematical paper, with accessible introductions and precise theorem statements, rather than verbose text full of non-standard terminology. The authors concede current LLMs may not match mathematicians' standards of attribution and exposition, but say this does not excuse labs from doing their best.
  • Results should be deposited promptly in scholarly repositories that no AI lab controls, with persistent citable identifiers and recorded modifications; comment functionality would be particularly useful. Labs are strongly urged not to treat releases as marketing vehicles for their models, given the negative externalities this imposes on the community.
  • For each result, labs should publish the model's name, the prompts used, a summarized chain of thought, the time taken, and the estimated computational cost. Releasing raw outputs from before cleanup is encouraged.
  • Proofs should be formalized as far as possible, with artifacts meeting community standards, including copyright headers, a challenge file for a comparator, and a formalization.yaml file, plus machine-readable metadata correlating natural-language and formal versions. Where formalization would cause unacceptable delays, its status should be clearly stated.
  • Each release should document how AI came to be used on that problem, and bulk releases should include a further document referencing all results, reporting how many comparable problems the models attempted and failed to solve, and explaining how the problems were chosen.

The support obligation

For unassimilated output, the statement's second step calls on labs to support the additional mathematical work needed for humans to understand the results, absorb them, and identify possible applications. The overarching principles put it more bluntly: that support should include significant funding. The published page also notes a further section commenting on access to powerful models.

Why it matters

Mathematics has no experimental fallback for verification; correctness rests almost entirely on humans working through arguments. If labs release large volumes of proofs that no one can vouch for, the discipline either absorbs a costly verification workload or lets unverified claims circulate. This statement shifts that cost back to its source: labs that generate results must fund the human understanding that legitimizes them. The disclosure requirements, especially publishing failed attempts alongside successes, would also make it harder to create inflated impressions of model capability through selective announcements. And the insistence on lab-independent repositories with persistent identifiers treats AI-generated proofs as scholarly records rather than product demos, a framing other fields confronting AI-generated research may end up copying.

  • #mathematics
  • #research-ethics
  • #open-science
  • #peer-review
  • #ai-labs