· via Hacker News – Front Page (native)
FutureHouse's Millennium Problems for Biology pair grand challenges with AI-ready success criteria
FutureHouse has published 'Millennium Problems for Biology': precisely specified challenges from origin-of-life chemistry to zero-shot protein design, several structured as benchmarks that rule out lab screening.

FutureHouse has published a set of "Millennium Problems for Biology": grand-challenge research targets, each paired with unusually concrete pass/fail criteria. The page at millenniumproblems.bio, credited to Edison Scientific at FutureHouse, drew broad attention after reaching the Hacker News front page on 20 September 2026. Like the Clay Mathematics Institute's Millennium Prize Problems that the name echoes, the idea is to name hard, verifiable goals — but here success is defined through numbers, assays and preregistered targets rather than proofs, and no monetary prize is mentioned.
What the list asks for
The published problems span biology from its origin to its engineering. One asks for the unassisted emergence of self-replicating, RNA- and protein-based cells from a plausible primordial soup and energy source. To count, the emergent cells must multiply a millionfold — roughly 20 generations — plausibly keep dividing indefinitely without shrinking away across generations, and carry heritable genetic information that causally shapes their molecular makeup. Proposals where the existence of heredity is debatable are rejected outright.
Two problems concern whole organisms. One demands reversible whole-body cryopreservation of live, wild-type adult mice: frozen or vitrified for at least 24 hours, recovered with over 99% viability and no permanent organ damage, with ethics approval required. The other asks for reproducible regeneration of amputated limbs in adult wild-type mice, where motor function and sensory perception match control animals and blinded observers cannot tell which limb regrew.
On the molecular side, the list calls for an enzyme that can "reverse translate" — reading an untagged polypeptide and writing out a nucleic acid strand that encodes its residue sequence under a preregistered codon convention, with no nucleic-acid template, barcodes or lookup tables involved. Validation requires inferring the sequences of 100 preregistered random peptides, each at least 50 amino acids long, at 90% accuracy or better, with average reads of 25 residues and quality scores of at least Q10. Another problem asks for a Rubisco that beats nature's trade-off between carbon-dioxide specificity and catalytic turnover, matching or exceeding Galdieria partita Rubisco on specificity and maize Rubisco on kcat in paired assays. A fifth, contributed by Erika Alden DeBenedictis, asks for a living, replicating cell in which every protein-coding sequence — including the translation machinery itself — is written in genuine four-base codons, with no detectable triplet decoding.
The remaining problems lean toward biotechnology. One asks for replication-incompetent yet infectious AAV and lentivirus particles produced in bacteria, carrying pre-specified genomes with quality ratios comparable to mammalian cell culture; producing AAV alone counts as partial success. The final two, described below, are framed explicitly as design benchmarks. Note that the retrieved text of the page is truncated, so the list may include further problems beyond those quoted here.
Rules built for computational design
Two problems read less like biology goals than like benchmarks for AI-driven design. The protease challenge requires producing, within 24 hours of receiving 20 preregistered target sites on endogenous folded proteins, enzymes that cleave those sites efficiently inside living cells with a success rate above 80% — and no wet-lab screening or target-specific evolution is permitted once the targets are announced. The protein-binder challenge is similar: zero-shot binders against 20 preregistered intracellular targets, delivered extracellularly at pharmacologically plausible concentrations, again with an 80% success bar and a 24-hour design window that rules out screening. Both effectively test whether a system can predict protein behavior from sequence and structure alone, which is precisely the capability that current AI protein-design systems are being built around.
Why it matters
Biology rarely gets this kind of scoreboard. By expressing success as measurable thresholds — a millionfold expansion, 99% viability, 90% read accuracy, 80% hit rates on preregistered targets — the list turns sprawling research ambitions into auditable milestones that a human lab or an AI system can be graded against. The no-screening clauses matter most: they separate genuine predictive design from brute-force directed evolution, a distinction that becomes critical as AI agents take on scientific work. Several problems would also unlock real capabilities if solved, from demonstrable origin-of-life chemistry and organ banking to programmable proteolysis, better photosynthesis, an expanded genetic code and cheaper gene-therapy manufacturing. Whether or not any are solved soon, they now exist as named, checkable targets — and as a template for how other fields might specify what progress actually looks like.
- #biology
- #protein-design
- #synthetic-biology
- #benchmarks
- #ai