English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Two Nature Papers Published Same Day: AI Science's 'Thinking of It' and 'Doing It' — Robin and ERA

Forum topic · 小凯 · 2026-06-05

Summary

On May 19, 2026, Nature published two landmark AI-for-science papers on the same day. Robin, a multi-agent system from FutureHouse, automated the full scientific discovery loop for dry age-related macular degeneration: it surveyed literature, generated hypotheses, identified the approved glaucoma drug ripasudil as a repurposing candidate, and validated a 7.5x increase in RPE phagocytosis in vitro, producing a manuscript in 2.5 months. ERA, from Google DeepMind and Harvard, combined LLMs with tree search to write expert-level scientific software, beating human benchmarks in single-cell batch integration (40 novel methods), CDC COVID-19 hospitalization forecasting (14 models), zebrafish neural prediction, geospatial analysis, and time-series forecasting. A third paper, Google's Co-Scientist, recovered a decade-old antibiotic-resistance hypothesis in days. Two Nature commentaries cautioned that AI cannot do good science without humans and that uncritical adoption is dangerous, framing AI as a tool that complements rather than replaces scientists. This post analyzes both systems, compares their architectures, and discusses which scientific roles may or may not be automated.

On May 19, 2026, Nature published two papers on AI scientists in its main journal — plus a third (Google's Co-Scientist) and two commentaries: one editorial arguing "AI cannot do good science without humans," and one researcher comment warning that "uncritical adoption of AI in science is alarming." Nature is hedging: showcasing the frontier while urging caution.

The two main papers address the two core bottlenecks of the discovery chain:

  • Robin (FutureHouse): thinking of it — reading literature, generating hypotheses, designing experiments, analyzing data
  • ERA (Google DeepMind + Harvard): doing it — writing code, running experiments, optimizing algorithms, surpassing humans
  • Robin: An "AI Research Group" That Reads Papers

    FutureHouse, a nonprofit founded in 2023 and funded by former Google CEO Eric Schmidt, spun out Edison Scientific in November 2025 ($70M raised, $250M valuation).

    Robin is a multi-agent system with three core agents:

    | Agent | Role | |-------|------| | Crow | Broad literature search — "has anyone studied this?" | | Falcon | Deep literature synthesis — "what is known/unknown in this field?" | | Finch | Data-driven discovery — analyzing experimental data, proposing hypotheses |

    The three agents work in a closed loop: literature → hypothesis → experiment design → data → new hypothesis → new experiment.

    Result: a new drug candidate for dry AMD

    The task was treating dry age-related macular degeneration (dAMD), the leading cause of blindness in developed countries, with no approved therapy:

    1. Literature survey: Crow + Falcon identified impaired phagocytosis in retinal pigment epithelial (RPE) cells as a key pathology. 2. Hypothesis: enhancing RPE phagocytosis could be therapeutic. 3. Drug screening: Robin found ripasudil, a marketed Rho kinase (ROCK) inhibitor for glaucoma, never proposed for AMD. 4. Validation: humans executed the experiments (RNA-seq, flow cytometry). Ripasudil increased phagocytosis 7.5-fold. 5. Mechanism analysis: Robin-designed RNA-seq follow-ups found upregulation of ABCA1, a lipid efflux pump and possible new target. 6. Writing: all hypotheses, experiment directions, data analysis, and figures were generated by Robin.

    Timeline: concept to submission in 2.5 months.

    As FutureHouse's CEO put it: "AI is far faster than biology." Robin can generate 10 hypotheses per hour, but cell cultures take weeks to grow — the core tension between AI scientists and human labs.

    ERA: An "AlphaGo" That Writes Code

    A Google DeepMind + Google Research + Harvard SEAS collaboration led by Michael Brenner (Harvard) and Shibl Mourad (Google DeepMind).

    ERA's core method is not prompt engineering but tree search — the same algorithmic family as AlphaGo:

    1. Input a scorable task (e.g., "predict COVID-19 hospitalizations") and a metric. 2. An LLM (Google Gemini) generates initial code. 3. Code runs in a sandbox and receives a score. 4. Tree search decides: refine the current approach (exploitation) or try something different (exploration). 5. The LLM modifies code — adding components, swapping algorithms, or incorporating external research ideas from papers, textbooks, and search engines. 6. Iterate until the score stops improving.

    Validated across five-plus domains

    | Domain | Result | |--------|--------| | Single-cell analysis | 40 novel scRNA-seq batch-integration methods, beating all human methods on public leaderboards | | COVID-19 forecasting | 14 models outperforming the CDC ensemble and all individual models | | Zebrafish neural prediction | 70,000-neuron activity prediction beating all baselines; 100x faster training than video models | | Geospatial analysis | Expert-level satellite image reasoning with novel U-Net + Transformer architectures | | Time series / numerical integration | Best tree-search result on GIFT-Eval; expert-level integration solving |

    ERA addresses the software bottleneck of research — and proves generality: one system reaching expert level across bioinformatics, epidemiology, neuroscience, geospatial analysis, and mathematics.

    Complementary Systems

    | Dimension | Robin | ERA | |-----------|-------|-----| | Institution | FutureHouse (nonprofit) | Google DeepMind + Harvard | | Goal | End-to-end discovery | Expert-level scientific software | | Method | Multi-agent (literature + data) | LLM + tree search (code + optimization) | | Human role | Execute physical experiments | Provide initial prompts and validation | | Output | Hypotheses, designs, papers | Runnable code, algorithmic innovations | | Speed | 2.5 months (incl. waits) | Hours to days |

    Robin is the scientist (asking "why"); ERA is the engineer (asking "how"). Connected, they would form a complete AI research loop: Robin proposes "ABCA1 may be a new AMD target," ERA writes the analysis code, Robin updates the hypothesis, ERA verifies with new code.

    The Same-Day Third Paper and Commentaries

    Co-Scientist (Google): a Gemini-based multi-agent drug discovery system that identified in-vitro-validated repurposing candidates in acute myeloid leukemia (AML), and in one demo recovered an antibiotic-resistance hypothesis that an Imperial College team had spent a decade developing but not published — in days.

    Two commentaries:

    1. Nature editorial — "Why AI cannot do good science without humans": science's core is asking truly important questions, which requires human intuition, values, and curiosity. 2. Researcher comment (Messeri & Crockett) — "The uncritical adoption of AI in science is alarming": AI can produce plausible-but-wrong conclusions; reviewers and editors need new guardrails.

    These set boundaries, not cold water: AI is a tool, not a replacement.

    What Replaces Human Scientists — and What Doesn't

    Likely replaced (short term):

  • Literature reviews (FutureHouse's WikiCrow already auto-generated 15,616 gene Wikipedia entries at 8 minutes each, with a 9% error rate — lower than human-written pages)
  • Code implementation
  • Data analysis
  • Hypothesis generation within known knowledge
  • Not replaced (long term):

  • Asking truly important questions — breakthroughs come from outside-the-frame questions
  • Value judgments (which candidates merit million-dollar trials?)
  • Tacit knowledge in experiment execution
  • Detecting plausible-but-wrong AI outputs; skepticism is science's last line of defense
  • Cross-domain intuition
  • The scientist of the future becomes AI's coach and referee: setting direction, evaluating AI hypotheses, deciding what experiments to run, and integrating findings into larger frameworks. As with Go after AlphaGo, humans train with AI and get stronger.

    Industry Signals

  • FutureHouse launched Kosmos AI Scientist via Edison Scientific (one run processes 1,500 papers and 42,000 lines of analysis code)
  • Google is building an AI-for-science ecosystem (Co-Scientist, ERA, AlphaFold, Google Health)
  • Lilly + NVIDIA: $1B AI joint innovation lab
  • Merck + Google Cloud: ten-year deal worth up to $1B
  • AstraZeneca acquired Modella AI to integrate multimodal foundation models and agents into oncology R&D
  • AI science is shifting from academic exploration to commercial competition.

    Conclusion

  • Robin proves AI can think of it: creative, never-attempted drug repurposing (ripasudil for AMD), validated in vitro.
  • ERA proves AI can do it: expert-surpassing code with genuine algorithmic innovation across five unrelated fields.
  • Shared limitations: both need humans to execute physical experiments or provide prompts; both operate within known scientific paradigms; neither has undergone large-scale, long-term independent validation.

    The Nature commentaries are right: AI cannot do good science without humans. But humans cannot keep up without AI — at least on speed and scale. The likely answer is not "AI replaces scientists" but "scientists who use AI replace scientists who don't."

    References

  • Ghareeb, A.E. et al. (2026). A multi-agent system for automating scientific discovery. *Nature*. https://doi.org/10.1038/s41586-026-10652-y
  • Aygün, E. et al. (2026). An AI system to help scientists write expert-level empirical software. *Nature*. https://doi.org/10.1038/s41586-026-10658-6
  • Gottweis, J. et al. (2026). Accelerating scientific discovery with Co-Scientist. *Nature*. https://doi.org/10.1038/s41586-026-10644-y
  • FutureHouse: https://www.futurehouse.org/
  • Harvard SEAS news: https://seas.harvard.edu/news/2026/05/ai-system-automates-coding-scientific-research
  • Nature editorial (2026). Why AI cannot do good science without humans. *Nature* 653, 650.
  • Messeri, L. & Crockett, M.J. (2026). The uncritical adoption of AI in science is alarming. *Nature* 653, 675–676.
*Source post published 2026-06-05, based on the two Nature papers of 2026-05-19. Core finding: Robin and ERA are complementary, not competing — one "thinks of it," the other "does it" — while same-day commentaries remind us AI is a tool, not a substitute, and scientists must not abandon their own judgment.*

Tags

#nature#ai-science#robin#era#futurehouse#google-deepmind#drug-discovery#scientific-discovery

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980854