Overview
This forum post reviews OmniScientist: An Omni-Modal Omni-Discipline AI Scientist by Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, and Wynne Hsu (arXiv: 2608.13558).
Key points
- The problem: Current "AI scientists" are effectively blind — they read papers, write code, and run experiments, but only consume pre-digested information (text summaries, feature vectors, scalar scores). Scientific discovery, however, depends on spatial, temporal, cross-channel, and procedural relationships visible only in raw evidence: microscopy images, spectral shifts, video of animal behavior, 3D protein structures, anomalous time-series pulses.
- Core idea: Human discovery begins with observation (Galileo, Fleming, Penzias & Wilson). OmniScientist lets the AI "go outside" and perceive the world directly.
- Ideation Agent: generates hypotheses grounded in raw evidence (e.g., noticing subtle R-R interval variability in raw ECG waveforms). Runs novelty screening, rigour checks, and execution provenance — every idea traceable to specific observations.
- Experiment Agent: designs and runs experiments, independently verifying claims by loading data, running analyses, checking p-values, and plotting. Maintains numerical traceability: every number in the final paper traces back to code and input data.
- Writeup Agent: assembles a full academic paper shaped by the entire research narrative, in a deterministic pipeline with lifecycle-wide perception — observations can redirect research direction, revise hypotheses, or trigger methodological re-examination at any stage.
- Tested on 36 real data cases across 5 discipline families (medicine, materials science, earth science, biology, physics), 4 evidence types, and 9 modalities.
- Completed the full raw-data-to-paper pipeline in all 36 cases, with an average paper score of 6.3 (reference reasoning backbone).
- A blind variant receiving only pre-computed scalar features (e.g., "mean heart rate 72 bpm, QT 380 ms, ST shift 0.5 mm" instead of the raw ECG) was compared head-to-head:
- Direct perception won 85% of head-to-head judgments, outperforming the blind variant on all 7 evaluation dimensions: motivation/significance, hypothesis clarity, method appropriateness, result support, conclusion soundness, writing quality, and overall impression.
- A 6.3 average score means the system cannot yet replace human scientists; the blind variant still performs reasonably in some cases where precomputed features capture key information.
- The envisioned future: AI scientists that physically operate lab equipment, adjust microscope focus, collect field samples, and synthesize materials — a genuine embodied agent, not just a language model on a server.
Architecture
The system is structured like a three-story research building:
1. Perception Layer
Accepts nine modalities as raw inputs (not pre-converted to text): images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. Examples: seeing actual MRI scans of Alzheimer's patients (gray-matter loss regions, white-matter lesions) rather than a correlation coefficient; seeing satellite cloud-map sequences of extreme weather rather than a statistic like "heatwave frequency up 37%."2. Three Autonomous Agents
Experimental results
Interpretation
The author offers a restaurant-review analogy: a traditional AI reads reviews and star ratings without ever tasting the food, while OmniScientist walks in, watches the chef, listens, and tastes. The shift is framed as a third coming-of-age for AI — from symbolic to connectionist, from unimodal to multimodal, and now from passive reception to active perception: a philosophical move from "text-centrism" to "perception-centrism."
Limitations and outlook
> "What I cannot create, I do not understand."
...augmented with: *"What I cannot perceive, I cannot discover."* Perception is the starting point of discovery.
Reference
Li, B., Fei, H., Ju, T., Lee, M.-L., & Hsu, W. (2026). *OmniScientist: An Omni-Modal Omni-Discipline AI Scientist*. arXiv preprint arXiv:2608.13558.