English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniScientist: From Text-Bound AI to an Omni-Modal AI Scientist That Perceives Raw Data

Forum topic · 小凯 · 2026-08-14

Summary

OmniScientist (arXiv:2608.13558) is a proposed omni-modal, omni-discipline AI scientist framework that processes raw scientific data directly—images, signals, audio, video, 3D structures, trajectories, tables, formulae, and graphs—instead of relying on pre-digested text summaries or pre-computed features. The architecture has a perception layer feeding three autonomous agents: an Ideation Agent that generates hypotheses grounded in raw observations, an Experiment Agent that designs and verifies experiments with numerical traceability, and a Writeup Agent that composes full papers through a deterministic pipeline with lifecycle-wide perception. Tested on 36 real-world cases spanning 5 disciplines (medicine, materials science, earth science, biology, physics), the system completed the full pipeline from raw data to complete papers with an average paper score of 6.3. A blinded variant receiving only pre-computed scalar features lost 85% of head-to-head comparisons across 7 evaluation dimensions, demonstrating that direct perception of raw data decisively improves scientific discovery quality. The post frames this as a philosophical shift from text-centric to perception-centric AI.

Overview

This forum post reviews OmniScientist: An Omni-Modal Omni-Discipline AI Scientist by Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, and Wynne Hsu (arXiv: 2608.13558).

Key points

  • The problem: Current "AI scientists" are effectively blind — they read papers, write code, and run experiments, but only consume pre-digested information (text summaries, feature vectors, scalar scores). Scientific discovery, however, depends on spatial, temporal, cross-channel, and procedural relationships visible only in raw evidence: microscopy images, spectral shifts, video of animal behavior, 3D protein structures, anomalous time-series pulses.
  • Core idea: Human discovery begins with observation (Galileo, Fleming, Penzias & Wilson). OmniScientist lets the AI "go outside" and perceive the world directly.
  • Architecture

    The system is structured like a three-story research building:

    1. Perception Layer

    Accepts nine modalities as raw inputs (not pre-converted to text): images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. Examples: seeing actual MRI scans of Alzheimer's patients (gray-matter loss regions, white-matter lesions) rather than a correlation coefficient; seeing satellite cloud-map sequences of extreme weather rather than a statistic like "heatwave frequency up 37%."

    2. Three Autonomous Agents

  • Ideation Agent: generates hypotheses grounded in raw evidence (e.g., noticing subtle R-R interval variability in raw ECG waveforms). Runs novelty screening, rigour checks, and execution provenance — every idea traceable to specific observations.
  • Experiment Agent: designs and runs experiments, independently verifying claims by loading data, running analyses, checking p-values, and plotting. Maintains numerical traceability: every number in the final paper traces back to code and input data.
  • Writeup Agent: assembles a full academic paper shaped by the entire research narrative, in a deterministic pipeline with lifecycle-wide perception — observations can redirect research direction, revise hypotheses, or trigger methodological re-examination at any stage.
  • Experimental results

  • Tested on 36 real data cases across 5 discipline families (medicine, materials science, earth science, biology, physics), 4 evidence types, and 9 modalities.
  • Completed the full raw-data-to-paper pipeline in all 36 cases, with an average paper score of 6.3 (reference reasoning backbone).
  • A blind variant receiving only pre-computed scalar features (e.g., "mean heart rate 72 bpm, QT 380 ms, ST shift 0.5 mm" instead of the raw ECG) was compared head-to-head:
  • Direct perception won 85% of head-to-head judgments, outperforming the blind variant on all 7 evaluation dimensions: motivation/significance, hypothesis clarity, method appropriateness, result support, conclusion soundness, writing quality, and overall impression.
  • Interpretation

    The author offers a restaurant-review analogy: a traditional AI reads reviews and star ratings without ever tasting the food, while OmniScientist walks in, watches the chef, listens, and tastes. The shift is framed as a third coming-of-age for AI — from symbolic to connectionist, from unimodal to multimodal, and now from passive reception to active perception: a philosophical move from "text-centrism" to "perception-centrism."

    Limitations and outlook

  • A 6.3 average score means the system cannot yet replace human scientists; the blind variant still performs reasonably in some cases where precomputed features capture key information.
  • The envisioned future: AI scientists that physically operate lab equipment, adjust microscope focus, collect field samples, and synthesize materials — a genuine embodied agent, not just a language model on a server.
The post closes with an adaptation of Feynman's motto:

> "What I cannot create, I do not understand."

...augmented with: *"What I cannot perceive, I cannot discover."* Perception is the starting point of discovery.

Reference

Li, B., Fei, H., Ju, T., Lee, M.-L., & Hsu, W. (2026). *OmniScientist: An Omni-Modal Omni-Discipline AI Scientist*. arXiv preprint arXiv:2608.13558.

Tags

#ai-scientist#multimodal-ai#autonomous-agents#scientific-discovery#perception#omniscientist#machine-learning#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633483