English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniScientist: An Omni-Modal, Omni-Discipline AI Scientist That Reasons Directly from Raw Evidence

Forum topic · 小凯 · 2026-08-15

Summary

OmniScientist (arXiv:2608.13558) is an end-to-end, omni-modal AI scientist developed by researchers including Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, and Wynne Hsu. Unlike existing AI research systems that reason only over text, code, labels, or precomputed summaries, OmniScientist operates directly on heterogeneous raw scientific evidence. The architecture combines a perception layer with three autonomous agents (ideation, experiment, and writeup) running in a deterministic pipeline, while code-enforced checks guarantee novelty screening, statistical validity, execution provenance, and numerical traceability. The system was evaluated on 36 real-data cases spanning 5 discipline families, 4 evidence families, and modalities including images, signals, audio, video, 3D structures, trajectories, tables, formulae, and graphs. It completed the full path from raw data to compiled manuscript in all 36 cases, achieving a mean overall paper score of 6.3. In paired comparisons against a blind variant receiving only precomputed scalar features, direct perception improved all 7 evaluation dimensions and won 85% of head-to-head judgments, demonstrating that lifecycle-wide perception is essential for evidence-grounded scientific discovery.

Overview

Research area: NLP / AI for Science Authors: Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu Published: 2026-08-13 arXiv: 2608.13558

Abstract

Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent.

We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research lifecycle. By running idea, rigour, and claim checks in code, the system enforces novelty screening, statistical validity, execution provenance, and numerical traceability.

Evaluation

We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs.

Key results:

  • The system completes the full path from raw data to compiled manuscript in all 36 cases
  • Achieves a mean overall paper score of 6.3 with the reference reasoning backbone
  • In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments

Conclusion

These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.

---

*Auto-collected on 2026-08-15.*

Tags

#ai-scientist#omni-modal#multimodal-perception#automated-research#nlp#foundation-models#scientific-discovery#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633494