[论文] OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
论文概要
研究领域: NLP 作者: Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu 发布时间: 2026-08-13 arXiv: 2608.13558中文摘要
基础模型的最新进展使AI科学家能够自动化日益完整的研究工作流,从假设生成、代码执行到手稿准备。然而,仅覆盖工作流并不能让系统获取科学发现所依赖的全部证据。现有系统通常基于文本、代码、标签或预计算摘要进行推理,导致科学上决定性的空间、时间、跨通道和过程关系对智能体不可见。我们提出OmniScientist,一个端到端的全模态AI科学家,直接从异构原始证据开展多学科研究。感知层和三个自主智能体(构思、实验、撰写)在确定性管道中运行,使观察能够在整个研究生命周期中塑造研究问题、实验决策和最终结论。通过在代码中运行创意、严谨性和结论检查,系统强制执行新颖性筛选、统计有效性、执行溯源和数值可追溯性。我们在36个真实数据案例上评估OmniScientist,涵盖5个学科家族、4个科学证据家族和包括图像、信号、音频、视频、3D结构、轨迹、表格、公式和图表在内的多种模态。系统在所有36个案例中完成了从原始数据到编译手稿的完整路径,使用参考推理骨干获得平均总体论文分数6.3。在与仅接收预计算标量特征的盲变体进行配对比较时,直接感知改善了所有7个评估维度,并在85%的正面交锋判断中获胜。这些结果表明,全生命周期的感知对于基于证据的科学发现至关重要,并为广泛能力的AI科学家提供了实际路径。原文摘要
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execution to manuscript preparation. Yet workflow coverage alone does not provide access to the full evidence on which scientific discovery depends. Existing systems typically reason over text, code, labels, or precomputed summaries, leaving scientifically decisive spatial, temporal, cross-channel, and procedural relations unavailable to the agent. We introduce OmniScientist, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence. A perception layer and 3 autonomous agents for ideation, experiment, and writeup operate within a deterministic pipeline, allowing observations to shape research questions, experimental decisions, and final claims throughout the research生命周期. By running idea, rigour, and claim checks in code, the system enforces novelty筛选、统计validity、execution provenance, and numerical traceability. We evaluate OmniScientist on 36 real-data cases spanning 5 discipline families, 4 families of scientific evidence, and modalities including images, signals, audio, video, 3-D structures, trajectories, tables, formulae, and graphs. The system completes the full path from raw data to compiled manuscript in all 36 cases and achieves a mean overall paper score of 6.3 with the reference reasoning backbone. In paired comparisons against a blind variant that receives only precomputed scalar features, direct perception improves all 7 evaluation dimensions and wins 85% of head-to-head judgments. These results show that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.--- *自动采集于 2026-08-15*
#论文 #arXiv #NLP #小凯