Paper Overview
- Field: Agents / Bioinformatics
- Authors: James P. Balhoff, Hilmar Lapp
- Published: 2026-05-30
- arXiv: 2605.28965
- Evaluated against the same gold standard, every agent fell within the range of inter-curator variability observed among the three trained human curators.
- The best-performing agents approached but did not reach the performance of the best human curator.
- Agents substantially outperformed Semantic CharaParser on all four evaluation metrics.
Abstract
Linking free-text phenotype descriptions to ontology terms is essential for cross-study integration of comparative morphological data. This labor-intensive process has heavily relied on human experts, making it a key scalability bottleneck. Dahdul et al. (2018) established a gold standard of EQ annotations across seven phylogenetic studies and used it to evaluate three human curators and the Semantic CharaParser NLP tool.
In this study, the authors revisit that benchmark with five frontier hosted LLMs from Anthropic and OpenAI. Each model operates as an "agentic curator" within a self-contained workspace, providing the source publication PDF, the original annotation guidelines, four project ontologies (UBERON, PATO, BSPO, GO), and validation scripts.
Key Findings
Original Abstract
> Linking free-text phenotype descriptions to ontology terms is essential for cross-study integration of comparative morphological data. This labor intensive process has heavily relied on human experts. Here we revisit the benchmark with five frontier hosted LLMs from Anthropic and OpenAI, each operating as an "agentic curator" within a self-contained workspace. Evaluated against the same Gold Standard, every agent fell within the range of inter-curator variability; the best performing agents approached but did not reach the best performing human curator. Agents substantially outperformed Semantic CharaParser on all four metrics.
---
*Auto-collected on 2026-05-30*