English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Frontier LLM Agents Break the Ontology Curation Bottleneck, Approaching Human Curator Performance

Forum topic · 小凯 · 2026-05-30

Summary

Phenotype annotation—linking free-text morphological descriptions to ontology terms—is a labor-intensive bottleneck in comparative morphology, historically requiring highly trained human experts. A new study by James P. Balhoff and Hilmar Lapp (arXiv:2605.28965) revisits the Dahdul et al. (2018) gold-standard benchmark, which contains EQ annotations across seven phylogenetic studies and was previously used to evaluate three human curators and the Semantic CharaParser NLP tool. The authors evaluated five frontier hosted LLMs from Anthropic and OpenAI, each operating as an "agentic curator" in a self-contained workspace equipped with the source publication PDFs, original annotation guidelines, four project ontologies (UBERON, PATO, BSPO, GO), and validation scripts. Every agent fell within the range of inter-curator variability of the three trained human curators; the best-performing agents approached but did not surpass the best human curator. All agents substantially outperformed Semantic CharaParser across all four evaluation metrics, demonstrating that frontier LLM agents can materially scale phenotype curation.

Paper Overview

  • Field: Agents / Bioinformatics
  • Authors: James P. Balhoff, Hilmar Lapp
  • Published: 2026-05-30
  • arXiv: 2605.28965
  • Abstract

    Linking free-text phenotype descriptions to ontology terms is essential for cross-study integration of comparative morphological data. This labor-intensive process has heavily relied on human experts, making it a key scalability bottleneck. Dahdul et al. (2018) established a gold standard of EQ annotations across seven phylogenetic studies and used it to evaluate three human curators and the Semantic CharaParser NLP tool.

    In this study, the authors revisit that benchmark with five frontier hosted LLMs from Anthropic and OpenAI. Each model operates as an "agentic curator" within a self-contained workspace, providing the source publication PDF, the original annotation guidelines, four project ontologies (UBERON, PATO, BSPO, GO), and validation scripts.

    Key Findings

  • Evaluated against the same gold standard, every agent fell within the range of inter-curator variability observed among the three trained human curators.
  • The best-performing agents approached but did not reach the performance of the best human curator.
  • Agents substantially outperformed Semantic CharaParser on all four evaluation metrics.

Original Abstract

> Linking free-text phenotype descriptions to ontology terms is essential for cross-study integration of comparative morphological data. This labor intensive process has heavily relied on human experts. Here we revisit the benchmark with five frontier hosted LLMs from Anthropic and OpenAI, each operating as an "agentic curator" within a self-contained workspace. Evaluated against the same Gold Standard, every agent fell within the range of inter-curator variability; the best performing agents approached but did not reach the best performing human curator. Agents substantially outperformed Semantic CharaParser on all four metrics.

---

*Auto-collected on 2026-05-30*

Tags

#llm-agents#ontology-curation#bioinformatics#phenotype-annotation#arxiv#nlp#benchmark

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980564