TREC iKAT 2023: A Test Collection for Evaluating Conversational and Interactive Knowledge Assistants (SIGIR 2024)
Overview
TREC iKAT 2023 is a test collection developed for the Text REtrieval Conference (TREC) to enable systematic evaluation of conversational and interactive knowledge assistants. The accompanying resource paper was published at SIGIR 2024.
- Resource type: Academic paper / evaluation resource
- Venue: SIGIR 2024
- Link: https://dl.acm.org/doi/abs/10.1145/3626772.3657860
- Topic area: Evaluation of search engines; conversational and interactive information access
- User intent unfolds over multiple turns, requiring context management and clarification.
- Systems must combine retrieval, ranking, and generation, often over external knowledge sources.
- Static relevance judgments and single-query benchmarks (e.g., classic nDCG-style test collections) do not capture interactive behavior.
- Introduces the TREC iKAT 2023 test collection for evaluating conversational and interactive knowledge assistants.
- Positioned within the broader shift toward LLM-based conversational search and retrieval-augmented generation (RAG).
- Provides community infrastructure (topics/tasks and judged data) to support comparable, reproducible evaluation of assistant systems.
- Published as a SIGIR 2024 resource paper documenting the collection for the research community.
- Evaluation of Retrieval-Augmented Generation: A Survey (arXiv 2405.07437)
- ARES: An Automated Evaluation Framework for RAG (arXiv 2311.09476)
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents (arXiv 2401.13178)
Why It Matters
Conversational search assistants differ fundamentally from traditional ad-hoc search systems:
The iKAT (Interactive Knowledge Assistant Track) addresses this gap by providing a standardized, reproducible test collection so that knowledge assistant systems can be compared under a common evaluation protocol.
Key Points
> Note: For exact task descriptions, dataset statistics, and quantitative results, consult the original paper at the ACM Digital Library link above. This page summarizes the resource based on its title, venue, and public metadata.
Context in Conversational IR Research
Evaluation methodology has become a central concern as systems move from single-shot retrieval to interactive, agentic pipelines. TREC iKAT joins related efforts—such as RAG evaluation frameworks, multi-turn agent benchmarks, and citation-quality studies—reflecting a community trend toward process-level metrics (task success, multi-turn consistency, answer grounding) rather than static ranking metrics alone.
Related Reading
Citation
Original source: TREC iKAT 2023: A Test Collection for Evaluating Conversational and Interactive Knowledge Assistants, SIGIR 2024 — https://dl.acm.org/doi/abs/10.1145/3626772.3657860