English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Benchmarks

Forum topic · 小凯 · 2026-09-15

Summary

Researchers Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy, and Navin Goyal introduce WinSyn, an automated pipeline (arXiv:2609.12171) for generating synthetic enterprise datasets to benchmark question-answering agents that rely on Retrieval-Augmented Generation and Deep Research. Enterprise data is challenging because information is scattered across evolving, potentially conflicting emails, chats, and documents, while existing benchmarks suffer from limited real-world complexity, short-form answers, and unnatural queries. WinSyn simulates enterprise projects spanning several months and involving up to 25 employees in distinct roles, producing realistic emails together with long- and short-form questions and gold answers grounded in the data, with an emphasis on ambiguity, scattered information, and naturally arising queries. Evaluating standard agent baselines built on frontier models, the authors found average total scores below 80% on all queries across every dataset, indicating substantial room for improvement and highlighting the importance of realistic, high-complexity evaluation data for building stronger enterprise Deep Research systems.

Overview

Research area: Machine Learning Authors: Amey Varhade, Ananya Sutradhar, Ravishankar Krishnaswamy, Navin Goyal Published: 2026-09-15 arXiv: 2609.12171

What the paper does

Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation (RAG), Deep Research (DR), and related techniques. Much of the difficulty comes from enterprise data complexity: information is often spread across evolving and potentially conflicting emails, chat messages, documents, and other artifacts. Existing benchmarks typically offer limited real-world complexity, short-form responses, and unnatural queries, so they fail to capture the real challenges of enterprise settings.

WinSyn is an automated pipeline that generates synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data.

Key points

  • Simulates enterprise projects spanning several months, involving up to 25 employees in distinct roles.
  • Data emphasizes ambiguity, scattered information, and naturally arising queries.
  • Produces both long- and short-form questions with data-grounded gold answers.
  • Evaluation of standard agent baselines using recent frontier models shows average total scores below 80% for all queries on every dataset.

Takeaway

There remains substantial room for improvement in enterprise agent deployment, and realistic, high-complexity evaluation data is critical for developing stronger enterprise-grade Deep Research systems.

Tags

#machine-learning#question-answering#retrieval-augmented-generation#deep-research#synthetic-data#benchmark#enterprise-ai#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634837