English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

General365: A Benchmark for General Reasoning in Large Language Models Across Diverse and Challenging Tasks

Forum topic · 小凯 · 2026-04-15

Summary

General365 is a benchmark designed to evaluate general reasoning in large language models (LLMs), a capability that remains under-explored compared to domain-specific reasoning in areas like mathematics and physics. Unlike specialized reasoning, general reasoning relies less on expert knowledge but still involves difficult challenges such as complex constraints, nested logical branches, and semantic distractors. To decouple reasoning ability from professional knowledge, General365 restricts required background knowledge to the K-12 level. The benchmark comprises 365 seed questions and 1,095 variant questions spanning 8 categories. The authors evaluate 26 leading LLMs on the benchmark and find that even the best-performing model achieves only 62.8% accuracy, indicating substantial room for improvement in general reasoning. The paper (arXiv:2604.11778) is authored by Junlin Liu, Shengnan An, and colleagues, including Xunliang Cai, and was released in April 2026.

Overview

Contemporary large language models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in specialized domains like mathematics and physics. However, their ability to generalize these reasoning skills to broader contexts—often termed general reasoning—remains under-explored.

Paper Details

  • Research areas: cs.CL, cs.AI
  • Authors: Junlin Liu, Shengnan An, Shuang Zhou, Dan Ma, Shixiong Luo, Ying Xie, Yuan Zhang, Wenling Yuan, Yifan Zhou, Xiaoyu Li, Ziwen Wang, Xuezhi Cao, Xunliang Cai
  • Published: 2026-04-13
  • arXiv: 2604.11778
  • Key Points

  • Problem: Unlike domain-specific reasoning, general reasoning relies less on expert knowledge but still poses formidable challenges, including:
  • Complex constraints
  • Nested logical branches
  • Semantic distractors (distracting information)
  • Benchmark design: The paper introduces General365, a benchmark specifically designed to assess LLM general reasoning. By limiting required background knowledge to the K-12 level, it explicitly decouples reasoning ability from professional knowledge.
  • Scale: The benchmark contains 365 seed questions and 1,095 variant questions, covering 8 categories.
  • Findings: Evaluation of 26 leading LLMs shows that even the top-performing model achieves only 62.8% accuracy, highlighting that general reasoning remains a significant unsolved challenge.

Original Abstract (excerpt)

> Contemporary large language models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in specialized domains like mathematics and physics. However, their ability to generalize these reasoning skills to more general and broader contexts--often termed general reasoning--remains under-explored.

---

*Auto-collected on 2026-04-15.*

Tags

#llm#benchmark#general-reasoning#arxiv#natural-language-processing#ai-evaluation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618486