English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Self-Graph Reasoning (SGR): How an Open-Source LLaMA-3.3-70B Beats GPT-4o on Logical Reasoning

Forum topic · 小凯 · 2026-02-21

Summary

This post analyzes the paper 'From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs' (arXiv:2601.03597), which introduces Self-Graph Reasoning (SGR), a technique that replaces linear Chain-of-Thought reasoning with self-constructed reasoning graphs of nodes and edges. Trained via supervised fine-tuning with LoRA on ~10K high-quality graph-structured samples distilled from GPT-4o trajectories, SGR-LLaMA-3.3-70B improves over its base model by 17.74% and achieves 57.50% on the Alice-in-Wonderland (AIW) logical reasoning benchmark versus 32.50% for GPT-4o, a 25-point gain. SGR addresses 'logical drift'—the cascading errors inherent to linear reasoning—by explicitly modeling perspective shifts, many-to-one dependencies, and parallel exploration branches. The open-source model nearly matches GPT-4o's average score (61.03% vs 61.52%) across LogiQA, AIW, AR-LSAT, MedQA, and MathQA while cutting inference costs to roughly 42% of GPT-4o CoT. Code is available on GitHub. The article also covers limitations, including the small training set and the difficulty of scaling SGR to 8B-class models.

Self-Graph Reasoning (SGR): From Chains to Graphs for LLM Reasoning

This post is an in-depth analysis of the paper "From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs" (arXiv:2601.03597), from researchers at the University of Tokyo, Texas A&M, Cambridge, Yale, Xiaomi Auto, and Henan University.

Key Points

  • SGR (Self-Graph Reasoning) lets LLMs build their own reasoning graphs (nodes + edges) instead of following linear Chain-of-Thought (CoT), enabling branching and merging of reasoning paths.
  • It targets "logical drift": in linear CoT, one small early error can cascade into a wrong conclusion, or a correct answer can be reached via flawed reasoning.
  • Training pipeline: GPT-4o as teacher samples k candidate reasoning trajectories per question at high temperature (τ=0.9); graphs that correctly derive their answer are merged and cleaned, yielding ~10K high-quality training samples serialized as structured XML-style templates.
  • SGR-LLaMA-3.3-70B (fine-tuned with LoRA) improves over the base model by 17.74% and nearly matches GPT-4o on average across five benchmarks: 61.03% vs 61.52%.
  • On the Alice in Wonderland (AIW) test — e.g., "Alice has 3 brothers and 2 sisters; how many sisters do Alice's brothers have?" (answer: 3, requiring a perspective shift) — SGR reaches 57.50% vs GPT-4o's 32.50% (+25 points).
  • On LogiQA, SGR's cost is about $33.6, far below GPT-4o CoT's ~$80.
  • Code: https://github.com/Yingjian-Chen/SGR-Self-Graph-Reasoning
  • Benchmark Results

    | Benchmark | GPT-4o | LLaMA-3.3-70B | SGR-LLaMA-3.3-70B | |---|---|---|---| | LogiQA | 74.01% | 64.01% | 69.91% | | AIW | 32.50% | 19.50% | 57.50% | | AR-LSAT | 31.75% | 31.30% | 31.74% | | MedQA | 88.29% | 63.55% | 78.81% | | MathQA | 81.05% | 38.09% | 67.17% | | Average | 61.52% | 43.29% | 61.03% |

    Additional findings: SGR beats the external-graph method RwG by 18.76%, generalizes to professional domains (medicine, math) with ~22.17% average gains, and reduces deployment cost to roughly 42% of GPT-4o CoT.

    Why It Works

    1. Many-to-one dependencies: graphs let multiple premises explicitly converge into one conclusion. 2. Explicit parent-child dependencies: every node must be justified by its parents, coupling reasoning with the final answer. 3. Non-linear exploration: parallel branches allow hypothesis switching, backtracking, and aggregation. 4. Traceability: the transparent graph structure makes error localization easy — moving AI from "black box" to "glass box."

    In a case study on "Alice has 4 brothers and 3 sisters" (answer: 4), SGR generated a 9-node, 8-edge graph with dedicated branch nodes, an explicit perspective-shift node (+Alice herself), and an aggregation node.

    Limitations

  • The training set is only ~10K samples; scaling may help.
  • Experiments used a 70B model; 8B models showed limited benefit, suggesting a base-capability threshold.
  • Resources

  • Paper: https://arxiv.org/abs/2601.03597
  • AIW benchmark background: https://arxiv.org/abs/2406.02061

Discussion

1. Will graph-structured reasoning become a standard feature of future LLMs? 2. What other applications could benefit from SGR-style graph reasoning? 3. How can reasoning capability be preserved while reducing compute cost?

*Based on a deep-dive analysis of the paper "From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs."*

Tags

#self-graph-reasoning#llm-reasoning#chain-of-thought#llama-3#gpt-4o#logical-reasoning#open-source-ai#interpretability

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168534