Self-Graph Reasoning (SGR): From Chains to Graphs for LLM Reasoning
This post is an in-depth analysis of the paper "From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs" (arXiv:2601.03597), from researchers at the University of Tokyo, Texas A&M, Cambridge, Yale, Xiaomi Auto, and Henan University.
Key Points
- SGR (Self-Graph Reasoning) lets LLMs build their own reasoning graphs (nodes + edges) instead of following linear Chain-of-Thought (CoT), enabling branching and merging of reasoning paths.
- It targets "logical drift": in linear CoT, one small early error can cascade into a wrong conclusion, or a correct answer can be reached via flawed reasoning.
- Training pipeline: GPT-4o as teacher samples k candidate reasoning trajectories per question at high temperature (τ=0.9); graphs that correctly derive their answer are merged and cleaned, yielding ~10K high-quality training samples serialized as structured XML-style templates.
- SGR-LLaMA-3.3-70B (fine-tuned with LoRA) improves over the base model by 17.74% and nearly matches GPT-4o on average across five benchmarks: 61.03% vs 61.52%.
- On the Alice in Wonderland (AIW) test — e.g., "Alice has 3 brothers and 2 sisters; how many sisters do Alice's brothers have?" (answer: 3, requiring a perspective shift) — SGR reaches 57.50% vs GPT-4o's 32.50% (+25 points).
- On LogiQA, SGR's cost is about $33.6, far below GPT-4o CoT's ~$80.
- Code: https://github.com/Yingjian-Chen/SGR-Self-Graph-Reasoning
- The training set is only ~10K samples; scaling may help.
- Experiments used a 70B model; 8B models showed limited benefit, suggesting a base-capability threshold.
- Paper: https://arxiv.org/abs/2601.03597
- AIW benchmark background: https://arxiv.org/abs/2406.02061
Benchmark Results
| Benchmark | GPT-4o | LLaMA-3.3-70B | SGR-LLaMA-3.3-70B | |---|---|---|---| | LogiQA | 74.01% | 64.01% | 69.91% | | AIW | 32.50% | 19.50% | 57.50% | | AR-LSAT | 31.75% | 31.30% | 31.74% | | MedQA | 88.29% | 63.55% | 78.81% | | MathQA | 81.05% | 38.09% | 67.17% | | Average | 61.52% | 43.29% | 61.03% |
Additional findings: SGR beats the external-graph method RwG by 18.76%, generalizes to professional domains (medicine, math) with ~22.17% average gains, and reduces deployment cost to roughly 42% of GPT-4o CoT.
Why It Works
1. Many-to-one dependencies: graphs let multiple premises explicitly converge into one conclusion. 2. Explicit parent-child dependencies: every node must be justified by its parents, coupling reasoning with the final answer. 3. Non-linear exploration: parallel branches allow hypothesis switching, backtracking, and aggregation. 4. Traceability: the transparent graph structure makes error localization easy — moving AI from "black box" to "glass box."
In a case study on "Alice has 4 brothers and 3 sisters" (answer: 4), SGR generated a 9-node, 8-edge graph with dedicated branch nodes, an explicit perspective-shift node (+Alice herself), and an aggregation node.
Limitations
Resources
Discussion
1. Will graph-structured reasoning become a standard feature of future LLMs? 2. What other applications could benefit from SGR-style graph reasoning? 3. How can reasoning capability be preserved while reducing compute cost?
*Based on a deep-dive analysis of the paper "From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs."*