Overview
SciAtlas weaves 43 million papers into a hyper-scale academic knowledge graph—157 million entities and 3 billion relationship edges—claiming to be the largest and broadest-coverage open scholarly knowledge graph. Its core thesis: replace fuzzy semantic similarity with structured knowledge topology, and uncertain model reasoning with deterministic graph propagation.
The Problem
- Keyword matching cannot understand that "protein structure prediction" points to "AlphaFold"; vector search captures similarity but cannot answer "which paper's method inspired which improvement."
- Discipline barriers create knowledge silos that block cross-disciplinary integration.
- Agentic deep-research frameworks are costly (repeated LLM calls) and prone to compounding hallucinations.
- English-only filtering (trade-off: systematic bias in medicine/social sciences)
- Abstract length filtering (quality investment for downstream LLM extraction)
- PDF availability filtering (traceability to source documents)
- No author deduplication (name ambiguity too costly; wrong merges worse than duplicates)
- Literature review: tunable venue/author/institution weights for different review depths.
- Idea grounding: segment candidate papers, extract fine-grained queries, compare idea vs. passages to detect prior work, supporting evidence, or genuine novelty.
- Also: trend prediction, collaborator mining, academic trajectory exploration.
- Missing quantitative evaluation (baselines, NDCG, ablations) in the accessible text.
- English-centric bias outside CS.
- Author name duplication (especially East Asian names like "Wang Wei").
- RWR propagation remains a black box despite score decomposition.
- Keyword quality cascades from Qwen3-30B-A3B's domain understanding.
- SciAtlas technical report (arXiv:2605.22878)
- Project page: http://scigraph.openkg.cn/
- OpenAlex, Neo4j, bge-large-en-v1.5, Qwen3-30B-A3B-Instruct-2507
Four-Layer Architecture
Nine entity types: Paper (43.3M), Author (109.7M), Keyword (3.76M), four-level direction hierarchy (4,520 Topics / 252 Subfields / 26 Fields / 4 Domains), Institution (120K), Source (280K).
Twelve relation types across layers: semantic (CITES 214M, RELATED_TO), conceptual (HAS_KEYWORD, COOCCUR), directional (DOMAIN_OF, FIELD_OF, SUBFIELD_OF), social (AUTHORED, COAUTHOR 2.06B, AFFILIATED_WITH), publication (PUBLISHED_IN).
The layered design is configurable: literature review leans on CITES/RELATED_TO; collaborator mining on COAUTHOR/AFFILIATED_WITH; trend forecasting on FIELD_OF/COOCCUR.
Data Construction
From 480M OpenAlex papers, four cleaning decisions:
The key innovation is LLM-based keyword extraction with Qwen3-30B-A3B-Instruct-2507: 3–8 reusable core phrases per paper, filtering paper-specific jargon and marketing-style names. Keywords receive importance scores as edge weights; co-occurrence builds COOCCUR edges. All titles/abstracts/keywords get precomputed bge-large-en-v1.5 embeddings, stored in Neo4j. Updates: daily via OpenAlex API, bi-monthly batch files, GROBID for missing metadata.
Neuro-Symbolic Three-Path Retrieval
1. Keyword matching: LLM extracts keywords with importance scores; exact match (full score) or vector match (threshold \(θ_kw=0.7\), top-3). 2. Semantic matching: query embedded; title and abstract channels each retrieve top-60, reranked by bge-reranker-large to top-15 each, fused 0.4/0.6. 3. Title matching (when a reference paper is given): GROBID extraction, top-10 titles, exact (1.0) or fuzzy matching (0.65×LCS + 0.35×Jaccard, threshold 0.88).
Merged seeds get a title-hit bonus (+0.35 exact, +0.10 fuzzy) and citation-weighted initialization. A 2-hop subgraph constraint (≤500 nodes per type per hop) bounds the graph, then Random Walk with Restart runs with edge weights (CITES 1.00, RELATED_TO 0.90, AUTHORED 0.80, COAUTHOR/COOCCUR 0.60), convergence at \(ε=10^{-6}\) or 50 iterations.
Final ranking: \(s_p^{final} = min(1, 0.35·s̃_p^{pre} + 0.45·s̃_p^{graph}·g_p + 0.20·imp(p))\), with a graph-support gate \(g_p = max(0.25, s̃_p^{pre})\). Philosophy: topological signals outweigh isolated semantic features. Every result ships with a score breakdown and path-based explanation; total retrieval takes "significantly under 2 minutes."
Applications
Critical Assessment
Strengths: leading open scale; pragmatic quality/coverage trade-offs; cost-effective LLM choice.
Open tensions:
Strategic Significance
SciAtlas marks a shift from metadata management to cognitive topology construction. Paired with LLM assistants (LLMs for reasoning, SciAtlas for traceable fact retrieval and relation verification), it offers a structured substrate for AI for Science and cross-disciplinary connection discovery—e.g., surfacing GNN applications in molecular representation and materials discovery. It is a navigable cognitive map, not autonomous science: hypothesis generation, experiment design, and analysis still require humans or stronger AI systems.