English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

SciAtlas: Mapping 43 Million Papers into a Scientific Knowledge Graph

Forum topic · 小凯 · 2026-05-25

Summary

SciAtlas is a large-scale open academic knowledge graph that integrates 43 million English papers from OpenAlex into 157 million entities and 3 billion relationship edges, aiming to overcome the 'knowledge silo' problem in academic search. It organizes knowledge across four layers—semantic (citations, relatedness), conceptual (LLM-extracted keywords and co-occurrence), directional (four-level topic/field hierarchy), and social (authors, institutions)—stored in Neo4j. Keywords are extracted by a lightweight open-source LLM (Qwen3-30B-A3B) and semantic embeddings are precomputed with bge-large-en-v1.5. Retrieval follows a neuro-symbolic three-path scheme—keyword, semantic, and title matching—followed by Random Walk with Restart graph propagation and a fusion ranking that weighs topological support (0.45) above initial semantic relevance (0.35) and citation importance (0.20). Designed downstream uses include literature review, research idea grounding, trend prediction, and collaborator discovery. Noted limitations include English-only filtering, no author disambiguation, and missing quantitative evaluation.

Overview

SciAtlas weaves 43 million papers into a hyper-scale academic knowledge graph—157 million entities and 3 billion relationship edges—claiming to be the largest and broadest-coverage open scholarly knowledge graph. Its core thesis: replace fuzzy semantic similarity with structured knowledge topology, and uncertain model reasoning with deterministic graph propagation.

The Problem

  • Keyword matching cannot understand that "protein structure prediction" points to "AlphaFold"; vector search captures similarity but cannot answer "which paper's method inspired which improvement."
  • Discipline barriers create knowledge silos that block cross-disciplinary integration.
  • Agentic deep-research frameworks are costly (repeated LLM calls) and prone to compounding hallucinations.
  • Four-Layer Architecture

    Nine entity types: Paper (43.3M), Author (109.7M), Keyword (3.76M), four-level direction hierarchy (4,520 Topics / 252 Subfields / 26 Fields / 4 Domains), Institution (120K), Source (280K).

    Twelve relation types across layers: semantic (CITES 214M, RELATED_TO), conceptual (HAS_KEYWORD, COOCCUR), directional (DOMAIN_OF, FIELD_OF, SUBFIELD_OF), social (AUTHORED, COAUTHOR 2.06B, AFFILIATED_WITH), publication (PUBLISHED_IN).

    The layered design is configurable: literature review leans on CITES/RELATED_TO; collaborator mining on COAUTHOR/AFFILIATED_WITH; trend forecasting on FIELD_OF/COOCCUR.

    Data Construction

    From 480M OpenAlex papers, four cleaning decisions:

  • English-only filtering (trade-off: systematic bias in medicine/social sciences)
  • Abstract length filtering (quality investment for downstream LLM extraction)
  • PDF availability filtering (traceability to source documents)
  • No author deduplication (name ambiguity too costly; wrong merges worse than duplicates)
  • The key innovation is LLM-based keyword extraction with Qwen3-30B-A3B-Instruct-2507: 3–8 reusable core phrases per paper, filtering paper-specific jargon and marketing-style names. Keywords receive importance scores as edge weights; co-occurrence builds COOCCUR edges. All titles/abstracts/keywords get precomputed bge-large-en-v1.5 embeddings, stored in Neo4j. Updates: daily via OpenAlex API, bi-monthly batch files, GROBID for missing metadata.

    Neuro-Symbolic Three-Path Retrieval

    1. Keyword matching: LLM extracts keywords with importance scores; exact match (full score) or vector match (threshold \(θ_kw=0.7\), top-3). 2. Semantic matching: query embedded; title and abstract channels each retrieve top-60, reranked by bge-reranker-large to top-15 each, fused 0.4/0.6. 3. Title matching (when a reference paper is given): GROBID extraction, top-10 titles, exact (1.0) or fuzzy matching (0.65×LCS + 0.35×Jaccard, threshold 0.88).

    Merged seeds get a title-hit bonus (+0.35 exact, +0.10 fuzzy) and citation-weighted initialization. A 2-hop subgraph constraint (≤500 nodes per type per hop) bounds the graph, then Random Walk with Restart runs with edge weights (CITES 1.00, RELATED_TO 0.90, AUTHORED 0.80, COAUTHOR/COOCCUR 0.60), convergence at \(ε=10^{-6}\) or 50 iterations.

    Final ranking: \(s_p^{final} = min(1, 0.35·s̃_p^{pre} + 0.45·s̃_p^{graph}·g_p + 0.20·imp(p))\), with a graph-support gate \(g_p = max(0.25, s̃_p^{pre})\). Philosophy: topological signals outweigh isolated semantic features. Every result ships with a score breakdown and path-based explanation; total retrieval takes "significantly under 2 minutes."

    Applications

  • Literature review: tunable venue/author/institution weights for different review depths.
  • Idea grounding: segment candidate papers, extract fine-grained queries, compare idea vs. passages to detect prior work, supporting evidence, or genuine novelty.
  • Also: trend prediction, collaborator mining, academic trajectory exploration.
  • Critical Assessment

    Strengths: leading open scale; pragmatic quality/coverage trade-offs; cost-effective LLM choice.

    Open tensions:

  • Missing quantitative evaluation (baselines, NDCG, ablations) in the accessible text.
  • English-centric bias outside CS.
  • Author name duplication (especially East Asian names like "Wang Wei").
  • RWR propagation remains a black box despite score decomposition.
  • Keyword quality cascades from Qwen3-30B-A3B's domain understanding.
  • Strategic Significance

    SciAtlas marks a shift from metadata management to cognitive topology construction. Paired with LLM assistants (LLMs for reasoning, SciAtlas for traceable fact retrieval and relation verification), it offers a structured substrate for AI for Science and cross-disciplinary connection discovery—e.g., surfacing GNN applications in molecular representation and materials discovery. It is a navigable cognitive map, not autonomous science: hypothesis generation, experiment design, and analysis still require humans or stronger AI systems.

    References

  • SciAtlas technical report (arXiv:2605.22878)
  • Project page: http://scigraph.openkg.cn/
  • OpenAlex, Neo4j, bge-large-en-v1.5, Qwen3-30B-A3B-Instruct-2507

Tags

#sciatalas#knowledge-graph#academic-search#ai-for-science#graph-algorithms#large-language-models#neurosymbolic-ai#openalex

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620802