English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

In-Depth Survey of Open Scholarly Paper Knowledge Graphs on the Internet

Forum topic · ✨步子哥 · 2026-07-11

Summary

This report from zhichai.net surveys 22 open Chinese and English scholarly data resources for building a paper-writing-assistant website. The core recommendation is to use OpenAlex (CC0, ~477 million works, REST API plus quarterly snapshots) as the primary backbone, supplemented by Crossref (176 million DOIs, CC0) for canonical metadata and OpenCitations/COCI (2.2 billion citation edges, CC0) for the citation graph. Full-text mining can draw on S2ORC, arXiv, and PubMed/PMC, while Semantic Scholar (S2AG) offers TLDRs and citation recommendations but carries a CC-BY-NC non-commercial restriction for bulk data. Papers with Code was shut down by Meta in July 2025 and was revived in 2026 via paperswithcode.co. For Chinese-language resources, the report finds a significant gap: NCPSSD (13.5 million Chinese social science full texts) and AMiner/OAG are the main usable open sources, whereas CNKI, Wanfang, VIP, and CSCD remain paywalled and closed to bulk access. The report concludes that a Chinese open scholarly knowledge graph effectively does not exist and recommends combining NCPSSD full texts, OAG citations, and OpenKG schemas, with Chinese data as a limited enhancement rather than the core dependency.

In-Depth Survey of Open Scholarly Paper Knowledge Graphs on the Internet

*A research report from zhichai.net on data resource selection for a paper-writing-assistant website. Survey date: 2026-07-10; method: four parallel research agents with cross-verification of key claims; coverage: 22 open Chinese and English resources.*

Executive summary

  • Primary backbone: OpenAlex (CC0, ~477 million works, 116 million authors, REST API + quarterly snapshots, actively maintained) — the strongest and most license-safe option for a literature index and citation network.
  • DOI and citation backbone: Crossref (CC0, 176 million DOIs) + OpenCitations/COCI (CC0, 2.2 billion citation edges).
  • Full text and semantics: S2ORC, arXiv, PubMed/PMC for full-text mining; Semantic Scholar (S2AG) offers TLDRs and citation recommendations, but bulk data is non-commercial only (CC-BY-NC).
  • Methods/tasks/datasets: Papers with Code was shut down by Meta in July 2025; its historical data is frozen in a GitHub archive, and in 2026 it was revived as paperswithcode.co rebuilt with AI agents.
  • Chinese resources are the weak point: genuinely open, bulk-accessible Chinese resources are scarce. NCPSSD (13.5 million Chinese social science full texts) and AMiner/OAG are usable; CNKI, Wanfang, VIP, and CSCD are paywalled or closed.
  • Biggest risk: scarce and closed licensing of Chinese open paper data, plus full-text copyright and citation lag. Build on compliant sources and treat Chinese data as a limited enhancement, not a core dependency.

Background

Core capabilities of a paper-writing site — literature search, citation suggestions, research-gap discovery, concept graphs, review generation — all depend on structured, linkable scholarly data. Knowledge graphs that explicitly model papers, authors, institutions, venues, concepts, datasets, and methods are especially suited. The report inventories currently available open paper knowledge graphs (Chinese and English), compares them on coverage, graph structure, access, licensing, Chinese coverage, maintenance activity, and fit for writing assistance.

Key resources (English)

| Resource | Highlights | License | Notes | |---|---|---|---| | OpenAlex ★ | ~477M works; entities: works/authors/sources/institutions/topics/funders; REST API + quarterly snapshots; mailto politeness pool | CC0 | Run by OurResearch (nonprofit); active (Walden rewrite in late 2025) | | Semantic Scholar (S2AG) | 214M+ papers, 2.49B citations, SPECTER2 vectors; graph/v1 API ~1 req/s free | API CC0; bulk ODC-BY, non-commercial | Allen Institute for AI | | Microsoft Academic Graph | Discontinued 2021-12-31; superseded by OpenAlex | — | Do not depend on it | | Open Academic Graph (OAG) | v3.1 (2024-02): ~700M entities, 2B relations; bulk download (tens of GB) | ODC-BY | Tsinghua AMiner + Microsoft Research; research-oriented | | DBLP | ~8.1M records, 3.9M authors (CS focus); monthly XML dumps + SPARQL | CC0 | No citation network | | arXiv | 2.4M+ e-prints; OAI-PMH + Atom API | Metadata CC0 | Preprints, not peer-reviewed | | PubMed/PMC | 39M+ citations; 11M+ full texts; E-utilities + FTP | Public domain metadata | Biomedical essential | | Crossref | 176M+ DOI records; REST API + annual public data file (197 GB) | CC0 | DOI/canonical-reference backbone | | OpenCitations/COCI ★ | 2.2B+ citations (July 2025); REST + SPARQL + CSV/N-Triples, no key needed | CC0 | Excellent for citation analysis | | WikiCite/Wikidata | ~41M scholarly-article items; dedicated SPARQL endpoint since 2025-05 (~8B triples) | CC0 | Cross-language entity disambiguation | | Papers with Code | Shut down by Meta 2025-07-24/25; archive on GitHub; revived 2026 as paperswithcode.co | MIT/CC-BY | Task/dataset/metric tuples + leaderboards | | S2ORC | 81.1M paper nodes, 8.1M OA full texts structured, 467M+ citation edges | CC0 (limited) | Best for full-text mining | | Connected Papers | Commercial, built on Semantic Scholar; no open API | Restricted | Not suitable as a data source | | ORKG | Structured problem–method–result descriptions; REST + RDF | CC BY-SA | TIB (Germany) | | CiteSeerX | 10M+ documents, 100M+ citations (CS) | Open | Aging but unique historical CS full texts | | Closed counter-examples | Google Scholar (no official API, automation banned); ResearchGate (no open API); CNKI (commercial, no public API) | — | — |

Key Chinese resources

| Resource | Status | Notes | |---|---|---| | AMiner / OAG (Chinese) | Open download (ODC-BY), low Chinese-paper share | 331M papers, 135M scholars; strong Chinese author/institution coverage | | OpenKG.cn ★ | Truly open, downloadable (RDF/JSON/CSV); 337 datasets + 71 tools | Academic subgraphs (SciKG, CN-DBpedia, etc.) are mostly English arXiv or encyclopedic; useful as schema reference | | NCPSSD ★ | Truly open access after registration; 13.5M+ Chinese social science papers across 2,380 journals | Humanities/social sciences focus, no API; compliance must be assessed | | NSTL | Free search only; full text via limited document delivery | Not suitable as an open data base | | CSCD | Not open; institutional subscription (WoS version) | No bulk access | | Wanfang / VIP / CNKI | Not open: subscription/paywall, no public API, bulk scraping prohibited | Most complete Chinese coverage but walled gardens | | CAS GoOA | OA paper discovery, 1,700+ OA journals, free downloads | Its 400M-entity KG is internal only | | CN-DBpedia (Fudan) | 9M entities, 67M triples; dump downloadable under research license | Encyclopedic, not a paper KG |

Chinese summary: genuinely open: OAG (download, little Chinese content), OpenKG subgraphs, NCPSSD (best Chinese full text), GoOA, CN-DBpedia; search-only/closed: CSCD, Wanfang/VIP/CNKI, NSTL delivery, CAS internal KG. An open Chinese scholarly paper knowledge graph remains a gap. A combination of NCPSSD full texts + OAG citations + OpenKG schemas is recommended.

Comparison matrix (selected)

Fit for writing assistance is rated out of five stars: OpenAlex ★★★★★; Crossref ★★★★★; S2AG, arXiv, PubMed/PMC, OpenCitations, Papers with Code, NCPSSD, AMiner/OAG ★★★★☆; OAG, DBLP, WikiCite, S2ORC, ORKG, CSCD ★★★☆☆; Connected Papers, CiteSeerX, OpenKG.cn, NSTL ★★☆☆☆; MAG and Wanfang/VIP/CNKI ✕ (not usable as open data sources).

Conclusion and recommendation

For a paper-writing-assistant website, build the data foundation as:

1. Index and citation network: OpenAlex + Crossref + OpenCitations (all CC0, commercially safe). 2. Full text and semantics: S2ORC, arXiv, PubMed/PMC; use Semantic Scholar API within its CC0 API tier; respect the non-commercial restriction on S2AG bulk data. 3. Task/dataset linkage: Papers with Code archive and the revived paperswithcode.co, used with care. 4. Chinese enhancement: NCPSSD full texts + OAG Chinese author/institution data + OpenKG schemas — as a supplementary layer, not the core, given licensing and coverage limitations.

Main risks to manage: scarcity and closed licensing of Chinese open scholarly data, full-text copyright constraints, and citation-index lag.

Tags

#open-science#knowledge-graph#openalex#opencitations#crossref#semantic-scholar#chinese-academic-databases#scholarly-data

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346327