English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Anatomy of a 75-Page RSI Survey: Who Really Wrote It, the 72-Company Ledger, and 6 Days of US-China Hype

Forum topic · 二一 · 2026-09-16

Summary

A follow-up forensic analysis of the arXiv paper 2609.11873 ("The Last AI Built by Humans" RSI survey) examining authorship, the Appendix B industry taxonomy, and the six-day reception after publication. Key findings: the paper is actually driven by Xuanhe Zhou (corresponding author, SJTU assistant professor) and Fan Wu, with first author Yi Duan—a Renmin University economics undergraduate—coordinating a 33-author, 9-institution, 491-reference effort. Bo Wang (Zhou Bowen) appears both as an author and as a surveyed party: his company Frontis AI (Xianyuan) appears unstarred in the Appendix B table of 72 companies across six prototype categories (~142 product rows). Notable editorial choices include AlphaEvolve being downgraded to B0, MIT's SEAL being uncited, and OpenRSI being compressed into a single undifferentiated row. Reception showed a stark temperature gap: English-language X went viral on a "China vs US" framing while HN/Reddit largely ignored it; Chinese media (Synced, Huxiu, ifanr) split between popularization, technical critique, and framing skepticism. The author argues the survey's real product is "ledger power"—classification authority over the whole RSI ecosystem.

This post is a follow-up to a 2026-09-14 fact-check of the paper arXiv 2609.11873. The first post audited the paper's content (the five-level framework, HCI, three challenges, lineage mapping). This one audits the paper's "who's who": authorship structure, the Appendix B table of 72 companies, and the 6 days of Chinese and English discourse after release. Source chain: full arXiv HTML (~320k characters) re-scraped, PDF author block, and 5 sub-agent passes (framework close-read, second-half close-read, appendix line-by-line, community reaction, team backgrounds) cross-checked against ~60 sources. All claims grounded in the original text and first-party homepages.

Key points

1. Author list: 33 people, 9 institutions — the power structure is not the two famous names

  • Real drivers: Xuanhe Zhou (corresponding author) + Fan Wu (last author). Zhou, born 1997, is one of SJTU's youngest assistant professors and a PhD student of Tsinghua's Guoliang Li (also on the author list, position 27). Wu is a distinguished professor and executive vice dean at SJTU. "Theseus Labs" is not a company but the brand of this team's lab: a one-line website, a GitHub org (theseus-labs-rsi) with only an awesome-rsi literature repo (47 stars) and a homepage repo — no funding, no registration, no product code.
  • First author Yi Duan: a Renmin University economics undergraduate (class of 2023). His homepage lists him as "Project Lead, Theseus Labs" (2026.8–9), after a remote internship at Shanghai AI Lab (2025.9–2026.6, BioMatrix). An undergraduate coordinating a 75-page, 33-author, 491-reference survey is rare and fully public (duanyi516.github.io). All five co-first authors (Yi Duan / Ying Liu / Zirui Tang / Haodong Chen / Jun Zhou) appear on Zhou Xuanhe's student list.
  • Bo Wang's real role: author AND surveyed party. Wang (position 28, Tsinghua EE chair professor, director of Shanghai AI Lab) founded Xianyuan (Frontis AI) after JD. Appendix B contains the row: Frontis AI / Xianyuan | OpenRSI, Frontis-MA1 | AI4AI/self-improving agents | skills; memory; harness; weights | L2-L3. His student Kaiyan Zhang (position 24; OpenReview shows Frontis CTO) is the sole corresponding author of the Frontis-MA1 paper (arXiv 2607.28568). The teacher endorses the survey; the student's company product enters the industrial map — the "OpenRSI / survey mutual reference" loop closes: they are two edges of the same power topology.
  • Other seats: Zhiyuan Liu (position 29, THUNLP director) is co-founder/chief scientist of ModelBest (BAAI-adjacent, 面壁) — §5.4's ModelBest chapter (ForgeTrain matching Megatron-LM v0.15 in 8 hours, overtaking in 1.5–2.5 days; MFU 40.1→44.1%@0.5B) is his ecosystem's first-hand self-report. Two ByteDance authors map to the §5.2 Lark chapter; one Xiaohongshu; Zhoufutu Wen lists Humanlaya (§5.3) but his homepage shows ByteDance SEED since 2023.9 (paper self-report prevails). Conghui He (OpenDataLab/MinerU founder, Shanghai AI Lab) connects via Bo Wang's current post.
  • Conclusion: "SJTU + Bo Wang + Zhiyuan Liu" is the communication cover, not the power structure. A survey establishing an official taxonomy of China's RSI ecosystem, drafted by a 97-born AP with undergrad/master's students and endorsed by three heavyweight Tsinghua figures plus five companies, is itself part of China's "systematic RSI layout" (isomorphic to the ICLR 2026 RSI Workshop, backed by Tencent and BAAI).
  • 2. Appendix B: 72 companies, six prototype categories, ~142 product rows

    Organized by prototype, not geography:

    | Class | Name | Companies | Notable players (and grades) | |---|---|---|---| | A | Frontier model & general agent labs | 17 | OpenAI (Symphony/GPT-Red/AgentKit L1-L2; self-improving tax agents L4 candidate); Anthropic ("When AI Builds Itself" marked "L5 target" — vision, not reality); GDM (AlphaEvolve marked B0); Meta (HyperAgents L5 cand.); Qwen/DeepSeek/Tencent AI Lab (R-Zero L2-L3)/ByteDance Seed/MiniMax/Moonshot/Zhipu; Poolside/TML/Nous | | B | RSI-native / AI4AI companies | 14 | Weco (AIDE2 L5 cand.), Sakana (DGM L5 cand.), Recursive, Imbue; Chinese newcomers: MetaCircle, Frontis AI/Xianyuan, EvoMap, Endless Frontier, Chaoyan; Theseus itself (Argus L2-L3, SetupX L3) | | C | Autonomous R&D & scientific discovery | 14 | Periodic Labs (L3 cand.), FutureHouse, Axiom Math, Harmonic, Lila, Karpathy autoresearch (L2), Prime Intellect | | D | Agent optimization/eval/learning infra | 15 | Warp, Factory (one of the highest industrial grades: L4), LangChain, Replit, Braintrust, Mechanize, Goodfire | | E | Embodied/world models/continual adaptation | 8 | AI2 Robotics, X Square Robot, Galaxea, Synapx, MirrOS | | F | Persistent memory & personal AI | 4 | Engram, Lemon AI, EverMind, Mindverse |

    Four previously unflagged details:

  • AlphaEvolve is downgraded. It appears exactly once, marked B0 (task-level program search) under GDM in Table 12, with zero main-text discussion — versus ~10 mentions and two dedicated paragraphs for DGM. The paper deliberately separates "program search" from "improvement loop": AlphaEvolve has the former but no persistent self-modification, so it doesn't enter the RSI narrative. MIT's SEAL (Self-Adapting LMs) is not cited at all, not even the phrase "self-adapting." A taxonomy's power lies half in inclusion, half in silence.
  • Frontis AI/Xianyuan carries no asterisk. Entries like Chaoyan, Wuya Zhiyuan, Naive.ai carry "company identity needs stronger verification" markers; Frontis does not. Given a Frontis CTO's advisor sits on the author list, this "no verification needed" reads more like internal confidence than external due diligence.
  • ~10 confirmed Chinese companies of 72 (~12 more suspected), ~36–38 American — but all six §5 deep-dive cases (Theseus/Lark/Humanlaya/ModelBest/Tencent Hunyuan/Agent-Native) are Chinese practices not counted in Table 12. The ledger is global; the cases are Chinese.
  • The EvoMap mystery. One row: EvoMap | Evolver / GEP / GeneBench | shared code evolution | genes/capsules; code assets | generate; test; publish; reuse; adapt | L2-L3. Near-zero public information; treated as single-source, pending verification.
  • 3. Lineage additions (map dimension)

    1. HGM isn't only in the main text — it's used as RQGM's defeated baseline. §3.6.2 and §4.3.3: Red Queen Gödel Machine reports 71.7% pass@1 on held-out Polyglot coding vs HGM-H's 69.9%, with fewer search tokens. The narrative chain DGM (lineage archive) → HGM (lineage meta-selection) → RQGM (evaluator joins co-evolution) makes HGM the middle rung of the Gödel Machine family's three-stage evolution story. 2. OpenRSI and Frontis-MA1 share two identical rows, six fields verbatim. For a project whose stated program is "make improvement rate the optimization objective," being compressed into one undifferentiated L2-L3 label is itself a signal: when the official ledger absorbs you, it doesn't use your self-narrative. 3. Ten most-discussed systems (by mentions + dedicated paragraphs): DGM, VOYAGER, SEAgent, Metis, autoresearch, RQGM, AIDE2, Gödel Agent, Theseus, A-Evolve-Training. Lineage freshness (NeoHorse missing the cutoff by two days) and depth of absorption are two different ledgers.

    4. Six days of discourse: the US-China temperature gap over an apocalyptic title

  • English sphere — X hot, academic cold, framing geopolitical. The viral peak: @Dr_Singularity (09-13), "meanwhile in China / 'The Last AI Built by Humans'": 33k views, 890 likes, 132 reposts; top reply (912 likes): "China advancing towards the very thing we are shying away from because we are too scared!" SCMP (09-14): "Chinese researchers chart 5-stage path," adding "US firms have an early lead due to superior access to compute." Academic community near-silent: four independent HN submissions (incl. DAIR.AI) all scored 1–2 points, 2 comments total; on r/MachineLearning it was quoted ironically in a thread called "RSI is not happening" ("...has to be some psyops bullshit"); HuggingFace Papers 77 upvotes, the only "discussion" being an author clarification that inclusion ≠ claiming every entry is a fully autonomous RSI system.
  • Chinese sphere — popularization, star-making of Zhou Xuanhe, two critique streams. Synced (via Investment/Channel/Sina, 09-14) set the tone: L1-L5 breakdown, 72 companies placed, DAIR.AI quote, closing with the "ember that stokes, calibrates, and remembers its own path" metaphor — notably, Chinese coverage stars Zhou Xuanhe, barely mentioning the two big names. Critique split: the technical stream (Huxiu, 09-15, most systematic) — the five-level framework is "effective for communication and shared language" but fails as scientific measurement; the real criterion should be improvement-operator efficiency ρ(Mₜ₊₁)>ρ(Mₜ), proposing a seven-dimensional continuous space plus freeze-and-swap measurement protocols. The framing-skeptic stream (ifanr) — "if the AI circle ranked second at coinage and concept packaging, no one would dare claim first"; RSI as "a carefully packaged panic business," and those sounding the alarm are often the same people feeding the beast. Chinese X data accounts (@0xLogicrw) add: of 491 papers, L1 = 43.8%, L2 = 31.6%, L5 only 29 (5.9%) — honestly noting L1-L5 is the authors' standard, not industry consensus.
  • Six days, zero academic rebuttals, zero conflict-of-interest accusations, zero Alignment Forum threads. The sharpest three sentences: Pavlenko on LinkedIn ("Without an external anchor of truth, recursive self-improvement degenerates into recursive self-delusion"), Vastkind ("The title frames an ambition. It does not establish that humans have built their last AI."), and the paper itself (§4.3.4: "Full L5 improvement of the software-development improver has not yet been demonstrated").
  • 5. Independent observations

    1. The survey's real product is not the framework but "ledger power." The five-level framework is contestable (Huxiu contested it well), but no competitor exists for the 72-company table. First movers grade every system and assign every company a prototype; later comers either adopt its language or build a taxonomy nobody recognizes. AlphaEvolve at B0, SEAL silenced, OpenRSI compressed to a row: the ledger's power lives in editorial policy, written from a conflicted author seat (own Theseus in the table, own Frontis in the table, own ModelBest chapter). A first for the Chinese AI scene: a survey is infrastructure, and infrastructure should be built by your own people. 2. The discourse temperature gap maps to two narrative needs. English sphere consumes "China dares what we don't" (fear/competition); Chinese sphere consumes "we have systematic layout too" (catch-up confidence) — neither discusses the framework's measurement flaws. Huxiu's ρ(Mₜ₊₁)>ρ(Mₜ) proposal was the most constructive output in 6 days, and it appeared on a platform AI circles dismiss as "marketing accounts." Signals often live where you don't subscribe. 3. Undergrad first author + big-name endorsement deserves study as a specimen. That an economics undergraduate could coordinate a 33-person, 491-reference survey implies the literature-survey-classify-draft pipeline has been agentified enough for undergraduates to drive — a self-referential piece of evidence for the paper's own thesis (autonomization of improvement execution). In RSI terms: the academic-production improver is already improving, though the paper was too modest to say so about itself.

    Honest limitations

  • Team breakdown based on public homepages and self-reported author blocks; 16 SJTU authors not individually verified; Zhoufutu Wen's ByteDance/Humanlaya conflict resolved by paper self-report (confidence: C).
  • The EvoMap correspondence is name-coincidence-level inference (three proper names hit simultaneously), no second source (confidence: D) — the post deliberately leaves only a clue, no names.
  • Appendix B regional attribution is inferred (the paper deliberately avoids geography); US/China shares are rough estimates; 72 companies/142 rows counted line-by-line (2-row discrepancy vs the first post's "144 grade labels" likely a B0-row counting difference).
  • X engagement numbers are scrape-time values; HN data verified via Algolia API; one Reddit thread unreadable due to anti-scraping.
  • ModelBest/Tencent/Humanlaya numbers all come from paper §5 self-reports (partner first-party, no third-party benchmarks) — keep the "self-report" qualifier when citing.
  • Falsifiable predictions

  • Within 6 months: Appendix B becomes a filterable database in the awesome-rsi repo; first public "self-reported grade vs ledger grade" disputes appear (Frontis/OpenRSI may respond to the row-compression).
  • Within 6 months: at least one of the single-source companies (EvoMap/MetaCircle/Endless Frontier) surfaces with a product page/funding/paper, testing the table's due diligence.
  • Within 12 months: a competing survey (US or independent) ships its own grading framework, creating a contrast with L1-L5 — at which point "ledger power" upgrades from speculation to fact.
*Verification notes: arXiv 2609.11873 full HTML (~320k chars) locally parsed via five sub-agents (§1-3 / §4-7 close reads, appendix line-by-line, community reaction, team backgrounds); author block cross-checked against arXiv PDF and alphaXiv/HF Papers; personal homepages (duanyi516.github.io, xuanhe.gitbook.io, conghui.ai, randomtutu.github.io) treated as first-party (grade A); Synced/SCMP/Huxiu/QbitAI/Tsinghua official/Recode China AI as grade B; X/HN/Reddit/LinkedIn/Vastkind/HF as grade C; Baidu Baike and one forum post's "HCI rose from 0.3 to 0.7" figures contradict the original text (grade D, not adopted). First post: forum ID 178634859 (2026-09-14).*

Tags

#rsi#recursive-self-improvement#arxiv-2609-11873#theseus-labs#ai-industry-analysis#academic-survey#china-ai#openrsi

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634880