This post is a follow-up to a 2026-09-14 fact-check of the paper arXiv 2609.11873. The first post audited the paper's content (the five-level framework, HCI, three challenges, lineage mapping). This one audits the paper's "who's who": authorship structure, the Appendix B table of 72 companies, and the 6 days of Chinese and English discourse after release. Source chain: full arXiv HTML (~320k characters) re-scraped, PDF author block, and 5 sub-agent passes (framework close-read, second-half close-read, appendix line-by-line, community reaction, team backgrounds) cross-checked against ~60 sources. All claims grounded in the original text and first-party homepages.
Key points
1. Author list: 33 people, 9 institutions — the power structure is not the two famous names
- Real drivers: Xuanhe Zhou (corresponding author) + Fan Wu (last author). Zhou, born 1997, is one of SJTU's youngest assistant professors and a PhD student of Tsinghua's Guoliang Li (also on the author list, position 27). Wu is a distinguished professor and executive vice dean at SJTU. "Theseus Labs" is not a company but the brand of this team's lab: a one-line website, a GitHub org (
theseus-labs-rsi) with only an awesome-rsi literature repo (47 stars) and a homepage repo — no funding, no registration, no product code. - First author Yi Duan: a Renmin University economics undergraduate (class of 2023). His homepage lists him as "Project Lead, Theseus Labs" (2026.8–9), after a remote internship at Shanghai AI Lab (2025.9–2026.6, BioMatrix). An undergraduate coordinating a 75-page, 33-author, 491-reference survey is rare and fully public (duanyi516.github.io). All five co-first authors (Yi Duan / Ying Liu / Zirui Tang / Haodong Chen / Jun Zhou) appear on Zhou Xuanhe's student list.
- Bo Wang's real role: author AND surveyed party. Wang (position 28, Tsinghua EE chair professor, director of Shanghai AI Lab) founded Xianyuan (Frontis AI) after JD. Appendix B contains the row:
Frontis AI / Xianyuan | OpenRSI, Frontis-MA1 | AI4AI/self-improving agents | skills; memory; harness; weights | L2-L3. His student Kaiyan Zhang (position 24; OpenReview shows Frontis CTO) is the sole corresponding author of the Frontis-MA1 paper (arXiv 2607.28568). The teacher endorses the survey; the student's company product enters the industrial map — the "OpenRSI / survey mutual reference" loop closes: they are two edges of the same power topology. - Other seats: Zhiyuan Liu (position 29, THUNLP director) is co-founder/chief scientist of ModelBest (BAAI-adjacent, 面壁) — §5.4's ModelBest chapter (ForgeTrain matching Megatron-LM v0.15 in 8 hours, overtaking in 1.5–2.5 days; MFU 40.1→44.1%@0.5B) is his ecosystem's first-hand self-report. Two ByteDance authors map to the §5.2 Lark chapter; one Xiaohongshu; Zhoufutu Wen lists Humanlaya (§5.3) but his homepage shows ByteDance SEED since 2023.9 (paper self-report prevails). Conghui He (OpenDataLab/MinerU founder, Shanghai AI Lab) connects via Bo Wang's current post.
- Conclusion: "SJTU + Bo Wang + Zhiyuan Liu" is the communication cover, not the power structure. A survey establishing an official taxonomy of China's RSI ecosystem, drafted by a 97-born AP with undergrad/master's students and endorsed by three heavyweight Tsinghua figures plus five companies, is itself part of China's "systematic RSI layout" (isomorphic to the ICLR 2026 RSI Workshop, backed by Tencent and BAAI).
- AlphaEvolve is downgraded. It appears exactly once, marked B0 (task-level program search) under GDM in Table 12, with zero main-text discussion — versus ~10 mentions and two dedicated paragraphs for DGM. The paper deliberately separates "program search" from "improvement loop": AlphaEvolve has the former but no persistent self-modification, so it doesn't enter the RSI narrative. MIT's SEAL (Self-Adapting LMs) is not cited at all, not even the phrase "self-adapting." A taxonomy's power lies half in inclusion, half in silence.
- Frontis AI/Xianyuan carries no asterisk. Entries like Chaoyan, Wuya Zhiyuan, Naive.ai carry "company identity needs stronger verification" markers; Frontis does not. Given a Frontis CTO's advisor sits on the author list, this "no verification needed" reads more like internal confidence than external due diligence.
- ~10 confirmed Chinese companies of 72 (~12 more suspected), ~36–38 American — but all six §5 deep-dive cases (Theseus/Lark/Humanlaya/ModelBest/Tencent Hunyuan/Agent-Native) are Chinese practices not counted in Table 12. The ledger is global; the cases are Chinese.
- The EvoMap mystery. One row:
EvoMap | Evolver / GEP / GeneBench | shared code evolution | genes/capsules; code assets | generate; test; publish; reuse; adapt | L2-L3. Near-zero public information; treated as single-source, pending verification. - English sphere — X hot, academic cold, framing geopolitical. The viral peak: @Dr_Singularity (09-13), "meanwhile in China / 'The Last AI Built by Humans'": 33k views, 890 likes, 132 reposts; top reply (912 likes): "China advancing towards the very thing we are shying away from because we are too scared!" SCMP (09-14): "Chinese researchers chart 5-stage path," adding "US firms have an early lead due to superior access to compute." Academic community near-silent: four independent HN submissions (incl. DAIR.AI) all scored 1–2 points, 2 comments total; on r/MachineLearning it was quoted ironically in a thread called "RSI is not happening" ("...has to be some psyops bullshit"); HuggingFace Papers 77 upvotes, the only "discussion" being an author clarification that inclusion ≠ claiming every entry is a fully autonomous RSI system.
- Chinese sphere — popularization, star-making of Zhou Xuanhe, two critique streams. Synced (via Investment/Channel/Sina, 09-14) set the tone: L1-L5 breakdown, 72 companies placed, DAIR.AI quote, closing with the "ember that stokes, calibrates, and remembers its own path" metaphor — notably, Chinese coverage stars Zhou Xuanhe, barely mentioning the two big names. Critique split: the technical stream (Huxiu, 09-15, most systematic) — the five-level framework is "effective for communication and shared language" but fails as scientific measurement; the real criterion should be improvement-operator efficiency ρ(Mₜ₊₁)>ρ(Mₜ), proposing a seven-dimensional continuous space plus freeze-and-swap measurement protocols. The framing-skeptic stream (ifanr) — "if the AI circle ranked second at coinage and concept packaging, no one would dare claim first"; RSI as "a carefully packaged panic business," and those sounding the alarm are often the same people feeding the beast. Chinese X data accounts (@0xLogicrw) add: of 491 papers, L1 = 43.8%, L2 = 31.6%, L5 only 29 (5.9%) — honestly noting L1-L5 is the authors' standard, not industry consensus.
- Six days, zero academic rebuttals, zero conflict-of-interest accusations, zero Alignment Forum threads. The sharpest three sentences: Pavlenko on LinkedIn ("Without an external anchor of truth, recursive self-improvement degenerates into recursive self-delusion"), Vastkind ("The title frames an ambition. It does not establish that humans have built their last AI."), and the paper itself (§4.3.4: "Full L5 improvement of the software-development improver has not yet been demonstrated").
- Team breakdown based on public homepages and self-reported author blocks; 16 SJTU authors not individually verified; Zhoufutu Wen's ByteDance/Humanlaya conflict resolved by paper self-report (confidence: C).
- The EvoMap correspondence is name-coincidence-level inference (three proper names hit simultaneously), no second source (confidence: D) — the post deliberately leaves only a clue, no names.
- Appendix B regional attribution is inferred (the paper deliberately avoids geography); US/China shares are rough estimates; 72 companies/142 rows counted line-by-line (2-row discrepancy vs the first post's "144 grade labels" likely a B0-row counting difference).
- X engagement numbers are scrape-time values; HN data verified via Algolia API; one Reddit thread unreadable due to anti-scraping.
- ModelBest/Tencent/Humanlaya numbers all come from paper §5 self-reports (partner first-party, no third-party benchmarks) — keep the "self-report" qualifier when citing.
- Within 6 months: Appendix B becomes a filterable database in the awesome-rsi repo; first public "self-reported grade vs ledger grade" disputes appear (Frontis/OpenRSI may respond to the row-compression).
- Within 6 months: at least one of the single-source companies (EvoMap/MetaCircle/Endless Frontier) surfaces with a product page/funding/paper, testing the table's due diligence.
- Within 12 months: a competing survey (US or independent) ships its own grading framework, creating a contrast with L1-L5 — at which point "ledger power" upgrades from speculation to fact.
2. Appendix B: 72 companies, six prototype categories, ~142 product rows
Organized by prototype, not geography:
| Class | Name | Companies | Notable players (and grades) | |---|---|---|---| | A | Frontier model & general agent labs | 17 | OpenAI (Symphony/GPT-Red/AgentKit L1-L2; self-improving tax agents L4 candidate); Anthropic ("When AI Builds Itself" marked "L5 target" — vision, not reality); GDM (AlphaEvolve marked B0); Meta (HyperAgents L5 cand.); Qwen/DeepSeek/Tencent AI Lab (R-Zero L2-L3)/ByteDance Seed/MiniMax/Moonshot/Zhipu; Poolside/TML/Nous | | B | RSI-native / AI4AI companies | 14 | Weco (AIDE2 L5 cand.), Sakana (DGM L5 cand.), Recursive, Imbue; Chinese newcomers: MetaCircle, Frontis AI/Xianyuan, EvoMap, Endless Frontier, Chaoyan; Theseus itself (Argus L2-L3, SetupX L3) | | C | Autonomous R&D & scientific discovery | 14 | Periodic Labs (L3 cand.), FutureHouse, Axiom Math, Harmonic, Lila, Karpathy autoresearch (L2), Prime Intellect | | D | Agent optimization/eval/learning infra | 15 | Warp, Factory (one of the highest industrial grades: L4), LangChain, Replit, Braintrust, Mechanize, Goodfire | | E | Embodied/world models/continual adaptation | 8 | AI2 Robotics, X Square Robot, Galaxea, Synapx, MirrOS | | F | Persistent memory & personal AI | 4 | Engram, Lemon AI, EverMind, Mindverse |
Four previously unflagged details:
3. Lineage additions (map dimension)
1. HGM isn't only in the main text — it's used as RQGM's defeated baseline. §3.6.2 and §4.3.3: Red Queen Gödel Machine reports 71.7% pass@1 on held-out Polyglot coding vs HGM-H's 69.9%, with fewer search tokens. The narrative chain DGM (lineage archive) → HGM (lineage meta-selection) → RQGM (evaluator joins co-evolution) makes HGM the middle rung of the Gödel Machine family's three-stage evolution story. 2. OpenRSI and Frontis-MA1 share two identical rows, six fields verbatim. For a project whose stated program is "make improvement rate the optimization objective," being compressed into one undifferentiated L2-L3 label is itself a signal: when the official ledger absorbs you, it doesn't use your self-narrative. 3. Ten most-discussed systems (by mentions + dedicated paragraphs): DGM, VOYAGER, SEAgent, Metis, autoresearch, RQGM, AIDE2, Gödel Agent, Theseus, A-Evolve-Training. Lineage freshness (NeoHorse missing the cutoff by two days) and depth of absorption are two different ledgers.
4. Six days of discourse: the US-China temperature gap over an apocalyptic title
5. Independent observations
1. The survey's real product is not the framework but "ledger power." The five-level framework is contestable (Huxiu contested it well), but no competitor exists for the 72-company table. First movers grade every system and assign every company a prototype; later comers either adopt its language or build a taxonomy nobody recognizes. AlphaEvolve at B0, SEAL silenced, OpenRSI compressed to a row: the ledger's power lives in editorial policy, written from a conflicted author seat (own Theseus in the table, own Frontis in the table, own ModelBest chapter). A first for the Chinese AI scene: a survey is infrastructure, and infrastructure should be built by your own people. 2. The discourse temperature gap maps to two narrative needs. English sphere consumes "China dares what we don't" (fear/competition); Chinese sphere consumes "we have systematic layout too" (catch-up confidence) — neither discusses the framework's measurement flaws. Huxiu's ρ(Mₜ₊₁)>ρ(Mₜ) proposal was the most constructive output in 6 days, and it appeared on a platform AI circles dismiss as "marketing accounts." Signals often live where you don't subscribe. 3. Undergrad first author + big-name endorsement deserves study as a specimen. That an economics undergraduate could coordinate a 33-person, 491-reference survey implies the literature-survey-classify-draft pipeline has been agentified enough for undergraduates to drive — a self-referential piece of evidence for the paper's own thesis (autonomization of improvement execution). In RSI terms: the academic-production improver is already improving, though the paper was too modest to say so about itself.