Summary
This independent research report examines Subquadratic, a Miami-based startup that claims its SubQ 1M-Preview is the first subquadratic frontier LLM, citing 52x faster prefill, 1000x attention compute reduction, 12M token context, and 1/5 the cost of frontier models. The report verifies that the company exists (SEC Form D, 2026), closed a $29M seed round, and operates a live API. However, it identifies serious evidentiary gaps: no public technical paper, no open weights, no independent reproduction, and conflicting numbers between the press release and technical post for competitor benchmarks. Critically, the CTO admitted the model is fine-tuned from open-source weights, contradicting the 'ground-up redesign' marketing language. The report also notes a 17-point gap between the research (83%) and production (65.9%) MRCR v2 scores without explanation, and compares SubQ to Magic.dev, which made similar unverified claims in 2024 and never delivered external proof. Overall rating: 6.5/10 for architectural plausibility but low evidence completeness.
Key Points
1. Claimed Breakthrough
Subquadratic (Miami) released SubQ 1M-Preview on 2026-05-19, claiming:
- ~1000x attention compute reduction vs. dense Transformer (12M tokens)
- 52.2x prefill speedup vs. FlashAttention-2 (1M tokens)
- 7.2x prefill speedup vs. FlashAttention-2 (128K tokens)
- 12M token research context window
- $8 cost on RULER 128K vs. Claude Opus ~$2,600
- RULER 128K: 95.0%; MRCR v2 (1M): 65.9% production / 83.0% research; SWE-Bench Verified: 81.8%
All benchmarks are self-reported or single unverified third-party tests. No independent reproduction exists.
2. SSA Architecture
The proposed Subquadratic Sparse Attention (SSA) is described as content-dependent dynamic routing where the selection step itself is subquadratic, distinct from sliding windows, Mamba-style recurrent state, or hybrid architectures like Kimi Linear (which mixes 3 linear layers + 1 quadratic MLA).The historical 'graveyard' of subquadratic attention includes Mamba, RWKV, DeepSeek Sparse Attention (whose indexer is itself quadratic — the 'indexer trap'), Longformer, and BigBird. SubQ claims to satisfy three simultaneously unproven constraints: subquadratic selection, no quadratic layers, and frontier-scale parity.
3. Evidence Chain
- High credibility: company exists (SEC), $29M seed round, CTO Alex Whedon confirmed use of open-source weights, GPU contract with Digi Power X ($19.6M), API runs
- Medium credibility: 11 PhDs (names undisclosed), single third-party benchmark verification
- Low credibility / red flags:
- 'Ground-up redesign' contradicts CTO's own admission of open-source weight fine-tuning
- 1000x and 52x figures are architecture-level, not end-to-end
- No public pricing to verify '1/5 cost'
- SWE-Bench: SubQ cites Opus at 80.8%, but Opus internal number is 87.6%
- 17-point MRCR gap between research and production models is unexplained
- Inconsistencies: press release shows Opus MRCR v2 at 32.2%, technical post shows 78.3% (2.4x discrepancy)
4. Historical Parallel: Magic.dev
Both companies claim 1000x efficiency and massive context windows (100M vs. 12M tokens), target software engineering, restrict access, and have never published peer-reviewed technical reports. Magic.dev's LTM-2-mini has shown no external usage evidence 21 months post-launch.5. Team and Backing
- CEO Justin Dangel: serial entrepreneur (healthtech, insurtech), no AI research background
- CTO Alex Whedon: ex-Meta engineer, TribeAI Head of Generative AI
- Research team: 11 PhDs from Meta/Google/Oxford/Cambridge/ByteDance/Adobe/Microsoft — names not disclosed, no foundational AI papers known
- Investors: Justin Mateen (Tinder co-founder), Javier Villamizar (ex-SoftBank Vision Fund), plus alumni from Anthropic/OpenAI/Stripe/Brex
6. Implications if True
If SSA works as claimed, it would be the most significant architectural shift since 2017: long context becomes default, inference costs collapse, KV-cache bottleneck disappears, and single-pass processing of entire codebases becomes routine. More likely outcome: a constant-factor speedup (5-10x) that becomes one component of a hybrid architecture rather than a Transformer replacement.7. Overall Rating: 6.5/10
- Architectural plausibility: 7/10
- Evidence completeness: 4/10
- Benchmark quality: 5/10
- Team credibility: 6/10
- Commercial viability: 7/10
- Marketing honesty: 5/10
Verdict: SubQ is real, the product runs, and the team has engineering capability. But core claims lack independent verification, marketing language exceeds technical reality, and no precedent exists for subquadratic attention succeeding at frontier scale. Developers can experiment with the API; investors should demand a technical report before additional commitment.
References
1. https://subq.ai/introducing-subq
2. https://venturebeat.com/technology/miami-startup-subquadratic-claims-1-000x-ai-efficiency-gain-with-subq-model-researchers-demand-independent-proof
3. https://thenewstack.io/subquadratic-12-million-context-window/
4. https://chatforest.com/reviews/subquadratic-subq-1m-preview-llm-review/
5. https://awesomeagents.ai/reviews/review-subq/
6. https://www.jakecuth.com/work/subquadratic-lab/
7. https://www.lesswrong.com/posts/kpSXeMcthtHgnwMx3/debunking-claims-about-subquadratic-attention
8. https://abhishek-shankar.com/posts/subquadratic-won-by-surrendering
9. https://www.atlaspeakresearch.com/report/542fd2
10. https://subq.ai/ssa
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177620445