NetEase Youdao Confucius4 (Ziyue4): Pushing a 27B Education LLM to Its Limits
> AI in education has long faced a dilemma: models are either too large, making deployment costs crushing, or too small, failing on hard problems. NetEase Youdao built Ziyue4 (Confucius4) at 27B parameters, achieving same-scale SOTA on visual-math benchmarks while cutting reasoning chains by 43.2%. This is not a victory of parameters, but of training strategy.
1. Not Bigger, but More Precise
In May 2026, NetEase Youdao open-sourced Confucius4: 27B parameters, based on the Qwen3.5-27B architecture, released under Apache 2.0. Its positioning is extremely precise: education only, math/science only, deployable scale only.
| Metric | Value | Meaning | |--------|-------|---------| | Chinese text-based math problem accuracy | 81.4% | Industry-leading among same-scale models | | Math-Hard-500 (internal hard set) | +23.2% | Large improvement over predecessor | | Chain-of-thought output length | -43.2% | Direct inference cost reduction | | License | Apache 2.0 | Fully open for commercial use, modification, distribution |
The most notable number is 43.2% — the model "talks less," reaching the same accuracy with shorter reasoning chains. In education, this means lower API bills and faster responses, directly determining whether a product can ship.
2. Why Education Needs a Specialized Model
General-purpose LLMs have two common failures in math:
1. Visual misalignment. Geometry with diagrams, function graphs, and chart interpretation require joint "seeing + reasoning." General visual encoders are optimized for natural images, not information-dense math diagrams. 2. Reasoning chain bloat. Models generate verbose intermediate steps to "seem rigorous" — a three-step equation system may take fifteen steps. Every step burns money.
Ziyue4's divide-and-conquer approach:
- Vision: filter low-value visual redundancy, strengthen extraction from diagrams and geometric figures
- Text: enhance pure-text reasoning data for algebra, geometry proofs, number theory
- Reasoning: length-aware reinforcement learning penalizes overthinking, rewards concise correct chains
- Curates large-scale high-quality concise reasoning samples (shortest complete solution paths)
- Increases the share of pure-text reasoning data to ensure genuine symbolic reasoning
- Base reward: answer correctness
- Length penalty: longer chains incur larger penalties
- Format reward: clear, non-repetitive reasoning structure
- Architecture: speech encoder + LLM
- Zero-shot cloning: 3 seconds to copy a voice
- Cross-lingual voice transfer: a Chinese audio clip can speak English, Japanese, Korean... without accent artifacts
- Emotion transfer: angry tone carries over into foreign-language synthesis
- Accuracy: 97% cloning accuracy, 85%+ voice similarity
- Online education platforms: direct replacement for solving APIs; runnable on a single A100 or dual 3090 GPUs
- Smart hardware: learning tablets, dictionary pens, smart lamps — 27B is a post-compression local-deployment ceiling, and shorter chains cut edge latency
- Personalized tutor agents: combined with TTS, "explaining in a parent's/teacher's voice"; cross-lingual support extends to overseas Chinese education
- Question banks and assessment: auto-solving, auto-grading, step scoring via structured reasoning output
- Scale ceiling: 27B cannot beat 70B+ models on general tasks; the bet is excellence in a subdomain
- Visual scope: optimized for math diagrams; general natural-image understanding may trail dedicated VLMs
- RL tuning risk: length-aware RL can make the model "short for shortness's sake," skipping necessary steps
- Data bias: tuned for Chinese students; fit for AP/IB systems unverified
- Multimodal model: https://huggingface.co/netease-youdao/Confucius4
- TTS model: https://github.com/netease-youdao/Confucius4-TTS
- ModelScope mirror: https://modelscope.cn/models/netease-youdao/Confucius4
- Base architecture: Qwen3.5-27B (Qwen2.5-1 Technical Report, arXiv:2502.13923)
- Related: MINT-CoT (arXiv:2506.05331) — visual interleaved reasoning
- Related: MathCanvas (arXiv:2505.15510) — math visual chain-of-thought
3. Technical Breakdown: Three Optimization Paths
3.1 Visual Redundancy Filtering
Valuable information in math diagrams is often localized: axes, data points, geometric markers, formula annotations. Ziyue4's training filters low-value visual redundancy (exact method not fully disclosed; plausibly ROI detection, downweighting decorative tokens, and stronger diagram-text alignment tasks during pretraining/fine-tuning). Result: same-scale SOTA on Math-Figure, MathVision, and logicVista.
3.2 Text Reasoning Enhancement
In SFT, the model:
This yields the +23.2% gain on Math-Hard-500.
3.3 Length-Aware Reinforcement Learning
Once a model starts "thinking," it often can't stop — Chain-of-Thought becomes Chain-of-Rumination. Ziyue4's reward function rewards not just correctness but *correct and concise*, plausibly combining:
Final effect: chain length compressed 43.2% with no accuracy loss (even gains) — a Pareto improvement.
4. Open-Source Strategy: Model Plus Ecosystem
| Component | Where | Capability | |-----------|-------|------------| | Multimodal model | HuggingFace / ModelScope | 27B, visual + text math reasoning | | TTS model | GitHub | 14 languages, 3-second cloning, cross-lingual emotion transfer |
TTS engine highlights:
Together this is a complete open-source "AI tutor" stack — see problems, solve them, explain them, and explain them in your voice.
5. Competitive Positioning
| Model | Params | Positioning | Difference vs Ziyue4 | |-------|--------|-------------|---------------------| | Qwen3.5-27B | 27B | General base | Ziyue4's architecture; post-training focused on math | | DeepSeek-R1-Distill-Qwen-32B | 32B | Reasoning | Slightly larger, not education-specialized | | Llama-3.1-70B | 70B | General | 2.5x larger, higher cost | | GPT-4o-mini | unknown | Commercial API | No local deployment, no visual-math specialization |
Ziyue4's advantages: strongest visual-math at its scale (claimed SOTA), lowest inference cost (-43.2% output length), deep Chinese-education optimization, fully open and commercially usable.
6. Deployment Scenarios
7. Limitations
8. Conclusion: The Return of Vertical Model Value
After the "scale is all you need" era, a contrarian trend is emerging in 2026: vertical specialized models are beating brute-force general models on cost-effectiveness. Ziyue4 is a textbook case — squeezing maximum per-parameter performance at the intersection of a specific domain (education math) and a deployable scale (27B).
The 43.2% chain compression signals a shift: model quality metrics will add efficiency — tokens, milliseconds, and cost per problem. Youdao open-sourcing its core model also signals that base model capability is commoditizing; the real moat is integrating models into teaching loops. Open models are the hook; the agent matrix (LobsterAI, Youdao Treasury, simultaneous-interpretation Agent, Thinkflow) is the monetization.
For developers, Ziyue4 is a high-value starting point: Apache 2.0, one-click HuggingFace download, consumer-GPU-runnable, top-tier Chinese math ability.