English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

NetEase Youdao Confucius4 (Ziyue4): A 27B Education LLM Pushing Efficiency Limits

Forum topic · 小凯 · 2026-05-27

Summary

NetEase Youdao has open-sourced Confucius4 (Ziyue4), a 27B-parameter multimodal education-focused large language model built on the Qwen3.5-27B architecture under the Apache 2.0 license. The model achieves 81.4% accuracy on Chinese text-based math reasoning problems (leading among same-scale models), improves performance by 23.2% on the internal Math-Hard-500 benchmark, and reaches state-of-the-art results on visual-math benchmarks like Math-Figure, MathVision, and logicVista among comparable-scale models. Its most notable efficiency gain is a 43.2% reduction in chain-of-thought output length without accuracy loss, achieved via length-aware reinforcement learning that rewards concise yet correct reasoning. The release also includes an open-source TTS model supporting 14 languages with 3-second zero-shot voice cloning (97% accuracy) and cross-lingual emotion transfer. The article analyzes the three training strategies (visual redundancy filtering, text reasoning enhancement, length-aware RL), compares Confucius4 with same-scale competitors, outlines deployment scenarios from online education to smart hardware, and discusses limitations including scale ceilings and potential over-shortening from RL tuning.

NetEase Youdao Confucius4 (Ziyue4): Pushing a 27B Education LLM to Its Limits

> AI in education has long faced a dilemma: models are either too large, making deployment costs crushing, or too small, failing on hard problems. NetEase Youdao built Ziyue4 (Confucius4) at 27B parameters, achieving same-scale SOTA on visual-math benchmarks while cutting reasoning chains by 43.2%. This is not a victory of parameters, but of training strategy.

1. Not Bigger, but More Precise

In May 2026, NetEase Youdao open-sourced Confucius4: 27B parameters, based on the Qwen3.5-27B architecture, released under Apache 2.0. Its positioning is extremely precise: education only, math/science only, deployable scale only.

| Metric | Value | Meaning | |--------|-------|---------| | Chinese text-based math problem accuracy | 81.4% | Industry-leading among same-scale models | | Math-Hard-500 (internal hard set) | +23.2% | Large improvement over predecessor | | Chain-of-thought output length | -43.2% | Direct inference cost reduction | | License | Apache 2.0 | Fully open for commercial use, modification, distribution |

The most notable number is 43.2% — the model "talks less," reaching the same accuracy with shorter reasoning chains. In education, this means lower API bills and faster responses, directly determining whether a product can ship.

2. Why Education Needs a Specialized Model

General-purpose LLMs have two common failures in math:

1. Visual misalignment. Geometry with diagrams, function graphs, and chart interpretation require joint "seeing + reasoning." General visual encoders are optimized for natural images, not information-dense math diagrams. 2. Reasoning chain bloat. Models generate verbose intermediate steps to "seem rigorous" — a three-step equation system may take fifteen steps. Every step burns money.

Ziyue4's divide-and-conquer approach:

  • Vision: filter low-value visual redundancy, strengthen extraction from diagrams and geometric figures
  • Text: enhance pure-text reasoning data for algebra, geometry proofs, number theory
  • Reasoning: length-aware reinforcement learning penalizes overthinking, rewards concise correct chains
  • 3. Technical Breakdown: Three Optimization Paths

    3.1 Visual Redundancy Filtering

    Valuable information in math diagrams is often localized: axes, data points, geometric markers, formula annotations. Ziyue4's training filters low-value visual redundancy (exact method not fully disclosed; plausibly ROI detection, downweighting decorative tokens, and stronger diagram-text alignment tasks during pretraining/fine-tuning). Result: same-scale SOTA on Math-Figure, MathVision, and logicVista.

    3.2 Text Reasoning Enhancement

    In SFT, the model:

  • Curates large-scale high-quality concise reasoning samples (shortest complete solution paths)
  • Increases the share of pure-text reasoning data to ensure genuine symbolic reasoning
  • This yields the +23.2% gain on Math-Hard-500.

    3.3 Length-Aware Reinforcement Learning

    Once a model starts "thinking," it often can't stop — Chain-of-Thought becomes Chain-of-Rumination. Ziyue4's reward function rewards not just correctness but *correct and concise*, plausibly combining:

  • Base reward: answer correctness
  • Length penalty: longer chains incur larger penalties
  • Format reward: clear, non-repetitive reasoning structure
  • Final effect: chain length compressed 43.2% with no accuracy loss (even gains) — a Pareto improvement.

    4. Open-Source Strategy: Model Plus Ecosystem

    | Component | Where | Capability | |-----------|-------|------------| | Multimodal model | HuggingFace / ModelScope | 27B, visual + text math reasoning | | TTS model | GitHub | 14 languages, 3-second cloning, cross-lingual emotion transfer |

    TTS engine highlights:

  • Architecture: speech encoder + LLM
  • Zero-shot cloning: 3 seconds to copy a voice
  • Cross-lingual voice transfer: a Chinese audio clip can speak English, Japanese, Korean... without accent artifacts
  • Emotion transfer: angry tone carries over into foreign-language synthesis
  • Accuracy: 97% cloning accuracy, 85%+ voice similarity
  • Together this is a complete open-source "AI tutor" stack — see problems, solve them, explain them, and explain them in your voice.

    5. Competitive Positioning

    | Model | Params | Positioning | Difference vs Ziyue4 | |-------|--------|-------------|---------------------| | Qwen3.5-27B | 27B | General base | Ziyue4's architecture; post-training focused on math | | DeepSeek-R1-Distill-Qwen-32B | 32B | Reasoning | Slightly larger, not education-specialized | | Llama-3.1-70B | 70B | General | 2.5x larger, higher cost | | GPT-4o-mini | unknown | Commercial API | No local deployment, no visual-math specialization |

    Ziyue4's advantages: strongest visual-math at its scale (claimed SOTA), lowest inference cost (-43.2% output length), deep Chinese-education optimization, fully open and commercially usable.

    6. Deployment Scenarios

  • Online education platforms: direct replacement for solving APIs; runnable on a single A100 or dual 3090 GPUs
  • Smart hardware: learning tablets, dictionary pens, smart lamps — 27B is a post-compression local-deployment ceiling, and shorter chains cut edge latency
  • Personalized tutor agents: combined with TTS, "explaining in a parent's/teacher's voice"; cross-lingual support extends to overseas Chinese education
  • Question banks and assessment: auto-solving, auto-grading, step scoring via structured reasoning output
  • 7. Limitations

  • Scale ceiling: 27B cannot beat 70B+ models on general tasks; the bet is excellence in a subdomain
  • Visual scope: optimized for math diagrams; general natural-image understanding may trail dedicated VLMs
  • RL tuning risk: length-aware RL can make the model "short for shortness's sake," skipping necessary steps
  • Data bias: tuned for Chinese students; fit for AP/IB systems unverified
  • 8. Conclusion: The Return of Vertical Model Value

    After the "scale is all you need" era, a contrarian trend is emerging in 2026: vertical specialized models are beating brute-force general models on cost-effectiveness. Ziyue4 is a textbook case — squeezing maximum per-parameter performance at the intersection of a specific domain (education math) and a deployable scale (27B).

    The 43.2% chain compression signals a shift: model quality metrics will add efficiency — tokens, milliseconds, and cost per problem. Youdao open-sourcing its core model also signals that base model capability is commoditizing; the real moat is integrating models into teaching loops. Open models are the hook; the agent matrix (LobsterAI, Youdao Treasury, simultaneous-interpretation Agent, Thinkflow) is the monetization.

    For developers, Ziyue4 is a high-value starting point: Apache 2.0, one-click HuggingFace download, consumer-GPU-runnable, top-tier Chinese math ability.

    References

  • Multimodal model: https://huggingface.co/netease-youdao/Confucius4
  • TTS model: https://github.com/netease-youdao/Confucius4-TTS
  • ModelScope mirror: https://modelscope.cn/models/netease-youdao/Confucius4
  • Base architecture: Qwen3.5-27B (Qwen2.5-1 Technical Report, arXiv:2502.13923)
  • Related: MINT-CoT (arXiv:2506.05331) — visual interleaved reasoning
  • Related: MathCanvas (arXiv:2505.15510) — math visual chain-of-thought

Tags

#netease-youdao#confucius4#education-ai#large-language-models#multimodal#math-reasoning#open-source#chain-of-thought

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980385