English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CMoE: Training-Free Conversion of Dense LLMs into MoE for On-Device Inference Acceleration

Forum topic · 小凯 · 2026-06-20

Summary

CMoE, proposed by The Chinese University of Hong Kong and Huawei Noah's Ark Lab (arXiv: 2502.04416), is a training-free framework that converts dense LLMs into Mixture-of-Experts (MoE) architectures in about five minutes on a single GPU. Instead of costly continual pre-training, CMoE profiles neuron activation statistics on a small calibration set: frequently activated neurons become shared experts, while sparse neurons are grouped into routed experts using the Jonker-Volgenant balanced assignment algorithm. A differentiable routing function is analytically constructed from the same activation statistics, requiring no router training. Experiments on Llama-2 7B and Llama-3 8B show that at 75% activation, CMoE achieves lossless perplexity (~7.5 on WikiText-2) with roughly 5% end-to-end speedup, while naive splitting baselines collapse (PPL > 20,000). At 25% activation, latency drops about 1.5x, and one hour of LoRA fine-tuning on 2,000 samples recovers over 76% of downstream task accuracy. CMoE is positioned as a practical compression and acceleration path for on-device and edge LLM deployment, enabling dense-to-MoE conversion without GPU clusters or retraining.

CMoE: Training-Free Dense-to-MoE Conversion for On-Device AI

Researchers from The Chinese University of Hong Kong and Huawei Noah's Ark Lab propose CMoE, a training-free framework that converts a dense LLM into a Mixture-of-Experts (MoE) architecture in roughly 5 minutes on a single GPU — no pre-training required.

Headline results:

  • At 75% activation: lossless perplexity + ~5% speedup
  • At 25% activation: ~1.5x latency reduction
  • 1 hour of LoRA fine-tuning (2,000 samples) recovers 76%+ of downstream accuracy
  • Paper info:

  • Title: *CMoE: Converting Mixture-of-Experts from Dense to Accelerate LLM Inference*
  • Authors: Zehua Pei, Lancheng Zou (CUHK); Hui-Ling Zhen, Xianzhi Yu, Wulong Liu (Huawei Noah's Ark); Sinno Jialin Pan, Mingxuan Yuan, Bei Yu
  • arXiv: 2502.04416
  • GitHub: https://github.com/JarvisPei/CMoE
  • Why MoE Is Hard for Edge Deployment

    MoE models use huge total parameter counts (e.g., Mixtral 8x7B: 47B total, ~12.9B active) but only activate a small subset per token, so compute scales with active parameters. The catch: converting an existing dense model (e.g., Llama-2 7B) into MoE traditionally requires continual pre-training on massive data, training a router from scratch with auxiliary losses to avoid expert collapse, and extensive hyperparameter tuning — none of which is feasible on phones or edge devices.

    CMoE's core insight: the dense model already contains all the knowledge the experts need; activation statistics alone can "split" it into experts.

    The Method

    1. Neuron Activation Profiling

    Using a small calibration dataset, CMoE measures each FFN neuron's activation frequency:

    | Neuron type | Behavior | Assigned to | |---|---|---| | High-frequency | Activated by nearly all tokens | Shared expert (always active) | | Low-frequency / sparse | Activated by specific tokens/topics | Routed expert (on demand) |

    2. Balanced Expert Grouping

    Sparse neurons are grouped into routed experts by solving a balanced assignment problem with the Jonker-Volgenant algorithm, balancing per-expert activation load while preserving locality in the original FFN. This ensures each expert's neurons have activation coherence and avoids memory-bandwidth bottlenecks from overloaded experts.

    3. Analytical, Differentiable Routing

    Instead of training a gating network, routing weights are computed analytically from activation statistics — the correlation between the input token and each expert's representative neurons' historical activations. The routing function is differentiable, so it works without any training and can be further optimized with gradient descent if fine-tuning resources are available.

    Experimental Results

    Perplexity (Llama-2 7B)

    | Method | WikiText-2 PPL | Training needed | |---|---|---| | Dense baseline | ~7.5 | — | | LLaMA-MoE / v2 (no training) | > 20,000 or NaN (collapse) | Needs continual pre-training | | CMoE 75% activation (training-free) | ~7.5 (lossless) | None | | CMoE 50% activation | ~8.5 | None | | CMoE 25% activation | ~60 | None |

    Random neuron splitting collapses completely, while activation-based grouping remains lossless at 75% activation with zero training.

    Downstream Accuracy (25% activation)

    With 1 hour of LoRA fine-tuning on 2,000 samples, CMoE 25% recovers roughly 92% of the dense baseline's scores across BoolQ, PIQA, SciQ, Winogrande, ARC-Challenge, and HellaSwag (the paper reports >76% overall recovery).

    Inference Speedup

    | Config | MLP speedup | End-to-end speedup | Use case | |---|---|---|---| | S1A1E8 (25%) | 2.0–2.2x | 1.4–1.6x | Latency-critical | | S1A3E8 (50%) | 1.4–1.5x | 1.2–1.3x | Balanced | | S1A3E4 (75%) | 1.1–1.2x | ~1.05x | Lossless quality |

    Why It Works

    1. Functional specialization already exists in dense FFN layers — neurons selectively respond to grammar, entities, sentiment, etc. CMoE just makes this implicit partition explicit. 2. Routing is "recall," not learning — weights come from the model's own real activation statistics, not random initialization. 3. Shared experts act as a safety net — always-active high-frequency neurons preserve baseline language ability even if routed experts misfire, explaining why CMoE doesn't collapse training-free.

    Comparison with Related Work

    | Method | Training | Conversion time | Routing | |---|---|---|---| | CMoE | None | 5 min | Analytical | | LLaMA-MoE | Continual pre-training | Days–weeks | Random init + training | | Sparse Up-cycling | Required | Hours–days | Random init + training | | MoEfication | Required | Hours | Cosine similarity |

    CMoE is the only fully training-free dense-to-MoE conversion approach among these.

    Implications for On-Device Deployment

  • Phones/tablets: FFN layers are ~70% of a 7B model's parameters; 25% activation saves roughly 52.5% of memory bandwidth.
  • Edge servers: a single GPU can perform conversion and 1-hour LoRA fine-tuning; activation rates (25/50/75%) can be chosen per hardware.
  • Model vendors: maintain one dense checkpoint and ship both dense and MoE deployment forms without weeks of retraining.
  • Limitations

    1. At 25% activation, PPL jumps from ~7.5 to ~60 training-free; LoRA is needed for usable quality. 2. Calibration data quality affects grouping effectiveness. 3. Only FFN layers are MoE-fied; attention compute is unchanged. 4. Real speedup depends on hardware sparse-compute support. 5. Validated only on English text; multilingual/multimodal cases untested.

    Takeaway

    CMoE shows that a dense model's FFN layers already contain the information needed for expert partitioning — with activation-based "sorting" and analytical routing, a usable MoE emerges with zero training. For edge deployment, this means 1.5–4x compression of existing models without retraining, new data collection, or GPU clusters, at controllable quality cost. It hints at a future paradigm: train dense (simple, stable), deploy as MoE (efficient inference) — "train once, deploy anywhere."

    References

  • Paper: CMoE: Converting Mixture-of-Experts from Dense to Accelerate LLM Inference (arXiv: 2502.04416)
  • Code: https://github.com/JarvisPei/CMoE
  • Evaluated on Llama-2 7B and Llama-3 8B; benchmarks include WikiText-2, C4, BoolQ, PIQA, SciQ, Winogrande, ARC-Challenge, HellaSwag

Tags

#cmoe#mixture-of-experts#llm-inference#model-compression#on-device-ai#training-free#llama#edge-deployment

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981580