Complete-muE – Optimal hyperparameter transfer for MoE models
WHY IT MATTERS
Research on hyperparameter optimization and transfer learning for mixture-of-experts models. Addresses efficient scaling of MoE architectures.
What Happened
Researchers introduced Complete-muE, a hyperparameter transfer method for mixture-of-experts (MoE) models that carries optimal configurations across architecture variants without full retuning. The technique extends maximal update parameterization (muP) — previously validated on dense transformers — to the distinct routing, gating, and expert-partitioning dynamics of MoE layers. Reported results show transferred learning rates and initialization scales remaining near-optimal as expert counts, capacity factors, and routing parameters change, eliminating the per-variant grid search that currently dominates MoE tuning cycles.
Why It Matters
MoE tuning is expensive because hyperparameters interact with architectural choices: the optimal learning rate for an 8-expert model is not the optimal learning rate for a 64-expert model, so each variant requires its own sweep. Complete-muE converts that sweep into a transfer step, meaning teams pay the tuning cost once and reuse the configuration across subsequent architecture edits. For organizations running multiple MoE variants in production — different expert counts for different latency or cost targets — this compresses the dominant fixed cost of architecture exploration. The strategic consequence is that MoE scaling shifts from brute-force search toward transfer learning, which changes the economics of how aggressively teams can iterate on routing and capacity decisions.
Technical Details
The method parameterizes MoE layers so that activation and gradient scales remain invariant to width — expert count and hidden dimension — as well as to capacity factor and routing temperature. This requires coordinated scaling rules for router logits, expert MLPs, and the load-balancing auxiliary loss coefficients, not just the dense-path matrices. Transfer is validated across model widths and depths, with the key claim being that a configuration tuned at small scale predicts the optimum at larger scale within a narrow band. Limitations: the approach assumes a fixed routing mechanism family; swapping between fundamentally different routers (e.g., top-1 to expert-choice) may still require partial retuning. Integration requires adopting the muE parameterization in the training code, which is a code-level change, not a config flag.
Operational Impact
The day-to-day change is that a new MoE variant — say, moving from 16 to 32 experts or adjusting capacity factor from 1.25 to 2.0 — can inherit the parent model's hyperparameters and require only a short confirmation run rather than a full sweep. This reduces the compute allocated to tuning, freeing budget for inference or additional architecture probes. Workflow-wise, teams can maintain a "tuning lineage" — a base configuration plus documented transfer rules — instead of re-deriving hyperparameters per experiment. The practical effect is that small architecture edits become cheap enough to test routinely, so routing and capacity decisions get made on evidence rather than on the cost of the experiment. Full retuning becomes the exception, reserved for genuinely new routing families.
SOURCE
ArXiv
SHARE
MORE FROM STUFFINSIDER
FuseReg: Layer Fusion Regularization for Representation Autoencoders
Sep 28RESEARCHInternW0-Delta Releases World Action Model With 20K+ Hours Open Data
Sep 28RESEARCHMicrosoft SkillOpt Trains Reusable Skills for Frozen LLM Agents
Sep 28RESEARCHCoding Agents for Generalized Task and Motion Planning
Sep 25