Tencent Releases Hy3 Model – 295B Total with 21B Active Parameters
WHY IT MATTERS
Tencent released Hy3, a new open model with 295B total parameters and 21B active parameters under Apache 2.0 license. Addresses efficiency in large model deployment.
What Happened
Tencent released Hy3, an open-weight mixture-of-experts model with 295B total parameters and 21B active parameters per token, under an Apache 2.0 license. The release places Hy3 in the same architectural class as other large sparse open-weight models but with a markedly lower activation ratio, roughly 7% of total capacity engaged per forward pass. Tencent positions the model for cost-efficient deployment of large-scale inference rather than frontier capability leadership.
Why It Matters
MoE architectures decouple total parameter count from per-token compute, which means capacity can scale without proportional increases in inference cost. A 295B model with 21B active parameters consumes compute closer to a 21B dense model while retaining the representational breadth of a much larger one. Apache 2.0 removes the commercial-use ambiguity that has historically slowed enterprise adoption of restricted-license weights, and it places Hy3 in direct competition with other permissively licensed open models. For organizations currently routing workloads to commercial APIs or maintaining dense models for compliance reasons, this adds a credible self-hosted option. The cumulative effect of repeated releases at this scale is continued downward pressure on closed-model licensing economics.
Technical Details
Hy3 uses sparse expert routing, activating approximately 21B of 295B parameters per token, which yields an activation ratio near 7%. The Apache 2.0 license permits commercial deployment, modification, and redistribution without royalty obligations or usage caps. Memory footprint is dominated by total parameter storage at inference time—295B parameters require substantially more VRAM than the active set implies—so the 30-40% savings cited versus dense equivalents depend on whether the comparison target is a dense model with similar active compute or similar total capacity. Quantization, expert offloading, and paged attention are effectively required to fit the model across realistic cluster sizes. Throughput per dollar scales with active parameters; memory cost scales with total parameters. Benchmark figures were not independently verified at the time of release, and published evaluations should be treated as vendor-reported until reproduced.
Operational Impact
The 21B active footprint means per-token inference latency and FLOPs land in the range operators already provision for mid-size dense models, so existing serving stacks (vLLM, TensorRT-LLM, SGLang) can host Hy3 without architectural redesign provided memory capacity is sized for 295B weights. Teams running RAG systems, classification pipelines, and enterprise inference can now benchmark a permissively licensed model against commercial APIs on the same hardware profile. The binding constraint shifts from compute to memory: clusters that were compute-limited may now be memory-limited, which favors nodes with high VRAM-per-GPU counts and fast interconnect for expert parallelism. On-premise inference becomes economically defensible at scales where API costs previously dominated, particularly for steady-state, high-volume workloads. Cost curves compress, but capacity planning and quantization work become the new gating tasks.
SOURCE
Reddit r/LocalLLaMA
SHARE
MORE FROM STUFFINSIDER