NVIDIA Open-Sources Model-Optimizer for LLM Compression
NVIDIA released a unified library bundling SOTA compression techniques — quantization, distillation, pruning, NAS, speculative decoding — with direct export paths to TensorRT-LLM, TensorRT, and vLLM.