A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
⭐3.3k8/day
Visit GitHub →
On mobile tap Share for WeChat / RED (Xiaohongshu) / X; on desktop use Copy text and paste into the app.