A flexible framework for optimizing heterogeneous LLM inference and fine-tuning, built for developers and researchers working on efficient LLM deployment
1.5k upvotes
Visit GitHub →
On mobile tap Share for WeChat / RED (Xiaohongshu) / X; on desktop use Copy text and paste into the app.