~/home / Hacker News
Hacker News · Show HN

We quantized Qwen3.6-35B-A3B to 2-bit: 12.3 GB, 225 tok/s on a 4090

将Qwen3.6-35B模型量化至2比特,仅需12.3GB显存,在RTX 4090上实现225 tokens/秒的推理速度,适合本地部署大模型的开发者与AI研究者,显著降低硬件门槛。
▲21/天
为什么值得关注

在消费级显卡上实现大模型高效推理,突破了大模型本地化部署的性能瓶颈。

AI开发者工具模型优化效率
访问 Hacker News 页面 →
分享到 X
手机端点「分享」直达微信/朋友圈/小红书;桌面端用「复制文案」后到 App 内粘贴发布

同类产品

I was tired of opening 2 tabs for every HN link, so I made a userscriptOpen-source engine running Gemma 4 26B in 2 GB RAM on any M-series MacCheapFoodMap – A map of good meals under $10Firemaps Spain – Live wildfire map with wind flow for ES and PTXY – A Fast, composable, GPU-accelerated interactive plotting libraryRun Full Kimi K3 with 29 GB of RAM