最高第 2 名 现在第 2 名
AirLLM低显存跑大模型
AirLLM 用逐层推理让 4GB 显卡跑 70B、8GB 跑 405B,支持 Llama/Qwen/Mistral,已开源。
要点
- AirLLM 采用"Layer-wise Inference",一次只加载、计算并清除一层,而不是加载整个模型。
- 它能在单张 4GB GPU 上运行 70B 模型,在 8GB VRAM 上可扩展到 Llama 3.1 405B。
- 默认无需量化,支持 Llama、Qwen 和 Mistral,可在 Linux、Windows 和 macOS 上运行。
- 它 100% 开源。
要点和反应摘要由 AI 依据本页推文整理,请以原推为准。 我们怎么用 AI
原推
"I don't have a GPU" is officially dead 🤯
You can now run 70B model on a single 4GB GPU and it even scales up to the colossal Llama 3.1 405B on just 8GB of VRAM.
AirLLM uses "Layer-wise Inference." Instead of loading the whole model, it loads, computes, and flushes one layer at a time
→ No quantization needed by default
→ Supports Llama, Qwen, and Mistral
→ Works on Linux, Windows, and macOS
100% Open Source.
X 上的热门回复
X 上的反应: 回复基本是调侃,集中吐槽速度极慢——有人开玩笑说每 10 分钟才出 1 个 token,还有人拿 SSD 损耗、龟速和用 CPU 跑说事。也有回复承认这原本不可能,若不考虑能耗和时间,仍有潜在用途。原推作者没有回应。
讨论 0
注册 登录 后参与讨论
还没有人讨论,来说第一句。