由于 trendforge.devlive.top 访问受限,请切换到 trendforge.devlive.org 域名
Project Intro
使用单张4GB GPU实现AirLLM 70B推理
AirLLM 70B inference with single 4GB GPU
AI Interpretation
AirLLM大幅降低大模型推理内存占用,让70B参数模型可运行在单张4GB GPU上,无需量化、蒸馏或剪枝。通过稀疏MoE逐专家流式加载机制,还能运行405B Llama 3.1、DeepSeek-V3(671B)甚至2.8T参数的Kimi K3模型,显存占用分别仅需8GB、12GB和不到4GB。支持FP8量化及Qwen3、Llama等主流开源模型,适合资源有限的个人开发者进行大模型推理和微调实验。
Original Tags
AI Tags
Use Cases
Project Health Score
1 days since last update
Platform Star TOP 2% · Forks 3,162
This week +1,588 ⭐ · This month +3,621 ⭐
9 contributors · 0 platform comments
1 item(s) missing
Project Info
Support
If this site has been helpful, feel free to support us
Alipay
Widget Badge
Related Projects
jackfrued/Python-100-Days
Python - 100天从新手到大师
microsoft/generative-ai-for-beginners
21堂课带你入门生成式AI开发
rasbt/LLMs-from-scratch
使用PyTorch从零开始逐步实现类ChatGPT大型语言模型
microsoft/ML-For-Beginners
12周、26节课、52个测验,面向所有人的经典机器学习课程
openai/openai-cookbook
OpenAI API 使用示例与指南
CompVis/stable-diffusion
潜在文本到图像扩散模型
Loading comments...