由于 trendforge.devlive.top 访问受限,请切换到 trendforge.devlive.org 域名

lyogavin

lyogavin/airllm

Jupyter Notebook Active
692
2026-08-06
29k
+603
#2
3.2k

Project Intro

使用单张4GB GPU实现AirLLM 70B推理

AirLLM 70B inference with single 4GB GPU

AI Interpretation

AI Interpretation Auto generated

AirLLM大幅降低大模型推理内存占用,让70B参数模型可运行在单张4GB GPU上,无需量化、蒸馏或剪枝。通过稀疏MoE逐专家流式加载机制,还能运行405B Llama 3.1、DeepSeek-V3(671B)甚至2.8T参数的Kimi K3模型,显存占用分别仅需8GB、12GB和不到4GB。支持FP8量化及Qwen3、Llama等主流开源模型,适合资源有限的个人开发者进行大模型推理和微调实验。

AI Tags

Generate multi-dimensional tags for tech stack, use cases, features, and audience

Use Cases

Discover what problems this project can solve for you

Project Health Score

A
87/100
优秀
ActivityPopularityGrowthCommunityDocs
Activity 23/23

1 days since last update

Popularity 25/25

Platform Star TOP 2% · Forks 3,162

Growth 23/25

This week +1,588 ⭐ · This month +3,621 ⭐

Community 7/17

9 contributors · 0 platform comments

Docs 9/10

1 item(s) missing

View on GitHub

Project Info

Author lyogavin
Source GitHub
Period Daily
Repo Id lyogavin/airllm
Latest Commit 2026-08-06 08:40:04
First Collected 2026-08-05 08:00:54
Last Updated 2026-08-07 08:00:32

Support

If this site has been helpful, feel free to support us

WeChat donation QR code

WeChat

Alipay donation QR code

Alipay

Comments 0

Log in to post a comment

Loading comments...

Assistant