What if your GPU has only 4 GB of VRAM, but the LLM you want to run needs tens or even hundreds of gigabytes?
Normally, the answer sounds simple:
You need a bigger GPU.
But what if we don’t load the entire model into GPU memory in the first place?
That’s the interesting idea behind AirLLM.
AirLLM is an open-source inference framework designed to dramatically reduce the GPU memory required to run large language models. Its core approach is to load model layers incrementally instead of keeping the entire model in GPU memory at the same time.
The project’s current repository advertises running models such as 70B-class models with very small GPU memory, and more recent versions extend this approach to extremely large models, including DeepSeek-V3 and large Mixture-of-Experts models. These numbers are project-reported measurements, so they should be treated as demonstrations of what the implementation can achieve rather than universal hardware requirements.
Let’s understand how it works. 🧠
1. The GPU memory problem
Let’s start with a simple example.
Imagine we want to run a 70-billion-parameter model.
A rough FP16 calculation is:
70 billion parameters × 2 bytes
≈ 140 GB
That’s already far beyond the VRAM of a typical consumer GPU.
Even a GPU with 24 GB of VRAM cannot simply load a 140 GB FP16 model into memory.
And this is only the model weights.
During inference, additional memory can be required for things such as:
- Activations
- KV cache
- Temporary tensors
- Framework overhead
So the actual memory requirement can be considerably more complicated than simply calculating:
parameters × bytes per parameter
Traditionally, we solve this problem using techniques such as:
- Quantization
- Model sharding across multiple GPUs
- CPU offloading
- Model compression
- Smaller models
These approaches are useful, but AirLLM attacks the problem from another direction:
What if we only put the part of the model that we currently need into GPU memory?
2. The key idea: don’t load the whole model 🚪
A Transformer model isn’t one giant mathematical operation.
It is composed of many layers.
A simplified model might look like this:
Input
│
▼
Embedding
│
▼
Layer 1
│
▼
Layer 2
│
▼
Layer 3
│
▼
...
│
▼
Layer N
│
▼
Output
A conventional inference implementation may keep most or all model weights in GPU memory.
AirLLM takes a different approach.
Conceptually:
Storage
│
│ load
▼
┌─────────────┐
│ Layer 1 │
└─────────────┘
│
▼
GPU
│
▼
Compute
│
▼
unload / replace
│
▼
┌─────────────┐
│ Layer 2 │
└─────────────┘
│
▼
GPU
│
▼
Compute
│
▼
...
Only a portion of the model needs to occupy GPU memory at any one time.
This is the fundamental trick.
3. Think of it like paging in an operating system
If you’re a software engineer, there is a familiar analogy:
Virtual memory.
An operating system doesn’t necessarily keep every page of a process in physical RAM simultaneously.
Instead, it can move pages between:
Disk / SSD
↕
RAM
AirLLM applies a somewhat similar idea to LLM inference:
Model storage
↕
System memory
↕
GPU memory
Instead of asking:
“How much GPU memory does the entire model need?”
we can ask:
“How much GPU memory does the currently active portion of the model need?”
That’s a much smaller number.
4. A simplified AirLLM inference flow
Imagine a model containing 80 transformer layers.
Instead of:
GPU:
Layer 1
Layer 2
Layer 3
...
Layer 80
❌ Everything must fit
the concept is closer to:
GPU:
Layer 1
Layer 2
Layer 3
Compute
↓
Layer 4
Layer 5
Layer 6
Compute
↓
Layer 7
Layer 8
Layer 9
Compute
↓
...
The exact implementation and scheduling are more sophisticated than this simplified diagram, but this mental model captures the important idea.
The project also introduced prefetching to overlap model loading and computation, reducing some of the performance penalty associated with moving layers.
5. Why this dramatically reduces VRAM requirements
Suppose a model contains:
70B parameters
and its complete FP16 weights occupy roughly:
140 GB
A conventional implementation might need GPU resources capable of holding a substantial portion of those weights.
With layer-wise loading, we don’t necessarily need:
140 GB VRAM
Instead, we may only need enough GPU memory for:
Current layer(s)
+
Activations
+
KV cache
+
Runtime overhead
This is why AirLLM can make surprisingly large models executable on hardware with very limited VRAM.
The AirLLM project reports, for example, 70B inference on a single 4 GB GPU without requiring quantization, distillation, or pruning in its original approach.
6. Does AirLLM magically compress the model?
No. And this is an important distinction.
There are several different techniques for making large models fit into limited memory.
Quantization
For example:
FP16 → INT8 → INT4
The weights require fewer bits.
This reduces memory consumption, but it changes the numerical representation of the model.
Pruning
Remove some parameters or connections.
Large model
↓
Remove less-important components
↓
Smaller model
Distillation
Train a smaller model to reproduce the behavior of a larger model.
Teacher model
↓
Knowledge
↓
Student model
AirLLM’s original core approach
Instead of fundamentally making the model smaller:
Huge model
↓
Keep model weights
↓
Stream portions into GPU
↓
Compute
↓
Move to next portion
That’s why AirLLM is interesting.
It changes where the model lives during inference, rather than relying solely on making the model smaller.
The current project also supports compression and quantization options, so it’s not correct to say that modern AirLLM never uses quantization. The “no quantization” description applies to its core layer-wise inference approach and specific configurations.
7. Let’s try it
AirLLM exposes a relatively simple Python interface.
The project uses Hugging Face model repositories and provides an AutoModel interface in its current codebase.
A simplified example looks like:
from airllm import AutoModel
model = AutoModel.from_pretrained(
"meta-llama/Llama-3.1-8B"
)
input_text = "Explain how MQTT works in IoT."
input_tokens = model.tokenizer(
input_text,
return_tensors="pt"
)
generation_output = model.generate(
input_tokens["input_ids"].cuda(),
max_new_tokens=100
)
output = model.tokenizer.decode(
generation_output[0]
)
print(output)
For a larger model, the model identifier can be replaced with a supported Hugging Face model.
For example, conceptually:
model = AutoModel.from_pretrained(
"meta-llama/Llama-3.1-70B"
)
However, model compatibility, dependencies, GPU backend, model format and available system memory matter. You should always check the current AirLLM documentation and the specific model’s requirements rather than assuming every Hugging Face model will work unchanged.
8. Installation
The project provides a Python package:
pip install airllm
The current package declares dependencies including PyTorch, Transformers, Accelerate, Safetensors and Hugging Face Hub.
A typical environment might start with:
python -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install airllm
Then:
from airllm import AutoModel
and start loading your selected model.
9. But there is a catch… 🐌
There is no free lunch.
If the entire model doesn’t fit into GPU memory, AirLLM has to move model data around.
And moving data costs time.
Consider:
SSD
│
│ Read
▼
RAM
│
│ Transfer
▼
GPU VRAM
│
│
▼
Compute
If we repeatedly perform this process:
Load
Compute
Unload
Load
Compute
Unload
...
the storage and memory bandwidth become important parts of inference performance.
This creates a fundamental trade-off:
Lower VRAM requirement
↓
More data movement
↓
Potentially higher latency
So AirLLM doesn’t make a 4 GB GPU suddenly perform like a 96 GB high-end GPU.
It makes a workload that previously could not fit become possible.
That’s a very different claim.
10. SSD speed suddenly becomes important
This is one of the most interesting aspects of AirLLM.
When a huge model is being streamed layer by layer, storage performance matters.
Compare:
HDD
↓
Slow random/sequential access
↓
😴
with:
NVMe SSD
↓
High bandwidth
↓
Much better streaming
↓
🚀
This means a system designed for AirLLM should not only ask:
“How much VRAM do I have?”
It should also consider:
- SSD speed
- PCIe bandwidth
- RAM capacity
- RAM bandwidth
- GPU memory bandwidth
- GPU compute performance
- CPU performance
- Model architecture
In other words:
AirLLM turns LLM inference into more of a system-engineering problem.
And that’s exactly what makes it interesting for edge engineers.
11. AirLLM and edge computing
This is where I think the concept becomes particularly interesting.
Imagine an industrial edge computer:
┌───────────────────────────────┐
│ Edge Computer │
│ │
│ CPU │
│ RAM │
│ Small GPU │
│ │
│ NVMe SSD │
│ │
│ ┌─────────────────────────┐ │
│ │ AirLLM │ │
│ │ │ │
│ │ Layer-wise inference │ │
│ └─────────────────────────┘ │
│ │
└───────────────────────────────┘
You might not have enough VRAM to run a large model traditionally.
But you could potentially trade:
GPU memory
↕
Storage + RAM + latency
This is an interesting architecture for:
- Edge AI
- Industrial gateways
- Development workstations
- Local AI experimentation
- Privacy-sensitive workloads
- Offline environments
- Research systems
- Hardware with limited GPU memory
The project itself supports CPU inference and has added support for increasingly large models over time.
12. AirLLM isn’t just for 70B anymore
This is where the current project becomes particularly interesting.
The original AirLLM story was essentially:
“Run a 70B model on a tiny GPU.”
But the project has continued evolving.
According to the current repository, recent versions support models including:
Llama
Qwen
DeepSeek
Phi
Gemma
and others
The project also reports:
DeepSeek-V3
671B parameters
~12 GB VRAM
and:
Qwen3-235B
~3 GB VRAM
for specific configurations.
The repository also reports support for Kimi K3, a 2.8T-parameter MoE model, using per-expert streaming so that only the experts selected for a token need to be loaded. Its reported end-to-end measurement uses about 3.72 GB of VRAM on an RTX 6000 Ada.
That’s a fascinating evolution.
But remember:
Parameter count alone does not determine inference cost.
A dense 100B model and a 100B MoE model can have very different compute and memory behavior.
13. MoE makes the idea even more interesting
Mixture-of-Experts models work differently from traditional dense Transformers.
Conceptually:
Input
│
▼
Router / Gate
/ | \
/ | \
Expert A Expert B Expert C
│
▼
Selected
experts
│
▼
Output
The router decides which experts should process a token.
That creates another opportunity for memory optimization.
Instead of loading every expert:
Expert A
Expert B
Expert C
Expert D
...
Expert N
the runtime can potentially stream only the experts needed for the current computation.
AirLLM’s current documentation describes this per-expert streaming approach for Kimi K3.
So the combination becomes:
MoE sparsity
+
Layer streaming
+
Expert streaming
↓
Much smaller active memory footprint
That’s a powerful idea.
14. AirLLM vs quantization
So which approach should you use?
There isn’t one universal answer.
|
Approach |
Main idea |
Memory |
Speed |
Model changes |
|---|---|---|---|---|
|
Full FP16 |
Keep original weights |
🔴 High |
🟢 Fast |
No |
|
INT8 |
Reduce weight precision |
🟡 Lower |
🟢 Usually good |
Yes |
|
INT4 |
Aggressive compression |
🟢 Much lower |
🟢/🟡 |
Yes |
|
CPU offload |
Keep some weights in RAM |
🟡 |
🟡 |
No |
|
AirLLM layer streaming |
Load layers progressively |
🟢 Very low VRAM |
🔴 Often slower |
Core approach preserves weights |
|
AirLLM + compression |
Combine approaches |
🟢 Very low |
🟡 |
Depends |
The interesting point is that these techniques don’t necessarily have to be mutually exclusive.
Modern AirLLM itself includes compression/quantization capabilities in addition to its layer-wise inference mechanism.
15. Pros 👍
Very low GPU memory requirement
This is obviously the headline feature.
A model that would otherwise require a multi-GPU system may become executable on much smaller hardware.
Preserves access to large models
If you don’t want to aggressively quantize a model, layer-wise inference provides another way to reduce the GPU memory requirement.
Useful for experimentation
You can experiment with models that your GPU normally cannot fit.
That’s particularly useful for developers and researchers.
Interesting for edge computing
Instead of assuming:
“AI = expensive GPU server”
we can consider:
Small GPU
+
Large RAM
+
Fast NVMe
+
Smart model execution
Supports very large models
The current project has moved well beyond its original 70B use case, with support claims covering models in the hundreds of billions and even trillion-plus parameter range for suitable architectures/configurations.
16. Cons and limitations 👎
It can be slow
This is probably the biggest trade-off.
If your inference repeatedly moves model data between storage, RAM and GPU, data movement becomes a bottleneck.
So:
Less VRAM
↓
More streaming
↓
Potentially more latency
Storage becomes part of the inference pipeline
A slow disk can seriously hurt performance.
NVMe is much more appropriate than relying on slow storage.
More complicated runtime behavior
Traditional inference can be conceptually:
Load model
↓
GPU
↓
Generate
Layer streaming introduces:
Storage
↓
RAM
↓
GPU
↓
Compute
↓
Next layer
↓
...
There are simply more moving parts.
Not every model is automatically supported
Model architectures, Transformers versions, checkpoint formats and dependencies matter.
You should verify compatibility before building a production system.
Not a replacement for high-end GPUs
If you have a powerful GPU with enough VRAM to comfortably run the model, traditional inference engines may provide much better latency and throughput.
AirLLM’s biggest advantage is often:
“I can run it at all.”
rather than:
“I can run it faster than a large GPU.”
17. The important engineering lesson
AirLLM demonstrates a broader principle that goes beyond LLMs.
When a workload doesn’t fit into memory, there are two fundamentally different questions:
Question 1:
How do I make the workload smaller?
versus
Question 2:
How do I avoid keeping the entire workload in memory?
The first question leads to:
Quantization
Compression
Pruning
Distillation
The second leads to:
Streaming
Paging
Offloading
Caching
Prefetching
AirLLM is an interesting example of the second category.
And this is a concept that appears everywhere in computing.
Operating systems do it.
Databases do it.
Video processing pipelines do it.
Storage systems do it.
And now we’re doing something similar with LLM inference.
18. AirLLM is really a systems-engineering story
At first glance, AirLLM looks like an AI framework.
But underneath, the interesting problem is actually:
┌──────────────┐
│ Huge Model │
└──────┬───────┘
│
▼
┌───────────────┐
│ Storage I/O │
└───────┬───────┘
│
▼
┌───────────────┐
│ System RAM │
└───────┬───────┘
│
▼
┌───────────────┐
│ GPU VRAM │
└───────┬───────┘
│
▼
┌───────────────┐
│ GPU Compute │
└───────────────┘
The optimization problem becomes:
How do we keep the GPU busy while minimizing the amount of data that has to move?
That brings us directly into:
- Memory hierarchy
- DMA
- PCIe
- NVMe
- Prefetching
- Caching
- Tensor computation
- Model architecture
- GPU utilization
- I/O scheduling
That’s why AirLLM is particularly interesting from an edge-computing perspective.
19. The big trade-off
The entire concept can be summarized with one equation-like idea:
Traditional approach:
More model
↓
More VRAM
↓
More expensive hardware
AirLLM-style approach:
More model
↓
More streaming
↓
More I/O + latency
↓
Less VRAM required
We are essentially exchanging:
GPU memory → time + bandwidth
And sometimes that’s a very good trade.
For example:
If your requirement is:
“Generate an answer in 100 ms.”
AirLLM may not be the right architecture.
But if your requirement is:
“I have a 4–8 GB GPU and I really want to experiment with a model that normally needs much more memory.”
then the trade-off becomes much more interesting.
20. Final thoughts 🚀
AirLLM doesn’t violate the laws of computer architecture.
It simply changes the question.
Instead of asking:
“How can I fit this enormous model into my GPU?”
it asks:
“Why do I need the entire model in my GPU at the same time?”
That small change in perspective leads to a completely different architecture.
The model can live primarily outside the GPU:
Huge LLM
│
┌─────┴─────┐
│ │
Storage RAM
│ │
└─────┬─────┘
│
Stream what
we currently need
│
▼
GPU VRAM
│
▼
Compute
│
▼
Next portion
And that is the real lesson behind AirLLM:
When hardware resources are limited, don’t always try to make the workload smaller. Sometimes, redesign the way the workload moves through the system.
For edge AI, local AI, and resource-constrained computing, that’s a very powerful idea. 💡
🔗 Resources
- AirLLM GitHub repository — Official project repository and current implementation.
- AirLLM package on PyPI — Python package installation.
- Hugging Face — Model repository commonly used with AirLLM.
#airllm #LLM #GenerativeAI #AI #MachineLearning #EdgeAI #EdgeComputing #LocalAI #GPU #CUDA #Python #OpenSource #AIEngineering #LLMInference #Transformers