AirLLM: How to Run Giant LLMs When Your GPU Is Way Too Small 🚀

What if your GPU has only 4 GB of VRAM, but the LLM you want to run needs tens or even hundreds of gigabytes?

Normally, the answer sounds simple:

You need a bigger GPU.

But what if we don’t load the entire model into GPU memory in the first place?

That’s the interesting idea behind AirLLM.

AirLLM is an open-source inference framework designed to dramatically reduce the GPU memory required to run large language models. Its core approach is to load model layers incrementally instead of keeping the entire model in GPU memory at the same time.

 

The project’s current repository advertises running models such as 70B-class models with very small GPU memory, and more recent versions extend this approach to extremely large models, including DeepSeek-V3 and large Mixture-of-Experts models. These numbers are project-reported measurements, so they should be treated as demonstrations of what the implementation can achieve rather than universal hardware requirements.

Let’s understand how it works. 🧠


1. The GPU memory problem

Let’s start with a simple example.

Imagine we want to run a 70-billion-parameter model.

A rough FP16 calculation is:

70 billion parameters × 2 bytes
≈ 140 GB

That’s already far beyond the VRAM of a typical consumer GPU.

Even a GPU with 24 GB of VRAM cannot simply load a 140 GB FP16 model into memory.

And this is only the model weights.

During inference, additional memory can be required for things such as:

  • Activations
  • KV cache
  • Temporary tensors
  • Framework overhead

So the actual memory requirement can be considerably more complicated than simply calculating:

parameters × bytes per parameter

Traditionally, we solve this problem using techniques such as:

  • Quantization
  • Model sharding across multiple GPUs
  • CPU offloading
  • Model compression
  • Smaller models

These approaches are useful, but AirLLM attacks the problem from another direction:

What if we only put the part of the model that we currently need into GPU memory?


2. The key idea: don’t load the whole model 🚪

A Transformer model isn’t one giant mathematical operation.

It is composed of many layers.

A simplified model might look like this:

Input
  │
  ▼
Embedding
  │
  ▼
Layer 1
  │
  ▼
Layer 2
  │
  ▼
Layer 3
  │
  ▼
...
  │
  ▼
Layer N
  │
  ▼
Output

A conventional inference implementation may keep most or all model weights in GPU memory.

AirLLM takes a different approach.

Conceptually:

              Storage
                 │
                 │ load
                 ▼
          ┌─────────────┐
          │   Layer 1   │
          └─────────────┘
                 │
                 ▼
              GPU
                 │
                 ▼
             Compute
                 │
                 ▼
        unload / replace
                 │
                 ▼
          ┌─────────────┐
          │   Layer 2   │
          └─────────────┘
                 │
                 ▼
              GPU
                 │
                 ▼
             Compute
                 │
                 ▼
                ...

Only a portion of the model needs to occupy GPU memory at any one time.

This is the fundamental trick.


3. Think of it like paging in an operating system

If you’re a software engineer, there is a familiar analogy:

Virtual memory.

An operating system doesn’t necessarily keep every page of a process in physical RAM simultaneously.

Instead, it can move pages between:

Disk / SSD
     ↕
RAM

AirLLM applies a somewhat similar idea to LLM inference:

Model storage
     ↕
System memory
     ↕
GPU memory

Instead of asking:

“How much GPU memory does the entire model need?”

we can ask:

“How much GPU memory does the currently active portion of the model need?”

That’s a much smaller number.


4. A simplified AirLLM inference flow

Imagine a model containing 80 transformer layers.

Instead of:

GPU:

Layer 1
Layer 2
Layer 3
...
Layer 80

❌ Everything must fit

the concept is closer to:

GPU:

Layer 1
Layer 2
Layer 3

Compute
   ↓

Layer 4
Layer 5
Layer 6

Compute
   ↓

Layer 7
Layer 8
Layer 9

Compute
   ↓

...

The exact implementation and scheduling are more sophisticated than this simplified diagram, but this mental model captures the important idea.

The project also introduced prefetching to overlap model loading and computation, reducing some of the performance penalty associated with moving layers.


5. Why this dramatically reduces VRAM requirements

Suppose a model contains:

70B parameters

and its complete FP16 weights occupy roughly:

140 GB

A conventional implementation might need GPU resources capable of holding a substantial portion of those weights.

With layer-wise loading, we don’t necessarily need:

140 GB VRAM

Instead, we may only need enough GPU memory for:

Current layer(s)
+
Activations
+
KV cache
+
Runtime overhead

This is why AirLLM can make surprisingly large models executable on hardware with very limited VRAM.

The AirLLM project reports, for example, 70B inference on a single 4 GB GPU without requiring quantization, distillation, or pruning in its original approach.


6. Does AirLLM magically compress the model?

No. And this is an important distinction.

There are several different techniques for making large models fit into limited memory.

Quantization

For example:

FP16 → INT8 → INT4

The weights require fewer bits.

This reduces memory consumption, but it changes the numerical representation of the model.

Pruning

Remove some parameters or connections.

Large model
    ↓
Remove less-important components
    ↓
Smaller model

Distillation

Train a smaller model to reproduce the behavior of a larger model.

Teacher model
      ↓
Knowledge
      ↓
Student model

AirLLM’s original core approach

Instead of fundamentally making the model smaller:

Huge model
     ↓
Keep model weights
     ↓
Stream portions into GPU
     ↓
Compute
     ↓
Move to next portion

That’s why AirLLM is interesting.

It changes where the model lives during inference, rather than relying solely on making the model smaller.

The current project also supports compression and quantization options, so it’s not correct to say that modern AirLLM never uses quantization. The “no quantization” description applies to its core layer-wise inference approach and specific configurations.


7. Let’s try it

AirLLM exposes a relatively simple Python interface.

The project uses Hugging Face model repositories and provides an AutoModel interface in its current codebase.

A simplified example looks like:

from airllm import AutoModel

model = AutoModel.from_pretrained(
    "meta-llama/Llama-3.1-8B"
)

input_text = "Explain how MQTT works in IoT."

input_tokens = model.tokenizer(
    input_text,
    return_tensors="pt"
)

generation_output = model.generate(
    input_tokens["input_ids"].cuda(),
    max_new_tokens=100
)

output = model.tokenizer.decode(
    generation_output[0]
)

print(output)

For a larger model, the model identifier can be replaced with a supported Hugging Face model.

For example, conceptually:

model = AutoModel.from_pretrained(
    "meta-llama/Llama-3.1-70B"
)

However, model compatibility, dependencies, GPU backend, model format and available system memory matter. You should always check the current AirLLM documentation and the specific model’s requirements rather than assuming every Hugging Face model will work unchanged.


8. Installation

The project provides a Python package:

pip install airllm

The current package declares dependencies including PyTorch, Transformers, Accelerate, Safetensors and Hugging Face Hub.

A typical environment might start with:

python -m venv .venv

source .venv/bin/activate

pip install --upgrade pip

pip install airllm

Then:

from airllm import AutoModel

and start loading your selected model.


9. But there is a catch… 🐌

There is no free lunch.

If the entire model doesn’t fit into GPU memory, AirLLM has to move model data around.

And moving data costs time.

Consider:

SSD
 │
 │  Read
 ▼
RAM
 │
 │  Transfer
 ▼
GPU VRAM
 │
 │
 ▼
Compute

If we repeatedly perform this process:

Load
Compute
Unload
Load
Compute
Unload
...

the storage and memory bandwidth become important parts of inference performance.

This creates a fundamental trade-off:

Lower VRAM requirement
        ↓
More data movement
        ↓
Potentially higher latency

So AirLLM doesn’t make a 4 GB GPU suddenly perform like a 96 GB high-end GPU.

It makes a workload that previously could not fit become possible.

That’s a very different claim.


10. SSD speed suddenly becomes important

This is one of the most interesting aspects of AirLLM.

When a huge model is being streamed layer by layer, storage performance matters.

Compare:

HDD
 ↓
Slow random/sequential access
 ↓
😴

with:

NVMe SSD
 ↓
High bandwidth
 ↓
Much better streaming
 ↓
🚀

This means a system designed for AirLLM should not only ask:

“How much VRAM do I have?”

It should also consider:

  • SSD speed
  • PCIe bandwidth
  • RAM capacity
  • RAM bandwidth
  • GPU memory bandwidth
  • GPU compute performance
  • CPU performance
  • Model architecture

In other words:

AirLLM turns LLM inference into more of a system-engineering problem.

And that’s exactly what makes it interesting for edge engineers.


11. AirLLM and edge computing

This is where I think the concept becomes particularly interesting.

Imagine an industrial edge computer:

┌───────────────────────────────┐
│        Edge Computer          │
│                               │
│  CPU                         │
│  RAM                         │
│  Small GPU                   │
│                               │
│  NVMe SSD                    │
│                               │
│  ┌─────────────────────────┐ │
│  │       AirLLM            │ │
│  │                         │ │
│  │ Layer-wise inference    │ │
│  └─────────────────────────┘ │
│                               │
└───────────────────────────────┘

You might not have enough VRAM to run a large model traditionally.

But you could potentially trade:

GPU memory
      ↕
Storage + RAM + latency

This is an interesting architecture for:

  • Edge AI
  • Industrial gateways
  • Development workstations
  • Local AI experimentation
  • Privacy-sensitive workloads
  • Offline environments
  • Research systems
  • Hardware with limited GPU memory

The project itself supports CPU inference and has added support for increasingly large models over time.


12. AirLLM isn’t just for 70B anymore

This is where the current project becomes particularly interesting.

The original AirLLM story was essentially:

“Run a 70B model on a tiny GPU.”

But the project has continued evolving.

According to the current repository, recent versions support models including:

Llama
Qwen
DeepSeek
Phi
Gemma
and others

The project also reports:

DeepSeek-V3
671B parameters
~12 GB VRAM

and:

Qwen3-235B
~3 GB VRAM

for specific configurations.

The repository also reports support for Kimi K3, a 2.8T-parameter MoE model, using per-expert streaming so that only the experts selected for a token need to be loaded. Its reported end-to-end measurement uses about 3.72 GB of VRAM on an RTX 6000 Ada.

That’s a fascinating evolution.

But remember:

Parameter count alone does not determine inference cost.

A dense 100B model and a 100B MoE model can have very different compute and memory behavior.


13. MoE makes the idea even more interesting

Mixture-of-Experts models work differently from traditional dense Transformers.

Conceptually:

                 Input
                   │
                   ▼
             Router / Gate
             /     |      \
            /      |       \
        Expert A Expert B Expert C
                    │
                    ▼
                Selected
                 experts
                    │
                    ▼
                  Output

The router decides which experts should process a token.

That creates another opportunity for memory optimization.

Instead of loading every expert:

Expert A
Expert B
Expert C
Expert D
...
Expert N

the runtime can potentially stream only the experts needed for the current computation.

AirLLM’s current documentation describes this per-expert streaming approach for Kimi K3.

So the combination becomes:

MoE sparsity
      +
Layer streaming
      +
Expert streaming
      ↓
Much smaller active memory footprint

That’s a powerful idea.


14. AirLLM vs quantization

So which approach should you use?

There isn’t one universal answer.

Approach

Main idea

Memory

Speed

Model changes

Full FP16

Keep original weights

🔴 High

🟢 Fast

No

INT8

Reduce weight precision

🟡 Lower

🟢 Usually good

Yes

INT4

Aggressive compression

🟢 Much lower

🟢/🟡

Yes

CPU offload

Keep some weights in RAM

🟡

🟡

No

AirLLM layer streaming

Load layers progressively

🟢 Very low VRAM

🔴 Often slower

Core approach preserves weights

AirLLM + compression

Combine approaches

🟢 Very low

🟡

Depends

The interesting point is that these techniques don’t necessarily have to be mutually exclusive.

Modern AirLLM itself includes compression/quantization capabilities in addition to its layer-wise inference mechanism.


15. Pros 👍

Very low GPU memory requirement

This is obviously the headline feature.

A model that would otherwise require a multi-GPU system may become executable on much smaller hardware.

Preserves access to large models

If you don’t want to aggressively quantize a model, layer-wise inference provides another way to reduce the GPU memory requirement.

Useful for experimentation

You can experiment with models that your GPU normally cannot fit.

That’s particularly useful for developers and researchers.

Interesting for edge computing

Instead of assuming:

“AI = expensive GPU server”

we can consider:

Small GPU
+
Large RAM
+
Fast NVMe
+
Smart model execution

Supports very large models

The current project has moved well beyond its original 70B use case, with support claims covering models in the hundreds of billions and even trillion-plus parameter range for suitable architectures/configurations.


16. Cons and limitations 👎

It can be slow

This is probably the biggest trade-off.

If your inference repeatedly moves model data between storage, RAM and GPU, data movement becomes a bottleneck.

So:

Less VRAM
    ↓
More streaming
    ↓
Potentially more latency

Storage becomes part of the inference pipeline

A slow disk can seriously hurt performance.

NVMe is much more appropriate than relying on slow storage.

More complicated runtime behavior

Traditional inference can be conceptually:

Load model
     ↓
GPU
     ↓
Generate

Layer streaming introduces:

Storage
  ↓
RAM
  ↓
GPU
  ↓
Compute
  ↓
Next layer
  ↓
...

There are simply more moving parts.

Not every model is automatically supported

Model architectures, Transformers versions, checkpoint formats and dependencies matter.

You should verify compatibility before building a production system.

Not a replacement for high-end GPUs

If you have a powerful GPU with enough VRAM to comfortably run the model, traditional inference engines may provide much better latency and throughput.

AirLLM’s biggest advantage is often:

“I can run it at all.”

rather than:

“I can run it faster than a large GPU.”


17. The important engineering lesson

AirLLM demonstrates a broader principle that goes beyond LLMs.

When a workload doesn’t fit into memory, there are two fundamentally different questions:

Question 1:
How do I make the workload smaller?

             versus

Question 2:
How do I avoid keeping the entire workload in memory?

The first question leads to:

Quantization
Compression
Pruning
Distillation

The second leads to:

Streaming
Paging
Offloading
Caching
Prefetching

AirLLM is an interesting example of the second category.

And this is a concept that appears everywhere in computing.

Operating systems do it.

Databases do it.

Video processing pipelines do it.

Storage systems do it.

And now we’re doing something similar with LLM inference.


18. AirLLM is really a systems-engineering story

At first glance, AirLLM looks like an AI framework.

But underneath, the interesting problem is actually:

                 ┌──────────────┐
                 │ Huge Model   │
                 └──────┬───────┘
                        │
                        ▼
                ┌───────────────┐
                │ Storage I/O   │
                └───────┬───────┘
                        │
                        ▼
                ┌───────────────┐
                │ System RAM    │
                └───────┬───────┘
                        │
                        ▼
                ┌───────────────┐
                │ GPU VRAM      │
                └───────┬───────┘
                        │
                        ▼
                ┌───────────────┐
                │ GPU Compute   │
                └───────────────┘

The optimization problem becomes:

How do we keep the GPU busy while minimizing the amount of data that has to move?

That brings us directly into:

  • Memory hierarchy
  • DMA
  • PCIe
  • NVMe
  • Prefetching
  • Caching
  • Tensor computation
  • Model architecture
  • GPU utilization
  • I/O scheduling

That’s why AirLLM is particularly interesting from an edge-computing perspective.


19. The big trade-off

The entire concept can be summarized with one equation-like idea:

Traditional approach:

More model
      ↓
More VRAM
      ↓
More expensive hardware


AirLLM-style approach:

More model
      ↓
More streaming
      ↓
More I/O + latency
      ↓
Less VRAM required

We are essentially exchanging:

GPU memory → time + bandwidth

And sometimes that’s a very good trade.

For example:

If your requirement is:

“Generate an answer in 100 ms.”

AirLLM may not be the right architecture.

But if your requirement is:

“I have a 4–8 GB GPU and I really want to experiment with a model that normally needs much more memory.”

then the trade-off becomes much more interesting.


20. Final thoughts 🚀

AirLLM doesn’t violate the laws of computer architecture.

It simply changes the question.

Instead of asking:

“How can I fit this enormous model into my GPU?”

it asks:

“Why do I need the entire model in my GPU at the same time?”

That small change in perspective leads to a completely different architecture.

The model can live primarily outside the GPU:

             Huge LLM
                │
          ┌─────┴─────┐
          │           │
        Storage      RAM
          │           │
          └─────┬─────┘
                │
         Stream what
         we currently need
                │
                ▼
             GPU VRAM
                │
                ▼
             Compute
                │
                ▼
          Next portion

And that is the real lesson behind AirLLM:

When hardware resources are limited, don’t always try to make the workload smaller. Sometimes, redesign the way the workload moves through the system.

For edge AI, local AI, and resource-constrained computing, that’s a very powerful idea. 💡


🔗 Resources


#airllm #LLM #GenerativeAI #AI #MachineLearning #EdgeAI #EdgeComputing #LocalAI #GPU #CUDA #Python #OpenSource #AIEngineering #LLMInference #Transformers

Post a Comment

Previous Post Next Post