LLM VRAM requirements: what fits on 8, 16, 24, 48 and 80GB

LLM VRAM requirements: what fits on 8, 16, 24, 48 and 80GB

Have you ever found yourself stuck in this question: "will this model fit on my GPU?". A lot of us have. The honest answer is always "it depends on quantization and context length," which is true but not actually useful when someone just wants to know if their card can run the model they want to try.

So here is the practical version. A real, tier-by-tier breakdown of what actually fits on the VRAM sizes people are most commonly working with, based on typical 4-bit quantization with reasonable context headroom.

Why does the same model need different amounts of VRAM on different GPUs?

Because model size is not the only variable. Two things change the actual footprint dramatically.

  • Quantization level: since dropping from full precision to 4-bit roughly quarters the memory a model's weights need

  • Context length: since every additional token of context adds to the KV cache sitting on top of the base weights

This guide assumes standard 4-bit quantization, commonly written as Q4 or Q4_K_M, with moderate context length, since that is the realistic default for most people running models locally or in production on a budget.

What actually fits on 8GB of VRAM?

Small, genuinely capable models, comfortably.

  • Llama 3.1 8B, which fits on any 8GB GPU with real room to spare

  • Qwen 3.5 9B, fitting in around 6.6GB, leaving headroom for context

  • Most 7B to 9B class models at Q4 quantization, generally landing in the 5 to 6GB range for weights alone

8GB is genuinely usable territory now, mostly thanks to newer, more efficient models in this size class rather than any change in how quantization works.

What actually fits on 16GB of VRAM?

A clear step up, mainly in model class rather than just headroom.

  • 13B to 14B class models, comfortably fitting at Q4 with room left for context

  • 8B class models, but now with substantially more headroom for longer conversations or larger context windows

  • This is generally considered the point where local LLM use starts feeling genuinely practical rather than tightly constrained

What actually fits on 24GB of VRAM?

This tier opens up a meaningfully larger class of model.

  • Dense 30B to 32B models, fitting at Q4 quantization, though somewhat tightly with limited context headroom

  • Mixture-of-experts models like Qwen3 30B-A3B, which carry 30B total parameters but only activate around 3B per token, running noticeably faster than a dense model of the same size while still fitting comfortably

Nvidia’s L4 GPU memory capacity is 24GB. It lands exactly in this range, which is part of why it has become such a common default for cost-efficient serving of 30B-class models in production, particularly MoE architectures that deliver strong performance without demanding the full compute of a much larger dense model.

What actually fits on 48GB of VRAM?

This is genuinely the practical entry point for 70B-class models, though not with much room to spare.

  • 70B models at Q4 quantization, typically landing around 40 to 42GB, fitting on a single 48GB card with modest headroom

  • The same 70B models do not fit on a single 24GB card even at aggressive quantization, making 48GB the real dividing line for this model class

  • An alternative path is combining two 24GB GPUs, though multi-GPU setups add interconnect overhead that a single larger card avoids

What actually fits on 80GB of VRAM?

This tier gives you genuine breathing room, and access to model classes that simply do not fit anywhere smaller.

  • 70B models at higher precision, such as FP8 or INT8, with real headroom left over instead of running right at the edge

  • Large mixture-of-experts models, like a 117B total parameter model with roughly 5B active parameters per token, which runs comfortably on a single 80GB GPU

  • Models like Llama 4 Scout, which need around 67GB at Q4 quantization since all experts stay resident in memory, fitting 80GB with room but not fitting a single 24GB or even 48GB card at all

Even at 80GB, this is not a limitless tier. Models in the 400 billion parameter range still require multi-GPU setups, commonly four or more 80GB cards working together, regardless of quantization.

How do you double-check these numbers for a specific model before committing to hardware?

The tiers above hold up well as a general reference, but individual models vary a little based on architecture details like the number of attention heads and layers. Before committing budget to a GPU, it is worth checking the specific model card or quantization repository you plan to use, since most popular open-weight models now list exact file sizes for each quantization level directly. That number, plus a reasonable buffer for context and runtime overhead, is a more reliable figure than a general size class alone.

This matters most right at the edge of a tier. A 70B model comfortably fitting 48GB with a short prompt can behave very differently once you push context length up toward 32K tokens or run several requests concurrently, which is exactly the kind of gap a quick model-specific check will catch before it becomes a production problem.

Quick reference: VRAM tier and practical model ceiling

VRAM tier

Practical model ceiling

Example

8GB

7B to 9B dense models

Llama 3.1 8B, Qwen 3.5 9B

16GB

13B to 14B dense models

Comfortable fit with context room

24GB

30B to 32B dense, or 30B-class MoE

Qwen3 30B-A3B

48GB

70B models at Q4

Llama 3.3 70B

80GB

70B at higher precision, or large MoE

117B total MoE, Llama 4 Scout


Where this leaves you

VRAM tiers map fairly predictably to model classes once you account for realistic quantization, which is exactly why this kind of lookup is more useful day to day than working through the full memory formula every time. Match your actual model choice to the tier that comfortably fits it, factor in context length honestly, and you will avoid both the common mistake of undersizing hardware and the equally common one of buying more memory than your actual workload will ever use.

Frequently asked question

Does more context length change these numbers significantly? 

Yes, and this is the part people forget most often. These figures assume moderate context. Long context windows, especially past 32K tokens, can add a meaningful amount of memory on top of the base weight footprint, sometimes enough to push a model that "fits" right up against the ceiling.

Is it better to run a smaller model at higher precision or a bigger model quantized down? 

Generally, a well-chosen quantized larger model outperforms a smaller model at full precision, especially at Q4 and above. The quality loss from quantization is usually smaller than the capability gap between meaningfully different model sizes.

Can I combine multiple smaller GPUs instead of buying one bigger card? 

Yes, but it is not a perfectly clean substitute. Combined VRAM across multiple cards works for many setups, but interconnect speed between GPUs adds overhead that a single larger card with the same total memory does not have, which can noticeably affect throughput on larger models.

Should I always use the most aggressive quantization to fit a bigger model?

Not automatically. Q4 quantization typically costs only 1 to 2 percent in quality compared to full precision, which is usually a fine trade-off. Going more aggressive than that, down to Q2 or Q3, starts to noticeably hurt output quality. Fitting a model technically is not the same as fitting it well.


0 Comments

No comments yet — be the first to respond.