🌱⚡
GPU Garden
Grow Your Own Compute
Link copied to clipboard! Ready to share.

Can My GPU Run Open-Weight AI?

Accurate, client-side VRAM estimation for LLaMA 3.1, Qwen 2.5, DeepSeek, and FLUX. Accounts for quantized weights, Grouped Query Attention (GQA) KV-cache, and CUDA driver buffers.

Quick Scenarios:

GGUF / Ollama / vLLM
8k tokens
KV Cache Precision:
Hardware Verdict
Runs Comfortably

Total Required VRAM
7.64 GB
GPU VRAM Capacity
8.00 GB
Available Headroom
+0.36 GB
Estimated Throughput
~68 tok/s
VRAM Allocation Breakdown 100% capacity
Weights: 4.74 GB
KV Cache: 1.00 GB
CUDA/Buffer: 1.90 GB
Headroom: 0.36 GB
1. Model Weights Formula 4.74 GB
Params * (BPW / 8) * overhead
8.03B params × (4.50 BPW / 8) × 1.05 overhead = 4.74 GB
2. KV-Cache Formula 1.00 GB
2 * Layers * KV_Heads * Head_Dim * Context_Tokens * Precision_Bytes
2 × 32 layers × 8 KV heads × 128 dim × 8,192 ctx × 2B (FP16) = 1.00 GB
3. Transient Activation Buffer & Driver Baseline (1.50 GB) 1.90 GB
Transient Activation Buffer & Driver Baseline (1.50 GB)
Activations (0.40 GB) + Driver Baseline (1.50 GB) = 1.90 GB
4. Total Required VRAM vs Available Capacity & Headroom 7.64 GB
Total Required VRAM vs Available Capacity & Headroom
Required: 4.74 GB + 1.00 GB + 1.90 GB = 7.64 GB
Available: 8.00 GB Headroom: +0.36 GB
Reference Cheatsheet

Popular GPU VRAM Compatibility & Model Sweet Spots

Find the exact memory requirements for open-weights AI on your hardware. Click any card below to test its parameters and KV-cache headroom instantly in the interactive matchmaker.

Updated for LLaMA 3.2, Qwen 2.5 & FLUX.1
24GB GDDR6X 1,008 GB/s

NVIDIA GeForce RTX 4090

The consumer flagship champion. Runs Qwen 2.5 Coder 32B at Q4 with 32k context, FLUX.1 [dev] at uncompressed FP16, and full LLaMA 3.1 8B FP16 at 120+ tokens/sec.

Amazon →
48GB VRAM (Dual) 1,872+ GB/s

Dual NVIDIA RTX 3090 / 4090 Rig

The home-lab gold standard for r/LocalLLaMA. 48GB combined VRAM comfortably fits Meta LLaMA 3.3 70B and Qwen 2.5 72B at Q4_K_M with 32k context without CPU offloading.

Amazon →
16GB GDDR6X 672 GB/s

NVIDIA GeForce RTX 4070 Ti Super

The top 16GB price-to-performance card. Features a 256-bit memory bus ideal for Qwen 2.5 14B, LLaMA 3.2 11B Vision, Mistral Small 24B (Q4), and FLUX Schnell.

Amazon →
12GB GDDR6 360 GB/s

NVIDIA GeForce RTX 3060 12GB

The undisputed budget king under $300. Its generous 12GB VRAM buffer handles LLaMA 3.1 8B with huge 64k context windows and 14B models at Q4_K_M with zero OOM errors.

Amazon →
~48GB Allocatable 400 GB/s

Apple M3 / M4 Max (64GB Unified)

The silent developer workstation. macOS Metal architecture allocates ~48GB unified RAM directly to llama.cpp and MLX, executing 70B models at Q4_K_M with full 128k context support.

Amazon →
16GB GDDR6 288 GB/s

NVIDIA GeForce RTX 4060 Ti 16GB

The lowest-cost ticket to 16GB modern VRAM (~$449). Operates at an ultra-cool 165W TDP, running FLUX.1 Schnell NF4 and Qwen 2.5 14B effortlessly.

Amazon →
Reference Matrix

AI Model VRAM Requirements Guide & Sweet-Spot Matrix

Exact VRAM requirements across 4-bit (Q4_K_M), 8-bit (Q8_0), and 16-bit (FP16) precisions with recommended hardware sweet-spots.

Quantized GGUF • AWQ • FP16 / bfloat16
Model & Family Architecture & Parameters 4-Bit (Q4_K_M) VRAM 8-Bit (Q8_0) VRAM 16-Bit (FP16) VRAM Min Recommended VRAM (8k context) Recommended Hardware Sweet Spot Action
LLaMA 3.3 70B
Meta AI • Flagship LLM
70.6B Dense GQA
128k context
~41.7 GB
46.1 GB w/ 8k KV
~77.3 GB
81.7 GB w/ 8k KV
~144.0 GB
148.4 GB w/ 8k KV
48 GB
Dual RTX 3090 / 4090
48GB VRAM (Dual GPU)
Qwen 2.5 Coder 32B
Alibaba Cloud • SOTA Coding
32.5B Dense GQA
128k context
~19.2 GB
23.1 GB w/ 8k KV
~35.6 GB
39.5 GB w/ 8k KV
~66.3 GB
70.2 GB w/ 8k KV
24 GB
RTX 4090 24GB
1,008 GB/s (Consumer King)
DeepSeek V2.5 / V3 MoE
DeepSeek • MLA MoE
236B MoE MLA
21B active params
~139.4 GB
142.5 GB w/ 8k KV
~258.3 GB
261.4 GB w/ 8k KV
~481.4 GB
484.6 GB w/ 8k KV
64 GB+ (Offload)
Apple M3 Max 64GB
Unified RAM / Dual 3090
Mistral Small 24B
Mistral AI • Math & Reasoning
23.6B Dense
128k context
~14.2 GB
17.8 GB w/ 8k KV
~26.3 GB
29.9 GB w/ 8k KV
~49.0 GB
52.6 GB w/ 8k KV
16 GB
RTX 4070 Ti Super 16GB
672 GB/s (256-bit bus)
LLaMA 3.2 11B Vision
Meta AI • Multimodal Vision
13.8B Multimodal
128k context
~8.2 GB
11.1 GB w/ 8k KV
~15.1 GB
18.0 GB w/ 8k KV
~28.2 GB
31.1 GB w/ 8k KV
12 GB - 16 GB
RTX 4070 Ti Super 16GB
Full 16GB Headroom
LLaMA 3.1 8B
Meta AI • Open Benchmark
8.03B Dense
128k context
~4.7 GB
7.6 GB w/ 8k KV
~8.8 GB
11.7 GB w/ 8k KV
~16.4 GB
19.3 GB w/ 8k KV
8 GB
RTX 3060 12GB
Budget King (<$300)
FLUX.1 [dev]
Black Forest Labs • SOTA DiT
12.0B DiT
+4.9B T5-XXL / CLIP
~10.2 GB
15.5 GB w/ Act
~18.7 GB
24.0 GB w/ Act
~34.6 GB
39.9 GB w/ Act
16 GB (Q4) / 24 GB
RTX 4090 24GB
Native FP16 / FP8
SDXL 1.0
Stability AI • UNet 1024px
2.6B UNet
+0.8B CLIP-L/G
~3.3 GB
7.0 GB w/ Act
~4.7 GB
8.4 GB w/ Act
~7.1 GB
10.8 GB w/ Act
8 GB - 12 GB
RTX 3060 12GB
Full LoRA + ControlNet Headroom