What can a NVIDIA T4 run locally?
The NVIDIA T4 has 16 GB of VRAM. At 4-bit quantization it can run LLMs up to roughly 21B parameters. Below, every popular model scored against it — verdict, best quantization, and how to run it.
Popular AI models on a T4
| Model | Size | Verdict | Est. speed | Best way to run |
|---|---|---|---|---|
| Qwen2.5-0.5B | 0.5B | Runs | ~150 tok/s | Run on your GPU with FP16 |
| GPT-2 | 124M | Runs | ~320 tok/s | Run on your GPU with FP16 |
| Phi-3-mini | 3.8B | Runs | ~28 tok/s | Run on your GPU with FP16 |
| Whisper-large-v3 | 1.5B | Runs | — | Run on CPU or GPU with the standard library |
| all-MiniLM-L6 | 22M | Runs | — | Run on CPU or GPU with the standard library |
| SDXL | 3.5B | Runs | — | Run on your GPU with Diffusers / ComfyUI |
| Qwen2.5-Coder | 7B | Runs | ~30 tok/s | Run on your GPU with Q8 |
| Mistral-7B | 7B | Runs | ~30 tok/s | Run on your GPU with Q8 |
| Llama-3.1-8B | 8B | Runs | ~27 tok/s | Run on your GPU with Q8 |
| Gemma-2-9B | 9B | Runs | ~24 tok/s | Run on your GPU with Q8 |
| Qwen2.5-14B | 14B | Runs | ~27 tok/s | Run on your GPU with Q4 |
| DeepSeek-R1-7B | 7B | Runs | ~30 tok/s | Run on your GPU with Q8 |
| Mixtral-8x7B | 47B | Workarounds | ~2.5 tok/s | Split across GPU + RAM with Q4 (offload) |
| Qwen2.5-32B | 32B | Workarounds | ~1.9 tok/s | Split across GPU + RAM with Q8 (offload) |
| Llama-3.1-70B | 70B | Workarounds | ~2.1 tok/s | Split across GPU + RAM with Q3 (offload) |
| FLUX.1-dev | 12B | Too big | — | Use CPU offload or a cloud GPU |
Est. speed = rough single-stream generation (tokens/sec) at the best-fitting quant, based on the NVIDIA T4's 320 GB/s memory bandwidth. Real speed varies with the runtime, context length, and batch size.
NVIDIA T4 — frequently asked questions
Can a NVIDIA T4 run a 7B LLM?
Yes, comfortably. A 7–8B model at 4-bit needs only ~5 GB, well within the NVIDIA T4's 16 GB.
How fast will a 7B model run on a NVIDIA T4?
Roughly ~30 tok/s for a 7–8B model at 4-bit, and about ~29 tok/s for a 13B — these are estimates for single-stream generation and vary with the runtime and context length. Speed scales with the NVIDIA T4's 320 GB/s of memory bandwidth.
Can a NVIDIA T4 run a 13B model?
Yes, comfortably on a 16 GB card. 13B at 4-bit needs ~9 GB.
Can a NVIDIA T4 run a 70B model like Llama 3 70B?
Yes, with quantization or CPU offload. A 70B model needs ~40 GB even at 4-bit; on 16 GB you'd need GPU+CPU offload or a smaller quant.
What's the largest LLM a NVIDIA T4 can run?
Roughly a 21B-parameter model at 4-bit quantization fits in 16 GB of VRAM. Larger models still run via GPU+CPU offload, just slower.