What can a NVIDIA GTX 1080 Ti run locally?
The NVIDIA GTX 1080 Ti has 11 GB of VRAM. At 4-bit quantization it can run LLMs up to roughly 15B parameters. Below, every popular model scored against it — verdict, best quantization, and how to run it.
Popular AI models on a GTX 1080 Ti
| Model | Size | Verdict | Est. speed | Best way to run |
|---|---|---|---|---|
| Qwen2.5-0.5B | 0.5B | Runs | ~200 tok/s | Run on your GPU with FP16 |
| GPT-2 | 124M | Runs | ~370 tok/s | Run on your GPU with FP16 |
| Whisper-large-v3 | 1.5B | Runs | — | Run on CPU or GPU with the standard library |
| all-MiniLM-L6 | 22M | Runs | — | Run on CPU or GPU with the standard library |
| Phi-3-mini | 3.8B | Runs | ~76 tok/s | Run on your GPU with Q8 |
| Qwen2.5-Coder | 7B | Runs | ~44 tok/s | Run on your GPU with Q8 |
| Mistral-7B | 7B | Runs | ~44 tok/s | Run on your GPU with Q8 |
| Llama-3.1-8B | 8B | Runs | ~67 tok/s | Run on your GPU with Q4 |
| Gemma-2-9B | 9B | Runs | ~60 tok/s | Run on your GPU with Q4 |
| Qwen2.5-14B | 14B | Runs | ~52 tok/s | Run on your GPU with Q3 |
| DeepSeek-R1-7B | 7B | Runs | ~44 tok/s | Run on your GPU with Q8 |
| Qwen2.5-32B | 32B | Workarounds | ~3.9 tok/s | Split across GPU + RAM with Q4 (offload) |
| SDXL | 3.5B | Workarounds | — | Use CPU offload or a cloud GPU |
| Mixtral-8x7B | 47B | Workarounds | ~2.2 tok/s | Split across GPU + RAM with Q4 (offload) |
| Llama-3.1-70B | 70B | Too big | — | Use a cloud GPU or a smaller model |
| FLUX.1-dev | 12B | Too big | — | Use CPU offload or a cloud GPU |
Est. speed = rough single-stream generation (tokens/sec) at the best-fitting quant, based on the NVIDIA GTX 1080 Ti's 484 GB/s memory bandwidth. Real speed varies with the runtime, context length, and batch size.
NVIDIA GTX 1080 Ti — frequently asked questions
Can a NVIDIA GTX 1080 Ti run a 7B LLM?
Yes, comfortably. A 7–8B model at 4-bit needs only ~5 GB, well within the NVIDIA GTX 1080 Ti's 11 GB.
How fast will a 7B model run on a NVIDIA GTX 1080 Ti?
Roughly ~44 tok/s for a 7–8B model at 4-bit, and about ~43 tok/s for a 13B — these are estimates for single-stream generation and vary with the runtime and context length. Speed scales with the NVIDIA GTX 1080 Ti's 484 GB/s of memory bandwidth.
Can a NVIDIA GTX 1080 Ti run a 13B model?
Yes, comfortably on a 11 GB card. 13B at 4-bit needs ~9 GB.
Can a NVIDIA GTX 1080 Ti run a 70B model like Llama 3 70B?
No, it's too big. A 70B model needs ~40 GB even at 4-bit; on 11 GB you'd need GPU+CPU offload or a smaller quant.
What's the largest LLM a NVIDIA GTX 1080 Ti can run?
Roughly a 15B-parameter model at 4-bit quantization fits in 11 GB of VRAM. Larger models still run via GPU+CPU offload, just slower.