canirunthismodel

Can I run gpt2 on a AMD RX 7900 XT?

Yes.Runs

You can run this comfortably. On a AMD RX 7900 XT (20 GB VRAM), gpt2 needs about 1.3 GB (FP16) and should generate at roughly ~400 tok/s. Recommended: Run on your GPU with FP16.

openai-community/gpt2

Why

How to run gpt2 on a RX 7900 XT

Serve with vLLM (OpenAI-compatible API)
# Fast production serving on http://localhost:8000/v1
vllm serve openai-community/gpt2
Python (Transformers)
from transformers import pipeline
pipe = pipeline("text-generation", model="openai-community/gpt2", device_map="auto")
print(pipe("Hello", max_new_tokens=50))

Want it to run comfortably at higher precision? The smallest GPU that runs gpt2 well is the NVIDIA RTX 4090 (24 GB).

Try the full analyzer

gpt2 on other GPUs

Other models on a RX 7900 XT