Tools · Model fit
Will this model fit?
The question everyone asks before buying a card. Pick your VRAM, a model and a quantisation, and see what actually loads, with headroom for context included.
Verdict
-
-
Model weights
-
+ context/kv
~2 GB
Headroom
-
A rule-of-thumb. Real usage depends on context length, the runtime (llama.cpp, vLLM, Ollama) and kv-cache settings, but it's close enough to know before you buy.