Can I run it?

Check if Llama, Qwen, Gemma or Bielik will run on your graphics card, how much memory it takes and roughly how fast it answers.

Used when the model does not fit in the card and part of it moves to RAM.

Google: Gemma 3 12B · Q4_K_M · NVIDIA GeForce RTX 4090

Not enough data

Memory needed
—
Expected speed
—

We do not have the layer data of this model yet, so only part of the estimate is shown.

Estimate for llama.cpp and similar apps (LM Studio, Ollama). Speed is for writing the answer; real results depend on drivers and settings.

Quantization shrinks the model. Q4_K_M keeps most of the quality at about a third of the full size. Context is how much text the model keeps in mind. More context needs more memory. Sources bits per weight from llama.cpp, card memory and bandwidth from Wikipedia (CC BY-SA 4.0) and maker specs.