Can I run it?
Check if Llama, Qwen, Gemma or Bielik will run on your graphics card, how much memory it takes and roughly how fast it answers.
Google: Gemma 3 12B · Q4_K_M · NVIDIA GeForce RTX 4090
Not enough data
- Memory needed
- —
- Expected speed
- —
We do not have the layer data of this model yet, so only part of the estimate is shown.
Estimate for llama.cpp and similar apps (LM Studio, Ollama). Speed is for writing the answer; real results depend on drivers and settings.
Quantization shrinks the model. Q4_K_M keeps most of the quality at about a third of the full size.
Context is how much text the model keeps in mind. More context needs more memory.
Sources bits per weight from llama.cpp, card memory and bandwidth from Wikipedia (CC BY-SA 4.0) and maker specs.