Can I run it?
Check if Llama, Qwen, Gemma or Bielik will run on your graphics card, how much memory it takes and roughly how fast it answers.
Qwen: Qwen3 8B · Q4_K_M · NVIDIA GeForce RTX 3060
Runs
- Model weights 4.7 GB
- Context memory 1.1 GB
- Buffers 0.8 GB
- Card memory you can use 12.0 GB
- Memory needed
- 6.5 GB
- Expected speed
- 43 to 61 tokens/s
- Layers on the card
- 36 / 36
Estimate for llama.cpp and similar apps (LM Studio, Ollama). Speed is for writing the answer; real results depend on drivers and settings.
Quantization shrinks the model. Q4_K_M keeps most of the quality at about a third of the full size.
Context is how much text the model keeps in mind. More context needs more memory.
Sources bits per weight from llama.cpp, card memory and bandwidth from Wikipedia (CC BY-SA 4.0) and maker specs.