Can I run it?
Check if Llama, Qwen, Gemma or Bielik will run on your graphics card, how much memory it takes and roughly how fast it answers.
Choose a model and your graphics card
You will see right away whether it fits, how much memory it needs and roughly how fast it answers.
- Pick a model with public weights (18 in the list).
- Pick your card, or "No graphics card" to use the processor and RAM.
- Change quantization or context to see how the memory changes.
Quantization shrinks the model. Q4_K_M keeps most of the quality at about a third of the full size.
Context is how much text the model keeps in mind. More context needs more memory.
Sources bits per weight from llama.cpp, card memory and bandwidth from Wikipedia (CC BY-SA 4.0) and maker specs.