Can I run it?
Check if Llama, Qwen, Gemma or Bielik will run on your graphics card, how much memory it takes and roughly how fast it answers.
No local models yet
The list of models with public weights fills in after the daily refresh from Hugging Face.
Quantization shrinks the model. Q4_K_M keeps most of the quality at about a third of the full size.
Context is how much text the model keeps in mind. More context needs more memory.
Sources bits per weight from llama.cpp, card memory and bandwidth from Wikipedia (CC BY-SA 4.0) and maker specs.