Can I run it?

Check if Llama, Qwen, Gemma or Bielik will run on your graphics card, how much memory it takes and roughly how fast it answers.

Used when the model does not fit in the card and part of it moves to RAM.

Meta: Llama 3.1 8B Instruct · Q4_K_M · NVIDIA GeForce RTX 4060

Runs

  • Model weights 4.6 GB
  • Context memory 1.0 GB
  • Buffers 0.8 GB
  • Card memory you can use 8.0 GB
Memory needed
6.3 GB
Expected speed
33 to 47 tokens/s
Layers on the card
32 / 32

Estimate for llama.cpp and similar apps (LM Studio, Ollama). Speed is for writing the answer; real results depend on drivers and settings.

Quantization shrinks the model. Q4_K_M keeps most of the quality at about a third of the full size. Context is how much text the model keeps in mind. More context needs more memory. Sources bits per weight from llama.cpp, card memory and bandwidth from Wikipedia (CC BY-SA 4.0) and maker specs.