Meta: Llama 3.3 70B Instruct
Made by meta-llama. Open weights
- Input price
- $0.10
- per 1M tokens
- Output price
- $0.32
- per 1M tokens
- Context window
- 131K
- tokens
- Quality
- 1,274
- #273 on LMArena
- Size
- 71B
- parameters
- Downloads
- 354.9K
- on Hugging Face, last 30 days
Can I run it locally?
Estimate for the Q4_K_M version with an 8K context. Pick your own card and settings in the calculator.
- NVIDIA GeForce RTX 40608 GBRuns, but slowly1.0 to 1.4 tokens/s
- NVIDIA GeForce RTX 306012 GBRuns, but slowly1.1 to 1.5 tokens/s
- NVIDIA GeForce RTX 4060 Ti 16 GB16 GBRuns, but slowly1.2 to 1.6 tokens/s
- NVIDIA GeForce RTX 409024 GBRuns, but slowly1.7 to 2.4 tokens/s
- NVIDIA GeForce RTX 509032 GBRuns, but slowly2.8 to 3.9 tokens/s
- Apple M1 Pro 32 GB32 GBWill not run—
Prices: OpenRouter, refreshed 8 Oct 2026. Quality: LMArena (CC BY 4.0), leaderboard of 2 Oct 2026. Model page on Hugging Face