Llama 3.3 70B — Apple Silicon Benchmarks
Measured inference speed for Llama 3.3 70B across 1 Apple Silicon chip. Tokens per second at multiple quantization levels. Real runs, not estimates.
Quantizations measured: Q5_K_M
1
Benchmark rows
1
Chip tiers covered
7.1
Fastest avg tok/s (M4 Max (24-core GPU))
50 GB
Minimum RAM observed
Benchmark results for Llama 3.3 70B
Rows sorted by avg tok/s descending. Click source badge to see original measurement page.
| Chip | Quant | Avg tok/s | Runtime | Source |
|---|---|---|---|---|
| M4 Max (24-core GPU) | Q5_K_M | 7.1 tok/s | — | ref |
Chips with published results for Llama 3.3 70B
Data
benchmarks.json — full dataset · models.json — model summaries · benchmarks.csv — CSV export