Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
🏆
Rankings

LLM Leaderboard

Real measurements of every tested language model on real, named hardware – rated by raw speed (performance in tokens per second, prefill and time to first token) and by practical task quality in complete agent and chat runs (harness). Pick a benchmark type below or filter by model, maker and hardware to see exactly what is tested and how the results are produced.

⚡ Performance (tok/s)🤖 Harness quality👥 Concurrency🖥️ real hardware
Reset
Metric:
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
226,59 tok/s TG
Prefill 9.526 · TTFT 3.488 ms
10×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ4_K_M
2Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
215,97 tok/s TG
Prefill 9.775 · TTFT 3.283 ms
10×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ8_0
3Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
173,08 tok/s TG
Prefill 7.158 · TTFT 2.024 ms
5×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ4_K_M
4Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
158,90 tok/s TG
Prefill 7.616 · TTFT 1.908 ms
5×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ8_0
5Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
100,62 tok/s TG
Prefill 3.646 · TTFT 677 ms
1×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ4_K_M
6Llama-3.2-3B-Instruct3BMeta Performance benchmarkTimebench 3 - Kombi (Prefill + Generation)
70,69 tok/s TG
Prefill 3.840 · TTFT 645 ms
1×2x NVIDIA GeForce RTX 20604x Intel(R) Xeon(R) CPU E5-4657L v2 @ 2.40GHz · 504 GB RAMllama.cppgodclawQ8_0