Meta
Meta-Llama-3.1-8B-Instruct
Benchmark profile and published results.
Dense8B
Position in the field
Bestwerte im Vergleich
Best metrics of this model against the minimum, average and maximum of all published systems.
Generation83,4 tok/s
Min 0,0Ø 244,1Max 1.443,9
Unter dem Durchschnitt · 596 Systeme im Feld
Prefill2.658 tok/s
Min 6Ø 5.354Max 39.029
Unter dem Durchschnitt · 595 Systeme im Feld
Time to First Token1.773 ms
Min 91Ø 13.148Max 299.591
Ueber dem Durchschnitt · 597 Systeme im Feld
Performance profile
Throughput by hardware & engine
Each bubble is a GPU, CPU or engine – position shows prompt processing (X) and output speed (Y), bubble size the number of runs.
GPUnach Grafikkarte
AMD Radeon 8060S Graphics 83,4 tok/sCPU-only 3,6 tok/s
CPUnach Prozessor
AMD RYZEN AI MAX+ 395 w/ Radeon 8060S 83,4 tok/sIntel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz 3,6 tok/s
ENGnach Engine
llama.cpp 83,4 tok/s
Agent & chat rating
Harness benchmark
Noch keine Harnessbenchmarks fuer dieses Modell.
Throughput & latency
Performance benchmark
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Meta-Llama-3.1-8B-InstructMeta PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 83,37 tok/s TG Prefill 2.658 · TTFT 10.905 ms | 10× | AMD Radeon 8060S Graphics32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 2 | Meta-Llama-3.1-8B-InstructMeta PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 76,24 tok/s TG Prefill 1.951 · TTFT 6.429 ms | 5× | AMD Radeon 8060S Graphics32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 3 | Meta-Llama-3.1-8B-InstructMeta PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 40,39 tok/s TG Prefill 1.237 · TTFT 1.773 ms | 1× | AMD Radeon 8060S Graphics32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 4 | Meta-Llama-3.1-8B-InstructMeta PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 3,55 tok/s TG Prefill 36 · TTFT 59.044 ms | 1× | Keine GPU (CPU-only)51x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → |
