🎮
Grafikkarte
Keine GPU (CPU-only)
Other
Beste Generation
70,9 tok/s
gpt-oss-20b
∅ Generation
8,3 tok/s
Schnitt aus 43 Läufen
Bester Prefill
218 tok/s
Prompt-Verarbeitung
Bester TTFT
569 ms
Time to First Token
Beste Harness-Quote
–%
noch kein Harness
🏆
gpt-oss-20b läuft am schnellsten auf dieser GPU · 70,9 tok/s
Bester Prefill: 218 tok/s · 18 Modelle getestet
Bester Prefill: 218 tok/s · 18 Modelle getestet
Top-Modelle
Bester gemessener Durchsatz je Modell auf Keine GPU (CPU-only) – umschaltbar nach Generation, Prefill oder Kombiwert.
Performanceprofil
Durchsatz auf Keine GPU (CPU-only)
Jede Blase steht für ein Modell, eine Engine bzw. eine CPU – Position: Prompt-Verarbeitung (X) × Ausgabe (Y), Blasengröße: Anzahl Messläufe.
MODnach Modell
gpt-oss-20b 70,9 tok/sQwen3-30B-A3B-Instruct-2507 26,0 tok/sQwen3-30B-A3B-Thinking-2507 8,5 tok/sNemotron-3-Nano-4B 8,0 tok/sQwen3-Coder-30B-A3B-Instruct 6,7 tok/sNemotron-3-Nano-30B-A3B 6,6 tok/sNemotron-Cascade-2-30B-A3B 5,9 tok/sgpt-oss-120b 5,3 tok/sNorth-Mini-Code-1.0 5,1 tok/sQwen3.6-35B-A3B 4,3 tok/sMeta-Llama-3.1-8B-Instruct 3,6 tok/sMiniMax-M2.7 2,6 tok/sQwen2.5-72B-Instruct 2,3 tok/sGLM-4.5-Air 2,1 tok/sDeepSeek-V4-Flash-284B-A13B 1,2 tok/sOpenReasoning-Nemotron-32B 1,1 tok/sSeed-OSS-36B-Instruct 1,0 tok/sMiniMax-M2.5 0,8 tok/s
Durchsatz & Latenz
Im Leaderboard →Performancebenchmark
43 veroeffentlichte Performance-Läufe auf Keine GPU (CPU-only).
| # | Modell / Hersteller | Messwerte | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 70,85 tok/s TG Prefill 154 · TTFT 129.121 ms | 10× | Keine GPU (CPU-only)AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclawMXFP4 | Details → | |
| 2 | gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 53,48 tok/s TG Prefill 153 · TTFT 64.583 ms | 5× | Keine GPU (CPU-only)AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclawMXFP4 | Details → | |
| 3 | gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 28,43 tok/s TG Prefill 143 · TTFT 13.839 ms | 1× | Keine GPU (CPU-only)AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclawMXFP4 | Details → | |
| 4 | Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 25,98 tok/s TG Prefill 218 · TTFT 59.755 ms | 5× | Keine GPU (CPU-only)AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 5 | Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 24,50 tok/s TG Prefill 152 · TTFT 16.071 ms | 1× | Keine GPU (CPU-only)AMD Ryzen Threadripper PRO 7955WX 16-Cores · 184 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 6 | gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 14,47 tok/s TG Prefill 97 · TTFT 222.547 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 7 | gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 9,77 tok/s TG Prefill 89 · TTFT 119.492 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 8 | Qwen3-30B-A3B-Thinking-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 8,52 tok/s TG Prefill 71 · TTFT 152.830 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 9 | Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 7,97 tok/s TG Prefill 72 · TTFT 153.653 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 10 | Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 7,14 tok/s TG Prefill 69 · TTFT 35.318 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 11 | Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 6,74 tok/s TG Prefill 74 · TTFT 32.942 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 12 | Qwen3-30B-A3B-Thinking-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 6,64 tok/s TG Prefill 73 · TTFT 33.548 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 13 | Nemotron-3-Nano-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 6,60 tok/s TG Prefill 73 · TTFT 150.849 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 14 | gpt-oss-20b20BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 6,46 tok/s TG Prefill 79 · TTFT 25.000 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 15 | Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 6,09 tok/s TG Prefill 100 · TTFT 121.223 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 16 | Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 6,05 tok/s TG Prefill 104 · TTFT 116.729 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 17 | Nemotron-3-Nano-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 5,88 tok/s TG Prefill 61 · TTFT 35.422 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 18 | Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 5,86 tok/s TG Prefill 61 · TTFT 35.481 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 19 | Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 5,33 tok/s TG Prefill 74 · TTFT 150.609 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 20 | gpt-oss-120b120BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 5,29 tok/s TG Prefill 56 · TTFT 189.966 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 21 | North-Mini-Code-1.0Cohere PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 5,05 tok/s TG Prefill 71 · TTFT 28.002 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 22 | Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 4,94 tok/s TG Prefill 41 · TTFT 62.079 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 23 | gpt-oss-120b120BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 4,46 tok/s TG Prefill 50 · TTFT 39.304 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 24 | Qwen3.6-35B-A3B35BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 4,29 tok/s TG Prefill 62 · TTFT 31.435 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawUD-Q4_K_M | Details → | |
| 25 | North-Mini-Code-1.0Cohere PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 4,20 tok/s TG Prefill 86 · TTFT 127.307 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 26 | Meta-Llama-3.1-8B-Instruct8BMeta PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 3,55 tok/s TG Prefill 36 · TTFT 59.044 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 27 | North-Mini-Code-1.0Cohere PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 3,16 tok/s TG Prefill 109 · TTFT 186.084 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 28 | MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 2,62 tok/s TG Prefill 15 · TTFT 156.050 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawUD-Q4_K_M | Details → | |
| 29 | Qwen2.5-72B-Instruct72BQwen (Alibaba) Performancebenchmark | 2,28 tok/s TG Prefill 81 · TTFT 569 ms | 1× | Keine GPU (CPU-only)51x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | godclaw | Details → | |
| 30 | GLM-4.5-Air106BZ.ai (Zhipu) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 2,13 tok/s TG Prefill 18 · TTFT 117.963 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 31 | DeepSeek-V4-Flash-284B-A13B284BDeepSeek PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 1,20 tok/s TG Prefill 14 · TTFT 148.315 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawUD-Q4_K_XL | Details → | |
| 32 | OpenReasoning-Nemotron-32B32BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 1,10 tok/s TG Prefill 8 · TTFT 266.668 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 33 | Qwen3-30B-A3B-Instruct-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 1,04 tok/s TG Prefill 137 · TTFT 171.473 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 34 | Qwen3-Coder-30B-A3B-Instruct30BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,99 tok/s TG Prefill 130 · TTFT 172.354 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 35 | Seed-OSS-36B-Instruct36BByteDance Seed PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,97 tok/s TG Prefill 8 · TTFT 240.790 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 36 | MiniMax-M2.5230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,84 tok/s TG Prefill 22 · TTFT 98.716 ms | 1× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 37 | Qwen3-30B-A3B-Thinking-250730BQwen (Alibaba) PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,58 tok/s TG Prefill 77 · TTFT 216.585 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 38 | Nemotron-3-Nano-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,27 tok/s TG Prefill 81 · TTFT 206.882 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 39 | Nemotron-Cascade-2-30B-A3B30BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,27 tok/s TG Prefill 82 · TTFT 204.776 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 40 | Nemotron-3-Nano-4B4BNVIDIA PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,25 tok/s TG Prefill 75 · TTFT 202.135 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 41 | gpt-oss-120b120BOpenAI PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,05 tok/s TG Prefill 20 · TTFT 267.436 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawQ4_K_M | Details → | |
| 42 | MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,04 tok/s TG Prefill 29 · TTFT 211.132 ms | 5× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawUD-Q4_K_M | Details → | |
| 43 | MiniMax-M2.7230BMiniMax PerformancebenchmarkTimebench 3 - Kombi (Prefill + Generation) | 0,03 tok/s TG Prefill 28 · TTFT 212.753 ms | 10× | Keine GPU (CPU-only)4x Intel(R) Xeon(R) CPU E5-4620 v2 @ 2.60GHz · 504 GB RAM | llama.cppgodclawUD-Q4_K_M | Details → |
Agent- & Chat-Bewertung
Im Leaderboard →Harnessbenchmark
0 veroeffentlichte Harness-Läufe auf Keine GPU (CPU-only).
Noch keine Harnessbenchmarks auf dieser Hardware.
Technische Daten
Technische Daten – Keine GPU (CPU-only)
| Hersteller | Other |
|---|---|
| Beste Generation | 70,9 tok/s |
| Bester Prefill | 218 tok/s |
| Modelle | 18 |
| Messläufe (Performance) | 43 |

