Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
Hersteller

NVIDIA

Nemotron-Modellfamilie von NVIDIA.

11Modelle
8Benchmarks
8Performance
0Harness
Beste Token-Generierung
107,9 tok/s
Nemotron-Cascade-2-30B-A3B
Bester Prefill
3.141,4 tok/s
Nemotron-Cascade-2-30B-A3B
∅ Token-Generierung
53,5 tok/s
Schnitt aus 8 Laeufen
∅ Prefill
1.746,0 tok/s
Schnitt aus 8 Laeufen
Bester TTFT
690 ms
Nemotron-Cascade-2-30B-A3B
Beste Harness-Quote
0%
noch kein Harness-Lauf
🏆
Nemotron-Cascade-2-30B-A3B · schnellstes Modell mit 107,9 tok/s Token-Generierung
Bester Prefill des Herstellers: 3.141,4 tok/s (Nemotron-Cascade-2-30B-A3B)

Top-Modelle nach Token-Generierung

Bester gemessener Generierungs-Durchsatz je Modell (tok/s), veroeffentlichte Performancebenchmarks.

Nemotron-Cascade-2-30B-A3B
107,9 tok/s
Nemotron-3-Nano-Omni-30B-A3B-Reasoning
102,9 tok/s
NVIDIA-Nemotron-3-Nano-4B
93,4 tok/s
NVIDIA-Nemotron-3-Nano-30B-A3B
78,3 tok/s
OpenReasoning-Nemotron-32B
23,3 tok/s
Llama-3_3-Nemotron-Super-49B-v1_5
12,3 tok/s
Nemotron-H-47B-Reasoning-128K
8,0 tok/s
NVIDIA-Nemotron-3-Super-120B-A12B
2,0 tok/s