Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
🤖
Harness-Benchmark

Wie gut löst das Modell echte Aufgaben?

Qualität in echten Agenten- und Chat-Aufgaben – das Modell muss reale Aufgaben mit Tools und mehreren Schritten bearbeiten. Bewertet nach Erfolgsquote und erreichten Aufgabenpunkten.

🎯 Erfolgsquote🏆 Aufgabenpunkte🔧 Tool-Nutzung🧠 echte Aufgaben
Reset
Messwert:
Größe
#Modell / HerstellerMesswerteGPU / CPU / RAMRuntime
1Gemma-4-26B-A4B-it26BGoogle HarnessbenchmarkTool Usage Standard 1.0
2.701 Pkt
90,6% · 16/16 Aufg. · 118,0 s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMopenclaw_cli
2Qwen3.6-27B27BQwen (Alibaba) HarnessbenchmarkTool Usage Standard 1.0
1.220 Pkt
40,9% · 16/16 Aufg. · 1.036,0 s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMvLLMgodclaw
3Qwen3.6-27B27BQwen (Alibaba) HarnessbenchmarkTool-Parcours Dossier 40 V1.0
2.267 Pkt
38,0% · 38/40 Aufg. · 10.331,0 s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation EditionAMD Ryzen 9 9950X 16-Core Processor · 92 GB RAMvLLMgodclaw