Created bymario-alka.dePowered bygodcore.denoob2claw.detricoma.de
🤖
Harness benchmark

How well does the model solve real tasks?

Quality on real agent and chat tasks – the model must handle real tasks with tools and multiple steps. Scored by success rate and task points achieved.

🎯 Success rate🏆 Task points🔧 Tool use🧠 real tasks
Reset
#Model / MakerMetricsParallelGPU / CPU / RAMRuntime
1Qwen2.5-32B-Instruct-AWQQwen (Alibaba) HarnessbenchmarkTool Usage V 1.0
13,4 %
235/1.760 Pkt · 15/15 Aufg.
1×AMD Radeon 8060S Graphics32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAMopenclaw_cliAWQ