🤖
Harness benchmark
How well does the model solve real tasks?
Quality on real agent and chat tasks – the model must handle real tasks with tools and multiple steps. Scored by success rate and task points achieved.
🎯 Success rate🏆 Task points🔧 Tool use🧠 real tasks
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
| 1 | Qwen2.5-32B-Instruct-AWQQwen (Alibaba) HarnessbenchmarkTool Usage V 1.0 | 13,4 % 235/1.760 Pkt · 15/15 Aufg. | 1× | AMD Radeon 8060S Graphics32x AMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 31 GB RAM | openclaw_cliAWQ | Details → |
