🤖
Harness benchmark
How well does the model solve real tasks?
Quality on real agent and chat tasks – the model must handle real tasks with tools and multiple steps. Scored by success rate and task points achieved.
🎯 Success rate🏆 Task points🔧 Tool use🧠 real tasks
| # | Model / Maker | Metrics | Parallel | GPU / CPU / RAM | Runtime | ||
|---|---|---|---|---|---|---|---|
No matching results yet. | |||||||
