mupt-ai/ self-bench
Benchmark coding agents and models on your own repository with private Harbor-formatted evals built from its merged pull requests.
Run by mupt-ai · Released 2026-10-04 · 36 tasks · 6 settings
Which coding agent works best on mupt-ai/self-bench? Most accurate: GPT-6.1 Sol (High), 77.8% at $0.42 per task. Best for less: GPT-6.1 Sol (Low), 69.4% at $0.20. 6 model settings scored on 36 tasks from its merged pull requests.
All Settings
| Model | Harness | Reasoning | Access | Accuracy | Cost / Task | Frontier |
|---|---|---|---|---|---|---|
| GPT-6.1 Sol | Codex | High | API Key | 77.8% | $0.42 | Yes |
| Claude Opus 5.5 | Claude Code | High | API Key | 72.2% | $1.47 | No |
| GPT-6.1 Sol | Codex | Low | API Key | 69.4% | $0.20 | Yes |
| Kimi K3 | Codex | High | Vercel AI Gateway | 63.9% | $1.44 | No |
| GLM 5.3 | Codex | High | Vercel AI Gateway | 55.6% | $1.11 | No |
| GPT-6 Luna | Codex | High | API Key | 47.2% | $0.024 | Yes |
Tasks
All 36 tasks, each with its instruction, tests, and solution to browse and download.