mupt-ai/self-bench

Benchmark coding agents and models on your own repository with private Harbor-formatted evals built from its merged pull requests.

Run by mupt-ai · Released 2026-10-04 · 36 tasks · 6 settings

Which coding agent works best on mupt-ai/self-bench? Most accurate: GPT-6.1 Sol (High), 77.8% at $0.42 per task. Best for less: GPT-6.1 Sol (Low), 69.4% at $0.20. 6 model settings scored on 36 tasks from its merged pull requests.

All Settings

ModelHarnessReasoningAccessAccuracyCost / TaskFrontier
GPT-6.1 SolCodexHighAPI Key77.8%$0.42Yes
Claude Opus 5.5Claude CodeHighAPI Key72.2%$1.47No
GPT-6.1 SolCodexLowAPI Key69.4%$0.20Yes
Kimi K3CodexHighVercel AI Gateway63.9%$1.44No
GLM 5.3CodexHighVercel AI Gateway55.6%$1.11No
GPT-6 LunaCodexHighAPI Key47.2%$0.024Yes

Tasks

All 36 tasks, each with its instruction, tests, and solution to browse and download.

All Repositories on SelfBench