harbor-framework/harbor

Framework for evaluating and improving agents

Run by mupt-ai · Released 2026-10-04 · 42 tasks · 5 settings

Which coding agent works best on harbor-framework/harbor? Most accurate: GPT-6.1 Sol, 83.3% at $0.26 per task. Best for less: GPT-6 Luna, 61.9% at $0.024. 5 model settings scored on 42 tasks from its merged pull requests.

All Settings

ModelHarnessReasoningAccessAccuracyCost / TaskFrontier
GPT-6.1 SolCodexLowAPI Key83.3%$0.26Yes
Claude Opus 5.5Claude CodeHighAPI Key76.2%$1.67No
GPT-6.1 SolCodexHighAPI Key71.4%$0.45No
Claude Opus 5.5Claude CodeLowAPI Key71.4%$0.59No
GPT-6 LunaCodexHighAPI Key61.9%$0.024Yes

Tasks

All 42 tasks, each with its instruction, tests, and solution to browse and download.

All Repositories on SelfBench