harbor-framework/ harbor
Framework for evaluating and improving agents
Run by mupt-ai · Released 2026-10-04 · 42 tasks · 5 settings
Which coding agent works best on harbor-framework/harbor? Most accurate: GPT-6.1 Sol, 83.3% at $0.26 per task. Best for less: GPT-6 Luna, 61.9% at $0.024. 5 model settings scored on 42 tasks from its merged pull requests.
All Settings
| Model | Harness | Reasoning | Access | Accuracy | Cost / Task | Frontier |
|---|---|---|---|---|---|---|
| GPT-6.1 Sol | Codex | Low | API Key | 83.3% | $0.26 | Yes |
| Claude Opus 5.5 | Claude Code | High | API Key | 76.2% | $1.67 | No |
| GPT-6.1 Sol | Codex | High | API Key | 71.4% | $0.45 | No |
| Claude Opus 5.5 | Claude Code | Low | API Key | 71.4% | $0.59 | No |
| GPT-6 Luna | Codex | High | API Key | 61.9% | $0.024 | Yes |
Tasks
All 42 tasks, each with its instruction, tests, and solution to browse and download.