Evals for Coding Agents with Harbor
For my October 5, 2026 Austin AI MUG presentation, I explored a question that came out of building RobotBridge.
How do I measure the effectiveness of coding agent (Claude Code/Codex) instructions on MCP tool use?
You can download the presentation slides as a PDF.
RobotBridge is my open source macOS agent workspace, built on a Zebric blueprint. It brings a task board together with Claude Code and Codex in terminal tabs, with MCP tools connecting the agents to the board.
Coding Agents don’t use the tools without instructions
When I used RobotBridge initially, there were no AGENTS.md or CLAUDE.md instructions - so the coding agents didn’t pull tasks off the task board. I added a feature to RobotBridge to add task tracker tool descriptions into the agent instructions in the repo, and saw some success, but not as much as I wanted.
Now I know where to put the improvements, but I didn’t know what instructions would be best. The solution to that would be setting up an experiment.
Testing instructions with Harbor
I used the Harbor framework to run repeatable tasks in sandboxed environments. Each task included a prompt, a starting repository and tracker state, and a verifier that checked the results.
My initial experiment covered five fixtures, two agents, two model tiers, three instruction conditions, and two repetitions, for 120 runs. The conditions were:
- A control with no tracker instructions.
- A short instruction telling the agent to list tasks first, mark the task
in_progress, and mark itdonewhen finished. - A longer instruction from RobotBridge 0.2.0 describing the tools and task board.
Harbor measured coding success with hidden pytest tests. On the Zebric side (the MCP server), Zebric maintains an audit log, along with a final database state. The primary RobotBridge workflow outcome was whether the seeded task ended in done.
Results
All 120 runs passed the coding tests.
Task completion in the tracker was the interesting part of the experiment:
| Agent | Model tier | Control | Short instruction | Product instruction |
|---|---|---|---|---|
| Claude Code | Small | 0/10 | 0/10 | 2/10 |
| Claude Code | Medium | 0/10 | 10/10 | 3/10 |
| Codex | Small | 0/10 | 2/10 | 2/10 |
| Codex | Medium | 0/10 | 8/10 | 0/10 |
| Total | 0/40 | 20/40 | 7/40 |
The short instruction worked best in this experiment, especially with the medium-tier models. Making the tools available on its own wasn’t enough to get the agents to complete the tracker workflow (the control).
This was a small experiment costing about $20, with two repetitions per combination.
The results gave me a starting point for improving the instructions. it’s not really enough runs to draw any agent or model-specific conclusions.
Conclusion
Out of this initial experiment run, I’ve got enough results to update the agent instructions for RobotBridge 0.2.0. I also had enough to explore a few more improvements to the prompt.
For the experiment itself - I’d like to simplify it a little bit, specifically for RobotBridge, where I’m interested in the MCP tools, because I have control over those. I may not need to spend API tokens on coding challenges, because they all passed.