Project

Coding-Agent Evaluation Harness

A personal experiment comparing coding-agent results against frozen acceptance checks, with explicit limits on what a small provider comparison can establish. Built AI-assisted.

  • Python
  • Coding agents
  • Evaluation
  • Acceptance testing

The project, internally named deadbolt, evaluates coding-agent results against acceptance checks fixed before the run. I developed it with AI-assisted implementation, concentrating on whether the evidence supports the claimed result.

Fixed criteria

The harness treats the acceptance checks as a frozen reference. Ambiguous criteria or altered checks invalidate a comparison rather than becoming a convenient pass. This makes it possible to distinguish completing the task from changing the test.

Results and limits

A small comparison used real providers to demonstrate the evaluation workflow. It is a bounded experiment, not a general ranking of models or proof that one agent is superior across software-engineering tasks. The useful artifact is a repeatable comparison with explicit criteria and inspectable outcomes.