Project
Coding-Agent Evaluation Harness
A personal experiment comparing coding-agent results against frozen acceptance checks, with explicit limits on what a small provider comparison can establish. Built AI-assisted.
The project, internally named deadbolt, evaluates coding-agent results against
acceptance checks fixed before the run. I developed it with AI-assisted
implementation, concentrating on whether the evidence supports the claimed result.
Fixed criteria
The harness treats the acceptance checks as a frozen reference. Ambiguous criteria or altered checks invalidate a comparison rather than becoming a convenient pass. This makes it possible to distinguish completing the task from changing the test.
Results and limits
A small comparison used real providers to demonstrate the evaluation workflow. It is a bounded experiment, not a general ranking of models or proof that one agent is superior across software-engineering tasks. The useful artifact is a repeatable comparison with explicit criteria and inspectable outcomes.