I run a couple of agentic CLI tools daily, and the failure modes are weirdly consistent: unasked-for refactors, invented APIs, passing tests by weakening them. Prompt tweaks help for a day, then regress. So I built a small harness with 30 real tasks from our repo plus a checker that runs the
Do you write evals for your coding agent?
Replies
No comments yet — be the first to share your thoughts.