When AI Writes the Code, the Engineer's Work Moves
The API contract allowed INACTIVE as a customer status. I asked an AI assistant to propose how to add an inactive customer to a small Mule API. It suggested returning 200 with the new customer in the shape already defined by the RAML. The change looked straightforward. What the contract did not say was whether an inactive customer should be visible through that lookup at all.
The assistant noticed the gap and called it an assumption. That is what makes this case useful: its answer was plausible, not nonsensical. It answered a question I had not settled. The schema said how to represent INACTIVE; it did not establish the behaviour the business wanted.
The decision exposed an older error
For this API, I decided that an inactive customer should produce a business error. The convention I use in these projects and examples assigns 400 to business errors, reserves 404 for an unknown endpoint, and uses 500 for system failures. This is a contract choice, not the only valid interpretation of HTTP: another API might return 404 for a missing resource. Here, I want consumers to be able to distinguish a domain outcome from a routing error.
Applying that rule exposed a problem that predated the AI proposal. The original API returned 404 for a well-formed customer ID with no matching customer. I changed that to 400 alongside the new inactive-customer case. The request was to add CUST-002; deciding its behaviour also meant revisiting CUST-999, which was already in the tests. The work moved from writing a new branch in the flow to defining the rule and finding the existing behaviour it affected.
The recorded AI proposal shows what the assistant suggested before any code change. The experiment record documents the decision, the implementation and the results. In that recorded round, the assistant worked in read-only mode and identified the ambiguity. I made the contract decision; the code and tests were changed afterwards.
Six parameterised MUnit cases passed. They include inactive CUST-002 and unknown CUST-999, both returning 400, and a simulated system failure returning 500. This verifies the handler for those cases. The tests call the handler directly, so they do not verify the HTTP listener or APIKit routing; the end-to-end response for an unknown route remains untested. The step-by-step tutorial shows how to inspect and reproduce the evidence.
What the assistant does, and what I do
I can use AI to inspect a contract, surface unanswered questions, suggest alternatives, write code and propose tests. In this exercise, its recorded contribution was narrower: it proposed 200 and flagged the question of whether inactive customers should be visible.
My role starts before accepting a proposal. I need to understand the problem with the people responsible for the product, choose the rule for this API and make its consequences explicit. It continues afterwards: review the diff, choose cases that might disprove the solution, and decide whether the evidence is enough to deliver. AI can help with each step, but the people responsible for the system still decide which behaviour and risks to accept. GitHub’s guide to reviewing AI-generated code likewise calls for checking tests, context and intent, not just whether the code looks convincing.
Speed needs the right rule
I use AI as an assistant to turn an idea into an early version I can test and improve. For me, the useful part of that speed is being able to try an approach while there is still time to change direction.
That speed creates value only when the solution follows the right rule. In this exercise, AI made a missing requirement visible; accepting 200 simply because it fitted the RAML would have carried the gap into the code. The 2025 DORA report describes AI as an amplifier of an organisation’s existing practices. This episode illustrates the point: the tool accelerated one reasonable interpretation, while the quality of the decision depended on the process around it.
What I observed was a plausible proposal that required judgement. The rule was missing from the requirement, and settling it uncovered an older error. As the first version of code becomes easier to produce, the engineer has to be better at saying what it must do, how to check it, and which outcomes they are prepared to stand behind.