Back

Before AI-generated code reaches customers: a release review checklist

October 7, 2026

AI coding assistants have made it quick to produce a working feature. A developer describes what is needed, the tool writes the code, and within an hour there is something on screen that behaves as the ticket described. The speed is real, and so is the question it raises: what has to be true before this reaches a customer?

Most teams already have acceptance criteria written in business language. "Customers can download their invoices." "A failed payment shows a clear message." These describe what the feature should do when everything goes right. A release review adds a second layer, which asks whether the feature still behaves sensibly when something goes wrong, when someone else uses it, or when someone has to change it in six months. This article translates business acceptance criteria into concrete checks across four areas: access, error handling, maintainability and deployment.

Why the code needs a different kind of review

Reviewing code written by a colleague gives you a lot of context for free. You know their habits, you know what they were asked to build, and they can explain their reasoning. Code produced by an AI tool comes with none of that. It tends to look tidy and often passes the obvious tests, which makes it easy to approve on a quick read.

The practical difference is that the person submitting it may not have written, or fully read, every line. That is not a criticism of the developer, since reading every generated line closely takes nearly as long as writing it. It means the review has to ask questions that the author would normally answer without being asked, such as which cases were considered and what happens at the edges. A checklist helps because it makes those questions routine instead of dependent on whoever happens to be reviewing that day.

Start with the acceptance criteria you already have

The easiest way to build a release review is to work from the criteria the business already agreed. Each one hides several technical questions. Take an illustrative example, with invented details. A company of around eighty people adds a customer portal feature, drafted with an AI assistant, so that customers can download their own invoices as PDFs. The acceptance criterion reads: "A customer can log in and download any of their invoices."

That sentence covers the success path. It leaves open whether one customer could download another customer's invoice by changing a number in the address. It leaves open what happens when the PDF generator times out, whether the code can be understood by the next developer, and whether the new feature can be switched off quickly if something is wrong after launch. Those four gaps map directly onto the four areas below.

Access: who can see and do what

Access is where a feature that works for the person testing it can behave very differently for everyone else. Generated code often checks that a user is logged in without checking that the user is allowed to see the specific record requested. In the invoice example, the difference is between "any logged in customer can fetch invoice 1042" and "a customer can fetch invoice 1042 only if it belongs to them".

The checks we would run for access are these:

  • Try to reach another customer's record by changing the identifier in the request, and confirm it is refused.
  • Confirm that permissions are enforced on the server, and not only hidden in the interface.
  • Check that the feature uses the access rules the rest of the application already has, rather than a new set written for this one feature.
  • Look at what the feature logs, and confirm that personal data, tokens and passwords do not end up in log files.
  • If the code calls an external service or an AI model, confirm what data is sent and that it is limited to what the task needs.

The business criterion behind all of this is simple: customers see only what belongs to them. A tester who is asked to break that rule on purpose, using a second account, will find most problems within an afternoon.

Error handling: what happens when things go wrong

Generated code tends to be written for the path where everything works. That is natural, because the prompt usually describes the path where everything works. Real systems meet slow networks, missing files, unexpected input and services that are briefly unavailable.

The business criterion here is that the customer is never left confused, and that nobody on your side discovers a problem from a customer complaint. In the invoice example, imagine the PDF generator takes thirty seconds on a large invoice. Does the customer see a spinner forever, an unexplained error, or a clear message with a way to try again? Does anyone on the team get alerted?

The checks we would run:

  • Trigger failures deliberately: switch off a dependency, send an empty or oversized input, and interrupt a request halfway.
  • Read the message the customer would see and confirm it is clear and does not reveal internal details such as file paths or database names.
  • Check that errors are recorded with enough context to diagnose them later, such as which request and which customer account.
  • Confirm that a failure in this feature does not take down the pages around it.
  • Look for places where the code catches an error and carries on silently, since these hide problems that surface much later.

Maintainability: whether someone else can change it

A feature that works today still has to be understood by whoever touches it next. This matters more with generated code, because the original author may not be able to explain why a particular approach was chosen. Code that nobody on the team can confidently explain becomes expensive the first time it needs to change.

The business criterion is that the feature can be supported and extended without depending on one person's memory. For a reviewer, this translates into a handful of practical questions. Can you describe what each part does in a sentence? Does the code follow the structure and naming the rest of the application uses, or does it introduce a new pattern for the sake of one feature? Are there tests that describe the intended behaviour, including the failure cases from the previous section?

We would also check what the feature brought with it. Generated code sometimes adds a new library for something the application can already do, or duplicates logic that exists elsewhere. Each new dependency has to be kept up to date and checked for known vulnerabilities, so it is worth confirming that every one is needed, actively maintained and acceptable under your licensing rules. A short note in the pull request explaining what was generated, what was changed by hand and why is a small habit that saves a lot of time later.

Deployment: how it reaches customers and how it comes back

The last area covers the moment of release. A feature can pass every earlier check and still cause trouble if it goes live in a way that cannot be reversed quickly. The business criterion is that a problem found after release is a small, short incident, not a long one.

For the invoice feature, that means asking whether it can be released to a small group of customers first, or switched off with a setting instead of a new deployment. It also means confirming that the configuration, such as secrets and connection details, is handled properly and is not written into the code. If the feature changes the database, there should be a tested way back, and the team should know who is watching the system in the first hours after release.

The checks we would run:

  • Confirm the feature can be disabled quickly without redeploying the application.
  • Check that secrets and environment settings are kept outside the code and differ properly between test and production.
  • Run the same build that will go live through a test environment that resembles production.
  • Rehearse the rollback at least once, particularly for any change to stored data.
  • Agree in advance what will be monitored after release and who will look at it.

Putting it together

The four areas fit on a single page, which is the point. A checklist that takes an hour to complete tends to be skipped, so it helps to keep it short and tied to the acceptance criteria the business has already agreed.

Business acceptance criterion Area What the reviewer confirms
Customers see only their own data Access Server side permission checks, tested with a second account
Customers are never left confused when something fails Error handling Failures triggered on purpose, clear messages, alerts reaching the team
The feature can be supported by the whole team Maintainability Code is explained, follows existing structure, has tests and justified dependencies
A problem after release stays small Deployment Switch off option, secrets handled properly, rollback rehearsed

A reasonable way to start is to attach this table to the definition of done for any change where AI tools were used. The developer fills in the right hand column before asking for review, and the reviewer spot checks it. Over a few releases you will see which rows catch the most issues, and you can add detail where it is needed instead of everywhere.

Who should own the review

The review works best when the person who owns the business outcome is involved at the start, not at the end. A product owner or the person who wrote the acceptance criteria can say which failures would matter most to customers, and that helps the technical reviewer decide where to spend time. Where the feature touches personal data or sits in a regulated area, someone responsible for security or compliance should see the access and logging checks before release.

None of this needs to slow delivery noticeably. For a small change, the whole review can take well under an hour once the team has done it a few times. The time goes into the questions that matter, and the answers are written down for the next person.

The takeaway

AI tools shorten the time between an idea and working code, which makes the step between working code and a safe release more important. Business acceptance criteria describe what the feature should do. A release review translates them into checks for who can access it, how it fails, whether the team can maintain it and how it is released and reversed. Keeping that review short, written down and tied to what the business agreed gives you the speed of the tools without handing customers the first test.

If you would like to look at how AI generated code moves through your own review and release process, let's talk about how this framework applies to your product.

Author

Chairunnisa Irianto

Nisa is a Marketing Manager at Itsavirus, a strategic software development partner working with companies across Europe and Southeast Asia. She writes about AI, application modernisation, and how businesses turn technology into practical results.

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
How to build a knowledge base that gets smarter over time with Obsidian and Claude Code
Your AI keeps forgetting. Here's how to stop repeating yourself
OpenClaw is exciting. But, here's what you need to secure before you experiment

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
Shadow AI: the tools your team already uses, and who is accountable for them
Switzerland is building a sovereign cloud. Here is why every European business should be paying attention
AI can write the code, but your company still owns the consequences.

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
The cost of leaving your cloud provider ends on 12 January 2027
Choosing an AI model just stopped being simple, and 25 US tech companies are fighting to keep it that way
Claude Fable 5: Launched, praised, then pulled within 3 days

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
Workshop : From Idea to MVP
Webinar: What’s next for NFT’s?
Webinar: finding opportunities in chaos

Latest insights

A sharp lens on what we’re building and our take on what comes next.

See more
How we helped Ecologies to turn survey results into reliable, faster reports using AI
How to deal with 1,000 shiny new tools
Develop AI Integrations with Itsavirus
No items found.