08 Oct 2026 AI in software testing: testing with AI, and testing AI itself
In short: AI in software testing means two things. You can use AI to help you test, for example to draft test cases or repair broken scripts. And you can test software that has AI inside, like a chatbot. The first saves time. The second needs a different way of testing, because AI doesn’t always give the same answer twice.
Ask ten people what “AI testing” means and you get two kinds of answers. Testers talk about tools that write and fix test scripts. Product owners talk about the chatbot their company just launched, and how to know it gives the right answers. Both are right. Most teams will need both soon: they start using AI in their testing, and they ship features with AI inside.
What does AI testing mean?
AI testing has two meanings that are easy to mix up. The first is testing with AI: using AI tools to support your testing work. They draft test cases from a user story and suggest test data. They repair a script when a button gets a new name, or summarise why a run failed. The second is testing of AI: checking software that has AI inside it. Think of a customer service chatbot, a tool that summarises documents, or a model that scores loan applications. Here the question is whether the AI gives correct, safe and fair answers. The testing world now treats these as separate subjects. The ISTQB AI Testing certification (CT-AI), in its current version 2.0, covers testing AI-based systems. Using generative AI in your test work has its own certification, CT-GenAI. For a team lead, the useful question is simple: which of the two do we need now, and which next?
How can AI help your software testing today?
AI is good at the repetitive parts of testing work, the parts that take time but little judgment. At Nekst we use AI where it helps our testers, and these are the places it pays off most:
- Drafting test cases. Give it a user story and acceptance criteria, and it suggests test cases, including edge cases you might skip. A tester still checks and trims the list.
- Creating test data. Realistic names, addresses and orders in the right format, without copying production data.
- Repairing scripts. Many automation tools now suggest a fix when a screen element changes. This is often called self-healing.
- Reading failures. Summarising long logs and grouping failures with the same cause, so you know where to look first.
- Writing test code faster. Coding assistants turn a described test into a first version of the script.
The common thread: AI makes the first draft faster. A person still decides whether the draft is right.
What can’t AI testing tools do for you?
AI writes test scripts faster. It doesn’t decide what matters. That part of testing stays with people, and it becomes more important as development speeds up with AI. When developers produce code faster, testing has to keep up, and more tests are not the same as better coverage. Five things AI tools don’t solve:
- Choosing what to test. Which business process can never break? Which risk costs the most? That comes from knowing the business, not from the code.
- Owning the suite. Someone still has to run the tests every night, fix what breaks and remove what no longer matters. A suite full of generated tests that nobody maintains decays just as fast as a hand-written one.
- Managing test data. Generated data helps, but tests still fail when the data in a shared environment changes underneath them.
- Judging a result. Is this failure a real bug, a test problem or a changed requirement? That call needs context.
- Testing like a user. AI can turn a user story into test cases. But it can’t tell you whether a real person will understand the screen, find the button or trust the result. That takes people: exploratory testing by testers who think for themselves, and acceptance testing by the business users who will work with it every day.
How do you test software that has AI inside?
Software with AI inside needs a different approach, because its answers aren’t always the same. Ask a chatbot the same question twice and the wording changes. Sometimes the content changes too. A simple pass or fail on one exact answer doesn’t work. You test against criteria instead. Is the answer correct? Is it complete? Does it stay within what the company actually offers? Does it refuse what it should refuse? That is how you test AI in practice. The steps are much the same for most applications:
- Build a set of reference questions. Start with 30 to 50 real questions from your customers or users, including awkward ones. For each, write down what a good answer must contain and must never contain.
- Score the answers, don’t just compare them. Check each answer against the criteria. Some checks can be automated, like whether a price or date is right. Others need a human reviewer or a second model that grades the answer. This is often called LLM evaluation.
- Test misuse. Try to make the AI do what it shouldn’t: share data, ignore its instructions or promise things the company can’t deliver. Attempts to trick a model with clever input are called prompt injection.
- Run the set again after every change. A new model version, a changed prompt or new source documents can change the answers. This is regression testing for AI.
- Keep monitoring, and be able to show it. Sample real conversations and check them against the same criteria, also when nothing has changed. Answers can drift as data and usage change. Record what you checked and what you found. For high-risk AI systems, the EU AI Act requires exactly this: documented monitoring after release, for the whole lifetime of the system.
Frameworks help you cover the risks. The OWASP AI Testing Guide describes how to test AI systems for security and trustworthiness. Many companies are already at this point. In the 2026 survey by Applause, 54.5% of 1,636 respondents said their organisation has already released AI features.
Which AI testing tools are worth a look?
There is no single best AI testing tool, because the tools do very different things. It helps to think in three groups. First, AI features inside test automation tools you may already use, like suggested fixes for broken scripts. These are the easiest start, with the least risk. Second, AI-first test platforms that write and run tests from plain-language instructions. They can be fast to start with. Check how you keep control over what gets tested, and what happens if you ever want to leave. Third, evaluation tools for applications with AI inside, which run your reference questions and score the answers. Pick the group that matches your question first, and only then compare tools.
What this means for your team
Start with the question you have now. If you want to use AI in your testing, pick one task that costs your testers the most time, like drafting test cases. Try AI there for a month and measure what it saves. Keep the thinking with your testers: they review what AI drafts, and they keep exploring the application themselves. Building a feature with AI inside? Start with 30 to 50 real questions and clear criteria before go-live, and decide who reviews the answers.
The second path doesn’t end at go-live. An AI feature keeps changing underneath you: a new model version, new source documents, new ways people use it. So you keep testing its answers in production. You also need to show what you tested, to your management and, for high-risk systems, to regulators. With a spreadsheet of questions and answers, that falls apart within a few months.
That is why we built VeritasAI, our in-house tool for testing AI applications. We use it to test your AI application as a service: it runs the test questions continuously and reports on the results. A management dashboard shows how your AI application is doing over time. An operational dashboard shows which questions it fails and which law, regulation or framework each failure relates to. Every finding comes with the prompt, the answer and a priority score, so your team knows what to fix first. You can read how that works on our AI Compliance Testing page.
Find the wrong answers in your AI before your customers do.
Frequently asked questions
Will AI replace software testers?
Not the part of the job that matters most. AI speeds up repetitive work like drafting test cases and fixing scripts. Deciding what to test, judging results and owning the quality of a release still need people who know the business. So do exploratory testing and acceptance testing by real users. The work shifts; it doesn’t disappear.
What is the difference between AI testing and AI in testing?
People use both terms loosely. “AI in testing” usually means using AI tools to support testing. “AI testing” often means testing AI systems themselves, but it is also used for the first meaning. When it matters, ask which one someone means.
How do you test a chatbot?
Write a set of real questions, decide what a good answer must and must not contain, and score the answers against those criteria. Add questions that try to misuse the chatbot. Run the set again after every change to the model, prompt or source documents.
Is there a certification for AI testing?
Yes. ISTQB offers Certified Tester AI Testing (CT-AI) for testing AI-based systems. For using generative AI in your test work, there is Certified Tester Testing with Generative AI (CT-GenAI).