
AI Hallucinations in Test Automation: 8 Fixes
Liviu Lupei
Founder & Solutions Architect, Endtest · July 21, 2026
AI hallucinations are usually discussed as a chatbot problem.
You ask a question, the model invents a fact, and someone notices that the answer sounds confident but wrong.
In test automation, the failure mode is more dangerous.
The AI can create a test that looks reasonable, runs successfully, and still verifies the wrong thing. It can click the wrong button, extract a locator for a nearby element, accept an invalid visual state, or quietly replace a broken selector with one that points somewhere else.
The result is often not a red build but a green one that gives the team false confidence.
That is why reducing AI hallucinations in test automation is a system-design problem more than a prompt-writing exercise. The model, the context, the output format, the validation layer, and the review process all matter.
Here are the practical ways to make AI-assisted testing more reliable.
What AI hallucinations look like in test automation
A hallucination does not always mean the AI invents an entire feature that does not exist. More often, it makes one plausible but incorrect decision inside a larger workflow.
| AI task | Example hallucination | Why it is risky |
|---|---|---|
| Creating tests from specifications | The model invents a confirmation message or skips a required precondition | The generated test no longer represents the actual requirement |
| Extracting element locators | The model selects a similar button, label, or hidden element | The test may interact with the wrong part of the page |
| Performing AI assertions | The model decides that a page is correct based on incomplete visual or textual evidence | A real defect can be classified as a pass |
| Generating random values | The model creates data with the wrong format, length, locale, or business rules | The test fails for irrelevant reasons or exercises an unrealistic path |
| Finding alternative locators | A self-healing system replaces a broken locator with a nearby but incorrect element | The test continues passing while testing a different behavior |
These mistakes become more likely when the model receives vague instructions, excessive context, no permission to express uncertainty, and no deterministic checks around its output.
1. Give clear and granular instructions
Do not assume the AI will infer the test intent from a short feature description.
A human tester may understand that "test the checkout flow" means using an existing customer, selecting a specific shipping method, applying a valid coupon, confirming the tax calculation, and verifying the order in the account history.
An AI model sees many possible checkout flows. Unless you define the important details, it has to fill the gaps.
This instruction is too broad:
Create a test for checkout.A better instruction defines the preconditions, actions, boundaries, and expected outcome:
Create a test for a signed-in customer purchasing one in-stock product.
Preconditions:
- The customer has a saved delivery address.
- The product price is $49.00.
- The coupon SAVE10 is valid and provides a 10% discount.
Actions:
- Add the product to the cart.
- Apply the coupon.
- Select standard shipping.
- Complete the purchase with the test card.
Assertions:
- The subtotal is $49.00.
- The discount is $4.90.
- The order confirmation page displays an order number.
- The same order appears in the customer's order history.
Do not add steps that are not supported by these instructions.Nobody wants to write a novel for every test. The point is to remove the places where the model would otherwise have to guess.
Smaller tasks also help. Asking AI to identify one locator, generate one assertion, or convert one requirement at a time is usually more reliable than asking it to design an entire regression suite in a single response.
2. Use a model specialized for test automation
A more expensive model is not automatically the best model for every testing task.
General-purpose foundation models are built to handle a huge range of requests: writing, research, programming, summarization, brainstorming, and many other activities. That breadth is useful, but it does not make them specialists in test intent, locator quality, browser behavior, assertion design, or self-healing decisions.
Test automation has its own rules.
A useful model needs to understand that a locator should be stable and unique, that a hidden element should not be preferred over an interactable one, that an assertion must verify business intent rather than merely confirm that some text exists, and that changing a locator can silently change what a test is testing.
At Endtest, we do not simply route every AI task through a generic model and hope that a larger prompt will solve the problem. We use models specialized for test automation. In our experience, that is one of the main reasons the hallucination rate is extremely low.
Specialization matters more than price.
Specialized models are also not always available as a simple option inside general AI tools such as ChatGPT or Claude. A team building an internal framework may have access to excellent general-purpose models while still lacking a model trained or optimized for the narrow decisions required by reliable test automation.
This is also why "Which AI model is smartest?" is the wrong question. A better one is "Which model is most reliable for this exact testing operation?" We discussed that distinction in more detail in our article about the best AI model for test automation.
3. Send less data, not more
When an AI model makes a mistake, the instinct is often to give it more context.
Sometimes that helps. Often it creates more opportunities for confusion.
A full page source can contain thousands of elements, hidden templates, duplicate mobile and desktop navigation, old modal content, analytics markup, SVG paths, dynamic IDs, and components that are not relevant to the current test step.
If the task is to find the locator for the visible Submit Order button, the model probably does not need the entire HTML document.
It may only need:
- The relevant form or component.
- A small amount of parent context.
- Nearby labels and attributes.
- Information about visibility and interactability.
- A screenshot or accessibility representation when useful.
This reduces both token consumption and ambiguity.
The same principle applies to test generation. Do not send a 70-page specification when only two paragraphs describe the workflow being automated. Retrieve the relevant section, remove unrelated material, and send the smallest context that still contains the required evidence.
4. Ask the AI to reason through a defined process
"Think step by step" can help, but it is more useful when the steps are specific to the task.
For locator generation, the model can be instructed to follow a process such as:
1. Identify the intended element from the instruction.
2. List the visible candidate elements that could match.
3. Reject candidates that are hidden, disabled, duplicated, or semantically different.
4. Prefer stable attributes over generated classes or dynamic IDs.
5. Verify that the proposed locator matches exactly one interactable element.
6. Return the locator and a short explanation of the supporting evidence.For an AI assertion, the process may be:
1. Restate the expected condition.
2. Identify the evidence required to prove that condition.
3. Check whether that evidence is present in the supplied data.
4. Distinguish between a pass, a failure, and insufficient evidence.
5. Do not infer missing information.This reduces internal logic gaps and makes the output easier to inspect.
The explanation should be concise. You do not need pages of reasoning. You need enough structure to verify that the model evaluated the right evidence and did not jump from "this looks similar" to "the test passed."
5. Give the model permission to be uncertain
Many AI workflows accidentally reward guessing.
The system expects a locator, a pass/fail answer, or a generated value. The model therefore produces one, even when the evidence is weak.
Explicitly give it another valid response:
If you are not completely certain about something, say you don't know rather than guessing.For production systems, make uncertainty part of the output schema:
{
"status": "uncertain",
"confidence": 0.54,
"reason": "Two visible buttons match the supplied instruction.",
"result": null
}The surrounding automation must respect that answer.
If the model reports insufficient evidence, the system should request more context, use a deterministic fallback, or send the decision for review. It should not quietly convert "uncertain" into "good enough."
This matters especially for self-healing. A system that always returns an alternative locator is not necessarily intelligent. It may simply be incapable of admitting that it cannot identify the original element safely.
6. Separate generation from validation
Do not ask the same AI response to generate an answer and certify that the answer is correct.
Add a validation layer after generation.
Some checks can be deterministic:
- Does the locator match exactly one visible element?
- Is the element interactable?
- Does the generated step type exist in the automation platform?
- Does a generated email address match the required format?
- Is a random date inside the allowed range?
- Does the expected value appear in the specification?
- Did a self-healed locator preserve the element's role, label, and surrounding context?
Other checks can use a second model pass with a narrower instruction. The validator should receive the proposed output and ask whether it is fully supported by the available evidence.
The principle is the same one used in reliable software systems: do not trust input merely because it came from a sophisticated component.
The AI proposes, and the validation layer decides whether the proposal is safe to use.
7. Constrain the output
Free-form text gives the model many ways to be creative. Test automation usually benefits from the opposite.
Use structured outputs with:
- A fixed list of allowed step types.
- Required fields for each action.
- Explicit assertion operators.
- Defined locator strategies.
- Valid data types and ranges.
- Rejection of unknown properties.
- Confidence and uncertainty fields.
For example, a test creation model should not be able to invent a new step called VerifyEverythingLooksGood. It should select from known actions such as click, write text, send an API request, switch to an iframe, or assert an exact value.
This is one reason editable, structured test steps are safer than receiving a large block of generated code. The system can validate each step independently, and a human can see what the test is supposed to do.
Endtest's AI Test Creation Agent creates tests as structured, editable actions rather than leaving the team with an opaque conversation and a pile of generated automation code.
8. Never let AI change test intent silently
AI can help repair a locator. It should not silently redefine the test.
A self-healing system should preserve evidence about:
- The original locator.
- The replacement locator.
- The old and new elements.
- The attributes and context used for matching.
- The confidence score.
- The screenshot or page state at the time of healing.
- Whether the change was applied automatically or requires approval.
High-confidence repairs can be useful for predictable changes, such as a generated ID being replaced while the element's label, role, and position remain consistent.
Low-confidence repairs should stop.
Consider a checkout page with two buttons: Continue Shopping and Continue to Payment. A self-healing system that selects the first visible element containing "Continue" has technically fixed the broken locator. It has also changed the test.
The requirement is to preserve the original test intent, and "something similar" does not meet it.
For more detail, see our guide to self-healing test automation.
The Playwright + Claude maintenance trap
There is nothing wrong with using Playwright or Selenium. Both can be useful foundations for browser automation.
The problem starts when a team assumes that adding Claude or another general-purpose AI model will turn an internal framework into an autonomous testing platform.
The first demo can look impressive. The AI reads a requirement, writes a Playwright test, fixes a selector, and explains a failure. It feels like the maintenance problem has disappeared.
Then the suite grows.
The model needs more repository context. It needs page objects, fixtures, helper methods, test data, environment rules, previous failures, application HTML, screenshots, and parts of the product specification. Token consumption increases, prompts become infrastructure, and every generated change still requires review.
Three problems appear quickly.
Token consumption becomes an operating cost
Large codebases and page sources must be repeatedly selected, sent, and interpreted. Even when model prices decrease, someone still has to build and maintain the context pipeline, control retries, store traces, and investigate inconsistent output.
Nobody is completely sure what the tests are testing
Generated code can be syntactically valid while representing the wrong business intent. When the AI changes selectors, assertions, waits, or helper methods, reviewers must determine whether it repaired the implementation or changed the meaning of the test.
That uncertainty compounds over time.
A general model is being asked to make specialist decisions
Claude and other foundation models are capable programming assistants. But test maintenance is more than code generation. It requires narrow judgment about intent, evidence, element identity, browser behavior, and the acceptable boundary of a repair.
A Playwright + Claude or Selenium + Claude setup can therefore become expensive and unreliable: high token usage, heavy review requirements, uncertain coverage, and more hallucinations than teams expect from the initial prototype.
These models are not bad. The problem is using a broad model as though it were a specialized test automation system.
We covered the broader maintenance tradeoff in AI Playwright Testing: Useful Shortcut or Maintenance Trap?.
How we reduce hallucinations at Endtest
Our approach is to keep AI powerful but bounded.
That means using specialized models for test automation, giving each model a narrow task, trimming the supplied context, validating structured outputs, allowing uncertainty, and keeping the resulting tests readable and editable.
The AI should accelerate test creation and maintenance without becoming an invisible layer that continuously rewrites the suite.
A useful rule is this:
The more important the decision, the less freedom the model should have to improvise.
Generating a realistic first name can tolerate more variation than changing the locator for a Delete Account button. A visual suggestion can tolerate more uncertainty than declaring a payment flow correct. The system should treat those operations differently.
A practical hallucination-reduction checklist
Before letting an AI-generated decision affect a test, ask:
- Is the instruction specific enough to define the expected behavior?
- Is the model specialized for this testing task?
- Did we send only the relevant context?
- Is the model following a defined evaluation process?
- Can it return "uncertain" or "insufficient evidence"?
- Is the output constrained by a schema?
- Can deterministic checks validate the result?
- Will a human be able to understand what changed?
- Can a self-healing action be reversed and audited?
- Did the AI preserve the original test intent?
If several answers are no, the system is not reducing maintenance so much as moving it somewhere less visible.
Frequently asked questions
What causes AI hallucinations in test automation?
The most common causes are vague instructions, missing requirements, excessive or irrelevant context, general-purpose models being used for specialized tasks, unconstrained output, and systems that force the model to return an answer even when evidence is insufficient.
Can AI-generated tests be trusted?
They can be useful and reliable when the output is structured, reviewable, validated, and tied to clear requirements. AI-generated tests should not be trusted merely because they compile or pass once.
Do larger and more expensive models hallucinate less?
Not necessarily for every task. Model quality matters, but specialization, context selection, prompting, output constraints, and validation often matter more than simply selecting the most expensive general-purpose model.
Is Playwright with Claude enough for AI test automation?
It can be useful for experiments and developer assistance. It is not automatically a complete maintenance system. Teams still need context management, validation, browser infrastructure, reporting, review workflows, and safeguards that prevent generated changes from altering test intent.
The goal is not blind trust
AI can make test automation dramatically faster. It can turn instructions into tests, identify elements, create realistic data, evaluate complex conditions, and reduce repetitive maintenance.
But reliability comes from boundaries: clear instructions, a specialized model, a small context, a defined reasoning process, permission to be uncertain, a schema on the output, deterministic validation, and an audit trail for every self-healing change. If a decision matters, the model should have less room to improvise, and someone on the team should be able to see exactly what it changed.