How We Use Claude AI in QA Test Execution

Ewertton Souza

Ewertton Souza

Posted on Oct 01, 2026
SHARE

When developers build tickets with Claude Code, QA has to use AI too — from a testing angle, with a human making the final call. Between 17 August and 30 September 2026, that workflow produced 109 QA comments citing a Playwright run, across 89 tickets on a legal-payments fintech platform SWARECO builds and supports. Manual testing did not go away. It moved to the part of the job only a person can do: judgment.

Why AI in QA testing matters once developers use AI

You can't check AI-written code well by reading it from the same point of view that wrote it.

More of our tickets are now resolved or built with help from Claude Code. If QA only re-reads that code the way the model wrote it, QA inherits the model's assumptions. So at SWARECO we use Claude for QA testing a second time, from a testing angle, and always pair it with a human view — because the human view is what reflects the real user.

Where Claude fits in our QA testing workflow

We use Claude for four specific jobs in our QA pipeline.

The product we test integrates closely with a third-party platform. Most of the work is validating UI flows, finding edge cases, and writing bug reports clear enough for engineering to act on fast.

1. Cross-checking AI-generated code before QA sign-off

A ticket built partly or fully with Claude Code is not "done" when the pull request opens.

We run Claude again to review the logic, flag assumptions the original generation made, and compare the implementation against the ticket's actual acceptance criteria. For specific tickets, we also use Claude with Playwright to write new tests and run them against our staging environment.

The rule that matters most: QA generates its own numbers instead of re-running the developer's.

On one performance ticket, the developer's change took the login page from a Lighthouse performance score of 37 to 90, and the CI size gate passed. Our QA pass measured independently — a Playwright network trace, a fresh production build, and two Lighthouse runs. It found an eager transfer of 552.7 KB against a ~400 kB acceptance criterion, and an index.html that was never compressed.

Both findings were real. Re-running the developer's own measurement would have missed them.

2. Testing the seams unit tests cannot reach

Some bugs live between systems, where a unit test passes while the app is broken.

One login bug hid every feature-flagged screen until the user refreshed the page. The relevant unit test passed the whole time the app was broken. We verified the fix on staging with 4 Playwright scenarios, each one a real login round-trip that checks a gated settings tab appears immediately, with no refresh.

3. Turning bug recordings into structured Jira tickets

We built a small internal tool that connects Jam screen recordings to Claude:

  • A Jam screen recording goes into the tool.
  • Claude extracts the bug details: steps to reproduce, console errors, and network failures.
  • It formats them as a structured Jira ticket, ready to file.

It shortens the time between "I found a bug" and "there is a clear, actionable ticket for it."

4. Reviewing UI and flow documentation

For larger initiatives, such as a new integration portal, we give Claude the wireframes and flow specs. It reads both, surfaces inconsistencies between them, and drafts test cases that map back to real user journeys rather than isolated screens.

What the numbers look like

Every QA pass ends as a comment on the Jira ticket, so the volume is countable. From the first Playwright-backed report on 17 August to 30 September 2026:

Measure Count
QA comments citing a Playwright run 109
Tickets those comments covered 89
Tickets per week 9 to 21
Busiest week (from 21 September) 21 tickets, 22 comments

Comments outnumber tickets, and that is expected. It happens when a ticket fails QA, gets a fix, and is tested again. A second report on the same ticket means the process is working.

The human layer still matters

None of this replaces manual testing. Our automated regression testing runs on Playwright with Claude, and it still sits underneath human checks, not in place of them.

AI test automation with Claude speeds up the repetitive parts of the job: writing tickets, summarizing screen recordings, writing and running Playwright specs, and spotting obvious gaps. The final judgment on whether something works the way a user expects stays human. We still validate business workflows, functional behavior and UI by hand, and those checks come first.

That is the balance: AI for speed and consistency, people for context and judgment. Claude can tell you a test passed. It cannot tell you whether the flow makes sense to the person who has to use it.

What's next

We are tightening the Jam-to-Jira pipeline and expanding the automated checks around the integration. The goal of AI in QA testing is not to remove QA from the loop. It is to give QA more time for the things that need a human eye.

If you are building your own QA checks, start small. Our guide to smoke testing in software development covers the first layer, and our post on Claude Code headless mode covers running Claude non-interactively inside a pipeline.

Other Articles

We build the engineering. You build the business.

If you are trying to figure out whether SWARECO is the right fit for what you are building, the best way to find out is to talk. Tell us what you have. We will be direct about what we can do and how we would approach it.