Functional tests check that a page works; they don't notice a broken layout or a button screen-reader users can't reach. This guide adds two kinds of checks to Playwright suites: visual regression (comparing screenshots against approved baselines) and accessibility (automated axe-core scans, ARIA snapshots and keyboard checks). Examples use Java, with notes for TypeScript.
PLAYWRIGHT (JAVA)
Visual Regression & Accessibility Testing
Screenshot comparison, masking, Applitools, axe-core & ARIA snapshots
The two specialized testing types most Java notes skip
The TypeScript runner has a built-in visual assertion, expect(page).toHaveScreenshot(), that manages baselines for you. The JAVA binding does NOT have that runner feature. In Java you either (a) capture screenshots and compare them yourself against stored baselines, or (b) use a service like Applitools. This document shows both, done properly.
Part A — Visual Regression Testing
What It Is & When To Use It
Visual regression testing catches unintended UI changes — shifted layouts, broken CSS, missing images, wrong fonts — that functional assertions miss. You capture a screenshot, compare it to an approved baseline, and fail on meaningful differences.
- Great for: design systems, component libraries, marketing pages, responsive layouts, theme/dark-mode.
- Not for: highly dynamic data screens (use masking) or anything where pixels legitimately change every run.
- Two approaches in Java: (1) self-managed screenshot diffing, (2) Applitools (AI-based, cloud baselines).
Capturing Stable Screenshots
Before comparing anything, make the screenshot deterministic. Playwright’s screenshot options let you disable animations, mask dynamic regions and clip to an element — all available in Java.
import com.microsoft.playwright.options.ScreenshotAnimations;
import java.util.List;
// Full-page, animations frozen, dynamic areas masked out
byte[] png = page.screenshot(new Page.ScreenshotOptions()
.setFullPage(true)
.setAnimations(ScreenshotAnimations.DISABLED) // freeze CSS animations
.setMask(List.of( // hide volatile regions
page.locator(".timestamp"),
page.getByTestId("live-ads")))
.setMaskColor("#FF00FF")
.setPath(java.nio.file.Paths.get("target/shots/home.png")));
// Element-only screenshot (a single component)
byte[] card = page.getByTestId("product-card")
.screenshot(new Locator.ScreenshotOptions()
.setAnimations(ScreenshotAnimations.DISABLED));
- Animations/transitions — disable with setAnimations(DISABLED).
- Dynamic content (dates, ads, counters) — hide with setMask(...).
- Fonts not yet loaded — wait for a web-first assertion on visible text first.
- Viewport/scale differences — pin setViewportSize and run baselines on the same OS (ideally the Docker image).
Self-Managed Baseline Comparison
The DIY approach: on first run, save the screenshot as a baseline; on later runs, compare the new capture to it and fail if the difference exceeds a threshold. Use a small image-diff helper (any Java image-comparison library, or a pixel-by-pixel BufferedImage compare).
// Pseudocode for a VisualUtil.compare(...) helper you own
Path baseline = Paths.get("src/test/resources/baselines/home.png");
byte[] actual = page.screenshot(stableOptions());
if (!Files.exists(baseline)) {
Files.write(baseline, actual); // first run: seed the baseline
return; // (optionally fail to force review)
}
double diffRatio = ImageDiff.compare(Files.readAllBytes(baseline), actual);
Assertions.assertTrue(diffRatio < 0.02, // allow < 2% pixel difference
"Visual diff " + diffRatio + " exceeded threshold");
- Store baselines per browser/OS — a Chromium-on-Linux baseline will not match WebKit-on-Mac.
- Generate baselines in CI (the Docker image) so local machine fonts don’t pollute them.
- Keep a threshold (e.g. 1-2%) to absorb anti-aliasing noise; save the diff image on failure for review.
- Commit an "update baselines" switch so intentional UI changes are easy to re-approve.
Self-managed pixel diffing is fragile: anti-aliasing, sub-pixel font rendering and OS differences cause false positives. For anything beyond a few screens, a purpose-built visual tool (below) pays for itself.
Applitools Eyes — AI Visual Testing (recommended at scale)
Applitools adds an AI visual layer over Playwright: it manages baselines in the cloud, ignores rendering noise, supports regions/ignore-areas, and does cross-browser visual validation. It has an official Java SDK for Playwright.
<!-- pom.xml -->
<dependency>
<groupId>com.applitools</groupId>
<artifactId>eyes-playwright-java5</artifactId>
<version>LATEST</version>
</dependency>
import com.applitools.eyes.playwright.Eyes;
Eyes eyes = new Eyes();
eyes.setApiKey(System.getenv("APPLITOOLS_API_KEY"));
eyes.open(page, "SauceDemo", "Login page layout");
page.navigate("/");
eyes.check("Login", Target.window().fully()); // full-page visual checkpoint
// Login and check the inventory grid, ignoring a dynamic region
new LoginPage(page).loginAs("standard_user", "secret_sauce");
eyes.check("Inventory", Target.window()
.ignore(page.getByTestId("promo-banner")));
eyes.closeAsync(); // results appear in the Applitools dashboard
- AI ignores rendering noise (anti-aliasing) but flags real, human-visible changes.
- Baselines live in the cloud with a review/approve workflow — no images in your repo.
- One capture can be validated across many browser/OS/viewport combos (Ultrafast Grid).
- Alternatives: Percy (BrowserStack), or the TypeScript runner’s built-in toHaveScreenshot.
Part B — Accessibility (a11y) Testing
Why Automate Accessibility
- Legal/compliance: WCAG 2.1/2.2 AA is required for many public and enterprise apps.
- It catches real defects: missing labels, poor contrast, no keyboard access, bad ARIA.
- Automated checks catch ~30-50% of issues cheaply; the rest still need manual/AT testing.
- Bonus: if your tests use getByRole, you are already exercising the accessibility tree.
axe-core — The Industry-Standard Engine
axe-core is the de-facto accessibility rules engine. In Java you inject it into the page and run it via page.evaluate(), then parse the violations JSON with Jackson.
// Inject axe-core (from CDN, or setPath to a bundled axe.min.js)
page.addScriptTag(new Page.AddScriptTagOptions()
.setUrl("https://cdnjs.cloudflare.com/ajax/libs/axe-core/4.10.2/axe.min.js"));
// Run axe against the whole page and get the results as JSON
Object raw = page.evaluate("async () => await axe.run()");
String json = new ObjectMapper().writeValueAsString(raw);
JsonNode result = new ObjectMapper().readTree(json);
JsonNode violations = result.get("violations");
int count = violations.size();
// Log each violation for the report
for (JsonNode v : violations) {
System.out.println(v.get("impact").asText() + " | "
+ v.get("id").asText() + " | " + v.get("help").asText());
}
Assertions.assertEquals(0, count,
"Found " + count + " accessibility violations");
Wrap the inject + run + parse into one helper, e.g. A11yUtil.analyze(page) returning a list of violations, and A11yUtil.assertNoViolations(page). Attach the violations JSON to Allure so failures are actionable. Scope a scan to one component with axe.run(document.querySelector("#checkout")).
Tuning axe — rules, tags and scope
// Only WCAG 2.1 AA rules, and only inside the main content area
String script = "async () => await axe.run(document.querySelector('main'), "
+ "{ runOnly: { type: 'tag', values: ['wcag21aa'] } })";
Object res = page.evaluate(script);
| axe tag | Meaning |
|---|---|
| wcag2a / wcag2aa | WCAG 2.0 Level A / AA rules |
| wcag21a / wcag21aa | WCAG 2.1 Level A / AA rules |
| best-practice | Common a11y best practices beyond WCAG |
| cat.forms / cat.color | Category filters (forms, colour-contrast, etc.) |
ARIA Snapshots — Playwright-Native a11y
Playwright can serialize the accessibility tree of an element to a readable YAML "aria snapshot". It is a lightweight way to assert structure and accessible names without axe.
// Capture the ARIA snapshot (YAML string) — great for seeing the a11y tree
String yaml = page.getByRole(AriaRole.NAVIGATION).ariaSnapshot();
System.out.println(yaml);
// e.g. - navigation:
// - link "Home"
// - link "Cart"
// Assert the tree matches an expected snapshot (recent Playwright versions)
assertThat(page.getByRole(AriaRole.NAVIGATION)).matchesAriaSnapshot(
"- navigation:\n - link \"Home\"\n - link \"Cart\"");
ariaSnapshot checks STRUCTURE and accessible names (is the tree what you expect?). axe checks RULES/violations (contrast, labels, roles). Use ariaSnapshot for stable structural assertions, axe for compliance scanning — they complement each other.
Keyboard & Focus Accessibility
Much of accessibility is keyboard operability — something you can assert directly with Playwright.
// Tab through the form and assert focus order
page.getByPlaceholder("Username").focus();
page.keyboard().press("Tab");
assertThat(page.getByPlaceholder("Password")).isFocused();
page.keyboard().press("Tab");
assertThat(page.getByRole(AriaRole.BUTTON,
new Page.GetByRoleOptions().setName("Login"))).isFocused();
// Operate a control with the keyboard only
page.keyboard().press("Enter");
Weaving a11y Into Existing Tests
The cheapest strategy: add a lightweight axe scan at key states inside functional tests you already have (after login, on the cart, on checkout). One extra line per test grows coverage fast.
@Test
void inventoryIsAccessible() {
new LoginPage(page).open().loginAs("standard_user", "secret_sauce");
A11yUtil.assertNoViolations(page, "wcag21aa"); // your helper
}
Interview Q&A — Visual & Accessibility
Does Playwright Java have built-in visual assertions like the TS runner?
No — toHaveScreenshot() is a TypeScript test-runner feature.
In Java you capture screenshots and either compare them yourself against baselines or use a tool like Applitools/Percy. Be ready to explain this distinction — it is a common trap.
How do you make visual tests stable?
- Disable animations (setAnimations(DISABLED)); mask dynamic regions (setMask).
- Pin the viewport; generate baselines on one environment (Docker image).
- Wait for fonts/content via web-first assertions before capturing; allow a small diff threshold.
How do you run accessibility checks with Playwright + Java?
Inject axe-core with addScriptTag, run axe.run() via page.evaluate(), parse the violations JSON with Jackson, and assert zero violations (optionally filtered to wcag21aa). Attach the results to Allure.
What is an ARIA snapshot and when do you use it?
A YAML serialization of an element’s accessibility tree (roles + accessible names). Use ariaSnapshot()/matchesAriaSnapshot for stable structural assertions; use axe for rule-based WCAG compliance scanning.
Can automated tools find all accessibility issues?
No — automation catches roughly a third to half (contrast, labels, roles). Keyboard-operability and screen-reader experience still need manual and assistive-technology testing. Automation is a fast first line, not the whole story.
Visual + a11y covered — the specialized layers most notes skip. ✅
FAQs
How does visual regression testing work in Playwright?
Take a deterministic screenshot (animations disabled, dynamic areas masked), compare it with a stored baseline image within a tolerance, and fail the test when the difference exceeds the threshold. TypeScript has this built in with toHaveScreenshot().
Can automated tools find all accessibility issues?
No. Tools such as axe-core catch a good share of WCAG issues (missing labels, contrast, ARIA errors), but keyboard flow, meaningful text and screen-reader experience still need manual checks.
Why do visual tests become flaky?
Animations, dynamic data, fonts and rendering differences between machines. Freeze animations, mask dynamic regions, use fixed test data and generate baselines in the same environment as CI, for example the official Playwright Docker image.