Playwright removes most timing flakiness, but flaky tests still appear from shared data, test order, environments and network dependencies. This guide covers retries (and why they are a signal, not a fix), how to detect flaky tests, the common root causes, and a real-world troubleshooting playbook.

Retries

Level: L4.

What is it?

Automatically re-running a failed test a set number of times before marking it failed.

Why do we need it?

CI infrastructure has transient hiccups (network blips, slow cold starts). Retries absorb genuine flakiness from the environment — but must not become a way to hide real bugs.

How does it work?

retries: N re-runs a failing test up to N times. If any attempt passes, the test passes (often flagged “flaky”). Combine with trace: 'on-first-retry' to capture evidence only when a retry occurs.

Syntax

export default defineConfig({  retries: process.env.CI ? 2 : 0,  use: { trace: 'on-first-retry' },});
npx playwright test --retries=1
Per-test:
test.describe.configure({ retries: 1 });

Basic Example

retries: process.env.CI ? 2 : 0

Practical Example — detect vs mask flakiness

// In the report, retried-but-passed tests show as "flaky".// Treat "flaky" as a bug to fix, not a green light.retries: 2,reporter: [['html'], ['list']],use: { trace: 'on-first-retry', video: 'retain-on-failure' },

Line-by-Line Explanation

  • 2 retries in CI absorb transient infra failures.
  • on-first-retry trace captures exactly the run that flaked, giving evidence to diagnose the root cause.
  • The “flaky” status in the report is your signal to investigate (Module 47).

Common Mistakes

  • High retries globally to force green suites → real bugs slip through.
  • No trace on retry → flakiness invisible and undiagnosable.
  • Ignoring the “flaky” tag instead of fixing the underlying issue.

Best Practices

  • Retry in CI only, small counts (1–2).
  • Always capture a trace on retry.
  • Track flaky tests and fix root causes; don’t let retries be permanent.

Interview Questions

  • Q: When should you use retries? A: To absorb transient CI/infra flakiness — not to mask product/test bugs.
  • Q: Downside of high retries? A: They hide real intermittent defects and slow the suite.
  • Q: How to investigate a retried failure? A: Enable trace: 'on-first-retry' and inspect the flaking run.

Practice Exercise

Set retries: 2 and trace: 'on-first-retry'. Introduce a flaky condition (e.g., a race), observe the “flaky” report status, then fix the race so retries aren’t needed.

Real-World Scenario

A team hid chronic flakiness behind retries: 5. When they capped it at 1 and treated every “flaky” as a ticket, they found and fixed a real race condition — the suite got faster and more trustworthy.

Advertisement

Handling Flaky Tests

Level: L4–L5. A signature senior/SDET topic.

What is it?

A flaky test passes sometimes and fails other times without code changes. This module is about finding and eliminating the causes.

Why do we need it?

Flaky suites destroy trust (“just re-run it”), hide real bugs, and slow delivery. Reliability is a core quality-engineering responsibility.

Root causes & fixes

Cause Fix
--------------------------------------------------------------
Timing / races web-first assertions; waitForResponse; no hard waits
Shared/duplicate test data unique data per test (Faker/UUID); isolation
Test interdependence make each test independent; fresh context
Animations/transitions assert final state; disable animations if possible
Network/3rd-party flakiness mock/route non-essential calls; retry infra only
Auto-generated locators role/text/testid instead of volatile ids
Time/locale/timezone fix TZ/locale in config; compute dates dynamically
Under-provisioned CI fewer workers; realistic timeouts

How to diagnose

npx playwright test flaky.spec.ts --repeat-each=20   # reproducenpx playwright test --retries=0                      # see raw failure rate

Enable trace: 'on-first-retry' and read the failing run in the trace viewer.

Basic Example — the classic race fix

// ❌ flaky: reads before value settlesexpect(await page.locator('.total').textContent()).toBe('$30');// ✅ stable: retries until settledawait expect(page.locator('.total')).toHaveText('$30');

Practical Example — isolation + unique data

test('register unique user', async ({ page }) => {  const email = `qa+${crypto.randomUUID()}@example.com`; // unique per run  // ...register with `email`...});

Line-by-Line Explanation

  • Web-first assertions remove timing races (the #1 cause).
  • Unique data removes cross-test collisions under parallelism.
  • --repeat-each reproduces intermittent failures so you can confirm a fix.

Common Mistakes

  • “Fixing” flakiness by adding waitForTimeout (masks, doesn’t cure).
  • Bumping retries instead of finding root cause.
  • Leaving flaky tests quarantined forever.

Best Practices

  • Reproduce with --repeat-each; read traces.
  • Prefer web-first assertions and network-based sync.
  • Ensure isolation + unique data.
  • Quarantine with a ticket, fix promptly, then un-quarantine.
  • Track a flakiness metric over time.

Interview Questions

  • Q: How do you handle flaky tests? A: Reproduce (--repeat-each), diagnose via traces, fix root cause (waits/data/isolation), retries only for infra.
  • Q: Most common cause? A: Timing/races — fixed by web-first assertions and event-based waits.
  • Q: How to prevent data-related flakiness? A: Unique data per test and strict isolation.
  • Q: Is increasing retries a fix? A: No — it masks the problem; use it only for transient infra.

Practice Exercise

Take a flaky test (add an artificial race), reproduce it with --repeat-each=20, diagnose via trace, and fix it with a web-first assertion. Prove it passes 20/20.

Real-World Scenario

A “10% flaky” checkout test failed randomly in CI. The trace revealed a toast animation intercepting the click. Waiting for the toast to disappear (final-state assertion) fixed it permanently — no retries required.

Real-World Troubleshooting

Level: L4–L5. Format: Problem → Cause → Debug → Solution → Prevention.

What is it?

A field guide to the failures you’ll actually hit, each dissected the way a senior engineer debugs them.

Why do we need it?

Fixing real failures fast is the day-to-day of automation. Pattern-recognition here separates seniors from juniors.

  • Cause: wrong/volatile selector, element in an iframe/shadow root, or content not loaded yet.
  • Debug: open the trace; check the DOM snapshot at that step; try the locator in Playwright Inspector/--debug.
  • Solution: switch to a semantic locator (getByRole/getByTestId); scope into the frame; assert a load signal first.
  • Prevention: semantic locators + data-testid; sync on real signals.
  • Cause: locator matches multiple elements.
  • Debug: run it; Playwright lists all matches.
  • Solution: narrow with getByRole({ name }), .filter({ hasText }), or scope to a container; use .first() only when truly appropriate.
  • Prevention: design unique, scoped locators; embrace strict mode as a correctness guard.
  • Cause: element never becomes actionable (hidden, disabled, covered by overlay), or wrong locator.
  • Debug: trace shows it waiting; inspect if it’s covered (toast/modal) or off-screen.
  • Solution: wait for the blocker to clear; assert enabled/visible; fix the locator; raise timeout only if genuinely slow.
  • Prevention: assert preconditions; handle overlays; avoid racing the UI.
  • Cause: covered by an overlay/animation, zero-size, or outside viewport.
  • Debug: screenshot/trace at the step; check for animations or sticky banners.
  • Solution: wait for animation/toast to finish; dismiss overlays; scrollIntoViewIfNeeded.
  • Prevention: assert final state; consider disabling animations in test builds.
  • Cause: using a page after its context closed, or a popup you closed; unhandled async after test end.
  • Debug: check where the context/page is closed vs used; look for missing await.
  • Solution: keep references valid; await all actions; don’t use a closed popup.
  • Prevention: correct fixture scoping; always await.
  • Cause: expired/missing storageState, wrong env credentials.
  • Debug: check auth/*.json freshness; log the target env; watch the network for 401s.
  • Solution: regenerate storageState in setup each run; verify env creds.
  • Prevention: setup project regenerates auth per run; validate env vars.
  • Cause: missing browser deps, different timezone/locale, headless-only rendering, slower runner, missing env var.
  • Debug: download the CI trace/video; compare env; run headless locally; check TZ/locale.
  • Solution: playwright install --with-deps; pin TZ/locale in config; fix env vars; realistic timeouts.
  • Prevention: run in the Playwright Docker image locally too; parity between local and CI.
  • Cause: timing races, shared data, animations, network flakiness. (See Module 47.)
  • Debug: --repeat-each=20; read trace of the failing run.
  • Solution: web-first assertions, unique data, isolation, mock non-essential calls.
  • Prevention: the reliability practices in Module 59.
  • Cause: shared/duplicate data or shared mutable state across workers.
  • Debug: run serially (passes) vs parallel (fails) → confirms data collision.
  • Solution: unique data per test; remove shared mutable singletons.
  • Prevention: Faker/UUID data; strict isolation.
  • Cause: slow/failing backend or third-party (ads/analytics) hanging.
  • Debug: trace network tab; identify the slow/failed call.
  • Solution: mock/abort non-essential calls; sync via waitForResponse; raise timeout for genuinely slow endpoints.
  • Prevention: block third-party noise; deterministic mocks for unstable deps.
  • Cause: contract drift, auth header missing, wrong env/data.
  • Debug: log status + body; check auth; verify env base URL.
  • Solution: fix auth/data; update expectations if the contract legitimately changed; add schema validation.
  • Prevention: JSON-schema checks; centralized auth; env validation.
  • Cause: unset/incorrect BASE_URL/TEST_ENV.
  • Debug: log the env at run start; inspect the config resolution.
  • Solution: set the correct env var; add a startup validator.
  • Prevention: validate required env vars; PROD guardrails (Module 55).
  • Cause: browsers not installed on the runner/container.
  • Debug: error explicitly says the executable is missing.
  • Solution: npx playwright install --with-deps (or use the Playwright Docker image).
  • Prevention: bake install into every pipeline/image; pin versions.
  • Cause: wrong relative path, missing dir, read-only location.
  • Debug: log the resolved absolute path; check the dir exists.
  • Solution: path.join(__dirname, ...); ensure output dirs exist; use download.saveAs.
  • Prevention: absolute paths; create/clean output dirs in setup.

Debugging toolkit (memorize)

--debug (Inspector) • --ui (UI mode) • show-trace (trace viewer) • --repeat-each=N (repro flakiness) • page.pause() • trace: 'on-first-retry' • --headed.

Interview Questions

  • Q: “Passes locally, fails in CI” — how do you debug? A: Pull the CI trace/video; check browser deps (--with-deps), TZ/locale, headless rendering, env vars; achieve local/CI parity via Docker.
  • Q: Strict mode violation fix? A: Narrow/scope the locator so it’s unique; don’t blindly .first().
  • Q: Element “not actionable” — likely cause? A: An overlay/animation covering it or it’s off-screen — wait for final state / dismiss the overlay.

Practice Exercise

Reproduce three of these (strict-mode violation, a race-based flake, a CI-only failure via a missing env var). For each, write out the Problem→Cause→Debug→Solution→Prevention and fix it.

Real-World Scenario

A high-severity “checkout broken in CI only” turned out to be P7 + P4: a CI-locale cookie banner (only shown in that timezone) covered the pay button. The trace made it obvious in minutes — exactly the payoff of disciplined troubleshooting.

FAQs

Should I enable retries in Playwright?

In CI, one or two retries keep pipelines moving, and Playwright marks tests that pass on retry as flaky. Treat every flaky result as a bug to fix, using the trace from the first failure.

What causes flaky Playwright tests?

Shared test data, dependence on test order, real third-party services, missing awaits, assertions on values instead of web-first assertions, animations and slow CI environments.