Running Playwright in CI
A Playwright suite delivers value when it runs automatically on every pull request and blocks regressions. CI environments add their own challenges, though: browsers and system libraries must be installed, the app must be started and reachable, large suites need sharding across machines, and failures must be debuggable without reproducing them locally, which is what traces, screenshots, videos, and HTML reports are for.
A well-tuned pipeline runs quickly, fails for real reasons, and gives developers everything they need to fix a failure in minutes.
TL;DR
- Install browsers with
npx playwright install --with-deps, or use the official Playwright Docker image. - Start the app with
webServerin the config (or a service container) and setbaseURL. - Use
retries: 2in CI withtrace: 'on-first-retry', and upload the HTML report and traces as artifacts. - Shard large suites (
--shard=1/4) across parallel jobs, with blob reports merged viamerge-reports. - Treat flaky tests as bugs: track them, fix the cause, and quarantine temporarily if necessary.
- Keep visual snapshots in a consistent environment (the same OS and image), since fonts and rendering differ across platforms.
Quick Example
A sharded GitHub Actions workflow with merged reports:
Core Concepts
Browsers and System Dependencies
Playwright downloads specific browser builds matching its version. In CI:
npx playwright install --with-deps [chromium]installs browsers and required OS libraries. Install only the browsers you test to save time.- The official Docker image (
mcr.microsoft.com/playwright:v<version>-noble) ships with browsers and dependencies preinstalled. Pin the tag to your@playwright/testversion. - Caching browsers between runs saves download time, but invalidate the cache when Playwright's version changes.
Starting the Application
webServerin the config builds and starts the app and waits for a URL before tests run, and it can start several servers (API plus frontend).- Alternatively, run against a preview deployment (Vercel, Netlify, or ephemeral environments) by setting
BASE_URL. - Backing services (databases, Redis) can run as CI service containers or via Docker Compose. See Docker Compose.
Retries, Traces, and Artifacts
- Retries absorb rare infrastructure flakes, and tests passing only on retry are reported as flaky.
- Trace Viewer (
trace: 'on-first-retry') records a timeline of actions, DOM snapshots, network calls, console logs, and source for failed tests. Open it withnpx playwright show-trace trace.zip, or on trace.playwright.dev. It's the single most useful debugging artifact. - Upload the HTML report, traces, screenshots, and videos as artifacts, and use the
githubreporter for inline annotations on pull requests.
Sharding
--shard=x/y splits tests across y machines. Each shard writes a blob report, and a final job merges them into one HTML report with merge-reports. Combine it with a CI matrix (see matrix builds). Balance shards by keeping test files similarly sized, since sharding splits by test and file, and with fullyParallel: true it splits more evenly.
Workers
CI machines often have few CPUs, and too many workers cause resource contention and timeouts. Start with 1–2 workers per shard on standard runners and tune from there. Scale horizontally with shards rather than vertically with workers.
Flaky Test Management
- Treat flakiness as a defect: find the root cause, usually missing web-first assertions, shared state between tests, animations, or real-time data.
- Use the report's flaky list and history tools to track offenders.
- Quarantine with
test.fixmeor a tag excluded from blocking runs, with an owner and a deadline, rather than living with random red builds. --repeat-each=10and--fail-on-flaky-testshelp verify fixes.
Visual Regression Testing
await expect(page).toHaveScreenshot() compares against baselines. Rendering differs across OSes, fonts, and GPUs, so generate and compare snapshots in the same environment, typically the Docker image, both locally and in CI. Mask dynamic regions (timestamps, ads), disable animations, and set thresholds thoughtfully.
Speeding Up Suites
- Log in once with storage state, and seed data via API.
- Run only affected tests on PRs (
--only-changedagainst the base branch, or path filters), and the full suite on main or nightly. - Block third-party scripts, and mock slow external services (see network mocking).
- Test the most critical journeys on all browsers, and the rest on Chromium only.
- Fail fast with
--max-failureson PRs when the build is clearly broken.
Best Practices
Make CI Match Local
Pin Playwright and browser versions, use the same Docker image for visual tests, and provide a script (npm run test:e2e) that developers run locally the same way CI does.
Keep E2E Focused on Critical Journeys
End-to-end tests are the most expensive tests. Cover key user journeys, and push detailed logic coverage down to unit and integration tests.
Gate Merges on E2E Results
Make the E2E job a required status check. A suite that can be ignored will be ignored, especially when it's flaky.
Protect Secrets in Artifacts
Traces and videos can contain credentials and personal data from test accounts. Mask inputs, use test-only accounts, and set short artifact retention.
Common Mistakes
Committing test.only
A stray test.only makes CI run a single test and report green. Set forbidOnly: !!process.env.CI.
Too Many Workers on Small Runners
Running 8 workers on a 2-vCPU runner makes every test slower and timeouts frequent, which looks like flakiness. Reduce workers and add shards.
Comparing Screenshots Across Operating Systems
Baselines generated on a Mac fail on Linux CI because of font rendering differences. Generate baselines in the CI environment (or the Docker image), and commit those.
FAQ
How do I install Playwright browsers in CI?
Run npx playwright install --with-deps (optionally naming specific browsers) after installing npm dependencies, or use the official Playwright Docker image, which includes browsers and system libraries. Pin the image version to match your @playwright/test version.
How do I debug a test that only fails in CI?
Enable traces (trace: 'on-first-retry' or 'retain-on-failure'), upload them as artifacts, and open them in Trace Viewer to inspect each action, DOM snapshot, network request, and console message. Screenshots and videos help too. Reproducing locally in the same Docker image often reveals environment differences.
How does sharding work in Playwright?
--shard=1/4 runs the first quarter of tests, and so on. Run each shard in a parallel CI job with the blob reporter, then merge the blob reports into one HTML report with npx playwright merge-reports.
Should I use retries in CI?
Yes, a small number (1–2) to absorb rare infrastructure hiccups, while tracking tests that pass only on retry as flaky and fixing them. Retries shouldn't hide persistent flakiness.
Related Topics
- Playwright — The testing framework overview
- Playwright Fixtures — Configuration and projects
- GitHub Actions — Running the pipeline
- GitHub Actions Matrix Builds — Sharding with matrices
- CI/CD — Continuous integration practices
- Playwright Authentication — Faster suites with storage state