We had a test suite that ran. It just couldn't be trusted. Flaky failures, slow CI pipelines, and painful debugging had become normal. Switching to Playwright improved all three: regression time went from ~6 hours to ~3.5, and the flaky rate went from ~35% to ~10%.
I work in quality engineering on a large, data-heavy enterprise web app: dashboards, sortable tables, interactive charts, and lots of asynchronous loading. Our end-to-end suite was built on Selenium + Java. It worked on paper. In practice, three problems kept coming back.
The Problem with Our Selenium Suite
1. Flakiness
Classic Selenium sends a command to the browser and hopes the page is ready. When it isn't, you get a NoSuchElementException or a stale element. We patched over it the usual way:
Thread.sleep(3000); // "should be enough"
driver.findElement(By.id("save")).click();
The sleeps made tests slower, and they were still unreliable. Three seconds is too long on a fast machine and too short on a busy CI runner.
2. Speed
Tests ran one after another, and each one paid a high browser startup cost. A full regression run took far longer than a CI pipeline could comfortably absorb, so people stopped running it on every PR.
3. Debugging
A failing test gave you a stack trace, plus a screenshot if someone had remembered to capture one. Reproducing intermittent failures was mostly guesswork.
To be fair, a lot of this was how we used Selenium, not only Selenium itself. Explicit waits help, and Selenium 4 has improved a lot. But Playwright makes the right approach the default, and that turned out to matter more than anything else.
Why Playwright Fixed It
Playwright controls browsers through a protocol connection (the Chrome DevTools Protocol for Chromium, and equivalent protocols for Firefox and WebKit) instead of the classic WebDriver request/response model. It sees what the browser is doing and checks that an element is attached, visible, stable, enabled, and able to receive events before it acts. You don't write the wait logic yourself.
The same interaction from above, in Playwright:
page.get_by_role("button", name="Save").click()
expect(page.get_by_text("Changes saved")).to_be_visible()
No sleeps. click() waits until the button can actually be clicked, and expect() retries until the assertion passes or the timeout is reached.
What changed for us:
- Auto-waiting. Every action waits for actionability. We deleted every hardcoded sleep on day one.
-
Parallel execution. With
pytest-xdist, tests run concurrently, and full regression time dropped from ~6 hours to ~3.5 hours. - Traces, videos, screenshots. On failure, Playwright can save a trace file: a full timeline of actions, DOM snapshots, network calls, and console logs. Average debugging time per failure dropped from about 2 hours to about 30 minutes.
- CI-friendly by default. Headless is the default. A single flag switches to headed mode for local debugging.
How We Structured the Framework
We paired Playwright with Python + Pytest (via the official pytest-playwright plugin) and built it around a clean Page Object Model. Each layer has one job:
| Layer | What lives here |
|---|---|
tests/ |
Test files grouped by feature area, e.g. dashboard/, reports/, settings/
|
pages/ |
All browser interactions. Tests never touch the DOM directly. |
components/ |
Reusable UI pieces: data tables, cards, chart tooltips, modals |
data/ |
Test data factories. Nothing is hardcoded inside a test file. |
utils/ |
HTML report enrichment, artifact embedding, failure screenshots |
A page object
# pages/login_page.py
from playwright.sync_api import Page, expect
class LoginPage:
def __init__(self, page: Page):
self.page = page
self.username = page.get_by_label("Username")
self.password = page.get_by_label("Password")
self.submit = page.get_by_role("button", name="Sign in")
def open(self):
self.page.goto("/login")
def login(self, user: str, password: str):
self.username.fill(user)
self.password.fill(password)
self.submit.click()
expect(self.page.get_by_role("navigation")).to_be_visible()
A test that reads like a spec
# tests/auth/test_login.py
from pages.login_page import LoginPage
from data.users import standard_user
def test_user_can_log_in(page):
login = LoginPage(page)
login.open()
login.login(standard_user.name, standard_user.password)
The test describes behavior. The page object owns the locators. When the UI changes, you fix it in one place.
Config: parallel runs and artifacts on failure
# pytest.ini
[pytest]
base_url = https://staging.example.com
addopts =
-n auto
--browser chromium
--tracing retain-on-failure
--video retain-on-failure
--screenshot only-on-failure
-n auto comes from pytest-xdist and spreads tests across all CPU cores. The other flags come from pytest-playwright, so every failed test leaves behind a trace, a video, and a screenshot automatically. Nobody has to remember to capture them.
Debugging locally:
pytest tests/auth --headed --slowmo 500 -n 0
playwright show-trace test-results/<test-name>/trace.zip
The trace viewer is the feature that won over the rest of the team. You can step through each action, see the DOM before and after it, and inspect every network request that happened along the way.
The Results
| Metric | Selenium + Java | Playwright + Pytest | Change |
|---|---|---|---|
| Full regression run time | ~6 hours | ~3.5 hours | ~42% faster |
| Flaky test rate | ~35% | ~10% | ~71% fewer flaky failures |
| Average time to debug a failure | ~2 hours | ~30 minutes | 4x faster |
- ✅ Flakiness dropped from ~35% to ~10%. Auto-waiting removed the timing guesses behind most of our failures.
- ✅ Regression runs went from ~6 hours to ~3.5 hours. Parallel execution with
pytest-xdistremoved the serial bottleneck, which means we can now run the full suite well within a working day. - ✅ Debugging went from ~2 hours to ~30 minutes per failure. Trace files show the root cause right away instead of leaving us to dig through logs.
- ✅ Lower onboarding cost. Python + Pytest has a gentler learning curve than a Java Selenium stack with its own build toolchain.
Things to Watch Out For
The migration was worth it, but it isn't free:
- It's a rewrite, not a port. Locator strategies, waits, and fixtures all change. Budget for it, and migrate feature by feature instead of all at once.
- Parallel tests need isolated data. Once tests run concurrently, any shared state (the same user, the same record) will cause collisions. Data factories that generate unique data per test fixed this for us.
-
Prefer user-facing locators.
get_by_role,get_by_label, andget_by_texthold up much better than long CSS or XPath chains, and they nudge your app toward better accessibility.
Final Thoughts
If your team maintains a Selenium suite that feels more like a liability than an asset, the move to Playwright is worth it. The reliability gains alone justify the effort, and faster CI and easier debugging come with it.
Have you made the switch, or are you still on the fence? I'd like to hear what's holding you back, or what surprised you, in the comments. 👇
Top comments (0)