DEV Community

Cover image for Ditching Selenium for Playwright + Pytest: What We Gained (and Why We Never Looked Back)
Khanjan Rathi
Khanjan Rathi

Posted on

Ditching Selenium for Playwright + Pytest: What We Gained (and Why We Never Looked Back)

We had a test suite that ran. It just couldn't be trusted. Flaky failures, slow CI pipelines, and painful debugging had become normal. Switching to Playwright improved all three: regression time went from ~6 hours to ~3.5, and the flaky rate went from ~35% to ~10%.

I work in quality engineering on a large, data-heavy enterprise web app: dashboards, sortable tables, interactive charts, and lots of asynchronous loading. Our end-to-end suite was built on Selenium + Java. It worked on paper. In practice, three problems kept coming back.

The Problem with Our Selenium Suite

1. Flakiness
Classic Selenium sends a command to the browser and hopes the page is ready. When it isn't, you get a NoSuchElementException or a stale element. We patched over it the usual way:

Thread.sleep(3000); // "should be enough"
driver.findElement(By.id("save")).click();
Enter fullscreen mode Exit fullscreen mode

The sleeps made tests slower, and they were still unreliable. Three seconds is too long on a fast machine and too short on a busy CI runner.

2. Speed
Tests ran one after another, and each one paid a high browser startup cost. A full regression run took far longer than a CI pipeline could comfortably absorb, so people stopped running it on every PR.

3. Debugging
A failing test gave you a stack trace, plus a screenshot if someone had remembered to capture one. Reproducing intermittent failures was mostly guesswork.

To be fair, a lot of this was how we used Selenium, not only Selenium itself. Explicit waits help, and Selenium 4 has improved a lot. But Playwright makes the right approach the default, and that turned out to matter more than anything else.

Why Playwright Fixed It

Playwright controls browsers through a protocol connection (the Chrome DevTools Protocol for Chromium, and equivalent protocols for Firefox and WebKit) instead of the classic WebDriver request/response model. It sees what the browser is doing and checks that an element is attached, visible, stable, enabled, and able to receive events before it acts. You don't write the wait logic yourself.

The same interaction from above, in Playwright:

page.get_by_role("button", name="Save").click()
expect(page.get_by_text("Changes saved")).to_be_visible()
Enter fullscreen mode Exit fullscreen mode

No sleeps. click() waits until the button can actually be clicked, and expect() retries until the assertion passes or the timeout is reached.

What changed for us:

  • Auto-waiting. Every action waits for actionability. We deleted every hardcoded sleep on day one.
  • Parallel execution. With pytest-xdist, tests run concurrently, and full regression time dropped from ~6 hours to ~3.5 hours.
  • Traces, videos, screenshots. On failure, Playwright can save a trace file: a full timeline of actions, DOM snapshots, network calls, and console logs. Average debugging time per failure dropped from about 2 hours to about 30 minutes.
  • CI-friendly by default. Headless is the default. A single flag switches to headed mode for local debugging.

How We Structured the Framework

We paired Playwright with Python + Pytest (via the official pytest-playwright plugin) and built it around a clean Page Object Model. Each layer has one job:

Layer What lives here
tests/ Test files grouped by feature area, e.g. dashboard/, reports/, settings/
pages/ All browser interactions. Tests never touch the DOM directly.
components/ Reusable UI pieces: data tables, cards, chart tooltips, modals
data/ Test data factories. Nothing is hardcoded inside a test file.
utils/ HTML report enrichment, artifact embedding, failure screenshots

A page object

# pages/login_page.py
from playwright.sync_api import Page, expect


class LoginPage:
    def __init__(self, page: Page):
        self.page = page
        self.username = page.get_by_label("Username")
        self.password = page.get_by_label("Password")
        self.submit = page.get_by_role("button", name="Sign in")

    def open(self):
        self.page.goto("/login")

    def login(self, user: str, password: str):
        self.username.fill(user)
        self.password.fill(password)
        self.submit.click()
        expect(self.page.get_by_role("navigation")).to_be_visible()
Enter fullscreen mode Exit fullscreen mode

A test that reads like a spec

# tests/auth/test_login.py
from pages.login_page import LoginPage
from data.users import standard_user


def test_user_can_log_in(page):
    login = LoginPage(page)
    login.open()
    login.login(standard_user.name, standard_user.password)
Enter fullscreen mode Exit fullscreen mode

The test describes behavior. The page object owns the locators. When the UI changes, you fix it in one place.

Config: parallel runs and artifacts on failure

# pytest.ini
[pytest]
base_url = https://staging.example.com
addopts =
    -n auto
    --browser chromium
    --tracing retain-on-failure
    --video retain-on-failure
    --screenshot only-on-failure
Enter fullscreen mode Exit fullscreen mode

-n auto comes from pytest-xdist and spreads tests across all CPU cores. The other flags come from pytest-playwright, so every failed test leaves behind a trace, a video, and a screenshot automatically. Nobody has to remember to capture them.

Debugging locally:

pytest tests/auth --headed --slowmo 500 -n 0
playwright show-trace test-results/<test-name>/trace.zip
Enter fullscreen mode Exit fullscreen mode

The trace viewer is the feature that won over the rest of the team. You can step through each action, see the DOM before and after it, and inspect every network request that happened along the way.

The Results

Metric Selenium + Java Playwright + Pytest Change
Full regression run time ~6 hours ~3.5 hours ~42% faster
Flaky test rate ~35% ~10% ~71% fewer flaky failures
Average time to debug a failure ~2 hours ~30 minutes 4x faster
  • ✅ Flakiness dropped from ~35% to ~10%. Auto-waiting removed the timing guesses behind most of our failures.
  • ✅ Regression runs went from ~6 hours to ~3.5 hours. Parallel execution with pytest-xdist removed the serial bottleneck, which means we can now run the full suite well within a working day.
  • ✅ Debugging went from ~2 hours to ~30 minutes per failure. Trace files show the root cause right away instead of leaving us to dig through logs.
  • ✅ Lower onboarding cost. Python + Pytest has a gentler learning curve than a Java Selenium stack with its own build toolchain.

Things to Watch Out For

The migration was worth it, but it isn't free:

  • It's a rewrite, not a port. Locator strategies, waits, and fixtures all change. Budget for it, and migrate feature by feature instead of all at once.
  • Parallel tests need isolated data. Once tests run concurrently, any shared state (the same user, the same record) will cause collisions. Data factories that generate unique data per test fixed this for us.
  • Prefer user-facing locators. get_by_role, get_by_label, and get_by_text hold up much better than long CSS or XPath chains, and they nudge your app toward better accessibility.

Final Thoughts

If your team maintains a Selenium suite that feels more like a liability than an asset, the move to Playwright is worth it. The reliability gains alone justify the effort, and faster CI and easier debugging come with it.

Have you made the switch, or are you still on the fence? I'd like to hear what's holding you back, or what surprised you, in the comments. 👇

Top comments (0)