Do not start a test automation strategy by choosing Selenium, Playwright, or another framework. Start with the failure you need to prevent. Then find the smallest part of the system that can reveal it with a fast, reliable test.
Browser automation still matters. Use it as one part of a test strategy, not as the starting point for every test.
TLDR
- Start with risks to users and the business, not a preferred test tool.
- Test each risk through the smallest part of the system that gives enough confidence.
- Keep only reliable checks on the continuous integration (CI) path, which runs for each code change.
- Treat an inconsistent test as defective evidence, even when a retry passes.
- Link pre-release tests to signals from the running system.
Choose the test before the tool
Selenium controls a web browser. It can prove that an important user journey works with real pages, navigation, and code that runs in the browser.
That does not make it the best tool for every rule in a system. The Selenium project's own guidance says that Selenium makes functional user interaction easier but does not help a team write a well-architected test suite.
A browser test is usually a poor first choice for proving that:
- A money calculation handles rounding correctly.
- An access rule rejects a forbidden action.
- A system can handle a new field from an application programming interface (API).
- A retry does not create a duplicate transaction.
- Two systems still agree on a message format.
- A service stays within its response-time goal under load.
These risks live in different parts of the system. Running every check through a browser adds unrelated infrastructure, data, network, and user interface behavior. Those extra parts make failures harder to understand.
Do not ask only, "Can Selenium automate this?" Ask, "What is the least costly reliable evidence that this failure will not reach a user?"
Cost includes run time, test setup, maintenance, false alarms, and time spent finding the cause. A fast test that takes an hour to explain is not cheap.
Match each risk to a test boundary
A test layer is the boundary through which a test observes or controls the system. A test can check code, one component, a service API, an agreement between systems, a browser journey, or the deployed system.
No test layer is always right. The National Institute of Standards and Technology guidance, NIST IR 8397 recommends adapting software checks to the product, technology, tools, and development process.
This matters in large or regulated systems. Financial, security, privacy, audit, and resilience risks may need different evidence. Resilience means the ability to keep working through failures.
Use this worksheet before selecting a framework.
Use a risk-to-test worksheet
Create one entry for each failure that would cause real harm:
- Business effect: What happens to a user or the organization?
- Cause: What system behavior could create that effect?
- Test boundary: Where can you control and observe the behavior with the fewest unrelated parts?
- Main check: What is the fastest test that behaves enough like the real system?
- Broader check: What could the main check miss because it uses stand-ins or covers less of the system?
- Deadline: Must the result arrive while coding, before merge, before release, or after deployment?
- Pipeline rule: Does the check block a change, report only, run on a schedule, or run only for relevant changes?
- Owner: Which team fixes the product or test when the check fails?
- Failure details: Which logs, request traces, random seeds, screenshots, and environment details will you save?
- Production signal: Which measurement, log, trace, probe, or business rule would reveal a missed failure?
This worksheet is a guide, not a fixed classification. One risk may need a fast main check and a small set of broader checks that use more of the real system.
Synthetic example: a regulated transfer application
The application and every detail below are synthetic. They do not come from an employer, customer, or production system.
Assume an invented money transfer system with:
- A web application for creating and approving transfers.
- Services that identify users and enforce access rules.
- A transfer API.
- A database with a transactional outbox, which saves outgoing messages in the same database transaction as the related change.
- A message broker that routes messages to a settlement system.
- Logs, measurements, and traces that follow a request across services.
The system requires separation between the person creating a transfer and the person approving it. It must also prevent duplicate settlement when a request or message is retried.
Risk 1: an invalid amount is accepted
Test calculation and policy rules with small code-level tests. Add property-based tests, which generate many inputs and check that a rule always holds. These tests can cover edge cases without starting a browser or database.
Add a small set of service-level checks. They should prove that the service reads requests, handles currencies, and calls the policy code correctly. In production, count accepted and rejected decisions by safe reason categories.
Risk 2: a consumer and provider disagree
Use contract tests where services or messages meet. A contract test checks that separately tested systems agree on the data they exchange. The Pact documentation provides one public description of the pattern.
Keep a small integration smoke test, which is a basic check across real components. It can catch problems that a contract cannot, such as routing, credentials, or message broker setup.
Risk 3: a retry creates a duplicate transfer
Test the transfer component with a temporary database. Check idempotency, which means that repeating the same request does not repeat its effect. Include repeated requests, requests that run at the same time, and a forced failure between saving data and publishing a message.
Run slower failure scenarios after merge or against a release candidate. Track duplicate record conflicts, repeated request keys, unpublished message age, settlement results, and request traces across retries.
Risk 4: an unauthorized user approves a transfer
Test access decisions in the policy code. Then prove that the service enforces those decisions. Include denied cases, not only successful ones.
Keep a small set of browser-level journeys to prove that the user interface and service enforce the same rule. Use a clear standard for security checks. One example is the Application Security Verification Standard from the Open Worldwide Application Security Project (OWASP).
Risk 5: correct code is misconfigured after deployment
Run a safe test journey after deployment. It must not change real customer data. Compare the outside result with internal logs and measurements from the services involved.
The Google monitoring guidance explains why outside symptoms and internal signals answer different questions.
Set a deadline for feedback
A useful test reports a problem while the developer still remembers the change.
DORA, a software delivery research program, recommends automated test feedback in less than ten minutes on developer computers and in CI. I treat ten minutes as a design goal, not a universal limit. Performance, resilience, security, and full-system tests may take longer.
Measure two times:
- Time to signal: How long does it take to get a result after a change?
- Time to diagnosis: How long does the owner need to find the cause and choose the next action?
Run checks when their results are most useful:
- Local checks cover focused code and component behavior.
- Pull request checks block changes that fail reliable safety checks.
- Checks after merge cover broader connections and compatibility.
- Scheduled or release checks cover costly performance, resilience, and security scenarios.
- Checks after deployment confirm the live configuration and behavior.
Do not fix a slow pipeline only by adding more machines. First ask whether each test uses more of the system than it needs.
Treat inconsistent tests as defects
A flaky test produces inconsistent outcomes without a relevant change to the system. The cause may be the test, infrastructure, shared data, an uncontrolled dependency, or a timing bug in the product.
Retries can help identify the cause and collect details. They should not turn a first-run failure into a normal pass. Playwright, for example, distinguishes tests that pass initially from those that pass only after a retry.
Use an explicit policy:
- Save the first-run result.
- Label tests that pass only after a retry as flaky.
- Save the random seed, environment, versions, timings, logs, and traces.
- Remove a flaky test from the blocking path only with an owner, issue, and review date.
- Keep results from non-blocking flaky tests visible.
- Delete tests that no longer protect against real harm.
Test isolation prevents false alarms. Selenium's shared-state guidance recommends isolated data and a new browser instance for each test.
Give every check an owner
A central test automation team should not permanently own tests for code it cannot change.
DORA reports better outcomes when developers are primarily responsible for creating and maintaining automated tests, while testers work alongside them and continue exploratory, usability, and acceptance testing.
Use clear ownership:
- Product or service teams own tests for their code and interfaces.
- Platform teams own test machines, environments, templates, saved results, and shared reliability tools.
- Security, risk, and control specialists define required evidence and review controls independently when needed.
- Test specialists improve test design, hands-on exploration, and suite health without becoming a handoff queue.
Every required check needs a named owner and a clear next step when it fails.
Connect tests to the running system
Observability means understanding a running system from the data it produces. Tests provide controlled evidence before and during a release. Observability shows what happens with the real configuration and workload. You need both.
Save enough details to find the cause of a failed test:
- Exact application and dependency versions.
- Test data identifiers and random seeds.
- Structured logs and an identifier that links events from one request.
- Relevant request and response details with sensitive fields removed.
- Screenshots or browser traces for user interface failures.
- Time spent in each test stage and dependency.
For the running system, connect each worksheet entry to telemetry, which is data that the system sends about its health and behavior. OpenTelemetry defines common signals such as traces, metrics, and logs. Metrics are numeric measurements over time. Add safe business rules when a successful request does not prove a correct result.
Use the test pyramid as a guide
The test pyramid is a rule of thumb that favors many focused tests and fewer broad tests. It remains useful because browser and full-system tests are often slower and harder to understand.
It becomes too simple when a team requires fixed percentages of code-level, component, and user interface tests. Martin Fowler's description of the test pyramid states the main assumption: broad tests usually cost more than focused tests, but exceptions exist.
Some work does not fit a simple triangle:
- Front-end components can behave like the real product without using the full system.
- Data contracts, database changes, and messages cross test layers.
- Performance, resilience, accessibility, and security need specialized methods.
- Older systems may first allow only broad tests.
- You can verify the production configuration only after deployment.
Selenium also discourages using WebDriver for performance testing because browser and external variation make the results hard to interpret.
Use the pyramid to spot a possible problem. Use the risk worksheet to decide what to change.
What I would do
I would teach the framework last.
- Draw the architecture of one critical user journey.
- Name a failure that matters to a user or the organization.
- Find the smallest part of the system that can reveal that failure.
- Choose the fastest reliable test that behaves enough like the real system.
- Add one broader check for facts the focused test cannot prove.
- Run each check at the right stage of CI or deployment.
- Assign an owner and save useful failure details.
- Link the risk to signals from the running system.
- Learn the tool needed for that test.
An engineer who understands these choices can learn Selenium when browser automation is the right answer. An engineer who knows only the framework may automate the wrong part of the system.