Using Claude Code CLI for Flaky Test Triage

Every test suite past a certain size accumulates flaky tests, ones that fail occasionally for reasons that have nothing to do with the code under test. A slow CI runner, a race condition in test setup, a shared resource another job happened to be using at the same time. The problem is never that flaky tests exist. The problem is triage time. Someone has to look at a failure, decide whether it is real or noise, and that decision eats far more time across a team than it should, especially on a suite with hundreds of tests running on every merge. ...

April 22, 2026

Load, Stress, Soak, and Spike Testing: Choosing the Right Injection Strategy

In the previous post, we made scenarios behave more like real users. With everything we have built so far, injection profiles, feeders, correlation, checks, and pacing, we finally have enough pieces to talk about the different types of performance tests properly, since “run a load test” actually covers several genuinely different testing goals. Load Testing: Expected Traffic A load test answers a simple question. Does the system perform acceptably under the traffic level you actually expect. This is the baseline test you should have running regularly, ideally on every significant release. ...

April 16, 2026

OpenSpec - Spec driven development for brownfield projects

Spec-Driven Development for the Rest of Us: OpenSpec Last month I found GitHub’s Spec Kit and how it flips the traditional dev workflow — spec as the durable source of truth, code as the disposable output. Spec Kit’s /constitution-first, plan-then-build ceremony assumes you’re starting mostly from a blank slate. Almost nothing I touch day-to-day looks like that. It’s always years old services, inherited conventions nobody remembers agreeing to, and test suites that are more archaeology than architecture. ...

April 4, 2026

Modeling Realistic User Journeys: Pacing, Think Time, and Scenario Design

In the previous post, we made sure our checks actually catch real failures. Today we look at something just as important but easier to overlook, whether your scenario actually behaves like a real user in the first place. A scenario that fires request after request with zero delay between them does not represent any real visitor to your site. It represents a script racing through steps as fast as the network allows. That produces load numbers, but not necessarily useful ones, since real traffic has gaps in it while people actually read a page, think about what to click next, or get distracted by something else entirely. ...

April 2, 2026

Playwright vs Selenium: Test Isolation and Parallelization

In the previous post, we looked at controlling network traffic during a test. This is the last post in this short series comparing Playwright and Selenium, and it covers what happens once your suite grows from ten tests to a thousand and you need them to run fast without stepping on each other. Isolation in Selenium Is Something You Build A single WebDriver instance is a single browser session, with its own cookies and storage. If two tests reuse that same session one after another, state can leak between them without you noticing. ...

March 24, 2026

Checks and Validations: Making Sure Your Load Test Catches Real Failures

In the previous post, we used checks to extract tokens for correlation. This time we look at checks purely from a validation angle, because a surprising number of load tests quietly pass while the application under test is actually broken. The Trap of Checking Status Codes Only Here is a scenario I have seen play out more than once. A test checks only that every response comes back with a 200 status. Midway through the run, the application starts returning a generic error page, but that error page itself happens to render with a 200 status code, because the server is misconfigured to return success even for its own error pages. The load test finishes green. The report shows a great response time. Meanwhile every single user in that window got an error page instead of what they asked for. ...

March 19, 2026

Playwright vs Selenium: Network Interception and API Mocking

In the previous post, we compared how each tool handles setup and teardown. Today we look at something that trips up a lot of teams moving from Selenium to Playwright for the first time. Controlling what actually happens on the network during a test. Why Selenium Was Never Built for This It helps to understand where Selenium comes from. The WebDriver protocol, the actual W3C specification Selenium implements, was designed to simulate a real user driving a real browser. Click here, type there, read what is on the page. Network traffic was never part of that picture. ...

March 13, 2026

Correlation and Authentication: Extracting Tokens and Handling Login Flows

In the previous post, we fed real data into our scenarios. Today we cover something almost every real application needs before you can test anything interesting behind a login screen. Correlation. Correlation just means grabbing a value out of one response and reusing it in a later request. The most common example by far is authentication. You log in once, get back a token, and then attach that token to every request that follows. ...

March 5, 2026

Playwright vs Selenium: Fixtures vs Manual Setup and Teardown

Most comparisons between Playwright and Selenium start with API syntax. Click this way, find an element that way. That is not where the real difference shows up day to day. The real difference shows up the moment you write your tenth test and notice how much setup code you are copying between files. This is the first of three posts comparing Playwright and Selenium on things that actually matter once a suite grows past a handful of tests. Today it is setup and teardown. The examples use Java for Selenium and TypeScript for Playwright, since that is a pairing a lot of teams genuinely work with side by side. ...

March 3, 2026

Spec Kit- Spec Driven development

Flipping the Script: A Look at GitHub’s Spec Kit I’ve spent the last few posts obsessing over agents that write, run, and heal tests. This time I want to zoom out one level, because I think I’ve been skipping over the artifact that actually matters most to an AI coding agent: the spec itself. That’s what pulled me into Spec Kit, GitHub’s open-source toolkit for what they’re calling “Spec-Driven Development” (SDD). The pitch is deceptively simple but genuinely inverts how most of us have worked for the last twenty years. In traditional development, the spec (if one even exists past the kickoff meeting) is a scaffold — you lean on it briefly, then throw it away the moment code starts shipping. Code becomes the source of truth, and the spec rots in Confluence somewhere, quietly lying to whoever reads it next. ...

March 2, 2026