<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://www.ontestautomation.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://www.ontestautomation.com/" rel="alternate" type="text/html" /><updated>2026-09-26T17:58:17+00:00</updated><id>https://www.ontestautomation.com/feed.xml</id><title type="html">On Test Automation</title><subtitle>I provide test automation training and consultancy, teaching you how to write tests that deliver valuable product feedback, fast</subtitle><author><name>Bas Dijkstra</name></author><entry><title type="html">An update on public speaking</title><link href="https://www.ontestautomation.com/an-update-on-public-speaking/" rel="alternate" type="text/html" title="An update on public speaking" /><published>2026-08-22T00:00:00+00:00</published><updated>2026-08-22T00:00:00+00:00</updated><id>https://www.ontestautomation.com/an-update-on-public-speaking</id><content type="html" xml:base="https://www.ontestautomation.com/an-update-on-public-speaking/"><![CDATA[<p>Earlier this year, I wrote a blog post <a href="/on-increasing-focus-in-my-career/">about increasing focus in my career</a>. If you haven’t read and/or don’t feel like reading that post, the tl;dr of it is that for a couple of years, I’ve been working on too many things at the same time, which led to</p>

<ul>
  <li>not delivering all of my work at the quality that I want to and can deliver, and</li>
  <li>not having enough time for things outside of work, most notably cycling</li>
</ul>

<p>It is now five months since I wrote that post, and things have changed slightly for the better. I feel like I am working on fewer things in parallel, and, as I expressed, most of my time these days is spent on building <a href="/training/">my training business</a> and on reading, writing for this blog and especially for <a href="/newsletter/">my newsletter</a> and studying.</p>

<p>I still do a bit of consulting, but my current engagement is ending soon, and after that, I’m aiming for the consulting that I will do to be <a href="/solutions/">solutions-based</a>, instead of <a href="/on-ditching-hourly-and-productizing-my-services/">based on hourly billing</a>.</p>

<p>One activity on which I expected to spend less time is public speaking. As I said in the blog post:</p>

<blockquote>
  <p>The public speaking will stay as well, but here too, I will consider and take what comes my way, instead of actively pursuing speaking gigs at meetups and conferences, especially those abroad. I do have a few in-company speaking gigs coming up in the next couple of months, as well as a keynote at an online conference. I enjoy speaking, so I will continue doing talks, but again, pretty much only when I am invited to do so.</p>
</blockquote>

<p>I’m not actively pursuing speaking opportunities anymore, which means that it has been quite some time since I submitted a proposal to a Call for Papers. I have made and will keep making an exception and submit proposals for a very short list of conferences, namely</p>

<ul>
  <li><a href="https://conference.eurostarsoftwaretesting.com/" target="_blank">EuroSTAR</a> and <a href="https://automation.eurostarsoftwaretesting.com/" target="_blank">AutomationSTAR</a> - because they’re the largest and therefore provide me with a podium to showcase what I do in front of a large audience, and because I really enjoy their events</li>
  <li><a href="https://kwsqa.org/" target="_blank">KWSQA’s</a> Targeting Quality - because I really like their community and because I’ll take every excuse to travel to Canada (I hope to be back in Cambridge, ON, soon)</li>
</ul>

<p>Even then, my submissions will probably be workshops rather than talks, because workshops are what I do for a living anyway in my training business, and the compensation will often be much better, too.</p>

<p>However, even now that I’m really no longer actively seeking out speaking opportunities, those opportunities still present themselves on a regular basis. I might not do as many talks in 2026 as I did in 2025 (so far, I did 6 this year, where I did 21 in total in 2025), but still, every now and then people and organizations reach out to me to have me as a speaker at their event. And not just online, in person, too.</p>

<p>Now, this might sound strange to some of you, but I’ve never really seen myself as a typical conference speaker. I enjoy doing talks, sure, but I’m not the ‘conference tiger’ taking their talk from one conference to the next as some of my fellow folks in the industry do (and all the best to them!). So, every time people do reach out and ask if I can do a talk at their event, I’m honoured, and often a little surprised, too. Even more so when ask me if I want to do a keynote.</p>

<p>I’ve done a few keynotes so far (more than 5 but less than 10, I lost count), from my first one ever at the <a href="https://testdag.nl/" target="_blank">Dutch Testing Day</a> in 2018, my first one abroad at UKSTAR in 2019, to the most recent one at the in-company developer conference at a large Dutch bank. And guess what? I actually enjoy doing keynotes. I never really thought I would, but I do.</p>

<p>That wasn’t always the case, by the way. In the past, after my first few keynotes, every time I thought ‘well that surely was the last one ever’. I’d seen so many other people deliver brilliant, inspiring, clever, moving keynotes, I thought I could never do that. Still don’t think I can. And that’s fine with me.</p>

<p>The talks I do, and enjoy doing, are about how we can and should do things better, mostly in test automation. Often, that will include some code or even a live demo. My talks are distilled from 20 years of experience in the field, and all the lessons I learned along the way. You might not get life-changing insights, but you’ll get practical advice, wrapped in ugly retro slides, with some references to 90s pop culture and bad jokes thrown in.</p>

<p>Turns out, there are events that are happy to have such a talk as a keynote. Next to the in-company developer conference I mentioned, I’m also doing a keynote at the online day of <a href="https://www.testear.la/en" target="_blank">Testear.la</a> and at this year’s <a href="https://cypress.registration.goldcast.io/events/670deb6c-06ee-4ce8-858b-8a4db3a62eb1" target="_blank">CypressConf</a>. I’ve even got my first keynote for 2027 lined up already, and I’m really looking forward to speaking at the inaugural <a href="https://ttc.taqelah.sg/" target="_blank">Taqelah Testing Conference</a> in Singapore in March.</p>

<p>Because I don’t actively pursue speaking opportunities, but will happily accept what comes my way, I have more time to create new, original, customized talks for these events, too. Two out of the three keynotes I do this year are developed specifically for the event. Come to think of it, that’s exactly what I do in my training business as well: no off-the-shelf standardized courses and workshops, but tailored engagements pretty much every time.</p>

<p>And that’s exactly how I like it best.</p>

<p><em>If your event is looking for a keynote speaker bringing in a customized, brand new talk with practical advice and lessons learned in 20 years of test automation, and if you think your audience can deal with my ugly retro slides and lame jokes, <a href="/contact/">let’s talk</a>.</em></p>]]></content><author><name>Bas Dijkstra</name></author><category term="Public speaking" /><category term="career" /><category term="skills" /><summary type="html"><![CDATA[Earlier this year, I wrote a blog post about increasing focus in my career. If you haven’t read and/or don’t feel like reading that post, the tl;dr of it is that for a couple of years, I’ve been working on too many things at the same time, which led to]]></summary></entry><entry><title type="html">Should testers write unit tests?</title><link href="https://www.ontestautomation.com/should-testers-write-unit-tests/" rel="alternate" type="text/html" title="Should testers write unit tests?" /><published>2026-08-20T00:00:00+00:00</published><updated>2026-08-20T00:00:00+00:00</updated><id>https://www.ontestautomation.com/should-testers-write-unit-tests</id><content type="html" xml:base="https://www.ontestautomation.com/should-testers-write-unit-tests/"><![CDATA[<p><em>This post was previously published through my <a href="/newsletter/">newsletter</a> on May 25, 2026. From time to time, I will republish newsletter issues on my blog here if I think people (and search engines) might benefit from it. If you want to read everything I’ve written once it is posted, I recommend signing up for the newsletter.</em></p>

<p>A few months ago, I had the pleasure of delivering a keynote at an internal developer conference for one of the largest banks in the Netherlands. While I still don’t really see myself as a keynote speaker - I enjoy doing practical, hands-on sessions much more - I had a great time talking about challenges of E2E testing, breaking down E2E tests and <a href="/the-test-automation-quadrant/">the test automation quadrant model</a> that I use in my thinking, speaking and teaching about test automation these days.</p>

<p><img src="/images/blog/bas_keynote.jpg" alt="bas_keynote" title="Bas on stage at the ABN AMRO developer conference in May 2026" /></p>

<p>However, the keynote or the contents of it are not what I want to talk about this week. Instead, it was a question that came up during the Q&amp;A after the talk that triggered me to write this post. That question was</p>

<blockquote>
  <p>“Do you think that testers should be writing unit tests?”</p>
</blockquote>

<p>It’s not the first time I heard that question. In fact, I’ve seen and heard whether or not testers should be involved in unit testing being discussed regularly in the past, with arguments for and against the various standpoints. However, I was under the impression that we, collectively, had found some sort of answer to the question and moved past this point by now. Guess I was mistaken.</p>

<p>I tried to give the person asking the question an answer as well as I could, but given that there is quite a bit to unpack around the topic, I’m not sure if I gave them the entire story. So, that’s what I’ll try and do here. Who knows they might even read it…</p>

<p>So, should testers be writing unit tests? Well, my answer is either a ‘yes’ or a ‘no’, depending on how you interpret the question.</p>

<p>Before we explore these various interpretations, what is a unit test anyway? Well, by now, I don’t really know anymore, as there are so many definitions floating around, some of them contradicting each other. This is one of the reasons I came up with the test automation quadrant model as an alternative to the well-known automation pyramid model, but again, that’s not the topic of this post.</p>

<p>For the sake of the argument, let’s define a unit test as a test that verifies a very small piece of behaviour of our product, without relying on external interfaces like APIs, databases or file systems, for example.</p>

<p>With that definition in place, let’s look at two different interpretations of the question of ‘should testers be writing unit tests?’.</p>

<h3 id="interpretation-1-testers-should-be-responsible-for-unit-testing">Interpretation 1: Testers should be responsible for unit testing</h3>
<p>Well, no, I don’t think they should, no matter how good of a tester they are, and no matter what their coding skills are. Writing unit tests is an activity that should be performed in support of and lockstep with software development.</p>

<p>Often, especially in practices like test-driven development, tests are written first, and they drive the design and development of the product. Leaving the writing of these tests to testers, especially if it is done after the product code itself is written, is both inefficient and a potential source of problems.</p>

<p>Inefficient, because the product has already been written, yet we only know whether its behaviour matches expectations once the tests are written and run. Also, because there’s a handoff happening between ‘development’ and ‘testing’, which takes up valuable time as the developer will switch to a different task while the tester writes the unit tests.</p>

<p>If there’s a problem with the product that is discovered during unit testing, the developer needs to make a context switch back to the original task, and the more context switching you do during the day, the less efficiently you will work and the less you will get done.</p>

<p>A potential source of problems, because when the developer only focuses on shipping a potentially working product, they will likely not spend too much time thinking about what to test for, or how to make the product (the code, in this case) easily testable in the first place.</p>

<p>Also, the handoff from ‘development’ to ‘test’ I described before leaves open room for different interpretations of what the software should do, leading to potential bugs slipping through and the resulting back-and-forth discussions after the fact.</p>

<p>So, no, I don’t think we should leave unit testing to testers. I still sometimes hear about developer who don’t want to or do not know how to write (decent) unit tests, and I think that’s a problem that needs to be fixed at the source, instead of trying to patch it up by someone else writing the unit tests for the already-created product.</p>

<h3 id="interpretation-2-testers-should-be-involved-in-unit-testing">Interpretation 2: testers should be involved in unit testing</h3>
<p>If we interpret the question this way, I think the answer should be a resounding ‘yes’, testers should be involved in unit testing. I mean, there’s the word ‘testing’ in the name, why shouldn’t we involve the people who specialize in testing in the process?</p>

<p>What that involvement looks like, exactly, depends on the context, of course. Again, ‘being involved’ and ‘being responsible’ are two entirely different things. While I believe that writing unit tests is a development activity, there are a couple of ways in which testers can add value, too:</p>

<ul>
  <li>They can review the unit tests to learn about what has been covered already, so that they do not repeat that testing later on</li>
  <li>They can review the unit tests to identify what has not yet been covered, and either give that back as feedback to a developer, or add the missing tests themselves</li>
  <li>They can suggest and use techniques like mutation testing to test the tests and find out whether the tests that were written before are actually able to catch meaningful problems</li>
</ul>

<p>I’m sure there are a few more benefits, but these alone should, in my opinion, be enough for any team to not exclude testers from the process of writing unit tests from here on. And yes, that will require some additional skills from both testers and developers.</p>

<p>This is where the power of collaboration comes in. I’m a big fan of pair programming and testing, and I’ve seen a lot of good things come from testers and developers pairing up to write, review and improve unit tests. The developer improves their testing skills, the tester learns more about both the product they’re testing and the development process, and in the end, both the entire team and the product itself reaps the benefits.</p>

<p>That’s a long answer to a simple question, and I’m sure there are some nuances that I didn’t yet unpack, but generally speaking, these are my views on the question of whether testers should be writing unit tests.</p>

<p>Oh, and before you ask ‘but what about integration / end-to-end / performance / security / … tests’? That’s easy. Simply replace ‘unit tests’ and ‘unit testing’ with ‘integration / end-to-end / performance / security / … tests’ and ‘… testing’, and you’ll have my answer.</p>

<p>I’ve always found it a little strange that unit tests have traditionally been seen as ‘different’ from other types of tests, as if they’re some kind of special artefact that only developers know how to write. If that’s what you think, too, let me let you in on a little secret: they aren’t special.</p>

<p>Unit tests are simply tests that test a small piece of the behaviour of the product that we write, written against and invoking a specific interface of the product: the source code. Other types of tests do exactly the same thing: verify parts of our product behaviour by invoking one or more specific interfaces (APIs, the UI, a database, a queue, …). They just have a different scope, and with that, they verify behaviour at a different scope. That’s all.</p>

<p>This, again, is the reason where I think the traditional test automation pyramid model lacks: I really don’t care that much about what is a unit / integration / end-to-end test. All I really care about is <a href="/training/valuable-feedback-fast/">tests that produce valuable information about the state and the behaviour of our product in as efficient a way as possible</a>.</p>

<p>Unit tests really aren’t any different.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Test automation" /><category term="unit tests" /><category term="skills" /><summary type="html"><![CDATA[This post was previously published through my newsletter on May 25, 2026. From time to time, I will republish newsletter issues on my blog here if I think people (and search engines) might benefit from it. If you want to read everything I’ve written once it is posted, I recommend signing up for the newsletter.]]></summary></entry><entry><title type="html">I’m (re-)starting a newsletter</title><link href="https://www.ontestautomation.com/im-restarting-a-newsletter/" rel="alternate" type="text/html" title="I’m (re-)starting a newsletter" /><published>2026-05-11T00:00:00+00:00</published><updated>2026-05-11T00:00:00+00:00</updated><id>https://www.ontestautomation.com/im-restarting-a-newsletter</id><content type="html" xml:base="https://www.ontestautomation.com/im-restarting-a-newsletter/"><![CDATA[<p>Just a quick update to let those of you who bookmarked this blog or who have subscribed to my <a href="/feed.xml">RSS feed</a> know that I have (re-)started a newsletter.</p>

<h3 id="why-a-newsletter">Why a newsletter?</h3>
<p>As you might know (or not), while I’ve been pretty active on <a href="https://www.linkedin.com/in/basdijkstra">LinkedIn</a> over the years, I do have a love-hate (or rather an appreciate-hate) relationship with that platform.</p>

<p>Lately, I’ve been noticing that the pendulum is swinging in the ‘hate’ direction more often, mainly because the ever-changing algorithm used by LinkedIn makes it incredibly hard to predict if people are even going to see what I write.</p>

<p>I’d rather publish my thoughts, ideas and other ramblings via a platform that I do control, and that platform will be a newsletter.</p>

<p>I’ve had a newsletter in the past, but that only lived for about three months. This time, I intend to keep writing and publishing a new issue every week. The first edition goes out a few hours after I’m writing this, and a new issue will be sent to subscribers every Monday morning around 11 AM CET.</p>

<h3 id="but-what-about-the-blog">But what about the blog?</h3>
<p>I’ll still publish to the blog, too, but that will be on a much less regular basis. Just like it has been for a while, really. The idea is to post the more ‘technical’ posts, that is, the ones including code, directly to my blog, whereas the ‘text-and-images-only’ posts go through my blog post first.</p>

<p>My priority is with the newsletter, though.</p>

<h3 id="how-to-subscribe">How to subscribe</h3>
<p>That’s easy, just go to the <a href="/newsletter/">subscription page</a>, leave your email address, click the button on the confirmation email and you’re in.</p>

<p>I promise I won’t use the newsletter or your email to spam or sell to you. Ever.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="General" /><category term="newsletter" /><summary type="html"><![CDATA[Just a quick update to let those of you who bookmarked this blog or who have subscribed to my RSS feed know that I have (re-)started a newsletter.]]></summary></entry><entry><title type="html">My thoughts on ‘self-healing’ in test automation</title><link href="https://www.ontestautomation.com/my-thoughts-on-self-healing-in-test-automation/" rel="alternate" type="text/html" title="My thoughts on ‘self-healing’ in test automation" /><published>2026-04-09T00:00:00+00:00</published><updated>2026-04-09T00:00:00+00:00</updated><id>https://www.ontestautomation.com/my-thoughts-on-self-healing-in-test-automation</id><content type="html" xml:base="https://www.ontestautomation.com/my-thoughts-on-self-healing-in-test-automation/"><![CDATA[<p>From what I’ve seen over the years, tests that exercise a system through the graphical user interface are disproportionally likely to fail for reasons other than an actual, genuine product failure. Call them false positives, call them flaky tests, call them whatever you want.</p>

<p>There are good reasons for this: graphical user interfaces are interfaces optimized for consumption by humans, not by machines or code. There is a lot of (often asynchronous) processing going on in your browser, all to create an application that looks good, is pleasant to work with, and generally provides a good end user experience. And remember, that end user is a human being, not a computer.</p>

<h2 id="the-problem">The problem</h2>
<p>The dynamic nature of a modern GUI often leads to test failures, even when the behaviour of the application remains unchanged. For example, if the text on a button that submits a form for a loan application used to be <code class="language-plaintext highlighter-rouge">Submit</code>, but changes to <code class="language-plaintext highlighter-rouge">Apply</code>, and you have Playwright tests that first locate and then click the button using</p>

<figure class="highlight"><pre><code class="language-typescript" data-lang="typescript"><span class="k">await</span> <span class="nx">page</span><span class="p">.</span><span class="nx">getByRole</span><span class="p">(</span><span class="dl">'</span><span class="s1">button</span><span class="dl">'</span><span class="p">,</span> <span class="p">{</span> <span class="na">name</span><span class="p">:</span> <span class="dl">'</span><span class="s1">Submit</span><span class="dl">'</span> <span class="p">}).</span><span class="nx">click</span><span class="p">();</span></code></pre></figure>

<p>your test is going to fail to locate the button, even when you’re using <code class="language-plaintext highlighter-rouge">getByRole()</code>, a Playwright-recommended locator.</p>

<p>Another example: let’s say you check that the table of accounts in an online banking platform is visible by using its HTML <code class="language-plaintext highlighter-rouge">id</code> attribute, like this:</p>

<figure class="highlight"><pre><code class="language-typescript" data-lang="typescript"><span class="k">await</span> <span class="nx">expect</span><span class="p">(</span><span class="nx">page</span><span class="p">.</span><span class="nx">locator</span><span class="p">(</span><span class="dl">'</span><span class="s1">#accounts</span><span class="dl">'</span><span class="p">)).</span><span class="nx">toBeVisible</span><span class="p">();</span></code></pre></figure>

<p>When the value of this <code class="language-plaintext highlighter-rouge">id</code> attribute changes, for example because of a new version of the framework used to build the frontend is used, your test is going to fail to locate the table, and subsequently tell you the table is not visible, even when it is.</p>

<h2 id="self-healing-algorithms---the-solution">Self-healing algorithms - the solution?</h2>
<p>To avoid unwanted rework, many tools these days offer a ‘self-healing’ feature: whenever a test fails to locate an element, it will try and identify the element that you <em>intended</em> to locate, often using a probabilistic algorithm that is sold to you as ‘AI’.</p>

<p>What this means is that it consults a large collection of training data, finds similar occurrences of changes in UI layout, either from your own history of changes or from other applications, checks those against what it sees on the screen, and from there it selects the candidate element that is most likely to be the one that you intended to locate. These tools will also often report on the results of the changes they made in a test run, so that a human being can review these changes afterwards and approve or reject them.</p>

<p>The bigger the size of the training database, the higher the quality of the training data, and the higher the sophistication of the algorithm, the higher the probability that the tool identifies the correct element, e.g., the button with the <code class="language-plaintext highlighter-rouge">Apply</code> text that replaced the <code class="language-plaintext highlighter-rouge">Submit</code> text. Still, trying to find the ‘right’ element <em>is</em> a probabilistic process, which means that mistakes will be made. That’s fine, as long as you’re aware of the risk, and you’re willing to accept it.</p>

<h2 id="why-i-think-self-healing-is-a-bad-idea">Why I think self-healing is a bad idea</h2>
<p>So, would I recommend using self-healing frameworks as a solution to the challenge of often-changing graphical user interfaces and the test code maintenance that is the result thereof?</p>

<p>Well, no.</p>

<p>Sure, using these tools might reduce the time required to maintain your tests in the case of changing HTML and the need to update the corresponding element locators. And I don’t even have a problem with the fact that the algorithms they use are probabilistic, which means they can make mistakes. I mean, there’s a lot of fairly useless and downright bad human-written test code out there, too, so I don’t think that problem is unique to LLM-powered test tools or LLM-generated test code.</p>

<p>No, the biggest problem I have with ‘self-healing’ test tools and frameworks is that they are, to me, nothing more than band-aids, hiding the real problem that is underneath.</p>

<p>That real problem, to me, is the fact that the people writing the tests weren’t aware that the locators changed in the first place. Why did the text on the button change from <code class="language-plaintext highlighter-rouge">Submit</code> to <code class="language-plaintext highlighter-rouge">Apply</code>? Who committed that change in the product code? And why didn’t they either change the corresponding test code, too, or informed someone in their team that the test might break because of this? Or, in case of the second example, why didn’t we know that the framework update might lead to changing <code class="language-plaintext highlighter-rouge">id</code> values? Or, as is often the case, that the update happened in the first place?</p>

<p><em>That’s</em> the problem that we need to address. We need to close the communication and collaboration gaps in our teams, instead of trying to patch them up with an algorithm. Test automation isn’t the deliberate chase of green checkmarks, it’s using tools to efficiently detect changes in the behaviour of our product. Self-healing, to me, feels like an attempt to <a href="/stop-sweeping-your-failing-tests-under-the-rug/">sweep these changes under the RUG</a>.</p>

<p>I don’t know about you, but I’d rather address the real problem than applying a band-aid and hoping the problem will remain out of sight.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Test automation" /><category term="self-healing" /><category term="graphical user interfaces" /><summary type="html"><![CDATA[From what I’ve seen over the years, tests that exercise a system through the graphical user interface are disproportionally likely to fail for reasons other than an actual, genuine product failure. Call them false positives, call them flaky tests, call them whatever you want.]]></summary></entry><entry><title type="html">The ‘valuable’ in valuable feedback, fast</title><link href="https://www.ontestautomation.com/the-valuable-in-valuable-feedback-fast/" rel="alternate" type="text/html" title="The ‘valuable’ in valuable feedback, fast" /><published>2026-04-01T00:00:00+00:00</published><updated>2026-04-01T00:00:00+00:00</updated><id>https://www.ontestautomation.com/the-valuable-in-valuable-feedback-fast</id><content type="html" xml:base="https://www.ontestautomation.com/the-valuable-in-valuable-feedback-fast/"><![CDATA[<p>When I talk about the goals and the purpose of test automation, I often use the phrase ‘<a href="/training/valuable-feedback-fast/">valuable feedback, fast</a>’: we use tools to support our testing to help us get valuable information about the state of our product in the most efficient manner possible.</p>

<p>The ‘fast’ part of ‘valuable feedback, fast’ is pretty self-explanatory for most people: as build and release cycles are becoming shorter, teams want to be informed timely about any unexpected changes in behaviour of their product, often after every change they make to that product.</p>

<p>Tools can help them achieve that by running <a href="/breaking-down-your-e2e-tests-an-example/">quick, focused tests</a> automatically when a change is made or committed to version control. Of course, it takes plenty of hard work to write those tests to be fast, but that’s not what I wanted to talk about here.</p>

<p>The ‘valuable’ in ‘valuable feedback, fast’ is a much more ambiguous term, and one that deserves some more explanation. To me, there are multiple dimensions to what makes a test valuable, and in this post, I want to unpack and address them one by one.</p>

<h2 id="valuable--important-to-someone-who-matters">Valuable = important to someone who matters</h2>
<p>Borrowing from the classic definition of ‘quality’ as defined by Jerry Weinberg and further refined by James Bach and Michael Bolton, this is where it all starts. The information presented by a test should be important to someone who matters. That someone could be a member of the development team, a stakeholder such as a product owner or business analyst, the end user of the product, or a combination of those.</p>

<p>Without that importance, a test is meaningless, dead weight. It could be the most reliable, best-written test ever, but if the information that is provided by it is not important to someone who matters in the context of the product, why bother writing, running and maintaining the test?</p>

<h2 id="valuable--covering-what-matters">Valuable = covering what matters</h2>
<p>Test coverage is a tricky subject, and I want to steer clear of the discussion on what ‘coverage’ means exactly in this blog post. The only realistic answer is ‘it depends’, anyway, as there are so many ways to define coverage (line, branch, requirements, mutation, …).</p>

<p>Having said that, for the information provided by our tests to be valuable, teams should invest time in making sure that the tests cover the parts of the product behaviour that are deemed ‘important enough’ in a sufficient manner. What exactly constitutes ‘sufficient’ here depends on, you guessed it, the context.</p>

<p>Some products require deeper, more thorough coverage than others. The same applies to individual parts of the same product. It all depends on the acceptable amount of risk a team is willing to take before putting a product in the hands of their users. Teams would do well to have a continual discussion about these risks and the extent to which they are covered by the tests that accompany and scrutinize the product.</p>

<h2 id="valuable--trustworthy">Valuable = trustworthy</h2>
<p>The higher the degree of automation in the build and delivery process of a product, and that includes testing, the more teams will rely (and have to rely) on the results of the execution of that automation. Concerning tests, that means that teams need to be able to rely on the information presented by the tests, because they will make decisions based on that information.</p>

<p>The nature of that decision might vary from anywhere between ‘this build seems sufficiently stable to warrant deeper testing’ to ‘this change is ready to be put in the hands of our users’. No matter what the specific decision is, if teams make it based on the results of your test automation, even in part, they can only confidently do so if the information provided by the tests is trustworthy.</p>

<p>In practice, that means that when a test emits a signal indicating a problem with the product, the team can safely conclude that there is a problem <em>with the product</em>, not with the test, the data it uses or the environment it runs in (no false positives). It also means that when a test does not emit such a signal, the team can trust that the particular piece of behaviour exercised by the test is working according to expectations expressed in the test (no false negatives).</p>

<h2 id="valuable--actionable">Valuable = actionable</h2>
<p>Another dimension of the value of the feedback provided by a test is that it should be actionable. This applies specifically to those situations where a test ‘fails’, i.e., it indicates a problem with the product I have put ‘fails’ between quotes here, because the test didn’t fail, the product failed the test. There’s a difference.</p>

<p>Anyway, when a test result indicates a (potential) problem with the product, teams need to able to act on that information as soon as possible, spending as little time digging deeper into the product or into the test as possible to identify the root cause of the problem. Some practices that might help here are:</p>

<ul>
  <li>Making your test scope as small as possible - the fewer moving parts your test has, the easier it will be to identify which of those parts made a move that was unexpected</li>
  <li>Have good test names - A descriptive test name that tells you what part of the behaviour your product verifies and what the expected behaviour is helps in finding out where exactly the problem might be found</li>
  <li>Use custom assertion messages - Many test frameworks allow you to specify custom, descriptive error messages in case of assertion failures (something RestAssured.Net <a href="https://github.com/basdijkstra/rest-assured-net/wiki/Usage-Guide#custom-error-messages" target="_blank">supports as of version 5.0.0</a>, too)</li>
</ul>

<p>So, is this a complete and final definition of what ‘valuable’ means to me when I talk about ‘valuable feedback, fast’ as the goal of test automation? I don’t think so. I don’t know if it is complete, but it definitely is a good reflection of my current thoughts on ‘value’ in test automation right now. Those thoughts are definitely not ‘final’, and I would appreciate your takes on what I wrote here.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Valuable feedback fast" /><category term="Test automation" /><category term="Training" /><summary type="html"><![CDATA[When I talk about the goals and the purpose of test automation, I often use the phrase ‘valuable feedback, fast’: we use tools to support our testing to help us get valuable information about the state of our product in the most efficient manner possible.]]></summary></entry><entry><title type="html">On increasing focus in my career</title><link href="https://www.ontestautomation.com/on-increasing-focus-in-my-career/" rel="alternate" type="text/html" title="On increasing focus in my career" /><published>2026-03-13T00:00:00+00:00</published><updated>2026-03-13T00:00:00+00:00</updated><id>https://www.ontestautomation.com/on-increasing-focus-in-my-career</id><content type="html" xml:base="https://www.ontestautomation.com/on-increasing-focus-in-my-career/"><![CDATA[<p><em>Note: this is not the most well-structured blog post I’ve written, but rather a recording of my not (yet) completely organized thoughts.</em></p>

<p>Last week, I was at the <a href="https://www.testautomationdays.com" target="_blank">Test Automation Days</a> conference, both as a member of the program committee, and as a last-minute substitute workshop facilitator. As is typically the case with conferences, I had a lot of conversations, catching up with people I’ve met before, as well as getting to know people I haven’t. One of the questions that invariably comes up in these conversations is</p>

<blockquote>
  <p>“So, what are you up to these days?”</p>
</blockquote>

<p>I typically start my answer with a lame attempt at humour by saying “well, do you have half an hour?”. Poor jokes aside, it <em>is</em> true, though: I am working on a lot of different things these days. So much so that I’m sometimes feeling like I am spreading myself a little thin. For example, at the moment of writing this I am involved in</p>

<ul>
  <li>working part time with a <a href="/consulting/">consulting</a> client here in the Netherlands</li>
  <li>running <a href="/training/">workshops and training courses</a> for several clients</li>
  <li>in the final stages of a <a href="/mentoring/">corporate mentoring</a> engagement</li>
  <li>wrapping up the creation of a video course for a new platform</li>
  <li>talking about and preparing for several public speaking gigs</li>
  <li>reviewing submissions for several conferences</li>
</ul>

<p>And then there’s the writing, the reading, the watching videos, the conversations to help out fellow testers, and more. Now, I enjoy doing all of these things (well, except for editing videos of me talking, that is painful), but it is <em>a lot</em>. In fact, I think I would be better off doing fewer things, going deeper on those, spending more time on fewer things and ultimately, raising the quality of the work I deliver.</p>

<p>Which means that something, or rather several things, will have to go. First on the list: creating video courses. I’m absolutely going to finish the one I’m working on now, but that will be the last one I do. Creating video content takes a lot of time, and in all honesty, the return on investment isn’t great. Plus, while I enjoy seeing the result go live, I could do without the process of writing transcripts and recording and then editing video. Did I tell you I despise editing video footage of me talking already?</p>

<p>As of today, my main focus will be on further building my training business. Why training? Well, for several reasons. First of all, it’s the kind of work I like doing best. That in itself is reason enough, but there’s more… It also gives me the most flexibility, and that’s something I’ll really need in the near future, but more on that in a moment. Another reason for focusing on training is that it puts me in front of lots of different people, teams and companies, which is a great way to build my network. And finally, and as I’m running a business, not insignificantly, running training courses, especially in-company, also pays the best, by a long shot.</p>

<p>I will keep working on some of the other activities, but it looks like I’ll have to practice saying ‘no’ a little more often. For example, I enjoy reviewing conference submissions, but it takes time I could spend on other things, too.</p>

<p>I am going to continue my current consulting engagement, because it is interesting and there’s plenty of opportunity to run in-company workshops, but it’s not very likely I’ll take on additional consulting gigs in the near future.</p>

<p>If a corporate mentoring gig comes up, I’ll happily take it, but I am not going to go out of my way to find one.</p>

<p>The public speaking will stay as well, but here too, I will consider and take what comes my way, instead of actively pursuing speaking gigs at meetups and conferences, <a href="/on-working-and-contributing-to-conferences-abroad/">especially those abroad</a>. I do have a few in-company speaking gigs coming up in the next couple of months, as well as a keynote at an online conference. I enjoy speaking, so I will continue doing talks, but again, pretty much only when I am invited to do so.</p>

<p>Finally, I will keep spending some time every week reading and writing, reading because I need to and want to stay on top of what is happening, writing because it is both a great way to organize my thoughts and to share those same thoughts with the world. Depending on opportunities, public speaking will be part of this, too.</p>

<p>In short, the only part of my business and service offerings that I will actively build and grow is my training business. What I’m looking for is spending most of my working time preparing for or running training sessions, with a few hours a week spent on outreach, coordination, and other logistical matters that come with running a training business. Ideally, running these sessions will also give me the opportunity to visit different places in the world, something I am able to do already, but I would love to go to more places, especially outside Europe (most training work I do is within the EU right now).</p>

<p>If all goes well, and that’s what I’m slowly working towards and have been working towards for a while now, that will leave a decent amount of time during the week for things outside of work. Most notably, for cycling.</p>

<p>I have always enjoyed cycling, but until recently, I only really brought out my racing bike every now and then. A few months ago, that changed, and I am now working my way towards riding ever longer distances (I’ll never be a fast rider, and that’s fine). I’ll ride my first century next month. In May, I hope to complete the cycling version of one of the <a href="https://www.fietselfstedentocht.frl/en/" target="_blank">most famous Dutch sporting events of all time</a>. If that goes well, I have my eyes set on a much longer event in 2027.</p>

<p>The only problem with cycling, and training for long-distance cycling events in particular, is that it takes time, a LOT of time. I go out for a bike ride three times a week, with my shortest rides being around 2 hours. A typical weekend ride, right now, is 5-6 hours, and that will only grow as I get closer to the events I would like to complete.</p>

<p>I don’t want to spend all my hours outside of work on a bike, and neither does my family. So, I decided to slowly start changing the way I work and the way I spend my time to make room for spending hours on my bike. That, to me, is the ‘freedom’ I talk about when I tell others why I am an independent consultant.</p>

<p>Anyway, ramblings over. If you’re looking to bring in an experienced trainer to help your team grow their test automation skills, let’s talk. I’d be happy to discuss options.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Career" /><category term="Focus" /><category term="Training" /><summary type="html"><![CDATA[Note: this is not the most well-structured blog post I’ve written, but rather a recording of my not (yet) completely organized thoughts.]]></summary></entry><entry><title type="html">Writing tests with Claude Code - part 1 - initial results</title><link href="https://www.ontestautomation.com/writing-tests-with-claude-code-part-1-initial-results/" rel="alternate" type="text/html" title="Writing tests with Claude Code - part 1 - initial results" /><published>2026-03-09T00:00:00+00:00</published><updated>2026-03-09T00:00:00+00:00</updated><id>https://www.ontestautomation.com/writing-tests-with-claude-code-part-1-initial-results</id><content type="html" xml:base="https://www.ontestautomation.com/writing-tests-with-claude-code-part-1-initial-results/"><![CDATA[<p>In <a href="/refactoring-the-rest-assured-net-code-with-claude-code/">a recent post</a>, I wrote about how I used Claude Code to analyze the code for <a href="https://github.com/basdijkstra/rest-assured-net" target="_blank">RestAssured.Net</a> and then perform a refactoring action, using handwritten tests as the safety net. In that post, I wrote that I didn’t want Claude to touch the tests themselves, and why. I was still curious, though, to find out for myself what Claude was capable of in terms of writing tests, and to what extent the trust that more and more people put into the tests written by LLMs is warranted.</p>

<p>In this blog post, I’ll share with you some first steps in doing exactly that, and you’ll read about my thoughts and my thought process along the way. You’ll see how I create an initial suite of tests for a small Spring Boot-based API that I wrote for use in several of my <a href="/training/">workshops</a>, and how I assessed the results. In a follow-up blog post, I’ll show you how I improved the test suite based on my observations, again using Claude Code.</p>

<h3 id="the-starting-point">The starting point</h3>
<p>As a starting point, I created a new repository containing the code for the API I wanted Claude to write tests for. Obviously, I removed the existing tests, and I also removed the README and the GitHub Actions build pipeline definition, as I want Claude to write tests based only on the product code itself, without being primed by other artifacts in the codebase. I left in are the dependencies used to write and run the tests, in this case REST Assured and JUnit.</p>

<p>After installing and <a href="https://github.com/basdijkstra/writing-tests-with-claude-code/blob/main/CLAUDE.md" target="_blank">initializing Claude</a>, I asked it to generate tests using this prompt:</p>

<blockquote>
  <p>“Add acceptance tests for the endpoints exposed by the AccountController to this project. Cover all the logic in the AccountService class. Use REST Assured as the tool to interact with the API. Use JUnit 5 as the test runner. Both libraries are already part of the project, see the pom.xml. Assert status codes and relevant response body elements as part of the tests. Extract common request properties into a RequestSpecification.”</p>
</blockquote>

<p>After some deliberation, Claude added a new test file to the project, containing 23 tests, all of them passing. <a href="https://github.com/basdijkstra/writing-tests-with-claude-code/blob/main/src/test/java/com/ontestautomation/mutationbank/controllers/AccountControllerTest.java" target="_blank">You can see these tests here</a>. The code in this file is the raw output from the above prompt, I haven’t changed anything.</p>

<p>It took Claude only a minute or two to write these tests, which definitely is a lot faster than what I could have done myself. But how good are they, really?</p>

<h3 id="a-first-look-at-the-tests">A first look at the tests</h3>
<p>Let’s look at the coding style first. I’m seeing people argue that code quality is not really all that important anymore once AI will write most of our code, but I beg to differ, especially when it concerns our tests. Tests are documentation of the intended behaviour of our code, and I would say that being able to read that documentation as a human being, without too much effort, remains very important. So, is the generated code easy to read?</p>

<p>There’s a <code class="language-plaintext highlighter-rouge">@BeforeEach</code> hook creating the <code class="language-plaintext highlighter-rouge">RequestSpecification</code> (an object in REST Assured containing shared HTTP request properties). There’s a helper method to create a new account passing in the <code class="language-plaintext highlighter-rouge">AccountType</code> and a predefined balance. There’s the aforementioned 23 tests that, especially at first glance, seem to verify things that are valuable.</p>

<p>What Claude did not do, very likely because I didn’t explicitly ask for it, is adding an abstraction layer to make the code easier to read, <a href="/using-the-client-test-model-in-rest-assured-net/">similar to the approach described here</a>. We’ll see how Claude does in this area in the next blog post, as I want to stick to assessing the quality of the initial output from Claude in this one.</p>

<p>And I have to say, all in all, for a first try, I’m not unhappy with what I’m seeing. Yes, there’s room for improvement, but I have seen humans do far worse than this. Of course, this is only a small and simple API, but that hasn’t stopped people (including myself) in the past from writing ugly test code for it.</p>

<p>When we look at what the tests seem to cover, at first glance it looks like they hit all the endpoints defined in the API controller, and most, if not all paths in the business logic defined in the service layer.</p>

<p>I should note here that I was able to fairly quickly and confidently come to these conclusions only because:</p>

<ul>
  <li>I wrote the code for the API, so I have knowledge of the inner workings and the intent of the API, and</li>
  <li>I have plenty of experience writing tests for APIs and writing tests in REST Assured, so I’d like to think I know what ‘good’ looks like</li>
</ul>

<p>If you don’t have that prior knowledge and experience, it will be harder to draw meaningful conclusions from just looking at what Claude coughs up. And that comes with a significant risk: the risk of saying ‘looks good to me’ without actually understanding what you’re approving and what you’re putting your trust into, and then ending up with a safety net of tests full of holes.</p>

<h3 id="testing-the-generated-tests-with-mutation-testing">Testing the generated tests with mutation testing</h3>
<p>To further increase our understanding of the value of the tests that were generated for me, let’s see if these tests can fail. If they can’t, the fact that we have generated 23 passing tests in two minutes flat is nothing more than an example of productivity theater.</p>

<p>My preferred method of finding out whether a test can actually fail is to <a href="/on-getting-started-with-mutation-testing/">use a mutation testing tool</a>. In this case, because we’re working with Java code, I’ll use <a href="https://pitest.org" target="_blank">PITest</a> as my mutation testing tool of choice. I configured the tool to mutate all the code in the project and run all the tests, to get a complete overview of the quality of the test suite generated. Note that in a real life-sized project, you probably want to start by mutating only part of the code base and run part of the tests to get mutation testing feedback within a reasonable amount of time.</p>

<p>After about a minute, PITest reports back that the initial test suite achieves 95% line coverage. This looks impressive, but it doesn’t really tell me anything. The much more valuable metric here is the number of mutants that were killed by the test suite. PITest reports that this is 91%, which, again is pretty good. In absolute numbers, out of 55 mutants generated by PITest, 50 were detected by the initial test suite.</p>

<p>From this number, two follow-up questions came up in my head immediately:</p>

<ul>
  <li>Which mutants were missed by the tests, and what is the impact of that? What are we not covering yet?</li>
  <li>Could we have achieved the same amount of (line and mutation) coverage with fewer tests? In other words, do we have tests that are dead weight?</li>
</ul>

<h3 id="looking-at-the-surviving-mutants">Looking at the surviving mutants</h3>

<p>First, let’s have a look at the mutants that survived, i.e., changes in the API code that were not detected by any of the tests.</p>

<p><img src="/images/blog/claude_mt_http500.png" alt="claude_mt_http500" title="Mutation testing: surviving mutants for HTTP 500 path" /></p>

<p>To start, in the <code class="language-plaintext highlighter-rouge">CustomizedResponseEntityExceptionHandler</code>, the HTTP 500 path isn’t covered in any of the tests, and that causes a surviving mutant. By design, the API returns an HTTP 500 when an <code class="language-plaintext highlighter-rouge">Exception</code> occurs that isn’t a <code class="language-plaintext highlighter-rouge">ResourceNotFoundException</code> (returning an HTTP 404) or a <code class="language-plaintext highlighter-rouge">BadRequestException</code> (returning an HTTP 400). This looks like a useful path to cover in a test.</p>

<p><img src="/images/blog/claude_mt_http204.png" alt="claude_mt_http204" title="Mutation testing: surviving mutants for HTTP 204 path" /></p>

<p>Second, the API returns an HTTP 204 in response to a GET call to <code class="language-plaintext highlighter-rouge">/accounts</code> when there are no accounts in the database. That path isn’t covered in any of the tests. This, too, seems like a useful path to test, because it is part of the intentional behaviour of the API.</p>

<p><img src="/images/blog/claude_mt_boundaries.png" alt="claude_mt_boundaries" title="Mutation testing: surviving mutants for boundary values" /></p>

<p>Finally, the tests that were written do not properly cover some of the boundary conditions, both in the logic that implements the business rule of ‘you cannot overdraw on a savings account’ and in the interest calculation logic. Once more, I would like to have these situations covered by tests.</p>

<p>These results, to me, indicate that mutation testing is a useful method of assessing what is tested and what isn’t. It doesn’t matter if you wrote the tests yourself, or you had them write by an LLM.</p>

<p><em>Note: it definitely helped here that I can confidently and quickly perform this analysis of the signals produced by PITest, and of the quality of my tests, because I know that mutation testing as a technique exists, and because I know how it works.</em></p>

<p>More importantly, these results strengthen my belief that it is my moral obligation to closely watch the output of an LLM. I deeply value writing tests that test meaningful things and that are actually able to detect changes in product behaviour, and I don’t think that should change when I have a tool writing the tests for me. If all I cared about was having <em>some</em> tests to cover the API and declared, for example, 90% line coverage as ‘good enough’, I would be done by now.</p>

<p>However, I am not OK with this result yet, for the reasons I gave just now. In the next blog post, I want to return this feedback to Claude and see how well it does in updating the existing test suite based on my observations. I also want to see if I can add mutation testing to the test generation loop, and have Claude achieve better mutation coverage without my interfering, but that’s something for another time.</p>

<p>For now, I’ll conclude that when I ask Claude to generate tests in the way I have done, it produces pretty good results in terms of both line and mutation coverage, but that it missed certain key paths in my application code.</p>

<h3 id="identifying-dead-weight-in-our-test-suite">Identifying dead weight in our test suite</h3>
<p>Next, I want to find out if the test suite that was generated by Claude contains dead weight, that is, do we have any tests that do not uniquely contribute to either line or mutation coverage? To do so, I asked PITest to generate a report in XML format next to the HTML report, as (for some reason) only the XML report contains information about which specific test killed a specific mutant.</p>

<p>Performing this analysis required a bit of elbow grease, as I had to manually search the XML test report for occurrences of the test name for every test in the test suite. This, too, is probably a process that can be automated, but for now, I’m OK with doing this the manual way. I don’t want to run the risk of hallucinations here, and besides, there are only 23 tests in the suite anyway.</p>

<p>My search tells me that four tests that were generated by Claude were not mentioned as a test killing a mutant in the results file. In all four cases, the reason behind this is that the exact same code path was already exercised in another test. For example, one of the generated tests performs a withdrawal on a checking account and verifies that the balance is updated accordingly:</p>

<figure class="highlight"><pre><code class="language-java" data-lang="java"><span class="nd">@Test</span>
<span class="kt">void</span> <span class="nf">withdraw_positiveAmount_fromCheckingAccount_updatesBalance</span><span class="o">()</span> <span class="o">{</span>
    <span class="kt">long</span> <span class="n">id</span> <span class="o">=</span> <span class="n">createAccount</span><span class="o">(</span><span class="nc">AccountType</span><span class="o">.</span><span class="na">CHECKING</span><span class="o">,</span> <span class="mf">500.0</span><span class="o">);</span>

    <span class="n">given</span><span class="o">(</span><span class="n">requestSpec</span><span class="o">)</span>
        <span class="o">.</span><span class="na">post</span><span class="o">(</span><span class="s">"/{id}/withdraw/{amount}"</span><span class="o">,</span> <span class="n">id</span><span class="o">,</span> <span class="mf">200.0</span><span class="o">)</span>
    <span class="o">.</span><span class="na">then</span><span class="o">()</span>
        <span class="o">.</span><span class="na">statusCode</span><span class="o">(</span><span class="mi">200</span><span class="o">)</span>
        <span class="o">.</span><span class="na">body</span><span class="o">(</span><span class="s">"balance"</span><span class="o">,</span> <span class="n">equalTo</span><span class="o">(</span><span class="mf">300.0f</span><span class="o">));</span>
<span class="o">}</span></code></pre></figure>

<p>Another test in the suite, however, does the exact same thing, but for a savings account:</p>

<figure class="highlight"><pre><code class="language-java" data-lang="java"><span class="nd">@Test</span>
<span class="kt">void</span> <span class="nf">withdraw_positiveAmount_fromSavingsAccount_withSufficientFunds_updatesBalance</span><span class="o">()</span> <span class="o">{</span>
    <span class="kt">long</span> <span class="n">id</span> <span class="o">=</span> <span class="n">createAccount</span><span class="o">(</span><span class="nc">AccountType</span><span class="o">.</span><span class="na">SAVINGS</span><span class="o">,</span> <span class="mf">500.0</span><span class="o">);</span>

    <span class="n">given</span><span class="o">(</span><span class="n">requestSpec</span><span class="o">)</span>
        <span class="o">.</span><span class="na">post</span><span class="o">(</span><span class="s">"/{id}/withdraw/{amount}"</span><span class="o">,</span> <span class="n">id</span><span class="o">,</span> <span class="mf">200.0</span><span class="o">)</span>
    <span class="o">.</span><span class="na">then</span><span class="o">()</span>
        <span class="o">.</span><span class="na">statusCode</span><span class="o">(</span><span class="mi">200</span><span class="o">)</span>
        <span class="o">.</span><span class="na">body</span><span class="o">(</span><span class="s">"balance"</span><span class="o">,</span> <span class="n">equalTo</span><span class="o">(</span><span class="mf">300.0f</span><span class="o">));</span>
<span class="o">}</span></code></pre></figure>

<p>Since the logic for handling withdrawals in the case of a high enough balance is exactly the same, no matter what the account type is, one test is enough to properly cover this path in the API code.</p>

<p>After removing these four tests from the suite and running mutation testing again, as expected, I can see that the impact on both line and mutation coverage is 0, meaning that these four tests can indeed be classified as ‘dead weight’.</p>

<p>Why go through this exercise, you might think? You practically got those extra tests for free when Claude generated them, isn’t it?</p>

<p>Well, sort of, but tests take time to run and to maintain, and most importantly, it takes time to ingest and process the information produced by them, especially in case of potential problems. The less the time I have to spend on that, the better. If I can safely remove tests from my test suite, and by safely, I mean without apparent adversarial impact on coverage, I would say that doing so is the prudent thing to do.</p>

<h3 id="conclusions">Conclusions</h3>

<p>So, after completing the analysis of the results of asking Claude Code to generate tests for a new code base, what do I think? Well, while I am impressed, I think a couple of words of warning are in order.</p>

<p>I am positively surprised by the quality and the coverage of the initial test suite. 95% line coverage and 91% mutation coverage are good numbers, and all that coverage was generated in a few minutes, definitely a lot less time than it would have taken me to write these tests myself.</p>

<p>There is some room for improvement in terms of readability of the tests, but that can probably be resolved by being more specific in my prompt and / or using dedicated <a href="https://code.claude.com/docs/en/skills" target="_blank">Claude Code skills</a>. I’ll explore that in more detail and write about that soon.</p>

<p>Now to the words of warning. While Claude achieved a pretty decent mutation coverage, it did oversee a few critical paths in the code. Maybe I was simply ‘unlucky’, and another attempt with the same prompt would have given better results. I don’t know, but these results do tell me not to simply accept what Claude gives me at face value.</p>

<p>The same applies to the tests that Claude <em>did</em> generate. 4 out of the 23 tests generated were dead weight, which equates to 17% of the test suite. Now, n = 1, and this is a small codebase and test suite, so the numbers might be skewed, but again, if you want your test suite to be as efficient and effective as possible, these are numbers that you probably don’t want to ignore.</p>

<p>Finally, there are of course many things that Claude did <em>not</em> do, mainly because I didn’t ask it to. An example of that would be telling me that since we’re working with a banking API, it probably would be a good idea to add some form of authentication to the endpoints. There’s a lot more to unpack about what Claude does and does not do, and I will probably write about that in more detail somewhere in the future, too.</p>

<p>First, in a follow-up blog post, I’ll document the process of improving the existing test suite that Claude generated, both in terms of coverage and of coding style. I will once again be using Claude and mutation testing to do that. You can find the code for the API that was used in this blog post, as well as the initial suite of tests generated by Claude, <a href="https://github.com/basdijkstra/writing-tests-with-claude-code/blob/main/src/test/java/com/ontestautomation/mutationbank/controllers/AccountControllerTest.java" target="_blank">here</a>.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Artificial Intelligence" /><category term="Claude Code" /><category term="test automation" /><category term="mutation testing" /><summary type="html"><![CDATA[In a recent post, I wrote about how I used Claude Code to analyze the code for RestAssured.Net and then perform a refactoring action, using handwritten tests as the safety net. In that post, I wrote that I didn’t want Claude to touch the tests themselves, and why. I was still curious, though, to find out for myself what Claude was capable of in terms of writing tests, and to what extent the trust that more and more people put into the tests written by LLMs is warranted.]]></summary></entry><entry><title type="html">Refactoring the RestAssured.Net code with Claude Code</title><link href="https://www.ontestautomation.com/refactoring-the-rest-assured-net-code-with-claude-code/" rel="alternate" type="text/html" title="Refactoring the RestAssured.Net code with Claude Code" /><published>2026-02-27T00:00:00+00:00</published><updated>2026-02-27T00:00:00+00:00</updated><id>https://www.ontestautomation.com/refactoring-the-rest-assured-net-code-with-claude-code</id><content type="html" xml:base="https://www.ontestautomation.com/refactoring-the-rest-assured-net-code-with-claude-code/"><![CDATA[<p>As some of you might know, the workshops and training courses I run, the talks I do and the blog posts I write tend to focus on fundamental sofware testing, software development and test automation skills, rather than <a href="/some-perspective-on-testing-trends/">focusing on the latest technology and trends</a>. I just don’t ‘do’ trends very well. I do, however, read a lot of what others think about and write about with regards to these technologies, as I am an independent consultant and therefore I simply cannot afford <em>not</em> to stay on top of what is happening in tech.</p>

<p>Until now, I have pretty much avoided spending a lot of time using AI tools myself, other than using ChatGPT to help me build a custom training plan for my endurance cycling training. That is, until I heard more and more folks talking about <a href="https://code.claude.com/docs/en/overview" target="_blank">Claude Code</a>, and how good they thought it was, especially with the new Opus 4.6 model. That triggered me to find out for myself if it was truly useful, and this blog post is probably the first of a few where I document my thoughts and findings.</p>

<p>Of course, I could start with building (or ‘vibe coding’, as the cool kids say) something from scratch, but I don’t think that is a particularly good way to find out what AI can really do for me. After all, I am not in the business of building new things that didn’t exist before, I’m in the business of helping people perform existing and important tasks, in my case, testing and automation, in a better and more efficient way.</p>

<h2 id="the-goal">The goal</h2>
<p>I’ve been working on <a href="https://github.com/basdijkstra/rest-assured-net" target="_blank">RestAssured.Net</a> for some 4 years now, and over time, the number of features has grown, as has the code base itself. As a result of tacking on more and more features, some of the classes in the project have grown to be very large, way too large, even, and those classes typically also have many different responsibilities.</p>

<p>I have been meaning to refactor them and extract different pieces of logic into their own classes to make the code easier to understand and to maintain, but in all honesty, I sometimes can’t see the forest for the trees anymore. So, it would be great if Claude Code could help me do that.</p>

<h2 id="the-guardrails">The guardrails</h2>
<p>One thing I’ve learned from my limited use of AI and from a lot of reading about other people’s experience with AI is to have strict guardrails in place before you’re letting an AI agent loose on your codebase. Since I don’t have a lot of experience with AI under my belt yet, I reckon it is a good idea to take small steps and remain in control of the process. I mean, <em>I</em> am responsible for the RestAssured.Net codebase, so I want to know what happens to it, and make sure I understand everything that is happening to it.</p>

<p>So, as a first guardrail, I’m not letting Claude Code touch the tests. The <a href="https://github.com/basdijkstra/rest-assured-net/tree/main/RestAssured.Net.Tests" target="_blank">acceptance tests for RestAssured.Net</a> act as a safety net for me when I fix bugs and add new features, and I write them with intent and <a href="/improving-the-tests-for-rest-assured-net-with-mutation-testing-and-stryker-net/">take good care of them</a>. As the task I set out to do is a ‘pure’ refactoring task, that is, I only want to change the code structure, not its behaviour, the tests should remain intact to prove that the refactoring was successful.</p>

<p>Second, I am doing a thorough review of every change that is made by Claude before I bring it under version control. My goal with this experiment, and with using AI in general, is not to outsource my thinking (thank you, <a href="https://www.linkedin.com/in/fionacharles/" target="_blank">Fiona Charles</a>), but rather to enhance my capabilities. As I said earlier, ultimately, I am the one responsible for the code and the changes made to it, not Claude. This is also why I don’t offload running the tests and bringing the code under version control to Claude.</p>

<p>Third and final, I’m going to take small steps. I have seen plenty of (horror) stories of people letting LLMs go for hours writing code, without intermediate scrutinizing of the changes and suggestions made, with results that ranged from the mildly amusing to the outright horrifying. I don’t claim that this code base is as important as, for example, the code in an online banking platform, but it has grown to a decent user base of the years. I don’t want to let those people down. And who knows, some of them might use the library to test those online banking platforms, so it’s my responsibility to only publish versions of a product that I feel is fit for that job. There’s no place there for code I don’t understand and that might give people a false sense of security.</p>

<h2 id="the-first-prompt">The first prompt</h2>
<p>After purchasing a Pro subscription to Claude Code, setting it up and having it initialize <a href="https://github.com/basdijkstra/rest-assured-net/blob/main/CLAUDE.md" target="_blank">the CLAUDE.md file</a> for the project, I first asked Claude to analyze the <code class="language-plaintext highlighter-rouge">ExecutableRequest</code> class (a class that’s been bothering me for a while because of its size and complexity) and suggest me improvements. I did that with this specific prompt:</p>

<blockquote>
  <p>The ExecutableRequest class is quite long, with a number of different responsibilities. Analyze it and suggest improvements to the code structure without changing the behaviour. List your top 5 recommendations, together with impact on code quality and reasons for your prioritization.</p>
</blockquote>

<p>I don’t want Claude to make changes yet, I want it to suggest me changes, and also, and more importantly perhaps, tell me <em>why</em> it thinks these improvements make sense. This should give me the information I need to make an informed decision on whether to proceed with the suggestion.</p>

<p>Opus 4.6 is slower than many other models out there, but the output is supposed to be much better. And indeed, after some deliberation, Claude produced a list of suggestions for improvement that at first glance, made a lot of sense. It’s #1 recommendation was to extract the logic to create a request body into a separate class called <code class="language-plaintext highlighter-rouge">RequestBodyFactory</code>. Not necessarily perfect as a class name, but as I couldn’t think of anything better at that moment, I asked Claude to proceed with doing the actual refactoring.</p>

<p>And so it did, and I have to say, it did not disappoint. It created the new class without problems, moved the logic in there, changed the <code class="language-plaintext highlighter-rouge">ExecutableRequest</code> class to use the methods in the new <code class="language-plaintext highlighter-rouge">RequestBodyFactory</code> class, and it even followed all the styling and formatting <a href="https://github.com/basdijkstra/rest-assured-net/blob/main/CLAUDE.md#code-style" target="_blank">requirements set by StyleCop</a>. That last point is very important, because I set the styling rules to level ‘nuclear’, as in, even the slightest infraction will make the code fail to compile.</p>

<p>The ultimate test of the work done by Claude, though, was running the tests, and those all passed, too. Which makes sense, as I asked for a change of the code structure, without changing the behaviour. <a href="https://github.com/basdijkstra/rest-assured-net/tree/main/RestAssured.Net.Tests" target="_blank">My tests</a> test the library for behaviour, not implementation, so this is exactly what I expected.</p>

<p>The only thing left for me to do was to review the changes made by Claude before committing them to version control. As I said at the start of this post, ultimately it is me who bears responsibility for the code, so I want to be able to read and understand it, even when I didn’t write it myself.</p>

<p>Overall, the changes made by Claude looked pretty good, but there was one thing I didn’t like. The newly created <code class="language-plaintext highlighter-rouge">Create()</code> method in the <code class="language-plaintext highlighter-rouge">RequestBodyFactory</code> class had a <em>lot</em> of arguments (9, if I remember correctly). So, naturally, I asked Claude if it could further improve that. It came back with a suggestion to group several properties related to request body settings in a custom <code class="language-plaintext highlighter-rouge">RequestBodySettings</code> type and then pass that object in as an argument. As I thought this did improve the readability of the code, I asked Claude to proceed with the change. Again, code compiled, tests passed, all good.</p>

<p>After that, I saw no further reasons to change or improve the work done by Claude, so I thought the code was safe to <a href="https://github.com/basdijkstra/rest-assured-net/commit/ceb79e644902984361303da6e15137f858f459cb">commit</a> and push. My build pipeline then took care of verifying that all tests run and pass on all .NET versions that RestAssured.Net supports. Again, no issues here.</p>

<h2 id="so-what-did-i-learn">So, what did I learn?</h2>
<p>Did this experiment teach me a lot of new things? Well, not necessarily. What it did, though, was reinforce my initial thoughts on what constitutes prudent use of AI in software development and software testing. It also showed me that Claude is both a really powerful and a user-friendly tool, at least when used on a codebase and for a task of this, admittedly small, size.</p>

<p>In short:</p>
<ul>
  <li>Software like Claude is a great tool for refactoring and code improvement tasks like the one described in this blog post</li>
  <li>Having solid guardrails in place (linters, tests, review before commit) is essential if you want to retain control of the software that is being written by AI</li>
  <li>At least initially, I would lean towards having AI support you in writing product code, and leave writing tests and reviewing the results to human beings</li>
</ul>

<p>With regards to that last point, I know that not everybody will agree with this. I have seen a lot of examples of AI writing tests, and of AI taking care of your code reviews. Personally, I’m not ready yet to hand over that kind of control to a piece of software that I do not fully understand or trust. At least, not without keeping a human being (myself, for example) in the loop. If you decide otherwise, that’s fine with me, as long as you can handle the responsibility, and are ready to deal with the potential fallout, that comes along with it…</p>

<p>In the meantime, I will continue improving the RestAssured.Net code using Claude and other tools, as there’s a lot of room for improvement left. And I think I’ll stick to keeping the current guardrails in place, too. It might be slightly slower, but it will be a lot safer, too.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Artificial Intelligence" /><category term="RestAssured.Net" /><category term="refactoring" /><category term="Claude Code" /><summary type="html"><![CDATA[As some of you might know, the workshops and training courses I run, the talks I do and the blog posts I write tend to focus on fundamental sofware testing, software development and test automation skills, rather than focusing on the latest technology and trends. I just don’t ‘do’ trends very well. I do, however, read a lot of what others think about and write about with regards to these technologies, as I am an independent consultant and therefore I simply cannot afford not to stay on top of what is happening in tech.]]></summary></entry><entry><title type="html">My LinkedIn break - six weeks in</title><link href="https://www.ontestautomation.com/my-linkedin-break-six-weeks-in/" rel="alternate" type="text/html" title="My LinkedIn break - six weeks in" /><published>2026-02-01T00:00:00+00:00</published><updated>2026-02-01T00:00:00+00:00</updated><id>https://www.ontestautomation.com/my-linkedin-break-six-weeks-in</id><content type="html" xml:base="https://www.ontestautomation.com/my-linkedin-break-six-weeks-in/"><![CDATA[<p>About six weeks ago, I decided to take some <a href="/time-for-another-linkedin-break/">time away from LinkedIn</a>. I won’t go into the reasons behind this decision again, you can read all about that in the post I just linked to, but I do want to take some time and look back on the past six weeks and the things that not spending so much time on LinkedIn have brought me.</p>

<p>First of all, moving away from spending an hour or two on LinkedIn wasn’t as easy as I thought. Especially early on, thoughts of ‘am I missing something?’ were in my head pretty much all the time, and yes, that led to me logging in and checking to see whether there was something I needed to address - DMs to answer, invites to accept - a few times.</p>

<p>After a few weeks, though, and after seeing that there wasn’t much of importance, or really anything at all, that I missed, that feeling slowly faded. It’s still there, sometimes, but I don’t feel the ‘need’ (it’s more of a ‘want’, really) to log in as often as I did in the beginning.</p>

<p>When I started my break, I was hoping for several positive side effects. It’s still early, too early to draw conclusions, but so far, things are looking pretty good:</p>

<ul>
  <li>The contents for my brand new ‘<a href="/training/valuable-feedback-fast/">Valuable Feedback, Fast</a>’ course is coming together slowly but surely, and I’ve got 3-4 companies interested in booking me to deliver this course in 2026, with one of them confirmed. My target is to deliver it at least three times in 2026, so things are looking good.</li>
  <li>My <a href="/training/">training</a> business is off to a good start, too, with 5 full days of training already delivered in January. I was off to a much slower start in 2025, so this makes me happy. I’m also working on different ways to bring my training offerings to the attention of potential clients. One thing I’m thinking about is starting a YouTube channel with instructional videos, to give people an idea of my teaching style and the type of content they can expect when they book me to teach a course.</li>
  <li>I’ve already published 4 blog posts, including this one, where I only wrote 13 in all of 2025. I really hope to keep up this pace.</li>
  <li>I’ve definitely been reading more, too. Mostly fiction, including the fantastic <a href="https://www.goodreads.com/book/show/112525.Winter_s_Bone" target="_blank">Winter’s Bone</a> by Daniel Woodrell, but I’m also slowly working my way through <a href="https://www.goodreads.com/book/show/201116293-taking-testing-seriously" target="_blank">Taking Testing Seriously</a> by James Bach and Michael Bolton.</li>
  <li>Outside of work, I’m definitely taking my cycling much more seriously. Even though January wasn’t the best cycling month, because of snow, the flu and some other things that got in the way, I have a training plan in place, I have set some ambitious but not-too-crazy goals, and I hope those will lead to my first century as well as my first 200k and even a 300k before the end of the year.</li>
</ul>

<p>Because of that last bullet point, and especially the time that training for long-distance cycling events takes, I have decided to pretty much entirely focus on my training business in 2026, as that is simply more flexible than consulting. This doesn’t mean I will not take on any consulting gigs at all, but the ones that I do will be of a very part-time nature. I don’t want to do all the cycling I want to do in the evenings and weekends only, if only because there’s no better feeling than going for a long bike ride at a time you know most other people are in the office ;)</p>

<p>Needless to say, I’ll remain absent from LinkedIn for the foreseeable future. Yes, I might log in once every other week to quickly check DMs and invites. You might see a <em>very</em> occasional post from me promoting my training services, but that will be once a month, tops, and probably even less than that.</p>

<p>I have zero interest at the moment in getting back to regular posting, commenting, sharing and liking. My brain relishes the quiet, the absence of that background noise that was present when I was spending a lot of time on LinkedIn. Plus, it has given me the time and attention it needs to do more valuable things, and I’m not ready to give that up again. Don’t expect me to write an update like this every month, either. As the months go by, I’ll probably think about LinkedIn even less, especially now that I’ve seen that I’m not really missing out on a lot.</p>

<p>As I said before, email is a much better way to get or stay in touch, so if you have a question, something to share, or you just want to catch up, email me at bas@ontestautomation.com, and I’ll happily talk to you. I’ve had some great conversations already, and I’m looking for many more of those.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Career" /><category term="social media" /><summary type="html"><![CDATA[About six weeks ago, I decided to take some time away from LinkedIn. I won’t go into the reasons behind this decision again, you can read all about that in the post I just linked to, but I do want to take some time and look back on the past six weeks and the things that not spending so much time on LinkedIn have brought me.]]></summary></entry><entry><title type="html">On the first training sessions in 2026</title><link href="https://www.ontestautomation.com/on-the-first-training-sessions-in-2026/" rel="alternate" type="text/html" title="On the first training sessions in 2026" /><published>2026-01-18T00:00:00+00:00</published><updated>2026-01-18T00:00:00+00:00</updated><id>https://www.ontestautomation.com/on-the-first-training-sessions-in-2026</id><content type="html" xml:base="https://www.ontestautomation.com/on-the-first-training-sessions-in-2026/"><![CDATA[<p>Earlier this week, I ran my first two <a href="/training/">training</a> sessions of 2026. Running these sessions reminded me (once again) that training really is what I enjoy the most and what I do best. Working with a group of engineers, introducing them to new techniques, patterns, principles and tools, and exploring and discussing how these could help them improve their automation, testing and development efforts to get closer to the end goal of <a href="/training/valuable-feedback-fast/">valuable feedback, fast</a> is a very rewarding thing to do.</p>

<p>On Tuesday, I facilitated a full-day course covering the ‘textbook’ Behaviour-Driven Development process. In this course, we took a change to an existing feature for <a href="https://parabank.parasoft.com" target="_blank">the world’s least safe online bank</a> from idea to working to automated acceptance tests. Along the way, we discussed and practiced Example Mapping, reviewed and wrote Gherkin scenarios, and took a quick look at the tools that support BDD, in this case <a href="https://cucumber.io/docs" target="_blank">Cucumber</a>.</p>

<p>As a course topic, BDD is a bit of an outlier for me. Most of the workshops I run cover topics that are related to test automation, yet in this workshop we cover the BDD process and methodology as a whole, and not <em>just</em> talk about Cucumber and Gherkin. Tuesday reminded me how fun it can be to talk about and practice BDD, though, so I do hope I get to run this workshop a couple more times in 2026. I’ve got at least one more planned, so that’s a start.</p>

<p>The next day, I ran a full-day workshop on <a href="/training/contract-testing/">contract testing</a> with Pact in Java. I quite recently redesigned this workshop, because I felt that the way I taught it before didn’t really do justice to the complex topic that is contract testing. Before, I often tried to cram it into a half-day session, which is just too short, and therefore had to limit myself mostly to writing and running consumer-driven contract generation tests on the consumer side and verify contracts at the provider side. If there was time left, we talked a little bit about <a href="/approaches-to-contract-testing/">challenges of consumer-driven contract testing and how bidirectional contract testing tries to address those challenges</a>, but that was about it.</p>

<p>In the new format, I’m taking more time, using more and smaller steps, to take people through the flow of both consumer-driven and bidirectional contract testing. There’s some coding involved, of course, but not an awful lot. The biggest change, I think, is that I now focus on the process flows rather than just the tools. For example, we now use a Pact Broker from the very start, and we use tools like <a href="https://docs.pact.io/pact_broker/can_i_deploy" target="_blank">can-i-deploy</a> throughout the workshop, too. From the initial feedback I gathered, this helps people a lot in understanding what contract testing really is, how it works and how it tries to answer the question of</p>

<blockquote>
  <p>“Are all individual components and services able to communicate with one another?”</p>
</blockquote>

<p>So, in short, 2026 is off to a good start when it comes to my training business. As my goal for the year is to grow that training business revenue with 20% compared to 2025, <a href="/time-for-another-linkedin-break/">while staying away from LinkedIn and social media in general</a>, I’m happy to see that next to the two I delivered this week, I have three more days of training scheduled for January, all three on <a href="/training/workshop-playwright/">Playwright</a>. That’s a big improvement over last year, when my training sessions didn’t really start until early March. Needless to say, I hope that this trend will continue throughout the year. I’ll keep you updated.</p>]]></content><author><name>Bas Dijkstra</name></author><category term="Training" /><category term="BDD" /><category term="contract testing" /><category term="Playwright" /><summary type="html"><![CDATA[Earlier this week, I ran my first two training sessions of 2026. Running these sessions reminded me (once again) that training really is what I enjoy the most and what I do best. Working with a group of engineers, introducing them to new techniques, patterns, principles and tools, and exploring and discussing how these could help them improve their automation, testing and development efforts to get closer to the end goal of valuable feedback, fast is a very rewarding thing to do.]]></summary></entry></feed>