
My case for putting software engineering back at the center of the AI conversation.
I am excited about what AI can bring to software development. When I read about teams using agents to implement features, test their work, and open merge requests, I can see why we want to move quickly.
But there is a question I keep coming back to: how quickly can we deliver a change today, without AI?
Because that answer tells me a lot more about our readiness than the coding assistant we have chosen.
We hear two very different stories. Companies like Anthropic describe substantial productivity gains. Other teams struggle to turn their AI experiments into faster delivery or better business outcomes. I think we spend too much time discussing the tools and too little time examining the system those tools are entering.
So here is my deliberately provocative position: if you want to do AI, start by not doing AI.
Put the tooling discussion aside for a moment. Look at how your team actually delivers software.
I would start with something almost embarrassingly simple. Take a web application. Agree with your business counterpart that you want to change the title of the index page and add an emoji.
Then do it. How long before users see the result?
For a change that small, I want the answer to be seconds. Depending on the stack, the build, checks, and rollout may take a few minutes. I can understand that. What I struggle to accept is a change like this spending days moving through our organization.
A day waiting for review. Another day waiting for an environment. Two weeks before somebody from the business can run the regression tests. Then a release slot.
What exactly would AI accelerate here?
It can help us edit the title faster. We have barely touched the delivery time.
That is why this little example matters to me. There is almost no implementation difficulty to hide behind. It exposes the workflow. Once the intent is agreed and the change satisfies our established controls, I want the technical path to production to run automatically.
The point is to understand every delay. A database migration deserves different evidence from a title change, and some releases need human judgment. But I want us to be able to explain the time we spend, rather than accept it because that is how the process has always worked.
DORA’s 2025 research describes AI as an amplifier of an organization’s strengths and weaknesses. That is very close to how I think about this problem. If our delivery system works well, AI has something useful to amplify. If our system is full of queues, unreliable checks, and handoffs, producing more code can put more pressure on those constraints. 1
I also think we should be careful with the success stories. Anthropic’s December 2025 research reported a self-assessed 50% productivity increase among its engineers and researchers. Most respondents still said they could fully delegate only 0-20% of their work. Those are encouraging results, with supervision still very much part of the picture. 2
On the other side, METR found that experienced developers working in familiar repositories took 19% longer with early-2025 AI tools. Its February 2026 update pointed toward improved results with newer tools, while explaining that selection effects made the size of the gains difficult to establish. I would use these findings as a reason to measure our own results carefully. The tools, tasks, and teams matter. 3 4
What interests me most in Anthropic’s work is the engineering behind the autonomy.
In its C compiler experiment, 16 agents worked across nearly 2,000 sessions. Nicholas Carlini explains that much of his effort went into the tests, environment, and feedback that allowed the agents to make progress. When new features started breaking existing behavior, he strengthened continuous integration. The compiler had impressive capabilities and real limitations. 5
That is the part I want teams to pay attention to.
The same theme appears in Anthropic’s work on long-running agents: clear requirements, a repeatable environment, incremental changes, Git history, progress notes, and end-to-end verification. These helped address unfinished work and premature claims of success. 6
My reading is that we can learn a great deal from how these teams organize the work around the model. They make progress observable. They give the agent useful feedback. They invest in a way to establish whether a change actually works.
Those are software engineering concerns. We can start addressing them before deploying an agent.
For me, the first practical step is value stream mapping.
Get the people involved in delivery into the same discussion: business, product, developers, testing, operations, and support. Follow a few actual changes from the original need to the result in production. Measure where people work, where work waits, and where it comes back because something was missing or incorrect. DORA’s guidance gives a useful basis for this exercise. 7 8
An illustrative map for our title change could look like this:
| Step | Active work | Waiting |
|---|---|---|
| Edit and commit | 2 minutes | 0 |
| Review | 5 minutes | 1 day |
| Automated checks | 4 minutes | 20 minutes |
| Manual regression testing | 10 minutes | 2 days |
| Production deployment | 3 minutes | 1 day |
Looking at that table, I know where I would put the team’s effort. I would work on the waiting, the repeated manual validation, and the release dependency.
The arithmetic is straightforward. If implementation represents 10% of total elapsed delivery time, making it twice as fast saves 5% overall, assuming the rest stays unchanged. We should welcome that improvement, while being honest about what it achieves.
I want us to understand delivery as a whole. Otherwise, we can improve one activity and leave the business waiting almost as long as before.
This is also why I am a big fan of Chuck Rossi’s work on release engineering at Facebook.
His 2012 account describes investment in test automation, more frequent releases, and engineers taking responsibility for their changes all the way to production. 9
Facebook’s later account explains how increasing release batches and manual coordination became unsustainable. The move toward quasi-continuous delivery involved automated tests, progressive rollouts, monitoring, and feature controls. It took serious infrastructure work and a dedicated release engineering team. 10
What I take from that is a commitment to making delivery scale. They worked on the system that allowed engineers to release safely and frequently.
That is the ambition I want us to bring to our own teams.
It starts with a version control workflow people understand and can use efficiently. Small changes. Clear review expectations. Frequent integration. A way to recover when something goes wrong. I would challenge long-lived branches and lengthy stabilization phases. DORA’s guidance on trunk-based development supports this direction; a complex Gitflow model is not a prerequisite for maturity. 11
Then there is code quality. I want tools such as SonarQube in the delivery workflow, with relevant rules and quality gates that influence whether a change can proceed. Its focus on new and modified code provides a practical way to improve an existing codebase gradually. 12
But the team has to own the findings. Installing the tool and looking at a dashboard once a month will not create the discipline we need. And a green quality gate cannot tell us whether we have implemented the right business rule.
The same applies to tests. I care about what they protect.
If we retry an order request, can we create a duplicate transaction? What happens at a quantity boundary? Can an order enter a state the business does not allow? Those are the kinds of questions I want our tests to answer.
Fast unit tests, relevant integration tests, and selected end-to-end tests should give us trustworthy feedback throughout development. DORA emphasizes reliable automated suites and continuous testing across the lifecycle. 13
For applications connected to other services, I would go further into contract testing.
Imagine our order service consuming availability from an inventory service. We can write a mock that returns exactly what we expect. Our unit tests will pass. But who checks that the real service still behaves that way?
Consumer-driven contract testing makes those expectations executable on both sides. Pact supports testing HTTP and message interactions in this way. 14
When versioned contracts, provider verification results, and deployment records are maintained, Pact’s can-i-deploy can check compatibility with application versions in the target environment. That can help us avoid coordinating every dependent service into the same release. 15
I would still test the stock calculation, persistence, and critical business journey separately. Contract tests establish agreement about interactions; they do not establish all the functional behavior behind them. 16
The environment and pipeline need the same attention. I want a team to be able to build, test, and deploy without a sequence of undocumented steps. Known inputs should give us repeatable results. An unchanged commit that passes on the third attempt is a problem we should investigate.
DORA recommends deploying the same packages through the same process across environments, separating configuration, and keeping the information needed to recreate environments in version control. 17
None of this removes the business from delivery. I want business counterparts involved early, helping us define the intended behavior and acceptance examples. Their judgment matters. What I want to reduce is the recurring wait for someone to repeat a regression script we already understand.
Continuous delivery gives us software we can release on demand. Continuous deployment takes the next step for eligible changes. We can build the first capability even when some releases require explicit approval. 18
Feature flags also help us separate deploying code from exposing a feature to users. The business can decide when to release a capability while technical deployment continues regularly. We need to manage the flags, test their relevant states, and remove them when their job is done. 19
The ownership question is fundamental to me.
I want the product team to own its code, its tests, its quality, and its production behavior. When the pipeline fails, the team needs the knowledge and access to fix it. DORA’s research highlights the value of developers maintaining automated tests and being able to resolve acceptance failures themselves. 13
If every routine action requires another team to pick up a ticket, we have built dependencies into our delivery model. DORA’s work on loosely coupled teams makes independent testing and deployment a central objective. 20
As a platform manager, I see a clear responsibility here. Shared platforms should make that autonomy easier: reusable pipelines, accessible test environments, observability, and controls that teams can use directly. DORA’s platform guidance describes this role in terms of self-service and reducing underlying complexity. 21
The squad should have support and a dependable platform. It should also be able to move.
Once we have that foundation, adding an agent becomes much more concrete.
We give it a business request with an intended outcome, acceptance examples, and relevant constraints. We make the business process documentation, architecture decisions, repository instructions, and service contracts available. It implements a focused change and opens a merge request.
The pipeline evaluates that change. Code quality. Business behavior. Service compatibility. Integration. Relevant nonfunctional requirements, such as response time, authorization, transaction integrity, or recovery.
The team reviews it according to the risk, and the established delivery path takes over.
That is a workflow I can understand, assess, and improve.
The context layer matters here too. An agent needs to understand what an order state means, when cancellation is allowed, and which exceptions apply. We need to maintain that knowledge as the system changes. A ticket alone will often leave too much unsaid.
I would also be careful about letting an agent define both the implementation and the entire basis for judging it. If it misunderstands a rule and writes matching tests, everything can be green and still be wrong. The team remains responsible for checking the acceptance expectations.
There is another reason I think these foundations matter: security patching.
We hear growing excitement, and concern, about AI finding vulnerabilities that people have missed. I take that seriously. Anthropic’s February 2026 account describes tools that identify vulnerabilities and propose patches for human review. Faster discovery can increase the volume of findings teams must investigate and fix. I expect that to put more pressure on remediation, although findings do not all become CVEs and a rise in CVE counts would not, by itself, tell us how much is caused by AI. 23
But I want us to look at what the same technology can do inside our software delivery factory. If AI helps us find problems faster, our engineering system should also help us apply fixes faster.
Sometimes the response to a CVE is a small dependency upgrade, a lockfile update, and perhaps a few lines of adaptation. With a maintained codebase and a dependable pipeline, that change can run through unit tests, contract and integration tests, regression checks, and the relevant security checks before deployment. We can confirm that the affected version has been replaced and that the fix addresses the vulnerability, alongside checking that the application still works.
We already have part of this workflow with dependency automation. GitHub documents how Dependabot updates can trigger tests, enter review, or merge automatically under configured rules. An agent can contribute analysis or remediation recommendations while deterministic pipeline checks continue to control whether the change proceeds. 24
Other fixes are harder. A secure version may require a major library upgrade, changes to our code, or updates to other dependencies. GitHub’s troubleshooting guidance describes cases where a vulnerable dependency cannot be upgraded without breaking the dependency graph. 25
That is precisely where I want meaningful regression tests. They give us evidence about what broke and what still behaves as expected. An agent could read the migration notes, propose the upgrade, run the checks, investigate failures, and adapt the affected code. We would still evaluate the security fix and the remaining risks. Passing functional tests alone does not prove that an application is secure.
For an eligible, well-understood update, I can imagine an automatic path from alert to proposed change, validation, deployment, and production verification, within the team’s agreed policy. For a larger migration, I would expect more review and narrower steps. A fully automated pipeline makes either path easier to execute repeatedly; the degree of agent autonomy can grow as the evidence supports it.
This is where the title-and-emoji example connects to a much more demanding problem. The delivery capability that handles a tiny change efficiently can also help us absorb a stream of urgent security updates. As complexity grows, we need stronger evidence, but we should not have to rebuild the delivery process each time.
That is the kind of augmentation I want: a team with the foundations to keep up, using agents to extend its capacity. Faster vulnerability discovery becomes much more useful when we can shorten the time from finding a problem to getting a verified fix into production.
We do not need a perfect engineering organization before experimenting. But I would give the next few weeks a clear sequence for one application:
- Week 1: Follow recent changes through delivery and run the title-and-emoji exercise. Make the constraints visible.
- Week 2: Work on the biggest recurring constraint: unreliable checks, review delays, manual steps, or environment dependencies.
- Week 3: Demonstrate a repeatable route for an agreed low-risk change, with meaningful validation and a tested recovery path.
- Week 4: Give an agent a bounded task through that route. Measure the result, including review effort, rework, and production behavior.
I would measure success through the delivery flow and the business outcome. DORA’s delivery metrics provide a useful starting point: change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. Add the time from business request to user outcome and the waiting at each stage. 22
I would also make room for this work in the roadmap. We cannot keep the same feature commitments, add an AI initiative, and expect teams to repair years of delivery friction in whatever time remains.
This is what I mean by starting without AI. Make the engineering problem the priority. We can use AI to help write tests, improve documentation, or propose pipeline changes along the way, with proper verification. The work still needs to address the constraint.
My concern is that AI becomes a smokescreen: another exciting initiative that consumes our attention while the same delivery problems remain.
I want us to use it well. That means being willing to work on the less glamorous foundations that make the impressive results possible.
So my first challenge to a team would be very simple.
Agree on the title. Add the emoji. Get it into production.
Then look honestly at what happened between those steps.
That is where I would start our AI journey.
Sources and further reading
Primary research, engineering accounts, and official documentation consulted on October 7-8, 2026. The numerical value stream example and four-week sequence are the author’s illustrations and recommendations.
-
Anthropic, How AI is transforming work at Anthropic, December 2, 2025.
-
METR, We are Changing our Developer Productivity Experiment Design, February 24, 2026.
-
Anthropic, Building a C compiler with a team of parallel Claudes, February 5, 2026.
-
Anthropic, Effective harnesses for long-running agents, November 26, 2025.
-
Facebook Engineering, Release engineering and push karma: Chuck Rossi, April 5, 2012.
-
Facebook Engineering, Rapid release at massive scale, August 31, 2017.
-
SonarSource, SonarQube Server 2026.1: Quality standards and new code.
-
Anthropic, Making frontier cybersecurity capabilities available to defenders, February 20, 2026.