Climbing the Path and Closing the Loop

This blog post is a summary of my takeaways from GOTO Copenhagen 2026, regarding the question that ran through almost every talk I attended: how do we introduce AI into the way we build software in a way that makes us not only faster, but better too? These are my own takeaways and readings of the talks, and not necessarily the official positions of the respective speakers.

TL;DR: AI amplifies both the strengths and the weaknesses of whatever system it is introduced into. What it provides should be seen primarily as increased optionality, through the capacity it unlocks, as opposed to merely speed. If we want to maximize the return on our investment in AI, that capacity needs to be reinvested in how our teams build software. The teams that stand to gain the most are the ones willing to invest in building and maintaining the system that builds the system: the software factory, progressing from initial one-shot prompting towards supervising teams of agents that can increasingly verify and correct their own work.

What We Don’t Yet Know

One thing that united both speakers and attendees was that nobody felt they could say for sure what software development will look like in five or ten years.

And it’s not just a question of the tools we use. How we build software is an open question too: who writes the code, who reads it, and how are development teams composed to begin with?

There are no turnkey solutions yet, let alone silver bullets. Every speaker I listened to acknowledged that the field is simply too young for its tools, practices and organizational patterns to have settled. What we do have, however, are some increasingly consistent signals about what does help.

What We Do Know

AI Is an Amplifier

Nathen Harvey’s keynote drew on DORA’s 2025 State of AI-assisted Software Development report, which he co-authored. It’s based on survey responses from nearly 5,000 technology professionals. Harvey’s talk left me with three closely related takeaways:

  • First, AI is an amplifier: it magnifies the strengths and weaknesses of the systems into which it is introduced.
  • Second, the real return on AI is optionality, in the form of the capacity it unlocks and the choices that capacity gives us.
  • Finally, realizing that return depends on reinvesting some of that capacity in the team and in the system around it. Without that investment, localized productivity gains are simply absorbed by downstream bottlenecks.

Three years of DORA reports illustrate why: the gains show up at an individual level, but they do not automatically survive the journey to production.

DORA reportIndividual-level findingsDelivery performance
2023Slightly improved well-beingNeutral or perhaps negative for team and delivery performance
202475% feel more productive
39% have little or no trust in AI-generated code
Per 25% increase in AI adoption:
Throughput -1.5%
Stability -7.2%
202590% use AI at work
Over 80% feel more productive
30% have little or no trust in AI-generated code
Throughput up
Instability still up

NB: The 2024 delivery figures are modeled estimates per 25% rise in AI adoption; the 2025 report instead gives the direction of each relationship.

The picture looks better in 2025. In 2024, increased AI adoption was associated with both lower throughput and lower stability. By 2025, the relationship with throughput had turned positive, but instability was still increasing. This suggests that the individual productivity gains are beginning to make their way through the delivery system, but the system is not yet able to absorb them cleanly.

This is where the amplifier metaphor matters. Increasing the rate at which developers can produce changes also increases the pressure on everything downstream: review, testing, deployment and feedback. Without the right foundations, local productivity gains can simply shift the bottleneck downstream. As Harvey’s slides put it, those “localized pockets of productivity” are often lost to “downstream chaos.”

Tellingly, the seven capabilities in the report’s AI Capabilities Model are mostly not about AI: small batches, strong version control, quality internal platforms, a focus on users. One of Harvey’s slides quoted Deming: “A bad system will beat a good person every time.” It would seem that a bad system will beat a good agent too.

Speed Now, Debt Later

Cole Medin’s talk on the Validation-First Loop described one of the traps waiting for us early in AI adoption: generating code is becoming a lot faster than validating that code. As he put it, the problem is increasingly not that agents write bad code, but that they write good code that simply “misses the point.” The failure happens at “intent time”: whatever the team hasn’t made explicit becomes an assumption the agent has to make, effectively making decisions on the team’s behalf.

That creates what Medin calls a “productivity mirage”. His slides contrasted a METR study showing developers feeling faster while actually completing real tasks 19% more slowly. Faros’s 2026 AI Engineering Report found 66% more epics completed per developer, but also 242% more incidents per pull request. More output and more activity are not necessarily the same thing as better software.

The longer-term evidence points in the same direction. A 2026 study by He et al. compared open-source projects that adopted Cursor with matched projects that did not. Development velocity increased significantly, but only temporarily; static-analysis warnings and code complexity increased and remained elevated. The researchers found those increases in warnings and complexity to be major factors driving the longer-term slowdown. One important caveat is that the study measures static-analysis warnings and complexity, not production defects.

Adam Tornhill, in his talk on AI-Friendly Code, explained why allowing this kind of complexity to accumulate is especially dangerous once agents enter the development loop. His argument was that technical debt becomes a multiplier: poor-quality code does not just slow humans down, it also makes the agents working on it less reliable. Research he presented found a 60% higher risk of defects when AI was applied to problematic code, while other work showed initial velocity gains quickly being cancelled out as complexity accumulated. In his words, “the consequences of past decisions accumulate.”

Tornhill’s proposed way out was itself agentic: give the agent an objective quality signal, require it to correct unhealthy code, and re-measure after every iteration.

The answer, then, is not simply to generate code faster, but to build the system that keeps that speed from turning into debt.

What Helps Humans Helps Agents

If nearly every speaker agreed on one thing, it was this: whatever makes development easier for humans makes agents more effective too: fast feedback, error messages that state how to recover, rules that are written down instead of living in someone’s head.

That makes it a no-regret move. Even if agents somehow look completely different in a few years, the humans still benefit. It works the other way round too: the tacit rules you write down for an agent become the onboarding guide you never got round to writing.

Ironically, for systems built around inference, agents are often worse than humans at inferring the intent behind a task. Gojko Adzic, whose performance path I’ll return to shortly, said the fitness functions we design for agents need to be “much more detailed than with humans,” and also argued for putting recovery instructions directly into error messages. That, of course, helps human developers too.

How We Move Forward

Taken together, these talks tell me we should be building self-reinforcing systems around both human developers and agents rather than viewing AI adoption as a one-off transformation effort. Process rather than project, if you will.

So what does building that system look like?

Climbing the Path

Gojko Adzic’s talk gave the climb a shape. His performance path, which his abstract says comes from research with early adopters, has five levels:

Every step up the path moves more knowledge out of human developers’ heads and into the system, while shifting humans from directing the work toward supervising it.

The job of the human changes on the way up the path. At the Dictating level you are steering every step of the process. At Regulating, the plans state what to do, while common context and rules say how. At Supervising, agents bring evidence for humans to certify rather than changes to review. Skip this structure and, in his abstract’s words, “speed quickly turns into long-term maintenance risk.”

TransitionDescription
Dictating to CommandingPay attention to repetitive tasks
Start documenting the process
Create commands, skills and task context files
Implement the Research-Plan-Implement loop
Commanding to RegulatingPay attention to repetitive parts of tasks
Separate contextual and task rules
Turn large commands into sequences
Shift feedback left and evolve process commands to prevent recurring problems
Regulating to OrchestratingPay attention to sequences of commands
Iteratively refine human specifications with agents
Stabilize and generalize process files
Move common feedback into a computational layer
Replace confidence with evidence
Orchestrating to SupervisingPay attention to workflow variance
Use multiple evidence signals
Improve feedback speed and reliability
Move mechanical orchestration into a computational layer
Introduce an observability platform

Importantly, Adzic’s destination is explicitly a lights-on factory, not a fully autonomous one. It still aims for a high degree of automation, but the work is driven by specifications and plans rather than by “vibes”, orchestration belongs to the surrounding platform rather than to the agents themselves, deterministic guardrails constrain what agents can do, and humans supervise throughout. That means the platform itself increasingly has to serve agents as well as humans. This ties back to Harvey’s earlier slide listing agents alongside developers as customers of the internal platform.

Two parts of that progression struck me as especially useful:

The transition from personal to team. Early on, prompt files and commands are usually personal and not shared. His fix is unglamorous: a common development container, a shared process and decision log in version control, and the same deterministic checks reproduced in CI. DORA’s Version Control capability recommends the same approach: keep prompts and agent configuration files in version control. Its AI Capabilities Model report found that only about 21% of respondents reported doing so with prompts.

The transition from inference to computation. Steps an agent keeps re-deciding can be moved out of model inference and into the computational layer. In other words, once a decision becomes sufficiently mechanical and repeatable, we stop paying a model to infer it. We encode it directly into the system instead. In Adzic’s example, this resulted in 80% fewer tokens, 6-10x faster execution, no testing gaps and no approval requests.

Closing the Loop

If nobody reads every line of code anymore, what keeps quality up? Vilhelm von Ehrenheim’s answer, in a talk titled Nobody Reads the Code Anymore, is that the rules and judgment developers traditionally apply during review have to be encoded into the system itself. Invariants capture what must always hold, judges assess what is allowed to vary, and failures return evidence the agent can act on: “proving the behavior instead of reading the diff,” as his abstract puts it.

The problem is that spec-driven development on its own is still an open loop: intent goes in and code comes out. Agents optimize for observable success criteria, meaning what your checks actually measure, rather than the intent behind it. Closing the loop means feeding a verifiable outcome back into the process, so the agent can tell whether what it built actually satisfies what was asked for.

The closed loop is simple. The agent runs the checks, fixes what fails, and then runs them again, with no human involved. By the time the change reaches a person, the agent has already iterated against the verifier and can surface the evidence it collected, allowing the human to merge it with more confidence.

That confidence matters more than it sounds. The system of checks, judges and feedback that closes this loop is what von Ehrenheim calls the verifier. DORA’s ROI report calls the effort spent checking AI output a “verification tax.” Low trust makes that tax expensive: time saved during generation is simply re-spent auditing the result. A verifier the team trusts is how you stop paying that tax manually, one change at a time.

What goes into the verifier? Two kinds of checks. Invariants cover what must always hold: decisions someone senior made, written so a machine can enforce them. For example, new database migrations must be ordered after all migrations already on the main branch. Judges cover what is allowed to vary, such as risk scoring, goal-based user journeys and agentic change review.

A failing check also has to be useful. A non-zero exit code tells the agent almost nothing. A good failure gives it evidence and shows it where to look: expected versus actual results, a screenshot at the break, a step trace. Adzic made the same point in his talk. His slide on automating the automation asks for agents that self-evaluate and self-correct, and for recovery instructions in error messages.

The verifier therefore needs its own feedback loop. Its benchmark, built from previously discovered bugs and failures, has to be rerun regularly as the product, models and prompts change, so that we keep measuring whether the verifier still catches the failures we care about.

Von Ehernheim also named four ways the verifier itself can fail while its dashboard stays green: the harness can overfit its benchmark, a judge can drift and become too lenient, the benchmark can rot as the product changes, or the score itself can become the target through Goodhart’s law. In every case, the numbers can improve while the verifier gets worse. That is why its feedback loop matters so much.

Von Ehrenheim’s final advice was to “work on the factory, not in it.” Make the checks agent-runnable. Turn senior engineers’ rules into executable invariants. Give every pull request a preview environment and let an agent use it first. Then keep improving the verifier itself using past bugs and feedback from the agents using it.

That, to me, is what reinvesting the capacity that AI unlocks ultimately looks like: spending less time manually pushing each change through the system, and more time improving the system that can build, verify and correct those changes for us.

Importantly, closing the loop does not mean removing the human from the process. Like Adzic’s lights-on factory, von Ehrenheim’s model keeps the human as the final approver. What changes is where the human spends their attention: the agent is expected to run the checks, correct its own failures and surface evidence first, so the human supervises the result rather than manually driving every iteration.

What Still Earns a Human Read?

As far as I could tell, the conference was split roughly down the middle on whether humans should still read agent-written code. I’ve come to think the better question might be how you decide what earns human attention.

Adzic’s performance path suggests why both camps can be right. At Regulating, detailed code review is part of the process. At Orchestrating it becomes high-level review. By the time we reach Supervising, agents bring evidence and humans certify it. Even von Ehrenheim, whose talk title says nobody reads the code anymore, still keeps a human as the final approver.

Risk should set the dial. von Ehrenheim scores the blast radius of every change, and the depth of verification follows the risk. The same principle can govern how much human attention a change receives: decide what earns a read based on the team’s position on the performance path, the blast radius of the change and the strength of the evidence.

I wouldn’t personally stop reading code altogether, though. A verifier can only reason from the signals and evidence available to it. Human review still has an important role in identifying what the system does not yet know how to check, and in preserving the judgment needed to write the next invariant.

That also means I don’t necessarily see the choice between lights-on and dark software factories as a strict dichotomy. How much human attention the process needs depends on the maturity of the system, the risk of the change and the strength of the evidence. The factory can get progressively darker as confidence grows, but I believe we should always leave a light on.

The Forest and the Desert

Kent Beck’s closing keynote, The Augmented Forest, added one final constraint to all of this: the software factory is a socio-technical system, not merely a technical one. Its performance depends as much on incentives, trust and how people are organized as it does on code, tools and platforms.

In Harvey’s terms, AI will amplify that socio-technical system too. In Adzic’s terms, an organization has to create the conditions in which teams are allowed to climb the performance path. And in von Ehrenheim’s terms, teams need to be empowered to work on the factory rather than merely inside it.

Beck uses the following metaphor to describe two very different environments for building software: in the desert, resources are scarce and treated as something to conserve. In the forest, resources are plentiful, replenishing and available to invest in making the system healthier. AI, Beck argued, is likely to intensify whichever environment it enters.

The distinction matters enormously once AI starts freeing capacity. In a desert, the obvious response is to harvest the gain: generate more with fewer developers. In a forest, that capacity instead becomes something to reinvest: better tests, healthier code, faster feedback, stronger platforms and a better verifier.

This is also where DORA’s advice becomes explicitly socio-technical. Its 2026 ROI guidance is clear: reinvest capacity rather than reduce headcount. Earlier DORA work framed much the same distinction as treating technology as a value driver rather than a cost center.

To me, that is the final condition on everything else in this post. We can climb the performance path, build the platform and close the verification loop, but whether the optionality AI creates becomes better software or merely lower cost is ultimately an organizational choice.

Sources

Talks at GOTO Copenhagen 2026

References


Nästa inlägg ("Remissvar SOU 2019:14: Ett säkert statligt ID-kort – med e-legitimation") >>


Dela på:    
Oscar Jacobsson
Oscar Jacobsson

Oscar bor i Stockholm med sin fru och två katter. När han inte är upptagen med att valla katter eller kodande agenter hittar man honom oftast på fotbollsläktaren eller nere på puben.