Skip to content
RJJ Software Jamie Taylor · fractional CTOBook a call
Menu

Writing · 19 December 2025

When Agile Isn’t Agile: How DORA Metrics Reveal What’s Really Broken

Your team runs every Agile ceremony and a small change still takes weeks. The DORA delivery measures show where that time actually goes.

A colourful chain of a paper clips, laid out on a surface. The paper clips alternate in colour between red and blue. The central paper clip in the link is broken. The paper clips are laid on an asphalt looking background, implying that they are laying in the street.

ℹ️ Rewritten 18 September 2026

I first published this post on 19 December 2025, and I have since rewritten it. It now uses the DevOps Research and Assessment (DORA) programme’s five current measures, including the one renamed in 2023 and the one added in 2024, and it corrects a claim in the original that these measures cannot be gamed. They can. The client examples that remain carry only the figures I can stand behind.

“We’re Agile, but it takes six weeks to get a simple change to production.”

The tech lead who told me that was not exaggerating for effect, and they were not describing a team that had skipped the reading. They had the daily standups, the sprint planning, the retrospectives and the story points. Their burndown charts looked exactly as the training material said they should. Velocity was consistent enough to forecast with.

Delivering anything to a customer still felt like pushing a boulder uphill, and nobody in the organisation could say why. That is the situation the DORA metrics were built for, because the ceremonies tell you what the team does on a Tuesday, and the measures tell you what reaches customers.

The Ceremonies Are Not the Point

A great many teams describe themselves as Agile when what they are running is a waterfall chopped into two-week lengths. The requirements still arrive fully formed at the start, the work still moves through the same phases in the same order, and the release still happens at the end. The iteration is real; the ability to change direction is not.

The symptoms are easy to recognise once you know to look for them. Work is declared done when it merges, then waits weeks for a release window. Retrospectives produce actions that nobody has time to take, so the same three complaints reappear every fortnight. Standups become a status report delivered standing up, because the people who need to coordinate are not the people in the room. None of that is a failure of discipline, and adding another ceremony will not shift any of it.

The reason this persists is that it is invisible from the inside. Every local signal says the process is healthy. Velocity is stable, tickets close, the board moves. What nobody is measuring is the thing the customer experiences, which is how long it takes to ask for something and get it.

The Five DORA Metrics, and What They Are Called Now

The DORA programme has spent more than a decade researching what separates teams that deliver well from teams that do not. Its measures describe delivery rather than effort, which is what makes them worth the trouble of collecting.

There are five of them, and the list has changed since most of us learned it. Deployment frequency is how often you put a change into production, and change lead time is how long a change takes to get from commit to production. Failed deployment recovery time is how quickly you recover when a change breaks something; that is the measure that used to be called time to restore service, or mean time to recover, until DORA renamed and redefined it in 2023. Change fail rate is the proportion of deployments that cause a failure needing immediate attention. Deployment rework rate, added in 2024, is the ratio of deployments that are unplanned and happen because of an incident in production; DORA introduced it because change fail rate on its own had been standing in for how much rework a team was doing. Both dates and the reasoning are DORA’s own, set out in their history of the metrics.

When I first wrote this post I claimed these measures could not be gamed. That was wrong, and I have written the opposite elsewhere, in the companion post on reporting to a board. Any measure can be gamed. Deployment frequency rewards empty deployments. A low change fail rate rewards not looking too hard at what broke.

What is true is narrower and still useful: these five are harder to inflate than velocity, because they describe things that either happened in production or did not, and because they pull against each other. Deploying more often is easy until the failure rate follows it upward. That tension is the point, and it is why you watch all five as a set.

The Sprint Waterfall Pattern

What you see from inside the team is a sprint that completes cleanly, every sprint, with a velocity you could set your watch by.

What the measures show is a deployment frequency of once per sprint at best, and a change lead time somewhere between two and six weeks. The work is being batched into sprint-sized releases, so finished code sits waiting to be deployed, accumulating risk and delaying feedback. Every day a completed change waits is a day the team cannot learn whether it was the right change, and a day the eventual release grows larger and harder to reason about.

The fix is to stop coupling deployment to the sprint boundary. A sprint is a planning rhythm and has no reason to be a release rhythm as well; you can plan in two-week blocks and still deploy several times a day. Teams often discover that the only real obstacle is a manual step somebody added years ago for a reason nobody can now reconstruct.

The Fear of Production Pattern

This one looks like diligence. There is a testing phase, a release checklist, a change advisory board, and a deployment window at a quiet time on a Thursday. Everybody is careful, and everybody can point at the process they followed.

The measures tell a different story. The care is not working. The change fail rate stays stubbornly where it was, and failed deployment recovery time is counted in hours or days. The elaborate process has not made failure less likely; it has made it rarer and larger, and it has made recovery slower, because nobody deploys often enough to be practised at it.

Early in my career I watched a deployment take a system back about six months. Features that had shipped were suddenly missing, pages were rendering in a design that had been replaced, and nobody could account for any of it. We worked through the last three deployments, oldest first, looking for the one that had done the damage. It took an hour of reading logs, commit messages and pull requests to find the real answer: somebody had pushed straight to the main branch, which was called master at the time, and that push had gone round the pull request, QA, staging and live workflow entirely. They had been working from a copy of the branch that was six months old, had not fully understood the warning their development environment gave them, and had taken the option offered to force the push through.

It was an easy mistake to make, particularly in a company where admitting uncertainty was not especially welcome, and it was straightforward to undo: re-run the previous deployment, with a little git surgery afterwards. The harder thing for the team to accept was that this was not one developer’s failure. The main branch was not protected, and the deployment process had gaps wide enough for a six-month-old codebase to go through unnoticed. The workflow everybody trusted was an agreement, not a control.

The fix is the opposite of what the instinct suggests: deploy more often, in smaller pieces, with the ability to roll back quickly and to turn a change off without a release. Feature flags, automated tests that run on every commit, and a rollback you have rehearsed will do more for reliability than a review board that meets weekly.

The rehearsing matters more than it sounds, because the fear lives in people rather than in the pipeline. A junior developer I paired with had picked up a ticket the team had sized at eight points, against a backlog that mostly ran three to five, and it changed the behaviour of several application programming interfaces (APIs). By the time we finished the work they were genuinely anxious about the size of it. What if QA sends it back? What if it gets through and breaks something that matters?

Most of what fixed that was ordinary. I took them to meet the QA engineer assigned to the ticket, so that the person who would review their work stopped being an abstraction with a veto. The change bounced back twice, and both times we read it as the requirements getting clearer rather than as a fight between two teams. We built a local environment as close to live as we could manage, deployed into it, and spent an hour using the feature until we had found a couple of problems and fixed them.

When the deployment went out the following week, we watched it together, talking through the backout steps and how we would restore the database if we needed to. We were so deep in that conversation over coffee that we both missed the deployment completing. They tested the feature first, once I had the rollback ready, noting down the test data we would need to remove afterwards. Nothing went wrong, and the relief was visible.

What I wanted them to take from it was that a deployment is something you can ask for help with. Systems are not infallible, and that was never the claim; the claim is that a good pipeline lets you undo a mistake in minutes rather than in an afternoon of archaeology. They joined the next four deployments, and went on to specialise in DevOps and site reliability.

The Fake Definition of Done Pattern

Here the board is full of completed stories and the velocity chart is the healthiest it has ever been. Ask when customers will see any of it and the answer involves a date some weeks away.

In the measures, that shows up as a lead time dwarfing the time anybody spent writing the code, and a deployment frequency that bears no relationship to how much work is being finished. The definition of done stops at “merged”, so everything after that point is invisible to the process that claims to be tracking the work.

The fix is to move the finish line to production and accept what that does to your numbers in the first month. Velocity will look worse, because work that used to count as done now waits for deployment before it counts at all. That drop is the first truthful measurement the team has had.

Getting a Baseline Without a Platform

You do not need to buy anything to start. The two that matter most at the beginning, deployment frequency and change lead time, can be recorded by the pipeline you already run.

Record every deployment with its commit, then calculate lead time against the oldest commit in that deployment rather than the newest:

- name: Record deployment
  run: |
    echo "$(date -u +%FT%TZ),$(git rev-parse HEAD),production" >> deployments.csv

- name: Calculate lead time
  run: |
    PREVIOUS=$(tail -n 2 deployments.csv | head -n 1 | cut -d, -f2)
    OLDEST=$(git log --format=%ct "${PREVIOUS}..HEAD" | sort -n | head -n 1)
    if [ -z "$OLDEST" ]; then echo "Nothing new in this deployment"; exit 0; fi
    echo "Lead time: $(( ( $(date +%s) - OLDEST ) / 3600 )) hours"

Three things about that snippet are worth saying before you copy it. The file has to outlive the run: a workspace on a hosted runner is destroyed when the job ends, so commit deployments.csv back to the repository, keep it as an artifact between runs, or read the previous deployment from your forge’s deployments API instead. The recording step has to run before the calculation, or every deployment measures itself and reports nothing. And the first recorded deployment has nothing behind it to measure against, which is what the empty check is for.

The original version of this snippet timed from the most recent commit, which reports minutes for work that took a fortnight. Taking the oldest commit since the previous deployment gets much closer to the truth, sorting the timestamps rather than trusting git’s ordering, and the checkout step needs full history for any of it to work. One honest caveat remains: if you squash on merge, or rebase before merging, the commit timestamps on your main branch are the timestamps of the rewrite rather than of the work. Where that is how your team merges, take the time the pull request was opened from your forge’s API instead, and treat the git-only version as a lower bound.

Then resist the urge to compare yourself to anybody. DORA does publish performance clusters, but they are redrawn from report to report, and a number quoted without its report year tells you very little. The comparison that matters is your own baseline from last quarter, because the question worth answering is whether this team is getting faster at delivering to these customers.

Change One Thing, Then Measure Again

Once you have a baseline, pick the single measure that hurts most and leave the others alone. Teams that try to improve all five at once end up unable to say which change did what, which is how organisations acquire processes that everybody follows and nobody can justify.

Lead time is usually the place to start, because queues are what show up in it. Look for where work waits: for review, for a test environment, for an approval, for a release window. Shrinking the work helps too, since a change that takes two days to write cannot spend three weeks in review without somebody noticing.

Two objections come up whenever I suggest this, and both deserve an answer. The first is that sprints are needed for planning, which is true and not in conflict with any of it; plan in sprints, deploy continuously, and let the two rhythms run at different speeds. The second is that compliance requires controlled releases, which is also true, and is an argument for automating the controls rather than for performing them by hand. An automated check that runs on every deployment and leaves an audit trail is more defensible to an auditor than a signature collected at the end of a long week, which is the argument I made at more length about attestations and the link between your code and production.

What It Looks Like When It Works

Improvement rarely arrives as every number moving at once. What usually happens is that one constraint gives way and another becomes visible, which is the clearest sign the measurement is working.

I saw exactly that during a twelve-month engagement with Ligentia, a globally distributed logistics business, where feature delivery improved by 75% on measurements the engineering organisation was already running.

The interesting part came afterwards. The team began producing pull requests faster than their reviewer rotation could clear them, so the bottleneck moved from writing code to reviewing it. That is the kind of problem worth having, and they could see it happening because delivery was already what they measured.

Where those numbers go next is a separate skill. I have written about showing this kind of measurement to a board without either overselling it or drowning the room, which is a different job from using the numbers inside the team.

Ceremonies Are Cheap, Delivery Is Not

Finding out that your Agile process is mostly ceremonial is uncomfortable, and it is also the most useful thing you can learn about it. The measures do not care how well your standups run. They describe how long it takes an idea to become something a customer can use, and how often the attempt goes wrong.

Start by measuring two of them honestly for a month. Fix the queue that the numbers point at. Then measure again, and let the process follow the evidence instead of the training material.

If your ceremonies are in good order and delivery still hurts, let’s talk; working out which queue is the expensive one is usually a short piece of work with a long payoff.

More from writing

Next step

Bring me the decision you keep deferring

A discovery call costs nothing and commits you to nothing. You'll leave with an honest read on your situation and a clear next step, whether or not that step involves me.

Book a discovery call