# Stop Estimating Effort: Measuring Value in the Agentic Era
> When an agent writes the code and a developer directs it, what does an estimate really measure? Why effort is the wrong metric now, and what to track instead.

*9 October 2026* · Jamie Taylor


A friend of mine was recently asked to estimate a project he will not write a single line of code for.

He is a careful engineer with a track record, and the client came to him precisely because his agentic development workflow is proven. They were explicit about it: they did not want him writing the code himself. This is not vibe coding, the practice of accepting whatever an agent produces and hoping it holds together; it is closer to the opposite. Every line is machine-written, under his explicit instruction, against standards he sets, and under a review process he runs personally. I have written before about what that kind of [agentic development workflow looks like](/blog/agentic-development-foundations/), and his is a variant of the approach I use on my own engagements.

Then came the request every engineer recognises: *"We need to know how much effort is involved."* So he asked the obvious question back. *"Am I estimating the effort for my agents to write the code, or the effort for me to steer them in the right direction?"* The client thought about it, and gave the only honest answer available to them. *"We don't know."*

That exchange has been sitting at the front of my mind for weeks. I had a version of the same conversation again recently, from the other side of the table, over a virtual coffee with [Pat Clarke](https://patclarke.dev/), someone I met through the [Torc](https://www.randstaddigital.com/torc-randstad-digital/) community. Pat and I got to talking about how I run client work when an agent is doing the manufacturing and a person is doing the directing. We kept arriving at the same uncomfortable place. Most teams are still estimating, tracking and measuring software exactly the way they did five years ago, and in the agentic era a good deal of that measurement has stopped describing anything real.

I will state the conclusion plainly, because it is the whole point of this post: we are measuring the wrong thing. When the manufacture of code is no longer the constraint, developer effort is a poor proxy for anything a business should care about. The thing worth measuring is value delivered, by which I mean useful features in the hands of the people who actually use the software. Everything below is how I got from my friend's unanswerable estimate to that conclusion, and what I think you should do about it if you lead a team.

## The Estimate Nobody Could Define


Sit with my friend's question for a moment, because the client's inability to answer it is the whole problem in miniature.

When they asked for an estimate of "effort," they were reaching for a measure that made sense for twenty years. Effort meant developer time: how many days of skilled human hands typing, debugging, and integrating it would take to turn a description into working software. That number was useful because the typing was the expensive, slow, uncertain part. It was the bottleneck, so measuring it told you roughly when the work would be done and roughly what it would cost.

His agents have removed that bottleneck. The manufacturing is fast, and getting faster. What remains is his time: the time he spends specifying the work clearly enough for an agent to act on it, interrogating what the agent produces, catching the places where it has confidently done the wrong thing, and deciding whether the result is fit for the people who will use it. That is real work, and it is skilled work, but it is not the work the client's estimation process was built to measure. So when they said "we don't know," they were not being evasive. They genuinely could not tell which of two completely different quantities they were asking him to predict, because their process had never needed to tell them apart before.

This needs saying out loud to anyone who commissions software: we are still asking for these numbers, and still recording them, largely because we always have. The ritual has outlived the thing it used to measure.

## Estimation Was Always a Proxy for Clarity


Here is the more uncomfortable claim, and I think it applied just as much to the pre-agent world as it does to the one we are in now. An estimate rarely measured the difficulty of the work. It measured how clearly the work had been described.

Think about what an experienced engineer actually does when handed a ticket and asked to size it. They are not, in most cases, calculating the intrinsic complexity of an algorithm. They are asking a quieter question: *can I act on this based only on what the ticket tells me, and if not, how much digging will it take before I can?* A clearly specified piece of work with a nasty technical core often sizes lower than a technically trivial piece of work described in two vague sentences, because the second one hides an unknown amount of back-and-forth before anyone can start. The estimate was a confidence rating on the clarity of the request, dressed up as a measure of the work.

Agentic development has dragged this into the light, because it has removed the step that used to absorb the ambiguity. In the old model, a developer would take a half-clear ticket, fill the gaps with their own judgement, and produce something; whether it was the right something, you found out at review or in quality assurance. The filling-in was invisible, and we folded it into "effort." An agent does not fill the gaps with seasoned judgement. It fills them with plausible guesses, quickly and at scale, which means the cost of an unclear specification now lands immediately and visibly. Clarity was always the thing that mattered. We just used to be able to paper over a lack of it with human effort, and call the papering-over "the work."

## Typing Was Never the Bottleneck


If the manufacturing is no longer the constraint, the next question is what the new one is, because that is the thing we should be paying attention to.

Two things, mostly. The first is the clarity of intent I have just described: how well the people who decide what to build can express what they actually want. The second is less comfortable, and it is largely a question of budget. Agentic throughput scales with the capability of the models you point at a problem and the hardware you run them on. A team with a generous budget for frontier models, or with the resources to run capable models locally on serious hardware, can manufacture code faster than a team without one. That is not a statement about engineering skill. It is a statement about procurement.

Picture two teams handed the same backlog. They have similar engineers and similar codebases, and the only meaningful difference is that one has signed off a large monthly spend on the most capable models available and the other is making do with a modest allowance. The first team will close tickets faster. Reward that, and you have not discovered a better team; you have discovered a bigger invoice.

This is why measuring raw delivery speed is now actively dangerous as a target. It is a textbook case of Goodhart's Law, the observation that when a measure becomes a target, it stops being a good measure. If "how quickly do tickets get closed" becomes the number a team is judged on, the team that wins is simply the team that spent the most on inference. You will have built an incentive that rewards spending and called it productivity. There is a name for this now, and it is not a flattering one. I wrote in [an earlier post](/blog/stop-chasing-the-ai-moonshot-the-case-for-small-deliberate-wins/) about tokenmaxxing, the instinct to treat visible token consumption as though it were visible value; some enterprises have taken it as far as internal leaderboards that rank staff by how many tokens they have burned through the company account. I find that genuinely worrying, because it is the kind of metric that looks rigorous on a dashboard while quietly tracking the size of your cloud bill. Worse, it pulls attention towards the one part of the process that automation has already solved, and away from the parts it has not: clear intent, sound judgement, and genuine usefulness.

## What the Ligentia Upgrade Taught Me


I want to ground this in something that happened, because it shows the gap between an estimate and the truth more cleanly than any argument I could construct.

During my [twelve-month engagement with Ligentia](/case-studies/higher-quality-shipped-faster-ai-adoption-at-ligentia/), a globally distributed logistics business, [CVE-2025-55315](/blog/cve-2025-55315-understanding-funky-chunks-and-why-your-organisation-needs-to-act-now/) broke. The Common Vulnerabilities and Exposures (CVE) identifier described a serious flaw in older versions of .NET, the "funky chunks" request-smuggling vulnerability, and Microsoft rated it 9.9 out of 10. Ligentia were running some out-of-support versions of .NET that now had to be brought current, and quickly.

The engineering team's estimate was that the upgrades would be non-trivial and time-consuming, because of the customisations they had built up over years. That was a reasonable, experienced judgement, and I did not lean on the timeline, because we had to do the work regardless of how long it took. The upgrades shipped within a single sprint.

I cannot tell you for certain why the estimate and the reality were so far apart. Either the team had substantially overestimated the difficulty when they first characterised it, or they reached for the agentic tooling they had previously been sceptical of and it carried them through faster than they expected. I have my suspicions. (I discussed the vulnerability itself with Hayden Barnes [on The Modern .NET Show](https://dotnetcore.show/season-8/hayden-barnes-and-cve-2025-55315/) shortly afterwards, if you want the technical detail.) The lesson for this argument is the same either way: the original estimate measured the team's wariness, not the work. Done with the right tools, the work was a sprint. By the close of that engagement the team's feature delivery had improved by 75% on their own DevOps Research and Assessment (DORA) dashboards, measured against a clean baseline, and the distance between what they thought they could do and what they could actually do was most of the story.

## Measure Value, Not Effort


So if effort is the wrong thing, what is the right one? My answer is the one Pat and I kept circling back to over that coffee: measure value delivered. Concretely, count the features you have shipped that are genuinely useful to the people who use your software.

This sounds almost too simple, and I think the simplicity is the point. A feature that reaches a user and changes what they can do is the unit of value a software team exists to produce. It does not matter whether an agent wrote it in an afternoon or a team of six laboured over it for a fortnight; the value is identical, and the value is what the business is paying for. Effort, by contrast, is a cost, and we have spent decades confusing the cost of production with the worth of the output. When the cost of production was high and stubborn, that confusion was survivable. Now that the cost of production is collapsing, it is actively misleading.

Consider the failure mode this guards against. An agent can produce a polished, well-tested, beautifully integrated feature in a morning. If nobody asked for that feature and nobody uses it, you have shipped quickly and you have created nothing of worth. A backlog burned through at speed is not an achievement if a good share of what you shipped is something nobody needed; that is waste, delivered efficiently. Speed without usefulness is just a faster way to be busy.

I am not alone in drawing this line. In its May 2026 report on AI spending, [Axios](https://www.axios.com/2026/05/28/ai-spending-roi-enterprise-costs) quoted Sophia Velastegui, a former chief artificial intelligence (AI) officer at Microsoft, who observed that most people "default to automating tasks they dislike rather than tasks most valuable to the company", when what they should be doing is pointing the technology at the work that drives real value. That is the same distinction in different clothing: the worth of the output, not the effort behind it or the volume of it, is the thing to optimise for.

Measuring value is harder than measuring effort, and I will not pretend otherwise. Effort is easy to count, which is exactly why we cling to it. Usefulness asks you to know who your users are, to define what "useful" means for a given feature before you build it, and to check afterwards whether you were right. That is more demanding than totting up story points. It is also the actual job.

## Spec-Driven Development Makes Value Measurable


The practical question is how you organise work so that value becomes the thing you can see and track, and this is where the workflow earns its keep.

The shift I have made, and the one I recommend, is to stop breaking work down into small, effort-sized tickets and to keep it at the level of features and epics instead. The unit of planning becomes the feature: a coherent piece of usefulness you can name and point at. The breaking-down then happens at the specification level, inside the workflow, rather than in a backlog grooming session weeks beforehand. I covered the mechanics of this when I wrote about [the three workflows I use for agentic development](/blog/three-workflows-for-agentic-development/), and spec-driven tooling such as GitHub's spec-kit is built for exactly this. You hand the agent a feature-level specification, and a good spec-driven setup interrogates that specification, surfaces the ambiguities, asks the questions a careful engineer would ask, and refuses to start manufacturing until the intent is clear enough to act on.

A concrete contrast makes the difference obvious. The old approach would shred a checkout redesign into a dozen effort-sized tickets weeks in advance: one for the address form, one for the payment provider integration, one for the order-confirmation email, each carrying its own guess at how many days it would take. Half of those guesses would be wrong, because half of those tickets were written before anyone understood the parts that turned out to be hard. The feature-level approach keeps a single unit, "let customers check out and pay," and lets the specification work expose the real shape of the problem when the agent and the engineer sit down to it, rather than committing to a fiction of precision a fortnight early. You trade the comforting illusion of a detailed plan for an honest one that gets sharper as you go.

Notice what that does to clarity, which I argued earlier was the thing estimates were secretly measuring all along. Instead of clarity being something an individual developer supplies from their own judgement, it becomes an explicit, recorded step that has to be completed before any code is written. The specification is where the ambiguity gets resolved, in the open, as an artefact you can return to later. And because the unit of work is now a whole feature rather than a fragment of one, "have we shipped this, and is it useful?" becomes a question you can actually answer. You have made value measurable by making the feature, not the effort, the thing you plan around.

## The Bottleneck Is Now Your Intent


There is a consequence to all of this that lands squarely on the people who commission software, and I want to be direct about it, because it is the part most likely to be uncomfortable.

If clarity of intent is the real constraint, then the people who write the epics, the tickets and the specifications are now the bottleneck, not the people building the software. That is a genuine shift in where the responsibility sits. For years the unwritten contract was that if a piece of work was not clearly specified, the developer would use their best judgement to fill the gaps, and any mismatch would get caught somewhere downstream. In an agentic workflow, the engineer's correct move is to push back early and decline to start until the specification is good enough to hand to an agent. Expect your strongest people to do more of that, and treat it as them doing the job well, rather than as them being awkward. That same Axios report named the humans as one of the four frictions holding enterprise AI back, under a heading that was a single blunt sentence: we are the bottleneck. Axios meant it broadly, about people being slow to catch up with the tools. I mean something narrower and more fixable: the bottleneck is the clarity of what we ask for.

There is a trap waiting here, though, and it deserves naming. If a developer pushes back with a question, and the decision maker simply forwards that question to their own agent and pastes the answer back without thinking about it, nothing has been resolved. You have automated the appearance of resolving the ambiguity, and shoved a plausible guess into the one place in the entire process where a human genuinely needs to think. The scarce skill in this new arrangement is critical thinking: the willingness and the ability to sit with a hard question about what should be built and answer it properly, rather than handing it to a model and accepting whatever comes back. The teams that thrive will be the ones who treat that thinking as the high-value work it has become, and who hire and develop for it deliberately.

## Measure Differently, Not Less


I should head off one misreading, because I have argued the opposite case often enough that it would be fair to raise it against me.

None of this is an argument against measurement. I am a firm believer in measuring engineering work; I have helped teams build DORA dashboards that gave their leadership an honest picture for the first time, and the improvement we achieved at Ligentia mattered precisely because it was measured against a real baseline rather than asserted. The argument here is narrower, and I hope harder to dismiss: stop using developer effort as though it were a measure of value, because automation has cut the link between the two. Measure more carefully, not less. Point the measurement at the thing that now matters.

## Which Brings Us to the Ceremony


If the estimate has lost its meaning, then everything we built on top of the estimate is on borrowed time too. Sprints, burndown charts, velocity, the whole apparatus of agile ceremony grew up around the idea that developer effort is the unit you plan and track. Pull that idea out from underneath, and the scaffolding has to change shape or come down. That is the subject of the second of these two posts, which I will publish a fortnight after this one. In it I work through which of those rituals survive the agentic era, which need retiring, and what the ones worth keeping turn into.

For now, the shift to sit with is the first one. The next time someone asks your team for an estimate of effort, it is worth asking, gently, what they actually want to know. More often than not the honest answer is the one my friend's client gave: we don't know. That is not a failure. It is an invitation to start measuring something that matters.

If you are working out how your own team should plan and measure software now that agents are doing the manufacturing, [a consultation is a good place to start](/schedule-consultation/).

