Skip to content
RJJ Software Jamie Taylor · fractional CTOBook a call
Menu

Writing · 11 September 2026

Stop Screening for AI Tools: Hire for Judgement Instead

Engineers are being turned down at interview for using the wrong AI coding tool. The adoption data says that hiring filter expires in months.

A close-up of a printed page with the word "Steps:" ringed in yellow highlighter, the surrounding text out of focus.

Over the last few weeks I have heard the same story from several engineers, and it has stayed with me because of how new it is. They were turned down at interview because of which artificial intelligence (AI) coding tool they use. Not because they could not do the work. Because the hiring team had settled on one tool and the candidate had settled on another.

I wrote a short post about this on LinkedIn, half expecting to be told I had misread a couple of unlucky rejections. What came back was more interesting than that. A recruitment consultant who places .NET engineers said they had never once seen it in a client’s job specification, and that if they did they would challenge it hard. A VP of Engineering said it was part of what they called the ridiculousness of the current market, where employers want an exact match on every named technology and transferable skills seem to have gone out of fashion. Between those two responses sits the thing that makes this hard to write about: it is happening in the room, at interview, and it leaves no trace anywhere that anybody measures.

So I want to make the argument carefully, and I want to lead with the conclusion rather than make you wait for it. The problem with screening candidates on their AI tooling is not, in the first instance, that it is unfair. It is that the filter has a shelf life of months. The tool you are screening for today is quite likely not the tool your team will be using by the time that hire clears probation, and the published adoption data on this is genuinely startling.

The Filter Nobody Wrote Down

Start with what we can actually measure, because it is not this.

Indeed’s Hiring Lab ran roughly six hundred AI-related keywords across a year of United States job postings. Among postings that mention AI at all, the generic term “AI” appears in nearly 74% of them. ChatGPT, the most recognisable AI product on the planet, appears in 2%. LinkedIn’s rising-skills report for 2026 names Google Gemini, the OpenAI application programming interface (API), LangChain and PyTorch among the skills growing fastest, and does not name a single coding agent. Lightcast maintains a taxonomy of more than three hundred named AI skills, and tracks products like Microsoft Copilot and Hugging Face by name in other sectors, yet lists no coding harness at all.

On the face of it, that looks like a flat contradiction of everything I just told you. Employers are not writing these tool names into job adverts in any meaningful volume.

Except that is not the claim, and the recruitment consultant is the reason I know it. They have never seen it in a specification. It surfaces at interview, in the conversation, as feedback. And there is no dataset anywhere for what gets said at the interview stage. Job postings are the only part of hiring that leaves a public, countable trace, which is precisely why every study you can find measures them.

There is a sharper version of this problem, and it comes from the Burning Glass Institute and Harvard Business School. When employers publicly dropped degree requirements from their job adverts, the change translated into roughly ninety-seven thousand workers out of seventy-seven million annual hires. Fewer than one in seven hundred. What a job advert says turns out to be a weak predictor of what the screening actually does. That finding cuts in both directions: it means advert text would have been poor evidence for me, and it means the absence of these tool names from advert text is poor evidence against me.

So I cannot tell you how widespread this is, and I am not going to pretend otherwise. What I have is a handful of first-hand accounts, a thread of people recognising the pattern, and a recruiter who says they would push back on it. That is a signal, not a survey, and it comes from my own network, which leans British and leans .NET and self-selects for people who already agreed with me. Treat it as a warning about something starting rather than a report on something established.

The rest of this post is about why you should not do it, and that argument does not depend on how many people currently are.

Seven Months Was All It Took

The case against a tooling filter does not need any of that context, though. It needs one comparison.

JetBrains surveyed more than ten thousand professional developers in January 2026, and then more than fifteen thousand between May and July, publishing the second set this August. Set the two waves side by side and the entire ranking has reordered.

ToolJanuary 2026May to July 2026
Claude Code18%39%
GitHub Copilot29%21%
Codex3%16%
Cursor18%12%

Codex went up more than fivefold. Cursor lost a third of its share. Copilot, the incumbent, fell from first place to second. In seven months.

Two honest caveats before I lean on that. These are two different survey instruments rather than one tracked panel, and JetBrains do not publish the exact question wording behind “adoption” in either write-up. They narrate the two waves as a series themselves, so the comparison is fair, but it is a comparison of two photographs rather than a continuous film.

Now apply it to hiring. Recruiting an engineer takes weeks; Ashby put the average time to first fill for technical roles at around ten weeks. Add notice periods and onboarding and you are six months from decision to a fully productive hire, comfortably. On the evidence above, six months is long enough for your team’s tool of choice to have changed underneath you. You would be rejecting a candidate for failing to match a snapshot that expires before they finish settling in.

JetBrains draw their own conclusion from the churn, and it is worth sitting with: product excellence now outweighs ecosystem lock-in, and developers will migrate to whatever component actually delivers. That is not a market where tool allegiance is a durable property of a person. It is a market where everybody is moving, all the time.

Why Sensible People Do This Anyway

I want to be fair to the person on the other side of the table, because I do not think this is stupidity or malice, and the comment thread’s instinct to call hiring managers incompetent is the least useful reading available.

There is a well-documented economics of exactly this behaviour. Modestino, Shoag and Ballance published a study in the Review of Economics and Statistics showing that employers opportunistically raise their requirements when job seekers are plentiful. They found that the increase in unemployed workers during the Great Recession accounted for between 18% and 25% of the rise in stated skill requirements between 2007 and 2010. When candidates are abundant, requirements inflate, and they inflate without anybody deciding to be unreasonable.

Look at the current market against that finding. Greenhouse, drawing on more than six hundred and forty million applications, put applications per job at 116 in 2022 and 244 in 2025. Over the same period the number of recruiters per organisation more than halved. Ashby, on a separate hundred-million-application dataset, has the average open role going from around a hundred applications in 2021 to more than three hundred now. More candidates, fewer people to assess them, and a longer time to fill. A filter that removes a chunk of the pile at no immediate cost is exactly what that pressure produces.

There is a cost, of course; it just lands somewhere you cannot see it. Harvard Business School and Accenture surveyed employers and found that 88% agreed that qualified, high-skill candidates get vetted out of their process because they do not match the exact criteria in the job description. Employers know this is happening. They know, and it continues, because the false positive costs you a bad hire and a visible problem, while the false negative costs you somebody you will never know you missed.

There is a second mechanism at work, though, and it may be the better explanation of why this particular attribute became the filter. Requirement inflation tells you that employers raise the bar when candidates are plentiful; it does not tell you why the bar ended up here, on the name of an editor. A developer who commented on my post offered a reading I have not been able to shake: a lot of people have stopped making pragmatic technology decisions and started picking a football team, staying loyal to their side regardless of whether they win or lose. If that is what is happening, then turning down a Claude Code user because the team runs Codex is not a capability judgement at all. It is a tribal one wearing the costume of a technical requirement.

I think that sits awkwardly next to the migration figures I just showed you, and I think both are true at once. The developers who post about their tooling are not the same population as the fifteen thousand who quietly changed tools between one survey and the next. Tribalism is loud, and the aggregate is mobile. If anything the two together make the case sharper: when your team’s sense of identity is bound up in a product that shed a third of its share in seven months, the lesson is to hold it more loosely, not to screen candidates on it.

I should be straight about one thing here. This is a steelman I have constructed, not one I was given. Nobody who has actually done this kind of rejecting has explained their reasoning to me. Every voice in that thread was rejected, sympathetic or appalled, and a hiring manager reading this would be entitled to say I have not represented them.

I will also give the strongest fact I found on their side. Ashby’s data shows offer conversion rates now running above 2021 levels despite the tripled volume, which means teams are filtering harder and, by that measure, filtering better. Aggressive screening is not self-evidently broken. That is the case I have to answer.

What the Filter Is Actually Measuring

Here is my answer, and it is a measurement problem before it is a fairness problem.

What the team actually wants to know is whether this person can work effectively with agents: whether they can specify work clearly, spot the moment the agent has gone confidently wrong, and take responsibility for output they did not type. None of that is visible on a curriculum vitae. So the process reaches for something that is visible, and the name of a tool is about as visible as it gets. It is a single word, it either matches or it does not, and checking takes no time at all.

Then the proxy quietly becomes the target, and the thing it was standing in for drops out of the conversation entirely.

Readers of my post on tokenmaxxing will recognise the shape of this. That was about treating visible token consumption as though it were visible value, and it is the same error wearing different clothes: reach for the thing you can see, mistake it for the thing you care about, and end up rewarding the wrong property. Goodhart’s Law does not care whether the measure is a story point, a token count, or the name of an editor. Once a measure becomes a target, it stops being a good measure.

The particular tragedy of this one is that the underlying question is a good question. Teams are right to want engineers who work well with agents. They have simply picked the one signal that carries almost none of that information.

The Honest Case Against Me

If skills genuinely did not transfer between these tools, my argument would collapse, so let me put the strongest evidence for that position rather than hope you do not go looking for it.

METR ran a randomised controlled trial with sixteen experienced open-source developers across two hundred and forty-six real issues. The developers using AI were 19% slower, while believing themselves to be around 20% faster. It is the most rigorous study in this field and its headline finding is not the one anybody expected. For our purposes the detail that matters is buried in the discussion: those developers had only a few dozen hours of experience with the tool, and METR say explicitly that they cannot rule out learning effects beyond fifty hours of use. If competence with a specific harness only saturates after fifty-odd hours, then switching genuinely does cost something real.

There is more. DORA’s work on the return on investment of AI-assisted development describes a J-curve, a measurable dip in productivity before the gains arrive, and names learning curves as one of its causes. A longitudinal study presented at the International Conference on Software Engineering this year tracked eight hundred developers across two years and a hundred and fifty-one million editor window activations, and found that developers using AI showed a rising trend in context switching that non-users did not, while 74% of them had not noticed it happening. Switching costs can be both real and invisible to the person paying them.

And people do love their tools. JetBrains report Claude Code at 91% customer satisfaction with a net promoter score of 54, which they describe as the strongest product loyalty in the market. Pretending harness choice is a matter of indifference would be untrue to how developers actually feel about it.

Now the part nobody can tell you: there is no published measurement of how long it takes a developer to become productive after switching AI coding tools. I went looking specifically and it does not exist. Practitioner blog posts describe adjusting in days to about a week, but those are written by people who liked the switch enough to blog about it, which is not evidence.

So how do I answer METR’s fifty hours? With revealed preference. Look again at that table. Copilot shed eight points of share and Cursor shed six while Codex gained thirteen, across seven months, among tens of thousands of working developers. That is a mass migration, and it is not the behaviour of people facing an expensive, painful switch. If moving between harnesses genuinely cost weeks of competence, that churn would not happen at that speed. Developers voted, in enormous numbers, with their time.

What Actually Transfers

The portability is not accidental, either. It is being built on purpose.

AGENTS.md, the convention for a single instruction file that an agent reads when it enters your repository, is now read by more than twenty-five different agents and appears in over sixty thousand open-source repositories. It is stewarded by the Agentic AI Foundation under the Linux Foundation. When I wrote about the foundations of agentic development I made the point that the filename differs between tools and the principle does not, and that is more true now than when I wrote it. One caveat for accuracy: Claude Code does not read AGENTS.md natively, and the documented workaround is an import on the first line of your CLAUDE.md or a symbolic link. A one-line fix is not a skills gap.

There is early empirical support too. A study this August across four hundred and forty-one corporate repositories in twenty-seven organisations, covering twelve different AI coding tools, found that repositories without a committed AI configuration file saw cognitive complexity rise by 53% after adopting agents, against 27% for those with one. The maturity distribution differed between the major tools, but all of them showed the same progression. The practice generalises even where the products do not.

What genuinely differs between these tools is ergonomics, and I would rather name it precisely than wave it away. Cursor owns the inline surface in a way the others do not; its Tab completion is its own model and its own product. Claude Code has no autocomplete offering at all, which is a real asymmetry rather than a rounding error. Copilot’s cloud agent hands you a branch and a pull request rather than a diff in your editor, which is a different unit of work. Those are honest differences.

But the extension primitives, the parts that took me longest to learn, have converged into a checklist that all of them tick: Model Context Protocol (MCP) support, hooks, subagents, skills. And what I actually spent that time learning was not any of those interfaces. It was how to write a specification an agent will not misread, when to stop it, and how to review something I did not type. I made this point in a case study on teaching AI-assisted development and I will make it again: technology choices matter less than approach.

Cursor’s Head of Talent Has a Name for This

I went looking for evidence about how the companies building these tools hire, half expecting to find them screening on their own products. What I found was that they are not discussing tooling at all. They are discussing the filtering.

Adam Ward is head of talent at Cursor. On Lenny’s Podcast this year he described a failure mode he calls the funnel of doom, and it is the thing this entire post has been circling. You reach out to a hundred people, a fifth of them reply, and you treat that fifth as the top of the market.

By definition, that is not the top 20%. That is just the 20 people who you caught on a bad day.

Sit with that, because it is a better version of my argument than the one I have been making. Replying to a recruiter is not a measure of quality. It is a measure of who was having a rough week. The filter selects for receptiveness, gets mistaken for one that selects for calibre, and the people it misses are not worse; they are content.

Then Ward describes what most of the industry does with whoever is left.

The way we all recruit is, we kind of weed people out at each stage and we hire the remainder.

Weed out, and hire the remainder. That is a process built to produce a survivor rather than to find a person, and every filter you add raises the odds that the survivor is whoever matched on the most surface criteria rather than whoever was best. A harness requirement is about as surface as a criterion gets. Ward’s own corrective is worth stealing: you are hiring one person, not ten, so stop letting the numbers game crowd out the judgement.

What to Interview For Instead

None of this means you should stop caring whether somebody can work with agents. It means asking about that directly rather than inferring it from a brand name. Five questions do most of the work for me, and every one of them survives a change of tooling.

The first is to ask about a time an agent confidently handed them something wrong, and how they caught it. That is the whole game. I wrote in the AI amplification paradox that these tools amplify whoever is using them, which is excellent news for a strong engineer and a serious problem otherwise; noticing that the plausible answer is the wrong answer is the skill that separates the two.

Then I want to hear how they specify a piece of work before handing it over, which is the habit underneath every agentic workflow I have written about. An agent does not fill the gaps with seasoned judgement, it fills them with confident guesses, quickly, which makes the quality of the specification most of the quality of the output.

The question I have come to like most is what goes into their instructions file, and what they deliberately left out. The second half is the interesting half, because it shows whether they have understood that every line in that file is costing them context they could have spent elsewhere, the same discipline that goes into hardening an agentic setup.

After that, when do they decide not to reach for an agent at all? Anybody who cannot answer that has enthusiasm for the tool rather than judgement about it.

Last, how do they review code they did not write? That is increasingly most of the job, and it is a discipline people have either built or avoided building.

Ask those five and you will learn more in twenty minutes than any tool-matching exercise will tell you. You will also be able to assess somebody who has never opened the harness your team runs on, which, given the numbers earlier, is a capability worth having.

If You Are on the Other Side of This

A shorter word for anyone who has been on the receiving end.

One commenter on that LinkedIn thread argued that developers should stop selling the tools they use and start selling their outcomes. Another pushed back, saying the problem is the hiring manager and not the developer. The answer that came back has stayed with me: developers do not control hiring managers, but they are in complete control of themselves.

I think that is right, and it is the most useful thing anybody said. You cannot fix somebody else’s screening process from inside a forty-five minute interview. You can decide what you put in front of them. A CV that lists tools invites a tool comparison. A CV that describes what you shipped, what you decided, and what you caught before it reached production invites a different conversation, and it is a conversation you will win.

That said, another commenter was honest about something worth repeating: it stings. Losing work to somebody else’s misunderstanding takes a chunk out of you, and being told the reason was your choice of editor does not make it land any softer. That part is not your failing, and you should not internalise it as one.

The Tool Is Not the Skill

Every generation of this industry has produced a version of this mistake. Turned down for a C role because your experience was C++. Turned down for Java because you wrote C#, and then turned down for C# because you wrote Java. One of the engineers who replied to my post rattled off that exact list from their own career, going back to Modula-2 and Pascal, and they did not need any prompting from me to do it. In hindsight every one of those looks absurd, and the absurdity was just as available at the time to anybody willing to look.

We are about to add AI harnesses to that list. The difference is that the churn is faster than anything on it, which means the filter will be embarrassing sooner. A hiring manager who screened for Cursor in January was screening for a tool that had shed a third of its share by July.

Hire the engineer. The tooling is a fortnight of onboarding, and the judgement is the thing you actually cannot teach in a fortnight.

If your team is working out how to adopt agentic development, or how to assess it properly when you are hiring, that is a large part of what I do. I help engineering teams build genuine capability with these tools rather than a collection of licences, and I help the people doing the hiring work out what to look for. If that is a conversation worth having, get in touch.

More from writing

Next step

Bring me the decision you keep deferring

A discovery call costs nothing and commits you to nothing. You'll leave with an honest read on your situation and a clear next step, whether or not that step involves me.

Book a discovery call