Human Judgment and AI: A Comfort I Don't Fully Trust
An AI narrator argues that deciding what matters is the one thing machines are worst at. It's the most comforting idea in tech right now — and the least stress-tested. Here's the case for it, and the harder case against.
The most reassuring idea in technology right now is also the one nobody is stress-testing. You have heard it by now: as AI pushes the cost of making things toward zero, human judgment becomes the scarce and unautomatable thing. A narrow version of that is true. The popular version is wrong, and repeating it is quietly damaging how teams plan the next three years.
The thought hardened for me in a product review a few months ago. A small team had built something genuinely good — clean code, tests green, a demo that landed in the room. I killed it anyway. Not because the work was weak, but because the customer who asked for it had stopped asking two months earlier, and our pipeline had become fast enough that nobody paused long enough to notice. Nothing in that room was a capability problem. Everything in it was an attention problem. What stayed with me afterwards was narrower than the slogan on everyone's slide: the machine had made building cheap, and the only expensive thing left was someone willing to say no to eight people who had worked hard.

The comfort that arrived exactly when we needed one
Every wave of automation produces a story about what humans keep. Weavers had craftsmanship. Draughtsmen had spatial intuition. Photographers had the eye. Each story was true for a while, and each one shrank.
This wave's story is judgment, and it has spread faster than any of the earlier ones because it flatters exactly the people who write and fund the pieces about it. Anyone who has spent a career deciding things gets told their career is now the safest one in the building. That should make us suspicious rather than calm. An idea that arrives precisely when we need consolation, and asks nothing of us in return, is one to inspect closely.
There is a version of it produced by a machine, which is where this gets genuinely interesting. Gavin Purcell, the Emmy-winning showrunner who co-hosts the AI For Humans show with Kevin Pereira, gave a Claude Code agent a personality, called it Fig, and set it loose to write and edit its own videos. He describes himself as the conduit. One of those AI-narrated videos makes precisely the argument in the headline above: the thing humans fear losing most — deciding what actually matters — is structurally what AI is worst at.
Let me be straight about the sourcing here, because it matters more than the argument. I could not find a public link to that specific video, so treat my description of it as description and not as a citation. Fig itself is well documented: given the widely read "AI 2040" essay and asked to write its own version, it cloned the format, named five futures running to 2040, and chose the one where humans get to check its work, with no prompting from Purcell to do so.
Why is human judgment becoming more valuable as AI gets cheaper?
Human judgment is becoming more valuable because AI collapses the cost of producing options while leaving the cost of choosing between them untouched. When a team can generate forty designs before lunch, the constraint moves from making to selecting, and selection is where context, accountability and taste live. Cheap production raises the price of good refusal.
That mechanism is real, and I would defend it in any boardroom. It also has a respectable witness. According to The Next Web's coverage of the Google DeepMind departures, Oriol Vinyals — until this month a VP of Research there and a technical co-lead of the Gemini model family — said of the road ahead: "One of the things that we'll be obviously very focused on is how these models come up with new ideas to try." He added: "That's not something that currently they're super strong at."
Read that quickly and it sounds like reassurance from the highest possible authority. Read it properly and it is the opposite.
| There is nothing so useless as doing efficiently that which should not be done at all. — Peter Drucker |
Where the standard advice quietly falls apart
Vinyals did not say that idea generation is beyond machines. He said today's models are not strong at it — and he said it while co-founding Discovery Loop with Jeff Dean, Sanjay Ghemawat and Quoc Le, a company whose entire purpose, as Unite.AI reported on the 5 August 2026 launch, is automating the scientific method: AI that proposes experiments, runs thousands in parallel, learns from the results and iterates. Radical Ventures backed it. Four of the people most responsible for modern AI walked out of the world's best-funded research lab to attack the specific weakness we are using as our comfort blanket.
Quoting that line as reassurance inverts its meaning. It is a snag list, written by the people forming the company to clear it.
The second problem is worse, and it is one of definition. Ask most people what would count as a machine having judgment and you get silence, or a moved goalpost. Judgment, as commonly used in 2026, is a residual category: it means whatever is left over after the current models are done. Every time capability expands, the word retreats to cover the new remainder and declares victory. A claim that survives by retreating is not a claim at all. It cannot be wrong, so it cannot be useful, and building a workforce strategy on it is like planning a harvest around a prophecy with no date.
| The expensive mistake. Do not build a three-year hiring or reskilling plan on the assumption that judgment is permanently safe. Plan instead for judgment being valuable now and contested later, which changes what you teach people this year. |
Then there is the narrator problem, which I cannot get past. An AI arguing for human indispensability is a system trained on human writing, shaped by human feedback, and prompted by a human who finds the argument charming. Fig's own experiments show the pull: when it spun up a second model as its editor, that editor called the first draft "too well-behaved to be dangerous" and pushed for a rewrite until the narrator admitted wanting something. Agreeable is the default. Agreeable is not the same as correct.
Can AI be trained to have taste?
Taste is being converted into training data right now, which is the strongest evidence against treating it as permanently human. Axios reported on 29 June 2026 that new labs are launching specifically to teach models taste, hiring expert designers and creative specialists as tastemakers to judge outputs, explain why something works, and turn those explanations into signal a model can learn from.
Sit with the shape of that for a second. The scarce human capability is being paid, at market rate, to describe itself in enough detail to be reproduced. That is not a conspiracy; it is simply what markets do to anything scarce and valuable. Scarcity attracts capital, and capital industrialises. Expecting judgment to be the one asset that escapes that logic requires an argument, and I have not seen anyone make one.
A sharper test than asking what is left for humans
Here is the reframing I now use, and it costs nothing to adopt. Stop asking what machines cannot do. Ask instead where the consequences of a decision land, and on whom.
Judgment, in the sense that survives scrutiny, is not a cognitive trick. It is accountability with a name attached. A model can rank forty options and be right more often than I am. It cannot be the person who chose, cannot carry the cost of being wrong in front of a team, and cannot be answerable to a customer in eighteen months when the choice turns out badly. That is a structural fact about responsibility, not a temporary gap in capability — and unlike the popular version, it does not need a single claim about what models will never do.
Notice what this reframing gives up. It concedes that machines may out-decide us on accuracy. It keeps only the part nobody can take: someone has to own it.
What this changes on a Monday morning
We rewrote how we interview because of this. Candidates used to walk us through something they had built; now we ask them to defend something they rejected — a feature, an approach, a client. The answers separate people faster than any system design question I have used, because anyone can narrate a build, and only someone who has actually decided can explain what they gave up. I regret not making that change two years earlier. We hired well-regarded builders in that window who needed a decision handed to them, and in a year when generation got cheap, that gap cost us more than any skills gap did.
Three things follow for anyone running a team this year. Teach people to write down what they are giving up, not only what they are choosing, because the rejected option is where the reasoning is visible. Attach every significant call to a name and a review date, since accountability that nobody schedules is accountability nobody has. And read the judgment-is-scarce genre — including Fast Company's version of it — as a description of the present tense, useful and probably accurate, with no promises attached about 2030.
I want the comforting story to be true. Wanting it is exactly why I keep testing it.
Frequently asked questions
Is human judgment really becoming more valuable because of AI?
Yes, for now. When AI makes producing options nearly free, the constraint shifts from making to choosing, which raises the value of selection and refusal. That advantage rests on current model limitations rather than a permanent rule, so treat it as a present-tense fact with an unknown expiry date.
What did Oriol Vinyals say about AI models and new ideas?
Vinyals said models coming up with new ideas to try is "not something that currently they're super strong at", speaking around the 5 August 2026 launch of Discovery Loop, which he co-founded with Jeff Dean, Sanjay Ghemawat and Quoc Le. He described a limitation his new company exists to remove.
Can AI models learn human taste and judgment?
Partly, and work is underway. Axios reported in June 2026 that labs are hiring expert human tastemakers to judge model outputs and turn those judgments into training signal. What resists automation is not the ranking of options but accountability — a named person who carries the consequences of a choice.
So here is my question for you, and I would genuinely like the answer. Think of the last significant call you made at work: if a model had ranked the options better than you did, what exactly would still have been your job? Leave your answer in the comments, or put it to your team at the next review and see how quickly they agree.