From Sanskriti Khandelwal | Product & Market Analysis
The Human Skills That Just Became More Valuable, and the Ones That Did Not
On this page
Employment for workers aged 22 to 25 in the most AI-exposed occupations sits 19% below where it would otherwise be, on Stanford's payroll analysis through mid-2026. Employment for experienced workers in the same occupations did not fall. The dividing line is not seniority. It is whether the knowledge a job runs on had already been written down. That is what makes certain human skills more valuable now, and the list is shorter than the reassurance pieces suggest.
Key takeaways
- The split is codified against tacit, not human against machine. Stanford's payroll data shows employment falling for young workers in occupations built on documented, standardised knowledge, and rising for experienced workers in occupations that run on knowledge acquired through practice.
- Judgement is worth most where the model is confident and wrong. In a field experiment with 758 consultants, those using GPT-4 on a task outside its capability were 19 percentage points less likely to reach a correct solution than colleagues using no AI at all.
- You cannot feel the difference from the inside. A randomised trial found 16 experienced developers took 19% longer with AI tools, while believing afterwards that the tools had made them 20% faster.
- The pay signal is real, and it is not paying for warmth. Postings naming AI skills advertise 28% higher salaries across 1.3 billion postings. Communication earns when it is attached to a decision somebody owns.
What a human skill means once execution is cheap
Most writing on this subject uses "human skills" to mean warmth. Empathy, collaboration, communication, listening. That definition is comfortable and it does not survive contact with the employment data.
A more useful test has two parts. A skill holds its value if exercising it requires information that was never written down anywhere. It holds value again if the correctness of the call cannot be checked at the moment you make it.
Both conditions describe one underlying property. The work resists being turned into a specification. Anything that can be specified precisely enough to hand to a competent stranger can now be handed to a model instead, at a fraction of the cost.
That is why the reassuring lists fail. Communication appears on every one of them, and communication as a category is not scarce. Writing a clear status update is not scarce. Telling a client that the plan they signed off is wrong, and keeping the account, is scarce.
Hold that distinction through the rest of this piece. The unit that gained value is not the skill in the abstract. It is the skill attached to a decision somebody is accountable for.
The labour data: a 19% gap, and where there is none
Start with the only large-scale measurement of what has actually happened to employment rather than to sentiment.
Stanford's Digital Economy Lab tracks payroll records from ADP. Its August 2026 update reports that employment for workers aged 22 to 25 in the most AI-exposed occupations now runs 19% below where it would sit had it kept pace with less-exposed peers. In July 2025 that same gap was 15%.
Two details matter more than the headline number. The divergence runs through reduced hiring rather than through increased separations. And there is no comparable gap for experienced workers in the same occupations.
So this is not a story about roles vanishing. It is a story about the entry point closing, which is the same mechanism examined in the piece on the junior developer pipeline.
Codified knowledge lost. Tacit knowledge gained.
The researchers name the mechanism directly, and it is the most useful sentence in the whole literature on this question.
Employment declined for young workers in roles that rely on codified knowledge: formal, standardised, documented knowledge. Employment rose among experienced workers in occupations that rely more heavily on tacit knowledge acquired through practice and mentorship.
Read that as a definition rather than as a finding. Codified knowledge is knowledge already written down somewhere a model has read it. Tacit knowledge is the residue: what a practitioner knows and cannot fully explain.
The same split shows up in where the declines cluster. Stanford found them concentrated in occupations where AI usage tends to automate tasks. In occupations where AI is used to complement the worker, employment is flat or rising.
That is the argument of this post in two sentences, measured rather than asserted. Everything after it is working out which specific skills sit on the tacit side of the line.
The four skills that actually gained value
Four, not twelve. A longer list is a comfort exercise. These are the ones where the evidence and the mechanism both hold up.
Judgement: catching the confident wrong answer
Judgement here has a narrow meaning. It is the ability to look at a plausible output and know it is wrong before anyone downstream finds out for you.
This is the skill that scales worst and pays best. It scales badly because it cannot be delegated to the thing being checked. It pays well because the cost of a missed error rises with output volume, and output volume has gone up sharply.
Note what it is not. It is not scepticism. A reviewer who rejects everything is as useless as one who approves everything, and slower. The analysis of the review bottleneck covers what happens to a team that adds generation capacity without adding judgement capacity.
Problem framing: choosing the question
Framing is deciding what should be built before anything gets built. When production was expensive, framing errors were caught by the cost of the build itself. Nobody spent six weeks on the wrong feature without somebody asking why.
Cheap execution removes that check. A team can now produce the wrong thing quickly and well, and the polish of the artefact disguises the error in the brief.
This is the skill the market is quietly repricing under a different name. What gets sold as context engineering is mostly problem framing with a technical vocabulary bolted on.
Negotiation: pricing work nobody can verify
Buyers now assume a proposal may have taken an afternoon. That assumption changes the negotiation whether or not either side says it out loud.
The answer is not to hide the tooling. It is to move the priced object from effort to outcome, and to defend that move in the room. That is a negotiation skill, and it has become load-bearing for anyone who sells work rather than hours.
My own position: this is the most underrated item on the list. Judgement and taste get written about constantly. The ability to hold a price while the buyer believes your costs collapsed gets written about almost never, and it decides more incomes than either.
Taste: the standard you hold when nobody is checking
Taste is the least respectable word here and the hardest to fake. It is the ability to tell good from adequate in a specific domain, quickly, without a rubric to lean on.
It matters more now for a mechanical reason. Models produce adequate output reliably and good output inconsistently. Somebody has to know the difference at the moment of shipping, because volume makes after-the-fact correction impractical.
Taste is also the clearest case of tacit knowledge in the Stanford sense. Nobody wrote the rule down. You acquire it by seeing several thousand examples with feedback attached, which is precisely the apprenticeship the entry-level squeeze is interrupting.
| Skill | Why it resists specification | What it looks like when present | How to test it |
|---|---|---|---|
| Judgement | The error is only visible against context the model does not hold | Rejections that turn out to have been right, at a steady rate | Hand over a good-looking artefact with one buried flaw |
| Problem framing | The brief is written before the evidence exists | Work that gets narrower, not broader, as it starts | Ask what the candidate would refuse to build, and why |
| Negotiation | Value is argued, not computed, and the counterparty adapts | Prices that survive a cost-transparency conversation | Role-play a client who says the work looks automated |
| Taste | The standard was never written down in the first place | Fast, consistent ranking of near-identical options | Blind-rank four outputs and explain the ordering |
The test column is our own construction, not a validated instrument. It exists because the alternative in most hiring processes is asking candidates to describe their judgement, which measures fluency rather than judgement. There is a fuller question set in the AI fluency interview guide.
What the field experiments actually show
Two experiments carry most of the weight in this debate. Both are worth reading in full, and both used models older than the ones your team runs today.
The frontier is jagged, and invisible from the inside
Dell'Acqua and colleagues ran a pre-registered field experiment with 758 Boston Consulting Group consultants, roughly 7% of the firm's individual contributors. The work was published in Organization Science.
Inside the model's capability, the results were large. Consultants using GPT-4 completed 12.2% more tasks, worked 25.1% faster and produced output rated more than 40% higher in quality by blind graders.
Outside it, the sign flipped. On a task the model handled badly, consultants using AI were 19 percentage points less likely to reach a correct solution than colleagues working without it.
The finding that matters is neither number on its own. It is that participants could not tell which side of the line they were standing on. The boundary is jagged, the model sounds equally confident on both sides, and confidence is the only signal the user receives.
That is a precise description of what judgement is now for. Not doing the work. Locating the edge of the tool.
People are poor judges of their own uplift
METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real tasks, in repositories where they averaged five years of prior work.
Developers took 19% longer to complete issues when AI tools were allowed. They had forecast a 24% speed-up beforehand. After finishing, they still believed AI had made them 20% faster.
The distance between the measurement and the perception is roughly 39 percentage points. Self-report is not evidence about productivity, and a version of the same disconnect appears in the gap between how much developers use these tools and how far they trust them.
METR is unusually careful about what the result does not establish. The authors state that it does not show AI fails to speed up most developers, that it does not extend beyond software, and that it may not hold with better tools.
The skills that did not gain, and the one that lost
A list of winners with no losers is a horoscope. Here is the other side of the ledger.
Prompt craft was a two-year skill
Prompt engineering was a real skill with a short half-life. Its value came from knowing the quirks of one model generation, and those quirks were exactly what the labs were working hardest to remove.
What survived is not the phrasing. It is knowing which information a system needs and where that information lives inside your organisation. That is not a prompting skill, and calling it one has cost people two years of misdirected practice.
Tool fluency has followed the same curve, one step behind. Being the person who knows the interface is worth a lot in year one and very little in year three, because the interface gets simpler and everybody else catches up.
The genuinely exposed class is codified procedural knowledge. Knowing the standard handling of a standard case, applied by somebody who cannot say why the standard exists. That is close to a definition of what Stanford found employers hiring less of.
The World Economic Forum's employer survey adds an uncomfortable data point. Employers reported manual dexterity, endurance and precision in net decline, with 24% expecting a decrease. Dependability and attention to detail also slipped, as did reading, writing and mathematics.
I think employers are wrong about attention to detail, and it is worth saying so plainly. Verification load rises with output volume. Reporting that detail matters less, in the same period that machine-written output needs more checking, reads as a survey artefact rather than a finding.
Where this argument is weakest
Four objections, in descending order of how much they should worry you.
None of the employment data is causal. Stanford says so in its own summary: these are descriptive patterns, not causal estimates of the effect of AI. The period from 2022 also contains a broad technology hiring correction that had nothing to do with models.
The pay premium is for AI skills, not for judgement. Lightcast counts named AI skills in postings. No dataset prices taste or problem framing directly, because postings do not name them in any countable way. The link from the employment data to the four skills above is our inference, not a measurement, and you should treat it as one.
The experiments are dated and small. The BCG study used GPT-4 in 2023. The METR trial used early-2025 tools, 16 developers and a single domain, and METR has since changed its experiment design. These are the best available evidence and neither describes the current frontier.
The strongest counter-case is that verification itself gets automated. If models become reliable at checking their own output, the judgement premium compresses quickly and four skills become three. I do not expect that soon, because the failure mode here is confident wrongness rather than expressed uncertainty. That is a belief, not a finding, and the honest version of this post says so.
One more caution about the survey evidence in general. Anthropic's own June 2026 index notes that experienced workers describe AI as lacking the judgement and situational reasoning their work requires. That is people describing their own value, which is the least reliable form of testimony available.
How to build these, concretely
Development paths, not affirmations. Each of these produces a record you can inspect later, which is the only way to tell whether the skill actually moved.
| Skill | Practice to run this quarter | How you know it worked |
|---|---|---|
| Judgement | Keep a decision log on AI output: what you accepted, what you rejected, and why | Your rejection rate stabilises and your reversals fall over three months |
| Problem framing | Write a one-page brief before any build, naming what you will not do | Fewer scope reversals after work starts, measured against last quarter |
| Negotiation | Reprice one engagement on outcome rather than effort, and hold it | The deal closes without an effort-based discount being conceded |
| Taste | Blind-rank five competing outputs weekly, then read the expert ranking | Your ordering converges with the reference ordering over eight weeks |
Three notes on running this with a team rather than alone.
First, judgement needs volume and feedback inside the same loop. A reviewer who never learns which of their rejections were correct does not improve at all. Log the calls, then revisit them at a fixed interval.
Second, the apprenticeship problem is structural and no amount of individual effort fixes it. If juniors never do the codified work, they never build the pattern library that tacit judgement is made from. That is a staffing design question, and it sits next to the one raised in the analysis of middle management after agents.
Third, if you write this into a review cycle, the criteria have to change with it. Measuring output volume in a period when output is cheap rewards the wrong behaviour, which is the subject of the piece on performance review criteria.
Frequently asked questions
What human skills are most valuable in the age of AI?
Four hold up against the evidence: judgement about when an output is wrong, problem framing before work begins, negotiation over price and expectations, and taste in a specific domain. What they share is that each depends on knowledge that was never written down, and each attaches to a decision somebody is accountable for. Generic communication and collaboration do not clear that bar on their own.
Are soft skills more valuable than technical skills now?
That framing does not match the data. Job postings naming AI skills carry a 28% salary premium, close to $18,000 a year, so technical fluency is still being paid for. The change is that technical fluency alone has a short half-life, because each new model generation removes the quirks the skill was built around. The pairing is what pays: a technical floor plus judgement about when to override the tool.
Is AI really replacing entry-level jobs?
The measured effect is on hiring, not on firing. Stanford's payroll analysis found employment for workers aged 22 to 25 in highly AI-exposed occupations running 19% below where it would otherwise be, driven by fewer new hires rather than more separations. Experienced workers in the same occupations show no comparable gap. The authors describe these as descriptive patterns, not causal estimates, so treat the number as a strong signal rather than proof.
How do I develop judgement if AI does most of the work?
Judgement comes from making calls and finding out whether they were right. Set up a loop you can inspect: record the decisions you make on AI output, note why you accepted or rejected each one, and review that record monthly against what actually happened. Volume without feedback builds nothing. The practice matters more than the reading, because the knowledge you are trying to acquire has never been written down.
Do AI skills actually pay more?
Yes, on the posting data. Lightcast analysed over 1.3 billion job postings and found those requiring AI skills advertised 28% higher salaries, close to $18,000 more per year, with 51% of those postings outside IT and computer science roles. US postings naming AI skills grew 144% in the year to April 2026, against 7% growth for postings overall. That is advertised pay, not paid pay.
Which skills are losing value because of AI?
Codified procedural knowledge is the exposed class: knowing the standard handling of a standard case without knowing why the standard exists. Prompt craft has already faded, because its value came from model quirks the labs then removed. Employers surveyed by the World Economic Forum also reported declining importance for manual dexterity, with 24% expecting a decrease, and small net declines for reading, writing and mathematics.
Where to start this week
Pick the one that matches your position rather than doing all four.
If you manage people, open your last ten hires and mark which roles you filled for codified knowledge. Those are the roles where the Stanford gap will show up in your own numbers first, and knowing the count is worth more than any reskilling plan you could write this month.
If you do the work yourself, start the decision log today. One line per accepted or rejected AI output, with the reason. In eight weeks you will have the only evidence anyone has ever offered you about your own judgement, and the METR result suggests your intuition about it is not reliable.
Related analysis
The staffing side of this argument is covered in the piece on the breaking junior pipeline, and the hiring side in the AI fluency interview questions.
References
- Stanford Digital Economy Lab, No Widespread Displacement, but the AI Employment Gap for Young Workers Has Widened to 19%, August 2026. Used for the employment gap, the codified versus tacit knowledge finding, and the hiring mechanism.
- Dell'Acqua, McFowland, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon and Lakhani, Navigating the Jagged Technological Frontier, 2023, later published in Organization Science. Used for all 758-consultant figures.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Used for the 19% slowdown and the perception gap.
- Lightcast, Beyond the Buzz: Developing the AI Skills Employers Actually Need, 23 July 2025. Used for the 28% salary premium and the share of AI postings outside IT.
- Bipartisan Policy Center, Navigating Skills Trends: Data Dashboard Analysis, April 2026, built on Lightcast data. Used for posting growth rates.
- World Economic Forum, Future of Jobs Report 2025, Skills Outlook. Used for declining skill importance and the 39% skill-change figure.
- Anthropic, Economic Index report: Cadences, June 2026. Used for the worker survey responses on judgement and relational work.
The weakest thing about this source base: none of it prices judgement, taste or framing directly. The employment and posting data measure occupations and advertised roles. Mapping those onto four named skills is this post's inference, not a finding any of the sources make.
Related reading