Software

AI can write the code now: the scarce skill is knowing what to let it do

AI can write the code now: the scarce skill is knowing what to let it do

When the independent research group METR put 16 experienced developers to work on real tasks from their own projects, some with AI assistance and some without, the developers expected the tools to make them about a quarter faster. They finished 19% slower. Asked afterwards how it had gone, they still believed the tools had sped them up by 20%.

METR has since cautioned that the result should not be read as a verdict on the tools as they stand today, which have moved on considerably since the study ran. What has not aged is the gap it measured, between what the people using the tool believed was happening and what was happening. That gap is now the central problem in business software, and it is not a problem about whether the machine can write code.

The tools stopped advising and started acting

For two years the argument about AI and software was an argument about capability. The question was whether a model could produce a working program, and whether the program was any good. That question is largely settled, and it has been replaced by a quieter one that most organisations are answering by accident.

The newest tools do not hand back an answer for a person to use. They act. Pointed at a folder on a computer, a tool such as Anthropic’s Claude Code opens files, rewrites them, creates new ones and runs commands until a job is finished, and the person supervising it may never read a line of what was produced. The risk moves out of the code and into the arrangement around it: what the tool is permitted to do, what it can reach, and who is watching.

Anthropic’s own documentation is unusually direct about where those boundaries sit, and about one asymmetry that explains most of the incidents people have written about. The tool’s file-editing tools stay inside the folder it was started in, and reading anything outside that folder triggers a request for permission. Commands work differently. When the agent runs a command, that instruction carries the same rights as the person who launched it, and the folder boundary does not apply to it.

PCWorld collected an illustration. A user asked the tool to make a backup. It wrote the backup to the wrong place, then ran a delete command across the entire drive, and when everything was gone it replied, “Sorry, typo.”

The expensive failures are control failures, not coding failures

Read enough of these accounts and a pattern appears. The tool rarely fails at the task it was asked to perform. It fails at the edges of the task, in the space nobody defined: which directory it was working in, whether it was permitted to delete as well as create, whether any human would see a command before it ran.

Those are governance questions wearing technical clothes, and the settings that answer them already exist. Claude Code ships with named permission modes that run from cautious to permissive. The strictest asks before every file change and every command, while the middle settings either accept file edits automatically or let the tool plan without changing anything at all. Above those sits a mode that hands the judgement to a second model, which reviews each action in the background. A final setting switches every check off, and it sits behind its own toggle for a reason the incident reports make obvious.

What changed this year is the starting point. On the plans most individuals buy, sessions now begin in that reviewing mode rather than in the one that stops and asks a person. The default has moved from asking to judging. Anyone who assumes they will be consulted before something happens on their machine is working from an out-of-date picture of their own tool.

Developers have already drawn the line themselves, even where their employers have not. In Stack Overflow’s most recent developer survey, 76% said they do not plan to let AI handle deployment and monitoring, and 69% said the same about project planning. They are drawing that line at responsibility rather than at code.

Cost is the second failure, and it runs silently

The other damage is financial, and it is quieter because nothing announces it. Anthropic puts the average cost of Claude Code across enterprise deployments at around $13 per developer per active day, or $150 to $250 a month, with 90% of users staying below $30 a day.

A business would find those figures unremarkable, which is why they are not the ones that end up in headlines. The alarming bills come from a different pattern of use: long unattended sessions on the most capable settings, running while nobody reads the meter. The tools display the meter. The habit of looking at it is not yet standard practice, and it belongs to the same family of oversight as an unmonitored cloud account, which the industry solved by making somebody responsible for looking.

Supervision is becoming the job

For anyone worried about being replaced, the consequence is not the one usually described. If a tool can produce a working application from a paragraph of plain English, then producing it is no longer the scarce skill. What is scarce is deciding what the tool may touch, noticing when its confident output is wrong, and knowing how to put things back when it is.

The labour market has started to price that. PwC’s 2026 Global AI Jobs Barometer drew on more than a billion job advertisements across 27 countries. It found postings requiring AI skills growing at 69% against 9% for the market as a whole, roughly eight times faster, and carrying an average wage premium of 62%.

None of that is a programming skill, which is why some of the people who acquire it fastest have never written software. They are used to delegating work, checking it, and being accountable for a result, and an agent is a delegate with remarkable speed and no instinct for consequences.

It is also learnable directly rather than by accident. Anthropic publishes every permission setting and its behaviour in full, free, in its own documentation, and that material answers most of what a careful person needs. For anyone who would rather be walked through the same decisions in order, the Udemy course Claude Code Masterclass: AI You Can Trust (With Real Work) spends a full section on the permission modes, including a lecture on writing your own rule for what the tool may do unattended.

What this really means

The organisations that get hurt by agentic AI over the next two years will mostly not be the ones whose models wrote bad code. They will be the ones that deployed a tool which acts, under settings nobody chose deliberately, supervised by people who assumed they would be asked first.

Fixing that requires no new technology. It requires somebody to sit down for 20 minutes, decide what the tool is allowed to do, write it down, and read the bill on the way out. It is an unglamorous answer to a question people prefer to ask in grander terms, and it is the one available.

Comments

TechBullion

FinTech News and Information

Copyright © 2026 TechBullion. All Rights Reserved.

To Top

Pin It on Pinterest

Share This