AI Is Getting More Autonomy. What Leaders Need to Redesign Before Scaling It
Audio Commentary
Having a human in the loop sounds reassuring. But what happens when that human thinks the AI is wrong?
This commentary looks at what happens between formal AI governance and the reality of a consequential decision: who can challenge, pause, require reassessment or change what happens next, and what organisations may miss when accountability exists on paper but authority is harder to exercise in practice.
AI Is Getting More Autonomy. What Leaders Need to Redesign Before Scaling It
# AI Is Getting More Autonomy. What Leaders Need to Redesign Before Scaling It
We’ve spent quite a lot of time wondering whether AI will replace managers. Something slightly stranger is already happening. Managers are getting AI employees.
But calling someone the human manager doesn’t necessarily tell us what they’re managing. And here’s the question I’ve become much more interested in:
**What happens when the human isn’t actually the best person to make the decision?**
I’m Bess Obarotimi. I spend a lot of my time looking at decision environments — particularly what has to be true around a decision for people to exercise judgement, authority and accountability when it actually matters.
And this week I came across something that made that question feel considerably less theoretical.
BNY already has what it calls digital employees. These are AI agents with employee IDs, assigned work and human managers. Microsoft describes around 150 of them operating across the bank, initiating work and responding to client needs and system triggers.
One of them works in payments. And I rather like this. He’s called Payment Pete.
Which does make him sound slightly more approachable than some people I’ve worked with.
But forget Pete for a moment. Imagine you’re his manager. You’ve just been told that this digital employee reports to you.
What do you expect that to mean? Do you allocate its work? Judge its performance? Correct it? Change what it’s permitted to do? Pause it? Override it? Decide when another human needs to become involved?
And here’s the one I’d really want answered:
**What are you responsible for that you cannot personally change?**
Don’t answer that too quickly.
Because when I first started thinking about this, I thought the interesting issue was authority. Can the human stop the AI? Who can override it? Who has the final word?
And then I began to wonder whether I was making the question too easy.
Of course banks think about permissions. Of course they think about controls. These aren’t organisations that have somehow forgotten governance exists.
BNY’s own programme includes extensive governance, compliance and model-risk review around how agents are developed and deployed.
So perhaps the interesting question isn’t: **Do you have a human overseeing the AI?**
Perhaps it’s: **What should that human still be doing once the AI becomes very good at the work?**
Actually, before I go further, I want to ask you to do something.
If you work around AI, risk or governance and you know somebody currently trying to work out how humans and agents should operate together, **send them this episode**.
I’m testing these questions against what organisations are actually encountering. And if this is your kind of work, follow the programme as well.
Not simply because I’d like another follower — although obviously I won’t object — but because I want more of the people actually doing this work in the conversation.
There are questions here that I don’t think any one person, framework or organisation is going to settle on its own. If you’re dealing with them in practice, you’re part of working out what good human–AI decision-making is going to look like.
And that brings me to another organisation.
Kantar now has a **Chief People and Agent Officer**.
Andy Doyle has talked about a workforce increasingly made up of people and AI agents, and the organisational questions that follow once you start treating those two groups as part of the same operating environment.
Molly Wood asked him:
> “Before we get further into the transformation at Kantar, when you think about being the chief people and the chief agent officer, you know, sort of again, going back to those signals that those two things are really interrelated, but also that agents are now employees. Do you think of yourself now as the person who’s in charge of all of the employees, whether human or digital?”
Andy Doyle replied:
> “We do, but it’s a bit of a mind shift for us all. That’s something that if you’d have asked me a year ago, I’d have gone, that’s weird. Today, I think we’re still trying to work that out, and the technology is catching up with that thinking and giving us better opportunities and better way to monitor and create visibility of agents across our organisation.”
Which raises what might sound like a slightly ridiculous question:
**What exactly are we importing when we call an AI system an employee?**
If I were sitting on a panel, I can imagine somebody stopping me there.
“Bess, aren’t we just anthropomorphising software?”
Maybe.
And perhaps that’s exactly why the language matters.
Because managers normally manage people partly because people exercise judgement. They misunderstand things. They make mistakes. They learn. They disagree. Sometimes they know something their manager doesn’t know. Sometimes they refuse.
An AI agent isn’t simply another slightly unusual employee.
So which parts of management transfer? And which don’t?
Because calling somebody the manager can give us a comforting answer to a question we haven’t actually asked.
Then I heard something completely different that made me rethink this again.
I was listening to Michael Saylor talking to Steven Bartlett on *The Diary of a CEO*. They were discussing self-driving cars, and Saylor described what he called the **“big inversion”**:
> “We used to get in traffic accidents, right? Like, the big reveal, right? The big inversion is when people realise that the self-driving cars are safer than the person-driven cars.”
And that caused a problem for my original question.
Because I’ve spent quite a bit of time asking:
**Does the human have enough authority to challenge the machine?**
But suppose the machine becomes demonstrably better at the decision.
Now what?
Imagine an AI-supported process has made a particular judgement hundreds of thousands of times. Its error rate is lower than the human team’s. It’s faster. It’s more consistent.
And today a human disagrees.
Who should win?
Would you allow the human to override it simply because they’re human? Would they need evidence? Would there be particular circumstances where their authority automatically takes precedence?
And if the human intervenes and makes the outcome worse, who owns that?
Take it one step further.
Suppose you’re the senior person accountable for the outcome. But you’ve deliberately designed the system so that the AI is trusted to act because its performance is materially better than your human team.
Then one of your people overrides it.
The AI was right. The human was wrong.
**Was the failure too much automation — or too much human authority?**
That’s a much harder question.
And I think it’s where the sentence “human in the loop” starts becoming surprisingly unhelpful.
It tells me somebody is there. It doesn’t tell me **why they’re there**. And it certainly doesn’t tell me when their judgement should prevail.
The Bank of England has been discussing related problems with industry.
Its AI Consortium has been looking at how increased automation and speed may change financial-system behaviour and existing controls. The Bank has also raised the possibility of guardrails such as circuit breakers or kill switches around AI-driven activity.
But here’s another uncomfortable question:
**What if stopping the AI isn’t always the safe decision either?**
In a real operating environment, the control has consequences too.
Stopping a payments process isn’t neutral. Delaying a decision isn’t neutral. Requiring human approval for everything isn’t neutral.
And this is why I don’t think the answer can simply be: put more humans in.
If I were on that imaginary panel again, I think the next challenge would be:
“Isn’t this just operational resilience or model risk?”
Possibly.
And I don’t want to invent a new problem simply because my own work needs something to solve.
If existing model-risk, operational-resilience, access-control and accountability arrangements already tell people precisely what to do, at the speed required, brilliant.
I genuinely want to find that out.
The cases I’m more interested in are the ones where they don’t. Particularly when an AI-supported action crosses systems, teams and different kinds of authority.
One person owns the model. Another owns the process. Someone else controls permissions. Another person owns the customer outcome. Someone else can suspend the service.
**The organisational chart can be perfectly clear. The actual decision can still be rather less so.**
And this is where I want to briefly step outside the commentary.
Because I’m not trying to build an argument in isolation and then spend six months finding evidence to agree with me.
I’m testing this.
I’m speaking to people working inside AI, risk, governance, operations and regulated decision-making, and I’m particularly interested in the awkward examples.
Maybe your problem isn’t that humans can’t intervene. Maybe **too many** people can intervene.
Maybe the AI isn’t slowing you down at all — perhaps the second line is.
Maybe you’re discovering that the real difficulty is giving the technology enough freedom to produce the return you bought it for.
Or perhaps all of this is already handled perfectly well in your organisation.
I want those cases as much as the ones that support my thinking.
Because we’re collectively having to work out something quite significant here. Not simply how to adopt another technology.
**How humans and increasingly capable systems should divide judgement, action and responsibility inside real organisations.**
If you’re one of the people having to make that work, I’d like you in the conversation.
Message me. Tell me what I’m missing. Tell me what happens where you work.
Or give me one decision that made you think, *this isn’t quite as straightforward as the governance document makes it look.*
The link to my work and contact details are with this episode.
Now, here’s something you can actually take into work.
Don’t audit your entire AI estate. Choose **one** sufficiently autonomous AI-supported process.
One.
And ask six questions:
**What is the AI allowed to do without a human?**
**What must cause it to come back to a human?**
**What can that human change without asking somebody else?**
**What happens when the human and the AI disagree?**
**What happens when intervention itself creates another risk?**
And finally:
**Who has the final word — and is there actually enough time for them to use it?**
I’d be particularly interested in what happens when you answer those questions against the real workflow rather than the governance diagram.
And this brings me back to Andy Doyle. Later in that same conversation, he put the decision rather simply:
> “And where do you deploy more technology solutions or super agents to help you? Where do you put people? Where do you make that choice? And up until now, that choice has been forced on you by the availability of technology and the skill sets of the people you have. Now, I think you have some choices to make about how do you deploy people in the best way.”
Because there is another side to this.
The temptation with AI governance is to imagine that safety always means more intervention.
I don’t think that’s necessarily true.
If the system is better at a particular activity, inserting human judgement everywhere might make the process **less** reliable. It can certainly make it slower.
So perhaps the objective isn’t maximum human control.
Perhaps it is much more precise.
**Know where human judgement adds value.**
Know where a human genuinely needs authority. Know what the system should be free to do without them.
And then design the decision environment accordingly.
That could mean **fewer** human interventions.
But better ones.
And someone listening might reasonably say:
“But Bess, isn’t that just sensible delegation?”
Yes.
Perhaps.
But delegation normally assumes we’re delegating to another accountable human being.
What’s changing is that organisations are beginning to delegate parts of real work to systems that can initiate activity, respond to events and increasingly operate across workflows.
That’s why I think the management question is interesting.
Not because AI needs a boss.
But because we’re having to become much more explicit about what management was accomplishing in the first place.
What are you supervising? What are you deciding? What are you accountable for? What authority do you retain?
And where are you now simply getting in the way?
And this is where the work becomes commercial for me as well.
I’m opening a small number of **AI Decision Authority Reviews**.
It’s a paid ten-day review of **one real AI-supported decision or agentic process**.
We look at what the system can actually do. Where human judgement enters. Who can intervene. Where escalation sits. Who owns the consequence.
And importantly, whether humans are sitting in parts of the decision where they no longer need to be.
The fee is **£3,500**.
If you’re responsible for an environment where AI is beginning to move from assisting people towards actually doing the work, and there’s one decision you’d like tested properly, message me.
I don’t need your entire organisation.
**Give me one decision.**
Because I started this week wondering:
**The AI has a manager. Now what?**
I’m ending it slightly less interested in whether it has a manager at all.
The more important question may be what that human is **uniquely there to do**.
Where should they intervene? Where should their judgement prevail? Where should they have the authority to challenge?
And where, perhaps, should they get out of the way?
The answer won’t simply be “human in the loop”.
It will be much more specific than that.
And I suspect that’s where the interesting work begins.
Why Decision Frameworks Cannot Guarantee ResponsibilityYou can now listen to The Decision Environment on Spotify and Apple Podcasts.
TAKE PART IN THE AI DECISION AUTHORITY STUDY
I’m researching how human authority works inside real AI-supported decisions. If you know one of these decision processes reasonably well, your experience can help build the evidence.
The questionnaire takes approximately 3–4 minutes and does not ask for your employer’s name.
