Skip to content
LIVE
The Executives BriefThe Executives BriefBeta

Managers made 18% more errors when they treated AI as an “employee” instead of a chatbot

A Boston University study suggests the more “human” you make agent tools, the worse people perform.

ByYousef Al-ZahraniTechnology Correspondent, The Executives Brief
·3 min read
Managers made 18% more errors when they treated AI as an “employee” instead of a chatbot
Executive summary

Boston University professor Emma Wiles and collaborators found managers made 18% fewer errors when work was attributed to an agentic “AI employee” rather than a chatbot. The finding challenges the push by Microsoft, OpenAI, Anthropic, and Google to sell AI agents as digital colleagues.

Your company is about to “hire” a new coworker. Not a person, an AI tool. It shows up with a name like Alex, a title, and defined responsibilities. And the scary part is not that it can do work. It’s that people will behave differently around it.

Boston University professor Emma Wiles studied how managers interacted with this kind of setup, and the results are oddly uncomfortable: managers caught 18% fewer errors when the work was attributed to an agentic “AI employee” rather than a chatbot. In plain English, when teams treated the system like a worker they could defer to, the human quality-control instincts loosened. They made more mistakes, and then they were less likely to catch them.

That’s a big deal because the market is moving hard in the opposite direction of “be careful what you call it.” Microsoft, OpenAI, Anthropic, and Google have all released tools for managing teams of AI agents, and many of them are advertised as digital colleagues. These products are built on a human-centric premise: if agents can coordinate, justify, and execute tasks, then giving them coworker status will accelerate adoption. Wiles’s study suggests a catch. Titles, implied authority, and role clarity can change behavior, sometimes toward worse outcomes.

This is also why the “AI coworker” pitch is more than branding. When executives roll out agentic systems, they are not only deploying software. They are redesigning accountability. If an agent is treated as an employee, what does that mean for escalation paths, review processes, audit trails, and sign-off? Your workflows do not remain neutral. Humans adjust how they trust, how they verify, and when they stop digging.

For decision-makers, the operational question becomes sharper: do you want your organization to treat the system as a system, or as a colleague? The study implies that “colleague” status can reduce error detection, even when the same underlying capability exists. That matters for everything from customer support and coding assistants to internal compliance checks. And it is exactly the kind of second-order effect that shows up only after teams start using the tools at scale, when the costs are no longer theoretical.

There’s also a governance angle, because “AI agents” are entering the regulatory conversation as if they are semi-autonomous actors. The newsletter notes that Senator Mark Warren is set to introduce a bill to regulate AI agents, with rules for agent permissions and verification. In other words, lawmakers are starting with the mechanism that the Boston University study hints at: verification. If organizations already struggle to catch errors when humans are psychologically “deferential,” then agent permissioning and verification rules are not just legal theater. They are operational necessities.

Meanwhile, the broader tech landscape is full of parallel stories where capability advances outpace social and safety frameworks. One of today’s highlights: Sceye, a New Mexico-based company, plans to launch a roughly 200-foot-long high-altitude platform station in a test that includes beaming data from about 18 kilometers above the ocean’s surface. It is targeting a supplementary role to a 5G network, and the run begins as soon as August, with a path from the southwestern US to Japan. The underlying theme is similar: new systems are being designed to blend into existing infrastructure, and success depends on how well the interfaces and assumptions work.

Zoom back to the enterprise and the strategic stakes sharpen. If AI agents are positioned as coworker-like employees, boards and audit committees will need to ask whether oversight processes actually match the way humans behave. Because if teams catch 18% fewer errors under that framing, then “deployment” is not a checkbox. It is a behavioral change management problem with measurable failure modes.

And this is where it gets real for everyone running product, ops, compliance, or AI governance. The next wave of “agent management” tools will likely keep emphasizing collaboration and colleagues. But if your organization wants real reliability, you will need to treat the system as high-leverage software whose outputs must be verified, not as a coworker whose role makes people assume it has the judgment to catch its own mistakes. That’s the difference between faster automation and an avoidable quality incident.

Executive ActionsLocked

This story's Key Insights and Take-aways are locked.

Create a free account to unlock Executive Actions for one credit.

Register to Unlock

Always free for Executives Club members. Join the Club

More in Technology