Falling asleep at the wheel: When good ai makes bad decisions

A recent study found that professional consultants working with AI assistance made worse decisions than consultants working alone.

It wasn’t a small study. Nearly 800 consultants at Boston Consulting Group, split into two groups. One worked normally. The other got access to GPT-4. They tackled eighteen realistic tasks—the kind BCG consultants do every day.

On seventeen of those tasks, the consultants using AI crushed it. They were faster, more creative, more analytical. The AI group won by every measure the researchers could devise.

Then came the eighteenth task, carefully designed so the AI would fail—a problem combining tricky statistics with misleading data. Without AI, consultants got it right 84% of the time. With AI assistance, accuracy dropped to 60-70%.

The people with the better tool performed worse.

Ethan Mollick and his research team at Wharton, working with colleagues from Harvard, MIT, and BCG, documented this in Co-Intelligence: Living and Working with AI. What they found explains something important about how AI works—and why the usual questions about whether it will replace your job miss the point.

The Invisible Wall

AI can pass the neurosurgery qualifying exam. It can write working JavaScript code for a tic-tac-toe game in seconds. But ask it to play tic-tac-toe—to identify the next move in a simple grid—and it gives you the wrong answer with complete confidence. Ask it to write a sonnet, and you’ll get something publishable. Ask for exactly fifty words, and it consistently gives you forty-seven or fifty-three, because it processes language in chunks (called tokens) rather than counting words the way we do.

Mollick and his colleagues call this the Jagged Frontier. Picture a fortress wall where some sections jut way out while others fold back in. That’s AI capability. Tasks inside the wall, AI handles. Tasks outside, it doesn’t. The catch is you can’t see the wall. Tasks that seem equally difficult to us—that appear to be the same distance from the center—sit on opposite sides.

When Good Tools Make Us Worse

Most of those BCG consultants just pasted questions into the AI and hit send. They barely edited what came back. And that pattern shows up across professions.

In a separate study, researcher Fabrizio Dell’Acqua tested 181 professional recruiters evaluating job applications. He gave them résumés where math ability wasn’t obvious and AI tools of varying quality to help. Recruiters with better AI made worse decisions. They spent less time per résumé. They followed the AI blindly. They missed brilliant candidates they would have spotted on their own. Meanwhile, recruiters with lower-quality AI stayed sharp, questioned the system, and improved their skills over time.

Dell’Acqua called this “falling asleep at the wheel.” When AI is very good, we stop paying attention. We let it drive. And when the road gets tricky—when the task crosses into territory the AI can’t handle—we crash.

Map Your Own Boundary

What matters for your career isn’t whether AI can do your job in theory. It’s where the boundary runs through your specific work.

You need to map it yourself, because no one else will do it for you. Every role has its own jagged frontier. The only way to find yours is to use AI on real tasks and pay attention to what happens.

Start anywhere. Mollick’s advice is simple: invite AI into everything, not to replace you, but to discover where it helps and where it doesn’t. Most people paste questions directly into AI and accept whatever comes back. That works until it doesn’t. The question is whether you’ll notice when you cross from territory where the AI excels into territory where it fails—and whether you’ll notice before the failure matters.

Mollick describes two approaches. The first he calls working like a Centaur: dividing tasks cleanly between human and machine. You handle strategy, AI handles execution. You decide which statistical analyses to run, AI produces the graphs. The second is working like a Cyborg: weaving your work and the AI’s together continuously, handing small pieces back and forth. You write part of a sentence, AI completes it, you revise, AI suggests alternatives.

When Mollick wrote his book, he used AI to generate ten different versions of paragraphs when he got stuck—not to use them directly, but to break through writer’s block. He asked AI to summarize technical papers to check his understanding, knowing it could get him only partway there. He created AI personas to give feedback on drafts: one pompous editor focused on cutting complexity, one creative connector suggesting unusual angles, one average reader flagging confusing sections. None of them wrote the book. They made it possible for him to write better than he could alone.

The work you hate but can easily verify? Probably safe to hand off. Boilerplate reports. Low-priority emails. First drafts you’ll substantially revise. These likely sit inside the frontier for your role.

Work requiring judgment about subtle context? Problems with misleading data? Situations where confidently wrong is worse than uncertain? These are where you can’t afford to coast.

Why the Failures Matter More

The frontier isn’t static. It shifts as models improve. Tasks you delegate today might be fully automated next year. Work you’re certain requires human judgment might not require it for long. Which means this isn’t a one-time exercise. The skill is learning continuously where the boundary has moved, what new capabilities have emerged, what persistent weaknesses remain.

But right now, knowing where AI fails is more useful than knowing where it succeeds. The successes are obvious. You can see them, use them, build on them. The failures arrive wrapped in confident, articulate explanations that pass the plausibility test. They sound right. They look professional. Unless you know to check—and check carefully—you won’t catch them.

The people who figure this out won’t be those who learn to write better prompts. They’ll be those who understand exactly where the invisible wall runs through their domain. And who stay alert enough to notice when they’re about to cross it, even when the AI is giving them exactly what they asked for.

Related Posts