AI Tools That Feel Smart but Make You Slower
Productivity Analysis

AI Tools That Feel Smart but Make You Slower

Prompting feels like delegation. Waiting feels like free time. Reading the output feels like review rather than work. None of that is true.

I watched a colleague spend most of an hour saving time on an email. The email would have taken three minutes. Instead she prompted, read, re-prompted, edited the output, re-prompted again, gave up, and wrote it herself. Her experience of that hour was cutting-edge productivity. The result was a slightly worse email and a morning gone.

This happens constantly, and it isn’t a story about bad users. “AI-powered” is now a marketing requirement rather than a description of function, tools arrive faster than anyone can evaluate them, and the appearance of intelligence reliably beats actual utility in a purchasing decision. The dark side of AI everywhere isn’t Skynet. It’s death by a thousand helpful suggestions.

My British lilac cat, Pixel, has better judgment about assistance than most product teams. When I try to help her through a door she’s perfectly capable of pushing, she stares at me with open contempt. She knows when help isn’t help.

Why the arithmetic hides

The promise is that you offload cognitive work and spend the reclaimed time on something better. Occasionally that happens. More often the tool consumes more time than it returns, and three things conspire to keep the loss invisible.

Prompting is a skill with a real learning curve, so anyone who hasn’t put in the hours spends enormous effort getting mediocre output. Reviewing generated text takes as much cognitive effort as writing it, sometimes more, but it feels lighter because you’re editing rather than creating. And the switching cost between your work and the tool accumulates invisibly, a few seconds and a chunk of attention at a time.

Run the numbers on one interaction. You switch context out of your work, which costs half a minute plus attention residue. You write a prompt, which is forty-five seconds if you’re practised and several minutes if you’re not. You wait. You read the output, which costs in proportion to its length. You decide whether it’s what you needed. If it isn’t, and often it isn’t, you either re-prompt and repeat the whole cycle or start editing, which adds time to time already spent.

The total frequently exceeds writing the thing yourself. It doesn’t feel that way because every individual step feels efficient. The aggregate is where the loss lives, and nobody measures the aggregate.

Four ways interfaces manufacture the impression of intelligence

The gap between capability and appearance is engineered, and knowing the mechanisms makes it easier to see past them.

Fluency masks uncertainty. Language models produce grammatically perfect, stylistically polished output whether or not the content is accurate or relevant. That polish reads as competence and is entirely disconnected from it. A tool that said “I don’t know” would be more useful and would feel considerably less smart.

Speed reads as authority. When something arrives in two seconds that would take you ten minutes, you unconsciously credit the speed as expertise. But generation speed says nothing about output quality, and a confident wrong answer delivered instantly is worse than a slow correct one, not better.

Length reads as thoroughness. More words feel like more thought, when in fact padding is cheap and filtering is expensive. The comprehensiveness that feels like value is usually work transferred from the machine to you.

And personalisation reads as understanding. A tool that uses your name and adapts to your patterns creates a strong impression of a mental model that isn’t there. That’s effective interface design doing exactly its job.

The categories that reliably cost more than they return

Predictive typing is the most insidious, because it interrupts continuously. Every suggestion is a decision, including the ones you reject, and the evaluation happens below conscious awareness while still consuming resources. The suggestions feel helpful. The interruptions cost more than the keystrokes save, and you’ll never notice because the cost is distributed across every sentence you write.

Summarisation promises to compress a long document into something digestible. Sometimes it does. More often you prompt for the summary, read the summary, wonder what it dropped, and skim the original anyway. You’ve done the work twice and neither pass was your best.

Meeting transcription creates a particular trap. The transcript is too long to read, the extracted action items are too generic to act on, and the belief that the content is captured reduces the pressure to pay attention while it’s happening. The tool makes the meeting worse while appearing to improve it.

Scheduling assistants demonstrate how automation adds steps. You describe your availability. It proposes times. The other person responds ambiguously. It interprets. Clarifications bounce. Something that would have taken two messages takes six, and the machine handled the typing while you handled all the thinking.

How to tell whether a tool is earning its place

Five checks, and the first one does most of the work.

Time the complete cycle honestly, from the moment you decide to use the tool until you have something usable. Not the generation time. Everything: the context switch, the prompt, the wait, the read, the re-prompt, the edit. Then compare against your honest estimate of doing it directly. The tool usually loses, and the gap is usually larger than you’d like.

Judge the output as though you hadn’t seen where it came from. Would you send this as it stands? If not, the tool produced a rough draft, which can be genuinely valuable but is not the same as finished work, and the editing time belongs in the tool’s column rather than yours.

Look for friction that doesn’t show up in a single interaction. Does it need context re-supplied every session? Does it integrate badly with everything else? Does it fail without a connection? Those costs are small individually and enormous across a year.

Then run the removal test, which is the one people skip. Take the tool away for two weeks. If you barely notice, it wasn’t helping. If your work gets noticeably harder, it earned its place. This catches the tools that feel productive and contribute nothing, and it catches them in a way that no amount of introspection will.

Finally, count the learning investment. Hours spent understanding prompting, quirks and workarounds, amortised across the tool’s expected lifetime, which for this category is short. If that per-session cost exceeds the per-session saving, the tool fails economically even when every individual session feels good.

flowchart diagram

The interruption business model

Tools increasingly push suggestions rather than waiting to be asked, and that shift creates a cost that’s genuinely hard to measure.

Every suggestion is a decision point. Accept, modify, ignore. Each consumes something even when the answer is automatic rejection, and the research on interruption and deep work has been consistent for two decades: brief interruptions do disproportionate damage to sustained concentration.

The design isn’t accidental. These products measure success in interactions, acceptances and time in tool. A tool that correctly identified when not to suggest anything would show terrible engagement numbers. A tool that interrupts constantly demonstrates value through activity, and the business incentive points away from your productivity.

The skill worth developing is telling the difference between a feature that exists because it helps you and one that exists because it generates a number for someone else. Sometimes those align. Frequently they don’t.

Where it genuinely works

Having spent two thousand words on the failure cases, the honest other half. There’s a pattern to when this stuff earns its keep, and it’s fairly tight.

It works when the task is precisely defined and quality is objective. Formatting, syntax validation, spell checking. The code compiles or it doesn’t. There’s no ambiguity requiring you to judge whether the output is any good.

It works when the volume exceeds human scale. Finding themes across ten thousand support tickets, searching a document repository nobody has read. The machine isn’t replacing judgment, it’s shrinking the haystack until judgment becomes possible.

It works when you’re stuck and want constraint-breaking. The lack of domain expertise occasionally produces something you’d never have considered. Real value, entirely unpredictable, and you can’t schedule it.

And it works when errors are cheap and iteration is cheaper. Brainstorming, early exploration, throwaway prototypes. Confident incorrectness stops mattering when everything is provisional anyway.

The principle underneath: the tool adds value when you can evaluate quality easily, when the task exceeds your scale, or when being wrong costs nothing. Outside those three conditions it usually subtracts.

Two costs that show up later

The first is attention residue, and it’s specific to partial delegation. Finish a task yourself and your mind releases it. Hand it to a colleague and your mind releases it, because it lives in someone else’s queue now. Hand it half to a machine and it sits in an unresolved state your brain refuses to let go of: something happened, and you haven’t checked whether it’s acceptable.

Stack four of those and you’ve acquired the cognitive weight of four unfinished items plus the time cost of four interactions. You feel busy because you’re tracking a lot. Your output is low because none of it is done.

The antidote is completing the human half immediately, before starting anything else, which is precisely the workflow these tools are designed to discourage.

The second is skill decay. Writing improves by writing. Analysis sharpens by analysing. When the tool handles those, even partially, you get less practice, and the convenience is immediate while the cost arrives years later. That would be fine if the tools were perfectly reliable and permanent. They’re neither. The capability you let atrophy is needed exactly when the tool fails, or changes, or isn’t available, which tends to be an interview, a meeting, or any conversation happening in real time.

I wrote this without assistance, not out of principle but because organising an argument is the part I don’t want to lose. It took longer. The capability stayed mine.

Building a bit of resistance

Default to doing things directly unless you have specific evidence the tool helps with that specific task. That isn’t Luddism, it’s the correct placement of the burden of proof.

Keep some friction in the way. Don’t wire these things deeply into your primary workflow; leave them available and slightly inconvenient, so invoking one is a decision rather than a reflex.

Practise unassisted regularly, partly to maintain the skill and partly because it gives you a recent baseline. Without one you can’t run the comparison at all, and the comparison is the whole argument.

Audit quarterly and be willing to drop things. The sunk cost is real and it’s also irrelevant: what matters is what the tool will cost you from here, not what you already spent learning it.

My own two-year tally, for what it’s worth. Genuinely useful: code completion in languages I already know, spell checking, transcription of voice notes, placeholder image generation. Roughly five tools. Neutral: general writing help, email drafting, meeting summaries. Clearly negative: note organisation, scheduling, research tools that produced confident nonsense, presentation generators.

The pattern is narrow tools sometimes work and broad ones promising help with complex thinking almost never do.

Pixel is currently sitting on the keyboard, having correctly identified that presence cannot be automated. Her success rate at obtaining treats this way is a hundred percent, and no software was involved.

Get the next live webinar in your inbox

One email a month: the upcoming live event + free recording access for subscribers. No spam, unsubscribe anytime.