Tool of the Week: The week I let an AI use my computer

For two years, "AI agent" mostly meant a chatbot that could look things up and write things. This month it started meaning something different: an AI that opens your apps and does the task — reads the screen, moves the cursor, clicks the buttons, fills the forms. Not through a special integration. The same way you would.

Meta shipped this into the mainstream on July 9 with Muse Spark 1.1, which added computer-use across desktop, browser, and mobile. Anthropic's been doing it too — you can now drive your desktop from your phone. The category has a name now: computer-use agents. And it changes what "hand it off" means.

So I handed one a real job this week. I had an AI drive the web editor for this very newsletter — open a new post, set the title, paste in the body, fill the subject line, and schedule it. It did about 90% of it, correctly, while I watched. Then it hit one thing it couldn't do — move text onto the clipboard — and it stopped and asked me to do that single step by hand.

That is exactly the current state of the art, and it's worth understanding precisely, because the hype and the reality are far apart.

What computer-use agents are genuinely good at right now: supervised, one-off, reversible tasks. Pulling data out of a web app that has no export button. Filling a repetitive form. Reformatting something across two tools that don't talk to each other. The kind of tedious clicking you'd hand a sharp intern — while you're sitting next to them.

What they are not good at yet: running unattended. Independent testing this year is blunt about it — computer-use agents "time out, miss consent banners, and lose state on multi-page forms." They're reliable enough for low-to-medium volume with a human ready to catch the edge cases, and too fragile for high-volume, mission-critical, lights-out work. Anyone selling you a fully autonomous computer-using employee is selling you next year's product at this year's reliability.

The honest mental model: you've hired a fast, capable intern who can use any software on the screen — and who occasionally freezes, clicks the wrong thing, or needs you for one weird step. That intern is incredibly useful. You just don't hand them the company checkbook and leave for the weekend.

What to actually do with this:

  • Point it at the tasks stuck in apps with no automation. The real unlock isn't speed — it's reach. The tedious jobs you couldn't automate before because the vendor had no API or integration are now automatable, because the agent just uses the app like a person. That's a genuinely new capability.

  • Keep a human on anything irreversible. Sending, publishing, paying, deleting — those are the steps to supervise or approve, exactly like you would with a new hire. (When I ran the newsletter task, the actual "schedule and send to subscribers" click was mine, on purpose.)

  • Start with one supervised task, not a workflow. Pick one annoying, low-stakes, repetitive thing you do in a browser this week and watch an agent do it once. You'll learn more in ten minutes than from any demo.

Who this is for: anyone whose week is full of manual clicking in software that won't automate — the exports, the copy-paste-between-tools, the form-filling. That work is now delegable for the first time. Just delegate it the way you'd delegate to a person you're still training: give it the reversible stuff, watch the first few runs, and keep your hand on anything that can't be undone.

Quick Hits

Computer-use went from "Anthropic experiment" to "category" this month. Meta shipped Muse Spark 1.1 on July 9 with computer-use across desktop, browser, and mobile, a 1-million-token context window, and parallel sub-agents — Meta claims top benchmark scores for tool use. Why it matters: when Meta ships a capability into the mainstream, it stops being a curiosity and starts showing up in the tools you already use. Expect "let the AI just do it on the screen" to become a normal button in your software over the next year.

The fine print the demos skip: these agents are still fragile. Independent 2026 testing found computer-use agents routinely time out, miss cookie/consent banners, and lose their place on multi-page forms — with a real sweet spot of low-to-medium volume and a human ready for edge cases. Why it matters: this is the single most important thing to internalize before you rely on one. It's a supervised intern, not an autonomous employee. Plan for the human-in-the-loop, and you'll love it; expect lights-out reliability, and it'll burn you.

The real unlock isn't speed — it's the apps that never had automation. For years, "automate it" meant "if there's an API or a Zapier connector." Computer-use agents don't need one; they operate the app like a person, so the tools that never integrated with anything are suddenly reachable. Why it matters: go look at the one piece of software where you do the most manual clicking and that never had an integration. That's now your highest-value automation target — the thing that was impossible last year.

Prompt of the Week: The Delegation Line

Before you point a computer-using agent at anything, sort the work. Paste this:

Act as an operations risk advisor. I'm deciding which of my computer tasks
to hand to an AI "computer-use" agent that operates my apps like a person
(clicks, types, navigates) but is only reliable when supervised, on
reversible, low-to-medium-volume tasks.

I'll describe the tasks I do in software each week. For each one, sort it
into exactly one bucket:
  1. HAND IT OVER (supervised) — repetitive, low-stakes, reversible; safe
     to let the agent do while I watch the first runs.
  2. APPROVE-ONLY — the agent can prepare it, but a human must click the
     final irreversible step (send, publish, pay, delete).
  3. KEEP IT HUMAN — high-stakes, irreversible, or needs judgment the agent
     can't be trusted with yet.

For each "hand it over," name the one thing most likely to go wrong and how
I'd catch it. Finish with the single best task to try first — the one with
the most tedium and the least risk.

Here are my tasks:
[list the software tasks you do manually each week]

Most people find their list is mostly bucket 2 and 3 — and that's the point. The value isn't handing everything over. It's finding the two or three genuinely safe, genuinely tedious tasks and getting those off your plate this week.

Like what you're reading? Forward it to someone who'd get value from it. And if you're curious what AI could actually do inside your business, book a free 15-minute audit — no pitch, just a look at where you're leaving time on the table.

Keep reading