Sidekick Orchestration

Why AI does more of the work for some people

If AI gives you a draft you have to rebuild, start with the brief, the checks and the corrections you keep.

// In this article

It took me months to get AI doing about 95% of my website design. That's my estimate of the labour, not a measured time saving. Most of those months went into teaching it my taste.

Design is about as subjective as work gets. At the start, AI got me maybe 20% of the way there: pages that were fine and completely forgettable, and constant feedback from me to turn them into something I liked. Over time that feedback got written down as rules and examples, and it grew into a design system the AI works from: the art I love and the things I never want to see.

Now I steer a little and review. The models got better over those months too, and that helped. But a better model won't create my brand. The design system helps make the work look like mine.

On everyday work, I see a big gap between people using the same tools. Picture one person getting 90% of the labour done by AI. Another gets 30%: a draft that's sort of right, an hour of fixing, and a decision that it's faster to do it themselves. Those numbers illustrate the gap. They're not study results or a target every task can reach.

How you set the work up can make a big difference. Give AI a useful brief, tell it what done looks like, and keep the corrections that should apply next time. The tool and the task still matter. Some work will need more of you.

The brilliant new hire with amnesia

Picture a capable new hire who can draft and research quickly, but doesn't yet know your clients, your standards, what you mean by "short," or that you've banned the phrase "unlock value."

A fresh AI chat can feel like that. Projects and memory can carry some information forward, but they don't guarantee that the right detail will be used. Put the standards that matter where the AI can find them.

Otherwise, it's easy to hand over a one-line task, fix the result ourselves at 10pm, and never tell the AI what we changed. Tuesday's report comes back with Monday's problems.

Three habits help. You can save much of the setup so you don't have to explain it all again.

Brief it fully, and let it interview you

A useful place to look when AI work misses the mark is a detail it had to guess: who it's for, what it's for, what you've already tried, what you'd hate to see.

For a substantial piece of work, let the AI pull the missing details out of you. Ask something like:

Before you start, read what I've provided. Ask me one question at a time about anything missing that would change the result. Then play back the goal, who it's for, and what done looks like in a few lines so I can confirm it.

Skip the interview for a small, clear task. You don't need a briefing meeting to shorten a sentence.

In Carnegie Mellon's Ambig-SWE software experiment, coding agents rarely asked for clarification without prompting. Requiring questions helped capable models recover much of their performance on incomplete briefs. For your own work, invite the questions that matter before the work begins.

The playback is where you catch a wrong assumption before it becomes a finished document.

It also helps when a conversation gets tangled. In a study of simulated conversations, researchers at Microsoft and Salesforce found an average 39% performance drop across six tasks when requirements arrived over several messages rather than in a complete prompt. Putting the same information together in one message removed most of that loss.

When a chat keeps building on a wrong assumption, try a fresh one with the corrected brief. Give it the relevant source material too, so it doesn't have to guess again.

You don't need to paste in your whole drive. Use the few things the task needs, within your organisation's rules for sharing information with AI.

Write down what done looks like, then have it checked

If you can't say what done looks like, the AI has little to check against. "Done" can end up meaning "whenever you personally stop being annoyed with it."

People who build AI systems call a test of the work an eval, short for evaluation. For a document, you can start with a short list of questions it should pass.

Some checks are straightforward: under 800 words, includes the required sections, every number matches the spreadsheet. Others need judgment. "Is this clear?" gives the checker little to work with. "Can the reader tell what we need from them and when?" gives it something specific to inspect. An example you loved and one you didn't help explain your standard.

For a follow-up email after a first call with a prospect, an illustrative list might be:

  1. Is it under 150 words?
  2. Does it name the problem they told us about, in their words?
  3. Does every name, date and price match the call notes?
  4. Does it state the agreed next step and any agreed date, without inventing either?
  5. Is the requested action easy to find and understand?
  6. Is it free of every phrase on our banned list?

AI can help check the draft against these questions if it has the call notes. A word counter or exact comparison is better for checks that can be settled mechanically. An AI opinion about clarity can reveal a weak sentence; it can't prove the prospect will understand it.

You still decide whether this is the right moment to ask for a meeting or give them room. Keep the factual checks and that relationship judgment distinct.

For work that matters, give the draft, source material and checklist to a fresh chat for review. Ask it to mark each item pass, fail or unknown, show the reason, and repair supported problems. "Unknown" matters when the evidence is missing. The checker should flag the gap rather than invent a fact to pass the list.

A fresh chat separates the review from the drafting conversation. It can still make the same mistake, so treat its verdict as help with your review, not a guarantee. Read the result before you send it.

Before relying on the checker, try it on a few pieces you've already judged, including one with a known error. Where it misses something or rejects good work, improve the question or give it a better example. A few trials are a starting check, not proof it will catch every error. Anthropic's evaluation guidance explains why AI judges need comparison with human judgment.

I started checking work this way after a slip I wasn't comfortable with. I ran a session too long, the AI mixed up two similar ideas, my review wasn't tight enough, and an asset went to a client with an error in it. Minor, but it shouldn't have happened. The checking step I built now catches issues constantly and flags the work it isn't sure about.

Keep your fixes

Your corrections tell AI something useful about your standards: "Too long." "We never say leverage." "This client hates being sold to." "Lead with the number."

If you only fix the document, that lesson may never reach the next task.

Keep reusable corrections where the AI can find them next time. In ChatGPT or Claude, that can be a Project's instructions or a short style guide. If you've fixed the same thing twice, write down the rule and its reason. "No bullet points in this client's emails, because they read like a form letter" teaches more than "no bullet points."

That is an example of one client's preference, not a rule against lists everywhere. Keep client-specific notes with that client's work. Save a general rule only when it should apply more broadly.

Same with my website. The other day I called a new page "clean, but boring," a six out of ten, and said I hate those faint grid lines. Notes like that go into the design system, so I don't have to say them twice.

Keep the list short enough to use and remove contradictions. In a test using large sets of keyword rules, models missed more requirements as the list grew. Save the rules that keep helping, and remove the ones that no longer apply.

Save the setup you want to repeat

Written out like this, it sounds like a lot of work on every task. Much of it belongs in the setup: the questions worth asking, the definition of done, examples and review instructions.

A Project can keep those instructions and files available when you return. You still start the work and provide what has changed. Saving instructions alone doesn't create a separate reviewer or an unattended process. Those steps need to be set up and checked in the tool you're using.

The review can go beyond the checklist. Ask AI to find the strongest objection, then look at the work from the reader's situation. What would a skeptical finance lead question? What would frustrate the person who has to use this? What would a customer struggle to understand? What's missing?

These are simulated perspectives. They can expose a gap before a real person sees it, but they aren't customer feedback.

Ask for one ranked list of material issues, with supported fixes applied and unresolved questions left visible. Stop when no worthwhile, well-supported improvement remains. The point is to bring you better work, not an endless series of review passes.

Where your taste still comes in

Deciding what to act on is the part that stays yours. Some work will always need your judgment: whether a design feels right, a message respects the relationship, or a recommendation makes sense for the business.

Writing down your preferences, examples and corrections gives AI more to work with. It can get closer to what you want. You decide whether it has, and whether the time spent setting up and checking the work is paying off.

The checklist

For a substantial or repeated task:

  1. Have we confirmed the goal, reader and what done looks like?
  2. Does the AI have the relevant information it's allowed to use?
  3. Are the important checks written down, with an example of good work?
  4. Have I tried the checker on work I've already judged, including a known error?
  5. Has the draft been checked against the sources and the list, with gaps left visible?
  6. Have I reviewed the facts and decisions that matter before using or sending it?
  7. Did a correction reveal a reusable rule, and have I saved it in the right place?
  8. Is the setup saved so I can use it again?

Pick one piece of work you hand to AI every week and try this. Notice how much fixing it still needs and whether the finished work is better. The setup earns its place when it improves the result or reduces the effort it takes to get there.

If you want help setting this up, bring one piece of work that keeps coming back for revision. Talk through what's stuck.