Sidekick Orchestration

Should AI Do This Work?

A practical scorecard for deciding what AI may run, what it should prepare, and where people must stay in control.

// In this article

Take a weekly report rebuilt from project notes every Friday. AI may be able to draft it quickly. That still does not tell you whether AI should run the task, prepare a draft, or stay out of the way.

In this article, run the task means AI completes the routine steps. Prepare the work means AI creates a draft that a person checks and approves.

Sidekick's Work Delegation Screen helps you choose between those options. Pick one recurring task. Score five factors from 0 to 2. The total gives you one of three starting routes.

An earlier Sidekick guide, When NOT to Use AI, made the simpler case: split a job into parts, let AI handle suitable labour, and keep people responsible for judgment, relationships, promises, and final decisions. This screen adds a practical way to score one task and decide how far AI should go.

5factors scored from how the work happens today
0 to 2points for each factor
3possible starting routes
AutomateAI or a simple automation may be able to run the task.
AccelerateAI prepares the work. A person steers and decides.
Stay HumanA person keeps responsibility for the task.

The screen helps you decide what deserves a closer look. It does not approve automation. It also does not replace privacy, security, legal, ethics, or risk review.

The score suggests where to start.Use the safety flags to decide how much of the task AI may actually do.
Use the worksheetScore one real task while you read.

Record the evidence, safety flags, and next decision on two printable pages.

Download the worksheet

Start with one task

Do not score a department such as sales, finance, or customer service. Choose one task or one useful step in a larger process.

Good examples include:

  • preparing a weekly project update;
  • sorting new support requests;
  • drafting an overdue invoice reminder;
  • comparing a renewal with the current contract; or
  • turning meeting notes into decisions, owners, and follow-up.

Before you score the task, write down:

  1. What starts the work?
  2. What useful result should it produce?
  3. Who owns the result?
  4. How often does it happen?
  5. What information does it use?
  6. Who is affected if it goes wrong?
  7. What is one approved example you can inspect?

If the task is not clear enough to describe, it is not clear enough to score.

Score the five factors

Score what happens in the real work today. Do not score the task based on what you hope AI will be able to do.

Factor012
Repetition and workloadRare, trivial, or too variable to justify changingRecurring with moderate effortFrequent, high-volume, or takes meaningful time
Process clarityDepends on judgment that has not been explainedThe normal path is clear, but important exceptions remain unstatedA capable person can explain the normal path and examples in about ten minutes
ReviewabilityAn important error is hard to find before harmReview is possible but needs substantial expert workA reviewer or fixed check can find important errors quickly
Cost of errorSerious, hard to reverse, or involves a decision only a person may makeContainable with a clear human checkpointLow impact, easy to reverse, and no unapproved external action
Human valueThe person's relationship, identity, accountability, or discretion is the workContext and judgment matter, but AI can prepare a useful first passThe result depends mainly on stated rules, evidence, or format

Add the five scores for a total out of 10.

1. Repetition and workload

This asks whether the work deserves redesign. A task that takes five minutes once a year may be annoying, but it is unlikely to justify a new system. A task that consumes several hours every week, delays customers, or appears across hundreds of cases may.

Record frequency, volume, and time when you can. If those facts are unknown, write down what you need to measure rather than inventing a number.

2. Process clarity

Could a capable person explain the normal path and show useful examples in about ten minutes? The steps, inputs, expected result, and common exceptions should be clear enough for another capable person to follow.

If the work depends on hidden judgment, AI will not repair the missing operating method. Clarify the process or keep the task with an experienced person.

3. Reviewability

Can someone find an important error quickly? Easy review supports more delegation. If review needs substantial expert work, use the Accelerate route and keep a person in charge. If an error would be hard to catch before it causes harm, keep the task with a person.

A fast draft is not useful when checking it takes as long as doing the work.

4. Cost of error

Consider the size of the harm, who is affected, whether the error creates an external action, and how easily the result can be corrected.

Do not let the total score hide this factor. A serious consequence can limit what AI may do even when the total score is high.

5. Human value

Is the person part of the value of the work? Trust, negotiation, leadership, care, discretion, and accountability may be essential to the result.

AI can often prepare context or a first draft without taking over the relationship or final call.

Read the result

The total suggests a starting route. It does not make the final decision.

8 to 10

Automate

AI or a simple automation handles the task after the team confirms it is ready and safe.

4 to 7

Accelerate

AI prepares or analyzes the work. A specific person reviews and decides.

0 to 3

Stay Human

A person handles the task. AI may prepare a short brief.

The score and safety checks answer different questions.A task can score as an Automate candidate and still begin in Accelerate because risk, permission, or human judgment requires review. The score suggests a starting point. It does not overrule common sense.

Before automating a task, ask whether ordinary automation would be more reliable. Use AI when the work needs interpretation, comparison, sorting, or a flexible draft. Use ordinary automation when the steps and rules do not change.

Check the safety flags

Before you plan a small test, mark every item that applies:

  • sensitive, personal, confidential, or restricted information;
  • customer, employee, partner, or public communication;
  • official records or regulated work;
  • money, pricing, access, permissions, signatures, or commitments;
  • advice or decisions that could materially affect someone;
  • errors that are difficult to find, contain, or reverse; or
  • no named owner, review step, reason to stop, or way to undo the change.

Any flag can limit what AI may do, regardless of the total.

A high score and a safety flag may mean:

  • keep the task draft-only;
  • remove sensitive information;
  • require a qualified person to review every result;
  • test with made-up examples instead of real business or customer information;
  • repair the process before adding AI; or
  • stop because the task is not responsible or worthwhile.

Example: a weekly leadership update

Consider a weekly leadership update prepared from approved project notes.

FactorScoreEvidence
Repetition and workload2It happens every week and takes meaningful preparation time.
Process clarity2The normal format, sources, and examples are clear.
Reviewability2A manager can compare the draft with the source notes quickly.
Cost of error1A poor internal draft can be contained before it is shared.
Human value1Leadership judgment shapes the final interpretation and message.
Total8Initial route: Automate candidate.

The score creates an initial route. The flags and the work itself set the final route.

Initial score8 out of 10: Automate candidate
Safety checkThe report shapes leadership communication.
Final route for nowAccelerate: AI drafts. A manager approves.

A sensible first design is:

  1. AI gathers approved notes and prepares a draft.
  2. The accountable manager checks facts, context, and message.
  3. The manager decides what is shared.
  4. The team records preparation time and correction effort, then improves, expands, or stops.

The score found a strong opportunity. The override kept a human checkpoint where meaning and accountability still matter.

This is an illustrative example, not a client result.

Decide what happens next

When the screen is complete, record five things:

  1. Initial route: Automate, Accelerate, or Stay Human.
  2. Safety check: Which safety flags change the route?
  3. Final route for now: What may AI actually do?
  4. Next evidence: What fact or example would improve the decision?
  5. Human control: What must remain with a named person?

Do not begin with software. First decide whether the task is clear, worthwhile, checkable, and safe enough to test.

If the task moves into a test, keep a basic record of what AI did, what information it used, who reviewed the result, and what happened next. AI Audit Logs explains the minimum practical system.

Where the Sidekick Assessment begins

The screen creates a list of tasks worth examining. The broader Sidekick Assessment helps the team choose priorities, set safeguards, run small tests, and build a practical plan.

Your team does not need to diagnose itself. It needs to describe the work accurately.

Your team describes the work
  • the task, trigger, and useful result;
  • the owner, frequency, and information used;
  • approved examples, exceptions, and pain points.
Sidekick applies the judgment
  • validate what actually happens;
  • compare value, readiness, and risk;
  • design the test, safeguards, and measures.

That is the division of work. Your team supplies the facts about how the work happens. Sidekick helps make the important decisions about what to change and how to test it safely.

The Government of Canada's AI Readiness Scorecard is another useful way to examine one problem. It asks whether AI suits the work, whether the current problem is worth solving, whether the process and information are ready, and whether early compliance concerns are visible. It does not replace formal privacy, risk, ethics, security, or legal review.

Start with what is stuck

Use the worksheet on one recurring task. If the score and flags point in the same direction, you have a practical next step. If they conflict, that conflict is exactly what a deeper assessment should resolve.

The goal is not a longer list of AI ideas. It is a clear, bounded decision about the work AI may run, the work it may prepare, and the work that stays with a person.

The five-factor logic and score bands are adapted from The AI Daily Brief's AI Deputization Audit. Sidekick adds the route names, safety flags, final decision about what AI may do, and the broader readiness check. Sidekick is not affiliated with The AI Daily Brief. The broader readiness check is informed by the Government of Canada's AI Readiness Scorecard. This is a working Sidekick method. It is not a client result or permission to automate.