Sidekick Orchestration
// The Sidekick Skim · August 29 - September 4, 2026

Retry the hard research job you stopped giving AI

Three providers announced stronger models for complex work. The useful business test is whether one research or analysis job that used to need too much correction can now produce a reviewable first pass.

In this edition
Read this if you read nothing else
This week's decision

Retry one job your team abandoned because AI needed too much correction, and compare the checked result with the old approach before changing a working process.

What changed

OpenAI announced Astra for a forthcoming release, while Anthropic and Google announced models they say improve long-running, tool-using and multi-step work.

What it means

A first pass that was previously too shallow, slow or expensive may now be useful for a document-heavy brief, supplier comparison or technical problem a person can check.

Safe to ignore

Do not rework every workflow around a release. A one-off chat or a task where a wrong answer has high cost can stay with the method that already works.

Only where it applies
  1. If your team has a research or analysis job it stopped giving AI15 minutes

    Retry one real job

    Choose a bounded job, such as comparing three suppliers from public materials or preparing a cited first draft from a fixed document set. Keep the source set and review standard unchanged. Compare correction time and decision usefulness with the prior approach.

    Limit: Start with a result a person can review. Do not give a newly tested system authority to send, publish, change records or make a commitment.
01

The new opportunity is a better first pass on difficult work

For owners who have a real job in mind, not a reason to chase a model release.

Capability updateProvider product and safety announcementsThree providers pointed to harder, longer workThe announcements describe systems that can keep working through more steps, use tools more effectively and handle more demanding reasoning than prior versions.Applies to: Owners with a recurring brief, comparison, proposal, investigation or analysis task that still begins from scattered material

OpenAI's September 1 Astra post described a forthcoming model for computer use, browsing, professional work and software engineering. Anthropic said Fable 5.1 improves coding, knowledge work and long-running problem solving. Google said Gemini 3.8 Flash improves software engineering, agentic tasks and multi-step reasoning. These are provider claims, not evidence that any one business task will now work well.

Take a job AI nearly handled before, give it stable source material and a clear review standard, then see whether it now produces a usable first pass.

Why it matters

Keep a short list of jobs AI did not handle well enough. When a material model release lands, retry the best candidate on the same inputs and judge the finished work, including correction time.

Sources: OpenAI: Path to Astra, September 1, 2026, Anthropic: Claude Fable 5.1 and Mythos 5.1, September 1, 2026, Google: Gemini 3.8 Flash and Flash Cyber, September 2, 2026

Evidence note: Sources checked September 14, 2026. This catch-up edition covers announcements made August 29 through September 4 and was source-checked September 14, 2026. Availability, eligibility, pricing and product behaviour are provider-described and can change. This edition does not establish that any product is suitable for a particular business or that a provider claim replaces a person's review.

// Get the next edition

Get the next Skim.

New editions by email when they clear the evidence and usefulness checks.

// Need help applying this decision?

Bring one decision. Leave with a next step.

Use the 30-minute call to work through that specific model, workflow, or risk choice.