The capability register · Autumn 2026

What AI can take off your desk.

A dated account of the work today’s AI systems can take from a 2–20 person business, how far to trust each job, and the decision that stays with you. Every entry names its evidence. None of it names a product to buy.

Verified September 29, 2026 · next review by December 31, 2026

Set the dial before the tool.

Each entry names the furthest setting we would defend today for a small business. Start one notch lower, and turn it only after a measured run of real cases says the work holds.

Act within limits
The system acts on routine cases inside written rules, spending limits, and permissions, and keeps a receipt for every action. Exceptions come to a person.
Draft for approval
The system prepares the work. A person approves it before anything reaches a customer, a record, or a payment.
Shadow test first
The system runs beside the current process on completed cases. Nothing it produces leaves the test.
Keep the decision human
A person makes the decision. AI may prepare only the administrative work around it.

Front desk

The phone, the inbox, and the calendar: the work that decides whether a customer feels answered.

01

Answer the phone and book the work

Voice agents answer every call, take messages, qualify callers, and book into open slots from your own script and calendar.

Act within limits

Inbound answering, messages, and bookings into open slots can run inside a written script with a transfer path. Outbound calls and texts stay at draft-for-approval until consent is documented.

Today
AI receptionists are a mature, commercially available category. They answer around the clock, take simultaneous calls, answer from the business's own information, take messages, transfer to a person, and book or reschedule through calendar integrations. Many also text a booking link while the caller is still on the line.
Where it breaks
A stale script produces confident wrong answers, calendar sync errors double-book, noisy lines garble details, and a badly built flow loops instead of transferring. No independent study yet measures failure rates for small-business receptionists, and the FTC has already settled with a seller that claimed its AI voice agent could replace customer-service staff.
Human gate
Prices beyond the published list, refunds and complaints, emergencies and safety calls, any outbound calling, and every change to the script the agent speaks from.
First test
Route only after-hours or overflow calls for two weeks. Read every transcript against what a person would have done, and call back anyone the agent mishandled.
Scoreboard
Answered calls, bookings kept, completed transfers, callbacks needed, and complaints, compared with the two weeks before.
Watch
A 2024 FCC ruling treats AI-generated voices as artificial voices under the TCPA, so outbound AI calls need prior express consent. Many states require every party's consent to record a call, and a growing number require telling people they are dealing with AI. Disclose both in the first sentence, and keep a path to a person.
02

Draft replies from your own records

Inbox and chat agents draft answers from order, account, and policy data, and can resolve routine questions inside written rules.

Draft for approval

Drafting from records saves real time today. Sending is earned one question type at a time, after a measured run needs no correction; order-status and opening-hours questions usually earn it first.

Today
Office suites draft replies with context from mail and documents, and several include it in plans a business already pays for. Customer-service agents resolve routine questions in chat and email from an approved set of answers, and the major help desks now price them per resolution instead of per seat.
Where it breaks
Drafts pull the wrong price, date, or promise from the wrong thread. Outdated answers produce confident wrong policy—the root of the Air Canada ruling, where the company was held to its chatbot's answer. Vendor resolution counts include customers who simply left, so they overstate real fixes. Email is also the main route for prompt injection.
Human gate
Refunds, credits, discounts, and policy exceptions; legal threats, chargebacks, and safety issues; any answer that is not in the approved set; and every change to that set.
First test
Draft replies to twenty already-answered messages and compare them with what was actually sent: facts, tone, and minutes of editing.
Scoreboard
Editing minutes per reply, factual errors, reopened conversations, and customer satisfaction—not the vendor's resolution count.
Watch
An instruction hidden inside an email is not a command, and an agent must never act on one. Keep an always-open path to a person; the business, not the bot, answers for what the bot says.
Evidence
  1. BC tribunal confirms companies remain liable for information provided by AI chatbotAmerican Bar Association, Business Law Today, accessed September 29, 2026
  2. Fin AI Agent outcomesIntercom, accessed September 29, 2026
  3. HubSpot's Customer Agent and Prospecting Agent: now you pay when the task is completeHubSpot, accessed September 29, 2026
03

Coordinate calendars and reminders

Scheduling assistants propose times, book, and remind inside the rules you set for hours, buffers, and travel.

Act within limits

Booking open slots under written rules is low-risk and reversible. Cancellations, exceptions, and anything that moves a client's committed time stay with a person.

Today
Booking pages and email scheduling assistants propose times, book, and reschedule within stated availability, and office suites now offer a booking page while the sender is still writing. Email reminders are routine.
Where it breaks
Hidden constraints—travel, staff skills, preparation time, a customer's preferences—make valid-looking slots unusable. Scheduling products also launch and close quickly, so the rules belong in your own documents, not only in the tool.
Human gate
Cancellations, conflicts, travel and staffing exceptions, and any change to time a client has already committed.
First test
Replay five completed scheduling threads and compare the proposed slots with the final calendar, including every constraint a person caught.
Scoreboard
Messages per booking, no-shows, conflicts, and missed constraints.
Watch
Automated texts from a business number generally require carrier registration, and marketing texts require consent. Honor opt-outs promptly and in whatever reasonable way the customer asks.
Evidence
  1. Gmail is entering the Gemini eraGoogle, accessed September 29, 2026
  2. Callie, the AI scheduling assistant: overviewCalendly, accessed September 29, 2026

Back office

Documents, books, and portals: the retyping and reconciling that happens after the sale.

04

Read documents into your systems

Models pull the fields out of invoices, orders, receipts, and forms, and hand anything uncertain to a person.

Act within limits

Extraction can write to a holding record when totals reconcile and every field clears a confidence check. Anything that fails a check goes to a person before it reaches the system of record.

Today
Current models read typed documents and many forms well, extract fields into structured records, and can be told to flag what they are unsure of. Accounting and document tools increasingly do this inside the product, though specialized document readers still outscore general chat assistants on complex layouts and tables in independent benchmarks.
Where it breaks
Poor scans, rotated or tiny images, unusual layouts, and handwriting produce confident-looking errors. Model makers themselves advise against tasks that need perfect precision without human oversight.
Human gate
Low-confidence fields, totals that do not reconcile, new vendors, and anything written to the ledger or the system of record.
First test
Extract ten already-processed documents and compare every field with what was entered by hand.
Scoreboard
Field accuracy on totals, dates, and identifiers; exceptions per batch; review minutes per document.
Watch
Invoices and vendor emails are a favored fraud route. A changed bank detail or an urgent invoice must never flow straight to payment.
Evidence
  1. Vision: limitationsAnthropic, accessed September 29, 2026
  2. OmniDocBench document-parsing benchmarkOpenDataLab (GitHub), accessed September 29, 2026
05

Prepare the books for review

Agents categorize transactions, match receipts, draft reconciliations and payment reminders, and stage it all for approval.

Draft for approval

Preparation is reliable and batchable. Money movement, bank-detail changes, and anything a tax professional signs stay with a person.

Today
Major accounting systems now offer official connections that let an agent read and write the books, and small-business assistants run month-end routines—categorizing, matching, reconciling, drafting reminders—while staging every send, post, or payment for the owner's approval.
Where it breaks
No independent study yet measures small-business AI bookkeeping accuracy, so vendor accuracy claims are not evidence. The same connections can delete invoices and journal entries when permissions are broad, and inbox-driven payment requests remain a classic fraud route.
Human gate
Moving money, changing any bank details, journal entries and period close, deleting or voiding transactions, tax filings, and new vendors.
First test
Run last month's close against a copy or test company, then compare every categorization and match with the version you accepted.
Scoreboard
Agreement with accepted categorizations, unreconciled items, reminder accuracy, and owner review minutes per week.
Watch
Grant read and draft permissions first and add write access one record type at a time. Keep a log of every change the agent proposed and every change you approved.
Evidence
  1. QuickBooks Online MCP serverIntuit (GitHub), accessed September 29, 2026
  2. Claude for Small Business launches new workflows, integrations, and training programsAnthropic, accessed September 29, 2026
06

Work a website or portal for you

Computer-use agents fill forms, pull reports, and work portals that have no connection—and stop before anything is submitted.

Shadow test first

Run it beside a person on completed work first. Once field-by-field accuracy holds, let it prepare the entry and stop at the submit button. It never pays, posts, or confirms on its own.

Today
Browser agents from several major AI companies can read a page, click, type, and fill forms with the owner's logins, and the leading ones are generally available on paid plans. Vendors describe them for portals with no integration, dashboard pulls, research across tabs, and CRM updates. By default they ask before purchases, posts, and other hard-to-reverse steps.
Where it breaks
Prompt injection is unsolved: text planted on a page or in an email can redirect an agent, and every major vendor says so. Complex tasks often need a second attempt, logins and security checks interrupt the work, and the products change fast—one major AI company retired its general agent and its browser within about a year of launch.
Human gate
Submitting orders or forms, payments, posts, bookings, anything that touches banking, legal, or medical data, and every login or password entry.
First test
Record three completed portal tasks, then have the agent prepare a fourth beside a person, in a browser profile with no saved payment method.
Scoreboard
Field accuracy, retries, minutes of supervision, and any instruction it tried to follow from a page.
Watch
Give the agent a separate browser profile with only the logins the job needs, keep automatic approval off except on sites you trust, and never paste a password into a chat.
Evidence
  1. Claude in Chrome is now generally availableAnthropic, accessed September 29, 2026
  2. Mitigating the risk of prompt injections in browser useAnthropic, accessed September 29, 2026
  3. Continuously hardening ChatGPT Atlas against prompt injectionOpenAI, accessed September 29, 2026

Growth

Marketing, reputation, and the new ways customers—and the agents working for them—find a business.

07

Draft marketing from real work

Models draft posts, newsletters, product copy, review replies, and images from your notes; a person approves every claim.

Draft for approval

Drafting is fast and safe. Publishing carries the business's name and legal exposure, so a person approves every claim, testimonial, price, and public reply.

Today
Drafting copy and images in a house voice is one of the most dependable uses of current AI, and office suites, design tools, and commerce platforms build it in.
Where it breaks
Fluent drafts invent results, statistics, and endorsements. Voice drifts without examples, generated images can misrepresent the product, and a public reply to a complaint can concede liability or reveal private details.
Human gate
Every factual claim, price, testimonial, before-and-after result, and image before publication, and every reply to a public complaint.
First test
Draft three posts from one real, approved customer outcome and compare editing time with writing from scratch.
Scoreboard
Editing minutes per piece, claims needing correction, and qualified replies or inquiries—not posting volume.
Watch
A 2024 federal rule bars fake reviews and testimonials, including AI-generated ones attributed to people who do not exist or never used the product, with civil penalties for knowing violations.
Evidence
  1. Federal Trade Commission announces final rule banning fake reviews and testimonialsFederal Trade Commission, accessed September 29, 2026
08

Get described accurately by AI assistants

Assistants now answer many buying questions before anyone clicks. Plain answers on your own pages decide what they say about you.

Act within limits

Monitoring what assistants say about you is read-only and safe to automate. Changing what your site says stays with a person.

Today
Search engines and chat assistants answer many buying questions directly and cite a handful of sources. Google says its AI features need no special files or markup—a page must simply be indexed and eligible for a snippet—and some assistants now call local businesses on a searcher's behalf to check prices, a setting owners can switch off in their business listing.
Where it breaks
No assistant publishes how it chooses what to cite, answers vary from one run to the next, and outdated facts get repeated. Thin pages written for machines can cost the search traffic a business already has.
Human gate
Every factual claim, price, and policy the site publishes, and the decision to opt in or out of automated calls and agent channels.
First test
Ask three assistants the ten questions customers ask before buying. Record whether and how the business appears, then fix the one page with the clearest missing answer.
Scoreboard
Questions answered accurately at the next monthly check, referral visits from assistants, and qualified inquiries.
Watch
Search crawlers and training crawlers are separate: blocking an assistant's search crawler can remove you from its answers, while refusing training does not. Keep prices and hours identical everywhere they are published, and treat any 'AI visibility score' as a noisy sample, not a ranking.
Evidence
  1. AI features and your websiteGoogle Search Central, accessed September 29, 2026
  2. Does Anthropic crawl data from the web, and how can site owners block the crawler?Anthropic, accessed September 29, 2026
  3. Manage automated calls and texts from GoogleGoogle Business Profile Help, accessed September 29, 2026
09

Sell to agents that shop

Shopping agents read product catalogs and check out through new open protocols. Clean product data is the low-regret move.

Shadow test first

Prepare the product data and let your commerce platform carry the protocols. Treat agent orders as a pilot channel with its own review of returns and disputes until volume proves itself.

Today
Open protocols for agent checkout now exist—one stewarded by OpenAI and Stripe, another co-developed by Google, Shopify, and major retailers—and commerce platforms expose store catalogs to agents. New payment credentials keep a person approving each purchase. Real volume is still thin: the most visible in-chat checkout was reportedly scaled back in early 2026 after weak conversion.
Where it breaks
Prices and stock drift between the feed and the store, agents pick the wrong variant or shipping option, the protocols are still changing, and product pages and reviews can carry text written to manipulate shopping agents.
Human gate
Opting into agent channels and their terms, agent-specific prices or discounts, return exceptions, and chargeback responses.
First test
Audit your twenty best sellers for complete titles, identifiers, price, stock, shipping, and returns, then ask three assistants to find and compare them.
Scoreboard
Catalog completeness, assistant-referred visits and orders, and return and dispute rates on agent orders.
Watch
Do not forecast revenue from a channel without history, and do not accidentally block the shopping agents you want through robots or bot-protection settings.
Evidence
  1. Agentic Commerce ProtocolOpenAI and Stripe (GitHub), accessed September 29, 2026
  2. Universal Commerce ProtocolUCP contributors (GitHub), accessed September 29, 2026
  3. OpenAI shifts checkout plans in its agentic commerce strategyDigital Commerce 360, accessed September 29, 2026

Management

Meetings, research, memory, and small tools: the work that keeps one operator from becoming the message bus.

10

Turn meetings into owned tasks

Notetakers transcribe, summarize, and draft follow-ups within minutes; a person confirms who owns what.

Act within limits

Capturing and summarizing is low-risk when everyone knows the notetaker is present. Decisions, owners, and anything sent to a client are confirmed by a person.

Today
Major video platforms include AI notes in paid plans, notify participants when notes are being taken, and some now capture in-person meetings. Summaries and action lists arrive minutes after the call.
Where it breaks
Transcription can invent whole phrases—a peer-reviewed study found fabricated text in about 1% of transcriptions from one widely used model—and speakers and owners get misattributed.
Human gate
Recording outside calls, sending AI-written follow-ups to clients, keeping or deleting recordings, and keeping notetakers out of HR, legal, and medical conversations.
First test
Run the notetaker on one internal, non-sensitive meeting and compare its decisions, owners, and dates with the team's own notes.
Scoreboard
Action items captured correctly, misattributions, and follow-up time.
Watch
Announce the notetaker and get consent at the start; many states require every party's consent to record. Lawsuits filed in 2025 target notetakers that recorded people who never signed up, so turn off vendor training on your data and speaker identification unless you have consent.
Evidence
  1. Use Gemini to take notes in Google MeetGoogle Meet Help, accessed September 29, 2026
  2. Careless Whisper: speech-to-text hallucination harmsACM FAccT 2024, accessed September 29, 2026
  3. Class-action lawsuit accuses Otter AI of secretly recording private work conversationsNPR, accessed September 29, 2026
11

Research and brief you on a schedule

Research agents read the public web on a schedule—prices, suppliers, grants, competitors—and deliver a sourced brief.

Act within limits

Reading and summarizing is read-only. The risk is acting on a wrong conclusion, so decisions stay with a person and material claims are checked against the source.

Today
The major assistants run multi-step research that reads many sources and returns a cited summary, and they can now run tasks on a schedule and deliver the result without being asked again.
Where it breaks
Summaries overstate what a source says, sources go stale, and a page written to manipulate agents can steer the brief.
Human gate
Any decision, purchase, or public claim based on the brief, and any change to suppliers or prices.
First test
Ask for a brief on one question you already researched by hand and compare what it found, missed, and misread.
Scoreboard
Claims that hold up against their sources, useful items per brief, and research minutes saved.
Watch
Require a link for every claim and read the source before acting on anything material.
Evidence
  1. Cowork is now ClaudeAnthropic, accessed September 29, 2026
12

Answer from your own playbooks

Assistants connected to your documents answer staff and agent questions from your own procedures, with the source attached.

Act within limits

Answering internal questions with a cited source is low-risk. Deciding what becomes official procedure—and what enters company memory at all—stays with a person.

Today
Assistants connect to shared drives, email, and business systems through permissioned connectors. A procedure can now be written once as a skill—an open format that dozens of assistants and coding tools can load—so agents follow the same method every time.
Where it breaks
Old or conflicting documents produce confident wrong answers, a connector can expose files to people who should not see them, and a skill keeps running yesterday's process after the business changes. Skills are instructions an agent will follow, so install them only from sources you trust.
Human gate
What enters the playbooks, what is retired, who can see what, and any answer that goes to a customer.
First test
Load one written procedure and ask it the ten questions a new hire asks. Check the source behind every answer.
Scoreboard
Answers with a correct source, questions it could not answer, and stale documents found.
Watch
Connect only the folders the job needs. A connector can see everything its account can see.
Evidence
  1. Equipping agents for the real world with Agent SkillsAnthropic, accessed September 29, 2026
  2. Agent Skills open standardagentskills (GitHub), accessed September 29, 2026
13

Build the small tool you keep wishing for

Coding agents and app builders let an owner build a calculator, tracker, or dashboard in an afternoon.

Draft for approval

Build freely for internal, read-only use. Anything that touches customer data, payments, or the public internet gets a security review before launch.

Today
App builders and coding agents turn a plain description into a working internal tool—a quote calculator, an order tracker, a dashboard over a spreadsheet—and can deploy it the same day.
Where it breaks
Generated apps can ship with open databases, missing access controls, or secrets in the code: one popular app builder produced apps whose data anyone could read until a 2025 fix. Telling an agent not to touch production is not a control; permissions are.
Human gate
Access controls, customer data, payments, public launch, and where secrets are stored.
First test
Rebuild one spreadsheet you maintain by hand as a small internal tool, with no customer data, and use it for a week.
Scoreboard
Minutes saved per week, defects found, and whether anyone still uses it a month later.
Watch
Never paste API keys or passwords into a builder prompt, and keep production data out of prototypes.
Evidence
  1. CVE-2025-48757: insufficient row-level security in generated appsCVE Program, accessed September 29, 2026
  2. OWASP GenAI Top 10 for LLM Applications, 2026 editionOWASP GenAI Security Project (GitHub), accessed September 29, 2026

Keep these decisions human.

Some decisions carry legal duties, a person's livelihood, or money that cannot be recovered. AI may prepare the work around them—scheduling, document assembly, reminders, summaries—but the decision and the accountability stay with a person. Hiring is the fastest-moving area: New York City requires bias audits of automated hiring tools, Illinois has required notice since January 2026 for employers of any size, and Colorado's replacement law adds notice and explanation duties from January 2027.

  • Hiring, firing, promotion, and pay.
  • Credit, lending, insurance, and housing eligibility.
  • Medical, legal, and tax judgments.
  • Moving money and changing bank details.
  • Anything irreversible done in a customer's name.
  1. Automated employment decision tools (Local Law 144)NYC Department of Consumer and Worker Protection, accessed September 29, 2026
  2. Illinois anti-discrimination law to address AI goes into effect January 1, 2026National Law Review, accessed September 29, 2026
  3. SB26-189: Automated decision-making technologyColorado General Assembly, accessed September 29, 2026
  4. Careful adoption of agentic AI servicesCISA and international partners, accessed September 29, 2026

Rules that hold at every desk.

  1. 01

    Start with the AI you already pay for.

    Check what your email suite, accounting software, CRM, and phone system already include before buying another subscription. Several now bundle drafting, summaries, and agents into existing plans.

  2. 02

    Grant the least access that finishes the job.

    Read before write, one system at a time, a named owner who can stop the agent, and no shared passwords. Joint guidance from CISA and its Five Eyes partners in 2026 asks for the same: least privilege, human approval at decision points, and readable logs.

  3. 03

    Treat outside content as evidence, not orders.

    Emails, web pages, documents, and reviews can carry planted instructions, and no current system is immune. Limit what an agent can reach, and keep sends and payments behind a person.

  4. 04

    Keep the receipt.

    Every action an agent takes should leave evidence in the system of record—a message ID, an order number, a ledger entry—and a log a person can read.

  5. 05

    Measure before you believe a number.

    Resolution rates, answer rates, and hours-saved figures are marketing until your own baseline confirms them. Self-reported speedups are unreliable too: in a controlled 2025 trial, experienced developers using AI took longer while believing they were faster.

  6. 06

    Disclose, and offer a person.

    Tell customers when they are dealing with AI and when a call is recorded, and keep a quick route to a human. Disclosure rules are spreading state by state and have applied to EU customers since August 2, 2026; honesty is the durable default.

  7. 07

    Write the method around the job, not the product.

    AI products launch, merge, and close within a year. Keep your procedure, examples, and quality bar in your own documents so the next tool can pick them up.

What changed this quarter.

  1. August 2, 2026

    AI disclosure became law for EU customers

    The EU AI Act's transparency duties took effect: people must be told when they are interacting with an AI system, and deepfakes must be labeled. A July amendment delayed the high-risk rules, including hiring tools, to December 2027, but not disclosure.

    European Commission: Transparency obligations under Article 50 of the AI Act
  2. August 4, 2026

    Customers asked for a human exit

    Gartner reported that 87% of customers say companies using generative AI for service must provide access to a human; a September follow-up found only 27% would try a chatbot again after a bad experience.

    Gartner: Gartner survey finds 87% of customers say companies using GenAI for customer service must provide access to a human agent
  3. August 9, 2026

    A major agent browser shut down

    OpenAI retired its standalone agent browser and moved browser work into its main apps, after also withdrawing its general agent mode. Procedures written around one product had to be rebuilt; procedures written around the job did not.

    OpenAI Help Center: Evolving Atlas into ChatGPT for browser-based agentic work
  4. August 26, 2026

    Browser agents reached general availability

    Anthropic made its Chrome agent available on every paid plan, with published prompt-injection results and confirmation before purchases and posts. Screen-driving agents are now a mainstream tool, not a demo.

    Anthropic: Claude in Chrome is now generally available
  5. September 15, 2026

    Small-business agents arrived in the stack

    Within one week, assistant makers and CRM vendors shipped agent workflows built for small teams, with connectors to accounting, payroll, payments, and commerce systems and with sends and payments staged for owner approval by default.

    Anthropic: Claude for Small Business launches new workflows, integrations, and training programs

Pick one job. Run the test.

Start the free chatOpen the loop library →