The capability register · Autumn 2026
What AI can take off your desk.
A dated account of the work today’s AI systems can take from a 2–20 person business, how far to trust each job, and the decision that stays with you. Every entry names its evidence. None of it names a product to buy.
Verified September 29, 2026 · next review by December 31, 2026
Set the dial before the tool.
Each entry names the furthest setting we would defend today for a small business. Start one notch lower, and turn it only after a measured run of real cases says the work holds.
- Act within limits
- The system acts on routine cases inside written rules, spending limits, and permissions, and keeps a receipt for every action. Exceptions come to a person.
- Draft for approval
- The system prepares the work. A person approves it before anything reaches a customer, a record, or a payment.
- Shadow test first
- The system runs beside the current process on completed cases. Nothing it produces leaves the test.
- Keep the decision human
- A person makes the decision. AI may prepare only the administrative work around it.
Front desk
The phone, the inbox, and the calendar: the work that decides whether a customer feels answered.
Answer the phone and book the work
Voice agents answer every call, take messages, qualify callers, and book into open slots from your own script and calendar.
Inbound answering, messages, and bookings into open slots can run inside a written script with a transfer path. Outbound calls and texts stay at draft-for-approval until consent is documented.
- Today
- AI receptionists are a mature, commercially available category. They answer around the clock, take simultaneous calls, answer from the business's own information, take messages, transfer to a person, and book or reschedule through calendar integrations. Many also text a booking link while the caller is still on the line.
- Where it breaks
- A stale script produces confident wrong answers, calendar sync errors double-book, noisy lines garble details, and a badly built flow loops instead of transferring. No independent study yet measures failure rates for small-business receptionists, and the FTC has already settled with a seller that claimed its AI voice agent could replace customer-service staff.
- Human gate
- Prices beyond the published list, refunds and complaints, emergencies and safety calls, any outbound calling, and every change to the script the agent speaks from.
- First test
- Route only after-hours or overflow calls for two weeks. Read every transcript against what a person would have done, and call back anyone the agent mishandled.
- Scoreboard
- Answered calls, bookings kept, completed transfers, callbacks needed, and complaints, compared with the two weeks before.
- Watch
- A 2024 FCC ruling treats AI-generated voices as artificial voices under the TCPA, so outbound AI calls need prior express consent. Many states require every party's consent to record a call, and a growing number require telling people they are dealing with AI. Disclose both in the first sentence, and keep a path to a person.
- Evidence
- FCC confirms that TCPA applies to AI technologies that generate human voicesFederal Communications Commission, accessed September 29, 2026
- Air AI and its owners will be banned from marketing business opportunities to settle FTC chargesFederal Trade Commission, accessed September 29, 2026
- Gartner survey finds 87% of customers say companies using GenAI for customer service must provide access to a human agentGartner, accessed September 29, 2026
Draft replies from your own records
Inbox and chat agents draft answers from order, account, and policy data, and can resolve routine questions inside written rules.
Drafting from records saves real time today. Sending is earned one question type at a time, after a measured run needs no correction; order-status and opening-hours questions usually earn it first.
- Today
- Office suites draft replies with context from mail and documents, and several include it in plans a business already pays for. Customer-service agents resolve routine questions in chat and email from an approved set of answers, and the major help desks now price them per resolution instead of per seat.
- Where it breaks
- Drafts pull the wrong price, date, or promise from the wrong thread. Outdated answers produce confident wrong policy—the root of the Air Canada ruling, where the company was held to its chatbot's answer. Vendor resolution counts include customers who simply left, so they overstate real fixes. Email is also the main route for prompt injection.
- Human gate
- Refunds, credits, discounts, and policy exceptions; legal threats, chargebacks, and safety issues; any answer that is not in the approved set; and every change to that set.
- First test
- Draft replies to twenty already-answered messages and compare them with what was actually sent: facts, tone, and minutes of editing.
- Scoreboard
- Editing minutes per reply, factual errors, reopened conversations, and customer satisfaction—not the vendor's resolution count.
- Watch
- An instruction hidden inside an email is not a command, and an agent must never act on one. Keep an always-open path to a person; the business, not the bot, answers for what the bot says.
- Evidence
- BC tribunal confirms companies remain liable for information provided by AI chatbotAmerican Bar Association, Business Law Today, accessed September 29, 2026
- Fin AI Agent outcomesIntercom, accessed September 29, 2026
- HubSpot's Customer Agent and Prospecting Agent: now you pay when the task is completeHubSpot, accessed September 29, 2026
Coordinate calendars and reminders
Scheduling assistants propose times, book, and remind inside the rules you set for hours, buffers, and travel.
Booking open slots under written rules is low-risk and reversible. Cancellations, exceptions, and anything that moves a client's committed time stay with a person.
- Today
- Booking pages and email scheduling assistants propose times, book, and reschedule within stated availability, and office suites now offer a booking page while the sender is still writing. Email reminders are routine.
- Where it breaks
- Hidden constraints—travel, staff skills, preparation time, a customer's preferences—make valid-looking slots unusable. Scheduling products also launch and close quickly, so the rules belong in your own documents, not only in the tool.
- Human gate
- Cancellations, conflicts, travel and staffing exceptions, and any change to time a client has already committed.
- First test
- Replay five completed scheduling threads and compare the proposed slots with the final calendar, including every constraint a person caught.
- Scoreboard
- Messages per booking, no-shows, conflicts, and missed constraints.
- Watch
- Automated texts from a business number generally require carrier registration, and marketing texts require consent. Honor opt-outs promptly and in whatever reasonable way the customer asks.
- Evidence
- Gmail is entering the Gemini eraGoogle, accessed September 29, 2026
- Callie, the AI scheduling assistant: overviewCalendly, accessed September 29, 2026
Back office
Documents, books, and portals: the retyping and reconciling that happens after the sale.
Read documents into your systems
Models pull the fields out of invoices, orders, receipts, and forms, and hand anything uncertain to a person.
Extraction can write to a holding record when totals reconcile and every field clears a confidence check. Anything that fails a check goes to a person before it reaches the system of record.
- Today
- Current models read typed documents and many forms well, extract fields into structured records, and can be told to flag what they are unsure of. Accounting and document tools increasingly do this inside the product, though specialized document readers still outscore general chat assistants on complex layouts and tables in independent benchmarks.
- Where it breaks
- Poor scans, rotated or tiny images, unusual layouts, and handwriting produce confident-looking errors. Model makers themselves advise against tasks that need perfect precision without human oversight.
- Human gate
- Low-confidence fields, totals that do not reconcile, new vendors, and anything written to the ledger or the system of record.
- First test
- Extract ten already-processed documents and compare every field with what was entered by hand.
- Scoreboard
- Field accuracy on totals, dates, and identifiers; exceptions per batch; review minutes per document.
- Watch
- Invoices and vendor emails are a favored fraud route. A changed bank detail or an urgent invoice must never flow straight to payment.
- Evidence
- Vision: limitationsAnthropic, accessed September 29, 2026
- OmniDocBench document-parsing benchmarkOpenDataLab (GitHub), accessed September 29, 2026
Prepare the books for review
Agents categorize transactions, match receipts, draft reconciliations and payment reminders, and stage it all for approval.
Preparation is reliable and batchable. Money movement, bank-detail changes, and anything a tax professional signs stay with a person.
- Today
- Major accounting systems now offer official connections that let an agent read and write the books, and small-business assistants run month-end routines—categorizing, matching, reconciling, drafting reminders—while staging every send, post, or payment for the owner's approval.
- Where it breaks
- No independent study yet measures small-business AI bookkeeping accuracy, so vendor accuracy claims are not evidence. The same connections can delete invoices and journal entries when permissions are broad, and inbox-driven payment requests remain a classic fraud route.
- Human gate
- Moving money, changing any bank details, journal entries and period close, deleting or voiding transactions, tax filings, and new vendors.
- First test
- Run last month's close against a copy or test company, then compare every categorization and match with the version you accepted.
- Scoreboard
- Agreement with accepted categorizations, unreconciled items, reminder accuracy, and owner review minutes per week.
- Watch
- Grant read and draft permissions first and add write access one record type at a time. Keep a log of every change the agent proposed and every change you approved.
- Evidence
- QuickBooks Online MCP serverIntuit (GitHub), accessed September 29, 2026
- Claude for Small Business launches new workflows, integrations, and training programsAnthropic, accessed September 29, 2026
Work a website or portal for you
Computer-use agents fill forms, pull reports, and work portals that have no connection—and stop before anything is submitted.
Run it beside a person on completed work first. Once field-by-field accuracy holds, let it prepare the entry and stop at the submit button. It never pays, posts, or confirms on its own.
- Today
- Browser agents from several major AI companies can read a page, click, type, and fill forms with the owner's logins, and the leading ones are generally available on paid plans. Vendors describe them for portals with no integration, dashboard pulls, research across tabs, and CRM updates. By default they ask before purchases, posts, and other hard-to-reverse steps.
- Where it breaks
- Prompt injection is unsolved: text planted on a page or in an email can redirect an agent, and every major vendor says so. Complex tasks often need a second attempt, logins and security checks interrupt the work, and the products change fast—one major AI company retired its general agent and its browser within about a year of launch.
- Human gate
- Submitting orders or forms, payments, posts, bookings, anything that touches banking, legal, or medical data, and every login or password entry.
- First test
- Record three completed portal tasks, then have the agent prepare a fourth beside a person, in a browser profile with no saved payment method.
- Scoreboard
- Field accuracy, retries, minutes of supervision, and any instruction it tried to follow from a page.
- Watch
- Give the agent a separate browser profile with only the logins the job needs, keep automatic approval off except on sites you trust, and never paste a password into a chat.
- Evidence
- Claude in Chrome is now generally availableAnthropic, accessed September 29, 2026
- Mitigating the risk of prompt injections in browser useAnthropic, accessed September 29, 2026
- Continuously hardening ChatGPT Atlas against prompt injectionOpenAI, accessed September 29, 2026
Growth
Marketing, reputation, and the new ways customers—and the agents working for them—find a business.
Draft marketing from real work
Models draft posts, newsletters, product copy, review replies, and images from your notes; a person approves every claim.
Drafting is fast and safe. Publishing carries the business's name and legal exposure, so a person approves every claim, testimonial, price, and public reply.
- Today
- Drafting copy and images in a house voice is one of the most dependable uses of current AI, and office suites, design tools, and commerce platforms build it in.
- Where it breaks
- Fluent drafts invent results, statistics, and endorsements. Voice drifts without examples, generated images can misrepresent the product, and a public reply to a complaint can concede liability or reveal private details.
- Human gate
- Every factual claim, price, testimonial, before-and-after result, and image before publication, and every reply to a public complaint.
- First test
- Draft three posts from one real, approved customer outcome and compare editing time with writing from scratch.
- Scoreboard
- Editing minutes per piece, claims needing correction, and qualified replies or inquiries—not posting volume.
- Watch
- A 2024 federal rule bars fake reviews and testimonials, including AI-generated ones attributed to people who do not exist or never used the product, with civil penalties for knowing violations.
- Evidence
- Federal Trade Commission announces final rule banning fake reviews and testimonialsFederal Trade Commission, accessed September 29, 2026
Get described accurately by AI assistants
Assistants now answer many buying questions before anyone clicks. Plain answers on your own pages decide what they say about you.
Monitoring what assistants say about you is read-only and safe to automate. Changing what your site says stays with a person.
- Today
- Search engines and chat assistants answer many buying questions directly and cite a handful of sources. Google says its AI features need no special files or markup—a page must simply be indexed and eligible for a snippet—and some assistants now call local businesses on a searcher's behalf to check prices, a setting owners can switch off in their business listing.
- Where it breaks
- No assistant publishes how it chooses what to cite, answers vary from one run to the next, and outdated facts get repeated. Thin pages written for machines can cost the search traffic a business already has.
- Human gate
- Every factual claim, price, and policy the site publishes, and the decision to opt in or out of automated calls and agent channels.
- First test
- Ask three assistants the ten questions customers ask before buying. Record whether and how the business appears, then fix the one page with the clearest missing answer.
- Scoreboard
- Questions answered accurately at the next monthly check, referral visits from assistants, and qualified inquiries.
- Watch
- Search crawlers and training crawlers are separate: blocking an assistant's search crawler can remove you from its answers, while refusing training does not. Keep prices and hours identical everywhere they are published, and treat any 'AI visibility score' as a noisy sample, not a ranking.
- Evidence
- AI features and your websiteGoogle Search Central, accessed September 29, 2026
- Does Anthropic crawl data from the web, and how can site owners block the crawler?Anthropic, accessed September 29, 2026
- Manage automated calls and texts from GoogleGoogle Business Profile Help, accessed September 29, 2026
Sell to agents that shop
Shopping agents read product catalogs and check out through new open protocols. Clean product data is the low-regret move.
Prepare the product data and let your commerce platform carry the protocols. Treat agent orders as a pilot channel with its own review of returns and disputes until volume proves itself.
- Today
- Open protocols for agent checkout now exist—one stewarded by OpenAI and Stripe, another co-developed by Google, Shopify, and major retailers—and commerce platforms expose store catalogs to agents. New payment credentials keep a person approving each purchase. Real volume is still thin: the most visible in-chat checkout was reportedly scaled back in early 2026 after weak conversion.
- Where it breaks
- Prices and stock drift between the feed and the store, agents pick the wrong variant or shipping option, the protocols are still changing, and product pages and reviews can carry text written to manipulate shopping agents.
- Human gate
- Opting into agent channels and their terms, agent-specific prices or discounts, return exceptions, and chargeback responses.
- First test
- Audit your twenty best sellers for complete titles, identifiers, price, stock, shipping, and returns, then ask three assistants to find and compare them.
- Scoreboard
- Catalog completeness, assistant-referred visits and orders, and return and dispute rates on agent orders.
- Watch
- Do not forecast revenue from a channel without history, and do not accidentally block the shopping agents you want through robots or bot-protection settings.
- Evidence
- Agentic Commerce ProtocolOpenAI and Stripe (GitHub), accessed September 29, 2026
- Universal Commerce ProtocolUCP contributors (GitHub), accessed September 29, 2026
- OpenAI shifts checkout plans in its agentic commerce strategyDigital Commerce 360, accessed September 29, 2026
Management
Meetings, research, memory, and small tools: the work that keeps one operator from becoming the message bus.
Turn meetings into owned tasks
Notetakers transcribe, summarize, and draft follow-ups within minutes; a person confirms who owns what.
Capturing and summarizing is low-risk when everyone knows the notetaker is present. Decisions, owners, and anything sent to a client are confirmed by a person.
- Today
- Major video platforms include AI notes in paid plans, notify participants when notes are being taken, and some now capture in-person meetings. Summaries and action lists arrive minutes after the call.
- Where it breaks
- Transcription can invent whole phrases—a peer-reviewed study found fabricated text in about 1% of transcriptions from one widely used model—and speakers and owners get misattributed.
- Human gate
- Recording outside calls, sending AI-written follow-ups to clients, keeping or deleting recordings, and keeping notetakers out of HR, legal, and medical conversations.
- First test
- Run the notetaker on one internal, non-sensitive meeting and compare its decisions, owners, and dates with the team's own notes.
- Scoreboard
- Action items captured correctly, misattributions, and follow-up time.
- Watch
- Announce the notetaker and get consent at the start; many states require every party's consent to record. Lawsuits filed in 2025 target notetakers that recorded people who never signed up, so turn off vendor training on your data and speaker identification unless you have consent.
- Evidence
- Use Gemini to take notes in Google MeetGoogle Meet Help, accessed September 29, 2026
- Careless Whisper: speech-to-text hallucination harmsACM FAccT 2024, accessed September 29, 2026
- Class-action lawsuit accuses Otter AI of secretly recording private work conversationsNPR, accessed September 29, 2026
Research and brief you on a schedule
Research agents read the public web on a schedule—prices, suppliers, grants, competitors—and deliver a sourced brief.
Reading and summarizing is read-only. The risk is acting on a wrong conclusion, so decisions stay with a person and material claims are checked against the source.
- Today
- The major assistants run multi-step research that reads many sources and returns a cited summary, and they can now run tasks on a schedule and deliver the result without being asked again.
- Where it breaks
- Summaries overstate what a source says, sources go stale, and a page written to manipulate agents can steer the brief.
- Human gate
- Any decision, purchase, or public claim based on the brief, and any change to suppliers or prices.
- First test
- Ask for a brief on one question you already researched by hand and compare what it found, missed, and misread.
- Scoreboard
- Claims that hold up against their sources, useful items per brief, and research minutes saved.
- Watch
- Require a link for every claim and read the source before acting on anything material.
- Evidence
- Cowork is now ClaudeAnthropic, accessed September 29, 2026
Answer from your own playbooks
Assistants connected to your documents answer staff and agent questions from your own procedures, with the source attached.
Answering internal questions with a cited source is low-risk. Deciding what becomes official procedure—and what enters company memory at all—stays with a person.
- Today
- Assistants connect to shared drives, email, and business systems through permissioned connectors. A procedure can now be written once as a skill—an open format that dozens of assistants and coding tools can load—so agents follow the same method every time.
- Where it breaks
- Old or conflicting documents produce confident wrong answers, a connector can expose files to people who should not see them, and a skill keeps running yesterday's process after the business changes. Skills are instructions an agent will follow, so install them only from sources you trust.
- Human gate
- What enters the playbooks, what is retired, who can see what, and any answer that goes to a customer.
- First test
- Load one written procedure and ask it the ten questions a new hire asks. Check the source behind every answer.
- Scoreboard
- Answers with a correct source, questions it could not answer, and stale documents found.
- Watch
- Connect only the folders the job needs. A connector can see everything its account can see.
- Evidence
- Equipping agents for the real world with Agent SkillsAnthropic, accessed September 29, 2026
- Agent Skills open standardagentskills (GitHub), accessed September 29, 2026
Build the small tool you keep wishing for
Coding agents and app builders let an owner build a calculator, tracker, or dashboard in an afternoon.
Build freely for internal, read-only use. Anything that touches customer data, payments, or the public internet gets a security review before launch.
- Today
- App builders and coding agents turn a plain description into a working internal tool—a quote calculator, an order tracker, a dashboard over a spreadsheet—and can deploy it the same day.
- Where it breaks
- Generated apps can ship with open databases, missing access controls, or secrets in the code: one popular app builder produced apps whose data anyone could read until a 2025 fix. Telling an agent not to touch production is not a control; permissions are.
- Human gate
- Access controls, customer data, payments, public launch, and where secrets are stored.
- First test
- Rebuild one spreadsheet you maintain by hand as a small internal tool, with no customer data, and use it for a week.
- Scoreboard
- Minutes saved per week, defects found, and whether anyone still uses it a month later.
- Watch
- Never paste API keys or passwords into a builder prompt, and keep production data out of prototypes.
- Evidence
- CVE-2025-48757: insufficient row-level security in generated appsCVE Program, accessed September 29, 2026
- OWASP GenAI Top 10 for LLM Applications, 2026 editionOWASP GenAI Security Project (GitHub), accessed September 29, 2026
Keep these decisions human.
Some decisions carry legal duties, a person's livelihood, or money that cannot be recovered. AI may prepare the work around them—scheduling, document assembly, reminders, summaries—but the decision and the accountability stay with a person. Hiring is the fastest-moving area: New York City requires bias audits of automated hiring tools, Illinois has required notice since January 2026 for employers of any size, and Colorado's replacement law adds notice and explanation duties from January 2027.
- Hiring, firing, promotion, and pay.
- Credit, lending, insurance, and housing eligibility.
- Medical, legal, and tax judgments.
- Moving money and changing bank details.
- Anything irreversible done in a customer's name.
- Automated employment decision tools (Local Law 144)NYC Department of Consumer and Worker Protection, accessed September 29, 2026
- Illinois anti-discrimination law to address AI goes into effect January 1, 2026National Law Review, accessed September 29, 2026
- SB26-189: Automated decision-making technologyColorado General Assembly, accessed September 29, 2026
- Careful adoption of agentic AI servicesCISA and international partners, accessed September 29, 2026
Rules that hold at every desk.
- 01
Start with the AI you already pay for.
Check what your email suite, accounting software, CRM, and phone system already include before buying another subscription. Several now bundle drafting, summaries, and agents into existing plans.
- 02
Grant the least access that finishes the job.
Read before write, one system at a time, a named owner who can stop the agent, and no shared passwords. Joint guidance from CISA and its Five Eyes partners in 2026 asks for the same: least privilege, human approval at decision points, and readable logs.
- 03
Treat outside content as evidence, not orders.
Emails, web pages, documents, and reviews can carry planted instructions, and no current system is immune. Limit what an agent can reach, and keep sends and payments behind a person.
- 04
Keep the receipt.
Every action an agent takes should leave evidence in the system of record—a message ID, an order number, a ledger entry—and a log a person can read.
- 05
Measure before you believe a number.
Resolution rates, answer rates, and hours-saved figures are marketing until your own baseline confirms them. Self-reported speedups are unreliable too: in a controlled 2025 trial, experienced developers using AI took longer while believing they were faster.
- 06
Disclose, and offer a person.
Tell customers when they are dealing with AI and when a call is recorded, and keep a quick route to a human. Disclosure rules are spreading state by state and have applied to EU customers since August 2, 2026; honesty is the durable default.
- 07
Write the method around the job, not the product.
AI products launch, merge, and close within a year. Keep your procedure, examples, and quality bar in your own documents so the next tool can pick them up.
What changed this quarter.
- August 2, 2026
AI disclosure became law for EU customers
The EU AI Act's transparency duties took effect: people must be told when they are interacting with an AI system, and deepfakes must be labeled. A July amendment delayed the high-risk rules, including hiring tools, to December 2027, but not disclosure.
European Commission: Transparency obligations under Article 50 of the AI Act - August 4, 2026
Customers asked for a human exit
Gartner reported that 87% of customers say companies using generative AI for service must provide access to a human; a September follow-up found only 27% would try a chatbot again after a bad experience.
Gartner: Gartner survey finds 87% of customers say companies using GenAI for customer service must provide access to a human agent - August 9, 2026
A major agent browser shut down
OpenAI retired its standalone agent browser and moved browser work into its main apps, after also withdrawing its general agent mode. Procedures written around one product had to be rebuilt; procedures written around the job did not.
OpenAI Help Center: Evolving Atlas into ChatGPT for browser-based agentic work - August 26, 2026
Browser agents reached general availability
Anthropic made its Chrome agent available on every paid plan, with published prompt-injection results and confirmation before purchases and posts. Screen-driving agents are now a mainstream tool, not a demo.
Anthropic: Claude in Chrome is now generally available - September 15, 2026
Small-business agents arrived in the stack
Within one week, assistant makers and CRM vendors shipped agent workflows built for small teams, with connectors to accounting, payroll, payments, and commerce systems and with sends and payments staged for owner approval by default.
Anthropic: Claude for Small Business launches new workflows, integrations, and training programs