The creation number is not the portfolio

On July 7, 2026, OpenAI published a MUFG customer story describing a phased ChatGPT Enterprise rollout to approximately 35,000 Mitsubishi UFJ Bank employees. After custom GPT training, bank employees created more than 1,800 GPTs in four months. That is evidence that creation became accessible. It is not evidence that 1,800 workflows reached production, remained in use, passed an acceptance test, or paid for themselves.

Anthropic supplied a louder creation signal on June 23. Its Claude Tag announcement says an internal version now creates 65% of its product team's code. The figure is self-reported in a vendor launch post. Anthropic does not define how the 65% share was measured or report acceptance, rewrite, rejection, or repair rates, and the commercial Claude Tag product launched in beta for Enterprise and Team customers on Slack.

The point is not that either figure is useless. The point is that creation has stopped being the scarce event. A company can now accumulate agents, custom GPTs, skills, scheduled tasks, and generated software faster than it can decide which ones deserve maintenance. The new operational smell is not an empty AI catalogue. It is a crowded one.

Every agent is a small operating liability

An agent does not need to be customer-facing to become an operating dependency. It may carry instructions, company knowledge, credentials, connectors, scheduled work, model assumptions, evaluation cases, and a route into another system. Someone eventually changes a policy, data source, API, model, or workflow. The agent remains cheerfully configured for the old company.

On July 14, OpenAI's guidance on managing AI investments told enterprise leaders to treat AI work as a portfolio. It recommends clear ownership, real-task evaluations, cost per accepted outcome, governance before scaling, and funding that follows maturity from exploration through validation to production. This is investment language because proliferation creates capital-allocation questions, even when the initial artifact cost eleven minutes and no purchase order.

An unowned agent is therefore not free. It consumes review attention, security surface, duplicated context, support time, and occasionally a budget. It can also preserve an obsolete answer with excellent consistency. The portfolio needs a default state beyond still visible in the sidebar.

Build the registry before the museum

The minimum control is an agent registry. One row should identify the agent or skill, the funded job it performs, the internal or external customer, the accountable owner, the approving human, and its lifecycle stage: experiment, validation, production, quarantined, or retired.

The same row should record its active version, deployment surfaces, permitted tools and data, credentials or service identities, model route, dependencies, scheduled triggers, evaluation suite, last passing evaluation, last meaningful use, accepted-job volume, cost per accepted outcome, exception rate, human review minutes, and next review or expiry date. Add fields for overlapping agents and the artifact this one supersedes. Otherwise new version becomes a synonym for additional version.

Finally, give retirement its own evidence. A tombstone should record why the agent was removed, when triggers and credentials were disabled, what replaced it, where its last known-good version and evaluations live, and who can authorize restoration. A graveyard is not a folder full of mystery ZIP files. It is a controlled record proving that dead machinery no longer has keys.

Evals decide what earns another quarter

Anthropic's living Skills for enterprise documentation, accessed July 16, makes the lifecycle unusually explicit. It recommends an internal registry with purpose, owner, version, dependencies, and evaluation status. It calls for evaluations in isolation and alongside existing skills, periodic reruns, version pinning, rollback, consolidation when skills overlap, and deprecation when failures persist or the workflow disappears.

The guide recommends three to five representative queries per skill, including cases where it should trigger, should not trigger, and faces ambiguity. That is a useful minimum, not proof of production quality. A portfolio review should add cases from real exceptions and score whether the output cleared the workflow's external acceptance standard.

The same documentation notes that API requests support at most eight Skills per request. That is a surface-specific request limit, not a universal recommendation that an enterprise own only eight skills. The broader warning is more durable: skill metadata competes for attention, overlapping descriptions can damage selection, and consolidation should be earned through coexistence evaluations rather than performed because the catalogue looks untidy.

Hold the cull and keep the tombstone

On a cadence matched to the work, freeze a registry extract and require each owner to bring evidence. Daily operating agents may deserve a monthly review; a quarterly reporting agent can be judged after several cycles. For every agent, choose one disposition: keep because it clears the quality and economics bar; repair because the job remains valuable but performance has drifted; consolidate because another agent covers the same work; quarantine because permissions or outputs are unsafe; or retire because demand, ownership, or the workflow itself has disappeared.

The portfolio scorecard should show active agents, dormant agents, percentage with named owners, percentage with current evaluations, accepted jobs, first-pass acceptance, cost and review minutes per accepted job, exception rate, duplicate-agent clusters, expired credentials, overdue reviews, and retirements completed. Creation count may remain on the page if someone enjoys decorative numbers.

For retirement, stop schedules and inbound routes first. Revoke credentials and connector access. Remove the agent from role bundles and approved catalogues. Preserve the last version, evaluation results, decision record, and replacement path. Then watch for failed calls or abandoned downstream work that proves the dependency map was incomplete. Deletion without dependency observation is how a quiet cleanup becomes a lively incident.

The operator takeaway

Before the next agent receives shared credentials, a recurring schedule, or production data, require a registry row and an expiry date. Experiments can be easy to create and deliberately temporary. Production status should require a named customer, an acceptance test, an owner, a permission boundary, and evidence that the workflow earns its cost.

For a first audit, inventory every shared agent, GPT, skill, scheduled task, and agent-built internal app. Start with anything unowned, unevaluated, unused, duplicated, or connected to sensitive systems. Classify each item, revoke what no longer needs access, and assign the survivors a review date. The job is not complete when the catalogue is shorter. It is complete when the remaining work has owners, current evidence, and a known way to stop.

A buyer should ask for both the active registry and the graveyard. Compare created agents with agents used on their intended cadence, agents with current evaluations, duplicate workflows, unowned credentials, and retired artifacts whose triggers were actually disabled. If nobody can explain the job, show the accepted result, and name the next review date, the agent has not earned another quarter.

Cheap creation is useful. Cheap permanence is not.

If nobody can explain why an agent is alive, this is not a philosophical problem. It is a retirement ticket.

sources
  1. How to manage AI investments in the agentic eraOpenAI, accessed July 16, 2026
  2. MUFG aims to become AI-native with OpenAIOpenAI, accessed July 16, 2026
  3. Introducing Claude TagAnthropic, accessed July 16, 2026
  4. Skills for enterpriseAnthropic / Claude Platform Docs, accessed July 16, 2026