AI Engineering
tutorial
Featured

Nobody Can Tell You How Many Agents You Are Running

A year of shipping agents leaves most teams with a fleet nobody can inventory: no owner, no declared budget, no record of which tools each one can reach. That is agent sprawl, and it shows up as a surprise invoice and an incident with no name on it. The fix is boring and it works: an agent registry. A typed manifest every agent must declare, a CI gate that blocks anything unregistered or over-budget, and a reconciliation job that flags an agent doing more than it said it would.

Viral Ruparel
11 min read
Share:

You can usually find the exact moment a team gets religion about this. Someone in finance forwards the monthly model-provider invoice, it is up forty percent, and the question comes back down the chain: which agent did that. Nobody can answer. There is a Slack channel where people announced agents when they launched them, a couple of repos, a shared spreadsheet that stopped being accurate in the spring, and a strong collective sense that there are "about a dozen" agents in production. The real number turns out to be thirty-one, four of them are still running against a deprecated internal API, and two of them nobody will claim.

That is agent sprawl, and if your team shipped agents through 2026 you almost certainly have it. It is not a moral failing. It is the predictable result of a technology that made it a one-afternoon job to stand up something autonomous that calls tools and spends money, combined with the fact that nobody's first agent came with a place to register it. The industry noticed too: a wave of "agent management" products launched this quarter whose entire pitch is inventorying the agents you already run and attaching cost and ownership to each one. You do not need to buy one of those to fix this. The core of it is a registry you can build yourself in an afternoon, and building it yourself means it sits inside your own deploy path instead of scanning from the outside.

The business problem: a fleet you cannot enumerate

Sprawl is not a cost problem or a security problem or a reliability problem. It is all three, wearing the same disguise, which is that you cannot produce a list.

The cost side is the one that gets attention because it arrives as a number. When spend is not attributed per agent, you cannot tell a runaway retry loop from legitimate growth, you cannot chargeback to the team that caused the spend, and you cannot set a limit because you do not know what normal looks like for any individual agent. The invoice is a single blob.

The reliability side is quieter and worse. When an agent misbehaves at three in the morning, the first thing the on-call engineer needs is a name to page. An agent with no registered owner is an incident with no responder, so it either gets left running because nobody dares turn off something they do not understand, or it gets killed and takes a workflow somebody depended on down with it.

The security side is the one that ends careers. A reasonable question from any auditor is "list every automated identity that can read the customers table." If your answer involves grepping through repos and asking around, you do not have a governance story, you have an archaeology project. Agents are exactly the identities that question is about, because each one holds credentials and calls tools on its own initiative.

The fix for all three is the same artifact: a registry that can produce the list, kept honest by the fact that nothing reaches production without being in it.

Step one: make every agent declare a manifest

The registry is only as good as what agents are forced to tell it, so start with the schema. A manifest is a small, typed record that each agent ships alongside its code. Keep it minimal at first, because a manifest nobody can fill out in two minutes is a manifest people route around.

from dataclasses import dataclass, field
from enum import Enum

class Lifecycle(str, Enum):
    ACTIVE = "active"
    DEPRECATED = "deprecated"   # still runs, but flagged for removal
    RETIRED = "retired"         # must not run; kept for audit history

@dataclass(frozen=True)
class AgentManifest:
    id: str                     # stable, e.g. "billing-dunning-agent"
    owner: str                  # a real person or team email, not "platform"
    purpose: str                # one line: what it does and why it exists
    model: str                  # the model id it runs on
    tools: frozenset[str]       # tool / scope names it is permitted to call
    daily_budget_usd: float     # hard ceiling on spend per day
    lifecycle: Lifecycle = Lifecycle.ACTIVE

    def validate(self) -> list[str]:
        """Return a list of problems; empty means the manifest is well-formed."""
        problems = []
        if "@" not in self.owner:
            problems.append("owner must be a routable email, not a vague label")
        if len(self.purpose.split()) < 4:
            problems.append("purpose is too thin to be useful in an incident")
        if self.daily_budget_usd <= 0:
            problems.append("daily_budget_usd must be a positive ceiling")
        if not self.tools:
            problems.append("declare tools explicitly, even if the set is empty")
        return problems

The fields are chosen to answer the three questions that come up in real incidents and reviews: who owns this, what can it touch, and what is it allowed to spend. The owner check being an email address is not pedantry. The single most common way a registry rots is a column full of "platform team" and "data" that page nobody, so the schema refuses those on the way in.

Step two: a CI gate so the registry cannot go stale

A registry that is a document is a spreadsheet with extra steps, and it will drift within a month. What makes it real is that it is enforced at the one chokepoint every agent passes through, which is your deploy pipeline. The rule is simple and unpopular for about a week: if an agent is not in the registry with a valid manifest, it does not ship.

import sys

def load_registry(path: str) -> dict[str, AgentManifest]:
    """Your registry can be a JSON/YAML file in a repo or a table in a DB.
    A file in version control is the cheapest thing that works, and it gives
    you history and code review on every change for free."""
    ...

def gate(deployed_agent_ids: set[str], registry: dict[str, AgentManifest]) -> int:
    failures = []

    # 1. Nothing deploys unless it is registered.
    for agent_id in deployed_agent_ids:
        if agent_id not in registry:
            failures.append(f"{agent_id}: deploying but not in the registry")

    # 2. Every manifest in the registry must be well-formed.
    for agent_id, manifest in registry.items():
        for problem in manifest.validate():
            failures.append(f"{agent_id}: {problem}")

    # 3. A retired agent must not be in the deploy set at all.
    for agent_id in deployed_agent_ids:
        m = registry.get(agent_id)
        if m and m.lifecycle is Lifecycle.RETIRED:
            failures.append(f"{agent_id}: marked retired but still being deployed")

    for f in failures:
        print(f"BLOCKED  {f}", file=sys.stderr)
    return 1 if failures else 0

# In CI: sys.exit(gate(discover_deployed_agents(), load_registry("registry.yaml")))

Two things make this hold up in practice. First, keep the registry in version control as a plain file rather than a database to start with, so every change to it is a pull request with an owner and a reviewer, and you get the full history of who added which agent for free. Second, the gate blocks the deploy, it does not file a ticket. A warning that scrolls past in a log is not a control. A red build is.

The pushback you will hear is that this slows people down. It adds about ninety seconds to the first deploy of a new agent, which is the exact ninety seconds during which someone writes down who owns the thing and what it is for, while they still remember. That is the cheapest that information will ever be to capture. This is the same principle behind a policy-as-code gate on agent actions, moved one level up from the individual action to the agent itself.

Step three: attribute spend and enforce the budget at runtime

The manifest declares a daily_budget_usd, but a declaration is a wish until something checks it. The registry becomes load-bearing when your runtime reads the budget from it and refuses to keep spending past the ceiling. The mechanism is a thin wrapper that tags every model call with the agent id, records the cost, and fails closed when an agent blows through its limit.

import time

class BudgetExceeded(Exception):
    pass

class SpendTracker:
    """Per-agent daily spend, checked before each model call. In production
    this lives in Redis or your metrics store so it is shared across every
    process running the agent, not per-instance memory."""
    def __init__(self, registry: dict[str, AgentManifest]):
        self.registry = registry
        self._spent: dict[tuple[str, str], float] = {}  # (agent_id, day) -> usd

    def _day(self) -> str:
        return time.strftime("%Y-%m-%d", time.gmtime())

    def check_and_record(self, agent_id: str, cost_usd: float) -> None:
        manifest = self.registry.get(agent_id)
        if manifest is None:
            # An unregistered agent spending money is the exact thing the
            # registry exists to make impossible. Fail closed and page someone.
            raise BudgetExceeded(f"{agent_id} is not registered; refusing the call")

        key = (agent_id, self._day())
        projected = self._spent.get(key, 0.0) + cost_usd
        if projected > manifest.daily_budget_usd:
            raise BudgetExceeded(
                f"{agent_id} would exceed {manifest.daily_budget_usd:.2f}/day "
                f"(already spent {self._spent.get(key, 0.0):.2f})"
            )
        self._spent[key] = projected

# Wrap your model client so no call can skip the check:
# tracker.check_and_record(agent_id, estimate_cost(model, prompt_tokens))

Now the invoice is not a blob. Every dollar has an agent id on it, the runaway agent trips its own ceiling instead of the company's, and the number in the manifest stops being aspirational. This is the fleet-level version of the same discipline as per-tenant cost guardrails inside a single agent; the registry is just where the ceilings for every agent get to live in one place.

Step four: reconcile what agents declared against what they actually did

The last gap is the sneaky one. An agent can be registered, owned, and under budget, and still be lying, because the tools list in its manifest is what someone typed six months ago and the agent has since been extended to call three tools nobody added to the manifest. This is drift, and it defeats the security story quietly. The fix is a reconciliation job that compares the declared tool set against the tools the agent was actually observed calling, using the telemetry you already emit if you have any tracing at all.

def reconcile(registry: dict[str, AgentManifest],
              observed_tools: dict[str, set[str]]) -> list[str]:
    """observed_tools maps agent_id -> the set of tool names seen in the
    last window (pulled from your traces or logs). Flag any agent using a
    tool it never declared, and any declared tool that is never used."""
    findings = []
    for agent_id, manifest in registry.items():
        seen = observed_tools.get(agent_id, set())
        undeclared = seen - manifest.tools
        unused = manifest.tools - seen
        if undeclared:
            # Security-relevant: the agent can reach things it never disclosed.
            findings.append(f"{agent_id}: UNDECLARED tools in use: {sorted(undeclared)}")
        if unused:
            # Hygiene: over-broad grant, tighten the manifest and the credentials.
            findings.append(f"{agent_id}: declared but unused: {sorted(unused)}")
    return findings

Undeclared tool use is the finding that matters, because it is the difference between "this agent can read the customers table" being true in the registry and being true in reality. Run this on a schedule, route the undeclared cases to the owner the manifest names, and you have closed the loop: the registry is not just what people said, it is checked against what happened. If you are already emitting traces through OpenTelemetry's GenAI conventions, the observed-tools map is a single query away.

Pitfalls that will bite you

A few things go wrong often enough to name. The registry becoming a bureaucracy is the first: if the manifest grows to thirty fields and a review board, people will build agents that route around it, and you will be back to sprawl with extra paperwork. Keep the required set small and let the gate, not a meeting, do the enforcing.

The second is treating the file as the source of truth about what is running. The registry is the source of truth about what is allowed and owned. What is actually running comes from your deploy system and your telemetry, and the point of the CI gate and the reconciliation job is to force those two views to agree. When they disagree, that disagreement is the alert.

The third is forgetting retirement. Agents do not get decommissioned, they get forgotten, and a forgotten agent still holds credentials and still might wake up on a cron. A lifecycle field is only useful if something acts on retired, which is why the CI gate refuses to deploy a retired agent and your credential system should refuse to issue it a token.

The takeaway

Agent sprawl is not exotic. It is the same thing that happened with microservices and unmanaged cloud resources, and it has the same fix, which is a registry that nothing gets to skip. The engineering is deliberately unglamorous: a small typed manifest so every agent declares an owner, a budget, and the tools it can reach; a CI gate so the registry can never quietly go stale; a runtime that reads the budget and fails closed; and a reconciliation job that checks what agents actually did against what they said they would. Put those four together and you can answer the three questions that sprawl makes unanswerable, which are who owns this, what can it touch, and what is it allowed to spend. You will get asked all three, and the difference between a good day and a bad one is whether the answer is a query or an investigation.

If you have shipped more agents than you can currently list, that gap is usually mappable in a week and closeable soon after, without a platform migration. Book a consultation call and we can inventory what you are actually running and stand up a registry that keeps it honest.

Viral Ruparel

Generative AI consultant helping teams ship reliable LLM and agent systems in production.

Contact Viral about your AI project →

Frequently Asked Questions

What is agent sprawl?+

Agent sprawl is what happens when a team ships LLM agents faster than it catalogs them. Each agent starts as a script or a service someone stood up to solve one problem, and because there is no central record, the fleet grows without anyone tracking who owns each agent, what it is allowed to do, what it costs, or whether it is still in use. The symptoms are a cloud bill nobody can attribute, an incident where the offending agent has no owner, and a security review that cannot enumerate which agents can reach a given system. It is the same failure mode as microservice sprawl or unmanaged cloud resources, except the units are autonomous and spend money on every run.

What should an agent registry track?+

At minimum: a stable id, a human owner, a one-line purpose, the model and tools or scopes the agent is permitted to use, a cost budget per day, and a lifecycle status such as active, deprecated, or retired. Those fields are enough to answer the three questions that actually come up in an incident or a finance review: who owns this, what can it touch, and what is it allowed to spend. You add more later, but a registry that only answers those three is already worth building.

Is an agent registry the same as an agent framework?+

No. A framework such as LangChain or CrewAI helps you build a single agent or a multi-agent system. A registry sits above all of them and describes your fleet regardless of how each agent was built. You can run agents in three different frameworks and two languages and still have one registry that knows they exist, who owns them, and what they cost, because the registry keys off a manifest each agent declares, not off any framework's internals.

Related Articles

AI Engineering

Your Agent's Memory Is Full and Most of It Is Junk

You gave your agent persistent memory and it worked. Six months later the store is full of duplicates, contradictions, and vague paraphrases that crowd out the facts you actually need, and retrieval quietly gets worse every week. This is a garbage-collection problem, not a storage problem. Here is how to build the curation layer: gate what gets written, deduplicate and reconcile on the way in, decay what stops earning its slot, and measure whether the store is still healthy.

AI Engineering

Your Agent Fetches the Same Row Ten Times a Session

A long-running agent calls the same read tool over and over inside a single session, paying full latency and quota for answers it already had a few turns ago. A tool result cache fixes it, but the naive version ships stale data and quiet correctness bugs. Here is how to build one that classifies which tools are cacheable, deduplicates concurrent calls with singleflight, and invalidates reads the moment a write touches the same data.

AI Engineering

Your Agent Is Idle Most of the Time It's Working

A tool-using agent spends a surprising share of its wall clock doing nothing, just waiting for a network round trip while the model has already stalled. CPUs solved this problem decades ago with branch prediction. You can borrow the same trick: predict the next tool call, run it while the model is still reasoning, and commit the result if the guess was right. Done carefully it cuts latency by a third with zero effect on correctness. Done carelessly it fires off writes nobody asked for.