How an Enterprise Deployed 10,000 AI Agents: A Practitioner Interview
A Fortune 500 platform lead explains how 10,000 AI agents run in production: orchestration, model routing, AgentOps, evals, and the failures that shaped it.
Running 10,000 AI agents in production is not a model problem. It is an operations problem wearing a model costume.
That is the single most consistent message from the platform lead at a Fortune 500 services company who agreed to walk us through their two-year journey from a 40-agent pilot to a fleet of more than 10,000 persistent AI agents handling internal operations and customer-facing workflows.
We are calling them "Dana," a director of AI platform engineering, at their request; the company is anonymized, but the architecture, numbers, and failures below are theirs. Figures are self-reported and directionally accurate rather than audited. What follows is a practitioner interview about the unglamorous work — orchestration, routing economics, evaluation gates, and incident response — that determined whether the deployment survived contact with reality.
"Nobody warns you that the LLM is the easy part. The hard part is that you now operate a distributed system where every node is nondeterministic, occasionally wrong, and occasionally lies — and you can't just restart it when it misbehaves. You have to design so that a wrong answer is contained, detected, and reversible."
The 10,000-Agent Deployment at a Glance
The fleet reached 10,000 agents in month 19. It runs across three business units: a customer support triage layer, an internal IT and HR service desk, and a back-office document processing pipeline. Roughly 60% of agents are long-lived and stateful (about 6,000), while the rest are ephemeral task runners spun up per job and torn down within minutes.
Peak concurrency sits between 2,800 and 3,400 simultaneous agent executions — a burstiness ratio of roughly one active execution per three governed units at peak, which reflects how many long-lived agents are idle at any moment and how short-lived the ephemeral runners are. The company tracks "agents" as configured, versioned, and governed units — not as raw API calls — which is why the count is stable rather than ballooning every sprint.
What 10,000 AI Agents Actually Means
Dana is blunt that the headline number is a governance artifact as much as a technical one. "If you count every request, we're at millions. If you count every prompt variant, it's unmanageable. We count what we can name, version, and page someone about." The operative test: an agent is a unit that has its own owner, a declared toolset, a model policy, an SLO, and a version lineage — and that gets invoked many times. That definition is what makes the fleet governable; it is also why the number is a floor, not a ceiling.
Each agent has an owner, a declared toolset, a model policy, and an SLO. Ownership is enforced, not aspirational: an agent without a named owner cannot be promoted to production.
Human-in-the-loop coverage varies by risk tier: low-risk internal agents run fully autonomous, while anything touching money, legal commitments, or external customer communication routes through human approval for a defined percentage of runs — between 5% and 100% depending on the workflow. The approval percentage is treated as a tunable cost lever, not just a safety control: raising it reduces error exposure but adds human minutes per run, so each workflow has a target coverage that balances the two.
The Problem They Set Out to Solve
The original mandate was unglamorous: reduce cost per resolution and stop the growth of headcount in tier-1 support. Baseline metrics before the program, measured on the same definitions used afterward: a median 14-hour first-response time, a 38% deflection rate on the existing rules-based chatbot, and a fully loaded cost per ticket of roughly $9.40.
The company set three targets: cut cost per resolution by half, raise self-service resolution above 60%, and redeploy — not eliminate — 200 support staff into higher-value roles. Inside 18 months they hit two of the three: cost per resolution and the self-service target. The redeployment target lagged, largely because retraining and role redesign took longer than the technology.
[ILLUSTRATION: A layered architecture diagram showing the orchestration layer at the top, model routing in the middle, and memory/tool/data boundaries at the bottom, with 10,000 agent nodes rendered as a dense grid.]
Inside the Architecture of an Enterprise AI Agent Fleet
The stack is deliberately boring. Dana's team chose proven distributed-systems primitives over AI-native frameworks wherever possible, because "we needed to debug at 3 a.m., not admire the abstraction."
Orchestration and Runtime Layer
A durable workflow engine handles scheduling, retries, and state persistence. Every agent invocation is a workflow step with idempotency keys, so a retried tool call cannot double-charge a customer or file a duplicate ticket — provided the downstream tool honors the key. Where a tool is not idempotent, the team wraps it with a reconciliation or compensating action rather than trusting the retry. Per-agent isolation runs in lightweight sandboxes with hard CPU, memory, and wall-clock ceilings.
Rate limiting is enforced at three levels: per agent, per tool, and per downstream system. "The first outage we caused wasn't the model," Dana recalls. "It was 400 agents hammering an internal API that was never designed for that concurrency." The fix was not throttling the model — it was a token-bucket limiter at the tool boundary plus backpressure that queues agent steps instead of dropping them.
Model Routing and Inference Economics
The fleet runs a tiered model mix: small models for classification, extraction, and routing; mid-tier models for drafting and summarization; frontier models reserved for complex reasoning and high-stakes decisions. A routing policy scores each request on complexity, risk, and latency budget before selecting a model. The routing decision itself is logged, so cost regressions can be traced to a policy change rather than guessed at.
Aggressive caching cut repeated-context inference by an estimated 30%, measured as the cache-hit rate on repeated prefixes and retrieved context blocks rather than asserted as a headline. Batch inference handles overnight document processing. The result: blended cost per agent-hour — including inference plus the platform overhead of orchestration, memory, sandboxing, and observability — dropped from about $1.80 in the pilot to $0.34 at fleet scale. Dana flags this as directional and highly workload-dependent, and points to the routing mix, caching, and batching as the dominant drivers rather than any single model change. As a sanity anchor, the team also tracks blended cost per resolved ticket, which fell in step with the agent-hour figure.
Memory, Tools, and Data Boundaries
Retrieval is tenant-scoped by default, with vector stores partitioned per business unit. Tool permissions are allow-listed per agent — no agent can call a tool it was not explicitly granted — and every tool call is logged with the agent ID, inputs, outputs, and a trace link.
PII handling runs through a redaction layer before anything reaches an external model provider; self-hosted inference paths skip redaction where policy allows. Dana is candid that redaction has a recall cost — some sensitive strings slip through — so the layer is evaluated like any other component, and its misses feed the same regression suite as the agents.
Data residency requirements forced a split between US and EU inference paths. Audit trails are immutable and retained for seven years, implemented as an append-only, hash-chained log store rather than a mutable database table, so tampering is detectable.
Security as a First-Class Concern
Prompt injection and tool abuse are treated as primary threats, not edge cases. Because agents read retrieved content and external inputs, any of those inputs can attempt to hijack behavior. Mitigations are layered: retrieved content is treated as untrusted, tool calls are validated against the agent's declared allow-list, high-impact tools require explicit confirmation or human approval, and outbound actions are checked against policy before execution. "Assume something will try to make your agent do the wrong thing," Dana says. "Then make sure the wrong thing is boring."
AgentOps: How to Run 10,000 Agents in Production
This is where Dana's team spends most of its engineering time. "We have more people on agentops than on agent development. That ratio surprised our leadership, and it's the correct one."
Observability and Tracing at Fleet Scale
Every agent run emits a trace with spans for planning, retrieval, tool calls, and generation, plus token accounting per span. Per-agent SLOs track success rate, p50/p95 latency, and cost per run. The on-call dashboard is intentionally sparse: error budget burn, cost anomalies, and stuck-workflow counts — three panels, not thirty. Everything else is a drill-down.
Evaluation, Regression, and Release Gates
Offline eval sets are curated per agent and expanded from production failures. Each failure that reaches production becomes a labeled test case, so the suite grows with the fleet rather than being authored once and forgotten. Evaluation blends human labels for high-stakes behavior with LLM-as-judge scoring for volume, and the team periodically audits judge agreement against human raters to keep the automated signal honest. To avoid overfitting to the eval set, they hold out a rotating sample of recent production traffic that is never used for tuning.
Canary rollouts send 5% of traffic to a new agent version, gated on statistical thresholds for success rate, latency, and cost. A regression beyond the gate triggers automatic rollback to the prior version; a human can also roll back manually. "We ship agent versions like software, because they are software," Dana says. "The prompt is part of the artifact."
The Failure Taxonomy
Dana's team keeps a running list of incident classes, and it is shorter than you would expect:
- Tool storms — many agents hitting one downstream system at once. Mitigated by tool-level rate limits and backpressure.
- Context blowup — retrieval returning too much, inflating cost and degrading quality. Mitigated by strict context budgets and relevance thresholds.
- Prompt injection via retrieved content — mitigated by treating retrieved text as untrusted and validating tool calls.
- Silent quality drift — outputs degrade without erroring, often after a model or data change. Caught by continuous eval on production samples, not by error rates.
"Most of our incidents are boring distributed-systems problems," Dana says. "The interesting ones are the silent ones, because nothing pages you until a customer does."
Cost and Capacity Governance
Cost is governed like any other resource: per-agent budgets, anomaly alerts, and a routing policy that is reviewed on a schedule. Capacity planning is driven by peak concurrency rather than agent count, since the two move independently. When a business unit wants to add agents, the question is not "how many?" but "what is the peak concurrent load, and who owns the downstream systems they will hit?"
What They Would Do Differently
Dana's advice to teams starting now is unromantic. Build the governance layer before the fleet, not after. Instrument cost and quality from the first agent. Treat evaluation as a product, not a chore. And resist the temptation to count agents as a proxy for value — the number that matters is cost per resolution, not headcount of bots.
"The fleet is the easy part to brag about. The eval suite, the on-call rotation, and the rollback path are what let you sleep. Nobody puts those in the press release, and they're the whole job."