Inside an Enterprise AI Platform Team: How One Company Scaled Agentic Automation Across the Business
Step inside one company's enterprise AI platform team and learn how they scaled agentic automation from scattered pilots to 50 governed, production workflows.
The Setup: Why One Company Built a Central AI Platform Team
I sat down with the platform lead at a mid-market enterprise we'll call Meridian — a ~4,000-person financial services and logistics firm. Eighteen months ago, the company had no AI platform team. It had scattered pilots, shadow AI, and no clear owner.
The change started with three uncomfortable questions. Who chooses the models? Who guards the data? Who pays the bill?
The answers were the same: nobody. That was the problem.
Meridian decided to stand up a central enterprise AI platform team. Their goal was to move agentic automation from experiments into daily operations. Today, more than 50 business processes run on governed AI agents.
This is their story. It is a representative case, built from the practices we saw work across several enterprises.
Key insight — the enterprise AI platform team centralizes agentic automation into one governed, measurable layer rather than leaving it scattered across pilots.
The Operating Model: Who Sits on the Platform Team
The platform lead was clear about one thing. The team's shape mattered more than the models. They built a hub-and-spoke model.
At the center sits the AI platform team. It has six roles:
- Platform lead — owns the roadmap and the budget.
- ML engineer — manages models, serving, and cost.
- Agent engineer — builds and tunes the agents and prompts.
- Governance lead — owns policy, guardrails, and audits.
- Product owner — translates business needs into workflows.
- Finance liaison — tracks spend and reports ROI.
Around the hub, domain squads in each business unit nominate agent owners. These owners understand their process deeply. They act as the bridge between the platform and their team.
The split prevented the platform team from becoming a bottleneck. It also prevented each business unit from building its own silo.
Key insight — the platform team owns the rails, domain squads own the use cases, and neither can go without the other.
The Reference Architecture for Agentic Automation
Meridian's stack follows a pattern that many teams repeat. It has four layers.
Layer one is the model gateway. It routes each request to the right model based on cost, latency, and capability. It also caches repeated calls.
Layer two is the orchestrator. It plans the steps an agent takes and coordinates tools and sub-agents. The orchestration layer routes tasks across models and tools in a single governed flow.
Layer three is the guardrail layer. This is where policy lives. It masks PII before data reaches a model. It blocks high-risk actions. agent governance enforces guardrails and approvals at every step.
Layer four is observability. Every step is logged. Every outcome is measured. Nothing runs blind.
The key design choice was placement. The platform sits between the models and the business tools. Nothing calls a model directly. Nothing touches a business system without passing the platform.
This one decision gave Meridian control. It turned the platform into a chokepoint for governance and a single place to measure everything. the platform team reduces fragmented shadow AI pilots into one audited path.
Governance and Evaluation: Keeping Agents Honest
The hardest part was not building agents. It was trusting them. Meridian built two mechanisms to earn that trust.
The first was evaluation. Before any agent shipped, it had to pass a golden set of test cases. These were real, historical requests with known correct answers. The team tracked accuracy, cost, and latency against a baseline. evaluation measures golden-set accuracy and cost against a release bar.
Scores were enforced, not suggested. If an agent failed the bar, it did not go live.
The second was human-in-the-loop approval. High-risk actions needed a person to click confirm. Sending money, contacting a customer, or touching records all required approval.
Key insight — half the win is the guardrail. The other half is knowing when to ask a human. Set the line high for irreversible actions.
Every agent also kept an audit trail. The team could replay any decision and see exactly which prompts and tools ran. That traceability built trust with compliance.
Scaling From 5 to 50 Agents
Scaling exposed real failures. Meridian's first few agents ran in isolation. When they grew the fleet, problems appeared.
The biggest failure was overlapping agents. Two agents sometimes handled the same ticket. Work got duplicated. Sometimes it got dropped.
The fix was a gateway and registry pattern. Every agent registered its purpose and scope. The orchestrator used that registry to route work to exactly one owner. scaling agents requires a reusable gateway and registry so each unit has a single owner.
They also added explicit ownership rules. Each process had a single responsible agent. Edge cases routed to a human, not to guesswork.
Key insight — do not scale agents by copying them. Scale by giving each one a clear job and a single owner. Reuse shrinks cost.
The result was a clear growth path. Meridian went from 5 tightly supervised agents to 50 governed ones. Reliability, not raw speed, was the goal.
The ROI That Got Measured
Meridian measured ROI carefully. They refused to report inflated numbers. They counted only what they could verify.
Across the governed workflows, they tracked several metrics. Monthly routine hours fell by roughly a third. Average resolution time per request dropped. Manual handoffs shrank. ROI counts hours saved and cycle time rather than hype.
Cost was controlled too. The model gateway routed cheaper models to easy tasks. Caching cut repeated token spend. The platform reported cost per resolved request.
The honest caveat: ROI was uneven. Some workflows paid for themselves in weeks. Others never justified their run cost and were retired.
Key insight — measure per workflow, not per platform. Some agents are winners. Some are experiments that should end. Kill the losers fast.
The practice of measuring per workflow became a discipline. It kept the platform credible with the finance team. It also prevented "zombie agents" that ran forever without value.
Advice for Teams Just Starting
The platform lead offered five pieces of advice for teams beginning this journey.
Start small. Pick one high-value, low-risk workflow. Prove the pattern before scaling.
Measure before and after. You cannot show ROI without a baseline.
Build governance early. It is painful to add after the fact.
Own the rails, not the use cases. Let domain teams drive adoption.
Expect uneven outcomes. Some agents will fail. That is the cost of learning.
Perhaps the most important lesson was about patience. This was not a two-week project. It took a quarter to build trust, and a year to see broad impact.
Expert Q&A
Q: What roles do you need on an AI platform team, in order of importance? A: Start with a platform lead and an agent engineer. Add a governance lead before you grow beyond a few agents. A finance liaison matters the moment costs are visible. Scale the rest as your fleet grows.
Q: How do you stop agents from overlapping? A: Use a registry where every agent declares its scope and owner. Route work through an orchestrator that reads that registry. Add a rule: each process has one responsible agent. Edge cases go to a human, never to guesswork.
Q: What is the single most common mistake teams make? A: Skipping evaluation before release. Teams ship an agent, watch it fail in production, and blame the model. A golden-set evaluation and an audit trail catch most issues before they reach a real customer.
Q: How do you keep ROI honest? A: Measure per workflow, not per platform. Count only what you can verify — hours saved, cycle time, error rates. Retire any workflow whose run cost exceeds its value. That discipline is what keeps finance on your side.
Q: How often should you re-evaluate an agent that is already live? A: Re-run the golden set after every prompt, model, or tool change. Schedule a light monthly pass for drift. Agent behavior degrades silently, so continuous evaluation beats a one-time check.
Q: What is the best first workflow to automate? A: Pick one that is high-frequency, low-risk, and easy to measure. A routine data-entry or triage task is ideal. Avoid anything irreversible while your guardrails are still young.
Q: How do you control LLM costs as the fleet grows? A: Route cheap models to easy tasks through the gateway. Cache repeated calls. Set per-workflow cost budgets and alert on outliers. Retire workflows whose cost exceeds their value.
If you are building agentic automation inside your own enterprise, follow the work on algorithmine.com. We publish practical playbooks on platform teams, governance, and scaling AI agents every week.