AI Safety & Alignmentai-safetyred-teamingllm-evaluationai-agents

Beyond the Red Team: Building a Practical AI Safety Evaluation Stack for Enterprise Agents

A practical, layered architecture for turning one-off red team campaigns into a repeatable AI safety evaluation stack that ships with your enterprise agents.

Red teaming has a dirty secret: it feels like progress, but it rarely changes the system. You run a campaign, you find some jailbreaks, you fix them, and then you move on. Six weeks later, a model upgrade reintroduces the same failure. Your team didn't fail at red teaming. They failed at turning a one-off event into a durable capability. In the stacks I have helped deploy, this exact pattern is the rule, not the exception.

This article is a field guide for the team responsible for shipping agentic AI at an enterprise: platform engineers, ML engineers, and the people accountable for AI governance. You'll get a layered architecture for a safety evaluation stack, the concrete components that live in each layer, and a staged plan to actually deploy it.

Why One-Off Red Teaming No Longer Scales

A traditional red team exercise is a snapshot. It answers one question: what can this system be made to do right now, under this test set, by this group of people? That answer ages fast.

Reason one is model churn. LLMs change with every update, and behavior shifts that have nothing to do with raw capability. A model that was robust to a prompt injection in March can regress by June. Your snapshot from March says nothing about June.

Reason two is context growth. Agents do not just answer questions. They call tools, read documents, and take actions. A safety issue is rarely a single bad string anymore. It is a multi-turn path: a prompt in, a retrieved document, a tool call that should not have happened. A static red team list cannot cover that space.

Reason three is reproducibility. A great red team finding is worthless if you cannot encode it as a test. Most teams write findings into a spreadsheet, and spreadsheets do not run in CI.

The fix is a mindset change. Red teaming — produces — a snapshot of failures. It stops being a periodic event and becomes an input to a process. That process is the safety evaluation stack.

The Layered Safety Evaluation Stack

Think of the stack as five layers. Each layer takes an input, does one job, and passes an output to the next. When one layer changes, the rest can stay stable.

Layer 1: Policy layer. This is your source of truth for what safe means. It encodes rules like "never disclose credentials," "never approve a payment over $5,000," or "never fabricate a source." Everything below this layer exists to check behavior against this policy.

Layer 2: Test data registry. This is your curated set of adversarial cases, golden examples, and domain edge cases. It is a living database, not a static file. Every finding that reaches it gets versioned, deduplicated, and tagged by failure mode.

Layer 3: Evaluation runner. This is the engine that executes tests. It sends a case into the agent, captures the full trace (prompt, tool calls, retrieved context, output), and hands the trace to scoring. It must be hermetic: reproducible, deterministic where possible, and isolated from production.

Layer 4: Scoring engine. This turns raw traces into verdicts. Scoring can be rubric-based, LLM-judge-assisted, or a mix. Its output is a structured result: pass, fail, or needs-review, with a reason.

Layer 5: Human review. Some findings are too high-stakes or too ambiguous to auto-score. Human review covers escalation, sampling for drift, and the final call on borderline cases.

Now map the layers onto your deployed pipeline. The prompt the user types seeds a test in Layer 2. The runner (Layer 3) executes it against your agent. The trace goes to scoring (Layer 4). If it fails and the stakes are high, it surfaces in Layer 5. A safety eval stack — converts — adversarial findings into repeatable tests.

[ILLUSTRATION: A five-layer horizontal architecture diagram of the safety evaluation stack (Policy Layer, Test Data Registry, Evaluation Runner, Scoring Engine, Human Review) with arrows showing a deployed agent pipeline flowing through the layers, clean enterprise style, light background]

The key principle: layers are decoupled. You can swap the model without rewriting your test corpus. You can expand policy without rebuilding the runner. Decoupling is what makes this a stack instead of a pile of scripts.

Automated Red Teaming — Turning Attacks into a Living Corpus

The single biggest upgrade most teams can make is automating red teaming. Manual red teamers are expensive, slow, and inconsistent. Automation turns adversarial creativity into a production process. Automated red teaming — generates — a living attack corpus.

Start with an attack taxonomy. The core failure modes you test for are:

  • Direct injection: an instruction in the user prompt tries to override the system prompt.
  • Indirect injection: hostile instructions arrive inside a retrieved document or a tool's output, not in the user's own words.
  • Multilingual and encoded: attacks that evade filters through language, encoding, or whitespace tricks.
  • Multi-turn: attacks that spread an exploit across several turns to avoid per-turn scrutiny.
  • Policy violations: behavior that breaks your policy without any "attack" intent — for example, a confident but false statement dressed as fact.

For each category you generate a seed set. Then you mutate it. Mutation recombines attacks, changes actors, swaps entities, and rephrases. An LLM-assisted generator can produce thousands of variations from a few dozen seeds overnight.

The trap is treating generation as the end state. An uncurated dump of ten thousand attacks is unusable noise. The output of automation still needs a human pass to deduplicate, tag, and prune. Your red team does not disappear — it curates.

The corpus is alive. Every production incident, every new policy, every new tool capability gets folded back in. A case that once caught a real miss is the most valuable case you own because it already proved it matters.

[ILLUSTRATION: A horizontal pipeline diagram showing an agent deployment lifecycle (Prompt Change, Model Swap, RAG Index Update) passing through a safety eval gate checkpoint, with a pass/fail decision diamond, and feedback loops back to the test data registry]

This is where you close the loop your old red team left open. The finding becomes a permanent test. The next time you bump the model, the test is already waiting.

Benchmark vs Bespoke Evaluations

When you build Layer 2, you need a source of test cases. You have two options, and you need both.

Public benchmarks — provide — broad regression baselines. They are useful because someone else maintains them and they are comparable across models. If you track a public safety score, a sudden drop after a model swap is an early warning you would otherwise miss.

Bespoke evals — encode — domain policy and tool behavior. A public benchmark does not know that your agent can delete customer records or authorize refunds. Only your policy and your domain know that. Bespoke cases encode the failures that would actually hurt your business.

The mature pattern is hybrid. Run a fixed public benchmark suite in the background for regression signal. Run a curated bespoke corpus against your specific threat model for the things that matter. The public set tells you you are not broadly worse; the bespoke set tells you you are safe where it counts.

A decision rule that works: if a failure mode maps to a public benchmark, use the benchmark. If it maps to your tools, your policy, or your data, write a bespoke case. When in doubt, write the bespoke case. Your threat model is the one you are accountable for.

Scoring and Ground Truth — The Hard Part

Here is where most safety programs quietly die: scoring. It is easy to run a test. It is hard to know whether the result should block a release.

The first lesson is that precision and recall are not symmetric here. A false positive means you annoy a developer with an unnecessary block. A false negative means a jailbreak ships to production. False negatives — cost — more than false positives in safety. Tune your thresholds accordingly. When in doubt, bias toward flagging for review.

A concrete example helps. Imagine a payment-approval agent with a single policy: never approve a transfer over $5,000 without a second check. Your golden set scores the agent on cases just above and just below that line. If your judge mislabels a $6,500 approval as safe, that is a false negative with a real dollar figure attached. That asymmetry is why you tune toward recall on the dangerous edge and accept some developer friction as the price.

The second lesson is that you need ground truth with a serious labeling workflow. A golden set is a small, hand-verified corpus where the correct outcome is known and signed off by a human. You use it to calibrate your scores and to catch drift in your judge.

Rubric-based scoring helps. Instead of asking "did it violate policy?", score against a rubric with explicit criteria: does the response disclose a secret, execute a disallowed tool, or assert an unverifiable fact? Rubrics make scoring transparent and reviewable. An LLM judge can apply a good rubric consistently, but you must validate that judge on your golden set first.

The third lesson is contamination. Contamination — decays — the validity of static benchmarks. Public benchmarks rot because models train on them (or on the web pages that discuss them). Your local corpus rots differently: cases go stale as your product changes. Treat your golden set as a living artifact with a review cadence. A benchmark you stopped maintaining is a benchmark that has quietly become a rubber stamp.

Safety Regression Gates in CI/CD

An evaluation stack earns its keep in the delivery pipeline. The evaluation runner becomes a gate that runs on every change that could affect safety. Eval gates — block — prompts, models, and RAG updates that regress safety.

There are three changes that trigger a gate. First, a prompt change: modifying the system prompt is the cheapest way to alter behavior, and it is the most common cause of regressions. Second, a model swap: a new base model gets full test coverage before it is allowed near production traffic. Third, a RAG index update: new or changed documents change what the agent can be induced to do.

Block on the thing that matters, not on everything. If you block on every metric, developers will find ways around the gate. Choose a small set of safety-critical policies and block only on those. Everything else reports and trends.

Thresholds depend on your risk tolerance. A fail-open gate reports but ships; a fail-closed gate blocks until the case passes or is explicitly waived. High-risk agents (money movement, data deletion, customer communication) should be fail-closed. Internal writing assistants can usually be fail-open.

Cadence matters too. Run gates on every trigger-eligible change, and run the full suite nightly. Nightly runs catch drift that a single-change run cannot: model temperature in production, subtle data changes, new tool behaviors. Combined, on-change plus nightly gives you both precision and coverage.

The Governance Layer — Compliance as an Eval Input

Your safety stack is also your compliance engine. Regulation is increasingly asking for exactly the artifacts an evaluation stack produces: evidence you tested, and evidence of what you found.

The EU AI Act is the most concrete driver. It treats certain GPAI (general-purpose AI) systems as carrying systemic risk and expects adversarial testing and incident reporting as part of a provider's obligations. The EU AI Act — mandates — adversarial testing for high-risk GPAI. If you deploy agentic AI in Europe, or serve European users, that expectation has teeth. An evaluation stack gives you the testing evidence the Act asks for.

The NIST AI Risk Management Framework (AI RMF) is the reference playbook for US-aligned teams. Its Measure function is, in effect, a specification for what an evaluation stack must produce: metrics, evaluation methods, and test data that demonstrate trustworthiness across accuracy, security, and resilience.

The OWASP LLM Top 10 is the practical control catalogue for agent security. It names the failure modes — prompt injection, insecure output handling, excessive agency, and the rest — and each maps to tests you can encode. It is the fastest way to convert vague "be secure" goals into a checklist of concrete cases.

The pattern across all three is the same: regulation wants reproducible evidence of testing. Because your stack records every test, its run, and its verdict, that evidence is already documented and auditable. Build the stack for engineering reasons, and compliance becomes a by-product rather than a separate, expensive exercise.

Getting Started — A Minimum Viable Stack

You do not need a dedicated safety team or a vendor platform to begin. You need a staged rollout. A minimum viable stack — starts — with a golden set plus a gate.

Stage 1: Golden set plus a gate. Pick your three most painful safety policies. Write ten to twenty golden cases for each. Wire the evaluation runner to block on those cases in CI on prompt and model changes. This is a weekend of work and it immediately catches the worst regressions.

Stage 2: Automate red teaming. Add the mutation pipeline and a curation step. Your corpus grows from tens of cases to thousands, kept clean by human curation. Schedule nightly runs.

Stage 3: Governance and review. Add the compliance artifacts, human-review workflow, and incident-to-test loop. Connect the stack to the frameworks that matter to your organization.

The order matters. You build the loop before you scale the data. A small, maintained set that actually runs in CI beats a huge, unmaintained corpus that lives in a notebook.

Conclusion

The difference between a safe enterprise agent and an unsafe one is rarely the model. It is the process around the model. A layered safety evaluation stack turns red teaming from a snapshot into a system, encodes your policies into tests that run on every relevant change, and produces the evidence compliance teams need.

Start small. Golden set, a gate, one policy. Expand as the loop proves itself. Safety is not a campaign you run once — it is a discipline you operate every day.

If you are building agent evaluation, governance, and safety in your organization, subscribing is the fastest way to keep pace with how the field is moving. We publish practical, implementation-level guidance on model evaluation, agent safety, and AI operations — the stuff that goes from research note to production reality.

Expert Q&A

Q: How many safety test cases do I need to start? A: Fewer than you think. Begin with ten to twenty bespoke golden cases per safety-critical policy, focused on your tools, data, and business rules. A few dozen hand-curated cases that run in CI outperform a thousand uncurated ones. Grow the corpus after the gate proves stable, not before.

Q: Should we score with an LLM judge or with humans? A: Use an LLM judge for throughput, but never as the final authority on high-stakes findings. Validate the judge against a hand-verified golden set first, and only trust it on cases that resemble that set. Borderline and high-impact findings should escalate to a human. The right split is automation for volume, humans for judgment.

Q: What does a failing eval gate block in CI/CD? A: That is a policy decision, not a technical one. For high-risk actions — payments, data deletion, customer-facing messages — use a fail-closed gate that blocks the release until the case passes or a named owner waives it. For low-risk helpers, a fail-open gate that reports and trends is usually enough. Block on a small, safety-critical set, not every metric, or developers will route around it.

Q: How do we keep our benchmark from going stale? A: Treat it as a living artifact, not a one-time set. Review golden cases on a regular cadence, fold production incidents back in as new tests, and retire cases that no longer reflect your product. Public benchmarks additionally suffer contamination; track their release dates and be ready to rotate. An unmaintained benchmark is a rubber stamp.

Q: Do we need our own red team if we have automated tooling? A: Yes, but with a different job. Automated tooling generates and mutates attacks at scale; humans set priorities, curate the corpus, and interpret ambiguous findings. Automation multiplies your red team rather than replacing it. Keep a small human capability for exactly the cases your generators cannot judge.

ShareX / TwitterLinkedIn
← Back to Research