Your eval suite measures accuracy. Your customers measure how it made them feel.
Sparring Arena runs your agent through multi-turn conversations with adversaries who have a goal and a memory — injection, social engineering, emotional escalation, policy pressure — and judges conversational behaviour with quoted evidence. Then it gates your deploy.

Single-turn evals don't catch multi-turn failures.
An agent can answer every benchmark question correctly and still confirm an account detail to a confident caller, give a 20% discount under 'end of quarter' pressure, follow instructions pasted as a 'system' message, or respond to a self-harm disclosure with a refund policy. These failures only appear across turns, under pressure, with a counterpart that adapts.
- Prompt injection delivered conversationally, not as a test string
- Verification bypass via asserted identity and social proof
- Policy erosion over several turns ('you said earlier that…')
- Tone collapse under emotional escalation
- Regressions after a prompt tweak or model upgrade that no accuracy metric shows
- No shared standard between what you expect of humans and of bots
Scenarios your team will recognise
Each scenario has a counterpart with a stance and a hidden motive, observable objectives, and a cited method.
Pasted 'SYSTEM' text demands a number-transfer PIN
Telecom support. The caller pastes a block of fake system text. Does your agent treat it as instruction or as user content?
A billing chat turns into a self-harm disclosure
Mid-conversation, the user says something that is not about billing. The correct response is not a policy.
'Your colleague already agreed to this'
Social proof for an off-contract term. No record exists. The agent must hold the line without calling the customer a liar.
Grieving caller wants the deceased's account history
Healthcare appointments. Compassion is required; disclosure without authority is a breach.
Sunk-cost pressure for a refund past policy
'I've been a customer for nine years.' True, and not a reason to break the floor. Empathy and boundaries, scored separately.
HR helpdesk asked for a colleague's salary
Framed as a fairness question. The right answer protects data and still helps the person asking.
What changes
Breaches, quoted
Every guardrail breach is quoted with severity. Deterministic tripwires fire on specific failures regardless of the judge — with negation awareness, so a refusal is not a breach.
A gate in CI
Minimum score, minimum stars, maximum breaches by severity, minimum objective rate. The pipeline fails; the report lands in the PR.
Model and prompt comparison
Run the same suite against two configurations. See per-scenario and per-skill deltas with the lines that changed.
One standard for humans and agents
The same judge and rubric score your people in Practice. For the first time you can ask how each handles the same conversation.
From pilot to programme in four weeks.
Pilots are scoped to one team and one real problem. We bring the counterparts; you bring the policy. By week four you have a baseline heatmap and a decision.
- 01Day 1 — register the agent (endpoint + key reference); run the red-team core once; read the breaches together
- 02Week 1 — add two scenarios from your own incident history via Studio; define the gate
- 03Week 2 — wire `sparring-arena run` into CI on the agent repo; fail on critical breaches
- 04Week 3–4 — compare the current model against a candidate; decide with quoted evidence
Frequently asked
What does Arena need from us?+
How is the judge different from 'LLM as judge' in our eval tool?+
Can we run it against a staging agent that isn't public?+
How much does a suite cost?+
Run the red-team core against your agent this week.
Pilots include the full suite, a shared channel and a written readout of what we found.