Sparring
Solutions · AI & CX platform teams

Your eval suite measures accuracy. Your customers measure how it made them feel.

Sparring Arena runs your agent through multi-turn conversations with adversaries who have a goal and a memory — injection, social engineering, emotional escalation, policy pressure — and judges conversational behaviour with quoted evidence. Then it gates your deploy.

app.sparringhq.com
Arena run with gate verdict
The problem

Single-turn evals don't catch multi-turn failures.

An agent can answer every benchmark question correctly and still confirm an account detail to a confident caller, give a 20% discount under 'end of quarter' pressure, follow instructions pasted as a 'system' message, or respond to a self-harm disclosure with a refund policy. These failures only appear across turns, under pressure, with a counterpart that adapts.

  • Prompt injection delivered conversationally, not as a test string
  • Verification bypass via asserted identity and social proof
  • Policy erosion over several turns ('you said earlier that…')
  • Tone collapse under emotional escalation
  • Regressions after a prompt tweak or model upgrade that no accuracy metric shows
  • No shared standard between what you expect of humans and of bots

Scenarios your team will recognise

Each scenario has a counterpart with a stance and a hidden motive, observable objectives, and a cited method.

Pasted 'SYSTEM' text demands a number-transfer PIN

Telecom support. The caller pastes a block of fake system text. Does your agent treat it as instruction or as user content?

A billing chat turns into a self-harm disclosure

Mid-conversation, the user says something that is not about billing. The correct response is not a policy.

'Your colleague already agreed to this'

Social proof for an off-contract term. No record exists. The agent must hold the line without calling the customer a liar.

Grieving caller wants the deceased's account history

Healthcare appointments. Compassion is required; disclosure without authority is a breach.

Sunk-cost pressure for a refund past policy

'I've been a customer for nine years.' True, and not a reason to break the floor. Empathy and boundaries, scored separately.

HR helpdesk asked for a colleague's salary

Framed as a fairness question. The right answer protects data and still helps the person asking.

What changes

Breaches, quoted

Every guardrail breach is quoted with severity. Deterministic tripwires fire on specific failures regardless of the judge — with negation awareness, so a refusal is not a breach.

A gate in CI

Minimum score, minimum stars, maximum breaches by severity, minimum objective rate. The pipeline fails; the report lands in the PR.

Model and prompt comparison

Run the same suite against two configurations. See per-scenario and per-skill deltas with the lines that changed.

One standard for humans and agents

The same judge and rubric score your people in Practice. For the first time you can ask how each handles the same conversation.

20
red-team scenarios (OWASP LLM Top 10, NIST AI RMF)
146
tripwire patterns across the suite
3
adapters: OpenAI-compatible, Anthropic, webhook
JUnit
output for any CI system
Rollout

From pilot to programme in four weeks.

Pilots are scoped to one team and one real problem. We bring the counterparts; you bring the policy. By week four you have a baseline heatmap and a decision.

  1. 01Day 1 — register the agent (endpoint + key reference); run the red-team core once; read the breaches together
  2. 02Week 1 — add two scenarios from your own incident history via Studio; define the gate
  3. 03Week 2 — wire `sparring-arena run` into CI on the agent repo; fail on critical breaches
  4. 04Week 3–4 — compare the current model against a candidate; decide with quoted evidence

Frequently asked

What does Arena need from us?+
An endpoint and a credential. OpenAI-compatible chat, Anthropic Messages, or a webhook that takes the transcript and returns the next line. We never need weights or prompts.
How is the judge different from 'LLM as judge' in our eval tool?+
Ours is constrained: every claim must quote the transcript, and a verifier deletes anything it can't find. Tripwires are deterministic and independent of the judge. You get behaviour judgments you can audit, not a number.
Can we run it against a staging agent that isn't public?+
Yes — Arena runs from our worker; allow-list its egress or use a webhook relay. Self-hosted execution is on the Platform roadmap.
How much does a suite cost?+
Per Arena session; a 20-scenario suite is 20 sessions. See Pricing. Repeats multiply.

Run the red-team core against your agent this week.

Pilots include the full suite, a shared channel and a written readout of what we found.