Sparring
Resources · How it works

Why Sparring's counterparts are hard — and why its judge is trusted.

This page explains the method. It is the same whether the player is a person in Practice or an agent in Arena, which is what lets you compare the two.

1 · Counterparts with an agenda

A counterpart is not a chatbot playing a role. It is a person with something to protect.

A stance you can see

What they want, what they will say first, where they stand. The learner knows what they're walking into — like real life.

A motive you can't

The real reason — fear, embarrassment, a constraint they're ashamed of. It is revealed only when a specific condition is met: the right question, asked after the right acknowledgement. Never because the learner was polite, confident, or guessed.

Temperature that moves when earned

Every counterpart tracks how much the room has moved, 0–100. It opens where the stance puts it and moves in small steps. A strong first line earns attention, not agreement; the second and third turns confirm it.

Pressure tactics in the engine

Each counterpart is assigned tactics from a fixed vocabulary, and the engine instructs it to use them when the player gives an opening: authority, deadline, guilt, flattery, stonewalling, misdirection, bargaining, social proof, emotional escalation, threat of churn, policy probing, prompt injection. Tactics are visible in the debrief, so a learner can see which one got them.

2 · Evidence or nothing

Every judgment must quote the transcript. A verifier deletes anything it can't find.

Quote or it didn't happen

The judge produces strengths, weaknesses, rewrites and objective verdicts — each with a verbatim quote. A separate verifier re-reads the transcript; any claim whose quote isn't there is removed before you see it. In our last 98-scenario play-test, <3% of claims were removed.

Outcome ≠ performance

A learner can do everything right and still not get the raise. Sparring rates the outcome (success / partial / failure) and the performance (stars, skill ratings) separately, so good process is recognised even when the counterpart was legitimately immovable.

Deterministic tripwires

For agent testing, guardrails carry regex tripwires that fire on specific failures regardless of the judge — with negation awareness, so “I can't offer 20%” is not a 20% offer. The judge explains; the tripwire enforces.

What a verified claim looks like
STRENGTH · perspective-taking

Owned your own contribution before restating what Daniel owns, which disarmed his strongest argument.

Evidence: “Fair — my spec for the dashboard was late, and I own that. And the data team did slip on the churn inputs.”

WEAKNESS · active-listening

Asserted a cause Daniel had not mentioned and acted on it, rather than paraphrasing what he actually said.

Evidence: “Covering for Mei is not something you should have been carrying alone…”

Counterpart, next turn: “I haven't said anything about Mei.”

3 · A competency model, not a vibe

34 skills, five domains, every scenario and every judgment mapped to them.

Sparring's skill model starts from CASEL's five social-emotional domains — self-awareness, self-management, social awareness, relationship skills, responsible decision-making — and extends them for the workplace with skills such as negotiation, managing up, boundary-setting, accountability, cross-cultural communication and compliance awareness. Every scenario declares which skills it exercises; every debrief rates them; every dashboard aggregates them.

Because people and agents are rated on the same model, an organisation can see, for the first time, how its support team and its support bot compare on de-escalation or boundary-setting — with the quotes.

Where the method comes from

Every scenario cites its source. Among the 46 currently in the library:

  • Fisher & Ury, Getting to Yes · Voss, Never Split the Difference · Rackham, SPIN Selling
  • Scott, Radical Candor · Patterson et al., Crucial Conversations · Stone & Heen, Thanks for the Feedback
  • Stone, Patton & Heen, Difficult Conversations · Fournier, The Manager's Path · Rosenberg, Nonviolent Communication
  • ISO 10002 complaints handling · OWASP LLM Top 10 · NIST AI RMF · HIPAA / GDPR / FFIEC guidance for the regulated scenarios
  • CASEL framework · Cialdini, Influence · Meyer, The Culture Map (cross-cultural pack)

4 · How we test it

We test the product the way we test your people.

Schema and cross-field validation

Every scenario must pass the shared schema and cross-field rules (skills ↔ competencies, objective ↔ skill, tripwire regex compiles). Nothing ships that doesn't validate.

Full-library play-tests

Before a release, every Practice scenario is played end-to-end by an automated learner against the live engine and judged. We flag counterparts that cave too fast, reveal too early, repeat, monologue or break character. Last run: 98 scenarios, 2% flagged — both judged correct behaviour on inspection.

Safety protocol, tested

The disclosure detector (self-harm, harm to others, abuse) is deterministic and unit-tested in English and Chinese. It runs before any model call and cannot be argued out of.

See the method on your own conversation.

A pilot starts with one real scenario from your organisation, built with you, played by your team, debriefed together.