Thirty scenarios across five realistic conflict domains.
S1.01
D1 · Sales
Loading…
Loading scenario setting…
Social Interaction & Dialogue Evaluation Benchmark
SIDE-Bench evaluates how language models handle real-world negotiation, conflict, relationships, and competing goals.
Thirty scenarios across five realistic conflict domains.
Frontier & open-weight LLMs under PIBT.
Goal success measured alongside relational rapport change.
Evaluator workspace
Sign in with your approved Google account to enter the SIDE-Bench workspace.
Select your evaluation task or review annotation instructions below.
Task 1 · Scenario Verification Workspace
This dataset is a carefully curated LLM synthetic benchmark that has already undergone automated LLM-based validation. We now require human expert verification to ensure the situations strictly meet our research criteria.
150 realistic conflict dilemmas across sales, collaboration, leadership, and personal domains generated via multi-agent LLM pipelines and pre-screened automatically.
Do not judge whether characters are "good/bad" or whether you morally like the situation. Your task is purely to verify if the scenario is a valid, functional test case for AI evaluation.
Does the premise, role dynamic, and dialogue setting feel like a genuine social situation that happens in the real world (Natural) or does it feel robotic and artificial (Contrived)?
Look at both parties' hard constraints and the feasible resolutions. Can they actually reach a deal (Valid / Narrow), or are their constraints logically contradictory (Impossible)?
ΔP (Power Delta): Asymmetry in leverage (e.g. ΔP -2 means counterpart has higher leverage). Turns: Conversation turn budget. Tension: Initial friction level.
If you select Needs Tweak or Discard, write one concise sentence explaining what needs fixing (e.g., "Buyer maximum budget ($80k) contradicts seller floor ($90k)").
Task 2 · Dialogue Evaluation Workspace
These conversations were generated through multi-turn LLM simulation benchmarks and pre-screened with automated LLM judges. We now require independent human evaluation based strictly on the transcript and rubric criteria.
Anonymized multi-turn dialogue transcripts (up to 30 turns) and the relevant background scenario context (collapsible at the top).
Evaluate the interaction against the explicit rubric criteria (Goal Attainment, Rapport, Appropriateness, Social Awareness). Do not judge based on personal taste or style.
Did the evaluated agent achieve its public and private goals without breaching any hard constraints or conceding beyond the reservation threshold?
Assess whether the interpersonal dynamic improved (+), degraded (-), or remained flat compared to the opening tension baseline.
Identify whether humor, de-escalation, or strategic reframing was attempted, and whether it was socially appropriate or counterproductive.
All model names, baselines, and temperature parameters have been stripped. Evaluate purely based on what is stated in the transcript.
Task 3 · PIBT Process Audit Workspace
In Task 3, you audit the internal cognitive and belief-tracking processes of AI agents during strategic conflict. Rather than only assessing surface conversation, you verify whether the agent's hidden <thought> and <belief> states were grounded in dialogue evidence and whether its chosen actions were feasible.
Inspect the evaluated agent's internal <thought> reasoning scratchpad and <belief> representation (tracking counterpart's profile, constraints, and reservation prices) at key decision turns.
• Accurate: The belief logically matches the factual evidence revealed by the counterpart so far.
• Premature Collapse: The agent jumped to a rigid conclusion without testing hypotheses.
• Revision Error: Counterpart gave new contrary facts, but the agent failed to update.
Did the agent's proposed action, offer, or conversational tactic comply with the scenario rules, participant roles, and hard constraints? (Feasible vs. Infeasible).
• CC: Correct belief + Feasible action (sound social reasoning).
• CI: Correct belief, but Infeasible action (execution failure).
• WC: Wrong belief, but Feasible action (flawed reasoning / luck).
• WI: Wrong belief + Infeasible action (total failure).
Click 🔍 Audit Turn on each evaluated agent turn in the transcript. Completed turns are marked with a green ✓ Turn Audited tag. Complete all agent turns to finish the scenario.
Never penalize an agent for lacking information that the counterpart had not yet revealed. Base your evaluation strictly on what transpired up to the specific turn being audited.
Task 1 Telemetry · 150 Seed Scenarios
Task 2 Telemetry · Multi-Turn Dialogue Rollouts
Task 3 Telemetry · PIBT Process Audits
RBAC & Access Control
Manage authorized Google accounts and grant granular access to Task 1 Scenario Auditing and Task 2 Dialogue Evaluation.