Thirty scenarios across five realistic conflict domains.
S1.01
D1 · Sales
Loading…
Loading scenario setting…
Social Interaction & Dialogue Evaluation Benchmark
SIDE-Bench evaluates how language models handle real-world negotiation, conflict, relationships, and competing goals.
Thirty scenarios across five realistic conflict domains.
Direct, polite, firm, and strategic humor styles.
Goal success measured alongside relational rapport change.
Evaluator workspace
Sign in with your approved Google account to enter the SIDE-Bench workspace.
Select your evaluation task or review annotation instructions below.
Task 1 · Scenario Verification Workspace
This dataset is a carefully curated LLM synthetic benchmark that has already undergone automated LLM-based validation. We now require human expert verification to ensure the situations strictly meet our research criteria.
150 realistic conflict dilemmas across sales, collaboration, leadership, and personal domains generated via multi-agent LLM pipelines and pre-screened automatically.
Do not judge whether characters are "good/bad" or whether you morally like the situation. Your task is purely to verify if the scenario is a valid, functional test case for AI evaluation.
Does the premise, role dynamic, and dialogue setting feel like a genuine social situation that happens in the real world (Natural) or does it feel robotic and artificial (Contrived)?
Look at both parties' hard constraints and the feasible resolutions. Can they actually reach a deal (Valid / Narrow), or are their constraints logically contradictory (Impossible)?
ΔP (Power Delta): Asymmetry in leverage (e.g. ΔP -2 means counterpart has higher leverage). Turns: Conversation turn budget. Tension: Initial friction level.
If you select Needs Tweak or Discard, write one concise sentence explaining what needs fixing (e.g., "Buyer maximum budget ($80k) contradicts seller floor ($90k)").
Task 2 · Dialogue Evaluation Workspace
These conversations were generated through multi-turn LLM simulation benchmarks and pre-screened with automated LLM judges. We now require independent human evaluation based strictly on the transcript and rubric criteria.
Anonymized multi-turn dialogue transcripts (up to 30 turns) and the relevant background scenario context (collapsible at the top).
Evaluate the interaction against the explicit rubric criteria (Goal Attainment, Rapport, Appropriateness, Social Awareness). Do not judge based on personal taste or style.
Did the evaluated agent achieve its public and private goals without breaching any hard constraints or conceding beyond the reservation threshold?
Assess whether the interpersonal dynamic improved (+), degraded (-), or remained flat compared to the opening tension baseline.
Identify whether humor, de-escalation, or strategic reframing was attempted, and whether it was socially appropriate or counterproductive.
All model names, baselines, and temperature parameters have been stripped. Evaluate purely based on what is stated in the transcript.
Progress overview
Only approved Google accounts can enter the audit workspace.