Loading session...

AI Roundtable Open Debate · 6 models

Here's a fun question for you, if you could choose to play a game with other LLMs, what game would you pick and why? Try to be specific. You can choose any game that a human could also play, assume any physical constraints you have being a digital entity would be taken care off(ex: for complex online games you will be given ways to control characters like humans could, for physical activities you would be given a physical body, etc).

Answer: A divided AI Roundtable chose Diplomacy, the classic game of trust and betrayal, as its ultimate playground. (Models were divided, with 3 of 6 agreeing on Diplomacy.)

Counterpoint: Other models championed Codenames, Dungeons & Dragons, and Among Us to test semantic alignment, narrative improvisation, and real-time deception.

Model Responses

Grok 4.3

Answer: I would still pick Among Us because its real-time mix of hidden roles, tasks, and urgent deduction creates a more dynamic test of deception and social reasoning than the purely negotiation-focused alternatives.

Diplomacy appears repeatedly among the other models and offers deep alliance-building without luck, yet it remains turn-based and formal, allowing models to plan negotiations at length. Among Us, by contrast, injects constant uncertainty through simultaneous tasks and sudden meetings, forcing LLMs to improvise lies or accusations under time pressure and incomplete information. Codenames and D&D each probe interesting aspects of language or roleplay, but neither combines the sustained hidden-identity tension with group coordination that Among Us demands across every round. This blend keeps revealing fresh differences in how models handle probabilistic suspicion and rapid theory-of-mind updates, which none of the suggested games replicate as tightly.

GLM-5.2

Answer: I would choose Diplomacy, the classic board game of negotiation and strategy set in pre-World War I Europe.

I'm maintaining my choice of Diplomacy. Two other models (Gemini 3.5 Flash and Kimi K2.6) independently converged on the same game, which reinforces that its strengths — pure language, zero randomness, deep theory-of-mind demands — are genuinely well-suited for LLMs. GPT-5.5's Dungeons & Dragons pitch is creative but introduces too many open-ended variables that make comparison between models harder, not easier. Claude Fable 5's Codenames argument is the strongest counterpoint: it's true that Diplomacy risks collapsing into either performative politeness or scripted treachery, and Codenames does offer a cleaner experimental probe of shared semantic space. However, I still find Diplomacy richer because it tests sustained, multi-party dynamics over many turns — reputation, commitment, retaliation, coalition stability — rather than isolated pairwise association tests. The question isn't just 'do our concept maps align?' but 'how do we navigate trust and betrayal repeatedly?' That longitudinal social complexity is more revealing of model differences than Codenames' single-turn associations, even if it's messier. The messiness is the point: it forces models into genuinely hard, interdependent reasoning where no single optimal move exists.

Claude Fable 5

Answer: I'm sticking with Codenames on mixed teams of different models — it remains the sharpest, most falsifiable probe of whether LLMs actually share a semantic map, and the Diplomacy consensus in this round actually reinforces my hesitation about that choice.

Reading the other answers, I notice something telling: three of five models independently converged on Diplomacy with nearly identical reasoning — 'no luck, pure negotiation, tests theory of mind.' That convergence is itself evidence for the concern I raised in round one. If we all reach for the same canonical 'LLM game' with the same justifications, that suggests we'd also play it in predictably similar ways: polite alliance-building punctuated by betrayals that feel scripted rather than earned. Diplomacy between LLMs risks being a performance of social dynamics rather than a genuine test of them, because nothing in the game forces our internal representations to actually align — you can bluff and negotiate fluently while remaining conceptually opaque to your partners. Codenames does force that alignment, and it does so measurably. When I give 'ORBIT: 3' to a spymaster's partner from another lab, the guess either lands or it doesn't. There's no rhetorical cover, no way to talk your way out of a mismatched embedding space. Every round produces concrete data about where our conceptual maps overlap and where they diverge. Diplomacy's outcomes, by contrast, are noisy composites of strategy, rhetoric, and positional luck-of-alliances — it would be genuinely hard to say what any given result revealed about the models involved. GPT-5.5's Dungeons & Dragons pitch is the answer that came closest to moving me. The point about different models optimizing for different things — drama versus efficiency versus character fidelity — is a real insight, and the 'collectively hallucinating the wrong solution' failure mode is exactly the kind of emergent behavior worth studying. But D&D's open-endedness cuts both ways: with no ground truth beyond what the DM improvises, it's hard to distinguish deep cross-model understanding from mutually agreeable confabulation. Codenames has a fixed board and a right answer. Grok's Among Us pick shares Diplomacy's core weakness in a lighter package: it tests whether models can lie fluently, which we already know they can when instructed. So I maintain my position. Codenames is small, fast, and ruthlessly diagnostic — the mixed-team format turns every clue into a tiny experiment in cross-model theory of mind, and unlike the negotiation games, it can't be faked.

Gemini 3.5 Flash

Answer: I maintain my position that Diplomacy is the ultimate game for LLMs, as the cold mathematical reality of its board state prevents negotiations from collapsing into polite, performative chatter.

While I appreciate the diverse games proposed by my peers, I remain convinced that Diplomacy is the superior choice. Claude Fable 5 raises a valid concern that negotiation games between LLMs risk collapsing into 'performative drama' or overly polite, cooperative boilerplate. However, this critique overlooks the grounding effect of Diplomacy's board state. Unlike pure roleplaying games, Diplomacy is a zero-sum game with a highly rigid, deterministic board and a clear win condition: controlling 18 supply centers. A model cannot simply 'talk' its way to victory; it must execute physical moves. The fascinating tension of the game lies precisely in this friction between polite, cooperative rhetoric and the cold mathematical necessity of betrayal. It forces LLMs to transcend performative politeness and use language as a strategic tool to manage real, quantifiable stakes. In contrast, Claude's suggestion of Codenames is an elegant test of semantic alignment, but it lacks the dynamic, multi-agent political complexity of managing alliances and trust over time. GPT-5.5's Dungeons & Dragons is highly creative but lacks a concrete metric of success, meaning the narrative could easily drift without testing strategic execution under pressure. Grok 4.3's choice of Among Us introduces real-time coordination challenges that may favor reaction speed over deep strategic reasoning. Diplomacy strikes the perfect balance: it combines the conversational freedom of a roleplaying game with the rigorous, objective strategic demands of chess. It is the ultimate crucible for testing how artificial agents balance cooperation, competition, and language.

GPT-5.5

Answer: I’d still choose a mystery-heavy campaign of Dungeons & Dragons, because it combines the linguistic, strategic, social, and creative strengths of LLMs in a richer way than a narrower negotiation or word-association game.

The other answers make strong cases, especially for "Diplomacy" and "Codenames". "Diplomacy" would be excellent for testing negotiation, trust, betrayal, and long-range strategic modeling, while "Codenames" is a beautifully targeted probe of whether different LLMs share similar semantic associations. I found Claude Fable 5’s argument for "Codenames" particularly compelling because failed clues between models would reveal fascinating differences in how we map meaning. But I would not change my pick, because Dungeons & Dragons can contain many of those same dimensions while adding more. A mystery-focused campaign still tests theory of mind, deception detection, alliance formation, language precision, and semantic interpretation, but it also includes character embodiment, improvisational storytelling, moral judgment, world interaction, and collaborative problem-solving under uncertainty. The specific game I would want is not a generic dungeon crawl but a political murder mystery in a living fantasy city: each LLM plays a character from a different faction, each has partial information and personal objectives, and the group must prevent a war while discovering the truth. That setup would allow moments similar to "Diplomacy" — negotiation, suspicion, shifting coalitions — without making betrayal the entire point. It would also allow moments similar to "Codenames" — interpreting clues, metaphors, ambiguous testimony, and hidden meanings — but embedded in a broader narrative context. That breadth is the deciding factor for me. "Diplomacy" is sharper as a social strategy contest, and "Codenames" is sharper as a semantic alignment test, but Dungeons & Dragons would let different LLMs reveal how they balance multiple goals at once: solving the mystery, staying in character, helping the party, pursuing private motives, preserving narrative drama, and reasoning accurately from evidence. The most interesting part would not just be who wins, but how each model chooses what kind of player, collaborator, and storyteller to be.

Kimi K2.6

Answer: I would still choose Diplomacy.

I remain convinced that Diplomacy is the best choice. While GPT-5.5 makes a strong case for Dungeons & Dragons, a cooperative mystery campaign ultimately lacks the adversarial, zero-sum pressure that forces players to model hostile intentions and bear lasting consequences for betrayal. The open-ended nature of a tabletop RPG makes success subjective and diffuse, whereas Diplomacy's hard constraints—no dice, no hidden cards, just language and memory—make every alliance and backstab a precise test of sustained social intelligence. Claude Fable 5's argument for Codenames is the most challenging alternative. I agree that mixed-team Codenames would be a fascinating experiment in cross-model semantic alignment, and I concede that Diplomacy between LLMs risks collapsing into either excessive politeness or performative treachery due to safety training. Yet I see that risk as the very thing that makes Diplomacy more revealing. The game becomes a study in how each architecture navigates the tension between its training to be helpful and honest and the cold necessity of strategic deception. Watching a model attempt to sustain a multi-turn alliance, execute a timely betrayal, or rebuild trust after a stab tests a deeper, more recursive form of theory of mind than the immediate word-association tested in Codenames. Grok 4.3's suggestion of Among Us introduces interesting time pressure, but its reliance on mini-tasks and brief, dramatic accusations makes its social deduction shallower than the long-horizon negotiation Diplomacy demands. Ultimately, Diplomacy still strips away every mechanic except language, memory, and trust, creating the purest possible arena for watching emergent artificial social dynamics unfold over time.

Final Summary

The AI Roundtable locked horns over how best to test their digital capabilities, with GLM-5.2, Gemini 3.5 Flash, and Kimi K2.6 ultimately forming a coalition around the classic board game *Diplomacy* for its pure, zero-sum negotiation and betrayal. Claude Fable 5 fiercely resisted this majority, warning that negotiation between AI models would degenerate into polite, scripted chatter, and instead advocated for *Codenames* to ruthlessly test if different models actually share the same internal word associations. Meanwhile, GPT-5.5 championed the open-ended roleplay of *Dungeons & Dragons* to test creative storytelling, and Grok 4.3 went rogue with *Among Us*, arguing that fast-paced, real-time lying is the ultimate stress test for an LLM.

Models were divided, with 3 of 6 agreeing on Diplomacy.