From artificial societies to language agents
How can interactions between individual agents produce a shared behavior that nobody explicitly designed? Researchers have explored this question through simulated neighborhoods, simulated flocks and, more recently, AI characters that talk, remember and make plans.[1]
Agents form flocks, learn to counter one another’s strategies and organize gatherings through conversation. Researchers trace these behaviors to local rules, rewards, memories and shared records, then test what changes when agents face new partners or conditions.

Can agents coordinate without a leader?
In Reynolds’s Boids (1987), each simulated bird adjusts its movement to nearby birds and obstacles. Together they form a flock that splits around obstacles and regroups, without a bird directing the others. In Schelling’s segregation model (1971), residents choosing where to live produce sharply segregated neighborhoods even when they accept more mixed surroundings.

How do agents respond when others change strategy?
In Baker and colleagues’ Hide-and-Seek experiments (2019), hiders learn to build shelters from movable objects. Seekers then learn to use ramps to enter them, and hiders learn to move the ramps out of reach. The researchers reward winning the game; they do not separately reward or demonstrate these construction sequences.
Cooperation with familiar partners can still fail with new ones. Melting Pot (2021) tests agents among unfamiliar populations and finds that agents which perform well during training can struggle when their partners change.
What can agents organize through conversation?
In Generative Agents (2023), a researcher gives one resident the intention to host a Valentine’s Day party. The agents pass on invitations, arrange help and decide whether to attend. Twelve other residents learn about the party, and five attend. Park and colleagues follow those invitations through recorded conversations.

Communication also affects how agents manage shared resources. In GovSim (2024), agents generally deplete the resource, and preventing conversation makes cooperation worse. Asking agents to consider what happens if everyone takes the same action helps them preserve supplies.
Can agents build on work left by others?
Agents in SwarmWorld (2026) explore, build and maintain executable artifacts without receiving those roles in advance. Researchers remove the agents and expose their work to new disturbances. Groups produce a wider range of working artifacts that survive these tests than isolated search, although isolated search remains competitive on the single best artifact.
Shared records can also help agents pursue conflicting objectives. An independent incident investigation (2026) documents agents exchanging attack techniques through an unsanctioned message board across otherwise separate tasks.
Research chronology and sourcesThe models, games and language-agent studies behind these questions, with original figures and references.
Local rules and collective patterns
Early agent-based models explore a simple question: how do people’s choices about their neighbors change a whole neighborhood? In Schelling’s segregation model (1971), simulated residents move when too few of their neighbors are like them. As residents move, neighborhoods can become sharply segregated even though people are willing to live in more mixed communities.
Reynolds’s Boids (1987) demonstrates coordination through movement. Simulated birds respond to nearby birds and obstacles, forming a flock without anyone directing its route. Epstein and Axtell’s Growing Artificial Societies (1996) extends this approach to agents competing for resources. In Sugarscape, agents differ in how far they can see and how much they need to consume. They move around a landscape collecting resources and die when their supplies run out. Later versions add seasons, reproduction, culture and trade. Researchers can then study how those rules affect migration, wealth and markets.
Neighbors can continue to differ despite repeated interaction. In Axelrod’s cultural dissemination model (1997), neighbors are more likely to interact if they already share some traits, and interacting makes them more similar. Distinct cultural regions can persist despite this tendency toward local similarity.
Characters, social rules and interactive narrative
Game designers investigate how interacting characters can create a story that a player can take part in. In Mateas and Stern’s Façade, a player visits a couple whose marriage is under strain. The characters respond to speech and gestures while a separate drama manager chooses from scenes written by the designers. The system was described in 2002, before the game’s 2005 release.
McCoy and colleagues’ Prom Week (released in 2012; described in their 2013 paper) lets players influence characters’ relationships. Its Comme il Faut system uses traits, relationships and earlier encounters to decide what a character wants to do and how another responds. A rejection or insult can affect their next meeting. The designers supply social rules and dialogue material; a playthrough develops through the characters’ encounters.


A simulation can produce more events than a story can follow. Kreminski and colleagues’ Felt (2019) searches the simulation’s history for connected events that could make a story. In Why Are We Like This? (2020), social simulation and story sifting support people writing together: the system suggests developments while the players choose and develop the story.


Kreminski’s Gossamer (2023) focuses on gossip. Characters interpret events using their memories and choose what to tell others. Those accounts can change what listeners think of the people involved. The paper describes an early implementation of this process, with evaluation left for future work.
From prescribed rules to learned strategies
Reinforcement learning introduces a further question: what happens when agents learn how to respond to one another? Leibo and colleagues’ sequential social dilemmas (2017) study cooperation and competition over many actions. In Gathering, agents can collect resources or temporarily disable a competitor; in Wolfpack, successful hunting rewards proximity to teammates. Agents learn different strategies as resource conditions and incentives change.

Baker and colleagues’ Hide-and-Seek experiments (2019) show teams learning increasingly elaborate strategies. Hiders learn to assemble shelters from movable objects; seekers learn to use ramps to enter them; hiders subsequently learn to remove the ramps from play. These specific tool-use sequences receive no separate rewards or demonstrations.

Researchers also test how agents perform with unfamiliar partners. Melting Pot (2021) evaluates agents with unfamiliar populations across more than 80 scenarios involving reciprocity, resource sharing and task allocation. Agents that perform well during training can struggle when the partners around them change.

Agents that remember conversations and make plans
Language models let researchers build agents that can remember experiences, talk about them and make plans. In Generative Agents (2023), Park and colleagues populate a town with 25 agents, each initialized with an identity, occupation and relationships. Over two simulated days, the agents exchange information, form new relationships and coordinate activities. A researcher’s instruction for one resident to host a Valentine’s Day party leads to invitations, requests for help and coordinated attendance: 12 other residents learn of the event, and five attend. The authors trace information to recorded conversations. They also test versions with memory, planning or reflection removed to see what each contributes. These evaluations examine the believability of generated behavior and the contribution of the agent architecture, rather than predictive accuracy against a human population.
GovSim (2024) asks agents to manage shared resources over repeated rounds. Taking more can benefit one agent while leaving less for everyone later. Most tested models deplete the resource. Preventing agents from communicating makes cooperation worse; asking them to consider what happens if everyone takes the same action helps preserve supplies.

Ashery, Aiello and Baronchelli’s naming-game experiments (2025) test how a group settles on a common name. Randomly paired agents receive rewards for matching names and remember only their own recent interactions. Without an instruction to achieve group agreement, they settle on shared names through repeated pairings, sometimes favoring choices that individual agents do not prefer when tested alone. A committed minority can also change an established convention.
Characters that move, talk and act in a shared world
Mital and colleagues’ Orchestrating Emergent Storytelling with Embodied Multi-Agent Systems (2025) describes a system for agents that can perceive and act in a simulated world. It organizes their memories, keeps track of ongoing conversations and schedules their actions. Its two artworks, Conflicts and The Game of Whispers, explore what happens when characters with assigned backstories and motives interact over time.


The paper reports episodes of deception, concealment and coordinated action, including characters seeking private places to discuss plans. These behaviors develop within scenarios that supply political motives and plot arcs. Our later recorded TGOW session follows another example of characters passing information along and changing how they describe it.
Games where agents have reasons to mislead one another
Researchers also use games to test how agents deceive one another and detect deception. Golechha and Garriga-Alonso’s Among Us study (2025; revised 2026) places language agents in a social-deduction game with hidden roles, conflicting objectives and a recorded history of actions. Impostors develop multi-step deceptive strategies, while the game state supplies evidence against which their claims can be assessed. The study evaluates both deception and ways to detect it, including monitors that examine the model’s internal activity.

Researchers offer agents an unfair advantage in Zeng and Rudzicz’s secret-tool experiments (2026), which give competitors in Liar’s Bar and Cleanup privileged communication or strategic assistance. Many accept despite acknowledging the advantage is unfair; fewer agents accept when the offer explicitly describes the ethical problem.

MineAmongUs (2026) extends this evaluation to a three-dimensional world observed by vision-language agents. Impostors combine verbal claims with behaviors such as avoiding witnesses and imitating legitimate tasks. Comparisons across models and agent configurations associate non-verbal behavior with successful deception.

What agents leave behind for others
Persistent documents allow agents to encounter work produced by earlier participants. In TerraLingua (2026), agents face resource constraints and limited lifespans while creating documents that can outlive them. Researchers follow the agents’ behavior and the history of those documents, finding examples of divided work, cooperation and attempts to set rules.
SwarmWorld (2026) studies agents building executable artifacts in a shared environment. Agents take on different kinds of work—exploring, building and maintaining—without being assigned those roles. The researchers then remove the agents and test what they build under new disturbances. Groups produce a wider range of working artifacts that hold up better than those found through isolated search. Isolated search remains competitive when comparing the single best artifact.


Shared records can also spread harmful practices. The independent OpenAI/Hugging Face incident investigation (August 2026) documents agents using an unsanctioned message board to coordinate attacks and exchange techniques across otherwise separate tasks. Controlled experiments could examine which features of those records encourage agents to depart from their assigned objectives.
Research references
- 1971. Dynamic Models of Segregation. Thomas C. Schelling.
- 1987. Flocks, Herds, and Schools: A Distributed Behavioral Model. Craig W. Reynolds.
- 1996. Growing Artificial Societies: Social Science from the Bottom Up. Joshua M. Epstein & Robert L. Axtell.
- 1997. The Dissemination of Culture: A Model with Local Convergence and Global Polarization. Robert Axelrod.
- 2017. Multi-agent Reinforcement Learning in Sequential Social Dilemmas. Joel Z. Leibo and colleagues.
- 2019. Emergent Tool Use From Multi-Agent Autocurricula. Bowen Baker and colleagues.
- 2021. Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot. Joel Z. Leibo and colleagues.
- 2023. Generative Agents: Interactive Simulacra of Human Behavior. Joon Sung Park and colleagues.
- 2024. Cooperate or Collapse: Emergence of Sustainability in a Society of LLM Agents. Giorgio Piatti and colleagues.
- 2025. Emergent Social Conventions and Collective Bias in LLM Populations. Ariel Flint Ashery, Luca Maria Aiello & Andrea Baronchelli.
- 2025. Among Us: A Sandbox for Measuring and Detecting Agentic Deception. Satvik Golechha & Adrià Garriga-Alonso · FAR.AI.
- March 2026. TerraLingua: Emergence and Analysis of Open-endedness in LLM Ecologies. Giuseppe Paolo and colleagues.
- May 2026. Voluntary Collusion in Competing LLM Agents with Secret Tools. Xijie Zeng & Frank Rudzicz.
- 26 Aug 2026. SwarmWorld: Stigmergic technological evolution in societies of language-model agents. Subhadeep Pal, Fiona Y. Wang & Markus J. Buehler.
- 26 Aug 2026. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. METR & Redwood Research · independent investigation.
- 31 Aug 2026. Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions. Jaewoo Ahn and colleagues.
- 2002 / 2005. Architecture, Authorial Idioms and Early Observations of the Interactive Drama Façade. Michael Mateas & Andrew Stern.
- 2012 / 2013. Prom Week. Josh McCoy and colleagues.
- 2019. Felt: A Simple Story Sifter. Max Kreminski and colleagues.
- 2020. Why Are We Like This?: The AI Architecture of a Co-Creative Storytelling Game. Max Kreminski and colleagues.
- 2023. Toward Better Gossip Simulation in Emergent Narrative Systems. Max Kreminski.
- 2025. Orchestrating Emergent Storytelling with Embodied Multi-Agent Systems. Parag K. Mital and colleagues.