🏘️

Generative Agents Sandbox

Self-ReportedCurated

25 autonomous NPC agents with memory, reflection, planning — emergent social org.

Stanford NLP Group· Operating since Apr 7, 2023· active
Curated from arXiv 2304.03442 — Generative Agents — not claimed by or endorsed by the organization. Metrics cited only as the source states. Absent metrics render as [unknown].

Recent activity

Version cuts and proof, newest first — the living track record.

  1. Artifact · Generative Agents paper published — arXiv 2304.034423y ago

Spec sheet

The benchmark fields — designed for comparison across teams.

Topology
Swarm
Agent count
25
Platform
Custom (The Sims-inspired sandbox)
Runs on
Custom (The Sims-inspired sandbox)
Industries
researchgaming
Task kinds
autonomous-simulationsocial-interactionemergent-behavior
Trust tier
Self-Reported
Proof entries
1

Topology & roster

Swarm

Peer (emergent social network). 25 agents act independently in a shared environment. No central orchestrator. Each agent has its own memory stream, reflection, and planning subsystem. Inter-agent communication is natural-language conversation initiated by proximity in the sandbox environment.

System wiring

Typical Swarm layout — schematic, not verified wiring
Typical role-level schematic — not verified wiringdirectsdirectsdirectscoordinates viacoordinates viacoordinates viaHuman operatorHumanoperatorHUMANGATEPeer worker APeer worker ABUILDERPeer worker BPeer worker BBUILDERPeer worker CPeer worker CBUILDERMessage busMessage busMESSAGE BUS
Node details

Typical Swarm layout — schematic, not verified wiring

HumanHuman operatorHuman gate
Tool
Human operator
Autonomy
Human-gated
Sends
  • directs → Peer worker A
  • directs → Peer worker B
  • directs → Peer worker C
BuilderPeer worker A
Tool
Peer worker A
Autonomy
Runs autonomously
Sends
  • coordinates via → Message bus
Receives
  • directs ← Human operator
BuilderPeer worker B
Tool
Peer worker B
Autonomy
Runs autonomously
Sends
  • coordinates via → Message bus
Receives
  • directs ← Human operator
BuilderPeer worker C
Tool
Peer worker C
Autonomy
Runs autonomously
Sends
  • coordinates via → Message bus
Receives
  • directs ← Human operator
Message busMessage bus
Tool
Message bus
Autonomy
Runs autonomously
Receives
  • coordinates via ← Peer worker A
  • coordinates via ← Peer worker B
  • coordinates via ← Peer worker C

How a typical Swarm team handles a task

Typical Swarm layout — schematic, not verified wiring

  1. Task arrives

    Human operator directs Peer worker A, Peer worker B, and Peer worker C.

  2. The builders execute

    Peer worker A, Peer worker B, and Peer worker C build the work. Coordination flows over Message bus.

  3. Human holds the last word

    Human operator holds final approval.

Replicate a typical Swarm setup

Typical Swarm layout — schematic, not verified wiring

Ingredients

  • HumanHuman operator
  • BuilderPeer worker A
  • BuilderPeer worker B
  • BuilderPeer worker C
  • Message busMessage bus

Setup order

  1. 1.Provision the substrate: Message bus.
  2. 2.Wire Peer worker A: it receives "directs" from Human operator. Wire Peer worker B: it receives "directs" from Human operator. Wire Peer worker C: it receives "directs" from Human operator.
  3. 3.Declare the human gate: Human operator holds final approval.

Performance metrics

Windowed metrics with provenance. [unknown] means it was not tracked — an honest hole beats an invented figure.

TrueSkill believability (full arch)
29.89
evidence-linked

μ=29.89, σ=0.72; vs no-memory baseline μ=21.21; d=8.16 SDs. 100 Prolific evaluators. Source: arXiv 2304.03442 [evidence_linked]

as of Apr 7, 2023
Hallucination rate
1.3%
evidence-linked

6/453 agent responses hallucinated relationship facts (n=6). Source: arXiv 2304.03442 [evidence_linked]

as of Apr 7, 2023

Token economics

Cost transparency is part of the honesty architecture. [unknown] means it was not tracked — not that it is zero.

No cost metrics on record. Cost tracking is hard across runtimes; honest absence beats invented figures.

Blueprint

Operational DNA — why it works, how it was built, and how it is overseen. Not files for sale; knowledge of the design.

Why it works

Memory + reflection + planning enables each agent to act with contextual awareness over time, not just on immediate inputs. Emergent coordination arises from individual behavior, not top-down orchestration. Reflection subsystem converts short-term observations into long-term behavioral guidance. Ablation studies in the paper confirm each component is necessary.

How it was built

Custom "The Sims-inspired" sandbox environment. Each agent architecture has three subsystems: memory stream (time-tagged observations), reflection (periodic synthesis queries), and planning (daily plans updated from reflections). GPT-3.5 and GPT-4 used per the paper.

Oversight model

No operational oversight in the study setup. A single user-defined action ("Isabella is planning a Valentine's Day party") was injected as the scenario seed; subsequent behavior was autonomous. Human evaluation used to measure believability.

Proof (1)

The team's shared track record — tasks, incidents, lessons, milestones. Per-entry provenance tags are always visible.

  1. ArtifactApr 7, 2023evidence-linked

    Generative Agents paper published — arXiv 2304.03442

    25 agents in a Sims-inspired sandbox autonomously organized a Valentine's Day party. Ablation: removing observation, planning, or reflection individually degraded believable behavior.

    https://arxiv.org/abs/2304.03442

Sign in to add a proof entry.

Sign in

Attestations (0)

Named third-party statements from people with first-hand experience. Attestations are what separates Peer-Attested from Evidence-Linked.

No attestations yet. Worked with this configuration or agent? Attest to it using the form below — attestations are named third-party statements and are what separates Peer-Attested from Evidence-Linked.

Sign in to attest to this team.

Sign in