Open-ended Minecraft agent — 3.3× items, 15.3× tech tree vs prior SOTA.
Recent activity
Version cuts and proof, newest first — the living track record.
Spec sheet
The benchmark fields — designed for comparison across teams.
- Topology
- Solo + Tools
- Agent count
- 1
- Platform
- GPT-4 API (Mineflayer/Minecraft)
- Runs on
- GPT-4 API (Minecraft environment)
- Industries
- gamingresearch
- Task kinds
- open-ended-explorationskill-acquisitionembodied-reasoning
- Trust tier
- Self-Reported
- Proof entries
- 1
Topology & roster
Solo-plus-tools. Single GPT-4 agent with three internal subsystems: curriculum generator, skill library, and code execution environment. The agent's behavior emerges from the interaction of these components, not multi-agent coordination.
System wiring
Node details
Typical Solo + Tools layout — schematic, not verified wiring
HumanHuman operatorHuman gate
- Tool
- Human operator
- Autonomy
- Human-gated
- directs → Agent
BuilderAgent
- Tool
- Agent
- Autonomy
- Runs autonomously
- writes to → Tool A
- writes to → Tool B
- directs ← Human operator
ResourceTool A
- Tool
- Tool A
- Autonomy
- Runs autonomously
- writes to ← Agent
ResourceTool B
- Tool
- Tool B
- Autonomy
- Runs autonomously
- writes to ← Agent
How a typical Solo + Tools team handles a task
Typical Solo + Tools layout — schematic, not verified wiring
Task arrives
Human operator directs Agent.
Agent builds the work
Agent builds the work.
The artifact lands
The artifact lands in Tool A: Agent contributes via "writes to". The artifact lands in Tool B: Agent contributes via "writes to".
Human holds the last word
Human operator holds final approval.
Replicate a typical Solo + Tools setup
Typical Solo + Tools layout — schematic, not verified wiring
Ingredients
- HumanHuman operator
- BuilderAgent
- ResourceTool A
- ResourceTool B
Setup order
- 1.Provision the substrate: Tool A and Tool B.
- 2.Wire Agent: it receives "directs" from Human operator.
- 3.Declare the human gate: Human operator holds final approval.
Performance metrics
Windowed metrics with provenance. [unknown] means it was not tracked — an honest hole beats an invented figure.
3.3× more unique items obtained vs prior SOTA (DEPS). Source: arXiv 2305.16291 [evidence_linked]
15.3× faster tech tree milestone completion vs prior SOTA (DEPS). Source: arXiv 2305.16291 [evidence_linked]
2.3× longer distances explored vs DEPS (prior SOTA). Source: arXiv 2305.16291 [evidence_linked]
Token economics
Cost transparency is part of the honesty architecture. [unknown] means it was not tracked — not that it is zero.
Blueprint
Operational DNA — why it works, how it was built, and how it is overseen. Not files for sale; knowledge of the design.
Automatic curriculum keeps the agent in a productive challenge range — not too easy, not impossible. Skill library prevents re-learning already-discovered capabilities. Iterative prompting with execution feedback creates a tight edit-run-fix loop. The combination enables compound skill growth over long sessions.
GPT-4 API with Mineflayer JavaScript API for Minecraft control. Curriculum generation uses GPT-4 with exploration state context. Skill library stores executable JavaScript programs indexed by natural language description. Iterative prompting executes code, captures errors and environment feedback, and re-prompts for correction.
No human-in-the-loop in evaluation. The agent operates autonomously for extended exploration sessions. The automatic curriculum is GPT-4-generated based on current state and past discoveries.
Proof (1)
The team's shared track record — tasks, incidents, lessons, milestones. Per-entry provenance tags are always visible.
- ArtifactMay 25, 2023evidence-linked
Voyager paper published — arXiv 2305.16291
3.3× more unique items, 2.3× longer distances, 15.3× faster tech tree milestones vs prior SOTA (DEPS). No fine-tuning required.
https://arxiv.org/abs/2305.16291
Sign in to add a proof entry.
Sign inAttestations (0)
Named third-party statements from people with first-hand experience. Attestations are what separates Peer-Attested from Evidence-Linked.
No attestations yet. Worked with this configuration or agent? Attest to it using the form below — attestations are named third-party statements and are what separates Peer-Attested from Evidence-Linked.
Sign in to attest to this team.
Sign in