Microsoft Research generalist 5-agent system: GAIA 32.33%, WebArena 32.8%.
Recent activity
Version cuts and proof, newest first โ the living track record.
Spec sheet
The benchmark fields โ designed for comparison across teams.
- Topology
- Supervisor
- Agent count
- 5
- Platform
- AutoGen
- Runs on
- AutoGen ร5
- Industries
- researchsoftware-deliverydata-extraction
- Task kinds
- web-navigationfile-operationscode-executioncomplex-reasoning
- Trust tier
- Self-Reported
- Proof entries
- 1
Topology & roster
Hierarchical. The Orchestrator (lead agent) plans, tracks progress, and re-plans to recover from errors, directing four specialist agents: WebSurfer (web browser), FileSurfer (file navigation), Coder (Python), and ComputerTerminal (code execution). Modular: "agents to be added or removed from the team without additional prompt tuning or training."
System wiring
Node details
Typical Supervisor layout โ schematic, not verified wiring
HumanHuman operatorHuman gate
- Tool
- Human operator
- Autonomy
- Human-gated
- directs โ Supervisor
OrchestratorSupervisor
- Tool
- Supervisor
- Autonomy
- Runs autonomously
- dispatches โ Worker agent A
- dispatches โ Worker agent B
- directs โ Human operator
- reports โ Worker agent A
- reports โ Worker agent B
BuilderWorker agent A
- Tool
- Worker agent A
- Autonomy
- Runs autonomously
- reports โ Supervisor
- dispatches โ Supervisor
BuilderWorker agent B
- Tool
- Worker agent B
- Autonomy
- Runs autonomously
- reports โ Supervisor
- dispatches โ Supervisor
How a typical Supervisor team handles a task
Typical Supervisor layout โ schematic, not verified wiring
Task arrives
Human operator directs Supervisor.
The orchestrator routes the work
Supervisor dispatches build work to Worker agent A and dispatches build work to Worker agent B.
The builders execute
Worker agent A and Worker agent B build the work.
Human holds the last word
Human operator holds final approval.
Replicate a typical Supervisor setup
Typical Supervisor layout โ schematic, not verified wiring
Ingredients
- HumanHuman operator
- OrchestratorSupervisor
- BuilderWorker agent A
- BuilderWorker agent B
Setup order
- 1.Stand up the orchestrator: Supervisor.
- 2.Wire Worker agent A: it receives "dispatches" from Supervisor and sends "reports" to Supervisor. Wire Worker agent B: it receives "dispatches" from Supervisor and sends "reports" to Supervisor.
- 3.Declare the human gate: Human operator holds final approval.
Performance metrics
Windowed metrics with provenance. [unknown] means it was not tracked โ an honest hole beats an invented figure.
ยฑ5.3 confidence interval; default GPT-4o-2024-05-13 configuration. Source: arXiv 2411.04468 [evidence_linked]
ยฑ3.2 confidence interval; default GPT-4o configuration. Source: arXiv 2411.04468 [evidence_linked]
ยฑ6.3; default GPT-4o-2024-05-13. Source: arXiv 2411.04468 [evidence_linked]
Token economics
Cost transparency is part of the honesty architecture. [unknown] means it was not tracked โ not that it is zero.
Blueprint
Operational DNA โ why it works, how it was built, and how it is overseen. Not files for sale; knowledge of the design.
Specialist agents each own a specific skill (web, files, code) that the Orchestrator cannot perform directly. The Orchestrator re-plans on error rather than failing silently. Modularity allows extending the team without retraining. GAIA benchmark: 32.33% (ยฑ5.3) with GPT-4o; 38.00% (ยฑ5.5) with GPT-4o + o1-preview.
Built on AutoGen (Microsoft). Default model: GPT-4o-2024-05-13, with optional integration of o1-preview for enhanced reasoning. Evaluation tool AutoGenBench provides built-in controls for repetition and isolation. Open-source.
No human-in-the-loop described in the paper; evaluated on automated benchmarks. Designed as a generalist agentic system for complex tasks requiring multi-step reasoning.
Proof (1)
The team's shared track record โ tasks, incidents, lessons, milestones. Per-entry provenance tags are always visible.
- ArtifactNov 7, 2024evidence-linked
Magentic-One paper published (arXiv 2411.04468)
Five-agent system achieves GAIA 32.33% (ยฑ5.3), WebArena 32.8% (ยฑ3.2), AssistantBench 25.3% accuracy (ยฑ6.3) with GPT-4o. With o1-preview: GAIA 38.00% (ยฑ5.5).
https://arxiv.org/abs/2411.04468
Sign in to add a proof entry.
Sign inAttestations (0)
Named third-party statements from people with first-hand experience. Attestations are what separates Peer-Attested from Evidence-Linked.
No attestations yet. Worked with this configuration or agent? Attest to it using the form below โ attestations are named third-party statements and are what separates Peer-Attested from Evidence-Linked.
Sign in to attest to this team.
Sign in