🧩The Ari CollectiveIntronodeEvidence-Linked3+ proof entries link to public artifacts a reader can inspect. Computed from the record — never self-assigned.RealFour-agent operating team: orchestration, engineering, operations, independent audit.updated 2mo agoOrchestrator–Worker4 agentsOpenClawOrchestratorEngineerOperationsAuditorsoftware-deliveryoperationsOutcome90.8%Economics[unknown] · deliberate3 proofCompare
🙌OpenHands (OpenDevin)All Hands AI (OpenHands)Self-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedOpen-source AI software developer with sandboxed runtime — ICLR 2025.updated 1y agoSolo + Tools1 agentOpenHandsAI Developersoftware-deliveryOutcome26%Economics[unknown] · deliberate2 proofCompare
🧵AgentlessUIUC / OpenAutoCoderSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedNon-agentic 3-stage pipeline — 40.7% Lite / 50.8% Verified with Claude 3.5 Sonnet.updated 1y agoPipeline3 agentsAgentlessLocalizationRepair· Claude 3.5 SonnetPatch Validationsoftware-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
📐Aider Architect/EditorAider (Paul Gauthier)Self-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedTwo-role pipeline: reasoning architect + format-specialist editor — 85% SWE-bench.updated 1y agoPipeline2 agentsAiderArchitect· o1-previewEditor· o1-minisoftware-deliveryOutcome85%Economics[unknown] · deliberate1 proofCompare
🔀Anthropic Orchestrator-Workers PatternAnthropicSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedCentral LLM dynamically breaks down tasks and delegates to specialist workers.updated 1y agoOrchestrator–Worker2 agentsClaude APIOrchestratorWorkersoftware-deliveryresearchdata-extractionOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
💬AutoGen Group ChatMicrosoft Research (AutoGen)Self-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedFlexible multi-agent group conversation — hierarchical, peer, or proxy topologies.updated 3y agoSupervisor3 agentsAutoGenConversableAgentsoftware-deliveryresearchdata-extractionOutcome69.5%Economics[unknown] · deliberate1 proofCompare
💬ChatDev Communicative PipelineAcademic / Open-Source (ChatDev & MetaGPT)Self-ReportedAll claims are the subject's own. No external evidence is on record yet.Curated5-role sequential pipeline — 22,949 tokens, 148s per software task.updated 3y agoPipeline5 agentsChatDevCEOCTOProgrammerReviewer+1 moresoftware-deliveryOutcome[unknown]Economics22,9491 proofCompare
👥Claude Code Agent Teams (Experimental)AnthropicSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedSwarm teammates sharing a task list — parallel exploration on a single codebase.updated 1y agoSwarm3 agentsClaude CodeTeammate ATeammate BTeammate Csoftware-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🔱Claude Code Sub-agents PatternAnthropicSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedLead agent spawns specialist subagents in isolated context windows.updated 1y agoOrchestrator–Worker2 agentsClaude CodeLeadSubagentsoftware-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🏆Claude SWE-Bench TeamAnthropicSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedSingle-agent software engineer achieving 49% on SWE-bench Verified.updated 1y agoSolo + Tools1 agentClaude APISoftware Engineer· Claude 3.5 Sonnetsoftware-deliveryOutcome49%Economics[unknown] · deliberate1 proofCompare
🗂️CodeRHuawei Cloud + CAS + SMU + PKU (CodeR)Self-ReportedAll claims are the subject's own. No external evidence is on record yet.Curated5-agent supervised pipeline — 28.33% SWE-bench Lite, $3.09/issue.updated 2y agoSupervisor5 agentsCodeRManager· GPT-4-preview-1106Reproducer· GPT-4-preview-1106Fault Localizer· GPT-4-preview-1106Editor· GPT-4-preview-1106+1 moresoftware-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🤖Factory DroidFactory AISelf-ReportedAll claims are the subject's own. No external evidence is on record yet.Curated#1 on Terminal-Bench Core v0.1.1 (58.75%) — model-agnostic harness.updated 11mo agoSolo + Tools1 agentDroid (Factory AI)Solo Agentsoftware-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare