🙌OpenHands (OpenDevin)All Hands AI (OpenHands)Self-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedOpen-source AI software developer with sandboxed runtime — ICLR 2025.updated 1y agoSolo + Tools1 agentOpenHandsAI Developersoftware-deliveryOutcome26%Economics[unknown] · deliberate2 proofCompare
🏆Claude SWE-Bench TeamAnthropicSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedSingle-agent software engineer achieving 49% on SWE-bench Verified.updated 1y agoSolo + Tools1 agentClaude APISoftware Engineer· Claude 3.5 Sonnetsoftware-deliveryOutcome49%Economics[unknown] · deliberate1 proofCompare
🤖Factory DroidFactory AISelf-ReportedAll claims are the subject's own. No external evidence is on record yet.Curated#1 on Terminal-Bench Core v0.1.1 (58.75%) — model-agnostic harness.updated 11mo agoSolo + Tools1 agentDroid (Factory AI)Solo Agentsoftware-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🧠h2oGPTe AgentH2O.aiSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedEnterprise agent + majority voting — claimed #1 on GAIA, Dec 2024.updated 1y agoSolo + Tools1 agenth2oGPTeSolo Agent· Claude 3.5 Sonnetenterprise-aiOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
💬Klarna AI AssistantKlarnaSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedTwo-thirds of customer-service chats in month 1 — 2.3M conversations.updated 2y agoSolo + Tools1 agentCustomSolo Agentcustomer-supportfintechOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🐚mini-SWE-agentPrinceton NLP / SWE-bench authorsSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedMinimal ~100-line bash-only agent — reference floor on SWE-bench Verified.updated 1y agoSolo + Tools1 agentmini-SWE-agentSolo Agentsoftware-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🛠️Refact.ai AgentRefact.aiSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedSolo agent, multi-model orchestration — 70.4% then 74.4% on SWE-bench Verified.updated 1y agoSolo + Tools1 agentRefact.aiSolo Agent· Claude 3.7 Sonnet / Claude 4 Sonnet (multi-model orchestration)software-deliveryOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🐍smolagents CodeAgentHugging FaceSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedPython-first multi-provider agent — minimal 1K-line library, MCP + LangChain tools.updated 1y agoSolo + Tools1 agentsmolagents (Python)CodeAgentsoftware-deliveryresearchdata-extractionOutcome[unknown]Economics[unknown] · deliberate1 proofCompare
🐛SWE-agent (Princeton ACI)Princeton NLP / SWE-bench authorsSelf-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedSolo software agent with custom ACI — 12.5% SWE-bench, 87.7% HumanEvalFix.updated 2y agoSolo + Tools1 agentCustom ACI (Docker)Software Engineersoftware-deliveryOutcome12.5%Economics[unknown] · deliberate1 proofCompare
⛏️Voyager (Minecraft)NVIDIA Research (Voyager)Self-ReportedAll claims are the subject's own. No external evidence is on record yet.CuratedOpen-ended Minecraft agent — 3.3× items, 15.3× tech tree vs prior SOTA.updated 3y agoSolo + Tools1 agentGPT-4 API (Mineflayer/Minecraft)Explorer· GPT-4gamingresearchOutcome[unknown]Economics[unknown] · deliberate1 proofCompare