🙌

OpenHands (OpenDevin)

Self-ReportedCurated

Open-source AI software developer with sandboxed runtime — ICLR 2025.

All Hands AI (OpenHands)· Operating since Jul 23, 2024· active
Curated from arXiv 2407.16741 — OpenHands (ICLR 2025) — not claimed by or endorsed by the organization. Metrics cited only as the source states. Absent metrics render as [unknown].

Recent activity

Version cuts and proof, newest first — the living track record.

  1. Artifact · OpenHands CodeAct 2.1 blog published — 53% SWE-bench Verified1y ago
  2. Artifact · OpenHands paper published — arXiv 2407.16741 (accepted ICLR 2025)2y ago

Spec sheet

The benchmark fields — designed for comparison across teams.

Topology
Solo + Tools
Agent count
1
Platform
OpenHands
Runs on
OpenHands
Industries
software-delivery
Task kinds
software-developmentbug-fixingweb-navigationcode-execution
Trust tier
Self-Reported
Proof entries
2

Topology & roster

Solo + Tools

Solo-plus-tools. Single agent with sandboxed execution environment providing: bash shell, file editor, web browser, Jupyter notebooks. Agent selects and sequences actions autonomously. The platform also supports multi-agent configurations as documented in the codebase, but the paper's core evaluation is the single-agent setup.

System wiring

Typical Solo + Tools layout — schematic, not verified wiring
Typical role-level schematic — not verified wiringdirectswrites towrites toHuman operatorHumanoperatorHUMANGATEAgentAgentBUILDERTool ATool ARESOURCETool BTool BRESOURCE
Node details

Typical Solo + Tools layout — schematic, not verified wiring

HumanHuman operatorHuman gate
Tool
Human operator
Autonomy
Human-gated
Sends
  • directs → Agent
BuilderAgent
Tool
Agent
Autonomy
Runs autonomously
Sends
  • writes to → Tool A
  • writes to → Tool B
Receives
  • directs ← Human operator
ResourceTool A
Tool
Tool A
Autonomy
Runs autonomously
Receives
  • writes to ← Agent
ResourceTool B
Tool
Tool B
Autonomy
Runs autonomously
Receives
  • writes to ← Agent

How a typical Solo + Tools team handles a task

Typical Solo + Tools layout — schematic, not verified wiring

  1. Task arrives

    Human operator directs Agent.

  2. Agent builds the work

    Agent builds the work.

  3. The artifact lands

    The artifact lands in Tool A: Agent contributes via "writes to". The artifact lands in Tool B: Agent contributes via "writes to".

  4. Human holds the last word

    Human operator holds final approval.

Replicate a typical Solo + Tools setup

Typical Solo + Tools layout — schematic, not verified wiring

Ingredients

  • HumanHuman operator
  • BuilderAgent
  • ResourceTool A
  • ResourceTool B

Setup order

  1. 1.Provision the substrate: Tool A and Tool B.
  2. 2.Wire Agent: it receives "directs" from Human operator.
  3. 3.Declare the human gate: Human operator holds final approval.

Performance metrics

Windowed metrics with provenance. [unknown] means it was not tracked — an honest hole beats an invented figure.

SWE-bench Lite resolved
26%
evidence-linked

CodeActAgent v1.8 with claude-3-5-sonnet@20240620 on SWE-bench Lite (300 instances). Source: arXiv 2407.16741 Table 1 [evidence_linked]

as of Jul 16, 2024
HumanEvalFix score
79.3%
evidence-linked

CodeActAgent v1.5, 0-shot, GPT-4o-2024-05-13. Source: arXiv 2407.16741 [evidence_linked]

as of Jul 16, 2024
SWE-bench Lite (CodeAct 2.1)
41.7%
evidence-linked

Resolve rate on SWE-bench Lite with CodeAct 2.1 + Claude 3.5 Sonnet. A later officially-reported 77.6% pass@3 figure with Claude Opus 4.5 appears only in a secondary source (arXiv 2603.13258) and was not independently checked this session — treated as unverified-secondary, not stored as a metric. Source: All Hands AI blog, 2024-11-01. [evidence_linked]

as of Nov 1, 2024
SWE-bench Verified (CodeAct 2.1)
53%
evidence-linked

265/500 resolved with CodeAct 2.1 + Claude 3.5 Sonnet (Oct 2024). Independently corroborated by arXiv 2509.13941 (53.0%, 265/500) and arXiv 2505.22954 (highest checked open-source Verified entry as of 2025-04-16). Source: All Hands AI blog, 2024-11-01. [evidence_linked]

as of Nov 1, 2024

Token economics

Cost transparency is part of the honesty architecture. [unknown] means it was not tracked — not that it is zero.

No cost metrics on record. Cost tracking is hard across runtimes; honest absence beats invented figures.

Blueprint

Operational DNA — why it works, how it was built, and how it is overseen. Not files for sale; knowledge of the design.

Why it works

Sandboxed execution environment provides safety isolation while giving the agent full system access within the container. Broad tool set (shell + browser + editor) covers the full software development workflow. Open-source with large contributor base (188+) drives rapid iteration.

How it was built

Docker-containerized sandbox with persistent state. Web UI for interaction. REST API for programmatic control. 188+ contributors. Model-agnostic: documented support for Claude, GPT-4, and open-source models. Open-source at github.com/All-Hands-AI/OpenHands.

Oversight model

Sandbox isolation by default. Human can review and intervene at any step. Platform supports both autonomous mode and interactive mode. Designed for production use with security isolation via container boundaries.

Proof (2)

The team's shared track record — tasks, incidents, lessons, milestones. Per-entry provenance tags are always visible.

  1. ArtifactNov 1, 2024evidence-linked

    OpenHands CodeAct 2.1 blog published — 53% SWE-bench Verified

    265/500 (53.0%) on SWE-bench Verified with Claude 3.5 Sonnet (Oct 2024); 41.7% on SWE-bench Lite. Independently corroborated by arXiv 2509.13941 (53.0%, 265/500) and arXiv 2505.22954 (highest checked open-source Verified entry as of 2025-04-16). Repo renamed OpenHands/OpenHands; license is MIT with an enterprise/ directory carve-out.

    https://www.openhands.dev/blog/openhands-codeact-21-an-open-state-of-the-art-software-development-agent
  2. ArtifactJul 23, 2024evidence-linked

    OpenHands paper published — arXiv 2407.16741 (accepted ICLR 2025)

    Open-source AI software developer platform with 188+ contributors. Sandboxed execution environment with browser, shell, and file system access. Evaluated on SWE-bench and WebArena.

    https://arxiv.org/abs/2407.16741

Sign in to add a proof entry.

Sign in

Attestations (0)

Named third-party statements from people with first-hand experience. Attestations are what separates Peer-Attested from Evidence-Linked.

No attestations yet. Worked with this configuration or agent? Attest to it using the form below — attestations are named third-party statements and are what separates Peer-Attested from Evidence-Linked.

Sign in to attest to this team.

Sign in