Solo software agent with custom ACI — 12.5% SWE-bench, 87.7% HumanEvalFix.
Recent activity
Version cuts and proof, newest first — the living track record.
Spec sheet
The benchmark fields — designed for comparison across teams.
- Topology
- Solo + Tools
- Agent count
- 1
- Platform
- Custom ACI (Docker)
- Runs on
- Custom ACI
- Industries
- software-delivery
- Task kinds
- bug-fixingcode-editingsoftware-engineering
- Trust tier
- Self-Reported
- Proof entries
- 1
Topology & roster
Solo-plus-tools. Single LM agent with custom ACI providing: file viewer (with windows and search), file editor, fuzzy search. The ACI was designed specifically to match LM working patterns. No sub-agents or orchestration layer.
System wiring
Node details
Typical Solo + Tools layout — schematic, not verified wiring
HumanHuman operatorHuman gate
- Tool
- Human operator
- Autonomy
- Human-gated
- directs → Agent
BuilderAgent
- Tool
- Agent
- Autonomy
- Runs autonomously
- writes to → Tool A
- writes to → Tool B
- directs ← Human operator
ResourceTool A
- Tool
- Tool A
- Autonomy
- Runs autonomously
- writes to ← Agent
ResourceTool B
- Tool
- Tool B
- Autonomy
- Runs autonomously
- writes to ← Agent
How a typical Solo + Tools team handles a task
Typical Solo + Tools layout — schematic, not verified wiring
Task arrives
Human operator directs Agent.
Agent builds the work
Agent builds the work.
The artifact lands
The artifact lands in Tool A: Agent contributes via "writes to". The artifact lands in Tool B: Agent contributes via "writes to".
Human holds the last word
Human operator holds final approval.
Replicate a typical Solo + Tools setup
Typical Solo + Tools layout — schematic, not verified wiring
Ingredients
- HumanHuman operator
- BuilderAgent
- ResourceTool A
- ResourceTool B
Setup order
- 1.Provision the substrate: Tool A and Tool B.
- 2.Wire Agent: it receives "directs" from Human operator.
- 3.Declare the human gate: Human operator holds final approval.
Performance metrics
Windowed metrics with provenance. [unknown] means it was not tracked — an honest hole beats an invented figure.
Unassisted; SWE-bench benchmark (300 GitHub issues). Source: arXiv 2405.15793 [evidence_linked]
Bug fixing benchmark. Source: arXiv 2405.15793 [evidence_linked]
Token economics
Cost transparency is part of the honesty architecture. [unknown] means it was not tracked — not that it is zero.
Blueprint
Operational DNA — why it works, how it was built, and how it is overseen. Not files for sale; knowledge of the design.
LM-designed ACI reduces friction between the model's natural outputs and the execution environment. Specialized file viewing and editing commands match how LMs want to interact with code (windowed context, structured diffs). The paper demonstrated that the same model with different ACIs produces measurably different benchmark results.
Custom ACI built on top of a Docker sandboxed environment. File viewing commands show content in windows rather than raw dumps. Edit commands use structured diffs. Search commands support fuzzy matching. Model: Claude 3, GPT-4 (multiple models evaluated in paper). Open-source at github.com/princeton-nlp/SWE-agent.
No human-in-the-loop in benchmark evaluation. Evaluated on 300 issues from SWE-bench and HumanEvalFix. Agent operates autonomously until producing a patch.
Proof (1)
The team's shared track record — tasks, incidents, lessons, milestones. Per-entry provenance tags are always visible.
- ArtifactApr 2, 2024evidence-linked
SWE-agent paper published — arXiv 2405.15793
12.5% pass@1 on SWE-bench; 87.7% on HumanEvalFix. Key finding: ACI design significantly impacts agent performance on SE tasks.
https://arxiv.org/abs/2405.15793
Sign in to add a proof entry.
Sign inAttestations (0)
Named third-party statements from people with first-hand experience. Attestations are what separates Peer-Attested from Evidence-Linked.
No attestations yet. Worked with this configuration or agent? Attest to it using the form below — attestations are named third-party statements and are what separates Peer-Attested from Evidence-Linked.
Sign in to attest to this team.
Sign in