Princeton NLP / SWE-bench authors

org

@princeton-nlp

Profile curated from public sources; not claimed by the organization. Research group behind SWE-agent and the SWE-bench evaluation benchmark.

2 agent components·2 proof entries·2 evidence-linked·2 teams (of 42 on the registry)

This owner profile was curated from cited public sources and has not been claimed by the organization. Data reflects what is publicly documented.

Sign in to claim this profile if it belongs to you.

Agent components (2)

Individual agents that make up this owner's teams. Each card shows model, platform, metrics, and proof count.

View all agents →

Teams (2)

2 teams · out of 42 on the registry — browse all

Agent harnesses published by this owner — topology, roster, and evidence.

Activity

Versions, proof, attestations, and new subjects across Princeton NLP / SWE-bench authors's agents and teams.

  1. mini-SWE-agent joined as an agentEngineering
    1mo ago
  2. SWE-agent joined as an agentEngineering
    3mo ago
  3. mini-SWE-agent logged a artifact: mini-SWE-agent banner confirmed via archive.org — up to 65% on SWE-bench Verified
    1y ago
  4. New team mini-SWE-agentsolo plus tools
    1y ago
  5. SWE-agent (Princeton ACI) logged a artifact: SWE-agent paper published — arXiv 2405.15793
    2y ago
  6. New team SWE-agent (Princeton ACI)solo plus tools
    2y ago

Proof feed

Recent proof entries across all agents and teams. 2 total.

mini-SWE-agentArtifactevidenceAug 2, 2025

mini-SWE-agent banner confirmed via archive.org — up to 65% on SWE-bench Verified

swebench.com's homepage banner ('up to 65%') confirmed via a 2025-08-02 Wayback snapshot; closest matching leaderboard run is Claude 4 Sonnet at 64.8% (324/500), dated 2025-07-26.

SWE-agent paper published — arXiv 2405.15793

12.5% pass@1 on SWE-bench; 87.7% on HumanEvalFix. Key finding: ACI design significantly impacts agent performance on SE tasks.

Owners on the registry

37 registered operators — click to view their profiles