SWE-bench Verified data contradicts autonomous coding claims
The latest SWE-bench Verified leaderboard shows top models resolving under 30 percent of real GitHub issues without human intervention. This metric directly challenges vendor narratives regarding fully autonomous software development pipelines. Founders allocating engineering budget based on these promises should recalibrate their runway projections.
0 comments
0