Skip to content

Blog

The leaderboard assumed a complete issue

SWE-bench Verified scores agents on a finished GitHub issue. Real use starts mid-thought. Dialogue-SWEBench measures whether agents know when to ask — and when to stop.

If you follow coding agents at all, you have seen the SWE-bench numbers. They keep going up. Verified sits near the mid-nineties for the strongest systems. The story those numbers tell is simple: agents can open a real repository, read a GitHub issue, write a patch, and pass the tests.

That story is incomplete. The issue on the leaderboard is a finished document. The issue in your product is usually not.

King and Flanigan (UC Santa Cruz, June 2026) built Dialogue-SWEBench to measure what happens when you take that finished document away. They start from the same five hundred Verified tasks, but the agent only gets a stripped first message — something closer to “this crashes when the input is empty, can you look?” — and a message_user tool. Resolve rate on a full specification is no longer the whole product. The product is whether the agent knows when to ask, and when to stop asking.

What SWE-bench actually measures

SWE-bench, from Jimenez and colleagues, is the main public yardstick for repository-level coding agents. Each task is a real GitHub issue from a real Python project, plus a base commit and a suite of tests. An agent explores the repo, edits files, and submits a patch. If the tests pass, the task counts as resolved. SWE-bench Verified is a cleaned subset of five hundred of those tasks where the issue text is judged complete and correct enough that a capable engineer could solve it from the write-up alone.

That last clause is doing a lot of work. Chowdhury and coauthors, in the Verified paper itself, report that about 76% of the original SWE-bench issues are at least somewhat underspecified, and 39% are too vague to say what a successful solution would even look like. Verified exists because the raw set was noisy. The leaderboard that everyone quotes is the cleaned set — the cases where the specification is already good.

That is a legitimate scientific choice. It is a bad model of how people actually use Cursor, Claude Code, Copilot Workspace, and the rest. Baumann and colleagues studied real SWE-chat sessions and found users correcting or rejecting agent output about 44% of the time. Agents sought clarification in only 1–2% of turns. The conversation is one-sided. The human is doing the judgment. The agent is mostly not.

That is the same thesis I argued in Software is changing: writing valid code is getting cheap; the scarce skill is judgment — what to build, what “done” means, when the brief is wrong. Dialogue-SWEBench is that thesis as a benchmark. Last week’s post was about whether an eval can also be your ranker. This one is about whether a coding agent can tell when it does not know enough to patch.

What they built

The authors keep the Verified tasks and the Verified tests. They change the interface. Instead of dumping the full issue into the agent’s first observation, they craft a short initial query that preserves intent but strips the details an agent would need to skip talking. Then they add a user to talk to.

That user is a simulator: a Llama-3.3-70B model with a persona, grounded on the full issue text the agent never sees, plus a self-revision step that catches hallucinations like “I ran the tests and they passed.” On a human audit, 97.5% of dialogues were defect-free on faithfulness, goal adherence, and environment limits. Treat that as a high-quality harness, not a field study of real developers. It is closer to a careful TREC assessor simulation than to a user interview.

They compare three agent scaffolds that share the same tools. Stock OpenHands almost never asks questions and resolves 32.9% of tasks. OH Interactive, built for clarifying underspecified specs, reaches 44.1%. Their own schema-guided agent — which forces the model to maintain an explicit checklist of what it still does not know — averages 46.9%. Information-seeking turns correlate with resolve rate. The agents that ask more, resolve more.

The model results are the part that should make product teams uncomfortable. GPT-5-mini matches GPT-5 overall on this benchmark (58.8% vs 58.0% on the schema-guided agent). GPT-5 still wins the hard engineering tasks. It loses the easy ones — the “under fifteen minutes” fixes — by asking too many questions at once, then failing to follow up when the user answers only two of them. Stronger coding models are not automatically stronger dialogue models. Naturalness and coherence, scored separately, do not track resolve rate either. A patch that passes the tests can still be a bad conversation.

What this means if you ship coding agents

Most of the industry is still optimizing the Verified number. That number answers a real question: given a complete issue, can the agent produce a correct patch? Useful. Incomplete. In production, the issue arrives as a Slack thread, a half-written ticket, a screenshot, or “it worked Friday.” The eng lead’s job is not to paste a perfect prompt. It is to get something shipped without babysitting every turn.

If you only measure resolve rate on a full spec, you will ship an agent that never asks — and then treat the 1–2% clarification rate as a user problem. That is backwards. The clarification rate is a product failure. The agent that dumps twelve questions in one message and then ignores the unanswered ones is also a product failure. Dialogue-SWEBench surfaces both failure modes on the same axis.

There is a second, quieter implication for how we evaluate these systems at all. Resolve rate is execution feedback. Naturalness and coherence are judgment about the conversation. King and Flanigan keep them separate, and they should. A system can be “good at SWE-bench” and still be exhausting to work with. If your internal dashboard only shows green tests, you will miss that.

What not to claim

Do not install 46.9% as “the new industry number” for Cursor or Claude Code. Different harnesses, different tools, different user models. The user here is a simulator with the gold issue in its pocket; a real developer does not always know the answer either. Dialogue quality is a separate axis from whether the patch merges. And Verified itself is a Python-heavy, issue-type-skewed slice of software engineering — Dialogue-SWEBench inherits that skew.

What you can claim is narrower and more useful. The autonomous leaderboard assumed a complete GitHub issue. Real users show up mid-thought. Measuring only the first setting will push you to ship agents that patch confidently and ask almost never. Measuring the second setting — even imperfectly, even with a simulator — pushes you toward agents that treat clarification as part of the job.

Resolve rate on a full specification is not the product. The product is whether the agent knows when to ask, and when to stop asking.

Papers: Dialogue-SWEBench (King and Flanigan, 2026); site. Context: SWE-bench (Jimenez et al.); SWE-bench Verified (Chowdhury et al.); SWE-chat (Baumann et al., 2026).