~/shanegraffiti.com/research/dialogue-swebench Semantic Adversarial Research Division 2026

DIALOGUE—SWEBENCH CODING AGENTS

AI coding agents are benchmarked as fully-autonomous systems but real-world use is interactive. Users correct and reject agent outputs 44% of the time. Agents seek clarification 1–2% of the time. Dialogue-SWEBench closes that gap: 500 real SWE-Bench problems, no pre-given specification, resolved entirely through multi-turn dialogue with a persona-grounded user simulator. Better coding models are not always better dialogue models.

Division
Semantic Adversarial Research
Domain
Coding Agents / Dialogue Eval
Published
arXiv 2606.13995
Key Result
Schema agent: +3–14% over baseline
Dialogue-Driven SWE◆Schema-Guided Agents◆User Simulation◆Self-Revision◆Information Seeking◆Persona Grounding◆OpenHands◆SWE-Bench Verified◆Resolve Rate◆Dialogue Naturalness◆Dialogue Coherence◆LLM-as-a-Judge◆
§ 1.0

The Dual-Environment Setup

The benchmark splits the coding agent's world into two environments. The user never touches the code. The agent lives in both — it must coordinate dialogue state and repository state simultaneously.

User Channel ENV 01

Dialogue Environment

The agent communicates with a simulated user via message_user actions. The task starts with a vague initial query — no issue text, no specification document. All problem details must be elicited through conversation.

User corrects/rejects agent output 44% of the time in real sessions. Agents seek clarification 1–2% of the time.
Code Shell ENV 02

Repo Environment

The agent operates in a programming environment — editing files, running bash, executing tests — and submits a git patch via the finish action. Evaluated against execution tests from SWE-Bench Verified.

500 problems · 100 step limit · scored on patch resolution rate
§ 2.0

Schema-Guided Agent

Off-the-shelf coding agents almost never seek information from users. The schema-guided agent addresses this by building and maintaining a structured dialogue state representation that drives what to ask next and when to code.

{{ step.num }} {{ step.tag }}

{{ step.title }}

{{ step.body }}

// dialogue_state — step {{ activeStepNum }}
{ "issue_type": {{ schemaType }}, "observed": {{ schemaObserved }}, "expected": {{ schemaExpected }}, "repro": {{ schemaRepro }} }▋
§ 3.0

User Simulator

The user simulator (LLaMA 3.3 70B) is grounded in the full issue text but never shares it directly. A self-revision step validates each candidate reply for hallucinations, environment boundary violations, and length — then revises if needed. 97.5% of dialogues are defect-free.

{{ stat.label }}
{{ stat.caption }}
Simulator ablation — schema-guided agent, n=50
{{ row.model }} {{ row.drop }}
Full simulator u₁ only, no follow-up
Persona types assigned per problem
Persona {{ activePersonaLetter }}

{{ activePersonaName }}

{{ activePersonaBody }}

§ 4.0

Experimental Results

500 SWE-Bench Verified problems. 4 models × 3 agents = 12 systems. Information-seeking dialogue moves are the strongest predictor of resolve rate — not coding model size.

Chart metric
ModelAgent{{ activeMetricLabel }}% ResTurnsStepsCost
{{ row.model }} {{ row.agent }} {{ row.barLabel }} {{ row.resolved }} {{ row.turns }} {{ row.steps }} {{ row.cost }}
§ 5.0

Dialogue Quality — LLM-as-a-Judge

Task resolution alone does not measure usability. Two automatic quality dimensions — Naturalness and Coherence — evaluated by Gemma 4 31B-IT and validated against human annotation of 360 dialogues.

Dimension 01

Naturalness

Degree to which the agent is easy to understand and converse with. Scored 1–3. More variance from model choice than agent design. GPT-5 suffers from failure to close dialogues and leaking internal system prompt text to the user.

κ = 0.70 · 100% rank accuracy
Dimension 02

Coherence

Degree to which dialogue moves guide conversation toward resolving the task. Scored via local coherence (each turn follows logically) and global coherence (conversation arc steers toward resolution). Stronger differentiation between agents than naturalness — agent design matters more here.

κ = 0.51 · 84.3% rank accuracy
Cross-cut

Task vs. Quality

No clear relationship between naturalness and resolution rate. Usability of a coding agent in dialogue cannot be measured by task success alone — a model can solve the task with poor dialogue or communicate naturally while failing to resolve it.

Independent signal
§ 6.0

Key Findings

Four findings that distinguish Dialogue-SWEBench from fully-autonomous SWE evaluation. [click a row to expand]

{{ f.num }}

{{ f.title }}

{{ f.tag }} +

{{ f.body }}

~/conclusion
$ query: what does Dialogue-SWEBench change
// Coding agents are not evaluated on what they actually do.
// Real use is interactive. Benchmarks have not been.
// Dialogue capability is a distinct, currently understudied dimension.
$ query: what is the cost of better dialogue
// Fewer steps, not more. Schema agents reduce total agent steps.
// Targeted questions are cheaper than exploratory dead-ends.
// Best average resolve rate at the lowest average cost.
$ query: what is the actual result
// GPT-5 mini matches GPT-5. Coding ≠ Dialogue.
// +3–14% over baselines. One schema. Any coding agent.▋

BETTER CODING IS NOT BETTER DIALOGUE.

Shane Graffiti Inc. — Semantic Adversarial Research Division Top ↑arXiv 2606.139952026