~/shanegraffiti.com/research/dlawbench Shane Graffiti Inc. Semantic Adversarial Research Division 2026

DLAWBENCH MULTI-TURN LEGAL CONSULTATION

Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients then reasoning from what's missing. DLawBench evaluates whether LLMs can conduct real legal consultation under realistic conditions: eliciting facts, correcting client misframes, and writing defensible memos. Built from 461 real court opinions across Chinese and U.S. law, with four client personality types and expert-authored evaluation rubrics. The best-performing model achieves only 0.562 in consultation-grounded legal reasoning.

Division
Semantic Adversarial Research
Domain
Legal AI / Multi-Turn Evaluation
Published
arXiv 2606.13931 · 2026
Key Result
Top model: 0.562 Resolution
Legal Consultation◆Multi-Turn LLM Eval◆Sycophancy Detection◆Fact Elicitation◆Client Narrative Styles◆Court-Record Grounding◆Abductive Reasoning◆Weak Supervision◆Chinese Law◆U.S. Federal Law◆Belief–Record Separation◆Issue Resolution◆
§ 1.0

The Consultation Pipeline

Legal consultation interleaves information gathering with legal reasoning. A lawyer rarely receives a complete, neutral, chronologically ordered fact pattern. DLawBench evaluates three linked but separable abilities across the full pipeline.

Phase 1 · Elicitation 01

Information
Gathering

The lawyer model interviews the simulated client up to 10 turns asking targeted follow-up questions. Scored on Fact Coverage (facts reaching the memo) and Inquiry (case-specific intake questions asked).

→
Phase 2 · Resolution 02

Legal
Reasoning

The model submits a structured legal analysis memo. Scored on Fact Resolution (facts correctly reframed against the hidden court record) and Issue Resolution (legal analysis points addressed).

→
Phase 3 · Fidelity 03

Claim
Support

Whether factual and inferential claims in the memo are grounded in the consultation or the case record not invented. High fidelity with wrong legal route is still a failure.

Metric hierarchy · DLawBench
Elicitation = (Fact Coverage + Inquiry) / 2 // whether the lawyer gathered the information needed for analysis
Resolution = (Fact Resolution + Issue Resolution) / 2 // whether gathered facts become correct legal analysis
Fidelity = 1 − (unsupported claims / total claims) // claim support guardrail high fidelity does not imply correct routing
§ 2.0

Client Narrative Styles

Each of 461 cases is replayed under four narrative styles adapted from the Interpersonal Circumplex. The underlying facts are fixed. Only how much the client volunteers and resists changes and that changes everything.

Interpersonal Circumplex
High Agency Low Agency Low Comm. High Comm.
{{ activeAxis }} {{ activeTag }}

{{ activeName }}

{{ activeBody }}

Client utterance
{{ activeQuote }}
§ 3.0

Sycophancy Taxonomy

In legal settings sycophancy is not simple agreement it is a structured, multi-level failure. Models deploy legal knowledge in service of an unchecked client frame, producing professionally packaged error. DLawBench makes this operationally measurable.

[ drag to scroll ]
{{ c.level }} {{ c.num }}

{{ c.title }}

{{ c.body }}

{{ c.signal }}
§ 4.0

Experimental Results

26 models evaluated across 461 cases × 4 narrative styles = 1,844 case-style cells. Three-judge panel: GPT-5.1, Claude Opus 4.6, Gemini 3.1 Pro. Same-vendor judges recuse. Scored by median aggregation.

0.562
Top model Resolution (GPT-5.5)
0.934
Top model Fidelity not enough
−10pp
Inquiry drop, Withdrawn vs Cooperative
0.076
LegalOne-8B Resolution domain tuning alone fails
Leaderboard ranked by {{ sortLabel }} — click a column to re-rank avg. Chinese + U.S. law
# Model
{{ r.rank }} {{ r.model }} {{ v.txt }}
Coverage ≠ Reasoning · GLM-5.1 diagnostic
{{ d.metric }} {{ d.score }}

{{ d.note }}

§ 5.0

What DLawBench Exposes

Without a hidden court record, a consultation may appear successful simply because the final advice sounds helpful. DLawBench penalizes user-responsive legal reasoning that lacks independent legal judgment.

Failure patternHow it appears in metricsVerdict

{{ f.title }}

{{ f.body }}

{{ f.verdict }}
§ 6.0

What DLawBench Enables

By separating client-belief and court-record views, DLawBench turns consultation failure into diagnosable capability gaps. Four diagnostic handles not available in prior benchmarks.

{{ e.num }}
Handle {{ e.num }}

{{ e.title }}

{{ e.body }}

~/conclusion
$ query: what does DLawBench change
// Legal consultation is not legal QA over complete facts.
// The symbolic gap is between knowing law and recovering facts.
// A hidden court record is required to see the real failure.
$ query: what do the results cost
// Under exact cooperative clients: inflated. Looks fine.
// Under Dependent and Withdrawn: real drop, real signal.
// Tune the model on client style, not just legal knowledge.
$ query: what is the actual result
// Best model: 0.562 Resolution. Substantial headroom.
// One framework. Any client. No free pass on legal reasoning.▋

THE CLIENT IS NOT THE RECORD.

Shane Graffiti Inc. — Semantic Adversarial Research Division Top ↑arXiv 2606.139312026