0:00 on page · 0% scrolled
↓
~/shanegraffiti.com/research/garmentsketch Shane Graffiti Inc. Semantic Adversarial Research Division 2026

GARMENT SKETCH A SKETCH-TO-FASHION BENCHMARK

Fashion sketching lets designers visualise a concept long before any fabric is cut, yet sketch-based fashion image synthesis has stalled for want of large-scale, high-quality paired data. GarmentSketch closes that gap: 26,249 fashion sketches across 21 garment categories, each paired with a detailed textual description. Captions were produced through a multi-stage pipeline combining several multimodal language models with human-in-the-loop refinement, balancing semantic accuracy against descriptive richness. Benchmarking state-of-the-art generators on the set exposes both the promise and the present limits of sketch-guided text-to-image generation and a clean trade-off between photorealism and faithfulness to the drawn line.

Division
Semantic Adversarial Research
Domain
Sketch-Guided T2I / Fashion
Published
arXiv 2606.14025 2026
Scale
26,249 sketches · 21 categories
Sketch-to-Fashion◆Sketch-Guided T2I◆Rich Captions◆Informative Drawings◆26,249 Sketches◆21 Categories◆FID · LPIPS · CLIPScore◆ControlNet◆T2I-Adapter◆Photorealism vs Structure◆Human-in-the-Loop◆Design-Oriented Generation◆
§ 1.0

The Modality Gap

A sketch is sparse: a few abstract lines, no texture, no colour, no material cues. A garment photograph is dense with exactly those signals. Generators trained on general-purpose data struggle to span this gap they lose fabric draping, silhouette, and decorative detail, and lack the fashion-specific knowledge a professional workflow needs.

The central hypothesis: aligning rich textual semantics with sparse sketches bridges the modality gap, letting a model preserve structure and synthesise complex detail at the same time. That requires paired data which did not previously exist at scale.

{{ side.tag }}
{{ side.name }}
{{ side.desc }}
+ caption bridges
§ 2.0

Where It Sits

Prior fashion datasets advance recognition, retrieval, and virtual try-on, but omit the sketch-plus-caption pairings needed for early-stage design. General sketch datasets, in turn, are either too simplistic or aimed at generic object retrieval rather than fine-grained garment structure.

Fashion & sketch datasets coverage — click Size to sort
DatasetModalitySize {{ sortArrow }}SketchesRich captions
{{ d.name }} {{ d.modality }} {{ d.size }} {{ d.sketches }} {{ d.captions }}
§ 3.0

The Construction Pipeline

Two parallel workflows turn a source image into a sketch–caption pair. Sketches come from an anime-style Informative-Drawings model about 6 seconds per image, versus roughly 10 minutes for stroke-optimisation tools, and far better at capturing intricate garment detail. Captions are synthesised by three multimodal models, then consolidated and human-verified.

Branch A sketch generation
Source garment image
↓
Informative-Drawings (anime style)
↓
Informative sketch (~6s)
{{ nodeADetail }}
Branch B caption synthesis
Source image + brief caption
↓
LlamaGemmaQwen
↓
Consolidate & polish (Gemma) → human verify
↓
Rich merged caption
{{ nodeBDetail }}
A ⊕ B → 26,249 sketch–caption pairs · 70 / 30 train–test split per category
§ 4.0

Inside the Dataset

Sources span Western e-commerce imagery, an upperwear top-up, and 650 hand-collected images of Eastern traditional dress such as the Vietnamese áo dài for cultural balance. The authors are candid about two biases: the set skews toward accessories, and remains Western-centric overall.

26,249
Sketch–caption pairs
21
Garment categories
27.9%
Shoes the largest single class
~2.5%
Áo dài cultural-inclusion subset

Shoes (27.9%) and bags (11.6%) together make up nearly 40% of the data, while core apparel like upperwear (8.06%) and bottomwear (10.2%) is thinner. The skew means models may learn the fixed structures of accessories more readily than clothing a limitation the authors flag as motivation for more globally balanced curation.

{{ cat.name }} {{ cat.pct }}
§ 5.0

Benchmark Results

Four sketch-to-image models were evaluated zero-shot on the test set: Gemini 2.5 Nano Banana, ControlNet Scribble SDXL, ControlNet Scribble SD1.5, and T2I-Adapter Sketch SDXL. FID measures quality and diversity, LPIPS perceptual similarity to ground truth, CLIPScore semantic agreement with the prompt.

Sorted by FID ↓ — lower is better
ModelFIDValue
Gemini 2.5 Nano BananaLPIPS 0.40 · CLIPScore 29.50 17.50
T2I-Adapter Sketch SDXLLPIPS 0.49 · CLIPScore 28.30 23.50
ControlNet Scribble SDXLLPIPS 0.60 · CLIPScore 30.10 29.18
ControlNet Scribble SD1.5LPIPS 0.70 · CLIPScore 29.02 36.08

Per category, Gemini takes best FID in 17 of 21 classes and best LPIPS in 20 of 21. CLIPScore is more balanced: ControlNet Scribble SDXL leads 8 categories, Gemini 7, T2I-Adapter 4. SD1.5 trails throughout and stumbles hardest on culturally specific garments like the áo dài, reflecting limited diversity in its training data.

§ 6.0

The Core Trade-off

No model wins on every axis. The benchmark surfaces one clean tension: the more photorealistic a model's output, the more it tends to drift from the drawn structure and vice versa. Each model lands somewhere along this line.

Structural
FidelityFaithful to the sketch's lines
Photo-
realismRich texture, lighting, polish
{{ mk.name }}{{ mk.caption }}

ControlNet and T2I-Adapter hold the sketch's shape but lack polish; Gemini produces the most appealing images yet frequently diverges from the intended silhouette and fine pattern. The open problem: a model that earns both ends of the line at once.

{{ activeMarkerName }}

{{ activeMarkerDetail }}

"A sparse line and a rich word bridge the gap." — §6.0, The Core Trade-off
~/conclusion
$ query: why did sketch-to-fashion stall
// no large-scale data pairing sketch, caption, and image.
$ query: what does GarmentSketch supply
// 26,249 sketches · 21 categories · rich captions.
// curated by an MLLM pipeline with human verification.
$ query: what did benchmarking reveal
// a hard trade-off: photorealism vs structural fidelity.
// no current model holds both. that is the next problem.▋

A SPARSE LINE AND A RICH WORD BRIDGE THE GAP.

Shane Graffiti Inc. Semantic Adversarial Research Division Top ↑arXiv 2606.140252026