<!-- Generated from /benchmarks at build time. Do not edit by hand. -->
<!-- Source of truth: the rendered page. A hand edit here is overwritten on the next build. -->

# Designesy — /benchmarks

Canonical page: https://www.designesy.org/benchmarks

---

## Benchmarks

Three tools, three questions. They do not compete — they answer different questions about design quality.

Hallmark prevents slop at generation time. Slop-eval scores slop after the fact. Designesy verifies contract conformance. This page maps what each catches that the others do not.

## The three tools

## They do not compete

A complete design quality pipeline uses all three: hallmark to generate, slop-eval to catch slop, designesy to verify the contract. They occupy different positions on the design-verification spectrum — not the same position.

Generation → hallmark (prevent slop at emit time, 57 gates)

Evaluation → slop-eval (score existing designs, 108 tells + 2 positive axes)

Verification → designesy (verify contract conformance, 42 checks + MCP delivery)

## Shared checks

Where the three tools overlap. The intersection is narrow — most checks are unique to each tool.

## What Designesy catches that they do not

20 checks with no hallmark or slop-eval equivalent. These are the uncontested layers — contract conformance, Unicode security, Core Web Vitals, forced-colors, AI disclosure, and the Cadence typography suite.

## What they catch that Designesy does not

Designesy gaps. These are the AI-slop aesthetic detection, structural-variety, signature, and cohesion checks. Designesy does not ask “does this look AI-generated” — it verifies contract conformance.

## The moat

The layers no competitor approaches. These are not features — they are the category-of-one positioning.

## Self-score on designesy.org

hallmark and slop-eval are agent skills, not URL-based APIs. They cannot score a URL without an LLM agent loading the skill and evaluating. Designesy is the only one with a programmatic URL-based scoring endpoint — a structural advantage for CI/CD integration, leaderboards, and automated pipelines.

## The wider landscape

The three tools above are the closest comparators. The field is growing — these are newer entrants worth tracking.

81 design-related MCP servers are now indexed at mcpservers.org (2026-08-01). The category is expanding from skills into agent-invocable tooling — the space designesy pioneered with its 17-tool MCP server and URL-based scoring API.

## Provenance

Live research via AnySearch MCP (2026-08-01). Designesy live score verified via POST to https://www.designesy.org/api/score (93% A,42 checks). hallmark taxonomy from Nutlope/hallmark/skills/hallmark/references/slop-test.md. slop-eval taxonomy from fabricioctelles/skills/skills/slop-eval/references/tells.md.
