Three tools, three questions. They do not compete — they answer different questions about design quality.
Hallmark prevents slop at generation time. Slop-eval scores slop after the fact. Designesy verifies contract conformance. This page maps what each catches that the others do not.
Score on designesy.org93% A (39/42 PASS)Not run (no URL API)Not run (no URL API)
They do not compete
The positioning
A complete design quality pipeline uses all three: hallmark to generate, slop-eval to catch slop, designesy to verify the contract. They occupy different positions on the design-verification spectrum — not the same position.
01
Generation → hallmark (prevent slop at emit time, 57 gates)
20 checks with no hallmark or slop-eval equivalent. These are the uncontested layers — contract conformance, Unicode security, Core Web Vitals, forced-colors, AI disclosure, and the Cadence typography suite.
What they catch that Designesy does not
Designesy gaps. These are the AI-slop aesthetic detection, structural-variety, signature, and cohesion checks. Designesy does not ask “does this look AI-generated” — it verifies contract conformance.
The moat
Designesy’s uncontested layers
The layers no competitor approaches. These are not features — they are the category-of-one positioning.
Self-score on designesy.org
ToolScoreWhat it caught
designesy93% A (39 PASS / 0 FAIL / 0 WARN / 0 SKIP / 3 MANUAL)0 FAIL, 0 WARN. 3 MANUAL: viewport overflow, sound toggle, Core Web Vitals — browser-only probes the static engine cannot run, so they are excluded from the score rather than counted against it.
hallmarkNot run (skill, not URL API)Would flag: pure #000/#fff (near-pure), zero-chroma neutrals (graphite is chromatic), nav structure
slop-evalNot run (skill, not URL API)Would flag: cool blue-charcoal dark (warm graphite, may pass), mono house voice (minor), focus states (should pass)
hallmark and slop-eval are agent skills, not URL-based APIs. They cannot score a URL without an LLM agent loading the skill and evaluating. Designesy is the only one with a programmatic URL-based scoring endpoint — a structural advantage for CI/CD integration, leaderboards, and automated pipelines.
The wider landscape
The three tools above are the closest comparators. The field is growing — these are newer entrants worth tracking.
design-slop-cop14 anti-slop patterns (rule-based)Lightweight, fast scan, focused pattern set
anti-slop-designDesign quality guardrailsPre-emit lint, not post-hoc scoring
Atlassian ADS MCPDesign-system MCP server (v0.21.1)Benchmarked against design.md — fewer tokens, lower variance
81 design-related MCP servers are now indexed at mcpservers.org (2026-08-01). The category is expanding from skills into agent-invocable tooling — the space designesy pioneered with its 17-tool MCP server and URL-based scoring API.
Provenance
Sources
Live research via AnySearch MCP (2026-08-01). Designesy live score verified via POST to https://www.designesy.org/api/score (93% A,42 checks). hallmark taxonomy from Nutlope/hallmark/skills/hallmark/references/slop-test.md. slop-eval taxonomy from fabricioctelles/skills/skills/slop-eval/references/tells.md.
Competitive benchmark — designesy vs hallmark vs slop-eval. The three tools answer different questions: hallmark prevents slop, slop-eval scores slop, designesy verifies contract conformance. A complete pipeline uses all three. /methodology · /leaderboard · /contracts/design-system