Annual report · Edition 1
State of Design Compliance
The first deterministic report on design-system contract compliance across the web. 30 sites scored against a 40-check engine. No surveys. No votes. No self-reported data. Every score is computed from the live CSS the site ships at :root.
compliance_index_version: 1.0contract v0.4.0last scored 2026-08-03
The trust contract
Independence firewall
Designesy does not accept payment for scores, methodology changes, or leaderboard placement. Every score is computed by the same deterministic 40-check engine against the same published contract. Enterprise customers pay for private scoring, custom contracts, and CI integration — never for public leaderboard placement. If a scored site is also an enterprise customer, their public score is computed identically to any non-customer’s score. No pre-release optimization. No score suppression. The engine is open: run it yourself with npx designesy-score or the MCP server.
This is the structural separation Artificial Analysis and Arena pioneered: the public trust asset is non-monetizable; the consulting and private-scoring layer around it is. The difference is that Designesy’s data is deterministic — there is no vote to manipulate, no survey to game, no subjective judge to influence.
The cohort
30 sites across five tiers — frontier references, competitors, design-system exemplars, inspiration, and high-traffic surfaces. Scored against contract v0.4.0 with the 40-check engine. Re-scored weekly.
Grade distribution
One A-grade site in a cohort of 30. The contract is demanding — most sites land in D or F because they don’t ship the primitives (token systems, reduced-motion blocks, font-synthesis rules) at :root. That is the point: the gap between what a site documents and what it ships is exactly what this report surfaces.
Where the cohort struggles
Average category scores across the scored cohort. The lowest-scoring categories reveal which contract primitives the industry has not yet adopted at the shipped-surface level.
The flagship finding
Material Design 3 — Google’s design system, the most influential on Earth — has no public conformance, verification, or certification tool. None. m3.material.io provides guidelines but no automated conformance checker. We scored it.
M3’s guidelines specify accessibility, motion, and token architecture in prose. The contract verifies whether those guidelines are actually shipped on the live surface. The gap between documented and shipped is the finding.
Best category: identity at 75.0%. Worst scored category: interaction at 0.0%. M3’s token architecture (DSP) was archived October 2024 and does not emit W3C DTCG format. M3 Expressive ships spring-based motion with no published reduced-motion token. The contract catches what the guidelines leave as prose.
This is the demonstration. The system that wrote the guidelines doesn’t have a tool to verify its own output — we do. The same engine that scores M3 scores designesy.org (self-score: 100.0% / A), in public, with the same 40 checks. Transparency earns trust.
Framework rankings
Design-system frameworks and documentation platforms in the cohort, ranked by compliance score. Each framework is a potential case study — each score is a piece of content.
| Framework | Grade | Score | Checks |
|---|---|---|---|
| Designesyself designesy.org | A | 100.0 % | 36p · 0f · 0w |
| GitHub Primer primer.style | C | 77.6 % | 26p · 5f · 5w |
| zeroheight zeroheight.com | C | 73.8 % | 16p · 4f · 15w |
| Atlassian Design System atlassian.design | C | 70.2 % | 14p · 3f · 19w |
| Designesy AI Studio designesy.ai.studio | D | 67.8 % | 10p · 1f · 23w |
| Desy Guard getdesy.com | D | 66.0 % | 16p · 2f · 17w |
| Radix Colors radix-ui.com | D | 60.7 % | 15p · 5f · 15w |
| Roast by AI roastbyai.com | F | 59.3 % | 16p · 3f · 16w |
| Material 3 m3.material.io | F | 59.0 % | 6p · 4f · 22w |
| Mozaika mozaika.design | F | 52.6 % | 6p · 3f · 23w |
| Google Stitch stitch.withgoogle.com | F | 52.5 % | 7p · 3f · 23w |
| IBM Carbon carbondesignsystem.com | F | 50.8 % | 14p · 6f · 15w |
| Vercel Geist geist.dev | F | 50.3 % | 7p · 3f · 21w |
| Adobe Spectrum spectrum.adobe.com | F | 49.4 % | 7p · 4f · 21w |
The spread is 51 points between the highest and lowest design-system framework. Primer leads the non-self cohort at 77.6/C. Material 3 — the most influential system on Earth — scores 59/F. The contract does not grade on a curve.
Methodology
Every score in this report is computed by the same deterministic 40-check engine — no LLM, no human judgment, no subjective vote. Each check is a regex, token-resolution, or spec-linter test against the live fetched CSS and HTML. The engine extracts CSS from the URL, parses :root custom properties, and runs 40 checks across 14 weighted categories.
The score is a weighted average of PASS/WARN/FAIL results, with an accessibility floor: if the accessibility category scores below 60%, the overall grade is capped at C. Twelve anti-slop rules subtract up to 20 points. Seven originality signals add up to 8 points. Taste is part of the number.
The full methodology — every check, its category weight, the scoring math, the grade bands, and what the engine cannot measure — is documented in full on the methodology page. The engine is open: score any URL at /score, run it locally with npx designesy-score, or integrate it in CI with the GitHub Action.
The cadence
24-hour SLA
New framework releases scored within 24 hours
When Radix, shadcn/ui, Mantine, Park UI, or Ark UI ship a new version, Designesy re-scores their default theme against the contract within 24 hours and publishes the result.
Weekly re-score
Every site re-scored weekly with delta badges
The leaderboard is re-scored every week via a GitHub Action (Mondays 10:00 UTC). Each site shows a delta badge — up, down, or flat — since the previous week’s score.
Annual report
State of Design Compliance published yearly with YoY trends
Each annual edition adds year-over-year trend tables: which categories improved, which frameworks moved, which primitives the industry adopted. Edition 1 establishes the baseline.
This is the content engine. Each scored site is a data point. Each framework release is a scoring event. Each annual report is a link magnet. The longer the leaderboard runs, the more unreplicable the dataset becomes — competitors can build a verification engine; they cannot replicate years of accumulated scores and trust.
The expanding surface
Scoring is the foundation. But verification is bigger than a single number. Edition 1 ships with five tools that expand what Designesy verifies — from a score to a maturity profile, a per-framework evaluation, a format bridge, and frontier physics validation.
Maturity
Compliance maturity self-assessment
24 questions across 6 axes — token discipline, motion consistency, accessibility readiness, platform fit, identity & copy, verification maturity. Shareable radar chart. The self-diagnostic that complements the deterministic score.
Evaluations
Per-framework evaluation pages
Every scored site gets a dedicated page with a score dial, per-category breakdown, auto-generated narrative, peer comparison, and cohort context. 30 evaluations — each one a piece of content.
Changelog
Contract changelog by dimension
16 entries across 11 design dimensions — tokens, motion, cadence, accessibility, takt, poise, acoustics, copywriting, identity, security, verification. Every contract change is traceable to a version bump and a dimension.
M3 Bridge
M3’s DSP export was archived October 2024 and doesn’t emit W3C DTCG. This tool converts M3 token CSS or JSON to DTCG 2025.10 format with validation. The neutral bridge between Google’s two non-interoperating design-data initiatives.
Frontier
Spring physics accessibility validator
No one validates spring-based motion against a reduced-motion contract. M3 Expressive ships springs with no published reduced-motion token. This tool simulates the physics, computes overshoot, and renders an accessibility verdict. Green-field.
The pattern: score → evaluate → bridge → validate. Each tool deepens the verification surface. The maturity assessment turns a score into a diagnostic. The evaluations turn a score into a per-framework article. The M3 bridge turns a format gap into a tool. The spring validator turns a frontier into a check. The scoring engine remains the foundation — but verification is now a platform, not a single number.