Design & engineering

drift measures the system a site ships

A crawler, a queue and an audit. Every colour, type, spacing, radius and shadow, attributed back to the pages using them.

  • TypeScript
  • Playwright
  • Express
  • BullMQ
  • WebSockets
  • CIEDE2000

What drift is

drift crawls a website and reports the design values it uses: every colour, type size, spacing step, radius, shadow and border, deduplicated, grouped perceptually and attributed to the pages they appear on. A health line and a verdict per category sit on top of that inventory, each measured against a stated reference.

A crawl drives headless Chromium over a site of unknown size, so it is slow, can fail part-way and outlives the request that started it. It runs as a job on a BullMQ queue backed by Redis, in an Express service that reports progress over a WebSocket and publishes its API as an OpenAPI document. The next three sections cover that service, its pipeline and its tests; the analysis comes after.

The crawl, aggregation, verdicts and export are computed, with no API key and no model, so the same pages give the same audit. Judging whether a change between two versions was intended is out of scope.

picocss.com

Design Health

7 of 29 colours are near-duplicates, 6 of 9 type sizes fall off the scale, and 14 of 21 spacing values miss the 4px grid. No issues found in radius, shadows, and contrast.

Colours297 indistinguishable
Contrast21all pairs pass AA
Type96 off-scale
Spacing2114 off a 4px grid
Radius22 values
Shadows22 values
each clause is a count against a stated reference.

Colours less than CIEDE2000 ΔE 2 apart count as near-duplicates, the same colour maths vault and haus use. On this capture that is 29 distinct colours in 6 families, with 7 near-duplicates. Four of them are a colour beside the same colour at 13% alpha (hex 20), which the audit lists as related.

The service contract

The backend is a standalone service. The client is one consumer of the same API a CI job would use. It is plain Express because it needs a durable process to own a persistent WebSocket, the Playwright workers and the BullMQ queue, none of which fit a request-scoped, serverless target.openapi.yaml is an OpenAPI 3.1 document covering every endpoint, and a contract test in the repo checks the app's responses against it. The acceptance suite two sections down tests the running service from outside and generates its TypeScript types from the same document.

Progress arrives over a WebSocket for liveness. /crawl/:id/result is the authoritative completion signal, so a dropped socket degrades to polling and the UI does not hang. An unusable URL is rejected at the edge with a 422; the job is never queued. A crawl that reached no page fails with the worker's reason instead of returning an empty audit. The caller-supplied callbackUrl is treated as an SSRF vector.

// src/queue/webhook.ts: callbackUrl is caller-supplied, so it is an SSRF vector
const resolved = await lookup(host, { all: true })

// a public NAME can still resolve to a private ADDRESS
// (localtest.me → 127.0.0.1), so resolve BEFORE the check
if (!allowlisted && resolved.some(r => isPrivateAddress(r.address, r.family))) {
  throw new Error("callbackUrl must not point at a private or loopback address.")
}

// ...and at delivery, because fetch follows redirects by default:
// a checked host answering 302 sends the request somewhere unchecked
await fetch(url, { method: "POST", redirect: "manual", body, signal })

The host is resolved before the check, because a public name can point at the loopback interface. Private, loopback, link-local, carrier-grade NAT and IPv4-mapped IPv6 addresses are refused. Node's fetch follows redirects by default, so a checked host answering 302 with a private Location would get the request delivered to an address nothing checked. Delivery sets redirect: "manual" and treats a 3xx as a permanent failure.

Delivery is best-effort. A webhook the receiver cannot accept does not fail a crawl that succeeded, because the audit is still served over the API. The worker writes a line to stderr when a callback was not delivered, so a caller waiting on a notification that never arrives is visible in the logs.

The pipeline

Inside the service, a crawl is a queued job running one chain: discover, crawl, extract, normalise, aggregate, audit, export.

the second readThe CSSOM, for the unitsand token names thecomputed pass discarded.discoverFinds the sitemap, orsame-origin links fromthe root if none.crawlDrives headlessChromium, same-origin,to a page cap of ten.extractReads every elementin-page, to a ceiling of12,000 so one heavy pagecannot OOM it.normaliseTurns raw CSS stringsinto typed values,Node-side and pure.aggregateFolds each page intotoken tallies, after thecrawl, not per page.auditProduces the inventory,the contrast findings,a verdict per category.exportWrites one JSON artefactthat leads with thediagnosis, from the client.

Most values come from shared stylesheets, so a few pages find most of a site's values and more pages add attribution. The cap is ten same-origin pages.

A crawl once ran the backend out of memory on a content site, then crashed it again each time BullMQ retried the job. The pipeline keeps every element of every page until the audit runs, so one animation-heavy page was enough. Three changes shipped: a limit of 12,000 elements per page, the page cap cut from forty to ten, and no retries. Folding each page into tallies as it arrives, so memory grows with distinct values, is not built.

Proving the contract

A separate repository, drift-tests, holds acceptance tests for the API: six feature files and eighteen scenarios that call the running service over HTTP and assert on the responses. It imports nothing from drift. Inside drift, 99 unit tests cover the analysis, discovery and crawl functions and 18 contract tests check the app's responses against openapi.yaml, but none of them starts the service with Redis and Playwright behind it. drift-tests does. It follows the BDD regression practice from pendula, applied to my own project.

The suite never touches the public internet. It serves a small, deliberately-inconsistent fixture site on 127.0.0.1, three same-origin pages seeded with near-duplicate blues, off-grid spacing, off-scale type and a failing-contrast pair, and points drift at that. Same input, same audit, every run.

HTTPcrawlsthe suite serves the site drift crawlsdrift-testscucumber-jsdriftthe running backendfixture site127.0.0.1 · 3 pagesHTTPcrawlsthe suite serves the site drift crawlsdrift-testscucumber-jsdriftthe running backendfixture site127.0.0.1 · 3 pages

The fixture is wrong in four measured ways, each chosen to trip one signal: #3366cc beside #3467cc at ΔE 0.3, padding of 13px and 7px off a 4px grid, font sizes of 15, 23 and 31px off the closest modular scale, and #999 on #fff at 2.85:1, which fails AA. A run fails if the audit misses any of the four.

Eleven of the eighteen scenarios assert a failure path: five missing or unusable URLs, an unreachable site, an unknown job id, and four refused callbacks.

drift-tests

6 feature files · 18 scenarios · 11 assert a failure path

The audit finds the seeded faults, and cites only pages the crawl visited.

  • The audit reports its structure and summary
  • The audit surfaces the fixture's seeded inconsistencies
  • Colour and contrast findings cite only pages that were crawled
eleven of the eighteen scenarios assert a failure.

One scenario checks page attribution. Every colour swatch and contrast pair must cite only pages the crawl visited, and no more of them than were crawled. A finding citing five pages after a four-page crawl means a page was counted twice, which no single-page test can show.

# features/audit.feature
Feature: The audit
  A completed crawl yields a deterministic audit of every design token actually
  shipped: colours, type and spacing, with WCAG contrast findings and a summary.
  The fixture site seeds known inconsistencies, so the audit's verdicts are
  predictable run to run.

  Background:
    Given a completed crawl of the fixture site

  Scenario: Colour and contrast findings cite only pages that were crawled
    Then every colour and contrast finding cites only pages that were crawled

The lifecycle scenarios assert that an unreachable target ends failed with a reason and its audit is a 409, rather than a 200 carrying an all-zeros audit. The webhook scenarios assert the SSRF guard refuses a loopback and a private callback while the test backend has 127.0.0.1 allowlisted for its receiver.

# features/webhooks.feature
Feature: Webhook callbacks
  A crawl can POST its finished audit to a callback URL. The URL is validated
  when the crawl is enqueued. A loopback, private or non-HTTP callback is a 422,
  unless its host is in DRIFT_WEBHOOK_ALLOWED_HOSTS.

  # The test backend allowlists the host 127.0.0.1 for webhook-delivery.feature.
  # The allowlist matches the host as written, so localhost, which resolves to
  # loopback, is still refused.
  Scenario: A loopback callback URL that is not allowlisted is refused
    When I enqueue a crawl of the fixture site with callback "http://localhost/hook"
    Then the response status is 422
    And the response carries an error message

  Scenario: A non-allowlisted private callback URL is refused
    When I enqueue a crawl of the fixture site with callback "http://10.0.0.1/hook"
    Then the response status is 422
    And the response carries an error message

CI runs on every push to drift-tests: a Redis service container, drift checked out beside the suite, Playwright's Chromium, the backend started with the webhook variables set, then the eighteen scenarios against it. Nothing is mocked.

A second kind of test runs inside the repo and asserts things about the code itself. One reads every stylesheet and fails if a var(--x) names a property nothing defines. CSS has no error for that: the declaration is dropped and the element inherits, so a missing token reads as a design decision. It found five broken references, including a --duration-default that never existed and had left five animations running with no duration at all. Another compares the client's hand-written copy of the wire types against the service's and fails when a field is optional on one side and required on the other. That had already drifted twice, and the symptom both times was a lint warning nobody could account for.

Neither tests behaviour. Both catch failures that raise no error and leave CI green.

What the audit contains

Under the diagnosis is the evidence: every distinct value the site ships, ranked by how often it is used and attributed to the pages it appears on. Usage is the ranking signal. A value used once is noise; the 16px in this capture is used 100 times against 54 for the next most common size.

Two pages populated 11 of the twelve categories, down to one gradient and two shadow stacks.

Inventory

picocss.com

29 values in 6 families

  • #373c44Neutral170
  • #5d6b89Blue124
  • #5c6370Neutral80
  • #181c25Neutral59
  • #0172adBlue50
  • #ffffffNeutralΔE 0.831
  • #646b79Neutral30
  • #934dafPink30
  • #2e685bCyan30
  • #8b4f00Orange26
  • #982e79PinkΔE 1.110
  • #424751Neutral6
  • #2a628aBlue6
  • #a83410Redopacity variant of #a83410204
  • #a8341020Redopacity variant of #a834104
  • #e7eaf0Neutral3
  • #962d7cPinkopacity variant of #962d7c203
  • #962d7c20Pinkopacity variant of #962d7c3
  • #825400Orangeopacity variant of #825400203
  • #82540020Orangeopacity variant of #8254003
  • #f3f5f7NeutralΔE 1.82
  • #2d3138Neutral2
  • #23728fCyanopacity variant of #23728f202
  • #23728f20Cyanopacity variant of #23728f2
  • #fbfcfcNeutralΔE 0.81
  • #cfd5e2Neutral1
  • #23262cNeutral1
  • #000000Neutral1
  • #48536bBlue1
ranked by usage: 170 uses at the top, one at the bottom.

Categories that are absent do not appear: a site with no gradients gets no gradient section.

Reading what was authored

getComputedStyle returns one resolved pixel number per property. It is accurate for what rendered, and it has already collapsed whatever was authored, rem, em, % or calc(), into that one number. Three things go with it: the unit the author wrote, the site's own token names, and any arithmetic the value was built from. picocss.com writes most of its spacing as calc(var(--pico-spacing) * 2) and the like; read as computed styles alone, each of those is a bare number with nothing left to say which token it came from.

So each page is read twice in the browser. haus-style-probe walks the elements, skipping scripts, styles and hidden nodes, stops after 12,000, and reads each with getComputedStyle. drift's own extract.ts then walks the CSSOM.

// haus-style-probe: runs in the browser, per element, up to 12,000
const cs = window.getComputedStyle(el)   // what RENDERED: resolved px

// drift src/crawler/extract.ts: a walk over the CSSOM for what was AUTHORED
for (let i = 0; i < style.length; i++) {
  const prop = style[i]
  const value = style.getPropertyValue(prop).trim()
  if (prop.startsWith("--"))
    customProperties.push({ name: prop, value })   // the site’s OWN token names
  else if (PROP_CATEGORY[prop])
    declarations.push({ category: PROP_CATEGORY[prop], value })
}

Both reads were also, for a while, blind to modern colour. Pointed at a site authored in OKLCH, drift read every colour as null and reported none: getComputedStyle returns oklch(0.52 0.138 300) verbatim, and both halves of the probe parsed rgb() alone. The walk that resolves an element's effective background treated an unparsed colour as transparent, so it fell through to the page canvas and every contrast pair on such a site was measured against the wrong backdrop. The fix went into haus-colour-utils as a toHex that handles both, which is where the rest of drift's colour maths already lived. The integration suite caught it; every unit test on both sides was written in rgb().

The second read recovers the unit and the token names both. The panel below takes spacing alone and runs both readings over the same 196 declarations: what each API returns, then what each one holds. Not one of those declarations was written in px, and the resolved side has no way to say so.

getComputedStylewhat rendered

One resolved px number per property. All 196 spacing declarations on the two crawled pages arrive here as px, whatever they were written as, and collapse to 21 distinct values.

CSSOMwhat was authored

The same 196 declarations, by the unit actually written. Not one of them was authored in px.

calc 144rem 50em 2
the px reading

21 bare numbers, with nothing left to say which of them came from the same token.

2.5px1645px757.5px1110px21815px1620px4826.4px430px635px2439.6px240px5645px155px360px980px590px6110px1125px2135px1180px6250px4
the CSSOM reading

The 16 most-used authored values. One token, --pico-spacing, and a handful of multipliers over it account for most of the set.

0.25rem22calc(var(--pico-spacing) * 2)18calc(var(--pico-homepage-spacing-vertical)/ 2)14calc(var(--pico-spacing) * .25)10calc(var(--pico-spacing)/ 2)100.125rem80.375rem80.5rem8calc(var(--pico-homepage-spacing-horizontal)/ 2)8calc(var(--pico-spacing) * .5)8calc(var(--pico-spacing) * 1)6calc(var(--pico-spacing) * 4)6calc(var(--pico-spacing)/ 4)6calc(var(--pico-block-spacing-horizontal) * -1)4calc(var(--pico-block-spacing-vertical) * -1)4calc(var(--pico-form-element-spacing-horizontal) + 1.5rem)4
the same 196 declarations, read two ways.

Why the unit is a finding

getComputedStyle cannot tell px from rem, because both arrive as the same resolved number, but they behave differently for a reader who has set a larger default font size: px stays fixed, rem scales. The audit flags font-size authored in px for that reason, which is only possible because the unit was recovered from the CSSOM. picocss.com authors its type in rem, so the flag does not fire here.

The other recovery is naming. getComputedStyle never sees a custom property: by the time it runs, every var() has resolved to a value on some element. A site's real design vocabulary lives only in the stylesheet, so the authored pass reads it off the :root rules. drift can then report a system in its own token names.

Recovered from the CSSOM

170 custom properties, 40 of them aliases onto another declared token. getComputedStyle discards every one. Six of the chains:

  • --pico-accordion-active-summary-color--pico-primary-hover#79c0ff
  • --pico-accordion-border-color--pico-muted-border-color#202632
  • --pico-accordion-close-summary-color--pico-color#c2c7d0
  • --pico-accordion-open-summary-color--pico-muted-color#7b8495
  • --pico-block-spacing-horizontal--pico-spacing1rem
  • --pico-block-spacing-vertical--pico-spacing1rem
the aliases come back with the names.

Measuring against a reference

A size is off-scale only relative to a scale, so the reference is selectable. Type is compared against any named modular ratio; spacing against a 4px or 8px grid. Every option carries its own off-count, so the row answers which scale the system is on before anything is picked.

6 of 9 type sizes fall off this reference · base 16px, the most-used size.

on the reference off itthe Overview verdict stays pinned to the closest fit, so exploring a hypothesis never rewrites the diagnosis
change the reference and the off-count changes with it.

The automatic pick is ranked by fewest values off, with mean relative error as the tiebreak. Ranking by error alone can crown a ratio that fits most sizes tightly but tips a couple over tolerance, which would leave the option labelled closest showing a higher off-count than its neighbour and read as a bug.

// the closest scale is the one the FEWEST sizes miss,
// mean relative error as the tiebreak
for (const r of RATIOS) {
  const scale = buildScaleToCover(basePx, r.ratio, min, max)
  const off = classifyAgainstScale(sizes, scale).filter(m => !m.onScale).length
  const err = meanError(sizes, basePx, r.ratio)
  if (off < bestOff || (off === bestOff && err < bestErr)) best = r
}

The selection drives that ruler and its table, but never the overview verdict, which stays pinned to the automatic best fit. Otherwise a reader who tried the golden ratio out of curiosity would be told their type system is failing.

Design decisions

The colour science is one published dependency. Perceptual near-duplicate clustering and WCAG contrast are real colour maths that would be error-prone to reimplement, so the audit consumes haus-colour-utils from npm. It ships ESM, CommonJS and types, with one dependency, and the backend imports deltaE and clusterByPerceptualDistance from it.

The extractor was lifted out into a package. The per-element measuring code was drift's own, and the same measurements are wanted over a single mounted component rather than a whole crawl. Keeping a second copy would let the two drift apart, so it was published as haus-style-probe with a root option, and drift now consumes it: crawler/types.ts re-exports the package's shapes and drift's own normalise.ts is deleted. The service installs two haus packages, the colour maths and the probe, and the probe started here. The client used haus-tokens and two haus-components from 2026-08-30 and removed both on 2026-09-09, so drift's own interface does not use the design system it would be asked to audit.

BullMQ over pg-boss and an in-memory queue. A crawl is slow and failure-prone and must outlive a restart, so an in-memory queue was out. pg-boss would add Postgres beside Redis. BullMQ provides concurrency control and progress events that map onto the WebSocket frames. Retries are off, because a job that ran the worker out of memory would crash it again on each retry.

Build each piece standalone, one new dependency at a time. drift combines Playwright, Redis, BullMQ and WebSockets, and the failure mode is integrating them together and being unable to tell which layer broke. The rule was to add at most one new infrastructure dependency per step, so a regression points at exactly one piece. Docker is the next step in that order: planned, multi-stage to contain the Chromium binary, and not yet built.

The export leads with the diagnosis. Its audience is a machine: something to assert on, two runs to diff, a model to reason over. Shipping raw counts made the consumer re-derive the judgement drift had already made, so the export leads with the health line, the typed findings and the verdicts, then a rules block stating the ΔE threshold, grid base, detected ratio and WCAG standard. The full inventory sits underneath as evidence.

The export is assembled in the client and produced by the Export button. No endpoint serves it, so a CI job gets the audit and not the diagnosis, and the acceptance suite cannot test the export. Serving it from the API was issue #3, closed while the public deployment only replays a capture. The health line and verdict helpers are separate tested functions; the findings are assembled inside the audit screen component and would have to move with them.

The proposals layer was cut. A second layer once projected the audited tokens onto known-good structures: a consolidated palette, a modular type scale, a spacing grid. An audit is a claim drift can defend from the evidence it collected. A proposal is a recommendation, and the only warrant drift had for one was that the result came out arithmetically tidier. Three parts survived the cut because they measure: the selectable references, the perceptual clustering, and the export.

Where it stands

Four things are not built. The page cap of ten is doing work that belongs to the pipeline: until each page is folded into tallies as it arrives, memory scales with elements times pages, and raising the cap means doing that refactor first. There is no Dockerfile yet, so deployment is a set of instructions. The crawl reads a single viewport, which makes the whole responsive system invisible to the audit, and it reads the resting state only, so hover and focus styles are never seen.

The service has 149 tests in 15 files: 99 over the analysis, discovery and crawl functions, 18 checking responses against the published schema, and 32 over the app, the queue and the realtime layer. The client, which once had no tests, has 231 in 12 files over the screens, the flow, the audit model and the token layer; the largest, 55 tests, covers the model behind the audit screen. drift-tests adds eighteen scenarios over the running service, in six feature files and 58 steps.

The deployed site does not crawl on demand. A Playwright crawler behind a Redis queue is not safe to leave open to anonymous callers, so the public build replays an audit captured from a real crawl, and the configure screen says so. The inventory, verdicts and export are computed from that capture; only the API calls are stubbed.

The capture is picocss.com over two pages, in client/src/demo/audit.json. Every figure on this page is generated from that file, so recapturing the site regenerates the page. Running npm run capture in drift rewrites the artefact and prints the diagnosis figures the documentation quotes.

One issue is open: splitting the audit stylesheet, which went from 1,542 lines to 694 and is parked. Authentication and rate limiting, the export on the API, and a timestamp and delivery id on webhook signatures were closed as not planned while the deployment only replays a capture; each is needed before the crawler runs on a public URL. The two endpoints with no consumer were removed, and the type-tier debt is 42 reads, 33 of them font-family.

The decision log, including the layer that was cut, in DESIGN.md →