awwwards-mcp 1.5.0 → 1.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.dev.md ADDED
@@ -0,0 +1,177 @@
1
+ # awwwards-mcp — project & architecture guide
2
+
3
+ Developer-facing guide to this repository. The [main README](README.md) covers
4
+ installation and user-facing features; this file explains how the codebase is
5
+ structured, how the data flows, and how to work on it without breaking it.
6
+
7
+ ## What this is
8
+
9
+ A free, open-source MCP server that gives AI agents design inspiration from
10
+ [Awwwards](https://www.awwwards.com/). No API key, no account: it politely
11
+ scrapes public pages (1 request/second, robots.txt-compliant paths), caches
12
+ them in local SQLite, and serves them through 7 MCP tools — screenshots,
13
+ design DNA, band maps and motion recordings included.
14
+
15
+ ```
16
+ ┌─────────────────────────────────────────────┐
17
+ awwwards.com ─────►│ src/awwwards.ts HTTP client (1 req/s, │
18
+ │ BlockedError on 403/429) │
19
+ └───────────────┬─────────────────────────────┘
20
+ │ HTML
21
+ ┌───────────────▼───────────────┐
22
+ │ src/parsers.ts regex parsers │
23
+ │ parseListing / parseDetail / │
24
+ │ parseJuryDimensions / │
25
+ │ parseCategories / parseElements
26
+ └───────────────┬───────────────┘
27
+ │ rows
28
+ ┌───────────────────────────▼───────────────────────────┐
29
+ │ src/cache.ts SQLite (node:sqlite) at ~/.awwwards-mcp │
30
+ │ sites table + FTS5 (probe-guarded, derived) + asset │
31
+ │ image cache (sha1 of CDN path, no TTL) │
32
+ └───────────────────────────┬───────────────────────────┘
33
+ │
34
+ ┌───────────────────────────▼───────────────────────────┐
35
+ │ src/server.ts 7 MCP tools over createHandlers(deps) │
36
+ └───────────────────────────┬───────────────────────────┘
37
+ │
38
+ src/cli.ts ── stdio McpServer (awwwards-mcp)
39
+ src/index-cli.ts ── crawl CLI (awwwards-index)
40
+ ```
41
+
42
+ ## Module map (`src/`)
43
+
44
+ | Module | Lines | Role |
45
+ |---|---|---|
46
+ | `awwwards.ts` | ~134 | HTTP client. `BASE_URL` / `ASSETS_URL`, `AWARD_FILTERS`, `buildFilterUrl`, `RateLimiter` (promise-chained), `AwwwardsClient` with one retry on transient failures — **never** retries 403/429 (throws `BlockedError`). All binary CDN assets flow through one `getAsset` path. |
47
+ | `parsers.ts` | ~250 | Dependency-free regex parsers. See [parsers & drift](#parsers--drift-the-safety-net) below. |
48
+ | `cache.ts` | ~262 | SQLite storage. FTS5 layer is **derived and probe-guarded**: if the Node build lacks FTS5, `searchSites` returns null and callers fall back to substring search. Per-operation DB open/close (an open handle makes cache files undeletable on Windows). Batched-transaction upserts (per-insert autocommit pays ~4 ms disk sync on Windows). |
49
+ | `indexer.ts` | ~153 | `runIndexer`: meta-keyed lock (`index:lock`, 30 min stale), incremental crawl via `index:progress`, abort-safe status writes. `isIndexStale` gates the 7-day background re-index. |
50
+ | `server.ts` | ~700 | All MCP tool handlers + search merge semantics. The biggest and most rule-dense module — see [search semantics](#search_sites-semantics) before touching it. |
51
+ | `motion.ts` | ~429 | `recordSiteMotion`: webm recording + ffmpeg filmstrip under a strict 30 s budget (documented timing table at the top of the file). Virtual SVG cursor, per-call `mkdtemp` isolation, 10-min stale-tmp sweep, 60 s ffmpeg deadline. |
52
+ | `structure.ts` | ~249 | Band map. `SCAN_SNIPPET` runs in-page via `page.evaluate` (self-contained, reaches DOM through `globalThis` — the Node tsconfig has no DOM lib). `collapseBands` is pure: 8 px sweep, smallest-covering candidate wins, sliver absorb, max-40 cap. Owns the ONE `preScroll` shared with capture. |
53
+ | `capture.ts` | ~66 | Full-page screenshot. Intentionally circular import with `structure.ts` (call-time-only bindings; ESM resolves it). |
54
+ | `cli.ts` | ~179 | `McpServer` wiring, zod schemas, lazy playwright/ffmpeg defaults (optional deps are only touched when a tool actually runs), background auto-reindex, daily version check — stdout stays clean for JSON-RPC. |
55
+ | `version-check.ts` | ~131 | Daily npm-registry check, stderr-only, 3 s timeout, all failures swallowed. Self-update is opt-in via `AWWWARDS_AUTO_UPDATE=1`. |
56
+ | `viewport.ts` | ~20 | Desktop (1440×900) / mobile (390×844, DPR 3, touch) profiles. Width/height go in playwright's `viewport` key; mobile flags spread as sibling context options. |
57
+ | `types.ts`, `index-cli.ts` | — | Shared types (`SiteSummary`, `SiteDetails`, `ElementMedia`); thin crawl-CLI entry. |
58
+
59
+ ## Search semantics (read before touching `server.ts`)
60
+
61
+ `search_sites` invariants, each pinned by tests in `test/server.test.ts`:
62
+
63
+ - **One filter per URL.** Combined filter URLs return 404 on awwwards.com, so
64
+ `buildFilterUrl` picks the highest-priority filter (color > award >
65
+ technology > first tag) and the rest are verified client-side by
66
+ `matchesFilters`. `honorUrlSource` skips re-checking the URL filter for
67
+ freshly scraped rows only.
68
+ - **FTS skip-query-check.** FTS matches prefix+porter-stem; re-checking the
69
+ query with substring semantics would wrongly drop stem matches
70
+ ("magazines" → Magazine). FTS-sourced rows skip the query check.
71
+ - **Color is never served from cache.** Site rows carry no colors, so a color
72
+ search always scrapes its filter page.
73
+ - **Page-window merge.** An empty requested window → scrape replaces. A
74
+ *partially* filled window → top-up scrape MERGES: dedupe by slug, verified
75
+ cache rows win duplicates, bm25 rank order is preserved, scraped-only rows
76
+ append newest-first.
77
+ - **Top-up scrapes are best-effort.** A failed fetch must not escape to the
78
+ stale-fallback catch, which would silently discard the cached page window.
79
+ - **Never cache an all-empty parse.** `get_site_details` /
80
+ `get_site_elements` treat an all-empty detail as layout drift and leave it
81
+ uncached so the mismatch keeps surfacing. One fetch feeds both caches.
82
+ - **Zero results get hints, not live fetches.** Loose-match OR hint (up to 3
83
+ slugs) + taxonomy suggestions computed from cached categories only.
84
+ - **Stale fallback.** On live-request failure, serve stale cache — and the
85
+ cache lookup itself is guarded (the store may be the failure source).
86
+
87
+ ## Parsers & drift: the safety net
88
+
89
+ awwwards.com is a moving target, so the parsers follow strict conventions:
90
+
91
+ - Every regex / `split` / `indexOf` anchor carries a date-stamped
92
+ verification note ("verified live 2026-09-XX").
93
+ - **`parseJuryDimensions` refuses to guess**: if the chartbar labels are not
94
+ exactly Design/Usability/Creativity/Content, it returns undefined instead
95
+ of mapping by position.
96
+ - `parseElements`: `null` = no Elements section (legitimate); empty array =
97
+ markup drift (never cached).
98
+ - Description: curated `>Description</h2>` section primary (~16 % coverage),
99
+ `og:description` meta fallback (~50 %).
100
+ - Elements media paths come in two schemes — modern `element/YYYY/MM/…` and
101
+ 2018-era `external/YYYY/MM/…` — both with `_static.jpeg` posters on the CDN.
102
+
103
+ Three layers guard these anchors:
104
+
105
+ 1. **Fixtures** — real saved awwwards pages in `test/fixtures/` (listing +
106
+ four detail variants, 300–600 KB each). All parser tests are offline.
107
+ 2. **Drift probe** — `scripts/parser-drift-probe.mjs` (`npm run drift`)
108
+ checks every parser anchor against the live listing and detail pages
109
+ (2 fetches at 1 req/s) and against the committed fixtures
110
+ (`npm run drift -- --fixture`, offline). Writes `.drift/status.json`.
111
+ 3. **CI** — a daily GitHub Action (`.github/workflows/parser-drift.yml`)
112
+ runs the live probe and opens/updates a single tracking issue on drift
113
+ (auto-closes when green).
114
+
115
+ When a drift issue lands, the fresh HTML snapshot becomes the new fixture.
116
+
117
+ ## Tests
118
+
119
+ ```bash
120
+ npm test # offline, vitest, fixtures only — no network
121
+ npm run smoke # manual live smoke test (test/live-smoke.ts)
122
+ npm run build # tsc → dist/
123
+ ```
124
+
125
+ Suites: `parsers` (fixtures pin every anchor), `cache` (FTS on/off, TTLs),
126
+ `server` (the merge semantics above — the largest suite), `indexer` (lock,
127
+ resume, abort), `motion` / `capture` / `structure` (fake playwright +
128
+ injectable ffmpeg — nothing launches a browser), `awwwards` (fetch fake),
129
+ `version-check`. `vitest.config.ts` raises the timeout to 30 s for the
130
+ multi-page indexer crawls.
131
+
132
+ Dependency injection is the pattern everywhere a hard dependency exists:
133
+ `AwwwardsClient.fetchFn`, `Cache` paths + clock, playwright loaders,
134
+ `ffmpegFn`. Write new tests against fakes, not networks.
135
+
136
+ ## Scripts & tooling
137
+
138
+ | Script | Purpose |
139
+ |---|---|
140
+ | `scripts/doctor.mjs` | `npm run doctor` — checks network, drift, deps, cache, boot; applies fixes. |
141
+ | `scripts/parser-drift-probe.mjs` | `npm run drift` — live + fixture anchor probe (see above). |
142
+ | `scripts/enrich-styles.mjs` | Phase-1 style-classification experiment via a self-hosted simple-jev classifier (spec: `docs/superpowers/specs/2026-09-20-jev-style-enrichment-design.md`). |
143
+ | `scripts/make-demo.mjs` | Builds the README demo assets. |
144
+ | `scripts/record-scrollthrough.mjs` | The validated recording script `motion.ts` was ported from. |
145
+ | `scripts/skill-memory.mjs` | Journal → techniques distillation for the skills flywheel (`skills/_memory/techniques.json`). |
146
+
147
+ ## Skills & showcase
148
+
149
+ - `skills/` ships three agent skills (`awwwards-inspiration`,
150
+ `awwwards-motion-study`, `awwwards-doctor`). Shipped copies change only via
151
+ human PRs; local copies self-improve through the skill-memory flywheel.
152
+ - `showcase/` holds real sites built through the loop (`ridge`,
153
+ `afjal-portfolio`, `fallow-press`), each with `_qa/` capture evidence. They
154
+ double as regression demos for the capture/structure/motion tools.
155
+
156
+ ## Conventions worth preserving
157
+
158
+ These are hard-won; check for them before "simplifying":
159
+
160
+ - `page.evaluate` snippets are self-contained and reach browser globals via
161
+ `globalThis` (no DOM lib in tsconfig).
162
+ - Per-operation SQLite open/close; batched transaction upserts.
163
+ - Optional heavy deps (playwright, ffmpeg-static) stay unresolved at compile
164
+ time and resolve lazily at runtime with in-band install hints.
165
+ - stdout is the MCP JSON-RPC channel — every log/notice goes to stderr.
166
+ - New crawler code must keep the politeness contract: `/websites/…` paths
167
+ only, 1 req/s, no query-string pagination (`/tag/`, `/search-websites`,
168
+ `/elements/*` etc. are robots-disallowed and never fetched).
169
+
170
+ ## Contributing
171
+
172
+ PRs welcome — parser-drift fixes especially (a fresh HTML snapshot attached
173
+ to a drift issue is the fastest merged PR). See [CONTRIBUTING.md](CONTRIBUTING.md)
174
+ for branch naming, the PR template and the politeness constraints. Security
175
+ issues go privately via [SECURITY.md](SECURITY.md).
176
+
177
+ MIT — see [LICENSE](LICENSE).
package/README.md CHANGED
@@ -4,31 +4,101 @@ Free, open-source MCP server that gives AI agents design inspiration from
4
4
  [Awwwards](https://www.awwwards.com/) — the Mobbin-style visual reference loop,
5
5
  sourced from the web's best award-winning websites.
6
6
 
7
+ <a href="https://github.com/INSANE0777/Awwwards-mcp"><img src="assets/demo.gif" alt="awwwards-mcp in action: search results, design DNA, motion filmstrip, band map — real tool output" width="480"></a>
8
+
7
9
  Your agent searches in natural language ("dark 3D portfolio sites", "soft pastel
8
10
  e-commerce"), sees **real screenshots inline**, and can pull the **design DNA**
9
11
  of any site: color palette, tech stack, design elements, award history.
12
+ Free-text queries run on a porter-stemmed, prefix-matching FTS5 index with BM25
13
+ ranking — "magazines" now finds Magazine-tagged sites (68 on the live index),
14
+ best matches first, where the old substring path returned zero. Multi-word
15
+ queries keep AND semantics: every token must hit the same site.
10
16
 
11
17
  ## Tools
12
18
 
13
19
  | Tool | What it does |
14
20
  |------|--------------|
15
- | `search_sites` | Search by color, tags, technology or award type. Returns site cards with inline screenshots. |
21
+ | `search_sites` | Search by color, tags, technology, award type or free-text query. Multi-word queries match against the local FTS5 index and rank BM25 (title hits lead); zero results come with loose-match and taxonomy-tag hints. Default full results include each site's screenshot; use `responseMode: "compact"` for concise cards and inline previews of only the first two results on the requested page. |
16
22
  | `get_site_details` | Full design DNA for one site: palette, technologies, elements, awards, description. |
23
+ | `compare_sites` | Compare 2–3 sites' design DNA and jury scores as text-only JSON. Uses cached details or fetches missing detail pages. |
24
+ | `get_index_status` | Read local index count, crawl progress, last success/error, lock state and freshness without network requests. |
17
25
  | `get_site_elements` | Component-level visuals for one site: each element's poster image inline (3D models, video content, mobile layouts, microcopy…) + video URLs. |
18
26
  | `list_categories` | Every filter the agent can search by (200+ tags, 27 colors). |
19
- | `capture_live_site` | Optional: fresh full-page screenshot of any live URL. Waits for `load` + a settle window with a bounded pre-scroll, so heavy sites work (`waitStrategy: "networkidle"` available). (needs [playwright](https://playwright.dev)). |
20
- | `analyze_page_structure` | Section band map of any page (live URL or local file:// build): tag, background, offset, height per band. Compare a reference site's structure against your build. Same heavy-site-friendly wait (`waitStrategy: "networkidle"` available). (needs [playwright](https://playwright.dev)). |
21
- | `record_site_motion` | Optional: short motion-through video of a live URL — preloader, scroll-triggered and hover/cursor animations. Returns an inline filmstrip JPEG plus the saved .webm path. (needs [playwright](https://playwright.dev) + ffmpeg-static). |
27
+ | `capture_live_site` | Optional: fresh full-page screenshot of any live URL. Waits for `load` + a settle window with a bounded pre-scroll, so heavy sites work (`waitStrategy: "networkidle"` available). Pass `viewport: "mobile"` for the 390×844 iPhone-class render (`"desktop"` 1440×900 default). (needs [playwright](https://playwright.dev)). |
28
+ | `analyze_page_structure` | Section band map of any page (live URL or local file:// build): tag, background, offset, height per band. Compare a reference site's structure against your build. Same heavy-site-friendly wait (`waitStrategy: "networkidle"` available); `viewport: "mobile"` analyzes the phone-class layout (`"desktop"` default). (needs [playwright](https://playwright.dev)). |
29
+ | `record_site_motion` | Optional: short motion-through video of a live URL — preloader, scroll-triggered and hover/cursor animations. Returns an inline filmstrip JPEG plus the saved .webm path. `viewport: "mobile"` records at phone size — the filmstrip renders at the selected viewport, no pillarboxing (`"desktop"` default). (needs [playwright](https://playwright.dev) + ffmpeg-static). |
30
+ | `search_elements` | Search the inspiration-elements gallery (footer, hero, pricing, 404…) by free text; ranks BM25 over title/author/category. |
31
+ | `get_element` | One element record: title, category, author, built-with stack, related elements and its media URL (image or video) pointing at awwwards' CDN. |
32
+ | `get_motion_dna` | Runtime motion fingerprint of a live URL: animation libraries, render engines, ScrollTrigger stats (trigger count, scrub ratio), tween easing/duration vocab and the scroll model. Fresh capture or cached capture with timestamp. |
33
+ | `search_motion` | Search previously captured motion-DNA scans by library, scroll model or easing vocabulary — find references by how a site moves. |
34
+ | `new_winners` | Poll today's freshly-crowned winners (SOTD / Developer Award / Honorable Mention) against a persisted baseline. First call seeds and dumps the listing; later calls report the delta. Each first-seen winner's Elements section is backfilled into the searchable element corpus, so new winners are element-searchable immediately. |
35
+ | `watch_site` | Persistent watches over studios, tags or specific sites (add/list/remove). `list` matches each watch against the freshest cached listing and reports per-watch `NEW since last check` deltas — it never fetches; `new_winners`/`search_sites` keep the pool fresh. |
36
+
37
+ **Data posture**: element records store metadata + media URLs pointing at
38
+ awwwards' own CDN — nothing is mirrored. Motion DNA records are local
39
+ captures, each stamped with the time it was taken.
22
40
 
23
41
  ## Setup
24
42
 
43
+ **v1.0.0** — the first stable release. Any MCP-compatible coding agent can use awwwards-mcp — no API key, no account.
44
+ Requires Node ≥ 22.13 (`node -v` to check). Pick your agent:
45
+
46
+ **Updates**: the server checks the npm registry once a day and prints an
47
+ stderr notice when a newer `awwwards-mcp` exists (stdout stays clean for the
48
+ JSON-RPC channel — your agent sees the notice as a log line). Set
49
+ `AWWWARDS_AUTO_UPDATE=1` in the server's `env` to opt into background
50
+ self-update; restart your agent afterwards to load it. Nothing is fetched
51
+ more than once a day and serving never waits on the check.
52
+
25
53
  **Claude Code**
26
54
 
27
55
  ```bash
28
56
  claude mcp add awwwards -- npx -y awwwards-mcp
29
57
  ```
30
58
 
31
- **Claude Desktop / Cursor / Windsurf** (`mcpServers` in the config):
59
+ **Codex CLI** (ChatGPT desktop app and the IDE extension share this config)
60
+
61
+ ```bash
62
+ codex mcp add awwwards -- npx -y awwwards-mcp
63
+ ```
64
+
65
+ or in `~/.codex/config.toml` (project-scoped: `.codex/config.toml`):
66
+
67
+ ```toml
68
+ [mcp_servers.awwwards]
69
+ command = "npx"
70
+ args = ["-y", "awwwards-mcp"]
71
+ ```
72
+
73
+ **OpenCode** (`opencode.json` — note the command is an array)
74
+
75
+ ```json
76
+ {
77
+ "$schema": "https://opencode.ai/config.json",
78
+ "mcp": {
79
+ "awwwards": {
80
+ "type": "local",
81
+ "command": ["npx", "-y", "awwwards-mcp"]
82
+ }
83
+ }
84
+ }
85
+ ```
86
+
87
+ **ZCode** (`~/.zcode/cli/config.json` — note servers nest under `"mcp": { "servers": ... }`)
88
+
89
+ ```json
90
+ {
91
+ "mcp": {
92
+ "servers": {
93
+ "awwwards": { "command": "npx", "args": ["-y", "awwwards-mcp"], "env": {} }
94
+ }
95
+ }
96
+ }
97
+ ```
98
+
99
+ **Claude Desktop / Cursor / Windsurf / Gemini CLI / Cline / Continue** — anything
100
+ reading the common `mcpServers` JSON shape (e.g. `~/.claude/claude_desktop_config.json`
101
+ or `~/.gemini/settings.json`):
32
102
 
33
103
  ```json
34
104
  {
@@ -38,25 +108,59 @@ claude mcp add awwwards -- npx -y awwwards-mcp
38
108
  }
39
109
  ```
40
110
 
41
- Optional full-page captures:
111
+ **Anything else** — awwwards-mcp is a plain stdio MCP server: point your client
112
+ at `npx -y awwwards-mcp` and it works. To pin a version, use
113
+ `npx -y awwwards-mcp@1.0.0`.
114
+
115
+ **pi coding agent** has no built-in MCP by design — it uses skills and
116
+ extensions instead. Two options:
117
+
118
+ 1. Install the awwwards-inspiration skill (below). pi reads skills from
119
+ `~/.pi/agent/skills/` or `~/.agents/skills/` (the latter is shared across
120
+ agents following the Agent Skills standard). The skill teaches the workflow;
121
+ for it to reach the live data, add an MCP-supporting pi extension, or run
122
+ the queries in another agent and paste results.
123
+ 2. Skip MCP entirely: ask pi to build you a small CLI wrapper around
124
+ awwwards.com, or use a shared skills directory (`~/.agents/skills/`) so the
125
+ same skill file serves pi and every other agent.
126
+
127
+ Optional full-page captures (needed by `capture_live_site`,
128
+ `analyze_page_structure`, `record_site_motion`):
42
129
 
43
130
  ```bash
44
131
  npm install -g playwright && npx playwright install chromium
45
132
  ```
46
133
 
134
+ `record_site_motion` additionally uses ffmpeg; it resolves the `ffmpeg-static`
135
+ package automatically if present.
136
+
47
137
  ## Skills
48
138
 
49
- This package ships an agent skill that teaches the inspiration workflow —
50
- search, judge from screenshots, pull design DNA, state a design direction —
51
- using the awwwards MCP tools. Copy it into your agent's skills directory:
139
+ This package ships four agent skills. Any agent that follows the
140
+ [Agent Skills standard](https://agentskills.io) can load them; copy them into
141
+ your agent's skills directory:
52
142
 
53
143
  ```bash
54
144
  npm install awwwards-mcp
55
- mkdir -p ~/.claude/skills && cp -r node_modules/awwwards-mcp/skills/awwwards-inspiration ~/.claude/skills/
145
+ mkdir -p ~/.agents/skills && cp -r node_modules/awwwards-mcp/skills/awwwards-inspiration node_modules/awwwards-mcp/skills/awwwards-setup node_modules/awwwards-mcp/skills/awwwards-doctor node_modules/awwwards-mcp/skills/awwwards-motion-study ~/.agents/skills/
56
146
  ```
57
147
 
58
- For ZCode, copy to `~/.zcode/skills/` instead of `~/.claude/skills/`.
59
- Windows: run this from Git Bash, or copy `node_modules\awwwards-mcp\skills\awwwards-inspiration` manually.
148
+ | Skill | What it teaches |
149
+ |-------|-----------------|
150
+ | `awwwards-setup` | First-time onboarding: asks the user's preferences (result density, viewport, captures, local index, winner watches), persists them to `~/.awwwards-mcp/preferences.json`, and runs any one-time installs they opt into. |
151
+ | `awwwards-inspiration` | The inspiration loop: search, judge from screenshots, pull design DNA, state a design direction, capture/motion-first builds; staying current with `new_winners` and `watch_site`. |
152
+ | `awwwards-motion-study` | The full video chain: what to record from a live site (and what to skip), frame-by-frame review (video input or tile-per-element), the motion inventory, and build verification by re-recording. |
153
+ | `awwwards-doctor` | Repair: run `npm run doctor`, apply its fixes, re-anchor parsers after real awwwards.com drift, recover the in-flight task that surfaced the failure. |
154
+
155
+ | Agent | Skills directory |
156
+ |-------|------------------|
157
+ | Claude Code | `~/.claude/skills/` |
158
+ | pi | `~/.pi/agent/skills/` (also reads `~/.agents/skills/`) |
159
+ | ZCode | `~/.zcode/skills/` |
160
+ | Agent Skills-standard agents | `~/.agents/skills/` |
161
+
162
+ Windows: run this from Git Bash, or copy
163
+ `node_modules\awwwards-mcp\skills\awwwards-setup` manually.
60
164
 
61
165
  ## Indexing (recommended)
62
166
 
@@ -74,10 +178,43 @@ npm run index
74
178
  SQLite cache at `~/.awwwards-mcp/`.
75
179
  - Resumable: interrupt it and re-run — completed pages are skipped.
76
180
  - The MCP server re-indexes automatically in the background whenever the
77
- index is older than 7 days (never blocking your session).
181
+ index is older than 7 days (never blocking your session). Completed crawl
182
+ checkpoints are cleared so each refresh actually revisits the tag pages.
183
+ Run `get_index_status` to inspect progress or the last crawl error.
184
+
185
+ ### Elements index (optional)
186
+
187
+ The `search_elements` tool auto-indexes the first gallery page (~48 items)
188
+ on first use. To build a full corpus (~1,500+ items, ~15 pages at 48/page):
189
+
190
+ ```bash
191
+ npm run index -- --elements all # follow pagination until exhausted
192
+ npm run index -- --elements 10 # first 10 listing pages
193
+ ```
194
+
195
+ Any crawl beyond the first page also runs the **taxonomy pass**: every
196
+ facet page (`/elements/footer/`, `/elements/cta/`, … 46 categories) is
197
+ fetched once and each element under it is stamped with that category, so
198
+ `search_elements`' `category` filter works across the corpus. Element
199
+ pages carry no breadcrumb, so this listing-side pass is the only category
200
+ source; elements seen on no facet page stay `unsorted`.
201
+
202
+ Element rows are searched by title, author, category **and slug tokens**
203
+ (slug is an FTS5-indexed column; a cache opened from an older schema
204
+ version rebuilds its search index automatically on first open). Each
205
+ element also carries the slug of the award-winning site it came from
206
+ (`siteSlug`/`siteUrl` in `search_elements`/`get_element` results). Two
207
+ record sources share this corpus: gallery records (indexed from the public
208
+ elements listing) and `source:"site"` records backfilled from each new
209
+ SOTD winner's own Elements section by `new_winners` — their slugs are
210
+ namespaced `site-<siteslug>-<title>` so the two never collide. The
211
+ elements index has the same 7-day freshness gate as the sites index —
212
+ a re-run inside the window skips itself.
78
213
 
79
214
  Site details (palettes, tech stacks) are still fetched on demand and cached
80
- for 7 days.
215
+ for 7 days. Awwwards page and CDN requests have a 10-second deadline per
216
+ attempt, including response-body reading; transient page failures are retried
217
+ once, while blocks (403/429) and CDN failures are not retried.
81
218
 
82
219
  ## How it works
83
220
 
@@ -94,6 +231,185 @@ and content remain the property of Awwwards and the credited creators — don't
94
231
  bulk-scrape, redistribute, or republish them. If you use this commercially,
95
232
  review awwwards.com's terms yourself.
96
233
 
234
+ ## Built with awwwards-mcp: four real sites
235
+
236
+ Four complete sites were built through the full inspiration loop this MCP
237
+ enables, using nothing but the server's tools plus the shipped
238
+ `awwwards-inspiration` skill. Each one exercised a different corner of the
239
+ loop — and every correction the loop caught on the way became doctrine in the
240
+ skill.
241
+
242
+ > **Built in one shot, by a model that can't watch video.** All three sites
243
+ > were built in a single prompt run on **GLM 5.3-flash** — which does not
244
+ > support video input. The loop's motion study worked entirely from
245
+ > frame-tiled filmstrips (ffmpeg, 1–2 fps per element) instead of watching
246
+ > the recordings. With a video-native model, those same `get_site_elements`
247
+ > videos and `record_site_motion` .webm files could be watched directly —
248
+ > timing, easing and overlap read at full fidelity — and the motion-true
249
+ > results would be better still. The skill's frame-tile doctrine is what
250
+ > closes that gap today.
251
+
252
+ **Watch the whole loop run (1:50):**
253
+
254
+ <video src="assets/demo-loop.mp4" controls muted playsinline></video>
255
+
256
+ *Screen recording of the agent running the `awwwards-inspiration` loop end to
257
+ end with the awwwards MCP tools — searching SOTD references with inline
258
+ screenshots, pulling design DNA, frame-studying element videos, building, and
259
+ verifying with band maps + motion recording. If your client doesn't render
260
+ the player, [watch the file directly](assets/demo-loop.mp4).*
261
+
262
+ **1. [Ridge](showcase/ridge/index.html)**
263
+ ([source](showcase/ridge/)) — a Swiss-minimal single-page showcase for a fictional engineering-talent studio, direction **Aspen Search** (SOTD + Developer Award, jury 7.48): monochrome `#FAFAF8`/`#1A1A1A` + mint, giant grotesque section markers, halftone grain, asymmetric panel grid, dark discipline panels in an interior **horizontal pin passage**, count-up stats, client rows, theme toggle, cursor-follower. Built with the v1.4.0 toolkit: FTS5-ranked direction search, **both-viewports** reference captures and QA (desktop 7,849px + mobile 390×844), overflow audit (0px both), pin-center shots, film verification — and the skill-memory flywheel recorded the findings. QA evidence: `showcase/ridge/_qa/`.
264
+
265
+ **The grain-panel hover**, studied from aspensearch.com's recording and rebuilt as a canvas dither-dissolve — dots flip to mint around the mouse, the trail elongates, the boundary dissolves:
266
+
267
+ ![Ridge grain-panel hover — canvas dither-dissolve following the mouse](assets/ridge-dither.gif)
268
+
269
+ | Panel grid (desktop) | Horizontal discipline passage | Mobile 390×844 |
270
+ |---|---|---|
271
+ | ![Ridge desktop — Swiss panel grid with mint and grain](assets/ridge-home.jpg) | ![Ridge disciplines — pinned horizontal passage mid-slide](assets/ridge-disciplines.jpg) | ![Ridge mobile — stacked grid, zero overflow](assets/ridge-mobile.jpg) |
272
+
273
+ **2. [Fallow Press](fallow-press/index.html)**
274
+ ([source](fallow-press/)) — a flat-2D editorial journal, direction
275
+ **Emergence Magazine** (SOTD): pink `#FF9398` on cream and black, torn-paper
276
+ masthead (pure CSS `clip-path`, zero WebGL), giant grotesque display over
277
+ grayscale photography, serif-italic brand, three pages with **separate
278
+ horizontal** projects/about pages (GSAP ScrollTrigger pin +
279
+ `containerAnimation`).
280
+
281
+ | Torn-paper masthead (home) | Horizontal gallery (Fields) | Horizontal chapters (Practices) |
282
+ |---|---|---|
283
+ | ![Fallow Press home — torn-paper masthead over grayscale photography](assets/fallow-home.jpg) | ![Fallow Press Fields — pinned horizontal gallery panel](assets/fallow-fields.jpg) | ![Fallow Press Practices — pink quote chapter](assets/fallow-practices.jpg) |
284
+
285
+ The loop as it ran:
286
+
287
+ 1. `search_sites` (magazine filters) → shortlist judged from inline
288
+ screenshots → `get_site_details` on Emergence Magazine.
289
+ 2. **Capture before building**: `capture_live_site` + `record_site_motion`
290
+ on the live site *first*; full-page PNG and motion .webm kept in
291
+ [fallow-press/ref-motion/](fallow-press/ref-motion/) as the evidence trail.
292
+ 3. Build, then verify: full-page capture plus **panel-center pin shots** of
293
+ both horizontal pages (13 stops each, in
294
+ [fallow-press/_qa/](fallow-press/_qa/) — `capture-qa.mjs` is reusable).
295
+ 4. The pin shots caught a real bug: horizontal-panel entrances used
296
+ `toggleActions: "play none none reverse"`, and 100vw panels hide content
297
+ at midpoints on the way back — copy disappeared mid-view. Fix
298
+ (one-shot play entrances) is now doctrine: **full-viewport panels get
299
+ one-shot entrances**; QA pin shots land at panel **centers**, not uniform
300
+ fractions, or you photograph empty transition zones.
301
+
302
+ **3. Cerebrium recreation** (`C:/Users/Afjal/cerebrium-recreation/`) — a
303
+ fidelity-first recreation of cerebrium.ai, pixel-checked against the live
304
+ reference: full-page captures of both sides, `analyze_page_structure` band
305
+ compare, and SVG icon/legend fixes until the build matched the reference to
306
+ within 1px of total page height (10,871px vs 10,870px). This is the
307
+ **structure-before-pixels** doctrine at its strictest — band maps compared,
308
+ never just totals.
309
+
310
+ ![Cerebrium recreation — full-page build capture](assets/cerebrium-build.jpg)
311
+
312
+ **4. The Meridian** (`C:/Users/Afjal/editorial-site/`) — an editorial journal
313
+ built from ORDR/Hearst references: the first build to run the whole loop
314
+ end-to-end. `analyze_page_structure` caught a masthead band bug by comparing
315
+ the build's band map against the reference's; the reference captures,
316
+ motion film, and the reusable pre-scroll capture script live in
317
+ `editorial-site/_qa/`.
318
+
319
+ ![The Meridian editorial journal — full-page build capture](assets/meridian-build.jpg)
320
+
321
+ ![The Meridian — motion filmstrip from record_site_motion](assets/meridian-filmstrip.jpg)
322
+
323
+ ## Skills used to build these
324
+
325
+ | Skill | Role in the builds |
326
+ |---|---|
327
+ | `awwwards-inspiration` | The 8-step loop itself (ships with this package): search → judge from screenshots → design DNA → capture/motion study → state direction → build → band-map verify. |
328
+ | `gsap-scrolltrigger` | The horizontal pin + `containerAnimation` pattern (ease `"none"`, one-shot entrances) driving both Fallow Press horizontal pages. |
329
+ | `gsap-core` / `gsap-timeline` | Tween composition and sequenced hero entrances (torn-paper drop, panel copy rises). |
330
+ | `frontend-design` | Typography, palette and layout judgment applied when translating reference DNA into original pages. |
331
+ | `lenis` (library, via skill guidance) | smooth scrolling synced to ScrollTrigger on the Fallow Press home page. |
332
+ | `tailwindcss` / plain CSS | All builds are plain hand-rolled CSS — flat 2D, no frameworks needed. |
333
+
334
+ **The skills self-improve:** every loop pass records what verification caught (`scripts/skill-memory.mjs record`), and a deterministic distiller folds rules seen 2+ times into your installed skill copy — while the shipped copies only change via human PR. A techniques registry (`skills/_memory/techniques.json`) catalogs researched how-tos per domain (video understanding, motion detection, UI structure, micro-interactions, images).
335
+
336
+ Reduced-motion, JS-less visits, and capture tools all get graceful fallbacks
337
+ (vertical stacks; progressive-enhancement reveals).
338
+
339
+ **What the verification loop caught** — proof the structure-before-pixels
340
+ doctrine is load-bearing:
341
+
342
+ - Element **posters lie**: the first showcase build was designed from poster
343
+ frames alone and rendered a spinning 3D ring as *floating static cards*.
344
+ Downloading the element videos (`get_site_elements`) and frame-tiling them
345
+ revealed the motion truth — now the skill mandates studying motion before
346
+ animating.
347
+ - Full-page captures of reveal-on-scroll builds showed blank sections: `.reveal`
348
+ animation state vs capture's no-scroll reality. Builds ship
349
+ content-visible-without-JS progressive enhancement.
350
+ - Horizontal-panel copy vanished **mid-view** on the Fallow Press pages:
351
+ `toggleActions` reverse reverts entrances while a 100vw panel is still
352
+ holding the viewport (see above).
353
+ - Band-map compare kept the references' rhythm instead of drifting on
354
+ section heights (Cerebrium, The Meridian).
355
+
356
+ Prompt counts: **3** for the original showcase build (the build ask, the
357
+ motion correction that exposed the poster-lie, the structure pass) and
358
+ **1** for Fallow Press ("create a new website using our MCP and skills… no
359
+ 3D websites") — its two follow-ups were caught by the QA loop, not by the
360
+ user. Each correction became doctrine in the shipped `awwwards-inspiration`
361
+ skill: frame-study element videos before animating; judge page architecture
362
+ from the studied passages; tile per element, not one giant filmstrip;
363
+ capture live sites and animation **before** building.
364
+
365
+ ## Can awwwards-mcp crawl the sitemap? (robots.txt notes)
366
+
367
+ The awwwards.com `robots.txt` advertises
368
+ `Sitemap: https://www.awwwards.com/sitemap.xml` and — verified live
369
+ 2026-09-18 — **that sitemap URL returns a soft-404 HTML page** (as do common
370
+ child names like `/sitemap-websites.xml`). So sitemap discovery isn't
371
+ currently a path to more data; the polite crawl surface is exactly what the
372
+ indexer uses:
373
+
374
+ - **Allowed and used**: `/websites/`, `/websites/<filter>/`, `/sites/<slug>`
375
+ (one filter per URL; deep pagination stays un-crawled).
376
+ - **Disallowed and never fetched**: `/tag/`, `/search-websites`,
377
+ `/websites/?` (query-string pagination), `/elements/*`, `/vote/`,
378
+ favourites/likes/follows, and the rest of the 33 rules.
379
+ - Our client (`src/awwwards.ts` `buildFilterUrl`) constructs **only**
380
+ `/websites/…` paths at 1 request/second — the loop stays inside the
381
+ published rules by construction, not by convention.
382
+
383
+ ## Contributing
384
+
385
+ PRs welcome! The project especially needs **parser-drift fixes** — when live
386
+ awwwards.com markup changes, a fresh HTML snapshot attached to an issue often
387
+ becomes the new test fixture and the fastest merged PR. See
388
+ [CONTRIBUTING.md](CONTRIBUTING.md) for the full guide:
389
+
390
+ **Parser-drift is monitored automatically.** A probe script
391
+ ([scripts/parser-drift-probe.mjs](scripts/parser-drift-probe.mjs),
392
+ `npm run drift`) checks every markup anchor the parsers depend on — the
393
+ `split`/`indexOf`/regex literals in `src/parsers.ts` — against the live
394
+ listing and detail pages (2 fetches, 1 request/second, same politeness as
395
+ the client). A daily GitHub Action ([.github/workflows/parser-drift.yml](.github/workflows/parser-drift.yml))
396
+ runs it and, on drift, opens/updates a single tracking issue with the exact
397
+ anchors that changed (and auto-closes it when a later run is green). To run
398
+ it yourself: `npm run drift` (live, exit code 0/1/2) or `npm run drift -- --fixture`
399
+ (offline, checks the committed fixtures still feed every anchor). Raw HTML
400
+ is never diffed or stored — anchors only fire when the parsers actually
401
+ break, so there are no false alarms from cosmetic tweaks.
402
+
403
+ - Development setup & project layout (offline fixture-tested, no network in tests)
404
+ - How to create a PR: fork → `fix/`/`feat/`/`docs/` branch → typecheck + tests → PR template
405
+ - The politeness constraints new code must keep (1 req/s, robots.txt paths, light runtime deps)
406
+
407
+ Bugs and feature ideas start as
408
+ [issues](https://github.com/INSANE0777/Awwwards-mcp/issues/new/choose) with
409
+ templates. Security problems go privately — see
410
+ [SECURITY.md](SECURITY.md). By participating you agree to the
411
+ [Code of Conduct](CODE_OF_CONDUCT.md).
412
+
97
413
  ## Development
98
414
 
99
415
  ```bash