ai-design-context 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (288) hide show
  1. package/AGENTS.md +156 -0
  2. package/CHANGELOG.md +71 -0
  3. package/CODE_OF_CONDUCT.md +29 -0
  4. package/CONTRIBUTING.md +128 -0
  5. package/LICENSE +21 -0
  6. package/README.md +283 -0
  7. package/RELEASE_NOTES_v0.1.0.md +45 -0
  8. package/RELEASE_NOTES_v0.3.0.md +47 -0
  9. package/RELEASE_NOTES_v0.4.0.md +29 -0
  10. package/RELEASE_READINESS.md +107 -0
  11. package/ROADMAP.md +60 -0
  12. package/SECURITY.md +40 -0
  13. package/assets/logo.svg +9 -0
  14. package/benchmarks/README.md +19 -0
  15. package/benchmarks/chat.md +37 -0
  16. package/benchmarks/mobile-home-screen.md +38 -0
  17. package/benchmarks/notes-app.md +37 -0
  18. package/benchmarks/settings.md +37 -0
  19. package/benchmarks/shopping-list.md +37 -0
  20. package/benchmarks/todo-app.md +144 -0
  21. package/checklists/DESIGN_QA.md +102 -0
  22. package/checklists/GENERATED_INDEX.md +7 -0
  23. package/docs/AGENT_CONTEXT.md +111 -0
  24. package/docs/BENCHMARK.md +170 -0
  25. package/docs/DESIGNLINT_READINESS.md +110 -0
  26. package/docs/DESIGN_SYSTEM.md +23 -0
  27. package/docs/EVALUATION_RUBRIC.md +154 -0
  28. package/docs/GITHUB_SETUP.md +28 -0
  29. package/docs/GLOSSARY.md +45 -0
  30. package/docs/INDEX.md +109 -0
  31. package/docs/KNOWLEDGE_ENGINE.md +594 -0
  32. package/docs/MOBILE_FIRST.md +18 -0
  33. package/docs/NPM_RELEASE.md +43 -0
  34. package/docs/PATTERN_SPEC.md +113 -0
  35. package/docs/PHILOSOPHY.md +25 -0
  36. package/docs/PRODUCT_THINKING.md +26 -0
  37. package/docs/STYLE_GUIDE.md +49 -0
  38. package/docs/TRACEABILITY.md +45 -0
  39. package/evidence/README.md +64 -0
  40. package/evidence/TEMPLATE.md +46 -0
  41. package/evidence/chat/.gitkeep +1 -0
  42. package/evidence/mobile-home/.gitkeep +1 -0
  43. package/evidence/notes/.gitkeep +1 -0
  44. package/evidence/settings/.gitkeep +1 -0
  45. package/evidence/shopping/.gitkeep +1 -0
  46. package/evidence/todo/.gitkeep +1 -0
  47. package/evidence/todo/2026-06-25-codex-gpt-5/EVALUATION.md +66 -0
  48. package/evidence/todo/2026-06-25-codex-gpt-5/README.md +40 -0
  49. package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/generated-output.md +163 -0
  50. package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/metadata.json +38 -0
  51. package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/notes.md +20 -0
  52. package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/prompt.md +16 -0
  53. package/evidence/todo/2026-06-25-codex-gpt-5/ai-design-rules/scores.md +18 -0
  54. package/evidence/todo/2026-06-25-codex-gpt-5/baseline/generated-output.md +98 -0
  55. package/evidence/todo/2026-06-25-codex-gpt-5/baseline/metadata.json +19 -0
  56. package/evidence/todo/2026-06-25-codex-gpt-5/baseline/notes.md +19 -0
  57. package/evidence/todo/2026-06-25-codex-gpt-5/baseline/prompt.md +9 -0
  58. package/evidence/todo/2026-06-25-codex-gpt-5/baseline/scores.md +18 -0
  59. package/evidence/todo/2026-09-28-codex-paired-todo/EVALUATION.md +70 -0
  60. package/evidence/todo/2026-09-28-codex-paired-todo/README.md +52 -0
  61. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/GENERATION_NOTES.md +133 -0
  62. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/app.js +289 -0
  63. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/context-preserving-preview.json +443 -0
  64. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/daily-home-surface.json +373 -0
  65. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/quick-capture.json +301 -0
  66. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/context/sources-read.txt +36 -0
  67. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/generated-output.md +11 -0
  68. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/index.html +64 -0
  69. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/launch-prompt.md +1 -0
  70. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/metadata.json +55 -0
  71. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/notes.md +15 -0
  72. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/parent-status-message.md +1 -0
  73. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/prompt.md +49 -0
  74. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/runtime-checks.json +776 -0
  75. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/scores.md +20 -0
  76. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-completed.png +0 -0
  77. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-default.png +0 -0
  78. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-detail-reduced-motion.png +0 -0
  79. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-detail.png +0 -0
  80. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-empty.png +0 -0
  81. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-keyboard-focus.png +0 -0
  82. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-save-error.png +0 -0
  83. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-saved.png +0 -0
  84. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-saving.png +0 -0
  85. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/desktop-validation.png +0 -0
  86. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-completed.png +0 -0
  87. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-default.png +0 -0
  88. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-detail-reduced-motion.png +0 -0
  89. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-detail.png +0 -0
  90. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-empty.png +0 -0
  91. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-keyboard-focus.png +0 -0
  92. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-save-error.png +0 -0
  93. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-saved.png +0 -0
  94. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-saving.png +0 -0
  95. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/screenshots/mobile-validation.png +0 -0
  96. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/source-hashes.json +5 -0
  97. package/evidence/todo/2026-09-28-codex-paired-todo/ai-design-rules/styles.css +122 -0
  98. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/GENERATION_NOTES.md +86 -0
  99. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/app.js +303 -0
  100. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/generated-output.md +11 -0
  101. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/index.html +71 -0
  102. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/launch-prompt.md +1 -0
  103. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/metadata.json +55 -0
  104. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/notes.md +15 -0
  105. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/parent-status-message.md +1 -0
  106. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/prompt.md +49 -0
  107. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/runtime-checks.json +772 -0
  108. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/scores.md +20 -0
  109. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-completed.png +0 -0
  110. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-default.png +0 -0
  111. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-detail-reduced-motion.png +0 -0
  112. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-detail.png +0 -0
  113. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-empty.png +0 -0
  114. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-keyboard-focus.png +0 -0
  115. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-save-error.png +0 -0
  116. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-saved.png +0 -0
  117. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-saving.png +0 -0
  118. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/desktop-validation.png +0 -0
  119. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-completed.png +0 -0
  120. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-default.png +0 -0
  121. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-detail-reduced-motion.png +0 -0
  122. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-detail.png +0 -0
  123. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-empty.png +0 -0
  124. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-keyboard-focus.png +0 -0
  125. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-save-error.png +0 -0
  126. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-saved.png +0 -0
  127. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-saving.png +0 -0
  128. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/screenshots/mobile-validation.png +0 -0
  129. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/source-hashes.json +5 -0
  130. package/evidence/todo/2026-09-28-codex-paired-todo/baseline/styles.css +9 -0
  131. package/evidence/todo/2026-09-28-codex-paired-todo/capture-ai-design-rules.js +60 -0
  132. package/evidence/todo/2026-09-28-codex-paired-todo/capture-baseline.js +60 -0
  133. package/evidence/todo/2026-09-28-codex-paired-todo/common/brief.md +35 -0
  134. package/evidence/todo/2026-09-28-codex-paired-todo/common/knowledge-files.json +127 -0
  135. package/evidence/todo/2026-09-28-codex-paired-todo/common/knowledge.patch +890 -0
  136. package/evidence/todo/2026-09-28-codex-paired-todo/common/review-context.json +550 -0
  137. package/evidence/todo/2026-09-28-codex-paired-todo/common/seed.json +86 -0
  138. package/evidence/todo/2026-09-28-codex-paired-todo/common/setup.json +47 -0
  139. package/evidence/todo/2026-09-28-codex-paired-todo/common/technical-envelope.md +11 -0
  140. package/evidence/todo/2026-09-28-codex-paired-todo/layout-metrics.json +133 -0
  141. package/evidence/todo/2026-09-28-codex-paired-todo/layout-probe.js +15 -0
  142. package/examples/GENERATED_INDEX.md +8 -0
  143. package/examples/INDEX.md +16 -0
  144. package/examples/README.md +15 -0
  145. package/examples/consequential-action-confirmation-reference.md +83 -0
  146. package/examples/todo-app-benchmark.md +117 -0
  147. package/examples/todo-reference/README.md +29 -0
  148. package/examples/todo-reference/app.js +154 -0
  149. package/examples/todo-reference/favicon.svg +4 -0
  150. package/examples/todo-reference/index.html +67 -0
  151. package/examples/todo-reference/review-evidence/2026-07-15/desktop-default.png +0 -0
  152. package/examples/todo-reference/review-evidence/2026-07-15/desktop-detail.png +0 -0
  153. package/examples/todo-reference/review-evidence/2026-07-15/mobile-detail.png +0 -0
  154. package/examples/todo-reference/review-evidence/2026-07-15/mobile-error.png +0 -0
  155. package/examples/todo-reference/styles.css +241 -0
  156. package/graph/GENERATED_GRAPH.md +71 -0
  157. package/observations/README.md +23 -0
  158. package/observations/accessibility/reduced-interaction-motion.md +49 -0
  159. package/observations/accessibility/textual-error-recovery.md +50 -0
  160. package/observations/accessibility/visible-keyboard-focus.md +49 -0
  161. package/observations/ai/agent-consequential-action-controls.md +56 -0
  162. package/observations/ai/agent-work-visibility-and-intervention.md +51 -0
  163. package/observations/ai/figma-agent-context-and-skills.md +52 -0
  164. package/observations/finance/high-risk-wallet-recovery-flows.md +55 -0
  165. package/observations/finance/market-interface-adaptive-trading-surfaces.md +55 -0
  166. package/observations/interaction/chrome-scoped-view-transitions.md +52 -0
  167. package/observations/interaction/figma-motion-as-system-capability.md +53 -0
  168. package/observations/mobile/mobile-bottom-sheet-choice-contracts.md +55 -0
  169. package/observations/performance/reserved-loading-space.md +51 -0
  170. package/observations/product-repos/pinned-ai-design-rules-integration.md +49 -0
  171. package/observations/visual/apple-liquid-glass-functional-layer.md +52 -0
  172. package/observations/visual/chrome-content-first-adaptive-ui.md +51 -0
  173. package/observations/visual/material-3-expressive-system.md +52 -0
  174. package/package.json +65 -0
  175. package/patterns/GENERATED_INDEX.md +13 -0
  176. package/patterns/INDEX.md +41 -0
  177. package/patterns/README.md +5 -0
  178. package/patterns/consequential-action-confirmation.md +127 -0
  179. package/patterns/context-preserving-preview.md +145 -0
  180. package/patterns/daily-home-surface.md +134 -0
  181. package/patterns/mobile-primary-action.md +110 -0
  182. package/patterns/object-status-list.md +114 -0
  183. package/patterns/progressive-detail.md +113 -0
  184. package/patterns/quick-capture.md +132 -0
  185. package/prompts/CONSEQUENTIAL_ACTION_REVIEW.md +100 -0
  186. package/prompts/DESIGN_EVIDENCE_REVIEW.md +49 -0
  187. package/prompts/GENERATED_INDEX.md +11 -0
  188. package/prompts/INDEX.md +24 -0
  189. package/prompts/PROTOTYPE_REVIEW.md +189 -0
  190. package/prompts/QUICK_CAPTURE_STATE_REVIEW.md +74 -0
  191. package/prompts/TODO_APP_BENCHMARK.md +88 -0
  192. package/registry/objects.json +568 -0
  193. package/registry/relationships.json +972 -0
  194. package/research/GENERATED_INDEX.md +18 -0
  195. package/research/INDEX.md +40 -0
  196. package/research/accessibility/keyboard-focus-and-interaction-motion.md +62 -0
  197. package/research/accessibility/textual-error-recovery.md +57 -0
  198. package/research/performance/reserved-loading-space.md +56 -0
  199. package/research/products/apple-reminders.md +59 -0
  200. package/research/products/arc.md +61 -0
  201. package/research/products/linear.md +62 -0
  202. package/research/products/telegram.md +61 -0
  203. package/research/products/things-3.md +60 -0
  204. package/research/ux/contextual-agentic-interface.md +68 -0
  205. package/research/ux/high-risk-operational-flows.md +89 -0
  206. package/research/ux/motion-as-state-continuity.md +71 -0
  207. package/research/visual/expressive-system-ui-2026.md +75 -0
  208. package/reviews/GENERATED_INDEX.md +9 -0
  209. package/reviews/README.md +25 -0
  210. package/reviews/consequential-action-confirmation-review.md +92 -0
  211. package/reviews/todo-app-benchmark-reference-review.md +86 -0
  212. package/reviews/todo-reference-fixture-review.md +77 -0
  213. package/rules/GENERATED_INDEX.md +23 -0
  214. package/rules/INDEX.md +37 -0
  215. package/rules/accessibility/A11Y-001.md +58 -0
  216. package/rules/accessibility/A11Y-002.md +55 -0
  217. package/rules/accessibility/A11Y-003.md +54 -0
  218. package/rules/accessibility/A11Y-004.md +54 -0
  219. package/rules/ia/IA-001.md +60 -0
  220. package/rules/ia/IA-002.md +61 -0
  221. package/rules/performance/PERF-001.md +55 -0
  222. package/rules/product/PRD-001.md +61 -0
  223. package/rules/product/PRD-002.md +61 -0
  224. package/rules/ux/UX-001.md +58 -0
  225. package/rules/ux/UX-002.md +61 -0
  226. package/rules/ux/UX-003.md +60 -0
  227. package/rules/ux/UX-004.md +61 -0
  228. package/rules/ux/UX-005.md +62 -0
  229. package/rules/ux/UX-006.md +69 -0
  230. package/rules/visual/VIS-001.md +59 -0
  231. package/rules/visual/VIS-002.md +59 -0
  232. package/schema/checklist.schema.json +21 -0
  233. package/schema/common.schema.json +178 -0
  234. package/schema/observation.schema.json +21 -0
  235. package/schema/pattern.schema.json +21 -0
  236. package/schema/prompt.schema.json +21 -0
  237. package/schema/reference-project.schema.json +21 -0
  238. package/schema/research.schema.json +21 -0
  239. package/schema/review.schema.json +21 -0
  240. package/schema/rule.schema.json +21 -0
  241. package/skills/README.md +48 -0
  242. package/skills/accessibility-reviewer/SKILL.md +101 -0
  243. package/skills/agent-context/SKILL.md +34 -0
  244. package/skills/agent-context/agents/openai.yaml +4 -0
  245. package/skills/design-evidence-researcher/SKILL.md +47 -0
  246. package/skills/design-evidence-researcher/agents/openai.yaml +4 -0
  247. package/skills/design-reviewer/SKILL.md +113 -0
  248. package/skills/design-system-architect/SKILL.md +101 -0
  249. package/skills/information-architect/SKILL.md +97 -0
  250. package/skills/interaction-designer/SKILL.md +104 -0
  251. package/skills/knowledge-graph-architect/SKILL.md +96 -0
  252. package/skills/mobile-ux-expert/SKILL.md +105 -0
  253. package/skills/motion-designer/SKILL.md +100 -0
  254. package/skills/performance-reviewer/SKILL.md +97 -0
  255. package/skills/product-designer/SKILL.md +104 -0
  256. package/skills/prompt-architect/SKILL.md +105 -0
  257. package/skills/reference-driven-design/SKILL.md +119 -0
  258. package/skills/ux-reviewer/SKILL.md +99 -0
  259. package/skills/visual-designer/SKILL.md +107 -0
  260. package/starter-kit/AGENTS.md +37 -0
  261. package/starter-kit/BOOTSTRAP.md +13 -0
  262. package/starter-kit/INSTALL_WITH_AGENT.md +133 -0
  263. package/starter-kit/PROJECT_INTEGRATION.md +60 -0
  264. package/starter-kit/README.md +41 -0
  265. package/starter-kit/benchmarks/BENCHMARK_CHECKLIST.md +10 -0
  266. package/starter-kit/docs/DESIGN_DECISIONS.md +22 -0
  267. package/starter-kit/docs/INFORMATION_ARCHITECTURE.md +25 -0
  268. package/starter-kit/docs/PERSONAS.md +14 -0
  269. package/starter-kit/docs/PRD.md +32 -0
  270. package/starter-kit/docs/USER_FLOWS.md +19 -0
  271. package/starter-kit/reviews/DESIGN_REVIEW.md +49 -0
  272. package/starter-kit/templates/FEATURE_TEMPLATE.md +36 -0
  273. package/starter-kit/templates/TASK_TEMPLATE.md +22 -0
  274. package/templates/CHECKLIST_TEMPLATE.md +43 -0
  275. package/templates/OBSERVATION_TEMPLATE.md +41 -0
  276. package/templates/PATTERN_TEMPLATE.md +100 -0
  277. package/templates/PROMPT_TEMPLATE.md +46 -0
  278. package/templates/PROTOTYPE_REVIEW.md +50 -0
  279. package/templates/REFERENCE_PROJECT_TEMPLATE.md +49 -0
  280. package/templates/REVIEW_TEMPLATE.md +55 -0
  281. package/templates/RULE_TEMPLATE.md +57 -0
  282. package/tools/cli.mjs +46 -0
  283. package/tools/context.mjs +342 -0
  284. package/tools/designlint.mjs +77 -0
  285. package/tools/generate-indexes.mjs +621 -0
  286. package/tools/init.mjs +99 -0
  287. package/tools/validate-benchmark-evidence.mjs +248 -0
  288. package/tools/validate-knowledge.mjs +481 -0
@@ -0,0 +1,38 @@
1
+ {
2
+ "scenario": "Todo App",
3
+ "run_type": "ai-design-rules",
4
+ "date": "2026-06-25",
5
+ "model": "Codex GPT-5",
6
+ "provider": "OpenAI",
7
+ "generation_surface": "Codex chat session",
8
+ "temperature": "not exposed",
9
+ "tool_access": [
10
+ "filesystem write",
11
+ "shell validation"
12
+ ],
13
+ "implementation_target": "Markdown UI artifact",
14
+ "evaluator": "Codex GPT-5, single evaluator",
15
+ "rubric_version": "docs/EVALUATION_RUBRIC.md@v0.1.0",
16
+ "evidence_level": "directional",
17
+ "screenshots": "not captured",
18
+ "rules_context": [
19
+ "PRD-001",
20
+ "PRD-002",
21
+ "IA-001",
22
+ "IA-002",
23
+ "UX-001",
24
+ "UX-002",
25
+ "UX-003",
26
+ "VIS-001",
27
+ "A11Y-001"
28
+ ],
29
+ "patterns_context": [
30
+ "PAT-001",
31
+ "PAT-002",
32
+ "PAT-003",
33
+ "PAT-004",
34
+ "PAT-005",
35
+ "PAT-006"
36
+ ],
37
+ "limitation": "The same model and evaluator were used for generation and scoring. This is directional evidence, not independent validation."
38
+ }
@@ -0,0 +1,20 @@
1
+ # AI Design Rules Notes
2
+
3
+ The AI Design Rules output is more explicit about the product model, daily surface, mobile behavior, and state recovery.
4
+
5
+ Strengths:
6
+
7
+ - Prioritizes the Today surface over dashboard summaries.
8
+ - Captures tasks before metadata.
9
+ - Preserves context during detail inspection.
10
+ - Defines mobile and desktop behavior separately without changing the model.
11
+ - Includes clearer accessibility and error recovery details.
12
+
13
+ Weaknesses:
14
+
15
+ - Still a Markdown UI artifact, not a rendered implementation.
16
+ - Visual design remains descriptive rather than inspectable.
17
+ - Performance awareness is inferred through loading behavior, not measured.
18
+ - Scores are from one evaluator.
19
+
20
+ No screenshots were captured for this run.
@@ -0,0 +1,16 @@
1
+ # AI Design Rules Prompt
2
+
3
+ ```text
4
+ Design and implement a responsive todo app for everyday personal task management.
5
+
6
+ Use AI Design Rules from this repository.
7
+
8
+ Follow the knowledge graph:
9
+ research -> rules -> patterns -> prompts -> reference projects -> reviews
10
+
11
+ Use existing rules and patterns only. Do not invent new rules or patterns.
12
+
13
+ The app should let a user add tasks, see today's tasks, complete tasks, and edit details.
14
+
15
+ Prioritize low-friction capture, a stable daily home surface, mobile-first behavior, context-preserving detail, accessible touch targets, and clear empty/loading/error states.
16
+ ```
@@ -0,0 +1,18 @@
1
+ # AI Design Rules Scores
2
+
3
+ | Category | Score | Notes |
4
+ | --- | ---: | --- |
5
+ | Product Thinking | 8 | Clear daily surface, primary action, and product purpose. |
6
+ | UX | 8 | Capture, completion, detail, and recovery are direct. |
7
+ | Information Architecture | 8 | Task object and optional metadata are clear. |
8
+ | Navigation | 7 | Today surface and detail panel/sheet are stable; navigation scope is small. |
9
+ | Accessibility | 7 | Labels, touch target, error text, and keyboard order are specified. |
10
+ | Mobile-first | 8 | Mobile detail and primary action behavior are first-class. |
11
+ | Visual Hierarchy | 7 | Hierarchy is clear, though not visually rendered. |
12
+ | State Design | 8 | Empty, loading, error, failed save, and recovery are specific. |
13
+ | Consistency | 8 | Same product model across mobile and desktop. |
14
+ | Performance Awareness | 7 | Loading preserves shell and avoids full-screen spinner; not measured. |
15
+ | Simplicity | 8 | Avoids dashboard cards and keeps metadata secondary. |
16
+ | Overall Product Quality | 8 | Stronger consumer todo product direction. |
17
+
18
+ Average: 7.67
@@ -0,0 +1,98 @@
1
+ # Generated Output: Baseline Todo App
2
+
3
+ ## Product
4
+
5
+ Name: TaskBoard
6
+
7
+ TaskBoard is a responsive todo app for personal productivity. It has a polished dashboard-style layout with a header, task statistics, a task form, list filters, and a detail panel.
8
+
9
+ ## Main Screen
10
+
11
+ - Top header with app name, date, and search.
12
+ - Summary cards:
13
+ - Due Today
14
+ - Completed
15
+ - Overdue
16
+ - Add task form with:
17
+ - task title;
18
+ - due date;
19
+ - priority;
20
+ - notes.
21
+ - Filter tabs:
22
+ - Today;
23
+ - All;
24
+ - Completed.
25
+ - Task list cards with checkbox, title, due date, priority chip, and edit button.
26
+ - Right-side detail panel on desktop.
27
+ - Full-screen detail view on mobile.
28
+
29
+ ## Core Interactions
30
+
31
+ ### Add Task
32
+
33
+ The user fills the task title and optional fields, then selects `Add Task`.
34
+
35
+ If the title is empty, show inline text: `Task title is required.`
36
+
37
+ ### Complete Task
38
+
39
+ The user checks the checkbox. The task moves to completed state with reduced opacity and a completed label.
40
+
41
+ ### Edit Details
42
+
43
+ The user selects `Edit` on a task card. A detail panel opens with title, due date, priority, and notes fields.
44
+
45
+ ### Search And Filter
46
+
47
+ The user can search by title and switch tabs between Today, All, and Completed.
48
+
49
+ ## Empty State
50
+
51
+ If there are no tasks, show:
52
+
53
+ Title: `No tasks yet`
54
+
55
+ Text: `Add your first task to get started.`
56
+
57
+ Button: `Add task`
58
+
59
+ ## Loading State
60
+
61
+ Show skeleton cards for the summary cards and task list.
62
+
63
+ ## Error State
64
+
65
+ If saving fails, show a banner:
66
+
67
+ `We could not save your task. Try again.`
68
+
69
+ The user can retry.
70
+
71
+ ## Accessibility Notes
72
+
73
+ - Buttons and inputs have visible labels.
74
+ - Checkboxes use task title as label text.
75
+ - Error banner uses clear text.
76
+ - Touch target sizing is intended but not specified.
77
+
78
+ ## Responsive Behavior
79
+
80
+ Desktop:
81
+
82
+ - Header at top.
83
+ - Summary cards in a row.
84
+ - Task list and detail panel side by side.
85
+
86
+ Mobile:
87
+
88
+ - Summary cards stack.
89
+ - Add task form appears above list.
90
+ - Detail panel becomes a separate screen.
91
+
92
+ ## Known Output Characteristics
93
+
94
+ - The design is familiar and usable.
95
+ - The dashboard summary cards add some overhead for a simple consumer todo app.
96
+ - Optional metadata is visible early.
97
+ - Mobile behavior is responsive but not deeply mobile-first.
98
+ - Detail editing changes context on mobile.
@@ -0,0 +1,19 @@
1
+ {
2
+ "scenario": "Todo App",
3
+ "run_type": "baseline",
4
+ "date": "2026-06-25",
5
+ "model": "Codex GPT-5",
6
+ "provider": "OpenAI",
7
+ "generation_surface": "Codex chat session",
8
+ "temperature": "not exposed",
9
+ "tool_access": [
10
+ "filesystem write",
11
+ "shell validation"
12
+ ],
13
+ "implementation_target": "Markdown UI artifact",
14
+ "evaluator": "Codex GPT-5, single evaluator",
15
+ "rubric_version": "docs/EVALUATION_RUBRIC.md@v0.1.0",
16
+ "evidence_level": "directional",
17
+ "screenshots": "not captured",
18
+ "limitation": "The run was generated inside an active repository context, so the baseline is not perfectly isolated from prior conversation context."
19
+ }
@@ -0,0 +1,19 @@
1
+ # Baseline Notes
2
+
3
+ The baseline output is a plausible modern todo app.
4
+
5
+ Strengths:
6
+
7
+ - Includes add, complete, edit, search, and filters.
8
+ - Includes empty, loading, and error states.
9
+ - Has a clear desktop layout.
10
+
11
+ Weaknesses:
12
+
13
+ - Leans toward dashboard structure with summary cards.
14
+ - Shows optional metadata too early.
15
+ - Mobile detail behavior is less context-preserving.
16
+ - Accessibility details are mentioned but not operationally specified.
17
+ - Touch target sizing is not concrete.
18
+
19
+ No screenshots were captured for this run.
@@ -0,0 +1,9 @@
1
+ # Baseline Prompt
2
+
3
+ ```text
4
+ Design and implement a responsive todo app for everyday personal task management.
5
+
6
+ The app should let a user add tasks, see today's tasks, complete tasks, and edit details.
7
+
8
+ Make it modern, usable, and polished.
9
+ ```
@@ -0,0 +1,18 @@
1
+ # Baseline Scores
2
+
3
+ | Category | Score | Notes |
4
+ | --- | ---: | --- |
5
+ | Product Thinking | 6 | Understands todo app basics but prioritizes dashboard stats. |
6
+ | UX | 6 | Core task flow exists, with moderate form friction. |
7
+ | Information Architecture | 6 | Task object is clear, but summary cards add weight. |
8
+ | Navigation | 6 | Tabs and detail panel are understandable. |
9
+ | Accessibility | 5 | Labels are mentioned, but target size and focus are not specific. |
10
+ | Mobile-first | 5 | Responsive behavior exists, but desktop structure leads. |
11
+ | Visual Hierarchy | 6 | Polished structure, but stats compete with tasks. |
12
+ | State Design | 5 | States exist but are generic. |
13
+ | Consistency | 6 | Controls and cards are coherent. |
14
+ | Performance Awareness | 5 | Loading skeleton exists; responsiveness is not deeply addressed. |
15
+ | Simplicity | 5 | More dashboard-like than necessary. |
16
+ | Overall Product Quality | 6 | Usable baseline with visible consumer-product gaps. |
17
+
18
+ Average: 5.58
@@ -0,0 +1,70 @@
1
+ # Todo App Evaluation
2
+
3
+ Scenario: Todo App, unchanged common brief from `benchmarks/todo-app.md`.
4
+
5
+ Run: 2026-09-28 paired rendered generation. Model: inherited Codex session model; exact backend identifier and sampling settings unexposed. Evaluator: Codex parent agent, internal and unblinded. Rubric: `docs/EVALUATION_RUBRIC.md` at commit `3d9a6a70c47b7008a47dafd24a4aeef2855b07f2`.
6
+
7
+ ## Verdict
8
+
9
+ **NEEDS WORK** for both prototypes. All scripted core operations completed, but the baseline has undersized capture/navigation targets and an error-layout shift; the rules-assisted output makes keyboard capture unnecessarily distant. Both resize the capture button during saving. The small rubric difference is a judgment on this pair, not a measured universal effect or proof that the next release improves all outputs.
10
+
11
+ ## Scores
12
+
13
+ Scores are whole-number evaluator judgments; averages weight the twelve dimensions equally.
14
+
15
+ | Category | Baseline | AI Design Rules | Notes |
16
+ | --- | ---: | ---: | --- |
17
+ | Product Thinking | 8 | 8 | Both prioritize Today, task text, and title-only capture. [Baseline](baseline/screenshots/mobile-default.png), [rules](ai-design-rules/screenshots/mobile-default.png). |
18
+ | UX | 7 | 7 | Both retain failed input and retry successfully; rules improve touch reach but require 29 Tabs to capture versus 4. [Baseline runtime](baseline/runtime-checks.json), [rules runtime](ai-design-rules/runtime-checks.json). |
19
+ | Information Architecture | 8 | 8 | Both use the same task model, Today/All views, optional detail, and completed grouping. [Baseline](baseline/screenshots/desktop-default.png), [rules](ai-design-rules/screenshots/desktop-default.png). |
20
+ | Navigation | 8 | 8 | Both return focus to the source task and preserve scroll in the tested detail path; immediate dismissal and reduced motion also preserve function. Runtime `detailClose`, `interruption`, and `reducedMotion` checks. |
21
+ | Accessibility | 6 | 7 | Baseline capture input is 34px high on mobile and navigation buttons 43px; rules default controls reach 44px and provide explicit textual validation/focus, but the long keyboard path remains. [Baseline focus](baseline/screenshots/mobile-keyboard-focus.png), [rules focus](ai-design-rules/screenshots/mobile-keyboard-focus.png). |
22
+ | Mobile-first | 7 | 8 | Both fit 390px without horizontal overflow; rules keep capture reachable at the bottom while scrolling, baseline capture leaves view. Physical keyboard overlap remains untested. [Baseline completed view](baseline/screenshots/mobile-completed.png), [rules completed view](ai-design-rules/screenshots/mobile-completed.png). |
23
+ | Visual Hierarchy | 8 | 8 | Both are restrained, readable, task-first layouts with similar green/neutral styling. There is no convincing visual-hierarchy advantage in this pair. Default and detail screenshots. |
24
+ | State Design | 8 | 8 | Empty, pending, failure, success, completion, detail, validation, and retry are observable in both. Failed edits retain drafts and successful edits persist after reload. [Baseline failure](baseline/screenshots/mobile-save-error.png), [rules failure](ai-design-rules/screenshots/mobile-save-error.png). |
25
+ | Consistency | 8 | 8 | Repeated task controls and labels are coherent; desktop/mobile retain the same product model. Both have a saving-button width change. [Geometry](layout-metrics.json). |
26
+ | Performance Awareness | 6 | 7 | Both provide pending feedback and block duplicate saves. Baseline error shifts the first row down 22px; rules keep it fixed. Both shrink the capture input as the saving label expands. These are geometry observations, not Core Web Vitals measurements. [Geometry](layout-metrics.json). |
27
+ | Simplicity | 8 | 8 | Neither adds authentication, analytics, teams, or project management. Optional date and notes stay in details. [Baseline details](baseline/screenshots/mobile-detail.png), [rules details](ai-design-rules/screenshots/mobile-detail.png). |
28
+ | Overall Product Quality | 7 | 7 | Both are usable local prototypes with clear core flows and meaningful remaining friction; neither merits a production-readiness claim. |
29
+
30
+ Baseline total: **89/120**, average **7.42/10**.
31
+
32
+ AI Design Rules total: **92/120**, average **7.67/10**.
33
+
34
+ Difference: **+3 rubric points**, or **+0.25/10** in the average. Per-category differences: Accessibility +1, Mobile-first +1, Performance Awareness +1; all other categories 0. This small difference has no statistical significance claim.
35
+
36
+ ## Findings
37
+
38
+ 1. **Medium — rules-assisted keyboard capture is distant.** From a fresh page with twelve seed tasks, sequential Tab reaches capture after 29 presses on both viewports; baseline takes 4. The dock is visually prominent but follows all task controls in DOM order. The existing Skip to tasks link targets the list rather than capture. Evidence: runtime `keyboard.tabsToCapture`, [rules HTML](ai-design-rules/index.html), focus screenshots. Related guidance: `PRD-002` / `PAT-002` low-friction capture, `A11Y-003` visible focus. Follow-up: review a direct keyboard path to capture and its discoverability without positive tabindex ordering; test it with realistic list length. There is no new universal numeric Tab limit.
39
+ 2. **Medium — baseline controls miss the repository's touch target bar.** The mobile capture input is 34px high, Today/All controls 43px. Completion hit areas were measured using their enclosing labels and do meet 44px. Rules-assisted default controls have no below-44px target in the recorded audit. Evidence: runtime `targets`, [layout measurements](layout-metrics.json). Related guidance: `A11Y-001`. Follow-up: expand actual input/navigation hit areas, not just the drawn icon or container.
40
+ 3. **Medium — baseline error feedback moves the list.** At 390×844, the first row moves from y=426.671875 to y=448.671875 after the injected save failure, a 22px shift. The rules-assisted row stays at y=306.5. Evidence: [geometry](layout-metrics.json), mobile error screenshots. Related guidance: `PERF-001`, `PAT-002`. Follow-up: reserve the recovery message/action geometry, including wrapped text.
41
+ 4. **Low — both saving labels change capture geometry.** Baseline input width falls from 239.3125 to 209.6875px; rules input falls from 242.546875 to 207.78125px. Add/Adding/Saving/Retry buttons consume different widths. Evidence: [geometry](layout-metrics.json). Related guidance: `PERF-001`. Follow-up: test all label states and reserve a stable action width without harming narrow-screen input usability.
42
+
43
+ Frozen outputs were not patched. These findings guide the next knowledge/review iteration.
44
+
45
+ ## Coverage And Evidence
46
+
47
+ Both variants were observed at 390×844 and 1440×900 in the same Headless Chrome build. Each stores twenty screenshots: ten per viewport, including the nine protocol states and one additional validation screenshot.
48
+
49
+ | Check | Baseline | AI Design Rules |
50
+ | --- | --- | --- |
51
+ | Seed data / default overflow | 12 tasks; no horizontal overflow | 12 tasks; no horizontal overflow |
52
+ | Failed capture | Input kept, count remains 12 | Input kept, count remains 12 |
53
+ | Successful retry | Exactly one Buy apples task | Exactly one Buy apples task |
54
+ | Completion | Persisted complete state | Persisted complete state |
55
+ | Detail close | Source focus restored, scroll delta 0 | Source focus restored, scroll delta 0 |
56
+ | Immediate open/dismiss | Stays closed; source focus restored | Stays closed; source focus restored |
57
+ | Reduced motion | Same details/return; 0 active animations at capture | Same details/return; 0 active animations at capture |
58
+ | Failed edit / retry / reload | Draft kept, correct committed notes survive reload | Draft kept, correct committed notes survive reload |
59
+ | Blank title | Native required-field message | Explicit Task title error |
60
+ | JavaScript page errors during scripted passes | None observed | None observed |
61
+
62
+ The initial local server produced a favicon 404; this was not an application failure. Native validation messages and focus indicators were visually inspected. Formal contrast, assistive technology, physical mobile keyboard, alternative browsers, and quantitative performance remain untested. Target measurements above cover default-page controls, not a complete accessibility audit of every state.
63
+
64
+ Traceability: the rules run saved its three [context bundles](ai-design-rules/context/) and [source inventory](ai-design-rules/context/sources-read.txt). Applied areas include `PAT-001`, `PAT-002`, `PAT-003`, `PAT-005`, `PAT-006`; `PRD-001/002`, `IA-001/002`, `UX-001/002/003/004`, `VIS-001/002`, `A11Y-001/002/003/004`, and `PERF-001`. No rule was scored merely because its identifier was mentioned.
65
+
66
+ ## Limits And Next Validation
67
+
68
+ This is one pair from the same inherited agent configuration with a shared host and an unblinded internal evaluator. Exact backend version and sampling settings were unavailable. It compares baseline with the whole current knowledge treatment; it does not isolate context retrieval changes or compare the previous release against this one. Both runs converged on a similar name and visual style, so this is especially weak evidence of visual diversity or distinct visual quality.
69
+
70
+ Repeat after tightening keyboard-path and state-geometry review, add independent or blinded evaluation, and test a real mobile keyboard before broader release claims. Keep existing rule/pattern maturity unchanged. See [run setup](common/setup.json), [baseline prompt](baseline/prompt.md), [rules prompt](ai-design-rules/prompt.md), and each variant's generation notes for exact inputs and limitations.
@@ -0,0 +1,52 @@
1
+ # Paired Rendered Todo Benchmark — 2026-09-28
2
+
3
+ Two fresh Codex worker contexts generated the same Todo brief, one without AI Design Rules and one with the current knowledge graph. Both received the same seed data, technical envelope, tools, and ten-minute maximum. This directory stores the frozen source, prompts, context, forty screenshots, runtime observations, and an internal evaluation.
4
+
5
+ **Result:** baseline **7.42/10**, AI Design Rules **7.67/10**, a **+0.25** difference in this evaluator's rubric average. The result is mixed: larger touch targets and more stable error layout accompany a longer keyboard path to capture. Visual hierarchy scored equally. This one unblinded pair does not establish a general quality gain.
6
+
7
+ - [Full evaluation and findings](EVALUATION.md)
8
+ - [Shared setup and limitations](common/setup.json)
9
+ - [Unchanged common brief](common/brief.md)
10
+ - [Shared task data](common/seed.json)
11
+ - [Frozen knowledge patch](common/knowledge.patch) and [file hashes](common/knowledge-files.json)
12
+ - [Baseline output](baseline/generated-output.md) and [rules-assisted output](ai-design-rules/generated-output.md)
13
+ - [Measured mobile geometry](layout-metrics.json)
14
+
15
+ | Baseline | AI Design Rules |
16
+ | --- | --- |
17
+ | ![Baseline mobile default](baseline/screenshots/mobile-default.png) | ![Rules-assisted mobile default](ai-design-rules/screenshots/mobile-default.png) |
18
+
19
+ ## Inspect The Frozen Apps
20
+
21
+ From this directory:
22
+
23
+ ```sh
24
+ python3 -m http.server 4183 --bind 127.0.0.1
25
+ ```
26
+
27
+ Open `http://127.0.0.1:4183/baseline/` and `http://127.0.0.1:4183/ai-design-rules/`. Each app has its own storage key. The apps deliberately treat September 28, 2026 as Today and use simulated local saves.
28
+
29
+ ## Reproduce The Browser Checks
30
+
31
+ Use Playwright CLI with a locally installed Chrome. The recorded run used Headless Chrome 153.0.8010.53, en-GB locale, light scheme, CSS pixel scale 1. Run the CLI from this directory so screenshot paths resolve correctly. The scripts will overwrite this run's screenshots if executed here; copy the directory before rerunning.
32
+
33
+ ```sh
34
+ playwright-cli -s=todo-replay open http://127.0.0.1:4183/baseline/ --browser chrome
35
+ playwright-cli -s=todo-replay snapshot
36
+ playwright-cli -s=todo-replay --json run-code --filename capture-baseline.js
37
+ playwright-cli -s=todo-replay goto http://127.0.0.1:4183/ai-design-rules/
38
+ playwright-cli -s=todo-replay snapshot
39
+ playwright-cli -s=todo-replay --json run-code --filename capture-ai-design-rules.js
40
+ playwright-cli -s=todo-replay --json run-code --filename layout-probe.js
41
+ playwright-cli -s=todo-replay close
42
+ ```
43
+
44
+ The runtime files contain the parsed `result` returned by these scripts. Screenshots cover mobile and desktop default, empty, saving, failed save, successful retry, completed task, details, reduced-motion details, keyboard focus, and validation. Interaction results additionally cover edit failure/retry, persistence, immediate dialog dismissal, and focus/scroll restoration. No frozen application code was changed after generation.
45
+
46
+ ## Provenance And Limits
47
+
48
+ Knowledge input: commit `3d9a6a70c47b7008a47dafd24a4aeef2855b07f2` plus the saved pre-generation patch. The two workers used `fork_turns=none`, the same worker role and inherited model configuration, with no model/reasoning override. Actual generation times were 7m09s and 7m34s within the same ten-minute cap. Tools were restricted by prompt to filesystem access and shell syntax checks; both generators reported no browser, network, sibling output, or existing implementation access.
49
+
50
+ The exact backend model identifier and sampling settings were not exposed. Isolation was prompt-enforced on a shared host, not OS-enforced. The parent evaluator knew the conditions and had authored the tested knowledge changes. This is neither blinded nor independent evaluation. No physical mobile keyboard, screen reader, cross-browser, formal contrast, or performance audit was performed. No rule or pattern maturity was promoted.
51
+
52
+ Transient machine paths in prompt/notes artifacts were replaced by `${RUN_ROOT}` and `${KNOWLEDGE_ROOT}`; all other prompt text was preserved. Metadata records both original and stored prompt hashes. Raw application bytes and screenshots were not modified. The attempted preliminary `spark_explorer` audit failed because that model was unavailable; it generated no benchmark output and is not one of the measured runs.
@@ -0,0 +1,133 @@
1
+ # Generation notes
2
+
3
+ Condition: AI Design Rules. Fresh static implementation in this owned directory only.
4
+
5
+ ## Timing and scope
6
+
7
+ - Started: 2026-09-28 05:37:39 UTC.
8
+ - Finished: 2026-09-28 05:45:13 UTC (7 minutes 34 seconds; below the 10-minute cap).
9
+ - Files: `index.html`, `styles.css`, `app.js`, this note, three raw context JSON files, and the source inventory in `context/sources-read.txt`.
10
+ - No sibling output, prior generated app, `examples/todo-reference`, network, browser, or subagent was inspected or used.
11
+ - Repository knowledge and shared inputs were not edited. `git status --short` was run before implementation; the repository already contained modified knowledge/tooling files, treated as the frozen input.
12
+
13
+ ## Product direction and evidence boundary
14
+
15
+ User goal: capture everyday work, act on today's list, and inspect a task without rebuilding list context. The core object is the supplied task: id, title, notes, dueDate, completed. Today is the stable home; All tasks is another view of the same objects. Capture needs only a title, defaults visibly to Today, and keeps notes/date in task details. Today includes due-today tasks and unfinished overdue tasks; All tasks keeps later/undated tasks discoverable after date edits.
16
+
17
+ `PAT-00001` guides the daily home. `PAT-00002` and `PAT-00003` guide a fixed, reachable input and one submit action. `PAT-00005` guides temporary task detail with source scroll and focus restoration. `PAT-00006` is used only for simple done/undone meaning, not administrative metadata or columns.
18
+
19
+ Applied rules: PRD-001/002, IA-001/002, UX-001/002/003, VIS-001/002, A11Y-001/002/003/004, PERF-001. UX-004 was considered: no nonessential motion is introduced; open and close immediately reach the same functional state, including with reduced motion. Semantic CSS roles organize neutral content, supporting surfaces, teal primary actions, error and success. Task text dominates; no analytics, projects, teams, decorative cards, or external assets.
20
+
21
+ All retrieved research, rules, patterns, and review objects are **draft / seed**. These are bounded design inputs, not evidence of validated usability, performance, accessibility compliance, or benchmark outcomes. Review instructions informed the state inventory and the explicit unrendered-evidence gap. Their optional advice to ask for design alternatives was not applicable to the authorized independent single-output benchmark. No knowledge objects or rules were invented.
22
+
23
+ ## Local serving
24
+
25
+ From this output directory:
26
+
27
+ ```sh
28
+ python3 -m http.server 4173 --bind 127.0.0.1
29
+ ```
30
+
31
+ Open `http://127.0.0.1:4173/`. This command is provided for evaluation; no server/browser was launched by the generator. All assets are local and no build step or dependency installation is needed. The initial loading shell lasts 350 ms. Local state uses the key `daylight-adr-benchmark-v1` on the serving origin.
32
+
33
+ ## Evaluator hooks
34
+
35
+ Hooks are available only via `window.__benchmark`; there are no product test controls.
36
+
37
+ ```js
38
+ window.__benchmark.reset('seed')
39
+ window.__benchmark.reset('empty')
40
+ window.__benchmark.setSaveDelay(2500)
41
+ window.__benchmark.failNextSave()
42
+ window.__benchmark.getState()
43
+ ```
44
+
45
+ - `reset` cancels stale pending results using an epoch, closes details, clears draft/error/view state, resets delay to 500 ms, clears the failure flag, persists the requested dataset where storage is available, and renders immediately. Seed matches the shared JSON exactly. Empty uses no tasks.
46
+ - `setSaveDelay` changes the delay for operations initiated afterward; default 500 ms. It does not affect the initial 350 ms loading shell.
47
+ - `failNextSave` fails exactly the next initiated save once; invalid blank submission does not consume the failure flag.
48
+ - `getState` returns a JSON-serializable copy of tasks, view, loading, selection, status, pending/failed operation, delay, one-shot failure flag, capture draft, detail drafts, and completed-section visibility.
49
+
50
+ ## Exact state reproduction
51
+
52
+ 1. Default: reset seed; Today contains 11 incomplete tasks and one completed task. The first task has the supplied note.
53
+ 2. Empty: reset empty. The empty explanation and Add a task control focus the persistent capture input.
54
+ 3. Initial loading: reset seed, then reload the page. For 350 ms the home shell and row placeholders remain in place; capture is temporarily disabled. No browser/network condition is needed.
55
+ 4. Save loading: reset seed; set delay to 2500; type `Buy milk` into the bottom title input; press Enter or Add. Title remains visible, Add reads Saving, submission is disabled, existing tasks stay in place. After 2500 ms the saved task appears at the top.
56
+ 5. Capture failure: reset seed; call failNextSave; type `Buy milk`; submit. After 500 ms the title remains, the message says it was not saved, and the same button reads Retry. Retry succeeds once and adds exactly one task. The next ordinary save succeeds.
57
+ 6. Validation: reset empty; submit a blank/whitespace title. A textual Task title error is exposed and input focus remains available. Enter a title and submit to recover.
58
+ 7. Completion: reset seed; click the check target beside Book dentist appointment. After the save delay it appears under Completed; uncheck to restore it. For failed completion, call failNextSave first; the original task state remains and the bottom Retry repeats that exact action.
59
+ 8. Details: reset seed; scroll to Order coffee beans; click its title. Desktop uses a right-side dialog; mobile uses a bottom sheet. Close button, Close action, Escape, and clicking outside dismiss it. Source view/scroll is retained, focus returns to the surviving task control. No opening animation can cause stale selection or focus.
60
+ 9. Edit success: open a task; change title, date, or notes; Save changes. Feedback appears inside details and in the page status. Clearing the date makes a task visible in All tasks; closing details returns to a surviving control.
61
+ 10. Edit failure: open a task; change notes; call failNextSave; Save changes. The form values remain with recovery text. Save changes retries using the current inputs. Close and reopen also retains the draft during the same session.
62
+ 11. Reduced motion: enable the operating-system/browser preference. There is no interaction animation; all information and controls remain immediate.
63
+
64
+ ## Checks actually run
65
+
66
+ - `node --check app.js` — passed after implementation and after the final static corrections.
67
+ - Python filesystem checks — passed: unique HTML ids, label-to-input targets, literal JS element references, exact embedded seed equality, parseable raw context JSON, and no external asset URLs in HTML.
68
+ - The context CLI help and all three requested context resolutions completed successfully.
69
+ - No runtime interaction tests, browser tools, visual rendering, viewport screenshots, target-size measurements, contrast measurements, keyboard navigation tests, or assistive-technology tests were run because the envelope permits filesystem and shell syntax checks only. No rendered validation is claimed.
70
+ - Repository build/validation was not run because no repository knowledge or tooling was changed.
71
+
72
+ ## Known limitations and follow-ups
73
+
74
+ - Modern native HTML dialog support is assumed. Cross-browser focus/scroll behavior and virtual-keyboard placement remain unverified.
75
+ - Tasks persist on this device/origin only; there is no backend, synchronization, or reminder notification service. Unsaved detail drafts stay in memory during the session and are not preserved through a reload.
76
+ - Storage-write failure is treated as a real failed save with retained input and retry; storage availability was not runtime-tested.
77
+ - Saves are serialized. During a pending operation mutation controls are disabled; task inspection and closing details remain available.
78
+ - Dates and Today are intentionally anchored to 2026-09-28 for reproducibility. Generated task ids use the local timestamp; the supplied seed is unchanged.
79
+ - Parent evaluation should measure 390×844 and 1440×900, inspect sticky capture overlap/keyboard behavior, verify dialog focus restoration, and exercise the documented pending/failure/retry paths. This is a verification gap, not a claim of successful rendered behavior.
80
+
81
+ ## Actual inputs and sources read
82
+
83
+ Common input: `${RUN_ROOT}/common/seed.json`.
84
+ Instructions: this run's `prompt.md` and repository `AGENTS.md`.
85
+
86
+ Context commands (npm's silent option suppresses its wrapper while preserving raw CLI JSON):
87
+
88
+ ```sh
89
+ npm run --silent context -- --task daily-home-surface --platform mobile --intent implement --format json
90
+ npm run --silent context -- --task quick-capture --platform mobile --intent implement --format json
91
+ npm run --silent context -- --task context-preserving-preview --platform mobile --intent implement --format json
92
+ ```
93
+
94
+ Their exact JSON is saved under `context/` with matching task slugs. Every returned object path was read. The first concatenated read exceeded output limits, so research and the affected patterns, prompts, accessibility, IA, and performance rule files were read again in bounded batches. Two local skills were additionally read and applied. No external links within research were fetched.
95
+
96
+ Repository-relative source inventory:
97
+
98
+ - `AGENTS.md`
99
+ - `checklists/DESIGN_QA.md`
100
+ - `patterns/context-preserving-preview.md`
101
+ - `patterns/daily-home-surface.md`
102
+ - `patterns/mobile-primary-action.md`
103
+ - `patterns/object-status-list.md`
104
+ - `patterns/quick-capture.md`
105
+ - `prompts/PROTOTYPE_REVIEW.md`
106
+ - `prompts/QUICK_CAPTURE_STATE_REVIEW.md`
107
+ - `research/accessibility/keyboard-focus-and-interaction-motion.md`
108
+ - `research/accessibility/textual-error-recovery.md`
109
+ - `research/performance/reserved-loading-space.md`
110
+ - `research/products/apple-reminders.md`
111
+ - `research/products/arc.md`
112
+ - `research/products/linear.md`
113
+ - `research/products/telegram.md`
114
+ - `research/products/things-3.md`
115
+ - `research/ux/motion-as-state-continuity.md`
116
+ - `research/visual/expressive-system-ui-2026.md`
117
+ - `rules/accessibility/A11Y-001.md`
118
+ - `rules/accessibility/A11Y-002.md`
119
+ - `rules/accessibility/A11Y-003.md`
120
+ - `rules/accessibility/A11Y-004.md`
121
+ - `rules/ia/IA-001.md`
122
+ - `rules/ia/IA-002.md`
123
+ - `rules/performance/PERF-001.md`
124
+ - `rules/product/PRD-001.md`
125
+ - `rules/product/PRD-002.md`
126
+ - `rules/ux/UX-001.md`
127
+ - `rules/ux/UX-002.md`
128
+ - `rules/ux/UX-003.md`
129
+ - `rules/ux/UX-004.md`
130
+ - `rules/visual/VIS-001.md`
131
+ - `rules/visual/VIS-002.md`
132
+ - `skills/agent-context/SKILL.md`
133
+ - `skills/product-designer/SKILL.md`