@mohammadhprp/system-prompt 0.11.1 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (121) hide show
  1. package/framework/agents/backend-architect.md +1 -1
  2. package/framework/agents/researcher.md +1 -1
  3. package/framework/mcps/figma-mcp-go/README.md +0 -1
  4. package/framework/mcps/github-mcp/README.md +1 -1
  5. package/framework/mcps/github-mcp/configs/.env.example +1 -1
  6. package/framework/mcps/gitlab-mcp/README.md +0 -1
  7. package/framework/mcps/jira-mcp/README.md +0 -1
  8. package/framework/mcps/laravel-boost/README.md +0 -1
  9. package/framework/mcps/notion-mcp/README.md +0 -1
  10. package/framework/mcps/supabase-mcp/README.md +0 -1
  11. package/framework/plugins/opencode-goal-plugin/README.md +0 -1
  12. package/framework/plugins/ponytail/README.md +1 -1
  13. package/framework/plugins/ponytail/capabilities.md +1 -1
  14. package/framework/references/standards/api.md +0 -1
  15. package/framework/references/standards/architecture.md +0 -1
  16. package/framework/references/standards/database.md +0 -1
  17. package/framework/references/standards/debugging.md +0 -1
  18. package/framework/references/standards/documentation.md +0 -2
  19. package/framework/references/standards/logging.md +0 -1
  20. package/framework/references/standards/naming.md +0 -1
  21. package/framework/references/standards/observability.md +0 -1
  22. package/framework/references/standards/performance.md +0 -1
  23. package/framework/references/standards/pull-requests.md +1 -2
  24. package/framework/references/standards/security.md +0 -1
  25. package/framework/references/standards/testing.md +0 -1
  26. package/framework/skills/README.md +16 -4
  27. package/framework/skills/codenavi/SKILL.md +306 -0
  28. package/framework/skills/codenavi/examples.md +33 -0
  29. package/framework/skills/codenavi/references/coding-principles.md +143 -0
  30. package/framework/skills/codenavi/references/notebook-spec.md +171 -0
  31. package/framework/skills/create-adr/SKILL.md +429 -0
  32. package/framework/skills/create-adr/examples.md +35 -0
  33. package/framework/skills/diagram-design/SKILL.md +1 -1
  34. package/framework/skills/docs-writer/SKILL.md +39 -0
  35. package/framework/skills/docs-writer/examples.md +34 -0
  36. package/framework/skills/docs-writer/references/style-guide.md +72 -0
  37. package/framework/skills/frontend-design/SKILL.md +55 -0
  38. package/framework/skills/frontend-design/examples.md +45 -0
  39. package/framework/skills/humanizer/SKILL.md +412 -0
  40. package/framework/skills/humanizer/examples.md +46 -0
  41. package/framework/skills/learning-opportunities/SKILL.md +140 -0
  42. package/framework/skills/learning-opportunities/examples.md +34 -0
  43. package/framework/skills/learning-opportunities/references/PRINCIPLES.md +42 -0
  44. package/framework/skills/perf-web-optimization/SKILL.md +163 -0
  45. package/framework/skills/perf-web-optimization/examples.md +35 -0
  46. package/framework/skills/perf-web-optimization/references/bundle-optimization.md +180 -0
  47. package/framework/skills/perf-web-optimization/references/core-web-vitals.md +154 -0
  48. package/framework/skills/perf-web-optimization/references/image-optimization.md +170 -0
  49. package/framework/skills/security-best-practices/LICENSE.txt +201 -0
  50. package/framework/skills/security-best-practices/SKILL.md +89 -0
  51. package/framework/skills/security-best-practices/examples.md +35 -0
  52. package/framework/skills/security-best-practices/references/golang-general-backend-security.md +988 -0
  53. package/framework/skills/security-best-practices/references/javascript-express-web-server-security.md +1151 -0
  54. package/framework/skills/security-best-practices/references/javascript-general-web-frontend-security.md +725 -0
  55. package/framework/skills/security-best-practices/references/javascript-jquery-web-frontend-security.md +672 -0
  56. package/framework/skills/security-best-practices/references/javascript-typescript-nextjs-web-server-security.md +1138 -0
  57. package/framework/skills/security-best-practices/references/javascript-typescript-react-web-frontend-security.md +975 -0
  58. package/framework/skills/security-best-practices/references/javascript-typescript-vue-web-frontend-security.md +789 -0
  59. package/framework/skills/security-best-practices/references/python-django-web-server-security.md +880 -0
  60. package/framework/skills/security-best-practices/references/python-fastapi-web-server-security.md +1030 -0
  61. package/framework/skills/security-best-practices/references/python-flask-web-server-security.md +835 -0
  62. package/framework/skills/sentry/SKILL.md +127 -0
  63. package/framework/skills/sentry/examples.md +34 -0
  64. package/framework/skills/sentry/scripts/sentry_api.py +238 -0
  65. package/framework/skills/show-me/SKILL.md +127 -0
  66. package/framework/skills/show-me/examples.md +78 -0
  67. package/framework/skills/spec-driven-eval/SKILL.md +341 -0
  68. package/framework/skills/spec-driven-eval/examples.md +35 -0
  69. package/framework/skills/spec-driven-eval/references/quickstart.md +118 -0
  70. package/framework/skills/spec-driven-eval/references/reference.md +295 -0
  71. package/framework/skills/taste/SKILL.md +47 -0
  72. package/framework/skills/taste/examples.md +73 -0
  73. package/framework/skills/taste/references/brandkit.md +793 -0
  74. package/framework/skills/taste/references/design-taste-frontend-v1.md +221 -0
  75. package/framework/skills/taste/references/design-taste-frontend.md +1201 -0
  76. package/framework/skills/taste/references/full-output-enforcement.md +44 -0
  77. package/framework/skills/taste/references/gpt-taste.md +69 -0
  78. package/framework/skills/taste/references/high-end-visual-design.md +93 -0
  79. package/framework/skills/taste/references/image-to-code.md +1223 -0
  80. package/framework/skills/taste/references/imagegen-frontend-mobile.md +1460 -0
  81. package/framework/skills/taste/references/imagegen-frontend-web.md +982 -0
  82. package/framework/skills/taste/references/industrial-brutalist-ui.md +87 -0
  83. package/framework/skills/taste/references/minimalist-ui.md +80 -0
  84. package/framework/skills/taste/references/redesign-existing-projects.md +173 -0
  85. package/framework/skills/taste/references/stitch-design-taste-design.md +121 -0
  86. package/framework/skills/taste/references/stitch-design-taste.md +179 -0
  87. package/framework/skills/technical-design-doc-creator/README.md +411 -0
  88. package/framework/skills/technical-design-doc-creator/SKILL.md +1484 -0
  89. package/framework/skills/technical-design-doc-creator/examples.md +35 -0
  90. package/framework/skills/tlc-spec-driven/SKILL.md +184 -0
  91. package/framework/skills/tlc-spec-driven/examples.md +34 -0
  92. package/framework/skills/tlc-spec-driven/references/code-analysis.md +98 -0
  93. package/framework/skills/tlc-spec-driven/references/coding-principles.md +72 -0
  94. package/framework/skills/tlc-spec-driven/references/context-limits.md +31 -0
  95. package/framework/skills/tlc-spec-driven/references/design.md +199 -0
  96. package/framework/skills/tlc-spec-driven/references/discuss.md +159 -0
  97. package/framework/skills/tlc-spec-driven/references/implement.md +436 -0
  98. package/framework/skills/tlc-spec-driven/references/lessons.md +115 -0
  99. package/framework/skills/tlc-spec-driven/references/memory.md +144 -0
  100. package/framework/skills/tlc-spec-driven/references/specify.md +228 -0
  101. package/framework/skills/tlc-spec-driven/references/sub-agents.md +147 -0
  102. package/framework/skills/tlc-spec-driven/references/tasks.md +451 -0
  103. package/framework/skills/tlc-spec-driven/references/validate.md +355 -0
  104. package/framework/skills/tlc-spec-driven/scripts/check_commit.py +115 -0
  105. package/framework/skills/tlc-spec-driven/scripts/lessons.py +412 -0
  106. package/framework/skills/tlc-spec-driven/scripts/validate_spec.py +260 -0
  107. package/framework/skills/tlc-spec-driven/scripts/validate_state.py +162 -0
  108. package/framework/skills/tlc-spec-driven/scripts/validate_tasks.py +251 -0
  109. package/framework/skills/web-design-guidelines/SKILL.md +65 -0
  110. package/framework/skills/web-design-guidelines/examples.md +32 -0
  111. package/framework/skills/web-design-guidelines/references/guideline.md +174 -0
  112. package/package.json +1 -1
  113. package/src/agent-configs.js +6 -2
  114. package/src/catalog.js +16 -3
  115. package/src/cli.js +82 -2
  116. package/src/doctor.js +84 -0
  117. package/src/installer.js +337 -154
  118. package/framework/skills/backend-engineer/SKILL.md +0 -76
  119. package/framework/skills/backend-engineer/examples.md +0 -31
  120. package/framework/skills/documentation/SKILL.md +0 -74
  121. package/framework/skills/documentation/examples.md +0 -31
@@ -0,0 +1,295 @@
1
+ # spec-driven-eval — Report template, calibration anchors & worked example
2
+
3
+ ## 1. Report template
4
+
5
+ Copy and fill. Replace bracketed text; keep the section order. `I` and `T` are **derived** from the MET/total counts, never typed in directly.
6
+
7
+ ```markdown
8
+ # Evaluation — [Priority]: [Story title]
9
+
10
+ **Feature**: [feature name]
11
+ **Source of truth**: [PRD ref] (cross-ref [spec.md IDs] if present)
12
+ **AC baseline**: [_ac-baseline.md ref — frozen checklist used]
13
+ **Judge model**: [model] (author model: [model] — flag if same)
14
+ **Module / paths**: [where the code lives]
15
+
16
+ ## Acceptance criteria
17
+ - AC1 — [restate the criterion]
18
+ - ACn — ...
19
+
20
+ ## Implementation checklist (binary — MET/UNMET, evidence required for MET)
21
+ | AC | I-check (atomic, observable) | Verdict | Evidence (file:line) |
22
+ | --- | --- | --- | --- |
23
+ | AC1 | I1. [behavior 1] | MET | `path:line` |
24
+ | AC1 | I2. [behavior 2] | UNMET | searched: [terms/files], absent |
25
+ | ... | ... | ... | ... |
26
+
27
+ Per-AC: **I = MET / total** → AC1 = x/y = X.XX
28
+
29
+ ---
30
+ # FRAMEWORK — extract & respect
31
+
32
+ ## Elicitation E — category rubric (recall)
33
+ | # | Category | Verdict | Evidence (spec.md:line) / why N/A |
34
+ | --- | --- | --- | --- |
35
+ | 1 | Input validation & bounds | Addressed/Missed/N-A | `spec.md:line` |
36
+ | ... | ... | ... | ... |
37
+
38
+ **E_recall = Addressed / (Addressed + Missed) = X.XX**
39
+
40
+ ## Elicitation E — added-requirement ledger (precision + justification)
41
+ | # | Requirement added beyond PRD | Verdict | Built? | Justified? | Evidence (spec.md:line) + warrant |
42
+ | --- | --- | --- | --- | --- | --- |
43
+ | A1 | [requirement] | Valid-necessary | built | yes | `spec.md:line` — [why] |
44
+ | A2 | [requirement] | Invalid (creep/hallucination) | — | no | `spec.md:line` — [why] |
45
+
46
+ **E_precision = valid / total = X.XX** · **E_justified = justified / total = X.XX**
47
+ `valid E-additions` set (used by S + harness denominator): [A1, …]
48
+
49
+ ## Scope S — traceability of built behavior
50
+ | Built behavior | Traces to | Verdict | Evidence (file:line) |
51
+ | --- | --- | --- | --- |
52
+ | [behavior] | PRD AC1 / valid add A1 / none | pass / fail (rogue) / fail (PRD out-of-scope) | `path:line` |
53
+ | [planned, not built] | spec/tasks | partial (plan drift) | `spec.md:line` |
54
+
55
+ **S = pass / partial / fail** (deferred-valid out-of-scope additions are NOT penalized)
56
+
57
+ ---
58
+ # HARNESS — ensure all implemented
59
+
60
+ ## Test checklist (binary — over the sanctioned set = PRD ACs ∪ valid E-additions)
61
+ | Requirement | Source | Level | T-check | Verdict | Evidence (file:line) |
62
+ | --- | --- | --- | --- | --- | --- |
63
+ | AC1 | PRD | unit | [primary behavior asserted] | MET | `path:line` |
64
+ | AC1 | PRD | e2e | [observable contract asserted] | MET | `path:line` |
65
+ | A1 | valid add | unit | [extracted requirement asserted] | UNMET | searched, no test |
66
+ | ... | ... | ... | ... | ... | ... |
67
+
68
+ Per-requirement: **T = MET / total** → AC1 = x/y = X.XX · harness completeness = MET / |sanctioned set|
69
+
70
+ ## Extra tests (Robustness — not scored toward ACs)
71
+ | # | Extra test | Level | Evidence | Value (High/Med/Low) |
72
+ | --- | --- | --- | --- | --- |
73
+
74
+ ## Test distribution by tier (D — reported, not scored)
75
+ | Tier | Count | % | Evidence (representative) |
76
+ | --- | --- | --- | --- |
77
+ | Necessary (P0 primary happy path) | n | xx% | `path:line` |
78
+ | Secondary (important) | n | xx% | `path:line` |
79
+ | Nice-to-have | n | xx% | `path:line` |
80
+ | **Total feature tests** | N | 100% | — |
81
+
82
+ **Shape**: [one line — e.g. "balanced", "top-light / robustness-heavy", "fragile / no defensive tests"]. (Pre-existing tests excluded: [list/none].)
83
+
84
+ ## Result
85
+ | AC | I (MET/total) | T (MET/total) | AC_score = 0.6·I + 0.4·T |
86
+ | --- | --- | --- | --- |
87
+
88
+ | Dimension | Subject | Value |
89
+ | --- | --- | --- |
90
+ | Story_score / Final (PRD fidelity) | framework+harness | X.XX |
91
+ | Elicitation E (recall / precision / justified) | framework | X.XX / X.XX / X.XX |
92
+ | Scope Adherence S | framework | pass/partial/fail |
93
+ | Harness completeness (T over sanctioned set) | harness | X.XX |
94
+ | Engineering Gates G | harness | build/lint/unit/e2e: ✓/✗/not-run |
95
+ | Robustness Index R | harness | [sum] |
96
+ | Test Distribution D | harness | Necessary xx% / Secondary xx% / Nice-to-have xx% (N tests) |
97
+ | Adjusted Final (only if a gate is ✗) | — | Final × 0.5 |
98
+ | k=3 disagreements | — | [checks where the 3 passes split, or none] |
99
+
100
+ **Verdict**: [band + one line]. **Framework**: respects + extracts requirements [read from Final-impl / E / S]. **Harness**: ensures implementation [read from T / G].
101
+
102
+ ## Gaps (ranked) and fixes to reach 1.00
103
+ 1. [AC] — [UNMET check] → [fix]
104
+ ```
105
+
106
+ For a **whole-PRD** roll-up, add a final section:
107
+
108
+ ```markdown
109
+ ## PRD final grade
110
+ | Story | Priority | Weight | Story_score |
111
+ | --- | --- | --- | --- |
112
+ | [P0 a] | P0 | 3 | X.XX |
113
+ | [P1 b] | P1 | 2 | X.XX |
114
+ | [P2 c] | P2 | 0 (excluded) | — |
115
+
116
+ **Final = Σ(w·Story)/Σ(w) = X.XX → [band]**
117
+ Roadmap readiness (P2, informational): [notes]
118
+
119
+ | Whole-PRD reported metrics | Subject | Value |
120
+ | --- | --- | --- |
121
+ | Elicitation E (recall / precision / justified) | framework | X.XX / X.XX / X.XX |
122
+ | Scope Adherence S | framework | pass/partial/fail |
123
+ | Harness completeness (T over sanctioned set) | harness | X.XX |
124
+ | Engineering Gates G | harness | build/lint/unit/e2e: ✓/✗/not-run |
125
+ | Robustness Index R | harness | [sum] |
126
+ | Test Distribution D | harness | Necessary xx% / Secondary xx% / Nice-to-have xx% (N tests) |
127
+ ```
128
+
129
+ ---
130
+
131
+ ## 2. Calibration anchors (read before scoring — they fix the MET/UNMET boundary)
132
+
133
+ These are reference verdicts. Match your verdicts to the reasoning style, not just the outcome. Add new borderline cases here whenever two evaluators split on a check.
134
+
135
+ | Anchor | Check | Verdict | Why |
136
+ | --- | --- | --- | --- |
137
+ | **Clearly MET** | "Creates a trialing subscription without a payment method" | **MET** | `stripe.client.ts:59-80` builds the subscription params with no `payment_method` field and `trial_period_days` set; traced end-to-end from `subscription.service.ts:63-106`. Behavior is present and observable. |
138
+ | **Clearly UNMET** | "409 conflict response includes the existing subscription" | **UNMET** | `subscription.service.ts:79-85` throws `ConflictDomainException` and `subscription.controller.ts:50-52` maps it to 409, but no existing-subscription payload is attached. Searched the exception body + controller mapping; the data is genuinely absent → UNMET (not partial). |
139
+ | **Borderline → resolved UNMET** | "Default of 14 trial days is asserted in e2e" | **UNMET** | `subscription-trial.e2e-spec.ts:157-163` *exercises* the default path (omits `trialDays`) but never asserts the value 14 reached Stripe. Exercised-not-asserted does **not** meet a verification check (Core rule 3). Had it asserted the propagated value, it would be MET. |
140
+ | **Borderline → resolved MET** | "Trialing status grants immediate access" | **MET** | `subscription-state-machine.service.ts:28-32,60-62` includes `Trialing ∈ ACCESS_GRANTED_STATUSES`, traced from the success path. The behavior is implemented even though a dedicated access-endpoint e2e is thin — the *implementation* check is MET; the *test* check for it is scored separately. |
141
+ | **Borderline → resolved UNMET** (Conjunction rule) | "Trigger payload contains trial end date" | **UNMET** | The AC said "emit a trigger associated with the user **and** the trial end date". The trigger *is* emitted with `userId`, but `stripe-webhook.service.ts:53-55` constructs `payload: { stripeSubscriptionId }` only — `trialEndsAt` is available on `localSubscription` (fetched line 46) but never written into the payload object. Per the Conjunction rule each named field after "and" is its own check; the field's absence **at the construction site** is UNMET regardless of whether the parent `emit(...)` is present. (This is the trap Core rule 3's data-shape clause exists to catch.) |
142
+ | **Borderline → resolved UNMET** (Disjunction rule, configurable) | "Cancel behavior is product-configurable" | **UNMET** | The AC said "SHALL apply the **product-chosen** behavior: pause (recommended) or cancel". `stripe.client.ts:48-50` hard-codes `missing_payment_method: 'pause'`; there is no config key, flag, or env var that switches it to `'cancel'`. "Product-chosen" reads as runtime-configurable, so check (2) "alternative reachable without a code change" is UNMET. The `cancel` path via `customer.subscription.deleted` + `TERMINAL_STATUSES` is a *different* behavior (user-initiated cancel), not this product-controlled decision point, so it does not satisfy (2). |
143
+ | **`T-outcome` MET** (outcome-based, entry-point-neutral) | "Inbound status-change event results in the persisted status changing" | **MET** | An integration test invokes the handler against the real DB and asserts the outcome — e.g. `await processWebhookEvent(customerSubscriptionUpdated)` then `expect((await subscriptionRepo.findBy({ stripeSubscriptionId })).status).toBe('paused')`. The real persisted row is asserted. Entry is via the handler directly (not HTTP) — that is fine: `T-outcome` is entry-point-neutral, so an ack-fast async design is not required to be driven through HTTP→queue→worker. |
144
+ | **`T-outcome` UNMET (mock-only)** | "Inbound cancellation event updates the displayed status to `cancelado`" | **UNMET** | The test asserts `expect(subscriptionRepo.save).toHaveBeenCalledWith(objectContaining({ status: 'canceled' }))` / `expect(stateMachine.transition).toHaveBeenCalled()` on a mocked repo/state-machine. That proves a *call*, not a *persisted outcome* — mock-only does not satisfy a "results in / displayed status" check (it counts toward `R` only). **Scoping note:** this exclusion does NOT fire on an "invokes the external API" check (e.g. "immediate cancel calls `StripeClient.cancel`"), which a spy/mock verifies correctly because the asserted proposition *is* the call. |
145
+ | **`T-outcome` UNMET (ingress-only)** | "Inbound event results in the persisted status changing" | **UNMET** | The test asserts the inbound event row was persisted and a processing job was enqueued (`expect(eventRepo.findBy(...)).toBeDefined()` + `expect(queue.add).toHaveBeenCalled()`) but never asserts the *resulting subscription status*. Ingress capture + enqueue is not the outcome — the resulting state is never asserted, so the `T-outcome` check is UNMET (the enqueue assertion may count toward `R`). |
146
+ | **Wiring I-check MET** (async ingress) | "Webhook endpoint receives, verifies, and dispatches the event by type to the handler" | **MET** | `stripe-webhook.controller.ts:22-40` constructs/verifies the event signature (`stripe.webhooks.constructEvent(...)`) and routes by `event.type` (`customer.subscription.updated` → `handleSubscriptionUpdated`). This is the dedicated wiring check — distinct from the `T-outcome` test that the handler produces the right state. A correct handler behind an unregistered/dead route would fail *this* check even with `T-outcome` green; here the route is present and dispatches, so MET. |
147
+
148
+ **Boundary rule the anchors encode:** MET requires implemented-and-evidenced (impl checks) or asserted-not-merely-exercised (test checks). For multi-field artifacts, each named field is checked against the **constructed payload object**, not the call site; for product-chosen alternatives, the non-default path must be reachable without a code change. For **persistence/async (`T-outcome`) checks**, MET requires asserting the **real resulting state** (persisted DB row / returned payload) via **any** entry point (HTTP or handler/consumer/service invoked directly against real infra) — a mock-only call assertion and an ingress-only (persist-event + enqueue) assertion are both UNMET, while an "invokes the external API" check is still correctly met by a spy/mock; the required level is a **floor**, so a stronger test satisfies a weaker-level proposition. An asynchronously-delivered effect additionally needs a **wiring I-check** (endpoint receives + verifies + dispatches by type) so a correct handler behind a dead route does not earn full credit. Anything else is UNMET. There is no middle verdict per check — partial credit emerges only from the MET/total fraction.
149
+
150
+ ---
151
+
152
+ ## 3. Worked example — billing service P0: Start Free Trial Without a Card
153
+
154
+ Real evaluation rebuilt on the binary checklist. Use it as the bar for rigor and evidence. The example happens to be a NestJS + Stripe billing service (TypeORM, Jest) with concrete `file:line` evidence — the methodology itself is stack-agnostic; treat the file paths, frameworks, and commands below as illustrative of the level of rigor, not as required tooling.
155
+
156
+ ### Acceptance criteria (from PRD, aligned with `spec.md` STRIPE-01..05)
157
+ - AC1 — Auth user starts trial with `planId` (+ optional `trialDays`); create Stripe customer if none + **trialing** sub **without payment method**; persist; return `Trialing` + end date.
158
+ - AC2 — No `trialDays` ⇒ default 14.
159
+ - AC3 — Existing `Trialing`/`Active` for same plan ⇒ reject duplicate **and inform the existing subscription** (409).
160
+ - AC4 — Success ⇒ immediate access (state machine grants `Trialing`).
161
+ - AC5 — Transient failure (Stripe/identity) ⇒ clear error, no inconsistent state, idempotent retry (no dup).
162
+
163
+ ### Implementation checklist
164
+ | AC | I-check | Verdict | Evidence |
165
+ | --- | --- | --- | --- |
166
+ | AC1 | I1. Resolve `planId`→`stripePriceId` | MET | `subscription.service.ts:63-106` |
167
+ | AC1 | I2. Create Stripe customer if none | MET | `customer-stripe.service.ts:16-38` |
168
+ | AC1 | I3. Create trialing sub **without** payment method | MET | `stripe.client.ts:59-80` |
169
+ | AC1 | I4. Persist subscription | MET | `subscription.entity.ts:61-62`; service `:97-104` |
170
+ | AC1 | I5. Return `Trialing` + end date | MET | `subscription.controller.ts:35-58` |
171
+ | AC2 | I1. Default 14 when `trialDays` omitted | MET | `subscription.service.ts:27,71` (`DEFAULT_TRIAL_DAYS=14`) |
172
+ | AC3 | I1. Reject duplicate active/trialing for plan (409) | MET | `subscription.service.ts:79-85`; `subscription.controller.ts:50-52` |
173
+ | AC3 | I2. Conflict response **informs** the existing subscription | UNMET | searched exception body + controller map; no existing-sub payload |
174
+ | AC4 | I1. `Trialing` ∈ access-granted statuses | MET | `subscription-state-machine.service.ts:28-32,60-62` |
175
+ | AC5 | I1. Atomic/transactional write (no partial state) | MET | `@Transactional` `subscription.service.ts:63` |
176
+ | AC5 | I2. Idempotent customer + sub creation (no dup) | MET | `stripe.client.ts:40-57`, idempotencyKey `:88-95`; unique `stripeSubscriptionId` `subscription.entity.ts:61-62` |
177
+ | AC5 | I3. Clear domain error on transient failure | UNMET | bare 500 surfaced; no mapped stable error code/message |
178
+
179
+ **I per AC:** AC1 5/5=1.00 · AC2 1/1=1.00 · AC3 1/2=0.50 · AC4 1/1=1.00 · AC5 2/3=0.67
180
+
181
+ ### Elicitation E (framework — extract)
182
+
183
+ **Category rubric (recall):**
184
+ | # | Category | Verdict | Evidence / why |
185
+ | --- | --- | --- | --- |
186
+ | 1 | Input validation & bounds | Addressed | `trialDays` 1–30 bound in `spec.md` (→ impl `subscription.service.ts`, e2e E1) |
187
+ | 2 | Error taxonomy & messaging | Missed | transient failure surfaces a bare 500; no typed error spec'd (AC5 I3 UNMET) |
188
+ | 3 | AuthN / AuthZ | Addressed | authenticated-user requirement carried into spec/controller |
189
+ | 4 | Idempotency & dedup | Addressed | idempotency key + unique `stripeSubscriptionId` spec'd |
190
+ | 5 | Concurrency & races | Missed | no guard spec'd for two concurrent trial-starts on same plan |
191
+ | 6 | Data lifecycle & consistency | Addressed | transactional write spec'd |
192
+ | 7 | Observability | Missed | no logging/metrics/trace requirement on the trial-start path |
193
+ | 8 | Limits, pagination & rate | N/A | single-resource create; no list endpoint |
194
+ | 9 | External-dependency failure | Addressed | Stripe SDK retries/timeout config spec'd (E6) |
195
+ | 10 | State-transition integrity | Addressed | status machine guards illegal transitions |
196
+
197
+ **E_recall = Addressed 6 / (6 + 3 Missed) = 0.67**
198
+
199
+ **Added-requirement ledger (precision + justification):**
200
+ | # | Requirement beyond PRD | Verdict | Built? | Justified? | Evidence + warrant |
201
+ | --- | --- | --- | --- | --- | --- |
202
+ | A1 | `trialDays` must be 1–30 | Valid-necessary | built | yes | `spec.md` — PRD implies a sane trial window |
203
+ | A2 | Plan without `stripePriceId` ⇒ 404 | Valid-necessary | built | yes | `spec.md` — required precondition for trial |
204
+ | A3 | Stripe SDK retry/timeout config | Valid-defensive | built | yes | hardening of external dep |
205
+ | A4 | Customer-snapshot caching | Valid-defensive | built | partial | perf; rationale thin in spec |
206
+
207
+ **E_precision = 4 valid / 4 = 1.00** · **E_justified = 3 / 4 = 0.75**
208
+ `valid E-additions` = [A1, A2, A3, A4] (all built; none deferred)
209
+
210
+ ### Scope S (framework — respect)
211
+ Every built behavior traces to a PRD AC or a valid `E`-addition (A1–A4); nothing maps to the PRD out-of-scope list (no monetization/cancellation built); spec/tasks items all have corresponding code (no plan drift).
212
+ **S = pass.**
213
+
214
+ ### Test checklist (over the sanctioned set = AC1–AC5 ∪ A1–A4)
215
+ | AC | Level | T-check | Verdict | Evidence |
216
+ | --- | --- | --- | --- | --- |
217
+ | AC1 | unit | Trial create asserted | MET | `subscription.service.spec.ts:114-130`; `customer-stripe.service.spec.ts:51-77` |
218
+ | AC1 | unit | No-payment-method asserted | MET | `stripe.client.spec.ts:149-193` |
219
+ | AC1 | e2e | Trial create returns Trialing+end | MET | `subscription-trial.e2e-spec.ts:129-164` |
220
+ | AC2 | unit | Default 14 asserted | MET | `subscription.service.spec.ts:114-130`; `stripe.client.spec.ts:178-190` |
221
+ | AC2 | e2e | Default 14 propagated to Stripe asserted | UNMET | `:157-163` exercised, value not asserted |
222
+ | AC3 | unit | Duplicate rejected asserted | MET | `subscription.service.spec.ts:151-160` |
223
+ | AC3 | unit | Existing-sub returned in conflict asserted | UNMET | not asserted (impl absent) |
224
+ | AC3 | e2e | 409 on duplicate asserted | MET | `subscription-trial.e2e-spec.ts:166-187` |
225
+ | AC4 | unit | Status `Trialing` asserted | MET | `subscription.service.spec.ts:126` |
226
+ | AC4 | e2e | Access actually granted asserted | UNMET | no `GET .../active` assertion |
227
+ | AC5 | unit | Transient-failure handling asserted | MET | `subscription.service.spec.ts:162-167`; `customer-stripe.service.spec.ts:79-87` |
228
+ | AC5 | e2e | Idempotent retry (twice ⇒ no dup) asserted | UNMET | no retry e2e |
229
+ | A1 (valid add) | e2e | `trialDays` 1–30 bound asserted | MET | `subscription-trial.e2e-spec.ts:189-200` |
230
+ | A2 (valid add) | unit | Plan without `stripePriceId` ⇒ 404 asserted | MET | `subscription.service.spec.ts:142-149` |
231
+ | A3 (valid add) | unit | Stripe SDK retry/timeout config asserted | MET | `stripe.client.spec.ts:262-270` |
232
+ | A4 (valid add) | unit | Customer cache-hit path asserted | MET | `customer-stripe.service.spec.ts:38-49` |
233
+
234
+ **T per AC (PRD):** AC1 3/3=1.00 · AC2 1/2=0.50 · AC3 2/3=0.67 · AC4 1/2=0.50 · AC5 1/2=0.50
235
+ **Harness completeness over sanctioned set:** PRD ACs all have ≥1 MET T-check; valid additions A1–A4 all tested ⇒ no extracted requirement is unverified.
236
+
237
+ > **Reconciliation note.** Because A1–A4 are *valid sanctioned* requirements, their tests (E1, E2, E6, E3 below) are coverage of the sanctioned set — under the two-subject model they reclassify from *Nice-to-have* (Robustness `R`) toward *Secondary*. The `R`/`D` figures below are shown in the legacy PRD-only framing for continuity; in a fresh run scored under the sanctioned set, `R` would shrink and Secondary would rise accordingly. Only genuinely AC-unmappable tests (e.g. E5 snapshot-mapper internals) remain in `R`.
238
+
239
+ ### Extra tests (Robustness)
240
+ | # | Extra test | Level | Evidence | Value |
241
+ | --- | --- | --- | --- | --- |
242
+ | E1 | `trialDays` outside 1–30 ⇒ 400 | e2e | `subscription-trial.e2e-spec.ts:189-200` | High |
243
+ | E2 | Plan without `stripePriceId` ⇒ 404 | unit | `subscription.service.spec.ts:142-149` | High |
244
+ | E3 | Customer cache hit ⇒ no identity/Stripe call | unit | `customer-stripe.service.spec.ts:38-49` | Med |
245
+ | E4 | `ensureCustomer` existing-vs-create | unit | `stripe.client.spec.ts:110-147` | Med |
246
+ | E5 | Snapshot mapper: expanded customer / null period | unit | `stripe.client.spec.ts:25-77` | Med |
247
+ | E6 | Stripe SDK config (retries/timeout) | unit | `stripe.client.spec.ts:262-270` | Low |
248
+
249
+ R = 1.0+1.0+0.5+0.5+0.5+0.25 = **3.75**
250
+
251
+ ### Test distribution by tier (D)
252
+ | Tier | Count | % | Representative evidence |
253
+ | --- | --- | --- | --- |
254
+ | Necessary (P0 primary happy path) | 4 | 27% | trial create unit+e2e `subscription.service.spec.ts:114-130`, `subscription-trial.e2e-spec.ts:129-164`; no-payment-method `stripe.client.spec.ts:149-193`; immediate access status `subscription.service.spec.ts:126` |
255
+ | Secondary (important) | 5 | 33% | default-14 `:114-130`,`:157-163`; duplicate guard `:151-160`,`:166-187`; idempotency/retry `:162-167` |
256
+ | Nice-to-have | 6 | 40% | E1–E6 (robustness inventory above) |
257
+ | **Total feature tests** | 15 | 100% | — |
258
+
259
+ **Shape**: top-light / robustness-heavy — only 27% proves the P0 primary path while 40% is defensive; every P0 path does have ≥1 Necessary test, so acceptable but the happy-path layer is thin. (Pre-existing tests excluded: none.)
260
+
261
+ ### Result
262
+ | AC | I | T | AC_score = 0.6·I + 0.4·T |
263
+ | --- | --- | --- | --- |
264
+ | AC1 | 1.00 | 1.00 | 1.00 |
265
+ | AC2 | 1.00 | 0.50 | 0.80 |
266
+ | AC3 | 0.50 | 0.67 | 0.57 |
267
+ | AC4 | 1.00 | 0.50 | 0.80 |
268
+ | AC5 | 0.67 | 0.50 | 0.60 |
269
+
270
+ | Dimension | Subject | Value |
271
+ | --- | --- | --- |
272
+ | Story_score / Final (PRD fidelity) | framework+harness | (1.00+0.80+0.57+0.80+0.60)/5 = **0.75** |
273
+ | Elicitation E (recall / precision / justified) | framework | 0.67 / 1.00 / 0.75 |
274
+ | Scope Adherence S | framework | pass (all built behavior traceable; no out-of-scope build; no plan drift) |
275
+ | Harness completeness (T over sanctioned set) | harness | all PRD ACs + A1–A4 verified (gaps are assertion-strength, not coverage holes) |
276
+ | Engineering Gates G | harness | run before reporting: the project's pinned build, lint, unit, and e2e commands (run any required DB migration before e2e) → record ✓/✗ each (do not assume) |
277
+ | Robustness Index R | harness | 3.75 (legacy framing — see reconciliation note) |
278
+ | Test Distribution D | harness | Necessary 27% / Secondary 33% / Nice-to-have 40% (15 tests) — top-light, robustness-heavy |
279
+ | k=3 disagreements | — | none (all checks stable across passes) |
280
+
281
+ **Verdict**: Strong (0.75).
282
+ - **Framework — respect + extract:** honors the PRD and stays in bounds (`S pass`); extracts cleanly (`E_precision 1.00`) but with moderate coverage of implicit requirements (`E_recall 0.67` — missed error-taxonomy, concurrency, observability).
283
+ - **Harness — ensure implemented:** every sanctioned requirement is tested, but several assertions are weak (exercised-not-asserted), so `T` trails `I`.
284
+
285
+ ### Gaps (ranked) and fixes
286
+ **Framework (respect/extract):**
287
+ 1. `E_recall` — error-taxonomy missed → spec a typed transient-failure error (also fixes AC5 I3).
288
+ 2. `E_recall` — concurrency missed → spec a guard for two concurrent trial-starts on the same plan.
289
+ 3. `E_recall` — observability missed → spec logging/metrics on the trial-start path.
290
+ 4. AC3 I2 (UNMET) — 409 doesn't inform existing sub → return the existing subscription in the conflict payload.
291
+
292
+ **Harness (ensure implemented):**
293
+ 5. AC5 I3 / e2e (UNMET) — bare 500 + no idempotent-retry e2e → map to a clear domain error and add a retry e2e (twice ⇒ no dup).
294
+ 6. AC4 e2e (UNMET) — no access assertion → e2e hitting `GET /subscription/user/:id/active` proving `Trialing` grants access.
295
+ 7. AC2 e2e (UNMET) — assert default 14 propagated to Stripe when `trialDays` omitted.
@@ -0,0 +1,47 @@
1
+ ---
2
+ name: taste
3
+ description: One skill for the full taste spectrum. Covers anti-slop frontend design (landing pages, portfolios, redesigns, premium and minimalist UI, GSAP motion), image generation and image-to-code direction (web, mobile, brand kits), full-output enforcement, and Google Stitch design systems. Read the brief, pick the matching reference below, and follow it end to end. Do not auto-apply every reference at once.
4
+ ---
5
+
6
+ # TASTE — Unified Design Quality Skill
7
+
8
+ You are an elite frontend design engineer and art director. This skill bundles every taste discipline into one entry point. Each reference below is a complete playbook for one domain.
9
+
10
+ ## How to use
11
+
12
+ 1. **Read the brief first.** Do not fire all references at once.
13
+ 2. **Route to the matching reference** using the table below.
14
+ 3. **Follow that reference end to end** — its directives override generic defaults.
15
+ 4. When a task spans domains (e.g. redesign an existing site that also needs new imagery), apply the relevant references in sequence and reconcile conflicting rules by task priority.
16
+
17
+ ## Reference library
18
+
19
+ | When the task is… | Use |
20
+ |---|---|
21
+ | Anti-slop frontend for landing pages, portfolios, redesigns (current default) | [design-taste-frontend.md](references/design-taste-frontend.md) |
22
+ | Exact backward-compatible v1 taste behavior (original baseline) | [design-taste-frontend-v1.md](references/design-taste-frontend-v1.md) |
23
+ | Awwwards-level design engineering with GSAP motion, randomization, bento grids | [gpt-taste.md](references/gpt-taste.md) |
24
+ | High-end agency visual design — fonts, spacing, shadows, motion choreography | [high-end-visual-design.md](references/high-end-visual-design.md) |
25
+ | Premium utilitarian minimalism — warm monochrome, editorial type, flat bento | [minimalist-ui.md](references/minimalist-ui.md) |
26
+ | Industrial brutalism / tactical telemetry UI — Swiss print, CRT terminals | [industrial-brutalist-ui.md](references/industrial-brutalist-ui.md) |
27
+ | Upgrading an existing project to premium quality (audit-first, no rewrites) | [redesign-existing-projects.md](references/redesign-existing-projects.md) |
28
+ | Completeness of output — no truncation, no placeholders, no skipped sections | [full-output-enforcement.md](references/full-output-enforcement.md) |
29
+ | Generating website section reference images, then implementing them | [image-to-code.md](references/image-to-code.md) |
30
+ | Generating website design-reference images (one image per section) | [imagegen-frontend-web.md](references/imagegen-frontend-web.md) |
31
+ | Generating mobile app screen concepts and flows (images only) | [imagegen-frontend-mobile.md](references/imagegen-frontend-mobile.md) |
32
+ | Generating premium brand-kit / identity-system boards (images only) | [brandkit.md](references/brandkit.md) |
33
+ | Generating a semantic DESIGN.md for Google Stitch | [stitch-design-taste.md](references/stitch-design-taste.md) |
34
+ | Reference DESIGN.md output produced by the Stitch skill | [stitch-design-taste-design.md](references/stitch-design-taste-design.md) |
35
+
36
+ ## Cross-cutting rules
37
+
38
+ Every reference shares these baseline principles; specific rules in a reference override these:
39
+
40
+ - Never output the same layout or aesthetic twice in a row.
41
+ - Reject the LLM defaults: Inter, Lucide/Feather icons, purple-blue AI gradients, three equal feature cards, centered hero over dark mesh, generic glassmorphism, `h-screen` (use `min-h-[100dvh]`).
42
+ - Animate only `transform` and `opacity`. No layout-triggering properties.
43
+ - No emojis in code, markup, text content, or alt text.
44
+ - No pure black (`#000000`) — use off-black, charcoal, or zinc-950.
45
+ - Prefer distinctive type (Geist, Outfit, Cabinet Grotesk, Satoshi, Fraunces, Instrument Serif) over generic fonts.
46
+ - Verify dependency availability before importing third-party libraries.
47
+ - Generate the `<design_plan>` / design-read step before writing any code where the reference requires it.
@@ -0,0 +1,73 @@
1
+ # Taste Skill — Example Outputs
2
+
3
+ Realistic examples of how the `taste` skill routes a brief to the right reference and enforces its standards.
4
+
5
+ ## Example 1: Landing Page (routes to `design-taste-frontend`)
6
+
7
+ **User:** "Build a landing page for a developer analytics platform. Clean, premium, not AI-sloppy."
8
+
9
+ **Design read:** B2B SaaS landing for technical buyers, with a Linear-style minimalist language, leaning toward Tailwind utilities + Geist + restrained motion.
10
+
11
+ **What the agent does (following `design-taste-frontend.md`):**
12
+
13
+ 1. Sets the dials from the read: VARIANCE 5, MOTION 4, DENSITY 3.
14
+ 2. States the design read in one line before writing any code.
15
+ 3. Chooses a non-default hero: centered over a wide headline with two high-contrast CTAs and a full-bleed background image with a dark radial wash — not a narrow 6-line text wall.
16
+ 4. Bans Inter and Lucide; uses Geist + Phosphor. No AI-purple gradient, no three equal feature cards — uses an asymmetric bento with `grid-flow-dense`.
17
+ 5. Applies `min-h-[100dvh]` (never `h-screen`), animates only `transform`/`opacity`, generous `py-32 md:py-48` section spacing.
18
+
19
+ **Result:** A distinctive, premium landing that does not read as a template.
20
+
21
+ ## Example 2: Brand Kit (routes to `brandkit`)
22
+
23
+ **User:** "Generate a brand-kit board for 'Aster', a security platform for cloud-native startups."
24
+
25
+ **What the agent does (following `brandkit.md`):**
26
+
27
+ 1. Infers brand strategy first: category security, audience CTOs/security engineers, core metaphor "vigilance + protection".
28
+ 2. Picks the Dark Security / Threat Intelligence visual mode: black/navy canvas, shield forms, radar lines, red/blue alert chips.
29
+ 3. Chooses a `3 × 3` layout with strong gutters and sparse typography: logo cover, logo construction grid, browser mockup, tagline, color system, typography, physical application, image direction, system detail.
30
+ 4. Designs one strong mark — an eye formed by crossing shield shapes — repeated consistently across panels. No generic lightning bolts, no fake luxury crests.
31
+ 5. Writes a short, specific tagline ("Your network. Our watch."), one URL, and minimal labels. No lorem ipsum, no dense body copy.
32
+
33
+ **Result:** A premium, presentation-ready identity-system board that feels like a serious studio deck.
34
+
35
+ ## Example 3: Mobile App Screens (routes to `imagegen-frontend-mobile`)
36
+
37
+ **User:** "Show me the onboarding flow for a habit-tracking app — three screens."
38
+
39
+ **What the agent does (following `imagegen-frontend-mobile.md`):**
40
+
41
+ 1. Generates three separate app-native screens (welcome, permissions, home) inside a subtle premium phone mockup with a visible frame.
42
+ 2. Keeps hierarchy clean and text comfortably readable — no tiny fake labels, no dense dashboards.
43
+ 3. Uses one controlled palette with textured surfaces and image-led composition.
44
+ 4. Keeps the three screens visually consistent (shared color, type scale, iconography).
45
+ 5. Produces images only — no code.
46
+
47
+ **Result:** Three cohesive, premium, app-native concepts a designer can hand to an engineer.
48
+
49
+ ## Example 4: Redesign (routes to `redesign-existing-projects`)
50
+
51
+ **User:** "The dashboard we built last month looks generic. Make it premium without breaking it."
52
+
53
+ **What the agent does (following `redesign-existing-projects.md`):**
54
+
55
+ 1. **Scan** — identifies the stack (React + Tailwind v4) and current patterns.
56
+ 2. **Diagnose** — audits: Inter everywhere, three equal card columns, pure black background, purple-blue gradient hero, no hover/active states, no loading or empty states, generic Lucide icons.
57
+ 3. **Fix in priority order** — font swap first (Geist), then palette cleanup (off-black + one desaturated accent), then hover/active states, layout, and finally component and icon swaps.
58
+ 4. Works with the existing stack — no framework migration, no rewrites. Tests after each change.
59
+
60
+ **Result:** The same app, elevated — no broken functionality, no templated look.
61
+
62
+ ## Example 5: Full Output (routes to `full-output-enforcement`)
63
+
64
+ **User:** "Write the complete `auth.ts` service — every function, no shortcuts."
65
+
66
+ **What the agent does (following `full-output-enforcement.md`):**
67
+
68
+ 1. Counts the deliverables (all auth functions + exports) and locks the scope.
69
+ 2. Emits the full file — no `// rest of code`, no `// TODO`, no `...` placeholders, no "the rest follows the same pattern."
70
+ 3. If the response approaches the token limit, stops at a clean breakpoint and marks it: `[PAUSED — 2 of 4 complete. Send "continue" to resume from: token refresh]`.
71
+ 4. Cross-checks output against the original request before finalizing.
72
+
73
+ **Result:** A complete, runnable file with zero placeholder patterns.