mustflow 2.116.4 → 2.117.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. package/package.json +1 -1
  2. package/templates/default/i18n.toml +145 -7
  3. package/templates/default/locales/en/.mustflow/skills/INDEX.md +121 -11
  4. package/templates/default/locales/en/.mustflow/skills/agent-eval-integrity-review/SKILL.md +62 -34
  5. package/templates/default/locales/en/.mustflow/skills/agent-execution-control-review/SKILL.md +102 -31
  6. package/templates/default/locales/en/.mustflow/skills/agent-memory-context-governance-review/SKILL.md +163 -0
  7. package/templates/default/locales/en/.mustflow/skills/agent-planning-recovery-review/SKILL.md +180 -0
  8. package/templates/default/locales/en/.mustflow/skills/agent-release-bundle-rollout-review/SKILL.md +181 -0
  9. package/templates/default/locales/en/.mustflow/skills/agent-runtime-isolation-review/SKILL.md +196 -0
  10. package/templates/default/locales/en/.mustflow/skills/agent-runtime-multi-worker-review/SKILL.md +180 -0
  11. package/templates/default/locales/en/.mustflow/skills/automation-investment-case-review/SKILL.md +173 -0
  12. package/templates/default/locales/en/.mustflow/skills/client-platform-strategy-review/SKILL.md +236 -0
  13. package/templates/default/locales/en/.mustflow/skills/credit-ledger-integrity-review/SKILL.md +8 -4
  14. package/templates/default/locales/en/.mustflow/skills/credit-monetization-integrity-review/SKILL.md +283 -0
  15. package/templates/default/locales/en/.mustflow/skills/desktop-commercial-distribution-review/SKILL.md +225 -0
  16. package/templates/default/locales/en/.mustflow/skills/external-prompt-injection-defense/SKILL.md +49 -3
  17. package/templates/default/locales/en/.mustflow/skills/freemium-ad-monetization-review/SKILL.md +196 -0
  18. package/templates/default/locales/en/.mustflow/skills/game-economy-monetization-review/SKILL.md +208 -0
  19. package/templates/default/locales/en/.mustflow/skills/game-liveops-commerce-integrity-review/SKILL.md +237 -0
  20. package/templates/default/locales/en/.mustflow/skills/growth-distribution-integrity-review/SKILL.md +247 -0
  21. package/templates/default/locales/en/.mustflow/skills/idempotency-integrity-review/SKILL.md +20 -2
  22. package/templates/default/locales/en/.mustflow/skills/llm-model-routing-integrity-review/SKILL.md +183 -0
  23. package/templates/default/locales/en/.mustflow/skills/llm-product-monetization-review/SKILL.md +311 -0
  24. package/templates/default/locales/en/.mustflow/skills/llm-token-cost-control-review/SKILL.md +15 -1
  25. package/templates/default/locales/en/.mustflow/skills/localization-market-expansion-review/SKILL.md +224 -0
  26. package/templates/default/locales/en/.mustflow/skills/multi-agent-work-coordination/SKILL.md +7 -2
  27. package/templates/default/locales/en/.mustflow/skills/pricing-model-integrity-review/SKILL.md +288 -0
  28. package/templates/default/locales/en/.mustflow/skills/product-engagement-retention-review/SKILL.md +234 -0
  29. package/templates/default/locales/en/.mustflow/skills/product-onboarding-activation-review/SKILL.md +269 -0
  30. package/templates/default/locales/en/.mustflow/skills/product-portfolio-integrity-review/SKILL.md +233 -0
  31. package/templates/default/locales/en/.mustflow/skills/prompt-contract-quality-review/SKILL.md +1 -0
  32. package/templates/default/locales/en/.mustflow/skills/referral-incentive-integrity-review/SKILL.md +206 -0
  33. package/templates/default/locales/en/.mustflow/skills/retry-policy-integrity-review/SKILL.md +45 -2
  34. package/templates/default/locales/en/.mustflow/skills/routes.toml +329 -7
  35. package/templates/default/locales/en/.mustflow/skills/service-portfolio-capital-allocation-review/SKILL.md +245 -0
  36. package/templates/default/locales/en/.mustflow/skills/subscription-retention-profit-review/SKILL.md +217 -0
  37. package/templates/default/manifest.toml +86 -1
@@ -0,0 +1,247 @@
1
+ ---
2
+ mustflow_doc: skill.growth-distribution-integrity-review
3
+ locale: en
4
+ canonical: true
5
+ revision: 1
6
+ lifecycle: mustflow-owned
7
+ authority: procedure
8
+ name: growth-distribution-integrity-review
9
+ description: Apply this skill when a product changes free-result watermarking or embedded attribution, public-share branding, affiliate or influencer compensation, partner attribution and clawbacks, recurring or lifetime commission, owned-product cross-promotion, portfolio promotion frequency, brand relationship disclosure, incremental acquisition, channel cannibalization, or retained portfolio contribution and must grow distribution without degrading the result users share, paying for natural demand, confusing product identity, or damaging the source product.
10
+ metadata:
11
+ mustflow_schema: "1"
12
+ mustflow_kind: procedure
13
+ pack_id: mustflow.core
14
+ skill_id: mustflow.core.growth-distribution-integrity-review
15
+ command_intents:
16
+ - changes_status
17
+ - changes_diff_summary
18
+ - lint
19
+ - build
20
+ - test_related
21
+ - test
22
+ - docs_validate_fast
23
+ - test_release
24
+ - mustflow_check
25
+ ---
26
+
27
+ # Growth Distribution Integrity Review
28
+
29
+ <!-- mustflow-section: purpose -->
30
+ ## Purpose
31
+
32
+ Review embedded result attribution, commercial partners, and owned-product cross-promotion as
33
+ different distribution mechanisms with one economic standard: incremental retained contribution
34
+ after source-product harm, partner cost, fraud, support, privacy, and brand confusion. Do not turn
35
+ useful output into an ad, grant perpetual commission for nonincremental demand, or call internal
36
+ inventory free when it damages the product that owns the user relationship.
37
+
38
+ <!-- mustflow-section: use-when -->
39
+ ## Use When
40
+
41
+ - A free result adds, removes, resizes, relocates, or gates a visible service name, watermark, footer,
42
+ end card, attribution link, metadata mark, removal entitlement, or branded export.
43
+ - A product recruits influencers, creators, affiliates, comparison publishers, agencies, resellers,
44
+ integration partners, or other commercial acquisition partners and changes fixed fees, CPA,
45
+ first-payment share, recurring commission, lifetime commission, attribution, or clawbacks.
46
+ - One owned service recommends another through a dashboard, success screen, result surface, account
47
+ hub, email, notification, modal, interstitial, or contextual next step.
48
+ - A report compares organic acquisition, affiliate CAC, external paid acquisition, cross-promotion,
49
+ brand lift, portfolio revenue, fatigue, cannibalization, or retained contribution.
50
+
51
+ <!-- mustflow-section: do-not-use-when -->
52
+ ## Do Not Use When
53
+
54
+ - The task is a customer-to-customer referral reward, invite code, dual-sided incentive, reward tier,
55
+ valid-referral event, or self-referral control; use
56
+ `referral-incentive-integrity-review`.
57
+ - The task is a paid ad or rewarded-ad placement inside a free tier, result gate, ad SDK, premium ad
58
+ removal, or ad-funded access decision; use `freemium-ad-monetization-review`.
59
+ - The task only changes trademark ownership, generic visual identity, SEO, ad creative, marketing
60
+ copy, public-relations messaging, or analytics implementation with no distribution policy; use the
61
+ narrower brand, content, attribution, privacy, security, or data procedure.
62
+ - The task requests jurisdiction-specific affiliate, advertising, tax, contract, privacy, or AI
63
+ marking advice. Use current qualified authority; this skill supplies the product evidence packet.
64
+
65
+ <!-- mustflow-section: required-inputs -->
66
+ ## Required Inputs
67
+
68
+ - Artifact ledger: result type, ownership, public or private use, export path, share rate, recipient
69
+ reach, attribution visibility and persistence, placement, removal, editability, accessibility,
70
+ quality, privacy, professional-use risk, and AI-origin or provenance requirement.
71
+ - Partner ledger: partner identity and type, audience, channel, content lifetime, claimed reach,
72
+ attribution window and priority, click, signup, qualification, payment, refund, chargeback,
73
+ retained revenue, commission, cap, duration, termination, clawback, disclosure, approval, brand
74
+ safety, prohibited claim, support, tax, contract, and jurisdiction review.
75
+ - Cross-promotion ledger: source and target product, user and job adjacency, product relationship,
76
+ placement, success state, eligibility, exposure sequence, frequency, suppression, account and data
77
+ continuity, consent, external-acquisition alternative, and source-product opportunity cost.
78
+ - Causal ledger: pre-exposure assignment, persistent holdout, natural discovery, existing intent,
79
+ channel overlap, incremental activation and payment, retained contribution, cannibalization,
80
+ source-product completion and retention, complaints, hides, unsubscribes, and brand-understanding.
81
+ - Economics ledger: partner and creative cost, net revenue, variable service and payment cost,
82
+ refund, chargeback, support, fraud, commission liability, source-product lost value, external CAC,
83
+ portfolio contribution, horizon, and uncertainty.
84
+ - Version ledger: artifact policy, campaign, partner contract, attribution rule, disclosure, source
85
+ and target eligibility, frequency, suppression, experiment, and current authority version.
86
+
87
+ <!-- mustflow-section: preconditions -->
88
+ ## Preconditions
89
+
90
+ - Separate embedded attribution, partner compensation, and owned cross-promotion. A shared growth
91
+ goal does not make their rights, risks, or attribution rules interchangeable.
92
+ - Define the source product's completed value and the incremental target outcome before adding a
93
+ brand mark, partner payment, or cross-promotion exposure.
94
+ - Assign experiments before the mark or promotion is visible where feasible, and preserve eligible
95
+ nonsharers, nonclickers, nonbuyers, and source-product abandoners in the relevant denominator.
96
+ - Treat copied watermark sizes, share rates, commission percentages, attribution windows, lifetime
97
+ terms, exposure counts, fit scores, CAC, and conversion thresholds as hypotheses, not defaults.
98
+ - Refresh current endorsement, advertising, consumer, privacy, accessibility, platform, tax,
99
+ contract, trademark, and AI-origin marking rules for the relevant role, content, channel,
100
+ geography, and date.
101
+ - This skill does not authorize live marks, partner recruitment, contracts, payments, tracking,
102
+ cross-product data sharing, campaigns, experiments, messages, or production changes.
103
+
104
+ <!-- mustflow-section: allowed-edits -->
105
+ ## Allowed Edits
106
+
107
+ - Add or refine artifact branding eligibility, placement and removal, share and recipient events,
108
+ partner qualification and compensation, attribution, disclosure, fraud and clawback controls,
109
+ cross-promotion eligibility and suppression, relationship explanations, experiment assignment,
110
+ contribution metrics, fixtures, tests, docs, route metadata, and synchronized templates.
111
+ - Replace universal watermarking, gross attributed revenue, last-click poaching, uncapped lifetime
112
+ commission, source-product interruption, or raw cross-product clicks with bounded rights and
113
+ causal portfolio economics.
114
+ - Do not insert private identifiers or covert user tracking into exported artifacts, misrepresent
115
+ authorship, suppress required disclosure, transfer account data without authority, or impair the
116
+ result users completed merely to make brand removal a paid feature.
117
+
118
+ <!-- mustflow-section: procedure -->
119
+ ## Procedure
120
+
121
+ 1. Split the review into artifact attribution, partner compensation, and owned cross-promotion.
122
+ Evaluate each independently before comparing their contribution as acquisition channels.
123
+ 2. Define the headline result as incremental retained contribution per eligible source user or
124
+ acquisition opportunity over a declared horizon. Include target revenue and service cost,
125
+ partner cost, refunds, disputes, fraud, support, source-product lost value, and cannibalization.
126
+ 3. Classify the artifact. Estimate whether it is privately stored, internally circulated, sent to a
127
+ small known group, published repeatedly, or distributed to a broad public audience.
128
+ Do not infer reach from export count alone.
129
+ 4. Build the artifact acquisition path from created result to external share, unique recipient,
130
+ visible attribution, attributable visit, activated user, payer, and retained contribution. Keep
131
+ original-user conversion and retention loss in the same calculation.
132
+ 5. Use visible branding only when attribution can survive a real public distribution path and its
133
+ incremental value exceeds quality, sharing, conversion, accessibility, privacy, and support harm.
134
+ A watermark on a private or low-reach artifact is not a growth engine by default.
135
+ 6. Preserve the completed result. Keep branding outside critical content where possible, readable
136
+ but not dominant, accessible, and compatible with professional use. Never obscure evidence,
137
+ damage legibility, imply false authorship, or block export solely to force removal payment.
138
+ 7. Choose the least harmful attribution carrier that can work. A small footer, end card, adjacent
139
+ share link, or approved metadata can be a candidate; visible and machine-readable marks solve
140
+ different problems and neither proves compliance with every current AI-origin rule.
141
+ 8. Do not embed user, recipient, prompt, tenant, or private tracking data in exported files. Define
142
+ whether attribution survives editing, screenshots, transcoding, printing, and platform stripping,
143
+ and measure the actual recipient path rather than assumed virality.
144
+ 9. Classify the partner's continuing work. Short-lived sponsored reach, durable search or tutorial
145
+ content, ongoing implementation, reseller support, and customer-success work justify different
146
+ compensation shapes.
147
+ 10. Compare fixed production cost, qualified CPA, first-payment share, and bounded recurring share at
148
+ equal expected budget and qualification quality. Do not call a higher-spend arm more persuasive
149
+ without separating spend from contract shape.
150
+ 11. Calculate partner economics from incremental customers. Exclude or discount existing demand,
151
+ branded-search interception, leaked coupons, self-referral, duplicate partners, pre-existing
152
+ accounts, refund-window purchases, and customers who would have converted through another owned
153
+ or paid channel.
154
+ 12. Bind compensation to a qualified event and net revenue. Version attribution priority, window,
155
+ product scope, refund and chargeback treatment, tax, currency, minimum payout, cap, termination,
156
+ inactivity, content removal, and clawback before value vests.
157
+ 13. Treat lifetime commission as a long-lived liability, not a free acquisition rate. Use it only
158
+ when the partner continues to create attributable value or service after acquisition and the
159
+ agreement has explicit scope, activity, cap or review, termination, transfer, audit, and margin
160
+ protection. A bounded recurring term is a candidate, not a universal duration.
161
+ 14. Make commercial relationships clear and conspicuous near the endorsement or link as required by
162
+ current authority. Train and monitor partners, preserve content evidence, prohibit unsupported
163
+ claims, and define brand-safety escalation rather than relying only on a platform disclosure UI.
164
+ 15. Prevent attribution abuse with deterministic rules and multiple signals. Cover cookie stuffing,
165
+ forced redirects, toolbar or extension injection, trademark bidding, last-click poaching,
166
+ coupon leakage, fake leads, duplicate identities, self-dealing, and partner collusion while
167
+ preserving review and appeal for legitimate conflicts.
168
+ 16. Classify source-target fit before cross-promotion. Use audience overlap, adjacent user job,
169
+ brand promise, account and data relationship, and target readiness as evidence; ownership by the
170
+ same operator alone does not make two products relevant.
171
+ 17. Deliver the source product's primary value before promotion. Prefer a contextual next step after
172
+ success or in a clearly separate portfolio surface. Do not interrupt first value, active work,
173
+ checkout, cancellation, safety, privacy, or accessibility flows.
174
+ 18. Explain the relationship honestly. State whether the target is another product, an optional
175
+ companion, or part of a suite; do not imply shared entitlements, identity, data, support, or
176
+ warranties that the current contract does not provide.
177
+ 19. Make cross-product data movement explicit. Define consent, purpose, tenant and account boundaries,
178
+ deletion, suppression, and access before transferring user input or state. A click may deep-link
179
+ without silently provisioning or copying private data.
180
+ 20. Derive fatigue from marginal value by exposure sequence. Track incremental activation and target
181
+ contribution against source completion, retention, hides, complaints, unsubscribes, and support
182
+ after each exposure. Stop when marginal portfolio value is nonpositive; do not copy a universal
183
+ session, week, or month cap.
184
+ 21. Suppress irrelevant repetition. Stop or cool down after conversion, purchase, explicit rejection,
185
+ repeated nonresponse, source-product distress, target ineligibility, or relationship mismatch.
186
+ Keep campaign and suppression versions reconstructable across channels.
187
+ 22. Compare owned inventory with external acquisition using opportunity cost. Internal exposure is
188
+ not free: include source-product lost value, displaced product messages, creative and operations,
189
+ and portfolio cannibalization alongside the target's external incremental CAC.
190
+ 23. Use persistent holdouts where feasible to estimate natural sharing, partner incrementality, cross-
191
+ product discovery, and source-product damage. Observed clicks, attributed revenue, or users of two
192
+ products do not prove the channel caused portfolio growth.
193
+ 24. Promote only a reversible policy that improves incremental portfolio contribution while
194
+ preserving useful artifacts, honest commercial disclosure, deterministic attribution, partner
195
+ quality, source-product value, brand clarity, privacy, accessibility, and current authority.
196
+
197
+ <!-- mustflow-section: postconditions -->
198
+ ## Postconditions
199
+
200
+ - Artifact attribution, partner compensation, and owned cross-promotion remain distinct contracts.
201
+ - Visible branding is limited to an evidence-backed distribution path and does not degrade or
202
+ misrepresent the completed result.
203
+ - Partner rewards use deterministic attribution, qualified net value, bounded liability, disclosure,
204
+ fraud controls, clawbacks, and review.
205
+ - Cross-promotion follows source value, explains product relationships, respects account and data
206
+ boundaries, suppresses fatigue, and includes source-product opportunity cost.
207
+ - Headline growth uses causal portfolio contribution rather than exports, impressions, clicks,
208
+ attributed gross revenue, or users observed in multiple products.
209
+
210
+ <!-- mustflow-section: verification -->
211
+ ## Verification
212
+
213
+ Use configured oneshot command intents when available: `changes_status`, `changes_diff_summary`,
214
+ `lint`, `build`, `test_related`, `test`, `docs_validate_fast`, `test_release`, and
215
+ `mustflow_check`. Do not infer live export, affiliate, payment, tracking, campaign, analytics,
216
+ cross-product data, messaging, deployment, or production commands.
217
+
218
+ <!-- mustflow-section: failure-handling -->
219
+ ## Failure Handling
220
+
221
+ - If recipient reach, partner qualification, or source-product damage cannot be measured, keep the
222
+ policy reversible and report association rather than claiming incremental acquisition.
223
+ - If branding reduces utility, accessibility, privacy, sharing, conversion, or professional use more
224
+ than attributable retained contribution, remove or narrow it instead of optimizing removal sales.
225
+ - If partner attribution cannot exclude existing intent, refunds, self-dealing, or channel overlap,
226
+ keep compensation pending or cap it under the current contract; do not grant unbounded liability.
227
+ - If commercial disclosure, claim substantiation, tax, contract, or jurisdiction authority is
228
+ unresolved, do not launch or expand the affected partner treatment.
229
+ - If cross-promotion improves the target while reducing source or portfolio contribution, reject or
230
+ narrow the placement even when click-through and target revenue rise.
231
+ - If account or data relationships are ambiguous, use a separate optional link and do not transfer
232
+ identity or private state.
233
+
234
+ <!-- mustflow-section: output-format -->
235
+ ## Output Format
236
+
237
+ - Artifact type, ownership, distribution path, recipient reach, attribution, placement, removal,
238
+ quality, accessibility, privacy, AI-origin, and contribution decision
239
+ - Partner type, continuing work, qualification, compensation shape, equal-budget comparison,
240
+ attribution, disclosure, cap, duration, termination, clawback, fraud, and incrementality decision
241
+ - Source and target product fit, relationship, placement, success state, account and data boundary,
242
+ frequency, suppression, source-product harm, external-CAC comparison, and portfolio decision
243
+ - Experiment assignment, holdout, horizon, natural demand, cannibalization, retained contribution,
244
+ guardrails, and uncertainty
245
+ - Files changed
246
+ - Command intents run and skipped checks
247
+ - Remaining growth-distribution risk
@@ -2,7 +2,7 @@
2
2
  mustflow_doc: skill.idempotency-integrity-review
3
3
  locale: en
4
4
  canonical: true
5
- revision: 3
5
+ revision: 4
6
6
  lifecycle: mustflow-owned
7
7
  authority: procedure
8
8
  name: idempotency-integrity-review
@@ -60,6 +60,9 @@ deduplicates the operation or rejects the stale result?"
60
60
  ## Required Inputs
61
61
 
62
62
  - Operation identity ledger: logical operation, attempt, effect, actor, tenant, target resource, business operation type, canonical payload hash, idempotency key, event ID, message ID, provider object ID, batch key, scheduler run key, and retry source.
63
+ - Workflow identity ledger when plans can change: `workflow_id`, `plan_id` or plan version,
64
+ `step_id`, stable `effect_id`, per-try `attempt_id`, causation ID, correlation ID, and the rule that
65
+ decides whether a replan represents the same admitted business intent or a new one.
63
66
  - Ordering and authority ledger: aggregate key, ordering scope, observed state, expected state, aggregate version, sequence, generation, fencing token, gap policy, stale-result decision, and authoritative conditional write.
64
67
  - Side-effect ledger: every charge, refund, balance change, stock change, coupon issue, shipment, email, notification, entitlement, status transition, file write, queue publish, cache change, provider call, and audit entry.
65
68
  - Durable dedupe evidence: unique constraints, idempotency table, inbox table, outbox table, applied event table, ledger source key, conditional update, state guard, provider idempotency key, and response record.
@@ -93,7 +96,11 @@ deduplicates the operation or rejects the stale result?"
93
96
  1. Name the logical operation, not only the endpoint.
94
97
  - POST, PATCH, DELETE, webhook, queue consumer, scheduler, batch replay, retry wrapper, and callback handlers can all carry the same duplicate-intent risk.
95
98
  - Ask what "same intent" means in business terms: same order payment attempt, same refund, same point grant, same coupon redemption, same shipment, same notification, same state transition, same imported row, or same provider event.
96
- - Separate the stable operation ID from attempt, message, event, effect, trace, worker, and provider IDs. Generate a new attempt ID for a retry without generating a new business operation.
99
+ - Separate the stable operation ID from attempt, message, event, effect, trace, worker, and provider IDs. Generate a new `attempt_id` for a retry without generating a new business operation.
100
+ - In a replanning workflow, keep `workflow_id` stable for the workflow, `step_id` stable for the
101
+ logical step when appropriate, `effect_id` stable for the same admitted business intent across
102
+ plan versions, and `attempt_id` new for each execution try. A new plan version is not permission
103
+ to create a duplicate effect.
97
104
  2. Start at the side effect.
98
105
  - Find calls and writes named like `charge`, `capture`, `refund`, `withdraw`, `grantPoint`, `deductStock`, `issueCoupon`, `sendEmail`, `createShipment`, `publish`, `ack`, `markPaid`, `fulfill`, and local equivalents.
99
106
  - For each side effect, ask what happens if the exact surrounding handler runs twice, runs concurrently, or runs after the first response is lost.
@@ -102,6 +109,12 @@ deduplicates the operation or rejects the stale result?"
102
109
  - Duplicate key plus changed amount, resource, user, tenant, product, operation, or payload should become a stable mismatch response, not a new operation or silent old response.
103
110
  - Do not accept memory-only stores, process-local maps, or Redis TTL alone for operations that can be retried after restart, failover, delayed callback, or TTL expiry.
104
111
  - Keep the operation key stable across request, durable operation record, inbox or outbox, external effect, result lookup, and response replay.
112
+ - Derive or mint keys at a trusted boundary. Do not place raw personal data, secrets, mutable
113
+ display names, timestamps, or model-generated prose in a key. Prefer opaque random or keyed
114
+ deterministic identifiers whose canonical request fingerprint is stored separately.
115
+ - Retain durable idempotency evidence for at least the longest supported client retry, queue
116
+ redelivery, delayed callback, workflow recovery, manual replay, and business dispute window.
117
+ Do not copy one universal TTL; name the owning lifecycle and expiry consequence.
105
118
  4. Require durable uniqueness at the operation boundary.
106
119
  - App-only `if exists return` followed by `insert` is a race. Use a unique constraint, atomic insert, upsert, conditional write, or durable ledger key.
107
120
  - Review migrations or schema definitions, not just service code, for keys such as `order_id`, `payment_id`, `refund_id`, `event_id`, `message_id`, `idempotency_key`, `campaign_id + user_id`, and `settlement_date + merchant_id`.
@@ -161,6 +174,9 @@ deduplicates the operation or rejects the stale result?"
161
174
  21. Review `PROCESSING` and lease recovery.
162
175
  - A `PROCESSING` row without lease, heartbeat, timeout, owner, or reconciliation can permanently block retries after a crash.
163
176
  - A stale processing record should recover by verifying side effects and either completing, failing safely, or allowing one fenced owner to continue.
177
+ - Preserve an explicit `UNKNOWN` or `RECONCILING` state when the external result cannot yet be
178
+ proved. Do not expire the idempotency record or mint a new effect while reconciliation or
179
+ manual review still owns the outcome.
164
180
  22. Review locks as helpers, not proof.
165
181
  - Distributed locks can expire, split ownership, or pause through stop-the-world events.
166
182
  - Keep durable uniqueness, conditional writes, state guards, idempotency records, or fencing tokens as the final defense.
@@ -176,6 +192,8 @@ deduplicates the operation or rejects the stale result?"
176
192
  ## Postconditions
177
193
 
178
194
  - The logical operation and attempt identities, duplicate sources, ordering aggregate, authority token, gap policy, stale-result decision, side effects, durable operation key, payload binding, response replay contract, timeout recovery, processing recovery, queue or webhook dedupe, scheduler or batch dedupe, outbox or inbox boundary, and adversarial tests are explicit.
195
+ - Workflow, plan, step, effect, and attempt identities are separated where replanning exists, and the
196
+ same admitted business intent keeps one effect identity across plan versions.
179
197
  - Duplicate business requests, changed-payload key reuse, concurrent `exists` then `insert`, per-retry operation IDs, duplicate increments, timestamp ordering, sequence gaps, stale state overwrites, cancellation-as-rollback, single-flight-as-proof, provider timeout retries, pre-record external calls, duplicate webhook delivery, queue redelivery, scheduler reruns, double compensation, stuck processing rows, weak locks, and frontend-only guards are fixed or reported.
180
198
  - Duplicate safety claims are backed by configured tests, schema evidence, framework evidence, provider documentation matched to current code, or labeled as static review risk.
181
199
 
@@ -0,0 +1,183 @@
1
+ ---
2
+ mustflow_doc: skill.llm-model-routing-integrity-review
3
+ locale: en
4
+ canonical: true
5
+ revision: 2
6
+ lifecycle: mustflow-owned
7
+ authority: procedure
8
+ name: llm-model-routing-integrity-review
9
+ description: Apply this skill when an LLM system selects, cascades, escalates, falls back, or switches models by task, stage, confidence, verifier result, distribution shift, cost, latency, quality, safety, or context handoff and must optimize total cost per accepted outcome without hiding route-specific failures.
10
+ metadata:
11
+ mustflow_schema: "1"
12
+ mustflow_kind: procedure
13
+ pack_id: mustflow.core
14
+ skill_id: mustflow.core.llm-model-routing-integrity-review
15
+ command_intents:
16
+ - changes_status
17
+ - changes_diff_summary
18
+ - lint
19
+ - build
20
+ - test_related
21
+ - test
22
+ - docs_validate_fast
23
+ - test_release
24
+ - mustflow_check
25
+ ---
26
+
27
+ # LLM Model Routing Integrity Review
28
+
29
+ <!-- mustflow-section: purpose -->
30
+ ## Purpose
31
+
32
+ Choose the least expensive model path that still satisfies measured outcome, latency, and safety
33
+ requirements. Prevent a cheap per-call route from becoming expensive after retries and repairs, and
34
+ prevent a strong-model default from becoming an unevaluated permanent tax.
35
+
36
+ <!-- mustflow-section: use-when -->
37
+ ## Use When
38
+
39
+ - A system adds or changes a model router, small-to-large cascade, escalation threshold, fallback,
40
+ provider or region route, stage-specific model, ensemble, speculative path, or confidence gate.
41
+ - Planning, retrieval, coding, validation, summarization, classification, or tool-use stages may use
42
+ different models and context can be lost when the route changes.
43
+ - Routing depends on cost, latency, accepted-outcome probability, external verifier results,
44
+ uncertainty, risk, task value, traffic volume, or distribution shift.
45
+ - A claim says a smaller model is sufficient, a stronger model is worth the premium, or a routing
46
+ system pays for its own fixed evaluation and operating cost.
47
+
48
+ <!-- mustflow-section: do-not-use-when -->
49
+ ## Do Not Use When
50
+
51
+ - The main risk is prompt payload size, cacheability, token budgets, reasoning spend, or retry replay
52
+ rather than route correctness; use `llm-token-cost-control-review`.
53
+ - The main risk is response speed, streaming, speculative cancellation, or time to first useful
54
+ output; use `llm-response-latency-review`.
55
+ - The main risk is agent tool authority, approval, side effects, or autonomy; use
56
+ `agent-execution-control-review`.
57
+ - The main risk is selling standard versus premium outcomes, charging for automatic escalation,
58
+ task-unit price, managed provider cost, or BYOK packaging; use
59
+ `llm-product-monetization-review`. Keep this skill for route correctness and total route cost.
60
+ - The task chooses a provider or technology stack for strategic reasons rather than routing requests
61
+ among already supported models; use `technology-stack-selection`.
62
+
63
+ <!-- mustflow-section: required-inputs -->
64
+ ## Required Inputs
65
+
66
+ - Outcome contract: accepted task result, deterministic and semantic validators, safety floor,
67
+ latency objective, abstain or escalation state, and value or loss of a wrong result.
68
+ - Route ledger: task and stage types, candidate models and versions, route features, baseline route,
69
+ escalation and fallback rules, context handoff, retry owner, route version, and stable assignment.
70
+ - Evidence ledger: per-route accepted-outcome rate, confidence interval or conservative bound,
71
+ false-accept and false-reject rates, route coverage, sample size, task mix, risk mix, and evaluator
72
+ independence.
73
+ - Cost and latency ledger: model, router, verifier, retry, repair, handoff, supervision, and fixed
74
+ evaluation, adapter, monitoring, migration, and maintenance cost.
75
+ - Distribution ledger: training and evaluation distribution, production feature distribution,
76
+ missing-feature behavior, out-of-distribution detector, drift monitor, and safe fallback.
77
+ - Safety ledger: tasks that require a fixed strong route, deterministic path, human decision, or
78
+ hard refusal regardless of expected financial utility.
79
+
80
+ <!-- mustflow-section: preconditions -->
81
+ ## Preconditions
82
+
83
+ - Establish a capable-model baseline and accepted-outcome evaluator before claiming a cheaper route
84
+ preserves quality.
85
+ - Treat model self-confidence, verbal certainty, token probability, and router score as uncalibrated
86
+ until matched to external outcome evidence on the current task distribution.
87
+ - Refresh provider model, price, latency, retention, and endpoint claims before embedding exact
88
+ values. Do not copy benchmark percentages or universal thresholds into reusable policy.
89
+ - Keep command execution under `.mustflow/config/commands.toml`; this skill does not authorize live
90
+ model calls, traffic routing, billing access, or production rollout.
91
+
92
+ <!-- mustflow-section: allowed-edits -->
93
+ ## Allowed Edits
94
+
95
+ - Add or refine route schemas, features, baseline and candidate policies, external verifiers,
96
+ escalation and abstain states, context handoff, OOD fallback, metrics, eval fixtures, shadow
97
+ comparisons, tests, docs, route metadata, and synchronized templates.
98
+ - Move deterministic classification, validation, policy, arithmetic, schema checking, and database
99
+ postconditions out of model routing when code can decide them.
100
+ - Do not weaken quality or safety validators to make a cheaper route appear successful.
101
+ - Do not route high-impact decisions by average utility when a hard safety or authorization gate
102
+ fails.
103
+
104
+ <!-- mustflow-section: procedure -->
105
+ ## Procedure
106
+
107
+ 1. Define the accepted outcome before selecting a model. Include final-state correctness, safety,
108
+ latency, abstention, and repair rules rather than counting a syntactically valid response.
109
+ 2. Establish the baseline with the most capable practical route and current evals. Only then replace
110
+ bounded task classes or stages with smaller models where accepted outcomes remain sufficient.
111
+ 3. Route deterministic work to code first. Do not pay a model or router to decide values that a
112
+ schema, compiler, test, database lookup, permission engine, or exact rule can determine.
113
+ 4. Choose route granularity from real context boundaries. Keep tightly coupled planning and
114
+ execution on one model when switching would lose hidden constraints or require replaying most of
115
+ the context. Split classification, retrieval-query generation, formatting, or verification when
116
+ their inputs and outputs are explicit and independently checkable.
117
+ 5. Measure route quality with external evidence. Calibrate router scores against tests, schema and
118
+ business validation, source agreement, database postconditions, human adjudication, or another
119
+ independent oracle. Preserve false-accept cost separately from ordinary failure.
120
+ 6. Optimize total cost per accepted outcome. Include router, verifier, failed cheap attempts,
121
+ escalation, retries, repair, context transfer, supervision, latency loss, and fixed operating
122
+ cost. A lower model price is not a saving when acceptance falls or repair rises.
123
+ 7. Apply hard constraints before utility. Authorization, privacy, irreversible effects, safety
124
+ invariants, and maximum tolerable loss can force a deterministic, strong-model, human, or refusal
125
+ route even when a cheaper path has positive expected value.
126
+ 8. Build cheap-first cascades only when the acceptance gate can reject bad cheap outputs with known
127
+ false-accept behavior. Without an external verifier, repeated sampling and self-critique do not
128
+ prove which candidate is correct.
129
+ 9. Escalate from observed evidence, not model preference. Use validator failure, missing required
130
+ facts, task complexity, risk class, OOD state, repeated error signature, or explicit uncertainty
131
+ features. Do not escalate privileges when escalating model capability.
132
+ 10. Handle distribution shift explicitly. Compare production features and outcomes with the route's
133
+ evaluation distribution. Send missing, novel, drifted, or low-support cases to the safe baseline
134
+ route, deterministic review, or human owner.
135
+ 11. Account for context handoff loss. Version the handoff schema, preserve verified facts and hard
136
+ constraints, avoid summary-on-summary drift, and measure whether route switches cause omissions
137
+ or contradictory plans.
138
+ 12. Keep route and retry ownership separate. A route change may be a recovery action, but it must
139
+ share the workflow's attempt budget and idempotency state rather than restarting the task as a
140
+ new operation.
141
+ 13. Compare candidates in shadow or replay before promotion. Pin model and route versions, use the
142
+ same representative task mix, preserve delayed outcomes, and route rollout details through
143
+ `agent-release-bundle-rollout-review` when the router is part of an agent bundle.
144
+ 14. Recalculate routing economics when traffic, task distribution, models, prices, verifier quality,
145
+ repair rate, or maintenance burden changes. Route broad investment decisions to
146
+ `automation-investment-case-review`.
147
+
148
+ <!-- mustflow-section: postconditions -->
149
+ ## Postconditions
150
+
151
+ - Every route has an accepted-outcome contract, external evidence, safe fallback, and versioned
152
+ decision reason.
153
+ - Cost comparisons include failed attempts, verification, escalation, repair, context handoff, and
154
+ fixed operating cost.
155
+ - OOD and hard-safety cases cannot silently enter a cheap route.
156
+ - Model self-confidence is not treated as outcome probability without calibration.
157
+
158
+ <!-- mustflow-section: verification -->
159
+ ## Verification
160
+
161
+ Use configured oneshot command intents when available: `changes_status`, `changes_diff_summary`,
162
+ `lint`, `build`, `test_related`, `test`, `docs_validate_fast`, `test_release`, and `mustflow_check`.
163
+ Do not infer raw provider, model, billing, traffic, eval, or rollout commands.
164
+
165
+ <!-- mustflow-section: failure-handling -->
166
+ ## Failure Handling
167
+
168
+ - If no accepted-outcome evaluator exists, keep the baseline route and report the missing gate.
169
+ - If route evidence is sparse or distribution support is missing, abstain from the cheaper route.
170
+ - If false accepts can create high-impact harm, treat the route as unsafe until a hard gate or human
171
+ decision closes the boundary.
172
+ - If route savings depend on stale price or benchmark data, mark the economic claim unverified.
173
+
174
+ <!-- mustflow-section: output-format -->
175
+ ## Output Format
176
+
177
+ - Outcome contract and baseline route
178
+ - Candidate route, features, version, and context boundary
179
+ - External calibration, false-accept, OOD, and drift evidence
180
+ - Total cost-per-accepted-outcome and latency findings
181
+ - Hard safety, escalation, fallback, and abstain decisions
182
+ - Command intents run and skipped checks
183
+ - Remaining model-routing risk