mustflow 2.116.4 → 2.117.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. package/package.json +1 -1
  2. package/templates/default/i18n.toml +145 -7
  3. package/templates/default/locales/en/.mustflow/skills/INDEX.md +121 -11
  4. package/templates/default/locales/en/.mustflow/skills/agent-eval-integrity-review/SKILL.md +62 -34
  5. package/templates/default/locales/en/.mustflow/skills/agent-execution-control-review/SKILL.md +102 -31
  6. package/templates/default/locales/en/.mustflow/skills/agent-memory-context-governance-review/SKILL.md +163 -0
  7. package/templates/default/locales/en/.mustflow/skills/agent-planning-recovery-review/SKILL.md +180 -0
  8. package/templates/default/locales/en/.mustflow/skills/agent-release-bundle-rollout-review/SKILL.md +181 -0
  9. package/templates/default/locales/en/.mustflow/skills/agent-runtime-isolation-review/SKILL.md +196 -0
  10. package/templates/default/locales/en/.mustflow/skills/agent-runtime-multi-worker-review/SKILL.md +180 -0
  11. package/templates/default/locales/en/.mustflow/skills/automation-investment-case-review/SKILL.md +173 -0
  12. package/templates/default/locales/en/.mustflow/skills/client-platform-strategy-review/SKILL.md +236 -0
  13. package/templates/default/locales/en/.mustflow/skills/credit-ledger-integrity-review/SKILL.md +8 -4
  14. package/templates/default/locales/en/.mustflow/skills/credit-monetization-integrity-review/SKILL.md +283 -0
  15. package/templates/default/locales/en/.mustflow/skills/desktop-commercial-distribution-review/SKILL.md +225 -0
  16. package/templates/default/locales/en/.mustflow/skills/external-prompt-injection-defense/SKILL.md +49 -3
  17. package/templates/default/locales/en/.mustflow/skills/freemium-ad-monetization-review/SKILL.md +196 -0
  18. package/templates/default/locales/en/.mustflow/skills/game-economy-monetization-review/SKILL.md +208 -0
  19. package/templates/default/locales/en/.mustflow/skills/game-liveops-commerce-integrity-review/SKILL.md +237 -0
  20. package/templates/default/locales/en/.mustflow/skills/growth-distribution-integrity-review/SKILL.md +247 -0
  21. package/templates/default/locales/en/.mustflow/skills/idempotency-integrity-review/SKILL.md +20 -2
  22. package/templates/default/locales/en/.mustflow/skills/llm-model-routing-integrity-review/SKILL.md +183 -0
  23. package/templates/default/locales/en/.mustflow/skills/llm-product-monetization-review/SKILL.md +311 -0
  24. package/templates/default/locales/en/.mustflow/skills/llm-token-cost-control-review/SKILL.md +15 -1
  25. package/templates/default/locales/en/.mustflow/skills/localization-market-expansion-review/SKILL.md +224 -0
  26. package/templates/default/locales/en/.mustflow/skills/multi-agent-work-coordination/SKILL.md +7 -2
  27. package/templates/default/locales/en/.mustflow/skills/pricing-model-integrity-review/SKILL.md +288 -0
  28. package/templates/default/locales/en/.mustflow/skills/product-engagement-retention-review/SKILL.md +234 -0
  29. package/templates/default/locales/en/.mustflow/skills/product-onboarding-activation-review/SKILL.md +269 -0
  30. package/templates/default/locales/en/.mustflow/skills/product-portfolio-integrity-review/SKILL.md +233 -0
  31. package/templates/default/locales/en/.mustflow/skills/prompt-contract-quality-review/SKILL.md +1 -0
  32. package/templates/default/locales/en/.mustflow/skills/referral-incentive-integrity-review/SKILL.md +206 -0
  33. package/templates/default/locales/en/.mustflow/skills/retry-policy-integrity-review/SKILL.md +45 -2
  34. package/templates/default/locales/en/.mustflow/skills/routes.toml +329 -7
  35. package/templates/default/locales/en/.mustflow/skills/service-portfolio-capital-allocation-review/SKILL.md +245 -0
  36. package/templates/default/locales/en/.mustflow/skills/subscription-retention-profit-review/SKILL.md +217 -0
  37. package/templates/default/manifest.toml +86 -1
@@ -0,0 +1,206 @@
1
+ ---
2
+ mustflow_doc: skill.referral-incentive-integrity-review
3
+ locale: en
4
+ canonical: true
5
+ revision: 2
6
+ lifecycle: mustflow-owned
7
+ authority: procedure
8
+ name: referral-incentive-integrity-review
9
+ description: Apply this skill when a product changes inviter-only, invitee-only, dual-sided, tiered, or milestone referral rewards; referral attribution; invite links or codes; valid-referral qualification; pending, vested, or reversed rewards; self-referral and reward farming controls; rolling thresholds; referral incrementality; or referred-user contribution and must grow without paying for natural signups, fraudulent identities, refundable purchases, or vanity referral counts.
10
+ metadata:
11
+ mustflow_schema: "1"
12
+ mustflow_kind: procedure
13
+ pack_id: mustflow.core
14
+ skill_id: mustflow.core.referral-incentive-integrity-review
15
+ command_intents:
16
+ - changes_status
17
+ - changes_diff_summary
18
+ - lint
19
+ - build
20
+ - test_related
21
+ - test
22
+ - docs_validate_fast
23
+ - test_release
24
+ - mustflow_check
25
+ ---
26
+
27
+ # Referral Incentive Integrity Review
28
+
29
+ <!-- mustflow-section: purpose -->
30
+ ## Purpose
31
+
32
+ Reward incremental, durable referrals without confusing reward direction with tiering, paying both
33
+ sides for disposable signup identities, crediting users who would have joined anyway, or treating
34
+ referred-user revenue as causal growth. Preserve transparent eligibility, reversible reward state,
35
+ privacy, appeals, and unit economics.
36
+
37
+ <!-- mustflow-section: use-when -->
38
+ ## Use When
39
+
40
+ - A program changes inviter-only, invitee-only, dual-sided, shared-budget, tiered, milestone, cash,
41
+ credit, discount, premium-time, access, badge, or partner rewards.
42
+ - Invite links, codes, attribution windows, late code entry, pre-existing accounts, channel conflicts,
43
+ last-click or first-touch rules, or referral ownership change.
44
+ - A valid referral depends on signup, verification, first owned value, retained use, purchase, refund
45
+ window, subscription survival, or another downstream event.
46
+ - Self-referral, account farms, shared devices, households, payment overlap, refund or chargeback,
47
+ reward reversal, manual review, appeal, or referral incrementality is implemented or reported.
48
+
49
+ <!-- mustflow-section: do-not-use-when -->
50
+ ## Do Not Use When
51
+
52
+ - The task only changes promotional-credit balance rights, expiry, spend order, or settlement; use
53
+ `credit-monetization-integrity-review` and `credit-ledger-integrity-review`.
54
+ - The task only changes signup, authentication, identity linking, or onboarding without a referral
55
+ reward or attribution policy; use `product-onboarding-activation-review` and the identity owner.
56
+ - The task is affiliate, influencer, reseller, creator, sales-partner, recurring-commission, or
57
+ owned-product cross-promotion policy rather than a customer referral reward; use
58
+ `growth-distribution-integrity-review` plus the matching payment, tax, contract, and legal owners.
59
+ - The task requests jurisdiction-specific sweepstakes, marketing, tax, privacy, anti-spam, or
60
+ consumer-law advice. Use current qualified authority; this skill supplies the product evidence.
61
+
62
+ <!-- mustflow-section: required-inputs -->
63
+ ## Required Inputs
64
+
65
+ - Actor ledger: inviter, invitee, account age, identity assurance, customer type, geography,
66
+ household or organization relationship, prior account, prior invitation, eligibility, and consent.
67
+ - Attribution ledger: code or link, issuer, channel, creation and click time, signup and qualifying
68
+ time, window, priority, late entry, conflicting claims, pre-existing intent, and version.
69
+ - Qualification ledger: verification, first owned value, retained activity, purchase, refund and
70
+ chargeback horizon, subscription survival, excluded behavior, invalidation, and evidence owner.
71
+ - Reward ledger: recipient, type, face value, expected cost, transferability, expiry, cap, pending,
72
+ vested, granted, consumed, reversed, appealed, and accounting, tax, platform, or legal review.
73
+ - Tier ledger: valid-referral count, rolling window, thresholds, marginal and cumulative reward,
74
+ reset, cap, manual review, downgrade, and campaign version.
75
+ - Abuse ledger: identity, device, payment, billing, phone, network, address, behavior and velocity
76
+ signals, graph links, false-positive risk, reason code, reviewer, appeal, and retention limits.
77
+ - Experiment ledger: eligible population, assignment, no-program or business-as-usual holdout,
78
+ reward-budget parity, exposure, natural signup, channel cannibalization, horizon, contribution,
79
+ and promotion rule.
80
+
81
+ <!-- mustflow-section: preconditions -->
82
+ ## Preconditions
83
+
84
+ - Separate who receives a reward, when it qualifies, when it vests, and how tiers accumulate. These
85
+ are independent axes.
86
+ - Define the incremental growth outcome and valid-referral event before issuing reward value.
87
+ - Preserve a stable program and attribution version across invite, qualification, vesting, reversal,
88
+ and appeal.
89
+ - Treat copied split ratios, reward amounts, signup conditions, day counts, tier thresholds, rolling
90
+ windows, and benchmark lifts as hypotheses, not defaults.
91
+ - Refresh current privacy, identity, messaging, platform, tax, accounting, and jurisdiction rules;
92
+ this skill does not authorize outbound invitations, live rewards, payments, experiments, or bans.
93
+
94
+ <!-- mustflow-section: allowed-edits -->
95
+ ## Allowed Edits
96
+
97
+ - Add or refine referral eligibility, attribution, valid-referral qualification, pending and vesting
98
+ states, reversal, tiering, caps, anti-abuse evidence, appeals, experiment assignment, contribution
99
+ metrics, fixtures, tests, docs, route metadata, and synchronized templates.
100
+ - Replace signup-only cash-like grants, permanent lifetime tier accumulation, single-signal bans,
101
+ or gross referred revenue claims with bounded downstream qualification and causal evidence.
102
+ - Do not silently reassign an earned referral, expose private abuse signals, auto-ban from one weak
103
+ signal, or reverse consumed value without an explicit lawful negative-balance or recovery policy.
104
+
105
+ <!-- mustflow-section: procedure -->
106
+ ## Procedure
107
+
108
+ 1. Split four decisions: reward direction, qualifying event, vesting and reversal, and tier schedule.
109
+ Dual-sided and tiered rewards are not competing alternatives.
110
+ 2. Define the headline result as incremental contribution per eligible participant or acquired user
111
+ over a declared horizon. Include reward cost, variable service cost, refunds, chargebacks,
112
+ support, fraud loss, messaging cost, channel cannibalization, retained value, and later revenue.
113
+ 3. Define referral eligibility before attribution. State whether existing accounts, prior visitors,
114
+ employees, partners, same organization, same household, minors, prior payers, or previously invited
115
+ users can participate and which actor is ineligible.
116
+ 4. Make attribution deterministic. Version code and link ownership, time window, channel priority,
117
+ late code entry, multiple inviters, cross-device joins, account merges, and pre-existing account
118
+ behavior. Do not let support choose winners without a recorded rule and audit trail.
119
+ 5. Do not equate signup with a valid referral. Choose a downstream event that represents real product
120
+ value and resists cheap identity creation, such as verified first owned value, retained use, or a
121
+ payment that has survived the applicable refund and dispute conditions.
122
+ 6. Give invitee value when it improves honest activation, but keep it bounded, nontransferable where
123
+ appropriate, and tied to the product. Delay inviter vesting until the invitee reaches the declared
124
+ durable condition. This sequence is a candidate, not a universal reward direction.
125
+ 7. Keep reward states explicit: proposed, pending, qualified, vested, granted, partly consumed,
126
+ reversed, expired, appealed, restored, and terminal. Use idempotent identities and preserve the
127
+ event and policy version that authorized each transition.
128
+ 8. Define reversal before launch. Cover failed verification, duplicate identity, refund, chargeback,
129
+ subscription cancellation, account deletion, policy violation, provider reversal, and later fraud
130
+ evidence. Separate unvested cancellation from clawing back vested or consumed rights.
131
+ 9. Hold economic value only as long as qualification needs. State the expected delay before invite,
132
+ explain pending status to both actors, and avoid an indefinite fraud-review state.
133
+ 10. Compare reward directions at equal expected total budget when the question is who should receive
134
+ value. If one arm spends more, label it a budget-and-direction test rather than a framing test.
135
+ 11. Add tiers only to valid referrals. Use a declared rolling or campaign window, cap, reset, and
136
+ marginal reward schedule. Do not let lifetime signup count create permanent high-yield farming.
137
+ 12. Keep higher tiers economically bounded. Prefer product value with controlled expected cost when
138
+ appropriate, but include service liability, credit consumption, subscription displacement,
139
+ transferability, and tax or accounting treatment rather than calling it free.
140
+ 13. Detect abuse from multiple consistent signals and behavior. Use identity assurance, device,
141
+ payment, billing, phone, network, address, velocity, graph, refund, and usage evidence
142
+ proportionately; one shared IP, device, or household is not sufficient proof by itself.
143
+ 14. Design for legitimate collisions. Cover families, schools, offices, shared devices, travel,
144
+ recycled phone numbers, privacy relays, payment by another household member, and accessibility
145
+ support. Record reason codes, manual-review thresholds, and an appeal path.
146
+ 15. Limit disclosure. Tell users the eligibility or review outcome and actionable remedy without
147
+ revealing thresholds, graph features, private identifiers, or a recipe for evasion.
148
+ 16. Separate invitation delivery from referral attribution. Respect consent, anti-spam, suppression,
149
+ sender identity, frequency, and channel rules; do not import contacts or message third parties
150
+ merely because a referral reward exists.
151
+ 17. Measure incrementality with a no-program, no-reward, or business-as-usual comparison where
152
+ feasible. Track natural signup, users who add a code after deciding to join, organic and paid
153
+ channel displacement, and inviter behavior. Referred-versus-organic revenue alone is observational.
154
+ 18. Preserve experiment assignment before reward exposure and keep nonclickers, nonjoiners,
155
+ unqualified invites, and invalid referrals in the relevant intent-to-treat denominator.
156
+ 19. Predeclare guardrails: spam complaints, blocks, privacy requests, false-positive reviews, appeal
157
+ reversals, account farms, reward cost, refund and chargeback, low-quality activity, support,
158
+ concentration, and downstream retained value.
159
+ 20. Promote only a reversible, versioned program whose incremental contribution and quality improve
160
+ without paying for disposable identities, natural signups, refundable transactions, or hidden
161
+ messaging and privacy costs.
162
+
163
+ <!-- mustflow-section: postconditions -->
164
+ ## Postconditions
165
+
166
+ - Reward direction, valid-referral qualification, vesting, reversal, and tiering remain distinct.
167
+ - Attribution, conflicts, late entry, account merges, and pre-existing users have deterministic rules.
168
+ - Reward states are reconstructable and downstream qualification precedes durable inviter value.
169
+ - Abuse decisions use multiple signals, legitimate-collision handling, bounded review, and appeals.
170
+ - Headline growth uses causal eligible-population contribution rather than raw invites, signups,
171
+ referred-user revenue, or reward claims.
172
+
173
+ <!-- mustflow-section: verification -->
174
+ ## Verification
175
+
176
+ Use configured oneshot command intents when available: `changes_status`, `changes_diff_summary`,
177
+ `lint`, `build`, `test_related`, `test`, `docs_validate_fast`, `test_release`, and `mustflow_check`.
178
+ Do not infer live reward, credit, payment, identity, messaging, analytics, experiment, deployment, or
179
+ production commands.
180
+
181
+ <!-- mustflow-section: failure-handling -->
182
+ ## Failure Handling
183
+
184
+ - If attribution or qualification cannot be reconstructed, pause new vesting rather than guessing
185
+ the inviter or granting duplicate value.
186
+ - If only signup or referred-versus-organic revenue exists, report acquisition association and do not
187
+ claim incremental growth or profitable referrals.
188
+ - If abuse evidence relies on one weak signal, keep the reward pending for bounded review or allow it;
189
+ do not auto-ban without corroboration and an appeal path.
190
+ - If reversal rights for vested or consumed rewards are unresolved, preserve the current right and
191
+ route future policy to credit, payment, accounting, tax, platform, or legal owners.
192
+ - If higher tiers increase low-quality or fraudulent referrals faster than incremental contribution,
193
+ cap, reset, or remove the tier rather than optimizing claimed referral count.
194
+
195
+ <!-- mustflow-section: output-format -->
196
+ ## Output Format
197
+
198
+ - Eligible inviter and invitee, attribution, conflicts, window, and program version
199
+ - Valid-referral event, pending period, vesting, grant, consumption, reversal, and appeal
200
+ - Reward direction, equal-budget comparison, cost, transferability, and rights
201
+ - Tier window, thresholds, marginal rewards, cap, reset, and review
202
+ - Abuse signals, legitimate collisions, privacy, messaging, false positives, and reason codes
203
+ - Incremental contribution, natural signup, channel displacement, quality, and guardrails
204
+ - Files changed
205
+ - Command intents run and skipped checks
206
+ - Remaining referral-incentive risk
@@ -2,11 +2,11 @@
2
2
  mustflow_doc: skill.retry-policy-integrity-review
3
3
  locale: en
4
4
  canonical: true
5
- revision: 1
5
+ revision: 2
6
6
  lifecycle: mustflow-owned
7
7
  authority: procedure
8
8
  name: retry-policy-integrity-review
9
- description: Apply this skill when code is created, changed, reviewed, or reported and retry loops, SDK or client retry configs, backoff, jitter, timeout, deadline, Retry-After, retry predicates, layered retries, circuit breakers, bulkheads, token buckets, queue redelivery, broker retries, cancellation-aware sleeps, or retry observability can amplify failure, duplicate side effects, hide permanent errors, exhaust pools, or overload dependencies.
9
+ description: Apply this skill when code or an agent runtime repeats, repairs, reroutes, replans, or stops after failures and retry loops, prompt correction, model or tool fallback, SDK configs, backoff, jitter, timeout, deadline, Retry-After, retry predicates, layered retries, circuit breakers, queue redelivery, cancellation, or retry observability can amplify failure, duplicate effects, hide permanent errors, or repeat the same unsupported recovery action.
10
10
  metadata:
11
11
  mustflow_schema: "1"
12
12
  mustflow_kind: procedure
@@ -39,6 +39,9 @@ The review question is not "does this code retry?" It is "when the dependency is
39
39
  - Code creates, changes, reviews, or reports retry loops, SDK retry options, client retry middleware, `while true`, `for (;;)`, recursive retry, `maxAttempts`, `maxRetries`, `maxElapsedTime`, `deadline`, `timeout`, `sleep`, `delay`, backoff, jitter, `Retry-After`, circuit breaker, bulkhead, token bucket, rate limiter, or cancellation-aware retry behavior.
40
40
  - HTTP, database, Redis, queue, stream, object storage, payment, email, notification, file upload, webhook, scheduler, batch, worker, or provider calls can be repeated after failure, timeout, cancellation, rate limit, lock conflict, transaction retry, or redelivery.
41
41
  - A workflow has retries in more than one layer, such as caller retry plus service retry plus SDK retry plus load balancer retry plus queue redelivery plus broker retry.
42
+ - An agent can repeat the same prompt, add validator feedback, choose another model or tool, replan,
43
+ request more privilege, or stop, and the correct recovery depends on whether new trustworthy
44
+ information changes the failure mechanism.
42
45
  - A review or final report claims a path is resilient, retry-safe, backoff-protected, rate-limit-aware, idempotent, transient-only, bounded, cancellation-safe, or protected by a circuit breaker, bulkhead, pool, timeout, or queue retry policy.
43
46
 
44
47
  <!-- mustflow-section: do-not-use-when -->
@@ -57,10 +60,17 @@ The review question is not "does this code retry?" It is "when the dependency is
57
60
  - Layered retry ledger: caller, API handler, service, SDK, driver, proxy, load balancer, queue, worker, scheduler, and provider retries with per-layer max attempts and elapsed-time budget.
58
61
  - Attempt budget: max attempts, max elapsed time, per-attempt timeout, total deadline, cancellation signal, sleep behavior, and whether DNS, TLS, connection checkout, pool wait, request body upload, response streaming, and response parsing are inside the budget.
59
62
  - Retry predicate: exception types, status codes, provider errors, timeout classes, rate-limit responses, transient versus permanent failure, unknown outcome, cancellation, validation errors, authorization errors, and programmer bugs.
63
+ - Recovery-action ledger: normalized failure class and signature, new external evidence, same-request
64
+ retry, corrected-input or prompt repair, alternate model, alternate tool or data source, local or
65
+ workflow replan, stop or human handoff, expected recovery value, added latency and cost, duplicate
66
+ risk, remaining budget, and why the action changes the failure mechanism.
60
67
  - Side-effect and idempotency ledger: logical operation, idempotency key, key scope, request body hash, conditional write, transaction boundary, lock boundary, provider call, stream or upload body replayability, queue ack or commit, and committed response boundary.
61
68
  - Backoff and jitter policy: fixed sleep, exponential backoff, cap, jitter type, `Retry-After` parsing, clock skew, maximum delay, per-key backoff, global backoff, and reset after success.
62
69
  - Overload and throttling evidence: pool sizes, concurrency limit, bulkhead, token bucket, circuit breaker state, retry queue size, rate-limit budget, per-tenant or per-key fairness, and dependency-specific limits.
63
70
  - Observability and test evidence: attempt logs, per-attempt spans, retry metrics, retry exhaustion metrics, `Retry-After` values, request id, correlation id, cancellation tests, permanent-failure tests, timeout-budget tests, and configured command intents.
71
+ - For model-directed recovery, external verifier evidence, repeated state-action-error signature,
72
+ route version, tool capability ceiling, plan version, and whether a self-critique or resampling path
73
+ has an independent way to select the correct result.
64
74
 
65
75
  <!-- mustflow-section: preconditions -->
66
76
  ## Preconditions
@@ -149,6 +159,31 @@ The review question is not "does this code retry?" It is "when the dependency is
149
159
  - Good tests cover permanent failure not retried, max attempts, max elapsed time, total deadline, cancellation during sleep, retry-after parsing and clamping, jitter boundedness, idempotency key reuse, unknown outcome, pool or concurrency limit, and wrapper cause preservation.
150
160
  - If deterministic timing tests are hard, use fake clocks, injected sleeper, injected retry policy, or local test helpers already present in the repository.
151
161
  - If configured integration evidence is missing, report static risk and the missing manual or integration proof instead of approving the retry path.
162
+ 18. Choose among recovery actions, not only retry or fail.
163
+ - Compare same-request retry, corrected input or prompt, alternate model, alternate tool or source,
164
+ local replan, workflow replan, stop, and human handoff for the normalized failure class.
165
+ - Continue only when current evidence says the action has positive remaining value after latency,
166
+ cost, risk, duplicate-effect exposure, and consumed budgets. Use repository policy rather than a
167
+ universal attempt count.
168
+ 19. Require new trustworthy information for repair loops.
169
+ - Correct a prompt or model output with compiler errors, failed tests, schema paths,
170
+ postcondition mismatches, source conflicts, provider status, or another external signal.
171
+ - Do not treat the same model rereading its own answer, generic "think harder" text, or an
172
+ unverified critique as new evidence.
173
+ 20. Repeat the identical request only when the failure is plausibly transient and the operation is
174
+ replay-safe, or when independent sampling has an external verifier that can choose a valid
175
+ result. If no verifier can distinguish candidates, resampling does not establish correctness.
176
+ 21. Change model, tool, source, or plan only when it changes the failure mechanism. Switching models
177
+ cannot repair missing evidence; repeating search cannot repair a stale source without a changed
178
+ query or source; replanning cannot erase an UNKNOWN external effect.
179
+ 22. Stop repeated failure signatures. Hash or normalize the relevant state, proposed action, error
180
+ class, validator output, and route or tool version. When the same signature recurs without new
181
+ evidence, stop the local loop and take one declared higher-level recovery or handoff rather than
182
+ spending the remaining cap blindly.
183
+ 23. Keep safety failures attenuation-only. Prompt injection, policy denial, secret exposure,
184
+ unauthorized tools, and insufficient privilege do not justify a stronger model or wider
185
+ capability. Isolate the input, narrow tools, return to read or draft mode, request explicit
186
+ authority, or stop.
152
187
 
153
188
  <!-- mustflow-section: postconditions -->
154
189
  ## Postconditions
@@ -156,6 +191,8 @@ The review question is not "does this code retry?" It is "when the dependency is
156
191
  - Retry layers, attempt multiplication, max attempts, max elapsed time, per-attempt timeout, total deadline, retry predicate, backoff, jitter, `Retry-After`, cancellation behavior, side-effect replay safety, idempotency key reuse, transaction or lock placement, pool pressure, throttling, circuit-breaker ordering, wrapper diagnostics, queue overlap, dependency policy, observability, and retry tests are explicit.
157
192
  - Infinite retry, broad catch-and-retry, permanent-error retry, unknown-outcome replay, new idempotency key per attempt, fixed-sleep herd behavior, retry inside long transaction or lock, pool exhaustion, app-plus-broker retry multiplication, stale failure counters, wrapper cause loss, committed-response retry, non-replayable body retry, cancellation-ignoring sleep, and missing retry metrics are fixed or reported.
158
193
  - Retry-safety claims are backed by configured tests, dependency or framework evidence matched to current code, schema or idempotency evidence, static review evidence, or labeled as manual-only or missing.
194
+ - Model and tool recovery uses external evidence, changes the failure mechanism, preserves the
195
+ capability ceiling, and stops repeated signatures instead of relying on self-critique loops.
159
196
 
160
197
  <!-- mustflow-section: verification -->
161
198
  ## Verification
@@ -180,6 +217,10 @@ Prefer the narrowest configured test, build, docs, release, or mustflow intent t
180
217
  - If a configured command fails, preserve the failing intent, failing assertion or output tail, and the retry invariant it exercised before editing again.
181
218
  - If retry layers cannot be enumerated, report that the path is not reviewable for retry amplification yet.
182
219
  - If the retry predicate cannot distinguish transient from permanent failure, report the missing classification instead of accepting a broad retry.
220
+ - If no external signal can distinguish a repaired or resampled output from the original failure,
221
+ stop automatic self-correction and report the missing verifier.
222
+ - If a proposed recovery widens privilege after a tool or policy failure, reject that recovery and
223
+ require a separate authority decision.
183
224
  - If safe repair requires provider idempotency, durable operation records, queue or broker configuration, circuit-breaker architecture, dependency-specific client policy, fake-clock test infrastructure, load testing, or integration replay outside the current scope, report the missing boundary.
184
225
  - If deterministic retry proof is not configured, complete available verification and report the missing manual or integration evidence.
185
226
 
@@ -188,6 +229,8 @@ Prefer the narrowest configured test, build, docs, release, or mustflow intent t
188
229
 
189
230
  - Retry policy boundary reviewed
190
231
  - Retry surface, layered retry ledger, attempt budget, timeout and deadline, retry predicate, unknown outcome, side-effect and idempotency key, transaction or lock placement, pool and concurrency pressure, backoff and jitter, `Retry-After`, global or per-key throttling, resilience-tool ordering, wrapper diagnostics, committed response or streaming body, queue or broker overlap, dependency policy, observability, and test evidence findings
232
+ - Recovery actions, new external evidence, repeated failure signature, model/tool/source/replan or
233
+ stop decision, capability ceiling, and self-correction verifier findings
191
234
  - Retry-policy fixes made or recommended
192
235
  - Evidence level: configured-test evidence, dependency or framework evidence, schema or idempotency evidence, static review risk, manual-only, missing, or not applicable
193
236
  - Command intents run