dsh-logicprobe 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (33) hide show
  1. package/README.en-US.md +7 -1
  2. package/README.md +7 -1
  3. package/lib/concurrency-tool.js +34 -0
  4. package/lib/concurrency.js +76 -0
  5. package/lib/data-engine.js +930 -0
  6. package/lib/data-tool.js +61 -0
  7. package/lib/engine.js +1622 -1364
  8. package/lib/index.js +291 -277
  9. package/lib/tool.js +57 -57
  10. package/lib/types/concurrency-tool.d.ts +8 -0
  11. package/lib/types/concurrency.d.ts +20 -0
  12. package/lib/types/data-engine.d.ts +199 -0
  13. package/lib/types/data-tool.d.ts +10 -0
  14. package/lib/types/engine.d.ts +171 -151
  15. package/package.json +79 -78
  16. package/skills/logicprobe/SKILL.md +285 -268
  17. package/skills/logicprobe/references/__pycache__/verification-harness.cpython-312.pyc +0 -0
  18. package/skills/logicprobe/references/concurrency-risk-guide.md +54 -0
  19. package/skills/logicprobe/references/dsh-model-schema.md +145 -129
  20. package/skills/logicprobe/references/logic-verification-guide.md +463 -413
  21. package/skills/logicprobe/references/verification-harness.py +806 -582
  22. package/skills/logicprobe-datamodel/SKILL.md +124 -0
  23. package/skills/logicprobe-datamodel/references/__pycache__/data-model-harness.cpython-312.pyc +0 -0
  24. package/skills/logicprobe-datamodel/references/data-model-guide.md +62 -0
  25. package/skills/logicprobe-datamodel/references/data-model-harness.py +528 -0
  26. package/skills/logicprobe-datamodel/references/data-model-schema.md +128 -0
  27. package/src/concurrency-tool.ts +37 -0
  28. package/src/concurrency.ts +102 -0
  29. package/src/data-engine.ts +1001 -0
  30. package/src/data-tool.ts +65 -0
  31. package/src/engine.ts +234 -0
  32. package/src/index.ts +315 -301
  33. package/src/tool.ts +60 -60
@@ -1,413 +1,463 @@
1
- # Logic Verification Guide
2
-
3
- Deep reference for Phase 2a/2b of logicprobe. Load this when the skill triggers escalation to logic-primitive verification. Covers model extraction methodology, probe design patterns, counter-example interpretation, and known limitations.
4
-
5
- ## Model Extraction Methodology
6
-
7
- ### Step 1: Identify the State Machine Boundary
8
-
9
- Before extracting states, determine what IS and IS NOT part of the machine:
10
-
11
- - **In**: states the plan explicitly names, events the plan describes, guards the plan states
12
- - **Out**: hardware-level behavior (register writes, DMA transfers), OS-level scheduling (task switches, priority inversion), external system interactions not described in the plan
13
-
14
- If the plan doesn't clearly define the boundary, flag it: "Plan does not define state machine scope — model may be incomplete."
15
-
16
- ### Step 2: Extract States
17
-
18
- Scan the plan for:
19
-
20
- - Enumerated state types (`typedef enum { ... } xxx_state_t`)
21
- - Named phases in prose ("the system then enters the RECOVERING phase")
22
- - Implicit states (error paths described but not named — give them explicit names in the model)
23
-
24
- **Rule**: If a behavioral section describes a distinct set of actions before a transition, it's a state — even if the plan doesn't label it as one.
25
-
26
- ### Step 3: Extract Transitions
27
-
28
- For each state, list:
29
-
30
- - **Explicit transitions**: `state = NEXT_STATE` assignments, `→` arrows in diagrams, prose like "then transitions to X"
31
- - **Guard conditions**: `if/else` branches that split a single event into multiple next states
32
- - **Timer-driven transitions**: timeouts, retry delays, cooldown periods
33
-
34
- **Edge case**: A transition mentioned in prose but missing from the state diagram. Flag it — "Plan describes transition X but state diagram omits it."
35
-
36
- ### Step 4: Identify Invariants
37
-
38
- Extract every claim that uses absolute language:
39
-
40
- | Pattern | Example | Invariant to Check |
41
- |---------|---------|-------------------|
42
- | "always ..." | "always returns to IDLE" | From every state reachable in ≤K steps, IDLE is reachable |
43
- | "never ..." | "never enters ACTIVE without power_ready" | ACTIVE is not reachable from any path that hasn't passed through power_ready |
44
- | "guaranteed ..." | "guaranteed cleanup within 500ms" | Every ERROR→RECOVERING path includes a cooldown transition |
45
- | "cannot ..." | "cannot deadlock" | No absorbing cycle exists |
46
- | "all paths ..." | "all paths lead to ERROR on failure" | Every failure event reaches ERROR (not FATAL, not stuck) |
47
-
48
- ### Step 5: Confirm with User
49
-
50
- Show the extracted transition table BEFORE writing the harness. Ask:
51
-
52
- 1. Are all states captured?
53
- 2. Are all transitions and guards correct?
54
- 3. Are there undocumented transitions not in the plan?
55
- 4. Is the initial state correct?
56
-
57
- Only proceed after confirmation. A wrong model produces wrong counter-examples, which wastes more time than no verification at all.
58
-
59
- ## Probe Design Patterns
60
-
61
- ### A1: Unexpected Event Injection
62
-
63
- ```python
64
- # For each state S, list all events defined anywhere in the machine
65
- # For each event E not handled by S, test: what happens?
66
- def unexpected_event_probe(states):
67
- all_events = set()
68
- for trans in states.values():
69
- all_events.update(trans.keys())
70
-
71
- findings = []
72
- for state, transitions in states.items():
73
- unhandled = all_events - set(transitions.keys())
74
- if unhandled:
75
- findings.append({
76
- "state": state,
77
- "unhandled_events": list(unhandled),
78
- "risk": "Silent ignore or undefined behavior"
79
- })
80
- return findings
81
- ```
82
-
83
- ### A2: Race Interleaving
84
-
85
- For pairs of events that can arrive within the same tick (e.g., timer expiry + message reception):
86
-
87
- ```python
88
- def race_interleaving(states, init, event_pairs):
89
- # event_pairs = [("timeout", "ack"), ("error", "done"), ...]
90
- findings = []
91
- for e1, e2 in event_pairs:
92
- # Path 1: e1 then e2
93
- s1 = step(init, [e1, e2])
94
- # Path 2: e2 then e1
95
- s2 = step(init, [e2, e1])
96
- if s1 != s2:
97
- findings.append({
98
- "events": (e1, e2),
99
- "order_e1_e2_ends_in": s1,
100
- "order_e2_e1_ends_in": s2,
101
- "risk": "Order-dependent outcome not documented"
102
- })
103
- return findings
104
- ```
105
-
106
- ### A3: Order Permutation
107
-
108
- Extends A2 to N events. For small N (≤5), do full permutation. For larger N, sample:
109
-
110
- ```python
111
- from itertools import permutations
112
-
113
- def order_permutation_probe(states, init, events, invariant_check=None):
114
- findings = []
115
- terminals = set()
116
- for perm in permutations(events):
117
- final = step(init, list(perm))
118
- terminals.add(final)
119
- if invariant_check and not invariant_check(final):
120
- findings.append({
121
- "sequence": list(perm),
122
- "final_state": final,
123
- "violates": "invariant"
124
- })
125
- if len(terminals) > 1:
126
- findings.append({
127
- "terminal_states": list(terminals),
128
- "risk": f"Same events produce {len(terminals)} different outcomes depending on order"
129
- })
130
- return findings
131
- ```
132
-
133
- ### A4: Pair Symmetry
134
-
135
- ```python
136
- # Define paired operations
137
- PAIRS = [
138
- ("lock", "unlock"),
139
- ("start", "stop"),
140
- ("alloc", "free"),
141
- ("enable_irq", "disable_irq"),
142
- ("open", "close"),
143
- ]
144
-
145
- def pair_symmetry_probe(states, init):
146
- findings = []
147
- for acquire, release in PAIRS:
148
- if acquire not in all_events(states) and release not in all_events(states):
149
- continue # This pair type is not used
150
- # DFS from init: every path that calls acquire must eventually call release
151
- # before reaching a terminal state (or acquire again)
152
- unbalanced = find_unbalanced_paths(states, init, acquire, release)
153
- if unbalanced:
154
- findings.append({
155
- "pair": (acquire, release),
156
- "unbalanced_paths": unbalanced,
157
- "risk": "Resource leak or deadlock"
158
- })
159
- return findings
160
- ```
161
-
162
- ### A5: Boundary Blast
163
-
164
- ```python
165
- def boundary_blast(counters, timestamps):
166
- findings = []
167
- for name, val in counters.items():
168
- test_values = [0, 1, val-1, val, val+1, 2**32-1, 2**32]
169
- for tv in test_values:
170
- if tv < 0 or tv > val:
171
- findings.append({
172
- "variable": name,
173
- "value": tv,
174
- "risk": f"Counter overflow/underflow at {tv} (max defined: {val})"
175
- })
176
- # Timestamp wraparound (millis() / sys_tick())
177
- # On 32-bit ARM: wraparound at ~49.7 days for 1ms tick
178
- for name, val in timestamps.items():
179
- findings.append({
180
- "variable": name,
181
- "risk": f"Timestamp wraparound not handled — check elapsed_ms() / elapsed_ticks() pattern"
182
- })
183
- return findings
184
- ```
185
-
186
- ### A6: Resource Injection
187
-
188
- ```python
189
- RESOURCE_FAILURES = {
190
- "malloc": "returns NULL",
191
- "queue_send": "queue full",
192
- "semaphore_take": "timeout",
193
- "message_alloc": "pool exhausted",
194
- }
195
-
196
- def resource_injection_probe(states):
197
- findings = []
198
- for state, transitions in states.items():
199
- for event in transitions:
200
- for resource, failure in RESOURCE_FAILURES.items():
201
- if resource in event.lower() or resource in state.lower():
202
- # Simulate: what if this resource call fails in this state?
203
- # Does the machine have a recovery transition?
204
- has_recovery = any(
205
- "error" in t.lower() or "fail" in t.lower() or "retry" in t.lower()
206
- for t in transitions.values()
207
- )
208
- if not has_recovery:
209
- findings.append({
210
- "state": state,
211
- "resource": resource,
212
- "failure_mode": failure,
213
- "risk": f"No recovery path if {resource} {failure} in state {state}"
214
- })
215
- return findings
216
- ```
217
-
218
- ### A7: Minimal Counter-Example
219
-
220
- ```python
221
- from collections import deque
222
-
223
- def shortest_violating_path(states, init, invariant_check):
224
- """BFS to find the shortest event sequence that violates an invariant."""
225
- queue = deque([(init, [])])
226
- visited = set()
227
- while queue:
228
- state, path = queue.popleft()
229
- if state in visited:
230
- continue
231
- visited.add(state)
232
-
233
- if not invariant_check(state):
234
- return path # Shortest violating path found
235
-
236
- for event, next_state in states.get(state, {}).items():
237
- queue.append((next_state, path + [event]))
238
-
239
- return None # Invariant holds for all reachable states
240
- ```
241
-
242
- ## Counter-Example Interpretation
243
-
244
- When a probe finds a counter-example, classify it:
245
-
246
- ### True Positive (plan has a real gap)
247
-
248
- - The counter-example uses only events/states defined in the plan
249
- - The plan's own rules would agree this is a violation
250
- - **Action**: Report as a finding with severity based on impact
251
-
252
- ### Model Error (extraction was wrong)
253
-
254
- - The counter-example depends on a transition the plan doesn't actually define
255
- - The plan explicitly handles this case but the model extraction missed it
256
- - **Action**: Fix the model and re-run; do NOT report as a finding
257
-
258
- ### Acceptable Risk (gap is intentional)
259
-
260
- - The plan acknowledges the gap explicitly ("not handled, system resets")
261
- - The counter-example requires physically impossible event sequences
262
- - **Action**: Note in findings but mark as "acknowledged — no fix needed"
263
-
264
- ## Known Limitations
265
-
266
- ### What This Method CAN Detect
267
-
268
- - Missing transitions and states
269
- - Deadlock and livelock
270
- - Violated invariants (the plan's own claims disproven)
271
- - Race conditions between documented events
272
- - Unsymmetrical resource pairs
273
- - Counter/timer boundary issues
274
-
275
- ### What This Method CANNOT Detect
276
-
277
- - **Implementation bugs**: The C code may have errors not present in the model
278
- - **Timing-dependent bugs**: Python model does not simulate real-time constraints
279
- - **Compiler/optimization issues**: Volatile omission, reordering, inlining effects
280
- - **Hardware-specific behavior**: Memory-mapped I/O timing, DMA races, cache coherency
281
- - **Undocumented behavior**: If the plan doesn't describe a transition, the model can't either
282
- - **Real concurrency**: The model is single-threaded; true preemptive multitasking bugs are out of scope
283
- - **Entry/exit actions**: State entry/exit side effects (e.g., `lock()` on enter, `unlock()` on exit) are not modeled as events. A4 Pair Symmetry may miss unbalanced pairs that exist only in entry/exit actions. If the plan describes these, manually extract them as pseudo-events (`ENTER_state`, `EXIT_state`) before running verification.
284
- - **Guard variable scope**: The model treats guard variables as global to the machine. If a guard variable's lifetime is state-scoped (reset on entry) but the model assumes it accumulates globally, boundary-blast results will be wrong. Verify guard variable scope during extraction.
285
- - **Nested/hierarchical states (Harel statecharts)**: The model only supports flat state machines. Parent/child state nesting, history pseudostates, and orthogonal regions are not supported — flatten them manually before verification.
286
- - **Cross-machine protocols**: Two interacting state machines are verified independently. Composition bugs (e.g., Machine A sends event E to Machine B, but B is in a state that doesn't handle E) are invisible to single-machine verification. If the plan describes multi-machine interaction, document the protocol contract separately.
287
-
288
- ### Model Fidelity Warning
289
-
290
- The Python model is an APPROXIMATION. It models state transitions, not execution semantics. A model that passes all 14 checks means the plan's LOGIC is consistent — NOT that the implementation will work. Always follow logic verification with code-level review.
291
-
292
- ## Refactoring Verification
293
-
294
- When the document is a refactoring plan (modifying existing logic, not greenfield design), the pipeline adapts to compare BEFORE and AFTER models.
295
-
296
- ### Extraction for Refactoring
297
-
298
- The BEFORE model comes from **code, not the plan**. The plan may describe the current state inaccurately — verify against the actual implementation:
299
-
300
- 1. Read the relevant source files (state machine dispatch, state enum, handler functions)
301
- 2. Extract the ACTUAL transition logic from code, not from the plan's description of "current behavior"
302
- 3. Extract the AFTER model from the plan as usual
303
- 4. Display both tables side by side
304
-
305
- ### Comparison Methodology
306
-
307
- **Behavioral preservation**: For each event sequence accepted by BEFORE, trace the same sequence in AFTER. If AFTER ends in a different state (or rejects the sequence), flag as BEHAVIORAL DELTA. The plan must explicitly document this delta — if it doesn't, it's a regression.
308
-
309
- **Invariant continuity**: Re-run S7 invariant checks on both models. Any invariant that passes on BEFORE but fails on AFTER is a regression — the refactoring broke an existing guarantee.
310
-
311
- **Complexity delta**: Count objectively:
312
-
313
- - Number of states (BEFORE vs AFTER)
314
- - Number of transitions
315
- - Number of guard conditions
316
- - Maximum path length from INIT to any terminal state
317
-
318
- If the plan claims "simplification" but the numbers don't decrease, flag as UNSUBSTANTIATED CLAIM.
319
-
320
- **Deadlock regression**: A refactoring that splits one state into two should not introduce a deadlock path that didn't exist before. Run S2 on AFTER and compare to BEFORE's deadlock report.
321
-
322
- ### Common Refactoring Bugs Caught by This Method
323
-
324
- - **State split introduces dead edge**: Splitting ERROR into ERROR_TRANSIENT and ERROR_PERMANENT creates a new state with no transition from ERROR_TRANSIENT back to RECOVERING
325
- - **Guard inversion**: Changing `retry < 3` to `retry <= 3` silently adds one more retry — the plan doesn't mention it
326
- - **Orphaned event**: A transition removed in the refactoring was the only path that handled `power_lost` — the event is now silently ignored in some states
327
- - **False simplification**: The plan claims "simplified error handling" but merges two states with different recovery paths into one, losing the distinction
328
-
329
- ## Manual Verification Mode
330
-
331
- When Python is NOT available (air-gapped embedded dev machine, locked-down Windows), execute each check manually. The agent performs the verification using its own reasoning — the methodology is identical, only the execution engine changes.
332
-
333
- ### Size Limit
334
-
335
- Manual verification is reliable for state machines with **≤ 10 states and ≤ 30 transitions**. For larger machines, manual BFS and cycle detection become error-prone. If the machine exceeds this threshold and Python is unavailable, either:
336
-
337
- - Decompose the machine into sub-machines and verify each independently, then check cross-machine contracts manually
338
- - Flag the size limitation as a finding and recommend the user run the Python harness offline
339
-
340
- ### Prerequisites
341
-
342
- - Transition table has been extracted and confirmed with the user
343
- - You have the full table in context (from Phase 2 extraction step)
344
- - **Harness validation** (Python mode only): After filling in `verification-harness.py`, translate the Python `STATES` dict BACK into a transition table and compare it against the confirmed extraction table. If they differ, fix the harness. This catches typo and whitespace errors in manual dict construction.
345
-
346
- ### Phase 2a — Manual Structural Checks
347
-
348
- **S1 Reachability**: Start from INIT. For each state, ask: "Is there any sequence of events from INIT that reaches this state?" Mark each state as reachable or unreachable. Unreachable states are dead code — flag them.
349
-
350
- **S2 Deadlock**: For each non-terminal state, count outgoing transitions. If count = 0, flag as DEADLOCK. Terminal states are exempt.
351
-
352
- **S3 Liveness**: Look for cycles in the transition graph that have no exit to a terminal/recovery state. Example: ERROR → RECOVERING → ERROR forms a cycle. If RECOVERING has a transition to IDLE, it's fine. If every transition from the cycle stays in the cycle with no terminal exit, flag as LIVELOCK.
353
-
354
- **S4 Determinism**: For each state, group transitions by base event name (strip guard suffixes). If two transitions share the same base event and have different targets without different guard conditions, flag as AMBIGUOUS.
355
-
356
- **S5 Event Completeness**: Collect all unique event names. For each non-terminal state, list which events have no defined transition. Flag any state missing handlers for events that other states define — these will be silently ignored at runtime.
357
-
358
- **S6 Guard Completeness**: For each transition with a guard (e.g., `timeout (retry<3)` → RETRY), check if the complementary branch is defined (e.g., `timeout (retry≥3)` → FATAL). If only one branch exists with no else/default, flag as INCOMPLETE GUARD.
359
-
360
- **S7 Invariants**: For each "always/never/guaranteed" claim in the plan, trace every reachable path and verify the claim holds. If you find a counter-example path, quote the plan claim, show the violating path, flag as INVARIANT VIOLATION.
361
-
362
- ### Phase 2b — Manual Adversarial Probes
363
-
364
- **A1 Unexpected Event**: For each state, list all events defined ANYWHERE in the machine that this state does NOT handle. For each unhandled (state, event) pair, assess: is this event physically possible in this state? If yes, what happens? Silence = undefined behavior. Flag if the plan doesn't document the behavior.
365
-
366
- **A2 Race Interleaving**: Identify pairs of events that could arrive in the same tick (e.g., timer expiry + message arrival). Trace both orders through the machine. If the final state differs, flag as RACE — the outcome depends on event ordering.
367
-
368
- **A3 Order Permutation**: Take 3-5 key events and trace all permutations. If different orders lead to different terminal states AND the plan claims independence, flag as ORDER DEPENDENT.
369
-
370
- **A4 Pair Symmetry**: Identify all `lock/unlock`, `start/stop`, `alloc/free` pairs. Check every path: if a path includes `lock` without a subsequent `unlock`, or `alloc` without `free`, flag as ASYMMETRIC. A quick heuristic: if a transition name contains "lock" (or "start", "alloc"), check that every path from its target state eventually hits the corresponding "unlock" before a terminal state.
371
-
372
- **A5 Boundary Blast**: For each counter variable mentioned in guards, test values: 0, 1, max-1, max, max+1. Flag any that would cause overflow or undefined behavior. For timestamps, check if `elapsed = now - start` handles wraparound correctly.
373
-
374
- **A6 Resource Injection**: For each state that appears to allocate (names containing "alloc", "init", "start", "open", "connect", "begin"), check if there is a recovery/error path if the allocation fails. If no recovery exists, flag as RESOURCE VULNERABLE.
375
-
376
- **A7 Minimal Counter-Example**: For each invariant that failed in S7, find the SHORTEST event sequence that violates it. This is the most useful output — it gives the plan author an exact repro.
377
-
378
- ### Manual Mode Output Format
379
-
380
- For each finding, output the same structured format as the Python harness:
381
-
382
- ```text
383
- [CHECK] S1 Reachability
384
- RESULT: PASS — All 7 states reachable from INIT
385
-
386
- [CHECK] S3 Liveness
387
- RESULT: FAIL — Absorbing cycle: ERROR → RECOVERING → ERROR
388
- EVIDENCE: RECOVERING has only one transition (→ ERROR after cooldown).
389
- No path from RECOVERING back to IDLE exists.
390
- SEVERITY: Error — machine can never recover from ERROR state.
391
- ```
392
-
393
- Manual verification takes longer but produces identical-quality findings. The key advantage: it requires zero dependencies and works on any platform.
394
-
395
- ## Integration with the Review Pipeline
396
-
397
- ```text
398
- 1. Load logicprobe skill
399
- 2. Read plan → Phase 1 (enumerate claims)
400
- 3. Detect behavioral claims → trigger logic-primitive escalation
401
- 4. Load this guide
402
- 5. Check Python: `python3 --version` or `python --version`
403
- 6a. Python ≥ 3.6 → load verification-harness.py → fill in MODEL → run
404
- 6b. No Python → use Manual Verification Mode (see above) — execute each check step by step
405
- 7. Extract model → show transition table → GET USER CONFIRMATION
406
- 8. Run Phase 2a (7 structural primitives) → log results
407
- 9. Run Phase 2b (7 adversarial probes) → log results
408
- 10. Classify counter-examples (true positive / model error / acceptable risk)
409
- 11. Feed confirmed findings into Phase 3 (gap analysis)
410
- 12. Include in Phase 5 (structured output)
411
- ```
412
-
413
- Never skip step 5. A verified model based on wrong extraction is worse than no verification — it creates false confidence.
1
+ # Logic Verification Guide
2
+
3
+ Deep reference for Phase 2a/2b of logicprobe. Load this when the skill triggers escalation to logic-primitive verification. Covers model extraction methodology, probe design patterns, counter-example interpretation, and known limitations.
4
+
5
+ ## Model Extraction Methodology
6
+
7
+ ### Step 1: Identify the State Machine Boundary
8
+
9
+ Before extracting states, determine what IS and IS NOT part of the machine:
10
+
11
+ - **In**: states the plan explicitly names, events the plan describes, guards the plan states
12
+ - **Out**: hardware-level behavior (register writes, DMA transfers), OS-level scheduling (task switches, priority inversion), external system interactions not described in the plan
13
+
14
+ If the plan doesn't clearly define the boundary, flag it: "Plan does not define state machine scope — model may be incomplete."
15
+
16
+ ### Step 2: Extract States
17
+
18
+ Scan the plan for:
19
+
20
+ - Enumerated state types (`typedef enum { ... } xxx_state_t`)
21
+ - Named phases in prose ("the system then enters the RECOVERING phase")
22
+ - Implicit states (error paths described but not named — give them explicit names in the model)
23
+
24
+ **Rule**: If a behavioral section describes a distinct set of actions before a transition, it's a state — even if the plan doesn't label it as one.
25
+
26
+ ### Step 3: Extract Transitions
27
+
28
+ For each state, list:
29
+
30
+ - **Explicit transitions**: `state = NEXT_STATE` assignments, `→` arrows in diagrams, prose like "then transitions to X"
31
+ - **Guard conditions**: `if/else` branches that split a single event into multiple next states
32
+ - **Timer-driven transitions**: timeouts, retry delays, cooldown periods
33
+
34
+ **Edge case**: A transition mentioned in prose but missing from the state diagram. Flag it — "Plan describes transition X but state diagram omits it."
35
+
36
+ ### Step 4: Identify Invariants
37
+
38
+ Extract every claim that uses absolute language:
39
+
40
+ | Pattern | Example | Invariant to Check |
41
+ |---------|---------|-------------------|
42
+ | "always ..." | "always returns to IDLE" | From every state reachable in ≤K steps, IDLE is reachable |
43
+ | "never ..." | "never enters ACTIVE without power_ready" | ACTIVE is not reachable from any path that hasn't passed through power_ready |
44
+ | "guaranteed ..." | "guaranteed cleanup within 500ms" | Every ERROR→RECOVERING path includes a cooldown transition |
45
+ | "cannot ..." | "cannot deadlock" | No absorbing cycle exists |
46
+ | "all paths ..." | "all paths lead to ERROR on failure" | Every failure event reaches ERROR (not FATAL, not stuck) |
47
+
48
+ ### Step 5: Confirm with User
49
+
50
+ Show the extracted transition table BEFORE writing the harness. Ask:
51
+
52
+ 1. Are all states captured?
53
+ 2. Are all transitions and guards correct?
54
+ 3. Are there undocumented transitions not in the plan?
55
+ 4. Is the initial state correct?
56
+
57
+ Only proceed after confirmation. A wrong model produces wrong counter-examples, which wastes more time than no verification at all.
58
+
59
+ ## Probe Design Patterns
60
+
61
+ ### A1: Unexpected Event Injection
62
+
63
+ ```python
64
+ # For each state S, list all events defined anywhere in the machine
65
+ # For each event E not handled by S, test: what happens?
66
+ def unexpected_event_probe(states):
67
+ all_events = set()
68
+ for trans in states.values():
69
+ all_events.update(trans.keys())
70
+
71
+ findings = []
72
+ for state, transitions in states.items():
73
+ unhandled = all_events - set(transitions.keys())
74
+ if unhandled:
75
+ findings.append({
76
+ "state": state,
77
+ "unhandled_events": list(unhandled),
78
+ "risk": "Silent ignore or undefined behavior"
79
+ })
80
+ return findings
81
+ ```
82
+
83
+ ### A2: Race Interleaving
84
+
85
+ For pairs of events that can arrive within the same tick (e.g., timer expiry + message reception):
86
+
87
+ ```python
88
+ def race_interleaving(states, init, event_pairs):
89
+ # event_pairs = [("timeout", "ack"), ("error", "done"), ...]
90
+ findings = []
91
+ for e1, e2 in event_pairs:
92
+ # Path 1: e1 then e2
93
+ s1 = step(init, [e1, e2])
94
+ # Path 2: e2 then e1
95
+ s2 = step(init, [e2, e1])
96
+ if s1 != s2:
97
+ findings.append({
98
+ "events": (e1, e2),
99
+ "order_e1_e2_ends_in": s1,
100
+ "order_e2_e1_ends_in": s2,
101
+ "risk": "Order-dependent outcome not documented"
102
+ })
103
+ return findings
104
+ ```
105
+
106
+ ### A3: Order Permutation
107
+
108
+ Extends A2 to N events. For small N (≤5), do full permutation. For larger N, sample:
109
+
110
+ ```python
111
+ from itertools import permutations
112
+
113
+ def order_permutation_probe(states, init, events, invariant_check=None):
114
+ findings = []
115
+ terminals = set()
116
+ for perm in permutations(events):
117
+ final = step(init, list(perm))
118
+ terminals.add(final)
119
+ if invariant_check and not invariant_check(final):
120
+ findings.append({
121
+ "sequence": list(perm),
122
+ "final_state": final,
123
+ "violates": "invariant"
124
+ })
125
+ if len(terminals) > 1:
126
+ findings.append({
127
+ "terminal_states": list(terminals),
128
+ "risk": f"Same events produce {len(terminals)} different outcomes depending on order"
129
+ })
130
+ return findings
131
+ ```
132
+
133
+ ### A4: Pair Symmetry
134
+
135
+ ```python
136
+ # Define paired operations
137
+ PAIRS = [
138
+ ("lock", "unlock"),
139
+ ("start", "stop"),
140
+ ("alloc", "free"),
141
+ ("enable_irq", "disable_irq"),
142
+ ("open", "close"),
143
+ ]
144
+
145
+ def pair_symmetry_probe(states, init):
146
+ findings = []
147
+ for acquire, release in PAIRS:
148
+ if acquire not in all_events(states) and release not in all_events(states):
149
+ continue # This pair type is not used
150
+ # DFS from init: every path that calls acquire must eventually call release
151
+ # before reaching a terminal state (or acquire again)
152
+ unbalanced = find_unbalanced_paths(states, init, acquire, release)
153
+ if unbalanced:
154
+ findings.append({
155
+ "pair": (acquire, release),
156
+ "unbalanced_paths": unbalanced,
157
+ "risk": "Resource leak or deadlock"
158
+ })
159
+ return findings
160
+ ```
161
+
162
+ ### A5: Boundary Blast
163
+
164
+ ```python
165
+ def boundary_blast(counters, timestamps):
166
+ findings = []
167
+ for name, val in counters.items():
168
+ test_values = [0, 1, val-1, val, val+1, 2**32-1, 2**32]
169
+ for tv in test_values:
170
+ if tv < 0 or tv > val:
171
+ findings.append({
172
+ "variable": name,
173
+ "value": tv,
174
+ "risk": f"Counter overflow/underflow at {tv} (max defined: {val})"
175
+ })
176
+ # Timestamp wraparound (millis() / sys_tick())
177
+ # On 32-bit ARM: wraparound at ~49.7 days for 1ms tick
178
+ for name, val in timestamps.items():
179
+ findings.append({
180
+ "variable": name,
181
+ "risk": f"Timestamp wraparound not handled — check elapsed_ms() / elapsed_ticks() pattern"
182
+ })
183
+ return findings
184
+ ```
185
+
186
+ ### A6: Resource Injection
187
+
188
+ ```python
189
+ RESOURCE_FAILURES = {
190
+ "malloc": "returns NULL",
191
+ "queue_send": "queue full",
192
+ "semaphore_take": "timeout",
193
+ "message_alloc": "pool exhausted",
194
+ }
195
+
196
+ def resource_injection_probe(states):
197
+ findings = []
198
+ for state, transitions in states.items():
199
+ for event in transitions:
200
+ for resource, failure in RESOURCE_FAILURES.items():
201
+ if resource in event.lower() or resource in state.lower():
202
+ # Simulate: what if this resource call fails in this state?
203
+ # Does the machine have a recovery transition?
204
+ has_recovery = any(
205
+ "error" in t.lower() or "fail" in t.lower() or "retry" in t.lower()
206
+ for t in transitions.values()
207
+ )
208
+ if not has_recovery:
209
+ findings.append({
210
+ "state": state,
211
+ "resource": resource,
212
+ "failure_mode": failure,
213
+ "risk": f"No recovery path if {resource} {failure} in state {state}"
214
+ })
215
+ return findings
216
+ ```
217
+
218
+ ### A7: Minimal Counter-Example
219
+
220
+ ```python
221
+ from collections import deque
222
+
223
+ def shortest_violating_path(states, init, invariant_check):
224
+ """BFS to find the shortest event sequence that violates an invariant."""
225
+ queue = deque([(init, [])])
226
+ visited = set()
227
+ while queue:
228
+ state, path = queue.popleft()
229
+ if state in visited:
230
+ continue
231
+ visited.add(state)
232
+
233
+ if not invariant_check(state):
234
+ return path # Shortest violating path found
235
+
236
+ for event, next_state in states.get(state, {}).items():
237
+ queue.append((next_state, path + [event]))
238
+
239
+ return None # Invariant holds for all reachable states
240
+ ```
241
+
242
+ ### S8: Monotonic Variables
243
+
244
+ For counters or progress variables that must only move in one direction, list the events that increase/decrease them. Any event that moves against the declared direction is a finding.
245
+
246
+ ```python
247
+ MONOTONIC_VARS = [
248
+ {"name": "retry_count", "direction": "inc", "increase_events": ["retry"], "decrease_events": []},
249
+ ]
250
+ ```
251
+
252
+ ### A8: Idempotent Replay
253
+
254
+ For events that must be safe to retry/replay, apply the event twice from every reachable state. If the second application changes state or is not possible, flag it.
255
+
256
+ ```python
257
+ IDEMPOTENT_EVENTS = {"retry", "sync", "webhook_delivery"}
258
+ ```
259
+
260
+ ### A9: Leads-To
261
+
262
+ For progress claims ("from MIGRATING, eventually DONE"), check every path from the source state. If a path dead-ends or loops before reaching the target, flag it.
263
+
264
+ ```python
265
+ LEADS_TO = [
266
+ ("MIGRATING", "DONE"),
267
+ ]
268
+ ```
269
+
270
+ ### A10: Sequence Order
271
+
272
+ For ordered event sequences ("backup before modify before commit"), search for any path where a later event occurs before an earlier one.
273
+
274
+ ```python
275
+ SEQUENCES = [
276
+ ["backup", "modify", "commit"],
277
+ ]
278
+ ```
279
+
280
+ ### A11: Atomicity
281
+
282
+ For all-or-nothing groups, track whether an atomic event has started and whether commit/rollback has occurred. Leaving the atomic scope or reaching a terminal state before commit/rollback is a violation.
283
+
284
+ ```python
285
+ ATOMIC_GROUPS = [
286
+ {"events": ["write"], "commit": "commit", "rollback": "rollback"},
287
+ ]
288
+ ```
289
+
290
+ ## Counter-Example Interpretation
291
+
292
+ When a probe finds a counter-example, classify it:
293
+
294
+ ### True Positive (plan has a real gap)
295
+
296
+ - The counter-example uses only events/states defined in the plan
297
+ - The plan's own rules would agree this is a violation
298
+ - **Action**: Report as a finding with severity based on impact
299
+
300
+ ### Model Error (extraction was wrong)
301
+
302
+ - The counter-example depends on a transition the plan doesn't actually define
303
+ - The plan explicitly handles this case but the model extraction missed it
304
+ - **Action**: Fix the model and re-run; do NOT report as a finding
305
+
306
+ ### Acceptable Risk (gap is intentional)
307
+
308
+ - The plan acknowledges the gap explicitly ("not handled, system resets")
309
+ - The counter-example requires physically impossible event sequences
310
+ - **Action**: Note in findings but mark as "acknowledged — no fix needed"
311
+
312
+ ## Known Limitations
313
+
314
+ ### What This Method CAN Detect
315
+
316
+ - Missing transitions and states
317
+ - Deadlock and livelock
318
+ - Violated invariants (the plan's own claims disproven)
319
+ - Race conditions between documented events
320
+ - Unsymmetrical resource pairs
321
+ - Counter/timer boundary issues
322
+
323
+ ### What This Method CANNOT Detect
324
+
325
+ - **Implementation bugs**: The C code may have errors not present in the model
326
+ - **Timing-dependent bugs**: Python model does not simulate real-time constraints
327
+ - **Compiler/optimization issues**: Volatile omission, reordering, inlining effects
328
+ - **Hardware-specific behavior**: Memory-mapped I/O timing, DMA races, cache coherency
329
+ - **Undocumented behavior**: If the plan doesn't describe a transition, the model can't either
330
+ - **Real concurrency**: The model is single-threaded; true preemptive multitasking bugs are out of scope
331
+ - **Entry/exit actions**: State entry/exit side effects (e.g., `lock()` on enter, `unlock()` on exit) are not modeled as events. A4 Pair Symmetry may miss unbalanced pairs that exist only in entry/exit actions. If the plan describes these, manually extract them as pseudo-events (`ENTER_state`, `EXIT_state`) before running verification.
332
+ - **Guard variable scope**: The model treats guard variables as global to the machine. If a guard variable's lifetime is state-scoped (reset on entry) but the model assumes it accumulates globally, boundary-blast results will be wrong. Verify guard variable scope during extraction.
333
+ - **Nested/hierarchical states (Harel statecharts)**: The model only supports flat state machines. Parent/child state nesting, history pseudostates, and orthogonal regions are not supported — flatten them manually before verification.
334
+ - **Cross-machine protocols**: Two interacting state machines are verified independently. Composition bugs (e.g., Machine A sends event E to Machine B, but B is in a state that doesn't handle E) are invisible to single-machine verification. If the plan describes multi-machine interaction, document the protocol contract separately.
335
+
336
+ ### Model Fidelity Warning
337
+
338
+ The Python model is an APPROXIMATION. It models state transitions, not execution semantics. A model that passes all 14 checks means the plan's LOGIC is consistent — NOT that the implementation will work. Always follow logic verification with code-level review.
339
+
340
+ ## Refactoring Verification
341
+
342
+ When the document is a refactoring plan (modifying existing logic, not greenfield design), the pipeline adapts to compare BEFORE and AFTER models.
343
+
344
+ ### Extraction for Refactoring
345
+
346
+ The BEFORE model comes from **code, not the plan**. The plan may describe the current state inaccurately — verify against the actual implementation:
347
+
348
+ 1. Read the relevant source files (state machine dispatch, state enum, handler functions)
349
+ 2. Extract the ACTUAL transition logic from code, not from the plan's description of "current behavior"
350
+ 3. Extract the AFTER model from the plan as usual
351
+ 4. Display both tables side by side
352
+
353
+ ### Comparison Methodology
354
+
355
+ **Behavioral preservation**: For each event sequence accepted by BEFORE, trace the same sequence in AFTER. If AFTER ends in a different state (or rejects the sequence), flag as BEHAVIORAL DELTA. The plan must explicitly document this delta — if it doesn't, it's a regression.
356
+
357
+ **Invariant continuity**: Re-run S7 invariant checks on both models. Any invariant that passes on BEFORE but fails on AFTER is a regression — the refactoring broke an existing guarantee.
358
+
359
+ **Complexity delta**: Count objectively:
360
+
361
+ - Number of states (BEFORE vs AFTER)
362
+ - Number of transitions
363
+ - Number of guard conditions
364
+ - Maximum path length from INIT to any terminal state
365
+
366
+ If the plan claims "simplification" but the numbers don't decrease, flag as UNSUBSTANTIATED CLAIM.
367
+
368
+ **Deadlock regression**: A refactoring that splits one state into two should not introduce a deadlock path that didn't exist before. Run S2 on AFTER and compare to BEFORE's deadlock report.
369
+
370
+ In the Python harness, set `BEFORE_STATES`, `BEFORE_INIT`, `BEFORE_TERMINALS`, and `STATE_MAPPING` to run D1-D4 automatically after the standard checks.
371
+
372
+ ### Common Refactoring Bugs Caught by This Method
373
+
374
+ - **State split introduces dead edge**: Splitting ERROR into ERROR_TRANSIENT and ERROR_PERMANENT creates a new state with no transition from ERROR_TRANSIENT back to RECOVERING
375
+ - **Guard inversion**: Changing `retry < 3` to `retry <= 3` silently adds one more retry — the plan doesn't mention it
376
+ - **Orphaned event**: A transition removed in the refactoring was the only path that handled `power_lost` — the event is now silently ignored in some states
377
+ - **False simplification**: The plan claims "simplified error handling" but merges two states with different recovery paths into one, losing the distinction
378
+
379
+ ## Manual Verification Mode
380
+
381
+ When Python is NOT available (air-gapped embedded dev machine, locked-down Windows), execute each check manually. The agent performs the verification using its own reasoning — the methodology is identical, only the execution engine changes.
382
+
383
+ ### Size Limit
384
+
385
+ Manual verification is reliable for state machines with **≤ 10 states and ≤ 30 transitions**. For larger machines, manual BFS and cycle detection become error-prone. If the machine exceeds this threshold and Python is unavailable, either:
386
+
387
+ - Decompose the machine into sub-machines and verify each independently, then check cross-machine contracts manually
388
+ - Flag the size limitation as a finding and recommend the user run the Python harness offline
389
+
390
+ ### Prerequisites
391
+
392
+ - Transition table has been extracted and confirmed with the user
393
+ - You have the full table in context (from Phase 2 extraction step)
394
+ - **Harness validation** (Python mode only): After filling in `verification-harness.py`, translate the Python `STATES` dict BACK into a transition table and compare it against the confirmed extraction table. If they differ, fix the harness. This catches typo and whitespace errors in manual dict construction.
395
+
396
+ ### Phase 2a — Manual Structural Checks
397
+
398
+ **S1 Reachability**: Start from INIT. For each state, ask: "Is there any sequence of events from INIT that reaches this state?" Mark each state as reachable or unreachable. Unreachable states are dead code — flag them.
399
+
400
+ **S2 Deadlock**: For each non-terminal state, count outgoing transitions. If count = 0, flag as DEADLOCK. Terminal states are exempt.
401
+
402
+ **S3 Liveness**: Look for cycles in the transition graph that have no exit to a terminal/recovery state. Example: ERROR → RECOVERING → ERROR forms a cycle. If RECOVERING has a transition to IDLE, it's fine. If every transition from the cycle stays in the cycle with no terminal exit, flag as LIVELOCK.
403
+
404
+ **S4 Determinism**: For each state, group transitions by base event name (strip guard suffixes). If two transitions share the same base event and have different targets without different guard conditions, flag as AMBIGUOUS.
405
+
406
+ **S5 Event Completeness**: Collect all unique event names. For each non-terminal state, list which events have no defined transition. Flag any state missing handlers for events that other states define — these will be silently ignored at runtime.
407
+
408
+ **S6 Guard Completeness**: For each transition with a guard (e.g., `timeout (retry<3)` → RETRY), check if the complementary branch is defined (e.g., `timeout (retry≥3)` → FATAL). If only one branch exists with no else/default, flag as INCOMPLETE GUARD.
409
+
410
+ **S7 Invariants**: For each "always/never/guaranteed" claim in the plan, trace every reachable path and verify the claim holds. If you find a counter-example path, quote the plan claim, show the violating path, flag as INVARIANT VIOLATION.
411
+
412
+ ### Phase 2b — Manual Adversarial Probes
413
+
414
+ **A1 Unexpected Event**: For each state, list all events defined ANYWHERE in the machine that this state does NOT handle. For each unhandled (state, event) pair, assess: is this event physically possible in this state? If yes, what happens? Silence = undefined behavior. Flag if the plan doesn't document the behavior.
415
+
416
+ **A2 Race Interleaving**: Identify pairs of events that could arrive in the same tick (e.g., timer expiry + message arrival). Trace both orders through the machine. If the final state differs, flag as RACE — the outcome depends on event ordering.
417
+
418
+ **A3 Order Permutation**: Take 3-5 key events and trace all permutations. If different orders lead to different terminal states AND the plan claims independence, flag as ORDER DEPENDENT.
419
+
420
+ **A4 Pair Symmetry**: Identify all `lock/unlock`, `start/stop`, `alloc/free` pairs. Check every path: if a path includes `lock` without a subsequent `unlock`, or `alloc` without `free`, flag as ASYMMETRIC. A quick heuristic: if a transition name contains "lock" (or "start", "alloc"), check that every path from its target state eventually hits the corresponding "unlock" before a terminal state.
421
+
422
+ **A5 Boundary Blast**: For each counter variable mentioned in guards, test values: 0, 1, max-1, max, max+1. Flag any that would cause overflow or undefined behavior. For timestamps, check if `elapsed = now - start` handles wraparound correctly.
423
+
424
+ **A6 Resource Injection**: For each state that appears to allocate (names containing "alloc", "init", "start", "open", "connect", "begin"), check if there is a recovery/error path if the allocation fails. If no recovery exists, flag as RESOURCE VULNERABLE.
425
+
426
+ **A7 Minimal Counter-Example**: For each invariant that failed in S7, find the SHORTEST event sequence that violates it. This is the most useful output — it gives the plan author an exact repro.
427
+
428
+ ### Manual Mode Output Format
429
+
430
+ For each finding, output the same structured format as the Python harness:
431
+
432
+ ```text
433
+ [CHECK] S1 Reachability
434
+ RESULT: PASS — All 7 states reachable from INIT
435
+
436
+ [CHECK] S3 Liveness
437
+ RESULT: FAIL — Absorbing cycle: ERROR → RECOVERING → ERROR
438
+ EVIDENCE: RECOVERING has only one transition (→ ERROR after cooldown).
439
+ No path from RECOVERING back to IDLE exists.
440
+ SEVERITY: Error — machine can never recover from ERROR state.
441
+ ```
442
+
443
+ Manual verification takes longer but produces identical-quality findings. The key advantage: it requires zero dependencies and works on any platform.
444
+
445
+ ## Integration with the Review Pipeline
446
+
447
+ ```text
448
+ 1. Load logicprobe skill
449
+ 2. Read plan → Phase 1 (enumerate claims)
450
+ 3. Detect behavioral claims → trigger logic-primitive escalation
451
+ 4. Load this guide
452
+ 5. Check Python: `python3 --version` or `python --version`
453
+ 6a. Python ≥ 3.6 → load verification-harness.py → fill in MODEL → run
454
+ 6b. No Python → use Manual Verification Mode (see above) — execute each check step by step
455
+ 7. Extract model → show transition table → GET USER CONFIRMATION
456
+ 8. Run Phase 2a (7 structural primitives) → log results
457
+ 9. Run Phase 2b (7 adversarial probes) → log results
458
+ 10. Classify counter-examples (true positive / model error / acceptable risk)
459
+ 11. Feed confirmed findings into Phase 3 (gap analysis)
460
+ 12. Include in Phase 5 (structured output)
461
+ ```
462
+
463
+ Never skip step 5. A verified model based on wrong extraction is worse than no verification — it creates false confidence.