dsh-logicprobe 0.4.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.en-US.md +7 -1
- package/README.md +7 -1
- package/lib/concurrency-tool.js +34 -0
- package/lib/concurrency.js +76 -0
- package/lib/data-engine.js +930 -0
- package/lib/data-tool.js +61 -0
- package/lib/engine.js +1622 -1364
- package/lib/index.js +291 -277
- package/lib/tool.js +57 -57
- package/lib/types/concurrency-tool.d.ts +8 -0
- package/lib/types/concurrency.d.ts +20 -0
- package/lib/types/data-engine.d.ts +199 -0
- package/lib/types/data-tool.d.ts +10 -0
- package/lib/types/engine.d.ts +171 -151
- package/package.json +79 -78
- package/skills/logicprobe/SKILL.md +285 -268
- package/skills/logicprobe/references/__pycache__/verification-harness.cpython-312.pyc +0 -0
- package/skills/logicprobe/references/concurrency-risk-guide.md +54 -0
- package/skills/logicprobe/references/dsh-model-schema.md +145 -129
- package/skills/logicprobe/references/logic-verification-guide.md +463 -413
- package/skills/logicprobe/references/verification-harness.py +806 -582
- package/skills/logicprobe-datamodel/SKILL.md +124 -0
- package/skills/logicprobe-datamodel/references/__pycache__/data-model-harness.cpython-312.pyc +0 -0
- package/skills/logicprobe-datamodel/references/data-model-guide.md +62 -0
- package/skills/logicprobe-datamodel/references/data-model-harness.py +528 -0
- package/skills/logicprobe-datamodel/references/data-model-schema.md +128 -0
- package/src/concurrency-tool.ts +37 -0
- package/src/concurrency.ts +102 -0
- package/src/data-engine.ts +1001 -0
- package/src/data-tool.ts +65 -0
- package/src/engine.ts +234 -0
- package/src/index.ts +315 -301
- package/src/tool.ts +60 -60
|
@@ -1,413 +1,463 @@
|
|
|
1
|
-
# Logic Verification Guide
|
|
2
|
-
|
|
3
|
-
Deep reference for Phase 2a/2b of logicprobe. Load this when the skill triggers escalation to logic-primitive verification. Covers model extraction methodology, probe design patterns, counter-example interpretation, and known limitations.
|
|
4
|
-
|
|
5
|
-
## Model Extraction Methodology
|
|
6
|
-
|
|
7
|
-
### Step 1: Identify the State Machine Boundary
|
|
8
|
-
|
|
9
|
-
Before extracting states, determine what IS and IS NOT part of the machine:
|
|
10
|
-
|
|
11
|
-
- **In**: states the plan explicitly names, events the plan describes, guards the plan states
|
|
12
|
-
- **Out**: hardware-level behavior (register writes, DMA transfers), OS-level scheduling (task switches, priority inversion), external system interactions not described in the plan
|
|
13
|
-
|
|
14
|
-
If the plan doesn't clearly define the boundary, flag it: "Plan does not define state machine scope — model may be incomplete."
|
|
15
|
-
|
|
16
|
-
### Step 2: Extract States
|
|
17
|
-
|
|
18
|
-
Scan the plan for:
|
|
19
|
-
|
|
20
|
-
- Enumerated state types (`typedef enum { ... } xxx_state_t`)
|
|
21
|
-
- Named phases in prose ("the system then enters the RECOVERING phase")
|
|
22
|
-
- Implicit states (error paths described but not named — give them explicit names in the model)
|
|
23
|
-
|
|
24
|
-
**Rule**: If a behavioral section describes a distinct set of actions before a transition, it's a state — even if the plan doesn't label it as one.
|
|
25
|
-
|
|
26
|
-
### Step 3: Extract Transitions
|
|
27
|
-
|
|
28
|
-
For each state, list:
|
|
29
|
-
|
|
30
|
-
- **Explicit transitions**: `state = NEXT_STATE` assignments, `→` arrows in diagrams, prose like "then transitions to X"
|
|
31
|
-
- **Guard conditions**: `if/else` branches that split a single event into multiple next states
|
|
32
|
-
- **Timer-driven transitions**: timeouts, retry delays, cooldown periods
|
|
33
|
-
|
|
34
|
-
**Edge case**: A transition mentioned in prose but missing from the state diagram. Flag it — "Plan describes transition X but state diagram omits it."
|
|
35
|
-
|
|
36
|
-
### Step 4: Identify Invariants
|
|
37
|
-
|
|
38
|
-
Extract every claim that uses absolute language:
|
|
39
|
-
|
|
40
|
-
| Pattern | Example | Invariant to Check |
|
|
41
|
-
|---------|---------|-------------------|
|
|
42
|
-
| "always ..." | "always returns to IDLE" | From every state reachable in ≤K steps, IDLE is reachable |
|
|
43
|
-
| "never ..." | "never enters ACTIVE without power_ready" | ACTIVE is not reachable from any path that hasn't passed through power_ready |
|
|
44
|
-
| "guaranteed ..." | "guaranteed cleanup within 500ms" | Every ERROR→RECOVERING path includes a cooldown transition |
|
|
45
|
-
| "cannot ..." | "cannot deadlock" | No absorbing cycle exists |
|
|
46
|
-
| "all paths ..." | "all paths lead to ERROR on failure" | Every failure event reaches ERROR (not FATAL, not stuck) |
|
|
47
|
-
|
|
48
|
-
### Step 5: Confirm with User
|
|
49
|
-
|
|
50
|
-
Show the extracted transition table BEFORE writing the harness. Ask:
|
|
51
|
-
|
|
52
|
-
1. Are all states captured?
|
|
53
|
-
2. Are all transitions and guards correct?
|
|
54
|
-
3. Are there undocumented transitions not in the plan?
|
|
55
|
-
4. Is the initial state correct?
|
|
56
|
-
|
|
57
|
-
Only proceed after confirmation. A wrong model produces wrong counter-examples, which wastes more time than no verification at all.
|
|
58
|
-
|
|
59
|
-
## Probe Design Patterns
|
|
60
|
-
|
|
61
|
-
### A1: Unexpected Event Injection
|
|
62
|
-
|
|
63
|
-
```python
|
|
64
|
-
# For each state S, list all events defined anywhere in the machine
|
|
65
|
-
# For each event E not handled by S, test: what happens?
|
|
66
|
-
def unexpected_event_probe(states):
|
|
67
|
-
all_events = set()
|
|
68
|
-
for trans in states.values():
|
|
69
|
-
all_events.update(trans.keys())
|
|
70
|
-
|
|
71
|
-
findings = []
|
|
72
|
-
for state, transitions in states.items():
|
|
73
|
-
unhandled = all_events - set(transitions.keys())
|
|
74
|
-
if unhandled:
|
|
75
|
-
findings.append({
|
|
76
|
-
"state": state,
|
|
77
|
-
"unhandled_events": list(unhandled),
|
|
78
|
-
"risk": "Silent ignore or undefined behavior"
|
|
79
|
-
})
|
|
80
|
-
return findings
|
|
81
|
-
```
|
|
82
|
-
|
|
83
|
-
### A2: Race Interleaving
|
|
84
|
-
|
|
85
|
-
For pairs of events that can arrive within the same tick (e.g., timer expiry + message reception):
|
|
86
|
-
|
|
87
|
-
```python
|
|
88
|
-
def race_interleaving(states, init, event_pairs):
|
|
89
|
-
# event_pairs = [("timeout", "ack"), ("error", "done"), ...]
|
|
90
|
-
findings = []
|
|
91
|
-
for e1, e2 in event_pairs:
|
|
92
|
-
# Path 1: e1 then e2
|
|
93
|
-
s1 = step(init, [e1, e2])
|
|
94
|
-
# Path 2: e2 then e1
|
|
95
|
-
s2 = step(init, [e2, e1])
|
|
96
|
-
if s1 != s2:
|
|
97
|
-
findings.append({
|
|
98
|
-
"events": (e1, e2),
|
|
99
|
-
"order_e1_e2_ends_in": s1,
|
|
100
|
-
"order_e2_e1_ends_in": s2,
|
|
101
|
-
"risk": "Order-dependent outcome not documented"
|
|
102
|
-
})
|
|
103
|
-
return findings
|
|
104
|
-
```
|
|
105
|
-
|
|
106
|
-
### A3: Order Permutation
|
|
107
|
-
|
|
108
|
-
Extends A2 to N events. For small N (≤5), do full permutation. For larger N, sample:
|
|
109
|
-
|
|
110
|
-
```python
|
|
111
|
-
from itertools import permutations
|
|
112
|
-
|
|
113
|
-
def order_permutation_probe(states, init, events, invariant_check=None):
|
|
114
|
-
findings = []
|
|
115
|
-
terminals = set()
|
|
116
|
-
for perm in permutations(events):
|
|
117
|
-
final = step(init, list(perm))
|
|
118
|
-
terminals.add(final)
|
|
119
|
-
if invariant_check and not invariant_check(final):
|
|
120
|
-
findings.append({
|
|
121
|
-
"sequence": list(perm),
|
|
122
|
-
"final_state": final,
|
|
123
|
-
"violates": "invariant"
|
|
124
|
-
})
|
|
125
|
-
if len(terminals) > 1:
|
|
126
|
-
findings.append({
|
|
127
|
-
"terminal_states": list(terminals),
|
|
128
|
-
"risk": f"Same events produce {len(terminals)} different outcomes depending on order"
|
|
129
|
-
})
|
|
130
|
-
return findings
|
|
131
|
-
```
|
|
132
|
-
|
|
133
|
-
### A4: Pair Symmetry
|
|
134
|
-
|
|
135
|
-
```python
|
|
136
|
-
# Define paired operations
|
|
137
|
-
PAIRS = [
|
|
138
|
-
("lock", "unlock"),
|
|
139
|
-
("start", "stop"),
|
|
140
|
-
("alloc", "free"),
|
|
141
|
-
("enable_irq", "disable_irq"),
|
|
142
|
-
("open", "close"),
|
|
143
|
-
]
|
|
144
|
-
|
|
145
|
-
def pair_symmetry_probe(states, init):
|
|
146
|
-
findings = []
|
|
147
|
-
for acquire, release in PAIRS:
|
|
148
|
-
if acquire not in all_events(states) and release not in all_events(states):
|
|
149
|
-
continue # This pair type is not used
|
|
150
|
-
# DFS from init: every path that calls acquire must eventually call release
|
|
151
|
-
# before reaching a terminal state (or acquire again)
|
|
152
|
-
unbalanced = find_unbalanced_paths(states, init, acquire, release)
|
|
153
|
-
if unbalanced:
|
|
154
|
-
findings.append({
|
|
155
|
-
"pair": (acquire, release),
|
|
156
|
-
"unbalanced_paths": unbalanced,
|
|
157
|
-
"risk": "Resource leak or deadlock"
|
|
158
|
-
})
|
|
159
|
-
return findings
|
|
160
|
-
```
|
|
161
|
-
|
|
162
|
-
### A5: Boundary Blast
|
|
163
|
-
|
|
164
|
-
```python
|
|
165
|
-
def boundary_blast(counters, timestamps):
|
|
166
|
-
findings = []
|
|
167
|
-
for name, val in counters.items():
|
|
168
|
-
test_values = [0, 1, val-1, val, val+1, 2**32-1, 2**32]
|
|
169
|
-
for tv in test_values:
|
|
170
|
-
if tv < 0 or tv > val:
|
|
171
|
-
findings.append({
|
|
172
|
-
"variable": name,
|
|
173
|
-
"value": tv,
|
|
174
|
-
"risk": f"Counter overflow/underflow at {tv} (max defined: {val})"
|
|
175
|
-
})
|
|
176
|
-
# Timestamp wraparound (millis() / sys_tick())
|
|
177
|
-
# On 32-bit ARM: wraparound at ~49.7 days for 1ms tick
|
|
178
|
-
for name, val in timestamps.items():
|
|
179
|
-
findings.append({
|
|
180
|
-
"variable": name,
|
|
181
|
-
"risk": f"Timestamp wraparound not handled — check elapsed_ms() / elapsed_ticks() pattern"
|
|
182
|
-
})
|
|
183
|
-
return findings
|
|
184
|
-
```
|
|
185
|
-
|
|
186
|
-
### A6: Resource Injection
|
|
187
|
-
|
|
188
|
-
```python
|
|
189
|
-
RESOURCE_FAILURES = {
|
|
190
|
-
"malloc": "returns NULL",
|
|
191
|
-
"queue_send": "queue full",
|
|
192
|
-
"semaphore_take": "timeout",
|
|
193
|
-
"message_alloc": "pool exhausted",
|
|
194
|
-
}
|
|
195
|
-
|
|
196
|
-
def resource_injection_probe(states):
|
|
197
|
-
findings = []
|
|
198
|
-
for state, transitions in states.items():
|
|
199
|
-
for event in transitions:
|
|
200
|
-
for resource, failure in RESOURCE_FAILURES.items():
|
|
201
|
-
if resource in event.lower() or resource in state.lower():
|
|
202
|
-
# Simulate: what if this resource call fails in this state?
|
|
203
|
-
# Does the machine have a recovery transition?
|
|
204
|
-
has_recovery = any(
|
|
205
|
-
"error" in t.lower() or "fail" in t.lower() or "retry" in t.lower()
|
|
206
|
-
for t in transitions.values()
|
|
207
|
-
)
|
|
208
|
-
if not has_recovery:
|
|
209
|
-
findings.append({
|
|
210
|
-
"state": state,
|
|
211
|
-
"resource": resource,
|
|
212
|
-
"failure_mode": failure,
|
|
213
|
-
"risk": f"No recovery path if {resource} {failure} in state {state}"
|
|
214
|
-
})
|
|
215
|
-
return findings
|
|
216
|
-
```
|
|
217
|
-
|
|
218
|
-
### A7: Minimal Counter-Example
|
|
219
|
-
|
|
220
|
-
```python
|
|
221
|
-
from collections import deque
|
|
222
|
-
|
|
223
|
-
def shortest_violating_path(states, init, invariant_check):
|
|
224
|
-
"""BFS to find the shortest event sequence that violates an invariant."""
|
|
225
|
-
queue = deque([(init, [])])
|
|
226
|
-
visited = set()
|
|
227
|
-
while queue:
|
|
228
|
-
state, path = queue.popleft()
|
|
229
|
-
if state in visited:
|
|
230
|
-
continue
|
|
231
|
-
visited.add(state)
|
|
232
|
-
|
|
233
|
-
if not invariant_check(state):
|
|
234
|
-
return path # Shortest violating path found
|
|
235
|
-
|
|
236
|
-
for event, next_state in states.get(state, {}).items():
|
|
237
|
-
queue.append((next_state, path + [event]))
|
|
238
|
-
|
|
239
|
-
return None # Invariant holds for all reachable states
|
|
240
|
-
```
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
245
|
-
|
|
246
|
-
|
|
247
|
-
|
|
248
|
-
|
|
249
|
-
|
|
250
|
-
|
|
251
|
-
|
|
252
|
-
###
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
-
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
|
|
307
|
-
|
|
308
|
-
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
-
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
|
|
324
|
-
|
|
325
|
-
- **
|
|
326
|
-
- **
|
|
327
|
-
- **
|
|
328
|
-
|
|
329
|
-
|
|
330
|
-
|
|
331
|
-
|
|
332
|
-
|
|
333
|
-
|
|
334
|
-
|
|
335
|
-
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
|
|
350
|
-
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
|
|
357
|
-
|
|
358
|
-
|
|
359
|
-
|
|
360
|
-
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
368
|
-
**
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
**
|
|
375
|
-
|
|
376
|
-
**
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
399
|
-
|
|
400
|
-
|
|
401
|
-
|
|
402
|
-
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
|
|
1
|
+
# Logic Verification Guide
|
|
2
|
+
|
|
3
|
+
Deep reference for Phase 2a/2b of logicprobe. Load this when the skill triggers escalation to logic-primitive verification. Covers model extraction methodology, probe design patterns, counter-example interpretation, and known limitations.
|
|
4
|
+
|
|
5
|
+
## Model Extraction Methodology
|
|
6
|
+
|
|
7
|
+
### Step 1: Identify the State Machine Boundary
|
|
8
|
+
|
|
9
|
+
Before extracting states, determine what IS and IS NOT part of the machine:
|
|
10
|
+
|
|
11
|
+
- **In**: states the plan explicitly names, events the plan describes, guards the plan states
|
|
12
|
+
- **Out**: hardware-level behavior (register writes, DMA transfers), OS-level scheduling (task switches, priority inversion), external system interactions not described in the plan
|
|
13
|
+
|
|
14
|
+
If the plan doesn't clearly define the boundary, flag it: "Plan does not define state machine scope — model may be incomplete."
|
|
15
|
+
|
|
16
|
+
### Step 2: Extract States
|
|
17
|
+
|
|
18
|
+
Scan the plan for:
|
|
19
|
+
|
|
20
|
+
- Enumerated state types (`typedef enum { ... } xxx_state_t`)
|
|
21
|
+
- Named phases in prose ("the system then enters the RECOVERING phase")
|
|
22
|
+
- Implicit states (error paths described but not named — give them explicit names in the model)
|
|
23
|
+
|
|
24
|
+
**Rule**: If a behavioral section describes a distinct set of actions before a transition, it's a state — even if the plan doesn't label it as one.
|
|
25
|
+
|
|
26
|
+
### Step 3: Extract Transitions
|
|
27
|
+
|
|
28
|
+
For each state, list:
|
|
29
|
+
|
|
30
|
+
- **Explicit transitions**: `state = NEXT_STATE` assignments, `→` arrows in diagrams, prose like "then transitions to X"
|
|
31
|
+
- **Guard conditions**: `if/else` branches that split a single event into multiple next states
|
|
32
|
+
- **Timer-driven transitions**: timeouts, retry delays, cooldown periods
|
|
33
|
+
|
|
34
|
+
**Edge case**: A transition mentioned in prose but missing from the state diagram. Flag it — "Plan describes transition X but state diagram omits it."
|
|
35
|
+
|
|
36
|
+
### Step 4: Identify Invariants
|
|
37
|
+
|
|
38
|
+
Extract every claim that uses absolute language:
|
|
39
|
+
|
|
40
|
+
| Pattern | Example | Invariant to Check |
|
|
41
|
+
|---------|---------|-------------------|
|
|
42
|
+
| "always ..." | "always returns to IDLE" | From every state reachable in ≤K steps, IDLE is reachable |
|
|
43
|
+
| "never ..." | "never enters ACTIVE without power_ready" | ACTIVE is not reachable from any path that hasn't passed through power_ready |
|
|
44
|
+
| "guaranteed ..." | "guaranteed cleanup within 500ms" | Every ERROR→RECOVERING path includes a cooldown transition |
|
|
45
|
+
| "cannot ..." | "cannot deadlock" | No absorbing cycle exists |
|
|
46
|
+
| "all paths ..." | "all paths lead to ERROR on failure" | Every failure event reaches ERROR (not FATAL, not stuck) |
|
|
47
|
+
|
|
48
|
+
### Step 5: Confirm with User
|
|
49
|
+
|
|
50
|
+
Show the extracted transition table BEFORE writing the harness. Ask:
|
|
51
|
+
|
|
52
|
+
1. Are all states captured?
|
|
53
|
+
2. Are all transitions and guards correct?
|
|
54
|
+
3. Are there undocumented transitions not in the plan?
|
|
55
|
+
4. Is the initial state correct?
|
|
56
|
+
|
|
57
|
+
Only proceed after confirmation. A wrong model produces wrong counter-examples, which wastes more time than no verification at all.
|
|
58
|
+
|
|
59
|
+
## Probe Design Patterns
|
|
60
|
+
|
|
61
|
+
### A1: Unexpected Event Injection
|
|
62
|
+
|
|
63
|
+
```python
|
|
64
|
+
# For each state S, list all events defined anywhere in the machine
|
|
65
|
+
# For each event E not handled by S, test: what happens?
|
|
66
|
+
def unexpected_event_probe(states):
|
|
67
|
+
all_events = set()
|
|
68
|
+
for trans in states.values():
|
|
69
|
+
all_events.update(trans.keys())
|
|
70
|
+
|
|
71
|
+
findings = []
|
|
72
|
+
for state, transitions in states.items():
|
|
73
|
+
unhandled = all_events - set(transitions.keys())
|
|
74
|
+
if unhandled:
|
|
75
|
+
findings.append({
|
|
76
|
+
"state": state,
|
|
77
|
+
"unhandled_events": list(unhandled),
|
|
78
|
+
"risk": "Silent ignore or undefined behavior"
|
|
79
|
+
})
|
|
80
|
+
return findings
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
### A2: Race Interleaving
|
|
84
|
+
|
|
85
|
+
For pairs of events that can arrive within the same tick (e.g., timer expiry + message reception):
|
|
86
|
+
|
|
87
|
+
```python
|
|
88
|
+
def race_interleaving(states, init, event_pairs):
|
|
89
|
+
# event_pairs = [("timeout", "ack"), ("error", "done"), ...]
|
|
90
|
+
findings = []
|
|
91
|
+
for e1, e2 in event_pairs:
|
|
92
|
+
# Path 1: e1 then e2
|
|
93
|
+
s1 = step(init, [e1, e2])
|
|
94
|
+
# Path 2: e2 then e1
|
|
95
|
+
s2 = step(init, [e2, e1])
|
|
96
|
+
if s1 != s2:
|
|
97
|
+
findings.append({
|
|
98
|
+
"events": (e1, e2),
|
|
99
|
+
"order_e1_e2_ends_in": s1,
|
|
100
|
+
"order_e2_e1_ends_in": s2,
|
|
101
|
+
"risk": "Order-dependent outcome not documented"
|
|
102
|
+
})
|
|
103
|
+
return findings
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
### A3: Order Permutation
|
|
107
|
+
|
|
108
|
+
Extends A2 to N events. For small N (≤5), do full permutation. For larger N, sample:
|
|
109
|
+
|
|
110
|
+
```python
|
|
111
|
+
from itertools import permutations
|
|
112
|
+
|
|
113
|
+
def order_permutation_probe(states, init, events, invariant_check=None):
|
|
114
|
+
findings = []
|
|
115
|
+
terminals = set()
|
|
116
|
+
for perm in permutations(events):
|
|
117
|
+
final = step(init, list(perm))
|
|
118
|
+
terminals.add(final)
|
|
119
|
+
if invariant_check and not invariant_check(final):
|
|
120
|
+
findings.append({
|
|
121
|
+
"sequence": list(perm),
|
|
122
|
+
"final_state": final,
|
|
123
|
+
"violates": "invariant"
|
|
124
|
+
})
|
|
125
|
+
if len(terminals) > 1:
|
|
126
|
+
findings.append({
|
|
127
|
+
"terminal_states": list(terminals),
|
|
128
|
+
"risk": f"Same events produce {len(terminals)} different outcomes depending on order"
|
|
129
|
+
})
|
|
130
|
+
return findings
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
### A4: Pair Symmetry
|
|
134
|
+
|
|
135
|
+
```python
|
|
136
|
+
# Define paired operations
|
|
137
|
+
PAIRS = [
|
|
138
|
+
("lock", "unlock"),
|
|
139
|
+
("start", "stop"),
|
|
140
|
+
("alloc", "free"),
|
|
141
|
+
("enable_irq", "disable_irq"),
|
|
142
|
+
("open", "close"),
|
|
143
|
+
]
|
|
144
|
+
|
|
145
|
+
def pair_symmetry_probe(states, init):
|
|
146
|
+
findings = []
|
|
147
|
+
for acquire, release in PAIRS:
|
|
148
|
+
if acquire not in all_events(states) and release not in all_events(states):
|
|
149
|
+
continue # This pair type is not used
|
|
150
|
+
# DFS from init: every path that calls acquire must eventually call release
|
|
151
|
+
# before reaching a terminal state (or acquire again)
|
|
152
|
+
unbalanced = find_unbalanced_paths(states, init, acquire, release)
|
|
153
|
+
if unbalanced:
|
|
154
|
+
findings.append({
|
|
155
|
+
"pair": (acquire, release),
|
|
156
|
+
"unbalanced_paths": unbalanced,
|
|
157
|
+
"risk": "Resource leak or deadlock"
|
|
158
|
+
})
|
|
159
|
+
return findings
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
### A5: Boundary Blast
|
|
163
|
+
|
|
164
|
+
```python
|
|
165
|
+
def boundary_blast(counters, timestamps):
|
|
166
|
+
findings = []
|
|
167
|
+
for name, val in counters.items():
|
|
168
|
+
test_values = [0, 1, val-1, val, val+1, 2**32-1, 2**32]
|
|
169
|
+
for tv in test_values:
|
|
170
|
+
if tv < 0 or tv > val:
|
|
171
|
+
findings.append({
|
|
172
|
+
"variable": name,
|
|
173
|
+
"value": tv,
|
|
174
|
+
"risk": f"Counter overflow/underflow at {tv} (max defined: {val})"
|
|
175
|
+
})
|
|
176
|
+
# Timestamp wraparound (millis() / sys_tick())
|
|
177
|
+
# On 32-bit ARM: wraparound at ~49.7 days for 1ms tick
|
|
178
|
+
for name, val in timestamps.items():
|
|
179
|
+
findings.append({
|
|
180
|
+
"variable": name,
|
|
181
|
+
"risk": f"Timestamp wraparound not handled — check elapsed_ms() / elapsed_ticks() pattern"
|
|
182
|
+
})
|
|
183
|
+
return findings
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
### A6: Resource Injection
|
|
187
|
+
|
|
188
|
+
```python
|
|
189
|
+
RESOURCE_FAILURES = {
|
|
190
|
+
"malloc": "returns NULL",
|
|
191
|
+
"queue_send": "queue full",
|
|
192
|
+
"semaphore_take": "timeout",
|
|
193
|
+
"message_alloc": "pool exhausted",
|
|
194
|
+
}
|
|
195
|
+
|
|
196
|
+
def resource_injection_probe(states):
|
|
197
|
+
findings = []
|
|
198
|
+
for state, transitions in states.items():
|
|
199
|
+
for event in transitions:
|
|
200
|
+
for resource, failure in RESOURCE_FAILURES.items():
|
|
201
|
+
if resource in event.lower() or resource in state.lower():
|
|
202
|
+
# Simulate: what if this resource call fails in this state?
|
|
203
|
+
# Does the machine have a recovery transition?
|
|
204
|
+
has_recovery = any(
|
|
205
|
+
"error" in t.lower() or "fail" in t.lower() or "retry" in t.lower()
|
|
206
|
+
for t in transitions.values()
|
|
207
|
+
)
|
|
208
|
+
if not has_recovery:
|
|
209
|
+
findings.append({
|
|
210
|
+
"state": state,
|
|
211
|
+
"resource": resource,
|
|
212
|
+
"failure_mode": failure,
|
|
213
|
+
"risk": f"No recovery path if {resource} {failure} in state {state}"
|
|
214
|
+
})
|
|
215
|
+
return findings
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
### A7: Minimal Counter-Example
|
|
219
|
+
|
|
220
|
+
```python
|
|
221
|
+
from collections import deque
|
|
222
|
+
|
|
223
|
+
def shortest_violating_path(states, init, invariant_check):
|
|
224
|
+
"""BFS to find the shortest event sequence that violates an invariant."""
|
|
225
|
+
queue = deque([(init, [])])
|
|
226
|
+
visited = set()
|
|
227
|
+
while queue:
|
|
228
|
+
state, path = queue.popleft()
|
|
229
|
+
if state in visited:
|
|
230
|
+
continue
|
|
231
|
+
visited.add(state)
|
|
232
|
+
|
|
233
|
+
if not invariant_check(state):
|
|
234
|
+
return path # Shortest violating path found
|
|
235
|
+
|
|
236
|
+
for event, next_state in states.get(state, {}).items():
|
|
237
|
+
queue.append((next_state, path + [event]))
|
|
238
|
+
|
|
239
|
+
return None # Invariant holds for all reachable states
|
|
240
|
+
```
|
|
241
|
+
|
|
242
|
+
### S8: Monotonic Variables
|
|
243
|
+
|
|
244
|
+
For counters or progress variables that must only move in one direction, list the events that increase/decrease them. Any event that moves against the declared direction is a finding.
|
|
245
|
+
|
|
246
|
+
```python
|
|
247
|
+
MONOTONIC_VARS = [
|
|
248
|
+
{"name": "retry_count", "direction": "inc", "increase_events": ["retry"], "decrease_events": []},
|
|
249
|
+
]
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
### A8: Idempotent Replay
|
|
253
|
+
|
|
254
|
+
For events that must be safe to retry/replay, apply the event twice from every reachable state. If the second application changes state or is not possible, flag it.
|
|
255
|
+
|
|
256
|
+
```python
|
|
257
|
+
IDEMPOTENT_EVENTS = {"retry", "sync", "webhook_delivery"}
|
|
258
|
+
```
|
|
259
|
+
|
|
260
|
+
### A9: Leads-To
|
|
261
|
+
|
|
262
|
+
For progress claims ("from MIGRATING, eventually DONE"), check every path from the source state. If a path dead-ends or loops before reaching the target, flag it.
|
|
263
|
+
|
|
264
|
+
```python
|
|
265
|
+
LEADS_TO = [
|
|
266
|
+
("MIGRATING", "DONE"),
|
|
267
|
+
]
|
|
268
|
+
```
|
|
269
|
+
|
|
270
|
+
### A10: Sequence Order
|
|
271
|
+
|
|
272
|
+
For ordered event sequences ("backup before modify before commit"), search for any path where a later event occurs before an earlier one.
|
|
273
|
+
|
|
274
|
+
```python
|
|
275
|
+
SEQUENCES = [
|
|
276
|
+
["backup", "modify", "commit"],
|
|
277
|
+
]
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
### A11: Atomicity
|
|
281
|
+
|
|
282
|
+
For all-or-nothing groups, track whether an atomic event has started and whether commit/rollback has occurred. Leaving the atomic scope or reaching a terminal state before commit/rollback is a violation.
|
|
283
|
+
|
|
284
|
+
```python
|
|
285
|
+
ATOMIC_GROUPS = [
|
|
286
|
+
{"events": ["write"], "commit": "commit", "rollback": "rollback"},
|
|
287
|
+
]
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
## Counter-Example Interpretation
|
|
291
|
+
|
|
292
|
+
When a probe finds a counter-example, classify it:
|
|
293
|
+
|
|
294
|
+
### True Positive (plan has a real gap)
|
|
295
|
+
|
|
296
|
+
- The counter-example uses only events/states defined in the plan
|
|
297
|
+
- The plan's own rules would agree this is a violation
|
|
298
|
+
- **Action**: Report as a finding with severity based on impact
|
|
299
|
+
|
|
300
|
+
### Model Error (extraction was wrong)
|
|
301
|
+
|
|
302
|
+
- The counter-example depends on a transition the plan doesn't actually define
|
|
303
|
+
- The plan explicitly handles this case but the model extraction missed it
|
|
304
|
+
- **Action**: Fix the model and re-run; do NOT report as a finding
|
|
305
|
+
|
|
306
|
+
### Acceptable Risk (gap is intentional)
|
|
307
|
+
|
|
308
|
+
- The plan acknowledges the gap explicitly ("not handled, system resets")
|
|
309
|
+
- The counter-example requires physically impossible event sequences
|
|
310
|
+
- **Action**: Note in findings but mark as "acknowledged — no fix needed"
|
|
311
|
+
|
|
312
|
+
## Known Limitations
|
|
313
|
+
|
|
314
|
+
### What This Method CAN Detect
|
|
315
|
+
|
|
316
|
+
- Missing transitions and states
|
|
317
|
+
- Deadlock and livelock
|
|
318
|
+
- Violated invariants (the plan's own claims disproven)
|
|
319
|
+
- Race conditions between documented events
|
|
320
|
+
- Unsymmetrical resource pairs
|
|
321
|
+
- Counter/timer boundary issues
|
|
322
|
+
|
|
323
|
+
### What This Method CANNOT Detect
|
|
324
|
+
|
|
325
|
+
- **Implementation bugs**: The C code may have errors not present in the model
|
|
326
|
+
- **Timing-dependent bugs**: Python model does not simulate real-time constraints
|
|
327
|
+
- **Compiler/optimization issues**: Volatile omission, reordering, inlining effects
|
|
328
|
+
- **Hardware-specific behavior**: Memory-mapped I/O timing, DMA races, cache coherency
|
|
329
|
+
- **Undocumented behavior**: If the plan doesn't describe a transition, the model can't either
|
|
330
|
+
- **Real concurrency**: The model is single-threaded; true preemptive multitasking bugs are out of scope
|
|
331
|
+
- **Entry/exit actions**: State entry/exit side effects (e.g., `lock()` on enter, `unlock()` on exit) are not modeled as events. A4 Pair Symmetry may miss unbalanced pairs that exist only in entry/exit actions. If the plan describes these, manually extract them as pseudo-events (`ENTER_state`, `EXIT_state`) before running verification.
|
|
332
|
+
- **Guard variable scope**: The model treats guard variables as global to the machine. If a guard variable's lifetime is state-scoped (reset on entry) but the model assumes it accumulates globally, boundary-blast results will be wrong. Verify guard variable scope during extraction.
|
|
333
|
+
- **Nested/hierarchical states (Harel statecharts)**: The model only supports flat state machines. Parent/child state nesting, history pseudostates, and orthogonal regions are not supported — flatten them manually before verification.
|
|
334
|
+
- **Cross-machine protocols**: Two interacting state machines are verified independently. Composition bugs (e.g., Machine A sends event E to Machine B, but B is in a state that doesn't handle E) are invisible to single-machine verification. If the plan describes multi-machine interaction, document the protocol contract separately.
|
|
335
|
+
|
|
336
|
+
### Model Fidelity Warning
|
|
337
|
+
|
|
338
|
+
The Python model is an APPROXIMATION. It models state transitions, not execution semantics. A model that passes all 14 checks means the plan's LOGIC is consistent — NOT that the implementation will work. Always follow logic verification with code-level review.
|
|
339
|
+
|
|
340
|
+
## Refactoring Verification
|
|
341
|
+
|
|
342
|
+
When the document is a refactoring plan (modifying existing logic, not greenfield design), the pipeline adapts to compare BEFORE and AFTER models.
|
|
343
|
+
|
|
344
|
+
### Extraction for Refactoring
|
|
345
|
+
|
|
346
|
+
The BEFORE model comes from **code, not the plan**. The plan may describe the current state inaccurately — verify against the actual implementation:
|
|
347
|
+
|
|
348
|
+
1. Read the relevant source files (state machine dispatch, state enum, handler functions)
|
|
349
|
+
2. Extract the ACTUAL transition logic from code, not from the plan's description of "current behavior"
|
|
350
|
+
3. Extract the AFTER model from the plan as usual
|
|
351
|
+
4. Display both tables side by side
|
|
352
|
+
|
|
353
|
+
### Comparison Methodology
|
|
354
|
+
|
|
355
|
+
**Behavioral preservation**: For each event sequence accepted by BEFORE, trace the same sequence in AFTER. If AFTER ends in a different state (or rejects the sequence), flag as BEHAVIORAL DELTA. The plan must explicitly document this delta — if it doesn't, it's a regression.
|
|
356
|
+
|
|
357
|
+
**Invariant continuity**: Re-run S7 invariant checks on both models. Any invariant that passes on BEFORE but fails on AFTER is a regression — the refactoring broke an existing guarantee.
|
|
358
|
+
|
|
359
|
+
**Complexity delta**: Count objectively:
|
|
360
|
+
|
|
361
|
+
- Number of states (BEFORE vs AFTER)
|
|
362
|
+
- Number of transitions
|
|
363
|
+
- Number of guard conditions
|
|
364
|
+
- Maximum path length from INIT to any terminal state
|
|
365
|
+
|
|
366
|
+
If the plan claims "simplification" but the numbers don't decrease, flag as UNSUBSTANTIATED CLAIM.
|
|
367
|
+
|
|
368
|
+
**Deadlock regression**: A refactoring that splits one state into two should not introduce a deadlock path that didn't exist before. Run S2 on AFTER and compare to BEFORE's deadlock report.
|
|
369
|
+
|
|
370
|
+
In the Python harness, set `BEFORE_STATES`, `BEFORE_INIT`, `BEFORE_TERMINALS`, and `STATE_MAPPING` to run D1-D4 automatically after the standard checks.
|
|
371
|
+
|
|
372
|
+
### Common Refactoring Bugs Caught by This Method
|
|
373
|
+
|
|
374
|
+
- **State split introduces dead edge**: Splitting ERROR into ERROR_TRANSIENT and ERROR_PERMANENT creates a new state with no transition from ERROR_TRANSIENT back to RECOVERING
|
|
375
|
+
- **Guard inversion**: Changing `retry < 3` to `retry <= 3` silently adds one more retry — the plan doesn't mention it
|
|
376
|
+
- **Orphaned event**: A transition removed in the refactoring was the only path that handled `power_lost` — the event is now silently ignored in some states
|
|
377
|
+
- **False simplification**: The plan claims "simplified error handling" but merges two states with different recovery paths into one, losing the distinction
|
|
378
|
+
|
|
379
|
+
## Manual Verification Mode
|
|
380
|
+
|
|
381
|
+
When Python is NOT available (air-gapped embedded dev machine, locked-down Windows), execute each check manually. The agent performs the verification using its own reasoning — the methodology is identical, only the execution engine changes.
|
|
382
|
+
|
|
383
|
+
### Size Limit
|
|
384
|
+
|
|
385
|
+
Manual verification is reliable for state machines with **≤ 10 states and ≤ 30 transitions**. For larger machines, manual BFS and cycle detection become error-prone. If the machine exceeds this threshold and Python is unavailable, either:
|
|
386
|
+
|
|
387
|
+
- Decompose the machine into sub-machines and verify each independently, then check cross-machine contracts manually
|
|
388
|
+
- Flag the size limitation as a finding and recommend the user run the Python harness offline
|
|
389
|
+
|
|
390
|
+
### Prerequisites
|
|
391
|
+
|
|
392
|
+
- Transition table has been extracted and confirmed with the user
|
|
393
|
+
- You have the full table in context (from Phase 2 extraction step)
|
|
394
|
+
- **Harness validation** (Python mode only): After filling in `verification-harness.py`, translate the Python `STATES` dict BACK into a transition table and compare it against the confirmed extraction table. If they differ, fix the harness. This catches typo and whitespace errors in manual dict construction.
|
|
395
|
+
|
|
396
|
+
### Phase 2a — Manual Structural Checks
|
|
397
|
+
|
|
398
|
+
**S1 Reachability**: Start from INIT. For each state, ask: "Is there any sequence of events from INIT that reaches this state?" Mark each state as reachable or unreachable. Unreachable states are dead code — flag them.
|
|
399
|
+
|
|
400
|
+
**S2 Deadlock**: For each non-terminal state, count outgoing transitions. If count = 0, flag as DEADLOCK. Terminal states are exempt.
|
|
401
|
+
|
|
402
|
+
**S3 Liveness**: Look for cycles in the transition graph that have no exit to a terminal/recovery state. Example: ERROR → RECOVERING → ERROR forms a cycle. If RECOVERING has a transition to IDLE, it's fine. If every transition from the cycle stays in the cycle with no terminal exit, flag as LIVELOCK.
|
|
403
|
+
|
|
404
|
+
**S4 Determinism**: For each state, group transitions by base event name (strip guard suffixes). If two transitions share the same base event and have different targets without different guard conditions, flag as AMBIGUOUS.
|
|
405
|
+
|
|
406
|
+
**S5 Event Completeness**: Collect all unique event names. For each non-terminal state, list which events have no defined transition. Flag any state missing handlers for events that other states define — these will be silently ignored at runtime.
|
|
407
|
+
|
|
408
|
+
**S6 Guard Completeness**: For each transition with a guard (e.g., `timeout (retry<3)` → RETRY), check if the complementary branch is defined (e.g., `timeout (retry≥3)` → FATAL). If only one branch exists with no else/default, flag as INCOMPLETE GUARD.
|
|
409
|
+
|
|
410
|
+
**S7 Invariants**: For each "always/never/guaranteed" claim in the plan, trace every reachable path and verify the claim holds. If you find a counter-example path, quote the plan claim, show the violating path, flag as INVARIANT VIOLATION.
|
|
411
|
+
|
|
412
|
+
### Phase 2b — Manual Adversarial Probes
|
|
413
|
+
|
|
414
|
+
**A1 Unexpected Event**: For each state, list all events defined ANYWHERE in the machine that this state does NOT handle. For each unhandled (state, event) pair, assess: is this event physically possible in this state? If yes, what happens? Silence = undefined behavior. Flag if the plan doesn't document the behavior.
|
|
415
|
+
|
|
416
|
+
**A2 Race Interleaving**: Identify pairs of events that could arrive in the same tick (e.g., timer expiry + message arrival). Trace both orders through the machine. If the final state differs, flag as RACE — the outcome depends on event ordering.
|
|
417
|
+
|
|
418
|
+
**A3 Order Permutation**: Take 3-5 key events and trace all permutations. If different orders lead to different terminal states AND the plan claims independence, flag as ORDER DEPENDENT.
|
|
419
|
+
|
|
420
|
+
**A4 Pair Symmetry**: Identify all `lock/unlock`, `start/stop`, `alloc/free` pairs. Check every path: if a path includes `lock` without a subsequent `unlock`, or `alloc` without `free`, flag as ASYMMETRIC. A quick heuristic: if a transition name contains "lock" (or "start", "alloc"), check that every path from its target state eventually hits the corresponding "unlock" before a terminal state.
|
|
421
|
+
|
|
422
|
+
**A5 Boundary Blast**: For each counter variable mentioned in guards, test values: 0, 1, max-1, max, max+1. Flag any that would cause overflow or undefined behavior. For timestamps, check if `elapsed = now - start` handles wraparound correctly.
|
|
423
|
+
|
|
424
|
+
**A6 Resource Injection**: For each state that appears to allocate (names containing "alloc", "init", "start", "open", "connect", "begin"), check if there is a recovery/error path if the allocation fails. If no recovery exists, flag as RESOURCE VULNERABLE.
|
|
425
|
+
|
|
426
|
+
**A7 Minimal Counter-Example**: For each invariant that failed in S7, find the SHORTEST event sequence that violates it. This is the most useful output — it gives the plan author an exact repro.
|
|
427
|
+
|
|
428
|
+
### Manual Mode Output Format
|
|
429
|
+
|
|
430
|
+
For each finding, output the same structured format as the Python harness:
|
|
431
|
+
|
|
432
|
+
```text
|
|
433
|
+
[CHECK] S1 Reachability
|
|
434
|
+
RESULT: PASS — All 7 states reachable from INIT
|
|
435
|
+
|
|
436
|
+
[CHECK] S3 Liveness
|
|
437
|
+
RESULT: FAIL — Absorbing cycle: ERROR → RECOVERING → ERROR
|
|
438
|
+
EVIDENCE: RECOVERING has only one transition (→ ERROR after cooldown).
|
|
439
|
+
No path from RECOVERING back to IDLE exists.
|
|
440
|
+
SEVERITY: Error — machine can never recover from ERROR state.
|
|
441
|
+
```
|
|
442
|
+
|
|
443
|
+
Manual verification takes longer but produces identical-quality findings. The key advantage: it requires zero dependencies and works on any platform.
|
|
444
|
+
|
|
445
|
+
## Integration with the Review Pipeline
|
|
446
|
+
|
|
447
|
+
```text
|
|
448
|
+
1. Load logicprobe skill
|
|
449
|
+
2. Read plan → Phase 1 (enumerate claims)
|
|
450
|
+
3. Detect behavioral claims → trigger logic-primitive escalation
|
|
451
|
+
4. Load this guide
|
|
452
|
+
5. Check Python: `python3 --version` or `python --version`
|
|
453
|
+
6a. Python ≥ 3.6 → load verification-harness.py → fill in MODEL → run
|
|
454
|
+
6b. No Python → use Manual Verification Mode (see above) — execute each check step by step
|
|
455
|
+
7. Extract model → show transition table → GET USER CONFIRMATION
|
|
456
|
+
8. Run Phase 2a (7 structural primitives) → log results
|
|
457
|
+
9. Run Phase 2b (7 adversarial probes) → log results
|
|
458
|
+
10. Classify counter-examples (true positive / model error / acceptable risk)
|
|
459
|
+
11. Feed confirmed findings into Phase 3 (gap analysis)
|
|
460
|
+
12. Include in Phase 5 (structured output)
|
|
461
|
+
```
|
|
462
|
+
|
|
463
|
+
Never skip step 5. A verified model based on wrong extraction is worse than no verification — it creates false confidence.
|