dsh-logicprobe 0.7.1 → 0.8.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,166 @@
1
+ # UML Modelling Guide
2
+
3
+ This guide covers two jobs: drawing a code flow as UML, and reviewing that drawing. Load it when the task is "show me the flow" rather than "check this claim". Typical cases are reverse-engineering a handler, documenting a protocol, or auditing a diagram somebody else drew.
4
+
5
+ One rule runs through the whole guide. **A diagram is a model, and a model can be wrong.** A tidy flow chart is not evidence about the code. Worse, a diagram that disagrees with the model it came from will mislead every reader who trusts it.
6
+
7
+ ## When to use this
8
+
9
+ | Situation | View that answers it |
10
+ |---|---|
11
+ | "What are the states of this handler, and what moves between them?" | State machine |
12
+ | "What does this function actually do, including its error paths?" | Activity (flow) |
13
+ | "In what order do these messages arrive, and what is the state after each?" | Sequence |
14
+ | "Somebody gave me a diagram. Is it right?" | Any view. Feed it to the review and compare it against the code. |
15
+ | "This diagram and this model disagree." | Feed both to the review. The round-trip check names every difference. |
16
+
17
+ Do not use this guide for a behavioural claim about a specific machine. That is the job of `logicprobe_verify` (S1-S8 and A1-A14). The review here covers the modelling, not the behaviour. Every finding it produces names the engine check that settles the behavioural half.
18
+
19
+ ## The three views
20
+
21
+ One machine, three projections. They are not interchangeable. The review reports which view it saw, and the differences matter.
22
+
23
+ **State machine** (`stateDiagram-v2`, or a PlantUML state diagram). This view shows the whole topology: every state, every event-and-guard branch, and every terminal state as `--> [*]`. Model code flow into this view, and use it for review. It is the only view that round-trips without loss.
24
+
25
+ **Activity** (`flowchart TD` in Mermaid). This view draws the same machine as work rather than as states. Each edge carries `event [guard] / actions`. It suits a reader who thinks in steps. It holds the same information as the state view. PlantUML activity is **not** generated: its structured-flowchart syntax needs a while/if reconstruction for any graph with a merge or a cycle. A diagram that quietly reshapes the machine is exactly the failure this feature exists to prevent, so the tool refuses that combination.
26
+
27
+ **Sequence** (`sequenceDiagram`). This view shows one BFS trace: the messages in the order a walker meets them, with a note per state change. A trace is not a machine. Branches appear as separate guarded messages, paths the walk never took are absent, and the trace is capped. Use it to show a protocol exchange to a human. It cannot be the review's fidelity input, because `parse` refuses it and `review` reports the check as inapplicable (`UML019`).
28
+
29
+ ## Modelling a code flow from source
30
+
31
+ Extraction follows the same discipline as `logic-verification-guide.md`, applied to code instead of a plan.
32
+
33
+ 1. **Fix the boundary.** Decide which function, task, or module is the machine. In scope: state the code holds across calls, such as statics, fields, enums, and task state variables. Out of scope: the call stack inside one invocation, hardware behaviour, and scheduler preemption.
34
+ 2. **Name the states from the code.** A state exists if something survives a call boundary and is tested later. An enum, a `state` field, or a task-local variable that gates the next entry all qualify. A local variable inside one function does not.
35
+ 3. **Name the events from the code.** Every edge must trace to a call site, a message, a timer expiry, or an ISR. An edge with no call site is a guess. Put it in the review findings instead of the model.
36
+ 4. **Record guards and actions verbatim.** A guard must be the real condition, with the real variable and the real constant. A retry limit of 3 modelled as "a few" is a model of your assumption, not of the code.
37
+ 5. **Mark the terminals.** An absorbing state is `terminal: true`. Examples are power-off, fatal, and done. If nothing is terminal, say so. The review will flag it either way.
38
+ 6. **Write the narrative as you go.** The `narrative` block holds natural-language meanings for states and events, plus a scenario for each state-and-event pair. It is what lets a reader check the diagram against the code without re-deriving every symbol. Once the block is present, the schema requires all three parts and full coverage, so it cannot rot half-way.
39
+
40
+ ### Evidence rule
41
+
42
+ Every state, event, guard and action needs a citation. Use `file:line` for code and a section reference for a document. A UML model without citations is a drawing. It cannot be reviewed, only admired. Present the citations together with the model.
43
+
44
+ ### Confirming the model
45
+
46
+ The usual gate applies: show the extracted transition table or the diagram, and get confirmation before treating the model as fact. In `logicprobe interaction=auto`, skip the question. Instead, cite evidence for every element, round-trip the model (render, parse, compare), and mark the result `UNCONFIRMED`.
47
+
48
+ ## Rendering
49
+
50
+ In DSH:
51
+
52
+ ```json
53
+ { "action": "render", "model": { "...LogicModelV1..." }, "notation": "mermaid", "kind": "state" }
54
+ ```
55
+
56
+ `logicprobe_uml` with `action=render` returns the diagram source plus `warnings`. Read those warnings. They list every construct the notation could not carry verbatim: a state id that needed an alias, a label containing `[` or `/`, a trace that was capped.
57
+
58
+ Without the DSH tool, run the standalone engine:
59
+
60
+ ```bash
61
+ python skills/logicprobe/references/logicprobe-engine.py uml-render model.json --notation mermaid --diagram state
62
+ ```
63
+
64
+ You can also write the diagram by hand. The generator is a convenience, not a requirement. A hand-written diagram is parsed and reviewed exactly like a generated one.
65
+
66
+ ### Directives in generated diagrams
67
+
68
+ Generated text carries comment lines. Mermaid and PlantUML ignore them. The parser reads them:
69
+
70
+ ```text
71
+ %%logicprobe:uml v1 notation=mermaid diagram=state
72
+ %%logicprobe:init INIT
73
+ %%logicprobe:terminal FATAL
74
+ %%logicprobe:alias S_1 1st state
75
+ %%logicprobe:variable retry integer
76
+ ```
77
+
78
+ These lines exist because the notation cannot express everything the model knows. `[*]` marks an initial state, but a flowchart has no such marker. A state id may contain characters the notation cannot spell. A boolean assignment (`armed := 1`) looks exactly like an integer one. Pinning those facts in comments is what makes the round-trip check exact instead of approximate. PlantUML uses `'` instead of `%%`.
79
+
80
+ A hand-written diagram needs none of these directives. Adding `%%logicprobe:init` and `%%logicprobe:terminal` to a flowchart is how you say which node starts and which one ends.
81
+
82
+ ## Reviewing the modelling
83
+
84
+ ```json
85
+ { "action": "review", "model": { "...LogicModelV1..." } }
86
+ ```
87
+
88
+ Four input shapes, four different questions:
89
+
90
+ | Input | Question answered |
91
+ |---|---|
92
+ | `model` only | Is this machine well-modelled? The review checks its structure, then renders and re-parses it to prove the diagram carries it. |
93
+ | `diagram` only | What does this diagram actually say? The diagram is parsed into a model, and that model is reviewed. Fidelity to code is **unchecked** (`UML018`). |
94
+ | `model` + `diagram` | Does the diagram match the model? Every structural difference is a modelling defect (`UML017`). |
95
+ | `model` + `kind: "sequence"` | The trace is rendered and the structure is reviewed, but the fidelity check **does not apply** (`UML019`). A trace cannot be parsed back into a machine, so the review says so instead of pretending it verified the diagram. `maxSteps` caps the trace. |
96
+
97
+ ### Findings
98
+
99
+ | Code | Severity | What it means |
100
+ |---|---|---|
101
+ | `UML001_DIAGRAM_UNREADABLE` | error | The text is not readable as a Mermaid or PlantUML state or activity diagram. |
102
+ | `UML002_UNREACHABLE_STATE` | error | No transition can enter this state from init. The diagram draws flow nobody can reach. The check is structural and ignores guards; S1 is the guard-aware check. |
103
+ | `UML003_DEAD_END_STATE` | error | A non-terminal state has no outgoing transition. Either it is terminal, or the outgoing flow was never modelled. |
104
+ | `UML004_AMBIGUOUS_BRANCH` | error | Two unconditional arrows share one state and event. No reader and no implementation can resolve that. S4 is the authoritative check. |
105
+ | `UML005_OVERLAPPING_GUARD` | warning | The same guard text appears twice in one branch group. |
106
+ | `UML006_INEXHAUSTIVE_BRANCH` | warning or info | The branch group has only guarded branches and no default. The severity is `warning` when the guards do not look complementary. It is `info` when a complementary pair such as `k < 3` and `k >= 3` is present, which is probably exhaustive. Only S6 can settle it. |
107
+ | `UML007_UNUSED_EVENT` | warning | The event fires only from unreachable states. The diagram shows messages that never arrive. |
108
+ | `UML008_SELF_LOOP_NO_EXIT` | warning | An unguarded self-loop has no other exit. The diagram presents it as progress, but the flow never leaves. S3 reports absorbing cycles. |
109
+ | `UML009_DUPLICATE_TRANSITION` | warning | The same from, event, guard, actions and target row appears twice. |
110
+ | `UML010_UNUSED_VARIABLE` | warning | No guard reads this variable and no action writes it. It is a symbol with no source. |
111
+ | `UML011_UNBOUNDED_VARIABLE` | info | An integer variable has no min or max, so no range invariant can be checked and A5 has no declared domain. |
112
+ | `UML012_NO_TERMINAL` | warning | No state is terminal, so completion, failure and a stuck flow all look the same. |
113
+ | `UML013_NO_NARRATIVE` | info | The model carries no meanings, so a reader must re-derive every symbol from the source. |
114
+ | `UML014_UNDOCUMENTED_STATE` | info | These states render as their bare id, so the diagram cannot be read against the code. |
115
+ | `UML015_LABEL_DRIFT` | warning | A diagram label disagrees with the model narrative. One of the two is stale, and the review cannot tell which. |
116
+ | `UML016_DIAGRAM_PARSE_NOTES` | info | Notes collected while rendering or reading the diagram. They mark information the notation could not carry. |
117
+ | `UML017_ROUND_TRIP_MISMATCH` | error | The diagram does not carry the model. Transitions were lost or invented, or init, terminals or variables differ. The `roundTrip.diffs` array lists each one. |
118
+ | `UML018_FIDELITY_UNCHECKED` | info | Only a diagram was supplied, so nothing here proves it matches the code. |
119
+ | `UML019_ROUND_TRIP_SKIPPED` | warning | The fidelity check could not run, either because the view is not round-trippable or because it was switched off. |
120
+
121
+ `ok: true` means the review ran. It does not mean the model is good. Read the findings and the `summary` counts.
122
+
123
+ ### What the round trip proves
124
+
125
+ It proves the diagram is a faithful rendering of the model:
126
+
127
+ - the same initial state
128
+ - the same set of states and terminal states
129
+ - the same transitions, with the same guards and actions
130
+ - the same variables and kinds
131
+
132
+ It does **not** prove the model matches the code. Only a citation-per-element comparison does that, as described under "Evidence rule". It also does not prove the machine is correct. That is the job of `logicprobe_verify`.
133
+
134
+ A failing round trip means the notation lost something. In practice that is a real finding: an unlabelled arrow, a guard the parser could not read, or a diagram that was hand-edited away from its model.
135
+
136
+ ## Worked example
137
+
138
+ The code under review is a handshake that retries on timeout and gives up.
139
+
140
+ ```text
141
+ INIT --power_ready--> STARTING
142
+ STARTING --ack--> ACTIVE
143
+ STARTING --timeout [retry < 3] / retry := retry + 1--> ERROR
144
+ STARTING --timeout [retry >= 3]--> FATAL (terminal)
145
+ ERROR --cooldown--> STARTING
146
+ ```
147
+
148
+ Render it as a state machine, then review it. The review reports two modelling gaps:
149
+
150
+ ```text
151
+ UML006_INEXHAUSTIVE_BRANCH (info) guards look complementary on retry
152
+ UML013_NO_NARRATIVE (info) no state, event or scenario meanings
153
+ ```
154
+
155
+ Neither is a behaviour bug. Add the narrative so a reader can check the diagram against the code, then run `logicprobe_verify` for S1-S8 and A1-A14.
156
+
157
+ Now delete the `ERROR --cooldown--> STARTING` arrow and review again. `ACTIVE` and `ERROR` become dead ends (`UML003`), and `cooldown` becomes an event that only fires from an unreachable state (`UML007`). The diagram still looks plausible. That is the whole point.
158
+
159
+ ## Limits
160
+
161
+ - The review is **structural**. It evaluates no guard over any valuation. "Probably exhaustive" is a shape test, not a proof. S6 is the check that decides.
162
+ - Reachability ignores guards (`UML002`). A state reachable only under an unsatisfiable guard is a behaviour finding (S1 and S6), not a modelling one.
163
+ - Sequence diagrams are traces. They are neither round-tripped nor parsed. Reviewing one renders the trace, reports the structure findings, and marks the fidelity check inapplicable (`UML019`).
164
+ - PlantUML activity is refused by design, as described above.
165
+ - Renaming a state in the diagram does not rename it in the code. The review can only tell you the two disagree. Resolving it needs the source.
166
+ - The generated diagram is a faithful view of the **model**, never of the **program**. If the model is wrong, the diagram is wrong in exactly the same way.
@@ -1,4 +1,5 @@
1
- import { defineTool, type JsonValue } from '@deepseek-ai/dsh-tools'
1
+ import { defineTool } from '@deepseek-ai/dsh-tools'
2
+ import type { JsonValue } from './json-value.js'
2
3
  import { runCompositionVerification } from './engine.js'
3
4
 
4
5
  export const LOGICPROBE_COMPOSE_TOOL_NAME = 'logicprobe_compose_verify'
@@ -1,4 +1,5 @@
1
- import { defineTool, type JsonValue } from '@deepseek-ai/dsh-tools'
1
+ import { defineTool } from '@deepseek-ai/dsh-tools'
2
+ import type { JsonValue } from './json-value.js'
2
3
  import { runConcurrencyScan } from './concurrency.js'
3
4
 
4
5
  export const LOGICPROBE_CONCURRENCY_SCAN_TOOL_NAME = 'logicprobe_concurrency_scan'
package/src/data-tool.ts CHANGED
@@ -1,4 +1,5 @@
1
- import { defineTool, type JsonValue } from '@deepseek-ai/dsh-tools'
1
+ import { defineTool } from '@deepseek-ai/dsh-tools'
2
+ import type { JsonValue } from './json-value.js'
2
3
  import { runDataVerification, DATA_ENGINE_SCHEMA_VERSION } from './data-engine.js'
3
4
 
4
5
  export const LOGICPROBE_DATAMODEL_VERIFY_TOOL_NAME = 'logicprobe_datamodel_verify'
package/src/engine.ts CHANGED
@@ -1824,26 +1824,68 @@ function S8_monotonicVariables(model: LogicModelV1): CheckResult {
1824
1824
  return checkResult('S8', 'Monotonic Variables', findings, findings.length === 0 ? 'Monotonic variables are respected' : 'Monotonic findings: ' + findings.length)
1825
1825
  }
1826
1826
 
1827
+ /**
1828
+ * Find a run from `start` that never reaches `target`, or return undefined when
1829
+ * every run does. The property is universal (see the A9 row in `SKILL.md`): one
1830
+ * branch that loops forever, or that stops before the target, refutes it.
1831
+ *
1832
+ * The search is a depth-first walk of the run graph with three colours, and target
1833
+ * runs are success leaves that are never expanded:
1834
+ *
1835
+ * - a node reached while it is on the current walk (GRAY) closes a cycle that avoids
1836
+ * the target, so some run never reaches it;
1837
+ * - a node with no outgoing step at all stops the machine where it stands, which is
1838
+ * the same violation for a different reason;
1839
+ * - a node whose walk completed without a violation is BLACK, and reaching it again
1840
+ * from another branch is a shared sub-graph, not a cycle.
1841
+ *
1842
+ * That last point is why the colour map cannot be a plain visited set: a diamond
1843
+ * (two branches rejoining) is acyclic and must pass, while a genuine cycle must not.
1844
+ * A run is identified by its state plus its variable values, so the same state id
1845
+ * with different values is a different node.
1846
+ */
1827
1847
  function findLeadsToBadPath(model: LogicModelV1, start: RuntimeState, target: string): { path: PathStep[]; reason: string } | undefined {
1828
1848
  if (start.state === target) return undefined
1829
- const visited = new Set<string>()
1830
- const queue: Array<{ runtime: RuntimeState; path: PathStep[] }> = [{ runtime: start, path: [] }]
1831
- while (queue.length > 0) {
1832
- const entry = queue.shift()!
1833
- const key = runtimeKey(entry.runtime)
1834
- if (entry.runtime.state === target) continue
1835
- if (visited.has(key)) return { path: entry.path, reason: 'Cycle avoids target ' + target }
1836
- visited.add(key)
1837
- const nexts: Array<{ next: RuntimeState; event: string }> = []
1838
- for (const event of allEvents(model)) {
1839
- for (const next of stepRuntime(model, entry.runtime, event)) nexts.push({ next, event })
1849
+ const GRAY = 1
1850
+ const BLACK = 2
1851
+ interface LeadFrame {
1852
+ runtime: RuntimeState
1853
+ path: PathStep[]
1854
+ /** Outgoing steps, filled on first expansion; `undefined` means "not expanded yet". */
1855
+ nexts: Array<{ next: RuntimeState; event: string }> | undefined
1856
+ index: number
1857
+ }
1858
+ const color = new Map<string, number>()
1859
+ const stack: LeadFrame[] = [{ runtime: start, path: [], nexts: undefined, index: 0 }]
1860
+ color.set(runtimeKey(start), GRAY)
1861
+ while (stack.length > 0) {
1862
+ const frame = stack[stack.length - 1]
1863
+ if (frame.nexts === undefined) {
1864
+ const nexts: Array<{ next: RuntimeState; event: string }> = []
1865
+ for (const event of allEvents(model)) {
1866
+ for (const next of stepRuntime(model, frame.runtime, event)) nexts.push({ next, event })
1867
+ }
1868
+ frame.nexts = nexts
1869
+ if (nexts.length === 0) return { path: frame.path, reason: 'Dead end before target ' + target }
1840
1870
  }
1841
- if (nexts.length === 0) return { path: entry.path, reason: 'Dead end before target ' + target }
1842
- for (const { next, event } of nexts) {
1843
- queue.push({ runtime: next, path: [...entry.path, { from: entry.runtime.state, event, to: next.state }] })
1871
+ if (frame.index >= frame.nexts.length) {
1872
+ color.set(runtimeKey(frame.runtime), BLACK)
1873
+ stack.pop()
1874
+ continue
1844
1875
  }
1876
+ const { next, event } = frame.nexts[frame.index]
1877
+ frame.index += 1
1878
+ const step: PathStep = { from: frame.runtime.state, event, to: next.state }
1879
+ if (next.state === target) continue
1880
+ const key = runtimeKey(next)
1881
+ const seen = color.get(key)
1882
+ if (seen === GRAY) return { path: [...frame.path, step], reason: 'Cycle avoids target ' + target }
1883
+ if (seen === BLACK) continue
1884
+ color.set(key, GRAY)
1885
+ stack.push({ runtime: next, path: [...frame.path, step], nexts: undefined, index: 0 })
1845
1886
  }
1846
- return { path: [], reason: 'No path reaches target ' + target }
1887
+ // Every branch either reached the target or rejoined a branch that did.
1888
+ return undefined
1847
1889
  }
1848
1890
 
1849
1891
  function A9_leadsTo(model: LogicModelV1, exploration: Exploration): CheckResult {
@@ -1,4 +1,5 @@
1
- import { defineTool, type JsonValue } from '@deepseek-ai/dsh-tools'
1
+ import { defineTool } from '@deepseek-ai/dsh-tools'
2
+ import type { JsonValue } from './json-value.js'
2
3
  import { exportModel, type ExportFormat } from './exporters.js'
3
4
 
4
5
  export const LOGICPROBE_EXPORT_TOOL_NAME = 'logicprobe_export'
package/src/index.ts CHANGED
@@ -39,6 +39,7 @@ import { logicProbeDataModelVerifyTool, DATA_ENGINE_SCHEMA_VERSION } from './dat
39
39
  import { logicProbeConcurrencyScanTool } from './concurrency-tool.js'
40
40
  import { logicProbeComposeTool } from './compose-tool.js'
41
41
  import { logicProbeExportTool } from './export-tool.js'
42
+ import { logicProbeUmlTool } from './uml-tool.js'
42
43
  import { ENGINE_SCHEMA_VERSION } from './engine.js'
43
44
 
44
45
  // DSH 0.1.7-alpha.1 (session format v4) retires the shared
@@ -85,7 +86,7 @@ Plugin logicprobe is active. Documents are not truth — code is. Verify every v
85
86
  | "I'll verify while implementing" | Verification happens before implementation, not during. |
86
87
  | "I can check this with reasoning alone" | Behavioral claims are verified with code/models, not intuition. One counter-example refutes a universal claim. |
87
88
 
88
- **Native verification path**: In dsh, prefer the \`logicprobe_verify\` tool for state-machine checks and \`logicprobe_datamodel_verify\` for data-model/schema migration checks. Both support before/after regression and common domain constraints (idempotency, monotonic, sequence, leads-to, atomicity). Python harnesses remain the fallback for non-dsh hosts.
89
+ **Native verification path**: In dsh, prefer the \`logicprobe_verify\` tool for state-machine checks and \`logicprobe_datamodel_verify\` for data-model/schema migration checks. Both support before/after regression and common domain constraints (idempotency, monotonic, sequence, leads-to, atomicity). To model a code flow as UML, or to audit such a modelling, use \`logicprobe_uml\` (render | parse | review). Python harnesses remain the fallback for non-dsh hosts.
89
90
 
90
91
  **Proactive suggestion**: When a user asks code-level behavioral questions — "could this state machine deadlock", "is this retry limit safe", "check this timing sequence for bugs", "is this migration non-breaking", "does this copy cover all required fields" — suggest logicprobe as an optional verification pass (do not auto-escalate).
91
92
  </EXTREMELY_IMPORTANT>`
@@ -211,6 +212,7 @@ function modeContextText(config: Config, session: Session): string {
211
212
  const interaction = resolveInteraction(config, session)
212
213
  const lines = [
213
214
  'logicprobe: use `logicprobe_verify` for state machines and `logicprobe_datamodel_verify` for data models; both cover before/after regression and common domain constraints.',
215
+ 'logicprobe: use `logicprobe_uml` to model a code flow as UML (render), to read a UML diagram back into a model (parse), or to audit the modelling (review: structural defects, documentation gaps, diagram-vs-model round-trip fidelity).',
214
216
  interaction === 'auto'
215
217
  ? 'logicprobe interaction=auto: do NOT call ask_user_question for model confirmation; run round-trip validation of the extracted transition table and mark the result UNCONFIRMED.'
216
218
  : 'logicprobe interaction=ask: show the extracted transition table and get user confirmation before running verification.',
@@ -234,7 +236,7 @@ interface SystemPromptLike {
234
236
  * lets the model read this plugin's runtime status without guessing. Mirrors
235
237
  * the registration pattern of the official dsh-tool-cordis host providers.
236
238
  */
237
- function inspectProvider(config: Config, isToolRegistered: () => boolean, isDataToolRegistered: () => boolean, isConcurrencyToolRegistered: () => boolean, isComposeToolRegistered: () => boolean, isExportToolRegistered: () => boolean): HostCordisInspectProviderRegistration {
239
+ function inspectProvider(config: Config, isToolRegistered: () => boolean, isDataToolRegistered: () => boolean, isConcurrencyToolRegistered: () => boolean, isComposeToolRegistered: () => boolean, isExportToolRegistered: () => boolean, isUmlToolRegistered: () => boolean): HostCordisInspectProviderRegistration {
238
240
  return {
239
241
  manifest: {
240
242
  id: 'logicprobe',
@@ -260,10 +262,11 @@ function inspectProvider(config: Config, isToolRegistered: () => boolean, isData
260
262
  concurrencyToolRegistered: { type: 'boolean', description: 'Whether the logicprobe_concurrency_scan tool is registered on ctx.tools.' },
261
263
  composeToolRegistered: { type: 'boolean', description: 'Whether the logicprobe_compose_verify tool is registered on ctx.tools.' },
262
264
  exportToolRegistered: { type: 'boolean', description: 'Whether the logicprobe_export tool is registered on ctx.tools.' },
265
+ umlToolRegistered: { type: 'boolean', description: 'Whether the logicprobe_uml tool (UML modelling + modelling review) is registered on ctx.tools.' },
263
266
  engineSchemaVersion: { type: 'integer', description: 'Model schema version the bundled state-machine verification engine accepts.' },
264
267
  dataEngineSchemaVersion: { type: 'integer', description: 'Model schema version the bundled data-model verification engine accepts.' },
265
268
  },
266
- required: ['enabled', 'gateContentLength', 'interaction', 'toolRegistered', 'dataToolRegistered', 'concurrencyToolRegistered', 'composeToolRegistered', 'exportToolRegistered', 'engineSchemaVersion', 'dataEngineSchemaVersion'],
269
+ required: ['enabled', 'gateContentLength', 'interaction', 'toolRegistered', 'dataToolRegistered', 'concurrencyToolRegistered', 'composeToolRegistered', 'exportToolRegistered', 'umlToolRegistered', 'engineSchemaVersion', 'dataEngineSchemaVersion'],
267
270
  additionalProperties: false,
268
271
  },
269
272
  },
@@ -280,6 +283,7 @@ function inspectProvider(config: Config, isToolRegistered: () => boolean, isData
280
283
  concurrencyToolRegistered: isConcurrencyToolRegistered(),
281
284
  composeToolRegistered: isComposeToolRegistered(),
282
285
  exportToolRegistered: isExportToolRegistered(),
286
+ umlToolRegistered: isUmlToolRegistered(),
283
287
  engineSchemaVersion: ENGINE_SCHEMA_VERSION,
284
288
  dataEngineSchemaVersion: DATA_ENGINE_SCHEMA_VERSION,
285
289
  }
@@ -299,13 +303,14 @@ export function apply(ctx: Context, config: Config): void {
299
303
  let concurrencyToolRegistered = false
300
304
  let composeToolRegistered = false
301
305
  let exportToolRegistered = false
306
+ let umlToolRegistered = false
302
307
  let modeContextRegistered = false
303
308
  const registerProvider = (): void => {
304
309
  if (providerRegistered) return
305
310
  const inspect = ctx.get('cordisInspect')
306
311
  if (inspect === undefined) return
307
312
  try {
308
- ctx.effect(() => inspect.register(inspectProvider(config, () => toolRegistered, () => dataToolRegistered, () => concurrencyToolRegistered, () => composeToolRegistered, () => exportToolRegistered)), 'logicprobe: inspect provider')
313
+ ctx.effect(() => inspect.register(inspectProvider(config, () => toolRegistered, () => dataToolRegistered, () => concurrencyToolRegistered, () => composeToolRegistered, () => exportToolRegistered, () => umlToolRegistered)), 'logicprobe: inspect provider')
309
314
  providerRegistered = true
310
315
  } catch (err) {
311
316
  console.warn('[logicprobe] inspect provider registration failed', err)
@@ -321,11 +326,13 @@ export function apply(ctx: Context, config: Config): void {
321
326
  ctx.effect(() => tools.register(logicProbeConcurrencyScanTool), 'logicprobe: concurrency scan tool')
322
327
  ctx.effect(() => tools.register(logicProbeComposeTool), 'logicprobe: compose tool')
323
328
  ctx.effect(() => tools.register(logicProbeExportTool), 'logicprobe: export tool')
329
+ ctx.effect(() => tools.register(logicProbeUmlTool), 'logicprobe: uml tool')
324
330
  toolRegistered = true
325
331
  dataToolRegistered = true
326
332
  concurrencyToolRegistered = true
327
333
  composeToolRegistered = true
328
334
  exportToolRegistered = true
335
+ umlToolRegistered = true
329
336
  } catch (err) {
330
337
  console.warn('[logicprobe] logicprobe_verify/logicprobe_datamodel_verify tool registration failed', err)
331
338
  }
@@ -0,0 +1,20 @@
1
+ /**
2
+ * Structural mirror of the JSON value union the dsh tool registry accepts.
3
+ *
4
+ * `@deepseek-ai/dsh-tools` exported this type through the 0.1.x line and stopped
5
+ * exporting it in 0.2.x, where the same type now lives in
6
+ * `@deepseek-ai/dsh-util-values`. This package declares peer support for both
7
+ * lines, so it cannot import the type from either one without breaking the other
8
+ * build. Declaring the one-line union here keeps `npm run typecheck` green against
9
+ * both, and it costs nothing at runtime: every use is a type-only cast on a value
10
+ * this package already produced.
11
+ *
12
+ * Keep it structurally identical to the official type. A cast through `unknown` to
13
+ * this alias is only a way to satisfy `defineTool`'s return-type constraint; if the
14
+ * official union ever widens, this alias must widen with it.
15
+ *
16
+ * @module logicprobe-json-value
17
+ */
18
+
19
+ /** JSON-serializable value: null, boolean, number, string, or an array/object of those. */
20
+ export type JsonValue = null | boolean | number | string | JsonValue[] | { [key: string]: JsonValue }
package/src/tool.ts CHANGED
@@ -1,4 +1,5 @@
1
- import { defineTool, type JsonValue } from '@deepseek-ai/dsh-tools'
1
+ import { defineTool } from '@deepseek-ai/dsh-tools'
2
+ import type { JsonValue } from './json-value.js'
2
3
  import { runVerification } from './engine.js'
3
4
 
4
5
  export const LOGICPROBE_VERIFY_TOOL_NAME = 'logicprobe_verify'
@@ -0,0 +1,106 @@
1
+ import { defineTool } from '@deepseek-ai/dsh-tools'
2
+ import type { JsonValue } from './json-value.js'
3
+ import { renderUml, parseUml, reviewUml, UmlError, type UmlDiagram, type UmlNotation } from './uml.js'
4
+
5
+ export const LOGICPROBE_UML_TOOL_NAME = 'logicprobe_uml'
6
+
7
+ /**
8
+ * DSH tool wrapping the UML front end: model a code flow as a UML diagram
9
+ * (`action: "render"`), read a UML diagram back into a LogicModelV1
10
+ * (`action: "parse"`), or review the modelling itself (`action: "review"`).
11
+ *
12
+ * review is the half that makes the feature a check rather than a drawing
13
+ * utility: it reports structural defects the diagram would present as valid
14
+ * flow (unreachable states, dead ends, ambiguous and non-exhaustive branches,
15
+ * self-loops with no exit), documentation gaps (a symbol no reader can map back
16
+ * to code), and — the fidelity check — whether the rendered diagram reads back
17
+ * as the model it was drawn from.
18
+ */
19
+ export const logicProbeUmlTool = defineTool({
20
+ name: LOGICPROBE_UML_TOOL_NAME,
21
+ description:
22
+ 'Model a code flow as UML, and review the modelling. action="render" turns a LogicModelV1 (schemaVersion=1) into diagram source: notation mermaid (state | activity flowchart | sequence) or plantuml (state | sequence). action="parse" reads Mermaid/PlantUML state or activity text back into a LogicModelV1 (plus the display labels it found), so a hand-drawn diagram can be verified with logicprobe_verify; a sequence diagram is refused because a trace cannot reconstruct a machine. action="review" audits the modelling: structural defects (UML002 unreachable state, UML003 dead end, UML004 ambiguous branch, UML005 overlapping guard, UML006 inexhaustive branch, UML007 unused event, UML008 self-loop with no exit, UML009 duplicate transition), documentation gaps (UML010 unused variable, UML011 unbounded variable, UML012 no terminal, UML013 no narrative, UML014 undocumented state, UML015 label drift), and the fidelity check — the diagram is rendered and re-parsed and any structural difference is reported as UML017 round-trip mismatch. Give review a model (checks it, renders and re-parses it), a diagram (parses and reviews that), or both (checks the diagram against the model). Rendering never invents structure and never silently drops a construct the notation cannot express: those become warnings. Review does NOT replace logicprobe_verify — it covers the modelling, not behaviour; findings name the engine check to run next.',
23
+ parameters: {
24
+ action: {
25
+ type: 'string',
26
+ required: true,
27
+ enum: ['render', 'parse', 'review'],
28
+ description: 'render = model → UML source; parse = UML source → model; review = audit the modelling (and its fidelity to the model).',
29
+ },
30
+ model: {
31
+ type: 'json',
32
+ description: 'LogicModelV1 machine. Required for render; accepted by review (alone, or next to diagram to check the two against each other).',
33
+ },
34
+ diagram: {
35
+ type: 'string',
36
+ description: 'UML source text. Required for parse; accepted by review.',
37
+ },
38
+ notation: {
39
+ type: 'string',
40
+ enum: ['auto', 'mermaid', 'plantuml'],
41
+ description: 'Diagram language. Default auto: detected from the text for parse/review, mermaid for render.',
42
+ },
43
+ kind: {
44
+ type: 'string',
45
+ enum: ['state', 'activity', 'sequence'],
46
+ description: 'Diagram kind to render. Default state. sequence is one BFS trace, not the whole machine. PlantUML activity is refused (its structured-flowchart syntax cannot faithfully carry a graph with merges or cycles) — use mermaid for that view.',
47
+ },
48
+ roundTrip: {
49
+ type: 'boolean',
50
+ description: 'review only: render and re-parse the model to prove the diagram carries it. Default true.',
51
+ },
52
+ maxSteps: {
53
+ type: 'integer',
54
+ description: 'Cap on the sequence trace length when rendering kind=sequence. Default 60.',
55
+ },
56
+ },
57
+ output: {
58
+ schema: {
59
+ type: 'json',
60
+ description: 'Render result (diagram source), parse result (LogicModelV1 + labels), or the modelling review report with findings, round-trip diff and next steps.',
61
+ },
62
+ render(_args, value) {
63
+ const record = value as unknown as Record<string, unknown>
64
+ // A rendered diagram is source text: print it verbatim so it can be copied
65
+ // into a file, then the structured result (warnings, notation, kind).
66
+ if (typeof record.diagram === 'string') {
67
+ return [
68
+ { type: 'text' as const, text: record.diagram },
69
+ { type: 'text' as const, text: JSON.stringify(value, null, 2) },
70
+ ]
71
+ }
72
+ return [{ type: 'text' as const, text: JSON.stringify(value, null, 2) }]
73
+ },
74
+ },
75
+ timeoutMs: 10_000,
76
+ isConcurrencySafe: () => true,
77
+ async execute(args) {
78
+ try {
79
+ if (args.action === 'render') {
80
+ if (args.model === undefined) return errorResult('action=render needs `model` (a LogicModelV1 object)')
81
+ const result = renderUml(args.model, (args.notation === undefined || args.notation === 'auto' ? 'mermaid' : args.notation) as UmlNotation, (args.kind ?? 'state') as UmlDiagram, args.maxSteps)
82
+ return { ok: true, action: 'render', notation: result.notation, kind: result.diagram, diagram: result.primary, warnings: result.warnings } as unknown as JsonValue
83
+ }
84
+ if (args.action === 'parse') {
85
+ if (args.diagram === undefined) return errorResult('action=parse needs `diagram` (Mermaid or PlantUML text)')
86
+ const result = parseUml(args.diagram, args.notation ?? 'auto')
87
+ return { ok: true, action: 'parse', notation: result.notation, kind: result.diagram, model: result.model, labels: result.labels, warnings: result.warnings } as unknown as JsonValue
88
+ }
89
+ const report = reviewUml({
90
+ model: args.model,
91
+ diagram: args.diagram,
92
+ notation: args.notation ?? 'auto',
93
+ diagramKind: (args.kind ?? 'state') as UmlDiagram,
94
+ roundTrip: args.roundTrip,
95
+ maxSteps: args.maxSteps,
96
+ })
97
+ return report as unknown as JsonValue
98
+ } catch (error) {
99
+ return errorResult(error instanceof Error ? error.message : String(error), error instanceof UmlError) as unknown as JsonValue
100
+ }
101
+ },
102
+ })
103
+
104
+ function errorResult(message: string, isUmlError = false): JsonValue {
105
+ return { ok: false, ...(isUmlError ? { errorCode: 'UML_INPUT' } : {}), error: message } as unknown as JsonValue
106
+ }