@cntxt-labs/medha-cli 0.5.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,10 +16,57 @@ decides an action; you and your agent decide what to do with it.
16
16
  treated like 200 out of 200.
17
17
  - **Explainable.** Every number can be traced: `medha show` splits trust into its components and
18
18
  `medha explain-threshold` says which bars were cleared and which were not.
19
- - **Safe to try.** `medha simulate` shows what a signal would do without recording anything.
20
19
  - **One binary, no server.** CLI and MCP server in a single executable. State is an append-only
21
20
  episode log you can back up, compact, sync, or replay.
22
21
 
22
+ > [!NOTE]
23
+ > **Rust Core & TypeScript Hybrid Acceleration**: Medha provides both an active, full-featured TypeScript CLI/library ecosystem and a high-performance native Rust core. The native algorithms and trust computations are directly bridged to TypeScript and Node.js/Bun through our NAPI-RS adapter (`@cntxt-labs/medha-napi` / `crates/medha-napi`), providing native execution speeds while keeping the TypeScript CLI and npm packages fully supported. Standalone Rust binaries are also available (`crates/medha-cli`).
24
+
25
+ ## Mental Model: The Dual-Loop Architecture
26
+
27
+ Most memory tools store unstructured chat history or flat key-value assertions. Medha acts as an **evidential calibration loop**:
28
+
29
+ ```text
30
+ THE DUAL-LOOP MENTAL MODEL
31
+
32
+ ┌──────────────────────────────────────────────────────────────┐
33
+ │ FAST INNER LOOP: EXECUTION │
34
+ │ │
35
+ │ Agent Task ──► Query Trust Hints ──► Context Injection │
36
+ │ │ │
37
+ │ ▼ │
38
+ │ Should I apply this rule? │
39
+ │ (Agent / Human Decision) │
40
+ └────────────────────────┬─────────────────────────────────────┘
41
+ │ Outcomes observed
42
+ ▼
43
+ ┌──────────────────────────────────────────────────────────────┐
44
+ │ SLOW OUTER LOOP: EVIDENCE │
45
+ │ │
46
+ │ Record Signals & Guard Checks (APPLY, REJECT, PASS/FAIL) │
47
+ │ │ │
48
+ │ ▼ │
49
+ │ Evidential Trust Engine │
50
+ │ T = min(Ceiling, L × G × R × D) │
51
+ │ Wilson Lower Bound (L) × Guard Factor (G) │
52
+ │ × Recency Decay (R) × Durability (D) │
53
+ │ │ │
54
+ │ ▼ │
55
+ │ Calibrated Status: Probation ──► Active ──► Trusted │
56
+ │ └──► Quarantined / Retired │
57
+ └──────────────────────────────────────────────────────────────┘
58
+ ```
59
+
60
+ - **The Fast Inner Loop**: When starting a task, agents query `hints` or `medha show`. Entities with high trust are injected into active context; probation or quarantined entities are discounted or ignored. **Medha reports evidence; you decide.**
61
+ - **The Slow Outer Loop**: As actions execute, the agent or test runner reports ground truth: did the rule work (`APPLY`), did a human reject it (`REJECT_RULE`), did an automated test pass (`guard --ok`)?
62
+ - **Trust Formula ($T$)**:
63
+ $$T = \min\big(\text{ceiling}, L \times G \times R \times D\big)$$
64
+ - **$L$ (Wilson Lower Bound)**: 95% confidence interval on success rate $k/n$. Protects against small-sample overconfidence ($2/2 \ne 200/200$).
65
+ - **$G$ (Guard Factor)**: 1.0 if verified by test/AST guard; penalized if failing or unverified.
66
+ - **$R$ (Recency Decay)**: Exponential decay based on time elapsed since last use (default 30-day half-life, floor 0.20).
67
+ - **$D$ (Durability Factor)**: Logarithmic bonus for rules validated across multiple git commits, branches, or weeks.
68
+ - **Ceiling**: Unguarded entities cannot exceed 0.85, preventing unverified heuristics from becoming `trusted`.
69
+
23
70
  ## Install
24
71
 
25
72
  ```sh
@@ -115,6 +162,32 @@ Report the result with `medha guard --ok` or `--fail`.
115
162
  Entity state is a fold over that log, so it can always be rebuilt, audited, or corrected
116
163
  (`medha retract`, `medha remove-episode`).
117
164
 
165
+ **Decision tree.** An entity can carry a tree of *branches*, each a condition plus a decision
166
+ (`apply`, `ignore`, or a probability). Branches earn their own trust from their own evidence, which
167
+ is what you want when a rule only holds in one situation: "always skip migration backups" can be
168
+ excellent advice in CI and wrong on a laptop.
169
+
170
+ ```sh
171
+ medha decision --id migration-notes --condition "in CI" --apply
172
+ medha decision --id migration-notes --condition "on a laptop" --ignore --parent <case-id>
173
+ ```
174
+
175
+ Evidence is then attributed to a branch as well as to the entity, so a branch can read `trusted`
176
+ while the entity as a whole reads `probation`:
177
+
178
+ ```sh
179
+ medha record --id migration-notes --signal APPLY --case-id <case-id>
180
+ ```
181
+
182
+ `caseId` must name a branch of that entity — a typo is rejected rather than silently filed against
183
+ the entity. Attributing evidence to a branch never *replaces* the entity-level fold; it is a second
184
+ view of the same signal.
185
+
186
+ Editing a branch (`medha decision --id .. --condition .. --case-id <case-id>`) changes its decision
187
+ and leaves it where it is. Pass `--parent <case-id>` to move it, or `--detach` to promote it to the
188
+ top level. Writes that name an unknown parent, a descendant of themselves, or a condition a sibling
189
+ already uses are rejected.
190
+
118
191
  ### How trust is computed
119
192
 
120
193
  ```
@@ -163,7 +236,8 @@ Then copy [`SKILL.md`](SKILL.md) to
163
236
  | `hints` | Batch-fetch trust hints for the entities you are about to rely on. Returns `{ hints, unknown }`; treat `unknown` as probation. |
164
237
  | `list_entities` | Paginated search by kind, status, namespace or drift. |
165
238
  | `show_entity` | Full detail for one entity: components, temporal state, recent episodes. |
166
- | `record_signal` | Record `APPLY`, `REJECT_RULE`, `SKIP`, and so on. Returns `recorded: false` for an unknown entity unless `ensure` is set. |
239
+ | `record_signal` | Record `APPLY`, `REJECT_RULE`, `SKIP`, and so on. Returns `recorded: false` for an unknown entity unless `ensure` is set. Pass `caseId` to also teach a decision-tree branch. |
240
+ | `record_decision` | Create or edit a branch of an entity's decision tree. Pass `parentId` to nest it, or `detach: true` to promote it to the top level. Returns the minted `caseId`. |
167
241
  | `report_guard` | Record a guard result. |
168
242
  | `propose` | Submit a candidate entity. |
169
243
  | `drift` | List drifting entities. |
@@ -179,6 +253,7 @@ host can inject the most trusted guidance without overrunning its context.
179
253
 
180
254
  - **`medha ui`** launches a local web dashboard over the store.
181
255
  - **`medha report`** writes a standalone, offline HTML snapshot you can attach to a review.
256
+ - **`medha issue [title]`** prepares a GitHub issue prefilled with sanitized runtime and store diagnostics.
182
257
  - **`medha sync status|pull|push`** shares evidence between machines through a git ref or a file.
183
258
  Registries travel with the episodes, so custom kinds and signals do not have to be copied by hand.
184
259
 
@@ -186,9 +261,57 @@ host can inject the most trusted guidance without overrunning its context.
186
261
 
187
262
  - **Kinds and signals.** Register your own in `.medha/config.json`. A kind can set its own trust
188
263
  thresholds and recency half-life, and can weight evidence by a signal's value (for example, a
189
- timeout counts less against a tool than a crash).
264
+ timeout counts less against a tool than a crash). See [Kind specs](#kind-specs) below.
190
265
  - **Weight updaters.** Swap how evidence moves trust with `medha updater list` and `medha updater fork <name>`, which scaffolds a custom one.
191
266
 
267
+ ### Kind specs
268
+
269
+ A **kind spec** gives one kind its own policy. Pass them to `medha init --config <file>`:
270
+
271
+ ```json
272
+ {
273
+ "kindSpecs": [
274
+ {
275
+ "name": "lint",
276
+ "description": "linter rules: frequent, cheap evidence",
277
+ "thresholds": { "active": 0.15, "trusted": 0.4, "minUsesForTrusted": 3 },
278
+ "recency": { "halfLifeDays": 20, "floor": 0.2 },
279
+ "evidenceWeighting": "signal-value"
280
+ },
281
+ {
282
+ "name": "rule",
283
+ "signalLimits": { "maxSuccessesPerAuthor": 3 },
284
+ "decisionPolicy": { "requireHumanFor": ["apply"] }
285
+ }
286
+ ]
287
+ }
288
+ ```
289
+
290
+ ```sh
291
+ medha init --config medha.config.json
292
+ ```
293
+
294
+ `--config` adds to the built-in kinds (`rule`, `recipe`, `tool`) rather than replacing them. A spec
295
+ naming a new kind (`lint` above) also registers it, and a spec naming a built-in kind (`rule`) sets
296
+ that kind's policy. Every field except `name` is optional, and a field you leave out keeps its
297
+ default.
298
+
299
+ | Field | Keys (default) | What it does |
300
+ |---|---|---|
301
+ | `thresholds` | `active` (0.25), `trusted` (0.6), `minUsesForTrusted` (5), `retiredTrustThreshold` (0.1), `minUsesForRetired` (3), `unguardedCeiling` (0.5) | Where the lifecycle transitions sit for this kind. `unguardedCeiling` caps trust while no guard has passed; no threshold lets an unguarded entity become `trusted`. |
302
+ | `recency` | `halfLifeDays` (45), `floor` (0.3) | How fast old evidence fades, and the least it fades to. |
303
+ | `evidenceWeighting` | `"count"` (default) or `"signal-value"` | Count each signal as one trial, or weight it by the signal's value. |
304
+ | `signalLimits` | `minIntervalMs`, `maxSuccessesPerAuthor` (both off) | Limit how much one author can raise trust. Negative evidence is never limited. |
305
+ | `decisionPolicy` | `requireHumanFor`: `"apply"` or a list of `apply`/`ignore`/`probability` (off) | Decision-tree branches of these types need a `human:` author. This is a label the caller sets, not verification: it stops an agent that follows instructions, not one that lies. |
306
+
307
+ Config files are strict: an unknown key, such as `kindPolicies` or `thresholds.bogus`, is refused with
308
+ the list of allowed keys, so a typo cannot quietly leave a policy switched off.
309
+
310
+ To change specs after `init`, edit `.medha/config.json`. There they live under
311
+ `registries.kindSpecs`, and a new kind's name must also be added to `registries.kinds`. Run
312
+ `medha maintain preflight` to check the result. [Extending medha](https://nimishph.github.io/medha/guide/extending)
313
+ covers kinds, signals and updaters in depth.
314
+
192
315
  ## Maintain it
193
316
 
194
317
  ```sh
package/SKILL.md ADDED
@@ -0,0 +1,135 @@
1
+ ---
2
+ name: medha
3
+ description: Evidential memory for rules, recipes and tools. Use when deciding how far to trust a review rule, recipe or tool based on how it has actually performed — record uses, rejections and guard results, and read back trust hints. Medha reports evidence; it never decides.
4
+ ---
5
+
6
+ # medha
7
+
8
+ Medha remembers how well rules, recipes and tools have actually worked and returns **trust hints**.
9
+ You record evidence; **you** decide what to do with it.
10
+
11
+ ## Mental Model: Dual-Loop Architecture
12
+
13
+ ```text
14
+ THE DUAL-LOOP MENTAL MODEL
15
+
16
+ ┌──────────────────────────────────────────────────────────────┐
17
+ │ FAST INNER LOOP: EXECUTION │
18
+ │ │
19
+ │ Agent Task ──► Query Trust Hints ──► Context Injection │
20
+ │ │ │
21
+ │ ▼ │
22
+ │ Should I apply this rule? │
23
+ │ (Agent / Human Decision) │
24
+ └────────────────────────┬─────────────────────────────────────┘
25
+ │ Outcomes observed
26
+ ▼
27
+ ┌──────────────────────────────────────────────────────────────┐
28
+ │ SLOW OUTER LOOP: EVIDENCE │
29
+ │ │
30
+ │ Record Signals & Guard Checks (APPLY, REJECT, PASS/FAIL) │
31
+ │ │ │
32
+ │ ▼ │
33
+ │ Evidential Trust Engine │
34
+ │ T = min(Ceiling, L × G × R × D) │
35
+ │ Wilson Lower Bound (L) × Guard Factor (G) │
36
+ │ × Recency Decay (R) × Durability (D) │
37
+ │ │ │
38
+ │ ▼ │
39
+ │ Calibrated Status: Probation ──► Active ──► Trusted │
40
+ │ └──► Quarantined / Retired │
41
+ └──────────────────────────────────────────────────────────────┘
42
+ ```
43
+
44
+ - **Fast Inner Loop**: Read `hints` (or `medha show`) before applying rules. Treat unknown entities as `probation`.
45
+ - **Slow Outer Loop**: Report ground truth as events occur (`record_signal`, `report_guard`).
46
+ - **Trust Formula**: $T = \min(\text{ceiling}, L \times G \times R \times D)$. Usage alone never exceeds 0.85; a passing guard is required to achieve `trusted`.
47
+
48
+ ## Setup
49
+
50
+ Run once per project (creates `.medha/`; add it to `.gitignore` or commit it deliberately):
51
+
52
+ ```sh
53
+ medha init
54
+ ```
55
+
56
+ MCP server (register in `.mcp.json`): `{ "mcpServers": { "medha": { "command": "medha", "args": ["mcp", "serve"] } } }`
57
+
58
+ ## Entities
59
+
60
+ An entity is addressed by `namespace` (default empty), `kind` (`rule` | `recipe` | `tool`, default
61
+ `rule`) and `id`. New entities start on **probation**.
62
+
63
+ ## Per-kind policy (kindSpecs)
64
+
65
+ A kind can carry its own policy. Set it at init with `medha init --config <file>`:
66
+
67
+ ```json
68
+ { "kindSpecs": [
69
+ { "name": "lint", "thresholds": { "active": 0.15, "trusted": 0.4, "minUsesForTrusted": 3 },
70
+ "recency": { "halfLifeDays": 20 }, "evidenceWeighting": "signal-value" },
71
+ { "name": "rule", "signalLimits": { "maxSuccessesPerAuthor": 3 },
72
+ "decisionPolicy": { "requireHumanFor": ["apply"] } }
73
+ ] }
74
+ ```
75
+
76
+ - Allowed keys: `name`, `description`, `thresholds` (`active`, `trusted`, `minUsesForTrusted`,
77
+ `retiredTrustThreshold`, `minUsesForRetired`, `unguardedCeiling`), `recency` (`halfLifeDays`,
78
+ `floor`), `evidenceWeighting` (`count` | `signal-value`), `signalLimits` (`minIntervalMs`,
79
+ `maxSuccessesPerAuthor`), `decisionPolicy` (`requireHumanFor`). An unknown key is an error.
80
+ - A spec for a new kind registers it too; built-in kinds stay. Omitted fields keep their defaults.
81
+ - After init, specs live in `.medha/config.json` under `registries.kindSpecs`, and a new kind's
82
+ name must also be listed in `registries.kinds`. Check with `medha maintain preflight`.
83
+ - `decisionPolicy.requireHumanFor` means **a human must run that command**. If you get
84
+ `CORE_PERMISSION_DENIED`, do not retry with a `human:` author. You can set that label yourself,
85
+ so passing it proves nothing. Escalate to a human instead.
86
+
87
+ ## Workflow
88
+
89
+ 1. **Read** before relying on a rule: MCP `hints` (pass `compact: true` for cheap output) or
90
+ `medha show --id <id>`. `hints` returns `{ hints, unknown }` — treat `unknown` keys as probation.
91
+ 2. **Propose** a new rule: `propose` / `medha propose --id <id> --source <who> [--text ..]`.
92
+ 3. **Record** evidence as it happens: `record_signal` / `medha record --id <id> --signal APPLY --ensure`.
93
+ Signals: `APPLY` (used and worked), `REJECT_RULE` (a human rejected it), `SKIP` (not applicable).
94
+ If the entity is unknown and `ensure` is not set, the response says `recorded: false` — nothing was written.
95
+ 4. **Report guards**: `report_guard` / `medha guard --id <id> --ok|--fail --guard review`.
96
+ Without a passing guard, trust is capped at 0.5 — usage alone never makes a rule `trusted`.
97
+ 5. **Explain**: `medha explain-threshold --id <id>` lists each threshold with ok/no and the numbers.
98
+ 6. **Preview** with `simulate` (persists nothing) before recording a consequential signal.
99
+
100
+ ## Branching a rule (only when a rule is only good sometimes)
101
+
102
+ A rule that holds in one situation and not another belongs in a decision tree, not in a single
103
+ trust number. `record_decision` mints a branch and returns its `caseId`:
104
+
105
+ ```
106
+ record_decision { id, condition: "in CI", decision: { type: "apply" } }
107
+ record_decision { id, condition: "on a laptop", decision: { type: "ignore" }, parentId: <first> }
108
+ ```
109
+
110
+ Then attribute evidence to the branch, not just the entity:
111
+
112
+ ```
113
+ record_signal { id, signal: "APPLY", caseId: <branch> }
114
+ ```
115
+
116
+ Read a branch's own trust with `show_entity` — `decisionTree` lists each branch with its `parentId`,
117
+ condition, evidence and status. Branch evidence is *in addition to* the entity's aggregate, never a
118
+ replacement for it, so an entity can be `probation` while one branch is `trusted`.
119
+
120
+ - `caseId` must be a real branch of that entity; a typo is rejected, not filed against the entity.
121
+ - Editing a branch by `caseId` without `parentId` leaves it exactly where it is. Pass `parentId` to
122
+ move it, or `detach: true` to make it a root.
123
+ - Two branches with the same condition under the same parent are rejected — the same condition once
124
+ under each of two roots is fine.
125
+
126
+ ## Lifecycle
127
+
128
+ `probation → active → trusted`, and down to `quarantined` / `retired` on repeated rejection.
129
+ `drift` lists entities whose recent behavior diverges from their baseline.
130
+
131
+ ## Rules of thumb
132
+
133
+ - Trust is a hint, not a verdict: Wilson lower bound over successes/trials, decayed, gated by guards.
134
+ - Use `--json` on any CLI command for machine output; `--home <dir>` to point at a shared home.
135
+ - `medha maintain preflight` checks store integrity; `medha maintain backup <path>` exports a snapshot.
package/bin/medha.cjs CHANGED
@@ -6,6 +6,7 @@
6
6
  // that matches the machine. This file finds it and runs it.
7
7
 
8
8
  const { spawnSync } = require('node:child_process');
9
+ const { chmodSync } = require('node:fs');
9
10
  const path = require('node:path');
10
11
 
11
12
  const SUPPORTED = ['linux-x64', 'linux-arm64', 'darwin-arm64', 'darwin-x64', 'win32-x64'];
@@ -37,6 +38,29 @@ function locate(platform, cpu, resolve) {
37
38
  }
38
39
  }
39
40
 
41
+ /**
42
+ * Restore the executable bit on the program, and say whether it worked.
43
+ *
44
+ * The published tarball is supposed to carry mode 0755, but a package can still arrive without it:
45
+ * an install that strips permissions, a copy through a filesystem or archive tool that does not keep
46
+ * them, or a version published before the mode was fixed. The program is always a program, so the
47
+ * bit is safe to set here — this is the difference between a one-line `chmod` the user has to guess
48
+ * and a command that just runs.
49
+ */
50
+ function makeExecutable(program) {
51
+ try {
52
+ chmodSync(program, 0o755);
53
+ return true;
54
+ } catch {
55
+ return false;
56
+ }
57
+ }
58
+
59
+ /** True when a spawn failed only because the file was not executable. */
60
+ function isPermissionFailure(error) {
61
+ return error !== undefined && error !== null && error.code === 'EACCES';
62
+ }
63
+
40
64
  function main() {
41
65
  const found = locate(process.platform, process.arch, require.resolve);
42
66
  if (found.problem !== undefined) {
@@ -44,14 +68,27 @@ function main() {
44
68
  process.exitCode = 1;
45
69
  return;
46
70
  }
47
- const child = spawnSync(found.program, process.argv.slice(2), {
48
- stdio: 'inherit',
49
- cwd: process.cwd(),
50
- });
71
+ const run = () =>
72
+ spawnSync(found.program, process.argv.slice(2), { stdio: 'inherit', cwd: process.cwd() });
73
+
74
+ let child = run();
75
+ if (
76
+ isPermissionFailure(child.error) &&
77
+ process.platform !== 'win32' &&
78
+ makeExecutable(found.program)
79
+ ) {
80
+ child = run();
81
+ }
51
82
  if (child.error) {
52
83
  process.stderr.write(
53
84
  `medha: could not start ${path.basename(found.program)}: ${child.error.message}\n`,
54
85
  );
86
+ if (isPermissionFailure(child.error)) {
87
+ process.stderr.write(
88
+ `hint: ${found.program} is not executable and medha could not make it so. ` +
89
+ `Run: chmod +x "${found.program}"\n`,
90
+ );
91
+ }
55
92
  process.exitCode = 1;
56
93
  return;
57
94
  }
@@ -62,6 +99,6 @@ function main() {
62
99
  process.exitCode = child.status ?? 1;
63
100
  }
64
101
 
65
- module.exports = { SUPPORTED, platformPackage, locate };
102
+ module.exports = { SUPPORTED, platformPackage, locate, makeExecutable, isPermissionFailure };
66
103
 
67
104
  if (require.main === module) main();
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@cntxt-labs/medha-cli",
3
- "version": "0.5.0",
3
+ "version": "0.7.0",
4
4
  "description": "Evidential memory for rules, recipes and tools: records what happened and returns trust hints. CLI and MCP server.",
5
5
  "keywords": [
6
6
  "mcp",
@@ -40,11 +40,11 @@
40
40
  "node": ">=18"
41
41
  },
42
42
  "optionalDependencies": {
43
- "@cntxt-labs/medha-linux-x64": "0.5.0",
44
- "@cntxt-labs/medha-linux-arm64": "0.5.0",
45
- "@cntxt-labs/medha-darwin-arm64": "0.5.0",
46
- "@cntxt-labs/medha-darwin-x64": "0.5.0",
47
- "@cntxt-labs/medha-win32-x64": "0.5.0"
43
+ "@cntxt-labs/medha-linux-x64": "0.7.0",
44
+ "@cntxt-labs/medha-linux-arm64": "0.7.0",
45
+ "@cntxt-labs/medha-darwin-arm64": "0.7.0",
46
+ "@cntxt-labs/medha-darwin-x64": "0.7.0",
47
+ "@cntxt-labs/medha-win32-x64": "0.7.0"
48
48
  },
49
49
  "devDependencies": {
50
50
  "@modelcontextprotocol/sdk": "^1.30.0",