@clien-ai/mcp 0.10.9 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +152 -0
- package/README.md +33 -10
- package/dist/hypothesis-semantics.js +230 -0
- package/dist/hypothesis-semantics.js.map +1 -0
- package/dist/tools/hypothesis-final.js +55 -0
- package/dist/tools/hypothesis-final.js.map +1 -0
- package/dist/tools/interview-scripts.js +50 -21
- package/dist/tools/interview-scripts.js.map +1 -1
- package/dist/tools/market-sizing-proof.js +69 -4
- package/dist/tools/market-sizing-proof.js.map +1 -1
- package/dist/tools/personas.js +9 -2
- package/dist/tools/personas.js.map +1 -1
- package/dist/tools/receipt-children.js +127 -0
- package/dist/tools/receipt-children.js.map +1 -0
- package/dist/tools/registry.js +49 -11
- package/dist/tools/registry.js.map +1 -1
- package/dist/tools/render-safety.js +80 -0
- package/dist/tools/render-safety.js.map +1 -1
- package/dist/tools/report-digest.js +560 -46
- package/dist/tools/report-digest.js.map +1 -1
- package/dist/tools/research.js +28 -3
- package/dist/tools/research.js.map +1 -1
- package/dist/tools/scoped-research.js +40 -11
- package/dist/tools/scoped-research.js.map +1 -1
- package/dist/tools/status.js +46 -2
- package/dist/tools/status.js.map +1 -1
- package/dist/types/report.js +452 -24
- package/dist/types/report.js.map +1 -1
- package/package.json +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -13,6 +13,158 @@ upgrade here does not imply the deployed app speaks the same shape, and the
|
|
|
13
13
|
client strips unknown fields by construction — so a field described below only
|
|
14
14
|
reaches you once both sides have shipped.
|
|
15
15
|
|
|
16
|
+
## 0.12.0
|
|
17
|
+
|
|
18
|
+
**Hypothesis confidence now means how reliable the final verdict is, not the
|
|
19
|
+
probability that the hypothesis is true.** The report and interview-script
|
|
20
|
+
tools expose the same effective verdict after robustness checks, with separate
|
|
21
|
+
evidence sufficiency and robustness outcomes. This is a minor release because
|
|
22
|
+
the `structuredContent.interviewScripts[].confidence` field changes from a
|
|
23
|
+
number to `low` / `medium` / `high`; integrations reading structured output
|
|
24
|
+
must handle the new categorical shape.
|
|
25
|
+
|
|
26
|
+
### Changed
|
|
27
|
+
|
|
28
|
+
- **Interview scripts rank what most needs a real answer.** Product priority,
|
|
29
|
+
lower verdict reliability, and thinner evidence replace the old distance from
|
|
30
|
+
0.5 ordering, which incorrectly treated a model-written score as a posterior
|
|
31
|
+
probability.
|
|
32
|
+
- **Every report-facing MCP reader resolves one final hypothesis result.** A
|
|
33
|
+
robustness downgrade or flip cannot retain contradictory high confidence,
|
|
34
|
+
and malformed robustness output cannot upgrade a result.
|
|
35
|
+
- **Legacy numeric confidence remains explicit and non-comparable.** In
|
|
36
|
+
`list_interview_scripts`, an older app response's 0–1 or 0–100 value moves to
|
|
37
|
+
`legacyConfidence` rather than mixing with the categorical measure. Report
|
|
38
|
+
data keeps that compatibility value at `hypothesisResults[].confidence` and
|
|
39
|
+
treats it as legacy beside `hypothesisResults[].final.confidence`.
|
|
40
|
+
- **Research tools drain every terminal event page before returning.** Full and
|
|
41
|
+
scoped runs no longer stop at the first bounded snapshot when later terminal
|
|
42
|
+
events carry the authoritative outcome.
|
|
43
|
+
|
|
44
|
+
### Added
|
|
45
|
+
|
|
46
|
+
- **Final hypothesis semantics in structured report output:** categorical
|
|
47
|
+
verdict reliability, evidence sufficiency, robustness outcome, and preserved
|
|
48
|
+
original/final state for audit.
|
|
49
|
+
- **Literal trust states and denominators.** Successful checks that found
|
|
50
|
+
nothing now say so, while report-prose and digest claim populations are
|
|
51
|
+
labelled separately instead of appearing to contradict one another.
|
|
52
|
+
- **Bounded status-snapshot disclosure.** `clien_research_status` adds
|
|
53
|
+
`_meta.snapshot_truncated` when newer events exist beyond the returned page.
|
|
54
|
+
|
|
55
|
+
### Fixed
|
|
56
|
+
|
|
57
|
+
- **Markdown and digest readers no longer disagree with stored or visible
|
|
58
|
+
hypothesis outcomes.** Shared resolution and position-specific escaping keep
|
|
59
|
+
the final verdict coherent without corrupting receipt text.
|
|
60
|
+
- **Failed and cancelled runs surface the latest observed refund status with
|
|
61
|
+
bounded guidance.** Full and scoped research now use the drained terminal
|
|
62
|
+
events for error, refunded-credit, and refund-pending guidance.
|
|
63
|
+
`refund_pending` means the run is refund-eligible but no refund is visible
|
|
64
|
+
yet; it is not a guarantee that a refund will occur.
|
|
65
|
+
|
|
66
|
+
## 0.11.0
|
|
67
|
+
|
|
68
|
+
**A claim grounded on the second excerpt of a discussion now shows the second
|
|
69
|
+
excerpt.** One community permalink can legitimately carry several bounded
|
|
70
|
+
excerpts, and a `GROUNDED` claim records which one its span lives in. Every
|
|
71
|
+
receipt reader in this package still resolved positionally — claim → `sourceId`
|
|
72
|
+
→ `personas[i].sources[j].quote` — and that parent row holds the FIRST excerpt.
|
|
73
|
+
So a claim grounded on the second displayed a clip of the first, under a
|
|
74
|
+
`[GROUNDED]` badge, with no way for a reader to tell. A minor rather than a
|
|
75
|
+
patch because the digest gained lines and a section, and because an older
|
|
76
|
+
installed client cannot honour `receiptId` at all.
|
|
77
|
+
|
|
78
|
+
### Fixed
|
|
79
|
+
|
|
80
|
+
- **The trust digest resolves `claim.receiptId` to the excerpt it names.** The
|
|
81
|
+
source's pool row is unchanged and still one row per discussion — receipt
|
|
82
|
+
count must never inflate how much community evidence a run appears to have —
|
|
83
|
+
with an indented `↳ rr1:… — "…"` line for each excerpt a rendered claim
|
|
84
|
+
actually addresses. The claim's own line carries both ids (`← RCP-p0-s0 ·
|
|
85
|
+
rr1:…`), so two claims on one discussion can be told apart.
|
|
86
|
+
- **The window is chosen per excerpt, not per source.** A span verified inside
|
|
87
|
+
the second excerpt is not findable in the first, so the previous single span
|
|
88
|
+
bucket would have windowed the parent line by a span it does not contain —
|
|
89
|
+
locating nothing, taking the head clip, and looking like a window.
|
|
90
|
+
- **`methodology.redditRetrieval`'s `outcome` and `reason` are bounded code
|
|
91
|
+
tokens.** They are approved keys, so whatever they carried was copied verbatim
|
|
92
|
+
into `structuredContent.report_data` and `_meta.report_data` — typed channels
|
|
93
|
+
the digest's own truncation never touches. Raw text in either field is now
|
|
94
|
+
refused rather than forwarded, at the schema and again in the projection that
|
|
95
|
+
runs when a report fails to parse. An unknown token still parses, because a
|
|
96
|
+
client older than the app that answered it must report an outcome it does not
|
|
97
|
+
recognise rather than fall silent.
|
|
98
|
+
- **An explicit `receiptId: null` fails closed instead of borrowing the parent's
|
|
99
|
+
excerpt.** It used to read as "no receipt named", which licenses showing the
|
|
100
|
+
source's first excerpt as the claim's evidence. Our own producer omits this
|
|
101
|
+
optional key rather than writing `null`, so a stored `null` has unknown
|
|
102
|
+
provenance and cannot claim that licence — it now loses the receipt, exactly
|
|
103
|
+
as an unparseable id does. The app spine, the public share payload and this
|
|
104
|
+
digest give the same answer, pinned by a coupling test that fails if either
|
|
105
|
+
side changes alone.
|
|
106
|
+
|
|
107
|
+
- **A whitespace-only receipt excerpt fails closed.** `" "` cleared every
|
|
108
|
+
producer and database check, so the digest rendered `↳ rr1:… — ""` under a
|
|
109
|
+
`[GROUNDED]` badge — a claim whose evidence is the source saying nothing. Such
|
|
110
|
+
a claim now renders "receipt unavailable; do not cite".
|
|
111
|
+
- **`completed` with zero accepted sources is treated as inconsistent input.**
|
|
112
|
+
The rail's contract makes `completed` mean at least one source was accepted, so
|
|
113
|
+
a note reporting both states coverage the run does not have and points at a
|
|
114
|
+
pool the next section may declare empty. It renders as coverage UNKNOWN. An
|
|
115
|
+
unreadable count is not zero and keeps the completed wording.
|
|
116
|
+
- **Malformed acceptance counters no longer print impossible evidence counts.**
|
|
117
|
+
Values that are not whole non-negative numbers now render as "not recorded
|
|
118
|
+
readably" instead of being treated as a real number of accepted discussions or
|
|
119
|
+
bounded excerpts.
|
|
120
|
+
- **The scoped primitives' `report_data` meets the same projection the full
|
|
121
|
+
report does.** `scan_competitors`, `search_forums` and `scan_market` forwarded
|
|
122
|
+
the server's payload straight into `_meta.report_data` and
|
|
123
|
+
`structuredContent`, bypassing the one projection that bounds the Reddit note's
|
|
124
|
+
code fields and removes an internal task id — and the scoped rail is the path
|
|
125
|
+
that writes a Reddit note today. The values were code-stamped, so nothing
|
|
126
|
+
escaped; the check now runs where the note actually travels.
|
|
127
|
+
|
|
128
|
+
### Added
|
|
129
|
+
|
|
130
|
+
- **`personas[].sources[].receipts[]`** — the receipt children of a source whose
|
|
131
|
+
one URL yielded several excerpts. Absent on every other source, which is what
|
|
132
|
+
makes a legacy report render byte-identically.
|
|
133
|
+
- **`methodology.redditRetrieval`** — the frozen, code-stamped outcome of the
|
|
134
|
+
Reddit community-evidence pass, mirroring the producer's own summary shape
|
|
135
|
+
(typed outcome, bounded reason code, the plan/completed/failed intent split,
|
|
136
|
+
request and evidence counts, `costUsd`, `latencyMs`), rendered as a short
|
|
137
|
+
`### Reddit community evidence` block. Five typed states; `failed` is never presented as an empty
|
|
138
|
+
search, and **absent means unrecorded** rather than disabled or empty. The
|
|
139
|
+
digest says nothing at all for an absent note or `not_run`. Cost, latency and
|
|
140
|
+
request counts are typed for auditing and deliberately not rendered — none of
|
|
141
|
+
them changes what an agent may cite.
|
|
142
|
+
|
|
143
|
+
### Changed
|
|
144
|
+
|
|
145
|
+
- **A `receiptId` that does not resolve under its own source loses the receipt**
|
|
146
|
+
rather than falling back to the parent excerpt. Malformed, dangling,
|
|
147
|
+
cross-source and duplicate ids all degrade to the existing "receipt
|
|
148
|
+
unavailable; do not cite" state. Falling back would restore the original defect
|
|
149
|
+
as an error path, firing only on the inputs nobody is watching.
|
|
150
|
+
- **An outcome this client does not know reads as UNKNOWN, not as silence.**
|
|
151
|
+
`methodology.redditRetrieval.outcome` is typed loosely on purpose: this
|
|
152
|
+
package and the app ship on independent release cycles, so a client older
|
|
153
|
+
than the deploy answering it is the normal pairing. A strict enum would have
|
|
154
|
+
turned a future sixth state into an absent note — and absent means
|
|
155
|
+
*unrecorded* — so the digest would have fallen silent about a run that really
|
|
156
|
+
did search Reddit. It now names the state it cannot read and tells you to
|
|
157
|
+
upgrade.
|
|
158
|
+
- **One malformed receipt child no longer costs the whole report its typing.**
|
|
159
|
+
`personas[].sources[].receipts` degrades to absent rather than failing the
|
|
160
|
+
`report_data` parse, matching how the app stores it. A claim naming one of
|
|
161
|
+
those children then fails closed, which is the only cost it should have.
|
|
162
|
+
- **"Verbatim" retired where the text is an excerpt.** `relevantQuotes[]` on
|
|
163
|
+
`search_forums` / `list_community_signals`, and `quoteSpan`'s description, no
|
|
164
|
+
longer promise a speaker's exact words: these fields are fed by more than one
|
|
165
|
+
retrieval rail and only "cached excerpt" is true of all of them. The span is
|
|
166
|
+
still exact against the text we cached — that is what verification means.
|
|
167
|
+
|
|
16
168
|
## 0.8.0
|
|
17
169
|
|
|
18
170
|
**The persona WRITE tools get the guards the read tools already had.** 0.6.0 and
|
package/README.md
CHANGED
|
@@ -8,7 +8,7 @@ competitors, and key findings) with a typed
|
|
|
8
8
|
[trust layer](#trust-signals) — per-claim GROUNDED/SPECULATION/NO_RECEIPT
|
|
9
9
|
classifications, source receipts, and robustness survival counts.
|
|
10
10
|
|
|
11
|
-
Alongside it, [
|
|
11
|
+
Alongside it, [the rest of the toolset](#tools) lets an agent read and edit what a project
|
|
12
12
|
already holds — projects, personas, interviews, reports, competitors, and
|
|
13
13
|
community signals — most of them free.
|
|
14
14
|
|
|
@@ -86,7 +86,24 @@ Restart Claude Code. First call triggers OAuth.
|
|
|
86
86
|
{ "name": "...", "sources": [{ "url": "...", "platform": "Hacker News", "quote": "...", "retrievedAt": "..." }] }
|
|
87
87
|
],
|
|
88
88
|
"hypothesisResults": [
|
|
89
|
-
{
|
|
89
|
+
{
|
|
90
|
+
"hypothesisId": "h1",
|
|
91
|
+
// Legacy compatibility fields: preserve, but do not rank or compare them.
|
|
92
|
+
"status": "validated", "confidence": 0.82,
|
|
93
|
+
// The authoritative effective result on current reports.
|
|
94
|
+
"final": {
|
|
95
|
+
"version": 1, "verdict": "validated", "originalVerdict": "validated",
|
|
96
|
+
"confidence": "high", "confidenceBasis": "evidence",
|
|
97
|
+
"sufficiency": "substantial", "robustness": "held", "priority": "high"
|
|
98
|
+
},
|
|
99
|
+
// Raw re-ask detail retained for audit; `final` above owns the effective result.
|
|
100
|
+
"robustness": {
|
|
101
|
+
"originalStatus": "validated", "survived": 1, "total": 1, "flipped": false,
|
|
102
|
+
"variants": [
|
|
103
|
+
{ "framing": "...", "returnedStatus": "validated", "returnedConfidence": 0.84, "agreed": true, "reasoning": "..." }
|
|
104
|
+
]
|
|
105
|
+
}
|
|
106
|
+
}
|
|
90
107
|
],
|
|
91
108
|
"sycophancySignals": { "lowDiscriminationCount": 0, "totalRejections": 2, "personaSignals": [...], "disagreements": [...] },
|
|
92
109
|
"openQuestions": [
|
|
@@ -160,7 +177,10 @@ Every field below reaches the agent-visible result text on any host, not only
|
|
|
160
177
|
`structuredContent`/`_meta`. Most are **summarised into the appended trust
|
|
161
178
|
digest**; `openQuestions[]` is the exception — it comes through instead as its
|
|
162
179
|
own `## Open Questions` section in the primary markdown report, ahead of the
|
|
163
|
-
digest.
|
|
180
|
+
digest. `methodology.redditRetrieval` is the other exception, in the opposite
|
|
181
|
+
direction: two of its states (`not_run`, and the field being absent entirely)
|
|
182
|
+
render **nothing at all**, because saying anything would claim the run searched
|
|
183
|
+
Reddit. The typed fields themselves are in `structuredContent.report_data`
|
|
164
184
|
(and `_meta.report_data`) when you want to read them programmatically.
|
|
165
185
|
|
|
166
186
|
When a claim/receipt/robustness/sycophancy section is missing, the digest says
|
|
@@ -176,10 +196,13 @@ a finding is, rather than taking the prose at face value:
|
|
|
176
196
|
|
|
177
197
|
| Field | What it tells you |
|
|
178
198
|
|------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------|
|
|
179
|
-
| `claims[]` | Every persona/interview claim with a `state`: `GROUNDED` (span-verified against a cached source
|
|
180
|
-
| `personas[].sources[]` | The cached forum posts (`url`, `platform`, `quote`, `retrievedAt`) a claim's `sourceId` (`RCP-p{i}-s{j}`) points at — the receipts. In the digest text each
|
|
199
|
+
| `claims[]` | Every persona/interview claim with a `state`: `GROUNDED` (span-verified against a cached source excerpt), `SPECULATION`, or `NO_RECEIPT` (evidence-shaped but unsourced). A `GROUNDED` claim carries a `sourceId` receipt and the `quoteSpan` that grounds it, plus a `receiptId` when one discussion yielded several excerpts (see the row below). |
|
|
200
|
+
| `personas[].sources[]` | The cached forum posts (`url`, `platform`, `quote`, `retrievedAt`) a claim's `sourceId` (`RCP-p{i}-s{j}`) points at — the receipts. In the digest text each excerpt is shown as a 200-character window **around the span that grounded the claim**, with a leading `…` when the window starts mid-excerpt; the untruncated text is in the structured payload. |
|
|
201
|
+
| `personas[].sources[].receipts[]` | Present only when **one discussion yielded several excerpts**. Each child is `{ receiptId, excerpt, retrievedAt }`, and `receiptId` is content-addressed from the normalized URL plus the exact excerpt. A claim addressing one shows a second id after `·` on its digest line, and the receipt pool hangs that excerpt off the source's single row on its own `↳` line — read that, not the row's first excerpt, as what the claim was checked against. The source still counts as **one** discussion; several excerpts never inflate the breadth. A `receiptId` that does not resolve under its own source loses the receipt arrow entirely rather than falling back. |
|
|
202
|
+
| `methodology.redditRetrieval` | Present only on runs that recorded it. The frozen, code-stamped outcome of the Reddit community-evidence pass: `completed`, `completed_empty`, `completed_partial`, `failed` or `not_run`, with a bounded reason code and accepted discussion / excerpt counts. **Absent means unrecorded** — not disabled, not empty, not failed. The digest renders a sentence for the four states a reader has to act on and says nothing for `not_run` or unrecorded. |
|
|
181
203
|
| `personas[].qaFlagged` / `insufficientEvidence` / `sourcesFound` | Per-persona quality flags: `qaFlagged` means the persona failed a QA check; `insufficientEvidence` / `sourcesFound` tell you whether a `NO_RECEIPT` reflects a weak claim or just a thin evidence base. |
|
|
182
|
-
| `hypothesisResults[].
|
|
204
|
+
| `hypothesisResults[].final` | The authoritative effective result on current reports: `verdict`, categorical reliability in `confidence`, evidence coverage in `sufficiency`, and the `robustness` outcome. `confidence` may be absent when reliability was withheld; read `confidenceBasis`. Raw top-level `status` and numeric `confidence` are legacy compatibility fields and must not be ranked, trended, or compared with `final.confidence`. |
|
|
205
|
+
| `hypothesisResults[].robustness` | Raw re-ask detail retained for audit, including survival counts. Read `final.verdict` and `final.robustness` for the effective result; never independently promote `downgradedStatus` or the top-level `status`. |
|
|
183
206
|
| `sycophancySignals` | Anti-sycophancy readout: broken-persona flags (`lowDiscrimination`), per-hypothesis disagreement splits, and explicit `totalRejections`. |
|
|
184
207
|
| `openQuestions[]` | What the run did **not** settle. Each row carries the hypothesis and a `status` telling you why it's open — `never_asked`, `nobody_answered`, or `voices_disagreed` — plus `personasAsked`, how many distinct personas it was put to. Deterministically derived, never model-opined. Reaches the result text via the report's own `## Open Questions` section (see above), not via the trust digest; absent on reports from before this field shipped, an empty array means nothing was left open. |
|
|
185
208
|
|
|
@@ -191,7 +214,7 @@ simply omits them. The full Zod schema lives in
|
|
|
191
214
|
|
|
192
215
|
## Tools
|
|
193
216
|
|
|
194
|
-
|
|
217
|
+
`clien_research` is the hero — the recommended default for
|
|
195
218
|
validating an idea end to end. Everything else is either a **free read** for
|
|
196
219
|
orienting on what a project already holds, or a **targeted primitive** for a
|
|
197
220
|
narrow, cheaper probe.
|
|
@@ -262,9 +285,9 @@ name — the aliases exist for callers written before the rename.
|
|
|
262
285
|
| `list_reports` | — | free | Past `clien_research` runs, newest first. Optionally scoped to one `project_id`. |
|
|
263
286
|
| `get_report` | `job_id` | free | A past run's full report + `report_data`, same shape `clien_research` returns. Use this instead of re-running. |
|
|
264
287
|
| `list_competitors` | `project_id` | free | The project's deduped, accumulated competitor set, most-recently-seen first. Each of the four company facts carries a sibling `<fact>Grounding` — key on its `state`. |
|
|
265
|
-
| `list_community_signals` | `project_id` | free | The project's deduped, accumulated thread set, newest-**discovered** first (a stable order — a thread seen again does not move back to page one). Carries the
|
|
288
|
+
| `list_community_signals` | `project_id` | free | The project's deduped, accumulated thread set, newest-**discovered** first (a stable order — a thread seen again does not move back to page one). Carries the cached `relevantQuotes[]` excerpts captured at collection; they are user-voice evidence but are not guaranteed to preserve a speaker's exact words. |
|
|
266
289
|
| `get_market` | `project_id` | free | The project's market snapshot — size, growth trend, key trends, positioning gaps. A **singleton**, not a page: every field was written together by the one run named in `sourceRunId`. No snapshot is reported as "no run has produced market data", never as an empty or unmeasured market. The figures carry no trust state here; read the source run with `get_report` for their claim states. |
|
|
267
|
-
| `list_interview_scripts` | `project_id` | free | The interview scripts a research run wrote for the project — one per hypothesis, ranked
|
|
290
|
+
| `list_interview_scripts` | `project_id` | free | The interview scripts a research run wrote for the project — one per hypothesis, ranked by product priority, then lower verdict reliability, then thinner evidence. Each carries its `hypothesisId`, statement, verdict, questions, categorical `confidence`, `sufficiency`, and `robustness`; older runs expose the retired numeric measure as `legacyConfidence`, which is not comparable with `low` / `medium` / `high`. **Stored output of one run, never generated on read.** Call it before `interview_persona` so the answers anchor to a hypothesis; no scripts is reported as "no run has written any", never as "nothing worth asking". |
|
|
268
291
|
|
|
269
292
|
### Account
|
|
270
293
|
|
|
@@ -320,7 +343,7 @@ same payload as `structuredContent`; see
|
|
|
320
343
|
| `job_id` | UUID | The validation job's ID |
|
|
321
344
|
| `status` | string | Terminal status: `complete`, `failed`, or `cancelled` |
|
|
322
345
|
| `project_id` | UUID \| `null` | The project this run is attached to. Pass back as `project_id` on follow-up runs. |
|
|
323
|
-
| `report_data` | object \| `null`| Structured report framing + findings, including typed `personasSynthesis`, `researchBrief`, `methodology`, safe `interviewHighlights[]`, competitors, hypotheses, and the trust layer — `claims[]`, `personas[].sources[]`, `hypothesisResults[].robustness`, `sycophancySignals`, `openQuestions[]` (see [Trust signals](#trust-signals)) |
|
|
346
|
+
| `report_data` | object \| `null`| Structured report framing + findings, including typed `personasSynthesis`, `researchBrief`, `methodology`, safe `interviewHighlights[]`, competitors, hypotheses, and the trust layer — `claims[]`, `personas[].sources[]`, authoritative `hypothesisResults[].final`, raw `hypothesisResults[].robustness`, `sycophancySignals`, `openQuestions[]` (see [Trust signals](#trust-signals)) |
|
|
324
347
|
| `total_cost_cents` | number | Optional. Total cost of the run, in cents. |
|
|
325
348
|
|
|
326
349
|
---
|
|
@@ -0,0 +1,230 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Pure FUL-358 hypothesis semantics for the published MCP package.
|
|
3
|
+
*
|
|
4
|
+
* This module deliberately has no Zod or tool imports. Both the transport projection and the
|
|
5
|
+
* human-readable digest call the same normalizer, so a raw-drift payload cannot be trusted by one
|
|
6
|
+
* exit and rejected by another. The returned object is an allowlist: producer-private or future
|
|
7
|
+
* fields never cross the current contract boundary merely because an enclosing schema is
|
|
8
|
+
* passthrough-compatible.
|
|
9
|
+
*/
|
|
10
|
+
export const VERDICT_STATUSES = ['validated', 'invalidated', 'inconclusive'];
|
|
11
|
+
export const CONFIDENCE_LEVELS = ['low', 'medium', 'high'];
|
|
12
|
+
export const EVIDENCE_SUFFICIENCY_LEVELS = ['none', 'thin', 'moderate', 'substantial'];
|
|
13
|
+
export const ROBUSTNESS_OUTCOMES = ['held', 'flipped', 'not_retested'];
|
|
14
|
+
export const CONFIDENCE_BASES = ['evidence', 'recomputed_after_flip', 'withheld_no_evidence'];
|
|
15
|
+
export const HYPOTHESIS_PRIORITIES = ['high', 'medium', 'low'];
|
|
16
|
+
/** The semantics contract version this MCP build knows how to read. */
|
|
17
|
+
export const HYPOTHESIS_SEMANTICS_VERSION = 1;
|
|
18
|
+
const VERDICT_STRENGTH = {
|
|
19
|
+
validated: 2,
|
|
20
|
+
inconclusive: 1,
|
|
21
|
+
invalidated: 0,
|
|
22
|
+
};
|
|
23
|
+
function asRecord(value) {
|
|
24
|
+
return typeof value === 'object' && value !== null && !Array.isArray(value)
|
|
25
|
+
? value
|
|
26
|
+
: null;
|
|
27
|
+
}
|
|
28
|
+
function asMember(value, members) {
|
|
29
|
+
return typeof value === 'string' && members.includes(value)
|
|
30
|
+
? value
|
|
31
|
+
: null;
|
|
32
|
+
}
|
|
33
|
+
/** The weaker of the original and candidate verdicts; an attempted upgrade is ignored. */
|
|
34
|
+
export function resolveMonotoneVerdict(rawOriginal, rawCandidate) {
|
|
35
|
+
const original = asMember(rawOriginal, VERDICT_STATUSES);
|
|
36
|
+
const candidate = asMember(rawCandidate, VERDICT_STATUSES);
|
|
37
|
+
if (!candidate)
|
|
38
|
+
return original;
|
|
39
|
+
if (!original)
|
|
40
|
+
return candidate;
|
|
41
|
+
return VERDICT_STRENGTH[candidate] < VERDICT_STRENGTH[original] ? candidate : original;
|
|
42
|
+
}
|
|
43
|
+
/**
|
|
44
|
+
* Read a rephrasing ratio only when it can describe a real finite count whose `flipped` flag
|
|
45
|
+
* agrees with whether every framing survived. A malformed ratio still cannot erase a separately
|
|
46
|
+
* readable weaker recorded verdict.
|
|
47
|
+
*/
|
|
48
|
+
export function readRobustnessCounts(value) {
|
|
49
|
+
const robustness = asRecord(value);
|
|
50
|
+
if (!robustness)
|
|
51
|
+
return null;
|
|
52
|
+
const survived = robustness.survived;
|
|
53
|
+
const total = robustness.total;
|
|
54
|
+
const flipped = robustness.flipped;
|
|
55
|
+
if (typeof survived !== 'number' ||
|
|
56
|
+
!Number.isInteger(survived) ||
|
|
57
|
+
survived < 0 ||
|
|
58
|
+
typeof total !== 'number' ||
|
|
59
|
+
!Number.isInteger(total) ||
|
|
60
|
+
total <= 0 ||
|
|
61
|
+
survived > total ||
|
|
62
|
+
typeof flipped !== 'boolean' ||
|
|
63
|
+
flipped !== (survived < total)) {
|
|
64
|
+
return null;
|
|
65
|
+
}
|
|
66
|
+
return { survived, total };
|
|
67
|
+
}
|
|
68
|
+
/** A typed robustness block is trustworthy only when its detailed working proves its summary. */
|
|
69
|
+
export function hasCoherentRobustnessWorking(value) {
|
|
70
|
+
const robustness = asRecord(value);
|
|
71
|
+
const counts = readRobustnessCounts(robustness);
|
|
72
|
+
const originalStatus = asMember(robustness?.originalStatus, VERDICT_STATUSES);
|
|
73
|
+
const variants = robustness?.variants;
|
|
74
|
+
if (!robustness || !counts || !originalStatus || !Array.isArray(variants))
|
|
75
|
+
return false;
|
|
76
|
+
if (variants.length !== counts.total)
|
|
77
|
+
return false;
|
|
78
|
+
let agreedCount = 0;
|
|
79
|
+
for (const raw of variants) {
|
|
80
|
+
const variant = asRecord(raw);
|
|
81
|
+
const returnedStatus = asMember(variant?.returnedStatus, VERDICT_STATUSES);
|
|
82
|
+
if (!variant || !returnedStatus || typeof variant.agreed !== 'boolean')
|
|
83
|
+
return false;
|
|
84
|
+
if (variant.agreed !== (returnedStatus === originalStatus))
|
|
85
|
+
return false;
|
|
86
|
+
if (variant.agreed)
|
|
87
|
+
agreedCount += 1;
|
|
88
|
+
}
|
|
89
|
+
return agreedCount === counts.survived;
|
|
90
|
+
}
|
|
91
|
+
/**
|
|
92
|
+
* Whether present working positively disproves its own summary. Absent working and individually
|
|
93
|
+
* unreadable variant rows do not by themselves disprove a separately valid count. A present
|
|
94
|
+
* non-array `variants` container does: it is an attempted shape that cannot authenticate dependent
|
|
95
|
+
* final semantics or cross the raw projection boundary.
|
|
96
|
+
*/
|
|
97
|
+
export function robustnessWorkingContradictsSummary(value) {
|
|
98
|
+
const robustness = asRecord(value);
|
|
99
|
+
if (!robustness)
|
|
100
|
+
return false;
|
|
101
|
+
const counts = readRobustnessCounts(robustness);
|
|
102
|
+
const originalStatus = asMember(robustness.originalStatus, VERDICT_STATUSES);
|
|
103
|
+
const variants = robustness.variants;
|
|
104
|
+
if (!Object.prototype.hasOwnProperty.call(robustness, 'variants'))
|
|
105
|
+
return false;
|
|
106
|
+
if (!Array.isArray(variants))
|
|
107
|
+
return true;
|
|
108
|
+
if (!originalStatus)
|
|
109
|
+
return true;
|
|
110
|
+
let readableCount = 0;
|
|
111
|
+
let agreedCount = 0;
|
|
112
|
+
let disagreedCount = 0;
|
|
113
|
+
for (const raw of variants) {
|
|
114
|
+
const variant = asRecord(raw);
|
|
115
|
+
const returnedStatus = asMember(variant?.returnedStatus, VERDICT_STATUSES);
|
|
116
|
+
if (!variant || !returnedStatus || typeof variant.agreed !== 'boolean')
|
|
117
|
+
continue;
|
|
118
|
+
readableCount += 1;
|
|
119
|
+
if (variant.agreed !== (returnedStatus === originalStatus))
|
|
120
|
+
return true;
|
|
121
|
+
if (variant.agreed)
|
|
122
|
+
agreedCount += 1;
|
|
123
|
+
else
|
|
124
|
+
disagreedCount += 1;
|
|
125
|
+
}
|
|
126
|
+
if (!counts)
|
|
127
|
+
return false;
|
|
128
|
+
if (variants.length > counts.total)
|
|
129
|
+
return true;
|
|
130
|
+
if (agreedCount > counts.survived || disagreedCount > counts.total - counts.survived)
|
|
131
|
+
return true;
|
|
132
|
+
return (readableCount === variants.length &&
|
|
133
|
+
variants.length === counts.total &&
|
|
134
|
+
agreedCount !== counts.survived);
|
|
135
|
+
}
|
|
136
|
+
/** Read robustness in the context of its owning row, including parent-verdict coherence. */
|
|
137
|
+
export function readRowRobustnessOutcome(rawRow) {
|
|
138
|
+
const row = asRecord(rawRow);
|
|
139
|
+
if (!row)
|
|
140
|
+
return null;
|
|
141
|
+
if (row.robustness === undefined || row.robustness === null)
|
|
142
|
+
return 'not_retested';
|
|
143
|
+
const robustness = asRecord(row.robustness);
|
|
144
|
+
if (!robustness || !readRobustnessCounts(robustness))
|
|
145
|
+
return null;
|
|
146
|
+
const rowStatus = asMember(row.status, VERDICT_STATUSES);
|
|
147
|
+
if (!rowStatus)
|
|
148
|
+
return null;
|
|
149
|
+
if (Object.prototype.hasOwnProperty.call(robustness, 'originalStatus') &&
|
|
150
|
+
asMember(robustness.originalStatus, VERDICT_STATUSES) !== rowStatus)
|
|
151
|
+
return null;
|
|
152
|
+
if (Object.prototype.hasOwnProperty.call(robustness, 'variants') &&
|
|
153
|
+
robustnessWorkingContradictsSummary(robustness))
|
|
154
|
+
return null;
|
|
155
|
+
return robustness.flipped ? 'flipped' : 'held';
|
|
156
|
+
}
|
|
157
|
+
/**
|
|
158
|
+
* Read a legacy robustness outcome without treating a contradicted measurement as truth.
|
|
159
|
+
* Old rows may carry only `flipped`; once either count is present, the whole measured triple
|
|
160
|
+
* must agree before its outcome can be rendered.
|
|
161
|
+
*/
|
|
162
|
+
export function readLegacyRobustnessOutcome(value) {
|
|
163
|
+
if (value === undefined || value === null)
|
|
164
|
+
return 'not_retested';
|
|
165
|
+
const robustness = asRecord(value);
|
|
166
|
+
if (!robustness || typeof robustness.flipped !== 'boolean')
|
|
167
|
+
return null;
|
|
168
|
+
const carriesMeasuredCounts = Object.prototype.hasOwnProperty.call(robustness, 'survived') ||
|
|
169
|
+
Object.prototype.hasOwnProperty.call(robustness, 'total');
|
|
170
|
+
if (carriesMeasuredCounts && readRobustnessCounts(robustness) === null)
|
|
171
|
+
return null;
|
|
172
|
+
return robustness.flipped ? 'flipped' : 'held';
|
|
173
|
+
}
|
|
174
|
+
function finalAxesAreCoherent(final) {
|
|
175
|
+
const confidence = final.confidence ?? null;
|
|
176
|
+
if (final.confidenceBasis === 'withheld_no_evidence') {
|
|
177
|
+
return confidence === null && final.sufficiency === 'none';
|
|
178
|
+
}
|
|
179
|
+
if (final.confidenceBasis === 'recomputed_after_flip') {
|
|
180
|
+
return confidence === 'low' && final.sufficiency !== 'none' && final.robustness === 'flipped';
|
|
181
|
+
}
|
|
182
|
+
return confidence !== null && final.sufficiency !== 'none' && final.robustness !== 'flipped';
|
|
183
|
+
}
|
|
184
|
+
/**
|
|
185
|
+
* Authenticate and normalize a producer-owned `final` block against its raw row.
|
|
186
|
+
*
|
|
187
|
+
* A readable row status is mandatory: it is the original verdict the code-derived block claims
|
|
188
|
+
* to reconcile. The final robustness outcome must also agree with the raw pass (or its absence),
|
|
189
|
+
* otherwise a stale pre-flip reliability block could caption a flipped row. Optional priority
|
|
190
|
+
* degrades field-locally; every unknown key is removed from the returned current-version block.
|
|
191
|
+
*/
|
|
192
|
+
export function normalizeFinalHypothesisState(rawRow) {
|
|
193
|
+
const row = asRecord(rawRow);
|
|
194
|
+
const block = asRecord(row?.final);
|
|
195
|
+
if (!row || !block || block.version !== HYPOTHESIS_SEMANTICS_VERSION)
|
|
196
|
+
return null;
|
|
197
|
+
const rowVerdict = asMember(row.status, VERDICT_STATUSES);
|
|
198
|
+
const originalVerdict = asMember(block.originalVerdict, VERDICT_STATUSES);
|
|
199
|
+
const verdict = asMember(block.verdict, VERDICT_STATUSES);
|
|
200
|
+
if (!rowVerdict || originalVerdict !== rowVerdict || !verdict)
|
|
201
|
+
return null;
|
|
202
|
+
const rawRobustness = asRecord(row.robustness);
|
|
203
|
+
const weakestRecorded = resolveMonotoneVerdict(rowVerdict, rawRobustness?.downgradedStatus);
|
|
204
|
+
if (resolveMonotoneVerdict(weakestRecorded, verdict) !== verdict)
|
|
205
|
+
return null;
|
|
206
|
+
const confidence = asMember(block.confidence, CONFIDENCE_LEVELS);
|
|
207
|
+
if (block.confidence !== undefined && confidence === null)
|
|
208
|
+
return null;
|
|
209
|
+
const confidenceBasis = asMember(block.confidenceBasis, CONFIDENCE_BASES);
|
|
210
|
+
const sufficiency = asMember(block.sufficiency, EVIDENCE_SUFFICIENCY_LEVELS);
|
|
211
|
+
const robustness = asMember(block.robustness, ROBUSTNESS_OUTCOMES);
|
|
212
|
+
if (!confidenceBasis || !sufficiency || !robustness)
|
|
213
|
+
return null;
|
|
214
|
+
if (readRowRobustnessOutcome(row) !== robustness)
|
|
215
|
+
return null;
|
|
216
|
+
const normalized = {
|
|
217
|
+
version: HYPOTHESIS_SEMANTICS_VERSION,
|
|
218
|
+
verdict,
|
|
219
|
+
originalVerdict,
|
|
220
|
+
...(confidence ? { confidence } : {}),
|
|
221
|
+
confidenceBasis,
|
|
222
|
+
sufficiency,
|
|
223
|
+
robustness,
|
|
224
|
+
};
|
|
225
|
+
const priority = asMember(block.priority, HYPOTHESIS_PRIORITIES);
|
|
226
|
+
if (priority)
|
|
227
|
+
normalized.priority = priority;
|
|
228
|
+
return finalAxesAreCoherent(normalized) ? normalized : null;
|
|
229
|
+
}
|
|
230
|
+
//# sourceMappingURL=hypothesis-semantics.js.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"hypothesis-semantics.js","sourceRoot":"","sources":["../src/hypothesis-semantics.ts"],"names":[],"mappings":"AAAA;;;;;;;;GAQG;AAEH,MAAM,CAAC,MAAM,gBAAgB,GAAG,CAAC,WAAW,EAAE,aAAa,EAAE,cAAc,CAAU,CAAA;AAGrF,MAAM,CAAC,MAAM,iBAAiB,GAAG,CAAC,KAAK,EAAE,QAAQ,EAAE,MAAM,CAAU,CAAA;AAGnE,MAAM,CAAC,MAAM,2BAA2B,GAAG,CAAC,MAAM,EAAE,MAAM,EAAE,UAAU,EAAE,aAAa,CAAU,CAAA;AAG/F,MAAM,CAAC,MAAM,mBAAmB,GAAG,CAAC,MAAM,EAAE,SAAS,EAAE,cAAc,CAAU,CAAA;AAG/E,MAAM,CAAC,MAAM,gBAAgB,GAAG,CAAC,UAAU,EAAE,uBAAuB,EAAE,sBAAsB,CAAU,CAAA;AAGtG,MAAM,CAAC,MAAM,qBAAqB,GAAG,CAAC,MAAM,EAAE,QAAQ,EAAE,KAAK,CAAU,CAAA;AAGvE,uEAAuE;AACvE,MAAM,CAAC,MAAM,4BAA4B,GAAG,CAAC,CAAA;AAiB7C,MAAM,gBAAgB,GAAkC;IACtD,SAAS,EAAE,CAAC;IACZ,YAAY,EAAE,CAAC;IACf,WAAW,EAAE,CAAC;CACf,CAAA;AAED,SAAS,QAAQ,CAAC,KAAc;IAC9B,OAAO,OAAO,KAAK,KAAK,QAAQ,IAAI,KAAK,KAAK,IAAI,IAAI,CAAC,KAAK,CAAC,OAAO,CAAC,KAAK,CAAC;QACzE,CAAC,CAAE,KAAiC;QACpC,CAAC,CAAC,IAAI,CAAA;AACV,CAAC;AAED,SAAS,QAAQ,CAAmB,KAAc,EAAE,OAAqB;IACvE,OAAO,OAAO,KAAK,KAAK,QAAQ,IAAK,OAA6B,CAAC,QAAQ,CAAC,KAAK,CAAC;QAChF,CAAC,CAAE,KAAW;QACd,CAAC,CAAC,IAAI,CAAA;AACV,CAAC;AAED,0FAA0F;AAC1F,MAAM,UAAU,sBAAsB,CACpC,WAAoB,EACpB,YAAqB;IAErB,MAAM,QAAQ,GAAG,QAAQ,CAAC,WAAW,EAAE,gBAAgB,CAAC,CAAA;IACxD,MAAM,SAAS,GAAG,QAAQ,CAAC,YAAY,EAAE,gBAAgB,CAAC,CAAA;IAC1D,IAAI,CAAC,SAAS;QAAE,OAAO,QAAQ,CAAA;IAC/B,IAAI,CAAC,QAAQ;QAAE,OAAO,SAAS,CAAA;IAC/B,OAAO,gBAAgB,CAAC,SAAS,CAAC,GAAG,gBAAgB,CAAC,QAAQ,CAAC,CAAC,CAAC,CAAC,SAAS,CAAC,CAAC,CAAC,QAAQ,CAAA;AACxF,CAAC;AAED;;;;GAIG;AACH,MAAM,UAAU,oBAAoB,CAAC,KAAc;IACjD,MAAM,UAAU,GAAG,QAAQ,CAAC,KAAK,CAAC,CAAA;IAClC,IAAI,CAAC,UAAU;QAAE,OAAO,IAAI,CAAA;IAC5B,MAAM,QAAQ,GAAG,UAAU,CAAC,QAAQ,CAAA;IACpC,MAAM,KAAK,GAAG,UAAU,CAAC,KAAK,CAAA;IAC9B,MAAM,OAAO,GAAG,UAAU,CAAC,OAAO,CAAA;IAClC,IACE,OAAO,QAAQ,KAAK,QAAQ;QAC5B,CAAC,MAAM,CAAC,SAAS,CAAC,QAAQ,CAAC;QAC3B,QAAQ,GAAG,CAAC;QACZ,OAAO,KAAK,KAAK,QAAQ;QACzB,CAAC,MAAM,CAAC,SAAS,CAAC,KAAK,CAAC;QACxB,KAAK,IAAI,CAAC;QACV,QAAQ,GAAG,KAAK;QAChB,OAAO,OAAO,KAAK,SAAS;QAC5B,OAAO,KAAK,CAAC,QAAQ,GAAG,KAAK,CAAC,EAC9B,CAAC;QACD,OAAO,IAAI,CAAA;IACb,CAAC;IACD,OAAO,EAAE,QAAQ,EAAE,KAAK,EAAE,CAAA;AAC5B,CAAC;AAED,iGAAiG;AACjG,MAAM,UAAU,4BAA4B,CAAC,KAAc;IACzD,MAAM,UAAU,GAAG,QAAQ,CAAC,KAAK,CAAC,CAAA;IAClC,MAAM,MAAM,GAAG,oBAAoB,CAAC,UAAU,CAAC,CAAA;IAC/C,MAAM,cAAc,GAAG,QAAQ,CAAC,UAAU,EAAE,cAAc,EAAE,gBAAgB,CAAC,CAAA;IAC7E,MAAM,QAAQ,GAAG,UAAU,EAAE,QAAQ,CAAA;IACrC,IAAI,CAAC,UAAU,IAAI,CAAC,MAAM,IAAI,CAAC,cAAc,IAAI,CAAC,KAAK,CAAC,OAAO,CAAC,QAAQ,CAAC;QAAE,OAAO,KAAK,CAAA;IACvF,IAAI,QAAQ,CAAC,MAAM,KAAK,MAAM,CAAC,KAAK;QAAE,OAAO,KAAK,CAAA;IAElD,IAAI,WAAW,GAAG,CAAC,CAAA;IACnB,KAAK,MAAM,GAAG,IAAI,QAAQ,EAAE,CAAC;QAC3B,MAAM,OAAO,GAAG,QAAQ,CAAC,GAAG,CAAC,CAAA;QAC7B,MAAM,cAAc,GAAG,QAAQ,CAAC,OAAO,EAAE,cAAc,EAAE,gBAAgB,CAAC,CAAA;QAC1E,IAAI,CAAC,OAAO,IAAI,CAAC,cAAc,IAAI,OAAO,OAAO,CAAC,MAAM,KAAK,SAAS;YAAE,OAAO,KAAK,CAAA;QACpF,IAAI,OAAO,CAAC,MAAM,KAAK,CAAC,cAAc,KAAK,cAAc,CAAC;YAAE,OAAO,KAAK,CAAA;QACxE,IAAI,OAAO,CAAC,MAAM;YAAE,WAAW,IAAI,CAAC,CAAA;IACtC,CAAC;IACD,OAAO,WAAW,KAAK,MAAM,CAAC,QAAQ,CAAA;AACxC,CAAC;AAED;;;;;GAKG;AACH,MAAM,UAAU,mCAAmC,CAAC,KAAc;IAChE,MAAM,UAAU,GAAG,QAAQ,CAAC,KAAK,CAAC,CAAA;IAClC,IAAI,CAAC,UAAU;QAAE,OAAO,KAAK,CAAA;IAC7B,MAAM,MAAM,GAAG,oBAAoB,CAAC,UAAU,CAAC,CAAA;IAC/C,MAAM,cAAc,GAAG,QAAQ,CAAC,UAAU,CAAC,cAAc,EAAE,gBAAgB,CAAC,CAAA;IAC5E,MAAM,QAAQ,GAAG,UAAU,CAAC,QAAQ,CAAA;IACpC,IAAI,CAAC,MAAM,CAAC,SAAS,CAAC,cAAc,CAAC,IAAI,CAAC,UAAU,EAAE,UAAU,CAAC;QAAE,OAAO,KAAK,CAAA;IAC/E,IAAI,CAAC,KAAK,CAAC,OAAO,CAAC,QAAQ,CAAC;QAAE,OAAO,IAAI,CAAA;IACzC,IAAI,CAAC,cAAc;QAAE,OAAO,IAAI,CAAA;IAEhC,IAAI,aAAa,GAAG,CAAC,CAAA;IACrB,IAAI,WAAW,GAAG,CAAC,CAAA;IACnB,IAAI,cAAc,GAAG,CAAC,CAAA;IACtB,KAAK,MAAM,GAAG,IAAI,QAAQ,EAAE,CAAC;QAC3B,MAAM,OAAO,GAAG,QAAQ,CAAC,GAAG,CAAC,CAAA;QAC7B,MAAM,cAAc,GAAG,QAAQ,CAAC,OAAO,EAAE,cAAc,EAAE,gBAAgB,CAAC,CAAA;QAC1E,IAAI,CAAC,OAAO,IAAI,CAAC,cAAc,IAAI,OAAO,OAAO,CAAC,MAAM,KAAK,SAAS;YAAE,SAAQ;QAChF,aAAa,IAAI,CAAC,CAAA;QAClB,IAAI,OAAO,CAAC,MAAM,KAAK,CAAC,cAAc,KAAK,cAAc,CAAC;YAAE,OAAO,IAAI,CAAA;QACvE,IAAI,OAAO,CAAC,MAAM;YAAE,WAAW,IAAI,CAAC,CAAA;;YAC/B,cAAc,IAAI,CAAC,CAAA;IAC1B,CAAC;IAED,IAAI,CAAC,MAAM;QAAE,OAAO,KAAK,CAAA;IACzB,IAAI,QAAQ,CAAC,MAAM,GAAG,MAAM,CAAC,KAAK;QAAE,OAAO,IAAI,CAAA;IAC/C,IAAI,WAAW,GAAG,MAAM,CAAC,QAAQ,IAAI,cAAc,GAAG,MAAM,CAAC,KAAK,GAAG,MAAM,CAAC,QAAQ;QAAE,OAAO,IAAI,CAAA;IACjG,OAAO,CACL,aAAa,KAAK,QAAQ,CAAC,MAAM;QACjC,QAAQ,CAAC,MAAM,KAAK,MAAM,CAAC,KAAK;QAChC,WAAW,KAAK,MAAM,CAAC,QAAQ,CAChC,CAAA;AACH,CAAC;AAED,4FAA4F;AAC5F,MAAM,UAAU,wBAAwB,CAAC,MAAe;IACtD,MAAM,GAAG,GAAG,QAAQ,CAAC,MAAM,CAAC,CAAA;IAC5B,IAAI,CAAC,GAAG;QAAE,OAAO,IAAI,CAAA;IACrB,IAAI,GAAG,CAAC,UAAU,KAAK,SAAS,IAAI,GAAG,CAAC,UAAU,KAAK,IAAI;QAAE,OAAO,cAAc,CAAA;IAClF,MAAM,UAAU,GAAG,QAAQ,CAAC,GAAG,CAAC,UAAU,CAAC,CAAA;IAC3C,IAAI,CAAC,UAAU,IAAI,CAAC,oBAAoB,CAAC,UAAU,CAAC;QAAE,OAAO,IAAI,CAAA;IACjE,MAAM,SAAS,GAAG,QAAQ,CAAC,GAAG,CAAC,MAAM,EAAE,gBAAgB,CAAC,CAAA;IACxD,IAAI,CAAC,SAAS;QAAE,OAAO,IAAI,CAAA;IAC3B,IACE,MAAM,CAAC,SAAS,CAAC,cAAc,CAAC,IAAI,CAAC,UAAU,EAAE,gBAAgB,CAAC;QAClE,QAAQ,CAAC,UAAU,CAAC,cAAc,EAAE,gBAAgB,CAAC,KAAK,SAAS;QACnE,OAAO,IAAI,CAAA;IACb,IACE,MAAM,CAAC,SAAS,CAAC,cAAc,CAAC,IAAI,CAAC,UAAU,EAAE,UAAU,CAAC;QAC5D,mCAAmC,CAAC,UAAU,CAAC;QAC/C,OAAO,IAAI,CAAA;IACb,OAAO,UAAU,CAAC,OAAO,CAAC,CAAC,CAAC,SAAS,CAAC,CAAC,CAAC,MAAM,CAAA;AAChD,CAAC;AAED;;;;GAIG;AACH,MAAM,UAAU,2BAA2B,CAAC,KAAc;IACxD,IAAI,KAAK,KAAK,SAAS,IAAI,KAAK,KAAK,IAAI;QAAE,OAAO,cAAc,CAAA;IAChE,MAAM,UAAU,GAAG,QAAQ,CAAC,KAAK,CAAC,CAAA;IAClC,IAAI,CAAC,UAAU,IAAI,OAAO,UAAU,CAAC,OAAO,KAAK,SAAS;QAAE,OAAO,IAAI,CAAA;IACvE,MAAM,qBAAqB,GACzB,MAAM,CAAC,SAAS,CAAC,cAAc,CAAC,IAAI,CAAC,UAAU,EAAE,UAAU,CAAC;QAC5D,MAAM,CAAC,SAAS,CAAC,cAAc,CAAC,IAAI,CAAC,UAAU,EAAE,OAAO,CAAC,CAAA;IAC3D,IAAI,qBAAqB,IAAI,oBAAoB,CAAC,UAAU,CAAC,KAAK,IAAI;QAAE,OAAO,IAAI,CAAA;IACnF,OAAO,UAAU,CAAC,OAAO,CAAC,CAAC,CAAC,SAAS,CAAC,CAAC,CAAC,MAAM,CAAA;AAChD,CAAC;AAED,SAAS,oBAAoB,CAAC,KAAqC;IACjE,MAAM,UAAU,GAAG,KAAK,CAAC,UAAU,IAAI,IAAI,CAAA;IAC3C,IAAI,KAAK,CAAC,eAAe,KAAK,sBAAsB,EAAE,CAAC;QACrD,OAAO,UAAU,KAAK,IAAI,IAAI,KAAK,CAAC,WAAW,KAAK,MAAM,CAAA;IAC5D,CAAC;IACD,IAAI,KAAK,CAAC,eAAe,KAAK,uBAAuB,EAAE,CAAC;QACtD,OAAO,UAAU,KAAK,KAAK,IAAI,KAAK,CAAC,WAAW,KAAK,MAAM,IAAI,KAAK,CAAC,UAAU,KAAK,SAAS,CAAA;IAC/F,CAAC;IACD,OAAO,UAAU,KAAK,IAAI,IAAI,KAAK,CAAC,WAAW,KAAK,MAAM,IAAI,KAAK,CAAC,UAAU,KAAK,SAAS,CAAA;AAC9F,CAAC;AAED;;;;;;;GAOG;AACH,MAAM,UAAU,6BAA6B,CAC3C,MAAe;IAEf,MAAM,GAAG,GAAG,QAAQ,CAAC,MAAM,CAAC,CAAA;IAC5B,MAAM,KAAK,GAAG,QAAQ,CAAC,GAAG,EAAE,KAAK,CAAC,CAAA;IAClC,IAAI,CAAC,GAAG,IAAI,CAAC,KAAK,IAAI,KAAK,CAAC,OAAO,KAAK,4BAA4B;QAAE,OAAO,IAAI,CAAA;IAEjF,MAAM,UAAU,GAAG,QAAQ,CAAC,GAAG,CAAC,MAAM,EAAE,gBAAgB,CAAC,CAAA;IACzD,MAAM,eAAe,GAAG,QAAQ,CAAC,KAAK,CAAC,eAAe,EAAE,gBAAgB,CAAC,CAAA;IACzE,MAAM,OAAO,GAAG,QAAQ,CAAC,KAAK,CAAC,OAAO,EAAE,gBAAgB,CAAC,CAAA;IACzD,IAAI,CAAC,UAAU,IAAI,eAAe,KAAK,UAAU,IAAI,CAAC,OAAO;QAAE,OAAO,IAAI,CAAA;IAE1E,MAAM,aAAa,GAAG,QAAQ,CAAC,GAAG,CAAC,UAAU,CAAC,CAAA;IAC9C,MAAM,eAAe,GAAG,sBAAsB,CAAC,UAAU,EAAE,aAAa,EAAE,gBAAgB,CAAC,CAAA;IAC3F,IAAI,sBAAsB,CAAC,eAAe,EAAE,OAAO,CAAC,KAAK,OAAO;QAAE,OAAO,IAAI,CAAA;IAE7E,MAAM,UAAU,GAAG,QAAQ,CAAC,KAAK,CAAC,UAAU,EAAE,iBAAiB,CAAC,CAAA;IAChE,IAAI,KAAK,CAAC,UAAU,KAAK,SAAS,IAAI,UAAU,KAAK,IAAI;QAAE,OAAO,IAAI,CAAA;IACtE,MAAM,eAAe,GAAG,QAAQ,CAAC,KAAK,CAAC,eAAe,EAAE,gBAAgB,CAAC,CAAA;IACzE,MAAM,WAAW,GAAG,QAAQ,CAAC,KAAK,CAAC,WAAW,EAAE,2BAA2B,CAAC,CAAA;IAC5E,MAAM,UAAU,GAAG,QAAQ,CAAC,KAAK,CAAC,UAAU,EAAE,mBAAmB,CAAC,CAAA;IAClE,IAAI,CAAC,eAAe,IAAI,CAAC,WAAW,IAAI,CAAC,UAAU;QAAE,OAAO,IAAI,CAAA;IAChE,IAAI,wBAAwB,CAAC,GAAG,CAAC,KAAK,UAAU;QAAE,OAAO,IAAI,CAAA;IAE7D,MAAM,UAAU,GAAmC;QACjD,OAAO,EAAE,4BAA4B;QACrC,OAAO;QACP,eAAe;QACf,GAAG,CAAC,UAAU,CAAC,CAAC,CAAC,EAAE,UAAU,EAAE,CAAC,CAAC,CAAC,EAAE,CAAC;QACrC,eAAe;QACf,WAAW;QACX,UAAU;KACX,CAAA;IACD,MAAM,QAAQ,GAAG,QAAQ,CAAC,KAAK,CAAC,QAAQ,EAAE,qBAAqB,CAAC,CAAA;IAChE,IAAI,QAAQ;QAAE,UAAU,CAAC,QAAQ,GAAG,QAAQ,CAAA;IAE5C,OAAO,oBAAoB,CAAC,UAAU,CAAC,CAAC,CAAC,CAAC,UAAU,CAAC,CAAC,CAAC,IAAI,CAAA;AAC7D,CAAC"}
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Defensive reader for the versioned FUL-358 hypothesis result.
|
|
3
|
+
*
|
|
4
|
+
* The MCP package cannot import the app or agent resolver, so this is its local mirror of the
|
|
5
|
+
* same contract: recognise only the semantics version this build understands, preserve every
|
|
6
|
+
* legacy number verbatim, and never let either a `final.verdict` or a raw robustness downgrade
|
|
7
|
+
* strengthen the hypothesis's original status.
|
|
8
|
+
*/
|
|
9
|
+
import { VERDICT_STATUSES, } from '../types/report.js';
|
|
10
|
+
import { readRowRobustnessOutcome, normalizeFinalHypothesisState, resolveMonotoneVerdict, } from '../hypothesis-semantics.js';
|
|
11
|
+
function asRecord(value) {
|
|
12
|
+
return typeof value === 'object' && value !== null && !Array.isArray(value)
|
|
13
|
+
? value
|
|
14
|
+
: null;
|
|
15
|
+
}
|
|
16
|
+
function asMember(value, members) {
|
|
17
|
+
return typeof value === 'string' && members.includes(value)
|
|
18
|
+
? value
|
|
19
|
+
: null;
|
|
20
|
+
}
|
|
21
|
+
export { resolveMonotoneVerdict } from '../hypothesis-semantics.js';
|
|
22
|
+
/** Read one raw `hypothesisResults[]` row without assuming its Zod parse succeeded. */
|
|
23
|
+
export function readFinalHypothesisState(rawHypothesis) {
|
|
24
|
+
const row = asRecord(rawHypothesis) ?? {};
|
|
25
|
+
const originalVerdict = asMember(row.status, VERDICT_STATUSES);
|
|
26
|
+
const legacyConfidence = typeof row.confidence === 'number' && Number.isFinite(row.confidence)
|
|
27
|
+
? row.confidence
|
|
28
|
+
: null;
|
|
29
|
+
const robustness = asRecord(row.robustness);
|
|
30
|
+
const weakestRecordedVerdict = resolveMonotoneVerdict(originalVerdict, robustness?.downgradedStatus);
|
|
31
|
+
const final = normalizeFinalHypothesisState(row);
|
|
32
|
+
if (final) {
|
|
33
|
+
return {
|
|
34
|
+
verdict: final.verdict,
|
|
35
|
+
originalVerdict: final.originalVerdict,
|
|
36
|
+
confidence: final.confidence ?? null,
|
|
37
|
+
confidenceBasis: final.confidenceBasis,
|
|
38
|
+
sufficiency: final.sufficiency,
|
|
39
|
+
robustness: final.robustness,
|
|
40
|
+
legacyConfidence,
|
|
41
|
+
fromContract: true,
|
|
42
|
+
};
|
|
43
|
+
}
|
|
44
|
+
return {
|
|
45
|
+
verdict: weakestRecordedVerdict,
|
|
46
|
+
originalVerdict,
|
|
47
|
+
confidence: null,
|
|
48
|
+
confidenceBasis: null,
|
|
49
|
+
sufficiency: null,
|
|
50
|
+
robustness: readRowRobustnessOutcome(row),
|
|
51
|
+
legacyConfidence,
|
|
52
|
+
fromContract: false,
|
|
53
|
+
};
|
|
54
|
+
}
|
|
55
|
+
//# sourceMappingURL=hypothesis-final.js.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"hypothesis-final.js","sourceRoot":"","sources":["../../src/tools/hypothesis-final.ts"],"names":[],"mappings":"AAAA;;;;;;;GAOG;AAEH,OAAO,EACL,gBAAgB,GAMjB,MAAM,oBAAoB,CAAA;AAC3B,OAAO,EACL,wBAAwB,EACxB,6BAA6B,EAC7B,sBAAsB,GACvB,MAAM,4BAA4B,CAAA;AAanC,SAAS,QAAQ,CAAC,KAAc;IAC9B,OAAO,OAAO,KAAK,KAAK,QAAQ,IAAI,KAAK,KAAK,IAAI,IAAI,CAAC,KAAK,CAAC,OAAO,CAAC,KAAK,CAAC;QACzE,CAAC,CAAE,KAAiC;QACpC,CAAC,CAAC,IAAI,CAAA;AACV,CAAC;AAED,SAAS,QAAQ,CAAmB,KAAc,EAAE,OAAqB;IACvE,OAAO,OAAO,KAAK,KAAK,QAAQ,IAAK,OAA6B,CAAC,QAAQ,CAAC,KAAK,CAAC;QAChF,CAAC,CAAE,KAAW;QACd,CAAC,CAAC,IAAI,CAAA;AACV,CAAC;AAED,OAAO,EAAE,sBAAsB,EAAE,MAAM,4BAA4B,CAAA;AAEnE,uFAAuF;AACvF,MAAM,UAAU,wBAAwB,CAAC,aAAsB;IAC7D,MAAM,GAAG,GAAG,QAAQ,CAAC,aAAa,CAAC,IAAI,EAAE,CAAA;IACzC,MAAM,eAAe,GAAG,QAAQ,CAAC,GAAG,CAAC,MAAM,EAAE,gBAAgB,CAAC,CAAA;IAC9D,MAAM,gBAAgB,GACpB,OAAO,GAAG,CAAC,UAAU,KAAK,QAAQ,IAAI,MAAM,CAAC,QAAQ,CAAC,GAAG,CAAC,UAAU,CAAC;QACnE,CAAC,CAAC,GAAG,CAAC,UAAU;QAChB,CAAC,CAAC,IAAI,CAAA;IACV,MAAM,UAAU,GAAG,QAAQ,CAAC,GAAG,CAAC,UAAU,CAAC,CAAA;IAC3C,MAAM,sBAAsB,GAAG,sBAAsB,CACnD,eAAe,EACf,UAAU,EAAE,gBAAgB,CAC7B,CAAA;IAED,MAAM,KAAK,GAAG,6BAA6B,CAAC,GAAG,CAAC,CAAA;IAChD,IAAI,KAAK,EAAE,CAAC;QACV,OAAO;YACL,OAAO,EAAE,KAAK,CAAC,OAAO;YACtB,eAAe,EAAE,KAAK,CAAC,eAAe;YACtC,UAAU,EAAE,KAAK,CAAC,UAAU,IAAI,IAAI;YACpC,eAAe,EAAE,KAAK,CAAC,eAAe;YACtC,WAAW,EAAE,KAAK,CAAC,WAAW;YAC9B,UAAU,EAAE,KAAK,CAAC,UAAU;YAC5B,gBAAgB;YAChB,YAAY,EAAE,IAAI;SACnB,CAAA;IACH,CAAC;IAED,OAAO;QACL,OAAO,EAAE,sBAAsB;QAC/B,eAAe;QACf,UAAU,EAAE,IAAI;QAChB,eAAe,EAAE,IAAI;QACrB,WAAW,EAAE,IAAI;QACjB,UAAU,EAAE,wBAAwB,CAAC,GAAG,CAAC;QACzC,gBAAgB;QAChB,YAAY,EAAE,KAAK;KACpB,CAAA;AACH,CAAC"}
|