forrt-research-mcp 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- forrt_research_mcp-0.1.0/.gitignore +7 -0
- forrt_research_mcp-0.1.0/LICENSE +21 -0
- forrt_research_mcp-0.1.0/PKG-INFO +479 -0
- forrt_research_mcp-0.1.0/README.md +453 -0
- forrt_research_mcp-0.1.0/pyproject.toml +59 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/__init__.py +18 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/api.py +123 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/chain.py +439 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/chain_draft.py +347 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/constellation.py +346 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/data/fields.snapshot.json +698 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/data/registry.json +53 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/drafts.py +478 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/grounding.py +255 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/quotes.py +309 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/server.py +415 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/template_spec.py +243 -0
- forrt_research_mcp-0.1.0/src/forrt_research_mcp/templates.py +250 -0
- forrt_research_mcp-0.1.0/tests/conftest.py +52 -0
- forrt_research_mcp-0.1.0/tests/fixtures/constellation-mhw.json +1 -0
- forrt_research_mcp-0.1.0/tests/fixtures/constellation-sado.json +1 -0
- forrt_research_mcp-0.1.0/tests/fixtures/question-pcc-aida-entry.json +93 -0
- forrt_research_mcp-0.1.0/tests/fixtures/question-pcc-full-chain.json +370 -0
- forrt_research_mcp-0.1.0/tests/fixtures/question-pcc.json +56 -0
- forrt_research_mcp-0.1.0/tests/fixtures/question-pico-aida-entry.json +98 -0
- forrt_research_mcp-0.1.0/tests/fixtures/question-pico-full-chain.json +384 -0
- forrt_research_mcp-0.1.0/tests/fixtures/question-pico.json +61 -0
- forrt_research_mcp-0.1.0/tests/fixtures/wikidata-marine-heatwave.json +110 -0
- forrt_research_mcp-0.1.0/tests/test_chain.py +434 -0
- forrt_research_mcp-0.1.0/tests/test_chain_draft.py +278 -0
- forrt_research_mcp-0.1.0/tests/test_constellation.py +159 -0
- forrt_research_mcp-0.1.0/tests/test_constellation_shapes.py +128 -0
- forrt_research_mcp-0.1.0/tests/test_drafts.py +320 -0
- forrt_research_mcp-0.1.0/tests/test_grounding.py +239 -0
- forrt_research_mcp-0.1.0/tests/test_question_anchors.py +163 -0
- forrt_research_mcp-0.1.0/tests/test_quotes.py +244 -0
- forrt_research_mcp-0.1.0/tests/test_server_tools.py +137 -0
- forrt_research_mcp-0.1.0/tests/test_templates.py +197 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Anne Fouilloux / Science Live
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,479 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: forrt-research-mcp
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: MCP server for producing verifiable FORRT nanopublication chains — grounded quote verification and compact Science Live constellation projections, for replications and new research alike.
|
|
5
|
+
Project-URL: Homepage, https://github.com/ScienceLiveHub/forrt-research-mcp
|
|
6
|
+
Project-URL: Repository, https://github.com/ScienceLiveHub/forrt-research-mcp
|
|
7
|
+
Project-URL: Replication template, https://github.com/ScienceLiveHub/forrt-replication-template
|
|
8
|
+
Project-URL: Science Live, https://sciencelive4all.org
|
|
9
|
+
Author: Anne Fouilloux
|
|
10
|
+
License: MIT
|
|
11
|
+
License-File: LICENSE
|
|
12
|
+
Keywords: FORRT,mcp,nanopublication,open-science,replication,reproducibility,science-live
|
|
13
|
+
Classifier: Development Status :: 3 - Alpha
|
|
14
|
+
Classifier: Intended Audience :: Science/Research
|
|
15
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Topic :: Scientific/Engineering
|
|
18
|
+
Classifier: Topic :: Scientific/Engineering :: Information Analysis
|
|
19
|
+
Requires-Python: >=3.10
|
|
20
|
+
Requires-Dist: mcp<2,>=1.2.0
|
|
21
|
+
Requires-Dist: pypdf>=4
|
|
22
|
+
Requires-Dist: rdflib>=7
|
|
23
|
+
Provides-Extra: dev
|
|
24
|
+
Requires-Dist: pytest>=8; extra == 'dev'
|
|
25
|
+
Description-Content-Type: text/markdown
|
|
26
|
+
|
|
27
|
+
# forrt-research-mcp
|
|
28
|
+
|
|
29
|
+
An MCP server for **producing** verifiable FORRT nanopublication chains — the
|
|
30
|
+
tools a researcher needs while doing the work, not while looking for it. It
|
|
31
|
+
serves all three shapes of study on the same rails: reproduction, replication,
|
|
32
|
+
and research that starts from scratch.
|
|
33
|
+
|
|
34
|
+
It is the third of three servers that compose:
|
|
35
|
+
|
|
36
|
+
| Server | Question it answers |
|
|
37
|
+
|---|---|
|
|
38
|
+
| [OpenAIRE MCP](https://github.com/ScienceLiveHub/replication-radar/blob/main/docs/openaire-mcp.md) | What is in the literature? |
|
|
39
|
+
| [`replication-radar`](https://github.com/ScienceLiveHub/replication-radar) | What is worth replicating, and has it been done? |
|
|
40
|
+
| **`forrt-research-mcp`** | **How do I produce a chain that is correct and verifiable?** |
|
|
41
|
+
|
|
42
|
+
Built for the workflow in
|
|
43
|
+
[`forrt-replication-template`](https://github.com/ScienceLiveHub/forrt-replication-template),
|
|
44
|
+
but it needs nothing from that repo — any agent can call it.
|
|
45
|
+
|
|
46
|
+
## Scope: reproduction, replication, and research from scratch
|
|
47
|
+
|
|
48
|
+
All three work here, and they use the same chain:
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
anchor → AIDA → Claim → Study → Outcome → CiTO
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
**Step 01 is one of three anchors, and they are not variants of one form.**
|
|
55
|
+
Which anchor fits is independent of whether the study is a reproduction, a
|
|
56
|
+
replication or original research:
|
|
57
|
+
|
|
58
|
+
| Anchor | Required fields | Notes |
|
|
59
|
+
|---|---|---|
|
|
60
|
+
| `01_quote` Quote-with-comment | `paper`, `quotation`, `comment` | no type, no label. `quotation` is capped at 500 characters and `comment` at 800 — in the template's own regex, not in prose |
|
|
61
|
+
| `01_pico` PICO question | `label`, `description`, `type`, + P/I/C/O descriptions | the only anchor with a question-type vocabulary (5 terms) |
|
|
62
|
+
| `01_pcc` PCC question | `label`, `description`, + P/C/C descriptions | **no `type` field at all** |
|
|
63
|
+
|
|
64
|
+
The three share **no field ids whatsoever**, so call `template_fields` on the
|
|
65
|
+
anchor you are actually using rather than assuming they match. The
|
|
66
|
+
`pico_question_type` vocabulary is named for PICO deliberately: a PCC question
|
|
67
|
+
has no type to look up.
|
|
68
|
+
|
|
69
|
+
`verify_quote` applies to the Quote anchor only — and that is the one anchor
|
|
70
|
+
with a dedicated tool, because a quotation is the one field whose correctness
|
|
71
|
+
can be *proved* rather than reviewed.
|
|
72
|
+
|
|
73
|
+
`constellation` and `prior_work` surface a question anchor with its `framework`,
|
|
74
|
+
`label`, full question text and its framework's `components` — PICO's
|
|
75
|
+
population / intervention / comparator / outcome, PCC's population / concept /
|
|
76
|
+
context. Question nodes keep all of that in a `question` object and leave their
|
|
77
|
+
top-level `label` empty, so anything reading them like a quote returns nothing.
|
|
78
|
+
Tested against two real published question nanopubs.
|
|
79
|
+
|
|
80
|
+
For a study starting from scratch, the Claim is **your own hypothesis**, derived
|
|
81
|
+
from your own question or from the work you are building on. The FORRT Claim
|
|
82
|
+
template's `source` is optional, so it needs no external paper — and the
|
|
83
|
+
Claim-before-Study order then reads as
|
|
84
|
+
**pre-registration**, not as a mismatch. `verify_chain` takes a `mode`
|
|
85
|
+
(`auto` / `replication` / `reproduction` / `new_research`) that changes only what
|
|
86
|
+
is *required*: from-scratch research has no existing work to cite, so no CiTO
|
|
87
|
+
step and no cited DOI are expected.
|
|
88
|
+
|
|
89
|
+
**The one place the templates still assume an original** is the `study_type`
|
|
90
|
+
vocabulary on `04_study`, whose three terms are all replication-flavoured
|
|
91
|
+
(Replication Study, Reproduction Study, or both). There is no term for *an
|
|
92
|
+
original study testing its own claim*. Three field prompts read oddly too —
|
|
93
|
+
`scope` and `methodology` say "is reproduced/replicated", and the Outcome's
|
|
94
|
+
`conclusion` says "about the original claim" — but they are only wording;
|
|
95
|
+
`validationStatus` (validated / partially supported / contradicted /
|
|
96
|
+
inconclusive / not tested) describes testing your own hypothesis perfectly well.
|
|
97
|
+
|
|
98
|
+
So closing the gap is plausibly **one added vocabulary term and three reworded
|
|
99
|
+
prompts**, not a new template family. When that lands, this server picks it up
|
|
100
|
+
with no code change: `template_fields` and `vocabulary` fetch live, and
|
|
101
|
+
`driftedFromSnapshot` flags the supersession so the vendored copy gets re-cut.
|
|
102
|
+
|
|
103
|
+
> **A note on using "Reproduction Study" for from-scratch work.** It is a
|
|
104
|
+
> reasonable workaround while the vocabulary lacks a better term, but the
|
|
105
|
+
> template means the replication-science sense — *"direct reproduction: same
|
|
106
|
+
> methodology, same tools"*, i.e. re-running someone else's analysis — not the
|
|
107
|
+
> RSE sense of "my work is reproducible". Downstream consumers read it the first
|
|
108
|
+
> way: `replication-radar`'s verdict overlay and `verified_claims` will present
|
|
109
|
+
> the study as verification of an existing claim.
|
|
110
|
+
|
|
111
|
+
## Why a server and not a prompt
|
|
112
|
+
|
|
113
|
+
Two jobs here look like reasoning but are not, and doing them in an agent loop
|
|
114
|
+
makes them unreproducible and expensive:
|
|
115
|
+
|
|
116
|
+
**1. Reading a published chain.** `/np/constellation` walks the FORRT citation
|
|
117
|
+
graph bidirectionally and returns everything it reaches. For a single
|
|
118
|
+
marine-heatwave chain that is **330 KB, 98 nodes and 1012 edges — of which 64
|
|
119
|
+
nodes are AIDA statements belonging to entirely different studies**. Handing
|
|
120
|
+
that to an agent burns its context and invites it to reason over another paper's
|
|
121
|
+
claims. `constellation` returns the same chain in ~16 KB (4.7 %).
|
|
122
|
+
|
|
123
|
+
**2. Checking a quotation is real.** A FORRT Quote nanopub must be verbatim. That
|
|
124
|
+
is a string search, not a judgement — so it should be a tool that cannot be
|
|
125
|
+
talked out of its answer, and that anyone can re-run to get the same result.
|
|
126
|
+
|
|
127
|
+
## What it does not do
|
|
128
|
+
|
|
129
|
+
It does not search papers (that is the OpenAIRE MCP), rank replication targets
|
|
130
|
+
(that is `replication-radar`), or write nanopub content. It also does not
|
|
131
|
+
*extract* claims: choosing which sentence carries a paper's headline claim is a
|
|
132
|
+
judgement, and wrapping a judgement in a tool would not make it reproducible — a
|
|
133
|
+
model behind a tool boundary is exactly as non-deterministic as one in the agent
|
|
134
|
+
loop. The division this server is built around:
|
|
135
|
+
|
|
136
|
+
| The model proposes | This server disposes |
|
|
137
|
+
|---|---|
|
|
138
|
+
| Which sentence is the claim | Whether that sentence is in the PDF, where, under what SHA-256 |
|
|
139
|
+
| How to phrase the AIDA | What fields the template actually has, and their caps |
|
|
140
|
+
| Which claim type or CiTO relation applies | Which values the form will actually accept |
|
|
141
|
+
| Which Wikidata topic is meant | Whether that QID exists and is of the right type |
|
|
142
|
+
| Which paper to cite | Whether that DOI resolves, and to what |
|
|
143
|
+
| Whether to extend or dispute prior work | What prior chains already claimed |
|
|
144
|
+
|
|
145
|
+
## Install
|
|
146
|
+
|
|
147
|
+
```bash
|
|
148
|
+
pipx install forrt-research-mcp
|
|
149
|
+
claude mcp add forrt-research -s user -- forrt-research-mcp
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
User-scoped (`-s user`) so it is available in every session and folder, and does
|
|
153
|
+
not clash with a per-repo config. For other agents, the stdio command is
|
|
154
|
+
`forrt-research-mcp`.
|
|
155
|
+
|
|
156
|
+
```bash
|
|
157
|
+
# optional
|
|
158
|
+
export SCIENCELIVE_API_BASE="https://api-dev.sciencelive4all.org" # default
|
|
159
|
+
export SCIENCELIVE_API_KEY="sl_…" # /np/constellation is a public read
|
|
160
|
+
```
|
|
161
|
+
|
|
162
|
+
> **The default base is `api-dev` deliberately, for now.** `/np/constellation`
|
|
163
|
+
> is newer than the current production deployment, so as of 2026-09-02
|
|
164
|
+
> production answers HTTP 500 on known-good URIs while `api-dev` serves them
|
|
165
|
+
> with 200. This is a deployment lag, not a fault. Once the release reaches
|
|
166
|
+
> production, switch `DEFAULT_API_BASE` in `api.py` — `SCIENCELIVE_API_BASE`
|
|
167
|
+
> already overrides it in the meantime.
|
|
168
|
+
|
|
169
|
+
## Tools
|
|
170
|
+
|
|
171
|
+
### `verify_quote(pdf_path, quotation)`
|
|
172
|
+
|
|
173
|
+
Proves a candidate quotation is in a source PDF before it is published as
|
|
174
|
+
verbatim. Returns a graded verdict with page, character offsets and the file's
|
|
175
|
+
SHA-256:
|
|
176
|
+
|
|
177
|
+
| Verdict | Meaning |
|
|
178
|
+
|---|---|
|
|
179
|
+
| `exact` | byte-identical to the extracted page text |
|
|
180
|
+
| `normalized` | matched after whitespace / ligature / typographic punctuation / line-break-hyphen repair |
|
|
181
|
+
| `extraction_tolerant` | additionally ignored hyphens and punctuation spacing |
|
|
182
|
+
| `whitespace_insensitive` | additionally deleted every space — for text layers that break words apart |
|
|
183
|
+
| `not_found` | **not in this PDF — do not publish it** |
|
|
184
|
+
|
|
185
|
+
Every tier canonicalises *formatting only* — never words, digits or order. A
|
|
186
|
+
quotation with one digit changed scores ~0.91 similarity against the source and
|
|
187
|
+
is still `not_found`; on a miss, `closest.text_in_pdf` shows what the paper
|
|
188
|
+
actually says there.
|
|
189
|
+
|
|
190
|
+
A tier below `exact` is normal. PDF extraction inserts line breaks, loses
|
|
191
|
+
hyphens, and sometimes breaks words apart entirely:
|
|
192
|
+
|
|
193
|
+
- the quotation published in the real marine-heatwave chain matches only at
|
|
194
|
+
`extraction_tolerant`, because `pypdf` reads the paper's `35-year` as `35year`
|
|
195
|
+
and `(p <` as `( p <`;
|
|
196
|
+
- Hundhausen et al. 2024 renders its own abstract as `CP en semble`, `largest c
|
|
197
|
+
hanges`, `prec ipitation` — spaces *inside* words, which no punctuation rule
|
|
198
|
+
can undo. That needs `whitespace_insensitive`, which deletes every space and
|
|
199
|
+
therefore does not verify word boundaries; it declines to run below 40
|
|
200
|
+
characters, where two short texts could collide.
|
|
201
|
+
|
|
202
|
+
**"Character-for-character" is not literally achievable against extracted PDF
|
|
203
|
+
text**, which is why the result is graded rather than boolean.
|
|
204
|
+
|
|
205
|
+
Words split across a line break (`convection-\npermitting`) are rejoined at the
|
|
206
|
+
`normalized` tier. A *suspended* hyphen (`15- and 30-minute`) is deliberately
|
|
207
|
+
NOT rejoined — the rule requires a line break, because joining those would
|
|
208
|
+
corrupt the text.
|
|
209
|
+
|
|
210
|
+
### `constellation(uri, depth=5, max_nodes=80)`
|
|
211
|
+
|
|
212
|
+
The FORRT chain(s) reachable from a published nanopub URI, projected to
|
|
213
|
+
agent-size: chains with their steps in chain order, the apex CiTO, any Research
|
|
214
|
+
Synthesis, and the Quote/Question anchors attributable to this paper.
|
|
215
|
+
|
|
216
|
+
Read `citedPaper`, not the API's top-level `paperDoi` — see *Upstream
|
|
217
|
+
quirks* below.
|
|
218
|
+
|
|
219
|
+
`stepsPresent` legitimately omits steps: a CiTO at the apex of a constellation
|
|
220
|
+
is hoisted out of its chain, and Quote/AIDA anchors are often not enumerated.
|
|
221
|
+
Missing does not mean unpublished.
|
|
222
|
+
|
|
223
|
+
### `prior_work(uri)`
|
|
224
|
+
|
|
225
|
+
What has already been claimed about a paper — the starting point for new work,
|
|
226
|
+
replication or otherwise. Per completed chain: claim type, `scope`, `method`,
|
|
227
|
+
`deviations`, verdict, confidence, repository, and `limitations`.
|
|
228
|
+
|
|
229
|
+
`limitations` is the field to read most carefully. It is where previous authors
|
|
230
|
+
stated, in signed and immutable words, what their study did *not* cover — which
|
|
231
|
+
is often exactly where the next study begins. In the marine-heatwave chain it
|
|
232
|
+
records that the paper's better-known 54 % headline describes a different
|
|
233
|
+
analysis over a different period and was not tested.
|
|
234
|
+
|
|
235
|
+
### `constellation_raw(uri, …)`
|
|
236
|
+
|
|
237
|
+
The unprojected response, for debugging the graph or an upstream data problem.
|
|
238
|
+
Large; prefer `constellation`.
|
|
239
|
+
|
|
240
|
+
### `template_fields(step, live=True)` · `vocabulary(name, live=True)` · `list_schemas()`
|
|
241
|
+
|
|
242
|
+
A nanopub template **is** the schema for its chain step, so these turn "never
|
|
243
|
+
invent a field name" and "never invent a claim type" from rules an agent has to
|
|
244
|
+
remember into a lookup that can only return real values.
|
|
245
|
+
|
|
246
|
+
`template_fields` returns the real field ids, prompts, required/repeatable
|
|
247
|
+
flags, and the `regex` / `prefix` / `datatype` constraints — which is where the
|
|
248
|
+
Quote template's character cap actually lives, rather than in a hand-written
|
|
249
|
+
doc that drifts. `step` accepts `05_outcome`, `05`, or `outcome`.
|
|
250
|
+
|
|
251
|
+
`vocabulary` returns the allowed values of a controlled vocabulary, taken from
|
|
252
|
+
the real restricted-choice field:
|
|
253
|
+
|
|
254
|
+
| Name | From | Terms |
|
|
255
|
+
|---|---|---|
|
|
256
|
+
| `claim_type` | `03_claim.forrtType` | 7 |
|
|
257
|
+
| `study_type` | `04_study.type` | 3 — the Reproduction vs Replication distinction |
|
|
258
|
+
| `validation_status` | `05_outcome.validationStatus` | 5 |
|
|
259
|
+
| `confidence_level` | `05_outcome.confidenceLevel` | 5 |
|
|
260
|
+
| `cito_relation` | `06_citation.cites` | 43, via a separate value-list nanopub |
|
|
261
|
+
| `question_type` | `01_pico.type` | 5 |
|
|
262
|
+
|
|
263
|
+
Both fetch live by default and fall back to a **bundled snapshot** when the
|
|
264
|
+
network is unavailable — always reporting which was used, and setting
|
|
265
|
+
`driftedFromSnapshot` when the upstream template has been superseded. A drifted
|
|
266
|
+
result means the live values win and this package needs re-vendoring; it is a
|
|
267
|
+
loud, reviewable event rather than slow silent divergence.
|
|
268
|
+
|
|
269
|
+
`cito_relation` is the one vocabulary that cannot be resolved offline: it lives
|
|
270
|
+
in a separate value-list nanopub. Offline it returns `source:
|
|
271
|
+
"unavailable-offline"` with a warning rather than an empty list, because an
|
|
272
|
+
empty list reads as "no valid values."
|
|
273
|
+
|
|
274
|
+
### `verify_chain(published_path, repo_url="")`
|
|
275
|
+
|
|
276
|
+
The final check before announcing a chain. Point it at a `nanopubs/PUBLISHED.md`
|
|
277
|
+
ledger; `green` is true only when nothing failed. Read-only — it never edits,
|
|
278
|
+
retracts or supersedes.
|
|
279
|
+
|
|
280
|
+
| Check | What it means |
|
|
281
|
+
|---|---|
|
|
282
|
+
| ledger | every required step (01–06) has a URI |
|
|
283
|
+
| reachable | every URI is really published |
|
|
284
|
+
| repository | the Outcome's archived version DOI resolves |
|
|
285
|
+
| cited-doi | every DOI the chain cites resolves |
|
|
286
|
+
| verdict-relation | **the CiTO relation agrees with the Outcome's verdict** |
|
|
287
|
+
|
|
288
|
+
That last one is the failure most worth catching: `Validated` implies
|
|
289
|
+
`confirms`, `PartiallySupported` implies `qualifies`, `Contradicted` implies
|
|
290
|
+
`disputes`. A mismatch means the Outcome and the Citation disagree about what
|
|
291
|
+
the replication actually found — a chain that cites `confirms` for a paper it
|
|
292
|
+
contradicted is worse than no chain at all.
|
|
293
|
+
|
|
294
|
+
**A constellation alone cannot verify a chain**, which is why this does not try.
|
|
295
|
+
On both real chains the walk stops short of the upstream anchors, so anything it
|
|
296
|
+
does not enumerate is checked by fetching its TriG from the bare
|
|
297
|
+
`w3id.org/np/` resolver — the `/sciencelive/np/` form serves an HTML viewer that
|
|
298
|
+
answers 200 and would pass a status-only check while serving no nanopub. A row
|
|
299
|
+
reading *"not enumerated by the walk but its TriG resolves"* is a pass, not a
|
|
300
|
+
warning.
|
|
301
|
+
|
|
302
|
+
The Outcome's repository is expected to be a **Zenodo version DOI**, not a
|
|
303
|
+
GitHub URL: a version DOI pins the archived state the outcome was computed
|
|
304
|
+
from, where `github.com/ORG/REPO` is a moving target. Both are accepted; a DOI
|
|
305
|
+
is checked by resolving it, a URL by comparing it to `repo_url`.
|
|
306
|
+
|
|
307
|
+
Both published chains verify green, matching the hand-run verification recorded
|
|
308
|
+
in marine-heatwave's ledger.
|
|
309
|
+
|
|
310
|
+
### `validate_chain_draft(path)`
|
|
311
|
+
|
|
312
|
+
Checks `nanopubs/chain-draft.json` — **the artifact that actually gets
|
|
313
|
+
published** — before it is handed to the Science Live chain wizard. Run it at the
|
|
314
|
+
end of Phase 5b, after `pixi run build-chain-draft` and before pushing the file
|
|
315
|
+
and opening the wizard URL.
|
|
316
|
+
|
|
317
|
+
The markdown drafts are the authoring format; `build_chain_draft.py` turns them
|
|
318
|
+
plus `CITATION.cff` and the templates into this file, which the wizard pre-fills
|
|
319
|
+
each step from and a human reviews and signs. `validate_draft` checks the input;
|
|
320
|
+
this checks the artifact.
|
|
321
|
+
|
|
322
|
+
What it catches that reading the file cannot:
|
|
323
|
+
|
|
324
|
+
- a **superseded `template_uri`** — invisible in the JSON, but it makes the
|
|
325
|
+
wizard pre-fill the old form;
|
|
326
|
+
- a `prefill` key that is neither a template field nor a known platform
|
|
327
|
+
form-field, which the wizard silently drops;
|
|
328
|
+
- a complex field in the wrong shape — `06_citation.st02` must be
|
|
329
|
+
`[{cites, cited}]` with at least one entry, and `04_study.disciplineSelection`
|
|
330
|
+
is a **single object, not an array** (the one asymmetry in the contract);
|
|
331
|
+
- a required field neither prefilled nor carried forward;
|
|
332
|
+
- a value over the template's own cap, an invalid vocabulary term, a malformed
|
|
333
|
+
date, an unresolved `{{TOKEN}}`, or a DOI that does not resolve;
|
|
334
|
+
- a `carry_forward` edge running backwards through the chain.
|
|
335
|
+
|
|
336
|
+
Fields the wizard fills itself are exempt rather than reported missing:
|
|
337
|
+
`02_aida` has no `project`, `03_claim` no `aida`, `04_study` no `claim`.
|
|
338
|
+
|
|
339
|
+
Both of this project's real chain drafts validate clean, and 13 deliberate
|
|
340
|
+
mutations of one are each caught.
|
|
341
|
+
|
|
342
|
+
### `validate_draft(path)` · `validate_drafts(directory)`
|
|
343
|
+
|
|
344
|
+
The pre-flight checklist in `docs/forrt-form-fields.md`, actually executed. One
|
|
345
|
+
call checks a drafted nanopub end to end: field ids against the live template,
|
|
346
|
+
choice values against its enumeration, length caps against its regex, DOIs
|
|
347
|
+
against the registry, QIDs against Wikidata. `publishable` is true only when
|
|
348
|
+
nothing came back as an error.
|
|
349
|
+
|
|
350
|
+
This is the composition the other tools exist for — individually they answer
|
|
351
|
+
"is this value real?", together they answer "is this draft publishable, and if
|
|
352
|
+
not, exactly where is it wrong?"
|
|
353
|
+
|
|
354
|
+
**Three placeholder conventions, and only one is a problem.** Getting this wrong
|
|
355
|
+
made the checker useless on real drafts, so each is handled explicitly:
|
|
356
|
+
|
|
357
|
+
| In a draft | Meaning | Severity |
|
|
358
|
+
|---|---|---|
|
|
359
|
+
| `«URI of step 05 …»` | back-reference; the chain wizard fills it from the published step | info |
|
|
360
|
+
| `{{ZENODO_VERSION_DOI}}` | release-time token; the release workflow substitutes it | warning |
|
|
361
|
+
| anything else standing in for a value | genuinely unfilled | **error** |
|
|
362
|
+
|
|
363
|
+
A draft whose required fields are nearly all empty is reported once as an
|
|
364
|
+
unfilled skeleton — "it has not been drafted yet" — rather than as a wall of
|
|
365
|
+
per-field errors.
|
|
366
|
+
|
|
367
|
+
**What it cannot see.** Values a draft puts in prose or a markdown table rather
|
|
368
|
+
than behind a `<!-- field: … -->` marker are reported as `coverage` warnings,
|
|
369
|
+
never as missing. The CiTO step's citation list is the known case: `cites` and
|
|
370
|
+
`cited` live in a table in both repos checked.
|
|
371
|
+
|
|
372
|
+
Validated against two real published chains (marine-heatwave and Sado estuary):
|
|
373
|
+
all six published steps pass in both, with the only remaining flags being an
|
|
374
|
+
unpublished step 07 and the release tokens above.
|
|
375
|
+
|
|
376
|
+
### `resolve_doi(doi)`
|
|
377
|
+
|
|
378
|
+
Does this DOI resolve, and to what? `resolves: false` means it is not
|
|
379
|
+
registered — **a well-formed DOI is not a real one**, and a fabricated one is
|
|
380
|
+
indistinguishable from a genuine one until something asks the registry. Returns
|
|
381
|
+
the registered title, authors, year and container so you can confirm it is the
|
|
382
|
+
paper you meant, not merely a paper that exists. Accepts bare, URL, and `doi:`
|
|
383
|
+
forms. A 5xx is reported as transient rather than as a bad DOI.
|
|
384
|
+
|
|
385
|
+
### `wikidata_lookup(query, expected_type="", limit=5)`
|
|
386
|
+
|
|
387
|
+
Real Wikidata candidates for a term, with their actual P31/P279 types. Pass
|
|
388
|
+
`expected_type` as a QID and each candidate is marked `typeMatches`.
|
|
389
|
+
|
|
390
|
+
Checked against the ten Wikidata terms in this project's published chains — the
|
|
391
|
+
QIDs `build_chain_draft.py` resolved and signed — it returns the published QID
|
|
392
|
+
at rank 1 for all ten (marine heatwave, sea surface temperature, climate change,
|
|
393
|
+
time series analysis, chlorophyll a, remote sensing, estuary, Sentinel-2, water
|
|
394
|
+
quality, atmospheric correction).
|
|
395
|
+
|
|
396
|
+
It deliberately **does not choose** — picking the right sense of an ambiguous
|
|
397
|
+
label is a judgement. Searching `Bombus` with `Q16521` (taxon) returns the
|
|
398
|
+
insect genus as a match and *the album of the same name* as not, which is
|
|
399
|
+
exactly the failure worth catching before a QID gets signed into a nanopub.
|
|
400
|
+
Candidates are annotated, never filtered: a near miss is often the informative
|
|
401
|
+
result. Zero candidates means leave the field empty — never fall back to a QID
|
|
402
|
+
from memory.
|
|
403
|
+
|
|
404
|
+
## Upstream quirks this handles
|
|
405
|
+
|
|
406
|
+
Found by probing live data rather than reading the API contract, and each one is
|
|
407
|
+
pinned by a test against a real recorded payload (two constellations, from
|
|
408
|
+
unrelated studies):
|
|
409
|
+
|
|
410
|
+
1. **Cloudflare rejects urllib's default User-Agent.** The API runs on
|
|
411
|
+
Cloudflare Workers, whose bot protection answers `Python-urllib/3.x` with
|
|
412
|
+
HTTP 403 — same URI, same key, 200 under any normal agent. Every request this
|
|
413
|
+
client makes sets an explicit `User-Agent`. Do the same in your own client;
|
|
414
|
+
no hermetic test can catch it.
|
|
415
|
+
2. **`paperDoi` can name the wrong paper.** It is chosen by frequency across the
|
|
416
|
+
whole walk, so a neighbouring study's Quote nanopubs can outvote the chain's
|
|
417
|
+
own citation. On the marine-heatwave chain it reports the LifeWatch ERIC
|
|
418
|
+
paper (4 unrelated quotes) instead of Oliver et al. 2018, which the chain's
|
|
419
|
+
CiTO actually cites — and it names **the same wrong paper** from the
|
|
420
|
+
unrelated Sado-estuary entry point, so this is systemic, not one bad record.
|
|
421
|
+
`citedPaper` derives the paper from the CiTO and sets
|
|
422
|
+
`disagreesWithReported` when the two differ.
|
|
423
|
+
3. **The walk stops short, and missing does not mean unpublished.** Two real
|
|
424
|
+
chains enumerated `[Claim, Study, Outcome]` and `[Study, Outcome, CiTO]`,
|
|
425
|
+
while the Quote, AIDA and Claim nanopubs listed in those repos'
|
|
426
|
+
`PUBLISHED.md` were published and merely unreachable. `stepsNotEnumerated`
|
|
427
|
+
reports the gap. **A constellation alone cannot verify a complete chain** —
|
|
428
|
+
fetch those URIs' TriG directly.
|
|
429
|
+
4. **`quote.quotedText` comes back empty** on nanopubs using the current quote
|
|
430
|
+
template. The text survives in `label` and in two untagged excerpts — the
|
|
431
|
+
quotation and the annotator's comment, in no guaranteed order. We match them
|
|
432
|
+
against the label stem rather than assuming a position, and report the
|
|
433
|
+
recovery in `text_source`.
|
|
434
|
+
5. **The apex CiTO is hoisted** out of `chains[].steps[]` to top level.
|
|
435
|
+
6. **The `/sciencelive/np/` URI form serves an HTML viewer** that answers 200, so
|
|
436
|
+
a status-only reachability check passes even when no nanopub is served. We
|
|
437
|
+
normalise to the bare `w3id.org/np/` resolver and assert the body is RDF.
|
|
438
|
+
|
|
439
|
+
## Checking against every real artefact
|
|
440
|
+
|
|
441
|
+
The test suite is hermetic: it pins behaviour against recorded data, which
|
|
442
|
+
prevents regressions but **cannot discover a wrong premise**. Every serious bug
|
|
443
|
+
in this server was a wrong premise, and each surfaced only when the tools met
|
|
444
|
+
real data they had not seen. So "have we tested comprehensively?" is not a
|
|
445
|
+
judgement call here — it is a computation with a visible denominator:
|
|
446
|
+
|
|
447
|
+
```bash
|
|
448
|
+
python scripts/check_corpus.py /path/to/your/study/repos
|
|
449
|
+
```
|
|
450
|
+
|
|
451
|
+
It discovers every repository with a `nanopubs/` directory, runs the applicable
|
|
452
|
+
tools against each artefact, and prints what passed out of what was found.
|
|
453
|
+
Findings are *not* automatically failures — an empty required field is a true
|
|
454
|
+
statement about a draft. What it exits non-zero for is a **crash or a parse
|
|
455
|
+
failure**, because those mean a tool could not read a real artefact at all.
|
|
456
|
+
|
|
457
|
+
The denominator is only as good as what is on disk. A chain that is not under
|
|
458
|
+
that root is not covered, and a green run does not say otherwise.
|
|
459
|
+
|
|
460
|
+
## Development
|
|
461
|
+
|
|
462
|
+
```bash
|
|
463
|
+
pip install -e '.[dev]'
|
|
464
|
+
pytest
|
|
465
|
+
```
|
|
466
|
+
|
|
467
|
+
All 287 tests are hermetic and need no network, no API key, and no live
|
|
468
|
+
service: the constellation fixtures are real recorded `/np/constellation`
|
|
469
|
+
responses, quote tests build minimal PDFs in-process that reproduce the
|
|
470
|
+
extraction artifacts deliberately, and the template/DOI/Wikidata tests stub HTTP
|
|
471
|
+
with recorded payloads. The suite stays green through an upstream outage.
|
|
472
|
+
|
|
473
|
+
The bundled template snapshot under `src/forrt_research_mcp/data/` is vendored
|
|
474
|
+
from `ScienceLiveHub/forrt-replication-template`, as is `template_spec.py`.
|
|
475
|
+
Re-vendor when `template_fields` reports `driftedFromSnapshot`.
|
|
476
|
+
|
|
477
|
+
## License
|
|
478
|
+
|
|
479
|
+
MIT — see [LICENSE](LICENSE).
|