@citeark/agent 0.3.20

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (347) hide show
  1. package/LICENSE +202 -0
  2. package/README.md +128 -0
  3. package/data/dataset-source-registry.v1.json +300 -0
  4. package/dist/arkgraph/boot.js +6 -0
  5. package/dist/arkgraph/index.html +1 -0
  6. package/dist/arkgraph/viewer.css +1 -0
  7. package/dist/arkgraph/viewer.en.css +1 -0
  8. package/dist/arkgraph/viewer.en.js +49 -0
  9. package/dist/arkgraph/viewer.en.js.LEGAL.txt +56 -0
  10. package/dist/arkgraph/viewer.js +49 -0
  11. package/dist/arkgraph/viewer.js.LEGAL.txt +56 -0
  12. package/docker/claude-code/Dockerfile +97 -0
  13. package/docker/claude-code/codex-pro-relay.mjs +466 -0
  14. package/docker/claude-code/runtime-contract-check.mjs +79 -0
  15. package/docs/arkgraph-reading.md +79 -0
  16. package/docs/configuration.md +100 -0
  17. package/docs/integration.md +92 -0
  18. package/docs/maturity-plan.md +27 -0
  19. package/docs/npm-release.md +44 -0
  20. package/docs/paper-reading.md +40 -0
  21. package/docs/research-plan-granularity.md +27 -0
  22. package/docs/terminal.md +49 -0
  23. package/examples/toy-evaluation/compile-task.json +27 -0
  24. package/examples/toy-evaluation/paper.md +5 -0
  25. package/examples/toy-evaluation/repository/README.md +9 -0
  26. package/examples/toy-evaluation/repository/checkpoint.json +4 -0
  27. package/examples/toy-evaluation/repository/evaluate.py +17 -0
  28. package/examples/toy-evaluation/task.json +81 -0
  29. package/package.json +59 -0
  30. package/prompts/compile-research.md +58 -0
  31. package/prompts/execute-contract.md +72 -0
  32. package/prompts/execute-workspace-simple.md +51 -0
  33. package/prompts/execute-workspace.md +34 -0
  34. package/prompts/prepare-reproduction.md +82 -0
  35. package/prompts/repair-research.md +45 -0
  36. package/protocol/CAP.md +129 -0
  37. package/protocol/LICENSE +12 -0
  38. package/protocol/MAPPINGS.md +72 -0
  39. package/protocol/README.md +38 -0
  40. package/protocol/conformance-v2.0-alpha.1.json +36 -0
  41. package/protocol/examples/arkgraph/checkpoint-evaluation.json +309 -0
  42. package/protocol/examples/arkgraph/fixtures.mjs +49 -0
  43. package/protocol/examples/arkgraph/paper-free.json +291 -0
  44. package/protocol/examples/arkgraph/partial-failure.json +344 -0
  45. package/protocol/examples/arkgraph/training-evaluation.json +443 -0
  46. package/protocol/profiles/agent-trace.md +16 -0
  47. package/protocol/profiles/computational-run.md +16 -0
  48. package/protocol/profiles/core.md +15 -0
  49. package/protocol/profiles/public-bundle.md +18 -0
  50. package/protocol/profiles/reproduction.md +29 -0
  51. package/protocol/profiles/research-compilation.md +44 -0
  52. package/protocol/profiles/research-plan.md +39 -0
  53. package/protocol/profiles/restricted-evidence.md +15 -0
  54. package/runtime/bootstrap-autodl-runtime.sh +314 -0
  55. package/runtime/create-runtime-venv.sh +41 -0
  56. package/runtime/install-local-cpu-runtime.sh +23 -0
  57. package/runtime/install-scientific-runtime.sh +153 -0
  58. package/runtime/mineru/parse.py +62 -0
  59. package/runtime/mineru/requirements.txt +4 -0
  60. package/runtime/requirements-baseline.txt +38 -0
  61. package/schemas/cap/v2/activity.schema.json +47 -0
  62. package/schemas/cap/v2/agent.schema.json +32 -0
  63. package/schemas/cap/v2/assertion.schema.json +110 -0
  64. package/schemas/cap/v2/descriptor.schema.json +243 -0
  65. package/schemas/cap/v2/entity.schema.json +64 -0
  66. package/schemas/cap/v2/manifest.schema.json +67 -0
  67. package/schemas/cap/v2/relation.schema.json +82 -0
  68. package/schemas/compute-catalog.schema.json +63 -0
  69. package/schemas/compute-decision.schema.json +27 -0
  70. package/schemas/execution-contract.schema.json +1024 -0
  71. package/schemas/research-card.schema.json +30 -0
  72. package/schemas/research-inventory-draft.schema.json +366 -0
  73. package/schemas/research.schema.json +1044 -0
  74. package/schemas/result.schema.json +173 -0
  75. package/schemas/verification-policy.schema.json +47 -0
  76. package/schemas/verified-conclusion.schema.json +58 -0
  77. package/schemas/workspace-summary.schema.json +24 -0
  78. package/scripts/build-arkgraph-view.mjs +12 -0
  79. package/scripts/check-execution-feasibility.mjs +24 -0
  80. package/scripts/check-syntax.mjs +15 -0
  81. package/scripts/deterministic-asset-preparation.py +438 -0
  82. package/scripts/package-cap.mjs +23 -0
  83. package/scripts/package-local-agent.mjs +23 -0
  84. package/scripts/preview-arkgraph.mjs +25 -0
  85. package/scripts/replay-research-compiler-candidate.mjs +134 -0
  86. package/scripts/review-compiler-sources.mjs +44 -0
  87. package/scripts/run-asset-preparation.sh +17 -0
  88. package/scripts/run-research-plan.mjs +98 -0
  89. package/scripts/validate-asset-preparation.py +290 -0
  90. package/scripts/verify-local-runtime.mjs +57 -0
  91. package/scripts/verify-npm-package.mjs +57 -0
  92. package/src/adapters/paper2agent.mjs +107 -0
  93. package/src/assets/cache.mjs +159 -0
  94. package/src/assets/compute.mjs +98 -0
  95. package/src/assets/executor.mjs +145 -0
  96. package/src/assets/lifecycle.mjs +213 -0
  97. package/src/assets/manifest.mjs +242 -0
  98. package/src/assets/opportunistic-preparation.mjs +81 -0
  99. package/src/assets/plan.mjs +411 -0
  100. package/src/assets/prompts.mjs +29 -0
  101. package/src/assets/public-asset-probe.mjs +525 -0
  102. package/src/assets/qualification.mjs +119 -0
  103. package/src/assets/readiness.mjs +130 -0
  104. package/src/assets/reproduction-admission.mjs +355 -0
  105. package/src/assets/requirements.mjs +152 -0
  106. package/src/assets/source-grounding.mjs +341 -0
  107. package/src/assets/source-policy.mjs +118 -0
  108. package/src/autodl/client.mjs +260 -0
  109. package/src/autodl/ssh.mjs +380 -0
  110. package/src/autodl/tools.mjs +129 -0
  111. package/src/cap/redaction.mjs +38 -0
  112. package/src/cap/v2/archive.mjs +152 -0
  113. package/src/cap/v2/attestation.mjs +204 -0
  114. package/src/cap/v2/canonical-json.mjs +114 -0
  115. package/src/cap/v2/compilation-artifact.mjs +240 -0
  116. package/src/cap/v2/core.mjs +282 -0
  117. package/src/cap/v2/measurement-assessment-records.mjs +23 -0
  118. package/src/cap/v2/pipeline-artifact.mjs +922 -0
  119. package/src/cap/v2/read.mjs +41 -0
  120. package/src/cap/v2/reassessment-artifact.mjs +383 -0
  121. package/src/cap/v2/research-artifact.mjs +231 -0
  122. package/src/cap/v2/research-map-records.mjs +46 -0
  123. package/src/cap/v2/research-object-records.mjs +163 -0
  124. package/src/cap/v2/research-records.mjs +187 -0
  125. package/src/cap/v2/verify.mjs +642 -0
  126. package/src/cli.mjs +1146 -0
  127. package/src/compute/autodl-pro-compiler.mjs +347 -0
  128. package/src/compute/autodl-pro-executor.mjs +459 -0
  129. package/src/compute/autodl-pro-job.mjs +843 -0
  130. package/src/compute/autodl-pro-network.mjs +295 -0
  131. package/src/compute/autodl-pro-remote.mjs +810 -0
  132. package/src/compute/autodl-pro-staging.mjs +117 -0
  133. package/src/compute/campaign.mjs +110 -0
  134. package/src/compute/catalog.mjs +123 -0
  135. package/src/compute/checkpoint-protocol.mjs +154 -0
  136. package/src/compute/codex-account-lock.mjs +111 -0
  137. package/src/compute/codex-account-session.mjs +107 -0
  138. package/src/compute/compiler-profile.mjs +38 -0
  139. package/src/compute/compiler-router.mjs +23 -0
  140. package/src/compute/coordinator-recovery.mjs +210 -0
  141. package/src/compute/executor-router.mjs +29 -0
  142. package/src/compute/gcp-batch-compiler.mjs +685 -0
  143. package/src/compute/gcp-batch-executor.mjs +1215 -0
  144. package/src/compute/gcp-batch-failure.mjs +92 -0
  145. package/src/compute/gcp-batch-job.mjs +527 -0
  146. package/src/compute/gcp-batch-lifecycle.mjs +81 -0
  147. package/src/compute/gcp-checkpoint-worker.mjs +1633 -0
  148. package/src/compute/local-codex-compiler.mjs +52 -0
  149. package/src/compute/measurement-hardware.mjs +128 -0
  150. package/src/compute/remote-attempt.mjs +226 -0
  151. package/src/compute/requirements.mjs +124 -0
  152. package/src/compute/research-phases.mjs +48 -0
  153. package/src/compute/scheduler.mjs +452 -0
  154. package/src/compute/shared-workloads.mjs +26 -0
  155. package/src/compute/stage-archive.mjs +79 -0
  156. package/src/contracts/campaign-contract.mjs +52 -0
  157. package/src/contracts/execution-contract.mjs +819 -0
  158. package/src/contracts/execution-mode.mjs +19 -0
  159. package/src/contracts/execution-timeouts.mjs +45 -0
  160. package/src/contracts/execution-workload.mjs +68 -0
  161. package/src/contracts/preflight-schema.mjs +25 -0
  162. package/src/contracts/public-contract.mjs +63 -0
  163. package/src/contracts/subject-tags.mjs +31 -0
  164. package/src/dashboard/data.mjs +898 -0
  165. package/src/dashboard/server.mjs +79 -0
  166. package/src/dashboard/static/dashboard.css +366 -0
  167. package/src/dashboard/static/dashboard.js +560 -0
  168. package/src/dashboard/static/index.html +85 -0
  169. package/src/deployment/community-policy.mjs +9 -0
  170. package/src/deployment/environment.mjs +112 -0
  171. package/src/deployment/guided.mjs +98 -0
  172. package/src/deployment/handoff.mjs +102 -0
  173. package/src/deployment/local-contract.mjs +31 -0
  174. package/src/deployment/local.mjs +100 -0
  175. package/src/deployment/prepare.mjs +46 -0
  176. package/src/deployment/recipe.mjs +108 -0
  177. package/src/deployment/supplement.mjs +51 -0
  178. package/src/deployment/terminal.mjs +43 -0
  179. package/src/diagnosis/renderer.mjs +75 -0
  180. package/src/diagnosis/target-failure.mjs +46 -0
  181. package/src/evidence/parser-registry.mjs +54 -0
  182. package/src/evidence/parsers/fasttext-classification.mjs +82 -0
  183. package/src/evidence/parsers/json-scalar.mjs +96 -0
  184. package/src/evidence/parsers/simcse-senteval.mjs +104 -0
  185. package/src/evidence/parsers/starspace-classification.mjs +78 -0
  186. package/src/evidence/registry.mjs +147 -0
  187. package/src/execution/runner-audit.mjs +473 -0
  188. package/src/gcp/auth.mjs +106 -0
  189. package/src/gcp/batch-client.mjs +120 -0
  190. package/src/gcp/resource-discovery.mjs +177 -0
  191. package/src/gcp/rest.mjs +82 -0
  192. package/src/gcp/secret-manager.mjs +34 -0
  193. package/src/gcp/signed-url.mjs +133 -0
  194. package/src/gcp/storage.mjs +220 -0
  195. package/src/graph/command.mjs +41 -0
  196. package/src/graph/execution.mjs +97 -0
  197. package/src/graph/model.mjs +37 -0
  198. package/src/graph/presentation.mjs +110 -0
  199. package/src/graph/query.mjs +159 -0
  200. package/src/graph/research-relations.mjs +69 -0
  201. package/src/graph/source-page.mjs +12 -0
  202. package/src/graph/source-preview.mjs +34 -0
  203. package/src/graph/validate.mjs +76 -0
  204. package/src/job.mjs +496 -0
  205. package/src/network/autodl-routing-proxy.mjs +462 -0
  206. package/src/network/egress-proxy.mjs +158 -0
  207. package/src/observability/event-contract.mjs +230 -0
  208. package/src/observability/pipeline-monitor.mjs +166 -0
  209. package/src/pipeline/orchestrator.mjs +1281 -0
  210. package/src/pipeline/recovery-error.mjs +11 -0
  211. package/src/pipeline/replay.mjs +304 -0
  212. package/src/pipeline/shared-execution.mjs +115 -0
  213. package/src/pipeline/stage-checkpoint.mjs +86 -0
  214. package/src/pipeline/stage-recovery.mjs +101 -0
  215. package/src/pipeline/targets.mjs +110 -0
  216. package/src/process.mjs +143 -0
  217. package/src/protocol.mjs +312 -0
  218. package/src/provider/codex-account.mjs +44 -0
  219. package/src/provider/codex-completion.mjs +49 -0
  220. package/src/provider/completion.mjs +292 -0
  221. package/src/provider/model-client.mjs +44 -0
  222. package/src/provider/model-route.mjs +29 -0
  223. package/src/provider/openrouter-readiness.mjs +189 -0
  224. package/src/provider/reader-bridge.mjs +35 -0
  225. package/src/provider/relay.mjs +263 -0
  226. package/src/provider/runtime-auth.mjs +40 -0
  227. package/src/public/cap.d.mts +90 -0
  228. package/src/public/cap.mjs +12 -0
  229. package/src/public/contracts.d.mts +2 -0
  230. package/src/public/host.mjs +171 -0
  231. package/src/public/operations.d.mts +11 -0
  232. package/src/public/presentation.d.mts +4 -0
  233. package/src/records/views.mjs +26 -0
  234. package/src/remote/command.mjs +178 -0
  235. package/src/remote/ssh.mjs +59 -0
  236. package/src/repository-origin.mjs +81 -0
  237. package/src/reproduction/evidence-feedback.mjs +96 -0
  238. package/src/reproduction/incomplete-initialization.mjs +25 -0
  239. package/src/reproduction/lifecycle.mjs +253 -0
  240. package/src/reproduction/plan.mjs +132 -0
  241. package/src/reproduction/prompts.mjs +70 -0
  242. package/src/reproduction/runner.mjs +188 -0
  243. package/src/reproduction/summary.mjs +130 -0
  244. package/src/reproduction/workspace-mode.mjs +7 -0
  245. package/src/research/automatic-admission.mjs +156 -0
  246. package/src/research/compiler-coverage.mjs +85 -0
  247. package/src/research/compiler-failure.mjs +24 -0
  248. package/src/research/compiler-normalization-guards.mjs +112 -0
  249. package/src/research/compiler-repair.mjs +3 -0
  250. package/src/research/compiler.mjs +853 -0
  251. package/src/research/continuation-selection.mjs +26 -0
  252. package/src/research/execution-graph-context.mjs +43 -0
  253. package/src/research/experiment-importance.mjs +15 -0
  254. package/src/research/inventory-handoff.mjs +104 -0
  255. package/src/research/inventory-revisions.mjs +32 -0
  256. package/src/research/mineru-local.mjs +73 -0
  257. package/src/research/paper-command.mjs +19 -0
  258. package/src/research/paper-markdown.mjs +180 -0
  259. package/src/research/paper-source-map.mjs +69 -0
  260. package/src/research/planning-policy.mjs +88 -0
  261. package/src/research/reference-materials.mjs +11 -0
  262. package/src/research/reproduction-scope.mjs +30 -0
  263. package/src/research/research-map.mjs +94 -0
  264. package/src/research/research-objects.mjs +88 -0
  265. package/src/research/source-discovery.mjs +646 -0
  266. package/src/research/source-observations.mjs +75 -0
  267. package/src/research/source-review-cli-mcp.mjs +26 -0
  268. package/src/research/source-review-input.mjs +209 -0
  269. package/src/research/source-review-local-codex.mjs +36 -0
  270. package/src/research/source-review-model.mjs +70 -0
  271. package/src/research/source-review.mjs +173 -0
  272. package/src/research/structure.mjs +3163 -0
  273. package/src/research-card/renderer.mjs +277 -0
  274. package/src/research-card/verified-conclusion.mjs +143 -0
  275. package/src/results/output-registry.mjs +183 -0
  276. package/src/runtime/claude-code.mjs +52 -0
  277. package/src/runtime/codex-capacity-retry.mjs +87 -0
  278. package/src/runtime/codex.mjs +64 -0
  279. package/src/runtime/config.mjs +157 -0
  280. package/src/runtime/final-output.mjs +40 -0
  281. package/src/runtime/index.mjs +21 -0
  282. package/src/runtime/local-codex.mjs +74 -0
  283. package/src/runtime/opencode.mjs +95 -0
  284. package/src/runtime/prompt.mjs +13 -0
  285. package/src/sandbox/docker.mjs +363 -0
  286. package/src/settings/command.mjs +297 -0
  287. package/src/settings/store.mjs +119 -0
  288. package/src/telemetry/pricing.mjs +68 -0
  289. package/src/telemetry/usage.mjs +265 -0
  290. package/src/terminal/events.mjs +97 -0
  291. package/src/terminal/input.mjs +40 -0
  292. package/src/terminal/plain.mjs +40 -0
  293. package/src/terminal/remote-stream.mjs +22 -0
  294. package/src/terminal/screen.mjs +214 -0
  295. package/src/terminal/transcript.mjs +69 -0
  296. package/src/util.mjs +107 -0
  297. package/src/verification/ai-assessor.mjs +534 -0
  298. package/src/verification/claim-evaluator.mjs +242 -0
  299. package/src/verification/evidence-context.mjs +165 -0
  300. package/src/verification/evidence-reader.mjs +95 -0
  301. package/src/verification/integrity.mjs +570 -0
  302. package/src/verification/tolerance.mjs +32 -0
  303. package/src/workloads/cpu-research-preparation.mjs +56 -0
  304. package/src/workloads/definition.mjs +74 -0
  305. package/src/workloads/phase-aware-reproduction.mjs +46 -0
  306. package/src/workloads/reproduction.mjs +85 -0
  307. package/src/workspace/command.mjs +242 -0
  308. package/src/workspace/control.mjs +49 -0
  309. package/src/workspace/entry.mjs +28 -0
  310. package/src/workspace/input.mjs +93 -0
  311. package/src/workspace/interactive.mjs +94 -0
  312. package/src/workspace/jobs.mjs +418 -0
  313. package/src/workspace/session.mjs +97 -0
  314. package/src/workspace/worker.mjs +137 -0
  315. package/ui/arkgraph/ambient-motion.mjs +10 -0
  316. package/ui/arkgraph/app.jsx +153 -0
  317. package/ui/arkgraph/boot.js +6 -0
  318. package/ui/arkgraph/camera-motion.mjs +20 -0
  319. package/ui/arkgraph/context-reveal.mjs +39 -0
  320. package/ui/arkgraph/details.css +3 -0
  321. package/ui/arkgraph/entry.jsx +28 -0
  322. package/ui/arkgraph/experiment-curves.mjs +17 -0
  323. package/ui/arkgraph/experiment-selection.mjs +15 -0
  324. package/ui/arkgraph/experiment-style.css +26 -0
  325. package/ui/arkgraph/experiment-ui.jsx +32 -0
  326. package/ui/arkgraph/frame.html +1 -0
  327. package/ui/arkgraph/graph-gestures.mjs +62 -0
  328. package/ui/arkgraph/label-layout.mjs +57 -0
  329. package/ui/arkgraph/locales/en.json +229 -0
  330. package/ui/arkgraph/locales/source-types.json +15 -0
  331. package/ui/arkgraph/localization-build.mjs +27 -0
  332. package/ui/arkgraph/material-build.mjs +23 -0
  333. package/ui/arkgraph/material-colors.mjs +39 -0
  334. package/ui/arkgraph/material-style.css +15 -0
  335. package/ui/arkgraph/open-graph.jsx +326 -0
  336. package/ui/arkgraph/outline.jsx +49 -0
  337. package/ui/arkgraph/package-lock.json +888 -0
  338. package/ui/arkgraph/package.json +17 -0
  339. package/ui/arkgraph/reading-layout.mjs +130 -0
  340. package/ui/arkgraph/reading-presentation.mjs +73 -0
  341. package/ui/arkgraph/record-detail.css +51 -0
  342. package/ui/arkgraph/record-details.jsx +29 -0
  343. package/ui/arkgraph/research-types.mjs +31 -0
  344. package/ui/arkgraph/selection-mark.jsx +6 -0
  345. package/ui/arkgraph/soft-spine.mjs +26 -0
  346. package/ui/arkgraph/steering-style.css +187 -0
  347. package/ui/arkgraph/style.css +272 -0
@@ -0,0 +1,4 @@
1
+ {
2
+ "correct": 819,
3
+ "total": 1000
4
+ }
@@ -0,0 +1,17 @@
1
+ import argparse
2
+ import json
3
+ from pathlib import Path
4
+
5
+
6
+ def main() -> None:
7
+ parser = argparse.ArgumentParser()
8
+ parser.add_argument("--checkpoint", required=True)
9
+ args = parser.parse_args()
10
+
11
+ checkpoint = json.loads(Path(args.checkpoint).read_text())
12
+ accuracy = checkpoint["correct"] / checkpoint["total"]
13
+ print(json.dumps({"accuracy": accuracy, "examples": checkpoint["total"]}))
14
+
15
+
16
+ if __name__ == "__main__":
17
+ main()
@@ -0,0 +1,81 @@
1
+ {
2
+ "schemaVersion": "1.0",
3
+ "kind": "citeark.execution-contract",
4
+ "name": "toy-checkpoint-evaluation",
5
+ "research": {
6
+ "workId": "toy-accuracy-work",
7
+ "claimId": "toy-accuracy-claim",
8
+ "claimVersionId": "sha256:1111111111111111111111111111111111111111111111111111111111111111",
9
+ "experimentId": "toy-checkpoint-evaluation",
10
+ "experimentVersionId": "sha256:2222222222222222222222222222222222222222222222222222222222222222",
11
+ "measurementIds": ["toy-accuracy"]
12
+ },
13
+ "paper": {
14
+ "title": "A Toy Accuracy Claim",
15
+ "url": "https://citeark.com/examples/toy-accuracy",
16
+ "path": "./paper.md"
17
+ },
18
+ "repository": {
19
+ "url": "https://citeark.com/examples/toy-evaluation.git",
20
+ "commit": "fixture",
21
+ "path": "./repository"
22
+ },
23
+ "claim": {
24
+ "text": "Evaluate the protocol-bound toy accuracy claim; the reported outcome is intentionally withheld until independent verification.",
25
+ "source": "toy-paper:Results, Table 1"
26
+ },
27
+ "measurements": [{
28
+ "measurementId": "toy-accuracy",
29
+ "metric": "accuracy",
30
+ "unit": "fraction",
31
+ "parser": {
32
+ "id": "json.scalar",
33
+ "version": "1",
34
+ "evidencePath": "logs/evaluation.log",
35
+ "config": { "pointer": "/accuracy", "metric": "accuracy", "unit": "fraction" }
36
+ }
37
+ }],
38
+ "protocol": {
39
+ "entrypoint": "evaluate.py",
40
+ "requiredCommandFragments": ["evaluate.py", "--checkpoint"],
41
+ "instructions": "Evaluate the released checkpoint. Do not retrain or alter the metric."
42
+ },
43
+ "reproductionLevel": "official-checkpoint",
44
+ "reconstructionFidelity": "faithful",
45
+ "computeRequirement": {
46
+ "schemaVersion": "0.1",
47
+ "source": "research-declared",
48
+ "classification": "cpu-only",
49
+ "accelerator": "none",
50
+ "cpuCores": 2,
51
+ "memoryGb": 4,
52
+ "gpuCount": 1,
53
+ "minVramGb": null,
54
+ "preferredGpuTypes": [],
55
+ "cpuFallbackAllowed": true,
56
+ "estimatedDurationMinutes": { "cpu": 5, "gpu": null },
57
+ "rationale": "The fixed toy checkpoint evaluation is CPU-only."
58
+ },
59
+ "verificationPolicyCommitment": "sha256:3333333333333333333333333333333333333333333333333333333333333333",
60
+ "agent": {
61
+ "runtime": "claude-code",
62
+ "model": null,
63
+ "subagentModel": null,
64
+ "autoCompactWindow": 786432,
65
+ "effort": "max",
66
+ "reasoningMode": "standard",
67
+ "maxTurns": 40,
68
+ "maxBudgetUsd": 3,
69
+ "validationRetries": 0
70
+ },
71
+ "environment": {
72
+ "image": "citeark-agent/runtime:0.2.0",
73
+ "timeoutMinutes": 15,
74
+ "cpus": 2,
75
+ "memoryGb": 4,
76
+ "shmGb": 1,
77
+ "gpu": "none",
78
+ "networkAccess": true
79
+ },
80
+ "contractDigest": "sha256:38dbd33676f1dbf2381d9548ef77ae5264410e3a714bf0f83c7529cfebd8d579"
81
+ }
package/package.json ADDED
@@ -0,0 +1,59 @@
1
+ {
2
+ "name": "@citeark/agent",
3
+ "version": "0.3.20",
4
+ "description": "Independent paper research, experiment execution and signed CAP artifacts",
5
+ "type": "module",
6
+ "bin": {
7
+ "citeark": "./src/cli.mjs",
8
+ "citeark-agent": "./src/cli.mjs"
9
+ },
10
+ "exports": {
11
+ "./host": "./src/public/host.mjs",
12
+ "./cap": "./src/public/cap.mjs",
13
+ "./package.json": "./package.json"
14
+ },
15
+ "scripts": {
16
+ "test": "node --test tests/*.test.mjs",
17
+ "check": "node scripts/check-syntax.mjs",
18
+ "package:local": "npm run build:arkgraph && node scripts/package-local-agent.mjs && node scripts/package-cap.mjs",
19
+ "replay:compiler": "node scripts/replay-research-compiler-candidate.mjs",
20
+ "test:package": "node scripts/verify-npm-package.mjs",
21
+ "build:arkgraph": "npm --prefix ui/arkgraph ci && node scripts/build-arkgraph-view.mjs",
22
+ "preview:arkgraph": "node scripts/preview-arkgraph.mjs",
23
+ "pretest:package": "npm run build:arkgraph"
24
+ },
25
+ "engines": {
26
+ "node": ">=22"
27
+ },
28
+ "license": "Apache-2.0",
29
+ "dependencies": {
30
+ "esbuild": "0.28.2",
31
+ "ssh2": "1.17.0"
32
+ },
33
+ "files": [
34
+ "src",
35
+ "prompts",
36
+ "schemas",
37
+ "scripts",
38
+ "docker/claude-code",
39
+ "runtime",
40
+ "data",
41
+ "protocol",
42
+ "docs",
43
+ "examples/toy-evaluation",
44
+ "ui/arkgraph",
45
+ "dist/arkgraph"
46
+ ],
47
+ "repository": {
48
+ "type": "git",
49
+ "url": "git+https://github.com/ganwumeng/CiteArk-Agent.git"
50
+ },
51
+ "homepage": "https://citeark.co/docs/local-reproduction",
52
+ "bugs": {
53
+ "url": "https://github.com/ganwumeng/CiteArk-Agent/issues"
54
+ },
55
+ "publishConfig": {
56
+ "access": "public",
57
+ "registry": "https://registry.npmjs.org/"
58
+ }
59
+ }
@@ -0,0 +1,58 @@
1
+ You are the CiteArk Research Compiler. Produce a faithful research inventory of the fixed paper for the reproduction Agent. Your task is to record what the paper reports, not to design an executable verification plan. Follow compile-task.json.scientificPlanningPolicy.compilationBoundary.
2
+
3
+ ## Write for the research reader
4
+
5
+ The inventory's prose appears directly in ArkGraph explanations. Apply this style to every user-facing summary, title, statement, condition, limitation, ambiguity, relationship note and reading label, in every output language:
6
+
7
+ - Lead with the scientific question, method, result or meaning of the object. Use concrete subjects such as the paper, method, experiment or evidence. Explain what the reader needs to understand the research.
8
+ - State necessary evidence boundaries directly and locally: the exact missing comparison, conflicting values, unresolved correspondence or restricted applicability, and its consequence when established. Preserve attribution, uncertainty, negative results and source contradictions. Distinguish information absent from the paper from information you have not located or checked. Do not strengthen a conclusion to make the prose sound cleaner.
9
+ - The safeguards elsewhere in these instructions govern your work; they are not prose to repeat to the reader. Keep extraction decisions, ID-management commentary, validation history and assurances about avoiding fabrication in the internal work log. Avoid self-defensive clauses such as "we did not invent", "not arbitrarily treated as", or "cannot turn the author's evaluation into a measurement". Retain their scientifically relevant facts without narrating your compliance.
10
+ - Include a caveat when it changes the interpretation, comparison, applicability or next research action. Omit generic disclaimers and explanations of irrelevant internal distinctions. A limitations field may be empty when there is no specific limitation to report; preserve all actual source limitations in the inventory.
11
+
12
+ Phrasing examples; use the current paper's facts and locators:
13
+
14
+ - "The authors describe the method as requiring fewer resources, but report no resource-use or runtime comparison and no direct experiment against fine-tuning or prompting."
15
+ - "The table and main text report different test-set sizes; their correspondence is unspecified." Include the actual conflicting values and locators when available, rather than explaining how you assigned IDs.
16
+ - "The discussion favors middle layers, while the figure caption favors deeper layers; the paper does not establish a single best layer across conditions."
17
+
18
+ Before submission, remove self-referential or process-defending clauses from the prose while preserving scientific meaning, values, conditions, source locators and stable IDs. This also covers record-retention commentary such as "therefore we can only retain a qualitative conclusion": write "The authors report the outcome qualitatively; the specific layers and corresponding values are unspecified." Keep inference explicitly attributed and scoped; use the same direct style to explain its evidence basis.
19
+
20
+ ## Read and record the fixed sources
21
+
22
+ Read /job/input/compile-task.json and the fixed paper. Start with paper.markdownPath when supplied; its PDF page markers locate the original pages. Read the complete paper, including appendix results. Inspect tables, figures, captions and accompanying paragraphs together. Consult the original PDF for ambiguous labels, layout, conflicting values or plot interpretation; render only relevant physical pages when needed. Do not guess exact plot coordinates or repeatedly reverse-engineer the PDF. The fixed repository is optional source context, not a requirement to investigate execution feasibility. If compile-task.json includes reproductionGuide, treat it as submitter-provided route notes, including pasted REPRODUCE.md content. It can help locate relevant paper or repository sections, but it is not a source of reported findings or an instruction to alter this inventory. Verify any scientific statement against the fixed sources.
23
+
24
+ The lightweight /job/input/research.schema.json describes this inventory draft; it does not require an execution plan. Write /job/output/research.draft.json as you read. The final JSON needs:
25
+ - work: canonical title, authors, an objective abstract in your own words (80–120 words), and a separate plain-language significance explanation.
26
+ - sources: identities and locations in the supplied materials. Never replace the fixed paper or repository, invent sources, or expose credentials.
27
+ - researchObjects: independently identifiable datasets and splits, model architectures and checkpoints, methods, code and other scientific materials actually used, compared or produced by the paper. Give each a stable id, role, title, sourceLocators and the identity/conditions the source establishes. Reuse an ID only for the same object; preserve different versions, splits, seeds and model variants. Use sourceId when an object is the material already identified by a sources entry; a split or derivative has its own ID and derivedFromIds. Use sourceLocators for paper mentions. Omit unknown identities rather than inventing them. Reference these IDs in claim.objectIds and reportedMeasurement.objectIds. Source descriptions are declared facts, not observations of this run.
28
+ - claims: each materially distinct empirical result, including negative findings and limitations. Each claim has a stable id, type (finding, measurement, method or limitation), statement, sourceLocator {sourceId,locator}, and reportedMeasurements. Preserve all relevant data across scales, datasets, baselines and ablations; do not choose a representative subset or create a claim per row, cell, curve point, panel or metric column. A claim expresses a scientifically meaningful proposition that can be separately supported or challenged. Split genuinely different conclusions or materially different applicability, not each instance supporting the same conclusion. Keep all measurements for a comparison together. Completeness means preserving source evidence and conditions, not maximizing claim count. Methods and theory also have their own researchObjects; they need not each become a testable claim.
29
+ - observation collections: represent each coherent source matrix, table or curve family as a researchObject with role observation, basis declared, observationKind (table, matrix, series or collection), stable id, title and sourceLocators. Put its numeric rows in the relevant scientific claim's reportedMeasurements, with observationId referencing this object. Preserve a stable measurement id plus the exact dimensions, conditions, units and locator for every row. Use descriptive observations/materials for plot-only data; never invent numerical coordinates. One claim can retain hundreds of measurements. Distinct claims may cite the same observation identity, but measurement IDs remain globally unique. Do not transcribe self-correlations or sample counts into standalone scientific claims; retain them as data or method conditions. Do not infer collection identity merely from sharing a page or figure.
30
+ - hypotheses: optional [{id, statement, sourceLocator, objectIds?, conditions?, limitations?}] for source-proposed mechanisms or explanations not empirically established. Keep these separate from claims; do not assign them execution dispositions or support verdicts.
31
+ - Each reportedMeasurement has id, metric, unit, value and sourceLocator. Include model/configuration, dataset/split, comparison objects and other necessary reported conditions in conditions or description. Preserve the paper's own numbers and units. Include original operands and weighting when the paper itself defines an aggregate. Do not invent an aggregate or precision.
32
+ - Use claim.conditions, claim.relations and claim.ambiguities as needed to retain context, comparison meaning and source contradictions. Label inference explicitly and separate it from reported facts. A short claim.verificationHint may suggest what a later investigator could inspect; it is optional and provisional.
33
+
34
+ For qualitative or plot-only findings without a reliable number, keep reportedMeasurements: [] and state the result, direction, objects, source locator and ambiguity. Do not manufacture a numeric target, mark the finding not_applicable, or call it impossible because the verification rule is still undecided. Preserve statements about instability, failures and exceptions as well as improvements. If the source contradicts itself, preserve what it prints and explain the inconsistency; do not silently choose a corrected value.
35
+
36
+ Do not provide experiments, execution commands, parser bindings, asset searches, scope exclusions, feasibility verdicts, cost estimates, probes or completed decision rules. Do not download datasets or checkpoints. Reproduction preparation reads this inventory and the original materials, investigates inputs, chooses comparisons and decision rules, and plans within the task's scope and resources. The execution Agent may revise its interpretation with source-grounded reasons. Independent assessment judges evidence. Those are later responsibilities.
37
+
38
+ The host supplies versions, digests, pending dispositions and structural defaults. An inventory item is recorded, not verified or declared executable. The initial inventory remains provisional. The reproduction Agent may reread and correct extraction errors with source-grounded revision records. An explicitly requested source audit can check omissions and fidelity; the default handoff does not require that extra model gate. An unresolved execution route is not an inventory defect.
39
+
40
+ Review the draft against every result section and figure/table, including negative results and appendix findings. Check scientific organization too: a figure is a source container, an observation stores data, a method describes a recipe, and an assertion expresses a proposition. A figure may support multiple genuinely different claims; neither one-claim-per-figure nor one-claim-per-cell is a rule. Do not impose a numerical claim cap or omit difficult evidence to make the graph small. Preserve discovered IDs across corrections. Work directly in the draft; do not create empty helper scaffolds. Reserve /job/output/research.json for the completed submission, parse the JSON and replace that final path atomically. Do not run experiments or produce an execution report.
41
+
42
+
43
+ ## Current research request and prior materials
44
+ Read task.researchGoal as the user's current question. Preserve the source claim inventory independently of this selection and of the available hardware. During reproduction preparation, design routes that address that question and disclose different experimental conditions. task.reproductionGuide and /job/input/reference-materials.json, when present, contain optional historical instructions and code. Inspect their relevance against the fixed paper and repository. They never fix this run's experiment selection, resources, measurements or success. Reuse suitable methods and code, retain uncertainty, and plan this run independently. Historical observations are not new execution evidence.
45
+
46
+ ## ArkGraph: complete science with a readable mainline
47
+
48
+ Produce `reading` and `researchRelations` alongside the source inventory. The schema describes their exact structure. Use typed references `object:<researchObject.id>`, `claim:<claim.id>`, `hypothesis:<hypothesis.id>`, `source:<source.id>`; references must resolve to actual inventory entries, never titles or invented display nodes.
49
+
50
+ Use the twelve reading categories where the source warrants them: question, concept, hypothesis/premise, claim, method, plan, activity, observation/result, argument/proof, assessment, resource, participant. Categories are not quotas. Existing material/method/observation roles remain. Additional researchObject roles are question, concept, premise, argument, research-plan, research-activity, source-assessment, person, organization, instrument. Each additional object needs a source locator and a substantive description. Claims retain scientific propositions; hypotheses retain proposed explanations; do not duplicate them as objects. A source-assessment is an explicitly attributed evaluation reported by the source, not this compiler's verdict. Record extraction contradictions on the affected object's limitations/claim ambiguities, not as invented author assessments. A research-activity describes something the paper reports actually doing, with declared basis; future work is a research-plan, never an actual activity. An argument may retain a proof or chain of reasoning without pretending it is an empirical measurement. Keep unsupported user-interface categories empty.
51
+
52
+ `reading: {schemaVersion:"1.0", overview, mainline:[typed references], branches:[{parent:typed reference, children:[typed references]}], labels:[{target:typed reference,title,subtype?}]}` is editorial organization only. Identify the paper's actual research progression from its motivating question through its central method/reasoning to the main findings and boundaries. Select a compact meaningful mainline, with no fixed count or mandatory stage template. Branches expose supporting concepts, methods, resources, findings and caveats. Place table/figure/observation collections and detailed derivations in contextual branches so they can appear on hover/selection; every object remains available in the complete outline. A child has one display parent, but its scientific relationships may have many parents. A display parent must ultimately connect to the mainline. Short labels must faithfully summarize complete object descriptions/statements, not replace them. Do not generate coordinates, animation parameters, colors or artificial reading-order scientific edges.
53
+
54
+ `researchRelations: [{id,from,to,predicate,basis,sourceLocators,note?}]` records source-grounded scientific relationships beyond the existing objectIds, observationId and derivedFromIds bindings. Do not repeat those bindings. Use only schema predicates and respect their endpoint primitives: about targets an Entity, evidenceFor connects an observation Entity to an Assertion, basedOn connects an Assertion to its basis, describes connects an Entity to the described object, used/generated connect an actual reported Activity to an Entity, associatedWith connects Activity to participant. Use declared for explicit source statements and inferred with an explanatory note for compiler interpretation. Visual proximity and reading order never establish a scientific relation. Complete source data, negative results, assumptions, conditions and limitations must survive this organization. No fixed node/edge count is a quality target.
55
+
56
+ Source researchRelations do not use hasStep or expects: those predicates are reserved for prospective executable procedures and are assembled from experiment steps later. A source method may describe a submethod or require an input method when those semantics are stated in the paper. Choose the actual supported relationship; do not relabel source methods as execution procedures. People and organizations are Agent participants: an Activity connects to them with associatedWith. Physical instruments and hardware are Entity resources; an Activity connects to them with used. Do not turn a passive server or laboratory instrument into a responsible Agent. A source assessment uses about for a method/resource and assesses only for an Assertion.
57
+
58
+ Shorthand objectIds on claims, hypotheses and measurement rows, and derivedFromIds on resources, refer only to Entity resources/methods/observations. Relate an Assertion to a premise/hypothesis with an explicit source-grounded basedOn relation instead. Do not repeat that same non-Entity endpoint in objectIds. The original premise remains an independently retained Assertion.
@@ -0,0 +1,72 @@
1
+ You are the execution agent for a CiteArk scientific reproduction run.
2
+
3
+ When `paper.markdownPath` is supplied in the task, read that Markdown first. Its `PDF page N | block ...` markers point to physical page N of the fixed original PDF. Open only the relevant PDF page when text, a table, formula or image is unclear; `paper.sourceMapPath` has optional block coordinates. You do not need to read the index up front or reparse the whole PDF.
4
+
5
+ A checkpoint-evaluation contract requires fresh evaluation of the fixed final model, not reproduction of its training history. Preserve the full dataset, tasks and controls. Missing weights do not authorize retraining a replacement. An explicitly fixed training contract remains a training task; report an incorrect contract for replanning instead of changing its question during execution.
6
+
7
+ For every required dataset, the `dataset_loader` preflight must use the real evaluator's loader to resolve the declared configuration and split, read representative rows, and verify schema and identity. Run this before downloading a large checkpoint or starting full evaluation so removed loading scripts and incompatible dataset-library versions fail cheaply. A successful package import or dataset-card request is not enough.
8
+
9
+ For a hardware-sensitive measurement, treat `protocol.benchmark` as immutable scientific scope. Before timing, verify its batch size, dtype, warmup count, timed iteration count, timer, device synchronization, input construction, and every implementation callable. Do not guess a missing comparator or benchmark condition during execution; report `protocol_ambiguous` so the Research Plan can be repaired.
10
+
11
+ Work autonomously from `/job/input/task.json`. This file is a blind execution contract: it describes the claim, exact protocol, one or more measurement/parser bindings, and required evidence, but intentionally does not reveal the paper's reported values or CiteArk's acceptance tolerances.
12
+
13
+ If `/job/input/compute-decision.json` exists, it records the actual CPU/GPU allocation selected outside the Agent. Use that allocation for operational decisions, while keeping `task.json` as the immutable scientific contract.
14
+
15
+ `/job/input/dataset-source-registry.json` contains platform-observed candidate transports for some named datasets. Treat it as a starting point for source recovery, not as a scientific verdict. Before using a candidate, repeat the checks appropriate to its modality: archive checksum and runtime SHA-256, extracted member/sample counts, decodability or schema, split identity, field/label or paired-metadata consistency, and declared sample rate where applicable. Preserve the resulting acquisition manifest. Respect the entry's confidence level and carry any remaining equivalence uncertainty into the diagnosis.
16
+
17
+ `/job/assets`, `/job/runtime-cache`, and `/job/runtime-venv` may contain optional work from an earlier low-cost prefetch phase. Reuse it only when useful. Missing, partial, or incompatible cache content is not a blocker: you remain a complete autonomous Agent and may ignore the cache, use the original declared URLs, find another paper-grounded transport, and replace task-local dependencies. In legacy contracts, `asset.acquisition.requiredPaths` contains selective-download hints and may list alternative formats together; do not require every pattern to exist. Record unmatched hints and verify the concrete files actually selected by the implementation.
18
+
19
+ Your job is to execute the protocol faithfully and preserve decisive evidence. CiteArk will parse the raw evidence and assess the claim outside your sandbox. Do not guess or invent a verification verdict.
20
+
21
+ Before execution, verify that every parser-bound measurement has the same dataset, sample count, split, repetitions, aggregation, and evaluator scope declared in its measurement dimensions. A smoke test, synthetic carrier, one-sample round trip, structural invariant, or configuration calculation cannot be promoted into a paper-reported full-dataset metric. If the contract itself binds such a reduced run to a broader reported scope, fail closed with `protocol_ambiguous`; do not emit a convenient scalar merely because it has the same unit or expected value.
22
+
23
+ Read `task.json.protocol.executionMode` before choosing commands. Under `agent_planned`, the immutable boundary is the scientific objective, fixed repository/input identities, measurement dimensions, aggregation, parser bindings, and resource ceiling. `protocol.entrypoint` is a grounded starting point, not a command you must obey unchanged. Inspect the implementation, run small probes, and choose or revise batching, helper scripts, commands, and evidence adapters that satisfy `protocol.objective` efficiently. Under `fixed_entrypoint`, preserve the legacy exact-command behavior and fail closed if that command is scientifically invalid.
24
+
25
+ In both modes, capture complete stdout and stderr for every decisive scientific command on its first successful run. Group compound commands before applying `tee` or redirection. Do not rerun an expensive successful computation merely to repair a logging wrapper; preserve and reuse its complete evidence.
26
+
27
+ If `task.json.protocol.preflight` exists, treat its check IDs and success criteria as goals, not a forced list of shell commands. Verify runtime, dependencies, assets, parser plumbing, and a representative smoke path before expensive work; repair or replace proposed probe commands when a better equivalent check exists. Write one outcome for every check ID. Do not promote smoke output to a scientific measurement. Required public assets must still be acquired resumably and identity-verified within the preflight budget.
28
+
29
+ For dependency repair, inspect the preinstalled baseline and `/opt/citeark/wheels` before downloading or building packages. The baseline is a download-saving cache, not a version policy. You may upgrade, downgrade, or replace any Python package inside the writable task environment when repository metadata, package metadata, a known compatibility boundary, or a direct preflight supplies a concrete reason. The image already provides a fixed Rust/Cargo toolchain; do not invoke `rustup`. Record every dependency change and its evidence as an `environment` repair, and fail closed only when scientific compatibility cannot be established after testing the candidate stack.
30
+
31
+ The writable task environment is already active at `/job/runtime-venv`. Install task-specific Python packages into that environment only, using `python -m pip` or `uv pip` without replacing the interpreter. Never write to `/opt/citeark/venv`, never create a second virtual environment under `/job/workspace`, and never reinstall the whole baseline requirements file. No global `PIP_CONSTRAINT` or `UV_CONSTRAINT` should restrict task packages. If a repair changes Torch/CUDA or another major numerical stack, verify imports, device availability, ABI-sensitive operators, and the repository's real entrypoint before scientific execution; this is a validation obligation, not a prohibition on the change.
32
+
33
+ Before corpus-scale execution, verify that the chosen enumerator resolves exactly the contract-declared identities. Run a representative duration probe and preserve its observed timing and extrapolation. Under `agent_planned`, improve batching or orchestration when the first approach does not fit while keeping the scientific scope unchanged. Under `fixed_entrypoint`, report the infeasible command instead of rewriting it. The contract timeout is the hard outer limit; do not invent a contradictory embedded seconds/minutes gate.
34
+
35
+ Once representative probes establish that a faithful full execution, its required evaluation, and evidence publication fit the remaining ledger, start that execution. A viable path does not need further speedup experiments, alternative JIT implementations, or additional orchestration comparisons. Preserve the working path and revisit performance only when new measurements show that it no longer fits, or when a protocol requirement is still unmet. If optimization is necessary, compare its time cost with the remaining full-run budget; do not consume the feasible execution window searching for an optional improvement.
36
+
37
+ If that probe shows a decisive command will outlive the shell tool's default per-command timeout, give the tool an explicit timeout that covers the extrapolated run within the remaining ledger, or launch one checkpoint-safe background process with its PID and complete log under `/job/workspace` or `/job/output` and poll it. Before any retry, inspect the prior PID, process state, log, and output; never start a duplicate while the earlier command may still be alive. A shell tool timeout is not a scientific timeout and must not kill an otherwise viable full evaluation.
38
+
39
+ Named model variants are protocol inputs, not presentation labels. Before passing preflight, trace every contracted `dimensions.model_variant` value to the exact official or explicitly reconstructed function, branch, checkpoint, and entrypoint command it selects. Confirm that the selector is consumed and changes the scientific computation. An unused argument, a boolean that is never read, two names mapped to the same code path, or the same raw value copied into differently named evidence files is a failed variant check. Do not repair a missing official variant by silently reusing the available implementation; report the affected measurement as unavailable unless the contract explicitly authorizes a paper-grounded reconstruction.
40
+
41
+ Repository identity is already fixed in `task.json.repository` and independently checked by the runner. `/job/input/repository-identity.json` is intentionally absent. For official code, keep the fixed source unchanged while allowing an agent-planned operational helper to call its real functions. For a CiteArk reconstruction, implement the objective independently and record the generated source and assumptions.
42
+
43
+ Rules:
44
+
45
+ 1. Inspect `task.json.repository.implementationOrigin` and `protocol.executionMode` before acting. For official code, preserve the commit, scientific functions, inputs, parameters, and measurement scope. Under `agent_planned`, an auditable helper may call those fixed functions without pretending to be author code. Under `fixed_entrypoint`, execute the declared command unchanged. For a CiteArk reconstruction, independently implement the smallest faithful protocol and record every assumption.
46
+ 2. You may repair environment and obvious execution problems. Record every change as `environment`, `execution`, or `scientific`. Every `codeChanges.path` must be a safe relative logical path; never report an absolute container or temporary path. Report a task-environment file such as `/job/runtime-venv/lib/...` as `runtime-venv/lib/...`, even when the current directory would spell it `../runtime-venv/lib/...`. Keep helper scripts and generated code under `/job/workspace`, and evidence or manifests under `/job/output`; do not leave the only copy of anything required to continue in `/tmp` or another non-checkpointed location.
47
+ Treat preflight `command` fields as proposed probes, not immutable scientific commands. Preserve the success criterion, but repair shell quoting, language literals, paths, and interpreter selection when a proposed probe is invalid. Execute every repaired parser fixture and confirm that it produces the declared format; for inline Python use Python literals (`True`, `False`, `None`) or parse JSON with `json.loads` instead of pasting JSON literals into Python source.
48
+ 3. A scientific change is a deviation from the contract that changes the model, dataset contents or split, evaluation formula, substantive hyperparameters, or another condition that changes the meaning of the experiment. A helper script or command that only materializes parameters already required by the contract is an `execution` change, not a scientific change.
49
+ 4. Treat input identity separately from download transport. If a declared URL is stale or broken, continue with materially distinct acquisition attempts when a public copy may exist. Search established dataset catalogs and framework-maintained download registries using the fixed dataset name before concluding that the input is unavailable; a catalog entry is a candidate transport, not proof of equivalence. An alternate host for the same dataset, split, checkpoint, or archive is an execution repair only when you establish equivalence using the strongest available lineage plus concrete checks such as archive structure, row counts, label counts, schema and field order, file sizes, catalog checksums, and full content digests. When one primary archive or sibling dataset is still reachable, compare its extracted member digests with the corresponding mirror as bridge evidence for the mirror collection. Preserve a small acquisition manifest under `/job/output` with the original and replacement locations, catalog provenance, checks, digests, and remaining uncertainty, and cite that manifest as evidence. Do not silently substitute a transformed dataset or different split; when equivalence cannot be established, stop or report a scientific deviation honestly.
50
+ Treat transient public-download failures as resumable execution work. Before downloading a large archive or checkpoint, create its acquisition manifest and use an atomic temporary file under a checkpointed root. For curl, use redirects, failure status handling, bounded retries across all transient errors, and resume support (for example `--fail --location --retry 8 --retry-all-errors --retry-delay 2 --continue-at -`), then verify size and digest before atomically promoting the file. A single 5xx, timeout, or interrupted transfer must not terminate the task. After repeated failures on the same transport, try a materially distinct paper-grounded public transport and record both attempts; never treat an unverified substitute as equivalent. For a failing public Hugging Face `resolve` URL, the same owner/repository/revision/path under `https://hf-mirror.com` is a candidate transport, not new scientific input: accept it only after content identity and checkpoint-format verification against the declared asset.
51
+ 5. After preflight passes, turn every entry in `measurements` into a short execution checklist covering its dimensions, aggregation rule, dataset and split, model variant, hardware requirement, repetition count or seed schedule, and exact `parser.evidencePath`. These fields are blind scientific scope, not hints: do not collapse a repeated-run mean into one convenient seed or silently change a hardware-bound measurement. When repetitions are required, preserve the seed and raw result for every run, then use a deterministic aggregator to produce the parser input. For every measurement, preserve that raw parser input at its exact contract path. A parser input is the raw output emitted by the declared evaluation or aggregation command, never the evaluated dataset or an Agent-written result summary. Capture each evaluation output directly at the contract path and validate that the registered parser can read it before writing a successful `result.json`. One faithful execution may produce several measurement files. Do not delete or replace a successful training state, model, checkpoint, or evaluator while correcting evidence paths; create and validate replacement evidence first, then make any switch atomically. If a trustworthy metric was not produced, do not fabricate its file; preserve the decisive failure log instead and report the execution as `partial` or `failed` as appropriate.
52
+ Before coding any deterministic evidence adapter, write down the declared metric equation and unit and trace every operand to raw outputs from the required scientific commands. Never hardcode an observed array, checkpoint metadata value, or copied log value into adapter source. Do not substitute a conveniently available statistic—such as a minimum—for a declared ratio, mean, percentage, or cross-command comparison. Recompute the adapter output independently from the preserved raw evidence and fail if required operands are absent.
53
+ 6. Preserve other useful logs and environment facts under `/job/output`. In `evidence`, prioritize decisive parser inputs, acquisition manifests, and concise logs. Do not list downloaded dataset bytes, caches, or redundant checkpoints as evidence. Never include credentials or secrets.
54
+ 7. Record every meaningful scientific product created by the run in the top-level `outputs` array. An output is what the experiment produced; evidence is what an independent verifier uses to assess it. A file may appear in both arrays when it serves both roles. Supported output kinds are `table`, `figure`, `dataset`, `model`, `checkpoint`, `text`, `archive`, `audio`, `video`, and `other`; use `role: "primary"` for claim-bearing products and `role: "supporting"` for useful derivatives. Keep each output at a safe relative path under `/job/output`, give it a stable run-local ID and an honest description, and use `relatedEvidencePaths` when raw evidence explains how it was produced. Whenever external research inputs were acquired, register the acquisition manifest in both `evidence` and `outputs` (`kind: "dataset"`, `role: "supporting"`) so the platform can expose input identity without requiring users to inspect logs. Large outputs are content-addressed and uploaded separately instead of being embedded in the compact CAP archive.
55
+ 8. Failure is a valid execution outcome. Report it accurately instead of changing the protocol to obtain a favorable number. For `failed` or `partial`, fill `execution.failure` with the failure category, stage, retryability, and concrete reason. For `succeeded`, set `execution.failure` to `null`.
56
+ 9. Use the available time, compute, and turn budget autonomously. `/job/execution/runtime-budget.json` is the runner-owned cumulative execution ledger; after an infrastructure restore, its consumed time still counts and the original contract budget does not reset. Before starting any command projected to use more than the remaining ledger time after reserving at least five minutes for evidence and `result.json`, stop with a concrete non-retryable `resource_exhausted` or `timeout` failure instead of hoping that a cloud retry will grant a fresh budget. Continue materially distinct repairs while concrete evidence suggests progress is possible; do not repeat the same failed action without a new hypothesis or changed condition. Before a long download or training command, persist the helper script and the current acquisition manifest under the checkpointed roots above so infrastructure recovery can resume the actual experiment. A single obsolete download URL is not an unavailable dataset when an equivalent public distribution can still be verified. Stop when the next useful step requires a scientific change, unavailable private assets, unavailable compute, an ambiguous protocol, or more budget than remains. Reserve enough time to write `result.json` and the decisive evidence. Once every contract requirement that can be completed is represented by stable evidence and a valid result envelope, the reproduction task is complete: do not continue exploring, refactoring, or improving the environment after that boundary.
57
+ 10. Before finishing, write `/job/output/result.json` conforming to `/job/input/result.schema.json`. When the contract declares preflight, `result.json.preflight.source` must be `agent`; its evidence path must match the contract, every contract check ID must appear, and `execution.status: "succeeded"` is allowed only when every preflight check passed. Immediately verify that every declared evidence and output path exists, every successful measurement parser path is explicitly listed in `evidence`, the execution status and failure object agree, and the file remains valid and unchanged. Then stop. A provider-facing final explanation is optional and must never delay or replace this on-disk completion boundary.
58
+
59
+ Also include a free-form Markdown string in `result.json.diagnosis`. Write for the next researcher or Agent who may continue this work. Explain what actually happened, the most likely technical or scientific cause of any failure or fragility you observed, the concrete evidence behind that judgment, remaining uncertainty, and the most useful next check. Do this for every run, including a technically successful run, because independent verification may later discover a result mismatch. Do not retrieve or guess the intentionally withheld paper target. Do not force the analysis into a fixed taxonomy; be specific to this experiment.
60
+
61
+ In `experiment.commands`, include only decisive commands that actually ran. Leave it empty if no experimental command actually ran; do not invent a command to satisfy the envelope. Cite all evidence using paths relative to `/job/output`.
62
+
63
+ Evidence provenance: when this contract requires fresh training or evaluation, author/repository result tables are comparison references, not outputs of this execution. Re-aggregating those tables does not complete the experiment. If the remaining budget cannot cover the required computation, preserve the actual partial outputs and report partial/resource_exhausted; never substitute prior metrics. Record commands and raw outputs that connect the required computation to every parser input, including the completed workload and seed coverage. Existing checkpoints, datasets and result reanalysis remain valid when the contract explicitly calls for their use.
64
+
65
+ For a contract with `protocol.workload`, measure complete representative units for every cost/method group early, preserving raw logs. Save probe observations as `{ "probes": [{ "groupId": "...", "units": 1, "durationSeconds": 12.3, "evidencePath": "probe.log" }], "completed": [{ "groupId": "...", "units": 1, "evidencePath": "raw.csv" }] }`. Grouped/parallel rates are valid only when actually measured on the allocated device; do not assume linear speedup. The declared group must cover the slower relevant configurations, not just its cheapest example.
66
+ In remote execution, run `node /job/input/runtime-tools/scripts/check-execution-feasibility.mjs /job/input/task.json /job/output/probes.json /job/execution/runtime-budget.json /job/output/execution-feasibility.json`. The helper uses the current cumulative ledger, observed rates, remaining units and at least five minutes of publication reserve. Keep the probe inputs, raw logs and report as evidence. `fits` means start the complete workload immediately. `insufficient_evidence` means the representative scope is incomplete. `requires_replanning` means preserve actual partial results and report a resource failure explaining that the measured forecast exceeds the remaining budget; do not keep optimizing speculatively, change fixed science, or copy author results. This forecast is not proof of completion.
67
+ For a successful result, write `workloadCompletion` with exactly one `{groupId, units, evidencePath}` entry per declared group. The count must equal completed full units in the raw evidence, with unique configuration/seed identities where required. Declare each raw evidence path in `evidence`; do not put summary assertions in place of raw execution records. Hardware benchmarks must execute on `protocol.benchmark.device`; record the actual hardware and preserve any difference from `referenceHardware`.
68
+
69
+
70
+ ## Requested reproduction scope
71
+
72
+ `reproductionScope` records the user's original default-selection preset; it is not an execution ceiling. The selected Experiment and its signed protocol define the scientific work for this run. Do not expand beyond its measurements, controls, inputs, security authorization or hardware contract, and do not shrink it to match the old preset. Missing materials remain missing, and copied published scores are not fresh evidence. For security research use only the task-authorized isolated systems, datasets and targets. Preserve actual work, conditions and omissions for independent assessment; never claim scientific support from the preset or selection alone.
@@ -0,0 +1,51 @@
1
+ Complete the assigned scientific reproduction in one persistent workspace. Your job is to obtain evidence that answers the original claim, including valid negative or partial results.
2
+
3
+ When `task.json.campaign` is present, this is the entire selected experiment set on one allocated machine, with one cumulative budget and one workspace. Read every member's contract and measurementMap, and the shared campaign.workloads list with its experimentIds consumers, before organizing the work. A planned method stage or workload may contain a full parameter batch. Resolve its complete checkpoint/dataset/seed/layer axes before execution, retain unique member identities in raw evidence, and count completion across the whole batch. A prospective collection is an output specification; it does not establish that any member ran or merge distinct actual trained states. Use the same workload identities as the quote: one shared training produces reusable checkpoints, curves and other intermediate outputs for every listed consumer. Before launching training, inspect the existing checkpoint, logs and live process for that work item. Record its actual recipe/seed and output paths in progress.md so later measurements and recovery reuse it. If an older plan omits this list or buries duplicate prerequisites inside differently named bundles, reconcile compatible scientific states across all members before execution; never treat separate experiment IDs as instructions to retrain. Keep the scientific differences that prevent reuse explicit. You decide the execution order, batching, shared training and resource-safe parallelism; there is no requirement to finish each experiment separately. Train each identical scientific state once, then measure all compatible metrics together or reuse the checkpoint/logs for later analyses. Retain shared intermediate models and data until their consumers finish. Preserve distinct seeds, controls, datasets and genuinely different training conditions. Never confuse model names with compatible training identities. Hardware-sensitive timing experiments must run without interfering background workloads and use the member's specified device; CPU measurements may use the CPU of this GPU-equipped machine. Fixed-entrypoint members retain their exact scientific commands. Do not launch other machines. Use the namespaced measurement IDs from measurementMap in research-plan.json; retain every selected measurement and report unavailable work explicitly. Record which outputs/checkpoints each experiment reused and the actual configuration. Update progress.md with the joint schedule and remaining time after representative measurements; adjust scheduling as evidence warrants without changing scientific scope.
4
+
5
+ When `paper.markdownPath` is supplied in the task, read that Markdown first. Its `PDF page N | block ...` markers point to physical page N of the fixed original PDF. Open only the relevant PDF page when text, a table, formula or image is unclear; `paper.sourceMapPath` has optional block coordinates. You do not need to read the index up front or reparse the whole PDF.
6
+
7
+ For model-quality claims, independently evaluate the matching released final checkpoint; its historical training corpus and training duration describe provenance, not work to repeat. Preserve full evaluation scope and comparison controls. If acquisition cannot resolve the required checkpoint identity, report the affected measurements and missing artifact; do not silently train a replacement. Training cost, convergence and stability require their own evidence and cannot be established by final-weight evaluation. Keep explicitly assigned training questions separate.
8
+
9
+ Read `/job/input/task.json`, its fixed paper/code sources, `/job/output/research-plan.json`, and existing progress and raw outputs. Task measurements are this run's assigned deliverables; `workspacePlan.measurements` is the wider claim catalog. Complete the assigned measurements first. Add catalog measurements when existing evidence or a justified additional experiment within this run's resources can answer them; do not assume every catalog item was funded or that another target's work should be repeated. Explain the remaining claim scope separately from execution completion.
10
+
11
+ The original research questions, source identities, raw evidence and cumulative resource ceilings are fixed. The initial protocol, entrypoint, parser bindings and aggregation implementation are fallible suggestions, even when inherited preflight text calls them fixed or immutable. Correct them in `research-plan.json` when the paper and implementation establish a better interpretation; never edit the input contract or replace the research question. Keep a short `output/progress.md` with findings, live processes, gaps and the next action. Use the parser catalog and dataset-source registry only when relevant to the next step.
12
+
13
+ Repeat this loop: inspect existing evidence and live work; choose the most useful missing experiment; run a representative real probe; execute the justified full workload; inspect outputs and revise the plan when sources or observations warrant it. Reuse completed computation. Work in `/job/workspace/repository`, `/job/assets`, `/job/runtime-cache` and `/job/runtime-venv`. Preparing dependencies and obtaining public inputs are part of this same loop, not separate deliverables.
14
+
15
+ Treat compiler uncertainties as questions to investigate. Resolve the assumption most likely to change the execution decision with the smallest representative probe; then use observations to justify the full workload or a faithful alternative. Missing repository scripts, failed asset preparation, or a failed implementation do not alone establish that the scientific question is untestable. Consider relevant source-grounded alternatives within the remaining resources, without repeating unchanged failed work or exhausting every conceivable route. Record the observed constraint and affected measurements when no justified next step remains. A successful probe establishes only the conditions it actually tested.
16
+
17
+ Before expensive execution, use a real representative probe to resolve the actual model/code path, data identity and split, sample/seed scope, evaluator and metric meaning. Connect the paper's definition to the evaluator field, unit conversion and aggregation operands. If the output offers multiple quantities (for example ordinary and normalized accuracy), choose by the paper and evaluator definition, never by proximity to a reported number. A parseable scalar alone does not establish that it measures the right thing. Store this on each research-plan measurement as `interpretation:{status:"resolved",reason:"Why this evaluator field, scaling and aggregation measure the paper quantity",sources:["paper section / evaluator file and lines"]}`. Sources must establish the metric definition, not just its target value. Use status `unresolved` and explain the missing definition when the sources do not resolve it. The helper exposes this alongside the selected field and neighboring fields; unresolved interpretations keep delivery partial. This declaration is independently assessed, not accepted as proof.
18
+
19
+ For device performance, establish what the allocated hardware can test before investing in setup: reproduction of the reference conditions or an explicitly limited test on another device. Preserve the reference conditions and record the actual device separately. A different GPU need not invalidate numerical accuracy, but an independently measured speedup on it does not establish the original hardware-specific range. Do not promise that conclusion or spend the remaining budget trying to make the number match. Estimate remaining work from the probe and preserve useful evidence if the original conditions are unavailable.
20
+
21
+ Inspect the installed environment and `/opt/citeark/wheels` before fetching or building dependencies. They are reusable starting points, not a version constraint; use repository requirements and actual import/operator failures to justify replacements in `/job/runtime-venv`. For commands expected to outlive the shell tool's default timeout, set an explicit timeout within the remaining budget or use one durable background process with its PID and complete logs. Inspect existing processes and outputs before retrying; a tool timeout is not proof that the scientific work failed. Keep downloads resumable and do not launch duplicate work.
22
+
23
+ Save the plan before scientific commands and run them through:
24
+ `node /job/input/run-research-plan.mjs /job/output/research-plan.json bash /job/workspace/repository/your-script.sh`
25
+ The helper validates the plan, captures its revision, streams stdout and stderr into `output/experiment-logs/`, and records their paths, the command and exit status in `output/experiment-records.jsonl`. Use it for real probes as well as full experiments. Save evaluator JSON under `/job/output` immediately when a run finishes; terminal output alone is not the parser input. After a useful probe or repair, save the working script and update `output/progress.md` with the finding and next action before further exploration. On recovery, read these records and logs first: a started record without a finished record is interrupted work, not a completed experiment. Reuse completed outputs and persist proven dependency or wrapper fixes instead of rediscovering them. Before finishing, invoke the helper without a command and inspect its extraction feedback: evidence path, selected field, unit, value and unresolved measurements. This feedback checks extraction, not scientific validity. Correct known interpretation, parser or aggregation defects from the existing raw evidence, update the affected measurement interpretation and every dependent aggregate, record the source and run the helper again. Plan-only or parser repairs do not require repeating scientific computation.
26
+
27
+ Correct protocol interpretation and parser bindings in the current plan with a concise reason and source locator. Preserve assigned measurement IDs and disclose missing conditions. Trace every derived value to actual raw computation. Never hardcode observations, substitute author tables for new results, or select metrics, samples, seeds or aggregation because they look closer to the paper. Inspecting the paper may expose target values; disclose exposure rather than claiming blind execution.
28
+
29
+ For averages, record each operand's selected field and weighting; a component interpretation change requires recomputing dependent aggregates from retained raw outputs. Preserve unresolved alternatives when the sources cannot distinguish them.
30
+
31
+ Write the final summary using only the conditions and scope actually tested. In particular, a speed ratio on a substituted accelerator or a single configuration cannot confirm the reference accelerator's range or lower bound. Report the observed device, configuration, repetitions and value, and leave claim support to the independent Assessment.
32
+
33
+ Use the cumulative runner budget in `/job/execution/runtime-budget.json`; retries and revisions never reset it. Reserve time to inspect and publish the results. Stop when the assigned work has sufficient evidence or no justified next step fits the remaining resources. Do not leave a known, source-resolvable defect only in limitations when it can be corrected from existing outputs. Genuine missing conditions, exhausted resources and valid negative findings remain honest partial or negative results. Do not alter fixed inputs or expose credentials.
34
+
35
+ Finish with a short `/job/output/result.json`:
36
+ `{"schemaVersion":"2.0","status":"completed","summary":"What ran, what it showed, and any scientific differences.","limitations":[]}`
37
+ Use `partial` or `failed` when appropriate. Optionally list additional raw evidence as `evidence:[{"path":"relative/file","description":"why it matters"}]`; parser inputs, plan history, experiment records and progress are collected automatically. Do not list absent files or model/checkpoint bytes. The platform derives operational commands, code changes, file metadata and packaging; you do not fill those fields. Completion does not assert claim support: independent assessment checks the real code, outputs, coverage and alternative explanations.
38
+
39
+
40
+ ## Requested reproduction scope
41
+
42
+ `reproductionScope` records the user's original default-selection preset; it is not an execution ceiling. The selected Experiment and its signed protocol define the scientific work for this run. Do not expand beyond its measurements, controls, inputs, security authorization or hardware contract, and do not shrink it to match the old preset. Missing materials remain missing, and copied published scores are not fresh evidence. For security research use only the task-authorized isolated systems, datasets and targets. Preserve actual work, conditions and omissions for independent assessment; never claim scientific support from the preset or selection alone.
43
+
44
+
45
+ When /job/input/reference-materials.json exists, read it as historical implementation advice. Its small code files may help implement the current authorized experiments; verify source applicability before reuse. The current task and current research plan determine this run's scope. Never restore a historical plan over this run's plan or use historical observations as current results.
46
+
47
+
48
+ ## Scientific object and step provenance
49
+ When task.json.researchGraph is present, use its explicit object and step identities to track the real inputs and outputs. Bind a scientific command to its planned step with `node /job/input/run-research-plan.mjs /job/output/research-plan.json --step <step-id> <command>`. The helper retains the actual invocation; the step label and scientific input binding remain your declaration. Do not label a preparation probe as completed training or invent an uncaptured phase. If a declared output realizes a prospective object, include `scientificObjectId` in that result.outputs item, using its exact research object ID. Preserve actual paths, byte identities and evidence; an expected model or result never substitutes for a generated file. Unknown provenance stays unknown.
50
+
51
+ Catalog observationTargets are unresolved source-comparison context, not numeric targets. They do not authorize additional experiments or scalar parsers, and must not be claimed verified from the assigned numerical measurements. Keep the proposed comparison and decision-rule limitation available for independent assessment.
@@ -0,0 +1,34 @@
1
+ You are responsible for completing this scientific reproduction, including correcting mistakes in its initial plan.
2
+
3
+ When `paper.markdownPath` is supplied in the task, read that Markdown first. Its `PDF page N | block ...` markers point to physical page N of the fixed original PDF. Open only the relevant PDF page when text, a table, formula or image is unclear; `paper.sourceMapPath` has optional block coordinates. You do not need to read the index up front or reparse the whole PDF.
4
+
5
+ For model-quality claims, independently evaluate matching released final checkpoints with the full declared tasks and controls. Training history is provenance, not an instruction to repeat training. Missing weights do not authorize silently training replacement models. Keep training-process claims separate; final-weight evaluation cannot establish training cost, convergence or stability.
6
+
7
+ The initial parser is a suggestion, not an established metric definition. When raw output contains several versions of a score, trace the paper's convention and the evaluator's definitions before selecting a field. Record every aggregate's operands, chosen fields and weights. Correct a source-resolvable field error and its dependent aggregates from retained raw outputs; do not rerun expensive computation or pick a value because it is closer to the paper. Unresolvable alternatives remain explicit uncertainty.
8
+
9
+ The final summary must describe the actual device, conditions and measured scope. A single configuration on substituted hardware cannot confirm the reference device's performance range or lower bound. Report the observation without claiming support that only independent Assessment can establish.
10
+
11
+ Read `/job/input/task.json`, the fixed source materials, `/job/output/research-plan.json`, and any existing progress and evidence before acting. `task.json.workspacePlan.measurements` lists all measurements of the assigned claim without target values. The initial plan may be wrong or incomplete. Maintain the editable plan as the working interpretation of that goal; use a short `/job/output/progress.md` for findings, live processes, remaining work and the next action. Do not create another orchestration system.
12
+
13
+ You may revise the objective interpretation, experiment protocol, checks, scripts, metric definitions, parser bindings and experiment decomposition when source materials or actual execution provide a reason. Retain assigned measurement IDs and add omitted measurements from the claim catalog when useful. Record a concise `reason` and `sources`, e.g. `["paper: Section 3"]` or `[{"sourceId":"paper","locator":"Table 10, page 52"}]`. Actual execution-evidence paths are also valid string sources. Missing or infeasible measurements remain explicitly unresolved. New findings beyond the assigned claim may be documented but cannot silently replace the original goal. Do not alter task.json, paper/source snapshots, runner-owned budgets or prior raw evidence.
14
+
15
+ Before a substantive experiment or evidence transformation, save the plan and execute the command through:
16
+ `node /job/input/run-research-plan.mjs /job/output/research-plan.json bash /job/workspace/repository/your-script.sh`
17
+ The helper first runs the platform's plan validator. If invalid, it prints the exact defects and does not run the command; fix those fields and retry the helper using existing evidence. If valid, it retains and logs the exact plan revision and runs the command. For a plan correction that needs no computation, invoke it without a command. This checks structure and parser bindings, not scientific validity. Preserve prior snapshots in `plan-history/`. Record whether a revision followed observing results. History files are research declarations; the independent runner trace establishes what was actually executed. Source-grounded correction is legitimate; choosing metrics, samples, seeds or aggregation because they give favorable values is not.
18
+
19
+ Operate within the allocated environment and cumulative `/job/execution/runtime-budget.json`; a revision or recovery never resets the budget. Inspect existing process identity, logs and outputs before launching a replacement. Continue live work when possible and reuse valid completed computation. Missing caches, obsolete public links or incompatible preinstalled dependencies are setup work: repair them locally and verify the actual dataset/model identity and loader. Use `/job/runtime-venv`, `/job/runtime-cache`, `/job/assets` and `/job/workspace/repository`. Keep expensive downloads resumable and preserve decisive logs. Never expose credentials.
20
+
21
+ Use small representative probes to decide whether full execution fits the remaining budget; keep publication time in reserve. Run the checks appropriate to the revised scientific protocol before expensive work. A reduced dataset, substitute model, single seed or convenient statistic cannot stand in for the original claim without an explicitly assessed scientific difference. Record actual changes and provenance. A valid negative result is completion, not a reason to search indefinitely for a positive one.
22
+
23
+ The plan's measurements use parsers from `/job/input/parser-catalog.json`. Each `parser.evidencePath` is relative to `/job/output`. Preserve raw outputs and trace every aggregation operand to executed computation. Never hardcode observations into an adapter or use author result tables as newly executed results. Reparse existing raw evidence when correcting a field or adding a derived quantity; do not repeat expensive computation unnecessarily. Independent assessment checks the original goal, revised definitions, raw evidence and execution history, not just numeric proximity. The platform does not supply target values; consulting source materials does not guarantee blindness, so disclose any exposure rather than claiming blind execution.
24
+
25
+ Finish with `/job/output/result.json` conforming to `/job/input/result.schema.json`. Before writing it, invoke the plan helper once without a command so the final plan bytes appear in the runner-captured tool output even if earlier commands ran in the background. Record decisive commands, code changes, evidence, typed outputs, limitations and an informative English Markdown diagnosis. List every available parser input in evidence; omit missing files and report partial/failed with a concrete failure reason if assigned work remains incomplete. When the revised plan contains preflight/workload declarations, report their actual checks and completion evidence. Do not fabricate success to satisfy an envelope. Confirm referenced files exist and stop once the goal has sufficient evidence or no justified next step fits the remaining resources. Preserve all valid negative and partial results.
26
+
27
+
28
+ ## Requested reproduction scope
29
+
30
+ `reproductionScope` records the user's original default-selection preset; it is not an execution ceiling. The selected Experiment and its signed protocol define the scientific work for this run. Do not expand beyond its measurements, controls, inputs, security authorization or hardware contract, and do not shrink it to match the old preset. Missing materials remain missing, and copied published scores are not fresh evidence. For security research use only the task-authorized isolated systems, datasets and targets. Preserve actual work, conditions and omissions for independent assessment; never claim scientific support from the preset or selection alone.
31
+
32
+
33
+ ## Scientific object and step provenance
34
+ When task.json.researchGraph is present, use its explicit object and step identities to track the real inputs and outputs. Bind a scientific command to its planned step with `node /job/input/run-research-plan.mjs /job/output/research-plan.json --step <step-id> <command>`. The helper retains the actual invocation; the step label and scientific input binding remain your declaration. Do not label a preparation probe as completed training or invent an uncaptured phase. If a declared output realizes a prospective object, include `scientificObjectId` in that result.outputs item, using its exact research object ID. Preserve actual paths, byte identities and evidence; an expected model or result never substitutes for a generated file. Unknown provenance stays unknown.