medsci-skills 5.3.0 → 5.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +13 -13
- package/metadata/distribution_files.json +40 -10
- package/metadata/distribution_manifest.json +1 -1
- package/package.json +1 -1
- package/skills/check-reporting/SKILL.md +6 -2
- package/skills/check-reporting/references/checklists/CROSS.md +78 -0
- package/skills/check-reporting/references/checklists/RECORD.md +66 -0
- package/skills/check-reporting/scripts/check_checklist_exists.py +2 -0
- package/skills/make-figures/references/reporting_guideline_figure_map.md +2 -2
- package/skills/orchestrate/SKILL.md +1 -1
- package/skills/peer-review/SKILL.md +12 -0
- package/skills/peer-review/references/domain-probes/record_routinely_collected_data.md +49 -0
- package/skills/peer-review/references/domain-probes/survey_research.md +51 -0
- package/skills/self-review/SKILL.md +2 -0
- package/skills/self-review/references/domain-probes/record_routinely_collected_data.md +49 -0
- package/skills/self-review/references/domain-probes/survey_research.md +51 -0
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
**51 skills that actually work.** Built by a physician-researcher, tested on real publications.
|
|
6
6
|
|
|
7
|
-
*MedSci Skills is an end-to-end research tool for physician and medical-engineering researchers — design → scaffold → validate → publish — for the clinical manuscript and the medical-AI model behind it. Its moat is the compliance layer —
|
|
7
|
+
*MedSci Skills is an end-to-end research tool for physician and medical-engineering researchers — design → scaffold → validate → publish — for the clinical manuscript and the medical-AI model behind it. Its moat is the compliance layer — 41 reporting guidelines and risk-of-bias tools, reference/citation verification, and deterministic integrity gates before peer review — now extended by a model-engineering lane that scaffolds reproducible, leakage-safe training repos and audits model validation. Clinical AI model research engineering is in scope; a general AI-scientist platform is not. It competes on clinical submission reliability, not skill count.*
|
|
8
8
|
|
|
9
9
|
[](LICENSE)
|
|
10
10
|
[](https://github.com/Aperivue/medsci-skills/releases/latest)
|
|
@@ -281,54 +281,54 @@ The E2E pipeline (`orchestrate --e2e`) produces everything up to `qc/`. The `sub
|
|
|
281
281
|
|
|
282
282
|
## What's New
|
|
283
283
|
|
|
284
|
-
**v4.10** — reviewer-coverage expansion reverse-engineered from high-IF, CC-BY papers (learn-only under the `reverse_engineer/` license firewall), plus a clinician-friendly update path. Additive and backward-compatible; 45 skills / **
|
|
284
|
+
**v4.10** — reviewer-coverage expansion reverse-engineered from high-IF, CC-BY papers (learn-only under the `reverse_engineer/` license firewall), plus a clinician-friendly update path. Additive and backward-compatible; 45 skills / **41 guidelines** / 36 detectors / **15 domain-probe modules** (was 12):
|
|
285
285
|
|
|
286
286
|
- **Three new reviewer domain-probe modules** (`/peer-review` + `/self-review`, vendored byte-identical): **Mendelian randomization** (MR1–MR8 — IV assumptions, pleiotropy-robust sensitivity suite, Steiger, sample overlap, NLMR, drug-target colocalization), **polygenic risk score** (PG1–PG8 — ancestry portability, base/target leakage, incremental value over the clinical model, screening-vs-discrimination, calibration), and **network meta-analysis** (NM1–NM8 — transitivity, incoherence, SUCRA over-interpretation, CINeMA/GRADE-NMA, component-NMA additivity). Plus observational **O17** (agnostic many-exposure-scan multiplicity: ExWAS/EWAS/MWAS).
|
|
287
287
|
- **Two reporting-guideline checklists** (36 → 38): **STROBE-MR** and **PGS-RS / PRS-RS**, with study-type routing. Four new `/analyze-stats` analysis guides (multiplicity, MR, PRS, NMA) and a `/clean-data` implausible-value + cross-field validity reference.
|
|
288
288
|
- **Clinician-friendly update reminders** — the classroom installers enable the in-app "update available" notice + one-click Desktop updater by default; the `npx`/manual paths print how to turn it on; the install guide recommends `npx medsci-skills install --enable-update-notify`.
|
|
289
289
|
|
|
290
|
-
**v4.9** — analysis-integrity hardening promoted from real review cycles, plus journal-mechanics additions. Additive and backward-compatible; still 45 skills /
|
|
290
|
+
**v4.9** — analysis-integrity hardening promoted from real review cycles, plus journal-mechanics additions. Additive and backward-compatible; still 45 skills / 41 guidelines, analysis-integrity detectors **32 → 36**:
|
|
291
291
|
|
|
292
292
|
- **Four new gates** — a **duplicate-bibliography** check (`check_reference_duplication.py`) for the hybrid `[@key]` + hand-typed `## References` build that renders the list twice; a **cross-script binning / composite-indicator** consistency check (`check_binning_consistency.py`, `BINNING_DRIFT` / `DERIVED_DEF_DRIFT`) for a derived categorical or composite indicator defined inconsistently across analysis scripts; a **float citation-order** check (`check_citation_order.py`) for numbered Tables/Figures not first cited in ascending order per series; and an **audit-dump leak** gate (`/sync-submission`) that blocks a `/check-reporting` output mistakenly attached as a submission file.
|
|
293
293
|
- **KJR technical-check conventions + percentage-decimal style**, reader-allocation-under-burden and generative-image-as-study-object reporting (`/design-ai-benchmarking`, `/check-reporting`), and a **Liver International** CSL with that journal's submission mechanics (`/manage-refs`).
|
|
294
294
|
|
|
295
|
-
**v4.8** is the **review-harvest batch** — deterministic detector hardening promoted from real-manuscript review cycles. Additive and backward-compatible; still 45 skills /
|
|
295
|
+
**v4.8** is the **review-harvest batch** — deterministic detector hardening promoted from real-manuscript review cycles. Additive and backward-compatible; still 45 skills / 41 guidelines, analysis-integrity detectors **30 → 32**:
|
|
296
296
|
|
|
297
297
|
- **Two new gates** — `check_supplement_hygiene.py` lints the rendered supplement / tables / caption files (not just the manuscript) for §-labels, placeholders, build markers, response-letter framing, and unresolved body↔supplement cross-references; `check_null_calibration.py` flags a headline negative/equivalence claim made without a minimum-detectable-effect / power / equivalence statement.
|
|
298
298
|
- **Four detector false-positive fixes** — gates no longer fire on a recommended colorblind-safe palette, author-footnote `§` daggers, a correctly-hedged disclaimer, or a tier-label digit; each with a regression fixture and three newly CI-wired test suites.
|
|
299
299
|
- **Nine reviewer-side domain probes** (SR/MA, observational, diagnostic, AI-overclaiming, survival) plus a `/design-study` design-stage ceiling gate for perceptual/reader-AI studies and a reusable confidence-weighted-rating→AUC monotonicity template.
|
|
300
300
|
|
|
301
|
-
**v4.7** is the **self-update foundation** — physician-researchers stay current without GitHub, git, or a terminal. Additive and backward-compatible; still 45 skills /
|
|
301
|
+
**v4.7** is the **self-update foundation** — physician-researchers stay current without GitHub, git, or a terminal. Additive and backward-compatible; still 45 skills / 41 guidelines / 30 detectors:
|
|
302
302
|
|
|
303
303
|
- **Transactional, crash-recoverable installer.** Each install runs through a durable journal state machine recovered on the next run (roll back / forward-clean / fail-closed), with per-target SHA-256 inventories — your modified or third-party skills are backed up and never clobbered or auto-deleted.
|
|
304
304
|
- **One-click self-updater** (`~/.medsci-skills/updater/`, `install.py --check-update`). Verifies the download against the github.com API digest and **never `extractall()`s** (per-entry rejection of traversal / symlink / duplicate / zip-bomb + an allowlist & per-file hash). The release pipeline injects a verified `provenance.json`, attests build provenance, runs on a protected `release` environment, and verifies each ZIP round-trips through the updater's own safe-extract before publishing.
|
|
305
305
|
- **Opt-in update notice (off by default):** `install.py --enable-update-notify` shows a one-line "update available" message at Claude Code session start — no telemetry, reads nothing about your session, installs nothing. `--disable-update-notify` / `MEDSCI_NO_UPDATE_CHECK=1` turn it off. *(Honest scope: the digest/attestation detect transport tampering, not a compromised publisher account — see `SECURITY.md`.)*
|
|
306
306
|
|
|
307
|
-
**v4.6** is a maintainability, governance, and review-depth release — still 45 skills /
|
|
307
|
+
**v4.6** is a maintainability, governance, and review-depth release — still 45 skills / 41 guidelines; analysis-integrity detectors **28 → 30**, domain probes 11 → 12:
|
|
308
308
|
|
|
309
309
|
- **Fairness / equity / subgroup-performance probe (EQ0–EQ6)** for AI/prediction/diagnostic studies that claim cross-population performance, plus two new detectors: an **AI-disclosure + data/code-availability** check (`/sync-submission`) and a **structured-summary-box conformance** check (`/academic-aio`).
|
|
310
310
|
- **Governance + answer-engine layer:** `ROADMAP.md`, `MAINTAINERS.md`, `SECURITY.md`, a maintainer workflow + release checklist, an AEO/GEO `docs/faq.md`, a "Start here: 3 workflows" + "Validation status" section in this README, and a new `maturity` field (official / experimental / community) on every skill.
|
|
311
311
|
- **Token diet (pilot):** `write-paper` Phase 7 integrity audits moved to a load-on-demand reference (~2,559 tokens saved per invocation). Positioning now leads with the compliance moat rather than skill count.
|
|
312
312
|
|
|
313
|
-
**v4.5** deepens the review + submission surface with no new skill or reporting-guideline count (still 45 skills /
|
|
313
|
+
**v4.5** deepens the review + submission surface with no new skill or reporting-guideline count (still 45 skills / 41 guidelines); analysis-integrity detectors **27 → 28**:
|
|
314
314
|
|
|
315
315
|
- **`/clean-data` + `/analyze-stats` — reverse-coded-item / negative-alpha detector.** A multi-item Likert scale with a negatively-worded item must be recoded `(min+max) − x` before the scale total or Cronbach's alpha is computed; left un-recoded, the item correlates negatively with the rest of the scale and alpha collapses (often negative). A negative alpha is a coding bug, not a "multidimensional construct." New stdlib-only `check_reverse_coding.py` returns `REVERSE_CODING_LIKELY` / `REVERSE_CODING_SUSPECT` / `OK` from per-item item-rest correlations + raw alpha; the Likert summary template gains a `--reverse-items` recode flag.
|
|
316
316
|
- **`/peer-review` + `/self-review` — SR/MA + DTA + prediction-model probe batch.** `sr_ma.md` **P12** risk-of-bias table row-sum ↔ traffic-light figure-matrix reconciliation and **P13** included-study ↔ reference-list completeness; `diagnostic_accuracy.md` **D7** index-test-as-enrollment-criterion circularity; `clinical_prediction_model.md` **CP5** intended-use horizon leakage and **CP6** development/CV vs held-out/external validation-nomenclature conflation. Vendored byte-identical into `/self-review`.
|
|
317
317
|
- **`/sync-submission` — embedded absolute-path leak scan.** A `word/*.xml` attribute (e.g. a pandoc-embedded image's `<pic:cNvPr descr="…">`) carrying an absolute home-dir path (`/Users/…`, `/home/…`) is a username leak invisible to a rendered-text scan; now flagged as `docx_embedded_abs_path` under `check_asset_anonymization.py`.
|
|
318
318
|
|
|
319
|
-
**v4.4** adds reviewer/analysis depth with no new skill or reporting-guideline count (still 45 skills /
|
|
319
|
+
**v4.4** adds reviewer/analysis depth with no new skill or reporting-guideline count (still 45 skills / 41 guidelines / 27 detectors):
|
|
320
320
|
|
|
321
321
|
- **`/author-strategy` — trajectory-archetype classification (optional).** Classifies a queried author's PubMed trajectory into abstract career archetypes (A1 infrastructure builder, A2 methodology rule-maker, A3 clinical→AI hybrid, A4 SR/MA volume engine, A5 large-consortium participation, A6 device/technique depth, + a computed composite) as an **explainable, multi-label, confidence-scored heuristic — not an objective verdict**. The rubric is a single canonical YAML (the narrative doc is generated from it); scores exclude `unavailable` signals (h-index/citation/venue-tier → `[VERIFY]`, never fabricated); a **disambiguation gate** binds an approved `corpus_manifest.json` to the CSV (csv + PMID-set hashes) so a surname alone never classifies, and target-author attribution never borrows a co-author's ORCID/affiliation.
|
|
322
322
|
- **`/peer-review` + `/self-review` — Image-Synthesis / cross-modality probe (IS1–IS4)** for studies that synthesize one imaging modality from another and claim the output carries the target's information, plus a reviewer-side reference-integrity spot-check.
|
|
323
323
|
- **`/verify-refs` — OpenAlex tertiary index** recovers conference-proceedings / non-DOI citations (NeurIPS/ICLR/ACL) that fall through PubMed and CrossRef, the free analogue of a portal's second index.
|
|
324
324
|
|
|
325
|
-
**v4.3** hardens the **cross-sectional / observational cohort** review surface end-to-end, much of it reverse-engineered from real CC-BY cohort papers (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills /
|
|
325
|
+
**v4.3** hardens the **cross-sectional / observational cohort** review surface end-to-end, much of it reverse-engineered from real CC-BY cohort papers (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills / 41 guidelines); analysis-integrity detectors **25 → 27**:
|
|
326
326
|
|
|
327
327
|
- **Observational probes O1 → O14** (`/peer-review` + `/self-review`, vendored) — over-adjustment / analysis-unit clustering / outcome construct-validity (O7–O9), overlapping-subset gradient (O10), **complex-survey design & weighting** for NHANES/KNHANES (O11), **data-driven threshold / "inflection-point" mining** (O12), **cross-sectional mediation** temporal-order & sequential-ignorability (O13), and **interaction scale** — additive RERI/AP/S vs multiplicative (O14). Plus a new **clinical-prediction-model** probe module **CP1–CP4** and survival **S9** (panel-data / multistate variance).
|
|
328
328
|
- **Two new detectors (25 → 27)** — `check_wordcount_cap.py` (the revision-inflation trap: body vs journal cap) and `check_paren_spans.py` (em-dash→paren conversions that wrap a whole sentence). Plus a `check_confounding_completeness.py` upgrade (DB-code↔prose alias map, SMD-from-mean±SD, exposure-defining-covariate exemption), a `check_cohort_arithmetic.py` `ANALYSIS_UNIT_UNDISCLOSED` check, a `check_scope_coherence.py` cross-sectional-yield lexicon, and a verify-refs corporate/collective-author render-abort fix.
|
|
329
329
|
- **Analysis & submission tooling** — `/analyze-stats` gains **mediation** and **interaction & effect-modification** guides; `/sync-submission` gains `assemble_supplement.py` (S{N} index↔file integrity) and a `/revise` body-word-count exit gate; `/render-pdf-doc` gains a `scan_glyph_coverage.py` xelatex silent-glyph-drop scan.
|
|
330
330
|
|
|
331
|
-
**v4.2** builds out the case-report capability end-to-end, grounded in real CC-BY case reports (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills /
|
|
331
|
+
**v4.2** builds out the case-report capability end-to-end, grounded in real CC-BY case reports (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills / 41 guidelines); journal profiles **68 → 73**:
|
|
332
332
|
|
|
333
333
|
- **Case-report + case-series writing** — `/write-paper` gains a CARE narrative + 150-word-abstract case-report exemplar, a **case-series** paper type (methods-light mini-cohort, all-cases summary table, counts-not-rates), and **adverse-event/pharmacovigilance** (Naranjo/WHO-UMC causality) and **diagnostic-pitfall/mimic** subtypes.
|
|
334
334
|
- **Radiology / imaging-led track** — a dedicated `exemplar_case_report_radiology.md` (per-modality technique→findings→impression, structured-reporting lexicons BI-RADS/LI-RADS/PI-RADS/TI-RADS/Lung-RADS/O-RADS, quantitative threshold honesty, an interventional-radiology procedure/complication subtype, DICOM de-identification) plus a `/make-figures` annotated multimodality imaging-panel exemplar.
|
|
@@ -452,7 +452,7 @@ ma-scout -> search-lit -> fulltext-retrieval -> design-study ──> write-proto
|
|
|
452
452
|
| **search-lit** | PubMed + Semantic Scholar + bioRxiv search with anti-hallucination citation verification. Token-efficient error handling -- CrossRef failures are silently batched, not repeated. BibTeX output tags each entry with `verified`/`verified_by`/`verified_on` fields so downstream skills can trust the citation provenance. |
|
|
453
453
|
| **verify-refs** | Pre-submission reference audit for `.md`, `.docx`, `.bib`, or `.tsv` inputs. Extracts references, verifies DOI/PMID via CrossRef/PubMed when available, and writes `qc/reference_audit.json` as the sole output — row-level status (OK / MISMATCH / UNVERIFIED / FABRICATED) lives inside the JSON `records[]` block. `/search-lit` produces candidate BibTeX; `/lit-sync` owns `manuscript/_src/refs.bib`. |
|
|
454
454
|
| **fulltext-retrieval** | Batch open-access PDF downloader. Unpaywall → PMC → OpenAlex → CrossRef pipeline. OA-only -- no paywall bypass. Input: DOI list or TSV. Optional PDF→Markdown conversion via [pymupdf4llm](https://pymupdf.readthedocs.io/en/latest/pymupdf4llm/) for token-efficient LLM analysis of academic papers. |
|
|
455
|
-
| **check-reporting** | Manuscript compliance audit against
|
|
455
|
+
| **check-reporting** | Manuscript compliance audit against 41 reporting guidelines and risk of bias tools (STROBE, STROBE-MR, RECORD, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, CHEERS 2022, CROSS, PRISMA, PRISMA-DTA, PRISMA-P, MOOSE, ARRIVE, CONSORT, CONSORT-AI, CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, SQUIRE 2.0, CLEAR, GRRAS, MI-CLEAR-LLM, SWiM, AMSTAR 2, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Machine-readable JSON summary with `compliance_pct` and `fixable_by_ai` flags for automated pipeline integration. |
|
|
456
456
|
| **analyze-stats** | Statistical analysis code generation (Python/R) for diagnostic accuracy, DTA meta-analysis (bivariate/HSROC), inter-rater agreement, survival analysis, demographics tables, regression (logistic/linear), propensity score (matching/IPTW/overlap weighting), and repeated measures (RM ANOVA/GEE/mixed models). Calibration mandatory for prediction models. |
|
|
457
457
|
| **meta-analysis** | Full systematic review and meta-analysis pipeline (8 phases). DTA (bivariate/HSROC) and intervention meta-analysis. Protocol to submission-ready manuscript with PRISMA-DTA compliance. |
|
|
458
458
|
| **make-figures** | Publication-ready figures and visual abstracts: ROC curves, forest plots, PRISMA/CONSORT/STARD flow diagrams, Kaplan-Meier curves, Bland-Altman plots, confusion matrices, and journal-specific visual/graphical abstracts (python-pptx template-based). Communication-first design principles (Nat Hum Behav 2026 — key message, audience, cognitive load, figure-vs-table decision) and five flow-diagram production lessons (official-template fidelity, VML fallback PDF export, docx XML escape, sequential placeholder mapping, version freeze); critic rubric Section G adds 5 communication-first checks. `--study-type` auto-generates the full required figure set; structured `_figure_manifest.md` output for downstream pipeline consumption; D2 enforced as default for flow diagrams. |
|
|
@@ -621,8 +621,8 @@ Projects declare their source-of-truth layout in `SSOT.yaml`, and a `qc/migratio
|
|
|
621
621
|
### Meta-Analysis Failure Modes
|
|
622
622
|
`/meta-analysis` ships empirical failure-mode references (data integrity, review orchestration, submission package drift, post-submission release ops) with four automation hooks: `scripts/prisma_5way_consistency.py` (DI-6 PRISMA number consistency), `scripts/extraction_consensus_log_init.py` (DI-1 dual-extraction scaffold), `scripts/tag_cleanup_gate.sh` (DI-8 placeholder tag gate), and `scripts/verify_package_integrity.py` (SPD SHA-256 manifest for submission bundles).
|
|
623
623
|
|
|
624
|
-
###
|
|
625
|
-
`check-reporting` includes bundled checklists for
|
|
624
|
+
### 41 Reporting Guidelines & RoB Tools Built-in
|
|
625
|
+
`check-reporting` includes bundled checklists for 41 guidelines and risk-of-bias tools: STROBE, STROBE-MR, RECORD, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, CHEERS 2022, CROSS, PRISMA 2020, PRISMA-DTA, PRISMA-P, MOOSE, ARRIVE, CONSORT, CONSORT-AI, CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, SQUIRE 2.0, CLEAR, GRRAS, MI-CLEAR-LLM, SWiM, AMSTAR 2, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA. Includes Results/Discussion section boundary checks and machine-readable JSON summary for pipeline integration.
|
|
626
626
|
|
|
627
627
|
### Publication-Ready Output
|
|
628
628
|
`analyze-stats` generates reproducible Python/R code for 13 analysis types -- including regression, propensity score, and repeated measures -- with mandatory calibration for prediction models. `make-figures` produces journal-specification figures (300 DPI, colorblind-safe palettes, proper dimensions), visual/graphical abstracts, and a tool selection guide (D2 for flow diagrams, matplotlib for data plots). `--study-type` auto-generates the complete figure set for each study design.
|
|
@@ -538,8 +538,8 @@
|
|
|
538
538
|
},
|
|
539
539
|
{
|
|
540
540
|
"path": "skills/check-reporting/SKILL.md",
|
|
541
|
-
"size":
|
|
542
|
-
"sha256": "
|
|
541
|
+
"size": 37775,
|
|
542
|
+
"sha256": "dd025dea034a0ba6ea900d501d182513a2c1f3cbe95a32efbcb701224962c1d6"
|
|
543
543
|
},
|
|
544
544
|
{
|
|
545
545
|
"path": "skills/check-reporting/references/LICENSES.md",
|
|
@@ -601,6 +601,11 @@
|
|
|
601
601
|
"size": 5535,
|
|
602
602
|
"sha256": "d5c677b029a9f7a8ecff6cd2cb9d532149d29de6026b89bc7d4d4d8e1571be94"
|
|
603
603
|
},
|
|
604
|
+
{
|
|
605
|
+
"path": "skills/check-reporting/references/checklists/CROSS.md",
|
|
606
|
+
"size": 7242,
|
|
607
|
+
"sha256": "2db92417cd7d5ffed623d8e53ac5b06e8f4d57bbc07a3c8ce900851c3e6aaca4"
|
|
608
|
+
},
|
|
604
609
|
{
|
|
605
610
|
"path": "skills/check-reporting/references/checklists/DECIDE_AI.md",
|
|
606
611
|
"size": 5321,
|
|
@@ -666,6 +671,11 @@
|
|
|
666
671
|
"size": 7571,
|
|
667
672
|
"sha256": "f3527f32851fbf4d446e7f9961b134d455c2fa19d7615ea65168ae0479fb6685"
|
|
668
673
|
},
|
|
674
|
+
{
|
|
675
|
+
"path": "skills/check-reporting/references/checklists/RECORD.md",
|
|
676
|
+
"size": 6530,
|
|
677
|
+
"sha256": "186000d9cd72e025ce483995735c05400450d9356f89f8f95b855d5815e925fa"
|
|
678
|
+
},
|
|
669
679
|
{
|
|
670
680
|
"path": "skills/check-reporting/references/checklists/ROBINS_E.md",
|
|
671
681
|
"size": 8206,
|
|
@@ -773,8 +783,8 @@
|
|
|
773
783
|
},
|
|
774
784
|
{
|
|
775
785
|
"path": "skills/check-reporting/scripts/check_checklist_exists.py",
|
|
776
|
-
"size":
|
|
777
|
-
"sha256": "
|
|
786
|
+
"size": 6722,
|
|
787
|
+
"sha256": "eaa77ce04947716984b049a85d9f94e04e4e03247e65bf5b1e08b5d591dac75a"
|
|
778
788
|
},
|
|
779
789
|
{
|
|
780
790
|
"path": "skills/check-reporting/scripts/check_checklist_version.py",
|
|
@@ -2014,7 +2024,7 @@
|
|
|
2014
2024
|
{
|
|
2015
2025
|
"path": "skills/make-figures/references/reporting_guideline_figure_map.md",
|
|
2016
2026
|
"size": 6862,
|
|
2017
|
-
"sha256": "
|
|
2027
|
+
"sha256": "1af49856cdd22f52c063a7bf2319835b5ce922074849855933ecc24a7f25ab2b"
|
|
2018
2028
|
},
|
|
2019
2029
|
{
|
|
2020
2030
|
"path": "skills/make-figures/references/visual_abstract_templates/european_radiology.pptx",
|
|
@@ -2804,7 +2814,7 @@
|
|
|
2804
2814
|
{
|
|
2805
2815
|
"path": "skills/orchestrate/SKILL.md",
|
|
2806
2816
|
"size": 35203,
|
|
2807
|
-
"sha256": "
|
|
2817
|
+
"sha256": "01e2e9c98985730f2df623a634233eefb0deeab155681abb12a6b0957ca0c0bc"
|
|
2808
2818
|
},
|
|
2809
2819
|
{
|
|
2810
2820
|
"path": "skills/orchestrate/references/dialogue_nodes.md",
|
|
@@ -2828,8 +2838,8 @@
|
|
|
2828
2838
|
},
|
|
2829
2839
|
{
|
|
2830
2840
|
"path": "skills/peer-review/SKILL.md",
|
|
2831
|
-
"size":
|
|
2832
|
-
"sha256": "
|
|
2841
|
+
"size": 64095,
|
|
2842
|
+
"sha256": "0848971688fb2d7bd115adf2b4fa0a5129cc979703fa0da2a680c9883644ac83"
|
|
2833
2843
|
},
|
|
2834
2844
|
{
|
|
2835
2845
|
"path": "skills/peer-review/references/aczel_2021_reviewer2_patterns.md",
|
|
@@ -2916,11 +2926,21 @@
|
|
|
2916
2926
|
"size": 9293,
|
|
2917
2927
|
"sha256": "23bdfcd9899f02792f4afcd28f1301822c6a5da8fc9223353b3354ca010913cf"
|
|
2918
2928
|
},
|
|
2929
|
+
{
|
|
2930
|
+
"path": "skills/peer-review/references/domain-probes/record_routinely_collected_data.md",
|
|
2931
|
+
"size": 7574,
|
|
2932
|
+
"sha256": "04174053da808ee99f72f21d41225d76e856e94aff5088dffd8040ec40a4aa4e"
|
|
2933
|
+
},
|
|
2919
2934
|
{
|
|
2920
2935
|
"path": "skills/peer-review/references/domain-probes/sr_ma.md",
|
|
2921
2936
|
"size": 17239,
|
|
2922
2937
|
"sha256": "486d569f559f16d62b882f7256ccfe7443e3e2b4cc72a407a5cd046162a98b25"
|
|
2923
2938
|
},
|
|
2939
|
+
{
|
|
2940
|
+
"path": "skills/peer-review/references/domain-probes/survey_research.md",
|
|
2941
|
+
"size": 7241,
|
|
2942
|
+
"sha256": "597e317780afdcc01bd929a7b902df678ff2e16ea43cf2cbfe5e5cace6495d78"
|
|
2943
|
+
},
|
|
2924
2944
|
{
|
|
2925
2945
|
"path": "skills/peer-review/references/domain-probes/survival_prognostic.md",
|
|
2926
2946
|
"size": 13765,
|
|
@@ -3373,8 +3393,8 @@
|
|
|
3373
3393
|
},
|
|
3374
3394
|
{
|
|
3375
3395
|
"path": "skills/self-review/SKILL.md",
|
|
3376
|
-
"size":
|
|
3377
|
-
"sha256": "
|
|
3396
|
+
"size": 95312,
|
|
3397
|
+
"sha256": "ff45218dcd5cc5233787b95c8f6ce9168dc2ff27db1ce98d4b121884e13ae526"
|
|
3378
3398
|
},
|
|
3379
3399
|
{
|
|
3380
3400
|
"path": "skills/self-review/references/domain-probes/ai_overclaiming.md",
|
|
@@ -3456,11 +3476,21 @@
|
|
|
3456
3476
|
"size": 9293,
|
|
3457
3477
|
"sha256": "23bdfcd9899f02792f4afcd28f1301822c6a5da8fc9223353b3354ca010913cf"
|
|
3458
3478
|
},
|
|
3479
|
+
{
|
|
3480
|
+
"path": "skills/self-review/references/domain-probes/record_routinely_collected_data.md",
|
|
3481
|
+
"size": 7574,
|
|
3482
|
+
"sha256": "04174053da808ee99f72f21d41225d76e856e94aff5088dffd8040ec40a4aa4e"
|
|
3483
|
+
},
|
|
3459
3484
|
{
|
|
3460
3485
|
"path": "skills/self-review/references/domain-probes/sr_ma.md",
|
|
3461
3486
|
"size": 17239,
|
|
3462
3487
|
"sha256": "486d569f559f16d62b882f7256ccfe7443e3e2b4cc72a407a5cd046162a98b25"
|
|
3463
3488
|
},
|
|
3489
|
+
{
|
|
3490
|
+
"path": "skills/self-review/references/domain-probes/survey_research.md",
|
|
3491
|
+
"size": 7241,
|
|
3492
|
+
"sha256": "597e317780afdcc01bd929a7b902df678ff2e16ea43cf2cbfe5e5cace6495d78"
|
|
3493
|
+
},
|
|
3464
3494
|
{
|
|
3465
3495
|
"path": "skills/self-review/references/domain-probes/survival_prognostic.md",
|
|
3466
3496
|
"size": 13765,
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "medsci-skills",
|
|
3
|
-
"version": "5.
|
|
3
|
+
"version": "5.5.0",
|
|
4
4
|
"description": "MedSci Skills — a medical/scientific research skill suite for AI coding agents (Claude Code, Codex, Cursor, Copilot). The npm package is a terminal-friendly installer shortcut; the canonical distribution remains the GitHub repository and the Claude Code plugin marketplace.",
|
|
5
5
|
"license": "SEE LICENSE IN LICENSE",
|
|
6
6
|
"homepage": "https://github.com/Aperivue/medsci-skills#readme",
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: check-reporting
|
|
3
|
-
description: Check manuscript compliance with medical research reporting guidelines. Supports
|
|
4
|
-
triggers: checklist, reporting guideline, STROBE, STROBE-MR, Mendelian randomization, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD-LLM, PGS-RS, PRS-RS, polygenic risk score, polygenic score, PRISMA, PRISMA-DTA, PRISMA-P, ARRIVE, CARE, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SPIRIT, SPIRIT-AI, QUADAS, QUADAS-C, RoB, ROBINS, ROBINS-E, ROBIS, ROB-ME, PROBAST, NOS, COSMIN, AMSTAR, SWiM, CHEERS, economic evaluation, cost-effectiveness, cost-utility, QALY, ICER, risk of bias, compliance check, LLM accuracy, large language model, clinical deployment
|
|
3
|
+
description: Check manuscript compliance with medical research reporting guidelines. Supports 41 guidelines including STROBE, STROBE-MR, RECORD, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, ARRIVE, PRISMA, PRISMA-DTA, PRISMA-P, CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SQUIRE 2.0, CLEAR, MOOSE, GRRAS, SWiM, AMSTAR 2, CHEERS 2022, CROSS (survey studies), and risk of bias tools (QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Generates item-by-item assessment with PRESENT/MISSING/PARTIAL status.
|
|
4
|
+
triggers: checklist, reporting guideline, STROBE, STROBE-MR, Mendelian randomization, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD-LLM, PGS-RS, PRS-RS, polygenic risk score, polygenic score, PRISMA, PRISMA-DTA, PRISMA-P, ARRIVE, CARE, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SPIRIT, SPIRIT-AI, QUADAS, QUADAS-C, RoB, ROBINS, ROBINS-E, ROBIS, ROB-ME, PROBAST, NOS, COSMIN, AMSTAR, SWiM, CHEERS, economic evaluation, cost-effectiveness, cost-utility, QALY, ICER, RECORD, RECORD-PE, routinely-collected data, registry, claims, electronic health records, EHR, real-world data, CROSS, CHERRIES, survey, questionnaire, KAP, e-survey, response rate, risk of bias, compliance check, LLM accuracy, large language model, clinical deployment
|
|
5
5
|
tools: Read, Write, Edit, Bash, Grep, Glob
|
|
6
6
|
model: inherit
|
|
7
7
|
---
|
|
@@ -30,6 +30,8 @@ compliance report suitable for journal submission.
|
|
|
30
30
|
- `TRIPOD_LLM.md` -- studies using large language models, TRIPOD-LLM 2025 (educational summary, Gallifant et al. Nat Med 2025)
|
|
31
31
|
- `PGS_RS.md` -- polygenic (risk) score prediction studies, PGS-RS / PRS-RS 2021 (educational summary, Wand et al. Nature 2021)
|
|
32
32
|
- `CHEERS_2022.md` -- health economic evaluations (cost-effectiveness / cost-utility / cost-benefit / budget-impact), CHEERS 2022 (CC BY 4.0, Husereau et al. BMJ 2022)
|
|
33
|
+
- `RECORD.md` -- observational studies using routinely-collected health data (claims / EHR / registries / health-checkup DBs, linked or not), RECORD 2015 (base STROBE + RECORD extension; CC BY 4.0, Benchimol et al. PLoS Med 2015; RECORD-PE for drug studies)
|
|
34
|
+
- `CROSS.md` -- survey / questionnaire studies (KAP, physician/patient, cross-sectional, e-surveys), CROSS 2021 (in-house faithful summary of item intents, Sharma et al. JGIM 2021) + CHERRIES (CC BY, Eysenbach JMIR 2004) for internet surveys
|
|
33
35
|
- `PRISMA_2020.md` -- systematic reviews (CC BY)
|
|
34
36
|
- `ARRIVE_2.md` -- animal studies (CC0)
|
|
35
37
|
- `PRISMA_DTA.md` -- DTA systematic reviews (CC BY, McInnes et al. JAMA 2018)
|
|
@@ -90,6 +92,8 @@ user specification.
|
|
|
90
92
|
| Observational study | STROBE | -- |
|
|
91
93
|
| Mendelian randomization study | STROBE-MR (base STROBE + MR extension) | -- |
|
|
92
94
|
| Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact) | CHEERS 2022 | -- |
|
|
95
|
+
| Observational study using routinely-collected data (claims / EHR / registry / health-checkup DB) | RECORD (base STROBE + RECORD extension; RECORD-PE for drug studies) | -- |
|
|
96
|
+
| Survey / questionnaire study (KAP, physician/patient, cross-sectional, e-survey) | CROSS (+ CHERRIES for internet surveys) | -- |
|
|
93
97
|
| Randomized controlled trial | CONSORT 2025 | CONSORT-AI |
|
|
94
98
|
| Diagnostic accuracy study | STARD 2015 | STARD-AI |
|
|
95
99
|
| Prediction model (development/validation) | TRIPOD | TRIPOD+AI |
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# CROSS Checklist (survey studies)
|
|
2
|
+
|
|
3
|
+
**Consensus-Based Checklist for Reporting of Survey Studies**
|
|
4
|
+
Version: CROSS 2021 (40 reportable elements across ~7 sections). For internet/electronic surveys, pair with **CHERRIES** (Checklist for Reporting Results of Internet E-Surveys).
|
|
5
|
+
Sources: Sharma A, Minh Duc NT, Luu Lam Thang T, et al. *J Gen Intern Med* 2021;36:3179–3187 (the CROSS statement; DOI 10.1007/s11606-021-06737-1). Eysenbach G. *J Med Internet Res* 2004;6(3):e34 (CHERRIES; CC BY). EQUATOR Network.
|
|
6
|
+
|
|
7
|
+
Apply when the manuscript is a **self-report survey / questionnaire study** — knowledge-attitudes-practices (KAP), physician or patient surveys, cross-sectional questionnaires, and web/e-surveys. CROSS covers the reportable elements of design, sampling, instrument development, administration, and analysis; for an **internet/e-survey**, the CHERRIES items (open-vs-closed sample, denominator/completion definition, voluntariness/incentives, duplicate-submission control) also apply. For the design/validity review of the same study, pair with the SV1–SV8 domain probes in `peer-review` / `self-review` `references/domain-probes/survey_research.md`; for scale reliability, see `analyze-stats` `Survey/Likert` + `survey_weighted.md`.
|
|
8
|
+
|
|
9
|
+
> Licensing note: the CROSS statement is © Society of General Internal Medicine (not a Creative Commons licence). The items below are an **in-house, faithful summary of the reportable elements (facts/intents, paraphrased — not the verbatim CROSS wording)** for item-by-item assessment; consult the published CROSS article for exact item text. The CHERRIES e-survey items are grounded in the CC BY JMIR source.
|
|
10
|
+
|
|
11
|
+
## Reportable elements (grouped by section)
|
|
12
|
+
|
|
13
|
+
### Title and Abstract
|
|
14
|
+
| # | Element | What to check is reported |
|
|
15
|
+
|---|---------|---------------------------|
|
|
16
|
+
| 1 | Title/abstract | The study is identified as a survey/questionnaire study; the abstract summarises objectives, design, sample, response, and key findings. |
|
|
17
|
+
|
|
18
|
+
### Introduction
|
|
19
|
+
| # | Element | What to check is reported |
|
|
20
|
+
|---|---------|---------------------------|
|
|
21
|
+
| 2 | Background & objectives | Rationale for the survey and explicit objectives / research questions or hypotheses. |
|
|
22
|
+
|
|
23
|
+
### Methods — Study design
|
|
24
|
+
| # | Element | What to check is reported |
|
|
25
|
+
|---|---------|---------------------------|
|
|
26
|
+
| 3 | Design & timeframe | Survey design (cross-sectional, repeated/longitudinal), and the dates/period of data collection. |
|
|
27
|
+
| 4 | Data sources/setting | Setting and the source/channel through which respondents were reached. |
|
|
28
|
+
|
|
29
|
+
### Methods — Sample
|
|
30
|
+
| # | Element | What to check is reported |
|
|
31
|
+
|---|---------|---------------------------|
|
|
32
|
+
| 5 | Target population & sampling frame | The target population and the **sampling frame** used to reach it, with comment on how well the frame covers the population (coverage). |
|
|
33
|
+
| 6 | Sample selection | The **sampling method** (probability vs non-probability/convenience) and selection procedure; eligibility criteria. |
|
|
34
|
+
| 7 | Sample size | An a-priori **sample-size or precision justification** (not a post-hoc rationalisation of whoever responded). |
|
|
35
|
+
|
|
36
|
+
### Methods — Survey administration
|
|
37
|
+
| # | Element | What to check is reported |
|
|
38
|
+
|---|---------|---------------------------|
|
|
39
|
+
| 8 | Administration mode | Mode(s) of administration (web, email, postal, telephone, in-person) and the implications for coverage/selection. |
|
|
40
|
+
| 9 | E-survey specifics (CHERRIES) | For internet surveys: open vs closed (invited) survey; how the **denominator and completion/response** were defined; voluntariness and any incentive; **duplicate-submission control** (IP/cookie/log-in); use of adaptive/mandatory questions and completeness. |
|
|
41
|
+
|
|
42
|
+
### Methods — Study preparation (instrument)
|
|
43
|
+
| # | Element | What to check is reported |
|
|
44
|
+
|---|---------|---------------------------|
|
|
45
|
+
| 10 | Instrument development | Whether the questionnaire was newly developed or adopted/adapted from a validated instrument (with citation). |
|
|
46
|
+
| 11 | Pre-testing/piloting | Pilot testing / cognitive pre-testing of the instrument before fielding. |
|
|
47
|
+
| 12 | Validity & reliability | Evidence of **validity** (content/construct) and **reliability** (e.g. Cronbach's α, test–retest) for multi-item scales. |
|
|
48
|
+
|
|
49
|
+
### Methods — Ethics & analysis
|
|
50
|
+
| # | Element | What to check is reported |
|
|
51
|
+
|---|---------|---------------------------|
|
|
52
|
+
| 13 | Ethics & data protection | Informed consent, ethics-committee approval/exemption, and respondent data protection/anonymity. |
|
|
53
|
+
| 14 | Statistical methods | Analysis methods, handling of **missing/incomplete responses**, and any **weighting / post-stratification** for representativeness. |
|
|
54
|
+
|
|
55
|
+
### Results
|
|
56
|
+
| # | Element | What to check is reported |
|
|
57
|
+
|---|---------|---------------------------|
|
|
58
|
+
| 15 | Response & representativeness | **Response rate with a defined denominator** (e.g. an AAPOR/CASRO definition), respondent flow, and an assessment of **non-response/representativeness** (responders vs non-responders or vs the population). |
|
|
59
|
+
| 16 | Descriptive results | Respondent characteristics and results with appropriate denominators (per-item N where it varies); uncertainty (CIs) for key estimates. |
|
|
60
|
+
|
|
61
|
+
### Discussion
|
|
62
|
+
| # | Element | What to check is reported |
|
|
63
|
+
|---|---------|---------------------------|
|
|
64
|
+
| 17 | Limitations | Coverage, sampling, non-response, and self-report/social-desirability biases; their likely direction. |
|
|
65
|
+
| 18 | Interpretation & generalisability | Conclusions matched to the sampled population — no over-generalisation from a convenience/low-response/single-setting sample. |
|
|
66
|
+
|
|
67
|
+
### Other
|
|
68
|
+
| # | Element | What to check is reported |
|
|
69
|
+
|---|---------|---------------------------|
|
|
70
|
+
| 19 | Funding, COI, availability | Funding source and role, conflicts of interest, and availability of the instrument/data where possible. |
|
|
71
|
+
|
|
72
|
+
---
|
|
73
|
+
|
|
74
|
+
## Notes for Assessors
|
|
75
|
+
|
|
76
|
+
- The **highest-yield** elements (where surveys most often fail review): **5/6** (a named sampling frame and a probability-vs-convenience statement — a self-selected web panel generalised to "clinicians"/"patients" is the commonest over-reach), **15** (a **response rate with a defined denominator** plus a **non-response/representativeness** assessment — "N responded" with no denominator and no non-response analysis is non-compliant), **10–12** (instrument **development, piloting, and validity/reliability** — a novel unvalidated questionnaire carrying the headline), **9** (for e-surveys, CHERRIES denominator/duplicate-control reporting), and **14** (weighting for a skewed sample; consistent denominators).
|
|
77
|
+
- A survey reporting only a raw count of respondents, with no sampling frame, no defined-denominator response rate, no non-response assessment, and an unvalidated instrument, is non-compliant on elements 5/6/10–12/15 and its population-level claims should be downgraded to "among respondents."
|
|
78
|
+
- This is an **in-house faithful summary of the CROSS reportable elements (paraphrased intents, not verbatim)** complemented by the CHERRIES e-survey items; map the manuscript's content to the elements rather than to exact CROSS item numbering, and verify against the published CROSS (Sharma et al. *JGIM* 2021) and CHERRIES (Eysenbach *JMIR* 2004, CC BY) sources. Verified 2026-06-29.
|
|
@@ -0,0 +1,66 @@
|
|
|
1
|
+
# RECORD Checklist
|
|
2
|
+
|
|
3
|
+
**REporting of studies Conducted using Observational Routinely-collected health Data**
|
|
4
|
+
Version: RECORD 2015 (13 items; a STROBE extension). The pharmacoepidemiology extension is **RECORD-PE** (Langan et al. *BMJ* 2018).
|
|
5
|
+
Source: Benchimol EI, Smeeth L, Guttmann A, et al. *PLoS Medicine* 2015;12(10):e1001885 (the RECORD statement). CC BY 4.0. https://www.record-statement.org/ · EQUATOR Network.
|
|
6
|
+
|
|
7
|
+
Apply when the manuscript is an **observational study conducted using routinely-collected health data** — administrative claims, electronic health records (EHR), disease/population registries, health-administrative or health-checkup databases, or linked versions of these — i.e. data **not collected for the purpose of the specific study**. RECORD extends the base **STROBE** items with reporting specific to secondary-use data: database identity, the codes/algorithms used to define the population and the variables, data linkage and its quality, and the limitations of analysing data collected for another purpose. For a drug safety/effectiveness study in such data, also apply **RECORD-PE**. For the design/validity review of the same study, pair with the RD1–RD8 domain probes in `peer-review` / `self-review` `references/domain-probes/record_routinely_collected_data.md`, and with the observational-confounding probes (`observational_confounding.md`).
|
|
8
|
+
|
|
9
|
+
## Checklist Items (13 items, extending STROBE)
|
|
10
|
+
|
|
11
|
+
### Title and Abstract (STROBE item 1)
|
|
12
|
+
|
|
13
|
+
| # | Item | Description |
|
|
14
|
+
|---|------|-------------|
|
|
15
|
+
| 1.1 | Data type | The type of data used should be specified in the title or abstract. When possible, the name(s) of the database(s) used should be stated. |
|
|
16
|
+
| 1.2 | Geography and timeframe | The geographic region and timeframe within which the study took place should be reported in the title or abstract. |
|
|
17
|
+
| 1.3 | Linkage | If linkage between databases was conducted for the study, this should be clearly stated in the title or abstract. |
|
|
18
|
+
|
|
19
|
+
### Methods — Setting / Participants (STROBE item 6)
|
|
20
|
+
|
|
21
|
+
| # | Item | Description |
|
|
22
|
+
|---|------|-------------|
|
|
23
|
+
| 6.1 | Population selection | The methods of study population selection (such as codes or algorithms used to identify subjects) should be listed in detail. If this is not possible, an explanation should be provided. |
|
|
24
|
+
| 6.2 | Validation of codes | Any validation studies of the codes or algorithms used to select the population should be referenced. If validation was conducted for this study and not published elsewhere, detailed methods and results should be provided. |
|
|
25
|
+
| 6.3 | Linkage diagram | If the study involved linkage of databases, consider use of a flow diagram or other graphical display to demonstrate the data linkage process, including the number of individuals with linked data at each stage. |
|
|
26
|
+
|
|
27
|
+
### Methods — Variables (STROBE item 7)
|
|
28
|
+
|
|
29
|
+
| # | Item | Description |
|
|
30
|
+
|---|------|-------------|
|
|
31
|
+
| 7.1 | Codes for variables | A complete list of codes and algorithms used to classify exposures, outcomes, confounders, and effect modifiers should be provided. If these cannot be reported, an explanation should be provided. |
|
|
32
|
+
|
|
33
|
+
### Methods — Data access and cleaning (STROBE item 12)
|
|
34
|
+
|
|
35
|
+
| # | Item | Description |
|
|
36
|
+
|---|------|-------------|
|
|
37
|
+
| 12.1 | Data access | Authors should describe the extent to which the investigators had access to the database population used to create the study population. |
|
|
38
|
+
| 12.2 | Data cleaning | Authors should provide information on the data cleaning methods used in the study. |
|
|
39
|
+
| 12.3 | Linkage methods | State whether the study included person-level, institutional-level, or other data linkage across two or more databases. The methods of linkage and methods of linkage quality evaluation should be provided. |
|
|
40
|
+
|
|
41
|
+
### Results — Participants (STROBE item 13)
|
|
42
|
+
|
|
43
|
+
| # | Item | Description |
|
|
44
|
+
|---|------|-------------|
|
|
45
|
+
| 13.1 | Selection of included persons | Describe in detail the selection of the persons included in the study (i.e. study population selection) including filtering based on data quality, data availability and linkage. The selection of included persons can be described in the text and/or by means of the study flow diagram. |
|
|
46
|
+
|
|
47
|
+
### Discussion — Limitations (STROBE item 19)
|
|
48
|
+
|
|
49
|
+
| # | Item | Description |
|
|
50
|
+
|---|------|-------------|
|
|
51
|
+
| 19.1 | Secondary-data limitations | Discuss the implications of using data that were not created or collected to answer the specific research question(s). Include discussion of misclassification bias, unmeasured confounding, missing data, and changing eligibility over time, as they pertain to the study being reported. |
|
|
52
|
+
|
|
53
|
+
### Other Information — Data access / cleaning (STROBE item 22)
|
|
54
|
+
|
|
55
|
+
| # | Item | Description |
|
|
56
|
+
|---|------|-------------|
|
|
57
|
+
| 22.1 | Supplemental access | Authors should provide information on how to access any supplemental information such as the study protocol, raw data, or programming code. |
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## Notes for Assessors
|
|
62
|
+
|
|
63
|
+
- RECORD is an **extension of STROBE**; for the non-RECORD-specific items the base `STROBE.md` guidance also applies. Report both the base instrument and the extension when describing methods (do not cite RECORD as if it replaced STROBE).
|
|
64
|
+
- The **highest-yield** items are **6.1 / 7.1** (the actual code lists / phenotype algorithms used to define the population, exposures, outcomes, and confounders — the single most common omission; "we identified diabetes from the database" with no codes is non-compliant), **6.2** (whether those algorithms were *validated*, and where), **12.3 / 6.3** (linkage method and linkage-quality evaluation, with a person-flow at each linkage stage), **13.1** (a participant-selection flow that includes data-quality/availability/linkage filtering — not only clinical eligibility), and **19.1** (the limitations specific to secondary-use data: misclassification from codes, unmeasured confounding, informative missingness, and eligibility drift over time).
|
|
65
|
+
- For a **drug safety/effectiveness** study in routinely-collected data, also apply **RECORD-PE** (Langan et al. *BMJ* 2018;363:k3532), which adds items on exposure definition (drug codes, dose, duration, exposure windows), the comparator and new-user/active-comparator design, and immortal-time/protopathic bias.
|
|
66
|
+
- This checklist was authored as a faithful summary of the RECORD statement (Benchimol EI, et al. *PLoS Med* 2015;12(10):e1001885, **CC BY 4.0**) for item-by-item assessment; verify against the published statement and its explanation-and-elaboration document for full item wording. Verified 2026-06-29.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Reporting Guideline → Figure Requirements Map
|
|
2
2
|
|
|
3
3
|
> **Bridge**: this file connects `/make-figures` to `/check-reporting`
|
|
4
|
-
> (
|
|
4
|
+
> (41 reporting guidelines). Each row tells you which figures the guideline
|
|
5
5
|
> **mandates** and how this skill currently supports them. Use during
|
|
6
6
|
> Step 1 (Specify) once the study type is known.
|
|
7
7
|
|
|
@@ -95,7 +95,7 @@ flags them:
|
|
|
95
95
|
|
|
96
96
|
## Cross-references
|
|
97
97
|
|
|
98
|
-
- `/check-reporting` skill — supports all
|
|
98
|
+
- `/check-reporting` skill — supports all 41 guidelines, item-level audit
|
|
99
99
|
- `flow_diagram_lessons.md` — production lessons that apply across all flows
|
|
100
100
|
- `pipeline_concepts_medical_ai.md` — DICOM / annotation / federated /
|
|
101
101
|
architecture diagram conventions
|
|
@@ -48,7 +48,7 @@ You do NOT do the work yourself. You classify, plan, and delegate.
|
|
|
48
48
|
| **meta-analysis** | Systematic review | Full MA pipeline: protocol, search, screening, extraction, synthesis, PRISMA-DTA |
|
|
49
49
|
| **write-paper** | Writing | IMRAD manuscript drafting (8-phase pipeline), any section writing |
|
|
50
50
|
| **self-review** | Quality | Pre-submission self-check with domain probes (Survival / SR-MA / Radiomics / Narrative); optional `--panel` for a high-stakes final QC pass |
|
|
51
|
-
| **check-reporting** | Compliance | Audit against
|
|
51
|
+
| **check-reporting** | Compliance | Audit against 41 reporting guidelines and risk-of-bias tools |
|
|
52
52
|
| **revise** | Revision | Parse reviewer comments, generate point-by-point response, track changes |
|
|
53
53
|
| **grant-builder** | Funding | Structure grant proposals: significance, innovation, approach, milestones |
|
|
54
54
|
| **present-paper** | Presentation | Prepare academic talks: analyze paper, draft scripts, inject slide notes, Q&A prep |
|
|
@@ -274,6 +274,18 @@ Apply this 8-probe checklist (HE1–HE8) **only when the manuscript is a health
|
|
|
274
274
|
|
|
275
275
|
**Probe detail (HE1–HE8), with output templates and the leads-vs-findings discipline:** `${CLAUDE_SKILL_DIR}/references/domain-probes/health_economic_evaluation.md`. Load it and apply each probe when the trigger fires. In this skill, map each probe finding to a Major / Minor comment; a missing/obsolete comparator or perspective inconsistent with the costs counted (HE1), a time horizon truncated below the point where costs and effects diverge or asymmetric/absent discounting (HE2), an unjustified/unvalidated model structure (HE5), a point-estimate ICER with no probabilistic sensitivity analysis / CEAC (HE6), or a "cost-effective" claim with no stated willingness-to-pay threshold and mishandled dominance (HE7) are design-level — surface them in the Confidential Comments to the Editor and place the strongest as the Major #1 candidate. A weak effectiveness source / unstated utility instrument (HE3), perspective-inconsistent costs or a missing price year (HE4), and an undisclosed industry-funder role on a threshold-hugging result (HE8) are validity/framing-level.
|
|
276
276
|
|
|
277
|
+
### Phase 2Q: Routinely-Collected-Data (RWD) Extension
|
|
278
|
+
|
|
279
|
+
Apply this 8-probe checklist (RD1–RD8) **only when the manuscript is an observational study conducted using routinely-collected health data** — administrative claims, electronic health records (EHR), disease/population registries, or health-administrative / health-checkup databases, linked or not. These probes complement (do not replace) the generic Phase 2 checklist, the STROBE + **RECORD** reporting items (**RECORD-PE** for drug studies), and the observational-confounding probes (`observational_confounding.md`). They target what secondary-use data add: whether the database can observe the question, whether phenotype code-lists and linkage are evidenced rather than asserted, and whether data-collected-for-another-purpose limitations are confronted.
|
|
280
|
+
|
|
281
|
+
**Probe detail (RD1–RD8), with output templates and the leads-vs-findings discipline:** `${CLAUDE_SKILL_DIR}/references/domain-probes/record_routinely_collected_data.md`. Load it and apply each probe when the trigger fires. In this skill, map each probe finding to a Major / Minor comment; missing phenotype code-lists / unvalidated algorithms for the population, exposure or outcome (RD2), undisclosed linkage method or linkage-quality evaluation (RD3), a source→analytic selection with no data-quality/availability/linkage flow (RD4), naive complete-case on informatively-missing fields (RD6), or an RWD drug-effect design exposed to immortal-time / prevalent-user bias with no mitigation (RD7) are design-level — surface them in the Confidential Comments to the Editor and place the strongest as the Major #1 candidate. A database that structurally cannot capture the exposure/outcome (RD1), unquantified coding misclassification (RD5), and unacknowledged coding/eligibility drift or no code/protocol availability (RD8) are validity/framing-level. Run the adjustment/collider/analysis-unit machinery via `observational_confounding.md`.
|
|
282
|
+
|
|
283
|
+
### Phase 2R: Survey / Questionnaire Study Extension
|
|
284
|
+
|
|
285
|
+
Apply this 8-probe checklist (SV1–SV8) **only when the manuscript is a self-report survey / questionnaire study** — KAP, physician/patient surveys, cross-sectional questionnaires, or web/e-surveys. These probes complement (do not replace) the generic Phase 2 checklist, the **CROSS** reporting items (**CHERRIES** for internet surveys), and the scale-reliability guidance. They target whether the sample can support a population claim at all: representativeness, the response-rate denominator and non-response bias, and whether the instrument measures what it claims — the most common failure being generalisation from a self-selected convenience sample.
|
|
286
|
+
|
|
287
|
+
**Probe detail (SV1–SV8), with output templates and the leads-vs-findings discipline:** `${CLAUDE_SKILL_DIR}/references/domain-probes/survey_research.md`. Load it and apply each probe when the trigger fires. In this skill, map each probe finding to a Major / Minor comment; a convenience/self-selected sample generalised to a population with no representativeness assessment (SV1), a non-probability sample presented as representative (SV2), a response rate with no defined denominator or no non-response analysis behind a population estimate (SV3), or a novel unvalidated/un-piloted instrument carrying the headline (SV4) are design-level — surface them in the Confidential Comments to the Editor and place the strongest as the Major #1 candidate. Missing CHERRIES e-survey reporting (SV5), biased question design / unavailable instrument (SV6), unweighted estimates from a skewed sample or shifting denominators (SV7), and over-generalisation or missing ethics/consent (SV8) are validity/framing-level. For multi-item-scale reliability (incl. the reverse-coded-item α trap), pair with the analyze-stats Survey/Likert guidance.
|
|
288
|
+
|
|
277
289
|
### Phase 3: Draft Review
|
|
278
290
|
|
|
279
291
|
Before writing comments, skim the relevant model in `references/exemplar_reviews/` for the
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
<!-- Domain probe module — shared, vendored BYTE-IDENTICAL by /peer-review and /self-review.
|
|
2
|
+
Severity words below (MAJOR / MINOR / major / minor) denote finding severity, NOT a journal
|
|
3
|
+
recommendation. Each consuming skill maps findings to its own output:
|
|
4
|
+
- peer-review: Major / Minor comments + Confidential Comments to the Editor; a
|
|
5
|
+
code-list / linkage-quality / selection-flow / RWD-bias design flaw is placed as Major #1.
|
|
6
|
+
- self-review: Anticipated Major / Minor Comments (Fatal / Fixable) mapped to category letters.
|
|
7
|
+
Do NOT edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`. -->
|
|
8
|
+
|
|
9
|
+
# Routinely-collected-data study probes (RD1–RD8)
|
|
10
|
+
|
|
11
|
+
An 8-probe checklist for observational studies conducted using **routinely-collected health data** — administrative claims, electronic health records (EHR), disease/population registries, health-administrative and health-checkup databases, and linked versions of these (data **not collected for the study's purpose**). These probes complement (do not replace) the generic Phase 2 checklist, the **STROBE** + **RECORD** reporting items (and **RECORD-PE** for drug studies), and the observational-confounding probes (`observational_confounding.md`, which cover adjustment/collider/analysis-unit issues). They target what secondary-use data add: whether the database can even observe the question, whether the phenotype code-lists and linkage are evidenced rather than asserted, and whether the limitations endemic to data collected for another purpose are confronted. RD2 (phenotype code-lists & validation), RD3 (linkage quality), and RD4 (participant-selection flow) are the highest-yield; run them first.
|
|
12
|
+
|
|
13
|
+
**RD1 — Database identity, provenance, and fitness-for-purpose**:
|
|
14
|
+
- Is the **database named** and its **type** (claims / EHR / registry / health-checkup), provenance, coverage population, and capture window described — and can it **structurally observe** the exposure and outcome? A claims database cannot see out-of-network or cash-pay care; an EHR cannot see care at other systems; neither reliably sees OTC drugs, over-the-counter outcomes, or death out of hospital.
|
|
15
|
+
- Is the **timeframe and geographic setting** stated (title/abstract per RECORD 1.1–1.2)?
|
|
16
|
+
- An unnamed/undescribed database, or one whose structure cannot capture the named exposure/outcome (so the measure is systematically incomplete), → MAJOR.
|
|
17
|
+
|
|
18
|
+
**RD2 — Phenotype definitions: code-lists and algorithms, evidenced not asserted**:
|
|
19
|
+
- Are the **codes / algorithms** (ICD-9/10, CPT/HCPCS, ATC/NDC drug codes, Read/SNOMED, lab thresholds) used to define the **population, exposure, outcome, confounders, and effect modifiers** provided in full (text or supplement), per RECORD 6.1 / 7.1? "Diabetes/MI/the cohort was identified from the database" with **no code list** is the single most common RECORD failure.
|
|
20
|
+
- Are those definitions **validated** — a referenced validation study, or a PPV/sensitivity estimate — or at least is the lack of validation acknowledged (RECORD 6.2)? An unvalidated outcome algorithm presented as if it were a gold-standard diagnosis is a misclassification risk (RD5).
|
|
21
|
+
- Missing code-lists, or validated-sounding phenotypes with no validation reference/acknowledgement → MAJOR.
|
|
22
|
+
|
|
23
|
+
**RD3 — Data linkage and linkage-quality evaluation**:
|
|
24
|
+
- If two or more databases were **linked**, is the **linkage method** (deterministic on a unique identifier vs **probabilistic**, and on which fields) and the **linkage-quality evaluation** (match/linkage rate, handling of non-matches and false matches, any bias in who links) reported (RECORD 12.3)? Is a **person-flow at each linkage stage** shown (RECORD 6.3)?
|
|
25
|
+
- Are individuals who **failed to link** characterised (linkage is often differential by age/region/insurance), and is the impact on selection considered?
|
|
26
|
+
- Undisclosed linkage method or quality, no linkage-stage flow, or treating the linked subset as representative without examining non-linkage → MAJOR.
|
|
27
|
+
|
|
28
|
+
**RD4 — Participant-selection flow including data-quality filtering**:
|
|
29
|
+
- Is there a **selection/flow** from the source database to the analytic cohort that includes filtering on **data quality, data availability, and linkage** — not only clinical eligibility — with the **N at each step** (RECORD 13.1)? A jump from "the database contains N million records" straight to an analytic N, with the exclusions opaque, hides selection bias.
|
|
30
|
+
- Is the **analysis unit** (persons vs records/encounters/claims) explicit and consistent (cross-link `observational_confounding.md` O8)?
|
|
31
|
+
- No data-driven selection flow, or an unexplained gap between source and analytic N → MAJOR.
|
|
32
|
+
|
|
33
|
+
**RD5 — Misclassification of exposure and outcome**:
|
|
34
|
+
- Are **exposure and outcome misclassification** (from coding/recording, not clinical adjudication) acknowledged and, where possible, **quantified** (validation PPV/sensitivity, quantitative bias analysis, or a sensitivity analysis under alternative definitions)? Are **proxy/surrogate** measures (a prescription ≠ ingestion; a code ≠ the disease) flagged as such?
|
|
35
|
+
- Coded variables treated as gold-standard with no misclassification discussion, or a single rigid definition with no sensitivity to a broader/narrower one → MAJOR (or MINOR if non-differential and acknowledged).
|
|
36
|
+
|
|
37
|
+
**RD6 — Missing data and informative missingness**:
|
|
38
|
+
- Secondary data are frequently **missing-not-at-random** — a lab not ordered is not a normal lab, an unrecorded covariate is not absence of the condition. Is missingness **characterised** (extent, pattern) and handled appropriately (not a naive complete-case that assumes MCAR when missingness is informative; multiple imputation or a sensitivity analysis where justified)?
|
|
39
|
+
- Naive complete-case analysis on informatively-missing EHR fields, or treating "no record of X" as "X absent" without justification → MAJOR.
|
|
40
|
+
|
|
41
|
+
**RD7 — Unmeasured confounding and RWD-specific design bias**:
|
|
42
|
+
- Is **unmeasured/residual confounding** confronted — secondary data often lack lifestyle, disease severity, frailty, or over-the-counter exposures — with a negative-control, E-value, or sensitivity analysis, rather than asserting "adjusted for available confounders" (cross-link `observational_confounding.md`)?
|
|
43
|
+
- For an exposure/drug study, are the biases endemic to RWD addressed by **design**: **immortal-time bias** (time-fixed exposure misclassified person-time), **prevalent-user bias** (new-user / active-comparator design), **protopathic/reverse-causation bias** (a lag/induction window), and confounding by indication? (These are the core of **RECORD-PE**.)
|
|
44
|
+
- An effect estimate with no engagement with unmeasured confounding, or a drug-effect design exposed to immortal-time / prevalent-user bias with no mitigation → MAJOR.
|
|
45
|
+
|
|
46
|
+
**RD8 — Eligibility drift, data access, and reproducibility**:
|
|
47
|
+
- Over the study window, did **coding systems or eligibility/enrolment rules change** (ICD-9→10 transition, formulary or coverage changes), and is that acknowledged (RECORD 19.1)?
|
|
48
|
+
- Are the **extent of data access**, the **data-cleaning methods**, and the **availability of the protocol, derived-variable definitions / code-lists, and analysis code** stated (RECORD 12.1 / 12.2 / 22.1)? Reproducibility in RWD studies rests on the published phenotype definitions and code.
|
|
49
|
+
- Unacknowledged coding/eligibility drift over a multi-year window, or no availability of protocol/code-lists/code for a non-public database, → MAJOR (drift) / MINOR (availability), per centrality.
|
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
<!-- Domain probe module — shared, vendored BYTE-IDENTICAL by /peer-review and /self-review.
|
|
2
|
+
Severity words below (MAJOR / MINOR / major / minor) denote finding severity, NOT a journal
|
|
3
|
+
recommendation. Each consuming skill maps findings to its own output:
|
|
4
|
+
- peer-review: Major / Minor comments + Confidential Comments to the Editor; a
|
|
5
|
+
representativeness / response-rate / instrument-validity flaw is placed as Major #1.
|
|
6
|
+
- self-review: Anticipated Major / Minor Comments (Fatal / Fixable) mapped to category letters.
|
|
7
|
+
Do NOT edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`. -->
|
|
8
|
+
|
|
9
|
+
# Survey / questionnaire study probes (SV1–SV8)
|
|
10
|
+
|
|
11
|
+
An 8-probe checklist for **self-report survey / questionnaire studies** — knowledge-attitudes-practices (KAP), physician and patient surveys, cross-sectional questionnaires, and web/e-surveys. These probes complement (do not replace) the generic Phase 2 checklist, the **CROSS** reporting items (and **CHERRIES** for internet surveys), and the scale-reliability guidance (`analyze-stats` Survey/Likert). They target the gap between a clean results table and whether the sample can support a population claim at all: representativeness, the response-rate denominator and non-response bias, and whether the instrument measures what it claims. The most common failure is generalising from a self-selected convenience sample. SV1 (representativeness), SV3 (response rate & non-response), and SV4 (instrument validity) are the highest-yield; run them first.
|
|
12
|
+
|
|
13
|
+
**SV1 — Target population, sampling frame, and representativeness**:
|
|
14
|
+
- Is the **target population** defined and the **sampling frame** (the actual list/channel from which respondents were drawn) stated, with comment on how well the frame **covers** the target population (coverage error)? A survey distributed via a society listserv, social media, or a conference app reaches a self-selected slice, not "clinicians" or "the public."
|
|
15
|
+
- Is there any evidence the respondents **resemble the target population** (compare respondent demographics to a known population, or to the frame)?
|
|
16
|
+
- A convenience / self-selected / undisclosed-frame sample whose results are generalised to a population, with no representativeness assessment → MAJOR (downgrade claims to "among respondents").
|
|
17
|
+
|
|
18
|
+
**SV2 — Sampling method and sample-size justification**:
|
|
19
|
+
- Is the **sampling method** stated and correctly characterised — **probability** (random/systematic/stratified/cluster) vs **non-probability** (convenience/snowball/quota)? A non-probability sample cannot yield design-unbiased population estimates and should not be presented as if it could.
|
|
20
|
+
- Is there an **a-priori sample-size or precision justification** (for a prevalence/estimate: target margin of error; for a comparison: power), rather than a post-hoc rationalisation of however many happened to respond?
|
|
21
|
+
- A non-probability sample presented as representative, or no sample-size rationale for a precision / comparison claim → MAJOR / MINOR per the claim.
|
|
22
|
+
|
|
23
|
+
**SV3 — Response rate (defined denominator) and non-response bias**:
|
|
24
|
+
- Is a **response rate reported with an explicit, defensible denominator** (an AAPOR/CASRO-style definition: completed responses ÷ eligible invitees), not just "N people responded"? For an **open** web survey where the denominator is unknowable, is that limitation stated (a view/participation/completion rate per CHERRIES instead of a true response rate)?
|
|
25
|
+
- Is **non-response bias** assessed — responders vs non-responders, early vs late responders, or respondents vs the population? A low response rate is not fatal, but an **unassessed** low/undefined response rate carrying a population estimate is.
|
|
26
|
+
- No defined denominator, or a low response rate with no non-response analysis behind a population claim → MAJOR.
|
|
27
|
+
|
|
28
|
+
**SV4 — Instrument development, validity, and reliability**:
|
|
29
|
+
- Was the questionnaire **previously validated** (cited) or **newly developed**? For a new/adapted instrument, was it **pre-tested / piloted** (cognitive interviewing, a pilot sample)?
|
|
30
|
+
- For **multi-item scales**, are **validity** (content/construct, factor structure) and **reliability** (Cronbach's α / McDonald's ω, test–retest) reported? (A negative or implausibly low α usually signals a reverse-coded item not re-scored — see the scale-reliability guidance, not a multidimensionality story.)
|
|
31
|
+
- A novel, unvalidated, un-piloted instrument carrying the headline, or multi-item scales with no reliability evidence → MAJOR (or MINOR if the instrument is established and cited).
|
|
32
|
+
|
|
33
|
+
**SV5 — Administration mode, coverage, and e-survey (CHERRIES) reporting**:
|
|
34
|
+
- Is the **mode** (web, email, postal, telephone, in-person) and its **coverage/selection implications** stated (a web survey excludes the digitally excluded; a clinic survey excludes non-attenders)?
|
|
35
|
+
- For an **internet survey**, are the CHERRIES specifics reported: **open vs closed** (invited) survey; how the **denominator and completion** were computed; **voluntariness and any incentive**; **duplicate-submission control** (IP/cookie/log-in); and use of mandatory/adaptive questions and completeness?
|
|
36
|
+
- A web survey with no CHERRIES reporting (unknown denominator, no duplicate control, undisclosed incentive) → MAJOR / MINOR per centrality.
|
|
37
|
+
|
|
38
|
+
**SV6 — Question design and measurement**:
|
|
39
|
+
- Is the **instrument available** (appended or referenced) so wording can be judged? Are there **leading, double-barrelled, or ambiguous** questions, and is the handling of **neutral / "don't know" / not-applicable** options appropriate (forced-choice can manufacture opinion)?
|
|
40
|
+
- Are **Likert / ordinal** items treated appropriately (ordinal vs assumed-interval), and composite scores justified?
|
|
41
|
+
- An unavailable instrument, biased item wording, or inappropriate scale treatment driving a conclusion → MAJOR / MINOR.
|
|
42
|
+
|
|
43
|
+
**SV7 — Analysis, weighting, denominators, and missing data**:
|
|
44
|
+
- For a sample that under-represents parts of the target population, were **design weights / post-stratification** applied (and the weighting described), or are unweighted estimates presented as population figures?
|
|
45
|
+
- Are **per-item denominators** explicit and consistent (completers vs all respondents; the denominator should not silently shift across items), and is **item-level missingness / partial completion** handled and reported?
|
|
46
|
+
- Unweighted estimates from a skewed sample presented as population values, or shifting/opaque denominators → MAJOR.
|
|
47
|
+
|
|
48
|
+
**SV8 — Interpretation, generalisability, ethics, and reporting**:
|
|
49
|
+
- Are conclusions **matched to the sampled population** (no over-generalisation from a single-setting / low-response / convenience sample to "physicians" or "patients" broadly), and are self-report and **social-desirability** biases acknowledged?
|
|
50
|
+
- Are **ethics** (consent, IRB approval/exemption, data protection/anonymity) reported, and is the study mapped to **CROSS** (and **CHERRIES** for e-surveys) with instrument/data availability where possible?
|
|
51
|
+
- Over-generalisation beyond the sampled population, or missing ethics/consent reporting for an identifiable-respondent survey → MAJOR (generalisation) / MINOR (reporting), per centrality.
|
|
@@ -301,6 +301,8 @@ These modules carry the same domain-specific critique probes used by `/peer-revi
|
|
|
301
301
|
| Polygenic risk score / polygenic score (PRS / PGS) developed, validated, or applied as a predictor or risk-stratifier | `references/domain-probes/polygenic_risk_score.md` (PG1–PG8) |
|
|
302
302
|
| Network meta-analysis (≥3 interventions via direct + indirect evidence, treatment ranking, incl. component NMA) | `references/domain-probes/network_meta_analysis.md` (NM1–NM8) |
|
|
303
303
|
| Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact; trial-based or decision-model-based — decision tree, Markov, DES) | `references/domain-probes/health_economic_evaluation.md` (HE1–HE8) |
|
|
304
|
+
| Observational study using routinely-collected health data (administrative claims / EHR / disease or population registry / health-checkup DB, linked or not) | `references/domain-probes/record_routinely_collected_data.md` (RD1–RD8) |
|
|
305
|
+
| Self-report survey / questionnaire study (KAP, physician/patient survey, cross-sectional questionnaire, web/e-survey) | `references/domain-probes/survey_research.md` (SV1–SV8) |
|
|
304
306
|
|
|
305
307
|
When the manuscript matches a row, read `${CLAUDE_SKILL_DIR}/references/domain-probes/<module>.md` and apply each probe as an additional source of Anticipated Major / Minor Comments. The module severity words (MAJOR / MINOR) map to this skill's framing as follows: a conclusion-threatening or design-level finding becomes a **Fatal** Anticipated Major Comment, a reporting-level finding becomes a **Fixable** Anticipated Minor Comment, and each is tagged with the closest category letter (A–K). These probes **complement** categories A–K above; they do not replace them. (The modules are vendored byte-identical from `/peer-review`; do not edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`.)
|
|
306
308
|
|
|
@@ -0,0 +1,49 @@
|
|
|
1
|
+
<!-- Domain probe module — shared, vendored BYTE-IDENTICAL by /peer-review and /self-review.
|
|
2
|
+
Severity words below (MAJOR / MINOR / major / minor) denote finding severity, NOT a journal
|
|
3
|
+
recommendation. Each consuming skill maps findings to its own output:
|
|
4
|
+
- peer-review: Major / Minor comments + Confidential Comments to the Editor; a
|
|
5
|
+
code-list / linkage-quality / selection-flow / RWD-bias design flaw is placed as Major #1.
|
|
6
|
+
- self-review: Anticipated Major / Minor Comments (Fatal / Fixable) mapped to category letters.
|
|
7
|
+
Do NOT edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`. -->
|
|
8
|
+
|
|
9
|
+
# Routinely-collected-data study probes (RD1–RD8)
|
|
10
|
+
|
|
11
|
+
An 8-probe checklist for observational studies conducted using **routinely-collected health data** — administrative claims, electronic health records (EHR), disease/population registries, health-administrative and health-checkup databases, and linked versions of these (data **not collected for the study's purpose**). These probes complement (do not replace) the generic Phase 2 checklist, the **STROBE** + **RECORD** reporting items (and **RECORD-PE** for drug studies), and the observational-confounding probes (`observational_confounding.md`, which cover adjustment/collider/analysis-unit issues). They target what secondary-use data add: whether the database can even observe the question, whether the phenotype code-lists and linkage are evidenced rather than asserted, and whether the limitations endemic to data collected for another purpose are confronted. RD2 (phenotype code-lists & validation), RD3 (linkage quality), and RD4 (participant-selection flow) are the highest-yield; run them first.
|
|
12
|
+
|
|
13
|
+
**RD1 — Database identity, provenance, and fitness-for-purpose**:
|
|
14
|
+
- Is the **database named** and its **type** (claims / EHR / registry / health-checkup), provenance, coverage population, and capture window described — and can it **structurally observe** the exposure and outcome? A claims database cannot see out-of-network or cash-pay care; an EHR cannot see care at other systems; neither reliably sees OTC drugs, over-the-counter outcomes, or death out of hospital.
|
|
15
|
+
- Is the **timeframe and geographic setting** stated (title/abstract per RECORD 1.1–1.2)?
|
|
16
|
+
- An unnamed/undescribed database, or one whose structure cannot capture the named exposure/outcome (so the measure is systematically incomplete), → MAJOR.
|
|
17
|
+
|
|
18
|
+
**RD2 — Phenotype definitions: code-lists and algorithms, evidenced not asserted**:
|
|
19
|
+
- Are the **codes / algorithms** (ICD-9/10, CPT/HCPCS, ATC/NDC drug codes, Read/SNOMED, lab thresholds) used to define the **population, exposure, outcome, confounders, and effect modifiers** provided in full (text or supplement), per RECORD 6.1 / 7.1? "Diabetes/MI/the cohort was identified from the database" with **no code list** is the single most common RECORD failure.
|
|
20
|
+
- Are those definitions **validated** — a referenced validation study, or a PPV/sensitivity estimate — or at least is the lack of validation acknowledged (RECORD 6.2)? An unvalidated outcome algorithm presented as if it were a gold-standard diagnosis is a misclassification risk (RD5).
|
|
21
|
+
- Missing code-lists, or validated-sounding phenotypes with no validation reference/acknowledgement → MAJOR.
|
|
22
|
+
|
|
23
|
+
**RD3 — Data linkage and linkage-quality evaluation**:
|
|
24
|
+
- If two or more databases were **linked**, is the **linkage method** (deterministic on a unique identifier vs **probabilistic**, and on which fields) and the **linkage-quality evaluation** (match/linkage rate, handling of non-matches and false matches, any bias in who links) reported (RECORD 12.3)? Is a **person-flow at each linkage stage** shown (RECORD 6.3)?
|
|
25
|
+
- Are individuals who **failed to link** characterised (linkage is often differential by age/region/insurance), and is the impact on selection considered?
|
|
26
|
+
- Undisclosed linkage method or quality, no linkage-stage flow, or treating the linked subset as representative without examining non-linkage → MAJOR.
|
|
27
|
+
|
|
28
|
+
**RD4 — Participant-selection flow including data-quality filtering**:
|
|
29
|
+
- Is there a **selection/flow** from the source database to the analytic cohort that includes filtering on **data quality, data availability, and linkage** — not only clinical eligibility — with the **N at each step** (RECORD 13.1)? A jump from "the database contains N million records" straight to an analytic N, with the exclusions opaque, hides selection bias.
|
|
30
|
+
- Is the **analysis unit** (persons vs records/encounters/claims) explicit and consistent (cross-link `observational_confounding.md` O8)?
|
|
31
|
+
- No data-driven selection flow, or an unexplained gap between source and analytic N → MAJOR.
|
|
32
|
+
|
|
33
|
+
**RD5 — Misclassification of exposure and outcome**:
|
|
34
|
+
- Are **exposure and outcome misclassification** (from coding/recording, not clinical adjudication) acknowledged and, where possible, **quantified** (validation PPV/sensitivity, quantitative bias analysis, or a sensitivity analysis under alternative definitions)? Are **proxy/surrogate** measures (a prescription ≠ ingestion; a code ≠ the disease) flagged as such?
|
|
35
|
+
- Coded variables treated as gold-standard with no misclassification discussion, or a single rigid definition with no sensitivity to a broader/narrower one → MAJOR (or MINOR if non-differential and acknowledged).
|
|
36
|
+
|
|
37
|
+
**RD6 — Missing data and informative missingness**:
|
|
38
|
+
- Secondary data are frequently **missing-not-at-random** — a lab not ordered is not a normal lab, an unrecorded covariate is not absence of the condition. Is missingness **characterised** (extent, pattern) and handled appropriately (not a naive complete-case that assumes MCAR when missingness is informative; multiple imputation or a sensitivity analysis where justified)?
|
|
39
|
+
- Naive complete-case analysis on informatively-missing EHR fields, or treating "no record of X" as "X absent" without justification → MAJOR.
|
|
40
|
+
|
|
41
|
+
**RD7 — Unmeasured confounding and RWD-specific design bias**:
|
|
42
|
+
- Is **unmeasured/residual confounding** confronted — secondary data often lack lifestyle, disease severity, frailty, or over-the-counter exposures — with a negative-control, E-value, or sensitivity analysis, rather than asserting "adjusted for available confounders" (cross-link `observational_confounding.md`)?
|
|
43
|
+
- For an exposure/drug study, are the biases endemic to RWD addressed by **design**: **immortal-time bias** (time-fixed exposure misclassified person-time), **prevalent-user bias** (new-user / active-comparator design), **protopathic/reverse-causation bias** (a lag/induction window), and confounding by indication? (These are the core of **RECORD-PE**.)
|
|
44
|
+
- An effect estimate with no engagement with unmeasured confounding, or a drug-effect design exposed to immortal-time / prevalent-user bias with no mitigation → MAJOR.
|
|
45
|
+
|
|
46
|
+
**RD8 — Eligibility drift, data access, and reproducibility**:
|
|
47
|
+
- Over the study window, did **coding systems or eligibility/enrolment rules change** (ICD-9→10 transition, formulary or coverage changes), and is that acknowledged (RECORD 19.1)?
|
|
48
|
+
- Are the **extent of data access**, the **data-cleaning methods**, and the **availability of the protocol, derived-variable definitions / code-lists, and analysis code** stated (RECORD 12.1 / 12.2 / 22.1)? Reproducibility in RWD studies rests on the published phenotype definitions and code.
|
|
49
|
+
- Unacknowledged coding/eligibility drift over a multi-year window, or no availability of protocol/code-lists/code for a non-public database, → MAJOR (drift) / MINOR (availability), per centrality.
|
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
<!-- Domain probe module — shared, vendored BYTE-IDENTICAL by /peer-review and /self-review.
|
|
2
|
+
Severity words below (MAJOR / MINOR / major / minor) denote finding severity, NOT a journal
|
|
3
|
+
recommendation. Each consuming skill maps findings to its own output:
|
|
4
|
+
- peer-review: Major / Minor comments + Confidential Comments to the Editor; a
|
|
5
|
+
representativeness / response-rate / instrument-validity flaw is placed as Major #1.
|
|
6
|
+
- self-review: Anticipated Major / Minor Comments (Fatal / Fixable) mapped to category letters.
|
|
7
|
+
Do NOT edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`. -->
|
|
8
|
+
|
|
9
|
+
# Survey / questionnaire study probes (SV1–SV8)
|
|
10
|
+
|
|
11
|
+
An 8-probe checklist for **self-report survey / questionnaire studies** — knowledge-attitudes-practices (KAP), physician and patient surveys, cross-sectional questionnaires, and web/e-surveys. These probes complement (do not replace) the generic Phase 2 checklist, the **CROSS** reporting items (and **CHERRIES** for internet surveys), and the scale-reliability guidance (`analyze-stats` Survey/Likert). They target the gap between a clean results table and whether the sample can support a population claim at all: representativeness, the response-rate denominator and non-response bias, and whether the instrument measures what it claims. The most common failure is generalising from a self-selected convenience sample. SV1 (representativeness), SV3 (response rate & non-response), and SV4 (instrument validity) are the highest-yield; run them first.
|
|
12
|
+
|
|
13
|
+
**SV1 — Target population, sampling frame, and representativeness**:
|
|
14
|
+
- Is the **target population** defined and the **sampling frame** (the actual list/channel from which respondents were drawn) stated, with comment on how well the frame **covers** the target population (coverage error)? A survey distributed via a society listserv, social media, or a conference app reaches a self-selected slice, not "clinicians" or "the public."
|
|
15
|
+
- Is there any evidence the respondents **resemble the target population** (compare respondent demographics to a known population, or to the frame)?
|
|
16
|
+
- A convenience / self-selected / undisclosed-frame sample whose results are generalised to a population, with no representativeness assessment → MAJOR (downgrade claims to "among respondents").
|
|
17
|
+
|
|
18
|
+
**SV2 — Sampling method and sample-size justification**:
|
|
19
|
+
- Is the **sampling method** stated and correctly characterised — **probability** (random/systematic/stratified/cluster) vs **non-probability** (convenience/snowball/quota)? A non-probability sample cannot yield design-unbiased population estimates and should not be presented as if it could.
|
|
20
|
+
- Is there an **a-priori sample-size or precision justification** (for a prevalence/estimate: target margin of error; for a comparison: power), rather than a post-hoc rationalisation of however many happened to respond?
|
|
21
|
+
- A non-probability sample presented as representative, or no sample-size rationale for a precision / comparison claim → MAJOR / MINOR per the claim.
|
|
22
|
+
|
|
23
|
+
**SV3 — Response rate (defined denominator) and non-response bias**:
|
|
24
|
+
- Is a **response rate reported with an explicit, defensible denominator** (an AAPOR/CASRO-style definition: completed responses ÷ eligible invitees), not just "N people responded"? For an **open** web survey where the denominator is unknowable, is that limitation stated (a view/participation/completion rate per CHERRIES instead of a true response rate)?
|
|
25
|
+
- Is **non-response bias** assessed — responders vs non-responders, early vs late responders, or respondents vs the population? A low response rate is not fatal, but an **unassessed** low/undefined response rate carrying a population estimate is.
|
|
26
|
+
- No defined denominator, or a low response rate with no non-response analysis behind a population claim → MAJOR.
|
|
27
|
+
|
|
28
|
+
**SV4 — Instrument development, validity, and reliability**:
|
|
29
|
+
- Was the questionnaire **previously validated** (cited) or **newly developed**? For a new/adapted instrument, was it **pre-tested / piloted** (cognitive interviewing, a pilot sample)?
|
|
30
|
+
- For **multi-item scales**, are **validity** (content/construct, factor structure) and **reliability** (Cronbach's α / McDonald's ω, test–retest) reported? (A negative or implausibly low α usually signals a reverse-coded item not re-scored — see the scale-reliability guidance, not a multidimensionality story.)
|
|
31
|
+
- A novel, unvalidated, un-piloted instrument carrying the headline, or multi-item scales with no reliability evidence → MAJOR (or MINOR if the instrument is established and cited).
|
|
32
|
+
|
|
33
|
+
**SV5 — Administration mode, coverage, and e-survey (CHERRIES) reporting**:
|
|
34
|
+
- Is the **mode** (web, email, postal, telephone, in-person) and its **coverage/selection implications** stated (a web survey excludes the digitally excluded; a clinic survey excludes non-attenders)?
|
|
35
|
+
- For an **internet survey**, are the CHERRIES specifics reported: **open vs closed** (invited) survey; how the **denominator and completion** were computed; **voluntariness and any incentive**; **duplicate-submission control** (IP/cookie/log-in); and use of mandatory/adaptive questions and completeness?
|
|
36
|
+
- A web survey with no CHERRIES reporting (unknown denominator, no duplicate control, undisclosed incentive) → MAJOR / MINOR per centrality.
|
|
37
|
+
|
|
38
|
+
**SV6 — Question design and measurement**:
|
|
39
|
+
- Is the **instrument available** (appended or referenced) so wording can be judged? Are there **leading, double-barrelled, or ambiguous** questions, and is the handling of **neutral / "don't know" / not-applicable** options appropriate (forced-choice can manufacture opinion)?
|
|
40
|
+
- Are **Likert / ordinal** items treated appropriately (ordinal vs assumed-interval), and composite scores justified?
|
|
41
|
+
- An unavailable instrument, biased item wording, or inappropriate scale treatment driving a conclusion → MAJOR / MINOR.
|
|
42
|
+
|
|
43
|
+
**SV7 — Analysis, weighting, denominators, and missing data**:
|
|
44
|
+
- For a sample that under-represents parts of the target population, were **design weights / post-stratification** applied (and the weighting described), or are unweighted estimates presented as population figures?
|
|
45
|
+
- Are **per-item denominators** explicit and consistent (completers vs all respondents; the denominator should not silently shift across items), and is **item-level missingness / partial completion** handled and reported?
|
|
46
|
+
- Unweighted estimates from a skewed sample presented as population values, or shifting/opaque denominators → MAJOR.
|
|
47
|
+
|
|
48
|
+
**SV8 — Interpretation, generalisability, ethics, and reporting**:
|
|
49
|
+
- Are conclusions **matched to the sampled population** (no over-generalisation from a single-setting / low-response / convenience sample to "physicians" or "patients" broadly), and are self-report and **social-desirability** biases acknowledged?
|
|
50
|
+
- Are **ethics** (consent, IRB approval/exemption, data protection/anonymity) reported, and is the study mapped to **CROSS** (and **CHERRIES** for e-surveys) with instrument/data availability where possible?
|
|
51
|
+
- Over-generalisation beyond the sampled population, or missing ethics/consent reporting for an identifiable-respondent survey → MAJOR (generalisation) / MINOR (reporting), per centrality.
|