medsci-skills 5.7.0 → 5.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  **51 skills that actually work.** Built by a physician-researcher, tested on real publications.
6
6
 
7
- *MedSci Skills is an end-to-end research tool for physician and medical-engineering researchers — design → scaffold → validate → publish — for the clinical manuscript and the medical-AI model behind it. Its moat is the compliance layer — 42 reporting guidelines and risk-of-bias tools, reference/citation verification, and deterministic integrity gates before peer review — now extended by a model-engineering lane that scaffolds reproducible, leakage-safe training repos and audits model validation. Clinical AI model research engineering is in scope; a general AI-scientist platform is not. It competes on clinical submission reliability, not skill count.*
7
+ *MedSci Skills is an end-to-end research tool for physician and medical-engineering researchers — design → scaffold → validate → publish — for the clinical manuscript and the medical-AI model behind it. Its moat is the compliance layer — 44 reporting guidelines and risk-of-bias tools, reference/citation verification, and deterministic integrity gates before peer review — now extended by a model-engineering lane that scaffolds reproducible, leakage-safe training repos and audits model validation. Clinical AI model research engineering is in scope; a general AI-scientist platform is not. It competes on clinical submission reliability, not skill count.*
8
8
 
9
9
  [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
10
10
  [![Release](https://img.shields.io/github/v/release/Aperivue/medsci-skills?style=flat-square&color=blue)](https://github.com/Aperivue/medsci-skills/releases/latest)
@@ -281,54 +281,54 @@ The E2E pipeline (`orchestrate --e2e`) produces everything up to `qc/`. The `sub
281
281
 
282
282
  ## What's New
283
283
 
284
- **v4.10** — reviewer-coverage expansion reverse-engineered from high-IF, CC-BY papers (learn-only under the `reverse_engineer/` license firewall), plus a clinician-friendly update path. Additive and backward-compatible; 45 skills / **42 guidelines** / 36 detectors / **15 domain-probe modules** (was 12):
284
+ **v4.10** — reviewer-coverage expansion reverse-engineered from high-IF, CC-BY papers (learn-only under the `reverse_engineer/` license firewall), plus a clinician-friendly update path. Additive and backward-compatible; 45 skills / **44 guidelines** / 36 detectors / **15 domain-probe modules** (was 12):
285
285
 
286
286
  - **Three new reviewer domain-probe modules** (`/peer-review` + `/self-review`, vendored byte-identical): **Mendelian randomization** (MR1–MR8 — IV assumptions, pleiotropy-robust sensitivity suite, Steiger, sample overlap, NLMR, drug-target colocalization), **polygenic risk score** (PG1–PG8 — ancestry portability, base/target leakage, incremental value over the clinical model, screening-vs-discrimination, calibration), and **network meta-analysis** (NM1–NM8 — transitivity, incoherence, SUCRA over-interpretation, CINeMA/GRADE-NMA, component-NMA additivity). Plus observational **O17** (agnostic many-exposure-scan multiplicity: ExWAS/EWAS/MWAS).
287
287
  - **Two reporting-guideline checklists** (36 → 38): **STROBE-MR** and **PGS-RS / PRS-RS**, with study-type routing. Four new `/analyze-stats` analysis guides (multiplicity, MR, PRS, NMA) and a `/clean-data` implausible-value + cross-field validity reference.
288
288
  - **Clinician-friendly update reminders** — the classroom installers enable the in-app "update available" notice + one-click Desktop updater by default; the `npx`/manual paths print how to turn it on; the install guide recommends `npx medsci-skills install --enable-update-notify`.
289
289
 
290
- **v4.9** — analysis-integrity hardening promoted from real review cycles, plus journal-mechanics additions. Additive and backward-compatible; still 45 skills / 42 guidelines, analysis-integrity detectors **32 → 36**:
290
+ **v4.9** — analysis-integrity hardening promoted from real review cycles, plus journal-mechanics additions. Additive and backward-compatible; still 45 skills / 44 guidelines, analysis-integrity detectors **32 → 36**:
291
291
 
292
292
  - **Four new gates** — a **duplicate-bibliography** check (`check_reference_duplication.py`) for the hybrid `[@key]` + hand-typed `## References` build that renders the list twice; a **cross-script binning / composite-indicator** consistency check (`check_binning_consistency.py`, `BINNING_DRIFT` / `DERIVED_DEF_DRIFT`) for a derived categorical or composite indicator defined inconsistently across analysis scripts; a **float citation-order** check (`check_citation_order.py`) for numbered Tables/Figures not first cited in ascending order per series; and an **audit-dump leak** gate (`/sync-submission`) that blocks a `/check-reporting` output mistakenly attached as a submission file.
293
293
  - **KJR technical-check conventions + percentage-decimal style**, reader-allocation-under-burden and generative-image-as-study-object reporting (`/design-ai-benchmarking`, `/check-reporting`), and a **Liver International** CSL with that journal's submission mechanics (`/manage-refs`).
294
294
 
295
- **v4.8** is the **review-harvest batch** — deterministic detector hardening promoted from real-manuscript review cycles. Additive and backward-compatible; still 45 skills / 42 guidelines, analysis-integrity detectors **30 → 32**:
295
+ **v4.8** is the **review-harvest batch** — deterministic detector hardening promoted from real-manuscript review cycles. Additive and backward-compatible; still 45 skills / 44 guidelines, analysis-integrity detectors **30 → 32**:
296
296
 
297
297
  - **Two new gates** — `check_supplement_hygiene.py` lints the rendered supplement / tables / caption files (not just the manuscript) for §-labels, placeholders, build markers, response-letter framing, and unresolved body↔supplement cross-references; `check_null_calibration.py` flags a headline negative/equivalence claim made without a minimum-detectable-effect / power / equivalence statement.
298
298
  - **Four detector false-positive fixes** — gates no longer fire on a recommended colorblind-safe palette, author-footnote `§` daggers, a correctly-hedged disclaimer, or a tier-label digit; each with a regression fixture and three newly CI-wired test suites.
299
299
  - **Nine reviewer-side domain probes** (SR/MA, observational, diagnostic, AI-overclaiming, survival) plus a `/design-study` design-stage ceiling gate for perceptual/reader-AI studies and a reusable confidence-weighted-rating→AUC monotonicity template.
300
300
 
301
- **v4.7** is the **self-update foundation** — physician-researchers stay current without GitHub, git, or a terminal. Additive and backward-compatible; still 45 skills / 42 guidelines / 30 detectors:
301
+ **v4.7** is the **self-update foundation** — physician-researchers stay current without GitHub, git, or a terminal. Additive and backward-compatible; still 45 skills / 44 guidelines / 30 detectors:
302
302
 
303
303
  - **Transactional, crash-recoverable installer.** Each install runs through a durable journal state machine recovered on the next run (roll back / forward-clean / fail-closed), with per-target SHA-256 inventories — your modified or third-party skills are backed up and never clobbered or auto-deleted.
304
304
  - **One-click self-updater** (`~/.medsci-skills/updater/`, `install.py --check-update`). Verifies the download against the github.com API digest and **never `extractall()`s** (per-entry rejection of traversal / symlink / duplicate / zip-bomb + an allowlist & per-file hash). The release pipeline injects a verified `provenance.json`, attests build provenance, runs on a protected `release` environment, and verifies each ZIP round-trips through the updater's own safe-extract before publishing.
305
305
  - **Opt-in update notice (off by default):** `install.py --enable-update-notify` shows a one-line "update available" message at Claude Code session start — no telemetry, reads nothing about your session, installs nothing. `--disable-update-notify` / `MEDSCI_NO_UPDATE_CHECK=1` turn it off. *(Honest scope: the digest/attestation detect transport tampering, not a compromised publisher account — see `SECURITY.md`.)*
306
306
 
307
- **v4.6** is a maintainability, governance, and review-depth release — still 45 skills / 42 guidelines; analysis-integrity detectors **28 → 30**, domain probes 11 → 12:
307
+ **v4.6** is a maintainability, governance, and review-depth release — still 45 skills / 44 guidelines; analysis-integrity detectors **28 → 30**, domain probes 11 → 12:
308
308
 
309
309
  - **Fairness / equity / subgroup-performance probe (EQ0–EQ6)** for AI/prediction/diagnostic studies that claim cross-population performance, plus two new detectors: an **AI-disclosure + data/code-availability** check (`/sync-submission`) and a **structured-summary-box conformance** check (`/academic-aio`).
310
310
  - **Governance + answer-engine layer:** `ROADMAP.md`, `MAINTAINERS.md`, `SECURITY.md`, a maintainer workflow + release checklist, an AEO/GEO `docs/faq.md`, a "Start here: 3 workflows" + "Validation status" section in this README, and a new `maturity` field (official / experimental / community) on every skill.
311
311
  - **Token diet (pilot):** `write-paper` Phase 7 integrity audits moved to a load-on-demand reference (~2,559 tokens saved per invocation). Positioning now leads with the compliance moat rather than skill count.
312
312
 
313
- **v4.5** deepens the review + submission surface with no new skill or reporting-guideline count (still 45 skills / 42 guidelines); analysis-integrity detectors **27 → 28**:
313
+ **v4.5** deepens the review + submission surface with no new skill or reporting-guideline count (still 45 skills / 44 guidelines); analysis-integrity detectors **27 → 28**:
314
314
 
315
315
  - **`/clean-data` + `/analyze-stats` — reverse-coded-item / negative-alpha detector.** A multi-item Likert scale with a negatively-worded item must be recoded `(min+max) − x` before the scale total or Cronbach's alpha is computed; left un-recoded, the item correlates negatively with the rest of the scale and alpha collapses (often negative). A negative alpha is a coding bug, not a "multidimensional construct." New stdlib-only `check_reverse_coding.py` returns `REVERSE_CODING_LIKELY` / `REVERSE_CODING_SUSPECT` / `OK` from per-item item-rest correlations + raw alpha; the Likert summary template gains a `--reverse-items` recode flag.
316
316
  - **`/peer-review` + `/self-review` — SR/MA + DTA + prediction-model probe batch.** `sr_ma.md` **P12** risk-of-bias table row-sum ↔ traffic-light figure-matrix reconciliation and **P13** included-study ↔ reference-list completeness; `diagnostic_accuracy.md` **D7** index-test-as-enrollment-criterion circularity; `clinical_prediction_model.md` **CP5** intended-use horizon leakage and **CP6** development/CV vs held-out/external validation-nomenclature conflation. Vendored byte-identical into `/self-review`.
317
317
  - **`/sync-submission` — embedded absolute-path leak scan.** A `word/*.xml` attribute (e.g. a pandoc-embedded image's `<pic:cNvPr descr="…">`) carrying an absolute home-dir path (`/Users/…`, `/home/…`) is a username leak invisible to a rendered-text scan; now flagged as `docx_embedded_abs_path` under `check_asset_anonymization.py`.
318
318
 
319
- **v4.4** adds reviewer/analysis depth with no new skill or reporting-guideline count (still 45 skills / 42 guidelines / 27 detectors):
319
+ **v4.4** adds reviewer/analysis depth with no new skill or reporting-guideline count (still 45 skills / 44 guidelines / 27 detectors):
320
320
 
321
321
  - **`/author-strategy` — trajectory-archetype classification (optional).** Classifies a queried author's PubMed trajectory into abstract career archetypes (A1 infrastructure builder, A2 methodology rule-maker, A3 clinical→AI hybrid, A4 SR/MA volume engine, A5 large-consortium participation, A6 device/technique depth, + a computed composite) as an **explainable, multi-label, confidence-scored heuristic — not an objective verdict**. The rubric is a single canonical YAML (the narrative doc is generated from it); scores exclude `unavailable` signals (h-index/citation/venue-tier → `[VERIFY]`, never fabricated); a **disambiguation gate** binds an approved `corpus_manifest.json` to the CSV (csv + PMID-set hashes) so a surname alone never classifies, and target-author attribution never borrows a co-author's ORCID/affiliation.
322
322
  - **`/peer-review` + `/self-review` — Image-Synthesis / cross-modality probe (IS1–IS4)** for studies that synthesize one imaging modality from another and claim the output carries the target's information, plus a reviewer-side reference-integrity spot-check.
323
323
  - **`/verify-refs` — OpenAlex tertiary index** recovers conference-proceedings / non-DOI citations (NeurIPS/ICLR/ACL) that fall through PubMed and CrossRef, the free analogue of a portal's second index.
324
324
 
325
- **v4.3** hardens the **cross-sectional / observational cohort** review surface end-to-end, much of it reverse-engineered from real CC-BY cohort papers (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills / 42 guidelines); analysis-integrity detectors **25 → 27**:
325
+ **v4.3** hardens the **cross-sectional / observational cohort** review surface end-to-end, much of it reverse-engineered from real CC-BY cohort papers (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills / 44 guidelines); analysis-integrity detectors **25 → 27**:
326
326
 
327
327
  - **Observational probes O1 → O14** (`/peer-review` + `/self-review`, vendored) — over-adjustment / analysis-unit clustering / outcome construct-validity (O7–O9), overlapping-subset gradient (O10), **complex-survey design & weighting** for NHANES/KNHANES (O11), **data-driven threshold / "inflection-point" mining** (O12), **cross-sectional mediation** temporal-order & sequential-ignorability (O13), and **interaction scale** — additive RERI/AP/S vs multiplicative (O14). Plus a new **clinical-prediction-model** probe module **CP1–CP4** and survival **S9** (panel-data / multistate variance).
328
328
  - **Two new detectors (25 → 27)** — `check_wordcount_cap.py` (the revision-inflation trap: body vs journal cap) and `check_paren_spans.py` (em-dash→paren conversions that wrap a whole sentence). Plus a `check_confounding_completeness.py` upgrade (DB-code↔prose alias map, SMD-from-mean±SD, exposure-defining-covariate exemption), a `check_cohort_arithmetic.py` `ANALYSIS_UNIT_UNDISCLOSED` check, a `check_scope_coherence.py` cross-sectional-yield lexicon, and a verify-refs corporate/collective-author render-abort fix.
329
329
  - **Analysis & submission tooling** — `/analyze-stats` gains **mediation** and **interaction & effect-modification** guides; `/sync-submission` gains `assemble_supplement.py` (S{N} index↔file integrity) and a `/revise` body-word-count exit gate; `/render-pdf-doc` gains a `scan_glyph_coverage.py` xelatex silent-glyph-drop scan.
330
330
 
331
- **v4.2** builds out the case-report capability end-to-end, grounded in real CC-BY case reports (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills / 42 guidelines); journal profiles **68 → 73**:
331
+ **v4.2** builds out the case-report capability end-to-end, grounded in real CC-BY case reports (learn-only under the license firewall) — no new skill or reporting-guideline count (still 45 skills / 44 guidelines); journal profiles **68 → 73**:
332
332
 
333
333
  - **Case-report + case-series writing** — `/write-paper` gains a CARE narrative + 150-word-abstract case-report exemplar, a **case-series** paper type (methods-light mini-cohort, all-cases summary table, counts-not-rates), and **adverse-event/pharmacovigilance** (Naranjo/WHO-UMC causality) and **diagnostic-pitfall/mimic** subtypes.
334
334
  - **Radiology / imaging-led track** — a dedicated `exemplar_case_report_radiology.md` (per-modality technique→findings→impression, structured-reporting lexicons BI-RADS/LI-RADS/PI-RADS/TI-RADS/Lung-RADS/O-RADS, quantitative threshold honesty, an interventional-radiology procedure/complication subtype, DICOM de-identification) plus a `/make-figures` annotated multimodality imaging-panel exemplar.
@@ -452,7 +452,7 @@ ma-scout -> search-lit -> fulltext-retrieval -> design-study ──> write-proto
452
452
  | **search-lit** | PubMed + Semantic Scholar + bioRxiv search with anti-hallucination citation verification. Token-efficient error handling -- CrossRef failures are silently batched, not repeated. BibTeX output tags each entry with `verified`/`verified_by`/`verified_on` fields so downstream skills can trust the citation provenance. |
453
453
  | **verify-refs** | Pre-submission reference audit for `.md`, `.docx`, `.bib`, or `.tsv` inputs. Extracts references, verifies DOI/PMID via CrossRef/PubMed when available, and writes `qc/reference_audit.json` as the sole output — row-level status (OK / MISMATCH / UNVERIFIED / FABRICATED) lives inside the JSON `records[]` block. `/search-lit` produces candidate BibTeX; `/lit-sync` owns `manuscript/_src/refs.bib`. |
454
454
  | **fulltext-retrieval** | Batch open-access PDF downloader. Unpaywall → PMC → OpenAlex → CrossRef pipeline. OA-only -- no paywall bypass. Input: DOI list or TSV. Optional PDF→Markdown conversion via [pymupdf4llm](https://pymupdf.readthedocs.io/en/latest/pymupdf4llm/) for token-efficient LLM analysis of academic papers. |
455
- | **check-reporting** | Manuscript compliance audit against 42 reporting guidelines and risk of bias tools (STROBE, STROBE-MR, RECORD, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, CHEERS 2022, CROSS, PRISMA, PRISMA-DTA, PRISMA-P, PRISMA-ScR, MOOSE, ARRIVE, CONSORT, CONSORT-AI, CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, SQUIRE 2.0, CLEAR, GRRAS, MI-CLEAR-LLM, SWiM, AMSTAR 2, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Machine-readable JSON summary with `compliance_pct` and `fixable_by_ai` flags for automated pipeline integration. |
455
+ | **check-reporting** | Manuscript compliance audit against 44 reporting guidelines and risk of bias tools (STROBE, STROBE-MR, RECORD, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, CHEERS 2022, CROSS, SRQR, COREQ, PRISMA, PRISMA-DTA, PRISMA-P, PRISMA-ScR, MOOSE, ARRIVE, CONSORT, CONSORT-AI, CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, SQUIRE 2.0, CLEAR, GRRAS, MI-CLEAR-LLM, SWiM, AMSTAR 2, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Machine-readable JSON summary with `compliance_pct` and `fixable_by_ai` flags for automated pipeline integration. |
456
456
  | **analyze-stats** | Statistical analysis code generation (Python/R) for diagnostic accuracy, DTA meta-analysis (bivariate/HSROC), inter-rater agreement, survival analysis, demographics tables, regression (logistic/linear), propensity score (matching/IPTW/overlap weighting), and repeated measures (RM ANOVA/GEE/mixed models). Calibration mandatory for prediction models. |
457
457
  | **meta-analysis** | Full systematic review and meta-analysis pipeline (8 phases). DTA (bivariate/HSROC) and intervention meta-analysis. Protocol to submission-ready manuscript with PRISMA-DTA compliance. |
458
458
  | **make-figures** | Publication-ready figures and visual abstracts: ROC curves, forest plots, PRISMA/CONSORT/STARD flow diagrams, Kaplan-Meier curves, Bland-Altman plots, confusion matrices, and journal-specific visual/graphical abstracts (python-pptx template-based). Communication-first design principles (Nat Hum Behav 2026 — key message, audience, cognitive load, figure-vs-table decision) and five flow-diagram production lessons (official-template fidelity, VML fallback PDF export, docx XML escape, sequential placeholder mapping, version freeze); critic rubric Section G adds 5 communication-first checks. `--study-type` auto-generates the full required figure set; structured `_figure_manifest.md` output for downstream pipeline consumption; D2 enforced as default for flow diagrams. |
@@ -621,8 +621,8 @@ Projects declare their source-of-truth layout in `SSOT.yaml`, and a `qc/migratio
621
621
  ### Meta-Analysis Failure Modes
622
622
  `/meta-analysis` ships empirical failure-mode references (data integrity, review orchestration, submission package drift, post-submission release ops) with four automation hooks: `scripts/prisma_5way_consistency.py` (DI-6 PRISMA number consistency), `scripts/extraction_consensus_log_init.py` (DI-1 dual-extraction scaffold), `scripts/tag_cleanup_gate.sh` (DI-8 placeholder tag gate), and `scripts/verify_package_integrity.py` (SPD SHA-256 manifest for submission bundles).
623
623
 
624
- ### 42 Reporting Guidelines & RoB Tools Built-in
625
- `check-reporting` includes bundled checklists for 42 guidelines and risk-of-bias tools: STROBE, STROBE-MR, RECORD, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, CHEERS 2022, CROSS, PRISMA 2020, PRISMA-DTA, PRISMA-P, PRISMA-ScR, MOOSE, ARRIVE, CONSORT, CONSORT-AI, CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, SQUIRE 2.0, CLEAR, GRRAS, MI-CLEAR-LLM, SWiM, AMSTAR 2, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA. Includes Results/Discussion section boundary checks and machine-readable JSON summary for pipeline integration.
624
+ ### 44 Reporting Guidelines & RoB Tools Built-in
625
+ `check-reporting` includes bundled checklists for 44 guidelines and risk-of-bias tools: STROBE, STROBE-MR, RECORD, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, CHEERS 2022, CROSS, SRQR, COREQ, PRISMA 2020, PRISMA-DTA, PRISMA-P, PRISMA-ScR, MOOSE, ARRIVE, CONSORT, CONSORT-AI, CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, SQUIRE 2.0, CLEAR, GRRAS, MI-CLEAR-LLM, SWiM, AMSTAR 2, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA. Includes Results/Discussion section boundary checks and machine-readable JSON summary for pipeline integration.
626
626
 
627
627
  ### Publication-Ready Output
628
628
  `analyze-stats` generates reproducible Python/R code for 13 analysis types -- including regression, propensity score, and repeated measures -- with mandatory calibration for prediction models. `make-figures` produces journal-specification figures (300 DPI, colorblind-safe palettes, proper dimensions), visual/graphical abstracts, and a tool selection guide (D2 for flow diagrams, matplotlib for data plots). `--study-type` auto-generates the complete figure set for each study design.
@@ -538,8 +538,8 @@
538
538
  },
539
539
  {
540
540
  "path": "skills/check-reporting/SKILL.md",
541
- "size": 38325,
542
- "sha256": "e4ffb4f0879cfbcff82e2bdd9601eeebbb4e31200b2ec1ca61d7ac46db3745e4"
541
+ "size": 39223,
542
+ "sha256": "2dec1d994bb37fb6091f755ab51b2f890eca2b3299e18ca247ae3c7e45be86bb"
543
543
  },
544
544
  {
545
545
  "path": "skills/check-reporting/references/LICENSES.md",
@@ -596,6 +596,11 @@
596
596
  "size": 4376,
597
597
  "sha256": "e923b3d2f550351f952334053978d5f0ca726ddb4798f3c1e4adba36939bd662"
598
598
  },
599
+ {
600
+ "path": "skills/check-reporting/references/checklists/COREQ.md",
601
+ "size": 6809,
602
+ "sha256": "9a8c10db7fa2859ab4d9b26e01a4a15b134e6b5f76c338911a5be118d41e8f98"
603
+ },
599
604
  {
600
605
  "path": "skills/check-reporting/references/checklists/COSMIN_RoB.md",
601
606
  "size": 5535,
@@ -726,6 +731,11 @@
726
731
  "size": 5730,
727
732
  "sha256": "d145e612c8fb7a182d8212257790f2599cd218810a57ae97c0c6309cd6c4f613"
728
733
  },
734
+ {
735
+ "path": "skills/check-reporting/references/checklists/SRQR.md",
736
+ "size": 6930,
737
+ "sha256": "27ecdb3835e122566b113dc9afc1de6d1a988820f45dabc34a9d56f4658016fd"
738
+ },
729
739
  {
730
740
  "path": "skills/check-reporting/references/checklists/STARD.md",
731
741
  "size": 6440,
@@ -788,8 +798,8 @@
788
798
  },
789
799
  {
790
800
  "path": "skills/check-reporting/scripts/check_checklist_exists.py",
791
- "size": 6817,
792
- "sha256": "c0a1eea2927b4b0edc2a4e9f15a3ac69a729fc8fa76d13de41f5519ecd782d2f"
801
+ "size": 6886,
802
+ "sha256": "c2ffe9761b4580e43a9ea7889b124a6591a6687c3cf939bc3ae1aad051e3040c"
793
803
  },
794
804
  {
795
805
  "path": "skills/check-reporting/scripts/check_checklist_version.py",
@@ -2029,7 +2039,7 @@
2029
2039
  {
2030
2040
  "path": "skills/make-figures/references/reporting_guideline_figure_map.md",
2031
2041
  "size": 7149,
2032
- "sha256": "f56ed1e6c77209cef20e82d1caa99aa65d7f41aac074198e217025f183bdee5d"
2042
+ "sha256": "81f2eff4969e055d886976f4c9e26bf20f1946048c2b61563a7c95b7ebec1d86"
2033
2043
  },
2034
2044
  {
2035
2045
  "path": "skills/make-figures/references/visual_abstract_templates/european_radiology.pptx",
@@ -2819,7 +2829,7 @@
2819
2829
  {
2820
2830
  "path": "skills/orchestrate/SKILL.md",
2821
2831
  "size": 35203,
2822
- "sha256": "78920458eece7f1d65bd376595f0c5f7a6d199514554da517242aff0837b62e5"
2832
+ "sha256": "7197cb70e6f8940653ab9adaad158167111a0478edad8a972e7cc8816ca85123"
2823
2833
  },
2824
2834
  {
2825
2835
  "path": "skills/orchestrate/references/dialogue_nodes.md",
@@ -2843,8 +2853,8 @@
2843
2853
  },
2844
2854
  {
2845
2855
  "path": "skills/peer-review/SKILL.md",
2846
- "size": 66025,
2847
- "sha256": "811ebdfbe6ddb04748cc57fd3e02922d07b3f508ec6307e0790f3b3793a0b827"
2856
+ "size": 68213,
2857
+ "sha256": "51feccb82dca34a28a0b9ace5467d23ddc654056af2cd9e152a4559cf074cfc3"
2848
2858
  },
2849
2859
  {
2850
2860
  "path": "skills/peer-review/references/aczel_2021_reviewer2_patterns.md",
@@ -2921,6 +2931,11 @@
2921
2931
  "size": 9718,
2922
2932
  "sha256": "2e383fec4cf2034c62f9e1a419a0ff61e27bc5b3c132f7115b2335bd72472452"
2923
2933
  },
2934
+ {
2935
+ "path": "skills/peer-review/references/domain-probes/qualitative_research.md",
2936
+ "size": 7014,
2937
+ "sha256": "fd82024939bbfc7f30f284a85a06712b69a269b89a2c8e3752484fb6fb23ebea"
2938
+ },
2924
2939
  {
2925
2940
  "path": "skills/peer-review/references/domain-probes/radiomics.md",
2926
2941
  "size": 5344,
@@ -3403,8 +3418,8 @@
3403
3418
  },
3404
3419
  {
3405
3420
  "path": "skills/self-review/SKILL.md",
3406
- "size": 104929,
3407
- "sha256": "085c13cd04dae19970ebb2de239fb6d1de9829ba95847e0e18c1e0bbea1e5edf"
3421
+ "size": 105186,
3422
+ "sha256": "202bf201cbe1b5dfc3703cddaaf2dfc025af5456d5446656132b1545243b0098"
3408
3423
  },
3409
3424
  {
3410
3425
  "path": "skills/self-review/references/domain-probes/ai_overclaiming.md",
@@ -3476,6 +3491,11 @@
3476
3491
  "size": 9718,
3477
3492
  "sha256": "2e383fec4cf2034c62f9e1a419a0ff61e27bc5b3c132f7115b2335bd72472452"
3478
3493
  },
3494
+ {
3495
+ "path": "skills/self-review/references/domain-probes/qualitative_research.md",
3496
+ "size": 7014,
3497
+ "sha256": "fd82024939bbfc7f30f284a85a06712b69a269b89a2c8e3752484fb6fb23ebea"
3498
+ },
3479
3499
  {
3480
3500
  "path": "skills/self-review/references/domain-probes/radiomics.md",
3481
3501
  "size": 5344,
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "schema_version": 1,
3
- "version": "5.7.0",
3
+ "version": "5.8.0",
4
4
  "owned_skills": [
5
5
  "academic-aio",
6
6
  "add-journal",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "medsci-skills",
3
- "version": "5.7.0",
3
+ "version": "5.8.0",
4
4
  "description": "MedSci Skills — a medical/scientific research skill suite for AI coding agents (Claude Code, Codex, Cursor, Copilot). The npm package is a terminal-friendly installer shortcut; the canonical distribution remains the GitHub repository and the Claude Code plugin marketplace.",
5
5
  "license": "SEE LICENSE IN LICENSE",
6
6
  "homepage": "https://github.com/Aperivue/medsci-skills#readme",
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  name: check-reporting
3
- description: Check manuscript compliance with medical research reporting guidelines. Supports 42 guidelines including STROBE, STROBE-MR, RECORD, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, ARRIVE, PRISMA, PRISMA-DTA, PRISMA-P, PRISMA-ScR (scoping reviews), CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SQUIRE 2.0, CLEAR, MOOSE, GRRAS, SWiM, AMSTAR 2, CHEERS 2022, CROSS (survey studies), and risk of bias tools (QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Generates item-by-item assessment with PRESENT/MISSING/PARTIAL status.
4
- triggers: checklist, reporting guideline, STROBE, STROBE-MR, Mendelian randomization, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD-LLM, PGS-RS, PRS-RS, polygenic risk score, polygenic score, PRISMA, PRISMA-DTA, PRISMA-P, PRISMA-ScR, scoping review, scoping, evidence map, ARRIVE, CARE, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SPIRIT, SPIRIT-AI, QUADAS, QUADAS-C, RoB, ROBINS, ROBINS-E, ROBIS, ROB-ME, PROBAST, NOS, COSMIN, AMSTAR, SWiM, CHEERS, economic evaluation, cost-effectiveness, cost-utility, QALY, ICER, RECORD, RECORD-PE, routinely-collected data, registry, claims, electronic health records, EHR, real-world data, CROSS, CHERRIES, survey, questionnaire, KAP, e-survey, response rate, risk of bias, compliance check, LLM accuracy, large language model, clinical deployment
3
+ description: Check manuscript compliance with medical research reporting guidelines. Supports 44 guidelines including STROBE, STROBE-MR, RECORD, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, ARRIVE, PRISMA, PRISMA-DTA, PRISMA-P, PRISMA-ScR (scoping reviews), CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SQUIRE 2.0, CLEAR, MOOSE, GRRAS, SWiM, AMSTAR 2, CHEERS 2022, CROSS (survey studies), SRQR and COREQ (qualitative research), and risk of bias tools (QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Generates item-by-item assessment with PRESENT/MISSING/PARTIAL status.
4
+ triggers: checklist, reporting guideline, STROBE, STROBE-MR, Mendelian randomization, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD-LLM, PGS-RS, PRS-RS, polygenic risk score, polygenic score, PRISMA, PRISMA-DTA, PRISMA-P, PRISMA-ScR, scoping review, scoping, evidence map, ARRIVE, CARE, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SPIRIT, SPIRIT-AI, QUADAS, QUADAS-C, RoB, ROBINS, ROBINS-E, ROBIS, ROB-ME, PROBAST, NOS, COSMIN, AMSTAR, SWiM, CHEERS, economic evaluation, cost-effectiveness, cost-utility, QALY, ICER, RECORD, RECORD-PE, routinely-collected data, registry, claims, electronic health records, EHR, real-world data, CROSS, CHERRIES, survey, questionnaire, KAP, e-survey, response rate, SRQR, COREQ, qualitative research, interviews, focus groups, thematic analysis, grounded theory, reflexivity, risk of bias, compliance check, LLM accuracy, large language model, clinical deployment
5
5
  tools: Read, Write, Edit, Bash, Grep, Glob
6
6
  model: inherit
7
7
  ---
@@ -33,6 +33,8 @@ compliance report suitable for journal submission.
33
33
  - `RECORD.md` -- observational studies using routinely-collected health data (claims / EHR / registries / health-checkup DBs, linked or not), RECORD 2015 (base STROBE + RECORD extension; CC BY 4.0, Benchimol et al. PLoS Med 2015; RECORD-PE for drug studies)
34
34
  - `CROSS.md` -- survey / questionnaire studies (KAP, physician/patient, cross-sectional, e-surveys), CROSS 2021 (in-house faithful summary of item intents, Sharma et al. JGIM 2021) + CHERRIES (CC BY, Eysenbach JMIR 2004) for internet surveys
35
35
  - `PRISMA_ScR.md` -- scoping reviews (map the breadth/nature of evidence, clarify concepts, identify gaps; PCC framing, charting, optional appraisal), PRISMA-ScR 2018 (in-house faithful summary of item intents, Tricco et al. Ann Intern Med 2018; DOI 10.7326/M18-0850)
36
+ - `SRQR.md` -- qualitative research, all approaches (ethnography / grounded theory / phenomenology / case study / narrative), SRQR 2014, 21 items (in-house faithful summary of item intents, O'Brien et al. Acad Med 2014; DOI 10.1097/ACM.0000000000000388)
37
+ - `COREQ.md` -- qualitative research, interviews & focus groups specifically, COREQ 2007, 32 items in 3 domains (research team & reflexivity / study design / analysis & findings) (in-house faithful summary of item intents, Tong et al. Int J Qual Health Care 2007; DOI 10.1093/intqhc/mzm042)
36
38
  - `PRISMA_2020.md` -- systematic reviews (CC BY)
37
39
  - `ARRIVE_2.md` -- animal studies (CC0)
38
40
  - `PRISMA_DTA.md` -- DTA systematic reviews (CC BY, McInnes et al. JAMA 2018)
@@ -96,6 +98,7 @@ user specification.
96
98
  | Observational study using routinely-collected data (claims / EHR / registry / health-checkup DB) | RECORD (base STROBE + RECORD extension; RECORD-PE for drug studies) | -- |
97
99
  | Survey / questionnaire study (KAP, physician/patient, cross-sectional, e-survey) | CROSS (+ CHERRIES for internet surveys) | -- |
98
100
  | Scoping review (maps breadth/nature of evidence, clarifies concepts, identifies gaps — not a focused effectiveness/accuracy question) | PRISMA-ScR (base PRISMA + scoping-review extension) | -- |
101
+ | Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis) | SRQR (all qualitative approaches); COREQ (interviews/focus groups specifically) | -- |
99
102
  | Randomized controlled trial | CONSORT 2025 | CONSORT-AI |
100
103
  | Diagnostic accuracy study | STARD 2015 | STARD-AI |
101
104
  | Prediction model (development/validation) | TRIPOD | TRIPOD+AI |
@@ -0,0 +1,86 @@
1
+ # COREQ Checklist (qualitative — interviews & focus groups)
2
+
3
+ **Consolidated criteria for reporting qualitative research**
4
+ Version: COREQ 2007 — 32 items in 3 domains. **Specific to in-depth interviews and focus groups** (the dominant qualitative data-collection methods in health research). For other qualitative approaches, use the broader **SRQR** (`SRQR.md`).
5
+ Source: Tong A, Sainsbury P, Craig J. *Int J Qual Health Care* 2007;19(6):349–357 (the COREQ statement; DOI 10.1093/intqhc/mzm042). EQUATOR Network.
6
+
7
+ Apply when the manuscript reports an **interview or focus-group** qualitative study. For the design/conduct review of the same study, pair with the QL1–QL8 domain probes in `peer-review` / `self-review` `references/domain-probes/qualitative_research.md`; for a non-interview qualitative approach (ethnography, document analysis, etc.), use `SRQR.md`.
8
+
9
+ > Licensing note: COREQ is published in the *International Journal for Quality in Health Care* (© Oxford University Press), with **no Creative Commons licence**. The items below are an **in-house, faithful summary of the criteria (facts/intents, paraphrased — not the verbatim COREQ wording)** for item-by-item assessment; consult the published article (DOI 10.1093/intqhc/mzm042) for exact item text.
10
+
11
+ ## Criteria (grouped by domain)
12
+
13
+ ### Domain 1 — Research team and reflexivity
14
+ **Personal characteristics**
15
+ | # | Item | What to check is reported |
16
+ |---|------|---------------------------|
17
+ | 1 | Interviewer / facilitator | Which author(s) conducted the interviews or focus groups. |
18
+ | 2 | Credentials | The interviewer's/facilitator's credentials (e.g. PhD, MD, RN). |
19
+ | 3 | Occupation | Their occupation/role at the time of the study. |
20
+ | 4 | Gender | The researcher's gender (where relevant to the dynamic). |
21
+ | 5 | Experience & training | The interviewer's experience and training in qualitative methods. |
22
+
23
+ **Relationship with participants**
24
+ | # | Item | What to check is reported |
25
+ |---|------|---------------------------|
26
+ | 6 | Relationship established | Whether a relationship existed between researcher and participants **before** the study. |
27
+ | 7 | Participant knowledge of the interviewer | What participants knew about the researcher (e.g. personal goals, reasons for the research). |
28
+ | 8 | Interviewer characteristics | Reported researcher characteristics — assumptions, interests, reasons for doing the research. |
29
+
30
+ ### Domain 2 — Study design
31
+ **Theoretical framework**
32
+ | # | Item | What to check is reported |
33
+ |---|------|---------------------------|
34
+ | 9 | Methodological orientation & theory | The methodological orientation/theory underpinning the study (e.g. grounded theory, content analysis, phenomenology). |
35
+
36
+ **Participant selection**
37
+ | # | Item | What to check is reported |
38
+ |---|------|---------------------------|
39
+ | 10 | Sampling | How participants were selected (purposive, convenience, consecutive, snowball). |
40
+ | 11 | Method of approach | How participants were approached (face-to-face, telephone, mail, email). |
41
+ | 12 | Sample size | The number of participants. |
42
+ | 13 | Non-participation | How many declined or dropped out, and why (where known). |
43
+
44
+ **Setting**
45
+ | # | Item | What to check is reported |
46
+ |---|------|---------------------------|
47
+ | 14 | Setting of data collection | Where the data were collected (home, clinic, workplace). |
48
+ | 15 | Presence of non-participants | Whether anyone besides participants and researchers was present. |
49
+ | 16 | Description of sample | The sample's key characteristics (e.g. demographics, dates). |
50
+
51
+ **Data collection**
52
+ | # | Item | What to check is reported |
53
+ |---|------|---------------------------|
54
+ | 17 | Interview guide | Whether questions/prompts/guides were provided, and whether they were pilot-tested. |
55
+ | 18 | Repeat interviews | Whether any interviews were repeated, and how many. |
56
+ | 19 | Audio/visual recording | Whether the data were audio- or video-recorded. |
57
+ | 20 | Field notes | Whether field notes were made during/after the interview or focus group. |
58
+ | 21 | Duration | The duration of the interviews or focus groups. |
59
+ | 22 | Data saturation | Whether **data saturation** was discussed/reached. |
60
+ | 23 | Transcripts returned | Whether participants received the transcripts to review, comment on, or correct. |
61
+
62
+ ### Domain 3 — Analysis and findings
63
+ **Data analysis**
64
+ | # | Item | What to check is reported |
65
+ |---|------|---------------------------|
66
+ | 24 | Number of data coders | How many coders coded the data. |
67
+ | 25 | Description of the coding tree | Whether a description of the coding tree/framework is provided. |
68
+ | 26 | Derivation of themes | Whether themes were identified in advance or derived from the data. |
69
+ | 27 | Software | Any software used to manage/analyse the data. |
70
+ | 28 | Participant checking | Whether participants provided feedback on the findings (**member checking**). |
71
+
72
+ **Reporting**
73
+ | # | Item | What to check is reported |
74
+ |---|------|---------------------------|
75
+ | 29 | Quotations presented | Whether participant **quotations** illustrate the themes/findings, and whether each is identified (e.g. participant number). |
76
+ | 30 | Data & findings consistent | Whether the reported findings cohere with the underlying data shown. |
77
+ | 31 | Clarity of major themes | Whether the major themes are presented clearly in the results. |
78
+ | 32 | Clarity of minor themes | Whether deviant/diverse cases and minor themes are addressed, not only the dominant ones. |
79
+
80
+ ---
81
+
82
+ ## Notes for Assessors
83
+
84
+ - The **highest-yield** checks (where interview/focus-group studies most often fail review): **Domain 1 reflexivity** (items 1–8 — who interviewed, their relationship to participants, and their assumptions; the single most-omitted COREQ domain), **item 9** (a named methodological orientation — not "themes emerged" with no method), **items 10/22** (a stated sampling approach and a **saturation** discussion), **items 24–28** (the coding/analysis process — how many coders, the coding framework, software, member checking), and **item 29** (themes substantiated by **identified participant quotations**).
85
+ - **Do not apply quantitative criteria**: a small purposive sample is appropriate, "generalizability" is **transferability**, and there are no power/effect-size/p-value requirements. A "sample too small / not generalizable" comment is mis-calibrated for an interview study.
86
+ - COREQ is interview/focus-group-specific; for ethnography, document analysis, or other qualitative approaches use `SRQR.md`. This is an **in-house faithful summary of the COREQ criteria (paraphrased intents, not verbatim)**; map the manuscript's content to the items rather than to exact wording, and verify against the published checklist (Tong et al. *Int J Qual Health Care* 2007; DOI 10.1093/intqhc/mzm042). Verified 2026-06-30.
@@ -0,0 +1,64 @@
1
+ # SRQR Checklist (qualitative research)
2
+
3
+ **Standards for Reporting Qualitative Research**
4
+ Version: SRQR 2014 — 21 items across Title/Abstract, Introduction, Methods, Results/Findings, Discussion, and Other. **Broad** — applies to all qualitative approaches (ethnography, grounded theory, phenomenology, case study, narrative research), not only interviews/focus groups.
5
+ Source: O'Brien BC, Harris IB, Beckman TJ, Reed DA, Cook DA. *Acad Med* 2014;89(9):1245–1251 (the SRQR statement; DOI 10.1097/ACM.0000000000000388). EQUATOR Network.
6
+
7
+ Apply when the manuscript is a **qualitative study** of any approach — interviews, focus groups, observation/ethnography, document analysis, grounded theory, phenomenology, narrative research. For **interview / focus-group** studies specifically, the more granular **COREQ** (`COREQ.md`) is the better-fit companion. For the design/conduct review of the same study, pair with the QL1–QL8 domain probes in `peer-review` / `self-review` `references/domain-probes/qualitative_research.md`.
8
+
9
+ > Licensing note: SRQR is published in *Academic Medicine* (© AAMC), with **no Creative Commons licence**. The items below are an **in-house, faithful summary of the reporting items (facts/intents, paraphrased — not the verbatim SRQR wording)** for item-by-item assessment; consult the published article (DOI 10.1097/ACM.0000000000000388) for exact item text.
10
+
11
+ ## Reporting items (grouped by section)
12
+
13
+ ### Title and Abstract
14
+ | # | Item | What to check is reported |
15
+ |---|------|---------------------------|
16
+ | 1 | Title | A concise description of the study's nature/topic that **identifies it as qualitative** and (recommended) names the approach (e.g. ethnography, grounded theory) or data-collection method (e.g. interviews, focus groups). |
17
+ | 2 | Abstract | A summary of the key elements in the journal's abstract format — typically background, purpose, methods, results, conclusions. |
18
+
19
+ ### Introduction
20
+ | # | Item | What to check is reported |
21
+ |---|------|---------------------------|
22
+ | 3 | Problem formulation | The problem/phenomenon studied and its significance, with a review of relevant theory and prior empirical work, and a problem statement. |
23
+ | 4 | Purpose or research question | The study's purpose and its specific objectives or questions. |
24
+
25
+ ### Methods
26
+ | # | Item | What to check is reported |
27
+ |---|------|---------------------------|
28
+ | 5 | Qualitative approach & research paradigm | The chosen approach (ethnography / grounded theory / case study / phenomenology / narrative) and the guiding paradigm (e.g. postpositivist, constructivist/interpretivist), **with a rationale**. |
29
+ | 6 | Researcher characteristics & **reflexivity** | The researchers' attributes, qualifications/experience, relationship with participants, and assumptions, and how these may have interacted with the questions, methods, findings, or transferability. |
30
+ | 7 | Context | The setting/site and salient contextual factors, with a rationale. |
31
+ | 8 | Sampling strategy | How and why participants/documents/events were selected, and the criterion for stopping sampling (e.g. **saturation**), with a rationale. |
32
+ | 9 | Ethical issues (human subjects) | Ethics-board approval and participant consent (or an explanation for their absence), and confidentiality / data-security handling. |
33
+ | 10 | Data collection methods | The data types and collection procedures — dates, iterative process, triangulation of sources/methods, and any procedure changes during the study — with a rationale. |
34
+ | 11 | Data collection instruments & technologies | The instruments (e.g. interview guides, questionnaires) and devices (e.g. audio recorders), and whether/how they changed over the study. |
35
+ | 12 | Units of study | The number and relevant characteristics of the participants/documents/events, and their level of participation. |
36
+ | 13 | Data processing | How data were processed before/during analysis — transcription, data entry/management/security, integrity checks, coding, and de-identification of excerpts. |
37
+ | 14 | Data analysis | The process by which inferences/themes were identified and developed, who analysed the data, and the referenced paradigm/approach, with a rationale. |
38
+ | 15 | Techniques to enhance **trustworthiness** | The techniques used to enhance trustworthiness/credibility (e.g. member checking, audit trail, triangulation), with a rationale. |
39
+
40
+ ### Results / Findings
41
+ | # | Item | What to check is reported |
42
+ |---|------|---------------------------|
43
+ | 16 | Synthesis & interpretation | The main findings (interpretations, inferences, themes), and any theory/model development or integration with prior research. |
44
+ | 17 | Links to empirical data | Evidence — participant **quotations**, field notes, text excerpts, images — that substantiates each analytic finding. |
45
+
46
+ ### Discussion
47
+ | # | Item | What to check is reported |
48
+ |---|------|---------------------------|
49
+ | 18 | Integration, implications, transferability, contribution | A short summary of findings; how they connect to / extend / challenge prior scholarship; the scope of application / **transferability**; and the study's unique contribution. |
50
+ | 19 | Limitations | The trustworthiness and limitations of the findings. |
51
+
52
+ ### Other
53
+ | # | Item | What to check is reported |
54
+ |---|------|---------------------------|
55
+ | 20 | Conflicts of interest | Potential sources of influence on the study and how they were managed. |
56
+ | 21 | Funding | Sources of funding/support and the role of the funders in data collection, interpretation, and reporting. |
57
+
58
+ ---
59
+
60
+ ## Notes for Assessors
61
+
62
+ - The **highest-yield** checks (where qualitative studies most often go wrong on review): **item 6** (researcher **reflexivity** — the researchers' position, assumptions, and relationship to participants is a core qualitative-rigor requirement, frequently omitted), **item 8** (a stated **sampling strategy and stopping criterion** — purposive logic and saturation, not a convenience sample with no rationale), **item 15** (explicit **trustworthiness** techniques — member checking / audit trail / triangulation), and **item 17** (analytic claims grounded in **quoted data**, not asserted).
63
+ - **Do not apply quantitative criteria** to a qualitative study: a small purposive sample is not a flaw, "generalizability" is **transferability** (not statistical external validity), and there are no power calculations, effect sizes, or p-values to demand. Mis-calibrated "the sample is too small / not generalizable / no p-value" comments are inappropriate here.
64
+ - This is an **in-house faithful summary of the SRQR items (paraphrased intents, not verbatim)**; map the manuscript's content to the items rather than to exact wording, and verify against the published statement (O'Brien et al. *Acad Med* 2014; DOI 10.1097/ACM.0000000000000388). For interview/focus-group designs, COREQ (`COREQ.md`) gives finer-grained items. Verified 2026-06-30.
@@ -87,6 +87,9 @@ ALIAS_TO_STEM = {
87
87
  "cheers2022": "CHEERS_2022",
88
88
  "record": "RECORD",
89
89
  "cross": "CROSS",
90
+ "srqr": "SRQR",
91
+ "coreq": "COREQ",
92
+ "qualitative": "SRQR",
90
93
  }
91
94
 
92
95
  EXIT_OK = 0
@@ -1,7 +1,7 @@
1
1
  # Reporting Guideline → Figure Requirements Map
2
2
 
3
3
  > **Bridge**: this file connects `/make-figures` to `/check-reporting`
4
- > (42 reporting guidelines). Each row tells you which figures the guideline
4
+ > (44 reporting guidelines). Each row tells you which figures the guideline
5
5
  > **mandates** and how this skill currently supports them. Use during
6
6
  > Step 1 (Specify) once the study type is known.
7
7
 
@@ -96,7 +96,7 @@ flags them:
96
96
 
97
97
  ## Cross-references
98
98
 
99
- - `/check-reporting` skill — supports all 42 guidelines, item-level audit
99
+ - `/check-reporting` skill — supports all 44 guidelines, item-level audit
100
100
  - `flow_diagram_lessons.md` — production lessons that apply across all flows
101
101
  - `pipeline_concepts_medical_ai.md` — DICOM / annotation / federated /
102
102
  architecture diagram conventions
@@ -48,7 +48,7 @@ You do NOT do the work yourself. You classify, plan, and delegate.
48
48
  | **meta-analysis** | Systematic review | Full MA pipeline: protocol, search, screening, extraction, synthesis, PRISMA-DTA |
49
49
  | **write-paper** | Writing | IMRAD manuscript drafting (8-phase pipeline), any section writing |
50
50
  | **self-review** | Quality | Pre-submission self-check with domain probes (Survival / SR-MA / Radiomics / Narrative); optional `--panel` for a high-stakes final QC pass |
51
- | **check-reporting** | Compliance | Audit against 42 reporting guidelines and risk-of-bias tools |
51
+ | **check-reporting** | Compliance | Audit against 44 reporting guidelines and risk-of-bias tools |
52
52
  | **revise** | Revision | Parse reviewer comments, generate point-by-point response, track changes |
53
53
  | **grant-builder** | Funding | Structure grant proposals: significance, innovation, approach, milestones |
54
54
  | **present-paper** | Presentation | Prepare academic talks: analyze paper, draft scripts, inject slide notes, Q&A prep |
@@ -292,6 +292,12 @@ Apply this 8-probe checklist (SC1–SC8) **only when the manuscript is a scoping
292
292
 
293
293
  **Probe detail (SC1–SC8), with output templates and the leads-vs-findings discipline:** `${CLAUDE_SKILL_DIR}/references/domain-probes/scoping_review.md`. Load it and apply each probe when the trigger fires. In this skill, map each probe finding to a Major / Minor comment; a focused effectiveness/accuracy question run as a scoping review to sidestep risk-of-bias and synthesis (SC1), or a scoping review reporting a pooled effect/accuracy estimate or definitive effectiveness conclusion (SC7) are design-level — surface them in the Confidential Comments to the Editor and place the strongest as the Major #1 candidate. Note the **asymmetric critical-appraisal calibration** (SC6): do **not** flag "no risk-of-bias assessment" as a deficiency for a scoping review, but do flag GRADE-style certainty claimed without appraisal. Wrong-registry (PROSPERO does not register scoping reviews) claims (SC2), narrow-search comprehensiveness claims (SC4), undocumented charting (SC5), and practice recommendations or mislabelling drawn from a map (SC8) are validity/framing-level.
294
294
 
295
+ ### Phase 2T: Qualitative Study Extension
296
+
297
+ Apply this 8-probe checklist (QL1–QL8) **only when the manuscript is a qualitative study** — in-depth interviews, focus groups, observation/ethnography, document analysis, grounded theory, phenomenology, narrative research. These probes complement (do not replace) the generic Phase 2 checklist and the qualitative reporting standards — **COREQ** (interviews/focus groups; Tong et al. 2007) and **SRQR** (all qualitative approaches; O'Brien et al. 2014). They target what makes qualitative rigour distinct from quantitative validity: researcher **reflexivity**, a transparent **analysis** process, **trustworthiness** (credibility/dependability/confirmability/transferability) rather than statistical validity, and findings **grounded in quoted data**.
298
+
299
+ **Probe detail (QL1–QL8), with output templates and the leads-vs-findings discipline:** `${CLAUDE_SKILL_DIR}/references/domain-probes/qualitative_research.md`. Load it and apply each probe when the trigger fires. In this skill, map each probe finding to a Major / Minor comment; a method–question mismatch (a quantitative question answered with a few interviews, QL1), absent **reflexivity** (QL2), an opaque "themes emerged" analysis with no coding process / audit trail (QL5), or interpretation not traceable to quoted data (QL7) are design-level — surface them in the Confidential Comments to the Editor and place the strongest as the Major #1 candidate. Note the **bidirectional calibration trap** (QL6): do **not** demand a power calculation, a "representative" sample, statistical generalizability, or treat inter-coder κ as the sole truth — these are quantitative yardsticks inappropriate to qualitative work (a small purposive sample is not a flaw; "generalizability" is **transferability**); but do flag authors who claim statistical generalizability or causal/prevalence/population over-reach (QL8) from qualitative data. Unjustified sampling / no saturation (QL3), thin data-collection reporting (QL4), and missing ethics/consent for identifiable quotes (QL8) are validity/framing-level. Map the study to **COREQ** (interviews/focus groups) or **SRQR** (broader).
300
+
295
301
  ### Phase 3: Draft Review
296
302
 
297
303
  Before writing comments, skim the relevant model in `references/exemplar_reviews/` for the
@@ -0,0 +1,50 @@
1
+ <!-- Domain probe module — shared, vendored BYTE-IDENTICAL by /peer-review and /self-review.
2
+ Severity words below (MAJOR / MINOR / major / minor) denote finding severity, NOT a journal
3
+ recommendation. Each consuming skill maps findings to its own output:
4
+ - peer-review: Major / Minor comments + Confidential Comments to the Editor; a
5
+ reflexivity / analysis-transparency / over-claim flaw is placed as Major #1.
6
+ - self-review: Anticipated Major / Minor Comments (Fatal / Fixable) mapped to category letters.
7
+ Do NOT edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`. -->
8
+
9
+ # Qualitative research probes (QL1–QL8)
10
+
11
+ An 8-probe checklist for **qualitative studies** — in-depth interviews, focus groups, observation/ethnography, document analysis, grounded theory, phenomenology, narrative research. These probes complement (do not replace) the generic Phase 2 checklist and the qualitative reporting standards — **COREQ** (interviews/focus groups; Tong et al. 2007) and **SRQR** (all qualitative approaches; O'Brien et al. 2014). They target what makes qualitative rigour distinct from quantitative validity: researcher **reflexivity**, a transparent **analysis** process, **trustworthiness** (credibility/dependability/confirmability/transferability) rather than statistical validity, and findings **grounded in quoted data** — and they guard the most common mis-calibration on both sides: applying quantitative yardsticks (sample-size power, statistical generalizability, p-values) to qualitative work. QL2 (reflexivity), QL5 (analysis transparency), and QL6 (trustworthiness, not statistical validity) are the highest-yield; run them first.
12
+
13
+ **QL1 — Approach/paradigm fit and research question**:
14
+ - Is a **qualitative** approach justified for the question — exploring meaning, experience, process, or context, rather than measuring frequency/effect (which would be a quantitative design)?
15
+ - Is a **named approach** (grounded theory, phenomenology, ethnography, case study, narrative) and guiding paradigm stated **with a rationale**, rather than an unspecified "qualitative study" / "thematic analysis" with no methodological orientation?
16
+ - A question that is really quantitative (prevalence/effect) answered with a few interviews, or no named methodological orientation behind the analysis → MAJOR (method–question mismatch) / MINOR (unnamed approach).
17
+
18
+ **QL2 — Reflexivity and researcher positioning**:
19
+ - Are the **researcher's characteristics** reported — who collected the data, their credentials/role, experience/training, and **prior relationship to participants** — and is there reflexive consideration of how their assumptions/position may have shaped data collection and interpretation?
20
+ - Reflexivity is a core qualitative-rigour requirement (COREQ Domain 1; SRQR item 6) and the **single most-omitted** element.
21
+ - An interview/focus-group study with no reflexivity / no statement of who interviewed and their relationship to participants → MAJOR (or MINOR if partially addressed).
22
+
23
+ **QL3 — Sampling logic and adequacy (purposive; information power / saturation)**:
24
+ - Is the **sampling strategy** stated and justified — **purposive / theoretical / maximum-variation** logic appropriate to the question, not an unexplained convenience sample — and is there an argument for **when sampling stopped** (data **saturation** or information power)?
25
+ - A **small sample is not a flaw** in qualitative work; the flaw is an *unjustified* sample with no purposive rationale and no saturation/information-power discussion.
26
+ - Convenience sample presented with no rationale, or no account of sampling adequacy/saturation behind broad thematic claims → MINOR / MAJOR per centrality.
27
+
28
+ **QL4 — Data-collection rigour**:
29
+ - Are the **data-collection methods** described in enough detail to judge them — the **interview/topic guide** (and whether it was piloted/iterated), the **setting**, **recording and transcription**, **field notes**, and interview/focus-group **duration**?
30
+ - Thinly reported data collection (no guide, unclear recording/transcription, no setting) that undercuts interpretability → MINOR.
31
+
32
+ **QL5 — Analysis transparency and audit trail**:
33
+ - Is the **analysis process** transparent — how many **coders**, the **coding framework/tree**, whether themes were **derived inductively or applied a priori**, any **software**, and an **audit trail**?
34
+ - "Themes **emerged** from the data" with no described coding/analytic process is a black box.
35
+ - An analysis with no described coding process / no audit trail behind the reported themes → MAJOR (analysis not reproducible/auditable).
36
+
37
+ **QL6 — Trustworthiness, NOT statistical validity (the calibration trap)**:
38
+ - Are **trustworthiness** techniques reported and matched to the four criteria — **credibility** (member checking, triangulation, prolonged engagement), **dependability** (audit trail), **confirmability** (reflexivity), **transferability** (thick description) — rather than quantitative "reliability/validity"?
39
+ - The trap is bidirectional: (a) a **reviewer** must **not** demand a power calculation, a "representative" sample, statistical generalizability, or treat inter-coder κ as the sole truth — these are quantitative yardsticks inappropriate to qualitative work; (b) **authors** must **not** claim statistical generalizability or dress qualitative findings in quantitative certainty.
40
+ - Missing trustworthiness strategies entirely, or quantitative-validity language misapplied (by either side) → MAJOR / MINOR per how load-bearing it is.
41
+
42
+ **QL7 — Findings grounded in data (quotations, thick description, deviant cases)**:
43
+ - Are the themes **substantiated by participant quotations / excerpts** (with participant identifiers), with enough **thick description** to let the reader judge the interpretation, and is there **consistency between the data and the findings**?
44
+ - Are **negative / deviant cases** and minor themes considered, not just confirmatory exemplars?
45
+ - Asserted themes with no quoted evidence, or interpretation not traceable to the data → MAJOR; cherry-picked confirmatory quotes with no deviant-case consideration → MINOR.
46
+
47
+ **QL8 — Ethics, interpretive scope, and reporting standard**:
48
+ - Are **ethics** reported (IRB approval/consent; confidentiality and de-identification of identifiable narrative quotes), and does interpretation **stay within what qualitative data support** — no **causal, effectiveness, prevalence, or population-level** claims, and no over-generalisation beyond the studied context (**transferability**, not generalizability)?
49
+ - Is the study mapped to the appropriate reporting standard — **COREQ** (interviews/focus groups) or **SRQR** (broader qualitative)?
50
+ - Causal/quantitative/population over-claiming from qualitative data, missing consent/de-identification for identifiable quotes, or no reporting-standard mapping → MAJOR (over-claim / ethics) / MINOR (reporting).
@@ -360,6 +360,7 @@ These modules carry the same domain-specific critique probes used by `/peer-revi
360
360
  | Observational study using routinely-collected health data (administrative claims / EHR / disease or population registry / health-checkup DB, linked or not) | `references/domain-probes/record_routinely_collected_data.md` (RD1–RD8) |
361
361
  | Self-report survey / questionnaire study (KAP, physician/patient survey, cross-sectional questionnaire, web/e-survey) | `references/domain-probes/survey_research.md` (SV1–SV8) |
362
362
  | Scoping review (maps the breadth/nature of evidence, clarifies concepts, identifies gaps; PCC framing, charting, optional appraisal — not a focused effectiveness/accuracy question) | `references/domain-probes/scoping_review.md` (SC1–SC8) |
363
+ | Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis; reflexivity, trustworthiness, thematic analysis — not quantitative validity) | `references/domain-probes/qualitative_research.md` (QL1–QL8) |
363
364
 
364
365
  When the manuscript matches a row, read `${CLAUDE_SKILL_DIR}/references/domain-probes/<module>.md` and apply each probe as an additional source of Anticipated Major / Minor Comments. The module severity words (MAJOR / MINOR) map to this skill's framing as follows: a conclusion-threatening or design-level finding becomes a **Fatal** Anticipated Major Comment, a reporting-level finding becomes a **Fixable** Anticipated Minor Comment, and each is tagged with the closest category letter (A–K). These probes **complement** categories A–K above; they do not replace them. (The modules are vendored byte-identical from `/peer-review`; do not edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`.)
365
366
 
@@ -0,0 +1,50 @@
1
+ <!-- Domain probe module — shared, vendored BYTE-IDENTICAL by /peer-review and /self-review.
2
+ Severity words below (MAJOR / MINOR / major / minor) denote finding severity, NOT a journal
3
+ recommendation. Each consuming skill maps findings to its own output:
4
+ - peer-review: Major / Minor comments + Confidential Comments to the Editor; a
5
+ reflexivity / analysis-transparency / over-claim flaw is placed as Major #1.
6
+ - self-review: Anticipated Major / Minor Comments (Fatal / Fixable) mapped to category letters.
7
+ Do NOT edit one copy only — run `python3 scripts/check_domain_probe_sync.py --sync`. -->
8
+
9
+ # Qualitative research probes (QL1–QL8)
10
+
11
+ An 8-probe checklist for **qualitative studies** — in-depth interviews, focus groups, observation/ethnography, document analysis, grounded theory, phenomenology, narrative research. These probes complement (do not replace) the generic Phase 2 checklist and the qualitative reporting standards — **COREQ** (interviews/focus groups; Tong et al. 2007) and **SRQR** (all qualitative approaches; O'Brien et al. 2014). They target what makes qualitative rigour distinct from quantitative validity: researcher **reflexivity**, a transparent **analysis** process, **trustworthiness** (credibility/dependability/confirmability/transferability) rather than statistical validity, and findings **grounded in quoted data** — and they guard the most common mis-calibration on both sides: applying quantitative yardsticks (sample-size power, statistical generalizability, p-values) to qualitative work. QL2 (reflexivity), QL5 (analysis transparency), and QL6 (trustworthiness, not statistical validity) are the highest-yield; run them first.
12
+
13
+ **QL1 — Approach/paradigm fit and research question**:
14
+ - Is a **qualitative** approach justified for the question — exploring meaning, experience, process, or context, rather than measuring frequency/effect (which would be a quantitative design)?
15
+ - Is a **named approach** (grounded theory, phenomenology, ethnography, case study, narrative) and guiding paradigm stated **with a rationale**, rather than an unspecified "qualitative study" / "thematic analysis" with no methodological orientation?
16
+ - A question that is really quantitative (prevalence/effect) answered with a few interviews, or no named methodological orientation behind the analysis → MAJOR (method–question mismatch) / MINOR (unnamed approach).
17
+
18
+ **QL2 — Reflexivity and researcher positioning**:
19
+ - Are the **researcher's characteristics** reported — who collected the data, their credentials/role, experience/training, and **prior relationship to participants** — and is there reflexive consideration of how their assumptions/position may have shaped data collection and interpretation?
20
+ - Reflexivity is a core qualitative-rigour requirement (COREQ Domain 1; SRQR item 6) and the **single most-omitted** element.
21
+ - An interview/focus-group study with no reflexivity / no statement of who interviewed and their relationship to participants → MAJOR (or MINOR if partially addressed).
22
+
23
+ **QL3 — Sampling logic and adequacy (purposive; information power / saturation)**:
24
+ - Is the **sampling strategy** stated and justified — **purposive / theoretical / maximum-variation** logic appropriate to the question, not an unexplained convenience sample — and is there an argument for **when sampling stopped** (data **saturation** or information power)?
25
+ - A **small sample is not a flaw** in qualitative work; the flaw is an *unjustified* sample with no purposive rationale and no saturation/information-power discussion.
26
+ - Convenience sample presented with no rationale, or no account of sampling adequacy/saturation behind broad thematic claims → MINOR / MAJOR per centrality.
27
+
28
+ **QL4 — Data-collection rigour**:
29
+ - Are the **data-collection methods** described in enough detail to judge them — the **interview/topic guide** (and whether it was piloted/iterated), the **setting**, **recording and transcription**, **field notes**, and interview/focus-group **duration**?
30
+ - Thinly reported data collection (no guide, unclear recording/transcription, no setting) that undercuts interpretability → MINOR.
31
+
32
+ **QL5 — Analysis transparency and audit trail**:
33
+ - Is the **analysis process** transparent — how many **coders**, the **coding framework/tree**, whether themes were **derived inductively or applied a priori**, any **software**, and an **audit trail**?
34
+ - "Themes **emerged** from the data" with no described coding/analytic process is a black box.
35
+ - An analysis with no described coding process / no audit trail behind the reported themes → MAJOR (analysis not reproducible/auditable).
36
+
37
+ **QL6 — Trustworthiness, NOT statistical validity (the calibration trap)**:
38
+ - Are **trustworthiness** techniques reported and matched to the four criteria — **credibility** (member checking, triangulation, prolonged engagement), **dependability** (audit trail), **confirmability** (reflexivity), **transferability** (thick description) — rather than quantitative "reliability/validity"?
39
+ - The trap is bidirectional: (a) a **reviewer** must **not** demand a power calculation, a "representative" sample, statistical generalizability, or treat inter-coder κ as the sole truth — these are quantitative yardsticks inappropriate to qualitative work; (b) **authors** must **not** claim statistical generalizability or dress qualitative findings in quantitative certainty.
40
+ - Missing trustworthiness strategies entirely, or quantitative-validity language misapplied (by either side) → MAJOR / MINOR per how load-bearing it is.
41
+
42
+ **QL7 — Findings grounded in data (quotations, thick description, deviant cases)**:
43
+ - Are the themes **substantiated by participant quotations / excerpts** (with participant identifiers), with enough **thick description** to let the reader judge the interpretation, and is there **consistency between the data and the findings**?
44
+ - Are **negative / deviant cases** and minor themes considered, not just confirmatory exemplars?
45
+ - Asserted themes with no quoted evidence, or interpretation not traceable to the data → MAJOR; cherry-picked confirmatory quotes with no deviant-case consideration → MINOR.
46
+
47
+ **QL8 — Ethics, interpretive scope, and reporting standard**:
48
+ - Are **ethics** reported (IRB approval/consent; confidentiality and de-identification of identifiable narrative quotes), and does interpretation **stay within what qualitative data support** — no **causal, effectiveness, prevalence, or population-level** claims, and no over-generalisation beyond the studied context (**transferability**, not generalizability)?
49
+ - Is the study mapped to the appropriate reporting standard — **COREQ** (interviews/focus groups) or **SRQR** (broader qualitative)?
50
+ - Causal/quantitative/population over-claiming from qualitative data, missing consent/de-identification for identifiable quotes, or no reporting-standard mapping → MAJOR (over-claim / ethics) / MINOR (reporting).