@pikaa-ai/pikaa 0.3.23 → 0.3.25

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +407 -219
  6. package/dist/index.js +7 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,222 +0,0 @@
1
- ---
2
- name: omero-integration
3
- description: Securely inspect and automate microscopy data workflows against OMERO.server with omero-py, BlitzGateway, OMERO CLI, tables, annotations, ROIs, rendering, and documented OMERO.web APIs. Use for scoped OMERO inventory, metadata export, import/export planning, or reviewed write workflows.
4
- license: MIT
5
- compatibility: >-
6
- Requires network access to a user-selected OMERO.server for remote operations.
7
- The 2026-07-23 snapshot uses OMERO.py 5.22.1 with ZeroC IcePy 3.6.5;
8
- OMERO supports Python 3.10-3.12 (3.12 recommended) while 3.13-3.14 remain
9
- upcoming in its support matrix. Bundled local planners require Python 3.10+
10
- and read only named OMERO_* variables; they never load .env files.
11
- metadata:
12
- version: "1.3"
13
- skill-author: K-Dense Inc.
14
- openclaw:
15
- envVars:
16
- - name: OMERO_HOST
17
- required: true
18
- description: OMERO.server hostname.
19
- - name: OMERO_PORT
20
- required: false
21
- description: OMERO SSL router port; default 4064.
22
- - name: OMERO_USER
23
- required: false
24
- description: Username when not reusing a session.
25
- - name: OMERO_PASSWORD
26
- required: false
27
- description: Password when not reusing a session.
28
- - name: OMERO_SESSION_KEY
29
- required: false
30
- description: Existing session key as an alternative credential.
31
- - name: OMERO_SECURE
32
- required: false
33
- description: Secure transport toggle; default true.
34
- ---
35
-
36
- # OMERO Integration
37
-
38
- Use current OME documentation and the smallest explicit data scope. OMERO data
39
- may contain unpublished images, identifiers, annotations, original files, and
40
- derived measurements.
41
-
42
- ## Verified Baseline
43
-
44
- This skill was refreshed on **2026-07-23**:
45
-
46
- - **OMERO.server 5.6.18** (May 2026) is the current documented stable server.
47
- - It was tested by OME with **OMERO.py/omero-py 5.22.1** and
48
- **OMERO.web 5.31.0**.
49
- - `omero-py==5.22.1` requires Python 3.10 or newer. The OMERO support matrix
50
- supports 3.10 and 3.11, recommends 3.12, and still labels 3.13/3.14
51
- “upcoming.”
52
- - OMERO 5.6 uses **IcePy 3.6**, with 3.6.5 prebuilt client wheels documented
53
- for Python versions through 3.12.
54
-
55
- The pin above is a reproducible skill snapshot, not a promise that every
56
- OMERO.server release accepts that client. For another server version, consult
57
- its release entry and use the OMERO.py version tested with it. See
58
- [`references/sources.md`](references/sources.md).
59
-
60
- ## Operating Contract
61
-
62
- 1. Start with local validation or a dry run. Do not connect until the user has
63
- selected the host, group, object type, IDs, and result limit.
64
- 2. Read credentials only from the named `OMERO_*` variables in the frontmatter.
65
- Never search parent directories or load `.env` files.
66
- 3. Never place a password or session key in command arguments, source code,
67
- output JSON, logs, tracebacks, or chat. A session key is a bearer credential.
68
- 4. Default to `secure=True`. OMERO encrypts login by default, but post-login
69
- data and the session ID may otherwise travel unencrypted. `secure=True` does
70
- not by itself guarantee certificate hostname verification.
71
- 5. Bound every list, page, ROI, shape, annotation, table row, pixel plane, and
72
- local file scan. Do not turn an object request into a group-wide or
73
- cross-group export without explicit approval.
74
- 6. Treat all writes separately: annotation/link creation, rendering-default
75
- saves, image creation, imports, script uploads, table writes, ownership or
76
- group changes, and deletion require an exact reviewed target.
77
- 7. Close `BlitzGateway`, table handles, raw stores, thumbnail stores, rendering
78
- engines, script clients, and other stateful services in `finally` blocks or
79
- documented context-manager patterns.
80
- 8. Never connect to a real server merely to “test” examples.
81
-
82
- ## Choose the Interface
83
-
84
- - **BlitzGateway (`omero-py`)**: primary Python client for object traversal,
85
- pixels, annotations, ROIs, rendering, and services.
86
- - **OMERO CLI**: sessions, import scanning/import, OME-TIFF or XML export,
87
- scripts, and administrative plugins. Most client commands are remote; import
88
- also needs the matching server-side Java libraries through `OMERODIR`.
89
- - **OMERO.web `api` and `webgateway`**: the only OMERO.web apps that official
90
- documentation calls stable public APIs. The documented JSON API is
91
- version-discovered and has limited object coverage; it is not evidence that
92
- every webclient URL is a supported REST endpoint.
93
- - **OMERO.server scripts**: uploaded plugins executed by server infrastructure.
94
- They are different from the bundled local client helpers in `scripts/`.
95
-
96
- ## Install a Reproducible Client
97
-
98
- Create a Python 3.12 environment:
99
-
100
- ```bash
101
- uv venv --python 3.12 .venv
102
- source .venv/bin/activate
103
- ```
104
-
105
- Install the exact IcePy 3.6.5 wheel matching the interpreter, OS, architecture,
106
- and wheel tags, then OMERO.py:
107
-
108
- ```bash
109
- # Download the matching 3.6.5 wheel from the official OMERO-linked matrix.
110
- uv pip install "/absolute/path/to/zeroc_ice-3.6.5-<matching-tags>.whl"
111
- uv pip install "omero-py==5.22.1"
112
- ```
113
-
114
- Do not substitute Ice 3.7: the OMERO 5.6 support matrix marks Ice 3.6 as
115
- recommended and 3.7 as unsupported. A plain install may attempt to compile
116
- IcePy from source; prefer a reviewed matching wheel. The upstream package is
117
- GPL-2.0-or-later; this skill’s own files are MIT.
118
-
119
- For import/admin commands only, `OMERODIR` must point to a compatible extracted
120
- OMERO.server directory. A normal remote BlitzGateway client does not require
121
- that server tree. Read [`references/connection.md`](references/connection.md)
122
- before installation or authentication work.
123
-
124
- ## Credentials and Connection
125
-
126
- Set named variables in the calling environment or secret manager. Do not put
127
- the password on an `omero` CLI command:
128
-
129
- ```bash
130
- export OMERO_HOST="omero.example.org"
131
- export OMERO_PORT="4064"
132
- export OMERO_USER="researcher"
133
- export OMERO_SECURE="true"
134
- # Supply OMERO_PASSWORD through the environment/secret manager, or use
135
- # OMERO_SESSION_KEY as an alternative. Do not echo either value.
136
- ```
137
-
138
- A password-authenticated, exception-safe read pattern is:
139
-
140
- ```python
141
- import os
142
- from omero.gateway import BlitzGateway
143
-
144
- conn = None
145
- try:
146
- conn = BlitzGateway(
147
- os.environ["OMERO_USER"],
148
- os.environ["OMERO_PASSWORD"],
149
- host=os.environ["OMERO_HOST"],
150
- port=int(os.environ.get("OMERO_PORT", "4064")),
151
- secure=True,
152
- )
153
- if not conn.connect():
154
- raise RuntimeError("OMERO connection failed")
155
-
156
- images = conn.getObjects(
157
- "Image",
158
- opts={"limit": 25, "offset": 0, "order_by": "obj.id"},
159
- )
160
- for image in images:
161
- print(image.getId()) # Do not print names unless requested.
162
- finally:
163
- if conn is not None:
164
- conn.close()
165
- ```
166
-
167
- For existing-session and CLI prompt patterns, certificate verification,
168
- group context, and cleanup details, read
169
- [`references/connection.md`](references/connection.md).
170
-
171
- ## Bundled Safe Helpers
172
-
173
- All helpers use `argparse`; `--help` works without OMERO installed. Remote
174
- helpers are dry-run by default and require `--execute`.
175
-
176
- ```bash
177
- python -B scripts/validate_config.py --help
178
- python -B scripts/inventory.py --help
179
- python -B scripts/export_image_metadata.py --help
180
- python -B scripts/plan_transfer.py --help
181
- ```
182
-
183
- - `validate_config.py`: validates only named endpoint/auth variables locally;
184
- optional DNS resolution still does not contact OMERO.
185
- - `inventory.py`: bounded, read-only object inventory with paged JSON output.
186
- - `export_image_metadata.py`: explicit-image annotation/ROI JSON export with
187
- redaction defaults and per-category limits; it never downloads file bytes or
188
- pixels.
189
- - `plan_transfer.py`: local-only import scan or per-image export plan; it never
190
- invokes OMERO and never emits credential flags.
191
-
192
- Read [`references/scripts.md`](references/scripts.md) before using them.
193
-
194
- ## Capability Guide
195
-
196
- - Connection, sessions, groups, TLS:
197
- [`references/connection.md`](references/connection.md)
198
- - Hierarchies, pagination, screening data, import/export:
199
- [`references/data_access.md`](references/data_access.md)
200
- - Tags, map/file/comment annotations, namespaces:
201
- [`references/metadata.md`](references/metadata.md)
202
- - Raw planes, tiles, thumbnails, rendering:
203
- [`references/image_processing.md`](references/image_processing.md)
204
- - ROI model, shape export, statistics caveat:
205
- [`references/rois.md`](references/rois.md)
206
- - Bounded table creation, paging, querying, closure:
207
- [`references/tables.md`](references/tables.md)
208
- - Local helpers and OMERO.server scripts:
209
- [`references/scripts.md`](references/scripts.md)
210
- - Permissions, filesets, web/public links, destructive operations:
211
- [`references/advanced.md`](references/advanced.md)
212
-
213
- ## Final Review Before Remote Work
214
-
215
- - Confirm server version and its tested OMERO.py pairing.
216
- - Confirm target host, SSL router port, user/session, and one group.
217
- - Confirm exact object IDs/types and hard limits.
218
- - Confirm whether names, annotation values, file names, ROI labels, owner names,
219
- pixels, or original files may leave the server.
220
- - Show the proposed output path and refuse overwrite unless explicitly allowed.
221
- - For a write, show the mutation and target IDs separately from any read plan.
222
- - Close every connection/service even after partial failure.
@@ -1,371 +0,0 @@
1
- ---
2
- name: onekgpd
3
- description: >
4
- Query the 1000 Genomes Project dataset (3,202 whole-genome-sequenced
5
- individuals, GRCh38) at the level of individual participants.
6
- Use when a question is about individuals or variants in the 1000 Genomes
7
- Project cohort: which individuals carry variants matching specific criteria
8
- in a gene or region, which individuals are homozygous-reference at a position,
9
- which variants exist in the dataset or carried by specified individuals
10
- in a gene or region, the relatedness between two specified individuals.
11
- Variants are returned with 1000 Genomes allele frequencies (AF),
12
- gnomAD v4.1 exome and genome AF, AlphaMissense score, and HGVSp annotations.
13
- license: MIT
14
- compatibility: Requires Python >=3.11. Variant and sample queries require outbound network access to the public 1000 Genomes query endpoint over TLS; the sample/population metadata commands run fully offline over a data file bundled in the skill. No credentials, API keys, or environment variables are used.
15
- allowed-tools: Write Bash
16
- metadata:
17
- version: "1.2"
18
- skill-author: Dnaerys
19
- ---
20
-
21
- # OneKGPd: Individual-Level Queries over the 1000 Genomes Project
22
-
23
- ## Scope
24
-
25
- This skill queries the 1000 Genomes Project dataset — the extended high-coverage cohort
26
- of 3,202 whole-genome-sequenced individuals, on the GRCh38 assembly. All results
27
- are drawn from this cohort, and sample names returned by the skill (for example
28
- `HG00096` or `NA21130`) identify its participants.
29
-
30
- Queries resolve against the cohort's per-individual genotype data. This supports
31
- two complementary classes of question: selecting **variants** carried within a
32
- region (across the whole cohort or within a specified set of individuals), and
33
- selecting the **individuals** who carry variants matching given criteria.
34
- Variant selection can be filtered by allele frequency, predicted consequence,
35
- clinical significance, AlphaMissense classification, and the other annotation
36
- axes listed below. Relatedness between two named individuals is also available.
37
-
38
- The genotype state in which a variant is carried — heterozygous or homozygous —
39
- is a criterion that queries may specify; results are returned as variants or as
40
- sample names, not as raw genotypes.
41
-
42
- ## When to Use
43
-
44
- **Use this skill when you need to:**
45
-
46
- - Find **variants** carried in a region or set of regions matching some criteria
47
- across the whole cohort (`select-variants`).
48
- - Find **variants** carried in a region or set of regions matching some criteria
49
- in specific set of individuals (`select-variants-in-samples`).
50
- - Find **which 1000 Genomes individuals** carry variants matching some criteria
51
- in a region or set of regions (`select-samples`).
52
- - Count how many individuals carry specific variants (`count-samples`).
53
- - Restrict any variant query to **heterozygous-only or homozygous-only**
54
- carriage, or query both together (default).
55
- - Identify which individuals are **homozygous reference** at a single position
56
- (`select-samples-hom-ref`).
57
- - Determine the **relatedness** between two named 1000 Genomes individuals —
58
- both the degree (twin / 1st / 2nd / 3rd / unrelated) and the KING kinship
59
- coefficient (`kinship`).
60
- - Get **dataset totals** — sample count, sex split, variant count, assembly
61
- (`dataset-info`).
62
- - Variant selection can be specified by KGP allele frequency, gnomAD 4.1 exome and
63
- gnomAD 4.1 genome allele frequency, AlphaMissense Score and AlphaMissense Class,
64
- ClinVar significance (202502), and VEP annotations (impact, biotype, feature type,
65
- variant class, consequences).
66
-
67
- **Do NOT use this skill for:**
68
-
69
- - Resolving a gene symbol, rsID, or transcript to coordinates, or fetching
70
- reference sequence. Resolve coordinates first (see Coordinate Provenance
71
- below), then query this skill with the resolved GRCh38 region.
72
- - Any cohort other than the 1000 Genomes Project — this skill serves only that
73
- dataset.
74
-
75
- ## Prerequisites
76
-
77
- 1. **`uv`**: This skill's script is run with `uv run`, which reads the script's
78
- inline dependency metadata and provisions an ephemeral environment. Ensure
79
- `uv` is installed and on PATH (https://docs.astral.sh/uv/).
80
- 2. **Data use terms**: The 1000 Genomes Project data is open; users should be
81
- aware of the 1000 Genomes Project / IGSR data-use terms
82
- (https://www.internationalgenome.org/data).
83
- 3. **Access constraints**: There is no API key, no `.env` file, and no
84
- rate-limit token to configure.
85
- 4. **No credentials required**
86
-
87
- ## Core Rules
88
-
89
- - **Use the Wrappers**: ALWAYS execute the provided helper scripts rather than
90
- constructing your own client calls or network requests. Use
91
- `scripts/onekgpd_api.py` for variant/sample/kinship queries (it handles the
92
- connection, streaming, pagination, and JSON serialization), and
93
- `scripts/onekgpd_meta.py` for sample/population metadata (offline, see
94
- [Sample & population metadata](#sample--population-metadata-offline)).
95
- - **Coordinates MUST be resolved against an authoritative source first** — see
96
- [Coordinate Provenance](#coordinate-provenance-mandatory-first-step). This
97
- is mandatory, not advisory.
98
- - **Count before you select**: every variant and sample selection has a paired
99
- counting command. Call the count command FIRST to size the result set, then
100
- select only if the count is manageable.
101
- - **Zygosity defaults to both**: selection and counting commands include both
102
- heterozygous and homozygous carriage by default. Narrow with `--het-only`
103
- or `--hom-only` when the question is specifically about one state. (You do
104
- not need to pass anything to get both.)
105
- - **Output**: scripts write full JSON to a file (`--output`, default under
106
- `/tmp/`) and print a concise summary to stdout. Do not read large JSON files
107
- into context — use `jq` or a small disposable `uv run python` snippet to
108
- extract fields.
109
-
110
- ## Coordinate Provenance (MANDATORY FIRST STEP)
111
-
112
- Before any region-based query, resolve the gene or feature to **GRCh38**
113
- coordinates against an authoritative source (for example Ensembl), and query
114
- with those resolved coordinates. The assembly must be explicit, and a gene-range
115
- must be resolved to precise positions before use. This is structural, not
116
- advisory: there is no source-side guardrail that would catch a misplaced region,
117
- so an unverified coordinate produces results for an unintended location with no
118
- error.
119
-
120
- ```bash
121
- # Resolve gene symbol -> GRCh38 region with an authoritative source FIRST,
122
- # then pass the verified coordinates to the OneKGPd query below.
123
- ```
124
-
125
- > [!CAUTION]
126
- > The dataset is GRCh38. A GRCh37 coordinate, or any region that does not
127
- > correctly correspond to the intended feature on GRCh38, will return
128
- > results for an unintended location without raising an error. Verify the
129
- > assembly and the resolved coordinates before querying.
130
-
131
- ## Command Selection Guide
132
-
133
- Match the question to the command. Counting commands are cheap and should
134
- precede their selection counterpart.
135
-
136
- - Which individuals carry matching variants in a region → `count-samples`
137
- then `select-samples`
138
- - Which variants are carried in a region, cohort-wide → `count-variants`
139
- then `select-variants`
140
- - Which variants are carried in a region, within a named set of individuals →
141
- `count-variants-in-samples` then `select-variants-in-samples`
142
- - Who is homozygous-reference at a single position → `count-samples-hom-ref`
143
- then `select-samples-hom-ref`
144
- - Relatedness (degree + coefficient) between two named individuals →
145
- `kinship`
146
- - Dataset totals (sample count, sex split, variant total, assembly) →
147
- `dataset-info`
148
-
149
- ## Annotation filters (shared across variant and sample selection/counting)
150
-
151
- All variant- and sample-selection commands (`count-variants`,
152
- `select-variants`, their `-in-samples` forms, `count-samples`, `select-samples`)
153
- accept the same annotation filters. Different filter fields are combined with
154
- **AND**; multiple values within one field are combined with **OR**. Enum values
155
- are case-insensitive (e.g. `missense_variant` or `MISSENSE_VARIANT`).
156
-
157
- These are selection criteria applied on the server. The fields returned on a
158
- selected variant are listed under
159
- [Variant-returning commands](#variant-returning-commands); a criterion used for
160
- filtering is not necessarily echoed back on the returned variant.
161
-
162
- - `--af-lt` / `--af-gt`: 1000 Genomes dataset allele frequency bounds
163
- - `--gnomad-exomes-af-lt` / `--gnomad-exomes-af-gt`: gnomAD v4.1 exome AF bounds
164
- - `--gnomad-genomes-af-lt` / `--gnomad-genomes-af-gt`: gnomAD v4.1 genome AF bounds
165
- - `--clin-significance`: ClinVar significance terms, CSV (e.g. `PATHOGENIC,LIKELY_PATHOGENIC`)
166
- - `--consequence`: Sequence Ontology consequence terms, CSV (e.g. `MISSENSE_VARIANT,STOP_GAINED`)
167
- - `--impact`: VEP impact, CSV (`HIGH,MODERATE,LOW,MODIFIER`)
168
- - `--variant-type`, `--feature-type`, `--bio-type`: SO variant class / VEP feature / VEP biotype, CSV
169
- - `--alpha-missense-class`: `AM_LIKELY_BENIGN,AM_LIKELY_PATHOGENIC,AM_AMBIGUOUS` (CSV)
170
- - `--alpha-missense-score-lt` / `--alpha-missense-score-gt`: AlphaMissense score bounds
171
- - `--biallelic-only` / `--multiallelic-only`
172
- - `--exclude-males` / `--exclude-females`
173
- - `--min-len-bp` / `--max-len-bp`: alternate-allele length bounds (bp)
174
-
175
- > [!NOTE]
176
- > `--alpha-missense-class` and `--alpha-missense-score-*` are mutually exclusive
177
- > (the engine ignores the class when a score bound is set). `--biallelic-only`
178
- > and `--multiallelic-only` are mutually exclusive. `--exclude-males` and
179
- > `--exclude-females` are mutually exclusive. Setting a `*-gt` bound greater than
180
- > or equal to its matching `*-lt` bound defines an empty range and will return
181
- > nothing.
182
-
183
- > [!NOTE]
184
- > Allele-frequency fields use `0.0` to mean "not present in that source." So
185
- > `--gnomad-exomes-af-gt 0` selects variants that *are* in gnomAD exomes; a
186
- > returned `gnomad_exomes_af` of `0.0` means the variant is absent from gnomAD
187
- > exomes. The same convention for gnomAD genomes AF.
188
- > Conversely, `--gnomad-exomes-af-lt` / `--gnomad-genomes-af-lt` bounds **include**
189
- unannotated variants: "AF < X in gnomAD" includes variants with gnomAD AF = 0,
190
- i.e. unannotated; pair it with `--gnomad-*-af-gt 0` to require presence in gnomAD.
191
-
192
- > [!NOTE]
193
- > `am_score` of `0.0` means not scored or not annotated by AlphaMissense - it does not mean `benign`.
194
- > A real AlphaMissense score is always greater than 0.
195
-
196
- ## Quick Start
197
-
198
- ```bash
199
- # Step 1. Resolve coordinates against an authoritative source — see Coordinate Provenance.
200
- # example: BRCA1: chr17:43044292-43170245
201
- # Step 2. Size the result set: how many individuals carry predicted likely-pathogenic
202
- # missense variants in this region?
203
- uv run scripts/onekgpd_api.py count-samples \
204
- --chrom chr17 --start 43044292 --end 43170245 \
205
- --consequence MISSENSE_VARIANT \
206
- --alpha-missense-class AM_LIKELY_PATHOGENIC \
207
- --output /tmp/count.json
208
- # Step 3. If the count is manageable, list those individuals.
209
- uv run scripts/onekgpd_api.py select-samples \
210
- --chrom chr17 --start 43044292 --end 43170245 \
211
- --consequence MISSENSE_VARIANT \
212
- --alpha-missense-class AM_LIKELY_PATHOGENIC \
213
- --output /tmp/samples.json
214
- # Step 4: For that set of individuals, see the actual variants they carry.
215
- uv run scripts/onekgpd_api.py select-variants-in-samples \
216
- --chrom chr17 --start 43044292 --end 43170245 \
217
- --samples HG03169,NA20506 \
218
- --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
219
- --output /tmp/variants.json
220
- ```
221
-
222
- ## Commands
223
-
224
- Each command writes full JSON to a file (`--output PATH`, default a temp file)
225
- and prints a concise stdout summary. All region/sample commands share: the
226
- region input (`--chrom`/`--start`/`--end` with optional `--ref`/`--alt`, or one
227
- or more repeated `--region CHR:START-END`), the zygosity flags
228
- (`--het-only`/`--hom-only`, default both), and the annotation filters above.
229
- The full per-flag tables live in
230
- [references/onekgpd_commands.md](references/onekgpd_commands.md).
231
-
232
- ### Variant-returning commands
233
-
234
- `select-*` return matching variants; `count-*` return an integer count.
235
-
236
- - `count-variants` — count variants in a region, cohort-wide.
237
- - `select-variants` — select variants in a region, cohort-wide. Use `--limit N`
238
- (hard cap, default 200) **or** `--page-size N` (retrieve the full set in
239
- pages); the two are mutually exclusive. The summary flags `truncated` when
240
- the cap is reached.
241
- - `count-variants-in-samples` — as `count-variants`, restricted to
242
- `--samples NAME1,NAME2,...` (required).
243
- - `select-variants-in-samples` — as `select-variants`, restricted to
244
- `--samples NAME1,NAME2,...` (required).
245
-
246
- Each returned variant carries these 22 keys: `chr`, `start`, `end`, `ref`,
247
- `alt`, `af`, `ac`, `an`, `hom_samples`, `het_samples`, `mis_samples`,
248
- `hom_samples_fx`, `het_samples_fx`, `mis_samples_fx`, `hom_samples_mxy`,
249
- `het_samples_mxy`, `mis_samples_mxy`, `gnomad_exomes_af`, `gnomad_genomes_af`,
250
- `am_score`, `amino_acids`, `biallelic`.
251
- ClinVar significance and VEP consequence are filter criteria only and are not
252
- returned. Full schema:
253
- [references/onekgpd_commands.md](references/onekgpd_commands.md).
254
-
255
- ### Sample-returning commands
256
-
257
- - `count-samples` — count individuals carrying a matching variant in a region.
258
- - `select-samples` — list the names of individuals carrying a matching variant.
259
- Supports `--skip N` and `--limit N`. Returns names only; to see which
260
- variants qualified an individual, feed the names into
261
- `select-variants-in-samples`.
262
-
263
- ### Homozygous-reference commands
264
-
265
- Single position via `--chrom` + `--position` (not a region).
266
-
267
- - `count-samples-hom-ref` — count individuals with a 0/0 call at the position.
268
- The count is a sentinel: `-1` = no variant exists at that position at all;
269
- `0` = a variant exists but no individual is homozygous reference; `>0` = the
270
- number of homozygous-reference individuals. The summary states which case.
271
- - `select-samples-hom-ref` — list the individuals with a 0/0 call at the position.
272
-
273
- ### Relatedness command
274
-
275
- - `kinship --sample1 NAME --sample2 NAME` — relatedness between two named
276
- individuals: the degree (`TWINS_MONOZYGOTIC` / `FIRST_DEGREE` /
277
- `SECOND_DEGREE` / `THIRD_DEGREE` / `UNRELATED`) and the KING kinship
278
- coefficient (`phi_bwf`).
279
-
280
- ### Dataset metadata command
281
-
282
- - `dataset-info` — dataset totals: `samples_total` (3,202), female/male split,
283
- `variants_total`, `assembly` (GRCh38), and the cohort breakdown. No region
284
- required; doubles as a connectivity check.
285
-
286
- ## Sample & population metadata (offline)
287
-
288
- Population, sex, pedigree, and superpopulation questions are answered by a second
289
- script, `scripts/onekgpd_meta.py`, from a data file bundled in the skill — **no
290
- network, no credentials, no coordinates**. The sample IDs are the same names the
291
- variant commands use, so the two layers compose (e.g. pick a cohort by population,
292
- then query its variants). Run `uv run scripts/onekgpd_meta.py <command>`.
293
-
294
- The cohort has 5 superpopulations (`AFR`, `AMR`, `EAS`, `EUR`, `SAS`) and 26
295
- populations. Population/superpopulation values match **case-insensitively** by
296
- short code or full name; **sample IDs are case-sensitive**.
297
-
298
- - `sample-metadata --samples NA19240,HG00096` — family, gender, parents,
299
- children, population, superpopulation, and phase3 status for the given samples.
300
- - `list-populations` — all 26 populations with superpopulation and sample count
301
- (use to discover valid values).
302
- - `list-superpopulations` — the 5 superpopulations with sample count and
303
- constituent populations.
304
- - `population-stats --populations YRI [--populations CHS …]` — per-population sex
305
- split, phase3 count, and trio membership. Repeat `--populations` for multiple
306
- values (full names contain commas, so they are not comma-separated).
307
- - `superpopulation-summary --superpopulations EAS [--superpopulations EUR …]` —
308
- per-superpopulation totals with a per-population breakdown.
309
- - `select-samples-by-population --population YRI` and/or `--superpopulation AFR`,
310
- with optional `--skip`/`--limit` (default 0 / 50, max 3202) — the sample IDs in
311
- a population and/or superpopulation; both given intersects. Feed the names into
312
- `select-variants-in-samples` to see their variants.
313
-
314
- See [references/onekgpd_commands.md](references/onekgpd_commands.md) for full
315
- argument tables and JSON output schemas.
316
-
317
- ## Typical Workflows
318
-
319
- ### Which individuals, then which variants they carry
320
-
321
- ```bash
322
- # Step 1: resolve gene -> verified GRCh38 region (authoritative source).
323
- # Step 2: count individuals carrying a qualifying variant in the region.
324
- uv run scripts/onekgpd_api.py count-samples \
325
- --chrom <chr> --start <start> --end <end> \
326
- --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
327
- --output /tmp/n.json
328
- # Step 3: list those individuals.
329
- uv run scripts/onekgpd_api.py select-samples \
330
- --chrom <chr> --start <start> --end <end> \
331
- --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
332
- --output /tmp/who.json
333
- # Step 4: for that set of individuals, see the actual variants they carry.
334
- uv run scripts/onekgpd_api.py select-variants-in-samples \
335
- --chrom <chr> --start <start> --end <end> \
336
- --samples <name1,name2,...> \
337
- --consequence MISSENSE_VARIANT --alpha-missense-class AM_LIKELY_PATHOGENIC \
338
- --output /tmp/variants.json
339
- ```
340
-
341
- ### Homozygous-reference carriers at a position of interest
342
-
343
- ```bash
344
- # After identifying a position of interest (verified coordinate):
345
- uv run scripts/onekgpd_api.py count-samples-hom-ref \
346
- --chrom <chr> --position <pos> --output /tmp/homref_n.json
347
- uv run scripts/onekgpd_api.py select-samples-hom-ref \
348
- --chrom <chr> --position <pos> --output /tmp/homref.json
349
- ```
350
-
351
- ## Common Mistakes
352
-
353
- - **Mistake:** Querying with an unverified coordinate.
354
- **Fix:** Always resolve gene/feature → GRCh38 against an authoritative
355
- source first.
356
- A misplaced region returns results for an unintended location without error.
357
- - **Mistake:** Calling a selection command before its counting command.
358
- **Fix:** Count first; selection result sets can be large.
359
- - **Mistake:** Assuming a GRCh37 coordinate will work.
360
- **Fix:** The dataset is GRCh38 only.
361
-
362
- ## References
363
-
364
- - [references/onekgpd_commands.md](references/onekgpd_commands.md) — full
365
- per-command argument tables and the returned-variant output schema.
366
- - [references/annotation_vocabularies.md](references/annotation_vocabularies.md)
367
- — the controlled-vocabulary terms accepted by the CSV filter flags
368
- (consequence, impact, biotype, feature type, ClinVar significance,
369
- AlphaMissense class, variant class).
370
- - 1000 Genomes Project / IGSR: https://www.internationalgenome.org/
371
- - 1000 Genomes Project dataset online: https://dnaerys.org/online/