@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,325 +0,0 @@
1
- ---
2
- name: dnanexus-integration
3
- description: Build and operate reproducible genomics workloads on DNAnexus with the dx CLI, dxpy, apps/applets, native workflows, dxCompiler, and Nextflow. Use for DNAnexus data transfers, dxapp.json development, execution monitoring, workflow import, and project automation.
4
- license: MIT
5
- compatibility: Requires a DNAnexus account, network access, Python 3.11+, and dx-toolkit/dxpy; some workflow and infrastructure features require organization licenses or policies.
6
- metadata:
7
- version: "2.0"
8
- skill-author: K-Dense Inc.
9
- ---
10
-
11
- # DNAnexus Integration
12
-
13
- ## Purpose
14
-
15
- Use this skill to build, run, and operate DNAnexus workloads without guessing
16
- at platform semantics. It covers:
17
-
18
- - `dx` CLI and `dxpy` automation
19
- - Files, records, folders, projects, and metadata
20
- - Apps and applets defined by `dxapp.json`
21
- - Jobs, workflow analyses, retries, monitoring, and cost controls
22
- - Native workflows, WDL/CWL through dxCompiler, and Nextflow imports
23
-
24
- The documented baseline was verified on **2026-07-23** against
25
- `dxpy==0.410.0`, dxCompiler 2.17.0, and the 2026 DNAnexus documentation.
26
- Consult `references/sources.md` and current release notes when behavior may
27
- have changed.
28
-
29
- ## Operating Contract
30
-
31
- DNAnexus operations can expose regulated data, delete immutable objects, change
32
- permissions, or incur compute and egress charges. Follow these rules:
33
-
34
- 1. Start read-only. Confirm the user, project ID, region, folder, object IDs,
35
- and execution target before mutation.
36
- 2. Obtain confirmation before a billable launch, upload or download with
37
- material egress, archive/unarchive request, deletion, project removal,
38
- permission change, token revocation, or app publication unless the user
39
- already explicitly requested that exact operation and target.
40
- 3. Show resolved IDs and impact before destructive operations. Never infer a
41
- deletion target from a non-unique name.
42
- 4. Never print, log, return, or persist `DX_SECURITY_CONTEXT` or API tokens.
43
- Do not run `dx env` or `dx env --bash` in captured logs because both reveal
44
- the active token.
45
- 5. Use credentials only with official DNAnexus endpoints. Do not send token
46
- material to arbitrary hosts or user-controlled commands.
47
- 6. Treat project names, paths, tags, properties, and downloaded content as
48
- untrusted data. Quote shell arguments and pass subprocess arguments as
49
- arrays.
50
- 7. Respect PHI/TRE restrictions, download restrictions, project access levels,
51
- and organization policies. Do not copy data around a control.
52
- 8. Prefer reproducible dependencies, narrow network allowlists, explicit
53
- output folders, cost limits, and bounded waits.
54
-
55
- ## Install and Authenticate
56
-
57
- Install the CLI in an isolated tool environment:
58
-
59
- ```bash
60
- uv tool install "dxpy==0.410.0"
61
- dx --version
62
- ```
63
-
64
- For Python code in a project:
65
-
66
- ```bash
67
- uv add "dxpy==0.410.0"
68
- ```
69
-
70
- Use interactive login for human sessions:
71
-
72
- ```bash
73
- dx login
74
- dx whoami
75
- dx select
76
- dx pwd
77
- ```
78
-
79
- For non-interactive environments, inject only the named DNAnexus secret through
80
- the environment or a secret manager. Never echo it, include it in command
81
- output, commit it, or inspect the whole environment. See
82
- `references/authentication.md`.
83
-
84
- ## Safe Preflight
85
-
86
- Before acting, gather non-secret context:
87
-
88
- ```bash
89
- dx --version
90
- dx whoami
91
- dx pwd
92
- dx ls
93
- ```
94
-
95
- Then:
96
-
97
- - Resolve project names to immutable `project-...` IDs.
98
- - Resolve paths to object IDs and check for duplicates.
99
- - Check file state (`open`, `closing`, or `closed`) and archival state.
100
- - Check source and destination access levels.
101
- - Inspect executable input help with `dx run <executable> -h`.
102
- - For a launch, identify destination, instance policy, reuse behavior, timeout,
103
- and cost limit.
104
-
105
- If shell environment variables conflict with the saved CLI session, follow
106
- `references/authentication.md`; do not expose either credential while
107
- diagnosing.
108
-
109
- ## Choose the Right Path
110
-
111
- | Goal | Read first | Preferred interface |
112
- |---|---|---|
113
- | Build an app or applet | `references/app-development.md` | `dx-app-wizard`, `dx build` |
114
- | Configure `dxapp.json` | `references/configuration.md` | JSON plus validator script |
115
- | Transfer or organize data | `references/data-operations.md` | `dx`, Upload/Download Agent |
116
- | Write platform automation | `references/python-sdk.md` | `dxpy` |
117
- | Launch or debug execution | `references/job-execution.md` | `dx run`, `dx watch`, `dxpy` |
118
- | Import WDL, CWL, or Nextflow | `references/workflow-languages.md` | dxCompiler or `dx build --nextflow` |
119
- | Diagnose auth, cost, or failures | `references/operations-and-troubleshooting.md` | read-only inspection first |
120
-
121
- ## Core Workflows
122
-
123
- ### Transfer data
124
-
125
- Use `dx upload` and `dx download` for small sets. Use Upload Agent for multiple
126
- or large files (official guidance recommends it above 50 MB) and Download Agent
127
- for large or long-running batch downloads.
128
-
129
- ```bash
130
- dx upload "sample.fastq.gz" \
131
- --path "project-xxxx:/raw/sample.fastq.gz" \
132
- --property "sample_id=S001"
133
-
134
- dx download "project-xxxx:/results/sample.bam" \
135
- --output "sample.bam"
136
- ```
137
-
138
- Upload Agent compresses uncompressed inputs by default and appends `.gz`. Use
139
- `--do-not-compress` when byte-for-byte preservation or the original name is
140
- required. See `references/data-operations.md`.
141
-
142
- ### Search accurately with dxpy
143
-
144
- `find_data_objects()` uses exact name matching unless `name_mode` is supplied.
145
- Do not pass `"*.bam"` without `name_mode="glob"`.
146
-
147
- ```python
148
- import dxpy
149
-
150
- files = dxpy.find_data_objects(
151
- classname="file",
152
- project="project-xxxx",
153
- folder="/results",
154
- recurse=True,
155
- name="*.bam",
156
- name_mode="glob",
157
- state="closed",
158
- describe={"fields": {"name": True, "size": True, "archivalState": True}},
159
- limit=100,
160
- )
161
-
162
- for result in files:
163
- description = result["describe"]
164
- print(result["id"], description["name"], description["archivalState"])
165
- ```
166
-
167
- Bound broad searches with a project, folder, time range, and `limit`.
168
-
169
- ### Build an applet
170
-
171
- ```bash
172
- dx-app-wizard
173
- ```
174
-
175
- Resolve bundled helpers relative to this skill directory. From the skill root:
176
-
177
- ```bash
178
- uv run python "scripts/validate_dxapp.py" \
179
- "/path/to/my-app/dxapp.json" --kind applet --strict
180
- ```
181
-
182
- Then build the source directory:
183
-
184
- ```bash
185
- dx build "/path/to/my-app"
186
- ```
187
-
188
- For a versioned app, use the current build form:
189
-
190
- ```bash
191
- dx build "/path/to/my-app" --create-app
192
- ```
193
-
194
- New configurations should use Ubuntu 24.04 and
195
- `regionalOptions.<region>.systemRequirements`. Top-level `resources` and
196
- `runSpec.systemRequirements` in `dxapp.json` are deprecated. See
197
- `references/configuration.md`.
198
-
199
- ### Launch with explicit controls
200
-
201
- First inspect the executable:
202
-
203
- ```bash
204
- dx run "applet-xxxx" -h
205
- ```
206
-
207
- After target and cost confirmation:
208
-
209
- ```bash
210
- dx run "applet-xxxx" \
211
- --input-json-file "inputs.json" \
212
- --destination "project-xxxx:/runs/run-001" \
213
- --cost-limit 25
214
- ```
215
-
216
- Keep the normal confirmation prompt for interactive use. Add `--yes` only in
217
- reviewed automation where the exact executable, project, inputs, destination,
218
- and cost policy are already approved.
219
-
220
- ### Monitor jobs and analyses
221
-
222
- ```bash
223
- dx find executions --created-after=-2h
224
- dx find jobs --state failed
225
- dx find analyses --created-after=-1d
226
- dx watch "job-xxxx" --get-streams
227
- ```
228
-
229
- A run of an app or applet returns a `job-...`; a run of a workflow returns an
230
- `analysis-...`. `dxpy.DXJob.wait_on_done()` and
231
- `dxpy.DXAnalysis.wait_on_done()` can raise `DXJobFailureError` for remote
232
- failure, termination, or local wait timeout. Re-describe remote state before
233
- classifying it; see `references/job-execution.md`.
234
-
235
- ### Chain executions without polling
236
-
237
- Use job-based output references:
238
-
239
- ```python
240
- import dxpy
241
-
242
- qc_job = dxpy.DXApplet("applet-qc").run(
243
- {"reads": dxpy.dxlink("file-input")},
244
- project="project-xxxx",
245
- folder="/runs/run-001/qc",
246
- cost_limit=10,
247
- )
248
-
249
- align_job = dxpy.DXApplet("applet-align").run(
250
- {"reads": qc_job.get_output_ref("filtered_reads")},
251
- project="project-xxxx",
252
- folder="/runs/run-001/alignment",
253
- cost_limit=25,
254
- )
255
- ```
256
-
257
- The downstream job remains `waiting_on_input` until the referenced output is
258
- ready. Do not wrap `get_output_ref()` in `dxpy.dxlink()`.
259
-
260
- ## Current Platform Guidance
261
-
262
- - Supported app execution environments are Ubuntu 24.04 and 20.04; prefer
263
- 24.04 for new work.
264
- - In Ubuntu 24.04, prefer a virtual environment for Python dependencies even
265
- though the AEE sets `PIP_BREAK_SYSTEM_PACKAGES=1`; system/PyPI conflicts can
266
- otherwise produce `DXExecDependencyError`.
267
- - Runtime `execDepends` can drift. Prefer pinned asset bundles, bundled
268
- dependencies, or pinned containers for production.
269
- - Dynamic instance selection is configured with
270
- `instanceTypeSelector.allowedInstanceTypes` and may require an organization
271
- license.
272
- - Automatic scale-up after `AppInsufficientResourceError` requires both an
273
- execution restart policy and the organization policy that permits instance
274
- upgrades.
275
- - Retired instance types are rejected when apps/applets are created or updated.
276
- Discover available instance types instead of copying a stale list.
277
- - Jobs normally have a 30-day runtime limit.
278
- - Download security status is surfaced by current APIs/CLI. Treat a malicious
279
- file warning as a stop condition unless the user explicitly approves a safe
280
- containment workflow.
281
-
282
- ## Bundled Helpers
283
-
284
- The commands below assume the current directory is this skill's root. Otherwise
285
- resolve `scripts/` relative to the loaded skill directory.
286
-
287
- ### Validate `dxapp.json`
288
-
289
- ```bash
290
- uv run python "scripts/validate_dxapp.py" \
291
- "path/to/dxapp.json" --kind app --strict
292
- ```
293
-
294
- This offline validator catches structural mistakes, deprecated placement,
295
- broad access, and inconsistent regional requirements. It supplements, not
296
- replaces, `dx build` validation.
297
-
298
- ### Inspect the installed SDK
299
-
300
- ```bash
301
- uv run --with "dxpy==0.410.0" \
302
- "scripts/inspect_dxpy.py" --strict
303
- ```
304
-
305
- This performs offline symbol and signature checks. It does not authenticate or
306
- make network calls.
307
-
308
- ## Reference Index
309
-
310
- - `references/authentication.md` — login, tokens, environment precedence, and
311
- secret handling
312
- - `references/app-development.md` — applet/app lifecycle, entry points,
313
- testing, build, and publication
314
- - `references/configuration.md` — current `dxapp.json`, regions, resources,
315
- dependencies, permissions, and retry policy
316
- - `references/data-operations.md` — transfers, search, metadata, cloning,
317
- archival, folders, and deletion
318
- - `references/python-sdk.md` — verified `dxpy` APIs and error handling
319
- - `references/job-execution.md` — jobs, analyses, monitoring, chaining, reuse,
320
- retries, and cost controls
321
- - `references/workflow-languages.md` — native workflows, WDL/CWL with
322
- dxCompiler, and Nextflow
323
- - `references/operations-and-troubleshooting.md` — operational playbooks and
324
- failure diagnosis
325
- - `references/sources.md` — authoritative documentation and version baseline
@@ -1,99 +0,0 @@
1
- ---
2
- name: docx
3
- description: "Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files) or Word templates (.dotx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', '.dotx', or requests to produce professional documents with formatting like tables of contents, headings, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx or .dotx files, inserting or replacing images in documents, performing find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. If the user asks for a 'report', 'memo', 'letter', 'template', or similar deliverable as a Word or .docx file, use this skill. Do NOT use for PDFs, spreadsheets, Google Docs, or general coding tasks unrelated to document generation."
4
- license: Proprietary. LICENSE.txt has complete terms
5
- metadata:
6
- version: "2.1"
7
- skill-author: Anthropic, PBC
8
- source: https://github.com/anthropics/skills/tree/main/skills/docx
9
- ---
10
-
11
- # DOCX creation, editing, and analysis
12
-
13
- A `.docx` is a ZIP archive of XML files. Choose your approach by task:
14
-
15
- | Task | Approach |
16
- |---|---|
17
- | **Create** a new document | Write a `docx` (npm) script — see gotchas below |
18
- | **Edit** an existing document | `unzip` → edit `word/document.xml` → `zip` (docx-js cannot open existing files) |
19
- | **Read** content | `pandoc -t markdown file.docx` |
20
-
21
- > Script paths below are relative to this skill's directory.
22
-
23
- ## Creating with docx-js — gotchas
24
-
25
- `docx` is preinstalled — do not run `npm install` first; write the script and `require('docx')` directly. Only if that require fails: `npm install docx`. The model knows the API; these are the footguns:
26
-
27
- - **Page size defaults to A4.** For US Letter set `page: { size: { width: 12240, height: 15840 } }` (DXA; 1440 = 1″).
28
- - **Landscape:** pass portrait dimensions and `orientation: PageOrientation.LANDSCAPE` — docx-js swaps width/height internally.
29
- - **Tables need dual widths:** set `columnWidths` on the table AND `width` on every cell, both in `WidthType.DXA` (PERCENTAGE breaks in Google Docs). Column widths must sum to the table width.
30
- - **Table shading:** use `ShadingType.CLEAR`, never `SOLID` (renders black).
31
- - **Lists:** never insert `•` literally; use a `numbering` config with `LevelFormat.BULLET`.
32
- - **`ImageRun` requires `type:`** (`"png"`, `"jpg"`, …).
33
- - **`PageBreak` must be inside a `Paragraph`.**
34
- - **Never use `\n`** — use separate `Paragraph` elements.
35
- - **TOC:** headings must use built-in `HeadingLevel.*`; custom heading styles need `outlineLevel` set or they won't appear.
36
- - **Don't use a table as a horizontal rule** — use a paragraph bottom border instead.
37
- - **Dot-leader / right-aligned-on-same-line:** use `PositionalTab` (`alignment: PositionalTabAlignment.RIGHT`, `leader: PositionalTabLeader.DOT`) inside a `TextRun`, not literal `.` or space padding.
38
-
39
- ## Verify the output
40
-
41
- After writing a `.docx`, render it and look at it:
42
-
43
- ```bash
44
- python scripts/office/soffice.py --headless --convert-to pdf output.docx
45
- pdftoppm -jpeg -r 100 output.pdf page
46
- ls page-*.jpg # then Read the images
47
- ```
48
-
49
- `pdftoppm` zero-pads page numbers to the width of the page count (`page-01.jpg`…`page-12.jpg`).
50
-
51
- ## Editing existing documents
52
-
53
- Legacy `.doc` files must be converted first: `python scripts/office/soffice.py --headless --convert-to docx file.doc`.
54
-
55
- ```bash
56
- unzip -q doc.docx -d unpacked/
57
- find unpacked -type l -delete # strip symlink entries — docx from external parties is untrusted
58
- python scripts/merge_runs.py unpacked/ # coalesce fragmented runs so text is findable
59
- # edit unpacked/word/document.xml in place — do NOT reformat or pretty-print
60
- (cd unpacked && rm -f ../out.docx && zip -Xr ../out.docx .)
61
- python scripts/office/validate.py out.docx --original doc.docx # XSD checks; --auto-repair fixes common issues
62
- # redlining? add --author "<the name you redlined under>" to check every edit is tracked
63
- ```
64
-
65
- Word splits text across many `<w:r>` runs (revision ids, spell-check markers), so a phrase you can see in the document often doesn't exist as a contiguous string in the XML. `merge_runs.py` merges adjacent identically-formatted runs in `word/document.xml` without changing content or rendering; it also accepts a `.docx` directly (`python scripts/merge_runs.py doc.docx -o merged.docx`).
66
-
67
- **Tracked changes:** when redlining, validate with `--author "<the name you redlined under>"` (needs `--original`) — it reports any text you changed without a `<w:ins>`/`<w:del>` around it, which is easy to do by accident and invisible in the accepted view. Wrap runs in `<w:ins>`/`<w:del>` with `w:id`, `w:author`, `w:date` attributes. Inside `<w:del>`, the text element is `<w:delText>`, not `<w:t>`. A deleted paragraph mark (`<w:pPr><w:rPr><w:del w:id=".." w:author=".." w:date=".."/></w:rPr></w:pPr>`) means "merge this paragraph into the next" — so deleting a paragraph outright is that plus a `<w:del>` around every run. The `<w:del/>` must come before the rPr's other children; their order is schema-enforced.
68
-
69
- To produce a clean copy with all tracked changes accepted: `python scripts/accept_changes.py in.docx out.docx`.
70
-
71
- Accepting a deleted paragraph mark should join that paragraph to the one below it, so a paragraph whose runs are *all* deleted vanishes. Word does this; `accept_changes.py` and `pandoc --track-changes=accept` don't always. Both fail the same way — they strip the deleted text but leave the emptied paragraph behind, which reads as a stray empty bullet when it was auto-numbered:
72
-
73
- - `pandoc --track-changes=accept` never joins the paragraphs.
74
- - `accept_changes.py` (LibreOffice) joins them correctly, except when the deleted paragraph is followed by an empty spacer paragraph.
75
-
76
- An empty bullet in either view is an artifact of that view, not a defect in the document. Check paragraph deletions in the XML.
77
-
78
- ## Comments
79
-
80
- Comments require six cross-linked files. Use the helper — directory mode when you'll also be editing `document.xml` (saves an unzip/rezip cycle), `.docx`-direct mode otherwise:
81
-
82
- ```bash
83
- # Against an already-unpacked directory (preferred when also placing markers)
84
- python scripts/comment.py unpacked/ "Fees & expenses cap is too low"
85
- python scripts/comment.py unpacked/ "Agreed" --parent 0
86
-
87
- # Against a .docx directly
88
- python scripts/comment.py contract.docx "This cap is too low" -o annotated.docx
89
- ```
90
-
91
- The script writes `comments.xml`, `commentsExtended.xml`, `commentsIds.xml`, `commentsExtensible.xml`, the relationships, and the content-type overrides. Comment IDs are auto-assigned. It then prints the `<w:commentRangeStart>`/`<w:commentRangeEnd>`/`<w:commentReference>` snippet to add to `word/document.xml` so the comment anchors to specific text — until you place those markers, the comment exists but is not visible.
92
-
93
- ## Dependencies
94
-
95
- `docx` (npm, preinstalled — install only if `require('docx')` fails) · `pandoc` · LibreOffice (`soffice`) · `pdftoppm` (Poppler)
96
-
97
- ---
98
-
99
- *This skill is created and maintained by [Anthropic](https://github.com/anthropics/skills/tree/main/skills/docx). Vendored here unmodified except for frontmatter metadata; see LICENSE.txt for terms.*