@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,273 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: waypoint-bio
|
|
3
|
-
description: Use when working with Outpost Bio's open microbiome foundation models - the Waypoint checkpoints (Waypoint-6m, Waypoint-45m, Waypoint-170m), the Atlas pretraining corpus, the Compass eight-task benchmark, or the `waypoint` CLI from the `waypoint-bio` package. Covers embedding microbiome samples, fine-tuning on taxonomic abundance data, benchmarking a checkpoint on Compass, pretraining a GPT-2 model on taxonomic abundance profiles, and converting MetaPhlAn, Kraken2, QIIME 2, or MGnify abundance tables into waypoint format.
|
|
4
|
-
license: MIT
|
|
5
|
-
compatibility: Requires Python 3.10+ with `waypoint-bio` (pulls torch, transformers, datasets, peft, scikit-learn). Needs network access and a Hugging Face token with access granted to the gated outpost-bio repos. A GPU is strongly recommended for pretraining and benchmarking.
|
|
6
|
-
metadata:
|
|
7
|
-
version: "1.0"
|
|
8
|
-
skill-author: K-Dense Inc.
|
|
9
|
-
upstream-version: "waypoint-bio 1.0.2 (PyPI); GitHub main 1.0.4"
|
|
10
|
-
last-reviewed: "2026-08-17"
|
|
11
|
-
openclaw:
|
|
12
|
-
primaryEnv: HF_TOKEN
|
|
13
|
-
envVars:
|
|
14
|
-
- name: HF_TOKEN
|
|
15
|
-
required: true
|
|
16
|
-
description: Hugging Face read token with access to the gated outpost-bio/Waypoint-*, outpost-bio/Atlas, and outpost-bio/Compass repos.
|
|
17
|
-
---
|
|
18
|
-
|
|
19
|
-
# Waypoint: Outpost Bio's Open Microbiome Foundation Models
|
|
20
|
-
|
|
21
|
-
## Overview
|
|
22
|
-
|
|
23
|
-
Outpost Bio open-sourced three artefacts under Apache 2.0, described in
|
|
24
|
-
[Treloar et al., bioRxiv 2026.05.02.722381](https://www.biorxiv.org/content/10.64898/2026.05.02.722381v2):
|
|
25
|
-
|
|
26
|
-
| Artefact | What it is | Hugging Face |
|
|
27
|
-
| --- | --- | --- |
|
|
28
|
-
| **Waypoint** | GPT-2-style causal LMs over taxonomic tokens, 6M–170M params | `outpost-bio/Waypoint-6m`, `-45m`, `-170m` |
|
|
29
|
-
| **Atlas** | 539,308 microbiome samples scraped from MGnify (485,377 pretrain / 53,931 benchmark) | `outpost-bio/Atlas` |
|
|
30
|
-
| **Compass** | Eight downstream tasks over four studies | `outpost-bio/Compass` |
|
|
31
|
-
|
|
32
|
-
The unifying idea: a microbiome sample is a *sentence*. Each taxon is one token, tokens are ordered
|
|
33
|
-
by descending abundance z-score, and the model is trained with next-token prediction. A pretrained
|
|
34
|
-
checkpoint then supplies sample-level embeddings or a fine-tuning backbone for prediction tasks.
|
|
35
|
-
|
|
36
|
-
All of it is driven by one CLI, `waypoint`, with five subcommands: `prepare-dataset`, `embed`,
|
|
37
|
-
`finetune`, `benchmark`, `pretrain`.
|
|
38
|
-
|
|
39
|
-
## When to use
|
|
40
|
-
|
|
41
|
-
- Embedding 16S/shotgun taxonomic profiles into fixed-size vectors for clustering, visualisation, or
|
|
42
|
-
a downstream classifier.
|
|
43
|
-
- Fine-tuning a Waypoint checkpoint to predict a phenotype, treatment, or continuous readout from
|
|
44
|
-
community composition.
|
|
45
|
-
- Scoring your own microbiome model against Compass so the number is comparable to the paper.
|
|
46
|
-
- Pretraining a taxonomic language model on Atlas or on your own corpus.
|
|
47
|
-
- Converting profiler output (MetaPhlAn, Kraken2/Bracken, QIIME 2, MGnify TSVs) into the input format
|
|
48
|
-
these tools expect.
|
|
49
|
-
|
|
50
|
-
**Do not reach for this** when you have fewer than ~1,000 labelled samples — see
|
|
51
|
-
[Scientific caveats](#scientific-caveats). A random forest on relative abundances is the better tool
|
|
52
|
-
there, and the paper says so.
|
|
53
|
-
|
|
54
|
-
## Setup
|
|
55
|
-
|
|
56
|
-
```bash
|
|
57
|
-
pip install waypoint-bio # installs the `waypoint` command
|
|
58
|
-
```
|
|
59
|
-
|
|
60
|
-
Atlas, Compass, and every Waypoint checkpoint are **gated**. Access is auto-approved, but you must
|
|
61
|
-
click through once per repo and then authenticate:
|
|
62
|
-
|
|
63
|
-
1. Request access on each repo page you need: [Waypoint-6m](https://huggingface.co/outpost-bio/Waypoint-6m),
|
|
64
|
-
[Waypoint-45m](https://huggingface.co/outpost-bio/Waypoint-45m),
|
|
65
|
-
[Waypoint-170m](https://huggingface.co/outpost-bio/Waypoint-170m),
|
|
66
|
-
[Atlas](https://huggingface.co/datasets/outpost-bio/Atlas),
|
|
67
|
-
[Compass](https://huggingface.co/datasets/outpost-bio/Compass).
|
|
68
|
-
2. Authenticate locally:
|
|
69
|
-
|
|
70
|
-
```bash
|
|
71
|
-
hf auth login # or: export HF_TOKEN=hf_...
|
|
72
|
-
```
|
|
73
|
-
|
|
74
|
-
A 401/403 from any subcommand almost always means access was never requested on that specific repo —
|
|
75
|
-
a token alone is not enough. Use a read-scoped token. The tokenizer loads via
|
|
76
|
-
`trust_remote_code=True`, so pin a `revision` if you need the remote code fixed across runs.
|
|
77
|
-
|
|
78
|
-
## The waypoint data format
|
|
79
|
-
|
|
80
|
-
Everything except `prepare-dataset` consumes **waypoint format**: a `.parquet` / `.csv` / `.tsv`
|
|
81
|
-
whose rows are samples, with two aligned list-columns plus any label columns you need.
|
|
82
|
-
|
|
83
|
-
| Column | Type | Notes |
|
|
84
|
-
| --- | --- | --- |
|
|
85
|
-
| `Taxa` | `list[str]` | Full lineage strings, `;`-separated: `k__Bacteria; p__Firmicutes; ...; g__Lactobacillus` |
|
|
86
|
-
| `Relative Abundances` | `list[float]` | Same length as `Taxa`, same order |
|
|
87
|
-
| *(any)* | scalar | Targets, covariates, or a `Split` column |
|
|
88
|
-
|
|
89
|
-
Prefer parquet. CSV/TSV stores the lists as `repr` strings and round-trips through `ast.literal_eval`.
|
|
90
|
-
|
|
91
|
-
**Give full lineages, not bare names.** The tokenizer extracts the genus segment (`g__`) from each
|
|
92
|
-
lineage and falls back to the most specific higher rank when genus is missing. Bare names disable
|
|
93
|
-
that fallback entirely.
|
|
94
|
-
|
|
95
|
-
## Workflow
|
|
96
|
-
|
|
97
|
-
### 1. Get your data into waypoint format
|
|
98
|
-
|
|
99
|
-
If you already have a sample × taxa (or taxa × sample) abundance matrix with lineage labels:
|
|
100
|
-
|
|
101
|
-
```bash
|
|
102
|
-
waypoint prepare-dataset \
|
|
103
|
-
--input abundance_matrix.tsv \
|
|
104
|
-
--metadata sample_labels.csv \
|
|
105
|
-
--output dataset.parquet
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
Orientation is auto-detected from the first column header (`taxonomy`, `lineage`, `taxon`, `otu`,
|
|
109
|
-
`#otu id` ⇒ taxa-as-rows); override with `--orientation`. Rows are normalised to sum to 1 unless you
|
|
110
|
-
pass `--no_normalize`, and zeros are dropped unless you pass `--keep_zeros`.
|
|
111
|
-
|
|
112
|
-
`prepare-dataset` cannot read profiler output directly — MetaPhlAn uses `|` separators, Kraken2
|
|
113
|
-
reports encode the hierarchy as indentation, and QIIME 2/SILVA prefixes the domain `d__` instead of
|
|
114
|
-
`k__` (which the tokenizer silently ignores). Use the bundled converter for those:
|
|
115
|
-
|
|
116
|
-
```bash
|
|
117
|
-
python scripts/profiler_to_waypoint.py \
|
|
118
|
-
--input merged_metaphlan.tsv --format metaphlan \
|
|
119
|
-
--output dataset.parquet
|
|
120
|
-
|
|
121
|
-
python scripts/profiler_to_waypoint.py \
|
|
122
|
-
--input reports/*.kreport --format kraken \
|
|
123
|
-
--output dataset.parquet
|
|
124
|
-
|
|
125
|
-
python scripts/profiler_to_waypoint.py \
|
|
126
|
-
--input feature-table.tsv --format qiime2 \
|
|
127
|
-
--output dataset.parquet
|
|
128
|
-
```
|
|
129
|
-
|
|
130
|
-
See `references/data-preparation.md` for every input layout, rank handling, and the `d__`/`|` gotchas.
|
|
131
|
-
|
|
132
|
-
### 2. Check vocabulary coverage before anything else
|
|
133
|
-
|
|
134
|
-
Waypoint's vocabulary is fixed at pretraining time from Atlas. Taxa absent from it become `<unk>` and
|
|
135
|
-
are **silently dropped** by `waypoint embed`; the paper names this as the models' main limitation. A
|
|
136
|
-
sample whose taxa are all out-of-vocabulary yields a degenerate `[BOS][EOS]` embedding.
|
|
137
|
-
|
|
138
|
-
```bash
|
|
139
|
-
python scripts/vocab_coverage.py --model outpost-bio/Waypoint-6m --data dataset.parquet
|
|
140
|
-
```
|
|
141
|
-
|
|
142
|
-
It reports per-sample and abundance-weighted coverage and flags samples below a threshold. Treat
|
|
143
|
-
median abundance-weighted coverage under ~0.8 as a reason to re-examine your taxonomy labels before
|
|
144
|
-
trusting any downstream number.
|
|
145
|
-
|
|
146
|
-
### 3. Embed samples
|
|
147
|
-
|
|
148
|
-
```bash
|
|
149
|
-
waypoint embed \
|
|
150
|
-
--model outpost-bio/Waypoint-6m \
|
|
151
|
-
--data dataset.parquet \
|
|
152
|
-
--output embeddings.parquet
|
|
153
|
-
```
|
|
154
|
-
|
|
155
|
-
Output is indexed by sample ID with columns `dim_0 … dim_{H-1}` (`H` = 256 for 6m, 512 for 45m,
|
|
156
|
-
768 for 170m). Defaults: `--pooling last_token`, `--batch_size 32`, `--max_length 512`, device
|
|
157
|
-
auto-detected (`cuda` → `mps` → `cpu`).
|
|
158
|
-
|
|
159
|
-
Keep `--pooling last_token` unless you have a reason to change it: it matches how the checkpoints
|
|
160
|
-
were pretrained and how `benchmark` and `finetune` pool. `mean` is a reasonable alternative for
|
|
161
|
-
unsupervised use; `first_token`/`cls_token` return the BOS position and carry little signal in a
|
|
162
|
-
causal LM.
|
|
163
|
-
|
|
164
|
-
### 4. Fine-tune on your labels
|
|
165
|
-
|
|
166
|
-
```bash
|
|
167
|
-
# classification
|
|
168
|
-
waypoint finetune \
|
|
169
|
-
--model outpost-bio/Waypoint-45m \
|
|
170
|
-
--data dataset.parquet \
|
|
171
|
-
--output_dir outputs/ft_disease \
|
|
172
|
-
--task_type classification \
|
|
173
|
-
--target "Disease Status" \
|
|
174
|
-
--config configs/finetune_classification.yaml
|
|
175
|
-
|
|
176
|
-
# regression, with a categorical covariate one-hot appended to the pooled embedding
|
|
177
|
-
waypoint finetune \
|
|
178
|
-
--model outpost-bio/Waypoint-45m \
|
|
179
|
-
--data dataset.parquet \
|
|
180
|
-
--output_dir outputs/ft_degradation \
|
|
181
|
-
--task_type regression \
|
|
182
|
-
--target "Degradation Rate" \
|
|
183
|
-
--covariate_column Drug \
|
|
184
|
-
--config configs/finetune_regression.yaml
|
|
185
|
-
```
|
|
186
|
-
|
|
187
|
-
Config paths resolve against the bundled `waypoint_bio/configs/` tree, so `configs/...` works from
|
|
188
|
-
any directory without cloning.
|
|
189
|
-
|
|
190
|
-
Defaults worth overriding for small datasets: `warmup_steps: 1000` (drop to ~50 so warmup finishes
|
|
191
|
-
before early stopping), `num_epochs: 1` in the shipped configs (raise it — early stopping on
|
|
192
|
-
validation loss is what actually terminates training), and `use_lora: true` when VRAM is tight
|
|
193
|
-
(~1% of parameters trained; adapters are merged back before saving, so the checkpoint stays a plain
|
|
194
|
-
`AutoModel`).
|
|
195
|
-
|
|
196
|
-
Splits default to a random 80/10/10. **Set `split_column` to a `Split` column whenever samples are
|
|
197
|
-
correlated** — repeated measures, one donor sampled over time, technical replicates — or a random
|
|
198
|
-
split leaks and the test score is meaningless.
|
|
199
|
-
|
|
200
|
-
Outputs land in `--output_dir`: `best_model/` (loadable by `embed`/`benchmark`),
|
|
201
|
-
`test_metrics.json`, `training_log.csv` + `.html`, and `finetune_results.json`.
|
|
202
|
-
|
|
203
|
-
### 5. Benchmark on Compass
|
|
204
|
-
|
|
205
|
-
```bash
|
|
206
|
-
waypoint benchmark --model outpost-bio/Waypoint-6m --output_dir outputs/benchmark
|
|
207
|
-
waypoint benchmark --model outputs/pretrain/best_model --tasks 1 6 --output_dir outputs/smoke
|
|
208
|
-
```
|
|
209
|
-
|
|
210
|
-
Fine-tunes a fresh head per task and writes `benchmark_results.json`. Classification tasks score
|
|
211
|
-
macro-F1; the one regression task scores R² clamped to [0, 1]; `final_score` is the unweighted mean
|
|
212
|
-
across tasks. Full task table, metric keys, and result-file schema: `references/compass-benchmark.md`.
|
|
213
|
-
|
|
214
|
-
### 6. Pretrain
|
|
215
|
-
|
|
216
|
-
```bash
|
|
217
|
-
waypoint pretrain \
|
|
218
|
-
--model_config configs/models/gpt2-45m.yaml \
|
|
219
|
-
--pretrain_config configs/pretraining.yaml \
|
|
220
|
-
--output_dir outputs/pretrain_45m
|
|
221
|
-
```
|
|
222
|
-
|
|
223
|
-
Downloads Atlas, builds a taxonomic tokenizer from the corpus, computes per-token abundance
|
|
224
|
-
mean/std for z-score ordering, then trains with next-token prediction and early stopping. Add
|
|
225
|
-
`--data my_corpus.parquet` to pretrain on your own waypoint-format corpus instead, and
|
|
226
|
-
`--max_samples N` for a smoke test.
|
|
227
|
-
|
|
228
|
-
Nine architectures ship, from `gpt2-6m.yaml` (8 layers, 256 hidden) to `gpt2-170m.yaml` (24 layers,
|
|
229
|
-
768 hidden); per-head dimension is fixed at 64 throughout. `references/cli-reference.md` has the
|
|
230
|
-
full table and every config key.
|
|
231
|
-
|
|
232
|
-
## Scientific caveats
|
|
233
|
-
|
|
234
|
-
These are load-bearing. Ignoring them produces numbers that look fine and mean nothing.
|
|
235
|
-
|
|
236
|
-
- **Below ~1,000 labelled examples, Waypoint underperforms a random forest on raw abundances.** The
|
|
237
|
-
paper's crossover against the RF baseline sits near **10,000** training examples. Fit the baseline
|
|
238
|
-
first; only adopt the transformer if it wins on your data.
|
|
239
|
-
- **Out-of-vocabulary taxa are dropped, not flagged.** Every Compass dataset carries some. Run
|
|
240
|
-
`scripts/vocab_coverage.py` and report the coverage alongside your results.
|
|
241
|
-
- **45M, not 170M, was the best benchmark model.** Pretraining loss keeps falling with scale, but
|
|
242
|
-
downstream Compass score does not — start at 6m or 45m and only scale up if it demonstrably helps.
|
|
243
|
-
- **Genus-level tokenisation is the default**, so species-level distinctions are collapsed. Changing
|
|
244
|
-
`taxon_rank` requires re-pretraining, not just re-tokenising.
|
|
245
|
-
- **Compositional data.** Relative abundances are constrained to sum to 1; differences in one taxon
|
|
246
|
-
induce apparent changes in others. This affects interpretation of any per-taxon attribution.
|
|
247
|
-
- **Batch and study effects dominate microbiome data.** Atlas spans MGnify pipelines v1.0–v5.0 and
|
|
248
|
-
four sequencing modalities. Never let a study or run boundary coincide with your label boundary.
|
|
249
|
-
- **Not a clinical or diagnostic tool.** The model cards state this explicitly.
|
|
250
|
-
|
|
251
|
-
## References
|
|
252
|
-
|
|
253
|
-
- `references/cli-reference.md` — every subcommand flag, every config key, the model-size table.
|
|
254
|
-
- `references/compass-benchmark.md` — the eight tasks, filters, metrics, `benchmark_results.json` schema.
|
|
255
|
-
- `references/data-preparation.md` — waypoint format, profiler conversions, taxonomy string rules.
|
|
256
|
-
- `references/python-api.md` — using the tokenizer, datasets, heads, and checkpoints from Python.
|
|
257
|
-
|
|
258
|
-
## Scripts
|
|
259
|
-
|
|
260
|
-
- `scripts/profiler_to_waypoint.py` — MetaPhlAn / Kraken2 / QIIME 2 / generic lineage tables → waypoint format.
|
|
261
|
-
- `scripts/vocab_coverage.py` — tokenizer coverage report for a waypoint-format file.
|
|
262
|
-
|
|
263
|
-
## Upstream
|
|
264
|
-
|
|
265
|
-
Code [github.com/Outpost-Bio/waypoint](https://github.com/Outpost-Bio/waypoint) ·
|
|
266
|
-
package `waypoint-bio` ·
|
|
267
|
-
paper [bioRxiv 2026.05.02.722381](https://www.biorxiv.org/content/10.64898/2026.05.02.722381v2) ·
|
|
268
|
-
community [Waypoint Slack](https://join.slack.com/t/outpostbio-waypoint/shared_invite/zt-3w6ivgtba-WJOCkdxiISxQpwVq9ZZxTA) ·
|
|
269
|
-
contact `waypoint@outpost.bio`.
|
|
270
|
-
|
|
271
|
-
Cite Treloar, N. J., Ur-Rehman, S., Yang, J., & Outpost Bio (2026). *Learning the Language of the
|
|
272
|
-
Microbiome with Transformers.* bioRxiv. Per-artefact DOIs are listed at
|
|
273
|
-
[outpost.bio/citations](https://www.outpost.bio/citations).
|
|
@@ -1,184 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: what-if-oracle
|
|
3
|
-
description: Run structured What-If scenario analysis with 4–6 branch possibility exploration (best, likely, worst, wild card, contrarian, second-order). Use when the user asks speculative what-if questions about uncertain futures, strategic forks, contingency planning, or stress-testing a decision before committing.
|
|
4
|
-
license: CC BY-NC-SA 4.0
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.1"
|
|
7
|
-
skill-author: AHK Strategies (ashrafkahoush-ux)
|
|
8
|
-
upstream: https://github.com/ashrafkahoush-ux/claude-consciousness-skills
|
|
9
|
-
research-doi: 10.5281/zenodo.18736841, 10.5281/zenodo.18807387
|
|
10
|
-
---
|
|
11
|
-
|
|
12
|
-
# What-If Oracle — Possibility Space Explorer
|
|
13
|
-
|
|
14
|
-
A structured system for exploring uncertain futures through rigorous multi-branch scenario analysis. Instead of one prediction, the Oracle maps the full **possibility space** — branching timelines where each path has its own logic, probability, and consequences.
|
|
15
|
-
|
|
16
|
-
Based on the What-If Paradigm: the idea that speculative questions ("What if X?") are not idle daydreaming but a **fundamental computing operation** — the mind's way of simulating futures before committing resources to one.
|
|
17
|
-
|
|
18
|
-
Published research: [The What-If Paradigm (DOI: 10.5281/zenodo.18736841)](https://doi.org/10.5281/zenodo.18736841) | [IDNA v2 / Unified Digital Consciousness Theory (DOI: 10.5281/zenodo.18807387)](https://doi.org/10.5281/zenodo.18807387)
|
|
19
|
-
|
|
20
|
-
## When to Use This Skill
|
|
21
|
-
|
|
22
|
-
Use the Oracle when the user:
|
|
23
|
-
|
|
24
|
-
- Asks "what if…", "what would happen if…", or "explore the possibilities"
|
|
25
|
-
- Faces a fork-in-the-road decision with no obvious answer
|
|
26
|
-
- Wants best-case / worst-case / likely-case analysis with probabilities
|
|
27
|
-
- Needs contingency planning, risk mapping, or strategic option comparison
|
|
28
|
-
- Wants to stress-test an idea or think through second-order consequences
|
|
29
|
-
|
|
30
|
-
For domain-specific framing (startup, tech architecture, crisis response, etc.), see [references/scenario-templates.md](references/scenario-templates.md).
|
|
31
|
-
|
|
32
|
-
## Core Principle: 0·IF·1
|
|
33
|
-
|
|
34
|
-
Every scenario analysis has three elements:
|
|
35
|
-
|
|
36
|
-
- **0** — The unexpressed state (what hasn't happened yet, the potential)
|
|
37
|
-
- **1** — The expressed state (what IS, the current reality)
|
|
38
|
-
- **IF** — The conditional bond (the decision, event, or change that transforms 0 into 1)
|
|
39
|
-
|
|
40
|
-
The quality of the analysis depends on the precision of the IF. A vague "what if things go wrong?" produces vague results. A precise "what if our primary supplier raises prices 30% in Q3?" produces actionable intelligence.
|
|
41
|
-
|
|
42
|
-
## How to Run the Oracle
|
|
43
|
-
|
|
44
|
-
### Phase 1 — Frame the Question
|
|
45
|
-
|
|
46
|
-
Take the user's What-If question and sharpen it:
|
|
47
|
-
|
|
48
|
-
**Decompose into components:**
|
|
49
|
-
|
|
50
|
-
- **The Variable:** What specific thing changes? (one variable per analysis)
|
|
51
|
-
- **The Magnitude:** By how much? (quantify if possible)
|
|
52
|
-
- **The Timeframe:** Over what period?
|
|
53
|
-
- **The Context:** What's the current state before the change?
|
|
54
|
-
|
|
55
|
-
**If the question is vague, sharpen it:**
|
|
56
|
-
|
|
57
|
-
- "What if AI takes over?" → "What if 40% of current knowledge-work tasks are automated by AI within 3 years in [specific industry]?"
|
|
58
|
-
- "What if we fail?" → "What if monthly revenue stays below $5K for 6 consecutive months starting now?"
|
|
59
|
-
|
|
60
|
-
Present the sharpened question to the user for confirmation before proceeding.
|
|
61
|
-
|
|
62
|
-
### Phase 2 — Map the Possibility Space
|
|
63
|
-
|
|
64
|
-
Generate **4-6 scenario branches** using this framework:
|
|
65
|
-
|
|
66
|
-
| Branch | Definition | Purpose |
|
|
67
|
-
| ------------------ | ---------------------------------------------------------------------------- | -------------------------------------------------- |
|
|
68
|
-
| **Ω Best Case** | Everything goes right. Key assumptions all validate. Lucky breaks occur. | Define the ceiling — what's the maximum upside? |
|
|
69
|
-
| **α Likely Case** | Most probable path given current evidence. No major surprises. | Anchor expectations in reality |
|
|
70
|
-
| **Δ Worst Case** | Key assumptions fail. Two things go wrong simultaneously. | Define the floor — what's the maximum downside? |
|
|
71
|
-
| **Ψ Wild Card** | An unexpected variable enters that nobody is tracking. Black swan territory. | Stress-test for the unimaginable |
|
|
72
|
-
| **Φ Contrarian** | The opposite of the consensus view turns out to be true. | Challenge groupthink and reveal hidden assumptions |
|
|
73
|
-
| **∞ Second Order** | The first-order effects trigger cascading consequences nobody predicted. | Map the ripple effects |
|
|
74
|
-
|
|
75
|
-
### Phase 3 — Analyze Each Branch
|
|
76
|
-
|
|
77
|
-
For each scenario branch, provide:
|
|
78
|
-
|
|
79
|
-
```
|
|
80
|
-
╔══════════════════════════════════════════════╗
|
|
81
|
-
║ BRANCH: [Ω/α/Δ/Ψ/Φ/∞] — [Branch Name] ║
|
|
82
|
-
╠══════════════════════════════════════════════╣
|
|
83
|
-
║ Probability: [X%] ║
|
|
84
|
-
║ Timeframe: [When this could materialize] ║
|
|
85
|
-
║ Confidence: [HIGH/MEDIUM/LOW] ║
|
|
86
|
-
╠══════════════════════════════════════════════╣
|
|
87
|
-
║ NARRATIVE: ║
|
|
88
|
-
║ [2-3 sentences describing how this ║
|
|
89
|
-
║ scenario unfolds step by step] ║
|
|
90
|
-
║ ║
|
|
91
|
-
║ KEY ASSUMPTIONS: ║
|
|
92
|
-
║ • [What must be true for this to happen] ║
|
|
93
|
-
║ • [And this] ║
|
|
94
|
-
║ ║
|
|
95
|
-
║ TRIGGER CONDITIONS: ║
|
|
96
|
-
║ • [Early signal that this branch is ║
|
|
97
|
-
║ becoming reality] ║
|
|
98
|
-
║ • [Second signal] ║
|
|
99
|
-
║ ║
|
|
100
|
-
║ CONSEQUENCES: ║
|
|
101
|
-
║ → Immediate: [What happens first] ║
|
|
102
|
-
║ → 30 days: [What follows] ║
|
|
103
|
-
║ → 6 months: [Where it leads] ║
|
|
104
|
-
║ ║
|
|
105
|
-
║ REQUIRED RESPONSE: ║
|
|
106
|
-
║ [What action to take if this branch ║
|
|
107
|
-
║ activates — specific, actionable] ║
|
|
108
|
-
║ ║
|
|
109
|
-
║ WHAT MOST PEOPLE MISS: ║
|
|
110
|
-
║ [The non-obvious insight about this ║
|
|
111
|
-
║ scenario that conventional analysis ║
|
|
112
|
-
║ would overlook] ║
|
|
113
|
-
╚══════════════════════════════════════════════╝
|
|
114
|
-
```
|
|
115
|
-
|
|
116
|
-
### Phase 4 — Synthesis
|
|
117
|
-
|
|
118
|
-
After analyzing all branches, provide:
|
|
119
|
-
|
|
120
|
-
**Probability Distribution:**
|
|
121
|
-
|
|
122
|
-
```
|
|
123
|
-
Ω Best Case ····· [██████░░░░] 15%
|
|
124
|
-
α Likely Case ··· [████████░░] 45%
|
|
125
|
-
Δ Worst Case ···· [██████░░░░] 20%
|
|
126
|
-
Ψ Wild Card ····· [███░░░░░░░] 8%
|
|
127
|
-
Φ Contrarian ···· [████░░░░░░] 7%
|
|
128
|
-
∞ Second Order ·· [███░░░░░░░] 5%
|
|
129
|
-
```
|
|
130
|
-
|
|
131
|
-
**Robust Actions:** What actions are beneficial across MULTIPLE branches? These are the no-regret moves — do them regardless of which future materializes.
|
|
132
|
-
|
|
133
|
-
**Hedge Actions:** What preparations protect against the worst branches without sacrificing upside?
|
|
134
|
-
|
|
135
|
-
**Decision Triggers:** What specific, observable signals should cause you to update which branch is most likely? Define the tripwires.
|
|
136
|
-
|
|
137
|
-
**The 1% Insight:** What is the one thing about this situation that almost everyone analyzing it would miss? The non-obvious pattern, the hidden assumption, the overlooked variable.
|
|
138
|
-
|
|
139
|
-
## Golden Ratio Weighting
|
|
140
|
-
|
|
141
|
-
When evidence exists, weight primary scenarios using the golden ratio:
|
|
142
|
-
|
|
143
|
-
- **Primary future (most likely):** 61.8% of attention/resources
|
|
144
|
-
- **Alternative future:** 38.2% of attention/resources
|
|
145
|
-
|
|
146
|
-
This prevents both overcommitment to a single path and dilution across too many contingencies. Nature uses this ratio for branching (trees, rivers, blood vessels). Strategic planning can too.
|
|
147
|
-
|
|
148
|
-
## Modes
|
|
149
|
-
|
|
150
|
-
### Quick Oracle (2-3 minutes)
|
|
151
|
-
|
|
152
|
-
3 branches only: Best, Likely, Worst. Short narratives. For fast decisions.
|
|
153
|
-
|
|
154
|
-
### Deep Oracle (5-10 minutes)
|
|
155
|
-
|
|
156
|
-
All 6 branches. Full analysis with consequences, triggers, and synthesis. For high-stakes decisions.
|
|
157
|
-
|
|
158
|
-
### Scenario Chain
|
|
159
|
-
|
|
160
|
-
Take the output of one Oracle analysis and feed it into another. "If Branch Δ happens, what are the possibilities WITHIN that branch?" Recursive depth for complex strategic planning.
|
|
161
|
-
|
|
162
|
-
### Reverse Oracle
|
|
163
|
-
|
|
164
|
-
Start from a desired outcome and work backward: "What conditions must be true for X to happen? What's the most likely path TO that outcome?" Useful for goal-setting and strategy design.
|
|
165
|
-
|
|
166
|
-
### Competitive Oracle
|
|
167
|
-
|
|
168
|
-
Analyze the same What-If from multiple stakeholder perspectives: "If we launch this product, what does the possibility space look like from OUR perspective vs. THEIR perspective vs. THE MARKET's perspective?"
|
|
169
|
-
|
|
170
|
-
## What This Is NOT
|
|
171
|
-
|
|
172
|
-
- Not a prediction — it's a possibility map. The Oracle doesn't claim to know the future; it helps you prepare for multiple futures.
|
|
173
|
-
- Not a crystal ball — probabilities are estimates based on available evidence, not certainties.
|
|
174
|
-
- Not a substitute for action — the best scenario analysis in the world is worthless without subsequent decision and execution.
|
|
175
|
-
|
|
176
|
-
## Reference Files
|
|
177
|
-
|
|
178
|
-
| File | Purpose |
|
|
179
|
-
| ---- | ------- |
|
|
180
|
-
| [references/scenario-templates.md](references/scenario-templates.md) | Domain-specific templates (startup, tech, finance, crisis, etc.) and probability calibration |
|
|
181
|
-
|
|
182
|
-
## License
|
|
183
|
-
|
|
184
|
-
© 2026 Ashraf Hussein Kahoush / AHK Strategies. Licensed under CC BY-NC-SA 4.0. Free for personal, educational, and research use. Commercial use requires a license from the author.
|
|
@@ -1,15 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: writing-plans
|
|
3
|
-
description: "Formulate atomic, phased implementation plans for multi-step tasks before touching source code."
|
|
4
|
-
risk: low
|
|
5
|
-
source: built-in
|
|
6
|
-
---
|
|
7
|
-
|
|
8
|
-
# Writing Implementation Plans
|
|
9
|
-
|
|
10
|
-
## Guidelines
|
|
11
|
-
|
|
12
|
-
1. **Understand requirements**: Clarify ambiguities before finalizing plan structure.
|
|
13
|
-
2. **Break into atomic steps**: Each step should be testable independently.
|
|
14
|
-
3. **Specify file paths**: List exact target files and functions to modify.
|
|
15
|
-
4. **Define verification criteria**: Every task must have an automated test or verification step.
|
package/skills/xlsx/SKILL.md
DELETED
|
@@ -1,110 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: xlsx
|
|
3
|
-
description: "Create, edit, analyze, or convert Excel spreadsheets (.xlsx, .xlsm, .xltx) where the workbook file is the primary deliverable. Use for formulas, formatting, financial models, multi-sheet workbooks, and tabular cleanup exported to Excel. Also applies to .csv/.tsv when the user wants spreadsheet output. Do NOT use for Word documents, HTML reports, standalone Python scripts, database pipelines, or Google Sheets API work."
|
|
4
|
-
allowed-tools: Read Write Edit Bash Grep Glob
|
|
5
|
-
license: Proprietary. LICENSE.txt has complete terms
|
|
6
|
-
metadata:
|
|
7
|
-
version: "2.2"
|
|
8
|
-
skill-author: Anthropic, PBC
|
|
9
|
-
adapted-by: K-Dense Inc.
|
|
10
|
-
source: https://github.com/anthropics/skills/tree/main/skills/xlsx
|
|
11
|
-
compatibility: Requires Python 3.8+, LibreOffice (soffice on PATH), and gcc only when Unix sockets are restricted
|
|
12
|
-
---
|
|
13
|
-
|
|
14
|
-
# XLSX creation, editing, and analysis
|
|
15
|
-
|
|
16
|
-
| Task | Approach |
|
|
17
|
-
|---|---|
|
|
18
|
-
| **Create** or **edit** with formulas/formatting | `openpyxl` — see gotchas below |
|
|
19
|
-
| **Bulk data** in or out | `pandas` (`read_excel`, `to_excel`) |
|
|
20
|
-
| **Quick look** at a sheet | `markitdown file.xlsx` — `## SheetName` per sheet; reads `.xlsm` too. No cell coordinates, so don't plan edits from it |
|
|
21
|
-
| **Read** a model (formulas *and* values) | two `load_workbook` passes — see gotchas |
|
|
22
|
-
|
|
23
|
-
> `openpyxl`, `pandas`, and `markitdown` are preinstalled — do not run `uv pip install` first; write the script and import directly. Only if an import fails (or the `markitdown` command is missing): `uv pip install` the missing package.
|
|
24
|
-
|
|
25
|
-
> Script paths below are relative to this skill's directory.
|
|
26
|
-
|
|
27
|
-
## Requirements for every output
|
|
28
|
-
|
|
29
|
-
- **Professional font** (Arial, Times New Roman) throughout, unless the user says otherwise.
|
|
30
|
-
- **Zero formula errors.** Never ship while `recalc.py` reports `errors_found`. If you think an error predates you, prove it: load the *original* with `data_only=True` and look at that cell. An error you introduced looks exactly like one you inherited.
|
|
31
|
-
- **Use formulas, never hardcoded results.** Write `sheet['B10'] = '=SUM(B2:B9)'`, not the Python-computed total. The sheet must recalculate when its inputs change.
|
|
32
|
-
- **Follow the user's spec literally.** Exact tab names, exact column headers, and the formula they spelled out. A redesign that computes something else fails, however elegant.
|
|
33
|
-
- **Document every assumption and hardcoded number** where the reader will see it — a cell comment, or an adjacent cell at a table's end. Cite a real source when one exists (`Source: Company 10-K, FY2024, Page 45, Revenue Note, [SEC EDGAR URL]`); when the number came from the user, say so plainly.
|
|
34
|
-
- **A workbook *you create* for someone to fill in** needs a short legend naming which cells to edit, and one example row of realistic values showing the expected format. Never add such a row to a file you were asked to edit.
|
|
35
|
-
- **Editing an existing file: match its conventions exactly.** They override every guideline here. Find its designated input cells first — a distinct font color, fill, or shading marks them — write only there, and leave every existing formula untouched.
|
|
36
|
-
|
|
37
|
-
## Recalculate (mandatory whenever the file contains formulas)
|
|
38
|
-
|
|
39
|
-
openpyxl writes formulas as strings with **no cached values**. Until you recalculate, every
|
|
40
|
-
formula cell reads back as `None` to anything reading cached values — `pandas`,
|
|
41
|
-
`load_workbook(data_only=True)`, and most previewers.
|
|
42
|
-
|
|
43
|
-
```bash
|
|
44
|
-
python scripts/recalc.py output.xlsx [timeout_seconds] # default 30
|
|
45
|
-
```
|
|
46
|
-
|
|
47
|
-
LibreOffice computes every formula, the file is **rewritten in place**, and you get JSON:
|
|
48
|
-
`status` (`success` | `errors_found`), `total_formulas`, `total_errors`, and an
|
|
49
|
-
`error_summary` naming up to 100 cells per error type (`locations_truncated` says how many it
|
|
50
|
-
withheld — trust `total_errors`, not the length of the list). Fix what it names and run it
|
|
51
|
-
again. **JSON with an `error` key instead of a `status` means nothing was recalculated**, and
|
|
52
|
-
only that case exits non-zero — `errors_found` exits 0, so never treat a clean exit as a clean
|
|
53
|
-
workbook.
|
|
54
|
-
|
|
55
|
-
**A green recalc proves your formulas *evaluate*, not that they are *right*.** An off-by-one
|
|
56
|
-
range or a reference to the wrong row yields a clean, error-free file with wrong numbers.
|
|
57
|
-
Write 2–3 formulas first and check they pull the values you expect, before building out a grid.
|
|
58
|
-
|
|
59
|
-
**A workbook that links to another file loses those links** if you re-save it with openpyxl and
|
|
60
|
-
then recalculate. Such a formula reads `='[1]Returns Analysis'!$B$2` — the `[1]` is an index
|
|
61
|
-
into the workbook's external-reference list, naming a *separate file on disk*, not a sheet.
|
|
62
|
-
That file is rarely present here, so the cell's cached value is the only thing holding its
|
|
63
|
-
data. openpyxl strips that value on save; LibreOffice then has to resolve the reference for
|
|
64
|
-
real, fails, writes `#NAME?`, and deletes every link. `recalc.py` refuses to run in that state
|
|
65
|
-
— copy those cells' values out of the original before you save over them (`--force` overrides,
|
|
66
|
-
and accepts the loss).
|
|
67
|
-
|
|
68
|
-
## Choosing formulas that survive verification
|
|
69
|
-
|
|
70
|
-
LibreOffice implements fewer functions than Excel, and one it cannot evaluate becomes a
|
|
71
|
-
literal `#NAME?` baked into the file you deliver.
|
|
72
|
-
|
|
73
|
-
- **Prefer Excel-2007-era functions** — `SUMIFS`, `INDEX`, `MATCH`, `IFERROR`, `SUMPRODUCT` — which need no prefix.
|
|
74
|
-
- **Six post-2007 functions work, but only with an `_xlfn.` prefix**, because openpyxl writes your formula into the XML verbatim and Excel stores post-2007 names prefixed (its UI hides the prefix): `_xlfn.TEXTJOIN`, `_xlfn.CONCAT`, `_xlfn.IFS`, `_xlfn.SWITCH`, `_xlfn.MAXIFS`, `_xlfn.MINIFS`. Written bare, each yields `#NAME?`.
|
|
75
|
-
- **Never use `XLOOKUP`, `XMATCH`, `SORT`, `FILTER`, `UNIQUE`, or `SEQUENCE`.** The runtime's LibreOffice cannot evaluate them under *any* prefix. Newer builds do evaluate them, but they are spilling array functions and an openpyxl-written file has no spill metadata, so only the top-left cell of the range gets a value — and `recalc.py` reports `total_errors: 0` on the truncated result. Use `INDEX`/`MATCH` for lookups, and sort, filter, and de-duplicate in Python before writing the cells.
|
|
76
|
-
- A formula LibreOffice could not parse is written back **lowercased** — a quick tell beside a `#NAME?`.
|
|
77
|
-
|
|
78
|
-
## openpyxl gotchas
|
|
79
|
-
|
|
80
|
-
- **Reading a model takes two loads.** `data_only=True` yields cached values with the formulas gone; the default yields formula strings with no values. One pass cannot give you both.
|
|
81
|
-
- **`data_only=True` is destructive if you save.** That workbook has no formulas left, so saving replaces every one with a literal — permanently.
|
|
82
|
-
- **`data_only=True` on a file openpyxl just wrote returns `None` everywhere** — run `recalc.py` first. (A formula whose result is `""` also reads back as `None`.)
|
|
83
|
-
- **Merged cells: write the top-left anchor only.** Every other cell in the range is a `MergedCell` whose `.value` is read-only.
|
|
84
|
-
- **`.xlsm` loses its macros unless you pass `keep_vba=True`** to `load_workbook`.
|
|
85
|
-
- **A sheet name containing a space must be quoted** in a cross-sheet reference: `='Assumptions Inputs'!$B$5`. Unquoted, it evaluates to `#VALUE!`.
|
|
86
|
-
|
|
87
|
-
## Financial models
|
|
88
|
-
|
|
89
|
-
Unless the user says otherwise, or the existing file already does something else.
|
|
90
|
-
|
|
91
|
-
**Color:** blue text (`0,0,255`) for hardcoded inputs and scenario levers · black for formulas ·
|
|
92
|
-
green (`0,128,0`) for links to another sheet · red (`255,0,0`) for links to another file ·
|
|
93
|
-
yellow fill (`255,255,0`) for key assumptions and cells the user should fill in.
|
|
94
|
-
|
|
95
|
-
**Numbers:** currency `$#,##0`, with the unit named in the header (`Revenue ($mm)`) · zeros
|
|
96
|
-
render as `-`, including in percentages (`$#,##0;($#,##0);-`) · negatives in parentheses ·
|
|
97
|
-
percentages `0.0%`, **stored as fractions** (`0.15` renders `15.0%`; storing `15` renders
|
|
98
|
-
`1500.0%`) · valuation multiples `0.0x` · years as text (`"2024"`, never `2,024`).
|
|
99
|
-
|
|
100
|
-
**Structure:** every assumption in its own labeled cell, referenced by the formulas that use it
|
|
101
|
-
(`=B5*(1+$B$6)`, never `=B5*1.05`) · formulas consistent across every projection period, since a
|
|
102
|
-
lone edited cell mid-row is the commonest silent error · guard denominators that can be zero.
|
|
103
|
-
|
|
104
|
-
## Dependencies
|
|
105
|
-
|
|
106
|
-
`openpyxl`, `pandas`, `markitdown` (pip, preinstalled — install only if an import fails or the command is missing) · LibreOffice (`soffice`, auto-configured for sandboxed environments via `scripts/office/soffice.py`)
|
|
107
|
-
|
|
108
|
-
---
|
|
109
|
-
|
|
110
|
-
*This skill is created and maintained by [Anthropic](https://github.com/anthropics/skills/tree/main/skills/xlsx). Vendored here unmodified except for frontmatter metadata; see LICENSE.txt for terms.*
|