@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,147 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: ontology-term-resolution
|
|
3
|
-
description: Resolve free-text scientific labels to ontology term IDs and validate existing CURIEs against the EBI Ontology Lookup Service (OLS4). Use whenever an ontology identifier must be produced or checked - annotating tissue, cell type, disease, phenotype, assay, chemical, organism, sex, or developmental stage fields; preparing metadata for GEO, ENA, BioSamples, CELLxGENE, HCA, or ISA-Tab submission; auditing a metadata table of term IDs; checking whether a term is obsolete and what replaced it; or mapping between ontologies. Triggers include "ontology term", "ontology ID", "CURIE", "controlled vocabulary", "UBERON", "CL:", "MONDO", "HPO", "EFO", "ChEBI", "NCBITaxon", "GO term", "PATO", "annotate this tissue/cell type/disease", and any request to emit or verify an identifier shaped like PREFIX:0001234.
|
|
4
|
-
license: MIT
|
|
5
|
-
compatibility: Requires Python 3.11+. Scripts use only the standard library - no third-party packages. Needs network access to https://www.ebi.ac.uk/ols4 (public, no API key).
|
|
6
|
-
allowed-tools: Read Write Edit Bash
|
|
7
|
-
metadata:
|
|
8
|
-
version: "1.0"
|
|
9
|
-
skill-author: K-Dense Inc.
|
|
10
|
-
---
|
|
11
|
-
|
|
12
|
-
# Ontology Term Resolution
|
|
13
|
-
|
|
14
|
-
## When to use
|
|
15
|
-
|
|
16
|
-
Any time an ontology identifier is about to be written down or trusted: annotating a metadata
|
|
17
|
-
column, filling a submission template, auditing a table someone else produced, or checking whether
|
|
18
|
-
an ID in an old file is still current.
|
|
19
|
-
|
|
20
|
-
## The rule
|
|
21
|
-
|
|
22
|
-
**Never write an ontology ID from memory, and never accept one without checking it.**
|
|
23
|
-
|
|
24
|
-
Ontology IDs are memorable in form and arbitrary in detail. A plausible-looking `UBERON:0002108`
|
|
25
|
-
is a real term (small intestine) that is not the liver, and nothing downstream will catch the
|
|
26
|
-
substitution — the ID is well-formed, the ontology is right, and the metadata is silently wrong.
|
|
27
|
-
Reviewers cannot spot it either, which is why these errors persist into published datasets.
|
|
28
|
-
|
|
29
|
-
Every ID this skill emits comes from a live OLS lookup. Every ID it is handed gets verified.
|
|
30
|
-
|
|
31
|
-
## Two directions
|
|
32
|
-
|
|
33
|
-
| Direction | Script | Question answered |
|
|
34
|
-
| --- | --- | --- |
|
|
35
|
-
| text → ID | `scripts/resolve_terms.py` | What is the term for "left ventricle"? |
|
|
36
|
-
| ID → verdict | `scripts/validate_terms.py` | Is `EFO:0001067` real, current, and labelled what this file claims? |
|
|
37
|
-
|
|
38
|
-
Both take single values or files, emit TSV or JSON, and need no packages beyond the standard
|
|
39
|
-
library.
|
|
40
|
-
|
|
41
|
-
## Resolve text to terms
|
|
42
|
-
|
|
43
|
-
```bash
|
|
44
|
-
cd skills/ontology-term-resolution/scripts
|
|
45
|
-
|
|
46
|
-
# one string, constrained to the ontology that should define it
|
|
47
|
-
python3 resolve_terms.py "liver" --ontology uberon
|
|
48
|
-
```
|
|
49
|
-
|
|
50
|
-
```
|
|
51
|
-
query rank curie label ontology match_type strategy defining_ontology
|
|
52
|
-
liver 1 UBERON:0002107 liver uberon exact_label exact true
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
```bash
|
|
56
|
-
# a column of tissue names; anything not an exact hit is reported, not guessed
|
|
57
|
-
python3 resolve_terms.py --input tissues.txt --ontology uberon \
|
|
58
|
-
--exact-only --format tsv -o resolved.tsv
|
|
59
|
-
|
|
60
|
-
# accept fuzzy fallbacks, then review the partial hits by hand
|
|
61
|
-
python3 resolve_terms.py "left ventrical of heart" --ontology uberon --top 3
|
|
62
|
-
```
|
|
63
|
-
|
|
64
|
-
The search escalates `exact` (label and synonym) → `token` → `fulltext` and stops at the first
|
|
65
|
-
strategy that returns anything, reporting which one fired. `--exact-only` disables the ladder.
|
|
66
|
-
`--branch UBERON:0000465` restricts candidates to descendants of a term.
|
|
67
|
-
|
|
68
|
-
**Read `match_type` before using a result.** `exact_label` and `exact_synonym` are safe;
|
|
69
|
-
`partial` means OLS returned its best guess for a string that does not exist as written, and
|
|
70
|
-
needs a human decision. `unresolved` is a legitimate output — see `references/curation-rules.md`
|
|
71
|
-
for the normalisations worth retrying first.
|
|
72
|
-
|
|
73
|
-
## Validate existing IDs
|
|
74
|
-
|
|
75
|
-
```bash
|
|
76
|
-
python3 validate_terms.py UBERON:0002107 EFO:0001067 UBERON:9999999
|
|
77
|
-
```
|
|
78
|
-
|
|
79
|
-
```
|
|
80
|
-
id status actual_label ontology replacement detail
|
|
81
|
-
UBERON:0002107 ok liver uberon
|
|
82
|
-
EFO:0001067 obsolete obsolete_parasitic infection efo MONDO:0005135 obsolete; replaced by MONDO:0005135
|
|
83
|
-
UBERON:9999999 not_found no such term in the ontology this prefix names
|
|
84
|
-
```
|
|
85
|
-
|
|
86
|
-
Exit code is 1 if anything failed, 0 otherwise, 2 on usage or network trouble — so it works as a
|
|
87
|
-
CI gate on a metadata file:
|
|
88
|
-
|
|
89
|
-
```bash
|
|
90
|
-
# id + label columns; catches IDs that exist but are labelled as something else
|
|
91
|
-
python3 validate_terms.py --input metadata.tsv --strict
|
|
92
|
-
|
|
93
|
-
# a tissue column must hold UBERON anatomical entities and nothing else
|
|
94
|
-
python3 validate_terms.py --input tissue_ids.tsv \
|
|
95
|
-
--branch UBERON:0000465 --expect-ontology uberon
|
|
96
|
-
```
|
|
97
|
-
|
|
98
|
-
| Status | Meaning | Verdict |
|
|
99
|
-
| --- | --- | --- |
|
|
100
|
-
| `ok` | Exists, current, consistent with everything asserted | pass |
|
|
101
|
-
| `matched_synonym` | Claimed label is a synonym; primary label differs | warn |
|
|
102
|
-
| `imported_only` | Home ontology no longer asserts this ID | warn |
|
|
103
|
-
| `not_a_class` | Term is a property or individual | warn |
|
|
104
|
-
| `not_found` | No such term | fail |
|
|
105
|
-
| `obsolete` | Obsoleted; `replacement` gives the successor when one exists | fail |
|
|
106
|
-
| `label_mismatch` | ID and claimed label describe different things | fail |
|
|
107
|
-
| `wrong_ontology` | Right kind of ID, wrong ontology for this column | fail |
|
|
108
|
-
| `wrong_branch` | Not a descendant of the required root | fail |
|
|
109
|
-
| `malformed_curie` | Not of the form `PREFIX:local` | fail |
|
|
110
|
-
|
|
111
|
-
`--strict` promotes warnings to failures.
|
|
112
|
-
|
|
113
|
-
## API behaviour that will mislead you
|
|
114
|
-
|
|
115
|
-
These are verified against the live service and are the reason this skill ships scripts rather
|
|
116
|
-
than a recipe. Full detail in `references/ols4-api.md`.
|
|
117
|
-
|
|
118
|
-
| Trap | Consequence |
|
|
119
|
-
| --- | --- |
|
|
120
|
-
| `exact=true` is exact **token** matching | `liver` returns 161 hits in UBERON; adding `queryFields=label` returns 1 |
|
|
121
|
-
| `/search` never returns `is_obsolete` or `term_replaced_by` | Named in `fieldList` they are dropped silently; only term detail can answer "is this ID still current" |
|
|
122
|
-
| `ontology=efo` returns MONDO and CL hits | Ontologies import each other; filter on the CURIE prefix yourself |
|
|
123
|
-
| The same term appears once per importing ontology | Deduplicate on `obo_id`, keep `is_defining_ontology: true` |
|
|
124
|
-
| The `obo_id` index has holes | `MONDO:0000001` is live but unindexed by `obo_id`; an IRI fallback is required to avoid a false `not_found` |
|
|
125
|
-
| IRIs are not all OBO PURLs | EFO and Orphanet use their own namespaces — resolve IRIs, do not template them |
|
|
126
|
-
| OxO is retired | Returns HTML with HTTP 200; use term cross-references or SSSOM instead |
|
|
127
|
-
| A branch check does not exclude cell types from anatomy | CARO puts `cell` under `anatomical structure`; constrain the prefix too |
|
|
128
|
-
|
|
129
|
-
## Choosing the ontology
|
|
130
|
-
|
|
131
|
-
MONDO for disease, HP for phenotype, UBERON for tissue, CL for cell type, EFO for assay, ChEBI for
|
|
132
|
-
compounds, NCBITaxon for organism, PATO for sex and for `normal`. Prefix-to-OLS-id mappings (`HP`
|
|
133
|
-
is served as `hp`, `Orphanet` as `ordo`), branch roots for `--branch`, and the overlapping-ontology
|
|
134
|
-
judgement calls are in `references/ontology-registry.md`.
|
|
135
|
-
|
|
136
|
-
## Reporting results
|
|
137
|
-
|
|
138
|
-
Give the ID **and** the label, and say how each was matched. A table of bare IDs cannot be
|
|
139
|
-
reviewed. State unresolved terms explicitly rather than filling them with the nearest hit.
|
|
140
|
-
|
|
141
|
-
## References
|
|
142
|
-
|
|
143
|
-
- `references/ols4-api.md` — endpoints, parameters, response fields, and every verified trap.
|
|
144
|
-
- `references/ontology-registry.md` — prefix/ontology-id table, branch roots, which ontology owns
|
|
145
|
-
which concept.
|
|
146
|
-
- `references/curation-rules.md` — candidate-selection procedure, normalisations to retry,
|
|
147
|
-
auditing an existing table, obsolete terms, cross-ontology mapping.
|
|
@@ -1,297 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: open-notebook
|
|
3
|
-
description: Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis. Use when organizing research materials into notebooks, ingesting diverse content sources (PDFs, videos, audio, web pages, Office documents), generating AI-powered notes and summaries, creating multi-speaker podcasts from research, chatting with documents using context-aware AI, searching across materials with full-text and vector search, or running custom content transformations. Supports 16+ AI providers including OpenAI, Anthropic, Google, Ollama, Groq, and Mistral with complete data privacy through self-hosting.
|
|
4
|
-
license: MIT
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.2"
|
|
7
|
-
skill-author: K-Dense Inc.
|
|
8
|
-
openclaw:
|
|
9
|
-
envVars:
|
|
10
|
-
- name: OPEN_NOTEBOOK_URL
|
|
11
|
-
required: true
|
|
12
|
-
description: Open Notebook server URL.
|
|
13
|
-
- name: OPEN_NOTEBOOK_PASSWORD
|
|
14
|
-
required: false
|
|
15
|
-
description: Open Notebook password, if auth is enabled.
|
|
16
|
-
- name: OPEN_NOTEBOOK_ENCRYPTION_KEY
|
|
17
|
-
required: false
|
|
18
|
-
description: Encryption key for stored content, if configured.
|
|
19
|
-
---
|
|
20
|
-
|
|
21
|
-
# Open Notebook
|
|
22
|
-
|
|
23
|
-
## Overview
|
|
24
|
-
|
|
25
|
-
Open Notebook is an open-source, self-hosted alternative to Google's NotebookLM that enables researchers to organize materials, generate AI-powered insights, create podcasts, and have context-aware conversations with their documents — all while maintaining complete data privacy.
|
|
26
|
-
|
|
27
|
-
Unlike Google's Notebook LM, which has no publicly available API outside of the Enterprise version, Open Notebook provides a comprehensive REST API, supports 16+ AI providers, and runs entirely on your own infrastructure.
|
|
28
|
-
|
|
29
|
-
**Key advantages over NotebookLM:**
|
|
30
|
-
- Full REST API for programmatic access and automation
|
|
31
|
-
- Choice of 16+ AI providers (not locked to Google models)
|
|
32
|
-
- Multi-speaker podcast generation with 1-4 customizable speakers (vs. 2-speaker limit)
|
|
33
|
-
- Complete data sovereignty through self-hosting
|
|
34
|
-
- Open source and fully extensible (MIT license)
|
|
35
|
-
|
|
36
|
-
**Repository:** https://github.com/lfnovo/open-notebook
|
|
37
|
-
|
|
38
|
-
## Quick Start
|
|
39
|
-
|
|
40
|
-
### Prerequisites
|
|
41
|
-
|
|
42
|
-
- Docker Desktop installed
|
|
43
|
-
- API key for at least one AI provider (or local Ollama for free local inference)
|
|
44
|
-
|
|
45
|
-
### Installation
|
|
46
|
-
|
|
47
|
-
Deploy Open Notebook using Docker Compose:
|
|
48
|
-
|
|
49
|
-
```bash
|
|
50
|
-
# Download the docker-compose file
|
|
51
|
-
curl -o docker-compose.yml https://raw.githubusercontent.com/lfnovo/open-notebook/main/docker-compose.yml
|
|
52
|
-
|
|
53
|
-
# Set the required encryption key
|
|
54
|
-
export OPEN_NOTEBOOK_ENCRYPTION_KEY="your-secret-key-here"
|
|
55
|
-
|
|
56
|
-
# Launch the services
|
|
57
|
-
docker-compose up -d
|
|
58
|
-
```
|
|
59
|
-
|
|
60
|
-
Access the application:
|
|
61
|
-
- **Frontend UI:** http://localhost:8502
|
|
62
|
-
- **REST API:** http://localhost:5055
|
|
63
|
-
- **API Documentation:** http://localhost:5055/docs
|
|
64
|
-
|
|
65
|
-
### Configure AI Provider
|
|
66
|
-
|
|
67
|
-
After startup, configure at least one AI provider:
|
|
68
|
-
|
|
69
|
-
1. Navigate to **Settings > API Keys** in the UI
|
|
70
|
-
2. Add credentials for your preferred provider (OpenAI, Anthropic, etc.)
|
|
71
|
-
3. Test the connection and discover available models
|
|
72
|
-
4. Register models for use across the platform
|
|
73
|
-
|
|
74
|
-
Or configure via the REST API:
|
|
75
|
-
|
|
76
|
-
```python
|
|
77
|
-
import requests
|
|
78
|
-
|
|
79
|
-
BASE_URL = "http://localhost:5055/api"
|
|
80
|
-
|
|
81
|
-
# Add a credential for an AI provider
|
|
82
|
-
response = requests.post(f"{BASE_URL}/credentials", json={
|
|
83
|
-
"provider": "openai",
|
|
84
|
-
"name": "My OpenAI Key",
|
|
85
|
-
"api_key": "sk-..."
|
|
86
|
-
})
|
|
87
|
-
credential = response.json()
|
|
88
|
-
|
|
89
|
-
# Discover available models
|
|
90
|
-
response = requests.post(
|
|
91
|
-
f"{BASE_URL}/credentials/{credential['id']}/discover"
|
|
92
|
-
)
|
|
93
|
-
discovered = response.json()
|
|
94
|
-
|
|
95
|
-
# Register discovered models
|
|
96
|
-
requests.post(
|
|
97
|
-
f"{BASE_URL}/credentials/{credential['id']}/register-models",
|
|
98
|
-
json={"model_ids": [m["id"] for m in discovered["models"]]}
|
|
99
|
-
)
|
|
100
|
-
```
|
|
101
|
-
|
|
102
|
-
## Core Features
|
|
103
|
-
|
|
104
|
-
### Notebooks
|
|
105
|
-
Organize research into separate notebooks, each containing sources, notes, and chat sessions.
|
|
106
|
-
|
|
107
|
-
```python
|
|
108
|
-
import requests
|
|
109
|
-
|
|
110
|
-
BASE_URL = "http://localhost:5055/api"
|
|
111
|
-
|
|
112
|
-
# Create a notebook
|
|
113
|
-
response = requests.post(f"{BASE_URL}/notebooks", json={
|
|
114
|
-
"name": "Cancer Genomics Research",
|
|
115
|
-
"description": "Literature review on tumor mutational burden"
|
|
116
|
-
})
|
|
117
|
-
notebook = response.json()
|
|
118
|
-
notebook_id = notebook["id"]
|
|
119
|
-
```
|
|
120
|
-
|
|
121
|
-
### Sources
|
|
122
|
-
Ingest diverse content types including PDFs, videos, audio files, web pages, and Office documents. Sources are processed for full-text and vector search.
|
|
123
|
-
|
|
124
|
-
```python
|
|
125
|
-
# Add a web URL source
|
|
126
|
-
response = requests.post(f"{BASE_URL}/sources", data={
|
|
127
|
-
"url": "https://arxiv.org/abs/2301.00001",
|
|
128
|
-
"notebook_id": notebook_id,
|
|
129
|
-
"process_async": "true"
|
|
130
|
-
})
|
|
131
|
-
source = response.json()
|
|
132
|
-
|
|
133
|
-
# Upload a PDF file
|
|
134
|
-
with open("paper.pdf", "rb") as f:
|
|
135
|
-
response = requests.post(
|
|
136
|
-
f"{BASE_URL}/sources",
|
|
137
|
-
data={"notebook_id": notebook_id},
|
|
138
|
-
files={"file": ("paper.pdf", f, "application/pdf")}
|
|
139
|
-
)
|
|
140
|
-
```
|
|
141
|
-
|
|
142
|
-
### Notes
|
|
143
|
-
Create and manage notes (human or AI-generated) associated with notebooks.
|
|
144
|
-
|
|
145
|
-
```python
|
|
146
|
-
# Create a human note
|
|
147
|
-
response = requests.post(f"{BASE_URL}/notes", json={
|
|
148
|
-
"title": "Key Findings",
|
|
149
|
-
"content": "TMB correlates with immunotherapy response in NSCLC...",
|
|
150
|
-
"note_type": "human",
|
|
151
|
-
"notebook_id": notebook_id
|
|
152
|
-
})
|
|
153
|
-
```
|
|
154
|
-
|
|
155
|
-
### Context-Aware Chat
|
|
156
|
-
Chat with your research materials using AI that cites sources.
|
|
157
|
-
|
|
158
|
-
```python
|
|
159
|
-
# Create a chat session
|
|
160
|
-
session = requests.post(f"{BASE_URL}/chat/sessions", json={
|
|
161
|
-
"notebook_id": notebook_id,
|
|
162
|
-
"title": "TMB Discussion"
|
|
163
|
-
}).json()
|
|
164
|
-
|
|
165
|
-
# Send a message with context from sources
|
|
166
|
-
response = requests.post(f"{BASE_URL}/chat/execute", json={
|
|
167
|
-
"session_id": session["id"],
|
|
168
|
-
"message": "What are the key biomarkers for immunotherapy response?",
|
|
169
|
-
"context": {"include_sources": True, "include_notes": True}
|
|
170
|
-
})
|
|
171
|
-
```
|
|
172
|
-
|
|
173
|
-
### Search
|
|
174
|
-
Search across all materials using full-text or vector (semantic) search.
|
|
175
|
-
|
|
176
|
-
```python
|
|
177
|
-
# Vector search across the knowledge base
|
|
178
|
-
results = requests.post(f"{BASE_URL}/search", json={
|
|
179
|
-
"query": "tumor mutational burden immunotherapy",
|
|
180
|
-
"search_type": "vector",
|
|
181
|
-
"limit": 10
|
|
182
|
-
}).json()
|
|
183
|
-
|
|
184
|
-
# Ask a question with AI-powered answer
|
|
185
|
-
answer = requests.post(f"{BASE_URL}/search/ask/simple", json={
|
|
186
|
-
"query": "How does TMB predict checkpoint inhibitor response?"
|
|
187
|
-
}).json()
|
|
188
|
-
```
|
|
189
|
-
|
|
190
|
-
### Podcast Generation
|
|
191
|
-
Generate professional multi-speaker podcasts from research materials with 1-4 customizable speakers.
|
|
192
|
-
|
|
193
|
-
```python
|
|
194
|
-
# Generate a podcast episode
|
|
195
|
-
job = requests.post(f"{BASE_URL}/podcasts/generate", json={
|
|
196
|
-
"notebook_id": notebook_id,
|
|
197
|
-
"episode_profile_id": episode_profile_id,
|
|
198
|
-
"speaker_profile_ids": [speaker1_id, speaker2_id]
|
|
199
|
-
}).json()
|
|
200
|
-
|
|
201
|
-
# Check generation status
|
|
202
|
-
status = requests.get(f"{BASE_URL}/podcasts/jobs/{job['job_id']}").json()
|
|
203
|
-
|
|
204
|
-
# Download audio when ready
|
|
205
|
-
audio = requests.get(
|
|
206
|
-
f"{BASE_URL}/podcasts/episodes/{status['episode_id']}/audio"
|
|
207
|
-
)
|
|
208
|
-
```
|
|
209
|
-
|
|
210
|
-
### Content Transformations
|
|
211
|
-
Apply custom AI-powered transformations to content for summarization, extraction, and analysis.
|
|
212
|
-
|
|
213
|
-
```python
|
|
214
|
-
# Create a custom transformation
|
|
215
|
-
transform = requests.post(f"{BASE_URL}/transformations", json={
|
|
216
|
-
"name": "extract_methods",
|
|
217
|
-
"title": "Extract Methods",
|
|
218
|
-
"description": "Extract methodology details from papers",
|
|
219
|
-
"prompt": "Extract and summarize the methodology section...",
|
|
220
|
-
"apply_default": False
|
|
221
|
-
}).json()
|
|
222
|
-
|
|
223
|
-
# Execute transformation on text
|
|
224
|
-
result = requests.post(f"{BASE_URL}/transformations/execute", json={
|
|
225
|
-
"transformation_id": transform["id"],
|
|
226
|
-
"input_text": "...",
|
|
227
|
-
"model_id": "model_id_here"
|
|
228
|
-
}).json()
|
|
229
|
-
```
|
|
230
|
-
|
|
231
|
-
## Supported AI Providers
|
|
232
|
-
|
|
233
|
-
Open Notebook supports 16+ AI providers through the Esperanto library:
|
|
234
|
-
|
|
235
|
-
| Provider | LLM | Embedding | Speech-to-Text | Text-to-Speech |
|
|
236
|
-
|----------|-----|-----------|----------------|----------------|
|
|
237
|
-
| OpenAI | Yes | Yes | Yes | Yes |
|
|
238
|
-
| Anthropic | Yes | No | No | No |
|
|
239
|
-
| Google GenAI | Yes | Yes | No | Yes |
|
|
240
|
-
| Vertex AI | Yes | Yes | No | Yes |
|
|
241
|
-
| Ollama | Yes | Yes | No | No |
|
|
242
|
-
| Groq | Yes | No | Yes | No |
|
|
243
|
-
| Mistral | Yes | Yes | No | No |
|
|
244
|
-
| Azure OpenAI | Yes | Yes | No | No |
|
|
245
|
-
| DeepSeek | Yes | No | No | No |
|
|
246
|
-
| xAI | Yes | No | No | No |
|
|
247
|
-
| OpenRouter | Yes | No | No | No |
|
|
248
|
-
| ElevenLabs | No | No | Yes | Yes |
|
|
249
|
-
| Perplexity | Yes | No | No | No |
|
|
250
|
-
| Voyage | No | Yes | No | No |
|
|
251
|
-
|
|
252
|
-
## Environment Variables
|
|
253
|
-
|
|
254
|
-
Key configuration variables for Docker deployment:
|
|
255
|
-
|
|
256
|
-
| Variable | Description | Default |
|
|
257
|
-
|----------|-------------|---------|
|
|
258
|
-
| `OPEN_NOTEBOOK_ENCRYPTION_KEY` | **Required.** Secret key for encrypting stored credentials | None |
|
|
259
|
-
| `SURREAL_URL` | SurrealDB connection URL | `ws://surrealdb:8000/rpc` |
|
|
260
|
-
| `SURREAL_NAMESPACE` | Database namespace | `open_notebook` |
|
|
261
|
-
| `SURREAL_DATABASE` | Database name | `open_notebook` |
|
|
262
|
-
| `OPEN_NOTEBOOK_PASSWORD` | Optional password protection for the UI | None |
|
|
263
|
-
|
|
264
|
-
## API Reference
|
|
265
|
-
|
|
266
|
-
The REST API is available at `http://localhost:5055/api` with interactive documentation at `/docs`.
|
|
267
|
-
|
|
268
|
-
Core endpoint groups:
|
|
269
|
-
- `/api/notebooks` - Notebook CRUD and source association
|
|
270
|
-
- `/api/sources` - Source ingestion, processing, and retrieval
|
|
271
|
-
- `/api/notes` - Note management
|
|
272
|
-
- `/api/chat/sessions` - Chat session management
|
|
273
|
-
- `/api/chat/execute` - Chat message execution
|
|
274
|
-
- `/api/search` - Full-text and vector search
|
|
275
|
-
- `/api/podcasts` - Podcast generation and management
|
|
276
|
-
- `/api/transformations` - Content transformation pipelines
|
|
277
|
-
- `/api/models` - AI model configuration and discovery
|
|
278
|
-
- `/api/credentials` - Provider credential management
|
|
279
|
-
|
|
280
|
-
For complete API reference with all endpoints and request/response formats, see `references/api_reference.md`.
|
|
281
|
-
|
|
282
|
-
## Architecture
|
|
283
|
-
|
|
284
|
-
Open Notebook uses a modern stack:
|
|
285
|
-
- **Backend:** Python with FastAPI
|
|
286
|
-
- **Database:** SurrealDB (document + relational)
|
|
287
|
-
- **AI Integration:** LangChain with the Esperanto multi-provider library
|
|
288
|
-
- **Frontend:** Next.js with React
|
|
289
|
-
- **Deployment:** Docker Compose with persistent volumes
|
|
290
|
-
|
|
291
|
-
## Important Notes
|
|
292
|
-
|
|
293
|
-
- Open Notebook requires Docker for deployment
|
|
294
|
-
- At least one AI provider must be configured for AI features to work
|
|
295
|
-
- For free local inference without API costs, use Ollama
|
|
296
|
-
- The `OPEN_NOTEBOOK_ENCRYPTION_KEY` must be set before first launch and kept consistent across restarts
|
|
297
|
-
- All data is stored locally in Docker volumes for complete data sovereignty
|