mortis-spatial 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (56) hide show
  1. mortis_spatial-0.1.0/CHANGELOG.md +217 -0
  2. mortis_spatial-0.1.0/CITATION.cff +42 -0
  3. mortis_spatial-0.1.0/CONTRIBUTING.md +123 -0
  4. mortis_spatial-0.1.0/LICENSE +21 -0
  5. mortis_spatial-0.1.0/MANIFEST.in +7 -0
  6. mortis_spatial-0.1.0/PKG-INFO +244 -0
  7. mortis_spatial-0.1.0/README.md +177 -0
  8. mortis_spatial-0.1.0/pyproject.toml +113 -0
  9. mortis_spatial-0.1.0/setup.cfg +4 -0
  10. mortis_spatial-0.1.0/src/mortis/__init__.py +395 -0
  11. mortis_spatial-0.1.0/src/mortis/analysis.py +2237 -0
  12. mortis_spatial-0.1.0/src/mortis/annotate.py +666 -0
  13. mortis_spatial-0.1.0/src/mortis/cli.py +352 -0
  14. mortis_spatial-0.1.0/src/mortis/compare.py +271 -0
  15. mortis_spatial-0.1.0/src/mortis/data/drug_names.db +0 -0
  16. mortis_spatial-0.1.0/src/mortis/exceptions.py +146 -0
  17. mortis_spatial-0.1.0/src/mortis/filter.py +355 -0
  18. mortis_spatial-0.1.0/src/mortis/image.py +416 -0
  19. mortis_spatial-0.1.0/src/mortis/interactive.py +117 -0
  20. mortis_spatial-0.1.0/src/mortis/io.py +815 -0
  21. mortis_spatial-0.1.0/src/mortis/organization.py +528 -0
  22. mortis_spatial-0.1.0/src/mortis/pathway.py +475 -0
  23. mortis_spatial-0.1.0/src/mortis/plotting.py +813 -0
  24. mortis_spatial-0.1.0/src/mortis/preprocessing.py +732 -0
  25. mortis_spatial-0.1.0/src/mortis/reproducibility.py +483 -0
  26. mortis_spatial-0.1.0/src/mortis/stats.py +610 -0
  27. mortis_spatial-0.1.0/src/mortis/viz.py +1880 -0
  28. mortis_spatial-0.1.0/src/mortis_spatial.egg-info/PKG-INFO +244 -0
  29. mortis_spatial-0.1.0/src/mortis_spatial.egg-info/SOURCES.txt +54 -0
  30. mortis_spatial-0.1.0/src/mortis_spatial.egg-info/dependency_links.txt +1 -0
  31. mortis_spatial-0.1.0/src/mortis_spatial.egg-info/entry_points.txt +2 -0
  32. mortis_spatial-0.1.0/src/mortis_spatial.egg-info/requires.txt +48 -0
  33. mortis_spatial-0.1.0/src/mortis_spatial.egg-info/top_level.txt +1 -0
  34. mortis_spatial-0.1.0/tests/__init__.py +0 -0
  35. mortis_spatial-0.1.0/tests/test_analysis_extended.py +176 -0
  36. mortis_spatial-0.1.0/tests/test_annotate.py +309 -0
  37. mortis_spatial-0.1.0/tests/test_cli.py +207 -0
  38. mortis_spatial-0.1.0/tests/test_compare.py +205 -0
  39. mortis_spatial-0.1.0/tests/test_core.py +270 -0
  40. mortis_spatial-0.1.0/tests/test_correctness_vs_reference.py +341 -0
  41. mortis_spatial-0.1.0/tests/test_filter.py +240 -0
  42. mortis_spatial-0.1.0/tests/test_image.py +171 -0
  43. mortis_spatial-0.1.0/tests/test_io.py +361 -0
  44. mortis_spatial-0.1.0/tests/test_messages.py +166 -0
  45. mortis_spatial-0.1.0/tests/test_new_analysis.py +320 -0
  46. mortis_spatial-0.1.0/tests/test_new_spatial_methods.py +464 -0
  47. mortis_spatial-0.1.0/tests/test_organization.py +290 -0
  48. mortis_spatial-0.1.0/tests/test_pathway.py +361 -0
  49. mortis_spatial-0.1.0/tests/test_plotting.py +349 -0
  50. mortis_spatial-0.1.0/tests/test_preprocessing.py +329 -0
  51. mortis_spatial-0.1.0/tests/test_reproducibility.py +378 -0
  52. mortis_spatial-0.1.0/tests/test_reproducibility_manifest.py +313 -0
  53. mortis_spatial-0.1.0/tests/test_spatial_structure.py +250 -0
  54. mortis_spatial-0.1.0/tests/test_stats.py +307 -0
  55. mortis_spatial-0.1.0/tests/test_viz.py +951 -0
  56. mortis_spatial-0.1.0/tools/build_drug_vocabulary.py +113 -0
@@ -0,0 +1,217 @@
1
+ # Changelog
2
+
3
+ Notable changes, in [Keep a Changelog](https://keepachangelog.com/en/1.1.0/)
4
+ order, versioned with [SemVer](https://semver.org/).
5
+
6
+ ## [0.1.0], unreleased
7
+
8
+ The first version worth putting a number on. Everything before this was me
9
+ finding out what the package needed to be; the version counter is starting here
10
+ because that is honest, even though the code has been through a lot more than
11
+ "0.1" usually implies.
12
+
13
+ It works, it is tested on real public data, and it is under review for
14
+ publication. What it is not yet is released.
15
+
16
+ ### The four things that were quietly wrong
17
+
18
+ Every one of these produced numbers. Wrong ones. Each was reproduced before it
19
+ was fixed, and each now has a test that fails on the old code.
20
+
21
+ - **Spatial coordinates collapsed when sections were combined.** Batches were
22
+ pushed apart by `i * 1e6` in float32, and float32 runs out of precision at
23
+ around 8.4 million, past which neighbouring pixels round onto each other. On
24
+ a real 52-section cohort, **35 sections lost roughly half their distinct pixel
25
+ positions**, which silently corrupted Moran's I, Geary's C, Gi\*, LISA,
26
+ co-occurrence, spatial domains and spatially-weighted NMF for all of them.
27
+ Now float64, with the offset derived from the actual coordinate range instead
28
+ of a hard-coded constant that assumed your microns were small.
29
+ - **A cache that never noticed the data had changed.** `_get_X` memoised a dense
30
+ copy into `.uns` and never invalidated it, so anything that replaced `.X`,
31
+ `scale()`, `correct_batches()`, sparse normalisation, left every later call
32
+ reading pre-correction data. Wiping `.X` to all zeros still returned the
33
+ original statistics, cheerfully. It also kept up to four full copies of the
34
+ matrix (~2.4 GB on a 100k × 2000 dataset) and wrote them into every saved
35
+ `.h5ad`. Deleted. Rebuilding costs about half a second.
36
+ - **Permutation tests were not reproducible.** `np.random.seed` inside a
37
+ `@njit(parallel=True)` function seeds exactly one worker thread and leaves the
38
+ rest to their own devices, so results depended on how many cores you had. Same
39
+ seed, 1 vs 4 threads, |Δz| up to 0.79, and not even repeatable twice in the
40
+ same process. Each permutation now draws its own independent seed stream.
41
+ - **`compare_groups` was testing pixels as if they were patients.** On simulated
42
+ data with six patients and *no group difference at all*, it called **183 of
43
+ 200** metabolites significant. It still exists for genuine within-section
44
+ comparisons, but it now warns and points at the sample-level path.
45
+
46
+ ### Added
47
+
48
+ - **`mortis.stats`**: the sample-level statistics that make everything else
49
+ legitimate. `pseudobulk()`, `differential_abundance()` (Cliff's δ first,
50
+ Mann-Whitney and BH-FDR as supporting detail, optional bootstrap interval),
51
+ `paired_differential_abundance()` for before/after designs, and
52
+ `cliffs_delta()` on its own.
53
+ - **`mortis.organization`**: differential spatial *organization*. Summarises
54
+ how each metabolite is arranged per section (Moran's I, normalised entropy,
55
+ Gi\* hotspot fraction, Gini), then tests those the same way abundance is
56
+ tested. `compare_abundance_and_organization()` labels which axis moved; the
57
+ "organization only" class is the one nothing else can find.
58
+
59
+ Benchmarked honestly: with 3 planted differences among 30 metabolites, effect
60
+ size alone over-called at every cohort size, while effect size plus FDR gave
61
+ exactly 3 true and 0 false **from six sections per group upward, and nothing
62
+ below it**. Six per arm is the floor, it is documented, and it is pinned by a
63
+ test.
64
+ - **`mortis.compare`**: `cross_cohort_profile()` and `track_flow()`: do two
65
+ drugs move the same metabolites, and does a signature persist, reorganise or
66
+ flip over time.
67
+ - **`mortis.annotate`**: layered chemical-class assignment, class-level
68
+ enrichment, and directional pathway over-representation. On a real
69
+ 2,231-compound untargeted panel the built-in name rules leave about 55%
70
+ unclassified; a `reference=` mapping from HMDB or LIPID MAPS takes that to
71
+ 28%. Both numbers are in the docstring rather than hidden, because a
72
+ classifier that quietly places two-thirds of a panel and says nothing about
73
+ the rest is the exact failure this module exists to avoid.
74
+ - **`mortis.pathway`**: compound names → HMDB/KEGG identifiers → pathways →
75
+ enrichment, in one call, cached on disk. The original plan was to hand
76
+ enrichment to MetaboAnalyst; it turns out to document exactly one REST
77
+ endpoint and enrichment is not it. Doing the statistics locally is better
78
+ anyway, because the background set decides the answer and a web service
79
+ cannot know which compounds *your* instrument saw.
80
+ - **`mortis.reproducibility`**: `export_manifest()` and `verify_manifest()`.
81
+ A sealed JSON recording the environment, every recorded step, and SHA-256
82
+ fingerprints of the inputs and every result table. Hand it to a reviewer:
83
+ they can confirm your analysis reproduces without you sending them any data,
84
+ because checksums only go one way. "Available on reasonable request" verifies
85
+ nothing; this verifies something.
86
+ - **`mortis.viz`**: publication figures. PDF text stays *text* (embedded
87
+ TrueType with a character map), so a co-author can retype a label in
88
+ Illustrator instead of emailing you about it. Only dense scatter interiors get
89
+ rasterised. Every figure carries its parameters and a hash in the PDF
90
+ metadata, so `pdfinfo` will tell you which run made it long after you have
91
+ forgotten. `theme="light"/"dark"` renders on a transparent ground for slides
92
+ and web, and `ion_cmap()` replaces viridis as the ion-image default, still
93
+ perceptually ordered, just less obviously the work of a plotting library.
94
+ - **`validation/run_validation.py`**: the package run end-to-end against a
95
+ public METASPACE study, no simulation and nothing else in the pipeline.
96
+ - **A documentation site** in `web/`, with the API reference generated from the
97
+ installed package at build time so it cannot drift from the code.
98
+
99
+ ### Changed
100
+
101
+ - **The bundled drug list is built from Wikidata instead of DrugBank.**
102
+ DrugBank releases its data under CC BY-NC, which does not permit shipping it
103
+ inside a wheel that anyone, including commercial users, can install from
104
+ PyPI. The vocabulary is now assembled from Wikidata, which is CC0: an entry
105
+ counts as a drug when Wikidata gives it a DrugBank or ATC identifier, or
106
+ files it under medication or pharmaceutical product. That comes to 20,181
107
+ names and 43,697 synonyms, against 17,430 and 45,731 before, and every drug
108
+ in a 19-compound spot check is still found. `tools/build_drug_vocabulary.py`
109
+ rebuilds it. `list_drug_matches()` returns `is_drug` in place of
110
+ `in_drugbank`, and adds an `endogenous` column, since taurine, cholesterol
111
+ and most amino acids carry drug identifiers and `remove_all=True` would
112
+ otherwise delete them from a metabolomics panel without comment.
113
+ - **`leidenalg` and `python-igraph` moved to a `[cluster]` extra.** Both are
114
+ GPL while MORTIS is MIT, so installing them makes the whole environment GPL.
115
+ That is a decision for whoever installs it, not something `pip install
116
+ mortis-spatial` should make on their behalf. `cluster()` and
117
+ `spatial_domains()` raise with the install command when they are missing;
118
+ nothing else in the package touches them.
119
+ - **The Leiden backend is named explicitly.** scanpy is switching its default
120
+ from `leidenalg` to `igraph`, and the two do not give the same partition, so
121
+ an unpinned call would have quietly changed everybody's clusters on a scanpy
122
+ upgrade. Minimum scanpy is now 1.10, which is where the argument appeared.
123
+ - **`correct_batches()` raises when ComBat fails** instead of mean-centring
124
+ each batch and printing a line about it. Substituting a weaker method
125
+ returns data corrected by something other than what was asked for, and other
126
+ than what the methods section will say.
127
+ - **Warnings go through `warnings.warn`.** Three modules printed them to
128
+ stdout, where they could not be filtered, caught or redirected, and were
129
+ invisible to `pytest.warns`.
130
+ - **`mortis.audit` removed.** `export_manifest()` records everything it did,
131
+ plus a verification step it never had. Two receipt systems in one package
132
+ was one too many.
133
+ - **`merge_samples()` no longer writes into the list it was given.** It
134
+ replaced the caller's elements with relabelled copies.
135
+
136
+ - **Spatial statistics stream over metabolite tiles.** The direct formulation
137
+ held several full pixels × metabolites matrices at once. Since the reduction
138
+ is over pixels and every output is one scalar per metabolite, metabolites can
139
+ be processed in tiles, exactly the same arithmetic, reassociated. Combined
140
+ with deriving `W @ (X − mean)` from `W @ X` algebraically, on a 95,751 × 2,231
141
+ dataset: **1.23 s → 0.62 s, working set 4.30 GB → 0.99 GB, bit-identical
142
+ checksums.**
143
+
144
+ Tile size is picked from free RAM rather than cache size, because
145
+ cache-sized tiles measured *slowest*: SciPy walks the whole sparse structure
146
+ of the weights matrix once per tile regardless of how many columns come along
147
+ for the ride, so amortising that beats locality. This was genuinely
148
+ counter-intuitive and the measurement is in the code.
149
+ - Sparse and disk-backed (`backed="r"`) inputs are now covered by tests proving
150
+ they give identical answers to dense ones.
151
+ - `metabolite_colocalization()` gained `metric="cosine_median"`, the measure
152
+ that actually won the ColocML benchmark. Opt-in, not default, because it
153
+ rasterises by coordinate span and scattered coordinates would ask for an
154
+ enormous empty grid, now guarded, after it hung the test suite once.
155
+ - The ColocML citation was wrong. It is Ovchinnikova, Stuart, Rakhlin,
156
+ Nikolenko & Alexandrov, *Bioinformatics* 2020;36(10):3215-3224, not
157
+ "Ryabchykov et al.", which is a paper about something else entirely.
158
+ - `local_moran()` gained a docstring, including an explicit warning that its
159
+ p-values are approximate rather than Anselin's conditional permutation. Fine
160
+ for ranking pixels and drawing a LISA map; not something to report as
161
+ calibrated inference, and it now says so instead of calling them "fast
162
+ analytical p-values".
163
+ - Comments and docstrings across the older modules were rewritten out of
164
+ marketing voice. "Seamless integration", "Intelligently detects",
165
+ "Ultra-fast index creation utilizing Python List Comprehensions" and a
166
+ section header reading `KILLER FEATURES` are all gone. None of it told a
167
+ reader anything they could use.
168
+ - **Minimum Python is now 3.10**, not 3.9. 3.9 was declared and never tested;
169
+ the CI matrix has always started at 3.10. Verified rather than assumed.
170
+ - `test_data/` is explicitly gitignored. The previous global `*.h5ad` rule
171
+ covered two fixtures and missed the patient CSV entirely.
172
+
173
+ ### Fixed
174
+
175
+ - Progress messages crashed on a Windows console using a legacy code page.
176
+ `->`, `>=` and similar characters cannot be encoded in cp1252, so a
177
+ `print()` containing one raised `UnicodeEncodeError` instead of reporting
178
+ progress. The package source is ASCII now, bar a plus-minus in the
179
+ stereodescriptor regex and a micrometre sign in one axis label, and a test
180
+ keeps it that way.
181
+ - Every public function has a docstring. Nineteen of them had none, so `help()`
182
+ and IDE tooltips came back blank even though the documentation site covered
183
+ them.
184
+ - `neighborhood_enrichment(n_jobs=...)` crashed with `ValueError: The number of
185
+ threads must be between 1 and N` when asked for more threads than the machine
186
+ has. Numba fixes its ceiling at import; asking for more is a wish, not an
187
+ error, so the request is clamped.
188
+ - Compatible with pandas 3 and anndata 0.13, both of which arrive by default on
189
+ Python 3.12. Two test assumptions broke there, pandas 3 string columns do not
190
+ support 2-D fancy indexing, and anndata 0.13 lists `.X` under a `None` key in
191
+ `layers`, while the package itself was already correct. Error messages that
192
+ list available layers now filter that `None` out, because showing it to
193
+ someone hunting for a metric name helps nobody.
194
+ - `save_figure()` used `Path.with_suffix("")`, which eats everything after the
195
+ last dot: `two_axis.dark` quietly became `two_axis`, and `figure_v1.2` would
196
+ have lost its version.
197
+ - Two runs of the same analysis produced different files. Result tables sorted
198
+ on effect size with no tiebreaker, so equally-ranked metabolites came out in
199
+ whatever order the sort happened to leave them; and figure PDFs carried the
200
+ wall-clock time. Ties now break on the compound name, and `SOURCE_DATE_EPOCH`
201
+ is honoured, so `diff` is a usable way to ask whether anything changed. The
202
+ public-data validation reproduces byte-for-byte apart from the manifest,
203
+ which records when the run happened on purpose.
204
+ - Error messages said what was wrong but not what to do about it.
205
+ `'sample' not found in adata.obs.` is technically accurate and practically
206
+ useless; it now lists the columns that do exist and, when the name looks like
207
+ a typo, guesses which one you meant. Several also told you to call
208
+ `MORTIS.preprocess()`, which is not how the package is imported.
209
+ - `plot_abundance_vs_organization` and `plot_signature_comparison` drew one dot
210
+ per coordinate, and on a small cohort Cliff's delta takes so few distinct
211
+ values that a whole panel collapses onto a handful of points. The figure
212
+ showed twelve dots while the legend said 160. Marker area now scales with how
213
+ many metabolites share a position, and the figure says so.
214
+ - Group labels on `plot_organization_heatmap` were rotated, so on a two-section
215
+ arm the text was taller than its own band and the group names printed over
216
+ each other.
217
+
@@ -0,0 +1,42 @@
1
+ cff-version: 1.2.0
2
+ title: "MORTIS: cohort-scale analysis for spatial metabolomics"
3
+ message: >-
4
+ Please cite the software using the metadata below, including the version
5
+ you actually ran. Once the paper is out the preferred citation will be the
6
+ article.
7
+ version: 0.1.0
8
+ date-released: "2026-09-30"
9
+ license: MIT
10
+ type: software
11
+ authors:
12
+ - given-names: Faris
13
+ family-names: Hrvat
14
+ email: farishrvatit@gmail.com
15
+ # orcid: "https://orcid.org/0000-0000-0000-0000"
16
+ repository-code: "https://github.com/FarisHrvat/mortis"
17
+ url: "https://farishrvat.github.io/mortis/"
18
+ abstract: >-
19
+ A Python package for downstream analysis of spatial metabolomics data
20
+ (MALDI-MSI, DESI and related imaging mass spectrometry). It tests at the
21
+ patient level rather than the pixel level, measures differential spatial
22
+ organization, whether a metabolite is arranged differently between groups
23
+ independently of how much of it there is, compares cohorts and timepoints,
24
+ and exports publication figures alongside a sealed manifest that lets a
25
+ reviewer verify an analysis reproduces without access to the underlying data.
26
+ keywords:
27
+ - spatial metabolomics
28
+ - imaging mass spectrometry
29
+ - MALDI-MSI
30
+ - pseudobulk
31
+ - spatial statistics
32
+ - reproducibility
33
+ - bioinformatics
34
+ - anndata
35
+ doi: 10.5281/zenodo.23056382
36
+ identifiers:
37
+ - type: doi
38
+ value: 10.5281/zenodo.23056382
39
+ description: Concept DOI, always resolves to the newest version.
40
+ - type: doi
41
+ value: 10.5281/zenodo.23056383
42
+ description: DOI for version 0.1.0.
@@ -0,0 +1,123 @@
1
+ # Contributing
2
+
3
+ Thanks for looking. This is a scientific package, which means a bug here does
4
+ not crash. It prints a number that is wrong, and somebody puts that number in a
5
+ paper. Most of what follows exists because of that.
6
+
7
+ ## Setting up
8
+
9
+ ```bash
10
+ git clone https://github.com/FarisHrvat/mortis.git && cd mortis
11
+ python -m venv .venv && source .venv/bin/activate # or conda/micromamba
12
+ pip install -e ".[dev]"
13
+ ```
14
+
15
+ Python 3.10 or newer. CI runs 3.10, 3.11 and 3.12 on Linux and macOS, so those
16
+ are the versions that are actually promised.
17
+
18
+ Worth knowing: on 3.12 pip resolves **pandas 3 and anndata 0.13**, which behave
19
+ differently from what 3.10 and 3.11 get. If a test passes locally and fails in
20
+ CI, that is the first thing to check.
21
+
22
+ ## Before you open a pull request
23
+
24
+ ```bash
25
+ pytest # 520 tests, about 25 seconds
26
+ ruff check .
27
+ ```
28
+
29
+ Both have to pass. CI runs exactly these.
30
+
31
+ ## Tests
32
+
33
+ The rule: **a test for a bug fix has to fail on the code before the fix.**
34
+
35
+ If it passes both before and after, it is not testing what you think it is. When
36
+ the four defects in `test_reproducibility.py` were fixed, the new tests were run
37
+ against the old code first, 9 of 18 failed, which is how anyone knows they mean
38
+ something.
39
+
40
+ Some practical consequences:
41
+
42
+ - **Check the artefact, not the setting.** "PDF text stays editable" is a
43
+ property of the bytes in the file, so the test looks for `/FontFile2` in the
44
+ PDF and round-trips it through `pdftotext`. Asserting that an rcParam was set
45
+ proves nothing about the file a co-author opens.
46
+ - **Look at figures.** Three real bugs: clipped axis labels, a title landing on
47
+ the panel labels, and a legend sitting on top of the bars, passed every
48
+ assertion and were found by rendering a PNG and looking at it.
49
+ - **Network tests are opt-in.** `tests/test_pathway.py` runs offline against
50
+ captured payloads. The live-service tests need `MORTIS_TEST_NETWORK=1`, so
51
+ the suite neither flakes nor hammers somebody else's free API.
52
+
53
+ ## Statistics
54
+
55
+ One non-negotiable, because it is the whole reason this package exists:
56
+ **pixels are not replicates.**
57
+
58
+ Anything that compares groups of samples goes through `pseudobulk()` first.
59
+ A section has tens of thousands of pixels and one patient; testing the pixels
60
+ inflates *n* by four orders of magnitude. On simulated null data that is the
61
+ difference between 0 findings and 183 of them.
62
+
63
+ If you add a test that compares groups, it needs a null-data case showing it
64
+ does not invent results.
65
+
66
+ ## Writing
67
+
68
+ Comments explain **why**, not what. The code already says what.
69
+
70
+ ```python
71
+ # Bad, restates the line below it
72
+ # Set the number of threads
73
+ nb.set_num_threads(n)
74
+
75
+ # Good, says the thing you cannot see
76
+ # Numba fixes its ceiling at import from the core count, so asking for more
77
+ # than the machine has raises rather than just using what is available.
78
+ nb.set_num_threads(int(np.clip(requested, 1, nb.config.NUMBA_NUM_THREADS)))
79
+ ```
80
+
81
+ Please avoid marketing voice. "Seamless integration", "intelligently detects"
82
+ and "ultra-fast" have all been removed from this codebase once already and none
83
+ of them told a reader anything actionable. Plain sentences, and a joke now and
84
+ then is fine, the package is named after rigor mortis.
85
+
86
+ Document limits where they exist. `classify_compounds` says out loud that it
87
+ leaves ~55% of an untargeted panel unclassified, and
88
+ `compare_abundance_and_organization` says it needs six sections per arm. A
89
+ number you are slightly embarrassed by is worth more than a claim nobody
90
+ checked.
91
+
92
+ ## Data
93
+
94
+ **Never commit patient data.** `test_data/` is gitignored and stays that way.
95
+ Derived figures are fine; the arrays that made them are not.
96
+
97
+ `validation/` runs on public METASPACE data. If you extend it, keep it that
98
+ way, so that anybody can run it.
99
+
100
+ ## Docs
101
+
102
+ The site lives in `web/`. `web/api.json` is generated from the installed
103
+ package, so do not hand-edit it:
104
+
105
+ ```bash
106
+ python web/build_api.py
107
+ ```
108
+
109
+ CI regenerates it on every deploy, and fails if an exported function is missing
110
+ from a group in `web/build_api.py`, which is how new functions avoid quietly
111
+ going undocumented.
112
+
113
+ ## Releasing
114
+
115
+ 1. Update `CHANGELOG.md`, real sentences, not a list of commit subjects.
116
+ 2. Bump `version` in `pyproject.toml` and `CITATION.cff`.
117
+ 3. Tag it. The publish workflow does the rest.
118
+
119
+ ## Anything else
120
+
121
+ Open an issue. A failing snippet and the output you expected is plenty, no
122
+ template to fill in.
123
+
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Faris Hrvat
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,7 @@
1
+ # The sdist should carry enough to rebuild and check the package, not just run
2
+ # it: the tests, and the script that regenerates the bundled drug vocabulary.
3
+ include CHANGELOG.md
4
+ include CITATION.cff
5
+ include CONTRIBUTING.md
6
+ include tools/build_drug_vocabulary.py
7
+ recursive-include tests *.py
@@ -0,0 +1,244 @@
1
+ Metadata-Version: 2.4
2
+ Name: mortis-spatial
3
+ Version: 0.1.0
4
+ Summary: Cohort-scale analysis for spatial metabolomics: patient-level statistics, differential spatial organization, and publication figures.
5
+ Author-email: Faris Hrvat <farishrvatit@gmail.com>
6
+ License: MIT
7
+ Project-URL: Homepage, https://farishrvat.github.io/mortis/
8
+ Project-URL: Documentation, https://farishrvat.github.io/mortis/
9
+ Project-URL: Repository, https://github.com/FarisHrvat/mortis
10
+ Project-URL: Changelog, https://github.com/FarisHrvat/mortis/blob/main/CHANGELOG.md
11
+ Keywords: spatial metabolomics,imaging mass spectrometry,MALDI,DESI,bioinformatics,anndata,scanpy
12
+ Classifier: Development Status :: 3 - Alpha
13
+ Classifier: Intended Audience :: Science/Research
14
+ Classifier: Programming Language :: Python :: 3
15
+ Classifier: Programming Language :: Python :: 3.10
16
+ Classifier: Programming Language :: Python :: 3.11
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: License :: OSI Approved :: MIT License
19
+ Classifier: Operating System :: OS Independent
20
+ Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
21
+ Classifier: Topic :: Scientific/Engineering :: Visualization
22
+ Requires-Python: >=3.10
23
+ Description-Content-Type: text/markdown
24
+ License-File: LICENSE
25
+ Requires-Dist: anndata>=0.10.0
26
+ Requires-Dist: numpy>=1.24.0
27
+ Requires-Dist: scipy>=1.10.0
28
+ Requires-Dist: pandas>=2.0.0
29
+ Requires-Dist: scanpy>=1.10.0
30
+ Requires-Dist: umap-learn>=0.5.0
31
+ Requires-Dist: scikit-learn>=1.3.0
32
+ Requires-Dist: matplotlib>=3.7.0
33
+ Requires-Dist: seaborn>=0.12.0
34
+ Requires-Dist: statsmodels>=0.14.0
35
+ Requires-Dist: numba>=0.58.0
36
+ Requires-Dist: threadpoolctl>=3.1.0
37
+ Requires-Dist: h5py>=3.9.0
38
+ Requires-Dist: openpyxl>=3.1.0
39
+ Requires-Dist: tifffile>=2023.1.0
40
+ Requires-Dist: psutil>=5.9.0
41
+ Provides-Extra: cluster
42
+ Requires-Dist: leidenalg>=0.10.0; extra == "cluster"
43
+ Requires-Dist: python-igraph>=0.10.0; extra == "cluster"
44
+ Provides-Extra: harmony
45
+ Requires-Dist: harmonypy>=0.0.9; extra == "harmony"
46
+ Provides-Extra: cli
47
+ Requires-Dist: pyyaml>=6.0; extra == "cli"
48
+ Provides-Extra: dev
49
+ Requires-Dist: pytest>=7.0; extra == "dev"
50
+ Requires-Dist: pytest-cov; extra == "dev"
51
+ Requires-Dist: ruff; extra == "dev"
52
+ Requires-Dist: build; extra == "dev"
53
+ Requires-Dist: esda; extra == "dev"
54
+ Requires-Dist: libpysal; extra == "dev"
55
+ Requires-Dist: pyyaml>=6.0; extra == "dev"
56
+ Requires-Dist: leidenalg>=0.10.0; extra == "dev"
57
+ Requires-Dist: python-igraph>=0.10.0; extra == "dev"
58
+ Provides-Extra: fast-io
59
+ Requires-Dist: pyarrow; extra == "fast-io"
60
+ Requires-Dist: python-calamine; extra == "fast-io"
61
+ Provides-Extra: rds
62
+ Requires-Dist: pyreadr>=0.5.0; extra == "rds"
63
+ Provides-Extra: image-network
64
+ Requires-Dist: networkx; extra == "image-network"
65
+ Requires-Dist: adjustText; extra == "image-network"
66
+ Dynamic: license-file
67
+
68
+ <h1 align="center">
69
+ <img src="https://raw.githubusercontent.com/FarisHrvat/mortis/main/web/assets/logo.svg" width="46" alt=""><br>
70
+ MORTIS
71
+ </h1>
72
+
73
+ <p align="center">
74
+ <b>Cohort-scale analysis for spatial metabolomics.</b><br>
75
+ Patient-level statistics, differential spatial organization, and figures you can submit.
76
+ </p>
77
+
78
+ <p align="center">
79
+ <a href="https://farishrvat.github.io/mortis/"><b>Read the documentation</b></a>
80
+ </p>
81
+
82
+ <p align="center">
83
+ <a href="https://github.com/FarisHrvat/mortis/actions/workflows/test.yml"><img src="https://github.com/FarisHrvat/mortis/actions/workflows/test.yml/badge.svg" alt="Tests"></a>
84
+ <a href="https://www.python.org/"><img src="https://img.shields.io/badge/python-3.10%2B-blue.svg" alt="Python 3.10+"></a>
85
+ <img src="https://img.shields.io/badge/tested-Linux%20%7C%20macOS%20%7C%20Windows-lightgrey.svg" alt="Linux, macOS, Windows">
86
+ <a href="https://github.com/FarisHrvat/mortis/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-MIT-green.svg" alt="MIT licence"></a>
87
+ <img src="https://img.shields.io/badge/version-0.1.0-orange.svg" alt="version 0.1.0">
88
+ <a href="https://doi.org/10.5281/zenodo.23056382"><img src="https://zenodo.org/badge/DOI/10.5281/zenodo.23056382.svg" alt="DOI"></a>
89
+ </p>
90
+
91
+ ---
92
+
93
+ ## What it does
94
+
95
+ Imaging mass spectrometry tells you *where* a metabolite is. Most analyses then
96
+ throw that away and ask only *how much*, which is the question bulk
97
+ metabolomics already answered, more cheaply.
98
+
99
+ MORTIS asks both, and asks them at the level where the statistics actually hold:
100
+
101
+ - **Patient-level testing.** A section has 30,000 pixels and one patient. Test
102
+ the pixels and you inflate your sample size by four orders of magnitude. On
103
+ simulated null data that turns 0 real findings into 183 significant ones.
104
+ `pseudobulk()` comes first here, and it is not optional.
105
+ - **Differential spatial *organization*.** A metabolite can sit at identical
106
+ abundance in two groups and be arranged completely differently, diffuse in
107
+ one, pooled into foci in the other. That finding is invisible to every
108
+ abundance test and to bulk metabolomics entirely.
109
+ - **Cohorts, not sections.** Compare two drugs, or the same patients before and
110
+ after treatment, and ask whether a signature persists, reorganises, or flips.
111
+ - **Figures and receipts.** Vector PDF, SVG and EPS whose text stays editable,
112
+ plus PNG, JPEG and TIFF at whatever DPI the journal asks for. Fonts, sizes,
113
+ colours and DPI are all yours to set. Every figure can carry a sealed
114
+ manifest a reviewer checks your re-run against, without you sending them a
115
+ single byte of patient data.
116
+
117
+ ## What it is not
118
+
119
+ It is not an acquisition or peak-picking tool. It starts from a feature table
120
+ or an `.h5ad`, so extraction, alignment and annotation happen upstream in
121
+ SCiLS, METASPACE, Cardinal or the vendor software.
122
+
123
+ It does not identify compounds. It takes the names your annotation pipeline
124
+ gave you, and it cannot tell a confident match from a shaky one beyond the
125
+ score your instrument software already wrote.
126
+
127
+ It is not built for a single section. Most of what it adds is about comparing
128
+ groups of patients, and on one section a good deal of it will refuse to run
129
+ rather than give you a p-value that counts pixels as replicates.
130
+
131
+ It does not read raw vendor formats or imzML. It starts from a peak-picked
132
+ table: `.csv`, `.tsv`, `.txt`, `.xlsx`, `.parquet`, `.rds` or `.h5ad`. The
133
+ delimiter and the decimal mark are worked out from the file, so a
134
+ semicolon-and-comma export out of a European Excel reads without editing.
135
+ For anything else, read it with whatever library does and hand the frame to
136
+ `mt.from_dataframe()`.
137
+
138
+ ## Install
139
+
140
+ ```bash
141
+ pip install mortis-spatial
142
+ ```
143
+
144
+ Leiden clustering needs two GPL packages, which are not installed by default
145
+ because MORTIS is MIT and the choice of pulling GPL code into your environment
146
+ should be yours:
147
+
148
+ ```bash
149
+ pip install "mortis-spatial[cluster]"
150
+ ```
151
+
152
+ Everything else works without them, and `spatial_domains_kmeans()` finds
153
+ tissue domains if you would rather not add GPL code at all.
154
+
155
+ ```python
156
+ import mortis as mt
157
+
158
+ adata = mt.preprocess(mt.read_metabolomics_data("section.h5ad"))
159
+
160
+ # collapse pixels to patients, then test, in that order
161
+ pb = mt.pseudobulk(adata, sample_key="patient")
162
+ ab = mt.differential_abundance(pb, "response", "R", "NR")
163
+
164
+ # and ask the question only imaging can answer
165
+ org = mt.spatial_organization(adata, sample_key="section")
166
+ do = mt.differential_spatial_organization(org, "response", "R", "NR")
167
+
168
+ mt.compare_abundance_and_organization(ab, do) # which axis actually moved?
169
+ ```
170
+
171
+ The [documentation](https://farishrvat.github.io/mortis/) has the guided tour,
172
+ every function with its parameters, the figure gallery, and a validation run on
173
+ a public METASPACE study using nothing but this package.
174
+
175
+ ## Or without writing any Python
176
+
177
+ ```bash
178
+ mortis template > analysis.yaml # a commented starting point
179
+ mortis run analysis.yaml # results, figures and a manifest
180
+ ```
181
+
182
+ The config that produced a result is a better methods section than one written
183
+ from memory, so every run writes it back out beside the results.
184
+
185
+ **In a container**, when you would rather not install anything:
186
+
187
+ ```bash
188
+ docker build -t mortis .
189
+ docker run --rm -u "$(id -u):$(id -g)" -v "$PWD:/work" mortis run /work/analysis.yaml
190
+ ```
191
+
192
+ **On a cluster**, [`hpc/`](https://github.com/FarisHrvat/mortis/blob/main/hpc) has Slurm and PBS templates, an Apptainer
193
+ definition for sites that will not permit `pip install`, and a conda
194
+ environment for the ones that will. The scripts derive thread limits from the
195
+ scheduler's allocation and set them before Python starts, which is the
196
+ difference between using your cores and oversubscribing a shared node.
197
+
198
+ ## Does it work?
199
+
200
+ `validation/run_validation.py` downloads a public imaging study, rebuilds the
201
+ pixel matrices from the ion images, and runs the whole pipeline, no simulation,
202
+ no private data, no other package:
203
+
204
+ ```bash
205
+ python validation/run_validation.py
206
+ ```
207
+
208
+ 24 sections, 250,514 pixels. The positive control (two different plant species)
209
+ separates on 4 ions at FDR < 0.05. The deliberately underpowered control returns
210
+ nothing, which is the right answer rather than a disappointing one.
211
+
212
+ ## Development
213
+
214
+ ```bash
215
+ git clone https://github.com/FarisHrvat/mortis.git && cd mortis
216
+ pip install -e ".[dev]"
217
+ pytest && ruff check .
218
+ ```
219
+
220
+ 610 tests, run against Python 3.10 to 3.14 on Linux, Windows and Apple
221
+ Silicon macOS. Intel macOS is not in CI because GitHub retired the last
222
+ Intel runner; the dependencies all ship x86-64 wheels, so it should work
223
+ there, but I have not tested it. See
224
+ [CONTRIBUTING.md](https://github.com/FarisHrvat/mortis/blob/main/CONTRIBUTING.md).
225
+
226
+ ## Citation
227
+
228
+ Cite the DOI, [10.5281/zenodo.23056382](https://doi.org/10.5281/zenodo.23056382),
229
+ which always resolves to the newest version. To pin the exact version you ran,
230
+ each release has its own DOI on the same record. Full metadata is in
231
+ [`CITATION.cff`](https://github.com/FarisHrvat/mortis/blob/main/CITATION.cff),
232
+ and the article will be the preferred citation once it is out.
233
+
234
+ `filter_drugs()` matches against a drug-name list built from Wikidata, which
235
+ is CC0. `run_harmony()` implements Korsunsky et al. 2019, and
236
+ `annotate_pathways()` calls MetaboAnalyst and KEGG. Each is linked from the
237
+ function's own documentation, and should be cited alongside MORTIS if you use
238
+ it.
239
+
240
+
241
+ ## Licence
242
+
243
+ MIT, see [LICENSE](https://github.com/FarisHrvat/mortis/blob/main/LICENSE).
244
+