package-doctor 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package_doctor-0.1.0/.github/ISSUE_TEMPLATE/bug.yml +29 -0
- package_doctor-0.1.0/.github/ISSUE_TEMPLATE/config.yml +5 -0
- package_doctor-0.1.0/.github/ISSUE_TEMPLATE/exposure-map.yml +56 -0
- package_doctor-0.1.0/.github/PULL_REQUEST_TEMPLATE.md +20 -0
- package_doctor-0.1.0/.github/workflows/ci.yml +45 -0
- package_doctor-0.1.0/.github/workflows/live-contracts.yml +28 -0
- package_doctor-0.1.0/.gitignore +15 -0
- package_doctor-0.1.0/CHANGELOG.md +118 -0
- package_doctor-0.1.0/CODE_OF_CONDUCT.md +52 -0
- package_doctor-0.1.0/CONTRIBUTING.md +117 -0
- package_doctor-0.1.0/LICENSE +21 -0
- package_doctor-0.1.0/PKG-INFO +561 -0
- package_doctor-0.1.0/README.md +526 -0
- package_doctor-0.1.0/SECURITY.md +79 -0
- package_doctor-0.1.0/pyproject.toml +70 -0
- package_doctor-0.1.0/research/analyse.py +216 -0
- package_doctor-0.1.0/research/benchmark_vs_pip_audit.py +119 -0
- package_doctor-0.1.0/research/bulk_scan.py +206 -0
- package_doctor-0.1.0/research/eval-repos.txt +62 -0
- package_doctor-0.1.0/research/evaluate_repos.py +211 -0
- package_doctor-0.1.0/research/suggest_map.py +220 -0
- package_doctor-0.1.0/src/package_doctor/__init__.py +3 -0
- package_doctor-0.1.0/src/package_doctor/__main__.py +6 -0
- package_doctor-0.1.0/src/package_doctor/analysis.py +141 -0
- package_doctor-0.1.0/src/package_doctor/cache.py +157 -0
- package_doctor-0.1.0/src/package_doctor/cli.py +378 -0
- package_doctor-0.1.0/src/package_doctor/data/exposure.toml +396 -0
- package_doctor-0.1.0/src/package_doctor/exposure.py +180 -0
- package_doctor-0.1.0/src/package_doctor/models.py +195 -0
- package_doctor-0.1.0/src/package_doctor/parsers/__init__.py +3 -0
- package_doctor-0.1.0/src/package_doctor/parsers/discovery.py +392 -0
- package_doctor-0.1.0/src/package_doctor/report.py +459 -0
- package_doctor-0.1.0/src/package_doctor/risk.py +254 -0
- package_doctor-0.1.0/src/package_doctor/sources/__init__.py +3 -0
- package_doctor-0.1.0/src/package_doctor/sources/client.py +227 -0
- package_doctor-0.1.0/src/package_doctor/sources/ecosystems.py +58 -0
- package_doctor-0.1.0/src/package_doctor/sources/exploitability.py +137 -0
- package_doctor-0.1.0/src/package_doctor/sources/osv.py +287 -0
- package_doctor-0.1.0/src/package_doctor/sources/pypi.py +200 -0
- package_doctor-0.1.0/src/package_doctor/sourcescan.py +280 -0
- package_doctor-0.1.0/tests/conftest.py +63 -0
- package_doctor-0.1.0/tests/test_analysis.py +133 -0
- package_doctor-0.1.0/tests/test_cache.py +224 -0
- package_doctor-0.1.0/tests/test_cli.py +266 -0
- package_doctor-0.1.0/tests/test_client.py +241 -0
- package_doctor-0.1.0/tests/test_exploitability.py +158 -0
- package_doctor-0.1.0/tests/test_exploitability_sources.py +180 -0
- package_doctor-0.1.0/tests/test_exposure.py +263 -0
- package_doctor-0.1.0/tests/test_hardening.py +386 -0
- package_doctor-0.1.0/tests/test_live_contracts.py +275 -0
- package_doctor-0.1.0/tests/test_osv.py +208 -0
- package_doctor-0.1.0/tests/test_parsers.py +305 -0
- package_doctor-0.1.0/tests/test_pypi.py +237 -0
- package_doctor-0.1.0/tests/test_report.py +218 -0
- package_doctor-0.1.0/tests/test_risk.py +271 -0
- package_doctor-0.1.0/tests/test_sourcescan.py +170 -0
|
@@ -0,0 +1,29 @@
|
|
|
1
|
+
name: Bug report
|
|
2
|
+
description: The tool reported something wrong, or crashed
|
|
3
|
+
labels: ["bug"]
|
|
4
|
+
body:
|
|
5
|
+
- type: textarea
|
|
6
|
+
id: what
|
|
7
|
+
attributes:
|
|
8
|
+
label: What happened?
|
|
9
|
+
description: |
|
|
10
|
+
If a finding looks wrong, paste the row and what you expected instead.
|
|
11
|
+
|
|
12
|
+
Two kinds are especially worth reporting:
|
|
13
|
+
- a package reported as fine that isn't (a false negative — the worst
|
|
14
|
+
failure this tool has)
|
|
15
|
+
- wording that reads as a judgement on a maintainer rather than a
|
|
16
|
+
description of risk
|
|
17
|
+
validations:
|
|
18
|
+
required: true
|
|
19
|
+
- type: textarea
|
|
20
|
+
id: output
|
|
21
|
+
attributes:
|
|
22
|
+
label: Output
|
|
23
|
+
description: "`package-doctor explain <name>` is usually the most useful thing to paste."
|
|
24
|
+
render: text
|
|
25
|
+
- type: input
|
|
26
|
+
id: version
|
|
27
|
+
attributes:
|
|
28
|
+
label: Version
|
|
29
|
+
placeholder: "package-doctor --version, and your Python version"
|
|
@@ -0,0 +1,56 @@
|
|
|
1
|
+
name: Exposure map correction
|
|
2
|
+
description: A package is in the wrong list, or missing from it
|
|
3
|
+
title: "[map] "
|
|
4
|
+
labels: ["exposure-map"]
|
|
5
|
+
body:
|
|
6
|
+
- type: markdown
|
|
7
|
+
attributes:
|
|
8
|
+
value: |
|
|
9
|
+
The map is ~710 judgement calls made by one person, so corrections are
|
|
10
|
+
the most useful thing this project receives.
|
|
11
|
+
|
|
12
|
+
Wrong entries matter more than missing ones: a bad entry makes the tool
|
|
13
|
+
lie, a gap only makes it quieter.
|
|
14
|
+
- type: input
|
|
15
|
+
id: package
|
|
16
|
+
attributes:
|
|
17
|
+
label: Package name
|
|
18
|
+
placeholder: e.g. diskcache
|
|
19
|
+
validations:
|
|
20
|
+
required: true
|
|
21
|
+
- type: dropdown
|
|
22
|
+
id: change
|
|
23
|
+
attributes:
|
|
24
|
+
label: What should change?
|
|
25
|
+
options:
|
|
26
|
+
- Add it — an attacker can reach this
|
|
27
|
+
- Remove it — it is not at a trust boundary
|
|
28
|
+
- Move it — it is in the wrong category
|
|
29
|
+
- Mark it mature — at a boundary, but finished by design
|
|
30
|
+
validations:
|
|
31
|
+
required: true
|
|
32
|
+
- type: textarea
|
|
33
|
+
id: reasoning
|
|
34
|
+
attributes:
|
|
35
|
+
label: What convinced you?
|
|
36
|
+
description: |
|
|
37
|
+
The part that actually matters. Ideally an advisory, a function, or an
|
|
38
|
+
API that shows the package handling data from outside the trust
|
|
39
|
+
boundary — or shows that it doesn't.
|
|
40
|
+
|
|
41
|
+
Please don't answer "it's popular" or "it sounds security-related".
|
|
42
|
+
`bandit` is a security scanner, not a target. `xxhash` is not
|
|
43
|
+
cryptographic. `num2words` has three unfixed advisories that turn out
|
|
44
|
+
to be a maintainer account compromise rather than anything it does
|
|
45
|
+
with input.
|
|
46
|
+
placeholder: |
|
|
47
|
+
diskcache stores values by pickling them to disk, and
|
|
48
|
+
GHSA-... says "DiskCache has unsafe pickle deserialization".
|
|
49
|
+
Anything that can write the cache directory gets code execution.
|
|
50
|
+
validations:
|
|
51
|
+
required: true
|
|
52
|
+
- type: input
|
|
53
|
+
id: evidence
|
|
54
|
+
attributes:
|
|
55
|
+
label: Link to evidence
|
|
56
|
+
placeholder: https://osv.dev/vulnerability/GHSA-... or a link to the source
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
## What this changes
|
|
2
|
+
|
|
3
|
+
<!-- One or two lines. -->
|
|
4
|
+
|
|
5
|
+
## Why
|
|
6
|
+
|
|
7
|
+
<!--
|
|
8
|
+
For exposure map changes, this is the part that gets reviewed. Quote the
|
|
9
|
+
advisory, function or API that convinced you — "adds foo, its load() unpickles
|
|
10
|
+
whatever you hand it, see GHSA-xxxx" takes ten seconds to check.
|
|
11
|
+
|
|
12
|
+
"It's popular" and "it sounds security-related" are not reasons. See
|
|
13
|
+
CONTRIBUTING.md.
|
|
14
|
+
-->
|
|
15
|
+
|
|
16
|
+
## Checklist
|
|
17
|
+
|
|
18
|
+
- [ ] `pytest` passes
|
|
19
|
+
- [ ] Behaviour changes have a test
|
|
20
|
+
- [ ] If this touches the exposure map, the reasoning is above
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
name: CI
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
push:
|
|
5
|
+
branches: [main]
|
|
6
|
+
pull_request:
|
|
7
|
+
|
|
8
|
+
jobs:
|
|
9
|
+
test:
|
|
10
|
+
runs-on: ubuntu-latest
|
|
11
|
+
strategy:
|
|
12
|
+
fail-fast: false
|
|
13
|
+
matrix:
|
|
14
|
+
python-version: ["3.10", "3.11", "3.12", "3.13", "3.14"]
|
|
15
|
+
steps:
|
|
16
|
+
- uses: actions/checkout@v4
|
|
17
|
+
- uses: actions/setup-python@v5
|
|
18
|
+
with:
|
|
19
|
+
python-version: ${{ matrix.python-version }}
|
|
20
|
+
- run: pip install -e ".[dev]"
|
|
21
|
+
- run: ruff check src tests research
|
|
22
|
+
# The suite is offline by design: HTTP is served by httpx.MockTransport,
|
|
23
|
+
# so a flaky upstream API can never turn into a red build.
|
|
24
|
+
#
|
|
25
|
+
# The coverage floor sits just under where the suite is today, to catch a
|
|
26
|
+
# drop rather than to chase a number. The decision logic - risk, exposure,
|
|
27
|
+
# sourcescan, osv - is well above it; the gap is mostly presentation.
|
|
28
|
+
- run: pytest -q --cov=package_doctor --cov-fail-under=85
|
|
29
|
+
|
|
30
|
+
# Installs the built wheel, not the editable checkout, and loads the exposure
|
|
31
|
+
# map from outside the repository. An editable install hides a wheel that is
|
|
32
|
+
# missing its data file - which is exactly what an over-broad .gitignore once
|
|
33
|
+
# produced - and this is the only step that would have caught it.
|
|
34
|
+
package:
|
|
35
|
+
runs-on: ubuntu-latest
|
|
36
|
+
steps:
|
|
37
|
+
- uses: actions/checkout@v4
|
|
38
|
+
- uses: actions/setup-python@v5
|
|
39
|
+
with:
|
|
40
|
+
python-version: "3.12"
|
|
41
|
+
- run: pip install build
|
|
42
|
+
- run: python -m build --wheel
|
|
43
|
+
- run: pip install dist/*.whl
|
|
44
|
+
- run: package-doctor --version
|
|
45
|
+
- run: cd / && python -c "from package_doctor.exposure import load_exposure_map; load_exposure_map(); print('exposure map loads from the installed wheel')"
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
name: Live API contracts
|
|
2
|
+
|
|
3
|
+
# These hit real PyPI, OSV, ecosyste.ms, CISA and FIRST endpoints to catch
|
|
4
|
+
# upstream drift - a renamed field, a changed envelope, a new pagination cap -
|
|
5
|
+
# which unit tests with stubbed transport cannot see.
|
|
6
|
+
#
|
|
7
|
+
# Deliberately NOT part of pull-request CI: four third-party services have no
|
|
8
|
+
# business deciding whether somebody's PR is red. A failure here means the tool
|
|
9
|
+
# is now misreading real data and needs a fix, not that a contributor did
|
|
10
|
+
# anything wrong.
|
|
11
|
+
|
|
12
|
+
on:
|
|
13
|
+
schedule:
|
|
14
|
+
- cron: "17 6 * * 1" # Mondays, early UTC
|
|
15
|
+
workflow_dispatch:
|
|
16
|
+
|
|
17
|
+
jobs:
|
|
18
|
+
contracts:
|
|
19
|
+
runs-on: ubuntu-latest
|
|
20
|
+
steps:
|
|
21
|
+
- uses: actions/checkout@v4
|
|
22
|
+
- uses: actions/setup-python@v5
|
|
23
|
+
with:
|
|
24
|
+
python-version: "3.12"
|
|
25
|
+
- run: pip install -e ".[dev]"
|
|
26
|
+
# -rs reports skips, so an outage is visible in the log rather than
|
|
27
|
+
# looking like a clean pass.
|
|
28
|
+
- run: pytest -m live -q -rs
|
|
@@ -0,0 +1,15 @@
|
|
|
1
|
+
__pycache__/
|
|
2
|
+
*.py[cod]
|
|
3
|
+
.venv/
|
|
4
|
+
venv/
|
|
5
|
+
build/
|
|
6
|
+
dist/
|
|
7
|
+
*.egg-info/
|
|
8
|
+
.pytest_cache/
|
|
9
|
+
.coverage
|
|
10
|
+
htmlcov/
|
|
11
|
+
.ruff_cache/
|
|
12
|
+
.DS_Store
|
|
13
|
+
# Research output only. Anchored so it cannot match src/package_doctor/data/,
|
|
14
|
+
# which hatchling would otherwise drop from the wheel.
|
|
15
|
+
/data/
|
|
@@ -0,0 +1,118 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented here.
|
|
4
|
+
|
|
5
|
+
The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
|
|
6
|
+
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
|
+
|
|
8
|
+
## [Unreleased]
|
|
9
|
+
|
|
10
|
+
## [0.1.0] - 2026-09-12
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
|
|
14
|
+
- Two-axis risk model: a dependency is only escalated when it sits at a trust
|
|
15
|
+
boundary **and** shows evidence that nobody is left to fix it.
|
|
16
|
+
- Exposure map of ~710 packages across 14 trust-boundary categories, plus
|
|
17
|
+
reviewed-and-cleared and mature-by-design lists.
|
|
18
|
+
- Advisory history from OSV, read as never-fixed / fixed-late / fixed-timely
|
|
19
|
+
rather than a naive time-to-fix.
|
|
20
|
+
- Exploitability ranking via CISA KEV and FIRST EPSS, scoped to the pinned
|
|
21
|
+
version — turns "affected by 35 advisories" into which one to read first.
|
|
22
|
+
- Reachability: AST-parses your own source and reports where each dependency is
|
|
23
|
+
imported.
|
|
24
|
+
- `scan`, `explain` and `cache` commands, JSON output, `--version`, and
|
|
25
|
+
`--fail-on` exit codes for CI.
|
|
26
|
+
- Research harness (`research/`) for bulk-scanning PyPI and for deciding what
|
|
27
|
+
the exposure map should cover next.
|
|
28
|
+
- `requirements/*.txt` is discovered, and `-r` includes are followed within
|
|
29
|
+
the project, so the pip-tools and Django layouts scan without flags.
|
|
30
|
+
- `explain` says when no pinned version was known and advisory matching was
|
|
31
|
+
therefore skipped; the version to match is given with `--pin`.
|
|
32
|
+
|
|
33
|
+
### Fixed
|
|
34
|
+
|
|
35
|
+
- Advisory ranges are read with OSV's semantics: `last_affected` closes a
|
|
36
|
+
range inclusively and an `introduced` with no closing event runs to
|
|
37
|
+
infinity. Only `fixed` was recognised before, so an advisory closed by
|
|
38
|
+
`last_affected` read as "no published fix" and an open-ended one never
|
|
39
|
+
matched the pinned version. Across eight real projects that produced 36 of
|
|
40
|
+
64 *act on these* verdicts, all wrong - Django 6 escalated for a CSRF bug
|
|
41
|
+
closed at 1.2.7 - and one missed advisory. "Never fixed" now means the
|
|
42
|
+
latest release is still affected; a range that ends before the latest
|
|
43
|
+
release with no fix named is reported as closed, not counted either way.
|
|
44
|
+
- GHSA and PYSEC records of the same CVE are merged before counting.
|
|
45
|
+
cryptography 46.0.7 read as "affected by 7 advisories"; it is four.
|
|
46
|
+
- Import sites are named relative to the project directory rather than to
|
|
47
|
+
the package directory being walked, so zulip's `analytics/models.py` and
|
|
48
|
+
`zerver/models.py` no longer both report as `models.py`.
|
|
49
|
+
- The project's own package, uv workspace members, and anything a lockfile
|
|
50
|
+
records with a local source are skipped and listed as such, instead of
|
|
51
|
+
being assessed as dependencies. mlflow's report led with mlflow.
|
|
52
|
+
- A wildcard such as `click==8.*` is read as a range, not stored as the
|
|
53
|
+
version "8.*", which OSV matched against nothing.
|
|
54
|
+
- `requirements/<env>/*.txt` is discovered, one level deep.
|
|
55
|
+
- The scan header and the JSON say how many packages had no pinned version
|
|
56
|
+
and so had no advisory matching, instead of implying "no advisories".
|
|
57
|
+
- An archived repository with a release inside the stale window is one weak
|
|
58
|
+
signal, worded "the code may have moved", rather than proof of abandonment.
|
|
59
|
+
google-cloud-bigquery's archived repo is a monorepo migration.
|
|
60
|
+
|
|
61
|
+
### Security
|
|
62
|
+
|
|
63
|
+
- Terminal output strips control and formatting characters, including whole
|
|
64
|
+
CSI and OSC sequences, from every string that originates outside the
|
|
65
|
+
process: lockfile versions, scanned file names, API-supplied URLs and
|
|
66
|
+
advisory ids. `rich` neutralises markup but passes a raw ESC through, so a
|
|
67
|
+
hostile lockfile could previously wrap a row in a hyperlink to somewhere
|
|
68
|
+
else or clear the screen.
|
|
69
|
+
- The source scanner and the dependency-file parsers read only regular files,
|
|
70
|
+
with a bounded read rather than a size check. A FIFO named `evil.py` or a
|
|
71
|
+
symlink to `/dev/zero` in a scanned checkout previously hung the scan.
|
|
72
|
+
- Dependency files are capped at 32 MB, and a refused file is reported rather
|
|
73
|
+
than silently treated as empty.
|
|
74
|
+
- Pathologically nested TOML and JSON is treated as malformed. `tomllib`
|
|
75
|
+
raises `RecursionError`, not `TOMLDecodeError`, on a lockfile a few thousand
|
|
76
|
+
brackets deep, which previously ended the scan with a traceback.
|
|
77
|
+
- CVE identifiers from OSV aliases are validated against their exact form
|
|
78
|
+
before being spliced into the EPSS query string.
|
|
79
|
+
- The response cache file is created `0600` before SQLite opens it, rather
|
|
80
|
+
than tightened afterwards.
|
|
81
|
+
|
|
82
|
+
### Changed
|
|
83
|
+
|
|
84
|
+
- PyPI responses are reduced to the fields the tool reads before they are
|
|
85
|
+
cached: the release timeline collapses to one upload time and one yanked
|
|
86
|
+
flag per release. The full body was over a megabyte per package on average,
|
|
87
|
+
which put any lockfile past about 230 packages over the cache's size limit
|
|
88
|
+
and into a cycle of pruning fresh entries and refetching them. Reduced
|
|
89
|
+
entries live under a new cache key; old ones expire on their own.
|
|
90
|
+
|
|
91
|
+
- A response over the size cap is now its own outcome. It was reported as
|
|
92
|
+
"not found on PyPI", which was untrue, and never cached, so every run
|
|
93
|
+
re-downloaded the body just to abandon it. It is now reported as too large
|
|
94
|
+
to read, remembered for the cache lifetime, and the package's advisories
|
|
95
|
+
are still counted from OSV. PyPI's cap is 64 MB, above the 32 MB default,
|
|
96
|
+
because its body is reduced the moment it is parsed; the largest today is
|
|
97
|
+
12.6 MB.
|
|
98
|
+
|
|
99
|
+
### Measured
|
|
100
|
+
|
|
101
|
+
On sixty open source repositories (`research/eval-repos.txt`, harness in
|
|
102
|
+
`research/evaluate_repos.py`), 12,973 packages:
|
|
103
|
+
|
|
104
|
+
- Advisory version matching agrees with OSV's own version-scoped query on
|
|
105
|
+
6,727 of 6,728 distinct pinned pairs; the one disagreement is a
|
|
106
|
+
SEMVER-typed range OSV ignores for PyPI and this tool reads.
|
|
107
|
+
- Against `pip-audit` on the same pins: 1,672 vulnerabilities found by both,
|
|
108
|
+
**zero missed**, one found additionally (the same range).
|
|
109
|
+
- Of 108,193 reported import sites, 108,162 verified against the source line;
|
|
110
|
+
the rest are pytest's `_pytest` and `py` modules, attributed correctly.
|
|
111
|
+
- Every act verdict resting on a maintenance signal was read; none was found
|
|
112
|
+
wrong on the tool's stated criteria.
|
|
113
|
+
- Exposure map tested for predictive validity across 3,000 packages: packages it
|
|
114
|
+
marks exposed carry advisories at ~2.6x the rate of packages it reviewed and
|
|
115
|
+
cleared.
|
|
116
|
+
|
|
117
|
+
[Unreleased]: https://github.com/binuka200/package-doctor/compare/v0.1.0...HEAD
|
|
118
|
+
[0.1.0]: https://github.com/binuka200/package-doctor/releases/tag/v0.1.0
|
|
@@ -0,0 +1,52 @@
|
|
|
1
|
+
# Code of Conduct
|
|
2
|
+
|
|
3
|
+
## Our pledge
|
|
4
|
+
|
|
5
|
+
We pledge to make participation in this project a harassment-free experience for
|
|
6
|
+
everyone, regardless of age, body size, visible or invisible disability,
|
|
7
|
+
ethnicity, sex characteristics, gender identity and expression, level of
|
|
8
|
+
experience, education, socio-economic status, nationality, personal appearance,
|
|
9
|
+
race, religion, or sexual identity and orientation.
|
|
10
|
+
|
|
11
|
+
## Our standards
|
|
12
|
+
|
|
13
|
+
Examples of behaviour that contributes to a positive environment:
|
|
14
|
+
|
|
15
|
+
- Demonstrating empathy and kindness toward other people
|
|
16
|
+
- Being respectful of differing opinions, viewpoints, and experiences
|
|
17
|
+
- Giving and gracefully accepting constructive feedback
|
|
18
|
+
- Accepting responsibility for our mistakes, and learning from them
|
|
19
|
+
- Focusing on what is best for the community
|
|
20
|
+
|
|
21
|
+
Examples of unacceptable behaviour:
|
|
22
|
+
|
|
23
|
+
- Sexualised language or imagery, and sexual attention or advances of any kind
|
|
24
|
+
- Trolling, insulting or derogatory comments, and personal or political attacks
|
|
25
|
+
- Public or private harassment
|
|
26
|
+
- Publishing others' private information without their explicit permission
|
|
27
|
+
- Other conduct which could reasonably be considered inappropriate in a
|
|
28
|
+
professional setting
|
|
29
|
+
|
|
30
|
+
## A note specific to this project
|
|
31
|
+
|
|
32
|
+
This tool makes public statements about whether open source packages are
|
|
33
|
+
maintained. Much of that work is unpaid, and many of the packages it names are
|
|
34
|
+
the work of volunteers who gave what they could.
|
|
35
|
+
|
|
36
|
+
Discuss packages, not the people who wrote them. "No releases since 2021,
|
|
37
|
+
repository archived" is a fact that helps a user. "This project is dead and the
|
|
38
|
+
maintainer abandoned it" is a judgement about a person, and it is not welcome
|
|
39
|
+
here — in issues, in pull requests, or in the tool's own output.
|
|
40
|
+
|
|
41
|
+
## Enforcement
|
|
42
|
+
|
|
43
|
+
Instances of abusive, harassing, or otherwise unacceptable behaviour may be
|
|
44
|
+
reported to the maintainer at **contact@binukajayaweera.dev**.
|
|
45
|
+
|
|
46
|
+
All complaints will be reviewed and investigated promptly and fairly. The
|
|
47
|
+
maintainer is obligated to respect the privacy and security of the reporter.
|
|
48
|
+
|
|
49
|
+
## Attribution
|
|
50
|
+
|
|
51
|
+
Adapted from the [Contributor Covenant](https://www.contributor-covenant.org),
|
|
52
|
+
version 2.1.
|
|
@@ -0,0 +1,117 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
Thanks for looking. Most of what this project needs is not code.
|
|
4
|
+
|
|
5
|
+
## The most useful thing you can do
|
|
6
|
+
|
|
7
|
+
**Argue with [`exposure.toml`](src/package_doctor/data/exposure.toml).**
|
|
8
|
+
|
|
9
|
+
That file is ~710 judgement calls about which Python packages sit somewhere an
|
|
10
|
+
attacker can reach. Every one of them was made by one person. Some are wrong,
|
|
11
|
+
and the wrong ones are worse than the missing ones — a bad entry makes the tool
|
|
12
|
+
lie, while a gap only makes it quieter.
|
|
13
|
+
|
|
14
|
+
If you know a corner of Python well — Django, ML, crypto, async, packaging —
|
|
15
|
+
twenty minutes reading the relevant category is worth more than a month of new
|
|
16
|
+
entries.
|
|
17
|
+
|
|
18
|
+
### What belongs in the map
|
|
19
|
+
|
|
20
|
+
A package belongs in a `[category.*]` list when it:
|
|
21
|
+
|
|
22
|
+
1. parses, decodes, renders, verifies or transports data that commonly comes
|
|
23
|
+
from outside your trust boundary, **or**
|
|
24
|
+
2. makes an authentication or authorisation decision, **or**
|
|
25
|
+
3. builds queries or commands from caller-supplied values.
|
|
26
|
+
|
|
27
|
+
### What does not
|
|
28
|
+
|
|
29
|
+
- **"It's popular."** Not a criterion. `six` is everywhere and belongs nowhere
|
|
30
|
+
near the map.
|
|
31
|
+
- **"It sounds security-related."** `bandit`, `semgrep` and `pip-audit` are
|
|
32
|
+
security *tools*, not trust boundaries. `xxhash` and `mmh3` are not
|
|
33
|
+
cryptographic despite the names.
|
|
34
|
+
- **"It has lots of CVEs."** A reason to look, never the answer. `num2words`
|
|
35
|
+
has three advisories with no fix — and reading them shows a maintainer
|
|
36
|
+
account compromise, not anything the library does with input. It is in
|
|
37
|
+
`[reviewed] not_exposed` for exactly that reason.
|
|
38
|
+
|
|
39
|
+
**Read the advisories before deciding.** Several entries look harmless from
|
|
40
|
+
their one-line description and turn out to be textbook: `diskcache` is "a disk
|
|
41
|
+
cache" whose advisories say *unsafe pickle deserialization*; `apscheduler` is
|
|
42
|
+
"a task scheduler" whose serializers had remote code execution.
|
|
43
|
+
|
|
44
|
+
### The three lists
|
|
45
|
+
|
|
46
|
+
| list | meaning |
|
|
47
|
+
| --- | --- |
|
|
48
|
+
| `[category.*]` | an attacker can reach this |
|
|
49
|
+
| `[reviewed] not_exposed` | checked, and they can't — with a note saying why the obvious guess is wrong |
|
|
50
|
+
| `[stable] mature` | they can, but the library is finished, so ignore its age |
|
|
51
|
+
|
|
52
|
+
The second list matters as much as the first. It is how "we checked and it's
|
|
53
|
+
fine" stays distinguishable from "nobody has looked", which is a distinction the
|
|
54
|
+
rest of the tool is built on.
|
|
55
|
+
|
|
56
|
+
### Finding something worth deciding
|
|
57
|
+
|
|
58
|
+
```bash
|
|
59
|
+
package-doctor scan . --json -o scan.json
|
|
60
|
+
python research/suggest_map.py --scan scan.json
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
That ranks the packages the map has no opinion about by how much the silence
|
|
64
|
+
costs, and prints what each one is for. Working the top of that list beats
|
|
65
|
+
reading down a popularity ranking.
|
|
66
|
+
|
|
67
|
+
### Opening the PR
|
|
68
|
+
|
|
69
|
+
One category change per PR where you can, and say **why** in the description —
|
|
70
|
+
ideally quoting the advisory or the API that convinced you. "Adds `foo`" is hard
|
|
71
|
+
to review; "adds `foo`, its `load()` unpickles whatever you hand it, see
|
|
72
|
+
GHSA-xxxx" takes ten seconds.
|
|
73
|
+
|
|
74
|
+
## Code
|
|
75
|
+
|
|
76
|
+
```bash
|
|
77
|
+
git clone https://github.com/binuka200/package-doctor
|
|
78
|
+
cd package-doctor
|
|
79
|
+
python -m venv .venv && source .venv/bin/activate
|
|
80
|
+
pip install -e ".[dev]"
|
|
81
|
+
pytest # offline, ~1s
|
|
82
|
+
pytest -m live # hits the real APIs; opt-in, not run in PR CI
|
|
83
|
+
ruff check src tests research
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
The suite is offline by design — HTTP is served by `httpx.MockTransport`, so no
|
|
87
|
+
upstream outage can redden a build. CI enforces 85% coverage and a clean
|
|
88
|
+
`ruff check`; the rule set is in `pyproject.toml` and is deliberately small.
|
|
89
|
+
|
|
90
|
+
### Principles worth knowing before you change behaviour
|
|
91
|
+
|
|
92
|
+
These are load-bearing, and there are tests asserting each of them:
|
|
93
|
+
|
|
94
|
+
- **Missing data is never a bad score.** No advisory history means *unknown*,
|
|
95
|
+
not *healthy*. No repository declared means *unknown*, not *abandoned*. If you
|
|
96
|
+
find code that turns a null into a zero, that's a bug.
|
|
97
|
+
- **Release age alone never produces a finding.** It cannot tell an abandoned
|
|
98
|
+
library from a finished one. Two weak signals must agree, or one
|
|
99
|
+
authoritative signal must fire.
|
|
100
|
+
- **A guess can never demand action.** An inferred category can raise something
|
|
101
|
+
to *watch*; only a curated one can reach *act on this*.
|
|
102
|
+
- **There is no aggregate health score, and there should not be.** A single
|
|
103
|
+
number is the thing users can't act on and maintainers can't argue with.
|
|
104
|
+
|
|
105
|
+
### Research scripts
|
|
106
|
+
|
|
107
|
+
`research/` holds the instruments used to study the ecosystem, not the product.
|
|
108
|
+
They write to `data/`, which is gitignored. Run `bulk_scan.py` before
|
|
109
|
+
`analyse.py` or `suggest_map.py --dataset`.
|
|
110
|
+
|
|
111
|
+
## Reporting a problem
|
|
112
|
+
|
|
113
|
+
A finding that reads as a judgement on a maintainer rather than a description of
|
|
114
|
+
risk **is a bug** — please report it. Most unmaintained packages are the work of
|
|
115
|
+
volunteers who gave what they could, and the tool's wording should reflect that.
|
|
116
|
+
|
|
117
|
+
For anything security-sensitive, see [SECURITY.md](SECURITY.md).
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Binuka Jayaweera
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|