goad-toolkit 0.2.7__tar.gz → 0.2.9__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511/.claude/settings.local.json +7 -0
- goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511/CHANGELOG.md +152 -0
- goad_toolkit-0.2.9/.remember/archive.md +3 -0
- goad_toolkit-0.2.9/.remember/logs/memory-2026-08-12.log +28 -0
- goad_toolkit-0.2.9/.remember/recent.md +7 -0
- goad_toolkit-0.2.9/.remember/tmp/capture-alive +1 -0
- goad_toolkit-0.2.9/.remember/tmp/capture-alive.d/90574e63-97cd-4c11-87e5-23b656c56fd9 +0 -0
- goad_toolkit-0.2.9/.remember/tmp/last-save-ts +1 -0
- goad_toolkit-0.2.9/.remember/tmp/post-tool-ran +0 -0
- goad_toolkit-0.2.9/.remember/tmp/save-session.pid +1 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/tmp/session-slug +1 -1
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/MCP_SERVER.md +7 -7
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/PKG-INFO +1 -1
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/README.md +30 -0
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/docs/02-pipelines.md +56 -35
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/docs/03-plot-composition.md +61 -3
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/docs/04-five-families.md +48 -14
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/docs/05-distributions.md +52 -7
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/docs/06-models-and-residuals.md +44 -0
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/docs/08-api-reference.md +137 -6
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/docs/09-analysis-method.md +53 -2
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/goad_mcp.py +77 -12
- goad_toolkit-0.2.9/img/null-distribution.png +0 -0
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/pyproject.toml +2 -1
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/src/goad_toolkit/analytics.py +230 -34
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/src/goad_toolkit/datatransforms.py +146 -3
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/src/goad_toolkit/distributions.py +16 -0
- goad_toolkit-0.2.9/src/goad_toolkit/visualizer.py +1635 -0
- goad_toolkit-0.2.9/tests/test_datatransforms.py +257 -0
- goad_toolkit-0.2.9/tests/test_distributions.py +195 -0
- goad_toolkit-0.2.9/tests/test_nulldistribution.py +156 -0
- goad_toolkit-0.2.9/tests/test_visualizer.py +782 -0
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/uv.lock +48 -1
- goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511/src/goad_toolkit/visualizer.py +0 -545
- goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511/tests/test_distributions.py +0 -72
- goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511/tests/test_visualizer.py +0 -113
- goad_toolkit-0.2.7/.remember/tmp/capture-alive +0 -1
- goad_toolkit-0.2.7/.remember/tmp/last-save-ts +0 -1
- goad_toolkit-0.2.7/.remember/tmp/save-session.pid +0 -1
- goad_toolkit-0.2.7/CHANGELOG.md +0 -37
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/settings.local.json +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/.gitignore +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/.python-version +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/MCP_SERVER.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/README.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/demo/linear.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/docs/01-goal-oriented-analysis.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/docs/02-pipelines.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/docs/03-plot-composition.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/docs/04-five-families.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/docs/05-distributions.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/docs/06-models-and-residuals.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/docs/07-visual-critique.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/docs/08-api-reference.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/docs/09-analysis-method.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/docs/10-teaching-path.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/docs/README.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/goad_mcp.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/img/distribution_fit.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/img/goaded.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/img/linear_results.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/img/null-distribution.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/img/residuals.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/img/zscores.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/pyproject.toml +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/src/goad_toolkit/__init__.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/src/goad_toolkit/analytics.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/src/goad_toolkit/cli.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/src/goad_toolkit/config.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/src/goad_toolkit/dataprocessor.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/src/goad_toolkit/datatransforms.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/src/goad_toolkit/distributions.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/src/goad_toolkit/filehandler.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/src/goad_toolkit/models.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/src/goad_toolkit/visualizer.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/tests/test_cli.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/tests/test_datatransforms.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/tests/test_distributions.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.claude/worktrees/quizzical-solomon-b62511/tests/test_filehandler.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/tests/test_nulldistribution.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/tests/test_visualizer.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9/.claude/worktrees/quizzical-solomon-b62511}/uv.lock +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.gitignore +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.python-version +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/.gitignore +0 -0
- /goad_toolkit-0.2.7/.remember/logs/autonomous/save-195516.log → /goad_toolkit-0.2.9/.remember/logs/autonomous/save-104759.log +0 -0
- /goad_toolkit-0.2.7/.remember/logs/autonomous/save-195720.log → /goad_toolkit-0.2.9/.remember/logs/autonomous/save-195516.log +0 -0
- /goad_toolkit-0.2.7/.remember/logs/autonomous/save-195933.log → /goad_toolkit-0.2.9/.remember/logs/autonomous/save-195720.log +0 -0
- /goad_toolkit-0.2.7/.remember/logs/autonomous/save-200831.log → /goad_toolkit-0.2.9/.remember/logs/autonomous/save-195933.log +0 -0
- /goad_toolkit-0.2.7/.remember/logs/autonomous/save-201034.log → /goad_toolkit-0.2.9/.remember/logs/autonomous/save-200831.log +0 -0
- /goad_toolkit-0.2.7/.remember/logs/hook-errors.log → /goad_toolkit-0.2.9/.remember/logs/autonomous/save-201034.log +0 -0
- /goad_toolkit-0.2.7/.remember/tmp/capture-alive.d/90574e63-97cd-4c11-87e5-23b656c56fd9 → /goad_toolkit-0.2.9/.remember/logs/hook-errors.log +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/logs/memory-2026-08-10.log +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/now.md +0 -0
- /goad_toolkit-0.2.7/.remember/tmp/post-tool-ran → /goad_toolkit-0.2.9/.remember/tmp/capture-alive.d/16b4788a-5507-4216-98ec-06bac13a9aba +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/tmp/case-divergence +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/tmp/last-ndc.ts +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/tmp/last-save.json +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/.remember/tmp/now-day +0 -0
- /goad_toolkit-0.2.7/.remember/today-2026-08-10.md → /goad_toolkit-0.2.9/.remember/today-2026-08-10.done.md +0 -0
- {goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/CHANGELOG.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/demo/linear.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/docs/01-goal-oriented-analysis.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/docs/07-visual-critique.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/docs/10-teaching-path.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/docs/README.md +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/img/distribution_fit.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/img/goaded.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/img/linear_results.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/img/residuals.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/img/zscores.png +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/src/goad_toolkit/__init__.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/src/goad_toolkit/cli.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/src/goad_toolkit/config.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/src/goad_toolkit/dataprocessor.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/src/goad_toolkit/filehandler.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/src/goad_toolkit/models.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/tests/test_cli.py +0 -0
- {goad_toolkit-0.2.7 → goad_toolkit-0.2.9}/tests/test_filehandler.py +0 -0
|
@@ -0,0 +1,152 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
## 0.2.7
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
|
|
7
|
+
- `ScatterPlot`, `RegPlot` and `CorrelationHeatmap`: the relation plots for two-variable
|
|
8
|
+
questions, in their order of use — a scatter first, a fit once the scatter has shown which
|
|
9
|
+
kind, a correlation matrix when there are too many pairs to plot by hand.
|
|
10
|
+
- `ScatterPlot` defaults `alpha` below 1, since overplotting is the standard way a scatter
|
|
11
|
+
misreports its own density.
|
|
12
|
+
- `RegPlot` wraps `sns.regplot`, which is axis-level and so can draw onto a layer's axis.
|
|
13
|
+
`scatter=False` lays a fit over an existing `ScatterPlot`; `lowess` together with
|
|
14
|
+
`order != 1` raises, since both answer "what shape" and only one can apply.
|
|
15
|
+
- `CorrelationHeatmap` pins the colour scale to `[-1, 1]` centred on 0 and defaults to a
|
|
16
|
+
diverging colormap, so the sign of a correlation stays readable. Non-numeric columns are
|
|
17
|
+
dropped, and a frame with nothing numeric left raises rather than drawing an empty grid.
|
|
18
|
+
- docs/04 §4.4 covers the three and their layering; docs/08 §8.10 documents the seaborn
|
|
19
|
+
axis-label override — a seaborn call with `x=`/`y=` writes the column names over
|
|
20
|
+
`PlotSettings`.
|
|
21
|
+
|
|
22
|
+
## 0.2.6
|
|
23
|
+
|
|
24
|
+
### Added
|
|
25
|
+
|
|
26
|
+
- `NullDistribution` and `NullResult`: shuffle one column, re-measure, and compare the
|
|
27
|
+
observed value against the cloud that produces. The statistic stays the caller's —
|
|
28
|
+
`NullDistribution` takes a callable and ships no ready-made statistics.
|
|
29
|
+
- `NullResult.p_value` counts rather than tests. Both one-sided counts include the observed
|
|
30
|
+
arrangement, so the smallest value it can return is `1 / (n_iter + 1)`.
|
|
31
|
+
- `NullPlot`: the null as a grey `HistogramPlot` with the observed value marked, its value
|
|
32
|
+
and p-value in the legend, stated as a bound (`p < …`) once the count floor is reached,
|
|
33
|
+
since an exact number there would overstate what a thousand shuffles support.
|
|
34
|
+
- `VerticalLine`, the numeric counterpart to `VerticalDate` and the layer `NullPlot`
|
|
35
|
+
composes.
|
|
36
|
+
- docs/06 §6.7 carries the prose, including the two things the shuffle test does not defend
|
|
37
|
+
against: the garden of forking paths, and a confounder that shuffling breaks along with
|
|
38
|
+
the association.
|
|
39
|
+
|
|
40
|
+
## 0.2.5
|
|
41
|
+
|
|
42
|
+
### Added
|
|
43
|
+
|
|
44
|
+
- `RegexFeature` transform: one new column derived from a regular expression over a text
|
|
45
|
+
column, in `count` (occurrences per row), `has` (bool) or `extract` (first capture group,
|
|
46
|
+
NaN where nothing matched) mode. The output column is named by `feature`, because
|
|
47
|
+
`Pipeline.add(name=...)` already claims `name` for the step.
|
|
48
|
+
- `mode="extract"` reports its match rate through `loguru`: it is the one mode that fails
|
|
49
|
+
silently, filling with NaN where a wrong pattern in the other two gives a visibly zero or
|
|
50
|
+
all-`False` column.
|
|
51
|
+
- docs/02 §2.5 documents the shipped transform with its mode table and coverage log; §2.4
|
|
52
|
+
keeps `IsWeekend` as the write-your-own example.
|
|
53
|
+
|
|
54
|
+
## 0.2.4
|
|
55
|
+
|
|
56
|
+
### Added
|
|
57
|
+
|
|
58
|
+
- `bernoulli`, `binomial`, `nbinom` and `beta` are default registry families. `pareto` is
|
|
59
|
+
deliberately not among them: lesson 4 has the student register it as the extensibility
|
|
60
|
+
exercise, and the class docstring shows the call.
|
|
61
|
+
- `QQPlot`: sample quantiles against a frozen distribution's theoretical quantiles with a
|
|
62
|
+
`y = x` reference line, for the tail mismatches a `PlotFits` histogram is too
|
|
63
|
+
bulk-dominated to show.
|
|
64
|
+
- `ECDFPlot`: one or two empirical CDFs, bin-free.
|
|
65
|
+
- `fit_table`: turns a `fit()` result list into a dataframe ranked by log-likelihood — one
|
|
66
|
+
row per family with params, log-likelihood, KS statistic and p-value, and which criteria
|
|
67
|
+
it won. `FailedFit` entries carry their message and sort last, since a missing likelihood
|
|
68
|
+
is not a small one.
|
|
69
|
+
|
|
70
|
+
### Fixed
|
|
71
|
+
|
|
72
|
+
- Log-likelihood uses `logpmf` for discrete families and `logpdf` for continuous ones, so
|
|
73
|
+
every discrete family reports a real likelihood and the comparison between them means
|
|
74
|
+
something.
|
|
75
|
+
- The lower bound for a single-parameter discrete family is `1e-3`, low enough for a
|
|
76
|
+
probability (bernoulli's `p`) as well as a count-scale rate (poisson's lambda).
|
|
77
|
+
- `best_likelihood` and `best_ks` are marked independently: a KS p-value that never clears
|
|
78
|
+
the `p > 0` threshold — routine for a discrete family on a large, tied sample — leaves
|
|
79
|
+
the likelihood winner marked. An unknown `criterion` raises `ValueError` whether or not
|
|
80
|
+
a winner was found.
|
|
81
|
+
|
|
82
|
+
## 0.2.3
|
|
83
|
+
|
|
84
|
+
### Added
|
|
85
|
+
|
|
86
|
+
- `TimeFeatures` transform: derives `date`, `hour`, `day_name`, `isoweek` and `year_week`
|
|
87
|
+
from a timestamp column, with a `features=[...]` selector. The default is the three
|
|
88
|
+
lesson 1 relies on, and an unknown feature name raises `ValueError`.
|
|
89
|
+
- `DecomposePlot`: `seasonal_decompose` as the four-panel observed/trend/seasonal/residual
|
|
90
|
+
view, composing a `LinePlot` per panel.
|
|
91
|
+
- `ACFPlot`: autocorrelations as bars with a Bartlett confidence band. NaN input raises,
|
|
92
|
+
since `statsmodels.acf` returns an all-NaN result instead of raising.
|
|
93
|
+
|
|
94
|
+
### Changed
|
|
95
|
+
|
|
96
|
+
- `statsmodels` is a dependency.
|
|
97
|
+
|
|
98
|
+
## 0.2.2
|
|
99
|
+
|
|
100
|
+
### Added
|
|
101
|
+
|
|
102
|
+
- Analysis-method rows on exploratory versus confirmatory questions and on mechanism: which
|
|
103
|
+
of the two a question is decides what would count as evidence, and a significant result
|
|
104
|
+
with no nameable mechanism points at a confounder rather than a discovery.
|
|
105
|
+
- Analysis-method rows on the unit of analysis — the unit a claim is about against the unit
|
|
106
|
+
of a row, and the number of independent units behind each group being compared — with
|
|
107
|
+
§9.3 naming pseudoreplication and showing the n=4000/n=5 bar chart.
|
|
108
|
+
- An analysis-method row on how many comparisons were considered, with §9.6 on the garden
|
|
109
|
+
of forking paths: the arithmetic of fifteen tries at a one-in-twenty threshold, and the
|
|
110
|
+
three defences (write the question down, split the data, report the count).
|
|
111
|
+
- A visual-critique row separating "A differs from B" from "A is the highest", since every
|
|
112
|
+
set of numbers has a maximum and a ranking claim needs the ranking to be stable.
|
|
113
|
+
- A counting step at the head of the distribution method: how many observations, and how
|
|
114
|
+
many independent ones.
|
|
115
|
+
|
|
116
|
+
These land in docs/09-analysis-method.md and in the matching `goad_mcp` checklists.
|
|
117
|
+
|
|
118
|
+
## 0.2.1
|
|
119
|
+
|
|
120
|
+
### Breaking
|
|
121
|
+
|
|
122
|
+
- **`FileConfig` gained `date_column` and `date_index`.** `FileHandler.load` reads csv and
|
|
123
|
+
parquet (reader picked from the suffix), parses `date_column` as dates and, when
|
|
124
|
+
`date_index` is true, moves it into the index. The defaults (`"date"`, `True`) keep the
|
|
125
|
+
covid pipeline working unchanged, but a `FileConfig` subclass that overrode `load` to
|
|
126
|
+
support another shape can now express it as configuration instead. A file whose suffix has
|
|
127
|
+
no reader, or a `date_column` that is not in the frame, raises `ValueError`. Parquet needs
|
|
128
|
+
`pyarrow`, available as the `parquet` extra.
|
|
129
|
+
- **`DistributionRegistry` is a normal class.** Constructing one no longer returns a shared
|
|
130
|
+
process-wide instance, so a registration affects only the registry it was made on. Code
|
|
131
|
+
that registered a family on one registry and relied on a separately constructed
|
|
132
|
+
`DistributionFitter()` picking it up must now pass it: `DistributionFitter(registry)`.
|
|
133
|
+
|
|
134
|
+
### Added
|
|
135
|
+
|
|
136
|
+
- `goad` console script: fits every family in the registry to one column of a csv or parquet
|
|
137
|
+
file and prints the fits ranked by log-likelihood (`goad data/residuals.csv residual`).
|
|
138
|
+
- `DistributionFitter` accepts an optional `DistributionRegistry`.
|
|
139
|
+
- `filehandler.reader_for`, `filehandler.write_frame` and `filehandler.parse_dates` as
|
|
140
|
+
reusable pieces; `FileHandler.save` writes parquet as well as csv.
|
|
141
|
+
- A pytest suite under `tests/`.
|
|
142
|
+
|
|
143
|
+
### Fixed
|
|
144
|
+
|
|
145
|
+
- `BarWithDates.build` and `VerticalDate.build` return `(fig, ax)` like every other `build`.
|
|
146
|
+
- Plots draw on `self.ax` / `self.fig` instead of module-level `plt` calls, so a layer placed
|
|
147
|
+
with `plot_on_axes` lands on the axis it was handed rather than on matplotlib's current one.
|
|
148
|
+
- The `goad` console script points at a module that exists.
|
|
149
|
+
|
|
150
|
+
### Changed
|
|
151
|
+
|
|
152
|
+
- `__version__` is read from the installed package metadata.
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
10:46:23 [hook] session-start: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
2
|
+
10:46:23 [hook] run-consolidation: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
3
|
+
10:46:24 [consolidation] start
|
|
4
|
+
10:46:34 [consolidation] tokens: 9+6254cache→792out ($0.013050)
|
|
5
|
+
10:46:34 [consolidation] done: 1 files consolidated
|
|
6
|
+
10:46:36 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
7
|
+
10:46:49 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
8
|
+
10:46:54 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
9
|
+
10:47:00 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
10
|
+
10:47:04 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
11
|
+
10:47:10 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
12
|
+
10:47:15 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
13
|
+
10:47:22 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
14
|
+
10:47:27 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
15
|
+
10:47:29 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
16
|
+
10:47:39 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
17
|
+
10:47:46 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
18
|
+
10:47:50 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
19
|
+
10:47:59 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
20
|
+
10:47:59 [hook] save-session: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
21
|
+
10:47:59 [extract] session 16b4788a-5507-4216-98ec-06bac13a9aba
|
|
22
|
+
10:47:59 [extract] 17 exchanges (1 human)
|
|
23
|
+
10:47:59 [extract] 1 human msgs < 3, skip
|
|
24
|
+
10:48:45 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
25
|
+
10:48:50 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
26
|
+
10:48:54 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
27
|
+
10:49:08 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
28
|
+
10:49:13 [hook] post-tool: PROJECT_DIR=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical-solomon-b62511 PIPELINE_DIR=/Users/rgrouls/.claude/plugins/cache/claude-plugins-official/remember/0.19.0 PYTHON=python3 REMEMBER_DIR=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
# Recent
|
|
2
|
+
|
|
3
|
+
## 2026-08-10
|
|
4
|
+
Fixed 5 defects in goad_toolkit (feat/docs-and-mcp), shipped v0.2.0. Removed DistributionRegistry singleton, parameterized FileHandler.load for csv/parquet with date col support, fixed BasePlot.build return contracts. Refactored visualizer.py from module-level plt.* to instance methods (self.ax/.fig), added CLI.
|
|
5
|
+
|
|
6
|
+
## Identity Candidates
|
|
7
|
+
- IDENTITY CANDIDATE: Architectural preference for instance methods over module-level state and singleton patterns
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
16b4788a-5507-4216-98ec-06bac13a9aba
|
|
File without changes
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
1786524479
|
|
File without changes
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
70230
|
|
@@ -4,4 +4,4 @@ project_dir=/Users/rgrouls/code/courses/goad_toolkit/.claude/worktrees/quizzical
|
|
|
4
4
|
slug=-Users-rgrouls-code-courses-goad-toolkit--claude-worktrees-quizzical-solomon-b62511
|
|
5
5
|
sessions_dir=/Users/rgrouls/.claude/projects/-Users-rgrouls-code-courses-goad-toolkit--claude-worktrees-quizzical-solomon-b62511
|
|
6
6
|
memory_dir=/Users/rgrouls/code/courses/goad_toolkit/.remember
|
|
7
|
-
session_id=
|
|
7
|
+
session_id=16b4788a-5507-4216-98ec-06bac13a9aba
|
{goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/MCP_SERVER.md
RENAMED
|
@@ -103,13 +103,13 @@ You pin a version with `GOAD_REF` — a git ref (branch, tag, or commit). That o
|
|
|
103
103
|
selects both the script (in the URL) and the docs it fetches (inside the script), so a whole
|
|
104
104
|
cohort runs exactly the same thing and the two cannot drift apart.
|
|
105
105
|
|
|
106
|
-
> **Note:** the docs and this server land in **v0.2.
|
|
107
|
-
> `GOAD_REF`
|
|
106
|
+
> **Note:** the docs and this server land in **v0.2.2**.
|
|
107
|
+
> You can also point `GOAD_REF` to a branch or work from a local clone (see below).
|
|
108
108
|
|
|
109
109
|
### Claude Code
|
|
110
110
|
|
|
111
111
|
```bash
|
|
112
|
-
claude mcp add goad -e GOAD_REF=v0.2.
|
|
112
|
+
claude mcp add goad -e GOAD_REF=v0.2.2 -- \
|
|
113
113
|
sh -c 'uv run --no-project https://raw.githubusercontent.com/raoulg/goad_toolkit/$GOAD_REF/goad_mcp.py'
|
|
114
114
|
```
|
|
115
115
|
|
|
@@ -123,7 +123,7 @@ Add to `.cursor/mcp.json` (per-project) or `~/.cursor/mcp.json` (global):
|
|
|
123
123
|
"goad": {
|
|
124
124
|
"command": "sh",
|
|
125
125
|
"args": ["-c", "uv run --no-project https://raw.githubusercontent.com/raoulg/goad_toolkit/$GOAD_REF/goad_mcp.py"],
|
|
126
|
-
"env": { "GOAD_REF": "v0.2.
|
|
126
|
+
"env": { "GOAD_REF": "v0.2.2" }
|
|
127
127
|
}
|
|
128
128
|
}
|
|
129
129
|
}
|
|
@@ -137,7 +137,7 @@ Add to `~/.codex/config.toml`:
|
|
|
137
137
|
[mcp_servers.goad]
|
|
138
138
|
command = "sh"
|
|
139
139
|
args = ["-c", "uv run --no-project https://raw.githubusercontent.com/raoulg/goad_toolkit/$GOAD_REF/goad_mcp.py"]
|
|
140
|
-
env = { GOAD_REF = "v0.2.
|
|
140
|
+
env = { GOAD_REF = "v0.2.2" }
|
|
141
141
|
```
|
|
142
142
|
|
|
143
143
|
### Claude Desktop
|
|
@@ -151,7 +151,7 @@ app:
|
|
|
151
151
|
"goad": {
|
|
152
152
|
"command": "sh",
|
|
153
153
|
"args": ["-c", "uv run --no-project https://raw.githubusercontent.com/raoulg/goad_toolkit/$GOAD_REF/goad_mcp.py"],
|
|
154
|
-
"env": { "GOAD_REF": "v0.2.
|
|
154
|
+
"env": { "GOAD_REF": "v0.2.2" }
|
|
155
155
|
}
|
|
156
156
|
}
|
|
157
157
|
}
|
|
@@ -187,7 +187,7 @@ add with the new tag:
|
|
|
187
187
|
|
|
188
188
|
```bash
|
|
189
189
|
claude mcp remove goad -s user
|
|
190
|
-
claude mcp add goad -s user -e GOAD_REF=v0.2.
|
|
190
|
+
claude mcp add goad -s user -e GOAD_REF=v0.2.2 -- \
|
|
191
191
|
sh -c 'uv run --no-project https://raw.githubusercontent.com/raoulg/goad_toolkit/$GOAD_REF/goad_mcp.py'
|
|
192
192
|
```
|
|
193
193
|
|
{goad_toolkit-0.2.7/.claude/worktrees/quizzical-solomon-b62511 → goad_toolkit-0.2.9}/README.md
RENAMED
|
@@ -183,6 +183,36 @@ For the [kstest](https://docs.scipy.org/doc/scipy/reference/generated/scipy.stat
|
|
|
183
183
|
The plots are sorted by log-likelihood, which means there is no good fit with a distribution in this case.
|
|
184
184
|

|
|
185
185
|
|
|
186
|
+
### 🎲 Is it real? The shuffle test
|
|
187
|
+
|
|
188
|
+
Shuffle the labels, measure again, and see where the real number lands. No test, no
|
|
189
|
+
assumptions, no table of critical values — just a picture of what your measurement does when
|
|
190
|
+
the label means nothing:
|
|
191
|
+
|
|
192
|
+
```python
|
|
193
|
+
from goad_toolkit.analytics import NullDistribution
|
|
194
|
+
from goad_toolkit.visualizer import NullPlot, PlotSettings
|
|
195
|
+
|
|
196
|
+
def gap(frame):
|
|
197
|
+
means = frame.groupby("is_bot")["length"].mean()
|
|
198
|
+
return means[True] - means[False]
|
|
199
|
+
|
|
200
|
+
result = NullDistribution(gap, n_iter=2000, seed=1).run(data, label="is_bot")
|
|
201
|
+
settings = PlotSettings(
|
|
202
|
+
title="Mean message length: bots minus humans",
|
|
203
|
+
xlabel="difference in mean length (characters)",
|
|
204
|
+
ylabel="density under shuffled labels",
|
|
205
|
+
)
|
|
206
|
+
NullPlot(settings).plot(result=result)
|
|
207
|
+
```
|
|
208
|
+
|
|
209
|
+

|
|
210
|
+
|
|
211
|
+
The grey cloud is the statistic under shuffled labels; the line is what the real data did.
|
|
212
|
+
Writing the statistic is your job — that is the claim. See
|
|
213
|
+
[Models and residuals](docs/06-models-and-residuals.md) §6.7 for what the p-value can and
|
|
214
|
+
cannot carry.
|
|
215
|
+
|
|
186
216
|
### 🧩 Extending with Custom Distributions
|
|
187
217
|
|
|
188
218
|
You can easily register new distributions:
|
|
@@ -58,6 +58,8 @@ input frame is safe, and the steps stay cheap.
|
|
|
58
58
|
| `SelectDataRange` | keep rows in a date range | `start_date`, `end_date` |
|
|
59
59
|
| `RollingAvg` | rolling mean, drops the leading NaNs | `column`, `window`, `rename` |
|
|
60
60
|
| `ZScaler` | standardise to mean 0, std 1 | `column`, `rename` |
|
|
61
|
+
| `TimeFeatures` | derive calendar columns from a timestamp | `column`, `features` |
|
|
62
|
+
| `RegexFeature` | count / flag / extract a pattern in a text column | `column`, `pattern`, `feature`, `mode` |
|
|
61
63
|
|
|
62
64
|
`rename=True` writes to a new column (`deaths_shifted`, `deaths_zscore`, …) instead of
|
|
63
65
|
overwriting. Prefer it. An overwritten column is a step you cannot debug, and the whole
|
|
@@ -75,20 +77,17 @@ Subclass `TransformBase` and implement `transform`. That is the entire contract.
|
|
|
75
77
|
from goad_toolkit.datatransforms import TransformBase
|
|
76
78
|
import pandas as pd
|
|
77
79
|
|
|
78
|
-
class
|
|
79
|
-
"""
|
|
80
|
+
class IsWeekend(TransformBase):
|
|
81
|
+
"""Flag Saturday/Sunday from a timestamp column."""
|
|
80
82
|
|
|
81
83
|
def transform(self, data: pd.DataFrame, column: str) -> pd.DataFrame:
|
|
82
84
|
ts = pd.to_datetime(data[column])
|
|
83
|
-
data["
|
|
84
|
-
data["hour"] = ts.dt.hour
|
|
85
|
-
data["day_name"] = ts.dt.day_name()
|
|
86
|
-
data["isoweek"] = ts.dt.isocalendar().week
|
|
85
|
+
data["is_weekend"] = ts.dt.dayofweek >= 5
|
|
87
86
|
return data
|
|
88
87
|
```
|
|
89
88
|
|
|
90
89
|
```python
|
|
91
|
-
pipeline.add(
|
|
90
|
+
pipeline.add(IsWeekend, column="timestamp")
|
|
92
91
|
```
|
|
93
92
|
|
|
94
93
|
What the base class does for you:
|
|
@@ -106,39 +105,39 @@ Two rules for your `transform`:
|
|
|
106
105
|
- **Name your parameters explicitly** in the signature. `def transform(self, data, column,
|
|
107
106
|
window)` documents itself; `**kwargs` does not, and the validation cannot help you.
|
|
108
107
|
|
|
109
|
-
## 2.5
|
|
108
|
+
## 2.5 `RegexFeature`: feature enrichment from text
|
|
110
109
|
|
|
111
|
-
The single most useful
|
|
110
|
+
The single most useful transform for text data, and the reason `TransformBase` is worth
|
|
111
|
+
subclassing at all. It ships:
|
|
112
112
|
|
|
113
113
|
```python
|
|
114
|
-
|
|
115
|
-
"""Add a feature extracted from a text column with a regular expression."""
|
|
116
|
-
|
|
117
|
-
def transform(
|
|
118
|
-
self,
|
|
119
|
-
data: pd.DataFrame,
|
|
120
|
-
column: str,
|
|
121
|
-
pattern: str,
|
|
122
|
-
feature: str,
|
|
123
|
-
mode: str = "count",
|
|
124
|
-
) -> pd.DataFrame:
|
|
125
|
-
text = data[column].fillna("")
|
|
126
|
-
if mode == "count":
|
|
127
|
-
data[feature] = text.str.count(pattern)
|
|
128
|
-
elif mode == "has":
|
|
129
|
-
data[feature] = text.str.contains(pattern, regex=True)
|
|
130
|
-
elif mode == "extract":
|
|
131
|
-
data[feature] = text.str.extract(pattern, expand=False)
|
|
132
|
-
else:
|
|
133
|
-
raise ValueError(f"mode must be count/has/extract, got {mode!r}")
|
|
134
|
-
return data
|
|
135
|
-
```
|
|
114
|
+
from goad_toolkit.datatransforms import RegexFeature
|
|
136
115
|
|
|
137
|
-
```python
|
|
138
116
|
pipeline.add(RegexFeature, name="url_flag",
|
|
139
117
|
column="message", pattern=r"https?://\S+", feature="has_url", mode="has")
|
|
140
118
|
```
|
|
141
119
|
|
|
120
|
+
Three modes, each writing one new column named by `feature`:
|
|
121
|
+
|
|
122
|
+
| `mode` | writes | use for |
|
|
123
|
+
|---|---|---|
|
|
124
|
+
| `"count"` | how many times the pattern occurs (int) | how many URLs, how many question marks |
|
|
125
|
+
| `"has"` | whether it occurs at all (bool) | flags you will group or filter on |
|
|
126
|
+
| `"extract"` | the first capture group, NaN where nothing matched | pulling a value *out* of the text |
|
|
127
|
+
|
|
128
|
+
`"extract"` needs exactly one capture group in `pattern`, and it is the mode worth being
|
|
129
|
+
careful with. `count` and `has` fail visibly when a pattern is wrong — a column of all zeros
|
|
130
|
+
or all `False` is hard to miss. Extraction fails *silently*, filling with NaN, so it reports
|
|
131
|
+
its own coverage through `loguru`:
|
|
132
|
+
|
|
133
|
+
```
|
|
134
|
+
mentions: extracted 'addressed_to' from 92,415/627,172 rows (14.7%); 534,757 rows had no match
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
That number is the point. A pattern that matches 15% of rows may be exactly right — on IRC,
|
|
138
|
+
most messages do not address anyone — or it may be silently broken. The log line makes you
|
|
139
|
+
decide which, instead of finding out four notebooks later.
|
|
140
|
+
|
|
142
141
|
Note the two different `name`s: `Pipeline.add(name=...)` names the *step*, and
|
|
143
142
|
`TransformBase.__init__` consumes it. So the new-column parameter has to be called something
|
|
144
143
|
else — `feature` here. Any transform that wants to name an output column hits this.
|
|
@@ -200,9 +199,31 @@ will not warn you. Write the steps in the order you would do them by hand.
|
|
|
200
199
|
you would want to run again on new data. A one-off `df[df.author == "Alice"]` while you are
|
|
201
200
|
poking around is a notebook cell, and should stay one.
|
|
202
201
|
|
|
203
|
-
**Debugging.** When a pipeline produces something surprising,
|
|
204
|
-
|
|
205
|
-
|
|
202
|
+
**Debugging.** When a pipeline produces something surprising, do not dismantle it into cells:
|
|
203
|
+
the moment it becomes cells again you have lost the record of what you did, which is the
|
|
204
|
+
thing you were trying to keep. Ask it what it did instead.
|
|
205
|
+
|
|
206
|
+
```python
|
|
207
|
+
result = pipeline.apply(data, keep_intermediate=True)
|
|
208
|
+
pipeline.report()
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
```
|
|
212
|
+
step rows columns added removed
|
|
213
|
+
calendar 4083 9 hour, day_name
|
|
214
|
+
weekend 4083 10 is_weekend
|
|
215
|
+
window 3961 10
|
|
216
|
+
```
|
|
217
|
+
|
|
218
|
+
`report()` gives one row per step: how many rows survived it, how many columns the frame had
|
|
219
|
+
afterwards, and which columns that step added or removed. The step where the row count falls
|
|
220
|
+
off a cliff is the one filtering more than you meant it to, and a step with an empty `added`
|
|
221
|
+
is a step that silently did nothing.
|
|
222
|
+
|
|
223
|
+
`pipeline.intermediate` holds the frame as it stood after each step, keyed by step name, so
|
|
224
|
+
you can look at the rows themselves once `report()` has told you where to look. Both are off
|
|
225
|
+
by default — keeping them costs one full copy of the frame per step — and turning them on
|
|
226
|
+
does not change what `apply` returns.
|
|
206
227
|
|
|
207
228
|
---
|
|
208
229
|
|
|
@@ -148,12 +148,29 @@ step 3.
|
|
|
148
148
|
| `LinePlot` | `sns.lineplot` | the primitive most compositions build on |
|
|
149
149
|
| `ComparePlot` | two lines with labels | `LinePlot` twice via `plot_on` |
|
|
150
150
|
| `VerticalDate` | a labelled vertical line at a date | a layer, not a plot |
|
|
151
|
+
| `VerticalLine` | the same at a number | a layer, not a plot |
|
|
151
152
|
| `ComparePlotDate` | `ComparePlot` + `VerticalDate` | |
|
|
153
|
+
| `BarPlot` | `sns.barplot`, one bar per category | see [Five families §4.1](04-five-families.md) |
|
|
154
|
+
| `GroupedBarPlot` | bars split by `hue` | two categorical variables at once |
|
|
155
|
+
| `HeatmapPlot` | `sns.heatmap`, pivoting the frame if asked | when both variables have many levels |
|
|
156
|
+
| `BarbellPlot` | before and after per category, joined | the change is the message |
|
|
157
|
+
| `HighlightCategory` | recolours bars: grey, plus the one that matters | a layer, and it needs no help from what drew them |
|
|
158
|
+
| `Annotate` | text on the plot, with an arrow to the point | a layer; the finding, written where it is read |
|
|
159
|
+
| `ScatterPlot` | `sns.scatterplot` | the first plot of any relation question |
|
|
160
|
+
| `RegPlot` | `sns.regplot` — `fit_reg`, `lowess`, `order` | `scatter=False` layers it over a `ScatterPlot` |
|
|
161
|
+
| `CorrelationHeatmap` | `.corr()` as a heatmap, scale pinned to [-1, 1] | where to look next, not a ranking |
|
|
152
162
|
| `BarWithDates` | bars with a month locator on the x axis | |
|
|
153
163
|
| `ResidualPlot` | `BarWithDates` + `VerticalDate` | the residual-over-time view |
|
|
164
|
+
| `DecomposePlot` | `seasonal_decompose`, one panel each | observed/trend/seasonal/residual, overrides `plot` |
|
|
165
|
+
| `ACFPlot` | `acf` as bars with a confidence band | see [Five families §4.2](04-five-families.md) |
|
|
154
166
|
| `HistogramPlot` | `sns.histplot`, `stat="density"` | density so a pdf can be overlaid |
|
|
155
167
|
| `DistPlot` | a scipy distribution's pdf or pmf | estimates its own x-range from `ppf` |
|
|
156
168
|
| `PlotFits` | histogram + fitted pdf, one panel per fit | see [Distributions](05-distributions.md) |
|
|
169
|
+
| `QQPlot` | sample quantiles vs. a distribution's theoretical quantiles | the tails, not the bulk — see [Distributions §5.5](05-distributions.md) |
|
|
170
|
+
| `ECDFPlot` | one or two empirical CDFs, bin-free | the right tool for comparing two samples |
|
|
171
|
+
| `NullPlot` | a shuffled statistic, with the observed value marked | `HistogramPlot` + `VerticalLine`, see [Models and residuals §6.7](06-models-and-residuals.md) |
|
|
172
|
+
| `ProjectionPlot` | already-fitted 2-D coordinates | takes an array, not a model — see [Five families §4.5](04-five-families.md) |
|
|
173
|
+
| `ScreePlot` | variance explained per component, plus the running total | the figure a projection has to come with |
|
|
157
174
|
|
|
158
175
|
`HistogramPlot` defaults to `sqrt(n)` bins capped at 50. That default is a starting point,
|
|
159
176
|
not an answer — bin count changes what a histogram appears to say, and choosing it is part of
|
|
@@ -164,6 +181,45 @@ falling back to `pmf` for discrete families. Left to itself it picks an x-range
|
|
|
164
181
|
0.1st to the 99.9th percentile, which is right for most things and wrong for anything with a
|
|
165
182
|
very heavy tail — pass `x_range` explicitly there.
|
|
166
183
|
|
|
184
|
+
`QQPlot` takes a sample and a *frozen* distribution and plots sorted data against the
|
|
185
|
+
distribution's quantiles at matching plotting positions, with a y=x reference line —
|
|
186
|
+
where `PlotFits`' histogram is dominated by the bulk of the data, this is where a tail
|
|
187
|
+
mismatch actually shows up.
|
|
188
|
+
|
|
189
|
+
`ECDFPlot` plots one sample, or two with `compare=`, as step functions of the empirical CDF —
|
|
190
|
+
no bin width to choose, so nothing about the shape is a plotting decision.
|
|
191
|
+
|
|
192
|
+
`HighlightCategory` is the clearest case for why layers beat settings. It recolours the bars
|
|
193
|
+
already on the axis — grey for everything, one colour for the categories you name — so it
|
|
194
|
+
works over `BarPlot`, over a bare `sns.barplot`, over pandas' `.plot.bar`, and over the
|
|
195
|
+
`BasePlot` subclass you wrote yourself. None of them need to know it exists:
|
|
196
|
+
|
|
197
|
+
```python
|
|
198
|
+
settings = PlotSettings(title="Survival rate by deck", highlight=["D"])
|
|
199
|
+
bars = BarPlot(settings)
|
|
200
|
+
bars.plot(data=decks, x="deck", y="survival")
|
|
201
|
+
bars.plot_on(HighlightCategory(settings))
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
`highlight`, `base_color` and `highlight_color` live on `PlotSettings` so that the call stays
|
|
205
|
+
about the data, but they are only defaults — pass `categories=` at the call to override.
|
|
206
|
+
Naming a category that is not on the axis raises, and lists the ones that are.
|
|
207
|
+
|
|
208
|
+
`Annotate` is the other layer worth reaching for by habit. A reader with thirty seconds reads
|
|
209
|
+
the annotation and not the axes, so the sentence you would have said out loud belongs on the
|
|
210
|
+
plot:
|
|
211
|
+
|
|
212
|
+
```python
|
|
213
|
+
line.plot_on(
|
|
214
|
+
Annotate(settings),
|
|
215
|
+
text="six-day gap: the logging host was down, not the channel",
|
|
216
|
+
xy=(gap_date, gap_value), # the point being explained
|
|
217
|
+
xytext=(label_date, label_value), # where the text sits; an arrow joins them
|
|
218
|
+
)
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
Give it `xy` alone and the text lands on the point with no arrow.
|
|
222
|
+
|
|
167
223
|
## 3.7 Writing your own plot
|
|
168
224
|
|
|
169
225
|
Two questions decide the shape:
|
|
@@ -216,9 +272,11 @@ lines of code.
|
|
|
216
272
|
`self.ax.axvline(...)`, `self.fig.suptitle(...)`. Keeping to them is what makes
|
|
217
273
|
`plot_on_axes` reliable.
|
|
218
274
|
|
|
219
|
-
`PlotFits`
|
|
220
|
-
create before the figure exists
|
|
221
|
-
|
|
275
|
+
`PlotFits` and `DecomposePlot` override `plot` rather than `build`, because both need to know
|
|
276
|
+
how many panels to create before the figure exists — variable for `PlotFits` (one per fit),
|
|
277
|
+
fixed at four for `DecomposePlot`. Either way, `build` alone cannot decide the panel count
|
|
278
|
+
before `create_figure` runs, so both leave it as a no-op. This is the exception, not a pattern
|
|
279
|
+
to copy for a plot that only ever needs one axis.
|
|
222
280
|
|
|
223
281
|
---
|
|
224
282
|
|