microdf-python 1.5.3__py3-none-any.whl → 1.5.5__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: microdf-python
3
- Version: 1.5.3
3
+ Version: 1.5.5
4
4
  Summary: Weighted pandas DataFrames and Series for survey microdata
5
5
  Author-email: Max Ghenis <max@policyengine.org>
6
6
  License: MIT
@@ -25,8 +25,28 @@ Dynamic: license-file
25
25
  # microdf
26
26
  Weighted pandas DataFrames and Series for survey microdata analysis.
27
27
 
28
- ## Overview
29
- microdf provides `MicroDataFrame` and `MicroSeries` classes that extend pandas functionality with integrated weighting support, essential for accurate survey data analysis.
28
+ ## Why this exists
29
+
30
+ Survey microdata comes with weights, and analysing it in pandas means getting two
31
+ things right that pandas will not do for you.
32
+
33
+ **The estimators are not the obvious ones.** A weighted median is not the median of
34
+ weighted values. Weighted variance requires choosing between treating weights as
35
+ frequencies or as precision. A top-1% share requires deciding what happens to a
36
+ record that straddles the cutoff. Each of these is a decision, and hand-rolling it
37
+ per analysis means making it differently each time. `microdf` makes each choice
38
+ once, documents it, and tests it — quantiles follow the inverse CDF so they can be
39
+ checked against R's `survey::svyquantile`, and variance treats weights as
40
+ frequencies so integer weights agree with `numpy` on the replicated sample.
41
+
42
+ **Weights have to survive the pipeline.** Before any estimator runs, weights must
43
+ stay aligned with their rows through merges, filters, grouping and reindexing. When
44
+ they do not, nothing raises. The pipeline completes and returns a plausible wrong
45
+ number. This is the harder of the two problems, and it is why `microdf` carries
46
+ weights inside the object rather than beside it.
47
+
48
+ If you are computing a poverty rate or a Gini on weighted survey data, those are
49
+ the two ways to get a believable-looking wrong answer.
30
50
 
31
51
  ## Key Features
32
52
  - **MicroDataFrame**: A pandas DataFrame with an integrated weight column
@@ -18,8 +18,8 @@ microdf/tests/test_ufunc_weight_dispatch.py,sha256=abQ5I4wFRM3poxO76nbR_1dQymHka
18
18
  microdf/tests/test_version_metadata.py,sha256=M1EabzHLKZZw3Djd6Zu2UuMQtDLV6rZ1zDrOU7W_jf0,227
19
19
  microdf/tests/test_weight_propagation.py,sha256=3odbufFZ2o1rnyjx6PJ7kn1TPO5K2ho3RErczI7mqO4,18761
20
20
  microdf/tests/test_weighted_cov_corr.py,sha256=LTnFhMWnb28f7L_9kOhPiV5UlLMAUbC9OlLIaZ1XgLs,13954
21
- microdf_python-1.5.3.dist-info/licenses/LICENSE,sha256=uPs-ASYnzlldpf2z8jeRgQFeEH3FLhSuX0rw0OKWoDU,1067
22
- microdf_python-1.5.3.dist-info/METADATA,sha256=Z-rZGyWz1EkjDrX7japyIO7fW4HnZNaGMWPy5xQT-M0,2305
23
- microdf_python-1.5.3.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
24
- microdf_python-1.5.3.dist-info/top_level.txt,sha256=T2WFPTygQQMdS3GF8YpZ12DKfMGrspbZ3r7z-e3KfiM,8
25
- microdf_python-1.5.3.dist-info/RECORD,,
21
+ microdf_python-1.5.5.dist-info/licenses/LICENSE,sha256=uPs-ASYnzlldpf2z8jeRgQFeEH3FLhSuX0rw0OKWoDU,1067
22
+ microdf_python-1.5.5.dist-info/METADATA,sha256=lB19BrXfqb_6vAEgjoO4EdfXIulI1rb60-9DOPW_Aao,3427
23
+ microdf_python-1.5.5.dist-info/WHEEL,sha256=YVMoNqKzERt-wjUZwJ33xBGAwnFl-4cqbYkTtWa4itE,91
24
+ microdf_python-1.5.5.dist-info/top_level.txt,sha256=T2WFPTygQQMdS3GF8YpZ12DKfMGrspbZ3r7z-e3KfiM,8
25
+ microdf_python-1.5.5.dist-info/RECORD,,