cfpb-complaints-analysis 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Bruno Novarini
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,66 @@
1
+ Metadata-Version: 2.4
2
+ Name: cfpb-complaints-analysis
3
+ Version: 0.1.0
4
+ Summary: CFPB Consumer Complaint Database as clean Parquet plus an MCP server: 18M complaints and 3.85M archived narratives
5
+ License: MIT
6
+ Requires-Python: >=3.10
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: duckdb>=1.0
10
+ Requires-Dist: mcp<2,>=1.2
11
+ Dynamic: license-file
12
+
13
+ # cfpb-complaints-analysis
14
+
15
+ The CFPB Consumer Complaint Database as clean Parquet, plus an MCP server so an AI assistant can query it: 18.2 million complaints (2011 to today) and 3.85 million consumer narratives.
16
+
17
+ ## What is in it
18
+
19
+ | Layer | Source | Coverage |
20
+ |---|---|---|
21
+ | Structured complaints | CFPB's public complaint file, `files.consumerfinance.gov/ccdb/complaints.csv.zip` | 18,239,728 complaints, received 2011-12-01 to 2026-10-07 (rebuild to refresh) |
22
+ | Narratives | CFPB FOIA Reading Room, "CFPB Consumer Complaint Database Narratives Archive" | 3,851,415 narratives, published through 2026-08-14, none before 2015 |
23
+
24
+ Both are CFPB's own files. Nothing comes from a third party. CFPB describes the narratives as public domain for FOIA purposes, and says on its data page that complaint data is "freely available for anyone to use, analyze, and build on."
25
+
26
+ ## The narrative freeze
27
+
28
+ CFPB stopped publishing complaint narratives in September 2026 (database release 24, and the Aug 14, 2026 announcement that it would cease discretionary publication of narratives and visualizations). It moved the previously published narratives to its FOIA Reading Room. The live complaint file has no narrative column any more.
29
+
30
+ This project joins the archived narratives back to the live complaints by Complaint ID. So:
31
+
32
+ - Narratives exist only for complaints CFPB had published with one through 2026-08-14. Complaints from after that have none, and it will stay that way.
33
+ - Fewer than half of complaints ever had a narrative: the consumer had to consent, and none were published before 2015.
34
+ - Narratives are scrubbed by CFPB (personal data shows as XXXX).
35
+ - Structured fields keep updating when you rebuild from CFPB's live file.
36
+
37
+ ## Use it
38
+
39
+ ```
40
+ pip install cfpb-complaints-analysis
41
+ cfpb-complaints-mcp
42
+ ```
43
+
44
+ The first run downloads the Parquet files (about 700 MB) from the hosted service to `~/.cache/cfpb-complaints-analysis`. Set `CFPB_DATA_URL` to fetch them from somewhere else, or `CFPB_DATA_DIR` to use a folder you built yourself. Add the server to your MCP client as a stdio command, or use the hosted endpoint listed in the MCP registry as `io.github.bnovarini/cfpb-complaints-analysis`.
45
+
46
+ Tools: `dataset_info`, `list_values`, `find_company`, `complaint_counts`, `trend`, `company_profile`, `compare_companies`, `search_narratives`, `count_narratives`, `get_complaint`.
47
+
48
+ Example questions: What do people complain about at Navy Federal versus PenFed? How did mortgage complaints about Rocket change by year? Find 2024 narratives that mention "overdraft fee" at a credit union.
49
+
50
+ Build it yourself from CFPB's sources: `cfpb-complaints --data data build`.
51
+
52
+ ## Read this before quoting numbers
53
+
54
+ - Complaints are unverified consumer allegations. CFPB says so itself.
55
+ - Counts are raw. A large company will have more complaints than a small one; nothing is scaled by customers or accounts.
56
+ - CFPB lists some firms under several names. Use `find_company` and `company_contains` to include all variants.
57
+ - The newest months are partial: complaints are still being sent to companies, and "In progress" is a status, not an outcome.
58
+ - Keyword narrative search scans text. Unfiltered searches take about ten seconds. A date, company or product filter makes them much faster.
59
+
60
+ ## Checks
61
+
62
+ Every count is reconciled to CFPB's own published numbers. See [docs/AUDIT.md](docs/AUDIT.md), including the figures that do not match and why.
63
+
64
+ ## License
65
+
66
+ MIT for the code. The data is CFPB's.
@@ -0,0 +1,54 @@
1
+ # cfpb-complaints-analysis
2
+
3
+ The CFPB Consumer Complaint Database as clean Parquet, plus an MCP server so an AI assistant can query it: 18.2 million complaints (2011 to today) and 3.85 million consumer narratives.
4
+
5
+ ## What is in it
6
+
7
+ | Layer | Source | Coverage |
8
+ |---|---|---|
9
+ | Structured complaints | CFPB's public complaint file, `files.consumerfinance.gov/ccdb/complaints.csv.zip` | 18,239,728 complaints, received 2011-12-01 to 2026-10-07 (rebuild to refresh) |
10
+ | Narratives | CFPB FOIA Reading Room, "CFPB Consumer Complaint Database Narratives Archive" | 3,851,415 narratives, published through 2026-08-14, none before 2015 |
11
+
12
+ Both are CFPB's own files. Nothing comes from a third party. CFPB describes the narratives as public domain for FOIA purposes, and says on its data page that complaint data is "freely available for anyone to use, analyze, and build on."
13
+
14
+ ## The narrative freeze
15
+
16
+ CFPB stopped publishing complaint narratives in September 2026 (database release 24, and the Aug 14, 2026 announcement that it would cease discretionary publication of narratives and visualizations). It moved the previously published narratives to its FOIA Reading Room. The live complaint file has no narrative column any more.
17
+
18
+ This project joins the archived narratives back to the live complaints by Complaint ID. So:
19
+
20
+ - Narratives exist only for complaints CFPB had published with one through 2026-08-14. Complaints from after that have none, and it will stay that way.
21
+ - Fewer than half of complaints ever had a narrative: the consumer had to consent, and none were published before 2015.
22
+ - Narratives are scrubbed by CFPB (personal data shows as XXXX).
23
+ - Structured fields keep updating when you rebuild from CFPB's live file.
24
+
25
+ ## Use it
26
+
27
+ ```
28
+ pip install cfpb-complaints-analysis
29
+ cfpb-complaints-mcp
30
+ ```
31
+
32
+ The first run downloads the Parquet files (about 700 MB) from the hosted service to `~/.cache/cfpb-complaints-analysis`. Set `CFPB_DATA_URL` to fetch them from somewhere else, or `CFPB_DATA_DIR` to use a folder you built yourself. Add the server to your MCP client as a stdio command, or use the hosted endpoint listed in the MCP registry as `io.github.bnovarini/cfpb-complaints-analysis`.
33
+
34
+ Tools: `dataset_info`, `list_values`, `find_company`, `complaint_counts`, `trend`, `company_profile`, `compare_companies`, `search_narratives`, `count_narratives`, `get_complaint`.
35
+
36
+ Example questions: What do people complain about at Navy Federal versus PenFed? How did mortgage complaints about Rocket change by year? Find 2024 narratives that mention "overdraft fee" at a credit union.
37
+
38
+ Build it yourself from CFPB's sources: `cfpb-complaints --data data build`.
39
+
40
+ ## Read this before quoting numbers
41
+
42
+ - Complaints are unverified consumer allegations. CFPB says so itself.
43
+ - Counts are raw. A large company will have more complaints than a small one; nothing is scaled by customers or accounts.
44
+ - CFPB lists some firms under several names. Use `find_company` and `company_contains` to include all variants.
45
+ - The newest months are partial: complaints are still being sent to companies, and "In progress" is a status, not an outcome.
46
+ - Keyword narrative search scans text. Unfiltered searches take about ten seconds. A date, company or product filter makes them much faster.
47
+
48
+ ## Checks
49
+
50
+ Every count is reconciled to CFPB's own published numbers. See [docs/AUDIT.md](docs/AUDIT.md), including the figures that do not match and why.
51
+
52
+ ## License
53
+
54
+ MIT for the code. The data is CFPB's.
@@ -0,0 +1,19 @@
1
+ [project]
2
+ name = "cfpb-complaints-analysis"
3
+ version = "0.1.0"
4
+ description = "CFPB Consumer Complaint Database as clean Parquet plus an MCP server: 18M complaints and 3.85M archived narratives"
5
+ readme = "README.md"
6
+ license = {text = "MIT"}
7
+ requires-python = ">=3.10"
8
+ dependencies = ["duckdb>=1.0", "mcp>=1.2,<2"]
9
+
10
+ [project.scripts]
11
+ cfpb-complaints = "cfpb_complaints.cli:main"
12
+ cfpb-complaints-mcp = "cfpb_complaints.mcp_server:main"
13
+
14
+ [build-system]
15
+ requires = ["setuptools>=61"]
16
+ build-backend = "setuptools.build_meta"
17
+
18
+ [tool.setuptools.packages.find]
19
+ where = ["src"]
@@ -0,0 +1,4 @@
1
+ [egg_info]
2
+ tag_build =
3
+ tag_date = 0
4
+
@@ -0,0 +1 @@
1
+ __version__ = "0.1.0"
@@ -0,0 +1,2 @@
1
+ from .cli import main
2
+ raise SystemExit(main())
@@ -0,0 +1,78 @@
1
+ """Build the Parquet files from CFPB's own downloads (live CSV plus the FOIA narratives archive)."""
2
+ from __future__ import annotations
3
+
4
+ import glob
5
+ import os
6
+ import re
7
+ import subprocess
8
+ import sys
9
+ import urllib.request
10
+ from pathlib import Path
11
+
12
+ import duckdb
13
+
14
+ LIVE_URL = "https://files.consumerfinance.gov/ccdb/complaints.csv.zip"
15
+ ARCHIVE_PAGE = "https://www.consumerfinance.gov/foia-requests/foia-electronic-reading-room/cfpb-consumer-complaint-database-narratives-archive/"
16
+ UA = {"User-Agent": "cfpb-complaints-analysis (open source research project)"}
17
+ _ZIP = re.compile(r'https://files\.consumerfinance\.gov/f/documents/CCDB_Export_(\d+)_[^"\']*?\.zip')
18
+
19
+
20
+ def _get(url: str, dest: Path | None = None) -> bytes | None:
21
+ req = urllib.request.Request(url, headers=UA)
22
+ with urllib.request.urlopen(req, timeout=300) as r:
23
+ if dest is None:
24
+ return r.read()
25
+ with open(dest, "wb") as f:
26
+ while chunk := r.read(1 << 20):
27
+ f.write(chunk)
28
+ return None
29
+
30
+
31
+ def archive_urls() -> list[tuple[int, str]]:
32
+ """Scrape the Reading Room page; file names changed over time, so no URL pattern is assumed."""
33
+ html = _get(ARCHIVE_PAGE).decode("utf-8", "replace")
34
+ found = {int(m.group(1)): m.group(0) for m in _ZIP.finditer(html)}
35
+ return sorted(found.items())
36
+
37
+
38
+ def _con() -> duckdb.DuckDBPyConnection:
39
+ c = duckdb.connect()
40
+ c.execute(f"SET memory_limit='{os.environ.get('CFPB_BUILD_MEMORY', '900MB')}'")
41
+ c.execute("SET threads=1")
42
+ c.execute("SET preserve_insertion_order=false")
43
+ return c
44
+
45
+
46
+ def build(data: Path, keep_raw: bool = False) -> None:
47
+ raw, out = data / "raw", data / "parquet"
48
+ raw.mkdir(parents=True, exist_ok=True)
49
+ out.mkdir(parents=True, exist_ok=True)
50
+ c = _con()
51
+ z = raw / "complaints.csv.zip"
52
+ print("downloading live database", file=sys.stderr)
53
+ _get(LIVE_URL, z)
54
+ subprocess.run(["unzip", "-q", "-o", str(z), "-d", str(raw)], check=True)
55
+ csv = raw / "complaints.csv"
56
+ c.execute(f"""COPY (SELECT cast("Date received" AS DATE) date_received, "Product" product, "Sub-product" sub_product,
57
+ "Issue" issue, "Sub-issue" sub_issue, "Company public response" company_public_response, "Company" company,
58
+ "State" state, "ZIP code" zip_code, "Tags" tags, "Submitted via" submitted_via,
59
+ cast("Date sent to company" AS DATE) date_sent_to_company, "Company response to consumer" company_response,
60
+ "Timely response?" timely_response, cast("Complaint ID" AS BIGINT) complaint_id
61
+ FROM read_csv('{csv}', header=true, all_varchar=true, max_line_size=10000000)
62
+ ORDER BY date_received, complaint_id) TO '{out}/complaints.parquet'
63
+ (FORMAT PARQUET, COMPRESSION ZSTD, ROW_GROUP_SIZE 100000)""")
64
+ csv.unlink()
65
+ for n, url in archive_urls():
66
+ print(f"narratives archive file {n}", file=sys.stderr)
67
+ zp = raw / f"archive_{n:02d}.zip"
68
+ _get(url, zp)
69
+ subprocess.run(["unzip", "-q", "-o", str(zp), "-d", str(raw / f"a{n:02d}")], check=True)
70
+ for f in glob.glob(str(raw / f"a{n:02d}" / "*.csv")):
71
+ c.execute(f"""COPY (SELECT cast("Complaint ID" AS BIGINT) complaint_id, "Consumer complaint narrative" narrative
72
+ FROM read_csv('{f}', header=true, all_varchar=true, max_line_size=10000000)
73
+ WHERE trim(coalesce("Consumer complaint narrative", '')) <> '')
74
+ TO '{out}/narratives_{n:02d}.parquet' (FORMAT PARQUET, COMPRESSION ZSTD, ROW_GROUP_SIZE 20000)""")
75
+ os.remove(f)
76
+ if not keep_raw:
77
+ zp.unlink()
78
+ print("done:", sorted(p.name for p in out.glob("*.parquet")), file=sys.stderr)
@@ -0,0 +1,26 @@
1
+ """Command line: build the Parquet files from CFPB's sources, or check an existing set."""
2
+ from __future__ import annotations
3
+
4
+ import argparse
5
+ from pathlib import Path
6
+
7
+
8
+ def main(argv=None) -> int:
9
+ ap = argparse.ArgumentParser(prog="cfpb-complaints")
10
+ ap.add_argument("--data", type=Path, default=Path("data"))
11
+ ap.add_argument("command", choices=["build", "info"])
12
+ a = ap.parse_args(argv)
13
+ if a.command == "build":
14
+ from .build import build
15
+ build(a.data)
16
+ else:
17
+ import os
18
+ os.environ["CFPB_DATA_DIR"] = str(a.data / "parquet")
19
+ from .mcp_server import dataset_info
20
+ import json
21
+ print(json.dumps(dataset_info(), indent=2, default=str))
22
+ return 0
23
+
24
+
25
+ if __name__ == "__main__":
26
+ raise SystemExit(main())
@@ -0,0 +1,447 @@
1
+ """MCP server over the CFPB Consumer Complaint Database.
2
+
3
+ Structured data: CFPB's public complaint file (2011 to present, refreshed by rebuilding).
4
+ Narratives: CFPB's FOIA Reading Room archive of narratives published through 2026-08-14.
5
+ CFPB stopped publishing narratives in September 2026, so the narrative layer is frozen at that date.
6
+ DuckDB queries the Parquet files directly.
7
+ """
8
+ from __future__ import annotations
9
+
10
+ import os
11
+ import re
12
+ import sys
13
+ import threading
14
+ import time
15
+ import urllib.request
16
+ from datetime import date
17
+ from pathlib import Path
18
+ from typing import Any, Optional
19
+
20
+ import duckdb
21
+ from mcp.server.fastmcp import FastMCP
22
+
23
+ RELEASE = "2026-10-07" # data snapshot; rebuild with `cfpb-complaints build` to refresh
24
+ BASE_URL = os.environ.get("CFPB_DATA_URL", "https://cfpb-complaints-analysis.fly.dev/data/")
25
+ FILES = ["complaints.parquet", *[f"narratives_{i:02d}.parquet" for i in range(1, 22)]]
26
+ MAX_ROWS = 100
27
+ QUERY_TIMEOUT_S = float(os.environ.get("CFPB_QUERY_TIMEOUT", "25"))
28
+ NARRATIVE_FREEZE = "2026-08-14"
29
+
30
+ NOTE = (
31
+ "Data: CFPB Consumer Complaint Database (complaints received 2011-12-01 onward, structured fields as CFPB publishes them) "
32
+ "plus consumer narratives from CFPB's FOIA Reading Room archive. CFPB stopped publishing narratives in September 2026: "
33
+ "narratives exist only for complaints CFPB had published with a narrative through 2026-08-14 (3.85M of them, none before 2015), so a complaint with no "
34
+ "narrative may be newer, may lack consumer consent, or may be pre-2015. Complaints are unverified consumer allegations, and a "
35
+ "company's count is not adjusted for its size. Narratives have personal data masked as XXXX by CFPB."
36
+ )
37
+ mcp = FastMCP("cfpb-complaints-analysis", instructions=NOTE)
38
+
39
+ _con: Optional[duckdb.DuckDBPyConnection] = None
40
+
41
+
42
+ def data_dir() -> Path:
43
+ env = os.environ.get("CFPB_DATA_DIR")
44
+ d = Path(env) if env else Path.home() / ".cache" / "cfpb-complaints-analysis" / RELEASE
45
+ d.mkdir(parents=True, exist_ok=True)
46
+ return d
47
+
48
+
49
+ def ensure_files() -> Path:
50
+ d = data_dir()
51
+ for f in FILES:
52
+ p = d / f
53
+ if not p.exists():
54
+ tmp = p.with_suffix(".part")
55
+ urllib.request.urlretrieve(BASE_URL + f, tmp)
56
+ tmp.rename(p)
57
+ return d
58
+
59
+
60
+ def con() -> duckdb.DuckDBPyConnection:
61
+ global _con
62
+ if _con is None:
63
+ d = ensure_files()
64
+ c = duckdb.connect()
65
+ c.execute(f"SET memory_limit='{os.environ.get('CFPB_MEMORY_LIMIT', '1GB')}'")
66
+ c.execute(f"SET threads={int(os.environ.get('CFPB_THREADS', '2'))}")
67
+ c.execute(f"CREATE VIEW c AS SELECT * FROM read_parquet('{d}/complaints.parquet')")
68
+ c.execute(f"CREATE VIEW n AS SELECT * FROM read_parquet('{d}/narratives_*.parquet')")
69
+ _con = c
70
+ return _con
71
+
72
+
73
+ def run(sql: str, params: list[Any] | None = None) -> list[dict]:
74
+ cur = con().cursor()
75
+ timer = threading.Timer(QUERY_TIMEOUT_S, cur.interrupt)
76
+ timer.start()
77
+ try:
78
+ cur.execute(sql, params or [])
79
+ rows = cur.fetchall()
80
+ except duckdb.InterruptException:
81
+ raise RuntimeError(f"Query took longer than {QUERY_TIMEOUT_S:.0f}s and was stopped. Add a date range, company or product filter.")
82
+ finally:
83
+ timer.cancel()
84
+ cols = [x[0] for x in cur.description]
85
+ return [{k: (round(v, 4) if isinstance(v, float) else (v.isoformat() if isinstance(v, date) else v)) for k, v in zip(cols, r)} for r in rows]
86
+
87
+
88
+ # ---------------- names and filters ----------------
89
+ STATES = set("AL AK AZ AR CA CO CT DE DC FL GA HI ID IL IN IA KS KY LA ME MD MA MI MN MS MO MT NE NV NH NJ NM NY NC ND OH OK OR PA RI SC SD TN TX UT VT VA WA WV WI WY AS GU MP PR VI UM FM MH PW AA AE AP".split())
90
+ GROUPS = {
91
+ "year": "year(c.date_received)", "quarter": "strftime(c.date_received, '%Y') || '-Q' || cast(quarter(c.date_received) AS VARCHAR)",
92
+ "month": "strftime(c.date_received, '%Y-%m')", "product": "c.product", "sub_product": "c.sub_product", "issue": "c.issue",
93
+ "sub_issue": "c.sub_issue", "company": "c.company", "state": "c.state", "company_response": "c.company_response",
94
+ "company_public_response": "c.company_public_response", "submitted_via": "c.submitted_via", "timely_response": "c.timely_response",
95
+ }
96
+ STATUS_HINT = "In progress and Untimely response are CFPB response categories, not outcomes."
97
+
98
+
99
+ def _d(v: Optional[str], name: str, end: bool = False) -> Optional[str]:
100
+ """YYYY-MM-DD passes through. YYYY-MM means the first day, or the last day when end=True (date_to)."""
101
+ if v is None or v == "":
102
+ return None
103
+ if not re.fullmatch(r"\d{4}-\d{2}(-\d{2})?", v):
104
+ raise ValueError(f"{name} must be YYYY-MM-DD or YYYY-MM, got {v!r}")
105
+ try:
106
+ if len(v) == 7:
107
+ y, m = int(v[:4]), int(v[5:])
108
+ if end:
109
+ import calendar
110
+ return f"{v}-{calendar.monthrange(y, m)[1]:02d}"
111
+ date(y, m, 1)
112
+ return v + "-01"
113
+ date.fromisoformat(v)
114
+ except ValueError:
115
+ raise ValueError(f"{name} is not a real date: {v!r}")
116
+ return v
117
+
118
+
119
+ def _filters(company=None, company_contains=None, product=None, sub_product=None, issue=None, state=None,
120
+ date_from=None, date_to=None, company_response=None, timely=None, submitted_via=None, tag=None,
121
+ has_narrative=None, prefix="c."):
122
+ w: list[str] = []
123
+ p: list[Any] = []
124
+ if company:
125
+ w.append(f"{prefix}company = ?"); p.append(company)
126
+ if company_contains:
127
+ w.append(f"{prefix}company ILIKE ?"); p.append(f"%{company_contains}%")
128
+ for col, v in (("product", product), ("sub_product", sub_product), ("issue", issue), ("company_response", company_response),
129
+ ("submitted_via", submitted_via)):
130
+ if v:
131
+ w.append(f"{prefix}{col} ILIKE ?"); p.append(v)
132
+ if state:
133
+ s = state.upper().strip()
134
+ if s not in STATES:
135
+ raise ValueError(f"state must be a two-letter code, got {state!r}")
136
+ w.append(f"{prefix}state = ?"); p.append(s)
137
+ f, t = _d(date_from, "date_from"), _d(date_to, "date_to", end=True)
138
+ if f:
139
+ w.append(f"{prefix}date_received >= CAST(? AS DATE)"); p.append(f)
140
+ if t:
141
+ w.append(f"{prefix}date_received <= CAST(? AS DATE)"); p.append(t)
142
+ if timely is not None:
143
+ w.append(f"{prefix}timely_response = ?"); p.append("Yes" if timely else "No")
144
+ if tag:
145
+ w.append(f"{prefix}tags ILIKE ?"); p.append(f"%{tag}%")
146
+ if has_narrative is True:
147
+ w.append(f"{prefix}complaint_id IN (SELECT complaint_id FROM n)")
148
+ elif has_narrative is False:
149
+ w.append(f"{prefix}complaint_id NOT IN (SELECT complaint_id FROM n)")
150
+ return (" AND ".join(w) or "TRUE"), p
151
+
152
+
153
+ FILTER_DOC = (
154
+ "Filters: company (exact name as in find_company), company_contains (substring), product, sub_product, issue, "
155
+ "company_response (exact CFPB value, case-insensitive), submitted_via, state (2 letters), tag (e.g. 'Servicemember', 'Older American'), "
156
+ "timely (true/false), date_from/date_to (YYYY-MM-DD or YYYY-MM, on date received)."
157
+ )
158
+
159
+ METRICS = (
160
+ "count(*) AS complaints, "
161
+ "round(avg(CASE WHEN c.timely_response='Yes' THEN 1.0 WHEN c.timely_response='No' THEN 0.0 END), 4) AS timely_rate, "
162
+ "round(avg(CASE WHEN c.company_response='Closed with monetary relief' THEN 1.0 ELSE 0.0 END), 4) AS monetary_relief_rate, "
163
+ "round(avg(CASE WHEN c.company_response='Closed with non-monetary relief' THEN 1.0 ELSE 0.0 END), 4) AS nonmonetary_relief_rate, "
164
+ "round(avg(CASE WHEN c.company_response='In progress' THEN 1.0 ELSE 0.0 END), 4) AS in_progress_rate"
165
+ )
166
+
167
+
168
+ def _page(limit: int, offset: int) -> tuple[int, int]:
169
+ return max(1, min(int(limit), MAX_ROWS)), max(0, int(offset))
170
+
171
+
172
+ # ---------------- tools ----------------
173
+ @mcp.tool(description="What this dataset is: row counts, date ranges, narrative coverage, and the narrative freeze. Call first when unsure about coverage.")
174
+ def dataset_info() -> dict:
175
+ r = run("SELECT count(*) AS complaints, min(date_received) AS first_received, max(date_received) AS last_received, "
176
+ "count(DISTINCT company) AS companies, count(DISTINCT product) AS products FROM c")[0]
177
+ nn = run("SELECT count(*) AS narratives, min(complaint_id) AS min_id FROM n")[0]
178
+ r["narratives"] = nn["narratives"]
179
+ r["narratives_note"] = (f"Narratives come from CFPB's FOIA Reading Room archive, frozen at {NARRATIVE_FREEZE}: CFPB stopped publishing "
180
+ "narratives in September 2026. Complaints published after that have none, and none exist before 2015.")
181
+ r["note"] = NOTE
182
+ r["latest_complaint_year_partial"] = True
183
+ return r
184
+
185
+
186
+ @mcp.tool(description="List the values of a field with complaint counts, to find exact product/issue/response names. "
187
+ "field is one of: product, sub_product, issue, sub_issue, company_response, company_public_response, submitted_via, state, timely_response. "
188
+ "Optional search narrows by substring; product narrows issues to that product. " + FILTER_DOC)
189
+ def list_values(field: str, search: Optional[str] = None, product: Optional[str] = None, limit: int = 50) -> list[dict]:
190
+ allowed = ["product", "sub_product", "issue", "sub_issue", "company_response", "company_public_response", "submitted_via", "state", "timely_response"]
191
+ if field not in allowed:
192
+ raise ValueError(f"field must be one of {allowed}. For company names use find_company.")
193
+ where, p = ["c." + field + " IS NOT NULL"], []
194
+ if search:
195
+ where.append(f"c.{field} ILIKE ?"); p.append(f"%{search}%")
196
+ if product:
197
+ where.append("c.product ILIKE ?"); p.append(product)
198
+ limit, _ = _page(limit, 0)
199
+ return run(f"SELECT c.{field} AS value, count(*) AS complaints FROM c WHERE {' AND '.join(where)} GROUP BY 1 ORDER BY 2 DESC LIMIT {limit}", p)
200
+
201
+
202
+ @mcp.tool(description="Find companies by name. CFPB lists the same firm under variants and subsidiaries (e.g. 'EQUIFAX, INC.', 'Equifax Information Services LLC'), "
203
+ "so this returns every matching name with complaint counts and first/last dates; pass the exact names you want to other tools via company, "
204
+ "or use company_contains for all variants at once.")
205
+ def find_company(name: str, limit: int = 25) -> list[dict]:
206
+ limit, _ = _page(limit, 0)
207
+ toks = [t for t in re.split(r"[^A-Za-z0-9]+", name) if t]
208
+ if not toks:
209
+ raise ValueError("name is empty")
210
+ where = " AND ".join("c.company ILIKE ?" for _ in toks)
211
+ return run(f"SELECT c.company, count(*) AS complaints, min(c.date_received) AS first_received, max(c.date_received) AS last_received "
212
+ f"FROM c WHERE {where} GROUP BY 1 ORDER BY 2 DESC LIMIT {limit}", [f"%{t}%" for t in toks])
213
+
214
+
215
+ @mcp.tool(description="Count complaints, optionally grouped. group_by is one or two of: " + ", ".join(GROUPS) + ". Returns complaints, timely_rate, "
216
+ "monetary_relief_rate, nonmonetary_relief_rate, in_progress_rate per group (rates are fractions of that group's complaints). "
217
+ "Counts are raw complaint counts, not adjusted for company size. " + STATUS_HINT + " " + FILTER_DOC)
218
+ def complaint_counts(group_by: Optional[list[str]] = None, company: Optional[str] = None, company_contains: Optional[str] = None,
219
+ product: Optional[str] = None, sub_product: Optional[str] = None, issue: Optional[str] = None, state: Optional[str] = None,
220
+ date_from: Optional[str] = None, date_to: Optional[str] = None, company_response: Optional[str] = None,
221
+ timely: Optional[bool] = None, submitted_via: Optional[str] = None, tag: Optional[str] = None,
222
+ has_narrative: Optional[bool] = None, order_by: str = "complaints", limit: int = 25, offset: int = 0) -> list[dict]:
223
+ gb = group_by or []
224
+ if len(gb) > 2:
225
+ raise ValueError("group_by takes at most two fields")
226
+ for g in gb:
227
+ if g not in GROUPS:
228
+ raise ValueError(f"unknown group_by {g!r}; choose from {list(GROUPS)}")
229
+ limit, offset = _page(limit, offset)
230
+ where, p = _filters(company, company_contains, product, sub_product, issue, state, date_from, date_to, company_response, timely, submitted_via, tag, has_narrative)
231
+ sel = ", ".join(f"{GROUPS[g]} AS {g}" for g in gb)
232
+ grp = f" GROUP BY {', '.join(str(i + 1) for i in range(len(gb)))}" if gb else ""
233
+ ob = order_by if order_by in ("complaints", "timely_rate", "monetary_relief_rate") else "complaints"
234
+ chrono = gb and gb[0] in ("year", "quarter", "month")
235
+ order = f" ORDER BY {'1 ASC' if chrono else ob + ' DESC'}" + (f", {ob} DESC" if chrono and len(gb) > 1 else "")
236
+ rows = run(f"SELECT {sel + ', ' if sel else ''}{METRICS} FROM c WHERE {where}{grp}{order} LIMIT {limit + 1} OFFSET {offset}", p)
237
+ if len(rows) > limit:
238
+ rows = rows[:limit]
239
+ rows.append({"truncated": True, "next_offset": offset + limit, "message": "More groups exist; raise offset or narrow filters."})
240
+ return rows
241
+
242
+
243
+ @mcp.tool(description="Complaints over time (month, quarter or year), for one filter set. Set period to month, quarter or year. "
244
+ "The latest period is partial. " + STATUS_HINT + " " + FILTER_DOC)
245
+ def trend(period: str = "month", company: Optional[str] = None, company_contains: Optional[str] = None, product: Optional[str] = None,
246
+ issue: Optional[str] = None, state: Optional[str] = None, date_from: Optional[str] = None, date_to: Optional[str] = None,
247
+ company_response: Optional[str] = None, tag: Optional[str] = None, submitted_via: Optional[str] = None) -> list[dict]:
248
+ if period not in ("month", "quarter", "year"):
249
+ raise ValueError("period must be month, quarter or year")
250
+ where, p = _filters(company, company_contains, product, None, issue, state, date_from, date_to, company_response, None, submitted_via, tag)
251
+ return run(f"SELECT {GROUPS[period]} AS {period}, {METRICS} FROM c WHERE {where} GROUP BY 1 ORDER BY 1 LIMIT 1000", p)
252
+
253
+
254
+ @mcp.tool(description="Profile of one company (or all name variants via company_contains): totals, first/last complaint, top products, issues, states, "
255
+ "response outcomes, yearly counts. Counts are not adjusted for company size. Optional date_from/date_to/product.")
256
+ def company_profile(company: Optional[str] = None, company_contains: Optional[str] = None, date_from: Optional[str] = None,
257
+ date_to: Optional[str] = None, product: Optional[str] = None) -> dict:
258
+ if not company and not company_contains:
259
+ raise ValueError("pass company (exact) or company_contains")
260
+ where, p = _filters(company, company_contains, product, None, None, None, date_from, date_to, None, None, None, None)
261
+ tot = run(f"SELECT {METRICS}, min(c.date_received) AS first_received, max(c.date_received) AS last_received, count(DISTINCT c.company) AS name_variants FROM c WHERE {where}", p)[0]
262
+ if not tot["complaints"]:
263
+ return {"complaints": 0, "message": "No complaints match. Use find_company to see exact names."}
264
+ def top(col, k=8):
265
+ return run(f"SELECT c.{col} AS value, count(*) AS complaints FROM c WHERE {where} AND c.{col} IS NOT NULL GROUP BY 1 ORDER BY 2 DESC LIMIT {k}", p)
266
+ return {**tot, "names_matched": run(f"SELECT c.company, count(*) AS complaints FROM c WHERE {where} GROUP BY 1 ORDER BY 2 DESC LIMIT 10", p),
267
+ "top_products": top("product"), "top_issues": top("issue"), "top_states": top("state"), "responses": top("company_response", 10),
268
+ "by_year": run(f"SELECT year(c.date_received) AS year, count(*) AS complaints FROM c WHERE {where} GROUP BY 1 ORDER BY 1", p),
269
+ "narratives_available": run(f"SELECT count(*) AS n FROM c WHERE {where} AND c.complaint_id IN (SELECT complaint_id FROM n)", p)[0]["n"]}
270
+
271
+
272
+ @mcp.tool(description="Side-by-side comparison of up to 6 companies on the same filters: complaints, timely_rate, relief rates, top product and top issue each. "
273
+ "Each entry in companies is matched as a substring on the company name (all variants combined). Counts are not adjusted for company size.")
274
+ def compare_companies(companies: list[str], product: Optional[str] = None, issue: Optional[str] = None, state: Optional[str] = None,
275
+ date_from: Optional[str] = None, date_to: Optional[str] = None) -> list[dict]:
276
+ if not companies or len(companies) > 6:
277
+ raise ValueError("pass 1 to 6 company names")
278
+ out = []
279
+ for name in companies:
280
+ where, p = _filters(None, name, product, None, issue, state, date_from, date_to, None, None, None, None)
281
+ r = run(f"SELECT {METRICS}, count(DISTINCT c.company) AS name_variants FROM c WHERE {where}", p)[0]
282
+ r = {"company_search": name, **r}
283
+ if r["complaints"]:
284
+ r["top_product"] = run(f"SELECT c.product AS v FROM c WHERE {where} GROUP BY 1 ORDER BY count(*) DESC LIMIT 1", p)[0]["v"]
285
+ r["top_issue"] = run(f"SELECT c.issue AS v FROM c WHERE {where} GROUP BY 1 ORDER BY count(*) DESC LIMIT 1", p)[0]["v"]
286
+ out.append(r)
287
+ return out
288
+
289
+
290
+ def _snippet(text: str, terms: list[str], width: int = 450) -> str:
291
+ low = text.lower()
292
+ pos = min([i for i in (low.find(t.lower()) for t in terms) if i >= 0] or [0])
293
+ s = max(0, pos - width // 3)
294
+ seg = text[s:s + width].strip()
295
+ return ("..." if s else "") + seg + ("..." if s + width < len(text) else "")
296
+
297
+
298
+ @mcp.tool(description="Keyword search over archived complaint narratives (frozen at " + NARRATIVE_FREEZE + "; CFPB stopped publishing them in September 2026). "
299
+ "query: words that must all appear (case-insensitive), or use phrase for an exact phrase, or any_of for alternatives. "
300
+ "Returns a snippet per complaint plus its structured fields; use get_complaint for the full text. A date, company or product filter makes it much faster. "
301
+ + FILTER_DOC)
302
+ def search_narratives(query: Optional[str] = None, phrase: Optional[str] = None, any_of: Optional[list[str]] = None,
303
+ company: Optional[str] = None, company_contains: Optional[str] = None, product: Optional[str] = None,
304
+ sub_product: Optional[str] = None, issue: Optional[str] = None, state: Optional[str] = None,
305
+ date_from: Optional[str] = None, date_to: Optional[str] = None, company_response: Optional[str] = None,
306
+ timely: Optional[bool] = None, tag: Optional[str] = None, limit: int = 10, offset: int = 0) -> list[dict]:
307
+ words = [w for w in re.split(r"\s+", (query or "").strip()) if w]
308
+ if not (words or phrase or any_of):
309
+ raise ValueError("pass query, phrase or any_of")
310
+ limit, offset = _page(limit, offset)
311
+ limit = min(limit, 25)
312
+ where, p = _filters(company, company_contains, product, sub_product, issue, state, date_from, date_to, company_response, timely, None, tag)
313
+ conds, tp = [], []
314
+ for w in words:
315
+ conds.append("contains(lower(n.narrative), ?)"); tp.append(w.lower())
316
+ if phrase:
317
+ conds.append("contains(lower(n.narrative), ?)"); tp.append(phrase.lower())
318
+ if any_of:
319
+ conds.append("(" + " OR ".join("contains(lower(n.narrative), ?)" for _ in any_of) + ")"); tp += [a.lower() for a in any_of]
320
+ sql = (f"SELECT c.complaint_id, c.date_received, c.company, c.product, c.issue, c.state, c.company_response, n.narrative "
321
+ f"FROM c JOIN n USING (complaint_id) WHERE {where} AND {' AND '.join(conds)} ORDER BY c.date_received DESC, c.complaint_id DESC "
322
+ f"LIMIT {limit + 1} OFFSET {offset}")
323
+ rows = run(sql, p + tp)
324
+ terms = words + ([phrase] if phrase else []) + (any_of or [])
325
+ more = len(rows) > limit
326
+ out = []
327
+ for r in rows[:limit]:
328
+ t = r.pop("narrative")
329
+ r["snippet"] = _snippet(t, terms)
330
+ out.append(r)
331
+ if more:
332
+ out.append({"truncated": True, "next_offset": offset + limit, "message": "More matches exist; raise offset, or narrow filters. Results are newest first."})
333
+ if not out:
334
+ out.append({"message": f"No narratives matched. Narratives exist only for complaints published by CFPB with a narrative through {NARRATIVE_FREEZE} (none for complaints received before 2015)."})
335
+ return out
336
+
337
+
338
+ @mcp.tool(description="Count how many archived narratives match a keyword search, optionally grouped by year, product, company, issue or state. "
339
+ "Same search arguments as search_narratives. Use this for 'how many complaints mention X'. Counts only complaints that have a narrative (through "
340
+ + NARRATIVE_FREEZE + "), not all complaints. " + FILTER_DOC)
341
+ def count_narratives(query: Optional[str] = None, phrase: Optional[str] = None, any_of: Optional[list[str]] = None, group_by: Optional[str] = None,
342
+ company: Optional[str] = None, company_contains: Optional[str] = None, product: Optional[str] = None,
343
+ issue: Optional[str] = None, state: Optional[str] = None, date_from: Optional[str] = None, date_to: Optional[str] = None,
344
+ limit: int = 25) -> list[dict]:
345
+ if group_by and group_by not in ("year", "product", "company", "issue", "state", "month"):
346
+ raise ValueError("group_by must be year, month, product, company, issue or state")
347
+ words = [w for w in re.split(r"\s+", (query or "").strip()) if w]
348
+ where, p = _filters(company, company_contains, product, None, issue, state, date_from, date_to, None, None, None, None)
349
+ conds, tp = [], []
350
+ for w in words:
351
+ conds.append("contains(lower(n.narrative), ?)"); tp.append(w.lower())
352
+ if phrase:
353
+ conds.append("contains(lower(n.narrative), ?)"); tp.append(phrase.lower())
354
+ if any_of:
355
+ conds.append("(" + " OR ".join("contains(lower(n.narrative), ?)" for _ in any_of) + ")"); tp += [a.lower() for a in any_of]
356
+ cond = (" AND " + " AND ".join(conds)) if conds else ""
357
+ limit, _ = _page(limit, 0)
358
+ sel = f"{GROUPS[group_by]} AS {group_by}, " if group_by else ""
359
+ grp = " GROUP BY 1 ORDER BY " + ("1" if group_by in ("year", "month") else "2 DESC") if group_by else ""
360
+ return run(f"SELECT {sel}count(*) AS narratives FROM c JOIN n USING (complaint_id) WHERE {where}{cond}{grp} LIMIT {limit}", p + tp)
361
+
362
+
363
+ @mcp.tool(description="One complaint by its CFPB Complaint ID: all structured fields and, when archived, the full narrative.")
364
+ def get_complaint(complaint_id: int) -> dict:
365
+ rows = run("SELECT * FROM c WHERE complaint_id = ?", [int(complaint_id)])
366
+ nr = run("SELECT narrative FROM n WHERE complaint_id = ?", [int(complaint_id)])
367
+ if not rows and not nr:
368
+ return {"error": f"Complaint {complaint_id} is not in the dataset. It may be newer than the data, or CFPB may have removed it."}
369
+ r = rows[0] if rows else {"complaint_id": int(complaint_id), "note": "Narrative is archived but the complaint is no longer in the live CFPB file."}
370
+ r["narrative"] = nr[0]["narrative"] if nr else None
371
+ if not nr:
372
+ r["narrative_note"] = (f"No archived narrative. Narratives exist only for complaints published by CFPB with a narrative through {NARRATIVE_FREEZE} (none for complaints received before 2015) where the consumer consented to publish.")
373
+ return r
374
+
375
+
376
+ # ---------------- hosting ----------------
377
+ class RateLimit:
378
+ def __init__(self, app, per_minute: int):
379
+ self.app, self.per_minute, self.hits = app, per_minute, {}
380
+
381
+ async def __call__(self, scope, receive, send):
382
+ if scope["type"] == "http" and scope["path"] != "/healthz":
383
+ h = dict(scope["headers"])
384
+ ip = (h.get(b"fly-client-ip") or h.get(b"x-forwarded-for", b"").split(b",")[0] or b"?").decode().strip()
385
+ now = time.time()
386
+ q_ = [t for t in self.hits.get(ip, []) if now - t < 60]
387
+ if len(q_) >= self.per_minute:
388
+ await send({"type": "http.response.start", "status": 429, "headers": [(b"content-type", b"application/json"), (b"retry-after", b"60")]})
389
+ await send({"type": "http.response.body", "body": b'{"error":"rate limit exceeded, try again in a minute"}'})
390
+ return
391
+ q_.append(now)
392
+ self.hits[ip] = q_
393
+ if len(self.hits) > 5000:
394
+ self.hits = {k: v for k, v in self.hits.items() if v and now - v[-1] < 60}
395
+ await self.app(scope, receive, send)
396
+
397
+
398
+ def _forbid_extra_arguments() -> None:
399
+ for t in mcp._tool_manager.list_tools():
400
+ model = t.fn_metadata.arg_model
401
+ model.model_config["extra"] = "forbid"
402
+ model.model_rebuild(force=True)
403
+
404
+
405
+ _forbid_extra_arguments()
406
+
407
+
408
+ def http_app():
409
+ from mcp.server.transport_security import TransportSecuritySettings
410
+ from starlette.responses import JSONResponse
411
+ hosts = [h for h in os.environ.get("CFPB_ALLOWED_HOSTS", "").split(",") if h]
412
+ mcp.settings.stateless_http = True
413
+ mcp.settings.json_response = True
414
+ mcp.settings.transport_security = TransportSecuritySettings(
415
+ enable_dns_rebinding_protection=bool(hosts), allowed_hosts=hosts, allowed_origins=["*"] if hosts else [])
416
+
417
+ @mcp.custom_route("/data/{name}", methods=["GET"])
418
+ async def data_file(request):
419
+ from starlette.responses import FileResponse
420
+ name = request.path_params["name"]
421
+ p = data_dir() / name
422
+ if name not in FILES or not p.exists():
423
+ return JSONResponse({"error": "not found"}, status_code=404)
424
+ return FileResponse(p, media_type="application/octet-stream", filename=name)
425
+
426
+ @mcp.custom_route("/healthz", methods=["GET"])
427
+ async def healthz(request):
428
+ return JSONResponse({"ok": True})
429
+
430
+ return RateLimit(mcp.streamable_http_app(), int(os.environ.get("CFPB_RATE_PER_MIN", "60")))
431
+
432
+
433
+ def main() -> None:
434
+ argv = sys.argv[1:]
435
+ if "--check" in argv:
436
+ ensure_files()
437
+ print(dataset_info())
438
+ return
439
+ if "--http" in argv:
440
+ import uvicorn
441
+ uvicorn.run(http_app(), host=os.environ.get("HOST", "0.0.0.0"), port=int(os.environ.get("PORT", "8080")), log_level="warning", timeout_keep_alive=5)
442
+ return
443
+ mcp.run()
444
+
445
+
446
+ if __name__ == "__main__":
447
+ main()
@@ -0,0 +1,66 @@
1
+ Metadata-Version: 2.4
2
+ Name: cfpb-complaints-analysis
3
+ Version: 0.1.0
4
+ Summary: CFPB Consumer Complaint Database as clean Parquet plus an MCP server: 18M complaints and 3.85M archived narratives
5
+ License: MIT
6
+ Requires-Python: >=3.10
7
+ Description-Content-Type: text/markdown
8
+ License-File: LICENSE
9
+ Requires-Dist: duckdb>=1.0
10
+ Requires-Dist: mcp<2,>=1.2
11
+ Dynamic: license-file
12
+
13
+ # cfpb-complaints-analysis
14
+
15
+ The CFPB Consumer Complaint Database as clean Parquet, plus an MCP server so an AI assistant can query it: 18.2 million complaints (2011 to today) and 3.85 million consumer narratives.
16
+
17
+ ## What is in it
18
+
19
+ | Layer | Source | Coverage |
20
+ |---|---|---|
21
+ | Structured complaints | CFPB's public complaint file, `files.consumerfinance.gov/ccdb/complaints.csv.zip` | 18,239,728 complaints, received 2011-12-01 to 2026-10-07 (rebuild to refresh) |
22
+ | Narratives | CFPB FOIA Reading Room, "CFPB Consumer Complaint Database Narratives Archive" | 3,851,415 narratives, published through 2026-08-14, none before 2015 |
23
+
24
+ Both are CFPB's own files. Nothing comes from a third party. CFPB describes the narratives as public domain for FOIA purposes, and says on its data page that complaint data is "freely available for anyone to use, analyze, and build on."
25
+
26
+ ## The narrative freeze
27
+
28
+ CFPB stopped publishing complaint narratives in September 2026 (database release 24, and the Aug 14, 2026 announcement that it would cease discretionary publication of narratives and visualizations). It moved the previously published narratives to its FOIA Reading Room. The live complaint file has no narrative column any more.
29
+
30
+ This project joins the archived narratives back to the live complaints by Complaint ID. So:
31
+
32
+ - Narratives exist only for complaints CFPB had published with one through 2026-08-14. Complaints from after that have none, and it will stay that way.
33
+ - Fewer than half of complaints ever had a narrative: the consumer had to consent, and none were published before 2015.
34
+ - Narratives are scrubbed by CFPB (personal data shows as XXXX).
35
+ - Structured fields keep updating when you rebuild from CFPB's live file.
36
+
37
+ ## Use it
38
+
39
+ ```
40
+ pip install cfpb-complaints-analysis
41
+ cfpb-complaints-mcp
42
+ ```
43
+
44
+ The first run downloads the Parquet files (about 700 MB) from the hosted service to `~/.cache/cfpb-complaints-analysis`. Set `CFPB_DATA_URL` to fetch them from somewhere else, or `CFPB_DATA_DIR` to use a folder you built yourself. Add the server to your MCP client as a stdio command, or use the hosted endpoint listed in the MCP registry as `io.github.bnovarini/cfpb-complaints-analysis`.
45
+
46
+ Tools: `dataset_info`, `list_values`, `find_company`, `complaint_counts`, `trend`, `company_profile`, `compare_companies`, `search_narratives`, `count_narratives`, `get_complaint`.
47
+
48
+ Example questions: What do people complain about at Navy Federal versus PenFed? How did mortgage complaints about Rocket change by year? Find 2024 narratives that mention "overdraft fee" at a credit union.
49
+
50
+ Build it yourself from CFPB's sources: `cfpb-complaints --data data build`.
51
+
52
+ ## Read this before quoting numbers
53
+
54
+ - Complaints are unverified consumer allegations. CFPB says so itself.
55
+ - Counts are raw. A large company will have more complaints than a small one; nothing is scaled by customers or accounts.
56
+ - CFPB lists some firms under several names. Use `find_company` and `company_contains` to include all variants.
57
+ - The newest months are partial: complaints are still being sent to companies, and "In progress" is a status, not an outcome.
58
+ - Keyword narrative search scans text. Unfiltered searches take about ten seconds. A date, company or product filter makes them much faster.
59
+
60
+ ## Checks
61
+
62
+ Every count is reconciled to CFPB's own published numbers. See [docs/AUDIT.md](docs/AUDIT.md), including the figures that do not match and why.
63
+
64
+ ## License
65
+
66
+ MIT for the code. The data is CFPB's.
@@ -0,0 +1,15 @@
1
+ LICENSE
2
+ README.md
3
+ pyproject.toml
4
+ src/cfpb_complaints/__init__.py
5
+ src/cfpb_complaints/__main__.py
6
+ src/cfpb_complaints/build.py
7
+ src/cfpb_complaints/cli.py
8
+ src/cfpb_complaints/mcp_server.py
9
+ src/cfpb_complaints_analysis.egg-info/PKG-INFO
10
+ src/cfpb_complaints_analysis.egg-info/SOURCES.txt
11
+ src/cfpb_complaints_analysis.egg-info/dependency_links.txt
12
+ src/cfpb_complaints_analysis.egg-info/entry_points.txt
13
+ src/cfpb_complaints_analysis.egg-info/requires.txt
14
+ src/cfpb_complaints_analysis.egg-info/top_level.txt
15
+ tests/test_server.py
@@ -0,0 +1,3 @@
1
+ [console_scripts]
2
+ cfpb-complaints = cfpb_complaints.cli:main
3
+ cfpb-complaints-mcp = cfpb_complaints.mcp_server:main
@@ -0,0 +1,29 @@
1
+ import os, pathlib, pytest
2
+
3
+ d = os.environ.get("CFPB_DATA_DIR")
4
+ pytestmark = pytest.mark.skipif(not d or not (pathlib.Path(d) / "complaints.parquet").exists(), reason="needs CFPB_DATA_DIR with the Parquet files")
5
+
6
+
7
+ def test_filters_validate():
8
+ from cfpb_complaints.mcp_server import _filters
9
+ with pytest.raises(ValueError):
10
+ _filters(state="Texas")
11
+ with pytest.raises(ValueError):
12
+ _filters(date_from="01/02/2024")
13
+ w, p = _filters(company_contains="navy", date_from="2024-01")
14
+ assert "company ILIKE" in w and p[-1] == "2024-01-01"
15
+
16
+
17
+ def test_counts_and_search():
18
+ from cfpb_complaints.mcp_server import complaint_counts, search_narratives, get_complaint
19
+ r = complaint_counts(group_by=["year"], date_from="2024-01", date_to="2024-12")
20
+ assert r[0]["year"] == 2024 and r[0]["complaints"] == 2734268
21
+ s = search_narratives(phrase="overdraft fee", date_from="2024-01", date_to="2024-01", limit=2)
22
+ assert s and "snippet" in s[0]
23
+ assert get_complaint(1)["error"]
24
+
25
+
26
+ def test_unknown_group_by_rejected():
27
+ from cfpb_complaints.mcp_server import complaint_counts
28
+ with pytest.raises(ValueError):
29
+ complaint_counts(group_by=["nope"])