ai-data-summarizer 0.7.1__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,200 @@
1
+ Metadata-Version: 2.5
2
+ Name: ai-data-summarizer
3
+ Version: 0.7.1
4
+ Summary: A local CLI tool that uses pandas and AI to summarize and interpret datasets.
5
+ Project-URL: Homepage, https://github.com/Bloomy52/ai-data-summarizer
6
+ Author: Louie Bloomberg
7
+ License-Expression: MIT
8
+ License-File: LICENSE
9
+ Requires-Python: >=3.12
10
+ Requires-Dist: anthropic
11
+ Requires-Dist: google-genai
12
+ Requires-Dist: openai
13
+ Requires-Dist: openpyxl
14
+ Requires-Dist: pandas
15
+ Requires-Dist: pandas-stubs>=3.0.5.260730
16
+ Requires-Dist: protobuf
17
+ Requires-Dist: pwinput>=1.0.3
18
+ Requires-Dist: readchar>=4.2.2
19
+ Requires-Dist: sentencepiece
20
+ Description-Content-Type: text/markdown
21
+
22
+ # AI Data Summarization Tool
23
+ A Python Tool that takes a dataset and uses pandas to generate a statistical summary and Artifical Intelligence (AI) to interpret the pandas summary and presents it to the user.
24
+
25
+ This project is a continuation of my CS178 Final Project. You can find the repo at [Bloomy52/cs178-data-summarizer-py](https://www.github.com/Bloomy52/cs178-data-summarizer-py)
26
+
27
+ ### Why This Exists
28
+ I created this project because I found that it is difficult to understand what an underlying dataset is and what it entails without reading and understanding the full dataset. I found that using a Large Language Model (LLM) to summarize the dataset made the dataset more approachable since I had a general understanding of what the dataset was and some features about said dataset I was analyzing. I specifically crafted the summary templates so they would help the user understand the dataset and its features. It can also give you a heads up if there are any concerns or anomalies before you start fully analyzing the data to prevent issues and to guide the user on the right path to analysis.
29
+
30
+ ## Repo Structure
31
+ The structure of this git repository is as follows:
32
+ ```text
33
+ ai-data-summarizer/
34
+ ├── main.py # CLI entry point and main application logic
35
+ ├── summarizer.py # Core data summarization functionality
36
+ ├── prompt.py # Prompt templates and management
37
+ ├── tokenizer.py # Token counting and management utilities
38
+ ├── .env.sample # Sample credentials file (rename to .env)
39
+ ├── apicheck.py # Checks API variables to prevent early issues
40
+ ├── envvar.py # Loads environment variables from .env file
41
+ ├── fileloader.py # Loads and parses CSV and Excel files
42
+ ├── profiler.py # Profiles the dataset for pandas summarization
43
+ ├── requirements.txt # Python dependencies
44
+ ├── README.md # Project documentation
45
+ ├── INSTALL.md # Project installation documentation
46
+ ├── CONFIGURATION.md # Project configuration documentation
47
+ ├── LICENSE # MIT License
48
+ ├── .gitignore # Git ignore rules
49
+ ├── pyproject.toml # Python project application configuration files
50
+ ├── uv.lock # Python/uv dependency lock file
51
+ ├── Dockerfile # Dockerfile for containerized deployment
52
+ ├── .dockerignore # Docker ignore rules
53
+ ├── .vscode/
54
+ │ └── settings.json # VSCode configuration
55
+ ├── summaries/ # Summary output folder -- gets created upon first output summary
56
+ ├── .devcontainer/
57
+ │ └── devcontainer.json # Dev container configuration for VSCode
58
+ └── examples/
59
+ ├── CTA_Ridership_RedLine_WrigleyField_DailyTotals.csv # Sample dataset
60
+ └── OverviewPrompt_CTA_Ridership_RedLine_WrigleyField_DailyTotals.txt # Sample summary
61
+ ```
62
+
63
+ ## Example
64
+ The example used here is the number of daily riders from the Addison 'L' Stop on the Chicago Transit Authority's (CTA) Red Line using the Data Overview prompt. More information about the original dataset can be found at the bottom of the README. Other prompts will be added to the example folder as well.
65
+
66
+ **CSV File Structure**
67
+ ```csv
68
+ "date","daytype","rides"
69
+ "01/01/2001","U","1,227"
70
+ "01/02/2001","W","3,937"
71
+ "01/03/2001","W","4,329"
72
+ "01/04/2001","W","4,607"
73
+ "01/05/2001","W","4,666"
74
+ ```
75
+ * `date` is when the data was recorded
76
+ * `daytype` is the type of day of which the data was recorded
77
+ - `W` is a weekday
78
+ - `A` is a Saturday
79
+ - `U` is Sunday/Holidays
80
+ * `rides` is the number of riders recorded on a given day
81
+
82
+ ```bash
83
+ sumdata
84
+ ```
85
+ ```text
86
+ Enter the path to your CSV file: examples/CTA_Ridership_RedLine_WrigleyField_DailyTotals.csv
87
+
88
+ Please Select a Summary:
89
+ 1. Just the Facts (AI is not used)
90
+ 2. TL;DR Summary
91
+ 3. Data Overview
92
+ 4. Deep Dive Analysis
93
+ Or enter 0 to Exit
94
+ Enter the number corresponding to your choice: 3
95
+ Your input text has 195174 tokens, which is within the free tier limit for Gemini.
96
+ Would you like to continue? (Y/n)
97
+ Y
98
+
99
+ Summary:
100
+
101
+ Overview:
102
+
103
+ I reviewed the dataset you shared. It contains daily transit ride tracking data spanning over 25 years, from
104
+ January 1, 2001, to March 31, 2026. This extensive daily record captures long-term ridership trends and reveals
105
+ how public transit usage fluctuates across weekdays, weekends, and holidays, alongside the long-term impact
106
+ of major external disruptions.
107
+
108
+ ```
109
+ The rest of the example can be found in the `examples` folder labeled [OverviewPrompt_CTA_Ridership_RedLine_WrigleyField_DailyTotals.txt](examples/OverviewPrompt_CTA_Ridership_RedLine_WrigleyField_DailyTotals.txt)
110
+
111
+ Summaries save to the `summaries` folder. It will be created automatically with the subfolder with the prompt and a header in the file.
112
+
113
+ > [!TIP]
114
+ > The capital `Y` means that it is the default option. You can click the `Enter`/`Return` key as a shortcut.
115
+
116
+
117
+ ## Project Requirements
118
+ > [!NOTE]
119
+ > This only supports Gemini at the moment. Support for the Anthropic and OpenAI APIs are coming soon.
120
+
121
+ - Python
122
+ - uv
123
+ - Git
124
+ - Gemini API Key from Google AI Studio
125
+
126
+ See [INSTALL.md](INSTALL.md) for detailed setup instructions, including how to obtain API keys.
127
+
128
+
129
+ ## How to Use/Installation
130
+ There are three options for running this program. Please choose one of these options.
131
+
132
+ ### Option A: Install as a package (recommended)
133
+ ```bash
134
+ uv tool install git+https://www.github.com/Bloomy52/ai-data-summarizer.git
135
+ ```
136
+ Then run:
137
+ ```bash
138
+ sumdata
139
+ ```
140
+
141
+ ### Option B: Run without installing using uv
142
+ 1. Clone the git repository and `cd` into it
143
+ ```bash
144
+ git clone https://www.github.com/Bloomy52/ai-data-summarizer.git
145
+ cd ai-data-summarizer
146
+ ```
147
+
148
+ 2. Rename `.env.sample` to `.env` and add your API key.
149
+ See [CONFIGURATION.md](CONFIGURATION.md) for details.
150
+
151
+ 3. Run the tool with uv
152
+ ```bash
153
+ uv run main.py
154
+ ```
155
+
156
+ ### Option C: Run without installing using pip
157
+ 1. Clone the git repository and `cd` into it
158
+ ```bash
159
+ git clone https://www.github.com/Bloomy52/ai-data-summarizer.git
160
+ cd ai-data-summarizer
161
+ ```
162
+
163
+ 2. Rename `.env.sample` to `.env` and add your API key.
164
+ See [CONFIGURATION.md](CONFIGURATION.md) for details.
165
+
166
+ 3. Install dependencies and run the tool according to your OS
167
+
168
+ **Linux & macOS Users**
169
+ ```bash
170
+ python3 -m venv .venv
171
+ source .venv/bin/activate
172
+ pip3 install -r requirements.txt
173
+ python3 main.py
174
+ ```
175
+ **Windows Users**
176
+ ```powershell
177
+ python -m venv .venv
178
+ .\.venv\Scripts\Activate.ps1
179
+ pip install -r requirements.txt
180
+ python main.py
181
+ ```
182
+
183
+ ### Option D: Run using a Docker container
184
+ This option requires that you have Docker installed on your computer. Information on how to install Docker is in [INSTALL.md](INSTALL.md#docker).
185
+ ```bash
186
+ docker build -t ai-data-summarizer .
187
+ docker run --rm -it --env-file .env ai-data-summarizer
188
+ ```
189
+
190
+
191
+
192
+ > [!NOTE]
193
+ > All four options run the same underlying code. Option A installs `sumdata` as a command while Options B & C run it directly via `main.py` inside a virtual environment.
194
+
195
+
196
+
197
+ ## License
198
+ This project is licensed under the MIT License. The full license text can be found [here](LICENSE)
199
+
200
+ *The original CTA example dataset came from the City of Chicago's Data Portal. You can find the original dataset link [here](https://data.cityofchicago.org/Transportation/CTA-Ridership-L-Station-Entries-Daily-Totals/5neh-572f/about_data)
@@ -0,0 +1,13 @@
1
+ apicheck.py,sha256=nyBXiEeg5BTW8PI78PgVzcmrwAZCw1QDgAGWkLVCeJQ,2842
2
+ envvar.py,sha256=ne2vizD19WupGvoeBaIw_6P26uKJobJndNOFZHxCNnU,3812
3
+ fileloader.py,sha256=y2B1GdwO6VaqqzUNjfOkwG0rKsfysSNbpyuYIquw6IY,4224
4
+ main.py,sha256=qgW6ezqWU47kvgePzs3kLUQLwXGgrO0IA5EhpXxrTfo,5309
5
+ profiler.py,sha256=GWQZKssqplad9LeRuKT9xD-RIcRzi-CRhyTtm9O-tYY,3795
6
+ prompt.py,sha256=W8jSf7K6DPdTyK9HO06fbNocmoHm1nfkpQd8lfN1wrU,7216
7
+ summarizer.py,sha256=71nRxw6aRPf0aA5gg45hQoIUSBa-MMOiuMlACld4amU,3534
8
+ tokenizer.py,sha256=4VLJ0umfy7r9Yj3QxK55B_Xq-SYM_N10PFE0CwQBYlo,1296
9
+ ai_data_summarizer-0.7.1.dist-info/METADATA,sha256=TEEJ2hmAXzHPvj4_8mVUS8tpbWx0_PsBqOItLyNO7os,8101
10
+ ai_data_summarizer-0.7.1.dist-info/WHEEL,sha256=zOwg4jB6zX2kU910N-cMawjivD6tO8NEWvE12je1bVk,87
11
+ ai_data_summarizer-0.7.1.dist-info/entry_points.txt,sha256=q7AKQVFhwm3ZFqGTzj6Se0vOEXSxBApuMA4xParXFWA,38
12
+ ai_data_summarizer-0.7.1.dist-info/licenses/LICENSE,sha256=NRgWTyaRalxqDDAqDxput6odIMzsPSV4RkkQh2cCoN0,1072
13
+ ai_data_summarizer-0.7.1.dist-info/RECORD,,
@@ -0,0 +1,4 @@
1
+ Wheel-Version: 1.0
2
+ Generator: hatchling 1.32.0
3
+ Root-Is-Purelib: true
4
+ Tag: py3-none-any
@@ -0,0 +1,2 @@
1
+ [console_scripts]
2
+ sumdata = main:main
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Louie Bloomberg
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
apicheck.py ADDED
@@ -0,0 +1,86 @@
1
+ # API Key Checker
2
+ # apicheck.py
3
+ # This file contains functions to check for API keys and token counts for the different model providers.
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ import os
7
+ import sys
8
+
9
+ import anthropic
10
+ from google import genai
11
+ from openai import OpenAI
12
+
13
+ from envvar import *
14
+
15
+ def check_api_keys(provider_choice):
16
+ if provider_choice == 1:
17
+ check_for_gemini_api_key()
18
+ check_valid_gemini_api_key()
19
+ elif provider_choice == 2:
20
+ check_for_openai_api_key()
21
+ check_valid_openai_api_key()
22
+ elif provider_choice == 3:
23
+ check_for_anthropic_api_key()
24
+ check_valid_anthropic_api_key()
25
+
26
+ def check_for_gemini_api_key():
27
+ if os.getenv("GEMINI_API_KEY") is None or os.getenv("GEMINI_API_KEY") == "":
28
+ print("Error: Gemini API key not found. Would you like to set it now? (Y/n)")
29
+ if input().lower() != 'n':
30
+ set_api_keys(1)
31
+ else:
32
+ sys.exit(1)
33
+ else:
34
+ return True
35
+
36
+ def check_valid_gemini_api_key():
37
+ try:
38
+ client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
39
+ # Attempt to list models to check if the API key is valid
40
+ models = client.models.list()
41
+ return True
42
+ except Exception as e:
43
+ print(f"Error: Invalid Gemini API key. Please check your .env file. Details: {e}")
44
+ sys.exit(1)
45
+
46
+
47
+ def check_for_openai_api_key():
48
+ if os.getenv("OPENAI_API_KEY") is None or os.getenv("OPENAI_API_KEY") == "":
49
+ print("Error: OpenAI API key not found. Would you like to set it now? (Y/n)")
50
+ if input().lower() != 'n':
51
+ set_api_keys(2)
52
+ else:
53
+ sys.exit(1)
54
+ else:
55
+ return True
56
+
57
+ def check_valid_openai_api_key():
58
+ try:
59
+ client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
60
+ # Attempt to list models to check if the API key is valid
61
+ # TODO: Implement correct method to validate OpenAI API key
62
+ return True
63
+ except Exception as e:
64
+ print(f"Error: Invalid OpenAI API key. Please check your .env file. Details: {e}")
65
+ sys.exit(1)
66
+
67
+
68
+ def check_for_anthropic_api_key():
69
+ if os.getenv("ANTHROPIC_API_KEY") is None or os.getenv("ANTHROPIC_API_KEY") == "":
70
+ print("Error: Anthropic API key not found. Would you like to set it now? (Y/n)")
71
+ if input().lower() != 'n':
72
+ set_api_keys(3)
73
+ else:
74
+ sys.exit(1)
75
+ else:
76
+ return True
77
+
78
+ def check_valid_anthropic_api_key():
79
+ try:
80
+ client = anthropic.Client(api_key=os.getenv("ANTHROPIC_API_KEY"))
81
+ # Attempt to list models to check if the API key is valid
82
+ # TODO: Implement correct method to validate Anthropic API key
83
+ return True
84
+ except Exception as e:
85
+ print(f"Error: Invalid Anthropic API key. Please check your .env file. Details: {e}")
86
+ sys.exit(1)
envvar.py ADDED
@@ -0,0 +1,99 @@
1
+ # Environment Variable Management
2
+ # envvar.py
3
+ # This file contains functions to manage environment variables, including reading and writing to the .env file
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ import os
7
+ import pwinput
8
+
9
+
10
+ def _strip_quotes(value):
11
+ # Remove a single matching pair of surrounding quotes (single or double), if present.
12
+ value = value.strip()
13
+ if len(value) >= 2 and value[0] == value[-1] and value[0] in ('"', "'"):
14
+ return value[1:-1]
15
+ return value
16
+
17
+
18
+ def write_env_value(env_path, key_name, value):
19
+ # Update key_name's value in the .env file in place, preserving other lines.
20
+ # If key_name isn't found in the file, it gets appended.
21
+ lines = []
22
+ if os.path.exists(env_path):
23
+ with open(env_path, "r") as env_file:
24
+ lines = env_file.readlines()
25
+
26
+ found = False
27
+ for i, line in enumerate(lines):
28
+ stripped = line.strip()
29
+ if "=" in stripped and not stripped.startswith("#"):
30
+ existing_key = stripped.split("=", 1)[0].strip()
31
+ if existing_key == key_name:
32
+ lines[i] = f"{key_name}={value}\n"
33
+ found = True
34
+ break
35
+
36
+ if not found:
37
+ lines.append(f"{key_name}={value}\n")
38
+
39
+ with open(env_path, "w") as env_file:
40
+ env_file.writelines(lines)
41
+
42
+
43
+ def create_env():
44
+ # Create a .env file if it doesn't exist
45
+ base_dir = os.path.dirname(os.path.abspath(__file__))
46
+ env_path = os.path.join(base_dir, ".env")
47
+ # Check if the .env file exists
48
+ if not os.path.exists(env_path):
49
+ # have system create a .env file with default values
50
+ with open(env_path, "w") as env_file:
51
+ env_file.write("GEMINI_API_KEY=\n")
52
+ env_file.write("OPENAI_API_KEY=\n")
53
+ env_file.write("ANTHROPIC_API_KEY=\n")
54
+ env_file.write("GEMINI_FREE_TIER=True\n") # TODO: Remove Free Tier Flag and make that hardcoded
55
+
56
+ return None
57
+
58
+ def read_env():
59
+ # read full .env and set environment variables
60
+ base_dir = os.path.dirname(os.path.abspath(__file__))
61
+ env_path = os.path.join(base_dir, ".env")
62
+ if not os.path.exists(env_path):
63
+ create_env()
64
+ if os.path.exists(env_path):
65
+ with open(env_path, "r") as env_file:
66
+ for line in env_file:
67
+ if "=" in line:
68
+ key, value = line.strip().split("=", 1)
69
+ value = _strip_quotes(value)
70
+ os.environ[key.strip()] = value
71
+ return None
72
+
73
+ def set_api_keys(model_provider_choice):
74
+ base_dir = os.path.dirname(os.path.abspath(__file__))
75
+ env_path = os.path.join(base_dir, ".env")
76
+
77
+ # Select the appropriate API key based on the model provider choice
78
+ if model_provider_choice == 1:
79
+ api_key_name = "GEMINI_API_KEY"
80
+ elif model_provider_choice == 2:
81
+ api_key_name = "OPENAI_API_KEY"
82
+ elif model_provider_choice == 3:
83
+ api_key_name = "ANTHROPIC_API_KEY"
84
+
85
+ if os.getenv(api_key_name): # Check to see if environment variables are already in RAM
86
+ return True
87
+
88
+ read_env() # Read the .env file to set environment variables
89
+
90
+ # If environment variable is not in .env nor in RAM, prompt the user to enter the API key and save it to .env securely
91
+ if not os.getenv(api_key_name):
92
+ api_key = pwinput.pwinput(prompt=f"Enter your {api_key_name} (or leave blank to skip): ", mask="*").strip()
93
+ api_key = _strip_quotes(api_key) # Allow the user to paste a key wrapped in quotes
94
+ if api_key:
95
+ os.environ[api_key_name] = api_key
96
+ write_env_value(env_path, api_key_name, api_key)
97
+ print(f"{api_key_name} has been set and saved to {env_path}.")
98
+ else:
99
+ print(f"Warning: {api_key_name} is not set. You will not be able to use the selected model provider.")
fileloader.py ADDED
@@ -0,0 +1,142 @@
1
+ # Dataset File Loader
2
+ # fileloader.py
3
+ # This module provides functions to load and read CSV files for processing.
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ import os
7
+ import sys
8
+ import pandas as pd
9
+
10
+
11
+ def detect_file_type(filepath):
12
+ """
13
+ Determines the normalized file type ("csv" or "excel") based on the file extension.
14
+ Raises a ValueError if the extension is not supported.
15
+
16
+ Returns: str: "csv" or "excel"
17
+ """
18
+ _, ext = os.path.splitext(filepath)
19
+ ext = ext.lower()
20
+
21
+ if ext == ".csv":
22
+ return "csv"
23
+ elif ext in [".xlsx", ".xls"]:
24
+ return "excel"
25
+ else:
26
+ print(f"Unsupported file extension '{ext}'. Supported extensions are: .csv, .xlsx, .xls")
27
+ sys.exit(1)
28
+
29
+
30
+ def get_sheet_names(filepath):
31
+ """
32
+ Returns a list of sheet names if the file is an Excel workbook.
33
+ Returns None for CSV files, since they do not have sheets.
34
+
35
+ Returns: list[str] | None
36
+ """
37
+ file_type = detect_file_type(filepath)
38
+
39
+ if file_type == "csv":
40
+ return None
41
+
42
+ try:
43
+ excel_file = pd.ExcelFile(filepath)
44
+ except Exception as e:
45
+ print(f"Could not read Excel file '{filepath}'. Details: {e}")
46
+ sys.exit(1)
47
+
48
+ return excel_file.sheet_names
49
+
50
+ def choose_sheet(filepath, file_type):
51
+ if file_type == "excel":
52
+ sheet_names = get_sheet_names(filepath)
53
+ if sheet_names is None:
54
+ print(f"Error: Could not retrieve sheet names from Excel file '{filepath}'.")
55
+ sys.exit(1)
56
+
57
+ print("\nAvailable sheets:")
58
+ for idx, sheet in enumerate(sheet_names):
59
+ print(f"{idx + 1}. {sheet}")
60
+
61
+ while True:
62
+ sheet_choice = input("Enter the number of the sheet you want to load: ").strip()
63
+ try:
64
+ sheet_index = int(sheet_choice) - 1
65
+ if 0 <= sheet_index < len(sheet_names):
66
+ selected_sheet = sheet_names[sheet_index]
67
+ break
68
+ else:
69
+ print("Invalid choice. Please enter a valid number.")
70
+ except Exception as e:
71
+ print(f"Invalid input. Please enter a number. Details: {e}")
72
+ sys.exit(1)
73
+
74
+ # Load the selected sheet as CSV text
75
+ try:
76
+ dataframe = pd.read_excel(filepath, sheet_name=selected_sheet)
77
+ csv_text = dataframe.to_csv(index=False)
78
+ return csv_text
79
+ except Exception as e:
80
+ print(f"Error: Could not read the selected sheet '{selected_sheet}' from Excel file '{filepath}'. Details: {e}")
81
+ sys.exit(1)
82
+
83
+
84
+ def load_csv(filepath):
85
+ """
86
+ Loads a CSV file and returns it as normalized CSV-formatted text.
87
+
88
+ Returns: str
89
+ """
90
+ try:
91
+ dataframe = pd.read_csv(filepath, thousands=',', decimal='.')
92
+ except Exception as e:
93
+ print(f"Could not read CSV file '{filepath}'. Details: {e}")
94
+ sys.exit(1)
95
+
96
+ if dataframe.empty:
97
+ print(f"CSV file '{filepath}' appears to be empty.")
98
+ sys.exit(1)
99
+
100
+ return dataframe
101
+
102
+
103
+ def load_excel(filepath, sheet_name=None):
104
+ """
105
+ Loads a single sheet from an Excel file (.xlsx or .xls) and returns it as
106
+ normalized CSV-formatted text. If sheet_name is not provided, the first
107
+ sheet in the workbook is used.
108
+
109
+ Returns: str
110
+ """
111
+ try:
112
+ if sheet_name is None:
113
+ available_sheets = get_sheet_names(filepath)
114
+ sheet_name = available_sheets[0]
115
+
116
+ dataframe = pd.read_excel(filepath, sheet_name=sheet_name)
117
+ except Exception as e:
118
+ print(f"Could not read Excel file '{filepath}'. Details: {e}")
119
+ sys.exit(1)
120
+
121
+ if dataframe.empty:
122
+ print(f"Sheet '{sheet_name}' in file '{filepath}' appears to be empty.")
123
+ sys.exit(1)
124
+
125
+ return dataframe
126
+
127
+
128
+ def load_file(filepath, file_type):
129
+ """
130
+ Dispatches to the correct loader based on the file's detected type, and
131
+ returns normalized CSV-formatted text regardless of the original source format.
132
+
133
+ sheet_name is only relevant for Excel files. It is ignored for CSV files.
134
+
135
+ Returns: str
136
+ """
137
+
138
+ if file_type == "csv":
139
+ return load_csv(filepath)
140
+ elif file_type == "excel":
141
+ return load_excel(filepath)
142
+ return None
main.py ADDED
@@ -0,0 +1,155 @@
1
+ # Main Code File
2
+ # main.py
3
+ # Code file containing the main function and the main logic for the AI Data Summarizer program.
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ # Import Statements
7
+ import os
8
+ import sys
9
+ import datetime
10
+ import csv
11
+ from readchar import readkey, key
12
+
13
+ # Import the functions from the summarizer.py helper file
14
+ from summarizer import *
15
+ from tokenizer import *
16
+ from prompt import *
17
+ from apicheck import *
18
+ from envvar import *
19
+ from fileloader import *
20
+ from profiler import *
21
+
22
+ # Function Definitions
23
+
24
+ def check_tokens_gemini(prompt, text):
25
+ # This function checks the number of tokens to make sure that they are within the Free Tier limits
26
+ google_tokens = google_tokenizer(prompt, text)
27
+ if os.getenv("GEMINI_FREE_TIER") == "True" and google_tokens > 250000:
28
+ print(f"Warning: Your input text has {google_tokens} tokens, which exceeds the free tier limit of 250,000 tokens per minute for Gemini. Consider reducing the input size or upgrading your plan. ")
29
+ sys.exit(1)
30
+ elif os.getenv("GEMINI_FREE_TIER") == "True":
31
+ print(f"Your input text has {google_tokens} tokens, which is within the free tier limit for Gemini.")
32
+ print("Would you like to continue? (Y/n)")
33
+ while True:
34
+ k = readkey()
35
+ if k == 'y' or k == 'Y' or k == key.ENTER:
36
+ return None
37
+ elif k == 'n' or k == 'N':
38
+ print("Exiting...")
39
+ sys.exit(1)
40
+ else:
41
+ print("Invalid input. Please enter y or n.")
42
+
43
+ # TODO (maybe): Add cost functionality to estimate cost of input response
44
+ return None
45
+
46
+ def check_tokens_openai(prompt, text):
47
+ tokens = openai_tokenizer(prompt, text)
48
+ print(f"Your input text has {tokens} tokens.")
49
+ return None
50
+
51
+ def check_tokens_anthropic(prompt, text):
52
+ tokens = anthropic_tokenizer(prompt, text)
53
+ print(f"Your input text has {tokens} tokens.")
54
+ return None
55
+
56
+ def get_model_provider():
57
+ while True:
58
+ print("\nSelect a model provider:")
59
+ print("1. Google Gemini")
60
+ print("2. OpenAI")
61
+ print("3. Anthropic")
62
+ print("Or press q to quit the program.")
63
+ k = readkey()
64
+
65
+ if k == '1':
66
+ return 1 # Gemini
67
+ elif k == '2':
68
+ return 2 # OpenAI
69
+ elif k == '3':
70
+ return 3 # Anthropic
71
+ elif k == 'q':
72
+ print("Exiting...")
73
+ sys.exit(1)
74
+ else:
75
+ print("Invalid choice. Please enter 1, 2, or 3.")
76
+
77
+ # Define Prompt Choosing Function
78
+ def get_prompt_type():
79
+ """
80
+ This function allows the user to select a prompt from a list of available prompts.
81
+ It returns the selected prompt as a string.
82
+ Returns: str: The selected prompt.
83
+ """
84
+ while True:
85
+ print("\nSelect a Summary:")
86
+ print("1. Just the Facts (AI is not used)")
87
+ print("2. TL;DR Summary")
88
+ print("3. Data Overview")
89
+ print("4. Deep Dive Analysis")
90
+ print("Or press q to quit the program.")
91
+ k = readkey()
92
+
93
+ if k == '2':
94
+ return "tldr"
95
+ elif k == '3':
96
+ return "overview"
97
+ elif k == '4':
98
+ return "deepdive"
99
+ elif k == '1':
100
+ return "facts"
101
+ elif k == 'q':
102
+ print("Exiting...")
103
+ sys.exit(1)
104
+ else:
105
+ print("Invalid choice. Please enter 1, 2, 3, 4, or q.")
106
+
107
+ # Main Function
108
+ def main():
109
+ read_env() # Read the .env file to set environment variables
110
+ # Main Interactive Loop:
111
+ # Get input file
112
+ while True:
113
+ filepath = input("\nEnter the path to your CSV file: ").strip()
114
+
115
+ # Remove quotes if user wrapped path in quotes
116
+ input_file = filepath.strip('"\'')
117
+
118
+
119
+ if not os.path.exists(input_file):
120
+ print(f"Error: File '{input_file}' not found. Please try again.")
121
+ continue
122
+
123
+ if not input_file.lower().endswith('.csv') and not input_file.lower().endswith('.xlsx') and not input_file.lower().endswith('.xls'):
124
+ print("Warning: file extension is not .csv, .xlsx, or .xls. Continue anyway? (y/N)")
125
+ if input().lower() != 'y':
126
+ continue
127
+
128
+ break
129
+
130
+ file_type = detect_file_type(input_file)
131
+
132
+ df = load_file(input_file, file_type)
133
+ profile = build_base_profile(df)
134
+ sample = get_sample(df)
135
+ csv_text = "\nData Profile:" + profile + "\nData Sample:\n" + sample
136
+
137
+ model_provider_choice = 1 # get_model_provider()
138
+ # Defaults to Gemini since it is the only one provided
139
+ check_api_keys(model_provider_choice)
140
+ prompt_type = get_prompt_type()
141
+ if prompt_type == "facts":
142
+ summary = "Dataset Statistical Summary:\n" + profile + "\nData Sample:\n" + sample
143
+ write_outfile(summary, os.path.basename(input_file).split(".")[0], prompt_type, "None")
144
+ else:
145
+ prompt = get_prompt(prompt_type)
146
+ check_tokens_gemini(prompt, csv_text)
147
+ summary = gemini_summarizer(prompt, csv_text, os.path.basename(input_file).split(".")[0], prompt_type)
148
+
149
+ print("\nSummary:\n")
150
+ print(summary)
151
+
152
+ return None
153
+
154
+ if __name__ == "__main__":
155
+ main()
profiler.py ADDED
@@ -0,0 +1,108 @@
1
+ # Pandas Pre-Summarization Profiler
2
+ # profiler.py
3
+ # This module provides functions to profile a DataFrame before summarization.
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ import pandas as pd
7
+
8
+
9
+ def _parse_dates(df):
10
+ # Returns a copy of df with any date-like object columns converted to datetime
11
+ df = df.copy()
12
+ for col in df.select_dtypes(include="object").columns:
13
+ try:
14
+ converted = pd.to_datetime(df[col], format="mixed", errors="coerce")
15
+ if converted.notna().sum() > len(df) * 0.9: # 90%+ parsed successfully
16
+ df[col] = converted
17
+ except Exception:
18
+ continue
19
+ return df
20
+
21
+
22
+ def build_base_profile(df):
23
+ """
24
+ Returns a plain-text statistical profile of the DataFrame.
25
+ Includes shape, dtypes, null counts, numeric stats, and top categorical values.
26
+ Does not mutate the input DataFrame.
27
+
28
+ Returns: str
29
+ """
30
+ df = _parse_dates(df)
31
+ lines = []
32
+
33
+ # Shape
34
+ num_rows, num_cols = df.shape
35
+ lines.append(f"Rows: {num_rows}")
36
+ lines.append(f"Columns: {num_cols}")
37
+ lines.append("")
38
+
39
+ # Dtypes and null counts
40
+ lines.append("Column Overview:")
41
+ for col in df.columns:
42
+ null_count = df[col].isna().sum()
43
+ lines.append(f" {col} ({df[col].dtype}) — {null_count} nulls")
44
+ lines.append("")
45
+
46
+ # Numeric stats
47
+ numeric_cols = df.select_dtypes(include="number").columns
48
+ if len(numeric_cols) > 0:
49
+ lines.append("Numeric Column Stats:")
50
+ lines.append(df[numeric_cols].describe().to_string())
51
+ lines.append("")
52
+
53
+ # Date column summary
54
+ date_cols = df.select_dtypes(include=["datetime"]).columns
55
+ if len(date_cols) > 0:
56
+ lines.append("Date Column Summary:")
57
+ for col in date_cols:
58
+ min_date = df[col].min()
59
+ max_date = df[col].max()
60
+ if pd.isna(min_date) or pd.isna(max_date):
61
+ lines.append(f" {col}: (no non-null datetimes)")
62
+ continue
63
+ span = (max_date - min_date).days + 1
64
+ duplicate_dates = df[col].duplicated().sum()
65
+ lines.append(f" {col}: {min_date.date()} to {max_date.date()} ({span} days)")
66
+ if duplicate_dates > 0:
67
+ lines.append(f" {col} has {duplicate_dates} duplicate date(s)")
68
+ lines.append("")
69
+
70
+ # Standout values with date context
71
+ if len(date_cols) > 0 and len(numeric_cols) > 0:
72
+ date_col = date_cols[0]
73
+ show_all_cols = len(df.columns) <= 5
74
+ lines.append("Standout Values:")
75
+ for col in numeric_cols:
76
+ if show_all_cols:
77
+ top5 = df.nlargest(5, col).to_string(index=False)
78
+ bottom5 = df.nsmallest(5, col).to_string(index=False)
79
+ else:
80
+ top5 = df[[date_col, col]].nlargest(5, col).to_string(index=False)
81
+ bottom5 = df[[date_col, col]].nsmallest(5, col).to_string(index=False)
82
+ lines.append(f" Top 5 highest {col}:")
83
+ lines.append(top5)
84
+ lines.append(f" Bottom 5 lowest {col}:")
85
+ lines.append(bottom5)
86
+ lines.append("")
87
+
88
+ # Categorical top values
89
+ categorical_cols = df.select_dtypes(include=["object", "category"]).columns
90
+ if len(categorical_cols) > 0:
91
+ lines.append("Categorical Column Summary:")
92
+ for col in categorical_cols:
93
+ unique_count = df[col].nunique()
94
+ top_values = df[col].value_counts().head(5).to_string()
95
+ lines.append(f" {col} ({unique_count} unique values):")
96
+ lines.append(f"{top_values}")
97
+ lines.append("")
98
+
99
+ return "\n".join(lines)
100
+
101
+
102
+ def get_sample(df, n=10):
103
+ """
104
+ Returns the first n rows of the DataFrame as CSV-formatted text.
105
+
106
+ Returns: str
107
+ """
108
+ return df.head(n).to_csv(index=False)
prompt.py ADDED
@@ -0,0 +1,144 @@
1
+ # Main Prompt File
2
+ # prompt.py
3
+ # File contains different prompts which can be used to help the user understand the file
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ # Define Prompt Choosing Functions
7
+ def get_prompt(prompt_type):
8
+ """
9
+ This function retrieves the appropriate prompt based on the user's selection.
10
+ Returns: str: The selected prompt.
11
+ """
12
+ if prompt_type == "tldr":
13
+ return get_tldr_prompt()
14
+ elif prompt_type == "overview":
15
+ return get_overview_prompt()
16
+ elif prompt_type == "deepdive":
17
+ return get_deepdive_prompt()
18
+
19
+
20
+ # Define Prompts that can be used
21
+
22
+ def get_tldr_prompt():
23
+ task_summary = f"""
24
+ ## Task Summary:
25
+ {{Produce a 'Too Long; Didn’t Read' (TL;DR) summary of the attached dataset and statistical summary. The TL;DR sentence should describe, in one high-level line, what the dataset is and what it covers. Then provide 2–3 bullets highlighting the most notable patterns, trends, or anomalies visible in the data. Use numeric ranges when helpful, but keep the focus on the biggest takeaways.}}
26
+ """
27
+
28
+ response_style = f"""
29
+ ## Response style and format requirements:
30
+ - {{Write in a sharp, punchy, and bottom-line-up-front (BLUF) style}}
31
+ - {{Format: Start with a single 'TL;DR:' sentence, followed by a short bulleted list}}
32
+ - {{Bullets must follow these rules:
33
+ - Start with a short, strong label (e.g., 'Trend shift:', 'Category contrast:', 'Peak anomaly:')
34
+ - Contain exactly one idea per bullet
35
+ - Use numeric anchors when possible
36
+ - Avoid hedging, filler, or multi-clause sentences
37
+ - Read like headlines, not explanations}}
38
+ - {{Numbers in the stats block are exact — use them directly, don't estimate from sample rows}}
39
+ - {{Strictly limit the response to 100 words or less}}
40
+ - {{Use plain text only — no Markdown. Bullets are to be noted with '-'}}
41
+ """
42
+
43
+ final_prompt = f"""{task_summary}
44
+ {response_style}"""
45
+
46
+ return final_prompt
47
+
48
+
49
+ def get_overview_prompt():
50
+ # Prompt for the Summarizer:
51
+ # Use this to clearly define the task and job needed by the model
52
+ task_summary = f"""
53
+ ## Task Summary:
54
+ {{Review the attached CSV and statistical summary and summarize what the data covers, including anything notable or unusual.}}
55
+ """
56
+
57
+ # Use this to provide contextual information related to the task
58
+ context_information = f"""
59
+ ## Context Information:
60
+ - {{Standard CSV format with headers in the first row}}
61
+ - {{Columns may include numbers, text, or dates}}
62
+ - {{Treat all dates in the data file as a recorded value and not predictions}}
63
+ - {{The dataset may cover any domain — do not assume a specific subject area}}
64
+ - {{The stats block (counts, mean/median, ranges, date span, category counts) is precomputed and exact — treat these numbers as ground truth}}
65
+ - {{Standout/extreme rows shown are the most unusual in the dataset, not typical examples}}
66
+ - {{Any raw sample rows are shown only to illustrate formatting and column meaning, not to infer statistics}}
67
+ """
68
+
69
+ # Use this to provide any model instructions that you want model to adhere to
70
+ model_instructions = f"""
71
+ ## Model Instructions:
72
+ - {{Explain what each column represents and flag anything out of the ordinary}}
73
+ - {{Base all observations only on the data provided}}
74
+ - {{Only use historical context or external knowledge where it clearly and directly explains a specific data pattern — do not force connections}}
75
+ - {{If no external context is relevant, rely entirely on what the data shows}}
76
+ """
77
+
78
+ # Use this to provide response style and formatting guidance
79
+ response_style = f"""
80
+ ## Response style and format requirements:
81
+ - {{Write a standalone written summary as if briefing a coworker}}
82
+ - {{Use three sections: overview, column breakdown, and key takeaways}}
83
+ - {{Limit the response to 500 words or less}}
84
+ - {{Use plain text only — no Markdown. Bullets are to be noted with '-'. No Markdown headings.}}
85
+ """
86
+
87
+ # Concatenate to final prompt
88
+ final_prompt = f"""{task_summary}
89
+ {context_information}
90
+ {model_instructions}
91
+ {response_style}"""
92
+
93
+ return final_prompt
94
+
95
+
96
+ def get_deepdive_prompt():
97
+ task_summary = f"""
98
+ ## Task Summary:
99
+ {{Perform a detailed deep dive analysis of the attached CSV dataset and statistical summary. Go beyond surface-level description to examine distributions, patterns, relationships, and notable characteristics in the data.}}
100
+ """
101
+
102
+ context_information = f"""
103
+ ## Context Information:
104
+ - {{Standard CSV format with headers in the first row}}
105
+ - {{Columns may include numbers, text, dates, or categorical values}}
106
+ - {{Treat all dates as recorded historical values}}
107
+ - {{The dataset may cover any domain — analyze based solely on what is present}}
108
+ """
109
+
110
+ input_composition = f"""
111
+ ## Input Composition:
112
+ - {{The stats block (counts, mean, median, std, min/max, date span, category counts) is precomputed directly from the full dataset and is exact — treat it as ground truth, never recompute or estimate these figures yourself}}
113
+ - {{The standout/extreme rows (top and bottom values) represent the most unusual points in the entire dataset, not a representative sample — use them only to discuss outliers, spikes, or anomalies}}
114
+ - {{The raw sample rows are a small illustrative slice included only to show formatting, column meaning, and qualitative texture — they are not necessarily representative of the full distribution and should not be used to infer statistics}}
115
+ """
116
+
117
+ model_instructions = f"""
118
+ ## Model Instructions:
119
+ - {{For numeric columns: report range (min-max), central tendency (mean/median if relevant), and distribution shape}}
120
+ - {{For numeric columns: also describe distribution shape based on the relationship between mean, median, and standard deviation}}
121
+ - {{For categorical/text columns: list top unique values with counts and note any dominant categories}}
122
+ - {{Identify any clear relationships or correlations between columns that stand out}}
123
+ - {{Highlight temporal patterns if dates are present, or geographic patterns if location data exists}}
124
+ - {{Flag outliers, unusual spikes/drops, or data quality concerns}}
125
+ - {{Base every observation strictly on the data provided — do not speculate beyond visible evidence}}
126
+ - {{Suggest 2–3 specific follow-up questions or analyses the user could explore next}}
127
+ """
128
+
129
+ response_style = f"""
130
+ ## Response style and format requirements:
131
+ - {{Write as if explaining the data in detail to a data-savvy coworker}}
132
+ - {{Use these four clear sections in order: Overview, Column Analysis, Key Patterns & Relationships, Takeaways & Next Steps}}
133
+ - {{Use plain text only — no Markdown. Bullets are to be noted with '-'. No Markdown headings in the final output.}}
134
+ - {{Keep the total response under 800 words}}
135
+ - {{Be specific and quantitative where possible}}
136
+ """
137
+
138
+ final_prompt = f"""{task_summary}
139
+ {context_information}
140
+ {input_composition}
141
+ {model_instructions}
142
+ {response_style}"""
143
+
144
+ return final_prompt
summarizer.py ADDED
@@ -0,0 +1,109 @@
1
+ # Dataset Summarizer
2
+ # summarizer.py
3
+ # Code File for calling the summarization functions -- Mainly a helper file to keep code organized
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ # Import Statements
7
+ import os
8
+ import sys
9
+ from zoneinfo import ZoneInfo
10
+ import datetime as dt
11
+ import csv
12
+
13
+ # Importing AI Libraries
14
+ from openai import OpenAI
15
+ import anthropic
16
+ from google import genai
17
+ from google.genai import types
18
+
19
+
20
+ # Function Definitions
21
+ # TODO: Implement summarization functions here
22
+
23
+ # Write File Function
24
+ def write_outfile(output, filename, prompt_type, model_id):
25
+ # Create summaries directory if it doesn't exist
26
+ summaries_dir = f"./summaries/{prompt_type}"
27
+ os.makedirs(summaries_dir, exist_ok=True)
28
+
29
+ now = dt.datetime.now()
30
+ date_time_string = now.strftime("%Y-%m-%d %H-%M-%S")
31
+ central_tz = ZoneInfo("America/Chicago")
32
+ date_written = dt.datetime.now(central_tz).strftime("%A, %B %d, %Y")
33
+ time_written = dt.datetime.now(central_tz).strftime("%I:%M %p %Z")
34
+ summary_local_path = f"{summaries_dir}/{date_time_string}.{filename}.{model_id}.txt"
35
+ with open(summary_local_path, "w", encoding="utf-8") as outfile:
36
+ # Add Header Section to the output file
37
+ outfile.write("=" * 10 + "BEGIN HEADER" + "=" * 10 + "\n")
38
+ outfile.write("Date Written: " + date_written + "\n")
39
+ outfile.write("Time Written: " + time_written + "\n")
40
+ outfile.write("Model Used: " + model_id + "\n")
41
+ outfile.write("Prompt Type: " + prompt_type + "\n")
42
+ outfile.write("File Name: " + filename + "\n")
43
+ outfile.write("=" * 10 + "END HEADER" + "=" * 10 + "\n\n")
44
+
45
+ # Add main summary output to the output file
46
+ outfile.write(output)
47
+
48
+ return None
49
+
50
+ # Gemini Summarizer
51
+ def gemini_summarizer(prompt, text, filename, prompt_type):
52
+ MODEL_ID = "gemini-3.5-flash"
53
+ try:
54
+ client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
55
+
56
+ response = client.models.generate_content(
57
+ model=MODEL_ID,
58
+ contents=[
59
+ text,
60
+ prompt,
61
+ ]
62
+ )
63
+ output = response.text
64
+ except Exception as e:
65
+ print(f"Error generating summary: {e}")
66
+ sys.exit(1)
67
+
68
+ write_outfile(output, filename, prompt_type, MODEL_ID)
69
+ return output
70
+
71
+
72
+ def openai_summarizer(prompt, text, filename, prompt_type):
73
+ MODEL_ID = "gpt-5.4-mini"
74
+ try:
75
+ client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
76
+
77
+ response = client.responses.create(
78
+ model=MODEL_ID,
79
+ messages=[
80
+ {"role": "system", "content": "You are a helpful assistant."},
81
+ {"role": "user", "content": prompt + "\n\n" + text},
82
+ ],
83
+ )
84
+ output = response.output_text
85
+ except Exception as e:
86
+ print(f"Error generating summary: {e}")
87
+ sys.exit(1)
88
+
89
+ write_outfile(output, filename, prompt_type, MODEL_ID)
90
+ return output
91
+
92
+ def anthropic_summarizer(prompt, text, filename, prompt_type):
93
+ MODEL_ID = "claude-haiku-4-5"
94
+ try:
95
+ client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
96
+
97
+ message = client.messages.create(
98
+ model=MODEL_ID,
99
+ messages=[
100
+ {"role": "user", "content": prompt + "\n\n" + text},
101
+ ],
102
+ )
103
+ output = message.content[0].text
104
+ except Exception as e:
105
+ print(f"Error generating summary: {e}")
106
+ sys.exit(1)
107
+
108
+ write_outfile(output, filename, prompt_type, MODEL_ID)
109
+ return output
tokenizer.py ADDED
@@ -0,0 +1,48 @@
1
+ # Dataset Tokenizer
2
+ # tokenizer.py
3
+ # Code File for calling the tokenization functions -- Mainly a helper file to keep code organized
4
+ # SPDX-License-Identifier: MIT
5
+
6
+ import os
7
+ import sys
8
+
9
+ from openai import OpenAI
10
+ import anthropic
11
+ from google import genai
12
+ from google.genai import local_tokenizer
13
+
14
+ # Tokenization Functions
15
+ # Google GenAI Tokenizer
16
+ def google_tokenizer(prompt, text):
17
+ MODEL_ID = "gemini-3.5-flash"
18
+ client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
19
+
20
+ response = client.models.count_tokens(
21
+ model=MODEL_ID,
22
+ contents=[
23
+ text,
24
+ prompt,
25
+ ]
26
+ )
27
+ tokens = response.total_tokens
28
+ return tokens
29
+
30
+ # OpenAI Tokenizer
31
+ def openai_tokenizer(prompt, text):
32
+ client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
33
+ response = client.responses.input_tokens.count(
34
+ model="gpt-5.4-mini",
35
+ instructions=prompt,
36
+ input=text,
37
+ )
38
+ return response.input_tokens
39
+
40
+ # Anthropic Tokenizer
41
+ def anthropic_tokenizer(prompt, text):
42
+ client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
43
+ response = client.messages.count_tokens(
44
+ model="claude-haiku-4-5",
45
+ system=prompt,
46
+ messages=[{"role": "user", "content": text}],
47
+ )
48
+ return response.get("input_tokens", 0)