ai-data-summarizer 0.7.1__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- ai_data_summarizer-0.7.1.dist-info/METADATA +200 -0
- ai_data_summarizer-0.7.1.dist-info/RECORD +13 -0
- ai_data_summarizer-0.7.1.dist-info/WHEEL +4 -0
- ai_data_summarizer-0.7.1.dist-info/entry_points.txt +2 -0
- ai_data_summarizer-0.7.1.dist-info/licenses/LICENSE +21 -0
- apicheck.py +86 -0
- envvar.py +99 -0
- fileloader.py +142 -0
- main.py +155 -0
- profiler.py +108 -0
- prompt.py +144 -0
- summarizer.py +109 -0
- tokenizer.py +48 -0
|
@@ -0,0 +1,200 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: ai-data-summarizer
|
|
3
|
+
Version: 0.7.1
|
|
4
|
+
Summary: A local CLI tool that uses pandas and AI to summarize and interpret datasets.
|
|
5
|
+
Project-URL: Homepage, https://github.com/Bloomy52/ai-data-summarizer
|
|
6
|
+
Author: Louie Bloomberg
|
|
7
|
+
License-Expression: MIT
|
|
8
|
+
License-File: LICENSE
|
|
9
|
+
Requires-Python: >=3.12
|
|
10
|
+
Requires-Dist: anthropic
|
|
11
|
+
Requires-Dist: google-genai
|
|
12
|
+
Requires-Dist: openai
|
|
13
|
+
Requires-Dist: openpyxl
|
|
14
|
+
Requires-Dist: pandas
|
|
15
|
+
Requires-Dist: pandas-stubs>=3.0.5.260730
|
|
16
|
+
Requires-Dist: protobuf
|
|
17
|
+
Requires-Dist: pwinput>=1.0.3
|
|
18
|
+
Requires-Dist: readchar>=4.2.2
|
|
19
|
+
Requires-Dist: sentencepiece
|
|
20
|
+
Description-Content-Type: text/markdown
|
|
21
|
+
|
|
22
|
+
# AI Data Summarization Tool
|
|
23
|
+
A Python Tool that takes a dataset and uses pandas to generate a statistical summary and Artifical Intelligence (AI) to interpret the pandas summary and presents it to the user.
|
|
24
|
+
|
|
25
|
+
This project is a continuation of my CS178 Final Project. You can find the repo at [Bloomy52/cs178-data-summarizer-py](https://www.github.com/Bloomy52/cs178-data-summarizer-py)
|
|
26
|
+
|
|
27
|
+
### Why This Exists
|
|
28
|
+
I created this project because I found that it is difficult to understand what an underlying dataset is and what it entails without reading and understanding the full dataset. I found that using a Large Language Model (LLM) to summarize the dataset made the dataset more approachable since I had a general understanding of what the dataset was and some features about said dataset I was analyzing. I specifically crafted the summary templates so they would help the user understand the dataset and its features. It can also give you a heads up if there are any concerns or anomalies before you start fully analyzing the data to prevent issues and to guide the user on the right path to analysis.
|
|
29
|
+
|
|
30
|
+
## Repo Structure
|
|
31
|
+
The structure of this git repository is as follows:
|
|
32
|
+
```text
|
|
33
|
+
ai-data-summarizer/
|
|
34
|
+
├── main.py # CLI entry point and main application logic
|
|
35
|
+
├── summarizer.py # Core data summarization functionality
|
|
36
|
+
├── prompt.py # Prompt templates and management
|
|
37
|
+
├── tokenizer.py # Token counting and management utilities
|
|
38
|
+
├── .env.sample # Sample credentials file (rename to .env)
|
|
39
|
+
├── apicheck.py # Checks API variables to prevent early issues
|
|
40
|
+
├── envvar.py # Loads environment variables from .env file
|
|
41
|
+
├── fileloader.py # Loads and parses CSV and Excel files
|
|
42
|
+
├── profiler.py # Profiles the dataset for pandas summarization
|
|
43
|
+
├── requirements.txt # Python dependencies
|
|
44
|
+
├── README.md # Project documentation
|
|
45
|
+
├── INSTALL.md # Project installation documentation
|
|
46
|
+
├── CONFIGURATION.md # Project configuration documentation
|
|
47
|
+
├── LICENSE # MIT License
|
|
48
|
+
├── .gitignore # Git ignore rules
|
|
49
|
+
├── pyproject.toml # Python project application configuration files
|
|
50
|
+
├── uv.lock # Python/uv dependency lock file
|
|
51
|
+
├── Dockerfile # Dockerfile for containerized deployment
|
|
52
|
+
├── .dockerignore # Docker ignore rules
|
|
53
|
+
├── .vscode/
|
|
54
|
+
│ └── settings.json # VSCode configuration
|
|
55
|
+
├── summaries/ # Summary output folder -- gets created upon first output summary
|
|
56
|
+
├── .devcontainer/
|
|
57
|
+
│ └── devcontainer.json # Dev container configuration for VSCode
|
|
58
|
+
└── examples/
|
|
59
|
+
├── CTA_Ridership_RedLine_WrigleyField_DailyTotals.csv # Sample dataset
|
|
60
|
+
└── OverviewPrompt_CTA_Ridership_RedLine_WrigleyField_DailyTotals.txt # Sample summary
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
## Example
|
|
64
|
+
The example used here is the number of daily riders from the Addison 'L' Stop on the Chicago Transit Authority's (CTA) Red Line using the Data Overview prompt. More information about the original dataset can be found at the bottom of the README. Other prompts will be added to the example folder as well.
|
|
65
|
+
|
|
66
|
+
**CSV File Structure**
|
|
67
|
+
```csv
|
|
68
|
+
"date","daytype","rides"
|
|
69
|
+
"01/01/2001","U","1,227"
|
|
70
|
+
"01/02/2001","W","3,937"
|
|
71
|
+
"01/03/2001","W","4,329"
|
|
72
|
+
"01/04/2001","W","4,607"
|
|
73
|
+
"01/05/2001","W","4,666"
|
|
74
|
+
```
|
|
75
|
+
* `date` is when the data was recorded
|
|
76
|
+
* `daytype` is the type of day of which the data was recorded
|
|
77
|
+
- `W` is a weekday
|
|
78
|
+
- `A` is a Saturday
|
|
79
|
+
- `U` is Sunday/Holidays
|
|
80
|
+
* `rides` is the number of riders recorded on a given day
|
|
81
|
+
|
|
82
|
+
```bash
|
|
83
|
+
sumdata
|
|
84
|
+
```
|
|
85
|
+
```text
|
|
86
|
+
Enter the path to your CSV file: examples/CTA_Ridership_RedLine_WrigleyField_DailyTotals.csv
|
|
87
|
+
|
|
88
|
+
Please Select a Summary:
|
|
89
|
+
1. Just the Facts (AI is not used)
|
|
90
|
+
2. TL;DR Summary
|
|
91
|
+
3. Data Overview
|
|
92
|
+
4. Deep Dive Analysis
|
|
93
|
+
Or enter 0 to Exit
|
|
94
|
+
Enter the number corresponding to your choice: 3
|
|
95
|
+
Your input text has 195174 tokens, which is within the free tier limit for Gemini.
|
|
96
|
+
Would you like to continue? (Y/n)
|
|
97
|
+
Y
|
|
98
|
+
|
|
99
|
+
Summary:
|
|
100
|
+
|
|
101
|
+
Overview:
|
|
102
|
+
|
|
103
|
+
I reviewed the dataset you shared. It contains daily transit ride tracking data spanning over 25 years, from
|
|
104
|
+
January 1, 2001, to March 31, 2026. This extensive daily record captures long-term ridership trends and reveals
|
|
105
|
+
how public transit usage fluctuates across weekdays, weekends, and holidays, alongside the long-term impact
|
|
106
|
+
of major external disruptions.
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
The rest of the example can be found in the `examples` folder labeled [OverviewPrompt_CTA_Ridership_RedLine_WrigleyField_DailyTotals.txt](examples/OverviewPrompt_CTA_Ridership_RedLine_WrigleyField_DailyTotals.txt)
|
|
110
|
+
|
|
111
|
+
Summaries save to the `summaries` folder. It will be created automatically with the subfolder with the prompt and a header in the file.
|
|
112
|
+
|
|
113
|
+
> [!TIP]
|
|
114
|
+
> The capital `Y` means that it is the default option. You can click the `Enter`/`Return` key as a shortcut.
|
|
115
|
+
|
|
116
|
+
|
|
117
|
+
## Project Requirements
|
|
118
|
+
> [!NOTE]
|
|
119
|
+
> This only supports Gemini at the moment. Support for the Anthropic and OpenAI APIs are coming soon.
|
|
120
|
+
|
|
121
|
+
- Python
|
|
122
|
+
- uv
|
|
123
|
+
- Git
|
|
124
|
+
- Gemini API Key from Google AI Studio
|
|
125
|
+
|
|
126
|
+
See [INSTALL.md](INSTALL.md) for detailed setup instructions, including how to obtain API keys.
|
|
127
|
+
|
|
128
|
+
|
|
129
|
+
## How to Use/Installation
|
|
130
|
+
There are three options for running this program. Please choose one of these options.
|
|
131
|
+
|
|
132
|
+
### Option A: Install as a package (recommended)
|
|
133
|
+
```bash
|
|
134
|
+
uv tool install git+https://www.github.com/Bloomy52/ai-data-summarizer.git
|
|
135
|
+
```
|
|
136
|
+
Then run:
|
|
137
|
+
```bash
|
|
138
|
+
sumdata
|
|
139
|
+
```
|
|
140
|
+
|
|
141
|
+
### Option B: Run without installing using uv
|
|
142
|
+
1. Clone the git repository and `cd` into it
|
|
143
|
+
```bash
|
|
144
|
+
git clone https://www.github.com/Bloomy52/ai-data-summarizer.git
|
|
145
|
+
cd ai-data-summarizer
|
|
146
|
+
```
|
|
147
|
+
|
|
148
|
+
2. Rename `.env.sample` to `.env` and add your API key.
|
|
149
|
+
See [CONFIGURATION.md](CONFIGURATION.md) for details.
|
|
150
|
+
|
|
151
|
+
3. Run the tool with uv
|
|
152
|
+
```bash
|
|
153
|
+
uv run main.py
|
|
154
|
+
```
|
|
155
|
+
|
|
156
|
+
### Option C: Run without installing using pip
|
|
157
|
+
1. Clone the git repository and `cd` into it
|
|
158
|
+
```bash
|
|
159
|
+
git clone https://www.github.com/Bloomy52/ai-data-summarizer.git
|
|
160
|
+
cd ai-data-summarizer
|
|
161
|
+
```
|
|
162
|
+
|
|
163
|
+
2. Rename `.env.sample` to `.env` and add your API key.
|
|
164
|
+
See [CONFIGURATION.md](CONFIGURATION.md) for details.
|
|
165
|
+
|
|
166
|
+
3. Install dependencies and run the tool according to your OS
|
|
167
|
+
|
|
168
|
+
**Linux & macOS Users**
|
|
169
|
+
```bash
|
|
170
|
+
python3 -m venv .venv
|
|
171
|
+
source .venv/bin/activate
|
|
172
|
+
pip3 install -r requirements.txt
|
|
173
|
+
python3 main.py
|
|
174
|
+
```
|
|
175
|
+
**Windows Users**
|
|
176
|
+
```powershell
|
|
177
|
+
python -m venv .venv
|
|
178
|
+
.\.venv\Scripts\Activate.ps1
|
|
179
|
+
pip install -r requirements.txt
|
|
180
|
+
python main.py
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
### Option D: Run using a Docker container
|
|
184
|
+
This option requires that you have Docker installed on your computer. Information on how to install Docker is in [INSTALL.md](INSTALL.md#docker).
|
|
185
|
+
```bash
|
|
186
|
+
docker build -t ai-data-summarizer .
|
|
187
|
+
docker run --rm -it --env-file .env ai-data-summarizer
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
|
|
191
|
+
|
|
192
|
+
> [!NOTE]
|
|
193
|
+
> All four options run the same underlying code. Option A installs `sumdata` as a command while Options B & C run it directly via `main.py` inside a virtual environment.
|
|
194
|
+
|
|
195
|
+
|
|
196
|
+
|
|
197
|
+
## License
|
|
198
|
+
This project is licensed under the MIT License. The full license text can be found [here](LICENSE)
|
|
199
|
+
|
|
200
|
+
*The original CTA example dataset came from the City of Chicago's Data Portal. You can find the original dataset link [here](https://data.cityofchicago.org/Transportation/CTA-Ridership-L-Station-Entries-Daily-Totals/5neh-572f/about_data)
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
apicheck.py,sha256=nyBXiEeg5BTW8PI78PgVzcmrwAZCw1QDgAGWkLVCeJQ,2842
|
|
2
|
+
envvar.py,sha256=ne2vizD19WupGvoeBaIw_6P26uKJobJndNOFZHxCNnU,3812
|
|
3
|
+
fileloader.py,sha256=y2B1GdwO6VaqqzUNjfOkwG0rKsfysSNbpyuYIquw6IY,4224
|
|
4
|
+
main.py,sha256=qgW6ezqWU47kvgePzs3kLUQLwXGgrO0IA5EhpXxrTfo,5309
|
|
5
|
+
profiler.py,sha256=GWQZKssqplad9LeRuKT9xD-RIcRzi-CRhyTtm9O-tYY,3795
|
|
6
|
+
prompt.py,sha256=W8jSf7K6DPdTyK9HO06fbNocmoHm1nfkpQd8lfN1wrU,7216
|
|
7
|
+
summarizer.py,sha256=71nRxw6aRPf0aA5gg45hQoIUSBa-MMOiuMlACld4amU,3534
|
|
8
|
+
tokenizer.py,sha256=4VLJ0umfy7r9Yj3QxK55B_Xq-SYM_N10PFE0CwQBYlo,1296
|
|
9
|
+
ai_data_summarizer-0.7.1.dist-info/METADATA,sha256=TEEJ2hmAXzHPvj4_8mVUS8tpbWx0_PsBqOItLyNO7os,8101
|
|
10
|
+
ai_data_summarizer-0.7.1.dist-info/WHEEL,sha256=zOwg4jB6zX2kU910N-cMawjivD6tO8NEWvE12je1bVk,87
|
|
11
|
+
ai_data_summarizer-0.7.1.dist-info/entry_points.txt,sha256=q7AKQVFhwm3ZFqGTzj6Se0vOEXSxBApuMA4xParXFWA,38
|
|
12
|
+
ai_data_summarizer-0.7.1.dist-info/licenses/LICENSE,sha256=NRgWTyaRalxqDDAqDxput6odIMzsPSV4RkkQh2cCoN0,1072
|
|
13
|
+
ai_data_summarizer-0.7.1.dist-info/RECORD,,
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Louie Bloomberg
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
apicheck.py
ADDED
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
# API Key Checker
|
|
2
|
+
# apicheck.py
|
|
3
|
+
# This file contains functions to check for API keys and token counts for the different model providers.
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
import os
|
|
7
|
+
import sys
|
|
8
|
+
|
|
9
|
+
import anthropic
|
|
10
|
+
from google import genai
|
|
11
|
+
from openai import OpenAI
|
|
12
|
+
|
|
13
|
+
from envvar import *
|
|
14
|
+
|
|
15
|
+
def check_api_keys(provider_choice):
|
|
16
|
+
if provider_choice == 1:
|
|
17
|
+
check_for_gemini_api_key()
|
|
18
|
+
check_valid_gemini_api_key()
|
|
19
|
+
elif provider_choice == 2:
|
|
20
|
+
check_for_openai_api_key()
|
|
21
|
+
check_valid_openai_api_key()
|
|
22
|
+
elif provider_choice == 3:
|
|
23
|
+
check_for_anthropic_api_key()
|
|
24
|
+
check_valid_anthropic_api_key()
|
|
25
|
+
|
|
26
|
+
def check_for_gemini_api_key():
|
|
27
|
+
if os.getenv("GEMINI_API_KEY") is None or os.getenv("GEMINI_API_KEY") == "":
|
|
28
|
+
print("Error: Gemini API key not found. Would you like to set it now? (Y/n)")
|
|
29
|
+
if input().lower() != 'n':
|
|
30
|
+
set_api_keys(1)
|
|
31
|
+
else:
|
|
32
|
+
sys.exit(1)
|
|
33
|
+
else:
|
|
34
|
+
return True
|
|
35
|
+
|
|
36
|
+
def check_valid_gemini_api_key():
|
|
37
|
+
try:
|
|
38
|
+
client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
|
|
39
|
+
# Attempt to list models to check if the API key is valid
|
|
40
|
+
models = client.models.list()
|
|
41
|
+
return True
|
|
42
|
+
except Exception as e:
|
|
43
|
+
print(f"Error: Invalid Gemini API key. Please check your .env file. Details: {e}")
|
|
44
|
+
sys.exit(1)
|
|
45
|
+
|
|
46
|
+
|
|
47
|
+
def check_for_openai_api_key():
|
|
48
|
+
if os.getenv("OPENAI_API_KEY") is None or os.getenv("OPENAI_API_KEY") == "":
|
|
49
|
+
print("Error: OpenAI API key not found. Would you like to set it now? (Y/n)")
|
|
50
|
+
if input().lower() != 'n':
|
|
51
|
+
set_api_keys(2)
|
|
52
|
+
else:
|
|
53
|
+
sys.exit(1)
|
|
54
|
+
else:
|
|
55
|
+
return True
|
|
56
|
+
|
|
57
|
+
def check_valid_openai_api_key():
|
|
58
|
+
try:
|
|
59
|
+
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
|
|
60
|
+
# Attempt to list models to check if the API key is valid
|
|
61
|
+
# TODO: Implement correct method to validate OpenAI API key
|
|
62
|
+
return True
|
|
63
|
+
except Exception as e:
|
|
64
|
+
print(f"Error: Invalid OpenAI API key. Please check your .env file. Details: {e}")
|
|
65
|
+
sys.exit(1)
|
|
66
|
+
|
|
67
|
+
|
|
68
|
+
def check_for_anthropic_api_key():
|
|
69
|
+
if os.getenv("ANTHROPIC_API_KEY") is None or os.getenv("ANTHROPIC_API_KEY") == "":
|
|
70
|
+
print("Error: Anthropic API key not found. Would you like to set it now? (Y/n)")
|
|
71
|
+
if input().lower() != 'n':
|
|
72
|
+
set_api_keys(3)
|
|
73
|
+
else:
|
|
74
|
+
sys.exit(1)
|
|
75
|
+
else:
|
|
76
|
+
return True
|
|
77
|
+
|
|
78
|
+
def check_valid_anthropic_api_key():
|
|
79
|
+
try:
|
|
80
|
+
client = anthropic.Client(api_key=os.getenv("ANTHROPIC_API_KEY"))
|
|
81
|
+
# Attempt to list models to check if the API key is valid
|
|
82
|
+
# TODO: Implement correct method to validate Anthropic API key
|
|
83
|
+
return True
|
|
84
|
+
except Exception as e:
|
|
85
|
+
print(f"Error: Invalid Anthropic API key. Please check your .env file. Details: {e}")
|
|
86
|
+
sys.exit(1)
|
envvar.py
ADDED
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
# Environment Variable Management
|
|
2
|
+
# envvar.py
|
|
3
|
+
# This file contains functions to manage environment variables, including reading and writing to the .env file
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
import os
|
|
7
|
+
import pwinput
|
|
8
|
+
|
|
9
|
+
|
|
10
|
+
def _strip_quotes(value):
|
|
11
|
+
# Remove a single matching pair of surrounding quotes (single or double), if present.
|
|
12
|
+
value = value.strip()
|
|
13
|
+
if len(value) >= 2 and value[0] == value[-1] and value[0] in ('"', "'"):
|
|
14
|
+
return value[1:-1]
|
|
15
|
+
return value
|
|
16
|
+
|
|
17
|
+
|
|
18
|
+
def write_env_value(env_path, key_name, value):
|
|
19
|
+
# Update key_name's value in the .env file in place, preserving other lines.
|
|
20
|
+
# If key_name isn't found in the file, it gets appended.
|
|
21
|
+
lines = []
|
|
22
|
+
if os.path.exists(env_path):
|
|
23
|
+
with open(env_path, "r") as env_file:
|
|
24
|
+
lines = env_file.readlines()
|
|
25
|
+
|
|
26
|
+
found = False
|
|
27
|
+
for i, line in enumerate(lines):
|
|
28
|
+
stripped = line.strip()
|
|
29
|
+
if "=" in stripped and not stripped.startswith("#"):
|
|
30
|
+
existing_key = stripped.split("=", 1)[0].strip()
|
|
31
|
+
if existing_key == key_name:
|
|
32
|
+
lines[i] = f"{key_name}={value}\n"
|
|
33
|
+
found = True
|
|
34
|
+
break
|
|
35
|
+
|
|
36
|
+
if not found:
|
|
37
|
+
lines.append(f"{key_name}={value}\n")
|
|
38
|
+
|
|
39
|
+
with open(env_path, "w") as env_file:
|
|
40
|
+
env_file.writelines(lines)
|
|
41
|
+
|
|
42
|
+
|
|
43
|
+
def create_env():
|
|
44
|
+
# Create a .env file if it doesn't exist
|
|
45
|
+
base_dir = os.path.dirname(os.path.abspath(__file__))
|
|
46
|
+
env_path = os.path.join(base_dir, ".env")
|
|
47
|
+
# Check if the .env file exists
|
|
48
|
+
if not os.path.exists(env_path):
|
|
49
|
+
# have system create a .env file with default values
|
|
50
|
+
with open(env_path, "w") as env_file:
|
|
51
|
+
env_file.write("GEMINI_API_KEY=\n")
|
|
52
|
+
env_file.write("OPENAI_API_KEY=\n")
|
|
53
|
+
env_file.write("ANTHROPIC_API_KEY=\n")
|
|
54
|
+
env_file.write("GEMINI_FREE_TIER=True\n") # TODO: Remove Free Tier Flag and make that hardcoded
|
|
55
|
+
|
|
56
|
+
return None
|
|
57
|
+
|
|
58
|
+
def read_env():
|
|
59
|
+
# read full .env and set environment variables
|
|
60
|
+
base_dir = os.path.dirname(os.path.abspath(__file__))
|
|
61
|
+
env_path = os.path.join(base_dir, ".env")
|
|
62
|
+
if not os.path.exists(env_path):
|
|
63
|
+
create_env()
|
|
64
|
+
if os.path.exists(env_path):
|
|
65
|
+
with open(env_path, "r") as env_file:
|
|
66
|
+
for line in env_file:
|
|
67
|
+
if "=" in line:
|
|
68
|
+
key, value = line.strip().split("=", 1)
|
|
69
|
+
value = _strip_quotes(value)
|
|
70
|
+
os.environ[key.strip()] = value
|
|
71
|
+
return None
|
|
72
|
+
|
|
73
|
+
def set_api_keys(model_provider_choice):
|
|
74
|
+
base_dir = os.path.dirname(os.path.abspath(__file__))
|
|
75
|
+
env_path = os.path.join(base_dir, ".env")
|
|
76
|
+
|
|
77
|
+
# Select the appropriate API key based on the model provider choice
|
|
78
|
+
if model_provider_choice == 1:
|
|
79
|
+
api_key_name = "GEMINI_API_KEY"
|
|
80
|
+
elif model_provider_choice == 2:
|
|
81
|
+
api_key_name = "OPENAI_API_KEY"
|
|
82
|
+
elif model_provider_choice == 3:
|
|
83
|
+
api_key_name = "ANTHROPIC_API_KEY"
|
|
84
|
+
|
|
85
|
+
if os.getenv(api_key_name): # Check to see if environment variables are already in RAM
|
|
86
|
+
return True
|
|
87
|
+
|
|
88
|
+
read_env() # Read the .env file to set environment variables
|
|
89
|
+
|
|
90
|
+
# If environment variable is not in .env nor in RAM, prompt the user to enter the API key and save it to .env securely
|
|
91
|
+
if not os.getenv(api_key_name):
|
|
92
|
+
api_key = pwinput.pwinput(prompt=f"Enter your {api_key_name} (or leave blank to skip): ", mask="*").strip()
|
|
93
|
+
api_key = _strip_quotes(api_key) # Allow the user to paste a key wrapped in quotes
|
|
94
|
+
if api_key:
|
|
95
|
+
os.environ[api_key_name] = api_key
|
|
96
|
+
write_env_value(env_path, api_key_name, api_key)
|
|
97
|
+
print(f"{api_key_name} has been set and saved to {env_path}.")
|
|
98
|
+
else:
|
|
99
|
+
print(f"Warning: {api_key_name} is not set. You will not be able to use the selected model provider.")
|
fileloader.py
ADDED
|
@@ -0,0 +1,142 @@
|
|
|
1
|
+
# Dataset File Loader
|
|
2
|
+
# fileloader.py
|
|
3
|
+
# This module provides functions to load and read CSV files for processing.
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
import os
|
|
7
|
+
import sys
|
|
8
|
+
import pandas as pd
|
|
9
|
+
|
|
10
|
+
|
|
11
|
+
def detect_file_type(filepath):
|
|
12
|
+
"""
|
|
13
|
+
Determines the normalized file type ("csv" or "excel") based on the file extension.
|
|
14
|
+
Raises a ValueError if the extension is not supported.
|
|
15
|
+
|
|
16
|
+
Returns: str: "csv" or "excel"
|
|
17
|
+
"""
|
|
18
|
+
_, ext = os.path.splitext(filepath)
|
|
19
|
+
ext = ext.lower()
|
|
20
|
+
|
|
21
|
+
if ext == ".csv":
|
|
22
|
+
return "csv"
|
|
23
|
+
elif ext in [".xlsx", ".xls"]:
|
|
24
|
+
return "excel"
|
|
25
|
+
else:
|
|
26
|
+
print(f"Unsupported file extension '{ext}'. Supported extensions are: .csv, .xlsx, .xls")
|
|
27
|
+
sys.exit(1)
|
|
28
|
+
|
|
29
|
+
|
|
30
|
+
def get_sheet_names(filepath):
|
|
31
|
+
"""
|
|
32
|
+
Returns a list of sheet names if the file is an Excel workbook.
|
|
33
|
+
Returns None for CSV files, since they do not have sheets.
|
|
34
|
+
|
|
35
|
+
Returns: list[str] | None
|
|
36
|
+
"""
|
|
37
|
+
file_type = detect_file_type(filepath)
|
|
38
|
+
|
|
39
|
+
if file_type == "csv":
|
|
40
|
+
return None
|
|
41
|
+
|
|
42
|
+
try:
|
|
43
|
+
excel_file = pd.ExcelFile(filepath)
|
|
44
|
+
except Exception as e:
|
|
45
|
+
print(f"Could not read Excel file '{filepath}'. Details: {e}")
|
|
46
|
+
sys.exit(1)
|
|
47
|
+
|
|
48
|
+
return excel_file.sheet_names
|
|
49
|
+
|
|
50
|
+
def choose_sheet(filepath, file_type):
|
|
51
|
+
if file_type == "excel":
|
|
52
|
+
sheet_names = get_sheet_names(filepath)
|
|
53
|
+
if sheet_names is None:
|
|
54
|
+
print(f"Error: Could not retrieve sheet names from Excel file '{filepath}'.")
|
|
55
|
+
sys.exit(1)
|
|
56
|
+
|
|
57
|
+
print("\nAvailable sheets:")
|
|
58
|
+
for idx, sheet in enumerate(sheet_names):
|
|
59
|
+
print(f"{idx + 1}. {sheet}")
|
|
60
|
+
|
|
61
|
+
while True:
|
|
62
|
+
sheet_choice = input("Enter the number of the sheet you want to load: ").strip()
|
|
63
|
+
try:
|
|
64
|
+
sheet_index = int(sheet_choice) - 1
|
|
65
|
+
if 0 <= sheet_index < len(sheet_names):
|
|
66
|
+
selected_sheet = sheet_names[sheet_index]
|
|
67
|
+
break
|
|
68
|
+
else:
|
|
69
|
+
print("Invalid choice. Please enter a valid number.")
|
|
70
|
+
except Exception as e:
|
|
71
|
+
print(f"Invalid input. Please enter a number. Details: {e}")
|
|
72
|
+
sys.exit(1)
|
|
73
|
+
|
|
74
|
+
# Load the selected sheet as CSV text
|
|
75
|
+
try:
|
|
76
|
+
dataframe = pd.read_excel(filepath, sheet_name=selected_sheet)
|
|
77
|
+
csv_text = dataframe.to_csv(index=False)
|
|
78
|
+
return csv_text
|
|
79
|
+
except Exception as e:
|
|
80
|
+
print(f"Error: Could not read the selected sheet '{selected_sheet}' from Excel file '{filepath}'. Details: {e}")
|
|
81
|
+
sys.exit(1)
|
|
82
|
+
|
|
83
|
+
|
|
84
|
+
def load_csv(filepath):
|
|
85
|
+
"""
|
|
86
|
+
Loads a CSV file and returns it as normalized CSV-formatted text.
|
|
87
|
+
|
|
88
|
+
Returns: str
|
|
89
|
+
"""
|
|
90
|
+
try:
|
|
91
|
+
dataframe = pd.read_csv(filepath, thousands=',', decimal='.')
|
|
92
|
+
except Exception as e:
|
|
93
|
+
print(f"Could not read CSV file '{filepath}'. Details: {e}")
|
|
94
|
+
sys.exit(1)
|
|
95
|
+
|
|
96
|
+
if dataframe.empty:
|
|
97
|
+
print(f"CSV file '{filepath}' appears to be empty.")
|
|
98
|
+
sys.exit(1)
|
|
99
|
+
|
|
100
|
+
return dataframe
|
|
101
|
+
|
|
102
|
+
|
|
103
|
+
def load_excel(filepath, sheet_name=None):
|
|
104
|
+
"""
|
|
105
|
+
Loads a single sheet from an Excel file (.xlsx or .xls) and returns it as
|
|
106
|
+
normalized CSV-formatted text. If sheet_name is not provided, the first
|
|
107
|
+
sheet in the workbook is used.
|
|
108
|
+
|
|
109
|
+
Returns: str
|
|
110
|
+
"""
|
|
111
|
+
try:
|
|
112
|
+
if sheet_name is None:
|
|
113
|
+
available_sheets = get_sheet_names(filepath)
|
|
114
|
+
sheet_name = available_sheets[0]
|
|
115
|
+
|
|
116
|
+
dataframe = pd.read_excel(filepath, sheet_name=sheet_name)
|
|
117
|
+
except Exception as e:
|
|
118
|
+
print(f"Could not read Excel file '{filepath}'. Details: {e}")
|
|
119
|
+
sys.exit(1)
|
|
120
|
+
|
|
121
|
+
if dataframe.empty:
|
|
122
|
+
print(f"Sheet '{sheet_name}' in file '{filepath}' appears to be empty.")
|
|
123
|
+
sys.exit(1)
|
|
124
|
+
|
|
125
|
+
return dataframe
|
|
126
|
+
|
|
127
|
+
|
|
128
|
+
def load_file(filepath, file_type):
|
|
129
|
+
"""
|
|
130
|
+
Dispatches to the correct loader based on the file's detected type, and
|
|
131
|
+
returns normalized CSV-formatted text regardless of the original source format.
|
|
132
|
+
|
|
133
|
+
sheet_name is only relevant for Excel files. It is ignored for CSV files.
|
|
134
|
+
|
|
135
|
+
Returns: str
|
|
136
|
+
"""
|
|
137
|
+
|
|
138
|
+
if file_type == "csv":
|
|
139
|
+
return load_csv(filepath)
|
|
140
|
+
elif file_type == "excel":
|
|
141
|
+
return load_excel(filepath)
|
|
142
|
+
return None
|
main.py
ADDED
|
@@ -0,0 +1,155 @@
|
|
|
1
|
+
# Main Code File
|
|
2
|
+
# main.py
|
|
3
|
+
# Code file containing the main function and the main logic for the AI Data Summarizer program.
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
# Import Statements
|
|
7
|
+
import os
|
|
8
|
+
import sys
|
|
9
|
+
import datetime
|
|
10
|
+
import csv
|
|
11
|
+
from readchar import readkey, key
|
|
12
|
+
|
|
13
|
+
# Import the functions from the summarizer.py helper file
|
|
14
|
+
from summarizer import *
|
|
15
|
+
from tokenizer import *
|
|
16
|
+
from prompt import *
|
|
17
|
+
from apicheck import *
|
|
18
|
+
from envvar import *
|
|
19
|
+
from fileloader import *
|
|
20
|
+
from profiler import *
|
|
21
|
+
|
|
22
|
+
# Function Definitions
|
|
23
|
+
|
|
24
|
+
def check_tokens_gemini(prompt, text):
|
|
25
|
+
# This function checks the number of tokens to make sure that they are within the Free Tier limits
|
|
26
|
+
google_tokens = google_tokenizer(prompt, text)
|
|
27
|
+
if os.getenv("GEMINI_FREE_TIER") == "True" and google_tokens > 250000:
|
|
28
|
+
print(f"Warning: Your input text has {google_tokens} tokens, which exceeds the free tier limit of 250,000 tokens per minute for Gemini. Consider reducing the input size or upgrading your plan. ")
|
|
29
|
+
sys.exit(1)
|
|
30
|
+
elif os.getenv("GEMINI_FREE_TIER") == "True":
|
|
31
|
+
print(f"Your input text has {google_tokens} tokens, which is within the free tier limit for Gemini.")
|
|
32
|
+
print("Would you like to continue? (Y/n)")
|
|
33
|
+
while True:
|
|
34
|
+
k = readkey()
|
|
35
|
+
if k == 'y' or k == 'Y' or k == key.ENTER:
|
|
36
|
+
return None
|
|
37
|
+
elif k == 'n' or k == 'N':
|
|
38
|
+
print("Exiting...")
|
|
39
|
+
sys.exit(1)
|
|
40
|
+
else:
|
|
41
|
+
print("Invalid input. Please enter y or n.")
|
|
42
|
+
|
|
43
|
+
# TODO (maybe): Add cost functionality to estimate cost of input response
|
|
44
|
+
return None
|
|
45
|
+
|
|
46
|
+
def check_tokens_openai(prompt, text):
|
|
47
|
+
tokens = openai_tokenizer(prompt, text)
|
|
48
|
+
print(f"Your input text has {tokens} tokens.")
|
|
49
|
+
return None
|
|
50
|
+
|
|
51
|
+
def check_tokens_anthropic(prompt, text):
|
|
52
|
+
tokens = anthropic_tokenizer(prompt, text)
|
|
53
|
+
print(f"Your input text has {tokens} tokens.")
|
|
54
|
+
return None
|
|
55
|
+
|
|
56
|
+
def get_model_provider():
|
|
57
|
+
while True:
|
|
58
|
+
print("\nSelect a model provider:")
|
|
59
|
+
print("1. Google Gemini")
|
|
60
|
+
print("2. OpenAI")
|
|
61
|
+
print("3. Anthropic")
|
|
62
|
+
print("Or press q to quit the program.")
|
|
63
|
+
k = readkey()
|
|
64
|
+
|
|
65
|
+
if k == '1':
|
|
66
|
+
return 1 # Gemini
|
|
67
|
+
elif k == '2':
|
|
68
|
+
return 2 # OpenAI
|
|
69
|
+
elif k == '3':
|
|
70
|
+
return 3 # Anthropic
|
|
71
|
+
elif k == 'q':
|
|
72
|
+
print("Exiting...")
|
|
73
|
+
sys.exit(1)
|
|
74
|
+
else:
|
|
75
|
+
print("Invalid choice. Please enter 1, 2, or 3.")
|
|
76
|
+
|
|
77
|
+
# Define Prompt Choosing Function
|
|
78
|
+
def get_prompt_type():
|
|
79
|
+
"""
|
|
80
|
+
This function allows the user to select a prompt from a list of available prompts.
|
|
81
|
+
It returns the selected prompt as a string.
|
|
82
|
+
Returns: str: The selected prompt.
|
|
83
|
+
"""
|
|
84
|
+
while True:
|
|
85
|
+
print("\nSelect a Summary:")
|
|
86
|
+
print("1. Just the Facts (AI is not used)")
|
|
87
|
+
print("2. TL;DR Summary")
|
|
88
|
+
print("3. Data Overview")
|
|
89
|
+
print("4. Deep Dive Analysis")
|
|
90
|
+
print("Or press q to quit the program.")
|
|
91
|
+
k = readkey()
|
|
92
|
+
|
|
93
|
+
if k == '2':
|
|
94
|
+
return "tldr"
|
|
95
|
+
elif k == '3':
|
|
96
|
+
return "overview"
|
|
97
|
+
elif k == '4':
|
|
98
|
+
return "deepdive"
|
|
99
|
+
elif k == '1':
|
|
100
|
+
return "facts"
|
|
101
|
+
elif k == 'q':
|
|
102
|
+
print("Exiting...")
|
|
103
|
+
sys.exit(1)
|
|
104
|
+
else:
|
|
105
|
+
print("Invalid choice. Please enter 1, 2, 3, 4, or q.")
|
|
106
|
+
|
|
107
|
+
# Main Function
|
|
108
|
+
def main():
|
|
109
|
+
read_env() # Read the .env file to set environment variables
|
|
110
|
+
# Main Interactive Loop:
|
|
111
|
+
# Get input file
|
|
112
|
+
while True:
|
|
113
|
+
filepath = input("\nEnter the path to your CSV file: ").strip()
|
|
114
|
+
|
|
115
|
+
# Remove quotes if user wrapped path in quotes
|
|
116
|
+
input_file = filepath.strip('"\'')
|
|
117
|
+
|
|
118
|
+
|
|
119
|
+
if not os.path.exists(input_file):
|
|
120
|
+
print(f"Error: File '{input_file}' not found. Please try again.")
|
|
121
|
+
continue
|
|
122
|
+
|
|
123
|
+
if not input_file.lower().endswith('.csv') and not input_file.lower().endswith('.xlsx') and not input_file.lower().endswith('.xls'):
|
|
124
|
+
print("Warning: file extension is not .csv, .xlsx, or .xls. Continue anyway? (y/N)")
|
|
125
|
+
if input().lower() != 'y':
|
|
126
|
+
continue
|
|
127
|
+
|
|
128
|
+
break
|
|
129
|
+
|
|
130
|
+
file_type = detect_file_type(input_file)
|
|
131
|
+
|
|
132
|
+
df = load_file(input_file, file_type)
|
|
133
|
+
profile = build_base_profile(df)
|
|
134
|
+
sample = get_sample(df)
|
|
135
|
+
csv_text = "\nData Profile:" + profile + "\nData Sample:\n" + sample
|
|
136
|
+
|
|
137
|
+
model_provider_choice = 1 # get_model_provider()
|
|
138
|
+
# Defaults to Gemini since it is the only one provided
|
|
139
|
+
check_api_keys(model_provider_choice)
|
|
140
|
+
prompt_type = get_prompt_type()
|
|
141
|
+
if prompt_type == "facts":
|
|
142
|
+
summary = "Dataset Statistical Summary:\n" + profile + "\nData Sample:\n" + sample
|
|
143
|
+
write_outfile(summary, os.path.basename(input_file).split(".")[0], prompt_type, "None")
|
|
144
|
+
else:
|
|
145
|
+
prompt = get_prompt(prompt_type)
|
|
146
|
+
check_tokens_gemini(prompt, csv_text)
|
|
147
|
+
summary = gemini_summarizer(prompt, csv_text, os.path.basename(input_file).split(".")[0], prompt_type)
|
|
148
|
+
|
|
149
|
+
print("\nSummary:\n")
|
|
150
|
+
print(summary)
|
|
151
|
+
|
|
152
|
+
return None
|
|
153
|
+
|
|
154
|
+
if __name__ == "__main__":
|
|
155
|
+
main()
|
profiler.py
ADDED
|
@@ -0,0 +1,108 @@
|
|
|
1
|
+
# Pandas Pre-Summarization Profiler
|
|
2
|
+
# profiler.py
|
|
3
|
+
# This module provides functions to profile a DataFrame before summarization.
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
import pandas as pd
|
|
7
|
+
|
|
8
|
+
|
|
9
|
+
def _parse_dates(df):
|
|
10
|
+
# Returns a copy of df with any date-like object columns converted to datetime
|
|
11
|
+
df = df.copy()
|
|
12
|
+
for col in df.select_dtypes(include="object").columns:
|
|
13
|
+
try:
|
|
14
|
+
converted = pd.to_datetime(df[col], format="mixed", errors="coerce")
|
|
15
|
+
if converted.notna().sum() > len(df) * 0.9: # 90%+ parsed successfully
|
|
16
|
+
df[col] = converted
|
|
17
|
+
except Exception:
|
|
18
|
+
continue
|
|
19
|
+
return df
|
|
20
|
+
|
|
21
|
+
|
|
22
|
+
def build_base_profile(df):
|
|
23
|
+
"""
|
|
24
|
+
Returns a plain-text statistical profile of the DataFrame.
|
|
25
|
+
Includes shape, dtypes, null counts, numeric stats, and top categorical values.
|
|
26
|
+
Does not mutate the input DataFrame.
|
|
27
|
+
|
|
28
|
+
Returns: str
|
|
29
|
+
"""
|
|
30
|
+
df = _parse_dates(df)
|
|
31
|
+
lines = []
|
|
32
|
+
|
|
33
|
+
# Shape
|
|
34
|
+
num_rows, num_cols = df.shape
|
|
35
|
+
lines.append(f"Rows: {num_rows}")
|
|
36
|
+
lines.append(f"Columns: {num_cols}")
|
|
37
|
+
lines.append("")
|
|
38
|
+
|
|
39
|
+
# Dtypes and null counts
|
|
40
|
+
lines.append("Column Overview:")
|
|
41
|
+
for col in df.columns:
|
|
42
|
+
null_count = df[col].isna().sum()
|
|
43
|
+
lines.append(f" {col} ({df[col].dtype}) — {null_count} nulls")
|
|
44
|
+
lines.append("")
|
|
45
|
+
|
|
46
|
+
# Numeric stats
|
|
47
|
+
numeric_cols = df.select_dtypes(include="number").columns
|
|
48
|
+
if len(numeric_cols) > 0:
|
|
49
|
+
lines.append("Numeric Column Stats:")
|
|
50
|
+
lines.append(df[numeric_cols].describe().to_string())
|
|
51
|
+
lines.append("")
|
|
52
|
+
|
|
53
|
+
# Date column summary
|
|
54
|
+
date_cols = df.select_dtypes(include=["datetime"]).columns
|
|
55
|
+
if len(date_cols) > 0:
|
|
56
|
+
lines.append("Date Column Summary:")
|
|
57
|
+
for col in date_cols:
|
|
58
|
+
min_date = df[col].min()
|
|
59
|
+
max_date = df[col].max()
|
|
60
|
+
if pd.isna(min_date) or pd.isna(max_date):
|
|
61
|
+
lines.append(f" {col}: (no non-null datetimes)")
|
|
62
|
+
continue
|
|
63
|
+
span = (max_date - min_date).days + 1
|
|
64
|
+
duplicate_dates = df[col].duplicated().sum()
|
|
65
|
+
lines.append(f" {col}: {min_date.date()} to {max_date.date()} ({span} days)")
|
|
66
|
+
if duplicate_dates > 0:
|
|
67
|
+
lines.append(f" {col} has {duplicate_dates} duplicate date(s)")
|
|
68
|
+
lines.append("")
|
|
69
|
+
|
|
70
|
+
# Standout values with date context
|
|
71
|
+
if len(date_cols) > 0 and len(numeric_cols) > 0:
|
|
72
|
+
date_col = date_cols[0]
|
|
73
|
+
show_all_cols = len(df.columns) <= 5
|
|
74
|
+
lines.append("Standout Values:")
|
|
75
|
+
for col in numeric_cols:
|
|
76
|
+
if show_all_cols:
|
|
77
|
+
top5 = df.nlargest(5, col).to_string(index=False)
|
|
78
|
+
bottom5 = df.nsmallest(5, col).to_string(index=False)
|
|
79
|
+
else:
|
|
80
|
+
top5 = df[[date_col, col]].nlargest(5, col).to_string(index=False)
|
|
81
|
+
bottom5 = df[[date_col, col]].nsmallest(5, col).to_string(index=False)
|
|
82
|
+
lines.append(f" Top 5 highest {col}:")
|
|
83
|
+
lines.append(top5)
|
|
84
|
+
lines.append(f" Bottom 5 lowest {col}:")
|
|
85
|
+
lines.append(bottom5)
|
|
86
|
+
lines.append("")
|
|
87
|
+
|
|
88
|
+
# Categorical top values
|
|
89
|
+
categorical_cols = df.select_dtypes(include=["object", "category"]).columns
|
|
90
|
+
if len(categorical_cols) > 0:
|
|
91
|
+
lines.append("Categorical Column Summary:")
|
|
92
|
+
for col in categorical_cols:
|
|
93
|
+
unique_count = df[col].nunique()
|
|
94
|
+
top_values = df[col].value_counts().head(5).to_string()
|
|
95
|
+
lines.append(f" {col} ({unique_count} unique values):")
|
|
96
|
+
lines.append(f"{top_values}")
|
|
97
|
+
lines.append("")
|
|
98
|
+
|
|
99
|
+
return "\n".join(lines)
|
|
100
|
+
|
|
101
|
+
|
|
102
|
+
def get_sample(df, n=10):
|
|
103
|
+
"""
|
|
104
|
+
Returns the first n rows of the DataFrame as CSV-formatted text.
|
|
105
|
+
|
|
106
|
+
Returns: str
|
|
107
|
+
"""
|
|
108
|
+
return df.head(n).to_csv(index=False)
|
prompt.py
ADDED
|
@@ -0,0 +1,144 @@
|
|
|
1
|
+
# Main Prompt File
|
|
2
|
+
# prompt.py
|
|
3
|
+
# File contains different prompts which can be used to help the user understand the file
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
# Define Prompt Choosing Functions
|
|
7
|
+
def get_prompt(prompt_type):
|
|
8
|
+
"""
|
|
9
|
+
This function retrieves the appropriate prompt based on the user's selection.
|
|
10
|
+
Returns: str: The selected prompt.
|
|
11
|
+
"""
|
|
12
|
+
if prompt_type == "tldr":
|
|
13
|
+
return get_tldr_prompt()
|
|
14
|
+
elif prompt_type == "overview":
|
|
15
|
+
return get_overview_prompt()
|
|
16
|
+
elif prompt_type == "deepdive":
|
|
17
|
+
return get_deepdive_prompt()
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
# Define Prompts that can be used
|
|
21
|
+
|
|
22
|
+
def get_tldr_prompt():
|
|
23
|
+
task_summary = f"""
|
|
24
|
+
## Task Summary:
|
|
25
|
+
{{Produce a 'Too Long; Didn’t Read' (TL;DR) summary of the attached dataset and statistical summary. The TL;DR sentence should describe, in one high-level line, what the dataset is and what it covers. Then provide 2–3 bullets highlighting the most notable patterns, trends, or anomalies visible in the data. Use numeric ranges when helpful, but keep the focus on the biggest takeaways.}}
|
|
26
|
+
"""
|
|
27
|
+
|
|
28
|
+
response_style = f"""
|
|
29
|
+
## Response style and format requirements:
|
|
30
|
+
- {{Write in a sharp, punchy, and bottom-line-up-front (BLUF) style}}
|
|
31
|
+
- {{Format: Start with a single 'TL;DR:' sentence, followed by a short bulleted list}}
|
|
32
|
+
- {{Bullets must follow these rules:
|
|
33
|
+
- Start with a short, strong label (e.g., 'Trend shift:', 'Category contrast:', 'Peak anomaly:')
|
|
34
|
+
- Contain exactly one idea per bullet
|
|
35
|
+
- Use numeric anchors when possible
|
|
36
|
+
- Avoid hedging, filler, or multi-clause sentences
|
|
37
|
+
- Read like headlines, not explanations}}
|
|
38
|
+
- {{Numbers in the stats block are exact — use them directly, don't estimate from sample rows}}
|
|
39
|
+
- {{Strictly limit the response to 100 words or less}}
|
|
40
|
+
- {{Use plain text only — no Markdown. Bullets are to be noted with '-'}}
|
|
41
|
+
"""
|
|
42
|
+
|
|
43
|
+
final_prompt = f"""{task_summary}
|
|
44
|
+
{response_style}"""
|
|
45
|
+
|
|
46
|
+
return final_prompt
|
|
47
|
+
|
|
48
|
+
|
|
49
|
+
def get_overview_prompt():
|
|
50
|
+
# Prompt for the Summarizer:
|
|
51
|
+
# Use this to clearly define the task and job needed by the model
|
|
52
|
+
task_summary = f"""
|
|
53
|
+
## Task Summary:
|
|
54
|
+
{{Review the attached CSV and statistical summary and summarize what the data covers, including anything notable or unusual.}}
|
|
55
|
+
"""
|
|
56
|
+
|
|
57
|
+
# Use this to provide contextual information related to the task
|
|
58
|
+
context_information = f"""
|
|
59
|
+
## Context Information:
|
|
60
|
+
- {{Standard CSV format with headers in the first row}}
|
|
61
|
+
- {{Columns may include numbers, text, or dates}}
|
|
62
|
+
- {{Treat all dates in the data file as a recorded value and not predictions}}
|
|
63
|
+
- {{The dataset may cover any domain — do not assume a specific subject area}}
|
|
64
|
+
- {{The stats block (counts, mean/median, ranges, date span, category counts) is precomputed and exact — treat these numbers as ground truth}}
|
|
65
|
+
- {{Standout/extreme rows shown are the most unusual in the dataset, not typical examples}}
|
|
66
|
+
- {{Any raw sample rows are shown only to illustrate formatting and column meaning, not to infer statistics}}
|
|
67
|
+
"""
|
|
68
|
+
|
|
69
|
+
# Use this to provide any model instructions that you want model to adhere to
|
|
70
|
+
model_instructions = f"""
|
|
71
|
+
## Model Instructions:
|
|
72
|
+
- {{Explain what each column represents and flag anything out of the ordinary}}
|
|
73
|
+
- {{Base all observations only on the data provided}}
|
|
74
|
+
- {{Only use historical context or external knowledge where it clearly and directly explains a specific data pattern — do not force connections}}
|
|
75
|
+
- {{If no external context is relevant, rely entirely on what the data shows}}
|
|
76
|
+
"""
|
|
77
|
+
|
|
78
|
+
# Use this to provide response style and formatting guidance
|
|
79
|
+
response_style = f"""
|
|
80
|
+
## Response style and format requirements:
|
|
81
|
+
- {{Write a standalone written summary as if briefing a coworker}}
|
|
82
|
+
- {{Use three sections: overview, column breakdown, and key takeaways}}
|
|
83
|
+
- {{Limit the response to 500 words or less}}
|
|
84
|
+
- {{Use plain text only — no Markdown. Bullets are to be noted with '-'. No Markdown headings.}}
|
|
85
|
+
"""
|
|
86
|
+
|
|
87
|
+
# Concatenate to final prompt
|
|
88
|
+
final_prompt = f"""{task_summary}
|
|
89
|
+
{context_information}
|
|
90
|
+
{model_instructions}
|
|
91
|
+
{response_style}"""
|
|
92
|
+
|
|
93
|
+
return final_prompt
|
|
94
|
+
|
|
95
|
+
|
|
96
|
+
def get_deepdive_prompt():
|
|
97
|
+
task_summary = f"""
|
|
98
|
+
## Task Summary:
|
|
99
|
+
{{Perform a detailed deep dive analysis of the attached CSV dataset and statistical summary. Go beyond surface-level description to examine distributions, patterns, relationships, and notable characteristics in the data.}}
|
|
100
|
+
"""
|
|
101
|
+
|
|
102
|
+
context_information = f"""
|
|
103
|
+
## Context Information:
|
|
104
|
+
- {{Standard CSV format with headers in the first row}}
|
|
105
|
+
- {{Columns may include numbers, text, dates, or categorical values}}
|
|
106
|
+
- {{Treat all dates as recorded historical values}}
|
|
107
|
+
- {{The dataset may cover any domain — analyze based solely on what is present}}
|
|
108
|
+
"""
|
|
109
|
+
|
|
110
|
+
input_composition = f"""
|
|
111
|
+
## Input Composition:
|
|
112
|
+
- {{The stats block (counts, mean, median, std, min/max, date span, category counts) is precomputed directly from the full dataset and is exact — treat it as ground truth, never recompute or estimate these figures yourself}}
|
|
113
|
+
- {{The standout/extreme rows (top and bottom values) represent the most unusual points in the entire dataset, not a representative sample — use them only to discuss outliers, spikes, or anomalies}}
|
|
114
|
+
- {{The raw sample rows are a small illustrative slice included only to show formatting, column meaning, and qualitative texture — they are not necessarily representative of the full distribution and should not be used to infer statistics}}
|
|
115
|
+
"""
|
|
116
|
+
|
|
117
|
+
model_instructions = f"""
|
|
118
|
+
## Model Instructions:
|
|
119
|
+
- {{For numeric columns: report range (min-max), central tendency (mean/median if relevant), and distribution shape}}
|
|
120
|
+
- {{For numeric columns: also describe distribution shape based on the relationship between mean, median, and standard deviation}}
|
|
121
|
+
- {{For categorical/text columns: list top unique values with counts and note any dominant categories}}
|
|
122
|
+
- {{Identify any clear relationships or correlations between columns that stand out}}
|
|
123
|
+
- {{Highlight temporal patterns if dates are present, or geographic patterns if location data exists}}
|
|
124
|
+
- {{Flag outliers, unusual spikes/drops, or data quality concerns}}
|
|
125
|
+
- {{Base every observation strictly on the data provided — do not speculate beyond visible evidence}}
|
|
126
|
+
- {{Suggest 2–3 specific follow-up questions or analyses the user could explore next}}
|
|
127
|
+
"""
|
|
128
|
+
|
|
129
|
+
response_style = f"""
|
|
130
|
+
## Response style and format requirements:
|
|
131
|
+
- {{Write as if explaining the data in detail to a data-savvy coworker}}
|
|
132
|
+
- {{Use these four clear sections in order: Overview, Column Analysis, Key Patterns & Relationships, Takeaways & Next Steps}}
|
|
133
|
+
- {{Use plain text only — no Markdown. Bullets are to be noted with '-'. No Markdown headings in the final output.}}
|
|
134
|
+
- {{Keep the total response under 800 words}}
|
|
135
|
+
- {{Be specific and quantitative where possible}}
|
|
136
|
+
"""
|
|
137
|
+
|
|
138
|
+
final_prompt = f"""{task_summary}
|
|
139
|
+
{context_information}
|
|
140
|
+
{input_composition}
|
|
141
|
+
{model_instructions}
|
|
142
|
+
{response_style}"""
|
|
143
|
+
|
|
144
|
+
return final_prompt
|
summarizer.py
ADDED
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# Dataset Summarizer
|
|
2
|
+
# summarizer.py
|
|
3
|
+
# Code File for calling the summarization functions -- Mainly a helper file to keep code organized
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
# Import Statements
|
|
7
|
+
import os
|
|
8
|
+
import sys
|
|
9
|
+
from zoneinfo import ZoneInfo
|
|
10
|
+
import datetime as dt
|
|
11
|
+
import csv
|
|
12
|
+
|
|
13
|
+
# Importing AI Libraries
|
|
14
|
+
from openai import OpenAI
|
|
15
|
+
import anthropic
|
|
16
|
+
from google import genai
|
|
17
|
+
from google.genai import types
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
# Function Definitions
|
|
21
|
+
# TODO: Implement summarization functions here
|
|
22
|
+
|
|
23
|
+
# Write File Function
|
|
24
|
+
def write_outfile(output, filename, prompt_type, model_id):
|
|
25
|
+
# Create summaries directory if it doesn't exist
|
|
26
|
+
summaries_dir = f"./summaries/{prompt_type}"
|
|
27
|
+
os.makedirs(summaries_dir, exist_ok=True)
|
|
28
|
+
|
|
29
|
+
now = dt.datetime.now()
|
|
30
|
+
date_time_string = now.strftime("%Y-%m-%d %H-%M-%S")
|
|
31
|
+
central_tz = ZoneInfo("America/Chicago")
|
|
32
|
+
date_written = dt.datetime.now(central_tz).strftime("%A, %B %d, %Y")
|
|
33
|
+
time_written = dt.datetime.now(central_tz).strftime("%I:%M %p %Z")
|
|
34
|
+
summary_local_path = f"{summaries_dir}/{date_time_string}.{filename}.{model_id}.txt"
|
|
35
|
+
with open(summary_local_path, "w", encoding="utf-8") as outfile:
|
|
36
|
+
# Add Header Section to the output file
|
|
37
|
+
outfile.write("=" * 10 + "BEGIN HEADER" + "=" * 10 + "\n")
|
|
38
|
+
outfile.write("Date Written: " + date_written + "\n")
|
|
39
|
+
outfile.write("Time Written: " + time_written + "\n")
|
|
40
|
+
outfile.write("Model Used: " + model_id + "\n")
|
|
41
|
+
outfile.write("Prompt Type: " + prompt_type + "\n")
|
|
42
|
+
outfile.write("File Name: " + filename + "\n")
|
|
43
|
+
outfile.write("=" * 10 + "END HEADER" + "=" * 10 + "\n\n")
|
|
44
|
+
|
|
45
|
+
# Add main summary output to the output file
|
|
46
|
+
outfile.write(output)
|
|
47
|
+
|
|
48
|
+
return None
|
|
49
|
+
|
|
50
|
+
# Gemini Summarizer
|
|
51
|
+
def gemini_summarizer(prompt, text, filename, prompt_type):
|
|
52
|
+
MODEL_ID = "gemini-3.5-flash"
|
|
53
|
+
try:
|
|
54
|
+
client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
|
|
55
|
+
|
|
56
|
+
response = client.models.generate_content(
|
|
57
|
+
model=MODEL_ID,
|
|
58
|
+
contents=[
|
|
59
|
+
text,
|
|
60
|
+
prompt,
|
|
61
|
+
]
|
|
62
|
+
)
|
|
63
|
+
output = response.text
|
|
64
|
+
except Exception as e:
|
|
65
|
+
print(f"Error generating summary: {e}")
|
|
66
|
+
sys.exit(1)
|
|
67
|
+
|
|
68
|
+
write_outfile(output, filename, prompt_type, MODEL_ID)
|
|
69
|
+
return output
|
|
70
|
+
|
|
71
|
+
|
|
72
|
+
def openai_summarizer(prompt, text, filename, prompt_type):
|
|
73
|
+
MODEL_ID = "gpt-5.4-mini"
|
|
74
|
+
try:
|
|
75
|
+
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
|
|
76
|
+
|
|
77
|
+
response = client.responses.create(
|
|
78
|
+
model=MODEL_ID,
|
|
79
|
+
messages=[
|
|
80
|
+
{"role": "system", "content": "You are a helpful assistant."},
|
|
81
|
+
{"role": "user", "content": prompt + "\n\n" + text},
|
|
82
|
+
],
|
|
83
|
+
)
|
|
84
|
+
output = response.output_text
|
|
85
|
+
except Exception as e:
|
|
86
|
+
print(f"Error generating summary: {e}")
|
|
87
|
+
sys.exit(1)
|
|
88
|
+
|
|
89
|
+
write_outfile(output, filename, prompt_type, MODEL_ID)
|
|
90
|
+
return output
|
|
91
|
+
|
|
92
|
+
def anthropic_summarizer(prompt, text, filename, prompt_type):
|
|
93
|
+
MODEL_ID = "claude-haiku-4-5"
|
|
94
|
+
try:
|
|
95
|
+
client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
|
|
96
|
+
|
|
97
|
+
message = client.messages.create(
|
|
98
|
+
model=MODEL_ID,
|
|
99
|
+
messages=[
|
|
100
|
+
{"role": "user", "content": prompt + "\n\n" + text},
|
|
101
|
+
],
|
|
102
|
+
)
|
|
103
|
+
output = message.content[0].text
|
|
104
|
+
except Exception as e:
|
|
105
|
+
print(f"Error generating summary: {e}")
|
|
106
|
+
sys.exit(1)
|
|
107
|
+
|
|
108
|
+
write_outfile(output, filename, prompt_type, MODEL_ID)
|
|
109
|
+
return output
|
tokenizer.py
ADDED
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# Dataset Tokenizer
|
|
2
|
+
# tokenizer.py
|
|
3
|
+
# Code File for calling the tokenization functions -- Mainly a helper file to keep code organized
|
|
4
|
+
# SPDX-License-Identifier: MIT
|
|
5
|
+
|
|
6
|
+
import os
|
|
7
|
+
import sys
|
|
8
|
+
|
|
9
|
+
from openai import OpenAI
|
|
10
|
+
import anthropic
|
|
11
|
+
from google import genai
|
|
12
|
+
from google.genai import local_tokenizer
|
|
13
|
+
|
|
14
|
+
# Tokenization Functions
|
|
15
|
+
# Google GenAI Tokenizer
|
|
16
|
+
def google_tokenizer(prompt, text):
|
|
17
|
+
MODEL_ID = "gemini-3.5-flash"
|
|
18
|
+
client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))
|
|
19
|
+
|
|
20
|
+
response = client.models.count_tokens(
|
|
21
|
+
model=MODEL_ID,
|
|
22
|
+
contents=[
|
|
23
|
+
text,
|
|
24
|
+
prompt,
|
|
25
|
+
]
|
|
26
|
+
)
|
|
27
|
+
tokens = response.total_tokens
|
|
28
|
+
return tokens
|
|
29
|
+
|
|
30
|
+
# OpenAI Tokenizer
|
|
31
|
+
def openai_tokenizer(prompt, text):
|
|
32
|
+
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))
|
|
33
|
+
response = client.responses.input_tokens.count(
|
|
34
|
+
model="gpt-5.4-mini",
|
|
35
|
+
instructions=prompt,
|
|
36
|
+
input=text,
|
|
37
|
+
)
|
|
38
|
+
return response.input_tokens
|
|
39
|
+
|
|
40
|
+
# Anthropic Tokenizer
|
|
41
|
+
def anthropic_tokenizer(prompt, text):
|
|
42
|
+
client = anthropic.Anthropic(api_key=os.getenv("ANTHROPIC_API_KEY"))
|
|
43
|
+
response = client.messages.count_tokens(
|
|
44
|
+
model="claude-haiku-4-5",
|
|
45
|
+
system=prompt,
|
|
46
|
+
messages=[{"role": "user", "content": text}],
|
|
47
|
+
)
|
|
48
|
+
return response.get("input_tokens", 0)
|