maltopic 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
maltopic-1.0.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2025 Yash Sharma
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,100 @@
1
+ Metadata-Version: 2.3
2
+ Name: maltopic
3
+ Version: 1.0.0
4
+ Summary: A multi-agent LLM topic modeling library.
5
+ License: MIT
6
+ Keywords: topic modeling,LLM,multi-agent,survey analysis,data enrichment
7
+ Author: Yash Sharma
8
+ Author-email: yash91sharma@gmail.com
9
+ Requires-Python: >=3.12,<4.0
10
+ Classifier: License :: OSI Approved :: MIT License
11
+ Classifier: Operating System :: OS Independent
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.12
14
+ Classifier: Programming Language :: Python :: 3.13
15
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
16
+ Requires-Dist: openai (>=1.79.0,<2.0.0)
17
+ Requires-Dist: pandas (>=2.2.3,<3.0.0)
18
+ Requires-Dist: tqdm (>=4.67.1,<5.0.0)
19
+ Project-URL: Homepage, https://github.com/yash91sharma/MALTopic-py
20
+ Project-URL: Repository, https://github.com/yash91sharma/MALTopic-py
21
+ Description-Content-Type: text/markdown
22
+
23
+ # MALTopic: Multi-Agent LLM Topic Modeling Library
24
+
25
+ MALTopic is a powerful library designed for topic modeling using a multi-agent approach. It leverages the capabilities of large language models (LLMs) to enhance the analysis of survey responses by integrating structured and unstructured data.
26
+
27
+ MALTopic as a research paper was published in 2025 World AI IoT Congress. Links here.
28
+
29
+ ## Features
30
+
31
+ - **Multi-Agent Framework**: Decomposes topic modeling into specialized tasks executed by individual LLM agents.
32
+ - **Data Enrichment**: Enhances textual responses using structured and categorical survey data.
33
+ - **Latent Theme Extraction**: Extracts meaningful topics from enriched responses.
34
+ - **Topic Deduplication**: Refines and consolidates identified topics for better interpretability.
35
+
36
+ ## Installation
37
+
38
+ To install the MALTopic library, you can use pip:
39
+
40
+ ```bash
41
+ pip install maltopic
42
+ ```
43
+
44
+ ## Usage
45
+
46
+ To use the MALTopic library, you need to initialize the main class with your API key and model name. You can choose between different LLMs such as OpenAI, Google Gemini (not supported yet), or Llama (not supported yet).
47
+
48
+ ```python
49
+ from maltopic import MALTopic
50
+
51
+ # Initialize the MALTopic class
52
+ mal_topic = client = MALTopic(
53
+ api_key="your_api_key",
54
+ default_model_name="gpt-4.1-nano",
55
+ llm_type="openai",
56
+ )
57
+
58
+ enriched_df = client.enrich_free_text_with_structured_data(
59
+ survey_context="context about survey, why, how of it...",
60
+ free_text_column="column_1",
61
+ structured_data_columns=["columns_2", "column_3"],
62
+ df=df,
63
+ examples=["free text response, category 1 -> free text response with additional context", "..."], # optional
64
+ )
65
+
66
+ topics = client.generate_topics(
67
+ topic_mining_context="context about what kind of topics you want to mine",
68
+ df=enriched_df,
69
+ enriched_column="column_1" + "_enriched", # MALTopic adds _enriched as the suffix.
70
+ )
71
+
72
+ print(topics)
73
+ ```
74
+
75
+ ## Agents
76
+
77
+ - **Enrichment Agent**: Enhances free-text responses using structured data.
78
+ - **Topic Modeling Agent**: Extracts latent themes from enriched responses.
79
+ - **Deduplication Agent**: Refines and consolidates the extracted topics. (not supported yet)
80
+
81
+ ## Contributing
82
+
83
+ Contributions are welcome! Please feel free to submit a pull request or open an issue for any enhancements or bug fixes.
84
+
85
+ ## License
86
+
87
+ This project is licensed under the MIT License. See the LICENSE file for more details.
88
+
89
+ ## Citation
90
+
91
+ If you use MALTopic in your research, please cite:
92
+
93
+ ```bibtex
94
+ @software{Sharma2025maltopic,
95
+ author = {Sharma, Yash},
96
+ title = {MALTopic: A library for topic modeling},
97
+ year = {2025},
98
+ url = {https://github.com/yash91sharma/MALTopic-py}
99
+ }
100
+
@@ -0,0 +1,77 @@
1
+ # MALTopic: Multi-Agent LLM Topic Modeling Library
2
+
3
+ MALTopic is a powerful library designed for topic modeling using a multi-agent approach. It leverages the capabilities of large language models (LLMs) to enhance the analysis of survey responses by integrating structured and unstructured data.
4
+
5
+ MALTopic as a research paper was published in 2025 World AI IoT Congress. Links here.
6
+
7
+ ## Features
8
+
9
+ - **Multi-Agent Framework**: Decomposes topic modeling into specialized tasks executed by individual LLM agents.
10
+ - **Data Enrichment**: Enhances textual responses using structured and categorical survey data.
11
+ - **Latent Theme Extraction**: Extracts meaningful topics from enriched responses.
12
+ - **Topic Deduplication**: Refines and consolidates identified topics for better interpretability.
13
+
14
+ ## Installation
15
+
16
+ To install the MALTopic library, you can use pip:
17
+
18
+ ```bash
19
+ pip install maltopic
20
+ ```
21
+
22
+ ## Usage
23
+
24
+ To use the MALTopic library, you need to initialize the main class with your API key and model name. You can choose between different LLMs such as OpenAI, Google Gemini (not supported yet), or Llama (not supported yet).
25
+
26
+ ```python
27
+ from maltopic import MALTopic
28
+
29
+ # Initialize the MALTopic class
30
+ mal_topic = client = MALTopic(
31
+ api_key="your_api_key",
32
+ default_model_name="gpt-4.1-nano",
33
+ llm_type="openai",
34
+ )
35
+
36
+ enriched_df = client.enrich_free_text_with_structured_data(
37
+ survey_context="context about survey, why, how of it...",
38
+ free_text_column="column_1",
39
+ structured_data_columns=["columns_2", "column_3"],
40
+ df=df,
41
+ examples=["free text response, category 1 -> free text response with additional context", "..."], # optional
42
+ )
43
+
44
+ topics = client.generate_topics(
45
+ topic_mining_context="context about what kind of topics you want to mine",
46
+ df=enriched_df,
47
+ enriched_column="column_1" + "_enriched", # MALTopic adds _enriched as the suffix.
48
+ )
49
+
50
+ print(topics)
51
+ ```
52
+
53
+ ## Agents
54
+
55
+ - **Enrichment Agent**: Enhances free-text responses using structured data.
56
+ - **Topic Modeling Agent**: Extracts latent themes from enriched responses.
57
+ - **Deduplication Agent**: Refines and consolidates the extracted topics. (not supported yet)
58
+
59
+ ## Contributing
60
+
61
+ Contributions are welcome! Please feel free to submit a pull request or open an issue for any enhancements or bug fixes.
62
+
63
+ ## License
64
+
65
+ This project is licensed under the MIT License. See the LICENSE file for more details.
66
+
67
+ ## Citation
68
+
69
+ If you use MALTopic in your research, please cite:
70
+
71
+ ```bibtex
72
+ @software{Sharma2025maltopic,
73
+ author = {Sharma, Yash},
74
+ title = {MALTopic: A library for topic modeling},
75
+ year = {2025},
76
+ url = {https://github.com/yash91sharma/MALTopic-py}
77
+ }
@@ -0,0 +1,30 @@
1
+ [tool.poetry]
2
+ name = "maltopic"
3
+ version = "1.0.0"
4
+ description = "A multi-agent LLM topic modeling library."
5
+ authors = ["Yash Sharma <yash91sharma@gmail.com>"]
6
+ license = "MIT"
7
+ readme = "README.md"
8
+ homepage = "https://github.com/yash91sharma/MALTopic-py"
9
+ repository = "https://github.com/yash91sharma/MALTopic-py"
10
+ keywords = ["topic modeling", "LLM", "multi-agent", "survey analysis", "data enrichment"]
11
+ classifiers = [
12
+ "Programming Language :: Python :: 3",
13
+ "License :: OSI Approved :: MIT License",
14
+ "Operating System :: OS Independent",
15
+ "Topic :: Scientific/Engineering :: Artificial Intelligence"
16
+ ]
17
+ packages = [{include = "maltopic", from = "src"}]
18
+ include = ["README.md", "LICENSE"]
19
+
20
+
21
+ [tool.poetry.dependencies]
22
+ python = "^3.12"
23
+ openai = "^1.79.0"
24
+ pandas = "^2.2.3"
25
+ tqdm = "^4.67.1"
26
+
27
+
28
+ [build-system]
29
+ requires = ["poetry-core>=1.0.0"]
30
+ build-backend = "poetry.core.masonry.api"
@@ -0,0 +1,16 @@
1
+ # maltopic package initialization
2
+
3
+ """
4
+ MALTopic: A Multi-Agent LLM Topic Modeling Library
5
+
6
+ This package provides a framework for topic modeling using multiple LLM agents.
7
+ Users can initialize the library with their API key and model name, and choose
8
+ between different LLM providers (OpenAI, Google Gemini, Llama) for enhanced topic
9
+ modeling capabilities.
10
+ """
11
+
12
+ from .core import MALTopic
13
+
14
+ __all__ = [
15
+ "MALTopic",
16
+ ]
@@ -0,0 +1,140 @@
1
+ import json
2
+
3
+ import pandas as pd
4
+ from tqdm import tqdm
5
+
6
+ from . import prompts, utils
7
+
8
+
9
+ class MALTopic:
10
+ def __init__(self, api_key: str, default_model_name: str, llm_type: str):
11
+ self.api_key = api_key
12
+ self.default_model_name = default_model_name
13
+ self.llm_type = llm_type.lower()
14
+
15
+ self.llm_client = self._select_agent(api_key, default_model_name)
16
+
17
+ def _select_agent(self, api_key: str, model_name: str):
18
+ if self.llm_type == "openai":
19
+ from .llms import openai
20
+
21
+ return openai.OpenAIClient(api_key, model_name)
22
+ raise ValueError("Invalid LLM api type. Choose 'openai'.")
23
+
24
+ def enrich_free_text_with_structured_data(
25
+ self,
26
+ *,
27
+ survey_context: str,
28
+ free_text_column: str,
29
+ structured_data_columns: list[str],
30
+ df: pd.DataFrame,
31
+ examples: list[str] = [],
32
+ ) -> pd.DataFrame:
33
+ """
34
+ Enrich free text responses with structured data from other columns.
35
+
36
+ Args:
37
+ survey_context: Context about the survey to provide to the LLM
38
+ free_text_column: Name of the column containing free text responses
39
+ free_text_definition: Description of what the free text represents
40
+ structured_data_columns: List of column names containing structured data
41
+ df: Pandas DataFrame containing the data
42
+ examples: Optional list of example enrichments for few-shot learning
43
+
44
+ Returns:
45
+ DataFrame with added column '{free_text_column}_enriched' containing enriched text
46
+ """
47
+ utils.validate_dataframe(df, [free_text_column] + structured_data_columns)
48
+
49
+ instructions: str = prompts.ENRICH_INST.format(
50
+ survey_context=survey_context,
51
+ free_text_column=free_text_column,
52
+ structured_data_columns=", ".join(structured_data_columns),
53
+ examples="\n".join(examples) if examples else "No examples provided.",
54
+ )
55
+
56
+ enriched_column = f"{free_text_column}_enriched"
57
+ results: list[str] = []
58
+
59
+ for _, row in tqdm(df.iterrows(), total=len(df), desc="Enriching free text"):
60
+ free_text = str(row[free_text_column])
61
+ if not free_text or pd.isna(free_text):
62
+ results.append("")
63
+ continue
64
+
65
+ # Format structured data
66
+ structured_data: list[str] = []
67
+ for col in structured_data_columns:
68
+ if not pd.isna(row[col]):
69
+ structured_data.append(f"{col}: {row[col]}")
70
+ structured_context: str = "\n".join(structured_data)
71
+
72
+ # Build complete prompt for this row
73
+ row_prompt: str = (
74
+ f"{free_text_column}: {free_text}\n\n"
75
+ f"Structured data:\n{structured_context}\n\n"
76
+ f"Enriched Response:"
77
+ )
78
+
79
+ try:
80
+ enriched_text: str = self.llm_client.generate(
81
+ instructions=instructions, input=row_prompt
82
+ )
83
+ results.append(enriched_text.strip())
84
+ except Exception as e:
85
+ error_msg = f"Error generating enriched text: {str(e)}"
86
+ results.append(error_msg)
87
+
88
+ df[enriched_column] = results
89
+ return df
90
+
91
+ def generate_topics(
92
+ self, *, topic_mining_context: str, df: pd.DataFrame, enriched_column: str
93
+ ) -> list[dict[str, str]]:
94
+ """
95
+ Generate topics from enriched text responses.
96
+
97
+ Args:
98
+ topic_mining_context: Context about the survey to provide to the LLM. Add the what and why of the topic mining to make this useful.
99
+ df: Pandas DataFrame containing the data
100
+ enriched_column: Name of the column containing enriched text responses
101
+
102
+ Returns:
103
+ List of dictionaries, each representing a topic with 'id', 'name', and 'description'
104
+ """
105
+ utils.validate_dataframe(df, [enriched_column])
106
+
107
+ instructions = prompts.TOPIC_INST.format(survey_context=topic_mining_context)
108
+
109
+ all_columns: list[str] = df[enriched_column].dropna().tolist()
110
+ labeled_columns = [
111
+ f"{i+1}: {response}" for i, response in enumerate(all_columns)
112
+ ]
113
+ input_text = "\n\n".join(labeled_columns)
114
+
115
+ try:
116
+ raw_response = self.llm_client.generate(
117
+ instructions=instructions, input=input_text
118
+ )
119
+ except Exception as e:
120
+ raise RuntimeError(f"Error generating topics: {str(e)}")
121
+
122
+ topics = []
123
+
124
+ try:
125
+ parsed_topics = json.loads(raw_response)
126
+ for topic in parsed_topics:
127
+ for key in topic:
128
+ if key != "representative_words" and not isinstance(
129
+ topic[key], str
130
+ ):
131
+ topic[key] = str(topic[key])
132
+ topics.append(topic)
133
+ except json.JSONDecodeError:
134
+ raise ValueError(
135
+ f"Failed to parse LLM response as JSON: {raw_response[:100]}..."
136
+ )
137
+ except Exception as e:
138
+ raise ValueError(f"Error processing topics: {str(e)}")
139
+
140
+ return topics
@@ -0,0 +1 @@
1
+ # This file initializes the llms subpackage, allowing for structured imports from the LLM modules.
@@ -0,0 +1,32 @@
1
+ from openai import OpenAI
2
+
3
+
4
+ class OpenAIClient:
5
+ def __init__(self, api_key: str, model_name: str):
6
+ self.api_key = api_key
7
+ self.model_name = model_name
8
+ self.client = OpenAI(api_key=api_key)
9
+ self.temperature = 0.2
10
+ self.seed = 12345
11
+ self.top_p = 0.9
12
+
13
+ def generate(self, *, instructions: str, input: str) -> str:
14
+ response = self.client.chat.completions.create(
15
+ model=self.model_name,
16
+ store=False,
17
+ messages=[
18
+ {"role": "system", "content": instructions},
19
+ {"role": "user", "content": input},
20
+ ],
21
+ temperature=self.temperature,
22
+ top_p=self.top_p,
23
+ seed=self.seed,
24
+ )
25
+ if not response or not response.choices:
26
+ raise ValueError("No response received from OpenAI API.")
27
+ if not response.choices[0].message:
28
+ raise ValueError("No message in the response from OpenAI API.")
29
+ content = response.choices[0].message.content
30
+ if content is None:
31
+ raise ValueError("No content in the response message from OpenAI API.")
32
+ return content
@@ -0,0 +1,56 @@
1
+ ENRICH_INST = (
2
+ "You are an AI language assistant. Your task is to help enrich the"
3
+ "free-text reponse with the structured data from survey responses."
4
+ "Enrich the free-text responses by summarizing them and injecting "
5
+ "crucial informationa and subtle details from the structured "
6
+ "columns. This should help with more context aware topic modeling "
7
+ "of the free-text responses. Survey context: {survey_context}. "
8
+ "Enrich the {free_text_column} with the following structured "
9
+ "columns: {structured_data_columns}. "
10
+ "Maintain the original sentiment and meaning of the response. Do "
11
+ "not introduce any new opinions, assumptions, conclusions or "
12
+ "extrapolations which were not present in the original response. "
13
+ "Keep the language generic and standardized. Only respond with "
14
+ "the enriched response. Some examples: {examples}."
15
+ )
16
+
17
+ TOPIC_INST = (
18
+ "You are an AI NLP data analyst. Your goal is to analyze a set of survey "
19
+ "responses and identify unique, exhaustive, and non-overlapping topics."
20
+ "Survey context: {survey_context}. "
21
+ "**Task:**"
22
+ "1. **Analyze the following survey responses.** For each response, "
23
+ "note the respondent profile, the key themes and issues raised."
24
+ "2. **Identify Potential Topics.** Based on your analysis of all "
25
+ "responses, generate a list of initial, potentially granular topics. "
26
+ "Account for the distinct perspectives, experiences and background "
27
+ "different respondent might have."
28
+ "3. **Synthesize and Refine Topics.** Review the initial list of "
29
+ "topics and synthesize them into a final set of topics that meet the "
30
+ "following criteria: "
31
+ "**Uniqueness:** Each topic should represent a distinct and clearly "
32
+ "defined area of impact. Avoid redundancy and overlapping concepts. "
33
+ "**Exhaustiveness within the dataset:** The set of topics should "
34
+ "comprehensively cover the range of issues and themes expressed in "
35
+ "the survey responses. It should aim to capture all significant "
36
+ "aspects of the survey. "
37
+ "**Non-Overlapping:** Topics should be conceptually distinct and "
38
+ "not simply rephrasing of the same underlying issue. Minimize "
39
+ "semantic overlap and ensure clear boundaries between topics. "
40
+ "**Respondent-Aware:** The topics should reflect the influence of "
41
+ "respondent profiles (if found). Indicate how different respondent "
42
+ "groups relate to or emphasize each topic, if applicable. "
43
+ "**Use all the data.** Use all the data provided to you. Do not "
44
+ "ignore any data. Do not make assumptions about the data. "
45
+ "4. **Output:** Present your final output as a JSON array of topics. "
46
+ "For each topic, provide the following fields: "
47
+ "* `name`: A concise and descriptive name for the topic. "
48
+ "* `description`: A one line summary of what the topic encompasses. "
49
+ "* `relevance`: Note which respondent profiles are particularly "
50
+ "relevant to this topic, if applicable. Also include other interesting "
51
+ "and subtle patterns which you see related to this topic. "
52
+ "* `representative_words`: Array of top words which would represent "
53
+ "this topic in an NLP analysis."
54
+ "Only respond with the JSON array, with no additional text, "
55
+ "explanations or markdown."
56
+ )
@@ -0,0 +1,24 @@
1
+ import pandas as pd
2
+
3
+
4
+ def validate_dataframe(df: pd.DataFrame, required_columns: list[str]) -> None:
5
+ """
6
+ Validates a Pandas DataFrame for required columns and non-emptiness.
7
+
8
+ Args:
9
+ df: The Pandas DataFrame to validate.
10
+ required_columns: A list of column names that must be present in the DataFrame.
11
+
12
+ Raises:
13
+ ValueError: If the DataFrame is empty or if any required columns are missing.
14
+ """
15
+ if df.empty:
16
+ raise ValueError("DataFrame is empty. It must contain at least one row.")
17
+
18
+ missing_columns = [col for col in required_columns if col not in df.columns]
19
+ if missing_columns:
20
+ raise ValueError(
21
+ f"DataFrame is missing required columns: {', '.join(missing_columns)}"
22
+ )
23
+
24
+ return None