spec-agents-text 0.2.1__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- spec_agents_text-0.2.1.dist-info/METADATA +55 -0
- spec_agents_text-0.2.1.dist-info/RECORD +10 -0
- spec_agents_text-0.2.1.dist-info/WHEEL +4 -0
- spec_agents_text-0.2.1.dist-info/entry_points.txt +2 -0
- spec_agents_text-0.2.1.dist-info/licenses/LICENSE +21 -0
- text_pack/__init__.py +166 -0
- text_pack/py.typed +0 -0
- text_pack/skills/document_digest/SKILL.md +25 -0
- text_pack/skills/keywords/SKILL.md +34 -0
- text_pack/skills/text_stats/SKILL.md +24 -0
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: spec-agents-text
|
|
3
|
+
Version: 0.2.1
|
|
4
|
+
Summary: Text domain pack for agent-fabric (text_stats, keywords, document_digest); proves the core is domain-agnostic
|
|
5
|
+
Project-URL: Homepage, https://github.com/marcionicolau/spec-agents
|
|
6
|
+
Project-URL: Repository, https://github.com/marcionicolau/spec-agents
|
|
7
|
+
Project-URL: Issues, https://github.com/marcionicolau/spec-agents/issues
|
|
8
|
+
Project-URL: Changelog, https://github.com/marcionicolau/spec-agents/blob/main/packages/text-pack/CHANGELOG.md
|
|
9
|
+
Project-URL: Documentation, https://github.com/marcionicolau/spec-agents/tree/main/docs
|
|
10
|
+
Author: Marcio Nicolau
|
|
11
|
+
License-Expression: MIT
|
|
12
|
+
License-File: LICENSE
|
|
13
|
+
Keywords: agents,nlp,text
|
|
14
|
+
Classifier: Development Status :: 2 - Pre-Alpha
|
|
15
|
+
Classifier: Intended Audience :: Developers
|
|
16
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
17
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
20
|
+
Classifier: Typing :: Typed
|
|
21
|
+
Requires-Python: >=3.12
|
|
22
|
+
Requires-Dist: spec-agents-core<1,>=0.0.1
|
|
23
|
+
Description-Content-Type: text/markdown
|
|
24
|
+
|
|
25
|
+
# spec-agents-text (`text_pack`)
|
|
26
|
+
|
|
27
|
+
Small text domain pack for [agent-fabric](../agent-fabric/README.md): components `text_stats` and `keywords`, pipeline `document_digest`.
|
|
28
|
+
It needs no pandas and no other pack, which makes it the proof that the core is domain-agnostic, and a minimal template for new packs.
|
|
29
|
+
|
|
30
|
+
> **Status:** pre-release. Until the first PyPI release, install a wheel from the
|
|
31
|
+
> [GitHub Releases](https://github.com/marcionicolau/spec-agents/releases) or use `uv sync --all-packages` in a checkout;
|
|
32
|
+
> the `pip install` lines below are the PyPI names.
|
|
33
|
+
|
|
34
|
+
## Names
|
|
35
|
+
|
|
36
|
+
| | |
|
|
37
|
+
| --- | --- |
|
|
38
|
+
| PyPI distribution | `spec-agents-text` |
|
|
39
|
+
| Import name | `text_pack` |
|
|
40
|
+
| Directory in the monorepo | `packages/text-pack` |
|
|
41
|
+
| Release tag / PR scope | `text-pack-vX.Y.Z` / `text-pack` |
|
|
42
|
+
|
|
43
|
+
The distribution is named `spec-agents-text` because the shorter names are taken on PyPI; the import name and the entry point do not change.
|
|
44
|
+
|
|
45
|
+
## Install
|
|
46
|
+
```bash
|
|
47
|
+
pip install spec-agents-text
|
|
48
|
+
```
|
|
49
|
+
```python
|
|
50
|
+
from agent_fabric import build_registry
|
|
51
|
+
from text_pack import register
|
|
52
|
+
|
|
53
|
+
registry = build_registry([register])
|
|
54
|
+
```
|
|
55
|
+
Start a new pack from this one: a `register(registry)` function, `skills/<name>/SKILL.md` contracts and an `agent_fabric.domains` entry point.
|
|
@@ -0,0 +1,10 @@
|
|
|
1
|
+
text_pack/__init__.py,sha256=Z_K5VIkZv8hhhO7_qxu5EKwd8tdS1mxx5NuV-2yde4Q,4959
|
|
2
|
+
text_pack/py.typed,sha256=47DEQpj8HBSa-_TImW-5JCeuQeRkm5NMpJWZG3hSuFU,0
|
|
3
|
+
text_pack/skills/document_digest/SKILL.md,sha256=sDLOPiHKS-dRyrE9sNjaBKpq44HhAXY2LnBMGYtEVCA,589
|
|
4
|
+
text_pack/skills/keywords/SKILL.md,sha256=SzxCi7fQ0CT-ttOue4dIiVYPFE-ER-QFYoy-YH6eOmo,1014
|
|
5
|
+
text_pack/skills/text_stats/SKILL.md,sha256=k9OKAgDpsmSnJyx0uzEO8E84PMYVJEXeGJJO9NLLGL0,739
|
|
6
|
+
spec_agents_text-0.2.1.dist-info/METADATA,sha256=9MC0GQFt-5Px96y32NNkZjW0Mg5xZ5-Jjabc8cN_Jog,2365
|
|
7
|
+
spec_agents_text-0.2.1.dist-info/WHEEL,sha256=W3fkpkm7-wf9vBI5Z-7s0eWkeM-spu78I8Neb98DeEg,87
|
|
8
|
+
spec_agents_text-0.2.1.dist-info/entry_points.txt,sha256=iEAYvK5UIC-Rt5l5nrrBDCSiU91QusF_OAe8eU3rveM,49
|
|
9
|
+
spec_agents_text-0.2.1.dist-info/licenses/LICENSE,sha256=PDXl-Nq_KQ-fBkmdUynB87KFuoDraUCnx6GOpXcXAAw,1071
|
|
10
|
+
spec_agents_text-0.2.1.dist-info/RECORD,,
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Marcio Nicolau
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
text_pack/__init__.py
ADDED
|
@@ -0,0 +1,166 @@
|
|
|
1
|
+
"""Text domain pack - proves the fabric is domain-agnostic (no pandas involved).
|
|
2
|
+
|
|
3
|
+
Components: text_stats, keywords. Pipeline: document_digest.
|
|
4
|
+
Load with ``build_registry([text_pack.register])`` or entry point ``agent_fabric.domains``.
|
|
5
|
+
"""
|
|
6
|
+
|
|
7
|
+
from __future__ import annotations
|
|
8
|
+
|
|
9
|
+
import re
|
|
10
|
+
from collections import Counter
|
|
11
|
+
from pathlib import Path
|
|
12
|
+
|
|
13
|
+
from pydantic import BaseModel, Field
|
|
14
|
+
|
|
15
|
+
from agent_fabric.compat import requires_api
|
|
16
|
+
from agent_fabric.component import Component, ComponentParams, ComponentResult, Num, StepContext
|
|
17
|
+
from agent_fabric.errors import ErrorDetail
|
|
18
|
+
from agent_fabric.registry import Registry, component
|
|
19
|
+
|
|
20
|
+
SPEC_DIR = Path(__file__).parent / "skills"
|
|
21
|
+
_WORD = re.compile(r"[A-Za-zÀ-ÿ]+(?:-[A-Za-zÀ-ÿ]+)*")
|
|
22
|
+
STOPWORDS = {
|
|
23
|
+
"the",
|
|
24
|
+
"a",
|
|
25
|
+
"an",
|
|
26
|
+
"and",
|
|
27
|
+
"or",
|
|
28
|
+
"of",
|
|
29
|
+
"to",
|
|
30
|
+
"in",
|
|
31
|
+
"on",
|
|
32
|
+
"for",
|
|
33
|
+
"is",
|
|
34
|
+
"are",
|
|
35
|
+
"was",
|
|
36
|
+
"were",
|
|
37
|
+
"with",
|
|
38
|
+
"by",
|
|
39
|
+
"de",
|
|
40
|
+
"da",
|
|
41
|
+
"do",
|
|
42
|
+
"das",
|
|
43
|
+
"dos",
|
|
44
|
+
"e",
|
|
45
|
+
"o",
|
|
46
|
+
"os",
|
|
47
|
+
"as",
|
|
48
|
+
"em",
|
|
49
|
+
"no",
|
|
50
|
+
"na",
|
|
51
|
+
"para",
|
|
52
|
+
"com",
|
|
53
|
+
"que",
|
|
54
|
+
"um",
|
|
55
|
+
"uma",
|
|
56
|
+
}
|
|
57
|
+
|
|
58
|
+
|
|
59
|
+
class TextStatsParams(ComponentParams):
|
|
60
|
+
lowercase: bool = True
|
|
61
|
+
|
|
62
|
+
|
|
63
|
+
class TextStatsResult(ComponentResult):
|
|
64
|
+
n_words: int
|
|
65
|
+
n_sentences: int
|
|
66
|
+
avg_word_length: Num
|
|
67
|
+
lexical_diversity: Num
|
|
68
|
+
|
|
69
|
+
|
|
70
|
+
@component("text_stats")
|
|
71
|
+
class TextStats(Component[TextStatsParams, TextStatsResult]):
|
|
72
|
+
Params = TextStatsParams
|
|
73
|
+
Result = TextStatsResult
|
|
74
|
+
|
|
75
|
+
def compute(self, inputs: dict, params: TextStatsParams, ctx: StepContext) -> TextStatsResult:
|
|
76
|
+
"""Count words and sentences and compute the average word length and lexical diversity (distinct words over all words).
|
|
77
|
+
|
|
78
|
+
Warns when the text has fewer than 50 words, where the statistics are unstable.
|
|
79
|
+
"""
|
|
80
|
+
text = inputs["text"]
|
|
81
|
+
words = [w.lower() if params.lowercase else w for w in _WORD.findall(text)]
|
|
82
|
+
sentences = [s for s in re.split(r"[.!?]+\s", text) if s.strip()]
|
|
83
|
+
warns = ["very short text; statistics are unstable"] if len(words) < 50 else []
|
|
84
|
+
return TextStatsResult(
|
|
85
|
+
n_words=len(words),
|
|
86
|
+
n_sentences=len(sentences),
|
|
87
|
+
avg_word_length=sum(map(len, words)) / len(words) if words else None,
|
|
88
|
+
lexical_diversity=len(set(words)) / len(words) if words else None,
|
|
89
|
+
warnings=warns,
|
|
90
|
+
)
|
|
91
|
+
|
|
92
|
+
def summarize(self, r: dict) -> tuple[str, list[str]]:
|
|
93
|
+
return (
|
|
94
|
+
f"{r['n_words']} words in {r['n_sentences']} sentences.",
|
|
95
|
+
[f"lexical diversity {r['lexical_diversity']:.2f}", f"average word length {r['avg_word_length']:.2f}"],
|
|
96
|
+
)
|
|
97
|
+
|
|
98
|
+
|
|
99
|
+
class KeywordsParams(ComponentParams):
|
|
100
|
+
top_k: int = Field(10, ge=1, le=50)
|
|
101
|
+
min_length: int = Field(4, ge=1)
|
|
102
|
+
extra_stopwords: list[str] = Field(default_factory=list)
|
|
103
|
+
|
|
104
|
+
|
|
105
|
+
class Keyword(BaseModel):
|
|
106
|
+
term: str
|
|
107
|
+
count: int
|
|
108
|
+
|
|
109
|
+
|
|
110
|
+
class KeywordsResult(ComponentResult):
|
|
111
|
+
keywords: list[Keyword]
|
|
112
|
+
|
|
113
|
+
|
|
114
|
+
@component("keywords")
|
|
115
|
+
class Keywords(Component[KeywordsParams, KeywordsResult]):
|
|
116
|
+
Params = KeywordsParams
|
|
117
|
+
Result = KeywordsResult
|
|
118
|
+
|
|
119
|
+
def extra_checks(self, inputs: dict, params: KeywordsParams) -> list[ErrorDetail]:
|
|
120
|
+
n = len([w for w in _WORD.findall(inputs["text"]) if len(w) >= params.min_length])
|
|
121
|
+
if n == 0:
|
|
122
|
+
return [
|
|
123
|
+
ErrorDetail(
|
|
124
|
+
loc=("params", "min_length"),
|
|
125
|
+
type="no_candidates",
|
|
126
|
+
msg=f"no words with >= {params.min_length} letters",
|
|
127
|
+
hint="lower min_length",
|
|
128
|
+
)
|
|
129
|
+
]
|
|
130
|
+
return []
|
|
131
|
+
|
|
132
|
+
def compute(self, inputs: dict, params: KeywordsParams, ctx: StepContext) -> KeywordsResult:
|
|
133
|
+
"""Most frequent content words: lowercased words of at least ``min_length`` letters minus stop words and ``extra_stopwords``.
|
|
134
|
+
|
|
135
|
+
Emits the ranked ``terms`` list; the result carries each term with its count.
|
|
136
|
+
"""
|
|
137
|
+
stop = STOPWORDS | {w.lower() for w in params.extra_stopwords}
|
|
138
|
+
words = [
|
|
139
|
+
w.lower() for w in _WORD.findall(inputs["text"]) if len(w) >= params.min_length and w.lower() not in stop
|
|
140
|
+
]
|
|
141
|
+
top = Counter(words).most_common(params.top_k)
|
|
142
|
+
ctx.emit("terms", [t for t, _ in top])
|
|
143
|
+
return KeywordsResult(keywords=[Keyword(term=t, count=c) for t, c in top])
|
|
144
|
+
|
|
145
|
+
def summarize(self, r: dict) -> tuple[str, list[str]]:
|
|
146
|
+
kws = r["keywords"]
|
|
147
|
+
return f"Top terms: {', '.join(k['term'] for k in kws[:5])}.", [f"{k['term']}: {k['count']}" for k in kws[:6]]
|
|
148
|
+
|
|
149
|
+
|
|
150
|
+
@requires_api(1)
|
|
151
|
+
def register(registry: Registry) -> list[str]:
|
|
152
|
+
"""Register the text domain (``text_stats``, ``keywords`` and the ``document_digest`` pipeline); returns the names registered."""
|
|
153
|
+
return registry.load_domain(SPEC_DIR, [TextStats, Keywords])
|
|
154
|
+
|
|
155
|
+
|
|
156
|
+
__all__ = [
|
|
157
|
+
"SPEC_DIR",
|
|
158
|
+
"Keyword",
|
|
159
|
+
"Keywords",
|
|
160
|
+
"KeywordsParams",
|
|
161
|
+
"KeywordsResult",
|
|
162
|
+
"TextStats",
|
|
163
|
+
"TextStatsParams",
|
|
164
|
+
"TextStatsResult",
|
|
165
|
+
"register",
|
|
166
|
+
]
|
text_pack/py.typed
ADDED
|
File without changes
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: document_digest
|
|
3
|
+
kind: pipeline
|
|
4
|
+
version: 1.0.0
|
|
5
|
+
domain: text
|
|
6
|
+
description: Text statistics plus top keywords of a document.
|
|
7
|
+
params:
|
|
8
|
+
top_k: {description: number of keywords, default: 8}
|
|
9
|
+
inputs:
|
|
10
|
+
text: {type: text, description: document to digest}
|
|
11
|
+
steps:
|
|
12
|
+
- {id: stats, component: text_stats}
|
|
13
|
+
- id: kw
|
|
14
|
+
component: keywords
|
|
15
|
+
params: {top_k: $params.top_k}
|
|
16
|
+
outputs: {terms: kw.terms}
|
|
17
|
+
---
|
|
18
|
+
# Document digest
|
|
19
|
+
|
|
20
|
+
## When to use
|
|
21
|
+
A quick overview of a report, abstract or field note.
|
|
22
|
+
|
|
23
|
+
## Procedure
|
|
24
|
+
1. `text_stats`: size and vocabulary.
|
|
25
|
+
2. `keywords`: top `top_k` terms, exposed as `terms`.
|
|
@@ -0,0 +1,34 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: keywords
|
|
3
|
+
version: 1.0.0
|
|
4
|
+
domain: text
|
|
5
|
+
category: extraction
|
|
6
|
+
description: Most frequent content words after stop-word removal.
|
|
7
|
+
runtime: code
|
|
8
|
+
params:
|
|
9
|
+
top_k: {description: how many terms to return, example: 10}
|
|
10
|
+
min_length: {description: minimum word length, example: 4}
|
|
11
|
+
extra_stopwords:
|
|
12
|
+
description: additional words to ignore
|
|
13
|
+
example: [wheat]
|
|
14
|
+
inputs:
|
|
15
|
+
text:
|
|
16
|
+
type: text
|
|
17
|
+
description: document text
|
|
18
|
+
constraints: {min_chars: 20}
|
|
19
|
+
outputs:
|
|
20
|
+
terms: {type: json, description: list of top terms}
|
|
21
|
+
---
|
|
22
|
+
# Keyword extraction
|
|
23
|
+
|
|
24
|
+
## When to use
|
|
25
|
+
Find what a document is about: the most frequent content words after stop-word removal.
|
|
26
|
+
|
|
27
|
+
## Interpreting
|
|
28
|
+
Counts are raw frequencies. With short texts, ties are common, so do not over-read the ranking.
|
|
29
|
+
Use `extra_stopwords` to remove words that are frequent but uninformative in the domain
|
|
30
|
+
(e.g. "plot", "field").
|
|
31
|
+
|
|
32
|
+
## Common mistakes
|
|
33
|
+
- Setting `min_length` so high that no words qualify.
|
|
34
|
+
- Expecting multi-word phrases: only single words are extracted.
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: text_stats
|
|
3
|
+
version: 1.0.0
|
|
4
|
+
domain: text
|
|
5
|
+
category: descriptive
|
|
6
|
+
description: Word and sentence counts, average word length and lexical diversity.
|
|
7
|
+
runtime: code
|
|
8
|
+
params:
|
|
9
|
+
lowercase: {description: lower-case words before counting, example: true}
|
|
10
|
+
inputs:
|
|
11
|
+
text:
|
|
12
|
+
type: text
|
|
13
|
+
description: document text
|
|
14
|
+
constraints: {min_chars: 1}
|
|
15
|
+
---
|
|
16
|
+
# Text statistics
|
|
17
|
+
|
|
18
|
+
## When to use
|
|
19
|
+
Describe the size and vocabulary richness of a document before summarising or comparing it.
|
|
20
|
+
|
|
21
|
+
## Interpreting
|
|
22
|
+
- `lexical_diversity` is the share of distinct words. Values near 1 are typical of short texts; long
|
|
23
|
+
technical reports usually fall between 0.3 and 0.5.
|
|
24
|
+
- Statistics for texts under ~50 words are unstable. The component warns when that happens.
|