@yottameta/yotta-intel 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,16 @@
1
+ # 更新日志
2
+
3
+ ## v0.1.0 (2026-08-27)
4
+
5
+ 初始发布:
6
+
7
+ - 引擎:零依赖(Python 3.8+ 标准库)威胁情报 IOC 提取与规范化。
8
+ - 七类 IOC:IPv4 / IPv6 / 域名 / URL / 邮箱 / 哈希(MD5/SHA1/SHA256/SHA512)/ CVE 编号。
9
+ - defang / refang:识别 `hxxp`、`[.]`、`(.)`、`[dot]`、`[:]`、`[@]`、`[/]` 等常见去活性写法并还原;
10
+ 每条结果自带统一 defang 安全形态。
11
+ - 归一化:域名小写 + IDN punycode、URL 去默认端口 / 去 fragment / 保留 userinfo、哈希小写、IPv6 压缩写法。
12
+ - 去重计数:`(类型, 规范值)` 为键合并,记录 count / first_line / snippet。
13
+ - 误报控制:域名 TLD 白名单 + 文件名过滤(README.md / test.py 不算域名)+ 中文标点截断 + 哈希长度校验。
14
+ - 四种输出:text / JSON / CSV / STIX-lite(STIX 2.1 Bundle + indicator pattern,uuid5 确定性)。
15
+ - 测试:103 个用例全绿;含 CLI 端到端(stdin / 文件 / 退出码)。
16
+ - 文档:SKILL.md + README 中英双版 + references(ioc-spec / defang-rules / stix-lite-spec)。
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 YottaMeta
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/NOTICE ADDED
@@ -0,0 +1,11 @@
1
+ # NOTICE — YottaMeta 品牌声明
2
+
3
+ 「YottaMeta」「元情」「yotta-intel」「元察」「yotta-logwatch」「元史」「yotta-logs」「元忆」「yotta-memory」以及本家族各技能名称(yotta-* 前缀)是 YottaMeta 的品牌与标识。
4
+
5
+ 本软件以 MIT 许可证开源,任何人均可自由使用、修改与分发。若你在其基础上制作派生作品:
6
+
7
+ 1. 不得继续使用 YottaMeta 或本家族名称(yotta-*、元情、元察、元史、元忆 等)作为派生作品的名称;
8
+ 2. 不得暗示派生作品由 YottaMeta 官方维护、认可或与之存在关联;
9
+ 3. 建议在派生作品中明确声明「与 YottaMeta 官方无关联」。
10
+
11
+ 上游来源致谢:本技能由 YottaMeta 全新实现(零依赖自研 + 中文教学),威胁情报 IOC 提取 / defang / STIX 映射方向参考开源社区 threat-intelligence / IOC-extraction 类技能思路,无上游代码。
package/README.md ADDED
@@ -0,0 +1,170 @@
1
+ <p align="center"><b>Language</b>: English · <a href="./README.zh-CN.md">中文</a></p>
2
+
3
+ <p align="center">
4
+ <img src="assets/banner.png" alt="yotta-intel banner" width="100%" />
5
+ </p>
6
+
7
+ <h1 align="center">yotta-intel · 元情 (Yuanqing)</h1>
8
+
9
+ <p align="center">YottaMeta's zero-dependency threat-intel IOC extraction & normalization engine: extracts <b>IP (IPv4/IPv6), domains, URLs, emails, hashes (MD5/SHA1/SHA256/SHA512) and CVE IDs</b> from threat reports, security write-ups, phishing emails and logs; recognizes and reverses defanged forms; deduplicates and normalizes; outputs <b>CSV / JSON / STIX-lite</b>.</p>
10
+ <p align="center">Activates when the user provides text with suspicious IPs / domains / URLs / hashes and needs IOC extraction, normalization, dedup, format conversion, or safe sharing — <b>fully local and offline: no reputation lookups, no sample downloads, no proactive scanning</b>.</p>
11
+ <p align="center">No external tools required (Python 3.8+ standard library); Windows + Linux + macOS; every result ships a defanged form for safe sharing plus a Chinese plain-language context line.</p>
12
+
13
+ <p align="center">
14
+ <a href="LICENSE"><img alt="License: MIT" src="https://img.shields.io/badge/license-MIT-blue" /></a>
15
+ <a href="https://agentskills.io/"><img alt="Standard: agentskills.io" src="https://img.shields.io/badge/standard-agentskills.io-orange" /></a>
16
+ <a href="https://www.npmjs.com/package/@yottameta/yotta-intel"><img alt="npm package" src="https://img.shields.io/npm/v/@yottameta/yotta-intel" /></a>
17
+ <a href="https://github.com/YottaMeta/yotta-intel"><img alt="GitHub stars" src="https://img.shields.io/github/stars/YottaMeta/yotta-intel" /></a>
18
+ <a href="https://github.com/YottaMeta/yotta-intel/commits/main"><img alt="last commit" src="https://img.shields.io/github/last-commit/YottaMeta/yotta-intel" /></a>
19
+ <a href="https://github.com/YottaMeta/yotta-intel"><img alt="PRs welcome" src="https://img.shields.io/badge/PRs-welcome-brightgreen" /></a>
20
+ </p>
21
+
22
+ ## What it is
23
+
24
+ Threat analysis often starts with "find the indicators in this blob of text": which suspicious IPs does this report mention? What are the domains / links in that phishing email? What hashes appear and which algorithm do they use? Yuanqing packages this into a zero-dependency engine — no MISP / OpenCTI / commercial intelligence platform needed. It extracts IOCs, handles defang/refang, deduplicates, normalizes, and converts formats using only the Python standard library.
25
+
26
+ It is not tied to any single platform: it is an agent-agnostic toolkit that works in any agent supporting Agent Skills. **Fully local and offline** — no reputation lookups, no sample downloads, no proactive scanning, no resident service.
27
+
28
+ ## Core value
29
+
30
+ - **Zero-dependency engine** — seven IOC types + defang/refang + normalization, built entirely with Python 3.8+ standard library.
31
+ - **Seven IOC types** — IPv4 / IPv6 / domain / URL / email / hash (MD5/SHA1/SHA256/SHA512) / CVE.
32
+ - **defang / refang** — recognizes common defanged forms (`hxxp`, `[.]`, `(.)`, `[dot]`, `[:]`, `[@]`, `[/]`) and reverses them; each result ships a canonical defanged form so sharing does not trigger accidental clicks.
33
+ - **Dedup + normalization** — one record per unique IOC with occurrence count, first-seen line and context; domains lowercased + IDN punycode, URLs stripped of default ports, hashes lowercased, IPv6 compressed.
34
+ - **Four output modes** — text / JSON / CSV / STIX-lite (STIX 2.1 Bundle + indicator patterns).
35
+ - **False-positive control** — TLD whitelist + filename filter (`README.md` / `test.py` are not domains) + CJK-punctuation trimming + hash length checks.
36
+
37
+ ## Why use it
38
+
39
+ | Advantage | Description |
40
+ |---|---|
41
+ | **Zero dependency** | Python 3.8+ standard library; no daemon / database / external scanner; Windows + Linux + macOS |
42
+ | **Fully local offline** | Processes existing text only; no reputation lookups, no sample downloads, no proactive scanning |
43
+ | **defang friendly** | Recognizes mainstream defanged forms and reverses them; outputs a unified defanged form for safe sharing |
44
+ | **Explainable** | Each result includes type, count, first-seen line and context; reports candidate indicators only, never verdicts |
45
+ | **Low noise** | Deterministic rules: TLD whitelist, filename filter, CJK-punctuation trimming |
46
+ | **Ecosystem distribution** | GitHub + npm + ClawHub synced; install via npx / install.sh / manual copy |
47
+
48
+ ## Commands
49
+
50
+ | Command | Description |
51
+ |---|---|
52
+ | extract | Extract IOCs and output structured results (text / json / csv / stix) |
53
+ | extract --path / --stdin | Input from a file / standard input |
54
+ | extract --types | Extract only the given types (comma-separated, e.g. ipv4,domain,hash) |
55
+ | extract --format | Switch output format (text / json / csv / stix) |
56
+ | extract --min-count | Keep only IOCs with occurrence count >= N |
57
+ | extract --output | Write results to a file (default: stdout) |
58
+ | defang | Replace IOCs in the text with safe defanged forms |
59
+ | refang | Reverse defanged text back to its raw form |
60
+ | --version | Print the version |
61
+
62
+ Exit codes: extract **0** = no IOCs; **1** = IOCs found; **4** = usage or read error; defang / refang return **0** on success.
63
+
64
+ ## Quick start
65
+
66
+ Use `python` on Windows and `python3` on Linux/macOS.
67
+
68
+ ```bash
69
+ # Extract IOCs from a text file (all types, text output)
70
+ python3 scripts/yotta_intel.py extract --path report.txt
71
+
72
+ # Read from stdin, output JSON
73
+ cat intel.txt | python3 scripts/yotta_intel.py extract --stdin --format json
74
+
75
+ # Domains and hashes only, occurrence count >= 2
76
+ python3 scripts/yotta_intel.py extract --path intel.md --types domain,hash --min-count 2
77
+
78
+ # CSV for spreadsheets / platform import
79
+ python3 scripts/yotta_intel.py extract --path intel.md --format csv --output iocs.csv
80
+
81
+ # STIX 2.1 Bundle
82
+ python3 scripts/yotta_intel.py extract --path intel.md --format stix --output iocs.json
83
+
84
+ # Turn a report into a safe-to-share defanged copy (prevents accidental clicks)
85
+ python3 scripts/yotta_intel.py defang --path report.txt --output safe.txt
86
+
87
+ # Reverse a defanged intelligence note back to its raw form
88
+ python3 scripts/yotta_intel.py refang --path safe.txt
89
+ ```
90
+
91
+ Sample text output:
92
+
93
+ ```
94
+ 元情 yotta-intel v0.1.0 —— IOC 提取结果
95
+ 共发现 2 个 IOC:
96
+
97
+ ■ IPv4 地址(ipv4)
98
+ 203.0.113.5 ×1 行 1
99
+ defang: 203[.]0[.]113[.]5
100
+ 上下文: 攻击者从 203.0.113.5 发起请求。
101
+ ```
102
+
103
+ ## Installation
104
+
105
+ Pick any of the three methods; skill files are always fetched from **npm** (GitHub can be slow without a proxy; npm supports mirrors).
106
+
107
+ ### Method 1: npm (recommended, one-liner)
108
+ ```bash
109
+ # Optional China mirror: npm config set registry https://registry.npmmirror.com
110
+ npx -y @yottameta/yotta-intel -g
111
+ npx -y @yottameta/yotta-intel --dir <your skills dir> # any agent: install to a custom directory
112
+ ```
113
+ > Agent not in the preset list? Use `--dir` to point at its skills directory, or copy manually (Method 3). `--list` shows the default directory of each agent.
114
+
115
+ ### Method 2: install.sh
116
+ After obtaining the skill folder (`npm pack` unpack or `git clone`), enter the folder:
117
+ ```bash
118
+ bash install.sh -g # user-level; bash install.sh --list shows all directories
119
+ bash install.sh --agent codex # a specific agent (see --list)
120
+ bash install.sh # project-level: auto-detect existing skills directories
121
+ bash install.sh --dir /path/to/skills
122
+ ```
123
+ > Covers 17 agent families, including Trae / Qwen / Comate / CodeBuddy / Kimi.
124
+
125
+ ### Method 3: manual copy
126
+ Copy the whole `yotta-intel` folder into the target agent's skills directory. Common user-level locations (`%USERPROFILE%` on Windows, `~` on Linux/macOS):
127
+
128
+ | Agent | User-level directory | Project-level directory |
129
+ |---|---|---|
130
+ | Codex | `%USERPROFILE%\.codex\skills\yotta-intel\` | `.codex\skills\` |
131
+ | Claude Code | `%USERPROFILE%\.claude\skills\yotta-intel\` | `.claude\skills\` |
132
+ | Cursor | `%USERPROFILE%\.cursor\skills\yotta-intel\` | `.cursor\skills\` |
133
+ | Windsurf | `%USERPROFILE%\.codeium\windsurf\skills\yotta-intel\` | `.windsurf\skills\` |
134
+ | opencode | `%USERPROFILE%\.config\opencode\skills\yotta-intel\` | `.opencode\skills\` |
135
+ | Gemini | `%USERPROFILE%\.gemini\skills\yotta-intel\` | `.gemini\skills\` |
136
+ | Goose | `%USERPROFILE%\.config\goose\skills\yotta-intel\` | `.goose\skills\` |
137
+ | Amp | `%USERPROFILE%\.config\agents\skills\yotta-intel\` | `.agents\skills\` |
138
+ | Kiro | `%USERPROFILE%\.kiro\skills\yotta-intel\` | `.kiro\skills\` |
139
+ | WorkBuddy | `%USERPROFILE%\.workbuddy\skills\yotta-intel\` | `.workbuddy\skills\` |
140
+ | Trae Code CLI | `%USERPROFILE%\.traecli\skills\yotta-intel\` | `.traecli\skills\` |
141
+ | Trae IDE (CN) | `%USERPROFILE%\.trae-cn\skills\yotta-intel\` | `.trae\skills\` |
142
+ | Qwen Code | `%USERPROFILE%\.qwen\skills\yotta-intel\` | `.qwen\skills\` |
143
+ | Comate | `%USERPROFILE%\.comate\skills\yotta-intel\` | `.comate\skills\` |
144
+ | CodeBuddy | `%USERPROFILE%\.codebuddy\skills\yotta-intel\` | `.codebuddy\skills\` |
145
+ | Kimi | `%USERPROFILE%\.kimi\skills\yotta-intel\` | `.kimi\skills\` |
146
+ | Generic AGENTS.md | `%USERPROFILE%\.agents\skills\yotta-intel\` | `.agents\skills\` |
147
+
148
+ > If Codex's `CODEX_HOME` is set, it overrides the default; the same applies to opencode's `XDG_CONFIG_HOME`. `.agents\skills` is not a universal directory — only OpenCode / Cursor / Cline / Amp / Kimi / Gemini CLI / GitHub Copilot etc. read it; **Claude Code and Codex do not read it by default**. When unsure, use `--dir` or let the agent install it.
149
+
150
+ ## Output formats
151
+
152
+ - **text** — a readable report grouped by type (defanged form + first-seen context included);
153
+ - **json** — `{tool, version, generated, source, summary, indicators[]}`; each `indicator` carries `type / value / defanged / count / first_line / snippet`;
154
+ - **csv** — `type,value,defanged,count,first_line,snippet`;
155
+ - **stix** — a STIX 2.1 Bundle; each IOC becomes an `indicator` (pattern + `x_yottameta_*` extension properties); see `references/stix-lite-spec.md`.
156
+
157
+ ## Development & validation
158
+
159
+ The package ships its own test script (included in the published package):
160
+
161
+ ```bash
162
+ # Run the full suite (103 cases) from the skill directory
163
+ python scripts/test_yotta_intel.py
164
+ ```
165
+
166
+ Spec details live in `references/`: ioc-spec.md (type rules), defang-rules.md (defang rules), stix-lite-spec.md (STIX mapping).
167
+
168
+ ## License
169
+
170
+ MIT © YottaMeta — see [LICENSE](./LICENSE).
@@ -0,0 +1,181 @@
1
+ <p align="center"><b>Language</b>: <a href="./README.md">English</a> · 中文</p>
2
+
3
+ <p align="center">
4
+ <img src="assets/banner.png" alt="yotta-intel banner" width="100%" />
5
+ </p>
6
+
7
+ <h1 align="center">yotta-intel · 元情 (Yuanqing)</h1>
8
+
9
+ <p align="center">YottaMeta 的零依赖威胁情报 IOC 提取与规范化引擎:从威胁情报文本 / 安全报告 / 钓鱼邮件 / 日志中提取
10
+ <b>IP(IPv4/IPv6)、域名、URL、邮箱、哈希(MD5/SHA1/SHA256/SHA512)与 CVE 编号</b>,
11
+ 识别并还原 defang(去活性)写法,去重、归一化后输出 <b>CSV / JSON / STIX-lite</b>。</p>
12
+ <p align="center">触发场景:用户给出含可疑 IP / 域名 / URL / 哈希的情报文本,要提取 IOC、规范化、去重、转格式或安全共享时 —
13
+ <b>纯本地离线,不联网查证、不下载样本、不主动扫描任何系统</b>。</p>
14
+ <p align="center">零外部依赖(Python 3.8+ 标准库);Windows + Linux + macOS;每条结果自带安全共享用的 defang 形态与中文说明。</p>
15
+
16
+ <p align="center">
17
+ <a href="LICENSE"><img alt="License: MIT" src="https://img.shields.io/badge/license-MIT-blue" /></a>
18
+ <a href="https://agentskills.io/"><img alt="Standard: agentskills.io" src="https://img.shields.io/badge/standard-agentskills.io-orange" /></a>
19
+ <a href="https://www.npmjs.com/package/@yottameta/yotta-intel"><img alt="npm package" src="https://img.shields.io/npm/v/@yottameta/yotta-intel" /></a>
20
+ <a href="https://github.com/YottaMeta/yotta-intel"><img alt="GitHub stars" src="https://img.shields.io/github/stars/YottaMeta/yotta-intel" /></a>
21
+ <a href="https://github.com/YottaMeta/yotta-intel/commits/main"><img alt="last commit" src="https://img.shields.io/github/last-commit/YottaMeta/yotta-intel" /></a>
22
+ <a href="https://github.com/YottaMeta/yotta-intel"><img alt="PRs welcome" src="https://img.shields.io/badge/PRs-welcome-brightgreen" /></a>
23
+ </p>
24
+
25
+ ## 这是什么
26
+
27
+ 威胁情报分析经常从「一堆文本里找指标」开始:这份报告里有哪些可疑 IP?钓鱼邮件里的域名 / 链接是什么?
28
+ 样本哈希是多少、属于哪种算法?元情把这些能力打包成零依赖引擎——不需要 MISP / OpenCTI / 商业情报平台,
29
+ 用纯 Python 标准库就能完成 IOC 提取、defang/refang、去重、归一化与格式转换。
30
+
31
+ 它不是任何单一平台的专属工具:它是一套与智能体无关的工具包,任何支持 Agent Skills 的智能体都能用。
32
+ **纯本地离线**——不联网查证、不下载样本、不主动扫描任何系统、无常驻服务。
33
+
34
+ ## 核心价值
35
+
36
+ - **零依赖引擎**——七类 IOC 提取 + defang/refang + 归一化,全部用 Python 3.8+ 标准库实现;
37
+ - **七类 IOC**——IPv4 / IPv6 / 域名 / URL / 邮箱 / 哈希(MD5/SHA1/SHA256/SHA512)/ CVE;
38
+ - **defang / refang**——识别 `hxxp`、`[.]`、`(.)`、`[dot]`、`[:]`、`[@]`、`[/]` 等常见去活性写法并还原;
39
+ 每条结果自带统一 defang 形态,共享时防误点;
40
+ - **去重 + 归一化**——同一 IOC 只保留一条,记录出现次数 / 首次行号 / 上下文;域名小写 + IDN punycode、
41
+ URL 去默认端口、哈希小写、IPv6 压缩写法;
42
+ - **四种输出**——text / JSON / CSV / STIX-lite(STIX 2.1 Bundle + indicator pattern);
43
+ - **误报控制**——域名 TLD 白名单 + 文件名过滤(`README.md` / `test.py` 不算域名)+ 中文标点截断 + 哈希长度校验。
44
+
45
+ ## 为什么用它
46
+
47
+ | 优势 | 说明 |
48
+ |---|---|
49
+ | **零依赖** | Python 3.8+ 标准库;无守护进程 / 数据库 / 外部扫描器;Windows + Linux + macOS |
50
+ | **纯本地离线** | 只处理已存在的文本内容;不联网查证、不下载样本、不主动扫描 |
51
+ | **defang 友好** | 识别主流去活性写法并还原;输出统一 defang 形态,安全共享防误点 |
52
+ | **可解释** | 每条结果带类型、出现次数、首次行号与上下文;只给「候选指标」,不给定性结论 |
53
+ | **低误报** | TLD 白名单 + 文件名过滤 + 中文标点截断等确定性规则 |
54
+ | **生态分发** | GitHub + npm + ClawHub 三源同步;npx / install.sh / 手动复制三种安装方式 |
55
+
56
+ ## 命令
57
+
58
+ | 命令 | 说明 |
59
+ |---|---|
60
+ | extract | 提取 IOC 并输出结构化结果(text / json / csv / stix) |
61
+ | extract --path / --stdin | 指定输入文件 / 从标准输入读取 |
62
+ | extract --types | 只提取指定类型(逗号分隔,如 ipv4,domain,hash) |
63
+ | extract --format | 切换输出格式(text / json / csv / stix) |
64
+ | extract --min-count | 只保留出现次数 >= N 的 IOC |
65
+ | extract --output | 结果写入文件(默认打印到 stdout) |
66
+ | defang | 把文本中识别到的 IOC 替换为安全 defang 形态 |
67
+ | refang | 把 defang 文本还原为原始形态 |
68
+ | --version | 打印版本 |
69
+
70
+ 退出码:extract **0** = 无 IOC;**1** = 发现 IOC;**4** = 用法或读取错误;defang / refang 成功均为 **0**。
71
+
72
+ ## 快速上手
73
+
74
+ Windows 用 python,Linux/macOS 用 python3。
75
+
76
+ ```bash
77
+ # 提取文本中的 IOC(默认全部类型,文本输出)
78
+ python3 scripts/yotta_intel.py extract --path report.txt
79
+
80
+ # 从标准输入读取,输出 JSON
81
+ cat intel.txt | python3 scripts/yotta_intel.py extract --stdin --format json
82
+
83
+ # 只提取域名与哈希,且出现次数 >= 2
84
+ python3 scripts/yotta_intel.py extract --path intel.md --types domain,hash --min-count 2
85
+
86
+ # 输出 CSV 供表格 / 平台导入
87
+ python3 scripts/yotta_intel.py extract --path intel.md --format csv --output iocs.csv
88
+
89
+ # 输出 STIX 2.1 Bundle
90
+ python3 scripts/yotta_intel.py extract --path intel.md --format stix --output iocs.json
91
+
92
+ # 把报告转成可安全共享的 defang 版(防误点)
93
+ python3 scripts/yotta_intel.py defang --path report.txt --output safe.txt
94
+
95
+ # 把 defang 情报还原成原始形态
96
+ python3 scripts/yotta_intel.py refang --path safe.txt
97
+ ```
98
+
99
+ 输出示例(text):
100
+
101
+ ```
102
+ 元情 yotta-intel v0.1.0 —— IOC 提取结果
103
+ 共发现 2 个 IOC:
104
+
105
+ ■ IPv4 地址(ipv4)
106
+ 203.0.113.5 ×1 行 1
107
+ defang: 203[.]0[.]113[.]5
108
+ 上下文: 攻击者从 203.0.113.5 发起请求。
109
+ ```
110
+
111
+ ## 安装
112
+
113
+ 三种方式任选其一,技能文件统一从 **npm** 获取(GitHub 无代理时较慢,npm 可配国内镜像加速)。
114
+
115
+ ### 方式一:npm(推荐,一行安装)
116
+ ```bash
117
+ # 国内加速(可选):npm config set registry https://registry.npmmirror.com
118
+ npx -y @yottameta/yotta-intel -g
119
+ npx -y @yottameta/yotta-intel --dir <你的技能目录> # 任意智能体:指定目录安装
120
+ ```
121
+ > 智能体不在预置列表里?用 `--dir` 指定它的 skills 目录,或手动复制(方式三)。`--list` 可查看各智能体对应的默认目录。
122
+
123
+ ### 方式二:install.sh 一键安装
124
+ 获取技能文件夹后(`npm pack` 解包或 `git clone`),进入技能文件夹:
125
+ ```bash
126
+ bash install.sh -g # 用户级;bash install.sh --list 查看全部目录
127
+ bash install.sh --agent codex # 指定智能体(--list 可查看可用项)
128
+ bash install.sh # 项目级:自动检测已存在的 skills 目录
129
+ bash install.sh --dir /path/to/skills
130
+ ```
131
+ > 覆盖 17 类智能体,含国内 Trae / Qwen / Comate / CodeBuddy / Kimi。
132
+
133
+ ### 方式三:手动复制
134
+ 把整个 `yotta-intel` 文件夹复制到目标智能体的 skills 目录。常见位置(用户级;Windows 用 `%USERPROFILE%`,Linux/macOS 用 `~`):
135
+
136
+ | 智能体 | 用户级目录 | 项目级目录 |
137
+ |---|---|---|
138
+ | Codex | `%USERPROFILE%\.codex\skills\yotta-intel\` | `.codex\skills\` |
139
+ | Claude Code | `%USERPROFILE%\.claude\skills\yotta-intel\` | `.claude\skills\` |
140
+ | Cursor | `%USERPROFILE%\.cursor\skills\yotta-intel\` | `.cursor\skills\` |
141
+ | Windsurf | `%USERPROFILE%\.codeium\windsurf\skills\yotta-intel\` | `.windsurf\skills\` |
142
+ | opencode | `%USERPROFILE%\.config\opencode\skills\yotta-intel\` | `.opencode\skills\` |
143
+ | Gemini | `%USERPROFILE%\.gemini\skills\yotta-intel\` | `.gemini\skills\` |
144
+ | Goose | `%USERPROFILE%\.config\goose\skills\yotta-intel\` | `.goose\skills\` |
145
+ | Amp | `%USERPROFILE%\.config\agents\skills\yotta-intel\` | `.agents\skills\` |
146
+ | Kiro | `%USERPROFILE%\.kiro\skills\yotta-intel\` | `.kiro\skills\` |
147
+ | WorkBuddy | `%USERPROFILE%\.workbuddy\skills\yotta-intel\` | `.workbuddy\skills\` |
148
+ | Trae Code CLI | `%USERPROFILE%\.traecli\skills\yotta-intel\` | `.traecli\skills\` |
149
+ | Trae IDE(国内) | `%USERPROFILE%\.trae-cn\skills\yotta-intel\` | `.trae\skills\` |
150
+ | Qwen Code | `%USERPROFILE%\.qwen\skills\yotta-intel\` | `.qwen\skills\` |
151
+ | Comate | `%USERPROFILE%\.comate\skills\yotta-intel\` | `.comate\skills\` |
152
+ | CodeBuddy | `%USERPROFILE%\.codebuddy\skills\yotta-intel\` | `.codebuddy\skills\` |
153
+ | Kimi | `%USERPROFILE%\.kimi\skills\yotta-intel\` | `.kimi\skills\` |
154
+ | 通用 AGENTS.md | `%USERPROFILE%\.agents\skills\yotta-intel\` | `.agents\skills\` |
155
+
156
+ > Codex 默认目录若设置了环境变量 `CODEX_HOME`,以该变量为准;opencode 若设置 `XDG_CONFIG_HOME` 同理。`.agents\skills` 并非通用目录,仅 OpenCode / Cursor / Cline / Amp / Kimi / Gemini CLI / GitHub Copilot 等会读取,**Claude Code 与 Codex 默认不读**。不确定时用 `--dir` 指定,或让该智能体自行安装。
157
+
158
+ ## 输出格式
159
+
160
+ - **text**:按类型分组的可读报告(含 defang 形态与首次出现的上下文);
161
+ - **json**:`{tool, version, generated, source, summary, indicators[]}`,`indicators` 每条含
162
+ `type / value / defanged / count / first_line / snippet`;
163
+ - **csv**:`type,value,defanged,count,first_line,snippet`;
164
+ - **stix**:STIX 2.1 Bundle,每条 IOC 生成一个 `indicator`(pattern + `x_yottameta_*` 扩展属性),
165
+ 详见 `references/stix-lite-spec.md`。
166
+
167
+ ## 开发与校验
168
+
169
+ 技能包内自带测试脚本(随包发布):
170
+
171
+ ```bash
172
+ # 在技能目录内运行全部测试(103 个用例)
173
+ python scripts/test_yotta_intel.py
174
+ ```
175
+
176
+ 规则与规范的细节见 `references/`:ioc-spec.md(类型判定)、defang-rules.md(defang 规则)、
177
+ stix-lite-spec.md(STIX 映射)。
178
+
179
+ ## 许可证
180
+
181
+ MIT © YottaMeta —— 详见 [LICENSE](./LICENSE)。
package/SKILL.md ADDED
@@ -0,0 +1,116 @@
1
+ ---
2
+ name: yotta-intel
3
+ version: 0.1.0
4
+ description: 元情 —— 跨智能体的威胁情报 IOC 提取与规范化技能:零依赖自研从文本 / 日志 / 报告中提取 IP(IPv4/IPv6)、域名、URL、邮箱、哈希(MD5/SHA1/SHA256/SHA512)与 CVE 编号,识别并还原 defang 写法,去重、归一化后输出 CSV / JSON / STIX-lite。触发:用户给出含可疑 IP / 域名 / URL / 哈希的威胁情报文本、恶意样本分析报告、钓鱼邮件或日志,要提取 IOC、规范化、去重、转格式、共享情报时。边界:纯本地离线提取与规范化;不联网查证、不下载样本、不主动扫描任何系统;仅用于已获授权 / 自有资产 / 教学环境的安全分析。
5
+ license: MIT
6
+ ---
7
+
8
+ # 元情(yotta-intel)
9
+
10
+ 跨智能体的威胁情报 IOC 提取与规范化技能:零依赖自研从**威胁情报文本 / 安全报告 / 钓鱼邮件 / 日志**中提取
11
+ **IP(IPv4/IPv6)、域名、URL、邮箱、哈希(MD5/SHA1/SHA256/SHA512)与 CVE 编号**,
12
+ 自动识别 defang(去活性)写法并还原,去重、归一化后输出 **CSV / JSON / STIX-lite**。
13
+
14
+ 纯 Python 3.8+ 标准库实现,零外部依赖;Windows + Linux + macOS 通用。
15
+ **纯本地离线处理**:不联网查证、不下载样本、不主动扫描任何系统。
16
+
17
+ ## 何时使用
18
+
19
+ - 用户给出一段威胁情报 / 恶意样本分析报告 / 钓鱼邮件 / 日志,要提取其中的 IOC;
20
+ - 要批量归一化、去重一堆 IP / 域名 / 哈希,或把 defang 文本还原成规范形态;
21
+ - 要把 IOC 整理成 CSV / JSON,或转成 STIX 2.1 Indicator 供平台导入;
22
+ - 要在邮件 / 群聊 / 工单里安全共享 IOC(defang 防误点)之前做一次清洗。
23
+
24
+ **Do NOT trigger**:
25
+ - 不联网查证 IOC 是否恶意、不查询威胁情报平台、不下载样本;
26
+ - 不主动扫描网络 / 主机,只处理**已存在**的文本内容;
27
+ - 不在无授权情况下用于对抗他人系统;仅用于已获授权 / 自有资产 / 教学环境的安全分析。
28
+
29
+ ## 快速使用
30
+
31
+ Windows 用 python,Linux/macOS 用 python3。
32
+
33
+ ```bash
34
+ # 提取文本中的 IOC(默认全部类型,文本输出)
35
+ python3 scripts/yotta_intel.py extract --path report.txt
36
+
37
+ # 从标准输入读取,输出 JSON
38
+ cat intel.txt | python3 scripts/yotta_intel.py extract --stdin --format json
39
+
40
+ # 只提取域名与哈希,且出现次数 >= 2
41
+ python3 scripts/yotta_intel.py extract --path intel.md --types domain,hash --min-count 2
42
+
43
+ # 输出 CSV 供表格 / 平台导入
44
+ python3 scripts/yotta_intel.py extract --path intel.md --format csv --output iocs.csv
45
+
46
+ # 输出 STIX 2.1 Bundle
47
+ python3 scripts/yotta_intel.py extract --path intel.md --format stix --output iocs.json
48
+
49
+ # 把报告转成可安全共享的 defang 版(防误点)
50
+ python3 scripts/yotta_intel.py defang --path report.txt --output safe.txt
51
+
52
+ # 把 defang 情报还原成原始形态
53
+ python3 scripts/yotta_intel.py refang --path safe.txt
54
+ ```
55
+
56
+ 退出码:**0** = 无 IOC(extract);**1** = 发现 IOC(extract);**4** = 用法或读取错误。
57
+ defang / refang 成功均为 **0**。
58
+
59
+ ## 工作流程(AI 智能体执行 IOC 提取时)
60
+
61
+ 1. **确认范围**:明确要处理的文本 / 文件与需要的 IOC 类型;纯本地文本,不做任何外部动作。
62
+ 2. **读取**:用 `extract --path` 指向文件,或 `--stdin` 从管道读取。
63
+ 3. **提取**:引擎自动 refang 预处理 → 逐行提取七类 IOC → 归一化(小写 / IDN punycode / 去默认端口等)。
64
+ 4. **去重**:同一 IOC 只保留一条,记录出现次数、首次行号与上下文;`--min-count` 过滤低频噪音。
65
+ 5. **输出**:text / JSON / CSV / STIX-lite 四种格式;`--output` 写文件,默认打印。
66
+ 6. **决策纪律**:结果只是「候选指标」,是否恶意需人工 / 其他情报源核实;共享时用 defang 形态防误点。
67
+
68
+ ## 功能
69
+
70
+ - **七类 IOC 提取**:IPv4、IPv6、域名、URL、邮箱、哈希(MD5/SHA1/SHA256/SHA512)、CVE 编号;
71
+ - **defang / refang**:识别 `hxxp`、`[.]`、`(.)`、`[dot]`、`[:]`、`[@]`、`[/]` 等常见去活性写法并还原;
72
+ 每条结果自带统一的 defang 安全形态;
73
+ - **归一化**:域名小写 + IDN punycode、URL 去默认端口 / 去 fragment / 保留 userinfo、哈希小写、IPv6 压缩写法、
74
+ IPv4-mapped IPv6 规范化;
75
+ - **误报控制**:域名 TLD 白名单 + 文件名过滤(`README.md` / `test.py` 不算域名)+ 中文标点截断 + 哈希长度校验;
76
+ - **去重计数**:以 `(类型, 规范值)` 为键合并,记录 `count` / `first_line` / `snippet`;
77
+ - **四种输出**:文本 / JSON(stdout 纯净)/ CSV / STIX-lite(STIX 2.1 Bundle + indicator pattern,uuid5 确定性)。
78
+
79
+ 详细的类型判定、defang 与 STIX 映射见 references/。
80
+
81
+ ## IOC 类型一览
82
+
83
+ | 类型 | 中文 | 示例 | 说明 |
84
+ |---|---|---|---|
85
+ | ipv4 | IPv4 地址 | `203.0.113.5` | 合法八位组;前导零归一 |
86
+ | ipv6 | IPv6 地址 | `2001:db8::1` | 压缩写法;IPv4-mapped 输出规范十六进制 |
87
+ | domain | 域名 | `evil.example.com` | TLD 白名单;IDN 转 punycode;`README.md` 不算域名 |
88
+ | url | URL | `http://evil.example.com/a` | http/https/ftp;去默认端口与 fragment |
89
+ | email | 邮箱 | `admin@example.com` | 域名段校验;defang 形态 `admin[@]example[.]com` |
90
+ | hash | 哈希 | `44d886…2f` | 仅 32/40/64/128 位十六进制 |
91
+ | cve | CVE 编号 | `CVE-2024-1234` | 统一大写 |
92
+
93
+ ## 输出格式
94
+
95
+ - **text**:按类型分组的可读报告(含 defang 形态与首次出现的上下文);
96
+ - **json**:`{tool, version, generated, source, summary, indicators[]}`,`indicators` 每条含
97
+ `type / value / defanged / count / first_line / snippet`;
98
+ - **csv**:`type,value,defanged,count,first_line,snippet`;
99
+ - **stix**:STIX 2.1 Bundle,每条 IOC 生成一个 `indicator`(pattern + `x_yottameta_*` 扩展属性)。
100
+
101
+ ## 边界(安全红线)
102
+
103
+ - **纯本地离线**:不联网查证、不下载样本、不主动扫描任何系统,只做文本提取与规范化;
104
+ - **不给定性**:所有 IOC 只是「候选指标」,是否恶意需人工 / 其他情报源核实;
105
+ - **授权**:仅用于已获明确授权 / 自有资产 / 教学环境的安全分析;未经授权分析他人数据违反法律,使用者自行承担责任。
106
+
107
+ ## 参考文档
108
+
109
+ - references/ioc-spec.md — IOC 类型与判定规则(归一化 / 误报控制 / 已知取舍)
110
+ - references/defang-rules.md — defang / refang 规则与安全共享建议
111
+ - references/stix-lite-spec.md — STIX-lite 输出规范与 pattern 映射
112
+
113
+ ## 法律声明
114
+
115
+ 本技能仅用于**已获明确授权**的安全分析(自有资产、授权测试、CTF 靶场、教学环境)。
116
+ 未经授权分析他人系统数据违反中国《网络安全法》与《刑法》相关条款,使用者自行承担法律责任。
Binary file