@yottameta/yotta-logs 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +10 -0
- package/LICENSE +21 -0
- package/NOTICE +11 -0
- package/README.md +163 -0
- package/SKILL.md +80 -0
- package/assets/banner.png +0 -0
- package/bin/install.js +163 -0
- package/install.sh +132 -0
- package/package.json +32 -0
- package/references/cli.md +60 -0
- package/references/format.md +40 -0
- package/references/security.md +42 -0
- package/scripts/test_yotta_logs.py +404 -0
- package/scripts/yotta_logs.py +877 -0
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# CLI 协议(cli)
|
|
2
|
+
|
|
3
|
+
入口:`scripts/yotta_logs.py`(Python 3.8+ 标准库,零依赖)。
|
|
4
|
+
|
|
5
|
+
## 通用选项
|
|
6
|
+
|
|
7
|
+
| 选项 | 说明 |
|
|
8
|
+
|---|---|
|
|
9
|
+
| --dir PATH | 会话日志目录;缺省读环境变量 YOTTA_LOGS_DIR,再自动定位首个已知目录 |
|
|
10
|
+
| --json | 输出纯净 JSON(stdout 无其它噪音) |
|
|
11
|
+
| --no-redact | 关闭默认脱敏 |
|
|
12
|
+
| --limit N | 最多返回 N 条(默认 50) |
|
|
13
|
+
| --version | 打印版本 |
|
|
14
|
+
|
|
15
|
+
## 子命令
|
|
16
|
+
|
|
17
|
+
### locate
|
|
18
|
+
自动发现本机常见会话日志目录(`~/.clawdbot/agents/*/sessions`、`~/.codex/sessions`、`~/.claude/projects/*`、`~/.config/opencode/sessions`、`~/.gemini/sessions`、`~/.agents/sessions`),只返回存在且含 `*.jsonl` 的目录。
|
|
19
|
+
|
|
20
|
+
### scan
|
|
21
|
+
列出目录下所有会话:ID / 日期(首条消息)/ 消息数 / 大小 / sessions.json 别名。支持 `--limit`、`--json`。空目录退出码 1。
|
|
22
|
+
|
|
23
|
+
### search <query>
|
|
24
|
+
跨会话检索,输出时间线命中(会话 / 时间 / 角色 / 原文片段)。
|
|
25
|
+
|
|
26
|
+
| 选项 | 说明 |
|
|
27
|
+
|---|---|
|
|
28
|
+
| --regex | 把 query 当正则(默认不区分大小写) |
|
|
29
|
+
| --date YYYY-MM-DD | 只检索指定日期(也支持 YYYY-MM) |
|
|
30
|
+
| -s / --session SID | 只检索指定会话 ID / 别名(可多次) |
|
|
31
|
+
| --role ROLE | user / assistant / tool / system |
|
|
32
|
+
| --context N | 命中上下文半径字符数(默认 40) |
|
|
33
|
+
| --limit N | 最多返回 N 条(默认 50) |
|
|
34
|
+
|
|
35
|
+
--json 输出:`{command, tool, version, query, regex, dir, total_matches, sessions_hit, truncated, matches[]}`,每条命中含 `session / timestamp / role / line / match / text`。
|
|
36
|
+
|
|
37
|
+
### session <sid>
|
|
38
|
+
提取单个会话原文(时间线 + 角色 + 文本)。`--role` 过滤、`--tools` 标注工具调用、`--limit`、`--json`。未知会话退出码 4。
|
|
39
|
+
|
|
40
|
+
### stats
|
|
41
|
+
会话统计:消息 / 角色分布 / token / 成本 / 时间范围;`--daily` 输出每日汇总;`-s/--session` 限定单会话。空目录退出码 1。
|
|
42
|
+
|
|
43
|
+
### tools
|
|
44
|
+
工具调用次数排行(`toolCall` / `toolResult` 的工具名计数),`-s/--session` 限定单会话。
|
|
45
|
+
|
|
46
|
+
### version
|
|
47
|
+
打印 `yotta-logs 0.1.0`。
|
|
48
|
+
|
|
49
|
+
## 退出码
|
|
50
|
+
|
|
51
|
+
| 码 | 含义 |
|
|
52
|
+
|---|---|
|
|
53
|
+
| 0 | 成功(检索到结果 / 操作完成) |
|
|
54
|
+
| 1 | 无匹配 / 空结果集 |
|
|
55
|
+
| 4 | 用法错误 / 目录不存在 / 未知会话 / 致命异常 |
|
|
56
|
+
|
|
57
|
+
## JSON 输出约定
|
|
58
|
+
|
|
59
|
+
- stdout 只输出 JSON,进度 / 提示走 stderr;
|
|
60
|
+
- `ensure_ascii=False`,中文原样输出;stdout 已重配 UTF-8(GBK 控制台不炸)。
|
|
@@ -0,0 +1,40 @@
|
|
|
1
|
+
# 会话日志格式(format)
|
|
2
|
+
|
|
3
|
+
元史按「目录 = 会话集」工作:一个目录下有若干 `*.jsonl` 会话文件,可选一个 `sessions.json` 索引。目录由 `--dir` 指定(缺省读环境变量 `YOTTA_LOGS_DIR`,再自动定位首个已知目录)。
|
|
4
|
+
|
|
5
|
+
## 会话文件(*.jsonl)
|
|
6
|
+
|
|
7
|
+
每行一个 JSON 对象,追加写入。常见字段:
|
|
8
|
+
|
|
9
|
+
| 字段 | 说明 |
|
|
10
|
+
|---|---|
|
|
11
|
+
| type | `session`(会话元数据)或 `message`(消息) |
|
|
12
|
+
| timestamp | ISO 时间戳(如 `2026-08-26T03:00:01+08:00`) |
|
|
13
|
+
| message.role | `user` / `assistant` / `toolResult`(归一为 `tool`) |
|
|
14
|
+
| message.content | 文本或列表;列表项 `type=text` 取文本,`type=toolCall` / `toolResult` 取工具名 |
|
|
15
|
+
| message.usage.cost.total | 单条成本 |
|
|
16
|
+
| message.usage.input_tokens / output_tokens | token 数 |
|
|
17
|
+
|
|
18
|
+
会话 ID = 文件名主干(去掉 `.jsonl`)。文件只被读取,绝不写入 / 修改 / 删除。
|
|
19
|
+
|
|
20
|
+
## 索引文件(sessions.json)
|
|
21
|
+
|
|
22
|
+
可选,把别名映射到会话 ID。两种形态均可:
|
|
23
|
+
|
|
24
|
+
```json
|
|
25
|
+
{ "微信-部署": "a1", "ci-排查": "b2" }
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
```json
|
|
29
|
+
[ { "key": "微信-部署", "sessionId": "a1" } ]
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
search / session / stats / tools 的 `--session` / `-s` 参数同时接受会话 ID 与别名。
|
|
33
|
+
|
|
34
|
+
## 容错规则
|
|
35
|
+
|
|
36
|
+
- content 为字符串时直接作为文本;为列表时只取 `type=text` 项;
|
|
37
|
+
- role 可在 message 内或顶层;`toolResult` / `tool_result` 统一归一为 `tool`;
|
|
38
|
+
- usage 可在 message 内或顶层;
|
|
39
|
+
- 坏行(非法 JSON / 非对象)跳过并计入 `invalid`,不中断检索;
|
|
40
|
+
- 非 `.jsonl` 文件(如 `sessions.json`、`notes.txt`)不会被当作会话。
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
# 安全边界(security)
|
|
2
|
+
|
|
3
|
+
## 只读保证
|
|
4
|
+
|
|
5
|
+
- 引擎只调用读取类操作(`open(path, "r")`),对会话目录不做任何写入 / 修改 / 删除;
|
|
6
|
+
- 检索、提取、统计全程无网络请求;不把日志内容上传到任何服务;
|
|
7
|
+
- 测试含只读回归:跑完 scan / search / session / stats / tools 后目录文件清单与大小不变。
|
|
8
|
+
|
|
9
|
+
## 默认脱敏
|
|
10
|
+
|
|
11
|
+
输出(search 命中片段、session 原文)默认把疑似密钥 / token / 口令打码,`--no-redact` 关闭。覆盖:
|
|
12
|
+
|
|
13
|
+
| 形态 | 示例 |
|
|
14
|
+
|---|---|
|
|
15
|
+
| API key 前缀 | sk-xxx / rk-xxx / pk-xxx |
|
|
16
|
+
| GitHub / Slack token | ghp_xxx / xoxb-xxx |
|
|
17
|
+
| AWS access key | AKIAxxx / ASIAxxx |
|
|
18
|
+
| JWT | eyJ...\\.eyJ...\\.eyJ... |
|
|
19
|
+
| Bearer token | Bearer xxx |
|
|
20
|
+
| URL 口令 | https://user:pass@host → https://user:***@host |
|
|
21
|
+
| 赋值式密钥 | token=xxx / password=xxx / secret=xxx / api_key=xxx |
|
|
22
|
+
| 超长 token | 40+ 位字母数字串 |
|
|
23
|
+
| PEM 私钥 | -----BEGIN ... PRIVATE KEY----- |
|
|
24
|
+
|
|
25
|
+
URL 路径(非凭据)原文保留,方便回溯链接。
|
|
26
|
+
|
|
27
|
+
## 边界
|
|
28
|
+
|
|
29
|
+
- 只检索 `*.jsonl` 会话文件与 `sessions.json` 索引;不读取、不分析其它文件;
|
|
30
|
+
- 检索范围限定用户显式传入的 `--dir`(或环境变量 / 自动定位结果),不主动扫描磁盘;
|
|
31
|
+
- 输出可能包含会话原文中的隐私,默认脱敏且仅用于本机回溯,请勿外传。
|
|
32
|
+
|
|
33
|
+
## 与元忆(yotta-memory)的分工
|
|
34
|
+
|
|
35
|
+
| 维度 | yotta-logs(元史) | yotta-memory(元忆) |
|
|
36
|
+
|---|---|---|
|
|
37
|
+
| 定位 | 原始会话日志(JSONL 事实) | 语义记忆(结构化条目) |
|
|
38
|
+
| 输出 | 原文片段 + 行号 + 时间戳 | 记忆条目 + 权限边界 + 画像 |
|
|
39
|
+
| 写操作 | 无(只读) | 支持(remember / forget / archive) |
|
|
40
|
+
| 权限 | 目录级只读 | 类型 / 属主级权限边界 |
|
|
41
|
+
|
|
42
|
+
回溯「原文」用元史;沉淀「长期知识 / 偏好 / 承诺」用元忆,二者互补。
|
|
@@ -0,0 +1,404 @@
|
|
|
1
|
+
#!/usr/bin/env python3
|
|
2
|
+
# -*- coding: utf-8 -*-
|
|
3
|
+
"""test_yotta_logs.py — 元史(yotta-logs)测试。
|
|
4
|
+
|
|
5
|
+
覆盖:JSONL 解析容错 / 会话发现 / sessions.json 索引 / 角色与文本提取 /
|
|
6
|
+
默认脱敏 / scan / search(关键词·正则·日期·会话·角色·截断)/ session 提取 /
|
|
7
|
+
stats(角色·成本·token·每日汇总)/ tools 排行 / CLI 退出码 / JSON 输出 /
|
|
8
|
+
GBK 控制台 / 只读保证。纯标准库,无 pytest 依赖。
|
|
9
|
+
|
|
10
|
+
运行:python scripts/test_yotta_logs.py
|
|
11
|
+
"""
|
|
12
|
+
import json
|
|
13
|
+
import os
|
|
14
|
+
import subprocess
|
|
15
|
+
import sys
|
|
16
|
+
import tempfile
|
|
17
|
+
from pathlib import Path
|
|
18
|
+
|
|
19
|
+
_HERE = Path(__file__).resolve().parent
|
|
20
|
+
sys.path.insert(0, str(_HERE))
|
|
21
|
+
|
|
22
|
+
import yotta_logs as YL # noqa: E402
|
|
23
|
+
|
|
24
|
+
PASS = 0
|
|
25
|
+
FAIL = 0
|
|
26
|
+
FAILED = []
|
|
27
|
+
|
|
28
|
+
|
|
29
|
+
def check(name, cond, detail=""):
|
|
30
|
+
global PASS, FAIL
|
|
31
|
+
if cond:
|
|
32
|
+
PASS += 1
|
|
33
|
+
else:
|
|
34
|
+
FAIL += 1
|
|
35
|
+
FAILED.append(name)
|
|
36
|
+
print(" FAIL: %s %s" % (name, detail))
|
|
37
|
+
|
|
38
|
+
|
|
39
|
+
def wjsonl(p, rows):
|
|
40
|
+
lines = []
|
|
41
|
+
for r in rows:
|
|
42
|
+
if isinstance(r, str):
|
|
43
|
+
lines.append(r)
|
|
44
|
+
else:
|
|
45
|
+
lines.append(json.dumps(r, ensure_ascii=False))
|
|
46
|
+
p.write_text("\n".join(lines) + "\n", encoding="utf-8")
|
|
47
|
+
|
|
48
|
+
|
|
49
|
+
def build_fixture(base):
|
|
50
|
+
"""构造一个真实形态的会话日志目录,返回目录 Path。"""
|
|
51
|
+
d = base / "sessions"
|
|
52
|
+
d.mkdir(parents=True, exist_ok=True)
|
|
53
|
+
a1 = [
|
|
54
|
+
{"type": "session", "timestamp": "2026-08-26T03:00:00+08:00",
|
|
55
|
+
"session_id": "a1", "title": "部署讨论"},
|
|
56
|
+
{"type": "message", "timestamp": "2026-08-26T03:00:01+08:00",
|
|
57
|
+
"message": {"role": "user", "content": [
|
|
58
|
+
{"type": "text", "text": "你好,部署方案定了吗?"}]}},
|
|
59
|
+
{"type": "message", "timestamp": "2026-08-26T03:00:05+08:00",
|
|
60
|
+
"message": {"role": "assistant", "content": [
|
|
61
|
+
{"type": "text",
|
|
62
|
+
"text": "定了,按灰度发布执行。密钥 sk-abcdef1234567890 已就位。"}],
|
|
63
|
+
"usage": {"cost": {"total": 0.01}, "input_tokens": 100,
|
|
64
|
+
"output_tokens": 50}}},
|
|
65
|
+
{"type": "message", "timestamp": "2026-08-26T03:01:00+08:00",
|
|
66
|
+
"message": {"role": "assistant", "content": [
|
|
67
|
+
{"type": "toolCall", "name": "read_file"},
|
|
68
|
+
{"type": "text", "text": "我读一下配置。"}]}},
|
|
69
|
+
{"type": "message", "timestamp": "2026-08-26T03:02:00+08:00",
|
|
70
|
+
"message": {"role": "toolResult", "content": [
|
|
71
|
+
{"type": "toolResult", "name": "read_file",
|
|
72
|
+
"content": "{\"ok\": true}"}]}},
|
|
73
|
+
"this line is not valid json",
|
|
74
|
+
]
|
|
75
|
+
b2 = [
|
|
76
|
+
{"type": "message", "timestamp": "2026-08-27T10:00:00+08:00",
|
|
77
|
+
"message": {"role": "user", "content": [
|
|
78
|
+
{"type": "text", "text": "CI 又失败了,看下日志。"}]}},
|
|
79
|
+
{"type": "message", "timestamp": "2026-08-27T10:01:00+08:00",
|
|
80
|
+
"message": {"role": "assistant", "content": [
|
|
81
|
+
{"type": "text",
|
|
82
|
+
"text": "好的,我去查。Bearer abcDEF123ghiJKL789 拿来用。"}]}},
|
|
83
|
+
{"type": "message", "timestamp": "2026-08-27T10:02:00+08:00",
|
|
84
|
+
"message": {"role": "assistant", "content": [
|
|
85
|
+
{"type": "toolCall", "name": "run_shell"},
|
|
86
|
+
{"type": "text", "text": "执行命令。"}]}},
|
|
87
|
+
{"type": "message", "timestamp": "2026-08-27T10:03:00+08:00",
|
|
88
|
+
"message": {"role": "user", "content": [
|
|
89
|
+
{"type": "text", "text": "hello, please retry with the new endpoint"}]}},
|
|
90
|
+
]
|
|
91
|
+
wjsonl(d / "a1.jsonl", a1)
|
|
92
|
+
wjsonl(d / "b2.jsonl", b2)
|
|
93
|
+
(d / "sessions.json").write_text(
|
|
94
|
+
json.dumps({"微信-部署": "a1", "ci-排查": "b2"}, ensure_ascii=False),
|
|
95
|
+
encoding="utf-8")
|
|
96
|
+
(d / "notes.txt").write_text("not a session file", encoding="utf-8")
|
|
97
|
+
return d
|
|
98
|
+
|
|
99
|
+
|
|
100
|
+
def test_parse_jsonl():
|
|
101
|
+
with tempfile.TemporaryDirectory() as td:
|
|
102
|
+
p = Path(td) / "x.jsonl"
|
|
103
|
+
p.write_text('{"a":1}\nnot json\n[1,2,3]\n{"b":2}\n',
|
|
104
|
+
encoding="utf-8")
|
|
105
|
+
records, invalid = YL.parse_jsonl(p)
|
|
106
|
+
check("parse_jsonl 记录数", len(records) == 2, "got %d" % len(records))
|
|
107
|
+
check("parse_jsonl 无效行计数", invalid == 2, "got %d" % invalid)
|
|
108
|
+
|
|
109
|
+
|
|
110
|
+
def test_list_sessions():
|
|
111
|
+
with tempfile.TemporaryDirectory() as td:
|
|
112
|
+
d = Path(td)
|
|
113
|
+
(d / "a.jsonl").write_text("\n", encoding="utf-8")
|
|
114
|
+
(d / "b.jsonl").write_text("\n", encoding="utf-8")
|
|
115
|
+
(d / "c.txt").write_text("x", encoding="utf-8")
|
|
116
|
+
(d / "sessions.json").write_text("{}", encoding="utf-8")
|
|
117
|
+
sess = YL.list_sessions(d)
|
|
118
|
+
check("list_sessions 只收 jsonl", len(sess) == 2, "got %s" % sess)
|
|
119
|
+
check("list_sessions 会话 ID = 文件名主干",
|
|
120
|
+
{s["session"] for s in sess} == {"a", "b"})
|
|
121
|
+
|
|
122
|
+
|
|
123
|
+
def test_load_index():
|
|
124
|
+
with tempfile.TemporaryDirectory() as td:
|
|
125
|
+
d = Path(td)
|
|
126
|
+
(d / "sessions.json").write_text(
|
|
127
|
+
json.dumps({"k1": "s1", "k2": "s2"}), encoding="utf-8")
|
|
128
|
+
idx = YL.load_index(d)
|
|
129
|
+
check("load_index dict 形态", idx == {"k1": "s1", "k2": "s2"}, str(idx))
|
|
130
|
+
(d / "sessions.json").write_text(
|
|
131
|
+
json.dumps([{"key": "k1", "sessionId": "s1"}]), encoding="utf-8")
|
|
132
|
+
idx = YL.load_index(d)
|
|
133
|
+
check("load_index list 形态", idx == {"k1": "s1"}, str(idx))
|
|
134
|
+
(d / "sessions.json").write_text("not json", encoding="utf-8")
|
|
135
|
+
check("load_index 坏文件容错", YL.load_index(d) == {})
|
|
136
|
+
|
|
137
|
+
|
|
138
|
+
def test_rec_parsing():
|
|
139
|
+
rec = {"type": "message", "timestamp": "2026-08-26T03:00:00Z",
|
|
140
|
+
"message": {"role": "user", "content": "直接字符串"}}
|
|
141
|
+
check("_rec_ts 顶层", YL._rec_ts(rec) == "2026-08-26T03:00:00Z")
|
|
142
|
+
check("_rec_role user", YL._rec_role(rec) == "user")
|
|
143
|
+
check("_rec_text 字符串", YL._rec_text(rec) == "直接字符串")
|
|
144
|
+
|
|
145
|
+
rec2 = {"type": "message", "message": {"role": "toolResult", "content": [
|
|
146
|
+
{"type": "text", "text": "A"}, {"type": "thinking", "text": "隐藏"},
|
|
147
|
+
{"type": "toolCall", "name": "run_shell"}]}}
|
|
148
|
+
check("_rec_role toolResult 归一为 tool", YL._rec_role(rec2) == "tool",
|
|
149
|
+
YL._rec_role(rec2))
|
|
150
|
+
check("_rec_text 只取 text", YL._rec_text(rec2) == "A",
|
|
151
|
+
repr(YL._rec_text(rec2)))
|
|
152
|
+
check("_rec_tool_names", YL._rec_tool_names(rec2) == ["run_shell"])
|
|
153
|
+
|
|
154
|
+
rec3 = {"type": "message", "message": {"role": "assistant", "content": [
|
|
155
|
+
{"type": "text", "text": "x"}]},
|
|
156
|
+
"usage": {"cost": {"total": 0.5}, "input_tokens": 10,
|
|
157
|
+
"output_tokens": 20}}
|
|
158
|
+
check("_rec_cost", YL._rec_cost(rec3) == 0.5)
|
|
159
|
+
check("_rec_tokens", YL._rec_tokens(rec3) == (10, 20))
|
|
160
|
+
check("_is_message 排除 session 元数据",
|
|
161
|
+
YL._is_message({"type": "session", "role": "session"}) is False)
|
|
162
|
+
|
|
163
|
+
|
|
164
|
+
def test_redact():
|
|
165
|
+
check("redact sk-", "sk-" not in YL.redact("密钥 sk-abcdef1234567890 已就位"))
|
|
166
|
+
check("redact ghp_", "ghp_" not in YL.redact("token ghp_ABCDEFGHIJKLMNOPQRST"))
|
|
167
|
+
check("redact AKIA", "AKIA" not in YL.redact("AKIA1234567890ABCDEF"))
|
|
168
|
+
check("redact JWT",
|
|
169
|
+
"eyJ" not in YL.redact("eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dozjgNryP4J3jVmNHl0w5N_XgL0n3I9PlFUP0THsR8U"))
|
|
170
|
+
check("redact Bearer",
|
|
171
|
+
YL.redact("Bearer abcDEF123ghiJKL789") == "Bearer ***")
|
|
172
|
+
check("redact URL 口令",
|
|
173
|
+
YL.redact("https://user:pass@example.com/path")
|
|
174
|
+
== "https://user:***@example.com/path")
|
|
175
|
+
check("redact 赋值",
|
|
176
|
+
"token=***" in YL.redact("token=sk-abcdef1234567890x"))
|
|
177
|
+
check("redact 长串",
|
|
178
|
+
"***" in YL.redact("abcdefghijklmnopqrstuvwxyz0123456789ABCDEFGH"))
|
|
179
|
+
check("redact URL 路径保留",
|
|
180
|
+
"https://example.com/api/v1/items" in YL.redact("看 https://example.com/api/v1/items 这里"))
|
|
181
|
+
check("redact 普通中文不动",
|
|
182
|
+
YL.redact("你好,今天天气不错。") == "你好,今天天气不错。")
|
|
183
|
+
check("redact PEM",
|
|
184
|
+
"PRIVATE KEY REDACTED" in YL.redact(
|
|
185
|
+
"-----BEGIN RSA PRIVATE KEY-----\nabc\n-----END RSA PRIVATE KEY-----"))
|
|
186
|
+
|
|
187
|
+
|
|
188
|
+
def test_scan(fx):
|
|
189
|
+
res = YL.scan_sessions(str(fx))
|
|
190
|
+
check("scan 会话数 2", res["total_sessions"] == 2,
|
|
191
|
+
str(res["rows"]))
|
|
192
|
+
check("scan 消息合计 8", res["total_messages"] == 8,
|
|
193
|
+
str(res["total_messages"]))
|
|
194
|
+
check("scan 无效行 1", res["total_invalid"] == 1)
|
|
195
|
+
by = {r["session"]: r for r in res["rows"]}
|
|
196
|
+
check("scan a1 消息 4", by["a1"]["messages"] == 4)
|
|
197
|
+
check("scan b2 日期", by["b2"]["date"] == "2026-08-27")
|
|
198
|
+
check("scan 别名映射", by["a1"]["alias"] == "微信-部署")
|
|
199
|
+
|
|
200
|
+
|
|
201
|
+
def test_search(fx):
|
|
202
|
+
r = YL.search_sessions(str(fx), "部署")
|
|
203
|
+
check("search 关键词命中", len(r["matches"]) == 1
|
|
204
|
+
and r["matches"][0]["session"] == "a1", str(r))
|
|
205
|
+
r = YL.search_sessions(str(fx), "HELLO")
|
|
206
|
+
check("search 不区分大小写", len(r["matches"]) == 1
|
|
207
|
+
and r["matches"][0]["session"] == "b2", str(r))
|
|
208
|
+
r = YL.search_sessions(str(fx), r"CI \w+", regex=True)
|
|
209
|
+
check("search 正则命中", len(r["matches"]) == 1
|
|
210
|
+
and r["matches"][0]["session"] == "b2", str(r))
|
|
211
|
+
r = YL.search_sessions(str(fx), "看下", date="2026-08-27")
|
|
212
|
+
check("search 日期过滤", len(r["matches"]) == 1
|
|
213
|
+
and r["matches"][0]["session"] == "b2", str(r))
|
|
214
|
+
r = YL.search_sessions(str(fx), "灰度", sessions=["微信-部署"])
|
|
215
|
+
check("search 会话别名过滤", len(r["matches"]) == 1
|
|
216
|
+
and r["matches"][0]["session"] == "a1", str(r))
|
|
217
|
+
r = YL.search_sessions(str(fx), "灰度", sessions=["b2"])
|
|
218
|
+
check("search 会话 ID 过滤无命中", r["matches"] == [], str(r))
|
|
219
|
+
r = YL.search_sessions(str(fx), "了", role="assistant")
|
|
220
|
+
check("search 角色过滤", r["matches"] and all(
|
|
221
|
+
m["role"] == "assistant" for m in r["matches"]), str(r))
|
|
222
|
+
r = YL.search_sessions(str(fx), "了", limit=1)
|
|
223
|
+
check("search limit 截断", r["truncated"] is True and len(r["matches"]) == 1
|
|
224
|
+
and r["sessions_hit"] == 1, str(r))
|
|
225
|
+
r = YL.search_sessions(str(fx), "绝不存在的词xyz")
|
|
226
|
+
check("search 无命中空列表", r["matches"] == [] and r["sessions_hit"] == 0)
|
|
227
|
+
r = YL.search_sessions(str(fx), "sk-abcdef1234567890")
|
|
228
|
+
check("search 命中脱敏打码", r["matches"] and
|
|
229
|
+
"sk-" not in r["matches"][0]["text"] and "***" in r["matches"][0]["text"],
|
|
230
|
+
str(r["matches"]))
|
|
231
|
+
|
|
232
|
+
|
|
233
|
+
def test_extract(fx):
|
|
234
|
+
r = YL.extract_session(str(fx), "a1")
|
|
235
|
+
check("extract 消息 4", len(r["messages"]) == 4, str(len(r["messages"])))
|
|
236
|
+
check("extract 首条", r["messages"][0]["role"] == "user"
|
|
237
|
+
and "部署方案" in r["messages"][0]["text"])
|
|
238
|
+
check("extract 脱敏", "sk-" not in r["messages"][1]["text"]
|
|
239
|
+
and "***" in r["messages"][1]["text"], repr(r["messages"][1]["text"]))
|
|
240
|
+
r2 = YL.extract_session(str(fx), "a1", role="assistant")
|
|
241
|
+
check("extract 角色过滤", len(r2["messages"]) == 2
|
|
242
|
+
and all(m["role"] == "assistant" for m in r2["messages"]))
|
|
243
|
+
r3 = YL.extract_session(str(fx), "ci-排查")
|
|
244
|
+
check("extract 别名解析", r3["session"] == "b2", str(r3["session"]))
|
|
245
|
+
r4 = YL.extract_session(str(fx), "b2", with_tools=True)
|
|
246
|
+
tools = [t for m in r4["messages"] for t in m["tools"]]
|
|
247
|
+
check("extract 工具标注", "run_shell" in tools, str(tools))
|
|
248
|
+
try:
|
|
249
|
+
YL.extract_session(str(fx), "nope")
|
|
250
|
+
check("extract 未知会话抛错", False)
|
|
251
|
+
except SystemExit:
|
|
252
|
+
check("extract 未知会话抛错", True)
|
|
253
|
+
|
|
254
|
+
|
|
255
|
+
def test_stats(fx):
|
|
256
|
+
r = YL.session_stats(str(fx))
|
|
257
|
+
check("stats 会话 2", r["sessions"] == 2)
|
|
258
|
+
check("stats 消息 8", r["messages"] == 8)
|
|
259
|
+
check("stats 角色分布", r["roles"] == {"user": 3, "assistant": 4, "tool": 1},
|
|
260
|
+
str(r["roles"]))
|
|
261
|
+
check("stats 成本", abs(r["cost"] - 0.01) < 1e-9, str(r["cost"]))
|
|
262
|
+
check("stats token", r["tokens_in"] == 100 and r["tokens_out"] == 50)
|
|
263
|
+
check("stats 首末时间", r["first"].startswith("2026-08-26")
|
|
264
|
+
and r["last"].startswith("2026-08-27"))
|
|
265
|
+
r2 = YL.session_stats(str(fx), daily=True)
|
|
266
|
+
check("stats 每日两天", set(r2["days"].keys()) == {"2026-08-26", "2026-08-27"},
|
|
267
|
+
str(r2["days"].keys()))
|
|
268
|
+
check("stats 每日成本", abs(r2["days"]["2026-08-26"]["cost"] - 0.01) < 1e-9)
|
|
269
|
+
r3 = YL.session_stats(str(fx), session_id="微信-部署")
|
|
270
|
+
check("stats 单会话", r3["sessions"] == 1 and r3["messages"] == 4,
|
|
271
|
+
str((r3["sessions"], r3["messages"])))
|
|
272
|
+
|
|
273
|
+
|
|
274
|
+
def test_tools(fx):
|
|
275
|
+
items = YL.tool_breakdown(str(fx))
|
|
276
|
+
by = dict(items)
|
|
277
|
+
check("tools 排行 read_file 2", by.get("read_file") == 2, str(items))
|
|
278
|
+
check("tools 排行 run_shell 1", by.get("run_shell") == 1, str(items))
|
|
279
|
+
items2 = YL.tool_breakdown(str(fx), session_id="a1")
|
|
280
|
+
by2 = dict(items2)
|
|
281
|
+
check("tools 单会话", by2.get("read_file") == 2 and "run_shell" not in by2,
|
|
282
|
+
str(items2))
|
|
283
|
+
|
|
284
|
+
|
|
285
|
+
def _run(args, inp=None, env=None, cwd=None):
|
|
286
|
+
e = dict(os.environ)
|
|
287
|
+
if env:
|
|
288
|
+
e.update(env)
|
|
289
|
+
return subprocess.run(
|
|
290
|
+
[sys.executable, str(_HERE / "yotta_logs.py")] + args,
|
|
291
|
+
input=inp, capture_output=True, text=True, encoding="utf-8",
|
|
292
|
+
errors="replace", env=e, cwd=cwd)
|
|
293
|
+
|
|
294
|
+
|
|
295
|
+
def test_cli(fx):
|
|
296
|
+
r = _run(["version"])
|
|
297
|
+
check("CLI version", r.returncode == 0 and YL.VERSION in r.stdout,
|
|
298
|
+
"rc=%d" % r.returncode)
|
|
299
|
+
|
|
300
|
+
r = _run(["scan", "--dir", str(fx), "--json"])
|
|
301
|
+
try:
|
|
302
|
+
obj = json.loads(r.stdout)
|
|
303
|
+
check("CLI scan --json", r.returncode == 0
|
|
304
|
+
and obj["total_sessions"] == 2, r.stdout[:120])
|
|
305
|
+
except Exception as e: # noqa: BLE001
|
|
306
|
+
check("CLI scan --json", False, str(e))
|
|
307
|
+
|
|
308
|
+
r = _run(["search", "部署", "--dir", str(fx), "--json"])
|
|
309
|
+
try:
|
|
310
|
+
obj = json.loads(r.stdout)
|
|
311
|
+
check("CLI search --json", r.returncode == 0
|
|
312
|
+
and obj["total_matches"] == 1 and len(obj["matches"]) == 1,
|
|
313
|
+
r.stdout[:120])
|
|
314
|
+
except Exception as e: # noqa: BLE001
|
|
315
|
+
check("CLI search --json", False, str(e))
|
|
316
|
+
|
|
317
|
+
r = _run(["search", "绝不存在的词xyz", "--dir", str(fx)])
|
|
318
|
+
check("CLI search 无命中退出码 1", r.returncode == 1,
|
|
319
|
+
"rc=%d" % r.returncode)
|
|
320
|
+
|
|
321
|
+
r = _run(["session", "a1", "--dir", str(fx)])
|
|
322
|
+
check("CLI session 文本输出", r.returncode == 0 and "部署方案" in r.stdout,
|
|
323
|
+
"rc=%d" % r.returncode)
|
|
324
|
+
|
|
325
|
+
r = _run(["stats", "--dir", str(fx), "--daily"])
|
|
326
|
+
check("CLI stats 每日", r.returncode == 0 and "每日汇总" in r.stdout,
|
|
327
|
+
"rc=%d" % r.returncode)
|
|
328
|
+
|
|
329
|
+
r = _run(["tools", "--dir", str(fx), "--json"])
|
|
330
|
+
try:
|
|
331
|
+
obj = json.loads(r.stdout)
|
|
332
|
+
by = {t["name"]: t["count"] for t in obj["tools"]}
|
|
333
|
+
check("CLI tools --json", r.returncode == 0 and by.get("read_file") == 2,
|
|
334
|
+
r.stdout[:120])
|
|
335
|
+
except Exception as e: # noqa: BLE001
|
|
336
|
+
check("CLI tools --json", False, str(e))
|
|
337
|
+
|
|
338
|
+
r = _run(["scan", "--dir", str(Path(fx).parent / "no-such-dir")])
|
|
339
|
+
check("CLI 目录不存在退出码 4", r.returncode == 4, "rc=%d" % r.returncode)
|
|
340
|
+
|
|
341
|
+
r = _run(["badcmd"])
|
|
342
|
+
check("CLI 未知子命令退出码 4", r.returncode == 4, "rc=%d" % r.returncode)
|
|
343
|
+
|
|
344
|
+
r = _run(["search", "Bearer", "--dir", str(fx), "--no-redact"])
|
|
345
|
+
check("CLI --no-redact 原文保留", r.returncode == 0
|
|
346
|
+
and "abcDEF123ghiJKL789" in r.stdout, repr(r.stdout[:120]))
|
|
347
|
+
|
|
348
|
+
r = _run(["scan"], env={"YOTTA_LOGS_DIR": str(fx)})
|
|
349
|
+
check("CLI YOTTA_LOGS_DIR 环境变量", r.returncode == 0
|
|
350
|
+
and "会话 2 个" in r.stdout, "rc=%d" % r.returncode)
|
|
351
|
+
|
|
352
|
+
|
|
353
|
+
def test_gbk_console(fx):
|
|
354
|
+
env = dict(os.environ)
|
|
355
|
+
env["PYTHONIOENCODING"] = "gbk"
|
|
356
|
+
r = subprocess.run(
|
|
357
|
+
[sys.executable, str(_HERE / "yotta_logs.py"),
|
|
358
|
+
"search", "部署", "--dir", str(fx)],
|
|
359
|
+
capture_output=True, text=True, encoding="gbk", errors="replace",
|
|
360
|
+
env=env)
|
|
361
|
+
check("GBK 控制台中文输出不炸", r.returncode == 0,
|
|
362
|
+
"rc=%d err=%r" % (r.returncode, r.stderr[:100]))
|
|
363
|
+
|
|
364
|
+
|
|
365
|
+
def test_readonly(fx):
|
|
366
|
+
before = sorted((p.name, p.stat().st_size) for p in fx.iterdir())
|
|
367
|
+
YL.scan_sessions(str(fx))
|
|
368
|
+
YL.search_sessions(str(fx), "部署")
|
|
369
|
+
YL.extract_session(str(fx), "a1")
|
|
370
|
+
YL.session_stats(str(fx), daily=True)
|
|
371
|
+
YL.tool_breakdown(str(fx))
|
|
372
|
+
after = sorted((p.name, p.stat().st_size) for p in fx.iterdir())
|
|
373
|
+
check("只读保证:目录内容不变", before == after,
|
|
374
|
+
"before=%s after=%s" % (before, after))
|
|
375
|
+
|
|
376
|
+
|
|
377
|
+
def main():
|
|
378
|
+
print("元史(yotta-logs)测试开始…")
|
|
379
|
+
with tempfile.TemporaryDirectory() as td:
|
|
380
|
+
fx = build_fixture(Path(td))
|
|
381
|
+
test_parse_jsonl()
|
|
382
|
+
test_list_sessions()
|
|
383
|
+
test_load_index()
|
|
384
|
+
test_rec_parsing()
|
|
385
|
+
test_redact()
|
|
386
|
+
test_scan(fx)
|
|
387
|
+
test_search(fx)
|
|
388
|
+
test_extract(fx)
|
|
389
|
+
test_stats(fx)
|
|
390
|
+
test_tools(fx)
|
|
391
|
+
test_cli(fx)
|
|
392
|
+
test_gbk_console(fx)
|
|
393
|
+
test_readonly(fx)
|
|
394
|
+
print("")
|
|
395
|
+
print("通过 %d 项,失败 %d 项" % (PASS, FAIL))
|
|
396
|
+
if FAILED:
|
|
397
|
+
print("失败清单:")
|
|
398
|
+
for name in FAILED:
|
|
399
|
+
print(" - " + name)
|
|
400
|
+
sys.exit(1 if FAIL else 0)
|
|
401
|
+
|
|
402
|
+
|
|
403
|
+
if __name__ == "__main__":
|
|
404
|
+
main()
|