vmware-debug 1.6.1__tar.gz → 1.8.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/PKG-INFO +43 -2
- vmware_debug-1.8.0/README-CN.md +79 -0
- vmware_debug-1.8.0/README.md +75 -0
- vmware_debug-1.8.0/RELEASE_NOTES.md +90 -0
- vmware_debug-1.8.0/mcp_server/server.py +150 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/pyproject.toml +2 -2
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/server.json +3 -3
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/skills/vmware-debug/SKILL.md +6 -2
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/skills/vmware-debug/references/capabilities.md +7 -1
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/skills/vmware-debug/references/setup-guide.md +8 -1
- vmware_debug-1.8.0/tests/eval/regression/test_declared_environment.py +181 -0
- vmware_debug-1.8.0/tests/eval/regression/test_read_only_mode.py +158 -0
- vmware_debug-1.8.0/tests/eval/regression/test_result_envelope.py +55 -0
- vmware_debug-1.8.0/uv.lock +1274 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/vmware_debug/__init__.py +1 -1
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/vmware_debug/cli.py +1 -1
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/vmware_debug/mcp/tools.py +11 -3
- vmware_debug-1.6.1/README-CN.md +0 -45
- vmware_debug-1.6.1/README.md +0 -34
- vmware_debug-1.6.1/RELEASE_NOTES.md +0 -26
- vmware_debug-1.6.1/mcp_server/server.py +0 -78
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/.gitignore +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/SECURITY.md +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/mcp_server/__init__.py +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/skills/vmware-debug/references/cli-reference.md +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/skills/vmware-debug/references/event-envelope.md +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/skills/vmware-debug/references/routing.md +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/tests/eval/regression/__init__.py +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/tests/eval/regression/test_debug_regressions.py +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/tests/test_timeline.py +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/vmware_debug/envelope.py +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/vmware_debug/mcp/__init__.py +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/vmware_debug/ops/__init__.py +0 -0
- {vmware_debug-1.6.1 → vmware_debug-1.8.0}/vmware_debug/ops/timeline.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: vmware-debug
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.8.0
|
|
4
4
|
Summary: VMware diagnostic brain — read-only incident triage, log/event correlation, and root-cause routing across the VMware skill family
|
|
5
5
|
Author-email: Wei Zhou <wei-wz.zhou@broadcom.com>
|
|
6
6
|
License-Expression: MIT
|
|
@@ -13,7 +13,7 @@ Requires-Python: >=3.10
|
|
|
13
13
|
Requires-Dist: mcp[cli]<2.0,>=1.10
|
|
14
14
|
Requires-Dist: rich<15.0,>=13.0
|
|
15
15
|
Requires-Dist: typer<1.0,>=0.12
|
|
16
|
-
Requires-Dist: vmware-policy<2.0,>=1.
|
|
16
|
+
Requires-Dist: vmware-policy<2.0,>=1.8.0
|
|
17
17
|
Description-Content-Type: text/markdown
|
|
18
18
|
|
|
19
19
|
<!-- mcp-name: io.github.zw008/vmware-debug -->
|
|
@@ -40,6 +40,8 @@ advisor/executor split.
|
|
|
40
40
|
See [`skills/vmware-debug/SKILL.md`](skills/vmware-debug/SKILL.md) for the full
|
|
41
41
|
methodology, the event-envelope contract, and symptom routing.
|
|
42
42
|
|
|
43
|
+
- **Read-only by design — and provable** (v1.8.0): both MCP tools are read, none write; set `VMWARE_READ_ONLY=true` (or the per-skill `VMWARE_DEBUG_READ_ONLY`) and the family read-only gate verifies that at startup instead of taking the docs' word for it — env vars are the only switch here, this skill has no config file. See [Read-Only Mode](#read-only-mode).
|
|
44
|
+
|
|
43
45
|
## MCP tools
|
|
44
46
|
|
|
45
47
|
| Tool | What |
|
|
@@ -47,6 +49,45 @@ methodology, the event-envelope contract, and symptom routing.
|
|
|
47
49
|
| `incident_timeline` | [READ] Correlate pre-fetched events → timeline + spikes + ranked hypotheses + next-check ideas |
|
|
48
50
|
| `list_symptom_categories` | [READ] List recognised symptom categories + what to check for each |
|
|
49
51
|
|
|
52
|
+
## Read-Only Mode
|
|
53
|
+
|
|
54
|
+
vmware-debug is read-only by design — both MCP tools carry the `[READ]` marker, take no
|
|
55
|
+
credentials, and make no network calls at all; they only correlate event dicts the
|
|
56
|
+
calling agent has already fetched with the other skills' read tools. Since v1.8.0 that
|
|
57
|
+
is **provable rather than merely documented**: set `VMWARE_READ_ONLY=true` and the
|
|
58
|
+
family read-only gate enumerates the registry at startup and verifies that zero write
|
|
59
|
+
tools are exposed — structural, not a prompt instruction a model can ignore. **Off by
|
|
60
|
+
default.** Fail-closed: if the mode is requested but cannot be guaranteed, the server
|
|
61
|
+
refuses to start rather than running open.
|
|
62
|
+
|
|
63
|
+
The same variable is family-wide: one env var also strips every write tool from the
|
|
64
|
+
write-capable siblings (aiops, storage, vks, nsx, ...), so a whole-estate read-only
|
|
65
|
+
posture is a single setting.
|
|
66
|
+
|
|
67
|
+
```json
|
|
68
|
+
{
|
|
69
|
+
"mcpServers": {
|
|
70
|
+
"vmware-debug": {
|
|
71
|
+
"command": "vmware-debug",
|
|
72
|
+
"args": ["mcp"],
|
|
73
|
+
"env": {
|
|
74
|
+
"VMWARE_READ_ONLY": "true"
|
|
75
|
+
}
|
|
76
|
+
}
|
|
77
|
+
}
|
|
78
|
+
}
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
- **Per-skill override**: `VMWARE_DEBUG_READ_ONLY` beats the family-wide
|
|
82
|
+
`VMWARE_READ_ONLY`. vmware-debug has no `config.yaml`, so the env vars are the only
|
|
83
|
+
switch. Precedence: per-skill env → family env → off.
|
|
84
|
+
- **Classification**: this skill registers its tools through a `build_server()` factory,
|
|
85
|
+
so the gate classifies from the `[READ]`/`[WRITE]` docstring marker rather than from
|
|
86
|
+
MCP annotations. Anything not provably read-only is treated as a write.
|
|
87
|
+
- **Startup log**: nothing is logged as withheld because nothing is — the gate's empty
|
|
88
|
+
result *is* the assertion (write-capable siblings log
|
|
89
|
+
`Read-only mode active ... withheld N write tool(s)` instead).
|
|
90
|
+
|
|
50
91
|
## License
|
|
51
92
|
|
|
52
93
|
MIT.
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
<!-- mcp-name: io.github.zw008/vmware-debug -->
|
|
2
|
+
|
|
3
|
+
# VMware Debug(中文)
|
|
4
|
+
|
|
5
|
+
> **声明**:本项目为社区维护的开源项目,**与 VMware, Inc. 或 Broadcom Inc. 无任何隶属、
|
|
6
|
+
> 背书或赞助关系。** "VMware"、"vSphere" 为 Broadcom 商标。源码以 MIT 许可证公开可审计。
|
|
7
|
+
|
|
8
|
+
VMware skill 家族的**诊断大脑**。你给出症状(报错、日志、变慢的 VM),它来跑系统化排查:
|
|
9
|
+
把其它 skill 取到的事件关联成一条时间线、检测突刺、给根因假设排序,并告诉你下一步该查什么。
|
|
10
|
+
**只读**——从不修改任何东西,也从不执行修复。修复一律路由给 vmware-aiops(单步)或
|
|
11
|
+
vmware-pilot(多步、带审批门控),完全复刻 vmware-harden → vmware-pilot 的「顾问/执行」分工。
|
|
12
|
+
|
|
13
|
+
- **设计上只读,且可证明**(v1.8.0)—— 2 个 MCP 工具全为只读、零写工具;设置 `VMWARE_READ_ONLY=true`(或按 skill 的 `VMWARE_DEBUG_READ_ONLY`),家族只读闸门会在启动时验证这一点,而不是让你相信文档——本 skill 没有配置文件,环境变量是唯一开关,详见[只读模式](#只读模式)
|
|
14
|
+
|
|
15
|
+
## 配套 Skill
|
|
16
|
+
|
|
17
|
+
| 需求 | Skill |
|
|
18
|
+
|---|---|
|
|
19
|
+
| 故障关联 / 根因 | **vmware-debug**(本项目) |
|
|
20
|
+
| 集中日志检索 | vmware-log-insight(把 `log_search` 结果喂给它) |
|
|
21
|
+
| vCenter 事件与告警 | vmware-monitor |
|
|
22
|
+
| 指标 / 异常 | vmware-aria |
|
|
23
|
+
| 执行修复 | vmware-aiops(单步)/ vmware-pilot(多步门控) |
|
|
24
|
+
|
|
25
|
+
## 安装
|
|
26
|
+
|
|
27
|
+
```bash
|
|
28
|
+
uv tool install vmware-debug
|
|
29
|
+
vmware-debug categories # 看它能诊断哪些症状类别
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
## MCP 工具(2 个,全只读)
|
|
33
|
+
|
|
34
|
+
- `incident_timeline`:把已取到的事件关联成 时间线 + 突刺 + 排序后的根因假设 + 下一步检查建议
|
|
35
|
+
- `list_symptom_categories`:症状类别及对应的排查路由(不知道查什么时用它)
|
|
36
|
+
|
|
37
|
+
**事件信封**:`{ts, source, severity, entity, text, fields}`。agent 把各源事件归一成此形状再交给
|
|
38
|
+
debug;debug 因此与其它包零运行时依赖。
|
|
39
|
+
|
|
40
|
+
## 只读模式
|
|
41
|
+
|
|
42
|
+
vmware-debug 结构上就是只读的——两个 MCP 工具都带 `[READ]` 标记,不接收任何凭据,也完全不发起
|
|
43
|
+
网络调用;它们只是对 agent 用其它 skill 的读工具**已经取到**的事件字典做关联分析。自 v1.8.0 起,
|
|
44
|
+
这一点从「文档承诺」变成了**可证明的事实**:设置 `VMWARE_READ_ONLY=true`,家族只读门控会在启动时
|
|
45
|
+
枚举注册表并验证暴露的写工具数为零——这是结构性保证,而非模型可以无视的提示词约束。**默认关闭**;
|
|
46
|
+
且为 fail-closed 设计:请求了只读模式但无法保证时,服务器直接拒绝启动,而不是敞开运行。
|
|
47
|
+
|
|
48
|
+
同一个变量是家族级的:一个环境变量同时会把有写能力的兄弟 skill(aiops、storage、vks、nsx 等)的
|
|
49
|
+
全部写工具剥离,因此「整个环境切只读」只需一处设置。
|
|
50
|
+
|
|
51
|
+
```json
|
|
52
|
+
{
|
|
53
|
+
"mcpServers": {
|
|
54
|
+
"vmware-debug": {
|
|
55
|
+
"command": "vmware-debug",
|
|
56
|
+
"args": ["mcp"],
|
|
57
|
+
"env": {
|
|
58
|
+
"VMWARE_READ_ONLY": "true"
|
|
59
|
+
}
|
|
60
|
+
}
|
|
61
|
+
}
|
|
62
|
+
}
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
- **按 skill 覆盖**:`VMWARE_DEBUG_READ_ONLY` 优先于家族级 `VMWARE_READ_ONLY`。vmware-debug
|
|
66
|
+
没有 `config.yaml`,因此环境变量是唯一的开关。优先级:按 skill 环境变量 → 家族环境变量 → 默认关闭。
|
|
67
|
+
- **分类依据**:本 skill 通过 `build_server()` 工厂注册工具,不传 annotations,因此门控依据
|
|
68
|
+
`[READ]`/`[WRITE]` docstring 标记分类。凡是无法证明为只读的,一律按写工具处理。
|
|
69
|
+
- **启动日志**:不会打印任何「被移除」的工具,因为确实一个都没有——门控返回空结果本身就是这个断言
|
|
70
|
+
(有写能力的兄弟 skill 则会打印 `Read-only mode active ... withheld N write tool(s)`)。
|
|
71
|
+
|
|
72
|
+
## 安全
|
|
73
|
+
|
|
74
|
+
结构上只读、离线、无凭据:不连任何 vCenter/NSX/Aria,没有可破坏面,也没有秘密可泄露。
|
|
75
|
+
详见 [SECURITY.md](SECURITY.md)。
|
|
76
|
+
|
|
77
|
+
## 许可证
|
|
78
|
+
|
|
79
|
+
MIT。
|
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
<!-- mcp-name: io.github.zw008/vmware-debug -->
|
|
2
|
+
|
|
3
|
+
# VMware Debug
|
|
4
|
+
|
|
5
|
+
> ⚠️ **Work in progress** — the core (event correlation engine, MCP tools, CLI)
|
|
6
|
+
> is built and tested; README, `server.json`, full reference docs, and packaging
|
|
7
|
+
> polish are still landing. Not yet published to PyPI.
|
|
8
|
+
|
|
9
|
+
> **Disclaimer**: Community-maintained open-source project, **not affiliated with,
|
|
10
|
+
> endorsed by, or sponsored by VMware, Inc. or Broadcom Inc.** "VMware" and
|
|
11
|
+
> "vSphere" are trademarks of Broadcom. Source is publicly auditable under the MIT
|
|
12
|
+
> license.
|
|
13
|
+
|
|
14
|
+
The diagnostic brain of the VMware skill family. You bring the symptom (an error,
|
|
15
|
+
a log dump, a slow VM); this skill runs a systematic investigation, correlates
|
|
16
|
+
events from the other skills into one timeline, ranks root-cause hypotheses, and
|
|
17
|
+
tells you what to check next. It is **read-only** — it never changes anything and
|
|
18
|
+
never executes fixes. Remediation is routed to `vmware-aiops` (single op) or
|
|
19
|
+
`vmware-pilot` (multi-step, gated), mirroring the `vmware-harden → vmware-pilot`
|
|
20
|
+
advisor/executor split.
|
|
21
|
+
|
|
22
|
+
See [`skills/vmware-debug/SKILL.md`](skills/vmware-debug/SKILL.md) for the full
|
|
23
|
+
methodology, the event-envelope contract, and symptom routing.
|
|
24
|
+
|
|
25
|
+
- **Read-only by design — and provable** (v1.8.0): both MCP tools are read, none write; set `VMWARE_READ_ONLY=true` (or the per-skill `VMWARE_DEBUG_READ_ONLY`) and the family read-only gate verifies that at startup instead of taking the docs' word for it — env vars are the only switch here, this skill has no config file. See [Read-Only Mode](#read-only-mode).
|
|
26
|
+
|
|
27
|
+
## MCP tools
|
|
28
|
+
|
|
29
|
+
| Tool | What |
|
|
30
|
+
|---|---|
|
|
31
|
+
| `incident_timeline` | [READ] Correlate pre-fetched events → timeline + spikes + ranked hypotheses + next-check ideas |
|
|
32
|
+
| `list_symptom_categories` | [READ] List recognised symptom categories + what to check for each |
|
|
33
|
+
|
|
34
|
+
## Read-Only Mode
|
|
35
|
+
|
|
36
|
+
vmware-debug is read-only by design — both MCP tools carry the `[READ]` marker, take no
|
|
37
|
+
credentials, and make no network calls at all; they only correlate event dicts the
|
|
38
|
+
calling agent has already fetched with the other skills' read tools. Since v1.8.0 that
|
|
39
|
+
is **provable rather than merely documented**: set `VMWARE_READ_ONLY=true` and the
|
|
40
|
+
family read-only gate enumerates the registry at startup and verifies that zero write
|
|
41
|
+
tools are exposed — structural, not a prompt instruction a model can ignore. **Off by
|
|
42
|
+
default.** Fail-closed: if the mode is requested but cannot be guaranteed, the server
|
|
43
|
+
refuses to start rather than running open.
|
|
44
|
+
|
|
45
|
+
The same variable is family-wide: one env var also strips every write tool from the
|
|
46
|
+
write-capable siblings (aiops, storage, vks, nsx, ...), so a whole-estate read-only
|
|
47
|
+
posture is a single setting.
|
|
48
|
+
|
|
49
|
+
```json
|
|
50
|
+
{
|
|
51
|
+
"mcpServers": {
|
|
52
|
+
"vmware-debug": {
|
|
53
|
+
"command": "vmware-debug",
|
|
54
|
+
"args": ["mcp"],
|
|
55
|
+
"env": {
|
|
56
|
+
"VMWARE_READ_ONLY": "true"
|
|
57
|
+
}
|
|
58
|
+
}
|
|
59
|
+
}
|
|
60
|
+
}
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
- **Per-skill override**: `VMWARE_DEBUG_READ_ONLY` beats the family-wide
|
|
64
|
+
`VMWARE_READ_ONLY`. vmware-debug has no `config.yaml`, so the env vars are the only
|
|
65
|
+
switch. Precedence: per-skill env → family env → off.
|
|
66
|
+
- **Classification**: this skill registers its tools through a `build_server()` factory,
|
|
67
|
+
so the gate classifies from the `[READ]`/`[WRITE]` docstring marker rather than from
|
|
68
|
+
MCP annotations. Anything not provably read-only is treated as a write.
|
|
69
|
+
- **Startup log**: nothing is logged as withheld because nothing is — the gate's empty
|
|
70
|
+
result *is* the assertion (write-capable siblings log
|
|
71
|
+
`Read-only mode active ... withheld N write tool(s)` instead).
|
|
72
|
+
|
|
73
|
+
## License
|
|
74
|
+
|
|
75
|
+
MIT.
|
|
@@ -0,0 +1,90 @@
|
|
|
1
|
+
## v1.8.0 (2026-07-18) — read-only mode, working policy defaults, declared environments
|
|
2
|
+
|
|
3
|
+
Family release driven by [VMware-AIops#31](https://github.com/zw008/VMware-AIops/issues/31),
|
|
4
|
+
where an operator running Llama 3.3 70B (Goose / OpenShift AI, on-prem H100) had to
|
|
5
|
+
hand-write 17 prompt guardrails to make tool calling reliable. A prompt is advisory — a
|
|
6
|
+
model can ignore it. Every guardrail that could move into the harness has.
|
|
7
|
+
|
|
8
|
+
### Added
|
|
9
|
+
- **Read-only mode.** Set `VMWARE_READ_ONLY=true` (or the per-skill
|
|
10
|
+
`VMWARE_DEBUG_READ_ONLY`) and every write tool is removed from the MCP registry at
|
|
11
|
+
start-up. `list_tools()` never offers them, so the model cannot call what it cannot
|
|
12
|
+
see. **Off by default** — nothing changes unless you turn it on. Fail-closed: if the
|
|
13
|
+
mode is requested but cannot be guaranteed, the server refuses to start rather than
|
|
14
|
+
running open. **vmware-debug has no `config.yaml`, so the environment variables are the
|
|
15
|
+
only switch** — precedence is per-skill env → family env → off. Both tools are `[READ]`,
|
|
16
|
+
so the gate withholds nothing here; its empty result *is* the assertion.
|
|
17
|
+
- **`environment:` scoping for policy rules.** Policy rules now scope by the environment
|
|
18
|
+
a target declares (production / staging / lab). Skills that connect to a VMware estate
|
|
19
|
+
declare it per target in their own `config.yaml`; vmware-debug connects to nothing and
|
|
20
|
+
has no config, so it reports a constant `local`.
|
|
21
|
+
|
|
22
|
+
### Added — list results now state whether they are complete
|
|
23
|
+
|
|
24
|
+
Every `[READ]` list tool returns the family envelope instead of a bare array:
|
|
25
|
+
|
|
26
|
+
{"items": [...], "returned": 50, "limit": 50, "total": 213,
|
|
27
|
+
"truncated": true, "hint": "Showing 50 of 213. Raise limit or narrow the query..."}
|
|
28
|
+
|
|
29
|
+
This closes the reported failure where long responses were summarised as "no data
|
|
30
|
+
returned": a bare list gives a model no way to tell a complete answer from page one, so
|
|
31
|
+
it guessed. `truncated: false` now positively states completeness — including when
|
|
32
|
+
`items` is empty, which means "checked, found none", not "the call failed".
|
|
33
|
+
|
|
34
|
+
- **1 tool(s) converted** across ops, MCP and CLI. The routing table is a fixed constant, so `total` is exact and `truncated` is always
|
|
35
|
+
false — there is no page two to go looking for.
|
|
36
|
+
|
|
37
|
+
### Changed — migration, read this
|
|
38
|
+
- **Approval tiers now actually run.** They shipped in v1.6.0 but the engine only ever
|
|
39
|
+
read `~/.vmware/rules.yaml`, and a fresh install has no such file — so every deny rule,
|
|
40
|
+
maintenance window and approval tier had been inert on every install that never
|
|
41
|
+
hand-authored one. A packaged baseline now loads when you have written no rules of your
|
|
42
|
+
own. Writes at medium risk and above are stamped with their tier in the audit log;
|
|
43
|
+
irreversible work and guest execution against a target declared `production` require a
|
|
44
|
+
named approver via `VMWARE_AUDIT_APPROVED_BY`.
|
|
45
|
+
- **`environment:` will become required for writes — but not here.** Across the family, a
|
|
46
|
+
state-changing operation against a target that declares no `environment:` still runs and
|
|
47
|
+
logs a warning, and **the next major release refuses it**. vmware-debug ships no tool
|
|
48
|
+
above read risk, holds no credentials and makes no network calls, so there is nothing to
|
|
49
|
+
migrate — it reports a constant `local` and the upgrade is a no-op for this package.
|
|
50
|
+
Read-only operations are never affected, in this release or the next. Check what applies
|
|
51
|
+
to your write-capable siblings before upgrading:
|
|
52
|
+
`vmware-audit policy --operation vm_delete --env <env>`.
|
|
53
|
+
|
|
54
|
+
### Fixed
|
|
55
|
+
- **Policy glob patterns with a leading wildcard silently matched nothing.** A rule written
|
|
56
|
+
`operations: ["*_delete"]` parsed fine, read correctly, and never fired — only a trailing
|
|
57
|
+
`*` was honoured. Now full glob matching, for operations and environments alike.
|
|
58
|
+
|
|
59
|
+
### Notes
|
|
60
|
+
- Requires `vmware-policy>=1.8.0`; publish that package first.
|
|
61
|
+
- `vmware-audit policy` reports which rules are in force and where they came from —
|
|
62
|
+
including the case where your rules file exists but failed to parse, which previously
|
|
63
|
+
looked identical to "policy is working".
|
|
64
|
+
|
|
65
|
+
## v1.6.1 (2026-06-24) — initial release
|
|
66
|
+
|
|
67
|
+
First release of **vmware-debug**: the read-only diagnostic brain of the VMware
|
|
68
|
+
skill family. You bring the symptom; it runs the investigation, correlates
|
|
69
|
+
events from the other skills into one timeline, ranks root-cause hypotheses, and
|
|
70
|
+
routes remediation to vmware-aiops / vmware-pilot. It never writes and never
|
|
71
|
+
executes fixes (advisor/executor split, mirroring vmware-harden → vmware-pilot).
|
|
72
|
+
|
|
73
|
+
### Added
|
|
74
|
+
- **2 read-only MCP tools**: `incident_timeline` (correlate pre-fetched events
|
|
75
|
+
into a timeline + z-score spikes + ranked hypotheses + next-check ideas) and
|
|
76
|
+
`list_symptom_categories` (the symptom→skill routing catalogue).
|
|
77
|
+
- **Unified event envelope** + tolerant normalizer so debug stays source-agnostic
|
|
78
|
+
with zero runtime dependency on the other skill packages — the agent fans out
|
|
79
|
+
to each skill's read tools and feeds events here (avoids cross-skill coupling,
|
|
80
|
+
踩坑 #21/#32).
|
|
81
|
+
- **Pure correlation engine** (timeline merge, time-binning, spike detection,
|
|
82
|
+
hypothesis ranking, symptom routing) — fully unit-tested offline.
|
|
83
|
+
- **Typer CLI**: `triage`, `categories`, `version`, `mcp`. The `mcp` entry point
|
|
84
|
+
needs no network at startup (proxy-safe, 踩坑 #25).
|
|
85
|
+
- SKILL.md + references (event-envelope contract, symptom routing, playbooks).
|
|
86
|
+
|
|
87
|
+
### Notes
|
|
88
|
+
- Read-only by construction; remediation is routed, never executed.
|
|
89
|
+
- `parse_timestamp` rejects implausible/garbage timestamps loudly rather than
|
|
90
|
+
silently landing at the epoch.
|
|
@@ -0,0 +1,150 @@
|
|
|
1
|
+
"""vmware-debug MCP server entry point.
|
|
2
|
+
|
|
3
|
+
Tools are defined in vmware_debug.mcp.tools (so audit logs see skill=debug).
|
|
4
|
+
This module wires them into a FastMCP server and provides the stdio entry point.
|
|
5
|
+
|
|
6
|
+
Note: signatures here use typing.Optional, never PEP 604 ``X | None`` — FastMCP
|
|
7
|
+
reflects these at registration and ``X | None`` crashes on Python 3.10 + older
|
|
8
|
+
mcp/pydantic (CLAUDE.md 踩坑 #33).
|
|
9
|
+
"""
|
|
10
|
+
|
|
11
|
+
import sys
|
|
12
|
+
from typing import Optional
|
|
13
|
+
|
|
14
|
+
from mcp.server.fastmcp import FastMCP
|
|
15
|
+
from vmware_policy import apply_read_only_gate, set_environment_resolver
|
|
16
|
+
|
|
17
|
+
from vmware_debug.mcp import tools as t
|
|
18
|
+
|
|
19
|
+
#: Names withheld by the most recent :func:`build_server` call. The gate runs
|
|
20
|
+
#: inside the factory (this server has no module-level instance), so the result
|
|
21
|
+
#: is recorded here for startup logging and tests.
|
|
22
|
+
WITHHELD_WRITE_TOOLS: list[str] = []
|
|
23
|
+
|
|
24
|
+
|
|
25
|
+
# ---------------------------------------------------------------------------
|
|
26
|
+
# Environment declaration
|
|
27
|
+
# ---------------------------------------------------------------------------
|
|
28
|
+
|
|
29
|
+
#: What this skill reports as the environment of everything it touches.
|
|
30
|
+
#:
|
|
31
|
+
#: Policy rules scope by environment, and the baseline treats a target that
|
|
32
|
+
#: declares none as unknown — today that warns on state-changing operations,
|
|
33
|
+
#: and the next major release refuses them. Every other skill answers this from
|
|
34
|
+
#: its own config, where an operator labels each target ``production`` /
|
|
35
|
+
#: ``staging`` / ``lab``.
|
|
36
|
+
#:
|
|
37
|
+
#: vmware-debug has no such config and no connection to declare one about: its
|
|
38
|
+
#: tools are pure correlation over event dicts the calling agent has already
|
|
39
|
+
#: fetched with other skills' read tools. There is no network access, no
|
|
40
|
+
#: writes, and in fact no @vmware_tool-decorated operation above read risk — so
|
|
41
|
+
#: nothing here is gated under either setting today. Requiring a declaration it
|
|
42
|
+
#: has no place to make would leave it permanently blocked for no gain.
|
|
43
|
+
#:
|
|
44
|
+
#: The constant is registered anyway so that the answer is explicit and stays
|
|
45
|
+
#: true if this skill ever grows a tool that writes to a local store.
|
|
46
|
+
LOCAL_ENVIRONMENT = "local"
|
|
47
|
+
|
|
48
|
+
#: Client-facing behaviour hints, matching the rest of the family. Both tools
|
|
49
|
+
#: are [READ]: pure correlation over dicts the caller already fetched, with the
|
|
50
|
+
#: same answer every time. These drive MCP client UI (e.g. whether a call needs
|
|
51
|
+
#: a confirmation prompt); the read-only gate classifies independently, from
|
|
52
|
+
#: the [READ]/[WRITE] docstring marker.
|
|
53
|
+
#:
|
|
54
|
+
#: ``openWorldHint`` is False rather than the family's usual True: this skill
|
|
55
|
+
#: has no network access at all, which is exactly the closed world the hint
|
|
56
|
+
#: describes. Copying True would contradict both docstrings below.
|
|
57
|
+
_READ = {
|
|
58
|
+
"readOnlyHint": True,
|
|
59
|
+
"destructiveHint": False,
|
|
60
|
+
"idempotentHint": True,
|
|
61
|
+
"openWorldHint": False,
|
|
62
|
+
}
|
|
63
|
+
|
|
64
|
+
|
|
65
|
+
def _environment_for(target: Optional[str]) -> str:
|
|
66
|
+
"""Report the environment for policy scoping. Always ``local`` — see above."""
|
|
67
|
+
return LOCAL_ENVIRONMENT
|
|
68
|
+
|
|
69
|
+
|
|
70
|
+
# Registered at import time rather than inside build_server(): the resolver is
|
|
71
|
+
# process-global state in vmware_policy, not per-server-instance, and every
|
|
72
|
+
# build_server() call would otherwise re-register the same constant.
|
|
73
|
+
set_environment_resolver(_environment_for)
|
|
74
|
+
|
|
75
|
+
|
|
76
|
+
def build_server() -> FastMCP:
|
|
77
|
+
"""Construct and configure the MCP server."""
|
|
78
|
+
server = FastMCP("vmware-debug")
|
|
79
|
+
|
|
80
|
+
@server.tool(name="incident_timeline", annotations=_READ)
|
|
81
|
+
def _incident_timeline_impl(
|
|
82
|
+
events: list[dict],
|
|
83
|
+
bin_seconds: Optional[float] = None,
|
|
84
|
+
z_threshold: float = 2.0,
|
|
85
|
+
top_n: int = 5,
|
|
86
|
+
) -> dict:
|
|
87
|
+
"""[READ] Correlate already-fetched VMware events into one incident view.
|
|
88
|
+
|
|
89
|
+
WHEN: after you've pulled events for an incident from the data-source
|
|
90
|
+
skills (vmware-monitor event_list/alarm_list, vmware-aria alerts/anomaly,
|
|
91
|
+
vmware-log-insight log_search/log_aggregate, vmware-nsx) — feed them all
|
|
92
|
+
here to find what correlates and where to look next. This tool does NOT
|
|
93
|
+
fetch anything itself; it has no vCenter/network access.
|
|
94
|
+
|
|
95
|
+
INPUT: events = list of event envelopes, each {ts, source, severity,
|
|
96
|
+
entity, text, fields} (ts may be ISO-8601, epoch seconds, or millis;
|
|
97
|
+
severity is normalised). Optional: bin_seconds (time-bin width; auto if
|
|
98
|
+
omitted), z_threshold (spike sensitivity, default 2.0), top_n (max
|
|
99
|
+
hypotheses, default 5).
|
|
100
|
+
|
|
101
|
+
RETURNS: {event_count, window, spikes (anomalous time bins), hypotheses
|
|
102
|
+
(ranked root-cause candidates, each with a suggested_check), next_checks
|
|
103
|
+
(concrete ideas for what to investigate next, including which skill/tool
|
|
104
|
+
to run)}.
|
|
105
|
+
|
|
106
|
+
GOTCHAS: read-only and stateless — nothing is executed. Remediation is
|
|
107
|
+
routed to vmware-aiops (single fix) or vmware-pilot (multi-step, gated).
|
|
108
|
+
A malformed event raises ValueError naming its index."""
|
|
109
|
+
return t.incident_timeline(events, bin_seconds, z_threshold, top_n)
|
|
110
|
+
|
|
111
|
+
@server.tool(name="list_symptom_categories", annotations=_READ)
|
|
112
|
+
def _list_symptom_categories_impl() -> dict:
|
|
113
|
+
"""[READ] List the symptom categories vmware-debug recognises and, for
|
|
114
|
+
each, example keywords and the suggested next check (which skill/tool to
|
|
115
|
+
run). Takes no parameters. Use this when you don't yet know what to look
|
|
116
|
+
at — it turns "something's wrong" into concrete investigation steps.
|
|
117
|
+
Returns the family list envelope {items, returned, limit, total,
|
|
118
|
+
truncated, hint}; each item is {category, example_keywords,
|
|
119
|
+
suggested_check}. The routing table is a fixed constant, so truncated is
|
|
120
|
+
always false and total is exact — this is every category there is, not a
|
|
121
|
+
page of them. Read-only; no network access."""
|
|
122
|
+
return t.list_symptom_categories()
|
|
123
|
+
|
|
124
|
+
# Applied after every tool above has registered and before the server is
|
|
125
|
+
# handed out. The [READ]/[WRITE] docstring marker is what the gate reads
|
|
126
|
+
# first, so the readOnlyHint annotations above inform client UI without
|
|
127
|
+
# changing this classification; both tools are [READ] and nothing is
|
|
128
|
+
# withheld. Wired anyway so that stays provable, and so the gate is already
|
|
129
|
+
# in place the day this skill grows a write tool.
|
|
130
|
+
global WITHHELD_WRITE_TOOLS # noqa: PLW0603 — factory has no module instance
|
|
131
|
+
WITHHELD_WRITE_TOOLS = apply_read_only_gate(
|
|
132
|
+
server, "vmware-debug", config_flag=None
|
|
133
|
+
)
|
|
134
|
+
|
|
135
|
+
return server
|
|
136
|
+
|
|
137
|
+
|
|
138
|
+
def main() -> None:
|
|
139
|
+
"""Entry point for `vmware-debug-mcp` (stdio transport)."""
|
|
140
|
+
if sys.version_info < (3, 11):
|
|
141
|
+
sys.exit(
|
|
142
|
+
"vmware-debug-mcp requires Python >= 3.11 (FastMCP schema reflection "
|
|
143
|
+
"is unreliable on 3.10). Reinstall under 3.11+: "
|
|
144
|
+
"uv tool install --python 3.11 vmware-debug"
|
|
145
|
+
)
|
|
146
|
+
build_server().run()
|
|
147
|
+
|
|
148
|
+
|
|
149
|
+
if __name__ == "__main__":
|
|
150
|
+
main()
|
|
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
|
|
|
4
4
|
|
|
5
5
|
[project]
|
|
6
6
|
name = "vmware-debug"
|
|
7
|
-
version = "1.
|
|
7
|
+
version = "1.8.0"
|
|
8
8
|
description = "VMware diagnostic brain — read-only incident triage, log/event correlation, and root-cause routing across the VMware skill family"
|
|
9
9
|
readme = "README.md"
|
|
10
10
|
license = "MIT"
|
|
@@ -21,7 +21,7 @@ dependencies = [
|
|
|
21
21
|
"typer>=0.12,<1.0",
|
|
22
22
|
"rich>=13.0,<15.0",
|
|
23
23
|
"mcp[cli]>=1.10,<2.0",
|
|
24
|
-
"vmware-policy>=1.
|
|
24
|
+
"vmware-policy>=1.8.0,<2.0",
|
|
25
25
|
]
|
|
26
26
|
|
|
27
27
|
[project.scripts]
|
|
@@ -2,17 +2,17 @@
|
|
|
2
2
|
"$schema": "https://static.modelcontextprotocol.io/schemas/2025-12-11/server.schema.json",
|
|
3
3
|
"name": "io.github.zw008/vmware-debug",
|
|
4
4
|
"title": "VMware Debug",
|
|
5
|
-
"description": "
|
|
5
|
+
"description": "Read-only VMware incident correlation: timeline, spikes, root-cause routing. 2 MCP tools.",
|
|
6
6
|
"repository": {
|
|
7
7
|
"url": "https://github.com/zw008/VMware-Debug",
|
|
8
8
|
"source": "github"
|
|
9
9
|
},
|
|
10
|
-
"version": "1.
|
|
10
|
+
"version": "1.8.0",
|
|
11
11
|
"packages": [
|
|
12
12
|
{
|
|
13
13
|
"registryType": "pypi",
|
|
14
14
|
"identifier": "vmware-debug",
|
|
15
|
-
"version": "1.
|
|
15
|
+
"version": "1.8.0",
|
|
16
16
|
"transport": {
|
|
17
17
|
"type": "stdio"
|
|
18
18
|
}
|
|
@@ -19,7 +19,7 @@ installer:
|
|
|
19
19
|
package: vmware-debug
|
|
20
20
|
allowed-tools:
|
|
21
21
|
- Bash
|
|
22
|
-
metadata: {"openclaw":{"requires":{"bins":["vmware-debug"]},"primaryEnv":"NONE"}}
|
|
22
|
+
metadata: {"openclaw":{"requires":{"bins":["vmware-debug"]},"optional":{"env":["VMWARE_READ_ONLY","VMWARE_DEBUG_READ_ONLY","VMWARE_AUDIT_APPROVED_BY","VMWARE_AUDIT_RATIONALE"],"bins":["vmware-policy"]},"primaryEnv":"NONE","homepage":"https://github.com/zw008/VMware-Debug","os":["macos","linux"]}}
|
|
23
23
|
---
|
|
24
24
|
|
|
25
25
|
# VMware Debug
|
|
@@ -107,6 +107,8 @@ a recommended plan.
|
|
|
107
107
|
| `incident_timeline` | [READ] Correlate pre-fetched events → timeline + spikes + ranked hypotheses + next-check ideas |
|
|
108
108
|
| `list_symptom_categories` | [READ] List recognised symptom categories + what to check for each |
|
|
109
109
|
|
|
110
|
+
**List envelope** (output of `list_symptom_categories`): `{items, returned, limit, total, truncated, hint}` — read the rows from `items`. `truncated` is always `false` here, which is the point: it states that the catalogue is complete instead of leaving you to infer it.
|
|
111
|
+
|
|
110
112
|
**Event envelope** (input to `incident_timeline`): `{ts, source, severity, entity, text, fields}`.
|
|
111
113
|
See `references/event-envelope.md`. The agent normalises each source's events into this
|
|
112
114
|
shape; debug stays source-agnostic and has no dependency on the other packages.
|
|
@@ -131,7 +133,9 @@ vmware-debug mcp # start stdio MCP server (proxy-
|
|
|
131
133
|
|
|
132
134
|
Read-only by construction: no write tools, no network, nothing executed. Remediation
|
|
133
135
|
is always routed to aiops/pilot, where the double-confirm / approval / audit gates live
|
|
134
|
-
(audit DB `~/.vmware/audit.db`).
|
|
136
|
+
(audit DB `~/.vmware/audit.db`). Policy rules scope by environment; debug has no config
|
|
137
|
+
and no connection to declare one about, so it reports a constant `local` — nothing here
|
|
138
|
+
touches a remote VMware estate. See `references/setup-guide.md`.
|
|
135
139
|
|
|
136
140
|
## License
|
|
137
141
|
|
|
@@ -5,7 +5,13 @@ Read-only, offline incident correlation. No network, no credentials, no writes.
|
|
|
5
5
|
| Tool | What it returns | Typical response tokens |
|
|
6
6
|
|---|---|---|
|
|
7
7
|
| `incident_timeline` | `{event_count, window, spikes:[{start,end,count,zscore}], hypotheses:[{category, score, summary, evidence_count, first_seen, last_seen, sample_text, suggested_check}], next_checks:[...]}` | 300–2000 (scales with hypotheses) |
|
|
8
|
-
| `list_symptom_categories` | `[{category, example_keywords, suggested_check}]` | ~400 |
|
|
8
|
+
| `list_symptom_categories` | `{items: [{category, example_keywords, suggested_check}], returned, limit, total, truncated, hint}` | ~400 |
|
|
9
|
+
|
|
10
|
+
`list_symptom_categories` returns the family list envelope — read the rows from
|
|
11
|
+
`items`. It has no `limit` parameter, which is exactly why the envelope matters:
|
|
12
|
+
`truncated: false` states that this is every category there is, rather than
|
|
13
|
+
leaving a model to guess whether it is holding page one. The catalogue is a
|
|
14
|
+
fixed in-process constant, so `total` is a real count and `limit` is `null`.
|
|
9
15
|
|
|
10
16
|
## Correlation engine
|
|
11
17
|
|
|
@@ -39,5 +39,12 @@ routes fixes to (vmware-aiops, vmware-pilot).
|
|
|
39
39
|
to vmware-aiops / vmware-pilot, where confirmation/approval/audit live.
|
|
40
40
|
5. **No cross-skill coupling** — events arrive as plain dicts (the event
|
|
41
41
|
envelope); debug imports no other skill package at runtime.
|
|
42
|
-
6. **
|
|
42
|
+
6. **Environment scoping** — policy rules scope by environment, and skills that
|
|
43
|
+
connect to a VMware estate declare `environment:` (`production` / `staging` /
|
|
44
|
+
`lab`) per target in their own `config.yaml`; a target that declares none is
|
|
45
|
+
treated as unknown, and state-changing operations against it currently log a
|
|
46
|
+
warning — the next major release will refuse them. debug has no config and
|
|
47
|
+
no connection to declare one about, so it reports a constant `local`. Since
|
|
48
|
+
it ships no operation above read risk, nothing here is gated either way.
|
|
49
|
+
7. **Static analysis** — `uvx bandit -r vmware_debug/ mcp_server/` (release bar:
|
|
43
50
|
0 Medium+).
|