niceeval 0.13.4-canary.49 → 0.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/INDEX.md +3 -3
- package/dist/inspection/overview.cjs +20 -4
- package/dist/inspection/overview.cjs.map +1 -1
- package/dist/inspection/overview.d.cts +5 -0
- package/dist/inspection/overview.d.mts +5 -0
- package/dist/inspection/overview.d.ts +5 -0
- package/dist/inspection/protocol.d.cts +238 -0
- package/dist/inspection/protocol.d.mts +238 -0
- package/dist/inspection/protocol.d.ts +238 -0
- package/dist/inspection/public.mjs +19 -16
- package/dist/inspection/results.cjs +11 -10
- package/dist/inspection/results.cjs.map +1 -1
- package/dist/inspection/results.d.cts +34 -0
- package/dist/inspection/results.d.mts +34 -0
- package/dist/inspection/results.d.ts +34 -0
- package/dist/show/contribution.cjs +11 -2
- package/dist/show/contribution.cjs.map +1 -1
- package/dist/show/contribution.d.cts +7 -0
- package/dist/show/contribution.d.mts +7 -0
- package/dist/show/contribution.d.ts +7 -0
- package/dist/show/model.cjs +1 -0
- package/dist/show/model.cjs.map +1 -1
- package/dist/show/model.d.cts +2 -0
- package/dist/show/model.d.mts +2 -0
- package/dist/show/model.d.ts +2 -0
- package/dist/show/render.cjs +149 -81
- package/dist/show/render.cjs.map +1 -1
- package/dist/show/render.d.cts +3 -1
- package/dist/show/render.d.mts +3 -1
- package/dist/show/render.d.ts +3 -1
- package/dist/view/app-dist/.vite/manifest.json +18 -18
- package/dist/view/app-dist/assets/{AttemptDialog-zlgkBA66.js → AttemptDialog-76dX1VQq.js} +1 -1
- package/dist/view/app-dist/assets/{AttemptRoute-DN6R4719.js → AttemptRoute-DRHVh2JA.js} +1 -1
- package/dist/view/app-dist/assets/{ResultsPage-BfH9vA6y.js → ResultsPage-9CSI8dvZ.js} +2 -2
- package/dist/view/app-dist/assets/{RunPage-XiJvNLB-.js → RunPage-BUF0REv7.js} +1 -1
- package/dist/view/app-dist/assets/{TableContentView-V8MYLnmk.js → TableContentView-_Sbg5wDr.js} +1 -1
- package/dist/view/app-dist/assets/index-BGr_6pGY.js +19 -0
- package/dist/view/app-dist/assets/{quality-cost-scatter-Bxyj_9oc.js → quality-cost-scatter-B5wDa_F1.js} +1 -1
- package/dist/view/app-dist/assets/useSuspenseQuery-D3W8Uy33.js +1 -0
- package/dist/view/app-dist/index.html +1 -1
- package/dist/view/function-runtime.mjs +35 -20
- package/docs-site/zh/examples/ai-agent-application.mdx +1 -1
- package/docs-site/zh/examples/index.mdx +1 -3
- package/docs-site/zh/explanation/evals.mdx +1 -1
- package/docs-site/zh/explanation/experiment.mdx +3 -1
- package/docs-site/zh/explanation/overview.mdx +6 -4
- package/docs-site/zh/explanation/record.mdx +9 -5
- package/docs-site/zh/explanation/runner.mdx +3 -3
- package/docs-site/zh/index.mdx +4 -8
- package/docs-site/zh/introduction.mdx +2 -2
- package/docs-site/zh/reference/cli.mdx +8 -7
- package/docs-site/zh/troubleshooting/debug-sandbox.mdx +5 -5
- package/docs-site/zh/troubleshooting/debugging.mdx +8 -8
- package/docs-site/zh/troubleshooting/recover-after-kill.mdx +35 -55
- package/docs-site/zh/tutorials/accept-results.mdx +4 -4
- package/docs-site/zh/tutorials/agent-onboarding.mdx +2 -2
- package/docs-site/zh/tutorials/authoring.mdx +1 -1
- package/docs-site/zh/tutorials/compare-runs.mdx +3 -3
- package/docs-site/zh/tutorials/concurrency.mdx +6 -6
- package/docs-site/zh/tutorials/connect-your-agent.mdx +2 -2
- package/docs-site/zh/tutorials/docker-in-docker.mdx +1 -1
- package/docs-site/zh/tutorials/experiments.mdx +1 -1
- package/docs-site/zh/tutorials/install-custom-sandbox-agent.mdx +1 -1
- package/docs-site/zh/tutorials/quickstart.mdx +2 -2
- package/docs-site/zh/tutorials/reporters.mdx +3 -3
- package/docs-site/zh/tutorials/sandbox-agent.mdx +2 -2
- package/docs-site/zh/tutorials/sandbox-reuse.mdx +3 -3
- package/docs-site/zh/tutorials/viewing-results.mdx +10 -11
- package/package.json +1 -1
- package/dist/view/app-dist/assets/index-BSJP_UAR.js +0 -19
- package/dist/view/app-dist/assets/useSuspenseQuery-Bdi9nQF1.js +0 -1
|
@@ -4,7 +4,7 @@ sidebarTitle: "CLI"
|
|
|
4
4
|
description: "用固定的 query、show 和 view 读取 canonical SQLite Record。"
|
|
5
5
|
---
|
|
6
6
|
|
|
7
|
-
运行评估使用 `niceeval exp
|
|
7
|
+
运行评估使用 `niceeval exp`。NiceEval 提供三条读取路径:给 AI、脚本和 CI 的 `niceeval query`,给终端中人读的 `niceeval show`,以及在浏览器中审阅的 `niceeval view`。三者都经固定 Inspection operation 读取一个 `PublicationCutoff` 内的已发布事实;origin Run 仍为 `active` 的 Attempt 也可见。
|
|
8
8
|
|
|
9
9
|
## `niceeval query`
|
|
10
10
|
|
|
@@ -22,7 +22,7 @@ niceeval query run [--record <file>] --request <file|->
|
|
|
22
22
|
|
|
23
23
|
先调用 `discover`,再从返回的 schema 构造 request。不要猜字段,也不要用 SQL、JSON path 或公式拼出另一种查询语言。
|
|
24
24
|
|
|
25
|
-
continuation token 绑定 operation、canonical request、
|
|
25
|
+
continuation token 绑定 operation、canonical request、source identity 和 `PublicationCutoff`。任一绑定变化都会得到 restart correction。
|
|
26
26
|
|
|
27
27
|
## `niceeval show`
|
|
28
28
|
|
|
@@ -38,7 +38,7 @@ niceeval show @<locator> --usage [--record <file>]
|
|
|
38
38
|
niceeval show @<locator> --diff [--record <file>]
|
|
39
39
|
```
|
|
40
40
|
|
|
41
|
-
`show` 是固定的终端排版入口。省略 selector 时显示默认 overview。`--run <run-id>`
|
|
41
|
+
`show` 是固定的终端排版入口。省略 selector 时显示默认 overview。`--run <run-id>` 显示一个精确 Run,可重复传入。`--experiment <experiment-id>` 显示一个精确 Experiment,也可重复传入。
|
|
42
42
|
|
|
43
43
|
`@<locator>` 选择一个 Attempt。它不能与 `--run` 或 `--experiment` 一起使用,后两者也不能一起使用。一个命令最多只能给出一个 Attempt locator。
|
|
44
44
|
|
|
@@ -54,11 +54,11 @@ niceeval show @<locator> --diff [--record <file>]
|
|
|
54
54
|
niceeval view [--run <run-id>...] [--no-open] [--port <port>] [--json]
|
|
55
55
|
```
|
|
56
56
|
|
|
57
|
-
`view` 是固定的第一方本地 Insight。省略 `--run` 时打开默认 overview;重复 `--run`
|
|
57
|
+
`view` 是固定的第一方本地 Insight。省略 `--run` 时打开默认 overview;重复 `--run` 可预选多个 exact Run。它不接受位置参数,也不装载项目页面、主题、组件或 renderer。
|
|
58
58
|
|
|
59
|
-
`--no-open` 不请求操作系统打开浏览器。`--port` 选择 loopback
|
|
59
|
+
`--no-open` 不请求操作系统打开浏览器。`--port` 选择 `127.0.0.1` listener 的端口。SPA assets、本机授权 session 与 generation-bound Inspection endpoint 就绪后,命令向 stdout 写出一次可打开的 loopback URL。传入 `--json` 时改为输出 `niceeval.view/v1` NDJSON 的 `ready` 与 `closed` lifecycle event。启动或运行失败通过退出码与 stderr 反馈。
|
|
60
60
|
|
|
61
|
-
|
|
61
|
+
View 只读取项目 canonical `.niceeval/record.sqlite`。新 publication 出现后,页面提示 refresh;用户确认后才切换到新的 cutoff。它不提供 `--record`、SQL、部署、分享、Report、Page、theme、renderer、route 或 operation 参数。
|
|
62
62
|
|
|
63
63
|
## 选择读取方式
|
|
64
64
|
|
|
@@ -66,7 +66,7 @@ niceeval view [--run <run-id>...] [--no-open] [--port <port>] [--json]
|
|
|
66
66
|
| --- | --- |
|
|
67
67
|
| 让脚本、AI 或 CI 发现并取得稳定 JSON | `niceeval query discover`,再用 `explain` / `run` |
|
|
68
68
|
| 在终端审阅 overview、Run、Experiment 或 Attempt 的固定摘要 | `niceeval show` |
|
|
69
|
-
| 在浏览器审阅 overview 或某个
|
|
69
|
+
| 在浏览器审阅 overview 或某个 Run | `niceeval view` |
|
|
70
70
|
| 把同一份事实交给另一台兼容 runtime | 复制 `.niceeval/record.sqlite`,再传入 `--record` |
|
|
71
71
|
|
|
72
72
|
## CLI flags
|
|
@@ -116,6 +116,7 @@ niceeval view [--run <run-id>...] [--no-open] [--port <port>] [--json]
|
|
|
116
116
|
| `--help` | `niceeval query` | boolean | 打印该命令的帮助。 |
|
|
117
117
|
| `--run` | `niceeval show` | string[] | 显示一个已封口的 Run;可重复传入。 |
|
|
118
118
|
| `--experiment` | `niceeval show` | string[] | 显示一个精确的 Experiment;可重复传入。 |
|
|
119
|
+
| `--all` | `niceeval show` | boolean | 显示紧凑 Results 视图中隐藏的 Attempt。 |
|
|
119
120
|
| `--source` | `niceeval show` | boolean | 显示一个 Attempt 已捕获的 source facts。 |
|
|
120
121
|
| `--execution` | `niceeval show` | boolean | 显示一个 Attempt 的有界 execution outline。 |
|
|
121
122
|
| `--timing` | `niceeval show` | boolean | 显示一个 Attempt 已捕获的 timing activities。 |
|
|
@@ -4,12 +4,12 @@ sidebarTitle: "保留 Sandbox 现场"
|
|
|
4
4
|
description: "用 --keep-sandbox 把失败 Attempt 的 Sandbox 保留成可随时唤醒的现场,用 niceeval sandbox enter 进去手动排查,用 sandbox list / stop 查看和清理。"
|
|
5
5
|
---
|
|
6
6
|
|
|
7
|
-
Sandbox 默认在每个 Attempt
|
|
7
|
+
Sandbox 默认在每个 Attempt 结束后销毁。先在 `niceeval view` 中打开对应 Run 或 Attempt,或通过 `query` 的固定 operation 读取判定、断言、diff、命令和事件流。完整路线见 [Debug 手册](/zh/troubleshooting/debugging)。
|
|
8
8
|
|
|
9
9
|
但有些问题只能进活的环境里看:
|
|
10
10
|
|
|
11
11
|
- **环境起不来**——setup 阶段装依赖失败、agent CLI 启动不了。这时 agent 还没开始跑,事件流是空的,最快的办法是进 Sandbox 手动重跑一遍安装命令。
|
|
12
|
-
- **改动落在 `git diff` 之外**——全局装了什么包、`$HOME` 下写了什么配置、`PATH`
|
|
12
|
+
- **改动落在 `git diff` 之外**——全局装了什么包、`$HOME` 下写了什么配置、`PATH` 实际是什么,已发布的 File Changes 不包含这些内容。
|
|
13
13
|
- **重跑太慢**——冷启动加安装要几分钟,想逐条验证猜测时,留着现场比每次重跑快得多。
|
|
14
14
|
|
|
15
15
|
## 跑的时候保留现场
|
|
@@ -30,13 +30,13 @@ Kept sandboxes (1)
|
|
|
30
30
|
Stop them with: niceeval sandbox stop --all
|
|
31
31
|
```
|
|
32
32
|
|
|
33
|
-
每行给三样东西:Attempt identity、Sandbox 实例 id、进入现场的命令。落盘证据从 `niceeval view
|
|
33
|
+
每行给三样东西:Attempt identity、Sandbox 实例 id、进入现场的命令。落盘证据从 `niceeval view --run <run-id>` 中选择对应 Attempt 查看。保留下来的 Sandbox 不会一直跑着烧资源——Docker 容器停在磁盘上,E2B 微 VM 暂停计费,Vercel 保存文件系统。
|
|
34
34
|
|
|
35
35
|
`niceeval sandbox enter` 会先唤醒再进入,在 workdir 打开 shell。退出 shell 后现场自动回到休眠;想让它保持运行,加 `--leave-running`。进去之后就是这次 Attempt 跑完时的环境,可以手动执行命令、翻文件、复现失败。
|
|
36
36
|
|
|
37
37
|
## 查看和清理
|
|
38
38
|
|
|
39
|
-
保留下来的 Sandbox
|
|
39
|
+
保留下来的 Sandbox 由项目的留存注册表记录,用 `niceeval sandbox` 管理:
|
|
40
40
|
|
|
41
41
|
```bash
|
|
42
42
|
niceeval sandbox list # 列出保留的沙箱和现场状态
|
|
@@ -62,4 +62,4 @@ niceeval sandbox stop --all # 全部销毁
|
|
|
62
62
|
|
|
63
63
|
## 边界
|
|
64
64
|
|
|
65
|
-
保留的 Sandbox
|
|
65
|
+
保留的 Sandbox 只用来排查,不能续跑或重新评分。判定、断言和 diff 等结论仍以已发布事实为准。查看结果的方法见[查看结果](/zh/tutorials/viewing-results)。
|
|
@@ -27,23 +27,23 @@ niceeval show @<attempt-locator> --diff
|
|
|
27
27
|
|
|
28
28
|
| 错误 | 下一步 |
|
|
29
29
|
| --- | --- |
|
|
30
|
-
|
|
|
31
|
-
|
|
|
32
|
-
|
|
|
33
|
-
| `
|
|
30
|
+
| 找不到项目 Record | 确认在项目目录中运行,并检查唯一的 `.niceeval/record.sqlite` 是否存在。 |
|
|
31
|
+
| `--record` 输入被拒绝 | 重新取得完整的当前 SQLite 文件;外部 Record 会经过完整性和领域不变量校验。 |
|
|
32
|
+
| 旧 schema 或数据库无效 | 在原项目用 current NiceEval 重新运行;读取不会迁移、修补或部分打开 Record。 |
|
|
33
|
+
| Run 仍为 `active`,但原进程已终止 | 先用 `niceeval run show <run-id>` 核对,再运行 `niceeval run recover <run-id> --yes`。 |
|
|
34
34
|
|
|
35
|
-
|
|
35
|
+
已发布事实没有手动编辑的修复路径。需要不同事实时创建新的 Run。`niceeval run recover` 只会收口可证明旧 owner 已终止的 Run,并保留已经发布的 Attempt。
|
|
36
36
|
|
|
37
37
|
## 结果不完整或无法比较
|
|
38
38
|
|
|
39
39
|
脚本、AI 或 CI 先运行 `niceeval query discover` 找到正确的固定 operation,再用 `explain` 检查 request,最后运行 `run`。结果会指出 missing、issues、Evidence 与 correction。`query` 只输出机器协议,不用来阅读终端排版。
|
|
40
40
|
|
|
41
|
-
需要浏览器界面时运行 `niceeval view --run <run-id> --no-open`。不要解析 View
|
|
41
|
+
需要浏览器界面时运行 `niceeval view --run <run-id> --no-open`。不要解析 View 页面,也不要读取或查询私有 SQLite 表。
|
|
42
42
|
|
|
43
43
|
比较使用 `runs.compare`。它的 `exact` 和 `paired` 模式会拒绝不兼容的 domain 或 pair,不会悄悄缩小分母。
|
|
44
44
|
|
|
45
|
-
## View
|
|
45
|
+
## View 显示旧结果
|
|
46
46
|
|
|
47
|
-
|
|
47
|
+
View 在一次读取开始时固定 PublicationCutoff。页面提示有新发布内容时,确认刷新即可读取新的闭合结果;刷新失败时,上一份可读结果会保留。
|
|
48
48
|
|
|
49
49
|
正常流程见[查看已封口的运行结果](/zh/tutorials/viewing-results),命令细节见 [CLI](/zh/reference/cli)。
|
|
@@ -1,95 +1,75 @@
|
|
|
1
1
|
---
|
|
2
2
|
title: "运行被强杀后恢复"
|
|
3
|
-
description: "进程被 kill -9、CI
|
|
3
|
+
description: "进程被 kill -9、CI 超时或断电中断后,保留已发布结果、收口失去 owner 的 Run,并收回遗留 Sandbox。"
|
|
4
4
|
---
|
|
5
5
|
|
|
6
|
-
`niceeval exp`
|
|
6
|
+
`niceeval exp` 被 `kill -9`、CI 时限或断电直接中断时,进程无法执行收尾。已经发布的 Attempt 仍在项目内唯一的 `.niceeval/record.sqlite` 中可读;尚未发布的工作不会变成半条公开事实。
|
|
7
7
|
|
|
8
|
-
正常的 Ctrl+C 或 SIGTERM
|
|
8
|
+
正常的 Ctrl+C 或 SIGTERM 走受控中断:NiceEval 发布已经完成的 Attempt,并把其余 slot 收口为中断原因。本页只处理进程没有机会收尾的情况。
|
|
9
9
|
|
|
10
|
-
##
|
|
10
|
+
## 先查看已发布结果
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
先列出本次 Invocation 创建过的 Run。若你有 Invocation ID,用它缩小范围:
|
|
13
13
|
|
|
14
14
|
```bash
|
|
15
|
-
niceeval
|
|
15
|
+
niceeval run list --invocation <invocation-id>
|
|
16
|
+
niceeval run show <run-id>
|
|
17
|
+
niceeval view --run <run-id>
|
|
16
18
|
```
|
|
17
19
|
|
|
18
|
-
|
|
19
|
-
- 被强杀的实验如果留了没做完的收尾,重跑会先补一次实验级 `teardown` 再开始,泄漏不会越积越多。
|
|
20
|
-
- 判定为 `errored` 的 Attempt 不复用,照常重跑。
|
|
21
|
-
- 想全部重来,加 `--rerun all`。
|
|
20
|
+
`active` Run 仍会显示已经发布的 Attempt,未发布的 expected slot 显示为 `pending`。不要根据运行时间、PID 过期或 heartbeat 停止自行判断它可以被接管。
|
|
22
21
|
|
|
23
|
-
|
|
22
|
+
## 收口失去 owner 的 Run
|
|
24
23
|
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
先核对有哪些实例属于已经死掉的运行:
|
|
24
|
+
只有确认原进程已经终止后,才执行恢复:
|
|
28
25
|
|
|
29
26
|
```bash
|
|
30
|
-
niceeval
|
|
27
|
+
niceeval run recover <run-id> --yes
|
|
31
28
|
```
|
|
32
29
|
|
|
33
|
-
|
|
34
|
-
ID PROVIDER OWNER STARTED STATE
|
|
35
|
-
f31b9a02 docker pid 4242@mbp dead 2026-07-20 14:02 orphan
|
|
36
|
-
```
|
|
30
|
+
`run recover` 会验证旧 owner 的精确 process identity。验证通过后,它把 Run 收口为 `interrupted`,并为未发布的 slot 写入 absence reason。它不会删除 Run 或已经发布的 Attempt,也不会让旧 generation 继续写入。
|
|
37
31
|
|
|
38
|
-
|
|
32
|
+
证据不足时命令会拒绝恢复。此时先确认原进程、CI job 和可能的远程 worker 都已停止,再重新检查;不要用超时或“等待太久”代替 owner 证据。
|
|
39
33
|
|
|
40
|
-
|
|
41
|
-
niceeval sandbox prune
|
|
42
|
-
```
|
|
34
|
+
需要重新取得结果时,重新运行原来的 Experiment。新的 Invocation 会创建新的 Run;是否沿用已有 Attempt 由当前 Experiment 的 eligibility 和 `--rerun` policy 决定。
|
|
43
35
|
|
|
44
|
-
|
|
45
|
-
- 多容器题目(Compose)的伴随容器和网络跟随主实例整组列出、整组销毁,不需要再手工 `docker rm` 收尾。
|
|
46
|
-
- 从别的机器创建、无法核对的实例标为 `unverified`,默认不动。确认后用 `niceeval sandbox prune --force`。
|
|
47
|
-
- Vercel Sandbox 无法按元数据核对,到 Provider 的保留期限后自动回收,不需要处理。
|
|
48
|
-
- `--keep-sandbox` 留存的现场不受 `prune` 影响,仍用 `niceeval sandbox stop` 管理。
|
|
36
|
+
## 收回强杀留下的 Sandbox
|
|
49
37
|
|
|
50
|
-
|
|
38
|
+
强杀时正在运行的 Sandbox 不会进入留存注册表。先只读核对孤儿实例:
|
|
51
39
|
|
|
52
|
-
|
|
40
|
+
```bash
|
|
41
|
+
niceeval sandbox list --orphans
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
确认后再收回:
|
|
53
45
|
|
|
54
46
|
```bash
|
|
55
|
-
niceeval
|
|
47
|
+
niceeval sandbox prune
|
|
56
48
|
```
|
|
57
49
|
|
|
58
|
-
-
|
|
59
|
-
-
|
|
60
|
-
-
|
|
50
|
+
- `orphan` 表示同一主机上的 owner 已被证明终止,可以安全销毁。
|
|
51
|
+
- `unverified` 表示来自另一台主机或无法确认的实例,默认不会删除;确认后才使用 `niceeval sandbox prune --force`。
|
|
52
|
+
- Compose 的伴随容器和网络会作为一组列出和销毁。
|
|
53
|
+
- Vercel Sandbox 没有可检索的孤儿通道,按 Provider 保留期限自行回收。
|
|
54
|
+
- 用 `--keep-sandbox` 留下的现场不受 `prune` 影响,仍用 `niceeval sandbox stop` 管理。
|
|
61
55
|
|
|
62
56
|
## 恢复共享状态租约
|
|
63
57
|
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
先只检查当前 owner evidence,不带 token 或确认参数:
|
|
58
|
+
声明 `sharedState.key` 的 Experiment 在原 owner 正常释放租约,或操作员完成显式恢复前不会启动相同 key 的新 Invocation。先运行不带 token 的检查命令,读取当前 owner evidence:
|
|
67
59
|
|
|
68
60
|
```bash
|
|
69
|
-
niceeval exp
|
|
70
|
-
--recover-shared-state
|
|
61
|
+
niceeval exp <experiment> --teardown \
|
|
62
|
+
--recover-shared-state <key>
|
|
71
63
|
```
|
|
72
64
|
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
核对完后,复制刚显示的完整 owner token,并显式确认两个条件:
|
|
65
|
+
确认原 owner 已终止,并确认它操作的远程服务和 checkpoint 都已静止后,复制命令显示的 token 并显式确认:
|
|
76
66
|
|
|
77
67
|
```bash
|
|
78
|
-
niceeval exp
|
|
79
|
-
--recover-shared-state
|
|
68
|
+
niceeval exp <experiment> --teardown \
|
|
69
|
+
--recover-shared-state <key> \
|
|
80
70
|
--owner-token <displayed-token> \
|
|
81
71
|
--confirm-owner-terminated \
|
|
82
72
|
--confirm-remote-quiesced
|
|
83
73
|
```
|
|
84
74
|
|
|
85
|
-
|
|
86
|
-
- 错误或过期的 token 不会修改租约,也不会释放后来 owner 的租约。
|
|
87
|
-
- 两条恢复命令同时竞争时,只有一条能取得恢复 ownership。
|
|
88
|
-
- 补偿 `teardown` 失败时,租约继续保留。解决 teardown 问题后重新检查,再用当前显示的 token 重试。
|
|
89
|
-
- 成功后,同 key 的等待 Invocation 才能取得租约并开始 Experiment setup 或创建 Sandbox。
|
|
90
|
-
|
|
91
|
-
## 预防:让下一次强杀无害
|
|
92
|
-
|
|
93
|
-
- 长任务放在有时限的环境(CI、外部看门狗)里跑时,要为 Run 留出到达 `complete` 发布点的时间。后续运行只能沿用完整 Run 已经发布的结果。
|
|
94
|
-
- 实验 `teardown` 写成幂等的:重复执行不报错、目标已经关掉也算成功。这是补收尾机制正确工作的前提。
|
|
95
|
-
- 使用 `sharedState.key` 时,让 checkpoint 回存具备原子性或可补偿性。显式恢复协调租约不能修复半次业务写入。
|
|
75
|
+
恢复只为该 Experiment 补跑 `teardown`。token 错误、旧 token 或 cleanup 失败都不会释放后来 owner 的租约;修复 teardown 后,用当前显示的 token 再试。
|
|
@@ -20,13 +20,13 @@ pnpm exec niceeval exp checkout --dry
|
|
|
20
20
|
|
|
21
21
|
## 核对历史 Attempt
|
|
22
22
|
|
|
23
|
-
使用完整 locator
|
|
23
|
+
使用完整 locator 在终端查看不可变结果:
|
|
24
24
|
|
|
25
25
|
```sh
|
|
26
|
-
pnpm exec niceeval
|
|
26
|
+
pnpm exec niceeval show @1K1P0VJAPVJ12
|
|
27
27
|
```
|
|
28
28
|
|
|
29
|
-
核对 Verdict、Assertion
|
|
29
|
+
核对 Verdict、Assertion、用量、时间和证据。需要回看来源 Run 的完整分母时,再使用:
|
|
30
30
|
|
|
31
31
|
```sh
|
|
32
32
|
pnpm exec niceeval view --run <source-run-id>
|
|
@@ -65,7 +65,7 @@ pnpm exec niceeval view --run <accepted-run-id>
|
|
|
65
65
|
| 错误 | 下一步 |
|
|
66
66
|
| --- | --- |
|
|
67
67
|
| `malformed-locator` | 从 `--dry` 或固定 query result 复制完整的大写 locator,不使用旧 UUID 或模糊前缀。 |
|
|
68
|
-
| `locator-not-found` |
|
|
68
|
+
| `locator-not-found` | 确认读取的是同一份 Record;读取复制出的 SQLite 文件时,为 `show` 显式传 `--record`。 |
|
|
69
69
|
| `accept-ineligible` | 按提示检查 Verdict、timeout、身份或 Observability;不能证明仍成立时重新运行。 |
|
|
70
70
|
| `duplicate-accept-member` | 每个目标 slot 只保留一个 locator,去掉重复或冲突输入。 |
|
|
71
71
|
|
|
@@ -23,7 +23,7 @@ description: "给 Coding Agent 的完整接入流程:探索项目、与用户
|
|
|
23
23
|
探完之后,向用户**介绍接入等级**并给出推荐(详见[接入等级](/zh/explanation/tier)):
|
|
24
24
|
|
|
25
25
|
- **Tier 1(只接 send)**:应用一行不改,全套断言(文本、Judge、多轮、工具、HITL)都在这一档。
|
|
26
|
-
- **Tier 2(send + OTel)**:应用把 OTel span 也发 NiceEval
|
|
26
|
+
- **Tier 2(send + OTel)**:应用把 OTel span 也发 NiceEval 一份,在 `niceeval view` 中查看调用瀑布图。已有埋点(第 3 点探到的)就零改动。
|
|
27
27
|
- **Tier 3(侵入改造 + flags)**:把应用内部变体暴露成 `flags` 做 feature A/B。已有 A/B 开关(第 4 点探到的)就是现成入口。
|
|
28
28
|
|
|
29
29
|
**默认推荐先 Tier 1 跑通,再升 Tier 2**——尤其当第 3 点探到应用已有 OTel 时,明确告诉用户「升 Tier 2 只是把 span 多发一份,成本接近零」。Tier 3 只在用户明确要做变体对比时提。
|
|
@@ -69,7 +69,7 @@ export default defineConfig({
|
|
|
69
69
|
按第 1 步选中的方向读完对应文档后,依次写:
|
|
70
70
|
|
|
71
71
|
1. **Adapter**(`agents/*.ts` 或用户项目里约定的目录)——用 `defineAgent` 实现 `send`,配置走工厂参数,不写死、不读 `process.env`。契约见 [Adapter](/zh/explanation/adapter),API 签名见 [Adapter 参考](/zh/reference/define-agent),事件映射见[事件参考](/zh/reference/events)。两个容易踩的点:**端点/模式要选被测系统核心能力的入口,不是最容易跑通的入口**——比如被测平台既有「纯 LLM 聊天」又有「连库执行」两种模式,接前者等于评了个底层模型代理,没评到产品本身。**`evidenceCoverage` 必须按实际映射如实声明**——只把最终文本映射出来就不要用 `completeEvidenceCoverage`,声明会影响断言完整性,虚报比保守更糟。
|
|
72
|
-
2. **Experiment**(`experiments/*.ts`)——引用上面的 Adapter,声明 `model`、`flags`、`attempts` 等。模型对比写两个实验文件,各自钉一个 `model`。`evals: (eval) => boolean` 决定各自本次运行哪些评估用例。路径只负责 id 和批量运行,固定 Inspection operation
|
|
72
|
+
2. **Experiment**(`experiments/*.ts`)——引用上面的 Adapter,声明 `model`、`flags`、`attempts` 等。模型对比写两个实验文件,各自钉一个 `model`。`evals: (eval) => boolean` 决定各自本次运行哪些评估用例。路径只负责 id 和批量运行,固定 Inspection operation 消费当前已发布的物理结果与覆盖事实。
|
|
73
73
|
3. **评估用例**(`evals/*.eval.ts`)——**先探明这个应用是干嘛的,再写一条贴着它真实功能的评估用例**。读它的 README、路由、工具定义或系统提示,找出它的核心用例(客服机器人就问一条真实的客服问题、SQL agent 就给一个真实的查询任务),拿这个用例做第一条评估用例的输入和断言。两类输入都不合格:「你好」这种和应用无关的占位输入,以及「你是什么/你能做什么」这种**问被测系统它自己的元问题**——那不是用户拿它干活的用例。形式上仍从最小写起:一句输入,`t.succeeded()` + 一个针对预期回答的内容断言——但最小形式只是调通的脚手架,不是交付标准,收尾前还要满足两条:
|
|
74
74
|
- **断言在被测系统胡编时要会变红**。不要断言输入里本来就有的词(问「X 是什么」再断言回答含「X」,被测方复读题目就能通过)。断言预期回答独有的实质内容——具体事实、结构(`hasSections()`)、真实链接(`includesUrl()`),或用 `check({ input, output }, closedQA(...))` 做语义判定。
|
|
75
75
|
- **至少一条负例**。喂一个被测系统应该答不了的输入(不存在的表、检索不到的主题),断言它明确说查不到/做不到,而不是编造一个看似合理的结果——对连着真实数据/检索源的 agent,这是最值得先测的失败形态。
|
|
@@ -168,7 +168,7 @@ export default defineEval({
|
|
|
168
168
|
<Lifecycle
|
|
169
169
|
title="一条 Attempt 从排队到收尾"
|
|
170
170
|
hint="悬停暂停"
|
|
171
|
-
note={<>收尾段不计入这条 Attempt 的耗时口径。用 <code>niceeval
|
|
171
|
+
note={<>收尾段不计入这条 Attempt 的耗时口径。用 <code>niceeval show @<attempt-locator></code> 审阅运行信息,或用固定 query operation 读取机器结果。</>}
|
|
172
172
|
phases={[
|
|
173
173
|
{ name: "experiment.setup", what: "起隧道、起共享服务", times: "整场一次" },
|
|
174
174
|
{ name: "sandbox.queue", what: "等一个并发位", times: "每 Attempt" },
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
title: "比较多个固定 Run"
|
|
3
3
|
sidebarTitle: "比较 Run"
|
|
4
|
-
description: "用 `runs.compare` operation 比较明确 Run
|
|
4
|
+
description: "用 `runs.compare` operation 比较明确 Run 的已发布事实,并在浏览器中审阅相同的结果。"
|
|
5
5
|
---
|
|
6
6
|
|
|
7
7
|
比较模型、提示词或 Agent 版本时,先保留每次运行 receipt 的 `runIds`。不要用最新结果代替明确的比较边界。
|
|
@@ -27,6 +27,6 @@ niceeval query run --request compare.json > compare-result.json
|
|
|
27
27
|
niceeval view --run <baseline-run-id> --run <candidate-run-id>
|
|
28
28
|
```
|
|
29
29
|
|
|
30
|
-
View
|
|
30
|
+
View 读取同样已发布的事实,帮助人审阅 Overview 与 Attempt detail。它不引入另一套统计、比较或页面组装逻辑。
|
|
31
31
|
|
|
32
|
-
当结果来自另一台兼容 runtime
|
|
32
|
+
当结果来自另一台兼容 runtime 时,复制 canonical `.niceeval/record.sqlite`。随后为 `show` 或 `query` 指定这份副本的 `--record` 输入;`view` 只读取当前项目的 Record。
|
|
@@ -199,23 +199,23 @@ export default defineExperiment({
|
|
|
199
199
|
});
|
|
200
200
|
```
|
|
201
201
|
|
|
202
|
-
##
|
|
202
|
+
## 多开终端时共用项目 Record
|
|
203
203
|
|
|
204
|
-
|
|
204
|
+
同一项目的 `.niceeval/record.sqlite` 允许多个 Invocation 发布结果,但不会让两条 Invocation 自动分工。确认本机、Provider 和 Agent 服务还有余量时,两个终端直接在同一项目运行:
|
|
205
205
|
|
|
206
206
|
```bash
|
|
207
|
-
niceeval exp compare --
|
|
208
|
-
niceeval exp compare --
|
|
207
|
+
niceeval exp compare --max-concurrency 2
|
|
208
|
+
niceeval exp compare --max-concurrency 2
|
|
209
209
|
```
|
|
210
210
|
|
|
211
|
-
两边各自规划完整选择,可能重复执行相同评估用例,因为双方都不会读取对方尚未发布的 Attempt
|
|
211
|
+
两边各自规划完整选择,可能重复执行相同评估用例,因为双方都不会读取对方尚未发布的 Attempt。每个 Attempt 发布后立刻可供读取,即使来源 Run 仍是 `active`。需要隔离结果时,在另一个项目目录运行;NiceEval 不会合并两份 Record。
|
|
212
212
|
|
|
213
213
|
三个边界:
|
|
214
214
|
|
|
215
215
|
- 两边的 CLI 上限会相加,Provider 容量不会。配额紧张时把两边都调低。
|
|
216
216
|
- 实验自己的 `maxConcurrency` 也是每个终端各算各的:两条命令都声明 3 时,合计最多跑六条。
|
|
217
217
|
- 两个终端的 Sandbox 池也各自独立。只有确实共享外部 checkpoint 时才声明 `sharedState.key`;它不会合并 Record。
|
|
218
|
-
- 并发
|
|
218
|
+
- 并发 Invocation 可以共用项目 Record。不要在运行期间复制 `.niceeval/record.sqlite`;等命令成功完成 portable gate 后再复制或归档这一个文件。
|
|
219
219
|
|
|
220
220
|
## 边界
|
|
221
221
|
|
|
@@ -104,7 +104,7 @@ export default defineEval({
|
|
|
104
104
|
```bash
|
|
105
105
|
npx niceeval exp my-agent # 跑这个 experiment 下的全部 eval
|
|
106
106
|
npx niceeval exp my-agent refund # 只跑 ID 以 refund 开头的
|
|
107
|
-
pnpm exec niceeval view --run <run-id> #
|
|
107
|
+
pnpm exec niceeval view --run <run-id> # 在本地查看器里查看一个 Run
|
|
108
108
|
```
|
|
109
109
|
|
|
110
110
|
**验证运行结果**:终端会显示动态 dashboard,完成和排队数量在原位更新。失败、错误和 warning 会保留在输出中。运行结束后会打印摘要和 receipt;用其中的 Run ID 也可以执行 `pnpm exec niceeval view --run <runId>`,查看每条评估用例逐轮的输入、事件和评分明细。
|
|
@@ -208,7 +208,7 @@ async send(input, ctx) {
|
|
|
208
208
|
|
|
209
209
|
`progress` 是可覆盖的短期状态。`diagnostic` 是运行结束后仍能回顾的有界记录。两者都不能指定 phase 或输出流,也不会自动改变 `Turn.status` 或 Attempt 判定。连接失败、解析无法继续等基础设施错误应抛出异常。正常收到的被测 Agent 失败通过 `Turn.status: "failed"` 表达。
|
|
210
210
|
|
|
211
|
-
终端只显示错误的一层摘要和 Attempt identity。完整 code、message、cause、stack 与 diagnostics 保存在 Attempt-owned 通道里,使用 `niceeval
|
|
211
|
+
终端只显示错误的一层摘要和 Attempt identity。完整 code、message、cause、stack 与 diagnostics 保存在 Attempt-owned 通道里,使用 `niceeval show @<attempt-locator>` 或固定 `query` operation 查看。OTel trace 只补充调用关系和耗时,不是错误数据的前提。
|
|
212
212
|
|
|
213
213
|
要分别评估本地和生产环境,创建两个 Experiment 文件,并传入不同的工厂参数:
|
|
214
214
|
|
|
@@ -33,7 +33,7 @@ export default defineExperiment({
|
|
|
33
33
|
npx niceeval exp models
|
|
34
34
|
```
|
|
35
35
|
|
|
36
|
-
另一个模型写另一个文件。固定 Inspection operation 在声明的 request 范围内比较合格的 Run;View
|
|
36
|
+
另一个模型写另一个文件。固定 Inspection operation 在声明的 request 范围内比较合格的 Run;View 打开选中的 Run,不接受 Experiment selector。
|
|
37
37
|
|
|
38
38
|
多层目录只负责 id 和批量选择:
|
|
39
39
|
|
|
@@ -95,7 +95,7 @@ npx niceeval exp my-agent fixtures/button
|
|
|
95
95
|
安装或复检失败时,Attempt 会在 `agent.ensure` 阶段变成 `errored`。终端会给出 Attempt 定位符。
|
|
96
96
|
|
|
97
97
|
```shell
|
|
98
|
-
pnpm exec niceeval
|
|
98
|
+
pnpm exec niceeval show @<attempt-locator>
|
|
99
99
|
```
|
|
100
100
|
|
|
101
101
|
先检查错误里的包名、版本和目标平台。修正 `identity`、npm 包或安装步骤后,重新运行原命令。
|
|
@@ -26,7 +26,7 @@ description: "安装 NiceEval,写三个文件,10 分钟内对你自己的应
|
|
|
26
26
|
<Step title="查看结果">
|
|
27
27
|
```bash
|
|
28
28
|
pnpm exec niceeval query discover # AI、脚本或 CI 先发现读取 operation
|
|
29
|
-
pnpm exec niceeval view #
|
|
29
|
+
pnpm exec niceeval view # 网页审阅已发布的结果
|
|
30
30
|
```
|
|
31
31
|
</Step>
|
|
32
32
|
</Steps>
|
|
@@ -103,7 +103,7 @@ export default defineEval({
|
|
|
103
103
|
```bash
|
|
104
104
|
pnpm exec niceeval exp my-agent # 跑起来
|
|
105
105
|
pnpm exec niceeval query discover # 让机器发现可用 operation
|
|
106
|
-
pnpm exec niceeval view #
|
|
106
|
+
pnpm exec niceeval view # 在网页审阅当前已发布的结果
|
|
107
107
|
```
|
|
108
108
|
|
|
109
109
|
到这里第一条评估用例已经跑通。
|
|
@@ -19,7 +19,7 @@ Reporter 面向当前 Invocation。它可以把当前运行的结果发送到外
|
|
|
19
19
|
|
|
20
20
|
## Reporter 适合什么
|
|
21
21
|
|
|
22
|
-
把 Reporter 配在需要观察的运行范围,让它处理当前进程可见的结果。外部系统可以接收运行摘要、Attempt 结果或链接,但要读取完整业务事实时,应使用 receipt 的 `runIds`
|
|
22
|
+
把 Reporter 配在需要观察的运行范围,让它处理当前进程可见的结果。外部系统可以接收运行摘要、Attempt 结果或链接,但要读取完整业务事实时,应使用 receipt 的 `runIds` 重新打开已发布事实。
|
|
23
23
|
|
|
24
24
|
Reporter 不能:
|
|
25
25
|
|
|
@@ -50,7 +50,7 @@ npx niceeval exp checkout --json
|
|
|
50
50
|
pnpm exec niceeval view --run 01J9ZK3M6P4T7V9X2C5N8QW0RY
|
|
51
51
|
```
|
|
52
52
|
|
|
53
|
-
不要把 progress 或 diagnostic 行当作长期结果格式。它们服务当前 Invocation
|
|
53
|
+
不要把 progress 或 diagnostic 行当作长期结果格式。它们服务当前 Invocation,完整业务事实属于已发布 Record。
|
|
54
54
|
|
|
55
55
|
## 接入外部系统时的原则
|
|
56
56
|
|
|
@@ -58,6 +58,6 @@ pnpm exec niceeval view --run 01J9ZK3M6P4T7V9X2C5N8QW0RY
|
|
|
58
58
|
2. 将凭据保留在运行变量中,不写入输出内容。
|
|
59
59
|
3. 用退出状态处理 CI 门禁。
|
|
60
60
|
4. 用 receipt 的 `runIds` 把外部工作流关联到可查看的结果。
|
|
61
|
-
5.
|
|
61
|
+
5. 需要审阅已发布事实时,使用 receipt 的 `runIds` 打开 `view`,或用 `query` 交给自动化流程。
|
|
62
62
|
|
|
63
63
|
有关 CI 命令与 JUnit,请阅读[CI 集成](/zh/tutorials/ci-integration)。
|
|
@@ -206,10 +206,10 @@ export default defineSandboxAgent({
|
|
|
206
206
|
```text
|
|
207
207
|
✗ @5TB8167MXJ30SYZCNAVRHPQ4D2 fixtures/button [local] errored · agent setup
|
|
208
208
|
agent-install-failed: npm install my-agent exited with code 1
|
|
209
|
-
Inspect: niceeval
|
|
209
|
+
Inspect: niceeval show @5TB8167MXJ30SYZCNAVRHPQ4D2
|
|
210
210
|
```
|
|
211
211
|
|
|
212
|
-
运行 `niceeval
|
|
212
|
+
运行 `niceeval show @5TB8167MXJ30SYZCNAVRHPQ4D2` 可查看结构化错误、diagnostics 和已完成的生命周期阶段。需要机器读取时,从 `query discover` 的固定 operation 选择 trace。它可直接看出错误或超时发生在哪一层。
|
|
213
213
|
|
|
214
214
|
树超过 80 个细节节点时会保留失败、慢点和首尾样本,并提示省略数量。需要逐节点审计时,使用 trace operation 的完整结果。Sandbox 创建失败可能发生在 telemetry 建立前,所以错误回顾不依赖 trace。
|
|
215
215
|
|
|
@@ -63,9 +63,9 @@ export default defineExperiment({
|
|
|
63
63
|
|
|
64
64
|
## 多开终端时 Sandbox 不共享
|
|
65
65
|
|
|
66
|
-
Sandbox 复用只发生在一次 Invocation 里。两个终端同时跑一个实验时,两边有各自的 Run 和 Sandbox 池。NiceEval 不把运行中的 Sandbox handle 交给另一个进程;两个 Invocation
|
|
66
|
+
Sandbox 复用只发生在一次 Invocation 里。两个终端同时跑一个实验时,两边有各自的 Run 和 Sandbox 池。NiceEval 不把运行中的 Sandbox handle 交给另一个进程;两个 Invocation 可以并发向同一项目的 `.niceeval/record.sqlite` 发布结果。
|
|
67
67
|
|
|
68
|
-
|
|
68
|
+
正常并发运行不需要为了避开 writer 冲突而拆成多个 Record。只在 Invocation 成功完成 portable gate 后复制或归档 `.niceeval/record.sqlite`;运行中的 SQLite 文件不是可搬运快照。
|
|
69
69
|
|
|
70
70
|
如果 Sandbox 只保留自己的临时状态,两边可以同时跑不同评估用例。动态 `.before()` 回调恢复、回存同一个外部 checkpoint 时,为 Experiment 声明稳定且不含秘密的 key:
|
|
71
71
|
|
|
@@ -86,7 +86,7 @@ export default defineExperiment({
|
|
|
86
86
|
|
|
87
87
|
NiceEval 会在 Experiment setup 或创建 Sandbox 前取得这个 key,直到 Sandbox cleanup、Provider finalizer 与 Experiment `teardown` 全部完成才释放。等待方不会创建 Sandbox;租约释放后,它继续自己的计划,不读取或采用另一条 Run 的结果。
|
|
88
88
|
|
|
89
|
-
key 会进入结果的配置身份。更换 key 就是更换状态 cohort,旧结果不会混进来。租约只保护外部 checkpoint
|
|
89
|
+
key 会进入结果的配置身份。更换 key 就是更换状态 cohort,旧结果不会混进来。租约只保护外部 checkpoint,与项目 Record 无关;普通 Invocation 仍可并发发布。不同机器或 working copy 仍需外部分布式锁。
|
|
90
90
|
|
|
91
91
|
## 生命周期:谁跑几次
|
|
92
92
|
|
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
---
|
|
2
|
-
title: "
|
|
2
|
+
title: "查看运行结果"
|
|
3
3
|
sidebarTitle: "查看结果"
|
|
4
4
|
description: "在终端用 show 逐层查看结果;在浏览器用 View 审阅,在自动化中用 query 读取 JSON。"
|
|
5
5
|
---
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
每个 Attempt 发布后立即可读,不必等待来源 Run 结束。运行结束后,receipt 中的 `runIds` 也是稳定的读取边界。人在终端用 `niceeval show`,在浏览器用 `niceeval view`。AI、脚本和 CI 用 `niceeval query` 读取 JSON。
|
|
8
8
|
|
|
9
9
|
## 先在终端查看 Overview
|
|
10
10
|
|
|
@@ -12,9 +12,9 @@ description: "在终端用 show 逐层查看结果;在浏览器用 View 审阅
|
|
|
12
12
|
niceeval show
|
|
13
13
|
```
|
|
14
14
|
|
|
15
|
-
不带 selector 时,`show` 输出当前
|
|
15
|
+
不带 selector 时,`show` 输出当前 PublicationCutoff 的 Overview。它按 Experiment、评估用例和 Attempt 展开汇总,并保留通过率、分数、coverage、缺失项和 issues。先从这里找出需要继续查看的 Experiment、Run 或 Attempt locator。
|
|
16
16
|
|
|
17
|
-
Overview
|
|
17
|
+
Overview 来自项目唯一的 `.niceeval/record.sqlite` 中已发布的结果。即使当前源码或当前安装候选的 identity 已经改变,
|
|
18
18
|
`show` 也会保留每个逻辑 slot 的最新结果,不会把存在的历史显示为 `Observed 0/0`。这表示结果可查看,
|
|
19
19
|
不表示它已通过当前 target 的复用资格检查。
|
|
20
20
|
|
|
@@ -25,7 +25,7 @@ niceeval show --experiment <experiment-id>
|
|
|
25
25
|
niceeval show --run <run-id>
|
|
26
26
|
```
|
|
27
27
|
|
|
28
|
-
`--experiment` 显示一个 Experiment 的聚合结果和各评估用例。`--run`
|
|
28
|
+
`--experiment` 显示一个 Experiment 的聚合结果和各评估用例。`--run` 显示一个精确 Run 的成员、判定、分数、coverage 和 usage 摘要;`active` Run 已发布的 Attempt 会出现,尚未发布的位置显示 `pending`。两者都可重复传入多个精确 ID;不能与彼此或 Attempt locator 混用。
|
|
29
29
|
|
|
30
30
|
## 查看一个 Attempt
|
|
31
31
|
|
|
@@ -49,7 +49,7 @@ niceeval show @<locator> --diff
|
|
|
49
49
|
niceeval show @<locator> --execution --expand <stable-id>
|
|
50
50
|
```
|
|
51
51
|
|
|
52
|
-
`--timing` 显示活动的顺序和耗时,`--usage` 显示 token、请求和成本总计,`--diff` 显示已捕获的文件变更窗口。详情只读取 Record
|
|
52
|
+
`--timing` 显示活动的顺序和耗时,`--usage` 显示 token、请求和成本总计,`--diff` 显示已捕获的文件变更窗口。详情只读取 Record 已发布的事实;`partial`、`not-recorded` 或 `unavailable` 表示实际可用范围,不能当作零值或失败。
|
|
53
53
|
|
|
54
54
|
## 正确解释完成、判定与分数
|
|
55
55
|
|
|
@@ -69,11 +69,11 @@ niceeval view
|
|
|
69
69
|
niceeval view --run <run-id>
|
|
70
70
|
```
|
|
71
71
|
|
|
72
|
-
View 是给人使用的固定第一方浏览器界面。它不读取项目自定义页面、主题或组件。不带 selector 时,View 打开项目当前
|
|
72
|
+
View 是给人使用的固定第一方浏览器界面。它不读取项目自定义页面、主题或组件。不带 selector 时,View 打开项目当前 PublicationCutoff 的 Overview;`--run` 选择一个或多个明确 Run。它同样显示 `active` Run 已发布的 Attempt,以及尚未发布 slot 的 `pending`。
|
|
73
73
|
|
|
74
74
|
可以使用 `--no-open` 保持浏览器不自动打开,或使用 `--port` 选择 loopback port。`--json` 只输出 lifecycle NDJSON;它不输出运行结果。
|
|
75
75
|
|
|
76
|
-
|
|
76
|
+
新事实发布后,项目 View 会提示用户确认 refresh;确认后才切换到新的 PublicationCutoff。
|
|
77
77
|
|
|
78
78
|
## 让机器读取 JSON
|
|
79
79
|
|
|
@@ -87,13 +87,12 @@ niceeval query run --request request.json
|
|
|
87
87
|
|
|
88
88
|
常见读取任务由固定 operation 覆盖:列出 Run、取得 Run summary、读取 Attempt detail、trace、diff、sources、artifacts 或比较两个 Run。查看 [CLI](/zh/reference/cli) 获取完整说明。
|
|
89
89
|
|
|
90
|
-
##
|
|
90
|
+
## 读取另一份已发布事实
|
|
91
91
|
|
|
92
92
|
```sh
|
|
93
93
|
cp .niceeval/record.sqlite ./baseline.record.sqlite
|
|
94
|
-
niceeval view --record ./baseline.record.sqlite
|
|
95
94
|
niceeval show --record ./baseline.record.sqlite
|
|
96
95
|
niceeval query explain --record ./baseline.record.sqlite --request request.json
|
|
97
96
|
```
|
|
98
97
|
|
|
99
|
-
`--record` 接受 canonical SQLite Record 的副本,并把它当 hostile input。精确 current schema、SQLite integrity 或领域 invariant
|
|
98
|
+
`show` 和 `query` 的 `--record` 接受 canonical SQLite Record 的副本,并把它当 hostile input。精确 current schema、SQLite integrity 或领域 invariant 任一校验失败,命令都会拒绝整个文件。`view` 只读取当前项目的 `.niceeval/record.sqlite`。
|