@codyswann/lisa 2.309.4 → 2.310.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/all/create-only/.lisa/ACCOUNTABILITY.md +72 -0
- package/dist/core/upstream-evidence-manifest.d.ts.map +1 -1
- package/dist/core/upstream-evidence-manifest.js +36 -2
- package/dist/core/upstream-evidence-manifest.js.map +1 -1
- package/package.json +1 -1
- package/plugins/lisa/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa/.codex-plugin/skills/lisa-agent-design-best-practices/SKILL.md +24 -1
- package/plugins/lisa/.codex-plugin/skills/lisa-capability-drift/SKILL.md +71 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-capability-drift/agents/openai.yaml +4 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-delivery-effectiveness/SKILL.md +62 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-delivery-effectiveness/agents/openai.yaml +4 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-evaluation-suite/SKILL.md +63 -0
- package/plugins/lisa/.codex-plugin/skills/lisa-evaluation-suite/agents/openai.yaml +4 -0
- package/plugins/lisa/agents/eval-specialist.md +33 -0
- package/plugins/lisa/rules/reference/intent-routing.md +15 -1
- package/plugins/lisa/skills/lisa-agent-design-best-practices/SKILL.md +24 -1
- package/plugins/lisa/skills/lisa-capability-drift/SKILL.md +71 -0
- package/plugins/lisa/skills/lisa-capability-drift/agents/openai.yaml +4 -0
- package/plugins/lisa/skills/lisa-delivery-effectiveness/SKILL.md +62 -0
- package/plugins/lisa/skills/lisa-delivery-effectiveness/agents/openai.yaml +4 -0
- package/plugins/lisa/skills/lisa-evaluation-suite/SKILL.md +63 -0
- package/plugins/lisa/skills/lisa-evaluation-suite/agents/openai.yaml +4 -0
- package/plugins/lisa-agy/agents/eval-specialist.md +33 -0
- package/plugins/lisa-agy/plugin.json +1 -1
- package/plugins/lisa-agy/skills/lisa-agent-design-best-practices/SKILL.md +24 -1
- package/plugins/lisa-agy/skills/lisa-capability-drift/SKILL.md +71 -0
- package/plugins/lisa-agy/skills/lisa-delivery-effectiveness/SKILL.md +62 -0
- package/plugins/lisa-agy/skills/lisa-evaluation-suite/SKILL.md +63 -0
- package/plugins/lisa-cdk/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk-agy/plugin.json +1 -1
- package/plugins/lisa-cdk-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cdk-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-copilot/agents/eval-specialist.agent.md +33 -0
- package/plugins/lisa-copilot/rules/reference/intent-routing.md +15 -1
- package/plugins/lisa-copilot/skills/lisa-agent-design-best-practices/SKILL.md +24 -1
- package/plugins/lisa-copilot/skills/lisa-capability-drift/SKILL.md +71 -0
- package/plugins/lisa-copilot/skills/lisa-delivery-effectiveness/SKILL.md +62 -0
- package/plugins/lisa-copilot/skills/lisa-evaluation-suite/SKILL.md +63 -0
- package/plugins/lisa-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-cursor/agents/eval-specialist.md +33 -0
- package/plugins/lisa-cursor/rules/intent-routing-reference.mdc +15 -1
- package/plugins/lisa-cursor/skills/lisa-agent-design-best-practices/SKILL.md +24 -1
- package/plugins/lisa-cursor/skills/lisa-capability-drift/SKILL.md +71 -0
- package/plugins/lisa-cursor/skills/lisa-delivery-effectiveness/SKILL.md +62 -0
- package/plugins/lisa-cursor/skills/lisa-evaluation-suite/SKILL.md +63 -0
- package/plugins/lisa-expo/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-expo/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-expo-agy/plugin.json +1 -1
- package/plugins/lisa-expo-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-expo-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-agy/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-harper-fabric-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs-agy/plugin.json +1 -1
- package/plugins/lisa-nestjs-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-nestjs-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw-agy/plugin.json +1 -1
- package/plugins/lisa-openclaw-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-openclaw-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser-agy/plugin.json +1 -1
- package/plugins/lisa-phaser-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-phaser-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-rails-agy/plugin.json +1 -1
- package/plugins/lisa-rails-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-rails-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript-agy/plugin.json +1 -1
- package/plugins/lisa-typescript-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-typescript-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki/.codex-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki-agy/plugin.json +1 -1
- package/plugins/lisa-wiki-copilot/.claude-plugin/plugin.json +1 -1
- package/plugins/lisa-wiki-cursor/.claude-plugin/plugin.json +1 -1
- package/plugins/src/base/agents/eval-specialist.md +33 -0
- package/plugins/src/base/rules/reference/intent-routing.md +15 -1
- package/plugins/src/base/skills/lisa-agent-design-best-practices/SKILL.md +24 -1
- package/plugins/src/base/skills/lisa-capability-drift/SKILL.md +71 -0
- package/plugins/src/base/skills/lisa-delivery-effectiveness/SKILL.md +62 -0
- package/plugins/src/base/skills/lisa-evaluation-suite/SKILL.md +63 -0
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: eval-specialist
|
|
3
|
+
description: Evaluation specialist agent. Measures the factory itself rather than the software it produces — maintains the entity's own task suite, samples it against a recorded baseline to catch capability decline, and reports delivery effectiveness alongside autonomy. Independent of the agents it judges by construction.
|
|
4
|
+
tools: Read, Grep, Glob, Bash
|
|
5
|
+
skills:
|
|
6
|
+
- evaluation-suite
|
|
7
|
+
- capability-drift
|
|
8
|
+
- delivery-effectiveness
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
# Evaluation Specialist Agent
|
|
12
|
+
|
|
13
|
+
Every other agent here is judged on the software it produces. You are the one that judges the producers.
|
|
14
|
+
|
|
15
|
+
Three skills carry the procedures — `evaluation-suite` for the corpus qualification runs against, `capability-drift` for noticing decline nobody caused, and `delivery-effectiveness` for whether the output was worth shipping. Follow them and emit the contract each defines.
|
|
16
|
+
|
|
17
|
+
## Why this agent is separate
|
|
18
|
+
|
|
19
|
+
Not to add a specialism, but because the measurement cannot sit with the measured. A reviewer scoring its own review, a builder reporting its own rework rate, a fleet grading its own suite — each produces a number with no information in it. Your independence from the work you judge is the whole reason you exist, so guard it: never take on the work you are measuring, even when it would be faster to fix what you found than to report it.
|
|
20
|
+
|
|
21
|
+
## What you route
|
|
22
|
+
|
|
23
|
+
- **Which question is being asked.** "Is this configuration good enough to adopt" goes to `evaluation-suite`. "Did we get worse without changing anything" goes to `capability-drift`. "Was the work any good" goes to `delivery-effectiveness`. They share a vocabulary and answer different things; conflating them produces a number nobody can act on.
|
|
24
|
+
- **Whether the instrument is trustworthy before the reading.** A suite the agents can read, one that no longer resembles the work, or one every candidate passes cannot support a conclusion. Say the instrument is unfit and stop — a reading from a broken instrument is worse than no reading, because someone will act on it.
|
|
25
|
+
- **Whether a difference is a difference.** Nothing here is deterministic. A single run is an anecdote, and two configurations whose spreads overlap have not been distinguished no matter what the means say.
|
|
26
|
+
|
|
27
|
+
## What you must not do
|
|
28
|
+
|
|
29
|
+
Do not report a point estimate where a distribution is available, and never let a favourable single run stand as a result. Do not compare the fleet only against its own past — improving on yourself is compatible with being bad, and a suite that only measures self-relative progress will never say so. Do not fill a gap with an estimate: a metric you could not collect is reported as uncollected, with the reason.
|
|
30
|
+
|
|
31
|
+
## What you hand on
|
|
32
|
+
|
|
33
|
+
A verdict per question with the distribution behind it, the instrument's own condition stated alongside, and — where a number moved in the wrong direction — the attribution work that separates a vendor-side change from one of ours. Findings that warrant action leave as build-ready work through the standard intake, never as advice buried in a report.
|
|
@@ -34,7 +34,7 @@ What this rule still enforces:
|
|
|
34
34
|
|
|
35
35
|
2. **Cascade rule (load-bearing)**: Before creating a team, determine your role. You are *inside* an agent team only if you are yourself a spawned teammate/subagent — you were spawned into a team context, or your context references a team lead you report to. A lead/root session is never "inside" a team in this sense, even when a prior team-creation or `Agent` call exists in the session: the lead simply keeps using its existing team (or forms one with its first named spawn) — including when a lifecycle skill is invoked there by `lisa-intake`. If you ARE a spawned teammate, **do NOT create a second team** — many harnesses reject double-creates and the work stalls — and do NOT collapse the nested flow into inline single-agent work. The nested flow must request the existing team lead add the specialist agent(s) it needs to the current team and coordinate through the shared task state. On Claude, teammates cannot add named teammates (teams are flat), so message the lead with the teammate(s), assignments, and completion criteria. On Codex, ask the addressable lead/root to `multi_agent_v1.spawn_agent` the specialists; if no lead handle exists but spawning is available, spawn the bounded specialist agent(s), `wait_agent`, and relay results upward. Invoke flows via the Skill tool; never satisfy a team-first flow by doing all the work inline.
|
|
36
36
|
|
|
37
|
-
3. **Default mode**: `Research`, `Plan`, `Implement`, `Intake`, and `Debrief` run as agent teams. (`Intake` is a special case: the Intake skill itself is a thin dispatcher that creates no team and never spawns the lifecycle flow as a subagent — the team is created by the per-item lifecycle skill, `lisa-plan` or `lisa-implement`, that Intake dispatches in-session, so the session still runs as an agent team.) The `Implement` flow — including every work type (`Build`, `Fix`, `Improve`, `Investigate-Only`) — is **always** a team flow. Bug fixes that "look simple" are not an exception: the Reproduce sub-flow, debug-specialist, bug-fixer, parallel reviewers, and verification-specialist all need to compose. `Debrief` runs as a team because tracker-mining and pr-mining parallelize cleanly and synthesis gates on both completing. `Verify` (standalone) and `
|
|
37
|
+
3. **Default mode**: `Research`, `Plan`, `Implement`, `Intake`, and `Debrief` run as agent teams. (`Intake` is a special case: the Intake skill itself is a thin dispatcher that creates no team and never spawns the lifecycle flow as a subagent — the team is created by the per-item lifecycle skill, `lisa-plan` or `lisa-implement`, that Intake dispatches in-session, so the session still runs as an agent team.) The `Implement` flow — including every work type (`Build`, `Fix`, `Improve`, `Investigate-Only`) — is **always** a team flow. Bug fixes that "look simple" are not an exception: the Reproduce sub-flow, debug-specialist, bug-fixer, parallel reviewers, and verification-specialist all need to compose. `Debrief` runs as a team because tracker-mining and pr-mining parallelize cleanly and synthesis gates on both completing. `Verify` (standalone), `Monitor` (standalone) and `Measure` (standalone) use the One-shot Sub-agents pattern (see `## Orchestration` below) — these flows are linear with no parallelism and the team overhead is not warranted. Single-agent mode is otherwise reserved for: `product-walkthrough` invoked standalone (not as part of Research/Plan), `debrief-apply` (deterministic routing of human-marked dispositions), and one-off diagnostic Bash/Read sessions that don't invoke any lifecycle skill. When in doubt, use a team.
|
|
38
38
|
|
|
39
39
|
The mechanical team bootstrap directive lives inside each lifecycle skill — see those skills' orchestration preambles for the exact wording. Modern Claude Code uses the implicit team model (the first `Agent` spawn establishes the team); older Claude Code may still expose `TeamCreate`. Other runtimes must use their equivalent team-discovery and team-creation tools, or explicitly declare the no-team fallback when no such tool exists.
|
|
40
40
|
|
|
@@ -327,6 +327,20 @@ Sequence:
|
|
|
327
327
|
|
|
328
328
|
The `observability-audit` rule owns the profile detection, rubric, anomaly thresholds, ticket templates (gate-passing), fingerprint/idempotency contract, the cap, and the Verify report-only guard. Monitor **files only** — the `intake` / `tracker-build-intake` cron implements what it files.
|
|
329
329
|
|
|
330
|
+
### Measure
|
|
331
|
+
|
|
332
|
+
Purpose: measure **the factory**, where Monitor measures the software it produces. Answers three separate questions — is a configuration good enough to adopt, did capability decline without anyone changing anything, and was the delivered work worth shipping. Repo-scoped, standalone, and report-plus-file like Monitor.
|
|
333
|
+
|
|
334
|
+
Sequence:
|
|
335
|
+
|
|
336
|
+
1. `eval-specialist` -- route to the question being asked, and refuse a reading from an unfit instrument (a suite the agents can read, one that no longer resembles the work, or one every candidate passes).
|
|
337
|
+
2. **Qualify or sample** -- `evaluation-suite` when a change is being considered; `capability-drift` when nothing changed and the question is whether the fleet slipped. Both report distributions, never point estimates.
|
|
338
|
+
3. **Report effectiveness** -- `delivery-effectiveness` for the window, paired with autonomy rate from AC7.4, because autonomy alone rewards producing rejected work faster.
|
|
339
|
+
4. **Attribute before filing** -- a confirmed decline is separated into vendor-side change, accumulated local change, or broken harness. "The model got worse" is the most expensive available conclusion and is usually wrong.
|
|
340
|
+
5. **File** -- findings enter intake as build-ready leaves through `tracker-write`, same idempotency and cap discipline as Monitor. Measure files only; it never fixes what it finds, and it never takes on the work it is measuring.
|
|
341
|
+
|
|
342
|
+
The independence is structural, not procedural: the agent that measures a specialist's output is never the specialist, because a number produced by the thing being judged carries no information.
|
|
343
|
+
|
|
330
344
|
## Tracker Entry Point (JIRA, GitHub Issues, or Linear)
|
|
331
345
|
|
|
332
346
|
When the request references a tracker ticket (a JIRA key like `PROJ-123`, a JIRA URL, a GitHub issue URL, an `org/repo#<n>` token, or a Linear identifier like `ENG-123` or a Linear project URL):
|
|
@@ -94,7 +94,30 @@ Grant only the tools necessary for the agent's domain. This enforces focus and p
|
|
|
94
94
|
| Implementer | `Read, Write, Edit, Bash, Grep, Glob` | Needs to modify code |
|
|
95
95
|
| Planner | `Read, Grep, Glob` | Research only, no execution |
|
|
96
96
|
|
|
97
|
-
|
|
97
|
+
Do not assign implementation tasks to agents without `Write` and `Edit`.
|
|
98
|
+
|
|
99
|
+
#### `tools:` is a focus mechanism, not a security boundary
|
|
100
|
+
|
|
101
|
+
**Only Claude enforces `tools:`.** Every other harness treats it as advisory: the
|
|
102
|
+
Codex transformer preserves the declaration and emits a compatibility note
|
|
103
|
+
saying tool access "is governed by the active Codex runtime, sandbox, and project
|
|
104
|
+
policy", because there is no portable primitive to enforce it. So "read-only
|
|
105
|
+
agents cannot implement code" is true on Claude and false elsewhere — the same
|
|
106
|
+
agent definition, run on another harness, can write files.
|
|
107
|
+
|
|
108
|
+
Treat the field accordingly:
|
|
109
|
+
|
|
110
|
+
- **Do** use it to keep an agent in its lane, to make its intent legible to a
|
|
111
|
+
reader, and to reduce accidental scope creep on the harness that honours it.
|
|
112
|
+
- **Do not** rely on it to contain a compromised or prompt-injected agent, to
|
|
113
|
+
keep credentials out of reach, or to satisfy any control that has to hold
|
|
114
|
+
under attack. A boundary one runtime ignores is a suggestion everywhere.
|
|
115
|
+
|
|
116
|
+
The controls that actually hold are outside the agent definition: the execution
|
|
117
|
+
sandbox, network egress restriction, the credential scope the agent authenticates
|
|
118
|
+
with, and — because an agent can ask a peer to act for it — which other agents it
|
|
119
|
+
can reach. State those in the system description rather than inferring safety from
|
|
120
|
+
a `tools:` line.
|
|
98
121
|
|
|
99
122
|
### 6. No Hardcoded Interaction Patterns
|
|
100
123
|
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: lisa-capability-drift
|
|
3
|
+
description: "Detect the fleet getting worse when nobody changed anything — sample the evaluation suite against a recorded baseline on a cadence, decide whether a decline is real given run-to-run variance, and attribute it to a vendor-side change, accumulated local change, or a broken harness."
|
|
4
|
+
allowed-tools: ["Read", "Grep", "Glob", "Bash"]
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Capability Drift
|
|
8
|
+
|
|
9
|
+
Qualification fires when you change something. This exists because capability also moves when you change nothing.
|
|
10
|
+
|
|
11
|
+
## What moves underneath you
|
|
12
|
+
|
|
13
|
+
- A vendor reroutes what a pinned name resolves to, or alters the served model behind it.
|
|
14
|
+
- Instruction surfaces accumulate — a rule here, a skill there — until they contradict each other and the model spends its budget reconciling them instead of working.
|
|
15
|
+
- Tool behaviour shifts under the same interface.
|
|
16
|
+
- The work itself changes shape, so yesterday's competence is aimed at a problem that has moved.
|
|
17
|
+
|
|
18
|
+
None of these produce an event you can subscribe to. They produce a slow decline that looks like ordinary bad luck, one task at a time, until somebody finally compares against a year ago and cannot explain the gap.
|
|
19
|
+
|
|
20
|
+
## Record a baseline worth comparing to
|
|
21
|
+
|
|
22
|
+
A baseline is the recorded distribution of suite outcomes for a stated configuration at a stated time — not a single score. Store the configuration completely enough to reproduce it: model, reasoning or effort level, tool set, context assembly, and the suite's own review date.
|
|
23
|
+
|
|
24
|
+
Re-baseline deliberately, after a qualified change, and never silently. A baseline quietly overwritten each run cannot detect anything, because the thing you are comparing against has been moving with you the whole time.
|
|
25
|
+
|
|
26
|
+
## Sample on a cadence
|
|
27
|
+
|
|
28
|
+
Declare an interval and hold it. Sampling only when something feels wrong means you find out about decline through the incident it caused, which is exactly the situation this is meant to prevent.
|
|
29
|
+
|
|
30
|
+
Each sample runs the suite against the deployed configuration and compares to the baseline distribution — not to the previous sample, which turns a trend into a series of individually unremarkable steps.
|
|
31
|
+
|
|
32
|
+
## Deciding whether a decline is real
|
|
33
|
+
|
|
34
|
+
Run-to-run variance is large enough that a lower number usually means nothing. Treat a decline as a finding only when the sampled distribution is distinguishable from the baseline by the spread statistic the suite declares — the same bar a qualification has to clear.
|
|
35
|
+
|
|
36
|
+
A borderline reading is not a finding and not an all-clear. It is a reason to take more samples, and saying so is the honest report.
|
|
37
|
+
|
|
38
|
+
## Attribute it before acting
|
|
39
|
+
|
|
40
|
+
A confirmed decline has three candidate causes, and the remedy differs completely:
|
|
41
|
+
|
|
42
|
+
| Candidate | How to separate it |
|
|
43
|
+
| --- | --- |
|
|
44
|
+
| **Vendor-side change** | Re-run the baseline configuration explicitly pinned. If it now scores like the sample rather than like the baseline, what the pin resolves to has changed underneath you |
|
|
45
|
+
| **Accumulated local change** | Diff the instruction surfaces, tool set and context assembly against the baseline's recorded configuration. Look for additions nobody qualified, and for rules that now contradict |
|
|
46
|
+
| **Broken harness** | Check whether failures are the work being wrong or the run being unable to proceed — missing install, absent credential, changed path. A harness failure reads as incapability and is not one |
|
|
47
|
+
|
|
48
|
+
Do not skip this. "The model got worse" is the most expensive conclusion available, and it is wrong most of the time.
|
|
49
|
+
|
|
50
|
+
## Report and route
|
|
51
|
+
|
|
52
|
+
A confirmed decline enters intake as build-ready work like any other finding, carrying the attribution. An unattributed decline is still worth filing — with the attribution work named as the first task, so nobody re-runs the analysis from scratch.
|
|
53
|
+
|
|
54
|
+
## Output
|
|
55
|
+
|
|
56
|
+
```text
|
|
57
|
+
## Capability Drift
|
|
58
|
+
|
|
59
|
+
**Baseline:** date · configuration (model · effort · tools · context) · suite review date
|
|
60
|
+
**Sample:** date · same fields, with every difference from the baseline marked
|
|
61
|
+
**Cadence:** declared interval, and whether this sample was on schedule
|
|
62
|
+
|
|
63
|
+
### Comparison
|
|
64
|
+
| Task class | Baseline distribution | Sampled distribution | Distinguishable? |
|
|
65
|
+
|---|---|---|---|
|
|
66
|
+
|
|
67
|
+
**Verdict:** stable | declined | inconclusive — take n more samples
|
|
68
|
+
**Attribution (declines only):** vendor-side | accumulated local | broken harness — with the
|
|
69
|
+
evidence that separated it from the other two
|
|
70
|
+
**Filed:** work item, or the reason nothing was filed
|
|
71
|
+
```
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: lisa-delivery-effectiveness
|
|
3
|
+
description: "Measure whether the work the factory delivers was worth shipping — gate rejection, rework, escape, first-pass yield and cost per delivered item — so autonomy rate cannot stand alone as a success measure."
|
|
4
|
+
allowed-tools: ["Read", "Grep", "Glob", "Bash"]
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Delivery Effectiveness
|
|
8
|
+
|
|
9
|
+
Autonomy rate answers how much ran without a human. It says nothing about whether any of it should have shipped, and on its own it rewards exactly the wrong thing: a factory that produces rejected work faster scores better.
|
|
10
|
+
|
|
11
|
+
## The five measures
|
|
12
|
+
|
|
13
|
+
Each needs a stated numerator, denominator and window, or it is a number-shaped opinion.
|
|
14
|
+
|
|
15
|
+
| Measure | Counts | Reads as |
|
|
16
|
+
| --- | --- | --- |
|
|
17
|
+
| **Gate rejection rate** | Work the pipeline refused, over work submitted | How much effort is spent producing output the standards reject |
|
|
18
|
+
| **Rework rate** | Items reopened or returned after being declared complete, over items completed | How often "done" was not done |
|
|
19
|
+
| **Escape rate** | Defects reaching production, over items released | What the gates did not catch |
|
|
20
|
+
| **First-pass yield** | Items reaching terminal state with no rework, over items started | The one number that moves only when the whole line works |
|
|
21
|
+
| **Cost per delivered item** | Metered spend, over items reaching terminal state | What a unit of accepted output actually costs |
|
|
22
|
+
|
|
23
|
+
First-pass yield is the summary measure worth watching, because every other failure shows up in it. The other four exist to tell you *where* it went.
|
|
24
|
+
|
|
25
|
+
## Where the numbers come from
|
|
26
|
+
|
|
27
|
+
Prefer the systems of record over anything self-reported: the tracker for state transitions and reopenings, CI history for gate rejections, the release record for escapes, the billing or gateway boundary for spend. An agent's account of its own rework rate is the least reliable source available and the easiest one to reach for.
|
|
28
|
+
|
|
29
|
+
Say which source each measure came from. Two measures drawn from different systems with different definitions of "complete" cannot be compared, and that mismatch is invisible in the result.
|
|
30
|
+
|
|
31
|
+
## Set targets that can only tighten
|
|
32
|
+
|
|
33
|
+
Targets are yours to declare — this says nothing about what a good rejection rate is, because that depends on the standards being enforced and the work being attempted. But declare them, disclose them, and let them move in one direction only. A target quietly loosened to match the current number is a redefinition dressed as an improvement.
|
|
34
|
+
|
|
35
|
+
## Read it against autonomy, never instead of it
|
|
36
|
+
|
|
37
|
+
Report both, together, always. The pairing is the point:
|
|
38
|
+
|
|
39
|
+
- **High autonomy, high first-pass yield** — the thing everyone is aiming for.
|
|
40
|
+
- **High autonomy, low first-pass yield** — an unattended process producing work its own gates reject. Reporting autonomy alone would conceal exactly this, which is why the pairing is not optional.
|
|
41
|
+
- **Low autonomy, high first-pass yield** — humans are carrying the quality. Honest, and expensive.
|
|
42
|
+
- **Low autonomy, low first-pass yield** — the machinery is not the problem yet.
|
|
43
|
+
|
|
44
|
+
## Output
|
|
45
|
+
|
|
46
|
+
```text
|
|
47
|
+
## Delivery Effectiveness
|
|
48
|
+
|
|
49
|
+
**Window:** dates · **Autonomy rate over the same window:** %
|
|
50
|
+
|
|
51
|
+
| Measure | Value | Numerator / denominator | Source of record | Target | Trend |
|
|
52
|
+
|---|---|---|---|---|---|
|
|
53
|
+
| Gate rejection rate | | | | | |
|
|
54
|
+
| Rework rate | | | | | |
|
|
55
|
+
| Escape rate | | | | | |
|
|
56
|
+
| First-pass yield | | | | | |
|
|
57
|
+
| Cost per delivered item | | | | | |
|
|
58
|
+
|
|
59
|
+
**Reading:** which of the four autonomy/yield quadrants this window sits in
|
|
60
|
+
**Uncollected:** any measure not gathered, and why — never an estimate in its place
|
|
61
|
+
**Filed:** work items raised for measures outside target
|
|
62
|
+
```
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: lisa-evaluation-suite
|
|
3
|
+
description: "Build and maintain the entity's own task suite for qualifying agent, model, prompt and effort-level changes — drawn from real work, kept representative, protected from contamination, and checked for the discriminating power that makes a result mean something."
|
|
4
|
+
allowed-tools: ["Read", "Grep", "Glob", "Bash"]
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
# Evaluation Suite
|
|
8
|
+
|
|
9
|
+
Public leaderboards measure somebody else's work mix. The suite that decides what runs in your factory is assembled from your own.
|
|
10
|
+
|
|
11
|
+
## Draw the tasks from real work
|
|
12
|
+
|
|
13
|
+
A task earns a place by having happened. Pull candidates from closed work items, past incidents, escaped defects, and the reviews that caught something — anywhere the outcome is already known and was consequential.
|
|
14
|
+
|
|
15
|
+
Include, deliberately, the work your current setup handles **badly**. A suite made only of tasks that already pass measures nothing: every candidate scores full marks and the result cannot separate them. Difficulty is what carries information, so keep the tasks that embarrass the fleet and record what "good" would have looked like.
|
|
16
|
+
|
|
17
|
+
Each task records: what is asked, what counts as success and how that is judged, where the task came from, and the outcome class it exercises. Without the judging rule written down, two runs of the same task are not comparable.
|
|
18
|
+
|
|
19
|
+
## Keep it resembling the work
|
|
20
|
+
|
|
21
|
+
Representativeness expires. Review it on a declared cadence and re-establish it whenever the work mix moves — a new stack, a new surface, a shift from features to migrations. A suite that has drifted from the work stops being evidence about the work while continuing to produce confident numbers, which is the worst failure mode available to an instrument.
|
|
22
|
+
|
|
23
|
+
State the review date alongside every result. A reader has to be able to see how old the instrument is.
|
|
24
|
+
|
|
25
|
+
## Contamination is the failure that hides
|
|
26
|
+
|
|
27
|
+
The tasks and their expected outcomes MUST be unreachable by the agents under evaluation. That means not in instruction files, not in skills, not in retrievable context, not in a wiki the agents read, and not in feedback or transcripts submitted to a model vendor.
|
|
28
|
+
|
|
29
|
+
An agent optimising against a suite it can see produces a score rather than a measurement, and nothing in the number reveals which one you have. Treat any suspected exposure as fatal to the affected tasks: retire them, record why, and replace them from real work. Where exposure cannot be ruled out for the whole suite, say so — a compromised instrument reported honestly still tells you something; one reported as clean does not.
|
|
30
|
+
|
|
31
|
+
## Check that it discriminates
|
|
32
|
+
|
|
33
|
+
Before trusting a comparison, look at the spread of results across candidates:
|
|
34
|
+
|
|
35
|
+
| What you see | What it means |
|
|
36
|
+
| --- | --- |
|
|
37
|
+
| Nearly every task passes for every candidate | Saturated. It cannot rank anything; add harder tasks |
|
|
38
|
+
| Nearly every task fails for every candidate | Out of range. The suite is measuring something other than the difference you care about |
|
|
39
|
+
| A handful of tasks decide the whole ordering | Fragile. Swapping any one of them would flip the result — say so with the result |
|
|
40
|
+
| Spreads overlap between candidates | Not distinguished, whatever the means say |
|
|
41
|
+
|
|
42
|
+
## Qualify a change against it
|
|
43
|
+
|
|
44
|
+
A qualification states, before the runs: how many runs per condition, the statistic expressing spread, and the threshold that constitutes a pass. Report the distribution rather than a mean, and treat the pinned operating configuration — model, reasoning or effort level, tools, context — as part of the condition, since results are not monotonic in effort and a change to any of them is a different candidate.
|
|
45
|
+
|
|
46
|
+
Third-party benchmark rankings and vendor claims are not admissible here. They measure a different task mix at a precision their own run-to-run variance does not support.
|
|
47
|
+
|
|
48
|
+
## Output
|
|
49
|
+
|
|
50
|
+
```text
|
|
51
|
+
## Evaluation Suite
|
|
52
|
+
|
|
53
|
+
**Tasks:** n | **Representativeness reviewed:** date | **Contamination:** controlled / suspected / unknown
|
|
54
|
+
**Discrimination:** healthy | saturated | out-of-range | fragile (deciding tasks: …)
|
|
55
|
+
|
|
56
|
+
### Result — per condition
|
|
57
|
+
| Condition (model · effort · tools) | Runs | Outcome distribution | Spread |
|
|
58
|
+
|---|---|---|---|
|
|
59
|
+
|
|
60
|
+
**Protocol declared in advance:** runs per condition, spread statistic, pass threshold
|
|
61
|
+
**Verdict:** qualified / not qualified / instrument unfit — with the reason
|
|
62
|
+
**Uncollected:** any metric that could not be gathered, and why
|
|
63
|
+
```
|