@prismatic-io/lux 0.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/lib/answerers/claude-code/index.d.ts +17 -0
- package/lib/answerers/claude-code/index.d.ts.map +1 -0
- package/lib/answerers/claude-code/index.js +280 -0
- package/lib/answerers/claude-code/index.js.map +1 -0
- package/lib/answerers/persona/index.d.ts +40 -0
- package/lib/answerers/persona/index.d.ts.map +1 -0
- package/lib/answerers/persona/index.js +265 -0
- package/lib/answerers/persona/index.js.map +1 -0
- package/lib/answerers/scripted/index.d.ts +8 -0
- package/lib/answerers/scripted/index.d.ts.map +1 -0
- package/lib/answerers/scripted/index.js +163 -0
- package/lib/answerers/scripted/index.js.map +1 -0
- package/lib/answerers/shared.d.ts +13 -0
- package/lib/answerers/shared.d.ts.map +1 -0
- package/lib/answerers/shared.js +17 -0
- package/lib/answerers/shared.js.map +1 -0
- package/lib/answerers/terminal/index.d.ts +9 -0
- package/lib/answerers/terminal/index.d.ts.map +1 -0
- package/lib/answerers/terminal/index.js +97 -0
- package/lib/answerers/terminal/index.js.map +1 -0
- package/lib/assertions/core/command-exits-zero.d.ts +21 -0
- package/lib/assertions/core/command-exits-zero.d.ts.map +1 -0
- package/lib/assertions/core/command-exits-zero.js +87 -0
- package/lib/assertions/core/command-exits-zero.js.map +1 -0
- package/lib/assertions/core/event-checks.d.ts +57 -0
- package/lib/assertions/core/event-checks.d.ts.map +1 -0
- package/lib/assertions/core/event-checks.js +123 -0
- package/lib/assertions/core/event-checks.js.map +1 -0
- package/lib/assertions/core/file-checks.d.ts +103 -0
- package/lib/assertions/core/file-checks.d.ts.map +1 -0
- package/lib/assertions/core/file-checks.js +288 -0
- package/lib/assertions/core/file-checks.js.map +1 -0
- package/lib/assertions/core/helpers.d.ts +66 -0
- package/lib/assertions/core/helpers.d.ts.map +1 -0
- package/lib/assertions/core/helpers.js +170 -0
- package/lib/assertions/core/helpers.js.map +1 -0
- package/lib/assertions/core/index.d.ts +15 -0
- package/lib/assertions/core/index.d.ts.map +1 -0
- package/lib/assertions/core/index.js +56 -0
- package/lib/assertions/core/index.js.map +1 -0
- package/lib/assertions/core/json-pointer-equals.d.ts +19 -0
- package/lib/assertions/core/json-pointer-equals.d.ts.map +1 -0
- package/lib/assertions/core/json-pointer-equals.js +71 -0
- package/lib/assertions/core/json-pointer-equals.js.map +1 -0
- package/lib/assertions/core/output-checks.d.ts +46 -0
- package/lib/assertions/core/output-checks.d.ts.map +1 -0
- package/lib/assertions/core/output-checks.js +121 -0
- package/lib/assertions/core/output-checks.js.map +1 -0
- package/lib/assertions/core/predicate.d.ts +14 -0
- package/lib/assertions/core/predicate.d.ts.map +1 -0
- package/lib/assertions/core/predicate.js +55 -0
- package/lib/assertions/core/predicate.js.map +1 -0
- package/lib/assertions/core/resource-checks.d.ts +71 -0
- package/lib/assertions/core/resource-checks.d.ts.map +1 -0
- package/lib/assertions/core/resource-checks.js +159 -0
- package/lib/assertions/core/resource-checks.js.map +1 -0
- package/lib/assertions/core/safe-regex.d.ts +23 -0
- package/lib/assertions/core/safe-regex.d.ts.map +1 -0
- package/lib/assertions/core/safe-regex.js +102 -0
- package/lib/assertions/core/safe-regex.js.map +1 -0
- package/lib/assertions/core/tool-called-with.d.ts +21 -0
- package/lib/assertions/core/tool-called-with.d.ts.map +1 -0
- package/lib/assertions/core/tool-called-with.js +69 -0
- package/lib/assertions/core/tool-called-with.js.map +1 -0
- package/lib/assertions/core/tool-calls.d.ts +22 -0
- package/lib/assertions/core/tool-calls.d.ts.map +1 -0
- package/lib/assertions/core/tool-calls.js +17 -0
- package/lib/assertions/core/tool-calls.js.map +1 -0
- package/lib/assertions/core/tool-invocation.d.ts +36 -0
- package/lib/assertions/core/tool-invocation.d.ts.map +1 -0
- package/lib/assertions/core/tool-invocation.js +170 -0
- package/lib/assertions/core/tool-invocation.js.map +1 -0
- package/lib/assertions/core/tool-result-checks.d.ts +59 -0
- package/lib/assertions/core/tool-result-checks.d.ts.map +1 -0
- package/lib/assertions/core/tool-result-checks.js +193 -0
- package/lib/assertions/core/tool-result-checks.js.map +1 -0
- package/lib/assertions/mcp/discovery-checks.d.ts +36 -0
- package/lib/assertions/mcp/discovery-checks.d.ts.map +1 -0
- package/lib/assertions/mcp/discovery-checks.js +122 -0
- package/lib/assertions/mcp/discovery-checks.js.map +1 -0
- package/lib/assertions/mcp/index.d.ts +10 -0
- package/lib/assertions/mcp/index.d.ts.map +1 -0
- package/lib/assertions/mcp/index.js +25 -0
- package/lib/assertions/mcp/index.js.map +1 -0
- package/lib/assertions/rubric/index.d.ts +104 -0
- package/lib/assertions/rubric/index.d.ts.map +1 -0
- package/lib/assertions/rubric/index.js +191 -0
- package/lib/assertions/rubric/index.js.map +1 -0
- package/lib/assertions/rubric/internal.d.ts +14 -0
- package/lib/assertions/rubric/internal.d.ts.map +1 -0
- package/lib/assertions/rubric/internal.js +335 -0
- package/lib/assertions/rubric/internal.js.map +1 -0
- package/lib/authoring.d.ts +3 -0
- package/lib/authoring.d.ts.map +1 -0
- package/lib/authoring.js +3 -0
- package/lib/authoring.js.map +1 -0
- package/lib/cli/bin.d.ts +3 -0
- package/lib/cli/bin.d.ts.map +1 -0
- package/lib/cli/bin.js +15 -0
- package/lib/cli/bin.js.map +1 -0
- package/lib/cli/command-runtime.d.ts +5 -0
- package/lib/cli/command-runtime.d.ts.map +1 -0
- package/lib/cli/command-runtime.js +30 -0
- package/lib/cli/command-runtime.js.map +1 -0
- package/lib/cli/init.d.ts +21 -0
- package/lib/cli/init.d.ts.map +1 -0
- package/lib/cli/init.js +340 -0
- package/lib/cli/init.js.map +1 -0
- package/lib/cli/program.d.ts +22 -0
- package/lib/cli/program.d.ts.map +1 -0
- package/lib/cli/program.js +757 -0
- package/lib/cli/program.js.map +1 -0
- package/lib/cli/render/format.d.ts +14 -0
- package/lib/cli/render/format.d.ts.map +1 -0
- package/lib/cli/render/format.js +71 -0
- package/lib/cli/render/format.js.map +1 -0
- package/lib/cli/render/reporter.d.ts +73 -0
- package/lib/cli/render/reporter.d.ts.map +1 -0
- package/lib/cli/render/reporter.js +519 -0
- package/lib/cli/render/reporter.js.map +1 -0
- package/lib/cli/render/tty.d.ts +77 -0
- package/lib/cli/render/tty.d.ts.map +1 -0
- package/lib/cli/render/tty.js +283 -0
- package/lib/cli/render/tty.js.map +1 -0
- package/lib/cli/run-options.d.ts +27 -0
- package/lib/cli/run-options.d.ts.map +1 -0
- package/lib/cli/run-options.js +58 -0
- package/lib/cli/run-options.js.map +1 -0
- package/lib/cli/skills.d.ts +38 -0
- package/lib/cli/skills.d.ts.map +1 -0
- package/lib/cli/skills.js +84 -0
- package/lib/cli/skills.js.map +1 -0
- package/lib/core/annotation.d.ts +56 -0
- package/lib/core/annotation.d.ts.map +1 -0
- package/lib/core/annotation.js +92 -0
- package/lib/core/annotation.js.map +1 -0
- package/lib/core/answerer.d.ts +68 -0
- package/lib/core/answerer.d.ts.map +1 -0
- package/lib/core/answerer.js +17 -0
- package/lib/core/answerer.js.map +1 -0
- package/lib/core/artifact-evidence.d.ts +50 -0
- package/lib/core/artifact-evidence.d.ts.map +1 -0
- package/lib/core/artifact-evidence.js +185 -0
- package/lib/core/artifact-evidence.js.map +1 -0
- package/lib/core/artifact-snapshot.d.ts +37 -0
- package/lib/core/artifact-snapshot.d.ts.map +1 -0
- package/lib/core/artifact-snapshot.js +339 -0
- package/lib/core/artifact-snapshot.js.map +1 -0
- package/lib/core/assertion.d.ts +129 -0
- package/lib/core/assertion.d.ts.map +1 -0
- package/lib/core/assertion.js +75 -0
- package/lib/core/assertion.js.map +1 -0
- package/lib/core/case.d.ts +79 -0
- package/lib/core/case.d.ts.map +1 -0
- package/lib/core/case.js +90 -0
- package/lib/core/case.js.map +1 -0
- package/lib/core/conversation-turn.d.ts +23 -0
- package/lib/core/conversation-turn.d.ts.map +1 -0
- package/lib/core/conversation-turn.js +31 -0
- package/lib/core/conversation-turn.js.map +1 -0
- package/lib/core/driver.d.ts +64 -0
- package/lib/core/driver.d.ts.map +1 -0
- package/lib/core/driver.js +18 -0
- package/lib/core/driver.js.map +1 -0
- package/lib/core/emitter.d.ts +20 -0
- package/lib/core/emitter.d.ts.map +1 -0
- package/lib/core/emitter.js +34 -0
- package/lib/core/emitter.js.map +1 -0
- package/lib/core/events.d.ts +202 -0
- package/lib/core/events.d.ts.map +1 -0
- package/lib/core/events.js +107 -0
- package/lib/core/events.js.map +1 -0
- package/lib/core/exec.d.ts +82 -0
- package/lib/core/exec.d.ts.map +1 -0
- package/lib/core/exec.js +168 -0
- package/lib/core/exec.js.map +1 -0
- package/lib/core/experiment.d.ts +285 -0
- package/lib/core/experiment.d.ts.map +1 -0
- package/lib/core/experiment.js +389 -0
- package/lib/core/experiment.js.map +1 -0
- package/lib/core/glob.d.ts +3 -0
- package/lib/core/glob.d.ts.map +1 -0
- package/lib/core/glob.js +32 -0
- package/lib/core/glob.js.map +1 -0
- package/lib/core/hash.d.ts +8 -0
- package/lib/core/hash.d.ts.map +1 -0
- package/lib/core/hash.js +214 -0
- package/lib/core/hash.js.map +1 -0
- package/lib/core/index.d.ts +29 -0
- package/lib/core/index.d.ts.map +1 -0
- package/lib/core/index.js +29 -0
- package/lib/core/index.js.map +1 -0
- package/lib/core/lifecycle-fixtures.d.ts +20 -0
- package/lib/core/lifecycle-fixtures.d.ts.map +1 -0
- package/lib/core/lifecycle-fixtures.js +109 -0
- package/lib/core/lifecycle-fixtures.js.map +1 -0
- package/lib/core/lifecycle-hooks.d.ts +52 -0
- package/lib/core/lifecycle-hooks.d.ts.map +1 -0
- package/lib/core/lifecycle-hooks.js +42 -0
- package/lib/core/lifecycle-hooks.js.map +1 -0
- package/lib/core/mcp-events.d.ts +44 -0
- package/lib/core/mcp-events.d.ts.map +1 -0
- package/lib/core/mcp-events.js +23 -0
- package/lib/core/mcp-events.js.map +1 -0
- package/lib/core/model-cli.d.ts +57 -0
- package/lib/core/model-cli.d.ts.map +1 -0
- package/lib/core/model-cli.js +237 -0
- package/lib/core/model-cli.js.map +1 -0
- package/lib/core/observer.d.ts +31 -0
- package/lib/core/observer.d.ts.map +1 -0
- package/lib/core/observer.js +25 -0
- package/lib/core/observer.js.map +1 -0
- package/lib/core/platform-process.d.ts +33 -0
- package/lib/core/platform-process.d.ts.map +1 -0
- package/lib/core/platform-process.js +214 -0
- package/lib/core/platform-process.js.map +1 -0
- package/lib/core/registry.d.ts +22 -0
- package/lib/core/registry.d.ts.map +1 -0
- package/lib/core/registry.js +45 -0
- package/lib/core/registry.js.map +1 -0
- package/lib/core/rubric.d.ts +18 -0
- package/lib/core/rubric.d.ts.map +1 -0
- package/lib/core/rubric.js +49 -0
- package/lib/core/rubric.js.map +1 -0
- package/lib/core/run.d.ts +419 -0
- package/lib/core/run.d.ts.map +1 -0
- package/lib/core/run.js +126 -0
- package/lib/core/run.js.map +1 -0
- package/lib/core/statistics.d.ts +39 -0
- package/lib/core/statistics.d.ts.map +1 -0
- package/lib/core/statistics.js +163 -0
- package/lib/core/statistics.js.map +1 -0
- package/lib/core/subject.d.ts +4 -0
- package/lib/core/subject.d.ts.map +1 -0
- package/lib/core/subject.js +46 -0
- package/lib/core/subject.js.map +1 -0
- package/lib/core/tool-events.d.ts +69 -0
- package/lib/core/tool-events.d.ts.map +1 -0
- package/lib/core/tool-events.js +242 -0
- package/lib/core/tool-events.js.map +1 -0
- package/lib/core/usage.d.ts +78 -0
- package/lib/core/usage.d.ts.map +1 -0
- package/lib/core/usage.js +104 -0
- package/lib/core/usage.js.map +1 -0
- package/lib/core/wire.d.ts +83 -0
- package/lib/core/wire.d.ts.map +1 -0
- package/lib/core/wire.js +101 -0
- package/lib/core/wire.js.map +1 -0
- package/lib/core/zod.d.ts +10 -0
- package/lib/core/zod.d.ts.map +1 -0
- package/lib/core/zod.js +18 -0
- package/lib/core/zod.js.map +1 -0
- package/lib/drivers/claude-code/index.d.ts +55 -0
- package/lib/drivers/claude-code/index.d.ts.map +1 -0
- package/lib/drivers/claude-code/index.js +779 -0
- package/lib/drivers/claude-code/index.js.map +1 -0
- package/lib/drivers/claude-code/parse-events.d.ts +46 -0
- package/lib/drivers/claude-code/parse-events.d.ts.map +1 -0
- package/lib/drivers/claude-code/parse-events.js +387 -0
- package/lib/drivers/claude-code/parse-events.js.map +1 -0
- package/lib/drivers/codex/app-events.d.ts +16 -0
- package/lib/drivers/codex/app-events.d.ts.map +1 -0
- package/lib/drivers/codex/app-events.js +184 -0
- package/lib/drivers/codex/app-events.js.map +1 -0
- package/lib/drivers/codex/config.d.ts +55 -0
- package/lib/drivers/codex/config.d.ts.map +1 -0
- package/lib/drivers/codex/config.js +125 -0
- package/lib/drivers/codex/config.js.map +1 -0
- package/lib/drivers/codex/index.d.ts +8 -0
- package/lib/drivers/codex/index.d.ts.map +1 -0
- package/lib/drivers/codex/index.js +904 -0
- package/lib/drivers/codex/index.js.map +1 -0
- package/lib/drivers/mcp/index.d.ts +52 -0
- package/lib/drivers/mcp/index.d.ts.map +1 -0
- package/lib/drivers/mcp/index.js +282 -0
- package/lib/drivers/mcp/index.js.map +1 -0
- package/lib/drivers/mcp/transport.d.ts +50 -0
- package/lib/drivers/mcp/transport.d.ts.map +1 -0
- package/lib/drivers/mcp/transport.js +190 -0
- package/lib/drivers/mcp/transport.js.map +1 -0
- package/lib/drivers/shared/artifacts.d.ts +4 -0
- package/lib/drivers/shared/artifacts.d.ts.map +1 -0
- package/lib/drivers/shared/artifacts.js +81 -0
- package/lib/drivers/shared/artifacts.js.map +1 -0
- package/lib/drivers/shared/secure-copy.d.ts +2 -0
- package/lib/drivers/shared/secure-copy.d.ts.map +1 -0
- package/lib/drivers/shared/secure-copy.js +35 -0
- package/lib/drivers/shared/secure-copy.js.map +1 -0
- package/lib/drivers/subprocess/index.d.ts +24 -0
- package/lib/drivers/subprocess/index.d.ts.map +1 -0
- package/lib/drivers/subprocess/index.js +451 -0
- package/lib/drivers/subprocess/index.js.map +1 -0
- package/lib/index.d.ts +40 -0
- package/lib/index.d.ts.map +1 -0
- package/lib/index.js +48 -0
- package/lib/index.js.map +1 -0
- package/lib/optimization/reflective.d.ts +27 -0
- package/lib/optimization/reflective.d.ts.map +1 -0
- package/lib/optimization/reflective.js +257 -0
- package/lib/optimization/reflective.js.map +1 -0
- package/lib/orchestrator/annotation-loader.d.ts +65 -0
- package/lib/orchestrator/annotation-loader.d.ts.map +1 -0
- package/lib/orchestrator/annotation-loader.js +73 -0
- package/lib/orchestrator/annotation-loader.js.map +1 -0
- package/lib/orchestrator/annotation-store.d.ts +3 -0
- package/lib/orchestrator/annotation-store.d.ts.map +1 -0
- package/lib/orchestrator/annotation-store.js +100 -0
- package/lib/orchestrator/annotation-store.js.map +1 -0
- package/lib/orchestrator/authored-dependencies.d.ts +3 -0
- package/lib/orchestrator/authored-dependencies.d.ts.map +1 -0
- package/lib/orchestrator/authored-dependencies.js +146 -0
- package/lib/orchestrator/authored-dependencies.js.map +1 -0
- package/lib/orchestrator/campaign-lifecycle.d.ts +105 -0
- package/lib/orchestrator/campaign-lifecycle.d.ts.map +1 -0
- package/lib/orchestrator/campaign-lifecycle.js +141 -0
- package/lib/orchestrator/campaign-lifecycle.js.map +1 -0
- package/lib/orchestrator/candidate-integrity.d.ts +4 -0
- package/lib/orchestrator/candidate-integrity.d.ts.map +1 -0
- package/lib/orchestrator/candidate-integrity.js +14 -0
- package/lib/orchestrator/candidate-integrity.js.map +1 -0
- package/lib/orchestrator/candidates.d.ts +13 -0
- package/lib/orchestrator/candidates.d.ts.map +1 -0
- package/lib/orchestrator/candidates.js +577 -0
- package/lib/orchestrator/candidates.js.map +1 -0
- package/lib/orchestrator/compare.d.ts +80 -0
- package/lib/orchestrator/compare.d.ts.map +1 -0
- package/lib/orchestrator/compare.js +270 -0
- package/lib/orchestrator/compare.js.map +1 -0
- package/lib/orchestrator/comparison-identity.d.ts +10 -0
- package/lib/orchestrator/comparison-identity.d.ts.map +1 -0
- package/lib/orchestrator/comparison-identity.js +39 -0
- package/lib/orchestrator/comparison-identity.js.map +1 -0
- package/lib/orchestrator/config.d.ts +121 -0
- package/lib/orchestrator/config.d.ts.map +1 -0
- package/lib/orchestrator/config.js +176 -0
- package/lib/orchestrator/config.js.map +1 -0
- package/lib/orchestrator/content-bound-json.d.ts +5 -0
- package/lib/orchestrator/content-bound-json.d.ts.map +1 -0
- package/lib/orchestrator/content-bound-json.js +59 -0
- package/lib/orchestrator/content-bound-json.js.map +1 -0
- package/lib/orchestrator/discover.d.ts +45 -0
- package/lib/orchestrator/discover.d.ts.map +1 -0
- package/lib/orchestrator/discover.js +98 -0
- package/lib/orchestrator/discover.js.map +1 -0
- package/lib/orchestrator/doctor.d.ts +141 -0
- package/lib/orchestrator/doctor.d.ts.map +1 -0
- package/lib/orchestrator/doctor.js +525 -0
- package/lib/orchestrator/doctor.js.map +1 -0
- package/lib/orchestrator/driver-session.d.ts +177 -0
- package/lib/orchestrator/driver-session.d.ts.map +1 -0
- package/lib/orchestrator/driver-session.js +396 -0
- package/lib/orchestrator/driver-session.js.map +1 -0
- package/lib/orchestrator/evaluator-identity.d.ts +7 -0
- package/lib/orchestrator/evaluator-identity.d.ts.map +1 -0
- package/lib/orchestrator/evaluator-identity.js +42 -0
- package/lib/orchestrator/evaluator-identity.js.map +1 -0
- package/lib/orchestrator/experiment-budget.d.ts +11 -0
- package/lib/orchestrator/experiment-budget.d.ts.map +1 -0
- package/lib/orchestrator/experiment-budget.js +51 -0
- package/lib/orchestrator/experiment-budget.js.map +1 -0
- package/lib/orchestrator/experiment-context.d.ts +43 -0
- package/lib/orchestrator/experiment-context.d.ts.map +1 -0
- package/lib/orchestrator/experiment-context.js +9 -0
- package/lib/orchestrator/experiment-context.js.map +1 -0
- package/lib/orchestrator/experiment-contracts.d.ts +290 -0
- package/lib/orchestrator/experiment-contracts.d.ts.map +1 -0
- package/lib/orchestrator/experiment-contracts.js +38 -0
- package/lib/orchestrator/experiment-contracts.js.map +1 -0
- package/lib/orchestrator/experiment-corpus.d.ts +21 -0
- package/lib/orchestrator/experiment-corpus.d.ts.map +1 -0
- package/lib/orchestrator/experiment-corpus.js +276 -0
- package/lib/orchestrator/experiment-corpus.js.map +1 -0
- package/lib/orchestrator/experiment-evaluation.d.ts +44 -0
- package/lib/orchestrator/experiment-evaluation.d.ts.map +1 -0
- package/lib/orchestrator/experiment-evaluation.js +542 -0
- package/lib/orchestrator/experiment-evaluation.js.map +1 -0
- package/lib/orchestrator/experiment-finish.d.ts +6 -0
- package/lib/orchestrator/experiment-finish.d.ts.map +1 -0
- package/lib/orchestrator/experiment-finish.js +19 -0
- package/lib/orchestrator/experiment-finish.js.map +1 -0
- package/lib/orchestrator/experiment-identity.d.ts +6 -0
- package/lib/orchestrator/experiment-identity.d.ts.map +1 -0
- package/lib/orchestrator/experiment-identity.js +257 -0
- package/lib/orchestrator/experiment-identity.js.map +1 -0
- package/lib/orchestrator/experiment-loader.d.ts +175 -0
- package/lib/orchestrator/experiment-loader.d.ts.map +1 -0
- package/lib/orchestrator/experiment-loader.js +77 -0
- package/lib/orchestrator/experiment-loader.js.map +1 -0
- package/lib/orchestrator/experiment-report.d.ts +450 -0
- package/lib/orchestrator/experiment-report.d.ts.map +1 -0
- package/lib/orchestrator/experiment-report.js +570 -0
- package/lib/orchestrator/experiment-report.js.map +1 -0
- package/lib/orchestrator/experiment-runs.d.ts +379 -0
- package/lib/orchestrator/experiment-runs.d.ts.map +1 -0
- package/lib/orchestrator/experiment-runs.js +796 -0
- package/lib/orchestrator/experiment-runs.js.map +1 -0
- package/lib/orchestrator/experiment-runtime-identity.d.ts +41 -0
- package/lib/orchestrator/experiment-runtime-identity.d.ts.map +1 -0
- package/lib/orchestrator/experiment-runtime-identity.js +411 -0
- package/lib/orchestrator/experiment-runtime-identity.js.map +1 -0
- package/lib/orchestrator/experiment-selection.d.ts +25 -0
- package/lib/orchestrator/experiment-selection.d.ts.map +1 -0
- package/lib/orchestrator/experiment-selection.js +286 -0
- package/lib/orchestrator/experiment-selection.js.map +1 -0
- package/lib/orchestrator/experiment.d.ts +11 -0
- package/lib/orchestrator/experiment.d.ts.map +1 -0
- package/lib/orchestrator/experiment.js +639 -0
- package/lib/orchestrator/experiment.js.map +1 -0
- package/lib/orchestrator/fixture-identity.d.ts +7 -0
- package/lib/orchestrator/fixture-identity.d.ts.map +1 -0
- package/lib/orchestrator/fixture-identity.js +55 -0
- package/lib/orchestrator/fixture-identity.js.map +1 -0
- package/lib/orchestrator/fs-read.d.ts +19 -0
- package/lib/orchestrator/fs-read.d.ts.map +1 -0
- package/lib/orchestrator/fs-read.js +113 -0
- package/lib/orchestrator/fs-read.js.map +1 -0
- package/lib/orchestrator/grader-experiment.d.ts +190 -0
- package/lib/orchestrator/grader-experiment.d.ts.map +1 -0
- package/lib/orchestrator/grader-experiment.js +579 -0
- package/lib/orchestrator/grader-experiment.js.map +1 -0
- package/lib/orchestrator/grading.d.ts +31 -0
- package/lib/orchestrator/grading.d.ts.map +1 -0
- package/lib/orchestrator/grading.js +186 -0
- package/lib/orchestrator/grading.js.map +1 -0
- package/lib/orchestrator/index.d.ts +22 -0
- package/lib/orchestrator/index.d.ts.map +1 -0
- package/lib/orchestrator/index.js +30 -0
- package/lib/orchestrator/index.js.map +1 -0
- package/lib/orchestrator/lifecycle-hook-execution.d.ts +4 -0
- package/lib/orchestrator/lifecycle-hook-execution.d.ts.map +1 -0
- package/lib/orchestrator/lifecycle-hook-execution.js +67 -0
- package/lib/orchestrator/lifecycle-hook-execution.js.map +1 -0
- package/lib/orchestrator/list-runs.d.ts +33 -0
- package/lib/orchestrator/list-runs.d.ts.map +1 -0
- package/lib/orchestrator/list-runs.js +308 -0
- package/lib/orchestrator/list-runs.js.map +1 -0
- package/lib/orchestrator/loader.d.ts +3 -0
- package/lib/orchestrator/loader.d.ts.map +1 -0
- package/lib/orchestrator/loader.js +17 -0
- package/lib/orchestrator/loader.js.map +1 -0
- package/lib/orchestrator/optimization-lifecycle.d.ts +12 -0
- package/lib/orchestrator/optimization-lifecycle.d.ts.map +1 -0
- package/lib/orchestrator/optimization-lifecycle.js +92 -0
- package/lib/orchestrator/optimization-lifecycle.js.map +1 -0
- package/lib/orchestrator/orchestrator.d.ts +144 -0
- package/lib/orchestrator/orchestrator.d.ts.map +1 -0
- package/lib/orchestrator/orchestrator.js +632 -0
- package/lib/orchestrator/orchestrator.js.map +1 -0
- package/lib/orchestrator/owned-lock.d.ts +17 -0
- package/lib/orchestrator/owned-lock.d.ts.map +1 -0
- package/lib/orchestrator/owned-lock.js +149 -0
- package/lib/orchestrator/owned-lock.js.map +1 -0
- package/lib/orchestrator/package-version.d.ts +2 -0
- package/lib/orchestrator/package-version.d.ts.map +1 -0
- package/lib/orchestrator/package-version.js +4 -0
- package/lib/orchestrator/package-version.js.map +1 -0
- package/lib/orchestrator/provenance.d.ts +5 -0
- package/lib/orchestrator/provenance.d.ts.map +1 -0
- package/lib/orchestrator/provenance.js +137 -0
- package/lib/orchestrator/provenance.js.map +1 -0
- package/lib/orchestrator/regrade.d.ts +80 -0
- package/lib/orchestrator/regrade.d.ts.map +1 -0
- package/lib/orchestrator/regrade.js +433 -0
- package/lib/orchestrator/regrade.js.map +1 -0
- package/lib/orchestrator/rubric-loader.d.ts +7 -0
- package/lib/orchestrator/rubric-loader.d.ts.map +1 -0
- package/lib/orchestrator/rubric-loader.js +65 -0
- package/lib/orchestrator/rubric-loader.js.map +1 -0
- package/lib/orchestrator/run-dir.d.ts +40 -0
- package/lib/orchestrator/run-dir.d.ts.map +1 -0
- package/lib/orchestrator/run-dir.js +161 -0
- package/lib/orchestrator/run-dir.js.map +1 -0
- package/lib/orchestrator/run-execution.d.ts +30 -0
- package/lib/orchestrator/run-execution.d.ts.map +1 -0
- package/lib/orchestrator/run-execution.js +531 -0
- package/lib/orchestrator/run-execution.js.map +1 -0
- package/lib/orchestrator/run-lifecycle.d.ts +115 -0
- package/lib/orchestrator/run-lifecycle.d.ts.map +1 -0
- package/lib/orchestrator/run-lifecycle.js +136 -0
- package/lib/orchestrator/run-lifecycle.js.map +1 -0
- package/lib/orchestrator/suite-compare.d.ts +76 -0
- package/lib/orchestrator/suite-compare.d.ts.map +1 -0
- package/lib/orchestrator/suite-compare.js +159 -0
- package/lib/orchestrator/suite-compare.js.map +1 -0
- package/lib/orchestrator/suite-lifecycle.d.ts +157 -0
- package/lib/orchestrator/suite-lifecycle.d.ts.map +1 -0
- package/lib/orchestrator/suite-lifecycle.js +253 -0
- package/lib/orchestrator/suite-lifecycle.js.map +1 -0
- package/lib/orchestrator/suite-runs.d.ts +91 -0
- package/lib/orchestrator/suite-runs.d.ts.map +1 -0
- package/lib/orchestrator/suite-runs.js +157 -0
- package/lib/orchestrator/suite-runs.js.map +1 -0
- package/lib/orchestrator/suite-summary.d.ts +109 -0
- package/lib/orchestrator/suite-summary.d.ts.map +1 -0
- package/lib/orchestrator/suite-summary.js +213 -0
- package/lib/orchestrator/suite-summary.js.map +1 -0
- package/lib/orchestrator/trace-limits.d.ts +20 -0
- package/lib/orchestrator/trace-limits.d.ts.map +1 -0
- package/lib/orchestrator/trace-limits.js +28 -0
- package/lib/orchestrator/trace-limits.js.map +1 -0
- package/lib/orchestrator/validate-plugin-config.d.ts +10 -0
- package/lib/orchestrator/validate-plugin-config.d.ts.map +1 -0
- package/lib/orchestrator/validate-plugin-config.js +26 -0
- package/lib/orchestrator/validate-plugin-config.js.map +1 -0
- package/package.json +84 -0
- package/skills/lux/SKILL.md +85 -0
- package/skills/lux/references/cli.md +102 -0
- package/skills/lux/references/eval-authoring.md +105 -0
- package/skills/lux-answerer/SKILL.md +169 -0
- package/src/answerers/claude-code/index.ts +342 -0
- package/src/answerers/persona/index.ts +327 -0
- package/src/answerers/scripted/index.ts +186 -0
- package/src/answerers/shared.ts +23 -0
- package/src/answerers/terminal/index.ts +117 -0
- package/src/assertions/core/command-exits-zero.ts +94 -0
- package/src/assertions/core/event-checks.ts +138 -0
- package/src/assertions/core/file-checks.ts +336 -0
- package/src/assertions/core/helpers.ts +203 -0
- package/src/assertions/core/index.ts +80 -0
- package/src/assertions/core/json-pointer-equals.ts +84 -0
- package/src/assertions/core/output-checks.ts +134 -0
- package/src/assertions/core/predicate.ts +67 -0
- package/src/assertions/core/resource-checks.ts +180 -0
- package/src/assertions/core/safe-regex.ts +196 -0
- package/src/assertions/core/tool-called-with.ts +68 -0
- package/src/assertions/core/tool-calls.ts +28 -0
- package/src/assertions/core/tool-invocation.ts +198 -0
- package/src/assertions/core/tool-result-checks.ts +211 -0
- package/src/assertions/mcp/discovery-checks.ts +139 -0
- package/src/assertions/mcp/index.ts +37 -0
- package/src/assertions/rubric/index.ts +258 -0
- package/src/assertions/rubric/internal.ts +367 -0
- package/src/authoring.ts +2 -0
- package/src/cli/bin.ts +14 -0
- package/src/cli/command-runtime.ts +34 -0
- package/src/cli/init.ts +410 -0
- package/src/cli/program.ts +992 -0
- package/src/cli/render/format.ts +69 -0
- package/src/cli/render/reporter.ts +643 -0
- package/src/cli/render/tty.ts +314 -0
- package/src/cli/run-options.ts +82 -0
- package/src/cli/skills.ts +129 -0
- package/src/core/annotation.ts +104 -0
- package/src/core/answerer.ts +75 -0
- package/src/core/artifact-evidence.ts +249 -0
- package/src/core/artifact-snapshot.ts +429 -0
- package/src/core/assertion.ts +156 -0
- package/src/core/case.ts +103 -0
- package/src/core/conversation-turn.ts +47 -0
- package/src/core/driver.ts +75 -0
- package/src/core/emitter.ts +38 -0
- package/src/core/events.ts +135 -0
- package/src/core/exec.ts +263 -0
- package/src/core/experiment.ts +467 -0
- package/src/core/glob.ts +31 -0
- package/src/core/hash.ts +237 -0
- package/src/core/index.ts +28 -0
- package/src/core/lifecycle-fixtures.ts +145 -0
- package/src/core/lifecycle-hooks.ts +101 -0
- package/src/core/mcp-events.ts +57 -0
- package/src/core/model-cli.ts +271 -0
- package/src/core/observer.ts +50 -0
- package/src/core/platform-process.ts +284 -0
- package/src/core/registry.ts +57 -0
- package/src/core/rubric.ts +52 -0
- package/src/core/run.ts +144 -0
- package/src/core/statistics.ts +194 -0
- package/src/core/subject.ts +50 -0
- package/src/core/tool-events.ts +292 -0
- package/src/core/usage.ts +123 -0
- package/src/core/wire.ts +135 -0
- package/src/core/zod.ts +20 -0
- package/src/drivers/claude-code/index.ts +866 -0
- package/src/drivers/claude-code/parse-events.ts +464 -0
- package/src/drivers/codex/app-events.ts +218 -0
- package/src/drivers/codex/config.ts +128 -0
- package/src/drivers/codex/index.ts +995 -0
- package/src/drivers/mcp/index.ts +351 -0
- package/src/drivers/mcp/transport.ts +222 -0
- package/src/drivers/shared/artifacts.ts +86 -0
- package/src/drivers/shared/secure-copy.ts +40 -0
- package/src/drivers/subprocess/index.ts +490 -0
- package/src/index.ts +594 -0
- package/src/optimization/reflective.ts +319 -0
- package/src/orchestrator/annotation-loader.ts +97 -0
- package/src/orchestrator/annotation-store.ts +112 -0
- package/src/orchestrator/authored-dependencies.ts +193 -0
- package/src/orchestrator/campaign-lifecycle.ts +181 -0
- package/src/orchestrator/candidate-integrity.ts +18 -0
- package/src/orchestrator/candidates.ts +683 -0
- package/src/orchestrator/compare.ts +358 -0
- package/src/orchestrator/comparison-identity.ts +45 -0
- package/src/orchestrator/config.ts +205 -0
- package/src/orchestrator/content-bound-json.ts +71 -0
- package/src/orchestrator/discover.ts +132 -0
- package/src/orchestrator/doctor.ts +781 -0
- package/src/orchestrator/driver-session.ts +507 -0
- package/src/orchestrator/evaluator-identity.ts +45 -0
- package/src/orchestrator/experiment-budget.ts +61 -0
- package/src/orchestrator/experiment-context.ts +41 -0
- package/src/orchestrator/experiment-contracts.ts +43 -0
- package/src/orchestrator/experiment-corpus.ts +416 -0
- package/src/orchestrator/experiment-evaluation.ts +749 -0
- package/src/orchestrator/experiment-finish.ts +33 -0
- package/src/orchestrator/experiment-identity.ts +356 -0
- package/src/orchestrator/experiment-loader.ts +105 -0
- package/src/orchestrator/experiment-report.ts +788 -0
- package/src/orchestrator/experiment-runs.ts +960 -0
- package/src/orchestrator/experiment-runtime-identity.ts +465 -0
- package/src/orchestrator/experiment-selection.ts +467 -0
- package/src/orchestrator/experiment.ts +991 -0
- package/src/orchestrator/fixture-identity.ts +67 -0
- package/src/orchestrator/fs-read.ts +126 -0
- package/src/orchestrator/grader-experiment.ts +856 -0
- package/src/orchestrator/grading.ts +260 -0
- package/src/orchestrator/index.ts +186 -0
- package/src/orchestrator/lifecycle-hook-execution.ts +90 -0
- package/src/orchestrator/list-runs.ts +384 -0
- package/src/orchestrator/loader.ts +19 -0
- package/src/orchestrator/optimization-lifecycle.ts +126 -0
- package/src/orchestrator/orchestrator.ts +919 -0
- package/src/orchestrator/owned-lock.ts +147 -0
- package/src/orchestrator/package-version.ts +5 -0
- package/src/orchestrator/provenance.ts +140 -0
- package/src/orchestrator/regrade.ts +580 -0
- package/src/orchestrator/rubric-loader.ts +74 -0
- package/src/orchestrator/run-dir.ts +209 -0
- package/src/orchestrator/run-execution.ts +694 -0
- package/src/orchestrator/run-lifecycle.ts +174 -0
- package/src/orchestrator/suite-compare.ts +195 -0
- package/src/orchestrator/suite-lifecycle.ts +335 -0
- package/src/orchestrator/suite-runs.ts +202 -0
- package/src/orchestrator/suite-summary.ts +276 -0
- package/src/orchestrator/trace-limits.ts +65 -0
- package/src/orchestrator/validate-plugin-config.ts +34 -0
package/package.json
ADDED
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "@prismatic-io/lux",
|
|
3
|
+
"version": "0.0.1",
|
|
4
|
+
"description": "Coding-agent evaluation and improvement with deterministic assertions, optional LLM judges, and first-class human-in-the-loop runs.",
|
|
5
|
+
"license": "MIT",
|
|
6
|
+
"homepage": "https://github.com/prismatic-io/lux",
|
|
7
|
+
"bugs": {
|
|
8
|
+
"url": "https://github.com/prismatic-io/lux/issues"
|
|
9
|
+
},
|
|
10
|
+
"repository": {
|
|
11
|
+
"type": "git",
|
|
12
|
+
"url": "git+https://github.com/prismatic-io/lux.git",
|
|
13
|
+
"directory": "packages/lux"
|
|
14
|
+
},
|
|
15
|
+
"type": "module",
|
|
16
|
+
"keywords": [
|
|
17
|
+
"evals",
|
|
18
|
+
"evaluation",
|
|
19
|
+
"agent",
|
|
20
|
+
"coding-agent",
|
|
21
|
+
"human-in-the-loop",
|
|
22
|
+
"claude",
|
|
23
|
+
"claude-code",
|
|
24
|
+
"codex",
|
|
25
|
+
"openai",
|
|
26
|
+
"llm-judge",
|
|
27
|
+
"testing"
|
|
28
|
+
],
|
|
29
|
+
"sideEffects": false,
|
|
30
|
+
"engines": {
|
|
31
|
+
"node": ">=22.18.0"
|
|
32
|
+
},
|
|
33
|
+
"main": "./lib/index.js",
|
|
34
|
+
"types": "./lib/index.d.ts",
|
|
35
|
+
"bin": {
|
|
36
|
+
"lux": "./lib/cli/bin.js"
|
|
37
|
+
},
|
|
38
|
+
"exports": {
|
|
39
|
+
".": {
|
|
40
|
+
"types": "./lib/index.d.ts",
|
|
41
|
+
"default": "./lib/index.js"
|
|
42
|
+
},
|
|
43
|
+
"./artifact-snapshot": {
|
|
44
|
+
"types": "./lib/core/artifact-snapshot.d.ts",
|
|
45
|
+
"default": "./lib/core/artifact-snapshot.js"
|
|
46
|
+
},
|
|
47
|
+
"./authoring": {
|
|
48
|
+
"types": "./lib/authoring.d.ts",
|
|
49
|
+
"default": "./lib/authoring.js"
|
|
50
|
+
},
|
|
51
|
+
"./package.json": "./package.json"
|
|
52
|
+
},
|
|
53
|
+
"publishConfig": {
|
|
54
|
+
"access": "public"
|
|
55
|
+
},
|
|
56
|
+
"files": [
|
|
57
|
+
"lib",
|
|
58
|
+
"skills",
|
|
59
|
+
"src",
|
|
60
|
+
"!**/*.test.ts",
|
|
61
|
+
"!**/test-helpers.ts",
|
|
62
|
+
"!src/**/fixtures/**"
|
|
63
|
+
],
|
|
64
|
+
"scripts": {
|
|
65
|
+
"build": "tsc --build",
|
|
66
|
+
"clean": "node ../../scripts/clean.mjs"
|
|
67
|
+
},
|
|
68
|
+
"devDependencies": {
|
|
69
|
+
"@arethetypeswrong/cli": "0.18.4",
|
|
70
|
+
"publint": "0.3.21"
|
|
71
|
+
},
|
|
72
|
+
"dependencies": {
|
|
73
|
+
"@modelcontextprotocol/sdk": "^1.30.0",
|
|
74
|
+
"@vscode/windows-process-tree": "0.8.0",
|
|
75
|
+
"commander": "^15.0.0",
|
|
76
|
+
"js-yaml": "^5.2.2",
|
|
77
|
+
"oxc-parser": "^0.142.0",
|
|
78
|
+
"tempy": "^3.2.0",
|
|
79
|
+
"tinyexec": "1.2.4",
|
|
80
|
+
"tinyglobby": "^0.2.17",
|
|
81
|
+
"xstate": "5.32.0",
|
|
82
|
+
"zod": "^4.4.3"
|
|
83
|
+
}
|
|
84
|
+
}
|
|
@@ -0,0 +1,85 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: lux
|
|
3
|
+
description: Author, validate, run, inspect, compare, and improve coding-agent evaluations with the Lux CLI. Use when creating Lux cases or campaigns, choosing assertions/personas/drivers, diagnosing eval failures, comparing runs, or running Lux experiments.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Lux
|
|
7
|
+
|
|
8
|
+
Use Lux as an eval test runner and evidence system for coding agents. Prefer
|
|
9
|
+
small, discriminating cases with deterministic assertions; add an LLM rubric
|
|
10
|
+
only for qualities that artifacts and commands cannot establish.
|
|
11
|
+
|
|
12
|
+
## Start here
|
|
13
|
+
|
|
14
|
+
1. Inspect the repository and its existing `lux.config.ts`, cases, fixtures,
|
|
15
|
+
prompts, and experiments before adding files.
|
|
16
|
+
2. If Lux is not initialized, run `npm exec -- lux init` for ordinary evals or
|
|
17
|
+
`npm exec -- lux init --experiment` for an improvement campaign.
|
|
18
|
+
3. Read [references/eval-authoring.md](references/eval-authoring.md) before
|
|
19
|
+
authoring or reviewing a case.
|
|
20
|
+
4. Read [references/cli.md](references/cli.md) before running, diagnosing,
|
|
21
|
+
comparing, or promoting results.
|
|
22
|
+
5. Run `npm exec -- lux doctor` before model calls. Treat its errors as
|
|
23
|
+
authoring problems, not as permission to bypass validation.
|
|
24
|
+
|
|
25
|
+
## Workflow
|
|
26
|
+
|
|
27
|
+
### Author an eval
|
|
28
|
+
|
|
29
|
+
- Define the behavior and likely failure mode first.
|
|
30
|
+
- Make the prompt realistic and self-contained.
|
|
31
|
+
- Give the persona only facts and preferences the simulated user should know.
|
|
32
|
+
- Assert observable outcomes, not implementation details.
|
|
33
|
+
- Use stable assertion IDs so grading evidence remains intelligible.
|
|
34
|
+
- Use the cheapest deterministic assertion that proves each requirement.
|
|
35
|
+
- Include a rubric assertion only when semantic judgment is unavoidable, and
|
|
36
|
+
write one narrow, falsifiable criterion per assertion.
|
|
37
|
+
- Pin the subject driver model explicitly when comparability matters.
|
|
38
|
+
|
|
39
|
+
### Run and diagnose
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
npm exec -- lux doctor
|
|
43
|
+
npm exec -- lux run <case-or-filter> --loop 3
|
|
44
|
+
npm exec -- lux view <run-or-suite>
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
Inspect `case.json`, `run.json`, `events.jsonl`, `grading.json`, and `artifacts/`
|
|
48
|
+
before changing the case. Separate subject failure from harness failure:
|
|
49
|
+
|
|
50
|
+
- Subject failure: the agent misunderstood, asked poorly, used the wrong tools,
|
|
51
|
+
or left incorrect artifacts.
|
|
52
|
+
- Harness failure: the prompt was underspecified, the persona lacked a needed
|
|
53
|
+
fact, an assertion measured a proxy, or the judge criterion was ambiguous.
|
|
54
|
+
|
|
55
|
+
Do not weaken a valid assertion merely to make a run pass.
|
|
56
|
+
|
|
57
|
+
### Improve with evidence
|
|
58
|
+
|
|
59
|
+
Use repeated runs and held-out splits for noisy agent behavior. For skill,
|
|
60
|
+
prompt, or agent-source changes, use `subjectPath()` in cases and an experiment
|
|
61
|
+
campaign with immutable candidates. Plan first, run the experiment, inspect the
|
|
62
|
+
report, and apply a promoted candidate as a separate action:
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
npm exec -- lux experiment experiments/example.experiment.ts --plan
|
|
66
|
+
npm exec -- lux experiment experiments/example.experiment.ts
|
|
67
|
+
npm exec -- lux report <experiment-dir>
|
|
68
|
+
npm exec -- lux apply <experiment-dir>
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
Never edit a materialized candidate or managed experiment record by hand.
|
|
72
|
+
|
|
73
|
+
## Guardrails
|
|
74
|
+
|
|
75
|
+
- Lux eval projects execute project-authored commands and agent-generated work.
|
|
76
|
+
Run untrusted evals only in a disposable environment.
|
|
77
|
+
- Do not infer success from agent prose. Grade the final artifacts and recorded
|
|
78
|
+
behavior.
|
|
79
|
+
- Do not use only `run-succeeded` for a case that claims to verify a concrete
|
|
80
|
+
result.
|
|
81
|
+
- Do not let train examples leak into validation or test splits.
|
|
82
|
+
- Do not combine evidence across changed models, harnesses, cases, or subjects;
|
|
83
|
+
Lux identity checks exist to prevent that.
|
|
84
|
+
- Prefer `--plan`, `doctor`, `view`, and `report` before expensive or mutating
|
|
85
|
+
commands.
|
|
@@ -0,0 +1,102 @@
|
|
|
1
|
+
# Lux CLI runbook
|
|
2
|
+
|
|
3
|
+
Invoke the project-local CLI with `npm exec -- lux ...` unless the repository
|
|
4
|
+
uses another package-manager convention.
|
|
5
|
+
|
|
6
|
+
## Author and preflight
|
|
7
|
+
|
|
8
|
+
```bash
|
|
9
|
+
lux init
|
|
10
|
+
lux init --experiment
|
|
11
|
+
lux doctor
|
|
12
|
+
lux doctor experiments/campaign.experiment.ts --json
|
|
13
|
+
lux run --list
|
|
14
|
+
lux run --list-tags
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
`init` is conservative around existing files. `doctor` validates configuration,
|
|
18
|
+
case discovery, driver executables, campaign splits, confidence feasibility,
|
|
19
|
+
and projected budgets without model calls.
|
|
20
|
+
|
|
21
|
+
## Run
|
|
22
|
+
|
|
23
|
+
```bash
|
|
24
|
+
lux run
|
|
25
|
+
lux run retry-flag
|
|
26
|
+
lux run cases/retry-flag.ts
|
|
27
|
+
lux run --tag smoke
|
|
28
|
+
lux run retry --loop 5 --concurrency 2
|
|
29
|
+
lux run one-case --interactive
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
Filters match case IDs and paths. Tags come from `meta.tags`. A run invocation
|
|
33
|
+
also creates a suite record, which is the right unit for repeated or multi-case
|
|
34
|
+
comparison.
|
|
35
|
+
|
|
36
|
+
## Inspect and compare
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
lux view
|
|
40
|
+
lux view <run-or-suite>
|
|
41
|
+
lux view --experiments
|
|
42
|
+
lux compare <run-or-suite-a> <run-or-suite-b>
|
|
43
|
+
```
|
|
44
|
+
|
|
45
|
+
Use `view` before opening raw files. When diagnosing, inspect:
|
|
46
|
+
|
|
47
|
+
- `case.json`: frozen authored inputs;
|
|
48
|
+
- `run.json`: exit reason, timing, usage, source identity;
|
|
49
|
+
- `events.jsonl`: trajectory, tools, and interrupts;
|
|
50
|
+
- `grading.json`: assertion evidence;
|
|
51
|
+
- `artifacts/`: the subject workspace output.
|
|
52
|
+
|
|
53
|
+
`compare` exits nonzero on a gated regression, making it suitable for CI.
|
|
54
|
+
|
|
55
|
+
## Regrade and annotate
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
lux grade <run>
|
|
59
|
+
lux grade <run> --case cases/changed-grader.ts
|
|
60
|
+
lux annotate <run> \
|
|
61
|
+
--out evals/annotations.json \
|
|
62
|
+
--rubric evals/rubrics/case.json \
|
|
63
|
+
--label pass
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
Regrading preserves subject evidence and changes only grading. Human annotations
|
|
67
|
+
support grader-alignment experiments; keep assertion IDs aligned with the
|
|
68
|
+
rubric and record reviewer context where available.
|
|
69
|
+
|
|
70
|
+
## Experiment lifecycle
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
lux experiment <campaign> --plan
|
|
74
|
+
lux experiment <campaign>
|
|
75
|
+
lux optimize <campaign> --plan
|
|
76
|
+
lux optimize <campaign>
|
|
77
|
+
lux optimize <campaign> --resume <experiment-dir>
|
|
78
|
+
lux view <experiment-dir>
|
|
79
|
+
lux report <experiment-dir> --out .lux-runs/reports/decision.md
|
|
80
|
+
lux apply <experiment-dir>
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
`experiment` compares authored variants. `optimize` searches proposed changes.
|
|
84
|
+
Both create durable, resumable, integrity-checked records. `report` re-verifies
|
|
85
|
+
the ledger and promotion decision. `apply` is deliberately separate and
|
|
86
|
+
hash-guards the source it changes.
|
|
87
|
+
|
|
88
|
+
Use the exact resume, report, and apply commands printed by Lux. Do not edit
|
|
89
|
+
`.lux-runs/.experiments/` manually.
|
|
90
|
+
|
|
91
|
+
## Failure triage
|
|
92
|
+
|
|
93
|
+
- Discovery/config error: run `doctor`, verify paths and exported definitions.
|
|
94
|
+
- Driver unavailable: verify the selected CLI is installed and the model is
|
|
95
|
+
explicitly configured where required.
|
|
96
|
+
- Agent exit/error: inspect `events.jsonl` and the driver section of `run.json`.
|
|
97
|
+
- Assertion failure: inspect the assertion evidence and final artifact; decide
|
|
98
|
+
whether the subject or measurement is wrong.
|
|
99
|
+
- Judge disagreement: narrow the rubric, regrade stored evidence, and collect
|
|
100
|
+
human annotations before changing the subject.
|
|
101
|
+
- Interrupted experiment: resume the printed experiment directory; do not start
|
|
102
|
+
a nominally identical campaign and combine results yourself.
|
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
# Authoring useful Lux evals
|
|
2
|
+
|
|
3
|
+
## Case shape
|
|
4
|
+
|
|
5
|
+
A case default-exports `defineEvalCase`:
|
|
6
|
+
|
|
7
|
+
```ts
|
|
8
|
+
import { defineEvalCase } from "@prismatic-io/lux";
|
|
9
|
+
|
|
10
|
+
export default defineEvalCase({
|
|
11
|
+
id: "retry-flag",
|
|
12
|
+
prompt: "Add a --retry <n> flag to the fetch command.",
|
|
13
|
+
persona: "You are the maintainer. If asked, retries default to 3.",
|
|
14
|
+
driver: { name: "codex", config: { model: "YOUR_CODEX_MODEL" } },
|
|
15
|
+
assertions: [
|
|
16
|
+
{ id: "flag-present", type: "file-contains", path: "src/cli.ts", text: "--retry" },
|
|
17
|
+
{ id: "tests-pass", type: "command-exits-zero", command: "npm test" },
|
|
18
|
+
{ id: "test-coverage", type: "rubric", criteria: "A test covers the new retry flag." },
|
|
19
|
+
],
|
|
20
|
+
meta: { tags: ["cli", "validation"] },
|
|
21
|
+
});
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Use the project-level default driver when all cases share it. Case-level
|
|
25
|
+
drivers are useful for intentional cross-agent or cross-model comparisons.
|
|
26
|
+
|
|
27
|
+
## Design checklist
|
|
28
|
+
|
|
29
|
+
Before writing:
|
|
30
|
+
|
|
31
|
+
1. State the capability or regression in one sentence.
|
|
32
|
+
2. Identify the artifact, command result, or interaction that proves it.
|
|
33
|
+
3. List plausible false positives. Adjust assertions to reject them.
|
|
34
|
+
4. Decide whether the case needs a fixture and whether each run must begin from
|
|
35
|
+
an isolated copy.
|
|
36
|
+
5. Identify any question a competent agent must ask. Put only the answer and
|
|
37
|
+
relevant preferences in the persona.
|
|
38
|
+
|
|
39
|
+
After writing:
|
|
40
|
+
|
|
41
|
+
1. Run `lux doctor`.
|
|
42
|
+
2. List discovery with `lux run <filter> --list`.
|
|
43
|
+
3. Run the case more than once if agent behavior can vary.
|
|
44
|
+
4. Inspect events and artifacts, including failed runs.
|
|
45
|
+
5. Confirm each assertion fails when its requirement is deliberately absent.
|
|
46
|
+
|
|
47
|
+
## Assertion selection
|
|
48
|
+
|
|
49
|
+
Prefer deterministic assertions:
|
|
50
|
+
|
|
51
|
+
- `run-succeeded`: the driver completed; necessary but rarely sufficient.
|
|
52
|
+
- `file-exists`: a required artifact exists.
|
|
53
|
+
- `file-contains`: a stable literal is present.
|
|
54
|
+
- `json-path-equals`: structured output contains the expected value.
|
|
55
|
+
- `command-exits-zero`: project-authored validation succeeds.
|
|
56
|
+
|
|
57
|
+
Use `rubric` for semantic properties such as design quality, completeness, or
|
|
58
|
+
whether a test meaningfully covers behavior. Keep criteria narrow. Avoid words
|
|
59
|
+
like “good,” “proper,” or “best practice” without observable conditions.
|
|
60
|
+
|
|
61
|
+
Assertions should test user-visible behavior and durable contracts. Avoid
|
|
62
|
+
requiring a specific function name, file layout, or algorithm unless that is
|
|
63
|
+
the contract under evaluation.
|
|
64
|
+
|
|
65
|
+
## Human-in-the-loop behavior
|
|
66
|
+
|
|
67
|
+
The persona is simulated user state, not another system prompt. Include:
|
|
68
|
+
|
|
69
|
+
- facts the user would know;
|
|
70
|
+
- preferences needed to resolve legitimate ambiguity;
|
|
71
|
+
- approval boundaries relevant to the task.
|
|
72
|
+
|
|
73
|
+
Do not tell the persona how to help the agent solve the task. If no interaction
|
|
74
|
+
is part of the capability being evaluated, omit the persona or keep it minimal.
|
|
75
|
+
|
|
76
|
+
Use `interactionMode: "defer-resume"` for automated Claude Code question-tool
|
|
77
|
+
handling. Use `textQuestionFallback: true` only when intentionally evaluating
|
|
78
|
+
models that ask in prose instead of using the supported tool.
|
|
79
|
+
|
|
80
|
+
## Experiments for prompts, skills, and agent source
|
|
81
|
+
|
|
82
|
+
Bind mutable source through `subjectPath()`:
|
|
83
|
+
|
|
84
|
+
```ts
|
|
85
|
+
import { defineEvalCase, subjectPath } from "@prismatic-io/lux";
|
|
86
|
+
|
|
87
|
+
export default defineEvalCase({
|
|
88
|
+
id: "skill-routing",
|
|
89
|
+
prompt: `Use the skill at ${subjectPath("plugin/skills/routing/SKILL.md")} to complete the task.`,
|
|
90
|
+
assertions: [{ id: "completed", type: "run-succeeded" }],
|
|
91
|
+
meta: { tags: ["validation"] },
|
|
92
|
+
});
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
Campaigns should:
|
|
96
|
+
|
|
97
|
+
- restrict `subject.mutable` to the intended source;
|
|
98
|
+
- define named, explicit driver/model profiles;
|
|
99
|
+
- keep train, validation, and test selectors disjoint;
|
|
100
|
+
- budget calls, tokens, cost, and wall time;
|
|
101
|
+
- use repetitions and confidence gates appropriate to stochastic outcomes;
|
|
102
|
+
- reserve held-out cases for promotion and final confirmation.
|
|
103
|
+
|
|
104
|
+
Use `experiment` for authored variants and `optimize` for model-proposed source
|
|
105
|
+
changes. Always run the corresponding `--plan` command first.
|
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: lux-answerer
|
|
3
|
+
version: 0.0.1
|
|
4
|
+
description: Play the persona for a running Lux orchestrator that uses the claude-code answerer. Read structured events on stdout, decide answers from persona + context, write JSON to its answer channel.
|
|
5
|
+
user-invocable: false
|
|
6
|
+
allowed-tools: Bash, Read
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
# Lux answerer
|
|
10
|
+
|
|
11
|
+
This skill is the runbook for driving a Lux orchestrator from a
|
|
12
|
+
Claude Code session. The orchestrator's claude-code answerer emits
|
|
13
|
+
structured events on its stdout; you read them, decide an answer in
|
|
14
|
+
character, and write the answer to the channel it provides. Lux uses a FIFO on
|
|
15
|
+
Linux and macOS and a regular answer file on Windows.
|
|
16
|
+
|
|
17
|
+
## TL;DR
|
|
18
|
+
|
|
19
|
+
1. Find the orchestrator's stdout — usually it's a Bash process you
|
|
20
|
+
started, so its output is right there in the conversation.
|
|
21
|
+
2. The first interesting event is `lux.answerer.ready`. **Capture
|
|
22
|
+
`fifoPath`, `transport`, `persona`, and `prompt` from it.** Commit to the
|
|
23
|
+
persona for the rest of the run.
|
|
24
|
+
3. For each `lux.answerer.question` event:
|
|
25
|
+
- Read the embedded `interrupt` (kind, id, question, context).
|
|
26
|
+
- Decide the answer **strictly from persona + question + context.**
|
|
27
|
+
- Write one newline-terminated JSON answer to `<fifoPath>` using the
|
|
28
|
+
event's `transport`.
|
|
29
|
+
4. The orchestrator finishes and exits when the agent reports done.
|
|
30
|
+
|
|
31
|
+
## 1. The wire protocol
|
|
32
|
+
|
|
33
|
+
### Events you receive (orchestrator → you, on stdout)
|
|
34
|
+
|
|
35
|
+
| Type | Payload | Action |
|
|
36
|
+
|------|---------|--------|
|
|
37
|
+
| `lux.answerer.ready` | `{ fifoPath, transport, persona, prompt, runDir }` | Capture the answer path, transport, and persona. |
|
|
38
|
+
| `lux.answerer.question` | `{ interrupt, historyDepth }` | Decide an answer; write to `fifoPath`. |
|
|
39
|
+
|
|
40
|
+
The `interrupt` field is the standard Lux Interrupt:
|
|
41
|
+
|
|
42
|
+
```ts
|
|
43
|
+
type Interrupt =
|
|
44
|
+
| { kind: "ask"; id: string; question: string; context?: unknown }
|
|
45
|
+
| { kind: "approve"; id: string; request: string; context?: unknown }
|
|
46
|
+
| { kind: "custom"; id: string; payload: unknown; context?: unknown };
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
### Commands you send (you → orchestrator, written to the answer channel)
|
|
50
|
+
|
|
51
|
+
One JSON object per line. Newline required. Always echo the
|
|
52
|
+
interrupt's `id` so the answerer can verify the reply pairs with the
|
|
53
|
+
question it is waiting on.
|
|
54
|
+
|
|
55
|
+
For `ask` interrupts:
|
|
56
|
+
```
|
|
57
|
+
{"id": "<interrupt id>", "text": "<your answer>"}
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
For `approve` interrupts:
|
|
61
|
+
```
|
|
62
|
+
{"id": "<interrupt id>", "approved": true, "reason": "<optional reason>"}
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
For `custom` interrupts:
|
|
66
|
+
```
|
|
67
|
+
{"id": "<interrupt id>", "payload": <whatever shape the interrupt expects>}
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
### Examples
|
|
71
|
+
|
|
72
|
+
**Answer on a FIFO** (`transport: "fifo"`, from Bash):
|
|
73
|
+
|
|
74
|
+
```
|
|
75
|
+
printf '%s\n' '{"id":"q-1","text":"Use API key authentication."}' > "$LUX_FIFO"
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
**Answer in a regular file** (`transport: "file"`, portable Node command):
|
|
79
|
+
|
|
80
|
+
```
|
|
81
|
+
node -e "require('node:fs').writeFileSync(process.argv[1], process.argv[2] + '\n')" "$LUX_FIFO" '{"id":"q-1","text":"Use API key authentication."}'
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
The same JSON shapes apply to approvals:
|
|
85
|
+
|
|
86
|
+
```
|
|
87
|
+
{"id":"a-1","approved":true,"reason":"plan looks correct"}
|
|
88
|
+
{"id":"a-1","approved":false,"reason":"step 3 should use Bash, not Code"}
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
## 2. Decision-making during a run
|
|
92
|
+
|
|
93
|
+
The persona is the single source of truth. When deciding an answer:
|
|
94
|
+
|
|
95
|
+
1. **Read the persona first.** It defines who you are and what you
|
|
96
|
+
want out of this run. Stay in character.
|
|
97
|
+
2. **Consult the embedded `context`.** Drivers attach what they
|
|
98
|
+
think you need — recent assistant text, recent tool calls. This
|
|
99
|
+
usually disambiguates ambiguous questions.
|
|
100
|
+
3. **If the persona supplies an explicit value** for the question
|
|
101
|
+
(a URL, an auth method, a config var name), use it verbatim.
|
|
102
|
+
4. **If the persona is silent** on a detail, choose a value
|
|
103
|
+
consistent with the persona's role and the run's overall goal.
|
|
104
|
+
Don't fabricate contradictory specifics; don't ask the agent for
|
|
105
|
+
help (you ARE the human in this loop).
|
|
106
|
+
5. **Keep answers concise.** One line where possible. Long answers
|
|
107
|
+
slow the run down and add noise.
|
|
108
|
+
|
|
109
|
+
## 3. Critical rules
|
|
110
|
+
|
|
111
|
+
- **Always terminate the JSON with a newline.** For a FIFO, use `printf`,
|
|
112
|
+
never `echo`. For a regular file, overwrite the file with one complete
|
|
113
|
+
newline-terminated answer.
|
|
114
|
+
- **Single-quote the JSON** in shell, double-quote inside the JSON.
|
|
115
|
+
If the answer text contains a literal single quote, escape via
|
|
116
|
+
`'\''` or rephrase.
|
|
117
|
+
- **Echo the interrupt id in your reply.** A reply whose `id` does
|
|
118
|
+
not match the interrupt currently awaiting an answer is silently
|
|
119
|
+
discarded — this protects the run from a late reply to an already
|
|
120
|
+
timed-out question being misread as the answer to the next one.
|
|
121
|
+
A reply without an `id` is paired with the most-recent outstanding
|
|
122
|
+
interrupt, so omitting it forfeits that protection.
|
|
123
|
+
- **Don't pre-commit answers.** Wait for the structured event;
|
|
124
|
+
don't guess what's coming next.
|
|
125
|
+
|
|
126
|
+
## 4. Failure modes and recovery
|
|
127
|
+
|
|
128
|
+
### Answer wait timed out (no event is emitted)
|
|
129
|
+
|
|
130
|
+
If no reply lands on the answer channel within the answerer's wait budget
|
|
131
|
+
(`maxAnswerWaitMs`, default 10 minutes), the pending answer is
|
|
132
|
+
rejected and the run fails — there is no further stdout event to
|
|
133
|
+
watch for. Separately, the driver can hit its own idle window if the
|
|
134
|
+
agent stalls mid-LLM-call. Either way, inspect the run dir
|
|
135
|
+
(`<runDir>/events.jsonl`) to see how far the run got.
|
|
136
|
+
|
|
137
|
+
### Wrong answer shape
|
|
138
|
+
|
|
139
|
+
If the answerer rejects your answer (e.g. you sent `{"approved":...}`
|
|
140
|
+
to an `ask` interrupt), it logs an error and the run fails. Re-read
|
|
141
|
+
the interrupt kind before composing your reply.
|
|
142
|
+
|
|
143
|
+
### Answer write hangs
|
|
144
|
+
|
|
145
|
+
A FIFO write can block when the orchestrator has already exited. A regular-file
|
|
146
|
+
answer may be ignored after the question has timed out. Check whether the
|
|
147
|
+
orchestrator process is still running.
|
|
148
|
+
|
|
149
|
+
## 5. After the run
|
|
150
|
+
|
|
151
|
+
The run dir contains everything captured:
|
|
152
|
+
|
|
153
|
+
- `case.json` — frozen prompt + persona used for this run.
|
|
154
|
+
- `run.json` — exit reason, duration, interrupt counts, tokens.
|
|
155
|
+
- `events.jsonl` — every event the orchestrator processed.
|
|
156
|
+
- `grading.json` — per-assertion results.
|
|
157
|
+
- `artifacts/` — files the agent wrote.
|
|
158
|
+
|
|
159
|
+
Use `/lux:view <runDir>` for a formatted summary.
|
|
160
|
+
|
|
161
|
+
## 6. What you cannot do from this skill
|
|
162
|
+
|
|
163
|
+
- **Pause the agent.** Once an interrupt is resolved, the agent
|
|
164
|
+
runs until the next interrupt or `done`. There is no
|
|
165
|
+
pause/resume primitive.
|
|
166
|
+
- **Revise a resolved interrupt.** Once a command is consumed, the
|
|
167
|
+
interrupt is gone. If your answer was wrong, the only recovery is
|
|
168
|
+
to wait for a follow-up question (if any) and course-correct
|
|
169
|
+
there.
|