value-stream 0.1.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- value_stream-0.1.1/.agent/PLANS.md +150 -0
- value_stream-0.1.1/.agent/SIMULATION_APP_EXECPLAN.md +353 -0
- value_stream-0.1.1/.agent/SIMULATION_SERVICE_EXECPLAN.md +114 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/README.md +15 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/chart.png +0 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/desktop.png +0 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/docker-validation.json +35 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/final-smoke.json +20 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/mobile.png +0 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/validate_runtime.py +146 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/wheel-validation.json +30 -0
- value_stream-0.1.1/.agent/artifacts/simulation-app/zoom-200.png +0 -0
- value_stream-0.1.1/.codex/config.toml +5 -0
- value_stream-0.1.1/.dockerignore +12 -0
- value_stream-0.1.1/.github/agents/Reviewer.agent.md +21 -0
- value_stream-0.1.1/.github/workflows/ci.yml +115 -0
- value_stream-0.1.1/.gitignore +185 -0
- value_stream-0.1.1/.vscode/extensions.json +12 -0
- value_stream-0.1.1/.vscode/launch.json +113 -0
- value_stream-0.1.1/.vscode/settings.json +25 -0
- value_stream-0.1.1/.vscode/tasks.json +219 -0
- value_stream-0.1.1/AGENTS.md +45 -0
- value_stream-0.1.1/CONTRIBUTING.md +187 -0
- value_stream-0.1.1/Dockerfile +9 -0
- value_stream-0.1.1/Dockerfile.app +15 -0
- value_stream-0.1.1/LICENSE +21 -0
- value_stream-0.1.1/PKG-INFO +288 -0
- value_stream-0.1.1/README.md +262 -0
- value_stream-0.1.1/docs/architecture.md +386 -0
- value_stream-0.1.1/docs/studio-workspace.md +118 -0
- value_stream-0.1.1/pyproject.toml +67 -0
- value_stream-0.1.1/src/doc/img/readme/01-define-the-work.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/02-shape-the-delivery-system.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/interactive-mode.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/landing-bottom.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/landing.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/preview.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-mean-stage-loss-d0t1.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-mean-stage-loss-d10t25.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-mean-stage-loss.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-overview.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-resource-backlog.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-resource-utilization-d0t1.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-resource-utilization-d10t25.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-resource-utilization.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-value-lost-vs-deployment-interval.png +0 -0
- value_stream-0.1.1/src/doc/img/readme/results-value-lost-vs-team-size.png +0 -0
- value_stream-0.1.1/src/doc/img/workflow.svg +67 -0
- value_stream-0.1.1/src/doc/workflow.mmd +14 -0
- value_stream-0.1.1/src/examples/demo.py +106 -0
- value_stream-0.1.1/src/value_stream/__init__.py +1 -0
- value_stream-0.1.1/src/value_stream/app/SPEC.md +79 -0
- value_stream-0.1.1/src/value_stream/app/__init__.py +1 -0
- value_stream-0.1.1/src/value_stream/app/__main__.py +28 -0
- value_stream-0.1.1/src/value_stream/app/contracts.py +18 -0
- value_stream-0.1.1/src/value_stream/app/coordinator.py +359 -0
- value_stream-0.1.1/src/value_stream/app/errors.py +17 -0
- value_stream-0.1.1/src/value_stream/app/exports.py +142 -0
- value_stream-0.1.1/src/value_stream/app/frontend/e2e/studio.spec.ts +218 -0
- value_stream-0.1.1/src/value_stream/app/frontend/index.html +16 -0
- value_stream-0.1.1/src/value_stream/app/frontend/package-lock.json +5385 -0
- value_stream-0.1.1/src/value_stream/app/frontend/package.json +38 -0
- value_stream-0.1.1/src/value_stream/app/frontend/playwright.config.ts +26 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/App.tsx +1180 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/api.generated.ts +2675 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/charts.tsx +340 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/editors.test.tsx +111 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/editors.tsx +492 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/main.tsx +4 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/plotly.d.ts +4 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/style.css +956 -0
- value_stream-0.1.1/src/value_stream/app/frontend/src/types.ts +47 -0
- value_stream-0.1.1/src/value_stream/app/frontend/tsconfig.json +25 -0
- value_stream-0.1.1/src/value_stream/app/frontend/vite.config.ts +21 -0
- value_stream-0.1.1/src/value_stream/app/gateway.py +130 -0
- value_stream-0.1.1/src/value_stream/app/generation.py +269 -0
- value_stream-0.1.1/src/value_stream/app/insights.py +75 -0
- value_stream-0.1.1/src/value_stream/app/metrics.py +132 -0
- value_stream-0.1.1/src/value_stream/app/openapi.json +3699 -0
- value_stream-0.1.1/src/value_stream/app/schemas.py +288 -0
- value_stream-0.1.1/src/value_stream/app/server.py +298 -0
- value_stream-0.1.1/src/value_stream/app/settings.py +46 -0
- value_stream-0.1.1/src/value_stream/app/storage.py +651 -0
- value_stream-0.1.1/src/value_stream/client/__init__.py +3 -0
- value_stream-0.1.1/src/value_stream/client/simulation_runner.py +83 -0
- value_stream-0.1.1/src/value_stream/client/views/__init__.py +4 -0
- value_stream-0.1.1/src/value_stream/client/views/metadata_viewer.py +235 -0
- value_stream-0.1.1/src/value_stream/client/views/result_viewer.py +76 -0
- value_stream-0.1.1/src/value_stream/client/views/viewer.py +8 -0
- value_stream-0.1.1/src/value_stream/client/web/__init__.py +10 -0
- value_stream-0.1.1/src/value_stream/client/web/client.py +115 -0
- value_stream-0.1.1/src/value_stream/client/web/errors.py +29 -0
- value_stream-0.1.1/src/value_stream/client/web/local_service.py +92 -0
- value_stream-0.1.1/src/value_stream/factory/__init__.py +5 -0
- value_stream-0.1.1/src/value_stream/factory/developer_factory.py +33 -0
- value_stream-0.1.1/src/value_stream/factory/factory.py +43 -0
- value_stream-0.1.1/src/value_stream/factory/task_factory.py +64 -0
- value_stream-0.1.1/src/value_stream/factory/task_generator.py +102 -0
- value_stream-0.1.1/src/value_stream/policy/__init__.py +3 -0
- value_stream-0.1.1/src/value_stream/policy/simulation_policy.py +20 -0
- value_stream-0.1.1/src/value_stream/resources/__init__.py +23 -0
- value_stream-0.1.1/src/value_stream/resources/developer.py +53 -0
- value_stream-0.1.1/src/value_stream/resources/qa_tester.py +51 -0
- value_stream-0.1.1/src/value_stream/resources/resource.py +162 -0
- value_stream-0.1.1/src/value_stream/resources/resource_metadata.py +24 -0
- value_stream-0.1.1/src/value_stream/resources/resource_policy.py +13 -0
- value_stream-0.1.1/src/value_stream/resources/resource_pool.py +39 -0
- value_stream-0.1.1/src/value_stream/resources/resource_tracker.py +160 -0
- value_stream-0.1.1/src/value_stream/resources/toolchain.py +43 -0
- value_stream-0.1.1/src/value_stream/service/SPEC.md +46 -0
- value_stream-0.1.1/src/value_stream/service/__init__.py +1 -0
- value_stream-0.1.1/src/value_stream/service/app.py +218 -0
- value_stream-0.1.1/src/value_stream/service/codec.py +237 -0
- value_stream-0.1.1/src/value_stream/service/job_store.py +322 -0
- value_stream-0.1.1/src/value_stream/service/openapi.json +1283 -0
- value_stream-0.1.1/src/value_stream/service/scheduler.py +234 -0
- value_stream-0.1.1/src/value_stream/service/schemas.py +248 -0
- value_stream-0.1.1/src/value_stream/service/settings.py +37 -0
- value_stream-0.1.1/src/value_stream/service/worker.py +82 -0
- value_stream-0.1.1/src/value_stream/simulation/__init__.py +16 -0
- value_stream-0.1.1/src/value_stream/simulation/default_simulation_policy.py +68 -0
- value_stream-0.1.1/src/value_stream/simulation/model.py +30 -0
- value_stream-0.1.1/src/value_stream/simulation/model_factory.py +54 -0
- value_stream-0.1.1/src/value_stream/simulation/simulation.py +144 -0
- value_stream-0.1.1/src/value_stream/simulation/simulation_metadata.py +12 -0
- value_stream-0.1.1/src/value_stream/simulation/simulation_result.py +20 -0
- value_stream-0.1.1/src/value_stream/task/__init__.py +23 -0
- value_stream-0.1.1/src/value_stream/task/epoch.py +21 -0
- value_stream-0.1.1/src/value_stream/task/event_status.py +8 -0
- value_stream-0.1.1/src/value_stream/task/task.py +260 -0
- value_stream-0.1.1/src/value_stream/task/task_event.py +97 -0
- value_stream-0.1.1/src/value_stream/task/task_history.py +208 -0
- value_stream-0.1.1/src/value_stream/task/task_router.py +70 -0
- value_stream-0.1.1/src/value_stream/task/task_state.py +7 -0
- value_stream-0.1.1/src/value_stream/task/task_type.py +8 -0
- value_stream-0.1.1/src/value_stream/workflow/__init__.py +18 -0
- value_stream-0.1.1/src/value_stream/workflow/assignment_strategy.py +10 -0
- value_stream-0.1.1/src/value_stream/workflow/pool_manager.py +136 -0
- value_stream-0.1.1/src/value_stream/workflow/resource_operator.py +204 -0
- value_stream-0.1.1/src/value_stream/workflow/sdlc_workflow.py +99 -0
- value_stream-0.1.1/src/value_stream/workflow/support_workflow.py +23 -0
- value_stream-0.1.1/src/value_stream/workflow/task_store.py +125 -0
- value_stream-0.1.1/src/value_stream/workflow/workflow_policy.py +12 -0
- value_stream-0.1.1/tests/fixtures/service_job.json +34 -0
- value_stream-0.1.1/tests/integration/test_app_roundtrip.py +128 -0
- value_stream-0.1.1/tests/integration/test_local_runner.py +66 -0
- value_stream-0.1.1/tests/integration/test_service_roundtrip.py +226 -0
- value_stream-0.1.1/tests/unit/__init__.py +2 -0
- value_stream-0.1.1/tests/unit/app/test_boundaries.py +247 -0
- value_stream-0.1.1/tests/unit/app/test_generation_metrics.py +196 -0
- value_stream-0.1.1/tests/unit/app/test_lifecycle.py +303 -0
- value_stream-0.1.1/tests/unit/client/test_simulation_runner.py +73 -0
- value_stream-0.1.1/tests/unit/client/views/test_metadata_viewer.py +72 -0
- value_stream-0.1.1/tests/unit/client/views/test_result_viewer.py +62 -0
- value_stream-0.1.1/tests/unit/client/web/test_local_service.py +32 -0
- value_stream-0.1.1/tests/unit/client/web/test_web_client.py +227 -0
- value_stream-0.1.1/tests/unit/factory/test_developer_factory.py +33 -0
- value_stream-0.1.1/tests/unit/factory/test_task_factory.py +68 -0
- value_stream-0.1.1/tests/unit/factory/test_task_generator.py +124 -0
- value_stream-0.1.1/tests/unit/resources/__init__.py +0 -0
- value_stream-0.1.1/tests/unit/resources/test_developer.py +278 -0
- value_stream-0.1.1/tests/unit/resources/test_qa_tester.py +101 -0
- value_stream-0.1.1/tests/unit/resources/test_qa_tester_pool.py +43 -0
- value_stream-0.1.1/tests/unit/resources/test_resource.py +84 -0
- value_stream-0.1.1/tests/unit/resources/test_resource_tracker.py +136 -0
- value_stream-0.1.1/tests/unit/resources/test_toolchain.py +73 -0
- value_stream-0.1.1/tests/unit/resources/test_toolchain_pool.py +58 -0
- value_stream-0.1.1/tests/unit/service/test_app_extensions.py +82 -0
- value_stream-0.1.1/tests/unit/service/test_contract_alignment.py +145 -0
- value_stream-0.1.1/tests/unit/service/test_job_store.py +41 -0
- value_stream-0.1.1/tests/unit/simulation/test_model.py +30 -0
- value_stream-0.1.1/tests/unit/simulation/test_model_factory.py +55 -0
- value_stream-0.1.1/tests/unit/simulation/test_simulation.py +218 -0
- value_stream-0.1.1/tests/unit/simulation/test_simulation_policy.py +54 -0
- value_stream-0.1.1/tests/unit/task/__init__.py +0 -0
- value_stream-0.1.1/tests/unit/task/test_epoch.py +64 -0
- value_stream-0.1.1/tests/unit/task/test_support_task.py +51 -0
- value_stream-0.1.1/tests/unit/task/test_task.py +260 -0
- value_stream-0.1.1/tests/unit/task/test_task_history.py +269 -0
- value_stream-0.1.1/tests/unit/task/test_task_router.py +93 -0
- value_stream-0.1.1/tests/unit/testutils.py +40 -0
- value_stream-0.1.1/tests/unit/workflow/__init__.py +0 -0
- value_stream-0.1.1/tests/unit/workflow/test_developer_manager.py +46 -0
- value_stream-0.1.1/tests/unit/workflow/test_pool_manager.py +380 -0
- value_stream-0.1.1/tests/unit/workflow/test_qa_manager.py +89 -0
- value_stream-0.1.1/tests/unit/workflow/test_resource_operator.py +121 -0
- value_stream-0.1.1/tests/unit/workflow/test_sdlc_workflow.py +66 -0
- value_stream-0.1.1/tests/unit/workflow/test_support_workflow.py +39 -0
- value_stream-0.1.1/tests/unit/workflow/test_task_store.py +235 -0
- value_stream-0.1.1/tests/unit/workflow/test_toolchain_manager.py +174 -0
- value_stream-0.1.1/uv.lock +2246 -0
- value_stream-0.1.1/value-stream.code-workspace +50 -0
|
@@ -0,0 +1,150 @@
|
|
|
1
|
+
# Codex Execution Plans (ExecPlans):
|
|
2
|
+
|
|
3
|
+
This document describes the requirements for an execution plan ("ExecPlan"), a design document that a coding agent can follow to deliver a working feature or system change. Treat the reader as a complete beginner to this repository: they have only the current working tree and the single ExecPlan file you provide. There is no memory of prior plans and no external context.
|
|
4
|
+
|
|
5
|
+
## How to use ExecPlans and PLANS.md
|
|
6
|
+
|
|
7
|
+
When authoring an executable specification (ExecPlan), follow PLANS.md _to the letter_. If it is not in your context, refresh your memory by reading the entire PLANS.md file. Be thorough in reading (and re-reading) source material to produce an accurate specification. When creating a spec, start from the skeleton and flesh it out as you do your research.
|
|
8
|
+
|
|
9
|
+
When implementing an executable specification (ExecPlan), do not prompt the user for "next steps"; simply proceed to the next milestone. Keep all sections up to date, add or split entries in the list at every stopping point to affirmatively state the progress made and next steps. Resolve ambiguities autonomously, and commit frequently.
|
|
10
|
+
|
|
11
|
+
When discussing an executable specification (ExecPlan), record decisions in a log in the spec for posterity; it should be unambiguously clear why any change to the specification was made. ExecPlans are living documents, and it should always be possible to restart from _only_ the ExecPlan and no other work.
|
|
12
|
+
|
|
13
|
+
When researching a design with challenging requirements or significant unknowns, use milestones to implement proof of concepts, "toy implementations", etc., that allow validating whether the user's proposal is feasible. Read the source code of libraries by finding or acquiring them, research deeply, and include prototypes to guide a fuller implementation.
|
|
14
|
+
|
|
15
|
+
## Requirements
|
|
16
|
+
|
|
17
|
+
NON-NEGOTIABLE REQUIREMENTS:
|
|
18
|
+
|
|
19
|
+
* Every ExecPlan must be fully self-contained. Self-contained means that in its current form it contains all knowledge and instructions needed for a novice to succeed.
|
|
20
|
+
* Every ExecPlan is a living document. Contributors are required to revise it as progress is made, as discoveries occur, and as design decisions are finalized. Each revision must remain fully self-contained.
|
|
21
|
+
* Every ExecPlan must enable a complete novice to implement the feature end-to-end without prior knowledge of this repo.
|
|
22
|
+
* Every ExecPlan must produce a demonstrably working behavior, not merely code changes to "meet a definition".
|
|
23
|
+
* Every ExecPlan must define every term of art in plain language or do not use it.
|
|
24
|
+
|
|
25
|
+
Purpose and intent come first. Begin by explaining, in a few sentences, why the work matters from a user's perspective: what someone can do after this change that they could not do before, and how to see it working. Then guide the reader through the exact steps to achieve that outcome, including what to edit, what to run, and what they should observe.
|
|
26
|
+
|
|
27
|
+
The agent executing your plan can list files, read files, search, run the project, and run tests. It does not know any prior context and cannot infer what you meant from earlier milestones. Repeat any assumption you rely on. Do not point to external blogs or docs; if knowledge is required, embed it in the plan itself in your own words. If an ExecPlan builds upon a prior ExecPlan and that file is checked in, incorporate it by reference. If it is not, you must include all relevant context from that plan.
|
|
28
|
+
|
|
29
|
+
## Formatting
|
|
30
|
+
|
|
31
|
+
Format and envelope are simple and strict. Each ExecPlan must be one single fenced code block labeled as `md` that begins and ends with triple backticks. Do not nest additional triple-backtick code fences inside; when you need to show commands, transcripts, diffs, or code, present them as indented blocks within that single fence. Use indentation for clarity rather than code fences inside an ExecPlan to avoid prematurely closing the ExecPlan's code fence. Use two newlines after every heading, use # and ## and so on, and correct syntax for ordered and unordered lists.
|
|
32
|
+
|
|
33
|
+
When writing an ExecPlan to a Markdown (.md) file where the content of the file *is only* the single ExecPlan, you should omit the triple backticks.
|
|
34
|
+
|
|
35
|
+
Write in plain prose. Prefer sentences over lists. Avoid checklists, tables, and long enumerations unless brevity would obscure meaning. Checklists are permitted only in the `Progress` section, where they are mandatory. Narrative sections must remain prose-first.
|
|
36
|
+
|
|
37
|
+
## Guidelines
|
|
38
|
+
|
|
39
|
+
Self-containment and plain language are paramount. If you introduce a phrase that is not ordinary English ("daemon", "middleware", "RPC gateway", "filter graph"), define it immediately and remind the reader how it manifests in this repository (for example, by naming the files or commands where it appears). Do not say "as defined previously" or "according to the architecture doc." Include the needed explanation here, even if you repeat yourself.
|
|
40
|
+
|
|
41
|
+
Avoid common failure modes. Do not rely on undefined jargon. Do not describe "the letter of a feature" so narrowly that the resulting code compiles but does nothing meaningful. Do not outsource key decisions to the reader. When ambiguity exists, resolve it in the plan itself and explain why you chose that path. Err on the side of over-explaining user-visible effects and under-specifying incidental implementation details.
|
|
42
|
+
|
|
43
|
+
Anchor the plan with observable outcomes. State what the user can do after implementation, the commands to run, and the outputs they should see. Acceptance should be phrased as behavior a human can verify ("after starting the server, navigating to [http://localhost:8080/health](http://localhost:8080/health) returns HTTP 200 with body OK") rather than internal attributes ("added a HealthCheck struct"). If a change is internal, explain how its impact can still be demonstrated (for example, by running tests that fail before and pass after, and by showing a scenario that uses the new behavior).
|
|
44
|
+
|
|
45
|
+
Specify repository context explicitly. Name files with full repository-relative paths, name functions and modules precisely, and describe where new files should be created. If touching multiple areas, include a short orientation paragraph that explains how those parts fit together so a novice can navigate confidently. When running commands, show the working directory and exact command line. When outcomes depend on environment, state the assumptions and provide alternatives when reasonable.
|
|
46
|
+
|
|
47
|
+
Be idempotent and safe. Write the steps so they can be run multiple times without causing damage or drift. If a step can fail halfway, include how to retry or adapt. If a migration or destructive operation is necessary, spell out backups or safe fallbacks. Prefer additive, testable changes that can be validated as you go.
|
|
48
|
+
|
|
49
|
+
Validation is not optional. Include instructions to run tests, to start the system if applicable, and to observe it doing something useful. Describe comprehensive testing for any new features or capabilities. Include expected outputs and error messages so a novice can tell success from failure. Where possible, show how to prove that the change is effective beyond compilation (for example, through a small end-to-end scenario, a CLI invocation, or an HTTP request/response transcript). State the exact test commands appropriate to the project’s toolchain and how to interpret their results.
|
|
50
|
+
|
|
51
|
+
Capture evidence. When your steps produce terminal output, short diffs, or logs, include them inside the single fenced block as indented examples. Keep them concise and focused on what proves success. If you need to include a patch, prefer file-scoped diffs or small excerpts that a reader can recreate by following your instructions rather than pasting large blobs.
|
|
52
|
+
|
|
53
|
+
## Milestones
|
|
54
|
+
|
|
55
|
+
Milestones are narrative, not bureaucracy. If you break the work into milestones, introduce each with a brief paragraph that describes the scope, what will exist at the end of the milestone that did not exist before, the commands to run, and the acceptance you expect to observe. Keep it readable as a story: goal, work, result, proof. Progress and milestones are distinct: milestones tell the story, progress tracks granular work. Both must exist. Never abbreviate a milestone merely for the sake of brevity, do not leave out details that could be crucial to a future implementation.
|
|
56
|
+
|
|
57
|
+
Each milestone must be independently verifiable and incrementally implement the overall goal of the execution plan.
|
|
58
|
+
|
|
59
|
+
## Living plans and design decisions
|
|
60
|
+
|
|
61
|
+
* ExecPlans are living documents. As you make key design decisions, update the plan to record both the decision and the thinking behind it. Record all decisions in the `Decision Log` section.
|
|
62
|
+
* ExecPlans must contain and maintain a `Progress` section, a `Surprises & Discoveries` section, a `Decision Log`, and an `Outcomes & Retrospective` section. These are not optional.
|
|
63
|
+
* When you discover optimizer behavior, performance tradeoffs, unexpected bugs, or inverse/unapply semantics that shaped your approach, capture those observations in the `Surprises & Discoveries` section with short evidence snippets (test output is ideal).
|
|
64
|
+
* If you change course mid-implementation, document why in the `Decision Log` and reflect the implications in `Progress`. Plans are guides for the next contributor as much as checklists for you.
|
|
65
|
+
* At completion of a major task or the full plan, write an `Outcomes & Retrospective` entry summarizing what was achieved, what remains, and lessons learned.
|
|
66
|
+
|
|
67
|
+
# Prototyping milestones and parallel implementations
|
|
68
|
+
|
|
69
|
+
It is acceptable—-and often encouraged—-to include explicit prototyping milestones when they de-risk a larger change. Examples: adding a low-level operator to a dependency to validate feasibility, or exploring two composition orders while measuring optimizer effects. Keep prototypes additive and testable. Clearly label the scope as “prototyping”; describe how to run and observe results; and state the criteria for promoting or discarding the prototype.
|
|
70
|
+
|
|
71
|
+
Prefer additive code changes followed by subtractions that keep tests passing. Parallel implementations (e.g., keeping an adapter alongside an older path during migration) are fine when they reduce risk or enable tests to continue passing during a large migration. Describe how to validate both paths and how to retire one safely with tests. When working with multiple new libraries or feature areas, consider creating spikes that evaluate the feasibility of these features _independently_ of one another, proving that the external library performs as expected and implements the features we need in isolation.
|
|
72
|
+
|
|
73
|
+
## Skeleton of a Good ExecPlan
|
|
74
|
+
|
|
75
|
+
# <Short, action-oriented description>
|
|
76
|
+
|
|
77
|
+
This ExecPlan is a living document. The sections `Progress`, `Surprises & Discoveries`, `Decision Log`, and `Outcomes & Retrospective` must be kept up to date as work proceeds.
|
|
78
|
+
|
|
79
|
+
If PLANS.md file is checked into the repo, reference the path to that file here from the repository root and note that this document must be maintained in accordance with PLANS.md.
|
|
80
|
+
|
|
81
|
+
## Purpose / Big Picture
|
|
82
|
+
|
|
83
|
+
Explain in a few sentences what someone gains after this change and how they can see it working. State the user-visible behavior you will enable.
|
|
84
|
+
|
|
85
|
+
## Progress
|
|
86
|
+
|
|
87
|
+
Use a list with checkboxes to summarize granular steps. Every stopping point must be documented here, even if it requires splitting a partially completed task into two (“done” vs. “remaining”). This section must always reflect the actual current state of the work.
|
|
88
|
+
|
|
89
|
+
- [x] (2025-10-01 13:00Z) Example completed step.
|
|
90
|
+
- [ ] Example incomplete step.
|
|
91
|
+
- [ ] Example partially completed step (completed: X; remaining: Y).
|
|
92
|
+
|
|
93
|
+
Use timestamps to measure rates of progress.
|
|
94
|
+
|
|
95
|
+
## Surprises & Discoveries
|
|
96
|
+
|
|
97
|
+
Document unexpected behaviors, bugs, optimizations, or insights discovered during implementation. Provide concise evidence.
|
|
98
|
+
|
|
99
|
+
- Observation: …
|
|
100
|
+
Evidence: …
|
|
101
|
+
|
|
102
|
+
## Decision Log
|
|
103
|
+
|
|
104
|
+
Record every decision made while working on the plan in the format:
|
|
105
|
+
|
|
106
|
+
- Decision: …
|
|
107
|
+
Rationale: …
|
|
108
|
+
Date/Author: …
|
|
109
|
+
|
|
110
|
+
## Outcomes & Retrospective
|
|
111
|
+
|
|
112
|
+
Summarize outcomes, gaps, and lessons learned at major milestones or at completion. Compare the result against the original purpose.
|
|
113
|
+
|
|
114
|
+
## Context and Orientation
|
|
115
|
+
|
|
116
|
+
Describe the current state relevant to this task as if the reader knows nothing. Name the key files and modules by full path. Define any non-obvious term you will use. Do not refer to prior plans.
|
|
117
|
+
|
|
118
|
+
## Plan of Work
|
|
119
|
+
|
|
120
|
+
Describe, in prose, the sequence of edits and additions. For each edit, name the file and location (function, module) and what to insert or change. Keep it concrete and minimal.
|
|
121
|
+
|
|
122
|
+
## Concrete Steps
|
|
123
|
+
|
|
124
|
+
State the exact commands to run and where to run them (working directory). When a command generates output, show a short expected transcript so the reader can compare. This section must be updated as work proceeds.
|
|
125
|
+
|
|
126
|
+
## Validation and Acceptance
|
|
127
|
+
|
|
128
|
+
Describe how to start or exercise the system and what to observe. Phrase acceptance as behavior, with specific inputs and outputs. If tests are involved, say "run <project’s test command> and expect <N> passed; the new test <name> fails before the change and passes after>".
|
|
129
|
+
|
|
130
|
+
## Idempotence and Recovery
|
|
131
|
+
|
|
132
|
+
If steps can be repeated safely, say so. If a step is risky, provide a safe retry or rollback path. Keep the environment clean after completion.
|
|
133
|
+
|
|
134
|
+
## Artifacts and Notes
|
|
135
|
+
|
|
136
|
+
Include the most important transcripts, diffs, or snippets as indented examples. Keep them concise and focused on what proves success.
|
|
137
|
+
|
|
138
|
+
## Interfaces and Dependencies
|
|
139
|
+
|
|
140
|
+
Be prescriptive. Name the libraries, modules, and services to use and why. Specify the types, traits/interfaces, and function signatures that must exist at the end of the milestone. Prefer stable names and paths such as `crate::module::function` or `package.submodule.Interface`. E.g.:
|
|
141
|
+
|
|
142
|
+
In crates/foo/planner.rs, define:
|
|
143
|
+
|
|
144
|
+
pub trait Planner {
|
|
145
|
+
fn plan(&self, observed: &Observed) -> Vec<Action>;
|
|
146
|
+
}
|
|
147
|
+
|
|
148
|
+
If you follow the guidance above, a single, stateless agent -- or a human novice -- can read your ExecPlan from top to bottom and produce a working, observable result. That is the bar: SELF-CONTAINED, SELF-SUFFICIENT, NOVICE-GUIDING, OUTCOME-FOCUSED.
|
|
149
|
+
|
|
150
|
+
When you revise a plan, you must ensure your changes are comprehensively reflected across all sections, including the living document sections, and you must write a note at the bottom of the plan describing the change and the reason why. ExecPlans must describe not just the what but the why for almost everything.
|
|
@@ -0,0 +1,353 @@
|
|
|
1
|
+
# Build the interactive value-stream simulation application
|
|
2
|
+
|
|
3
|
+
|
|
4
|
+
This ExecPlan is a living document maintained according to `.agent/PLANS.md`. Keep `Progress`, `Surprises & Discoveries`, `Decision Log`, and `Outcomes & Retrospective` current throughout implementation. The agreed product behavior is in `src/value_stream/app/SPEC.md`. This document contains implementation decisions, proposed defaults, and executable validation milestones.
|
|
5
|
+
|
|
6
|
+
Implementation was authorized by the user on 2026-09-25 after review of this plan. Progress and evidence below distinguish completed work from planned validation.
|
|
7
|
+
|
|
8
|
+
## Purpose / Big Picture
|
|
9
|
+
|
|
10
|
+
|
|
11
|
+
An individual demonstrating value-stream simulation will start one local command, open a browser, generate a reusable task set, and compare named scenarios without writing Python. A scenario is one concrete simulation model applied to a particular task set. A parameter sweep generates scenarios from every combination of selected property values. Results appear as individual scenarios finish, and an interactive comparison lets the user change a property, retain a baseline, and pin interesting alternatives. Clear loss, stage, resource activity, and backlog plots support explainable suggestions and separate CSV exports.
|
|
12
|
+
|
|
13
|
+
Version 1 is a single-instance application with temporary server memory. Browser refreshes retain saved inputs and comparisons. A Docker image provides the same behavior as local command-line execution. Storage and execution boundaries must permit future shared storage and distributed workers, but this version has no horizontal scaling, authentication, persistent database, automatic optimization, or machine-learning dependency.
|
|
14
|
+
|
|
15
|
+
The observable completion example is: run `python -m value_stream.app`, open `http://127.0.0.1:8081`, create 100 tasks, sweep team sizes 2, 4, and 6 at cadence 5, run three scenarios, and see three points in Value Lost vs. Team Size. Enable interactive mode and change cadence to 1. The original series remains and a second series appears as three further results arrive. Pin it, refresh the browser, and find both comparisons still available. Download CSVs, then stop the server; temporary workspace state is gone after restart.
|
|
16
|
+
|
|
17
|
+
## Progress
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
- [x] (2026-09-25 16:55Z) Read the app spec, repository guidelines, plan format, current service, simulation resource telemetry, viewers, packaging, and CI configuration.
|
|
21
|
+
- [x] (2026-09-25 16:55Z) Record the user's decisions on inputs, distributions, sweeps, incremental results, interactive comparisons, exports, deployment, and deferring horizontal scaling.
|
|
22
|
+
- [x] (2026-09-25 16:55Z) Resolve resource metric scope with the user: preserve existing metrics with accurate labels in V1.
|
|
23
|
+
- [x] (2026-09-25 16:55Z) Update the application specification and draft this implementation plan (planning checkpoint).
|
|
24
|
+
- [x] (2026-09-25) Check required plan sections, seven milestones, source references, and whitespace; review consistency of seeds, comparison grouping, lifecycle, and metric scope. Runtime validation was deferred until implementation authorization at that checkpoint.
|
|
25
|
+
- [x] (2026-09-25 19:30Z) Milestone 1: establish generated app/service contracts and retain real-browser plot, zoom-preservation, incremental-result, and metric proofs.
|
|
26
|
+
- [x] (2026-09-25 19:30Z) Milestone 2: reproducible task/team generation, stable scenarios, explicit execution seeds, atomic service submission deduplication, and deterministic resource enumeration. Committed the service changes separately as `75d36b6`.
|
|
27
|
+
- [x] (2026-09-25 19:30Z) Milestone 3: bounded in-memory storage, HTTP coordination, partial failure/cancellation, retry/resume, supersession, and protected retention, with focused lifecycle/boundary tests.
|
|
28
|
+
- [x] (2026-09-25 19:30Z) Milestone 4: task/model editors, multi-property preview, incremental progress, accessible result tables, and five Plotly views.
|
|
29
|
+
- [x] (2026-09-25 19:30Z) Milestone 5: immutable baseline/pins, latest interactive comparison, committed slider/numeric edits, cached reuse, and refresh recovery.
|
|
30
|
+
- [x] (2026-09-25 19:30Z) Milestone 6: evidence-based suggestions, observed baseline differences with changed settings, and separate provenance-rich CSV exports.
|
|
31
|
+
- [x] (2026-09-25 19:40Z) Milestone 7: final wheel and Docker asset hashes match; both packaged runtimes completed real simulations without Node, 100-model/admission/cancellation checks passed, CI/docs/evidence are present, and temporary services shut down cleanly.
|
|
32
|
+
|
|
33
|
+
## Surprises & Discoveries
|
|
34
|
+
|
|
35
|
+
|
|
36
|
+
The HTTP service already exposes batch submission, polling, indexed completed outcomes, partial failures, and whole-job cancellation. `src/value_stream/service/app.py` defines these routes and `schemas.py` supplies explicit JSON schemas. Workers return only complete model results in `service/worker.py`; there is no need to introduce within-model streaming, WebSockets, or server-pushed events.
|
|
37
|
+
|
|
38
|
+
The existing job seed is combined with the model's batch index in `service/worker.py:derive_seed`. Submitting a changed subset therefore changes randomness unless the contract supports seeds independent of positions. Add explicit per-model seeds without changing legacy job-seed behavior.
|
|
39
|
+
|
|
40
|
+
`src/value_stream/workflow/pool_manager.py` keeps registered resources in a set and constructs the random support-assignment candidate list from that set. Resource hashes depend on UUIDs in `resources/resource.py`. Consequently a fixed random seed alone may not reproduce support assignments across processes. Verify this with a targeted support/rework scenario, then preserve resource registration order while retaining membership checks. UUIDs themselves remain nonreproducible identity metadata.
|
|
41
|
+
|
|
42
|
+
`client/views/metadata_viewer.py:resource_capacity` cumulatively sums resource `waiting` deltas. `workflow/resource_operator.py:_execute` increments that counter once per resource request, which can represent a batch of tasks. It also excludes time spent waiting for a deployment cadence trigger before the resource request. Call this Resource Backlog, with an axis of Waiting resource requests and a visible explanation of batching and scope.
|
|
43
|
+
|
|
44
|
+
The existing resource-utilization viewer normalizes recorded idle, successful-work, failed-work, and interruption durations. These durations can overlap under interruptions, and they do not cover all unused capacity to the simulation horizon. Preserve that calculation for comparison with existing viewers, but label its measure Recorded activity share. It is not a measured fraction of total available capacity. The user explicitly chose this V1 scope rather than expanded telemetry.
|
|
45
|
+
|
|
46
|
+
Task event loss is a signed fraction of that task's initial value; summary loss is a signed fraction of total initial value. Events lack a task identifier and initial-value weight in the wire result. A mean of stage visits therefore cannot be advertised as an additive, value-weighted decomposition of the overall loss. No new task-level attribution model is required for V1.
|
|
47
|
+
|
|
48
|
+
The simulation service already expires completed jobs after one hour by default. Application comparisons must copy completed results into the application's bounded workspace store as they arrive; otherwise refreshing a saved comparison after service expiry would fail during the same server lifetime.
|
|
49
|
+
|
|
50
|
+
The repository has a Python package and a service Dockerfile, but no frontend build. `pyproject.toml` uses Hatchling and `uv.lock` is used by CI. The existing local service helper owns a loopback Uvicorn server with reference counting. Reuse it through an adapter instead of creating a second simulation implementation.
|
|
51
|
+
|
|
52
|
+
Plotly uses trace identifiers in CSS selectors. The first implementation put a serialized grouping key into `uid`, causing a selector error during cleanup and subsequent redraw failures. Stable scenario UUIDs fix the selector issue. The retained production browser test switches all five views rapidly, resets zoom, resizes the page, and verifies no plot errors. A small serialized Plotly rendering adapter also replaces queued updates with the newest view and copies input data because Plotly mutates it. The extra React wrapper dependency was removed.
|
|
53
|
+
|
|
54
|
+
Native number-input minimum/step alignment initially rejected valid default efficiencies. Aligned the UI bounds/steps, while retaining positive-value validation on the server. A finite-value edge test also found that averaging very large efficiency bounds could overflow; midpoint/interpolation arithmetic now avoids that overflow.
|
|
55
|
+
|
|
56
|
+
The host has Node 23, outside the locked Vitest version's supported versions. All final local frontend validation uses a task-local Node 22 runtime; Docker and CI use Node 22. The locked frontend audit reports zero vulnerabilities. Chromium requires macOS sandbox escalation; ordinary Python and HTTP checks run in the workspace sandbox, except process-tree memory sampling.
|
|
57
|
+
|
|
58
|
+
## Decision Log
|
|
59
|
+
|
|
60
|
+
|
|
61
|
+
- Decision: task size means task count; task effort is story points, with no separate complexity field. Initial value and story points support constants or bounded uniform ranges; depreciation rate is an exposed scalar. Rationale: match the simulation model and avoid redundant controls. Date/Author: 2026-09-25, user.
|
|
62
|
+
- Decision: team generation accepts size, efficiency minimum/maximum, and linear or normal distribution. Linear is evenly spaced including endpoints, with the midpoint for one developer. Rationale: satisfy repeatable team comparisons with understandable controls. Date/Author: 2026-09-25, user.
|
|
63
|
+
- Decision: support individual definitions and multi-property sweeps, making sweeps the main creation workflow. Rationale: larger model sets are the common demonstration use case. Date/Author: 2026-09-25, user.
|
|
64
|
+
- Decision: update progress and plots when each model finishes; cancellation stops the entire job. Rationale: the existing service already supplies the needed granularity. Date/Author: 2026-09-25, user.
|
|
65
|
+
- Decision: interactive mode retains an immutable baseline, the latest alternative, and explicitly pinned alternatives. Commit sliders on release; replace superseded interactive work. Rationale: compare one changed property without retaining every intermediate slider position. Date/Author: 2026-09-25, user.
|
|
66
|
+
- Decision: use positive value-lost percentages, labeled accessible plot tabs, stable scenario names/colors, separate CSV exports, and rule-based suggestions. Rationale: make demonstration results understandable without implying proven causal improvements. Date/Author: 2026-09-25, user.
|
|
67
|
+
- Decision: keep resource metrics compatible with existing viewers, with accurate qualifiers. Rationale: precise capacity utilization and task backlog would require a separate telemetry change. Date/Author: 2026-09-25, user response during planning.
|
|
68
|
+
- Decision: use TypeScript, React, Vite, and Plotly.js in the browser, with Python FastAPI/Uvicorn on the server. Rationale: rich interactive controls and charts fit a browser application while keeping model generation and execution integration in the existing language. Date/Author: 2026-09-25, Codex recommendation accepted in the design discussion.
|
|
69
|
+
- Decision: run one application server process, reuse or connect to one simulation service, and store workspaces in memory behind an interface. Rationale: the user deferred horizontal scaling while retaining extensibility and local/Docker startup. Date/Author: 2026-09-25, user.
|
|
70
|
+
- Decision: expose asynchronous service operations through an application-owned HTTP gateway, not `SimulationRunner.execute`. Rationale: the runner waits for the entire batch, whereas the application needs polling, cancellation, and persistent run identity. Date/Author: 2026-09-25, Codex design choice.
|
|
71
|
+
- Decision: normal efficiencies use a truncated normal with midpoint mean and standard deviation `(maximum - minimum) / 6`, drawing again if outside the bounds. Equal bounds yield constant efficiencies. A one-person normal team is sampled normally. Rationale: positive bounded efficiencies and no additional distribution parameter in V1. Date/Author: 2026-09-25, Codex design choice.
|
|
72
|
+
- Decision: add optional explicit model seeds and an optional idempotent submission identifier to the simulation service. An idempotent submission returns the original job when the same request is retried. Rationale: preserve subset reproducibility and avoid duplicate jobs after lost submission responses. Date/Author: 2026-09-25, Codex design choice.
|
|
73
|
+
- Decision: preserve deterministic registration order for support resource selection if the focused test confirms the set-order problem. Rationale: user-requested numerical reproducibility includes models with support work, not only simple models. Date/Author: 2026-09-25, Codex design choice.
|
|
74
|
+
- Decision: use one sampled run per scenario in V1 and display that fact. Rationale: seed preservation enables repeatability, but it does not provide uncertainty estimates or statistical significance; replicate sampling and confidence intervals are outside the accepted scope. Date/Author: 2026-09-25, Codex design choice.
|
|
75
|
+
|
|
76
|
+
- Decision: save task settings and model definitions together through an atomic versioned editor endpoint. Rationale: preview must describe one consistent input revision; separate partial writes would create avoidable stale combinations. Date/Author: 2026-09-25, implementation refinement.
|
|
77
|
+
- Decision: use the concrete schema names `ModelSettings`, `RunRecord`, `RunStatus`, and `OutcomeSummary`, with comparison flags on immutable run snapshots; use React hooks and generated types with guarded obsolete reads. Rationale: no duplicate domain objects or client-side authoritative state machine is needed in V1. Date/Author: 2026-09-25, implementation refinement.
|
|
78
|
+
- Decision: call Plotly directly through a small serialized adapter and use CSS-safe trace IDs. Rationale: predictable coalescing, resize, cleanup, and zoom behavior; no extra wrapper package. Date/Author: 2026-09-25, browser validation.
|
|
79
|
+
- Decision: keep plot-tab preference in browser session storage and rebuild visible comparisons from server baseline/latest/pins. Rationale: server state retains scientific inputs/results, while view-only choices stay local. Date/Author: 2026-09-25, implementation refinement.
|
|
80
|
+
|
|
81
|
+
## Outcomes & Retrospective
|
|
82
|
+
|
|
83
|
+
|
|
84
|
+
The application is implemented end to end. Users can generate and reuse seeded task sets, compare multi-property model sweeps, inspect five incrementally updated plots and accessible tables, retain baseline/pinned alternatives, test suggestions, cancel whole jobs, and export three CSVs. The app uses the existing HTTP simulation service, with two backward-compatible request extensions for stable seeds and duplicate-safe submission. Storage, execution gateway, and coordination remain separable for later distributed implementations.
|
|
85
|
+
|
|
86
|
+
Final validation: 177 Python tests passed, frontend unit tests passed three cases, and production Chromium passed three end-to-end cases. A subsequent focused 22-test app run verified the final finite-limit/extreme-range guards. App Python coverage is 88%; generation 92%, metrics 98%, gateway 98%, insights 100%, storage 89%, coordinator 81%. CLI/contracts are additionally exercised by build/runtime checks. The existing Python 3.11 environment was tested locally; the repository's wider Python matrix is retained in CI, not claimed as locally executed.
|
|
87
|
+
|
|
88
|
+
A fresh wheel environment with Node absent from PATH served the UI and ran 100 three-task scenarios in 5.821 seconds. The largest sampled status request took 0.0184 seconds; peak sampled app/service/worker RSS was 315,424,768 bytes (including previously retained validation runs). Docker completed the same-sized workload in 7.559 seconds with a maximum sampled status-request duration of 0.0745 seconds. Both rejected 101 scenarios before execution. Docker whole-job cancellation retained five successes and cancelled 95 remaining scenarios. These are small local validation workloads, not throughput guarantees. Docker memory after the batch and cancellation checks was 533.7 MiB, a post-test observation rather than a measured peak. Serialized-byte limits do not equal resident memory limits.
|
|
89
|
+
|
|
90
|
+
Screenshots and machine-readable runtime evidence are under `.agent/artifacts/simulation-app/`. Browser validation uses Chromium at 1440px desktop, 390px mobile, and a 720-CSS-pixel viewport representing reflow at 200% zoom on a 1440px display. No Safari/Firefox or native browser text-zoom run is claimed. The wheel includes compiled assets and excludes frontend sources; Docker uses a Node build stage and a Python runtime, removing the build source tree. Existing broad Python package dependencies are preserved, including its declared test/notebook dependencies; changing the package's dependency model is outside this feature.
|
|
91
|
+
|
|
92
|
+
V1 intentionally remains single-process, server-memory-only, with existing qualified resource telemetry and one sampled execution per scenario. Horizontal scaling, persistence, expanded telemetry, and ML optimization remain deferred as agreed.
|
|
93
|
+
|
|
94
|
+
## Context and Orientation
|
|
95
|
+
|
|
96
|
+
|
|
97
|
+
Core simulation code is in `src/value_stream/simulation/`, including `Simulation.execute`, `Model`, and summary/metadata result types. Inputs include `task/task.py`, `resources/developer.py`, `resources/qa_tester.py`, and `resources/toolchain.py`. `factory/factory.py` currently has a shared NumPy random generator; do not reuse that mutable generator for application reproducibility. `service/schemas.py` and `service/codec.py` define and translate the language-neutral API. `service/job_store.py`, `scheduler.py`, and `worker.py` own temporary jobs and isolated model processes. The application must submit through HTTP rather than call `Simulation.execute` directly.
|
|
98
|
+
|
|
99
|
+
Existing simulation routes are `POST /v1/simulation-jobs`, `GET /v1/simulation-jobs/{job_id}`, `GET /v1/simulation-jobs/{job_id}/outcomes?after={cursor}`, and `DELETE /v1/simulation-jobs/{job_id}`. Submission returns 202 and a job ID. Outcomes carry model indexes in completion order. A cursor is a monotonically increasing position allowing later polls to fetch only newer outcomes. Terminal states are completed, completed_with_errors, and cancelled. Repeated cancellation preserves completed results.
|
|
100
|
+
|
|
101
|
+
Create Python application files under `src/value_stream/app/`, frontend sources and npm configuration under `src/value_stream/app/frontend/`, and compiled browser assets under `src/value_stream/app/static/`. New Python tests belong in `tests/unit/app/` and `tests/integration/test_app_*.py`. Frontend unit tests and browser tests belong under `frontend/src/` and `frontend/e2e/`, respectively. The service-only root `Dockerfile` remains supported; add `Dockerfile.app` for the combined application.
|
|
102
|
+
|
|
103
|
+
A workspace is a server-stored collection of task sets, editable scenario definitions, immutable run snapshots, comparisons, and results. A task set is a saved specification plus its materialized task array, meaning the actual sampled values submitted to the service. A scenario definition may expand into many concrete scenarios. A run is an application-owned request and stable mapping from concrete scenario identities to service model indexes. A comparison is the immutable group of scenario revisions plotted together. A revision is a monotonically increasing version of an editable object. A cache key describes exact inputs and seeds, allowing an already completed result to be reused safely.
|
|
104
|
+
|
|
105
|
+
## Interfaces and Dependencies
|
|
106
|
+
|
|
107
|
+
|
|
108
|
+
Use React and TypeScript with Vite, a small rendering adapter around `plotly.js`, and CSS maintained inside the frontend. Begin with semantic HTML controls; do not introduce a design-system dependency merely for cosmetic consistency. Use Vitest and React Testing Library for frontend behavior, and Playwright for browser integration. Commit `package-lock.json` and pin an interoperable dependency set during Milestone 1; Node 22.12 or newer in the Node 22 line is the initial build environment. Node is a build/development dependency only. Python remains compatible with the repository's declared Python 3.11+ support and CI matrix. Use the already declared `httpx` package's `AsyncClient` for the new gateway; preserve the existing client's use of `httpx2` without unrelated dependency migration.
|
|
109
|
+
|
|
110
|
+
Add `app/schemas.py` with strict Pydantic models forbidding extra fields. Define `ValueSpec` as a discriminated constant or uniform-range object; `TaskSetSpec` with count, story points, initial value, depreciation rate, and seed; `ModelSettings` with team size, efficiency bounds, distribution, and model properties; `ScenarioDefinition` with a name, base model settings, team specification, and sweep values; `ConcreteScenario` with identity, revision, name, concrete `ModelInput`, and execution seed; and `Workspace`, `RunStatus`, `OutcomeSummary`, `PlotData`, and `Observation`. `RunRecord` in storage holds the immutable run inputs; comparison flags and identifiers refer to those runs. Reuse service input schemas for concrete model validation and service error/result schemas in the gateway, rather than copying their constraints.
|
|
111
|
+
|
|
112
|
+
Use UUIDs represented as JSON strings for workspace, task-set, scenario, run, and comparison identities. Store a distinct integer revision and exact numerical configuration with each saved object. Editing never mutates a baseline snapshot. Names are user-editable with generated defaults and a maximum of 120 characters. Frontend links retain the workspace ID in the URL and can remember the last workspace ID locally; local browser storage is not the authoritative workspace database or an authentication mechanism.
|
|
113
|
+
|
|
114
|
+
In `app/storage.py`, define a `WorkspaceStore` protocol with operations to create/read/delete workspaces, save versioned definitions, atomically register runs, append deduplicated outcomes, retain/release comparisons, and enforce byte/count reservations. Provide a lock-protected `InMemoryWorkspaceStore`. Use expected revision on mutations and return a conflict for stale edits, including edits from a second browser tab. Store serialized/value objects, not worker process handles or HTTP connections.
|
|
115
|
+
|
|
116
|
+
In `app/gateway.py`, define an asynchronous `SimulationGateway` protocol with `submit(request)`, `status(job_id)`, `outcomes(job_id, after)`, `cancel(job_id)`, and `close()`. Its implementation talks to the configured HTTP service using explicit timeouts. In `app/coordinator.py`, define `RunCoordinator.start_run(workspace_id, request)`, `cancel_run(...)`, and `close()`. The coordinator owns background polling tasks and cancellation/recovery transitions; the store owns the bounded result cache; routes only validate and delegate. An in-memory coordinator implements this protocol for V1; a future distributed dispatcher can replace it without changing browser contracts.
|
|
117
|
+
|
|
118
|
+
In `app/generation.py`, define pure `materialize_tasks(spec)`, `materialize_team(settings, seed)`, and `expand_scenarios(definitions, limits)` functions. In `app/metrics.py`, implement pure result-to-plot calculations and provenance. In `app/insights.py`, return structured rule identifiers, evidence values, explanation text, and proposed property edits. In `app/exports.py`, stream CSV from immutable stored snapshots. In `app/server.py`, expose `create_app(settings=None, store=None, gateway=None)` with dependency injection for tests. In `app/__main__.py`, implement the standalone command and lifecycle ownership.
|
|
119
|
+
|
|
120
|
+
Publish the app's OpenAPI 3.1 document as `app/openapi.json` and serve it at `/openapi.json`. Generate TypeScript wire types from that document using a locked `openapi-typescript` development dependency. Keep application state types separate from generated wire types. CI must fail if the generated contract or generated TypeScript types differ from the checked-in versions. Service OpenAPI remains separately checked at `service/openapi.json`.
|
|
121
|
+
|
|
122
|
+
## Inputs, Identity, and Reproducibility
|
|
123
|
+
|
|
124
|
+
|
|
125
|
+
Default task count is 100, story points are constant 1, initial value is constant 1, and depreciation rate is 0.005. Depreciation is displayed as 0.5% per simulation time unit and converted to a fraction at the API boundary. Constant/range story points and values must be finite and nonnegative; count is a positive integer. A range with equal bounds acts as a constant. All generated tasks are development tasks created at simulation time zero. Editing a task specification creates a new task-set revision rather than changing tasks used by earlier runs. Regenerate explicitly chooses a new generation seed and creates a new task set; it starts a separate comparison context rather than overlaying unlike task sets.
|
|
126
|
+
|
|
127
|
+
Use local `random.Random` instances with explicit seeds and versioned, domain-separated seed derivation using SHA-256 of canonical JSON. Domain separation means task generation, team generation, and simulation execution include different fixed tags, so changes in one generator do not consume another generator's random stream. Store generated arrays and a generator version alongside the seeds. Generated numeric results are reproducible for the same supported engine/generator version; UUIDs and timings of HTTP completion are not part of that guarantee.
|
|
128
|
+
|
|
129
|
+
Team efficiency minimum and maximum must be finite, positive, and ordered. Linear distribution produces the midpoint for count one, otherwise both endpoints and evenly spaced interior values. Normal distribution uses the formula in the Decision Log and samples without clipping out-of-bounds values onto endpoints. Equal bounds always produce the constant value. Name developers in deterministic input order. For a team-size sweep, generate a distinct team for each size; linear teams are resampled across their full bounds at each size, not formed by appending to the smaller team. Normal teams use a stable team-family seed and deterministic draws so increasing size extends the same draw sequence. In both cases, changing unrelated settings such as cadence or QA capacity reuses the materialized team at each size. Explain the linear resizing rule in the UI.
|
|
130
|
+
|
|
131
|
+
The default model has team size 4, efficiency bounds 0.5 to 1.5 with linear distribution, cadence 5, QA pool size 2, QA time cost 0.1, QA failure rate/cost 0, toolchain pool size 1, deployment duration 0.25, toolchain failure rate 0, support disabled, and support story points 1. Efficiency means story points processed per simulation time unit. QA time cost means time units per story point. Cadence zero means deployment whenever work is available. Support disabled maps to null, not zero; enabling support requires a strictly positive interval. Show rates as percentages and duration units explicitly.
|
|
132
|
+
|
|
133
|
+
Each numeric model property can be fixed or swept: team size, efficiency bounds, cadence, QA capacity/time cost/failure rate/rework fraction, toolchain capacity/duration/failure rate, support interval, and support story points. Distribution choice also accepts an explicit list if swept. Task-generation properties are not model sweep axes because every model within a job must receive identical tasks. Support either explicit values or inclusive start/end/step inputs, using decimal arithmetic for expansion and enforcing integer-only count/cadence values. Reject invalid combinations, such as an efficiency minimum exceeding a maximum, with field/combination details; do not silently omit them. Deduplicate repeated values and identical concrete model inputs within a definition, reporting the final count. Concatenate separately named definitions without merging their identities.
|
|
134
|
+
|
|
135
|
+
Compute cardinality before generating task/team arrays or a Cartesian product, meaning all combinations of sweep values. Reject expansion above the configured model limit; do not truncate or silently split into multiple jobs. Include a compact preview of names and differing settings. Store concrete scenario IDs so reordering definitions or sweep values preserves surviving scenario identities. Removed scenarios retain their historical snapshots. Display duplicate configuration scenarios distinctly, while sharing numerical results when the exact cache key matches.
|
|
136
|
+
|
|
137
|
+
Add `JobRequest.model_seeds: list[int] | None` to `service/schemas.py`. Its length must match models and values must be nonnegative. Reject requests supplying both job `seed` and `model_seeds`; omitting both remains valid for legacy unseeded callers. Without model_seeds, retain the current `derive_seed(job_seed, model_index)` behavior unchanged. The scheduler extracts the indexed explicit seed into a one-element model_seeds array in each one-model payload; the worker uses that explicit seed directly, while retaining the original model_index for result association. The app always submits model_seeds and stores their mapping by scenario identity, independent of service model index. Preserve an execution seed when an interactive scenario changes configuration; assign independent reproducible seeds to new scenario identities. This supports controlled repeatable comparisons, but does not promise statistically coupled random draws when workflow paths differ.
|
|
138
|
+
|
|
139
|
+
Cache numerical results by canonical materialized task content, concrete model content, explicit execution seed, generator version, and engine/contract version. Names, submission index, and batch order do not affect this key. Cache only successful complete results; never cache a failure as a successful simulation. Keep the original service result/index as provenance, and map it into a separate application scenario outcome so cached results do not masquerade as another service job's raw outcome.
|
|
140
|
+
|
|
141
|
+
## Application HTTP Contract and Lifecycle
|
|
142
|
+
|
|
143
|
+
|
|
144
|
+
All application routes begin `/api/v1`; static browser assets use `/assets`. `/health` reports process liveness and `/ready` reports whether the configured service is reachable and supports required optional request fields. Check service compatibility through its OpenAPI schema at startup. An older service lacking model_seeds or idempotent submission support produces `SERVICE_INCOMPATIBLE`; never silently fall back to nonreproducible submissions. Failed readiness should leave existing workspace reading and exports available.
|
|
145
|
+
|
|
146
|
+
Provide `GET /api/v1/config` for documented limits/defaults and service readiness. Workspace routes are `POST /api/v1/workspaces`, `GET /api/v1/workspaces/{id}`, and `DELETE /api/v1/workspaces/{id}`. An atomic `PUT /api/v1/workspaces/{id}/editor` saves task specifications and scenario definitions with an expected workspace revision and returns the preview. A nested task-set DELETE route removes unused retained sets; definition additions, edits, and deletions use the editor request. Add `POST /api/v1/workspaces/{id}/preview` returning validated expansion count, concrete scenario summaries, and a preview digest. `POST /api/v1/workspaces/{id}/runs` accepts a unique client request ID, task-set revision, definition revisions or committed preview digest, and run intent (manual or interactive). It returns 202 with an app run ID promptly; it must reject a stale preview rather than execute inputs the user has not previewed. Baseline selection, pin/unpin, and comparison deletion are explicit workspace mutations. Deleting a workspace with running work returns a conflict until the user cancels its jobs.
|
|
147
|
+
|
|
148
|
+
`GET /api/v1/workspaces/{id}/runs/{run_id}` returns state, per-scenario outcomes (from which progress counts are derived), service-job identity if known, scenario revisions, and last outcome cursor. `GET .../runs/{run_id}/outcomes?after={cursor}` returns small indexed outcome summaries; detail/plot data is fetched separately by scenario outcome ID. `DELETE .../runs/{run_id}` cancels the whole app run and its outstanding service job while retaining completed results. States include submitting, queued, running, reconnecting, cancelling, completed, completed_with_errors, cancelled, and failed. A superseded run is a cancelled historical run with a superseded reason; it is never the latest comparison. Outcome cursors advance on terminal model outcomes, including cached outcomes, and are deduplicated on repeated polls. A fully cached run finishes without a new service job.
|
|
149
|
+
|
|
150
|
+
Expose immutable result details and plot-ready data by comparison/outcome ID. Never send all raw metadata on every status poll or workspace load. Completed result details are loaded when needed for a plot, insight, or export. Poll lightweight workspace status about once per second to retain cross-tab consistency; raw results are never included in these polls. Server harvesting polls upstream about every 500 ms while running, backing off to five seconds on connection errors. Backend harvesting continues when the browser closes so completed service outcomes are retained before upstream expiry. Permit one large upstream outcome download at a time with a semaphore. Stream into a byte-limited buffer before decoding, capped at the configured 64 MiB job-outcome budget plus 1 MiB envelope allowance by default; do not use an unbounded body read. App detail endpoints return one model's reduced plot data and observations per response; full raw metadata is available through CSV, and browser code fetches details only for the selected scenarios.
|
|
151
|
+
|
|
152
|
+
Add optional `JobRequest.submission_id: UUID | None` to the service. In `job_store.py`, atomically associate this identifier with a canonical request fingerprint and the created job. Repeated identical submissions return the original job without scheduling again, even when admission is otherwise full. Reusing the identifier for different content returns 409 `SUBMISSION_CONFLICT`. Keep this mapping as long as the job exists and prune it with job expiry; its semantics do not survive a service restart. Extend the storage protocol with atomic submission reservation and test that two concurrent identical submissions start only one batch. The scheduler must receive work only for newly reserved jobs. Preserve existing callers that omit the identifier. Document the extension in `service/SPEC.md` and regenerate its OpenAPI when implementing it.
|
|
153
|
+
|
|
154
|
+
App run submissions are likewise idempotent per workspace/client request ID. Derive one fixed upstream submission_id from the app run and retry uncertain submissions only with that identifier during a bounded reconnect window. If a known upstream job becomes unknown/expired, mark the run failed with partial outcomes preserved; do not automatically recreate known missing jobs. A user-requested retry is a new run with clear provenance. Cancellation during uncertain submission must resolve the same idempotent submission and cancel it; if service communication cannot be restored, show cancellation unconfirmed rather than claiming the processes stopped. Temporary idempotency records do not provide exactly-once execution across an undetected service restart during an uncertain submission; document that boundary, and stop automatic retries when a restart is known. Log identifiers and transitions without dumping full result bodies.
|
|
155
|
+
|
|
156
|
+
The standalone launcher defaults to host 127.0.0.1 and port 8081. It creates one app process. If an explicit `--service-url` or `VALUE_STREAM_SERVICE_URL` is supplied, use it; otherwise acquire the existing managed loopback service once through `client/web/local_service.py`. Wrap acquisition/release behind the gateway's lifecycle adapter, performing blocking startup outside the app event loop. Release only a locally owned service at shutdown; never shut down an external service. Stop accepting runs, request cancellation of app-owned active jobs, stop harvest tasks, close the gateway, and release local service ownership in a bounded orderly shutdown. Report cancellation failures for an unreachable remote service. Bind Docker to 0.0.0.0 internally while publishing host loopback by default. Running more than one app server worker is explicitly unsupported with temporary memory.
|
|
157
|
+
|
|
158
|
+
## Interactive Comparison Behavior
|
|
159
|
+
|
|
160
|
+
|
|
161
|
+
The first manually completed run becomes the baseline comparison. A partially successful completed run can be a baseline, with failed scenarios clearly marked as gaps; an entirely failed run cannot enable interactive mode. Baseline and pinned comparison content never changes after creation. Users may explicitly choose a different retained comparison as baseline. Task-set revisions cannot be mixed inside a comparison context.
|
|
162
|
+
|
|
163
|
+
For Value Lost vs. Team Size, the team-size sweep supplies points along the x-axis; other differing properties define separate series. For Value Lost vs. Deployment Interval, cadence supplies the x-axis. Sort numeric axes ascending, with cadence zero labeled Continuous. Group lines by comparison identity, definition/team-family identity, and all non-axis generator/model settings. For a team-size sweep, group on the team's generation family and distribution settings rather than its changing materialized developer list; for a cadence sweep, retain the concrete team identity. Never collapse different QA/toolchain/support settings merely because team size and cadence match. When arbitrary individual scenarios do not form a comparable series, show labeled points without misleading connections. An accessible results table and filters support large sweep sets, with a default maximum of 12 visible series and an explicit selector for more.
|
|
164
|
+
|
|
165
|
+
When interactive mode is enabled, choose a property currently held fixed across the selected comparison family, such as cadence. The x-axis sweep remains unchanged. In V1 one interactive property is active at a time; other changes use the normal editor/run workflow. A property already varying across multiple series must first be narrowed to one selected value/family before it becomes an interactive slider. Discrete and integer fields use appropriate steps; support enabled/disabled uses a select or toggle, not an invalid numeric zero. Provide a numeric entry alongside sliders.
|
|
166
|
+
|
|
167
|
+
Commit pointer sliders on release, numeric fields on Enter/blur, and keyboard slider changes after a 400 ms idle delay. Enforce at least one second between automatic submissions. While superseded work is being cancelled, retain only the newest pending edit, then submit that edit after the prior service job is terminal; do not queue every change. Associate every response with the immutable run revision and ignore it for latest-plot selection if superseded. Completed results from superseded runs may remain cache entries but cannot replace baseline or latest labels.
|
|
168
|
+
|
|
169
|
+
Keep previous plots visible with a named Updating comparison indicator until new outcomes arrive. Each completed point adds to the new series; pending/failed/cancelled points remain gaps with status details. The latest unpinned comparison is replaced by the next committed alternative. Pin moves that immutable comparison into the retained set, including a partial-comparison label when applicable. Pinning is available after terminal completion to avoid mutable pinned content. Cancel stops the whole current job and pauses automatic submissions until another explicit user edit or Resume action; cancellation itself must not immediately trigger a replacement run.
|
|
170
|
+
|
|
171
|
+
Refresh restores workspace definitions, baseline/latest/pinned comparisons, and current run identities from the server. Restore visible comparisons from baseline/latest/pins and plot selection from browser session storage; pixel-level zoom may stay browser-local. Browser disconnection does not cancel simulation work. Baseline and pinned comparisons count toward retention limits and are never silently evicted.
|
|
172
|
+
|
|
173
|
+
## Metric Definitions, Observations, and Exports
|
|
174
|
+
|
|
175
|
+
|
|
176
|
+
Overall Value lost (%) is `100 * (sum(initial_value) - total_delivered_value) / sum(initial_value)` for the task set. It equals negative 100 times the existing summary loss for supported inputs. A zero denominator produces null/N/A, not zero loss. Preserve the service's raw signed loss unchanged in stored data. Do not take absolute value indiscriminately or silently clamp unexpected discrepancies. A hand-calculated fixture with initial values 1 and 3 and delivered total 2 must show 50% loss. The formula weights tasks by initial value naturally.
|
|
177
|
+
|
|
178
|
+
Mean Stage Loss uses end events for development tasks, including rework visits, grouped by workflow stage and scenario. Display the mean of `-100 * event.loss` and the number of contributing stage visits. Each visit is weighted equally, not by initial value or task count; missing visits produce N/A. Include failed visits, label that choice, and keep support work out of this loss calculation. Workflow stages must follow `SDLCWorkflow.WorkflowState` order with friendly labels such as Waiting for development, Development, Waiting for QA, QA testing, Waiting for deployment, and Deployment. Supporting data preserves raw state names. Explain that unequal task values and repeated visits make this chart non-additive; suggestions based on it describe the largest observed mean stage-visit loss, not the largest proven contributor to overall monetary loss.
|
|
179
|
+
|
|
180
|
+
Resource Utilization is the plot-tab name, with the visible subtitle Recorded activity share. Sum each stage's non-null idle_t, success_t, failure_t, and interruption_t, and divide each by their combined sum, matching the existing viewer. Display categories Recorded idle, Successful work, Failed work, and Interruption duration. A zero total yields N/A. Explain that durations can overlap and do not measure utilization of all configured capacity. Show the raw recorded duration totals alongside the proportions so normalized shares do not suggest more precision than exists.
|
|
181
|
+
|
|
182
|
+
Resource Backlog sums waiting deltas by workflow stage and time, then forms a cumulative count in time order. Aggregate equal timestamps before cumulative summation, begin with zero, and show step lines through completion time on a linear y-axis so zero remains visible. Label the y-axis Waiting resource requests and explain that requests may contain task batches and exclude pre-request cadence waiting. Do not rename these counts tasks. Detect unexplained negative cumulative counts as a data-quality error. Export full resolution. For browser display above 5,000 points per trace, implement and test an aggregation preserving each time bucket's first, last, minimum, and maximum in temporal order, and clearly label a reduced-resolution view.
|
|
183
|
+
|
|
184
|
+
Keep insight rules small and deterministic. Identify the stage with the highest observed mean stage-visit loss, the stage with the highest observed request backlog, and observed failed-work activity when present. Rules may suggest reducing cadence for deployment waiting, trying more QA capacity for QA backlog, or testing a lower failure rate when failed-work duration is material. Show evidence, applicable settings, and an explicit untested suggestion label. Do not promise a quantified improvement before running it. The Test suggestion action creates an editable comparison proposal preserving task/team/seed context; execution is a deliberate action. If another retained comparable scenario has lower observed total loss, report the measured difference in percentage points, listing all changed settings rather than claiming one setting caused it. Show no optimization or significance claim from a single run.
|
|
185
|
+
|
|
186
|
+
Expose separate CSV downloads for summaries, stage events, and resource history from a selected immutable comparison. Summary rows include every scenario including failure/cancellation, scenario ID/name/revision, task-set ID/content hash, run/comparison ID, execution and generation seeds, generator/engine version, concrete model settings, team efficiencies as a quoted JSON cell, status/error, completion time, initial/delivered value, and positive loss percentage. Event/resource rows repeat identifying provenance and carry all raw service fields, including original signed event loss; event rows add a clearly named positive stage_loss_percent. Use stable headers, UTF-8, proper CSV quoting, full numeric precision, blank unavailable fields, and deterministic scenario/event ordering. Do not invent per-task IDs that are absent from current telemetry. Quote and neutralize spreadsheet-formula prefixes in user-supplied text cells only, without altering negative numeric data. Downloads of partial comparisons carry status information in summary rows and include only available event/resource data. Stream from immutable references so downloading does not duplicate all stored results in memory.
|
|
187
|
+
|
|
188
|
+
## Limits and Error Contract
|
|
189
|
+
|
|
190
|
+
|
|
191
|
+
Retain current service defaults: 1,000 tasks per job, 100 models per job, 2 MiB submission body, four active jobs plus four queued jobs, at most min(4, CPU count) model workers, 120 seconds per model, 8 MiB per model outcome, 64 MiB per job outcomes, 64 retained jobs, 256 MiB retained outcome bytes, and one-hour terminal job lifetime. The app's validated defaults must not exceed the managed service settings. For remote services, app limits are preliminary and service rejections remain authoritative. Expose both relevant local limits and any discovered remote mismatch without implying local limits reconfigure a remote service.
|
|
192
|
+
|
|
193
|
+
Add configurable `VALUE_STREAM_APP_` settings for four workspaces, 20 task sets and 100 definitions per workspace, 100 models per run, 100 developers or pooled resources per concrete model, 32 values per sweep axis, 12 pinned comparisons per workspace, one active app run per workspace, four active app runs globally, 32 retained terminal runs per workspace, 2 MiB per mutation request, 16 MiB total stored input definitions/task arrays, and 256 MiB total retained raw application results. Reserve at most 64 MiB of result budget for each admitted run and account for existing cached references only once. When actual data exceeds its reservation or model result limit, preserve acquired results, fail the affected acquisition with `RESULT_TOO_LARGE`, and cancel outstanding work where necessary. These byte limits bound serialized payload storage, not exact Python resident memory; document that distinction and observe process memory during the large-batch validation.
|
|
194
|
+
|
|
195
|
+
Workspace data survives until server shutdown unless explicitly deleted or bounded unpinned history is evicted. Baseline, latest comparison, and pinned comparison references protect their results. Evict least recently used unreferenced cache entries and oldest unpinned terminal history first. Never evict active runs or protected inputs/results to admit new work. If protected data fills the budget, reject new work with a clear action to unpin/delete comparisons. New capacity limits are checked before allocating expanded models or starting service work. Numeric/configuration length limits prevent expansion attacks even when the resulting JSON body is small.
|
|
196
|
+
|
|
197
|
+
Use the service's error shape `{code, message, details}` for request errors, with field locations where applicable. App routes return 422 `INVALID_INPUT` or `LIMIT_EXCEEDED`; 413 `BODY_TOO_LARGE`; 404 `WORKSPACE_NOT_FOUND`, `RUN_NOT_FOUND`, or `RESULT_NOT_FOUND`; 409 `REVISION_CONFLICT`, `RUN_ACTIVE`, `SUBMISSION_CONFLICT`, or `COMPARISON_LIMIT`; 429 `APP_CAPACITY`; 503 `SERVICE_UNAVAILABLE` or `SERVICE_INCOMPATIBLE`; and 500 `INTERNAL_ERROR` without a traceback. Preserve model-scoped service errors such as MODEL_FAILED, MODEL_TIMEOUT, WORKER_FAILED, and RESULT_TOO_LARGE with scenario identity added. Service JOB_NOT_FOUND after restart/expiry becomes a run-level SERVICE_JOB_LOST condition with successful app results preserved.
|
|
198
|
+
|
|
199
|
+
Connection errors during polling move a run to reconnecting without declaring its models failed. Retry safe reads with exponential backoff capped at five seconds, for up to 60 seconds by default. Retain an explicit Resume connection action after the automatic retry budget ends. Submission retries use the same idempotent ID only; unresolved submission and cancellation display unconfirmed states and relevant details. Browser request cancellation is distinct from simulation cancellation. Render inline field errors, job banners, model-level failures, and accessible retry controls; a failed model must not blank successful plots or exports.
|
|
200
|
+
|
|
201
|
+
## Plan of Work
|
|
202
|
+
|
|
203
|
+
|
|
204
|
+
### Milestone 1: establish contracts and verify frontend/metric integration
|
|
205
|
+
|
|
206
|
+
|
|
207
|
+
Create the application module skeleton, schema document, frontend package/lockfile, and dependency-injected FastAPI application factory. Add a small retained frontend fixture that renders one summary series and updates it without resetting zoom. Use Plotly's stable `uirevision` for the same axes and stable trace IDs; reset only when the task set or axis meaning changes or Reset zoom is requested. Verify keyboard tabs, numeric inputs, plot resizing, and a text/table alternative. This is a future implementation proof, not work performed during planning.
|
|
208
|
+
|
|
209
|
+
Add `tests/unit/app/test_generation_metrics.py` with hand-calculated loss, stage visits, zero denominators, backlog batches, equal timestamps, missing categories, and overlapping activity-duration examples. Implement only the necessary pure metric calculations for the fixture. Confirm the displayed labels and computations agree with the definitions above. The proof is accepted when the browser shows two progressively added scenario points, preserves a selected zoom, and displays accurate activity/backlog qualifiers. Retain this fixture as a regression test rather than leaving a throwaway prototype in the product.
|
|
210
|
+
|
|
211
|
+
Define app routes and Pydantic schemas enough to generate `app/openapi.json`, and generate frontend types. Add npm scripts for dev, build, typecheck, test, test:e2e, and generate:api. Run the focused metric tests, frontend unit tests, and typecheck/build. Record pinned versions, proof screenshots, and any compatibility findings in this plan. Do not adopt a framework change without updating the Decision Log and affected packaging instructions.
|
|
212
|
+
|
|
213
|
+
### Milestone 2: reproducible generation and service extensions
|
|
214
|
+
|
|
215
|
+
|
|
216
|
+
Implement generation/expansion functions and stable revision/identity mapping. Add meaningful tests for task distribution bounds, count-one/equal-bound teams, fixed seeds, deterministic normal sampling, linear resizing, cartesian cardinality, integer steps, invalid combinations, model-limit rejection before allocation, and changes to cadence preserving tasks/teams. Avoid tests that simply repeat implementation lines; assert materialized behavioral examples and invariants.
|
|
217
|
+
|
|
218
|
+
Modify `service/schemas.py`, `scheduler.py`, and `worker.py` for optional model_seeds. Modify `service/job_store.py` and `app.py` for optional atomic submission_id behavior. Update `service/SPEC.md`, OpenAPI, and existing contract/roundtrip tests, preserving old request forms and runner behavior. Add support-work reproducibility tests; if confirmed, replace nondeterministic registered-resource enumeration with insertion order in `workflow/pool_manager.py` without changing assignment strategies. Run existing workflow and service tests as these are simulation/service behavior changes required by repository policy.
|
|
219
|
+
|
|
220
|
+
Acceptance: materialize the same configuration twice and obtain the same tasks/efficiencies; execute a stochastic scenario within a batch, in a reordered batch, and alone with its explicit seed and obtain the same numerical summary/event/resource measurements, excluding UUIDs. Include support and rework in this check. Legacy job-seed behavior stays unchanged except for the documented resource-order determinism correction. Two identical concurrent submission IDs produce one job and one execution; a conflicting body returns 409.
|
|
221
|
+
|
|
222
|
+
### Milestone 3: workspace and job lifecycle
|
|
223
|
+
|
|
224
|
+
|
|
225
|
+
Implement store, gateway, coordinator, settings, and routes. Reuse managed local service ownership through an adapter while keeping the gateway independently testable against a real remote HTTP endpoint. Store immutable submissions and index-to-scenario mappings before execution; harvest and reserve results without browser participation. Implement idempotent app runs, conflict checks, byte/count limits, safe retries, and cancellation propagation. Add `/health`, `/ready`, and `/api/v1/config`.
|
|
226
|
+
|
|
227
|
+
Use a controlled fake gateway in unit tests to force out-of-order success/failure, delayed cancellation, duplicate outcomes, loss of the submission response, reconnecting status, stale revisions, and result overflow. Add `tests/integration/test_app_roundtrip.py` using a real loopback simulation service for successful batch execution, partial failure, cancellation, and refresh-equivalent workspace retrieval. Shorten limits in tests instead of sleeping for real TTL/timeouts. Verify that server-side result copying keeps a completed comparison available after upstream expiry, and that app shutdown cancels its owned work without stopping an external service.
|
|
228
|
+
|
|
229
|
+
Acceptance: ordinary HTTP requests create a workspace, materialize a task set, preview scenarios, submit a run, fetch incremental outcomes, cancel it, and export/read prior results using stable identities. Invalid input returns the documented envelope; there is no duplicate execution after a retried submission and no hidden unlimited background queue.
|
|
230
|
+
|
|
231
|
+
### Milestone 4: editors, progress, and plotting
|
|
232
|
+
|
|
233
|
+
|
|
234
|
+
Implement editor components in `frontend/src/editors.tsx`, chart rendering in `charts.tsx`, and workspace/progress/comparison/result controls in `App.tsx`. Use React hooks for transient editor/view state, server snapshots for authoritative run state, and fetch helpers with generated types. Guard obsolete asynchronous detail reads and abort obsolete workspace initialization reads. Keep backend calculations authoritative, with lightweight client validation for immediate feedback. Provide labels/help for units, support disabled, deployment continuous, distribution bounds, and limits. Default to a clean desktop layout with an input panel and results area; keep forms usable on narrow screens without requiring chart interactions.
|
|
235
|
+
|
|
236
|
+
Render all five plot views using plot-ready data and friendly workflow labels. Keep a consistent scenario/comparison legend across views. Show progress counts for succeeded, failed, cancelled, running, and queued models; do not imply a percentage of CPU work completed. Provide a visible whole-job Cancel button, empty/loading/error states, a selectable scenario table, and accessible data alternatives. A partially filled sweep must show gaps rather than fabricate data or connect across failed points.
|
|
237
|
+
|
|
238
|
+
Acceptance: browser automation creates a multi-property sweep, verifies preview count, launches it, observes plot/result rows before the final scenario finishes, changes tabs, zooms/resets, cancels, and sees retained successes with failure details. Test two scenarios sharing cadence/team size but differing QA settings and verify distinct series/identities. Use controlled gateway timings for deterministic incremental UI checks, plus a real-service browser smoke run.
|
|
239
|
+
|
|
240
|
+
### Milestone 5: baseline and interactive comparisons
|
|
241
|
+
|
|
242
|
+
|
|
243
|
+
Implement `ComparisonControls` and `InteractiveControls`, the server comparison mutations, and result reuse. Enforce the baseline/latest/pinned lifecycle and one automatic in-flight run per workspace. Preserve configuration revision and task/team/seed snapshots so superseded responses cannot win a race. Limit submissions using commit semantics, delay, and newest-pending-edit replacement as specified above.
|
|
244
|
+
|
|
245
|
+
Acceptance: a three-team-size baseline at cadence 5 remains visible while cadence 1 creates a second series; returning to an already completed configuration reuses its result without a worker launch. Pinning, another change, and refresh preserve the baseline and pin. Rapid slider edits launch only committed values and retain only the latest pending edit; late results never replace the newest selected comparison. Cancel pauses further automatic execution. Verify plot series group correctly when starting from a multi-series sweep and selecting one family for interactive control.
|
|
246
|
+
|
|
247
|
+
### Milestone 6: observations and exports
|
|
248
|
+
|
|
249
|
+
|
|
250
|
+
Implement the rule-based observations and explicit Test suggestion workflow in `insights.py` and the browser. Each rule must have evidence and a meaningful behavioral test with a crafted result. Do not make a suggestion when relevant data is unavailable, all value is zero, or the suggested input would violate limits. Suggestions are proposals, never an automatic search or silent job submission.
|
|
251
|
+
|
|
252
|
+
Implement CSV streaming endpoints and frontend downloads for the selected comparison. Test quoting, Unicode, user-text formula prefixes, raw signed numeric fields, deterministic ordering, all status rows, and partial comparisons. Read generated CSVs with Python's csv reader and compare their values to stored results and plotted metrics.
|
|
253
|
+
|
|
254
|
+
Acceptance: a controlled backlog result produces a clearly qualified capacity suggestion, a user can prepare/run its comparison, and observed differences are reported only after completion. Each CSV opens as a valid table and contains the scenario IDs, seeds, settings, status, and measurements required to interpret it independently.
|
|
255
|
+
|
|
256
|
+
### Milestone 7: startup, packaging, documentation, and full validation
|
|
257
|
+
|
|
258
|
+
|
|
259
|
+
Complete `app/__main__.py` with `--host`, `--port`, and `--service-url`, and graceful signal handling. Build the frontend into `app/static` and include those assets in the Python wheel through Hatchling configuration. Ensure API 404s and errors are not swallowed by the single-page-app fallback route. Serve hashed assets with long-lived cache headers and the HTML shell with revalidation. If assets are missing in a source checkout, fail with the exact frontend build command; do not download/build dependencies automatically on startup.
|
|
260
|
+
|
|
261
|
+
Add `Dockerfile.app` with a Node build stage and a Python runtime stage. Build/install the Python package after copying compiled assets into the package tree. Do not include node_modules, development servers, tests, or compilers in the runtime image. Preserve the root service-only Dockerfile, update `.dockerignore` for frontend artifacts carefully, and update `.gitignore` so build output is not committed. Add explicit wheel inclusions/exclusions so static assets are present but frontend development artifacts are absent. Update README with source-build prerequisites, one-command runtime startup, service URL precedence, Docker examples, limits, metric meanings, CSV descriptions, reproducibility boundaries, and the single-server limitation.
|
|
262
|
+
|
|
263
|
+
Extend CI with a locked npm install, generated-contract/type checks, TypeScript checking, frontend tests/build, Playwright Chromium tests, and a built-wheel asset smoke test. Retain Python tests across the existing supported-version matrix. Add one browser test against the built production UI and a real service, not only Vite or mocked API responses. Use Playwright's managed web-server fixture so tests start/stop only their own server and never connect to an unrelated developer session.
|
|
264
|
+
|
|
265
|
+
Acceptance: install the built wheel in a fresh Python environment without Node, run the app, and perform the three-scenario demonstration. Build/run Docker and repeat health, readiness, UI asset loading, one real simulation, CSV download, cancellation, and clean shutdown. Inspect screenshots at a wide desktop viewport, a narrow viewport, and 200% text zoom; verify labels, errors, legends, focus, and controls are usable. Record actual screenshots/commands/results and any performance limitations in the final plan update.
|
|
266
|
+
|
|
267
|
+
## Concrete Steps
|
|
268
|
+
|
|
269
|
+
|
|
270
|
+
Run the following implementation/validation commands from repository root unless stated otherwise. Actual observed results appear in Outcomes and the evidence files. Install Python dependencies according to the existing uv lock or documented virtual-environment workflow. After adding frontend package files, install initially with `npm install --prefix src/value_stream/app/frontend`, review and commit the lockfile, then use `npm ci --prefix src/value_stream/app/frontend` for repeatable subsequent installs. Update uv.lock if Python dependency/build configuration changes require it. Do not claim a dependency version until the locked build has been tested.
|
|
271
|
+
|
|
272
|
+
Focused Python validation after the relevant milestones:
|
|
273
|
+
|
|
274
|
+
.venv/bin/pytest tests/unit/app/test_generation_metrics.py -q
|
|
275
|
+
.venv/bin/pytest tests/unit/app tests/unit/service tests/unit/workflow -q
|
|
276
|
+
.venv/bin/pytest tests/integration/test_app_roundtrip.py -q
|
|
277
|
+
|
|
278
|
+
Create npm scripts matching these commands before using them:
|
|
279
|
+
|
|
280
|
+
npm ci --prefix src/value_stream/app/frontend
|
|
281
|
+
npm run generate:api --prefix src/value_stream/app/frontend
|
|
282
|
+
npm run typecheck --prefix src/value_stream/app/frontend
|
|
283
|
+
npm test --prefix src/value_stream/app/frontend -- --run
|
|
284
|
+
npm run build --prefix src/value_stream/app/frontend
|
|
285
|
+
|
|
286
|
+
From `src/value_stream/app/frontend`, install test browsers and run integration tests:
|
|
287
|
+
|
|
288
|
+
npx playwright install chromium
|
|
289
|
+
npm run test:e2e
|
|
290
|
+
|
|
291
|
+
Use `npx playwright install --with-deps chromium` in the Linux CI job where system packages are needed. Browser installation and loopback tests may require environment/network permissions; report any actual restriction instead of claiming they passed. Frontend `test:e2e` must configure its working directory and Python executable explicitly so it starts the repository's environment reliably.
|
|
292
|
+
|
|
293
|
+
Full repository validation after implementation:
|
|
294
|
+
|
|
295
|
+
.venv/bin/pytest --cov=src --cov-report=xml tests -s
|
|
296
|
+
.venv/bin/python -m pip wheel --no-deps . -w /private/tmp/value-stream-app-wheel
|
|
297
|
+
|
|
298
|
+
Run the locally built application, then inspect it in a browser and query it from another terminal:
|
|
299
|
+
|
|
300
|
+
.venv/bin/python -m value_stream.app --host 127.0.0.1 --port 8081
|
|
301
|
+
curl -sS http://127.0.0.1:8081/health
|
|
302
|
+
curl -sS http://127.0.0.1:8081/ready
|
|
303
|
+
curl -sS http://127.0.0.1:8081/api/v1/config
|
|
304
|
+
|
|
305
|
+
Expected observations are HTTP 200 for healthy/ready local startup, visible configured limits, a rendered UI at `/`, and incremental successful outcome rows from a submitted three-model run. These observations were verified against the source app, installed wheel, and Docker runtime. Keep setup-generated workspace/run IDs in test variables rather than hardcoding random IDs.
|
|
306
|
+
|
|
307
|
+
For Docker validation:
|
|
308
|
+
|
|
309
|
+
docker build -f Dockerfile.app -t value-stream-app .
|
|
310
|
+
docker run --rm --name value-stream-app-check -p 127.0.0.1:8081:8081 value-stream-app
|
|
311
|
+
|
|
312
|
+
Run the browser/API checks against this container, then stop only this named test container from another terminal:
|
|
313
|
+
|
|
314
|
+
docker stop value-stream-app-check
|
|
315
|
+
|
|
316
|
+
For a separately managed simulation service, start the existing documented command on port 8080 and launch the app with `--service-url http://127.0.0.1:8080`. The app still uses one browser origin on 8081. Stop the app and verify the externally managed service remains healthy. For a wheel-only check, create a fresh temporary environment, install the newly built wheel using the actual produced filename, start its Python executable with `-m value_stream.app`, and confirm static assets and simulations work with Node absent from that runtime's path. Record the exact artifact filename when it exists.
|
|
317
|
+
|
|
318
|
+
## Validation and Acceptance
|
|
319
|
+
|
|
320
|
+
|
|
321
|
+
Python tests use unittest-style Test classes under pytest as in the repository. Add tests when simulation resource ordering, wire contracts, or service workflows change. Frontend tests exercise user-visible behavior rather than implementation details. No exact final test count is predicted; record observed passing counts, coverage, and browser results after execution. A failed new reproducibility or ordering test must be demonstrated against the old behavior when practical before applying the fix.
|
|
322
|
+
|
|
323
|
+
The feature is complete only when all seven milestones pass their observable acceptance criteria, the existing simulation/client/service tests still pass, both OpenAPI documents and generated types agree, production assets load from a wheel and Docker image, and the browser demonstrates incremental results, whole-job cancellation, baseline/pinned comparisons, refresh retention, error recovery, and valid exports. Include deterministic failures and cancellation tests as well as successful paths. Polling, retry, and supersession races must be forced in tests instead of relying on fast simulations to happen to overlap.
|
|
324
|
+
|
|
325
|
+
Test a model set at the default 100-scenario limit with short simulations, then a rejected 101-scenario expansion. Record responsiveness, peak observed app/service memory, outcome byte sizes, and total duration; distinguish expected workload runtime from UI response latency. Check that large result requests are bounded, status stays lightweight, and cache/retention accounting releases unreferenced results. This is a validation of configured protections, not a throughput claim or an invitation to broaden scope into benchmarking infrastructure.
|
|
326
|
+
|
|
327
|
+
UI verification includes keyboard-only task entry, tab switching, sliders and numeric alternatives, cancellation, pinning, accessible error summaries, status announcements that do not flood screen readers, plot data tables, and meaningful labels at 200% text zoom. Use shapes/dashes as well as colors for key comparisons. Include Chromium automation and a manual supported-browser smoke check where available; document what was actually tested. Mobile usability focuses on readable forms/tables and resizable charts, not a separate mobile product.
|
|
328
|
+
|
|
329
|
+
## Idempotence and Recovery
|
|
330
|
+
|
|
331
|
+
|
|
332
|
+
Keep changes additive where possible and preserve the standalone Python simulation runner and service image. Updating the app spec must not rewrite the completed service ExecPlan; service contract changes belong in its spec, tests, and a clearly scoped note referencing this feature. Do not overwrite unrelated user edits, particularly existing staged files. No database migration or persistent user-data conversion is required.
|
|
333
|
+
|
|
334
|
+
Repeated UI run submissions with the same request ID and payload return the same run. Repeated upstream submission IDs return the same job for its retention lifetime. Repeated polls deduplicate by cursor, and repeated cancellation preserves successes. Rebuilding frontend assets, reinstalling the wheel, and rerunning tests must be safe; clean only task-owned build/temp artifacts. Browser refresh recovers from server state without re-execution. A server restart explicitly loses workspaces; a stale saved browser link offers a new workspace and explains that the temporary server state ended.
|
|
335
|
+
|
|
336
|
+
If an implementation proof exposes a metric or reproducibility defect, record evidence and choose the smallest correction consistent with this plan. Do not expand telemetry, distributed infrastructure, or statistical analysis silently. If a new product decision materially changes the user's expectations, raise it while continuing independent work. Routine implementation decisions can proceed with an entry in the Decision Log. The in-memory interfaces are extension points; implementing them does not establish horizontal scalability.
|
|
337
|
+
|
|
338
|
+
## Artifacts and Notes
|
|
339
|
+
|
|
340
|
+
|
|
341
|
+
Planning evidence comes from the repository paths identified above and official documentation checked on 2026-09-25. React documents the TypeScript/Vite setup at https://react.dev/learn/build-a-react-app-from-scratch. Vite's guide at https://vite.dev/guide/ documents its static build and Node requirements. Plotly's https://plotly.com/javascript/uirevision/ documents preserving user interaction across plot updates. Playwright's https://playwright.dev/docs/test-webserver documents test-owned server startup. The necessary implementation behavior is specified in this plan; these references are supplemental rather than prerequisites for interpreting it.
|
|
342
|
+
|
|
343
|
+
Store implementation screenshots and concise validation transcripts under `.agent/artifacts/simulation-app/` when they are produced, excluding large/generated transient files from commits as appropriate. Record actual test counts, dependency versions, supported browser evidence, wheel contents, and Docker behavior here. The runtime JSON and screenshots now exist in that directory; large transient browser traces remain ignored.
|
|
344
|
+
|
|
345
|
+
There are no unanswered blocking product questions. Reviewable defaults selected by this plan include the normal-distribution spread, one active interactive property, linear team resizing, default admission/retention limits, and two small backward-compatible service request extensions. They are design choices, not claims that the user specified their exact implementation. The resource telemetry question raised during planning was resolved in favor of existing metrics with clear labels.
|
|
346
|
+
|
|
347
|
+
Revision note, 2026-09-25: created the initial application ExecPlan and updated the app spec to capture the design discussion. Added explicit reproducibility, idempotent submission, metric interpretation, retention, and packaging decisions after inspecting the current service and resource implementations. The consistency review clarified model-seed payload mapping, line grouping across generated team sizes, bounded outcome downloads, and the limits of temporary idempotency across restarts. Verified required sections, milestones, source references, and whitespace. Kept horizontal scaling and expanded telemetry deferred, and stopped before implementation as requested.
|
|
348
|
+
|
|
349
|
+
Implementation update, 2026-09-25: implemented generation, service seed/idempotency extensions, deterministic resource enumeration, workspace/coordinator/gateway, metrics, observations, CSVs, and the initial React UI. The full Python suite passed 165 tests; focused frontend editor tests passed two tests. Production assets and the first Docker/wheel builds succeeded. Browser verification found native number-input step validation rejecting default efficiencies; corrected the min/step alignment and rerunning. Upgraded Plotly to 4.1.1 and Vitest to 4.1.11 after dependency audit, with zero reported npm vulnerabilities. The host Node 23 installation is outside Vitest's supported range, so validation uses a task-local Node 22 runtime; no global runtime was replaced. Chromium requires sandbox escalation on macOS.
|
|
350
|
+
|
|
351
|
+
Implementation revision, 2026-09-25 19:35Z: completed the seven feature areas and retained behavioral tests, clarified concrete interfaces and atomic editor persistence, recorded the Plotly selector fix and dependency choices, and replaced planning-only statements with observed runtime evidence. Final packaging and cleanup were verified: the wheel and Docker image serve byte-identical production assets and complete fresh three-model jobs, and the temporary validation processes/containers were stopped. No new blocking product decisions were needed.
|
|
352
|
+
|
|
353
|
+
Completion revision, 2026-09-25 19:40Z: final wheel `value_stream-0.1.0-py3-none-any.whl` SHA-256 is `2ec1b8012a736f20ebf7c3f266630b15042ca8c042ae46e596d654eb085d1bc8`. `final-smoke.json` records byte-identical assets from installed wheel and Docker, readiness, and three successful scenarios in each. Full Python validation passed 177 tests (one upstream Starlette/AnyIO deprecation warning); three frontend tests and three production Chromium tests passed. Both OpenAPI snapshots match live schemas, generated types compile, the uv lock is unchanged/valid, and whitespace checks pass. No blocking product questions remain. The user's unrelated staged `.codex/config.toml` was preserved.
|