@noodleseed/agent-kit 0.34.0 → 0.35.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. package/manifest.json +57 -17
  2. package/package.json +1 -1
  3. package/skills/claude-code/SKILL.md +48 -29
  4. package/skills/claude-code/examples/acme-bistro/README.md +1 -1
  5. package/skills/claude-code/references/app-directory-compliance.md +59 -0
  6. package/skills/claude-code/references/build-an-mcp-app.md +52 -0
  7. package/skills/claude-code/references/build-an-mcp-server.md +54 -0
  8. package/skills/claude-code/references/connect-an-api.md +60 -20
  9. package/skills/claude-code/references/deploy-and-ops.md +15 -79
  10. package/skills/claude-code/references/experience-design.md +1 -1
  11. package/skills/claude-code/references/inspect-hosted.md +26 -0
  12. package/skills/claude-code/references/publishing.md +15 -17
  13. package/skills/claude-code/references/verify-and-recover.md +65 -0
  14. package/skills/codex/SKILL.md +48 -29
  15. package/skills/codex/examples/acme-bistro/README.md +1 -1
  16. package/skills/codex/references/app-directory-compliance.md +59 -0
  17. package/skills/codex/references/build-an-mcp-app.md +52 -0
  18. package/skills/codex/references/build-an-mcp-server.md +54 -0
  19. package/skills/codex/references/connect-an-api.md +60 -20
  20. package/skills/codex/references/deploy-and-ops.md +15 -79
  21. package/skills/codex/references/experience-design.md +1 -1
  22. package/skills/codex/references/inspect-hosted.md +26 -0
  23. package/skills/codex/references/publishing.md +15 -17
  24. package/skills/codex/references/verify-and-recover.md +65 -0
  25. package/skills/claude-code/references/chatgpt-compliance.md +0 -63
  26. package/skills/codex/references/chatgpt-compliance.md +0 -63
@@ -0,0 +1,52 @@
1
+ # Outcome
2
+
3
+ Deliver an MCP App whose visual interaction gives the user a concrete benefit beyond a good text response, while preserving useful model-visible output when the widget is unavailable.
4
+
5
+ ## Use when
6
+
7
+ - The user asks for an MCP App, widget, interactive card, visual workflow, or host-visible UI.
8
+ - Comparison, selection, progress, editing, confirmation, or another visual interaction materially improves the conversational job.
9
+
10
+ ## Do not use when
11
+
12
+ - A concise text or structured tool result fully serves the user. UI must earn its place.
13
+ - The requested task is a headless server, API connector, diagnosis, deployment, or publication with no UI change; select that route.
14
+ - The agent lacks the product inputs needed to explain who benefits, what action the UI enables, and what happens without it.
15
+
16
+ ## Required inputs
17
+
18
+ Before implementation, capture a short design spec: target user, conversational job, explicit user benefit, information hierarchy, primary interaction, states (loading/empty/error/success), model-visible result, widget-only data, and useful text fallback. Use `references/experience-design.md` for the deeper product-design questions only when needed.
19
+
20
+ ## Workflow
21
+
22
+ 1. **Pass the UI fit check.** State why a visual interaction is better than text for this request. If there is no defensible user benefit, keep the capability headless and stop the App route.
23
+ 2. **Agree on the design spec.** Describe the smallest complete experience and its states before writing the component. Avoid recreating a full dashboard or website inside the conversation.
24
+ 3. **Define the output boundary.** Keep concise facts and action results model-visible. Put presentation-heavy or interactive widget data in the widget-only channel. The model must not depend on opaque UI state to continue the conversation.
25
+ 4. **Preserve fallback.** Every tool that launches a widget must still return useful text without the widget, so unsupported hosts and failed rendering remain usable.
26
+ 5. **Author and wire the App contract.** Follow `references/widgets-and-apps.md` for the canonical component guidance, view registration, hooks, state, CSP, tool visibility, and output shaping. Keep tool effects and confirmation semantics correct independently of the UI.
27
+ 6. **Validate the local artifact.** Run `noodle validate --json`, `noodle test --json`, and `noodle check --json`. Repair failures at the layer that produced them.
28
+ 7. **Inspect the experience.** Run `noodle devtools` and verify loading, empty, error, success, responsive layout, focus/keyboard behavior, and the text fallback.
29
+ 8. **Escalate evidence only on request.** Run a host test only when the user requested host verification. Run host-specific compliance only when preparing that host submission; select the exact host-testing or compliance entry from the router lookup catalog only after that evidence level is explicitly requested.
30
+
31
+ ## Verification evidence
32
+
33
+ - **Product:** the design spec states the user benefit and the UI fit decision.
34
+ - **Server:** `noodle validate --json` and `noodle test --json` succeeded.
35
+ - **App contract:** `noodle check --json` succeeded.
36
+ - **Local UX:** `noodle devtools` exercised the relevant states and the useful text fallback without the widget.
37
+ - **Host/compliance:** report each requested host or compliance check with its evidence; report every unperformed higher level as not run.
38
+
39
+ ## Recovery paths
40
+
41
+ - Weak UI fit: remove the widget and ship the stronger headless result, or narrow the visual interaction to the one decision it improves.
42
+ - App check failure: repair the cited view, metadata, output, CSP, or accessibility issue and rerun `noodle check --json` before reopening devtools.
43
+ - Blank or stale widget: verify the tool returns the intended widget data, the view is registered, and state derives from supported hooks rather than hidden global state.
44
+ - Model cannot continue without UI: move the essential facts into model-visible output and keep only presentation data widget-only.
45
+ - Host-only mismatch: record local checks as passed, isolate the host symptom, and select the host-testing lookup only for that observed host; do not rewrite a working local contract without host evidence.
46
+
47
+ ## Stop conditions
48
+
49
+ - Stop complete at the locally requested boundary when product fit, server tests, App checks, devtools states, and text fallback are evidenced.
50
+ - Stop before host connection, deployment, or submission unless the user requested that next evidence level.
51
+ - Stop blocked when the required design decision, external data, credentials, or host access is unavailable; name the missing input and the exact next action.
52
+ - Never claim host compatibility, directory compliance, or production behavior from local devtools evidence alone.
@@ -0,0 +1,54 @@
1
+ # Outcome
2
+
3
+ Deliver the smallest useful Noodle Seed MCP server that turns a real user intent into a safe, typed result. Author only the configured TypeScript entrypoint, normally `src/server.ts`; keep the public authoring surface TypeScript-only and never hand-author generated manifests or connector IR.
4
+
5
+ ## Use when
6
+
7
+ - The user asks to create or extend a headless MCP server, tools, resources, prompts, or connector-backed behavior.
8
+ - The requested result is primarily model-facing and does not require a widget or host-visible UI.
9
+
10
+ ## Do not use when
11
+
12
+ - The primary outcome is an MCP App, widget, or visual interaction; select the App route.
13
+ - The task is only to diagnose existing failures, deploy, publish, embed, or report feedback; select that dedicated route.
14
+ - The idea has no conversational fit: static content, a dashboard, deep navigation, or a full existing app port should be narrowed to the few actions that are better said than clicked.
15
+
16
+ ## Required inputs
17
+
18
+ Establish only the inputs needed for the requested stopping point. Follow `references/authoring-workflow.md` for the canonical discovery paths. Do not guess or invent a private schema, endpoint, authentication model, eligibility rule, or approval flow. If a required input is unavailable, state exactly what evidence is missing and stop before fabricating behavior.
19
+
20
+ ## Workflow
21
+
22
+ 1. **Confirm conversational fit.** Name one to three focused jobs where saying the request is easier than navigating the underlying system, and identify the data or action the model cannot provide by itself.
23
+ 2. **Define the product contract.** For each job, write the user phrase, the intent-shaped tool or resource, its minimal typed input, the useful output, read/write effect, and backing operation. Design for user intent, not a 1:1 API endpoint wrapper.
24
+ 3. **Choose the smallest implementation.** Use native tools, resources, or prompts for local/static behavior; add a connector only when external data or actions are required. Keep response output small and model-readable.
25
+ 4. **Author in TypeScript.** Follow `references/authoring-workflow.md` for connector and flow patterns and `references/sdk-surface.md` for exact builders. These are this route’s complete canonical support set; use the router lookup catalog only when observed evidence names a different concern.
26
+ 5. **Validate and repair.** Run `noodle validate --json`. Parse `error.errors[]`, repair the cited `path`, and rerun validation. Consult the lookup catalog only for the specific reported error code; do not open another reference speculatively.
27
+ 6. **Run the local smoke.** After validation succeeds, run `noodle test --json` and repair any failure at that evidence layer.
28
+ 7. **Prove external behavior.** For connector-backed reads, set credentials through the effective local target and run a safe representative `noodle tools call`. Confirm populated mapped fields from real output, not merely successful registration.
29
+ 8. **Stop at the requested boundary.** Do not add an App, host test, hosted environment, publication work, or deployment unless the user requested that outcome. Deploy only when the selected route or the user explicitly requires it.
30
+
31
+ ## Verification evidence
32
+
33
+ Report evidence as a ladder and claim only levels actually exercised:
34
+
35
+ - **Authoring:** the requested TypeScript behavior exists with typed inputs and outputs.
36
+ - **Compilation:** `noodle validate --json` returned success.
37
+ - **Local smoke:** `noodle test --json` returned success.
38
+ - **Connector reality:** a representative safe read via `noodle tools call` returned populated mapped fields. This is required for connector-backed work.
39
+ - **Higher levels:** explicitly report host, deployment, and production checks as not run unless they were separately requested and evidenced.
40
+
41
+ ## Recovery paths
42
+
43
+ - Validation failure: fix each structured error at its reported path, rerun validation, then resume at the next unproven layer.
44
+ - Tool registers but returns empty or `undefined` fields: inspect one sanitized real response, correct `${response...}` mappings, and rerun the same read.
45
+ - Credential unavailable: verify `secret(...)` naming and the effective local target; never inline or print the secret.
46
+ - Missing product input: ask for the smallest concrete example, schema, or rule that unblocks the selected job. Do not widen the build to compensate.
47
+ - Repeated failure at the same layer: stop after two evidence-backed repair attempts with the same failure signature and report the command, sanitized error, evidence already proven, and exact next action.
48
+
49
+ ## Stop conditions
50
+
51
+ - Stop complete when the requested behavior passes validation and local smoke, and every connector-backed read has real-output evidence.
52
+ - Stop at the user's requested boundary; do not deploy unless the user requested deployment.
53
+ - Stop blocked when progress requires unavailable credentials, private schemas, external approval, or a live write the user has not approved.
54
+ - In the handoff, name what changed, what passed, what was not run, and any remaining risk without upgrading local evidence into a hosted or production claim.
@@ -1,10 +1,13 @@
1
- # Connect a live API (you were given a key)
1
+ # Outcome
2
2
 
3
- When the user hands you an API key or credentials, don't infer the data from documentation docs
4
- drift. Probe the live API, learn the real shape, then encode it as a `connector`. The loop:
3
+ Connect a real API to a focused MCP product using managed credentials, mappings derived from observed responses, and representative live-read evidence. A connector that merely compiles is not complete.
5
4
 
6
5
  ## Contents
7
6
 
7
+ - Use when
8
+ - Do not use when
9
+ - Required inputs
10
+ - Workflow
8
11
  - Secure the key first
9
12
  - Probe the live API
10
13
  - Model the connector from the observed shape
@@ -13,9 +16,30 @@ drift. Probe the live API, learn the real shape, then encode it as a `connector`
13
16
  - Design intent tools
14
17
  - Set the secret for local runs
15
18
  - Prove real output
16
- - Then build the app
19
+ - Verification evidence
20
+ - Recovery paths
21
+ - Stop conditions
17
22
 
18
- ## Secure the key first
23
+ ## Use when
24
+
25
+ - The user provides credentials, a reachable API, or an OpenAPI document and wants real MCP behavior backed by it.
26
+ - Existing connector behavior compiles but still needs proof against the actual service and data shape.
27
+
28
+ ## Do not use when
29
+
30
+ - The task is a local/static MCP capability with no external data source.
31
+ - The user only wants a widget, deployment, publication, or diagnosis unrelated to API behavior; select that route.
32
+ - Required credentials or authority are unavailable. Do not bypass authentication or substitute fabricated payloads for live evidence.
33
+
34
+ ## Required inputs
35
+
36
+ Identify the API base URL, authentication scheme, one representative safe read, the user intent it serves, and either an OpenAPI document or one sanitized example response. For writes, also establish the effect, a safe test target, and explicit user approval before any live write.
37
+
38
+ Do not guess or invent a field, schema, endpoint, pagination contract, or authentication behavior. Documentation is a hypothesis until a representative live read confirms the response actually returned.
39
+
40
+ ## Workflow
41
+
42
+ ### Secure the key first
19
43
 
20
44
  Never inline or log the key. Have the user put it in an environment variable, then store it as a
21
45
  managed secret and reference it only as `secret(...)`:
@@ -28,7 +52,7 @@ noodle secrets set SOME_API_KEY --runtime local --from-env SOME_API_KEY # same
28
52
  In `server.ts` the key is only ever `secret("SOME_API_KEY")` — keep the raw value out of code, tests,
29
53
  prompts, logs, and generated files.
30
54
 
31
- ## Probe the live API
55
+ ### Probe the live API
32
56
 
33
57
  Learn the actual response shape empirically. Two ways — capture one real example response per endpoint
34
58
  you will use, and read its field names, nesting, array shapes, pagination, and id-vs-label fields:
@@ -40,7 +64,7 @@ you will use, and read its field names, nesting, array shapes, pagination, and i
40
64
  '${response}' }`), `noodle secrets set` the key, then `noodle tools call` it to see the real payload
41
65
  in-process.
42
66
 
43
- ## Model the connector from the observed shape
67
+ ### Model the connector from the observed shape
44
68
 
45
69
  Encode the API as an HTTP connector, mapping only the fields you actually saw into a small typed
46
70
  `output`:
@@ -56,7 +80,7 @@ Encode the API as an HTTP connector, mapping only the fields you actually saw in
56
80
  The full connector shape, every `auth.kind`, and compute connectors are in
57
81
  `references/authoring-workflow.md`.
58
82
 
59
- ## Return a list
83
+ ### Return a list
60
84
 
61
85
  Most real tools return a variable-length list (search results, a user’s tasks). Bind the **whole array** — a single `${response.path}` returns the referenced value verbatim, arrays included:
62
86
 
@@ -87,7 +111,7 @@ pagination: {
87
111
  response: { tasks: '${response.items}' },
88
112
  ```
89
113
 
90
- ## Create, update, delete
114
+ ### Create, update, delete
91
115
 
92
116
  Pair the read/list with the mutations your intent tools need:
93
117
  - **Create / update** — `method: 'POST'` / `'PATCH'`; author the body as `request: { field: '${input.x}' }` (do not nest it under `body`). It is JSON by default; use `requestEncoding: 'form-urlencoded'` only when the API requires a URLSearchParams body. URL query params remain the operation-level `query: [...]` array.
@@ -117,14 +141,14 @@ close_task: {
117
141
  },
118
142
  ```
119
143
 
120
- ## Design intent tools
144
+ ### Design intent tools
121
145
 
122
- Shape tools around what the user says, not 1:1 around endpoints. Pair an id-taking action with a
146
+ Create intent-shaped tools around what the user says, not 1:1 around endpoints. Pair an id-taking action with a
123
147
  find/search operation that returns `{ id, label }` summaries so the model resolves text → id itself,
124
148
  and map each response to a few labelled fields the model can speak from. See the "Design tools for the
125
149
  model" section of `references/authoring-workflow.md`.
126
150
 
127
- ## Set the secret for local runs
151
+ ### Set the secret for local runs
128
152
 
129
153
  Local `dev`, smoke commands, secrets, and variables resolve one effective target: explicit flags, then the project link, then the saved CLI target, then local defaults. Set the secret through that same target:
130
154
 
@@ -137,17 +161,33 @@ noodle secrets set SOME_API_KEY --runtime local --scope env --org <org> --app <a
137
161
 
138
162
  Local secrets live in `./.env.noodle` (never commit it). A required `secret(...)` or `variable(...)` that cannot resolve fails boot closed. `noodle tools call` / `noodle test` / `noodle dev` / `noodle devtools` stop before exposing an empty endpoint and print the exact effective target plus recovery command.
139
163
 
140
- ## Prove real output
164
+ ### Prove real output
141
165
 
142
166
  `noodle validate` / `noodle test` prove a connector tool *compiles and registers* — not that its
143
167
  mapping returns data. With the secret set, run a live read: `noodle tools call <read_tool> --args
144
168
  '{…}'` executes the connector against the real API in-process. Confirm the mapped fields are populated,
145
- not `undefined`; if they are empty, fix the `${response…}` paths against the real payload and re-run.
146
- Only run a live write if it is safe or the user approved it.
169
+ not `undefined`; if they are empty, distinguish a legitimate empty result from a missing or incorrect mapping, fix `${response…}` paths against the real payload when needed, and re-run.
170
+ Only run a live write after explicit user approval and when a safe test target and expected effect are known.
171
+
172
+ ## Verification evidence
173
+
174
+ - **Credential path:** the raw credential remained in an environment variable and the managed `secret(...)` path for the same effective local target.
175
+ - **Observed shape:** a representative safe live read established the real fields, nesting, arrays, pagination, and empty-result behavior used by the mapping.
176
+ - **Local proof:** `noodle validate --json` and `noodle test --json` succeeded, then `noodle tools call` returned populated mapped fields or an intentionally verified empty result.
177
+ - **Writes:** name the approval and safe target used, or report writes as not run.
178
+ - **Hosted boundary:** local proof does not prove hosted credentials, deployment health, or host behavior. Report hosted checks as not run unless a separate requested route exercised them.
179
+
180
+ ## Recovery paths
181
+
182
+ - Authentication failure: verify the connector auth kind, managed secret name, and effective local target without printing the credential.
183
+ - Successful HTTP call with `undefined` fields: compare the mapping with one sanitized observed response, correct the path, and rerun the same read.
184
+ - Legitimate empty result: test a second known query or record the empty case as intentional; do not rewrite a correct mapping merely to manufacture data.
185
+ - Response too broad for the model: narrow it with response mapping, projection, or a separate compute connector; do not rely on a Zod output to strip runtime fields.
186
+ - Repeated external failure: stop after bounded attempts and report the sanitized status, endpoint class, evidence already proven, and exact external action needed.
147
187
 
148
- ## Then build the app
188
+ ## Stop conditions
149
189
 
150
- With real data flowing, design the experience (`references/experience-design.md`), add widgets where a
151
- UI genuinely helps (`references/widgets-and-apps.md`), and verify with `noodle check`. Deploy per
152
- `references/deploy-and-ops.md`, and set the same secret in the hosted environment with `noodle secrets
153
- set` before the first hosted call.
190
+ - Stop complete when the representative safe read returns populated mapped fields or an intentionally verified empty result through the same effective local target.
191
+ - Stop before a live write without explicit approval, a known effect, and a safe target.
192
+ - Stop blocked when credentials, a reachable service, a representative input, or a required private schema is unavailable.
193
+ - Do not continue into App design, deployment, or publication unless the user requested that next outcome; route to the corresponding primary playbook instead.
@@ -1,89 +1,25 @@
1
- # Deploy and operations
1
+ # Hosted mutation authorization
2
2
 
3
- ## Contents
3
+ > This route changes hosted or external state. Use it only when the current user request explicitly authorizes the exact mutation and target.
4
4
 
5
- - Authenticate
6
- - Link and target
7
- - Deploy and inspect
8
- - Installed Developer plugin handoff
9
- - Eject path (portable manifest)
10
- - Connect into a host
11
- - Access modes
12
- - Org and members
13
- - Config and observability
14
- - Agent-safe CLI recipes
15
- - Analytics
5
+ ## Route boundary
16
6
 
17
- ## Authenticate
7
+ - Select this route only for the exact hosted mutation the user requested.
8
+ - Route inspection, diagnosis, preparation, validation, testing, and other read-only work to their read-only references. Those requests do not authorize a mutation.
9
+ - Authentication, target binding, configuration, access changes, deployment, connection writes, and rollback are separate mutations. Authorization for one does not imply another.
18
10
 
19
- `noodle login` to authenticate with Noodle Seed Cloud, `noodle whoami` to confirm, `noodle logout` to clear credentials. Local commands (`dev`/`validate`/`test`) need none of this.
11
+ ## Authorization check
20
12
 
21
- ## Link and target
13
+ Before any mutation, require the current request to name both the action and its complete target. A mutation-capable target consists of an explicit organization, application, and environment. When the environment is absent, stop and ask for it instead of applying a default, reusing local state, or selecting a target implicitly.
22
14
 
23
- `noodle link --org <slug> --app <slug> [--env <slug>]` binds the directory to a target. Unlinked, `noodle deploy` uses your default org, the project name as the app, and `prod`. `noodle target show|set` inspects or changes the target.
15
+ Do not broaden a request to prepare, inspect, diagnose, or validate into permission to authenticate, bind a target, change configuration or access, deploy, connect, submit, or roll back.
24
16
 
25
- ## Deploy and inspect
17
+ ## Command and service contract
26
18
 
27
- `noodle deploy` deploys the server. Then `noodle open` (latest URL), `noodle status`, `noodle inspect` (metadata, no secrets), `noodle smoke` (readiness diagnostics), and `noodle rollback <deploymentId>` to revert.
19
+ Use `references/cli-commands.md` as the generated command, flag, and exit-code contract. Consult the live command catalog before acting, and treat the service response as the authority for resulting hosted state. This reference intentionally does not duplicate operational command sequences, defaults, or status semantics.
28
20
 
29
- ## Installed Developer plugin handoff
21
+ ## Evidence and stop conditions
30
22
 
31
- When this project was bootstrapped by the installed Noodle Developer plugin, preserve its managed launcher invocation for every local CLI command. After `deploy --json`, take the returned deployment ID to the connected remote `noodle-developer.inspect_deployment` tool. Use `noodle-developer.diagnose_app` only when the Cloud evidence needs diagnosis. Those tools inspect the selected organization and environment; they do not edit or replace the application source, which remains the coding agent's responsibility.
32
-
33
- ## Eject path (portable manifest)
34
-
35
- `noodle export manifest [--output <file>]` compiles the entrypoint locally and emits the portable, vendor-neutral manifest JSON — no service, no login, no account. A Noodle app is just `src/server.ts` plus this manifest: the user can read it, diff it, and keep it.
36
-
37
- ## Connect into a host
38
-
39
- Once deployed, register the server as a tool in a host with `noodle connect <host>` (`claude-code`, `codex`, `chatgpt`, `cursor`, `vscode`, `claude`, `inspector`) — it prints the exact config to paste.
40
-
41
- - **Claude Code / Claude Desktop** (verified) — add the `mcpServers` block, or one-shot `claude mcp add-json noodle-server '<json>'`:
42
-
43
- ```json
44
- {
45
- "mcpServers": {
46
- "noodle-server": { "type": "https", "url": "https://<app>.mcp.noodleseed.dev" }
47
- }
48
- }
49
- ```
50
-
51
- - **Codex / Cursor / VS Code** — the same `mcpServers` block is emitted as a starting point (these hosts' config formats are not officially documented). Wiring a deployed Noodle server into Codex means registering that block in Codex's MCP config.
52
- - **ChatGPT / Claude.ai** — no config file: open the host's Settings → Connectors → Add custom connector, paste the MCP URL, then authenticate.
53
- - `noodle connect codex|claude-code --write` writes the project-local agent files (only these two targets).
54
-
55
- ## Access modes
56
-
57
- `noodle access set owner-only|org-members|authenticated|customers` controls who can call the deployed server. Hosted access is identity-based; never add static data-plane keys.
58
-
59
- ## Org and members
60
-
61
- `noodle orgs list|create` and `noodle members list|add|remove --org <slug>` manage organizations and membership.
62
-
63
- ## Config and observability
64
-
65
- Manage runtime config with `noodle secrets` / `noodle variables` (scoped org/app/env). Operators use `noodle logs`, `noodle audit`, and `noodle policy` for logs, governance audit, and policy.
66
-
67
- ## Agent-safe CLI recipes
68
-
69
- Use explicit flags in headless runs so commands never wait for a prompt:
70
-
71
- ```sh
72
- noodle link --org acme --app support-assistant --env prod
73
- noodle secrets set CRM_TOKEN --scope env --org acme --app support-assistant --env prod --from-env CRM_TOKEN
74
- noodle secrets set CRM_CERT --scope env --org acme --app support-assistant --env prod --from-file ./cert.pem
75
- printf %s "$CRM_TOKEN" | noodle secrets set CRM_TOKEN --scope env --org acme --app support-assistant --env prod --from-stdin
76
- noodle variables set CRM_BASE_URL --scope env --org acme --app support-assistant --env prod --value https://crm.example.com
77
- noodle secrets list --scope env --org acme --app support-assistant --env prod --json
78
- noodle validate --json
79
- noodle test --json
80
- noodle deploy --json
81
- noodle smoke --json
82
- noodle agents doctor --json
83
- ```
84
-
85
- `secrets resolve` is for local diagnostics only; do not print resolved values into prompts, logs, tests, or docs. Prefer `--from-env`, `--from-file`, or `--from-stdin` over inline `--value` for sensitive values. Variables may use `--value` when the value is non-secret.
86
-
87
- ## Analytics (verify after deploy, debug errors)
88
-
89
- After a deploy gets traffic, verify with `noodle metrics --agent-output` — it returns a `health` verdict (`ok`/`attention`), a one-line summary, and `attention[]` items each carrying the exact next command. When a tool errors, drill in with `noodle events --tool <name> --json` (filters: `--status tool_error|mcp_error`, `--client <name>`); `noodle events --session <id> --json` replays one session chronologically. `--json` on both returns the full payload; human runs get the branded report. Two-tier errors: `tool_error` is recoverable (handed back to the model), `mcp_error` needs attention (protocol/timeout/internal). Wire edge-triggered webhooks on error share, error count, calls, or p95 latency with `noodle alerts add|list|remove|test`.
23
+ - Stop before execution when the action or complete target is missing.
24
+ - After an authorized mutation, report only the state evidenced by the command and service response.
25
+ - Do not claim host behavior, production health, or successful external registration without direct evidence at that layer.
@@ -126,7 +126,7 @@ The design phase produces up to three artifacts — worked gold-standard version
126
126
  architecture, demo scope, success metrics, and future enhancements — opening on the funnel-boundary
127
127
  line every scope debate resolves against.
128
128
  - **Wireframe** — the single-file HTML alignment artifact (anatomy above) with the embedded compliance
129
- audit; see `references/chatgpt-compliance.md`.
129
+ audit; see `references/app-directory-compliance.md`.
130
130
  - **API contract** — when the partner's backend must be built or wrapped. Escalate: (1) the MCP
131
131
  tool→call-sequence map (always); (2) "Recommended API Shapes" — concrete request/response JSON per
132
132
  tool, including the hardest nested case; (3) a full OpenAPI spec for transactional apps. Contract
@@ -0,0 +1,26 @@
1
+ # Inspect hosted state
2
+
3
+ Read hosted evidence without changing target, credentials, configuration, access, host wiring, revisions, or directory state.
4
+
5
+ ## Use when
6
+
7
+ - The user asks for hosted status, deployment metadata, health, logs, events, metrics, audit evidence, or diagnosis.
8
+ - The request is inspect-only, diagnose-only, or asks whether an existing deployment works.
9
+
10
+ ## Authority boundary
11
+
12
+ This route is read-only. It never authorizes `login`, `logout`, `link`, `target set`, hosted secret/variable/config/access changes, `deploy`, `rollback`, host configuration writes, or directory submission. If evidence shows one of those actions is needed, report the exact proposed action and target, then stop for a new explicit user request.
13
+
14
+ ## Workflow
15
+
16
+ 1. Resolve the requested org, app, environment, and deployment from existing non-secret context. Do not change the effective target to make inspection easier.
17
+ 2. Choose the narrowest read-only command: `noodle target show`, `noodle status`, `noodle inspect`, `noodle smoke`, `noodle metrics --agent-output`, `noodle events --json`, `noodle logs`, or `noodle audit`.
18
+ 3. Prefer machine output when the selected command supports it. Record the target, revision/deployment ID, timestamp, result, and any request ID without exposing secrets or customer payloads.
19
+ 4. When the installed Developer MCP is available, use its deployment inspection or diagnosis tool only for the selected org/app/env. Treat it as evidence gathering, not mutation authority.
20
+ 5. If a command fails, distinguish missing authentication/access from unhealthy application behavior. Do not repair, relink, redeploy, rotate config, or roll back under this route.
21
+
22
+ ## Stop conditions
23
+
24
+ - Stop complete when the requested hosted fact is supported by current evidence and higher untested levels are named.
25
+ - Stop blocked when existing access cannot read the target or the requested evidence requires a host/user journey unavailable in scope.
26
+ - Stop for authorization when the next useful action would mutate local targeting, hosted state, host configuration, or directory state.
@@ -1,31 +1,29 @@
1
1
  # Publish to app directories
2
2
 
3
- Directory requirements evolve treat this as the workflow map and verify against the host’s current submission docs before submitting.
3
+ > Preparation is read-only unless the current user request explicitly authorizes the exact deploy, access change, host write, or submission target. A request to prepare must report missing readiness work and stop before mutation.
4
+
5
+ Directory requirements evolve. Identify the requested directory first and verify its current official requirements before preparing directory-specific evidence.
4
6
 
5
7
  ## Contents
6
8
 
7
- - Readiness gate
8
- - ChatGPT apps directory
9
- - Claude connectors directory
9
+ - Shared readiness gate
10
+ - Directory-specific evidence
11
+ - Submission boundary
10
12
 
11
- ## Readiness gate
13
+ ## Shared readiness gate
12
14
 
13
15
  Before any submission:
14
16
 
15
- 1. `noodle check --target chatgpt` must be clean — every widget needs `domain` (one https origin per app) and an exact `csp` (hosts require the CSP to list precisely the domains you fetch from).
16
- 2. Audit tool responses in developer mode: run realistic prompts and strip anything not strictly needed — PII, internal identifiers (session/trace/request IDs, internal account IDs), and any secrets.
17
- 3. The server must be deployed and publicly reachable: `noodle deploy`, confirm with `noodle open --print` and `noodle smoke`. Reviewers connect to the real endpoint never submit a placeholder or loopback URL, and the access mode must not be `owner-only` (`noodle access set`).
18
- 4. Polish the listing surface: tool descriptions, widget titles, and the `server` branding tokens are what reviewers and users see.
17
+ Use `references/app-directory-compliance.md` as this route’s canonical shared compliance checklist.
18
+
19
+ Prepare evidence for a reachable production MCP endpoint, accurate capability descriptions and schemas, useful fallback behavior, realistic positive and negative tests, data minimization, privacy disclosures, support ownership, and any interactive surface the directory will review.
19
20
 
20
- ## ChatGPT apps directory
21
+ ## Directory-specific evidence
21
22
 
22
- Submit from the OpenAI developer dashboard (platform.openai.com Apps):
23
+ Read the selected directory’s current official submission documentation at review time. Record each additional requirement separately from the shared checklist, including listing fields, identity verification, test credentials, screenshots, policy declarations, review limits, and appeal or resubmission steps. Never project one directory’s requirements onto another.
23
24
 
24
- - Complete organization identity verification first (individual or business) it is enforced at review time.
25
- - The submission form asks for the app name, logo, description, company and privacy policy URLs, MCP server URL and tool information, screenshots, test prompts with expected responses, and localization details.
26
- - One version may be published and one in review at a time; to revise a pending submission, cancel the review and resubmit rather than creating a new app.
27
- - Review combines automated checks and manual evaluation; rejections come with feedback — fix and resubmit, or reply to appeal. An approved app is also distributed as a Codex plugin.
25
+ When a requirement cannot be verified from the selected directory’s current documentation or direct review evidence, mark it unknown instead of borrowing a rule from another host.
28
26
 
29
- ## Claude connectors directory
27
+ ## Submission boundary
30
28
 
31
- Anthropic runs a connectors directory for Claude; submission goes through Anthropic’s published process (see the Anthropic connectors directory FAQ on support.claude.com). The same readiness gate applies: deployed public endpoint, clean `noodle check`, and graceful degradation where Apps rendering is unavailable.
29
+ Preparation is read-only. Deployment, access changes, directory registration, and final submission each require explicit authorization for the exact target. Report remaining evidence gaps and stop when that authority or required directory access is absent.
@@ -0,0 +1,65 @@
1
+ # Outcome
2
+
3
+ Identify the first failing evidence layer, repair only that layer, rerun it, and report the highest level actually proven. A successful lower layer must never be presented as proof of a higher one.
4
+
5
+ ## Use when
6
+
7
+ - The user asks to validate, test, diagnose, recover, or establish whether a local or hosted Noodle Seed project works.
8
+ - A command, connector, App, host integration, deployment, or production check is failing or has uncertain evidence.
9
+
10
+ ## Do not use when
11
+
12
+ - The primary request is to design or build a new product capability; select its build route and use this playbook only if evidence fails.
13
+ - The user asks for a higher-risk external action rather than diagnosis. This route does not grant deployment, publication, live-write, or merge authority.
14
+
15
+ ## Required inputs
16
+
17
+ Capture the requested evidence level, the exact command or user-visible symptom, sanitized machine output, the environment/target, and the last known passing layer. Do not broaden the goal beyond the level the user asked to prove.
18
+
19
+ ## Workflow
20
+
21
+ Use this ordered evidence ladder. Start at the last known passing layer or the lowest plausible failure; never jump upward over an unproven dependency:
22
+
23
+ 1. **Compile** — the TypeScript build and authoring import surface are valid.
24
+ 2. **Validate** — `noodle validate --json` accepts the Noodle contract.
25
+ 3. **Local smoke** — `noodle test --json` starts the local runtime and exercises registration.
26
+ 4. **Real API** — a representative safe `noodle tools call` proves connector credentials, transport, observed mapping, and populated data.
27
+ 5. **App compliance** — `noodle check --json` and local devtools prove the App contract and intended states.
28
+ 6. **Host** — the requested host connects, invokes the expected capability, and renders useful fallback/UI behavior.
29
+ 7. **Deploy** — the requested hosted revision and configuration exist and report healthy at the deployment layer.
30
+ 8. **Production health** — the live production endpoint and requested user journey are observed on the intended revision.
31
+
32
+ For the first failing layer:
33
+
34
+ 1. Read the process exit code or status first. If machine JSON exists, parse it before reading human prose or editing files.
35
+ 2. For validation envelopes, inspect every `error.errors[]` item and repair the field at its reported `path`. Use `references/agent-contract.md` for the envelope and `references/compile-errors.md` for the named error code.
36
+ 3. Form one evidence-backed cause from the observed output. If the two canonical supports do not cover it, select one matching symptom from the router lookup catalog; do not scan every recovery path.
37
+ 4. Make the smallest in-scope repair. Do not freeform re-edit adjacent code, change credentials, redeploy, or add product behavior without evidence and authority.
38
+ 5. Rerun only the same evidence layer that failed. Once it passes, continue upward only to the user-requested level.
39
+ 6. Stop after two evidence-backed repair attempts with the same failure signature, or immediately when the next action requires new authority or external state.
40
+
41
+ ## Verification evidence
42
+
43
+ Report a compact ledger for every exercised layer: command/action, target, result, and the evidence it establishes. Claim only the highest contiguous passing layer.
44
+
45
+ - Compile success does not prove runtime behavior.
46
+ - Validation and local smoke do not prove a real API mapping or credential path.
47
+ - Local evidence does not prove hosted or host behavior.
48
+ - Deployment existence does not prove production health or a user journey.
49
+ - Report every requested but unperformed or blocked higher layer as not run, with the reason.
50
+
51
+ ## Recovery paths
52
+
53
+ - Compile/validation: repair the exact import, schema, or reported path, then rerun that command without freeform changes.
54
+ - Local boot/smoke: use the structured startup error to correct the effective target, config, or entrypoint before retrying.
55
+ - Real API: distinguish authentication, reachability, legitimate empty results, and broken response mappings before changing code.
56
+ - App: repair the cited contract or state in `noodle check --json`, then confirm it in devtools before attempting a host.
57
+ - Host/deployment/production: confirm revision, target, identity, and configuration independently; do not infer one from another.
58
+ - Repeated external failure: preserve passing evidence and report the sanitized failure, required authority or external state, owner, and exact next action.
59
+
60
+ ## Stop conditions
61
+
62
+ - Stop complete when the user-requested evidence level and every dependency below it pass in the current target.
63
+ - Stop blocked when progress requires credentials, approval, host access, deployment authority, production access, or an external-state change not available in scope.
64
+ - Stop after two evidence-backed repair attempts with the same failure signature at one layer; do not hide repetition behind unrelated edits.
65
+ - Never claim fixed or working without rerunning the failed layer, and never upgrade compile, local, deployment, or stale historical evidence into a stronger claim.
@@ -1,63 +0,0 @@
1
- # ChatGPT App compliance (pre-submission)
2
-
3
- `noodle check --target chatgpt` verifies the *metadata* prerequisites; app-store submission also faces a
4
- human review against OpenAI’s Apps SDK UX principles. Run this checklist against the built app before
5
- submitting, and render it as an audit table in the design wireframe (`design/wireframe.html` in the
6
- `acme-*` examples) so partners and reviewers see it up front.
7
-
8
- ## Contents
9
-
10
- - Metadata gate vs review
11
- - Pre-submission checklist
12
- - UI guidelines
13
- - Domain guardrails
14
- - Privacy and data
15
-
16
- ## Metadata gate vs review
17
-
18
- `noodle check --target chatgpt --json` returning `ok:true` means the widget is *metadata-ready* (widget
19
- `domain`, `openai/outputTemplate`, CSP, tool annotations, and `invoking`/`invoked` invocation copy are
20
- present) — it does NOT prove host rendering, conversation UX, or submission acceptance. Validate real
21
- rendering in ChatGPT Developer Mode / MCP Inspector, then run the checklist below.
22
-
23
- ## Pre-submission checklist (what review looks for)
24
-
25
- 1. **Conversational value** — at least one capability relies on ChatGPT’s strengths: natural-language
26
- actions no tap-driven app can do (e.g. "two margheritas and a lemon tart" parses into a cart). Cite
27
- concrete app behavior, not aspirations.
28
- 2. **Beyond base ChatGPT** — new knowledge, actions, or presentation (grounded partner data, live
29
- inventory, signed handoffs, real-world routing).
30
- 3. **Atomic, model-friendly actions** — self-contained tools with explicit input/output schemas, and an
31
- annotation on every tool (`annotations.readOnly()` / `.action()` / `.openAction()`).
32
- 4. **Helpful UI only** — justify each widget (would plain text degrade UX?), and note what you
33
- deliberately did NOT build a widget for (payment is off-app → no payment widget).
34
- 5. **In-chat task completion** — the user finishes a meaningful task in chat. For a top-of-funnel app,
35
- the task is the discovery/config loop completed in-chat with an intentional handoff.
36
- 6. **Performance** — tool calls scoped per step; response-time targets stated.
37
- 7. **Discoverability** — broad, natural trigger prompts listed; description keywords planned. Golden
38
- prompt sets and metadata optimization are a launch workstream, not polish.
39
- 8. **Platform fit** — multi-turn dialogue, conversation memory, and multimodality where genuinely useful.
40
-
41
- ## UI guidelines
42
-
43
- System fonts, monochrome outlined icons, WCAG AA contrast, at most two actions on inline cards, no nested
44
- scroll, and the right display mode per intent (inline by default; fullscreen only where browsing needs
45
- it; picture-in-picture only for live state). Brand only through `server` `branding` tokens — accent on
46
- the primary CTA, logo, and badges, nothing else; the compiler derives the palette. Never inject raw
47
- global CSS.
48
-
49
- ## Domain guardrails
50
-
51
- For regulated-adjacent apps, add app-specific trust behaviors and **show them in the rendered pixels**:
52
- cite the source and its revision for consequential lookups; frame regulated content as "considerations,
53
- not a ruling"; never invent compatibility, availability, or pricing; and always show the relevant
54
- caution/disclaimer. These are what make a regulated-adjacent app approvable.
55
-
56
- ## Privacy and data
57
-
58
- Data flows through OpenAI; tool payloads and whatever the server stores must match the partner’s privacy
59
- policy. No payment happens in chat (PCI stays off-app). Avoid per-user OAuth in a top-of-funnel v1 (use
60
- service credentials via a `connector`); add end-user auth only for two-way apps (`customerAuth`). Keep
61
- secrets out of tool output, widgets, and logs. If the partner’s published policy predates the app, flag a
62
- privacy gap for their counsel before submission. Re-run this checklist against the *built* app before
63
- every submission — not just the wireframe.