runbios-mcp 0.2.13 → 0.2.14-dev.244

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/server.js CHANGED
@@ -772,6 +772,7 @@ Use this by default for published Run BiOS model ids. There is no GPU selection,
772
772
  2. Read the published ids and capabilities from the serverless catalog; do not substitute a dedicated catalog repo id for a serverless slug.
773
773
  3. Call chat_with_inference with that model id, messages, and only capabilities the model advertises. Streaming is optional.
774
774
  4. The completion is the result. Do NOT call preflight_inference or create_inference for serverless.
775
+ 5. If this credential has analytics:read, use get_serverless_limits and get_serverless_usage_daily for workspace RPM/spend settings and full-history admitted usage; get_serverless_usage_overview, get_serverless_usage_by_model, list_serverless_requests and get_serverless_usage_timeseries show retained request detail. These never reveal another workspace or per-key usage.
775
776
 
776
777
  ## Dedicated deployment — reserved GPU, durable endpoint
777
778
  Choose this for dedicated capacity, custom/trained checkpoints, or a durable endpoint under your control. It creates a paid resource and bills GPU time per second.
@@ -847,6 +848,14 @@ get_inference_status instead of asserting them.
847
848
  10. chat_with_inference - Call a serverless catalog model by id or a dedicated deployment via the unified /v1 endpoint (optional streaming)
848
849
  11. delete_inference - Permanently remove a deployment after verified infrastructure teardown
849
850
 
851
+ ## Serverless Workspace Observability (analytics:read)
852
+ - get_serverless_limits - Read saved workspace RPM and monthly spend cap; the org's tier may clamp effective admission further
853
+ - get_serverless_usage_overview / get_serverless_usage_by_model / get_serverless_usage_timeseries - Read retained billed usage, per-model costs and per-bucket series without provider or prompt content
854
+ - list_serverless_requests - Read bounded per-request outcomes and platform failures, not per-key usage or customer prompts
855
+ - get_serverless_usage_daily - Read admitted request totals and per-day points from the same full-history workspace counters (including failed admitted requests)
856
+ - get_serverless_savings - Read the authenticated user and workspace's model-promotion savings
857
+ Only the current workspace is readable with analytics:read; serverless inference scope alone is not analytics. Key creation, per-key usage, org-wide totals, and all RPM/spend mutations remain authenticated-console-only.
858
+
850
859
  ## Conscious Loop MCP Tools
851
860
  The loop turns what a deployed model actually did into what it learns next.
852
861
  NOTHING IS RECORDED until loop_set_capture turns a source on — capture is a
@@ -870,7 +879,7 @@ consent decision, not a default, and these tools cannot bypass it.
870
879
  17. loop_preflight_training_rule - PRICE a standing training rule without creating anything: the pinned revision, the worst hourly price each GPU ladder can reach, how many hours each ceiling buys, the refusals, and terms_text — the sentence the member is agreeing to, with their own figures in it. Always first.
871
880
  18. loop_create_training_rule - Create it. SPEND CONSENT, and recorded: accept_terms is the member's signature on that sentence. Read the estimate and terms_text back VERBATIM, wait for an explicit yes, and never state a ceiling, a price cap or an hour count preflight did not return.
872
881
  19. loop_list_training_rules / loop_get_training_rule / loop_update_training_rule / loop_consent_training_rule / loop_run_training_rule / loop_delete_training_rule - Read, edit, re-consent, fire now, retire. A money-bearing edit bumps the rule's revision, drops the consent and STOPS FUTURE FIRINGS until somebody accepts the new terms.
873
- 20. loop_list_training_runs / loop_get_training_run - What each firing did, its timeline, and available_actions — which says what this reader may actually do next, rather than what the state looks like it allows.
882
+ 20. loop_list_training_runs / loop_get_training_run / loop_list_pipelines / loop_get_pipeline - What each firing did, its timeline, and available_actions — which says what this reader may actually do next, rather than what the state looks like it allows. The pipeline tools read a rule as the versions it produced: which one serves, how many of max_versions are made, the run in flight, and the month's spend against its limit.
874
883
  21. loop_get_evaluation / loop_list_evaluation_items / loop_get_judge_agreement - The comparison behind a verdict: the win rate and the per-dimension means, the paired conversations themselves, and how far the judge already agrees with this workspace's own reviewers. Read the warnings out, not just the verdict.
875
884
  22. loop_promote_training_run / loop_reject_training_run / loop_rollback_training_run / loop_cancel_training_run - The decision. Promotion re-points the serving name at the candidate and changes what real customers get; it is undoable for 30 days.
876
885
  23. loop_get_agent_settings / loop_update_agent_settings / loop_update_build_rule - The agent's default model, its system prompts and the monthly evaluation cap; and the standing curation rule a training rule draws from.
@@ -975,6 +984,14 @@ get_inference_status instead of asserting them.
975
984
  10. chat_with_inference - Call a serverless catalog model by id or a dedicated deployment via the unified /v1 endpoint (optional streaming)
976
985
  11. delete_inference - Permanently remove a deployment after verified infrastructure teardown
977
986
 
987
+ ## Serverless Workspace Observability (analytics:read)
988
+ - get_serverless_limits - Read saved workspace RPM and monthly spend cap; the org's tier may clamp effective admission further
989
+ - get_serverless_usage_overview / get_serverless_usage_by_model / get_serverless_usage_timeseries - Read retained billed usage, per-model costs and per-bucket series without provider or prompt content
990
+ - list_serverless_requests - Read bounded per-request outcomes and platform failures, not per-key usage or customer prompts
991
+ - get_serverless_usage_daily - Read admitted request totals and per-day points from the same full-history workspace counters (including failed admitted requests)
992
+ - get_serverless_savings - Read the authenticated user and workspace's model-promotion savings
993
+ Only the current workspace is readable with analytics:read; serverless inference scope alone is not analytics. Key creation, per-key usage, org-wide totals, and all RPM/spend mutations remain authenticated-console-only.
994
+
978
995
  ## Conscious Loop MCP Tools
979
996
  The loop turns what a deployed model actually did into what it learns next.
980
997
  NOTHING IS RECORDED until loop_set_capture turns a source on — capture is a
@@ -998,7 +1015,7 @@ consent decision, not a default, and these tools cannot bypass it.
998
1015
  17. loop_preflight_training_rule - PRICE a standing training rule without creating anything: the pinned revision, the worst hourly price each GPU ladder can reach, how many hours each ceiling buys, the refusals, and terms_text — the sentence the member is agreeing to, with their own figures in it. Always first.
999
1016
  18. loop_create_training_rule - Create it. SPEND CONSENT, and recorded: accept_terms is the member's signature on that sentence. Read the estimate and terms_text back VERBATIM, wait for an explicit yes, and never state a ceiling, a price cap or an hour count preflight did not return.
1000
1017
  19. loop_list_training_rules / loop_get_training_rule / loop_update_training_rule / loop_consent_training_rule / loop_run_training_rule / loop_delete_training_rule - Read, edit, re-consent, fire now, retire. A money-bearing edit bumps the rule's revision, drops the consent and STOPS FUTURE FIRINGS until somebody accepts the new terms.
1001
- 20. loop_list_training_runs / loop_get_training_run - What each firing did, its timeline, and available_actions — which says what this reader may actually do next, rather than what the state looks like it allows.
1018
+ 20. loop_list_training_runs / loop_get_training_run / loop_list_pipelines / loop_get_pipeline - What each firing did, its timeline, and available_actions — which says what this reader may actually do next, rather than what the state looks like it allows. The pipeline tools read a rule as the versions it produced: which one serves, how many of max_versions are made, the run in flight, and the month's spend against its limit.
1002
1019
  21. loop_get_evaluation / loop_list_evaluation_items / loop_get_judge_agreement - The comparison behind a verdict: the win rate and the per-dimension means, the paired conversations themselves, and how far the judge already agrees with this workspace's own reviewers. Read the warnings out, not just the verdict.
1003
1020
  22. loop_promote_training_run / loop_reject_training_run / loop_rollback_training_run / loop_cancel_training_run - The decision. Promotion re-points the serving name at the candidate and changes what real customers get; it is undoable for 30 days.
1004
1021
  23. loop_get_agent_settings / loop_update_agent_settings / loop_update_build_rule - The agent's default model, its system prompts and the monthly evaluation cap; and the standing curation rule a training rule draws from.
@@ -2641,6 +2658,24 @@ create (1-5 choices), and the queue itself stays opt-in.`,
2641
2658
  return json(data);
2642
2659
  });
2643
2660
  }
2661
+ const serverlessUsageWindow = z.enum(["1h", "24h", "7d", "30d", "90d"]);
2662
+ const serverlessUsageWindowInput = serverlessUsageWindow.optional().describe("Analytics window (default 24h; for full-history admitted totals use daily). Raw request details are retained for about seven days.");
2663
+ server.tool("get_serverless_limits", "Read the current workspace's saved RPM ceiling and monthly spend cap in cents. analytics:read and current workspace access are required. This does not change a limit or return key details; the org's tier may clamp the effective RPM further.", {}, async () => json(await client.serverlessRead("/api/serverless/limits", "limits")));
2664
+ server.tool("get_serverless_usage_overview", "Read successful serverless request/token counts, actual billed spend (including any charged failures), and latency in this workspace over a retained window. Requires analytics:read; does not expose organization totals.", { window: serverlessUsageWindowInput }, async ({ window }) => json(await client.serverlessRead("/api/serverless/usage/overview", "overview", { window: window ?? "24h" })));
2665
+ server.tool("get_serverless_usage_by_model", "Read this workspace's billed serverless usage by model, with no provider identity or prompts. Requires analytics:read.", { window: serverlessUsageWindowInput }, async ({ window }) => json(await client.serverlessRead("/api/serverless/usage/by-model", "rows", { window: window ?? "24h" })));
2666
+ server.tool("list_serverless_requests", "Read this workspace's bounded per-request serverless outcomes (including platform failures), without prompt/completion content or key usage details. Requires analytics:read. Raw detail is retained for about seven days.", {
2667
+ window: serverlessUsageWindowInput,
2668
+ outcome: z.enum(["all", "ok", "failed", "rejected"]).optional(),
2669
+ request_id: z.string().min(1).max(255).optional(),
2670
+ limit: z.number().int().min(1).max(500).optional().describe("Maximum log rows (default 100, hard maximum 500)."),
2671
+ }, async ({ window, outcome, request_id, limit }) => json(await client.serverlessRead("/api/serverless/usage/requests", "rows", {
2672
+ window: window ?? "24h", outcome: outcome ?? "all", request_id, limit: String(limit ?? 100),
2673
+ })));
2674
+ server.tool("get_serverless_usage_timeseries", "Read this workspace's per-bucket serverless requests, tokens, spend or streaming latency/throughput. Requires analytics:read. Missing streaming-only buckets are null, not zero.", { window: serverlessUsageWindowInput, metric: z.enum(["requests", "tokens", "spend", "ttft_p50", "ttft_p95", "tps"]) }, async ({ window, metric }) => json(await client.serverlessRead("/api/serverless/usage/timeseries", "timeseries", {
2675
+ window: window ?? "24h", metric,
2676
+ })));
2677
+ server.tool("get_serverless_usage_daily", "Read this workspace's full-history admitted requests, tokens and spend alongside its per-day series from the same counters (failed admitted requests count; synthetic probes do not). Requires analytics:read. Returns the server-clamped 1–92 day range.", { days: z.number().int().min(1).max(92).optional().describe("Number of UTC days (default 30, maximum 92).") }, async ({ days }) => json(await client.serverlessRead("/api/serverless/usage/daily", "daily", { days: String(days ?? 30) })));
2678
+ server.tool("get_serverless_savings", "Read what model promotions saved the signed-in key creator and their workspace, over the selected window, today, and over the retained lifetime. Requires analytics:read and current workspace access.", { window: serverlessUsageWindow.optional() }, async ({ window }) => json(await client.serverlessRead("/api/serverless/usage/savings", "savings", { window: window ?? "30d" })));
2644
2679
  /* ══════════════════════════════════════════════════════════════════════════ */
2645
2680
  /* TOOLS: Conscious Loop */
2646
2681
  /* */
@@ -2762,10 +2797,16 @@ create (1-5 choices), and the queue itself stays opt-in.`,
2762
2797
  ground_truth: z.string().optional().describe("The value or fact the answer can be checked against"),
2763
2798
  reason: z.string().optional().describe("Why, in your own words"),
2764
2799
  author: z.string().optional(),
2765
- }, async ({ trace_id, verdict, source, correction, score, ground_truth, reason, author }) => {
2800
+ labels: z.array(z.string()).optional().describe("Which training pipeline this feedback feeds, named HERE rather than in a second call. A build rule selects conversations by label and a training rule trains from that build rule, so a label is the pipeline a conversation goes down. Sent with the verdict it is written in one transaction: either both land or neither does. A separate labelling call is the one that gets skipped, and what it leaves is a reviewed conversation in no pipeline — counted in every 'reviewed' total and selected by nothing."),
2801
+ attributes: z.record(z.string(), z.string()).optional().describe("The same, by named dimension: {\"category\": \"billing\"}. Dimensions AND together when a set is built, which is what lets one corpus become a different dataset per task."),
2802
+ parent: z.string().optional().describe("The master group the bare labels belong to, so selecting the group picks them up without anybody maintaining a list."),
2803
+ }, async ({ trace_id, verdict, source, correction, score, ground_truth, reason, author, labels, attributes, parent }) => {
2766
2804
  const data = await client.api(`/api/loop/traces/${encodeURIComponent(trace_id)}/signals`, {
2767
2805
  method: "POST",
2768
- body: { verdict, source: source ?? "human", correction, score, ground_truth, reason, author },
2806
+ body: {
2807
+ verdict, source: source ?? "human", correction, score, ground_truth, reason, author,
2808
+ labels, attributes, parent,
2809
+ },
2769
2810
  });
2770
2811
  return json(data);
2771
2812
  });
@@ -3079,6 +3120,8 @@ create (1-5 choices), and the queue itself stays opt-in.`,
3079
3120
  cadence_weekday: z.number().int().min(0).max(6).nullable().optional().describe("0 is Sunday. null DROPS the weekday, which is what a weekly rule moved to daily needs."),
3080
3121
  combinator: z.enum(["and", "or"]).optional().describe("'and' fires only when the schedule is due AND enough new rows exist; 'or' fires on either"),
3081
3122
  min_new_rows: z.number().int().min(1).nullable().optional().describe("Rows reviewed SINCE THE LAST FIRING, not the whole corpus. null removes the row floor, leaving the schedule as the only trigger."),
3123
+ max_versions: z.number().int().min(1).max(10).optional().describe("How many versions this pipeline may make, 1 to 10; five if absent on a create, unchanged if absent on an update. A version is a trained model whose comparison scored conversations, or one a member put live. At the limit the pipeline stops and a manual run is refused with VERSION_LIMIT_REACHED until the limit is raised. Each version is a paid run, so the limit is also a bound on how many the rule will pay for."),
3124
+ explore_recipes: z.boolean().optional().describe("MONEY-BEARING. true lets the pipeline, once it has made a version and nothing reviewed since adds anything new to train on (nothing was reviewed, or the reviews give it no row the last version's set lacked, such as thumbs-down with no correction), train the SAME base model on the SAME conversations with ONE training setting changed (the learning rate, the epochs, the LoRA rank) and compare it like any other version; it only takes over if it beats the model serving the app. Every such variant is a FULL PAID RUN inside the rule's per-run ceilings and its monthly limit, so changing it in EITHER direction changes the terms: turning it on or off pauses the rule, and it makes no version of any kind until somebody accepts the terms again. Turning it off to save money stops the pipeline too. Say that before you set it. Absent on a create means off; absent on an update leaves it alone."),
3082
3125
  }).optional(),
3083
3126
  serving: z.object({
3084
3127
  kind: z.enum(["serverless_slug", "deployment"]).describe("What serves this model to customers today"),
@@ -3090,7 +3133,13 @@ create (1-5 choices), and the queue itself stays opt-in.`,
3090
3133
  model_revision: z.string().optional().describe("Pinned once at save so a rule cannot silently start training a different set of weights. Preflight returns the revision it pinned; pass it back rather than inventing one."),
3091
3134
  training_method: z.enum(["sft", "rlhf"]).optional().describe("'rlhf' is refused until the platform enables it; preflight says so rather than the rule failing at fire time"),
3092
3135
  rlhf_type: z.enum(["dpo", "kto"]).nullable().optional().describe("null clears it, which is what moving a rule back to sft means"),
3093
- train_type: z.enum(["lora", "qlora", "full"]).optional(),
3136
+ // lora or qlora only: loop-service refuses `full` on preflight, create AND
3137
+ // update for a standing rule, so advertising it here would hand an agent a
3138
+ // legal-looking choice that is guaranteed to come back a 400 — on the tool
3139
+ // its own description marks "Always first", with no estimate and no terms
3140
+ // text to show the person whose wallet pays. One-off full fine-tunes are
3141
+ // unaffected; they are a different tool.
3142
+ train_type: z.enum(["lora", "qlora"]).optional(),
3094
3143
  config: z.record(z.string(), z.any()).optional().describe("Hyperparameters, the same shape create_training_job takes"),
3095
3144
  train_gpu_priorities: z.array(trainingGPURungSchema).optional().describe("The training ladder, best rung first. The run takes the first rung in stock under the price cap."),
3096
3145
  train_max_price_hour_cents: z.number().int().min(1).optional().describe("The most this rule will pay per hour to TRAIN, in cents"),
@@ -3107,7 +3156,7 @@ create (1-5 choices), and the queue itself stays opt-in.`,
3107
3156
  training_ceiling_cents: z.number().int().min(1).optional().describe("The hard stop on one run's training spend, in cents"),
3108
3157
  candidate_ceiling_cents: z.number().int().min(1).optional().describe("The hard stop on the candidate's serving spend for one run, in cents"),
3109
3158
  eval_ceiling_cents: z.number().int().min(1).optional().describe("The hard stop on the comparison's model calls for one run, in cents"),
3110
- monthly_ceiling_cents: z.number().int().min(1).nullable().optional().describe("Everything this rule may spend in a calendar month. null REMOVES the monthly cap; absent leaves it exactly where it stands."),
3159
+ monthly_ceiling_cents: z.number().int().min(1).nullable().optional().describe("The monthly limit. A run starts only if the most this rule's runs this UTC calendar month can have cost (month_spent_cents, an upper bound, not an exact spend), plus the most one more run can cost, fits under it. null REMOVES the monthly cap; absent leaves it exactly where it stands."),
3111
3160
  eval_max_rows: z.number().int().min(1).optional().describe("How many held-out conversations the comparison scores"),
3112
3161
  }).optional(),
3113
3162
  evaluation: z.object({
@@ -3142,8 +3191,8 @@ create (1-5 choices), and the queue itself stays opt-in.`,
3142
3191
  }, async ({ enabled, limit, offset }) => json(await client.api("/api/loop/training-rules", {
3143
3192
  params: { enabled: enabled?.toString(), limit: limit?.toString(), offset: offset?.toString() },
3144
3193
  })));
3145
- server.tool("loop_get_training_rule", "One training rule in full: the consent it stands on, its recent runs, what it has spent this calendar month, and how far its judge agrees with this workspace's own reviewers. Report month_spent_cents against monthly_ceiling_cents when the user asks what this is costing them, and say so plainly if accepted_revision is behind revision — the rule is paused until somebody consents again.", { rule_id: z.string().describe("The rule's id, from loop_list_training_rules") }, async ({ rule_id }) => json(await client.api(`/api/loop/training-rules/${encodeURIComponent(rule_id)}`)));
3146
- server.tool("loop_update_training_rule", "Edit or pause a training rule. Send ONLY what is changing: an absent key leaves that setting alone and a present null clears the ones that can be cleared (monthly_ceiling_cents, cadence_weekday, min_new_rows, rlhf_type, judge_id, judge_model). A MONEY-BEARING CHANGE — any ceiling, either price cap, the model, the method or the ladders — bumps the rule's revision, drops the consent and STOPS FUTURE FIRINGS until somebody accepts the new terms; the response says so as consent_required, and the honest next step is loop_preflight_training_rule again, not loop_consent_training_rule on a sentence nobody re-read. Pass expected_revision to be refused rather than overwrite an edit somebody else made.", {
3194
+ server.tool("loop_get_training_rule", "One training rule in full: the consent it stands on, its recent runs, what its runs this UTC calendar month have cost at most, and how far its judge agrees with this workspace's own reviewers. Report month_spent_cents against monthly_ceiling_cents when the user asks what this is costing them, as 'up to' that amount: it is an upper bound (a run still in progress counts at its ceilings, and a comparison machine at the most it could have billed), never an exact spend, and say so plainly if accepted_revision is behind revision — the rule is paused until somebody consents again.", { rule_id: z.string().describe("The rule's id, from loop_list_training_rules") }, async ({ rule_id }) => json(await client.api(`/api/loop/training-rules/${encodeURIComponent(rule_id)}`)));
3195
+ server.tool("loop_update_training_rule", "Edit or pause a training rule. Send ONLY what is changing: an absent key leaves that setting alone and a present null clears the ones that can be cleared (monthly_ceiling_cents, cadence_weekday, min_new_rows, rlhf_type, judge_id, judge_model). A MONEY-BEARING CHANGE — any ceiling, either price cap, the model, the method, the ladders or changing trigger.explore_recipes in EITHER direction — bumps the rule's revision, drops the consent and STOPS FUTURE FIRINGS until somebody accepts the new terms. Turning recipe variants OFF does this too: asked to turn them off to save money, tell the user first that the pipeline then makes no version at all, variant or not, until the terms are accepted again; the response says so as consent_required, and the honest next step is loop_preflight_training_rule again, not loop_consent_training_rule on a sentence nobody re-read. Pass expected_revision to be refused rather than overwrite an edit somebody else made.", {
3147
3196
  rule_id: z.string().describe("The rule's id, from loop_list_training_rules"),
3148
3197
  ...trainingRuleInputShape,
3149
3198
  expected_revision: z.number().int().min(0).optional().describe("The revision you read. A stale value is refused with REVISION_MISMATCH instead of overwriting somebody else's edit."),
@@ -3160,7 +3209,7 @@ create (1-5 choices), and the queue itself stays opt-in.`,
3160
3209
  method: "POST",
3161
3210
  body: { terms_version, revision },
3162
3211
  })));
3163
- server.tool("loop_run_training_rule", "Fire a rule NOW, without waiting for its cadence or its row floor. SPEND CONSENT: this starts a real training job and boots a real candidate against the rule's ceilings, so say what it will cost before you call it. The floors that protect the result still apply — too few curated or held-out rows answers NOT_ENOUGH_ROWS with the counts, an unconsented rule answers CONSENT_REQUIRED, and a rule already mid-run answers RUN_ACTIVE rather than starting a second one.", { rule_id: z.string() }, async ({ rule_id }) => json(await client.api(`/api/loop/training-rules/${encodeURIComponent(rule_id)}/run`, {
3212
+ server.tool("loop_run_training_rule", "Fire a rule NOW, without waiting for its cadence or its row floor. SPEND CONSENT: this starts a real training job and boots a real candidate against the rule's ceilings, so say what it will cost before you call it. The floors that protect the result still apply — too few curated or held-out rows answers NOT_ENOUGH_ROWS with the counts, an unconsented rule answers CONSENT_REQUIRED, a switched-off or paused rule answers RULE_PAUSED, a rule already mid-run answers RUN_ACTIVE rather than starting a second one, a pipeline that has made max_versions versions answers VERSION_LIMIT_REACHED, a run whose worst case would take this month past the rule's monthly limit answers MONTHLY_LIMIT_REACHED with month_spent_cents (what this month's runs have cost at most, not an exact spend), monthly_ceiling_cents, run_max_cents and resumes_at — read those figures back rather than retrying — and a workspace whose Conscious Loop agent is not turned on answers AGENT_OFF, because every comparison runs under the agent's key and a model trained without one could not be compared: tell the user to turn the agent on in Agent settings, and do not retry until they have.", { rule_id: z.string() }, async ({ rule_id }) => json(await client.api(`/api/loop/training-rules/${encodeURIComponent(rule_id)}/run`, {
3164
3213
  method: "POST",
3165
3214
  body: {},
3166
3215
  })));
@@ -3174,6 +3223,8 @@ create (1-5 choices), and the queue itself stays opt-in.`,
3174
3223
  params: { rule_id, state, limit: limit?.toString(), offset: offset?.toString() },
3175
3224
  })));
3176
3225
  server.tool("loop_get_training_run", "One run in full, with the timeline of every state it passed through and why, links to the job, the candidate, the datasets and the comparison, and available_actions — which says what THIS reader may actually do next. Use available_actions rather than guessing from the state: a run can look promotable and not be, because the consent moved or the candidate is already gone.", { run_id: z.string() }, async ({ run_id }) => json(await client.api(`/api/loop/training-runs/${encodeURIComponent(run_id)}`)));
3226
+ server.tool("loop_list_pipelines", "Every training rule in the workspace seen as a PIPELINE: the series of versions it has produced, which one serves the app today (champion_version, null when what serves is not one of its versions), how each version did against the champion of its day and on the standing benchmark, versions_made against max_versions, the run in flight and which version it will become, and status. Read-only. When the user asks 'is this getting better' or 'what is it doing', answer from this, and read status back honestly: needs_review and needs_funds are waiting on THEM (a decision, a top-up), paused names why in paused_reason or needs_consent, complete means the version limit is reached, and running is the platform working. next_reason is the rule's own sentence about what it is waiting for. month_spent_cents against monthly_ceiling_cents is what decides whether the next run may start; month_spent_cents is what runs started this UTC month have cost at most (a run still in progress counts at its ceilings), so say 'up to', never 'spent'. recipe_exploration says whether a recipe variant can run next and, when not, why: off, waiting_for_v1 (no version to vary yet), available (recipe_variants_left untried), no_slots (untried variants left, recipe_variants_left above 0, but no version slot free; raise max_versions, or at 10 start a new pipeline, and one can run — but no_slots with recipe_variants_left 0 means all_tried, so never suggest raising the limit for it), all_tried (every variant that fits has been tried; raising max_versions does not start one, new reviews bring the next version) or none_fit (no variant fits the rule's settings). Read that, never recipe_variants_left alone: 0 does not mean every variant was tried. The list is the newest 100: total is how many pipelines the workspace has and truncated is true when more exist than were returned. When truncated, say it is the newest of total, and page through loop_list_training_rules for the rest.", {}, async () => json(await client.api("/api/loop/pipelines")));
3227
+ server.tool("loop_get_pipeline", "One pipeline in full, by its training rule's id: every version with its verdict, the decision a person made, its win rate and benchmark score, what it cost and the recipe it trained on (a version whose trigger is 'variant' retrained the same conversations with one setting changed, named in recipe_note), plus the run in flight, explore_recipes, recipe_variants_left and recipe_exploration (why a variant can or cannot run next; read it rather than the count), and what this month's runs have cost at most (month_spent_cents, an upper bound, not an exact spend) against the monthly limit. Read-only: promote, reject and roll back stay on the run tools, and the version limit and explore_recipes stay on loop_update_training_rule. Only compare benchmark scores whose benchmark_comparable is true; the others were measured on a different yardstick. A run with competed false is not a version yet, and says so.", { rule_id: z.string().describe("The training rule's id, from loop_list_pipelines or loop_list_training_rules") }, async ({ rule_id }) => json(await client.api(`/api/loop/pipelines/${encodeURIComponent(rule_id)}`)));
3177
3228
  server.tool("loop_promote_training_run", "Re-point the serving name at the candidate: from this moment the workspace's traffic is answered by the newly trained model. Read the comparison back to the user first (loop_get_evaluation) and get an explicit yes — this changes what real customers get, and it is only undoable for 30 days with loop_rollback_training_run. An INCONCLUSIVE comparison is refused unless force is set, and forcing one means promoting a model nothing showed to be better; say that out loud rather than setting force to get past the error.", {
3178
3229
  run_id: z.string(),
3179
3230
  expected_revision: z.number().int().min(0).optional().describe("The rule revision the decision was made against. Stale answers CONSENT_CHANGED instead of promoting under terms nobody agreed to."),