rag-wright 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (843) hide show
  1. rag_wright-0.1.0/.claude/settings.json +14 -0
  2. rag_wright-0.1.0/.claude/skills/authoring-a-capability/SKILL.md +106 -0
  3. rag_wright-0.1.0/.claude/skills/classifier-opportunity-analysis/SKILL.md +175 -0
  4. rag_wright-0.1.0/.claude/skills/creating-evals/SKILL.md +114 -0
  5. rag_wright-0.1.0/.claude/skills/laya/SKILL.md +119 -0
  6. rag_wright-0.1.0/.claude/skills/qwen-vllm-modal/SKILL.md +96 -0
  7. rag_wright-0.1.0/.claude/skills/setfit/SKILL.md +293 -0
  8. rag_wright-0.1.0/.claude/skills/using-the-rag-wright-engine/SKILL.md +90 -0
  9. rag_wright-0.1.0/.env.example +20 -0
  10. rag_wright-0.1.0/.github/workflows/publish.yml +56 -0
  11. rag_wright-0.1.0/.gitignore +53 -0
  12. rag_wright-0.1.0/CHANGELOG.md +48 -0
  13. rag_wright-0.1.0/CLAUDE.md +201 -0
  14. rag_wright-0.1.0/LICENSE +21 -0
  15. rag_wright-0.1.0/PKG-INFO +168 -0
  16. rag_wright-0.1.0/README.md +118 -0
  17. rag_wright-0.1.0/SPEC.md +267 -0
  18. rag_wright-0.1.0/conftest.py +41 -0
  19. rag_wright-0.1.0/docs/ARCHITECTURE_OVERVIEW.md +199 -0
  20. rag_wright-0.1.0/docs/ArcadeDB_Local.md +43 -0
  21. rag_wright-0.1.0/docs/OBSERVABILITY.md +88 -0
  22. rag_wright-0.1.0/docs/adr/0001-stack-and-library-choices.md +50 -0
  23. rag_wright-0.1.0/docs/adr/0002-validation-corpus-cuad-edgar.md +52 -0
  24. rag_wright-0.1.0/docs/adr/0003-ard-registration-and-capability-kinds.md +96 -0
  25. rag_wright-0.1.0/docs/adr/0004-entity-disambiguation.md +105 -0
  26. rag_wright-0.1.0/docs/adr/0005-relational-golden-set.md +89 -0
  27. rag_wright-0.1.0/docs/adr/0006-model-profile.md +66 -0
  28. rag_wright-0.1.0/docs/adr/0007-arcadedb-store-schema.md +110 -0
  29. rag_wright-0.1.0/docs/adr/0008-unified-framework-index-docs-grounding.md +111 -0
  30. rag_wright-0.1.0/docs/adr/0009-rlm-chunking-retained-as-configurable-capability.md +69 -0
  31. rag_wright-0.1.0/docs/adr/0010-transformers-pinned-below-5-for-flagembedding-reranker.md +46 -0
  32. rag_wright-0.1.0/docs/adr/0011-cuad-not-a-retrieval-benchmark-acord-for-queries.md +70 -0
  33. rag_wright-0.1.0/docs/adr/0012-entity-mention-confidence-no-proximity-edges.md +69 -0
  34. rag_wright-0.1.0/docs/adr/0013-entity-resolution-matching-strategy.md +48 -0
  35. rag_wright-0.1.0/docs/adr/0014-split-generation-and-vision-to-text.md +51 -0
  36. rag_wright-0.1.0/docs/adr/0015-rlm-dynamic-subagents-and-granted-subagents.md +107 -0
  37. rag_wright-0.1.0/docs/adr/0016-rlm-defined-by-required-capabilities-enforced-as-tests.md +42 -0
  38. rag_wright-0.1.0/docs/adr/0017-dynamic-dispatch-trigger-as-typed-flag.md +77 -0
  39. rag_wright-0.1.0/docs/adr/0018-rlm-method-is-the-orchestrator-system-prompt.md +65 -0
  40. rag_wright-0.1.0/docs/adr/0019-recursion-is-optional-for-chunking-required-for-synthesis.md +51 -0
  41. rag_wright-0.1.0/docs/adr/0020-serialize-interpreter-sessions-per-process-ki1.md +97 -0
  42. rag_wright-0.1.0/docs/adr/0021-capability-interface-governed-typed-io.md +106 -0
  43. rag_wright-0.1.0/docs/adr/0022-fr-k-embedding-free-okf-navigation-experimental.md +161 -0
  44. rag_wright-0.1.0/docs/adr/0023-cheap-model-for-okf-enrichment.md +51 -0
  45. rag_wright-0.1.0/docs/adr/0024-reader-parallelism-lives-in-python-not-the-interpreter.md +57 -0
  46. rag_wright-0.1.0/docs/adr/0025-retrieval-pivot-function-classify-property-graph-rerank.md +82 -0
  47. rag_wright-0.1.0/docs/adr/0026-property-schema-and-extended-function-taxonomy.md +68 -0
  48. rag_wright-0.1.0/docs/adr/0027-openrouter-provider-routing-by-throughput.md +46 -0
  49. rag_wright-0.1.0/docs/adr/0028-deterministic-grounding-judge-and-flash-pro-cascade.md +52 -0
  50. rag_wright-0.1.0/docs/adr/0029-retrieval-pipeline-domain-portability-and-adaptation.md +94 -0
  51. rag_wright-0.1.0/docs/adr/0030-model-training-standard-modal-reusable-checkpointed.md +63 -0
  52. rag_wright-0.1.0/docs/adr/0031-single-call-chunking-for-structured-contracts.md +57 -0
  53. rag_wright-0.1.0/docs/adr/0032-nl-to-type-two-step-reason-emit-on-gemma.md +58 -0
  54. rag_wright-0.1.0/docs/adr/0033-unified-contract-kg-three-legs-one-graph.md +56 -0
  55. rag_wright-0.1.0/docs/adr/0034-granite-json-schema-structured-output-profile.md +44 -0
  56. rag_wright-0.1.0/docs/adr/0035-graph-extraction-rebacked-with-gp1b-retire-hybrid.md +54 -0
  57. rag_wright-0.1.0/docs/adr/0036-party-clause-link-party-to-edge.md +62 -0
  58. rag_wright-0.1.0/docs/adr/0037-clause-template-is-authoritative-code-not-generated.md +59 -0
  59. rag_wright-0.1.0/docs/adr/0038-full-cuad-kg-on-gcp-gpu-vm-and-gcs-backup.md +54 -0
  60. rag_wright-0.1.0/docs/adr/0039-self-hosted-open-model-stack-on-modal-a100.md +59 -0
  61. rag_wright-0.1.0/docs/adr/0040-neuro-symbolic-extraction-fidelity-ontology-shacl-validation.md +82 -0
  62. rag_wright-0.1.0/docs/adr/0041-capability-kind-rubric-and-agent-skill-runtime-tiers.md +81 -0
  63. rag_wright-0.1.0/docs/adr/0042-clause-level-span-id-provenance-and-content-hash-backfill.md +56 -0
  64. rag_wright-0.1.0/docs/adr/0043-retire-cross-corpus-retrieval-standardize-on-leg-b.md +47 -0
  65. rag_wright-0.1.0/docs/adr/0044-is-exception-to-derived-carveout-relationship.md +122 -0
  66. rag_wright-0.1.0/docs/adr/0045-client-side-tag-parse-structured-output.md +82 -0
  67. rag_wright-0.1.0/docs/adr/0046-acord-unified-into-one-production-kg.md +50 -0
  68. rag_wright-0.1.0/docs/adr/0047-retire-precomputed-clause-function-gate.md +44 -0
  69. rag_wright-0.1.0/docs/adr/0048-ingest-llm-classifier-nondestructive-reclassify.md +187 -0
  70. rag_wright-0.1.0/docs/adr/0049-generic-customer-lens-for-ingestion.md +63 -0
  71. rag_wright-0.1.0/docs/adr/0050-async-langgraph-ingestion.md +87 -0
  72. rag_wright-0.1.0/docs/adr/0051-schema-bootstrap-and-feedback-driven-ontology-evolution.md +127 -0
  73. rag_wright-0.1.0/docs/adr/0052-engine-product-split-graphwright-parked.md +119 -0
  74. rag_wright-0.1.0/docs/adr/0053-answer-prose-output-hygiene.md +72 -0
  75. rag_wright-0.1.0/docs/adr/0054-remove-auto-tag-from-generator-evidence.md +75 -0
  76. rag_wright-0.1.0/docs/adr/0055-confidence-out-of-band.md +68 -0
  77. rag_wright-0.1.0/docs/adr/0056-bound-structured-retry-wall-clock.md +96 -0
  78. rag_wright-0.1.0/docs/adr/0057-async-engine-architecture.md +82 -0
  79. rag_wright-0.1.0/docs/adr/0058-structure-first-chunking.md +75 -0
  80. rag_wright-0.1.0/docs/adr/0059-recall-decoupled-from-classification-and-visible-partial-loss.md +102 -0
  81. rag_wright-0.1.0/docs/adr/0060-scope-compliance-check-to-named-policy-sources.md +83 -0
  82. rag_wright-0.1.0/docs/adr/0061-subject-document-compliance-per-section.md +59 -0
  83. rag_wright-0.1.0/docs/adr/0062-tiered-ocr-scan-quality-vlm-escalation.md +86 -0
  84. rag_wright-0.1.0/docs/adr/0063-per-sentence-compliance-subject-facts.md +61 -0
  85. rag_wright-0.1.0/docs/adr/0064-typed-properties-out-of-band-on-evidence.md +58 -0
  86. rag_wright-0.1.0/docs/adr/0065-deontic-applicability-gates-query-side.md +74 -0
  87. rag_wright-0.1.0/docs/adr/0066-ontology-ttl-single-source-of-truth.md +175 -0
  88. rag_wright-0.1.0/docs/adr/0067-domain-pack-retargeting.md +93 -0
  89. rag_wright-0.1.0/docs/adr/0068-recall-first-actor-gate-ontology-role-disjointness.md +73 -0
  90. rag_wright-0.1.0/docs/adr/0069-reading-order-chunking-tables-figures-retrievable.md +80 -0
  91. rag_wright-0.1.0/docs/adr/0070-text-layer-first-parsing-no-false-vlm-escalation.md +51 -0
  92. rag_wright-0.1.0/docs/adr/0071-defragmentation-reconstruct-paragraphs-from-line-items.md +57 -0
  93. rag_wright-0.1.0/docs/adr/0072-extract-guard-furniture-and-deterministic-failure.md +52 -0
  94. rag_wright-0.1.0/docs/adr/0073-born-digital-threshold-sparse-pages-no-vlm-escalation.md +50 -0
  95. rag_wright-0.1.0/docs/adr/0074-retry-transient-extraction-failures.md +44 -0
  96. rag_wright-0.1.0/docs/adr/0075-per-page-vlm-escalation.md +45 -0
  97. rag_wright-0.1.0/docs/adr/0076-thread-safe-shared-embedder-reranker.md +55 -0
  98. rag_wright-0.1.0/docs/adr/0077-concurrent-function-classification-across-chunks.md +45 -0
  99. rag_wright-0.1.0/docs/adr/0078-dedicated-extraction-executor.md +57 -0
  100. rag_wright-0.1.0/docs/adr/0079-product-default-granite-4.2-openrouter-routing.md +62 -0
  101. rag_wright-0.1.0/docs/adr/0080-nested-tag-parse-and-degrade.md +48 -0
  102. rag_wright-0.1.0/docs/adr/0081-function-independent-tagparse-clause-extraction.md +70 -0
  103. rag_wright-0.1.0/docs/adr/0082-symbolic-gate-function-independent.md +49 -0
  104. rag_wright-0.1.0/docs/adr/0083-executor-hop-trace-context-capture.md +31 -0
  105. rag_wright-0.1.0/docs/adr/0084-gleaning-off-on-the-query-leg.md +24 -0
  106. rag_wright-0.1.0/docs/adr/0085-query-constraint-extraction-tagparse.md +24 -0
  107. rag_wright-0.1.0/docs/adr/0086-streaming-cost-capture.md +29 -0
  108. rag_wright-0.1.0/docs/adr/0087-retrieval-relevance-score-on-rankedspan.md +22 -0
  109. rag_wright-0.1.0/docs/adr/0088-per-span-relevance-verdict.md +32 -0
  110. rag_wright-0.1.0/docs/adr/0089-structured-output-generation-tracing.md +26 -0
  111. rag_wright-0.1.0/docs/adr/0090-affiliate-of-extraction.md +25 -0
  112. rag_wright-0.1.0/docs/adr/0091-retire-partyto-edge.md +29 -0
  113. rag_wright-0.1.0/docs/adr/0092-idempotent-write-graph.md +30 -0
  114. rag_wright-0.1.0/docs/adr/0093-entities-by-name-seam.md +28 -0
  115. rag_wright-0.1.0/docs/adr/0094-workspace-document-scope.md +25 -0
  116. rag_wright-0.1.0/docs/adr/0095-span-page-provenance.md +24 -0
  117. rag_wright-0.1.0/docs/adr/0096-carveout-keyword-normalization.md +25 -0
  118. rag_wright-0.1.0/docs/adr/0097-caller-configurable-ingest-extraction-models.md +29 -0
  119. rag_wright-0.1.0/docs/adr/0098-invoke-time-document-scope.md +24 -0
  120. rag_wright-0.1.0/docs/adr/0099-mcp-tools-never-take-a-model-supplied-tenant.md +27 -0
  121. rag_wright-0.1.0/docs/adr/0100-profile-based-model-routing.md +30 -0
  122. rag_wright-0.1.0/docs/adr/0101-untagged-spans-reach-extraction-aspect-gate-removed.md +26 -0
  123. rag_wright-0.1.0/docs/adr/0102-open-descriptive-list-dims-retain-verbatim.md +33 -0
  124. rag_wright-0.1.0/docs/adr/0103-a-clause-is-a-provision-not-a-span.md +41 -0
  125. rag_wright-0.1.0/docs/adr/0104-dense-floor-protection-in-leg-b.md +41 -0
  126. rag_wright-0.1.0/docs/adr/0105-in-band-model-usage-accounting.md +26 -0
  127. rag_wright-0.1.0/docs/adr/0106-document-signals-off-the-citation.md +25 -0
  128. rag_wright-0.1.0/docs/adr/0107-requirement-page-provenance.md +25 -0
  129. rag_wright-0.1.0/docs/adr/0108-qwen3-27b-single-a100-serving-profile.md +47 -0
  130. rag_wright-0.1.0/docs/adr/0109-vllm-cold-start-reduction.md +54 -0
  131. rag_wright-0.1.0/docs/adr/0110-fp8-kv-cache-16k-single-a100.md +56 -0
  132. rag_wright-0.1.0/docs/adr/0111-pin-qwen3-27b-deepinfra-bf16.md +25 -0
  133. rag_wright-0.1.0/docs/adr/0112-relax-deepagents-pin-to-floor.md +26 -0
  134. rag_wright-0.1.0/docs/adr/0113-llm-span-instrumentation-for-latency-attribution.md +29 -0
  135. rag_wright-0.1.0/docs/adr/0114-setfit-default-clause-function-classifier.md +45 -0
  136. rag_wright-0.1.0/docs/adr/0115-classifier-only-step3a-property-extraction.md +61 -0
  137. rag_wright-0.1.0/docs/adr/0116-soft-function-scoping-classifier-lane.md +45 -0
  138. rag_wright-0.1.0/docs/adr/0117-engine-api-layer-and-capability-runtime.md +90 -0
  139. rag_wright-0.1.0/docs/adr/0118-engine-core-api-vs-ard-adapter-free-impl-ref-client.md +100 -0
  140. rag_wright-0.1.0/docs/adr/0119-jev-typed-decision-model-for-compliance-closed-set-fields.md +68 -0
  141. rag_wright-0.1.0/docs/adr/0120-rlm-sub-agent-identity-dynamic-dispatch-trigger.md +35 -0
  142. rag_wright-0.1.0/docs/adr/0121-spacy-optional-extra-model-runtime-download.md +44 -0
  143. rag_wright-0.1.0/docs/adr/0122-provision-boundary-deterministic-plus-decision-model-residue.md +60 -0
  144. rag_wright-0.1.0/docs/adr/README.md +181 -0
  145. rag_wright-0.1.0/docs/api/README.md +123 -0
  146. rag_wright-0.1.0/docs/architecture/foundations-and-adding-a-domain.md +115 -0
  147. rag_wright-0.1.0/docs/architecture.md +103 -0
  148. rag_wright-0.1.0/docs/archive/README.md +32 -0
  149. rag_wright-0.1.0/docs/archive/design/deontic-applicability-routing.md +146 -0
  150. rag_wright-0.1.0/docs/archive/design/ingestion-neuro-symbolic-gaps.md +113 -0
  151. rag_wright-0.1.0/docs/archive/design/semantic-subject-segmentation.md +120 -0
  152. rag_wright-0.1.0/docs/archive/design/unify-subject-preprocessing.md +135 -0
  153. rag_wright-0.1.0/docs/archive/engine-issues/0037-closed-vocab-drops-verbatim-values-to-other.md +44 -0
  154. rag_wright-0.1.0/docs/archive/engine-issues/0038-a-clause-node-is-now-created-per-sentence-so-98-percent-of-spans-become-clauses.md +109 -0
  155. rag_wright-0.1.0/docs/archive/engine-issues/0039-the-provision-detector-reads-text-but-docling-puts-the-section-number-in-marker.md +117 -0
  156. rag_wright-0.1.0/docs/archive/engine-issues/0040-covered-subject-is-asked-of-every-provision-so-verbatim-retention-fills-it-with-noise.md +114 -0
  157. rag_wright-0.1.0/docs/archive/engine-issues/0041-VERIFICATION-dense-floor.md +82 -0
  158. rag_wright-0.1.0/docs/archive/engine-issues/0041-leg-b-discards-the-best-dense-matches-on-some-queries.md +106 -0
  159. rag_wright-0.1.0/docs/archive/engine-issues/0042-usage-is-captured-per-call-but-never-returned-so-cost-needs-a-langfuse-round-trip.md +85 -0
  160. rag_wright-0.1.0/docs/archive/engine-issues/0043-a-requirement-carries-no-page-provenance-so-a-finding-cannot-point-into-its-policy.md +67 -0
  161. rag_wright-0.1.0/docs/archive/engine-issues/0044-document-signals-scaffolding-is-shown-to-the-user-as-a-quote-from-their-document.md +72 -0
  162. rag_wright-0.1.0/docs/archive/engine-issues/0045-a-fixed-money-cap-records-no-cap-quantum-so-the-amount-is-only-prose.md +82 -0
  163. rag_wright-0.1.0/docs/archive/engine-issues/0046-appending-a-natural-follow-up-to-a-question-drops-the-clause-type.md +72 -0
  164. rag_wright-0.1.0/docs/archive/engine-issues/0047-an-exact-deepagents-pin-transitively-pins-every-consumer.md +111 -0
  165. rag_wright-0.1.0/docs/archive/engine-issues/0048-the-pinned-endpoints-latency-tail-makes-agent-runs-undebuggable-from-either-side.md +89 -0
  166. rag_wright-0.1.0/docs/archive/handoffs/2026-07-20_graphwright_capability_interface_reply.md +136 -0
  167. rag_wright-0.1.0/docs/archive/handoffs/2026-07-20b_graphwright_interface_confirmation.md +101 -0
  168. rag_wright-0.1.0/docs/archive/handoffs/2026-07-20c_graphwright_ingestion_interfaces.md +101 -0
  169. rag_wright-0.1.0/docs/archive/handoffs/2026-08-30_subject_compliance_final_rulewright.md +111 -0
  170. rag_wright-0.1.0/docs/archive/handoffs/2026-09-01_issue-0013-recall-first-actor-gate_rulewright.md +41 -0
  171. rag_wright-0.1.0/docs/archive/handoffs/2026-09-01_ontology-source-of-truth_rulewright.md +80 -0
  172. rag_wright-0.1.0/docs/archive/handoffs/2026-09-03_issue-0014-table-retrievability_rulewright.md +58 -0
  173. rag_wright-0.1.0/docs/archive/handoffs/2026-09-04_bulk-ingestion-wall_rulewright.md +63 -0
  174. rag_wright-0.1.0/docs/archive/handoffs/2026-09-04_issue-0016-thread-safe-embedder_rulewright.md +36 -0
  175. rag_wright-0.1.0/docs/archive/handoffs/2026-09-05_observability-langfuse_rulewright.md +78 -0
  176. rag_wright-0.1.0/docs/archive/handoffs/2026-09-05_tagparse-ingestion-and-granite-4.2_rulewright.md +130 -0
  177. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_affiliate-of-extraction_rulewright.md +41 -0
  178. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_issue-0028-partyto-retired_rulewright.md +40 -0
  179. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_issue-0029-idempotent-write-graph_rulewright.md +38 -0
  180. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_issue-0030-entities-by-name_rulewright.md +42 -0
  181. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_observability-and-retrieval-floor_rulewright.md +26 -0
  182. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_relevance-verdict_rulewright.md +53 -0
  183. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_rename-run-ad-compliance-check_rulewright.md +30 -0
  184. rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_structured-output-cost-tracing_rulewright.md +44 -0
  185. rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_ingest-models-caller-configurable_rulewright.md +49 -0
  186. rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0031-followup-failclosed-and-nodrift_rulewright.md +28 -0
  187. rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0031-followup-validate-ingested-set_rulewright.md +31 -0
  188. rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0031-workspace-document-scope_rulewright.md +51 -0
  189. rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0032-span-page-provenance_rulewright.md +38 -0
  190. rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0033-carveout-normalization_rulewright.md +35 -0
  191. rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0034-invoke-time-document-scope_rulewright.md +35 -0
  192. rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_issue-0035-followup-derived-guard_rulewright.md +29 -0
  193. rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_issue-0035-mcp-no-model-tenant_rulewright.md +44 -0
  194. rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_profile-based-model-routing_rulewright.md +47 -0
  195. rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_qwen-default-modal-or_rulewright.md +53 -0
  196. rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_qwen-everywhere-config_rulewright.md +32 -0
  197. rag_wright-0.1.0/docs/archive/handoffs/2026-09-12_clause-granularity_0036-0039_rulewright.md +77 -0
  198. rag_wright-0.1.0/docs/archive/handoffs/2026-09-12_closed-vocab_0037-0040_rulewright.md +46 -0
  199. rag_wright-0.1.0/docs/archive/handoffs/2026-09-12_dense-floor-protection_0041_rulewright.md +31 -0
  200. rag_wright-0.1.0/docs/archive/handoffs/2026-09-13_in-band-usage_0042_rulewright.md +44 -0
  201. rag_wright-0.1.0/docs/archive/handoffs/2026-09-14_document-signals-off-citation_0044_rulewright.md +26 -0
  202. rag_wright-0.1.0/docs/archive/handoffs/2026-09-14_requirement-page-provenance_0043_rulewright.md +35 -0
  203. rag_wright-0.1.0/docs/archive/handoffs/2026-09-16_qwen3-27b-single-a100-serving-profile_rulewright.md +38 -0
  204. rag_wright-0.1.0/docs/archive/handoffs/2026-09-18_fp8-accuracy-eval-configB_rulewright.md +44 -0
  205. rag_wright-0.1.0/docs/archive/handoffs/2026-09-18_fp8-kv-16k-single-a100_rulewright.md +41 -0
  206. rag_wright-0.1.0/docs/archive/handoffs/2026-09-20_deepagents-pin-relaxed_0047_rulewright.md +22 -0
  207. rag_wright-0.1.0/docs/archive/handoffs/2026-09-20_llm-span-instrumentation_0048_rulewright.md +31 -0
  208. rag_wright-0.1.0/docs/archive/handoffs/2026-09-21_0048-part2-modal-answers_from_rulewright.md +131 -0
  209. rag_wright-0.1.0/docs/archive/handoffs/2026-09-21_0048-part2-modal-scoping-questions_rulewright.md +34 -0
  210. rag_wright-0.1.0/docs/archive/misc/.gitkeep +0 -0
  211. rag_wright-0.1.0/docs/archive/plans/Corpus_Acquisition.md +11 -0
  212. rag_wright-0.1.0/docs/archive/plans/async-migration.md +72 -0
  213. rag_wright-0.1.0/docs/archive/plans/demo_plan.md +103 -0
  214. rag_wright-0.1.0/docs/archive/plans/gp1b_docling_graph_plan.md +121 -0
  215. rag_wright-0.1.0/docs/archive/plans/unified_contract_kg_ontology_bridge.md +260 -0
  216. rag_wright-0.1.0/docs/archive/plans/unified_contract_kg_plan.md +104 -0
  217. rag_wright-0.1.0/docs/archive/results/2026-07-24-t58-property-graph-population.md +43 -0
  218. rag_wright-0.1.0/docs/archive/results/2026-07-25-t58b-full-pipeline-rerank.md +76 -0
  219. rag_wright-0.1.0/docs/archive/results/2026-07-25-t58b-topk-ordering-levers.md +80 -0
  220. rag_wright-0.1.0/docs/concepts.md +110 -0
  221. rag_wright-0.1.0/docs/configuration.md +91 -0
  222. rag_wright-0.1.0/docs/contract_pipeline_explainer.md +169 -0
  223. rag_wright-0.1.0/docs/corpus_ingest_recipe.md +56 -0
  224. rag_wright-0.1.0/docs/domain-adaptation/README.md +71 -0
  225. rag_wright-0.1.0/docs/domain-adaptation/_engine-gaps.md +44 -0
  226. rag_wright-0.1.0/docs/domain-adaptation/authoring-capabilities.md +85 -0
  227. rag_wright-0.1.0/docs/domain-adaptation/classification-and-decision-models.md +69 -0
  228. rag_wright-0.1.0/docs/domain-adaptation/entity-resolution.md +54 -0
  229. rag_wright-0.1.0/docs/domain-adaptation/kg-construction.md +66 -0
  230. rag_wright-0.1.0/docs/domain-adaptation/ontology-authoring.md +83 -0
  231. rag_wright-0.1.0/docs/eval/classifier_model_ab.md +61 -0
  232. rag_wright-0.1.0/docs/eval/compliance_demo.md +53 -0
  233. rag_wright-0.1.0/docs/eval/compliance_gate_cc7.md +68 -0
  234. rag_wright-0.1.0/docs/eval/compliance_rung2_cc7.md +156 -0
  235. rag_wright-0.1.0/docs/eval/cuad_highlighting_cu-d1.md +72 -0
  236. rag_wright-0.1.0/docs/eval/dg_model_ab.md +48 -0
  237. rag_wright-0.1.0/docs/eval/function_gate_recall.md +48 -0
  238. rag_wright-0.1.0/docs/eval/generation_robustness_b_vs_c.md +64 -0
  239. rag_wright-0.1.0/docs/eval/nl_to_type_cu-d2.md +90 -0
  240. rag_wright-0.1.0/docs/eval/ocr_benchmark.md +87 -0
  241. rag_wright-0.1.0/docs/eval/prod1_readiness.md +55 -0
  242. rag_wright-0.1.0/docs/eval/prod2_readiness.md +99 -0
  243. rag_wright-0.1.0/docs/eval/query_side_model_ab.md +63 -0
  244. rag_wright-0.1.0/docs/eval/silver_granite_vs_gemma4.md +47 -0
  245. rag_wright-0.1.0/docs/eval/silver_provider_routing.md +48 -0
  246. rag_wright-0.1.0/docs/eval/silver_selfhosted_gemma4_26b.md +53 -0
  247. rag_wright-0.1.0/docs/installation.md +93 -0
  248. rag_wright-0.1.0/docs/playbook.md +222 -0
  249. rag_wright-0.1.0/docs/product/capability_profiles.md +166 -0
  250. rag_wright-0.1.0/docs/product/contracts_product_roadmap.md +502 -0
  251. rag_wright-0.1.0/docs/product/engine-api-migration-handoff.md +278 -0
  252. rag_wright-0.1.0/docs/product/engine_async_api.md +150 -0
  253. rag_wright-0.1.0/docs/product/new-domain-build-sequence.md +10 -0
  254. rag_wright-0.1.0/docs/product/seam-adaptation-guide.md +64 -0
  255. rag_wright-0.1.0/docs/proposals/compliance-ingest-classifier-decomposition.md +127 -0
  256. rag_wright-0.1.0/docs/proposals/de-domaining-and-capability-runtime.md +213 -0
  257. rag_wright-0.1.0/docs/proposals/new-domain-developer-journey.md +263 -0
  258. rag_wright-0.1.0/docs/quickstart.md +106 -0
  259. rag_wright-0.1.0/docs/reference-pack.md +69 -0
  260. rag_wright-0.1.0/docs/specs/engine-platform/SPEC.md +81 -0
  261. rag_wright-0.1.0/docs/specs/engine-platform/TASKS.md +234 -0
  262. rag_wright-0.1.0/docs/specs/engine-prep/plan.md +525 -0
  263. rag_wright-0.1.0/docs/templates/product-starter/CLAUDE.md.template +193 -0
  264. rag_wright-0.1.0/docs/templates/product-starter/README.md +45 -0
  265. rag_wright-0.1.0/docs/templates/product-starter/playbook.md.template +174 -0
  266. rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-buckets-schema.md +82 -0
  267. rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-docs-extraction.json +2465 -0
  268. rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-vector-embeddings.md +696 -0
  269. rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-vector-search-tutorial.md +455 -0
  270. rag_wright-0.1.0/eval/__init__.py +8 -0
  271. rag_wright-0.1.0/eval/ablation.py +58 -0
  272. rag_wright-0.1.0/eval/acord.py +93 -0
  273. rag_wright-0.1.0/eval/acord_retrieval.py +75 -0
  274. rag_wright-0.1.0/eval/build_golden.py +51 -0
  275. rag_wright-0.1.0/eval/category_retrieval.py +118 -0
  276. rag_wright-0.1.0/eval/compliance_demo/policy/community_conduct_policy.md +24 -0
  277. rag_wright-0.1.0/eval/compliance_demo/subjects/post_borderline.txt +1 -0
  278. rag_wright-0.1.0/eval/compliance_demo/subjects/post_compliant.txt +1 -0
  279. rag_wright-0.1.0/eval/compliance_demo/subjects/post_violation.txt +1 -0
  280. rag_wright-0.1.0/eval/condensed_pipeline.py +284 -0
  281. rag_wright-0.1.0/eval/contractnli_judge.py +83 -0
  282. rag_wright-0.1.0/eval/cuad_highlight.py +218 -0
  283. rag_wright-0.1.0/eval/full_pipeline_rerank.py +159 -0
  284. rag_wright-0.1.0/eval/function_ceiling.py +75 -0
  285. rag_wright-0.1.0/eval/function_property_rerank.py +113 -0
  286. rag_wright-0.1.0/eval/function_rerank.py +107 -0
  287. rag_wright-0.1.0/eval/gate1_chunker_ab.py +150 -0
  288. rag_wright-0.1.0/eval/gate2_hybrid_rerank.py +301 -0
  289. rag_wright-0.1.0/eval/golden/relational/set.json +698 -0
  290. rag_wright-0.1.0/eval/golden.py +112 -0
  291. rag_wright-0.1.0/eval/ground_discriminator_rerank.py +193 -0
  292. rag_wright-0.1.0/eval/harness.py +83 -0
  293. rag_wright-0.1.0/eval/kg_primary.py +399 -0
  294. rag_wright-0.1.0/eval/kg_property_rerank.py +135 -0
  295. rag_wright-0.1.0/eval/listwise_rerank.py +254 -0
  296. rag_wright-0.1.0/eval/multihop.py +211 -0
  297. rag_wright-0.1.0/eval/nl_to_type.py +206 -0
  298. rag_wright-0.1.0/eval/okf_gold.py +214 -0
  299. rag_wright-0.1.0/eval/reachability.py +344 -0
  300. rag_wright-0.1.0/eval/relational_eval.py +91 -0
  301. rag_wright-0.1.0/eval/test_acord.py +60 -0
  302. rag_wright-0.1.0/eval/test_acord_retrieval.py +51 -0
  303. rag_wright-0.1.0/eval/test_category_retrieval.py +54 -0
  304. rag_wright-0.1.0/eval/test_harness.py +112 -0
  305. rag_wright-0.1.0/eval/test_multihop_set.py +164 -0
  306. rag_wright-0.1.0/eval/test_okf_gold.py +164 -0
  307. rag_wright-0.1.0/eval/test_reachability.py +134 -0
  308. rag_wright-0.1.0/examples/quickstart.py +94 -0
  309. rag_wright-0.1.0/plan.md +323 -0
  310. rag_wright-0.1.0/pyproject.toml +157 -0
  311. rag_wright-0.1.0/scripts/ab_model_overlap.py +110 -0
  312. rag_wright-0.1.0/scripts/acord_unify.py +223 -0
  313. rag_wright-0.1.0/scripts/acquire_acord.py +83 -0
  314. rag_wright-0.1.0/scripts/acquire_cuad.py +183 -0
  315. rag_wright-0.1.0/scripts/acquire_ecfr.py +72 -0
  316. rag_wright-0.1.0/scripts/acquire_edgar.py +105 -0
  317. rag_wright-0.1.0/scripts/acquire_ftc_255.py +64 -0
  318. rag_wright-0.1.0/scripts/acquire_prod1_corpus.py +78 -0
  319. rag_wright-0.1.0/scripts/audit_reclass_flips.py +127 -0
  320. rag_wright-0.1.0/scripts/author_leg_a_silver_key.py +151 -0
  321. rag_wright-0.1.0/scripts/backfill_affiliations.py +106 -0
  322. rag_wright-0.1.0/scripts/backfill_clause_span_id.py +95 -0
  323. rag_wright-0.1.0/scripts/backfill_edge_source_doc_id.py +69 -0
  324. rag_wright-0.1.0/scripts/backup_kg_to_gcs.sh +39 -0
  325. rag_wright-0.1.0/scripts/bootstrap_template_capture.py +72 -0
  326. rag_wright-0.1.0/scripts/bootstrap_typed_edges.py +54 -0
  327. rag_wright-0.1.0/scripts/build_api_docs.py +59 -0
  328. rag_wright-0.1.0/scripts/build_api_docs.sh +7 -0
  329. rag_wright-0.1.0/scripts/build_cuad_clause_cache.py +92 -0
  330. rag_wright-0.1.0/scripts/build_function_routing_map.py +50 -0
  331. rag_wright-0.1.0/scripts/cic1_assemble_v2.py +125 -0
  332. rag_wright-0.1.0/scripts/cic1_assemble_v3.py +126 -0
  333. rag_wright-0.1.0/scripts/cic1_generate_hard_negatives.py +100 -0
  334. rag_wright-0.1.0/scripts/cic1_generate_hard_positives.py +108 -0
  335. rag_wright-0.1.0/scripts/cic1_generate_training.py +154 -0
  336. rag_wright-0.1.0/scripts/cic1_hybrid.py +88 -0
  337. rag_wright-0.1.0/scripts/cic1_hybrid2.py +104 -0
  338. rag_wright-0.1.0/scripts/cic1_jev_actor.py +100 -0
  339. rag_wright-0.1.0/scripts/cic1_jev_claimtypes.py +104 -0
  340. rag_wright-0.1.0/scripts/cic1_jev_operative.py +80 -0
  341. rag_wright-0.1.0/scripts/cic1_label_spans.py +202 -0
  342. rag_wright-0.1.0/scripts/cic1_laya_prep.py +44 -0
  343. rag_wright-0.1.0/scripts/cic1_laya_train.py +81 -0
  344. rag_wright-0.1.0/scripts/cic1_prep_operative.py +83 -0
  345. rag_wright-0.1.0/scripts/cic1_relabel_rubric.py +107 -0
  346. rag_wright-0.1.0/scripts/cic1_relabel_rubric2.py +108 -0
  347. rag_wright-0.1.0/scripts/cic1_train_operative.py +175 -0
  348. rag_wright-0.1.0/scripts/compare_extraction_models.py +142 -0
  349. rag_wright-0.1.0/scripts/compliance_actor_gate_smoke.py +113 -0
  350. rag_wright-0.1.0/scripts/compliance_engine_smoke.py +65 -0
  351. rag_wright-0.1.0/scripts/compliance_policy_demo.py +69 -0
  352. rag_wright-0.1.0/scripts/curate_taxonomy_gaps.py +127 -0
  353. rag_wright-0.1.0/scripts/dg_model_ab.py +119 -0
  354. rag_wright-0.1.0/scripts/diagnose_acord_knn_expansion.py +106 -0
  355. rag_wright-0.1.0/scripts/diagnose_acord_knn_rerank.py +99 -0
  356. rag_wright-0.1.0/scripts/diagnose_acord_leg_pool.py +86 -0
  357. rag_wright-0.1.0/scripts/diagnose_acord_legs.py +112 -0
  358. rag_wright-0.1.0/scripts/diagnose_acord_pool.py +66 -0
  359. rag_wright-0.1.0/scripts/diagnose_acord_relstructure.py +106 -0
  360. rag_wright-0.1.0/scripts/diagnose_acord_retrieval.py +63 -0
  361. rag_wright-0.1.0/scripts/distill/eval_all.py +122 -0
  362. rag_wright-0.1.0/scripts/distill/eval_ce.py +107 -0
  363. rag_wright-0.1.0/scripts/distill/eval_listwise_b.py +152 -0
  364. rag_wright-0.1.0/scripts/distill/export_ce_dataset.py +69 -0
  365. rag_wright-0.1.0/scripts/distill/extract_features.py +85 -0
  366. rag_wright-0.1.0/scripts/distill/listwise_variants.py +213 -0
  367. rag_wright-0.1.0/scripts/distill/relational_features.py +81 -0
  368. rag_wright-0.1.0/scripts/distill/train_ce.py +74 -0
  369. rag_wright-0.1.0/scripts/distill/train_modal.py +121 -0
  370. rag_wright-0.1.0/scripts/enrich_edgar_candidates.py +114 -0
  371. rag_wright-0.1.0/scripts/eval_chunking_ab.py +176 -0
  372. rag_wright-0.1.0/scripts/eval_classifier_guided_vs_tagparse.py +156 -0
  373. rag_wright-0.1.0/scripts/eval_compliance_gold.py +91 -0
  374. rag_wright-0.1.0/scripts/extract_arcadedb_docs.py +74 -0
  375. rag_wright-0.1.0/scripts/generate_contract_python.py +30 -0
  376. rag_wright-0.1.0/scripts/git_post_commit_graphify.sh +25 -0
  377. rag_wright-0.1.0/scripts/granite_chunker_assess.py +124 -0
  378. rag_wright-0.1.0/scripts/graph_status.sh +35 -0
  379. rag_wright-0.1.0/scripts/ingest_acord.py +103 -0
  380. rag_wright-0.1.0/scripts/ingest_compliance_async_prod2.py +67 -0
  381. rag_wright-0.1.0/scripts/ingest_compliance_document_prod2.py +57 -0
  382. rag_wright-0.1.0/scripts/ingest_compliance_prod2.py +63 -0
  383. rag_wright-0.1.0/scripts/ingest_cuad.py +203 -0
  384. rag_wright-0.1.0/scripts/ingest_cuad_full.py +61 -0
  385. rag_wright-0.1.0/scripts/ingest_ftc_compliance.py +50 -0
  386. rag_wright-0.1.0/scripts/ingest_prod1.py +95 -0
  387. rag_wright-0.1.0/scripts/ingest_prod1_async.py +101 -0
  388. rag_wright-0.1.0/scripts/ingest_smoke.py +81 -0
  389. rag_wright-0.1.0/scripts/install_git_hooks.sh +13 -0
  390. rag_wright-0.1.0/scripts/label_new_functions.py +136 -0
  391. rag_wright-0.1.0/scripts/legb_function_gate_recall.py +121 -0
  392. rag_wright-0.1.0/scripts/make_compliance_fixtures.py +87 -0
  393. rag_wright-0.1.0/scripts/make_compliance_gold.py +159 -0
  394. rag_wright-0.1.0/scripts/mcp_compliance_agent_demo.py +78 -0
  395. rag_wright-0.1.0/scripts/mcp_intra_document_qa_smoke.py +65 -0
  396. rag_wright-0.1.0/scripts/mcp_query_legs_agent_demo.py +84 -0
  397. rag_wright-0.1.0/scripts/measure_generation_robustness.py +172 -0
  398. rag_wright-0.1.0/scripts/measure_silver.py +187 -0
  399. rag_wright-0.1.0/scripts/merge_docs_into_framework.py +43 -0
  400. rag_wright-0.1.0/scripts/migrate_entity_cik_to_canonical_id.py +47 -0
  401. rag_wright-0.1.0/scripts/migrate_silver_evidence_autotag.py +58 -0
  402. rag_wright-0.1.0/scripts/migrate_silver_evidence_remove_autotag.py +68 -0
  403. rag_wright-0.1.0/scripts/mine_scarce_functions.py +89 -0
  404. rag_wright-0.1.0/scripts/modal_arcadedb.py +80 -0
  405. rag_wright-0.1.0/scripts/modal_backfill.py +114 -0
  406. rag_wright-0.1.0/scripts/modal_exception_linking.py +109 -0
  407. rag_wright-0.1.0/scripts/modal_gemma4_vllm.py +105 -0
  408. rag_wright-0.1.0/scripts/modal_gemma4_vllm_snapshot.py +189 -0
  409. rag_wright-0.1.0/scripts/modal_granite_server.py +128 -0
  410. rag_wright-0.1.0/scripts/modal_granite_throughput.py +147 -0
  411. rag_wright-0.1.0/scripts/modal_granite_vllm_server.py +44 -0
  412. rag_wright-0.1.0/scripts/modal_query_app.py +103 -0
  413. rag_wright-0.1.0/scripts/modal_qwen3_27b_bench.py +234 -0
  414. rag_wright-0.1.0/scripts/modal_qwen3_27b_snapshot.py +158 -0
  415. rag_wright-0.1.0/scripts/modal_qwen3_vllm_server.py +96 -0
  416. rag_wright-0.1.0/scripts/modal_stack_a100.py +121 -0
  417. rag_wright-0.1.0/scripts/ocr_benchmark.py +268 -0
  418. rag_wright-0.1.0/scripts/ocr_preprocess.py +77 -0
  419. rag_wright-0.1.0/scripts/ontology_dimension_check.py +95 -0
  420. rag_wright-0.1.0/scripts/phase_a_leg_validate.py +120 -0
  421. rag_wright-0.1.0/scripts/populate_clause_kg.py +141 -0
  422. rag_wright-0.1.0/scripts/populate_entity_graph.py +107 -0
  423. rag_wright-0.1.0/scripts/populate_entity_graph_extracted.py +110 -0
  424. rag_wright-0.1.0/scripts/populate_property_store.py +202 -0
  425. rag_wright-0.1.0/scripts/prep_relational_verification.py +113 -0
  426. rag_wright-0.1.0/scripts/publish_manifests.py +30 -0
  427. rag_wright-0.1.0/scripts/reclassify_kg.py +322 -0
  428. rag_wright-0.1.0/scripts/refresh_framework_graph.sh +100 -0
  429. rag_wright-0.1.0/scripts/revert_reclass_flips.py +69 -0
  430. rag_wright-0.1.0/scripts/run_acord_retrieval.py +123 -0
  431. rag_wright-0.1.0/scripts/run_clause_exception_linking.py +38 -0
  432. rag_wright-0.1.0/scripts/semantic_judge_live_validate.py +131 -0
  433. rag_wright-0.1.0/scripts/semantic_judge_probe.py +84 -0
  434. rag_wright-0.1.0/scripts/snapshot_leg_a_eval.py +130 -0
  435. rag_wright-0.1.0/scripts/stack_correctness_validate.py +120 -0
  436. rag_wright-0.1.0/scripts/table_retrieval_smoke.py +118 -0
  437. rag_wright-0.1.0/scripts/train_function_classifier.py +95 -0
  438. rag_wright-0.1.0/scripts/train_legalbert_function.py +270 -0
  439. rag_wright-0.1.0/scripts/train_legalbert_modal.py +473 -0
  440. rag_wright-0.1.0/scripts/typed_rerank_validate.py +74 -0
  441. rag_wright-0.1.0/scripts/verify_wheel_install.sh +60 -0
  442. rag_wright-0.1.0/scripts/vllm_extraction_ab.py +111 -0
  443. rag_wright-0.1.0/scripts/vllm_kg_query_validate.py +83 -0
  444. rag_wright-0.1.0/scripts/vllm_raw_diagnostic.py +76 -0
  445. rag_wright-0.1.0/src/rag_wright/__init__.py +13 -0
  446. rag_wright-0.1.0/src/rag_wright/api/__init__.py +33 -0
  447. rag_wright-0.1.0/src/rag_wright/api/config.py +59 -0
  448. rag_wright-0.1.0/src/rag_wright/api/discover.py +70 -0
  449. rag_wright-0.1.0/src/rag_wright/api/documents.py +39 -0
  450. rag_wright-0.1.0/src/rag_wright/api/ids.py +31 -0
  451. rag_wright-0.1.0/src/rag_wright/api/invoke.py +99 -0
  452. rag_wright-0.1.0/src/rag_wright/api/kg.py +61 -0
  453. rag_wright-0.1.0/src/rag_wright/api/mcp.py +94 -0
  454. rag_wright-0.1.0/src/rag_wright/api/usage.py +30 -0
  455. rag_wright-0.1.0/src/rag_wright/api/workspace.py +85 -0
  456. rag_wright-0.1.0/src/rag_wright/capabilities/__init__.py +8 -0
  457. rag_wright-0.1.0/src/rag_wright/capabilities/answer_generator.py +427 -0
  458. rag_wright-0.1.0/src/rag_wright/capabilities/ard.py +286 -0
  459. rag_wright-0.1.0/src/rag_wright/capabilities/assertion_extraction.py +79 -0
  460. rag_wright-0.1.0/src/rag_wright/capabilities/chunk_read.py +58 -0
  461. rag_wright-0.1.0/src/rag_wright/capabilities/chunk_write.py +163 -0
  462. rag_wright-0.1.0/src/rag_wright/capabilities/claim_extraction.py +153 -0
  463. rag_wright-0.1.0/src/rag_wright/capabilities/clause_exception_linking.py +117 -0
  464. rag_wright-0.1.0/src/rag_wright/capabilities/compliance_judgment.py +322 -0
  465. rag_wright-0.1.0/src/rag_wright/capabilities/compliance_store.py +87 -0
  466. rag_wright-0.1.0/src/rag_wright/capabilities/contract_kg_serve.py +156 -0
  467. rag_wright-0.1.0/src/rag_wright/capabilities/contract_kg_store.py +251 -0
  468. rag_wright-0.1.0/src/rag_wright/capabilities/dg_extraction.py +585 -0
  469. rag_wright-0.1.0/src/rag_wright/capabilities/disambiguation.py +163 -0
  470. rag_wright-0.1.0/src/rag_wright/capabilities/document_parse.py +87 -0
  471. rag_wright-0.1.0/src/rag_wright/capabilities/document_scope.py +49 -0
  472. rag_wright-0.1.0/src/rag_wright/capabilities/embedding.py +164 -0
  473. rag_wright-0.1.0/src/rag_wright/capabilities/embedding_profiles.py +43 -0
  474. rag_wright-0.1.0/src/rag_wright/capabilities/entity_resolution.py +154 -0
  475. rag_wright-0.1.0/src/rag_wright/capabilities/fusion.py +64 -0
  476. rag_wright-0.1.0/src/rag_wright/capabilities/graph_extraction.py +243 -0
  477. rag_wright-0.1.0/src/rag_wright/capabilities/graph_query.py +73 -0
  478. rag_wright-0.1.0/src/rag_wright/capabilities/graph_storage.py +111 -0
  479. rag_wright-0.1.0/src/rag_wright/capabilities/highlight_serve.py +142 -0
  480. rag_wright-0.1.0/src/rag_wright/capabilities/hybrid_search.py +65 -0
  481. rag_wright-0.1.0/src/rag_wright/capabilities/invoke.py +31 -0
  482. rag_wright-0.1.0/src/rag_wright/capabilities/jev_decision.py +38 -0
  483. rag_wright-0.1.0/src/rag_wright/capabilities/manifests.py +872 -0
  484. rag_wright-0.1.0/src/rag_wright/capabilities/okf_navigate.py +456 -0
  485. rag_wright-0.1.0/src/rag_wright/capabilities/parsing.py +286 -0
  486. rag_wright-0.1.0/src/rag_wright/capabilities/property_boosted_retrieval.py +125 -0
  487. rag_wright-0.1.0/src/rag_wright/capabilities/query_function_classifier.py +94 -0
  488. rag_wright-0.1.0/src/rag_wright/capabilities/query_understanding.py +109 -0
  489. rag_wright-0.1.0/src/rag_wright/capabilities/registry.py +262 -0
  490. rag_wright-0.1.0/src/rag_wright/capabilities/remote_encoders.py +94 -0
  491. rag_wright-0.1.0/src/rag_wright/capabilities/requirement_extraction.py +247 -0
  492. rag_wright-0.1.0/src/rag_wright/capabilities/reranking.py +123 -0
  493. rag_wright-0.1.0/src/rag_wright/capabilities/retrieval_core.py +126 -0
  494. rag_wright-0.1.0/src/rag_wright/capabilities/rlm_chunking.py +808 -0
  495. rag_wright-0.1.0/src/rag_wright/capabilities/rlm_synthesis.py +316 -0
  496. rag_wright-0.1.0/src/rag_wright/capabilities/scan_quality.py +136 -0
  497. rag_wright-0.1.0/src/rag_wright/capabilities/span_relevance_judgment.py +191 -0
  498. rag_wright-0.1.0/src/rag_wright/capabilities/vision_to_text.py +85 -0
  499. rag_wright-0.1.0/src/rag_wright/capabilities/vlm_ocr.py +85 -0
  500. rag_wright-0.1.0/src/rag_wright/contracts/__init__.py +6 -0
  501. rag_wright-0.1.0/src/rag_wright/contracts/chunk.py +79 -0
  502. rag_wright-0.1.0/src/rag_wright/contracts/compliance.py +303 -0
  503. rag_wright-0.1.0/src/rag_wright/contracts/contract_meta.py +27 -0
  504. rag_wright-0.1.0/src/rag_wright/contracts/extraction.py +130 -0
  505. rag_wright-0.1.0/src/rag_wright/contracts/function.py +167 -0
  506. rag_wright-0.1.0/src/rag_wright/contracts/function_routing.py +91 -0
  507. rag_wright-0.1.0/src/rag_wright/contracts/highlight.py +74 -0
  508. rag_wright-0.1.0/src/rag_wright/contracts/identifiers.py +153 -0
  509. rag_wright-0.1.0/src/rag_wright/contracts/jurisdiction.py +96 -0
  510. rag_wright-0.1.0/src/rag_wright/contracts/ontology.py +142 -0
  511. rag_wright-0.1.0/src/rag_wright/contracts/property.py +201 -0
  512. rag_wright-0.1.0/src/rag_wright/contracts/provenance.py +78 -0
  513. rag_wright-0.1.0/src/rag_wright/contracts/query_intent.py +53 -0
  514. rag_wright-0.1.0/src/rag_wright/contracts/span.py +76 -0
  515. rag_wright-0.1.0/src/rag_wright/contracts/value_match.py +84 -0
  516. rag_wright-0.1.0/src/rag_wright/corpus/__init__.py +0 -0
  517. rag_wright-0.1.0/src/rag_wright/corpus/canonicalize.py +116 -0
  518. rag_wright-0.1.0/src/rag_wright/corpus/cuad.py +153 -0
  519. rag_wright-0.1.0/src/rag_wright/corpus/cuad_ingestion.py +72 -0
  520. rag_wright-0.1.0/src/rag_wright/corpus/document_parser.py +299 -0
  521. rag_wright-0.1.0/src/rag_wright/corpus/edgar.py +231 -0
  522. rag_wright-0.1.0/src/rag_wright/corpus/gcs_ingestion.py +120 -0
  523. rag_wright-0.1.0/src/rag_wright/corpus/http.py +110 -0
  524. rag_wright-0.1.0/src/rag_wright/corpus/selection.py +152 -0
  525. rag_wright-0.1.0/src/rag_wright/mcp/__init__.py +11 -0
  526. rag_wright-0.1.0/src/rag_wright/mcp/compliance_server.py +299 -0
  527. rag_wright-0.1.0/src/rag_wright/mcp/intra_document_qa_server.py +170 -0
  528. rag_wright-0.1.0/src/rag_wright/mcp/relational_qa_server.py +171 -0
  529. rag_wright-0.1.0/src/rag_wright/mcp/session_store.py +64 -0
  530. rag_wright-0.1.0/src/rag_wright/mcp/typed_property_retrieval_server.py +191 -0
  531. rag_wright-0.1.0/src/rag_wright/models/__init__.py +8 -0
  532. rag_wright-0.1.0/src/rag_wright/models/profiles.py +331 -0
  533. rag_wright-0.1.0/src/rag_wright/models/seam.py +497 -0
  534. rag_wright-0.1.0/src/rag_wright/models/tag_structured.py +285 -0
  535. rag_wright-0.1.0/src/rag_wright/models/tracing.py +179 -0
  536. rag_wright-0.1.0/src/rag_wright/models/usage.py +102 -0
  537. rag_wright-0.1.0/src/rag_wright/okf/__init__.py +11 -0
  538. rag_wright-0.1.0/src/rag_wright/okf/compile.py +292 -0
  539. rag_wright-0.1.0/src/rag_wright/okf/document.py +47 -0
  540. rag_wright-0.1.0/src/rag_wright/okf/enrich.py +176 -0
  541. rag_wright-0.1.0/src/rag_wright/okf/links.py +190 -0
  542. rag_wright-0.1.0/src/rag_wright/okf/lint.py +105 -0
  543. rag_wright-0.1.0/src/rag_wright/ontology/__init__.py +6 -0
  544. rag_wright-0.1.0/src/rag_wright/ontology/_generated_template_meta.py +60 -0
  545. rag_wright-0.1.0/src/rag_wright/ontology/_generated_vocab.py +52 -0
  546. rag_wright-0.1.0/src/rag_wright/ontology/clause_template.py +964 -0
  547. rag_wright-0.1.0/src/rag_wright/ontology/codegen.py +84 -0
  548. rag_wright-0.1.0/src/rag_wright/ontology/compliance_bridge.ttl +186 -0
  549. rag_wright-0.1.0/src/rag_wright/ontology/contract_bridge.ttl +2685 -0
  550. rag_wright-0.1.0/src/rag_wright/ontology/contract_taxonomy.py +24 -0
  551. rag_wright-0.1.0/src/rag_wright/ontology/derive.py +58 -0
  552. rag_wright-0.1.0/src/rag_wright/ontology/loader.py +435 -0
  553. rag_wright-0.1.0/src/rag_wright/ontology/packs/ftc_16cfr255.ttl +29 -0
  554. rag_wright-0.1.0/src/rag_wright/ontology/registry.py +87 -0
  555. rag_wright-0.1.0/src/rag_wright/ontology/template_introspect.py +100 -0
  556. rag_wright-0.1.0/src/rag_wright/py.typed +0 -0
  557. rag_wright-0.1.0/src/rag_wright/reference/__init__.py +2 -0
  558. rag_wright-0.1.0/src/rag_wright/reference/compliance.py +41 -0
  559. rag_wright-0.1.0/src/rag_wright/reference/contract_seam.py +123 -0
  560. rag_wright-0.1.0/src/rag_wright/skills/__init__.py +7 -0
  561. rag_wright-0.1.0/src/rag_wright/skills/claim_extraction/SKILL.md +47 -0
  562. rag_wright-0.1.0/src/rag_wright/skills/claim_extraction/__init__.py +1 -0
  563. rag_wright-0.1.0/src/rag_wright/skills/claim_extraction/template.py +50 -0
  564. rag_wright-0.1.0/src/rag_wright/skills/compliance_judgment/SKILL.md +59 -0
  565. rag_wright-0.1.0/src/rag_wright/skills/corpus_ingest/SKILL.md +106 -0
  566. rag_wright-0.1.0/src/rag_wright/skills/extraction_semantic_judge/SKILL.md +51 -0
  567. rag_wright-0.1.0/src/rag_wright/skills/extraction_semantic_judge/__init__.py +1 -0
  568. rag_wright-0.1.0/src/rag_wright/skills/generation/SKILL.md +64 -0
  569. rag_wright-0.1.0/src/rag_wright/skills/generation/__init__.py +1 -0
  570. rag_wright-0.1.0/src/rag_wright/skills/generic_compliance_judgment/SKILL.md +58 -0
  571. rag_wright-0.1.0/src/rag_wright/skills/okf_navigate/SKILL.md +137 -0
  572. rag_wright-0.1.0/src/rag_wright/skills/requirement_extraction/SKILL.md +47 -0
  573. rag_wright-0.1.0/src/rag_wright/skills/requirement_extraction/__init__.py +1 -0
  574. rag_wright-0.1.0/src/rag_wright/skills/requirement_extraction/template.py +50 -0
  575. rag_wright-0.1.0/src/rag_wright/skills/rlm/SKILL.md +186 -0
  576. rag_wright-0.1.0/src/rag_wright/skills/rlm/__init__.py +31 -0
  577. rag_wright-0.1.0/src/rag_wright/skills/rlm/agent.py +292 -0
  578. rag_wright-0.1.0/src/rag_wright/skills/span_relevance_judgment/SKILL.md +67 -0
  579. rag_wright-0.1.0/src/rag_wright/skills/vision_to_text/SKILL.md +36 -0
  580. rag_wright-0.1.0/src/rag_wright/skills/vision_to_text/__init__.py +1 -0
  581. rag_wright-0.1.0/src/rag_wright/spans/__init__.py +1 -0
  582. rag_wright-0.1.0/src/rag_wright/spans/boundary.py +78 -0
  583. rag_wright-0.1.0/src/rag_wright/spans/clause_function_classifier.py +490 -0
  584. rag_wright-0.1.0/src/rag_wright/spans/clause_kg_extractor.py +337 -0
  585. rag_wright-0.1.0/src/rag_wright/spans/cuad_labels.py +81 -0
  586. rag_wright-0.1.0/src/rag_wright/spans/dim_classifier.py +158 -0
  587. rag_wright-0.1.0/src/rag_wright/spans/dim_fleet.json +411 -0
  588. rag_wright-0.1.0/src/rag_wright/spans/function_classifier.py +77 -0
  589. rag_wright-0.1.0/src/rag_wright/spans/function_families.py +62 -0
  590. rag_wright-0.1.0/src/rag_wright/spans/hybrid_classifier.py +103 -0
  591. rag_wright-0.1.0/src/rag_wright/spans/legalbert_classifier.py +83 -0
  592. rag_wright-0.1.0/src/rag_wright/spans/model_capabilities.py +107 -0
  593. rag_wright-0.1.0/src/rag_wright/spans/new_function_labels.py +111 -0
  594. rag_wright-0.1.0/src/rag_wright/spans/page_map.py +68 -0
  595. rag_wright-0.1.0/src/rag_wright/spans/property_extractor.py +365 -0
  596. rag_wright-0.1.0/src/rag_wright/spans/property_grounding.py +182 -0
  597. rag_wright-0.1.0/src/rag_wright/spans/reclassify.py +77 -0
  598. rag_wright-0.1.0/src/rag_wright/spans/scarce_function_labels.py +105 -0
  599. rag_wright-0.1.0/src/rag_wright/spans/segment.py +341 -0
  600. rag_wright-0.1.0/src/rag_wright/spans/semantic_judge.py +197 -0
  601. rag_wright-0.1.0/src/rag_wright/spans/symbolic_validation.py +131 -0
  602. rag_wright-0.1.0/src/rag_wright/spans/tag_clause_extractor.py +182 -0
  603. rag_wright-0.1.0/src/rag_wright/store/__init__.py +6 -0
  604. rag_wright-0.1.0/src/rag_wright/store/arcadedb.py +1135 -0
  605. rag_wright-0.1.0/src/rag_wright/store/chunk_text.py +66 -0
  606. rag_wright-0.1.0/src/rag_wright/store/seam.py +213 -0
  607. rag_wright-0.1.0/src/rag_wright/subgraphs/__init__.py +0 -0
  608. rag_wright-0.1.0/src/rag_wright/subgraphs/async_ingestion.py +204 -0
  609. rag_wright-0.1.0/src/rag_wright/subgraphs/compliance_check.py +1042 -0
  610. rag_wright-0.1.0/src/rag_wright/subgraphs/compliance_ingestion.py +306 -0
  611. rag_wright-0.1.0/src/rag_wright/subgraphs/contract_ingestion_pipeline.py +999 -0
  612. rag_wright-0.1.0/src/rag_wright/subgraphs/graph_extraction.py +102 -0
  613. rag_wright-0.1.0/src/rag_wright/subgraphs/intra_document_qa.py +328 -0
  614. rag_wright-0.1.0/src/rag_wright/subgraphs/observability.py +140 -0
  615. rag_wright-0.1.0/src/rag_wright/subgraphs/query_constraint_extraction.py +73 -0
  616. rag_wright-0.1.0/src/rag_wright/subgraphs/relational_qa.py +165 -0
  617. rag_wright-0.1.0/src/rag_wright/subgraphs/requirement_extraction.py +137 -0
  618. rag_wright-0.1.0/src/rag_wright/subgraphs/scaffold.py +65 -0
  619. rag_wright-0.1.0/src/rag_wright/subgraphs/semantic_chunking.py +183 -0
  620. rag_wright-0.1.0/src/rag_wright/subgraphs/typed_clause_extraction.py +172 -0
  621. rag_wright-0.1.0/src/rag_wright/subgraphs/typed_property_retrieval.py +278 -0
  622. rag_wright-0.1.0/src/rag_wright/util/__init__.py +1 -0
  623. rag_wright-0.1.0/src/rag_wright/util/concurrent.py +153 -0
  624. rag_wright-0.1.0/src/rag_wright/util/spacy_model.py +45 -0
  625. rag_wright-0.1.0/tasks.md +5053 -0
  626. rag_wright-0.1.0/tests/__init__.py +0 -0
  627. rag_wright-0.1.0/tests/api/__init__.py +0 -0
  628. rag_wright-0.1.0/tests/api/test_capability_reexports.py +30 -0
  629. rag_wright-0.1.0/tests/api/test_discover.py +92 -0
  630. rag_wright-0.1.0/tests/api/test_documents.py +66 -0
  631. rag_wright-0.1.0/tests/api/test_e2e.py +94 -0
  632. rag_wright-0.1.0/tests/api/test_invoke.py +311 -0
  633. rag_wright-0.1.0/tests/api/test_kg.py +90 -0
  634. rag_wright-0.1.0/tests/api/test_mcp.py +108 -0
  635. rag_wright-0.1.0/tests/api/test_options.py +84 -0
  636. rag_wright-0.1.0/tests/api/test_usage.py +58 -0
  637. rag_wright-0.1.0/tests/api/test_workspace.py +84 -0
  638. rag_wright-0.1.0/tests/arch/test_import_contracts.py +32 -0
  639. rag_wright-0.1.0/tests/async_helpers.py +65 -0
  640. rag_wright-0.1.0/tests/capabilities/__init__.py +0 -0
  641. rag_wright-0.1.0/tests/capabilities/_fixtures/rlm_probe_skill/SKILL.md +8 -0
  642. rag_wright-0.1.0/tests/capabilities/test_answer_generator.py +521 -0
  643. rag_wright-0.1.0/tests/capabilities/test_answer_generator_async.py +65 -0
  644. rag_wright-0.1.0/tests/capabilities/test_assertion_extraction.py +62 -0
  645. rag_wright-0.1.0/tests/capabilities/test_authoring_contract.py +58 -0
  646. rag_wright-0.1.0/tests/capabilities/test_chunk_read.py +56 -0
  647. rag_wright-0.1.0/tests/capabilities/test_chunk_write.py +249 -0
  648. rag_wright-0.1.0/tests/capabilities/test_claim_extraction.py +151 -0
  649. rag_wright-0.1.0/tests/capabilities/test_clause_exception_linking.py +103 -0
  650. rag_wright-0.1.0/tests/capabilities/test_compliance_judgment.py +356 -0
  651. rag_wright-0.1.0/tests/capabilities/test_compliance_store.py +123 -0
  652. rag_wright-0.1.0/tests/capabilities/test_contract_kg_serve.py +193 -0
  653. rag_wright-0.1.0/tests/capabilities/test_contract_kg_store.py +138 -0
  654. rag_wright-0.1.0/tests/capabilities/test_contract_kg_store_reads.py +132 -0
  655. rag_wright-0.1.0/tests/capabilities/test_contract_taxonomy_and_spans.py +102 -0
  656. rag_wright-0.1.0/tests/capabilities/test_dg_adapter.py +62 -0
  657. rag_wright-0.1.0/tests/capabilities/test_dg_async.py +94 -0
  658. rag_wright-0.1.0/tests/capabilities/test_dg_extraction.py +182 -0
  659. rag_wright-0.1.0/tests/capabilities/test_dg_model_seam.py +43 -0
  660. rag_wright-0.1.0/tests/capabilities/test_dg_private.py +55 -0
  661. rag_wright-0.1.0/tests/capabilities/test_disambiguation.py +171 -0
  662. rag_wright-0.1.0/tests/capabilities/test_document_scope.py +70 -0
  663. rag_wright-0.1.0/tests/capabilities/test_embedding.py +177 -0
  664. rag_wright-0.1.0/tests/capabilities/test_embedding_profiles.py +26 -0
  665. rag_wright-0.1.0/tests/capabilities/test_entity_resolution.py +189 -0
  666. rag_wright-0.1.0/tests/capabilities/test_fusion.py +63 -0
  667. rag_wright-0.1.0/tests/capabilities/test_graph_extraction.py +186 -0
  668. rag_wright-0.1.0/tests/capabilities/test_graph_query.py +127 -0
  669. rag_wright-0.1.0/tests/capabilities/test_graph_storage.py +159 -0
  670. rag_wright-0.1.0/tests/capabilities/test_highlight_serve.py +118 -0
  671. rag_wright-0.1.0/tests/capabilities/test_hybrid_search.py +183 -0
  672. rag_wright-0.1.0/tests/capabilities/test_jev_decision.py +119 -0
  673. rag_wright-0.1.0/tests/capabilities/test_manifests.py +288 -0
  674. rag_wright-0.1.0/tests/capabilities/test_okf_compile.py +184 -0
  675. rag_wright-0.1.0/tests/capabilities/test_okf_links.py +92 -0
  676. rag_wright-0.1.0/tests/capabilities/test_okf_navigate.py +129 -0
  677. rag_wright-0.1.0/tests/capabilities/test_parsing.py +148 -0
  678. rag_wright-0.1.0/tests/capabilities/test_parties_extraction.py +71 -0
  679. rag_wright-0.1.0/tests/capabilities/test_property_boosted_retrieval.py +128 -0
  680. rag_wright-0.1.0/tests/capabilities/test_query_function_classifier.py +43 -0
  681. rag_wright-0.1.0/tests/capabilities/test_query_understanding.py +117 -0
  682. rag_wright-0.1.0/tests/capabilities/test_registry.py +283 -0
  683. rag_wright-0.1.0/tests/capabilities/test_remote_encoders.py +68 -0
  684. rag_wright-0.1.0/tests/capabilities/test_repair_partition.py +158 -0
  685. rag_wright-0.1.0/tests/capabilities/test_requirement_extraction.py +262 -0
  686. rag_wright-0.1.0/tests/capabilities/test_reranking.py +162 -0
  687. rag_wright-0.1.0/tests/capabilities/test_retrieval_core.py +113 -0
  688. rag_wright-0.1.0/tests/capabilities/test_rlm_chunking.py +624 -0
  689. rag_wright-0.1.0/tests/capabilities/test_rlm_chunking_async.py +53 -0
  690. rag_wright-0.1.0/tests/capabilities/test_rlm_method.py +504 -0
  691. rag_wright-0.1.0/tests/capabilities/test_rlm_synthesis.py +317 -0
  692. rag_wright-0.1.0/tests/capabilities/test_runtime_registry.py +32 -0
  693. rag_wright-0.1.0/tests/capabilities/test_scan_quality.py +90 -0
  694. rag_wright-0.1.0/tests/capabilities/test_span_relevance_judgment.py +116 -0
  695. rag_wright-0.1.0/tests/capabilities/test_structural_boundary_discoverer.py +75 -0
  696. rag_wright-0.1.0/tests/capabilities/test_structural_model_fallback.py +105 -0
  697. rag_wright-0.1.0/tests/capabilities/test_tag_boundary_discoverer.py +78 -0
  698. rag_wright-0.1.0/tests/capabilities/test_tiered_ocr.py +241 -0
  699. rag_wright-0.1.0/tests/capabilities/test_vlm_ocr.py +54 -0
  700. rag_wright-0.1.0/tests/conftest.py +8 -0
  701. rag_wright-0.1.0/tests/contracts/__init__.py +0 -0
  702. rag_wright-0.1.0/tests/contracts/test_canonical_source_doc_id.py +53 -0
  703. rag_wright-0.1.0/tests/contracts/test_chunk_record.py +141 -0
  704. rag_wright-0.1.0/tests/contracts/test_compliance.py +164 -0
  705. rag_wright-0.1.0/tests/contracts/test_cuad_highlight_contracts.py +109 -0
  706. rag_wright-0.1.0/tests/contracts/test_extraction.py +171 -0
  707. rag_wright-0.1.0/tests/contracts/test_function.py +119 -0
  708. rag_wright-0.1.0/tests/contracts/test_function_routing.py +69 -0
  709. rag_wright-0.1.0/tests/contracts/test_identifiers.py +161 -0
  710. rag_wright-0.1.0/tests/contracts/test_jurisdiction.py +76 -0
  711. rag_wright-0.1.0/tests/contracts/test_ontology.py +167 -0
  712. rag_wright-0.1.0/tests/contracts/test_property.py +107 -0
  713. rag_wright-0.1.0/tests/contracts/test_provenance.py +119 -0
  714. rag_wright-0.1.0/tests/contracts/test_value_match.py +43 -0
  715. rag_wright-0.1.0/tests/corpus/__init__.py +0 -0
  716. rag_wright-0.1.0/tests/corpus/test_canonicalize.py +96 -0
  717. rag_wright-0.1.0/tests/corpus/test_cuad.py +45 -0
  718. rag_wright-0.1.0/tests/corpus/test_cuad_ingestion.py +38 -0
  719. rag_wright-0.1.0/tests/corpus/test_document_parser.py +312 -0
  720. rag_wright-0.1.0/tests/corpus/test_edgar.py +136 -0
  721. rag_wright-0.1.0/tests/corpus/test_gcs_ingestion.py +138 -0
  722. rag_wright-0.1.0/tests/corpus/test_http.py +121 -0
  723. rag_wright-0.1.0/tests/corpus/test_parse_wire2.py +50 -0
  724. rag_wright-0.1.0/tests/corpus/test_selection.py +133 -0
  725. rag_wright-0.1.0/tests/eval/test_contractnli_judge.py +17 -0
  726. rag_wright-0.1.0/tests/eval/test_cuad_highlight_metrics.py +43 -0
  727. rag_wright-0.1.0/tests/eval/test_kg_property_rerank.py +30 -0
  728. rag_wright-0.1.0/tests/eval/test_nl_to_type_metrics.py +27 -0
  729. rag_wright-0.1.0/tests/eval/test_relational_eval.py +72 -0
  730. rag_wright-0.1.0/tests/fixtures/leg_a_silver/README.md +41 -0
  731. rag_wright-0.1.0/tests/fixtures/leg_a_silver/evidence_snapshot.json +1408 -0
  732. rag_wright-0.1.0/tests/fixtures/table-bearing-contract.pdf +74 -0
  733. rag_wright-0.1.0/tests/foundation/__init__.py +0 -0
  734. rag_wright-0.1.0/tests/foundation/test_arcadedb_hybrid.py +106 -0
  735. rag_wright-0.1.0/tests/foundation/test_model_seam_structured.py +50 -0
  736. rag_wright-0.1.0/tests/journey/incidents_domain.py +84 -0
  737. rag_wright-0.1.0/tests/journey/incidents_pack.ttl +14 -0
  738. rag_wright-0.1.0/tests/journey/test_incidents_journey.py +81 -0
  739. rag_wright-0.1.0/tests/journey/test_pack_schema.py +79 -0
  740. rag_wright-0.1.0/tests/mcp/__init__.py +0 -0
  741. rag_wright-0.1.0/tests/mcp/test_compliance_server.py +200 -0
  742. rag_wright-0.1.0/tests/mcp/test_intra_document_qa_server.py +64 -0
  743. rag_wright-0.1.0/tests/mcp/test_no_model_supplied_tenant.py +116 -0
  744. rag_wright-0.1.0/tests/mcp/test_relational_qa_server.py +63 -0
  745. rag_wright-0.1.0/tests/mcp/test_typed_property_retrieval_server.py +74 -0
  746. rag_wright-0.1.0/tests/models/__init__.py +0 -0
  747. rag_wright-0.1.0/tests/models/test_async_infra.py +58 -0
  748. rag_wright-0.1.0/tests/models/test_profile_routing.py +77 -0
  749. rag_wright-0.1.0/tests/models/test_profile_seam.py +264 -0
  750. rag_wright-0.1.0/tests/models/test_seam_async.py +92 -0
  751. rag_wright-0.1.0/tests/models/test_seam_retry.py +74 -0
  752. rag_wright-0.1.0/tests/models/test_seam_stream.py +189 -0
  753. rag_wright-0.1.0/tests/models/test_serving_seam.py +121 -0
  754. rag_wright-0.1.0/tests/models/test_serving_wiring.py +82 -0
  755. rag_wright-0.1.0/tests/models/test_tag_structured.py +320 -0
  756. rag_wright-0.1.0/tests/models/test_tag_structured_async.py +47 -0
  757. rag_wright-0.1.0/tests/models/test_tracing.py +134 -0
  758. rag_wright-0.1.0/tests/models/test_usage.py +87 -0
  759. rag_wright-0.1.0/tests/ontology/__init__.py +0 -0
  760. rag_wright-0.1.0/tests/ontology/test_clause_template.py +171 -0
  761. rag_wright-0.1.0/tests/ontology/test_compliance_ontology_authoritative.py +92 -0
  762. rag_wright-0.1.0/tests/ontology/test_derivation.py +120 -0
  763. rag_wright-0.1.0/tests/ontology/test_entity_taxonomy.py +33 -0
  764. rag_wright-0.1.0/tests/ontology/test_generated_template_meta_in_sync.py +21 -0
  765. rag_wright-0.1.0/tests/ontology/test_generated_vocab_in_sync.py +23 -0
  766. rag_wright-0.1.0/tests/ontology/test_template_captured_in_ttl.py +22 -0
  767. rag_wright-0.1.0/tests/ontology/test_ttl_is_source_of_truth.py +33 -0
  768. rag_wright-0.1.0/tests/reference/test_compliance_reference.py +146 -0
  769. rag_wright-0.1.0/tests/reference/test_contract_seam.py +116 -0
  770. rag_wright-0.1.0/tests/scripts/test_eval_chunking_ab.py +63 -0
  771. rag_wright-0.1.0/tests/scripts/test_eval_classifier_ab.py +38 -0
  772. rag_wright-0.1.0/tests/spans/test_boundary.py +35 -0
  773. rag_wright-0.1.0/tests/spans/test_classifier_property_extractor.py +96 -0
  774. rag_wright-0.1.0/tests/spans/test_clause_classifier_tags_0005.py +54 -0
  775. rag_wright-0.1.0/tests/spans/test_clause_function_classifier.py +185 -0
  776. rag_wright-0.1.0/tests/spans/test_clause_function_classifier_async.py +78 -0
  777. rag_wright-0.1.0/tests/spans/test_clause_kg_extractor.py +234 -0
  778. rag_wright-0.1.0/tests/spans/test_clause_kg_extractor_async.py +70 -0
  779. rag_wright-0.1.0/tests/spans/test_dim_classifier.py +43 -0
  780. rag_wright-0.1.0/tests/spans/test_dim_fleet_live.py +147 -0
  781. rag_wright-0.1.0/tests/spans/test_function_classifier.py +111 -0
  782. rag_wright-0.1.0/tests/spans/test_hybrid_classifier.py +86 -0
  783. rag_wright-0.1.0/tests/spans/test_hybrid_property_extractor.py +150 -0
  784. rag_wright-0.1.0/tests/spans/test_legalbert_classifier.py +56 -0
  785. rag_wright-0.1.0/tests/spans/test_model_capabilities.py +51 -0
  786. rag_wright-0.1.0/tests/spans/test_new_function_labels.py +51 -0
  787. rag_wright-0.1.0/tests/spans/test_page_map.py +79 -0
  788. rag_wright-0.1.0/tests/spans/test_property_extractor.py +108 -0
  789. rag_wright-0.1.0/tests/spans/test_property_grounding.py +108 -0
  790. rag_wright-0.1.0/tests/spans/test_reclassify.py +59 -0
  791. rag_wright-0.1.0/tests/spans/test_scarce_function_labels.py +64 -0
  792. rag_wright-0.1.0/tests/spans/test_segment.py +281 -0
  793. rag_wright-0.1.0/tests/spans/test_semantic_judge.py +143 -0
  794. rag_wright-0.1.0/tests/spans/test_setfit_clause_adapter.py +70 -0
  795. rag_wright-0.1.0/tests/spans/test_span_offsets.py +54 -0
  796. rag_wright-0.1.0/tests/spans/test_stage_labels_0005.py +41 -0
  797. rag_wright-0.1.0/tests/spans/test_symbolic_validation.py +197 -0
  798. rag_wright-0.1.0/tests/spans/test_tag_clause_extractor.py +140 -0
  799. rag_wright-0.1.0/tests/store/__init__.py +0 -0
  800. rag_wright-0.1.0/tests/store/test_affiliation_backfill.py +52 -0
  801. rag_wright-0.1.0/tests/store/test_arcadedb_clause_kg.py +175 -0
  802. rag_wright-0.1.0/tests/store/test_arcadedb_contract.py +72 -0
  803. rag_wright-0.1.0/tests/store/test_arcadedb_property.py +107 -0
  804. rag_wright-0.1.0/tests/store/test_arcadedb_requirement_sources.py +97 -0
  805. rag_wright-0.1.0/tests/store/test_arcadedb_schema.py +224 -0
  806. rag_wright-0.1.0/tests/store/test_arcadedb_span.py +107 -0
  807. rag_wright-0.1.0/tests/store/test_chunk_text.py +89 -0
  808. rag_wright-0.1.0/tests/store/test_document_scope.py +150 -0
  809. rag_wright-0.1.0/tests/store/test_engine_domain_neutral.py +48 -0
  810. rag_wright-0.1.0/tests/store/test_entities_by_name.py +60 -0
  811. rag_wright-0.1.0/tests/store/test_kg_edges.py +124 -0
  812. rag_wright-0.1.0/tests/store/test_kg_read.py +107 -0
  813. rag_wright-0.1.0/tests/store/test_kg_write.py +99 -0
  814. rag_wright-0.1.0/tests/store/test_write_graph_idempotent.py +112 -0
  815. rag_wright-0.1.0/tests/subgraphs/__init__.py +0 -0
  816. rag_wright-0.1.0/tests/subgraphs/test_async_ingestion.py +256 -0
  817. rag_wright-0.1.0/tests/subgraphs/test_chunk7_structure_carry.py +73 -0
  818. rag_wright-0.1.0/tests/subgraphs/test_compliance_check.py +1509 -0
  819. rag_wright-0.1.0/tests/subgraphs/test_compliance_ingestion.py +333 -0
  820. rag_wright-0.1.0/tests/subgraphs/test_contract_ingestion_pipeline.py +300 -0
  821. rag_wright-0.1.0/tests/subgraphs/test_contract_ingestion_pipeline_async.py +142 -0
  822. rag_wright-0.1.0/tests/subgraphs/test_contract_ingestion_pipeline_graph_async.py +223 -0
  823. rag_wright-0.1.0/tests/subgraphs/test_graph_extraction.py +60 -0
  824. rag_wright-0.1.0/tests/subgraphs/test_ingest_knobs.py +32 -0
  825. rag_wright-0.1.0/tests/subgraphs/test_ingest_segment_classify.py +121 -0
  826. rag_wright-0.1.0/tests/subgraphs/test_intra_document_qa.py +395 -0
  827. rag_wright-0.1.0/tests/subgraphs/test_observability.py +41 -0
  828. rag_wright-0.1.0/tests/subgraphs/test_partial_entry_contract.py +81 -0
  829. rag_wright-0.1.0/tests/subgraphs/test_query_constraint_extraction.py +47 -0
  830. rag_wright-0.1.0/tests/subgraphs/test_relational_qa.py +128 -0
  831. rag_wright-0.1.0/tests/subgraphs/test_requirement_extraction.py +127 -0
  832. rag_wright-0.1.0/tests/subgraphs/test_scaffold.py +82 -0
  833. rag_wright-0.1.0/tests/subgraphs/test_semantic_chunking.py +106 -0
  834. rag_wright-0.1.0/tests/subgraphs/test_typed_clause_extraction.py +117 -0
  835. rag_wright-0.1.0/tests/subgraphs/test_typed_property_retrieval.py +244 -0
  836. rag_wright-0.1.0/tests/test_backfill_clause_span_id.py +43 -0
  837. rag_wright-0.1.0/tests/test_dg_model_ab.py +37 -0
  838. rag_wright-0.1.0/tests/test_populate_clause_kg.py +70 -0
  839. rag_wright-0.1.0/tests/test_populate_entity_graph.py +43 -0
  840. rag_wright-0.1.0/tests/util/__init__.py +0 -0
  841. rag_wright-0.1.0/tests/util/test_concurrent.py +67 -0
  842. rag_wright-0.1.0/tests/util/test_spacy_model.py +53 -0
  843. rag_wright-0.1.0/uv.lock +4848 -0
@@ -0,0 +1,14 @@
1
+ {
2
+ "hooks": {
3
+ "SessionStart": [
4
+ {
5
+ "hooks": [
6
+ {
7
+ "type": "command",
8
+ "command": "bash \"$CLAUDE_PROJECT_DIR/scripts/graph_status.sh\""
9
+ }
10
+ ]
11
+ }
12
+ ]
13
+ }
14
+ }
@@ -0,0 +1,106 @@
1
+ ---
2
+ name: authoring-a-capability
3
+ description: >-
4
+ How to author a new RAG_Wright engine capability of any kind (subgraph, function, model, agent_skill, mcp_tool)
5
+ so it is registered, ARD-discoverable, and invokable by name through the engine API. Use it whenever you add a
6
+ new capability or a new-domain product/graph needs one: it gives the shared registration + ARD + invocation
7
+ contract (the four surfaces + the definition of done), the per-kind implementation specifics, and the
8
+ conformance guardrail that keeps the catalog honest. Grounded against the real code; keep it in step with it.
9
+ ---
10
+
11
+ # Authoring a capability
12
+
13
+ A **capability** is a named, ARD-registered unit of engine behavior (FR-C). Every capability has a `kind`
14
+ (`ard.py::EntryKind`): `subgraph | function | model | agent_skill | mcp_tool` (`dagster_asset` is reserved). The
15
+ contract below is the SAME for every kind; only the implementation differs. Capabilities compose — a subgraph calls
16
+ functions/models; a product invokes a capability by name through `rag_wright.api`.
17
+
18
+ **Ground every call before writing it** (CLAUDE.md library rule). The authoritative sources this skill summarizes:
19
+ `capabilities/registry.py` (the canonical-slug set + `CapabilityRegistry.register`), `capabilities/manifests.py`
20
+ (`CapabilityManifest` + `_SPECS`/`MANIFEST_SPECS`), `capabilities/ard.py` (`EntryKind`, `MEDIA_TYPE_BY_KIND`,
21
+ `CALLABLE_KINDS`), `scripts/publish_manifests.py`, `api/invoke.py` + `capabilities/invoke.py::capability_impl` (the adapter-free impl_ref invoker + drift guard),
22
+ `api/mcp.py` (generic MCP exposure). The guardrail test is `tests/capabilities/test_authoring_contract.py`.
23
+
24
+ ## The four surfaces (the definition of done)
25
+
26
+ A finished capability touches these: 1 (implementation) is for EVERY kind; 2 (the invoke factory + `impl_ref`) is
27
+ for invokable kinds (`subgraph`/`model`); 3 (manifest + register) is for every discoverable kind; 4 (invocable +
28
+ MCP) follows automatically for invokable kinds; 4b (a bespoke MCP server) is optional.
29
+
30
+ 1. **Implementation** — the real code, in that kind's home (see per-kind below).
31
+ 2. **The invoke factory + `impl_ref`** (invokable kinds: subgraph/model) — write a co-located
32
+ `async def ainvoke(resources, inputs)` (subgraph) / `def <name>(resources, inputs)` (model) in the capability's
33
+ own module (model factories ignore `resources`; they build over the opaque `WorkspaceHandle` — `resources._store`,
34
+ `resources.model_id(role)`, never env). The manifest's `impl_ref="module:attr"` points to it. The invoker imports
35
+ it LAZILY and calls it — there is **NO central adapter dict** (EP-CORE-2). A plain function/agent_skill/mcp_tool
36
+ declares no `impl_ref`.
37
+ 3. **ARD manifest + register it** — a `CapabilityManifest(slug, kind, display_name, description,
38
+ representative_queries=(2-5…), tags=…, impl_ref=…)`. **The catalog ships EMPTY (EP-CORE-3):** call
39
+ `register_capability(manifest)` at runtime to add it (a product registers its own; the engine's reference pack is
40
+ opt-in via `load_reference_pack()`). `representative_queries` is the field ARD discovery ranks on — write real,
41
+ specific queries. For the ENGINE's reference pack, the manifest is committed in `manifests.py::_SPECS` and the
42
+ slug is in `CANONICAL_CAPABILITY_SLUGS`; a downstream product registers freely (its slugs need not be canonical).
43
+ Publish to `~/.air/registry` (what GraphWright's store loads) with `uv run python scripts/publish_manifests.py`.
44
+ Callable kinds get `ResponseBounds` (defaulted); `agent_skill` must NOT declare bounds (loaded, not called).
45
+ 4. **Invocable + MCP for free** — once registered with an `impl_ref`, the capability is callable as
46
+ `ainvoke_subgraph(slug, inputs, resources=ws)` / `invoke_model(slug, inputs, resources=ws)` — the invoker resolves
47
+ the `impl_ref` via `capabilities.invoke.capability_impl` (the drift guard asserts it resolves to a callable of the
48
+ declared kind) — AND exposable over MCP (surface 4b), with **zero engine edits**.
49
+
50
+ 4b. **MCP exposure** (optional) — any invokable capability is already an MCP tool with zero extra code via
51
+ `api/mcp.py::build_capability_mcp(slug, resources=ws)` (EP-RT-2). Write a bespoke `mcp/<slug>_server.py` only
52
+ when you want a CURATED, typed tool signature instead of the generic opaque-`inputs` surface.
53
+
54
+ ## Per-kind specifics
55
+
56
+ ### subgraph — a compiled LangGraph `StateGraph`
57
+ - **Home:** `subgraphs/<slug>.py`. A `production_<slug>(*, store, ...) -> CompiledGraph` builder: `g = StateGraph(_State)`,
58
+ add nodes/edges with `START`/`END`, `return g.compile()`. Nodes call functions/models (compose).
59
+ - **Invoke:** the co-located `async def ainvoke(resources, inputs)` factory (impl_ref target) builds + awaits the graph.
60
+ - Retry/dead-letter come from the graph scaffold, not the invoker. The output contract is the registered `contract`.
61
+
62
+ ### function — a plain, typed callable
63
+ - **Home:** `capabilities/<slug>.py`. A deterministic or model-backed callable with a Pydantic in/out contract.
64
+ - Invoker adapters for `function` are not wired yet (EP-API-2c); until then functions are composed inside
65
+ subgraphs, not invoked standalone through the API. Still register + manifest it.
66
+
67
+ ### model — a trained checkpoint behind a seam
68
+ - **Home:** `spans/` or `capabilities/` wrapping the checkpoint (e.g. the SetFit clause classifier; the 29-dim
69
+ property fleet via `spans/property_extractor.py`). Load the checkpoint ONCE and cache it (the fleet is heavy) —
70
+ see `spans/model_capabilities.py::_dim_registry`.
71
+ - Serve behind the existing seam/adapter so nothing upstream changes (to FIND where a model cap belongs, use the
72
+ `classifier-opportunity-analysis` skill; to BUILD/train + checkpoint + serve it, the `setfit` skill). The impl_ref
73
+ factory is `def <slug>(resources, inputs)` for a SYNC impl (CPU-bound local inference — a classifier/XGBoost
74
+ checkpoint; `resources` ignored) or `async def <slug>(resources, inputs)` for an ASYNC impl (I/O-bound — an
75
+ LLM-backed model cap calling OpenRouter / a local vLLM client). `invoke_model` runs a sync impl and REFUSES an
76
+ async one; `ainvoke_model` (EP-API-7) off-loads a sync impl with `asyncio.to_thread` and awaits an async impl
77
+ directly, with an optional `sem` for fan-out backpressure.
78
+
79
+ ### agent_skill — authored SKILL.md + the Deep Agents runtime
80
+ - **Home:** `skills/<slug>/SKILL.md` + the agent runtime (e.g. `skills/rlm/`). It is LOADED (progressive
81
+ disclosure), not called: no `ResponseBounds`. Declare `requires=(...)` for a closure over other skills and
82
+ `skill_runtime` for its intrinsic runtime.
83
+
84
+ ### mcp_tool — a capability exposed over MCP
85
+ - A distinct ARD identity (`<slug>_mcp`) for the same underlying capability exposed as a cross-agent MCP tool.
86
+ Prefer the generic `build_capability_mcp` (surface 4b); author a bespoke `mcp/<slug>_server.py` only for a
87
+ curated typed signature. Bind the store server-side (issue 0035) — the tool never takes a tenant/store argument.
88
+
89
+ ## Verify (the guardrail)
90
+
91
+ Run `uv run pytest tests/capabilities/test_authoring_contract.py tests/capabilities/test_manifests.py
92
+ tests/capabilities/test_registry.py` after authoring. It pins the contract this skill teaches: no manifest under a
93
+ non-canonical slug; the reserved-without-manifest set is a fixed allowlist (so adding a slug but forgetting its
94
+ manifest FAILS here); every manifest kind is a real ARD kind; every invokable cap's impl_ref resolves to a callable whose
95
+ manifest declares the matching kind. If you deliberately add a reserved/internal slug (no manifest), add it to
96
+ `_RESERVED_WITHOUT_MANIFEST` with a one-line reason.
97
+
98
+ ## Common mistakes
99
+
100
+ - Adding the slug but forgetting the manifest (slug becomes silently un-discoverable) — the guardrail catches it.
101
+ - An impl_ref factory that reaches env/globals instead of the `WorkspaceHandle` — breaks multi-workspace use; build
102
+ everything from `h`.
103
+ - A non-lazy import at the top of an adapter — inflates the light index / `import rag_wright.api`; import inside
104
+ the adapter body.
105
+ - Declaring `response_bounds` on an `agent_skill`, or omitting the output `contract` on registration.
106
+ - Inventing a kind. If a capability fits none of the five, flag it — do not force-fit.
@@ -0,0 +1,175 @@
1
+ ---
2
+ name: classifier-opportunity-analysis
3
+ description: >-
4
+ Structured guide for ANALYZING a domain's ingestion + retrieval pipeline to find where an LLM call can be
5
+ replaced by a deterministic rule, a trained classifier, or a routing decision. Use it BEFORE building or
6
+ refactoring a domain pack, when an LLM is doing per-unit work that multiplies over a document, or when
7
+ onboarding a new domain — it is the identification/decision step upstream of `setfit` (which BUILDS the
8
+ classifier) and `authoring-a-capability` (which REGISTERS it as a capability). It captures the recipe applied
9
+ twice (contracts, then compliance) so the next domain is mapped the same way instead of re-derived.
10
+ ---
11
+
12
+ # Finding classifier / routing opportunities in a pipeline
13
+
14
+ This is an **analysis** skill: its output is a decision — an *opportunity list / decomposition plan*, not code and
15
+ not a trained model. It answers "where in this domain's ingestion and retrieval does a classifier, a routing
16
+ decision, or a deterministic rule belong, and where must the LLM stay?" Build what it identifies with the `setfit`
17
+ skill, serve the teacher with `qwen-vllm-modal`, and register each result as a capability with
18
+ `authoring-a-capability`.
19
+
20
+ ## The core pattern (why this works, and why we have done it twice)
21
+
22
+ A per-unit "do everything" LLM extraction is almost never one decision. It is a **bundle of separable decisions**
23
+ wearing one prompt. Decomposed, most of the bundle is not LLM-shaped work:
24
+
25
+ - **boundaries** (where does a unit start / is this span worth extracting) are usually structural → deterministic;
26
+ - **closed-vocab tags** (the type, the role, the dimension values) are classification → a classifier, or a
27
+ deterministic cue-rule when the cues are enumerable;
28
+ - only the **genuinely open part** (numbers, free text, synthesis) needs an LLM, and then only **one residual call
29
+ per unit**.
30
+
31
+ Worked precedent in this engine:
32
+
33
+ | | Contracts (decomposed) | Compliance (the CIC arc) |
34
+ |---|---|---|
35
+ | Unit boundaries | deterministic (section numbering / headings) + a soft type classifier | deterministic sub-section split + a span-level operative-cue gate |
36
+ | Closed-vocab tags | the 21-dim / 29-dim classifier fleet | deontic cue-rule + actor / claim-type / applicability classifiers |
37
+ | Open / numeric field | 1 residual LLM call / provision (7 fields) | 1 residual LLM call / section (evidence standard) |
38
+ | Text of the record | verbatim span | verbatim span (was a paraphrase) |
39
+ | Structure / graph extraction | once per document (parties), keep the LLM | once per document, keep the LLM |
40
+
41
+ The pattern is domain-independent. What changes per domain is the vocabulary and the document structure — which is
42
+ exactly what the phases below make you look at.
43
+
44
+ ## Phase A — Map the pipeline as per-unit decisions
45
+
46
+ Enumerate every LLM call and, for each, its **unit** and how the unit COUNT scales:
47
+
48
+ - per **document** (parse, a party/graph-structure pass) — count ≈ corpus size; cheap per doc.
49
+ - per **section / chunk / segment / span** — count scales with **document length**. A 100-page document is
50
+ thousands of these. **This is where the cost lives and where decomposition pays.**
51
+ - per **query** / per **candidate** / per **(claim, requirement) pair** — query-time; count ≈ traffic, usually a
52
+ few per request (see Phase D).
53
+
54
+ Two things to separate immediately:
55
+
56
+ 1. **Structure / connection extraction** (entities, parties, graph edges — "who/what is in this document and how is
57
+ it connected") is a **once-per-document** pass. It is NOT the per-unit cost; leave the LLM there. Do not mistake
58
+ it for the thing to decompose. (In this engine that is the docling-graph pass; it runs once per contract and once
59
+ per regulation.)
60
+ 2. **Per-unit semantic tagging** (what IS this unit, what are its typed properties) is the multiplying cost. This is
61
+ the target.
62
+
63
+ ## Phase B — Classify each decision by its shape, then pick the mechanism
64
+
65
+ For every per-unit decision, name its shape. The shape dictates the mechanism, in this order of preference (cheapest
66
+ and most robust first):
67
+
68
+ 1. **Boundary / segmentation** — "does a new unit start here?", "is this span operative / extractable?" →
69
+ **deterministic structure first** (numbering, enumeration `(a)(b)`, headings, list markers). Add a small
70
+ **binary classifier** only for the residue where structure is ambiguous (the analog of a `is_extractable_span`
71
+ model). Rarely needs an LLM.
72
+ 2. **Closed-vocab with an enumerable cue list** — a value the ontology can map from a fixed set of trigger phrases
73
+ (e.g. a deontic type from "must / shall / may not") → a **deterministic cue-rule, NO ML**. Do this before
74
+ training anything; it is free and exact.
75
+ 3. **Single-label routing** — one of a closed set, no clean cue list → a **classifier**.
76
+ 4. **Multi-label soft-tagging** — several of a closed set, used as *guidance* not a gate → a **soft-tag classifier**
77
+ (top-k). The most forgiving shape; a modest-accuracy model is still useful because wrong extra tags are cheap.
78
+ 5. **Verbatim vs generated text** — if the record just needs the unit's text, extract the **verbatim span**
79
+ (deterministic) rather than a generated paraphrase. Drops a generative LLM step and is more faithful for
80
+ citation. Keep a paraphrase only if a human-readable restatement is a real requirement.
81
+ 6. **Open / numeric / free-text / synthesis** — no closed set → keep the LLM, but reduce it to **one residual call
82
+ per unit** carrying only the fields that genuinely need it.
83
+ 7. **Pair / entailment** — rerank a candidate against a query, or a judge verdict over a (subject, rule) pair → a
84
+ **cross-encoder / NLI classifier** is a candidate (a verdict over a closed label set IS a classification). Often
85
+ query-time; see Phase D.
86
+
87
+ ## Phase C — What to look for in the documents themselves
88
+
89
+ Signals that make decomposition **feasible** (push toward rules + classifiers):
90
+
91
+ - explicit **numbering / enumeration / heading** structure → deterministic boundaries;
92
+ - a **closed, ontology-authored vocabulary** for the typed fields → classifiers + cue-rules;
93
+ - **repeated template structure** across documents → stable features;
94
+ - **enumerable linguistic cues** (deontic verbs, defined terms, standard phrasings) → cue-rules.
95
+
96
+ Signals that **resist** it (keep the LLM): genuinely open / unbounded values, cross-document or multi-hop
97
+ reasoning, long-prose synthesis, values that depend on interpretation rather than surface form.
98
+
99
+ ## Phase D — Ingestion vs retrieval: where the win actually is
100
+
101
+ - **Ingestion** per-unit work on long documents multiplies into thousands of calls. This is the **biggest, do-first**
102
+ opportunity. The whole Phase B decomposition applies.
103
+ - **Retrieval / query time** is typically a **few calls per request** (classify the query, extract its constraints,
104
+ rerank, judge). These ARE classifier/routing shapes (closed-set query routing, a pair-classifier reranker or
105
+ judge), but the volume is low and you often want the LLM's rationale or synthesis. It is legitimate to **live with
106
+ the LLM at query time** (as this engine does for the contract and compliance judges) and still decompose
107
+ ingestion fully. State this as a deliberate choice per decision; do not reflexively de-LLM query time.
108
+
109
+ ## Phase E — Soft-tag vs hard-gate: decide the role before committing
110
+
111
+ The same classifier is safe or dangerous depending on how its output is used:
112
+
113
+ - a **soft tag** that only augments / hints → safe even at modest accuracy; wrong extra tags are cheap.
114
+ - a **hard gate** that drops, blocks, or routes irreversibly → needs high accuracy AND a safe fallback.
115
+
116
+ Decide the role first. Keep a **graceful-degrade path**: a classifier abstention or a persistent rule-miss should
117
+ fall back to the residual LLM (or to an explicit "ambiguous"), never to a silent wrong answer.
118
+
119
+ ## Phase F — Make it ttl-driven, and know what transfers across domains
120
+
121
+ Author the closed vocabulary and the cue lists in the **ontology (`.ttl`), never in Python** (the engine's
122
+ knowledge-in-the-ontology rule). Then:
123
+
124
+ - the **deterministic mechanism** (structure split + cue-rule + verbatim extraction) is domain-generic and
125
+ **transfers to any pack for free** — it reads whatever vocab/cues the pack authors;
126
+ - a **trained classifier is vocabulary-specific** — a new-vocabulary domain pack trains **its own**. "Reuse across
127
+ products" therefore means the *same mechanism + per-pack models*, not one model everywhere.
128
+ - A pack that only re-routes an existing vocabulary at query time (an override overlay) is NOT a new-vocabulary pack
129
+ and needs no new ingestion classifier.
130
+
131
+ ## The output: the opportunity list
132
+
133
+ Produce a decomposition plan, not prose:
134
+
135
+ 1. **Headline the single biggest per-unit LLM cost** (the multiplying ingestion pass).
136
+ 2. For **each decision** give: its unit, its shape (Phase B), the chosen mechanism (deterministic / cue-rule /
137
+ classifier / residual-LLM / verbatim), and its role (soft-tag vs gate).
138
+ 3. Separate an **ingestion bucket** (do first) from a **query-time bucket** (decide case by case; often live with
139
+ the LLM).
140
+ 4. Note the **ttl + per-pack** generality (what transfers, what each pack re-trains).
141
+ 5. Exclude anything that is not actually a per-unit cost (once-per-document structure extraction) and anything
142
+ already settled (a decision an existing cue-rule covers).
143
+
144
+ ## Hand-off (what to do with the opportunities)
145
+
146
+ - **Deterministic rule / cue-rule / span-split** → plain code inside the ingestion subgraph. NOT a capability; it is
147
+ mechanism, and the knowledge it reads lives in the `.ttl`.
148
+ - **System-1 decision model (NO training)** → for a closed-set decision (yes/no, choice, score), A/B a decision
149
+ model — **Jev** (managed, OpenRouter Decisions API, zero/few-shot, calibrated) or **Laya** (open, fine-tuned) —
150
+ BEFORE committing to a trained classifier. It often wins when data is scarce or label-ambiguous, or when you need
151
+ calibrated uncertainty to route/gate (measured: RAG_Wright CIC-1c — Jev zero-shot 0.92 vs a trained SetFit 0.82;
152
+ ADR-0119). See `setfit` Phase 0.5 (the decision-vs-train A/B) and the `laya` skill; wire it as a `jev_decision`-style
153
+ model capability with a `DecisionModelProfile`. **Evaluate this first; it may remove the need to train at all.**
154
+ - **Trained classifier** → build it with the **`setfit`** skill (framing, symmetric leakage-safe eval, per-class
155
+ floor, soft-tag/top-k, rare-class curation, checkpointing). Serve the teacher / bulk-labeler with **`qwen-vllm-modal`**.
156
+ Then register it as a capability with **`authoring-a-capability`**: a `kind="model"` capability with an
157
+ `impl_ref` factory `def <slug>(resources, inputs)` over a cached checkpoint, invoked by name through the engine
158
+ API (`invoke_model` / `ainvoke_model`) and the pipeline (`dispatch_model` / `adispatch_model`) — **routed THROUGH
159
+ the capability layer, never hand-constructed around it.** A model impl may be sync (a classifier / XGBoost — run
160
+ off-loop by `ainvoke_model`) or async (an LLM-backed cap — awaited by `ainvoke_model`).
161
+
162
+ ## Anti-patterns (from the real sessions — do not repeat)
163
+
164
+ - **Letting classifier training cost/time decide whether an opportunity exists.** Identify opportunities by the
165
+ *pattern* (per-unit closed-set decision on a long document); training ROI is a separate, later question.
166
+ - **Mistaking once-per-document structure extraction for the per-unit cost.** It is cheap; leave the LLM.
167
+ - **Claiming a judge/verdict step "can't classify."** A verdict over a closed label set (compliant / violation /
168
+ needs-review, relevant / not) IS a classification; it is a legitimate pair-classifier candidate (usually
169
+ query-time).
170
+ - **Hardcoding the vocabulary or cues in Python.** They belong in the `.ttl`; the mechanism reads them.
171
+ - **Training a classifier where an enumerable cue-rule already settles the decision** (the deontic-type case).
172
+ - **Forcing a hard gate where a soft tag suffices** — it imposes an accuracy bar you did not need.
173
+ - **Reflexively de-LLM'ing query time** — low volume + wanted rationale often make the LLM the right call there.
174
+ - **A router/cascade of specialists** — it multiplies errors (router acc × specialist acc). A flat classifier +
175
+ multi-tag usually beats it (see `setfit`).
@@ -0,0 +1,114 @@
1
+ ---
2
+ name: creating-evals
3
+ description: >-
4
+ Domain-agnostic recipe for creating EVALS for engine/product capabilities, eval-first (TDD): write the eval as
5
+ soon as a capability is DEFINED (its contract + acceptance criterion), before it is implemented. Use it when
6
+ starting a new domain (right after capabilities are defined), when adding or changing a capability, or when A/B-ing
7
+ alternatives (e.g. a trained classifier vs a System-1 decision model). Covers gold-set design + reliability,
8
+ per-capability-KIND metrics (extraction / classification / retrieval / graph / generation / judgment), the
9
+ gate-vs-diagnostic split, building the gold cheaply, an isolated executable harness, and optional Langfuse
10
+ Datasets/Experiments/Scores automation. An eval is the executable acceptance criterion; training data IS an eval.
11
+ ---
12
+
13
+ # Creating evals (eval-first / TDD)
14
+
15
+ An eval is the **executable acceptance criterion** for a capability. Write it **as soon as the capability is
16
+ DEFINED** — its contract (typed in/out) and its acceptance criterion exist — **before it is implemented**. This is
17
+ TDD at the capability level: the eval fails (red) on the unbuilt/weak capability, you implement to green, then you
18
+ can A/B alternatives and catch regressions forever. Corollary observed repeatedly in this engine: **the training
19
+ data you build for a classifier/decision model IS an eval** (a labeled gold set + a metric) — so building the eval
20
+ first also gives you the data design for free.
21
+
22
+ ## When to use
23
+ - **Starting a new domain**: the FIRST build step after capabilities are defined (build-sequence step 4.5) — write
24
+ each capability's eval before/while you implement it.
25
+ - Adding or changing a capability, or tuning a threshold/prompt/model.
26
+ - **A/B-ing alternatives** on one capability (a deterministic rule vs a trained classifier vs a System-1 decision
27
+ model vs an LLM) — the eval is the neutral judge; select on the metric.
28
+
29
+ ## 1. Design the gold set (the foundation — get this right FIRST)
30
+ - **Real, in-domain items + expected outputs/labels.** Small but REPRESENTATIVE; never the easy cases only.
31
+ - **Reproducible + pinned + gitignored.** The gold is a generated artifact: pin its source snapshot + the selected
32
+ ids so it rebuilds identically; gitignore the data, commit the BUILDER. (Pattern: `eval/build_golden.py`.)
33
+ - **Leakage-safe + balanced.** Split by the natural grouping unit (document / source / record), not by row.
34
+ For classification, a SYMMETRIC per-class test and a per-class FLOOR (not overall accuracy — it hides dead
35
+ classes). "k-shot" = k per class.
36
+ - **The gold is the CEILING — measure its RELIABILITY.** For subjective/ambiguous labels, get a second
37
+ independent labeling and report inter-annotator (or inter-pass) agreement; adjudicate the disagreements and
38
+ document the calls. A model cannot beat the gold's own consistency, and label ambiguity in the gold shows up as a
39
+ classifier ceiling you cannot train past (CIC-1c: a two-pass gold agreed at 0.905 and a trained SetFit capped
40
+ ~0.82 — diagnose the gold before blaming the model). If you have no human expert yet, a documented rubric + a
41
+ two-pass consensus is
42
+ the honest proxy — say so.
43
+
44
+ ## 2. Pick the metric by capability KIND, and split GATE vs DIAGNOSTIC
45
+ Always set ONE pass/fail **gate** (from the acceptance criterion) and report **diagnostics** alongside (never gate
46
+ on a diagnostic).
47
+ - **Extraction** (section → records, clauses, requirements): recall / precision / F1 of extracted items vs gold,
48
+ reported SEPARATELY (under- vs over-extraction are different failures). Watch over-extraction on non-operative
49
+ input (definitions) and under-extraction on long input.
50
+ - **Classification / typed decision** (incl. SetFit, Laya, **Jev**): per-class recall + the per-class **floor**;
51
+ symmetric eval; top-k recall for multi-label (reported against tags-per-item). Prefer a soft-tag/calibrated
52
+ metric when the output routes rather than hard-gates.
53
+ - **Retrieval**: recall@k (binary, relevant = grade ≥ a floor) as the GATE; nDCG@k (graded, exp gain) as a
54
+ DIAGNOSTIC (do NOT threshold nDCG). Isolate the retrieval legs (score corpus-ids before rehydration). Pattern:
55
+ `eval/acord_retrieval.py`.
56
+ - **Graph / relational**: recall over the answer SET (reachability/traversal), k large enough to cover the set;
57
+ node ids match gold by construction. Pattern: `eval/relational_eval.py`.
58
+ - **Generation / QA**: citation-recall / groundedness / correct-abstention; LLM-as-judge for free-text, but anchor
59
+ with deterministic checks (a cited span must exist). Generation is non-deterministic near the abstain boundary —
60
+ measure over several runs.
61
+ - **Judgment / compliance verdicts**: report PRECISION and RECALL of the actionable class (e.g. violation)
62
+ SEPARATELY (alert-fatigue vs missed), and break out by provenance (real vs constructed). Pattern:
63
+ `scripts/eval_compliance_gold.py`.
64
+
65
+ ## 3. Build the gold cheaply (without faking it)
66
+ - **Silver bootstrapping**: a higher-capability teacher (LLM, or a decision model) labels candidate items; CURATE
67
+ with reject-rules; mark silver, never conflate with gold. (Teacher labeling runs on the flat-GPU substrate per
68
+ `qwen-vllm-modal`, not ad-hoc paid calls.)
69
+ - **Real public datasets** where they exist (e.g. labeled corpora for the task) — add to TRAIN; keep the TEST
70
+ in-corpus + human-adjudicated so it stays an honest transfer test.
71
+ - **Generate hard cases** (near-boundary positives/negatives) to stress the exact confusions — TRAIN-ONLY; the
72
+ gold test stays real. (Full data playbook: the `setfit` skill Phase 1/4 + `classifier-opportunity-analysis`.)
73
+
74
+ ## 4. Make it an executable, isolated harness
75
+ - **Isolate the capability under test**: inject it (a seam / `extract_override` / an injected `retrieve`) so the
76
+ eval measures ONE capability, not the whole pipeline. Invoke production code THROUGH the capability layer
77
+ (`ainvoke_subgraph`/`ainvoke_model`), never a hand-built copy.
78
+ - **Score from a RESULT ARTIFACT (JSON), not stdout scraping** (scraping truncates and silently drops rows).
79
+ - **Parallelize** model/LLM calls (async + semaphore) — same cost, far less wall-clock; order-preserving so it
80
+ stays deterministic.
81
+ - **Env-selected** so the SAME harness runs local (dev) or on Modal (full corpus + GPU).
82
+ - **Pre-flight paid bulk** (>~50 paid calls): print the exact count + cost and wait (`warn-before-bulk` rule).
83
+ - Stream `X/N` progress + actively monitor any run over ~30s (never launch-and-forget).
84
+
85
+ ## 5. Langfuse — optional eval automation (we already use it for tracing)
86
+ Langfuse has a first-class eval stack we are NOT yet using (we use it only for spans/usage today): a **Dataset**
87
+ (items = input + optional expected output) → a **Task** (your capability) run over the dataset as an **Experiment
88
+ Run** → **Evaluators** (deterministic checks or LLM-as-judge) producing **Scores** (numeric/categorical/boolean),
89
+ all linked to traces. Reach for it when you want **tracked, re-runnable, dashboarded** evals + regression tracking
90
+ across capability versions; a local JSON harness is enough for a quick one-off gate. Keep the gold BUILDER + a
91
+ committed snapshot in the repo (reproducibility); push the items to the Langfuse dataset so runs and scores land
92
+ next to the traces you already collect. Ground the exact Dataset/Experiment API before wiring
93
+ (https://langfuse.com/docs/evaluation/concepts, https://langfuse.com/docs/datasets/overview); confirm the installed
94
+ SDK version's surface, not prose.
95
+
96
+ ## 6. The TDD loop
97
+ 1. Write the eval from the contract + a handful of real gold items → it fails (red) on the unbuilt/weak capability.
98
+ 2. Implement the minimum to pass the gate (green).
99
+ 3. A/B alternatives (rule / trained classifier / decision model / LLM) on the SAME gold; select on gate + diagnostics + cost + calibration.
100
+ 4. Keep the eval; it is the regression guard (the Beyonce rule — if you shipped it, it has an eval).
101
+
102
+ ## Anti-patterns (seen, do not repeat)
103
+ - **Overall accuracy** instead of a per-class floor — hides dead classes.
104
+ - **Testing on teacher/generated data as if it were gold** — measures mimicry, not accuracy; keep the test real + adjudicated.
105
+ - **Gating on a diagnostic** (e.g. nDCG) — report it, don't threshold it.
106
+ - **No reliability check on a subjective gold** — you can't read a number off a gold whose own labels disagree.
107
+ - **Eval not isolated** to the capability (measures the whole pipeline) — you can't attribute a regression.
108
+ - **Scraping stdout** for metrics — write + read a JSON artifact.
109
+ - **Faking a win by shrinking the sample / picking easy items** — real, representative, honest.
110
+
111
+ ## Hand-off / where this sits
112
+ Chain: `classifier-opportunity-analysis` (identify the decision) → **`creating-evals` (THIS — write the eval FIRST)**
113
+ → build it (a deterministic rule, or `setfit`/`laya`/Jev for a decision, or an LLM) → `authoring-a-capability`
114
+ (register). In the new-domain build sequence, this is the step right after capabilities are defined.
@@ -0,0 +1,119 @@
1
+ ---
2
+ name: laya
3
+ description: Recipe for fine-tuning and serving a Laya (ModernBERT-large, RL-trained typed-decision) classifier -- the OPEN-weight System-1 decision model -- to replace an LLM decision that a SetFit/encoder classifier PLATEAUS on. Also explains Jev, the MANAGED zero-shot sibling (OpenRouter Decisions API), and when to A/B Jev first (usually) vs fine-tune Laya (on-prem / no managed API). Use when confusable/relational/subjective closed-vocab values won't clear the bar with any encoder backbone. Covers uv install + env-isolation gotchas, the JSONL schema, per-value criteria, single-T4 fine-tune (RLCD), few-shot-in-context, balance, OOM knobs, Modal harness, CPU/MPS/GPU serving, and the jev_decision/DecisionModelProfile wiring (ADR-0119).
4
+ ---
5
+
6
+ # Laya typed-decision classifier recipe
7
+
8
+ Laya (https://github.com/NandhaKishorM/laya) is a **non-autoregressive "System-1 decision engine"**: a
9
+ ModernBERT-large encoder + an **RL-trained decision head** that makes a typed `choice` / `score` / `noul` in one
10
+ forward pass, with per-option **criteria** (written descriptions) and calibrated confidence. It is the **escalation
11
+ past SetFit**: use it when the decision — not the representation — is the bottleneck.
12
+
13
+ **Every number below is a DIRECTION from one project, never a target — re-measure on your own data.**
14
+
15
+ ## Laya vs Jev (the managed sibling) — pick the decision model first
16
+ Laya is the **open-weight** System-1 decision model; **Jev** (TypeSafe, https://openrouter.ai, `typesafe/jev-1.13`
17
+ via the OpenRouter **Decisions API**) is the **managed** one — same idea (typed `noul`/`choice`/`score` + calibrated
18
+ probabilities), but **strong ZERO/few-shot with no fine-tuning**, whereas Laya scores near-random zero-shot and MUST
19
+ be fine-tuned. Measured (RAG_Wright CIC-1c, ADR-0119): on the operative-rule gate **Jev zero-shot hit 0.92** where a
20
+ trained SetFit capped ~0.82 and fine-tuned Laya reached ~0.70–0.74. So: **for a closed-set decision, A/B Jev
21
+ (zero-shot, instant) first**; reach for Laya when a **managed API is unacceptable** (on-prem / data-residency) and
22
+ you can fine-tune. In RAG_Wright a decision model is wired as a `kind="model"` capability (`jev_decision`) behind a
23
+ `DecisionModelProfile` (model id / endpoint / thresholds in config, swappable Jev ↔ Laya) — see `setfit` Phase 0.5 +
24
+ `authoring-a-capability`. The decision questions/criteria live in the `.ttl` (ADR-0066), not the capability.
25
+
26
+ ## When to use Laya (vs SetFit)
27
+ - A closed-vocab value **plateaus below the bar with EVERY encoder backbone you try** (we ruled it out with
28
+ LegalBERT, bge-large, all-mpnet, AND ModernBERT-large as SetFit bodies — all stuck). That means the bottleneck is
29
+ the *decision* ("who bears the obligation", "which side is favored", "may either party terminate?"), not the
30
+ embedding. Laya's RL-against-proper-scoring-rules training targets exactly that.
31
+ - Measured proof it's a different mechanism: `termination_right` cleared 0.68 with Laya where four encoders maxed at
32
+ 0.60–0.64. It is NOT magic — it did NOT rescue minority-data-starved or purely-numeric dims (see Limits).
33
+ - Default to SetFit first (cheaper, simpler, CPU-native). Reach for Laya only for the residual confusable/relational
34
+ dims SetFit can't crack.
35
+
36
+ ## Install — uv only, and beware env bleed
37
+ - **uv, never pip.** Locally `uv add laya`; on Modal `.uv_pip_install("laya", extra_index_url="https://download.pytorch.org/whl/cu124")` for CUDA torch wheels.
38
+ - **The env-bleed gotcha (cost real time):** a bare `uv run --with laya` on a machine with a system Anaconda picked
39
+ up conda's numpy/scipy/sklearn and crashed (`numpy.core.multiarray failed to import`). FIX: force uv's own managed
40
+ Python with a cleared path — `PYTHONPATH= PYTHONNOUSERSITE=1 uv run --python 3.11 --with laya python …` — or run
41
+ inside a proper isolated uv project. Never let a base conda leak in.
42
+ - Deps are compatible with our stack: `transformers>=4.48`, `torch>=2.0`, `huggingface_hub`, Python ≥3.10.
43
+
44
+ ## Checkpoints — pick the ENGLISH base, not multilingual
45
+ - English base = `convaiinnovations/laya` **root** (`laya.load("convaiinnovations/laya")`, subfolder=None) —
46
+ ModernBERT-large. `subfolder="multilingual"` = mmBERT; `subfolder="typed-decisions"` = a fine-tuned variant.
47
+ - **The fine-tune script defaults `--model-dir` to the MULTILINGUAL subfolder** — you must pass the English root
48
+ explicitly, e.g. `snapshot_download("convaiinnovations/laya", allow_patterns=["*.json","*.safetensors","encoder/*","tokenizer/*"])` and point `--model-dir` there. A checkpoint dir = `{encoder/, tokenizer/, model.safetensors, rl_agent_config.json}`.
49
+ - Base checkpoints score ~chance zero-shot (~0.36); **all the value is in fine-tuning.** (Zero-shot on our clauses
50
+ was directionally right but ~0.5 confidence — expected.)
51
+
52
+ ## Data — the JSONL schema + criteria (the differentiator)
53
+ One JSONL line per case (`research/scripts/finetune_single_device.py` reads this exact shape):
54
+ ```
55
+ {"state": "<clause text>",
56
+ "questions": {"<dim>": {"type":"choice","instructions":"<the question>","criteria":{"<value>":"<description>", ...}}},
57
+ "gold": {"<dim>": {"probabilities": {"<value>": <p>, ...}, "label": "<value>"}}}
58
+ ```
59
+ - **`criteria` = per-value written descriptions = the semantic-guidance lever SetFit never had.** Author them from
60
+ the ontology (`.ttl`); if the ontology is thin (ours had vocab but no per-value defs), author from the value
61
+ semantics and CONFIRM with a human before training. Wording matters — it directly shapes what the model learns.
62
+ - **`gold` is a DISTRIBUTION (soft targets), not a hard label** — RLCD trains on the teacher's probability per option.
63
+ One-hot (`{gold:1.0, other:0.0}`) works and is the pragmatic start; true soft targets (teacher per-option probs)
64
+ are the design-intended enhancement.
65
+ - Reuse your existing **contract-disjoint, symmetric splits** so floors are directly comparable to SetFit. Silver
66
+ goes in TRAIN only; val/test stay gold.
67
+
68
+ ## Fine-tune — single T4, RLCD, calibration built in
69
+ - `python research/scripts/finetune_single_device.py --data <train.jsonl> --model-dir <english-base> --output-dir <out> --epochs 4` (reproduces the 2×T4 notebook without DDP; CPU/one-GPU).
70
+ - Runs on **one 16 GB GPU (T4)** — ~1–2 h for large data, minutes for our small dims. RLCD = soft-CE + GRPO-style
71
+ policy gradient on proper scoring rules. Calibration (one temperature per type) is fitted inside the run on a
72
+ held-out slice; **argmax/accuracy unchanged, only confidence moves** — always fit before gating on confidence.
73
+ - **OOM knob (hit this):** ModernBERT-large (~400M) OOMs a T4/L4 at the default `batch_size=16, max_seq=256`. Clauses
74
+ are short → set `--batch-size 8 --max-seq 128` (the script/notebook also enables gradient checkpointing). That fit
75
+ comfortably.
76
+ - **Run preflight FIRST — see the [[setfit]] "Run preflight & monitoring" section.** It is framework-agnostic and
77
+ applies to Laya exactly as to SetFit, including when you GENERATE the Laya JSONL labels with a teacher: resolve
78
+ the model from the engine (never a hardcoded/stale id; bulk teacher labeling on Modal Qwen, not OpenRouter), load
79
+ `.env` by EXPLICIT path from an out-of-repo script, SMOKE one item before the fan-out, and stream X/N to a log you
80
+ actively monitor. Those exact mistakes cost runs on 2026-10-04.
81
+ - **Modal harness = reuse the [[setfit]] hardened pattern:** `.uv_pip_install("laya")` + `.add_local_file` the
82
+ finetune script; launcher BLOCKS on `.get()` per spawn (no spawn-and-return), stamps + verifies a `data_sha`,
83
+ writes a manifest, streams X/N; snapshot keepers server-side to a `/checkpoints/<name>` path (a small copy fn) —
84
+ the laya volume has no built-in snapshot, add one. Never lose a fine-tune.
85
+ - **ACCOUNT CONTAINER CAP = 10 (fzaidi2014).** Spawning more than 10 fine-tunes at once does NOT run them all —
86
+ Modal queues the rest and runs ~10 at a time (correct, but ~N/10 waves of wall-clock, and a "why only 10 running?"
87
+ surprise). Size a `groupbake`/dimbatch fan-out to ≤10 in flight (chunk into waves + gather between), or state the
88
+ wave count honestly. This is a DIFFERENT knob from the vLLM `max_containers=1` + `@modal.concurrent` batching in the
89
+ qwen-vllm-modal skill — do not conflate. (Ignored the stated cap once, spawned 29 → 3 waves.)
90
+
91
+ ## Levers that matter (measured, dim-dependent)
92
+ - **Few-shot-in-`state` is a big BUT dim-dependent lever.** Prepend a few labeled exemplars (from TRAIN, per class)
93
+ to the state. It lifted subjective/relational dims a lot (favorability 0.17→0.50, party_asymmetry 0.18→0.55) and
94
+ **HURT a numeric dim** (cap_basis 0.40→0.00, collapsed). **A/B few-shot per dim; never assume it helps.**
95
+ - **Balance: don't down-sample to tiny data.** Balancing favorability DOWN to 18/18 (36 rows) + soft targets
96
+ COLLAPSED it (0.00) — the larger imbalanced set did better (0.50). A starved minority value needs balance-**UP**
97
+ (mine more minority data), not down-sampling the majority.
98
+ - **The residual failures are minority-data-starvation, not a Laya ceiling** — the closest miss (party_asymmetry
99
+ 0.55) is one minority-silver top-up from the bar. Diagnose starvation before concluding "Laya can't."
100
+
101
+ ## Limits (honest)
102
+ - Laya rescues *decision-limited* confusable dims; it does **not** fix (a) **minority-data-starved** values (needs
103
+ more data), or (b) **numeric/structural** distinctions (cap_basis "fixed sum vs multiple-of-fees" — encoders and
104
+ Laya both failed; that belongs on the LLM).
105
+
106
+ ## Serving — device-agnostic, load once
107
+ - `agent = laya.load(<checkpoint>)` — device auto: CUDA → MPS → CPU. **Verified on a Mac (MPS): ~25 s load, ~1.8 s
108
+ first-call warmup, then ~80–140 ms/call.** Runs on CPU too. So it honors the engine's "use a GPU if present, else
109
+ CPU" philosophy — same as SetFit and the LLM profiles.
110
+ - **Load once / preload** (`Router(preload=True)`); never load per call. Serve behind the existing extractor seam so
111
+ nothing upstream changes. The checkpoint is ~820 MB (ModernBERT-large) — heavier than a SetFit body+joblib head;
112
+ budget the memory when co-loading with SetFit models.
113
+ - Serving location is NOT pinned: single T4, share the LLM's A100, or CPU/Mac — the seam + a device arg decide at
114
+ runtime, exactly like the model-profile seam for the LLM.
115
+
116
+ ## Reference implementation
117
+ `~/work/laya` (the cloned repo: `research/scripts/finetune_single_device.py`, `docs/finetune.md`, `examples/`) and
118
+ `~/work/clause-classifier-ab/{laya_modal.py, build_laya_data.py, build_laya_refine.py}` (our Modal fine-tune + eval
119
+ + data builders). Re-use the PATTERNS; the criteria, thresholds, and floors are specific to that problem.
@@ -0,0 +1,96 @@
1
+ ---
2
+ name: qwen-vllm-modal
3
+ description: Authoritative recipe for standing up the self-hosted Qwen3.8-27B vLLM server on Modal (the production LLM substrate). READ THIS + ADR-0110 before touching the deploy — it fixes the "we forgot the agreed config and rediscovered every vLLM/Modal gotcha the hard way" failure. Covers the locked config, the image-build fixes, the concurrency + timeout serving fixes, cold start, and cost control.
4
+ ---
5
+
6
+ # Qwen3.8-27B on Modal (vLLM) — the production LLM server
7
+
8
+ **Read the record, do NOT reconstruct from memory.** The config is decided and the gotchas are known. Rummaging
9
+ half-remembered fragments and iterating live cost a full painful session even though it had all been done, decided,
10
+ and tested before. The two sources of truth:
11
+ 1. **ADR-0110** (`docs/adr/0110-fp8-kv-cache-16k-single-a100.md`) — the LOCKED config + why.
12
+ 2. **`scripts/modal_qwen3_vllm_server.py`** — the deploy script.
13
+ This skill is the operational checklist that ties them together.
14
+
15
+ ## The LOCKED production config (ADR-0110) — do not re-litigate
16
+ **Config B: `Qwen/Qwen3.8-27B-FP8` weights + `--kv-cache-dtype fp8`, on 1× A100-80GB, TP=1, `--max-model-len 16384`, `--gpu-memory-utilization 0.95`.**
17
+ - Config B (FP8 weights) is the **locked deploy** — 52× concurrency @16K, ~23% higher throughput, ~27 GB weights;
18
+ accuracy measured **identical to bf16** on the real compliance workload (no FP8 penalty).
19
+ - **Config A** (`Qwen/Qwen3.8-27B`, **bf16 weights** + FP8 KV) is the **documented FALLBACK only** — 25× concurrency,
20
+ "zero weight-loss" — use it *only* if a later clean check ever shows an FP8-weight regression. Don't default to A.
21
+ - If anyone (including a future you) says "config A is production" — **check ADR-0110 first**; the ADR locks **B**.
22
+ - Served-model-name is `Qwen/Qwen3.8-27B` (what the engine profile `qwen3.8-27b-modal` asks for), even though the
23
+ weights are the `-FP8` repo.
24
+
25
+ **The exact working deploy (this succeeded):**
26
+ ```
27
+ MODAL_IMAGE_BUILDER_VERSION=2025.06 \
28
+ APP_NAME=rw-qwen3-modal MODEL=Qwen/Qwen3.8-27B-FP8 KV_CACHE_DTYPE=fp8 \
29
+ SERVED_NAME=Qwen/Qwen3.8-27B TOOL_PARSER=hermes REASONING_PARSER=qwen3 \
30
+ GPU=A100-80GB:1 TP=1 MAX_LEN=16384 GPU_UTIL=0.95 MAX_NUM_SEQS=<target> \
31
+ uv run --no-sync modal deploy scripts/modal_qwen3_vllm_server.py
32
+ ```
33
+ Then wire the engine: `.env` `VLLM_BASE_URL=https://<workspace>--rw-qwen3-modal-serve.modal.run/v1`,
34
+ `VLLM_API_KEY=rw-vllm-dev-key`; the profile `qwen3.8-27b-modal` routes there. **Stop billing when done:**
35
+ `uv run --no-sync modal app stop rw-qwen3-modal --yes` (A100 is expensive).
36
+
37
+ ## Image-build gotchas (the painful iterations — all fixed)
38
+ Building a fresh vLLM image on a clean Modal account exposed a chain of failures. **The reliable fix is to base off
39
+ the OFFICIAL prebuilt vLLM image** rather than pip-building vLLM:
40
+ 1. `pip_install("vllm")` (unpinned) on `cuda:12.8.1` → **`xformers` source-builds** → `ModuleNotFound: torch`
41
+ (pip build-isolation). Don't pip-build vLLM.
42
+ 2. `uv_pip_install("vllm", extra_index_url=pytorch-cu)` → resolver "no versions of vllm … unsatisfiable" — also fragile.
43
+ 3. **WORKING:** `modal.Image.from_registry("vllm/vllm-openai:latest")` — vllm+torch+xformers+flashinfer prebuilt,
44
+ nothing compiles. It needs two adjustments:
45
+ - `setup_dockerfile_commands=["RUN ln -sf $(command -v python3) /usr/local/bin/python"]` — the image ships
46
+ `python3` but not `python`; Modal's **own pip bootstrap runs right after FROM** as `python -m pip` → **exit 127**
47
+ without the symlink. It must be in `setup_dockerfile_commands` (runs BEFORE the bootstrap), not a later layer.
48
+ - `.entrypoint([])` — clear the image's ENTRYPOINT (the api_server) so Modal's runtime isn't hijacked.
49
+ 4. **Modal LEGACY image builder clobbers the stack**: it installs its old client deps OVER the vLLM image,
50
+ downgrading **pydantic→v1 and fastapi→old**, so vLLM crashes at startup (`cannot import name 'model_validator'`,
51
+ then `cannot import name 'Undefined' from pydantic.fields`). Patching pydantic just moves the break to fastapi
52
+ (whack-a-mole). **FIX = `MODAL_IMAGE_BUILDER_VERSION=2025.06`** (a modern builder whose client stack is pydantic-v2
53
+ compatible). Valid builder versions: `{2023.12, 2024.04, 2024.10, 2025.06}` — use the newest.
54
+ 5. `HF_HUB_DISABLE_XET=1` in the image env — Xet holds an open log handle under the HF cache so `vol.commit()` fails
55
+ on a fresh weight download.
56
+
57
+ ## Serving gotchas (equally painful — all fixed)
58
+ 1. **`@modal.concurrent(max_inputs=N)` is REQUIRED on the web_server function.** Without it Modal feeds the single
59
+ container ONE request at a time (`Running: 1 reqs` + a flood of "Received a cancellation signal", `/health`
60
+ blocked) — vLLM's continuous batching is starved. Concurrency comes from **KV-cache batching inside ONE container**
61
+ (`max_containers=1`), not from more containers. Set `max_inputs` = your `--max-num-seqs`.
62
+ 2. **The client structured-call timeout is too tight for reasoning-ON calls under batch.** At `--max-num-seqs` load,
63
+ reasoning calls take ~50–60 s; the seam's default 60 s timeout kills the ones just over the line → **retry pileup →
64
+ requests pile past `max_inputs` → instant `APIConnectionError` (0.0 s) cascade**. FIX: the seam reads
65
+ `RAG_STRUCTURED_TIMEOUT_S` (default 60) — raise it (e.g. `150`) for reasoning-ON bulk work. Server logs confirm the
66
+ calls actually complete 200 OK (~50 s); it was the client giving up.
67
+ 3. **Pick client concurrency empirically, below the ceiling.** 16 was stable; 20 (== `max-num-seqs`) sat at the exact
68
+ ceiling and risked queueing. `Running: N, Waiting: 0` + zero cancellations = healthy.
69
+ 4. **Reasoning ON vs OFF for bulk classification:** reasoning-OFF is ~30× faster but **changes the labels**
70
+ (over-keeps / different distribution) — do NOT swap it in for work that must match the reasoning-teacher's gold
71
+ (e.g. distillation silver). Keep reasoning ON when fidelity to the gold teacher matters; accept the cost.
72
+
73
+ ## Cold start & cost
74
+ - Cold start ~440–480 s (CPU-bound engine init + a 51-graph CUDA-graph capture + first-time weight download), NOT
75
+ weight I/O — no load trick alone fixes it. GPU-snapshot / sleep-mode is a DEFERRED todo (the one attempt failed
76
+ because it was CPU-only; GPU retry not done). Run long deploys detached; poll `/health` for readiness.
77
+ - **A100 is billed while up — `modal app stop <app> --yes` the moment you're done.** Verify with `modal app list`
78
+ (state `stopped`, 0 tasks) + endpoint 404.
79
+ - **Do NOT warm the GPU before the consumer is ready, and do NOT trust auto-scaledown to protect billing** (burned
80
+ 2026-10-04: a smoke-warmed A100 sat idle ~20 min while unrelated work ran, because "scaledown_window will handle
81
+ it" was assumed — it did not, fast enough). RULE: build the bulk/label/eval consumer FIRST (while the app is
82
+ stopped/cold), THEN warm → smoke → run → **`modal app stop` immediately after**, in one continuous go. If you warm
83
+ only to prove stand-up and the consumer is not next, STOP it right after the smoke. Treat the explicit stop as
84
+ mandatory, not the scaledown as sufficient; verify stopped with `modal app list`.
85
+
86
+ ## The meta-lesson
87
+ Config + recipe live in **ADR-0110 + the script + this skill**. Before any Qwen/Modal work: read them, don't
88
+ reconstruct. When something "was working and we changed accounts/rebuilt," expect the image-builder + concurrency +
89
+ timeout trio above — they are the recurring three.
90
+
91
+ ## Client side (when you use this server for bulk labeling/eval)
92
+ The SERVER config is here; the CLIENT preflight (resolve the model from the engine not a hardcoded id, load `.env`
93
+ by explicit path from an out-of-repo script, smoke ONE item before the bulk fan-out, mandatory X/N progress +
94
+ active monitoring) is in the **`setfit` skill's "Run preflight & monitoring"** section. Read it before driving a
95
+ bulk job against this server — those client gotchas (a stale model id, a silently-unloaded `.env`) cost a run each
96
+ on 2026-10-04.