rag-wright 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- rag_wright-0.1.0/.claude/settings.json +14 -0
- rag_wright-0.1.0/.claude/skills/authoring-a-capability/SKILL.md +106 -0
- rag_wright-0.1.0/.claude/skills/classifier-opportunity-analysis/SKILL.md +175 -0
- rag_wright-0.1.0/.claude/skills/creating-evals/SKILL.md +114 -0
- rag_wright-0.1.0/.claude/skills/laya/SKILL.md +119 -0
- rag_wright-0.1.0/.claude/skills/qwen-vllm-modal/SKILL.md +96 -0
- rag_wright-0.1.0/.claude/skills/setfit/SKILL.md +293 -0
- rag_wright-0.1.0/.claude/skills/using-the-rag-wright-engine/SKILL.md +90 -0
- rag_wright-0.1.0/.env.example +20 -0
- rag_wright-0.1.0/.github/workflows/publish.yml +56 -0
- rag_wright-0.1.0/.gitignore +53 -0
- rag_wright-0.1.0/CHANGELOG.md +48 -0
- rag_wright-0.1.0/CLAUDE.md +201 -0
- rag_wright-0.1.0/LICENSE +21 -0
- rag_wright-0.1.0/PKG-INFO +168 -0
- rag_wright-0.1.0/README.md +118 -0
- rag_wright-0.1.0/SPEC.md +267 -0
- rag_wright-0.1.0/conftest.py +41 -0
- rag_wright-0.1.0/docs/ARCHITECTURE_OVERVIEW.md +199 -0
- rag_wright-0.1.0/docs/ArcadeDB_Local.md +43 -0
- rag_wright-0.1.0/docs/OBSERVABILITY.md +88 -0
- rag_wright-0.1.0/docs/adr/0001-stack-and-library-choices.md +50 -0
- rag_wright-0.1.0/docs/adr/0002-validation-corpus-cuad-edgar.md +52 -0
- rag_wright-0.1.0/docs/adr/0003-ard-registration-and-capability-kinds.md +96 -0
- rag_wright-0.1.0/docs/adr/0004-entity-disambiguation.md +105 -0
- rag_wright-0.1.0/docs/adr/0005-relational-golden-set.md +89 -0
- rag_wright-0.1.0/docs/adr/0006-model-profile.md +66 -0
- rag_wright-0.1.0/docs/adr/0007-arcadedb-store-schema.md +110 -0
- rag_wright-0.1.0/docs/adr/0008-unified-framework-index-docs-grounding.md +111 -0
- rag_wright-0.1.0/docs/adr/0009-rlm-chunking-retained-as-configurable-capability.md +69 -0
- rag_wright-0.1.0/docs/adr/0010-transformers-pinned-below-5-for-flagembedding-reranker.md +46 -0
- rag_wright-0.1.0/docs/adr/0011-cuad-not-a-retrieval-benchmark-acord-for-queries.md +70 -0
- rag_wright-0.1.0/docs/adr/0012-entity-mention-confidence-no-proximity-edges.md +69 -0
- rag_wright-0.1.0/docs/adr/0013-entity-resolution-matching-strategy.md +48 -0
- rag_wright-0.1.0/docs/adr/0014-split-generation-and-vision-to-text.md +51 -0
- rag_wright-0.1.0/docs/adr/0015-rlm-dynamic-subagents-and-granted-subagents.md +107 -0
- rag_wright-0.1.0/docs/adr/0016-rlm-defined-by-required-capabilities-enforced-as-tests.md +42 -0
- rag_wright-0.1.0/docs/adr/0017-dynamic-dispatch-trigger-as-typed-flag.md +77 -0
- rag_wright-0.1.0/docs/adr/0018-rlm-method-is-the-orchestrator-system-prompt.md +65 -0
- rag_wright-0.1.0/docs/adr/0019-recursion-is-optional-for-chunking-required-for-synthesis.md +51 -0
- rag_wright-0.1.0/docs/adr/0020-serialize-interpreter-sessions-per-process-ki1.md +97 -0
- rag_wright-0.1.0/docs/adr/0021-capability-interface-governed-typed-io.md +106 -0
- rag_wright-0.1.0/docs/adr/0022-fr-k-embedding-free-okf-navigation-experimental.md +161 -0
- rag_wright-0.1.0/docs/adr/0023-cheap-model-for-okf-enrichment.md +51 -0
- rag_wright-0.1.0/docs/adr/0024-reader-parallelism-lives-in-python-not-the-interpreter.md +57 -0
- rag_wright-0.1.0/docs/adr/0025-retrieval-pivot-function-classify-property-graph-rerank.md +82 -0
- rag_wright-0.1.0/docs/adr/0026-property-schema-and-extended-function-taxonomy.md +68 -0
- rag_wright-0.1.0/docs/adr/0027-openrouter-provider-routing-by-throughput.md +46 -0
- rag_wright-0.1.0/docs/adr/0028-deterministic-grounding-judge-and-flash-pro-cascade.md +52 -0
- rag_wright-0.1.0/docs/adr/0029-retrieval-pipeline-domain-portability-and-adaptation.md +94 -0
- rag_wright-0.1.0/docs/adr/0030-model-training-standard-modal-reusable-checkpointed.md +63 -0
- rag_wright-0.1.0/docs/adr/0031-single-call-chunking-for-structured-contracts.md +57 -0
- rag_wright-0.1.0/docs/adr/0032-nl-to-type-two-step-reason-emit-on-gemma.md +58 -0
- rag_wright-0.1.0/docs/adr/0033-unified-contract-kg-three-legs-one-graph.md +56 -0
- rag_wright-0.1.0/docs/adr/0034-granite-json-schema-structured-output-profile.md +44 -0
- rag_wright-0.1.0/docs/adr/0035-graph-extraction-rebacked-with-gp1b-retire-hybrid.md +54 -0
- rag_wright-0.1.0/docs/adr/0036-party-clause-link-party-to-edge.md +62 -0
- rag_wright-0.1.0/docs/adr/0037-clause-template-is-authoritative-code-not-generated.md +59 -0
- rag_wright-0.1.0/docs/adr/0038-full-cuad-kg-on-gcp-gpu-vm-and-gcs-backup.md +54 -0
- rag_wright-0.1.0/docs/adr/0039-self-hosted-open-model-stack-on-modal-a100.md +59 -0
- rag_wright-0.1.0/docs/adr/0040-neuro-symbolic-extraction-fidelity-ontology-shacl-validation.md +82 -0
- rag_wright-0.1.0/docs/adr/0041-capability-kind-rubric-and-agent-skill-runtime-tiers.md +81 -0
- rag_wright-0.1.0/docs/adr/0042-clause-level-span-id-provenance-and-content-hash-backfill.md +56 -0
- rag_wright-0.1.0/docs/adr/0043-retire-cross-corpus-retrieval-standardize-on-leg-b.md +47 -0
- rag_wright-0.1.0/docs/adr/0044-is-exception-to-derived-carveout-relationship.md +122 -0
- rag_wright-0.1.0/docs/adr/0045-client-side-tag-parse-structured-output.md +82 -0
- rag_wright-0.1.0/docs/adr/0046-acord-unified-into-one-production-kg.md +50 -0
- rag_wright-0.1.0/docs/adr/0047-retire-precomputed-clause-function-gate.md +44 -0
- rag_wright-0.1.0/docs/adr/0048-ingest-llm-classifier-nondestructive-reclassify.md +187 -0
- rag_wright-0.1.0/docs/adr/0049-generic-customer-lens-for-ingestion.md +63 -0
- rag_wright-0.1.0/docs/adr/0050-async-langgraph-ingestion.md +87 -0
- rag_wright-0.1.0/docs/adr/0051-schema-bootstrap-and-feedback-driven-ontology-evolution.md +127 -0
- rag_wright-0.1.0/docs/adr/0052-engine-product-split-graphwright-parked.md +119 -0
- rag_wright-0.1.0/docs/adr/0053-answer-prose-output-hygiene.md +72 -0
- rag_wright-0.1.0/docs/adr/0054-remove-auto-tag-from-generator-evidence.md +75 -0
- rag_wright-0.1.0/docs/adr/0055-confidence-out-of-band.md +68 -0
- rag_wright-0.1.0/docs/adr/0056-bound-structured-retry-wall-clock.md +96 -0
- rag_wright-0.1.0/docs/adr/0057-async-engine-architecture.md +82 -0
- rag_wright-0.1.0/docs/adr/0058-structure-first-chunking.md +75 -0
- rag_wright-0.1.0/docs/adr/0059-recall-decoupled-from-classification-and-visible-partial-loss.md +102 -0
- rag_wright-0.1.0/docs/adr/0060-scope-compliance-check-to-named-policy-sources.md +83 -0
- rag_wright-0.1.0/docs/adr/0061-subject-document-compliance-per-section.md +59 -0
- rag_wright-0.1.0/docs/adr/0062-tiered-ocr-scan-quality-vlm-escalation.md +86 -0
- rag_wright-0.1.0/docs/adr/0063-per-sentence-compliance-subject-facts.md +61 -0
- rag_wright-0.1.0/docs/adr/0064-typed-properties-out-of-band-on-evidence.md +58 -0
- rag_wright-0.1.0/docs/adr/0065-deontic-applicability-gates-query-side.md +74 -0
- rag_wright-0.1.0/docs/adr/0066-ontology-ttl-single-source-of-truth.md +175 -0
- rag_wright-0.1.0/docs/adr/0067-domain-pack-retargeting.md +93 -0
- rag_wright-0.1.0/docs/adr/0068-recall-first-actor-gate-ontology-role-disjointness.md +73 -0
- rag_wright-0.1.0/docs/adr/0069-reading-order-chunking-tables-figures-retrievable.md +80 -0
- rag_wright-0.1.0/docs/adr/0070-text-layer-first-parsing-no-false-vlm-escalation.md +51 -0
- rag_wright-0.1.0/docs/adr/0071-defragmentation-reconstruct-paragraphs-from-line-items.md +57 -0
- rag_wright-0.1.0/docs/adr/0072-extract-guard-furniture-and-deterministic-failure.md +52 -0
- rag_wright-0.1.0/docs/adr/0073-born-digital-threshold-sparse-pages-no-vlm-escalation.md +50 -0
- rag_wright-0.1.0/docs/adr/0074-retry-transient-extraction-failures.md +44 -0
- rag_wright-0.1.0/docs/adr/0075-per-page-vlm-escalation.md +45 -0
- rag_wright-0.1.0/docs/adr/0076-thread-safe-shared-embedder-reranker.md +55 -0
- rag_wright-0.1.0/docs/adr/0077-concurrent-function-classification-across-chunks.md +45 -0
- rag_wright-0.1.0/docs/adr/0078-dedicated-extraction-executor.md +57 -0
- rag_wright-0.1.0/docs/adr/0079-product-default-granite-4.2-openrouter-routing.md +62 -0
- rag_wright-0.1.0/docs/adr/0080-nested-tag-parse-and-degrade.md +48 -0
- rag_wright-0.1.0/docs/adr/0081-function-independent-tagparse-clause-extraction.md +70 -0
- rag_wright-0.1.0/docs/adr/0082-symbolic-gate-function-independent.md +49 -0
- rag_wright-0.1.0/docs/adr/0083-executor-hop-trace-context-capture.md +31 -0
- rag_wright-0.1.0/docs/adr/0084-gleaning-off-on-the-query-leg.md +24 -0
- rag_wright-0.1.0/docs/adr/0085-query-constraint-extraction-tagparse.md +24 -0
- rag_wright-0.1.0/docs/adr/0086-streaming-cost-capture.md +29 -0
- rag_wright-0.1.0/docs/adr/0087-retrieval-relevance-score-on-rankedspan.md +22 -0
- rag_wright-0.1.0/docs/adr/0088-per-span-relevance-verdict.md +32 -0
- rag_wright-0.1.0/docs/adr/0089-structured-output-generation-tracing.md +26 -0
- rag_wright-0.1.0/docs/adr/0090-affiliate-of-extraction.md +25 -0
- rag_wright-0.1.0/docs/adr/0091-retire-partyto-edge.md +29 -0
- rag_wright-0.1.0/docs/adr/0092-idempotent-write-graph.md +30 -0
- rag_wright-0.1.0/docs/adr/0093-entities-by-name-seam.md +28 -0
- rag_wright-0.1.0/docs/adr/0094-workspace-document-scope.md +25 -0
- rag_wright-0.1.0/docs/adr/0095-span-page-provenance.md +24 -0
- rag_wright-0.1.0/docs/adr/0096-carveout-keyword-normalization.md +25 -0
- rag_wright-0.1.0/docs/adr/0097-caller-configurable-ingest-extraction-models.md +29 -0
- rag_wright-0.1.0/docs/adr/0098-invoke-time-document-scope.md +24 -0
- rag_wright-0.1.0/docs/adr/0099-mcp-tools-never-take-a-model-supplied-tenant.md +27 -0
- rag_wright-0.1.0/docs/adr/0100-profile-based-model-routing.md +30 -0
- rag_wright-0.1.0/docs/adr/0101-untagged-spans-reach-extraction-aspect-gate-removed.md +26 -0
- rag_wright-0.1.0/docs/adr/0102-open-descriptive-list-dims-retain-verbatim.md +33 -0
- rag_wright-0.1.0/docs/adr/0103-a-clause-is-a-provision-not-a-span.md +41 -0
- rag_wright-0.1.0/docs/adr/0104-dense-floor-protection-in-leg-b.md +41 -0
- rag_wright-0.1.0/docs/adr/0105-in-band-model-usage-accounting.md +26 -0
- rag_wright-0.1.0/docs/adr/0106-document-signals-off-the-citation.md +25 -0
- rag_wright-0.1.0/docs/adr/0107-requirement-page-provenance.md +25 -0
- rag_wright-0.1.0/docs/adr/0108-qwen3-27b-single-a100-serving-profile.md +47 -0
- rag_wright-0.1.0/docs/adr/0109-vllm-cold-start-reduction.md +54 -0
- rag_wright-0.1.0/docs/adr/0110-fp8-kv-cache-16k-single-a100.md +56 -0
- rag_wright-0.1.0/docs/adr/0111-pin-qwen3-27b-deepinfra-bf16.md +25 -0
- rag_wright-0.1.0/docs/adr/0112-relax-deepagents-pin-to-floor.md +26 -0
- rag_wright-0.1.0/docs/adr/0113-llm-span-instrumentation-for-latency-attribution.md +29 -0
- rag_wright-0.1.0/docs/adr/0114-setfit-default-clause-function-classifier.md +45 -0
- rag_wright-0.1.0/docs/adr/0115-classifier-only-step3a-property-extraction.md +61 -0
- rag_wright-0.1.0/docs/adr/0116-soft-function-scoping-classifier-lane.md +45 -0
- rag_wright-0.1.0/docs/adr/0117-engine-api-layer-and-capability-runtime.md +90 -0
- rag_wright-0.1.0/docs/adr/0118-engine-core-api-vs-ard-adapter-free-impl-ref-client.md +100 -0
- rag_wright-0.1.0/docs/adr/0119-jev-typed-decision-model-for-compliance-closed-set-fields.md +68 -0
- rag_wright-0.1.0/docs/adr/0120-rlm-sub-agent-identity-dynamic-dispatch-trigger.md +35 -0
- rag_wright-0.1.0/docs/adr/0121-spacy-optional-extra-model-runtime-download.md +44 -0
- rag_wright-0.1.0/docs/adr/0122-provision-boundary-deterministic-plus-decision-model-residue.md +60 -0
- rag_wright-0.1.0/docs/adr/README.md +181 -0
- rag_wright-0.1.0/docs/api/README.md +123 -0
- rag_wright-0.1.0/docs/architecture/foundations-and-adding-a-domain.md +115 -0
- rag_wright-0.1.0/docs/architecture.md +103 -0
- rag_wright-0.1.0/docs/archive/README.md +32 -0
- rag_wright-0.1.0/docs/archive/design/deontic-applicability-routing.md +146 -0
- rag_wright-0.1.0/docs/archive/design/ingestion-neuro-symbolic-gaps.md +113 -0
- rag_wright-0.1.0/docs/archive/design/semantic-subject-segmentation.md +120 -0
- rag_wright-0.1.0/docs/archive/design/unify-subject-preprocessing.md +135 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0037-closed-vocab-drops-verbatim-values-to-other.md +44 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0038-a-clause-node-is-now-created-per-sentence-so-98-percent-of-spans-become-clauses.md +109 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0039-the-provision-detector-reads-text-but-docling-puts-the-section-number-in-marker.md +117 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0040-covered-subject-is-asked-of-every-provision-so-verbatim-retention-fills-it-with-noise.md +114 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0041-VERIFICATION-dense-floor.md +82 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0041-leg-b-discards-the-best-dense-matches-on-some-queries.md +106 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0042-usage-is-captured-per-call-but-never-returned-so-cost-needs-a-langfuse-round-trip.md +85 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0043-a-requirement-carries-no-page-provenance-so-a-finding-cannot-point-into-its-policy.md +67 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0044-document-signals-scaffolding-is-shown-to-the-user-as-a-quote-from-their-document.md +72 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0045-a-fixed-money-cap-records-no-cap-quantum-so-the-amount-is-only-prose.md +82 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0046-appending-a-natural-follow-up-to-a-question-drops-the-clause-type.md +72 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0047-an-exact-deepagents-pin-transitively-pins-every-consumer.md +111 -0
- rag_wright-0.1.0/docs/archive/engine-issues/0048-the-pinned-endpoints-latency-tail-makes-agent-runs-undebuggable-from-either-side.md +89 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-07-20_graphwright_capability_interface_reply.md +136 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-07-20b_graphwright_interface_confirmation.md +101 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-07-20c_graphwright_ingestion_interfaces.md +101 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-08-30_subject_compliance_final_rulewright.md +111 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-01_issue-0013-recall-first-actor-gate_rulewright.md +41 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-01_ontology-source-of-truth_rulewright.md +80 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-03_issue-0014-table-retrievability_rulewright.md +58 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-04_bulk-ingestion-wall_rulewright.md +63 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-04_issue-0016-thread-safe-embedder_rulewright.md +36 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-05_observability-langfuse_rulewright.md +78 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-05_tagparse-ingestion-and-granite-4.2_rulewright.md +130 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_affiliate-of-extraction_rulewright.md +41 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_issue-0028-partyto-retired_rulewright.md +40 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_issue-0029-idempotent-write-graph_rulewright.md +38 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_issue-0030-entities-by-name_rulewright.md +42 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_observability-and-retrieval-floor_rulewright.md +26 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_relevance-verdict_rulewright.md +53 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_rename-run-ad-compliance-check_rulewright.md +30 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-09_structured-output-cost-tracing_rulewright.md +44 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_ingest-models-caller-configurable_rulewright.md +49 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0031-followup-failclosed-and-nodrift_rulewright.md +28 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0031-followup-validate-ingested-set_rulewright.md +31 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0031-workspace-document-scope_rulewright.md +51 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0032-span-page-provenance_rulewright.md +38 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0033-carveout-normalization_rulewright.md +35 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-10_issue-0034-invoke-time-document-scope_rulewright.md +35 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_issue-0035-followup-derived-guard_rulewright.md +29 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_issue-0035-mcp-no-model-tenant_rulewright.md +44 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_profile-based-model-routing_rulewright.md +47 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_qwen-default-modal-or_rulewright.md +53 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-11_qwen-everywhere-config_rulewright.md +32 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-12_clause-granularity_0036-0039_rulewright.md +77 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-12_closed-vocab_0037-0040_rulewright.md +46 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-12_dense-floor-protection_0041_rulewright.md +31 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-13_in-band-usage_0042_rulewright.md +44 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-14_document-signals-off-citation_0044_rulewright.md +26 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-14_requirement-page-provenance_0043_rulewright.md +35 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-16_qwen3-27b-single-a100-serving-profile_rulewright.md +38 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-18_fp8-accuracy-eval-configB_rulewright.md +44 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-18_fp8-kv-16k-single-a100_rulewright.md +41 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-20_deepagents-pin-relaxed_0047_rulewright.md +22 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-20_llm-span-instrumentation_0048_rulewright.md +31 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-21_0048-part2-modal-answers_from_rulewright.md +131 -0
- rag_wright-0.1.0/docs/archive/handoffs/2026-09-21_0048-part2-modal-scoping-questions_rulewright.md +34 -0
- rag_wright-0.1.0/docs/archive/misc/.gitkeep +0 -0
- rag_wright-0.1.0/docs/archive/plans/Corpus_Acquisition.md +11 -0
- rag_wright-0.1.0/docs/archive/plans/async-migration.md +72 -0
- rag_wright-0.1.0/docs/archive/plans/demo_plan.md +103 -0
- rag_wright-0.1.0/docs/archive/plans/gp1b_docling_graph_plan.md +121 -0
- rag_wright-0.1.0/docs/archive/plans/unified_contract_kg_ontology_bridge.md +260 -0
- rag_wright-0.1.0/docs/archive/plans/unified_contract_kg_plan.md +104 -0
- rag_wright-0.1.0/docs/archive/results/2026-07-24-t58-property-graph-population.md +43 -0
- rag_wright-0.1.0/docs/archive/results/2026-07-25-t58b-full-pipeline-rerank.md +76 -0
- rag_wright-0.1.0/docs/archive/results/2026-07-25-t58b-topk-ordering-levers.md +80 -0
- rag_wright-0.1.0/docs/concepts.md +110 -0
- rag_wright-0.1.0/docs/configuration.md +91 -0
- rag_wright-0.1.0/docs/contract_pipeline_explainer.md +169 -0
- rag_wright-0.1.0/docs/corpus_ingest_recipe.md +56 -0
- rag_wright-0.1.0/docs/domain-adaptation/README.md +71 -0
- rag_wright-0.1.0/docs/domain-adaptation/_engine-gaps.md +44 -0
- rag_wright-0.1.0/docs/domain-adaptation/authoring-capabilities.md +85 -0
- rag_wright-0.1.0/docs/domain-adaptation/classification-and-decision-models.md +69 -0
- rag_wright-0.1.0/docs/domain-adaptation/entity-resolution.md +54 -0
- rag_wright-0.1.0/docs/domain-adaptation/kg-construction.md +66 -0
- rag_wright-0.1.0/docs/domain-adaptation/ontology-authoring.md +83 -0
- rag_wright-0.1.0/docs/eval/classifier_model_ab.md +61 -0
- rag_wright-0.1.0/docs/eval/compliance_demo.md +53 -0
- rag_wright-0.1.0/docs/eval/compliance_gate_cc7.md +68 -0
- rag_wright-0.1.0/docs/eval/compliance_rung2_cc7.md +156 -0
- rag_wright-0.1.0/docs/eval/cuad_highlighting_cu-d1.md +72 -0
- rag_wright-0.1.0/docs/eval/dg_model_ab.md +48 -0
- rag_wright-0.1.0/docs/eval/function_gate_recall.md +48 -0
- rag_wright-0.1.0/docs/eval/generation_robustness_b_vs_c.md +64 -0
- rag_wright-0.1.0/docs/eval/nl_to_type_cu-d2.md +90 -0
- rag_wright-0.1.0/docs/eval/ocr_benchmark.md +87 -0
- rag_wright-0.1.0/docs/eval/prod1_readiness.md +55 -0
- rag_wright-0.1.0/docs/eval/prod2_readiness.md +99 -0
- rag_wright-0.1.0/docs/eval/query_side_model_ab.md +63 -0
- rag_wright-0.1.0/docs/eval/silver_granite_vs_gemma4.md +47 -0
- rag_wright-0.1.0/docs/eval/silver_provider_routing.md +48 -0
- rag_wright-0.1.0/docs/eval/silver_selfhosted_gemma4_26b.md +53 -0
- rag_wright-0.1.0/docs/installation.md +93 -0
- rag_wright-0.1.0/docs/playbook.md +222 -0
- rag_wright-0.1.0/docs/product/capability_profiles.md +166 -0
- rag_wright-0.1.0/docs/product/contracts_product_roadmap.md +502 -0
- rag_wright-0.1.0/docs/product/engine-api-migration-handoff.md +278 -0
- rag_wright-0.1.0/docs/product/engine_async_api.md +150 -0
- rag_wright-0.1.0/docs/product/new-domain-build-sequence.md +10 -0
- rag_wright-0.1.0/docs/product/seam-adaptation-guide.md +64 -0
- rag_wright-0.1.0/docs/proposals/compliance-ingest-classifier-decomposition.md +127 -0
- rag_wright-0.1.0/docs/proposals/de-domaining-and-capability-runtime.md +213 -0
- rag_wright-0.1.0/docs/proposals/new-domain-developer-journey.md +263 -0
- rag_wright-0.1.0/docs/quickstart.md +106 -0
- rag_wright-0.1.0/docs/reference-pack.md +69 -0
- rag_wright-0.1.0/docs/specs/engine-platform/SPEC.md +81 -0
- rag_wright-0.1.0/docs/specs/engine-platform/TASKS.md +234 -0
- rag_wright-0.1.0/docs/specs/engine-prep/plan.md +525 -0
- rag_wright-0.1.0/docs/templates/product-starter/CLAUDE.md.template +193 -0
- rag_wright-0.1.0/docs/templates/product-starter/README.md +45 -0
- rag_wright-0.1.0/docs/templates/product-starter/playbook.md.template +174 -0
- rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-buckets-schema.md +82 -0
- rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-docs-extraction.json +2465 -0
- rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-vector-embeddings.md +696 -0
- rag_wright-0.1.0/docs/vendor/arcadedb/arcadedb-vector-search-tutorial.md +455 -0
- rag_wright-0.1.0/eval/__init__.py +8 -0
- rag_wright-0.1.0/eval/ablation.py +58 -0
- rag_wright-0.1.0/eval/acord.py +93 -0
- rag_wright-0.1.0/eval/acord_retrieval.py +75 -0
- rag_wright-0.1.0/eval/build_golden.py +51 -0
- rag_wright-0.1.0/eval/category_retrieval.py +118 -0
- rag_wright-0.1.0/eval/compliance_demo/policy/community_conduct_policy.md +24 -0
- rag_wright-0.1.0/eval/compliance_demo/subjects/post_borderline.txt +1 -0
- rag_wright-0.1.0/eval/compliance_demo/subjects/post_compliant.txt +1 -0
- rag_wright-0.1.0/eval/compliance_demo/subjects/post_violation.txt +1 -0
- rag_wright-0.1.0/eval/condensed_pipeline.py +284 -0
- rag_wright-0.1.0/eval/contractnli_judge.py +83 -0
- rag_wright-0.1.0/eval/cuad_highlight.py +218 -0
- rag_wright-0.1.0/eval/full_pipeline_rerank.py +159 -0
- rag_wright-0.1.0/eval/function_ceiling.py +75 -0
- rag_wright-0.1.0/eval/function_property_rerank.py +113 -0
- rag_wright-0.1.0/eval/function_rerank.py +107 -0
- rag_wright-0.1.0/eval/gate1_chunker_ab.py +150 -0
- rag_wright-0.1.0/eval/gate2_hybrid_rerank.py +301 -0
- rag_wright-0.1.0/eval/golden/relational/set.json +698 -0
- rag_wright-0.1.0/eval/golden.py +112 -0
- rag_wright-0.1.0/eval/ground_discriminator_rerank.py +193 -0
- rag_wright-0.1.0/eval/harness.py +83 -0
- rag_wright-0.1.0/eval/kg_primary.py +399 -0
- rag_wright-0.1.0/eval/kg_property_rerank.py +135 -0
- rag_wright-0.1.0/eval/listwise_rerank.py +254 -0
- rag_wright-0.1.0/eval/multihop.py +211 -0
- rag_wright-0.1.0/eval/nl_to_type.py +206 -0
- rag_wright-0.1.0/eval/okf_gold.py +214 -0
- rag_wright-0.1.0/eval/reachability.py +344 -0
- rag_wright-0.1.0/eval/relational_eval.py +91 -0
- rag_wright-0.1.0/eval/test_acord.py +60 -0
- rag_wright-0.1.0/eval/test_acord_retrieval.py +51 -0
- rag_wright-0.1.0/eval/test_category_retrieval.py +54 -0
- rag_wright-0.1.0/eval/test_harness.py +112 -0
- rag_wright-0.1.0/eval/test_multihop_set.py +164 -0
- rag_wright-0.1.0/eval/test_okf_gold.py +164 -0
- rag_wright-0.1.0/eval/test_reachability.py +134 -0
- rag_wright-0.1.0/examples/quickstart.py +94 -0
- rag_wright-0.1.0/plan.md +323 -0
- rag_wright-0.1.0/pyproject.toml +157 -0
- rag_wright-0.1.0/scripts/ab_model_overlap.py +110 -0
- rag_wright-0.1.0/scripts/acord_unify.py +223 -0
- rag_wright-0.1.0/scripts/acquire_acord.py +83 -0
- rag_wright-0.1.0/scripts/acquire_cuad.py +183 -0
- rag_wright-0.1.0/scripts/acquire_ecfr.py +72 -0
- rag_wright-0.1.0/scripts/acquire_edgar.py +105 -0
- rag_wright-0.1.0/scripts/acquire_ftc_255.py +64 -0
- rag_wright-0.1.0/scripts/acquire_prod1_corpus.py +78 -0
- rag_wright-0.1.0/scripts/audit_reclass_flips.py +127 -0
- rag_wright-0.1.0/scripts/author_leg_a_silver_key.py +151 -0
- rag_wright-0.1.0/scripts/backfill_affiliations.py +106 -0
- rag_wright-0.1.0/scripts/backfill_clause_span_id.py +95 -0
- rag_wright-0.1.0/scripts/backfill_edge_source_doc_id.py +69 -0
- rag_wright-0.1.0/scripts/backup_kg_to_gcs.sh +39 -0
- rag_wright-0.1.0/scripts/bootstrap_template_capture.py +72 -0
- rag_wright-0.1.0/scripts/bootstrap_typed_edges.py +54 -0
- rag_wright-0.1.0/scripts/build_api_docs.py +59 -0
- rag_wright-0.1.0/scripts/build_api_docs.sh +7 -0
- rag_wright-0.1.0/scripts/build_cuad_clause_cache.py +92 -0
- rag_wright-0.1.0/scripts/build_function_routing_map.py +50 -0
- rag_wright-0.1.0/scripts/cic1_assemble_v2.py +125 -0
- rag_wright-0.1.0/scripts/cic1_assemble_v3.py +126 -0
- rag_wright-0.1.0/scripts/cic1_generate_hard_negatives.py +100 -0
- rag_wright-0.1.0/scripts/cic1_generate_hard_positives.py +108 -0
- rag_wright-0.1.0/scripts/cic1_generate_training.py +154 -0
- rag_wright-0.1.0/scripts/cic1_hybrid.py +88 -0
- rag_wright-0.1.0/scripts/cic1_hybrid2.py +104 -0
- rag_wright-0.1.0/scripts/cic1_jev_actor.py +100 -0
- rag_wright-0.1.0/scripts/cic1_jev_claimtypes.py +104 -0
- rag_wright-0.1.0/scripts/cic1_jev_operative.py +80 -0
- rag_wright-0.1.0/scripts/cic1_label_spans.py +202 -0
- rag_wright-0.1.0/scripts/cic1_laya_prep.py +44 -0
- rag_wright-0.1.0/scripts/cic1_laya_train.py +81 -0
- rag_wright-0.1.0/scripts/cic1_prep_operative.py +83 -0
- rag_wright-0.1.0/scripts/cic1_relabel_rubric.py +107 -0
- rag_wright-0.1.0/scripts/cic1_relabel_rubric2.py +108 -0
- rag_wright-0.1.0/scripts/cic1_train_operative.py +175 -0
- rag_wright-0.1.0/scripts/compare_extraction_models.py +142 -0
- rag_wright-0.1.0/scripts/compliance_actor_gate_smoke.py +113 -0
- rag_wright-0.1.0/scripts/compliance_engine_smoke.py +65 -0
- rag_wright-0.1.0/scripts/compliance_policy_demo.py +69 -0
- rag_wright-0.1.0/scripts/curate_taxonomy_gaps.py +127 -0
- rag_wright-0.1.0/scripts/dg_model_ab.py +119 -0
- rag_wright-0.1.0/scripts/diagnose_acord_knn_expansion.py +106 -0
- rag_wright-0.1.0/scripts/diagnose_acord_knn_rerank.py +99 -0
- rag_wright-0.1.0/scripts/diagnose_acord_leg_pool.py +86 -0
- rag_wright-0.1.0/scripts/diagnose_acord_legs.py +112 -0
- rag_wright-0.1.0/scripts/diagnose_acord_pool.py +66 -0
- rag_wright-0.1.0/scripts/diagnose_acord_relstructure.py +106 -0
- rag_wright-0.1.0/scripts/diagnose_acord_retrieval.py +63 -0
- rag_wright-0.1.0/scripts/distill/eval_all.py +122 -0
- rag_wright-0.1.0/scripts/distill/eval_ce.py +107 -0
- rag_wright-0.1.0/scripts/distill/eval_listwise_b.py +152 -0
- rag_wright-0.1.0/scripts/distill/export_ce_dataset.py +69 -0
- rag_wright-0.1.0/scripts/distill/extract_features.py +85 -0
- rag_wright-0.1.0/scripts/distill/listwise_variants.py +213 -0
- rag_wright-0.1.0/scripts/distill/relational_features.py +81 -0
- rag_wright-0.1.0/scripts/distill/train_ce.py +74 -0
- rag_wright-0.1.0/scripts/distill/train_modal.py +121 -0
- rag_wright-0.1.0/scripts/enrich_edgar_candidates.py +114 -0
- rag_wright-0.1.0/scripts/eval_chunking_ab.py +176 -0
- rag_wright-0.1.0/scripts/eval_classifier_guided_vs_tagparse.py +156 -0
- rag_wright-0.1.0/scripts/eval_compliance_gold.py +91 -0
- rag_wright-0.1.0/scripts/extract_arcadedb_docs.py +74 -0
- rag_wright-0.1.0/scripts/generate_contract_python.py +30 -0
- rag_wright-0.1.0/scripts/git_post_commit_graphify.sh +25 -0
- rag_wright-0.1.0/scripts/granite_chunker_assess.py +124 -0
- rag_wright-0.1.0/scripts/graph_status.sh +35 -0
- rag_wright-0.1.0/scripts/ingest_acord.py +103 -0
- rag_wright-0.1.0/scripts/ingest_compliance_async_prod2.py +67 -0
- rag_wright-0.1.0/scripts/ingest_compliance_document_prod2.py +57 -0
- rag_wright-0.1.0/scripts/ingest_compliance_prod2.py +63 -0
- rag_wright-0.1.0/scripts/ingest_cuad.py +203 -0
- rag_wright-0.1.0/scripts/ingest_cuad_full.py +61 -0
- rag_wright-0.1.0/scripts/ingest_ftc_compliance.py +50 -0
- rag_wright-0.1.0/scripts/ingest_prod1.py +95 -0
- rag_wright-0.1.0/scripts/ingest_prod1_async.py +101 -0
- rag_wright-0.1.0/scripts/ingest_smoke.py +81 -0
- rag_wright-0.1.0/scripts/install_git_hooks.sh +13 -0
- rag_wright-0.1.0/scripts/label_new_functions.py +136 -0
- rag_wright-0.1.0/scripts/legb_function_gate_recall.py +121 -0
- rag_wright-0.1.0/scripts/make_compliance_fixtures.py +87 -0
- rag_wright-0.1.0/scripts/make_compliance_gold.py +159 -0
- rag_wright-0.1.0/scripts/mcp_compliance_agent_demo.py +78 -0
- rag_wright-0.1.0/scripts/mcp_intra_document_qa_smoke.py +65 -0
- rag_wright-0.1.0/scripts/mcp_query_legs_agent_demo.py +84 -0
- rag_wright-0.1.0/scripts/measure_generation_robustness.py +172 -0
- rag_wright-0.1.0/scripts/measure_silver.py +187 -0
- rag_wright-0.1.0/scripts/merge_docs_into_framework.py +43 -0
- rag_wright-0.1.0/scripts/migrate_entity_cik_to_canonical_id.py +47 -0
- rag_wright-0.1.0/scripts/migrate_silver_evidence_autotag.py +58 -0
- rag_wright-0.1.0/scripts/migrate_silver_evidence_remove_autotag.py +68 -0
- rag_wright-0.1.0/scripts/mine_scarce_functions.py +89 -0
- rag_wright-0.1.0/scripts/modal_arcadedb.py +80 -0
- rag_wright-0.1.0/scripts/modal_backfill.py +114 -0
- rag_wright-0.1.0/scripts/modal_exception_linking.py +109 -0
- rag_wright-0.1.0/scripts/modal_gemma4_vllm.py +105 -0
- rag_wright-0.1.0/scripts/modal_gemma4_vllm_snapshot.py +189 -0
- rag_wright-0.1.0/scripts/modal_granite_server.py +128 -0
- rag_wright-0.1.0/scripts/modal_granite_throughput.py +147 -0
- rag_wright-0.1.0/scripts/modal_granite_vllm_server.py +44 -0
- rag_wright-0.1.0/scripts/modal_query_app.py +103 -0
- rag_wright-0.1.0/scripts/modal_qwen3_27b_bench.py +234 -0
- rag_wright-0.1.0/scripts/modal_qwen3_27b_snapshot.py +158 -0
- rag_wright-0.1.0/scripts/modal_qwen3_vllm_server.py +96 -0
- rag_wright-0.1.0/scripts/modal_stack_a100.py +121 -0
- rag_wright-0.1.0/scripts/ocr_benchmark.py +268 -0
- rag_wright-0.1.0/scripts/ocr_preprocess.py +77 -0
- rag_wright-0.1.0/scripts/ontology_dimension_check.py +95 -0
- rag_wright-0.1.0/scripts/phase_a_leg_validate.py +120 -0
- rag_wright-0.1.0/scripts/populate_clause_kg.py +141 -0
- rag_wright-0.1.0/scripts/populate_entity_graph.py +107 -0
- rag_wright-0.1.0/scripts/populate_entity_graph_extracted.py +110 -0
- rag_wright-0.1.0/scripts/populate_property_store.py +202 -0
- rag_wright-0.1.0/scripts/prep_relational_verification.py +113 -0
- rag_wright-0.1.0/scripts/publish_manifests.py +30 -0
- rag_wright-0.1.0/scripts/reclassify_kg.py +322 -0
- rag_wright-0.1.0/scripts/refresh_framework_graph.sh +100 -0
- rag_wright-0.1.0/scripts/revert_reclass_flips.py +69 -0
- rag_wright-0.1.0/scripts/run_acord_retrieval.py +123 -0
- rag_wright-0.1.0/scripts/run_clause_exception_linking.py +38 -0
- rag_wright-0.1.0/scripts/semantic_judge_live_validate.py +131 -0
- rag_wright-0.1.0/scripts/semantic_judge_probe.py +84 -0
- rag_wright-0.1.0/scripts/snapshot_leg_a_eval.py +130 -0
- rag_wright-0.1.0/scripts/stack_correctness_validate.py +120 -0
- rag_wright-0.1.0/scripts/table_retrieval_smoke.py +118 -0
- rag_wright-0.1.0/scripts/train_function_classifier.py +95 -0
- rag_wright-0.1.0/scripts/train_legalbert_function.py +270 -0
- rag_wright-0.1.0/scripts/train_legalbert_modal.py +473 -0
- rag_wright-0.1.0/scripts/typed_rerank_validate.py +74 -0
- rag_wright-0.1.0/scripts/verify_wheel_install.sh +60 -0
- rag_wright-0.1.0/scripts/vllm_extraction_ab.py +111 -0
- rag_wright-0.1.0/scripts/vllm_kg_query_validate.py +83 -0
- rag_wright-0.1.0/scripts/vllm_raw_diagnostic.py +76 -0
- rag_wright-0.1.0/src/rag_wright/__init__.py +13 -0
- rag_wright-0.1.0/src/rag_wright/api/__init__.py +33 -0
- rag_wright-0.1.0/src/rag_wright/api/config.py +59 -0
- rag_wright-0.1.0/src/rag_wright/api/discover.py +70 -0
- rag_wright-0.1.0/src/rag_wright/api/documents.py +39 -0
- rag_wright-0.1.0/src/rag_wright/api/ids.py +31 -0
- rag_wright-0.1.0/src/rag_wright/api/invoke.py +99 -0
- rag_wright-0.1.0/src/rag_wright/api/kg.py +61 -0
- rag_wright-0.1.0/src/rag_wright/api/mcp.py +94 -0
- rag_wright-0.1.0/src/rag_wright/api/usage.py +30 -0
- rag_wright-0.1.0/src/rag_wright/api/workspace.py +85 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/__init__.py +8 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/answer_generator.py +427 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/ard.py +286 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/assertion_extraction.py +79 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/chunk_read.py +58 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/chunk_write.py +163 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/claim_extraction.py +153 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/clause_exception_linking.py +117 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/compliance_judgment.py +322 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/compliance_store.py +87 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/contract_kg_serve.py +156 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/contract_kg_store.py +251 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/dg_extraction.py +585 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/disambiguation.py +163 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/document_parse.py +87 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/document_scope.py +49 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/embedding.py +164 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/embedding_profiles.py +43 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/entity_resolution.py +154 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/fusion.py +64 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/graph_extraction.py +243 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/graph_query.py +73 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/graph_storage.py +111 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/highlight_serve.py +142 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/hybrid_search.py +65 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/invoke.py +31 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/jev_decision.py +38 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/manifests.py +872 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/okf_navigate.py +456 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/parsing.py +286 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/property_boosted_retrieval.py +125 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/query_function_classifier.py +94 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/query_understanding.py +109 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/registry.py +262 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/remote_encoders.py +94 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/requirement_extraction.py +247 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/reranking.py +123 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/retrieval_core.py +126 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/rlm_chunking.py +808 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/rlm_synthesis.py +316 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/scan_quality.py +136 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/span_relevance_judgment.py +191 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/vision_to_text.py +85 -0
- rag_wright-0.1.0/src/rag_wright/capabilities/vlm_ocr.py +85 -0
- rag_wright-0.1.0/src/rag_wright/contracts/__init__.py +6 -0
- rag_wright-0.1.0/src/rag_wright/contracts/chunk.py +79 -0
- rag_wright-0.1.0/src/rag_wright/contracts/compliance.py +303 -0
- rag_wright-0.1.0/src/rag_wright/contracts/contract_meta.py +27 -0
- rag_wright-0.1.0/src/rag_wright/contracts/extraction.py +130 -0
- rag_wright-0.1.0/src/rag_wright/contracts/function.py +167 -0
- rag_wright-0.1.0/src/rag_wright/contracts/function_routing.py +91 -0
- rag_wright-0.1.0/src/rag_wright/contracts/highlight.py +74 -0
- rag_wright-0.1.0/src/rag_wright/contracts/identifiers.py +153 -0
- rag_wright-0.1.0/src/rag_wright/contracts/jurisdiction.py +96 -0
- rag_wright-0.1.0/src/rag_wright/contracts/ontology.py +142 -0
- rag_wright-0.1.0/src/rag_wright/contracts/property.py +201 -0
- rag_wright-0.1.0/src/rag_wright/contracts/provenance.py +78 -0
- rag_wright-0.1.0/src/rag_wright/contracts/query_intent.py +53 -0
- rag_wright-0.1.0/src/rag_wright/contracts/span.py +76 -0
- rag_wright-0.1.0/src/rag_wright/contracts/value_match.py +84 -0
- rag_wright-0.1.0/src/rag_wright/corpus/__init__.py +0 -0
- rag_wright-0.1.0/src/rag_wright/corpus/canonicalize.py +116 -0
- rag_wright-0.1.0/src/rag_wright/corpus/cuad.py +153 -0
- rag_wright-0.1.0/src/rag_wright/corpus/cuad_ingestion.py +72 -0
- rag_wright-0.1.0/src/rag_wright/corpus/document_parser.py +299 -0
- rag_wright-0.1.0/src/rag_wright/corpus/edgar.py +231 -0
- rag_wright-0.1.0/src/rag_wright/corpus/gcs_ingestion.py +120 -0
- rag_wright-0.1.0/src/rag_wright/corpus/http.py +110 -0
- rag_wright-0.1.0/src/rag_wright/corpus/selection.py +152 -0
- rag_wright-0.1.0/src/rag_wright/mcp/__init__.py +11 -0
- rag_wright-0.1.0/src/rag_wright/mcp/compliance_server.py +299 -0
- rag_wright-0.1.0/src/rag_wright/mcp/intra_document_qa_server.py +170 -0
- rag_wright-0.1.0/src/rag_wright/mcp/relational_qa_server.py +171 -0
- rag_wright-0.1.0/src/rag_wright/mcp/session_store.py +64 -0
- rag_wright-0.1.0/src/rag_wright/mcp/typed_property_retrieval_server.py +191 -0
- rag_wright-0.1.0/src/rag_wright/models/__init__.py +8 -0
- rag_wright-0.1.0/src/rag_wright/models/profiles.py +331 -0
- rag_wright-0.1.0/src/rag_wright/models/seam.py +497 -0
- rag_wright-0.1.0/src/rag_wright/models/tag_structured.py +285 -0
- rag_wright-0.1.0/src/rag_wright/models/tracing.py +179 -0
- rag_wright-0.1.0/src/rag_wright/models/usage.py +102 -0
- rag_wright-0.1.0/src/rag_wright/okf/__init__.py +11 -0
- rag_wright-0.1.0/src/rag_wright/okf/compile.py +292 -0
- rag_wright-0.1.0/src/rag_wright/okf/document.py +47 -0
- rag_wright-0.1.0/src/rag_wright/okf/enrich.py +176 -0
- rag_wright-0.1.0/src/rag_wright/okf/links.py +190 -0
- rag_wright-0.1.0/src/rag_wright/okf/lint.py +105 -0
- rag_wright-0.1.0/src/rag_wright/ontology/__init__.py +6 -0
- rag_wright-0.1.0/src/rag_wright/ontology/_generated_template_meta.py +60 -0
- rag_wright-0.1.0/src/rag_wright/ontology/_generated_vocab.py +52 -0
- rag_wright-0.1.0/src/rag_wright/ontology/clause_template.py +964 -0
- rag_wright-0.1.0/src/rag_wright/ontology/codegen.py +84 -0
- rag_wright-0.1.0/src/rag_wright/ontology/compliance_bridge.ttl +186 -0
- rag_wright-0.1.0/src/rag_wright/ontology/contract_bridge.ttl +2685 -0
- rag_wright-0.1.0/src/rag_wright/ontology/contract_taxonomy.py +24 -0
- rag_wright-0.1.0/src/rag_wright/ontology/derive.py +58 -0
- rag_wright-0.1.0/src/rag_wright/ontology/loader.py +435 -0
- rag_wright-0.1.0/src/rag_wright/ontology/packs/ftc_16cfr255.ttl +29 -0
- rag_wright-0.1.0/src/rag_wright/ontology/registry.py +87 -0
- rag_wright-0.1.0/src/rag_wright/ontology/template_introspect.py +100 -0
- rag_wright-0.1.0/src/rag_wright/py.typed +0 -0
- rag_wright-0.1.0/src/rag_wright/reference/__init__.py +2 -0
- rag_wright-0.1.0/src/rag_wright/reference/compliance.py +41 -0
- rag_wright-0.1.0/src/rag_wright/reference/contract_seam.py +123 -0
- rag_wright-0.1.0/src/rag_wright/skills/__init__.py +7 -0
- rag_wright-0.1.0/src/rag_wright/skills/claim_extraction/SKILL.md +47 -0
- rag_wright-0.1.0/src/rag_wright/skills/claim_extraction/__init__.py +1 -0
- rag_wright-0.1.0/src/rag_wright/skills/claim_extraction/template.py +50 -0
- rag_wright-0.1.0/src/rag_wright/skills/compliance_judgment/SKILL.md +59 -0
- rag_wright-0.1.0/src/rag_wright/skills/corpus_ingest/SKILL.md +106 -0
- rag_wright-0.1.0/src/rag_wright/skills/extraction_semantic_judge/SKILL.md +51 -0
- rag_wright-0.1.0/src/rag_wright/skills/extraction_semantic_judge/__init__.py +1 -0
- rag_wright-0.1.0/src/rag_wright/skills/generation/SKILL.md +64 -0
- rag_wright-0.1.0/src/rag_wright/skills/generation/__init__.py +1 -0
- rag_wright-0.1.0/src/rag_wright/skills/generic_compliance_judgment/SKILL.md +58 -0
- rag_wright-0.1.0/src/rag_wright/skills/okf_navigate/SKILL.md +137 -0
- rag_wright-0.1.0/src/rag_wright/skills/requirement_extraction/SKILL.md +47 -0
- rag_wright-0.1.0/src/rag_wright/skills/requirement_extraction/__init__.py +1 -0
- rag_wright-0.1.0/src/rag_wright/skills/requirement_extraction/template.py +50 -0
- rag_wright-0.1.0/src/rag_wright/skills/rlm/SKILL.md +186 -0
- rag_wright-0.1.0/src/rag_wright/skills/rlm/__init__.py +31 -0
- rag_wright-0.1.0/src/rag_wright/skills/rlm/agent.py +292 -0
- rag_wright-0.1.0/src/rag_wright/skills/span_relevance_judgment/SKILL.md +67 -0
- rag_wright-0.1.0/src/rag_wright/skills/vision_to_text/SKILL.md +36 -0
- rag_wright-0.1.0/src/rag_wright/skills/vision_to_text/__init__.py +1 -0
- rag_wright-0.1.0/src/rag_wright/spans/__init__.py +1 -0
- rag_wright-0.1.0/src/rag_wright/spans/boundary.py +78 -0
- rag_wright-0.1.0/src/rag_wright/spans/clause_function_classifier.py +490 -0
- rag_wright-0.1.0/src/rag_wright/spans/clause_kg_extractor.py +337 -0
- rag_wright-0.1.0/src/rag_wright/spans/cuad_labels.py +81 -0
- rag_wright-0.1.0/src/rag_wright/spans/dim_classifier.py +158 -0
- rag_wright-0.1.0/src/rag_wright/spans/dim_fleet.json +411 -0
- rag_wright-0.1.0/src/rag_wright/spans/function_classifier.py +77 -0
- rag_wright-0.1.0/src/rag_wright/spans/function_families.py +62 -0
- rag_wright-0.1.0/src/rag_wright/spans/hybrid_classifier.py +103 -0
- rag_wright-0.1.0/src/rag_wright/spans/legalbert_classifier.py +83 -0
- rag_wright-0.1.0/src/rag_wright/spans/model_capabilities.py +107 -0
- rag_wright-0.1.0/src/rag_wright/spans/new_function_labels.py +111 -0
- rag_wright-0.1.0/src/rag_wright/spans/page_map.py +68 -0
- rag_wright-0.1.0/src/rag_wright/spans/property_extractor.py +365 -0
- rag_wright-0.1.0/src/rag_wright/spans/property_grounding.py +182 -0
- rag_wright-0.1.0/src/rag_wright/spans/reclassify.py +77 -0
- rag_wright-0.1.0/src/rag_wright/spans/scarce_function_labels.py +105 -0
- rag_wright-0.1.0/src/rag_wright/spans/segment.py +341 -0
- rag_wright-0.1.0/src/rag_wright/spans/semantic_judge.py +197 -0
- rag_wright-0.1.0/src/rag_wright/spans/symbolic_validation.py +131 -0
- rag_wright-0.1.0/src/rag_wright/spans/tag_clause_extractor.py +182 -0
- rag_wright-0.1.0/src/rag_wright/store/__init__.py +6 -0
- rag_wright-0.1.0/src/rag_wright/store/arcadedb.py +1135 -0
- rag_wright-0.1.0/src/rag_wright/store/chunk_text.py +66 -0
- rag_wright-0.1.0/src/rag_wright/store/seam.py +213 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/__init__.py +0 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/async_ingestion.py +204 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/compliance_check.py +1042 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/compliance_ingestion.py +306 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/contract_ingestion_pipeline.py +999 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/graph_extraction.py +102 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/intra_document_qa.py +328 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/observability.py +140 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/query_constraint_extraction.py +73 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/relational_qa.py +165 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/requirement_extraction.py +137 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/scaffold.py +65 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/semantic_chunking.py +183 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/typed_clause_extraction.py +172 -0
- rag_wright-0.1.0/src/rag_wright/subgraphs/typed_property_retrieval.py +278 -0
- rag_wright-0.1.0/src/rag_wright/util/__init__.py +1 -0
- rag_wright-0.1.0/src/rag_wright/util/concurrent.py +153 -0
- rag_wright-0.1.0/src/rag_wright/util/spacy_model.py +45 -0
- rag_wright-0.1.0/tasks.md +5053 -0
- rag_wright-0.1.0/tests/__init__.py +0 -0
- rag_wright-0.1.0/tests/api/__init__.py +0 -0
- rag_wright-0.1.0/tests/api/test_capability_reexports.py +30 -0
- rag_wright-0.1.0/tests/api/test_discover.py +92 -0
- rag_wright-0.1.0/tests/api/test_documents.py +66 -0
- rag_wright-0.1.0/tests/api/test_e2e.py +94 -0
- rag_wright-0.1.0/tests/api/test_invoke.py +311 -0
- rag_wright-0.1.0/tests/api/test_kg.py +90 -0
- rag_wright-0.1.0/tests/api/test_mcp.py +108 -0
- rag_wright-0.1.0/tests/api/test_options.py +84 -0
- rag_wright-0.1.0/tests/api/test_usage.py +58 -0
- rag_wright-0.1.0/tests/api/test_workspace.py +84 -0
- rag_wright-0.1.0/tests/arch/test_import_contracts.py +32 -0
- rag_wright-0.1.0/tests/async_helpers.py +65 -0
- rag_wright-0.1.0/tests/capabilities/__init__.py +0 -0
- rag_wright-0.1.0/tests/capabilities/_fixtures/rlm_probe_skill/SKILL.md +8 -0
- rag_wright-0.1.0/tests/capabilities/test_answer_generator.py +521 -0
- rag_wright-0.1.0/tests/capabilities/test_answer_generator_async.py +65 -0
- rag_wright-0.1.0/tests/capabilities/test_assertion_extraction.py +62 -0
- rag_wright-0.1.0/tests/capabilities/test_authoring_contract.py +58 -0
- rag_wright-0.1.0/tests/capabilities/test_chunk_read.py +56 -0
- rag_wright-0.1.0/tests/capabilities/test_chunk_write.py +249 -0
- rag_wright-0.1.0/tests/capabilities/test_claim_extraction.py +151 -0
- rag_wright-0.1.0/tests/capabilities/test_clause_exception_linking.py +103 -0
- rag_wright-0.1.0/tests/capabilities/test_compliance_judgment.py +356 -0
- rag_wright-0.1.0/tests/capabilities/test_compliance_store.py +123 -0
- rag_wright-0.1.0/tests/capabilities/test_contract_kg_serve.py +193 -0
- rag_wright-0.1.0/tests/capabilities/test_contract_kg_store.py +138 -0
- rag_wright-0.1.0/tests/capabilities/test_contract_kg_store_reads.py +132 -0
- rag_wright-0.1.0/tests/capabilities/test_contract_taxonomy_and_spans.py +102 -0
- rag_wright-0.1.0/tests/capabilities/test_dg_adapter.py +62 -0
- rag_wright-0.1.0/tests/capabilities/test_dg_async.py +94 -0
- rag_wright-0.1.0/tests/capabilities/test_dg_extraction.py +182 -0
- rag_wright-0.1.0/tests/capabilities/test_dg_model_seam.py +43 -0
- rag_wright-0.1.0/tests/capabilities/test_dg_private.py +55 -0
- rag_wright-0.1.0/tests/capabilities/test_disambiguation.py +171 -0
- rag_wright-0.1.0/tests/capabilities/test_document_scope.py +70 -0
- rag_wright-0.1.0/tests/capabilities/test_embedding.py +177 -0
- rag_wright-0.1.0/tests/capabilities/test_embedding_profiles.py +26 -0
- rag_wright-0.1.0/tests/capabilities/test_entity_resolution.py +189 -0
- rag_wright-0.1.0/tests/capabilities/test_fusion.py +63 -0
- rag_wright-0.1.0/tests/capabilities/test_graph_extraction.py +186 -0
- rag_wright-0.1.0/tests/capabilities/test_graph_query.py +127 -0
- rag_wright-0.1.0/tests/capabilities/test_graph_storage.py +159 -0
- rag_wright-0.1.0/tests/capabilities/test_highlight_serve.py +118 -0
- rag_wright-0.1.0/tests/capabilities/test_hybrid_search.py +183 -0
- rag_wright-0.1.0/tests/capabilities/test_jev_decision.py +119 -0
- rag_wright-0.1.0/tests/capabilities/test_manifests.py +288 -0
- rag_wright-0.1.0/tests/capabilities/test_okf_compile.py +184 -0
- rag_wright-0.1.0/tests/capabilities/test_okf_links.py +92 -0
- rag_wright-0.1.0/tests/capabilities/test_okf_navigate.py +129 -0
- rag_wright-0.1.0/tests/capabilities/test_parsing.py +148 -0
- rag_wright-0.1.0/tests/capabilities/test_parties_extraction.py +71 -0
- rag_wright-0.1.0/tests/capabilities/test_property_boosted_retrieval.py +128 -0
- rag_wright-0.1.0/tests/capabilities/test_query_function_classifier.py +43 -0
- rag_wright-0.1.0/tests/capabilities/test_query_understanding.py +117 -0
- rag_wright-0.1.0/tests/capabilities/test_registry.py +283 -0
- rag_wright-0.1.0/tests/capabilities/test_remote_encoders.py +68 -0
- rag_wright-0.1.0/tests/capabilities/test_repair_partition.py +158 -0
- rag_wright-0.1.0/tests/capabilities/test_requirement_extraction.py +262 -0
- rag_wright-0.1.0/tests/capabilities/test_reranking.py +162 -0
- rag_wright-0.1.0/tests/capabilities/test_retrieval_core.py +113 -0
- rag_wright-0.1.0/tests/capabilities/test_rlm_chunking.py +624 -0
- rag_wright-0.1.0/tests/capabilities/test_rlm_chunking_async.py +53 -0
- rag_wright-0.1.0/tests/capabilities/test_rlm_method.py +504 -0
- rag_wright-0.1.0/tests/capabilities/test_rlm_synthesis.py +317 -0
- rag_wright-0.1.0/tests/capabilities/test_runtime_registry.py +32 -0
- rag_wright-0.1.0/tests/capabilities/test_scan_quality.py +90 -0
- rag_wright-0.1.0/tests/capabilities/test_span_relevance_judgment.py +116 -0
- rag_wright-0.1.0/tests/capabilities/test_structural_boundary_discoverer.py +75 -0
- rag_wright-0.1.0/tests/capabilities/test_structural_model_fallback.py +105 -0
- rag_wright-0.1.0/tests/capabilities/test_tag_boundary_discoverer.py +78 -0
- rag_wright-0.1.0/tests/capabilities/test_tiered_ocr.py +241 -0
- rag_wright-0.1.0/tests/capabilities/test_vlm_ocr.py +54 -0
- rag_wright-0.1.0/tests/conftest.py +8 -0
- rag_wright-0.1.0/tests/contracts/__init__.py +0 -0
- rag_wright-0.1.0/tests/contracts/test_canonical_source_doc_id.py +53 -0
- rag_wright-0.1.0/tests/contracts/test_chunk_record.py +141 -0
- rag_wright-0.1.0/tests/contracts/test_compliance.py +164 -0
- rag_wright-0.1.0/tests/contracts/test_cuad_highlight_contracts.py +109 -0
- rag_wright-0.1.0/tests/contracts/test_extraction.py +171 -0
- rag_wright-0.1.0/tests/contracts/test_function.py +119 -0
- rag_wright-0.1.0/tests/contracts/test_function_routing.py +69 -0
- rag_wright-0.1.0/tests/contracts/test_identifiers.py +161 -0
- rag_wright-0.1.0/tests/contracts/test_jurisdiction.py +76 -0
- rag_wright-0.1.0/tests/contracts/test_ontology.py +167 -0
- rag_wright-0.1.0/tests/contracts/test_property.py +107 -0
- rag_wright-0.1.0/tests/contracts/test_provenance.py +119 -0
- rag_wright-0.1.0/tests/contracts/test_value_match.py +43 -0
- rag_wright-0.1.0/tests/corpus/__init__.py +0 -0
- rag_wright-0.1.0/tests/corpus/test_canonicalize.py +96 -0
- rag_wright-0.1.0/tests/corpus/test_cuad.py +45 -0
- rag_wright-0.1.0/tests/corpus/test_cuad_ingestion.py +38 -0
- rag_wright-0.1.0/tests/corpus/test_document_parser.py +312 -0
- rag_wright-0.1.0/tests/corpus/test_edgar.py +136 -0
- rag_wright-0.1.0/tests/corpus/test_gcs_ingestion.py +138 -0
- rag_wright-0.1.0/tests/corpus/test_http.py +121 -0
- rag_wright-0.1.0/tests/corpus/test_parse_wire2.py +50 -0
- rag_wright-0.1.0/tests/corpus/test_selection.py +133 -0
- rag_wright-0.1.0/tests/eval/test_contractnli_judge.py +17 -0
- rag_wright-0.1.0/tests/eval/test_cuad_highlight_metrics.py +43 -0
- rag_wright-0.1.0/tests/eval/test_kg_property_rerank.py +30 -0
- rag_wright-0.1.0/tests/eval/test_nl_to_type_metrics.py +27 -0
- rag_wright-0.1.0/tests/eval/test_relational_eval.py +72 -0
- rag_wright-0.1.0/tests/fixtures/leg_a_silver/README.md +41 -0
- rag_wright-0.1.0/tests/fixtures/leg_a_silver/evidence_snapshot.json +1408 -0
- rag_wright-0.1.0/tests/fixtures/table-bearing-contract.pdf +74 -0
- rag_wright-0.1.0/tests/foundation/__init__.py +0 -0
- rag_wright-0.1.0/tests/foundation/test_arcadedb_hybrid.py +106 -0
- rag_wright-0.1.0/tests/foundation/test_model_seam_structured.py +50 -0
- rag_wright-0.1.0/tests/journey/incidents_domain.py +84 -0
- rag_wright-0.1.0/tests/journey/incidents_pack.ttl +14 -0
- rag_wright-0.1.0/tests/journey/test_incidents_journey.py +81 -0
- rag_wright-0.1.0/tests/journey/test_pack_schema.py +79 -0
- rag_wright-0.1.0/tests/mcp/__init__.py +0 -0
- rag_wright-0.1.0/tests/mcp/test_compliance_server.py +200 -0
- rag_wright-0.1.0/tests/mcp/test_intra_document_qa_server.py +64 -0
- rag_wright-0.1.0/tests/mcp/test_no_model_supplied_tenant.py +116 -0
- rag_wright-0.1.0/tests/mcp/test_relational_qa_server.py +63 -0
- rag_wright-0.1.0/tests/mcp/test_typed_property_retrieval_server.py +74 -0
- rag_wright-0.1.0/tests/models/__init__.py +0 -0
- rag_wright-0.1.0/tests/models/test_async_infra.py +58 -0
- rag_wright-0.1.0/tests/models/test_profile_routing.py +77 -0
- rag_wright-0.1.0/tests/models/test_profile_seam.py +264 -0
- rag_wright-0.1.0/tests/models/test_seam_async.py +92 -0
- rag_wright-0.1.0/tests/models/test_seam_retry.py +74 -0
- rag_wright-0.1.0/tests/models/test_seam_stream.py +189 -0
- rag_wright-0.1.0/tests/models/test_serving_seam.py +121 -0
- rag_wright-0.1.0/tests/models/test_serving_wiring.py +82 -0
- rag_wright-0.1.0/tests/models/test_tag_structured.py +320 -0
- rag_wright-0.1.0/tests/models/test_tag_structured_async.py +47 -0
- rag_wright-0.1.0/tests/models/test_tracing.py +134 -0
- rag_wright-0.1.0/tests/models/test_usage.py +87 -0
- rag_wright-0.1.0/tests/ontology/__init__.py +0 -0
- rag_wright-0.1.0/tests/ontology/test_clause_template.py +171 -0
- rag_wright-0.1.0/tests/ontology/test_compliance_ontology_authoritative.py +92 -0
- rag_wright-0.1.0/tests/ontology/test_derivation.py +120 -0
- rag_wright-0.1.0/tests/ontology/test_entity_taxonomy.py +33 -0
- rag_wright-0.1.0/tests/ontology/test_generated_template_meta_in_sync.py +21 -0
- rag_wright-0.1.0/tests/ontology/test_generated_vocab_in_sync.py +23 -0
- rag_wright-0.1.0/tests/ontology/test_template_captured_in_ttl.py +22 -0
- rag_wright-0.1.0/tests/ontology/test_ttl_is_source_of_truth.py +33 -0
- rag_wright-0.1.0/tests/reference/test_compliance_reference.py +146 -0
- rag_wright-0.1.0/tests/reference/test_contract_seam.py +116 -0
- rag_wright-0.1.0/tests/scripts/test_eval_chunking_ab.py +63 -0
- rag_wright-0.1.0/tests/scripts/test_eval_classifier_ab.py +38 -0
- rag_wright-0.1.0/tests/spans/test_boundary.py +35 -0
- rag_wright-0.1.0/tests/spans/test_classifier_property_extractor.py +96 -0
- rag_wright-0.1.0/tests/spans/test_clause_classifier_tags_0005.py +54 -0
- rag_wright-0.1.0/tests/spans/test_clause_function_classifier.py +185 -0
- rag_wright-0.1.0/tests/spans/test_clause_function_classifier_async.py +78 -0
- rag_wright-0.1.0/tests/spans/test_clause_kg_extractor.py +234 -0
- rag_wright-0.1.0/tests/spans/test_clause_kg_extractor_async.py +70 -0
- rag_wright-0.1.0/tests/spans/test_dim_classifier.py +43 -0
- rag_wright-0.1.0/tests/spans/test_dim_fleet_live.py +147 -0
- rag_wright-0.1.0/tests/spans/test_function_classifier.py +111 -0
- rag_wright-0.1.0/tests/spans/test_hybrid_classifier.py +86 -0
- rag_wright-0.1.0/tests/spans/test_hybrid_property_extractor.py +150 -0
- rag_wright-0.1.0/tests/spans/test_legalbert_classifier.py +56 -0
- rag_wright-0.1.0/tests/spans/test_model_capabilities.py +51 -0
- rag_wright-0.1.0/tests/spans/test_new_function_labels.py +51 -0
- rag_wright-0.1.0/tests/spans/test_page_map.py +79 -0
- rag_wright-0.1.0/tests/spans/test_property_extractor.py +108 -0
- rag_wright-0.1.0/tests/spans/test_property_grounding.py +108 -0
- rag_wright-0.1.0/tests/spans/test_reclassify.py +59 -0
- rag_wright-0.1.0/tests/spans/test_scarce_function_labels.py +64 -0
- rag_wright-0.1.0/tests/spans/test_segment.py +281 -0
- rag_wright-0.1.0/tests/spans/test_semantic_judge.py +143 -0
- rag_wright-0.1.0/tests/spans/test_setfit_clause_adapter.py +70 -0
- rag_wright-0.1.0/tests/spans/test_span_offsets.py +54 -0
- rag_wright-0.1.0/tests/spans/test_stage_labels_0005.py +41 -0
- rag_wright-0.1.0/tests/spans/test_symbolic_validation.py +197 -0
- rag_wright-0.1.0/tests/spans/test_tag_clause_extractor.py +140 -0
- rag_wright-0.1.0/tests/store/__init__.py +0 -0
- rag_wright-0.1.0/tests/store/test_affiliation_backfill.py +52 -0
- rag_wright-0.1.0/tests/store/test_arcadedb_clause_kg.py +175 -0
- rag_wright-0.1.0/tests/store/test_arcadedb_contract.py +72 -0
- rag_wright-0.1.0/tests/store/test_arcadedb_property.py +107 -0
- rag_wright-0.1.0/tests/store/test_arcadedb_requirement_sources.py +97 -0
- rag_wright-0.1.0/tests/store/test_arcadedb_schema.py +224 -0
- rag_wright-0.1.0/tests/store/test_arcadedb_span.py +107 -0
- rag_wright-0.1.0/tests/store/test_chunk_text.py +89 -0
- rag_wright-0.1.0/tests/store/test_document_scope.py +150 -0
- rag_wright-0.1.0/tests/store/test_engine_domain_neutral.py +48 -0
- rag_wright-0.1.0/tests/store/test_entities_by_name.py +60 -0
- rag_wright-0.1.0/tests/store/test_kg_edges.py +124 -0
- rag_wright-0.1.0/tests/store/test_kg_read.py +107 -0
- rag_wright-0.1.0/tests/store/test_kg_write.py +99 -0
- rag_wright-0.1.0/tests/store/test_write_graph_idempotent.py +112 -0
- rag_wright-0.1.0/tests/subgraphs/__init__.py +0 -0
- rag_wright-0.1.0/tests/subgraphs/test_async_ingestion.py +256 -0
- rag_wright-0.1.0/tests/subgraphs/test_chunk7_structure_carry.py +73 -0
- rag_wright-0.1.0/tests/subgraphs/test_compliance_check.py +1509 -0
- rag_wright-0.1.0/tests/subgraphs/test_compliance_ingestion.py +333 -0
- rag_wright-0.1.0/tests/subgraphs/test_contract_ingestion_pipeline.py +300 -0
- rag_wright-0.1.0/tests/subgraphs/test_contract_ingestion_pipeline_async.py +142 -0
- rag_wright-0.1.0/tests/subgraphs/test_contract_ingestion_pipeline_graph_async.py +223 -0
- rag_wright-0.1.0/tests/subgraphs/test_graph_extraction.py +60 -0
- rag_wright-0.1.0/tests/subgraphs/test_ingest_knobs.py +32 -0
- rag_wright-0.1.0/tests/subgraphs/test_ingest_segment_classify.py +121 -0
- rag_wright-0.1.0/tests/subgraphs/test_intra_document_qa.py +395 -0
- rag_wright-0.1.0/tests/subgraphs/test_observability.py +41 -0
- rag_wright-0.1.0/tests/subgraphs/test_partial_entry_contract.py +81 -0
- rag_wright-0.1.0/tests/subgraphs/test_query_constraint_extraction.py +47 -0
- rag_wright-0.1.0/tests/subgraphs/test_relational_qa.py +128 -0
- rag_wright-0.1.0/tests/subgraphs/test_requirement_extraction.py +127 -0
- rag_wright-0.1.0/tests/subgraphs/test_scaffold.py +82 -0
- rag_wright-0.1.0/tests/subgraphs/test_semantic_chunking.py +106 -0
- rag_wright-0.1.0/tests/subgraphs/test_typed_clause_extraction.py +117 -0
- rag_wright-0.1.0/tests/subgraphs/test_typed_property_retrieval.py +244 -0
- rag_wright-0.1.0/tests/test_backfill_clause_span_id.py +43 -0
- rag_wright-0.1.0/tests/test_dg_model_ab.py +37 -0
- rag_wright-0.1.0/tests/test_populate_clause_kg.py +70 -0
- rag_wright-0.1.0/tests/test_populate_entity_graph.py +43 -0
- rag_wright-0.1.0/tests/util/__init__.py +0 -0
- rag_wright-0.1.0/tests/util/test_concurrent.py +67 -0
- rag_wright-0.1.0/tests/util/test_spacy_model.py +53 -0
- rag_wright-0.1.0/uv.lock +4848 -0
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: authoring-a-capability
|
|
3
|
+
description: >-
|
|
4
|
+
How to author a new RAG_Wright engine capability of any kind (subgraph, function, model, agent_skill, mcp_tool)
|
|
5
|
+
so it is registered, ARD-discoverable, and invokable by name through the engine API. Use it whenever you add a
|
|
6
|
+
new capability or a new-domain product/graph needs one: it gives the shared registration + ARD + invocation
|
|
7
|
+
contract (the four surfaces + the definition of done), the per-kind implementation specifics, and the
|
|
8
|
+
conformance guardrail that keeps the catalog honest. Grounded against the real code; keep it in step with it.
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
# Authoring a capability
|
|
12
|
+
|
|
13
|
+
A **capability** is a named, ARD-registered unit of engine behavior (FR-C). Every capability has a `kind`
|
|
14
|
+
(`ard.py::EntryKind`): `subgraph | function | model | agent_skill | mcp_tool` (`dagster_asset` is reserved). The
|
|
15
|
+
contract below is the SAME for every kind; only the implementation differs. Capabilities compose — a subgraph calls
|
|
16
|
+
functions/models; a product invokes a capability by name through `rag_wright.api`.
|
|
17
|
+
|
|
18
|
+
**Ground every call before writing it** (CLAUDE.md library rule). The authoritative sources this skill summarizes:
|
|
19
|
+
`capabilities/registry.py` (the canonical-slug set + `CapabilityRegistry.register`), `capabilities/manifests.py`
|
|
20
|
+
(`CapabilityManifest` + `_SPECS`/`MANIFEST_SPECS`), `capabilities/ard.py` (`EntryKind`, `MEDIA_TYPE_BY_KIND`,
|
|
21
|
+
`CALLABLE_KINDS`), `scripts/publish_manifests.py`, `api/invoke.py` + `capabilities/invoke.py::capability_impl` (the adapter-free impl_ref invoker + drift guard),
|
|
22
|
+
`api/mcp.py` (generic MCP exposure). The guardrail test is `tests/capabilities/test_authoring_contract.py`.
|
|
23
|
+
|
|
24
|
+
## The four surfaces (the definition of done)
|
|
25
|
+
|
|
26
|
+
A finished capability touches these: 1 (implementation) is for EVERY kind; 2 (the invoke factory + `impl_ref`) is
|
|
27
|
+
for invokable kinds (`subgraph`/`model`); 3 (manifest + register) is for every discoverable kind; 4 (invocable +
|
|
28
|
+
MCP) follows automatically for invokable kinds; 4b (a bespoke MCP server) is optional.
|
|
29
|
+
|
|
30
|
+
1. **Implementation** — the real code, in that kind's home (see per-kind below).
|
|
31
|
+
2. **The invoke factory + `impl_ref`** (invokable kinds: subgraph/model) — write a co-located
|
|
32
|
+
`async def ainvoke(resources, inputs)` (subgraph) / `def <name>(resources, inputs)` (model) in the capability's
|
|
33
|
+
own module (model factories ignore `resources`; they build over the opaque `WorkspaceHandle` — `resources._store`,
|
|
34
|
+
`resources.model_id(role)`, never env). The manifest's `impl_ref="module:attr"` points to it. The invoker imports
|
|
35
|
+
it LAZILY and calls it — there is **NO central adapter dict** (EP-CORE-2). A plain function/agent_skill/mcp_tool
|
|
36
|
+
declares no `impl_ref`.
|
|
37
|
+
3. **ARD manifest + register it** — a `CapabilityManifest(slug, kind, display_name, description,
|
|
38
|
+
representative_queries=(2-5…), tags=…, impl_ref=…)`. **The catalog ships EMPTY (EP-CORE-3):** call
|
|
39
|
+
`register_capability(manifest)` at runtime to add it (a product registers its own; the engine's reference pack is
|
|
40
|
+
opt-in via `load_reference_pack()`). `representative_queries` is the field ARD discovery ranks on — write real,
|
|
41
|
+
specific queries. For the ENGINE's reference pack, the manifest is committed in `manifests.py::_SPECS` and the
|
|
42
|
+
slug is in `CANONICAL_CAPABILITY_SLUGS`; a downstream product registers freely (its slugs need not be canonical).
|
|
43
|
+
Publish to `~/.air/registry` (what GraphWright's store loads) with `uv run python scripts/publish_manifests.py`.
|
|
44
|
+
Callable kinds get `ResponseBounds` (defaulted); `agent_skill` must NOT declare bounds (loaded, not called).
|
|
45
|
+
4. **Invocable + MCP for free** — once registered with an `impl_ref`, the capability is callable as
|
|
46
|
+
`ainvoke_subgraph(slug, inputs, resources=ws)` / `invoke_model(slug, inputs, resources=ws)` — the invoker resolves
|
|
47
|
+
the `impl_ref` via `capabilities.invoke.capability_impl` (the drift guard asserts it resolves to a callable of the
|
|
48
|
+
declared kind) — AND exposable over MCP (surface 4b), with **zero engine edits**.
|
|
49
|
+
|
|
50
|
+
4b. **MCP exposure** (optional) — any invokable capability is already an MCP tool with zero extra code via
|
|
51
|
+
`api/mcp.py::build_capability_mcp(slug, resources=ws)` (EP-RT-2). Write a bespoke `mcp/<slug>_server.py` only
|
|
52
|
+
when you want a CURATED, typed tool signature instead of the generic opaque-`inputs` surface.
|
|
53
|
+
|
|
54
|
+
## Per-kind specifics
|
|
55
|
+
|
|
56
|
+
### subgraph — a compiled LangGraph `StateGraph`
|
|
57
|
+
- **Home:** `subgraphs/<slug>.py`. A `production_<slug>(*, store, ...) -> CompiledGraph` builder: `g = StateGraph(_State)`,
|
|
58
|
+
add nodes/edges with `START`/`END`, `return g.compile()`. Nodes call functions/models (compose).
|
|
59
|
+
- **Invoke:** the co-located `async def ainvoke(resources, inputs)` factory (impl_ref target) builds + awaits the graph.
|
|
60
|
+
- Retry/dead-letter come from the graph scaffold, not the invoker. The output contract is the registered `contract`.
|
|
61
|
+
|
|
62
|
+
### function — a plain, typed callable
|
|
63
|
+
- **Home:** `capabilities/<slug>.py`. A deterministic or model-backed callable with a Pydantic in/out contract.
|
|
64
|
+
- Invoker adapters for `function` are not wired yet (EP-API-2c); until then functions are composed inside
|
|
65
|
+
subgraphs, not invoked standalone through the API. Still register + manifest it.
|
|
66
|
+
|
|
67
|
+
### model — a trained checkpoint behind a seam
|
|
68
|
+
- **Home:** `spans/` or `capabilities/` wrapping the checkpoint (e.g. the SetFit clause classifier; the 29-dim
|
|
69
|
+
property fleet via `spans/property_extractor.py`). Load the checkpoint ONCE and cache it (the fleet is heavy) —
|
|
70
|
+
see `spans/model_capabilities.py::_dim_registry`.
|
|
71
|
+
- Serve behind the existing seam/adapter so nothing upstream changes (to FIND where a model cap belongs, use the
|
|
72
|
+
`classifier-opportunity-analysis` skill; to BUILD/train + checkpoint + serve it, the `setfit` skill). The impl_ref
|
|
73
|
+
factory is `def <slug>(resources, inputs)` for a SYNC impl (CPU-bound local inference — a classifier/XGBoost
|
|
74
|
+
checkpoint; `resources` ignored) or `async def <slug>(resources, inputs)` for an ASYNC impl (I/O-bound — an
|
|
75
|
+
LLM-backed model cap calling OpenRouter / a local vLLM client). `invoke_model` runs a sync impl and REFUSES an
|
|
76
|
+
async one; `ainvoke_model` (EP-API-7) off-loads a sync impl with `asyncio.to_thread` and awaits an async impl
|
|
77
|
+
directly, with an optional `sem` for fan-out backpressure.
|
|
78
|
+
|
|
79
|
+
### agent_skill — authored SKILL.md + the Deep Agents runtime
|
|
80
|
+
- **Home:** `skills/<slug>/SKILL.md` + the agent runtime (e.g. `skills/rlm/`). It is LOADED (progressive
|
|
81
|
+
disclosure), not called: no `ResponseBounds`. Declare `requires=(...)` for a closure over other skills and
|
|
82
|
+
`skill_runtime` for its intrinsic runtime.
|
|
83
|
+
|
|
84
|
+
### mcp_tool — a capability exposed over MCP
|
|
85
|
+
- A distinct ARD identity (`<slug>_mcp`) for the same underlying capability exposed as a cross-agent MCP tool.
|
|
86
|
+
Prefer the generic `build_capability_mcp` (surface 4b); author a bespoke `mcp/<slug>_server.py` only for a
|
|
87
|
+
curated typed signature. Bind the store server-side (issue 0035) — the tool never takes a tenant/store argument.
|
|
88
|
+
|
|
89
|
+
## Verify (the guardrail)
|
|
90
|
+
|
|
91
|
+
Run `uv run pytest tests/capabilities/test_authoring_contract.py tests/capabilities/test_manifests.py
|
|
92
|
+
tests/capabilities/test_registry.py` after authoring. It pins the contract this skill teaches: no manifest under a
|
|
93
|
+
non-canonical slug; the reserved-without-manifest set is a fixed allowlist (so adding a slug but forgetting its
|
|
94
|
+
manifest FAILS here); every manifest kind is a real ARD kind; every invokable cap's impl_ref resolves to a callable whose
|
|
95
|
+
manifest declares the matching kind. If you deliberately add a reserved/internal slug (no manifest), add it to
|
|
96
|
+
`_RESERVED_WITHOUT_MANIFEST` with a one-line reason.
|
|
97
|
+
|
|
98
|
+
## Common mistakes
|
|
99
|
+
|
|
100
|
+
- Adding the slug but forgetting the manifest (slug becomes silently un-discoverable) — the guardrail catches it.
|
|
101
|
+
- An impl_ref factory that reaches env/globals instead of the `WorkspaceHandle` — breaks multi-workspace use; build
|
|
102
|
+
everything from `h`.
|
|
103
|
+
- A non-lazy import at the top of an adapter — inflates the light index / `import rag_wright.api`; import inside
|
|
104
|
+
the adapter body.
|
|
105
|
+
- Declaring `response_bounds` on an `agent_skill`, or omitting the output `contract` on registration.
|
|
106
|
+
- Inventing a kind. If a capability fits none of the five, flag it — do not force-fit.
|
|
@@ -0,0 +1,175 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: classifier-opportunity-analysis
|
|
3
|
+
description: >-
|
|
4
|
+
Structured guide for ANALYZING a domain's ingestion + retrieval pipeline to find where an LLM call can be
|
|
5
|
+
replaced by a deterministic rule, a trained classifier, or a routing decision. Use it BEFORE building or
|
|
6
|
+
refactoring a domain pack, when an LLM is doing per-unit work that multiplies over a document, or when
|
|
7
|
+
onboarding a new domain — it is the identification/decision step upstream of `setfit` (which BUILDS the
|
|
8
|
+
classifier) and `authoring-a-capability` (which REGISTERS it as a capability). It captures the recipe applied
|
|
9
|
+
twice (contracts, then compliance) so the next domain is mapped the same way instead of re-derived.
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
# Finding classifier / routing opportunities in a pipeline
|
|
13
|
+
|
|
14
|
+
This is an **analysis** skill: its output is a decision — an *opportunity list / decomposition plan*, not code and
|
|
15
|
+
not a trained model. It answers "where in this domain's ingestion and retrieval does a classifier, a routing
|
|
16
|
+
decision, or a deterministic rule belong, and where must the LLM stay?" Build what it identifies with the `setfit`
|
|
17
|
+
skill, serve the teacher with `qwen-vllm-modal`, and register each result as a capability with
|
|
18
|
+
`authoring-a-capability`.
|
|
19
|
+
|
|
20
|
+
## The core pattern (why this works, and why we have done it twice)
|
|
21
|
+
|
|
22
|
+
A per-unit "do everything" LLM extraction is almost never one decision. It is a **bundle of separable decisions**
|
|
23
|
+
wearing one prompt. Decomposed, most of the bundle is not LLM-shaped work:
|
|
24
|
+
|
|
25
|
+
- **boundaries** (where does a unit start / is this span worth extracting) are usually structural → deterministic;
|
|
26
|
+
- **closed-vocab tags** (the type, the role, the dimension values) are classification → a classifier, or a
|
|
27
|
+
deterministic cue-rule when the cues are enumerable;
|
|
28
|
+
- only the **genuinely open part** (numbers, free text, synthesis) needs an LLM, and then only **one residual call
|
|
29
|
+
per unit**.
|
|
30
|
+
|
|
31
|
+
Worked precedent in this engine:
|
|
32
|
+
|
|
33
|
+
| | Contracts (decomposed) | Compliance (the CIC arc) |
|
|
34
|
+
|---|---|---|
|
|
35
|
+
| Unit boundaries | deterministic (section numbering / headings) + a soft type classifier | deterministic sub-section split + a span-level operative-cue gate |
|
|
36
|
+
| Closed-vocab tags | the 21-dim / 29-dim classifier fleet | deontic cue-rule + actor / claim-type / applicability classifiers |
|
|
37
|
+
| Open / numeric field | 1 residual LLM call / provision (7 fields) | 1 residual LLM call / section (evidence standard) |
|
|
38
|
+
| Text of the record | verbatim span | verbatim span (was a paraphrase) |
|
|
39
|
+
| Structure / graph extraction | once per document (parties), keep the LLM | once per document, keep the LLM |
|
|
40
|
+
|
|
41
|
+
The pattern is domain-independent. What changes per domain is the vocabulary and the document structure — which is
|
|
42
|
+
exactly what the phases below make you look at.
|
|
43
|
+
|
|
44
|
+
## Phase A — Map the pipeline as per-unit decisions
|
|
45
|
+
|
|
46
|
+
Enumerate every LLM call and, for each, its **unit** and how the unit COUNT scales:
|
|
47
|
+
|
|
48
|
+
- per **document** (parse, a party/graph-structure pass) — count ≈ corpus size; cheap per doc.
|
|
49
|
+
- per **section / chunk / segment / span** — count scales with **document length**. A 100-page document is
|
|
50
|
+
thousands of these. **This is where the cost lives and where decomposition pays.**
|
|
51
|
+
- per **query** / per **candidate** / per **(claim, requirement) pair** — query-time; count ≈ traffic, usually a
|
|
52
|
+
few per request (see Phase D).
|
|
53
|
+
|
|
54
|
+
Two things to separate immediately:
|
|
55
|
+
|
|
56
|
+
1. **Structure / connection extraction** (entities, parties, graph edges — "who/what is in this document and how is
|
|
57
|
+
it connected") is a **once-per-document** pass. It is NOT the per-unit cost; leave the LLM there. Do not mistake
|
|
58
|
+
it for the thing to decompose. (In this engine that is the docling-graph pass; it runs once per contract and once
|
|
59
|
+
per regulation.)
|
|
60
|
+
2. **Per-unit semantic tagging** (what IS this unit, what are its typed properties) is the multiplying cost. This is
|
|
61
|
+
the target.
|
|
62
|
+
|
|
63
|
+
## Phase B — Classify each decision by its shape, then pick the mechanism
|
|
64
|
+
|
|
65
|
+
For every per-unit decision, name its shape. The shape dictates the mechanism, in this order of preference (cheapest
|
|
66
|
+
and most robust first):
|
|
67
|
+
|
|
68
|
+
1. **Boundary / segmentation** — "does a new unit start here?", "is this span operative / extractable?" →
|
|
69
|
+
**deterministic structure first** (numbering, enumeration `(a)(b)`, headings, list markers). Add a small
|
|
70
|
+
**binary classifier** only for the residue where structure is ambiguous (the analog of a `is_extractable_span`
|
|
71
|
+
model). Rarely needs an LLM.
|
|
72
|
+
2. **Closed-vocab with an enumerable cue list** — a value the ontology can map from a fixed set of trigger phrases
|
|
73
|
+
(e.g. a deontic type from "must / shall / may not") → a **deterministic cue-rule, NO ML**. Do this before
|
|
74
|
+
training anything; it is free and exact.
|
|
75
|
+
3. **Single-label routing** — one of a closed set, no clean cue list → a **classifier**.
|
|
76
|
+
4. **Multi-label soft-tagging** — several of a closed set, used as *guidance* not a gate → a **soft-tag classifier**
|
|
77
|
+
(top-k). The most forgiving shape; a modest-accuracy model is still useful because wrong extra tags are cheap.
|
|
78
|
+
5. **Verbatim vs generated text** — if the record just needs the unit's text, extract the **verbatim span**
|
|
79
|
+
(deterministic) rather than a generated paraphrase. Drops a generative LLM step and is more faithful for
|
|
80
|
+
citation. Keep a paraphrase only if a human-readable restatement is a real requirement.
|
|
81
|
+
6. **Open / numeric / free-text / synthesis** — no closed set → keep the LLM, but reduce it to **one residual call
|
|
82
|
+
per unit** carrying only the fields that genuinely need it.
|
|
83
|
+
7. **Pair / entailment** — rerank a candidate against a query, or a judge verdict over a (subject, rule) pair → a
|
|
84
|
+
**cross-encoder / NLI classifier** is a candidate (a verdict over a closed label set IS a classification). Often
|
|
85
|
+
query-time; see Phase D.
|
|
86
|
+
|
|
87
|
+
## Phase C — What to look for in the documents themselves
|
|
88
|
+
|
|
89
|
+
Signals that make decomposition **feasible** (push toward rules + classifiers):
|
|
90
|
+
|
|
91
|
+
- explicit **numbering / enumeration / heading** structure → deterministic boundaries;
|
|
92
|
+
- a **closed, ontology-authored vocabulary** for the typed fields → classifiers + cue-rules;
|
|
93
|
+
- **repeated template structure** across documents → stable features;
|
|
94
|
+
- **enumerable linguistic cues** (deontic verbs, defined terms, standard phrasings) → cue-rules.
|
|
95
|
+
|
|
96
|
+
Signals that **resist** it (keep the LLM): genuinely open / unbounded values, cross-document or multi-hop
|
|
97
|
+
reasoning, long-prose synthesis, values that depend on interpretation rather than surface form.
|
|
98
|
+
|
|
99
|
+
## Phase D — Ingestion vs retrieval: where the win actually is
|
|
100
|
+
|
|
101
|
+
- **Ingestion** per-unit work on long documents multiplies into thousands of calls. This is the **biggest, do-first**
|
|
102
|
+
opportunity. The whole Phase B decomposition applies.
|
|
103
|
+
- **Retrieval / query time** is typically a **few calls per request** (classify the query, extract its constraints,
|
|
104
|
+
rerank, judge). These ARE classifier/routing shapes (closed-set query routing, a pair-classifier reranker or
|
|
105
|
+
judge), but the volume is low and you often want the LLM's rationale or synthesis. It is legitimate to **live with
|
|
106
|
+
the LLM at query time** (as this engine does for the contract and compliance judges) and still decompose
|
|
107
|
+
ingestion fully. State this as a deliberate choice per decision; do not reflexively de-LLM query time.
|
|
108
|
+
|
|
109
|
+
## Phase E — Soft-tag vs hard-gate: decide the role before committing
|
|
110
|
+
|
|
111
|
+
The same classifier is safe or dangerous depending on how its output is used:
|
|
112
|
+
|
|
113
|
+
- a **soft tag** that only augments / hints → safe even at modest accuracy; wrong extra tags are cheap.
|
|
114
|
+
- a **hard gate** that drops, blocks, or routes irreversibly → needs high accuracy AND a safe fallback.
|
|
115
|
+
|
|
116
|
+
Decide the role first. Keep a **graceful-degrade path**: a classifier abstention or a persistent rule-miss should
|
|
117
|
+
fall back to the residual LLM (or to an explicit "ambiguous"), never to a silent wrong answer.
|
|
118
|
+
|
|
119
|
+
## Phase F — Make it ttl-driven, and know what transfers across domains
|
|
120
|
+
|
|
121
|
+
Author the closed vocabulary and the cue lists in the **ontology (`.ttl`), never in Python** (the engine's
|
|
122
|
+
knowledge-in-the-ontology rule). Then:
|
|
123
|
+
|
|
124
|
+
- the **deterministic mechanism** (structure split + cue-rule + verbatim extraction) is domain-generic and
|
|
125
|
+
**transfers to any pack for free** — it reads whatever vocab/cues the pack authors;
|
|
126
|
+
- a **trained classifier is vocabulary-specific** — a new-vocabulary domain pack trains **its own**. "Reuse across
|
|
127
|
+
products" therefore means the *same mechanism + per-pack models*, not one model everywhere.
|
|
128
|
+
- A pack that only re-routes an existing vocabulary at query time (an override overlay) is NOT a new-vocabulary pack
|
|
129
|
+
and needs no new ingestion classifier.
|
|
130
|
+
|
|
131
|
+
## The output: the opportunity list
|
|
132
|
+
|
|
133
|
+
Produce a decomposition plan, not prose:
|
|
134
|
+
|
|
135
|
+
1. **Headline the single biggest per-unit LLM cost** (the multiplying ingestion pass).
|
|
136
|
+
2. For **each decision** give: its unit, its shape (Phase B), the chosen mechanism (deterministic / cue-rule /
|
|
137
|
+
classifier / residual-LLM / verbatim), and its role (soft-tag vs gate).
|
|
138
|
+
3. Separate an **ingestion bucket** (do first) from a **query-time bucket** (decide case by case; often live with
|
|
139
|
+
the LLM).
|
|
140
|
+
4. Note the **ttl + per-pack** generality (what transfers, what each pack re-trains).
|
|
141
|
+
5. Exclude anything that is not actually a per-unit cost (once-per-document structure extraction) and anything
|
|
142
|
+
already settled (a decision an existing cue-rule covers).
|
|
143
|
+
|
|
144
|
+
## Hand-off (what to do with the opportunities)
|
|
145
|
+
|
|
146
|
+
- **Deterministic rule / cue-rule / span-split** → plain code inside the ingestion subgraph. NOT a capability; it is
|
|
147
|
+
mechanism, and the knowledge it reads lives in the `.ttl`.
|
|
148
|
+
- **System-1 decision model (NO training)** → for a closed-set decision (yes/no, choice, score), A/B a decision
|
|
149
|
+
model — **Jev** (managed, OpenRouter Decisions API, zero/few-shot, calibrated) or **Laya** (open, fine-tuned) —
|
|
150
|
+
BEFORE committing to a trained classifier. It often wins when data is scarce or label-ambiguous, or when you need
|
|
151
|
+
calibrated uncertainty to route/gate (measured: RAG_Wright CIC-1c — Jev zero-shot 0.92 vs a trained SetFit 0.82;
|
|
152
|
+
ADR-0119). See `setfit` Phase 0.5 (the decision-vs-train A/B) and the `laya` skill; wire it as a `jev_decision`-style
|
|
153
|
+
model capability with a `DecisionModelProfile`. **Evaluate this first; it may remove the need to train at all.**
|
|
154
|
+
- **Trained classifier** → build it with the **`setfit`** skill (framing, symmetric leakage-safe eval, per-class
|
|
155
|
+
floor, soft-tag/top-k, rare-class curation, checkpointing). Serve the teacher / bulk-labeler with **`qwen-vllm-modal`**.
|
|
156
|
+
Then register it as a capability with **`authoring-a-capability`**: a `kind="model"` capability with an
|
|
157
|
+
`impl_ref` factory `def <slug>(resources, inputs)` over a cached checkpoint, invoked by name through the engine
|
|
158
|
+
API (`invoke_model` / `ainvoke_model`) and the pipeline (`dispatch_model` / `adispatch_model`) — **routed THROUGH
|
|
159
|
+
the capability layer, never hand-constructed around it.** A model impl may be sync (a classifier / XGBoost — run
|
|
160
|
+
off-loop by `ainvoke_model`) or async (an LLM-backed cap — awaited by `ainvoke_model`).
|
|
161
|
+
|
|
162
|
+
## Anti-patterns (from the real sessions — do not repeat)
|
|
163
|
+
|
|
164
|
+
- **Letting classifier training cost/time decide whether an opportunity exists.** Identify opportunities by the
|
|
165
|
+
*pattern* (per-unit closed-set decision on a long document); training ROI is a separate, later question.
|
|
166
|
+
- **Mistaking once-per-document structure extraction for the per-unit cost.** It is cheap; leave the LLM.
|
|
167
|
+
- **Claiming a judge/verdict step "can't classify."** A verdict over a closed label set (compliant / violation /
|
|
168
|
+
needs-review, relevant / not) IS a classification; it is a legitimate pair-classifier candidate (usually
|
|
169
|
+
query-time).
|
|
170
|
+
- **Hardcoding the vocabulary or cues in Python.** They belong in the `.ttl`; the mechanism reads them.
|
|
171
|
+
- **Training a classifier where an enumerable cue-rule already settles the decision** (the deontic-type case).
|
|
172
|
+
- **Forcing a hard gate where a soft tag suffices** — it imposes an accuracy bar you did not need.
|
|
173
|
+
- **Reflexively de-LLM'ing query time** — low volume + wanted rationale often make the LLM the right call there.
|
|
174
|
+
- **A router/cascade of specialists** — it multiplies errors (router acc × specialist acc). A flat classifier +
|
|
175
|
+
multi-tag usually beats it (see `setfit`).
|
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: creating-evals
|
|
3
|
+
description: >-
|
|
4
|
+
Domain-agnostic recipe for creating EVALS for engine/product capabilities, eval-first (TDD): write the eval as
|
|
5
|
+
soon as a capability is DEFINED (its contract + acceptance criterion), before it is implemented. Use it when
|
|
6
|
+
starting a new domain (right after capabilities are defined), when adding or changing a capability, or when A/B-ing
|
|
7
|
+
alternatives (e.g. a trained classifier vs a System-1 decision model). Covers gold-set design + reliability,
|
|
8
|
+
per-capability-KIND metrics (extraction / classification / retrieval / graph / generation / judgment), the
|
|
9
|
+
gate-vs-diagnostic split, building the gold cheaply, an isolated executable harness, and optional Langfuse
|
|
10
|
+
Datasets/Experiments/Scores automation. An eval is the executable acceptance criterion; training data IS an eval.
|
|
11
|
+
---
|
|
12
|
+
|
|
13
|
+
# Creating evals (eval-first / TDD)
|
|
14
|
+
|
|
15
|
+
An eval is the **executable acceptance criterion** for a capability. Write it **as soon as the capability is
|
|
16
|
+
DEFINED** — its contract (typed in/out) and its acceptance criterion exist — **before it is implemented**. This is
|
|
17
|
+
TDD at the capability level: the eval fails (red) on the unbuilt/weak capability, you implement to green, then you
|
|
18
|
+
can A/B alternatives and catch regressions forever. Corollary observed repeatedly in this engine: **the training
|
|
19
|
+
data you build for a classifier/decision model IS an eval** (a labeled gold set + a metric) — so building the eval
|
|
20
|
+
first also gives you the data design for free.
|
|
21
|
+
|
|
22
|
+
## When to use
|
|
23
|
+
- **Starting a new domain**: the FIRST build step after capabilities are defined (build-sequence step 4.5) — write
|
|
24
|
+
each capability's eval before/while you implement it.
|
|
25
|
+
- Adding or changing a capability, or tuning a threshold/prompt/model.
|
|
26
|
+
- **A/B-ing alternatives** on one capability (a deterministic rule vs a trained classifier vs a System-1 decision
|
|
27
|
+
model vs an LLM) — the eval is the neutral judge; select on the metric.
|
|
28
|
+
|
|
29
|
+
## 1. Design the gold set (the foundation — get this right FIRST)
|
|
30
|
+
- **Real, in-domain items + expected outputs/labels.** Small but REPRESENTATIVE; never the easy cases only.
|
|
31
|
+
- **Reproducible + pinned + gitignored.** The gold is a generated artifact: pin its source snapshot + the selected
|
|
32
|
+
ids so it rebuilds identically; gitignore the data, commit the BUILDER. (Pattern: `eval/build_golden.py`.)
|
|
33
|
+
- **Leakage-safe + balanced.** Split by the natural grouping unit (document / source / record), not by row.
|
|
34
|
+
For classification, a SYMMETRIC per-class test and a per-class FLOOR (not overall accuracy — it hides dead
|
|
35
|
+
classes). "k-shot" = k per class.
|
|
36
|
+
- **The gold is the CEILING — measure its RELIABILITY.** For subjective/ambiguous labels, get a second
|
|
37
|
+
independent labeling and report inter-annotator (or inter-pass) agreement; adjudicate the disagreements and
|
|
38
|
+
document the calls. A model cannot beat the gold's own consistency, and label ambiguity in the gold shows up as a
|
|
39
|
+
classifier ceiling you cannot train past (CIC-1c: a two-pass gold agreed at 0.905 and a trained SetFit capped
|
|
40
|
+
~0.82 — diagnose the gold before blaming the model). If you have no human expert yet, a documented rubric + a
|
|
41
|
+
two-pass consensus is
|
|
42
|
+
the honest proxy — say so.
|
|
43
|
+
|
|
44
|
+
## 2. Pick the metric by capability KIND, and split GATE vs DIAGNOSTIC
|
|
45
|
+
Always set ONE pass/fail **gate** (from the acceptance criterion) and report **diagnostics** alongside (never gate
|
|
46
|
+
on a diagnostic).
|
|
47
|
+
- **Extraction** (section → records, clauses, requirements): recall / precision / F1 of extracted items vs gold,
|
|
48
|
+
reported SEPARATELY (under- vs over-extraction are different failures). Watch over-extraction on non-operative
|
|
49
|
+
input (definitions) and under-extraction on long input.
|
|
50
|
+
- **Classification / typed decision** (incl. SetFit, Laya, **Jev**): per-class recall + the per-class **floor**;
|
|
51
|
+
symmetric eval; top-k recall for multi-label (reported against tags-per-item). Prefer a soft-tag/calibrated
|
|
52
|
+
metric when the output routes rather than hard-gates.
|
|
53
|
+
- **Retrieval**: recall@k (binary, relevant = grade ≥ a floor) as the GATE; nDCG@k (graded, exp gain) as a
|
|
54
|
+
DIAGNOSTIC (do NOT threshold nDCG). Isolate the retrieval legs (score corpus-ids before rehydration). Pattern:
|
|
55
|
+
`eval/acord_retrieval.py`.
|
|
56
|
+
- **Graph / relational**: recall over the answer SET (reachability/traversal), k large enough to cover the set;
|
|
57
|
+
node ids match gold by construction. Pattern: `eval/relational_eval.py`.
|
|
58
|
+
- **Generation / QA**: citation-recall / groundedness / correct-abstention; LLM-as-judge for free-text, but anchor
|
|
59
|
+
with deterministic checks (a cited span must exist). Generation is non-deterministic near the abstain boundary —
|
|
60
|
+
measure over several runs.
|
|
61
|
+
- **Judgment / compliance verdicts**: report PRECISION and RECALL of the actionable class (e.g. violation)
|
|
62
|
+
SEPARATELY (alert-fatigue vs missed), and break out by provenance (real vs constructed). Pattern:
|
|
63
|
+
`scripts/eval_compliance_gold.py`.
|
|
64
|
+
|
|
65
|
+
## 3. Build the gold cheaply (without faking it)
|
|
66
|
+
- **Silver bootstrapping**: a higher-capability teacher (LLM, or a decision model) labels candidate items; CURATE
|
|
67
|
+
with reject-rules; mark silver, never conflate with gold. (Teacher labeling runs on the flat-GPU substrate per
|
|
68
|
+
`qwen-vllm-modal`, not ad-hoc paid calls.)
|
|
69
|
+
- **Real public datasets** where they exist (e.g. labeled corpora for the task) — add to TRAIN; keep the TEST
|
|
70
|
+
in-corpus + human-adjudicated so it stays an honest transfer test.
|
|
71
|
+
- **Generate hard cases** (near-boundary positives/negatives) to stress the exact confusions — TRAIN-ONLY; the
|
|
72
|
+
gold test stays real. (Full data playbook: the `setfit` skill Phase 1/4 + `classifier-opportunity-analysis`.)
|
|
73
|
+
|
|
74
|
+
## 4. Make it an executable, isolated harness
|
|
75
|
+
- **Isolate the capability under test**: inject it (a seam / `extract_override` / an injected `retrieve`) so the
|
|
76
|
+
eval measures ONE capability, not the whole pipeline. Invoke production code THROUGH the capability layer
|
|
77
|
+
(`ainvoke_subgraph`/`ainvoke_model`), never a hand-built copy.
|
|
78
|
+
- **Score from a RESULT ARTIFACT (JSON), not stdout scraping** (scraping truncates and silently drops rows).
|
|
79
|
+
- **Parallelize** model/LLM calls (async + semaphore) — same cost, far less wall-clock; order-preserving so it
|
|
80
|
+
stays deterministic.
|
|
81
|
+
- **Env-selected** so the SAME harness runs local (dev) or on Modal (full corpus + GPU).
|
|
82
|
+
- **Pre-flight paid bulk** (>~50 paid calls): print the exact count + cost and wait (`warn-before-bulk` rule).
|
|
83
|
+
- Stream `X/N` progress + actively monitor any run over ~30s (never launch-and-forget).
|
|
84
|
+
|
|
85
|
+
## 5. Langfuse — optional eval automation (we already use it for tracing)
|
|
86
|
+
Langfuse has a first-class eval stack we are NOT yet using (we use it only for spans/usage today): a **Dataset**
|
|
87
|
+
(items = input + optional expected output) → a **Task** (your capability) run over the dataset as an **Experiment
|
|
88
|
+
Run** → **Evaluators** (deterministic checks or LLM-as-judge) producing **Scores** (numeric/categorical/boolean),
|
|
89
|
+
all linked to traces. Reach for it when you want **tracked, re-runnable, dashboarded** evals + regression tracking
|
|
90
|
+
across capability versions; a local JSON harness is enough for a quick one-off gate. Keep the gold BUILDER + a
|
|
91
|
+
committed snapshot in the repo (reproducibility); push the items to the Langfuse dataset so runs and scores land
|
|
92
|
+
next to the traces you already collect. Ground the exact Dataset/Experiment API before wiring
|
|
93
|
+
(https://langfuse.com/docs/evaluation/concepts, https://langfuse.com/docs/datasets/overview); confirm the installed
|
|
94
|
+
SDK version's surface, not prose.
|
|
95
|
+
|
|
96
|
+
## 6. The TDD loop
|
|
97
|
+
1. Write the eval from the contract + a handful of real gold items → it fails (red) on the unbuilt/weak capability.
|
|
98
|
+
2. Implement the minimum to pass the gate (green).
|
|
99
|
+
3. A/B alternatives (rule / trained classifier / decision model / LLM) on the SAME gold; select on gate + diagnostics + cost + calibration.
|
|
100
|
+
4. Keep the eval; it is the regression guard (the Beyonce rule — if you shipped it, it has an eval).
|
|
101
|
+
|
|
102
|
+
## Anti-patterns (seen, do not repeat)
|
|
103
|
+
- **Overall accuracy** instead of a per-class floor — hides dead classes.
|
|
104
|
+
- **Testing on teacher/generated data as if it were gold** — measures mimicry, not accuracy; keep the test real + adjudicated.
|
|
105
|
+
- **Gating on a diagnostic** (e.g. nDCG) — report it, don't threshold it.
|
|
106
|
+
- **No reliability check on a subjective gold** — you can't read a number off a gold whose own labels disagree.
|
|
107
|
+
- **Eval not isolated** to the capability (measures the whole pipeline) — you can't attribute a regression.
|
|
108
|
+
- **Scraping stdout** for metrics — write + read a JSON artifact.
|
|
109
|
+
- **Faking a win by shrinking the sample / picking easy items** — real, representative, honest.
|
|
110
|
+
|
|
111
|
+
## Hand-off / where this sits
|
|
112
|
+
Chain: `classifier-opportunity-analysis` (identify the decision) → **`creating-evals` (THIS — write the eval FIRST)**
|
|
113
|
+
→ build it (a deterministic rule, or `setfit`/`laya`/Jev for a decision, or an LLM) → `authoring-a-capability`
|
|
114
|
+
(register). In the new-domain build sequence, this is the step right after capabilities are defined.
|
|
@@ -0,0 +1,119 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: laya
|
|
3
|
+
description: Recipe for fine-tuning and serving a Laya (ModernBERT-large, RL-trained typed-decision) classifier -- the OPEN-weight System-1 decision model -- to replace an LLM decision that a SetFit/encoder classifier PLATEAUS on. Also explains Jev, the MANAGED zero-shot sibling (OpenRouter Decisions API), and when to A/B Jev first (usually) vs fine-tune Laya (on-prem / no managed API). Use when confusable/relational/subjective closed-vocab values won't clear the bar with any encoder backbone. Covers uv install + env-isolation gotchas, the JSONL schema, per-value criteria, single-T4 fine-tune (RLCD), few-shot-in-context, balance, OOM knobs, Modal harness, CPU/MPS/GPU serving, and the jev_decision/DecisionModelProfile wiring (ADR-0119).
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Laya typed-decision classifier recipe
|
|
7
|
+
|
|
8
|
+
Laya (https://github.com/NandhaKishorM/laya) is a **non-autoregressive "System-1 decision engine"**: a
|
|
9
|
+
ModernBERT-large encoder + an **RL-trained decision head** that makes a typed `choice` / `score` / `noul` in one
|
|
10
|
+
forward pass, with per-option **criteria** (written descriptions) and calibrated confidence. It is the **escalation
|
|
11
|
+
past SetFit**: use it when the decision — not the representation — is the bottleneck.
|
|
12
|
+
|
|
13
|
+
**Every number below is a DIRECTION from one project, never a target — re-measure on your own data.**
|
|
14
|
+
|
|
15
|
+
## Laya vs Jev (the managed sibling) — pick the decision model first
|
|
16
|
+
Laya is the **open-weight** System-1 decision model; **Jev** (TypeSafe, https://openrouter.ai, `typesafe/jev-1.13`
|
|
17
|
+
via the OpenRouter **Decisions API**) is the **managed** one — same idea (typed `noul`/`choice`/`score` + calibrated
|
|
18
|
+
probabilities), but **strong ZERO/few-shot with no fine-tuning**, whereas Laya scores near-random zero-shot and MUST
|
|
19
|
+
be fine-tuned. Measured (RAG_Wright CIC-1c, ADR-0119): on the operative-rule gate **Jev zero-shot hit 0.92** where a
|
|
20
|
+
trained SetFit capped ~0.82 and fine-tuned Laya reached ~0.70–0.74. So: **for a closed-set decision, A/B Jev
|
|
21
|
+
(zero-shot, instant) first**; reach for Laya when a **managed API is unacceptable** (on-prem / data-residency) and
|
|
22
|
+
you can fine-tune. In RAG_Wright a decision model is wired as a `kind="model"` capability (`jev_decision`) behind a
|
|
23
|
+
`DecisionModelProfile` (model id / endpoint / thresholds in config, swappable Jev ↔ Laya) — see `setfit` Phase 0.5 +
|
|
24
|
+
`authoring-a-capability`. The decision questions/criteria live in the `.ttl` (ADR-0066), not the capability.
|
|
25
|
+
|
|
26
|
+
## When to use Laya (vs SetFit)
|
|
27
|
+
- A closed-vocab value **plateaus below the bar with EVERY encoder backbone you try** (we ruled it out with
|
|
28
|
+
LegalBERT, bge-large, all-mpnet, AND ModernBERT-large as SetFit bodies — all stuck). That means the bottleneck is
|
|
29
|
+
the *decision* ("who bears the obligation", "which side is favored", "may either party terminate?"), not the
|
|
30
|
+
embedding. Laya's RL-against-proper-scoring-rules training targets exactly that.
|
|
31
|
+
- Measured proof it's a different mechanism: `termination_right` cleared 0.68 with Laya where four encoders maxed at
|
|
32
|
+
0.60–0.64. It is NOT magic — it did NOT rescue minority-data-starved or purely-numeric dims (see Limits).
|
|
33
|
+
- Default to SetFit first (cheaper, simpler, CPU-native). Reach for Laya only for the residual confusable/relational
|
|
34
|
+
dims SetFit can't crack.
|
|
35
|
+
|
|
36
|
+
## Install — uv only, and beware env bleed
|
|
37
|
+
- **uv, never pip.** Locally `uv add laya`; on Modal `.uv_pip_install("laya", extra_index_url="https://download.pytorch.org/whl/cu124")` for CUDA torch wheels.
|
|
38
|
+
- **The env-bleed gotcha (cost real time):** a bare `uv run --with laya` on a machine with a system Anaconda picked
|
|
39
|
+
up conda's numpy/scipy/sklearn and crashed (`numpy.core.multiarray failed to import`). FIX: force uv's own managed
|
|
40
|
+
Python with a cleared path — `PYTHONPATH= PYTHONNOUSERSITE=1 uv run --python 3.11 --with laya python …` — or run
|
|
41
|
+
inside a proper isolated uv project. Never let a base conda leak in.
|
|
42
|
+
- Deps are compatible with our stack: `transformers>=4.48`, `torch>=2.0`, `huggingface_hub`, Python ≥3.10.
|
|
43
|
+
|
|
44
|
+
## Checkpoints — pick the ENGLISH base, not multilingual
|
|
45
|
+
- English base = `convaiinnovations/laya` **root** (`laya.load("convaiinnovations/laya")`, subfolder=None) —
|
|
46
|
+
ModernBERT-large. `subfolder="multilingual"` = mmBERT; `subfolder="typed-decisions"` = a fine-tuned variant.
|
|
47
|
+
- **The fine-tune script defaults `--model-dir` to the MULTILINGUAL subfolder** — you must pass the English root
|
|
48
|
+
explicitly, e.g. `snapshot_download("convaiinnovations/laya", allow_patterns=["*.json","*.safetensors","encoder/*","tokenizer/*"])` and point `--model-dir` there. A checkpoint dir = `{encoder/, tokenizer/, model.safetensors, rl_agent_config.json}`.
|
|
49
|
+
- Base checkpoints score ~chance zero-shot (~0.36); **all the value is in fine-tuning.** (Zero-shot on our clauses
|
|
50
|
+
was directionally right but ~0.5 confidence — expected.)
|
|
51
|
+
|
|
52
|
+
## Data — the JSONL schema + criteria (the differentiator)
|
|
53
|
+
One JSONL line per case (`research/scripts/finetune_single_device.py` reads this exact shape):
|
|
54
|
+
```
|
|
55
|
+
{"state": "<clause text>",
|
|
56
|
+
"questions": {"<dim>": {"type":"choice","instructions":"<the question>","criteria":{"<value>":"<description>", ...}}},
|
|
57
|
+
"gold": {"<dim>": {"probabilities": {"<value>": <p>, ...}, "label": "<value>"}}}
|
|
58
|
+
```
|
|
59
|
+
- **`criteria` = per-value written descriptions = the semantic-guidance lever SetFit never had.** Author them from
|
|
60
|
+
the ontology (`.ttl`); if the ontology is thin (ours had vocab but no per-value defs), author from the value
|
|
61
|
+
semantics and CONFIRM with a human before training. Wording matters — it directly shapes what the model learns.
|
|
62
|
+
- **`gold` is a DISTRIBUTION (soft targets), not a hard label** — RLCD trains on the teacher's probability per option.
|
|
63
|
+
One-hot (`{gold:1.0, other:0.0}`) works and is the pragmatic start; true soft targets (teacher per-option probs)
|
|
64
|
+
are the design-intended enhancement.
|
|
65
|
+
- Reuse your existing **contract-disjoint, symmetric splits** so floors are directly comparable to SetFit. Silver
|
|
66
|
+
goes in TRAIN only; val/test stay gold.
|
|
67
|
+
|
|
68
|
+
## Fine-tune — single T4, RLCD, calibration built in
|
|
69
|
+
- `python research/scripts/finetune_single_device.py --data <train.jsonl> --model-dir <english-base> --output-dir <out> --epochs 4` (reproduces the 2×T4 notebook without DDP; CPU/one-GPU).
|
|
70
|
+
- Runs on **one 16 GB GPU (T4)** — ~1–2 h for large data, minutes for our small dims. RLCD = soft-CE + GRPO-style
|
|
71
|
+
policy gradient on proper scoring rules. Calibration (one temperature per type) is fitted inside the run on a
|
|
72
|
+
held-out slice; **argmax/accuracy unchanged, only confidence moves** — always fit before gating on confidence.
|
|
73
|
+
- **OOM knob (hit this):** ModernBERT-large (~400M) OOMs a T4/L4 at the default `batch_size=16, max_seq=256`. Clauses
|
|
74
|
+
are short → set `--batch-size 8 --max-seq 128` (the script/notebook also enables gradient checkpointing). That fit
|
|
75
|
+
comfortably.
|
|
76
|
+
- **Run preflight FIRST — see the [[setfit]] "Run preflight & monitoring" section.** It is framework-agnostic and
|
|
77
|
+
applies to Laya exactly as to SetFit, including when you GENERATE the Laya JSONL labels with a teacher: resolve
|
|
78
|
+
the model from the engine (never a hardcoded/stale id; bulk teacher labeling on Modal Qwen, not OpenRouter), load
|
|
79
|
+
`.env` by EXPLICIT path from an out-of-repo script, SMOKE one item before the fan-out, and stream X/N to a log you
|
|
80
|
+
actively monitor. Those exact mistakes cost runs on 2026-10-04.
|
|
81
|
+
- **Modal harness = reuse the [[setfit]] hardened pattern:** `.uv_pip_install("laya")` + `.add_local_file` the
|
|
82
|
+
finetune script; launcher BLOCKS on `.get()` per spawn (no spawn-and-return), stamps + verifies a `data_sha`,
|
|
83
|
+
writes a manifest, streams X/N; snapshot keepers server-side to a `/checkpoints/<name>` path (a small copy fn) —
|
|
84
|
+
the laya volume has no built-in snapshot, add one. Never lose a fine-tune.
|
|
85
|
+
- **ACCOUNT CONTAINER CAP = 10 (fzaidi2014).** Spawning more than 10 fine-tunes at once does NOT run them all —
|
|
86
|
+
Modal queues the rest and runs ~10 at a time (correct, but ~N/10 waves of wall-clock, and a "why only 10 running?"
|
|
87
|
+
surprise). Size a `groupbake`/dimbatch fan-out to ≤10 in flight (chunk into waves + gather between), or state the
|
|
88
|
+
wave count honestly. This is a DIFFERENT knob from the vLLM `max_containers=1` + `@modal.concurrent` batching in the
|
|
89
|
+
qwen-vllm-modal skill — do not conflate. (Ignored the stated cap once, spawned 29 → 3 waves.)
|
|
90
|
+
|
|
91
|
+
## Levers that matter (measured, dim-dependent)
|
|
92
|
+
- **Few-shot-in-`state` is a big BUT dim-dependent lever.** Prepend a few labeled exemplars (from TRAIN, per class)
|
|
93
|
+
to the state. It lifted subjective/relational dims a lot (favorability 0.17→0.50, party_asymmetry 0.18→0.55) and
|
|
94
|
+
**HURT a numeric dim** (cap_basis 0.40→0.00, collapsed). **A/B few-shot per dim; never assume it helps.**
|
|
95
|
+
- **Balance: don't down-sample to tiny data.** Balancing favorability DOWN to 18/18 (36 rows) + soft targets
|
|
96
|
+
COLLAPSED it (0.00) — the larger imbalanced set did better (0.50). A starved minority value needs balance-**UP**
|
|
97
|
+
(mine more minority data), not down-sampling the majority.
|
|
98
|
+
- **The residual failures are minority-data-starvation, not a Laya ceiling** — the closest miss (party_asymmetry
|
|
99
|
+
0.55) is one minority-silver top-up from the bar. Diagnose starvation before concluding "Laya can't."
|
|
100
|
+
|
|
101
|
+
## Limits (honest)
|
|
102
|
+
- Laya rescues *decision-limited* confusable dims; it does **not** fix (a) **minority-data-starved** values (needs
|
|
103
|
+
more data), or (b) **numeric/structural** distinctions (cap_basis "fixed sum vs multiple-of-fees" — encoders and
|
|
104
|
+
Laya both failed; that belongs on the LLM).
|
|
105
|
+
|
|
106
|
+
## Serving — device-agnostic, load once
|
|
107
|
+
- `agent = laya.load(<checkpoint>)` — device auto: CUDA → MPS → CPU. **Verified on a Mac (MPS): ~25 s load, ~1.8 s
|
|
108
|
+
first-call warmup, then ~80–140 ms/call.** Runs on CPU too. So it honors the engine's "use a GPU if present, else
|
|
109
|
+
CPU" philosophy — same as SetFit and the LLM profiles.
|
|
110
|
+
- **Load once / preload** (`Router(preload=True)`); never load per call. Serve behind the existing extractor seam so
|
|
111
|
+
nothing upstream changes. The checkpoint is ~820 MB (ModernBERT-large) — heavier than a SetFit body+joblib head;
|
|
112
|
+
budget the memory when co-loading with SetFit models.
|
|
113
|
+
- Serving location is NOT pinned: single T4, share the LLM's A100, or CPU/Mac — the seam + a device arg decide at
|
|
114
|
+
runtime, exactly like the model-profile seam for the LLM.
|
|
115
|
+
|
|
116
|
+
## Reference implementation
|
|
117
|
+
`~/work/laya` (the cloned repo: `research/scripts/finetune_single_device.py`, `docs/finetune.md`, `examples/`) and
|
|
118
|
+
`~/work/clause-classifier-ab/{laya_modal.py, build_laya_data.py, build_laya_refine.py}` (our Modal fine-tune + eval
|
|
119
|
+
+ data builders). Re-use the PATTERNS; the criteria, thresholds, and floors are specific to that problem.
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: qwen-vllm-modal
|
|
3
|
+
description: Authoritative recipe for standing up the self-hosted Qwen3.8-27B vLLM server on Modal (the production LLM substrate). READ THIS + ADR-0110 before touching the deploy — it fixes the "we forgot the agreed config and rediscovered every vLLM/Modal gotcha the hard way" failure. Covers the locked config, the image-build fixes, the concurrency + timeout serving fixes, cold start, and cost control.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Qwen3.8-27B on Modal (vLLM) — the production LLM server
|
|
7
|
+
|
|
8
|
+
**Read the record, do NOT reconstruct from memory.** The config is decided and the gotchas are known. Rummaging
|
|
9
|
+
half-remembered fragments and iterating live cost a full painful session even though it had all been done, decided,
|
|
10
|
+
and tested before. The two sources of truth:
|
|
11
|
+
1. **ADR-0110** (`docs/adr/0110-fp8-kv-cache-16k-single-a100.md`) — the LOCKED config + why.
|
|
12
|
+
2. **`scripts/modal_qwen3_vllm_server.py`** — the deploy script.
|
|
13
|
+
This skill is the operational checklist that ties them together.
|
|
14
|
+
|
|
15
|
+
## The LOCKED production config (ADR-0110) — do not re-litigate
|
|
16
|
+
**Config B: `Qwen/Qwen3.8-27B-FP8` weights + `--kv-cache-dtype fp8`, on 1× A100-80GB, TP=1, `--max-model-len 16384`, `--gpu-memory-utilization 0.95`.**
|
|
17
|
+
- Config B (FP8 weights) is the **locked deploy** — 52× concurrency @16K, ~23% higher throughput, ~27 GB weights;
|
|
18
|
+
accuracy measured **identical to bf16** on the real compliance workload (no FP8 penalty).
|
|
19
|
+
- **Config A** (`Qwen/Qwen3.8-27B`, **bf16 weights** + FP8 KV) is the **documented FALLBACK only** — 25× concurrency,
|
|
20
|
+
"zero weight-loss" — use it *only* if a later clean check ever shows an FP8-weight regression. Don't default to A.
|
|
21
|
+
- If anyone (including a future you) says "config A is production" — **check ADR-0110 first**; the ADR locks **B**.
|
|
22
|
+
- Served-model-name is `Qwen/Qwen3.8-27B` (what the engine profile `qwen3.8-27b-modal` asks for), even though the
|
|
23
|
+
weights are the `-FP8` repo.
|
|
24
|
+
|
|
25
|
+
**The exact working deploy (this succeeded):**
|
|
26
|
+
```
|
|
27
|
+
MODAL_IMAGE_BUILDER_VERSION=2025.06 \
|
|
28
|
+
APP_NAME=rw-qwen3-modal MODEL=Qwen/Qwen3.8-27B-FP8 KV_CACHE_DTYPE=fp8 \
|
|
29
|
+
SERVED_NAME=Qwen/Qwen3.8-27B TOOL_PARSER=hermes REASONING_PARSER=qwen3 \
|
|
30
|
+
GPU=A100-80GB:1 TP=1 MAX_LEN=16384 GPU_UTIL=0.95 MAX_NUM_SEQS=<target> \
|
|
31
|
+
uv run --no-sync modal deploy scripts/modal_qwen3_vllm_server.py
|
|
32
|
+
```
|
|
33
|
+
Then wire the engine: `.env` `VLLM_BASE_URL=https://<workspace>--rw-qwen3-modal-serve.modal.run/v1`,
|
|
34
|
+
`VLLM_API_KEY=rw-vllm-dev-key`; the profile `qwen3.8-27b-modal` routes there. **Stop billing when done:**
|
|
35
|
+
`uv run --no-sync modal app stop rw-qwen3-modal --yes` (A100 is expensive).
|
|
36
|
+
|
|
37
|
+
## Image-build gotchas (the painful iterations — all fixed)
|
|
38
|
+
Building a fresh vLLM image on a clean Modal account exposed a chain of failures. **The reliable fix is to base off
|
|
39
|
+
the OFFICIAL prebuilt vLLM image** rather than pip-building vLLM:
|
|
40
|
+
1. `pip_install("vllm")` (unpinned) on `cuda:12.8.1` → **`xformers` source-builds** → `ModuleNotFound: torch`
|
|
41
|
+
(pip build-isolation). Don't pip-build vLLM.
|
|
42
|
+
2. `uv_pip_install("vllm", extra_index_url=pytorch-cu)` → resolver "no versions of vllm … unsatisfiable" — also fragile.
|
|
43
|
+
3. **WORKING:** `modal.Image.from_registry("vllm/vllm-openai:latest")` — vllm+torch+xformers+flashinfer prebuilt,
|
|
44
|
+
nothing compiles. It needs two adjustments:
|
|
45
|
+
- `setup_dockerfile_commands=["RUN ln -sf $(command -v python3) /usr/local/bin/python"]` — the image ships
|
|
46
|
+
`python3` but not `python`; Modal's **own pip bootstrap runs right after FROM** as `python -m pip` → **exit 127**
|
|
47
|
+
without the symlink. It must be in `setup_dockerfile_commands` (runs BEFORE the bootstrap), not a later layer.
|
|
48
|
+
- `.entrypoint([])` — clear the image's ENTRYPOINT (the api_server) so Modal's runtime isn't hijacked.
|
|
49
|
+
4. **Modal LEGACY image builder clobbers the stack**: it installs its old client deps OVER the vLLM image,
|
|
50
|
+
downgrading **pydantic→v1 and fastapi→old**, so vLLM crashes at startup (`cannot import name 'model_validator'`,
|
|
51
|
+
then `cannot import name 'Undefined' from pydantic.fields`). Patching pydantic just moves the break to fastapi
|
|
52
|
+
(whack-a-mole). **FIX = `MODAL_IMAGE_BUILDER_VERSION=2025.06`** (a modern builder whose client stack is pydantic-v2
|
|
53
|
+
compatible). Valid builder versions: `{2023.12, 2024.04, 2024.10, 2025.06}` — use the newest.
|
|
54
|
+
5. `HF_HUB_DISABLE_XET=1` in the image env — Xet holds an open log handle under the HF cache so `vol.commit()` fails
|
|
55
|
+
on a fresh weight download.
|
|
56
|
+
|
|
57
|
+
## Serving gotchas (equally painful — all fixed)
|
|
58
|
+
1. **`@modal.concurrent(max_inputs=N)` is REQUIRED on the web_server function.** Without it Modal feeds the single
|
|
59
|
+
container ONE request at a time (`Running: 1 reqs` + a flood of "Received a cancellation signal", `/health`
|
|
60
|
+
blocked) — vLLM's continuous batching is starved. Concurrency comes from **KV-cache batching inside ONE container**
|
|
61
|
+
(`max_containers=1`), not from more containers. Set `max_inputs` = your `--max-num-seqs`.
|
|
62
|
+
2. **The client structured-call timeout is too tight for reasoning-ON calls under batch.** At `--max-num-seqs` load,
|
|
63
|
+
reasoning calls take ~50–60 s; the seam's default 60 s timeout kills the ones just over the line → **retry pileup →
|
|
64
|
+
requests pile past `max_inputs` → instant `APIConnectionError` (0.0 s) cascade**. FIX: the seam reads
|
|
65
|
+
`RAG_STRUCTURED_TIMEOUT_S` (default 60) — raise it (e.g. `150`) for reasoning-ON bulk work. Server logs confirm the
|
|
66
|
+
calls actually complete 200 OK (~50 s); it was the client giving up.
|
|
67
|
+
3. **Pick client concurrency empirically, below the ceiling.** 16 was stable; 20 (== `max-num-seqs`) sat at the exact
|
|
68
|
+
ceiling and risked queueing. `Running: N, Waiting: 0` + zero cancellations = healthy.
|
|
69
|
+
4. **Reasoning ON vs OFF for bulk classification:** reasoning-OFF is ~30× faster but **changes the labels**
|
|
70
|
+
(over-keeps / different distribution) — do NOT swap it in for work that must match the reasoning-teacher's gold
|
|
71
|
+
(e.g. distillation silver). Keep reasoning ON when fidelity to the gold teacher matters; accept the cost.
|
|
72
|
+
|
|
73
|
+
## Cold start & cost
|
|
74
|
+
- Cold start ~440–480 s (CPU-bound engine init + a 51-graph CUDA-graph capture + first-time weight download), NOT
|
|
75
|
+
weight I/O — no load trick alone fixes it. GPU-snapshot / sleep-mode is a DEFERRED todo (the one attempt failed
|
|
76
|
+
because it was CPU-only; GPU retry not done). Run long deploys detached; poll `/health` for readiness.
|
|
77
|
+
- **A100 is billed while up — `modal app stop <app> --yes` the moment you're done.** Verify with `modal app list`
|
|
78
|
+
(state `stopped`, 0 tasks) + endpoint 404.
|
|
79
|
+
- **Do NOT warm the GPU before the consumer is ready, and do NOT trust auto-scaledown to protect billing** (burned
|
|
80
|
+
2026-10-04: a smoke-warmed A100 sat idle ~20 min while unrelated work ran, because "scaledown_window will handle
|
|
81
|
+
it" was assumed — it did not, fast enough). RULE: build the bulk/label/eval consumer FIRST (while the app is
|
|
82
|
+
stopped/cold), THEN warm → smoke → run → **`modal app stop` immediately after**, in one continuous go. If you warm
|
|
83
|
+
only to prove stand-up and the consumer is not next, STOP it right after the smoke. Treat the explicit stop as
|
|
84
|
+
mandatory, not the scaledown as sufficient; verify stopped with `modal app list`.
|
|
85
|
+
|
|
86
|
+
## The meta-lesson
|
|
87
|
+
Config + recipe live in **ADR-0110 + the script + this skill**. Before any Qwen/Modal work: read them, don't
|
|
88
|
+
reconstruct. When something "was working and we changed accounts/rebuilt," expect the image-builder + concurrency +
|
|
89
|
+
timeout trio above — they are the recurring three.
|
|
90
|
+
|
|
91
|
+
## Client side (when you use this server for bulk labeling/eval)
|
|
92
|
+
The SERVER config is here; the CLIENT preflight (resolve the model from the engine not a hardcoded id, load `.env`
|
|
93
|
+
by explicit path from an out-of-repo script, smoke ONE item before the bulk fan-out, mandatory X/N progress +
|
|
94
|
+
active monitoring) is in the **`setfit` skill's "Run preflight & monitoring"** section. Read it before driving a
|
|
95
|
+
bulk job against this server — those client gotchas (a stale model id, a silently-unloaded `.env`) cost a run each
|
|
96
|
+
on 2026-10-04.
|