jupytermind 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.github/skills/ai-chemistry-scientist/SKILL.md +97 -0
- package/.github/skills/ai-chemistry-scientist/manifest.json +156 -0
- package/.github/skills/ai-data-scientist/SKILL.md +330 -0
- package/.github/skills/ai-genomics-scientist/SKILL.md +98 -0
- package/.github/skills/ai-genomics-scientist/manifest.json +93 -0
- package/.github/skills/ai-materials-scientist/SKILL.md +51 -0
- package/.github/skills/ai-materials-scientist/manifest.json +58 -0
- package/.github/skills/ai-scientist/SKILL.md +69 -0
- package/.github/skills/ai-scientist/manifest.json +61 -0
- package/.github/skills/ai-structural-biology-scientist/SKILL.md +67 -0
- package/.github/skills/ai-structural-biology-scientist/manifest.json +72 -0
- package/.github/skills/japanese-prose/NOTICE.md +17 -0
- package/.github/skills/japanese-prose/SKILL.md +111 -0
- package/.github/skills/japanese-prose/references/review-workflow.md +50 -0
- package/.github/skills/japanese-prose/references/scoring.md +24 -0
- package/.github/skills/japanese-prose/references/writing-guidelines.md +60 -0
- package/.github/skills/japanese-prose/scripts/core.py +192 -0
- package/.github/skills/japanese-prose/scripts/fixtures/natural.md +5 -0
- package/.github/skills/japanese-prose/scripts/fixtures/unnatural.md +5 -0
- package/.github/skills/japanese-prose/scripts/lint.py +378 -0
- package/.github/skills/japanese-prose/scripts/outline.py +68 -0
- package/.github/skills/japanese-prose/scripts/terms.py +112 -0
- package/.github/skills/japanese-prose/scripts/test_engine.py +117 -0
- package/.github/skills/presentation-planner/SKILL.md +257 -0
- package/.github/skills/presentation-planner/assets/design-templates/data-report.yaml +97 -0
- package/.github/skills/presentation-planner/assets/design-templates/executive-proposal.yaml +92 -0
- package/.github/skills/presentation-planner/assets/design-templates/technical-briefing.yaml +96 -0
- package/.github/skills/presentation-planner/assets/scenario-templates/data-report.md +47 -0
- package/.github/skills/presentation-planner/assets/scenario-templates/executive-decision.md +43 -0
- package/.github/skills/presentation-planner/assets/scenario-templates/technical-briefing.md +45 -0
- package/.github/skills/presentation-planner/references/customizing-design-templates.md +160 -0
- package/.github/skills/presentation-planner/references/design-spec-schema.md +72 -0
- package/.github/skills/presentation-planner/references/handoff-contract.md +49 -0
- package/.github/skills/presentation-planner/references/responsibility-boundary.md +32 -0
- package/.github/skills/presentation-planner/references/scenario-templates.md +55 -0
- package/.github/skills/tech-writer/SKILL.md +434 -0
- package/.github/skills/tech-writer/assets/templates/blueprint.md +187 -0
- package/.github/skills/tech-writer/assets/templates/design-doc.md +29 -0
- package/.github/skills/tech-writer/assets/templates/migration-plan.md +173 -0
- package/.github/skills/tech-writer/assets/templates/operations-runbook.md +202 -0
- package/.github/skills/tech-writer/assets/templates/pr-description.md +23 -0
- package/.github/skills/tech-writer/assets/templates/qiita.md +44 -0
- package/.github/skills/tech-writer/assets/templates/readme.md +38 -0
- package/.github/skills/tech-writer/assets/templates/requirements-definition.md +170 -0
- package/.github/skills/tech-writer/assets/templates/rfi.md +113 -0
- package/.github/skills/tech-writer/assets/templates/rfp.md +180 -0
- package/.github/skills/tech-writer/assets/templates/security-design.md +167 -0
- package/.github/skills/tech-writer/assets/templates/system-design.md +220 -0
- package/.github/skills/tech-writer/assets/templates/technical-proposal.md +112 -0
- package/.github/skills/tech-writer/assets/templates/test-plan.md +153 -0
- package/.github/skills/tech-writer/assets/templates/user-manual.md +22 -0
- package/.github/skills/tech-writer/assets/templates/white-paper.md +192 -0
- package/.github/skills/tech-writer/references/doctypes/api-docs.md +33 -0
- package/.github/skills/tech-writer/references/doctypes/blueprint.md +81 -0
- package/.github/skills/tech-writer/references/doctypes/code-comments.md +39 -0
- package/.github/skills/tech-writer/references/doctypes/design-doc.md +42 -0
- package/.github/skills/tech-writer/references/doctypes/migration-plan.md +63 -0
- package/.github/skills/tech-writer/references/doctypes/operations-runbook.md +63 -0
- package/.github/skills/tech-writer/references/doctypes/pr-commit.md +82 -0
- package/.github/skills/tech-writer/references/doctypes/qiita.md +75 -0
- package/.github/skills/tech-writer/references/doctypes/readme.md +43 -0
- package/.github/skills/tech-writer/references/doctypes/release-notes.md +30 -0
- package/.github/skills/tech-writer/references/doctypes/requirements-definition.md +61 -0
- package/.github/skills/tech-writer/references/doctypes/rfi.md +43 -0
- package/.github/skills/tech-writer/references/doctypes/rfp.md +46 -0
- package/.github/skills/tech-writer/references/doctypes/security-design.md +71 -0
- package/.github/skills/tech-writer/references/doctypes/system-design.md +74 -0
- package/.github/skills/tech-writer/references/doctypes/technical-proposal.md +49 -0
- package/.github/skills/tech-writer/references/doctypes/test-plan.md +67 -0
- package/.github/skills/tech-writer/references/doctypes/user-manual.md +58 -0
- package/.github/skills/tech-writer/references/doctypes/white-paper.md +84 -0
- package/.github/skills/tech-writer/references/doctypes/zenn.md +66 -0
- package/.github/skills/tech-writer/references/japanese-prose-optimization.md +110 -0
- package/.github/skills/tech-writer/references/style-constitution.md +104 -0
- package/.github/skills/tech-writer/scripts/lint.py +412 -0
- package/LICENSE +21 -0
- package/README.md +92 -0
- package/bin/ai-data-scientist.js +123 -0
- package/package.json +41 -0
- package/pyproject.toml +45 -0
- package/src/ai_chemistry_scientist/__init__.py +0 -0
- package/src/ai_chemistry_scientist/admet_prediction.py +71 -0
- package/src/ai_chemistry_scientist/bioactivity_classification.py +73 -0
- package/src/ai_chemistry_scientist/data/sample_molecules.csv +21 -0
- package/src/ai_chemistry_scientist/dispatch.py +369 -0
- package/src/ai_chemistry_scientist/docking_score.py +97 -0
- package/src/ai_chemistry_scientist/drug_likeness_rules.py +84 -0
- package/src/ai_chemistry_scientist/evidence.py +41 -0
- package/src/ai_chemistry_scientist/molecular_descriptors.py +97 -0
- package/src/ai_chemistry_scientist/molecular_formula_mass.py +40 -0
- package/src/ai_chemistry_scientist/molecular_similarity.py +78 -0
- package/src/ai_chemistry_scientist/qsar_modeling.py +105 -0
- package/src/ai_chemistry_scientist/salt_standardization.py +81 -0
- package/src/ai_chemistry_scientist/structural_alerts.py +76 -0
- package/src/ai_chemistry_scientist/structure_format_conversion.py +84 -0
- package/src/ai_chemistry_scientist/validation.py +70 -0
- package/src/ai_data_scientist/__init__.py +0 -0
- package/src/ai_data_scientist/analysis_assumptions.py +121 -0
- package/src/ai_data_scientist/anomaly_detection.py +39 -0
- package/src/ai_data_scientist/automl.py +109 -0
- package/src/ai_data_scientist/cleaning.py +56 -0
- package/src/ai_data_scientist/cli.py +90 -0
- package/src/ai_data_scientist/clustering.py +54 -0
- package/src/ai_data_scientist/dashboard.py +33 -0
- package/src/ai_data_scientist/data_definition.py +100 -0
- package/src/ai_data_scientist/data_quality.py +164 -0
- package/src/ai_data_scientist/dataset_validation.py +135 -0
- package/src/ai_data_scientist/dependency_pins.py +60 -0
- package/src/ai_data_scientist/eda.py +82 -0
- package/src/ai_data_scientist/experiment_evaluation.py +635 -0
- package/src/ai_data_scientist/explainability.py +340 -0
- package/src/ai_data_scientist/feature_engineering.py +163 -0
- package/src/ai_data_scientist/gate_config.py +32 -0
- package/src/ai_data_scientist/ingestion.py +127 -0
- package/src/ai_data_scientist/insight_engine.py +180 -0
- package/src/ai_data_scientist/japanese_nlp.py +43 -0
- package/src/ai_data_scientist/jupyter_launcher.py +137 -0
- package/src/ai_data_scientist/jupyter_mcp_client.py +94 -0
- package/src/ai_data_scientist/language_router.py +28 -0
- package/src/ai_data_scientist/lifecycle.py +221 -0
- package/src/ai_data_scientist/mcp_gateway.py +113 -0
- package/src/ai_data_scientist/mcp_runtime.py +194 -0
- package/src/ai_data_scientist/mcp_transport.py +53 -0
- package/src/ai_data_scientist/ml_modeling.py +451 -0
- package/src/ai_data_scientist/model_tuning.py +104 -0
- package/src/ai_data_scientist/notebook_audit.py +574 -0
- package/src/ai_data_scientist/project_manager.py +243 -0
- package/src/ai_data_scientist/report_export.py +73 -0
- package/src/ai_data_scientist/sensitivity.py +445 -0
- package/src/ai_data_scientist/signal_analysis.py +201 -0
- package/src/ai_data_scientist/skill_packaging.py +40 -0
- package/src/ai_data_scientist/stats_analysis.py +88 -0
- package/src/ai_data_scientist/text_nlp.py +44 -0
- package/src/ai_data_scientist/timeseries.py +68 -0
- package/src/ai_data_scientist/visualization.py +708 -0
- package/src/ai_genomics_scientist/__init__.py +1 -0
- package/src/ai_genomics_scientist/differential_expression.py +147 -0
- package/src/ai_genomics_scientist/dispatch.py +267 -0
- package/src/ai_genomics_scientist/evidence.py +45 -0
- package/src/ai_genomics_scientist/gene_set_enrichment.py +76 -0
- package/src/ai_genomics_scientist/sequence_alignment.py +97 -0
- package/src/ai_genomics_scientist/sequence_features.py +111 -0
- package/src/ai_genomics_scientist/splice_site_scoring.py +66 -0
- package/src/ai_genomics_scientist/validation.py +83 -0
- package/src/ai_genomics_scientist/variant_effect.py +147 -0
- package/src/ai_genomics_scientist/variant_pathogenicity.py +125 -0
- package/src/ai_materials_scientist/__init__.py +0 -0
- package/src/ai_materials_scientist/calphad.py +117 -0
- package/src/ai_materials_scientist/classical_monte_carlo.py +165 -0
- package/src/ai_materials_scientist/crystal_plasticity.py +184 -0
- package/src/ai_materials_scientist/dispatch.py +100 -0
- package/src/ai_materials_scientist/evidence.py +84 -0
- package/src/ai_materials_scientist/fem.py +279 -0
- package/src/ai_materials_scientist/kinetic_monte_carlo.py +145 -0
- package/src/ai_materials_scientist/molecular_dynamics.py +240 -0
- package/src/ai_materials_scientist/phase_field.py +167 -0
- package/src/ai_materials_scientist/validation.py +70 -0
- package/src/ai_scientist/__init__.py +1 -0
- package/src/ai_scientist/completion_gate.py +15 -0
- package/src/ai_scientist/data_analysis.py +46 -0
- package/src/ai_scientist/evidence_registry.py +99 -0
- package/src/ai_scientist/experimental_design.py +20 -0
- package/src/ai_scientist/language.py +14 -0
- package/src/ai_scientist/latex_renderer.py +41 -0
- package/src/ai_scientist/literature_review.py +37 -0
- package/src/ai_scientist/manifest.py +87 -0
- package/src/ai_scientist/manuscript.py +94 -0
- package/src/ai_scientist/mcp_config.py +76 -0
- package/src/ai_scientist/mcp_external.py +42 -0
- package/src/ai_scientist/mcp_failures.py +23 -0
- package/src/ai_scientist/mcp_gateway.py +38 -0
- package/src/ai_scientist/mcp_managed.py +180 -0
- package/src/ai_scientist/npm_packaging.py +49 -0
- package/src/ai_scientist/orchestrator.py +133 -0
- package/src/ai_scientist/peer_review.py +60 -0
- package/src/ai_scientist/phase_gate.py +74 -0
- package/src/ai_scientist/phase_state.py +230 -0
- package/src/ai_scientist/presentation.py +56 -0
- package/src/ai_scientist/project_config.py +31 -0
- package/src/ai_scientist/project_handle.py +74 -0
- package/src/ai_scientist/reproducibility.py +20 -0
- package/src/ai_scientist/research_planning.py +20 -0
- package/src/ai_scientist/skill_invocation.py +21 -0
- package/src/ai_scientist/tdd_gate.py +99 -0
- package/src/ai_structural_biology_scientist/__init__.py +0 -0
- package/src/ai_structural_biology_scientist/contact_map.py +87 -0
- package/src/ai_structural_biology_scientist/dispatch.py +269 -0
- package/src/ai_structural_biology_scientist/evidence.py +43 -0
- package/src/ai_structural_biology_scientist/hydrophobicity.py +101 -0
- package/src/ai_structural_biology_scientist/protein_docking_score.py +104 -0
- package/src/ai_structural_biology_scientist/secondary_structure.py +95 -0
- package/src/ai_structural_biology_scientist/structural_similarity.py +74 -0
- package/src/ai_structural_biology_scientist/validation.py +100 -0
|
@@ -0,0 +1,97 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ai-chemistry-scientist
|
|
3
|
+
description: "Use when a user asks, in Japanese or English, to run a cheminformatics module: molecular descriptor calculation, ADMET heuristic screening, QSAR linear-regression modeling, molecular similarity search, a simplified docking-score heuristic, drug-likeness rule screening, structural-alert screening, molecular formula/exact-mass calculation, heuristic bioactivity classification, SMILES salt removal/structure standardization, or chemical structure format conversion (SMILES/InChI/InChIKey/Molblock). 化学情報学固有の処理(分子記述子計算、ADMETヒューリスティックスクリーニング、QSARモデリング、分子類似性検索、ドッキングスコアヒューリスティック、薬物らしさルールスクリーニング、構造アラートスクリーニング、分子式・正確質量計算、生物活性ヒューリスティック分類、SMILES塩除去・構造標準化、化学構造フォーマット変換)を実行する際に使用。"
|
|
4
|
+
---
|
|
5
|
+
# AI Chemistry Scientist / AI化学科学者
|
|
6
|
+
|
|
7
|
+
Respond in the user's input language (日本語 / English) for every user-facing
|
|
8
|
+
message (REQ-ACHEM-001). Dispatch to exactly one of the 11 supported
|
|
9
|
+
cheminformatics modules per request; never mix modules in a single run
|
|
10
|
+
(REQ-ACHEM-002).
|
|
11
|
+
|
|
12
|
+
## Workflow / 手順
|
|
13
|
+
1. **Detect language and match the request** — call
|
|
14
|
+
`ai_chemistry_scientist.dispatch.dispatch(request_text)`, which loads
|
|
15
|
+
`.github/skills/ai-chemistry-scientist/manifest.json` and matches the
|
|
16
|
+
request text against each module's registered English/Japanese
|
|
17
|
+
name/synonym list.
|
|
18
|
+
- Exactly one match → invoke that module's handler function.
|
|
19
|
+
- Two or more distinct modules matched → ask a clarification question
|
|
20
|
+
listing every matched candidate; invoke no module.
|
|
21
|
+
- No match → return a rejection message; invoke no module.
|
|
22
|
+
2. **Validate parameters before any computation** — each module validates
|
|
23
|
+
its resolved parameters (e.g. SMILES must parse to a valid RDKit
|
|
24
|
+
molecule, training sets must have at least 5 compounds with a full-rank
|
|
25
|
+
descriptor matrix, `k` must be an integer in `[1, 20]`) before running
|
|
26
|
+
any computation (REQ-ACHEM-003); on violation it rejects the run naming
|
|
27
|
+
the parameter and the violated constraint. `molecular-descriptors`
|
|
28
|
+
instead validates and computes per-item, continuing past individual
|
|
29
|
+
rejected entries rather than failing the whole batch.
|
|
30
|
+
3. **Record reproducible run evidence** — every completed run returns a
|
|
31
|
+
`RunRecord` with exactly `metadata`, `parameters`, and `result`
|
|
32
|
+
(REQ-ACHEM-004), including the RDKit version used (and the
|
|
33
|
+
scikit-learn version for QSAR modeling).
|
|
34
|
+
|
|
35
|
+
## Supported modules / 対応モジュール
|
|
36
|
+
| Method | English | 日本語 |
|
|
37
|
+
| --- | --- | --- |
|
|
38
|
+
| Molecular descriptors | molecular descriptors / descriptor calculation | 分子記述子 / 記述子計算 |
|
|
39
|
+
| ADMET prediction | admet prediction / admet screening | ADMET予測 / ADMETスクリーニング |
|
|
40
|
+
| QSAR modeling | qsar modeling / qsar regression | QSARモデリング / QSAR回帰 |
|
|
41
|
+
| Molecular similarity | molecular similarity / similarity search | 分子類似性 / 類似性検索 |
|
|
42
|
+
| Docking score | docking score / docking simulation | ドッキングスコア / ドッキングシミュレーション |
|
|
43
|
+
| Drug-likeness rules | drug-likeness rules / drug-likeness screening | 薬物らしさルール / 薬物らしさルールスクリーニング |
|
|
44
|
+
| Structural alerts | structural alerts / structural alert screening | 構造アラート / 構造アラートスクリーニング |
|
|
45
|
+
| Formula & exact mass | molecular formula and exact mass / formula and exact mass | 分子式と正確質量 / 分子式・正確質量 |
|
|
46
|
+
| Bioactivity classification | target-class activity classification / bioactivity classification | 標的クラス活性分類 / 生物活性分類 |
|
|
47
|
+
| Salt removal / standardization | salt removal / structure standardization | 塩除去 / 構造標準化 |
|
|
48
|
+
| Structure format conversion | structure format conversion / chemical format conversion | 構造フォーマット変換 / 化学構造フォーマット変換 |
|
|
49
|
+
|
|
50
|
+
## Important limitations / 重要な制限
|
|
51
|
+
ADMET prediction, docking-score, structural-alert, bioactivity
|
|
52
|
+
classification, and salt-removal results always carry a fixed, verbatim
|
|
53
|
+
limitation label stating they are heuristic approximations, not validated
|
|
54
|
+
predictions, full structural-alert catalogs, physically accurate
|
|
55
|
+
simulations, or a curated salt/solvent lookup table:
|
|
56
|
+
- ADMET (en): "Heuristic only: not a physically or clinically validated
|
|
57
|
+
ADMET prediction."
|
|
58
|
+
- ADMET (ja): "ヒューリスティックのみ: 物理的または臨床的に検証されたADMET
|
|
59
|
+
予測ではありません。"
|
|
60
|
+
- Docking (en): "Heuristic only: not a physically accurate docking
|
|
61
|
+
simulation (no 3D conformer generation, no energy function)."
|
|
62
|
+
- Docking (ja): "ヒューリスティックのみ: 物理的に正確なドッキングシミュレー
|
|
63
|
+
ションではありません(3D配座生成・エネルギー関数なし)。"
|
|
64
|
+
- Structural alerts (en): "Heuristic only: a small fixed illustrative
|
|
65
|
+
SMARTS alert list, not the validated PAINS/Brenk filter catalog."
|
|
66
|
+
- Structural alerts (ja): "ヒューリスティックのみ: 固定の小規模な例示用SMARTS
|
|
67
|
+
アラート一覧であり、検証済みのPAINS/Brenkフィルタ・カタログではない。"
|
|
68
|
+
- Bioactivity classification (en): "Heuristic only: not a ChEMBL-trained
|
|
69
|
+
or experimentally validated bioactivity classifier."
|
|
70
|
+
- Bioactivity classification (ja): "ヒューリスティックのみ:
|
|
71
|
+
ChEMBLで学習済みでも実験的に検証済みでもない生物活性分類器ではない。"
|
|
72
|
+
- Salt removal (en): "Heuristic only: treats every disconnected fragment
|
|
73
|
+
except the one with the greatest heavy-atom count as removable
|
|
74
|
+
salt/solvent; not always chemically correct (e.g. for a genuine
|
|
75
|
+
covalent multi-component cocrystal)."
|
|
76
|
+
- Salt removal (ja): "ヒューリスティックのみ: 最大重原子数を持つフラグメント
|
|
77
|
+
以外のすべての分離フラグメントを除去可能な塩・溶媒として扱うが、常に
|
|
78
|
+
化学的に正しいとは限らない(例: 真の共有結合性多成分共結晶の場合)。"
|
|
79
|
+
|
|
80
|
+
Structure format conversion carries no limitation label: it is a
|
|
81
|
+
deterministic local RDKit parse/render operation (SMILES/InChI/Molblock as
|
|
82
|
+
input; SMILES/InChI/InChIKey/Molblock as output), not a predictive
|
|
83
|
+
heuristic. It does not promise bytewise, metadata, or coordinate
|
|
84
|
+
preservation across formats that do not share that information (e.g. a
|
|
85
|
+
Molblock's 2D/3D coordinates have no SMILES/InChI equivalent), and
|
|
86
|
+
performs no network call or PubChem/ChEMBL/DrugBank identifier lookup of
|
|
87
|
+
any kind.
|
|
88
|
+
|
|
89
|
+
## Scope boundary / 対象外
|
|
90
|
+
Generic, domain-agnostic analysis stays in `ai-data-scientist`; multi-phase
|
|
91
|
+
research orchestration stays in `ai-scientist`; materials-science
|
|
92
|
+
simulation stays in `ai-materials-scientist`. This skill owns
|
|
93
|
+
cheminformatics only, implemented with RDKit/numpy/scikit-learn (no
|
|
94
|
+
external docking engines, no 3D conformer generation, no network calls).
|
|
95
|
+
|
|
96
|
+
Traceability: REQ-ACHEM-001 through REQ-ACHEM-110, DES-ACHEM-001 through
|
|
97
|
+
DES-ACHEM-110 (`.musubix/features/ai-chemistry-scientist/`).
|
|
@@ -0,0 +1,156 @@
|
|
|
1
|
+
{
|
|
2
|
+
"molecular-descriptors": {
|
|
3
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
4
|
+
"functionName": "handle_molecular_descriptors",
|
|
5
|
+
"names": {
|
|
6
|
+
"en": [
|
|
7
|
+
"molecular descriptors",
|
|
8
|
+
"descriptor calculation"
|
|
9
|
+
],
|
|
10
|
+
"ja": [
|
|
11
|
+
"分子記述子",
|
|
12
|
+
"記述子計算"
|
|
13
|
+
]
|
|
14
|
+
}
|
|
15
|
+
},
|
|
16
|
+
"admet-prediction": {
|
|
17
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
18
|
+
"functionName": "handle_admet_prediction",
|
|
19
|
+
"names": {
|
|
20
|
+
"en": [
|
|
21
|
+
"admet prediction",
|
|
22
|
+
"admet screening"
|
|
23
|
+
],
|
|
24
|
+
"ja": [
|
|
25
|
+
"ADMET予測",
|
|
26
|
+
"ADMETスクリーニング"
|
|
27
|
+
]
|
|
28
|
+
}
|
|
29
|
+
},
|
|
30
|
+
"qsar-modeling": {
|
|
31
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
32
|
+
"functionName": "handle_qsar_modeling",
|
|
33
|
+
"names": {
|
|
34
|
+
"en": [
|
|
35
|
+
"qsar modeling",
|
|
36
|
+
"qsar regression"
|
|
37
|
+
],
|
|
38
|
+
"ja": [
|
|
39
|
+
"QSARモデリング",
|
|
40
|
+
"QSAR回帰"
|
|
41
|
+
]
|
|
42
|
+
}
|
|
43
|
+
},
|
|
44
|
+
"molecular-similarity": {
|
|
45
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
46
|
+
"functionName": "handle_molecular_similarity",
|
|
47
|
+
"names": {
|
|
48
|
+
"en": [
|
|
49
|
+
"molecular similarity",
|
|
50
|
+
"similarity search"
|
|
51
|
+
],
|
|
52
|
+
"ja": [
|
|
53
|
+
"分子類似性",
|
|
54
|
+
"類似性検索"
|
|
55
|
+
]
|
|
56
|
+
}
|
|
57
|
+
},
|
|
58
|
+
"docking-score": {
|
|
59
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
60
|
+
"functionName": "handle_docking_score",
|
|
61
|
+
"names": {
|
|
62
|
+
"en": [
|
|
63
|
+
"docking score",
|
|
64
|
+
"docking simulation"
|
|
65
|
+
],
|
|
66
|
+
"ja": [
|
|
67
|
+
"ドッキングスコア",
|
|
68
|
+
"ドッキングシミュレーション"
|
|
69
|
+
]
|
|
70
|
+
}
|
|
71
|
+
},
|
|
72
|
+
"drug-likeness-rules": {
|
|
73
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
74
|
+
"functionName": "handle_drug_likeness_rules",
|
|
75
|
+
"names": {
|
|
76
|
+
"en": [
|
|
77
|
+
"drug-likeness rules",
|
|
78
|
+
"drug-likeness screening"
|
|
79
|
+
],
|
|
80
|
+
"ja": [
|
|
81
|
+
"薬物らしさルール",
|
|
82
|
+
"薬物らしさルールスクリーニング"
|
|
83
|
+
]
|
|
84
|
+
}
|
|
85
|
+
},
|
|
86
|
+
"structural-alerts": {
|
|
87
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
88
|
+
"functionName": "handle_structural_alerts",
|
|
89
|
+
"names": {
|
|
90
|
+
"en": [
|
|
91
|
+
"structural alerts",
|
|
92
|
+
"structural alert screening"
|
|
93
|
+
],
|
|
94
|
+
"ja": [
|
|
95
|
+
"構造アラート",
|
|
96
|
+
"構造アラートスクリーニング"
|
|
97
|
+
]
|
|
98
|
+
}
|
|
99
|
+
},
|
|
100
|
+
"molecular-formula-mass": {
|
|
101
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
102
|
+
"functionName": "handle_molecular_formula_mass",
|
|
103
|
+
"names": {
|
|
104
|
+
"en": [
|
|
105
|
+
"molecular formula and exact mass",
|
|
106
|
+
"formula and exact mass"
|
|
107
|
+
],
|
|
108
|
+
"ja": [
|
|
109
|
+
"分子式と正確質量",
|
|
110
|
+
"分子式・正確質量"
|
|
111
|
+
]
|
|
112
|
+
}
|
|
113
|
+
},
|
|
114
|
+
"bioactivity-classification": {
|
|
115
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
116
|
+
"functionName": "handle_bioactivity_classification",
|
|
117
|
+
"names": {
|
|
118
|
+
"en": [
|
|
119
|
+
"target-class activity classification",
|
|
120
|
+
"bioactivity classification"
|
|
121
|
+
],
|
|
122
|
+
"ja": [
|
|
123
|
+
"標的クラス活性分類",
|
|
124
|
+
"生物活性分類"
|
|
125
|
+
]
|
|
126
|
+
}
|
|
127
|
+
},
|
|
128
|
+
"salt-removal": {
|
|
129
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
130
|
+
"functionName": "handle_salt_removal",
|
|
131
|
+
"names": {
|
|
132
|
+
"en": [
|
|
133
|
+
"salt removal",
|
|
134
|
+
"structure standardization"
|
|
135
|
+
],
|
|
136
|
+
"ja": [
|
|
137
|
+
"塩除去",
|
|
138
|
+
"構造標準化"
|
|
139
|
+
]
|
|
140
|
+
}
|
|
141
|
+
},
|
|
142
|
+
"structure-format-conversion": {
|
|
143
|
+
"modulePath": "ai_chemistry_scientist.dispatch",
|
|
144
|
+
"functionName": "handle_structure_format_conversion",
|
|
145
|
+
"names": {
|
|
146
|
+
"en": [
|
|
147
|
+
"structure format conversion",
|
|
148
|
+
"chemical format conversion"
|
|
149
|
+
],
|
|
150
|
+
"ja": [
|
|
151
|
+
"構造フォーマット変換",
|
|
152
|
+
"化学構造フォーマット変換"
|
|
153
|
+
]
|
|
154
|
+
}
|
|
155
|
+
}
|
|
156
|
+
}
|
|
@@ -0,0 +1,330 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ai-data-scientist
|
|
3
|
+
description: "Use when a user asks, in Japanese or English, to load, clean, explore, analyze, visualize, or draw insights from a dataset via Jupyter. データ分析・可視化・統計・Insight抽出をJupyter上で自然言語で行う際に使用。"
|
|
4
|
+
---
|
|
5
|
+
# AI Data Scientist / AIデータサイエンティスト
|
|
6
|
+
|
|
7
|
+
Respond in the user's input language (日本語 / English) for every user-facing
|
|
8
|
+
message, per REQ-AIDS-001. Route execution exclusively through the Jupyter
|
|
9
|
+
MCP (Datalayer `jupyter-mcp-server`) tools configured for this project;
|
|
10
|
+
never execute analysis code outside that path (REQ-AIDS-003). The `mcp` and
|
|
11
|
+
`jupyter-mcp-server` packages are pinned to a narrow, verified-interoperable
|
|
12
|
+
minor-version range (`>=2.2,<2.3` for both) rather than an open-ended range,
|
|
13
|
+
so client/server MCP protocol negotiation cannot silently drift to an
|
|
14
|
+
untested combination (REQ-AIDS-050 / DES-AIDS-038); see
|
|
15
|
+
`ai_data_scientist.dependency_pins.get_dependency_specifier` for the
|
|
16
|
+
programmatic check.
|
|
17
|
+
|
|
18
|
+
## Scope / 対象範囲 (MVP)
|
|
19
|
+
Data ingestion (CSV/Excel/DB/API), cleaning, exploratory data analysis,
|
|
20
|
+
statistical analysis, visualization, and reasoning-based insight generation
|
|
21
|
+
whose evidence is always recorded inside the project notebook. See
|
|
22
|
+
`.musubix/features/ai-data-scientist/requirements.md` for the full,
|
|
23
|
+
authoritative requirement set (REQ-AIDS-001–014, 027–032). ML-extension
|
|
24
|
+
capabilities (clustering, AutoML, GiNZA-based Japanese NLP, etc.) live in
|
|
25
|
+
the separate `ai-data-scientist-ml` feature and are out of scope here.
|
|
26
|
+
|
|
27
|
+
## Workflow / 手順
|
|
28
|
+
1. **Resolve project & notebook** — call
|
|
29
|
+
`ai_data_scientist.project_manager.resolve_project(name)` to validate the
|
|
30
|
+
project identifier (ADR-0005 slug policy), then
|
|
31
|
+
`ensure_notebook(handle)` to create or reuse
|
|
32
|
+
`projects/<name>/notebooks/<name>.ipynb` (REQ-AIDS-002/028). The default
|
|
33
|
+
`projects/` root is stable across kernel working-directory changes
|
|
34
|
+
(REQ-AIDS-044): set `AI_DATA_SCIENTIST_PROJECTS_ROOT` to an absolute path
|
|
35
|
+
to pin the workspace root explicitly, especially if the kernel may `cd`
|
|
36
|
+
into a dataset or notebook directory during the session.
|
|
37
|
+
2. **Detect instruction language** — call
|
|
38
|
+
`ai_data_scientist.language_router.detect_language(instruction_text)` and
|
|
39
|
+
use its result for every reply and inserted markdown cell in this turn
|
|
40
|
+
(REQ-AIDS-001).
|
|
41
|
+
3. **Execute analysis code via Jupyter MCP** — `run_and_record` is a
|
|
42
|
+
**host-side** API: it takes an in-process `MCPClient` object (anything
|
|
43
|
+
exposing `.execute(code) -> dict`) and is meant for a Python process that
|
|
44
|
+
holds its own direct connection to the Jupyter MCP server/kernel. If
|
|
45
|
+
your own process has such a client, call
|
|
46
|
+
`ai_data_scientist.mcp_gateway.run_and_record(client, handle, code,
|
|
47
|
+
timeout_ms=30000)`; it routes execution only through that client,
|
|
48
|
+
enforces the timeout, and appends the executed cell only on success —
|
|
49
|
+
never a partial/corrupted cell (REQ-AIDS-003/030/031).
|
|
50
|
+
|
|
51
|
+
If instead you are a Copilot CLI (or similar) agent that executes code
|
|
52
|
+
by calling Jupyter MCP tools directly (e.g. `insert_execute_code_cell`)
|
|
53
|
+
— with no in-process `MCPClient` instance of your own to pass in — do
|
|
54
|
+
**not** try to construct/pass a client whose transport would submit work
|
|
55
|
+
synchronously back into the same kernel that is executing it (risks
|
|
56
|
+
deadlock/reentrancy). `run_and_record` itself cannot run in that case, so
|
|
57
|
+
you are responsible for reproducing its guarantees yourself: no
|
|
58
|
+
partial/corrupted cell on failure or timeout, and the user is notified on
|
|
59
|
+
failure (REQ-AIDS-003/030/031) — lifecycle calls alone do **not** provide
|
|
60
|
+
this; they only track run/cancellation state. At minimum: call
|
|
61
|
+
`register_run(run_id, handle.notebook_path)`, then
|
|
62
|
+
`mark_execution_start(run_id)` before invoking the MCP tool. On success,
|
|
63
|
+
call `mark_execution_end(run_id)` and then `mark_completed(run_id)`. On a
|
|
64
|
+
failure that truly corresponds to the kernel execution having stopped, call
|
|
65
|
+
`mark_execution_end(run_id)` and `mark_failed(run_id)` **and** ensure the
|
|
66
|
+
MCP tool did not leave a partial/corrupted cell (e.g. delete it if it did)
|
|
67
|
+
and surface the failure to the user — do not rely on
|
|
68
|
+
`insert_execute_code_cell` alone to guarantee this (see step 11 for the
|
|
69
|
+
full lifecycle-call set, including `mark_write_start`/`mark_write_end`
|
|
70
|
+
if the same tool call also writes the notebook). If a direct-MCP-tool
|
|
71
|
+
timeout only means the host stopped waiting for output and the kernel may
|
|
72
|
+
still be running, do **not** immediately clear lifecycle execution state as
|
|
73
|
+
though the cell had finished; first confirm real completion by the methods
|
|
74
|
+
in step 11, then close the lifecycle record consistently with the actual
|
|
75
|
+
outcome. Note the fallback explicitly in the notebook or hand-off notes
|
|
76
|
+
(e.g. "executed via MCP tool, not run_and_record") so later audits are not
|
|
77
|
+
misled into assuming a host-side client was used.
|
|
78
|
+
|
|
79
|
+
**Known ordering constraint (jupyter-mcp-server cache coherency,
|
|
80
|
+
Issue #75)**: `insert_execute_code_cell` keeps its own in-memory cached
|
|
81
|
+
notebook model, while `insight_engine.record_insight`/
|
|
82
|
+
`visualization.record_chart` (step 9/8) write directly to the notebook
|
|
83
|
+
file on disk, bypassing that cache. If a direct-write call is
|
|
84
|
+
immediately followed by another `insert_execute_code_cell` call in the
|
|
85
|
+
same session, jupyter-mcp-server's cache can be stale relative to the
|
|
86
|
+
file it just missed, so the next `insert_execute_code_cell` call may
|
|
87
|
+
report success without actually appending a cell (silently dropped), and
|
|
88
|
+
a caller that retries by polling cell count can end up inserting a
|
|
89
|
+
genuine duplicate once the cache catches up. Until this cache-coherency
|
|
90
|
+
issue is fixed upstream, there is **no supported safe way to interleave
|
|
91
|
+
them**: in a given MCP/Jupyter session, complete every
|
|
92
|
+
`insert_execute_code_cell` call you need before making any
|
|
93
|
+
`record_insight`/`record_chart` direct write, and once a direct write has
|
|
94
|
+
occurred, do not issue another `insert_execute_code_cell` call in that
|
|
95
|
+
same session. Retrying an `insert_execute_code_cell` call after a direct
|
|
96
|
+
write (e.g. by polling cell count) does not reliably avoid the duplicate
|
|
97
|
+
described above and must not be used as a workaround. If an insight
|
|
98
|
+
needs to be recorded between two dependent live-execution steps,
|
|
99
|
+
restructure the work into a new session/kernel handle for the
|
|
100
|
+
remaining `insert_execute_code_cell` calls instead.
|
|
101
|
+
4. **Ingest data** with `ai_data_scientist.ingestion.ingest(source_spec,
|
|
102
|
+
fetcher=..., allowlist=..., row_limit=...)` for CSV, Excel, database, or
|
|
103
|
+
API sources; non-allowlisted hosts and over-limit responses are rejected
|
|
104
|
+
or truncated before load (REQ-AIDS-014/032). Write any downloaded/staged
|
|
105
|
+
dataset file under `ai_data_scientist.project_manager.ensure_data_dir(handle)`
|
|
106
|
+
(i.e. `handle.data_dir`), not a hand-rolled `projects/<name>/data` relative
|
|
107
|
+
path — `data_dir` is anchored to the same stable workspace root as
|
|
108
|
+
`notebook_path`, so it stays correct even if the kernel's working
|
|
109
|
+
directory has drifted into the notebook's own directory (REQ-AIDS-049).
|
|
110
|
+
5. **Clean data** with `ai_data_scientist.cleaning.clean_dataset(df,
|
|
111
|
+
operation=...)` (`drop_duplicates` | `drop_na` | `fillna`); report the
|
|
112
|
+
returned row/column impact to the user (REQ-AIDS-004).
|
|
113
|
+
6. **Explore data** with `ai_data_scientist.eda.explore(df)` for dtypes,
|
|
114
|
+
non-null counts, and summary statistics (REQ-AIDS-005). The returned
|
|
115
|
+
`EDAReport` object exposes two further attributes — `EDAReport.missing_summary`
|
|
116
|
+
(per-column count/ratio) and, for categorical columns,
|
|
117
|
+
`EDAReport.categorical_summary` (unique_count and top `top_n` values with
|
|
118
|
+
counts/ratios, `truncated` flag) (REQ-AIDS-043) — read them off the
|
|
119
|
+
object returned by `explore()`; there is no separate
|
|
120
|
+
`eda.missing_summary(...)` function to import.
|
|
121
|
+
7. **Run statistics** with
|
|
122
|
+
`ai_data_scientist.stats_analysis.correlation(df, col_a, col_b,
|
|
123
|
+
language=...)`; write the resulting statistic in a code cell and its
|
|
124
|
+
`interpretation` in an adjacent markdown cell (REQ-AIDS-006).
|
|
125
|
+
8. **Visualize** with `ai_data_scientist.visualization.render_chart(df,
|
|
126
|
+
kind=..., x=..., y=..., title=..., xlabel=..., ylabel=...)` and
|
|
127
|
+
`build_image_output(png_bytes)`, then append via
|
|
128
|
+
`project_manager.enqueue_write` so the image MIME bundle is persisted
|
|
129
|
+
in the notebook JSON (REQ-AIDS-007). `render_chart`/`record_chart` render
|
|
130
|
+
locally and never execute the stored code string against the live
|
|
131
|
+
Jupyter kernel (REQ-AIDS-040): only reference variables already
|
|
132
|
+
established by a prior successful MCP-routed execution (via
|
|
133
|
+
`run_and_record`, or the direct-MCP-tool fallback from step 3) in that
|
|
134
|
+
code string, so the notebook stays consistent if a human re-runs it
|
|
135
|
+
top-to-bottom later.
|
|
136
|
+
Pass Japanese (or other non-ASCII) text in `title`/`xlabel`/`ylabel`
|
|
137
|
+
freely: `render_chart` automatically switches to a bundled
|
|
138
|
+
Japanese-capable font the first time such text appears in a process, so
|
|
139
|
+
it renders as legible glyphs instead of mojibake/placeholder boxes,
|
|
140
|
+
regardless of what fonts are installed on the host (REQ-AIDS-046).
|
|
141
|
+
**Concurrent-write risk**: calling `enqueue_write` directly against a
|
|
142
|
+
notebook file that a human has simultaneously open in a Jupyter MCP
|
|
143
|
+
session can be silently overwritten when that MCP session later saves
|
|
144
|
+
its own in-memory copy (REQ-AIDS-059). Prefer routing writes through the
|
|
145
|
+
active MCP session instead, or ask the human to pause MCP-side saves
|
|
146
|
+
while this skill writes to the notebook file directly.
|
|
147
|
+
9. **Record insights with evidence** — only after executing the
|
|
148
|
+
evidentiary cell, derive `cited_value` from the executed cell's actual
|
|
149
|
+
output with `ai_data_scientist.insight_engine.extract_cited_value(result,
|
|
150
|
+
pattern)` (REQ-AIDS-039) rather than hand-transcribing or rounding a
|
|
151
|
+
number, then call
|
|
152
|
+
`ai_data_scientist.insight_engine.record_insight(handle, insight_text,
|
|
153
|
+
evidence_execution_count, cited_value, claim_type, language=...)`. It
|
|
154
|
+
verifies the cited value actually appears in that cell's output, embeds
|
|
155
|
+
a structured `evidence` manifest (execution count, cited value, claim
|
|
156
|
+
type — ADR-0003) in the markdown cell, and raises
|
|
157
|
+
`EvidenceMissingError` (which you must report to the user, not silently
|
|
158
|
+
swallow) instead of writing an insight with no notebook-backed rationale
|
|
159
|
+
(REQ-AIDS-009/010/027).
|
|
160
|
+
10. **Audit before handing off** — before telling the user the notebook is
|
|
161
|
+
complete, run `ai_data_scientist.notebook_audit.audit_notebook(path)`
|
|
162
|
+
(REQ-AIDS-045). It is read-only (never rewrites the notebook) and
|
|
163
|
+
reports unexecuted/error code cells and any insight-like markdown cell
|
|
164
|
+
whose evidence manifest is missing or no longer resolves to a real
|
|
165
|
+
executed output. `path` may be the same workspace-root-relative path
|
|
166
|
+
used elsewhere (e.g. `handle.notebook_path`); it resolves against the
|
|
167
|
+
stable workspace root even if the kernel's cwd has drifted into the
|
|
168
|
+
notebook's own directory, and reports a distinct, clearly-worded
|
|
169
|
+
finding instead of a parse-failure finding when the path genuinely
|
|
170
|
+
cannot be found (REQ-AIDS-047). Because the audit is typically run from
|
|
171
|
+
inside the notebook's own still-running final cell, that one trailing,
|
|
172
|
+
not-yet-executed cell is excluded from the unexecuted-cell findings
|
|
173
|
+
when its source invokes `audit_notebook` — only a different unexecuted
|
|
174
|
+
or error-producing cell will fail the audit (REQ-AIDS-048). If
|
|
175
|
+
`report.ok` is `False`, fix the underlying cells or insights before
|
|
176
|
+
finishing, rather than reporting success with unresolved findings. The
|
|
177
|
+
same check is available from the shell as
|
|
178
|
+
`ai-data-scientist validate-notebook <path>` for CI or manual spot
|
|
179
|
+
checks. For a stricter pre-handoff pass, call
|
|
180
|
+
`ai_data_scientist.notebook_audit.audit_notebook(path, visual_audit=True)`
|
|
181
|
+
(REQ-AIDS-053) to additionally flag charts with missing glyph metadata,
|
|
182
|
+
missing title/axis-label/legend metadata, or a near-empty-looking image
|
|
183
|
+
(a byte-density heuristic, not an exact pixel scan — see
|
|
184
|
+
`audit_visual_outputs`), surfaced in `report.visual_findings`.
|
|
185
|
+
11. **Manage long-running or cancellable work** — for any analysis that may
|
|
186
|
+
run long enough that a user wants to cancel it, or that writes to the
|
|
187
|
+
notebook from a background task, call
|
|
188
|
+
`ai_data_scientist.lifecycle.register_run(run_id, handle.notebook_path)`
|
|
189
|
+
first (both arguments are required), call
|
|
190
|
+
`mark_execution_start`/`mark_execution_end` (and
|
|
191
|
+
`mark_write_start`/`mark_write_end`) around the corresponding work, check
|
|
192
|
+
`is_cancel_requested(run_id)` between steps, and call
|
|
193
|
+
`mark_completed`/`mark_failed` at the end. A caller elsewhere can call
|
|
194
|
+
`request_cancel(run_id)` (cooperative only — it cannot interrupt a cell
|
|
195
|
+
already executing in the kernel) and `wait_for_quiescence(run_id)` to
|
|
196
|
+
wait until that run becomes quiescent **or** the supplied timeout elapses;
|
|
197
|
+
inspect the returned status before assuming execution/write activity has
|
|
198
|
+
fully settled (REQ-AIDS-051). For Copilot CLI / direct-Jupyter-MCP operation,
|
|
199
|
+
`insert_execute_code_cell`-style tools may stop waiting for output after
|
|
200
|
+
roughly 120 seconds even while the kernel keeps running the cell; treat
|
|
201
|
+
that as **execution still unconfirmed**, not as proof that the run is
|
|
202
|
+
finished. Therefore:
|
|
203
|
+
- Split any work likely to exceed that wait window into restartable units
|
|
204
|
+
such as fold-by-fold training, stage-by-stage preprocessing, or
|
|
205
|
+
checkpointed batch chunks, with each unit persisting its own durable
|
|
206
|
+
output before the next unit begins.
|
|
207
|
+
- After any host/tool timeout, **do not execute another cell yet**. First
|
|
208
|
+
confirm completion with an actual idle/quiescent signal: the active MCP
|
|
209
|
+
session reports the kernel as idle if that capability is truly available
|
|
210
|
+
in your environment, and/or your own orchestration still shows full
|
|
211
|
+
lifecycle quiescence. `lifecycle.get_run_status` /
|
|
212
|
+
`wait_for_quiescence` are useful only when your orchestration keeps the
|
|
213
|
+
lifecycle counters aligned with the kernel's real completion state; in
|
|
214
|
+
the simple direct-tool fallback where a timeout path immediately runs
|
|
215
|
+
`mark_execution_end`/`mark_failed`, those lifecycle calls alone are **not**
|
|
216
|
+
sufficient proof that the kernel is idle. When you do rely on lifecycle
|
|
217
|
+
state, require full quiescence (`active_cell_executions == 0`,
|
|
218
|
+
`pending_notebook_writes == 0`, and `locks_held == 0`), not just the
|
|
219
|
+
first two counters. A durable result/checkpoint file is valuable for
|
|
220
|
+
resume/restart, but by itself does **not** prove that the timed-out cell
|
|
221
|
+
has finished; if no real idle/quiescent signal is available yet,
|
|
222
|
+
completion remains unconfirmed and you still must not submit another
|
|
223
|
+
cell on that kernel. In that situation, treat the current kernel/session
|
|
224
|
+
as unsafe to reuse for follow-up cells and resume only from the latest
|
|
225
|
+
durable checkpoint/cache/result shard in a fresh kernel/session (or
|
|
226
|
+
another execution context whose clean idle state you can actually
|
|
227
|
+
verify).
|
|
228
|
+
- If the run was cancelled or interrupted, resume from the latest durable
|
|
229
|
+
cache/checkpoint/result shard (for example a completed fold's model or
|
|
230
|
+
metrics file) instead of restarting the whole job in one giant cell.
|
|
231
|
+
Prefer idempotent chunks that can be skipped safely when their output
|
|
232
|
+
already exists.
|
|
233
|
+
12. **Record what each field actually means** — before relying on a
|
|
234
|
+
column's unit or definition in an insight, build a
|
|
235
|
+
`ai_data_scientist.data_definition.DataDefinitionManifest` via
|
|
236
|
+
`build_manifest(...)`, recording each semantic field's value plus a
|
|
237
|
+
status of `verified` (confirmed against a data dictionary/source),
|
|
238
|
+
`inferred` (inferred from values/name), `reported` (as stated by the
|
|
239
|
+
user), or `unknown`. Call `manifest.unresolved_fields()` and surface any
|
|
240
|
+
`unknown`/`inferred` fields as caveats rather than presenting them as
|
|
241
|
+
verified facts (REQ-AIDS-052).
|
|
242
|
+
13. **Record assumptions and causal scope** — for any conclusion that
|
|
243
|
+
depends on a non-obvious analytical choice (sampling, preprocessing,
|
|
244
|
+
causal interpretation), build an
|
|
245
|
+
`ai_data_scientist.analysis_assumptions.AnalysisAssumptionManifest`
|
|
246
|
+
(one `Assumption` per choice, with a `status` of `verified`/`tested`/
|
|
247
|
+
`assumed`/`rejected`, and `causal_scope` of `descriptive`/
|
|
248
|
+
`associational`/`causal`). Call `check_manifest(manifest)` and report
|
|
249
|
+
every returned finding to the user — in particular, never present a
|
|
250
|
+
`causal` conclusion without a `tested`/`verified` identification
|
|
251
|
+
assumption, and flag any conclusion-critical assumption left `assumed`
|
|
252
|
+
or `rejected` (REQ-AIDS-054).
|
|
253
|
+
14. **Check for semantic data-quality issues** — beyond the statistical
|
|
254
|
+
z-score check in `anomaly_detection.detect_anomalies`, use
|
|
255
|
+
`ai_data_scientist.data_quality.detect_anomalies(df, schema={...})` to
|
|
256
|
+
check declarative per-column constraints (`min`/`max`, `allowed`
|
|
257
|
+
categories, `not_null`, `unique`). When an independent reference
|
|
258
|
+
dataset is available, call
|
|
259
|
+
`validate_anomalies(primary, reference, columns, tolerance=...)` to
|
|
260
|
+
confirm a detected anomaly isn't an artifact of the primary dataset
|
|
261
|
+
alone (REQ-AIDS-055).
|
|
262
|
+
15. **Test conclusion stability** — before stating a conclusion as robust,
|
|
263
|
+
define an `ai_data_scientist.sensitivity.SensitivityPlan(target_claim=
|
|
264
|
+
"...", parameter_grid={...}, max_runs=...)` covering the alternative
|
|
265
|
+
specifications that matter (model choice, subset, parameters), and
|
|
266
|
+
call `run_sensitivity(plan, analysis_fn, stability_tolerance=...)`.
|
|
267
|
+
Report the target claim together with `report.stable` and
|
|
268
|
+
`report.max_relative_deviation` to the user;
|
|
269
|
+
`SensitivityBudgetExceededError` means the grid must be narrowed rather
|
|
270
|
+
than silently truncated (REQ-AIDS-056).
|
|
271
|
+
16. **Compare against an independent dataset** — when the user supplies or
|
|
272
|
+
names a second, already-loaded dataset to validate findings against,
|
|
273
|
+
call `ai_data_scientist.dataset_validation.compare_datasets(primary,
|
|
274
|
+
candidate, key_mapping, value_mapping, candidate_relationship=...)` to
|
|
275
|
+
report key overlap and per-column agreement. Automated dataset
|
|
276
|
+
*discovery* (e.g. searching an external catalog such as Kaggle) is out
|
|
277
|
+
of scope for this module — the caller must load the candidate dataset
|
|
278
|
+
first (REQ-AIDS-057).
|
|
279
|
+
17. **Analyze spectral/signal peaks** — when the user has 1-D
|
|
280
|
+
spectrum-like data (x/y pairs such as wavelength/intensity or
|
|
281
|
+
time/amplitude), first remove the background with
|
|
282
|
+
`ai_data_scientist.signal_analysis.baseline_correct(x, y,
|
|
283
|
+
method="linear"|"asls")` (`"linear"` anchors the two endpoints;
|
|
284
|
+
`"asls"` runs a fixed-parameter Asymmetric Least Squares fit for
|
|
285
|
+
curved backgrounds), then call
|
|
286
|
+
`ai_data_scientist.signal_analysis.find_spectral_peaks(x, y,
|
|
287
|
+
prominence_frac=..., window=...)` to get a list of
|
|
288
|
+
`{"position", "fwhm", "prominence", "height"}` dicts, one per
|
|
289
|
+
detected peak, ordered by ascending `position` (set `window`, an odd
|
|
290
|
+
integer >= 5, to apply Savitzky-Golay smoothing before peak-finding
|
|
291
|
+
on noisy data) (REQ-AIDS-094, REQ-AIDS-095). To check whether a peak
|
|
292
|
+
count/position conclusion is robust to the detection parameters,
|
|
293
|
+
build a plan with
|
|
294
|
+
`ai_data_scientist.signal_analysis.build_peak_sensitivity_plan(x, y,
|
|
295
|
+
prominence_fracs=[...], windows=[...], target_claim="...")`, which
|
|
296
|
+
returns a ready-to-use `(SensitivityPlan, analysis_fn)` pair — pass
|
|
297
|
+
both straight into `sensitivity.run_sensitivity` exactly as in step 15
|
|
298
|
+
with no extra glue code required (REQ-AIDS-096).
|
|
299
|
+
|
|
300
|
+
## Constraints / 制約
|
|
301
|
+
- Every notebook write goes through
|
|
302
|
+
`project_manager.enqueue_write`/`ensure_notebook`, which serialize
|
|
303
|
+
concurrent writers and guarantee `nbformat`-valid output (REQ-AIDS-011/029).
|
|
304
|
+
- Never accept a project name outside the ADR-0005 slug pattern
|
|
305
|
+
(`^[a-z0-9]+(-[a-z0-9]+)*$`); ask the user to choose a valid slug instead
|
|
306
|
+
of silently transforming their input.
|
|
307
|
+
- Never fabricate an insight without a located, executed evidence cell.
|
|
308
|
+
- All modules are covered by TDD (pytest) per REQ-AIDS-013; run
|
|
309
|
+
`.venv/bin/pytest` before considering any change to this skill complete.
|
|
310
|
+
|
|
311
|
+
## Source layout / 実装
|
|
312
|
+
- `src/ai_data_scientist/language_router.py` — DES-AIDS-002
|
|
313
|
+
- `src/ai_data_scientist/project_manager.py` — DES-AIDS-003
|
|
314
|
+
- `src/ai_data_scientist/mcp_gateway.py` — DES-AIDS-004
|
|
315
|
+
- `src/ai_data_scientist/ingestion.py` — DES-AIDS-005
|
|
316
|
+
- `src/ai_data_scientist/cleaning.py` — DES-AIDS-006
|
|
317
|
+
- `src/ai_data_scientist/eda.py` — DES-AIDS-007
|
|
318
|
+
- `src/ai_data_scientist/stats_analysis.py` — DES-AIDS-008
|
|
319
|
+
- `src/ai_data_scientist/visualization.py` — DES-AIDS-009
|
|
320
|
+
- `src/ai_data_scientist/insight_engine.py` — DES-AIDS-010
|
|
321
|
+
- `src/ai_data_scientist/notebook_audit.py` — DES-AIDS-033, DES-AIDS-041
|
|
322
|
+
- `src/ai_data_scientist/gate_config.py` — DES-AIDS-011
|
|
323
|
+
- `src/ai_data_scientist/dependency_pins.py` — DES-AIDS-038
|
|
324
|
+
- `src/ai_data_scientist/lifecycle.py` — DES-AIDS-039
|
|
325
|
+
- `src/ai_data_scientist/data_definition.py` — DES-AIDS-040
|
|
326
|
+
- `src/ai_data_scientist/analysis_assumptions.py` — DES-AIDS-042
|
|
327
|
+
- `src/ai_data_scientist/data_quality.py` — DES-AIDS-043
|
|
328
|
+
- `src/ai_data_scientist/sensitivity.py` — DES-AIDS-044
|
|
329
|
+
- `src/ai_data_scientist/dataset_validation.py` — DES-AIDS-045
|
|
330
|
+
- `src/ai_data_scientist/signal_analysis.py` — DES-AIDS-094
|