quantdb-sdk 0.2.5__tar.gz → 0.2.7__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/CHANGELOG.md +24 -0
- {quantdb_sdk-0.2.5/quantdb_sdk.egg-info → quantdb_sdk-0.2.7}/PKG-INFO +10 -5
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/README.md +59 -54
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/pyproject.toml +1 -1
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk/__init__.py +1 -1
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk/_utils.py +25 -1
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk/async_client.py +74 -21
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk/client.py +117 -71
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7/quantdb_sdk.egg-info}/PKG-INFO +10 -5
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk.egg-info/SOURCES.txt +2 -1
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/tests/test_async_client.py +17 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/tests/test_client.py +29 -1
- quantdb_sdk-0.2.7/tests/test_update_pipeline.py +42 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/LICENSE +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/MANIFEST.in +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk/__main__.py +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk/errors.py +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk/py.typed +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk.egg-info/dependency_links.txt +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk.egg-info/entry_points.txt +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk.egg-info/requires.txt +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/quantdb_sdk.egg-info/top_level.txt +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/setup.cfg +0 -0
- {quantdb_sdk-0.2.5 → quantdb_sdk-0.2.7}/tests/test_technical_indicators.py +0 -0
|
@@ -3,6 +3,30 @@
|
|
|
3
3
|
所有 notable 变更都会记录在此文件。格式基于 [Keep a Changelog](https://keepachangelog.com/zh-CN/1.1.0/),
|
|
4
4
|
版本号遵循 [Semantic Versioning](https://semver.org/lang/zh-CN/)。
|
|
5
5
|
|
|
6
|
+
## [0.2.7] - 2026-07-29
|
|
7
|
+
|
|
8
|
+
### Fixed
|
|
9
|
+
- **V1 回退对纯 V2 数据集失效**:`query_kline` / `a_query_kline` 无日期范围时不再默认走 V1。`margin_trading`、`features_daily`、`l1_factors`、`l2_factors` 等 11 个纯 V2 数据集在 COS 上已无 V1 逐股票文件,V1 回退会 404。现在无范围时也走 V2 manifest 路径(manifest 列出的所有分区全量下载),仅在有日期范围时才做日历完整性校验。显式 `layout="v1"` 仍可直接走逐股票文件。
|
|
10
|
+
- **异步客户端时区处理不一致**:`AsyncQuantDBClient._normalise_kline` 缺少 `tz_convert` 逻辑,与同步客户端行为不一致(相同输入可能产生不同的 `trade_date` 字符串),已对齐。
|
|
11
|
+
|
|
12
|
+
## [0.2.6] - 2026-07-27
|
|
13
|
+
|
|
14
|
+
### Fixed
|
|
15
|
+
- **适配服务端 CDN 直连下载(302 跳转)**:异步客户端(httpx 默认不跟随重定向)此前在服务端开启 302 直连后所有下载/同步接口失效,现已修复。
|
|
16
|
+
- **凭证保护**:同步客户端此前跟随 302 时会把 `X-API-Key` 透传给 CDN 域;现两个客户端统一改为「手动处理 302 + 无鉴权头裸请求直连 CDN」,凭证只发给 QuantDB 网关。
|
|
17
|
+
- 302 响应中的 COS ETag 回填至 CDN 响应,保证进程内 ETag 缓存与 If-None-Match 304 逻辑不受影响。
|
|
18
|
+
|
|
19
|
+
### Security (后端)
|
|
20
|
+
- **一次性下载令牌**:302 直连模式改为两跳(`/download` → `/cdn-bridge?token=xxx` → CDN),令牌一次性消费,防止签名 URL 被重复使用或分享给他人免费下载。
|
|
21
|
+
- **移除 `?direct=1` 请求级覆盖**:用户无法再通过请求参数强制走 302 模式,仅全局开关 `QUANTDB_DOWNLOAD_REDIRECT` 可控制。
|
|
22
|
+
- 令牌有效期 90s,远小于 CDN 签名有效期 300s,窗口极小。
|
|
23
|
+
|
|
24
|
+
### Data Architecture (COS)
|
|
25
|
+
- **V2 数据集全面清理 V1 残留**:COS 上 11 个 V2 数据集已删除全部 V1 逐股票文件(~1.8 万对象),实现纯 V2 存储。ListObjects 分页效率提升 3-5 倍。
|
|
26
|
+
- **L1 日频因子 V2 全量上线**:`l1_factors` 2565 个交易日分区(2016-01-04 ~ 至今),支持按日范围高效查询。
|
|
27
|
+
- **L2 高频因子 V2 上线**:`l2_factors` 34 个交易日分区(2026-01-05 ~ 2026-02-27)。
|
|
28
|
+
- **融资融券纯 V2**:`margin_trading` 已从 V1 迁移至纯 V2 按日分区(2563 个分区)。
|
|
29
|
+
|
|
6
30
|
## [0.2.5] - 2026-07-26
|
|
7
31
|
|
|
8
32
|
### Security
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: quantdb-sdk
|
|
3
|
-
Version: 0.2.
|
|
3
|
+
Version: 0.2.7
|
|
4
4
|
Summary: QuantDB 量化数据平台官方 Python SDK
|
|
5
5
|
Author: QuantDB Team
|
|
6
6
|
License: MIT
|
|
@@ -90,12 +90,17 @@ client = QuantDBClient(username="admin", password="admin123")
|
|
|
90
90
|
|
|
91
91
|
## V1 / V2 数据布局
|
|
92
92
|
|
|
93
|
-
|
|
94
|
-
`layout="auto" | "v1" | "v2"
|
|
95
|
-
|
|
93
|
+
QuantDB 数据采用两种物理布局:V1(按股票的全历史文件 `{Symbol}.parquet`)和 V2(按交易日的全市场分区 `dt=YYYYMMDD/data.parquet`)。
|
|
94
|
+
所有下载相关接口均可传入 `layout="auto" | "v1" | "v2"`。
|
|
95
|
+
|
|
96
|
+
**V2 数据集**(COS 纯 V2,零 V1 残留):daily_unadjusted / daily_forward / daily_backward / index_daily / valuation / technical_indicators / market_sentiment / features_daily / l1_factors / l2_factors / margin_trading
|
|
97
|
+
|
|
98
|
+
**V1 数据集**(纯 V1,无 V2 分区):min1_kline / min5_kline / tick_data / 财务七表 / 基础板块
|
|
99
|
+
|
|
100
|
+
默认 `auto` 始终优先 V2:有日期范围时聚合 V2 多日分区,若覆盖不完整则回退 V1(仅对仍保留 V1 文件的数据集有效);无日期范围时走 V2 全量(manifest 列出的所有分区),跳过完整性校验。
|
|
96
101
|
|
|
97
102
|
```python
|
|
98
|
-
#
|
|
103
|
+
# 始终优先 V2 按日分区;有范围且覆盖不完整时自动回退 V1
|
|
99
104
|
df = client.query_kline("600519.SH", start_date="2026-07-01", end_date="2026-07-24")
|
|
100
105
|
|
|
101
106
|
# 强制指定物理布局;layout="v2" 缺日时会明确报错
|
|
@@ -43,65 +43,70 @@ print(df.head())
|
|
|
43
43
|
client = QuantDBClient(username="admin", password="admin123")
|
|
44
44
|
```
|
|
45
45
|
|
|
46
|
-
## 核心功能
|
|
46
|
+
## 核心功能
|
|
47
47
|
|
|
48
48
|
- **数据查询**:K 线、Tick 通过下载 Parquet 切片后客户端解析(消耗流量);股票列表、交易日历、元数据走网关 JSON(不计流量)。
|
|
49
49
|
- **数据下载**:Parquet 文件下载或直读 DataFrame,计入订阅流量。
|
|
50
50
|
- **本地分析**:基于 DuckDB 对本地 Parquet 执行 SQL。
|
|
51
51
|
- **账户管理**:查询用户信息、用量、API Key、订阅与订单。
|
|
52
|
-
- **异步客户端**:基于 httpx,适用于 asyncio 量化框架。
|
|
53
|
-
|
|
54
|
-
## V1 / V2 数据布局
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
`layout="auto" | "v1" | "v2"
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
73
|
-
#
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
52
|
+
- **异步客户端**:基于 httpx,适用于 asyncio 量化框架。
|
|
53
|
+
|
|
54
|
+
## V1 / V2 数据布局
|
|
55
|
+
|
|
56
|
+
QuantDB 数据采用两种物理布局:V1(按股票的全历史文件 `{Symbol}.parquet`)和 V2(按交易日的全市场分区 `dt=YYYYMMDD/data.parquet`)。
|
|
57
|
+
所有下载相关接口均可传入 `layout="auto" | "v1" | "v2"`。
|
|
58
|
+
|
|
59
|
+
**V2 数据集**(COS 纯 V2,零 V1 残留):daily_unadjusted / daily_forward / daily_backward / index_daily / valuation / technical_indicators / market_sentiment / features_daily / l1_factors / l2_factors / margin_trading
|
|
60
|
+
|
|
61
|
+
**V1 数据集**(纯 V1,无 V2 分区):min1_kline / min5_kline / tick_data / 财务七表 / 基础板块
|
|
62
|
+
|
|
63
|
+
默认 `auto` 始终优先 V2:有日期范围时聚合 V2 多日分区,若覆盖不完整则回退 V1(仅对仍保留 V1 文件的数据集有效);无日期范围时走 V2 全量(manifest 列出的所有分区),跳过完整性校验。
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
# 始终优先 V2 按日分区;有范围且覆盖不完整时自动回退 V1
|
|
67
|
+
df = client.query_kline("600519.SH", start_date="2026-07-01", end_date="2026-07-24")
|
|
68
|
+
|
|
69
|
+
# 强制指定物理布局;layout="v2" 缺日时会明确报错
|
|
70
|
+
latest = client.download_file("1", "daily_forward", trade_date="2026-07-24", layout="v2")
|
|
71
|
+
history = client.download_file("1", "daily_forward", symbol="600519.SH", layout="v1")
|
|
72
|
+
|
|
73
|
+
# 以发布清单为 cursor 做原子化增量同步(含 V2 patch)
|
|
74
|
+
result = client.sync_dataset("daily_forward", save_dir="D:/quantdb-data")
|
|
75
|
+
|
|
76
|
+
# 财务、ETF/可转债等尚无 V2 release 的数据集自动按 V1 Manifest 增量同步,
|
|
77
|
+
# 使用 ETag + size 校验;也可从命令行执行:
|
|
78
|
+
# quantdb sync qdb_xxx balance --save-dir D:/quantdb-data
|
|
79
|
+
financial = client.sync_dataset("balance", save_dir="D:/quantdb-data")
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
### 加速批量下载:建议从 8 个工作线程开始
|
|
83
|
+
|
|
84
|
+
`sync_dataset()` 单次调用会按发布顺序串行下载,以保证本地 SQLite 同步状态和 release cursor 一致。
|
|
85
|
+
当需要下载多个独立标的或文件时,可由调用方并行调度;建议先使用 **8 个工作线程**,再结合网络带宽、磁盘写入能力和账户流量配额调整。
|
|
86
|
+
|
|
87
|
+
```python
|
|
88
|
+
from concurrent.futures import ThreadPoolExecutor
|
|
89
|
+
from quantdb_sdk import QuantDBClient
|
|
90
|
+
|
|
91
|
+
API_KEY = "qdb_xxx..."
|
|
92
|
+
symbols = ["600519.SH", "000001.SZ", "600036.SH"]
|
|
93
|
+
|
|
94
|
+
def download_one(symbol: str) -> str:
|
|
95
|
+
# 每个 worker 使用独立客户端,避免跨线程共享 HTTP Session。
|
|
96
|
+
with_client = QuantDBClient(api_key=API_KEY)
|
|
97
|
+
return with_client.download_file(
|
|
98
|
+
"1", "daily_forward", symbol=symbol, save_dir="D:/quantdb-data"
|
|
99
|
+
)
|
|
100
|
+
|
|
101
|
+
with ThreadPoolExecutor(max_workers=8) as pool:
|
|
102
|
+
files = list(pool.map(download_one, symbols))
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
不要对相同 `save_dir` 并行调用多个 `sync_dataset()`:它们会共同写入 `quantdb_sync.sqlite`,可能产生锁竞争。批量同步本身仍建议一次一个数据集执行。
|
|
106
|
+
|
|
107
|
+
## 流量说明
|
|
108
|
+
|
|
109
|
+
免费注册用户获赠 100 MB 一次性体验流量;订阅用户每月含 30 GB 下载流量,超出部分按 ¥1/GB 从账户余额扣减。余额不足时下载会被拦截。
|
|
105
110
|
|
|
106
111
|
## 文档
|
|
107
112
|
|
|
@@ -5,6 +5,7 @@
|
|
|
5
5
|
|
|
6
6
|
import os
|
|
7
7
|
import re
|
|
8
|
+
from datetime import date
|
|
8
9
|
from typing import Dict, Optional
|
|
9
10
|
from urllib.parse import unquote, urlparse
|
|
10
11
|
|
|
@@ -20,7 +21,7 @@ SYNC_DATASET_CATEGORIES: Dict[str, str] = {
|
|
|
20
21
|
"holder_num": "3", "pershare_index": "3", "dividend_factors": "3",
|
|
21
22
|
"etf_pcf": "4", "convertible_bond": "4",
|
|
22
23
|
"valuation": "5", "technical_indicators": "5", "market_sentiment": "5",
|
|
23
|
-
"features_daily": "6", "l1_l2_factors": "6", "l2_factors": "6",
|
|
24
|
+
"features_daily": "6", "l1_factors": "6", "l1_l2_factors": "6", "l2_factors": "6",
|
|
24
25
|
}
|
|
25
26
|
|
|
26
27
|
|
|
@@ -97,6 +98,29 @@ def safe_join(root: str, relative_path: str) -> str:
|
|
|
97
98
|
return target
|
|
98
99
|
|
|
99
100
|
|
|
101
|
+
def normalise_tick_trade_date(trade_date: str) -> str:
|
|
102
|
+
"""规范化 Tick 交易日为 ``YYYY-MM-DD``,拒绝含糊或非法输入。"""
|
|
103
|
+
value = str(trade_date).strip()
|
|
104
|
+
if re.fullmatch(r"\d{8}", value):
|
|
105
|
+
value = f"{value[:4]}-{value[4:6]}-{value[6:]}"
|
|
106
|
+
if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", value):
|
|
107
|
+
raise ValueError("trade_date 必须为 YYYY-MM-DD 或 YYYYMMDD")
|
|
108
|
+
try:
|
|
109
|
+
date.fromisoformat(value)
|
|
110
|
+
except ValueError as exc:
|
|
111
|
+
raise ValueError("trade_date 不是有效交易日日期") from exc
|
|
112
|
+
return value
|
|
113
|
+
|
|
114
|
+
|
|
115
|
+
def tick_shard_relative_path(symbol: str, trade_date: str) -> str:
|
|
116
|
+
"""返回 Tick 平铺 shard 的规范相对路径。"""
|
|
117
|
+
clean_symbol = str(symbol).strip().upper()
|
|
118
|
+
if not re.fullmatch(r"[A-Z0-9]+\.[A-Z0-9]+", clean_symbol):
|
|
119
|
+
raise ValueError("Tick symbol 必须为类似 600519.SH 的代码")
|
|
120
|
+
yyyymmdd = normalise_tick_trade_date(trade_date).replace("-", "")
|
|
121
|
+
return f"1_kline_data/tick_data/{clean_symbol.replace('.', '_')}_{yyyymmdd}.parquet"
|
|
122
|
+
|
|
123
|
+
|
|
100
124
|
def validate_api_host(api_host: str) -> str:
|
|
101
125
|
"""Require HTTPS except for local development/test loopback endpoints."""
|
|
102
126
|
host = api_host.rstrip("/")
|
|
@@ -5,6 +5,7 @@ import io
|
|
|
5
5
|
import os
|
|
6
6
|
import re
|
|
7
7
|
import sqlite3
|
|
8
|
+
from contextlib import asynccontextmanager
|
|
8
9
|
from typing import Any, Dict, List, Optional, Literal
|
|
9
10
|
|
|
10
11
|
import httpx
|
|
@@ -12,7 +13,7 @@ import pandas as pd
|
|
|
12
13
|
|
|
13
14
|
from ._utils import (
|
|
14
15
|
SYNC_DATASET_CATEGORIES, bytes_to_gb, check_download_size, default_download_dir,
|
|
15
|
-
max_download_bytes, parse_filename_from_content_disposition, safe_filename, safe_join,
|
|
16
|
+
max_download_bytes, normalise_tick_trade_date, parse_filename_from_content_disposition, safe_filename, safe_join,
|
|
16
17
|
validate_api_host,
|
|
17
18
|
)
|
|
18
19
|
from .errors import (
|
|
@@ -25,6 +26,9 @@ from .errors import (
|
|
|
25
26
|
ValidationError,
|
|
26
27
|
)
|
|
27
28
|
|
|
29
|
+
# 维护提示:每次发版必须与 pyproject.toml 版本号同步
|
|
30
|
+
_USER_AGENT = "QuantDB-Python-SDK/0.2.7"
|
|
31
|
+
|
|
28
32
|
|
|
29
33
|
class AsyncQuantDBClient:
|
|
30
34
|
"""QuantDB 异步客户端。
|
|
@@ -41,7 +45,7 @@ class AsyncQuantDBClient:
|
|
|
41
45
|
):
|
|
42
46
|
self.api_host = validate_api_host(api_host)
|
|
43
47
|
self.timeout = timeout
|
|
44
|
-
headers = {"User-Agent":
|
|
48
|
+
headers = {"User-Agent": _USER_AGENT}
|
|
45
49
|
if api_key:
|
|
46
50
|
headers["X-API-Key"] = api_key
|
|
47
51
|
elif token:
|
|
@@ -53,17 +57,49 @@ class AsyncQuantDBClient:
|
|
|
53
57
|
headers=headers,
|
|
54
58
|
timeout=httpx.Timeout(timeout, connect=5.0),
|
|
55
59
|
)
|
|
60
|
+
# 裸客户端:跟随 302 直连 CDN 用,不携带任何鉴权头,避免凭证泄露给 CDN 域。
|
|
61
|
+
self._bare_client = httpx.AsyncClient(
|
|
62
|
+
headers={"User-Agent": _USER_AGENT},
|
|
63
|
+
timeout=httpx.Timeout(timeout, connect=5.0),
|
|
64
|
+
)
|
|
56
65
|
# 进程内 parquet ETag 缓存:同对象(ETag 未变)不重复下载,避免重复计费。
|
|
57
66
|
self._cache: Dict[str, Dict[str, Any]] = {}
|
|
58
67
|
|
|
59
68
|
async def close(self) -> None:
|
|
60
69
|
"""关闭底层 httpx 客户端。"""
|
|
61
70
|
await self.client.aclose()
|
|
71
|
+
await self._bare_client.aclose()
|
|
62
72
|
|
|
63
73
|
def clear_cache(self) -> None:
|
|
64
74
|
"""清空进程内 Parquet 缓存(强制下次重新下载最新数据)。"""
|
|
65
75
|
self._cache.clear()
|
|
66
76
|
|
|
77
|
+
@asynccontextmanager
|
|
78
|
+
async def _download_stream(self, params: Dict[str, Any], headers: Optional[Dict[str, str]] = None):
|
|
79
|
+
"""下载专用流式请求:手动处理服务端 302 CDN 直连跳转。
|
|
80
|
+
|
|
81
|
+
服务端开启 CDN 直连后,/api/v1/data/download 预扣流量成功时返回
|
|
82
|
+
302 -> CDN 签名 URL。httpx 默认不跟随重定向;这里用不带鉴权头的
|
|
83
|
+
裸客户端直连 Location,鉴权头不会泄露给 CDN 域。
|
|
84
|
+
"""
|
|
85
|
+
async with self.client.stream(
|
|
86
|
+
"GET", f"{self.api_host}/api/v1/data/download", params=params, headers=headers
|
|
87
|
+
) as resp:
|
|
88
|
+
if resp.status_code not in (301, 302, 303, 307, 308):
|
|
89
|
+
yield resp
|
|
90
|
+
return
|
|
91
|
+
location = resp.headers.get("Location", "")
|
|
92
|
+
etag = resp.headers.get("ETag", "")
|
|
93
|
+
if not location:
|
|
94
|
+
raise ServerError("下载重定向缺少 Location 头")
|
|
95
|
+
async with self._bare_client.stream("GET", location) as cdn_resp:
|
|
96
|
+
if cdn_resp.status_code != 200:
|
|
97
|
+
raise ServerError(f"CDN 直连下载失败:HTTP {cdn_resp.status_code}")
|
|
98
|
+
# 网关 302 响应携带 COS ETag;若 CDN 响应缺失则回填,保证缓存逻辑一致。
|
|
99
|
+
if etag and not cdn_resp.headers.get("ETag"):
|
|
100
|
+
cdn_resp.headers["ETag"] = etag
|
|
101
|
+
yield cdn_resp
|
|
102
|
+
|
|
67
103
|
@staticmethod
|
|
68
104
|
def _validate_layout(layout: str) -> Literal["auto", "v1", "v2"]:
|
|
69
105
|
if layout not in {"auto", "v1", "v2"}:
|
|
@@ -78,6 +114,8 @@ class AsyncQuantDBClient:
|
|
|
78
114
|
def _normalise_kline(df: pd.DataFrame, start_date: Optional[str], end_date: Optional[str], fields: str, limit: Optional[int]) -> pd.DataFrame:
|
|
79
115
|
if "time" in df.columns:
|
|
80
116
|
dt = pd.to_datetime(df["time"], errors="coerce")
|
|
117
|
+
if getattr(dt.dt, "tz", None) is not None:
|
|
118
|
+
dt = dt.dt.tz_convert(None)
|
|
81
119
|
elif "trade_date" in df.columns:
|
|
82
120
|
dt = pd.to_datetime(df["trade_date"], errors="coerce")
|
|
83
121
|
else:
|
|
@@ -242,22 +280,34 @@ class AsyncQuantDBClient:
|
|
|
242
280
|
limit: Optional[int] = None,
|
|
243
281
|
layout: Literal["auto", "v1", "v2"] = "auto",
|
|
244
282
|
) -> pd.DataFrame:
|
|
245
|
-
"""查询 K 线数据(下载 COS parquet 切片后客户端解析,消耗下载流量,异步)。
|
|
283
|
+
"""查询 K 线数据(下载 COS parquet 切片后客户端解析,消耗下载流量,异步)。
|
|
284
|
+
|
|
285
|
+
``auto`` 始终优先 V2 全市场日分区;仅当提供了日期范围且 V2 覆盖不完整时,
|
|
286
|
+
整次回退 V1 股票历史文件。未提供日期范围时走 V2 全量(manifest 列出的
|
|
287
|
+
所有分区),跳过日历完整性校验。显式 ``v1`` 直接走逐股票文件;显式
|
|
288
|
+
``v2`` 不会静默回退。
|
|
289
|
+
"""
|
|
246
290
|
layout = self._validate_layout(layout)
|
|
247
291
|
sub_category = f"daily_{adj_type}"
|
|
248
|
-
|
|
292
|
+
has_range = bool(start_date or end_date)
|
|
293
|
+
if layout == "v1":
|
|
249
294
|
return self._normalise_kline(await self.a_load_as_df("1", sub_category, symbol, layout="v1"), start_date, end_date, fields, limit)
|
|
250
295
|
files = await self.a_query_manifest("1", sub_category, layout="v2")
|
|
251
296
|
selected = [f for f in files if (not start_date or f.get("trade_date", "") >= start_date) and (not end_date or f.get("trade_date", "") <= end_date)]
|
|
252
|
-
|
|
253
|
-
|
|
254
|
-
if
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
if
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
297
|
+
# 仅在有日期范围时校验完整性;无范围时 manifest 列出什么就下什么。
|
|
298
|
+
complete = bool(selected)
|
|
299
|
+
if has_range:
|
|
300
|
+
calendar = await self.a_query_calendar(start_date, end_date)
|
|
301
|
+
expected = set()
|
|
302
|
+
if not calendar.empty:
|
|
303
|
+
date_col = next((c for c in ("trade_date", "date", "cal_date") if c in calendar.columns), None)
|
|
304
|
+
open_col = next((c for c in ("is_open", "is_trading_day", "open") if c in calendar.columns), None)
|
|
305
|
+
if date_col:
|
|
306
|
+
rows = calendar if not open_col else calendar[calendar[open_col].astype(str).isin(["1", "True", "true"])]
|
|
307
|
+
expected = set(pd.to_datetime(rows[date_col], errors="coerce").dropna().dt.strftime("%Y-%m-%d"))
|
|
308
|
+
found = {f.get("trade_date") for f in selected}
|
|
309
|
+
complete = bool(selected) and (not expected or expected.issubset(found))
|
|
310
|
+
if not complete:
|
|
261
311
|
if layout == "auto":
|
|
262
312
|
return self._normalise_kline(await self.a_load_as_df("1", sub_category, symbol, layout="v1"), start_date, end_date, fields, limit)
|
|
263
313
|
raise NotFoundError("V2 日切片在请求日期范围内覆盖不完整;请改用 layout='auto' 或 'v1'")
|
|
@@ -277,6 +327,7 @@ class AsyncQuantDBClient:
|
|
|
277
327
|
layout: Literal["auto", "v1", "v2"] = "auto",
|
|
278
328
|
) -> pd.DataFrame:
|
|
279
329
|
"""查询 Tick 分笔数据(下载 COS parquet 切片后客户端解析,消耗下载流量,异步)。"""
|
|
330
|
+
trade_date = normalise_tick_trade_date(trade_date)
|
|
280
331
|
df = await self.a_load_as_df("1", "tick_data", symbol, trade_date=trade_date, layout=layout)
|
|
281
332
|
ts_col = "ts" if "ts" in df.columns else ("time" if "time" in df.columns else None)
|
|
282
333
|
if ts_col and (start_ts or end_ts):
|
|
@@ -398,9 +449,7 @@ class AsyncQuantDBClient:
|
|
|
398
449
|
if object_key:
|
|
399
450
|
params["object_key"] = object_key
|
|
400
451
|
|
|
401
|
-
async with self.
|
|
402
|
-
"GET", f"{self.api_host}/api/v1/data/download", params=params
|
|
403
|
-
) as resp:
|
|
452
|
+
async with self._download_stream(params) as resp:
|
|
404
453
|
if resp.status_code != 200:
|
|
405
454
|
body = await resp.aread()
|
|
406
455
|
# 构造一个完整的 httpx.Response 用于错误解析,然后立即抛出
|
|
@@ -467,9 +516,7 @@ class AsyncQuantDBClient:
|
|
|
467
516
|
req_headers = {}
|
|
468
517
|
if cached and cached.get("etag"):
|
|
469
518
|
req_headers["If-None-Match"] = cached["etag"]
|
|
470
|
-
async with self.
|
|
471
|
-
"GET", f"{self.api_host}/api/v1/data/download", params=params, headers=req_headers
|
|
472
|
-
) as resp:
|
|
519
|
+
async with self._download_stream(params, headers=req_headers) as resp:
|
|
473
520
|
if resp.status_code == 304 and cached:
|
|
474
521
|
return cached["df"]
|
|
475
522
|
if resp.status_code != 200:
|
|
@@ -567,7 +614,7 @@ class AsyncQuantDBClient:
|
|
|
567
614
|
if old and old[0] == obj.get("etag") and old[1] == obj.get("sha256") and os.path.exists(old[3]) and (expected_size is None or os.path.getsize(old[3]) == expected_size): continue
|
|
568
615
|
os.makedirs(os.path.dirname(target), exist_ok=True)
|
|
569
616
|
tmp, digest = target + ".part", hashlib.sha256()
|
|
570
|
-
async with self.
|
|
617
|
+
async with self._download_stream({"category_id": SYNC_DATASET_CATEGORIES[dataset], "sub_category": dataset, "layout": "v2", "object_key": key}) as resp:
|
|
571
618
|
if resp.status_code != 200:
|
|
572
619
|
body = await resp.aread(); self._check_response(httpx.Response(resp.status_code, content=body)); raise QuantDBError("下载失败")
|
|
573
620
|
try:
|
|
@@ -592,10 +639,16 @@ class AsyncQuantDBClient:
|
|
|
592
639
|
expected_size = obj.get("size")
|
|
593
640
|
old = state.execute("SELECT etag,size,path FROM objects WHERE key=?", (key,)).fetchone()
|
|
594
641
|
if old and old[0] == obj.get("etag") and os.path.exists(old[2]) and (expected_size is None or os.path.getsize(old[2]) == expected_size): continue
|
|
642
|
+
download_params = {"category_id": SYNC_DATASET_CATEGORIES[dataset], "sub_category": dataset, "layout": "v1", "symbol": obj.get("symbol", "")}
|
|
643
|
+
if dataset == "tick_data":
|
|
644
|
+
trade_date = obj.get("trade_date")
|
|
645
|
+
if not trade_date:
|
|
646
|
+
raise ServerError("Tick manifest 缺少 trade_date")
|
|
647
|
+
download_params["trade_date"] = normalise_tick_trade_date(trade_date)
|
|
595
648
|
os.makedirs(os.path.dirname(target), exist_ok=True)
|
|
596
649
|
tmp, written_size = target + ".part", 0
|
|
597
650
|
try:
|
|
598
|
-
async with self.
|
|
651
|
+
async with self._download_stream(download_params) as resp:
|
|
599
652
|
if resp.status_code != 200:
|
|
600
653
|
body = await resp.aread(); self._check_response(httpx.Response(resp.status_code, content=body)); raise QuantDBError("下载失败")
|
|
601
654
|
with open(tmp, "wb") as fh:
|
|
@@ -13,11 +13,11 @@ import requests
|
|
|
13
13
|
from requests.adapters import HTTPAdapter
|
|
14
14
|
from urllib3.util.retry import Retry
|
|
15
15
|
|
|
16
|
-
from ._utils import (
|
|
17
|
-
SYNC_DATASET_CATEGORIES, bytes_to_gb, check_download_size, default_download_dir,
|
|
18
|
-
max_download_bytes, parse_filename_from_content_disposition, safe_filename, safe_join,
|
|
19
|
-
validate_api_host,
|
|
20
|
-
)
|
|
16
|
+
from ._utils import (
|
|
17
|
+
SYNC_DATASET_CATEGORIES, bytes_to_gb, check_download_size, default_download_dir,
|
|
18
|
+
max_download_bytes, normalise_tick_trade_date, parse_filename_from_content_disposition, safe_filename, safe_join,
|
|
19
|
+
validate_api_host,
|
|
20
|
+
)
|
|
21
21
|
from .errors import (
|
|
22
22
|
AuthError,
|
|
23
23
|
InsufficientTrafficError,
|
|
@@ -28,6 +28,9 @@ from .errors import (
|
|
|
28
28
|
ValidationError,
|
|
29
29
|
)
|
|
30
30
|
|
|
31
|
+
# 维护提示:每次发版必须与 pyproject.toml 版本号同步
|
|
32
|
+
_USER_AGENT = "QuantDB-Python-SDK/0.2.7"
|
|
33
|
+
|
|
31
34
|
|
|
32
35
|
class QuantDBClient:
|
|
33
36
|
"""QuantDB 同步客户端。
|
|
@@ -45,14 +48,12 @@ class QuantDBClient:
|
|
|
45
48
|
timeout: tuple = (5, 60),
|
|
46
49
|
max_retries: int = 2,
|
|
47
50
|
):
|
|
48
|
-
self.api_host = validate_api_host(api_host)
|
|
51
|
+
self.api_host = validate_api_host(api_host)
|
|
49
52
|
self.timeout = timeout
|
|
50
53
|
self.token: Optional[str] = None
|
|
51
54
|
self.headers: Dict[str, str] = {}
|
|
52
55
|
self.session = requests.Session()
|
|
53
|
-
|
|
54
|
-
# 维护提示:每次版本号变化必须同步改这里(init 里的 __version__ 走 metadata 自动同步)
|
|
55
|
-
self.session.headers.update({"User-Agent": "QuantDB-Python-SDK/0.2.5"})
|
|
56
|
+
self.session.headers.update({"User-Agent": _USER_AGENT})
|
|
56
57
|
|
|
57
58
|
if api_key:
|
|
58
59
|
self.headers = {"X-API-Key": api_key}
|
|
@@ -98,6 +99,7 @@ class QuantDBClient:
|
|
|
98
99
|
json: Optional[dict] = None,
|
|
99
100
|
stream: bool = False,
|
|
100
101
|
headers: Optional[dict] = None,
|
|
102
|
+
allow_redirects: bool = True,
|
|
101
103
|
) -> requests.Response:
|
|
102
104
|
url = f"{self.api_host}{path}"
|
|
103
105
|
resp = self.session.request(
|
|
@@ -108,6 +110,7 @@ class QuantDBClient:
|
|
|
108
110
|
stream=stream,
|
|
109
111
|
timeout=self.timeout,
|
|
110
112
|
headers=headers,
|
|
113
|
+
allow_redirects=allow_redirects,
|
|
111
114
|
)
|
|
112
115
|
return resp
|
|
113
116
|
|
|
@@ -115,6 +118,38 @@ class QuantDBClient:
|
|
|
115
118
|
resp = self._request("GET", path, params=params)
|
|
116
119
|
return self._check_response(resp)
|
|
117
120
|
|
|
121
|
+
def _download_stream(self, params: dict, headers: Optional[dict] = None) -> requests.Response:
|
|
122
|
+
"""下载专用流式请求:手动处理服务端 302 CDN 直连跳转。
|
|
123
|
+
|
|
124
|
+
服务端开启 CDN 直连后,/api/v1/data/download 预扣流量成功时返回
|
|
125
|
+
302 -> CDN 签名 URL。requests 自动跟随会把自定义 X-API-Key 头
|
|
126
|
+
透传给 CDN 域,故禁用自动跟随,改用不带鉴权头的裸请求直连。
|
|
127
|
+
"""
|
|
128
|
+
resp = self._request(
|
|
129
|
+
"GET", "/api/v1/data/download", params=params, stream=True,
|
|
130
|
+
headers=headers, allow_redirects=False,
|
|
131
|
+
)
|
|
132
|
+
if resp.status_code not in (301, 302, 303, 307, 308):
|
|
133
|
+
return resp
|
|
134
|
+
location = resp.headers.get("Location", "")
|
|
135
|
+
etag = resp.headers.get("ETag", "")
|
|
136
|
+
resp.close()
|
|
137
|
+
if not location:
|
|
138
|
+
raise ServerError("下载重定向缺少 Location 头")
|
|
139
|
+
cdn_resp = requests.get(
|
|
140
|
+
location,
|
|
141
|
+
stream=True,
|
|
142
|
+
timeout=self.timeout,
|
|
143
|
+
headers={"User-Agent": _USER_AGENT},
|
|
144
|
+
)
|
|
145
|
+
if cdn_resp.status_code != 200:
|
|
146
|
+
cdn_resp.close()
|
|
147
|
+
raise ServerError(f"CDN 直连下载失败:HTTP {cdn_resp.status_code}")
|
|
148
|
+
# 网关 302 响应携带 COS ETag;若 CDN 响应缺失则回填,保证客户端缓存逻辑一致。
|
|
149
|
+
if etag and not cdn_resp.headers.get("ETag"):
|
|
150
|
+
cdn_resp.headers["ETag"] = etag
|
|
151
|
+
return cdn_resp
|
|
152
|
+
|
|
118
153
|
def _post(self, path: str, json: Optional[dict] = None) -> dict:
|
|
119
154
|
resp = self._request("POST", path, json=json)
|
|
120
155
|
return self._check_response(resp)
|
|
@@ -302,29 +337,33 @@ class QuantDBClient:
|
|
|
302
337
|
) -> pd.DataFrame:
|
|
303
338
|
"""查询 K 线数据(下载 COS parquet 切片后客户端解析,消耗下载流量)。
|
|
304
339
|
|
|
305
|
-
``auto``
|
|
306
|
-
整次回退 V1
|
|
307
|
-
|
|
340
|
+
``auto`` 始终优先 V2 全市场日分区;仅当提供了日期范围且 V2 覆盖不完整时,
|
|
341
|
+
整次回退 V1 股票历史文件。未提供日期范围时走 V2 全量(manifest 列出的
|
|
342
|
+
所有分区),跳过日历完整性校验。显式 ``v1`` 直接走逐股票文件;显式
|
|
343
|
+
``v2`` 不会静默回退。
|
|
308
344
|
"""
|
|
309
345
|
layout = self._validate_layout(layout)
|
|
310
346
|
sub_category = f"daily_{adj_type}"
|
|
311
347
|
has_range = bool(start_date or end_date)
|
|
312
|
-
if layout == "v1"
|
|
348
|
+
if layout == "v1":
|
|
313
349
|
return self._normalise_kline(self.load_as_df("1", sub_category, symbol, layout="v1"), start_date, end_date, fields, limit)
|
|
314
350
|
|
|
315
351
|
files = self.query_manifest("1", sub_category, layout="v2")
|
|
316
352
|
selected = [f for f in files if (not start_date or f.get("trade_date", "") >= start_date) and (not end_date or f.get("trade_date", "") <= end_date)]
|
|
317
|
-
# 日历是 V2
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
if
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
if
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
353
|
+
# 日历是 V2 完整性的权威。仅在有日期范围时校验完整性;无范围时 manifest
|
|
354
|
+
# 列出什么就下什么,不要求日历覆盖。auto 在覆盖不完整时回退 V1;v2 明确报错。
|
|
355
|
+
complete = bool(selected)
|
|
356
|
+
if has_range:
|
|
357
|
+
calendar = self.query_calendar(start_date, end_date)
|
|
358
|
+
expected = set()
|
|
359
|
+
if not calendar.empty:
|
|
360
|
+
date_col = next((c for c in ("trade_date", "date", "cal_date") if c in calendar.columns), None)
|
|
361
|
+
open_col = next((c for c in ("is_open", "is_trading_day", "open") if c in calendar.columns), None)
|
|
362
|
+
if date_col:
|
|
363
|
+
rows = calendar if not open_col else calendar[calendar[open_col].astype(str).isin(["1", "True", "true"])]
|
|
364
|
+
expected = set(pd.to_datetime(rows[date_col], errors="coerce").dropna().dt.strftime("%Y-%m-%d"))
|
|
365
|
+
found = {f.get("trade_date") for f in selected}
|
|
366
|
+
complete = bool(selected) and (not expected or expected.issubset(found))
|
|
328
367
|
if not complete:
|
|
329
368
|
if layout == "auto":
|
|
330
369
|
return self._normalise_kline(self.load_as_df("1", sub_category, symbol, layout="v1"), start_date, end_date, fields, limit)
|
|
@@ -350,6 +389,7 @@ class QuantDBClient:
|
|
|
350
389
|
下载 trade_date 当日该 symbol 的 tick parquet,按 start_ts/end_ts 过滤时间、按 fields 选列。
|
|
351
390
|
start_ts/end_ts 可传完整时间戳或 "HH:MM:SS"(自动补 trade_date 日期)。
|
|
352
391
|
"""
|
|
392
|
+
trade_date = normalise_tick_trade_date(trade_date)
|
|
353
393
|
df = self.load_as_df("1", "tick_data", symbol, trade_date=trade_date, layout=layout)
|
|
354
394
|
# 时间过滤
|
|
355
395
|
ts_col = "ts" if "ts" in df.columns else ("time" if "time" in df.columns else None)
|
|
@@ -486,38 +526,36 @@ class QuantDBClient:
|
|
|
486
526
|
if object_key:
|
|
487
527
|
params["object_key"] = object_key
|
|
488
528
|
|
|
489
|
-
resp = self.
|
|
490
|
-
"GET", "/api/v1/data/download", params=params, stream=True
|
|
491
|
-
)
|
|
529
|
+
resp = self._download_stream(params)
|
|
492
530
|
if resp.status_code != 200:
|
|
493
531
|
self._check_response(resp)
|
|
494
532
|
|
|
495
533
|
fallback = f"{sub_category}.parquet"
|
|
496
534
|
if symbol:
|
|
497
535
|
fallback = f"{sub_category}_{symbol}.parquet"
|
|
498
|
-
filename = safe_filename(
|
|
499
|
-
parse_filename_from_content_disposition(resp.headers.get("Content-Disposition", ""), fallback),
|
|
500
|
-
safe_filename(fallback),
|
|
501
|
-
)
|
|
536
|
+
filename = safe_filename(
|
|
537
|
+
parse_filename_from_content_disposition(resp.headers.get("Content-Disposition", ""), fallback),
|
|
538
|
+
safe_filename(fallback),
|
|
539
|
+
)
|
|
502
540
|
|
|
503
541
|
save_path = os.path.join(save_dir, filename)
|
|
504
542
|
tmp_path = save_path + ".part"
|
|
505
|
-
maximum = max_download_bytes()
|
|
506
|
-
check_download_size(resp.headers.get("Content-Length"), maximum)
|
|
507
|
-
written = 0
|
|
508
|
-
try:
|
|
509
|
-
with open(tmp_path, "wb") as f:
|
|
510
|
-
for chunk in resp.iter_content(chunk_size=8192):
|
|
511
|
-
if chunk:
|
|
512
|
-
written += len(chunk)
|
|
513
|
-
if written > maximum:
|
|
514
|
-
raise ServerError(f"下载文件超过大小限制({maximum} 字节)")
|
|
515
|
-
f.write(chunk)
|
|
516
|
-
os.replace(tmp_path, save_path)
|
|
517
|
-
except Exception:
|
|
518
|
-
if os.path.exists(tmp_path):
|
|
519
|
-
os.remove(tmp_path)
|
|
520
|
-
raise
|
|
543
|
+
maximum = max_download_bytes()
|
|
544
|
+
check_download_size(resp.headers.get("Content-Length"), maximum)
|
|
545
|
+
written = 0
|
|
546
|
+
try:
|
|
547
|
+
with open(tmp_path, "wb") as f:
|
|
548
|
+
for chunk in resp.iter_content(chunk_size=8192):
|
|
549
|
+
if chunk:
|
|
550
|
+
written += len(chunk)
|
|
551
|
+
if written > maximum:
|
|
552
|
+
raise ServerError(f"下载文件超过大小限制({maximum} 字节)")
|
|
553
|
+
f.write(chunk)
|
|
554
|
+
os.replace(tmp_path, save_path)
|
|
555
|
+
except Exception:
|
|
556
|
+
if os.path.exists(tmp_path):
|
|
557
|
+
os.remove(tmp_path)
|
|
558
|
+
raise
|
|
521
559
|
return os.path.abspath(save_path)
|
|
522
560
|
|
|
523
561
|
def load_as_df(
|
|
@@ -554,7 +592,7 @@ class QuantDBClient:
|
|
|
554
592
|
headers = {}
|
|
555
593
|
if cached and cached.get("etag"):
|
|
556
594
|
headers["If-None-Match"] = cached["etag"]
|
|
557
|
-
resp = self.
|
|
595
|
+
resp = self._download_stream(params, headers=headers)
|
|
558
596
|
if resp.status_code == 304 and cached:
|
|
559
597
|
return cached["df"]
|
|
560
598
|
if resp.status_code != 200:
|
|
@@ -563,16 +601,16 @@ class QuantDBClient:
|
|
|
563
601
|
# 命中缓存(ETag 未变):复用已解析的 df,不重复消耗解析
|
|
564
602
|
if cached and etag and cached.get("etag") == etag:
|
|
565
603
|
return cached["df"]
|
|
566
|
-
maximum = max_download_bytes()
|
|
567
|
-
check_download_size(resp.headers.get("Content-Length"), maximum)
|
|
568
|
-
payload, written = io.BytesIO(), 0
|
|
569
|
-
for chunk in resp.iter_content(chunk_size=1024 * 1024):
|
|
570
|
-
if chunk:
|
|
571
|
-
written += len(chunk)
|
|
572
|
-
if written > maximum:
|
|
573
|
-
raise ServerError(f"下载文件超过大小限制({maximum} 字节)")
|
|
574
|
-
payload.write(chunk)
|
|
575
|
-
df = pd.read_parquet(payload)
|
|
604
|
+
maximum = max_download_bytes()
|
|
605
|
+
check_download_size(resp.headers.get("Content-Length"), maximum)
|
|
606
|
+
payload, written = io.BytesIO(), 0
|
|
607
|
+
for chunk in resp.iter_content(chunk_size=1024 * 1024):
|
|
608
|
+
if chunk:
|
|
609
|
+
written += len(chunk)
|
|
610
|
+
if written > maximum:
|
|
611
|
+
raise ServerError(f"下载文件超过大小限制({maximum} 字节)")
|
|
612
|
+
payload.write(chunk)
|
|
613
|
+
df = pd.read_parquet(payload)
|
|
576
614
|
if etag:
|
|
577
615
|
self._cache[cache_key] = {"etag": etag, "df": df}
|
|
578
616
|
return df
|
|
@@ -589,12 +627,12 @@ class QuantDBClient:
|
|
|
589
627
|
|
|
590
628
|
若文件不存在会先自动下载。
|
|
591
629
|
|
|
592
|
-
``sql`` 仅接受 WHERE 条件字符串;SDK 固定查询下载的单个 Parquet 文件。
|
|
630
|
+
``sql`` 仅接受 WHERE 条件字符串;SDK 固定查询下载的单个 Parquet 文件。
|
|
593
631
|
|
|
594
632
|
安全限制:
|
|
595
633
|
- 不允许分号(;),防止多语句注入
|
|
596
634
|
- 不允许注释(-- 或 /* */),防止注释注入
|
|
597
|
-
- 不允许子查询、JOIN、外部表函数或任何数据修改语句
|
|
635
|
+
- 不允许子查询、JOIN、外部表函数或任何数据修改语句
|
|
598
636
|
"""
|
|
599
637
|
file_path = self.download_file(
|
|
600
638
|
category_id=category_id,
|
|
@@ -611,19 +649,19 @@ class QuantDBClient:
|
|
|
611
649
|
|
|
612
650
|
# 安全检查:拒绝危险字符和关键字
|
|
613
651
|
dangerous_keywords = [
|
|
614
|
-
";", "--", "/*", "*/", "select", "from", "join", "union", "drop",
|
|
615
|
-
"delete", "insert", "update", "alter", "create", "exec", "execute",
|
|
616
|
-
"attach", "copy", "pragma", "install", "load", "read_", "http", "glob",
|
|
652
|
+
";", "--", "/*", "*/", "select", "from", "join", "union", "drop",
|
|
653
|
+
"delete", "insert", "update", "alter", "create", "exec", "execute",
|
|
654
|
+
"attach", "copy", "pragma", "install", "load", "read_", "http", "glob",
|
|
617
655
|
]
|
|
618
656
|
sql_upper = sql.upper()
|
|
619
657
|
for kw in dangerous_keywords:
|
|
620
658
|
if kw in sql_upper:
|
|
621
659
|
raise ValidationError(
|
|
622
660
|
f"SQL 包含危险关键字 '{kw}',已被拒绝。"
|
|
623
|
-
"query_local 仅支持针对已下载文件的 WHERE 条件。"
|
|
624
|
-
)
|
|
625
|
-
sql = f"SELECT * FROM '{clean_path}' WHERE {sql}"
|
|
626
|
-
return duckdb.query(sql).df()
|
|
661
|
+
"query_local 仅支持针对已下载文件的 WHERE 条件。"
|
|
662
|
+
)
|
|
663
|
+
sql = f"SELECT * FROM '{clean_path}' WHERE {sql}"
|
|
664
|
+
return duckdb.query(sql).df()
|
|
627
665
|
|
|
628
666
|
def get_local_warehouse(self, save_dir: Optional[str] = None) -> "DuckDBWarehouse":
|
|
629
667
|
"""获取 DuckDB 本地数据仓库实例,用于离线多表 SQL JOIN 查询。"""
|
|
@@ -656,12 +694,12 @@ class QuantDBClient:
|
|
|
656
694
|
for obj in release.get("objects", []):
|
|
657
695
|
key = self._normalise_release_key(obj["key"])
|
|
658
696
|
relative_path = obj.get("relative_path") or key
|
|
659
|
-
target = safe_join(root, relative_path)
|
|
697
|
+
target = safe_join(root, relative_path)
|
|
660
698
|
expected_size = obj.get("size")
|
|
661
699
|
old = state.execute("SELECT etag, sha256, size, path FROM objects WHERE key=?", (key,)).fetchone()
|
|
662
700
|
if old and old[0] == obj.get("etag") and old[1] == obj.get("sha256") and os.path.exists(old[3]) and (expected_size is None or os.path.getsize(old[3]) == expected_size):
|
|
663
701
|
continue
|
|
664
|
-
resp = self.
|
|
702
|
+
resp = self._download_stream({"category_id": SYNC_DATASET_CATEGORIES[dataset], "sub_category": dataset, "layout": "v2", "object_key": key})
|
|
665
703
|
if resp.status_code != 200:
|
|
666
704
|
self._check_response(resp)
|
|
667
705
|
os.makedirs(os.path.dirname(target), exist_ok=True)
|
|
@@ -694,12 +732,20 @@ class QuantDBClient:
|
|
|
694
732
|
manifest = self._get("/api/v1/data/download/manifest", {"category_id": SYNC_DATASET_CATEGORIES[dataset], "sub_category": dataset, "layout": "v1"})
|
|
695
733
|
for obj in manifest.get("files", []):
|
|
696
734
|
key, relative_path = obj["key"], obj.get("relative_path") or obj["key"]
|
|
697
|
-
target = safe_join(root, relative_path)
|
|
735
|
+
target = safe_join(root, relative_path)
|
|
698
736
|
expected_size = obj.get("size")
|
|
699
737
|
old = state.execute("SELECT etag, size, path FROM objects WHERE key=?", (key,)).fetchone()
|
|
700
738
|
if old and old[0] == obj.get("etag") and os.path.exists(old[2]) and (expected_size is None or os.path.getsize(old[2]) == expected_size):
|
|
701
739
|
continue
|
|
702
|
-
|
|
740
|
+
download_params = {"category_id": SYNC_DATASET_CATEGORIES[dataset], "sub_category": dataset, "layout": "v1", "symbol": obj.get("symbol", "")}
|
|
741
|
+
# Tick manifest 的每个对象都是独立交易日 shard。不能省略日期,
|
|
742
|
+
# 否则后端会回退下载该股票的最新 shard,造成历史同步错位。
|
|
743
|
+
if dataset == "tick_data":
|
|
744
|
+
trade_date = obj.get("trade_date")
|
|
745
|
+
if not trade_date:
|
|
746
|
+
raise ServerError("Tick manifest 缺少 trade_date")
|
|
747
|
+
download_params["trade_date"] = normalise_tick_trade_date(trade_date)
|
|
748
|
+
resp = self._download_stream(download_params)
|
|
703
749
|
if resp.status_code != 200: self._check_response(resp)
|
|
704
750
|
os.makedirs(os.path.dirname(target), exist_ok=True)
|
|
705
751
|
tmp = target + ".part"
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: quantdb-sdk
|
|
3
|
-
Version: 0.2.
|
|
3
|
+
Version: 0.2.7
|
|
4
4
|
Summary: QuantDB 量化数据平台官方 Python SDK
|
|
5
5
|
Author: QuantDB Team
|
|
6
6
|
License: MIT
|
|
@@ -90,12 +90,17 @@ client = QuantDBClient(username="admin", password="admin123")
|
|
|
90
90
|
|
|
91
91
|
## V1 / V2 数据布局
|
|
92
92
|
|
|
93
|
-
|
|
94
|
-
`layout="auto" | "v1" | "v2"
|
|
95
|
-
|
|
93
|
+
QuantDB 数据采用两种物理布局:V1(按股票的全历史文件 `{Symbol}.parquet`)和 V2(按交易日的全市场分区 `dt=YYYYMMDD/data.parquet`)。
|
|
94
|
+
所有下载相关接口均可传入 `layout="auto" | "v1" | "v2"`。
|
|
95
|
+
|
|
96
|
+
**V2 数据集**(COS 纯 V2,零 V1 残留):daily_unadjusted / daily_forward / daily_backward / index_daily / valuation / technical_indicators / market_sentiment / features_daily / l1_factors / l2_factors / margin_trading
|
|
97
|
+
|
|
98
|
+
**V1 数据集**(纯 V1,无 V2 分区):min1_kline / min5_kline / tick_data / 财务七表 / 基础板块
|
|
99
|
+
|
|
100
|
+
默认 `auto` 始终优先 V2:有日期范围时聚合 V2 多日分区,若覆盖不完整则回退 V1(仅对仍保留 V1 文件的数据集有效);无日期范围时走 V2 全量(manifest 列出的所有分区),跳过完整性校验。
|
|
96
101
|
|
|
97
102
|
```python
|
|
98
|
-
#
|
|
103
|
+
# 始终优先 V2 按日分区;有范围且覆盖不完整时自动回退 V1
|
|
99
104
|
df = client.query_kline("600519.SH", start_date="2026-07-01", end_date="2026-07-24")
|
|
100
105
|
|
|
101
106
|
# 强制指定物理布局;layout="v2" 缺日时会明确报错
|
|
@@ -83,6 +83,23 @@ async def test_a_sync_financial_v1_validates_manifest_size(tmp_path):
|
|
|
83
83
|
assert (tmp_path / "3_financial_data" / "balance" / "600000.SH.parquet").read_bytes() == payload
|
|
84
84
|
|
|
85
85
|
|
|
86
|
+
@pytest.mark.asyncio
|
|
87
|
+
@respx.mock
|
|
88
|
+
async def test_a_sync_tick_uses_manifest_trade_date(tmp_path):
|
|
89
|
+
payload = b"tick-parquet"
|
|
90
|
+
respx.get(f"{API_HOST}/api/v1/data/releases").mock(return_value=httpx.Response(200, json={"releases": []}))
|
|
91
|
+
respx.get(f"{API_HOST}/api/v1/data/download/manifest").mock(return_value=httpx.Response(200, json={"files": [{
|
|
92
|
+
"key": "1_kline_data/tick_data/600519_SH_20260720.parquet",
|
|
93
|
+
"relative_path": "1_kline_data/tick_data/600519_SH_20260720.parquet",
|
|
94
|
+
"symbol": "600519.SH", "trade_date": "2026-07-20", "etag": "tick-etag", "size": len(payload),
|
|
95
|
+
}]}))
|
|
96
|
+
route = respx.get(f"{API_HOST}/api/v1/data/download").mock(return_value=httpx.Response(200, content=payload))
|
|
97
|
+
async with AsyncQuantDBClient(api_host=API_HOST, api_key="test-key") as client:
|
|
98
|
+
result = await client.a_sync_dataset("tick_data", str(tmp_path))
|
|
99
|
+
assert result["downloaded"] == ["1_kline_data/tick_data/600519_SH_20260720.parquet"]
|
|
100
|
+
assert parse_qs(urlparse(str(route.calls[0].request.url)).query)["trade_date"] == ["2026-07-20"]
|
|
101
|
+
|
|
102
|
+
|
|
86
103
|
@pytest.mark.asyncio
|
|
87
104
|
@respx.mock
|
|
88
105
|
async def test_a_auth_error_raises_auth_error():
|
|
@@ -9,7 +9,7 @@ import responses
|
|
|
9
9
|
|
|
10
10
|
from quantdb_sdk import AuthError, InsufficientTrafficError, QuantDBClient
|
|
11
11
|
from quantdb_sdk.errors import ServerError
|
|
12
|
-
from quantdb_sdk._utils import safe_join
|
|
12
|
+
from quantdb_sdk._utils import normalise_tick_trade_date, safe_join, tick_shard_relative_path
|
|
13
13
|
|
|
14
14
|
|
|
15
15
|
API_HOST = "http://localhost:5000"
|
|
@@ -55,6 +55,15 @@ def test_safe_join_rejects_path_traversal(tmp_path):
|
|
|
55
55
|
safe_join(str(tmp_path), "../../outside.parquet")
|
|
56
56
|
|
|
57
57
|
|
|
58
|
+
def test_tick_shard_path_is_flat_and_date_is_normalised():
|
|
59
|
+
assert normalise_tick_trade_date("20260720") == "2026-07-20"
|
|
60
|
+
assert tick_shard_relative_path("600519.SH", "2026-07-20") == (
|
|
61
|
+
"1_kline_data/tick_data/600519_SH_20260720.parquet"
|
|
62
|
+
)
|
|
63
|
+
with pytest.raises(ValueError):
|
|
64
|
+
normalise_tick_trade_date("2026-99-20")
|
|
65
|
+
|
|
66
|
+
|
|
58
67
|
@responses.activate
|
|
59
68
|
def test_download_file_uses_safe_fallback_for_malicious_filename(tmp_path):
|
|
60
69
|
responses.get(
|
|
@@ -130,6 +139,25 @@ def test_sync_financial_v1_validates_manifest_size(tmp_path):
|
|
|
130
139
|
assert (tmp_path / "3_financial_data" / "balance" / "600000.SH.parquet").read_bytes() == payload
|
|
131
140
|
|
|
132
141
|
|
|
142
|
+
@responses.activate
|
|
143
|
+
def test_sync_tick_uses_manifest_trade_date(tmp_path):
|
|
144
|
+
payload = b"tick-parquet"
|
|
145
|
+
responses.get(f"{API_HOST}/api/v1/data/releases", json={"releases": []}, status=200)
|
|
146
|
+
responses.get(
|
|
147
|
+
f"{API_HOST}/api/v1/data/download/manifest",
|
|
148
|
+
json={"files": [{
|
|
149
|
+
"key": "1_kline_data/tick_data/600519_SH_20260720.parquet",
|
|
150
|
+
"relative_path": "1_kline_data/tick_data/600519_SH_20260720.parquet",
|
|
151
|
+
"symbol": "600519.SH", "trade_date": "2026-07-20", "etag": "tick-etag", "size": len(payload),
|
|
152
|
+
}]}, status=200,
|
|
153
|
+
)
|
|
154
|
+
responses.get(f"{API_HOST}/api/v1/data/download", body=payload, status=200)
|
|
155
|
+
result = QuantDBClient(api_host=API_HOST, api_key="test-key").sync_dataset("tick_data", str(tmp_path))
|
|
156
|
+
assert result["downloaded"] == ["1_kline_data/tick_data/600519_SH_20260720.parquet"]
|
|
157
|
+
params = parse_qs(urlparse(responses.calls[-1].request.url).query)
|
|
158
|
+
assert params["trade_date"] == ["2026-07-20"]
|
|
159
|
+
|
|
160
|
+
|
|
133
161
|
@responses.activate
|
|
134
162
|
def test_sync_v1_size_failure_does_not_keep_partial_file(tmp_path):
|
|
135
163
|
responses.get(f"{API_HOST}/api/v1/data/releases", json={"releases": []}, status=200)
|
|
@@ -0,0 +1,42 @@
|
|
|
1
|
+
import importlib.util
|
|
2
|
+
import json
|
|
3
|
+
import sys
|
|
4
|
+
from pathlib import Path
|
|
5
|
+
|
|
6
|
+
|
|
7
|
+
MODULE = Path(__file__).resolve().parents[1] / "pipeline" / "update_pipeline.py"
|
|
8
|
+
SPEC = importlib.util.spec_from_file_location("update_pipeline", MODULE)
|
|
9
|
+
pipeline = importlib.util.module_from_spec(SPEC)
|
|
10
|
+
assert SPEC and SPEC.loader
|
|
11
|
+
sys.modules[SPEC.name] = pipeline
|
|
12
|
+
SPEC.loader.exec_module(pipeline)
|
|
13
|
+
|
|
14
|
+
|
|
15
|
+
def test_manual_tasks_are_not_automatic():
|
|
16
|
+
names = {task.name for task in pipeline.AUTO_TASKS}
|
|
17
|
+
assert "tick" not in names
|
|
18
|
+
assert "l2" not in names
|
|
19
|
+
assert pipeline.MANUAL_TASKS == {"tick", "l2"}
|
|
20
|
+
|
|
21
|
+
|
|
22
|
+
def test_dry_run_writes_machine_readable_manifest(tmp_path, monkeypatch):
|
|
23
|
+
monkeypatch.setattr(pipeline, "DATA_ROOT", tmp_path)
|
|
24
|
+
monkeypatch.setattr(pipeline, "RUNS_DIR", tmp_path / "_meta" / "update_runs")
|
|
25
|
+
args = type("Args", (), {"tasks": "market", "resume": None, "dry_run": True, "no_publish": True})()
|
|
26
|
+
assert pipeline.run(args) == 0
|
|
27
|
+
manifests = list((tmp_path / "_meta" / "update_runs").glob("*.json"))
|
|
28
|
+
assert len(manifests) == 1
|
|
29
|
+
data = json.loads(manifests[0].read_text(encoding="utf-8"))
|
|
30
|
+
assert data["status"] == "dry_run"
|
|
31
|
+
assert data["tasks"][0]["name"] == "market"
|
|
32
|
+
assert data["tasks"][0]["strategy"] == "always_full"
|
|
33
|
+
|
|
34
|
+
|
|
35
|
+
def test_dry_run_keeps_dependencies_plannable(tmp_path, monkeypatch):
|
|
36
|
+
monkeypatch.setattr(pipeline, "DATA_ROOT", tmp_path)
|
|
37
|
+
monkeypatch.setattr(pipeline, "RUNS_DIR", tmp_path / "_meta" / "update_runs")
|
|
38
|
+
args = type("Args", (), {"tasks": None, "resume": None, "dry_run": True, "no_publish": True})()
|
|
39
|
+
assert pipeline.run(args) == 0
|
|
40
|
+
manifest = next((tmp_path / "_meta" / "update_runs").glob("*.json"))
|
|
41
|
+
statuses = {task["name"]: task["status"] for task in json.loads(manifest.read_text(encoding="utf-8"))["tasks"]}
|
|
42
|
+
assert "skipped_dependency" not in statuses.values()
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|