ferrox-py-utils 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,5 @@
1
+ venv/
2
+ __pycache__/
3
+ *.pyc
4
+ .env
5
+ .pytest_cache/
@@ -0,0 +1,67 @@
1
+ Metadata-Version: 2.5
2
+ Name: ferrox-py-utils
3
+ Version: 1.0.0
4
+ Summary: Data Engineering and ETL utilities for the Ferrox ecosystem.
5
+ Author: AI-Autistic-Intelligence
6
+ Requires-Python: >=3.11
7
+ Requires-Dist: ferrox-py>=1.0.0
8
+ Requires-Dist: pydantic>=2.0
9
+ Description-Content-Type: text/markdown
10
+
11
+ # 🛠️ Ferrox-Py-Utils (Data Engineering)
12
+
13
+ ## 1. Overview (What does this do?)
14
+ The `ferrox-py-utils` package is a specialized extension of the Ferrox ecosystem dedicated to data manipulation, ETL (Extract, Transform, Load) pipelines, and cross-platform data movement. It provides agnostic `Connectors` (e.g., for AWS S3, CSV files, or REST APIs) and a `PipelineOrchestrator` to seamlessly sequence data transformation jobs without writing monolithic scripts.
15
+
16
+ ## 2. Philosophy (Why does it exist?)
17
+ Data engineering often suffers from "wild copy-pasting" where scripts to import users or export CSVs are hastily written and tightly coupled to specific database schemas or cloud vendors. The philosophy here is absolute **agnosticism**. Instead of building monolithic ETL scripts, this package enforces a modular approach where small, isolated tasks are injected into an orchestrator. This allows data engineers to reuse extraction and validation logic across entirely different projects.
18
+
19
+ ## 3. Target Audience (Who is it for?)
20
+ This package is built for Data Engineers and backend developers who need to quickly stand up a Data Platform for ingestion, parsing, and bulk loading of large datasets (like CSV or JSON) into Data Lakes or Object Storage (such as Amazon S3, MinIO, or Google Cloud Storage).
21
+
22
+ ## 4. Architecture (How does it work?)
23
+ The data architecture rests on three pillars:
24
+ - **Connectors**: Classes inheriting from an abstract `BaseConnector` to standardize streaming I/O operations (Read/Write/Delete), entirely isolating the pipeline from the specific storage vendor.
25
+ - **PipelineOrchestrator**: A linear execution engine that passes a shared state (`context`) through a sequence of nodes (steps), ensuring proper error isolation and retry mechanics.
26
+ - **Schema Registry**: Native integration with Pydantic to register and enforce strict validation on datasets in transit, ensuring corrupted data never enters the database.
27
+
28
+ ## 5. Installation / Setup
29
+ Ensure you are using Python 3.11+. The package installs basic dependencies, but you may need to install specific data drivers (like `boto3` or `pandas`) depending on the connectors you intend to use.
30
+
31
+ ```bash
32
+ pip install ferrox-py-utils
33
+ # Optional extensions:
34
+ # pip install boto3 pandas
35
+ ```
36
+
37
+ ## 6. Quickstart (Usage)
38
+ ```python
39
+ from ferrox_py_utils.connectors.csv import CsvConnector
40
+ from ferrox_py_utils.pipelines.orchestrator import PipelineOrchestrator
41
+
42
+ # 1. Setup the agnostic connector
43
+ csv_connector = CsvConnector(file_path="/tmp/data.csv")
44
+
45
+ # 2. Define isolated pipeline steps
46
+ def step_read_csv(ctx):
47
+ ctx['data'] = csv_connector.read()
48
+ return ctx
49
+
50
+ def step_transform(ctx):
51
+ # Transform logic here...
52
+ ctx['data'] = [row for row in ctx['data'] if row.get("active")]
53
+ return ctx
54
+
55
+ # 3. Orchestrate and execute
56
+ orchestrator = PipelineOrchestrator()
57
+ orchestrator.add_step("Read Data", step_read_csv)
58
+ orchestrator.add_step("Clean Data", step_transform)
59
+
60
+ final_context = orchestrator.execute()
61
+ print(f"Processed {len(final_context['data'])} records.")
62
+ ```
63
+
64
+ ## 7. Ecosystem Integration
65
+ This module is fully integrated with the core `ferrox-py` ecosystem:
66
+ - **Core (IoC Container)**: The framework's Dependency Injection container is used to instantiate Connectors globally as Singletons, meaning you only initialize your S3 credentials once.
67
+ - **AuthModule (ferrox-py-auth)**: RBAC can be utilized to restrict which users or system roles are authorized to trigger specific data pipelines via the API Gateway.
@@ -0,0 +1,57 @@
1
+ # 🛠️ Ferrox-Py-Utils (Data Engineering)
2
+
3
+ ## 1. Overview (What does this do?)
4
+ The `ferrox-py-utils` package is a specialized extension of the Ferrox ecosystem dedicated to data manipulation, ETL (Extract, Transform, Load) pipelines, and cross-platform data movement. It provides agnostic `Connectors` (e.g., for AWS S3, CSV files, or REST APIs) and a `PipelineOrchestrator` to seamlessly sequence data transformation jobs without writing monolithic scripts.
5
+
6
+ ## 2. Philosophy (Why does it exist?)
7
+ Data engineering often suffers from "wild copy-pasting" where scripts to import users or export CSVs are hastily written and tightly coupled to specific database schemas or cloud vendors. The philosophy here is absolute **agnosticism**. Instead of building monolithic ETL scripts, this package enforces a modular approach where small, isolated tasks are injected into an orchestrator. This allows data engineers to reuse extraction and validation logic across entirely different projects.
8
+
9
+ ## 3. Target Audience (Who is it for?)
10
+ This package is built for Data Engineers and backend developers who need to quickly stand up a Data Platform for ingestion, parsing, and bulk loading of large datasets (like CSV or JSON) into Data Lakes or Object Storage (such as Amazon S3, MinIO, or Google Cloud Storage).
11
+
12
+ ## 4. Architecture (How does it work?)
13
+ The data architecture rests on three pillars:
14
+ - **Connectors**: Classes inheriting from an abstract `BaseConnector` to standardize streaming I/O operations (Read/Write/Delete), entirely isolating the pipeline from the specific storage vendor.
15
+ - **PipelineOrchestrator**: A linear execution engine that passes a shared state (`context`) through a sequence of nodes (steps), ensuring proper error isolation and retry mechanics.
16
+ - **Schema Registry**: Native integration with Pydantic to register and enforce strict validation on datasets in transit, ensuring corrupted data never enters the database.
17
+
18
+ ## 5. Installation / Setup
19
+ Ensure you are using Python 3.11+. The package installs basic dependencies, but you may need to install specific data drivers (like `boto3` or `pandas`) depending on the connectors you intend to use.
20
+
21
+ ```bash
22
+ pip install ferrox-py-utils
23
+ # Optional extensions:
24
+ # pip install boto3 pandas
25
+ ```
26
+
27
+ ## 6. Quickstart (Usage)
28
+ ```python
29
+ from ferrox_py_utils.connectors.csv import CsvConnector
30
+ from ferrox_py_utils.pipelines.orchestrator import PipelineOrchestrator
31
+
32
+ # 1. Setup the agnostic connector
33
+ csv_connector = CsvConnector(file_path="/tmp/data.csv")
34
+
35
+ # 2. Define isolated pipeline steps
36
+ def step_read_csv(ctx):
37
+ ctx['data'] = csv_connector.read()
38
+ return ctx
39
+
40
+ def step_transform(ctx):
41
+ # Transform logic here...
42
+ ctx['data'] = [row for row in ctx['data'] if row.get("active")]
43
+ return ctx
44
+
45
+ # 3. Orchestrate and execute
46
+ orchestrator = PipelineOrchestrator()
47
+ orchestrator.add_step("Read Data", step_read_csv)
48
+ orchestrator.add_step("Clean Data", step_transform)
49
+
50
+ final_context = orchestrator.execute()
51
+ print(f"Processed {len(final_context['data'])} records.")
52
+ ```
53
+
54
+ ## 7. Ecosystem Integration
55
+ This module is fully integrated with the core `ferrox-py` ecosystem:
56
+ - **Core (IoC Container)**: The framework's Dependency Injection container is used to instantiate Connectors globally as Singletons, meaning you only initialize your S3 credentials once.
57
+ - **AuthModule (ferrox-py-auth)**: RBAC can be utilized to restrict which users or system roles are authorized to trigger specific data pipelines via the API Gateway.
@@ -0,0 +1,50 @@
1
+ # Connectors
2
+
3
+ ## 1. Overview (What does this do?)
4
+ The `connectors` module standardizes read and write access to external data sources. It provides a set of pre-built, reusable classes that can interact with Object Storage (S3, MinIO), local file systems, and tabular formats (CSV, Excel) using a unified programming interface.
5
+
6
+ ## 2. Philosophy (Why does it exist?)
7
+ Interacting with external data sources usually involves writing boilerplate code using varying SDKs (e.g., `boto3` for S3, native `open()` for files, `requests` for APIs). This creates tightly coupled code that is difficult to test and impossible to swap out easily. By forcing all data I/O through a generic `BaseConnector` interface, the business logic remains entirely oblivious to where the data is actually coming from or going to, allowing developers to easily mock these interactions during unit testing.
8
+
9
+ ## 3. Target Audience (Who is it for?)
10
+ This module is intended for Data Engineers writing ETL pipelines or backend developers who need to generate, upload, or process reports, backups, and user-uploaded files without wrestling with raw provider SDKs.
11
+
12
+ ## 4. Architecture (How does it work?)
13
+ - **BaseConnector**: An abstract base class defining the strict I/O contract (`read` and `write` methods).
14
+ - **S3Connector**: Facilitates interaction with S3-compatible object storage (AWS, MinIO). It handles authentication securely and manages streaming byte uploads and downloads to prevent memory exhaustion on large files.
15
+ - **CsvConnector**: Simplifies handling CSV files, natively supporting chunking mechanisms (often wrapping Pandas or the native Python `csv` module) to process massive datasets sequentially.
16
+
17
+ ## 5. Installation / Setup
18
+ The `BaseConnector` requires no external libraries. However, to use the specific implementations, you must install the respective drivers.
19
+
20
+ ```bash
21
+ # For S3 connectivity
22
+ pip install boto3
23
+
24
+ # For advanced CSV/Excel manipulation
25
+ pip install pandas
26
+ ```
27
+
28
+ ## 6. Quickstart (Usage)
29
+ ```python
30
+ from ferrox_py_utils.connectors.s3 import S3Connector
31
+
32
+ # 1. Initialize the connector (usually done within the IoC Container)
33
+ s3_connector = S3Connector(
34
+ endpoint_url="http://localhost:9000",
35
+ access_key="minioadmin",
36
+ secret_key="minioadmin",
37
+ bucket_name="bronze-layer"
38
+ )
39
+
40
+ # 2. Use the unified interface
41
+ def backup_database_dump(data_bytes: bytes):
42
+ # The business logic only calls 'write'
43
+ s3_connector.write(file_path="backups/db_dump.sql", data=data_bytes)
44
+
45
+ # 3. Read data back
46
+ raw_data = s3_connector.read(file_path="backups/db_dump.sql")
47
+ ```
48
+
49
+ ## 7. Ecosystem Integration
50
+ Connectors are explicitly designed to be registered as **Providers** in the core `ferrox-py` Dependency Injection Container. This ensures that the application only maintains one active connection pool to S3 or a remote FTP server, injecting it seamlessly into the **PipelineOrchestrator** steps when needed.
@@ -0,0 +1,43 @@
1
+ # Ferrox-Py-Utils Overview
2
+
3
+ ## 1. Overview (What does this do?)
4
+ The `ferrox-py-utils` package implements a decoupled data architecture based on interchangeable connectors and a linear execution pipeline. It provides the scaffolding required to build robust Data Platforms entirely within the Ferrox ecosystem.
5
+
6
+ ## 2. Philosophy (Why does it exist?)
7
+ Traditional ETL (Extract, Transform, Load) scripts are often written as massive, procedural files (e.g., `import_users.py`). These scripts are incredibly fragile, hard to test, and impossible to reuse. This package mandates a functional, "matryoshka" approach: extraction is isolated to Connectors, transformation is isolated to pure functions (Steps), and execution is managed by an Orchestrator. This ensures that a bug in one transformation step doesn't crash the entire data ingestion process without proper logging.
8
+
9
+ ## 3. Target Audience (Who is it for?)
10
+ This overview is for Software Architects and Data Engineers who are designing the data ingestion, processing, and warehousing layers of their application and need a structured, maintainable framework rather than a collection of disorganized Python scripts.
11
+
12
+ ## 4. Architecture (How does it work?)
13
+ The data engineering process is strictly divided into three phases:
14
+ 1. **Isolated Extraction (I/O)**: Data is read from external sources (S3, Databases, APIs) exclusively through specialized *Connectors*.
15
+ 2. **In-Memory Transformation**: A *PipelineOrchestrator* executes a sequence of functional steps, mutating a shared state object (`context`) in a predictable order.
16
+ 3. **Guaranteed Validation**: The *Schema Registry* ensures that the data being transformed strictly adheres to predefined contracts (DTOs) before it is passed to the next step or loaded into the final Database (the Bronze/Silver/Gold layers).
17
+
18
+ ## 5. Installation / Setup
19
+ The package is installed via standard pip mechanics. It relies heavily on the core `ferrox-py` framework for architectural consistency.
20
+
21
+ ```bash
22
+ pip install ferrox-py-utils
23
+ ```
24
+
25
+ ## 6. Quickstart (Usage)
26
+ ```python
27
+ from ferrox_py.core.container import Container
28
+ from ferrox_py_utils.connectors.s3 import S3Connector
29
+
30
+ # Centralized setup of data dependencies
31
+ container = Container()
32
+ container.register("s3_connector", S3Connector(
33
+ endpoint_url="https://s3.amazonaws.com",
34
+ access_key="...",
35
+ secret_key="...",
36
+ bucket_name="my-data-lake"
37
+ ))
38
+
39
+ # The connector is now available globally for any Pipeline to use
40
+ ```
41
+
42
+ ## 7. Ecosystem Integration
43
+ The entire philosophy of this package is built upon the **Inversion of Control (IoC) Container** provided by the `ferrox-py` core. Connectors and Pipeline steps are resolved dynamically via Dependency Injection, ensuring that unit testing an ETL pipeline doesn't require spinning up an actual AWS environment.
@@ -0,0 +1,45 @@
1
+ # Pipelines
2
+
3
+ ## 1. Overview (What does this do?)
4
+ The `PipelineOrchestrator` is the execution engine at the heart of the `ferrox-py-utils` data engineering toolkit. It provides a linear, procedural runner that executes a chain of responsibility (a sequence of isolated tasks or "steps") to process data safely and predictably.
5
+
6
+ ## 2. Philosophy (Why does it exist?)
7
+ When data transformations are written as one large function, handling specific errors (like a network timeout on step 4 of 10) is a nightmare. The Orchestrator exists to enforce the single-responsibility principle. By breaking an ETL job down into small, functional steps that only mutate a shared `context` dictionary, developers can easily unit test individual transformations, insert retry logic automatically, and halt the pipeline gracefully upon failure.
8
+
9
+ ## 3. Target Audience (Who is it for?)
10
+ This module is for Data Engineers who are writing complex ETL or ELT jobs, data migration scripts, or nightly batch processes that require clear execution boundaries, error isolation, and detailed logging at each stage.
11
+
12
+ ## 4. Architecture (How does it work?)
13
+ - **Shared Context**: The orchestrator initializes a `context` (usually a dictionary) that is passed sequentially from one step to the next.
14
+ - **Adding Steps**: Tasks are registered in the orchestrator via `add_step(name, function)`. The function must accept the `context` as an argument and return it.
15
+ - **Execution Engine**: The `execute()` method runs the steps in order. Because it manages the execution loop, it can natively implement advanced features like automated retry policies, step-timing metrics, or alert dispatching if a specific node crashes.
16
+
17
+ ## 5. Installation / Setup
18
+ The `PipelineOrchestrator` is natively included in the `ferrox-py-utils` package. It uses standard Python synchronous or asynchronous mechanics and requires no external dependencies.
19
+
20
+ ## 6. Quickstart (Usage)
21
+ ```python
22
+ from ferrox_py_utils.pipelines.orchestrator import PipelineOrchestrator
23
+
24
+ # 1. Define isolated, pure functions (Steps)
25
+ def extract_data(ctx):
26
+ ctx['raw_data'] = [1, 2, 3]
27
+ return ctx
28
+
29
+ def transform_data(ctx):
30
+ # Mutates the shared context
31
+ ctx['transformed'] = [x * 10 for x in ctx['raw_data']]
32
+ return ctx
33
+
34
+ # 2. Build the Pipeline
35
+ pipeline = PipelineOrchestrator()
36
+ pipeline.add_step("Extraction Step", extract_data)
37
+ pipeline.add_step("Transformation Step", transform_data)
38
+
39
+ # 3. Execute the chain
40
+ final_context = pipeline.execute()
41
+ assert final_context['transformed'] == [10, 20, 30]
42
+ ```
43
+
44
+ ## 7. Ecosystem Integration
45
+ The `PipelineOrchestrator` relies heavily on the **Observability Component** from the core `ferrox-py` framework. Because the orchestrator controls the execution loop, it automatically injects logging trace IDs and emits structured logs (e.g., `Step 'Extraction Step' completed in 0.4s`) into the centralized logging system, ensuring complete visibility over background data jobs.
@@ -0,0 +1,47 @@
1
+ # Schema Registry
2
+
3
+ ## 1. Overview (What does this do?)
4
+ The `SchemaRegistry` module guarantees data quality during ETL processes. It provides a centralized dictionary that maps string identifiers to Pydantic models (or Dataclasses), allowing the system to perform rigorous data casting and validation on raw, untyped datasets.
5
+
6
+ ## 2. Philosophy (Why does it exist?)
7
+ When ingesting data from external sources (like a third-party CSV or a public JSON API), the data is inherently untrusted and untyped. Directly inserting this raw data into the Database (the Bronze/Silver/Gold layers) leads to corrupted records and fatal crashes. This module enforces the philosophy that data must be formally validated at the application boundary. If a CSV row says an ID is `"123"`, the Schema Registry mathematically proves and casts it to an `int(123)` before it is allowed further into the pipeline.
8
+
9
+ ## 3. Target Audience (Who is it for?)
10
+ This component is vital for Data Engineers who manage data lakes or data warehouses, and backend developers who need to ensure that bulk data imports adhere strictly to the internal domain models of the application.
11
+
12
+ ## 4. Architecture (How does it work?)
13
+ - **Registration**: At application boot, developers map string names (e.g., `"user_import"`) to corresponding Pydantic schemas.
14
+ - **Validation**: During a pipeline execution, a transformation step can invoke `SchemaRegistry.validate("user_import", raw_dict)`. The registry locates the schema and processes the dictionary.
15
+ - **Error Handling**: If the raw data violates the schema (e.g., a missing required field or wrong data type), Pydantic raises a detailed validation error, allowing the orchestrator to log the specific malformed row and continue or halt appropriately.
16
+
17
+ ## 5. Installation / Setup
18
+ The Schema Registry relies heavily on the `pydantic` library to perform its fast, strict type checking.
19
+
20
+ ```bash
21
+ pip install pydantic
22
+ ```
23
+
24
+ ## 6. Quickstart (Usage)
25
+ ```python
26
+ from ferrox_py_utils.schemas.registry import SchemaRegistry
27
+ from pydantic import BaseModel
28
+
29
+ # 1. Define the strict data contract
30
+ class UserImportSchema(BaseModel):
31
+ id: int
32
+ email: str
33
+
34
+ # 2. Register the schema centrally
35
+ registry = SchemaRegistry()
36
+ registry.register("user_import", UserImportSchema)
37
+
38
+ # 3. Automatic casting and validation inside a Pipeline step
39
+ raw_data = {"id": "123", "email": "test@test.com"}
40
+ valid_user = registry.validate("user_import", raw_data)
41
+
42
+ # The string "123" has been safely cast to an integer
43
+ print(type(valid_user.id)) # <class 'int'>
44
+ ```
45
+
46
+ ## 7. Ecosystem Integration
47
+ The Schema Registry serves the exact same purpose in the Data Engineering layer as the **Validation Pipe** (Layer 5) serves in the Web/HTTP layer of the core `ferrox-py` framework. It ensures that the **CQRS Bus** and **Data Repositories** only ever interact with strongly-typed, predictable Domain Objects.
@@ -0,0 +1 @@
1
+ # init
@@ -0,0 +1,19 @@
1
+ from abc import ABC, abstractmethod
2
+ from typing import AsyncGenerator, Any
3
+
4
+ class DataConnector(ABC):
5
+ @abstractmethod
6
+ async def connect(self):
7
+ pass
8
+
9
+ @abstractmethod
10
+ async def extract(self, query: str = None) -> AsyncGenerator[Any, None]:
11
+ pass
12
+
13
+ @abstractmethod
14
+ async def load(self, data: Any) -> bool:
15
+ pass
16
+
17
+ @abstractmethod
18
+ async def close(self):
19
+ pass
@@ -0,0 +1,35 @@
1
+ import csv
2
+ import aiofiles
3
+ from typing import AsyncGenerator, Any
4
+ from .base import DataConnector
5
+ from ferrox_py.core.provider import injectable
6
+
7
+ @injectable()
8
+ class CsvConnector(DataConnector):
9
+ def __init__(self, file_path: str):
10
+ self.file_path = file_path
11
+ self._file = None
12
+
13
+ async def connect(self):
14
+ # In a real app we'd keep it open for stream
15
+ pass
16
+
17
+ async def extract(self, query: str = None) -> AsyncGenerator[dict, None]:
18
+ async with aiofiles.open(self.file_path, mode='r', encoding='utf-8') as f:
19
+ header = None
20
+ async for line in f:
21
+ row = line.strip().split(",")
22
+ if not header:
23
+ header = row
24
+ continue
25
+ yield dict(zip(header, row))
26
+
27
+ async def load(self, data: Any) -> bool:
28
+ # Simplistic append
29
+ async with aiofiles.open(self.file_path, mode='a', encoding='utf-8') as f:
30
+ if isinstance(data, dict):
31
+ await f.write(",".join(str(v) for v in data.values()) + "\n")
32
+ return True
33
+
34
+ async def close(self):
35
+ pass
@@ -0,0 +1,26 @@
1
+ from typing import AsyncGenerator, Any
2
+ from .base import DataConnector
3
+ from ferrox_py.core.provider import injectable
4
+
5
+ @injectable()
6
+ class S3Connector(DataConnector):
7
+ def __init__(self, bucket: str, path: str):
8
+ self.bucket = bucket
9
+ self.path = path
10
+ # Would inject aioboto3 session here
11
+
12
+ async def connect(self):
13
+ print(f"Connecting to S3 Bucket: {self.bucket}...")
14
+ pass
15
+
16
+ async def extract(self, query: str = None) -> AsyncGenerator[Any, None]:
17
+ print(f"Extracting streaming chunks from s3://{self.bucket}/{self.path}")
18
+ # Mock streaming chunks
19
+ yield {"chunk_id": 1, "data": b"mock_data"}
20
+
21
+ async def load(self, data: Any) -> bool:
22
+ print(f"Uploading chunk to s3://{self.bucket}/{self.path}")
23
+ return True
24
+
25
+ async def close(self):
26
+ pass
@@ -0,0 +1,28 @@
1
+ from typing import Callable, List, Any
2
+ import asyncio
3
+ from ferrox_py.core.provider import injectable
4
+ from ferrox_py.core.errors import FerroxError
5
+
6
+ class PipelineStep:
7
+ def __init__(self, name: str, execute_fn: Callable[[Any], Any]):
8
+ self.name = name
9
+ self.execute_fn = execute_fn
10
+
11
+ @injectable()
12
+ class PipelineOrchestrator:
13
+ async def execute_pipeline(self, name: str, initial_data: Any, steps: List[PipelineStep]) -> Any:
14
+ print(f"Starting Pipeline: {name}")
15
+ current_data = initial_data
16
+
17
+ for step in steps:
18
+ print(f" -> Executing Step: {step.name}")
19
+ try:
20
+ if asyncio.iscoroutinefunction(step.execute_fn):
21
+ current_data = await step.execute_fn(current_data)
22
+ else:
23
+ current_data = step.execute_fn(current_data)
24
+ except Exception as e:
25
+ raise FerroxError(message=f"Pipeline '{name}' failed at step '{step.name}': {e}", status_code=500)
26
+
27
+ print(f"Pipeline '{name}' finished successfully.")
28
+ return current_data
@@ -0,0 +1,25 @@
1
+ from typing import Dict, Type
2
+ from pydantic import create_model, BaseModel, ValidationError
3
+ from ferrox_py.core.provider import injectable
4
+ from ferrox_py.core.errors import FerroxError
5
+
6
+ @injectable()
7
+ class SchemaRegistry:
8
+ def __init__(self):
9
+ self._schemas: Dict[str, Type[BaseModel]] = {}
10
+
11
+ def register_schema(self, name: str, schema_def: Dict[str, Type]):
12
+ model = create_model(name, **schema_def)
13
+ self._schemas[name] = model
14
+ print(f"Schema '{name}' registered.")
15
+
16
+ def validate(self, name: str, data: dict) -> dict:
17
+ if name not in self._schemas:
18
+ raise FerroxError(message=f"Schema {name} not found", status_code=404)
19
+
20
+ model = self._schemas[name]
21
+ try:
22
+ instance = model(**data)
23
+ return instance.model_dump()
24
+ except ValidationError as e:
25
+ raise FerroxError(message=f"Schema validation failed: {e.errors()}", status_code=400)
@@ -0,0 +1,15 @@
1
+ [build-system]
2
+ requires = ["hatchling"]
3
+ build-backend = "hatchling.build"
4
+
5
+ [project]
6
+ name = "ferrox-py-utils"
7
+ version = "1.0.0"
8
+ description = "Data Engineering and ETL utilities for the Ferrox ecosystem."
9
+ authors = [{ name = "AI-Autistic-Intelligence" }]
10
+ readme = "README.md"
11
+ requires-python = ">=3.11"
12
+ dependencies = [
13
+ "ferrox-py>=1.0.0",
14
+ "pydantic>=2.0"
15
+ ]