salesforce-data-customcode 6.1.0.dev7__tar.gz → 7.0.0rc1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/PKG-INFO +1 -1
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/pyproject.toml +1 -1
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/__init__.py +5 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/config.yaml +2 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/deploy.py +3 -5
- salesforce_data_customcode-7.0.0rc1/src/datacustomcode/spark/column_hints.py +133 -0
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/function/example/chunking_with_external_callout/README.md +0 -120
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/function/example/chunking_with_external_callout/entrypoint.py +0 -162
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/function/example/chunking_with_external_callout/external_callout_config.json +0 -11
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/function/example/chunking_with_external_callout/tests/test.json +0 -16
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/function/example/chunking_with_prediction/config.json +0 -3
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/script/examples/external_callout/README.md +0 -140
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/script/examples/external_callout/entrypoint.py +0 -144
- salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/script/examples/external_callout/external_callout_config.json +0 -11
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/LICENSE.txt +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/README.md +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/auth.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/cli.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/client.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/cmd.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/common_config.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/config.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/constants.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/credentials.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_platform_client.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_platform_config.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions/errors.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions/impl/default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions/spark_base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions/spark_default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions/types.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/einstein_predictions_config.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/file/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/file/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/file/path/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/file/path/default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/function/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/function/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/function/feature_types/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/function/feature_types/chunking.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/function/runtime.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/function_utils.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/reader/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/reader/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/reader/query_api.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/reader/sf_cli.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/reader/utils.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/writer/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/writer/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/writer/csv.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/io/writer/print.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/errors.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/spark_base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/spark_default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/types/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/types/generate_text_request.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/types/generate_text_request_builder.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/types/generate_text_response.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway/types/generate_text_response_builder.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/llm_gateway_config.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/mixin.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/direct/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/direct/auth.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/direct/credentials.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/direct/transport.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/direct/url_resolver.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/errors.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/spark_base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/spark_default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/types/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/types/http_method.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/types/http_request.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/types/http_request_builder.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/types/http_response.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential/types/http_response_builder.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/named_credential_config.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/py.typed +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/run.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/scan.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/spark/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/spark/base.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/spark/default.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/template.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/__init__.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/.devcontainer/devcontainer.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/Dockerfile.dependencies +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/README.md +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/build_native_dependencies.sh +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/chunking/payload/config.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/chunking/payload/entrypoint.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/chunking/requirements.txt +0 -0
- {salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/function/example/chunking_with_external_callout → salesforce_data_customcode-7.0.0rc1/src/datacustomcode/templates/function/example/chunking_with_llm}/config.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/example/chunking_with_llm/entrypoint.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/example/chunking_with_llm/files/chunking_prompt.txt +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/example/chunking_with_llm/tests/test.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7/src/datacustomcode/templates/function/example/chunking_with_llm → salesforce_data_customcode-7.0.0rc1/src/datacustomcode/templates/function/example/chunking_with_prediction}/config.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/example/chunking_with_prediction/entrypoint.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/example/chunking_with_prediction/tests/test.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/payload/config.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/payload/entrypoint.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/payload/utility.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/requirements-dev.txt +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/function/requirements.txt +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/.devcontainer/devcontainer.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/Dockerfile +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/Dockerfile.dependencies +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/README.md +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/account.ipynb +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/build_native_dependencies.sh +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/examples/employee_hierarchy/employee_data.csv +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/examples/employee_hierarchy/entrypoint.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/examples/streaming_deltas/entrypoint.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/jupyterlab.sh +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/payload/config.json +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/payload/entrypoint.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/requirements-dev.txt +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/templates/script/requirements.txt +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/token_provider.py +0 -0
- {salesforce_data_customcode-6.1.0.dev7 → salesforce_data_customcode-7.0.0rc1}/src/datacustomcode/version.py +0 -0
|
@@ -13,6 +13,11 @@
|
|
|
13
13
|
# See the License for the specific language governing permissions and
|
|
14
14
|
# limitations under the License.
|
|
15
15
|
|
|
16
|
+
from datacustomcode.spark.column_hints import install_column_casing_hints
|
|
17
|
+
|
|
18
|
+
# Friendlier errors for wrong-case column references. Idempotent.
|
|
19
|
+
install_column_casing_hints()
|
|
20
|
+
|
|
16
21
|
__all__ = [
|
|
17
22
|
"AuthType",
|
|
18
23
|
"Client",
|
|
@@ -18,6 +18,8 @@ spark_config:
|
|
|
18
18
|
spark.submit.deployMode: client
|
|
19
19
|
spark.sql.execution.arrow.pyspark.enabled: 'true'
|
|
20
20
|
spark.driver.extraJavaOptions: -Djava.security.manager=allow
|
|
21
|
+
# Resolve columns case-sensitively so wrong-case refs fail locally too.
|
|
22
|
+
spark.sql.caseSensitive: 'true'
|
|
21
23
|
|
|
22
24
|
einstein_predictions_config:
|
|
23
25
|
type_config_name: DefaultEinsteinPredictions
|
|
@@ -551,15 +551,13 @@ def create_data_transform(
|
|
|
551
551
|
"version": "56.0",
|
|
552
552
|
}
|
|
553
553
|
|
|
554
|
-
|
|
555
|
-
|
|
556
|
-
# an existing materialized table and must not include this field.
|
|
557
|
-
if isinstance(data_transform_config.permissions.write, DmoPermission):
|
|
558
|
-
if not data_transform_config.dataObjects:
|
|
554
|
+
if not data_transform_config.dataObjects:
|
|
555
|
+
if isinstance(data_transform_config.permissions.write, DmoPermission):
|
|
559
556
|
raise ValueError(
|
|
560
557
|
"DMO transforms require 'dataObjects' in config.json describing "
|
|
561
558
|
"the schema of each output DMO."
|
|
562
559
|
)
|
|
560
|
+
else:
|
|
563
561
|
definition["outputDataObjects"] = [
|
|
564
562
|
_data_object_to_output(obj) for obj in data_transform_config.dataObjects
|
|
565
563
|
]
|
|
@@ -0,0 +1,133 @@
|
|
|
1
|
+
# Copyright (c) 2025, Salesforce, Inc.
|
|
2
|
+
# SPDX-License-Identifier: Apache-2
|
|
3
|
+
#
|
|
4
|
+
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
5
|
+
# you may not use this file except in compliance with the License.
|
|
6
|
+
# You may obtain a copy of the License at
|
|
7
|
+
#
|
|
8
|
+
# http://www.apache.org/licenses/LICENSE-2.0
|
|
9
|
+
#
|
|
10
|
+
# Unless required by applicable law or agreed to in writing, software
|
|
11
|
+
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
12
|
+
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
13
|
+
# See the License for the specific language governing permissions and
|
|
14
|
+
# limitations under the License.
|
|
15
|
+
"""Friendlier errors when a script references a column by the wrong case.
|
|
16
|
+
|
|
17
|
+
Data Cloud column names are lowercase. A reference using another case (e.g.
|
|
18
|
+
``df["UnitPrice__c"]``) fails to resolve with a generic error. These hooks
|
|
19
|
+
detect when the lowercase form of a referenced name is a real column and
|
|
20
|
+
prepend an explicit hint, keeping the original message:
|
|
21
|
+
|
|
22
|
+
Column 'UnitPrice__c' not found. Did you mean 'unitprice__c'?
|
|
23
|
+
Data Cloud columns must be lowercase.
|
|
24
|
+
|
|
25
|
+
Two hooks cover all access paths: one for column-resolution errors, and one
|
|
26
|
+
for attribute access (``df.X``), which fails separately. Both only augment an
|
|
27
|
+
error that was already going to be raised, and only for a pure casing mismatch.
|
|
28
|
+
"""
|
|
29
|
+
from __future__ import annotations
|
|
30
|
+
|
|
31
|
+
import logging
|
|
32
|
+
|
|
33
|
+
logger = logging.getLogger(__name__)
|
|
34
|
+
|
|
35
|
+
# Marker on our wrappers so repeated installs never stack.
|
|
36
|
+
_WRAPPED_MARKER = "_datacustomcode_wrapped"
|
|
37
|
+
|
|
38
|
+
|
|
39
|
+
def _strip_backticks(name: str) -> str:
|
|
40
|
+
"""Return the bare column name from a Spark identifier token."""
|
|
41
|
+
name = name.strip()
|
|
42
|
+
if "." in name: # drop any table/alias qualifier
|
|
43
|
+
name = name.split(".")[-1]
|
|
44
|
+
return name.strip().strip("`")
|
|
45
|
+
|
|
46
|
+
|
|
47
|
+
def _casing_hint(bad: str, suggestion: str) -> str:
|
|
48
|
+
return (
|
|
49
|
+
f"Column '{bad}' not found. Did you mean '{suggestion}'? "
|
|
50
|
+
"Data Cloud columns must be lowercase."
|
|
51
|
+
)
|
|
52
|
+
|
|
53
|
+
|
|
54
|
+
def _lowercase_match(bad: str, available: list[str]) -> str | None:
|
|
55
|
+
"""Return the column matching *bad* case-insensitively, else ``None``.
|
|
56
|
+
|
|
57
|
+
Returns ``None`` when *bad* is already lowercase, so a real typo is not
|
|
58
|
+
reported as a casing issue.
|
|
59
|
+
"""
|
|
60
|
+
if not bad or bad == bad.lower():
|
|
61
|
+
return None
|
|
62
|
+
lowered = bad.lower()
|
|
63
|
+
for col in available:
|
|
64
|
+
if col.lower() == lowered:
|
|
65
|
+
return col
|
|
66
|
+
return None
|
|
67
|
+
|
|
68
|
+
|
|
69
|
+
def _install_analysis_exception_hook() -> None:
|
|
70
|
+
"""Add a casing hint to column-resolution errors."""
|
|
71
|
+
from pyspark.errors.exceptions import captured
|
|
72
|
+
|
|
73
|
+
original = captured.convert_exception
|
|
74
|
+
if getattr(original, _WRAPPED_MARKER, False):
|
|
75
|
+
return
|
|
76
|
+
|
|
77
|
+
def convert_exception_with_hint(e): # type: ignore[no-untyped-def]
|
|
78
|
+
exc = original(e)
|
|
79
|
+
try:
|
|
80
|
+
error_class = exc.getErrorClass()
|
|
81
|
+
except Exception: # pragma: no cover - defensive
|
|
82
|
+
return exc
|
|
83
|
+
if not error_class or not error_class.startswith("UNRESOLVED_COLUMN"):
|
|
84
|
+
return exc
|
|
85
|
+
|
|
86
|
+
params = exc.getMessageParameters() or {}
|
|
87
|
+
bad = _strip_backticks(params.get("objectName", ""))
|
|
88
|
+
proposal = params.get("proposal", "")
|
|
89
|
+
available = [
|
|
90
|
+
_strip_backticks(tok) for tok in proposal.split(",") if tok.strip()
|
|
91
|
+
]
|
|
92
|
+
|
|
93
|
+
suggestion = _lowercase_match(bad, available)
|
|
94
|
+
if suggestion is not None: # prepend hint; keep original detail
|
|
95
|
+
exc.desc = f"{_casing_hint(bad, suggestion)}\n{exc.desc}"
|
|
96
|
+
return exc
|
|
97
|
+
|
|
98
|
+
setattr(convert_exception_with_hint, _WRAPPED_MARKER, True)
|
|
99
|
+
captured.convert_exception = convert_exception_with_hint
|
|
100
|
+
|
|
101
|
+
|
|
102
|
+
def _install_getattr_hook() -> None:
|
|
103
|
+
"""Add a casing hint to attribute access (``df.X``)."""
|
|
104
|
+
from pyspark.sql import DataFrame
|
|
105
|
+
|
|
106
|
+
original = DataFrame.__getattr__
|
|
107
|
+
if getattr(original, _WRAPPED_MARKER, False):
|
|
108
|
+
return
|
|
109
|
+
|
|
110
|
+
def __getattr__with_hint(self, name): # type: ignore[no-untyped-def]
|
|
111
|
+
try:
|
|
112
|
+
return original(self, name)
|
|
113
|
+
except AttributeError as exc:
|
|
114
|
+
if name.startswith("_"): # leave dunder/private lookups alone
|
|
115
|
+
raise
|
|
116
|
+
suggestion = _lowercase_match(name, list(self.columns))
|
|
117
|
+
if suggestion is not None: # prepend hint; keep original text
|
|
118
|
+
raise AttributeError(
|
|
119
|
+
f"{_casing_hint(name, suggestion)}\n{exc}"
|
|
120
|
+
) from None
|
|
121
|
+
raise
|
|
122
|
+
|
|
123
|
+
setattr(__getattr__with_hint, _WRAPPED_MARKER, True)
|
|
124
|
+
DataFrame.__getattr__ = __getattr__with_hint # type: ignore[assignment]
|
|
125
|
+
|
|
126
|
+
|
|
127
|
+
def install_column_casing_hints() -> None:
|
|
128
|
+
"""Install the column-casing hint hooks. Idempotent; never raises."""
|
|
129
|
+
try:
|
|
130
|
+
_install_analysis_exception_hook()
|
|
131
|
+
_install_getattr_hook()
|
|
132
|
+
except Exception as exc: # pragma: no cover - defensive
|
|
133
|
+
logger.debug(f"Could not install column-casing hint hooks: {exc}")
|
|
@@ -1,120 +0,0 @@
|
|
|
1
|
-
# Chunking with a Gemini Named Credential Callout
|
|
2
|
-
|
|
3
|
-
Splits each input document into paragraph-sized chunks and calls Google's
|
|
4
|
-
**Gemini** `generateContent` API for every chunk. The model returns a summary,
|
|
5
|
-
category, sentiment, and topics, which are attached to the chunk as citations so
|
|
6
|
-
the search index can filter and rank on them. Gemini is reached through a
|
|
7
|
-
**Named Credential**, so this code never handles the endpoint URL or the API key.
|
|
8
|
-
|
|
9
|
-
## How the callout works
|
|
10
|
-
|
|
11
|
-
```python
|
|
12
|
-
CALLOUT_URL = "callout:gemini" # callout:<NC name>[/<path>]
|
|
13
|
-
|
|
14
|
-
request = (
|
|
15
|
-
HTTPRequestBuilder()
|
|
16
|
-
.set_url(CALLOUT_URL)
|
|
17
|
-
.set_method(HTTPMethod.POST)
|
|
18
|
-
.set_headers({"Content-Type": "application/json"})
|
|
19
|
-
.set_response_timeout_seconds(60) # optional: per-callout timeout (seconds)
|
|
20
|
-
.build()
|
|
21
|
-
)
|
|
22
|
-
# Body is sent verbatim (serialize it yourself); the response body is a raw string.
|
|
23
|
-
response = runtime.named_credential.request(request, json.dumps(payload))
|
|
24
|
-
if response.is_success:
|
|
25
|
-
envelope = json.loads(response.body)
|
|
26
|
-
text = envelope["candidates"][0]["content"]["parts"][0]["text"]
|
|
27
|
-
```
|
|
28
|
-
|
|
29
|
-
The request asks for `responseMimeType: application/json` with a `responseSchema`,
|
|
30
|
-
so Gemini returns the classification as a JSON string in
|
|
31
|
-
`candidates[0].content.parts[0].text` — decode it, then decode that text again.
|
|
32
|
-
|
|
33
|
-
The `gemini` Named Credential's URL already includes the full
|
|
34
|
-
`/v1beta/models/<model>:generateContent` path, so the callout is just
|
|
35
|
-
`callout:gemini` with **no path suffix** (anything after the name is appended to
|
|
36
|
-
the credential's URL).
|
|
37
|
-
|
|
38
|
-
## Configure the Named Credential
|
|
39
|
-
|
|
40
|
-
1. Create an **External Credential** (e.g. `google_api_key`) that injects your
|
|
41
|
-
Gemini API key as the `X-goog-api-key` header.
|
|
42
|
-
2. Create a **Named Credential** named `gemini`:
|
|
43
|
-
- **URL**: `https://generativelanguage.googleapis.com/v1beta/models/gemini-flash-latest:generateContent`
|
|
44
|
-
- **Enabled for Callouts** + **Generate Authorization Header**: on
|
|
45
|
-
- **External Credential**: `google_api_key`
|
|
46
|
-
|
|
47
|
-
## Test locally
|
|
48
|
-
|
|
49
|
-
```bash
|
|
50
|
-
DATACUSTOMCODE_EXTERNAL_CALLOUT_CONFIG=/abs/path/to/external_callout_config.json \
|
|
51
|
-
sf data-code-extension function run \
|
|
52
|
-
--entrypoint payload/entrypoint.py \
|
|
53
|
-
--test-with payload/tests/test.json \
|
|
54
|
-
--target-org <your-org-alias>
|
|
55
|
-
```
|
|
56
|
-
|
|
57
|
-
With `--target-org` the SDK fetches only the **URL** from the org's Named
|
|
58
|
-
Credential; **auth is always taken from `external_callout_config.json`** locally
|
|
59
|
-
(the org's External Credential is used only in the Data Cloud runtime). So the
|
|
60
|
-
`X-goog-api-key` must be in the local config for a local test. Omit
|
|
61
|
-
`--target-org` to run fully offline using `target_url`.
|
|
62
|
-
|
|
63
|
-
```json
|
|
64
|
-
{
|
|
65
|
-
"credentials": {
|
|
66
|
-
"callout:gemini": {
|
|
67
|
-
"auth_type": "Custom",
|
|
68
|
-
"custom_headers": { "X-goog-api-key": "YOUR_GEMINI_API_KEY" },
|
|
69
|
-
"target_url": "https://generativelanguage.googleapis.com/v1beta/models/gemini-flash-latest:generateContent"
|
|
70
|
-
}
|
|
71
|
-
}
|
|
72
|
-
}
|
|
73
|
-
```
|
|
74
|
-
|
|
75
|
-
Place `external_callout_config.json` in the **parent of your payload folder** (or
|
|
76
|
-
point `DATACUSTOMCODE_EXTERNAL_CALLOUT_CONFIG` at it). It is never packaged into
|
|
77
|
-
the deployment zip. Get a key from [Google AI Studio](https://aistudio.google.com/apikey);
|
|
78
|
-
**do not commit it.**
|
|
79
|
-
|
|
80
|
-
## Auth types
|
|
81
|
-
|
|
82
|
-
`auth_type` selects how auth is injected for local testing. It should mirror the
|
|
83
|
-
External Credential your Named Credential uses in the org, so local and deployed
|
|
84
|
-
runs behave the same. This example uses `Custom` (Gemini's `X-goog-api-key`);
|
|
85
|
-
all four supported types:
|
|
86
|
-
|
|
87
|
-
```json
|
|
88
|
-
{
|
|
89
|
-
"credentials": {
|
|
90
|
-
"callout:my_custom_api": {
|
|
91
|
-
"auth_type": "Custom",
|
|
92
|
-
"custom_headers": { "X-goog-api-key": "YOUR_API_KEY" }
|
|
93
|
-
},
|
|
94
|
-
"callout:my_basic_api": {
|
|
95
|
-
"auth_type": "Basic",
|
|
96
|
-
"username": "svc_user",
|
|
97
|
-
"password": "YOUR_PASSWORD"
|
|
98
|
-
},
|
|
99
|
-
"callout:my_oauth_api": {
|
|
100
|
-
"auth_type": "OAuth",
|
|
101
|
-
"access_token": "YOUR_ACCESS_TOKEN"
|
|
102
|
-
},
|
|
103
|
-
"callout:my_jwt_api": {
|
|
104
|
-
"auth_type": "Jwt",
|
|
105
|
-
"token": "YOUR_JWT"
|
|
106
|
-
}
|
|
107
|
-
}
|
|
108
|
-
}
|
|
109
|
-
```
|
|
110
|
-
|
|
111
|
-
| `auth_type` | Fields read | Header sent |
|
|
112
|
-
| ----------- | ------------------------------- | --------------------------------------- |
|
|
113
|
-
| `Basic` | `username`, `password` | `Authorization: Basic <base64 user:pw>` |
|
|
114
|
-
| `Custom` | `custom_headers` (sent verbatim)| the headers you list |
|
|
115
|
-
| `OAuth` | `access_token` or `token` | `Authorization: Bearer <token>` |
|
|
116
|
-
| `Jwt` | `access_token` or `token` | `Authorization: Bearer <token>` |
|
|
117
|
-
|
|
118
|
-
`OAuth`/`Jwt` take a token you supply for the local run — the SDK does not fetch
|
|
119
|
-
or refresh it. In the Data Cloud runtime the Named Credential handles token
|
|
120
|
-
acquisition; this local config only stands in for that during testing.
|
|
@@ -1,162 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env python3
|
|
2
|
-
# Copyright (c) 2025, Salesforce, Inc.
|
|
3
|
-
# SPDX-License-Identifier: Apache-2
|
|
4
|
-
|
|
5
|
-
"""
|
|
6
|
-
Document Chunking with a Gemini Named Credential Callout
|
|
7
|
-
|
|
8
|
-
Splits each input document into paragraph-sized chunks and classifies every
|
|
9
|
-
chunk via Google's Gemini ``generateContent`` API, reached through a Named
|
|
10
|
-
Credential (``callout:gemini``) so the endpoint URL and API key are resolved
|
|
11
|
-
outside this code. The classification is attached to each chunk as citations.
|
|
12
|
-
"""
|
|
13
|
-
|
|
14
|
-
import json
|
|
15
|
-
import logging
|
|
16
|
-
|
|
17
|
-
from datacustomcode.function import Runtime
|
|
18
|
-
from datacustomcode.function.feature_types.chunking import (
|
|
19
|
-
ChunkType,
|
|
20
|
-
SearchIndexChunkingV1Output,
|
|
21
|
-
SearchIndexChunkingV1Request,
|
|
22
|
-
SearchIndexChunkingV1Response,
|
|
23
|
-
)
|
|
24
|
-
from datacustomcode.named_credential.types.http_method import HTTPMethod
|
|
25
|
-
from datacustomcode.named_credential.types.http_request_builder import (
|
|
26
|
-
HTTPRequestBuilder,
|
|
27
|
-
)
|
|
28
|
-
|
|
29
|
-
logger = logging.getLogger(__name__)
|
|
30
|
-
logging.basicConfig(level=logging.INFO)
|
|
31
|
-
|
|
32
|
-
CALLOUT_URL = "callout:gemini"
|
|
33
|
-
|
|
34
|
-
_ANALYSIS_FIELDS = ("summary", "category", "sentiment")
|
|
35
|
-
|
|
36
|
-
_PROMPT = (
|
|
37
|
-
"Analyze the following document chunk and classify it. Respond with its "
|
|
38
|
-
"one-sentence summary, a single-word category, overall sentiment "
|
|
39
|
-
"(positive, negative, or neutral), and up to five key topics.\n\nChunk:\n"
|
|
40
|
-
)
|
|
41
|
-
|
|
42
|
-
# Force Gemini to return the classification as JSON in a fixed shape.
|
|
43
|
-
_GENERATION_CONFIG = {
|
|
44
|
-
"responseMimeType": "application/json",
|
|
45
|
-
"responseSchema": {
|
|
46
|
-
"type": "object",
|
|
47
|
-
"properties": {
|
|
48
|
-
"summary": {"type": "string"},
|
|
49
|
-
"category": {"type": "string"},
|
|
50
|
-
"sentiment": {"type": "string"},
|
|
51
|
-
"topics": {"type": "array", "items": {"type": "string"}},
|
|
52
|
-
},
|
|
53
|
-
"required": ["summary", "category", "sentiment", "topics"],
|
|
54
|
-
},
|
|
55
|
-
}
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
def _chunk_text(text: str, max_words: int = 80) -> list[str]:
|
|
59
|
-
"""Split text into paragraph-aligned chunks of at most ``max_words`` words."""
|
|
60
|
-
paragraphs = [p.strip() for p in text.split("\n\n") if p.strip()]
|
|
61
|
-
|
|
62
|
-
chunks: list[str] = []
|
|
63
|
-
current: list[str] = []
|
|
64
|
-
current_words = 0
|
|
65
|
-
|
|
66
|
-
for paragraph in paragraphs:
|
|
67
|
-
paragraph_words = len(paragraph.split())
|
|
68
|
-
if current and current_words + paragraph_words > max_words:
|
|
69
|
-
chunks.append("\n\n".join(current))
|
|
70
|
-
current = []
|
|
71
|
-
current_words = 0
|
|
72
|
-
current.append(paragraph)
|
|
73
|
-
current_words += paragraph_words
|
|
74
|
-
|
|
75
|
-
if current:
|
|
76
|
-
chunks.append("\n\n".join(current))
|
|
77
|
-
|
|
78
|
-
return chunks
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
def _extract_model_json(body: str) -> dict:
|
|
82
|
-
"""Decode the model's JSON classification from a Gemini response.
|
|
83
|
-
|
|
84
|
-
The generated text sits at ``candidates[0].content.parts[0].text`` and is
|
|
85
|
-
itself a JSON string, so decode twice. Any malformed layer yields ``{}``.
|
|
86
|
-
"""
|
|
87
|
-
try:
|
|
88
|
-
envelope = json.loads(body) if body else {}
|
|
89
|
-
except json.JSONDecodeError:
|
|
90
|
-
return {}
|
|
91
|
-
|
|
92
|
-
try:
|
|
93
|
-
text = envelope["candidates"][0]["content"]["parts"][0]["text"]
|
|
94
|
-
except (KeyError, IndexError, TypeError):
|
|
95
|
-
return {}
|
|
96
|
-
|
|
97
|
-
try:
|
|
98
|
-
payload = json.loads(text)
|
|
99
|
-
except json.JSONDecodeError:
|
|
100
|
-
return {}
|
|
101
|
-
return payload if isinstance(payload, dict) else {}
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
def _analyze_chunk(chunk_text: str, runtime: Runtime) -> dict[str, str]:
|
|
105
|
-
"""Classify one chunk via the Gemini callout and return it as citations."""
|
|
106
|
-
request = (
|
|
107
|
-
HTTPRequestBuilder()
|
|
108
|
-
.set_url(CALLOUT_URL)
|
|
109
|
-
.set_method(HTTPMethod.POST)
|
|
110
|
-
.set_headers({"Content-Type": "application/json", "Accept": "application/json"})
|
|
111
|
-
.set_response_timeout_seconds(60)
|
|
112
|
-
.build()
|
|
113
|
-
)
|
|
114
|
-
|
|
115
|
-
payload = {
|
|
116
|
-
"contents": [{"parts": [{"text": _PROMPT + chunk_text}]}],
|
|
117
|
-
"generationConfig": _GENERATION_CONFIG,
|
|
118
|
-
}
|
|
119
|
-
response = runtime.named_credential.request(request, json.dumps(payload))
|
|
120
|
-
|
|
121
|
-
# Don't raise: a single failed callout shouldn't abort the whole job.
|
|
122
|
-
if not response.is_success:
|
|
123
|
-
logger.error(f"Gemini callout failed with status {response.status_code}")
|
|
124
|
-
return {"analysis_status": "failed", "http_status": str(response.status_code)}
|
|
125
|
-
|
|
126
|
-
data = _extract_model_json(response.body)
|
|
127
|
-
citations = {"analysis_status": "success"}
|
|
128
|
-
for field in _ANALYSIS_FIELDS:
|
|
129
|
-
value = data.get(field)
|
|
130
|
-
citations[field] = str(value) if value is not None else "unavailable"
|
|
131
|
-
|
|
132
|
-
topics = data.get("topics")
|
|
133
|
-
if isinstance(topics, list):
|
|
134
|
-
citations["topics"] = ", ".join(str(topic) for topic in topics)
|
|
135
|
-
|
|
136
|
-
return citations
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
def function(
|
|
140
|
-
request: SearchIndexChunkingV1Request, runtime: Runtime
|
|
141
|
-
) -> SearchIndexChunkingV1Response:
|
|
142
|
-
"""Chunk each input document and classify every chunk via the Gemini API."""
|
|
143
|
-
logger.info(f"Received {len(request.input)} documents to chunk")
|
|
144
|
-
|
|
145
|
-
chunks = []
|
|
146
|
-
chunk_id = 1
|
|
147
|
-
|
|
148
|
-
for doc in request.input:
|
|
149
|
-
for chunk_text in _chunk_text(doc.text):
|
|
150
|
-
citations = _analyze_chunk(chunk_text, runtime)
|
|
151
|
-
|
|
152
|
-
chunk = SearchIndexChunkingV1Output(
|
|
153
|
-
text=chunk_text,
|
|
154
|
-
seq_no=chunk_id,
|
|
155
|
-
chunk_type=ChunkType.TEXT,
|
|
156
|
-
citations=citations,
|
|
157
|
-
)
|
|
158
|
-
chunks.append(chunk)
|
|
159
|
-
chunk_id += 1
|
|
160
|
-
|
|
161
|
-
logger.info(f"Produced {len(chunks)} classified chunks")
|
|
162
|
-
return SearchIndexChunkingV1Response(output=chunks)
|
|
@@ -1,16 +0,0 @@
|
|
|
1
|
-
{
|
|
2
|
-
"input": [
|
|
3
|
-
{
|
|
4
|
-
"text": "Product Review: Northstar Analytics\n\nWe rolled Northstar out to our whole revenue team last quarter and the difference has been night and day. Dashboards that used to take our analysts a full day to assemble now refresh in seconds, and the natural-language query box means our account executives can answer their own questions without filing a ticket.\n\nOnboarding was smoother than any tool we have adopted in years. The guided setup imported our Salesforce data on the first try and the sample templates gave us something useful on day one. Support answered our two questions within the hour. Easily the best purchase decision we made this year."
|
|
5
|
-
},
|
|
6
|
-
{
|
|
7
|
-
"text": "Support Ticket #48210: Repeated timeouts on scheduled exports\n\nFor the third week running our nightly export to the data warehouse has failed silently. There is no alert, no email, nothing in the activity log, and we only find out when the morning report is empty and the leadership meeting has no numbers.\n\nI have raised this twice already and both times the ticket was closed as resolved without anyone actually contacting me. This is costing us real credibility internally and I am extremely frustrated. If the connector cannot handle our volume we need to know now so we can plan a migration, because right now the product is not doing the one job we bought it for."
|
|
8
|
-
},
|
|
9
|
-
{
|
|
10
|
-
"text": "Renewal Feedback: mixed feelings heading into year two\n\nThe core product is genuinely good. The reporting engine is fast, the permissions model is granular enough for our compliance team, and our analysts like working in it. On the functionality alone I would renew without hesitation.\n\nWhat gives me pause is the pricing. The per-seat cost jumped noticeably at renewal and several add-ons that used to be included are now separate line items. The value is still there, but the conversation with my finance team was harder than it should have been, and I would like more transparency before the next cycle."
|
|
11
|
-
},
|
|
12
|
-
{
|
|
13
|
-
"text": "Feature Request: scheduled report subscriptions\n\nWe would like the ability to subscribe internal stakeholders to a report on a recurring schedule so a PDF lands in their inbox every Monday morning. Today we export manually and forward it, which is workable but easy to forget.\n\nA few teams have asked whether subscriptions could support filtered views per recipient, for example each regional manager receiving only their own territory. Not urgent for us, but it would remove a recurring bit of manual work and is something a couple of competing tools already offer."
|
|
14
|
-
}
|
|
15
|
-
]
|
|
16
|
-
}
|
|
@@ -1,140 +0,0 @@
|
|
|
1
|
-
# Transform with a Gemini Named Credential Callout
|
|
2
|
-
|
|
3
|
-
The **transform** calls Google's **Gemini** `generateContent` API to summarize text, and
|
|
4
|
-
write the result back to a DLO. Gemini is reached through a **Named Credential**
|
|
5
|
-
(`callout:gemini`), so this code never handles the endpoint URL or the API key.
|
|
6
|
-
|
|
7
|
-
It shows **both** callout paths against the same Named Credential.
|
|
8
|
-
|
|
9
|
-
## Shared request template
|
|
10
|
-
|
|
11
|
-
```python
|
|
12
|
-
from datacustomcode.client import Client, named_credential_request_col
|
|
13
|
-
|
|
14
|
-
# URL, method and headers apply to every callout on this template.
|
|
15
|
-
request = (
|
|
16
|
-
HTTPRequestBuilder()
|
|
17
|
-
.set_url("callout:gemini") # callout:<NC name>[/<path>]
|
|
18
|
-
.set_method(HTTPMethod.POST)
|
|
19
|
-
.set_headers({"Content-Type": "application/json"})
|
|
20
|
-
.set_response_timeout_seconds(60) # optional: per-callout timeout (seconds)
|
|
21
|
-
.build()
|
|
22
|
-
)
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
## Driver path — one-shot on the driver
|
|
26
|
-
|
|
27
|
-
`Client.named_credential_request` runs the callout **once** on the driver and
|
|
28
|
-
returns an `HTTPResponse` (`.status_code`, `.body`, `.headers`, `.is_success`).
|
|
29
|
-
Use it for a lookup or a shared value you reuse across the job.
|
|
30
|
-
|
|
31
|
-
```python
|
|
32
|
-
response = client.named_credential_request(request, body=json.dumps(payload))
|
|
33
|
-
if response.is_success:
|
|
34
|
-
envelope = json.loads(response.body)
|
|
35
|
-
text = envelope["candidates"][0]["content"]["parts"][0]["text"]
|
|
36
|
-
```
|
|
37
|
-
|
|
38
|
-
## Per-row path — fan out across the DataFrame
|
|
39
|
-
|
|
40
|
-
`named_credential_request_col` dispatches one callout per row; only the body
|
|
41
|
-
Column varies.
|
|
42
|
-
|
|
43
|
-
```python
|
|
44
|
-
# One callout per row; body is a Column built from the row's data.
|
|
45
|
-
df = df.withColumn("_callout", named_credential_request_col(request, body=body_col))
|
|
46
|
-
```
|
|
47
|
-
|
|
48
|
-
It returns a struct Column:
|
|
49
|
-
|
|
50
|
-
```
|
|
51
|
-
{status, response: {status_code, body, headers}, error_code, error_message}
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
- A non-2xx response is still `status = "SUCCESS"` with the HTTP code in
|
|
55
|
-
`response.status_code` — extracting the model text just yields null for that row.
|
|
56
|
-
- A transport failure sets `status = "ERROR"`; the row survives, the job does not
|
|
57
|
-
abort. Pull the model text out of `response.body` with `get_json_object(...)`.
|
|
58
|
-
|
|
59
|
-
Both paths resolve the same Named Credential and read auth from the same local
|
|
60
|
-
`external_callout_config.json`.
|
|
61
|
-
|
|
62
|
-
The `gemini` Named Credential's URL already includes the full
|
|
63
|
-
`/v1beta/models/<model>:generateContent` path, so the callout is just
|
|
64
|
-
`callout:gemini` with **no path suffix** (anything after the name is appended to
|
|
65
|
-
the credential's URL).
|
|
66
|
-
|
|
67
|
-
## Configure the Named Credential
|
|
68
|
-
|
|
69
|
-
1. Create an **External Credential** (e.g. `google_api_key`) that injects your
|
|
70
|
-
Gemini API key as the `X-goog-api-key` header.
|
|
71
|
-
2. Create a **Named Credential** named `gemini`:
|
|
72
|
-
- **URL**: `https://generativelanguage.googleapis.com/v1beta/models/gemini-flash-latest:generateContent`
|
|
73
|
-
- **Enabled for Callouts** + **Generate Authorization Header**: on
|
|
74
|
-
- **External Credential**: `google_api_key`
|
|
75
|
-
|
|
76
|
-
## Test locally
|
|
77
|
-
|
|
78
|
-
Copy `entrypoint.py` into your `payload/` folder (or point the run at it), then:
|
|
79
|
-
|
|
80
|
-
```bash
|
|
81
|
-
DATACUSTOMCODE_EXTERNAL_CALLOUT_CONFIG=/abs/path/to/external_callout_config.json \
|
|
82
|
-
sf data-code-extension script run --entrypoint entrypoint.py --target-org <your-org-alias>
|
|
83
|
-
```
|
|
84
|
-
|
|
85
|
-
With `----target-org` the SDK fetches only the **URL** from the org's Named
|
|
86
|
-
Credential; **auth is always taken from `external_callout_config.json`** locally
|
|
87
|
-
(the org's External Credential is used only in the Data Cloud runtime). So the
|
|
88
|
-
`X-goog-api-key` must be in the local config for a local test. Omit
|
|
89
|
-
`--target-org` to run fully offline using `target_url`.
|
|
90
|
-
|
|
91
|
-
```json
|
|
92
|
-
{
|
|
93
|
-
"credentials": {
|
|
94
|
-
"callout:gemini": {
|
|
95
|
-
"auth_type": "Custom",
|
|
96
|
-
"custom_headers": { "X-goog-api-key": "YOUR_GEMINI_API_KEY" },
|
|
97
|
-
"target_url": "https://generativelanguage.googleapis.com/v1beta/models/gemini-flash-latest:generateContent"
|
|
98
|
-
}
|
|
99
|
-
}
|
|
100
|
-
}
|
|
101
|
-
```
|
|
102
|
-
|
|
103
|
-
Place `external_callout_config.json` in the **parent of your payload folder** (or
|
|
104
|
-
point `DATACUSTOMCODE_EXTERNAL_CALLOUT_CONFIG` at it). It is never packaged into
|
|
105
|
-
the deployment zip. Get a key from [Google AI Studio](https://aistudio.google.com/apikey);
|
|
106
|
-
**do not commit it.**
|
|
107
|
-
|
|
108
|
-
## What it reads / writes
|
|
109
|
-
|
|
110
|
-
`config.json` declares the DLO permissions for deployment:
|
|
111
|
-
|
|
112
|
-
| | DLO | Notes |
|
|
113
|
-
| ------ | ------------------- | --------------------------------------------- |
|
|
114
|
-
| read | `Account_std__dll` | source rows; `description__c` is summarized |
|
|
115
|
-
| write | `Account_std_copy__dll` | adds `summary__c`, `callout_status__c`, `callout_http_code__c` |
|
|
116
|
-
|
|
117
|
-
Adjust `_TEXT_COLUMN`, `_SOURCE_DLO` and `_TARGET_DLO` in `entrypoint.py` (and the
|
|
118
|
-
matching entries in `config.json`) to point at your own DLOs.
|
|
119
|
-
|
|
120
|
-
## Auth types
|
|
121
|
-
|
|
122
|
-
`auth_type` selects how auth is injected for local testing. It should mirror the
|
|
123
|
-
External Credential your Named Credential uses in the org, so local and deployed
|
|
124
|
-
runs behave the same. This example uses `Custom` (Gemini's `X-goog-api-key`);
|
|
125
|
-
all supported types:
|
|
126
|
-
|
|
127
|
-
| `auth_type` | Fields read | Header sent |
|
|
128
|
-
| ----------- | ------------------------------- | --------------------------------------- |
|
|
129
|
-
| `Basic` | `username`, `password` | `Authorization: Basic <base64 user:pw>` |
|
|
130
|
-
| `Custom` | `custom_headers` (sent verbatim)| the headers you list |
|
|
131
|
-
| `OAuth` | `access_token` or `token` | `Authorization: Bearer <token>` |
|
|
132
|
-
| `Jwt` | `access_token` or `token` | `Authorization: Bearer <token>` |
|
|
133
|
-
| `AwsSv4` | `aws_access_key_id`, `aws_secret_access_key`, `aws_region`, `aws_service`, optional `aws_session_token` | `Authorization: AWS4-HMAC-SHA256 ...` plus `x-amz-date` / `x-amz-content-sha256` (and `x-amz-security-token` when a session token is set) |
|
|
134
|
-
|
|
135
|
-
`OAuth`/`Jwt` take a token you supply for the local run — the SDK does not fetch
|
|
136
|
-
or refresh it. In the Data Cloud runtime the Named Credential handles token
|
|
137
|
-
acquisition; this local config only stands in for that during testing.
|
|
138
|
-
|
|
139
|
-
`AwsSv4` signs the request with AWS Signature Version 4 using the keys you
|
|
140
|
-
supply, mirroring an AWS Signature Version 4 External Credential in the org.
|