diff --git a/umn/source/change_history.rst b/umn/source/change_history.rst index 098f7d7..e8d8d8c 100644 --- a/umn/source/change_history.rst +++ b/umn/source/change_history.rst @@ -6,8 +6,13 @@ Change History ============== +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ -| Released On | What's New | +| Released On | Change Description | +===================================+====================================================================================================================================================================================================================================================+ +| 2025-02-10 | Modified the following sections: | +| | | +| | - Added the descriptions of IPv6 enabling to relevant parameters in :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. | +| | - Added the descriptions of IPv6 enabling to relevant parameters in :ref:`Creating an Enhanced Datasource Connection `. | ++-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | 2025-09-25 | Modified the following sections: | | | | | | - Added the resource specification configuration descriptions of v1 and v2 to :ref:`Creating a Flink OpenSource SQL Job `. | @@ -39,7 +44,7 @@ Change History +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | 2023-11-01 | Modified the following content: | | | | -| | - Modified the link for obtaining service support during quota application in :ref:`Quotas `. | +| | - Modified the link for obtaining service support during quota application in :ref:`Quota Management `. | | | - Added the link to *Data Lake Insight API Reference* to :ref:`What Is Data Lake Insight `. | | | - Modified the method of obtaining host information in :ref:`Modifying Host Information in an Elastic Resource Pool `. | +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ @@ -55,9 +60,9 @@ Change History | | - :ref:`Enabling Dynamic Scaling for Flink Jobs ` | | | - :ref:`Setting the Priority for a SQL Job ` | | | | -| | Deleted the following section: | +| | Taken offline the following content: | | | | -| | - Deleted the content related to Flink job debugging. | +| | - Taken offline the relevant content for debugging Flink jobs. | | | | | | Modified the following section: | | | | diff --git a/umn/source/common_dli_management_operations/enhancing_the_job_runtime_environment_using_a_custom_image.rst b/umn/source/common_dli_management_operations/enhancing_the_job_runtime_environment_using_a_custom_image.rst index eb0e423..d50938b 100644 --- a/umn/source/common_dli_management_operations/enhancing_the_job_runtime_environment_using_a_custom_image.rst +++ b/umn/source/common_dli_management_operations/enhancing_the_job_runtime_environment_using_a_custom_image.rst @@ -34,7 +34,7 @@ Use Process #. Obtain DLI base images. #. Use Dockerfile to pack dependencies (files, JAR files, or software) required for job execution into the base image to create a custom image. -#. Publish the custom image to SoftWare Repository for Container (SWR). +#. Publish the custom image to SWR. #. On the DLI job editing page, select the created image and run the job. #. Check the job execution status. @@ -47,7 +47,7 @@ Contact the administrator to obtain DLI base images. Select the base image of the same type as the architecture of the queue. -For the CPU architecture type of a queue, see :ref:`Viewing Basic Information About a Queue `. +For details about the CPU architecture type of a queue, see :ref:`Viewing Basic Information About a Queue `. Creating a Custom Image ----------------------- diff --git a/umn/source/common_dli_management_operations/managing_dli_resource_quotas.rst b/umn/source/common_dli_management_operations/managing_dli_resource_quotas.rst index f831d18..52a4ed5 100644 --- a/umn/source/common_dli_management_operations/managing_dli_resource_quotas.rst +++ b/umn/source/common_dli_management_operations/managing_dli_resource_quotas.rst @@ -8,9 +8,9 @@ Managing DLI Resource Quotas What Is a Quota? ---------------- -A quota limits the quantity of a resource available to users, thereby preventing spikes in the usage of the resource. +Quotas are enforced for service resources on the platform to prevent unforeseen spikes in resource usage. Quotas can limit the quantity and capacity of resources available to users. -You can also request for an increased quota if your existing quota cannot meet your service requirements. +If your current quota does not meet your needs, you can apply for an increase. How Do I View My Quotas? ------------------------ @@ -21,18 +21,18 @@ How Do I View My Quotas? #. Click the **My Quota** icon |image2| in the upper right corner of the page. - The **Service Quota** page is displayed. + This will take you to the **Service Quota** page. -#. View the used and total quota of each type of resources on the displayed page. +#. Here, you can view the total quota and usage details for each resource. - If a quota cannot meet service requirements, increase a quota. + If your current quota is insufficient, follow the steps below to request an increase. -How Do I Apply for a Higher Quota? ----------------------------------- +How Do I Apply for a Quota Increase? +------------------------------------ -The system does not support online quota adjustment. To increase a resource quota, dial the hotline or send an email to the customer service. We will process your application and inform you of the progress by phone call or email. +The system currently does not support online quota adjustments. If you need to modify your quota, please contact our customer service team via phone or email. They will promptly process your request and keep you updated on its progress through either a call or email. -Before dialing the hotline number or sending an email, ensure that the following information has been obtained: +Before reaching out, ensure you have the following details ready: - Domain name, project name, and project ID @@ -42,7 +42,7 @@ Before dialing the hotline number or sending an email, ensure that the following - Service name - Quota type - - Required quota + - Desired quota value `Learn how to obtain the service hotline and email address. `__ diff --git a/umn/source/common_dli_management_operations/managing_program_packages_of_jar_jobs/dli_built-in_dependencies.rst b/umn/source/common_dli_management_operations/managing_program_packages_of_jar_jobs/dli_built-in_dependencies.rst index d6c4e8d..27f0707 100644 --- a/umn/source/common_dli_management_operations/managing_program_packages_of_jar_jobs/dli_built-in_dependencies.rst +++ b/umn/source/common_dli_management_operations/managing_program_packages_of_jar_jobs/dli_built-in_dependencies.rst @@ -318,7 +318,7 @@ Spark 3.1.1 Dependencies .. table:: **Table 2** Spark 3.1.1 dependencies +-----------------------------------------------------------------+--------------------------------------------------------------------------+---------------------------------------------------------------+ - | Dependency | _ | _ | + | Dependency | | | +=================================================================+==========================================================================+===============================================================+ | accessors-smart-1.2.jar | hive-shims-scheduler-3.1.0-h0.cbu.mrs.321.r10.jar | metrics-graphite-4.1.1.jar | +-----------------------------------------------------------------+--------------------------------------------------------------------------+---------------------------------------------------------------+ @@ -631,7 +631,7 @@ Spark 2.4.5 Dependencies .. table:: **Table 3** Spark 2.4.5 dependencies +------------------------------------------------------------------+-----------------------------------------------------------+------------------------------------------------------------------------+ - | Dependency | _ | _ | + | Dependency | | | +==================================================================+===========================================================+========================================================================+ | JavaEWAH-1.1.7.jar | httpclient-4.5.6.jar | lucene-queryparser-7.7.2.jar | +------------------------------------------------------------------+-----------------------------------------------------------+------------------------------------------------------------------------+ @@ -860,7 +860,7 @@ Spark 2.3.2 Dependencies .. table:: **Table 4** Spark 2.3.2 dependencies +-------------------------------------------------------+----------------------------------------------------------------+-------------------------------------------------------------------------+ - | Dependency | _ | _ | + | Dependency | | | +=======================================================+================================================================+=========================================================================+ | accessors-smart-1.2.jar | HikariCP-java7-2.4.12.jar | logging-interceptor-3.14.4.jar | +-------------------------------------------------------+----------------------------------------------------------------+-------------------------------------------------------------------------+ @@ -1082,7 +1082,7 @@ Flink 1.12 Dependencies .. table:: **Table 5** Flink 1.12 dependencies +----------------------------------------------------------+-----------------------------------------------------------------------+-----------------------------------------------------------+ - | Dependency | _ | _ | + | Dependency | | | +==========================================================+=======================================================================+===========================================================+ | bcpkix-jdk15on-1.60.jar | flink-json-1.12.2-ei-313001-dli-2022011002.jar | libtensorflow-1.12.0.jar | +----------------------------------------------------------+-----------------------------------------------------------------------+-----------------------------------------------------------+ @@ -1135,7 +1135,7 @@ Only queues created after December 2020 can use the Flink 1.10 dependencies. .. table:: **Table 6** Flink 1.10 dependencies +-------------------------------+-----------------------------------------------+--------------------------------+ - | Dependency | _ | _ | + | Dependency | | | +===============================+===============================================+================================+ | bcpkix-jdk15on-1.60.jar | esdk-obs-java-3.20.6.1.jar | java-xmlbuilder-1.1.jar | +-------------------------------+-----------------------------------------------+--------------------------------+ @@ -1176,7 +1176,7 @@ Flink 1.7.2 Dependencies .. table:: **Table 7** Flink 1.7.2 dependencies +-------------------------------+----------------------------------------------+--------------------------------+ - | Dependency | _ | _ | + | Dependency | | | +===============================+==============================================+================================+ | bcpkix-jdk15on-1.60.jar | esdk-obs-java-3.1.3.jar | httpcore-4.4.4.jar | +-------------------------------+----------------------------------------------+--------------------------------+ diff --git a/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/agency_permission_policies_in_common_scenarios.rst b/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/agency_permission_policies_in_common_scenarios.rst index 529ca1b..4751ba6 100644 --- a/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/agency_permission_policies_in_common_scenarios.rst +++ b/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/agency_permission_policies_in_common_scenarios.rst @@ -50,6 +50,7 @@ Application scenario: Data cleanup agency, which is used to clean up data accord "dli:table:showPartitions", "dli:table:select", "dli:table:dropTable", + "dli:table:alter", "dli:table:alterTableDropPartition" ] } @@ -61,7 +62,7 @@ Application scenario: Data cleanup agency, which is used to clean up data accord Permission Policies for Accessing and Using OBS ----------------------------------------------- -Application scenario: For DLI Flink jobs, the permissions include downloading OBS objects, obtaining OBS/GaussDB(DWS) data sources (foreign tables), transferring logs, using savepoints, and enabling checkpointing. For DLI Spark jobs, the permissions allow downloading OBS objects and reading/writing OBS foreign tables. +Application scenario: For DLI Flink jobs, the permissions include downloading OBS objects, obtaining OBS/DWS data sources (foreign tables), transferring logs, using savepoints, and enabling checkpointing. For DLI Spark jobs, the permissions allow downloading OBS objects and reading/writing OBS foreign tables. .. code-block:: @@ -180,6 +181,7 @@ Application scenario: DLI Flink and Spark jobs are authorized to access DLI meta "dli:table:alterTableRecoverPartition", "dli:table:dropTable", "dli:table:update", + "dli:table:alter", "dli:table:alterTableDropPartition" ] } diff --git a/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/creating_a_custom_dli_agency.rst b/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/creating_a_custom_dli_agency.rst index be6a11a..cb8aa5d 100644 --- a/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/creating_a_custom_dli_agency.rst +++ b/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/creating_a_custom_dli_agency.rst @@ -14,19 +14,19 @@ DLI Custom Agency Scenarios .. table:: **Table 1** DLI custom agency scenarios - +----------------------------------------------------------------------+-----------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ - | Scenario | Agency Name | Description | Permission Policy | - +======================================================================+=======================+===========================================================================================================================================================================================================================================================================================================+==========================================================================================+ - | Allowing DLI to clear data according to the lifecycle of a table | dli_data_clean_agency | Data cleanup agency, which is used to clean up data according to the lifecycle of a table and clean up lakehouse table data. | :ref:`Data Cleanup Agency Permission Configuration ` | - | | | | | - | | | You need to create an agency and customize permissions for it. However, the agency name is fixed to **dli_data_clean_agency**. | | - +----------------------------------------------------------------------+-----------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ - | Allowing DLI to read and write data from and to OBS to transfer logs | Custom | For DLI Flink jobs, the permissions include downloading OBS objects, obtaining OBS/GaussDB(DWS) data sources (foreign tables), transferring logs, using savepoints, and enabling checkpointing. For DLI Spark jobs, the permissions allow downloading OBS objects and reading/writing OBS foreign tables. | :ref:`Permission Policies for Accessing and Using OBS ` | - +----------------------------------------------------------------------+-----------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ - | Allowing DLI to obtain data access credentials by accessing DEW | Custom | DLI jobs use DEW-CSMS' secret management. | :ref:`Permission to Use DEW's Encryption Function ` | - +----------------------------------------------------------------------+-----------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ - | Allowing DLI to access DLI catalogs to retrieve metadata | Custom | DLI accesses catalogs to retrieve metadata. | :ref:`Permission to Access DLI Catalog Metadata ` | - +----------------------------------------------------------------------+-----------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ + +----------------------------------------------------------------------+-----------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ + | Scenario | Agency Name | Description | Permission Policy | + +======================================================================+=======================+==================================================================================================================================================================================================================================================================================================+==========================================================================================+ + | Allowing DLI to clear data according to the lifecycle of a table | dli_data_clean_agency | Data cleanup agency, which is used to clean up data according to the lifecycle of a table and clean up lakehouse table data. | :ref:`Data Cleanup Agency Permission Configuration ` | + | | | | | + | | | You need to create an agency and customize permissions for it. However, the agency name is fixed to **dli_data_clean_agency**. | | + +----------------------------------------------------------------------+-----------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ + | Allowing DLI to read and write data from and to OBS to transfer logs | Custom | For DLI Flink jobs, the permissions include downloading OBS objects, obtaining OBS/DWS data sources (foreign tables), transferring logs, using savepoints, and enabling checkpointing. For DLI Spark jobs, the permissions allow downloading OBS objects and reading/writing OBS foreign tables. | :ref:`Permission Policies for Accessing and Using OBS ` | + +----------------------------------------------------------------------+-----------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ + | Allowing DLI to obtain data access credentials by accessing DEW | Custom | DLI jobs use DEW-CSMS' secret management. | :ref:`Permission to Use DEW's Encryption Function ` | + +----------------------------------------------------------------------+-----------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ + | Allowing DLI to access DLI catalogs to retrieve metadata | Custom | DLI accesses catalogs to retrieve metadata. | :ref:`Permission to Access DLI Catalog Metadata ` | + +----------------------------------------------------------------------+-----------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------+ Procedure --------- @@ -81,7 +81,7 @@ Step 1: Create a Cloud Service Agency on the IAM Console and Grant Permissions c. In the **Policy Content** area, paste a custom policy. - In this example, the permissions allow access and usage of OBS in various scenarios. For DLI Flink jobs, this includes downloading OBS objects, obtaining OBS/GaussDB(DWS) data sources (foreign tables), transferring logs, using savepoints, and enabling checkpointing. For DLI Spark jobs, the permissions allow downloading OBS objects and reading/writing OBS foreign tables. + In this example, the permissions allow access and usage of OBS in various scenarios. For DLI Flink jobs, this includes downloading OBS objects, obtaining OBS/DWS data sources (foreign tables), transferring logs, using savepoints, and enabling checkpointing. For DLI Spark jobs, the permissions allow downloading OBS objects and reading/writing OBS foreign tables. For how to configure common agency permissions for Flink jobs, see :ref:`Agency Permission Policies in Common Scenarios `. @@ -148,7 +148,7 @@ Step 2: Set Agency Permissions for a Job When Spark 3.3.1, Flink 1.15, or a later version is used to execute jobs, you need to add information about the new agency to the job configuration. -Otherwise, If you do not specify an agency for Spark 3.3.1 jobs, the jobs cannot use OBS. If you do not specify an agency for a Flink 1.15 job, checkpointing cannot be enabled, savepoints cannot be used, logs cannot be transferred, and data sources such as OBS and GaussDB(DWS) cannot be used. +Otherwise, if you do not specify an agency for Spark 3.3.1 jobs, the jobs cannot use OBS. If you do not specify an agency for a Flink 1.15 job, checkpointing cannot be enabled, savepoints cannot be used, logs cannot be transferred, and data sources such as OBS and DWS cannot be used. .. caution:: diff --git a/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/dli_agency_overview.rst b/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/dli_agency_overview.rst index af67087..afbb25e 100644 --- a/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/dli_agency_overview.rst +++ b/umn/source/configuring_an_agency_to_allow_dli_to_access_other_cloud_services/dli_agency_overview.rst @@ -8,9 +8,9 @@ DLI Agency Overview What Is an Agency? ------------------ -Cloud services often interact with each other, with some of which dependent on other services. You can create an agency to delegate DLI to use other cloud services and perform resource O&M on your behalf. +Cloud services often need to interact with each other for business operations. In some cases, one cloud service may require collaboration with another. To facilitate this, you can create a cloud service agency, which grants operational permissions to DLI. This allows DLI to act on your behalf when using other cloud services, enabling it to perform resource management tasks for you. -For example, the AK/SK required by DLI Flink jobs is stored in DEW. To allow DLI to access DEW data during job execution, you need to provide an IAM agency to delegate the permissions to perform operations on DEW data to DLI. +For example, when creating a Flink job in DLI, the required AK/SK is stored in DEW. If you want DLI to access DEW data during job execution, you must provide an IAM agency. This grants the necessary permissions to DLI, allowing it to access DEW on your behalf. .. figure:: /_static/images/en-us_image_0000001742695104.png @@ -24,20 +24,7 @@ DLI Agencies Before using DLI, you are advised to set up DLI agency permissions to ensure the proper functioning of DLI. - By default, DLI provides the following agencies: **dli_admin_agency**, **dli_management_agency**, and **dli_data_clean_agency**. The names of these agencies are fixed, but the permissions contained in them can be customized. In other scenarios, you need to create custom agencies. For details about the agencies, see :ref:`Table 1 `. - - DLI upgrades **dli_admin_agency** to **dli_management_agency** to meet the demand for fine-grained agency permissions management. The new agency has the necessary permissions for datasource operations, notifications, and user authorization operations. For details, see :ref:`Configuring DLI Agency Permissions `. - -- To use Flink 1.15, Spark 3.3.1 (Spark general queue scenario), or a later version to execute jobs, perform the following operations: - - Create an agency on the IAM console and add the agency information to the job configuration. For details, see :ref:`Creating a Custom DLI Agency `. - - - Common scenarios for creating an agency: DLI is allowed to read and write data from and to OBS, dump logs, and read and write Flink checkpoints. DLI is allowed to access DEW to obtain data access credentials and access catalogs to obtain metadata. - - You cannot use the default agency names **dli_admin_agency**, **dli_management_agency**, or **dli_data_clean_agency**. It must be unique. - -- If the engine version is earlier than Flink 1.15, **dli_admin_agency** is used by default during job execution. If the engine version is earlier than Spark 3.3.1, user authentication information (AK/SK and security token) is used during job execution. - - This means that jobs whose engine versions are earlier than Flink 1.15 or Spark 3.3.1 are not affected by the update of agency permissions and do not require custom agencies. - - To maintain compatibility with existing job agency permission requirements, **dli_admin_agency** will still be listed in the IAM agency list even after the update. .. note:: @@ -86,3 +73,17 @@ Before using DLI, you are advised to set up DLI agency permissions to ensure the +------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | DLI Notification Agency Access | Permissions to send notifications through SMN when a job fails to be executed | +------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +Notes and Constraints +--------------------- + +- To use Flink 1.15, Spark 3.3.1 (Spark general queue scenario), or a later version to execute jobs, perform the following operations: + + Create an agency on the IAM console and add the agency information to the job configuration. For details, see :ref:`Creating a Custom DLI Agency `. + + - Common scenarios for creating an agency: DLI is allowed to read and write data from and to OBS, dump logs, and read and write Flink checkpoints. DLI is allowed to access DEW to obtain data access credentials and access catalogs to obtain metadata. + - You cannot use the default agency names **dli_admin_agency**, **dli_management_agency**, or **dli_data_clean_agency**. It must be unique. + +- If the engine version is earlier than Flink 1.15, **dli_admin_agency** is used by default during job execution. If the engine version is earlier than Spark 3.3.1, user authentication information (AK/SK and security token) is used during job execution. + + This means that jobs whose engine versions are earlier than Flink 1.15 or Spark 3.3.1 are not affected by the update of agency permissions and do not require custom agencies. diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/common_development_methods_for_dli_cross-source_analysis.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/common_development_methods_for_dli_cross-source_analysis.rst index 458abd0..60f39bf 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/common_development_methods_for_dli_cross-source_analysis.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/common_development_methods_for_dli_cross-source_analysis.rst @@ -21,7 +21,7 @@ Notes DLI Supported Data Sources -------------------------- -:ref:`Table 1 ` lists the data sources supported by DLI. For how to use the data sources, see *Data Lake Insight SQL Syntax Reference*. +:ref:`Table 1 ` lists the data sources supported by DLI. For details about how to use the data sources, see *Data Lake Insight SQL Syntax Reference*. .. _dli_01_0410__table1771918377534: @@ -35,7 +35,7 @@ DLI Supported Data Sources DCS Redis Y Y Y Y DDS Y Y Y Y DMS for Kafka x x Y Y - GaussDB(DWS) Y Y Y Y + DWS Y Y Y Y MRS HBase Y Y Y Y MRS Kafka x x Y Y MRS OpenTSDB Y Y x Y diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/creating_an_enhanced_datasource_connection.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/creating_an_enhanced_datasource_connection.rst index bd6ee80..23b516c 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/creating_an_enhanced_datasource_connection.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/configuring_the_network_connection_between_dli_and_data_sources_enhanced_datasource_connection/creating_an_enhanced_datasource_connection.rst @@ -8,11 +8,11 @@ Creating an Enhanced Datasource Connection Scenario -------- -Create an enhanced datasource connection for DLI to access, import, query, and analyze data of other data sources. +Before using DLI to access data from other sources, you need to create an enhanced datasource connection to enable network communication between DLI and the destination data source. This allows DLI to seamlessly access, import, query, and analyze data stored in external systems. -For example, to connect DLI to the MRS, RDS, CSS, Kafka, or GaussDB(DWS) data source, you need to enable the network between DLI and the VPC of the data source. +For example, when connecting DLI to services such as MRS, RDS, CSS, Kafka, or DWS, you must first configure the network connectivity between DLI and the corresponding VPC of these data sources to facilitate smooth data exchange. -Create an enhanced datasource connection on the console. +This section provides a step-by-step guide on how to create an enhanced datasource connection via the management console. Notes and Constraints --------------------- @@ -153,6 +153,8 @@ Step 1: Create an Enhanced Datasource Connection | VPC | VPC used by the data source. | +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Subnet | Subnet used by the data source. | + | | | + | | If the subnet of the selected data source has IPv6 enabled, the enhanced datasource connection you create will also support IPv6. For more information on using IPv6 for cross-source access, see :ref:`How Do I Configure a Network Connection with IPv6 Address Enabled? `. | +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Host Information | In this text field, you can configure the mapping between host IP addresses and domain names so that jobs can only use the configured domain names to access corresponding hosts. This parameter is optional. | | | | @@ -221,18 +223,31 @@ After an enhanced datasource connection is created, the subnet is automatically .. table:: **Table 4** Parameters for adding a custom route - +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parameter | Description | - +===================================+==========================================================================================================================================================================================================+ - | Route Name | Name of a custom route, which is unique in the same enhanced datasource connection. The name can contain up to 64 characters. Only digits, letters, underscores (_), and hyphens (-) are allowed. | - +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | IP Address | Custom route CIDR block. The CIDR blocks of different routes can overlap but cannot be identical. | - | | | - | | Do not add the **100.125.**\ *xx.xx* or **100.64.**\ *xx.xx* CIDR blocks to avoid conflicts with the internal CIDR blocks of services like SWR, which can cause enhanced datasource connections to fail. | - +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+===================================================================================================================================================================================================================+ + | Route Name | Name of a custom route, which is unique in the same enhanced datasource connection. The name can contain up to 64 characters. Only digits, letters, underscores (_), and hyphens (-) are allowed. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | IP Address Type | The options are **IPv4** and **IPv6**. | + | | | + | | If your data source has IPv6 enabled and the current enhanced datasource connection supports IPv6, you can select IPv6 routes when adding a route table. | + | | | + | | You can check whether the current enhanced datasource connection supports IPv6 in its basic information. For details, see :ref:`Viewing Basic Information About an Enhanced Datasource Connection `. | + | | | + | | The route IP address example is as follows: | + | | | + | | - IPv4 address: **192.168.2.0/24**. | + | | - IPv6 address: **2407:c080:802:be7::/64**. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | IP Address | Custom route CIDR block. The CIDR blocks of different routes can overlap but cannot be identical. | + | | | + | | Do not add the **100.125.**\ *xx.xx* or **100.64.**\ *xx.xx* CIDR blocks to avoid conflicts with the internal CIDR blocks of services like SWR, which can cause enhanced datasource connections to fail. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ #. After adding a route, you can view the route information on the route details page. +.. _dli_01_0006__section17945175313448: + Step 3: Test the Connectivity Between the Queue in the Elastic Resource Pool and the Data Source Address -------------------------------------------------------------------------------------------------------- @@ -248,10 +263,26 @@ Step 3: Test the Connectivity Between the Queue in the Elastic Resource Pool and - IPv4 + Port number: 192.168.x.x:8080 - Domain name: domain-xxxxxx.com - Domain name + Port number: domain-xxxxxx.com:8080 + - IPv6 address: 2001:0db8:XXXX:XXXX:XXXX:XXXX:XXXX:XXXX + - [IPv6] + Port number: [2001:0db8:XXXX:XXXX:XXXX:XXXX:XXXX:XXXX]:8080 #. Click **Test**. - If the test address is reachable, you will receive a message. - If the test address is unreachable, you will also receive a message. Check the network configurations and retry. Network configurations include the VPC peering and the datasource connection. Check whether they have been activated. +.. _dli_01_0006__section463114164718: + +How Do I Configure a Network Connection with IPv6 Address Enabled? +------------------------------------------------------------------ + +DLI resource networks support IPv4/IPv6 dual stack. When creating an enhanced datasource connection, you can choose to use an IPv6 address for communication to enhance network compatibility and security. + +A prerequisite for using IPv6 in datasource scenarios is that both the DLI elastic resource pool and the data source must have IPv6 enabled. + +- Elastic resource pool: Enable IPv6 when creating the resource pool. For details, see :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. +- Data source: The subnet where the data source is must have IPv6 enabled. Otherwise, the IPv6 network connection cannot be established. + +To verify if IPv6 communication is successful, use an IPv6 address to test the network connectivity between the queue and the data source by referring to :ref:`Step 3: Test the Connectivity Between the Queue in the Elastic Resource Pool and the Data Source Address `. + .. |image1| image:: /_static/images/en-us_image_0000002363080342.png diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/flink_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/flink_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs.rst new file mode 100644 index 0000000..908db3b --- /dev/null +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/flink_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs.rst @@ -0,0 +1,261 @@ +:original_name: dli_09_0211.html + +.. _dli_09_0211: + +Flink Jar Jobs Using DEW to Acquire Access Credentials for Reading and Writing Data from and to OBS +=================================================================================================== + +Scenario +-------- + +When writing output data from Flink Jar jobs to OBS, you need to configure an AK/SK for accessing OBS. To ensure the security of AK/SK data, you can use DEW and CSMS for centralized management of AK/SK. This approach effectively mitigates risks such as sensitive information leakage caused by hardcoding in programs or plaintext configurations, as well as potential business disruptions due to unauthorized access. + +This section walks you through on how a Flink Jar job acquires an AK/SK to read and write data from and to OBS. + +Notes and Constraints +--------------------- + +DEW can be used to manage access credentials only in Flink 1.15. When creating a Flink job, select version 1.15 and configure the information of the agency that allows DLI to access DEW for the job. + +Prerequisites +------------- + +- A shared secret has been created on the DEW console and the secret value has been stored. +- An agency has been created and authorized for DLI to access DEW. The agency must have been granted the following permissions: + + - Permission of the **ShowSecretVersion** interface for querying secret versions and secret values in DEW: **csms:secretVersion:get**. + - Permission of the **ListSecretVersions** interface for listing secret versions in DEW: **csms:secretVersion:list**. + - Permission to decrypt DEW secrets: **kms:dek:decrypt** + +- To use this function, you need to configure AK/SK for all OBS buckets. + +Syntax +------ + +On the Flink Jar job editing page, set **Runtime Configuration** as needed. The configuration information is as follows: + +Different OBS buckets use different AK/SK authentication information. You can use the following configuration method to specify the AK/SK information based on the bucket. For details about the parameters, see :ref:`Table 1 `. + +:: + + flink.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.access.key=USER_AK_CSMS_KEY + flink.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.secret.key=USER_SK_CSMS_KEY + flink.hadoop.fs.obs.security.provider=com.dli.provider.UserObsBasicCredentialProvider + flink.hadoop.fs.dew.csms.secretName=CredentialName + flink.hadoop.fs.dew.endpoint=ENDPOINT + flink.hadoop.fs.dew.csms.version=VERSION_ID + flink.hadoop.fs.dew.csms.cache.time.second=CACHE_TIME + flink.dli.job.agency.name=USER_AGENCY_NAME + +Parameter Description +--------------------- + +.. _dli_09_0211__en-us_topic_0000001836634008_en-us_topic_0000001794701748_table517231215112: + +.. table:: **Table 1** Parameter descriptions + + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Mandatory | Default Value | Data Type | Description | + +============================================================+=============+================+=============+=================================================================================================================================================================================================================================================================================================================+ + | flink.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.access.key | Yes | None | String | *USER_BUCKET_NAME* needs to be replaced with the user's OBS bucket name. | + | | | | | | + | | | | | The value of this parameter is the key defined by the user in the CSMS shared secret. The value corresponding to the key is the user's access key ID (AK). The user must have the permission to access the bucket on OBS. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.secret.key | Yes | None | String | *USER_BUCKET_NAME* needs to be replaced with the user's OBS bucket name. | + | | | | | | + | | | | | The value of this parameter is the key defined by the user in the CSMS shared secret. The value corresponding to the key is the user's secret access key (SK). The user must have the permission to access the bucket on OBS. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.hadoop.fs.obs.security.provider | Yes | None | String | OBS AK/SK authentication mechanism, which uses DEW-CSMS' secret management to obtain the AK and SK for accessing OBS. | + | | | | | | + | | | | | The default value is **com.dli.provider.UserObsBasicCredentialProvider**. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.hadoop.fs.dew.endpoint | Yes | None | String | Endpoint of the DEW service to be used. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.hadoop.fs.dew.projectId | No | Yes | String | ID of the project DEW belongs to. The default value is the ID of the project where the Flink job is. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.hadoop.fs.dew.csms.secretName | Yes | None | String | Name of the shared secret in DEW's secret management. | + | | | | | | + | | | | | Configuration example: **flink.hadoop.fs.dew.csms.secretName=**\ *secretInfo* | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.hadoop.fs.dew.csms.version | Yes | Latest version | String | Version number (secret version identifier) of the shared secret created in DEW CSMS. | + | | | | | | + | | | | | If the latest version number (secret version identifier) is not specified, the system will not be able to directly retrieve the most recent version of the secret, potentially leading to application access failures or the use of outdated secrets, thereby compromising data security and service stability. | + | | | | | | + | | | | | View the version information of the secret on the DEW management console and configure this parameter with the latest version number to ensure that applications can securely and reliably access the necessary data. | + | | | | | | + | | | | | Configuration example: **flink.hadoop.fs.dew.csms.version=v1** | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.hadoop.fs.dew.csms.cache.time.second | No | 3600 | Long | Cache duration after the CSMS shared secret is obtained during Flink job access. | + | | | | | | + | | | | | The unit is second. The default value is 3600 seconds. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | flink.dli.job.agency.name | Yes | ``-`` | String | Custom agency name. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +Sample Code +----------- + +This section describes how to write processed DataGen data to OBS. You need to modify the parameters in the sample Java code based on site requirements. + +#. Create an agency for DLI to access DEW and complete authorization. +#. Create a shared secret in DEW. + + a. Log in to the DEW management console. + b. In the navigation pane on the left, choose **Cloud Secret Management Service** > **Secrets**. + c. On the displayed page, click **Create Secret**. Set basic secret information. + +#. Set job parameters on the DLI Flink Jar job editing page. + + - Class name + + :: + + com.dli.demo.dew.DataGen2FileSystemSink + + - Parameters + + :: + + --checkpoint.path obs://test/flink/jobs/checkpoint/120891/ + --output.path obs://dli/flink.db/79914/DataGen2FileSystemSink + + - Runtime configuration: + + :: + + flink.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.access.key=USER_AK_CSMS_KEY + flink.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.secret.key=USER_SK_CSMS_KEY + flink.hadoop.fs.obs.security.provider=com.dli.provider.UserObsBasicCredentialProvider + flink.hadoop.fs.dew.csms.secretName=obsAksK + flink.hadoop.fs.dew.endpoint=kmsendpoint + flink.hadoop.fs.dew.csms.version=v6 + flink.hadoop.fs.dew.csms.cache.time.second=3600 + flink.dli.job.agency.name=*** + +#. Flink Jar job example. + + - **Environment preparation** + + Development tools such as IntelliJ IDEA and other development tools, JDK, and Maven have been installed and configured. + + Dependency package in POM file configurations + + :: + + + 1.15.0 + + + + + org.apache.flink + flink-statebackend-rocksdb + ${flink.version} + provided + + + + org.apache.flink + flink-streaming-java + ${flink.version} + provided + + + + + fastjson + 2.0.15 + + + + - **Example code** + + :: + + import org.apache.flink.api.common.serialization.SimpleStringEncoder; + import org.apache.flink.api.java.utils.ParameterTool; + import org.apache.flink.contrib.streaming.state.EmbeddedRocksDBStateBackend; + import org.apache.flink.core.fs.Path; + import org.apache.flink.streaming.api.datastream.DataStream; + import org.apache.flink.streaming.api.environment.CheckpointConfig; + import org.apache.flink.streaming.api.environment.StreamExecutionEnvironment; + import org.apache.flink.streaming.api.functions.sink.filesystem.StreamingFileSink; + import org.apache.flink.streaming.api.functions.sink.filesystem.rollingpolicies.OnCheckpointRollingPolicy; + import org.apache.flink.streaming.api.functions.source.ParallelSourceFunction; + import org.slf4j.Logger; + import org.slf4j.LoggerFactory; + + import java.time.LocalDateTime; + import java.time.ZoneOffset; + import java.time.format.DateTimeFormatter; + import java.util.Random; + + public class DataGen2FileSystemSink { + private static final Logger LOG = LoggerFactory.getLogger(DataGen2FileSystemSink.class); + + public static void main(String[] args) { + ParameterTool params = ParameterTool.fromArgs(args); + LOG.info("Params: " + params.toString()); + try { + StreamExecutionEnvironment streamEnv = StreamExecutionEnvironment.getExecutionEnvironment(); + + // set checkpoint + String checkpointPath = params.get("checkpoint.path", "obs://bucket/checkpoint/jobId_jobName/"); + LocalDateTime localDateTime = LocalDateTime.ofEpochSecond(System.currentTimeMillis() / 1000, + 0, ZoneOffset.ofHours(8)); + String dt = localDateTime.format(DateTimeFormatter.ofPattern("yyyyMMdd_HH:mm:ss")); + checkpointPath = checkpointPath + dt; + + streamEnv.setStateBackend(new EmbeddedRocksDBStateBackend()); + streamEnv.getCheckpointConfig().setCheckpointStorage(checkpointPath); + streamEnv.getCheckpointConfig().setExternalizedCheckpointCleanup( + CheckpointConfig.ExternalizedCheckpointCleanup.RETAIN_ON_CANCELLATION); + streamEnv.enableCheckpointing(30 * 1000); + + DataStream stream = streamEnv.addSource(new DataGen()) + .setParallelism(1) + .disableChaining(); + + String outputPath = params.get("output.path", "obs://bucket/outputPath/jobId_jobName"); + + // Sink OBS + final StreamingFileSink sinkForRow = StreamingFileSink + .forRowFormat(new Path(outputPath), new SimpleStringEncoder("UTF-8")) + .withRollingPolicy(OnCheckpointRollingPolicy.build()) + .build(); + + stream.addSink(sinkForRow); + + streamEnv.execute("sinkForRow"); + } catch (Exception e) { + LOG.error(e.getMessage(), e); + } + } + } + + class DataGen implements ParallelSourceFunction { + + private boolean isRunning = true; + + private Random random = new Random(); + + @Override + public void run(SourceContext ctx) throws Exception { + while (isRunning) { + JSONObject jsonObject = new JSONObject(); + jsonObject.put("id", random.nextLong()); + jsonObject.put("name", "Molly" + random.nextInt()); + jsonObject.put("address", "hangzhou" + random.nextInt()); + jsonObject.put("birthday", System.currentTimeMillis()); + jsonObject.put("city", "hangzhou" + random.nextInt()); + jsonObject.put("number", random.nextInt()); + ctx.collect(jsonObject.toJSONString()); + Thread.sleep(1000); + } + } + + @Override + public void cancel() { + isRunning = false; + } + } diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/flink_opensource_sql_jobs_using_dew_to_manage_access_credentials.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/flink_opensource_sql_jobs_using_dew_to_manage_access_credentials.rst new file mode 100644 index 0000000..fd367df --- /dev/null +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/flink_opensource_sql_jobs_using_dew_to_manage_access_credentials.rst @@ -0,0 +1,142 @@ +:original_name: dli_09_0210.html + +.. _dli_09_0210: + +Flink OpenSource SQL Jobs Using DEW to Manage Access Credentials +================================================================ + +Scenario +-------- + +When DLI writes the output data of Flink jobs to MySQL or DWS, you need to set sensitive parameters such as the username and password in the connector. Storing this information in plaintext poses significant security risks. To safeguard user data privacy, you are advised to encrypt these credentials. + +Data Encryption Workshop (DEW) and Cloud Secret Management Service (CSMS) offer a secure, reliable, and easy-to-use solution for encrypting and decrypting sensitive data. By hosting database account details (for example, username and password) as managed secrets within CSMS, you can securely reference these credentials in your Flink jobs. This ensures that sensitive information is retrieved through a secure channel during runtime. + +Additionally, CSMS offers comprehensive lifecycle management for credentials, enhancing both security and efficiency. It effectively mitigates risks associated with hardcoding sensitive information or storing it in plaintext configurations, thereby preventing unauthorized access and potential business disruptions. + +This section walks you through on how to use DEW to manage access credentials for Flink OpenSource SQL jobs. + +Notes and Constraints +--------------------- + +DEW can be used to manage access credentials only in Flink 1.15. When creating a Flink job, select version 1.15 and configure the information of the agency that allows DLI to access DEW for the job. + +Prerequisites +------------- + +- A shared secret has been created on the DEW console and the secret value has been stored. + +- An agency has been created and authorized for DLI to access DEW. The agency must have been granted the following permissions: + + - Permission of the **ShowSecretVersion** interface for querying secret versions and secret values in DEW: **csms:secretVersion:get**. + - Permission of the **ListSecretVersions** interface for listing secret versions in DEW: **csms:secretVersion:list**. + - Permission to decrypt DEW secrets: **kms:dek:decrypt** + +- On the DLI management console, create an enhanced datasource connection and configure the network connection between DLI and the data source. + +Syntax +------ + +.. code-block:: + + create table tableName( + attr_name attr_type + (',' attr_name attr_type)* + (',' WATERMARK FOR rowtime_column_name AS watermark-strategy_expression) + ) + with ( + ... + 'dew.endpoint'='', + 'dew.csms.secretName'='', + 'dew.csms.decrypt.fields'='', + 'dew.projectId'='', + 'dew.csms.version'='' + + ); + +Parameter Description +--------------------- + +.. table:: **Table 1** Parameter descriptions + + +-------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Mandatory | Default Value | Data Type | Description | + +=========================+=============+================+=============+=================================================================================================================================================================================================================================================================================================================+ + | dew.endpoint | Yes | None | String | Endpoint of the DEW service to be used. | + +-------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | dew.projectId | No | Yes | String | ID of the project DEW belongs to. The default value is the ID of the project where the Flink job is. | + +-------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | dew.csms.secretName | Yes | None | String | Name of the shared secret in DEW's secret management. | + | | | | | | + | | | | | Configuration example: **'dew.csms.secretName'='**\ *secretInfo*\ **'** | + +-------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | dew.csms.decrypt.fields | Yes | None | String | Specify which fields in the **connector with** attribute need to be decrypted using DEW's CSMS. | + | | | | | | + | | | | | Separate the field attributes with commas, for example, **'dew.csms.decrypt.fields'='field1,field2,field3'** | + +-------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | dew.csms.version | Yes | Latest version | String | Version number (secret version identifier) of the shared secret created in DEW CSMS. | + | | | | | | + | | | | | If the latest version number (secret version identifier) is not specified, the system will not be able to directly retrieve the most recent version of the secret, potentially leading to application access failures or the use of outdated secrets, thereby compromising data security and service stability. | + | | | | | | + | | | | | View the version information of the secret on the DEW management console and configure this parameter with the latest version number to ensure that applications can securely and reliably access the necessary data. | + | | | | | | + | | | | | Configuration example: **'dew.csms.version'='v1'** | + +-------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +Example +------- + +This example demonstrates how to configure Flink OpenSource SQL to manage access credentials using DEW by generating random data through a DataGen table and outputting it to a MySQL result table. + +#. Create a shared secret in DEW. + + a. Log in to the DEW management console. + b. In the navigation pane on the left, choose **Cloud Secret Management Service** > **Secrets**. + c. Click **Create Secret**. On the displayed page, configure basic secret information. + + - **Secret Name**: Enter a secret name. In this example, the name is **secretInfo**. + - **Secret Value**: Enter the username and password for logging in to the RDS for MySQL DB instance. + + - The key in the first line is **MySQLUsername**, and the value is the username for logging in to the DB instance. + - The key in the second line is **MySQLPassword**, and the value is the password for logging in to the DB instance. + + d. Set other parameters as required and click **OK**. + +#. Configure DEW to manage access credentials in the Flink OpenSource SQL job. + + Ensure that an enhanced datasource connection between DLI and MySQL has been created. + + Ensure that an agency has been created for DLI to access DEW and authorization has been completed. + + Here is an example configuration for a Flink OpenSource SQL job: + + :: + + create table dataGenSource( + user_id string, + amount int + ) with ( + 'connector' = 'datagen', + 'rows-per-second' = '1', --Generate a piece of data per second. + 'fields.user_id.kind' = 'random', --Specify a random generator for the user_id field. + 'fields.user_id.length' = '3' --Limit the length of user_id to 3. + ); + + CREATE TABLE jdbcSink ( + user_id string, + amount int + ) + WITH ( + 'connector' = 'jdbc', + 'url? = 'jdbc:mysql://MySQLAddress:MySQLPort/flink',--flink is the MySQL database where the orders table locates. + 'table-name' = 'orders', + 'username' = 'MySQLUsername', -- Shared secret in DEW whose name is secretInfo and version is v1. The key MySQLUsername defines the secret value. The value is the user's sensitive information. + 'password' = 'MySQLPassword', -- Shared secret in DEW whose name is secretInfo and version is v1. The key MySQLPassword defines the secret value. The value is the user's sensitive information. + 'sink.buffer-flush.max-rows' = '1', + 'dew.endpoint'='endpoint', --Endpoint information for the DEW service being used + 'dew.csms.secretName'='secretInfo', --Name of the DEW shared secret + 'dew.csms.decrypt.fields'='username,password', --The username and password field values must be decrypted and replaced using DEW secret management. + 'dew.csms.version'='v1' + ); + + insert into jdbcSink select * from dataGenSOurce; diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/index.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/index.rst index 84fe423..1ffd8c2 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/index.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/index.rst @@ -5,10 +5,16 @@ Managing Data Source Access Credentials Using DEW ================================================= -- :ref:`Overview ` +- :ref:`Using an Agency to Obtain Access Credentials in DLI ` +- :ref:`Flink OpenSource SQL Jobs Using DEW to Manage Access Credentials ` +- :ref:`Flink Jar Jobs Using DEW to Acquire Access Credentials for Reading and Writing Data from and to OBS ` +- :ref:`Spark Jar Jobs Using DEW to Acquire Access Credentials for Reading and Writing Data from and to OBS ` .. toctree:: :maxdepth: 1 :hidden: - overview + using_an_agency_to_obtain_access_credentials_in_dli + flink_opensource_sql_jobs_using_dew_to_manage_access_credentials + flink_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs + spark_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/overview.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/overview.rst deleted file mode 100644 index 87d20d8..0000000 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/overview.rst +++ /dev/null @@ -1,21 +0,0 @@ -:original_name: dli_01_0687.html - -.. _dli_01_0687: - -Overview -======== - -When submitting Flink or Spark jobs through DLI to access external data sources (such as OBS and Kafka), there is a risk of plaintext exposure if AK/SK, usernames/passwords are directly embedded in the job code or parameter configurations. - -To securely store data source access credentials, ensure data source authentication safety, and facilitate secure access to data sources by DLI, you are advised to use DEW for managing data source access credentials. DLI employs "agency + temporary credentials" to safely retrieve data source access credentials. - -DEW is a comprehensive cloud-based encryption service designed to address challenges related to data security, key security, and the complexities of key management. - -This section describes how to use DEW to store data source authentication information across various job types. - -Notes and Constraints ---------------------- - -You are advised to use DEW for storing data source authentication information exclusively when Spark 3.3.1 or later and Flink 1.15 or later jobs access data sources using datasoure connections. - -When SQL and Flink 1.12 jobs access data sources using datasource connections, use DLI's datasource authentication feature to manage data source access credentials. For details, see :ref:`Using DLI Datasource Authentication to Manage Access Credentials for Data Sources `. diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/spark_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/spark_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs.rst new file mode 100644 index 0000000..920f901 --- /dev/null +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/spark_jar_jobs_using_dew_to_acquire_access_credentials_for_reading_and_writing_data_from_and_to_obs.rst @@ -0,0 +1,125 @@ +:original_name: dli_09_0215.html + +.. _dli_09_0215: + +Spark Jar Jobs Using DEW to Acquire Access Credentials for Reading and Writing Data from and to OBS +=================================================================================================== + +What Is a Temporary Credential? +------------------------------- + +A temporary security credential grants temporary access rights. It includes a temporary AK/SK and a security token, both of which must be used together. + +Scenario +-------- + +When writing output data from Spark Jar jobs to OBS, you need to configure an AK/SK for accessing OBS. To ensure the security of AK/SK data, you can use DEW and CSMS for centralized management of AK/SK. This approach effectively mitigates risks such as sensitive information leakage caused by hardcoding in programs or plaintext configurations, as well as potential business disruptions due to unauthorized access. + +This section walks you through on how a Spark Jar job acquires an AK/SK to read and write data from and to OBS. + +Notes and Constraints +--------------------- + +- DEW can be used to manage access credentials only in Spark 3.3.1 (Spark general queue scenario) or later. When creating a Spark job, select version 3.3.1 and configure the information of the agency that allows DLI to access DEW for the job. + +- To use this function, you need to configure AK/SK for all OBS buckets. + +Prerequisites +------------- + +- A shared secret has been created on the DEW console and the secret value has been stored. +- An agency has been created and authorized for DLI to access DEW. The agency must have been granted the following permissions: + + - Permission of the **ShowSecretVersion** interface for querying secret versions and secret values in DEW: **csms:secretVersion:get**. + - Permission of the **ListSecretVersions** interface for listing secret versions in DEW: **csms:secretVersion:list**. + - Permission to decrypt DEW secrets: **kms:dek:decrypt** + +Syntax +------ + +On the Spark Jar job editing page, configure the **Spark Arguments(--conf)** parameter as needed. The configuration information is as follows: + +Different OBS buckets use different AK/SK authentication information. You can use the following configuration method to specify the AK/SK information based on the bucket. For details about the parameters, see :ref:`Table 1 `. + +:: + + spark.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.access.key= USER_AK_CSMS_KEY + spark.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.secret.key= USER_SK_CSMS_KEY + spark.hadoop.fs.obs.security.provider = com.dli.provider.UserObsBasicCredentialProvider + spark.hadoop.fs.dew.csms.secretName= CredentialName + spark.hadoop.fs.dew.endpoint=ENDPOINT + spark.hadoop.fs.dew.csms.version=VERSION_ID + spark.hadoop.fs.dew.csms.cache.time.second =CACHE_TIME + spark.dli.job.agency.name=USER_AGENCY_NAME + +Parameter Description +--------------------- + +.. _dli_09_0215__en-us_topic_0000001883313257_en-us_topic_0000001841630985_table517231215112: + +.. table:: **Table 1** Parameter descriptions + + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Mandatory | Default Value | Data Type | Description | + +============================================================+=============+================+=============+=================================================================================================================================================================================================================================================================================================================+ + | spark.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.access.key | Yes | None | String | *USER_BUCKET_NAME* needs to be replaced with the user's OBS bucket name. | + | | | | | | + | | | | | The value of this parameter is the key defined by the user in the CSMS shared secret. The value corresponding to the key is the user's access key ID (AK). The user must have the permission to access the bucket on OBS. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.secret.key | Yes | None | String | *USER_BUCKET_NAME* needs to be replaced with the user's OBS bucket name. | + | | | | | | + | | | | | The value of this parameter is the key defined by the user in the CSMS shared secret. The value corresponding to the key is the user's secret access key (SK). The user must have the permission to access the bucket on OBS. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.hadoop.fs.obs.security.provider | Yes | None | String | OBS AK/SK authentication mechanism, which uses DEW-CSMS' secret management to obtain the AK and SK for accessing OBS. | + | | | | | | + | | | | | The default value is **com.dli.provider.UserObsBasicCredentialProvider**. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.hadoop.fs.dew.csms.secretName | Yes | None | String | Name of the shared secret in DEW's secret management. | + | | | | | | + | | | | | Configuration example: **spark.hadoop.fs.dew.csms.secretName=secretInfo** | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.hadoop.fs.dew.endpoint | Yes | None | String | Endpoint of the DEW service to be used. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.hadoop.fs.dew.csms.version | Yes | Latest version | String | Version number (secret version identifier) of the shared secret created in DEW CSMS. | + | | | | | | + | | | | | If the latest version number (secret version identifier) is not specified, the system will not be able to directly retrieve the most recent version of the secret, potentially leading to application access failures or the use of outdated secrets, thereby compromising data security and service stability. | + | | | | | | + | | | | | View the version information of the secret on the DEW management console and configure this parameter with the latest version number to ensure that applications can securely and reliably access the necessary data. | + | | | | | | + | | | | | Configuration example: **spark.hadoop.fs.dew.csms.version=v1** | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.hadoop.fs.dew.csms.cache.time.second | No | 3600 | Long | Cache duration after the CSMS shared secret is obtained during Spark job access. | + | | | | | | + | | | | | The unit is second. The default value is 3600 seconds. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.hadoop.fs.dew.projectId | No | Yes | String | ID of the project DEW belongs to. The default value is the ID of the project where the Spark job is. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | spark.dli.job.agency.name | Yes | ``-`` | String | Custom agency name. | + +------------------------------------------------------------+-------------+----------------+-------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +Sample Code +----------- + +This section describes how to write processed DataGen data to OBS. You need to modify the parameters in the sample Java code based on site requirements. + +#. Create an agency for DLI to access DEW and complete authorization. + +#. Create a shared secret in DEW. + + a. Log in to the DEW management console. + b. In the navigation pane on the left, choose **Cloud Secret Management Service** > **Secrets**. + c. On the displayed page, click **Create Secret**. Set basic secret information. + +#. Set job parameters on the DLI Spark Jar job editing page. + + Spark Arguments + + :: + + spark.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.access.key= USER_AK_CSMS_KEY + spark.hadoop.fs.obs.bucket.USER_BUCKET_NAME.dew.secret.key= USER_SK_CSMS_KEY + spark.hadoop.fs.obs.security.provider=com.dli.provider.UserObsBasicCredentialProvider + spark.hadoop.fs.dew.csms.secretName=obsAkSk + spark.hadoop.fs.dew.endpoint=kmsendpoint + spark.hadoop.fs.dew.csms.version=v3 + spark.dli.job.agency.name=agency diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/using_an_agency_to_obtain_access_credentials_in_dli.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/using_an_agency_to_obtain_access_credentials_in_dli.rst new file mode 100644 index 0000000..732fd0c --- /dev/null +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_data_source_access_credentials_using_dew/using_an_agency_to_obtain_access_credentials_in_dli.rst @@ -0,0 +1,41 @@ +:original_name: dli_01_0687.html + +.. _dli_01_0687: + +Using an Agency to Obtain Access Credentials in DLI +=================================================== + +When you use DLI to perform big data processing or cross-service data queries, authentication is a critical prerequisite to ensure legitimate access and data security. Embedding AK/SK or usernames and passwords directly in job code or configuration files introduces the risk of plaintext credential leakage. + +To address diverse access scenarios and authentication protocol requirements, DLI provides the following solutions: + +**DLI Agency with DEW Temporary Credentials and Permanent AK/SK** + +This solution is suitable for services that support only fine-grained authorization (version = 1.1) or workloads that require highly stable credentials without frequent rotation. It provides reliable authentication for these scenarios. + +In this approach, DLI uses agencies and temporary credentials to access the DEW service. DEW then provides permanent AK/SK, which are used to securely access other cloud services. + +Related guidance: + +- :ref:`Flink OpenSource SQL Jobs Using DEW to Manage Access Credentials ` +- :ref:`Flink Jar Jobs Using DEW to Acquire Access Credentials for Reading and Writing Data from and to OBS ` +- :ref:`Spark Jar Jobs Using DEW to Acquire Access Credentials for Reading and Writing Data from and to OBS ` + +Use Cases +--------- + +- Address security risks caused by hard-coded credentials. +- Avoid embedding sensitive information, such as data source usernames and passwords, in job code. +- Enable dynamic credential acquisition and periodic credential rotation. + +Notes and Constraints +--------------------- + +You are advised to use DEW for storing data source authentication information exclusively when Spark 3.3.1 or later and Flink 1.15 or later jobs access data sources using datasource connections. + +When SQL and Flink 1.12 jobs access data sources using datasource connections, use DLI's datasource authentication feature to manage data source access credentials. For details, see :ref:`Overview `. + +Learn More: What Are Temporary Security Credentials? +---------------------------------------------------- + +A temporary security credential grants temporary access rights. It includes a temporary AK/SK and a security token, both of which must be used together. diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/adding_a_route_for_an_enhanced_datasource_connection.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/adding_a_route_for_an_enhanced_datasource_connection.rst index 1c95fb2..d3ea7cf 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/adding_a_route_for_an_enhanced_datasource_connection.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/adding_a_route_for_an_enhanced_datasource_connection.rst @@ -46,14 +46,25 @@ Procedure .. table:: **Table 1** Parameters for adding a custom route - +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parameter | Description | - +===================================+==========================================================================================================================================================================================================+ - | Route Name | Name of a custom route, which is unique in the same enhanced datasource connection. The name can contain up to 64 characters. Only digits, letters, underscores (_), and hyphens (-) are allowed. | - +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | IP Address | Custom route CIDR block. The CIDR blocks of different routes can overlap but cannot be identical. | - | | | - | | Do not add the **100.125.**\ *xx.xx* or **100.64.**\ *xx.xx* CIDR blocks to avoid conflicts with the internal CIDR blocks of services like SWR, which can cause enhanced datasource connections to fail. | - +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+===================================================================================================================================================================================================================+ + | Route Name | Name of a custom route, which is unique in the same enhanced datasource connection. The name can contain up to 64 characters. Only digits, letters, underscores (_), and hyphens (-) are allowed. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | IP Address Type | The options are **IPv4** and **IPv6**. | + | | | + | | If your data source has IPv6 enabled and the current enhanced datasource connection supports IPv6, you can select IPv6 routes when adding a route table. | + | | | + | | You can check whether the current enhanced datasource connection supports IPv6 in its basic information. For details, see :ref:`Viewing Basic Information About an Enhanced Datasource Connection `. | + | | | + | | The route IP address example is as follows: | + | | | + | | - IPv4 address: **192.168.2.0/24**. | + | | - IPv6 address: **2407:c080:802:be7::/64**. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | IP Address | Custom route CIDR block. The CIDR blocks of different routes can overlap but cannot be identical. | + | | | + | | Do not add the **100.125.**\ *xx.xx* or **100.64.**\ *xx.xx* CIDR blocks to avoid conflicts with the internal CIDR blocks of services like SWR, which can cause enhanced datasource connections to fail. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ #. After adding a route, you can view the route information on the route details page. diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/enhanced_datasource_connection_tag_management.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/enhanced_datasource_connection_tag_management.rst index 95b9415..cc066a8 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/enhanced_datasource_connection_tag_management.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/enhanced_datasource_connection_tag_management.rst @@ -25,7 +25,7 @@ DLI allows you to add, modify, or delete tags for datasource connections. Procedure --------- -#. In the left navigation pane of the DLI management console, choose **Datasource Connections**. +#. In the navigation pane of the DLI management console, choose **Datasource Connections**. #. In the **Operation** column of the link, choose **More** > **Tags**. #. The tag management page is displayed, showing the tag information about the current connection. #. Click **Add/Edit Tag**. The **Add/Edit Tag** dialog is displayed. Add or edit tag keys and values and click **OK**. diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/viewing_basic_information_about_an_enhanced_datasource_connection.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/viewing_basic_information_about_an_enhanced_datasource_connection.rst index e2fc334..7309056 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/viewing_basic_information_about_an_enhanced_datasource_connection.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/managing_enhanced_datasource_connections/viewing_basic_information_about_an_enhanced_datasource_connection.rst @@ -7,7 +7,7 @@ Viewing Basic Information About an Enhanced Datasource Connection After creating an enhanced datasource connection, you can view and manage it on the management console. -This section describes how to view basic information about an enhanced datasource connection on the management console, including the enhanced datasource connection's host information and more. +This section describes how to view basic information about an enhanced datasource connection on the management console, including the enhanced datasource connection's host information, IPv6 support, and more. Procedure --------- @@ -25,6 +25,7 @@ Procedure You can view the following information: + - **IPv6 Support**: If you selected a subnet with IPv6 enabled when creating the enhanced datasource connection, then your enhanced datasource connection will support IPv6. - **Host Information**: When accessing an MRS HBase cluster, you need to configure the host name (domain name) and the corresponding IP address of the instance. For details, see :ref:`Modifying Host Information in an Elastic Resource Pool `. .. |image1| image:: /_static/images/en-us_image_0000001891931040.png diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/creating_a_password_datasource_authentication.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/creating_a_password_datasource_authentication.rst index 760eb83..0235b6f 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/creating_a_password_datasource_authentication.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/creating_a_password_datasource_authentication.rst @@ -8,7 +8,7 @@ Creating a Password Datasource Authentication Scenario -------- -Create a password datasource authentication on the DLI console to store passwords of the GaussDB(DWS), RDS, DCS, and DDS data sources to DLI. This will allow you to access to the data sources without having to configure a username and password in SQL jobs. +Create a password datasource authentication on the DLI console to store passwords of the DWS, RDS, DCS, and DDS data sources to DLI. This will allow you to access to the data sources without having to configure a username and password in SQL jobs. Procedure --------- diff --git a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/overview.rst b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/overview.rst index 9ee32eb..c8faf24 100644 --- a/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/overview.rst +++ b/umn/source/configuring_dli_to_read_and_write_data_from_and_to_external_data_sources/using_dli_datasource_authentication_to_manage_access_credentials_for_data_sources/overview.rst @@ -10,9 +10,9 @@ What Is Datasource Authentication? When analyzing across multiple sources, you are advised not to configure authentication information directly in a job as it can lead to password leakage. Instead, you are advised to use either DEW or datasource authentication provided by DLI to securely store data source authentication information. -- DEW is a comprehensive cloud-based encryption service designed to address challenges related to data security, key security, and the complexities of key management. You are advised to use DEW to store authentication information for data sources. +- DEW is designed to address critical challenges such as data security, key security, and the complexities of key management. You are advised to use DEW to securely store authentication credentials for your data sources. - You are advised to use DEW to store authentication information of data sources when Spark 3.3.1 or later and Flink 1.15 or later jobs access data sources using datasource connections. This will help you address issues related to data security, key security, and complex key management. For details, see :ref:`Managing Data Source Access Credentials Using DEW `. + For cross-source access scenarios involving Spark 3.3.1 or later versions, as well as Flink 1.15 or later versions, we strongly advise using DEW to manage your data source authentication information. This approach ensures robust solutions for data security, key protection, and streamlined key management. For details, see :ref:`Managing Data Source Access Credentials Using DEW `. - Datasource authentication is used to manage authentication information for accessing specified data sources. After datasource authentication is configured, you do not need to repeatedly configure data source authentication information in jobs, improving data source authentication security while enabling DLI to securely access data sources. @@ -35,7 +35,7 @@ Notes and Constraints | | - CSS: applies to 6.5.4 or later CSS clusters with the security mode enabled. | | | - Kerberos: applies to MRS security clusters with Kerberos authentication enabled. | | | - Kafka_SSL: applies to Kafka with SSL enabled. | - | | - Password: applies to GaussDB(DWS), RDS, DDS, and DCS. | + | | - Password: applies to DWS, RDS, DDS, and DCS. | +-----------------------------------+----------------------------------------------------------------------------------------------------------------------+ Datasource Authentication Types @@ -46,7 +46,7 @@ DLI supports four types of datasource authentication. Select an authentication t - CSS: applies to 6.5.4 or later CSS clusters with the security mode enabled. During the configuration, you need to specify the username, password, and authentication certificate of the cluster and store the information in DLI through datasource authentication so that DLI can securely access CSS data sources. For details, see :ref:`Creating a CSS Datasource Authentication `. - Kerberos: applies to MRS security clusters with Kerberos authentication enabled. During the configuration, you need to specify MRS cluster authentication credentials, including the **krb5.conf** and **user.keytab** files. For details, see :ref:`Creating a Kerberos Datasource Authentication `. - Kafka_SSL: applies to Kafka with SSL enabled. During the configuration, you need to specify the KafkaTruststore path and password. For details, see :ref:`Creating a Kafka_SSL Datasource Authentication `. -- Password: applies to GaussDB(DWS), RDS, DDS, and DCS data sources. During the configuration, you need to store the passwords of the data sources in DLI. For details, see :ref:`Creating a Password Datasource Authentication `. +- Password: applies to DWS, RDS, DDS, and DCS data sources. During the configuration, you need to store the passwords of the data sources in DLI. For details, see :ref:`Creating a Password Datasource Authentication `. Jobs That Can Connect to Data Sources Through Datasource Authentication ----------------------------------------------------------------------- @@ -60,42 +60,42 @@ Different types of jobs can connect to data sources through different types of d .. table:: **Table 2** Data sources that Spark SQL jobs can connect to through datasource authentication - +--------------------------------+-----------------------------------+-----------------------------------------------------+ - | Datasource Authentication Type | Data Source | Notes and Constraints | - +================================+===================================+=====================================================+ - | CSS | CSS | The cluster version is 6.5.4 or later. | - | | | | - | | | The security mode has been enabled for the cluster. | - +--------------------------------+-----------------------------------+-----------------------------------------------------+ - | Password | GaussDB(DWS), RDS, DDS, and Redis | ``-`` | - +--------------------------------+-----------------------------------+-----------------------------------------------------+ + +--------------------------------+--------------------------+-----------------------------------------------------+ + | Datasource Authentication Type | Data Source | Notes and Constraints | + +================================+==========================+=====================================================+ + | CSS | CSS | The cluster version is 6.5.4 or later. | + | | | | + | | | The security mode has been enabled for the cluster. | + +--------------------------------+--------------------------+-----------------------------------------------------+ + | Password | DWS, RDS, DDS, and Redis | ``-`` | + +--------------------------------+--------------------------+-----------------------------------------------------+ .. _dli_01_0561__table208001745193719: .. table:: **Table 3** Data sources that Flink SQL jobs can connect to through datasource authentication - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ - | Table Type | Datasource Authentication Type | Data Source | Notes and Constraints | - +=================+================================+============================+===============================================================+ - | Source table | Kerberos | Kafka | Kerberos authentication has been enabled for MRS Kafka. | - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ - | | Kafka_SSL | Kafka | SASL_SSL authentication has been enabled for DMS Kafka. | - | | | | | - | | | | SASL authentication has been enabled for MRS Kafka. | - | | | | | - | | | | SSL authentication has been enabled for MRS Kafka. | - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ - | Result table | Kerberos | HBase | Kerberos authentication has been enabled for the MRS cluster. | - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ - | | | Kafka | Kerberos authentication has been enabled for MRS Kafka. | - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ - | | Kafka_SSL | Kafka | SASL_SSL authentication has been enabled for DMS Kafka. | - | | | | | - | | | | SASL authentication has been enabled for MRS Kafka. | - | | | | | - | | | | SSL authentication has been enabled for MRS Kafka. | - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ - | | Password | GaussDB(DWS), RDS, and CSS | ``-`` | - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ - | Dimension table | Password | RDS and Redis | ``-`` | - +-----------------+--------------------------------+----------------------------+---------------------------------------------------------------+ + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ + | Table Type | Datasource Authentication Type | Data Source | Notes and Constraints | + +=================+================================+===================+===============================================================+ + | Source table | Kerberos | Kafka | Kerberos authentication has been enabled for MRS Kafka. | + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ + | | Kafka_SSL | Kafka | SASL_SSL authentication has been enabled for DMS Kafka. | + | | | | | + | | | | SASL authentication has been enabled for MRS Kafka. | + | | | | | + | | | | SSL authentication has been enabled for MRS Kafka. | + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ + | Result table | Kerberos | HBase | Kerberos authentication has been enabled for the MRS cluster. | + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ + | | | Kafka | Kerberos authentication has been enabled for MRS Kafka. | + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ + | | Kafka_SSL | Kafka | SASL_SSL authentication has been enabled for DMS Kafka. | + | | | | | + | | | | SASL authentication has been enabled for MRS Kafka. | + | | | | | + | | | | SSL authentication has been enabled for MRS Kafka. | + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ + | | Password | DWS, RDS, and CSS | ``-`` | + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ + | Dimension table | Password | RDS and Redis | ``-`` | + +-----------------+--------------------------------+-------------------+---------------------------------------------------------------+ diff --git a/umn/source/creating_a_data_directory_database_and_table/creating_a_data_catalog_database_and_table_on_the_dli_console.rst b/umn/source/creating_a_data_directory_database_and_table/creating_a_data_catalog_database_and_table_on_the_dli_console.rst index 519684e..16f1556 100644 --- a/umn/source/creating_a_data_directory_database_and_table/creating_a_data_catalog_database_and_table_on_the_dli_console.rst +++ b/umn/source/creating_a_data_directory_database_and_table/creating_a_data_catalog_database_and_table_on_the_dli_console.rst @@ -88,7 +88,7 @@ Before creating a table, ensure that a database has been created. .. note:: - Datasource connection tables, such as View tables, HBase (MRS) tables, OpenTSDB (MRS) tables, GaussDB(DWS) tables, RDS tables, and CSS tables, cannot be created. You can use SQL to create views and datasource connection tables. For details, see "Creating a View" and "Creating a Datasource Connection Table" in *Data Lake Insight SQL Syntax Reference*. + Datasource connection tables, such as View tables, HBase (MRS) tables, OpenTSDB (MRS) tables, DWS tables, RDS tables, and CSS tables, cannot be created. You can use SQL to create views and datasource connection tables. For details, see "Creating a View" and "Creating a Datasource Connection Table" in the *Data Lake Insight SQL Syntax Reference*. - To create a table on the **Data Management** page: @@ -144,7 +144,7 @@ Before creating a table, ensure that a database has been created. | | - **date**: The value ranges from 0000-01-01 to 9999-12-31. | | | | - **double**: Each number is stored on eight bytes. | | | | - **boolean**: Each value is stored on one byte. | | - | | - **decimal**: The valid bits are positive integers between 1 to 38, including 1 and 38. The decimal digits are integers less than 10. | | + | | - **decimal**: The valid bits are positive integers ranging from 1 to 38. The decimal digits are integers less than 10. | | | | - **smallint/short**: The number is stored on two bytes. | | | | - **bigint/long**: The number is stored on eight bytes. | | | | - **timestamp**: The data indicates a date and time. The value can be accurate to six decimal points. | | diff --git a/umn/source/creating_a_data_directory_database_and_table/managing_table_resources_on_the_dli_console/exporting_dli_table_data_to_obs.rst b/umn/source/creating_a_data_directory_database_and_table/managing_table_resources_on_the_dli_console/exporting_dli_table_data_to_obs.rst index 2de1f3e..ec716f0 100644 --- a/umn/source/creating_a_data_directory_database_and_table/managing_table_resources_on_the_dli_console/exporting_dli_table_data_to_obs.rst +++ b/umn/source/creating_a_data_directory_database_and_table/managing_table_resources_on_the_dli_console/exporting_dli_table_data_to_obs.rst @@ -29,7 +29,7 @@ Procedure a. In the navigation pane on the left of the management console, choose **SQL Editor**. b. In the navigation tree on the left, click **Databases** to see all databases. Click the database name corresponding to the table to which data is to be exported. The tables are displayed. - c. Click |image1| on the right of the managed table (DLI table) whose data is to be exported, and choose **Export** from the shortcut menu. + c. Click |image1| next to the managed table (DLI table) whose data is to be exported, and choose **Export** from the shortcut menu. #. In the displayed **Export Data** dialog box, specify parameters by referring to :ref:`Table 1 `. diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_a_non-elastic_resource_pool_queue_deprecated_not_recommended.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_a_non-elastic_resource_pool_queue_deprecated_not_recommended.rst index 446fa0d..d4b98f3 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_a_non-elastic_resource_pool_queue_deprecated_not_recommended.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_a_non-elastic_resource_pool_queue_deprecated_not_recommended.rst @@ -7,7 +7,7 @@ Creating a Non-Elastic Resource Pool Queue (Deprecated, Not Recommended) Queues in the non-elastic resource pool mode are the previous-gen of resource management for DLI. It involved purchasing and releasing resources based on usage demands, requiring estimation of resource needs before making purchases. -Queues in an elastic resource pool are recommended, as they offer the flexibility to use resources with high utilization as needed. For how to buy an elastic resource pool and create queues within it, see :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. +Queues in an elastic resource pool are recommended, as they offer the flexibility to use resources with high utilization as needed. For details about how to buy an elastic resource pool and create queues within it, see :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. .. note:: @@ -46,7 +46,7 @@ Procedure #. You can create a queue on the **Overview**, **SQL Editor**, or **Queue Management** page. - - In the upper right corner of the **Overview** page, click Create Queue. + - In the upper right corner of the **Overview** page, click **Create Queue**. - To create a queue on the **Queue Management** page: a. In the navigation pane on the left of the DLI management console, choose **Resources** > **Queue Management**. @@ -57,7 +57,7 @@ Procedure a. In the navigation pane on the left of the DLI management console, choose **SQL Editor**. b. Click **Queues**. On the tab page displayed, click |image1| on the right to create a queue. -#. On the **Create Queue** page displayed, set the parameters according to :ref:`Table 2 `. +#. On the displayed **Create Queue** page, configure the parameters according to :ref:`Table 2 `. .. _dli_01_0363__table17301125219910: diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_an_elastic_resource_pool_and_creating_queues_within_it.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_an_elastic_resource_pool_and_creating_queues_within_it.rst index b78b2e9..2f4a68e 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_an_elastic_resource_pool_and_creating_queues_within_it.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/creating_an_elastic_resource_pool_and_creating_queues_within_it.rst @@ -22,31 +22,39 @@ Notes and Constraints .. table:: **Table 1** Notes and constraints on elastic resource pools - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Item | Description | - +===================================+============================================================================================================================================================================================================================================================================================================================================================================================+ - | Resource specifications | - An elastic resource pool currently supports up to 32,000 CUs. | - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Managing elastic resource pools | - You cannot change the region of an elastic resource pool once the pool is created. | - | | - Flink 1.10 or later jobs can run in elastic resource pools. | - | | - The CIDR block of an elastic resource pool cannot be changed once set. | - | | - You can view only the scaling history of an elastic resource pool within 30 days. | - | | - Elastic resource pools cannot directly access the Internet. | - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Elastic resource pool scaling | - Changes to elastic resource pool CUs can occur when setting the CU, adding or deleting queues in an elastic resource pool, or modifying the scaling policies of queues in an elastic resource pool, or when the system automatically triggers elastic resource pool scaling. However, in some cases, the system cannot guarantee that the scaling will reach the target CUs as planned. | - | | | - | | - If there are not enough physical resources, an elastic resource pool may not be able to scale out to the desired target size. | - | | | - | | - The system does not guarantee that an elastic resource pool will be scaled in to the desired target size. | - | | | - | | The system checks the resource usage before scaling in the elastic resource pool to determine if there is enough space for scaling in. If the existing resources cannot be scaled in according to the minimum scaling step, the pool may not be scaled in successfully or only partially. | - | | | - | | The scaling step may vary depending on the resource specifications, usually 16 CUs, 32 CUs, 48 CUs, 64 CUs, and more. | - | | | - | | For example, if the elastic resource pool has a capacity of 192 CUs and the queues in the pool are using 68 CUs due to running jobs, the plan is to scale in to 64 CUs. | - | | | - | | When executing a scaling in task, the system determines that there are 124 CUs remaining and scales in by the minimum step of 64 CUs. However, the remaining 60 CUs cannot be scaled in any further. Therefore, after the elastic resource pool executes the scaling in task, its capacity is reduced to 128 CUs. | - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Item | Description | + +===================================================+============================================================================================================================================================================================================================================================================================================================================================================================+ + | Resource specifications | - An elastic resource pool currently supports up to 32,000 CUs. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Managing elastic resource pools | - You cannot change the region of an elastic resource pool once the pool is created. | + | | - Flink 1.10 or later jobs can run in elastic resource pools. | + | | - The CIDR block of an elastic resource pool cannot be changed once set. | + | | - You can view only the scaling history of an elastic resource pool within 30 days. | + | | - Elastic resource pools cannot directly access the Internet. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Elastic resource pool scaling | - Changes to elastic resource pool CUs can occur when setting the CU, adding or deleting queues in an elastic resource pool, or modifying the scaling policies of queues in an elastic resource pool, or when the system automatically triggers elastic resource pool scaling. However, in some cases, the system cannot guarantee that the scaling will reach the target CUs as planned. | + | | | + | | - If there are not enough physical resources, an elastic resource pool may not be able to scale out to the desired target size. | + | | | + | | - The system does not guarantee that an elastic resource pool will be scaled in to the desired target size. | + | | | + | | The system checks the resource usage before scaling in the elastic resource pool to determine if there is enough space for scaling in. If the existing resources cannot be scaled in according to the minimum scaling step, the pool may not be scaled in successfully or only partially. | + | | | + | | The scaling step may vary depending on the resource specifications, usually 16 CUs, 32 CUs, 48 CUs, 64 CUs, and more. | + | | | + | | For example, if the elastic resource pool has a capacity of 192 CUs and the queues in the pool are using 68 CUs due to running jobs, the plan is to scale in to 64 CUs. | + | | | + | | When executing a scaling in task, the system determines that there are 124 CUs remaining and scales in by the minimum step of 64 CUs. However, the remaining 60 CUs cannot be scaled in any further. Therefore, after the elastic resource pool executes the scaling in task, its capacity is reduced to 128 CUs. | + | | | + | | - If jobs such as Flink or Spark Streaming jobs are continuously running on a node, that node cannot be scaled in as scheduled. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Constraints for setting elastic resource pool CUs | - The minimum CUs (minCU) of an elastic resource pool must be **less than or equal to** the actual CUs. If expanding minCU exceeds the current actual CUs, you must first increase the actual CUs. Otherwise, the modification will fail. | + | | - The sum of all queues' minimum CUs in an elastic resource pool must not exceed the pool's minCU. | + | | - Any single queue's maxCU cannot exceed the pool's maxCU. | + | | - Adjustments to a queue's CU range, changes to the pool's specifications, or modifications to the pool's CU settings take effect at the next full hour. | + | | - Increasing the number of queues to adjust the pool's actual CUs takes immediate effect. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ Creating an Elastic Resource Pool --------------------------------- @@ -108,6 +116,15 @@ Creating an Elastic Resource Pool | | | | | 192.168.0.0-192.168.0.0/16-19 | +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | IPv6 | After IPv6 is enabled, the elastic resource pool can communicate with data sources (the CIDR block of which have IPv6 enabled) using IPv6 addresses. | + | | | + | | If IPv6 is enabled, an IPv6 CIDR block will be automatically allocated to the elastic resource pool. | + | | | + | | - IPv6 can only be enabled during creation. | + | | - After enabling, both IPv4 and IPv6 addresses will be available, allowing for both private and public network access. IPv6 cannot be used independently. | + | | - The IPv6 CIDR block cannot be specified; it is allocated by the system. | + | | - IPv6 cannot be disabled after being enabled. | + +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Enterprise Project | If the created elastic resource pool belongs to an enterprise project, select the enterprise project. | | | | | | Enterprise projects let you manage cloud resources and users by project. | @@ -164,11 +181,7 @@ Creating a queue within an elastic resource pool will trigger changes of elastic | Type | - **For SQL**: The queue is used to run SQL jobs. | | | - **For general purpose**: The queue is used to run Spark and Flink jobs. | +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Engine | If **Type** is **For SQL**, the queue engine can be **Spark** or **HetuEngine**. | - | | | - | | If **HetuEngine** is selected, the minimum number of CUs of the SQL queue cannot be fewer than 96 CUs. | - | | | - | | To use HetuEngine to submit SQL jobs, you need to configure a DLI job bucket. For details, see :ref:`Configuring a DLI Job Bucket `. | + | Engine | If **Type** is **For SQL**, the queue engine can only be **Spark**. | +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Enterprise Project | Select the enterprise project the queue belongs to. Queues under different enterprise projects can be added to an elastic resource pool. | | | | @@ -207,39 +220,37 @@ Creating a queue within an elastic resource pool will trigger changes of elastic .. table:: **Table 4** Scaling policy parameters - +-----------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parameter | Description | - +===================================+===============================================================================================================================================================================================================================================+ - | Priority | Priority of the scaling policy in the current elastic resource pool. A larger value indicates a higher priority. You can set a number ranging from 1 to 100. | - +-----------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Period | Time segment when the policy takes effect. It can be set only by hour. The start time is on the left, and the end time is on the right. | - | | | - | | - The time range includes the start time but not the end time, that is, [start time, end time). | - | | | - | | For example, if you set **Period** to **01** and **17**, the scaling policy takes effect at 01:00 a.m. till 05:00 p.m. | - | | | - | | - The periods of scaling policies with different priorities should not overlap. | - +-----------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Min CU | Minimum number of CUs allowed by the scaling policy. | - | | | - | | - In any time segment of a day, the total minimum CUs of all queues in an elastic resource pool cannot be more than the minimum CUs of the pool. | - | | - If the minimum CUs of the queue are fewer than 16 CUs, both **Max. Spark Driver Instances** and **Max. Prestart Spark Driver Instances** set in the queue properties do not apply. Refer to :ref:`Setting Queue Properties `. | - | | | - | | For a HetuEngine SQL queue, there must be at least 96 CUs. | - +-----------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Max CU | Maximum number of CUs allowed by the scaling policy. | - | | | - | | In any time segment of a day, the maximum CUs of any queue in an elastic resource pool cannot be more than the maximum CUs of the pool. | - | | | - | | - The maximum CUs of a queue in the elastic resource pool of the basic edition must be a multiple of 4. | - | | - The maximum CUs of a queue in the elastic resource pool of the standard edition must be a multiple of 16. | - +-----------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Use Resource Pool's Max CU | When selected, the maximum CUs of queues equal the maximum CUs of the resource pool within the period specified by the current scaling policy. | - | | | - | | Increasing the maximum CUs of the elastic resource pool automatically adjusts the queue's maximum CUs to match without manual intervention. | - | | | - | | The setting takes effect only within the period specified by the current scaling policy. You need to manually configure CU limits for other periods. | - +-----------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+==========================================================================================================================================================================================================================================+ + | Priority | Priority of the scaling policy in the current elastic resource pool. A larger value indicates a higher priority. You can set a number ranging from 1 to 100. | + +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Period | Time segment when the policy takes effect. It can be set only by hour. The start time is on the left, and the end time is on the right. | + | | | + | | - The time range includes the start time but not the end time, that is, [start time, end time). | + | | | + | | For example, if you set **Period** to **01** and **17**, the scaling policy takes effect at 01:00 a.m. till 05:00 p.m. | + | | | + | | - The periods of scaling policies with different priorities should not overlap. | + +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Min CU | Minimum number of CUs allowed by the scaling policy. | + | | | + | | - In any time segment of a day, the total minimum CUs of all queues in an elastic resource pool cannot be more than the minimum CUs of the pool. | + | | - If the minimum CUs of the queue are fewer than 16 CUs, both **Max. Spark Driver Instances** and **Max. Prestart Spark Driver Instances** set in the queue properties do not apply. See :ref:`Setting Queue Properties `. | + +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Max CU | Maximum number of CUs allowed by the scaling policy. | + | | | + | | In any time segment of a day, the maximum CUs of any queue in an elastic resource pool cannot be more than the maximum CUs of the pool. | + | | | + | | - The maximum CUs of a queue in the elastic resource pool of the basic edition must be a multiple of 4. | + | | - The maximum CUs of a queue in the elastic resource pool of the standard edition must be a multiple of 16. | + +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Use Resource Pool's Max CU | When selected, the maximum CUs of queues equal the maximum CUs of the resource pool within the period specified by the current scaling policy. | + | | | + | | Increasing the maximum CUs of the elastic resource pool automatically adjusts the queue's maximum CUs to match without manual intervention. | + | | | + | | The setting takes effect only within the period specified by the current scaling policy. You need to manually configure CU limits for other periods. | + +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ .. note:: diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/example_use_case_creating_an_elastic_resource_pool_and_running_jobs.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/example_use_case_creating_an_elastic_resource_pool_and_running_jobs.rst index 75b7bed..5e8133a 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/example_use_case_creating_an_elastic_resource_pool_and_running_jobs.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/example_use_case_creating_an_elastic_resource_pool_and_running_jobs.rst @@ -15,26 +15,26 @@ This section walks you through the procedure of adding a queue to an elastic res .. table:: **Table 1** Procedure - +------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ - | Step | Description | Reference | - +======================================================+=======================================================================================================================================+========================================================================================+ - | Create an elastic resource pool | Create an elastic resource pool and configure basic information, such as the billing mode, CU range, and CIDR block. | :ref:`Creating an Elastic Resource Pool and Creating Queues Within It ` | - +------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ - | Add a queue to the elastic resource pool | Add the queue where your jobs will run on to the elastic resource pool. The operations are as follows: | :ref:`Creating an Elastic Resource Pool and Creating Queues Within It ` | - | | | | - | | #. Set basic information about the queue, such as the name and type. | :ref:`Adjusting Scaling Policies for Queues in an Elastic Resource Pool ` | - | | #. Configure the scaling policy of the queue, including the priority, period, and the maximum and minimum CUs allowed for scaling. | | - +------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ - | (Optional) Create an enhanced datasource connection. | If a job needs to access data from other data sources, for example, GaussDB(DWS) and RDS, you need to create a datasource connection. | :ref:`Creating an Enhanced Datasource Connection ` | - | | | | - | | The created datasource connection must be bound to the elastic resource pool. | | - +------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ - | Run a job. | Create and submit the job as you need. | :ref:`Managing SQL Jobs ` | - | | | | - | | | :ref:`Flink Job Overview ` | - | | | | - | | | :ref:`Creating a Spark Job ` | - +------------------------------------------------------+---------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ + +------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ + | Step | Description | Reference | + +======================================================+====================================================================================================================================+========================================================================================+ + | Create an elastic resource pool | Create an elastic resource pool and configure basic information, such as the billing mode, CU range, and CIDR block. | :ref:`Creating an Elastic Resource Pool and Creating Queues Within It ` | + +------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ + | Add a queue to the elastic resource pool | Add the queue where your jobs will run on to the elastic resource pool. The operations are as follows: | :ref:`Creating an Elastic Resource Pool and Creating Queues Within It ` | + | | | | + | | #. Set basic information about the queue, such as the name and type. | :ref:`Adjusting Scaling Policies for Queues in an Elastic Resource Pool ` | + | | #. Configure the scaling policy of the queue, including the priority, period, and the maximum and minimum CUs allowed for scaling. | | + +------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ + | (Optional) Create an enhanced datasource connection. | If a job needs to access data from other data sources, for example, DWS and RDS, you need to create a datasource connection. | :ref:`Creating an Enhanced Datasource Connection ` | + | | | | + | | The created datasource connection must be bound to the elastic resource pool. | | + +------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ + | Run a job. | Create and submit the job as you need. | :ref:`Managing SQL Jobs ` | + | | | | + | | | :ref:`Flink Job Overview ` | + | | | | + | | | :ref:`Creating a Spark Job ` | + +------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------+ .. _dli_01_0515__section20375737115219: diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/adjusting_scaling_policies_for_queues_in_an_elastic_resource_pool.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/adjusting_scaling_policies_for_queues_in_an_elastic_resource_pool.rst index 169da7e..12fe327 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/adjusting_scaling_policies_for_queues_in_an_elastic_resource_pool.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/adjusting_scaling_policies_for_queues_in_an_elastic_resource_pool.rst @@ -5,28 +5,28 @@ Adjusting Scaling Policies for Queues in an Elastic Resource Pool ================================================================= -Multiple queues can be added to an elastic resource pool. For how to add a queue, see :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. You can configure the number of CUs you want based on the compute resources used by DLI queues during peaks and troughs and set priorities for the scaling policies to ensure stable running of jobs. +Elastic resource pools support multiple queues to manage job execution efficiently. For details about how to create queues in a pool, see :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. Once configured, you can define scaling policies based on each queue's compute resource usage patterns, such as peak and off-peak periods, and priority levels. This ensures optimal allocation of CUs to maintain stable and efficient job performance. Precautions ----------- -- You are advised to implement fine-grained management of resource pools for stream and batch processing jobs by placing Flink real-time stream jobs and SQL batch processing jobs in separate elastic resource pools. +- You are advised to implement fine-grained management by separating Flink real-time streaming jobs from SQL batch processing jobs into distinct elastic resource pools. This approach offers two key benefits: - Flink real-time stream jobs can run stably without forced scale-in, thus avoiding job interruption and system instability. + Flink streaming jobs run continuously and require stability without forced scale-in, preventing interruptions and system instability. - SQL batch processing jobs are placed in independent resource pools, which can scale out and in more flexibly, significantly enhancing the success rate and operational efficiency of scaling operations. + SQL batch jobs benefit from flexible scaling within their dedicated pool, significantly improving scaling success rates and operational efficiency. -- In any time segment of a day, the total minimum CUs of all queues in an elastic resource pool cannot be more than the minimum CUs of the pool. +- At any time of day, the sum of minimum CUs across all queues in the elastic resource pool must not exceed the pool's minimum CUs. -- In any time segment of a day, the maximum CUs of any queue in an elastic resource pool cannot be more than the maximum CUs of the pool. +- Similarly, at any given time, no single queue's maximum CU allocation should exceed the pool's maximum CUs. -- The periods of scaling policies cannot overlap. +- The time intervals for different scaling policies within the same queue must not overlap. -- The period of a scaling policy can only be set by hour and specified by the start time and end time. For example, if you set the period to **00-09**, the time range when the policy takes effect is [00:00, 09:00). The period of the default scaling policy cannot be modified. +- Scaling policy periods are set in whole-hour increments, including the start time but excluding the end time. For example, a 00-09 interval covers [00:00, 09:00). Note that default scaling policies do not allow modifications to these time settings. -- In any period, compute resources are preferentially allocated to meet the minimum number of CUs of all queues. The remaining CUs (total CUs of the elastic resource pool - total minimum CUs of all queues) are allocated in accordance with the scaling policy priorities. +- During any period, the system first ensures all queues meet their minimum CU requirements. Any remaining CUs (calculated as the pool's maximum CUs minus the sum of all queues' minimum CUs) are allocated based on configured priorities until fully distributed. -- After the queue is scaled out, the system starts billing you for the added CUs. So, if you do not have sufficient requirements, scale in your queue to release unnecessary CUs to save cost. +- Once a queue successfully scales out, billing begins for the additional CUs and continues until scale-in occurs. To avoid unnecessary costs, promptly release resources when they are no longer needed, as unused CUs will still incur charges. .. table:: **Table 1** CU allocation (without jobs) diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/allocating_to_an_enterprise_project.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/allocating_to_an_enterprise_project.rst index 4a85234..a13ca7c 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/allocating_to_an_enterprise_project.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/allocating_to_an_enterprise_project.rst @@ -5,15 +5,15 @@ Allocating to an Enterprise Project =================================== -You can create enterprise projects matching the organizational structure of your enterprises to centrally manage cloud resources across regions by project. Then you can create user groups and users with different permissions and add them to enterprise projects. +An enterprise project is a cloud resource management approach that allows organizations to plan resources based on their organizational structure. It enables unified management of resources distributed across different regions under specific enterprise projects. Additionally, user groups and users with varying permissions can be assigned to each enterprise project. -DLI allows you to select an enterprise project when creating an elastic resource pool. This section describes how to bind an elastic resource pool to and modify an enterprise project. +DLI allows you to select an enterprise project when creating an elastic resource pool. This section explains how to bind or modify the enterprise project for a DLI elastic resource pool. .. note:: - Modifying the enterprise project of an elastic resource pool will modify the enterprise projects of the queues in the elastic resource pool. + Changing the enterprise project of an elastic resource pool will also update the enterprise project of its associated queue resources. - Only queues under the same enterprise project can be bound to an elastic resource pool. + Here, the elastic resource pool only supports adding queues from the same enterprise project. Prerequisites ------------- @@ -27,7 +27,7 @@ When creating an elastic resource pool, you can select a created enterprise proj Alternatively, you can click **Create Enterprise Project** to go to the Enterprise Project Management Service console to create an enterprise project and check existing ones. -For how to create a queue, see :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. +For details about how to create a queue, see :ref:`Creating an Elastic Resource Pool and Creating Queues Within It `. Modifying an Enterprise Project ------------------------------- diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/checking_basic_information.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/checking_basic_information.rst index b7d4c32..5c510c8 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/checking_basic_information.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/checking_basic_information.rst @@ -23,29 +23,29 @@ Procedure #. Click |image2| to expand the basic information card of the elastic resource pool and view detailed information about the pool. - These details include the name of the pool, the user who created it, the date it was created, and the VPC CIDR block. + These details include the name of the pool, the user who created it, the date it was created, whether IPv6 is enabled, and the VPC CIDR block. If IPv6 is enabled, the subnet's IPv6 CIDR block will also be displayed. - For the definitions of actual CUs, used CUs, CU range, and specifications of the elastic resource pool, refer to :ref:`Actual CUs, Used CUs, CU Range, and Specifications of an Elastic Resource Pool `. + For details about the definitions of actual CUs, used CUs, CU range, and yearly/monthly CUs (specifications) of the elastic resource pool, see :ref:`Actual CUs, Used CUs, CU Range, and Yearly/Monthly CUs (Specifications) of an Elastic Resource Pool `. .. _dli_01_0622__section3723248135610: -Actual CUs, Used CUs, CU Range, and Specifications of an Elastic Resource Pool ------------------------------------------------------------------------------- +Actual CUs, Used CUs, CU Range, and Yearly/Monthly CUs (Specifications) of an Elastic Resource Pool +--------------------------------------------------------------------------------------------------- -- **Actual CUs**: actual size of resources currently allocated to the elastic resource pool (in CUs). +- **Actual CUs**: The current allocated resource size of the elastic resource pool, measured in CUs. - - When there is no queue in the resource pool, the actual CUs are equal to the minimum CUs when the elastic resource pool is created. + - When no queues exist in the resource pool: The actual CUs equal the minimum CUs set during its creation. - - When there are queues in the resource pool, the formula for calculating actual CUs is: + - When there are queues in the resource pool, the formula to calculate actual CUs is: - Actual CUs = max{(min[sum(maximum CUs of queues), maximum CUs of the elastic resource pool]), minimum CUs of the elastic resource pool}. - - The calculation result must be a multiple of 16 CUs. If it cannot be exactly divided by 16 CUs, round up to the nearest multiple. + - The result must be a multiple of 16 CUs. If not divisible by 16, round up to the nearest multiple. - - Scaling out or in an elastic resource pool means adjusting the actual CUs of the resource pool. Refer to :ref:`Scaling Out or In an Elastic Resource Pool `. + - Scaling out or in an elastic resource pool means adjusting its actual CUs. See :ref:`Scaling Out or In an Elastic Resource Pool `. - Example of actual CU allocation: - In :ref:`Table 1 `, the calculation process for the actual allocation of CUs in an elastic resource pool is as follows: + Consider :ref:`Table 1 ` below, which illustrates the process of calculating actual CUs for an elastic resource pool: #. Calculate the sum of maximum CUs of the queues: sum(maximum CUs) = 32 + 56 = 88 CUs. @@ -74,15 +74,15 @@ Actual CUs, Used CUs, CU Range, and Specifications of an Elastic Resource Pool | | Queue B | 16-56CUS | +---------------------------------------------------------------------------------------------------+-----------------------+-----------------------+ -- **Used CUs**: CUs that have been used by jobs or tasks. These resources may be executing computing tasks. +- **Used CUs**: The portion of CUs currently occupied by jobs or tasks, which may be actively performing computations. - **CU range**: CU settings are used to control the maximum and minimum CU ranges for elastic resource pools to avoid unlimited resource scaling. - - The total minimum CUs of all queues in an elastic resource pool must be no more than the minimum CUs of the pool. - - The maximum CUs of any queue in an elastic resource pool must be no more than the maximum CUs of the pool. - - An elastic resource pool should at least ensure that all queues in it can run with the minimum CUs and should try to ensure that all queues in it can run with the maximum CUs. - - When expanding the specifications of an elastic resource pool, the minimum value of the CU range is linked to the specifications of the elastic resource pool. After changing the specifications of the elastic resource pool, the minimum value of the CU range is modified to match the specifications. + - The sum of all queues' minimum CUs in an elastic resource pool must not exceed the pool's minCU. + - Any single queue's maxCU cannot exceed the pool's maxCU. + - The resource pool ensures it meets the minCU requirements across all queues while striving to accommodate their maxCU demands. + - When expanding the specifications of an elastic resource pool, the minimum value of the CU range is linked to the yearly/monthly CUs (specifications) of the elastic resource pool. After changing the specifications of the elastic resource pool, the minimum value of the CU range is modified to match the yearly/monthly CUs (specifications). -- **Specifications**: The minimum CUs selected during elastic resource pool purchase are elastic resource pool specifications. +- **Yearly/monthly CUs (specifications)**: The minimum value of the CU range selected when purchasing an elastic resource pool is the elastic resource pool specifications. .. |image1| image:: /_static/images/en-us_image_0000001934973209.png .. |image2| image:: /_static/images/en-us_image_0000002129969110.png diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/managing_tags.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/managing_tags.rst index 645d45f..eb4d055 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/managing_tags.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_elastic_resource_pools/managing_tags.rst @@ -22,7 +22,7 @@ DLI supports the following two types of tags: DLI allows you to add, modify, or delete tags for queues. -#. In the left navigation pane of the DLI console, choose **Resources** > **Resource Pool**. +#. In the navigation pane of the DLI console, choose **Resources** > **Resource Pool**. #. In the **Operation** column of the queue, choose **More** > **Tags**. #. The tag management page is displayed, showing the tag information about the current queue. #. On the page that appears, click **Add/Edit Tag** in the upper left corner. Enter a tag and a value, and click **Add**. diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/deleting_a_queue.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/deleting_a_queue.rst index 34ab04d..f91554e 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/deleting_a_queue.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/deleting_a_queue.rst @@ -16,7 +16,7 @@ Procedure --------- #. In the navigation pane on the left of the DLI management console, choose **Resources** > **Queue Management**. -#. Locate the row where the target queue locates and click **Delete** in the **Operation** column. +#. Locate the row containing the queue to delete and click **Delete** in its **Operation** column. .. note:: diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/testing_the_network_connectivity_between_a_queue_and_a_data_source.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/testing_the_network_connectivity_between_a_queue_and_a_data_source.rst index 258404a..95c8757 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/testing_the_network_connectivity_between_a_queue_and_a_data_source.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/managing_queues/testing_the_network_connectivity_between_a_queue_and_a_data_source.rst @@ -24,6 +24,8 @@ Testing the Address Connectivity Between a Queue and the Data Source - IPv4 + Port number: 192.168.x.x:8080 - Domain name: domain-xxxxxx.com - Domain name + Port number: domain-xxxxxx.com:8080 + - IPv6 address: 2001:0db8:XXXX:XXXX:XXXX:XXXX:XXXX:XXXX + - [IPv6] + Port number: [2001:0db8:XXXX:XXXX:XXXX:XXXX:XXXX:XXXX]:8080 #. Click **Test**. diff --git a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/overview_of_dli_elastic_resource_pools_and_queues.rst b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/overview_of_dli_elastic_resource_pools_and_queues.rst index c8b0004..c8703fb 100644 --- a/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/overview_of_dli_elastic_resource_pools_and_queues.rst +++ b/umn/source/creating_an_elastic_resource_pool_and_creating_queues_within_it/overview_of_dli_elastic_resource_pools_and_queues.rst @@ -50,7 +50,7 @@ Before we dive into the compute resource modes of DLI, let us first understand t - **For SQL:** - For SQL queues are used to execute SQL jobs and supports specifying engine types including Spark and HetuEngine. + For SQL queues are designed to execute SQL jobs. They support the Spark engine. This type of queues is suitable for businesses that require fast data query and analysis, as well as regular cache clearing or environment resetting. @@ -83,7 +83,7 @@ DLI offers three compute resource management modes, each with unique advantages - Use cases: suitable for scenarios with significant fluctuations in business volume, such as periodic data batch processing tasks or real-time data processing needs. - - Supported queue types: for SQL (Spark), for SQL (HetuEngine), and for general purpose. For details about DLI queue types, see :ref:`Queue Types `. + - Supported queue types: for SQL (Spark) and for general purpose. For details about DLI queue types, see :ref:`Queue Types `. .. note:: @@ -139,9 +139,7 @@ DLI Compute Resource Modes and Supported Queue Types +====================================+======================+======================================================================+==========================================================================================================================================================+ | **Elastic resource pool mode** | For SQL (Spark) | Resources are shared among multiple queues for a single user. | Suitable for scenarios with significant fluctuations in business demand, where resources need to be flexibly adjusted to meet peak and off-peak demands. | | | | | | - | | For SQL (HetuEngine) | Resources are dynamically allocated and can be flexibly adjusted. | | - | | | | | - | | For general purpose | | | + | | For general purpose | Resources are dynamically allocated and can be flexibly adjusted. | | +------------------------------------+----------------------+----------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------+ | **Global sharing mode** | default queue | Resources are shared among multiple queues for multiple users. | Suitable for temporary or testing projects where data size is uncertain or data processing is only required occasionally. | | | | | | @@ -241,7 +239,7 @@ The quantities of compute resources required for jobs change in different time o - During the data period from approximately 04:00 to 07:00 in the morning, there are no other jobs running after ETL jobs complete. Since the resources remain constantly occupied, it leads to significant resource wastage. -- Between 09:00 a.m. to 12:00 p.m. and 14:00 p.m. to 16:00 p.m., the volume of ETL report and job query requests is high. Due to insufficient allocated resources, jobs end up queuing continuously. +- Between 09:00 a.m. to 12:00 p.m. and 14:00 to 16:00, the volume of ETL report and job query requests is high. Due to insufficient allocated resources, jobs end up queuing continuously. .. _dli_01_0504__fig6453203515012: diff --git a/umn/source/data_import_and_migration/migrating_data_from_external_data_sources_to_dli/overview_of_data_migration_scenarios.rst b/umn/source/data_import_and_migration/migrating_data_from_external_data_sources_to_dli/overview_of_data_migration_scenarios.rst index 9f9dd4a..5de806b 100644 --- a/umn/source/data_import_and_migration/migrating_data_from_external_data_sources_to_dli/overview_of_data_migration_scenarios.rst +++ b/umn/source/data_import_and_migration/migrating_data_from_external_data_sources_to_dli/overview_of_data_migration_scenarios.rst @@ -25,7 +25,7 @@ Refer to :ref:`Table 1 ` for data type mapping be .. table:: **Table 1** Data type mapping +----------------------------------+------------------------------------+----------------------------------+-------------------------------------+----------------------------------+----------------------------------+------------------------------------+ - | MySQL | Hive | GaussDB(DWS) | Oracle | PostgreSQL | Hologres | DLI Spark | + | MySQL | Hive | DWS | Oracle | PostgreSQL | Hologres | DLI Spark | +==================================+====================================+==================================+=====================================+==================================+==================================+====================================+ | CHAR | CHAR | CHAR | CHAR | CHAR | CHAR | CHAR | +----------------------------------+------------------------------------+----------------------------------+-------------------------------------+----------------------------------+----------------------------------+------------------------------------+ diff --git a/umn/source/dli_job_development_process.rst b/umn/source/dli_job_development_process.rst index 97d19a2..bace989 100644 --- a/umn/source/dli_job_development_process.rst +++ b/umn/source/dli_job_development_process.rst @@ -14,16 +14,16 @@ Creating an IAM User and Granting Permissions - When using DLI for the first time, you need to update the DLI agency according to the console's guidance so that DLI can use other cloud services and perform resource O&M operations on your behalf. The agency includes permissions to obtain IAM user information, access and use VPCs, CIDR blocks, routes, and peering connections, and send notifications via SMN in case of job execution failure. - For more information on the specific permissions included in the agency, refer to :ref:`Configuring DLI Agency Permissions `. + For more information on the specific permissions included in the agency, see :ref:`Configuring DLI Agency Permissions `. Creating Compute Resources and Metadata Required for Running Jobs ----------------------------------------------------------------- -- Before submitting a job using DLI, you need to create an elastic resource pool and create queues within it. This will provide the necessary compute resources for running the job. For how to create an elastic resource pool and create queues within it, see :ref:`Overview of DLI Elastic Resource Pools and Queues `. +- Before submitting a job using DLI, you need to create an elastic resource pool and create queues within it. This will provide the necessary compute resources for running the job. For details about how to create an elastic resource pool and create queues within it, see :ref:`Overview of DLI Elastic Resource Pools and Queues `. Alternatively, you can enhance DLI's computing environment by creating custom images. Specifically, to enhance the functions and performance of Spark and Flink jobs, you can create custom images by downloading the base images provided by DLI and adding dependencies (files, JAR files, or software) and private capabilities required for job execution. This changes the container runtime environment for the jobs. - For example, you can add a Python package or C library related to machine learning to a custom image to help you extend functions. For how to create a custom image, see :ref:`Enhancing the Job Runtime Environment Using a Custom Image `. + For example, you can add a Python package or C library related to machine learning to a custom image to help you extend functions. For details about how to create a custom image, see :ref:`Enhancing the Job Runtime Environment Using a Custom Image `. - DLI metadata is the basis for developing SQL and Spark jobs. Before executing a job, you need to define databases and tables based on your business scenario. @@ -104,4 +104,4 @@ For example, you can monitor the resource usage and job status of a DLI queue. F Using CTS to Audit DLI ---------------------- -With CTS, you can log operations related to DLI, making it easier to search, audit, and trace in the future. For the supported operations, see :ref:`Using CTS to Audit DLI `. +With CTS, you can log operations related to DLI, making it easier to search, audit, and trace in the future. For details about the supported operations, see :ref:`Using CTS to Audit DLI `. diff --git a/umn/source/dli_permission_management/iam_permission_management/overview.rst b/umn/source/dli_permission_management/iam_permission_management/overview.rst index ea0a7c9..fb609f6 100644 --- a/umn/source/dli_permission_management/iam_permission_management/overview.rst +++ b/umn/source/dli_permission_management/iam_permission_management/overview.rst @@ -14,16 +14,6 @@ IAM can authorize different enterprise users to access cloud service resources. For newly created users, they must first log in to DLI once to record metadata before being able to use DLI. -.. table:: **Table 1** IAM authorization types - - +---------------------------------+-------------------------------------+--------------------------+-------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Type | Core Relationship | Permission | Authorization Method | Use Case | - +=================================+=====================================+==========================+===========================================+==================================================================================================================================================================================================================================================================================================+ - | Role/Policy-based authorization | User-permission-authorization scope | - System-defined role | Assigning roles or policies to principals | To authorize a user, you need to add it to a user group first and then specify the scope of authorization. It provides a limited number of condition keys and cannot meet the requirements of fine-grained permissions control. This method is suitable for small- and medium-sized enterprises. | - | | | - System-defined policy | | | - | | | - Custom policy | | | - +---------------------------------+-------------------------------------+--------------------------+-------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - IAM is free to use, and you only need to pay for the resources in your account. If your account does not need individual IAM users for permission management, skip over this section. @@ -31,17 +21,18 @@ If your account does not need individual IAM users for permission management, sk DLI System Permissions ---------------------- -:ref:`Table 2 ` lists all system-defined permissions for DLI. +:ref:`Table 1 ` lists all system-defined permissions for DLI. .. _dli_01_0417__table14662032053: -.. table:: **Table 2** DLI system permissions +.. table:: **Table 1** DLI system permissions - +----------------------------------------------------------------+ - | Type | - +================================================================+ - | System-defined permissions for role/policy-based authorization | - +----------------------------------------------------------------+ + +----------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------+ + | Type | Link | + +================================================================+====================================================================================================================+ + | System-defined permissions for role/policy-based authorization | - :ref:`DLI System Permissions ` | + | | - :ref:`Common Operations Supported by DLI System Policy ` | + +----------------------------------------------------------------+--------------------------------------------------------------------------------------------------------------------+ Permission types: Based on the granularity of authorization, they are divided into roles and policies. diff --git a/umn/source/dli_permission_management/iam_permission_management/using_iam_roles_or_policies_to_grant_access_to_dli.rst b/umn/source/dli_permission_management/iam_permission_management/using_iam_roles_or_policies_to_grant_access_to_dli.rst index 7db6066..0e2aa4c 100644 --- a/umn/source/dli_permission_management/iam_permission_management/using_iam_roles_or_policies_to_grant_access_to_dli.rst +++ b/umn/source/dli_permission_management/iam_permission_management/using_iam_roles_or_policies_to_grant_access_to_dli.rst @@ -7,7 +7,7 @@ Using IAM Roles or Policies to Grant Access to DLI The role/policy-based authorization model provided by `Identity and Access Management (IAM) `__ lets you control access to DLI resources. With IAM, you can: -- Based on the organizational structure of your enterprise, create IAM users in you master account for employees from different departments within the company. This ensures that each employee has unique security credentials and can use DLI resources. +- Based on the organizational structure of your enterprise, create IAM users in your master account for employees from different departments within the company. This ensures that each employee has unique security credentials and can use DLI resources. - Grant users the minimum permissions required to perform specific tasks based on their job responsibilities. - Entrust an account or cloud service to perform professional, efficient O&M on your DLI resources. @@ -18,7 +18,7 @@ If your account does not need individual IAM users, you may skip over this secti Prerequisites ------------- -Before granting permissions to a user group, familiarize yourself with the DLI permissions that can be added to the user group and select them accordingly based on actual needs. +Before granting permissions to a user group, familiarize yourself with the DLI permissions that can be added to the user group and select them as needed. For details about the system permissions of other services, see `Permissions `__. @@ -55,7 +55,7 @@ Process Flow Creating a Custom DLI Policy ---------------------------- -If the predefined DLI permissions in the system do not meet your authorization requirements, you can create custom policies. For the actions that can be added to custom policies, refer to "Permission Policies and Supported Actions" in *Data Lake Insight API Reference*. +If the predefined DLI permissions in the system do not meet your authorization requirements, you can create custom policies. For details about the actions that can be added to custom policies, see "Permission Policies and Supported Actions" in the *Data Lake Insight API Reference*. You can create custom policies in either of the following two ways: @@ -109,7 +109,7 @@ In the following example, an IAM user is granted the permission to create tables A key in the **Condition** element of a statement. There are global and service-specific condition keys. - - Global-level condition key: The prefix is **g:**, which applies to all actions. For details, see the condition key description in Policy Syntax. + - Global-level condition key: The prefix is **g:**, which applies to all actions. For details, see the condition key description in "Policy Syntax". - Service-level condition key: applies only to actions of the specific service. An operator must be used together with a condition key to form a complete condition statement. For details, see :ref:`Table 2 `. @@ -197,35 +197,84 @@ You can set actions and resources at varying levels based on scenarios. .. table:: **Table 5** DLI resources and their paths - +---------------------+-------------------------------------------+------------------------------------------------+ - | Type | Resource | Path | - +=====================+===========================================+================================================+ - | elasticresourcepool | DLI elastic resource pool | elasticresourcepools.name | - +---------------------+-------------------------------------------+------------------------------------------------+ - | queue | DLI queue | queues.queuename | - +---------------------+-------------------------------------------+------------------------------------------------+ - | database | DLI database | databases.dbname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | table | DLI table | databases.dbname.tables.tbname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | column | DLI column | databases.dbname.tables.tbname.columns.colname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | jobs | DLI Flink job | jobs.flink.jobid | - +---------------------+-------------------------------------------+------------------------------------------------+ - | resource | DLI package | resources.resourcename | - +---------------------+-------------------------------------------+------------------------------------------------+ - | group | DLI package group | groups.groupname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | datasourceauth | DLI datasource authentication information | datasourceauth.name | - +---------------------+-------------------------------------------+------------------------------------------------+ - | edsconnections | Enhanced datasource connection | edsconnections.\ *Connection ID* | - +---------------------+-------------------------------------------+------------------------------------------------+ - | variable | DLI global variable | variables.name | - +---------------------+-------------------------------------------+------------------------------------------------+ - | sqldefendrule | SQL inspection rule | sqldefendes.\* | - +---------------------+-------------------------------------------+------------------------------------------------+ - | catalog | DLI data catalog | catalogs.name | - +---------------------+-------------------------------------------+------------------------------------------------+ + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Type | Resource | Path | Description | + +=====================+===========================================+===================================================================================+=========================================================================================================================================================================================================================================+ + | elasticresourcepool | DLI elastic resource pool | DLI:``*``:``*``:elasticresourcepool:elasticresourcepools\ *.name* | - The path prefix for **elasticresourcepool** is fixed as **DLI:*:*:elasticresourcepool:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:elasticresourcepool:elasticresourcepools.\*** indicates any DLI elastic resource pool. | + | | | | | + | | | | **DLI:*:*:elasticresourcepool:elasticresourcepools.pool01** indicates a specific DLI elastic resource pool named **pool01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | queue | DLI queue | DLI:``*``:``*``:queue:queues\ *.queuename* | - The path prefix for **queue** is fixed as **DLI:*:*:queue:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:queue:queues.\*** indicates any DLI queue. | + | | | | | + | | | | **DLI:*:*:queue:queues.queue01** indicates a specific DLI queue named **queue01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | database | DLI database | DLI:``*``:``*``:database:databases\ *.dbname* | - The path prefix for **database** is fixed as **DLI:*:*:database:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:database:databases.\*** indicates any DLI database. | + | | | | | + | | | | **DLI:*:*:database:databases.db01** indicates a specific DLI database named **db01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | table | DLI table | DLI:``*``:``*``:table:databases.\ *dbname*.tables\ *.tbname* | - The path prefix for **table** is fixed as **DLI:*:*:table:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:able:databases..tables.\*** indicates any DLI table. | + | | | | | + | | | | **DLI:*:*:table:databases.db01.tables.tb01** indicates a DLI table named **tb01** in the **db01** database. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | column | DLI column | DLI:``*``:``*``:column:databases.\ *dbname*.tables.\ *tbname*.columns\ *.colname* | - The path prefix for **column** is fixed as **DLI:*:*:column:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:column:databases..tables..columns.\*** indicates any DLI column. | + | | | | | + | | | | **DLI:*:*:column:databases.db01.tables.tb01.columns.col01** indicates a DLI column named **col01** in the **tb01** table of the **db01** database. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | jobs | DLI Flink job | DLI:``*``:``*``:jobs:jobs.flink\ *.jobid* | - The path prefix for **jobs** is fixed as **DLI:*:*:jobs:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:jobs:jobs.flink.\*** indicates any DLI Flink job. | + | | | | | + | | | | **DLI:*:*:jobs:jobs.flink.123456** indicates a DLI Flink job with ID of **123456**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | resource | DLI package | DLI:``*``:``*``:resource:resources.\ *resourcename* | - The path prefix for **resource** is fixed as **DLI:*:*:resource:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:resource:resources.\*** indicates any DLI package. | + | | | | | + | | | | **DLI:*:*:resource:resources.jar01** indicates a DLI package named **jar01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | group | DLI package group | DLI:``*``:``*``:group:groups.\ *groupname* | - The path prefix for **group** is fixed as **DLI:*:*:group:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:group:groups.\*** indicates any DLI package group. | + | | | | | + | | | | **DLI:*:*:group:groups.group01** indicates a DLI package group named **group01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | datasourceauth | DLI datasource authentication information | DLI:``*``:``*``:datasourceauth:datasourceauth.\ *name* | - The path prefix for **datasourceauth** is fixed as **DLI:*:*:datasourceauth:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:datasourceauth:datasourceauth.\*** indicates any DLI datasource authentication information. | + | | | | | + | | | | **DLI:*:*:datasourceauth:datasourceauth.auth01** indicates DLI datasource authentication information named **auth01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | edsconnections | Enhanced datasource connection | DLI:``*``:``*``:edsconnections:edsconnections.\ *Connection ID* | - The path prefix for **edsconnections** is fixed as **DLI:*:*:edsconnections:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:edsconnections:edsconnections.\*** indicates any DLI enhanced datasource connection. | + | | | | | + | | | | **DLI:*:*:edsconnections:edsconnections.conn01** indicates a DLI enhanced datasource connection with connection ID of **conn01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | variable | DLI global variable | DLI:``*``:``*``:variable:variables.\ *name* | - The path prefix for **variable** is fixed as **DLI:*:*:variable:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:variable:variables.\*** indicates any DLI global variable. | + | | | | | + | | | | **DLI:*:*:variable:variables.var01** indicates a DLI global variable named **var01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | sqldefendrule | SQL inspection rule | DLI:``*``:``*``:sqldefendrule:sqldefendes.\ ``*`` | - The path prefix for **sqldefendrule** is fixed as **DLI:*:*:sqldefendrule:**. | + | | | | - Supports wildcard (*): **DLI:*:*:sqldefendrule:sqldefendes.\*** indicates all SQL inspection rules. (The **sqldefendrule** resource path inherently includes wildcard characters, so there is no need to specify them additionally.) | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | catalog | DLI data catalog | DLI:``*``:``*``:catalog:catalogs.\ *name* | - The path prefix for **catalog** is fixed as **DLI:*:*:catalog:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:catalog:catalogs.\*** indicates any DLI data catalog. | + | | | | | + | | | | **DLI:*:*:catalog:catalogs.cat01** indicates a DLI data catalog named **cat01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ #. Combine all of the preceding fields into a JSON string to create a complete policy. You can set multiple actions and resources, and you also have the option to create policies through the IAM console. For example: @@ -352,35 +401,84 @@ Resources are objects that exist within a service. In DLI, resources include the .. table:: **Table 6** DLI resources and their paths - +---------------------+-------------------------------------------+------------------------------------------------+ - | Type | Resource | Path | - +=====================+===========================================+================================================+ - | elasticresourcepool | DLI elastic resource pool | elasticresourcepools.name | - +---------------------+-------------------------------------------+------------------------------------------------+ - | queue | DLI queue | queues.queuename | - +---------------------+-------------------------------------------+------------------------------------------------+ - | database | DLI database | databases.dbname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | table | DLI table | databases.dbname.tables.tbname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | column | DLI column | databases.dbname.tables.tbname.columns.colname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | jobs | DLI Flink job | jobs.flink.jobid | - +---------------------+-------------------------------------------+------------------------------------------------+ - | resource | DLI package | resources.resourcename | - +---------------------+-------------------------------------------+------------------------------------------------+ - | group | DLI package group | groups.groupname | - +---------------------+-------------------------------------------+------------------------------------------------+ - | datasourceauth | DLI datasource authentication information | datasourceauth.name | - +---------------------+-------------------------------------------+------------------------------------------------+ - | edsconnections | Enhanced datasource connection | edsconnections.\ *Connection ID* | - +---------------------+-------------------------------------------+------------------------------------------------+ - | variable | DLI global variable | variables.name | - +---------------------+-------------------------------------------+------------------------------------------------+ - | sqldefendrule | SQL inspection rule | sqldefendes.\* | - +---------------------+-------------------------------------------+------------------------------------------------+ - | catalog | DLI data catalog | catalogs.name | - +---------------------+-------------------------------------------+------------------------------------------------+ + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Type | Resource | Path | Description | + +=====================+===========================================+===================================================================================+=========================================================================================================================================================================================================================================+ + | elasticresourcepool | DLI elastic resource pool | DLI:``*``:``*``:elasticresourcepool:elasticresourcepools\ *.name* | - The path prefix for **elasticresourcepool** is fixed as **DLI:*:*:elasticresourcepool:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:elasticresourcepool:elasticresourcepools.\*** indicates any DLI elastic resource pool. | + | | | | | + | | | | **DLI:*:*:elasticresourcepool:elasticresourcepools.pool01** indicates a specific DLI elastic resource pool named **pool01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | queue | DLI queue | DLI:``*``:``*``:queue:queues\ *.queuename* | - The path prefix for **queue** is fixed as **DLI:*:*:queue:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:queue:queues.\*** indicates any DLI queue. | + | | | | | + | | | | **DLI:*:*:queue:queues.queue01** indicates a specific DLI queue named **queue01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | database | DLI database | DLI:``*``:``*``:database:databases\ *.dbname* | - The path prefix for **database** is fixed as **DLI:*:*:database:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:database:databases.\*** indicates any DLI database. | + | | | | | + | | | | **DLI:*:*:database:databases.db01** indicates a specific DLI database named **db01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | table | DLI table | DLI:``*``:``*``:table:databases.\ *dbname*.tables\ *.tbname* | - The path prefix for **table** is fixed as **DLI:*:*:table:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:able:databases..tables.\*** indicates any DLI table. | + | | | | | + | | | | **DLI:*:*:table:databases.db01.tables.tb01** indicates a DLI table named **tb01** in the **db01** database. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | column | DLI column | DLI:``*``:``*``:column:databases.\ *dbname*.tables.\ *tbname*.columns\ *.colname* | - The path prefix for **column** is fixed as **DLI:*:*:column:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:column:databases..tables..columns.\*** indicates any DLI column. | + | | | | | + | | | | **DLI:*:*:column:databases.db01.tables.tb01.columns.col01** indicates a DLI column named **col01** in the **tb01** table of the **db01** database. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | jobs | DLI Flink job | DLI:``*``:``*``:jobs:jobs.flink\ *.jobid* | - The path prefix for **jobs** is fixed as **DLI:*:*:jobs:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:jobs:jobs.flink.\*** indicates any DLI Flink job. | + | | | | | + | | | | **DLI:*:*:jobs:jobs.flink.123456** indicates a DLI Flink job with ID of **123456**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | resource | DLI package | DLI:``*``:``*``:resource:resources.\ *resourcename* | - The path prefix for **resource** is fixed as **DLI:*:*:resource:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:resource:resources.\*** indicates any DLI package. | + | | | | | + | | | | **DLI:*:*:resource:resources.jar01** indicates a DLI package named **jar01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | group | DLI package group | DLI:``*``:``*``:group:groups.\ *groupname* | - The path prefix for **group** is fixed as **DLI:*:*:group:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:group:groups.\*** indicates any DLI package group. | + | | | | | + | | | | **DLI:*:*:group:groups.group01** indicates a DLI package group named **group01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | datasourceauth | DLI datasource authentication information | DLI:``*``:``*``:datasourceauth:datasourceauth.\ *name* | - The path prefix for **datasourceauth** is fixed as **DLI:*:*:datasourceauth:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:datasourceauth:datasourceauth.\*** indicates any DLI datasource authentication information. | + | | | | | + | | | | **DLI:*:*:datasourceauth:datasourceauth.auth01** indicates DLI datasource authentication information named **auth01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | edsconnections | Enhanced datasource connection | DLI:``*``:``*``:edsconnections:edsconnections.\ *Connection ID* | - The path prefix for **edsconnections** is fixed as **DLI:*:*:edsconnections:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:edsconnections:edsconnections.\*** indicates any DLI enhanced datasource connection. | + | | | | | + | | | | **DLI:*:*:edsconnections:edsconnections.conn01** indicates a DLI enhanced datasource connection with connection ID of **conn01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | variable | DLI global variable | DLI:``*``:``*``:variable:variables.\ *name* | - The path prefix for **variable** is fixed as **DLI:*:*:variable:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:variable:variables.\*** indicates any DLI global variable. | + | | | | | + | | | | **DLI:*:*:variable:variables.var01** indicates a DLI global variable named **var01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | sqldefendrule | SQL inspection rule | DLI:``*``:``*``:sqldefendrule:sqldefendes.\ ``*`` | - The path prefix for **sqldefendrule** is fixed as **DLI:*:*:sqldefendrule:**. | + | | | | - Supports wildcard (*): **DLI:*:*:sqldefendrule:sqldefendes.\*** indicates all SQL inspection rules. (The **sqldefendrule** resource path inherently includes wildcard characters, so there is no need to specify them additionally.) | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | catalog | DLI data catalog | DLI:``*``:``*``:catalog:catalogs.\ *name* | - The path prefix for **catalog** is fixed as **DLI:*:*:catalog:**. | + | | | | | + | | | | - Supports wildcard (*): **DLI:*:*:catalog:catalogs.\*** indicates any DLI data catalog. | + | | | | | + | | | | **DLI:*:*:catalog:catalogs.cat01** indicates a DLI data catalog named **cat01**. | + +---------------------+-------------------------------------------+-----------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ DLI Request Condition --------------------- diff --git a/umn/source/faq/dli_apis/index.rst b/umn/source/faq/dli_apis/index.rst index 2d18a6f..6ce81fe 100644 --- a/umn/source/faq/dli_apis/index.rst +++ b/umn/source/faq/dli_apis/index.rst @@ -5,12 +5,12 @@ DLI APIs ======== -- :ref:`Why Is Error "unsupported media Type" Reported When I Subimt a SQL Job? ` +- :ref:`Why Am I Getting an "unsupported media Type" Error When Submitting a SQL Job? ` - :ref:`What Can I Do If an Error Is Reported When the Execution of the API for Creating a SQL Job Times Out? ` .. toctree:: :maxdepth: 1 :hidden: - why_is_error_unsupported_media_type_reported_when_i_subimt_a_sql_job + why_am_i_getting_an_unsupported_media_type_error_when_submitting_a_sql_job what_can_i_do_if_an_error_is_reported_when_the_execution_of_the_api_for_creating_a_sql_job_times_out diff --git a/umn/source/faq/dli_apis/why_is_error_unsupported_media_type_reported_when_i_subimt_a_sql_job.rst b/umn/source/faq/dli_apis/why_am_i_getting_an_unsupported_media_type_error_when_submitting_a_sql_job.rst similarity index 85% rename from umn/source/faq/dli_apis/why_is_error_unsupported_media_type_reported_when_i_subimt_a_sql_job.rst rename to umn/source/faq/dli_apis/why_am_i_getting_an_unsupported_media_type_error_when_submitting_a_sql_job.rst index 8241a29..cc0362b 100644 --- a/umn/source/faq/dli_apis/why_is_error_unsupported_media_type_reported_when_i_subimt_a_sql_job.rst +++ b/umn/source/faq/dli_apis/why_am_i_getting_an_unsupported_media_type_error_when_submitting_a_sql_job.rst @@ -2,8 +2,8 @@ .. _dli_03_0060: -Why Is Error "unsupported media Type" Reported When I Subimt a SQL Job? -======================================================================= +Why Am I Getting an "unsupported media Type" Error When Submitting a SQL Job? +============================================================================= In the REST API provided by DLI, the request header can be added to the request URI, for example, **Content-Type**. diff --git a/umn/source/faq/dli_databases_and_tables/do_i_need_to_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name.rst b/umn/source/faq/dli_databases_and_tables/do_i_need_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name.rst similarity index 75% rename from umn/source/faq/dli_databases_and_tables/do_i_need_to_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name.rst rename to umn/source/faq/dli_databases_and_tables/do_i_need_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name.rst index f35fbff..f24d441 100644 --- a/umn/source/faq/dli_databases_and_tables/do_i_need_to_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name.rst +++ b/umn/source/faq/dli_databases_and_tables/do_i_need_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name.rst @@ -2,8 +2,8 @@ .. _dli_03_0175: -Do I Need to to Regrant Permissions to Users and Projects After Deleting and Recreating a Table With the Same Name? -=================================================================================================================== +Do I Need to Regrant Permissions to Users and Projects After Deleting and Recreating a Table With the Same Name? +================================================================================================================ Scenario -------- @@ -13,7 +13,7 @@ User A created the **testTable** table in a database through a SQL job and grant Possible Causes --------------- -After a table is deleted, the table permissions are not retained. You need to grant permissions to a user or project. +After deleting a table and recreating one with the same name, table permissions do not automatically inherit in this scenario. You need to regrant permissions to the users or projects. Solution -------- diff --git a/umn/source/faq/dli_databases_and_tables/index.rst b/umn/source/faq/dli_databases_and_tables/index.rst index cdc6192..15afbff 100644 --- a/umn/source/faq/dli_databases_and_tables/index.rst +++ b/umn/source/faq/dli_databases_and_tables/index.rst @@ -8,7 +8,7 @@ DLI Databases and Tables - :ref:`Why Am I Unable to Query a Table on the DLI Console? ` - :ref:`How Do I Do If the Compression Rate of an OBS Table Is High? ` - :ref:`How Do I Do If Inconsistent Character Encoding Leads to Garbled Characters? ` -- :ref:`Do I Need to to Regrant Permissions to Users and Projects After Deleting and Recreating a Table With the Same Name? ` +- :ref:`Do I Need to Regrant Permissions to Users and Projects After Deleting and Recreating a Table With the Same Name? ` - :ref:`How Do I Do If Files Imported Into a DLI Partitioned Table Lack Data for the Partition Columns, Causing Query Failures After the Import Is Completed? ` - :ref:`How Do I Fix Incorrect Data in an OBS Foreign Table Caused by Newline Characters in OBS File Fields? ` - :ref:`How Do I Prevent a Cartesian Product Query and Resource Overload Due to Missing "ON" Conditions in Table Joins? ` @@ -26,7 +26,7 @@ DLI Databases and Tables why_am_i_unable_to_query_a_table_on_the_dli_console how_do_i_do_if_the_compression_rate_of_an_obs_table_is_high how_do_i_do_if_inconsistent_character_encoding_leads_to_garbled_characters - do_i_need_to_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name + do_i_need_to_regrant_permissions_to_users_and_projects_after_deleting_and_recreating_a_table_with_the_same_name how_do_i_do_if_files_imported_into_a_dli_partitioned_table_lack_data_for_the_partition_columns_causing_query_failures_after_the_import_is_completed how_do_i_fix_incorrect_data_in_an_obs_foreign_table_caused_by_newline_characters_in_obs_file_fields how_do_i_prevent_a_cartesian_product_query_and_resource_overload_due_to_missing_on_conditions_in_table_joins diff --git a/umn/source/faq/dli_databases_and_tables/why_am_i_unable_to_query_a_table_on_the_dli_console.rst b/umn/source/faq/dli_databases_and_tables/why_am_i_unable_to_query_a_table_on_the_dli_console.rst index f737a0d..1a3a37c 100644 --- a/umn/source/faq/dli_databases_and_tables/why_am_i_unable_to_query_a_table_on_the_dli_console.rst +++ b/umn/source/faq/dli_databases_and_tables/why_am_i_unable_to_query_a_table_on_the_dli_console.rst @@ -20,7 +20,7 @@ Solution Contact the user who creates the table and obtain the required permissions. To assign permissions, perform the following steps: -#. Log in to the DLI management console as the user who creates the table. Choose **Data Management** > **Databases and Tables** form the navigation pane on the left. +#. Log in to the DLI management console as the user who creates the table. In the navigation pane on the left, choose **Data Management** > **Databases and Tables**. #. Click the database name. The table management page is displayed. In the **Operation** column of the target table, click **Permissions**. The table permission management page is displayed. #. Click **Set Permission**. In the displayed dialog box, set **Authorization Object** to **User**, set **Username** to the name of the user that requires the permission, and select the required permissions. For example, **Select Table** and **Insert** permissions. #. Click **OK**. diff --git a/umn/source/faq/dli_elastic_resource_pools_and_queues/how_do_i_monitor_job_exceptions_on_a_dli_queue.rst b/umn/source/faq/dli_elastic_resource_pools_and_queues/how_do_i_monitor_job_exceptions_on_a_dli_queue.rst index 01a5e35..dd33ee8 100644 --- a/umn/source/faq/dli_elastic_resource_pools_and_queues/how_do_i_monitor_job_exceptions_on_a_dli_queue.rst +++ b/umn/source/faq/dli_elastic_resource_pools_and_queues/how_do_i_monitor_job_exceptions_on_a_dli_queue.rst @@ -9,4 +9,4 @@ DLI allows you to subscribe to an SMN topic for failed jobs. #. Log in to the DLI console. #. In the navigation pane on the left, choose **Queue Management**. -#. On the **Queue Management** page, click **Create SMN Topic** in the upper left corner. . +#. On the **Queue Management** page, click **Create SMN Topic** in the upper left corner. diff --git a/umn/source/faq/dli_permissions_management/index.rst b/umn/source/faq/dli_permissions_management/index.rst index 4ac88a9..07d4815 100644 --- a/umn/source/faq/dli_permissions_management/index.rst +++ b/umn/source/faq/dli_permissions_management/index.rst @@ -10,7 +10,7 @@ DLI Permissions Management - :ref:`Why Is Error "DLI.0003: Permission denied for resource..." Reported When I Run a SQL Statement? ` - :ref:`How Do I Do If I Can't Query Table Data After Being Granted Table Permissions? ` - :ref:`Will Granting Duplicate Permissions to a Table After Inheriting Database Permissions Cause an Error? ` -- :ref:`Why Can't I Query a View After I'm Granted the Select Table Permission on the View? ` +- :ref:`Why Can't I Query a View Despite Having the Select Permission? ` - :ref:`How Do I Do If I Receive a Message Saying I Don't Have Sufficient Permissions to Submit My Jobs to the Job Bucket? ` .. toctree:: @@ -22,5 +22,5 @@ DLI Permissions Management why_is_error_dli.0003_permission_denied_for_resource..._reported_when_i_run_a_sql_statement how_do_i_do_if_i_cant_query_table_data_after_being_granted_table_permissions will_granting_duplicate_permissions_to_a_table_after_inheriting_database_permissions_cause_an_error - why_cant_i_query_a_view_after_im_granted_the_select_table_permission_on_the_view + why_cant_i_query_a_view_despite_having_the_select_permission how_do_i_do_if_i_receive_a_message_saying_i_dont_have_sufficient_permissions_to_submit_my_jobs_to_the_job_bucket diff --git a/umn/source/faq/dli_permissions_management/why_cant_i_query_a_view_after_im_granted_the_select_table_permission_on_the_view.rst b/umn/source/faq/dli_permissions_management/why_cant_i_query_a_view_after_im_granted_the_select_table_permission_on_the_view.rst deleted file mode 100644 index 7be11db..0000000 --- a/umn/source/faq/dli_permissions_management/why_cant_i_query_a_view_after_im_granted_the_select_table_permission_on_the_view.rst +++ /dev/null @@ -1,25 +0,0 @@ -:original_name: dli_03_0067.html - -.. _dli_03_0067: - -Why Can't I Query a View After I'm Granted the Select Table Permission on the View? -=================================================================================== - -Symptom -------- - -User A created Table1. - -User B created View1 based on Table1. - -After the **Select Table** permission on Table1 is granted to user C, user C fails to query View1. - -Possible Causes ---------------- - -User B does not have the **Select Table** permission on Table1. - -Solution --------- - -Grant the **Select Table** permission on Table1 to user B. Then, query View1 as user C again. diff --git a/umn/source/faq/dli_permissions_management/why_cant_i_query_a_view_despite_having_the_select_permission.rst b/umn/source/faq/dli_permissions_management/why_cant_i_query_a_view_despite_having_the_select_permission.rst new file mode 100644 index 0000000..0167341 --- /dev/null +++ b/umn/source/faq/dli_permissions_management/why_cant_i_query_a_view_despite_having_the_select_permission.rst @@ -0,0 +1,44 @@ +:original_name: dli_03_0067.html + +.. _dli_03_0067: + +Why Can't I Query a View Despite Having the Select Permission? +============================================================== + +Symptom +------- + +User A created a table named **Table1**. + +User B created a view named **View1** based on **Table1**. + +After granting user C the select permission on **Table1**, user C failed to query the view. + +Current Permission Assignments +------------------------------ + +- User A already has: **admin** permission on **Table1**. +- User B already has: **admin** permission on **View1**. +- User C already has: **select** permission on **Table1**. + +Solution +-------- + +Different versions of the Spark engine have varying permission requirements for views: + +- **Spark 2.4.**\ *x*: User C needs the view query permission, and user B requires the select permission on **Table1**. + + Permission requirements: + + - User B should have: admin permission on **View1** and select permission on **Table1** (currently missing). + - User C already has: select permission on **Table1**. + + To resolve this, grant user B the select permission on **Table1**, then retry querying **View1** as user C. + +- **Spark 3.3.**\ *x*: User C needs the view query permission, and user C requires the select permission on **Table1**. + + Permission requirements: + + - User C should have: select permission on **Table1** and select permission on **View1** (currently missing). + + To resolve this, grant user C the select permission on **View1**, then retry querying **View1** as user C. diff --git a/umn/source/faq/dli_resource_quotas/how_do_i_apply_for_a_higher_quota.rst b/umn/source/faq/dli_resource_quotas/how_do_i_apply_for_a_higher_quota.rst index 6d71428..1e380eb 100644 --- a/umn/source/faq/dli_resource_quotas/how_do_i_apply_for_a_higher_quota.rst +++ b/umn/source/faq/dli_resource_quotas/how_do_i_apply_for_a_higher_quota.rst @@ -5,13 +5,12 @@ How Do I Apply for a Higher Quota? ================================== +How Do I Apply for a Quota Increase? +------------------------------------ -How Do I Apply for a Higher Quota? ----------------------------------- - -The system does not support online quota adjustment. To increase a resource quota, dial the hotline or send an email to the customer service. We will process your application and inform you of the progress by phone call or email. +The system currently does not support online quota adjustments. If you need to modify your quota, please contact our customer service team via phone or email. They will promptly process your request and keep you updated on its progress through either a call or email. -Before dialing the hotline number or sending an email, ensure that the following information has been obtained: +Before reaching out, ensure you have the following details ready: - Domain name, project name, and project ID @@ -21,6 +20,6 @@ Before dialing the hotline number or sending an email, ensure that the following - Service name - Quota type - - Required quota + - Desired quota value `Learn how to obtain the service hotline and email address. `__ diff --git a/umn/source/faq/dli_resource_quotas/how_do_i_view_my_quotas.rst b/umn/source/faq/dli_resource_quotas/how_do_i_view_my_quotas.rst index 9b41fbf..5c9c5e5 100644 --- a/umn/source/faq/dli_resource_quotas/how_do_i_view_my_quotas.rst +++ b/umn/source/faq/dli_resource_quotas/how_do_i_view_my_quotas.rst @@ -11,11 +11,11 @@ How Do I View My Quotas? #. Click the **My Quota** icon |image2| in the upper right corner of the page. - The **Service Quota** page is displayed. + This will take you to the **Service Quota** page. -#. View the used and total quota of each type of resources on the displayed page. +#. Here, you can view the total quota and usage details for each resource. - If a quota cannot meet service requirements, increase a quota. + If your current quota is insufficient, follow the steps below to request an increase. .. |image1| image:: /_static/images/en-us_image_0000001487683840.png .. |image2| image:: /_static/images/en-us_image_0000001488163536.png diff --git a/umn/source/faq/flink_jobs/flink_sql_jobs/why_does_dis_stream_not_exist_during_job_semantic_check.rst b/umn/source/faq/flink_jobs/flink_sql_jobs/why_does_dis_stream_not_exist_during_job_semantic_check.rst index d4ac151..a73533d 100644 --- a/umn/source/faq/flink_jobs/flink_sql_jobs/why_does_dis_stream_not_exist_during_job_semantic_check.rst +++ b/umn/source/faq/flink_jobs/flink_sql_jobs/why_does_dis_stream_not_exist_during_job_semantic_check.rst @@ -9,7 +9,7 @@ To rectify this fault, perform the following steps: #. Log in to the DIS management console. In the navigation pane, choose **Stream Management**. View the Flink job SQL statements to check whether the DIS stream exists. -#. If the DIS stream was not created, create a DIS stream by referring to "Creating a DIS Stream" in the Data Ingestion Service User Guide. +#. If the DIS stream in the Flink job has not been created yet, create one by referring to "Creating a DIS Stream" in *Data Lake Insight User Guide*. Ensure that the created DIS stream and Flink job are in the same region. diff --git a/umn/source/faq/spark_jobs/spark_job_development/how_do_i_use_python_scripts_to_access_the_mysql_database_if_the_pymysql_module_is_missing_from_the_spark_job_results_stored_in_mysql.rst b/umn/source/faq/spark_jobs/spark_job_development/how_do_i_use_python_scripts_to_access_the_mysql_database_if_the_pymysql_module_is_missing_from_the_spark_job_results_stored_in_mysql.rst index a4e4417..d1b5f75 100644 --- a/umn/source/faq/spark_jobs/spark_job_development/how_do_i_use_python_scripts_to_access_the_mysql_database_if_the_pymysql_module_is_missing_from_the_spark_job_results_stored_in_mysql.rst +++ b/umn/source/faq/spark_jobs/spark_job_development/how_do_i_use_python_scripts_to_access_the_mysql_database_if_the_pymysql_module_is_missing_from_the_spark_job_results_stored_in_mysql.rst @@ -21,6 +21,6 @@ How Do I Use Python Scripts to Access the MySQL Database If the pymysql Module I #. To interconnect PySpark jobs with MySQL, you need to create a datasource connection to enable the network between DLI and RDS. - For how to create a datasource connection on the management console, see "Enhanced Datasource Connections" in *Data Lake Insight User Guide*. + For details about how to create a datasource connection on the management console, see "Enhanced Datasource Connections" in *Data Lake Insight User Guide*. - For how to call an API to create a datasource connection, see "Creating an Enhanced Datasource Connection" in *Data Lake Insight API Reference*. + For details about how to call an API to create a datasource connection, see "Creating an Enhanced Datasource Connection" in *Data Lake Insight API Reference*. diff --git a/umn/source/faq/spark_jobs/spark_job_o_and_m/why_do_i_get_responsecode_403_and_responsestatus_forbidden_errors_when_a_spark_job_accesses_obs_data.rst b/umn/source/faq/spark_jobs/spark_job_o_and_m/why_do_i_get_responsecode_403_and_responsestatus_forbidden_errors_when_a_spark_job_accesses_obs_data.rst index 64054a2..97e92e4 100644 --- a/umn/source/faq/spark_jobs/spark_job_o_and_m/why_do_i_get_responsecode_403_and_responsestatus_forbidden_errors_when_a_spark_job_accesses_obs_data.rst +++ b/umn/source/faq/spark_jobs/spark_job_o_and_m/why_do_i_get_responsecode_403_and_responsestatus_forbidden_errors_when_a_spark_job_accesses_obs_data.rst @@ -12,7 +12,7 @@ The following error is reported when a Spark job accesses OBS data: .. code-block:: - Caused by: com.obs.services.exception.ObsException: Error message:Request Error.OBS servcie Error Message. -- ResponseCode: 403, ResponseStatus: Forbidden + Caused by: com.obs.services.exception.ObsException: Error message:Request Error.OBS service Error Message. -- ResponseCode: 403, ResponseStatus: Forbidden Solution -------- diff --git a/umn/source/faq/sql_jobs/sql_job_development/how_do_i_merge_small_files.rst b/umn/source/faq/sql_jobs/sql_job_development/how_do_i_merge_small_files.rst index 5965685..b95eda3 100644 --- a/umn/source/faq/sql_jobs/sql_job_development/how_do_i_merge_small_files.rst +++ b/umn/source/faq/sql_jobs/sql_job_development/how_do_i_merge_small_files.rst @@ -5,14 +5,271 @@ How Do I Merge Small Files? =========================== -If a large number of small files are generated during SQL execution, job execution and table query will take a long time. In this case, you should merge small files. +What Are Small Files? +--------------------- -You are advised to use temporary tables for data transfer. There is a risk of data loss in self-read and self-write operations during unexpected exceptional scenarios. +In distributed file systems, data is stored in blocks. A small file refers to a file whose size is significantly smaller than the block size of the storage system. Large numbers of small files can cause significant performance and management issues for big data systems. Merging small files is one of the key strategies to optimize system performance. -Run the following SQL statements: +This section explains how to use **DISTRIBUTE BY** in DLI to merge small files. + +Key Issues Caused by Small Files +-------------------------------- + +- **Metadata overhead**: In file systems like HDFS, the master node (NameNode) manages metadata for each file in memory. A massive number of small files drastically increases memory consumption, becoming a bottleneck for cluster scalability. +- **Poor computational performance**: Compute engines like Spark typically launch a task for each file or file block. Processing numerous small files leads to: + + - **High task scheduling overhead**: The time spent launching and scheduling thousands of tasks may exceed the actual data processing time. + - **Resource wastage**: Significant CPU and memory resources are consumed by task management rather than effective computation. + - **Query latency**: Overall job execution time increases, slowing down query responses. + +- **Low storage efficiency**: Small files often result in inefficient storage utilization, failing to leverage the advantages of distributed block sizes. +- **Risk of limits being reached**: They may trigger limits on the number of files in a directory or the number of tasks per job in compute engines or file systems, causing job failures. + +Basic Principles of Small File Merging +-------------------------------------- + +- **Basic Principles** + + In MapReduce-based engines like Spark or Hive, the number of output files depends on the number of tasks in the final stage (usually the Reduce phase). Each Reduce task generates one file by default. + + By using **DISTRIBUTE BY**, data is redistributed into a specified number of Reduce tasks, with each task producing one file, thereby merging files. + + - **Data redistribution**: Data is sent to different Reduce tasks based on the hash value of the expression that follows (e.g., **floor(rand()*N)**). + - **Controlling Reduce task count**: By setting the value of **N**, you indirectly and probabilistically control the number of Reduce tasks, which determines the number of output files. Typically, N represents the desired number of merged files per partition. + +Non-Partitioned Table Merging Example +------------------------------------- + +- **Using a temporary table for intermediate storage** (recommended) + + :: + + -- 1. Create a temporary table with the same structure as the source table. + CREATE TABLE temp_tablename LIKE tablename; + + -- 2. Write merged data into the temporary table. + INSERT OVERWRITE TABLE temp_tablename + SELECT * FROM tablename + DISTRIBUTE BY floor(rand()*20); -- Generate approximately 20 files. + + -- 3. Verify the data. + SELECT count(*) FROM tablename; -- Source table + SELECT count(*) FROM temp_tablename; -- Temporary table + + -- 4. Write the data back to the source table. + INSERT OVERWRITE TABLE tablename + SELECT + * + FROM temp_tablename; + + -- 5. (Optional) Delete the old table if no longer needed. + DROP TABLE table_name_old; + +- **Self read-write method** + + .. code-block:: + + INSERT OVERWRITE TABLE tablename + SELECT * + FROM tablename + DISTRIBUTE BY floor(rand() * 20) + + - **INSERT OVERWRITE TABLE tablename** + + Overwrite the target table with query results. + + .. caution:: + + This operation deletes the source table data, so consider using a temporary table to avoid data loss. + + - **SELECT \* FROM tablename** + + Read all data from the source table. + + - **DISTRIBUTE BY floor(rand() \* 20)** + + - Random data distribution: + + - **rand()** generates a random float between [0, 1). + - **rand() \* 20** generates a random float between [0, 20). + - **floor(...)** rounds down to integers in [0, 19]. + + - Distribution result: + + - Data is randomly and evenly distributed across up to 20 Reduce tasks. + - Up to 20 output files are generated (each Reduce task corresponds to one file). + +Partitioned Table Merging +------------------------- + +For partitioned tables, we generally prefer to merge files within each partition separately rather than mixing data across partitions. + +- **Parameter description:** + + - **N**: The desired number of files **per partition** after merging. For example, set it to **5** if you wish for each partition to ultimately contain only five files. + - **pt**: The partition field. + +- **Why do I need to add a partition field to DISTRIBUTE BY?** + + - This ensures **data from the same partition is sent to the same group of Reduce tasks** for processing. + + - Using **DISTRIBUTE BY pt, floor(rand() \* N)** first distributes data by the partition field (pt) and then further shuffles it within each partition based on randomness. + + Each partition's data is split into *N* files without affecting other partitions. + +- **Using a temporary table** (recommended) + + :: + + -- 1. Create a temporary table with the same structure as the source table. + CREATE TABLE table_tmp LIKE table_name; + + -- 2. Merge partitions into the temporary table. + INSERT OVERWRITE TABLE table_tmp PARTITION(pt) + SELECT col1, col2, ..., pt + FROM tablename + WHERE pt = * -- Partition conditions + DISTRIBUTE BY pt, floor(rand() * N); -- N indicates the number of target files per partition. + + -- 3. Verify the data. + SELECT COUNT(*) FROM tablename; -- Source table + SELECT COUNT(*) FROM temp_tablename; -- Temporary table + + -- 4. Write the data back to the source table. + INSERT OVERWRITE TABLE tablename PARTITION (pt) + SELECT + * + FROM temp_tablename + WHERE pt = *; + + -- 5. (Optional) Delete the old table if no longer needed. + DROP TABLE temp_tablename; + +- **Self read-write method** + + :: + + INSERT OVERWRITE TABLE target_table + PARTITION (pt) + SELECT + col1, + col2, + ..., + pt -- The partition field must appear last in the SELECT statement. + FROM + target_table + WHERE pt = * + DISTRIBUTE BY pt, floor(rand() * N) + +Example: Merging Partitions in a Student Table +---------------------------------------------- + +- **Scenario** + + Consider a partitioned table named **student**, with partition fields **facultyNo** (college ID) and **classNo** (class ID). The goal is to merge small files within each partition, limiting each partition to a maximum of 2 files. + + - Table name: **student** + + - Partition fields: **facultyNo** (college ID) and **classNo** (class ID) + + - Goal: Merge each partition into 2 files. + +- **Example code** + + .. code-block:: + + -- 1. Create a temporary table. + CREATE TABLE student_tmp LIKE student; + + -- 2. Merge specific partitions. + INSERT OVERWRITE TABLE student_tmp PARTITION(facultyNo, classNo) + SELECT + id, + name, + gender, + age, + birth_date, + phone, + email, + address, + enrollment_date, + major, + grade, + status, + facultyNo, + classNo + FROM student + WHERE facultyNo = 1 AND classNo = 7 -- Partition conditions + DISTRIBUTE BY facultyNo, classNo, floor(rand() * 2); -- Controlling the number of files + + After execution, the data under **facultyNo = 1, classNo = 7** will be consolidated into up to 2 files. + + .. code-block:: + + -- 3. Verify the partition (e.g., facultyNo = 1, classNo = 7): + SELECT count(*) FROM student + WHERE facultyNo=1 AND classNo=7; --> Original data volume + + SELECT count(*) FROM student_tmp + WHERE facultyNo=1 AND classNo=7; --> Merged data volume + + -- 4. Run INSERT OVERWRITE to write the merged data back to the source table. + INSERT OVERWRITE TABLE student PARTITION(facultyNo, classNo) + SELECT + * + FROM student_temp + WHERE facultyNo > 0; + +More Operation Recommendations +------------------------------ + +When directly overwriting the source table, data loss may occur if the task fails during execution. + +**To mitigate this risk, you are advised to use a temporary table as an intermediary:** + +#. **Create a temporary table** with the same structure as the source table. +#. **Insert merged data into the temporary table.** +#. **After verifying the data in the temporary table**, switch it to the production table using **ALTER TABLE ... RENAME TO** or **INSERT OVERWRITE**. + +For the **DISTRIBUTE BY** clause: Instead of relying solely on random numbers, you can use a naturally high-cardinality field like **user_id**. This approach not only merges files but also sorts the data based on that field, potentially improving query performance. .. code-block:: - INSERT OVERWRITE TABLE tablename - select * FROM tablename - DISTRIBUTE BY floor(rand()*20) + -- 1. Create a temporary table with the same structure as the source table table_name. + CREATE TABLE table_name_tmp LIKE table_name; + + -- 2. Merge the data from the source table and insert it into the temporary table. + INSERT OVERWRITE TABLE table_name_tmp PARTITION (pt) -- For a partitioned table + SELECT ... -- Selected fields + FROM table_name + WHERE pt = * + DISTRIBUTE BY pt, floor(rand() * N); -- Partitioned table syntax + -- For a non-partitioned table: DISTRIBUTE BY floor(rand() * N) + + -- 3. Verify data integrity (e.g., row counts, partitions). + SELECT count(*) FROM table_name; + SELECT count(*) FROM table_name_tmp; + + -- 4. Swap tables via renaming (a metadata operation completed instantly). + ALTER TABLE table_name RENAME TO table_name_old; + ALTER TABLE table_name_tmp RENAME TO table_name; + + -- 5. (Optional) Delete the old table if no longer needed. + DROP TABLE table_name_old; + +Summary of Small File Merging Methods +------------------------------------- + ++---------------------------+-------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------+ +| Scenario | Recommended SQL Syntax | Key Point | ++===========================+===============================================================================================================================+==============================================================================================+ +| **Non-partitioned table** | .. code-block:: | Use random numbers to control the number of files (N). | +| | | | +| | INSERT OVERWRITE TABLE table_tmp SELECT * FROM table DISTRIBUTE BY floor(rand()*N); | | ++---------------------------+-------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------+ +| **Partitioned table** | .. code-block:: | **Include the partition field in DISTRIBUTE BY** to ensure merging occurs within partitions. | +| | | | +| | INSERT OVERWRITE TABLE table_tmp PARTITION(pt) SELECT ..., pt FROM table WHERE pt = * DISTRIBUTE BY pt, floor(rand()*N); | | ++---------------------------+-------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------+ +| **General strategy** | **Use a temporary table for intermediate storage**. Verify data before swapping tables via renaming to avoid data loss risks. | Minimize the risk of data loss during the process. | ++---------------------------+-------------------------------------------------------------------------------------------------------------------------------+----------------------------------------------------------------------------------------------+ diff --git a/umn/source/faq/sql_jobs/sql_job_o_and_m/why_is_error_org.apache.hadoop.fs.obs.obsioexception_reported_when_i_run_dli_sql_scripts_on_dataarts_studio.rst b/umn/source/faq/sql_jobs/sql_job_o_and_m/why_is_error_org.apache.hadoop.fs.obs.obsioexception_reported_when_i_run_dli_sql_scripts_on_dataarts_studio.rst index 18d49f9..5537723 100644 --- a/umn/source/faq/sql_jobs/sql_job_o_and_m/why_is_error_org.apache.hadoop.fs.obs.obsioexception_reported_when_i_run_dli_sql_scripts_on_dataarts_studio.rst +++ b/umn/source/faq/sql_jobs/sql_job_o_and_m/why_is_error_org.apache.hadoop.fs.obs.obsioexception_reported_when_i_run_dli_sql_scripts_on_dataarts_studio.rst @@ -13,9 +13,9 @@ When you run a DLI SQL script on DataArts Studio, the log shows that the stateme .. code-block:: DLI.0999: RuntimeException: org.apache.hadoop.fs.obs.OBSIOException: initializing on obs://xxx.csv: status [-1] - request id - [null] - error code [null] - error message [null] - trace :com.obs.services.exception.ObsException: OBS servcie Error Message. Request Error: + [null] - error code [null] - error message [null] - trace :com.obs.services.exception.ObsException: OBS service Error Message. Request Error: ... - Cause by: ObsException: com.obs.services.exception.ObsException: OBSs servcie Error Message. Request Error: java.net.UnknownHostException: xxx: Name or service not known + Cause by: ObsException: com.obs.services.exception.ObsException: OBSs service Error Message. Request Error: java.net.UnknownHostException: xxx: Name or service not known Possible Causes --------------- diff --git a/umn/source/getting_started/submitting_a_flink_jar_job_using_dli.rst b/umn/source/getting_started/submitting_a_flink_jar_job_using_dli.rst index 46eaf76..8906a35 100644 --- a/umn/source/getting_started/submitting_a_flink_jar_job_using_dli.rst +++ b/umn/source/getting_started/submitting_a_flink_jar_job_using_dli.rst @@ -46,7 +46,7 @@ Develop a Flink Jar job program, compile it, and pack it into **flink-examples.j Before submitting a Flink job, upload data files to OBS. -#. Log in to the DLI console. +#. Log in to the management console. #. In the service list, click **Object Storage Service** under **Storage**. @@ -74,7 +74,7 @@ Before submitting a Flink job, upload data files to OBS. Step 2: Buy an Elastic Resource Pool and Create Queues Within It ---------------------------------------------------------------- -To execute SQL jobs in datasource scenarios, you must use your own SQL queue as the existing **default** queue cannot be used. In this example, create an elastic resource pool named **dli_resource_pool** and a queue named **dli_queue_01**. +When running a Flink job in datasource scenarios, you cannot use the existing **default** queue. You need to create a general-purpose queue. In this example, the elastic resource pool **dli_resource_pool** and queue **dli_queue_01** are created. #. Log in to the DLI management console. @@ -160,7 +160,7 @@ Step 3: Use DEW to Manage Access Credentials In cross-source analysis scenarios, you need to set attributes such as the username and password in the connector. However, information such as usernames and passwords is highly sensitive and needs to be encrypted to ensure user data privacy. -DEW is a secure, reliable, and easy-to-use solution for encrypting and decrypting private data while ensuring its security. +DEW offers a secure, reliable, and easy-to-use solution for encrypting and decrypting sensitive data. This example introduces how to create a shared secret in DEW. @@ -321,7 +321,7 @@ Step 5: Create a Flink Jar Job and Configure Job Information | | | .. note:: | | | | | | | | - The value cannot exceed four times the number of compute units (**CUs** - **Job Manager CUs**). | - | | | - You are advised to set this parameter to a value greater than that configured in the code. Otherwise, job submission may fail. | + | | | - Set this parameter to a value greater than that configured in the code. Otherwise, job submission may fail. | +---------------------------+-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Task Manager Config | No | Whether TaskManager resource parameters are set | | | | | @@ -342,7 +342,7 @@ Step 5: Create a Flink Jar Job and Configure Job Information | | | | | | | **SMN Topic** | | | | | - | | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | + | | | Select a custom SMN topic. For details about how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | +---------------------------+-----------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Auto Restart on Exception | No | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | | | | | diff --git a/umn/source/getting_started/submitting_a_flink_opensource_sql_job_to_query_rds_for_mysql_data_using_dli.rst b/umn/source/getting_started/submitting_a_flink_opensource_sql_job_to_query_rds_for_mysql_data_using_dli.rst index caaaef1..f474c31 100644 --- a/umn/source/getting_started/submitting_a_flink_opensource_sql_job_to_query_rds_for_mysql_data_using_dli.rst +++ b/umn/source/getting_started/submitting_a_flink_opensource_sql_job_to_query_rds_for_mysql_data_using_dli.rst @@ -56,14 +56,14 @@ Enable DIS to import Kafka data to DLI. For details, see "Buying a Kafka Instanc Before creating a Kafka instance, ensure the availability of resources, including a virtual private cloud (VPC), subnet, security group, and security group rules. - - For how to create a VPC and subnet, see "Creating a VPC and Subnet" in the *Virtual Private Cloud User Guide*. For how to create and use a subnet in an existing VPC, see "Create a Subnet for the VPC" in the *Virtual Private Cloud User Guide*. + - For details about how to create a VPC and subnet, see "Creating a VPC with a Subnet" in the *Virtual Private Cloud User Guide*. For details about how to create and use a subnet in an existing VPC, see "Creating a Subnet for an Existing VPC" in the *Virtual Private Cloud User Guide*. .. note:: - The created VPC and the Kafka instance you will create must be in the same region. - Retain the default settings unless otherwise specified. - - For how to create a security group, see "Creating a Security Group" in the *Virtual Private Cloud User Guide*. For how to add rules to a security group, see "Creating a Subnet for the VPC" in the *Virtual Private Cloud User Guide*. + - For details about how to create a security group, see "Creating a Security Group" in the *Virtual Private Cloud User Guide*. For details about how to add rules to a security group, see "Adding a Security Group Rule" in the *Virtual Private Cloud User Guide*. #. Create a Kafka premium instance as the job source stream. @@ -148,7 +148,7 @@ Use RDS for MySQL as the data sink stream and create an RDS for MySQL DB instanc #. Click **Buy DB Instance** in the upper right corner of the page and set related parameters. Retain the default values for other parameters. - For the parameters, see "RDS for MySQL Getting Started" in the *Relational Database Service Getting Started*. + For details about the parameters, see "RDS for MySQL Getting Started" in the *Relational Database Service Getting Started*. .. table:: **Table 3** RDS for MySQL instance parameters @@ -191,7 +191,7 @@ Use RDS for MySQL as the data sink stream and create an RDS for MySQL DB instanc +------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-----------------------------------------------------------------------------------------------------------+ | VPC and Subnet | Select an existing VPC and subnet. | Select the VPC and subnet created in :ref:`1 `. | | | | | - | | For how to recreate a VPC and subnet, refer to "Creating a VPC and Subnet" in the *Virtual Private Cloud User Guide*. | | + | | For details about how to recreate a VPC and subnet, see "Creating a VPC with a Subnet" in the *Virtual Private Cloud User Guide*. | | | | | | | | .. note:: | | | | | | @@ -256,10 +256,10 @@ Step 3: Create an OBS Bucket to Store Output Data In this example, you need to enable OBS for job **JobSample** to provide DLI Flink jobs with the functions of checkpointing, saving job logs, and commissioning test data. -For how to create a bucket, see "Creating a Bucket" in the *Object Storage Service Console Operation Guide*. +For details about how to create a bucket, see "Creating a Bucket" in the *Object Storage Service User Guide*. #. In the navigation pane on the OBS management console, choose **Object Storage**. -#. In the upper right corner of the page, click **Create Bucket** and set bucket parameters. +#. In the upper right corner of the page, click **Create Bucket** and configure bucket parameters. .. table:: **Table 4** OBS bucket parameters @@ -286,7 +286,7 @@ For how to create a bucket, see "Creating a Bucket" in the *Object Storage Servi Step 4: Create an Elastic Resource Pool and Create Queues Within It ------------------------------------------------------------------- -To create a Flink OpenSource SQL job, you must use your own queue as the existing **default** queue cannot be used. In this example, create an elastic resource pool named **dli_resource_pool** and a queue named **dli_queue_01**. +When creating a Flink OpenSource SQL job, you cannot use the existing **default** queue. You need to create a general-purpose queue. In this example, the elastic resource pool **dli_resource_pool** and queue **dli_queue_01** are created. #. Log in to the DLI management console. @@ -396,7 +396,7 @@ You need to create an enhanced datasource connection for the Flink OpenSource SQ a. Log in to the DLI management console. In the navigation pane on the left, choose **Datasource Connections**. On the displayed page, click **Create** in the **Enhanced** tab. - b. In the displayed dialog box, set the following parameters: For details, see the following section: + b. In the dialog box that appears, configure the parameters as follows: - **Connection Name**: Name of the enhanced datasource connection For this example, enter **dli_kafka**. - **Resource Pool**: Select the elastic resource pool created in :ref:`Step 4: Create an Elastic Resource Pool and Create Queues Within It `. @@ -431,7 +431,7 @@ Step 6: Create an Enhanced Datasource Connection Between DLI and RDS a. Log in to the DLI management console. In the navigation pane on the left, choose **Datasource Connections**. On the displayed page, click **Create** in the **Enhanced** tab. - b. In the displayed dialog box, set the following parameters: For details, see the following section: + b. In the dialog box that appears, configure the parameters as follows: - **Connection Name**: Name of the enhanced datasource connection For this example, enter **dli_rds**. - **Resource Pool**: Select the name of the queue created in :ref:`Step 4: Create an Elastic Resource Pool and Create Queues Within It `. @@ -454,7 +454,9 @@ In cross-source analysis scenarios, you need to set attributes such as the usern Flink 1.15 allows for the use of DEW to manage credentials. Before running a job, create a custom agency and configure agency information within the job. -Data Encryption Workshop (DEW) and Cloud Secret Management Service (CSMS) joint form a secure, reliable, and easy-to-use privacy data encryption and decryption solution. This example describes how a Flink OpenSource SQL job uses DEW to manage RDS access credentials. +Data Encryption Workshop (DEW) and Cloud Secret Management Service (CSMS) offer a secure, reliable, and easy-to-use solution for encrypting and decrypting sensitive data. + +This example describes how a Flink OpenSource SQL job uses DEW to manage RDS access credentials. #. Create an agency for DLI to access DEW and complete authorization. #. Create a shared secret in DEW. @@ -477,7 +479,7 @@ Step 8: Create a Flink OpenSource SQL Job After the source and sink streams are prepared, you can create a Flink OpenSource SQL job. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. In the upper right corner of the **Flink Jobs** page, click **Create Job**. diff --git a/umn/source/getting_started/submitting_a_spark_jar_job_using_dli.rst b/umn/source/getting_started/submitting_a_spark_jar_job_using_dli.rst index f75ad98..138c19e 100644 --- a/umn/source/getting_started/submitting_a_spark_jar_job_using_dli.rst +++ b/umn/source/getting_started/submitting_a_spark_jar_job_using_dli.rst @@ -156,7 +156,7 @@ In this example, the elastic resource pool **dli_resource_pool** and queue **dli Step 3: Use DEW to Manage Access Credentials -------------------------------------------- -To write the output data of a Spark Jar job to OBS, AK/SK is required for accessing OBS. To ensure the security of AK/SK data, you can use DEW and CSMS for unified management of AK/SK, effectively avoiding sensitive information leakage and business risks caused by hard-coded or plaintext configuration of programs. +When writing output data from Spark Jar jobs to OBS, you need to configure an AK/SK for accessing OBS. To ensure the security of AK/SK data, you can use DEW and CSMS for centralized management of AK/SK. This approach effectively mitigates risks such as sensitive information leakage caused by hardcoding in programs or plaintext configurations, as well as potential business disruptions due to unauthorized access. This part introduces how to create a shared secret in DEW. diff --git a/umn/source/getting_started/submitting_a_sql_job_to_query_rds_for_mysql_data_using_dli.rst b/umn/source/getting_started/submitting_a_sql_job_to_query_rds_for_mysql_data_using_dli.rst index 6bf8e53..da7d524 100644 --- a/umn/source/getting_started/submitting_a_sql_job_to_query_rds_for_mysql_data_using_dli.rst +++ b/umn/source/getting_started/submitting_a_sql_job_to_query_rds_for_mysql_data_using_dli.rst @@ -97,7 +97,7 @@ For details, see "RDS for MySQL Getting Started" in the *Relational Database Ser +------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+---------------------------+ | VPC | Select an existing VPC. | ``-`` | | | | | - | | For how to recreate a VPC and subnet, refer to "Creating a VPC and Subnet" in the *Virtual Private Cloud User Guide*. | | + | | For details about how to recreate a VPC and subnet, see "Creating a VPC with a Subnet" in the *Virtual Private Cloud User Guide*. | | | | | | | | .. note:: | | | | | | diff --git a/umn/source/monitoring_dli_using_cloud_eye.rst b/umn/source/monitoring_dli_using_cloud_eye.rst index a6f61de..0d79a7b 100644 --- a/umn/source/monitoring_dli_using_cloud_eye.rst +++ b/umn/source/monitoring_dli_using_cloud_eye.rst @@ -20,91 +20,92 @@ Metric .. table:: **Table 1** DLI metrics - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | Metric ID | Name | Description | Value Range | Unit | Conversion Rule | Monitored Object | Monitoring Period (Raw Data) | - +=================================+=========================================+=================================================================================================================+=============+==========+=================+===================================================================+==============================+ - | queue_cu_num | Queue CU Usage | Displays the number of CUs applied by the user queue | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_job_launching_num | Number of Jobs Being Submitted | Displays the number of jobs in the Submitting state in the user queue. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_job_running_num | Number of Running Jobs | Displays the number of running jobs in the user queue. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_job_succeed_num | Number of Finished Jobs | Displays the number of completed jobs in the user queue. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_job_failed_num | Failed Jobs | Displays the number of failed jobs in the user queue. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_job_cancelled_num | Number of Canceled Jobs | Displays the number of canceled jobs in the user queue. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_alloc_cu_num | Allocated CUs (queue) | Displays the CU allocation for user queues. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_min_cu_num | Minimum CUs for Queue | Displays the minimum number of CUs for a user queue. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_max_cu_num | Maximum CUs for Queue | Displays the maximum number of CUs for a user queue. | >= 0 | Count | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_priority | Queue Priority | Displays the priority of a user queue. | 1-100 | N/A | N/A | Queues | 5 minutes | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_cpu_usage | Queue CPU Usage | Displays the CPU usage of user queues. | 0-100 | % | N/A | Queues | 5 minutes | - | | | | | | | | | - | | | | | | | This metric applies only to queues in non-elastic resource pools. | | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_disk_usage | Queue Disk Usage | Displays the disk usage of user queues. | 0-100 | % | N/A | Queues | 5 minutes | - | | | | | | | | | - | | | | | | | This metric applies only to queues in non-elastic resource pools. | | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_disk_used | Max Disk Usage | Displays the maximum disk usage of user queues. | 0-100 | % | N/A | Queues | 5 minutes | - | | | | | | | | | - | | | | | | | This metric applies only to queues in non-elastic resource pools. | | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_mem_usage | Queue Memory Usage | Displays the memory usage of user queues. | 0-100 | % | N/A | Queues | 5 minutes | - | | | | | | | | | - | | | | | | | This metric applies only to queues in non-elastic resource pools. | | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | queue_mem_used | Used Memory | Displays the memory usage rate of the user queues. | >= 0 | MB | N/A | Queues | 5 minutes | - | | | | | | | | | - | | | | | | | This metric applies only to queues in non-elastic resource pools. | | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_read_records_per_second | Flink Job Data Read Rate | Displays the data input rate of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_write_records_per_second | Flink Job Data Write Rate | Displays the data output rate of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_read_records_total | Flink Job Total Data Read | Displays the total number of data inputs of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_write_records_total | Flink Job Total Data Write | Displays the total number of output data records of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_read_bytes_per_second | Flink Job Byte Read Rate | Displays the number of input bytes per second of a Flink job. | >= 0 | byte/s | 1024(IEC) | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_write_bytes_per_second | Flink Job Byte Write Rate | Displays the number of output bytes per second of a Flink job. | >= 0 | byte/s | 1024(IEC) | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_read_bytes_total | Flink Job Total Read Byte | Displays the total number of input bytes of a Flink job. | >= 0 | byte/s | 1024(IEC) | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_write_bytes_total | Flink Job Total Write Byte | Displays the total number of output bytes of a Flink job. | >= 0 | byte/s | 1024(IEC) | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_cpu_usage | Flink Job CPU Usage | Displays the CPU usage of Flink jobs. | 0-100 | % | N/A | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_mem_usage | Flink Job Memory Usage | Displays the memory usage of Flink jobs. | 0-100 | % | N/A | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_max_op_latency | Flink Job Max Operator Latency | Displays the maximum operator delay of a Flink job. The unit is **ms**. | >= 0 | ms | N/A | Flink jobs | 10 seconds | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ - | flink_max_op_backpressure_level | Flink Job Maximum Operator Backpressure | Displays the maximum operator backpressure value of a Flink job. A larger value indicates severer backpressure. | 0-100 | N/A | N/A | Flink jobs | 10 seconds | - | | | | | | | | | - | | | **0**: OK | | | | | | - | | | | | | | | | - | | | **50**: low | | | | | | - | | | | | | | | | - | | | **100**: high | | | | | | - +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+-------------------------------------------------------------------+------------------------------+ + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | Metric ID | Name | Description | Value Range | Unit | Conversion Rule | Dimension | Monitored Object | Monitoring Period (Raw Data) | + +=================================+=========================================+=================================================================================================================+=============+==========+=================+==============+===================================================================+==============================+ + | queue_cu_num | Queue CU Usage | Displays the number of CUs applied by the user queue | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_job_launching_num | Number of Jobs Being Submitted | Displays the number of jobs in the Submitting state in the user queue. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_job_running_num | Number of Running Jobs | Displays the number of running jobs in the user queue. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_job_succeed_num | Number of Finished Jobs | Displays the number of completed jobs in the user queue. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_job_failed_num | Failed Jobs | Displays the number of failed jobs in the user queue. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_job_cancelled_num | Number of Canceled Jobs | Displays the number of canceled jobs in the user queue. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_alloc_cu_num | Allocated CUs (queue) | Displays the CU allocation for user queues. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_min_cu_num | Minimum CUs for Queue | Displays the minimum number of CUs for a user queue. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_max_cu_num | Maximum CUs for Queue | Displays the maximum number of CUs for a user queue. | >= 0 | Count | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_priority | Queue Priority | Displays the priority of a user queue. | 1-100 | N/A | N/A | queue_id | Queues | 5 minutes | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_cpu_usage | Queue CPU Usage | Displays the CPU usage of user queues. | 0-100 | % | N/A | queue_id | Queues | 5 minutes | + | | | | | | | | | | + | | | | | | | | This metric applies only to queues in non-elastic resource pools. | | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_disk_usage | Queue Disk Usage | Displays the disk usage of user queues. | 0-100 | % | N/A | queue_id | Queues | 5 minutes | + | | | | | | | | | | + | | | | | | | | This metric applies only to queues in non-elastic resource pools. | | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_disk_used | Max Disk Usage | Displays the maximum disk usage of user queues. | 0-100 | % | N/A | queue_id | Queues | 5 minutes | + | | | | | | | | | | + | | | | | | | | This metric applies only to queues in non-elastic resource pools. | | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_mem_usage | Queue Memory Usage | Displays the memory usage of user queues. | 0-100 | % | N/A | queue_id | Queues | 5 minutes | + | | | | | | | | | | + | | | | | | | | This metric applies only to queues in non-elastic resource pools. | | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | queue_mem_used | Used Memory | Displays the memory usage rate of the user queues. | >= 0 | MB | N/A | queue_id | Queues | 5 minutes | + | | | | | | | | | | + | | | | | | | | This metric applies only to queues in non-elastic resource pools. | | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_read_records_per_second | Flink Job Data Read Rate | Displays the data input rate of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_write_records_per_second | Flink Job Data Write Rate | Displays the data output rate of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_read_records_total | Flink Job Total Data Read | Displays the total number of data inputs of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_write_records_total | Flink Job Total Data Write | Displays the total number of output data records of a Flink job for monitoring and debugging. | >= 0 | record/s | N/A | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_read_bytes_per_second | Flink Job Byte Read Rate | Displays the number of input bytes per second of a Flink job. | >= 0 | byte/s | 1024(IEC) | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_write_bytes_per_second | Flink Job Byte Write Rate | Displays the number of output bytes per second of a Flink job. | >= 0 | byte/s | 1024(IEC) | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_read_bytes_total | Flink Job Total Read Byte | Displays the total number of input bytes of a Flink job. | >= 0 | byte/s | 1024(IEC) | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_write_bytes_total | Flink Job Total Write Byte | Displays the total number of output bytes of a Flink job. | >= 0 | byte/s | 1024(IEC) | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_cpu_usage | Flink Job CPU Usage | Displays the CPU usage of Flink jobs. | 0-100 | % | N/A | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_mem_usage | Flink Job Memory Usage | Displays the memory usage of Flink jobs. | 0-100 | % | N/A | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_max_op_latency | Flink Job Max Operator Latency | Displays the maximum operator delay of a Flink job. The unit is **ms**. | >= 0 | ms | N/A | flink_job_id | Flink jobs | 10 seconds | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ + | flink_max_op_backpressure_level | Flink Job Maximum Operator Backpressure | Displays the maximum operator backpressure value of a Flink job. A larger value indicates severer backpressure. | 0-100 | N/A | N/A | flink_job_id | Flink jobs | 10 seconds | + | | | | | | | | | | + | | | **0**: OK | | | | | | | + | | | | | | | | | | + | | | **50**: low | | | | | | | + | | | | | | | | | | + | | | **100**: high | | | | | | | + +---------------------------------+-----------------------------------------+-----------------------------------------------------------------------------------------------------------------+-------------+----------+-----------------+--------------+-------------------------------------------------------------------+------------------------------+ Dimension --------- .. table:: **Table 2** Dimension - ============ ========= - Key Value - ============ ========= - queue_id Queue - flink_job_id Flink job - ============ ========= + ======================== ===================== + Key Value + ======================== ===================== + queue_id Queue + elastic_resource_pool_id Elastic resource pool + flink_job_id Flink job + ======================== ===================== Viewing DLI Monitoring Metrics on Cloud Eye ------------------------------------------- diff --git a/umn/source/preparations/configuring_dli_agency_permissions.rst b/umn/source/preparations/configuring_dli_agency_permissions.rst index 6cc24b6..8e38b25 100644 --- a/umn/source/preparations/configuring_dli_agency_permissions.rst +++ b/umn/source/preparations/configuring_dli_agency_permissions.rst @@ -89,9 +89,9 @@ In addition to the permissions provided by **dli_management_agency**, you need t - Data cleanup agency required for clearing data according to the lifecycle of a table and clearing lakehouse table data. You need to create a DLI agency named **dli_data_clean_agency** on IAM and grant permissions to it. You need to create an agency and customize permissions for it. However, the agency name is fixed to **dli_data_clean_agency**. - **Tenant Administrator** permissions are required to access data from OBS to execute Flink jobs on DLI, for example, obtaining OBS data sources, log dump (including bucket authorization), checkpointing enabling, and job import and export. -- The AK/SK required by DLI Flink jobs is stored in DEW. To allow DLI to access DEW data during job execution, you need to create an agency to delegate the permissions to operate on DEW data to DLI. +- The AK and SK required for DLI Flink jobs are securely stored in DEW. To enable DLI to access DEW data during job execution, you need to create an agency that grants DLI the necessary permissions to operate on DEW data. This agency allows DLI to access DEW on your behalf. - To allow DLI to access DLI catalogs to retrieve metadata when executing jobs, you need to create an agency that grants DLI catalog data operation permissions to DLI. This will enable DLI to access DLI catalogs on your behalf. -- Cloud data required by DLI Flink jobs is stored in LakeFormation. To allow DLI to access catalogs to retrieve metadata during job execution, you need to create an agency to delegate the permissions to operate on catalog data to DLI. +- Cloud data required for DLI Flink jobs is securely stored in LakeFormation. To enable DLI to access catalogs to retrieve metadata during job execution, you need to create an agency that grants DLI the necessary permissions to operate on catalog data. This agency allows DLI to access catalog metadata on your behalf. When creating an agency, you cannot use the default agency names **dli_admin_agency**, **dli_management_agency**, or **dli_data_clean_agency**. It must be unique. diff --git a/umn/source/service_overview/advantages.rst b/umn/source/service_overview/advantages.rst index 18a1f0d..2458bde 100644 --- a/umn/source/service_overview/advantages.rst +++ b/umn/source/service_overview/advantages.rst @@ -5,24 +5,29 @@ Advantages ========== -Full SQL Compatibility ----------------------- +Pure SQL Operations: Zero Learning Curve +---------------------------------------- -You do not need a background in big data to use DLI for data analysis. You only need to know SQL, and you are good to go. The SQL syntax is fully compatible with the standard ANSI SQL 2003. +- DLI offers standard SQL APIs, enabling you to perform massive data query and analysis using only SQL. Its syntax is fully compatible with ANSI SQL 2003 standards. +- This significantly lowers the barrier for data analysts and business professionals, enhancing overall efficiency in data analysis. -Decoupled Storage and Compute ------------------------------ +Decoupled Storage and Compute: Efficient Resource Utilization +------------------------------------------------------------- -DLI compute and storage loads are decoupled. This architecture allows you to flexibly configure storage and compute resources on demand, improving resource utilization and reducing costs. +- DLI decouples storage and compute workloads through its decoupling architecture, allowing flexible configuration of resources based on demand. This improves resource utilization and reduces costs. +- The elastic resource pool supports multiple engines like Flink and Spark, further optimizing resource allocation efficiency. -DLI's elastic resource pool feature effectively boosts compute resource utilization. The same compute resources can support multiple compute engines like Flink and Spark simultaneously. +Serverless Architecture: Full-Scenario Adaptability +--------------------------------------------------- -Serverless DLI --------------- +DLI is fully compatible with `Apache Spark `__ and `Apache Flink `__ ecosystems and APIs, providing a unified serverless big data computing service for real-time, offline, and interactive analytics. -DLI is fully compatible with `Apache Spark `__ and `Apache Flink `__ ecosystems and APIs. It is a serverless big data computing and analysis service that integrates real-time, offline, and interactive analysis. Offline applications can be seamlessly migrated to the cloud, reducing the migration workload. DLI provides a highly-scalable framework integrating batch and stream processing, allowing you to handle data analysis requests with ease. With a deeply optimized kernel and architecture, DLI delivers 100-fold performance improvement compared with the MapReduce model. Your analysis is backed by an industry-vetted 99.95% SLA. +- On-premises Spark/Flink applications can migrate to the cloud effortlessly, minimizing migration efforts and ensuring smooth transitions. +- A batch-stream fusion framework delivers scalable, high-performance processing for TB to EB-level data, meeting diverse big data needs. +- Deep optimizations in product core and architecture result in performance over 100x faster than traditional MapReduce models, with 99.95% SLA. -Cross-Source Analysis ---------------------- +Cross-Source Analysis: No Data Migration Required +------------------------------------------------- -Analyze your data across databases. No migration required. A unified view of your data gives you a comprehensive understanding of your data and helps you innovate faster. There are no restrictions on data formats, cloud data sources, or whether the database is created online or off. +- Supports multiple data formats and sources, including cloud-based (e.g., OBS, RDS, DWS, CSS, MongoDB, Redis), ECS-hosted databases, and on-premises databases. +- Enables unified cross-source analysis without data relocation, accelerating enterprise-wide data insights and innovation. diff --git a/umn/source/service_overview/basic_concepts.rst b/umn/source/service_overview/basic_concepts.rst index acb5018..88ad0fa 100644 --- a/umn/source/service_overview/basic_concepts.rst +++ b/umn/source/service_overview/basic_concepts.rst @@ -5,33 +5,23 @@ Basic Concepts ============== -Elastic Resource Pool ---------------------- +Actual CUs, Used CUs, CU Range, and Yearly/Monthly CUs (Specifications) of an Elastic Resource Pool +--------------------------------------------------------------------------------------------------- -An elastic resource pool consists of dedicated compute resources where the resources in different pools are completely isolated from each other. Within a single elastic resource pool, multiple queues can share the resources available in that pool. Additionally, you can configure policies based on the resource load of these queues to enable time-based elastic scaling to meet diverse business needs. +- **Actual CUs**: The current allocated resource size of the elastic resource pool, measured in CUs. -DLI Storage Resource --------------------- + - When no queues exist in the resource pool: The actual CUs equal the minimum CUs set during its creation. -DLI storage resources are the internal storage capacities of the DLI service. They are utilized for storing databases and DLI tables, playing a crucial role in data import into DLI. These resources also indicate the volume of data that users have stored within DLI. - -Actual CUs, Used CUs, CU Range, and Specifications of an Elastic Resource Pool ------------------------------------------------------------------------------- - -- **Actual CUs**: actual size of resources currently allocated to the elastic resource pool (in CUs). - - - When there is no queue in the resource pool, the actual CUs are equal to the minimum CUs when the elastic resource pool is created. - - - When there are queues in the resource pool, the formula for calculating actual CUs is: + - When there are queues in the resource pool, the formula to calculate actual CUs is: - Actual CUs = max{(min[sum(maximum CUs of queues), maximum CUs of the elastic resource pool]), minimum CUs of the elastic resource pool}. - - The calculation result must be a multiple of 16 CUs. If it cannot be exactly divided by 16 CUs, round up to the nearest multiple. + - The result must be a multiple of 16 CUs. If not divisible by 16, round up to the nearest multiple. - - Scaling out or in an elastic resource pool means adjusting the actual CUs of the resource pool. Refer to :ref:`Scaling Out or In an Elastic Resource Pool `. + - Scaling out or in an elastic resource pool means adjusting its actual CUs. See :ref:`Scaling Out or In an Elastic Resource Pool `. - Example of actual CU allocation: - In :ref:`Table 1 `, the calculation process for the actual allocation of CUs in an elastic resource pool is as follows: + Consider :ref:`Table 1 ` below, which illustrates the process of calculating actual CUs for an elastic resource pool: #. Calculate the sum of maximum CUs of the queues: sum(maximum CUs) = 32 + 56 = 88 CUs. @@ -60,73 +50,77 @@ Actual CUs, Used CUs, CU Range, and Specifications of an Elastic Resource Pool | | Queue B | 16-56CUS | +---------------------------------------------------------------------------------------------------+-----------------------+-----------------------+ -- **Used CUs**: CUs that have been used by jobs or tasks. These resources may be executing computing tasks. +- **Used CUs**: The portion of CUs currently occupied by jobs or tasks, which may be actively performing computations. - **CU range**: CU settings are used to control the maximum and minimum CU ranges for elastic resource pools to avoid unlimited resource scaling. - - The total minimum CUs of all queues in an elastic resource pool must be no more than the minimum CUs of the pool. - - The maximum CUs of any queue in an elastic resource pool must be no more than the maximum CUs of the pool. - - An elastic resource pool should at least ensure that all queues in it can run with the minimum CUs and should try to ensure that all queues in it can run with the maximum CUs. - - When expanding the specifications of an elastic resource pool, the minimum value of the CU range is linked to the specifications of the elastic resource pool. After changing the specifications of the elastic resource pool, the minimum value of the CU range is modified to match the specifications. + - The sum of all queues' minimum CUs in an elastic resource pool must not exceed the pool's minCU. + - Any single queue's maxCU cannot exceed the pool's maxCU. + - The resource pool ensures it meets the minCU requirements across all queues while striving to accommodate their maxCU demands. + - When expanding the specifications of an elastic resource pool, the minimum value of the CU range is linked to the yearly/monthly CUs (specifications) of the elastic resource pool. After changing the specifications of the elastic resource pool, the minimum value of the CU range is modified to match the yearly/monthly CUs (specifications). -- **Specifications**: The minimum CUs selected during elastic resource pool purchase are elastic resource pool specifications. +- **Yearly/monthly CUs (specifications)**: The minimum value of the CU range selected when purchasing an elastic resource pool is the elastic resource pool specifications. Database -------- -A database is a warehouse where data is organized, stored, and managed based on the data structure. DLI management permissions are granted on a per database basis. +A database is a structured repository designed to organize, store, and manage data efficiently. In DLI, databases serve as the fundamental unit for managing permissions, with access rights assigned at the database level. -In DLI, tables and databases are metadata containers that define underlying data. The metadata in the table shows the location of the data and specifies the data structure, such as the column name, data type, and table name. A database is a collection of tables. +Within DLI, both tables and databases act as metadata containers that define underlying data structures. Table metadata informs DLI about the location of the data and specifies its structure, such as column names, data types, and table names. Databases provide logical groupings for these tables. -OBS Table, DLI Table, and CloudTable Table ------------------------------------------- +OBS Tables, DLI Tables, CloudTable Tables +----------------------------------------- -The table type indicates the storage location of data. +Different table types indicate distinct storage locations: -- OBS table indicates that data is stored in the OBS bucket. -- DLI table indicates that data is stored in the internal table of DLI. -- CloudTable table indicates that data is stored in CloudTable. +- OBS table: Data is stored in buckets within OBS. -You can create a table on DLI and associate the table with other services to achieve querying data from multiple data sources. +- DLI table: Data is stored in tables internal to DLI. + + DLI storage resources are internal resources used to house databases and DLI tables, essential for importing data into DLI and reflecting the volume of user data stored within the service. + +- CloudTable table: Data is stored in tables managed by CloudTable. + +Tables can be created through DLI to establish connections with other services, enabling federated query and analysis across diverse data sources. Metadata -------- -Metadata is used to define data types. It describes information about the data, including the source, size, format, and other data features. In database fields, metadata interprets data content in the data warehouse. +Metadata refers to data that defines other data types. It primarily describes information about the data itself, including its source, size, format, or other characteristics. In database fields, metadata is used to interpret the contents of a data warehouse. -SQL Job -------- +SQL Jobs +-------- -SQL job refers to the SQL statement executed in the SQL job editor. It serves as the execution entity used for performing operations, such as importing and exporting data, in the SQL job editor. +A SQL job refers to the execution entity within the system that handles operations such as running SQL statements, importing data, and exporting data through the SQL job editor. -This type is suitable for scenarios where standard SQL statements are used for querying. It is typically used for querying and analyzing structured data. +It is ideal for scenarios involving structured data queries and analysis using standard SQL. -Flink Job ---------- +Flink Jobs +---------- -This type is specifically designed for real-time data stream processing, making it ideal for scenarios that require low latency and quick response. It is well-suited for real-time monitoring and online analysis. +Designed for real-time stream processing, Flink jobs are suited for low-latency applications requiring rapid responses, such as real-time monitoring and online analytics. -- Flink OpenSource job: When submitting jobs, you can quickly integrate with other data systems using DLI's standard connectors and various APIs. -- Flink Jar job: allows you to submit Flink jobs compiled into JAR files, providing greater flexibility and customization capabilities. It is suitable for complex data processing scenarios that require user-defined functions (UDFs) or specific library integration. The Flink ecosystem can be utilized to implement advanced stream processing logic and status management. +- Flink OpenSource jobs: These allow you to use DLI-provided connectors and APIs for seamless integration with other data systems during job submission. +- Flink Jar jobs: You can submit pre-compiled JAR files containing Flink jobs, offering greater flexibility and customization. This type is ideal for complex data processing tasks involving custom functions, UDFs, or specific library integrations, enabling advanced stream processing logic and state management using Flink's ecosystem. -Spark Job ---------- +Spark Jobs +---------- -Spark jobs are those submitted by users through visualized interfaces and RESTful APIs. Full-stack Spark jobs are allowed, such as Spark Core, DataSet, MLlib, and GraphX jobs. +Spark jobs refer to those submitted via visual interfaces or RESTful APIs, supporting full-stack Spark functionalities including Spark Core, DataSet, MLlib, and GraphX. CU -- -CU is the unit of compute resources in DLI, where 1 CU equals 1 vCPU and 4 GB of memory. The higher the specifications of compute resources, the better its computing power. +CU represents the unit of compute resources in DLI. One CU equals one vCPU paired with 4 GB of memory. Higher specifications correspond to increased computational power. Constants and Variables ----------------------- The differences between constants and variables are as follows: -- During the running of a program, the value of a constant cannot be changed. -- Variables are readable and writable, whereas constants are read-only. A variable is a memory address that contains a segment of data that can be changed during program running. For example, in **int a = 123**, **a** is an integer variable. +- Constants retain their value throughout program execution and cannot be altered. They are strictly read-only. +- Variables are both readable and writable. A variable represents a specific memory address where the stored value can be updated at any time during runtime. For example, in **int a = 123**, **a** is an integer variable. Table Lifecycle --------------- -The table lifecycle management feature in DLI refers to the automatic recycling of tables or partitions that have not been updated for a specified period of time since their last update. This specified period is known as the lifecycle. This feature simplifies the process of recycling data and frees up storage space. Additionally, it provides data backup and recovery functions to prevent data loss due to accidental operations. +The table lifecycle management feature in DLI automatically reclaims table (or partition) data if it remains unchanged after a specified period from its last update. This duration is termed the lifecycle. The feature simplifies storage space reclamation and data recycling processes while providing backup and recovery options to prevent accidental data loss. diff --git a/umn/source/service_overview/compute_resource_types_and_specifications.rst b/umn/source/service_overview/compute_resource_types_and_specifications.rst index 9532f50..1e2e6ac 100644 --- a/umn/source/service_overview/compute_resource_types_and_specifications.rst +++ b/umn/source/service_overview/compute_resource_types_and_specifications.rst @@ -5,7 +5,7 @@ Compute Resource Types and Specifications ========================================= -DLI compute resources are the foundation for job execution. Both DLI's elastic resource pools and queues fall under compute resources. This section introduces the types and specifications of DLI compute resources. +When performing big data analysis with DLI, you need to select the appropriate compute resources based on specific business needs. DLI offers a variety of resource types and product specifications, including elastic resource pools and queues. This section provides a detailed overview of these options to help you make informed decisions tailored to your requirements. What Are Elastic Resource Pools and Queues? ------------------------------------------- @@ -26,7 +26,7 @@ Before we dive into the compute resource modes of DLI, let us first understand t - An elastic resource pool can simultaneously support SQL, Spark, and Flink jobs. The specific job types supported depend on the queue types created within the elastic resource pool. - Refer to :ref:`DLI Compute Resource Modes and Supported Queue Types `. + See :ref:`DLI Compute Resource Modes and Supported Queue Types `. - .. _dli_07_0027__li1752524818338: @@ -48,7 +48,7 @@ Before we dive into the compute resource modes of DLI, let us first understand t - **For SQL:** - For SQL queues are used to execute SQL jobs and supports specifying engine types including Spark and HetuEngine. + For SQL queues are designed to execute SQL jobs. They support the Spark engine. This type of queues is suitable for businesses that require fast data query and analysis, as well as regular cache clearing or environment resetting. @@ -74,7 +74,7 @@ DLI offers three compute resource management modes, each with unique advantages A pooled management mode for compute resources that provides dynamic scaling capabilities. Queues within the same elastic resource pool share compute resources. By appropriately setting up compute resource allocation policies for queues, you can enhance compute resource utilization and handle peak business demands efficiently. - Use cases: suitable for scenarios with significant fluctuations in business volume, such as periodic data batch processing tasks or real-time data processing needs. - - Supported queue types: for SQL (Spark), for SQL (HetuEngine), and for general purpose. For details about DLI queue types, see :ref:`Queue Types `. + - Supported queue types: for SQL (Spark) and for general purpose. For details about DLI queue types, see :ref:`Queue Types `. .. note:: @@ -128,9 +128,7 @@ DLI Compute Resource Modes and Supported Queue Types +====================================+======================+======================================================================+==========================================================================================================================================================+ | **Elastic resource pool mode** | For SQL (Spark) | Resources are shared among multiple queues for a single user. | Suitable for scenarios with significant fluctuations in business demand, where resources need to be flexibly adjusted to meet peak and off-peak demands. | | | | | | - | | For SQL (HetuEngine) | Resources are dynamically allocated and can be flexibly adjusted. | | - | | | | | - | | For general purpose | | | + | | For general purpose | Resources are dynamically allocated and can be flexibly adjusted. | | +------------------------------------+----------------------+----------------------------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------+ | **Global sharing mode** | default queue | Resources are shared among multiple queues for multiple users. | Suitable for temporary or testing projects where data size is uncertain or data processing is only required occasionally. | | | | | | diff --git a/umn/source/service_overview/features.rst b/umn/source/service_overview/features.rst new file mode 100644 index 0000000..4014bd5 --- /dev/null +++ b/umn/source/service_overview/features.rst @@ -0,0 +1,109 @@ +:original_name: dli_01_0698.html + +.. _dli_01_0698: + +Features +======== + +This section outlines the key features supported by DLI. For detailed information on regional availability of each feature, you can refer to the console. + +Elastic Resource Pools and Queues +--------------------------------- + +Before submitting jobs using DLI, you need to prepare the necessary compute resources. + +- **Elastic Resource Pool** + + An elastic resource pool provides the required compute resources (CPU and memory) for running DLI jobs. It offers robust computing power, high availability, and flexible resource management capabilities. This makes it ideal for large-scale computing tasks and business use cases that require long-term resource planning. Additionally, it can dynamically adapt to changing demands for compute resources. + +- **Queue** + + After creating an elastic resource pool, you can create multiple queues within it. Each queue is associated with specific jobs and data processing tasks, serving as the fundamental unit for allocating and utilizing resources in the pool. In other words, a queue represents the actual compute resources needed to execute a job. + + Within the same elastic resource pool, compute resources are shared among queues. By configuring appropriate allocation policies for these queues, you can optimize the utilization of compute resources. + +- **default Queue** + + DLI comes pre-configured with a default queue named **default**, where resources are allocated on-demand. If you are unsure about the required queue capacity or lack available space to create queues, you can use this **default** queue to run your jobs. However, since the **default** queue is shared among all users, there may be instances of resource contention. As a result, access to resources cannot always be guaranteed for every operation. + +DLI Metadata Management +----------------------- + +DLI metadata serves as the foundation for developing SQL jobs and Spark jobs. Before executing a job, you need to define databases and tables based on your business requirements. + +In addition to managing its own metadata, DLI supports integration with LakeFormation for unified metadata management. This enables seamless connectivity with various compute engines and big data cloud services, making it efficient and convenient to build data lakes and operate related businesses. + +- **DLI Metadata** + + DLI metadata serves as the foundation for developing SQL jobs and Spark jobs. Prior to running a job, you must define databases and tables according to your specific use case. Data catalog: A data catalog is a metadata management object that can contain multiple databases. You can create and manage multiple catalogs in DLI to isolate different sets of metadata. + + - Database + + A database is a structured repository built on computer storage devices used to organize, store, and manage data. It typically stores, retrieves, and manages structured data, consisting of multiple interrelated data tables connected through keys and indexes. + + - Table + + Tables are one of the most critical components of a database, composed of rows and columns. Each row represents a data entry, while each column defines an attribute or characteristic of the data. Tables are used to organize and store specific types of data, enabling efficient query and analysis. While a database provides the framework, tables constitute its actual content—a single database may include one or more tables. + + - Metadata + + Metadata refers to data that describes other data. It primarily includes information about the source, size, format, or other characteristics of the data itself. In the context of database fields, metadata helps interpret the contents of a data warehouse. When creating a table, metadata is defined by specifying three elements: column names, data types, and column descriptions. + +- **Connecting DLI to LakeFormation for Metadata Management** + + After creating an elastic resource pool, you can create multiple queues within it. Each queue is associated with specific jobs and data processing tasks, serving as the fundamental unit for allocating and utilizing resources in the pool. In other words, a queue represents the actual compute resources needed to execute a job. Within the same elastic resource pool, compute resources are shared among queues. By configuring appropriate allocation policies for these queues, you can optimize the utilization of compute resources. + +DLI SQL Job +----------- + +DLI SQL jobs, also known as DLI Spark SQL jobs, allow you to execute data queries and other operations by executing SQL statements in the SQL editor. It supports SQL:2003 and is fully compatible with Spark SQL. + +DLI Spark Job +------------- + +Spark is a unified analytics engine designed for large-scale data processing, focusing on query, computation, and analysis. DLI has undergone extensive performance optimization and service-oriented enhancements over the open-source Spark, maintaining compatibility with the Apache Spark ecosystem and APIs while boosting performance by 2.5 times, enabling exabyte-scale data queries and analyses within hours. + +DLI Flink Job +------------- + +DLI Flink jobs are specifically designed for real-time data stream processing, making them ideal for scenarios that require low latency and quick response. They support cross-source connectivity with various cloud services, forming a robust streaming ecosystem. These jobs are ideal for applications such as real-time monitoring and online analytics. + +- **Flink OpenSource Job** + + DLI provides standard connectors and a rich set of APIs, facilitating seamless integration with other data systems. + +- **Flink Jar Job** + + You can submit Flink jobs compiled into JAR files, offering greater flexibility and customization capabilities. This option is well-suited for complex data processing scenarios requiring custom functions, user-defined functions (UDFs), or specific library integrations. Leveraging Flink's ecosystem, advanced stream processing logic and state management can be achieved. + +- **Flink Python Job** + + Starting from Flink 1.17, DLI introduces support for PyFlink jobs, providing you with more flexible and powerful data processing tools. You can directly submit PyFlink jobs through the DLI job management page. Additionally, it allows specifying third-party Python and Java dependencies, customizing Python virtual environments, and uploading compressed data files, significantly enhancing the convenience of job submission and execution flexibility. + + Additionally, the DLI Flink image comes pre-installed with default Python execution environments supporting versions 3.7, 3.8, 3.9, and 3.10, meeting diverse development needs. This empowers Python developers to efficiently use DLI for data processing and analysis tasks. + + Flink Python jobs are particularly suitable for scenarios involving customized stream processing logic, complex state management, or specific library integrations. You are required to write and build Python job packages independently. Before submitting a Flink Python job, upload the Jar job package to OBS and submit it along with the data and job parameters to execute the job. + +Datasource Connection +--------------------- + +Before performing cross-source analysis using DLI, you need to establish a datasource connection to enable network communication between data sources. + +DLI's enhanced datasource connections use VPC peering connections to directly connect the VPC networks of DLI queues and destination data sources. This point-to-point approach facilitates seamless data exchange, offering more flexible use cases and superior performance compared to basic datasource connections. + +Note: The system's **default** queue does not support creating datasource connections. Establishing such connections requires functionalities like VPCs, subnets, routing, and VPC peering connections. Therefore, you must have the **VPC Administrator** permission for VPC. You can set these permissions by referring to "Service Authorization". + +Permission Management +--------------------- + +DLI is equipped with a robust permission control mechanism. Additionally, DLI supports fine-grained authentication through Identity and Access Management (IAM). You can create IAM policies to manage access controls within DLI. Both permission control mechanisms can operate simultaneously without conflict. + +Custom DLI Agency +----------------- + +To perform cross-source analysis, DLI requires agency permissions to access other cloud services. This allows DLI to act on behalf of users or services in other cloud services, enabling it to read/write data and execute specific operations during job execution. Custom DLI agency ensures secure and efficient access to other cloud services during cross-source analysis. + +Custom Image +------------ + +DLI supports containerized cluster deployments. In these clusters, components related to Spark and Flink jobs run within containers. By downloading custom images provided by DLI, you can modify the runtime environment of Spark and Flink containers. For example, adding Python packages or C libraries for machine learning into the custom image enables easy functional expansion tailored to your needs. diff --git a/umn/source/service_overview/index.rst b/umn/source/service_overview/index.rst index c4c47ea..976fa5f 100644 --- a/umn/source/service_overview/index.rst +++ b/umn/source/service_overview/index.rst @@ -8,10 +8,11 @@ Service Overview - :ref:`What Is Data Lake Insight ` - :ref:`Advantages ` - :ref:`Use Cases ` -- :ref:`Notes and Constraints ` +- :ref:`Features ` - :ref:`Compute Resource Types and Specifications ` - :ref:`Permission Management ` -- :ref:`Quotas ` +- :ref:`Notes and Constraints ` +- :ref:`Quota Management ` - :ref:`Related Services ` - :ref:`Basic Concepts ` @@ -22,9 +23,10 @@ Service Overview what_is_data_lake_insight advantages use_cases - notes_and_constraints + features compute_resource_types_and_specifications permission_management - quotas + notes_and_constraints + quota_management related_services basic_concepts diff --git a/umn/source/service_overview/notes_and_constraints.rst b/umn/source/service_overview/notes_and_constraints.rst index d3b0e7b..ced68bb 100644 --- a/umn/source/service_overview/notes_and_constraints.rst +++ b/umn/source/service_overview/notes_and_constraints.rst @@ -12,31 +12,39 @@ For more notes and constraints on elastic resource pools, see :ref:`Notes and Co .. table:: **Table 1** Notes and constraints on elastic resource pools - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Item | Description | - +===================================+============================================================================================================================================================================================================================================================================================================================================================================================+ - | Resource specifications | - An elastic resource pool currently supports up to 32,000 CUs. | - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Managing elastic resource pools | - You cannot change the region of an elastic resource pool once the pool is created. | - | | - Flink 1.10 or later jobs can run in elastic resource pools. | - | | - The CIDR block of an elastic resource pool cannot be changed once set. | - | | - You can view only the scaling history of an elastic resource pool within 30 days. | - | | - Elastic resource pools cannot directly access the Internet. | - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Elastic resource pool scaling | - Changes to elastic resource pool CUs can occur when setting the CU, adding or deleting queues in an elastic resource pool, or modifying the scaling policies of queues in an elastic resource pool, or when the system automatically triggers elastic resource pool scaling. However, in some cases, the system cannot guarantee that the scaling will reach the target CUs as planned. | - | | | - | | - If there are not enough physical resources, an elastic resource pool may not be able to scale out to the desired target size. | - | | | - | | - The system does not guarantee that an elastic resource pool will be scaled in to the desired target size. | - | | | - | | The system checks the resource usage before scaling in the elastic resource pool to determine if there is enough space for scaling in. If the existing resources cannot be scaled in according to the minimum scaling step, the pool may not be scaled in successfully or only partially. | - | | | - | | The scaling step may vary depending on the resource specifications, usually 16 CUs, 32 CUs, 48 CUs, 64 CUs, and more. | - | | | - | | For example, if the elastic resource pool has a capacity of 192 CUs and the queues in the pool are using 68 CUs due to running jobs, the plan is to scale in to 64 CUs. | - | | | - | | When executing a scaling in task, the system determines that there are 124 CUs remaining and scales in by the minimum step of 64 CUs. However, the remaining 60 CUs cannot be scaled in any further. Therefore, after the elastic resource pool executes the scaling in task, its capacity is reduced to 128 CUs. | - +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Item | Description | + +===================================================+============================================================================================================================================================================================================================================================================================================================================================================================+ + | Resource specifications | - An elastic resource pool currently supports up to 32,000 CUs. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Managing elastic resource pools | - You cannot change the region of an elastic resource pool once the pool is created. | + | | - Flink 1.10 or later jobs can run in elastic resource pools. | + | | - The CIDR block of an elastic resource pool cannot be changed once set. | + | | - You can view only the scaling history of an elastic resource pool within 30 days. | + | | - Elastic resource pools cannot directly access the Internet. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Elastic resource pool scaling | - Changes to elastic resource pool CUs can occur when setting the CU, adding or deleting queues in an elastic resource pool, or modifying the scaling policies of queues in an elastic resource pool, or when the system automatically triggers elastic resource pool scaling. However, in some cases, the system cannot guarantee that the scaling will reach the target CUs as planned. | + | | | + | | - If there are not enough physical resources, an elastic resource pool may not be able to scale out to the desired target size. | + | | | + | | - The system does not guarantee that an elastic resource pool will be scaled in to the desired target size. | + | | | + | | The system checks the resource usage before scaling in the elastic resource pool to determine if there is enough space for scaling in. If the existing resources cannot be scaled in according to the minimum scaling step, the pool may not be scaled in successfully or only partially. | + | | | + | | The scaling step may vary depending on the resource specifications, usually 16 CUs, 32 CUs, 48 CUs, 64 CUs, and more. | + | | | + | | For example, if the elastic resource pool has a capacity of 192 CUs and the queues in the pool are using 68 CUs due to running jobs, the plan is to scale in to 64 CUs. | + | | | + | | When executing a scaling in task, the system determines that there are 124 CUs remaining and scales in by the minimum step of 64 CUs. However, the remaining 60 CUs cannot be scaled in any further. Therefore, after the elastic resource pool executes the scaling in task, its capacity is reduced to 128 CUs. | + | | | + | | - If jobs such as Flink or Spark Streaming jobs are continuously running on a node, that node cannot be scaled in as scheduled. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Constraints for setting elastic resource pool CUs | - The minimum CUs (minCU) of an elastic resource pool must be **less than or equal to** the actual CUs. If expanding minCU exceeds the current actual CUs, you must first increase the actual CUs. Otherwise, the modification will fail. | + | | - The sum of all queues' minimum CUs in an elastic resource pool must not exceed the pool's minCU. | + | | - Any single queue's maxCU cannot exceed the pool's maxCU. | + | | - Adjustments to a queue's CU range, changes to the pool's specifications, or modifications to the pool's CU settings take effect at the next full hour. | + | | - Increasing the number of queues to adjust the pool's actual CUs takes immediate effect. | + +---------------------------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ Queues ------ @@ -189,7 +197,7 @@ For more notes and constraints on datasource authentication, see :ref:`Datasourc | | - CSS: applies to 6.5.4 or later CSS clusters with the security mode enabled. | | | - Kerberos: applies to MRS security clusters with Kerberos authentication enabled. | | | - Kafka_SSL: applies to Kafka with SSL enabled. | - | | - Password: applies to GaussDB(DWS), RDS, DDS, and DCS. | + | | - Password: applies to DWS, RDS, DDS, and DCS. | +-----------------------------------+----------------------------------------------------------------------------------------------------------------------+ SQL Syntax @@ -211,14 +219,14 @@ Other .. table:: **Table 9** Other notes and constraints - +-----------------------------------+-------------------------------------------------------------------+ - | Item | Description | - +===================================+===================================================================+ - | Quota | For quota notes and constraints, see :ref:`Quotas `. | - +-----------------------------------+-------------------------------------------------------------------+ - | Browser version | Recommended browsers and their versions: | - | | | - | | - Google Chrome 43.0 or later | - | | - Mozilla Firefox 38.0 or later | - | | - Internet Explorer 9.0 or later | - +-----------------------------------+-------------------------------------------------------------------+ + +-----------------------------------+-----------------------------------------------------------------------------+ + | Item | Description | + +===================================+=============================================================================+ + | Quota | For quota notes and constraints, see :ref:`Quota Management `. | + +-----------------------------------+-----------------------------------------------------------------------------+ + | Browser version | Recommended browsers and their versions: | + | | | + | | - Google Chrome 43.0 or later | + | | - Mozilla Firefox 38.0 or later | + | | - Internet Explorer 9.0 or later | + +-----------------------------------+-----------------------------------------------------------------------------+ diff --git a/umn/source/service_overview/permission_management.rst b/umn/source/service_overview/permission_management.rst index d710b46..036fbe5 100644 --- a/umn/source/service_overview/permission_management.rst +++ b/umn/source/service_overview/permission_management.rst @@ -5,23 +5,23 @@ Permission Management ===================== -If you need to assign different permissions to employees in your enterprise to access your DLI resources, IAM is a good choice for fine-grained permissions management. IAM provides identity authentication, permissions management, and access control, helping you securely access to your cloud resources. +After purchasing DLI resources, you can use IAM to assign different access rights to employees, ensuring proper isolation between roles. IAM provides identity authentication, permission assignment, and access control, helping you securely manage access to cloud resources. -With IAM, you can use your account to create IAM users for your employees, and assign permissions to the users to control their access to specific resource types. For example, some software developers in your enterprise need to use DLI resources but must not delete them or perform any high-risk operations. To achieve this result, you can create IAM users for the software developers and grant them only the permissions required for using DLI resources. +With IAM, you can create users under your account and apply policies to define their access scope. For example, developers may need permission to use DLI but should not be allowed to delete it. In this case, you can create IAM users for developers and assign a policy that grants usage rights while restricting deletion. -If your account does not require individual IAM users for permissions management, you may skip over this section. +If your account already meets your needs, you can skip creating separate IAM users without affecting other DLI functions. DLI Permissions --------------- -New IAM users do not have any permissions assigned by default. You need to first add them to one or more groups and attach policies or roles to these groups. The users then inherit permissions from the groups and can perform specified operations on cloud services based on the permissions they have been assigned. +By default, newly created IAM users have no permissions. To enable access, you must add them to a user group and assign roles or policies. This process is called authorization. Once authorized, users can perform operations within the scope of their assigned permissions. -DLI is a project-level service deployed and accessed in specific physical regions. To assign ServiceStage permissions to a user group, specify the scope as region-specific projects and select projects for the permissions to take effect. If **All projects** is selected, the permissions will take effect for the user group in all region-specific projects. When accessing DLI, the users need to switch to a region where they have been authorized to use cloud services. +DLI is deployed by physical region and operates as a project-level service. When authorizing, select **Region-specific projects** for **Scope**, then assign permissions within the specific project. These permissions apply only to that project. If you assign permissions to **All projects**, they apply across all regions. To access DLI, users must first switch to the authorized region. -Permission types: Based on the granularity of authorization, they are divided into roles and policies. +Permission types: Permissions are categorized into roles and policies. -- Roles: A coarse-grained authorization strategy that defines permissions by job responsibility. This strategy offers limited service-level roles for authorization. If one role has a dependency role required for accessing SA, assign both roles to the users. Roles are not suitable for fine-grained authorization and least privilege access. -- Policies: A fine-grained authorization strategy that defines permissions required to perform operations on specific cloud resources under certain conditions. This type of authorization is more flexible and is ideal for least privilege access. For example, you can grant DLI users only the permissions for managing a certain type of ECSs. For the actions supported by DLI APIs, see "Permissions Policies and Supported Actions" in the *Data Lake Insight API Reference*. +- Roles: A coarse-grained authorization mechanism based on job responsibilities. This strategy offers limited service-level roles for authorization. Granting a role may require assigning additional dependency roles. Roles cannot fully meet fine-grained authorization or strict least-privilege access. +- Policies: A fine-grained authorization mechanism that defines permissions at the level of specific operations, resources, and conditions. Policies provide flexible control and support enterprise-level least-privilege security. For example, you can restrict IAM users to perform only certain management operations on specific DLI resources. For details about the actions supported by DLI APIs, see "Permissions Policies and Supported Actions" in the *Data Lake Insight API Reference*. .. table:: **Table 1** DLI system permissions @@ -50,7 +50,7 @@ Permission types: Based on the granularity of authorization, they are divided in | | - Scope: project-level service | | | +---------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-----------------------+------------------------------------------------------------------+ -:ref:`Table 2 ` lists the common operations supported by each system policy. You can choose required system policies according to this table. +:ref:`Table 2 ` lists the common operations supported by each system policy. You can use this table to select appropriate system policies. .. _dli_07_0006__en-us_topic_0206791772_table168060107500: diff --git a/umn/source/service_overview/quota_management.rst b/umn/source/service_overview/quota_management.rst new file mode 100644 index 0000000..556d1b8 --- /dev/null +++ b/umn/source/service_overview/quota_management.rst @@ -0,0 +1,50 @@ +:original_name: dli_07_0009.html + +.. _dli_07_0009: + +Quota Management +================ + +What Is a Quota? +---------------- + +Quotas are enforced for service resources on the platform to prevent unforeseen spikes in resource usage. Quotas can limit the quantity and capacity of resources available to users. + +If your current quota does not meet your needs, you can apply for an increase. + +How Do I View My Quotas? +------------------------ + +#. Log in to the management console. + +#. Click |image1| in the upper left corner and select a region and a project. + +#. Click the **My Quota** icon |image2| in the upper right corner of the page. + + This will take you to the **Service Quota** page. + +#. Here, you can view the total quota and usage details for each resource. + + If your current quota is insufficient, follow the steps below to request an increase. + +How Do I Apply for a Quota Increase? +------------------------------------ + +The system currently does not support online quota adjustments. If you need to modify your quota, please contact our customer service team via phone or email. They will promptly process your request and keep you updated on its progress through either a call or email. + +Before reaching out, ensure you have the following details ready: + +- Domain name, project name, and project ID + + To obtain the preceding information, log in to the management console, click the username in the upper right corner, and choose **My Credentials** from the drop-down list. + +- Quota information, including: + + - Service name + - Quota type + - Desired quota value + +`Learn how to obtain the service hotline and email address. `__ + +.. |image1| image:: /_static/images/en-us_image_0000001487683840.png +.. |image2| image:: /_static/images/en-us_image_0000001488163536.png diff --git a/umn/source/service_overview/quotas.rst b/umn/source/service_overview/quotas.rst deleted file mode 100644 index 5e1e5f9..0000000 --- a/umn/source/service_overview/quotas.rst +++ /dev/null @@ -1,50 +0,0 @@ -:original_name: dli_07_0009.html - -.. _dli_07_0009: - -Quotas -====== - -What Is a Quota? ----------------- - -A quota limits the quantity of a resource available to users, thereby preventing spikes in the usage of the resource. - -You can also request for an increased quota if your existing quota cannot meet your service requirements. - -How Do I View My Quotas? ------------------------- - -#. Log in to the management console. - -#. Click |image1| in the upper left corner and select a region and a project. - -#. Click the **My Quota** icon |image2| in the upper right corner of the page. - - The **Service Quota** page is displayed. - -#. View the used and total quota of each type of resources on the displayed page. - - If a quota cannot meet service requirements, increase a quota. - -How Do I Apply for a Higher Quota? ----------------------------------- - -The system does not support online quota adjustment. To increase a resource quota, dial the hotline or send an email to the customer service. We will process your application and inform you of the progress by phone call or email. - -Before dialing the hotline number or sending an email, ensure that the following information has been obtained: - -- Domain name, project name, and project ID - - To obtain the preceding information, log in to the management console, click the username in the upper right corner, and choose **My Credentials** from the drop-down list. - -- Quota information, including: - - - Service name - - Quota type - - Required quota - -`Learn how to obtain the service hotline and email address. `__ - -.. |image1| image:: /_static/images/en-us_image_0000001487683840.png -.. |image2| image:: /_static/images/en-us_image_0000001488163536.png diff --git a/umn/source/service_overview/related_services.rst b/umn/source/service_overview/related_services.rst index 61a369e..65b852e 100644 --- a/umn/source/service_overview/related_services.rst +++ b/umn/source/service_overview/related_services.rst @@ -69,17 +69,17 @@ Relational Database Service (RDS) works as the data source and data storage syst - Data source: DLI allows you to import RDS data using DataFrame or SQL. - Query result storage: DLI uses the SQL INSERT syntax to store query result data to RDS tables. -For how to access RDS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. +For details about how to access RDS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. -GaussDB(DWS) ------------- +DWS +--- -GaussDB(DWS) works as the data source and data storage system for DLI, and delivers the following capabilities: +DWS works as the data source and data storage system for DLI, and delivers the following capabilities: -- Data source: DLI allows you to import GaussDB(DWS) data using DataFrame or SQL. -- Query result storage: DLI uses the SQL INSERT syntax to store query result data to GaussDB(DWS) tables. +- Data source: DLI allows you to import DWS data using DataFrame or SQL. +- Query result storage: DLI uses the SQL INSERT syntax to store query result data to DWS tables. -For how to access GaussDB(DWS) data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. +For details about how to access DWS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. CSS --- @@ -89,7 +89,7 @@ CSS works as the data source and data storage system for DLI, and delivers the f - Data source: DLI allows you to import CSS data using DataFrame or SQL. - Query result storage: DLI uses the SQL INSERT syntax to store query result data to CSS tables. -For how to access CSS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. +For details about how to access CSS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. DCS --- @@ -99,7 +99,7 @@ Distributed Cache Service (DCS) works as the data source and data storage system - Data source: DLI allows you to import DCS data using DataFrame or SQL. - Query result storage: DLI uses the SQL INSERT syntax to store query result data to DCS tables. -For how to access DCS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. +For details about how to access DCS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. DDS --- @@ -109,7 +109,7 @@ Document Database Service (DDS) works as the data source and data storage system - Data source: DLI allows you to import DDS data using DataFrame or SQL. - Query result storage: DLI uses the SQL INSERT syntax to store query result data to DDS tables. -For how to access DDS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. +For details about how to access DDS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. MRS --- @@ -119,7 +119,7 @@ MapReduce Service (MRS) works as the data source and data storage system for DLI - Data source: DLI allows you to import MRS data using DataFrame or SQL. - Query result storage: DLI uses the SQL INSERT syntax to store query result data to MRS tables. -For how to access MRS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. +For details about how to access MRS data through a DLI datasource connection, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. DataArts Studio --------------- diff --git a/umn/source/service_overview/what_is_data_lake_insight.rst b/umn/source/service_overview/what_is_data_lake_insight.rst index f2a8d9d..baad370 100644 --- a/umn/source/service_overview/what_is_data_lake_insight.rst +++ b/umn/source/service_overview/what_is_data_lake_insight.rst @@ -8,13 +8,15 @@ What Is Data Lake Insight DLI Introduction ---------------- -Data Lake Insight (DLI) is a serverless data processing and analysis service fully compatible with `Apache Spark `__ and `Apache Flink `__ ecosystems. It frees you from managing any servers. +Data Lake Insight (DLI) is a fully-managed, serverless service for data processing and analytics. It seamlessly combines stream processing, batch processing, and interactive analysis into a single platform. DLI provides full support for `Apache Spark `__ and `Apache Flink `__. No server management is required. You can get started with DLI right away. -DLI supports multiple querying methods including standard SQL, Spark SQL, and Flink SQL, with compatibility with mainstream data formats. DLI supports SQL statements and Spark applications for heterogeneous data sources, including CloudTable, RDS, GaussDB(DWS), CSS, OBS, custom databases on ECSs, and offline databases. +DLI works with standard SQL, Spark SQL, and Flink SQL. It offers multiple access options and supports mainstream data formats. With DLI, you can explore diverse data sources such as CloudTable, RDS, DWS, CSS, OBS, ECS-hosted databases, or on-premises databases. Eliminate complex extraction, transformation, or loading processes. Analyze your data directly using SQL or programs. Core Functions -------------- +For details about DLI functions, see :ref:`Features `. + .. table:: **Table 1** DLI core functions +---------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ @@ -24,7 +26,7 @@ Core Functions | | | | | - Auto scaling: DLI ensures you always have enough capacity on hand to deal with any traffic spikes. | +---------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | DLI supports multiple compute engines. | DLI is fully compatible with ecosystems like Apache Spark and Apache Flink, and supports standard SQL, Spark SQL, and Flink SQL. It is compatible with mainstream data formats such as CSV, JSON, Parquet, and ORC. | + | DLI supports multiple compute engines. | DLI provides full support for ecosystems like Apache Spark and Apache Flink. DLI works with standard SQL, Spark SQL, and Flink SQL. It supports mainstream data formats such as CSV, JSON, Parquet, and ORC. | | | | | | - **Spark** is a unified analytics engine designed for large-scale data processing, focusing on query, compute, and analysis. DLI has undergone extensive performance optimization and service-oriented enhancements over the open-source Spark, maintaining compatibility with the Apache Spark ecosystem and APIs while boosting performance by 2.5 times, enabling exabyte-scale data queries and analyses within hours. | | | - **Flink** is a distributed compute engine that can be used for batch processing, which involves handling static datasets and historical datasets. It can also be used for stream processing, enabling the real-time processing of live data streams and the immediate generation of data results. DLI has enhanced features and security based on the open-source Flink and offers the Stream SQL feature needed for data processing. | @@ -40,7 +42,7 @@ Core Functions | | - Submitting DLI jobs using DataArts Studio | | | - Connecting to BI tools for visual analysis | +---------------------------------------------------------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | DLI can connect to multiple data sources for cross-source data analysis. | - Spark datasource connection: Data sources such as GaussDB(DWS), RDS, and CSS can be accessed through DLI. | + | DLI can connect to multiple data sources for cross-source data analysis. | - Spark datasource connection: Data sources such as DWS, RDS, and CSS can be accessed through DLI. | | | - Flink supports cross-source connectivity with various cloud services, forming a rich streaming ecosystem. DLI's streaming ecosystem is divided into cloud service ecosystems and open-source ecosystems: | | | | | | - Cloud service ecosystem: DLI supports connectivity with other services in Flink SQL. You can directly use SQL to read and write data from these cloud services. | @@ -69,32 +71,31 @@ DLI includes the following core modules: .. table:: **Table 2** DLI core modules - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Module | Description | - +===================================+=================================================================================================================================================================================================================================================================================================================================================================================================+ - | Ecosystem tools | DLI leverages its robust serverless architecture and multimodal engine support to fulfill the diverse needs of various industries, driving their digital transformation and fostering innovation. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Compute engine | - Spark: supports batch processing and interactive analysis of large-scale data and provides high-performance distributed computing capabilities. | - | | - Flink: supports real-time stream processing, capable of handling large-scale real-time data streams, with support for event time processing and state management. | - | | - HetuEngine: supports interactive data analysis, swiftly handles complex SQL queries, and facilitates connections and queries across various data sources. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Unified resource management | - Resource decoupling: DLI adopts a decoupled compute and storage architecture, decoupling compute resources from storage resources. This allows for flexible adjustment of the ratio between compute and storage resources based on actual needs, enhancing resource utilization and reducing costs. | - | | - Elastic scaling: DLI compute resources are based on containerized Kubernetes and possess ultimate elastic scaling capabilities. Resources can be automatically adjusted based on job demands. | - | | - Multi-tenant support: Compute resources can be isolated by tenant to ensure independence among different tenants. Each tenant can independently manage their own compute resources, enabling fine-grained resource management and facilitating inter-departmental data sharing and permissions management within enterprises. | - | | - Pay-per-use compute resources: You only pay for the compute resources you actually use, with no need to pre-purchase or manage servers, enhancing usage efficiency. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Unified metadata management | - Multi-source metadata integration: DLI supports centralized management of metadata from various data sources, including cloud-based data sources (such as OBS, RDS, GaussDB(DWS), and CSS) and on-premises data sources (such as self-built databases and Redis). You can manage and analyze metadata across different data sources without the need to migrate data to a unified data lake. | - | | - Metadata synchronization: DLI provides metadata management to ensure the timeliness and consistency of metadata. | - | | - Metadata query and management: DLI offers standard SQL APIs, enabling you to query and manage metadata using SQL statements. You can add, delete, modify, and query metadata to facilitate data governance and analysis. | - | | - Data security and permission management: Permissions on data catalogs, databases, and tables can be managed. You can assign different permissions to various tenants and user groups to ensure data security and compliance. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Storage service | OBS and databases are used to store structured or unstructured data for data analysis, providing persistent data storage services. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Data source connection | - Cloud data sources can be connected. For example, OBS can be used to store and manage unstructured data. Relational database service (RDS) can be used to store and manage structured data. GaussDB(DWS) can be used to efficiently query and analyze data. | - | | - On-premises data sources, such as self-built databases (MySQL, PostgreSQL, and HDFS), can be connected. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Data applications | DLI can connect to mainstream BI tools in the industry to flexibly meet data presentation needs. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Module | Description | + +===================================+========================================================================================================================================================================================================================================================================================================================================================================================+ + | Ecosystem tools | DLI leverages its robust serverless architecture and multimodal engine support to fulfill the diverse needs of various industries, driving their digital transformation and fostering innovation. | + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Compute engine | - Spark: supports batch processing and interactive analysis of large-scale data and provides high-performance distributed computing capabilities. | + | | - Flink: supports real-time stream processing, capable of handling large-scale real-time data streams, with support for event time processing and state management. | + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Unified resource management | - Resource decoupling: DLI adopts a decoupled compute and storage architecture, decoupling compute resources from storage resources. This allows for flexible adjustment of the ratio between compute and storage resources based on actual needs, enhancing resource utilization and reducing costs. | + | | - Elastic scaling: DLI compute resources are built upon containerized Kubernetes and possess elastic scaling capabilities. Resources can be automatically adjusted based on job demands. | + | | - Multi-tenant support: Compute resources can be isolated by tenant to ensure independence among different tenants. Each tenant can independently manage their own compute resources, enabling fine-grained resource management and facilitating inter-departmental data sharing and permissions management within enterprises. | + | | - Pay-per-use compute resources: You only pay for the compute resources you actually use, with no need to pre-purchase or manage servers, enhancing usage efficiency. | + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Unified metadata management | - Multi-source metadata integration: DLI supports centralized management of metadata from various data sources, including cloud-based data sources (such as OBS, RDS, DWS, and CSS) and on-premises data sources (such as self-built databases and Redis). You can manage and analyze metadata across different data sources without the need to migrate data to a unified data lake. | + | | - Metadata synchronization: DLI provides metadata management to ensure the timeliness and consistency of metadata. | + | | - Metadata query and management: DLI offers standard SQL APIs, enabling you to query and manage metadata using SQL statements. You can add, delete, modify, and query metadata to facilitate data governance and analysis. | + | | - Data security and permission management: Permissions on data catalogs, databases, and tables can be managed. You can assign different permissions to various tenants and user groups to ensure data security and compliance. | + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Storage service | OBS and databases are used to store structured or unstructured data for data analysis, providing persistent data storage services. | + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Data source connection | - Cloud data sources can be connected. For example, OBS can be used to store and manage unstructured data. Relational database service (RDS) can be used to store and manage structured data. DWS can be used to efficiently query and analyze data. | + | | - On-premises data sources, such as self-built databases (MySQL, PostgreSQL, and HDFS), can be connected. | + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Data applications | DLI can connect to mainstream BI tools in the industry to flexibly meet data presentation needs. | + +-----------------------------------+----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ Accessing DLI ------------- diff --git a/umn/source/submitting_a_flink_job_on_the_dli_management_console/adding_tags_to_a_flink_job.rst b/umn/source/submitting_a_flink_job_on_the_dli_management_console/adding_tags_to_a_flink_job.rst index 8f2936c..92ef0cb 100644 --- a/umn/source/submitting_a_flink_job_on_the_dli_management_console/adding_tags_to_a_flink_job.rst +++ b/umn/source/submitting_a_flink_job_on_the_dli_management_console/adding_tags_to_a_flink_job.rst @@ -29,7 +29,7 @@ Managing a Job Tag DLI allows you to add, modify, or delete tags for jobs. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. Click the name of the job to be viewed. The **Job Details** page is displayed. #. Click **Tags** to display the tag information about the current job. #. On the page that appears, click **Add/Edit Tag** in the upper left corner. @@ -78,7 +78,7 @@ Searching for a Job by Tag If tags have been added to a job, you can search for the job by setting tag filtering conditions to quickly find it. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. In the upper right corner of the page, click the search box and select **Tags**. #. Choose a tag key and value as prompted. If no tag key or value is available, create a tag for the job. For details, see :ref:`Managing a Job Tag `. #. Choose other tags to generate a tag combination for job search. diff --git a/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_jar_job.rst b/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_jar_job.rst index 9e4a6c7..307185a 100644 --- a/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_jar_job.rst +++ b/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_jar_job.rst @@ -14,11 +14,11 @@ This section describes how to create a Flink Jar job on the DLI management conso Prerequisites ------------- -- When you use a Flink Jar job to access other external data sources, such as OpenTSDB, HBase, Kafka, GaussDB(DWS), RDS, CSS, CloudTable, DCS Redis, and DDS, you need to create a datasource connection to connect the job running queue to the external data source. +- When you use a Flink Jar job to access other external data sources, such as OpenTSDB, HBase, Kafka, DWS, RDS, CSS, CloudTable, DCS Redis, and DDS, you need to create a datasource connection to connect the job running queue to the external data source. - For details about the external data sources that can be accessed by Flink jobs, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. - - For how to create a datasource connection, see :ref:`Configuring the Network Connection Between DLI and Data Sources (Enhanced Datasource Connection) `. + - For details about how to create a datasource connection, see :ref:`Configuring the Network Connection Between DLI and Data Sources (Enhanced Datasource Connection) `. On the **Resources** > **Queue Management** page, locate the queue you have created, click **More** in the **Operation** column, and select **Test Address Connectivity** to check if the network connection between the queue and the data source is normal. For details, see :ref:`Testing the Network Connectivity Between a Queue and a Data Source `. @@ -38,7 +38,7 @@ Before creating and submitting jobs, you are advised to enable CTS to record DLI Creating a Flink Jar Job ------------------------ -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. In the upper right corner of the **Flink Jobs** page, click **Create Job**. @@ -106,7 +106,7 @@ Creating a Flink Jar Job | | | | | There are the following ways to manage JAR files: | | | | - | | - Upload packages to OBS: Upload Jar packages to an OBS bucket in advance and select the corresponding OBS path. | + | | - Upload packages to OBS: Upload JAR files to an OBS bucket in advance and select the corresponding OBS path. | | | - Upload packages to DLI: Upload JAR files to an OBS bucket in advance and create a package on the **Data Management** > **Package Management** page of the DLI management console. For details, see :ref:`Creating a DLI Package `. | | | | | | For Flink 1.15 or later, you can only select packages from OBS, instead of DLI. | @@ -128,7 +128,7 @@ Creating a Flink Jar Job | | | | | There are the following ways to manage JAR files: | | | | - | | - Upload packages to OBS: Upload Jar packages to an OBS bucket in advance and select the corresponding OBS path. | + | | - Upload packages to OBS: Upload JAR files to an OBS bucket in advance and select the corresponding OBS path. | | | - Upload packages to DLI: Upload JAR files to an OBS bucket in advance and create a package on the **Data Management** > **Package Management** page of the DLI management console. For details, see :ref:`Creating a DLI Package `. | | | | | | For Flink 1.15 or later, you can only select packages from OBS, instead of DLI. | @@ -154,7 +154,7 @@ Creating a Flink Jar Job +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Agency | If you choose Flink 1.15 or later to execute your job, you can create a custom agency to allow DLI to access other services. | | | | - | | For how to create a custom agency, see :ref:`Creating a Custom DLI Agency `. | + | | For details about how to create a custom agency, see :ref:`Creating a Custom DLI Agency `. | +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Runtime Configuration | User-defined optimization parameters. The parameter format is **key=value**. | | | | @@ -195,6 +195,8 @@ Creating a Flink Jar Job | | - The parallelism degree of Spark resources is jointly determined by the number of Executors and the number of Executor CPU cores. | +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Job Manager CUs | Number of management unit CUs. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the JobManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Parallelism | Number of tasks concurrently executed by each operator in a job. | | | | @@ -208,6 +210,9 @@ Creating a Flink Jar Job | | - If this option is selected, you need to set the following parameters: | | | | | | - **CU(s) per TM**: Number of resources occupied by each TaskManager. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that a single TaskManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | + | | | | | - **Slot(s) per TM**: Number of slots contained in each TaskManager. | | | | | | - If not selected, the system automatically uses the default values. | @@ -259,7 +264,7 @@ Creating a Flink Jar Job | | | | | **SMN Topic** | | | | - | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | + | | Select a custom SMN topic. For details about how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Auto Restart upon Exception | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | | | | @@ -281,9 +286,9 @@ Creating a Flink Jar Job Compared with the v1 template, the v2 template does not support the setting of the number of CUs. The v2 template supports the setting of **Job Manager Memory** and **Task Manager Memory**. - **v1**: applicable to Flink 1.12, 1.13, and 1.15. + **V1**: applicable to Flink 1.12 and 1.15. - **v2**: applicable to Flink 1.13, 1.15, and 1.17. + **V2**: applicable to Flink 1.15 and 1.17. You are advised to use the parameter settings of v2. @@ -295,79 +300,50 @@ Creating a Flink Jar Job .. table:: **Table 4** Parameter descriptions of v1 - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parameter | Description | - +===================================+=================================================================================================================================================================================================================================+ - | CUs | One CU consists of one vCPU and 4 GB of memory. The number of CUs ranges from 2 to 10000. | - | | | - | | .. note:: | - | | | - | | When **Task Manager Config** is selected, elastic resource pool queue management is optimized by automatically adjusting **CUs** to match **Actual CUs** after setting **Slot(s) per TM**. | - | | | - | | **CUs = Actual number of CUs = max[Job Manager CPU + Task Manager CPU, (Job Manager Memory + Task Manager Memory/4)]** | - | | | - | | - Job Manager CPU + Task Manager CPU = Actual TMs x CU(s) per TM + Job Manager CUs. | - | | - Job Manager Memory + Task Manager Memory = Actual TMs x Memory per TM + Job Manager Memory | - | | - If **Slot(s) per TM** is set, then: Actual TMs = Parallelism/Slot(s) per TM. | - | | - If **Slot(s) per TM** is not set, then: Actual TMs = (CUs - Job Manager CUs)/CU(s) per TM. | - | | - If **Memory per TM** and **Job Manager Memory** in the optimization parameters are not set, then: Memory per TM = CU(s) per TM x 4. Job Manager Memory = Job Manager CUs x 4. | - | | - The parallelism degree of Spark resources is jointly determined by the number of Executors and the number of Executor CPU cores. | - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Job Manager CUs | Number of management unit CUs. | - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parallelism | Number of tasks concurrently executed by each operator in a job. | - | | | - | | .. note:: | - | | | - | | - The value cannot exceed four times the number of compute units (**CUs** - **Job Manager CUs**). | - | | - Set this parameter to a value greater than that configured in the code to avoid job submission failures. | - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Task Manager Config | Whether TaskManager resource parameters are set | - | | | - | | - If this option is selected, you need to set the following parameters: | - | | | - | | - **CU(s) per TM**: Number of resources occupied by each TaskManager. | - | | - **Slot(s) per TM**: Number of slots contained in each TaskManager. | - | | | - | | - If not selected, the system automatically uses the default values. | - | | | - | | - **CU(s) per TM**: The default value is **1**. | - | | - **Slot(s) per TM**: The default value is (Parallelism x CU(s) per TM)/(CUs - Job Manager CUs). | - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Save Job Log | Whether to save the job running logs to the OBS bucket. | - | | | - | | .. caution:: | - | | | - | | CAUTION: | - | | You are advised to select this parameter. Otherwise, no run logs will be generated after the job is executed. If the job runs abnormally later, you will be unable to obtain the run logs for troubleshooting. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | **OBS Bucket**: Select an OBS bucket to store job logs. If the selected OBS bucket is not authorized, click **Authorize**. | - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Alarm on Job Exception | Whether to notify users of any job exceptions, such as running exceptions or arrears, via SMS or email. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | **SMN Topic** | - | | | - | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Auto Restart upon Exception | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | - **Max. Retry Attempts**: maximum number of retries upon an exception. The unit is times/hour. | - | | | - | | - **Unlimited**: The number of retries is unlimited. | - | | - **Limited**: The number of retries is user-defined. | - | | | - | | - **Restore Job from Checkpoint**: Restore the job from the saved checkpoint. | - | | | - | | If you select this parameter, you also need to set **Checkpoint Path**. | - | | | - | | **Checkpoint Path**: Select a path for storing checkpoints. This path must match that configured in the application package. Each job must have a unique checkpoint path, or, you will not be able to obtain the checkpoint. | - +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+================================================================================================================================================================================================================================+ + | CUs | One CU consists of one vCPU and 4 GB of memory. The number of CUs ranges from 2 to 10000. | + | | | + | | .. note:: | + | | | + | | When **Task Manager Config** is selected, elastic resource pool queue management is optimized by automatically adjusting **CUs** to match **Actual CUs** after setting **Slot(s) per TM**. | + | | | + | | **CUs = Actual number of CUs = max[Job Manager CPU + Task Manager CPU, (Job Manager Memory + Task Manager Memory/4)]** | + | | | + | | - Job Manager CPU + Task Manager CPU = Actual TMs x CU(s) per TM + Job Manager CUs. | + | | - Job Manager Memory + Task Manager Memory = Actual TMs x Memory per TM + Job Manager Memory | + | | - If **Slot(s) per TM** is set, then: Actual TMs = Parallelism/Slot(s) per TM. | + | | - If **Slot(s) per TM** is not set, then: Actual TMs = (CUs - Job Manager CUs)/CU(s) per TM. | + | | - If **Memory per TM** and **Job Manager Memory** in the optimization parameters are not set, then: Memory per TM = CU(s) per TM x 4. Job Manager Memory = Job Manager CUs x 4. | + | | - The parallelism degree of Spark resources is jointly determined by the number of Executors and the number of Executor CPU cores. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Job Manager CUs | Number of management unit CUs. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the JobManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parallelism | Number of tasks concurrently executed by each operator in a job. | + | | | + | | .. note:: | + | | | + | | - The value cannot exceed four times the number of compute units (**CUs** - **Job Manager CUs**). | + | | - Set this parameter to a value greater than that configured in the code to avoid job submission failures. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Task Manager Config | Whether TaskManager resource parameters are set | + | | | + | | - If this option is selected, you need to set the following parameters: | + | | | + | | - **CU(s) per TM**: Number of resources occupied by each TaskManager. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that a single TaskManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | + | | | + | | - **Slot(s) per TM**: Number of slots contained in each TaskManager. | + | | | + | | - If not selected, the system automatically uses the default values. | + | | | + | | - **CU(s) per TM**: The default value is **1**. | + | | - **Slot(s) per TM**: The default value is (Parallelism x CU(s) per TM)/(CUs - Job Manager CUs). | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ .. _dli_01_0457__table155918376538: @@ -386,6 +362,8 @@ Creating a Flink Jar Job | Job Manager CPU | Number of vCPUs available for JobManager. | | | | | | The default value is **1**. The minimum value cannot be less than 0.5. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the JobManager's CPU value does not exceed 2 to avoid resource scheduling failures during job execution. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Job Manager Memory | Memory available for JobManager. | | | | @@ -394,6 +372,8 @@ Creating a Flink Jar Job | Task Manager CPU | Number of vCPUs available for TaskManager. | | | | | | The default value is **1**. The minimum value cannot be less than 0.5. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the TaskManager's CPU value does not exceed 2 to avoid resource scheduling failures during job execution. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Task Manager Memory | Memory available for TaskManager. | | | | @@ -405,42 +385,6 @@ Creating a Flink Jar Job | | | | | By default, a single TM slot is set to **1**. The minimum parallelism must not be less than 1. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Save Job Log | Whether to save the job running logs to the OBS bucket. | - | | | - | | .. caution:: | - | | | - | | CAUTION: | - | | You are advised to select this parameter. Otherwise, no run logs will be generated after the job is executed. If the job runs abnormally later, you will be unable to obtain the run logs for troubleshooting. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | **OBS Bucket**: Select an OBS bucket to store job logs. If the selected OBS bucket is not authorized, click **Authorize**. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | OBS Bucket | OBS bucket to store job logs and checkpoint information. If the OBS bucket you selected is unauthorized, click **Authorize**. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Alarm on Job Exception | Whether to notify users of any job exceptions, such as running exceptions or arrears, via SMS or email. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | **SMN Topic** | - | | | - | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Auto Restart upon Exception | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | - **Max. Retry Attempts**: maximum number of retries upon an exception. The unit is times/hour. | - | | | - | | - **Unlimited**: The number of retries is unlimited. | - | | - **Limited**: The number of retries is user-defined. | - | | | - | | - **Restore Job from Checkpoint**: Restore the job from the saved checkpoint. | - | | | - | | If you select this parameter, you also need to set **Checkpoint Path**. | - | | | - | | **Checkpoint Path**: Select a path for storing checkpoints. This path must match that configured in the application package. Each job must have a unique checkpoint path, or, you will not be able to obtain the checkpoint. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ You can set compute resource specification parameters on the **Runtime Configuration** tab of Flink jobs, and the parameter values have a higher priority than the specified values. @@ -474,6 +418,74 @@ Creating a Flink Jar Job | | | | The default value is 4 GB. The minimum size cannot be less than 2 GB (2,048 MB). The default unit is GB, which can be set to GB or MB. | +---------------------------------+------------------------------------------------+------------------------------------------------+----------------------------------------------------------------------------------------------------------------------------------------+ +#. Configure logs and status. + + .. table:: **Table 7** Configuring logs and status + + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+===================================================================================================================================================================================================================+ + | Save Job Log | Whether to save the job running logs to the OBS bucket. | + | | | + | | .. caution:: | + | | | + | | CAUTION: | + | | You are advised to select this parameter. Otherwise, no run logs will be generated after the job is executed. If the job runs abnormally later, you will be unable to obtain the run logs for troubleshooting. | + | | | + | | If this option is selected, you need to set the following parameters: | + | | | + | | **OBS Bucket**: Select an OBS bucket to store job logs. If the selected OBS bucket is not authorized, click **Authorize**. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | OBS Bucket | OBS bucket to store job logs and checkpoint information. If the OBS bucket you selected is unauthorized, click **Authorize**. | + +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + + .. table:: **Table 8** Root Log Level + + +-------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------------------------------------------------------------------------------+ + | Type | Description | Use Case | + +=======+=========================================================================================================================================================================+===================================================================================================================+ + | TRACE | The most granular level of logging, typically used during development and debugging. It captures all operational details, including variable values and function calls. | Primarily for the development phase to help developers understand code execution flow and state. | + +-------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------------------------------------------------------------------------------+ + | DEBUG | Less detailed than TRACE, intended for debugging purposes. It logs program runtime states and key variable values but omits exhaustive detail. | Mainly used during development and testing phases to assist developers in identifying issues. | + +-------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------------------------------------------------------------------------------+ + | INFO | Records important information during normal operations, useful for system operators and maintainers without impacting system functionality. | Tracks major system operations and statuses, such as startup, shutdown, and configuration changes. | + +-------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------------------------------------------------------------------------------+ + | WARN | Potential issues or anomalies that do not disrupt system operation but warrant attention from developers or operators. | Logs situations that may affect system performance or functionality, serving as a warning for possible problems. | + +-------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------------------------------------------------------------------------------+ + | ERROR | Highlights critical issues or exceptions that impair system functionality and require immediate resolution. | Captures errors and exceptions during system operation, enabling developers to quickly identify and address them. | + +-------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------+-------------------------------------------------------------------------------------------------------------------+ + +#. Configures error handling parameters. + + .. table:: **Table 9** Configuring exception handling parameters + + +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+=================================================================================================================================================================================================================================+ + | Alarm on Job Exception | Whether to notify users of any job exceptions, such as running exceptions or arrears, via SMS or email. | + | | | + | | If this option is selected, you need to set the following parameters: | + | | | + | | **SMN Topic** | + | | | + | | Select a custom SMN topic. For details about how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | + +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Auto Restart upon Exception | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | + | | | + | | If this option is selected, you need to set the following parameters: | + | | | + | | - **Max. Retry Attempts**: maximum number of retries upon an exception. The unit is times/hour. | + | | | + | | - **Unlimited**: The number of retries is unlimited. | + | | - **Limited**: The number of retries is user-defined. | + | | | + | | - **Restore Job from Checkpoint**: Restore the job from the saved checkpoint. | + | | | + | | If you select this parameter, you also need to set **Checkpoint Path**. | + | | | + | | **Checkpoint Path**: Select a path for storing checkpoints. This path must match that configured in the application package. Each job must have a unique checkpoint path, or, you will not be able to obtain the checkpoint. | + +-----------------------------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + #. Click **Save** in the upper right of the page. #. Click **Start** in the upper right corner. On the displayed **Start Flink Job** page, confirm the job specifications, and click **Start Now** to start the job. After the job is started, the system automatically switches to the **Flink Jobs** page, and the created job is displayed in the job list. You can view the job status in the **Status** column. diff --git a/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_opensource_sql_job.rst b/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_opensource_sql_job.rst index f3cadc4..7a3d264 100644 --- a/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_opensource_sql_job.rst +++ b/umn/source/submitting_a_flink_job_on_the_dli_management_console/creating_a_flink_opensource_sql_job.rst @@ -7,7 +7,7 @@ Creating a Flink OpenSource SQL Job This section describes how to create a Flink OpenSource SQL job. -DLI Flink OpenSource SQL jobs are fully compatible with the syntax of Flink provided by the community. In addition, Redis and GaussDB(DWS) data source types are added based on the community connector. For the syntax and constraints of Flink SQL DDL, DML, and functions, see `Table API & SQL `__. +DLI Flink OpenSource SQL jobs are fully compatible with the syntax of Flink provided by the community. In addition, Redis and DWS data source types are added based on the community connector. For details about the syntax and constraints of Flink SQL DDL, DML, and functions, see `Table API & SQL `__. Prerequisites ------------- @@ -17,7 +17,7 @@ Prerequisites - For details about the external data sources that can be accessed by Flink jobs, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. - - For how to create a datasource connection, see :ref:`Configuring the Network Connection Between DLI and Data Sources (Enhanced Datasource Connection) `. + - For details about how to create a datasource connection, see :ref:`Configuring the Network Connection Between DLI and Data Sources (Enhanced Datasource Connection) `. On the **Resources** > **Queue Management** page, locate the queue you have created, click **More** in the **Operation** column, and select **Test Address Connectivity** to check if the network connection between the queue and the data source is normal. For details, see :ref:`Testing the Network Connectivity Between a Queue and a Data Source `. @@ -25,7 +25,7 @@ Prerequisites Creating a Flink OpenSource SQL Job ----------------------------------- -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. In the upper right corner of the **Flink Jobs** page, click **Create Job**. @@ -111,22 +111,22 @@ Creating a Flink OpenSource SQL Job | | | | | There are the following ways to manage UDF JAR files: | | | | - | | - Upload packages to OBS: Upload Jar packages to an OBS bucket in advance and select the corresponding OBS path. | + | | - Upload packages to OBS: Upload JAR files to an OBS bucket in advance and select the corresponding OBS path. | | | - Upload packages to DLI: Upload JAR files to an OBS bucket in advance and create a package on the **Data Management** > **Package Management** page of the DLI management console. For details, see :ref:`Creating a DLI Package `. | | | | | | For Flink 1.15 or later, only OBS packages can be selected when creating jobs, and DLI packages are not supported. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Agency | If you choose Flink 1.15 or later to execute your job, you can create a custom agency to allow DLI to access other services. | | | | - | | For how to create a custom agency, see :ref:`Creating a Custom DLI Agency `. | + | | For details about how to create a custom agency, see :ref:`Creating a Custom DLI Agency `. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Resource Configuration Version | DLI offers various resource configuration templates based on different Flink engine versions. | | | | | | Compared with the v1 template, the v2 template does not support the setting of the number of CUs. The v2 template supports the setting of **Job Manager Memory** and **Task Manager Memory**. | | | | - | | **v1**: applicable to Flink 1.12, 1.13, and 1.15. | + | | **V1**: applicable to Flink 1.12 and 1.15. | | | | - | | **v2**: applicable to Flink 1.13, 1.15, and 1.17. | + | | **V2**: applicable to Flink 1.15 and 1.17. | | | | | | You are advised to use the parameter settings of v2. | | | | @@ -134,7 +134,7 @@ Creating a Flink OpenSource SQL Job | | | | | For details about the parameters of v2, see :ref:`Table 4 `. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | CUs | Sum of the number of compute units and JobManager CUs of DLI. One CU equals 1 vCPU and 4 GB of memory. | + | CUs | Sum of the number of compute units and JobManager CUs of DLI. One CU equals one vCPU and 4 GB of memory. | | | | | | The value is the number of CUs required for job running and cannot exceed the number of CUs in the bound queue. | | | | @@ -152,6 +152,8 @@ Creating a Flink OpenSource SQL Job | | - The parallelism degree of Spark resources is jointly determined by the number of Executors and the number of Executor CPU cores. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Job Manager CUs | Number of JobManager CUs. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the JobManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Parallelism | Number of tasks concurrently executed by each operator in a job. | | | | @@ -164,6 +166,9 @@ Creating a Flink OpenSource SQL Job | | - If selected, you need to set the following parameters: | | | | | | - **CU(s) per TM**: Number of resources occupied by each TaskManager. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that a single TaskManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | + | | | | | - **Slot(s) per TM**: Number of slots contained in each TaskManager. | | | | | | - If not selected, the system automatically uses the default values. | @@ -194,7 +199,7 @@ Creating a Flink OpenSource SQL Job | | | | | **SMN Topic** | | | | - | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | + | | Select a custom SMN topic. For details about how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Enable Checkpointing | Whether to enable job snapshots. If this function is enabled, jobs can be restored based on the checkpoints. | | | | @@ -240,108 +245,113 @@ Creating a Flink OpenSource SQL Job .. table:: **Table 3** Resource specification parameters of v1 - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parameter | Description | - +===================================+===================================================================================================================================================================================================================+ - | CUs | Sum of the number of compute units and JobManager CUs of DLI. One CU equals 1 vCPU and 4 GB of memory. | - | | | - | | The value is the number of CUs required for job running and cannot exceed the number of CUs in the bound queue. | - | | | - | | .. note:: | - | | | - | | When **Task Manager Config** is selected, elastic resource pool queue management is optimized by automatically adjusting **CUs** to match **Actual CUs** after setting **Slot(s) per TM**. | - | | | - | | **CUs = Actual number of CUs = max[Job Manager CPU + Task Manager CPU, (Job Manager Memory + Task Manager Memory/4)]** | - | | | - | | - Job Manager CPU + Task Manager CPU = Actual TMs x CU(s) per TM + Job Manager CUs. | - | | - Job Manager Memory + Task Manager Memory = Actual TMs x Memory per TM + Job Manager Memory | - | | - If **Slot(s) per TM** is set, then: Actual TMs = Parallelism/Slot(s) per TM. | - | | - If **Slot(s) per TM** is not set, then: Actual TMs = (CUs - Job Manager CUs)/CU(s) per TM. | - | | - If **Memory per TM** and **Job Manager Memory** in the optimization parameters are not set, then: Memory per TM = CU(s) per TM x 4. Job Manager Memory = Job Manager CUs x 4. | - | | - The parallelism degree of Spark resources is jointly determined by the number of Executors and the number of Executor CPU cores. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Job Manager CUs | Number of JobManager CUs. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parallelism | Number of tasks concurrently executed by each operator in a job. | - | | | - | | .. note:: | - | | | - | | This value cannot be greater than four times the compute units (**CUs** - **Job Manager CUs**). | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Task Manager Config | Whether TaskManager resource parameters are set | - | | | - | | - If selected, you need to set the following parameters: | - | | | - | | - **CU(s) per TM**: Number of resources occupied by each TaskManager. | - | | - **Slot(s) per TM**: Number of slots contained in each TaskManager. | - | | | - | | - If not selected, the system automatically uses the default values. | - | | | - | | - **CU(s) per TM**: The default value is **1**. | - | | - **Slot(s) per TM**: The default value is (Parallelism x CU(s) per TM)/(CUs - Job Manager CUs). | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | OBS Bucket | OBS bucket to store job logs and checkpoint information. If the OBS bucket you selected is unauthorized, click **Authorize**. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Save Job Log | Whether job running logs are saved to OBS. The logs are saved in the following path: *Bucket name*\ **/jobs/logs/**\ *Directory starting with the job ID*. | - | | | - | | .. caution:: | - | | | - | | CAUTION: | - | | You are advised to select this parameter. Otherwise, no run logs will be generated after the job is executed. If the job runs abnormally later, you will be unable to obtain the run logs for troubleshooting. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | **OBS Bucket**: Select an OBS bucket to store job logs. If the OBS bucket you selected is unauthorized, click **Authorize**. | - | | | - | | .. note:: | - | | | - | | If **Enable Checkpointing** and **Save Job Log** are both selected, you only need to authorize OBS once. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Alarm on Job Exception | Whether to notify users of any job exceptions, such as running exceptions or arrears, via SMS or email. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | **SMN Topic** | - | | | - | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Enable Checkpointing | Whether to enable job snapshots. If this function is enabled, jobs can be restored based on the checkpoints. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | - **Checkpoint Interval**: interval for creating checkpoints, in seconds. The value ranges from 1 to 999999, and the default value is **30**. | - | | - **Checkpoint Mode** can be set to either of the following values: | - | | | - | | - **At least once**: Events are processed at least once. | - | | - **Exactly once**: Events are processed only once. | - | | | - | | - **OBS Bucket**: Select an OBS bucket to store your checkpoints. If the OBS bucket you selected is unauthorized, click **Authorize**. | - | | | - | | The checkpoint path is *Bucket name*\ **/jobs/checkpoint/**\ *Directory starting with the job ID*. | - | | | - | | .. note:: | - | | | - | | If **Enable Checkpointing** and **Save Job Log** are both selected, you only need to authorize OBS once. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Auto Restart upon Exception | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | - | | | - | | If this option is selected, you need to set the following parameters: | - | | | - | | - **Max. Retry Attempts**: maximum number of retries upon an exception. The unit is times/hour. | - | | | - | | - **Unlimited**: The number of retries is unlimited. | - | | - **Limited**: The number of retries is user-defined. | - | | | - | | - **Restore Job from Checkpoint**: This parameter is available only when **Enable Checkpointing** is selected. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Idle State Retention Time | Clears intermediate states of operators such as **GroupBy**, **RegularJoin**, **Rank**, and **Depulicate** that have not been updated after the maximum retention time. The default value is 1 hour. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Dirty Data Policy | Policy for processing dirty data. The following policies are supported: **Ignore**, **Trigger a job exception**, and **Save**. | - | | | - | | If you set this field to **Save**, the **Dirty Data Dump Address** must be set. Click the address box to select the OBS path for storing dirty data. | - | | | - | | This parameter is available only when a DIS data source is used. | - +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+================================================================================================================================================================================================================================+ + | CUs | Sum of the number of compute units and JobManager CUs of DLI. One CU equals one vCPU and 4 GB of memory. | + | | | + | | The value is the number of CUs required for job running and cannot exceed the number of CUs in the bound queue. | + | | | + | | .. note:: | + | | | + | | When **Task Manager Config** is selected, elastic resource pool queue management is optimized by automatically adjusting **CUs** to match **Actual CUs** after setting **Slot(s) per TM**. | + | | | + | | **CUs = Actual number of CUs = max[Job Manager CPU + Task Manager CPU, (Job Manager Memory + Task Manager Memory/4)]** | + | | | + | | - Job Manager CPU + Task Manager CPU = Actual TMs x CU(s) per TM + Job Manager CUs. | + | | - Job Manager Memory + Task Manager Memory = Actual TMs x Memory per TM + Job Manager Memory | + | | - If **Slot(s) per TM** is set, then: Actual TMs = Parallelism/Slot(s) per TM. | + | | - If **Slot(s) per TM** is not set, then: Actual TMs = (CUs - Job Manager CUs)/CU(s) per TM. | + | | - If **Memory per TM** and **Job Manager Memory** in the optimization parameters are not set, then: Memory per TM = CU(s) per TM x 4. Job Manager Memory = Job Manager CUs x 4. | + | | - The parallelism degree of Spark resources is jointly determined by the number of Executors and the number of Executor CPU cores. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Job Manager CUs | Number of JobManager CUs. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the JobManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parallelism | Number of tasks concurrently executed by each operator in a job. | + | | | + | | .. note:: | + | | | + | | This value cannot be greater than four times the compute units (**CUs** - **Job Manager CUs**). | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Task Manager Config | Whether TaskManager resource parameters are set | + | | | + | | - If selected, you need to set the following parameters: | + | | | + | | - **CU(s) per TM**: Number of resources occupied by each TaskManager. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that a single TaskManager's CU count does not exceed 2 to avoid resource scheduling failures during job execution. | + | | | + | | - **Slot(s) per TM**: Number of slots contained in each TaskManager. | + | | | + | | - If not selected, the system automatically uses the default values. | + | | | + | | - **CU(s) per TM**: The default value is **1**. | + | | - **Slot(s) per TM**: The default value is (Parallelism x CU(s) per TM)/(CUs - Job Manager CUs). | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | OBS Bucket | OBS bucket to store job logs and checkpoint information. If the OBS bucket you selected is unauthorized, click **Authorize**. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Save Job Log | Whether job running logs are saved to OBS. The logs are saved in the following path: *Bucket name*\ **/jobs/logs/**\ *Directory starting with the job ID*. | + | | | + | | .. caution:: | + | | | + | | CAUTION: | + | | You are advised to select this parameter. Otherwise, no run logs will be generated after the job is executed. If the job runs abnormally later, you will be unable to obtain the run logs for troubleshooting. | + | | | + | | If this option is selected, you need to set the following parameters: | + | | | + | | **OBS Bucket**: Select an OBS bucket to store job logs. If the OBS bucket you selected is unauthorized, click **Authorize**. | + | | | + | | .. note:: | + | | | + | | If **Enable Checkpointing** and **Save Job Log** are both selected, you only need to authorize OBS once. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Alarm on Job Exception | Whether to notify users of any job exceptions, such as running exceptions or arrears, via SMS or email. | + | | | + | | If this option is selected, you need to set the following parameters: | + | | | + | | **SMN Topic** | + | | | + | | Select a custom SMN topic. For details about how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Enable Checkpointing | Whether to enable job snapshots. If this function is enabled, jobs can be restored based on the checkpoints. | + | | | + | | If this option is selected, you need to set the following parameters: | + | | | + | | - **Checkpoint Interval**: interval for creating checkpoints, in seconds. The value ranges from 1 to 999999, and the default value is **30**. | + | | - **Checkpoint Mode** can be set to either of the following values: | + | | | + | | - **At least once**: Events are processed at least once. | + | | - **Exactly once**: Events are processed only once. | + | | | + | | - **OBS Bucket**: Select an OBS bucket to store your checkpoints. If the OBS bucket you selected is unauthorized, click **Authorize**. | + | | | + | | The checkpoint path is *Bucket name*\ **/jobs/checkpoint/**\ *Directory starting with the job ID*. | + | | | + | | .. note:: | + | | | + | | If **Enable Checkpointing** and **Save Job Log** are both selected, you only need to authorize OBS once. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Auto Restart upon Exception | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | + | | | + | | If this option is selected, you need to set the following parameters: | + | | | + | | - **Max. Retry Attempts**: maximum number of retries upon an exception. The unit is times/hour. | + | | | + | | - **Unlimited**: The number of retries is unlimited. | + | | - **Limited**: The number of retries is user-defined. | + | | | + | | - **Restore Job from Checkpoint**: This parameter is available only when **Enable Checkpointing** is selected. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Idle State Retention Time | Clears intermediate states of operators such as **GroupBy**, **RegularJoin**, **Rank**, and **Depulicate** that have not been updated after the maximum retention time. The default value is 1 hour. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Dirty Data Policy | Policy for processing dirty data. The following policies are supported: **Ignore**, **Trigger a job exception**, and **Save**. | + | | | + | | If you set this field to **Save**, the **Dirty Data Dump Address** must be set. Click the address box to select the OBS path for storing dirty data. | + | | | + | | This parameter is available only when a DIS data source is used. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ .. _dli_01_0498__table174111633103818: @@ -360,6 +370,8 @@ Creating a Flink OpenSource SQL Job | Job Manager CPU | Number of vCPUs available for JobManager. | | | | | | The default value is **1**. The minimum value cannot be less than 0.5. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the JobManager's CPU value does not exceed 2 to avoid resource scheduling failures during job execution. | +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Job Manager Memory | Memory available for JobManager. | | | | @@ -368,6 +380,8 @@ Creating a Flink OpenSource SQL Job | Task Manager CPU | Number of vCPUs available for TaskManager. | | | | | | The default value is **1**. The minimum value cannot be less than 0.5. | + | | | + | | If the current job is running on a basic edition elastic resource pool (16-64 CUs), it is recommended that the TaskManager's CPU value does not exceed 2 to avoid resource scheduling failures during job execution. | +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Task Manager Memory | Memory available for TaskManager. | | | | @@ -402,7 +416,7 @@ Creating a Flink OpenSource SQL Job | | | | | **SMN Topic** | | | | - | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | + | | Select a custom SMN topic. For details about how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | +-----------------------------------+-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Enable Checkpointing | Whether to enable job snapshots. If this function is enabled, jobs can be restored based on the checkpoints. | | | | diff --git a/umn/source/submitting_a_flink_job_on_the_dli_management_console/flink_job_overview.rst b/umn/source/submitting_a_flink_job_on_the_dli_management_console/flink_job_overview.rst index 7859f82..6b84131 100644 --- a/umn/source/submitting_a_flink_job_on_the_dli_management_console/flink_job_overview.rst +++ b/umn/source/submitting_a_flink_job_on_the_dli_management_console/flink_job_overview.rst @@ -10,7 +10,7 @@ DLI supports two types of Flink jobs: - **Flink OpenSource SQL job:** - It is fully compatible with Flink of the community edition, ensuring that jobs can run smoothly on these Flink versions. - - DLI Flink has expanded the support for connectors based on Flink of the community edition, supporting Redis and GaussDB(DWS) as new data source types. With this expansion, you can now utilize a wider range of data source types, providing greater flexibility and convenience when working with datasets. + - DLI Flink has expanded the support for connectors based on Flink of the community edition, supporting Redis and DWS as new data source types. With this expansion, you can now utilize a wider range of data source types, providing greater flexibility and convenience when working with datasets. - Flink OpenSource SQL jobs are ideal for scenarios where stream processing logic can be defined and executed through SQL statements. This simplifies stream processing, allowing developers to focus more on implementing service logic. For how to create a Flink OpenSource SQL job, see :ref:`Creating a Flink OpenSource SQL Job `. diff --git a/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_job_templates.rst b/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_job_templates.rst index 818be6e..0948e73 100644 --- a/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_job_templates.rst +++ b/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_job_templates.rst @@ -95,7 +95,7 @@ You can create a template using any of the following methods: - Creating a template on the **Template Management** page - #. In the left navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. + #. In the navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. #. Click **Create Template** in the upper right corner of the page. The **Create Template** dialog box is displayed. @@ -148,43 +148,43 @@ You can create a template using any of the following methods: .. table:: **Table 5** Template parameters - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Parameter | Description | - +===================================+==================================================================================================================================================================+ - | Name | You can modify the template name. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Description | You can modify the template description. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Saving Mode | - **Save Here**: Save the modification to the current template. | - | | - **Save as New**: Save the modification as a new template. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | SQL statement editing area | In the area, you can enter detailed SQL statements to implement business logic. For how to compile SQL statements, see *Data Lake Insight SQL Syntax Reference*. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Save | Save the modifications. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Create Job | Use the current template to create a job. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Format | Format SQL statements. After SQL statements are formatted, you need to compile SQL statements again. | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Theme Settings | Change the font size, word wrap, and page style (black or white background). | - +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - - #. In the SQL statement editing area, enter SQL statements to implement service logic. For how to compile SQL statements, see *Data Lake Insight SQL Syntax Reference*. + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Parameter | Description | + +===================================+================================================================================================================================================================================+ + | Name | You can modify the template name. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Description | You can modify the template description. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Saving Mode | - **Save Here**: Save the modification to the current template. | + | | - **Save as New**: Save the modification as a new template. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | SQL statement editing area | In the area, you can enter detailed SQL statements to implement business logic. For details about how to compile SQL statements, see *Data Lake Insight SQL Syntax Reference*. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Save | Save the modifications. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Create Job | Use the current template to create a job. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Format | Format SQL statements. After SQL statements are formatted, you need to compile SQL statements again. | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + | Theme Settings | Change the font size, word wrap, and page style (black or white background). | + +-----------------------------------+--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ + + #. In the SQL statement editing area, enter SQL statements to implement service logic. For details about how to compile SQL statements, see *Data Lake Insight SQL Syntax Reference*. #. After the SQL statement is edited, click **Save** in the upper right corner to complete the template creation. - #. (Optional) If you do not need to modify the template, click **Create Job** in the upper right corner to create a job based on the current template. For how to create a job, see :ref:`Creating a Flink Jar Job `. + #. (Optional) If you do not need to modify the template, click **Create Job** in the upper right corner to create a job based on the current template. For details about how to create a job, see :ref:`Creating a Flink Jar Job `. - Creating a template based on an existing job template - #. In the left navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. Click the **Custom Templates** tab. + #. In the navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. Click the **Custom Templates** tab. #. In the row where the desired template is located in the custom template list, click **Edit** under **Operation** to enter the **Edit** page. #. After the modification is complete, set **Saving Mode** to **Save as New**. #. Click **Save** in the upper right corner to save the template as a new one. - Creating a template using a created job - #. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. + #. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. Click **Create Job** in the upper right corner. The **Create Job** page is displayed. #. Specify parameters as required. #. Click **OK** to enter the editing page. @@ -193,7 +193,7 @@ You can create a template using any of the following methods: - Creating a template based on the existing job - #. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. + #. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. In the job list, locate the row where the job that you want to set as a template resides, and click **Edit** in the **Operation** column. #. After the SQL statement is compiled, click **Set as Template**. #. In the **Set as Template** dialog box that is displayed, specify **Name** and **Description** and click **OK**. @@ -205,8 +205,8 @@ Creating a Job Based on a Template You can create jobs based on sample templates or custom templates. -#. In the left navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. -#. In the sample template list, click **Create Job** in the **Operation** column of the target template. For how to create a job, see :ref:`Creating a Flink OpenSource SQL Job ` and :ref:`Creating a Flink Jar Job `. +#. In the navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. +#. In the sample template list, click **Create Job** in the **Operation** column of the target template. For details about how to create a job, see :ref:`Creating a Flink OpenSource SQL Job ` and :ref:`Creating a Flink Jar Job `. .. _dli_01_0464__section735234815411: @@ -215,7 +215,7 @@ Modifying a Template After creating a custom template, you can modify it as required. The sample template cannot be modified, but you can view the template details. -#. In the left navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. Click the **Custom Templates** tab. +#. In the navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. Click the **Custom Templates** tab. #. In the row where the template you want to modify is located in the custom template list, click **Edit** in the **Operation** column to enter the **Edit** page. #. In the SQL statement editing area, modify the SQL statements as required. #. Set **Saving Mode** to **Save Here**. @@ -228,7 +228,7 @@ Deleting a Template You can delete a custom template as required. The sample templates cannot be deleted. Deleted templates cannot be restored. Exercise caution when performing this operation. -#. In the left navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. Click the **Custom Templates** tab. +#. In the navigation pane of the DLI management console, choose **Job Templates** > **Flink Templates**. Click the **Custom Templates** tab. #. In the custom template list, select the templates you want to delete and click **Delete** in the upper left of the custom template list. diff --git a/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_jobs/common_operations_of_flink_jobs.rst b/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_jobs/common_operations_of_flink_jobs.rst index 3de5f94..fed588b 100644 --- a/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_jobs/common_operations_of_flink_jobs.rst +++ b/umn/source/submitting_a_flink_job_on_the_dli_management_console/managing_flink_jobs/common_operations_of_flink_jobs.rst @@ -12,7 +12,7 @@ Editing a Job You can edit a created job, for example, by modifying the SQL statement, job name, job description, or job configurations. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. In the row where the job you want to edit locates, click **Edit** in the **Operation** column to switch to the editing page. @@ -25,7 +25,7 @@ Starting a Job You can start a saved or stopped job. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. Use either of the following methods to start jobs: @@ -50,7 +50,7 @@ Stopping a Job You can stop a job in the **Running** or **Submitting** state. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. Stop a job using either of the following methods: @@ -83,7 +83,7 @@ Deleting a Job If you do not need to use a job, perform the following operations to delete it. A deleted job cannot be restored. Therefore, exercise caution when deleting a job. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. 2. Perform either of the following methods to delete jobs: @@ -110,7 +110,7 @@ This mode is applicable to the scenario where a large number of jobs need to be When switching to another project or user, you need to grant permissions to the new project or user. For details, see :ref:`Configuring Flink Job Permissions `. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. 2. Click **Export Job** in the upper right corner. The **Export Job** dialog box is displayed. @@ -138,7 +138,7 @@ For details, see :ref:`Creating a Flink OpenSource SQL Job ` and :r - When switching to another project or user, you need to grant permissions to the new project or user. For details, see :ref:`Configuring Flink Job Permissions `. - Only jobs whose data format is the same as that of Flink jobs exported from DLI can be imported. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. 2. Click **Import Job** in the upper right corner. The **Import Job** dialog box is displayed. 3. Select the complete OBS path of the job configuration file to be imported. Click **Next**. @@ -154,7 +154,7 @@ Modifying the Name and Description of a Flink Job You can change the job name and description as required. -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. The **Flink Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Flink Jobs**. #. In the **Operation** column of the job whose name and description need to be modified, choose **More > Modify Name and Description**. The **Modify Name and Description** dialog box is displayed. Change the name or modify the description of a job. #. Click **OK**. @@ -206,7 +206,7 @@ You can configure job exception alarms and restart options by selecting **Runtim | | | | | **SMN Topic** | | | | - | | Select a custom SMN topic. For how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | + | | Select a custom SMN topic. For details about how to create a custom SMN topic, see "Creating a Topic" in the *Simple Message Notification User Guide*. | +-------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Auto Restart upon Exception | Whether automatic restart is enabled. If enabled, jobs will be automatically restarted and restored when exceptions occur. | | | | diff --git a/umn/source/submitting_a_spark_job_on_the_dli_management_console/creating_a_spark_job.rst b/umn/source/submitting_a_spark_job_on_the_dli_management_console/creating_a_spark_job.rst index 5af18fd..fa90055 100644 --- a/umn/source/submitting_a_spark_job_on_the_dli_management_console/creating_a_spark_job.rst +++ b/umn/source/submitting_a_spark_job_on_the_dli_management_console/creating_a_spark_job.rst @@ -17,22 +17,22 @@ Prerequisites ------------- - You have uploaded the dependencies to the corresponding OBS bucket on the **Data Management > Package Management** page. -- Before creating a Spark job to access other external data sources, such as OpenTSDB, HBase, Kafka, GaussDB(DWS), RDS, CSS, CloudTable, DCS Redis, and DDS, you need to create a datasource connection to enable the network between the job running queue and external data sources. +- Before creating a Spark job to access other external data sources, such as OpenTSDB, HBase, Kafka, DWS, RDS, CSS, CloudTable, DCS Redis, and DDS, you need to create a datasource connection to enable the network between the job running queue and external data sources. - For details about the external data sources that can be accessed by Spark jobs, see :ref:`Common Development Methods for DLI Cross-Source Analysis `. - - For how to create a datasource connection, see :ref:`Configuring the Network Connection Between DLI and Data Sources (Enhanced Datasource Connection) `. + - For details about how to create a datasource connection, see :ref:`Configuring the Network Connection Between DLI and Data Sources (Enhanced Datasource Connection) `. On the **Resources** > **Queue Management** page, locate the queue you have created, click **More** in the **Operation** column, and select **Test Address Connectivity** to check if the network connection between the queue and the data source is normal. For details, see :ref:`Testing the Network Connectivity Between a Queue and a Data Source `. Procedure --------- -#. In the left navigation pane of the DLI management console, choose **Job Management** > **Spark Jobs**. The **Spark Jobs** page is displayed. +#. In the navigation pane of the DLI management console, choose **Job Management** > **Spark Jobs**. Click **Create Job** in the upper right corner. In the job editing window, you can set parameters in **Fill Form** mode or **Write API** mode. - The following uses the **Fill Form** as an example. In **Write API** mode, refer to the *Data Lake Insight API Reference* for parameter settings. + The following uses the **Fill Form** as an example. In **Write API** mode, see *Data Lake Insight API Reference* for parameter settings. 2. Select a queue. @@ -121,7 +121,7 @@ Procedure | | | | | spark.executor.extraClassPath=/usr/share/extension/dli/spark-jar/datasource/css/\* | +-----------------------------------+-----------------------------------------------------------------------------------------+ - | GaussDB(DWS) | spark.driver.extraClassPath=/usr/share/extension/dli/spark-jar/datasource/dws/\* | + | DWS | spark.driver.extraClassPath=/usr/share/extension/dli/spark-jar/datasource/dws/\* | | | | | | spark.executor.extraClassPath=/usr/share/extension/dli/spark-jar/datasource/dws/\* | +-----------------------------------+-----------------------------------------------------------------------------------------+ @@ -195,7 +195,7 @@ Procedure +---------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Other Dependencies (--files) | Other files on which the Spark job depends. You can enter the name of the dependency file or the corresponding OBS path of the dependency file. The format is as follows: **obs://Bucket name/Folder name/File name**. | +---------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ - | Group Name | If you select a group when creating a package, you can select all the packages and files in the group. For how to create a package, see :ref:`Creating a DLI Package `. | + | Group Name | If you select a group when creating a package, you can select all the packages and files in the group. For details about how to create a package, see :ref:`Creating a DLI Package `. | | | | | | Spark 3.3.\ *x* or later does not support group information configuration. | +---------------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ diff --git a/umn/source/submitting_a_spark_job_on_the_dli_management_console/managing_spark_jobs.rst b/umn/source/submitting_a_spark_job_on_the_dli_management_console/managing_spark_jobs.rst index 3197de1..2233cf0 100644 --- a/umn/source/submitting_a_spark_job_on_the_dli_management_console/managing_spark_jobs.rst +++ b/umn/source/submitting_a_spark_job_on_the_dli_management_console/managing_spark_jobs.rst @@ -10,7 +10,7 @@ Managing Spark Jobs Viewing Basic Information ------------------------- -On the **Overview** page, click **Spark Jobs** to go to the SQL job management page. Alternatively, you can click **Job Management** > **Spark Jobs**. The page displays all Spark jobs. If there are a large number of jobs, they will be displayed on multiple pages. DLI allows you to view jobs in all statuses. +To access the Spark job management page, click **Spark Jobs** on the **Overview** page or choose **Job Management** > **Spark Jobs** in the navigation pane on the left. The page displays all Spark jobs. If there are a large number of jobs, they will be displayed on multiple pages. DLI allows you to view jobs in all statuses. .. table:: **Table 1** Job management parameters diff --git a/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_a_sql_inspection_rule.rst b/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_a_sql_inspection_rule.rst index 0b4e30b..278aa46 100644 --- a/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_a_sql_inspection_rule.rst +++ b/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_a_sql_inspection_rule.rst @@ -109,7 +109,7 @@ This part describes the system inspection rules supported by DLI. For details, s +==============+=============================+==============================================================================================================================+=========+===================+=================+===========================+=====================+========================================================================+==========================+ | dynamic_0001 | Scan files number | Maximum number of files to be scanned | Dynamic | Spark | Info | Value range: 1-2000000 | Yes | N/A | Spark 3.3.1 | | | | | | | | | | | | - | | | | | HetuEngine | Block | Default value: **200000** | | | | + | | | | | | Block | Default value: **200000** | | | | +--------------+-----------------------------+------------------------------------------------------------------------------------------------------------------------------+---------+-------------------+-----------------+---------------------------+---------------------+------------------------------------------------------------------------+--------------------------+ | dynamic_0002 | Scan partitions number | Maximum number of partitions involved in the operations (select, delete, update, and alter) that can be performed on a table | Dynamic | Spark | Info | Value range: 1-500000 | Yes | select \* from Partitioned table | Spark 3.3.1 | | | | | | | | | | | | diff --git a/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_managing_sql_job_templates/developing_and_submitting_a_sql_job_using_a_sql_job_template.rst b/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_managing_sql_job_templates/developing_and_submitting_a_sql_job_using_a_sql_job_template.rst index 3ccd5c6..52e6d4e 100644 --- a/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_managing_sql_job_templates/developing_and_submitting_a_sql_job_using_a_sql_job_template.rst +++ b/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_managing_sql_job_templates/developing_and_submitting_a_sql_job_using_a_sql_job_template.rst @@ -22,4 +22,4 @@ Procedure This example uses the **default** queue and database preset in the system as an example. You can also run the command in a self-created queue and database. -For details, see "Creating a Queue" in the *Data Lake Insight User Guide*. For how to create a database, see "Data Management" > "Databases and Tables" > "Creating a Database" in the *Data Lake Insight User Guide*. +For details, see "Creating a Queue" in the *Data Lake Insight User Guide*. For details about how to create a database, see "Data Management" > "Databases and Tables" > "Creating a Database" in the *Data Lake Insight User Guide*. diff --git a/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_submitting_a_sql_job.rst b/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_submitting_a_sql_job.rst index 5e8a0b9..7cca7f5 100644 --- a/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_submitting_a_sql_job.rst +++ b/umn/source/submitting_a_sql_job_on_the_dli_management_console/creating_and_submitting_a_sql_job.rst @@ -27,7 +27,7 @@ Notes When executing a job, the system accesses the job bucket using the identity credentials of the user who submitted the job. If permissions are insufficient, it results in failure to save or retrieve the job results. - For details, refer to :ref:`How Do I Check if Job Result Saving to a DLI Job Bucket Is Enabled for a SQL Queue? ` + For details, see :ref:`How Do I Check if Job Result Saving to a DLI Job Bucket Is Enabled for a SQL Queue? ` - On the OBS management console, you can configure lifecycle rules for a bucket to automatically delete objects within it or change object storage classes on a regular basis. @@ -38,11 +38,11 @@ Creating and Submitting a SQL Job Using the SQL Editor On the **SQL Editor** page, the system prompts you to create a temporary DLI data bucket to store temporary data generated by DLI. For details about how to configure a job bucket, see :ref:`Configuring a DLI Job Bucket `. -#. Above the SQL job editing window, set the parameters required for running a SQL job, such as the queue and database. For how to set the parameters, refer to :ref:`Table 1 `. +#. Above the SQL job editing window, set the parameters required for running a SQL job, such as the queue and database. Configure the parameters based on :ref:`Table 1 `. .. _dli_01_0320__table151023712614: - .. table:: **Table 1** Setting SQL job parameters + .. table:: **Table 1** Configuring SQL job parameters +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ | Button & Drop-Down List | Description | @@ -64,7 +64,7 @@ Creating and Submitting a SQL Job Using the SQL Editor | | | | | If no database is available, the **default** database is displayed. | | | | - | | For how to create a database, see :ref:`Creating a Data Catalog, Database, and Table on the DLI Console `. | + | | For details about how to create a database, see :ref:`Creating a Data Catalog, Database, and Table on the DLI Console `. | | | | | | If you have specified a database where tables are located in SQL statements, the database you choose here does not apply. | +-----------------------------------+------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+ diff --git a/umn/source/submitting_a_sql_job_on_the_dli_management_console/managing_sql_jobs.rst b/umn/source/submitting_a_sql_job_on_the_dli_management_console/managing_sql_jobs.rst index 12398b6..216ab68 100644 --- a/umn/source/submitting_a_sql_job_on_the_dli_management_console/managing_sql_jobs.rst +++ b/umn/source/submitting_a_sql_job_on_the_dli_management_console/managing_sql_jobs.rst @@ -84,13 +84,13 @@ The **SQL Jobs** page displays all SQL jobs, which may span multiple pages if th Viewing Job Details ------------------- -On the **SQL Jobs** page, you can click |image2| in front of a job record to view details about the job. +On the **SQL Jobs** page, click |image2| next to a job to view its details. -Job details vary with job types. The job details vary depending on the job types, status, and configuration options. The following describes how to load data, create a table, and select a job. For details about other job types, see the information on the management console. +The details displayed vary depending on the job type, status, and configuration options. The console will show the precise information based on these factors. Below are examples for three common job types: load data, create table, and select jobs. For other job types, refer to the console for supported details. -- **Load data** (job type: IMPORT) include the following information: queue, job ID, username, type, status, execution statement, running duration, creation time, end time, parameter settings, label, number of results, scanned data, number of scanned data, number of error records, storage path, data format, database, table, table header, separator, reference character, escape character, date format, timestamp format, total CPU used, and output bytes. -- **Create table** (job type: DDL) include the following information: queue, job ID, username, type, status, execution statement, running duration, creation time, end time, parameter settings, tags, number of results, scanned data, and database. -- **Select** (job type: QUERY) include the following information: queue, job ID, username, type, status, execution statement, running duration, creation time, end time, parameter setting, label, number of results (results of successful executions can be exported), and scanned data, username, result status (results of successful tasks can be viewed. Failure causes of failed tasks are displayed), database, total CPU used, and output bytes. +- For **load data** jobs (job type: **IMPORT**), the following details are included: queue, job ID, username, type, status, execution statement, runtime duration, creation time, end time, parameter settings, tags, result count, scanned data volume, scan record count, error record count, storage path, data format, database, table, table headers, delimiter, quote character, escape character, date format, timestamp format, cumulative CPU usage, and output bytes. +- For **create table** jobs (job type: **DDL**), the details include: queue, job ID, username, type, status, execution statement, runtime duration, creation time, end time, parameter settings, tags, result count, scanned data volume, and database. +- For **select** jobs (job type: **QUERY**), the details consist of: queue, job ID, username, type, status, execution statement, runtime duration, creation time, end time, parameter settings, tags, result count (exportable if successful), scanned data volume, executing user, result status (view results if successful; failure reason if failed), database, cumulative CPU usage, and output bytes. .. note:: diff --git a/umn/source/submitting_a_sql_job_on_the_dli_management_console/viewing_a_sql_execution_plan.rst b/umn/source/submitting_a_sql_job_on_the_dli_management_console/viewing_a_sql_execution_plan.rst index 24ffbad..fc186e8 100644 --- a/umn/source/submitting_a_sql_job_on_the_dli_management_console/viewing_a_sql_execution_plan.rst +++ b/umn/source/submitting_a_sql_job_on_the_dli_management_console/viewing_a_sql_execution_plan.rst @@ -12,7 +12,7 @@ This section describes how to view a SQL execution plan on the DLI management co Notes and Constraints --------------------- -- You can only view SQL execution plans for Spark 3.3.\ *x* or later queues and HetuEngine queues. +- You can only view SQL execution plans for Spark 3.3.\ *x* or later queues. - You can only view a SQL execution plan after a SQL job is executed. - You can only view the SQL execution plan for SQL jobs that have reached the **Finished** state. - Make sure you have authorized DLI to use OBS buckets for saving the SQL execution plans of user jobs.