DP-203: Data Engineering on Microsoft Azure 2025

via Udemy

Go to Course: https://www.udemy.com/course/dp-203-data-engineering-on-microsoft-azure-2025/

Overview

Launch Your Career in Data Engineering. Master designing and implementing data solutions that use Microsoft Azure data servicesThis Professional Certificate is intended for data engineers and developers who want to demonstrate their expertise in designing and implementing data solutions that use Microsoft Azure data services anyone interested in preparing for the Exam DP-203: Data Engineering on Microsoft Azure. This Professional Certificate will help you develop expertise in designing and implementing data solutions that use Microsoft Azure data services. You will learn how to integrate, transform, and consolidate data from various structured and unstructured data systems into structures that are suitable for building analytics solutions that use Microsoft Azure data services. This program consists of 10 courses to help prepare you to take Exam DP-203: Data Engineering on Microsoft Azure. Each course teaches you the concepts and skills that are measured by the exam. By the end of this Professional Certificate, you will be ready to take and sign-up for the Exam DP-203: Data Engineering on Microsoft Azure.Applied Learning ProjectLearners will engage in interactive exercises throughout this program that offers opportunities to practice and implement what they are learning. They use the Microsoft Learn Sandbox. This is a free environment that allows learners to explore Microsoft Azure and get hands-on with live Microsoft Azure resources and services.Skills measured on Microsoft Azure DP-203 ExamDesign and Implement Data Storage (40-45%)Design and implement data storage (40-45%)Design and develop data processing (25-30%)Design and implement data security (10-15%)Monitor and optimize data storage and data processing (10-15%)The exam measures your ability to accomplish the following technical tasks: design and implement data storage; design and develop data processing; design and implement data security; and monitor and optimize data storage and data processing.Functional groupsDesign and implement data storage (40-45%)Design a data storage structureDesign an Azure Data Lake solutionRecommend file types for storageRecommend file types for analytical queriesDesign for efficient queryingDesign for data pruningDesign a folder structure that represents the levels of data transformationDesign a distribution strategyDesign a data archiving solutionDesign a partition strategyDesign a partition strategy for filesDesign a partition strategy for analytical workloadsDesign a partition strategy for efficiency/performanceDesign a partition strategy for Azure Synapse AnalyticsIdentify when partitioning is needed in Azure Data Lake Storage Gen2Design the serving layerDesign star schemasDesign slowly changing dimensionsDesign a dimensional hierarchyDesign a solution for temporal dataDesign for incremental loadingDesign analytical storesDesign metastores in Azure Synapse Analytics and Azure DatabricksImplement physical data storage structuresImplement compressionImplement partitioning Implement shardingImplement different table geometries with Azure Synapse Analytics poolsImplement data redundancyImplement distributionsImplement data archivingImplement logical data structuresBuild a temporal data solutionBuild a slowly changing dimensionBuild a logical folder structureBuild external tablesImplement file and folder structures for efficient querying and data pruningImplement the serving layerDeliver data in a relational starDeliver data in Parquet filesMaintain metadataImplement a dimensional hierarchyDesign and develop data processing (25-30%)Ingest and transform dataTransform data by using Apache SparkTransform data by using Transact-SQLTransform data by using Data FactoryTransform data by using Azure Synapse PipelinesTransform data by using Stream AnalyticsCleanse dataSplit dataShred JSONEncode and decode dataConfigure error handling for the transformationNormalize and denormalize valuesTransform data by using ScalaPerform data exploratory analysisDesign and develop a batch processing solutionDevelop batch processing solutions by using Data Factory, Data Lake, Spark, Azure Synapse Pipelines, PolyBase, and Azure DatabricksCreate data pipelinesDesign and implement incremental data loadsDesign and develop slowly changing dimensionsHandle security and compliance requirementsScale resourcesConfigure the batch sizeDesign and create tests for data pipelinesIntegrate Jupyter/Python notebooks into a data pipelineHandle duplicate dataHandle missing dataHandle late-arriving dataUpsert dataRegress to a previous stateDesign and configure exception handlingConfigure batch retentionDesign a batch processing solutionDebug Spark jobs by using the Spark UIDesign and develop a stream processing solutionDevelop a stream processing solution by using Stream Analytics, Azure Databricks, and Azure Event HubsProcess data by using Spark structured streamingMonitor for performance and functional regressionsDesign and create windowed aggregatesHandle schema driftProcess time series dataProcess across partitionsProcess within one partitionConfigure checkpoints/watermarking during processingScale resourcesDesign and create tests for data pipelinesOptimize pipelines for analytical or transactional purposesHandle interruptionsDesign and configure exception handlingUpsert dataReplay archived stream dataDesign a stream processing solutionManage batches and pipelinesTrigger batchesHandle failed batch loadsValidate batch loadsManage data pipelines in Data Factory/Synapse PipelinesSchedule data pipelines in Data Factory/Synapse PipelinesImplement version control for pipeline artifactsManage Spark jobs in a pipelineDesign and implement data security (10-15%)Design security for data policies and standardsDesign data encryption for data at rest and in transitDesign a data auditing strategyDesign a data masking strategyDesign for data privacyDesign a data retention policyDesign to purge data based on business requirementsDesign Azure role-based access control (Azure RBAC) and POSIX-like Access Control List (ACL) for Data Lake Storage Gen2Design row-level and column-level securityImplement data securityImplement data maskingEncrypt data at rest and in motionImplement row-level and column-level securityImplement Azure RBACImplement POSIX-like ACLs for Data Lake Storage Gen2Implement a data retention policyImplement a data auditing strategyManage identities, keys, and secrets across different data platform technologiesImplement secure endpoints (private and public)Implement resource tokens in Azure DatabricksLoad a DataFrame with sensitive informationWrite encrypted data to tables or Parquet filesManage sensitive informationMonitor and optimize data storage and data processing (10-15%)Monitor data storage and data processingImplement logging used by Azure MonitorConfigure monitoring servicesMeasure performance of data movementMonitor and update statistics about data across a systemMonitor data pipeline performanceMeasure query performanceMonitor cluster performanceUnderstand custom logging optionsSchedule and monitor pipeline testsInterpret Azure Monitor metrics and logsInterpret a Spark directed acyclic graph (DAG)Optimize and troubleshoot data storage and data processingCompact small filesRewrite user-defined functions (UDFs)Handle skew in dataHandle data spillTune shuffle partitionsFind shuffling in a pipelineOptimize resource managementTune queries by using indexersTune queries by using cacheOptimize pipelines for analytical or transactional purposesOptimize pipeline for descriptive versus analytical workloadsTroubleshoot a failed spark jobTroubleshoot a failed pipeline run

Skills

Reviews