DP-750: Select and configure compute in a workspace

Select and configure compute in a workspace

1. Compute Types and Workspace Architecture

The “Select and configure compute in a workspace” topic within the “Set up and configure an Azure Databricks environment” section of DP-750 focuses on choosing the correct compute resources for different workloads within Azure Databricks. Candidates should understand the different compute options available, including all-purpose clusters, job clusters, SQL warehouses, and serverless compute where supported. Each compute type is designed for specific use cases such as interactive notebook development, scheduled ETL pipelines, business intelligence queries, machine learning workloads, or real-time streaming. The exam expects candidates to understand how compute resources integrate with Azure Databricks workspaces and support scalable distributed processing using Apache Spark.

2. Cluster Configuration and Optimization

A major focus area is configuring clusters for performance, scalability, and cost efficiency. Candidates should understand cluster modes, runtime versions, autoscaling settings, autotermination, and worker node configuration. Knowledge of Databricks Runtime selection is important, including runtimes optimized for machine learning, SQL analytics, or Photon acceleration. Candidates should also understand the impact of driver and worker node sizing, cluster policies, and spot instances on reliability and operational cost. The exam expects candidates to understand how cluster configuration affects Spark job execution, concurrency, and workload performance across notebooks, pipelines, and SQL workloads.

3. Workload Management and Performance Tuning

Candidates should understand how to select appropriate compute resources based on workload requirements. Interactive workloads may require all-purpose clusters for collaborative notebook development, while automated ETL workloads commonly use job clusters for improved isolation and cost control. SQL warehouses are optimized for BI and reporting scenarios including Power BI integration. Candidates should also understand serverless options where available to simplify operational management. Performance optimization topics include caching, partitioning, autoscaling, Photon acceleration, query optimization, and efficient resource allocation. The exam also expects knowledge of Spark UI concepts such as skew, shuffle, spill, and executor performance to support troubleshooting and tuning activities.

4. Governance, Security, and Operational Best Practices

The exam emphasizes secure and governed compute management within enterprise environments. Candidates should understand how cluster policies enforce organizational standards for security, cost management, and compliance. This includes restricting runtime versions, node types, public IP usage, and workspace permissions. Integration with Unity Catalog is important to ensure compute resources support centralized governance and secure data access. Candidates should also understand monitoring and operational management using Azure Monitor, logs, audit trails, and query history. Overall, this section of DP-750 validates that a candidate can select, configure, optimize, secure, and manage Azure Databricks compute resources to support reliable, scalable, and cost-effective enterprise analytics and AI workloads.

Links

Exam DP-750: Implementing Data Engineering Solutions Using Azure Databricks – Innovative Business Intelligence

Microsoft Certified: Azure Databricks Data Engineer Associate – Certifications | Microsoft Learn