Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free Microsoft Implementing Data Engineering Solutions Using Azure Databricks DP-750 Exam Questions

Page: 1 / 10 Total 91 questions

Want more questions? Get Premium Access.

Question 1

You have an Azure Databricks workspace that contains a Git folder and uses Azure Repos as the Git provider. From the main branch, you create a branch named Branch1. You commit changes to Branch1.

You need to incorporate the changes from Branch1 into main The solution must preserve the commit history in the repository. Which command should you run?

Correct Answer: A. merge
Explanation:

The correct answer is A --- merge.

A Git merge combines the histories of two branches by creating a merge commit that joins them. Every individual commit from Branch1 remains visible in the repository log --- the full development history is preserved. This is the requirement: 'the solution must preserve the commit history in the repository.'

Option C (rebase) moves Branch1's commits on top of main by replaying them as new commits with new hashes. The end result looks like a linear history, but the original commit hashes are rewritten --- the prior history is not preserved in its original form. For a shared repository, rebase rewrites public history, which is considered problematic.

Option B (pull) fetches remote changes and merges or rebases them into the current branch --- it's used to sync with a remote, not to incorporate a feature branch. Option D (push) sends local commits to the remote but doesn't incorporate any branch into another.


Question 2

You have an Azure Databricks workspace that is enabled for Unity Catalog.

You need to profile a table to meet the following requirements:

The count of null values per column must be evaluated repeatedly as new records are added to the table.

Changes in the count of null values must be observable over the progression of the dataset.

Which type of profile should you create?

Correct Answer: C. time series
Explanation:

A time series profile repeatedly calculates data-quality metrics as a table changes and preserves those measurements over time. This makes it possible to observe whether the number of null values in each column is increasing, decreasing, or remaining stable as new records are added. A snapshot profile evaluates the table at a particular point in time and is therefore insufficient when the requirement is to analyze metric progression across multiple updates. Inference is associated with deriving information such as a schema or statistical characteristics; it is not the profile type used to maintain a historical sequence of monitoring measurements. Because the question requires repeated evaluation and the ability to observe changes throughout the dataset's progression, the time series profile satisfies both requirements.


Question 3

You have an Azure Databricks workspace that is enabled for Unity Catalog.

You plan to ingest data from CSV files stored in Azure Data Lake Storage Gen2. New rows are appended frequently.

You need to implement a data ingestion solution that meets the following requirements:

* New data must be available in near-real-time (NRT).

* The data must be stored in managed Delta tables.

* The solution must minimize custom code and maintenance effort.

What should you include in the solution?

Correct Answer: D. Auto Loader
Explanation:

Auto Loader incrementally detects and processes new files arriving in Azure Data Lake Storage Gen2 through the cloudFiles Structured Streaming source. It supports near-real-time ingestion while automatically tracking processed files, reducing the custom state-management code required. Its output can be written to a managed Delta table, and built-in schema inference and evolution reduce ongoing maintenance. Scheduled Spark batch jobs introduce latency based on their schedule and usually require custom file-tracking logic. An external table over CSV files does not ingest the data into a managed Delta table. Azure Data Factory can orchestrate ingestion, but it introduces another service and more configuration than the native Databricks capability needed here. Auto Loader is therefore the most direct and maintainable solution for continuously arriving cloud files. Microsoft Learn


Question 4

You have an Azure Databricks workspace named Workspace1 that contains a lakehouse and is enabled for Unity Catalog.

You have a connection to a Microsoft SQL Server database named DB1.

You need to expose the schemas and tables of DB1 to meet the following requirements:

* The schemas and tables can be queried in Databricks.

* The schemas and tables appear alongside other Unity Catalog objects.

* The data is NOT copied into Databricks-managed storage.

Solution: You create a Lakeflow Connect pipeline and connect it to DB1. Does this meet the goal?

Correct Answer: B. No
Explanation:

The correct answer is B --- No.

Lakeflow Connect is an ingestion service that physically copies data from external databases into Delta tables managed by Databricks. It's designed for scenarios where you want a replicated, writable Delta copy of external data --- essentially a CDC-based ingestion pipeline.

That's the opposite of what's required here. The requirement states 'the data is NOT copied into Databricks-managed storage.' Lakeflow Connect would create Delta tables in Databricks and copy DB1's data into them --- a direct violation.

Additionally, Lakeflow Connect creates Databricks-native Delta tables rather than exposing DB1's original schemas and tables as virtual objects. Analysts querying through a Lakeflow Connect pipeline are querying a replicated copy, not the live source.

For zero-copy, live query federation of an external SQL Server into Unity Catalog, Lakehouse Federation (foreign catalog) is the correct tool.


Question 5

You use Databricks Asset Bundles to manage two jobs and an app.

You need to deploy the bundle to development and production environments. The solution must meet the following requirements

* Deploy the app to both environments.

* Deploy only one job to development.

* Minimize administrative effort.

What should you use?

Correct Answer: D. a targets node in a databricks.yml file
Explanation:

The correct answer is D --- a targets node in databricks.yml.

Databricks Asset Bundles use a single databricks.yml to define all resources (jobs, apps, pipelines) once, and a targets node to define per-environment overrides. Within the development target, you can use the include/exclude mechanism or resource-level overrides to deploy only one of the two jobs. The app and the second job are deployed to both environments through the shared resource definition.

Option B (separate databricks.yml files per environment) works technically but means duplicating the shared resource definitions across files --- any change to a shared resource requires edits in multiple places, which is exactly the administrative overhead the question wants to avoid.

Option A (resources node) defines resources globally across all targets --- it doesn't provide environment-specific filtering. Option C (variables node) parameterises values like cluster sizes or paths but doesn't control which resources are deployed to which environment.


Question 6

You have a Lakeflow Spark Declarative Pipelines {SDP) pipeline in Azure Databricks. The pipeline ingests transaction data into a table named Table1.

You need to ensure that in the event of an invalid record, the pipeline continues to run. The solution must meet the following requirements:

* Invalid records must NOT be written to Table 1.

* Invalid records must be preserved for review.

* Minimize development effort

What should you do?

Correct Answer: B. Define a pipeline expectation.
Explanation:

The correct answer is B --- define a pipeline expectation.

SDP expectations with @dlt.expect_or_drop are built precisely for this scenario: the pipeline keeps running, bad records are excluded from Table1, and those records are automatically captured in the pipeline's event log as expectation violations --- available for review without any extra code.

Option A (custom quarantine logic) would work but requires writing and maintaining additional pipeline tables and routing logic. The whole point of SDP expectations is to handle this pattern declaratively, with far less code.

Option C (WHERE clauses in downstream queries) is a read-time filter, not a write-time guard. Invalid records would still land in Table1 and would simply be hidden from downstream views --- they're not preserved for review in any structured way. Option D (check constraint on Table1) would throw an exception on write and halt the pipeline, violating the 'pipeline continues to run' requirement.


Question 7

You have an Azure Databricks workspace named Workspace! that uses a Git repository. The repository contains a Databricks notebook named Notebook1.

From the main branch, you create a feature branch named Branch! and commit changes to Notebooks Another user commits changes to Notebook1 in main.

When you attempt to merge Branch! into main, the merge fails due to conflicts.

You need to merge Branch! into the main branch. The solution must ensure that Notebook1 includes all the changes from both the branches.

What should you do?

Correct Answer: D. Apply the main branch changes to Branch! and resolve the conflicts.
Explanation:

The correct answer is D --- apply the main branch changes to Branch1 and resolve the conflicts.

When a merge fails due to conflicts, the right workflow is to bring main's changes into the feature branch, resolve conflicts there, and then merge the clean feature branch into main. This is the standard Git conflict resolution pattern --- resolve in the feature branch, not in main --- because it protects the main branch from partial or broken states during resolution.

Option A (clone Branch1 as a new repository) creates a disconnected copy; it doesn't resolve the conflict and breaks the relationship with the remote. Option B (apply changes directly to main) bypasses the feature branch entirely and risks overwriting the other developer's work. Option C (clone main as a new repository) again creates a disconnected copy --- none of Branch1's changes would be incorporated, and history would be lost.


Question 8

You have an Azure Databricks workspace that uses Unity Catalog.

You have a Lakeflow Spark Declarative Pipelines (SDP) pipeline that ingests data into a managed Delta table named Table1. Table! is used for analytics.

New columns are added to the source data, causing pipeline failures during writes to Table!

You need to prevent the pipeline failures. The solution must ensure that schema changes are detected and handled.

What should you do?

Correct Answer: C. Enable schema evolution.
Explanation:

The correct answer is C --- Enable schema evolution.

When new columns are added to the source data, a pipeline without schema evolution treats the unexpected columns as a schema mismatch and fails the write. Schema evolution, when enabled in an SDP pipeline, automatically adds those new columns to the target Delta table on the next pipeline run. The pipeline continues without intervention, and no historical data is lost.

Option A (disable schema enforcement) is the wrong lever --- it removes all schema validation, which could allow corrupt or mistyped data into Table1. Schema evolution is a targeted, safer response.

Option B (row filters to exclude records with new columns) would silently discard valid records just because they carry extra fields --- that's data loss. Option D (separate table per schema version) creates an explosion of tables as schemas evolve and makes downstream analytics significantly more complex. Schema evolution is the clean, built-in solution.


Question 9

You have an Azure Databricks workspace named Workspace1.

You create a compute cluster named Cluser1 that will be used to ingest data.

You need to install the required libraries on Cluster1. The solution must use Unity Catalog for access control.

What should you do?

Correct Answer: C. Upload the libraries to a Unity Catalog volume and install the libraries on Cluster1.
Explanation:

A Unity Catalog volume is the appropriate location because it provides governed file storage with permissions managed through Unity Catalog. After uploading the library package to the volume, it can be configured as a cluster library on Cluster1. This approach permits centralized access control, auditing, and lifecycle management. Option A installs the library without placing its source under Unity Catalog governance. Option B provides notebook-scoped dependency management but does not, by itself, satisfy the requirement that Unity Catalog control access to the library artifact. A schema is a logical container for tables, views, functions, models, and volumes; a library file cannot be uploaded directly to the schema itself, eliminating option D. Unity Catalog volumes explicitly support storing cluster libraries and job dependencies. Microsoft Learn


Question 10

You have an Azure Databricks workspace that contains a job in Lakeflow Jobs named Job1.

Job! processes raw data files stored in Azure Storage.

New files arrive at unpredictable intervals.

You need to ensure that Job1 starts automatically when new files arrive and does NOT consume compute resources when no data is available.

Which type of job trigger should you use?

Correct Answer: C. file arrival
Explanation:

The correct answer is C --- File Arrival trigger.

File Arrival monitors a specified Azure Storage path and fires a job run each time a new file lands there. This ticks both requirements: Job1 starts automatically in response to new data (no human involvement), and when no files arrive, no job runs --- no cluster spins up, no compute cost is incurred.

Option A (scheduled) runs at fixed intervals regardless of whether files are waiting. A quiet weekend still kicks off hourly (or daily) runs, burning compute for nothing. Option B (continuous) keeps the job running perpetually, consuming resources even during long gaps between file arrivals --- exactly what 'does NOT consume compute resources when no data is available' rules out. Option D (manual) requires a person to trigger every run, which is unsuitable for unpredictable arrival patterns.