Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free Databricks Certified Associate Developer for Apache Spark 3.5 - Python Databricks-Certified-Associate-Developer-for-Apache-Spark-3.5 Exam Questions

Page: 1 / 14 Total 135 questions

Want more questions? Get Premium Access.

Question 1

Which configuration can be enabled to optimize the conversion between Pandas and PySpark DataFrames using Apache Arrow?

Correct Answer: B. spark.conf.set('spark.sql.execution.arrow.pyspark.enabled', 'true')
Explanation:

Apache Arrow is used under the hood to optimize conversion between Pandas and PySpark DataFrames. The correct configuration setting is:

spark.conf.set('spark.sql.execution.arrow.pyspark.enabled', 'true')

From the official documentation:

''This configuration must be enabled to allow for vectorized execution and efficient conversion between Pandas and PySpark using Arrow.''

Option B is correct.

Options A, C, and D are invalid config keys and not recognized by Spark.

Final Answer: B


Question 2

36 of 55.

What is the main advantage of partitioning the data when persisting tables?

Correct Answer: D. It optimizes by reading only the relevant subset of data from fewer partitions.
Explanation:

Partitioning a dataset divides data into separate directories based on partition column values. When queries filter on partitioned columns, Spark can prune irrelevant partitions --- meaning it only reads files that match the filter criteria.

Advantage:

Reduces I/O and improves performance by scanning only relevant subsets of data.

Example:

/data/sales/year=2023/month=10/...

/data/sales/year=2024/month=01/...

A query filtering WHERE year = 2024 reads only the relevant partition.

Why the other options are incorrect:

A: Compression is independent of partitioning.

B: Spark does not automatically clean partitions unless managed manually.

C: Partitioning does not cause Spark to load entire data into memory.


Databricks Exam Guide (June 2025): Section ''Using Spark SQL'' --- partitioning and pruning for optimized data retrieval.

Spark SQL Documentation --- DataFrameWriter partitionBy() and query optimization.

Question 3

44 of 55. A data engineer is working on a real-time analytics pipeline using Spark Structured Streaming. They want the system to process incoming data in micro-batches at a fixed interval of 5 seconds.

Which code snippet fulfills this requirement?

A.

query = df.writeStream \

.outputMode("append") \

.trigger(processingTime="5 seconds") \

.start()

B.

query = df.writeStream \

.outputMode("append") \

.trigger(continuous="5 seconds") \

.start()

C.

query = df.writeStream \

.outputMode("append") \

.trigger(once=True) \

.start()

D.

query = df.writeStream \

.outputMode("append") \

.start()

Correct Answer: A. Option A
Explanation:

To process data in fixed micro-batch intervals, use the .trigger(processingTime='interval') option in Structured Streaming.

Correct usage:

query = df.writeStream \

.outputMode('append') \

.trigger(processingTime='5 seconds') \

.start()

This instructs Spark to process available data every 5 seconds.

Why the other options are incorrect:

B: continuous triggers are for continuous processing mode (different execution model).

C: once=True runs the stream a single time (batch mode).

D: Default trigger runs as fast as possible, not fixed intervals.


PySpark Structured Streaming Guide --- Trigger types: processingTime, once, continuous.

Databricks Exam Guide (June 2025): Section ''Structured Streaming'' --- controlling streaming triggers and batch intervals.

Question 4

16 of 55.

A data engineer is reviewing a Spark application that applies several transformations to a DataFrame but notices that the job does not start executing immediately.

Which two characteristics of Apache Spark's execution model explain this behavior? (Choose 2 answers)

Correct Answer: C. Transformations are evaluated lazily.; E. Only actions trigger the execution of the transformation pipeline.
Explanation:

Apache Spark follows a lazy evaluation model, meaning transformations (like filter(), select(), map()) are not executed immediately. Instead, they build a logical plan (lineage graph) that represents the sequence of operations to be applied.

Execution only begins when an action (e.g., count(), collect(), save(), show()) is called. At that point, Spark's engine:

Optimizes the logical plan into a physical plan.

Divides it into stages and tasks.

Executes them across the cluster.

This design helps Spark optimize execution paths and avoid unnecessary computations.

Why the other options are incorrect:

A: Transformations do not execute immediately; they are deferred.

B: Optimization happens during job execution (after an action), not during transformations.

D: Execution starts automatically once an action is triggered, no manual intervention needed.


Databricks Exam Guide (June 2025): Section ''Apache Spark Architecture and Components'' --- covers lazy evaluation, actions vs. transformations, and execution hierarchy.

Spark 3.5 Documentation --- Lazy Evaluation model and DAG scheduling.

Question 5

A data scientist is working on a project that requires processing large amounts of structured data, performing SQL queries, and applying machine learning algorithms. The data scientist is considering using Apache Spark for this task.

Which combination of Apache Spark modules should the data scientist use in this scenario?

Options:

Correct Answer: D. Spark DataFrames, Spark SQL, and MLlib
Explanation:

Comprehensive

To cover structured data processing, SQL querying, and machine learning in Apache Spark, the correct combination of components is:

Spark DataFrames: for structured data processing

Spark SQL: to execute SQL queries over structured data

MLlib: Spark's scalable machine learning library

This trio is designed for exactly this type of use case.

Why other options are incorrect:

A: GraphX is for graph processing --- not needed here.

B: Pandas API on Spark is useful, but MLlib is essential for ML, which this option omits.

C: Spark Streaming is legacy; GraphX is irrelevant here.


Question 6

20 of 55.

What is the difference between df.cache() and df.persist() in Spark DataFrame?

Correct Answer: D. cache() --- Persists the DataFrame with the default storage level (MEMORY_AND_DISK_DESER), and persist() --- Can be used to set different storage levels to persist the contents of the DataFrame.
Explanation:

Both cache() and persist() are Spark DataFrame storage operations that store computed results in memory (and optionally on disk) to speed up subsequent actions on the same DataFrame.

Key difference:

cache() is a shorthand for persist(StorageLevel.MEMORY_AND_DISK).

persist() allows specifying different storage levels, such as MEMORY_ONLY, DISK_ONLY, or MEMORY_AND_DISK_SER.

Example:

df.cache() # Uses MEMORY_AND_DISK by default

df.persist(StorageLevel.MEMORY_ONLY) # Custom storage level

Both trigger caching upon an action (e.g., count(), collect()).

Why the other options are incorrect:

A: persist() default is not DISK_ONLY; default storage level is MEMORY_AND_DISK.

B/C: cache() cannot set arbitrary levels; only persist() can.


PySpark API Reference --- DataFrame.cache() and DataFrame.persist().

Databricks Exam Guide (June 2025): Section ''Developing Apache Spark DataFrame/DataSet API Applications'' --- caching, persistence, and storage levels.

Question 7

A data engineer is reviewing a Spark application that applies several transformations to a DataFrame but notices that the job does not start executing immediately.

Which two characteristics of Apache Spark's execution model explain this behavior?

Choose 2 answers:

Correct Answer: B. Only actions trigger the execution of the transformation pipeline.; E. Transformations are evaluated lazily.
Explanation:

Apache Spark employs a lazy evaluation model for transformations. This means that when transformations (e.g., map(), filter()) are applied to a DataFrame, Spark does not execute them immediately. Instead, it builds a logical plan (lineage) of transformations to be applied.

Execution is deferred until an action (e.g., collect(), count(), save()) is called. At that point, Spark's Catalyst optimizer analyzes the logical plan, optimizes it, and then executes the physical plan to produce the result.

This lazy evaluation strategy allows Spark to optimize the execution plan, minimize data shuffling, and improve overall performance by reducing unnecessary computations.


Question 8

8 of 55.

A data scientist at a large e-commerce company needs to process and analyze 2 TB of daily customer transaction data. The company wants to implement real-time fraud detection and personalized product recommendations.

Currently, the company uses a traditional relational database system, which struggles with the increasing data volume and velocity.

Which feature of Apache Spark effectively addresses this challenge?

Correct Answer: B. In-memory computation and parallel processing capabilities
Explanation:

Apache Spark was designed for big data and high-velocity workloads. Its core strength lies in its in-memory computation and parallel distributed processing model.

These features allow Spark to:

Process large-scale datasets quickly across many nodes.

Support real-time and near--real-time analytics for tasks like fraud detection and recommendations.

Minimize disk I/O through caching and memory persistence.

Thus, the key advantage in this use case is Spark's ability to handle large data volumes efficiently using distributed, in-memory computation.

Why the other options are incorrect:

A: Spark is optimized for large, not small, datasets.

C: SQL support is useful but doesn't solve the scalability issue.

D: MLlib supports machine learning but relies on Spark's parallel computation for speed.


Databricks Exam Guide (June 2025): Section ''Apache Spark Architecture and Components'' --- identifies Spark's advantages: in-memory processing, distributed computation, and scalability.

Apache Spark 3.5 Overview --- Key design goals and cluster computation model.

Question 9

A Spark application is experiencing performance issues in client mode because the driver is resource-constrained.

How should this issue be resolved?

Correct Answer: C. Switch the deployment mode to cluster mode
Explanation:

In Spark's client mode, the driver runs on the local machine that submitted the job. If that machine is resource-constrained (e.g., low memory), performance degrades.

From the Spark documentation:

'In cluster mode, the driver runs inside the cluster, benefiting from cluster resources and scalability.'

Option A is incorrect --- executors do not help the driver directly.

Option B might help short-term but does not scale.

Option C is correct --- switching to cluster mode moves the driver to the cluster.

Option D (local mode) is for development/testing, not production.

Final Answer: C


Question 10

41 of 55. A data engineer is working on the DataFrame df1 and wants the Name with the highest count to appear first (descending order by count), followed by the next highest, and so on.

The DataFrame has columns:

id | Name | count | timestamp

---------------------------------

1 | USA | 10

2 | India | 20

3 | England | 50

4 | India | 50

5 | France | 20

6 | India | 10

7 | USA | 30

8 | USA | 40

Which code fragment should the engineer use to sort the data in the Name and count columns?

Correct Answer: A. df1.orderBy(col('count').desc(), col('Name').asc())
Explanation:

To sort a Spark DataFrame by multiple columns, use .orderBy() (or .sort()) with column expressions.

Correct syntax for descending and ascending mix:

from pyspark.sql.functions import col

df1.orderBy(col('count').desc(), col('Name').asc())

This sorts primarily by count in descending order and secondarily by Name in ascending order (alphabetically).

Why the other options are incorrect:

B/C: Default sort order is ascending; won't place highest counts first.

D: Reverses sorting logic --- sorts Name descending, not required.


PySpark DataFrame API --- orderBy() and col() for sorting with direction.

Databricks Exam Guide (June 2025): Section ''Using Spark DataFrame APIs'' --- sorting, ordering, and column expressions.