Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free CompTIA DataAI Certification Exam DY0-001 Exam Questions

Page: 1 / 9 Total 85 questions

Want more questions? Get Premium Access.

Question 1

A data scientist is attempting to identify sentences that are conceptually similar to each other within a set of text files. Which of the following is the best way to prepare the data set to accomplish this task after data ingestion?

Correct Answer: A. Embeddings
Explanation:

Generating embeddings transforms each sentence into a dense numerical vector in a semantic space, where conceptually similar sentences lie close together, enabling straightforward similarity calculations (e.g., cosine similarity) to group or identify related sentences.


Question 2

A data analyst is examining the correlation matrix of a new data set to identify issues that could adversely impact model performance. Which of the following is the analyst most likely checking for?

Correct Answer: B. Multicollinearity
Explanation:

Examining a correlation matrix helps identify predictors that are highly correlated with each other, which can inflate variance in coefficient estimates and degrade model reliability - i.e., multicollinearity.


Question 3

A data scientist is building an inferential model with a single predictor variable. A scatter plot of the independent variable against the real-number dependent variable shows a strong relationship between them. The predictor variable is normally distributed with very few outliers. Which of the following algorithms is the best fit for this model, given the data scientist wants the model to be easily interpreted?

Correct Answer: C. A linear regression

Question 4

A data analyst wants to save a newly analyzed data set to a local storage option. The data set must meet the following requirements:

Which of the following file types is the best to use?

Correct Answer: B. Parquet
Explanation:

Parquet is a columnar storage format that automatically includes schema (data types), uses efficient compression to minimize file size, and enables very fast reads for analytic workloads.


Question 5

Which of the following explains back propagation?

Correct Answer: D. The passage of errors backward through a neural network to update weights and biases
Explanation:

Back propagation computes the gradient of the loss (error) with respect to each weight by propagating the error signal backward through the network, then uses those gradients to adjust weights and biases.


Question 6

A computer vision model is trained to identify cats on a training set that is composed of both cat and dog images. The model predicts a picture of a cat is a dog. Which of the following describes this error?

Correct Answer: D. Type II error
Explanation:

Classifying an actual cat (positive instance) as a dog (negative prediction) is a false negative, which corresponds to a Type II error.


Question 7

Which of the following best describes the minimization of the residual term in a ridge linear regression?

Correct Answer: C. e2
Explanation:

Ridge regression extends ordinary least squares by adding an L2 penalty on the coefficients, but it still minimizes the sum of squared residuals (e) as its loss term.


Question 8

Which of the following best describes the minimization of the residual term in a LASSO linear regression?

Correct Answer: D. e2
Explanation:

LASSO regression retains the ordinary least squares loss by minimizing the sum of squared residuals (e), with an added L1 penalty on the coefficients, but the residual term itself remains squared.


Question 9

A data scientist is deploying a model that needs to be accessed by multiple departments with minimal development effort by the departments. Which of the following APIs would be best for the data scientist to use?

Correct Answer: D. REST
Explanation:

RESTful APIs use standard HTTP methods and lightweight data formats (typically JSON), making them easy for diverse teams to integrate with minimal effort and without heavy tooling.


Question 10

Which of the following environmental changes is most likely to resolve a memory constraint error when running a complex model using distributed computing?

Correct Answer: D. Adding nodes to a cluster deployment
Explanation:

Increasing the number of nodes in your cluster directly expands the total available memory across the distributed system, alleviating memoryconstraint errors without changing your code or deployment paradigm. Containerization or edge deployments don't inherently provide more memory, and migrating to the cloud alone doesn't guarantee additional nodes unless you explicitly scale out.