Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free Amazon AWS Certified Machine Learning Engineer - Associate MLA-C01 Exam Questions

Page: 1 / 17 Total 241 questions

Want more questions? Get Premium Access.

Question 1

An ML engineer at a credit card company built and deployed an ML model by using Amazon SageMaker AI. The model was trained on transaction data that contained very few fraudulent transactions. After deployment, the model is underperforming.

What should the ML engineer do to improve the model's performance?

Correct Answer: C. Use Synthetic Minority Oversampling Technique (SMOTE) to generate synthetic minority samples and retrain the model.
Explanation:

This is a classic class imbalance problem, where fraudulent transactions (minority class) are severely underrepresented. AWS documentation for SageMaker Data Wrangler recommends SMOTE (Synthetic Minority Oversampling Technique) as an effective approach for improving model performance in such scenarios.

SMOTE generates synthetic minority samples by interpolating between existing minority class examples. This improves the model's ability to learn decision boundaries without simply duplicating data, which can cause overfitting.

Random undersampling removes valuable majority class data, reducing overall model robustness. Random oversampling duplicates data and increases overfitting risk. Changing algorithms does not address the root cause.

AWS best practices highlight SMOTE as the preferred technique for fraud detection and other highly imbalanced datasets.

Therefore, Option C is the correct and AWS-verified answer.


Question 2

An ML engineer is building a logistic regression model to predict customer churn for subscription services. The dataset contains two string variables: location and job_seniority_level.

The location variable has 3 distinct values, and the job_seniority_level variable has over 10 distinct values.

The ML engineer must perform preprocessing on the variables.

Which solution will meet this requirement?

Correct Answer: B. Apply one-hot encoding to location. Apply ordinal encoding to job_seniority_level.
Explanation:

Logistic regression requires numeric input features and is sensitive to how categorical variables are encoded. AWS feature engineering best practices recommend one-hot encoding for low-cardinality categorical variables with no inherent order and ordinal encoding for categorical variables with a meaningful order.

The location feature has only three distinct values and no ordinal relationship, making one-hot encoding the most appropriate method. This prevents the model from inferring a false numerical relationship between locations.

The job_seniority_level feature typically has an inherent order (for example: junior, mid-level, senior, lead). Even with more than 10 categories, ordinal encoding preserves this natural hierarchy while keeping the feature dimensionality manageable.

Tokenization is used for unstructured text, not structured categorical variables. Standard scaling applies only to continuous numeric features and is not suitable for categorical string variables.

AWS documentation explicitly highlights using one-hot encoding for nominal features and ordinal encoding for ordered categorical features when preparing data for linear models such as logistic regression.

Therefore, Option B is the correct and AWS-aligned solution.


Question 3

A company is building a conversational AI assistant on Amazon Bedrock. The company is using Retrieval Augmented Generation (RAG) to reference the company's internal knowledge base. The AI assistant uses the Anthropic Claude 4 foundation model (FM).

The company needs a solution that uses a vector embedding model, a vector store, and a vector search algorithm.

Which solution will develop the AI assistant with the LEAST development effort?

Correct Answer: A. Use Amazon Kendra Experience Builder.
Explanation:

Amazon Kendra Experience Builder provides a fully managed, low-code solution for building conversational search and question-answering applications. AWS documentation states that Kendra natively supports semantic search, vector embeddings, and vector-based retrieval, making it well suited for RAG-style applications with minimal development effort.

When integrated with Amazon Bedrock, Kendra can act as the retrieval layer, handling document ingestion, indexing, embedding generation, and relevance ranking automatically. This eliminates the need to manually manage embedding models, vector databases, and search logic.

Options B and C require custom schema design, vector indexing, query logic, and operational management of PostgreSQL instances. Although pgvector supports vector search, it significantly increases development and maintenance effort. Option D is unrelated to vector search and is used only for metadata cataloging.

AWS explicitly positions Amazon Kendra as the fastest way to build enterprise-grade conversational assistants that integrate with foundation models.

Therefore, Option A is the correct and most AWS-aligned solution.


Question 4

A healthcare analytics company wants to segment patients into groups that have similar risk factors to develop personalized treatment plans. The company has a dataset that includes patient health records, medication history, and lifestyle changes. The company must identify the appropriate algorithm to determine the number of groups by using hyperparameters.

Which solution will meet these requirements?

Correct Answer: B. Use the Amazon SageMaker k-means clustering algorithm. Set k to specify the number of clusters.
Explanation:

The problem described is a patient segmentation use case, which is a classic example of unsupervised learning. The objective is to group patients with similar characteristics without predefined labels. AWS documentation clearly states that Amazon SageMaker k-means is designed specifically for clustering and segmentation tasks.

The SageMaker k-means algorithm groups data points into clusters based on feature similarity and requires the user to define the number of clusters using the k hyperparameter. This directly satisfies the requirement to ''determine the number of groups by using hyperparameters.'' AWS recommends k-means for applications such as customer segmentation, risk grouping, and pattern discovery in healthcare data.

Option A (XGBoost) is a supervised learning algorithm used for classification and regression. The max_depth hyperparameter controls tree complexity, not the number of groups, making it unsuitable for this task.

Option C (DeepAR) is a time-series forecasting algorithm optimized for predicting future values, not clustering patients.

Option D (Random Cut Forest) is an anomaly detection algorithm. While useful for identifying outliers or unusual patient behavior, it does not perform clustering or group segmentation.

AWS SageMaker documentation explicitly identifies k-means as the correct choice when the goal is to partition data into a predefined number of clusters using a tunable hyperparameter.

Therefore, Option B is the correct and AWS-verified answer.


Question 5

An ML engineer is training a simple neural network model. The model's performance improves initially and then degrades after a certain number of epochs.

Which solutions will mitigate this problem? (Select TWO.)

Correct Answer: A. Enable early stopping on the model.; B. Increase dropout in the layers.
Explanation:

The described behavior indicates overfitting, where the model starts to memorize training data instead of generalizing.

Early stopping halts training when validation performance stops improving, preventing the model from overfitting further. AWS documentation recommends early stopping as a primary regularization technique.

Dropout randomly disables neurons during training, forcing the model to learn robust representations and reducing reliance on specific neurons. Increasing dropout is a well-established method for improving generalization.

Increasing layers or neurons increases model capacity and worsens overfitting. Model bias is unrelated to epoch-based degradation.

Therefore, Options A and B are correct.


Question 6

A company uses Amazon SageMaker AI to create ML models. The data scientists need fine-grained control of ML workflows, DAG visualization, experiment history, and model governance for auditing and compliance.

Which solution will meet these requirements?

Correct Answer: C. Use SageMaker Pipelines with SageMaker Studio and SageMaker ML Lineage Tracking.
Explanation:

Amazon SageMaker Pipelines provides native orchestration of ML workflows with fine-grained control, DAG-based visualization, and seamless integration with SageMaker Studio. AWS documentation explicitly states that Pipelines is designed for end-to-end ML workflow automation and visualization.

SageMaker ML Lineage Tracking records relationships between datasets, models, training jobs, and endpoints, enabling full auditability and governance, which is essential for compliance.

SageMaker Experiments tracks experiment metrics but does not provide lineage-level governance. CodePipeline is a general CI/CD service and lacks ML-specific DAG visualization and lineage tracking.

AWS best practices recommend combining SageMaker Pipelines + SageMaker Studio + ML Lineage Tracking for enterprise-grade ML workflow management.

Therefore, Option C is the correct and AWS-verified solution.


Question 7

An ML engineer needs to use an ML model to predict the price of apartments in a specific location.

Which metric should the ML engineer use to evaluate the model's performance?

Correct Answer: D. Mean absolute error (MAE)
Explanation:

When predicting continuous variables, such as apartment prices, it's essential to evaluate the model's performance using appropriate regression metrics. The Mean Absolute Error (MAE) is a widely used metric for this purpose.

Understanding Mean Absolute Error (MAE):

MAE measures the average magnitude of errors in a set of predictions, without considering their direction. It calculates the average absolute difference between predicted values and actual values, providing a straightforward interpretation of prediction accuracy.

Advantages of MAE:

Interpretability: MAE is expressed in the same units as the target variable, making it easy to understand.

Robustness to Outliers: Unlike metrics that square the errors (e.g., Mean Squared Error), MAE does not disproportionately penalize larger errors, making it more robust to outliers.

Comparison with Other Metrics:

Accuracy, AUC, F1 Score: These metrics are designed for classification tasks, where the goal is to predict discrete labels. They are not suitable for regression problems involving continuous target variables.

Mean Squared Error (MSE): While MSE also measures prediction errors, it squares the differences, giving more weight to larger errors. This can be useful in certain contexts but may be sensitive to outliers.

Conclusion:

For evaluating the performance of a model predicting apartment prices---a continuous variable---MAE is an appropriate and effective metric. It provides a clear indication of the average prediction error in the same units as the target variable, facilitating straightforward interpretation and comparison.


Regression Metrics -- GeeksforGeeks

Evaluation Metrics for Your Regression Model -- Analytics Vidhya

Regression Metrics for Machine Learning -- Machine Learning Mastery

Question 8

A company has deployed an ML model that detects fraudulent credit card transactions in real time in a banking application. The model uses Amazon SageMaker Asynchronous Inference. Consumers are reporting delays in receiving the inference results.

An ML engineer needs to implement a solution to improve the inference performance. The solution also must provide a notification when a deviation in model quality occurs.

Which solution will meet these requirements?

Correct Answer: A. Use SageMaker real-time inference for inference. Use SageMaker Model Monitor for notifications about model quality.
Explanation:

SageMaker real-time inference is designed for low-latency, real-time use cases, such as detecting fraudulent transactions in banking applications. It eliminates the delays associated with SageMaker Asynchronous Inference, improving inference performance.

SageMaker Model Monitor provides tools to monitor deployed models for deviations in data quality, model performance, and other metrics. It can be configured to send notifications when a deviation in model quality is detected, ensuring the system remains reliable.


Question 9

A company must install a custom script on any newly created Amazon SageMaker AI notebook instances.

Which solution will meet this requirement with the LEAST operational overhead?

Correct Answer: A. Create a lifecycle configuration script to install the custom script when a new SageMaker AI notebook is created. Attach the lifecycle configuration to every new SageMaker AI notebook as part of the creation steps.
Explanation:

AWS recommends lifecycle configuration scripts as the simplest and most direct way to customize Amazon SageMaker Notebook Instances at creation time. Lifecycle configurations run automatically when a notebook instance is created or started, allowing scripts, packages, and system dependencies to be installed without manual intervention.

This approach is fully supported, requires no additional infrastructure, and integrates directly with the notebook creation workflow. The script can be reused across notebooks, ensuring consistency.

Options B, C, and D introduce unnecessary complexity, such as container management, private package repositories, or event-driven orchestration.

Therefore, lifecycle configuration scripts provide the least operational overhead solution.


Question 10

A company has historical data that shows whether customers needed long-term support from company staff. The company needs to develop an ML model to predict whether new customers will require long-term support.

Which modeling approach should the company use to meet this requirement?

Correct Answer: C. Logistic regression
Explanation:

Logistic regression is a suitable modeling approach for this requirement because it is designed for binary classification problems, such as predicting whether a customer will require long-term support ('yes' or 'no'). It calculates the probability of a particular class and is widely used for tasks like this where the outcome is categorical.


Question 11

A retail company is analyzing customer purchase data to develop personalized product recommendations. The company wants to use Amazon SageMaker Clarify to assess fairness metrics across different customer groups to avoid potential bias in the recommendation system.

The recommendation system needs to identify if certain customer segments are underrepresented in the training data. The company needs to choose a pre-training bias metric in SageMaker Clarify.

Which metric meets these requirements?

Correct Answer: C. Class imbalance ratio
Explanation:

The correct answer is C. Class imbalance ratio.

Amazon SageMaker Clarify provides both pre-training and post-training bias detection capabilities to help data scientists ensure fairness in ML models. Pre-training bias metrics specifically analyze the training dataset to detect underrepresentation of certain groups or features before a model is trained. The class imbalance ratio measures how evenly different classes or sensitive attributes are represented in the dataset. A high imbalance indicates that a particular segment of customers (e.g., certain age groups, genders, or ethnicities) may be underrepresented, which could lead to biased predictions when the model is deployed.

Feature attribution bias (Option B) and prediction distribution skew (Option A) are post-training metrics. Feature attribution bias evaluates whether the model relies disproportionately on certain features for predictions, while prediction distribution skew examines whether the model's outputs are skewed for specific groups. Model performance gap (Option D) is also a post-training measure, focusing on differences in accuracy or other performance metrics between groups.

By using the class imbalance ratio during pre-training analysis, the company can identify and mitigate underrepresentation before training the model. Actions may include resampling the dataset, weighting underrepresented classes, or applying synthetic data augmentation to ensure that the model is trained on a balanced, fair dataset. This aligns with AWS best practices for ML solution monitoring, maintenance, and security, particularly for responsible AI and fairness assessment.

In summary, the class imbalance ratio allows proactive bias detection at the dataset level, preventing unfair recommendations and ensuring equitable treatment across all customer segments in the recommendation system.


Question 12

Case Study

A company is building a web-based AI application by using Amazon SageMaker. The application will provide the following capabilities and features: ML experimentation, training, a

central model registry, model deployment, and model monitoring.

The application must ensure secure and isolated use of training data during the ML lifecycle. The training data is stored in Amazon S3.

The company needs to run an on-demand workflow to monitor bias drift for models that are deployed to real-time endpoints from the application.

Which action will meet this requirement?

Correct Answer: A. Configure the application to invoke an AWS Lambda function that runs a SageMaker Clarify job.
Explanation:

Monitoring bias drift in deployed machine learning models is crucial to ensure fairness and accuracy over time. Amazon SageMaker Clarify provides tools to detect bias in ML models, both during training and after deployment. To monitor bias drift for models deployed to real-time endpoints, an effective approach involves orchestrating SageMaker Clarify jobs using AWS Lambda functions.

Implementation Steps:

Set Up Data Capture:

Enable data capture on the SageMaker endpoint to record input data and model predictions. This captured data serves as the basis for bias analysis.

Develop a Lambda Function:

Create an AWS Lambda function configured to initiate a SageMaker Clarify job. This function will process the captured data to assess bias metrics.

Schedule or Trigger the Lambda Function:

Configure the Lambda function to run on-demand or at scheduled intervals using Amazon CloudWatch Events or EventBridge. This setup allows for regular bias monitoring as per the application's requirements.

Analyze and Respond to Results:

After each Clarify job completes, review the generated bias reports. If bias drift is detected, take appropriate actions, such as retraining the model or adjusting data preprocessing steps.

Advantages of This Approach:

Automation: Utilizing AWS Lambda for orchestrating Clarify jobs enables automated and scalable bias monitoring without manual intervention.

Cost-Effectiveness: AWS Lambda's serverless nature ensures that you only pay for the compute time consumed during the execution of the function, optimizing resource usage.

Flexibility: The solution can be tailored to specific monitoring needs, allowing for adjustments in monitoring frequency and analysis parameters.

By implementing this solution, the company can effectively monitor bias drift in real-time, ensuring that the AI application maintains fairness and accuracy throughout its lifecycle.


Bias drift for models in production - Amazon SageMaker

Schedule Bias Drift Monitoring Jobs - Amazon SageMaker

Question 13

A company's ML engineer has deployed an ML model for sentiment analysis to an Amazon SageMaker AI endpoint. The ML engineer needs to explain to company stakeholders how the model makes predictions.

Which solution will provide an explanation for the model's predictions?

Correct Answer: B. Use SageMaker Clarify on the deployed model.
Explanation:

Explaining how a model makes predictions is the domain of model interpretability and explainability. Amazon SageMaker Clarify is designed specifically to provide explanations for ML predictions using techniques such as SHAP (SHapley Additive exPlanations).

SageMaker Clarify can analyze deployed endpoints to show feature importance, explain individual predictions, and quantify how each input feature contributes to the model's output. This makes it ideal for communicating model behavior to non-technical stakeholders and meeting transparency requirements.

Model Monitor focuses on data and performance drift, not explanations. A/B testing and shadow endpoints compare performance but do not explain predictions.

Therefore, SageMaker Clarify is the correct solution for explaining model predictions.


Question 14

A company uses an ML model to recommend videos to users. The model is deployed on Amazon SageMaker AI. The model performed well initially after deployment, but the model's performance has degraded over time.

Which solution can the company use to identify model drift in the future?

Correct Answer: B. Create a baseline from the training dataset. Then create a monitoring job in SageMaker Model Monitor.
Explanation:

AWS recommends Amazon SageMaker Model Monitor for detecting data drift and model drift in deployed models. Model Monitor works by comparing live inference data against a baseline, which must first be created from the training dataset.

AWS documentation clearly specifies the required order:

Create a baseline using training data statistics

Create a monitoring schedule to compare incoming data against the baseline

Option A reverses this order and is therefore incorrect. Option C is incorrect because SageMaker Clarify focuses on bias and explainability, not ongoing drift detection. Option D is reactive and does not provide continuous monitoring.

Model Monitor integrates with Amazon CloudWatch, enabling automated alerts and downstream retraining workflows. This proactive approach allows companies to detect degradation early and maintain model quality.

Therefore, Option B is the correct and AWS-verified answer.


Question 15

A company is developing ML models by using PyTorch and TensorFlow estimators with Amazon SageMaker AI. An ML engineer configures the SageMaker AI estimator and now needs to initiate a training job that uses a training dataset.

Which SageMaker AI SDK method can initiate the training job?

Correct Answer: A. fit method
Explanation:

In the Amazon SageMaker Python SDK, the fit() method is used to start a training job after an estimator has been configured. AWS documentation explicitly states that once an estimator (such as PyTorch or TensorFlow) is defined with parameters like instance type, framework version, and hyperparameters, the fit() method is responsible for launching the training process.

The fit() method accepts the training data location (commonly an Amazon S3 URI) and initiates the managed training job on SageMaker infrastructure. SageMaker then provisions the required compute resources, stages the data, executes the training script, and stores model artifacts in Amazon S3.

The create_model() method is used after training to create a SageMaker model object from trained artifacts. The deploy() method deploys a trained model to an endpoint for inference. The predict() method is used only after deployment to request predictions from an endpoint.

AWS documentation clearly separates these lifecycle steps and identifies fit() as the correct method to initiate training.

Therefore, Option A is the correct and AWS-verified answer.