Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free Google Professional Cloud DevOps Engineer Professional-Cloud-DevOps-Engineer Exam Questions

Page: 1 / 14 Total 205 questions

Want more questions? Get Premium Access.

Question 1

You are managing an application that exposes an HTTP endpoint without using a load balancer. The latency of the HTTP responses is important for the user experience. You want to understand what HTTP latencies all of your users are experiencing. You use Stackdriver Monitoring. What should you do?

Correct Answer: C. * In your application, create a metric with a metricKind set to gauge and a valueType set to distribution.* In Stackdriver's Metrics Explorer, use a Heatmap graph to visualize the metric.
Explanation:

https://sre.google/workbook/implementing-slos/

https://cloud.google.com/architecture/adopting-slos/

Latency is commonly measured as a distribution. Given a distribution, you can measure various percentiles. For example, you might measure the number of requests that are slower than the historical 99th percentile.


Question 2

You need to create a Cloud Monitoring SLO for a service that will be published soon. You want to verify that requests to the service will be addressed in fewer than 300 ms at least 90% Of the time per calendar month. You need to identify the metric and evaluation method to use. What should you do?

Correct Answer: A. Select a latency metric for a request-based method of evaluation.
Explanation:

The correct answer is A. Select a latency metric for a request-based method of evaluation.

A latency metric measures how responsive your service is to users.For example, you can use thecloud.googleapis.com/http/server/response_latenciesmetric to measure the latency of HTTP requests to your service1. A request-based method of evaluation counts the number of successful requests that meet a certain criterion, such as being below a latency threshold, and compares it to the number of all requests.For example, you can define an SLI as the ratio of requests with latency below 300 ms to all requests2. A request-based method of evaluation is suitable for measuring performance over time, such as per calendar month.You can set an SLO for the SLI to be at least 90%, which means that you expect 90% of the requests to have latency below 300 ms in a month3.


Creating an SLO | Operations Suite | Google Cloud, Choosing a metric, Latency metric.Concepts in service monitoring | Operations Suite | Google Cloud, Service-level indicators, Request-based SLIs.Learn how to set SLOs -- SRE tips | Google Cloud Blog, Setting SLOs.

Question 3

You need to define Service Level Objectives (SLOs) for a high-traffic multi-region web application. Customers expect the application to always be available and have fast response times. Customers are currently happy with the application performance and availability. Based on current measurement, you observe that the 90th percentile of latency is 120ms and the 95th percentile of latency is 275ms over a 28-day window. What latency SLO would you recommend to the team to publish?

Correct Answer: C. 90th percentile -- 150ms95th percentile -- 300ms
Explanation:

https://sre.google/sre-book/service-level-objectives/


Question 4

You manage your company's primary revenue-generating application. You have an error budget policy in place that freezes production deployments when the application is close to breaching its SLO. A number of issues have recently occurred, and the application has exhausted its error budget. You need to deploy a new release to the application that includes a feature urgently required by your largest customer. You have been told that the release has passed all unit tests. What should you do?

Correct Answer: D. Deploy the feature to a subset of users, and gradually roll out to all users if there are no errors reported.
Explanation:

Comprehensive and Detailed Explanation From SRE Principles:

This scenario presents a classic SRE conflict: maintaining reliability (as dictated by the exhausted error budget and deployment freeze) versus delivering an urgent business requirement. The error budget policy is there for a reason -- to protect users from further instability.

A . Start the deployment of the feature immediately: This directly violates the established error budget policy and the deployment freeze. While the feature is urgent, deploying without caution when the system is already unstable (as indicated by the exhausted error budget) is highly risky and could exacerbate existing problems or introduce new ones, further impacting revenue and customer trust.

B . Delay the deployment of the feature until the error budget is replenished: This strictly adheres to the policy but might not be acceptable given the 'urgently required by your largest customer' clause. SRE principles allow for reasoned exceptions and risk management, not just blind adherence if the business context is compelling enough and risks are managed.

C . Re-run the unit tests, and start the deployment of the feature if the tests pass: Unit tests are foundational but insufficient to guarantee a complex application will perform reliably in production, especially when the system is already indicating instability (exhausted error budget). Passing unit tests doesn't negate the risk signaled by the depleted error budget.

D . Deploy the feature to a subset of users, and gradually roll out to all users if there are no errors reported: This is the most balanced SRE approach in this situation. It acknowledges the urgency while attempting to mitigate risk:Risk Mitigation: A canary release (deploying to a small subset of users) limits the potential negative impact if the new feature introduces new errors or worsens existing instability.

Observation: It allows for careful monitoring of the new release in the production environment with real users.

Data-Driven Decision: The decision to proceed with a wider rollout is based on observed behavior ('if there are no errors reported'), not just assumptions.

Controlled Rollout: A gradual rollout allows for quick rollback if issues arise.

While an exhausted error budget signals a deployment freeze, critical business needs can sometimes necessitate a carefully managed exception. A canary release is a standard SRE technique for deploying changes with reduced risk, making it the most appropriate course of action when faced with such conflicting priorities. The team would also need to communicate clearly about the risks and the rationale for this exception. It's implied that this urgent feature might also fix existing issues or is critical enough to warrant the carefully managed risk.

Reference (Based on SRE principles from Google's SRE books and general practices):

Error Budgets: 'The SRE Book' (Site Reliability Engineering: How Google Runs Production Systems) discusses error budgets and deployment freezes. An exhausted error budget typically means no more risky changes until reliability improves.

Canary Releases: This is a fundamental practice for safely deploying new versions. It's about testing in production with a small percentage of traffic.

Managing Risk: SRE is about managing risk, not eliminating it entirely. In situations like this, a calculated risk with strong mitigation (canary, monitoring, rollback plan) can be justified for critical business needs. The decision involves weighing the risk of deploying against the risk of not deploying the urgent feature.

Option D represents a pragmatic SRE approach to navigate this difficult situation by minimizing the blast radius of the change.


Question 5

You support a Node.js application running on Google Kubernetes Engine (GKE) in production. The application makes several HTTP requests to dependent applications. You want to anticipate which dependent applications might cause performance issues. What should you do?

Correct Answer: B. Instrument all applications with Stackdriver Trace and review inter-service HTTP requests.

Question 6

Your company is developing applications that are deployed on Google Kubernetes Engine (GKE). Each team manages a different application. You need to create the development and production environments for each team, while minimizing costs. Different teams should not be able to access other teams' environments. What should you do?

Correct Answer: D. Create a Development and a Production GKE cluster in separate projects. In each cluster, create a Kubernetes namespace per team, and then configure Kubernetes Role-based access control (RBAC) so that each team can only access its own namespace.
Explanation:

https://cloud.google.com/architecture/prep-kubernetes-engine-for-prod#roles_and_groups


Question 7

You support a service that recently had an outage. The outage was caused by a new release that exhausted the service memory resources. You rolled back the release successfully to mitigate the impact on users. You are now in charge of the post-mortem for the outage. You want to follow Site Reliability Engineering practices when developing the post-mortem. What should you do?

Correct Answer: B. Focus on identifying the contributing causes of the incident rather than the individual responsible for the cause.

Question 8

You are configuring a CI pipeline. The build step for your CI pipeline integration testing requires access to APIs inside your private VPC network. Your security team requires that you do not expose API traffic publicly. You need to implement a solution that minimizes management overhead. What should you do?

Correct Answer: A. Use Cloud Build private pools to connect to the private VPC.
Explanation:

Cloud Build Private Pools allow your builds to run in a secure, isolated environment with direct access to resources inside your private VPC network, without exposing them to the public internet. This is the Google-recommended approach for minimizing management overhead while maintaining strong security boundaries.

From the official documentation:

'Private pools run your builds in a dedicated and secure environment that can connect to private network resources.'

--- Cloud Build Private Pools Overview

'You can configure private pools to access resources in a VPC network, which enables secure access to services without using public IPs.'

--- Accessing Private Resources

This method eliminates the need to create and manage additional compute instances or complex load balancing setups, providing seamless integration with your private services during CI pipeline execution.


Question 9

You are building the Cl/CD pipeline for an application deployed to Google Kubernetes Engine (GKE) The application is deployed by using a Kubernetes Deployment, Service, and Ingress The application team asked you to deploy the application by using the blue'green deployment methodology You need to implement the rollback actions What should you do?

Correct Answer: C. Update the Kubernetes Service to point to the previous Kubernetes Deployment
Explanation:

The best option for implementing the rollback actions is to update the Kubernetes Service to point to the previous Kubernetes Deployment. A Kubernetes Service is a resource that defines how to access a set of Pods. A Kubernetes Deployment is a resource that manages the creation and update of Pods. By using the blue/green deployment methodology, you can create two Deployments, one for the current version (blue) and one for the new version (green), and use a Service to switch traffic between them. If you need to rollback, you can update the Service to point to the previous Deployment (blue) and stop sending traffic to the new Deployment (green).


Question 10

You support an e-commerce application that runs on a large Google Kubernetes Engine (GKE) cluster deployed on-premises and on Google Cloud Platform. The application consists of microservices that run in containers. You want to identify containers that are using the most CPU and memory. What should you do?

Correct Answer: A. Use Stackdriver Kubernetes Engine Monitoring.
Explanation:

https://cloud.google.com/anthos/clusters/docs/on-prem/1.7/concepts/logging-and-monitoring


Question 11

You are responsible for creating development environments for your company's development team. You want to create environments with identical IDEs for all developers while ensuring that these environments are not exposed to public networks. You need to choose the most cost-effective solution without impacting developer productivity. What should you do?

Correct Answer: A. Create a Cloud Workstations private cluster. Create a workstation configuration with a runningTimeout parameter.
Explanation:

Comprehensive and Detailed 150 to 200 words of Explanation From Google Cloud DevOps guides documents:

According to Google Cloud's documentation on Cloud Workstations, this service is specifically designed to provide managed, secure, and highly customizable development environments. By selecting a private cluster, you ensure that the workstations are not assigned public IP addresses, keeping them entirely off the public internet and satisfying the security requirement. This managed approach is superior to manual Compute Engine setups because it uses container-based configurations to provide identical IDEs and toolsets to every developer, which eliminates environment drift and boosts productivity.

Regarding cost-effectiveness, the runningTimeout parameter is a vital mechanism. While idleTimeout is excellent for short-term inactivity, runningTimeout provides a definitive 'hard stop' to ensure that workstations do not run indefinitely if a developer forgets to shut down their session at the end of a shift, thereby preventing runaway costs. This aligns with Google's SRE and DevOps best practices for cost optimization and resource management. Choosing Cloud Workstations over manual VM management (Options C and D) reduces the operational overhead of patching and maintaining individual machine images, allowing the team to focus on delivery rather than infrastructure maintenance.


Question 12

Your company is developing applications that are deployed on Google Kubernetes Engine (GKE) Each team manages a different application You need to create the development and production environments for each team while you minimize costs Different teams should not be able to access other teams environments You want to follow Google-recommended practices What should you do?

Correct Answer: D. Create a development and a production GKE cluster in separate projects In each cluster create a Kubernetes namespace per team and then configure Kubernetes role-based access control (RBAC) so that each team can only access its own namespace
Explanation:

The best option for creating the development and production environments for each team while minimizing costs and ensuring isolation is to create a development and a production GKE cluster in separate projects, in each cluster create a Kubernetes namespace per team, and then configure Kubernetes role-based access control (RBAC) so that each team can only access its own namespace. This option allows you to use fewer clusters and projects than creating one project or cluster per team, which reduces costs and complexity. It also allows you to isolate each team's environment by using namespaces and RBAC, which prevents teams from accessing other teams' environments.


Question 13

Your product is currently deployed in three Google Cloud Platform (GCP) zones with your users divided between the zones. You can fail over from one zone to another, but it causes a 10-minute service disruption for the affected users. You typically experience a database failure once per quarter and can detect it within five minutes. You are cataloging the reliability risks of a new real-time chat feature for your product. You catalog the following information for each risk:

* Mean Time to Detect (MUD} in minutes

* Mean Time to Repair (MTTR) in minutes

* Mean Time Between Failure (MTBF) in days

* User Impact Percentage

The chat feature requires a new database system that takes twice as long to successfully fail over between zones. You want to account for the risk of the new database failing in one zone. What would be the values for the risk of database failover with the new system?

Correct Answer: B. MTTD:5MTTR: 20MTBF: 90Impact: 33%
Explanation:

https://www.atlassian.com/incident-management/kpis/common-metrics

https://linkedin.github.io/school-of-sre/


Question 14

You use Google Cloud Managed Service for Prometheus with managed collection to gather metrics from your service running on Google Kubernetes Engine (GKE). After deploying the service, there is no metric data appearing in Cloud Monitoring, and you have not encountered any error messages. You need to troubleshoot this issue. What should you do?

Correct Answer: D. Verify that your PodMonitoring configuration references a valid port.
Explanation:

Comprehensive and Detailed

When using Managed Service for Prometheus, metrics may not appear in Cloud Monitoring if PodMonitoring is misconfigured. The most likely issue is that the PodMonitoring configuration does not reference a valid port.

PodMonitoring is required to collect metrics from workloads in GKE.

If the port is incorrect or missing, metrics won't be scraped.

Why not other options?

A (Check quota limits) If quotas were exceeded, you'd see explicit errors, but the question states no errors were encountered.

B (Grafana installation check) Grafana is not required for Prometheus metric collection.

C (monitoring.servicesViewer IAM role issue) This role allows viewing metrics, but does not affect metric collection.

Official Reference:

Managed Service for Prometheus - PodMonitoring

GKE Monitoring Troubleshooting


Question 15

You are using Stackdriver to monitor applications hosted on Google Cloud Platform (GCP). You recently deployed a new application, but its logs are not appearing on the Stackdriver dashboard.

You need to troubleshoot the issue. What should you do?

Correct Answer: A. Confirm that the Stackdriver agent has been installed in the hosting virtual machine.
Explanation:

https://cloud.google.com/monitoring/agent/monitoring/troubleshooting#checklist