Limited-Time Offer: Enjoy 50% Savings! Ends in 00h 00m 00s Coupon code: 50OFF
Skip to content

Free PeopleCert DEVOPS Institute Site Reliability Engineering Foundation DevOps-SRE Exam Questions

Page: 1 / 8 Total 80 questions

Want more questions? Get Premium Access.

Question 1

If SREs own some sections of a service, but not others, then this organizational approach is known as __________________

Correct Answer: C. Slice and dice
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

The Slice-and-Dice model is an SRE adoption pattern where the SRE team owns specific portions of a service---typically the most critical, complex, or high-risk components---while development teams own the rest.

From the SRE Workbook, Organizational Models section:

''In the slice-and-dice model, SREs take responsibility for particular portions of a service or system rather than owning the entire thing. This works well when parts of the system require stronger reliability engineering than others.''

This model is used when:

Services are large or complex

Only certain components need SRE-level reliability

Full SRE ownership is not feasible

Why the other options are incorrect:

A Consultant SREs advise; they do not own components

B Full SRE fully owns the entire service

D Platform SRE builds shared reliability tooling, not owning service slices

Thus, C. Slice and dice is the correct answer.


SRE Workbook, ''SRE Organizational Patterns''

Site Reliability Engineering Book, ''Engagement Models''

Question 2

Which type of engineering work will reduce toil within the service?

Correct Answer: D. Internal automation
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

Toil-reduction engineering focuses on making the service itself easier to operate. The most direct way to achieve this is through internal automation --- automation built into the service that eliminates repetitive, manual operational tasks.

The Site Reliability Engineering Book, Chapter ''Eliminating Toil,'' states:

''Automation that replaces manual, repetitive operational tasks is the primary mechanism for reducing toil. The most effective form of toil reduction is automation that is integrated directly into the service itself.''

The SRE Workbook reinforces:

''Internal automation contributes directly to service reliability and reduces the operational burden by ensuring that manual tasks are permanently removed.''

Why the other options are not the best answer:

A Continuous delivery pipelines reduce release friction but do not directly remove service-operational toil.

B External scripts and tools help but are less effective and harder to maintain than internal automation.

C Scalable infrastructure reduces linear-scaling toil but does not address broader operational burdens.

Thus, the correct answer is D.


Site Reliability Engineering Book, ''Eliminating Toil''

SRE Workbook, ''Toil Reduction Approaches''

Question 3

How does automation reduce toil?

Correct Answer: A. Automated releases can replace manual releases
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

Automation is the primary method of reducing toil in SRE. The Google Site Reliability Engineering Book, Chapter ''Eliminating Toil,'' states:

''Automation is the most effective tool for reducing toil. Any recurring, manual, automatable task should be automated to prevent it from consuming engineering time.''

Automated release systems directly eliminate toil by:

Removing manual deployment steps

Removing repeated, error-prone human processes

Increasing reliability and consistency

Freeing engineers for high-value project work

The SRE Workbook reinforces this:

''CI/CD pipelines and release automation remove significant operational toil by replacing manual processes with repeatable, reliable automation.''

Why the other answers are incorrect:

B AI is not required for toil reduction.

C Meeting travel is not an SRE toil concern.

D Incorrect; automation dramatically reduces long-term toil, even though initial setup requires effort.

Thus, A is the correct answer.


Site Reliability Engineering Book, ''Eliminating Toil''

SRE Workbook, ''Toil Reduction Strategies''

Question 4

What types of outages must fit into an Error Budget?

Correct Answer: C. Any planned or unplanned outage
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

An error budget accounts for all downtime, including both planned and unplanned outages. This is a critical SRE principle: the user does not distinguish between maintenance downtime and accidental downtime --- therefore, neither should the SLO nor the error budget.

The SRE Book, Chapter ''Service Level Objectives,'' states:

''From the user's perspective, availability is simply whether the service is working or not, regardless of whether the outage was planned or unplanned.''

This means all downtime counts toward the error budget.

Additionally, the SRE Workbook reinforces this point:

''Error budgets must include every form of unavailability --- maintenance events, configuration changes, emergency work, and unexpected incidents.''

This confirms that planned outages (maintenance windows) and unplanned outages (incidents) both consume error budget.

Why the other options are incorrect:

A Only includes unplanned incidents; SRE requires counting planned outages as well.

B Defect fixes may contribute to downtime, but ''defect fixes'' alone are not a downtime category.

D CAB approval has no bearing on whether outages count toward error budgets.

Thus, C is correct: any planned or unplanned outage must be included.


Site Reliability Engineering Book, ''Service Level Objectives''

SRE Workbook, ''Implementing SLOs''

Question 5

In a blameless post-mortem, those involved report

Correct Answer: C. Both A and B
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

A blameless post-mortem is a foundational SRE practice that encourages truthful, detailed reporting after an incident. The purpose is to learn, not punish. Google SRE emphasizes that engineers must feel psychologically safe to report what they did, what they assumed, and why they made those decisions.

From the Site Reliability Engineering Book, Chapter ''Postmortem Culture'':

''Blameless postmortems encourage engineers to share the full details of their actions and assumptions without fear of punishment, enabling learning and preventing repeated failures.''

The book further states:

''Understanding the assumptions made during an incident is critical to uncovering systemic issues.''

Thus:

Engineers must report without fear of retribution

They must report assumptions and decisions made during the incident

Therefore, the correct answer is C. Both A and B.

Why the other options are insufficient:

A Only partially correct

B Only partially correct

D Testing data may be included, but it is not the defining feature of blameless postmortems


Site Reliability Engineering Book, ''Postmortem Culture''

SRE Workbook, ''Learning from Incidents''

Question 6

Which of the following BEST illustrates the role of a launch coordination engineer?

Correct Answer: C. A software engineer who acts as a consultant and liaison between the parties involved in a launch
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

Google's SRE model includes the role of Launch Coordination Engineer (LCE), described in the SRE Book as: ''an engineer who serves as the central liaison between product teams, SRE, and other stakeholders to ensure safe and reliable launches.'' (SRE Book -- Chapter: Production Environment & Launch Coordination). Their responsibilities include assessing launch readiness, ensuring SLOs are defined, facilitating cross-team communication, and managing risk associated with new service rollouts.

Option C precisely reflects this role: acting as a consultant and liaison across all parties involved in a launch.

Option A focuses on server engineering, which is not the focus of LCE.

Option B describes application-level performance work, unrelated to cross-team launch facilitation.

Option D describes operational tuning, not coordination.

Thus, C is the correct answer, capturing the SRE-defined launch coordination function.


Site Reliability Engineering: How Google Runs Production Systems, Chapter: ''Handling Overload and Launch Coordination.''

The Site Reliability Workbook, Sections on production readiness and launch processes.

Question 7

Reliability is a key pillar of digital experience monitoring and incident management.

Which of the following describes the BEST type of reliability monitoring strategy in SRE?

Correct Answer: B. A strategy that instruments observability and provides monitoring insights across all components and layers
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

SRE defines effective monitoring as comprehensive observability across all layers of a system, including latency, traffic, errors, saturation, dependencies, and infrastructure. The SRE Book states: ''Monitoring must offer insight across all system components, enabling teams to rapidly detect and diagnose issues.'' (SRE Book -- Monitoring Distributed Systems). Observability instrumentation (logs, metrics, traces) provides the necessary depth for reliable digital experience monitoring.

Option B captures this exactly: broad observability across all components and layers.

Option A rejects modern observability practices---contradicting SRE guidance.

Option C is too narrow (network-only).

Option D focuses only on advanced technologies, not comprehensive coverage.

Thus, B is the best answer.


Site Reliability Engineering, Chapter: ''Monitoring Distributed Systems.''

The Site Reliability Workbook, Observability and Monitoring chapters.

Question 8

Which of the following BEST describes the engineering side of SRE?

Correct Answer: D. Applying software development best practices to solving operational problems and automating solutions
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

The foundational definition of SRE, as stated in Google's SRE Book, is that SRE uses software engineering as its primary tool to solve operational problems: ''SRE is fundamentally doing operations work using software engineering approaches.'' (SRE Book -- What Is SRE?). This includes building automation, writing tools, creating pipelines, and eliminating manual work. The ''engineering side'' focuses specifically on applying coding practices, testing, CI/CD, version control, and automation frameworks to operational domains such as deployment, monitoring, incident response, and capacity planning.

Option D captures this precisely: using software engineering best practices to solve operational issues and drive automation.

Options A, B, and C focus too narrowly on network or infrastructure engineering. While these can be components of SRE, they do not describe its engineering foundation as Google defines it.

Thus, D is the correct answer.


Site Reliability Engineering: How Google Runs Production Systems, Introduction & Chapter: ''What is SRE?''

The Site Reliability Workbook, Chapter: ''Eliminating Toil.''

Question 9

A bank has been using traditional monitoring tools for ensuring that their systems are available and operating as planned. Their strategic initiatives now include a renewed focus on customer experience as well as identifying ways to scale service.

Why would migrating to an observability approach be important now?

Correct Answer: D. All of the above
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

All the listed reasons correctly describe why observability becomes essential in modern, user-focused, dynamically scaling architectures.

The SRE Workbook and Google Observability guidance both emphasize that traditional monitoring is insufficient in environments where:

Services are distributed

Traffic is unpredictable

Customer experience is a priority

Cloud-native, containerized, or microservice architectures are used

Key excerpts:

From Google's Observability guidance:

''Monitoring relies on known failure modes; observability enables teams to explore unknown-unknowns and understand complex, dynamic systems.''

From the SRE Workbook:

''As systems scale and architectures shift toward microservices or containers, component-level monitoring provides an incomplete picture. Observability enables teams to understand user impact and system behavior holistically.''

Thus:

A Observability is critical for containerized and dynamic environments.

B Component monitoring alone cannot show customer experience or end-to-end reliability.

C Observability helps teams diagnose issues that could not be predicted in advance ('unknown unknowns').

All statements are correct, making D the correct answer.


SRE Workbook, ''Monitoring and Observability''

Google Cloud Architecture Framework: ''Observability vs Monitoring''

Site Reliability Engineering Book, Alerting & Monitoring chapters

Question 10

An organization is experiencing significant turnover of IT operational staff with most not staying more than one year. The HR Director and IT Director are trying to determine why they are having difficulty retaining IT operations professionals.

What could be one of the reasons?

Correct Answer: D. All of the above
Explanation:

Comprehensive and Detailed Explanation From Exact Extract:

High turnover in IT operations roles is often driven by a combination of factors, not just one. The Google SRE Book, Chapter ''Eliminating Toil,'' outlines that excessive toil, unpredictable work, and overload contribute to burnout and churn:

''Excessive operational workload and interrupt-driven work lead to burnout and high attrition among engineering and operational staff.''

The SRE Workbook adds:

''Teams overwhelmed with toil struggle to innovate, automate, or develop new skills, creating frustration and increasing turnover.''

Each option listed represents a recognized driver of burnout in SRE and operations environments:

Overload and disruptive work patterns are known contributors to burnout.

Lack of time for skills development demotivates engineers and prevents career growth.

Backlog-driven cultures force teams into reactive rather than proactive work.

The combination of these factors matches common causes of attrition in operations teams. Therefore, all of the above is the correct answer.


Site Reliability Engineering Book, ''Eliminating Toil''

SRE Workbook, ''Addressing Operational Overload''