Skip to content

ITIL 5 Master : High-Velocity Culture & Toil (Domain 6)

ITIL 5 – Master : Certified ITIL Master - Domain 6 - High Velocity Culture and Toil Reduction

25 questionsmedium

This study guide provides a comprehensive analysis of the core competencies required for the ITIL 5 Master designation, specifically focusing on the intersection of High-Velocity IT (HVIT), Create, Deliver and Support (CDS), and Digital and IT Strategy (DITS). As organizations transition to Digital Product and Service Management (DPSM), the ability to foster high-velocity cultures while maintaining operational resilience is paramount. This guide synthesizes the official ITIL 5 syllabus targets, including the execution of AIOps, the reduction of manual toil, and the establishment of psychological safety through blameless cultures.

1. The ITIL 5 Master Pathway and the Unified Lifecycle

The ITIL 5 Master designation represents the pinnacle of the qualification scheme, recognizing professionals who have mastered the full suite of ITIL competencies across strategic, tactical, and operational levels. To achieve this status, a candidate must successfully complete three primary streams:

  • ITIL Practice Manager: Focused on building practical, real-world capability across specific service management practices.
  • ITIL Managing Professional: Targeted at IT practitioners working in technology and digital teams to build advanced skills in product delivery and experience management.
  • ITIL Strategic Leader: Concentrated on aligning digital product and service management (DPSM) with enterprise strategy, governance, and investment.

The Master designation demonstrates the capability to apply these principles within the ITIL Product and Service Lifecycle Model. Unlike previous versions, ITIL 5 unifies traditional IT management and digital product management into a single, cohesive lifecycle. This unified approach is essential for competing in an AI-enabled, product-centric economy where customer and employee experiences serve as strategic differentiators.

2. High-Velocity Organizational Culture and Psychological Safety

At the heart of Domain 6 is the transition to a high-velocity organizational culture. High-velocity IT is not merely about speed; it is about the ability to move quickly and reliably in environments characterized by volatility, uncertainty, complexity, and ambiguity (VUCA).

Establishing Psychological Safety

A high-velocity culture depends on psychological safety—the belief that one will not be punished or humiliated for speaking up with ideas, questions, concerns, or mistakes. In the context of the Create, Deliver and Support (CDS) module, this culture is a prerequisite for high-performing teams. Without psychological safety, “Shift Left” initiatives fail because team members are hesitant to take on the responsibilities of earlier lifecycle stages or share the honest feedback necessary for continual improvement.

Team Topologies

The structure of teams influences the flow of value. ITIL 5 emphasizes modern team topologies that reduce handovers and cognitive load. By aligning team structures with the value stream, organizations can foster a sense of ownership and accountability. This alignment supports the Four Dimensions of Product and Service Management, ensuring that “Organizations and People” are optimized to support “Value Streams and Processes.”

3. Fostering Blameless Postmortems and Learning from Failure

In a high-velocity environment, failures are viewed as opportunities for systemic improvement rather than occasions for individual reprimand. ITIL 5 advocates for blameless postmortems as a critical practice within the High-Velocity IT (HVIT) syllabus.

The Mechanics of Blamelessness

  • Systemic Focus: Postmortems analyze the conditions, tools, and processes that allowed a failure to occur, rather than searching for a “person to blame.”
  • Transparency: Findings are shared broadly to prevent similar issues across different value streams.
  • Actionable Outcomes: Every postmortem must lead to concrete changes in the digital product or the underlying infrastructure to increase resilience.

Incident Command

For major disruptions, ITIL 5 introduces Incident Command structures. This provides a clear, hierarchical response mechanism during a crisis, ensuring that communication is streamlined and decision-making is centralized. Once the incident is resolved, the Incident Command transition back to a blameless postmortem ensures that the high-velocity culture remains intact even after high-pressure events.

4. Reducing Manual Toil: The Site Reliability Engineering (SRE) Approach

Toil is defined as manual, repetitive, automatable work that provides no long-term value and increases as a service grows. High-Velocity IT (HVIT) focuses heavily on toil reduction to free up engineering capacity for innovation.

The Role of SRE

Site Reliability Engineering (SRE) is a core component of the ITIL 5 Managing Professional stream. SRE applies engineering principles to operations, treating “operations as a software problem.” Key SRE concepts integrated into Domain 6 include:

  • Error Budgets: A defined amount of allowed unreliability. If the budget is exhausted, the team shifts focus from new features to reliability improvements.
  • Service Level Objectives (SLOs): Specific, measurable goals for service performance that define the threshold for the error budget.
  • Automation-First Mindset: Identifying and eliminating toil through the creation of “self-healing” systems and automated recovery scripts.

5. Executing AIOps and Intelligent Observability

As digital environments grow in complexity, human monitoring becomes insufficient. Domain 6 mandates the execution of AIOps (Artificial Intelligence for IT Operations) and the move toward observability.

Monitoring vs. Observability

While traditional monitoring tracks known failure modes (“is the server up?”), observability allows teams to understand the internal state of a system by looking at its external outputs (logs, metrics, and traces). This is crucial for debugging complex, distributed digital products where the failure mode may be “unknown-unknown.”

AIOps Integration

AIOps uses machine learning and data analytics to:

  • Detect Anomalies: Identifying patterns that precede failures before they impact the user.
  • Event Correlation: Reducing “alert fatigue” by grouping related events into a single actionable incident.
  • Automated Remediation: Executing low-risk changes or toil-reduction scripts in response to specific triggers.

The Monitoring and Event Management practice is thus elevated from a passive support function to an active, AI-driven component of the service value chain.

6. Continuous Delivery and the High-Velocity Pipeline

The HVIT syllabus emphasizes the Continuous Delivery Pipeline as the primary vehicle for value realization. Achieving high velocity requires the integration of Change Enablement directly into the automated pipeline.

Change Enablement at Velocity

Traditional, manual Change Advisory Boards (CABs) are often a bottleneck. In high-velocity environments, Change Enablement is “shifted left” and automated:

  • Standard Changes: High-frequency, low-risk changes are pre-approved and executed via the pipeline.
  • Peer Reviews: Coding and configuration changes are validated by peers within the team topology rather than an external board.
  • Automated Testing: The pipeline includes rigorous validation and testing stages, ensuring that only “known good” code reaches production.

Release and Deployment Management

In ITIL 5, release (making a feature available to users) is decoupled from deployment (moving code to a production environment). This allows for “dark launches” and “canary releases,” where new features are tested with a small subset of users before a full-scale rollout, further reducing risk in a high-velocity environment.

7. Platform Engineering and Infrastructure as Code

To support high-velocity teams, organizations are increasingly adopting Platform Engineering. This involves creating internal developer platforms that provide “self-service” capabilities for developers, further reducing manual toil and inter-team dependencies.

Key Concepts in Platform Engineering

  • Standardization: Providing a curated set of tools and infrastructure patterns.
  • Infrastructure as Code (IaC): Managing and provisioning infrastructure through machine-readable definition files, rather than physical hardware configuration or interactive configuration tools.
  • Chaos Engineering Awareness: Proactively introducing controlled failures into a system to test its resilience and identify weaknesses before they cause real-world outages.

Platform Engineering acts as a force multiplier for the Create, Deliver and Support (CDS) value streams, allowing product teams to focus on business logic rather than infrastructure management.

8. Strategic Direction and Digital Transformation

Domain 6 requires a deep understanding of how high-velocity operations align with the ITIL Strategic Leader stream. Strategy in ITIL 5 is defined as a set of decisions and plans that enable an organization to fulfill its purpose.

Strategy in a VUCA Environment

The Digital and IT Strategy (DITS) module explores how strategy evolves in volatile settings. Leaders must use tools like:

  • Wardley Mapping: To understand the landscape and evolution of components within a value chain.
  • Hoshin Kanri and OKR Cascading: To ensure that high-level strategic goals are translated into measurable Objectives and Key Results (OKRs) at the team level.
  • Scenario Planning: To prepare for digital disruption by imagining multiple potential futures.

Digital Vision and Business Alignment

Alignment is no longer about IT “supporting” the business; it is about IT being “in” the business. The Business Model Canvas is used to ensure that every digital product contributes directly to value co-creation and competitive positioning.

9. Value Stream Design and Mapping

The CDS and Managing Professional modules place heavy emphasis on Value Stream Mapping (VSM). A value stream represents the end-to-end flow from an initial idea to the final support of a product.

Optimizing the Flow of Value

ActivityDescription
Idea to SupportThe complete journey of a digital product, encompassing design, development, testing, and operation.
Waste EliminationUsing VSM to identify bottlenecks, unnecessary handovers, and manual toil.
Shift LeftMoving testing, security, and operational considerations to the earliest possible stages of the value stream.
Value Co-CreationEnsuring that every step in the stream contributes to the user experience (UX) and customer experience (CX).

10. Governance, Digital Ethics, and Responsible AI

As organizations adopt AIOps and high-velocity practices, governance must evolve to remain effective without hindering speed. The AI Governance extension module is a critical component of Domain 6.

Responsible AI Adoption

Governance in the AI era focuses on:

  • Digital Ethics: Ensuring that algorithms are transparent, accountable, and free from bias.
  • Transparency: Maintaining a clear “line of sight” into how AI-driven decisions are made within the digital operating model.
  • Regulatory Compliance: Addressing emerging laws regarding data strategy, carbon-aware strategy, and ESG (Environmental, Social, and Governance) reporting.

The Three Lines of Defense

Strategic leaders must implement the “Three Lines of Defense” model to manage risk:

  1. First Line: Operational management (the teams creating value).
  2. Second Line: Risk management and compliance functions.
  3. Third Line: Internal audit, providing independent assurance to the board.

By integrating these lines into the high-velocity pipeline, organizations achieve “Governance as Code,” where compliance checks are automated and continuous.


Short-Answer Questions

  1. What is the primary objective of a “blameless postmortem” in an HVIT environment?
  2. Define “manual toil” as used in the context of SRE and HVIT.
  3. How does ITIL 5 distinguish between “deployment” and “release”?
  4. What are the three required designations to achieve the ITIL 5 Master certification?
  5. What is the function of an “Error Budget” in Site Reliability Engineering?
  6. Name two strategic mapping tools mentioned in the ITIL Strategic Leader syllabus.
  7. What is “Shift Left” and how does it impact the Create, Deliver and Support (CDS) stream?
  8. What does the “VUCA” acronym stand for in the context of digital strategy?
  9. What is the core focus of the ITIL AI Governance extension module?
  10. Explain the purpose of “Incident Command” within a high-velocity culture.

Answer Key and Rationales

  1. To identify systemic failures and process improvements without individual blame. Rationale: Blameless postmortems focus on the “how” and “what” of a failure to prevent recurrence, rather than the “who.”
  2. Repetitive, manual tasks that provide no long-term value and scale linearly with service growth. Rationale: Toil reduction is essential in HVIT to free up resources for high-value engineering work.
  3. Deployment is the technical act of moving code to production; Release is making that functionality available to users. Rationale: Decoupling these allows for safer, more controlled feature launches (e.g., canary releases).
  4. ITIL Practice Manager, ITIL Managing Professional, and ITIL Strategic Leader. Rationale: The Master designation requires completion of all three advanced streams to prove total framework competency.
  5. It defines the acceptable level of unreliability, balancing the need for speed with the requirement for stability. Rationale: If the budget is spent, the team must prioritize reliability over new features.
  6. Wardley Mapping and Strategy Maps. Rationale: These tools help leaders visualize the competitive landscape and align technology with business outcomes.
  7. Moving activities (like testing or security) to earlier stages of the lifecycle. Rationale: This improves quality and reduces the cost of fixing defects found late in the value stream.
  8. Volatility, Uncertainty, Complexity, and Ambiguity. Rationale: VUCA describes the dynamic environment where high-velocity IT and agile strategies are most necessary.
  9. The responsible, ethical, and compliant adoption of artificial intelligence. Rationale: It addresses risks related to transparency, accountability, and regulatory considerations in AI.
  10. To provide a clear, structured response and communication hierarchy during major incidents. Rationale: Incident Command ensures decision-making remains effective and efficient during high-pressure disruptions.

Open-Ended and Design-Thinking Questions

  1. Cultural Transformation: Given an organization with a deeply ingrained “blame culture,” design a phased roadmap to transition the workforce toward a high-velocity culture that embraces blameless postmortems and psychological safety.
  2. Toil Elimination: Analyze a hypothetical “Idea to Support” value stream for a banking app. Identify three areas where manual toil likely exists and propose specific automation or SRE practices to eliminate them.
  3. AIOps Implementation: You are tasked with implementing AIOps for a global e-commerce platform. How would you structure the event management practice to balance AI-driven automated remediation with the need for human oversight and digital ethics?
  4. Strategic Alignment: Use the Hoshin Kanri or OKR cascading model to demonstrate how a corporate goal of “Reducing Carbon Footprint by 20%” would be reflected in the daily operations of a platform engineering team.
  5. Digital Disruption Response: Design a “Target Operating Model” for a traditional retailer facing a new digital-only competitor. How would you integrate SIAM (Service Integration and Management) and high-velocity IT principles to increase their agility?

Glossary of Key Terms

  1. AIOps (Artificial Intelligence for IT Operations): The application of machine learning and data science to IT operations to automate event correlation and anomaly detection.
  2. Blameless Postmortem: A post-incident review process that focuses on systemic causes rather than individual human error.
  3. Chaos Engineering: The practice of intentionally introducing failures into a system to verify its resilience and identify hidden vulnerabilities.
  4. Continuous Delivery Pipeline: A set of automated processes that allow for the rapid and reliable release of digital products.
  5. Digital Product and Service Management (DPSM): The unified approach in ITIL 5 that merges traditional ITSM with modern digital product management.
  6. Error Budget: The maximum amount of time or frequency a service is allowed to fail within a specific period, used to balance innovation and reliability.
  7. Hoshin Kanri: A strategic planning methodology that ensures the goals of a company are communicated and implemented at every level.
  8. Incident Command: A standardized approach to the command, control, and coordination of emergency response during major incidents.
  9. Infrastructure as Code (IaC): The management of infrastructure through machine-readable definition files rather than manual configuration.
  10. Observability: The ability to measure the internal state of a system by examining its external outputs, such as logs and metrics.
  11. Platform Engineering: The discipline of designing and building self-service internal platforms to improve developer productivity and system reliability.
  12. Psychological Safety: A shared belief that a team is safe for interpersonal risk-taking, essential for a high-velocity culture.
  13. Shift Left: The practice of moving testing, security, and operational checks to earlier stages in the development lifecycle.
  14. Site Reliability Engineering (SRE): A discipline that incorporates aspects of software engineering and applies them to IT operations problems.
  15. Toil: Repetitive, manual work associated with running a service that can be automated and provides no enduring value.
  16. Value Stream Mapping (VSM): A lean management tool used to analyze the current state and design a future state for the series of events that take a product from beginning to end-user.
  17. VUCA: An acronym (Volatility, Uncertainty, Complexity, Ambiguity) used to describe the challenging environments in which modern digital strategies must operate.
  18. Wardley Mapping: A strategic tool that maps components of a value chain against their evolutionary stage to help identify competitive advantages.
  19. XLA (Experience Level Agreement): A commitment to a specific level of experience for users or customers, moving beyond traditional technical SLAs.
  20. Target Operating Model (TOM): A description of the future state of an organization’s operations, designed to deliver its digital strategy.

Leaderboard

No scores saved yet. Be the first!

25 Questions — ITIL 5 – Master : Certified ITIL Master - Domain 6 - High Velocity Culture and Toil Reduction

Expand any question to reveal the correct answer and explanation.

  1. 1 An organization is transitionging to a high-velocity culture and implementing blameless postmortems. What is the primary objective of this shift regarding human error?

    Consider the difference between finding a person to blame and finding a process to fix.

    To treat human error as a starting point for investigation into systemic vulnerabilities rather than a conclusion.

    Blameless postmortems seek to understand the environmental and systemic factors that allowed an error to occur, fostering a learning-oriented culture.

    • To identify the specific individuals responsible for service outages to provide targeted retraining.

      Focusing on individual culpability contradicts the blameless philosophy, which views human error as a symptom of deeper systemic issues.

    • To eliminate the need for disciplinary action by categorizing all incidents as 'unavoidable technical failures'.

      Simply re-categorizing failures does not improve systemic resilience or address the underlying process gaps that blamelessness aims to fix.

    • To ensure that all postmortem documentation remains anonymous to protect the psychological safety of the workforce.

      Anonymity is not the goal; rather, the goal is a culture where individuals feel safe to openly discuss their actions to help the organization learn.

  2. 2 In the context of ITIL (Version 5) High Velocity IT, which characteristic distinguishes 'toil' from other forms of technical work?

    Focus on the repetitive and tactical nature of the activities.

    The work is manual, repetitive, automatable, and scales linearly with service growth.

    Toil represents administrative or operational tasks that do not provide enduring value and grow at the same rate as the service itself.

    • The work requires high levels of specialized human judgment and creative problem-solving.

      Toil is specifically characterized by its tactical, repetitive nature and lack of enduring value, whereas creative work is strategic.

    • The work is essential for long-term service strategy but is currently underfunded.

      Toil refers to the nature of the tasks performed, not their budget status or strategic importance.

    • The work involves complex cross-functional collaboration and stakeholder management.

      Complex collaboration typically requires the human judgment that toil lacks; toil is often rote and solitary.

  3. 3 A team uses AIOps to reduce the 'noise' generated by monitoring systems. According to the ITIL AI Capability Model ($6C$ Model), which capability is being primarily utilized?

    This capability involves recognizing patterns and detecting anomalies in data.

    Cognition

    Cognition encompasses pattern recognition and anomaly detection, which are essential for filtering noise and identifying significant events.

    • Creation

      Creation involves generating new designs or content, whereas noise reduction is about processing existing data streams.

    • Communication

      Communication refers to conversational interfaces and portals, not the backend processing of system alerts.

    • Coordination

      Coordination involves orchestrating workflows and integrations, which usually follows the initial detection of a relevant event.

  4. 4 Which component of a High Velocity Culture is most critical for allowing team members to respectfully disagree and take measured risks without fear of retribution?

    Think about the interpersonal environment necessary for innovation and learning from failure.

    Psychological Safety

    Psychological safety is the foundation that enables a safety culture where mistakes are viewed as learning opportunities and dissent is valued.

    • Automated Change Risk Scoring

      While this supports technical speed, it does not address the interpersonal dynamics required for risk-taking and healthy disagreement.

    • Incident Command Systems

      Incident Command is a structural approach to managing crises, not a cultural framework for daily interpersonal risk.

    • Standard Operating Procedures (SOPs)

      Rigid adherence to SOPs can often stifle the very innovation and risk-taking that a high-velocity culture seeks to promote.

  5. 5 When implementing 'Change Enablement at Velocity', what is the preferred method for managing low-risk, frequent changes in an ITIL (Version 5) environment?

    Consider how automation and peer-level checks replace traditional, slow approval gates.

    The use of automated peer reviews and integrated continuous delivery pipelines.

    Decentralizing authority through automation and peer review allows changes to flow quickly while maintaining rigorous quality standards.

    • Weekly review by a centralized Change Advisory Board (CAB) to ensure enterprise-wide alignment.

      Centralized boards often act as approval gates that slow down high-velocity teams without adding significant clarity.

    • Manual sign-off by a designated Product Owner for every individual code commit.

      Requiring manual intervention for every commit creates a bottleneck that contradicts the principles of high-velocity delivery.

    • A 'freeze-first' policy where changes are only permitted during designated maintenance windows.

      Change freezes are traditional stability mechanisms that high-velocity organizations seek to replace with continuous resilience.

  6. 6 An SRE team defines an Availability SLO of $99.9\%$. If the service is down for $60$ minutes in a $30$-day period, how does this affect their error budget?

    Calculate the total allowable downtime for $99.9\%$ over a month ($30 \times 24 \times 60 \times 0.001$).

    The error budget is completely exhausted and likely in a deficit, requiring a halt on new feature releases.

    Exceeding the allowable downtime defined by the SLO consumes the entire error budget, triggering a shift in focus toward stability.

    • The error budget is unaffected because $60$ minutes is within the allowable downtime for $99.9\%$ availability.

      At $99.9\%$, the total allowable downtime for $30$ days is approximately $43.2$ minutes; $60$ minutes exceeds this.

    • The error budget is increased automatically by the AIOps system to compensate for the unexpected volatility.

      Error budgets are fixed targets based on business requirements; they do not expand simply because a failure occurred.

    • The incident is excluded from the error budget calculation if it was caused by a third-party vendor.

      SLOs measure the user experience, which is impacted regardless of whether the root cause was internal or external.

  7. 7 In a High Velocity IT environment, what is the primary role of an 'Incident Commander' during a major service disruption?

    Think of this role as the conductor of an orchestra rather than a musician.

    To lead the response, set priorities, and coordinate the efforts of specialized technical teams.

    Incident Command provides a clear role-based structure to manage resources and decision-making during high-pressure events.

    • To personally perform the technical troubleshooting and root cause analysis.

      The commander should lead and orchestrate rather than performing the hands-on technical work.

    • To handle all external communications with customers and the media.

      Communications are typically handled by a Liaison or specialized role, allowing the commander to focus on incident strategy.

    • To document every action taken during the incident for later use in a 'blame-filled' report.

      While documentation is important, its purpose in high velocity IT is learning, not assigning blame.

  8. 8 Why does ITIL (Version 5) advocate for 'Team Topologies' that favor cross-functional product teams over functional silos?

    Consider the impact of organizational structure on the speed of value delivery.

    To reduce the number of handoffs and align teams with the end-to-end product lifecycle.

    Cross-functional teams have the collective skills to move a product from 'Discover' to 'Support' without external dependencies.

    • To ensure that all employees have identical skill sets to maximize resource fungibility.

      The goal is diverse, complementary skills within a team, not identical skills across the entire workforce.

    • To centralize specialized knowledge within 'centers of excellence' to improve technical depth.

      Centralizing knowledge in silos often creates the very bottlenecks that cross-functional teams are designed to eliminate.

    • To lower personnel costs by replacing senior specialists with generalist 'full-stack' developers.

      Team topologies are about optimizing flow and value delivery, not necessarily reducing headcounts or specialist roles.

  9. 9 Which term describes the practice of understanding the internal state of a system solely from the data it provides externally, such as logs, metrics, and traces?

    This concept is essential for managing complex, distributed systems where traditional monitoring is insufficient.

    Observability

    Observability allows teams to diagnose 'unknown unknowns' by analyzing the outputs and telemetry emitted by the system.

    • Monitoring

      Monitoring generally refers to checking for known 'failure' states, whereas this term refers to broader system understanding.

    • Introspection

      In computing, introspection often refers to a program examining its own type or properties, which is narrower than systemic observability.

    • Compliance Auditing

      Auditing is a governance activity focused on adherence to rules, not the real-time technical diagnosis of system health.

  10. 10 According to the ITIL (Version 5) Master Breakdown, how does 'Shift-Left' impact the workforce culture in high-performing teams?

    Think about the timing of testing and operational planning within the software lifecycle.

    It encourages operational and security considerations to be integrated earlier into the development process.

    Shifting left ensures that non-functional requirements are addressed during design and build, reducing late-stage failures.

    • It requires all developers to move their desks to the left side of the office to improve airflow.

      This is a literal and incorrect interpretation of a metaphorical term used for process optimization.

    • It mandates that all senior leadership roles be filled by former developers to ensure technical alignment.

      While technical leadership is valuable, 'shift-left' refers to the timing of activities in the lifecycle, not hiring policies.

    • It focuses on moving all 'low-value' customers to self-service portals to save support costs.

      This describes a support strategy, but 'shift-left' in HVIT primarily refers to moving testing, security, and ops 'earlier' in time.

  11. 11 An organization is using 'Chaos Engineering' as part of its High Velocity IT strategy. What is the ultimate goal of this practice?

    Focus on the relationship between controlled experimentation and systemic resilience.

    To build confidence in the system's ability to withstand turbulent and unexpected conditions in production.

    Chaos engineering uses controlled experiments to identify weaknesses and verify that resilience mechanisms work as intended.

    • To intentionally disrupt services during peak hours to test the patience of the customer base.

      Intentionally causing outages without a controlled experimental framework is reckless and not the goal of chaos engineering.

    • To identify which specific engineers are making the most errors in the codebase.

      Chaos engineering tests the system's architectural resilience, not individual performance or culpability.

    • To replace traditional testing phases with random automated failures to save time.

      Chaos engineering complements traditional testing; it does not replace it and requires a stable baseline to be effective.

  12. 12 Which of the following would be considered 'Toil' according to ITIL 5 principles?

    Identify the task that is repetitive, provides no long-term value, and could be automated.

    Manually resetting the same service account password three times a day because the automation tool is broken.

    This task is manual, repetitive, tactical, and could be eliminated through a permanent fix, making it classic toil.

    • Designing a new automated deployment pipeline for a cloud-native application.

      Designing an automated pipeline is a project-based activity with enduring value, which is the opposite of toil.

    • Facilitating a strategic workshop to define the product vision for the next fiscal year.

      Strategic workshops require high human judgment and have significant enduring value for the organization.

    • Writing a 'blameless postmortem' report after a major systemic outage.

      While it can be time-consuming, a postmortem report provides enduring value through organizational learning and improvement.

  13. 13 A team implementing High Velocity Culture uses 'Error Budgets' to manage the tension between speed and stability. If a team has consistently exceeded its error budget, what action should they take?

    Think about how a team should pivot its focus when service reliability targets are at risk.

    Prioritize reliability-focused tasks and toil reduction over new feature development until the budget is restored.

    Error budgets provide a data-driven signal to rebalance efforts toward stability when reliability targets are missed.

    • Increase the budget to $100\%$ to allow for more innovation without the fear of failure.

      A $100\%$ error budget would imply no reliability requirement, which is not a sustainable service management practice.

    • Disregard the budget and continue pushing features to meet the aggressive marketing deadline.

      Ignoring an error budget defeats its purpose as a governance mechanism for maintaining service health.

    • Fire the least productive member of the team to signal that performance must improve.

      This is a 'blame-based' reaction that destroys psychological safety and fails to address the systemic causes of instability.

  14. 14 In ITIL (Version 5), 'AIOps' is often used to execute 'Coordination' tasks. Which of the following is an example of this?

    Look for an action that involves active orchestration or cross-platform integration.

    Automatically triggering a failover to a redundant server when a primary server heartbeat is lost.

    Coordination involves orchestrating automated workflows and cross-platform integrations to maintain service availability.

    • Predicting that a hard drive will fail in the next $48$ hours based on vibration patterns.

      This is an example of 'Cognition' (predictive analytics), not the orchestration of an automated response.

    • A chatbot providing a link to the password reset portal for a user.

      This is an example of 'Communication', facilitating an interface between a human and a system.

    • Scanning an infrastructure environment to identify all servers missing a critical security patch.

      This is typically part of 'Curation' or 'Cognition' (filtering/analysis), rather than the coordination of a response.

  15. 15 High-velocity organizations often utilize 'Lean' principles to improve flow. What is 'waste' (Muda) in a service delivery context?

    Consider the core Lean definition of any step that doesn't add value.

    Any activity that does not directly contribute to the co-creation of value for the stakeholder.

    Lean focuses on eliminating non-value-adding steps to streamline flow and increase the speed of delivery.

    • The leftover budget at the end of a fiscal year that was not spent on hardware.

      While this might be 'wasteful' in a financial sense, it is not the definition of 'waste' in Lean process optimization.

    • Technical documentation that is more than six months old.

      Old documentation might be obsolete, but it is only 'waste' if the process of creating/maintaining it doesn't add value.

    • The energy consumed by servers that are running at less than $50\%$ capacity.

      This is environmental or resource waste, but Lean 'Muda' specifically refers to process-level inefficiencies.

  16. 16 The ITIL 5 Managing Professional (MP) syllabus includes 'High Velocity Culture & Toil Reduction' with a $12.5\%$ weighting. Which of these is a key concept in 'High Velocity Culture'?

    Identify the human element that allows for rapid adaptation and learning.

    Psychological safety and trust

    These cultural elements are necessary for teams to experiment, learn from failure, and move quickly with confidence.

    • Command-and-control leadership

      High velocity requires decentralized decision-making and autonomy, which is the opposite of command-and-control.

    • Rigid adherence to $v3$ lifecycle stages

      High velocity culture favors flexibility and integrated lifecycles over the rigid, sequential stages of older frameworks.

    • Annual performance reviews based on ticket volume

      Measuring activity (tickets) rather than outcomes (value/experience) is a traditional management practice that HVIT seeks to evolve.

  17. 17 How does 'Toil Reduction' directly contribute to organizational resilience in ITIL (Version 5)?

    Think about the opportunity cost of having experts perform rote, manual tasks.

    By freeing up technical talent to focus on high-value activities like reliability engineering and architectural improvements.

    Reducing toil allows engineers to invest time in the strategic work that prevents future incidents and builds systemic strength.

    • By ensuring that all employees work exactly $40$ hours per week to prevent burnout.

      Resilience is about system and team capability, not just enforcing specific work hours.

    • By automating every single task so that human intervention is never required.

      Automating everything is often impossible and risky; resilience requires human judgment for novel or complex situations.

    • By shifting all repetitive tasks to low-cost offshore providers.

      Outsourcing toil does not reduce it; it just moves it, and can actually decrease resilience due to increased handoff complexity.

  18. 18 A team holds a postmortem where they create a 'Timeline of Events'. Why is this important in a blameless culture?

    Consider how knowing what someone knew at the time helps explain why they made a specific choice.

    To understand the context and information available to the responders at each point in the incident.

    Understanding the 'local rationality' of responders helps identify systemic flaws in the tools, data, or processes they used.

    • To determine who made the first mistake and calculate the financial impact of their error.

      Blameless cultures avoid using timelines to isolate individuals for punishment or financial liability.

    • To provide evidence for the legal department to sue the vendor responsible for the outage.

      While legal evidence may be a byproduct, the primary cultural goal is internal learning and process improvement.

    • To prove that the incident could have been avoided if the team had followed the manual exactly.

      This 'hindsight bias' ignores the complexity of real-time operations and is counterproductive to systemic learning.

  19. 19 In the ITIL AI Capability Model, 'Curation' is described as filtering and structuring knowledge. How does this help reduce toil?

    Think about how well-organized information helps users help themselves.

    By maintaining accurate, searchable knowledge that allows users to self-solve issues before raising tickets.

    Good curation enables self-service, which prevents repetitive support requests (toil) from reaching the service desk.

    • By automatically writing new feature code so developers don't have to.

      This describes 'Creation' or 'Coordination', not the curation of knowledge bases.

    • By monitoring social media for negative comments about the IT department.

      This is a sentiment analysis task (Cognition) rather than the internal curation of service knowledge.

    • By ensuring that every employee has a physical copy of the ITIL 5 Foundation handbook.

      Distributing physical books is a logistical task, not the digital curation of dynamic operational knowledge.

  20. 20 What is the 'Second Story' in the context of blameless postmortems?

    This term refers to looking past the obvious surface-level cause of a failure.

    The deeper, systemic explanation of why an incident occurred, moving beyond 'human error'.

    The 'First Story' is often 'someone messed up'; the 'Second Story' investigates why the system allowed that mess-up to happen.

    • The official PR statement released to the public that omits the technical details.

      The 'Second Story' is a concept for internal learning, not a sanitized public relations narrative.

    • The backup plan that is executed only if the primary resolution fails.

      This refers to a contingency plan or technical redundancy, not a narrative approach to incident analysis.

    • A fictional scenario used during training to simulate a major service outage.

      While simulations are useful, the 'Second Story' specifically refers to the analysis of real, past incidents.

  21. 21 A high-velocity team uses 'Peer Review' as their primary change authority. Which organizational risk is this approach primarily designed to mitigate?

    Think about the tradeoff between slow, centralized control and fast, distributed expertise.

    The risk of a centralized Change Advisory Board (CAB) becoming a bottleneck that slows down delivery.

    Peer review provides specialized technical oversight at the speed of the development workflow.

    • The risk of employees talking to each other too much and wasting time.

      Collaboration is encouraged in HVIT; peer review is a structured way to leverage that collaboration for quality.

    • The risk of hackers infiltrating the office and changing the server configuration.

      Physical security risks are handled by different controls; peer review is about code quality and operational stability.

    • The risk of the CEO not knowing every single change that happens in the IT environment.

      It is impractical and undesirable for executives to oversee individual technical changes in a high-velocity environment.

  22. 22 In ITIL (Version 5), what is the relationship between 'AIOps' and 'Observability' in a high-velocity environment?

    Consider how one provides the 'eyes' and the other provides the 'brain' for system management.

    Observability provides the telemetry (data) that AIOps uses to detect patterns and automate responses.

    Without the rich data from observability, AI systems lack the context needed for accurate cognition and coordination.

    • AIOps is the brand name of a specific software, while observability is a legal requirement.

      AIOps is a general practice/technology category, and observability is a technical discipline, not typically a law.

    • They are the same thing and the terms can be used interchangeably in the exam.

      Interchanging these terms is incorrect; one is about data/transparency (observability) and the other is about processing/action (AIOps).

    • Observability replaces the need for AIOps by making all systems simple enough for humans to manage.

      Observability often reveals more complexity, which increases the need for AIOps to help humans process the information.

  23. 23 Which of these is a symptom of an organization that has 'automated away accountability'?

    Look for a scenario where humans are 'out of the loop' and unable to explain why things are happening.

    No human understands the logic behind an automated decision, and there is no override mechanism.

    Accountability requires that humans can explain, oversee, and intervene in automated processes when necessary.

    • Teams use automated testing for $100\%$ of their code deployments.

      High levels of automated testing are a best practice and do not inherently remove human accountability for the system.

    • The service desk uses a chatbot to answer common user questions.

      Using a chatbot for simple tasks is a standard efficiency gain and does not absolve the team of their service responsibilities.

    • The organization has a dedicated AI Governance module in their training scheme.

      Training in AI governance is a step toward *increasing* accountability, not removing it.

  24. 24 In the High Velocity IT domain, what is 'Shift-Left' applied to 'Security' often called?

    Think of the common portmanteau for Development, Security, and Operations.

    DevSecOps

    DevSecOps is the specific integration of security practices into the DevOps and continuous delivery lifecycle.

    • SafeMode

      SafeMode is usually a computer diagnostic mode, not an organizational process for shifting security left.

    • SecOpsShift

      While it uses the right words, 'DevSecOps' is the globally recognized industry term for this integration.

    • GuardRails

      Guardrails are the specific policies or automated checks used within a DevSecOps approach, not the approach itself.

  25. 25 A team wants to reduce 'Toil' by $50\%$ over the next six months. What is the first step they should take?

    Consider the importance of data-driven decision-making and establishing a baseline.

    Measure and categorize current activities to identify which tasks actually constitute toil.

    You cannot reduce what you haven't measured; a baseline is essential for identifying the most impactful automation opportunities.

    • Buy a new expensive AIOps tool and hope it solves the problem automatically.

      Tooling without process understanding and data baseline is rarely successful in reducing toil.

    • Cancel all meetings to give everyone more time to do manual work.

      Reducing meetings might increase capacity, but it doesn't change the nature of the work from toil to high-value.

    • Require every team member to work $2$ extra hours a day to 'catch up' on automation tasks.

      This approach leads to burnout and ignores the fact that toil scales linearly with the service; more hours just means more toil.