All writing

ITIL through an engineer's eyes

ITIL is often mistaken for ticket queues, change boards and process bureaucracy. Through an engineer's eyes, its more useful ideas concern services, value streams, operating responsibilities, incident learning, controlled change and continual improvement.

about 19 minutes min read

Mention ITIL to a room of engineers and several images may appear.

A queue containing twelve thousand tickets.

A change form asking whether the implementation has a rollback plan, followed by a box too small to describe one.

A weekly meeting where somebody reads deployment titles from a spreadsheet.

A configuration-management database containing an exquisite catalogue of servers that stopped existing three years ago.

The framework is then summarised as:

ITIL

the department of forms
that stands between code
and production

That reputation did not appear from nowhere.

Organisations have implemented service management through:

  • rigid process gates
  • centralised approval
  • excessive hand-offs
  • disconnected support queues
  • stale documentation
  • activity metrics based on ticket closure
  • tools treated as the operating model

But a poor implementation of a framework is not the complete meaning of the framework.

ITIL 4 describes a Service Value System, guiding principles, four dimensions of service management, a service value chain, management practices and continual improvement. Its official material presents ITIL as a flexible approach to creating, delivering and improving technology-enabled services rather than one mandatory organisational process.

Through an engineer’s eyes, the useful question is not:

How do we make engineering
obey ITIL?

It is:

Which service-management ideas help us
build, operate and improve
dependable technology?

ITIL becomes useful when it helps engineers see the service around the software: the consumers, outcomes, dependencies, operating responsibilities, failure paths and improvement work that continue after deployment.

An engineer’s service operating model

A service is larger than the thing we deployed.

A bfstore customer does not ask for:

order-service Pod availability

They expect to:

place an order

receive confirmation

track its progress

obtain support when something fails

A developer does not primarily want:

a Crossplane composite resource

They want:

a usable database
with approved access,
backups, status and support

The engineering system and the service outcome are connected.

COMPONENTS

code
infrastructure
identity
data
monitoring
documentation
support
        │
        ▼
SERVICE

a useful outcome
delivered under understood conditions

A practical operating model asks:

CONSUMER

Who depends on the service?


OUTCOME

Which useful result do they need?


PROMISE

Under which conditions
should the outcome be available?


VALUE STREAM

How does demand become
a usable result?


OPERATING CAPABILITIES

How do we change, support,
restore and secure it?


EVIDENCE

How do we know
the service is working?


IMPROVEMENT

What should become better next?

This wider view does not diminish engineering.

It reveals more of the system engineering is responsible for.

ITIL begins with services, value and flow

Value is the point

A process can run exactly as documented and still produce little value.

A deployment may collect five approvals and remain unsafe.

A service desk may close tickets quickly while users repeatedly experience the same failure.

A platform team may provision infrastructure rapidly while developers cannot understand its status.

ITIL 4 places value creation and co-creation at the centre of service management. The provider contributes capability, but the consumer experiences value through use, surrounding business activity and shared information. Official ITIL material describes the framework as flexible and focused on value rather than process compliance.

Weak measure:

Changes processed:
    480

Stronger measures:

Successful change rate:
    97%

Changes causing incidents:
    2%

Median recovery time:
    18 minutes

Checkout availability:
    99.95%

For an internal platform:

Valid database requests
completed within SLO:
    98.7%

Requests requiring
manual intervention:
    4.1%

Restore tests
meeting RTO:
    82%

Activity is evidence that work occurred.

Value asks whether the work improved an outcome somebody cares about.

The Service Value System connects the operating parts

The ITIL Service Value System brings together:

  • guiding principles
  • governance
  • the service value chain
  • practices
  • continual improvement

It provides an operating-model lens rather than a replacement for every product, engineering or business model.

An engineer can translate it like this:

DEMAND OR OPPORTUNITY
        │
        ▼
GOVERNANCE AND PRINCIPLES
        │
        ▼
VALUE STREAM
        │
        ├── people
        ├── practices
        ├── technology
        └── suppliers
        │
        ▼
SERVICE OUTCOME
        │
        ▼
FEEDBACK AND IMPROVEMENT

Suppose bfstore needs customer reviews.

The service may require:

  • product decisions
  • review-service development
  • moderation rules
  • database provisioning
  • authentication
  • deployment
  • monitoring
  • support
  • incident response
  • feedback

Optimising one part may damage the whole.

BUILD FASTER
    │
    ▼
deployments increase
    │
    ▼
operational evidence missing
    │
    ▼
incidents become harder to diagnose

Local speed has increased.

The value stream may have become slower.

The guiding principles resist bureaucracy

ITIL 4 includes seven guiding principles:

  1. focus on value
  2. start where you are
  3. progress iteratively with feedback
  4. collaborate and promote visibility
  5. think and work holistically
  6. keep it simple and practical
  7. optimise and automate

Through an engineering lens, they become:

FOCUS ON VALUE

Do not automate a workflow
merely because it exists.


START WHERE YOU ARE

Measure and understand
the current system first.


PROGRESS ITERATIVELY

Change a smaller blast radius,
observe and adjust.


PROMOTE VISIBILITY

Make work, risk,
ownership and outcomes visible.


WORK HOLISTICALLY

Trace failures across technical
and organisational boundaries.


KEEP IT PRACTICAL

Remove controls and metrics
that do not improve an outcome.


OPTIMISE AND AUTOMATE

Understand the work,
remove waste, then automate it.

Automating a broken process creates a high-throughput broken process.

Taken seriously, these principles resist bureaucracy rather than creating it.

The four dimensions prevent technology tunnel vision

ITIL 4 describes four dimensions:

  • organisations and people
  • information and technology
  • partners and suppliers
  • value streams and processes

A service may fail through any of them.

ORGANISATIONS AND PEOPLE

ownership
skills
authority
support


INFORMATION AND TECHNOLOGY

applications
infrastructure
data
telemetry
security


PARTNERS AND SUPPLIERS

cloud providers
identity services
payment providers
registries


VALUE STREAMS AND PROCESSES

queues
hand-offs
decisions
automation

An external dependency does not stop being an engineering concern because another company operates it.

Value streams combine activities and practices

The ITIL Service Value Chain includes:

  • plan
  • improve
  • engage
  • design and transition
  • obtain or build
  • deliver and support

These activities can be combined into different value streams rather than treated as one mandatory sequence.

A feature stream may move from customer need to delivery and feedback.

An incident stream may begin with support and restoration, then move into repair and improvement.

The activities are ingredients.

The value stream selects and combines them according to the outcome.

Practices are operating capabilities

ITIL 4 describes a practice as a set of organisational resources designed to perform work or accomplish an objective.

A practice may include:

  • people
  • responsibilities
  • workflows
  • information
  • tools
  • suppliers
  • measurements
  • knowledge

The tool may support the practice.

It is not the whole practice.

The most useful engineering translation is to group practices by the operating problem they solve.

Flow and change

Service requests become platform self-service

A service request is normal demand for an agreed service action.

Examples include:

  • request a development database
  • obtain repository access
  • create a test environment
  • rotate an application credential
  • retrieve approved software

This is different from an incident.

SERVICE REQUEST

normal demand for
an agreed service action


INCIDENT

unplanned interruption
or degradation of a service

For an internal platform:

DEVELOPER REQUEST
       │
       ▼
PLATFORM API
       │
       ▼
POLICY AND DEFAULTS
       │
       ▼
AUTOMATED FULFILMENT
       │
       ▼
VISIBLE STATUS

The ticket may disappear entirely for ordinary cases.

The service-management concern remains: discoverable capability, clear intent, consistent fulfilment and visible status.

Change enablement becomes risk-based delivery

ITIL 4 change enablement focuses on increasing successful change while assessing risk, authorising appropriately and managing schedules. PeopleCert’s practice material does not require one committee to approve every change.

LOW-RISK REPEATABLE CHANGE

reviewed source
automated tests
policy checks
progressive delivery
automatic evidence


HIGH-RISK CONTEXTUAL CHANGE

explicit review
decision authority
timing
recovery evidence

Change authority should match the change.

It may belong to:

  • an automated delivery controller
  • the service owner
  • a platform team
  • a security owner
  • an incident commander

ITIL does not become modern merely because an approval form is submitted through an API.

The practice becomes effective when control is proportionate, evidence is useful and delivery remains safe.

Deployment and release solve different problems

Deployment moves components into an environment.

Release makes functionality available for use.

DEPLOYMENT

new version in production
        │
        ▼
FEATURE FLAG OFF
        │
        ▼
not released

Later:

FEATURE FLAG ON
        │
        ▼
capability released

This distinction supports canaries, dark launches and controlled exposure without requiring one oversized release ticket.

Restore and learn

Observability helps engineers understand system state.

Monitoring and event management help decide which observed changes require which response.

SIGNAL
  │
  ▼
INTERPRETATION
  │
  ├── informational
  ├── warning
  ├── automated response
  └── incident

An event pipeline without ownership and response rules becomes an alert factory.

Incident management restores useful service

During an incident, the immediate objective is often:

restore useful service

not:

complete the root-cause analysis
before acting

Suppose checkout fails after a release.

rollback release
      │
      ▼
restore checkout
      │
      ▼
investigate safely

Incident management may include:

  • detection
  • triage
  • impact assessment
  • coordination
  • mitigation
  • communication
  • restoration verification

The service outcome determines restoration.

A Pod restart is not enough when customers still cannot place orders.

Problem management reduces recurrence

Problem management asks:

Why did this happen?

How might it happen again?

What should change?

It may use:

  • incident reviews
  • trend analysis
  • defect investigation
  • workarounds
  • known errors
  • resilience testing
  • reliability backlog
INCIDENT RESTORATION

return useful service now


PROBLEM MANAGEMENT

understand and reduce recurrence

Combining both into one frantic activity can produce weak restoration and weak learning.

Continuity prepares for serious disruption

Continuity connects business tolerance to tested engineering capability.

BACKUP
   │
   ▼
RESTORE
   │
   ▼
SERVICE VERIFICATION
   │
   ▼
RECOVERY OBJECTIVE

Useful distinctions are:

INCIDENT RESTORATION

return useful service now


RECOVERY

recreate service and state
after disruption


CONTINUITY

prepare to maintain or recover
agreed service under serious disruption

Restoring a database is not enough when identity, DNS, artefacts or operator access remain unavailable.

Understand and govern the service

Configuration management means trustworthy relationships

The useful objective is not to inventory everything that can receive a serial number.

It is to maintain the information required to manage services effectively.

ORDER SERVICE
    │
    ├── deployed version
    ├── database
    ├── Kafka topics
    ├── workload identity
    ├── cluster
    ├── owner
    └── runbook

Modern sources may include:

  • Git
  • cloud APIs
  • Kubernetes
  • service catalogues
  • infrastructure state
  • telemetry

A federated model is often stronger:

Git:
    desired state

cloud API:
    native resource state

Kubernetes:
    workload status

service catalogue:
    ownership and relationships

telemetry:
    runtime behaviour

Configuration management should connect those truths rather than manually retyping them into one polished mausoleum.

Freshness matters.

A populated ownership field, runbook link or recovery record is not automatically trustworthy.

Service levels connect users and engineers

Service level management helps establish business-based targets for service utility, warranty and experience. PeopleCert’s service-level material connects naturally to SLIs, SLOs and stakeholder expectations.

UTILITY

does the service provide
the needed capability?


WARRANTY

is it available, secure,
continuous and sufficiently capable?

A useful conversation asks:

Which user journey matters?

How do we measure it?

How much failure is tolerable?

What happens when
reliability falls?

A target without operational consequences is decorative mathematics.

Availability and capacity follow the service promise

Availability should measure the useful capability.

COMPONENT STATUS

Pod ready


SERVICE STATUS

customer can place order

Capacity includes more than CPU and memory.

A service can run out of:

  • database connections
  • Kafka partitions
  • IP addresses
  • API quota
  • build runners
  • telemetry capacity
  • operator attention

The component signals explain the constraint.

The service signal measures impact.

Security is part of service quality

A dependable service must protect:

  • identities
  • data
  • integrity
  • availability
  • evidence

Security should become:

  • platform defaults
  • delivery guardrails
  • service design
  • observable controls
  • recovery procedures

It should not remain a separate ticket queue attached at the end.

Suppliers remain inside the service architecture

Cloud providers, payment systems, identity services and registries operate parts of the machinery.

The organisation still owns the outcome delivered to its consumers.

Supplier questions include:

  • What does the provider promise?
  • Which limits apply?
  • How are incidents communicated?
  • How can data be exported?
  • Which recovery responsibilities remain ours?
  • Which alternative exists?

The phrase managed service does not mean the supplier owns every consequence.

Continual improvement closes the loop

Engineers already practise continual improvement through:

  • retrospectives
  • incident actions
  • refactoring
  • reliability work
  • platform feedback
  • cost optimisation
  • automation

The useful loop is:

WHERE ARE WE NOW?
        │
        ▼
WHERE DO WE WANT TO BE?
        │
        ▼
WHICH CHANGE WILL HELP?
        │
        ▼
IMPLEMENT
        │
        ▼
MEASURE
        │
        ▼
LEARN AND REPEAT

Improvement should remain connected to evidence.

Suppose database provisioning changes from forty minutes to twelve.

If manual intervention remains unchanged, the next improvement should target the hand-off rather than declaring victory.

Improvement is not a project with a finish banner.

It is the habit of comparing current outcomes with desired ones and changing the system deliberately.

Modern engineering provides the implementation mechanisms

ITIL, DevOps, SRE and platform engineering operate at different levels.

They do not need to compete.

ITIL

service-management frame


DEVOPS

delivery flow and shared responsibility


SRE

reliability methods and engineering discipline


PLATFORM ENGINEERING

internal services delivered as products

For example:

Service-management concern Engineering expression
Change enablement CI/CD controls, risk-based gates, progressive delivery
Incident management On-call response, incident command, restoration
Problem management Incident review, reliability backlog, resilience testing
Monitoring and event management Telemetry, alerts, event routing
Service request management Platform self-service and automated fulfilment
Configuration management Git, cloud state, catalogues, dependency metadata
Service level management SLIs, SLOs, error budgets
Continual improvement Retrospectives, experiments, measured backlog work

The framework and engineering methods can support each other.

A user does not need to say:

I would like to invoke
service request management.

They select Create database.

The practice exists behind the interface.

A selective bfstore service model

bfstore does not need every practice at maximum maturity.

It needs a small number of clearly defined services.

Examples include:

Customer checkout

Application deployment

Transactional database

Messaging

Workload identity

Observability

Backup and recovery

Each service should identify:

  • consumer
  • accountable owner
  • outcome
  • dependencies
  • service level
  • support path
  • recovery expectation
  • decision rights
  • evidence freshness

A lightweight service record might contain:

service: transactional-database
owner: platform-data

consumers:
  - catalog-service
  - order-service
  - review-service

service_level:
  provisioning: 99% within 15 minutes

recovery:
  rpo: 15 minutes
  rto: 60 minutes
  last_demonstrated: 2025-10-12

decision_rights:
  routine_change: automated-path
  high_risk_change: platform-data-owner
  incident_authority: incident-commander

support:
  runbook: /runbooks/database
  runbook_last_exercised: 2025-09-18

ownership_verified: 2025-11-30

The service owner is accountable for the outcome.

Component and supplier owners remain accountable for contributing capabilities.

The service owner needs coordination authority and escalation paths rather than direct control of every dependency.

Routine requests should be automated.

High-risk changes should use explicit authority.

Incidents should identify the affected service, user impact, restoration action and follow-up problem work.

SLOs should measure useful journeys.

Improvement should use evidence from:

  • incidents
  • failed changes
  • support requests
  • platform feedback
  • cost trends
  • restore drills
  • SLO breaches

This is enough to create a useful operating model.

It does not require bfstore to construct a miniature Ministry of Ticket Affairs.

How bureaucracy takes over

Service management becomes harmful when:

The tool becomes the framework

An ITSM platform is installed and assumed to be the operating model.

Every activity becomes a ticket

Deployments, automation events and ordinary requests enter one manual queue.

Approval replaces evidence

A person clicks Approve, but test results, risk, rollback and service health remain unclear.

Practice ownership becomes silo ownership

Separate teams optimise change, incidents and configuration while nobody improves the end-to-end service.

Metrics reward queue movement

Teams are measured on tickets closed rather than recurrence, waiting time, service impact and successful outcomes.

The framework is implemented wholesale

Every practice receives a process before the organisation understands which problems need solving.

The guiding principles suggest the opposite:

focus on value

start where you are

progress iteratively

keep it simple

A good ITIL adoption should be recognisable by those principles, not contradicted by them.

Questions I now ask

Service

Which service outcome
and consumer are we responsible for?

Dimensions

Which people, technology,
suppliers and workflows enable it?

Flow

How does demand move
through the value stream?

Operations

Which capabilities support
change, restoration and recovery?

Service level

Which user-centred measure
defines an acceptable outcome?

Information

Which relationships and records
must remain trustworthy?

Authority

Who owns decisions,
escalation and improvement?

Evidence

What tells us
what to improve next?

These questions feel more useful than asking whether the organisation has completed an ITIL implementation.

The mental model I am keeping

My earlier model was:

ITIL
  │
  ├── tickets
  ├── approvals
  ├── processes
  └── service desk

The stronger model is:

                         CONSUMER NEED
                               │
                               ▼
                        SERVICE OUTCOME
                               │
                               ▼
                         VALUE STREAM
                               │
              ┌────────────────┼────────────────┐
              │                │                │
              ▼                ▼                ▼
            PEOPLE         TECHNOLOGY       SUPPLIERS
              │                │                │
              └────────────────┼────────────────┘
                               ▼
                    OPERATING CAPABILITIES
                               │
              ┌────────────────┼────────────────┐
              │                │                │
              ▼                ▼                ▼
            CHANGE          RESTORE          SUPPORT
              │                │                │
              └────────────────┼────────────────┘
                               ▼
                            EVIDENCE
                               │
                               ▼
                         IMPROVEMENT

ITIL does not need to become an engineer’s primary identity.

I do not need to look at a failing Kubernetes workload and announce that I am activating a service-management practice.

But the framework offers a useful reminder.

The software is part of a service.

The service has consumers.

Consumers experience outcomes, not architecture diagrams.

Those outcomes depend on people, technology, suppliers and value streams working together.

Through an engineer’s eyes, ITIL is most useful not as a collection of processes to obey, but as a service operating model connecting consumer outcomes to flow, operating responsibility, restoration and improvement.

Its ideas become practical when they improve:

  • flow
  • ownership
  • reliability
  • communication
  • recovery
  • learning

They become harmful when the framework is used to justify hand-offs, delay and ceremonial control.

The difference is not whether an organisation says it uses ITIL.

The difference is whether service management helps technology create dependable value or merely creates more work around the work.

References and further reading

  1. PeopleCert: ITIL 4 Foundation
    Official overview of ITIL 4 concepts, including the Service Value System, guiding principles, four dimensions, service value chain, practices and continual improvement.

  2. Axelos and ISACA: Using ITIL 4 and COBIT 2019 for an integrated I&T framework
    Official paper covering the Service Value System, value chain, governance, continual improvement and the four dimensions.

  3. PeopleCert: ITIL 4 Practitioner, Incident Management
    Official practice overview focused on minimising incident impact and restoring normal service.

  4. PeopleCert: ITIL 4 Practitioner, Problem Management
    Official practice overview covering incident causes, workarounds, known errors and reduction of recurrence.

  5. PeopleCert: ITIL 4 Practitioner, Change Enablement
    Official practice overview covering risk assessment, authorisation and scheduling.

  6. PeopleCert: ITIL 4 Practitioner, Service Level Management
    Official overview of business-based service targets for utility, warranty and experience.

  7. ISO/IEC 20000-1:2018
    International requirements for establishing, implementing, maintaining and continually improving a service-management system.

  8. ISO/IEC TS 20000-11:2021
    ISO guidance on the relationship between ISO/IEC 20000-1 and ITIL 4.