Director of Site Reliability Engineering

Company: EPAM Systems
Apply for the Director of Site Reliability Engineering
Location: London
Job Description:

We’re looking for a Director of Site Reliability Engineering to join our team in London, United Kingdom in a hybrid working mode.

This role is responsible for driving reliability engineering and operational excellence across global technology platforms while leading the adoption of AI-enabled solutions for automation and efficiency. The position combines strategic leadership with hands‑on governance to ensure highly available, resilient systems that align with business and regulatory requirements. As a technology thought leader, you will influence engineering standards, enhance operational frameworks, and foster a culture of continuous improvement across mission‑critical environments.

Responsibilities

  • Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment
  • Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices
  • Define and monitor KPIs for system reliability, performance, and operational efficiency
  • Advance automation, Infrastructure as Code approaches, and promote self‑healing systems using AI/ML techniques
  • Develop robust incident management frameworks and lead major incident response activities for critical systems
  • Implement blameless postmortems and deliver systemic improvements across production environments
  • Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems
  • Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services
  • Drive resilience strategies with highly available architectures and disaster recovery readiness
  • Champion an automation‑first culture, leveraging CI/CD pipelines and operational tooling to reduce manual processes

Requirements

  • Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments
  • Expertise in observability platforms, troubleshooting distributed systems, and telemetry‑driven insights
  • Hands‑on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices
  • Deep understanding of incident management processes, ITSM standards, and ITIL principles
  • Knowledge of resilience design patterns, high availability, and fault‑tolerant architectures
  • Familiarity with AI/ML-driven approaches for operational efficiency and system reliability
  • Ability to lead transformation, influence across teams, and foster continuous improvement in culture

Nice to have

  • Experience in financial services or other highly regulated, mission‑critical environments
  • Certifications in cloud technologies, such as AWS
  • Exposure to AIOps platforms or advanced observability tooling

We offer

  • EPAM Employee Stock Purchase Plan (ESPP)
  • Protection benefits including life assurance, income protection and critical illness cover
  • Private medical insurance and dental care
  • Employee Assistance Program
  • Cyclescheme, Techscheme and season ticket loans
  • Various perks such as free Wednesday lunch in-office, on‑site massages and regular social events
  • Learning and development opportunities including in‑house training and coaching, professional certifications, and courses
  • If otherwise eligible, participation in the discretionary annual bonus program
  • If otherwise eligible and hired into a qualifying level, participation in the discretionary Long‑Term Incentive (LTI) Program

#J-18808-Ljbffr…

Posted: July 19th, 2026