Spectrum IT Recruitment

Spectrum IT Recruitment

Southampton, South East

Lead Site Reliability Engineer

Full-Time£75,000 - 85,000 per yearпозавчераUnited Kingdom
IT

Job Description

Salary: £75,000 - 85,000 per year

Requirements:
  • At least six years commercial experience in Site Reliability Engineering or a closely related cloud platform role
  • Demonstrable experience supporting business-critical cloud platforms and live production services
  • Strong hands-on knowledge of Microsoft Azure
  • Production experience with Kubernetes and containerised workloads, ideally using Azure Kubernetes Service (AKS)
  • Extensive experience in platform engineering, cloud provisioning and observability
  • Strong monitoring, alerting and dashboarding experience with technologies such as Azure Monitor, Grafana, Prometheus, OpenTelemetry or Elasticsearch
  • Experience creating custom metrics, queries, dashboards and alerts for microservices
  • Advanced scripting or software development skills using PowerShell, Python, C# or a comparable language
  • Strong Infrastructure as Code experience using Bicep, ARM or Terraform
  • Experience using Git or another version-control platform
  • Good knowledge of Microsoft SQL Server, Elasticsearch and structured data formats, including YAML, JSON and XML
  • Strong understanding of microservices architecture, cloud platforms and containerisation
  • Experience defining or working with SLOs, SLAs, SLIs and error budgets
  • Excellent troubleshooting and root-cause analysis skills
  • Experience designing scalable, secure and maintainable cloud solutions
  • Strong understanding of cybersecurity principles, governance and compliance
  • Experience working across transformation projects and live-service environments
  • Applicants must have lived continuously in the UK for the past five years and be eligible to obtain NPPV3 and UK Security Clearance
  • Desirable: Azure DevOps pipeline experience covering CI/CD and automated deployment
  • Desirable: Experience developing reusable infrastructure and monitoring modules
  • Desirable: Familiarity with AI-enabled engineering and automation tools
  • Desirable: Knowledge of security and compliance frameworks such as ISO 27001, Cyber Essentials Plus or FedRAMP
  • Desirable: Experience providing technical leadership across multidisciplinary cloud, engineering and support teams
  • Desirable: A background in regulated, public-sector or security-sensitive environments
Responsibilities:
  • Work as part of the Site Reliability Engineering team to protect and improve production environments
  • Manage and prioritise a technical backlog of reliability, scalability and operational improvements
  • Lead investigations into service outages, performance degradation, platform reliability and cloud expenditure
  • Conduct root-cause analysis and ensure corrective actions are implemented
  • Identify repetitive operational activities and replace them with sustainable automation
  • Provide technical leadership and guidance to Cloud Operations, Support, DevOps and Engineering teams
  • Establish and maintain service level objectives, service level agreements, service level indicators and error budgets
  • Design and implement monitoring, alerting and dashboards across cloud platforms and microservices
  • Deploy and configure observability technologies, including Grafana, Prometheus, Azure Monitor and OpenTelemetry
  • Develop custom application and platform metrics to improve operational visibility
  • Create advanced queries, dashboards and alerts for distributed microservices
  • Develop reusable Bicep or Terraform modules for monitoring and cloud infrastructure
  • Support and improve production Kubernetes environments, particularly Azure Kubernetes Service
  • Review and optimise platform performance, availability, security and cost
  • Contribute to cloud architecture, technical scoping and implementation of scalable platform solutions
  • Support continuous improvement across deployment, provisioning and operational processes
  • Help ensure platforms and working practices meet relevant security, governance and compliance requirements
  • Explore opportunities to use AI-assisted tools to improve automation, troubleshooting and engineering productivity
Technologies:
  • AI
  • ARM
  • Azure
  • C#
  • CI/CD
  • Cloud
  • DevOps
  • ElasticSearch
  • Git
  • Grafana
  • Support
  • JSON
  • Kubernetes
  • OpenTelemetry
  • PowerShell
  • Prometheus
  • Python
  • SQL
  • Security
  • Terraform
  • XML
  • microservices

More:

We provide advanced SaaS solutions to organisations in the public safety and justice sectors, supporting multimedia evidence management and emergency contact centre operations for customers worldwide. As our Cloud Platform Engineering function expands, we are seeking a Senior Site Reliability Engineer for a highly hands-on role focused on keeping business-critical cloud platforms observable, measurable, secure, scalable and reliable. The role is based in Southampton with hybrid working. We welcome candidates from Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering or Cloud Development backgrounds. Spectrum IT Recruitment (South) Limited is acting as an Employment Agency in relation to this vacancy.

last updated 40 week of 2026

Interested in this role?

Submit your application now

How to Apply

Ready to apply for this position? Here's what you need:

  • An updated resume highlighting relevant experience
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

About Spectrum IT Recruitment

Spectrum IT Recruitment

Spectrum IT Recruitment

Southampton

IT

Skills & Technologies

PythonC#GoScalaAzureKubernetesTerraformGitCI/CDElasticsearchAIDevOps

Inferred from job description

Salary Insight

£80,000

This role

£85,000

UK median

This salary is 6% below the UK median for Lead roles (£85,000/yr).

Based on 2024–2025 UK technology sector benchmarks

Explore More UK Opportunities

Thousands of tech jobs across the United Kingdom