JP Morgan Chase
Senior Lead Site Reliability / DevOps Engineer
Job Description
Salary: £62,000 - 102,000 per year
Requirements:- We require formal training or certification in software engineering concepts, along with advanced applied experience delivering system design, application development, testing, and operational stability.
- We require advanced knowledge of reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with in-depth expertise in one or more technical disciplines such as cloud, observability, or distributed systems.
- We require advanced proficiency in one or more programming languages such as Java, Python, or Go.
- We require advanced proficiency and experience in observability, including white-box and black-box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, and similar platforms.
- We require proficiency in continuous integration and continuous delivery tools such as Jenkins, GitLab, and Terraform.
- We require experience with containers and container orchestration technologies such as ECS, Kubernetes, and Docker.
- We require hands-on experience designing, deploying, and operating OpenTelemetry collectors in production, including configuring, optimizing, and troubleshooting OTLP endpoints and receivers.
- We require the ability to solve reliability design and functionality problems independently with little to no oversight.
- We require practical cloud-native experience.
- We require the ability to collaborate effectively across different levels and stakeholder groups.
- Preferred: knowledge of distributed tracing, metrics, and logging best practices.
- Preferred: certification in AWS, Kubernetes, or relevant technologies.
- Preferred: proven track record in system health monitoring, capacity management, and blameless postmortems for high-availability services.
- Preferred: deep understanding of distributed system design principles, networking concepts such as TCP/IP, DNS, and load balancing, and Linux internals.
- Preferred: contributions to open-source observability or telemetry projects.
- Preferred: experience with agent control planes and management protocols, with hands-on knowledge of OpAMP highly desirable.
- We provide technical guidance and direction on site reliability practices to support our business, technical teams, contractors, and vendors.
- We develop secure, high-quality production code for reliability tooling and telemetry pipelines, and review and debug code written by others.
- We drive decisions that influence reliability design, observability architecture, application functionality, and technical operations and processes.
- We serve as a subject matter expert in one or more areas of site reliability, observability, or telemetry engineering.
- We lead resiliency design reviews and break complex reliability problems into digestible work for other engineers, acting as a technical lead for large products.
- We act as the main point of contact during major incidents, identify and solve issues quickly to avoid financial losses, and champion a blameless postmortem culture.
- We collaborate with team members and stakeholders to define service level indicators, service level objectives, and error budgets.
- We design, implement, and maintain operational reliability for large-scale OpenTelemetry pipelines in hybrid on-prem and cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearch.
- We drive the assessment, refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability.
- We actively contribute to the engineering community as advocates of firmwide frameworks, tools, and practices, and influence peers and project decision-makers to adopt leading-edge observability and reliability technologies.
- We contribute to our culture of diversity, opportunity, inclusion, and respect.
- AWS
- OpenSearch
- Cloud
- Datadog
- Docker
- Dynatrace
- ElasticSearch
- GitLab
- Grafana
- Support
- Java
- Jenkins
- Kubernetes
- Linux
- Load Balancing
- OpenTelemetry
- Prometheus
- Python
- Security
- Splunk
- TCP/IP
- Terraform
- DevOps
More:
We are J.P. Morgan Chase, a global leader in financial services and a leader across banking, markets, securities services, and payments through our Commercial & Investment Bank. We provide strategic advice and products to major corporations, governments, wealthy individuals, and institutional investors in more than 100 countries. Our first-class business in a first-class way approach drives everything we do, and we build trusted, long-term partnerships to help clients achieve their business objectives. We value the diverse talents of our people, are committed to equal opportunity, and foster a culture of diversity, inclusion, and respect. This is a full-time role.
last updated 34 week of 2026
Interested in this role?
Submit your application now
How to Apply
About JP Morgan Chase
JP Morgan Chase
Glasgow
IT
Skills & Technologies
Inferred from job description
Salary Insight
£82,000
This role
£75,000
UK median
This salary is 9% above the UK median for Senior roles (£75,000/yr).
Based on 2024–2025 UK technology sector benchmarks