JP Morgan Chase

JP Morgan Chase

Glasgow, Scotland

Senior Lead Software Engineer - LLM Ops Platform

Full-Time£62,000 - 102,000 per yearanteayerUnited Kingdom
IT

Job Description

Salary: £62,000 - 102,000 per year

Requirements:
  • We have hands-on experience with system design, application development, testing, and operational stability in production environments.
  • We have advanced proficiency in Python for building production-grade services and tooling.
  • We have proficiency with automation and continuous delivery methods.
  • We have hands-on experience with cloud infrastructure platforms and infrastructure-as-code tooling for delivery and lifecycle management.
  • We have a strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns.
  • We have practical knowledge of observability and instrumentation across metrics, logs, and traces.
  • We have hands-on experience with Kubernetes and container-based orchestration platforms, including managed cloud variants.
  • We have experience hosting and serving large language models on cloud-based infrastructure and local GPU environments.
  • We have knowledge of large language model reliability and risk considerations, including latency and throughput trade-offs, model versioning, prompt and response logging, and safe rollout patterns.
  • We have hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment, with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
  • We understand responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs and outputs, and adherence to resiliency and security expectations.
  • We have the ability to guide peers on safe and effective usage within team practices.
Responsibilities:
  • We design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure.
  • We build backend services and APIs that enable reliable operation of AI infrastructure in production environments.
  • We operate and scale large language model serving infrastructure, including model hosting, request routing, continuous batching, and cache optimization.
  • We deploy, host, and lifecycle-manage open-source and proprietary large language models on cloud-based container orchestration platforms and on-premises GPU clusters using reproducible infrastructure as code and continuous delivery pipelines.
  • We implement observability across logs, metrics, and traces with dashboards and actionable alerting for large language model and GPU workloads.
  • We tune GPU and accelerator capacity, autoscaling, and cost efficiency for large language model inference workloads using performance optimization techniques such as quantization, parallelism, and speculative decoding.
  • We lead reliability engineering for large language model endpoints through capacity planning, load and soak testing, safe rollouts, failover, and incident response for outages and model-quality regressions.
  • We participate in on-call rotations, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-up actions.
  • We identify recurring operational issues and automate remediation to improve platform stability and developer experience.
  • We build and maintain multi-agent systems with strong orchestration, including planning, coordination, tool-calling, state and memory management, and workflow control where appropriate.
  • We drive team adoption of enterprise-authorized AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes, while establishing consistent validation standards and promoting reuse of effective patterns across the team.
Technologies:
  • AI
  • Backend
  • Cloud
  • Incident Management
  • Support
  • Kubernetes
  • Machine Learning
  • Model Serving
  • Python
  • Security
  • AI Agents
  • LLM
  • Marketing
  • vLLM

More:

We are JPMorganChase, a global leader in financial services providing strategic advice and products to prominent corporations, governments, wealthy individuals, and institutional investors. Our AI and Machine Learning Platform team is focused on building and scaling AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. We offer a full-time role with meaningful ownership of reliability, performance, and cost-efficiency for large language model inference platforms, along with deep hands-on work in cloud, Kubernetes, observability, and production AI systems. We value diversity and inclusion, support equal opportunity, and provide reasonable accommodations where needed.

last updated 34 week of 2026

Interested in this role?

Submit your application now

How to Apply

Ready to apply for this position? Here's what you need:

  • An updated resume highlighting relevant experience
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

About JP Morgan Chase

JP Morgan Chase

JP Morgan Chase

Glasgow

IT

Skills & Technologies

PythonGoKubernetesMachine LearningAILLMUI

Inferred from job description

Salary Insight

£82,000

This role

£75,000

UK median

This salary is 9% above the UK median for Senior roles75,000/yr).

Based on 2024–2025 UK technology sector benchmarks

Explore More UK Opportunities

Thousands of tech jobs across the United Kingdom