Hackajob Ltd

Hackajob Ltd

Glasgow, Scotland

Lead Software Engineer - LLM Ops Platform Reliability

Full-Time£100,000 - 100,000 per yearheuteUnited Kingdom
IT

Job Description

Salary: £100,000 - 100,000 per year

Requirements:
  • Formal training, certification, or equivalent practical experience in software engineering concepts
  • Hands-on experience with system design, application development, testing, and operational stability in production environments
  • Advanced proficiency in Python for building production-grade services and tooling
  • Proficiency with automation and continuous delivery methods
  • Hands-on experience with AWS and Terraform for infrastructure delivery and lifecycle management
  • Strong understanding of site reliability engineering practices, including incident management, root-cause analysis, runbooks, and reliability patterns
  • Practical knowledge of observability and instrumentation across metrics, logs, and traces
  • Comfort with on-call operations and production troubleshooting
  • Hands-on production experience operating LLM inference servers such as vLLM and llm-d, or directly equivalent serving stacks
  • Hands-on experience hosting and serving LLMs on Amazon EKS and/or Amazon SageMaker, and on local GPU infrastructure
  • Knowledge of LLM reliability and risk considerations, including latency/throughput trade-offs, model and weight versioning, prompt/response logging, and safe rollout patterns
  • Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns
  • Experience building AI agents using frameworks such as LangChain, CrewAI, LangGraph, or similar orchestration platforms
  • Experience operating or integrating serving platforms such as KServe, Ray Serve, NVIDIA Triton Inference Server, Text Generation Inference (TGI), alongside vLLM/llm-d
  • Familiarity with Amazon SageMaker JumpStart, SageMaker Endpoints, and Amazon Bedrock for managed model hosting
  • Experience with online LLM quality monitoring, such as hallucination, toxicity, and drift detection, and tracing via OpenTelemetry conventions
  • Contributions to open-source LLM serving or inference projects, such as vLLM, llm-d, Ray, KServe, or Triton
Responsibilities:
  • Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure
  • Build backend services and APIs that enable reliable operation of AI infrastructure in production
  • Operate and scale LLM serving infrastructure, including model hosting, request routing, continuous batching, and KV-cache optimization
  • Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon EKS, Amazon SageMaker, on-prem, and local GPU clusters using reproducible infrastructure as code and continuous delivery pipelines
  • Implement observability with logs, metrics, traces, dashboards, and actionable alerting for LLM and GPU workloads
  • Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization techniques
  • Lead reliability engineering for LLM endpoints through capacity planning, load and soak testing, safe rollouts, failover, and incident response for outages and model-quality regressions
  • Participate in an on-call rotation, lead incident triage and mitigation, and produce clear post-incident root-cause analyses and follow-ups
  • Identify recurring operational issues and automate remediation to improve platform stability and developer experience
  • Build and maintain multi-agent systems with strong orchestration where appropriate
  • Contribute to an inclusive team culture and help drive adoption of leading-edge technologies through communities of practice
Technologies:
  • AI
  • AI Agents
  • AWS
  • Backend
  • Cloud
  • Incident Management
  • Support
  • Kubernetes
  • LLM
  • Machine Learning
  • Marketing
  • OpenTelemetry
  • Python
  • Terraform
  • vLLM
  • Grafana
  • Model Serving
  • Prometheus

More:

We are partnering directly with JPMorgan Chase for this role on our AI and Machine Learning Platform team. We are building and scaling AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI, with a focus on reliable, cost-efficient LLM inference at scale. We work with cloud and Kubernetes-based deployments, deep observability, and production-grade operational rigor across AWS, Amazon EKS, Amazon SageMaker, on-prem, and local GPU environments. J.P. Morgan is a global leader in financial services, and our Corporate Functions teams support our businesses, clients, customers, and employees across finance, risk, human resources, marketing, and more. We value diversity, inclusion, and equal opportunity, and we make reasonable accommodations for applicants and employees religious practices and beliefs as well as mental health or physical disability needs.

last updated 35 week of 2026

Interested in this role?

Submit your application now

How to Apply

Ready to apply for this position? Here's what you need:

  • An updated resume highlighting relevant experience
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

About Hackajob Ltd

Hackajob Ltd

Hackajob Ltd

Glasgow

IT

Skills & Technologies

PythonGoAWSKubernetesTerraformMachine LearningAILLMUI

Inferred from job description

Salary Insight

£100,000

This role

£85,000

UK median

This salary is 18% above the UK median for Lead roles85,000/yr).

Based on 2024–2025 UK technology sector benchmarks

Explore More UK Opportunities

Thousands of tech jobs across the United Kingdom