Humanloop

Humanloop

London

Staff Software Engineer AI Reliability Engineering

Full-Time£57,000 - 73,000 per yearevvelsi günUnited Kingdom
IT

Job Description

Salary: £57,000 - 73,000 per year

Requirements:
  • We are looking for strong distributed systems, infrastructure, or reliability backgrounds, including reliability-minded software engineers and SREs.
  • We value candidates who are curious and brave, and who are comfortable jumping into unfamiliar systems during an incident to help drive resolution.
  • We look for people who think holistically about how systems compose and where the seams are.
  • We need someone who can build lasting relationships across teams and work as a trusted teammate.
  • We value people who care about users and feel ownership over outcomes, even for systems they do not own.
  • Excellent communication and collaboration skills are important, as you will partner across the company.
  • We value diverse experience across product stacks, databases, distributed systems, and related technical domains.
  • A bachelors degree or an equivalent combination of education, training, and/or experience is required.
  • A field relevant to the role through coursework, training, or professional experience is required.
  • Years of experience should align with the internal job level requirements for the position.
  • Strong candidates may also have experience as an SRE, Production Engineer, or in similar reliability-focused roles on large-scale systems.
  • Experience operating large-scale model serving or training infrastructure, including environments with more than 1,000 GPUs, is a plus.
  • Experience with ML hardware accelerators such as GPUs, TPUs, or Trainium is a plus.
  • Knowledge of ML-specific networking optimizations such as RDMA and InfiniBand is a plus.
  • Expertise in AI-specific observability tools and frameworks is a plus.
  • Experience with chaos engineering and systematic resilience testing is a plus.
  • Contributions to open-source infrastructure or ML tooling are a plus.
Responsibilities:
  • We develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity.
  • We design and implement monitoring and observability systems across the token path.
  • We assist in the design and implementation of high-availability serving infrastructure across multiple regions and cloud providers.
  • We lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements.
  • We support the reliability of safeguard model serving, which is critical for both site reliability and our safety commitments.
Technologies:
  • AI
  • API
  • Cloud
  • Hardware
  • InfiniBand
  • Support
  • Model Serving
  • Network
  • RDMA

More:

We are Anthropic, a public benefit corporation headquartered in San Francisco, and our mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for our users and for society. Our AIRE (AI Reliability Engineering) team works across Anthropic to improve reliability across our most critical serving paths, from the SDK through our network, API layers, serving infrastructure, and accelerators and back. We work in a highly collaborative environment with dynamic, cross-cutting exposure to the systems that matter most, and we value communication, teamwork, and impact. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, a lovely office space in which to collaborate with colleagues, and a hybrid policy with in-office presence expected at least 25% of the time. We also sponsor visas where possible and encourage applicants from underrepresented groups to apply.

last updated 34 week of 2026

Interested in this role?

Submit your application now

How to Apply

Ready to apply for this position? Here's what you need:

  • An updated resume highlighting relevant experience
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

About Humanloop

Humanloop

Humanloop

London

IT

Skills & Technologies

RustAISREUI

Inferred from job description

Salary Insight

£65,000

This role

£60,000

UK median

This salary is 8% above the UK median for Software Engineers60,000/yr).

Based on 2024–2025 UK technology sector benchmarks

Explore More UK Opportunities

Thousands of tech jobs across the United Kingdom