Humanloop

Humanloop

London

Staff Software Engineer, Observability & Profiling

Full-Time£57,000 - 73,000 per yearavant-hierUnited Kingdom
IT

Job Description

Salary: £57,000 - 73,000 per year

Requirements:
  • We have hands-on experience building and operating large-scale observability or monitoring infrastructure.
  • We have deep, hands-on experience with observability signals end to end, from instrumentation through ingest to query and analysis.
  • We understand high-throughput telemetry pipelines and the tradeoffs involved in collecting, storing, and querying operational data at scale.
  • We are comfortable digging below the application layer into the kernel, the network stack, or the hardware.
  • We have excellent communication skills and enjoy partnering with internal teams to improve operational visibility and incident response capabilities.
  • We are excited about building foundational infrastructure and comfortable navigating ambiguous, high-impact technical challenges, both independently and with a team.
  • We have a bachelors degree or an equivalent combination of education, training, and/or experience.
  • We have a field of study relevant to the role, as demonstrated through coursework, training, or professional experience.
  • We have years of experience aligned with the internal job level requirements for the position.
  • We have 10+ years of relevant industry experience, including building and operating large-scale observability or monitoring infrastructure.
  • We have experience building or operating eBPF-based observability in production, including tracing, profiling, or network visibility.
  • We have experience running continuous profiling at fleet scale, including managing overhead budgets and symbolization.
  • We have kernel- and syscall-level debugging experience and performance engineering craft.
  • We have experience profiling or instrumenting accelerator workloads.
  • We have experience operating metrics systems at very high cardinality, or large-scale telemetry storage backends.
  • We have experience with OpenTelemetry instrumentation, collector pipelines, and tail-based sampling strategies.
  • We are interested in applying AI/LLMs to operational workflows such as automated root cause analysis, anomaly detection, or intelligent alerting.
Responsibilities:
  • We design and build scalable telemetry ingest and storage pipelines for metrics, logs, traces, and error data across our multi-cluster infrastructure.
  • We build observability solutions that give engineers deep, low-overhead visibility into system behavior across the fleet.
  • We own and evolve core observability platforms, driving migrations and architectural improvements that improve reliability, reduce cost, and scale with organizational growth.
  • We build instrumentation libraries, SDKs, and eBPF-based auto-instrumentation to emit high-quality telemetry, with and without code changes.
  • We reduce mean time to detection and resolution by building cross-signal correlation, unified query interfaces, and AI-assisted diagnostic tooling.
  • We drive fleet-wide efficiency by turning continuous profiling and utilization telemetry into actionable optimization insights across CPU, memory, and accelerator fleets.
  • We partner with Research, Inference, Product, and Infrastructure teams to ensure observability solutions meet the unique needs of each organization.
Technologies:
  • AI
  • Hardware
  • Support
  • Network
  • OpenTelemetry

More:

We are Anthropic, a public benefit corporation headquartered in San Francisco, and our mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. Our Observability team sits within our Infrastructure organization and builds the monitoring and telemetry systems that support our researchers and engineers across metrics, logging, tracing, profiling, error analytics, alerting, dashboards, and query tools. As we scale across massive GPU, TPU, and Trainium clusters, we are developing next-generation observability capabilities to help teams detect, diagnose, and resolve issues quickly, even below the application layer. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office environment. We also support a location-based hybrid policy, expecting staff to be in one of our offices at least 25% of the time, and we make reasonable efforts to sponsor visas where possible.

last updated 37 week of 2026

Interested in this role?

Submit your application now

How to Apply

Ready to apply for this position? Here's what you need:

  • An updated resume highlighting relevant experience
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

About Humanloop

Humanloop

Humanloop

London

IT

Skills & Technologies

ScalaRESTAILLMUI

Inferred from job description

Salary Insight

£65,000

This role

£60,000

UK median

This salary is 8% above the UK median for Software Engineers60,000/yr).

Based on 2024–2025 UK technology sector benchmarks

Explore More UK Opportunities

Thousands of tech jobs across the United Kingdom