Hackajob Ltd

Hackajob Ltd

Charing Cross, London

Site Reliability Engineer / Production Support

Full-Time£100,000 - 100,000 per yearevvelsi günUnited Kingdom
IT

Job Description

Salary: £100,000 - 100,000 per year

Requirements:
  • Strong SRE or production support experience with accountability for incident response in a production environment.
  • Deep understanding of observability tools, alerting, logging, and distributed systems debugging.
  • Experience managing and working with offshore support teams.
  • Hands-on experience with reliability engineering, including resilience patterns, performance tuning, and capacity planning.
  • Active use of AI tools for incident triage, automation, and runbook generation.
  • Ability to understand complex system flows across multiple services and third-party integrations.
  • Experience in financial services or similarly regulated environments is a strong advantage.
Responsibilities:
  • Directly oversee the offshore Production Support team and be the single point person when incidents occur, escalating only to the Head of Cloud Operations when required.
  • Run on-call and incident response, ensuring fast detection, triage, and restoration of services.
  • Maintain observability standards, including logs, metrics, traces, and alert quality with low noise and high signal.
  • Understand at a working level the key system flows, services, partners, and teams involved, and actively debug incidents.
  • Lead reliability engineering work, including resilience patterns, performance tuning, and capacity planning.
  • Facilitate post-incident reviews and track actions to completion.
  • Use AI tools for automated alert correlation, root cause analysis, and runbook generation.
  • Identify routine and common tasks, create automation plans, and execute them.
  • Build automation that prevents incidents or resolves them faster next time.
Technologies:
  • AI
  • Cloud
  • Support

More:

hackajob is partnering directly with Monument to hire for this role. We are building something genuinely rare: a financial brand designed for the mass affluent, professionals, entrepreneurs, and ambitious savers that traditional banks have systematically underserved for decades. We exist to make managing wealth simpler, smarter, and more human, treating every clients wealth with the same care as if it were our own. We hold over £7 billion in client savings, serve more than 100,000 clients, and were named the UKs fastest growing fintech in 2025. This Site Reliability Engineer / Production Support role is based in London (Oxford Circus) with a hybrid working pattern of 2 days per week, and reports to the Head of Cloud Operations. It offers the opportunity to own production reliability at a pre-IPO challenger bank, working with modern AI tools as a core part of the daily workflow and making a direct impact on clients and the business.

last updated 34 week of 2026

Interested in this role?

Submit your application now

How to Apply

Ready to apply for this position? Here's what you need:

  • An updated resume highlighting relevant experience
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

About Hackajob Ltd

Hackajob Ltd

Hackajob Ltd

Charing Cross

IT

Skills & Technologies

ScalaRESTAISREUI

Inferred from job description

Salary Insight

£100,000

This role

£60,000

UK median

This salary is 67% above the UK median for Software Engineers60,000/yr).

Based on 2024–2025 UK technology sector benchmarks

Explore More UK Opportunities

Thousands of tech jobs across the United Kingdom