
Hackajob Ltd
Site Reliability Engineer / Production Support
Job Description
Salary: £100,000 - 100,000 per year
Requirements:- Strong SRE or production support experience with accountability for incident response in a production environment.
- Deep understanding of observability tools, alerting, logging, and distributed systems debugging.
- Experience managing and working with offshore support teams.
- Hands-on experience with reliability engineering, including resilience patterns, performance tuning, and capacity planning.
- Active use of AI tools for incident triage, automation, and runbook generation.
- Ability to understand complex system flows across multiple services and third-party integrations.
- Experience in financial services or similarly regulated environments is a strong advantage.
- Directly oversee the offshore Production Support team and be the single point person when incidents occur, escalating only to the Head of Cloud Operations when required.
- Run on-call and incident response, ensuring fast detection, triage, and restoration of services.
- Maintain observability standards, including logs, metrics, traces, and alert quality with low noise and high signal.
- Understand at a working level the key system flows, services, partners, and teams involved, and actively debug incidents.
- Lead reliability engineering work, including resilience patterns, performance tuning, and capacity planning.
- Facilitate post-incident reviews and track actions to completion.
- Use AI tools for automated alert correlation, root cause analysis, and runbook generation.
- Identify routine and common tasks, create automation plans, and execute them.
- Build automation that prevents incidents or resolves them faster next time.
- AI
- Cloud
- Support
More:
hackajob is partnering directly with Monument to hire for this role. We are building something genuinely rare: a financial brand designed for the mass affluent, professionals, entrepreneurs, and ambitious savers that traditional banks have systematically underserved for decades. We exist to make managing wealth simpler, smarter, and more human, treating every clients wealth with the same care as if it were our own. We hold over £7 billion in client savings, serve more than 100,000 clients, and were named the UKs fastest growing fintech in 2025. This Site Reliability Engineer / Production Support role is based in London (Oxford Circus) with a hybrid working pattern of 2 days per week, and reports to the Head of Cloud Operations. It offers the opportunity to own production reliability at a pre-IPO challenger bank, working with modern AI tools as a core part of the daily workflow and making a direct impact on clients and the business.
last updated 34 week of 2026
Interested in this role?
Submit your application now
How to Apply
About Hackajob Ltd

Hackajob Ltd
Charing Cross
IT
Skills & Technologies
Inferred from job description
Salary Insight
£100,000
This role
£60,000
UK median
This salary is 67% above the UK median for Software Engineers (£60,000/yr).
Based on 2024–2025 UK technology sector benchmarks