Biometric Talent
Site Reliability Engineer
Job Description
Salary: £40,000 - 65,000 per year
Requirements:- Strong software engineering experience, particularly with Python or Golang
- Experience with monitoring, alerting and observability
- Knowledge of OpenTelemetry and modern observability practices
- Experience establishing proactive monitoring and alerting for complex platforms
- Strong understanding of SRE principles, including SLIs and SLOs
- Experience with modern software development practices and lifecycles
- Proficiency in shell scripting
- Experience with Infrastructure as Code, automation and orchestration, ideally using Terraform and Ansible
- Experience with tools such as Grafana, Splunk, New Relic and PagerDuty
- Experience working within large-scale, 24/7 enterprise environments where availability and stability are critical
- Strong incident management, troubleshooting and root cause analysis experience
- Hands-on experience using LLM platforms and coding assistants to improve productivity and quality
- Experience or interest in using AI for telemetry, predictive insights and root-cause analysis
- Write and contribute to code that improves service reliability and observability
- Develop tools, operational APIs and automation to improve system management
- Establish proactive monitoring and alerting across complex platforms
- Implement service instrumentation using OpenTelemetry
- Build sophisticated dashboards using Grafana, Splunk and New Relic
- Automate manual processes and reduce operational toil
- Work with Infrastructure as Code and orchestration technologies
- Support live incident resolution and contribute to post-mortem analysis
- Carry out root cause analysis and implement effective remediation
- Drive initiatives to improve system reliability, performance and observability
- Maintain and administer existing monitoring and analytics platforms
- Work with IT Operations to provide critical tooling and capabilities
- Share knowledge and mentor colleagues on new technologies and practices
- Use AI tools, LLM platforms and coding assistants in day-to-day work to improve productivity, reduce toil and explore new approaches to autonomous operations, telemetry and system health
- AI
- Ansible
- Golang
- Grafana
- Incident Management
- Support
- LLM
- OpenTelemetry
- PagerDuty
- Python
- Splunk
- Terraform
- JavaScript
More:
We are supporting our client with the appointment of an experienced Site Reliability Engineer (SRE) to join their Platform Engineering team in Manchester, with 2 days per week onsite. This is a software-focused SRE role where software engineering skills are used to solve operational problems and improve reliability, observability and performance in a large-scale production environment. The role offers a salary of £40,000 to £65,000 per annum depending on experience and skillset, along with a performance-based bonus, pension scheme, hybrid working, flexible working hours, 25 days holiday plus your birthday off and bank holidays, the option to buy or sell up to 5 additional days, and free gym membership. We work closely with Development, Platform Delivery and IT Operations teams to build the tooling, automation, monitoring and practices that keep critical systems reliable and resilient, and we place a strong focus on AI-enabled engineering using AI tools, LLM platforms and coding assistants.
last updated 39 week of 2026
Interested in this role?
Submit your application now
How to Apply
About Biometric Talent
Biometric Talent
Manchester
IT
Skills & Technologies
Inferred from job description
Salary Insight
£52,500
This role
£60,000
UK median
This salary is 12% below the UK median for Software Engineers (£60,000/yr).
Based on 2024–2025 UK technology sector benchmarks