
A&O Shearman
Cloud Operations Service Reliability Engineer
Job Description
Salary: £45,000 - 73,000 per year
Requirements:- Strong analytical and problem-solving skills, with a logical approach to issue identification, diagnosis and service improvement.
- Technically curious, with enthusiasm for understanding a broad set of systems, technologies and operational domains.
- Ability to interpret monitoring data, identify patterns and translate operational insight into meaningful improvement activity.
- Ability to make sound decisions under pressure and support effective incident response.
- Strong commitment to service reliability, operational resilience and excellent customer service.
- Commercial acumen, including an understanding of IT service costs, cloud consumption and how technology adds value to the business.
- Ability to promote technical standards, automation and reliability practices using clear, business-friendly language.
- Highly self-motivated and able to undertake activities to the highest professional standards.
- Excellent communication skills, both oral and written.
- Ability to operate within a wider team where there may be ambiguity and conflicting priorities.
- Ability to build effective working relationships across diverse internal teams and influence adoption of monitoring, automation and cloud engineering standards.
- Experience of working in a global environment across international locations with an appreciation of multiple cultures.
- Practical knowledge of SRE principles, observability, incident response, problem management and operational resilience.
- Detailed practical knowledge of Microsoft Azure infrastructure and platform services, including monitoring, diagnostics, RBAC, networking and automation.
- Knowledge of Infrastructure as Code, source control and pipeline-based delivery using tools such as Bicep, Azure DevOps and GitHub.
- Knowledge of configuration management and automation tooling such as Ansible.
- Minimum 4–5 years IT experience with at least 2 years experience in a cloud operations, infrastructure, platform engineering, SRE or 3rd line support role.
- Experience of monitoring, alerting, incident investigation and operational issue identification in a complex technology environment.
- Experience using or implementing monitoring using Elastic is desirable.
- Experience using or supporting automation and delivery tooling such as Bicep, Azure DevOps, GitHub and Ansible.
- Experience working with diverse internal teams to improve service supportability, resilience and operational standards.
- Experience of working in an ITIL environment.
- Ideally, minimum A level standard education or equivalent.
- Accreditation in relevant technologies is preferred.
- ITIL Foundation is preferred.
- Improve the reliability, observability and operational resilience of our cloud-hosted services.
- Monitor cloud infrastructure, platform services and supported application environments for health, availability, performance signals and operational events.
- Perform cloud engineering and automation using Infrastructure as Code, deployment pipelines, configuration management and standards-led delivery.
- Promote service resiliency through proactive issue identification, operational insight, automation and continuous improvement.
- Support monitoring, observability and alerting across cloud-hosted services.
- Work with Azure public cloud engineering across IaaS, PaaS, networking, identity, RBAC and platform diagnostics.
- Use Bicep, Azure DevOps pipelines and GitHub-based source control and collaboration for Infrastructure as Code and automation.
- Use Ansible or equivalent tooling for configuration management and standards automation.
- Use or implement monitoring solutions using Elastic where applicable.
- Develop operational reporting, issue trend analysis and actionable dashboards to support service improvement.
- Ensure monitoring and operational insight are designed, implemented and understood so services can be supported, improved and made more resilient.
- Provide subject matter expertise in cloud operations, observability, automation and reliability engineering practices.
- Work globally across cloud-hosted services and platform capabilities, independent of location.
- Support our environmental goals and initiatives.
- Improve end-to-end observability for supported services with internal technology teams.
- Maintain documentation including monitoring standards, known issues, operational patterns, troubleshooting guidance and support handbooks.
- Diagnose and support resolution of incidents and problems by interpreting monitoring signals, operational telemetry and service behaviour.
- Improve the quality, relevance and routing of alerts so issues can be detected and acted on quickly.
- Contribute to root cause analysis, problem management and continuous improvement by identifying recurring patterns, observability gaps and automation opportunities.
- Provide specialist guidance on cloud engineering patterns, Infrastructure as Code, deployment pipelines and automated configuration management.
- Support implementation of monitoring and automation standards across new and existing services.
- Ensure operational documentation, handover materials and support guidance are created and suitable for BAU operation.
- Identify operational, reliability and supportability risks arising from gaps in monitoring, alerting, automation or cloud platform standards.
- Participate in recovery, resilience and operational readiness activities to help prove services can be supported effectively.
- Promote standardised patterns, automated controls, repeatable engineering practices and effective use of best practice.
- Advocate for source control, pipeline-based delivery, Infrastructure as Code and configuration management to improve quality, auditability and operational reliability.
- Ansible
- Azure
- Cloud
- DevOps
- GitHub
- IaaS
- Support
- ITIL
- PaaS
- RBAC
- Security
More:
We are A&O Shearman, and this role sits within our cloud operations and reliability function supporting a broad technology estate. You will work across international locations with a diverse set of internal teams, helping to improve service supportability, resilience and operational standards. We value technical curiosity, professional standards and collaborative working, and we encourage the use of monitoring, automation and cloud engineering practices to support continuous improvement.
last updated 35 week of 2026
Interested in this role?
Submit your application now
How to Apply
About A&O Shearman

A&O Shearman
Carrickfergus
IT
Skills & Technologies
Inferred from job description
Salary Insight
£59,000
This role
£60,000
UK median
This salary is 2% below the UK median for Software Engineers (£60,000/yr).
Based on 2024–2025 UK technology sector benchmarks