Fuse Energy Supply
AI Inference Engineer
Job Description
Salary: £63,000 - 103,000 per year
Requirements:- 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
- Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them, including batching, KV-cache management, quantisation, and speculative decoding.
- Strong systems thinking, with the ability to reason about the full path from incoming request to served response across a large cluster.
- Comfort working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
- A track record of making high-stakes architecture calls and owning the outcome.
- Comfort operating without a playbook in a founding role shaping a new function around early-stage architecture.
- Experience with Triton or custom ML inference/training frameworks is nice to have.
- Experience with autoscaling or capacity planning for large-scale inference workloads is nice to have.
- Exposure to multi-tenant serving or SLA-driven infrastructure is nice to have.
- Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system is nice to have.
- Familiarity with Kubernetes/Slurm for cluster orchestration is nice to have.
- Interest or experience in energy markets, grid systems, or sustainability-focused compute is nice to have.
- Define our inference serving strategy and architecture from first principles.
- Design and build the serving stack, including request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
- Own our model-level optimisation strategy for serving, deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with our CUDA/GPU engineers.
- Make the core software architecture decisions on serving frameworks and orchestration, such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents.
- Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
- Act as a direct technical owner of inference performance and reliability.
- Work closely with our CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
- Set the standards, tooling, and benchmarks this function will run on as it grows.
- AI
- CTO
- CUDA
- Hardware
- Kubernetes
- LLM
- vLLM
More:
We are Fuse Energy, a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system, and we have raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group, and strategic angels such as Nico Rosberg, the Co-Founder of Solana, and GPs behind Meta, Revolut, Spotify, Uber, and more. As data centres become one of the largest and fastest-growing sources of electricity demand, we are expanding into high-performance compute infrastructure at the intersection of energy and AI. We are building the GPU/CUDA performance layer and the inference serving layer from scratch, and this founding engineer role reports directly to our CTO. We offer a competitive salary, an equity sign-on bonus, a biannual bonus scheme, fully expensed tech to match your needs, and a breakfast and dinner allowance for office-based employees.
last updated 36 week of 2026
Interested in this role?
Submit your application now
How to Apply
About Fuse Energy Supply
Fuse Energy Supply
London
IT
Skills & Technologies
Inferred from job description
Salary Insight
£83,000
This role
£60,000
UK median
This salary is 38% above the UK median for Software Engineers (£60,000/yr).
Based on 2024–2025 UK technology sector benchmarks