Kraken

Kraken

united kingdom, London

Senior AI Compute Infrastructure Engineer

Full-Time£47,000 - 87,000 per yearأمسUnited Kingdom
IT

Job Description

Salary: £47,000 - 87,000 per year

Requirements:
  • 5+ years of infrastructure engineering experience, including significant experience with GPU compute, ML infrastructure, distributed systems, high-performance computing, or large-scale production platforms.
  • Hands-on experience operating GPU clusters or accelerator-backed infrastructure in production or production-like environments, including scheduling, orchestration, utilization monitoring, and cost optimization.
  • Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging.
  • Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalent systems.
  • Proficiency in Python for infrastructure automation, tooling, debugging, integration, and operational workflows.
  • Practical understanding of performance tradeoffs involving batching, concurrency, memory usage, GPU utilization, model size, latency, throughput, availability, and cost.
  • Track record of optimizing compute costs while maintaining performance, reliability, and availability expectations.
  • Experience building observable systems with useful metrics, logs, traces, dashboards, alerts, and incident workflows.
  • Comfort working in high-stakes, always-on environments where uptime, throughput, correctness, and operational discipline are critical.
  • Clear communication skills and ability to explain infrastructure tradeoffs to researchers, product teams, platform engineers, security stakeholders, and engineering leadership.
  • Nice to have: Experience at a frontier AI lab, hyperscaler, high-frequency trading firm, research platform, or high-scale ML organization.
  • Nice to have: Familiarity with custom silicon or specialized accelerators such as TPUs, AWS Trainium, or Gaudi.
  • Nice to have: Background in capacity planning, procurement input, reserved capacity strategy, cloud accelerator economics, or GPU fleet cost management.
  • Nice to have: Experience with distributed training frameworks such as DeepSpeed, Megatron-LM, FSDP, Ray, or equivalent systems.
  • Nice to have: Experience debugging CUDA, NCCL, kernel, driver, runtime, memory, networking, or low-level performance issues.
  • Nice to have: Experience with Rust, C++, Go, CUDA, or other systems languages used for performance-critical infrastructure.
  • Nice to have: Crypto, financial services, trading infrastructure, or security-sensitive production infrastructure experience.
Responsibilities:
  • Own and operate GPU and accelerator clusters for training, inference, evaluation, and experimentation, including drivers, runtimes, kernels, device plugins, node configuration, scheduling primitives, and workload isolation.
  • Design infrastructure that enables our teams to run models locally on GPUs where strategically and economically preferable, reducing unnecessary reliance on external providers and containing compute costs.
  • Build and improve scheduling, orchestration, placement, quota management, and utilization systems across heterogeneous accelerator environments.
  • Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using frameworks such as vLLM, Triton Inference Server, TensorRT, or equivalent serving stacks.
  • Partner with ML engineers and researchers to remove bottlenecks in training, evaluation, batch inference, online inference, deployment, and production debugging workflows.
  • Build observability for GPU utilization, memory pressure, queue depth, saturation, token throughput, request latency, failed workloads, capacity pressure, and spend.
  • Drive reliability, incident response, alerting, runbooks, and post-incident improvements for always-on AI compute infrastructure.
  • Evaluate and integrate new hardware, cloud instance families, specialized accelerators, runtimes, schedulers, and serving frameworks.
  • Build tooling that makes GPU usage visible and accountable, and easier for internal teams to consume without requiring them to become infrastructure experts.
  • Contribute to long-term architecture decisions that balance performance, cost efficiency, scalability, operational simplicity, and production safety.
Technologies:
  • AI
  • AWS
  • Cloud
  • CUDA
  • Hardware
  • Kubernetes
  • Linux
  • Machine Learning
  • Model Training
  • Python
  • Rust
  • Security
  • vLLM
  • NodeJS
  • Backbone
  • DevOps
  • Model Serving

More:

We are Payward, the parent company behind Kraken, NinjaTrader, Breakout, xStocks, Payward Services, and CF Benchmarks. For 15 years, we have been building globally accessible financial infrastructure to advance an open financial system. Founded in 2011, Kraken is a crypto platform trusted by more than 10 million individuals and institutions worldwide, offering spot trading, margin, futures, staking, and OTC services. Our AI Compute and Infrastructure team sits within engineering leadership and builds the infrastructure for AI model training, inference, evaluation, and experimentation. You will join a small, senior, high-impact team working with AI/ML researchers, platform engineers, security teams, and product teams. This is a full-time, remote Engineering role in AI & Machine Learning, open to applicants in the United Kingdom, Argentina, Brazil, Bulgaria, Canada, Costa Rica, Cyprus, Czech Republic, Estonia, Hungary, Ireland, Latvia, Lithuania, Mexico, Panama, Peru, Poland, Portugal, Romania, Slovenia, South Africa, and Spain. We hire based on merit and value diverse backgrounds and perspectives. We encourage applications from people who do not meet every listed requirement, particularly those passionate or knowledgeable about crypto. We are an equal opportunity employer and consider qualified applicants with criminal histories consistent with applicable requirements. Applications are accepted on an ongoing basis unless a deadline is stated; our hiring process may include job-related skills or work-style assessments.

last updated 40 week of 2026

Interested in this role?

Submit your application now

How to Apply

Ready to apply for this position? Here's what you need:

  • An updated resume highlighting relevant experience
  • A compelling cover letter (if required)
  • Portfolio or work samples (for relevant positions)

About Kraken

Kraken

Kraken

united kingdom

IT

Skills & Technologies

PythonC++GoRustScalaAWSKubernetesLinuxMachine LearningAILLMDevOps

Inferred from job description

Salary Insight

£67,000

This role

£75,000

UK median

This salary is 11% below the UK median for Senior roles (£75,000/yr).

Based on 2024–2025 UK technology sector benchmarks

Explore More UK Opportunities

Thousands of tech jobs across the United Kingdom