# Senior ML Infrastructure Engineer

**Company:** [Ellison Institute of Technology](null/companies/tYoBQmFCuLVKUygNZbHUKp.md)
**Location:** Oxford, United Kingdom
**Workplace:** hybrid
**Department:** Central Business Operations & Other

[Apply for this job](null/view/34fd225f-d5f3-4556-a2d1-a0523339a4b3)

## Description

**Join us at EIT:**

At the Ellison Institute of Technology (EIT), we’re on a mission to translate scientific discovery into real world impact. We bring together visionary scientists, technologists, policy makers, and entrepreneurs to tackle humanity’s greatest challenges in four transformative areas:

-   Health, Medical Science & Generative Biology
-   Food Security & Sustainable Agriculture
-   Climate Change & Managing CO₂
-   Artificial Intelligence & Robotics

This is ambitious work - work that demands curiosity, courage, and a relentless drive to make a difference. At EIT, you’ll join a community built on excellence, innovation, tenacity, trust, and collaboration, where bold ideas become real-world breakthroughs. Together, we push boundaries, embrace complexity, and create solutions to scale ideas for lab to society. Explore more at [www.eit.org](https://www.eit.or)

**Your Role:**

Join our SciComp team to build the cloud and compute foundation that enables scientific breakthroughs. Deliver reliable, secure platforms and self-service guardrails that accelerate experimentation and turn ideas into results - faster, at scale, and with confidence. 

**Your Responsibilities:**

-   Build, operate, and continuously optimise our high-performance GPU training and inference clusters, focusing on robust, high-availability scheduling, isolation, and automated lifecycle management. 
-   Drive systems design and implementation for high-throughput data paths, optimising I/O, caching, and data locality across compute and storage (including our current Lustre implementation). 
-   Proactively benchmark, profile, and resolve performance bottlenecks across the compute, network, and orchestration layers to maximise efficiency for distributed training and inference. 
-   Establish comprehensive observability, resilience, and automated security controls to ensure compliance and robust operation of sensitive research environments. 
-   Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs, setting quotas and streamlining ML experimentation pipelines.

## Requirements

**Essential Skills, Qualifications & Experience:**

-   Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale 
-   A proactive, autonomous approach to systems design and the proven ability and desire to ideate, co-create and implement optimal solutions 
-   Exposure to migrating or transforming ML infrastructure from traditional schedulers to modern, containerised systems 
-   Expertise with high-throughput storage systems for ML/HPC workloads 
-   Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling to resolve bottlenecks 
-   A solid grasp of IaC and CI/CD practices (e.g., Terraform, Argo CD)

## Benefits

**We offer the following salary and benefits:**

-   Competitive salary (dependent on experience) + travel allowance + bonus
-   Enhanced holiday. Our annual leave allowance is 25 days plus 8 bank holidays and an additional 3 days between Christmas and New Year. You will also have the opportunity to purchase an additional 5 days annual leave in January and July.
-   Pension - Employer contribution 7.5%, minimum employee contribution 5%
-   Life Assurance.
-   Income Protection
-   Private Medical Insurance as standard for you, your partner and any dependents. Including hospital Cash Plan
-   Employee discounts
-   Electric car scheme
-   Nursery Salary Sacrifice scheme
-   Cycle to Work Scheme
-   Family Planning
-   Neurodiversity support including advise and assessments
-   Coaching & Therapy services

**Why work for EIT:**

You must have the right to work permanently in the UK with a willingness to travel as necessary. In certain cases, we can consider sponsorship, and this will be assessed on a case-by-case basis.

You will live in, or within easy commuting distance of, Oxford (or be willing to relocate) and can commit to being onsite at our Oxford office, a minimum of 3 days per working week.
