Machine Learning Infrastructure Engineer Job Description Template

The machine learning infrastructure engineer builds and operates the compute, storage and networking that training and inference run on. The workloads are unusual: expensive accelerators, large data movement, long running jobs and inference services with tight latency requirements. The role combines systems engineering with careful attention to cost, since machine learning infrastructure is easy to overspend on quietly.

Typical Duties and Responsibilities

  • Build and operate compute infrastructure for training and inference
  • Manage GPU and accelerator resources, scheduling and utilization
  • Optimize data loading and storage throughput for training workloads
  • Build and tune inference serving infrastructure for latency and cost
  • Implement autoscaling appropriate to bursty machine learning workloads
  • Monitor utilization and drive down cost per training run and per inference
  • Support distributed training including networking and communication libraries
  • Automate provisioning and environment management
  • Investigate performance bottlenecks across the stack
  • Maintain reliability and capacity planning for the platform

Education

  • Bachelor’s degree in computer science, engineering or a related field

Required Skills and Experience

  • 4+ years in infrastructure engineering including machine learning workloads
  • Strong Linux, networking and cloud infrastructure skills
  • Experience with GPU infrastructure and scheduling
  • Understanding of distributed training communication patterns
  • Skill at profiling and removing data loading bottlenecks
  • Cost management for compute intensive workloads
  • Infrastructure as code and automation practice
  • Good incident response and reliability discipline

Preferred Qualifications

  • Experience with high performance computing environments
  • Familiarity with inference optimization such as quantisation or batching
Contact us

Recruit with Nexus IT Group