- Machine Learning Recruiters and Staffing SpecialistsMachine Learning Infrastructure Engineer
Machine Learning Jobs
- AI/ML Engineer
- Applied Scientist
- Computer Vision Engineer
- Computer Vision Research Scientist
- Generative AI Engineer
- Large Language Model (LLM) Engineer
- Lead Machine Learning Engineer
- Machine Learning Architect
- Machine Learning Infrastructure Engineer
- Machine Learning Platform Engineer
- Machine Learning Research Scientist
- ML Product Manager
- MLOps Engineer
- Natural Language Processing (NLP) Engineer
- Principal Machine Learning Engineer
- Recommendation Systems Engineer
- Research Engineer
- Senior Machine Learning Engineer
- VP of Machine Learning
The machine learning infrastructure engineer builds and operates the compute, storage and networking that training and inference run on. The workloads are unusual: expensive accelerators, large data movement, long running jobs and inference services with tight latency requirements. The role combines systems engineering with careful attention to cost, since machine learning infrastructure is easy to overspend on quietly.
Typical Duties and Responsibilities
- Build and operate compute infrastructure for training and inference
- Manage GPU and accelerator resources, scheduling and utilization
- Optimize data loading and storage throughput for training workloads
- Build and tune inference serving infrastructure for latency and cost
- Implement autoscaling appropriate to bursty machine learning workloads
- Monitor utilization and drive down cost per training run and per inference
- Support distributed training including networking and communication libraries
- Automate provisioning and environment management
- Investigate performance bottlenecks across the stack
- Maintain reliability and capacity planning for the platform
Education
- Bachelor’s degree in computer science, engineering or a related field
Required Skills and Experience
- 4+ years in infrastructure engineering including machine learning workloads
- Strong Linux, networking and cloud infrastructure skills
- Experience with GPU infrastructure and scheduling
- Understanding of distributed training communication patterns
- Skill at profiling and removing data loading bottlenecks
- Cost management for compute intensive workloads
- Infrastructure as code and automation practice
- Good incident response and reliability discipline
Preferred Qualifications
- Experience with high performance computing environments
- Familiarity with inference optimization such as quantisation or batching