Machine Learning Platform Engineer Job Description Template

The machine learning platform engineer builds the internal tooling that machine learning teams depend on: training orchestration, experiment tracking, feature pipelines, model registry and deployment. The role is infrastructure engineering with machine learning specific requirements layered on top, and its success is measured by how much faster the modeling teams can move once it exists. The work is largely invisible when it is going well, which is the usual difficulty in arguing for investment in it.

Typical Duties and Responsibilities

  • Build and operate training orchestration and job scheduling
  • Develop feature pipelines and the systems that serve features consistently
  • Maintain experiment tracking and model registry infrastructure
  • Build deployment pipelines for models including rollback
  • Manage compute resources including GPU allocation and cost
  • Build observability for pipeline health and model serving
  • Provide interfaces and libraries that modeling teams actually use
  • Ensure training and serving consistency to avoid skew
  • Support scaling of training workloads
  • Document platform capabilities and support their adoption

Education

  • Bachelor’s degree in computer science or a related field

Required Skills and Experience

  • 4+ years in infrastructure or platform engineering with machine learning workloads
  • Strong Python plus infrastructure skills
  • Experience with Kubernetes and containerised workloads
  • Familiarity with orchestration tools such as Airflow, Kubeflow or Argo
  • Understanding of training and serving skew and how to prevent it
  • Experience managing GPU compute and its cost
  • Good API and library design for internal users
  • Solid monitoring and reliability practice

Preferred Qualifications

  • Experience with feature store implementations
  • Exposure to distributed training frameworks
Contact us

Recruit with Nexus IT Group