BY Recruiting

BY Recruiting

Machine Learning Infrastructure Engineer

fast-scaling AI startup
San Francisco, CADirect HireOn-site
Apply Now

About the role

We are a fast-scaling AI startup building large foundation models for physical systems. Our training and inference challenges require deep expertise in distributed clusters, petabyte-scale data pipelines, and low-level performance optimization.

We look for infrastructure engineers who are excited to tackle unsolved problems. If you have experience building large-scale ML infrastructure in related fields such as language and vision models, robotics, or biology, we want to hear from you.

Responsibilities

Design, deploy, and maintain large distributed ML training and inference clusters

Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the entire ML lifecycle

Research and test various training approaches including parallelization techniques and numerical precision trade-offs across different model scales

Analyze, profile and debug low-level GPU operations to optimize performance

Stay up-to-date on research to bring new ideas to work

What we are looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.

Strong grasp of state-of-the-art techniques for optimizing training and inference workloads

Demonstrated proficiency with distributed training frameworks (e.g. FSDP, DeepSpeed) to train large foundation models

Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings

Familiarity with containerization and orchestration frameworks (e.g., Kubernetes, Docker)

Background working on distributed task management systems and scalable model serving and deployment architectures

Understanding of monitoring, logging, observability, and version control best practices for ML systems

Requirements

- 3-10 years of experience building large-scale ML infrastructure for core foundation models (not fine-tuning)
- Experience training large-scale foundation models and working with large GPU infrastructures
- Background in physics, robotics, biology, or AI at the frontier of these fields, ideally at a science or physical AI company (e.g., self-driving, robotics, biology)
- Proficiency with distributed training frameworks (e.g., FSDP, DeepSpeed)
- Experience with low-level GPU performance optimization and debugging (CUDA, JAX)
- Familiarity with containerization and orchestration (Kubernetes, Docker) and cloud ML/AI services (GCP, AWS, or Azure)
- Generalist mindset with experience across the ML lifecycle, able to work closely with researchers and engineers
- Mission-driven, intentional career focus with strong problem-solving and rapid execution