SF $1M Permanent
I’m working with a rapidly growing, well-funded AI company building some of the most technically ambitious real-world AI systems today.
They’re looking for an ML Infrastructure Engineer to join a highly technical team responsible for building the infrastructure that powers large-scale model training, experimentation, and deployment.
This is a hands-on engineering role for someone who enjoys solving difficult systems problems at the intersection of machine learning, distributed systems, and infrastructure.
What you’ll work on:
• Build and scale infrastructure for training large machine learning models
• Develop distributed training systems and improve training efficiency, reliability, and throughput
• Build data pipelines and infrastructure supporting large-scale ML workloads
• Improve GPU utilization, compute orchestration, checkpointing, and experiment management
• Develop tooling that enables researchers and ML engineers to iterate faster
• Diagnose performance bottlenecks across training, data, and compute systems
• Help take ML systems from experimentation through production and real-world deployment
What they’re looking for:
• Strong Python and software engineering fundamentals
• Experience building ML infrastructure, training infrastructure, or large-scale distributed systems
• Experience working with PyTorch and modern ML training stacks
• Strong understanding of distributed computing and GPU-based workloads
• Experience building reliable systems that support ML research or production ML
• Comfortable working in a fast-moving environment with significant technical ownership
Especially interesting backgrounds include:
• Distributed training and large-scale model training
• ML platforms / internal training infrastructure
• GPU infrastructure and compute orchestration
• Large-scale data infrastructure and pipelines
• Performance optimization for ML workloads
• Infrastructure supporting robotics, embodied AI, computer vision, or other real-world ML systems
• Experience taking ML systems beyond research and into production or physical-world environments
This is a great opportunity for an engineer who wants to work on hard ML systems problems at scale while being much closer to the models and real-world applications than you would be on a traditional infrastructure team.
📍 Bay Area - in person
If you have a strong ML infrastructure or distributed systems background and want to hear more, apply directly or message me.
Register for job alerts and be the first to hear about opportunities that match your search.
Finding your next role has never been so simple.
