goliathpatter.png

Machine Learning Engineer

SF $1M Permanent

I’m working with a rapidly growing, well-funded AI company building some of the most technically ambitious real-world AI systems today.


They’re looking for an ML Infrastructure Engineer to join a highly technical team responsible for building the infrastructure that powers large-scale model training, experimentation, and deployment.


This is a hands-on engineering role for someone who enjoys solving difficult systems problems at the intersection of machine learning, distributed systems, and infrastructure.


What you’ll work on:

• Build and scale infrastructure for training large machine learning models

• Develop distributed training systems and improve training efficiency, reliability, and throughput

• Build data pipelines and infrastructure supporting large-scale ML workloads

• Improve GPU utilization, compute orchestration, checkpointing, and experiment management

• Develop tooling that enables researchers and ML engineers to iterate faster

• Diagnose performance bottlenecks across training, data, and compute systems

• Help take ML systems from experimentation through production and real-world deployment


What they’re looking for:

• Strong Python and software engineering fundamentals

• Experience building ML infrastructure, training infrastructure, or large-scale distributed systems

• Experience working with PyTorch and modern ML training stacks

• Strong understanding of distributed computing and GPU-based workloads

• Experience building reliable systems that support ML research or production ML

• Comfortable working in a fast-moving environment with significant technical ownership


Especially interesting backgrounds include:

• Distributed training and large-scale model training

• ML platforms / internal training infrastructure

• GPU infrastructure and compute orchestration

• Large-scale data infrastructure and pipelines

• Performance optimization for ML workloads

• Infrastructure supporting robotics, embodied AI, computer vision, or other real-world ML systems

• Experience taking ML systems beyond research and into production or physical-world environments


This is a great opportunity for an engineer who wants to work on hard ML systems problems at scale while being much closer to the models and real-world applications than you would be on a traditional infrastructure team.


📍 Bay Area - in person


If you have a strong ML infrastructure or distributed systems background and want to hear more, apply directly or message me.

Share this job:

Apply now

Similar Jobs