Waabi
Senior / Staff ML Training Optimization Engineer
Remotefull timeunspecifiedtransportation
Python — C++ — PyTorch — CUDA — GPU Profiling — Distributed Training — Rust — Kubernetes
Description
You will... - Build standardized distributed training frameworks for research and production, drive our training towards new levels of stability and efficiency. - Comprehensively profile model runtime and memory to pinpoint performance bottlenecks. - Identify and evaluate emerging technologies that can be adopted into Waabi’s training and inference frameworks. Examples include designing new CUDA kernels, quantization-aware training and inference, and compilation/deployment techniques. - Work with researchers and ML engineers on best-practices for optimal resource usage. - Create and improve tooling and dashboards to ensure broad adoption of your work. Qualifications: - MS/PhD or Bachelors degree with a minimum of 4 years of industry experience in Computer Science, Robotics and/or similar technical field(s) of study. - Solid coding proficiency in a variety of coding languages including Python, C++ or Rust. - Experience in deep learning frameworks such as PyTorch or Jax. - Skilled in profiling CPU and GPU code using tools such as PyTorch Profiler and NVIDIA Nsight. - Open-minded and collaborative team player with willingness to help others. - Passionate about self-driving technologies, solving hard problems, and creating innovative solutions. Bonus/nice to have: - Experience in identifying when custom CUDA kernels are needed, and implementing them. - Experience in Bazel in a monorepo environment, and integrating third party packages into dev environments. - Experience with Kubernetes-based training platforms.