Zoox
Senior AI Inference Engineer - Model Optimization & Deployment
Foster City, CAfull timeunspecifiedtransportation$225k–$305k/yr
CUDA — Model Optimization — LLM Deployment — C++ — Machine Learning — Edge Computing — Inference Optimization — Python — Real-time Systems — Model Compression
Description
As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex models (LLMs, VLMs, or FMs) for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.