- Architect, build, and operate end-to-end ML pipelines for training, validation and deployment on Google Cloud and AWS.
- Define, instrument, and maintain logging, monitoring, and alerting for model performance and data drift.
- Automate CI/CD for ML artifacts and infrastructure using GitHub Actions or equivalent.
- Collaborate with cross-functional teams, including frontend engineers, backend engineers, research engineers, and infrastructure engineers.
- Write clean, well-documented, fast, and maintainable code.
- Help ensure our systems have high availability and performance.
- Experience in computer graphics or physics-based simulation.
- Background in setting up Prometheus/Grafana, ELK, or similar monitoring stacks.
- Experience with Vertex AI.
- Experience working with custom Domain-Specific Languages.
Foundation-llm-technologies
ML Ops Engineer (Boston, MA)
Remotefull timeunspecifiedmanufacturing
ML Pipeline Architecture — Google Cloud — AWS — CI/CD — GitHub Actions — Monitoring & Logging — Prometheus/Grafana — Python
Description