Senior / Staff Machine Learning Infrastructure Engineer
Waabi — Tracked from its lever job board
About the role
Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class team, we're unlocking the next era of autonomous transportation with technology that's powering commercial autonomous trucks and robotaxis. Waabi is backed by and partners with world leaders in AI, automotive, logistics, and deep tech.
With offices in Toronto, San Francisco, Dallas, and Pittsburgh, Waabi is growing quickly and looking for diverse, innovative and collaborative candidates who want to impact the world in a positive way. To learn more visit: www.waabi.ai
You will..
- Design, develop, and implement the machine learning platform for the continuous deployment and integration of machine learning models.
- Collaborate with data scientists and engineers to understand model requirements and optimize pipeline processes.
- Automate the training, testing and deployment processes for machine learning models.
- Continuously monitor and maintain model pipelines, ensuring optimal performance, accuracy and reliability.
- Optimize machine learning pipelines for scalability, efficiency and cost-effectiveness.
- Ensure compliance with security and data privacy standards in all MLOps activities.
Qualifications:
- 3-5 years of experience supporting machine learning training platforms.
- Bachelor’s degree in Computer Science, Data Science or a related field.
- Strong understanding of machine learning principles and model lifecycle management.
- Proficiency in programming languages such as Python, with hands-on experience in machine learning frameworks like TensorFlow or PyTorch.
- Experience with cloud platforms like AWS, Azure, or Google Cloud and their respective machine learning services.
- Experience managing technology such as JupyterHub and Kubeflow.
- Familiarity with containerization and orchestration tools such as Kubernetes and Docker.
- Strong problem-solving skills and ability to troubleshoot complex issues.
- Experience with monitoring tools and practices for model performance in production.
- Ability to work collaboratively in cross-functional teams.
Bonus/nice to have:
- Experience with infrastructure-as-code (IaC) tools such as Terraform or Crossplane.
- Knowledge of big data technologies like Apache Spark or Hadoop.
- Familiarity with data engineering practices and tools.
- Experience with A/B testing and model validation in production environments.
- Relevant MLOps certifications (e.g., AWS Certified Machine Learning – Specialty, DataRobot MLOps Certification) are a plus.
What the index says about this role
- First seen by JobLarper — Aug 2, 2026, 6 days ago. Older postings collect hundreds of applicants — a tailored résumé matters more the longer a role has been live.
- What DevOps & Site Reliability Engineer roles ask for — across 1,405 indexed openings: Cloud (50%), Python (41%), REST/APIs (33%), Go (31%), Backend (24%). This posting names Cloud, Python, ML.
- Waabi is hiring actively — 44 open roles indexed.
Derived from the 27,000 roles JobLarper indexes daily from official company boards — not from the job description above.
More open roles at Waabi
- Electrical Engineering Team LeadPittsburgh, PA · San Francisco, CA · Senior+
- Senior/Staff Electrical EngineerPittsburgh, PA · Senior+
- IT Systems AdministratorPittsburgh, PA · Mid
- Senior / Staff Research Engineer, Simulation Assets & Content SystemsToronto, ON · San Francisco, CA · Pittsburgh, PA · Remote US & Canada · Senior+
- Physical Infrastructure Engineer (On-premise)Dallas, TX · Pittsburgh, PA · Mid
- Senior IT AdministratorDallas, TX · Senior+
- Lead Technical Program Manager, Platform IntegrationToronto, ON · San Francisco, CA · Pittsburgh, PA · Remote US & Canada · Manager
- Platform Verification Engineer - Embedded SystemsSan Francisco, CA · Pittsburgh, PA · Mid
Similar DevOps & Site Reliability Engineer roles at other companies
- Senior Site Reliability Engineer (SRE & Platform Reliability)Affirm · Remote Poland · Remote
- Senior Systems EngineerVoyager Space · Pittsburgh, PA; Remote · Remote
- Sr. Systems Engineer (Internal Support)Atlas Technica · Kyiv, Ukraine · Remote
- Staff AI Platform Engineer, Infrastructure ServicesSentinelOne · United States - Remote · Remote
- Senior AI Platform Engineer, Infrastructure ServicesSentinelOne · United States - Remote · Remote
- Staff AI Platform EngineerRelevance AI · Sydney, Australia · Remote
- Senior DevOps EngineerViz.ai · Tel Aviv, Israel · Remote
- Site Reliability Engineer - ClickHousePostHog · Remote (EMEA) · Remote (UK) · Remote
Browse all devops & site reliability engineer jobs in canada — 38 open roles across 29 companies.