Machine Learning Infrastructure Engineer, GenAI Technology
Point72 Asset Management — Asset Management · Finance · Financial Services
About the role
A Career with Point72's Technology Team
As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications.
As a member of Point72’s Technology team, we encourage and support your professional development from day one—helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity—all while delivering real business impact for our multi-billion-dollar global business.
WHAT YOU'LL DO
Design and implement high-performance infrastructure to support large-scale generative AI and machine learning workloads, enabling faster model iteration and real business impact
Design and operate distributed systems for model training, hyperparameter tuning, inference, and data preprocessing pipelines to deliver reliable end-to-end machine learning (ML) workflows
Collaborate with ML researchers and engineers to produce models, optimizing compute utilization, training throughput, and inference latency
Develop and automate deployment, orchestration, and CI/CD pipelines for models and data workflows using container orchestration and infrastructure-as-code (IaC)
Implement observability, monitoring, and cost-management strategies for GPU and accelerator compute environments to maintain predictable performance and spend
Evaluate, integrate, and benchmark emerging hardware and software technologies across cloud and on-prem environments to improve scalability and throughput
Drive security, compliance, and operational runbooks for GenAI infrastructure including access controls, secrets management, and incident response procedures
Troubleshoot, profile, and optimize performance across GPU and CPU compute stacks to remove bottlenecks and increase reliability
Document architecture, operational practices, and mentor engineers to expand team capability and accelerate adoption of production-ready GenAI infrastructure
WHAT'S REQUIRED
Bachelor's or master's degree in computer science, electrical engineering, or a related technical field
3–7 years of experience building and maintaining scalable compute or machine learning infrastructure systems
Deep understanding of distributed systems, container orchestration (Kubernetes), and public cloud platforms such as AWS, Google Cloud Platform, or Azure
Hands-on experience with machine learning operations and infrastructure tools such as MLflow, Ray, Airflow, Kubeflow, and Terraform
Strong understanding of reinforcement learning concepts and their infrastructure implications
Proficiency in Python and systems-level programming in one or more languages such as Go, C++, or Rust
Strong debugging, performance profiling, and optimization skills across GPU and CPU compute stacks
Experience implementing monitoring, observability, and cost-optimization for GPU/accelerator-based compute environments
Excellent collaboration and communication skills with a systems-thinking mindset
Commitment to the highest ethical standards
WE TAKE CARE OF OUR PEOPLE
We invest in our people, their careers, their health, and their well-being. When you work here, we provide:
Fully-paid health care benefits
Generous parental and family leave policies
Volunteer opportunities
Support for employee-led affinity groups representing women, people of color and the LGBT+ community
Mental and physical wellness programs
Tuition assistance
A 401(k) savings program with an employer match and more
ABOUT POINT72
Point72 is a leading global alternative investment firm led by Steven A. Cohen. Building on more than 30 years of investing experience, Point72 seeks to deliver superior returns for its investors through fundamental and systematic investing strategies across asset classes and geographies. We aim to attract and retain the industry's brightest talent by cultivating an investor-led culture and committing to our people's long-term growth. For more information, visit https://point72.com/.
The annual base salary range for this role is $180,000-$300,000 (USD) , which does not include discretionary bonus compensation or our comprehensive benefits package. Actual compensation offered to the successful candidate may vary from posted hiring range based upon geographic location, work experience, education, and/or skill level, among other things.
What the index says about this role
- First seen by JobLarper — Aug 2, 2026, 6 days ago. Older postings collect hundreds of applicants — a tailored résumé matters more the longer a role has been live.
- No pay range in our index for this listing. Across 177 indexed DevOps & Site Reliability Engineer roles in US that do publish one, the middle half sits between $165k and $250k, median $204k — JobLarper's read of the market, not a figure from Point72 Asset Management.
- What DevOps & Site Reliability Engineer roles ask for — across 1,405 indexed openings: Cloud (50%), Python (41%), REST/APIs (33%), Go (31%), Backend (24%). This posting names Cloud, Python, Go, Backend, AI/LLM, ML.
- Point72 Asset Management is hiring actively — 95 open roles indexed.
Derived from the 27,000 roles JobLarper indexes daily from official company boards — not from the job description above.
More open roles at Point72 Asset Management
- Observability EngineerBengaluru, India · Mid
- Software Engineer, Treasury TechnologyWarsaw · Mid
- AI Data ScientistHong Kong, Singapore · Mid
- AI Solutions ArchitectUnited Kingdom · Senior+
- AI Product Analyst, Market Intelligence, SingaporeSingapore · Mid
- Machine Learning Engineer, GenAI TechnologyNew York, NY · Mid
- SDLC EngineerWarsaw · Mid
- Software Engineer, Treasury TechWarsaw · Mid
All 95 open roles at Point72 Asset Management →
Similar DevOps & Site Reliability Engineer roles at other companies
- Finance Systems Engineer, TaxAnthropic · San Francisco, CA | Seattle, WA
- Senior SRE - VOIPRingCentral · Bangalore, India
- Senior OT / Edge DevOps Engineer (m/f/x)Reverion · Eresing · München
- Principal AI Platform EngineerSentinelOne · Brno, South Moravian, Czech Republic
- Lead Cloud DevOps Engineer (Oakland, CA Office)Fictiv · Oakland, CA Office
- Principal Site Reliability EngineerDell Technologies · Bengaluru, Karnataka, India
- Senior Application Support Engineer / Site Reliability Engineer (SRE)DTCC · Boston, MA, United States
- Senior Application Support Engineer (SRE)DTCC · Tampa, FL, United States
Browse all devops & site reliability engineer jobs in united states — 721 open roles across 306 companies.