AI Infrastructure Engineer San Jose, CA
ESR Healthcare — Healthcare and IT staffing
About the role
AI Infrastructure Engineer San Jose, CA Duration: 6+ months Must have skills: AI, Kubernetes, Orchestration and DevOps ( All the four skills are mandatory)
Role Description: Architect and build custom Artificial Intelligence (AI) infrastructure solutions leveraging the Nutanix Kubernetes Platform and Nutanix AI. You will be responsible for designing high-performance computational stacks that integrate Nutanix AI, high-speed software-defined storage, and GPU-accelerated nodes. Your mission is to make AI infrastructure & ;invisible &; by optimizing for performance, power consumption, and seamless hybrid-multicloud scalability across on-prem.
Minimum Experience: 10 years Educational Qualification: 12 years full-time education Summary As an AI Infrastructure Engineer, you will design tailored AI solutions that bridge the gap between private data centers and public cloud. Your day-to-day will involve optimizing the Nutanix computational stack for large language models (LLMs) and generative AI workloads. You will serve as the SME for Nutanix AI, ensuring that compute, storage (Nutanix Objects/Files), and networking (Flow) are perfectly tuned for AI model training and inference.
Nutanix-Specific Responsibilities Hybrid Multicloud Architecture: Design seamless AI workflows using NC2 on Prem, allowing for rapid bursting of AI workloads from on-prem AHV clusters to the public cloud. Data Services for AI: Architect high-performance storage backends using Nutanix Objects (S3-compatible) to handle the massive datasets required for AI/ML. Kubernetes & Orchestration: Deploy and manage AI workloads using Nutanix Kubernetes Platform (NKP) to ensure containerized AI models are scalable and resilient. Infrastructure-as-Code: Implement IaC using Nutanix Calm or Terraform to automate the lifecycle of GPU-enabled nodes. Observability: Design frameworks (monitoring, logging, alerting) for proactive issue detection. Hands on experience on Prometheus, Grafana, ELK, and OpenTelemetry. Ensure high availability, disaster recovery, and fault tolerance across all systems. Networking & Security: Familiarity with Zero-Trust architectures, enterprise networking, storage, and virtualization. Invisible Infrastructure: Modernize legacy 3-tier AI silos into a unified, web-scale Nutanix environment. Public Professional & Technical Skills Nutanix Core: Deep proficiency in AOS (Acropolis Operating System) and AHV (Native Hypervisor). AI Performance: Experience with GPU Passthrough and vGPU configurations on Nutanix to optimize AI training performance. Security: Applying Nutanix Flow for micro segmentation to secure sensitive AI training data. Cost Management: Using Nutanix Cloud Manager (NCM) Cost Governance to monitor and optimize spend across hybrid environments. The "Hungry, Humble, Honest &; Expectations SME Leadership: Act as the primary technical authority for Nutanix AI integrations within the San Jose office. Collaboration: Work across teams to dismantle data silos, moving the organization toward a "One Platform" philosophy. Strategic Vision: Stay ahead of Nutanix product roadmaps to inform long-term AI infrastructure strategy.
What the index says about this role
- First seen by JobLarper — Aug 2, 2026, 6 days ago. Older postings collect hundreds of applicants — a tailored résumé matters more the longer a role has been live.
- No pay range in our index for this listing. Across 177 indexed DevOps & Site Reliability Engineer roles in US that do publish one, the middle half sits between $165k and $250k, median $204k — JobLarper's read of the market, not a figure from ESR Healthcare.
- What DevOps & Site Reliability Engineer roles ask for — across 1,405 indexed openings: Cloud (50%), Python (41%), REST/APIs (33%), Go (31%), Backend (24%). This posting names Cloud, Backend, AI/LLM, ML.
- ESR Healthcare is hiring actively — 90 open roles indexed.
Derived from the 27,000 roles JobLarper indexes daily from official company boards — not from the job description above.
More open roles at ESR Healthcare
- Amazon Connect Developer RemoteNYC, New York, United States · Mid
- Azure / Cloud Architect Data bricks LINUX Santa Clara, CaSanta Clara, California, United States · Senior+
- Azure Databricks Data Engineering Lead (17305-1) Seattle, WASeattle, WA, USA · Senior+
- Cybersecurity Technical Lead (14958-1) Houston, TXHouston, Texas, United States · Senior+
- AI Architect (12032-1) Houston, TXHouston, Texas, United States · Senior+
- cybersecurity Project Manager Bridgewater, NJBridgewater, New Jersey, United States · Manager
- Data Modeler & Architect remoteNYC, New York, United States · Senior+
- ai/ml phthon engineerSan Francisco, CA, USA · Mid
All 90 open roles at ESR Healthcare →
Similar DevOps & Site Reliability Engineer roles at other companies
- Finance Systems Engineer, TaxAnthropic · San Francisco, CA | Seattle, WA
- Senior SRE - VOIPRingCentral · Bangalore, India
- Senior OT / Edge DevOps Engineer (m/f/x)Reverion · Eresing · München
- Principal AI Platform EngineerSentinelOne · Brno, South Moravian, Czech Republic
- Lead Cloud DevOps Engineer (Oakland, CA Office)Fictiv · Oakland, CA Office
- Principal Site Reliability EngineerDell Technologies · Bengaluru, Karnataka, India
- Senior Application Support Engineer / Site Reliability Engineer (SRE)DTCC · Boston, MA, United States
- Senior Application Support Engineer (SRE)DTCC · Tampa, FL, United States
Browse all devops & site reliability engineer jobs in united states — 721 open roles across 306 companies.