Staff Systems Engineer
Graphcore — Tracked from its greenhouse job board
About the role
About us
Graphcore is one of the world’s leading innovators in Artificial Intelligence compute. It is developing hardware, software and systems infrastructure that will unlock the next generation of AI breakthroughs and power the widespread adoption of AI solutions across every industry.
As part of the SoftBank Group, Graphcore is a member of an elite family of companies responsible for some of the world’s most transformative technologies. Together, they share a bold vision: to enable Artificial Super Intelligence and ensure its benefits are accessible to everyone.
Graphcore’s teams are drawn from diverse backgrounds and bring a broad range of skills and perspectives. A melting pot of AI research specialists, silicon designers, software engineers and systems architects, Graphcore enjoys a culture of continuous learning and constant innovation.
Job Summary
We are seeking a Staff Hardware Engineer to provide advanced operational, diagnostic, and engineering support for Graphcore’s Arm-based hardware platforms across lab and data center environments.
This role focuses on supporting hardware bring-up, validation, and troubleshooting of complex AI compute platforms, including server blades, racks, and rack-scale infrastructure. The successful candidate will collaborate closely with engineering, platform, and data center teams to ensure the reliability and performance of next-generation AI systems.
The Team
The Systems Engineering and Hardware Engineering teams are responsible for enabling the bring-up, validation, and operational reliability of Graphcore’s AI infrastructure platforms.
The team works closely with server engineering, firmware teams, platform architects, and data center operations to support the development, testing, and deployment of next-generation AI compute systems.
This collaborative environment enables rapid problem-solving and continuous improvement of Graphcore’s hardware platforms from early development through production deployment.
Responsibilities and Duties
Lead advanced break-fix troubleshooting for server blades, motherboards, power systems, and rack-scale infrastructure.
Support engineering bring-up activities, including component validation and firmware interaction testing.
Diagnose system-level failures involving thermal behavior, power anomalies, network configuration, and BIOS/BMC issues.
Collaborate with server engineering teams to perform root cause analysis and propose corrective actions or design improvements.
Support deployment and rollout of next-generation hardware platforms through structured validation and qualification cycles.
Interface with facilities and infrastructure teams to understand environmental factors impacting system reliability.
Develop and maintain standard operating procedures (SOPs), troubleshooting guides, and validation documentation.
Provide guidance and mentorship to junior technicians and engineers on troubleshooting methodologies and hardware diagnostics.
Participate in on-call rotations or off-hours support during critical engineering milestones or hardware bring-up phases.
Candidate Profile
Essential
Bachelor’s degree in Electrical Engineering, Computer Engineering, Computer Science, or related discipline.
10 years experience with server hardware architectures and board-level debugging.
Experience analyzing system logs, hardware telemetry, and power/thermal metrics to isolate hardware failures.
Hands-on experience with HPC systems, AI compute platforms, or rack-scale infrastructure.
Strong collaboration skills and ability to work effectively in fast-paced engineering environments.
Excellent written and verbal communication skills.
Desirable
Experience supporting prototype or pre-production hardware bring-up.
Familiarity with data center facilities, including liquid cooling and power distribution systems.
Experience using Python, Bash, or automation tools for hardware validation or troubleshooting.
Exposure to structured failure analysis and reliability engineering methodologies.
USA Benefits In addition to a competitive salary, Graphcore offers flexible working and a comprehensive benefits package designed to support your health, wellbeing and financial future. Our benefits include medical, dental and vision coverage, Flexible Spending Accounts (FSAs), Health Savings Accounts (HSAs), disability and life insurance, a 401(k) retirement plan, commuter benefits, wellness services and an Employee Assistance Programme (EAP). We welcome people of different backgrounds and experiences; we're committed to building an inclusive work environment that makes Graphcore a great home for everyone. We offer an equal opportunity process and understand that there are visible and invisible differences in all of us. We can provide a flexible approach to interview and encourage you to chat to us if you require any reasonable adjustments.
What the index says about this role
- First seen by JobLarper — Aug 2, 2026, 5 days ago. Older postings collect hundreds of applicants — a tailored résumé matters more the longer a role has been live.
- No pay range in our index for this listing. Across 177 indexed DevOps & Site Reliability Engineer roles in US that do publish one, the middle half sits between $165k and $250k, median $204k — JobLarper's read of the market, not a figure from Graphcore.
- What DevOps & Site Reliability Engineer roles ask for — across 1,405 indexed openings: Cloud (50%), Python (41%), REST/APIs (33%), Go (31%), Backend (24%). This posting names Python, React Native.
- Graphcore is hiring actively — 158 open roles indexed.
Derived from the 27,000 roles JobLarper indexes daily from official company boards — not from the job description above.
More open roles at Graphcore
- Staff Engineering Operations Technical Program ManagerAustin, Texas, United States · Manager
- Principal Electrical EngineerAustin, Texas, United States · Senior+
- Staff Systems Bring Up EngineerAustin, Texas, United States · Senior+
- Lead Bring-up & Characterisation EngineerBristol, UK · Senior+
- Python EngineerGdańsk, Pomeranian Voivodeship, Poland · Mid
- Senior DevOps EngineerGdańsk, Pomeranian Voivodeship, Poland · Senior+
- Senior SLT Test EngineerAustin, Texas, United States · Senior+
- Staff System Software EngineerBristol, UK; Cambridge, UK · Senior+
All 158 open roles at Graphcore →
Similar DevOps & Site Reliability Engineer roles at other companies
- Finance Systems Engineer, TaxAnthropic · San Francisco, CA | Seattle, WA
- Senior SRE - VOIPRingCentral · Bangalore, India
- Senior OT / Edge DevOps Engineer (m/f/x)Reverion · Eresing · München
- Principal AI Platform EngineerSentinelOne · Brno, South Moravian, Czech Republic
- Lead Cloud DevOps Engineer (Oakland, CA Office)Fictiv · Oakland, CA Office
- Principal Site Reliability EngineerDell Technologies · Bengaluru, Karnataka, India
- Senior Application Support Engineer / Site Reliability Engineer (SRE)DTCC · Boston, MA, United States
- Senior Application Support Engineer (SRE)DTCC · Tampa, FL, United States
Browse all devops & site reliability engineer jobs in united states — 721 open roles across 306 companies.