Site Reliability Engineer - Dedicated Hosted Runners
GitLab — Artificial Intelligence · Automation · Business/Productivity Software
About the role
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.
The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.
* Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab.
An overview of this role We're looking for an Intermediate Site Reliability Engineer to join the Runners Platform team. In this role, you'll build and operate Hosted Runners for GitLab Dedicated, the managed CI/CD compute platform that runs our customers' pipelines inside single-tenant Amazon Web Services (AWS) environments. You'll own infrastructure automation across the full lifecycle: provisioning runner fleets with Terraform, extending the Go tooling and autoscaling stack, and making sure customer continuous integration and continuous delivery (CI/CD) jobs keep running reliably across the platform. What you’ll do
Design, build, and operate AWS infrastructure for Hosted Runners across many single-tenant environments, including Elastic Compute Cloud (EC2), Auto Scaling Groups, Virtual Private Cloud (VPC) networking, subnets, Network Address Translation (NAT), network access control lists, PrivateLink, Identity and Access Management (IAM), and Elastic Container Registry (ECR).
Develop and maintain infrastructure as code using Terraform, contributing to common modules and the deployment tooling that provisions and upgrades runner stacks.
Write Go code for our runner tooling and autoscaling components, including the fleeting instance-lifecycle plugins, zero-downtime deployment command-line interface, and reusable infrastructure toolkits.
Build and improve the GitLab CI/CD pipelines that orchestrate blue/green zero-downtime deployments, automated upgrades, quality assurance validation, and performance testing of runner stacks.
Define and monitor service level objectives for CI job execution, including queue times, job success rates, and fleet saturation, and build the Grafana dashboards, alerts, and runbooks behind them.
Participate in an on-call rotation, handle incidents affecting customer CI/CD workloads, and automate away recurring toil.
Run performance and scale testing that reflects real customer workloads, and tune autoscaling parameters for cost and reliability.
Write documentation and runbooks so the broader team can operate runner stacks consistently.
What you’ll bring
Professional experience operating production infrastructure on AWS at scale, including EC2, Auto Scaling Groups, IAM, and VPC networking.
Strong infrastructure-as-code experience with Terraform, including writing and refactoring modules used by other teams.
Proficiency in Go for building and debugging infrastructure tooling, or strong experience in another systems language and willingness to work in Go daily.
Practical knowledge of CI/CD systems and job execution: how pipelines schedule work, how ephemeral build environments get provisioned and torn down, and what makes CI workloads reliable.
Experience with observability practices such as metrics, dashboards, alerting, logging, and service-level-objective-based monitoring, using tools such as Prometheus, Grafana, and OpenSearch.
Experience with on-call rotations and incident management for customer-facing systems.
Strong problem-solving skills, excellent written communication, and comfort working asynchronously across Americas, Europe, Middle East, Africa, and Asia-Pacific time zones.
Direct GitLab Runner experience, familiarity with configuration management such as Ansible, and container tooling such as Docker are a plus.
About the team The Runners Platform team builds and operates GitLab Hosted Runners, the managed CI/CD compute behind GitLab.com and GitLab Dedicated. This role focuses on Hosted Runners for GitLab Dedicated: GitLab deploys and operates isolated, single-tenant runner fleets in each customer's AWS environment. We support zero-downtime deployments and continue to invest in high availability, performance at scale, and cost efficiency. We own everything from the Terraform modules and Go tooling through the service level objective dashboards and runbooks we use to operate the platform day to day.
How GitLab Supports Full-Time Employees
Benefits to support your health, finances, and well-being
Flexible Paid Time Off
Team Member Resource Groups
Equity Compensation & Employee Stock Purchase Plan
Growth and Development Fund
Parental Leave
Please note that we welcome interest from candidates with varying levels of experience; many successful candidates do not meet every single requirement. Additionally, studies have shown that people from underrepresented groups are less likely to apply to a job unless they meet every single qualification. If you're excited about this role, please apply and allow our recruiters to assess your application.
Country Hiring Guidelines: GitLab hires new team members in countries around the world. All of our roles are remote, however some roles may carry specific location-based eligibility requirements. Our Talent Acquisition team can help answer any questions about location after starting the recruiting process.
Privacy Policy: Please review our Recruitment Privacy Policy. Your privacy is important to us.
GitLab is proud to be an equal opportunity workplace and is an affirmative action employer. GitLab’s policies and practices relating to recruitment, employment, career development and advancement, promotion, and retirement are based solely on merit, regardless of race, color, religion, ancestry, sex (including pregnancy, lactation, sexual orientation, gender identity, or gender expression), national origin, age, citizenship, marital status, mental or physical disability, genetic information (including family medical history), discharge status from the military, protected veteran status (which includes disabled veterans, recently separated veterans, active duty wartime or campaign badge veterans, and Armed Forces service medal veterans), or any other basis protected by law. GitLab will not tolerate discrimination or harassment based on any of these characteristics. See also GitLab’s EEO Policy and EEO is the Law . If you have a disability or special need that requires accommodation , please let us know during the recruiting process .
What the index says about this role
- First seen by JobLarper — Aug 2, 2026, 6 days ago. Older postings collect hundreds of applicants — a tailored résumé matters more the longer a role has been live.
- What DevOps & Site Reliability Engineer roles ask for — across 1,405 indexed openings: Cloud (50%), Python (41%), REST/APIs (33%), Go (31%), Backend (24%). This posting names Cloud, REST/APIs, Go.
- GitLab is hiring actively — 94 open roles indexed.
Derived from the 27,000 roles JobLarper indexes daily from official company boards — not from the job description above.
More open roles at GitLab
- Principal Product Manager, AI Software FactoryRemote, Canada; Remote, United States · Manager
- Senior Backend Engineer, Platform ReadinessRemote, Canada; Remote, United States · Senior+
- Intermediate Backend Engineer, Platform ReadinessRemote, Canada; Remote, United States · Mid
- Manager, Solutions Architecture - EastRemote, United States · Manager
- Support Engineer (AMER)Remote, Canada; Remote, United States · Mid
- Senior Customer Success Architect - SingaporeRemote, Singapore · Senior+
- Solutions Architect - FranceRemote, France · Senior+
- Enterprise QA Engineer, SalesforceBangalore, India · Mid
Similar DevOps & Site Reliability Engineer roles at other companies
- Finance Systems Engineer, TaxAnthropic · San Francisco, CA | Seattle, WA
- Senior SRE - VOIPRingCentral · Bangalore, India
- Senior OT / Edge DevOps Engineer (m/f/x)Reverion · Eresing · München
- Principal AI Platform EngineerSentinelOne · Brno, South Moravian, Czech Republic
- Lead Cloud DevOps Engineer (Oakland, CA Office)Fictiv · Oakland, CA Office
- Principal Site Reliability EngineerDell Technologies · Bengaluru, Karnataka, India
- Senior Application Support Engineer / Site Reliability Engineer (SRE)DTCC · Boston, MA, United States
- Senior Application Support Engineer (SRE)DTCC · Tampa, FL, United States
Browse all devops & site reliability engineer jobs in india — 93 open roles across 57 companies.