GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation.An overview of this roleSite Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure.This is a single application for Site Reliability Engineering opportunities across our Infrastructure Platforms department. Rather than asking you to choose the right team or level upfront, we evaluate your skills holistically and match you to the opportunity that best aligns with your experience and our hiring needs. We hire Site Reliability Engineers from Intermediate through Senior Staff across multiple Infrastructure Platforms teams.We don't expect every candidate to have experience with every technology in our environment. We're looking for engineers with strong technical fundamentals, a growth mindset, and the ability to learn quickly. We'll support you in becoming successful with GitLab's tools, systems, and ways of working.How our SRE hiring worksRecruiter Screen: A conversation about your background, what you're looking for, and the level and teams that fit, so we can point your process in the right direction.Core Technical: The shared assessment every SRE candidate takes, regardless of eventual team. A low-stress, collaborative discussion covering source code, system architecture, and incident review.Peer Technical: Team-specific depth, run by SREs from the team you're most likely to join, focused on the problems that team actually works on.Hiring Manager Interview: A conversation about ownership, judgment, execution, collaboration, and growth, the non-technical signals that make an SRE effective at GitLab.Skip-Level Interview: A conversation with a senior leader on values alignment, and how you'll work across teams.After your interviews, we consider your performance alongside our current hiring needs to confirm the level and team where you'll do your best work. Interview results are a major factor, and final placement also reflects our active hiring priorities at the time.What level am I?We calibrate your level during the process, but here is roughly what each looks like so you know where you might land.IntermediateYou make meaningful contributions to reliability, automation, and operational efficiency, working independently within a scoped areaYou diagnose issues on your own, understand system dependencies, and can explain the tradeoffs you madeYou prioritize well, break work into manageable steps, and use automation to reduce toilYou document your work clearly and keep yourself moving without needing check-insSeniorYou drive reliability improvements across multiple projects or services and prioritize them based on real system needsYou lead investigations, anticipate cascading failures, and coordinate incident responseYou own delivery end to end, unblock others, and improve the patterns your team works byYou communicate complex ideas clearly, influence how work gets done, and enable coordination across teamsStaffYou shape reliability strategy across teams and services and define patterns that others reuseYou introduce prevention strategies, identify systemic weaknesses, and influence incident response practices beyond your immediate areaYou design execution and automation approaches that work at organizational scaleYou connect reliability work to platform and business needsSenior StaffYou set technical direction for reliability across a sub-department, not just a teamYou drive the hardest, most ambiguous systems problems and establish standards and guardrails that multiple teams adoptYou mentor Staff and Senior engineersYou align reliability strategy with long-range platform direction and represent Infrastructure's interests across the wider Engineering organizationWhat you'll doKeep user-facing services and production systems reliable, scalable, and efficientBuild automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflowsOperate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scalingWrite and mainta ship changes safely through CI/CD and GitOpsParticipate age alerts, follow and improve runbooks, and escalate appropriatelyContribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outagesTake part ning learnings into changes in automation and processDocument runbooks, architecture decisions, and reviews so your findings become repeatable practicesWhat you'll bringExperience keeping production systems reliable, combining an operations mindset with real software engineering practiceExperience building net-new infrastructure tooling and automation, not just configuring existing tools. For example, Terraform modules, Kubernetes operators or controllers, or production automation and services written from scratchThe ability to read, debug, and reason about code. Most of our teams work in Go; some work in Ruby. You can discuss a piece of code's behavior, performance, and failure modesExperience with infrastructure as code, and with Kubernetes and its ecosystem, at a depth appropriate to your levelHands-on experience with at least one major cloud provider (GCP or AWS)Familiarity with observability practices, including metrics, logging, alerting, and SLOs or SLIs, and using data to inform operational decisionsComfort participating h a structured approach to troubleshooting under pressureStrong written communication and the ability to operate as a manager-of-one tributed environmentA track record of using automation, and increasingly AI, to reduce toil and improve how you and your team workAlignment with GitLab's values and a commitment to working in accordance with themAbout the teamInfrastructure Platforms is responsible for the availability, reliability, performance, and scalability of GitLab's user-facing services, most notably GitLab.com. The department spans sub-departments including Production Engineering and Dedicated, and the teams withident response, and our single-tenant Dedicated offering. We are a globally distributed, all-remote group that works asynchronously, favors automation over toil, and closes the loop with monitoring and metrics to drive accountability.The base salary range for this role’s listed level is currently for residents of the United States only. The base salary range does not include any bonuses, equity, or benefits. Sales roles are also eligible for incentive pay targeted at up to 100% of the offered base salary.United States Salary Range$126,400 - $314,400 USDHow GitLab Supports Full-Time EmployeesBenefits to support your health, finances, and well-beingFlexible Paid Time OffTeam Member Resource GroupsEquity Compensation & Employee Stock Purchase PlanGrowth and Development FundParental LeavePlease note that we welcome interest from candidates with varying levels of experience; many successful candidates do not meet every single requirement. If you're excited about this role, please apply and allow our recruiters to assess your application.Country Hiring Guidelines: GitLab hires new team members in countries around the world. All of our roles are remote, however some roles may carry specific location-based eligibility requirements.