Senior Site Reliability Engineer at Playson | Torre

Senior Site Reliability Engineer

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Provide your expected compensation while applying
location_on
Remote (for Ukraine residents)
Shared by
Emma of Torre.ai
7 days ago

Responsibilities


About the RoleWe’re looking for a Senior Site Reliability Engineer to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale.If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.Key ResponsibilitiesOwn system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real timeParticipate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environmentInvestigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrenceBuild and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystemDeploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)Drive automation and proactively improve system resilience, reducing manual intervention and recurring issuesMaintain and evolve CI/CD pipelines and infrastructure-as-code practicesCollaborate closely with engineering teams to support deployments and minimise user impact in a live environmentIntroduce and integrate new tools and technologies to enhance scalability, reliability, and performanceHandle environment-specific requests and ensure smooth day-to-day platform operations under constant loadRequirementsStrong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environmentsExperience with GitOps tools such as FluxCD or ArgoCDProven experience in incident response, root cause analysis, and postmortems in production systemsSolid experience with AWS, Terraform, Docker, and CI/CD pipelinesExperience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatchStrong understanding of networking concepts and protocolsProficiency in at least one scripting language (e.g. Python, Go, Node.js)Experience working with version control systems (Git)Familiarity with incident management tools like PagerDuty, Opsgenie, or similarAbility to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountabilityProactive, resilient mindset with a focus on continuous improvement and system stabilityWhat We OfferCompetitive SalaryQuarterly BonusesUnlimited Paid Time OffUnlimited Paid Sick LeaveRemote & Flexible WorkingPrivate Medical InsuranceFinancial Support for Life EventsProfessional Development BudgetInternational ExposureRegular Company Events*Benefits may vary depending on location and contractual agreementRecruitment Process1. HR Interview (30-45 min)2. Technical interview (90 min)4. Final Interview with C-level (60 min)