Lead SRE - BeReal at Voodoo | Torre

Lead SRE - BeReal

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Provide your expected compensation while applying
location_on
Remote (for France residents)
Shared by
Emma of Torre.ai
3 months ago

Responsibilities


About BeRealAt BeReal, we are dedicated to authenticity in social media. By encouraging users to share unfiltered moments, we foster genuine connections and celebrate real life. We are now an international team of 100+ and have 40M+ monthly active users. Backed by Voodoo, our team is fully focused on scaling BeReal into an iconic social network used by hundreds of millions.The Infrastructure team provides the backbone that powers the company’s growth, ensuring the scalability, efficiency, and reliability of our platform. We design and operate our infrastructure on GCP. Working hand in hand with developers, we enable teams to ship fast and efficiently while maintaining a strong focus on costs and performance. Our mission is to create a developer-friendly, cost-effective, and highly automated infrastructure that supports innovation at scale.RoleDefine and drive SRE practices across the organization, including SLIs, SLOs, error budgets, incident management, postmortem processes, and long-term reliability improvements across the platformDesign, implement, and optimize infrastructure for availability, scalability, reliability, and cost efficiencyOwn and evolve our observability stack, improving monitoring, alerting, logging, and distributed tracingDrive automation of infrastructure and operational workflows (e.g., Terraform, Terragrunt, Kubernetes)Lead FinOps initiatives, developing tools and insights to optimize cloud costsPartner closely with development squads to improve service reliability, performance, and operational excellenceInfluence architectural decisions and establish best practices for building resilient distributed systemsMentor and support Infrastructure engineers, helping raise the bar on reliability, operational excellence, and technical executionAnalyze performance bottlenecks and work on solutions such as scaling strategies, service optimizations, and system debuggingProfileStrong knowledge of KubernetesExperience with high traffic, distributed systems architectures, and related tools (service discovery, config/secret management, etc.)Strong knowledge of one Cloud provider (AWS or GCP preferred)Proven experience defining and operating SRE practices (SLOs, incident management, observability, reliability engineering)Strong operational mindset with experience managing production incidents and driving reliability improvementsLeadership and mentoring experience, with the ability to influence technical decisions across teamsOwnership-driven – If something isn’t working, you don’t wait for instructions; you improve itPragmatic and impact-oriented – You balance reliability, delivery speed, and business prioritiesPerformance vs cost-conscious – You make decisions that align with both technical excellence and financial sustainabilityOur StackOperator: KubernetesCI/CD: Argocd, Github actionsCloud provider: GCPMonitoring: DatadogInfra as code: Terraform / TerragruntLanguages: golang / nodeDatastores: Spanner / PostgreSQL / RedisBenefitsCompetitive salary based on experienceSwile Lunch voucherGymlib (100% covered by Voodoo)Premium healthcare coverage with SideCare, 100% covered for you and your familyWellness activities in our Paris office