Staff Platform Engineer (IC-4) at Moxie | Torre

Staff Platform Engineer (IC-4)

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Provide your expected compensation while applying
location_on
Remote (for Colombia residents)
Remote (for Brazil residents)
Remote (for Dominican Republic residents)
Remote (for Chile residents)
Shared by
Emma of Torre.ai
24 days ago

Responsibilities


At Moxie, we empower ambitious aesthetic entrepreneurs to build profitable, independent practices—without burnout, overwhelm, or guesswork. In just a few years, we've grown from an idea to a global, remote-first team now supporting 700+ practices nationwide.Our purpose is simple: to unlock sustainable success for aesthetic entrepreneurs, at every stage of their journey.About MoxieMoxie empowers aesthetic industry professionals to become successful entrepreneurs. We provide a sophisticated SaaS platform that simplifies the operational complexities of running MedSpas, enabling nurses and medical professionals to launch, operate, and grow their businesses across the country.Hundreds of customers rely on Moxie Suite to run their MedSpas end-to-end: scheduling, medical purchasing, payments and invoicing, bookkeeping, analytics, and more. The platform operates at real-world scale and reliability requirements, integrating with systems such as AWS, Vercel, Cloudflare, Stripe, Twilio, Datadog, and others.We are a fast-growing company focused on building reliable infrastructure, strong developer experience, and operational excellence as we scale.The RoleWe are looking for a Staff Platform Engineer to help own and evolve the systems that enable Moxie engineers to ship safely, quickly, and reliably.This role sits at the intersection of DevOps and SRE. Your most important responsibility will be incident handling — you'll be the primary owner of detecting, responding to, and resolving production issues across our infrastructure. During US business hours, you'll typically be the sole platform/infra expert on call, though developers will be available to support you as needed. Given this, the role can be demanding at times and requires availability beyond standard hours when incidents arise.Beyond incident response, you'll work on cloud infrastructure, CI/CD pipelines, deployment workflows, local development environments, observability, and operational tooling. You'll partner closely with product engineering teams, but your primary responsibility is the health, reliability, and usability of the platform itself.This is a hands-on individual contributor role with meaningful ownership, but no people management responsibilities.Key ResponsibilitiesObservability & ReliabilityParticipate in incident response as neededOwn and improve monitoring, logging, and alerting using Datadog, AWS, Vercel and related toolsEnsure systems are observable and failure modes are well understoodHelp teams learn from incidents through postmortems and follow-upsBalance reliability with delivery speed through pragmatic SRE practicesPlatform & InfrastructureOwn and operate core platform systems across AWS, GCP, Vercel, Github, and CloudflareImprove reliability, scalability, and security of production and non-production environmentsMaintain and evolve infrastructure supporting multiple services and teamsCI/CD & DeploymentsOwn and improve CI/CD pipelines (GitHub Actions), focusing on speed, reliability, and clarityImprove deployment workflows, rollbacks, and environment consistencyReduce deployment-related risk and manual interventionPartner with engineers and our QA team to improve release confidence and velocityDeveloper Experience & Local DevelopmentImprove local development environments and onboarding experience for engineersReduce friction in common workflows (setup, testing, debugging)Maintain tooling and documentation that helps engineers move faster with confidenceCross-Team CollaborationWork closely with the product engineering team to understand platform pain points and improve local development experience.Provide guidance and support on infrastructure, deployments, and operational best practicesContribute to platform standards and shared tooling through collaboration, not mandatesQualificationsRequired5+ years of experience in platform, DevOps, or SRE-focused rolesStrong experience operating production systems on AWS (certifications strongly preferred), Vercel, and/or GCPExperience building and maintaining CI/CD pipelines (GitHub Actions or similar)Strong understanding of cloud networking, security fundamentals, and IAMExperience with observability tooling (Datadog preferred)Ability to troubleshoot production issues calmly and systematicallyExcellent written and verbal communication skills in English (C1 or higher).Nice to HaveExperience with Cloudflare (DNS, WAF, edge configuration)Experience with Agentic and LLM based tooling and automations (Cursor, Codex, Claude Code, etc)Experience improving local development tooling (ie Docker, Husky, bash)Familiarity with infrastructure-as-code (Terraform or similar)Experience supporting regulated or compliance-sensitive environmentExperience working with PII, PHI, and sensitive data systems in general.Our StackCloud & Infrastructure: AWS ECS/Fargate, GCP, and VercelEdge & Security: CloudflareCI/CD: GitHub ActionsObservability: DatadogDatabase: AWS RDS (postgres)Version Control: Git, GitHubAI / LLM Tooling: Claude Code, Gemini, Cursor, CodeRabbit, Glean, CodexBackend Services: Python, Django (operational ownership, not feature dev)At Moxie, we believe in creating a workplace where everyone feels valued, trusted, and included. Our team lives by our values: act as owners, give more than we take, move with speed and care, and simplify and learn every day.We welcome people of all backgrounds, experiences, and perspectives to apply. If you require any accommodations to fully participate in the interview process, please let us know, we’re happy to assist.Compensation Range: $97K - $170K