Senior Technical Product Manager, GPU Orchestration at Vultr | Torre

Senior Technical Product Manager, GPU Orchestration

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Compensation
USD130k - 165k/year
location_on
Remote (for United States residents)
Match
skeleton-gauges
You have opted out of job matches in .
To undo this, go to the 'Skills and Interests' section of your preferences.
Review preferences
Shared by
Emma of Torre.ai
about 17 hours ago

Requirements and responsibilities


Who We AreVultr is on a mission to make high-performance cloud infrastructure easy to use, affordable, and locally accessible for enterprises and AI innovators around the world. With 33 global cloud data center locations, Vultr is trusted by hundreds of thousands of active customers across 185 countries for its flexible, scalable, global Cloud Compute, Cloud GPU, Bare Metal, and Cloud Storage solutions. In December 2024 Vultr announced an equity financing at a $3.5 billion valuation. Founded by David Aninowsky and self-funded for over a decade, Vultr has grown to become the world’s largest privately-held cloud infrastructure company.Vultr Cares100% company-paid insurance premiums for employee medical, dental and vision plans.401(k) plan that matches 100% up to 4%, with immediate vestingProfessional Development Reimbursement of $2,500 each year11 Holidays + Paid Time Off Accrual + Rollover PlanCommitment matters to Vultr! Increased PTO at 3 year and 10 year anniversary + 1 month paid sabbatical every 5 years + Anniversary Bonus each year$500 stipend for remote office setup in first year + $400 each following yearInternet reimbursement up to $75 per monthGym membership reimbursement up to $50 per monthCompany paid Wellable subscriptionJoin VultrVultr is seeking a highly skilled and experienced Senior Techincal Product Manager to own the GPU Orchestration product line — the platform that powers managed Kubernetes, managed Slurm, SUNK, and Run:ai integration for GPU-based AI and HPC workloads. The ideal candidate brings deep technical fluency in container orchestration, HPC scheduling, and distributed systems, combined with a strong product instinct for developer and operator platforms. This is a highly visible role in a high-growth technology company, which will require close partnership with Infrastructure, Compute, Networking, and Platform teams to build a reliable, scalable, and cost-efficient orchestration platform. This is your opportunity to join our fast growing team and leave your mark on Vultr and the future of AI Infrastructure.Key ResponsibilitiesDefine and execute the roadmap for managed Kubernetes, managed Slurm services, SUNK, and Run:ai integrationOwn the end-to-end cluster lifecycle, including provisioning, configuration, upgrades, scaling, high availability, and decommissioningEstablish scheduling and resource management capabilities for GPU workloads, including quotas, fair-share policies, multi-tenant isolation, and priority handlingDrive integration between orchestration services and core infrastructure components, including networking, storage, identity, observability, and billing systemsDefine service-level objectives for control plane reliability, job scheduling latency, cluster availability, and upgrade stabilityDesign APIs, CLI tooling, and UI workflows that enable self-service cluster management and workload operationsPartner with customer-facing teams to understand training, inference, and HPC use cases, translating real workload requirements into product capabilitiesMonitor industry trends in container orchestration, HPC scheduling, distributed systems, and AI infrastructure to inform product directionQualifications7+ years of product management experience in cloud infrastructure, container orchestration, HPC, or developer platformsDeep understanding of Kubernetes, Slurm, or similar orchestration and scheduling systems, including GPU scheduling, resource management, and multi-tenant isolationExperience defining product strategy and roadmaps for platform or infrastructure products at scaleStrong technical background — ability to engage with engineering on cluster lifecycle, control plane reliability, API design, and distributed systemsExperience with AI/ML infrastructure, including training workloads, inference serving, and GPU resource optimizationTrack record of shipping developer- and operator-facing products with measurable impact on reliability, adoption, or operational efficiencyExperience working across cross-functional teams (engineering, design, marketing, sales) in a fast-paced environmentExcellent written and verbal communication skills, with the ability to translate complex technical concepts for diverse audiencesBachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience)Compensation$130,000 - $165,000Final compensation will vary depending on years of experience, background/skill set, location, and applicable laws.