Site Reliability Engineer (SRE) for AI Training and Inference Infrastructure at Boson AI | Torre

Site Reliability Engineer (SRE) for AI Training and Inference Infrastructure

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Provide your expected compensation while applying
location_on
Remote (for Canada residents)
Shared by
Emma of Torre.ai
6 days ago

Responsibilities


About The RoleBoson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from “it works” to dependable, observable, and scalable.You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area—networking, cluster scheduling, storage, GPU systems, or AI infrastructure— and the curiosity and judgment to collaborate across the rest.ResponsibilitiesDesign, operate, and improve reliable infrastructure for AI training and inference workloadsOwn and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platformsBuild monitoring, alerting, runbooks, and incident-response practices that make systems easier to operateDiagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloadsPartner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvementsImprove provisioning, configuration management, testing, and deployment automationHelp plan cluster growth, capacity allocation, upgrades, and lifecycle managementContribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standardsMinimum Qualifications4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations roleStrong hands-on expertise in at least one of the following:Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBandCluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platformsDistributed storage, particularly CephGPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshootingAI training or model-serving infrastructureExperience operating production systems with a focus on availability, performance, security, and automationStrong Linux administration and scripting skillsA systematic approach to troubleshooting across multiple layers of a complex systemClear written and verbal communication skills, including the ability to work effectively with a distributed teamPreferred QualificationsExperience supporting GPU-intensive AI or HPC environmentsExperience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ EthernetFamiliarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure toolingExperience operating or tuning Ceph clustersFamiliarity with observability tooling such as Prometheus, Grafana, and centralized logging systemsExperience with hardware provisioning, firmware management, and bare-metal automationExperience running large-scale distributed training or high-throughput inference workloadsFamiliarity with cloud and hybrid infrastructure across AWS, GCP, or AzureBoson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we’d love to hear from you.