Principal SRE at Merative | Torre

Principal SRE

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Compensation
USD174k - 262k/year
location_on
Remote (anywhere)
Shared by
Emma of Torre.ai
7 days ago

Responsibilities


Merge medical imaging solutions, offered by Merative, combine intelligent, scalable imaging workflow tools with deep and broad expertise to help healthcare organizations improve their confidence in patient outcomes and optimize care delivery.With minimal supervision, leverages deep technical expertise to set the reliability, observability, and operational automation direction for business-critical, PHI-bearing cloud platforms.Establishes reliability standards, service level objectives, and automation practices across multiple teams, and leads their implementation through technical influence rather than direct reporting authority.ResponsibilitiesPeopleProvidetechnical guidance, mentorship, and leadership to engineers across development, QA, and operations teams.Interact regularly with lower and/or senior management on matters concerning multiple functional areas, departments, and/or customers.Act as a senior point of escalation for complex or high-severity production issues.Build reliability capability in others through design reviews, pairing, and blameless post-incident learning.Foster a culture and develop approaches that generate innovative ideas, products, and services.Reliability EngineeringDefine, publish, and govern service levelobjectives(SLOs)management framework,help to defineservice level indicators, and error budgets for the platform and its shared services.Set observability standards for metrics, logging, tracing, and alerting, and drive consistent adoption across teams.Own the technical side of the incident lifecycle: detection, response, escalation, and post-incident review; drive corrective actions to closure.Lead capacity planning, performance engineering, failure-domain isolation, and disaster recovery design.Drive resilience patterns into product architecture in partnership with development teams.Monitor and act on incoming issues from support, customers, or other stakeholders.Operational Automation Design and ImplementationDesign the automation strategy for platform operations and set the standards, patterns, and tooling other engineers build against.Design and implement automated remediation for recurring failure modes so routine faults are resolved without human intervention.Build andmaintaininfrastructure-as-code, environment provisioning, and deployment automation.Automate recurring operational work including patching, scaling, certificate rotation, backup and restore validation, and disaster recovery exercises.Build self-service tooling that lets development and support teams perform routine operational tasks safely and without escalation.Automate the collection of evidence and the verification of security and compliance controls that apply to PHI-bearing workloads.Identify, measure, and reduce operational toil; set and report against measurable toil-reduction targets.ProcessProvide guidance on company processes.Liaise with cross functional teams (development, product, program management, support, implementation, security, etc.) in delivering and supporting their projects.Support cross functional teams in resolving customer concerns.Participate in the creation and/or review and/or approval of architecture, design, and project documents.Plan, track, and deliver assigned reliability and automation initiatives, on-timeand on-budget.Report initiative status andimmediatelyescalate when work is varying from commitments.Effectively represent the platform’s reliability posture to the customer and to auditors if/as needed.Adhere to Mergemethodology, quality system requirements, and good engineering practices.Provide input into applicable budgets, including cloud consumption and tooling spend.Pursue self-development as an employee to be better at their current role as well as to grow into their next role if/asappropriate.Adhere to all applicable legal requirements.Core CompetenciesOrganized: Demonstrates strong planning, coordination, and attention to detail while managing multiple priorities.Analytical: Uses data-driven thinking and sound judgment to solve complex problems and make informed decisions.Growth Mindset: Embraces feedback, continuous learning, and opportunities for personal and professional development.Quick Learner: Rapidly acquires new technical knowledge, processes, and business context.Systems Thinking: Understands and evaluates failures and performance across an entire distributed platform rather than focusing on isolated components.Technical Acumen: Possesses deep expertise in cloud infrastructure and distributed systems, with the credibility to influence technical strategy across teams.Automation-First Mindset: Continuously seeks opportunities to eliminate manual effort through scalable automation and process improvement.Communication Skills: Effectively communicates complex ideas through clear, concise verbal and written communication.Leadership & Influence: Builds alignment, drives outcomes, and influences stakeholders without relying on direct authority.Priority Management: Effectively balances competing demands and focuses efforts on the highest-impact work.Composure Under Pressure: Maintains sound judgment, clear decision-making, and calm leadership during production incidents and high-pressure situations.Collaboration: Builds strong relationships and works effectively with diverse stakeholders across teams, functions, and external partners.Technical Skills, (if applicable):Cloud infrastructure and architecture, with depth in Microsoft AzureInfrastructure as code and configuration management (e.g., Terraform, Bicep/ARM, Ansible)CI/CD pipeline design and release automationContainers and orchestration (Kubernetes/AKS); service mesh conceptsObservability and telemetry tooling (e.g., Azure Monitor/KQL, Prometheus, Grafana, distributed tracing)Proficiencyin at least one automation or systems language (e.g., Python, Go, Bash, Java)Linux administration, networking, and identity/authorization fundamentalsIncident management and post-incident analysis practiceAbility to understand software architecture and design patternsTechnical Project ManagementExperience in an Agile EnvironmentMicrosoft OfficePreferred SkillsService mesh implementation (Istio)Event streaming and messaging platforms (Kafka)Relational and NoSQL data platform operations at scaleCybersecurity and security engineeringChaos engineering and resilience testingData EngineeringMachine Learning and Artificial Intelligence applied to operationsSupervisory Skills, (if applicable):Individual contributor role; no direct reports.Possess some fundamental leadership skills including strategic thinking, team building,adaptabilityand conflict resolution.Ability to technically lead and influence experienced level professionals.Qualifications Required:a) Education Requirements:Degree in Computer Science,Engineeringor related field; or equivalent level of industry related experience.PreferredAzure certification (e.g., Azure Solutions Architect Expert or DevOps Engineer Expert).Experience Required10+ years’ experience in software engineering, systems engineering, or infrastructure operations withdemonstratedprogression into a principal or staff level technical role.Demonstrated experience informally leading teams,projectsor people to successful outcomes.Demonstrated experience operating production SaaS at scale against formal availability commitments.Demonstrated experience designing and implementing operational automation that measurably reduced manual effort or recovery time.Preferred:Experience in a regulated environment (HIPAA/HITRUST, ISO 13485, IEC 62304).Medical imaging experience: DICOM, HL7Agile, ScrumWork EnvironmentThe work environment characteristics here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made