Reliability Engineer (Remote) at Kohl's | Torre

Reliability Engineer (Remote)

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Provide your expected compensation while applying
location_on
Remote (for United States residents)
Shared by
Julien PETIT
5 days ago

Responsibilities


As Reliability Engineer, you will ensure the resilience and availability of Kohl’s systems and applications and collaborate closely with development teams to review designs, conduct risk assessments and implement robust monitoring and failover mechanisms.What You’ll DoDrive incident response efforts, perform root cause analysis and implement preventative measures to enhance system reliabilityEstablish consistent practices that elevate Kohl’s operational excellence through automation and process improvementsFollow software lifecycle and drive reliability, observability and efficiency across product teams within an assigned domainIdentify repeated toil and find opportunities for automation and risk reductionOn-call on a rotation to respond to production incidents and conduct blameless retros and root-cause analyses (RCAs) to drive a culture of continuous improvementsProactively identify failures before they cause outages using chaos engineering techniques such as edge cases, failure modes and design reviewAdvise on capacity planning and provide continuous assessments on systems behavior and consumptionWork with product managers to identify and prioritize work for reliability best practices (i.e., leveraging SLIs/SLOs/Error Budgets)Additional tasks may be assignedWhat Skills You HaveRequiredBachelor's Degree or equivalent puter Science or related field2+ years of experience in software developmentStrong programming skills in one or more languages (Java, Python, Go or Node.js)Working knowledge of systems architecture, operating system internals and network fundamentalsExperience working with one cloud platform (e.g., GCP, AWS, or Azure)PreferredExperience with monitoring techniques and tools (e.g., CloudWatch, Grafana, Prometheus, OpenTelemetry, Tracing)Working knowledge around containerization and container orchestration (e.g., Docker, Kubernetes, Rancher)