Site Reliability Engineer (SRE) / Monitoring Specialist at NEEPANLOK INFOTECH | Torre

Site Reliability Engineer (SRE) / Monitoring Specialist

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: To be defined

Provide your expected compensation while applying
location_on
Remote (anywhere)
Shared by
Raul Vasquez P.
1 day ago

Responsibilities


Key ResponsibilitiesMonitor production systems and applications using Prometheus & GrafanaDesign and maintain monitoring dashboards and alertsProactively identify performance, availability and reliability issuesHandle production incidents and ensure timely resolutionPerform troubleshooting and Root Cause Analysis (RCA)Coordinate with development, infrastructure and application teams during critical incidentsMonitor system health, capacity and overall platform reliabilityParticipate in incident response and problem-management activitiesIdentify opportunities to improve monitoring, alerting and system reliabilitySupport production deployments and post-deployment monitoringRequired Technical SkillsPrometheus – Monitoring & MetricsGrafana – Dashboards & VisualizationProduction Support & Application MonitoringIncident Management & TroubleshootingRoot Cause Analysis (RCA)System health, availability & performance monitoringAlert configuration, analysis and resolutionExperience working in production environmentsCandidate Profile6+ years of relevant experience in SRE / Production Support / MonitoringStrong hands-on experience with Prometheus & GrafanaGood understanding of production support and incident managementStrong troubleshooting and analytical skillsExperience with monitoring, alerting and performance analysisGood communication and stakeholder coordination skillsComfortable working independently in a remote environmentImmediate Joiners Only