Principal Cloud Engineer - Site Reliability Infrastructure at Couchbase | Torre
Principal Cloud Engineer - Site Reliability Infrastructure
Report

Principal Cloud Engineer - Site Reliability Infrastructure

Emma highlights
This highlight was written by Emma’s AI. Ask Emma to edit it.
Full-time

Legal agreement: Employment

Compensation
USD100k - 160k/year
location_on
Hybrid (Canada)
Hybrid (United States)
Hybrid (United Kingdom)
Hybrid (India)
Match
skeleton-gauges
You have opted out of job matches in .
To undo this, go to the 'Skills and Interests' section of your preferences.
Review preferences
Posted over 5 years ago

Requirements and responsibilities


• Design, creation, and provisioning of infrastructure. • Deploy and maintain applications. • Design, build, manage and operate the infrastructure and configuration of SaaS applications with a focus on automation and infrastructure as code. • Design, build, manage and operate the infrastructure as a service layer (hosted and cloud-based platforms) that supports the different platform services. • Develop comprehensive monitoring solutions to provide full visibility to the different platform components using tools and services like Kubernetes, Prometheus, Grafana, ELK, Datadog, New Relic and other similar tools. • Experience working within an Agile/Scrum SDLC • Integrate different components and develop new services with a focus on open source to allow a minimal friction developer interaction with the platform and application services. • Identify and troubleshoot any availability and performance issues at multiple layers of deployment, from hardware, operating environment, network, and application. • Evaluate performance trends and expected changes in demand and capacity, and establish the appropriate scalability plans • Troubleshoot and solve customer issues on production deployments • Ensure that SLAs are met in executing operational tasks • Collaborate with other engineers to implement operational solutions while defining, adhering to industry best practices. • Experience in Building and managing Virtualized systems (KVM, OVM, Containers/Docker) and ability to read and understand source code • Systematic problem-solving approach, combined with a strong sense of ownership and drive. • Conduct periodic on-call duties • Working knowledge of information security issues • Firm grasp of at least one modern programming language, beyond advanced scripting (Shell, Perl, Python) • Solid experience using configuration management frameworks (e.g. Chef, Puppet) • Working knowledge of web and network protocols and standards (HTTP, TLS, DNS, etc)