Senior Site Reliability Engineer (Sre) – Application Observability & Readiness (Azure)
Encora10
Mexico
Main Responsibilities
- Collaborate with development teams to design and implement monitoring, alerting, dashboards, and APM instrumentation across applications and services.
- Lead the implementation, configuration, and optimization of Application Performance Monitoring (APM) solutions.
- Apply observability best practices using tools such as Azure Monitor, Application Insights, New Relic, and Log Analytics (KQL).
- Enable code-level instrumentation, distributed tracing, and structured logging to improve application visibility and reliability.
- Design and maintain application-level monitoring dashboards and operational health metrics.
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and effective alerting strategies based on latency, error rates, traffic, and resource saturation.
- Continuously improve monitoring and alerting mechanisms through production insights and incident learnings.
- Participate in production readiness reviews, identifying operational risks, observability gaps, and potential failure scenarios before deployment.
- Support incident analysis and post-incident improvements through enhanced telemetry and monitoring practices.
- Partner with engineering teams to ensure applications are reliable, scalable, and production-ready.
Mandatory Requirements
- Strong experience supporting and operating applications in Microsoft Azure IaaS environments.
- Hands-on experience with application observability, monitoring, and reliability engineering practices.
- Mandatory experience with DBT, Databricks, and SQL (minimum 1 year of experience).
- Experience implementing and managing APM solutions such as Application Insights, New Relic, or similar platforms.
- Experience designing dashboards and monitoring solutions using Azure Monitor, Application Insights, and Log Analytics (KQL).
- Familiarity with CI/CD environments including Azure DevOps and GitHub Actions.
- Solid understanding of cloud-native architectures and distributed application systems.
- Practical SRE mindset with experience in incident analysis, root cause investigation, and proactive problem prevention.
- Strong verbal and written English communication skills, with the ability to collaborate effectively with global teams.
Preferred Requirements
- Experience with scripting and automation using PowerShell and/or Bash.
- Knowledge of scalability, availability, and resilience patterns in modern cloud environments.
- Experience driving production readiness and operational excellence initiatives.
- Exposure to reliability engineering best practices in enterprise-scale environments.
Similar jobs
Senior Cloud Architect
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As an AWS Pr...
Retail Media Search Account Executive
Pacvue is establishing a local presence in Mexico, building out a new hub in Mexico City, where we plan to open an office soon. Roles wil...
Senior Retail Media Search Account Executive
Pacvue is establishing a local presence in Mexico, building out a new hub in Mexico City, where we plan to open an office soon. Roles wil...
Senior Mobile Engineer (Ai-Assisted)
About Goods & Services Goods & Services is a product design and engineering company. We solve mission-critical challenges for some of the...
Senior Engineer
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As an AWS Pr...
Senior Test Environment Manager
Job Title: Senior Test Environment ManagerKey Skills: Test Environment Management, Test Data Management, Environment Planning, Release Ma...