Senior Site Reliability Engineer (Sre) – Application Observability & Readiness (Azure)
Encora10
Latin America
Main Responsibilities
- Collaborate with development teams to design and implement monitoring, alerting, dashboards, and APM instrumentation across applications and services.
- Lead the implementation, configuration, and optimization of Application Performance Monitoring (APM) solutions.
- Apply observability best practices using tools such as Azure Monitor, Application Insights, New Relic, and Log Analytics (KQL).
- Enable code-level instrumentation, distributed tracing, and structured logging to improve application visibility and reliability.
- Design and maintain application-level monitoring dashboards and operational health metrics.
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and effective alerting strategies based on latency, error rates, traffic, and resource saturation.
- Continuously improve monitoring and alerting mechanisms through production insights and incident learnings.
- Participate in production readiness reviews, identifying operational risks, observability gaps, and potential failure scenarios before deployment.
- Support incident analysis and post-incident improvements through enhanced telemetry and monitoring practices.
- Partner with engineering teams to ensure applications are reliable, scalable, and production-ready.
Mandatory Requirements
- Strong experience supporting and operating applications in Microsoft Azure IaaS environments.
- Hands-on experience with application observability, monitoring, and reliability engineering practices.
- Mandatory experience with DBT, Databricks, and SQL (minimum 1 year of experience).
- Experience implementing and managing APM solutions such as Application Insights, New Relic, or similar platforms.
- Experience designing dashboards and monitoring solutions using Azure Monitor, Application Insights, and Log Analytics (KQL).
- Familiarity with CI/CD environments including Azure DevOps and GitHub Actions.
- Solid understanding of cloud-native architectures and distributed application systems.
- Practical SRE mindset with experience in incident analysis, root cause investigation, and proactive problem prevention.
- Strong verbal and written English communication skills, with the ability to collaborate effectively with global teams.
Preferred Requirements
- Experience with scripting and automation using PowerShell and/or Bash.
- Knowledge of scalability, availability, and resilience patterns in modern cloud environments.
- Experience driving production readiness and operational excellence initiatives.
- Exposure to reliability engineering best practices in enterprise-scale environments.
Similar jobs
Senior Machine Learning Engineer
Job Title: Senior Machine Learning EngineerKey Skills: Python, SQL, PySpark, Machine Learning, Scikit-learn, PyTorch, XGBoost, TensorFlow...
Senior Full Stack Ai Engineer
Job Title:Senior Full Stack AI Engineer (Solution Architecture Experience)Key Skills:Solution Architecture, AI/ML, Generative AI, Agentic...
Senior Macos Developer
Job Title:Senior MacOS DeveloperKey Skills:Swift, Xcode, macOS Development, iOS Development, SwiftUI, AppKit, Authentication, OAuth 2.0, ...
Senior Data Engineer
About DonorboxDonorbox is a leading fundraising platform and donor management system for nonprofit organizations. Our mission is to accel...
Senior Data Engineer
Senior Data EngineerLocation/RegionLatin America | 100% RemoteAbout CodeRoadCodeRoad provides end-to-end software development services, h...
Team Lead, Deal Desk (Enterprise)
This is AdyenAdyen provides payments, data, and financial products in a single solution for customers like Meta, Uber, H&M, and Microsoft...