Senior Site Reliability Engineer (Sre) – Application Observability & Readiness (Azure)
Encora10
Brazil
Main Responsibilities
- Collaborate with development teams to design and implement monitoring, alerting, dashboards, and APM instrumentation across applications and services.
- Lead the implementation, configuration, and optimization of Application Performance Monitoring (APM) solutions.
- Apply observability best practices using tools such as Azure Monitor, Application Insights, New Relic, and Log Analytics (KQL).
- Enable code-level instrumentation, distributed tracing, and structured logging to improve application visibility and reliability.
- Design and maintain application-level monitoring dashboards and operational health metrics.
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and effective alerting strategies based on latency, error rates, traffic, and resource saturation.
- Continuously improve monitoring and alerting mechanisms through production insights and incident learnings.
- Participate in production readiness reviews, identifying operational risks, observability gaps, and potential failure scenarios before deployment.
- Support incident analysis and post-incident improvements through enhanced telemetry and monitoring practices.
- Partner with engineering teams to ensure applications are reliable, scalable, and production-ready.
Mandatory Requirements
- Strong experience supporting and operating applications in Microsoft Azure IaaS environments.
- Hands-on experience with application observability, monitoring, and reliability engineering practices.
- Mandatory experience with DBT, Databricks, and SQL (minimum 1 year of experience).
- Experience implementing and managing APM solutions such as Application Insights, New Relic, or similar platforms.
- Experience designing dashboards and monitoring solutions using Azure Monitor, Application Insights, and Log Analytics (KQL).
- Familiarity with CI/CD environments including Azure DevOps and GitHub Actions.
- Solid understanding of cloud-native architectures and distributed application systems.
- Practical SRE mindset with experience in incident analysis, root cause investigation, and proactive problem prevention.
- Strong verbal and written English communication skills, with the ability to collaborate effectively with global teams.
Preferred Requirements
- Experience with scripting and automation using PowerShell and/or Bash.
- Knowledge of scalability, availability, and resilience patterns in modern cloud environments.
- Experience driving production readiness and operational excellence initiatives.
- Exposure to reliability engineering best practices in enterprise-scale environments.
Similar jobs
Senior Software Engineer (Typescript, Node.js, React)
Why Join ExadelWe’re an AI-first global tech company with 25+ years of engineering leadership, 2,000+ team members, and 500+ active proje...
Senior Cloud Architect
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As an AWS Pr...
Senior Solutions Engineer
Sobre a BlipA Blip é a maior empresa de inteligência conversacional e mensageria do Brasil, com hubs de inovação e operação no Brasil, Mé...
Senior Engineer
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As an AWS Pr...
Senior Data Engineer (Remote In Brazil)
KnowBe4 empowers the modern workforce to make smarter security decisions every day. Trusted by more than 70,000 organizations worldwide, ...
Senior .Net Developer [Short-Term Contractor]
WANTED: Thinkers, Builders & DreamersFor over 15 years, ArcTouch has created lovable apps, websites, and connected experiences for world-...