Job detail for Senior Observability Analyst

E
Senior Observability Analyst
Experian
Todayvia fourdayweek

Use AI to assess how you fit

Company Description

Experian is a global data and technology company that powers opportunities for people and businesses around the world. We operate in diverse markets, such as financial services, healthcare, automotive, agribusiness, insurance, among others. Experian invests in people and new advanced technologies to unlock the power of data. We have an incredible team of 25,200 employees in 32 countries.

Our uniqueness is valuing yours. Experian's people-centric, inclusive, and purpose-driven culture is recognized by numerous awards—including World’s Best Workplaces™ 2025 (Fortune's Top 25 global) and Great Place To Work™ in 26 countries, among others. Check out Experian Life on social media or explore our careers site to understand why. Experian is also proud to be an equal opportunity employer and affirmative action employer.

Job Description

About the opportunity

We are looking for a Senior Observability Analyst to lead the evolution of the organization's observability strategy, with Datadog as the primary corporate monitoring and observability platform.

This position will be responsible for defining standards, promoting best practices, and ensuring the implementation of a proactive approach based on metrics, logs, traces, and business indicators. The professional will work on building a reliability-oriented culture, supporting Engineering, Architecture, Development, SRE, and Operations teams in the early identification of risks and the continuous reduction of incidents.

We are looking for someone with strong technical knowledge in Datadog and the ability to transform operational data into concrete actions to improve user experience, application stability, and operational efficiency.

**Key **

  • Act as the technical and functional reference for the Datadog platform within the organization.
  • Design, implement, and evolve observability solutions using APM, Infrastructure Monitoring, Logs Management, Dashboards, Real User Monitoring (RUM), Synthetic Monitoring, and Continuous Testing modules.
  • Define corporate observability standards for applications, APIs, microservices, and cloud workloads.
  • Structure and maintain executive, operational, and analytical dashboards to monitor platform health.
  • Develop proactive monitoring strategies based on SLIs, SLOs, SLAs, and business indicators.
  • Create, review, and optimize intelligent monitors and alerts using advanced Datadog features to reduce operational noise and false positives.
  • Support development teams in application instrumentation using OpenTelemetry and native Datadog integrations.
  • Conduct root cause analysis (RCA) of critical incidents and propose structural actions to prevent recurrences.
  • Identify opportunities for automation, failure prediction, and self-healing initiatives.
  • Develop training, playbooks, standards, and documentation related to the Datadog platform.
  • Perform periodic analyses of capacity, performance, availability, and end-user experience.
  • Act as a transformation agent in disseminating the culture of observability and operational excellence.
Qualifications

Mandatory Qualifications

  • Completed higher education degree in Computer Science, Engineering, Information Systems, or related fields.
  • Advanced experience with Datadog, working in implementation, administration, and platform evolution.
  • Practical experience with:
  • Datadog APM
  • Datadog Infrastructure Monitoring
  • Datadog Logs Management
  • Datadog Dashboards
  • Datadog Monitors
  • Monitoring Service Catalog
  • Advanced knowledge of observability pillars: Logs, Metrics, Distributed Tracing
  • Experience in API and microservices architecture observability.
  • Experience with AWS environments.
  • Experience in troubleshooting and investigating complex incidents.
  • Knowledge of OpenTelemetry and application instrumentation.
  • Knowledge of automation using Python, Shell Script, or PowerShell.
  • Experience with Kubernetes, Docker, and cloud-native ecosystems.
  • Solid knowledge of system availability, performance, scalability, and reliability.
  • Ability to translate technical indicators into business impacts.
  • Excellent communication and ability to work with multiple stakeholders.

Desirable

  • Datadog Certified Associate certification or higher.
  • Experience with SRE (Site Reliability Engineering) practices.
  • Experience in implementing observability strategies for large-scale distributed environments.
  • Knowledge of CI/CD and DevSecOps.
  • Experience with incident response automation and self-healing processes.
  • Experience in operational governance, incident management, Problem Management, and ITIL processes.
  • Previous experience with tools such as Dynatrace, Grafana, Prometheus, Elastic Stack, or Zabbix.

Expected profile

We are looking for someone who has: A passion for observability and system reliability. An investigative profile oriented toward solving complex problems. A strong sense of ownership and operational responsibility. Ability to influence technical teams in adopting best practices. A data-driven and continuous improvement mindset. Systemic vision to correlate technical events with business impacts. A collaborative, consultative profile with the ability to train other teams. Position highlights: Be a protagonist in defining the corporate observability strategy. Act directly on the evolution of the organization's operational maturity. Lead initiatives for automation, intelligent observability, and incident reduction. Influence engineering and reliability standards used by multiple teams. Work in a high-visibility environment with business impact. Participate in building an increasingly proactive, data-driven operation based on operational excellence.

Job summary in one sentence: We are looking for a Datadog expert capable of transforming observability into an operational advantage, reducing incidents, increasing platform reliability, and disseminating a proactive, data-driven monitoring culture.

Additional Information

Serasa Experian is much more than you imagine. With the purpose of creating a better future by expanding opportunities for people and businesses, in Brazil we are more than 4,000 people working in diverse teams and specialties. Here, every piece of knowledge and diversity complements each other and you can work on what you love most. We are committed to building an inclusive culture and an environment in which people can balance their careers with their personal commitments and interests, valuing well-being.

We are very dedicated to being one of the best and most innovative companies to work for in the country, enabling incredible experiences and careers for our people. Our strong people-first approach is recognized externally through various market certifications: we have been awarded by Great Place To Work™ in 24 countries and by the international Top Employers certification, in addition to being recognized as one of the best companies for young professionals and having a 4.6 rating on Glassdoor. Each recognition indicates that we are on the right track, providing an increasingly better work environment for our talent.

#LI-HOME

Experian Careers - Creating a better tomorrow together

Find out what its like to work for Experian by clicking here