Job detail for Mid-level SRE Analyst

E
Mid-level SRE Analyst
Experian
Todayvia fourdayweek

Use AI to assess how you fit

Company Description

Experian is a global data and technology company that powers opportunities for people and businesses around the world. We operate in diverse markets, such as financial services, healthcare, automotive, agribusiness, insurance, among others. Experian invests in people and new advanced technologies to unlock the power of data. We have an incredible team of 25,200 employees in 32 countries.

Our uniqueness is valuing yours. Experian's people-centric, inclusive, and purpose-driven culture is recognized by numerous awards — including World’s Best Workplaces™ 2025 (Fortune's Top 25 global) and Great Place To Work™ in 26 countries, among others. Check out Experian Life on social media or explore our careers site to understand why. Experian is also proud to be an equal opportunity and affirmative action employer.

Job Description

We are looking for a highly motivated Mid-level Site Reliability Engineer (SRE) to join our Cloud, Data & AI Platform team. In this role, you will be responsible for designing, operating, and continuously improving the reliability, scalability, observability, and performance of cloud-native platforms that support business-critical applications, data pipelines, and AI/ML workloads.

You will work closely with Software Engineering, Data Engineering, AI Engineering, and Platform teams to build resilient systems, automate operations, improve developer experience, and establish reliability best practices across the organization.

As a mid-level engineer, you will play a key role in advancing operational excellence through automation, observability, incident management, cost optimization, and infrastructure modernization.

Key :

  • Design and operate highly available, scalable, and secure cloud platforms on AWS.
  • Build and maintain Kubernetes-based infrastructure to support applications, data, and AI workloads.
  • Improve platform reliability through automation, Infrastructure as Code (IaC), and self-service capabilities.
  • Implement and enhance observability solutions using Datadog, including monitoring, logs, tracing, alerts, dashboards, and SLO management.
  • Support and optimize large-scale data processing environments using Airflow, Amazon EMR, S3, and other AWS data services.
  • Partner with Data and AI teams to improve the reliability, scalability, and operational maturity of Machine Learning and Artificial Intelligence platforms.
  • Lead incident response activities, root cause analysis, and post-incident reviews, promoting continuous improvement.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Enhance deployment processes, CI/CD pipelines, and release reliability.
  • Optimize cloud infrastructure utilization, performance, and costs.
  • Mentor team members and promote SRE best practices across the engineering organization.

What defines success in this role

  • Increased platform availability and reliability.
  • Improved observability and reduced incident resolution time.
  • Greater automation and reduction of manual, repetitive operational efforts.
  • Reliable, scalable, and cost-efficient Data and AI platforms.
  • Strong collaboration with Engineering teams to deliver resilient production systems.
Qualifications

Required Qualifications

  • Higher education in progress or completed.
  • Solid experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps roles.
  • Strong hands-on experience with AWS services and cloud-native architectures.
  • Deep knowledge of Kubernetes and containerized workloads in production environments.
  • Experience managing and troubleshooting large-scale distributed systems.
  • Solid experience with observability platforms, preferably Datadog.
  • Experience supporting platforms and data flows using technologies such as Airflow, EMR, Spark, and S3.
  • Expertise in Infrastructure as Code (IaC) using Terraform or similar tools.
  • Experience building and maintaining CI/CD pipelines and platform automation.
  • Solid knowledge of Linux, networking, and system performance troubleshooting.
  • Proficiency in scripting and automation using Python, Bash, or similar languages.

Desirable Qualifications

  • Experience supporting large-scale cloud-native platforms in AWS environments.
  • Experience with Kubernetes platform operations and cluster lifecycle management.
  • Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and operational excellence practices.
  • Experience implementing observability solutions using tools such as Datadog, Prometheus, Grafana, OpenTelemetry, or similar technologies.
  • Familiarity with data processing and workflow orchestration platforms, such as Airflow, Spark, or EMR.
  • Experience with Infrastructure as Code (IaC) and platform automation practices.
  • AWS, Kubernetes, Terraform, or Datadog certifications.
  • Experience working in large-scale, highly available, or mission-critical corporate environments.
  • Intermediate technical English.
Additional Information

This is an affirmative position for women, as part of our commitment to advancing gender equity in the workplace.

Serasa Experian invests in female leadership and has a partnership with the TODAS Group. Additionally, we adhere to the 'Elas Lideram 2030' movement and have joined UN Women to reduce gender inequality by 2030!!

We also have the Women in Experian group, which seeks to improve gender equity for women by creating development opportunities, different perspectives, skills, among others.

Come be part of this transformation!

If you do not identify with this group, we invite you to explore other opportunities on our careers page. We have several vacancies that may connect with your profile and interests.

#LI-HOME

Experian Careers - Creating a better tomorrow together

Find out what its like to work for Experian by clicking here