Job detail for Site Reliability Engineering Lead
Use AI to assess how you fit
Company Description
Experian is a global data and technology company, powering opportunities for people and businesses around the world. We operate across a range of markets, from financial services to healthcare, automotive, agribusiness, insurance, and many more. Experian invests in people and new advanced technologies to unlock the power of data. We have an amazing team of 25,200 people in 32 countries.
Job Description
- Leadership & Strategy
- Define and implement SRE best practices across the organization.
- Proven expertise in production support, engineering, disaster recovery (DCR), automation, and cloud operations
- Mentor and guide a team of SREs, fostering growth
- Collaborate with senior stakeholders to align reliability goals with business objectives.
- Reliability & Performance
- Establish SLIs, SLOs, and SLAs for critical services and ensure adherence.
- Drive initiatives to improve system and reduce operational toil.
- Excellent in designing systems that detect and remediate issues without manual intervention – Self Healing systems, Runbook automation
- Exposure to tools like Gremlin, Chaos Monkey, AWS FIS to simulate outages and improve fault tolerance
- Incident Management
- Act as the primary point of escalation for critical production issues and lead major incident response, root cause analysis, and postmortems.
- Perform detailed post-incident investigations to identify underlying causes. Document findings and share learnings to prevent recurrence.
- Implement preventive measures and continuous improvement processes.
- Observability
- Champion monitoring, logging, and alerting strategies using tools like Prometheus, Grafana, ELK, and AWS CloudWatch.
- Build real-time dashboards to visualize system health and reliability metrics.
- Configure intelligent alerting based on anomaly detection and thresholds.
- Combine metrics, logs, and traces to enable root cause analysis and reduce Mean Time to Resolution (MTTR).
- Knowledge of AIOps or ML-based anomaly detection for proactive reliability management.
- Collaboration
- Work closely with development teams to integrate reliability into application design and deployment
- Promote a culture of shared responsibility for uptime and performance across engineering teams.
Qualifications
- Qualified with a degree in B.Sc. in Computer Science, MCA in Computer Science, Bachelor of Technology in Engineering, or higher
- Hands on technologist with minimum 12 years of experience working in software development with at least 5 years of experience leading an SRE team currently
- Deep expertise with various AWS services. Advanced knowledge of monitoring and observability tools.
- Proven track record of building secure, mission-critical, high-volume transaction web-based software systems, in regulated environments (finance and insurance industries).
Additional Information
Our uniqueness is that we celebrate yours. Experian's people first, inclusive and purpose driven culture is multi award-winning; World's Best Workplaces™ 2025 (Fortune Global Top 25), Great Place To Work™ in 26 countries to name a few. Check out Experian Life on social or explore our Careers Site to understand why. Experian is also proud to be an Equal Opportunity and Affirmative Action employer. If you have a disability or special need that requires accommodation, please let us know at the earliest opportunity.
Experian Careers - Creating a better tomorrow together
Recruitment Fraud Awareness - Experian's recruitment process is conducted only through authorised channels. Recruitment communications will only be sent from an @experian.com email address. Experian will never ask candidates to make any payment as part of an application, interview, assessment, onboarding, or recruitment process. To apply for roles or verify opportunities, please visit experian.com/careers.
Experian Careers - Creating a better tomorrow together
Find out what its like to work for Experian by clicking here