Senior Site Reliability Engineer (5+ years)
We are seeking an experienced Senior Site Reliability Engineer to take a leadership role in optimizing our system observability and infrastructure performance. The successful candidate will drive peak readiness strategies, implement robust monitoring solutions, and collaborate closely with cross-functional teams to ensure system stability at scale.
Key Responsibilities
Lead the design, implementation, and optimization of enterprise-grade monitoring solutions for real-time system visibility.
Collaborate with the engineering team to troubleshoot complex system issues and implement scalable infrastructure solutions.
Manage end-to-end change management processes to improve system reliability with minimal disruption.
Develop and execute comprehensive strategies to ensure system performance and peak readiness during high-traffic periods.
Enhance system observability by implementing best-in-class logging, tracing, and alerting frameworks.
Required Skills & Experience
Minimum of 5 years of professional experience in Site Reliability Engineering.
Demonstrated leadership experience in mentoring technical teams and managing complex infrastructure projects.
Deep technical expertise in monitoring tools, change management protocols, and observability best practices.
Proven ability to develop strategies for handling high-load scenarios and peak traffic.
Strong problem-solving, communication, and innovative thinking skills.
Bachelor's degree in Computer Science, Information Technology, or a related field.
Relevant certifications such as AWS Certified DevOps Engineer or Certified Kubernetes Administrator are highly preferred.