Start Your Search Here

Job Search

IBM

Toronto / Global

Expert Reliability Engineering Role at IBM

  • Hybrid

Job Description

At IBM Software, take on an Expert Reliability Engineering role focused on incident command in multi-cloud settings. Drive improvements to enhance system reliability and incident response capabilities.

This hybrid role combines deep engineering work with strategic oversight. You'll engage in hands-on projects such as improving tooling and analyzing failures while also mentoring teams in incident response. Your goal will be to enhance reliability across IBM’s Cloud services.

Key Responsibilities:

• Analyze and design improvements to prevent incidents

• Oversee Rootly configurations and related integrations

• Maintain SLO/SLA standards for incident management

• Edit customer-facing incident documentation for clarity

• Coach teams through incident post-mortems and training

Requirements:

• 10+ years in SRE or reliability-focused roles

• Cloud expertise in major platforms like AWS, GCP, Azure

• Familiarity with incident management tools like PagerDuty

• In-depth knowledge of distributed systems performance

• Strong written and verbal communication skills

Join IBM Software to drive impactful changes in incident management and shape the future of technology reliability.

#J-18808-Ljbffr

Apply Now

Similar Opportunities

View all jobs