Senior Lead, Site Reliability Engineering (SRE) at BMO in Toronto, ON
- Company: BMO
- Location: Toronto, ON, CAN
- Salary: $70K – $150K
- Job type: full time
- Workplace: onsite
- Posted: 2026-09-29
Job description
Application Deadline: 10/30/2026 Address: 4100 Gordon Baker Road Job Family Group: Technology We are seeking a highly motivated and technically strong Senior Lead, Site Reliability Engineering (SRE) to provide technical leadership within the DCOE Reliability Engineering organization. This role is critical to advancing platform reliability, operational resilience, automation, and AI-driven operations across a portfolio of enterprise platforms. The successful candidate will serve as a senior technical leader and trusted advisor, driving reliability engineering practices, modernization, automation, observability, and operational excellence across the organization. The role requires deep expertise in SRE principles, cloud and infrastructure operations, AIOps, automation engineering, service resilience, and continuous improvement. This position is ideal for an experienced engineering leader who excels through influence, technical expertise, collaboration, and execution rather than direct people management. Key Responsibilities Reliability Engineering & Operational Excellence Lead and advance SRE practices across critical production platforms. Define, establish, and monitor platform reliability objectives, including availability, resiliency, performance, and recovery metrics. Champion improvements in incident management, problem management, capacity planning, and operational readiness. Develop proactive reliability strategies that reduce operational risk and improve customer experience. Serve as a subject matter expert in reliability engineering, providing guidance and recommendations across technology teams. Automation & AIOps Lead automation initiatives that reduce manual effort, eliminate operational toil, and improve service reliability. Design and implement automated remediation, self-healing capabilities, and operational workflows. Drive adoption of AI-driven operational capabilities to enhance event correlation, anomaly detection, predictive analytics, and intelligent incident response. Champion Infrastructure as Code (IaC), Configuration as Code (CaC), CI/CD integration, and automated governance controls. Develop and enhance enterprise observability solutions leveraging monitoring, logging, tracing, and analytics platforms. Establish standardized monitoring frameworks, operational dashboards, and reliability scorecards. Drive advancements in telemetry, operational insights, and performance engineering. Modernization & Transformation Support cloud modernization, platform simplification, and technology transformation initiatives. Partner with engineering, architecture, security, and operations teams to embed reliability requirements throughout the software delivery lifecycle. Influence the adoption of engineering standards, best practices, and operating models across enterprise platforms. Act as a key contributor to strategic technology transformation initiatives and reliability-focused roadmaps. Technical Leadership & Engineering Enablement Provide technical leadership, coaching, and mentorship to engineers and technical leads. Foster a culture of accountability, innovation, continuous learning, and operational excellence. Build organizational capability in automation, AI, cloud operations, observability, and reliability engineering. Lead communities of practice and drive knowledge-sharing across engineering teams. Influence senior stakeholders and technology partners to align on reliability, resilience, and modernization priorities. Required Qualifications Education University degree in Computer Science, Engineering, Information Technology, or a related discipline. Relevant industry certifications are considered an asset. Experience 8+ years of progressive experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, Cloud Operations, or DevOps. 3+ years of experience providing technical leadership for enterprise-scale platforms and complex engineering initiatives. Proven experience leading
Senior Lead, Site Reliability Engineering (SRE) on JobPost.