About this role
About the Role
The Senior Platform Engineer will play a critical role within the Enterprise Data & AI Technology organization — one of Scotiabank's most significant enterprise-wide strategic initiatives. This role focuses on building, tuning, and managing infrastructure, DevOps, platform site reliability, monitoring, troubleshooting, and enabling new features on Data & AI platforms in line with the bank's Data & AI strategy. You will work with cross-functional teams including IAM, Network, Cloud Ops, Security, and Client partners to drive integration, process automation, platform enhancements, and delivery of new projects. What You'll Do
-
Guidance and Direction: Provide clear direction to the team, set goals, and keep the team accountable for their deliverables; align team goals with the Azure & Databricks Platform roadmap and enterprise standards.
-
Technical Oversight: Own the technical direction across Azure and Databricks, including Azure networking and security architecture (VNets, Private Endpoints, NSGs, route tables, Azure Firewall), Azure Identity & Access Management (RBAC, PIM), and Databricks platform governance (Unity Catalog, workspace configuration, cluster policies).
-
Quality Assurance: Ensure high quality platform support and adherence to SLAs/SLOs and service objectives.
-
Process Improvements: Continually improve platform processes and SOPs for efficiency and automation; design and develop reusable Terraform modules for Azure native resources and Databricks (clusters, SQL warehouses, Unity Catalog objects) for automated deployments via Terraform Cloud/Enterprise and CI/CD.
-
Customer Relations: Build strong relationships with data engineers, analysts, and platform users; communicate proactively with stakeholders and cross-functional teams to align priorities and drive adoption of platform standards.
-
Advanced Monitoring and Troubleshooting: Troubleshoot performance issues across Databricks jobs, clusters, SQL warehouses, and Azure dependencies; implement Azure Monitor and Log Analytics-based observability with dashboards for cluster/job health, driver/executor metrics, and cost insights; establish proactive alerting and early issue detection.
-
Site Reliability: Analyze, triage, and resolve platform issues and ensure reliability and cost efficiency. What We're Looking For
-
Experience owning technical direction across Azure and Databricks, including Azure networking and security architecture (VNets, Private Endpoints, NSGs, route tables, Azure Firewall), Azure Identity & Access Management (RBAC, PIM), and Databricks governance (Unity Catalog, workspace configuration, cluster policies).
-
Strong ability to align team goals with platform roadmaps and enterprise standards; track record of delivering scalable, reliable platform solutions.
-
Proficiency in designing, developing, and using reusable Terraform modules for Azure native resources and Databricks (clusters, SQL warehouses, Unity Catalog objects) and deploying via Terraform Cloud/Enterprise and CI/CD.
-
Experience implementing observability with Azure Monitor and Log Analytics, building dashboards for cluster/job health, metrics, and cost insights, and setting up proactive alerting.
-
Excellent collaboration with cross-functional teams (Platform, Security, Cloud Ops, Networking, Data Governance) and ability to communicate with data engineers, analysts, and platform users.