Senior IT Cloud Operations Engineering Specialist
Who we are:
Who you are:
We are always looking for amazing talent who can contribute to our growth and deliver results! Geotab is seeking a Senior IT Cloud Operations Engineering Specialist who will lead the evolution of our globally distributed cloud infrastructure, driving automation, AI-powered monitoring, and operational excellence at scale. In this senior role, you will bridge innovation and stability — implementing cutting-edge infrastructure as code practices, self-healing systems, and intelligent cost optimization strategies while ensuring the reliability and security of Geotab’s SaaS platform. If you love cloud infrastructure, automation, and building systems that make a real difference — we would love to hear from you!
What you'll do:
As a Senior IT Cloud Operations Engineering Specialist, your key area of responsibility will be the ongoing development and reinforcement of Geotab’s complex, globally distributed cloud infrastructure. You will handle on-call pages and address operational tickets — approximately five to six per day, primarily provisioning requests — while continuously improving and expanding the team’s PowerShell-based automation stack. You will lead critical infrastructure projects including assuming ownership of the monitoring stack (Prometheus, Grafana, Victoria Metrics), driving end-to-end certificate lifecycle management automation, and implementing GCP cost management strategies to identify and eliminate stale or idle resources. You will also champion the adoption of AI and ML tools across cloud operations — from predictive monitoring to anomaly detection — and mentor teammates on AI-assisted scripting and infrastructure workflows.
To be successful in this role, you will need a builder’s mindset: someone who spots inefficiencies, takes ownership, and ships solutions without being asked. You will work closely with the Cloud Automation, Development, Security, Server Operations, and Solution Engineering teams, navigating complex technical relationships while maintaining clear, proactive communication. The ideal candidate is a self-starter who thrives in a flat, high-autonomy environment and brings a proven track record of taking complex infrastructure problems from diagnosis to resolution.
How you'll make an impact:
-
Implement automated deployment strategies for patches and upgrades to reduce downtime and ensure a seamless update process.
-
Develop predictive monitoring systems using AI/ML algorithms to proactively troubleshoot issues in high-availability environments.
-
Implement new and optimize existing infrastructure as code (IaC) practices to automate server provisioning and configuration, ensuring scalability and rapid deployment.
-
Establish self-healing mechanisms and fault-tolerant architecture to minimize manual intervention during outages and ensure quick recovery.
-
Apply AI-driven cost prediction models to optimize cloud spending, creating dynamic thresholds and proactive cost-saving strategies.
-
Introduce AI-driven anomaly detection systems to provide real-time insights into security controls and tool effectiveness, enhancing the overall security posture.
-
Lead the strategic implementation and integration of AI technologies across cloud operations, establishing best practices, frameworks, and standards for AI adoption.
-
Foster a culture of continuous improvement using collaboration tools to streamline communication and optimize operational processes between internal departments.
What you'll bring to the role:
-
Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
-
7+ years of experience in IT/Cloud Operations with a proven track record leading complex cloud infrastructure projects at a senior level.
-
Mastery of Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE); AWS experience is a plus.
-
Advanced proficiency in PowerShell (primary), Bash, and Python; hands-on experience with Terraform for infrastructure as code.
-
Strong SQL skills across multiple database platforms (MSSQL, MySQL, PostgreSQL, BigQuery) and experience with Prometheus, Grafana, and Victoria Metrics.
-
Associate-level cloud certification (GCP or AWS) required; Professional-level certification highly valued; Google Cloud Generative AI or Professional Machine Learning Engineer certification preferred.
-
Experience with AI/ML platforms (e.g., Vertex AI, BigQuery ML) and generative AI tools (e.g., Claude, Copilot) for operational automation and productivity.
Why job seekers choose Geotab:
Flex working arrangements
Home office reimbursement program
Baby bonus & parental leave top up program
Online learning and networking opportunities
Electric vehicle purchase incentive program
Competitive medical and dental benefits
Retirement savings program
*The above are offered to full-time permanent employees only
How we work:
The annual base salary for this position is the expected annual salary for this role, and may be subject to change. Geotab offers various perks and benefits and other compensation components that an individual may be eligible for. The actual base salary for this position depends on a variety of factors such as but not limited to skills, qualifications, education and overall experience, including the location the applicant lives while performing the job. This also includes equity with other team members and alignment with local market data. All offers of employment are contingent upon proof of eligibility to work and the individual's ability to pass a background check.
Hiring Range
$104,400 - $135,700 CAD
Apply for this job
*
indicates a required field
