Senior Cloud Engineer
About impact.com
impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey. From affiliates and influencers to content publishers, brand ambassadors, and customer advocates, impact.com empowers brands to drive trusted, performance-based growth through authentic relationships. Its award-winning products—Performance (affiliate), Creator (influencer), and Advocate (customer referral)—unify every type of partner into one integrated platform. As consumers increasingly rely on recommendations from people and communities they trust, impact.com helps brands show up where it matters most. Today, over 5,000 global brands, including Walmart, Uber, Shopify, Lenovo, L’Oréal, and Fanatics, rely on impact.com to power more than 225,000 partnerships that deliver measurable business results.
Position Overview
As impact.com scales, we require a Senior Cloud Engineer to serve as a technical specialist for our infrastructure substrate. You will be part of a specialised squad (IaC & Provisioning,
Observability, SRE, SecOps, or FinOps) focused on ensuring that production environments are stable, observable, secure, and cost-effective. In this role, you treat the infrastructure layer as a capability provider to the rest of the engineering organisation. You are responsible for executing the technical roadmap, covering everything from IaC evolution and GCP/AWS Kubernetes and compute management to vulnerability remediation and observability standards.
Key Responsibilities
Infrastructure & Compute Engineering
- Author, maintain, and version high-quality IaC modules (Terraform) for GCP and AWS services, ensuring they are consumable by platform engineering.
- Deploy and maintain foundational networking topology, VPCs, DNS, and IAM primitives.
- Manage the lifecycle of GKE/EKS clusters and compute instances, including the maintenance of secure base images and container standards.
- Drive drift detection and remediation, ensuring automated scans run daily with remediation completed within 24 hours.
Reliability & Observability
- Maintain observability backends (metrics pipelines, log storage, and alerting) to ensure infrastructure-layer monitoring standards are met.
- Ensure log pipeline uptime of 99.9% and maintain an alert accuracy rate of < 5% false positives for P1/P2 alerts.
- Participate in stress testing and capacity planning to ensure infrastructure can scale safely.
- Reduce toil through automation, contributing to at least one major toil reduction initiative per quarter.
Incident Response & SecOps
- Participate in the infrastructure on-call rotation, acknowledging P1 incidents within 5 minutes and aiming for an MTTR of < 30 minutes.
- Execute vulnerability remediation and security patching on advisories from the SOC: Critical (48 hours) and High (7 days).
- Ensure patch compliance for > 95% of systems and conduct quarterly access reviews.
- Draft and publish blameless post-mortems within 5 working days of resolving infrastructure incidents.
Cloud Economics (FinOps)
- Support cloud spend governance by identifying rightsizing opportunities and reserved instance strategies.
- Flag regional spend anomalies (> 10% deviation) within 24 hours of detection.
Required Experience & Skills
- 5+ years of experience in cloud infrastructure, SRE, or platform operations.
- Hands-on experience operating infrastructure at scale in GCP (preferred) and AWS.
- Proficiency with Terraform, Kubernetes (GKE), and observability stacks like Prometheus, Grafana, or Open Telemetry.
- Experience with security patching, access controls, and Google SREprinciples (error budgets, SLOs).
Key Performance Indicators (KPIs)
- Maintain 99.99% availability for cloud infrastructure per year.
- Target > 90% IaC module coverage for all provisioned infrastructure.
- Acknowledge SOC advisories within 4 hours and triage within 1 working day.
- Delivery of new standard environments within 2 working days.
Role Boundaries
To maintain focus, this role does not own:
- Application-level business logic or shared libraries.
- Developer-facing observability standards or CI/CD deployment tooling.
- Security policy and risk management frameworks (owned by the Security department).
Benefits:
- Hybrid, Casual work environment
- Unlimited PTO policy
- Take the time off that you need. We are truly committed to a positive work-life balance, recognising that it is important to be happy and fulfilled in both
- Training & Development
- Learning the advanced partnership automation products
- Medical Aid and Provident Fund
- Group schemes with Discovery & Bonitas for medical aid
- Group scheme with 10x for provident fund
- Restricted Stock Units
- 3-year vesting schedule pending Board approval
- Internet Allowance
- Fitness club fee reimbursements
- Technology Stipened
- Primary Caregiver Leave
- Mental Health and Wellness Benefit - Including 12 Therapy/Coaching sessions + Dependent coverage
impact.com is proud to be an equal opportunity workplace. All employees and applicants for employment shall be given fair treatment and equal employment opportunity regardless of their race, ethnicity or ancestry, color or caste, religion or belief, age, sex (including gender identity, gender reassignment, sexual orientation, pregnancy/maternity), national origin, weight, neurodivergence, disability, marital and civil partnership status, caregiving status, veteran status, genetic information, political affiliation, or other prohibited non-merit factors.
#LI-Hybrid
Create a Job Alert
Interested in building your career at Impact.com? Get future opportunities sent straight to your email.
Apply for this job
*
indicates a required field