Mgr. SRE
LivePerson (NASDAQ:LPSN) is a leading customer engagement company, creating digital experiences powered by Curiously Human AI. Every person is unique, and our technology makes it possible for companies, including leading brands like HSBC, Orange, and GM Financial, to treat their audiences that way at scale. Nearly a billion conversational interactions are powered by our Conversational Cloud each month.
You'll be successful at LivePerson if you are excited to build something from the ground up. You excel by finding daily opportunities to grow at the same pace as the technology we're building, and you build partnerships that improve our business. Likewise, you're someone who sees feedback as a chance to learn and grow and believes decisions powered by data are the norm. You care about the wellbeing of others and yourself.
Overview:
LivePerson transforms customer care from voice calls to mobile messaging. Our cloud-based software platform, LiveEngage, allows brands with millions of customers and tens of thousands of care agents to deliver digital experiences at scale. As the market leader in real-time intelligent customer engagement, we are a B2B SaaS company with 20 years of experience and the heart of a startup. We work day in and day out to help our customers live out our mission of creating lasting, meaningful connections with their customers.
The Cloud DevOps team at LivePerson is looking for an SRE Team Lead to lead a team of Site Reliability Engineers responsible for building, operating, and continuously improving highly reliable, scalable, and secure cloud infrastructure and services.
The ideal candidate is a strong technical leader who combines deep experience in Site Reliability Engineering and cloud infrastructure with a passion for developing people and building high-performing engineering teams. This role requires someone who can balance strategic thinking with hands-on technical leadership and is comfortable making decisions in complex and fast-changing environments.
As an SRE Team Lead, you will be responsible for the team's technical direction, execution, operational excellence, and engineering practices. You will work closely with engineering, product, security, networking, and other infrastructure teams to ensure our platforms and services meet the reliability, scalability, security, and performance expectations of our customers.
You will lead initiatives that improve system reliability, eliminate operational toil, strengthen automation and observability, and establish engineering best practices. You will also mentor engineers, provide technical guidance, support career development, and foster a culture of ownership, collaboration, continuous improvement, and operational excellence.
We will provide you with an environment where you can make a meaningful impact, develop talented engineers, and solve complex technical challenges at scale.
Role and Responsibilities:
- Lead, mentor, and develop a team of SREs, fostering a culture of ownership, collaboration, technical excellence, and continuous improvement.
- Provide technical leadership and direction for the design, implementation, and operation of highly available, scalable, secure, and resilient infrastructure and services.
- Own the team's technical roadmap and ensure alignment with broader engineering and business objectives.
- Plan and prioritize team initiatives, balancing new development, reliability improvements, technical debt, operational work, and business priorities.
- Design and maintain cloud infrastructure and services across cloud and hybrid environments, with a strong focus on Google Cloud Platform (GCP).
- Guide the development of automation and infrastructure-as-code solutions using Python, Terraform, Ansible, Bash, and other modern DevOps/SRE technologies.
- Lead the design, deployment, and operation of Kubernetes-based platforms and workloads, including troubleshooting complex production issues.
- Establish and maintain GitOps-based deployment workflows using Kubernetes, Helm, and FluxCD.
- Drive the design, implementation, and continuous improvement of CI/CD pipelines using GitLab CI/CD.
- Establish and improve observability practices using metrics, logs, traces, dashboards, and alerting to provide actionable insights into system health and performance.
- Define and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics for critical services.
- Lead incident response and ensure effective handling of critical production incidents, including root cause analysis and follow-up corrective actions.
- Drive initiatives to reduce operational toil, eliminate recurring incidents, and improve the overall reliability and operational maturity of the platform.
- Partner with software engineering, security, networking, product, and other infrastructure teams to influence architecture and deliver reliable and secure solutions.
- Lead capacity planning, performance analysis, scalability assessments, and reliability reviews for critical systems.
- Establish and promote engineering standards, operational processes, and best practices across the organization.
- Identify technical risks and dependencies and proactively develop mitigation strategies.
- Drive continuous improvement through post-incident reviews, reliability assessments, and engineering retrospectives.
- Support and participate in an on-call rotation and ensure the team has effective processes for handling production incidents.
- Provide technical mentorship and career guidance to engineers, helping them grow their technical expertise and ownership.
- Participate in hiring, interviewing, onboarding, performance management, and development of SRE team members.
- Build strong relationships with stakeholders and communicate technical risks, priorities, and progress clearly to engineering leadership and partner teams.
Requirements:
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 7+ years of experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Systems Engineering, or a related field.
- 2+ years of experience leading, mentoring, or managing engineers in a technical leadership or team lead capacity.
- Strong programming and scripting experience with Python, Bash, or similar languages, with a focus on automation and operational tooling.
- Strong hands-on experience with Google Cloud Platform (GCP) and cloud infrastructure concepts, including networking, IAM, compute, storage, and managed services.
- Extensive experience with Kubernetes and containerization technologies such as Docker.
- Strong experience with Infrastructure as Code, particularly Terraform, and configuration management and automation tools such as Ansible.
- Hands-on experience with GitOps practices and Kubernetes deployment technologies such as Helm and FluxCD.
- Strong experience designing, implementing, and maintaining CI/CD pipelines using GitLab CI/CD.
- Solid understanding of Linux systems administration, networking, DNS, TLS/SSL, authentication, and security fundamentals.
- Hands-on experience with monitoring and observability platforms such as Prometheus, Grafana, Alertmanager, or equivalent technologies.
- Proven experience designing and operating highly available, scalable, and distributed systems in production environments.
- Strong understanding of SRE principles, including SLOs, SLIs, error budgets, incident management, and operational excellence.
- Proven experience leading production incident response, root cause analysis, and post-incident reviews.
- Strong troubleshooting and problem-solving skills, with the ability to diagnose complex issues across multiple layers of a technology stack.
- Demonstrated ability to make sound technical decisions and influence architecture across multiple teams.
- Excellent communication and collaboration skills, with the ability to work effectively with both technical and non-technical stakeholders.
- Proven ability to prioritize competing initiatives and translate business requirements into actionable technical plans.
- Demonstrated experience mentoring engineers and building high-performing engineering teams.
- Strong ownership mindset with the ability to work effectively in a fast-paced, highly collaborative environment.
Nice to Have:
- Experience leading SRE or DevOps teams in a large-scale B2B SaaS environment.
- Experience with service meshes such as Istio.
- Experience with secrets management and security platforms such as HashiCorp Vault.
- Experience with PostgreSQL and other distributed data services.
- Experience with networking, load balancing, DNS, TLS/SSL certificates, and enterprise infrastructure.
- Experience with hybrid cloud or on-premises infrastructure environments.
- Experience developing internal platforms, automation, and self-service tooling for engineering teams.
- Experience driving cloud migrations or large-scale infrastructure modernization initiatives.
- Experience establishing organization-wide SRE practices, reliability standards, and operational processes.
Benefits:
- Health: medical, dental, and vision
- Development: Native AI learning
- Food vouchers
- Multisport card
- 28 days paid leave
- 5 care days
Why you'll love working here:
Your entrepreneurial spirit will be supported. We love team members who chase down their big ideas, become experts, help colleagues, and own their work. These four company values guide our continued, holistic growth as individuals, as teams, and as a global organization. And to further make our point, let's just say we're very proud to be on Fast Company's list of Most Innovative Companies and Newsweek's list of most-loved workplaces.
Belonging at LivePerson:
At LivePerson, people from diverse backgrounds come together to make an impact and be their authentic selves. One way we share and connect is through our employee resource groups such as: Live In Color, LP Proud, and Women In Tech. We are proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable laws, regulations and ordinances. We also consider qualified applicants with criminal histories, consistent with applicable federal, state, and local law.
Apply for this job
*
indicates a required field