Back to jobs
New

Site Reliability Engineer

Remote, US

Here at Ooma we empower people to connect in smarter ways. We do this by creating powerful communication experiences through our cloud-based platform to bring people together at work and at home. Our solutions help small business owners stay connected with their customers and manage their businesses from anywhere. For larger companies we provide customized unified communications solutions to meet their unique needs. At home, we help our customers connect with their loved ones by providing the #1 rated VoIP phone service available. We also provide them with peace of mind through our innovative smart home security solution. At Ooma, all our products and services are priced competitively, because we believe advanced technology should be accessible to all. 

About the Role: 

As a Site Reliability Engineer, you will leverage your extensive expertise in Linux systems, virtualization, containers, Kubernetes clusters, and CI/CD pipelines to ensure the stability and efficiency of our systems, collaborating across teams to implement best practices for infrastructure management, automated deployment, and application performance monitoring. 

Deep on-premises experience is a core requirement, not a secondary consideration. Our production environment runs on our own hardware — large data centers built on hundreds of bare metal servers and VMs, with our own storage, virtualization, and physical network beneath them. You will operate comfortably at the hardware and OS layer while also owning the container and delivery platform on top of it. 

**Location and Onsite Requirement: This role requires onsite work at least once per week at one of our designated data centers in Dallas, TX, Ashburn, VA or San Jose, CA. Candidates must be able to commute regularly to one of these locations. Relocation assistance and reimbursement for routine commuting, travel, or overnight lodging are not available.

What You’ll Do:

  • Provide expert guidance on managing large data centers, including hundreds of bare metal servers and virtual machines (VMs), ensuring optimal configuration and performance.
  • Monitor and troubleshoot system performance, reliability, and availability using modern observability tools and techniques, with strong emphasis on diagnosing and resolving issues in operating systems and bare metal environments.
  • Manage the server hardware lifecycle — firmware and BIOS baselines, out-of-band management, failure triage, vendor and RMA coordination, capacity forecasting, and refresh planning.
  • Administer virtualization platforms and manage storage appliances across all data center locations worldwide.
  • Automate bare metal  and VM provisioning and OS lifecycle at scale.  
  • Work hands-on with data center network infrastructure — VLANs, routing, link aggregation, load balancers, and firewall rules — troubleshooting latency, throughput, and packet loss from the Linux host outward, including multi-site and colocation resiliency. 
  • Design, implement, and maintain scalable, reliable infrastructure using containers, Kubernetes, and microservices architecture, including clusters on bare metal and on-premises VMs. 
  • Oversee configuration management for consistent, reliable releases across environments, using Ansible for system configuration, patch management, and provisioning across data center infrastructure, and eliminating configuration drift across the fleet. 
  • Design and operate high-throughput Kafka clusters for event streaming — topics, partitions, replication, consumer lag monitoring, and disaster recovery across data center infrastructure. 
  • Implement name services and server management practices, including DNS, DHCP, NTP, directory and authentication services, and certificate management. 
  • Collaborate with development teams to influence system design choices and operational policies, and continuously evaluate and integrate new technologies, hardware platforms, and automation approaches to improve operational efficiency and reliability. 
  • Participate in on-call rotations supporting production systems, conduct blameless post-mortems with root cause analysis, and maintain incident response runbooks and procedures. 
  • Create comprehensive technical documentation — runbooks, architectural diagrams, network topology maps, rack elevations, and capacity models — and maintain knowledge bases for operational procedures and best practices. 

Experience We’re Looking For: 

  • 8+ years of experience as an SRE or in a related field, with a strong focus on production systems, containers, microservices, and service delivery, required. A bachelor’s degree in Computer Science, Engineering, or a related field required; an advanced degree is strongly preferred.
  • Must have extensive on-premises data center experience managing large environments with hundreds of bare-metal servers and virtual machines. Experience with multi-site or colocation deployments, cross-site disaster recovery, and hybrid on-premises and cloud trade-offs is a plus.
  • Hands-on server hardware experience is essential, including firmware and BIOS management, out-of-band management, failure diagnosis, vendor coordination, and capacity and refresh planning.
  • Deep knowledge of Linux operating systems—including configuration, performance tuning, and troubleshooting—is expected.
  • A strong understanding of Linux networking concepts and protocols is necessary, along with familiarity with data center networking fundamentals such as VLANs, routing, load balancing, DNS, and DHCP.
  • Demonstrated experience with containers and orchestration technologies, particularly Kubernetes, is required, including running clusters on bare metal or on-premises virtual machines.
  • Must bring extensive experience managing and maintaining CI/CD pipelines and the supporting technologies, including GitOps workflows, Argo CD, and Helm charts.
  • Comprehensive knowledge of observability tools such as Prometheus, the ELK Stack, log collectors, and Grafana is key to success in this role.
  • Experience with configuration management tools—particularly Ansible—is important. Infrastructure as code applied to physical infrastructure using Terraform providers, MAAS, Foreman, or similar technologies, as well as scripting in Python, Bash, or Go, would be advantageous.
  • Proven ability to analyze complex systems, identify bottlenecks, troubleshoot issues, and implement effective solutions is critical. Experience operating Kafka or other high-throughput, stateful distributed systems on premises—including replication and disaster recovery—would be highly valued.
  • Excellent communication and cross-functional collaboration skills are essential, as is a willingness to participate in an on-call rotation and support work in physical data center environments when needed.  #LI-CC1

What We Offer: 

Working at Ooma means being a team player, while allowing your individual voice to come through. And, you'll receive competitive compensation, benefits and generous company perks. 

  • Comprehensive Medical/Dental/Vision insurance for you and eligible dependents
    • HMO, PPO’s or a PPO with a HDHP (including HSA, which Ooma helps fund) 
  • Employer Paid Income Protection Benefits (Basic Life and AD&D, Short- and Long-term disability)
  • FSA Healthcare & Dependent Care
  • Commuter Benefits
  • Voluntary Accident, Critical Illness, Hospital Indemnity and Legal
  • 401(k), including employer match, and Roth
  • Employee Stock Purchase Plan (ESPP)
  • Paid Time off, Sick Time, as well as corporate holidays observed
  • Employee Assistance Program
  • Life Balance benefits with Travel Assistance Services and Identity Theft 
  • Additional Benefits include a Discount Program, Credit Union, Medicare Assistance, etc.

Ooma is an equal-opportunity employer committed to recruiting, employing, retaining, promoting, and otherwise treating all employees on the basis of merit, qualifications, and competence. We do not discriminate on the basis of any trait or characteristic protected by applicable federal, state, or local laws.

We may utilize AI-enabled tools during the hiring process, including for resume review, scheduling, and interview note-taking or transcription. These tools are used solely to support our hiring team; all employment decisions are made by human reviewers.  

Where interviews are recorded or transcribed, candidates will be notified in advance and their consent will be obtained. 

The base salary range for candidates within the United States is listed below. Actual base pay will depend on a variety of factors such as education, skills, experience, specific location, etc. The base pay range is subject to change and may be modified in the future. Regular employees may also be eligible for bonus(es), sales incentive(s) (target included in OTE) and/or stock in the form of Restricted Stock Units (RSUs).

United States Pay Range

$120,000 - $170,000 USD

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf


Select...
How many years of professional experience do you have as a Site Reliability Engineer or in a closely related infrastructure or production engineering role? *
Have you personally managed or supported a production, on-premises data center environment containing hundreds of bare-metal servers and virtual machines? *
Does your hands-on server hardware experience include firmware or BIOS management, out-of-band management, hardware failure diagnosis, and vendor or RMA coordination? *
Select...
This role requires onsite work at least one day per week at one of our data center locations. The company does not provide relocation assistance or cover routine commuting, travel, or overnight lodging expenses associated with meeting this requirement. Are you able to meet this onsite requirement beginning on your start date, with or without reasonable accommodation? *
Select...
Select...
Select...

U.S. Standard Demographic Questions

We invite applicants to share their demographic background. If you choose to complete this survey, your responses may be used to identify areas of improvement in our hiring process.
Select...
Select...
Select...
Select...
Select...
Select...

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in Ooma’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Select...
Select...
Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Select...

Voluntary Self-Identification of Disability

Form CC-305
Page 1 of 1
OMB Control Number 1250-0005
Expires 07/31/2029

Why are you being asked to complete this form?

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

How do you know if you have a disability?

A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to:

  • Alcohol or other substance use disorder (not currently using drugs illegally)
  • Autoimmune disorder, for example, lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS
  • Blind or low vision
  • Cancer (past or present)
  • Cardiovascular or heart disease
  • Celiac disease
  • Cerebral palsy
  • Deaf or serious difficulty hearing
  • Diabetes
  • Disfigurement, for example, disfigurement caused by burns, wounds, accidents, or congenital disorders
  • Epilepsy or other seizure disorder
  • Gastrointestinal disorders, for example, Crohn's Disease, irritable bowel syndrome
  • Intellectual or developmental disability
  • Mental health conditions, for example, depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD
  • Missing limbs or partially missing limbs
  • Mobility impairment, benefiting from the use of a wheelchair, scooter, walker, leg brace(s) and/or other supports
  • Nervous system condition, for example, migraine headaches, Parkinson’s disease, multiple sclerosis (MS)
  • Neurodivergence, for example, attention-deficit/hyperactivity disorder (ADHD), autism spectrum disorder, dyslexia, dyspraxia, other learning disabilities
  • Partial or complete paralysis (any cause)
  • Pulmonary or respiratory conditions, for example, tuberculosis, asthma, emphysema
  • Short stature (dwarfism)
  • Traumatic brain injury
Select...

PUBLIC BURDEN STATEMENT: According to the Paperwork Reduction Act of 1995 no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.