Back to jobs
New

Site Reliability Engineer

New York, New York

About Onyx.

Onyx is building the market for people who enjoy being right. Our consumer platform lets people take positions on what happens next across sports, prediction markets, crypto, financial markets, politics, culture, and real-world events.

More than one million users have already found Onyx, driving over $2 billion in monthly notional volume across our social sports platform. Now, we’re bringing that same energy into regulated financial markets through Onyx Predictions, a CFTC-registered introducing broker and NFA member.

In June 2026, we raised a $20 million Series A led by Payward, the parent company of Kraken, valuing Onyx at $220 million less than one year after emerging from beta. Our team comes from Jane Street, DraftKings, Robinhood, xAI, SIG, HRT, and Harvard.

We’re building Onyx in New York with a small team, serious momentum, and no interest in moving slowly. The product is live. The market is growing. The category is still ours to define.

The Role.

We’re hiring a Site Reliability Engineer to build the infrastructure, systems, and operational practices that keep Onyx fast and reliable as we scale.

Markets move in real time. Traffic spikes without warning. Customers expect orders, balances, positions, and settlement to work every time. You’ll help ensure our platform is ready for all of it.

Our infrastructure runs primarily on AWS and is managed using Terraform, with Datadog supporting observability and incident response. You’ll work directly with application engineers, product, trading, security, and operations to improve reliability, automate infrastructure, strengthen production systems, and respond when something goes wrong.

This is a high-ownership role. You won’t just monitor infrastructure or maintain someone else’s playbook. You’ll help define how reliability engineering works at Onyx.

What You’ll Own:

  • Build, operate, and improve scalable production infrastructure in AWS.
  • Manage cloud infrastructure through Terraform and infrastructure as code.
  • Strengthen monitoring, logging, tracing, alerting, and observability using Datadog.
  • Improve the availability, performance, resilience, and security of critical services.
  • Build tooling and automation that make infrastructure safer and easier to operate.
  • Improve deployment systems, release processes, environment management, and rollback capabilities.
  • Partner with engineers to design reliable, observable, and scalable services.
  • Monitor production systems and lead incident investigation and resolution.
  • Build runbooks, escalation processes, and post-incident review practices.
  • Improve database reliability, backups, disaster recovery, and capacity planning.
  • Participate in an on-call rotation and provide support during high-volume market events.

You’ll Do Well Here If You:

  • Have experience operating production systems in AWS.
  • Are proficient with Terraform and infrastructure as code.
  • Have hands-on experience with observability, incident response, and production troubleshooting.
  • Understand distributed systems, networking, databases, and cloud architecture.
  • Have built automation using Python, TypeScript, Go, Bash, or a similar language.
  • Understand deployment pipelines, release strategies, and rollback procedures.
  • Can balance long-term infrastructure improvements with immediate production needs.
  • Remain calm during incidents and take ownership from the first alert through the permanent fix.

Even Better If You Have:

  • Operated infrastructure for trading, crypto, payments, gaming, or consumer fintech.
  • Supported real-time, transactional, or highly concurrent systems.
  • Worked with PostgreSQL in a high-availability production environment.
  • Experience with Docker, ECS, EKS, Kubernetes, or similar container systems.
  • Built CI/CD pipelines and automated deployment processes.
  • Worked with event-driven systems, WebSockets, streaming platforms, or live data feeds.
  • Experience with disaster recovery, IAM, secrets management, or cloud security.
  • Established SRE practices or scaled infrastructure at an early-stage company.

Location and Schedule:

  • This role is based in our NoHo, Manhattan office, with core in-office days Tuesday through Thursday.
  • This position includes participation in an on-call rotation and may require support during major sporting events, market-moving news, or production incidents.
  • Candidates must be authorized to work in the United States. Onyx is unable to provide employment sponsorship for this position.

Compensation and Benefits:

  • The anticipated base salary range for this position is $145,000 - $190,000 depending on experience, impact, and scope, plus target bonus and equity.

Benefits include:

  • Fully covered medical, dental, and vision insurance.
  • Meaningful equity ownership.
  • Commuter benefits.

Why Onyx?

Onyx is for people with conviction.

We hire talented people, give them meaningful ownership, and expect them to use it. Our teams stay small, the work stays visible, and the distance between an idea and production stays short.

Onyx moves quickly because our markets move quickly. Titles matter less than judgment, initiative, and the quality of what you ship. Good ideas can come from anywhere, but the person willing to make the call is expected to follow it through.

We value people who are ambitious without ego, intellectually honest, comfortable with uncertainty, and energized by hard problems. We debate openly, decide quickly, and take responsibility for the result.

Come build what’s next with us.

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf