Senior Site Reliability Engineer
Who We Are
At 2K, we create some of the most iconic and culture-shaping video games in entertainment, including NBA® 2K, one of the top-selling franchises in the world, and legendary titles like BioShock®, Borderlands®, Mafia, Sid Meier’s Civilization®, and XCOM®, as well as fan favorites WWE® 2K, TopSpin®, and PGA TOUR® 2K. We build unforgettable experiences by pushing the boundaries of creativity, authenticity and innovation across every genre.
Our portfolio is brought to life by some of the most influential game development studios in the world. Visual Concepts, Firaxis Games, Hangar 13, Cat Daddy Games, 31st Union, Cloud Chamber, Gearbox, HB Studios, and 2K SportsLab create world-class experiences across platforms.
But what truly powers 2K is our people.
We believe the best ideas come from teams that feel empowered, supported, and inspired. As an equal opportunity employer, we are committed to fostering a diverse, inclusive workplace where people are encouraged to come as they are and do their best work.
What We Need
The 2K SRE team owns the infrastructure behind every player connection. That covers all 2K game services, account platforms, CI/CD pipelines, databases and developer tooling, running on AWS and on VMware-based data centers in several global regions. Global launches and live-service events push these systems to their limits, and this team is expected to keep them running. Post-mortems here focus on systems, not people. Automation is the default answer to repetitive work. The infrastructure keeps millions of players connected, and the team takes that seriously.
The Senior SRE at 2K is a hands-on technical leader. You'll shape production infrastructure across AWS and on-premises VMware environments, working with network engineers, database administrators, systems architects and game studio developers. This is an ownership role. You'll set technical direction, improve reliability from architecture review through production, and close the gap between what engineering ships and what players experience.
Duties
Platform & Infrastructure
-
Design, build and operate scalable hybrid infrastructure across AWS and VMware vSphere, using Terraform for provisioning and Puppet for configuration management.
-
Own the full lifecycle of EC2 fleets and VMware clusters (ESXi, vCenter, vSAN, NSX): capacity planning, patching, image pipelines (Packer, AMIs, VM templates) and autoscaling.
-
Architect secure, resilient AWS networking: VPCs, Transit Gateway, Direct Connect, Route 53, ALB/NLB and CloudFront.
-
Operate and automate F5 BIG-IP load balancers and firewalls (LTM, AFM, DNS/GTM): virtual servers, pools, health monitors, iRules, SSL offload and global traffic management across data centers and AWS.
-
Use blue/green and canary releases for game service deployments, with tools like CodeDeploy, weighted target groups, F5 pool weighting and DNS-based traffic shifting.
Databases (MySQL & PostgreSQL)
-
Operate MySQL and PostgreSQL at production scale, both self-managed on VMware and EC2 and managed on Amazon RDS and Aurora.
-
Own replication topologies, failover, backup and point-in-time recovery, and disaster recovery testing.
-
Lead major-version upgrades, schema migration tooling and zero-downtime maintenance procedures.
-
Tune query performance, connection pooling (ProxySQL, PgBouncer) and resource sizing for high-concurrency game workloads.
Observability & Reliability
-
Build and run the full observability stack: Prometheus, Grafana, Datadog and OpenTelemetry, including database, VMware and F5 telemetry.
-
Define SLI/SLO/error budget policies and build alerting that cuts through the noise.
-
Lead chaos engineering and failover exercises (for example, AWS Fault Injection Service, F5 HA failover and database failover drills) to find failure modes before players do.
-
Drive incident response and post-mortems, with a focus on systemic fixes and real follow-through.
Automation, Security & Developer Experience
-
Eliminate toil through self-service provisioning, automated remediation and intelligent scaling.
-
Manage server configuration at scale with Puppet, Ansible and AWS Systems Manager, including module development, Hiera data design and code promotion workflows.
-
Automate F5 configuration through AS3, iControl REST and Terraform, and treat load balancer and firewall policy as code.
-
Harden CI/CD pipelines in GitHub Actions and Jenkins.
-
Build security into the platform layer through secrets management (PasswordState, 1Password, AWS Secrets Manager), IAM least privilege, F5 AFM firewall policy and policy-as-code (OPA, AWS Config, SCPs).
Leadership
-
Promote SRE practices across 2K studios through reliability reviews, runbooks and embedded collaboration.
-
Shape architectural decisions and write engineering RFCs that move the platform forward.
Must-Haves
-
5+ years in SRE, Platform Engineering or equivalent infrastructure work at production scale
-
Deep AWS experience: EC2, VPC networking, IAM, RDS/Aurora, S3, Route 53, ELB and CloudWatch
-
Strong VMware vSphere experience: ESXi, vCenter, vSAN and/or NSX, plus bare-metal server operations
-
Hands-on experience with F5 BIG-IP load balancers and firewalls (LTM, AFM, DNS/GTM), including iRules, SSL/TLS offload and HA pairs
-
Production experience running MySQL and PostgreSQL, including replication, HA/failover, backup and restore, upgrades and performance tuning
-
Strong IaC skills with Terraform; hands-on with Packer
-
Deep Puppet experience (modules, Hiera, r10k/Code Manager, PuppetDB), plus Ansible and AWS Systems Manager
-
Experience with observability tools: Datadog, Prometheus, Grafana and OpenTelemetry
-
Fluency with SLIs, SLOs and error budgets, including how to put them into practice inside engineering teams
-
Production-quality code in Go, Python or TypeScript for tools, automation and internal libraries
-
Knowledge of Linux internals, TCP/IP networking, DNS and TLS deep enough to debug at the system level
-
Incident response and post-mortem leadership, with a track record of systemic follow-through
Nice-to-Haves
-
Live-service game or large-scale consumer internet experience with millions of concurrent users
-
Database reliability engineering at scale: sharding, read-replica fleets and Aurora Global Database or cross-region replication
-
F5 automation experience (AS3, Declarative Onboarding, iControl REST) and multi-site global load balancing
-
Experience migrating workloads between VMware and AWS (for example, AWS Application Migration Service or VMware HCX)
-
FinOps and managing resources at cloud scale
-
Experience with AI and agentic development
-
Certifications such as AWS Solutions Architect, AWS Database Specialty, VMware VCP-DCV, F5 Certified Administrator, Puppet Professional or equivalent
-
Experience mentoring SREs or leading reliability working groups
As an equal opportunity employer, we are committed to ensuring that qualified individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform their essential job functions, and to receive other benefits and privileges of employment. Please contact us if you need reasonable accommodation.
Please note that 2K Games and its studios never uses instant messaging apps or personal email accounts to contact prospective employees or conduct interviews and when emailing, only use 2K.com accounts.
Please note that 2K Publishing is unable to provide visa sponsorship or assistance for this position. All candidates must be legally authorized to work in the United States without requiring current or future employer sponsorship.
#LI-Hybrid
Create a Job Alert
Interested in building your career at 2K? Get future opportunities sent straight to your email.
Apply for this job
*
indicates a required field
