Senior Detection & Response Engineer, Cyber Defense
About Nscale
Nscale is building the infrastructure platform for the AI era. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers, reducing the complexity of AI development and helping customers manage cost, innovate rapidly, and operate responsibly.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.
About the Role
We are hiring Detection & Response Engineers to decide what we detect, command the response when something is real, and automate everything in between, across enterprise, cloud, production, data centre, and operational technology environments.
You will work in a three-person regional rotation. Roughly one week in three, you hold the operational slot: taking escalations, deciding fast, and acting. The other two weeks are engineering. You find where our detection coverage is thin and work with our AI agents to close it, build automation that removes human time from the response path, and make sure we are ready for the next incident. When a serious incident lands, you command it.
If you have spent years closing the same ticket every Tuesday—and knowing exactly how to fix it for good, but never having the mandate—this is the job where that is the mandate.
How this function works
Read this section carefully. It is not a standard SOC, and the difference is the whole point.
- Agents hold the first tier. In-house security agents hold level 0 and level 1 around the clock. A contracted managed provider covers weekends and works alerts to conclusion, escalating to us only in defined circumstances. Work reaches you as an escalation with context already attached.
- One week in three on the operational rotation. Each regional team runs a three-person rotation in week-long blocks. One engineer holds the operational slot. The other two do detection engineering, automation, and incident preparedness, and both can jump in when escalations spike.
- Every escalation forks three ways: act, solve, or both. Act is the immediate containment. Solve is the engineering work that stops that alert class firing again: a new or tuned detection, an automation, or a control change. Most escalations are both.
- Your primary operational metric is the share of escalations permanently solved. Not tickets closed, and not mean time to close. If the same alert fires twice, we got it wrong the first time.
- Detection is this team’s job. There is no separate detection engineering team. Detection coverage is owned by this team working with our AI agents: you find the gaps and decide what good looks like, the agents draft the rule logic, and you assess, tune, and decide what goes live.
- Two classes, not five. We sort everything into risk indicator or actionable, using the SSVC model. If you have drowned in a five-tier severity scheme nobody trusted, you will like this.
- Building to 24 hours. We are building follow-the-sun coverage across three regional hubs. Each region hands the queue to the next as its working day ends, so everyone works conventional weekday hours in their own daylight. Until every hub is staffed, the rotation includes out-of-hours on-call. The end state is a team where nobody gets paged at night.
What you'll be doing
Detection engineering
- Map our detection coverage against attacker tradecraft and the real estate, and find where we are blind.
- Turn coverage gaps, incidents, threat research, and recurring escalations into clear detection ideas, and work with our AI agents to build them.
- Assess and tune the detection ideas and rule logic our agents propose, which run as shadow rules before they go live. Judge what each one catches, what it misses, and what its noise will cost.
- Make the human call on promotion. You will rarely write a query from scratch, but nothing goes live without your judgement, and you are accountable for whether it is right.
- Retire or retune live detections that are noisy, stale, or no longer relevant.
- Identify telemetry gaps and make the business case for closing them.
Automation
- Remove recurring toil from the response path: enrichment, triage, containment actions, case handling, and reporting.
- Improve the agents that hold the first tier, so they resolve more on their own and escalate with better context. Every gain here cuts the human time per alert and lets us cover more attack surface.
- Ship new capabilities in our internal security platform, such as response actions, integrations, and agent tooling, so every engineer and agent on the team gets faster, and prove each one works in production.
Incident command
- Command security incidents: set objectives, coordinate responders across engineering, infrastructure, and service owners, and keep leadership informed.
- Make containment calls under time pressure, knowing what you can do immediately, what requires authority, and what could disrupt production if handled incorrectly.
- Run post-incident reviews that end in engineering work, not just a document.
Incident preparedness
- Design and run tabletop exercises and recovery tests.
- Keep playbooks and runbooks honest, and turn the steps that do not need a human into automation.
- Make sure the evidence, access, knowledge, and authority we need in an incident are in place before we need them.
Your operational week
- Take escalations from the agents and the managed provider, scope them against real asset and business context, decide, and act.
- Execute approved containment, including isolating devices, revoking sessions, restricting access, blocking activity, and preserving evidence.
- Ensure every escalation exits with a solve where one is warranted. Where the solve is a detection or automation, you build it with the agents. Where it is a control change, specify it clearly enough that the owning team can act without a second conversation.
- Close every case with a security disposition, an owner, the evidence, actions taken, and any required follow-up.
- Hold the managed provider to our standard for evidence, analysis, routing, and closure quality.
First 90 days
- Learn the incident command, escalation, evidence, and handover model, and take a rotation slot.
- Independently own escalations across common classes, with defensible dispositions and evidence that stands up.
- Ship your first solve: a detection, automation, or control change that permanently retires a recurring alert class.
- Map coverage for one area of the estate, name its top gaps, and work with the agents to close at least one.
- Assess, tune, and promote your first agent-proposed detections to live.
- Add your first feature to the internal security platform.
- Shadow the incident commander on a real incident or exercise.
- Name one telemetry gap with a business case behind it. Extra credit if it is in operational technology or building management systems, both of which we currently cannot see.
About You
- 5+ years in detection engineering, incident response, security operations, threat hunting, security engineering, or related roles.
- Detection judgement. You can read a rule and say what it catches, what it misses, and how noisy it will be in production. You think in coverage, not in individual alerts.
- Fluency in modern attacker tradecraft, including credential theft, session abuse, phishing, malware execution, persistence, privilege escalation, lateral movement, command and control, and exfiltration, and in how each shows up in telemetry.
- Working knowledge of Windows, macOS, and Linux internals, including where attackers persist on each and what each one logs.
- Incident command experience. You have led, or been central to, a significant incident: running the bridge, setting priorities, and keeping leadership informed without losing the technical thread.
- Experience building in, and securing, public cloud environments like AWS and GCP
- You automate what you repeat, using query languages, scripting, APIs, or workflow automation. When you hit the same problem a third time, your instinct is to build—not add another step to a runbook.
- Hands-on investigation experience across endpoint, identity, cloud, SaaS, network, or production workloads.
- Sound judgement on containment under time pressure.
- Willingness to challenge automated conclusions. You will work alongside AI agents every day, and their analysis and detection logic are inputs, not authorities. We expect you to validate them and notice when they are confidently wrong.
- You write investigations another engineer can follow, escalations a leader can act on in thirty seconds, and case notes that still make sense to someone reading them two years later in an audit.
- Calm and methodical when facts are incomplete or contradict each other.
- Ability to work effectively with engineering, infrastructure, and service owners who do not report to you.
Strong pluses
- Operational technology, industrial control systems, building management systems, or critical facilities experience.
- Cloud infrastructure, AI infrastructure, data centres, HPC, or other availability-sensitive environments.
- Detection-as-code, detection testing, adversary emulation, or purple teaming.
- Building, tuning, or evaluating AI agents or LLM-based security tooling.
- Response experience involving ransomware, identity compromise, destructive attacks, cloud intrusion, insider threats, or supply-chain incidents.
- Experience holding a managed monitoring or response provider to a standard while keeping decisions in-house.
- Experience working as part of a Follow-the-sun rotation.
This role may not be a fit if you:
- Want a queue to work through all day. You hold it one week in three, and the job is making it smaller.
- Measure success in tickets closed.
- Close cases without a security disposition or evidence.
- Accept a vendor alert or an agent’s conclusion without validating it.
- Want to spend most of your time hand-writing detection queries. Here the agents draft the logic and you own whether it is right.
- Would rather not be the person in charge when an incident is live.
What we can offer you
At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
- Highly competitive US compensation package (base + bonus + equity), with performance reviews every 12 months.
- Join one of the fastest-growing AI infrastructure companies—your chance to directly shape how global AI capacity is planned and deployed. ✨
- Expect a dynamic progression plan tailored to your ambitions. Grow by leading critical cross-functional initiatives and shaping capital strategy—always with our full support.
- Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.
Equal Opportunities Statement
We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$160,000 - $190,000 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Nscale does not accept unsolicited candidate submissions from recruitment agencies.
Apply for this job
*
indicates a required field
