tags.new

Incident Operations Lead (EMEA/AMER)

Remote - (EMEA/AMER)

Who We Are:

Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.

Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.

Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.

Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.

Our Team Members:

We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!

We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.

Role

Lead the team that commands Alpaca's most critical incidents. You will build the function and then keep raising its bar: the severity model, the escalation and communication paths, 24x7 follow-the-sun coverage, and the KPIs that prove it is improving. You will do that across boundaries - with the engineering teams who own the services, with SRE on reliability standards and on-call readiness, with Risk on financial and regulatory materiality, with our partner communications teams on what reaches a customer, and with reliability programme management on what happens after.

You will own how well we respond. Not the fix, not the partner communication, and not the reliability standard. Holding that line is a deliberate part of the design and a core part of the job.

Things You Get To Do

  • Build the team and stand up 24x7 command. Recruit and certify Incident Commanders, build a follow-the-sun rotation across APAC, EMEA and AMER with warm handoffs at every regional boundary, and carry a rostered slot yourself. Keep the team sharp between real incidents with game days, tabletop exercises and simulations, and coach them through the live ones. Build a blameless review culture that treats an outlier as a process gap rather than a person's failure.
  • Own the process, and keep raising it. Drive severity maturity with Risk on financial and regulatory materiality - in a regulated brokerage a severity call can also start a reporting clock, so the model has to map cleanly onto those thresholds. Own the escalation path and what happens when a page goes unanswered, agree the thresholds for taking an incident to engineering leadership, and keep the service catalogue and its ownership current - time spent working out who owns a failing service is customer impact.
  • Own both bridges. Your team connects the engineers and technical support fixing the problem, who need uninterrupted focus, to the partner communications teams, who need a continuous and accurate feed. You open the channel, supply the facts and hold the update cadence to account. Afterwards, your team runs the retrospective with SRE - who own the technical depth - and builds the post-incident package while the room is still warm, every item ticketed, owned and tagged, delivered inside a service level you define and then hold, before reliability programme management drives it to closure. You then synthesise the discussion into short, digestible learnings and publish them to the whole engineering organisation, so one team's failure becomes everyone's lesson instead of a document three people read.
  • Own the KPIs. Time to respond and time to mitigate end to end, including the definitions and data hygiene beneath them: what separates mitigated from resolved, and whether a timestamp means what it claims. Establish a defensible baseline before committing to targets, then move them by severity. Review and approval, not authorship, is where postmortems stall, so report overdue reviews by team and incident with a next action against each.
  • Build it as a product, then automate it with AI. Everything is documented, versioned and deployable, so you can stand up command from the artefacts alone; the process needs to scale considerably faster than the team. The automation we are after is AI workflows and agents rather than scripts and dashboards - agents that set an incident up, assemble the timeline as it runs, draft the RCA and the action package, and chase the update that is due or the review that is overdue. You own that roadmap: what an agent may do unsupervised, what still needs a commander's judgement, and the decision-tree quality that makes either of them safe.

Who You Are (Must-Haves)

  • You have stood up an incident command or major-incident function, not only worked inside one - you have owned the severity model, built the roster and driven adoption across teams.
  • 5+ years in production engineering, SRE or technical operations, including hands-on command of high-severity incidents.
  • You have led a distributed team across time zones and run a 24x7 rotation.
  • You get engineers you do not manage to do things, and you can defend a severity call to someone who disagrees with it.
  • You have built reliability metrics people trust, and you know the difference between improving a number and improving reality.
  • You are disciplined about scope. You can say "that is not ours" and route it, in the middle of an outage, without leaving a gap.
  • You write well enough that your process documents actually get used, you can hold a bridge calm under pressure, and you can brief an executive mid-incident without either downplaying it or dramatising it.
  • You understand FinTech and the trust stakes of API-driven financial platforms.
  • You use AI and agentic automation to remove toil rather than to add tooling.

Who You Might Be (Nice-to-Haves)

  • Formal incident command training - ITIL, Major Incident Management or crisis management.
  • You have run a certification, game day or drill programme, or built a pool of certified responders beyond your own headcount.
  • Experience with modern incident management and on-call platforms.
  • You have built a service catalogue or ownership registry that people actually maintained.
  • You have worked with programme management or reliability functions to convert incident follow-ups into funded roadmap work.
  • Familiarity with incident reporting obligations in regulated financial services - DORA, Reg SCI, FINRA or equivalent.
  • Online securities trading or capital markets experience, or another regulated, market-hours-sensitive domain.
  • You have deployed the same operating model into a second region or entity.

How We Take Care of You:

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.

Recruitment Privacy Policy

Create a Job Alert

Interested in building your career at Alpaca ? Get future opportunities sent straight to your email.

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in Alpaca ’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Select...
Select...
Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Select...

Voluntary Self-Identification of Disability

Form CC-305
Page 1 of 1
OMB Control Number 1250-0005
Expires 07/31/2029

Why are you being asked to complete this form?

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

How do you know if you have a disability?

A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to:

  • Alcohol or other substance use disorder (not currently using drugs illegally)
  • Autoimmune disorder, for example, lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS
  • Blind or low vision
  • Cancer (past or present)
  • Cardiovascular or heart disease
  • Celiac disease
  • Cerebral palsy
  • Deaf or serious difficulty hearing
  • Diabetes
  • Disfigurement, for example, disfigurement caused by burns, wounds, accidents, or congenital disorders
  • Epilepsy or other seizure disorder
  • Gastrointestinal disorders, for example, Crohn's Disease, irritable bowel syndrome
  • Intellectual or developmental disability
  • Mental health conditions, for example, depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD
  • Missing limbs or partially missing limbs
  • Mobility impairment, benefiting from the use of a wheelchair, scooter, walker, leg brace(s) and/or other supports
  • Nervous system condition, for example, migraine headaches, Parkinson’s disease, multiple sclerosis (MS)
  • Neurodivergence, for example, attention-deficit/hyperactivity disorder (ADHD), autism spectrum disorder, dyslexia, dyspraxia, other learning disabilities
  • Partial or complete paralysis (any cause)
  • Pulmonary or respiratory conditions, for example, tuberculosis, asthma, emphysema
  • Short stature (dwarfism)
  • Traumatic brain injury
Select...

PUBLIC BURDEN STATEMENT: According to the Paperwork Reduction Act of 1995 no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.


We use Greenhouse’s AI-powered Talent Matching tool to compare your application against our job requirements.

Learn more