Engineering Manager, Platform
Firmus Technologies
Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure from model to grid. Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.
At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. This co-designed approach from model to grid allows us to make every watt count and deliver low-cost AI tokens globally.
Firmus AI Cloud
Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale. It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.
Why you’ll love working here
At Firmus, you’ll work at the intersection of sustainability and artificial intelligence in a fast-paced environment powered by next-generation technology. You’ll be helping to transform an entire industry — and you’ll feel it every day.
Our team is made up of true innovators and leaders in their fields, and as an emerging company, you won’t be lost in a crowd. You’ll work closely with the founders, build a strong network, and see the impact of your work first-hand as we democratise AI tools for everyone — more sustainably and more affordably.
We believe great things happen when people from diverse backgrounds come together to do their best work and be their authentic selves. We are proud to be an equal opportunity employer.
ROLE
Firmus Technologies is seeking an Engineering Manager to lead the Platform Engineering and Observability team. You are accountable for both the people and the delivery: the engineers you grow and the platform software they ship. Your team builds and operates the platform beneath the Firmus AI Cloud, from bare-metal GPU compute and high-performance networking to the internal platform services, self-service tooling, and the observability platform that our engineering teams and customers depend on. You run these as products for the engineering teams that build on them, investing in self-service, reliability, and developer experience. You are the escalation point above first-line operations and own the cross-cutting decisions your team shares. This is a build-and-grow leadership role: you build the team that scales our AI platform toward gigawatt-scale AI factories across the regions and grow your scope as you earn it.
KEY RESPONSIBILITIES
- Team Leadership & Growth
Be the Single-Threaded Owner (STO) for multiple squads: their engineers and tech leads report to you, and you work alongside product, architecture, and delivery peers. Hire strong engineers, develop them through coaching and clear feedback, and manage performance directly. Own the org design, headcount planning, and engineering culture for your group, building a team that raises its own bar rather than depending on you.
- Delivery & Execution
Own end-to-end delivery for your team. Turn the product roadmap into sequenced, predictable delivery, shaping prioritisation with product and working with team leads on sprint planning and release gates. Own the cross-squads' dependencies that determine whether the platform ships on time, keep delivery health visible through metrics, and balance velocity against reliability and technical debt.
- Technical Direction
Stay technically credible and set the engineering quality bar across your squads. You are accountable for making sure the right cross-cutting decisions get made and driven to closure, such as build, buy, or open source, and the product SLAs your teams commit to. Keep design and code review rigorous, champion AI-assisted development so your teams ship faster without lowering the design, review, or security bar, and unblock the hard problems that stall delivery.
- Operational Excellence & Governance
Own the reliability and security of the services your squads run in production. Act as the escalation point above the operations centre, accountable for L3 incident resolution, SLA-breach response, and post-mortems that convert into runbook and prevention work. Set the standards for SLOs, on-call, and change management, backed by the observability platform your team owns. Govern how AI-generated code and agentic workloads reach production, extend Firmus' existing SOC 2 Type 2 and ISO 27001 controls as the platform scales into new data centres, and own the cost efficiency of what your team operate, balancing reliability, performance, and speed against infrastructure spend.
- Stakeholder & Customer Communication
Represent your team to engineering leadership and the CTO, clear about progress, risk, and the decisions you need. Own the engineering side of the customer relationship: lead customer technical briefings, and represent engineering directly when escalations turn on delivery, reliability, or architecture. Align with product, who own the roadmap and customer outcomes, with the architects on technical direction, and with delivery and operations so the platform ships and runs as one system, not a set of parts.
SKILLS AND EXPERIENCE
Education
- Bachelor's degree in computer science or a related technical field.
- 4+ years managing engineers, with direct ownership of hiring, growth, and performance for a team of at least 8 engineers.
- Track record building high-performing, self-improving engineering team and a strong delivery culture.
- Proven track record delivering complex infrastructure or platform at scale, making delivery predictable across multiple teams and dependencies.
- You thrive in a fast-paced, high-growth environment and make sound decisions with incomplete information, reprioritising as the business requirements change.
- Experienced with AI-assisted development and agentic AI, and able to lift your teams' throughput and quality with these tools.
Platform and Technical Depth
- 10+ years in platform or infrastructure engineering, including a strong hands-on background as a senior or staff engineer before moving into management. You are still able to reason about architecture in depth and intend to stay close to the technology.
- Experience running a platform as a product: understanding your customers, curating a self-service engineer experience that lowers cognitive load, and measuring adoption and value rather than uptime alone.
- Strong technical foundation in cloud-native or bare-metal infrastructure: Kubernetes, infrastructure-as-code, CI/CD, and distributed systems, fluent in a modern backend stack such as Go and Python, with enough depth to hold your own in design reviews and make sound technical trade-offs.
- Hands-on experience operating an observability platform at scale: metrics, logs, and traces, and the streaming pipelines behind them as a product that customers or other teams rely on.
- Comfortable reviewing code and holding a high bar in design and code review.
Operations, Security and Cost
- Experience owning operational excellence for production systems: reliability and SLOs, on-call and incident management, and change control.
- Experience owning security and compliance for production systems, ideally against frameworks such as SOC 2 Type 2 or ISO 27001.
- Comfortable owning the path to production for AI or agentic workloads, including how AI-assisted development is reviewed and security-gated.
- Ownership of infrastructure, platform or software cost as part of running a team, with the judgement to balance cost against reliability and speed.
- Willing to take part in the incident-response on-call rotation for the services your team owns.
- Willing to travel overseas occasionally when the role requires it.
Stakeholder and Business
- Strong stakeholder management: partnering with product, presenting to senior technical leadership, and briefing enterprise or government customers on technical trade-offs.
- Able to connect engineering investment to business outcomes, weighing customer value, cost, risk, and platform scalability when you set priorities.
- Clear and effective written and verbal communication in English. You can explain a delivery or architectural trade-off to a CTO or customer without losing the substance.
Create a Job Alert
Interested in building your career at Firmus Technologies ? Get future opportunities sent straight to your email.
Apply for this job
*
indicates a required field