Senior/Staff Software Engineer, Distributed Systems
Why This Role Matters Right Now
The Atomic Machines robotics fleet has reached a level of maturity where it's ready to bring the Matter Compiler online. Now is the time manufacturing software must be brought up to par to leverage this fleet into a true fab. Concretely:
- Architectural decisions — cloud vs. on-prem, how far the control layer should extend — are currently being made in parallel by people with different mental models of the end state, without a shared, written-down source of truth.
- The cost is real: engineering time is being spent today on work that may be thrown away once the architecture actually converges.
We're hiring a seasoned engineer who can operate productively inside that ambiguity — someone who will write the design docs, drive the team toward a decision, and also model (and insist on) the testing, review, and traceability discipline that keeps this from happening again. This is as much a mandate to bring engineering rigor to the org as it is to build a platform that lets process developers — and eventually in-the-loop physical AI — safely program the fab. If that sounds like more org-building than you want in a "software engineer" role, this probably isn't the right fit — and that's a useful thing to know before either of us invests time in the process.
What You’ll Do:
- Design and build the distributed software systems that coordinate state, timing, and behavior across manufacturing hardware; write, test, and debug sensors, actuators, and process controllers under real-time and reliability constraints.
- Design workflows that bridge manual and automated process steps, and coordinate handoffs between production and process development.
- Evolve the existing system in place — including but not limited to API development to govern machine behavior across our fleet— as the architecture converges, without waiting for a clean-slate rewrite to start delivering value.
- Instrument systems so machine and process data is legible — not just to humans via logs, but structured for downstream AI/ML consumption.
- Investigate and resolve issues that span software, firmware, and physical systems.
- Establish and model software engineering practices — testing, code review, CI/CD, documentation — appropriate for a team that's outgrown its current ones.
- Contribute to system reliability through structured observability, fault handling, and graceful degradation.
- Collaborate closely with mechanical, electrical, and process engineers to translate physical constraints into resilient software behavior.
- Partner with engineering leadership to help converge competing architectural visions into one well-reasoned direction, rather than waiting for it to be handed down.
What You’ll Need:
- 5+ years building or debugging systems with real external dependencies: hardware, embedded devices, networked services, or similar.
- A track record of making and defending nontrivial architecture decisions — build vs. buy, deployment topology, service boundaries — not just implementing someone else's design.
- Strong Python skills for production systems, plus proficiency in at least one systems or strongly-typed language (C++, Rust, or Go) [confirm against actual stack].
- Solid grounding in distributed systems fundamentals: state coordination, consistency, failure modes, concurrency.
- Experience introducing or maintaining CI/CD, automated testing, or observability tooling in a codebase that didn't already have it.
- Comfort operating against an incomplete or contested spec — and a bias toward driving clarity rather than waiting for it.
- Bachelor's degree in Computer Science, Electrical Engineering, Robotics, or related field, or equivalent experience.
Bonus Points For:
- Experience with real-time or resource-constrained environments, or analogous domains (IoT, edge compute, industrial systems, robotics, warehouses, manufacturing lines, fabrication and automation facilities).
- You've thought seriously about observability: what to instrument, when logs aren’t enough, how to make failure legible, and you’ve instrumented systems specifically to make their data usable by ML/AI pipelines.
- You've been the person who introduced testing or CI discipline to a team that didn't have it — and can speak concretely about how you got buy-in, not just what you built.
- Experience evaluating cloud vs. on-prem/edge deployment tradeoffs for latency- or safety-sensitive systems.
- You've debugged issues that required reasoning across multiple system layers: application logic, transport, firmware, hardware.
- You're genuinely energized by translating physical constraints — latency, noise, mechanical tolerance, safety margins — into software behavior.
The compensation for this position also includes equity and benefits.
Salary Range
$180,000 - $230,000 USD
Apply for this job
*
indicates a required field