Back to jobs
New

Senior Software Engineer, Observability

Amsterdam

About the Role

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure.

Together AI is seeking a highly skilled Senior Storage Engineer to implement, optimize, and operate critical components of our distributed storage infrastructure that powers AI training and inference at scale. You will work on core systems including distributed file systems (WekaFS and Vast), object storage, and time-series databases, ensuring reliability, performance, and cost-efficiency for our rapidly growing GPU clusters and AI workloads. 

Responsibilities

  • Independently implement, optimize, and maintain high-performance storage solutions for AI/ML workloads; evaluate and integrate new storage technologies (e.g., WekaFS, Ceph, Lustre).
  • Implement and optimize storage network configurations (RDMA, InfiniBand, 100GbE/400GbE); tune network parameters for maximum storage throughput and minimum latency; troubleshoot network bottlenecks affecting storage performance.
  • Build Kubernetes storage operators/controllers; enable automated provisioning, self-service abstractions, multi-tenant isolation, quotas; create reusable Helm/Terraform patterns.
  • Optimize storage systems for GPU clusters (working towards 10-50 GB/s per-node throughput); implement and tune caching strategies (model weights, datasets, checkpoints); troubleshoot performance bottlenecks using profiling tools; contribute to scaling the storage infrastructure.
  • Build multi-tier caches (local NVMe, distributed, object); optimize data locality and model-weight distribution; implement smart prefetching/eviction.
  • Implement monitoring, alerting, SLOs; design DR/backups with runbooks; run chaos engineering; ensure 99.9%+ uptime via proactive/automated remediation.

Requirements

  • 5+ years of experience in storage engineering with significant hands-on experience operating distributed storage systems at scale.
  • Proven track record deploying and operating high-performance storage for GPU/HPC clusters.
  • Deep Kubernetes and cloud-native storage experience in production environments.
  • Strong coding skills, preferably in Go and Python with demonstrated ability to build production-grade tools.
  • BS/MS in Computer Science, Engineering, or equivalent practical experience.
  • Demonstrated ability to independently deliver complex technical projects that significantly improved performance, reliability, or cost efficiency.

Nice to Have Skills

  • ML/AI storage patterns (model weights, checkpointing, dataset caching)
  • Kubernetes operator development (controller-runtime, kubebuilder)
  • Storage snapshots, cloning, and thin provisioning
  • Backup and disaster recovery (Velero, Restic, cross-region replication)
  • Storage encryption (at-rest and in-transit), security and compliance
  • Storage benchmarking and profiling tools (fio, iperf3, iostat, blktrace)

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.

 

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy  




Create a Job Alert

Interested in building your career at Together AI? Get future opportunities sent straight to your email.

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf