Performance Engineer
Astera Labs (NASDAQ: ALAB) provides rack-scale AI infrastructure through purpose-built connectivity solutions. By collaborating with hyperscalers and ecosystem partners, Astera Labs enables organizations to unlock the full potential of modern AI. Astera Labs’ Intelligent Connectivity Platform integrates CXL®, Ethernet, NVLink, PCIe®, and UALink™ semiconductor-based technologies with the company’s COSMOS software suite to unify diverse components into cohesive, flexible systems that deliver end-to-end scale-up, and scale-out connectivity. The company’s custom connectivity solutions business complements its standards-based portfolio, enabling customers to deploy tailored architectures to meet their unique infrastructure requirements. Discover more at www.asteralabs.com.
Senior Performance Engineer
Location: San Jose, CA (On-site)
Role Overview
Astera Labs is a hyper-growth connectivity company enabling the rack-scale AI infrastructure powering the world's most advanced GPU clusters. Our Scorpio scale-up fabric switches are purpose-built to unlock the performance of next-generation AI workloads, and we're looking for a Senior Performance Engineer to demonstrate the real-world value of our silicon where it matters most: on real inference and training workloads running on GPUs at scale.
In this role, you will define how the world measures scale-up fabric performance. You'll build the roofline models, benchmarks, and end-to-end workload studies that quantify our performance leadership, expose bottlenecks, and drive performance fine-tuning of real AI workloads on our fabric to inform product direction. Your data will directly shape architecture, firmware, and product decisions — and fuel the marketing narrative that positions Astera Labs at the center of AI connectivity.
Key Responsibilities
- Performance Characterization & Benchmarking
- Establish theoretical and measured roofline models for Astera Labs' scale-up fabric across key performance metrics, defining the reference for all comparative testing.
- Build and maintain baseline performance benchmarks using industry-standard tools such as NVBandwidth and NCCL across a range of GPU configurations and switch topologies.
- Quantify the impact of differentiated Astera Labs AI fabric features (e.g., Hypercast, In-Network Computing) against baselines using both synthetic benchmarks and real inference workloads.
- Real Workload Analysis & Fabric Scalability
- Run end-to-end inference model workloads on target hardware to capture real-world performance beyond synthetic benchmarks, supporting architecture decisions and customer-facing demonstrations.
- Evaluate fabric performance as inference cluster size scales from 16 to 32 GPUs and beyond, identifying bottlenecks and building performance scaling models for state-of-the-art AI workloads.
- Design and execute head-to-head performance comparisons against competing fabric switch solutions to produce data-driven differentiation evidence.
- Test Infrastructure & Automation
- Design, build, and maintain automated lab infrastructure including test execution pipelines, traffic generation tooling, and data collection and reporting systems.
- Enable repeatable, high-quality, and scalable performance measurements across all hardware configurations, reducing manual effort and accelerating the test cycle.
- Share infrastructure and playbooks with the Product Applications team to accelerate customer application development and issue resolution.
- Cross-Functional Impact & Innovation
- Partner closely with ASIC architecture, firmware, software, Product Definition, Product Applications, and Product Marketing teams to communicate findings, influence design decisions, and resolve performance-impacting issues.
- Serve as a key technical resource in the early evaluation of new fabric architectures, interconnect technologies (UALink, PCIe Gen 6/Gen 7, Ethernet, UEC), and AI/ML communication paradigms.
- Provide performance data, analysis, and live benchmark support for key customer engagements and industry events; produce clear, audience-appropriate performance reports, technical briefs, and marketing collateral, and maintain living documentation in Confluence.
Basic Qualifications
- Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field. We welcome both recent graduates with strong, directly relevant project, research, or internship experience and candidates with 2–5 years of industry experience in performance or systems engineering.
- Hands-on experience running AI/ML workloads on GPU clusters — including benchmarking, performance analysis, and fine-tuning of workloads across clusters of GPUs or accelerators. This can come from industry, research, or substantial academic projects.
- Demonstrated ability to debug and root-cause system-level performance issues across hardware, firmware, software, and network boundaries.
- Excellent fundamental knowledge of compute algorithms, parallel algorithms, and AI/ML algorithms and workloads.
- Strong working knowledge of computer systems, GPU systems, and datacenter networking — including PCIe and Ethernet fundamentals.
- Working knowledge of GPU and CPU software stacks (e.g., CUDA, MPI, collective communication libraries, drivers, and OS-level performance tooling).
- Proficiency in scripting and automation (e.g., Python) to build test pipelines and analyze large performance datasets.
Preferred Qualifications
- MS or PhD in Computer Engineering, Computer Science, Electrical Engineering, or a related field.
- Experience with scale-up fabrics and next-generation interconnects such as UALink, PCIe Gen 6/Gen 7, Ethernet, or UEC.
- Deep understanding of modern inference and training workloads (LLMs, MoE, recommender systems) and their communication patterns.
- Experience developing roofline models and competitive performance analyses for switching, networking, or accelerator silicon.
- Excellent written and verbal communication skills, with the ability to translate deep technical findings into concise executive summaries and customer-facing narratives.
Salary range is $135,000 to $170,000 depending on experience, level, and business need. This role may be eligible for discretionary bonus, incentives and benefits.
We know that creativity and innovation happen more often when teams include diverse ideas, backgrounds, and experiences, and we actively encourage everyone with relevant experience to apply, including people of color, LGBTQ+ and non-binary people, veterans, parents, and individuals with disabilities.
Apply for this job
*
indicates a required field