
Senior Inference Optimization Engineer
About Us
Venice is the world's leading consumer AI company built on principles of privacy, free speech, and user sovereignty. We're building the Port City of AI, in which millions of individuals, third party apps, and AI agents gather, interact, and access sophisticated AI resources. Our mission is to make artificial intelligence approachable and useful in everyday work—bridging the gap between cutting-edge research and practical, real-world impact.
We're a fast-moving startup where every team member is expected to make a clear impact. Our culture is rooted in curiosity, ownership, ethical principle, philosophy, and collaboration—whether we're designing better AI workflows, supporting our growing community, or shaping the future of human-AI interaction. Joining Venice AI means joining a team of unorthodox builders who believe in moving quickly, delivering a beautiful and highly useful consumer product, and maintaining an edge in the rapidly evolving world of agentic machine intelligence. If you're energized by big ideas, entrepreneurial spirit, individual empowerment, and the opportunity to help shape a company from the ground up, you'll feel right at home here.
Why we are hiring
Venice is taking on the global market as the only AI infrastructure provider that prioritizes the privacy and personal sovereignty of users' personal data. This is an opportunity for you to be on the bleeding edge of privacy-focused AI with a unique and dedicated team of high-agency individuals doing truly novel work. You'll be instrumental in helping Venice achieve maximum inference performance at significant scale.
What you'll do
- Help stand up and optimize Venice's owned GPU infrastructure, including B300 nodes in our data centers
- Drive down latency (TTFT and TPOT), push throughput, and improve cost per token for LLM inference workloads
- Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang) to identify the optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
- Work with our inference routing system to optimize multivariate inference load-balancing algorithms
- Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels), novel attention variants, new quantization schemes, and compilation stack improvements. Hands-on kernel development experience is a strong plus.
- Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in Venice's stack.
Who you are
- 5+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge
- Proficiency in one of Python, Rust, or Go. Bonus: C++/CUDA
- Hands-on experience with at least one production LLM inference engine (e.g. vLLM, SGLang) running at high volume in production
- Demonstrated experience with LLM inference optimization techniques: continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile
- Fluency with quantization tradeoffs, both qualitative and quantitative
- Experience with distributed inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments
- Fluency with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler) and a bias toward measuring before optimizing
- Bonus: diffusion/image model inference optimization, custom Triton kernels, contributions to open-source inference frameworks
Full time employment offers from Venice.ai. include a variety of benefits, including medical, dental, vision, and 401(k), and may include an offer of restricted token units.
*Our perks and benefit packages are subject to change and may vary based on your location or employee status.
Apply for this job
*
indicates a required field