Senior Software Engineer, Gemini Audio Inference - Mountain View
Snapshot:
Our team is responsible for building the Productionization and Serving infrastructure for Gemini Audio Inference at Google DeepMind. We help land Gemini Audio capabilities into numerous clients including Astra, GeminiApp, YouTube, Search, Meet, Cloud, Geo, Assistant etc. Based on research Gemini model flavors, our tasks involve model latency/throughput optimizations, serializations, orchestration, evaluation & finally landing in production, at Google scale.
- Our team focuses deeply on Inference efficiencies for Gemini and its related components. We actively develop new infrastructure to make Gemini more accessible to new streaming use cases.
- Our team has deep collaboration with both the research team and production platform team, exposed with SOTA research work and their inference optimizations.
- Our team owns the infra to serve the Audio tokenization & Audio generation around the Gemini models.
Join us if you are interested in having a direct impact on making Google's products better for our users in over 100 languages! Here are showcase videos that directly used the infra built from our team: Gemini Audio, Project Astra, Real Time Translation.
About us:
Artificial Intelligence could be one of humanity’s most useful inventions. At Google DeepMind, we’re a team of scientists, engineers, machine learning experts and more, working together to advance the state of the art in artificial intelligence. We use our technologies for widespread public benefit and scientific discovery, and collaborate with others on critical challenges, ensuring safety and ethics are the highest priority.
The role:
You will work with the research teams to build the serving solution of Gemini models for different clients. Propose and test the best serving configurations based on client needs. Build new infrastructure to serve Gemini in different ways (e.g. streaming). Design and build audio-specific logic in the orchestration framework. Ensure the Gemini model quality in the production environment.
We are looking for highly skilled Software Engineers interested in improving Gemini serving performance. This is a fast-paced role, and deadlines may be aggressive, however we do value work-life balance. You should be prepared to wear many hats and be ready to solve a wide range of exciting technical problems.
Some of the things we are working on:
- Build and improve the new infrastructure to support the streaming Gemini prompt on a large scale.
- Bring the Gemini Audio related capabilities to clients: e.g. Long Context capabilities to Meet YouTube clients; Audio Out capabilities to Maps team.
- Embracing a new orchestration infra around Gemini model servings.
About you:
You are a talented and resourceful engineer/researcher capable of applying large language models, aka Gemini, at scale; conceiving of and executing experiments to improve them; and overcoming last mile issues to launch your model into production for millions of people to use. We seek out individuals who thrive in ambiguity and have a shipping mindset, willing to help out with whatever moves prototypes forward. We regularly need to invent novel solutions to problems, and often change course if our ideas don’t work out, so flexibility and adaptability to work on any project is a must.
In order to set you up for success as a Software Engineer at Google DeepMind, we look for the following skills and experience:
- BS, MS or PhD degree in computer science, mathematics, applied stats, machine learning or similar experience working in industry
- Experience working on software engineering projects from proof-of-concept through to implementation.
- Required programming languages: C++; Python a plus
- Experience in performance engineering; GPU/TPU a plus.
- Great communication skills and interpersonal skills
- Knowledge of machine learning and inference
- Experience productionizing state-of-the-art large language and multimodal models a plus
The US base salary range for this full-time position is between $166,000 - $244,000 + bonus + equity + benefits. Your recruiter can share more about the specific salary range for your targeted location during the hiring process.
Note: In the event your application is successful and an offer of employment is made to you, any offer of employment will be conditional on the results of a background check, performed by a third party acting on our behalf. For more information on how we handle your data, please see our Applicant and Candidate Privacy Policy.
At Google DeepMind, we value diversity of experience, knowledge, backgrounds and perspectives and harness these qualities to create extraordinary impact. We are committed to equal employment opportunities regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, pregnancy, or related condition (including breastfeeding) or any other basis as protected by applicable law. If you have a disability or additional need that requires accommodation, please do not hesitate to let us know.
Apply for this job
*
indicates a required field