
LLMops Engineer
At Coursera, we aim to allow employees to develop their skills and grow their careers. One way we do this is with internal transfers. In order to support open and fair transfers, we practice transparency and have set guidelines for who is eligible to request a transfer.
Anyone who has been in their role for more than 6 months and in good standing may apply for a new position at Coursera. For consideration outside this timeframe, individuals should talk to their manager and People Business Partner. Please be sure to read the our Internal Mobility Program Policy here.
We’re a globally distributed team that comes together intentionally for collaboration, complex problem-solving, and key milestones — creating opportunities for teams to do their best work together. Our virtual hiring and onboarding experience makes it easy to join us and start making an impact from anywhere. If you’re ready to make a global impact, help scale unique products across Coursera + Udemy, and grow your career, apply below.
As an AI Platform and LLMOps Engineer, you will join a fast-paced innovation team at Coursera working on AI enablement for our organization and customers. As an internal focus area, this team builds and maintains the infrastructure & tooling required to drive meaningful AI adoption across the organization, providing a platform for non-technical team members to build and deploy tools and transforming entire business functions into AI-native operating models. As a customer-facing focus area, this team builds custom AI-powered solutions tailored for our enterprise customers supporting them in their journey of AI adoption and upskilling beyond just the content on our platforms.
You will own the operational backbone that lets our agentic AI systems run reliably in production — from prompt and pipeline versioning to evaluation, monitoring, incident and cost management. You will work closely with a cross-functional team of Software Engineers, AI Specialists and Product Managers – helping turn promising AI prototypes into reliable, observable, and cost-effective production systems.
Key Responsibilities
- Build and maintain the operational tooling for our LLM-powered systems, including prompt/pipeline versioning, evaluation harnesses, and CI/CD workflows tailored to non-deterministic AI outputs
- Implement monitoring and observability for production LLM systems — tracking latency, token usage, cost per request, output quality, and drift over time
- Design and run automated evaluation suites to catch regressions, hallucinations, and quality degradation before they reach customers
- Manage RAG pipelines and vector store infrastructure, keeping retrieval sources fresh, accurate, and performant
- Implement safety and compliance guardrails — content filtering, PII redaction, and access controls — in line with enterprise data privacy and residency requirements
- Own cost governance for LLM usage: caching strategies, model routing, and usage reporting to keep spend predictable as adoption scales
- Collaborate closely with Product Managers and Senior Engineers to scope operational requirements and translate them into reliable systems
- Participate in sprint planning, technical grooming, and retrospective discussions
Education
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering or a related field
Basic Qualifications
- 4+ years of experience in a software engineering role, with a solid background in Backend Engineering and DevOps fundamentals
- Hands-on experience operating or supporting LLM-powered systems in production (via APIs, RAG pipelines, frontier/open weight/fine-tuned models)
- Proficiency in a backend language such as Python, Java, TypeScript or Go, and familiarity with Docker and container orchestration (Kubernetes)
- Experience implementing APIs, working with SQL and NoSQL databases, and writing automated tests
- Working knowledge of at least one cloud platform (AWS, GCP, or Azure), across both managed and self-hosted services
- Familiarity with infrastructure-as-code (e.g., Terraform) and CI/CD tooling (e.g., GitHub Actions, Jenkins)
- Comfort debugging production issues involving non-deterministic systems, with attention to detail and a bias towards reliability
- Strong belief in engineering quality and building tooling that creates leverage for others
Preferred Qualifications
- Experience with LLMOps-specific tooling such as LangFuse, LangSmith, Weights & Biases-style evaluation frameworks, or vector databases like Pinecone, Weaviate, pgvector
- Working knowledge of RAG architectures, MCP implementation & governance, agent orchestration frameworks like LangGraph and LLM Gateways such as LiteLLM, Open Router & enterprise AI ecosystems such as Vertex AI, Bedrock
- Exposure to prompt management and versioning practices treated as code along with model access management, deterministic guardrails and evals
- Understanding of AI governance considerations — data privacy, residency, and compliance in enterprise AI deployments
- Ability to work in a fast-paced, ambiguous environment with a proactive, ownership-driven mindset
- Strong communication skills and comfort collaborating across engineering, product, and strategy functions
Why Join Us?
- Work at the operational core of Coursera's AI transformation — solving problems that keep real production AI systems reliable, safe, and cost-effective
- Build hands-on expertise in one of the fastest-growing and most in-demand engineering disciplines
- Join a supportive, innovative team with a strong culture of continuous learning and improvement
- Be part of a mission-driven company transforming global access to education and upskilling in the AI era
Coursera is an Equal Opportunity Employer committed to building a welcoming and inclusive workplace. We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request at recruting@coursera.org.
Apply for this job
*
indicates a required field