Senior AI Engineer
The Job in short
The Data Enablement team is here to enable every team in the organisation with their data needs. Our job starts the moment that data enters our platform and ends when it reaches whoever needs it. We are a small, senior team of data and AI engineers, working across customer data, product knowledge, and the data the organisation runs on.
We bring data in, reconcile it into one version people can rely on, and make it available to the right audience. Increasingly that audience is agents as well as people, so everything we build has to work for both. Two things must hold at every step: that only the right people can see it, and that we can prove it is correct. The second is where this role sits.
As a Senior AI Engineer you own the promise that what we serve is right. You are the leading voice on assurance and evaluation across everything the team delivers, from the source through to the answer someone reads. Product calls this assurance. Engineers call it evaluation. It is the same job.
In practice it is two kinds of checks and one framework. Deterministic checks on structure, completeness and freshness. Probabilistic checks on whether a generated answer is grounded and correct, including the judges and the golden datasets behind them. Today those checks exist per project and per tool, built by whoever needed them. You turn them into one configurable framework, reusable across sources, reporting into a single view of platform health. Product knowledge is where it starts, because it is the one journey we run end to end ourselves, from ingesting the source to the answer a customer reads.
That framework is the first phase, not the ceiling. Once it is running and trusted, the same judgement applies to everything else we build: retrieval quality, agent behaviour, and the systems that serve them. We are hiring for the authority you bring on assurance and evaluation, because it is the capability that decides how far the rest can go.
Meet the job
The list below describes the AI Engineer discipline at Backbase. In Data Enablement your first focus is the evaluation and assurance responsibilities that follow. The wider discipline opens up once that foundation is in place.
- Agentic Orchestration: design and implement complex agentic workflows and assistant platforms using the LangChain ecosystem.
- Advanced Retrieval: develop and optimise RAG and GraphRAG pipelines to give agents deep, contextual domain knowledge.
- System Design: architect scalable, distributed AI services that integrate into our Kubernetes environments.
- Agentic Ops (AIOps): implement robust monitoring, tracing and evaluation frameworks (LangSmith, Langfuse, Promptfoo).
- Skill Integration: build and manage Skills and toolsets for agents, including the Model Context Protocol (MCP).
- Human-in-the-Loop: design HIL patterns so high-stakes financial decisions remain governed and accurate.
- Mentorship: provide technical leadership to junior engineers and contribute to internal AI strategy and best practices.
- Evaluation datasets: Build them from acceptance criteria and expert grounding, split dev and test on different samples so tuning and grading never share data, and freeze golden sets.
- LLM judges: Run failure-mode analysis on each acceptance criterion, map it to measurable metrics, and write and iterate the judge prompts.
- Calibration: Rate judges against human ratings and measure agreement as true positive and true negative rate. A judge does not gate a release until its agreement with human raters clears an agreed threshold.
- Retrieval evaluation: Judge chunk relevance, corpus coverage and freshness separately from answer quality, and run the quality gate that blocks a bad release.
How about you
- A Bachelor's or Master's degree in Computer Science, Data Science or a related field.
- 5+ years of professional engineering experience, including at least one LLM-based system you took to production.
- Excellent written and verbal communication skills in English.
- Agent frameworks: experience with an agent framework such as LangChain or LangGraph, or a credible equivalent, for production-grade LLM applications.
- Agentic experience: hands-on experience building agents or assistant platforms using Tools, Skills and MCP.
- AI infrastructure: strong knowledge of RAG architectures. Awareness of graph-based retrieval approaches is a plus.
- Python: strong Python, including clean, maintainable asynchronous code.
- Observability and Evaluation: experience with Agentic Ops platforms such as Langfuse, LangSmith or Promptfoo.
- Measurement discipline. Comfortable with true and false positive rates, thresholds and regression baselines, and able to explain why a one-sided judge is worse than none.
- Dataset construction discipline. Test cases from acceptance criteria, dev and test splits, frozen golden sets, and preventing leakage.
- The temperament to be the gate. Comfortable saying a result is not yet trustworthy, and making that case to the people who own the release. Proactive, autonomous and self-sufficient.
- Evaluations in CI as real release gates rather than advisory reports.
Strong pluses
- System architecture: designing and managing scalable distributed services in a Kubernetes ecosystem.
- Java Knowledge: experience with Java, particularly integrating AI services with Backbase's core Java backend.
- DevOps Culture: familiarity with CI/CD pipelines and cloud-native logging and monitoring.
- Fintech Experience: understanding of the security and compliance requirements unique to banking and financial services.
- A testing or quality engineering background. The instinct transfers better than the domain does.
- Production monitoring of AI systems: online judges over live traces, review sampling and alerting.
Our tech stack
- Languages: Python (primary), Java (secondary).
- Frameworks: LangChain, LangGraph, FastAPI.
- Ops and tools: LangSmith, Langfuse, Promptfoo, MCP.
- Deployment: Docker, Kubernetes, AWS/Azure.
Apply for this job
*
indicates a required field