New

(RD) Senior Data Engineer

Ho Chi Minh

The Job in short

The Data Enablement team is here to enable every team in the organisation with their data needs. Our job starts the moment that data enters our platform and ends when it reaches whoever needs it. We are a small, senior team of data and AI engineers, working across customer data, product knowledge, and the data the organisation runs on.

We bring data in, reconcile it into one version people can rely on, and make it available to the right audience. Increasingly that audience is agents as well as people, so everything we build has to work for both. Two things must hold at every step: that only the right people can see it, and that we can prove it is correct. You build the platform that makes both possible.

As a Senior Data Engineer you own the foundation: the lakehouse, the pipelines that fill it, and the layers that serve it. We treat data as a product, so each one has a named owner, a documented contract with the teams who consume it, and stated expectations on freshness and quality. That includes Data as a Service, where internal and external teams run analytics on data held in our AI-powered Banking OS.

This is not only a tabular data job. A large part of it is a versioned documentation corpus, kept current, deduplicated and traceable across many product versions, and much of it arrives semi-structured rather than clean. The consumers are not only dashboards either. You build and maintain the MCP tools that let AI agents query documentation, API specs and release history directly. That is a meaningful part of the role, not a side project.

We hold pipelines to the same standard as application code: tested, reviewed, deployed through CI/CD, observable in production, and promoted properly across environments. If that is already how you think about data engineering, you will recognise this team quickly.

Meet the job

  • Develop and maintain pipelines extracting data from many sources: RDBMS, Change Data Capture and streaming endpoints.
  • Process data ranging from observability metrics to bank accounts and transactions, in a highly secure and compliant manner.
  • Transform raw data into business-ready analytical models on a Data Lakehouse.
  • Guard data consistency, quality and end to end governance across the full lifecycle.
  • Unstructured data at scale. Ingest and maintain a large versioned documentation corpus: incremental sync, change tracking, deduplication and freshness. Much of it is public or semi-structured source material rather than clean tabular data.
  • Agent-facing data tools. Build and maintain MCP tools that let AI agents query documentation, API specs and release history. A meaningful part of this role.
  • Pipelines as software. GitHub Actions CI/CD, automated tests, observability and proper environment promotion across Fabric and Databricks. We automate through GitHub, and we expect the same engineering discipline from data pipelines as from application code.

How about you

  • A Bachelor's or Master's degree in Computer Science, Data Science or a related field.
  • 5+ years of experience in data engineering, preferably Microsoft Fabric, Azure Databricks or Spark. Strong experience on another major cloud is fine if you are ready to convert.
  • Expertise in Data Lakehouse concepts and Delta Lake.
  • Working knowledge of data regulation and compliance, and the judgement to know when to involve Legal.
  • Strong technical leadership. You can drive a piece of work end to end, take the calls with other value streams and stakeholders yourself, and keep it moving without a project manager.
  • Excellent written and verbal communication skills in English.
  • Proactive, autonomous and self-sufficient. You work independently on your own projects and take the calls with other value streams and stakeholders yourself.
  • Python and SQL depth, including pipeline and query performance work.

Strong pluses

  • Experience with data contracts and making a platform other teams can self-serve from.
  • A BI layer in production, such as Power BI.
  • Machine learning fundamentals, enough to work with AI engineers as a peer.
  • Retrieval and vector search. Building or maintaining the vector and hybrid search layer that AI products query.
  • Preparing data for model training or fine-tuning, including anonymisation of sensitive data.
  • Instrumenting retrieval so its quality can be measured by someone else.

Our tech stack

  • Languages: Python, SQL.
  • Platform: Microsoft Fabric, Azure Databricks, Delta Lake, Azure Blob, Cosmos DB, Azure Key Vault.
  • Processing and orchestration: PySpark, Fabric pipelines and notebooks.
  • Ops: GitHub Actions, Fabric deployment pipelines, Power BI.
  • AI surface: MCP, vector and hybrid search.

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf