Instructor-led course · Advanced

Enterprise RAG Engineering

Enterprise RAG Engineering is an instructor-led VDF AI course for data and AI engineers who build retrieval over documents and databases. In four live half-days you connect and profile sources, build and test vector indexes, keep retrieval inside source permissions and prepare fine-tuning datasets, then earn the VDF AI Certified RAG Engineer certificate through a capstone.

  • 4 live half-days
  • 6 modules + capstone
  • Remote or on-site
  • VDF AI Certified RAG Engineer
Level
Advanced
Format
Live and instructor-led, remote or on-site
Length
Four live half-day sessions (3.5 hours each)
Audience
Data engineers, AI engineers and data leads who build retrieval over enterprise documents and databases
Cost
Free for customers and partners; quoted for other teams
Certificate
VDF AI Certified RAG Engineer
Reply to applications
Within 2 business days
Labs
One after every module

What you will be able to do

  • Connect apps and databases with the narrowest useful scope and a dedicated read-only account
  • Judge whether a source deserves trust from discovery, quality scores and drift signals
  • Build vector indexes with row text and chunk settings that suit the content
  • Decide which table columns become searchable text and which data stays with live queries
  • Keep retrieval inside the permissions and visibility of every source
  • Test retrieval against real questions before and after each rebuild

Prerequisites

  • Completion of the Private RAG Engineering path, or equivalent experience indexing enterprise content
  • A test database and document set you may connect in a deployment with VDF AI Data enabled
  • Recommended free path: Private RAG Engineering

6 modules and a capstone

The modules run across the four sessions. Each one ends with a lab in VDF AI.

  1. Connecting sources and databases

    Bringing files, connected apps and databases in as governed sources, scoped for precise retrieval from the first day.

    • Choosing between file upload, connected app, database connection and pasted text
    • Scoping by project, team space or schema rather than a whole drive or server
    • Supported databases, the generic JDBC option and Jira as a structured source
    • Dedicated read-only accounts, secret references and the four connection states
    • Testing a connection, refreshing asset inventory and rotating credentials on a calendar
    Lab
    Connect one document source and one database schema through a dedicated read-only account, run Test connection, refresh the asset inventory and record the owner in each connection’s description.
    Outcome
    Two tightly scoped, tested sources, each with a named owner and a refresh routine.
  2. Discovery and data health

    Finding out what a source holds and whether it deserves trust before anything is indexed.

    • Running discovery: asset cards with size, freshness, owner, quality score and tags
    • Reading the 0–100 quality score and deciding when to investigate first
    • Exploratory data analysis: missing values, duplicate rows, outlier columns and class imbalance
    • Column profiles and low, medium and high drift signals
    • Tags, comments and asset history as context the whole team can find
    Lab
    Run discovery on your connected schema, tag the assets that matter, run exploratory analysis on the most important table and leave what you learn as comments on the asset.
    Outcome
    A short, evidence-based list of assets that are fit to index, with their risks written down.
  3. Vector indexes, semantic search and rows to text

    Building a search surface you own: what goes into it, how rows become text and how text is split.

    • Index sources: text-heavy table columns, feature lists, connected apps and file collections
    • Rows to text: which columns become searchable chunks and which data is better queried live
    • Chunk size and overlap for short snippets, long documents and mixed material
    • Row text and row limits at build time: textColumns, textTemplate and maxRows in the Data API
    • Build stages and states, from reading and chunking through embedding to Ready
    Lab
    Build one index over long-form documents and one over the text-heavy columns of a table, each with chunk settings suited to its content, then search both with the same set of questions.
    Outcome
    Index settings you can justify from search results rather than from defaults alone.
  4. Feature lists, relationships and fine-tuning datasets

    Curating the columns that matter for one use case, then reusing that curation for search and for training data.

    • Feature lists: one default per asset plus specialised lists for each use case
    • Feature discovery: derived features with confidence scores to accept or dismiss
    • Association analysis, and why a strong correlation is a question rather than a conclusion
    • Fine-tuning datasets: mapping templates, preview and the Draft, Ready and Exported states
    • Index or fine-tune: current facts to look up versus a change in behaviour or voice
    Lab
    Build a feature list for one table, triage the suggested derived features, scope an index to the list and preview a small fine-tuning dataset from the same source.
    Outcome
    One curated feature list behind both a focused index and a previewed training dataset.
  5. Permission-aware retrieval and data governance

    Keeping every answer inside what the person asking may see, and proving that it stays there.

    • Why a vector index flattens the permissions of its source
    • Letting the source account’s grants decide what an index can read
    • Keeping sensitive rows out of an index built for a wider audience
    • Who can search an index, and why each index needs a named owner account
    • Rebuilding after every permission change, and what a failed rebuild leaves behind
    Lab
    Index a source through a read-only account with narrow grants, keep sensitive rows out with a retrieval table, revoke a grant and watch what the index returns before and after a rebuild.
    Outcome
    Retrieval that stays inside source permissions, with tests that prove it.
  6. Testing retrieval and keeping it current

    Measuring retrieval against real questions, then keeping indexes and connections fresh once people depend on them.

    • A question set drawn from real questions, answered with citations
    • Diagnosing a wrong answer: the wrong document cited, paraphrase drift or contradictory sources
    • Narrowing or broadening scope when answers are too generic or too thin
    • Rebuild triggers, and comparing search results before and after a rebuild
    • Recording which agents and networks use each index so they are re-tested
    Lab
    Write a question set for your index, record which answers cite the right source, change one chunking setting, rebuild and compare the new results with the first run.
    Outcome
    A retrieval test you can repeat after every rebuild or change to a source.

Capstone: governed retrieval over your own data

Build retrieval over a document source and a database from your organisation, with discovery findings, a scoped index, enforced visibility and a repeatable question set, then take a VDF AI engineer through your design choices and test results.

VDF AI Certified RAG Engineer badge

VDF AI Certified RAG Engineer

Awarded to data and AI engineers who complete Enterprise RAG Engineering and pass the capstone review.

  • Connecting and assessing document and database sources
  • Building and tuning vector indexes for semantic search
  • Keeping retrieval inside source permissions and visibility
  • Testing retrieval quality and preparing fine-tuning datasets

How VDF AI certification works

What your team gets

  • Four live half-day sessions (3.5 hours each)
  • A hands-on lab after every module
  • A materials pack: session slides and lab guides
  • A capstone review with a VDF AI engineer
  • The VDF AI Certified RAG Engineer certificate on passing the capstone

Apply, agree dates, learn

  1. Send the application below. It takes two minutes.
  2. We reply within 2 business days. Then we agree dates.
  3. Your team gets the materials pack. Then the live sessions begin.

Related courses: Production Agentic Systems: Multi-Agent, RAG and Governance; VDF AI API and Integration Engineering.

Apply for Enterprise RAG Engineering

Free for VDF AI customers and partners. Other teams receive a quote after the application is reviewed. We reply within 2 business days.

Course *

Fields marked * are required.