Enterprise RAG Engineering
Enterprise RAG Engineering is an instructor-led VDF AI course for data and AI engineers who build retrieval over documents and databases. In four live half-days you connect and profile sources, build and test vector indexes, keep retrieval inside source permissions and prepare fine-tuning datasets, then earn the VDF AI Certified RAG Engineer certificate through a capstone.
- 4 live half-days
- 6 modules + capstone
- Remote or on-site
- VDF AI Certified RAG Engineer
- Level
- Advanced
- Format
- Live and instructor-led, remote or on-site
- Length
- Four live half-day sessions (3.5 hours each)
- Audience
- Data engineers, AI engineers and data leads who build retrieval over enterprise documents and databases
- Cost
- Free for customers and partners; quoted for other teams
- Certificate
- VDF AI Certified RAG Engineer
- Reply to applications
- Within 2 business days
- Labs
- One after every module
What you will be able to do
- Connect apps and databases with the narrowest useful scope and a dedicated read-only account
- Judge whether a source deserves trust from discovery, quality scores and drift signals
- Build vector indexes with row text and chunk settings that suit the content
- Decide which table columns become searchable text and which data stays with live queries
- Keep retrieval inside the permissions and visibility of every source
- Test retrieval against real questions before and after each rebuild
Prerequisites
- Completion of the Private RAG Engineering path, or equivalent experience indexing enterprise content
- A test database and document set you may connect in a deployment with VDF AI Data enabled
- Recommended free path: Private RAG Engineering
6 modules and a capstone
The modules run across the four sessions. Each one ends with a lab in VDF AI.
-
Connecting sources and databases
Bringing files, connected apps and databases in as governed sources, scoped for precise retrieval from the first day.
- Choosing between file upload, connected app, database connection and pasted text
- Scoping by project, team space or schema rather than a whole drive or server
- Supported databases, the generic JDBC option and Jira as a structured source
- Dedicated read-only accounts, secret references and the four connection states
- Testing a connection, refreshing asset inventory and rotating credentials on a calendar
- Lab
- Connect one document source and one database schema through a dedicated read-only account, run Test connection, refresh the asset inventory and record the owner in each connection’s description.
- Outcome
- Two tightly scoped, tested sources, each with a named owner and a refresh routine.
-
Discovery and data health
Finding out what a source holds and whether it deserves trust before anything is indexed.
- Running discovery: asset cards with size, freshness, owner, quality score and tags
- Reading the 0–100 quality score and deciding when to investigate first
- Exploratory data analysis: missing values, duplicate rows, outlier columns and class imbalance
- Column profiles and low, medium and high drift signals
- Tags, comments and asset history as context the whole team can find
- Lab
- Run discovery on your connected schema, tag the assets that matter, run exploratory analysis on the most important table and leave what you learn as comments on the asset.
- Outcome
- A short, evidence-based list of assets that are fit to index, with their risks written down.
-
Vector indexes, semantic search and rows to text
Building a search surface you own: what goes into it, how rows become text and how text is split.
- Index sources: text-heavy table columns, feature lists, connected apps and file collections
- Rows to text: which columns become searchable chunks and which data is better queried live
- Chunk size and overlap for short snippets, long documents and mixed material
- Row text and row limits at build time: textColumns, textTemplate and maxRows in the Data API
- Build stages and states, from reading and chunking through embedding to Ready
- Lab
- Build one index over long-form documents and one over the text-heavy columns of a table, each with chunk settings suited to its content, then search both with the same set of questions.
- Outcome
- Index settings you can justify from search results rather than from defaults alone.
-
Feature lists, relationships and fine-tuning datasets
Curating the columns that matter for one use case, then reusing that curation for search and for training data.
- Feature lists: one default per asset plus specialised lists for each use case
- Feature discovery: derived features with confidence scores to accept or dismiss
- Association analysis, and why a strong correlation is a question rather than a conclusion
- Fine-tuning datasets: mapping templates, preview and the Draft, Ready and Exported states
- Index or fine-tune: current facts to look up versus a change in behaviour or voice
- Lab
- Build a feature list for one table, triage the suggested derived features, scope an index to the list and preview a small fine-tuning dataset from the same source.
- Outcome
- One curated feature list behind both a focused index and a previewed training dataset.
-
Permission-aware retrieval and data governance
Keeping every answer inside what the person asking may see, and proving that it stays there.
- Why a vector index flattens the permissions of its source
- Letting the source account’s grants decide what an index can read
- Keeping sensitive rows out of an index built for a wider audience
- Who can search an index, and why each index needs a named owner account
- Rebuilding after every permission change, and what a failed rebuild leaves behind
- Lab
- Index a source through a read-only account with narrow grants, keep sensitive rows out with a retrieval table, revoke a grant and watch what the index returns before and after a rebuild.
- Outcome
- Retrieval that stays inside source permissions, with tests that prove it.
-
Testing retrieval and keeping it current
Measuring retrieval against real questions, then keeping indexes and connections fresh once people depend on them.
- A question set drawn from real questions, answered with citations
- Diagnosing a wrong answer: the wrong document cited, paraphrase drift or contradictory sources
- Narrowing or broadening scope when answers are too generic or too thin
- Rebuild triggers, and comparing search results before and after a rebuild
- Recording which agents and networks use each index so they are re-tested
- Lab
- Write a question set for your index, record which answers cite the right source, change one chunking setting, rebuild and compare the new results with the first run.
- Outcome
- A retrieval test you can repeat after every rebuild or change to a source.
Capstone: governed retrieval over your own data
Build retrieval over a document source and a database from your organisation, with discovery findings, a scoped index, enforced visibility and a repeatable question set, then take a VDF AI engineer through your design choices and test results.
VDF AI Certified RAG Engineer
Awarded to data and AI engineers who complete Enterprise RAG Engineering and pass the capstone review.
- Connecting and assessing document and database sources
- Building and tuning vector indexes for semantic search
- Keeping retrieval inside source permissions and visibility
- Testing retrieval quality and preparing fine-tuning datasets
What your team gets
- Four live half-day sessions (3.5 hours each)
- A hands-on lab after every module
- A materials pack: session slides and lab guides
- A capstone review with a VDF AI engineer
- The VDF AI Certified RAG Engineer certificate on passing the capstone
Apply, agree dates, learn
- Send the application below. It takes two minutes.
- We reply within 2 business days. Then we agree dates.
- Your team gets the materials pack. Then the live sessions begin.
Related courses: Production Agentic Systems: Multi-Agent, RAG and Governance; VDF AI API and Integration Engineering.