IBM

Vector Databases and Retrieval Data Engineering

IBM

Vector Databases and Retrieval Data Engineering

Antonio Cangiano
Ruslan Podgaets

Instructors: Antonio Cangiano

Included with Coursera PlusLearn more

Gain insight into a topic and learn the fundamentals.
Intermediate level

Recommended experience

2 weeks to complete
at 10 hours a week
Flexible schedule
Learn at your own pace
Gain insight into a topic and learn the fundamentals.
Intermediate level

Recommended experience

2 weeks to complete
at 10 hours a week
Flexible schedule
Learn at your own pace

What you'll learn

  • 1.Explain how embeddings and vector retrieval differ from keyword search.

  • 2.Design metadata-rich vector schemas and embedding pipelines.

  • 3.Evaluate retrieval systems using recall, latency, and drift signals.

  • 4. Apply security and governance controls to vector retrieval systems.

Details to know

Shareable certificate

Add to your LinkedIn profile

Recently updated!

July 2026

Assessments

32 assignments

Taught in English

See how employees at top companies are mastering in-demand skills

 logos of Petrobras, TATA, Danone, Capgemini, P&G and L'Oreal

Build your Machine Learning expertise

This course is part of the IBM AI-Native Data Engineering Professional Certificate
When you enroll in this course, you'll also be enrolled in this Professional Certificate.
  • Learn new concepts from industry experts
  • Gain a foundational understanding of a subject or tool
  • Develop job-relevant skills with hands-on projects
  • Earn a shareable career certificate from IBM

There are 9 modules in this course

This welcome module introduces Course 3, explains its role in the AI-Native Data Engineering Professional Certificate, and orients learners to the course journey ahead. Learners review the course purpose, major learning goals, recommended prerequisites, and where to find the full syllabus.

What's included

1 video2 plugins

This module introduces how embeddings power semantic retrieval and how vector retrieval differs from keyword search in practical data engineering workflows. Learners compare similarity metrics and retrieval patterns, then design an initial metadata-rich vector schema and retrieval use case that will serve as the foundation for later vector database design work.

What's included

4 videos5 assignments2 app items4 plugins

This module teaches learners how to design vector database storage patterns that connect embeddings, metadata, filters, source links, and source-of-truth systems. Learners build practical, project-ready artifacts for schema design, filtering strategy, integration boundaries, and storage tradeoff documentation in retrieval and RAG architectures.

What's included

4 videos5 assignments2 app items4 plugins

This module teaches you how to choose and justify vector index strategies by comparing exact and approximate retrieval, practical ANN structures, compression options, and scaling patterns. You will benchmark tradeoffs across recall, latency, memory, build time, and cost, then turn that evidence into a project ready index strategy and scaling memo.

What's included

4 videos5 assignments2 app items4 plugins

Learn how to keep vector indexes accurate and safe as source data changes by designing refresh, update, delete, rollback, and validation workflows. The module emphasizes operational reliability, governance, and auditability so retrieval systems stay consistent with source-of-truth data over time.

What's included

4 videos5 assignments2 app items4 plugins

Learn how to evaluate retrieval systems with golden query sets, relevance labels, and core retrieval metrics, then extend that evidence into production observability with latency, throughput, drift, dashboards, and alerts. The module emphasizes turning metric patterns into actionable engineering decisions, documented limitations, and improvement plans.

What's included

4 videos5 assignments2 app items4 plugins

This module focuses on securing and governing vector retrieval systems while preparing the final course project package. Learners design access controls, tenant isolation, permissions-aware retrieval, auditability, and governance evidence, then assemble and present a complete engineering handoff for a secure vector retrieval subsystem.

What's included

4 videos5 assignments4 plugins

This closing module helps learners reflect on the broad capabilities they developed in Course 3 and recognize their growth in thinking about vector retrieval as an engineered, governable, and operable subsystem. It also provides a high-level transition to Course 4 by previewing the shift from retrieval systems to Lakehouse architecture for AI-native data platforms.

What's included

1 video1 plugin

The Final Exam assesses your ability to apply vector database and retrieval engineering concepts across realistic production scenarios. It combines a quiz on architecture, operations, evaluation, and security tradeoffs with a case study focused on end to end retrieval system design decisions.

What's included

2 assignments1 plugin

Earn a career certificate

Add this credential to your LinkedIn profile, resume, or CV. Share it on social media and in your performance review.

Instructors

Antonio Cangiano
IBM
10 Courses750,442 learners
Ruslan Podgaets
IBM
0 Courses0 learners

Offered by

IBM

Why people choose Coursera for their career

Felipe M.

Learner since 2018
"To be able to take courses at my own pace and rhythm has been an amazing experience. I can learn whenever it fits my schedule and mood."

Jennifer J.

Learner since 2020
"I directly applied the concepts and skills I learned from my courses to an exciting new project at work."

Larry W.

Learner since 2021
"When I need courses on topics that my university doesn't offer, Coursera is one of the best places to go."

Chaitanya A.

"Learning isn't just about being better at your job: it's so much more than that. Coursera allows me to learn without limits."

Frequently asked questions