✓ Link copied to clipboard!
Multimodal models (VLM): AI that sees and reads | Sri AI
AI: Local AI & Model Types

Multimodal models (VLM): AI that sees and reads | Sri AI

(0 reviews)
Advanced 1 views

Course overview

Models that take an image and a question together: describing photographs, reading charts and answering questions about scanned documents.

Level: Advanced  ·  Mode: Part-time

Who this course is for

Developers building visual assistants, document tools and accessibility aids.

What you will learn

  • Run vision-language models locally
  • Build image question-answering tools
  • Read charts, receipts and forms with a VLM
  • Fine-tune a VLM on your own images

Syllabus

  1. Module 1: How vision encoders join language models
  2. Module 2: Running VLMs locally
  3. Module 3: Visual question answering
  4. Module 4: Documents and charts
  5. Module 5: Fine-tuning a small VLM

Final project

Every module ends in hands-on practice, and the course ends with a project you build and present. Your certificate names that project.

Before you start

Run your own AI and Deep learning.

Open-source tools you will use

⭐ Rate This Course