AI: Local AI & Model Types
Multimodal models (VLM): AI that sees and reads | Sri AI
Advanced
2 views
Course overview
Models that take an image and a question together: describing photographs, reading charts and answering questions about scanned documents.
Level: Advanced · Mode: Part-time
Who this course is for
Developers building visual assistants, document tools and accessibility aids.
What you will learn
- Run vision-language models locally
- Build image question-answering tools
- Read charts, receipts and forms with a VLM
- Fine-tune a VLM on your own images
Syllabus
- Module 1: How vision encoders join language models
- Module 2: Running VLMs locally
- Module 3: Visual question answering
- Module 4: Documents and charts
- Module 5: Fine-tuning a small VLM
Final project
Every module ends in hands-on practice, and the course ends with a project you build and present. Your certificate names that project.
Before you start
Run your own AI and Deep learning.