On-Device Generative AI: Small Language Models, Multimodal AI & Edge Inference
Build efficient generative AI workflows with small language models, multimodal AI, and on-device inference.
About This Course
This 3-day mentor-based workshop introduces practical approaches for running generative AI models beyond the cloud. Participants will explore Small Language Models (SLMs), multimodal AI, Vision-Language Models (VLMs), model quantization, optimization, and efficient on-device inference. Through one hands-on workflow each day, participants will learn how to build, optimize, and evaluate compact generative AI applications for resource-constrained environments.
Aim
To provide participants with practical knowledge of efficient generative AI, covering small language models, multimodal AI, model optimization, and on-device inference workflows.
Workshop Objectives
- Understand the fundamentals of on-device generative AI and Small Language Models.
- Analyze model size, memory, latency, and computational requirements.
- Apply basic quantization and model optimization techniques.
- Understand multimodal AI and Vision-Language Model workflows.
- Build simple image–text generative AI applications.
- Explore efficient inference using CPU, GPU, and NPU-based environments.
- Evaluate model performance using latency, memory, and output-quality metrics.
- Understand privacy, offline inference, and deployment considerations for local AI.
Workshop Structure
Day 1: Small Language Models & Efficient Generative AI
Focus: Understanding compact generative AI models and the techniques used to make them efficient for local inference.
Topics Covered
- Introduction to on-device generative AI and local inference.
- Understanding Small Language Models (SLMs) and their applications.
- SLM architecture, parameters, context length, and inference workflow.
- Model size, memory requirements, latency, and computational efficiency.
- Introduction to model compression and efficient inference.
- Quantization and reduced-precision model formats.
- Cloud-based vs. on-device generative AI.
- Applications of SLMs in private and low-latency AI systems.
🛠️ Hands-on
Run and benchmark a compact language model, comparing model size, inference latency, and response performance before and after quantization.
Tools:
Hugging Face | Google Colab | Python | Transformers | llama.cpp
Day 2: Multimodal AI & Vision-Language Models
Focus: Exploring how generative AI can combine text and visual information for intelligent reasoning.
Topics Covered
- Introduction to multimodal generative AI.
- Fundamentals of Vision-Language Models (VLMs).
- Understanding image–text inputs and multimodal prompting.
- Image understanding and visual question answering.
- Multimodal reasoning and response generation.
- Compact multimodal models for resource-constrained environments.
- Challenges of multimodal inference on edge devices.
- Applications in healthcare, robotics, manufacturing, and smart devices.
🛠️ Hands-on
Build a multimodal AI workflow that accepts an image and text prompt, generates an interpretation, and evaluates the model output.
Tools:
Hugging Face | Google Colab | Python | Transformers | Open-Source VLMs
Day 3: Generative AI Model Optimization & On-Device Inference
Focus: Optimizing generative AI models for efficient, low-latency, and privacy-preserving local inference.
Topics Covered
- End-to-end workflow for on-device generative AI deployment.
- Advanced quantization and efficient model formats.
- Model compression and inference optimization.
- CPU, GPU, and NPU-accelerated inference concepts.
- Memory, latency, and computational efficiency benchmarking.
- Local AI vs. cloud-based inference.
- Privacy, offline inference, and data security considerations.
- Practical deployment challenges and optimization strategies.
🛠️ Hands-on
Optimize and run a quantized generative AI model for local inference, then benchmark latency, memory requirements, and output performance.
Tools:
llama.cpp | GGUF | Hugging Face | Python | LiteRT / Google AI Edge
Who Should Enrol?
- Students and researchers in Computer Science, AI/ML, Data Science, Electronics, and Engineering.
- Developers and professionals working with Generative AI, Edge AI, IoT, or intelligent devices.
- Researchers interested in Small Language Models, multimodal AI, and efficient AI deployment.
- Beginners to intermediate learners with basic Python and machine-learning/AI knowledge.
Important Dates
Registration Ends
October 7, 2026
IST 4:30PM
Workshop Dates
October 7, 2026 – October 9, 2026
IST 5:00 PM
Workshop Outcomes
- Running and evaluating compact language models.
- Working with multimodal and Vision-Language AI models.
- Applying quantization and efficient model formats.
- Optimizing generative AI models for local inference.
- Benchmarking latency, memory usage, and model performance.
- Understanding the workflow from a cloud-based model to an efficient on-device AI application.
Fee Structure
Student
₹2499 | $65
Ph.D. Scholar / Researcher
₹3499 | $75
Academician / Faculty
₹4499 | $85
Industry Professional
₹5499 | $105
What You’ll Gain
- Live & recorded sessions
- e-Certificate upon completion
- Post-workshop query support
- Hands-on learning experience
View All Feedbacks →
