AI Tutorials
Computer Vision
Master computer vision from first principles — covering image preprocessing, CNNs, transfer learning, object detection (YOLO, Faster R-CNN, DETR), segmentation (U-Net, Mask R-CNN, SAM), Vision Transformers, GANs, diffusion models, vision-language models (CLIP, LLaVA), and video understanding.
10 chapters · 590 min
From pixel fundamentals to diffusion models and video understanding
- Ch. 01Read →
Image Fundamentals & Preprocessing
Pixels, color spaces, transforms, and augmentation pipelines
beginner · 45 min
- Ch. 02Read →
CNNs in Depth
Convolutions, pooling, BatchNorm, ResNet skip connections, and architecture evolution
intermediate · 55 min
- Ch. 03Read →
Transfer Learning & Pre-trained Models
Fine-tune ImageNet giants for your own vision tasks
intermediate · 55 min
- Ch. 04Read →
Object Detection
Localise and classify multiple objects with YOLO, Faster R-CNN, and DETR
advanced · 65 min
- Ch. 05Read →
Image Segmentation
Semantic, instance, and panoptic segmentation with U-Net, Mask R-CNN, and SAM
advanced · 60 min
- Ch. 06Read →
Vision Transformers
ViT, Swin Transformer, DINO, and the end of convolutional dominance
advanced · 60 min
- Ch. 07Read →
GANs & Variational Autoencoders
Generative models for image synthesis, style transfer, and latent space manipulation
advanced · 65 min
- Ch. 08Read →
Diffusion Models
DDPM, DDIM, and Stable Diffusion for high-quality image generation
advanced · 65 min
- Ch. 09Read →
Vision-Language Models
CLIP, BLIP-2, LLaVA, and multimodal understanding
advanced · 60 min
- Ch. 10Read →
Video Understanding
Temporal models, optical flow, and action recognition in video
advanced · 60 min