Computer Vision
Explore image understanding, object detection, visual intelligence, image generation, and where vision models earn their place.
This series covers how machines interpret images: core tasks, data preparation, generation, evaluation and running vision at scale.
Read AI Fundamentals first, particularly neural networks. References include the landmark papers for each technique.
What You’ll Learn
Images are grids of pixel values. Convolutional neural networks slide learned filters over the image to build feature maps, from edges to objects. Residual connections made very deep networks trainable, and vision transformers apply attention to image patches.
A vision model learns what to look for by seeing many labelled images.
ResNet won the 2015 ImageNet challenge using residual connections; the Vision Transformer later showed attention alone can match convolutional networks at scale.