Computer vision, made visual
CodeWithPurpose lessons that explain how a machine turns a photo into a label, a box, or a mask — through live demos and clear visuals. Click any topic to explore, no setup required.
Ready when you are
Go at your own pace. Your progress saves on this device.
Syllabus
What you'll learn
Take it one chapter at a time. A short quick check at the end of each one opens the next.
Part 1
Seeing as Numbers
What a computer actually receives when you hand it a photo, and the operation — sliding a small grid across every pixel — that almost everything else in this track builds on.
- What Is Computer Vision?A photograph looks obvious to you and is, to a computer, a wall of numbers with no idea what a face is. See the three different jobs hiding inside the phrase "computer vision", and why none of them were ever solved by writing rules.
- Images as NumbersZoom into any photo far enough and the picture disappears into a grid of numbers with nothing else hiding inside it. Click through a tiny image pixel by pixel and watch a face turn back into the grid it always was.
- Colour and ChannelsColour is not one number per pixel, it is three, stacked on top of each other and constantly disagreeing. Split a photo into its red, green, and blue channels and find out which one was quietly doing almost all of the work.
- Convolution and FiltersOne small grid of numbers, slid across every position in an image, is the operation underneath blur, sharpen, and every edge a neural network has ever found. Drag a 3×3 kernel across a photo and watch what each stop produces.
Part 2
Getting a Clean Signal
Edges, corners, and the preprocessing that turns an inconsistent pile of photos into something a model can actually learn from.
- Edge DetectionAn edge is not a line the camera drew, it is a place where brightness changes fast, and "fast" is a number you get to choose. Drag the threshold on a real photo and watch a coherent outline dissolve into noise, or vanish into nothing.
- Corners and KeypointsA corner survives a rotation, a crop, and a change of lighting in ways a raw pixel value never will, which is why decades of computer vision were built on finding them instead. See what makes one point in an image worth finding again in another.
- Resizing, Normalising, AugmentingA model trained on perfectly centred, evenly lit photos will fail the first time a real photo arrives sideways. Rotate, crop, and relight one training image and see how many new examples a single photo can honestly become.
Part 3
Teaching a Model to Recognise
From a weight on every pixel to a network that learns edges, then parts, then whole objects — and why you rarely train one from nothing.
- Image ClassificationNaming what is in a photo sounds like the whole problem, and is actually the easy version of it: one label, for the entire image, and nothing about where. See why that constraint is a feature, not a shortcut.
- From Pixels to a PredictionThe simplest possible image classifier is a weight for every single pixel, added up into one score, and it is worse than it sounds. Drag the weights yourself and watch how little a model that ignores position can actually learn.
- Convolutional Neural NetworksStack the sliding kernel from chapter four dozens of layers deep and something new happens: early layers learn edges, and later ones learn eyes, wheels, and faces, without anybody telling them to. Step through the layers of a small network and see what each one has learned to notice.
- Transfer LearningTraining a network like chapter ten's from nothing takes millions of photos you almost certainly do not have. Take one somebody else already trained on a different problem, keep everything but its last layer, and see how little new data it takes to repurpose it.
Part 4
Finding Things in a Scene
Naming what is in a photo is the easy version. This part finds where, how many, and where one object's pixels end and another's begin.
- Object DetectionClassification tells you a photo contains a dog. Detection tells you where, how many, and draws a box around each one — a harder problem, and the rest of this part exists to solve it.
- Bounding Boxes and IoUA predicted box that is "close" to the right one needs a number, not a shrug, and that number is Intersection over Union. Drag two boxes apart and watch a plausible-looking overlap score fall below 0.5.
- Non-Max SuppressionA detector rarely proposes one box per object, it proposes a dozen, all clustered around the same dog. Raise and lower a suppression threshold and watch eleven overlapping boxes collapse into one, or refuse to.
- Image SegmentationA box around a dog still includes fifteen pixels of fence behind it. Segmentation labels every pixel individually, and the moment two dogs overlap, decides whether it still knows they are different dogs.
Part 5
Making the Numbers Honest
Every way a vision model's score flatters it, from a confusion matrix that hides which classes fail to a benchmark that keeps rising while the dataset it was trained on quietly biases it.
- Confusion Matrix, One Class at a TimeOne accuracy number across a hundred classes hides which ones the model actually confuses with each other. Open the full matrix and find the one pair of classes responsible for most of the errors.
- Mean Average PrecisionDetection has no single "accuracy", because a prediction can be right about the class and wrong about the box, by varying amounts. Slide the confidence threshold across a set of detections and watch precision and recall trade against each other the way they did for a plain classifier.
- Dataset BiasA model trained mostly on photos taken in daylight, indoors, of one demographic, has not learned to see — it has learned that dataset. See how a benchmark can keep rising while the system quietly gets worse for everyone the dataset under-represented.
- Overfitting and AugmentationA network with millions of parameters can memorise ten thousand training photos outright, pixel quirks and all. Turn augmentation up and down and watch the gap between training and validation accuracy close, or refuse to.
Part 6
Vision in the World
What changes when a model has to run on a phone, in a car, or on a face, instead of in a notebook with all the time it wants.
- Face Detection and PrivacyThe same pipeline that unlocks your phone can identify a stranger in a crowd from a single frame, and the two uses are not equally consented to. See where face detection ends and face recognition begins, and why that line is the whole ethical argument.
- Vision in Self-Driving CarsA self-driving car cannot ask a model to be 99% sure before braking; it has to act inside a fraction of a second on whatever the camera and the other sensors currently agree on. See why no camera is ever trusted alone, and what happens when the sensors disagree.
- From Prototype to ProductionA model that scores 94% in a notebook is not a product; a camera feed, a latency budget, and a phone that is three years old are. Age a vision model against a slowly drifting camera and a shrinking latency budget and see where a demo actually breaks.
Open source
Your lesson goes here
Every lesson here was written by a student. Propose your own topic, or fix one that confused you. Drafts stay invisible until they are ready, so nothing you start can break the site.
Keep building your computer vision foundation
These lessons are part of CodeWithPurpose's free learning library — built by students, for students, everywhere.