From Prototype to Production
A model that scores 94% in a notebook is not a product; a camera feed, a latency budget, and a phone that is three years old are. Age a vision model against a slowly drifting camera and a shrinking latency budget and see where a demo actually breaks.
Read first: Mean Average Precision, Overfitting and Augmentation
A notebook cell that prints 94% accuracy feels finished. It is not a product — it is a measurement of one static test set, taken once, on hardware that never changes and a camera that never ages. A product is a camera bolted to a wall for three years, a phone that is not this year’s, and a latency budget that does not care how good your model tested.
The notebook number was never the product
Every metric this track has built up — accuracy, the confusion matrix, mAP — describes performance on a fixed test set, measured once, under conditions that held still for the duration of the measurement. That number is real and it is useful for comparing models. It is not a description of what happens when the same model faces a continuous stream of camera frames arriving from the actual world, which does not hold still.
Production input is not a bigger version of the test set. It is a different kind of thing entirely: unbounded, arriving in real time, drawn from conditions the test set could only sample a slice of. A 94% test accuracy is a fact about the test set, not a promise about tomorrow’s camera feed.
The camera you ship to is not the camera you trained on
A deployed camera drifts, in ways a training set fixed at one point in time cannot anticipate. A security camera’s lens picks up a scratch or a light film of dust over months of outdoor exposure. Its white balance shifts gradually with the seasons, as the colour temperature of daylight itself changes. A phone released two years after the training photos were captured carries a different camera sensor, a different default sharpening pipeline, sometimes a visibly different colour response out of the box.
None of this shows up as a dramatic failure. It shows up as a slow, unannounced decline — the model still runs, still returns confident-looking numbers, and is quietly working against slightly different input than the one it was tuned against.
The latency budget you do not get to ignore
A live video feed at 30 frames per second hands the model a new frame every 33 milliseconds, whether or not it has finished with the last one. Take a large, highly accurate model that needs 200 milliseconds per frame, running on a single worker. That worker cannot keep up with the frame rate, regardless of how the model scored offline: it has to skip five frames in every six, and whatever happens in those gaps goes unseen.
Two numbers are tangled together here. Throughput is how many frames per second the system can get through; latency is how long any one frame waits for its answer. Running six or seven copies of the model side by side, each taking every sixth or seventh frame, can bring throughput back up to 30 frames per second, at the cost of that much more hardware. It does nothing for latency. Every answer still arrives at least 200 milliseconds after its frame, which for a robot arm or a car can already be too late.
A three-second simulation of a 30 fps feed and a 200 ms model, with one worker and then seven.
frame_gap = 1000 / 30 # ms between frames at 30 fpsinference = 200 # ms per frame, on one workerframes = 90 # three seconds of video for workers in [1, 7]: # try 3 or 5 free_at = [0.0] * workers # when each worker can next start processed = 0 for i in range(frames): arrives = i * frame_gap w = min(range(workers), key=lambda j: free_at[j]) if free_at[w] > arrives: # every worker is busy, so this frame is skipped continue free_at[w] = arrives + inference processed += 1 print(f"{workers} worker(s): {processed} of {frames} frames processed, " f"each answer {inference} ms after its frame")One worker processes 15 of the 90 frames. Seven process all 90. Either way, every answer arrives 200 ms after its frame.
This is a genuine trade-off against accuracy, not an implementation footnote to solve later. A smaller, faster, slightly less accurate model that fits inside 33 milliseconds per frame on the hardware you actually have is often the only real option; a bigger model that cannot keep up does not get partial credit for its offline score.
Monitoring a model nobody is relabelling
A benchmark had labels: every image had a known right answer, so accuracy could be computed directly. Production traffic has no such thing — nobody is sitting behind the live camera feed manually labelling every frame the moment it arrives, and a clean accuracy number requires exactly that.
Monitoring in practice means watching proxies instead of the real metric. A model’s confidence-score distribution drifting lower over weeks, even without any single dramatic failure, is a signal something in the input has shifted. A sudden spike in low-confidence predictions on a specific camera is a signal worth investigating before it becomes a user complaint. User-reported corrections, sparse as they are, are often the closest thing to ground truth a production system gets in real time.
What to actually watch, absent ground truth
- →The shape of the confidence-score distribution over time, not just its average.
- →Sudden spikes in low-confidence predictions, broken out by camera or device if possible.
- →User-reported corrections, even at low volume — they are close to the only ground truth available.
- →Any known hardware or firmware change on the deployed cameras, correlated against the above.
Key takeaways
- A benchmark score is a fact about a static test set measured once; it is not a promise about a continuous, real-time production camera feed.
- Deployed cameras drift from what the model trained on through scratched lenses, seasonal white-balance shifts, and newer sensor generations, and the decline is usually gradual, not dramatic.
- At 30 frames per second a single worker has roughly 33 milliseconds per frame; a 200-millisecond model on one worker can only keep up by skipping most frames. Parallel workers can restore throughput, but each answer still arrives 200 milliseconds late.
- Choosing a smaller, faster, slightly less accurate model to fit the latency budget is a genuine engineering trade-off, not a footnote to accept reluctantly.
- Production traffic has no automatic ground truth, so monitoring means watching proxy signals — confidence-score drift, spikes in low-confidence predictions, user corrections — rather than a clean accuracy number.
Before you move on
Quick check
Answer these to unlock the next chapter: 3 of 4 to pass. You can retake it anytime.
Make a free account to read on
Every chapter is free — an account lets your course progress follow you from your laptop to your phone. No payment, no trial.